Benchmark Analysis Report

2026 CSAT English AI Performance

Evaluation of 24 Large Language Models across 45 questions in the Korean CSAT English Section without external web search.

DATE: 2026-08-01 ITEMS: 45 MODELS: 24
Perfect Score 100% Accuracy
6 / 24
Claude Fable 5, GPT-5.6 Terra/Sol, Gemini 3.6
Mean Accuracy Overall
95.4%
Average 42.9 / 45 items answered correctly
Hardest Question 50.0% Accuracy
Item #39
Sentence Insertion (12 models failed)
Thinking Mode Uplift Reasoning
+7.0%p
Claude Sonnet 5 (High 96% vs None 89%)

Model Accuracy Ranking

24 Models Benchmark

Vulnerability Ratio

Item Breakdown

Model Performance Directory

Complete list of evaluated models ordered by total correct answers and accuracy rate.

RANK AI MODEL NAME SCORE ACCURACY INCORRECT QUESTION NUMBERS

Reasoning Mode Impact Matrix

Precision Stepper & Horizontal Range Bar (0------0) comparing accuracy progression across Reasoning Tiers.

Interactive Range Stepper
SELECT MODEL FAMILY TO INSPECT
REASONING IMPACT OVERVIEW RANGE BARS (0------0 MATRIX)

Vulnerable Question Deep-Dive

Analysis of 10 items where incorrect answers occurred. (The remaining 35 items had 100% accuracy).

10 Error Items

Key Takeaways & Analytical Insights

Benchmark Summary
100% Perfect Score

6 Flagship AI Models Attain 45/45

Claude Fable 5 (high), GPT-5.6 Terra/Sol (high/Pro), and Gemini 3.6 Flash (high) successfully solved all 45 questions without errors, showcasing flawless reading comprehension in CSAT English.

Core Differentiator

Sentence Insertion & Paragraph Sequence

Item #39 (50% accuracy) and Item #37 (63% accuracy) served as the critical distinction barrier. Prominent models like Grok 4.5, Gemini 3.1 Pro, and Claude Opus 5 struggled with these items.

Reasoning Impact

Thinking Mode Substantially Boosts Accuracy

Enabling High / Thinking reasoning options yielded significant accuracy gains across identical model families: Claude Sonnet 5 (+7%p), GPT-5.6 Luna (+7%p), and Claude Opus 5 (+2%p).

Model Recommendations & Cost Efficiency

Optimal Model Selection Guide based on Accuracy Rate, Latency, and Compute Cost Efficiency.

Editor's Choice
Top Best Value Pick

Gemini 3.6 Flash (high)

Score: 45 / 45 (100% Perfect)
Delivers flagship-level 100% accuracy with ultra-fast Flash latency and low token costs. The absolute best cost-to-performance choice for automated English comprehension tasks.

COST: LOW SPEED: FAST ACCURACY: 100%
Top Open Weight Pick

Qwen3.7 Max & DeepSeek Flash

Score: 44 / 45 (98% Accuracy)
Qwen3.7 Max (Thinking) and DeepSeek V4 Flash missed only 1 item (#4 listening/vocab error). Exceptional performance for accessible open-weights and efficiency tiers.

COST: VERY LOW MISSED: #4 ONLY
Top Mission-Critical Pick

Claude Fable 5 & GPT-5.6 Terra

Score: 45 / 45 (100% Perfect)
Uncompromising reasoning reliability for critical document parsing where zero error is tolerated, backed by full multi-turn high reasoning capabilities.

REASONING: HIGH RELIABILITY: MAX