Evaluation of 24 Large Language Models across 45 questions in the Korean CSAT English Section without external web search.
Complete list of evaluated models ordered by total correct answers and accuracy rate.
| RANK | AI MODEL NAME | SCORE | ACCURACY | INCORRECT QUESTION NUMBERS |
|---|
Precision Stepper & Horizontal Range Bar (0------0) comparing accuracy progression across Reasoning Tiers.
Analysis of 10 items where incorrect answers occurred. (The remaining 35 items had 100% accuracy).
Claude Fable 5 (high), GPT-5.6 Terra/Sol (high/Pro), and Gemini 3.6 Flash (high) successfully solved all 45 questions without errors, showcasing flawless reading comprehension in CSAT English.
Item #39 (50% accuracy) and Item #37 (63% accuracy) served as the critical distinction barrier. Prominent models like Grok 4.5, Gemini 3.1 Pro, and Claude Opus 5 struggled with these items.
Enabling High / Thinking reasoning options yielded significant accuracy gains across identical model families: Claude Sonnet 5 (+7%p), GPT-5.6 Luna (+7%p), and Claude Opus 5 (+2%p).
Optimal Model Selection Guide based on Accuracy Rate, Latency, and Compute Cost Efficiency.
Score: 45 / 45 (100% Perfect)
Delivers flagship-level 100% accuracy with ultra-fast Flash latency and low token costs. The absolute best cost-to-performance choice for automated English comprehension tasks.
Score: 44 / 45 (98% Accuracy)
Qwen3.7 Max (Thinking) and DeepSeek V4 Flash missed only 1 item (#4 listening/vocab error). Exceptional performance for accessible open-weights and efficiency tiers.
Score: 45 / 45 (100% Perfect)
Uncompromising reasoning reliability for critical document parsing where zero error is tolerated, backed by full multi-turn high reasoning capabilities.