EXAONE Deep Released ━ Setting a New Standard for Reasoning AI

 

The era of Agentic AI, where AI independently formulates hypotheses, verifies them, and autonomously makes decisions without human instructions, is approaching. The development of enhanced reasoning models is essential for this transition, but securing high-performance reasoning models is no easy task. Globally, only a few companies with foundation models have been able to develop their own advanced reasoning models.

LG AI Research is now unveiling EXAONE Deep, a reasoning AI with enhanced reasoning capabilities capable of competing with these industry-leading models. EXAONE Deep excels in understanding mathematical logic, reasoning about scientific concepts, and solving programming problems, making it a high-performance model specialized in advanced reasoning.

To release EXAONE Deep, we focused on dramatically improving Reasoning performance in Math, Science, and Coding while ensuring the model’s ability to understand and apply knowledge across various domains.


  1. Math – EXAONE Deep 32B Outperforms Competitors in High-Difficulty Math Benchmarks Even at Just 5% of Their Size

  2. Science & Coding – EXAONE Deep 7.8B & 2.4B Achieve Dominant Performance, Ranking #1 Across All Key Benchmarks

  3. MMLU – EXAONE Deep 32B Achieves 83.0 score, Proving the Best Performance Among Domestic Models


Immediately after its release, the EXAONE Deep 32B model was listed in Epoch AI’s Notable AI Models, validating its performance. This marks another achievement following EXAONE 3.5, making EXAONE the only Korean model to be listed in the past two years.


Image 1. EXAONE Deep and EXAONE 3.5 listed in Epoch AI’s Notable AI Models
(Source: Epoch AI)


Now, let’s dive into the key features and performance of EXAONE Deep, supported by actual benchmark results. Explore the future and potential of AI with EXAONE Deep!



1. Math ━ Achieving the Highest Score on CSAT 2025, the 7.8B&2.4B models ranked first in all major benchmarks

All EXAONE Deep models (32B, 7.8B, and 2.4B) achieved the highest score on the CSAT 2025 Mathematics section, proving their superior performance compared to the latest global reasoning models.


Image 2. Performance Comparison in the Mathematics Category
 ※ The asterisk (*) indicates figures that have been officially reported. Scores highlighted in purple indicates the best performance. 


Image 3. CSAT 2025 Mathematics Evaluation Results



EXAONE Deep 32B ━ Equivalent Performance at 5% the Size of the Competing Model 

The EXAONE Deep 32B model achieved 94.5 points in the CSAT mathematics section and 90.0 points on the AIME 2024 (American Invitational Mathematics Examination), which serves as a qualifying standard for the U.S. Olympiad selection process, demonstrating the highest performance among competing models.

In AIME 2025, it achieved performance on par with the DeepSeek-R1 (671B) model, proving its exceptional problem-solving capabilities and logical reasoning skills in challenging mathematical evaluations.

Notably, in high-difficulty benchmarks like AIME, EXAONE Deep 32B outperformed significantly larger models, reaffirming its learning efficiency and cost-effectiveness, which are recognized as key strengths of EXAONE.


EXAONE 7.8B & 2.4B  ━ Ranked #1 Across All Key Benchmarks

The 7.8B and 2.4B models each ranked first in all major benchmarks within the lightweight model and on-device model categories, respectively proving their practical utility. The 7.8B model scored 94.8 in MATH-500 and 59.6 in AIME 2025. The 2.4B model achieved 92.3 in MATH-500 and 47.9 in AIME 2024.


2. Science & Coding ━ Exceptional Expertise in Scientific Reasoning and Software Coding 


Image 4. Performance Comparison in the Coding Category
※ The asterisk (*) indicates figures that have been officially reported. Scores highlighted in purple indicates the best performance. 


EXAONE Deep has demonstrated competitive superiority in Science and Coding domains.

The 32B model achieved 66.1 score on the GPQA Diamond test, which evaluates PhD-level problem-solving abilities in physics, chemistry, and biology. It also recorded a score of 59.5 on LiveCodeBench, a benchmark for evaluating coding capabilities, outperforming other reasoning AI models of similar size. This suggests that it has high applicability even in domains requiring specialized knowledge.

The 7.8B and 2.4B models also ranked first in GPQA Diamond and LiveCodeBench. In particular, following the EXAONE 3.5 2.4B model, released last December, securing the top spot in the edge category of Hugging Face's "LLM Leaderboard," EXAONE Deep has achieved the highest performance, proving itself as a globally leading lightweight and on-device model.


3. MMLU ━ Proving the highest performance among domestically developed models

Beyond specialized reasoning in Math, Science, and Coding, EXAONE Deep also delivers exceptional performance in general language understanding. It has secured the highest MMLU score among Korean models, setting a new benchmark for General Reasoning capabilities.

Image 5. Performance Comparison in the General Category 
※ The asterisk (*) indicates figures that have been officially reported. Scores highlighted in purple indicates the best performance. 


EXAONE Deep expands AI’s reasoning capabilities across various domains such as mathematics, science, and coding, and takes on even more complex problem-solving challenges. Through continuous research and innovation, we will continue to advance AI to contribute to enriching and enhancing the convenience of human life.