The three-year journey from version 1.0 to 3.0 of LG AI Research’s EXAONE has not been an easy one, as we have continued our research and released upgraded models and external commercialization results every year. From proving the development potential of AI technology in various industries one by one, we’ve created models that users can better utilize between the two pillars of performance and cost and developed expert-level AI that can be applied to real-world industrial fields.

EXAONE Milestone
Release of the EXAONE 3.0 7.8B Instruction Tuned Language Model
August 2024. Finally, we are excited to announce EXAONE 3.0. Among various EXAONE 3.0 language model lineups, we are releasing the 7.8B Instruction Tuned model as an open source for research. We hope that this model will help AI researchers in Korea and abroad to conduct more meaningful research and help the AI ecosystem move forward.
The 7.8B model released this time is based on the Decoder-only Transformer Architecture in line with recent trends, with 7.8B parameters and 8T training data (tokens). This post will introduce the main features, performance evaluation results, and insights of EXAONE 3.0 7.8B Instruction Tuned language model. For our performance evaluation, we utilized a combination of publicly available datasets and our own benchmark datasets to compare the performance of the 7.8B model with the latest AI models that support English and Korean, which are similar in size to the 7.8B model.
Key Takeaways
■ Achieved global top level in English : Ranked 1st average in real-world use cases and excellent performance in benchmarks
The English performance of the 7.8B model is at the Global Top-level compared to other models. EXAONE is aiming to be a high-level Expert AI that can be utilized in specialized industries. In order for AI models to be utilized in specialized industries and fields of expertise, they must perform well in real-world use cases, i.e., in a complex manner so humans can trust and use them. To evaluate this aspect, the Chatbot Arena method has recently been widely used, which is a method of directly using and evaluating models based on features that humans often use. While this evaluation is time-consuming, an accurate assessment of the real-world utility of the model is an advantage it provides. To confirm the English performance of the 7.8B model, we selected four key benchmarks that are similar to how Chatbot Arena is evaluated and evaluated the model on items with high human utilization. The results showed that EXAONE 7.8B model ranked first in most benchmarks, with the highest average score.
It also demonstrated superior performance on benchmarks. It ranked first in average scores for math and coding, demonstrating superiority over other models. And it also achieved strong performance results in reasoning.

Evaluation Results of Real-world Use Cases (English)

Benchmark – Math, Coding, Reasoning (English)
■ Clearly outstanding Korean language performance : Ranked first in average scores for both real-world use cases and benchmarks
EXAONE 7.8B model is a bilingual model that targets both English and Korean languages. For the Korean performance evaluation, we used two benchmarks to check the performance for real-world use cases, and configured multiple benchmarks to check general performance. As a result, we were able to see top overall results in both real-world use cases and general performance.

Evaluation Results of Real-world Use Cases (Korean)

Benchmark (Korean)
■ Securing economic feasibility : Reduced to 6% of the cost of the initially released model through three years of research and development
In order for AI to be applied to our lives, it is essential to improve performance as well as enhance economic feasibility. Since the release of EXAONE 1.0 in 2021, we have spent the past three years focusing on research and development in AI model compression technologies to achieve cost efficiency. As a result, the 7.8B model released shows a 56% reduction in inference processing time and a 72% reduction in cost compared to EXAONE 2.0. In particular, it is a significant reduction in cost, bringing it down to just 6% of the cost of the initially released EXAONE 1.0.

EXAONE 3.0 Performance Improvement
■ Ethical transparency : In addtion to excellent results, disclosure of areas requiring improvement
LG AI Research always considers AI ethics in the research and development process of AI models. EXAONE 3.0 7.8B Instruction Tuned language model also underwent a Red Teaming process to assess its ethics and security and was evaluated using both internal and external third-party datasets.
While the model released this time is excellent at providing non-sexually discriminatory and legal answers, there are areas that need to be improved. We disclosed the evaluation results as they are because we believe that transparent disclosure of information is a prerequisite for the development of AI ethics. We hope that researchers will conduct more active research on AI ethics based on this disclosure, and LG AI Research will also continue to research AI ethics.

Evaluation Results of Harmlessness (Korean Large Language Model Trustworthiness Benchmark Data)
You can view the detailed information, including the model's performance evaluation results, through the link below and directly download and use the 7.8B model. We hope that the release of this model will contribute to assisting various research and development by AI researchers and enhancing technological competitiveness.