The Story of Ahra Jo, Leading the DDU Research to Build an AI That Understands All Documents in the World

Image 1. Ahra Jo, at LG AI Research Vision Lab 


AI is now being applied to various industries, driving groundbreaking changes in our everyday lives. As AI becomes capable of understanding and analyzing specialized content within documents, it can suggest tailored applications for different job roles. It also processes diverse data from industrial sites and swiftly optimizes production scheduling. 

LG AI Research has unveiled Deep Document Understanding (DDU), a technology that enables AI to comprehend various types of information, including pictures, graphs, and tables in documents. We aim to apply this technology to conversational AI services and AI prediction services to expand its scope of use and develop unique solutions. 

Today, we’re interviewing Ahra Jo from the Vision Lab, who has been leading the development of DDU technology, to learn more about the research and life at LG AI Research.

  

The Beginning of My First Career


Image 2. The Beginning of My First Career


In my third year of undergrad, I developed an interest in HCI(Human-Computer Interaction), which naturally led me to study Vision on my own. I was particularly fascinated by the idea that our visual perception can be converted into RGB and that computers can use this information to provide valuable insights. This sparked my broader interest in AI.

While considering ways to make a positive impact on society with Vision AI, I came up with the idea of sign language recognition. I wondered—if sign language could be recognized and converted into text or speech, wouldn’t it be possible to facilitate communication with visually impaired individuals who don’t know sign language? Motivated by this thought, I aimed to present my work to the world as my undergraduate graduation project.

I pursued graduate school with a strong desire to conduct research on sign language recognition. However, in graduate school, I kept sign language recognition research close to my heart but spent my time working on other projects and staying busy. Instead, this journey allowed me to explore diverse areas of vision research, where I learned about the vast potential for expanding AI technology.


Image 3. Ahra Jo, deeply immersed in research


Right after receiving my PhD in 2016, I joined a startup, where I worked on a tax evasion prediction project for about a year and a half. I still remember the CEO saying, “As technology advances, it should serve a greater purpose. By identifying black money through tax evasion prediction, we can return it to the country and make the world a better place.” His vision of contributing to social justice through technology resonated with me, which is why I joined the company.

This was probably the only time I focused on AI models using data rather than images. I dedicated myself to the project for a year and a half, and while we made significant progress, I found myself longing to work with Vision again. That’s when I received an exciting offer from LG CNS and joined the team, marking my first connection with LG.


Image 4. Ahra Jo, sitting in front of the LG AI Research logo


At LG, I worked on the Smart Store project, where customers could place the products they wanted to purchase on a self-checkout counter, click the checkout button, and a camera would scan the items on the shelf within a second to process the payment. I handled data collection and model development myself, achieving a recognition accuracy of over 99%.

Later I was invited to join LG AI Research(AI Research of LG Science Park, formerly LG AI Research) in 2019. I was confident that at the institute, I would be able to introduce the AI models I developed to the world. In the early days of my tenure, I was in charge of a vision inspection project. After returning from parental leave in 2021, I became fully involved in the DDU project.


The Infinite Expansion Potential of DDU Technology

 

Image 5. The Infinite Expansion Potential of DDU Technology

 

Currently, the DDU technology is managed by the DDU Squad within the Vision Lab at LG AI Research. Initially, it worked on Basic Structure, Super Atom, and Stereo Chemistry in 2021, gradually expanding its scope to Polymers in 2023 and Vision Markush and Reaction in 2024.

Starting in 2024, we are broadening our research to general document-related areas, developing Vision-based models such as layout detection, Table to HTML, Chart to Table, and OCR. Our major achievements include delivering layout detection, table to HTML, and molecular formula conversion models to LG Chem, as well as layout detection, molecular formula conversion, and Reaction models to global publisher.

Additionally, we packaged all the models into APIs, making them easily accessible for projects that require them. For example, in ChatEXAONE, an AI assistant for enterprise professionals recently released by LG AI Research, the DDU technology has been partially applied to the process of understanding user documents, and we plan to gradually expand its use. In the Financial AI project, it is used to extract and organize numerical data from charts and tables in financial report documentss to generate new financial reports.

In EXAONE Discovery, a platform for discovering new materials, novel substances, and drug discovery, DDU technology is utilized to detect and convert molecular structures. To enable efficient application across various projects, we have developed all our DDU models as APIs.


LG's DDU: What Sets It Apart?


Image 6. An image explaining LG's DDU technology


LG AI Research's Deep Document Understanding (DDU) technology aims to understand all types of documents worldwide. To achieve this, we focused on developing technology that can analyze and understand all the contents of a document and extract the relevant information. Our competitive advantage lies in two key areas. 

The first is accurate information extraction for anti-hallucination, along with the understanding technology that supports it. Recently, the rapid development of LLM generative models has permeated our lives, but there are still hallucination issues based on the data that the model has already learned. To minimize this problem, our DDU Squad aims to extract precise data using a Vision-only model.

The second is independence from document extension. Since we process all documents as images, our model can seamlessly handle multiple formats, including PDF, PPT, DOC, HWP, PNG, and JPG, with the same model. In the case of regular LLM models that don't process as images, they require preprocessing with a parser for each extension at the beginning, but we are able to skip this step entirely.


DDU: Transforming Our Lives

Document AI technology is already being used in various fields of our lives, contributing to improved efficiency across industries such as finance, medicine, law, and logistics. Common applications today include automating the processing of documents like loan applications and insurance claims in finance and insurance, digitizing patient medical records to enhance administrative efficiency in healthcare, automating contract reviews and legal document analysis in legal and contract management to speed up workflows, and optimizing logistics operations by automating invoices and receipts in logistics and distribution.

In the future, We aim to lead new innovations in document work by utilizing Vision models such as DLA, Table to HTML, Chart to Table, OCR, and the Document Foundation Model(DocFM) currently under development. We are working to develop a model capable of understanding documents in depth, along with applications that can convert image-based documents into editable formats for easy modification, and tools that convert documents into desired formats (e.g., PPT Style, Blog Style). Our goal is to maximize the efficiency of document processing and contribute to the further evolution of the work environment.


How Vision Lab Works


Image 7. How Vision Lab Works


At the Vision Lab, we are not only conducting research in the field of vision but also expanding our studies to various modalities. Our goal is to further extend our research toward world-class foundation model-based Embodied AI in the future.

At the Vision Lab, open sharing of ideas and close collaboration are core values. We freely share our thoughts, methods, and ideas, leading to more creative and efficient solutions. We also recognize that our colleagues' work is the work of the team, and we don’t hesitate to ask for help when needed, discuss challenges, and work through them together. This collaborative approach allows us to maximize each other’s strengths and effectively tackle even the toughest challenges. 

Above all, we foster an atmosphere of trust— valuing each other’s expertise and opinions, and actively giving and receiving feedback to grow. Thanks to this culture, we maintain a flexible and collaborative research environment that is continuously evolving.


Work-Life Balance at LG AI Research

 

Image 8. Ahra Jo's child, who attends the company daycare

 

Life as a working mom at LG AI Research is quite fulfilling. My child is now 50 months old, and we've been commuting together since my child was 20 months old, using the company daycare for about 30 months. Having childcare nearby has allowed me to quickly pick up my child in case of sudden illnesses, and it's reassuring to see them running around through the window during playtime.

Seoul Botanic Park, nestled within the bustling metropolis of Seoul, provided an exceptional environment for urban children to experience the changing seasons and learn about nature firsthand. The D&O daycare center, supported by LG AI Research, is located right within this park. In spring, the children are surrounded by a vibrant array of flowers, while in summer, they run around and tan like kids from the countryside. Come autumn, they collect acorns and colorful leaves, and in winter, they joyfully play on untouched snow-covered lawns.

Among the many benefits offered by LG AI Research, I personally consider the daycare support to be the most outstanding. Now that my child has reached the age to transition to kindergarten, they are starting a new journey in a different place. However, after just one day at the new school, my child already misses the pristine environment of the daycare and the outdoor playtime at Seoul Botanic Park. IT made me reflect once again on how great the environment was. 

The flexible work system also allows me to adjust my commute when my child is sick or when other childcare issues arise, which is a huge advantage as a working mom. I try to live each day as best I can so that when my child grows up, they’ll be proud of the mom I am.

 

Descriptors Before My Name

 

Image 9. Descriptors Before My Name

 

I aim to create impactful research that makes LG AI Research the first name that comes to mind in the field of Document AI. In doing so, we want to not only lead the way in technological innovation but also conduct research that provides value to real customers. I continue my research with the goal of one day seeing my name proudly recognized in these achievements.

 

Our Goals for 2025

From a research perspective, we aim to build a Document Foundation Model through research and develop a highly effective model with minimal hallucination by integrating it with the Vision model. This will help enhance the precision of Document AI technology and create reliable AI models.

From a business perspective, we plan to focus on expanding the API business based on the models we have already developed, as they are currently available as APIs. Furthermore, we aim to integrate all APIs into ChatEXAONE to provide more powerful AI services, ultimately building our own Document AI solution. 

With these research and business achievements, we look forward to a year of delivering innovative value to our customers and establishing a unique position in the Document AI field.