son_profile_38479e3d1.png Sungryull Sohn 2025.02.17

조정원님_편집_b065f1471.jpg Jeongwon Jo 2025.02.17

최지훈님_편집_05b206241.jpg Jihoon Choi 2025.02.17

김재겸님_5263a36d1.jpg Jaekyeom Kim 2025.02.17

Unveiling NEXUS : The Agent AI System for Tracking Legal Risks in AI Training Datasets

When AI technology first emerged, people’s primary interest was, "How far can AI go?" AI researchers focused on developing more powerful AI models, and as a result, AI technology experienced explosive growth, proving its ability to turn various possibilities into reality.

However, this rapid advancement also brought several unintended consequences. Issues arose where AI models, due to indiscriminate data training, produced unreliable, unethical, or even infringement of rights outcomes. Consequently, the key discussion in AI technology has now expanded beyond just "How powerful is the AI model?" to "How trustworthy and safe are the AI-generated results and their underlying reasoning?"

To address these issues, LG AI Research has developed an Agent AI system that tracks the life cycle of training datasets used in AI models, comprehensively analyzes legal risks, and assesses potential threats related to dataset usage. Along with the research outcomes[1], LG AI Research is also introducing "NEXUS", where users can directly explore results generated by this Agent AI system. 

Now, here’s a look at LG AI Research’s Agent AI technology and its Data Compliance framework embedded in NEXUS.


LG’s Agent AI focuses on data compliance 


Image 1. NEXUS, unveiled by LG AI Research (nexus.lgresearch.ai)


Recent advances in artificial intelligence (AI) have fundamentally transformed our society, but the challenges we face with its use are still unprecedented. As AI expands into all areas of industry and everyday life, the biggest challenge is the advancement of this technology in a legal, safe, ethical, and responsible manner.

We focused on the training data that underlies AI models. We decided it was important to ensure legal transparency and stability of datasets and started thinking about how we could achieve that. One of the characteristics of modern AI datasets is that they have a unique and complex layer structure.  A dataset has a complex relationship with many sub-datasets and dependent datasets (Dependencies) that are the source of its data.

LG’s Agent AI technology found that key AI learning datasets were being redistributed so many times that it was impossible for humans to track their sources, and each redistribution was linked to dozens or even hundreds of datasets. This means that a single dataset can be linked to thousands or even tens of thousands of separate data sources. The complexity of a dataset can result in significant transparency and legal risks, as identifying these risks manually in a large number of intricately linked datasets is challenging.

That’s why we’ve been working on Agent AI technology that tracks the lifecycle of complex datasets to analyze and assess legal risk, and why we’ve released NEXUS, a project that enables direct use of this technology.


LG’s Agent AI automatically analyzes the layer structure of datasets


Image 2. NEXUS Analyzes the Layer Structure of Datasets


A key feature of the Agent AI technology is its ability to find and automatically analyze even complex layers in datasets. LG’s Agent AI tracks the changes in rights relationships that occur as datasets are redistributed and combined, with broad scalability. The system automatically detects possible license conflicts, rights relationship inconsistencies, and the inclusion of personal information at each step and provides a detailed risk assessment.

LG’s Agent AI recognizes a dataset as a complex tree structure with datasets and dependencies, rather than just a collection of information. It assesses the legal risks that may arise at each node separately and automatically analyzes the impact of these risks on the parent node.


Unrivaled Competitiveness : Accurate, fast, and cost-effective

The Agent AI technology applied to NEXUS is a comprehensive data compliance solution developed based on LG AI Research’s EXAONE 3.5 model. The Agent AI technology consists of three core modules: Navigation, QA, and Scoring, each of which has been fine-tuned in different ways.


  1. The navigation module : Intensively trained to navigate web documents and analyze text using synthetic (artificially generated) data. Specifically, the Navigation module performs navigation based on the name and type of a given entity (e.g., dataset, software, content, etc.) to find links to web pages or license documents related to that entity.

  2. The QA module : Then trained to take the collected documents as input and efficiently extract dependency and license information from within the documents.

  3. The scoring module : Trained using a refined dataset labeled by lawyers, which performs a complex process of understanding the details of the license terms and combining them with the entity’s typology-metadata to score and assess the legal risk the entity may pose.


Through the organic interplay of these components, Agent AI technology has achieved a high degree of accuracy and speed. It can complete tasks more than 45 times faster than a human expert, at a cost savings of more than 700 times. It’s also about 16.85%p more accurate in identifying sub-datasets and dependent datasets (Dependencies) and 8.1%p more accurate in identifying license documents than a human expert, delivering superior performance across time, cost, and accuracy.


Image 3. Accuracy of finding the sources and license terms of 216 datasets


In fact, when evaluating a randomly selected 216 of the top 1,000+ downloaded hugging face datasets, NEXUS achieved approximately 81.04% accuracy in identifying dependencies and 95.83% accuracy in identifying license documents (primarily in the form of URLs). These results not only outperform other large language model services for web search in similar environments, but also exceed the average scores of lawyers previously trained on similar tasks. This difference in performance is due to our Agent AI technology’s ability to retrieve specific information from the web, as well as the sophisticated algorithms in the scoring module that analyze license terms and convert them into quantitative scores.


Image 4. Efficiency Comparison between AutoCompliance Agent and Human expert


Also, in terms of speed and cost savings, the entire process of finding dependency and license information for a single dataset, synthesizing it, and assessing legal risk takes only about 53.1 seconds on average. This is significantly faster than the same task being performed by a human reviewer (which takes over 2,400 seconds on average), and at a labor cost of only $0.29 (around KRW 300) per review, it represents a significant cost savings over a professional review. As such, our Agent AI technology and NEXUS is an innovative technology that combines speed, accuracy, and cost-effectiveness for industrial field, research institutions, and legal departments that need to review licenses at scale.


LG’s Data Compliance Framework for Determining Legal Risks

The legal risk assessment for datasets is based on the data compliance framework developed by LG AI Research. This framework comprehensively assesses the legal safety of datasets through 18 key assessment items. These include whether a license is granted, the right to modify the data, the right to create derivative works, the potential for infringement of original copyright in the output, and privacy concerns. Each item has a designed weighting based on real-world legal dispute examples and case law analysis, resulting in a realistic and reliable risk assessment.

Data compliance assessment results are categorized into a seven-level risk rating system, which operates on the following detailed criteria.

Image 5. LG AI Research’s Data Compliance Framework

 

The A-1 rating is the highest and safest and requires two conditions. First, it must have explicit commercial use permission or be in the public domain. Second, all sub-datasets must also have the same level of rights permission.

A-2 to B-2 ratings indicate datasets that are safe to use within a limited scope. For example, they can be used freely for research purposes but have certain restrictions on commercial use.

C-1 through C-2 are datasets that fall into a “gray zone.” These are categories where there is a high risk of unclear license status, ambiguous rights relationships with lower-level datasets, potentially infringing copyright, data containing sensitive personal information, or outright violations of license terms, web crawls, disputes, etc.

 

The Future Envisioned by LG’s Agent AI and NEXUS

While LG’s Agent AI technology and NEXUS have set a new standard for the legal stability of AI datasets, we believe there’s still a long way to go. In fact, LG AI Research conducted an in-depth analysis of 3,612 major datasets through NEXUS and found that the inconsistency of rights relationships between datasets and dependencies is far higher than expected. 

Of particular concern is the fact that many of the datasets found to have these inconsistencies are learning data for major AI models in widespread use today. For example, of the 2,852 AI learning datasets that we determined to be commercially available, only 605 datasets, or 21.21%, were commercially available after accounting for the risk of dependencies.  Recognizing these real-world issues with the legal reliability of AI datasets, we will continue to work to improve our efforts to keep pace with the changing AI technology and legal environment.

Our first challenge is to dramatically expand the scope and depth of the datasets that our Agent AI technology analyzes. Beyond our current 3,612 datasets, we plan to grow to a scale that will allow us to understand the life cycle of all the data in the world. This will be an evolution that focuses on qualitative deepening, not just quantitative expansion.

The data compliance framework will also evolve into a global standard. The current 18 evaluation criteria will be continuously refined to reflect changes in AI models and data usage patterns. We plan to collaborate with the global AI community and legal experts to develop these criteria into an international standard. Through this initiative, we aim to ensure the legal stability of AI datasets, enabling developers to easily explore datasets suitable for their projects while receiving reliable risk assessments and mitigation strategies. In the long term, we will evolve NEXUS into a comprehensive legal risk management system for AI developer, contributing to the creation of a safe and responsible AI ecosystem.

 

참고
[1] Do Not Trust Licenses You See - Dataset Compliance Requires Massive-Scale AI-Powered Lifecycle Tracing