
|
The Journey Toward Trustworthy AI: Introducing LG AI Research’s “AI Ethics Seminar” |
As generative AI rapidly spreads, AI is no longer a specialized tool, but rather, is becoming fundamental infrastructure for work. With productivity in repetitive tasks now becoming more common, we frequently face the question: “Where should we apply AI?”
This post takes the discussion a step further from that question, beginning from the perspective of “To what extent should human judgment and ways of working be extended through AI?”
First, this post examines human thought structures and working methods through the theories of two psychologists. Next, it explores when human-AI collaboration is effective based on the latest research. Finally, it introduces UX principles for designing AI not merely as a “tool” but as a “colleague” to achieve effective collaboration, along with real-world service examples like LG AI Research products.
1. How do humans think and work?
1-1. Fast thinking and slow thinking

Image 1. Daniel Kahneman's System 1 and System 2[1]
Our everyday judgments and decision-making processes (approximately 35,000 per day according to some estimates[2]) can be personified and explained as two distinct systems based on psychologist Daniel Kahneman's dual-process theory[3].
System 1: Fast, automatic thinking based on intuition and experience, making immediate judgments (e.g., writing familiar report formats, handling repetitive decisions)
System 2: Slow, effort-intensive, and complex thinking that involves analyzing and validating problems (e.g., dealing with complex, high-cost, high-risk decisions such as entering new markets or changing strategies)
According to Kahneman, System 2 typically operates in a “comfortable low-effort mode” and exhibits a characteristic “laziness,” avoiding expenditure of energy beyond what is strictly necessary. As a result, System 2 tends to readily accept the intuitive impressions and judgments generated by System 1 without sufficient scrutiny, making it highly susceptible to various cognitive biases during this process[3].
This premise holds significant implications for collaborative environments with AI. The moment AI swiftly and fluently presents a “plausible answer,” human System 2 instinctively wants to linger longer in that comfortable, low-effort mode.
1-2. Humans think narratively
Another psychologist, Jerome Bruner, viewed the self not as a fixed entity that can be simply retrieved from memory, but as something constructed through various mental processes, including memory[4].
In particular, he explained that culturally shaped cognitive and linguistic processes influence how we structure our experiences and narratives, determining what we perceive as experience, how we organize it in memory, and which moments we recognize as distinct “events” among countless others.
Ultimately, humans are “narrative beings” who interpret experiences in our own unique ways and weave them into stories. Bruner emphasizes that “In the end, we become the autobiographical narratives by which we ‘tell about’ our lives[5]”.
According to this perspective, the unique capabilities of AI, which are rapidly and accurately extracting patterns from large-scale data, and humans, which are interpreting their experiences by weaving them into narratives, are distinguished as follows.
AI response: “This customer has a 73% chance of leaving”
Human question: “Why are they leaving? What did our service mean to them?”
2. When is human-AI collaboration effective?
2-1. Will the synergy of human-AI collaboration always be guaranteed?
Will collaboration between AI and humans yield better results? Intuitively, it seems like it should. Combining AI's data processing capabilities with human contextual judgment is expected to produce superior outcomes compared to when each works independently.
However, actual studies warn that this intuition is not always correct. Particularly, the more logical and plausible the answer presented by AI appears, the more counterproductive the effect observed, as human judgment becomes dulled. This is explained by automation bias (AB), a cognitive tendency to over-rely on automated recommendations, leading to misjudgment[6].
"The AI & Society”[7] review, which comprehensively analyzed 35 human-AI decision-making cases, and the individual experimental studies featured within it, consistently demonstrate this pattern of automation bias.
A study by Bond et al.(2018)[8] on electrocardiogram (ECG) diagnosis reported that physicians' accuracy dropped significantly when presented with inaccurate automated diagnoses (fellows: 84.86-41.66%, non-fellows: 86.38-27.43%).
In a study targeting radiologists, Dratsch et al. (2023)[9] also analyzed that incorrect interpretation suggestions from AI distorted the judgment of doctors regardless of their expertise, reducing accuracy (unexperienced : 79.7–19.8%, highly experienced: 82.3–45.5%), and in particular, the less experienced doctors were, the more likely they were to follow the incorrect suggestions from AI.
Similar patterns are observed in the realm of business decision-making. Keding and Meissner(2021)[10] conducted a scenario experiment involving 150 senior executives regarding R&D investment decisions, and reported that participants tended to trust AI-generated advice more than human advice, exhibiting overreliance on AI-based recommendations, including a higher likelihood of investment action when AI recommendations were provided.
Putting these findings together, they suggest that human-AI collaboration does not always guarantee synergy in high-risk decision-making, and when automation bias intervenes, it may actually lead to diminished performance.
2-2. Conditions under which collaboration performance declines as AI becomes more advanced
Interestingly, a meta-analysis comparing the performance of humans alone, AI alone, and human-AI collaborations, published in Nature Human Behaviour (2024)[11], demonstrates this point more clearly with numerical data.
The research team conducted a meta-analysis of 370 effect sizes reported across 106 experimental studies. They found that while an “augmentation” effect clearly exists where human-AI combinations outperform humans alone—no “synergy effect”, in which performance exceeds that of whichever performs better between humans or AI individually, was observed on average.

Image 2. Forest plots of all effect sizes (k = 370) included in the meta-analysis[11].
Furthermore, the higher the AI's standalone performance relative to humans, the more the average performance of human–AI collaboration tends to decline, and conversely, when humans outperform AI, collaboration performance tends to increase.
The results of two tasks by Cabrera et al. (2023)[12] support this.
“Fake hotel review detection” task: This is a language-based binary classification problem. While AI (73%) outperformed humans (55%), the collaborative model actually achieved lower accuracy (69%).
“Bird image classification” task: A multi-class classification problem based on visual recognition. Humans (81%) outperformed AI (73%), and collaboration resulted in a synergy effect, achieving 90% accuracy.
The performance of human-AI collaboration cannot be explained simply by adding or averaging the scores of humans and AI separately. Instead, the collaborative effect varies depending on which party holds a relative advantage in specific tasks. In other words, this suggests that the design that distinguishes which roles are left to humans and which to AI can determine the effectiveness of collaboration.
3. UX principles for designing AI as a colleague
Above, this post examined how humans think and make judgments through the theories of two psychologists, and through the latest research, explored when human-AI collaboration enhances performance and when it actually diminishes it.
The more AI does the thinking for us, the more our human System 2 wants to rest, and automation bias becomes easier to reinforce. Therefore, we must not only refine AI models but also consider how to design structures within UX that prompt users to think again.
AI & Society researchers[7] point out that simple explainability (XAI) or transparency alone is not enough to solve this problem. In fact, poorly designed explanations themselves can become another heuristic, reinforcing automation bias by fostering trust while reducing doubt. Accordingly, this post summarizes the five human–AI collaboration design implications presented in this review from a UX perspective and introduce services where these principles are well-demonstrated, along with examples from LG AI Research.
Balancing simplicity and detail in explanations: Explanations that are overly simplistic can induce overreliance, while excessively complex explanations increase cognitive load, thereby reducing users' ability to evaluate AI outputs effectively. Microsoft Copilot in Excel's "show explanation" feature exemplifies a well-balanced approach to reconciling simplicity and detail. At the formula suggestion stage, it first presents a concise result, then allows users to selectively access detailed rationale when needed.
Promoting verification engagement: System design that encourages users to actively engage with explanations and rationale is critical. This process stimulates users' critical thinking, helping them review AI outputs rather than accepting them uncritically.
EXAONE Data Foundry, LG AI Research's end-to-end LLM tuning platform based on domain-specific documents, is designed to naturally integrate domain experts' judgment throughout the data generation and validation process. During the generation stage, survey-based data profiling reflects expert preferences, while in the validation stage, quality evaluation results are presented alongside evaluation criteria, enabling users to review outputs while examining the underlying rationale.

Image 3. EXAONE Data Foundry’s data validation stage
Adaptive and personalized explanation design: Explanations cannot be uniform, they must be tailored to users’ diverse domains, expertise, trust preferences, and contexts. Claude AI's artifact customization feature[13] is a prime example of anticipating domain-specific use cases through example output documents, extending beyond simple chat windows. It lowers psychological barriers and enhances accessibility.
Guided training and mental model formation: Through training, users should be guided on the system's scenarios and limitations, and an honest and realistic understanding of AI during tasks should be fostered.

Image 4. Project Albert's molecule generation page
LG AI Research's Project Albert is a service that supports materials researchers across the R&D process. Its molecule generation UX is designed with principles that support realistic understanding of AI model behavior. Specifically, to generate molecules matching desired properties, the system explicitly communicates the process in which the AI first generates virtual molecules in data-rich regions, then progressively expands its exploration range. Each exploration stage presents distinct advantages and limitations: Narrow (high prediction accuracy and stability), Intermediate (balanced accuracy and diversity), and Broad (high diversity with uncertainty). By clearly articulating and streaming these accuracy differences across stages (Narrow → Intermediate → Broad), the system fosters honest mental model formation.
Human-centered AI: rethinking the oversight model and future directions: Finally, AI should function as an assistant to human judgment —a second opinion—and control over decisions must remain with humans. In every AI system we design, it must be clear where the final decision-making authority lies. When AI suggestions function as the “default,” human critical review and active judgment may be weakened. GitHub's code autocompletion, which provides UX cues such as “continue,” “keep,” or “discard” that signal users to “pause and choose,” is a prime example of explicitly creating decision points for human judgment intervention.
Beyond a tool, to a colleague who extends our judgment
AI should be designed not to solve everything quickly, but to leave room for users to think and judge for themselves.
Creating AI experiences that extend rather than replace human judgment remains our challenge.
LG AI Research will continue to pursue UX design that extends human judgment based on an understanding of people, and practice AI ethics, ensuring humans remain at the center of decision-making.
AI Ethics Seminar 2025 Series
#1. [AI Ethics Seminar 2025 EP.1] How AI Is Changing Human Critical Thinking
#2. [2025 AI Ethics Seminar EP.2] Beyond bias, the journey to fair AI
#3. [2025 AI Ethics Seminar EP.3] Agentic AI Threat Modeling and Guardrail Implementation Strategy
[1] Bergamaschi Ganapini, M., Campbell, M., et al. (2025). Fast, slow, and metacognitive thinking in AI. npj Artif. Intell., 1, 27. https://doi.org/10.1038/s44387-025-00027-5
[2] Frank, Graff. (2021). “How Many Decisions Do We Make in One Day?” PBS North Carolina. https://www.pbsnc.org/blogs/science/how-many-decisions-do-we-make-in-one-day/.
[3] Daniel, Kahneman. (2011). Thinking, fast and slow. Farrar, Straus and Giroux. https://ia800603.us.archive.org/10/items/DanielKahnemanThinkingFastAndSlow/Daniel%20Kahneman-Thinking%2C%20Fast%20and%20Slow%20%20.pdf
[4] 1. Bruner J. (1994). The “remembered” self. In: Neisser U, Fivush R, eds. The Remembering Self: Construction and Accuracy in the Self-Narrative. Emory Symposia in Cognition. Cambridge University Press; 1994:41-54.
[5] Bruner, Jerome. (2004). Life as Narrative. Social Research: An International Quarterly. 71. 691-710. 10.1353/sor.2004.0045.
[6] Mosier et al. (1996). Automation Bias, Accountability, and Verification Behaviors. Proceedings of the Human Factors and Ergonomics Society Annual Meeting, 40(4), 204-208. https://doi.org/10.1177/154193129604000413 (Original work published 1996)
[7] Romeo, G., & Conti, D. (2025). Exploring automation bias in human-AI collaboration: a review and implications for explainable AI. AI & Soc. https://doi.org/10.1007/s00146-025-02422-7
[8] Bond, Raymond R., et al. (2018). "Automation bias in medicine: the influence of automated diagnoses on interpreter accuracy and uncertainty when reading electrocardiograms." Journal of electrocardiology 51.6 (2018): S6-S11.
[9] Dratsch, Thomas, et al. (2023). "Automation bias in mammography: the impact of artificial intelligence BI-RADS suggestions on reader performance." Radiology 307.4 (2023): e222176.
[10] Keding, C., & Meissner, P. (2021). Managerial overreliance on AI-augmented decision-making processes: How the use of AI-based advisory systems shapes choice behavior in R&D investment decisions. Technological Forecasting and Social Change, 171, 120970.
[11] Vaccaro, M., et al. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nat Hum Behav, 8, 2293?2303. https://doi.org/10.1038/s41562-024-02024-1
[12] Cabrera, Angel Alexander, et al. (2023). Improving Human-AI Collaboration With Descriptions of AI Behavior. Proc. ACM Hum.-Comput. Interact., 7(CSCW1), Article 136, 21 pages. https://doi.org/10.1145/3579612
[13] Anthropic. (2024). “Build Customer Personas.” Anthropic. https://www.claude.com/resources/use-cases/build-customer-personas