Yongmin Shin Yongmin Shin 2026.10.02

[ICML2026] Application of Mechanistic Interpretability: Understanding Graph Transformers (TokenGT)

Modern AI models have enabled us to achieve human-level performance in problem-solving. The key idea behind AI systems is to design algorithms that enable models to learn patterns autonomously from large amounts of data. Despite the effectiveness of this approach, the internal structures of these models are still determined by the algorithms, not by humans. Since the development of AI, there has also been consistent demand for explainability for such models, from everyday users to machine learning researchers.
 

The field of explainable AI (XAI) has been steadily developing, from the early days of attribution methods—drawing a heatmap to show the most influential parts of the input data—to counterfactuals (“what if”-type explanations), among other approaches that provide rationale on model behavior. Although most XAI methods can provide intuitive explanations, this does not necessarily mean that they accurately reflect the model[1] or provide a specific, detailed description of how computation actually occurred in the model and led to a particular outcome.
 

A relatively new and frontier branch of XAI named mechanistic interpretability (Mech. interp. for short) was developed to address this issue. Beginning with early work on reverse-engineering neural networks by recovering specific computational paths (“circuits”)[2,3], the field has made critical advancements and now focuses mostly on transformers and, eventually, LLMs. Notable works include the identification of induction heads[4], insights into grokking[5], the use of sparse autoencoders[6], automated circuit discovery[7], activation patching[8], and related topics. Mech. interp. has shown that it is possible to understand neural networks at a granular level and even semantically interpret abstract representation spaces in modern LLMs[9].
 

However, most research in mech. interp. has focused on understanding LLMs. Due to the nature of mech. interp., the findings of each study are usually confined to the specific target model, which has often been a language model. Little work has been dedicated to applying mech. interp. techniques to graph transformers, a significant model architecture often used to process molecular data in chemistry and biology applications. Our work, “Discovering Mechanisms in Tokenized Graph Transformers”[10], is a first step toward applying mech. interp. to these types of models.
 

The field of mech. interp. has a highly active community, and the Mechanistic Interpretability Workshop has become one of the central venues that brings together researchers from around the globe. This is the third edition of the workshop, following previous events at ICML 2024 and NeurIPS 2025, and LG AI Research had the opportunity to present our work to the community for the first time.

1. Our work
1-1. Background and Motivation

The utility and benefits of XAI methods depend on the target audience. Specifically, because mech. interp. aims to reverse-engineer neural networks, it provides a highly fine-grained understanding of the target model. This benefits researchers who require deep knowledge of neural networks, while also enabling the extraction of actual knowledge learned by the model[11]. These benefits are also relevant to AI4Science research, especially because we often need both a full understanding of neural networks and the ability to gain scientific knowledge from them.

To achieve this, we explored how Mech. Interp could be adapted to models and data used in the chemical domain. Specifically, we need a basic understanding of models that process graph inputs, which are often used to represent chemical molecules. This remains a vastly underexplored problem, and we chose to initiate this research direction in a clean and minimal setup. We ultimately chose to study one of our previous graph transformer models, TokenGT[12], in a plain graph setting, without molecular graphs at this stage.

1-2. Setup to the problem – TokenGT and datasets
To investigate how a vanilla transformer such as TokenGT (specifically, a T5 encoder) processes graphs represented as sequences of node and edge tokens, lets first explain the setup. Specifically, we generated a synthetic graph dataset consisting of simple undirected, unweighted graphs with no node or edge features. The model learns to understand graphs independently of node indices by shuffling node IDs in each batch during training. See the figure for a detailed illustration.


Figure 1. (a) The same graph can be represented multiple ways by permuting node IDs.
(b) We consider the case where nodes and edges are tokenized by assembling type (node/edge) and ID vectors.
(c) Task-specific heads are applied to the node outputs to solve graph tasks.
Note that we train separate models per task although they are shown together in the figure for illustrative purposes.

 
We analyze models trained on three fundamental graph tasks: (1) degree calculation, (2) cycle detection, and (3) shortest-path distance calculation. Across all three tasks, the models achieve nearly 100% validation accuracy, demonstrating that they successfully learn each task.

1-3. Findings in Our Analysis
A central finding is that all three models begin with a shared local-structure computation. The most prominent example is encoding the degree of a node token. In the first transformer layer, node tokens attend more strongly to edge tokens whose endpoint identifiers (IDs) match the node identifier. This ID-matching behavior allows the model to recover which edges are incident to each node, even though the architecture contains no built-in graph operations (unlike graph neural networks). Our query-key analysis shows that this mechanism is distributed across nearly all attention heads in the first layer and remains consistent across the three task-specific models.

Figure 2. An illustration on how the model solves the degree counting task.
The attention in early layers (specifically layer 0 in our case) can identify node ID occurrences in the edge token (blue token in the middle) and assigns strong attention to them.
The degree information is then written in a specific direction in the residual stream, while the magnitude of the write vector is very proportional to the ground truth degree
(arrows in the bottom part of the figure).


As mentioned, this incident-edge retrieval produces an early degree-like representation. Linear probes show that node degree is poorly recoverable from the initial embeddings but can be accurately identified immediately after the first attention operation. We can also obtain causal evidence through direct intervention using activation patching: modifying the residual stream along the identified degree direction changes degree predictions and also affects performance on the more complex ring-membership and shortest-path tasks. This suggests that the models reuse the same early local feature rather than learning each task independently from scratch. Naturally, this easily solves the degree counting problem (see figure 2). Empirically, we observe additional evidence: the model allocates more attention mass to incident edges as node degree increases, and the combined output of several attention heads writes proportionally into a degree-related direction. 

Figure 3. An illustration on how the model solves the ring membership problem.
It first identifies all nodes as part of a ring structure and flips the prediction when it gathers evidence that it is not part of a ring.


For ring membership, the model does not appear to implement an explicit cycle-tracing algorithm. Instead, we need to examine the results of ablation experiments, which involve removing key model components one at a time to observe their effects (see figure 3). The results show that the model initially classifies virtually all nodes as the ring class and then progressively gathers evidence that  a node is not part of a ring. Information related to leaf nodes and leaf neighbors can be accurately identified using linear probes on an important second-layer attention head. Furthermore, removing the second layer mainly damages non-ring predictions, supporting the view that the model accumulates non-ring evidence rather than constructing a general cycle detector.

Figure 4. (a) The attention map of L2:H2, which shows that the model views the input graph as a rooted BFS tree.
(b) If we replace the attention map by a code-generated one-hot attention map mid-inference, the model performance largely stays the same. 

 

For shortest-path distance, the model uses a more sophisticated mechanism. Early layers first encode local structural features such as degree and neighborhood statistics. Then, the third attenion head in the third layer then concentrates on one adjacent node that frequently matches a parent in a breadth-first-search-like tree (in other words, the model views the input graph as a BFS tree that is rooted at a particular node. See figure 4.) As evidence of this mechanism, ablating this single head causes a large drop in accuracy, while replacing its attention target with the corresponding BFS parent largely preserves performance. Probe results further indicate that the head copies selected structural information from this source node. The final layer then refines the resulting representation, with ablations showing a much larger effect on long-distance predictions than on nearby node pairs.

Overall, the experiments reveal a shared computational pattern. The model first reconstructs graph incidence through node-ID and edge-endpoint matching, converts this information into local degree-like features, and then composes those features differently for each task. These findings provide an initial mechanistic account of how transformers perform graph computation without graph-specific architectural biases.
 

2. The Next Step for us

Naturally, the extenion of this line of work is towards understanding models that processes realistic inputs. Specifically, the materials intelligence lab in LG AI research often deal with molecular graphs, in order to solve real-world problems such as molecular property prediction models and reaction prediction models. In order to produce the correct prediction, we can expect that the model has to learn (1) how to understand molecular graphs and (2) chemical facts and rules, but explicitly revealing and confirming this is yet to be done. Although application of mech. interp. techniques in the AI4Science domain is a current and active subject, we have also found a concensus among researchers on the necessity to expand mech. interp. into various modalities (including graphs) during this event.
 

We expect that LG AI research will be a key player in the domain of (mechanistic) interpretabilitiy in AI4Science in the future. As we progress in our research, the materials intelligence lab aims to revisit popular chemical models and understand their inner structures. We hope that this will open up the dialogue on alignment between AI models and domain scientists, as well as research towards gathering scientific knowledge via mech. interp. analysis.

참고

[1] Adebayo et al., “Sanity Checks for Saliency Maps”, NeurIPS 2018

[2] “Thread: Circuits” (2020/03/10). Distill. https://distill.pub/2020/circuits/

[3] Elhage et al. (2021/12/22) “A Mathematical Framework for Transformer Circuits”, Transformer circuits thread. https://transformer-circuits.pub/2021/framework/index.html

[4] Olsson et al. (2022/03/08) “In-context Learning and Induction Heads”, Transformer circuits thread. https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html

[5] Nanda et al., “Progress measures for grokking via mechanistic interpretability”, ICLR 2023

[6] Cunningham et al., “Sparse autoencoders find highly interpretable features in language models”, arXiv 2309.08600

[7] Conmy et al., “Towards automated circuit discovery for mechanistic interpretability”, NeurIPS 2023

[8] Heimersheim & Nanda, “How to use and interpret activation patching”, arXiv 2404.15255.

[9] Lieberum et al., “Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2”, arXiv 2408.05147.

[10] Shin et al., “Discovering Mechanisms in Tokenized Graph Transformers”, ICML Workshop on Mechanistic Interpretability, 2026

[11] Schut et al., “Bridging the human?AI knowledge gap through concept discovery and transfer in AlphaZero”. PNAS, 122 (13) e2406675122, https://doi.org/10.1073/pnas.2406675122 (2025).

[12] Kim et al., “Pure Transformers are Powerful Graph Learners”, NeurIPS 2022