Attribution Graphs
Attribution graphs are a mechanistic interpretability technique developed by Anthropic that traces the internal "circuits" a large language model uses to turn a specific prompt into a specific output.
Explore Interpretability through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Interpretability.
Showing 1-8 of 8 articles
Attribution graphs are a mechanistic interpretability technique developed by Anthropic that traces the internal "circuits" a large language model uses to turn a specific prompt into a specific output.
Christopher Olah (commonly Chris Olah) is a Canadian machine learning researcher, a co-founder of Anthropic, and the researcher most often credited with founding mechanistic interpretability
A crosscoder is a mechanistic interpretability tool, introduced by Anthropic in October 2024, that generalizes the sparse autoencoder (SAE) and the transcoder by learning a single shared dictionary of sparse…
Golden Gate Claude was a temporary, research-oriented public demonstration released by Anthropic on May 23, 2024
On the Biology of a Large Language Model is a mechanistic interpretability paper published by Anthropic on March 27, 2025, in the Transformer Circuits Thread.
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet is the May 21
Towards Monosemanticity is an October 2023 mechanistic interpretability paper from Anthropic that used a sparse autoencoder to decompose the internal activations of a small language model into thousands of…
Toy Models of Superposition is a September 2022 mechanistic interpretability paper from Anthropic that shows how a neural network can represent more features than it has dimensions by packing them into…