Anthropic

Explore Anthropic through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Interpretability

Articles that also belong to these categories. Counts cover all of Anthropic.

Showing 1-8 of 8 articles

Attribution Graphs

Attribution graphs are a mechanistic interpretability technique developed by Anthropic that traces the internal "circuits" a large language model uses to turn a specific prompt into a specific output.

Interpretability

Christopher Olah

Christopher Olah (commonly Chris Olah) is a Canadian machine learning researcher, a co-founder of Anthropic, and the researcher most often credited with founding mechanistic interpretability

InterpretabilityPeople

Crosscoder

A crosscoder is a mechanistic interpretability tool, introduced by Anthropic in October 2024, that generalizes the sparse autoencoder (SAE) and the transcoder by learning a single shared dictionary of sparse…

Interpretability

Towards Monosemanticity

Towards Monosemanticity is an October 2023 mechanistic interpretability paper from Anthropic that used a sparse autoencoder to decompose the internal activations of a small language model into thousands of…

AI ResearchInterpretability