Interpretability

Explore Interpretability through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Transformer Models

Articles that also belong to these categories. Counts cover all of Interpretability.

Showing 1-2 of 2 articles

Induction Heads

Induction heads are a circuit pattern in Transformer language models in which a small set of attention heads, typically spread across two layers, perform an in-context "match and copy" operation that completes…

Transformer Models

Logit lens

The logit lens is a foundational technique in mechanistic interpretability for inspecting the intermediate computations of transformer language models.

Transformer Models