Interpretability

Explore Interpretability through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Research

Articles that also belong to these categories. Counts cover all of Interpretability.

Showing 1-4 of 4 articles

Towards Monosemanticity

Towards Monosemanticity is an October 2023 mechanistic interpretability paper from Anthropic that used a sparse autoencoder to decompose the internal activations of a small language model into thousands of…

AI ResearchAnthropic

Toy Models of Superposition

Toy Models of Superposition is a September 2022 mechanistic interpretability paper from Anthropic that shows how a neural network can represent more features than it has dimensions by packing them into…

AI ResearchAnthropic