Interpretability

Explore Interpretability through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Safety

Articles that also belong to these categories. Counts cover all of Interpretability.

Showing 1-6 of 6 articles

Activation steering

Activation steering is a family of inference-time techniques in mechanistic interpretability and AI safety that modify a neural network's internal activations to influence its behavior, without retraining the…

AI SafetyLarge Language Models

Goodfire AI

Goodfire AI is a San Francisco-based artificial intelligence research lab and public benefit corporation focused on mechanistic interpretability

AI CompaniesAI Safety

Mechanistic interpretability

Mechanistic interpretability (often abbreviated as mech interp or MI) is the field that reverse-engineers the internal computations of neural networks, particularly transformers

AI Safety

Persona vectors

Persona vectors are single linear directions in the activation space of a large language model that correspond to high level character traits such as evil, sycophancy, or a propensity to hallucinate.

AI SafetyLarge Language Models

Representation Engineering

Representation Engineering (often abbreviated RepE) is a top-down approach to artificial-intelligence transparency and control that reads and manipulates high-level concepts (such as honesty, harmlessness, and…

AI Safety