Large Language Models

Explore language models, how they work, and the techniques used to build applications with them.

Explore articles

Reset filters
Browse subtopics: Interpretability

Articles that also belong to these categories. Counts cover all of Large Language Models.

Showing 1-4 of 4 articles

Activation steering

Activation steering is a family of inference-time techniques in mechanistic interpretability and AI safety that modify a neural network's internal activations to influence its behavior, without retraining the…

AI SafetyInterpretability

Patchscopes

Patchscopes is an interpretability framework for inspecting hidden representations of large language models by patching an internal activation from a source computation into a separate target inference whose…

Interpretability

Persona vectors

Persona vectors are single linear directions in the activation space of a large language model that correspond to high level character traits such as evil, sycophancy, or a propensity to hallucinate.

AI SafetyInterpretability

Refusal direction

The refusal direction is a finding from mechanistic interpretability research that the refusal behavior of safety fine-tuned chat language models is mediated by a single

AI SafetyInterpretability