AI Safety

Explore AI Safety through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Interpretability

Articles that also belong to these categories. Counts cover all of AI Safety.

Showing 1-6 of 6 articles

Mechanistic interpretability

Mechanistic interpretability (often abbreviated as mech interp or MI) is the field that reverse-engineers the internal computations of neural networks, particularly transformers

Interpretability

Representation Engineering

Representation Engineering (often abbreviated RepE) is a top-down approach to artificial-intelligence transparency and control that reads and manipulates high-level concepts (such as honesty, harmlessness, and…

Interpretability