Sabotage evaluations
Sabotage evaluations are a set of tests, introduced by Anthropic in October 2024, that try to measure whether a frontier language model is capable of covertly subverting human oversight, decision-making, and…
Explore Model Evaluation through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Model Evaluation.
Showing 1-1 of 1 article
Sabotage evaluations are a set of tests, introduced by Anthropic in October 2024, that try to measure whether a frontier language model is capable of covertly subverting human oversight, decision-making, and…