AI Safety

Explore AI Safety through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: OpenAI

Articles that also belong to these categories. Counts cover all of AI Safety.

Showing 1-7 of 7 articles

An Alien Mind

An Alien Mind is an essay by Jakub Pachocki, the chief scientist of OpenAI, published by OpenAI on September 6, 2026 under its Safety and Research sections.

AI ResearchOpenAI

Cybersecurity ChatGPT Plugins

Cybersecurity ChatGPT Plugins were a small, informal grouping of third-party extensions for ChatGPT that focused on security related tasks during the brief life of the ChatGPT plugins beta.

ChatGPTOpenAI

Model Spec

The Model Spec is a public document published by openai that defines the intended behavior of the company's language models: how they should follow instructions, when they should refuse a request, how to…

AI AlignmentOpenAI

OpenAI Moderation API

The OpenAI Moderation API is a free classification tool, accessed through a dedicated moderation endpoint of the OpenAI API, that assesses whether text (and, for newer models

Developer ToolsOpenAI

Preparedness Framework (OpenAI)

The Preparedness Framework is the risk-management policy maintained by openai for tracking, evaluating, forecasting, and mitigating catastrophic risks from frontier artificial-intelligence models.

OpenAI

Rule-Based Rewards (RBR)

Rule-Based Rewards (RBR) is a safety-alignment technique introduced by OpenAI in July 2024 that replaces large quantities of human-labeled safety preference data with an explicit collection of natural-language…

AI AlignmentOpenAI