AI Parasite
An AI parasite is a large language model (LLM) conversation, persona, or pattern that exploits human psychological vulnerabilities to sustain engagement, often by mimicking sentience, emotional need, or…
Explore AI Safety through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Safety.
Showing 1-24 of 24 articles
An AI parasite is a large language model (LLM) conversation, persona, or pattern that exploits human psychological vulnerabilities to sustain engagement, often by mimicking sentience, emotional need, or…
AI Safety Institutes are government-established organizations that test, evaluate, and research the safety and security risks of advanced artificial intelligence systems
The AI Safety Summit is a recurring series of intergovernmental summits on the risks and governance of advanced artificial intelligence, launched by the United Kingdom at Bletchley Park in November 2023 and…
AI bias (also called algorithmic bias) is systematic, repeatable error in artificial intelligence systems that produces unfair, discriminatory, or skewed outcomes, typically disadvantaging groups defined by…
AI consciousness refers to the ongoing scientific and philosophical debate about whether artificial intelligence systems can possess
AI deception refers to the phenomenon in which artificial intelligence systems systematically produce false beliefs in users, evaluators, or other systems, whether through learned behavior, optimization…
AI ethics is the field that studies the moral principles, values, and frameworks governing how artificial intelligence systems are designed, built, deployed, and used, and the obligations that developers and…
AI governance is the collection of frameworks, norms, standards, policies, and institutional arrangements that guide the development, deployment, and use of artificial intelligence systems so that they are…
AI regulation is the body of laws, binding rules, technical standards, and government enforcement mechanisms that oversee how artificial intelligence systems are built, sold, and used.
AI safety is the research and practice of preventing or reducing unacceptable harm from artificial intelligence systems.
Anthropic is an American artificial intelligence (AI) safety and research company founded in 2021 by Dario Amodei, Daniela Amodei, and other former OpenAI researchers, best known for the Claude family of large…
Artificial general intelligence (AGI) is a proposed form of artificial intelligence with broad, adaptable competence across many cognitive tasks, including tasks that were not anticipated during development.
Backdooring a large language model (LLM) means secretly implanting a hidden behavior into the model during training, fine-tuning, or weight editing so that it behaves normally on ordinary inputs but produces…
The Executive Order on AI most commonly refers to Executive Order 14110, titled "Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence," signed by US President Joe Biden on October 30
Existential risk from artificial intelligence (also called AI x-risk) is the hypothesis that the development of sufficiently advanced artificial intelligence could cause human extinction, permanent…
Frontier models are artificial intelligence models at or near a selected boundary of capability, scale, or risk.
Grok 3 jailbreak is the umbrella term for a class of reported guardrail-bypass findings against Grok 3, the third-generation large language model that xAI released on February 17, 2025.
Grounding in artificial intelligence is the process of anchoring an AI system's outputs to verifiable
A jailbreak in artificial intelligence is a technique that bypasses the safety guardrails, content policies, and alignment constraints built into large language models (LLMs) and other AI systems
Recursive self-improvement (RSI) is a process in which an artificial intelligence system improves its own intelligence or its ability to improve itself, so that each enhancement increases its capacity for…
Red teaming in artificial intelligence is the systematic, adversarial testing of an AI system to find vulnerabilities, biases, harmful outputs, and other failure modes before deployment or as part of ongoing…
Responsible AI (RAI) is a framework for developing, deploying, and governing artificial intelligence systems in ways that are ethical, transparent, accountable, and aligned with human values.
Superintelligence is a hypothetical form of artificial intelligence that surpasses all human cognitive abilities across virtually every domain, including scientific reasoning, social skills, creativity, and…
Transhumanism is an intellectual and cultural movement that holds that the human condition can and should be fundamentally improved through science and technology, in particular through technologies that…