Anthropic Insights
Anthropic Insights, formerly called Clio, is a privacy-preserving conversation-analysis system developed by Anthropic. It uses Claude models and statistical clustering to turn large samples of conversations into aggregate descriptions, hierarchies, and measurements. Anthropic announced Clio in December 2024 and renamed the system and research tool Anthropic Insights in August 2026.[1][2]
In its research configuration, the system processes source conversations inside a restricted environment and exposes aggregate outputs rather than raw transcripts to analysts. A separate safety use can support review by a small group of authorized staff. These access models should not be conflated: external researchers in Anthropic's 2026 pilot did not receive raw conversations, user identifiers, or organization identifiers.[2][3][4]
Anthropic Insights is an analysis method rather than a survey instrument or a record of individual behavior. Its cluster labels are generated interpretations of conversation samples. Anthropic's own documentation says they can be incomplete, unstable, or misleading, and should not be treated as objective measurements without additional validation.[2][4][5]
Key facts
| Field | Detail |
|---|---|
| Developer | Anthropic[1][2] |
| Original name | Clio, short for "Claude insights and observations"[1][2] |
| Initial publication | December 2024[1][2] |
| Renamed | August 2026[1] |
| System type | Privacy-preserving aggregate analysis of AI conversation data[1][2] |
| Main operations | Facet extraction, semantic clustering, cluster description, hierarchy construction, and interactive exploration[2] |
| Formal privacy guarantee | None documented for the original system; the paper describes empirical and statistical defense in depth rather than differential privacy or k-anonymity[2] |
| 2026 external pilot | Studies led by Stanford SALT Lab, Oxford's Human Information Processing Lab, and METR[3][4] |
| Public pilot release | Aggregate cluster outputs under CC BY 4.0; no raw conversations or identifiers[5] |
Development and purpose
Anthropic developed Clio to find patterns in real-world large language model use without requiring analysts to read individual conversations. Traditional evaluations and red-team exercises begin with a behavior to test. Clio was designed as a complementary, bottom-up method that could surface recurring topics or behaviors not specified in advance. The authors proposed research, product analysis, and AI safety monitoring as applications.[1][2]
The 2024 paper demonstrated the system on one million Claude.ai Free and Pro conversations sampled from October 17 through October 24, 2024. It also described separate multilingual, safety-classifier, and privacy-evaluation runs. The published general-use analysis excluded Team, Enterprise, business, and Zero Retention traffic, so its results did not represent all Claude customers or all AI-assistant users.[2]
Clio later became part of the Anthropic Economic Index. The index initially used it to map approximately one million Claude.ai conversations to tasks in the United States Department of Labor's O*NET taxonomy.[9] Anthropic Insights is the general analysis system; the Economic Index is one research program that uses it.
How the system works
The 2024 design has five stages.[2]
- Facet extraction: Claude answers a question about each conversation, such as its topic, language, or number of turns. Some facets are computed directly, while others are model-generated summaries or classifications.
- Semantic clustering: For an open-ended facet, the system embeds the answers and uses k-means to group similar answers. Categorical, multiple-choice, and numeric facets can instead be counted or summarized.
- Cluster description: A model produces a name and description for each group from sampled answers, with instructions to omit private information.
- Hierarchy construction: Related clusters are grouped recursively so an analyst can move from broad categories to narrower ones.
- Exploration: A map and tree interface let analysts examine cluster size, language, safety scores, and other facets.
The word "facet" refers to the question or attribute applied to each conversation. In the 2026 external-research pilot, partner researchers wrote facets for their own studies. Anthropic tested them on the public WildChat corpus, ran the approved questions on private Claude data, and returned aggregate clusters or tallies. External researchers analyzed those outputs but did not run code against, inspect, or download the source conversations.[3][4]
This design separates open-ended discovery from predefined measurement. Open-ended answers can reveal themes that researchers did not anticipate, but their clusters are not mutually exclusive natural categories. A categorical facet is more suitable when the aim is to estimate the prevalence of defined behaviors, provided the classifier itself has been validated for that question.[4][5]
Privacy design and access boundaries
The original paper described four privacy layers. Claude is first instructed to omit private details from each conversation-level facet. Clusters below minimum thresholds for both conversations and distinct accounts are removed. The model is again told to exclude private details when describing retained clusters. A final model-based auditor screens the descriptions and removes clusters it judges too identifying.[2]
Anthropic explicitly did not claim a formal guarantee for this pipeline. The paper said rich, open-ended text made established methods such as differential privacy and k-anonymity difficult to apply, and characterized the design as an empirically tested set of overlapping defenses. It also stated that no single layer could perfectly guarantee privacy.[2]
Raw conversations are processed in a restricted environment. The 2024 paper said a small number of authorized employees could access individual records for limited purposes such as policy enforcement or auditing safety infrastructure, while aggregate research outputs were available more broadly. Anthropic's safety application can also identify accounts for further authorized review. This is different from the research interface and from the 2026 pilot, in which all raw data and computation stayed on Anthropic servers.[1][2][4]
Anthropic reported several internal evaluations. On 19,476 synthetic general conversations with known topic categories, the system reconstructed the categories with 94% accuracy. A model-based privacy auditor scored 98% on a hand-designed 1,237-example test set. In a 5,000-conversation pipeline run, Anthropic reported that none of the final cluster summaries fell into the three least-private levels of its five-level rubric. These results test particular synthetic data, prompts, models, and internal rubrics; they are not an independent privacy certification or a guarantee for every facet and deployment.[2]
External-research pilot
Anthropic began a pilot with three outside groups in February 2026 and published its account of the program in August. Stanford's Social and Language Technologies Lab studied task criticality, human agency, learning, and friction in Claude.ai conversations. Oxford's Human Information Processing Lab studied users' reported experience and model behavior. METR studied time savings and task outcomes in Claude Code sessions.[3][4]
Each study used a separate random sample of roughly 250,000 conversations from April and May 2026. Stanford and Oxford used Claude.ai data from Free, Pro, and Max users. METR used consumer Claude Code data from users who had opted in to model-improvement use. The samples excluded Team, Enterprise, and API customers. They were one-time snapshots, not representative samples of the public or Anthropic's full user base.[4][5]
The partners chose their questions and analysis. Anthropic handled sampling and computation and reviewed study plans for privacy, safety, confidential information, and methodological accuracy. The appendix says partners remained free to publish findings that were inconvenient to Anthropic. After the runs, every cluster name and description received manual review before release.[4]
That review removed or redacted a small portion of output. For Stanford, 1.9% of clusters were affected, representing 4.28% of conversations. For Oxford, 3.33% of clusters and 3.85% of conversations were affected, and Anthropic removed one entire facet because it suspected the prompt produced misleading descriptions. For METR, 1.8% of clusters and 2.96% of conversations were affected. These percentages describe the review process, not an estimate of policy violations among users.[4]
Anthropic also commissioned a two-week privacy red-team exercise by Imperial College London's AI Security and Privacy Lab. The team did not reidentify a person or otherwise violate Anthropic's stated threat model in the released pilot outputs. It did infer that one cluster concerned use of a popular open-source project, an organization-level inference outside that threat model. Anthropic said it would respond with higher aggregation minimums, less distinctive phrasing, and repeat testing. The exercise provides evidence about the tested release under a defined threat model, not proof against all attacks.[4]
Stanford SALT study
The Stanford team published a preprint based on 249,834 conversations. Its abstract reports that human-led collaboration dominated, Claude was classified as actively teaching in 67% of conversations, and some form of friction appeared in roughly half. It also reports that conversations classified as higher stakes were longer and involved more active user engagement.[6]
These are observational classifications produced through researcher-designed facets and aggregate analysis. They do not establish that Claude caused learning, improved capability, or produced better real-world outcomes. The paper was a preprint at the research cutoff, and Anthropic's appendix cautions that the original system's accuracy evaluation covered conversation topics, not facets about model behavior or emotional state.[4][6]
Oxford and METR had not completed public writeups when Anthropic announced the pilot. Anthropic summarized early patterns from their work, but labeled the METR findings preliminary and said both analyses were continuing. Their research questions can therefore be described, but their early results should not be treated as established conclusions.[3][5]
Released aggregate data
Anthropic released the exact aggregate outputs given to the three partners in a public dataset. It has stanford, oxford, metr, and metr_addendum subsets; the addendum used a fresh sample and stricter aggregation minimums. The release is licensed under CC BY 4.0.[5]
Each row describes a cluster for one facet. Fields include a Claude-generated cluster name and description, hierarchy level, conversation and organization counts, proportions and confidence intervals, optional numeric summaries, and cross-facet counts or ratios. Empty cross-facet cells can mean that an intersection was not computed or fell below a privacy threshold. The files contain no raw transcripts, user identifiers, or organization identifiers.[5]
The dataset card warns that labels are interpretations rather than validated findings. Approximately 3% of conversations in the original topic-validation study were not clearly described by their assigned cluster. A conversation can cover several topics even though the pipeline assigns it to one cluster, and randomized clustering can produce different groupings on repeated runs. Safety-oriented description prompts can also emphasize the most concerning examples in a group.[4][5]
Limitations and privacy research
Errors can enter at every stage. A model can hallucinate a facet, miss sarcasm or context, or force an ambiguous conversation into an ill-fitting category. Embeddings and k-means can divide one topic across clusters or merge different behaviors. Generated labels can overstate a pattern, and a hierarchy can imply boundaries that are not present in the source data. The system observes conversations, not user intent or downstream outcomes, and aggregation makes it poorly suited to rare behavior.[2][4]
The model-specific sample creates another limit. Claude users, plan eligibility, opt-in rules, and the sampled time window all shape the data. Findings from these runs do not automatically generalize to other assistants, organizations, countries, or time periods. The authors of the original paper recommended treating outputs as leads for further study rather than as a sole basis for decisions.[2][5]
Independent privacy research has challenged the heuristic design. The 2026 Cliopatra preprint constructed synthetic medical conversations, mixed them with WildChat data, and assumed an adversary who knew some target attributes and could insert malicious chats through multiple accounts. Its authors reported extracting the synthetic target's medical condition in up to 65% of tested cases with nearly 100% precision in some configurations, and found that their tested model-based auditors failed to flag successful attack clusters.[7]
Cliopatra demonstrates a targeted poisoning and prompt-injection risk under controlled assumptions. It is not evidence that an actual user was identified from Anthropic's external-pilot dataset. Attack success also fell as the test corpus grew in some configurations. The result nonetheless shows why an empirical auditor score should not be read as a universal privacy guarantee.[7]
Urania is a separate research framework that applies private clustering, partition selection, and private histograms to provide end-to-end record-level differential privacy for Clio-like analysis. Its evaluations describe a tradeoff between stronger formal protection and reduced utility. Urania is not the documented production implementation of Anthropic Insights, and record-level protection does not automatically cover a user who contributes multiple conversations.[8]
References
- ^Anthropic. "Clio: A system for privacy-preserving insights into real-world AI use." December 12, 2024; rename update August 24, 2026. anthropic.com/...clio
- ^Alex Tamkin et al. "Clio: Privacy-Preserving Insights into Real-World AI Use." arXiv:2412.13678, December 2024. arxiv.org/...2412.13678
- ^Anthropic. "Enabling independent research on how people use Claude." August 26, 2026. anthropic.com/...enabling-independent-research
- ^Anthropic. "Enabling independent research appendix." August 2026. www-cdn.anthropic.com/...6287b9255657016f50e29.pdf
- ^Anthropic. "Enabling independent research." Hugging Face dataset, accessed August 27, 2026. huggingface.co/...enabling-independent-research
- ^Yijia Shao, Dora Zhao, Vishakh Padmakumar, Jennifer Wang, and Diyi Yang. "Human-AI Collaboration at Scale: Task Criticality, Agency, and Friction Across 250,000 Conversations." Preprint submitted August 21, 2026. alphaxiv.org/....human-ai-collaboration-at-scalev1
- ^Meenatchi Sundaram Muthu Selva Annamalai, Emiliano De Cristofaro, and Peter Kairouz. "Cliopatra: Extracting Private Information from LLM Insights." arXiv:2603.09781, revised May 27, 2026. arxiv.org/...2603.09781
- ^Daogao Liu et al. "Urania: Differentially Private Insights into AI Use." arXiv:2506.04681, revised September 23, 2025. arxiv.org/...2506.04681
- ^Anthropic. "The Anthropic Economic Index." February 10, 2025. anthropic.com/...the-anthropic-economic-index
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 2,152 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independently fact-checked against the cited sources on Aug. 27, 2026; claims were limited to what those sources support.
Cite this page: AI Wiki. "Anthropic Insights." aiwiki.ai, updated 27 Aug 2026, fact-checked 27 Aug 2026. CC BY 4.0. https://aiwiki.ai/wiki/anthropic_insights