Thursday, July 23, 2026
- Beev2Bee is a San Francisco startup that made a low-cost, always-listening AI wearable: a small wristband (also wearable as a clip-on) that passively records the wearer's conversations and surroundings and uses AI…
- Base44v2Base44 is an artificial intelligence vibe coding platform that lets users build fully functional web applications by describing them in natural language, without writing code.
- Brexv2Brex is an American fintech company that builds a software platform for corporate spending, combining corporate charge cards, expense management, business accounts, bill pay, and travel into a single product…
- Deep Cogitov2Deep Cogito is a San Francisco artificial intelligence research lab that develops open-weight large language models under the Cogito name.
- Arc Institutev3The Arc Institute is an independent, nonprofit biomedical research organization, headquartered in Palo Alto, California, that funds long-horizon basic science to understand and cure complex human diseases.
- Rampv4Ramp is an American fintech company that runs an AI-powered finance and spend-management platform built around a corporate charge card, automating expense categorization, receipt matching, bill payment…
- Mercuryv2Mercury (legally Mercury Technologies, Inc.) is an American financial-technology company that provides online business banking and financial-operations software aimed primarily at startups and small and…
- Cluelyv2Cluely is an American artificial intelligence startup, founded in 2025, that makes a desktop AI assistant which watches a user's screen and listens to their audio in real time and supplies answers and talking…
- Moonvalleyv2Moonvalley is a generative-AI research company that builds video-generation models for professional filmmakers, studios, and brands.
- Plaudv2Plaud (styled PLAUD, and also known as Plaud.AI or Plaud.ai) is a consumer AI hardware company that makes compact voice-recording devices paired with cloud-based speech recognition and large-language-model…
- Roxv2Rox is a San Francisco based enterprise software company that builds AI agents for sales and revenue teams.
- Invisible Technologiesv2Invisible Technologies is an American artificial intelligence company that combines a global network of vetted human experts with its own orchestration software to provide AI training data, reinforcement…
- Axelera AIv2Axelera AI is a Dutch semiconductor company that designs AI chips for inference, the stage at which a trained model is run to produce predictions.
- MatXv2MatX is an American semiconductor startup that designs AI chips purpose-built for large language models (LLMs).
- Cradlev2Cradle (legally Cradle Bio B.V., styled cradle.bio) is a Dutch-Swiss software company that builds generative AI tools for protein engineering.
- Basecamp Researchv2Basecamp Research is a London-based artificial intelligence and biotechnology company that has assembled what it describes as the world's largest and most diverse database of biological sequences, sampled…
- Formation Biov2Formation Bio (formerly TrialSpark) is an American "AI-native" pharmaceutical company that in-licenses and acquires clinical-stage drug candidates and develops them more efficiently by applying artificial…
- Envedav2Enveda (Enveda Biosciences, legally Enveda Therapeutics) is an American AI drug discovery company that uses machine learning and high-resolution mass spectrometry to decode the chemistry of nature and turn…
- Chai Discoveryv3Chai Discovery is an American AI biotechnology company that builds foundation models for predicting and designing the three-dimensional structures of biomolecules to accelerate drug discovery.
- Profluentv3Profluent (legally Profluent Bio Inc.) is an American artificial intelligence company that uses large language models and generative AI to design novel proteins and genome editors.
- Insitrov3Insitro (styled "insitro") is a machine learning-driven drug discovery and development company founded in 2018 by Daphne Koller, the Stanford machine learning professor and Coursera co-founder, and…
- Tennrv2Tennr is an American healthcare AI company that builds software to automate the patient referral and intake workflows that move people from one healthcare provider to another.
- Qventusv3Qventus is a privately held healthcare AI company that develops software to automate hospital and health system operations.
- Innovaccerv2Innovaccer Inc. is a healthcare technology company that builds a data platform and an expanding suite of artificial intelligence tools for hospitals, health systems, payers, government agencies, and life…
- Sukiv2Suki is an American healthcare AI company that builds an AI-powered voice assistant for clinicians.
- Commurev2Commure is a United States healthcare technology company that builds an artificial intelligence operating system and a suite of AI tools for hospitals and health systems.
- Robin AIv3Robin AI is a British legal AI company, founded in London in 2019, that built software to review, draft, and negotiate contracts by pairing large language models with human lawyers in a "lawyer in the loop"…
- Spellbookv2Spellbook is a Canadian legal AI company that builds an artificial intelligence assistant for transactional and contract lawyers.
- Crosbyv2Crosby is an AI-native law firm based in New York City that pairs licensed attorneys with proprietary artificial intelligence to review and negotiate commercial contracts
- Eudiav3Eudia is an American legal AI company that builds an "augmented intelligence" software platform for the in-house legal departments of large enterprises.
- Legorav4Legora is a Swedish legal artificial-intelligence company that builds a collaborative legal AI platform for law firms and corporate in-house legal teams.
- Remote Labor Indexv2The Remote Labor Index (RLI) is an AI benchmark that measures how well AI agents can complete real, paid remote knowledge work end to end.
- Studio Ghibli ChatGPT momentv2The "Studio Ghibli ChatGPT moment" refers to a viral internet trend that erupted in late March 2025, after OpenAI released a new native image generation capability inside ChatGPT, powered by the GPT-4o model.
- MIT "GenAI Divide" report (2025)v2The "GenAI Divide" report, formally titled The GenAI Divide: State of AI in Business 2025, is a research report published in July 2025 by MIT NANDA, an initiative based at the MIT Media Lab.
- Grok "MechaHitler" incidentv2The Grok "MechaHitler" incident was a content-moderation and AI-safety failure that occurred on July 8 to 9, 2025, when Grok, the chatbot built by xAI and integrated into the social platform X (formerly…
- Reflection 70B controversyv2The Reflection 70B controversy was an open-source AI credibility episode that began on September 5, 2024, when Matt Shumer, co-founder and chief executive of OthersideAI (the company behind the AI writing…
- Perplexity AI copyright lawsuitsv2The Perplexity AI copyright lawsuits are a cluster of lawsuits, cease-and-desist demands, and public disputes brought against Perplexity AI by news publishers, reference publishers, and online platforms over…
- Disney & Universal v. Midjourneyv2Disney and Universal v. Midjourney is a copyright infringement lawsuit filed on June 11, 2025, in which The Walt Disney Company and Comcast's NBCUniversal jointly sued the generative AI image service…
- Florida v. OpenAIv2Florida v. OpenAI is a civil enforcement lawsuit filed on June 1, 2026, by the State of Florida, through Attorney General James Uthmeier, against OpenAI and its chief executive Sam Altman.
- Musk v. Altman (Musk v. OpenAI)v2Musk v. Altman, also referred to as Musk v. OpenAI, is a lawsuit filed by the entrepreneur Elon Musk accusing OpenAI and its chief executive Sam Altman of abandoning the company's founding mission as an open…
- Raine v. OpenAIv4Raine v. OpenAI, Inc. is a wrongful-death lawsuit filed on August 26, 2025, in the Superior Court of California for the County of San Francisco (case number CGC-25-628528) by Matthew and Maria Raine, the…
- AI at the 2025 ICPC World Finalsv3In September 2025, at the 49th International Collegiate Programming Contest (ICPC) World Finals in Baku, Azerbaijan, two leading artificial-intelligence laboratories reported that their general-purpose…
- AI gold medals at the 2025 IMOv3In July 2025, two leading artificial intelligence laboratories, OpenAI and Google DeepMind
- The Leaderboard Illusionv3The Leaderboard Illusion is a 2025 research paper, led by Cohere and Cohere Labs with academic collaborators, that argues the most influential public ranking of large language models, Chatbot Arena
- Scale SEAL Leaderboardsv2The SEAL Leaderboards are a set of expert-curated, contamination-resistant evaluation leaderboards for frontier large language models, produced by the Safety, Evaluations and Alignment Lab (SEAL) at Scale AI.
- Future of Life Institute AI Safety Indexv2The Future of Life Institute AI Safety Index (often shortened to the AI Safety Index or FLI AI Safety Index) is a periodic "report card" published by the Future of Life Institute (FLI) that grades the leading…
- State of AI Reportv2The State of AI Report is a free, annual review of progress in artificial intelligence, published every year since 2018.
- Stanford HAI AI Index Reportv2The AI Index Report is an annual, independent, data-driven publication that tracks, distills, and visualizes trends in artificial intelligence across research and development, technical performance, the…
- Data Provenance Initiativev2The Data Provenance Initiative (DPI) is a volunteer-led, multi-institution research collective that audits and documents the licenses, sources, creators, and consent status of the datasets most widely used to…
- OpenThoughtsv2OpenThoughts is an open-source initiative and a series of datasets of verified reasoning traces created to train open reasoning models.
- WildChatv3WildChat is a large public corpus of real conversations between human users and ChatGPT, released by researchers at the Allen Institute for AI (AI2) and Cornell University.
- The Stack v2v3The Stack v2 is a large open dataset of source code released by BigCode in February 2024 as the training dataset behind the StarCoder2 family of code models.
- FineWeb-2v2FineWeb-2 (also written FineWeb2) is a massively multilingual web pretraining dataset released by Hugging Face in December 2024.
- SIMPLERv3SIMPLER (Simulated Manipulation Policy Evaluation for Real Robot Setups) is a collection of simulated robot manipulation environments, released in 2024
- LIBEROv2LIBERO ("LIfelong learning BEnchmark on RObot manipulation tasks") is an AI benchmark for studying knowledge transfer in lifelong robot learning.
- MedHELMv2MedHELM (Holistic Evaluation of Large Language Models for Medical Tasks) is a benchmark and evaluation framework that measures how well large language models perform on realistic clinical work.
- FinanceBenchv2FinanceBench is an AI benchmark for open-book financial question answering, designed to test whether large language models can answer the kinds of questions a financial analyst asks about a publicly traded…
- MASKv3MASK (Model Alignment between Statements and Knowledge) is an AI safety benchmark that measures the honesty of large language models (LLMs) by testing whether a model will knowingly assert something it…
- OpenAI MRCR (Multi-Round Co-reference Resolution)v2OpenAI MRCR (Multi-Round Co-reference Resolution) is a long-context evaluation dataset published by OpenAI that measures a language model's ability to distinguish between multiple near-identical "needles"…
- LongFact / SAFEv2LongFact and SAFE are a paired benchmark and evaluation method for measuring the long-form factuality of large language models, introduced by researchers at Google DeepMind and Stanford University in the 2024…
- FActScorev2FActScore (Factual precision in Atomicity Score) is an evaluation method and metric, introduced in 2023, for measuring the factual precision of long-form text generated by large language models.
- WISEv3WISE (World Knowledge-Informed Semantic Evaluation) is an AI benchmark that tests whether a text-to-image model actually possesses and correctly applies real-world knowledge when it draws a scene
- GenAI-Benchv3GenAI-Bench is an AI benchmark for evaluating compositional text-to-image and text-to-video generation, introduced in 2024 by researchers from Carnegie Mellon University and Meta AI .
- T2I-CompBenchv3T2I-CompBench is an AI benchmark for evaluating compositional text-to-image generation, introduced in the paper "T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation"…
- BELEBELEv2Belebele is a multiple-choice machine reading comprehension (MRC) AI benchmark that is fully parallel across 122 language variants, meaning the same questions, passages, and answer choices are translated into…
- MM-Vetv2MM-Vet is an AI benchmark for evaluating large multimodal models (LMMs, also called multimodal large language models or MLLMs) on tasks that require combining several core vision-language skills at once.
- BLINKv2BLINK is an AI benchmark that evaluates the core visual perception abilities of multimodal large language models (MLLMs).
- Video-MMEv4Video-MME (Video Multi-Modal Evaluation) is a benchmark for testing how well multimodal large language models (MLLMs) understand video, built from 900 manually selected videos totaling 254 hours and 2,700…
- RepoBenchv2RepoBench is an AI benchmark for repository-level code auto-completion, introduced in the 2023 paper "RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems" by Tianyang Liu, Canwen Xu, and…
- GSO (Global Software Optimization Bench)v2GSO (Global Software Optimization), also styled as a benchmark of "Challenging Software Optimization Tasks for Evaluating SWE-Agents," is an AI benchmark that measures whether AI agents and language models can…
- SWE-bench Multilingualv3SWE-bench Multilingual is an AI benchmark of 300 real-world software bug-fixing tasks drawn from 42 open-source repositories across nine programming languages
- SWE-rebenchv2SWE-rebench is a continuously refreshed, contamination-resistant AI benchmark and public leaderboard for evaluating AI agents on real-world software engineering tasks.
- CRMArena / CRMArena-Prov2CRMArena is an AI benchmark for evaluating large language model agents on professional customer relationship management (CRM) tasks inside a realistic, schema-faithful Salesforce environment.
- MultiChallengev2MultiChallenge is an AI benchmark for evaluating large language models on realistic multi-turn conversations.
- TheAgentCompanyv2TheAgentCompany is an AI benchmark that evaluates AI agents on long-horizon, economically valuable knowledge work inside a self-hosted simulation of a small software company.
- LAB-Benchv2LAB-Bench (the Language Agent Biology Benchmark) is an AI benchmark of more than 2,400 multiple-choice questions built to measure how well language models and AI agents can perform practical biology research…
- ChemBenchv2ChemBench is an automated AI benchmark that measures the chemical knowledge, reasoning, and safety judgment of large language models and compares their performance against expert human chemists.
- ScienceAgentBenchv2ScienceAgentBench is an AI benchmark for evaluating whether language model agents can perform real, data-driven scientific analysis by writing and executing code.
- PutnamBenchv2PutnamBench is a benchmark for evaluating automated theorem proving systems, especially neural and large-language-model-based provers, on competition mathematics.
- Omni-MATHv2Omni-MATH is an AI benchmark of Olympiad-level competition mathematics, introduced in October 2024 to measure the mathematical reasoning ability of large language models on problems far harder than those in…
- OlympiadBenchv2OlympiadBench is an AI benchmark of 8,476 Olympiad-level mathematics and physics problems, designed to test the advanced scientific reasoning of large language models and large multimodal models.
- MMLU-Reduxv2MMLU-Redux is a manually re-annotated, error-corrected subset of the Massive Multitask Language Understanding (MMLU) benchmark
- AGIEvalv2AGIEval is an AI benchmark for evaluating foundation models on tasks that were originally designed for, and taken by, humans.
- EnigmaEvalv3EnigmaEval is an AI benchmark of long, complex multimodal puzzles drawn from real-world puzzle hunts, designed to measure the unstructured, creative, multi-step reasoning abilities of frontier AI models.
- ARC-AGI-2v4ARC-AGI-2 (Abstraction and Reasoning Corpus for Artificial General Intelligence 2) is an abstract reasoning benchmark for artificial intelligence, released on March 24, 2025 by the ARC Prize Foundation
- MoE load-balancing lossv3The MoE load balancing loss is an auxiliary training objective used in sparse mixture of experts (MoE) neural networks to keep work spread evenly across the expert sub-networks.
- DualPipev3DualPipe is a bidirectional pipeline parallelism scheduling algorithm created by DeepSeek-AI to almost fully overlap computation with communication and shrink the idle time, called the pipeline "bubble," that…
- NLWebv2NLWeb (short for Natural Language Web) is an open project from Microsoft that makes it easy to add a natural-language conversational interface to a website, drawing on the site's own existing structured data…
- RAPTORv2RAPTOR (Recursive Abstractive Processing for Tree-Organized Retrieval) is a retrieval method for retrieval-augmented generation introduced in a 2024 paper by researchers at Stanford University.
- Corrective RAG (CRAG)v2Corrective Retrieval Augmented Generation (CRAG) is a method for improving the robustness of retrieval-augmented generation (RAG) when the underlying retrieval step returns irrelevant, incomplete, or factually…
- Self-RAGv2Self-RAG (Self-Reflective Retrieval-Augmented Generation) is a framework that trains a single large language model to adaptively decide when to retrieve external passages, to generate text grounded in those…
- MemGPTv4MemGPT (short for Memory-GPT) is a system and agent design pattern that gives large language model agents long-term memory by managing data between the model's bounded context window and external storage
- Sleep-time computev2Sleep-time compute is a technique for large language model inference in which a model uses idle periods, before any user query has arrived, to "think" about a known context offline and pre-compute a richer…
- Coconut (Chain of Continuous Thought)v3Coconut (Chain of Continuous Thought) is a reasoning paradigm for large language models introduced by researchers at FAIR at Meta, Meta's Fundamental AI Research lab, and the University of California
- SDEditv2SDEdit (Stochastic Differential Editing) is a method for guided image synthesis and editing that turns a rough user guide, such as a stroke painting, a coarse collage, or a real photograph with edits pasted…
- Prompt-to-Promptv2Prompt-to-Prompt is a training-free image editing technique for text-conditioned diffusion models that edits a generated image by manipulating the model's cross-attention maps when the text prompt is changed .
- Textual Inversionv3Textual Inversion is a technique for personalizing text-to-image diffusion models that teaches a frozen model a new visual concept from only three to five example images by learning a single new "pseudo-word"…
- DreamBoothv3DreamBooth is a subject-driven fine-tuning method for text-to-image diffusion models that personalizes a pretrained model to a specific subject, for example a particular dog, toy, or person, from just 3 to 5…
- IP-Adapterv2IP-Adapter (short for Image Prompt Adapter) is a lightweight neural network module that adds image-prompt conditioning to a pretrained text-to-image diffusion model, allowing a reference image to guide…
- Diffusion Forcingv2Diffusion Forcing is a training paradigm for sequence generative modeling introduced in 2024 that assigns each token in a sequence its own independent, randomly sampled noise level during training .