BABILong
BABILong is a benchmark for testing how well a large language model can reason over facts scattered through very long text.
Explore Model Evaluation through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Model Evaluation.
Showing 1-12 of 12 articles
BABILong is a benchmark for testing how well a large language model can reason over facts scattered through very long text.
As of July 2026, the strongest general reasoning models are Anthropic's Claude Opus 4.8 and Claude Fable 5, OpenAI's GPT-5.5, and Google's Gemini 3.1 Pro and Gemini 3 Pro in its Deep Think mode
FACTS Grounding is a factuality benchmark from Google DeepMind and Google Research that measures whether a large language model answers a request using only the information in a provided source document
Every figure in this article is a US dollar list rate per 1,000,000 tokens, taken from the provider's own pricing page.
As of July 2026, no single model wins every LLM benchmark, but Anthropic's Claude Fable 5 tops the most: it currently leads MMLU-Pro (91.5%), SWE-bench Verified (95.0%), Humanity's Last Exam in both the…
As of July 2026, the largest context window ever announced belongs to Magic's LTM-2-mini at 100,000,000 tokens (100M), but it is a research prototype that has never been publicly released.
LLM evaluation is the practice of measuring what a large language model can do, how reliably it does it, and how it behaves under adversarial or high-stakes conditions.
LLM-as-a-Verifier is a probabilistic verification framework and open-source Python package for scoring and selecting large language model agent trajectories.
LLM-as-a-judge is the practice of using a strong large language model to evaluate the outputs of other models, or of itself, in place of a human annotator.
LongBench v2 is a benchmark for evaluating how well large language models understand and reason over long contexts.
MRCR (Multi-Round Co-reference Resolution) is a synthetic long-context evaluation that tests whether a large language model can locate and disambiguate among several near-identical "needles" buried inside a…
NoLiMa, short for "No Literal Matching," is a long-context benchmark for large language models that measures how well a model can find and use a single relevant fact buried in a long document when that fact…