Arize Phoenix
Arize Phoenix is an open-source AI observability and evaluation platform developed by Arize AI for tracing, evaluating, and debugging large language model (LLM) and agent applications.
Explore Model Evaluation through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of Model Evaluation.
Showing 1-14 of 14 articles
Arize Phoenix is an open-source AI observability and evaluation platform developed by Arize AI for tracing, evaluating, and debugging large language model (LLM) and agent applications.
DesignArena (also written Design Arena) is a crowdsourced benchmark and consumer creation platform for AI-generated design, built by the San Francisco startup Intelligence.
Epoch AI is a nonprofit research organization, founded in 2022 and directed by Jaime Sevilla, that studies the trajectory of artificial intelligence through quantitative analysis of compute, data, algorithms…
The Future of Life Institute AI Safety Index (often shortened to the AI Safety Index or FLI AI Safety Index) is a periodic "report card" published by the Future of Life Institute (FLI) that grades the leading…
Helicone is an open-source LLM observability platform and AI gateway founded in 2023 by Justin Torre, Cole Gottdank, Barak Oshri, and Scott Nguyen.
LMArena is a crowdsourced artificial intelligence evaluation platform and company that ranks large language models by having anonymous users vote on which of two blind, side-by-side model responses they prefer
LangSmith is a commercial observability, evaluation, and deployment platform for large language model (LLM) applications and AI agents, developed and operated by LangChain Inc. It provides developers and…
Langfuse is an open-source LLM engineering platform that provides observability, tracing, prompt management, evaluation, and dataset tooling for applications built on large language models.
Patronus AI is an automated LLM evaluation, observability, and guardrails platform founded in 2023 and headquartered in San Francisco.
The SEAL Leaderboards are a set of expert-curated, contamination-resistant evaluation leaderboards for frontier large language models, produced by the Safety, Evaluations and Alignment Lab (SEAL) at Scale AI.
The AI Index Report is an annual, independent, data-driven publication that tracks, distills, and visualizes trends in artificial intelligence across research and development, technical performance, the…
The State of AI Report is a free, annual review of progress in artificial intelligence, published every year since 2018.
The Leaderboard Illusion is a 2025 research paper, led by Cohere and Cohere Labs with academic collaborators, that argues the most influential public ranking of large language models, Chatbot Arena
Vals AI is an independent, third-party AI model evaluation company based in San Francisco that builds and publishes domain-specific benchmarks measuring how well large language models perform on real…