Model Evaluation

Explore Model Evaluation through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: AI Agents

Articles that also belong to these categories. Counts cover all of Model Evaluation.

Showing 1-6 of 6 articles

Agent evaluation

Agent evaluation is the systematic measurement of how well AI agents (LLM-based systems that plan and act over multiple steps using tools) perform on real-world tasks, using benchmarks, metrics, and testing…

AI AgentsAI Benchmarks

WebArena

WebArena is a realistic, self-hosted web environment and benchmark designed for developing and evaluating autonomous AI agents that perform tasks on the web.

AI AgentsAI Benchmarks

WebVoyager

WebVoyager is an end-to-end web agent and its companion benchmark, introduced by Hongliang He and seven coauthors in a paper accepted to ACL 2024 .

AI AgentsAI Benchmarks