SkillEvaluator
SkillEvaluator is an open-source command-line framework developed by NVIDIA for evaluating artifacts used by AI agents.
Explore AI Agents through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Agents.
Showing 241-272 of 272 articles
SkillEvaluator is an open-source command-line framework developed by NVIDIA for evaluating artifacts used by AI agents.
SkillsBench is a benchmark for measuring whether Agent Skills, the structured packages of procedural knowledge that augment AI agents at inference time, actually improve how well those agents do real work.
Sleep-time compute is a technique for large language model inference in which a model uses idle periods, before any user query has arrived, to "think" about a known context offline and pre-compute a richer…
Spec-driven development (SDD) is a software engineering methodology in which structured specification documents serve as the primary source of truth for a project, with code treated as a generated or verified…
Stilla is a Swedish workplace AI agent developed by Stilla Development AB. It is designed to retain shared organizational context across connected services, receive requests in team channels, and propose or…
Summation is an enterprise software company in Bellevue, Washington, that sells what it markets as an "AI analyst": a hosted system that connects to a company's data warehouse, ERP and spreadsheets and…
Swarm intelligence (SI) is a branch of artificial intelligence concerned with the collective behavior of decentralized, self-organized systems
Sweep (stylized with a broom emoji and reachable at sweep.dev) is an AI developer tool built by the San Francisco startup of the same name, founded in 2023 by Kevin Lu and William Zeng.
τ²-bench (also written Tau2-bench or τ^2-bench) is a benchmark for evaluating conversational AI agents in dual-control environments, where both the agent and a simulated user can call tools to read from and…
Terminal-Bench is an open benchmark for evaluating AI agents on complex, real-world tasks performed through command-line terminal interfaces.
Tim Rocktäschel is a German computer scientist known for his work on reinforcement learning, open-ended learning, and language-based AI agents.
Toast 1 is a hosted text model from Mixedbread that is specialized for multistep search and evidence synthesis.
Tool use in artificial intelligence is the ability of a model-based system to request, coordinate, and use capabilities outside the model's ordinary token-generation process.
Toolformer is a research language model from Meta AI that learns, in a self-supervised way, to call external software tools through simple text-based API calls.
Toolhouse is a cloud infrastructure platform that helps developers equip large language models with tools, actions, and knowledge.
UI-TARS is a native graphical user interface (GUI) agent model developed by ByteDance through its Seed research team.
Vapi is a voice AI orchestration platform that lets software developers build, deploy, and scale AI phone agents through a programmable API.
Vellum is an enterprise AI product development platform that helps engineering teams build, test, and deploy applications powered by large language models (LLMs).
Vending-Bench is a benchmark that measures the long-term coherence of large language model-based AI agents by tasking them with running a simulated vending-machine business over an extended horizon.
VisualWebArena (often abbreviated VWA) is a benchmark of 910 realistic, visually grounded web tasks for evaluating multimodal autonomous AI agents, released in January 2024 by researchers at Carnegie Mellon…
Voyager is an open-ended embodied agent that uses a large language model to play Minecraft by writing, executing, and storing JavaScript programs against the Mineflayer bot API.
Waddle Labs is a two-person American startup, founded in 2025 and based in Boston, that sells an API through which large language model agents write, run and revise the control code that drives a customer's…
WebArena is a realistic, self-hosted web environment and benchmark designed for developing and evaluating autonomous AI agents that perform tasks on the web.
WebGPT is a research model released by OpenAI in December 2021 that fine-tuned the 175-billion-parameter GPT-3 to answer long-form questions by operating a text-based web browser and citing the sources it…
WebVoyager is an end-to-end web agent and its companion benchmark, introduced by Hongliang He and seven coauthors in a paper accepted to ACL 2024 .
Windows Agent Arena (WAA) is a reproducible benchmark and evaluation environment created by Microsoft for testing multimodal computer-use agents on real Windows tasks at scale.
WindowsWorld is a process-aware benchmark for evaluating autonomous computer-use agents on professional workflows that cross several Windows applications.
Wolfram GPT is the integration of Wolfram|Alpha and the Wolfram Language with OpenAI's ChatGPT, giving the language model on-demand access to authoritative computational knowledge, real-time data feeds…
Zapier is a no-code web-automation platform, founded in 2011, that connects more than 8,000 software applications and lets people build automated workflows between them without writing code
n8n (pronounced "n-eight-n") is a source-available workflow automation platform built for technical teams, developed by the Berlin-based company n8n GmbH.
opencode is an open-source AI coding agent built for the terminal. Developed by Anomaly Innovations, formerly known as the SST (Serverless Stack) team, it provides a terminal user interface (TUI) through which…
smolagents is an open source Python library for building agents powered by large language models (LLMs), released by hugging face on 30 December 2024.