LLM Evaluation
LLM evaluation is the practice of measuring what a large language model can do, how reliably it does it, and how it behaves under adversarial or high-stakes conditions.
Explore AI Research through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Research.
Showing 1-2 of 2 articles
LLM evaluation is the practice of measuring what a large language model can do, how reliably it does it, and how it behaves under adversarial or high-stakes conditions.
LLM-as-a-Verifier is a probabilistic verification framework and open-source Python package for scoring and selecting large language model agent trajectories.