Model Evaluation

Explore Model Evaluation through related topics and the articles other pages reference most.

Explore articles

Reset filters
Browse subtopics: Multimodal AI

Articles that also belong to these categories. Counts cover all of Model Evaluation.

Showing 1-2 of 2 articles

GDP.pdf

GDP.pdf is a benchmark from Surge AI that tests whether multimodal language models can answer realistic professional questions about real PDF documents, such as benefits packets, leases, datasheets, clinical…

AI BenchmarksEnterprise AI