# Ragas

> Source: https://aiwiki.ai/wiki/ragas
> Updated: 2026-09-27
> Fact-checked: 2026-09-27
> Categories: Developer Tools, Information Retrieval, Model Evaluation
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Ragas." aiwiki.ai, 27 Sept 2026. https://aiwiki.ai/wiki/ragas
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

Ragas is an open-source Python framework for evaluating applications built with [large language models](https://aiwiki.ai/wiki/large_language_model). It provides evaluation metrics, test-data generation, and tools for comparing experiments. Its original research focused on [retrieval-augmented generation](https://aiwiki.ai/wiki/retrieval_augmented_generation) (RAG), while its documentation also covers conversational agents, tool use, and custom evaluation criteria. The project is distributed under the Apache License 2.0.[1][2]

Ragas is a collection of evaluators rather than a single universal quality score. A result needs the metric name, its configuration, the evaluation data, and the judging models to be interpretable. In particular, an answer can match its retrieved evidence while failing to answer the question, or answer the question directly while containing unsupported claims. Ragas exposes separate measurements for these properties.[3][4]

## Research origins

Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert presented *RAGAs: Automated Evaluation of Retrieval Augmented Generation* at the EACL 2024 system demonstrations track. The name expands to Retrieval Augmented Generation Assessment. The paper investigated reference-free measures of faithfulness, answer relevance, and context relevance: evaluation without requiring a human-written target answer for every question.[5]

The authors also introduced WikiEval, constructed using 50 Wikipedia pages and human comparisons of alternative answers or contexts. Those experiments tested agreement with annotators under a particular dataset, prompting procedure, and model setup. They were not a certification that later Ragas releases, other judges, or other domains achieve the same agreement. The paper's context-relevance procedure should not be assumed to be identical to every retrieval metric offered by the current library.[5]

## Evaluation inputs

Ragas separates the application being evaluated from the models used to judge its output. Its model factories allow evaluators and embedding models to be configured for different providers. Changing the judge is therefore an evaluation change, even if the application itself stays unchanged.[6]

The documented sample types include `SingleTurnSample` for one interaction and `MultiTurnSample` for a conversation. A sample can carry the question, answer, retrieved material, and reference information; it need not contain every possible field. The selected metric determines which inputs are necessary.[7]

| Field or input | Meaning in an evaluation |
|---|---|
| `user_input` | The question or instruction; in multi-turn evaluation, the conversation messages |
| `response` | The application's generated answer |
| `retrieved_contexts` | The passages supplied as retrieved evidence |
| `reference` | A target answer or expected outcome used by reference-based metrics |
| `reference_tool_calls` | Expected tool calls for relevant agent metrics |
| Rubric | The criteria against which a response is judged |

These distinctions prevent a common setup error: using the application's own answer as its supposed ground truth. A reference is a separate evaluation input. Reference-free scoring avoids that input for particular metrics; it does not remove the need for suitable questions, retrieved evidence, or validation of the evaluator.[7][11]

## Main RAG metrics

The following are distinct measurements, not interchangeable names for overall accuracy. The API names shown are the collections-based names documented in the Ragas migration guidance.[9]

| Metric | Main question | Principal inputs |
|---|---|---|
| `Faithfulness` | Are the answer's claims supported by the retrieved passages? | User input, response, retrieved contexts [3] |
| `AnswerRelevancy` | Does the response address the question? | User input and response; an LLM and embeddings [4] |
| `ContextPrecision` | Does retrieval place useful passages before unhelpful ones? | User input, retrieved contexts, reference answer [10] |
| `ContextPrecisionWithoutReference` (`ContextUtilization` is a compatibility wrapper) | Which retrieved passages support the generated response? | User input, retrieved contexts, response [9][10] |
| `ContextRecall` | Does the retrieved evidence cover the information in a reference answer? | User input, retrieved contexts, reference answer [11] |
| `FactualCorrectness` | How do the answer's claims compare with a reference? | Response and reference [8] |

### Faithfulness and correctness

Faithfulness decomposes the answer into claims and asks whether the retrieved context supports them. For a nonempty set of judged claims, its score is the number supported divided by the total number judged. It assesses consistency with the supplied material, not independent truth about the world.[3]

For a hypothetical illustration, suppose the only passage says that a fictional service accepts refunds for 14 days. An answer asserting both that deadline and free return shipping adds a claim that the passage does not establish. If the evaluator extracts exactly those two claims and supports only the first, the ratio is 1/2. This illustrates the arithmetic, not a promised output from a particular judge.[12]

The inspected collections implementation raises errors for missing required inputs and returns `NaN` when it generates no statements. A missing score is not evidence that an answer passed or failed. Reporting code needs to preserve the difference between a numeric judgment and an evaluation that produced no usable judgment.[12]

`FactualCorrectness` instead compares claims in the response with claims in a reference answer. It supports precision, recall, and F1 modes, with F1 as the documented default. Its `atomicity` and `coverage` settings affect claim decomposition. Even this metric depends on the reference's quality: agreement with an incorrect or incomplete reference is not an independent fact check.[8]

### Response relevance

The documented `AnswerRelevancy` procedure generates questions from the response and compares their [embeddings](https://aiwiki.ai/wiki/embeddings) with the original question using cosine similarity. It is intended to measure whether the answer addresses the request, not whether its statements are true. The documentation notes that cosine similarity is not mathematically restricted to the interval from zero to one, even though scores often fall there.[4]

The inspected implementation also checks generated judgments of whether the answer is noncommittal. This is a reason to test refusal and unanswerable-question cases explicitly: declining to answer can be the correct application behavior, even when a relevance metric does not reward it.[13]

### Retrieval precision and recall

LLM-based context precision first judges passage usefulness, then rewards useful passages appearing earlier in the retrieved list. The documentation distinguishes reference-answer-based scoring from context utilization, which uses the generated response. A high utilization result cannot establish that retrieval found information the answer omitted. Other variants compare retrieved passages with reference contexts or compare document IDs, so the evaluator variant must be reported.[10]

LLM-based context recall decomposes the reference answer and checks which claims the retrieved passages support. Its denominator comes from that reference, not from every relevant document that might exist in the corpus. Ragas also documents non-LLM and ID-based recall variants with different reference requirements. An incomplete reference can leave a retrieval omission unmeasured.[11]

## Experiments and interpretation

The experiment workflow runs an application or component on a test dataset and records evaluation outputs. Ragas recommends isolating changes so that comparisons remain interpretable. For example, replacing a retriever while keeping the questions, generator, and scoring setup fixed tests a narrower hypothesis than changing all of them together.[14]

A useful experiment record identifies the application configuration alongside its dataset and evaluator configuration. It also retains individual results, so an aggregate improvement can be traced to the examples that changed. If a run changes both the judge and the application, a score difference alone cannot identify which change caused it. These are experimental-design constraints, not problems that averaging more metrics automatically solves.[14]

The framework supports criteria beyond the built-in RAG measures. Its general-purpose documentation describes binary aspect judgments, custom categorical or numerical scoring, and rubric-based evaluation. The criterion still needs an operational definition. A label such as "good answer" does not specify whether completeness, source support, brevity, or policy compliance should decide the result.[15]

## Agent evaluation

For [AI agents](https://aiwiki.ai/wiki/ai_agents), Ragas provides metrics for tool-call accuracy, tool-call F1, goal achievement, and topic adherence. Tool-call accuracy examines recorded calls against expected calls, including arguments and, depending on configuration, their ordering. Goal-accuracy metrics judge outcomes with a supplied reference or infer the goal and outcome from a conversation.[16]

These assess different evidence. A matching tool-call sequence does not by itself establish that an external operation succeeded. Conversely, multiple valid tool sequences may reach the same outcome. An evaluation needs to distinguish a required procedure from an acceptable result. Conversation-based goal judgments concern the supplied record; they do not independently inspect every external system an agent used.[16]

## Synthetic test sets

Ragas can generate [synthetic data](https://aiwiki.ai/wiki/synthetic_data) for evaluation from documents. Its RAG test-generation design builds a [knowledge graph](https://aiwiki.ai/wiki/knowledge_graph) using document chunks, extracted properties, and relationships. Query generation can target single-hop or multi-hop questions and different levels of specificity. The graph helps select related source material for generating those cases.[17]

Generated tests are useful inputs to an evaluation process, but generation and validation remain separate steps. A document-derived question set reflects the documents and synthesis procedure that produced it. It does not establish how often real users ask those questions. Review should check that questions are answerable from their designated evidence and that generated references are correct before treating them as evaluation targets. This follows from the dependence of reference-based metrics on the supplied reference.[17][8]

## Reliability and limitations

Many Ragas metrics use the [LLM-as-a-judge](https://aiwiki.ai/wiki/llm_as_a_judge) approach. Research by Zheng and colleagues documented position, verbosity, and self-enhancement biases, as well as reasoning limitations, in language-model judges. The relevance of each bias depends on the judging task: a pairwise preference test and a claim-support classifier are not the same evaluation. General judge research does not establish an error rate for every Ragas metric.[18]

GroUSE, a COLING 2025 study, examined grounded-answer evaluators using targeted tests for generator failure modes. It found that agreement with another model's judgments was not a sufficient substitute for checking whether evaluators recognized particular failures. For a Ragas deployment, this supports checking the evaluator against manually reviewed cases rather than treating a plausible-looking numerical score as validation.[19]

Repeated model calls can also produce different judgments. Cached calls require special care when investigating that variation: Ragas supports exact-match caching of model and embedding outputs, so repeated cache hits reuse earlier results rather than independently resampling the evaluator. Caching can reduce repeated work; it does not measure the uncached judge's consistency.[20]

## Versions and operating considerations

Ragas APIs have changed across releases. Its v0.3-to-v0.4 migration guide recommends collections-based metrics, `llm_factory`, and experiment-oriented workflows, while older examples use wrappers, `evaluate()`, and sample-based scoring methods. Some legacy interfaces remain documented with deprecation notices. Code should follow the installed version's documentation rather than combine imports and method calls from different generations.[9]

The project can be installed with `pip install ragas`. Reproducible comparisons should record the installed release rather than treat that unpinned command as a version specification. The repository also documents usage analytics and an opt-out setting, `RAGAS_DO_NOT_TRACK=true`. That setting concerns Ragas analytics; it is not a substitute for understanding data handling by whichever model provider is configured.[1][6]

Evaluation has its own model-call cost. Ragas documents token-usage accounting for its `evaluate()` workflow through a provider-specific parser; this is not a claim that every API path automatically reports a complete bill. Test-data generation also consumes model resources. Budgets need to account for judging and generation separately from the application requests under test.[21]

Timeout and retry settings likewise depend on the interface. For collections-based metrics, the documentation directs configuration to the model client; the older evaluation path uses `RunConfig`. A timeout, parsing failure, or unavailable provider is an evaluation failure to investigate, not a low-quality answer score.[22]

## References

1. Ragas maintainers. [Ragas repository and README](https://github.com/vibrantlabsai/ragas). Accessed September 27, 2026.
2. Ragas documentation. [Introduction](https://docs.ragas.io/en/stable/). Accessed September 27, 2026.
3. Ragas documentation. [Faithfulness](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/faithfulness/). Accessed September 27, 2026.
4. Ragas documentation. [Response Relevancy](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/answer_relevance/). Accessed September 27, 2026.
5. Es, Shahul; James, Jithin; Espinosa Anke, Luis; and Schockaert, Steven. [RAGAs: Automated Evaluation of Retrieval Augmented Generation](https://aclanthology.org/2024.eacl-demo.16/). EACL 2024 System Demonstrations, pp. 150-158. DOI: 10.18653/v1/2024.eacl-demo.16. [Author manuscript](https://arxiv.org/html/2309.15217v2).
6. Ragas documentation. [Customise models](https://docs.ragas.io/en/stable/howtos/customizations/customize_models/). Accessed September 27, 2026.
7. Ragas documentation. [Evaluation Sample](https://docs.ragas.io/en/stable/concepts/components/eval_sample/). Accessed September 27, 2026.
8. Ragas documentation. [Factual Correctness](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/factual_correctness/). Accessed September 27, 2026.
9. Ragas documentation. [From v0.3 to v0.4](https://docs.ragas.io/en/stable/howtos/migrations/migrate_from_v03_to_v04/). Accessed September 27, 2026.
10. Ragas documentation. [Context Precision](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/context_precision/). Accessed September 27, 2026.
11. Ragas documentation. [Context Recall](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/context_recall/). Accessed September 27, 2026.
12. Ragas maintainers. [Collections Faithfulness implementation](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/collections/faithfulness/metric.py). Commit 298b68274234c060deacab3cf5fb52aa3a20e885, inspected September 27, 2026.
13. Ragas maintainers. [Collections AnswerRelevancy implementation](https://github.com/vibrantlabsai/ragas/blob/298b68274234c060deacab3cf5fb52aa3a20e885/src/ragas/metrics/collections/answer_relevancy/metric.py). Commit 298b68274234c060deacab3cf5fb52aa3a20e885, inspected September 27, 2026.
14. Ragas documentation. [Experimentation](https://docs.ragas.io/en/stable/concepts/experimentation/). Accessed September 27, 2026.
15. Ragas documentation. [General Purpose Metrics](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/general_purpose/). Accessed September 27, 2026.
16. Ragas documentation. [Agentic or Tool use](https://docs.ragas.io/en/stable/concepts/metrics/available_metrics/agents/). Accessed September 27, 2026.
17. Ragas documentation. [Testset Generation for RAG](https://docs.ragas.io/en/stable/concepts/test_data_generation/rag/). Accessed September 27, 2026.
18. Zheng, Lianmin, et al. [Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena](https://arxiv.org/abs/2306.05685). NeurIPS 2023 Datasets and Benchmarks.
19. Muller, Sacha; Loison, Antonio; Omrani, Bilel; and Viaud, Gautier. [GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering](https://aclanthology.org/2025.coling-main.304/). COLING 2025, pp. 4510-4534.
20. Ragas documentation. [Caching in Ragas](https://docs.ragas.io/en/stable/howtos/customizations/_caching/). Accessed September 27, 2026.
21. Ragas documentation. [How to estimate Cost and Usage of evaluations and testset generation](https://docs.ragas.io/en/stable/howtos/applications/_cost/). Accessed September 27, 2026.
22. Ragas documentation. [Customize Timeouts and Rate Limits](https://docs.ragas.io/en/stable/howtos/customizations/run_config/). Accessed September 27, 2026.

