This article surveys _Evaluation_ models to **automatically detect hallucinations in Retrieval-Augmented Generation** (RAG), and presents a comprehensive benchmark of their performance across six RAG applications. While _hallucination_ sometimes refers to specific types of LLM errors, we use this term synonymously with _incorrect response_ (i.e. what matters to users of your RAG system). Code to reproduce our benchmark is available [here](https://github.com/cleanlab/cleanlab-tools/tree/main/benchmarking_hallucination_model).

_Retrieval-Augmented Generation_ enables AI to rely on company-specific knowledge when answering user requests. While RAG reduces LLM hallucinations, they still remain a critical concern that limits trust in RAG-based AI. Unlike traditional Search systems, RAG systems occasionally generate misleading/incorrect answers. This _unreliability_ poses risks for companies that deploy RAG externally, and limits how much internal RAG applications get used.

Real-time _Evaluation Models_ offer a solution, by providing a confidence score for every RAG response. Recall that for every user query: a RAG system retrieves relevant context from its knowledge base and then feeds this context along with the query into a LLM that generates a response for the user. Evaluation models take in the response, query, and context (or equivalently, the actual LLM prompt used to generate the response), and then output a score between 0 and 1 indicating how confident we can be that the response is correct. The challenge for Evaluation models is to provide _reference-free_ evaluation (i.e. with no ground-truth answers/labels available) that runs in real-time.

## Popular Real-Time Evaluation Models

### LLM as a Judge

**LLM-as-a-judge** (also called Self-Evaluation) is a straightforward approach, in which the LLM is directly asked to evaluate the correctness/confidence of the response. Often, the same LLM model is used as the model that generated the response. One might run LLM-as-a-judge using a Likert-scale scoring prompt.

The main issue with LLM-as-a-judge is that hallucination stems from the unreliability of LLMs, so directly relying on the LLM once again might not close the reliability gap as much as we’d like. That said, simple approaches can be surprisingly effective and LLM capabilities are ever-increasing.

### HHEM

Vectara’s [Hughes Hallucination Evaluation Model (HHEM)](https://www.vectara.com/blog/hhem-2-1-a-better-hallucination-detection-model) focuses on factual consistency between the AI response and retrieved context. This model’s scores have a probabilistic interpretation, where a score of 0.8 indicates an 80% probability that the response is factually consistent with the context.

Example use of HHEM

In our study, including the question within the premise improved HHEM results over just using the context alone.

### Prometheus

Prometheus (specifically the recent [Prometheus 2](https://arxiv.org/abs/2405.01535) upgrade) is a fine-tuned LLM model that was trained on direct assessment annotations of responses from various LLMs over various context/response data. Here we focus on the [Prometheus2 8x7B](https://huggingface.co/prometheus-eval/prometheus-8x7b-v2.0) (mixture of experts) model, as the authors reported that it performs better in direct assessment tasks. One can think of Prometheus as an LLM-as-a-judge that has been fine-tuned to align with human ratings of LLM responses.

### Patronus Lynx

Like the Prometheus model, [Lynx](https://www.patronus.ai/blog/lynx-state-of-the-art-open-source-hallucination-detection-model) is a LLM fine-tuned by Patronus AI to generate a PASS/FAIL score for LLM responses, trained on datasets with annotated responses. Lynx utilizes chain-of-thought to enable better reasoning by the LLM when evaluating a response.

### Trustworthy Language Model (TLM)

Unlike HHEM, Prometheus, and Lynx: Cleanlab’s [Trustworthy Language Model](/content/blog/trustworthy-language-model/index.html) (TLM) does not involve a custom-trained model. The TLM system is more similar to LLM-as-a-judge in that it can utilize any LLM model, including the latest frontier LLMs as soon as they are released. TLM is a wrapper framework on top of any base LLM that uses an efficient combination of self-reflection, consistency across sampled responses, and probabilistic measures to comprehensively quantify the trustworthiness of a LLM response.

## Benchmark Methodology

To study the real-world hallucination detection performance of these Evaluation models/techniques, we apply them to RAG datasets from different domains. Each dataset is composed of entries containing: a user query, retrieved context that the LLM should rely on to answer the query, a LLM-generated response, and a binary annotation whether this response was actually correct or not.

We focus on the most important overall concern in RAG: how effectively does each detection method **flag responses that turned out to be incorrect**. This is quantified in terms of _precision/recall_ using the **Area under the Receiver Operating Characteristic curve (AUROC)**. A detector with high AUROC more consistently assigns lower scores to RAG responses that are incorrect than those which are correct. We run all methods at default recommended settings from the corresponding libraries. Recall that both LLM-as-a-judge and TLM can be powered by any LLM model; our benchmarks run them with OpenAI’s gpt-4o-mini LLM, a low latency/cost solution.

We report results over 6 datasets, each representing different challenges in RAG applications. Four datasets stem from the [HaluBench](https://huggingface.co/datasets/PatronusAI/HaluBench) benchmark suite. The other two datasets, **FinQA** and **ELI5**, cover more complex settings.

## Benchmark Results

### FinQA

[FinQA](https://huggingface.co/datasets/Cleanlab/FinQA-hallucination-detection) is a dataset of complex questions from financial experts pertaining to public financial reports, where responses stem from OpenAI’s GPT-4o LLM. In this benchmark, almost all Evaluation models are reassuringly able to detect incorrect AI responses better than random chance (which would be AUROC = 0.5).

### ELI5

[ELI5](https://facebookresearch.github.io/ELI5/index.html) is a dataset that captures the challenge of breaking down complicated concepts into understandable/simplified explanations without sacrificing accuracy. In this benchmark, no method manages to detect incorrect AI responses with very high precision/recall, but Prometheus and TLM are more effective than the other detectors.

### FinanceBench

[FinanceBench](https://huggingface.co/datasets/PatronusAI/HaluBench/viewer/default/test?f%5Bsource_ds%5D%5Bvalue%5D=%27FinanceBench%27) is a dataset reflecting the types of questions that financial analysts answer day-to-day, based on public filings from publicly traded companies including 10Ks, 10Qs, 8Ks, and Earnings Reports. For FinanceBench, TLM and LLM-as-a-judge detect incorrect AI responses with the highest precision and recall.

### PubmedQA

[PubmedQA](https://huggingface.co/datasets/PatronusAI/HaluBench/viewer/default/test?f%5Bsource_ds%5D%5Bvalue%5D=%27pubmedQA%27) is a dataset where context comes from PubMed (medical research publication) abstracts. In this benchmark, Prometheus and TLM detect incorrect AI responses with the highest precision and recall.

### CovidQA

[CovidQA](https://huggingface.co/datasets/PatronusAI/HaluBench/viewer/default/test?f%5Bsource_ds%5D%5Bvalue%5D=%27covidQA%27&views%5B%5D=test) is a dataset to help experts answer questions related to the Covid-19 pandemic based on the medical research literature. In this benchmark, TLM detects incorrect AI responses with the highest precision and recall, followed by Prometheus and LLM-as-a-judge.

### DROP

[Discrete reasoning over paragraphs (DROP)](https://huggingface.co/datasets/PatronusAI/HaluBench/viewer/default/test?f%5Bsource_ds%5D%5Bvalue%5D=%27DROP%27) consists of passages retrieved from Wikipedia articles and questions that require discrete operations and mathematical reasoning to answer. In this benchmark, TLM detects incorrect AI responses with the highest precision and recall, followed by LLM-as-a-judge.

## Discussion

Our study presented one of the first benchmarks of real-time Evaluation models in RAG. We observed that most of the Evaluation models could detect incorrect RAG responses significantly better than random chance on some of the datasets, but certain Evaluation models didn’t fare much better than random chance on other datasets. Thus carefully consider your domain when choosing an Evaluation model.

## Resources to learn more

- [Code to reproduce these benchmarks](https://github.com/cleanlab/cleanlab-tools/tree/main/benchmarking_hallucination_model) 
- [Quickstart tutorial](https://help.cleanlab.ai/tlm/use-cases/tlm_rag/)
