Trustworthy Language Model (TLM)
The Trustworthy Language Model (TLM) overcomes the biggest barrier to enterprise adoption of LLMs: hallucinations and reliability. By adding a trust score to every LLM response, TLM helps you automatically catch incorrect LLM outputs in real-time. This enables you to deploy generative AI for new use cases previously unsuitable for LLMs. Rigorous benchmarking shows that: TLM has better-calibrated trustworthiness scores (enabling greater cost/time savings) than existing approaches to detect LLM errors, and TLM can utilize these trustworthiness scores to produce more accurate responses than existing LLMs.
LLMs’ biggest challenge: hallucinations
A recent Gartner poll shows that while 55% of organizations are experimenting with generative AI, only 10% have put generative AI into production. A major barrier to productionizing LLMs is their occasional tendency to produce bogus outputs known as hallucinations, which precludes their use in applications where correct outputs are necessary (i.e., most applications)!
Despite their brittle nature, organizations have deployed LLMs, sometimes with catastrophic results. Air Canada’s chatbot hallucinated refund policies, resulting in the airline being held responsible for the misinformation and monetary penalties; the chatbot has since been taken down. A federal judge fined a law firm after their lawyers used ChatGPT to draft a brief full of fabricated citations. New York City’s “MyCity” chatbot has been hallucinating wrong answers to business owners’ questions about local laws.
Overcoming hallucinations with trustworthiness scores
LLMs will always exhibit occassional hallucinations and incorrect responses, but by providing a trustworthiness score with every output, TLM lets you identify when the LLM is hallucinating.
The TLM API can serve as:
- A drop-in replacement for your LLM. Like existing LLM APIs, TLM provides a
.prompt()method that will return a response along with a trustworthiness score, enabling more reliable AI deployments. - A layer of trust for your existing LLM outputs or human-generated data. TLM provides a
.get_trustworthiness_score()method that can score any prompt/response pair to detect bad/wrong responses in real-time.
TLM works by augmenting existing LLMs with a layer of trust. The generally-available version of TLM lets you choose between a number of popular base models, including GPT-4o, GPT-4o mini, GPT-4, GPT-3.5, o1-preview, Claude 3 and 3.5 Sonnet, but TLM can augment any LLM. For enterprise use cases, such as adding trustworthiness to your custom fine-tuned LLM, contact us.
Use cases enabled by TLM
Trustworthiness scores unlock new production use cases of LLMs, and any existing application of LLMs can also benefit by taking into account these scores.
Customer service chatbot
TLM powers trustworthy chatbots that answer the 80% of questions where they are confident, but escalate to a human if they’re unsure about a response rather than hallucinating one. This is done simply by routing the question to a human when the trustworthiness score falls below a chosen threshold. If human escalation is not possible, untrustworthy responses can at least visually be flagged.
Auto-labeling
LLMs are commonly used for auto-labeling data. With TLM, you can confidently auto-label a large fraction of your data and only have humans review a portion of the data where the LLM does not return trustworthy results.
template = '''
What type of compliance issue is most likely present in the following document?
Your answer should be selected from the following options: HIPAA, FERPA, GDPR, none.
Document below here:
{document}
'''
def classify(document) -> Tuple[str, float]:
answer = tlm.prompt(template.format(document=document))
return answer['response'], answer['trustworthiness_score']
Using this prompt to classify a large number of legal documents, we see that the documents with high trustworthiness scores were labeled correctly, while the documents with low scores often received erroneous labels that needed double-checking:
| document | response | trustworthiness |
|---|---|---|
| All medical health records will be accessed one way only. The patient’s medical data will be stored on unencrypted public servers at the discretion of the enterprise customer. | HIPAA | 0.984 |
| TechTarget’s Cookies Policy includes the following terminology: “By continuing to use the site, you agree to the use of cookies.” | FERPA | 0.426 |
Data extraction
TLM can also be used for open-domain data extraction.
If you were populating a parts catalog, you might be interested in extracting information like operating voltage from such documents, where TLM’s trustworthiness scores can automatically separate correctly extracted values from those that are wrong:
| part | operating voltage | trustworthiness |
|---|---|---|
| ATtiny44A | 1.8 - 5.5V | 0.937 |
| ZRE200GE | 1V - 15V DC | 0.567 |
Evaluating TLM Performance
We evaluate TLM’s ability to add trust to arbitrary LLMs by benchmarking TLM against OpenAI’s GPT-4 LLM (and many other models in the Appendix). Our comprehensive benchmarks investigate two questions to evaluate the reliability of TLM’s (1) responses, and (2) trustworthiness scores:
- How accurate are TLM responses compared to the baseline LLM?
- How much costs/time does a team save by scoring responses via TLM vs. existing confidence estimation approaches?
Benchmark datasets
Our study focuses on Q&A settings. Unlike other LLM benchmarks, we never measure benchmark performance using LLM-based evaluations. All of our benchmarks involve questions with a single correct answer. We consider these popular Q&A datasets:
- TriviaQA: Open-domain trivia questions.
- ARC: Grade school multiple-choice questions.
- SVAMP: Elementary-level math word problems.
- GSM8k: Grade school math problems.
- Diagnosis: Diagnosing medical conditions based on symptom descriptions from the patient.
Benchmark Results
The following table reports the accuracy of responses from TLM and GPT-4 across each benchmark dataset:
| Dataset | OpenAI GPT-4 API | Cleanlab TLM API |
|---|---|---|
| TriviaQA | 84.7% | 84.8% |
| ARC | 94.6% | 94.9% |
| SVAMP | 90.7% | 91.7% |
| GSM8k | 46.5% | 55.6% |
| Diagnosis | 67.4% | 68.0% |
Conclusion
This article shows how the TLM technology can boost the reliability of any LLM application. Use TLM trustworthiness scores to automatically catch bad outputs from any LLM in real-time. Additionally, use TLM to produce more accurate responses than any base LLM model.
Of course, there’s no free lunch. TLM uses extra computation to provide these benefits. It internally calls the underlying base LLM multiple times to self-reflect on candidate responses, compute probabilistic measures, and assess the semantic consistency between candidate responses.