# Learn more about Data-Centric AI

Data-Centric AI is the systematic engineering of better data (via AI and automation). Learn about key concepts, useful tricks, and helpful tools.

[**CROWDLAB: The Right Way to Combine Humans and AI for LLM Evaluation**  
CROWDLAB improves your team's LLM Evals process by automatically producing reliable ratings and flagging which outputs need further review.](/content/blog/learn/team-llm-evals/index.html)

[**How to detect bad data in your instruction tuning dataset (for better LLM fine-tuning)**  
Overview of automated tools for catching: low-quality responses, incomplete/vague prompts, and other problematic text (toxic language, PII, informal writing, bad grammar/spelling) lurking in a instruction-response dataset. Here we reveal findings for the Dolly dataset.](/content/blog/learn/filter-llm-tuning-data/index.html)

[**Automatically Detect Problematic Content in any Text Dataset**  
Introducing AI text audits for automated content moderation and curation, including the detection of: toxic, non-English, and informal language, as well as personally identifiable information.](/content/blog/learn/text-content-moderation/index.html)

[**Improving any OpenAI Language Model by Systematically Improving its Data**  
Reduce LLM prediction error by 37% via data-centric AI.](/content/blog/learn/fine-tune-LLM/index.html)

[**ActiveLab: Active Learning with Data Re-Labeling**  
ActiveLab helps you optimally choose which data to (re)label, lowering the cost to train an accurate ML model.](/content/blog/learn/active-learning/index.html)

[**CROWDLAB: Simple and effective algorithms to handle data labeled by multiple annotators**  
Understanding cleanlab's new methods for multi-annotator data and what makes them effective.](/content/blog/learn/multiannotator/index.html)

[**Automatically catching spurious correlations in ML datasets**  
An open-source module to detect spurious correlations between dataset labels and features that will not generalize to real-world deployment.](/content/blog/learn/spurious-correlations/index.html)

[**Announcing Auto-Labeling Agent: Your Assistant for Rapid and High Quality Labeling**  
Generate AI, not headaches. Automate annotation with AI.](/content/blog/learn/auto-labeling/index.html)

[**An open-source platform to catch all sorts of issues in all sorts of datasets**  
With cleanlab v2.6, the most popular library for Data-Centric AI now offers more comprehensive data audits including new checks for underperforming groups, null values, imbalanced classes, and more.](/content/blog/learn/cleanlab-2.6)

[**Comparing tools for Data Science, Data Quality, Data Annotation, and AI/ML**  
What's the next-generation platform for Data Science? A data-centric AI system that can automatically: find and fix data issues, label data, and train/deploy reliable models.](/content/blog/learn/tools/index.html)

[**How to Filter Unsafe and Low-Quality Images from any Dataset: A Product Catalog Case Study**  
Introducing an automated solution to ensure high-quality image data, for both content moderation and boosting engagement. Easily curate any product/content catalog or photo gallery to delight your customers.](/content/blog/learn/image-issues/index.html)

[**Detecting Annotation Errors in Semantic Segmentation Data**  
Introducing new methods for estimating labeling quality in image segmentation datasets.](/content/blog/learn/segmentation-errors/index.html)

Learn more from the first-ever [course on Data-Centric AI](https://dcai.csail.mit.edu/) taught at MIT by the Cleanlab team and made freely available.

Lecture 1: Data-Centric AI vs. Model-Centric AI - YouTube  
[Lecture 1: Data-Centric AI vs. Model-Centric AI](https://www.youtube.com/watch?v=ayzOzZGHZy4) [Introduction to Data-Centric AI](https://www.youtube.com/channel/UC0aqiaC_HliDfEDboDiDv9w)

Introduction to Data-Centric AI2.45K subscribers
