Definition
Data quality covers correctness, completeness, duplication, schema drift, freshness, and labeling consistency. Poor quality silently tanks model metrics and user trust.
Checks belong in pipelines: null rates, range checks, distribution drift, and referential integrity.
In simple terms
Cooking with spoiled ingredients — no recipe skill saves the meal. Models are the same with bad data.
Where you see it
- Deduplicating web text before LLM pretraining.
- Fixing mislabeled intent examples that confuse a chatbot.
How it works
1.Profile
Measure nulls, outliers, and class balance.
2.Validate
Automated tests on every pipeline run.
3.Remediate
Clean, relabel, or quarantine bad slices.
Why it matters
- "Garbage in, garbage out" is still the most expensive lesson in AI projects.
Often confused
More rows fix quality issues.
Scaling bad labels scales the mistakes — clean first.