The process of identifying and correcting errors, missing values, and inconsistencies in a dataset before analysis.
Real-world data is messy — duplicate records, missing fields, inconsistent formatting. Most practicing data professionals spend a significant portion of their time on cleaning before any actual analysis begins.