When the test labels are wrong: label errors, data cascades, and datasheets

Barakaeli Lawuo, Mar 4, 2026

In 2021 Curtis Northcutt, Anish Athalye and Jonas Mueller went looking for wrong labels in the test sets of ten of the most used datasets in machine learning, covering images, text and audio. They estimated an average of at least 3.3% label errors across the ten. In the ImageNet validation set, which most papers use as ImageNet's test set because the real test labels are not public, they found 2,916 errors, about 6% of its 50,000 images. For QuickDraw they estimated more than 5 million errors, about 10%.1

The more uncomfortable result came next. When they corrected the labels and looked again at which models were best, the answer depended on how much bad data was in the test set. On ImageNet with corrected labels, ResNet-18 outperforms ResNet-50 if the share of originally mislabeled test examples rises by just 6%. On CIFAR-10, VGG-11 outperforms VGG-19 if that share rises by just 5%.1 In both pairs the model with fewer parameters, the one that looks worse on the standard leaderboard, becomes the better choice. The authors' conclusion: lower capacity models may be practically more useful than higher capacity ones on real datasets with a high proportion of wrong labels.1

Redrawn from Table 1 of Northcutt, Athalye and Mueller, 2021.1 Values count only errors that humans confirmed, so they are lower bounds. QuickDraw and Amazon Reviews were checked on random samples and extrapolated; the others had every flagged example checked. The mean of the ten is the paper's 3.3%.

Simpler datasets, and ones built with careful labeling and curation, had fewer errors than ones collected more automatically, the paper notes.1

Two accuracies, and why they disagree

To talk about what a wrong label does to a benchmark, the paper separates a few quantities that normally get lumped together as "test accuracy."1

Original accuracy
the share of test examples where the model's prediction matches the label that shipped with the dataset. This is the number everyone reports.
Corrected accuracy
the same measurement after humans have fixed the wrong labels, and removed the examples that have no clear right label. The paper argues this is the one that reflects real use.
Correctable set
the flagged test examples where human reviewers agreed on a label different from the original one. These are clear mistakes with a clear fix.
Noise prevalence
the correctable set as a fraction of the test data that remains after the ambiguous examples are dropped. For ImageNet this starts at about 2.9%.

On the full ImageNet validation set, fixing the labels barely moves the leaderboard. The authors compared 34 pretrained models and found benchmark conclusions "largely unchanged" once errors were removed.1 The disagreement shows up on the correctable set, the examples whose labels were wrong. There, models that score best against the original labels score worst against the corrected ones. NASNet-large drops from rank 1 of 34 to rank 29. ResNet-18 rises from rank 34 to rank 1.1 The same pattern held across 13 pretrained CIFAR-10 models.1

On the mislabeled images, NASNet agrees with the wrong label more often than ResNet-18 does. It has learned to reproduce the mistakes the original annotators made. The paper points out that this is not overfitting in the usual sense, since every number comes from held-out test data.1 The authors offer two explanations: smaller models may resist learning the lopsided pattern of label noise, and newer, larger architectures were tuned for years against the original test accuracy, so they may have absorbed the original annotators' quirks.1 Mistakes like these are systematic rather than random. A tiger is more likely to be mislabeled a cheetah than a CD player, as the paper puts it,1 so a flexible enough model can learn them.

At 2.9%, the correctable set is too small to flip the overall ranking. To see what happens in a messier dataset, the authors built smaller test sets by randomly deleting correctly labeled examples, which raises the noise prevalence while keeping the same errors.1 On ImageNet, the ResNet-50 and ResNet-18 lines cross at a noise prevalence of 9% when scored on corrected labels. That crossing is the "6%" in the opening: from 2.9% to 9%.1 Past that point, a team picking a model by standard test accuracy would ship the one that is worse on the true labels. The authors add that datasets collected for a specific application are often noisier than these curated benchmarks,1 which suggests, in my reading, that many real projects are already past the crossing point.

Finding 2,916 bad labels without reading 50,000

Checking every label by hand is too expensive at this scale. The team used an algorithm to shortlist likely errors first, then sent only the shortlist to people. That filter often cut the number of examples needing human review by as much as 90%.1 The algorithm is confident learning, described in an earlier paper by Northcutt, Lu Jiang and Isaac Chuang.2

Confident learning needs two inputs for every example: the label it was given, and a model's predicted probability for each class. The probabilities must be out of sample, meaning the model never trained on that example. For test sets, the team pretrained on the training set and then fine-tuned on the test set with cross-validation, so each test example's probabilities came from a model that had not seen it.1 It also assumes the noise is class-conditional: the chance a label is wrong depends on the true class, not on the particular image.1

The first step sets a threshold for each class. Call the given (possibly wrong) label y~\tilde{y} and the unknown true label y∗y^*. For class jj, the threshold tjt_j is the model's average confidence in class jj over all examples labeled jj:2

tj=1∣Xy~=j∣∑x∈Xy~=jp^(y~=j; x,θ)t_j = \frac{1}{|X_{\tilde{y}=j}|} \sum_{x \in X_{\tilde{y}=j}} \hat{p}(\tilde{y}=j;\, x, \theta)
Per-class threshold, equation 2 of Northcutt, Jiang and Chuang.2

Term by term: Xy~=jX_{\tilde{y}=j} is the set of examples whose given label is jj, and the bars mean "how many." p^(y~=j;x,θ)\hat{p}(\tilde{y}=j; x, \theta) is the probability that model θ\theta assigns to class jj for example xx. So tjt_j is a plain average: how sure the model usually is about class jj when the label says jj. Each class gets its own bar, which the authors argue makes the method robust when a model is overconfident about some classes and timid about others, or when classes are imbalanced.2

The second step counts. An example labeled ii is counted as probably truly jj if its probability for jj clears tjt_j. If it clears the bar for several classes, it goes to the one with the highest probability. The counts form an m×mm \times m table for mm classes, called the confident joint:2

Cy~,y∗[i][j]=∣X^y~=i, y∗=j∣X^y~=i, y∗=j={ x∈Xy~=i:p^(y~=j; x,θ)≥tj }\begin{gathered} C_{\tilde{y},y^*}[i][j] = |\hat{X}_{\tilde{y}=i,\,y^*=j}| \\[4pt] \hat{X}_{\tilde{y}=i,\,y^*=j} = \{\, x \in X_{\tilde{y}=i} : \\ \hat{p}(\tilde{y}=j;\, x, \theta) \ge t_j \,\} \end{gathered}
The confident joint, simplified from equation 1 of Northcutt, Jiang and Chuang,2 which also resolves ties by taking the most probable class.

Row ii, column jj says how many examples labeled ii look confidently like jj. The diagonal holds examples whose label and prediction agree. The table is then rescaled so each row sums to that label's observed share of the data and the whole table sums to 1, which turns it into Q^\hat{Q}, an estimate of the joint distribution of given and true labels.2 The diagonal of Q^\hat{Q} is the probability that an example is labeled correctly, so the estimated fraction of errors is:1

ρ=1−∑i=1mQ^y~=i, y∗=i\rho = 1 - \sum_{i=1}^{m} \hat{Q}_{\tilde{y}=i,\,y^*=i}
Estimated error fraction, from Section 3 of Northcutt, Athalye and Mueller.1 The number of suspected errors is ρ⋅n\rho \cdot n for a test set of nn examples.

To decide which ρ⋅n\rho \cdot n examples to flag, they are ranked by normalized margin: the probability of the given label minus the largest probability of any other class. The most negative margins, where the model strongly prefers another class, come first.1

Then came people. On Mechanical Turk, each flagged example went to five workers, who saw the image along with the given label and the algorithm's suggested label and said which applied: one, the other, both, or neither. An example counted as an error if fewer than three of the five agreed with the given label.1 On average across datasets, 51% of the flagged candidates turned out to be real errors.1 The paper sorts the errors into four kinds: correctable, where most workers agreed on the suggested label; multi-label, where both labels fit; neither, where no offered label fit; and non-agreement, where there was no majority.1 Of ImageNet's 2,916 errors, 1,428 were correctable.1

The method has blind spots, and the paper shows them. It flagged some correctly labeled but hard images: a close crop of part of an old sewing machine, a view of a runway from inside an airplane cockpit.1 And it misses errors it never flags. For ImageNet the team ran an expert review of 1,934 images, one flagged and one unflagged image per class where possible. A flagged image was 2.6 times as likely to be mislabeled as an unflagged one. But about 16% of the unflagged images were also mislabeled, and since 89% of images were never flagged, the authors estimate the ImageNet validation set contains closer to 20% label errors, not 6%.1 The headline numbers are floors. The corrected labels are public, with a browsable gallery of the errors.5,6

How data problems compound after deployment

A benchmark with bad labels misleads you about which model to ship. In a deployed system the damage can run longer, and the same year a Google Research team documented how. Nithya Sambasivan and colleagues interviewed 53 AI practitioners between May and July 2020, working on high-stakes applications in India (23), the US (16) and East and West African countries (14), in areas like maternal health, cancer diagnosis, credit, landslide detection and wildlife conservation.3

They named what they found a data cascade: "compounding events causing negative, downstream effects from data issues, that result in technical debt over time."3 Of the 53 practitioners, 92% reported at least one cascade, and 45.3% reported two or more in a given project.3 Cascades usually started upstream, at data collection or labeling, and showed up downstream, in evaluation or deployment. They were opaque: there were no clear tools or metrics to detect them, so practitioners fell back on proxy metrics such as accuracy or F1, which score the whole system rather than the data.3 The worst took two to three years to surface.3

Redrawn from Table 2 of Sambasivan et al., 2021.3 Percentages are of participants who self-reported the cascade in interviews; one person could report several, so they do not sum to 100.

The four cascades have distinct triggers.3 The most common came from the physical world: models trained on clean data met messy live data from cameras and sensors. In a road safety project in India, the slightest movement of a camera due to weather caused failures in detecting traffic violations. These cascades took the longest to show up, almost always in production, and ended in complete model failures and abandoned projects.3

The second came from practitioners making data decisions outside their expertise: defining ground truth, choosing features, deciding what to discard. In one wildlife project, patrollers disputed a deployed model's predicted poaching locations, and the team learned that most poaching attacks were missing from the data.3 The third came from conflicting incentives. Data collection was often added on top of field partners' existing jobs, such as nurses, patrollers and farmers, without adequate pay for the new tasks, and some lost the motivation to do it well. One conservation dataset broke because collectors forgot to reset a GPS app, which then logged every hour instead of every five minutes. "Then it is useless, and it messes up my whole ML algorithm," the practitioner said.3

The fourth, at 20.8%, is the one this post cares most about: poor documentation across organizations. Inherited datasets lacked metadata, so practitioners guessed, and guesses led to discarded or re-collected data. In a US medical robotics project, missing metadata and collaborators changing the schema without context cost four months of data collection.3 Practitioners said metadata on equipment, origin, weather, time and collection process was what they needed to judge quality and fitness for a use. The counterexample came from an aquaculture team who wrote a data curation plan before a rare collection trip and kept detailed field notes. A note recording the time of a lunch break later saved a large part of their dataset when they were diagnosing a data problem.3

The paper's title quote comes from a healthcare practitioner in India: "Everyone wants to do the model work, not the data work."3 The authors link cascades to that attitude. Data work was seen as operational, hard to track and rarely rewarded, while models brought publications and careers.3 In my reading, the Northcutt result is the same pattern in the benchmark world: years of architecture work were measured against labels nobody had rechecked.

What a datasheet asks you to write down

The prevention with the clearest written form is older than both studies. In 2018 Timnit Gebru and six coauthors proposed datasheets for datasets.4 The idea comes from electronics, where every component, however simple, ships with a datasheet listing its operating characteristics, test results and recommended usage. A dataset, they argued, should ship with a document covering its motivation, composition, collection process, recommended uses, and so on.4 It has two audiences. For the people who create a dataset, writing it forces careful reflection on assumptions, risks and implications. For the people who use it, it provides what they need to pick an appropriate dataset and avoid misusing it.4

The questions were refined over roughly two years. The authors tested them on datasheets for two public datasets and with product teams at two large US technology companies, and had lawyers review them.4 Along the way they reworded questions to discourage yes or no answers.4 The final set is grouped into seven stages: motivation, composition, collection process, preprocessing and cleaning and labeling, uses, distribution, and maintenance.4 Several of the questions read like direct answers to the failures above:4

  • "Are there any errors, sources of noise, or redundancies in the dataset?" This is the question Northcutt's team answered for ten test sets after the fact.
  • "Are there recommended data splits (e.g., training, development/validation, testing)?" with the rationale behind them.
  • "What mechanisms or procedures were used to collect the data (e.g., hardware apparatuses or sensors, manual human curation, software programs, software APIs)? How were these mechanisms or procedures validated?"
  • "Who was involved in the data collection process (e.g., students, crowdworkers, contractors) and how were they compensated?"
  • "Over what timeframe was the data collected?"
  • "Is there an erratum?" and "Will the dataset be updated (e.g., to correct labeling errors, add new instances, delete instances)?" with how often, by whom and how users will hear about it.

The mapping is my reading, not a claim either paper makes. But the sensor and procedure question covers the metadata the robotics and clean energy teams in Sambasivan's study were missing.3,4 The compensation question touches the incentive cascade. And the erratum question gives a label correction somewhere to live. The authors say plainly that writing a datasheet is not meant to be automated, because automatic documentation defeats the point, which is to make creators reflect.4 The questions are also not prescriptive; some will not fit a given dataset.4

What the authors say is still missing

Each paper is careful about what it did not show. Northcutt and colleagues say their study does not settle why high-capacity models do worse on corrected labels: overfitting to training set noise, overfitting to validation noise during hyperparameter tuning, or sensitivity to the shift created when test labels are corrected. How to split a human verification budget between training and test data is also left open.1 Sambasivan's team could not observe anyone's work directly. Because of COVID-19, every interview happened over video or phone, so the findings rest on what practitioners reported about themselves, without shadowing or contextual inquiry.3

Gebru and colleagues note that datasheets do not provide a complete solution to unwanted bias or other harms, since creators cannot anticipate every use of a dataset, and that their workflow may not fit datasets that change often.4 Their last limitation connects the three papers. Creating a datasheet will always take time, they write, and organizational infrastructure and workflows, "not to mention incentives," will need to change to make room for it.4

Sources

  1. Northcutt, Athalye, and Mueller, Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks, NeurIPS Datasets and Benchmarks 2021
  2. Northcutt, Jiang, and Chuang, Confident Learning: Estimating Uncertainty in Dataset Labels, Journal of Artificial Intelligence Research, 2021
  3. Sambasivan et al., "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI, CHI 2021
  4. Gebru et al., Datasheets for Datasets, arXiv 1803.09010 (v8, 2021)
  5. Label Errors gallery: validated test set errors from Northcutt et al., 2021
  6. cleanlab/label-errors: corrected labels and reproduction code for Northcutt et al., 2021