Mislabeled Test Data can destabilize Machine Learning Benchmarks
An interestring new study by Northcutt et al shows that mislabeled test data is pervasive in ML benchmarking trials. By correcting test data labels, ML benchmarking can better predict real-world performance.
Small increases in the prevalence of originally mislabeled test data can destabilize ML benchmarks, indicating that low-capacity models may actually outperform high-capacity models in noisy real-world applications, even if their measured performance on the original test data may be worse.