\begin{table}[t]
\caption{Representative real-noise or real-source benchmarks, their diagnostic value, and comparison caveats.}\label{tab:benchmark_map}
\scriptsize
\setlength{\tabcolsep}{2.5pt}
\renewcommand{\arraystretch}{0.95}
\begin{tabular}{@{}L{0.18\textwidth}L{0.27\textwidth}L{0.25\textwidth}L{0.20\textwidth}@{}}
\toprule
Benchmark and task & Noise source and clean signal & Diagnostic value & Caveat and minimum reporting \\
\midrule
CIFAR-N: image classification \citep{wei2022cifarn} & Human annotations with original CIFAR labels retained as a reference & Tests whether human label errors can be identified without discarding hard clean samples & Low-resolution closed set; report accuracy, detection quality, and hard-clean retention \\
Clothing1M: clothing recognition \citep{xiao2015learning} & Shopping metadata with clean train, validation, and test subsets & Tests high-volume web supervision and class imbalance & State exactly which clean subsets and pre-training are used \\
WebVision: large-scale image classification \citep{li2017webvision} & Search-query labels with curated validation and test sets & Tests web-scale supervision, transfer, and dataset bias & Separate label-noise robustness from web-to-curated domain shift \\
Food-101N: food recognition \citep{lee2018cleannet} & Web recipe images with partially verified labels & Tests taxonomy noise under limited label verification & Report verification coverage and treatment of irrelevant images \\
ANIMAL-10N: animal recognition \citep{song2019selfie} & Human labels for visually similar animal classes with a clean test set & Tests fine-grained semantic confusion & Narrow domain; report class-wise effects rather than only mean accuracy \\
Controlled web labels: image recognition \citep{jiang2020beyond} & Professionally relabeled web subsets with controlled noise levels & Tests web noise while retaining an estimate of corruption rate & More curated than unconstrained web collection; document subset construction \\
WebFG and WebiNat: fine-grained recognition \citep{sun2021webly} & Web searches over fine-grained categories and hierarchies & Tests semantic hierarchy, imbalance, and irrelevant retrievals & No complete instance-level clean reference; state rejection and verification rules \\
NoisywikiHow: intent classification \citep{wu2023noisywikihow} & Heterogeneous instance-dependent text noise & Tests whether image-derived assumptions transfer to text & Constructed benchmark; report noise provenance and calibration \\
BeGIN: graph node classification \citep{kim2025begin} & Algorithmic and LLM-based corruptions embedded in graph structure & Tests node-specific noise and propagation through topology & Graph-specific; report topology effects separately from node accuracy \\
\botrule
\end{tabular}
\renewcommand{\arraystretch}{1}
\end{table}
