\begin{table}[t]
\caption{Metrics beyond final accuracy for LNL evaluation. No single row is mandatory for every study; the metric should match the claimed capability and the available reference information.}\label{tab:metrics_beyond_accuracy}
\footnotesize
\setlength{\tabcolsep}{3pt}
\begin{tabular}{@{}L{0.20\textwidth}L{0.25\textwidth}L{0.27\textwidth}L{0.18\textwidth}@{}}
\toprule
Metric or metric family & What it measures & Recommended reporting & Interpretation boundary \\
\midrule
Balanced accuracy, macro-F1, and worst-class accuracy & Robustness under class imbalance and heterogeneous class-wise corruption & Report with overall accuracy and per-class noise statistics when available & Improvements may reflect class reweighting rather than better noise handling \\
Noise-detection precision, recall, and AUPRC & Whether selected or rejected samples correspond to known label errors & Define the positive class and report the threshold or precision--recall curve & Requires instance-level clean reference labels; AUROC can obscure rare noisy cases \\
Hard-clean retention & Fraction of difficult but correctly labeled samples retained as trusted supervision & Predefine difficulty independently of the evaluated selector and report retention at a fixed selection budget & Otherwise ``hard'' can be defined circularly by the method's own loss \\
Unknown/OOD AUROC and FPR@95TPR & Separation of in-distribution mislabeled samples from examples outside the target label space & Report the unknown class construction and threshold-selection protocol & High OOD separation does not imply correct within-class relabeling \\
ECE, NLL, and Brier score & Calibration and probabilistic quality of predictions or refined targets & Report binning for ECE and distinguish test predictions from training pseudo-labels & ECE is bin-sensitive and can improve without better representation quality \\
Retained-data fraction and supervision-use breakdown & How much data is trusted, relabeled, treated as unlabeled, or rejected & Report fractions over time and per class, together with final performance & Similar accuracy may hide very different information and compute efficiency \\
Representation diagnostics & Preservation of semantic neighborhoods and transferable features & Use linear probes, $k$NN purity, prototype drift, or in-/out-of-domain transfer under a fixed protocol & A diagnostic correlation is not by itself a mechanism proof \\
Robustness profile and cost & Sensitivity across noise rates/seeds and the resources needed to obtain it & Report mean and variation, training time, model count, memory, and external pre-training access & A stronger average may rely on a larger information or compute budget \\
\botrule
\end{tabular}
\end{table}
