\begin{table}[t]
\caption{Evaluation axes for realistic noisy-supervision benchmarks.}\label{tab:evaluation_axes}
\small
\begin{tabular}{@{}L{0.20\textwidth}L{0.25\textwidth}L{0.25\textwidth}L{0.20\textwidth}@{}}
\toprule
Axis & Key question & Typical confound & Reporting expectation \\
\midrule
Noise source & Who or what generated the label? & Human ambiguity, web collection, weak rules, LLM prompts & Describe annotation or generation process \\
Noise structure & How does noise depend on data? & Class imbalance, instance difficulty, open-set samples & Report beyond a single global noise rate \\
Observability & What clean information is available? & Clean validation data, repeated labels, known priors & State information budget explicitly \\
Evaluation unit & What must be robust? & Accuracy, detection, calibration, representation, transfer & Use metrics aligned with the claimed problem \\
Transfer setting & Where is the model evaluated? & In-domain versus out-of-domain shift & Separate robustness from domain adaptation \\
\botrule
\end{tabular}
\end{table}
