Model evaluation · Global concepts
Compare models under the same conditions
A report approval is not an independently reproduced experiment.
Freeze the comparison
Record the task, dataset version, split dates, baseline, evaluation rules, budget and number of trials before comparing candidates. Keep a held-out set separate from model selection.
Make a result reproducible
Preserve code and model versions, random seeds, inputs, evaluation configuration and results. Report the metric definition and uncertainty. A favorable backtest can reflect chance, data leakage or repeated selection.
Understand the current lab
The workspace can import and compare JSON reports. Its verification action records an administrator’s review; it does not rerun a model. Publishing changes a version pointer, not a production inference service. Synthetic demo results are not measured model performance.