DecayBench: A Reference-Free Benchmark for Trustworthy Drift Detection
Abstract
1. Introduction
- DecayBench: A reference-free, calibrated benchmark for drift detection with conformal false-alarm guarantees (Proposition A1), statistical significance testing (Section 4.4), and cross-modal evaluation (Table 13).
- Alert: A label-free, self-configuring aggregation rule with finite-sample calibration (Proposition A1) and a no-regret guarantee (Proposition A2).
- Trustworthiness measurement: A definition of reference-free drift-detection trustworthiness (Definition 3), made measurable along five axes (calibrated false-alarm control, validity, timeliness, no-regret, and adaptivity) under the label-free precondition (Section 4.4).
- Empirical results: A significance-tested study showing that drift detectors are regime-dependent rather than uniformly optimal (Table 4), with consistent Alert gains across NLP (Table 9), vision (Table 12), and multimodal settings (Table 13), especially under low-severity (Table 9) and cross-modal drift (Table 13).
2. Related Work
2.1. Drift and OOD Detection
2.2. Evaluating Drift Detectors
2.3. Reference-Free Reliability Estimation and Score Combination
3. Problem Setup
3.1. Distribution Drift
3.2. Drift Detection
4. DecayBench: Evaluation Framework
4.1. Typed, Graded Drift Streams
4.2. No-Peeking Calibration
4.3. Accuracy Metrics
| Algorithm 1: DecayBench protocol: calibrated drift detection and evaluation |
![]() |
4.4. Trustworthiness Metrics
- Calibrated. Can an operator trust and budget the alarm? The clean false-alarm rate is held at the target,fixing on clean data alone (Section 4.2); we ground this in a finite-sample conformal guarantee (Proposition A1) that the combiner inherits (Corollary A1). Calibration plays a dual role, both an axis and the common operating point at which the others are read.
- Valid. Does an alert mean something has actually gone wrong? The flagged drift is consequential: deployed accuracy is strictly decreasing in severity,verified by the accuracy–decay study (Section 6.4.2).
- Timely. Does the alert arrive in time to act? Given a severity ramp (a stream of batches with drift onset at ) and a delay budget , the detection delay is bounded,measured in Section 6.4.3.
- No-regret. Is the deployed monitor a dependable default? For M constituent detectors , their combiner , and a tolerance , the regret (the AUC gap to the best constituent) is non-positive up to significance,with a paired bootstrap failing to reject for every k (Table 9; empirically ).
- Adaptive. Will it still work on the next task or backbone? A label-free selector S maps a deployment j’s clean data to a detector subset ; writing for detection ROC-AUC on deployment j, re-selecting is never worse than transferring a family chosen on a different deployment ,realised by the label-free contextual selection (Section 5.3).
5. Alert: A Calibrated Label-Free Combiner
5.1. Combining Complementary Detectors
5.2. Why Combining Helps
5.3. Label-Free Detector Selection
| Algorithm 2: Alert: label-free self-configuring drift alert |
![]() |
6. Experiments
6.1. Datasets and Deployed Models
6.2. Baseline Detectors
6.3. Label-Free Detector Performance
6.4. Trustworthiness
6.4.1. Calibrated
6.4.2. Valid
6.4.3. Timely
6.4.4. No-Regret
6.4.5. Adaptive
6.5. Generalisation
Is the Multimodal Gain the Combiner or the Alignment Feature?
6.6. Discussion
7. Conclusions
8. Limitations and Future Work
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Appendix A. Appendix Map (Guide to Supplementary Materials)
- What is defined (detectors and scoring functions). Appendix B formally defines all ten detector families used in the main experiments, including (i) training-free baselines (cosine, MMD, Mahalanobis, energy, MSP, ODIN, KNN, DriftLens, BBSD), and (ii) the trained reconstruction-based monitor (Recon). These definitions specify exactly what is compared in Table 3.
- What the detectors actually measure (behavioural validity). Appendix C analyzes failure modes and identifiability: structured reconstruction signals do not reliably infer drift type (Table A1), and performance is strongly batch-size dependent (Table A2). These results explain why aggregate ROC-AUC alone is insufficient to characterise detector quality.
- What happens under realistic drift regimes (robustness). Appendix D evaluates detectors under (i) graded natural-drift mixtures (Table A3) and (ii) temporal real-world drift (Amazon reviews). The key finding is that synthetic severity rankings generalise to real drift, while some real temporal shifts are benign (no accuracy decay, no meaningful alarms).
- What happens beyond the single-modality setting (extensions). Appendix E reports the supporting tables relocated from the main text, together with multimodal and compound-drift experiments: real image-caption pairs, cross-modal mismatch, and multi-channel monitoring with a deliberately useless signal. The results show that cross-modal alignment and robust aggregation dominate in heterogeneous settings.
- Operating-point validity (precision/recall and calibration). Appendix E also reports precision/recall at a fixed false-positive rate and AUPRC, confirming that ROC-AUC trends are not artefacts of thresholding and that training-free detectors remain competitive at calibrated operating points.
- Why the method works (theory). Appendix F provides two guarantees: (i) a conformal, distribution-free false-positive control result for calibrated thresholds, and (ii) a no-regret analysis showing that averaging standardised detector scores improves signal-to-noise under a location-shift model, explaining the empirical success of Alert.
- How to reproduce everything (benchmark release). Appendix G describes DecayBench, including frozen evaluation sets, calibration protocol, detector API, and full script list. All results are reproducible from public data with fixed seeds and deterministic batching.
- Notation and equations. Appendix H and Appendix I provide a complete reference for symbols and scoring functions used throughout the paper.
- How to read this appendix. If the reader is interested in:
- Definitions: Appendix B
- Failure modes and limitations: Appendix C
- Real-world robustness: Appendix D
- Multimodal behaviour: Appendix E
- Theory: Appendix F
- Reproducibility: Appendix G
Appendix B. Detector Definitions
Appendix B.1. Trained Reconstruction Monitor (Recon)
| Algorithm A1: Recon batch drift score |
![]() |
Appendix B.2. Training-Free Baselines
Appendix B.3. Drift-Type Typology Decomposition
Appendix C. Behavioural Analysis of Detectors
Appendix C.1. Typology: Structured Signals Do Not Reliably Diagnose Drift Type
| Task | Only | Only | Only | Energy Only | |
|---|---|---|---|---|---|
| SST-2 | |||||
| CoLA | |||||
| Mean |

Appendix C.2. Sensitivity to Batch Size
| Batch Size | |||||
|---|---|---|---|---|---|
| 16 | |||||
| 32 | |||||
| 64 | |||||
| 128 |
Appendix D. Robustness to Drift Regimes
Appendix D.1. Graded and Temporal Natural Drift
| Detector | |||||
|---|---|---|---|---|---|
| Energy | |||||
| MSP | |||||
| ODIN | |||||
| DriftLens | |||||
| MMD | |||||
| Recon | |||||
| SageMaker (cosine) | |||||
| Alert (ours) |
Temporal Drift in the Wild, and the Benign-Drift Check
Appendix E. Additional Experimental Results
Appendix E.1. Detection Severity Curves

Appendix E.2. Supervised Upper Bound
| Detector | SST-2 | CoLA | ||
|---|---|---|---|---|
| Supervised probe (trained) | ||||
| Deep-kernel MMD (trained) | ||||
| Energy (free) | ||||
| MMD (free) | ||||
| Alert (free, ours) | ||||
Appendix E.3. Additional Signals: DriftLens, ODIN, KNN
| Signal | SST-2 | CoLA | ||||
|---|---|---|---|---|---|---|
| DriftLens | ||||||
| ODIN | ||||||
| KNN | ||||||
Appendix E.4. Multimodal Supplementary Experiments
Appendix E.4.1. Real Image-Caption Pairs
Appendix E.4.2. Compound Drift and Robustness to a Dead Channel
Appendix E.5. Operating-Point Metrics
| Task | Recon | MMD | Energy | |
|---|---|---|---|---|
| SST-2 | ||||
| CoLA | ||||
Appendix F. Theoretical Analysis
Appendix F.1. Distribution-Free False-Positive Control
Appendix F.2. No-Regret Analysis of the Combiner
Empirical Validation
| Task | Severity | Observed | Pred (Indep.) | Pred (Corr.) |
|---|---|---|---|---|
| SST-2 | ||||
| CoLA | ||||
Appendix G. Reproducibility and Benchmark Release
- Frozen eval set. build_benchmark.py deterministically materialises, per task (SST-2, CoLA), the clean batches with a fixed calibration/evaluation split, graded data-drift batches at each severity , and natural OOD batches (IMDB, Yelp), with a manifest.json. Default batch size 64 (--batch_size 32 regenerates the finer split used for the powered significance test of Table 9); seeds ; threshold calibrated to the 95th percentile of clean scores (target FPR ). Source data are public, so the eval set regenerates exactly from one command.
- Harness and detector API. A detector is any callable plus a higher_is_drift flag; eval_benchmark.py calibrates its threshold on the clean calibration batches and reports all metrics on the frozen eval set. All ten reference detectors, the Alert combiner, and the significance, delay, natural-drift, and scope-expansion scripts (extended_experiment.py, stats_experiment.py, delay_experiment.py, natural_experiment.py, natural_graded_experiment.py, contextual_experiment.py, contextual_extra.py, contextual_roberta_experiment.py, mnli_experiment.py) reproduce the tables in this paper.
- Models. The deployed classifiers (bert-base-uncased per task, plus the three-class MNLI backbone of Section 6.5 and the roberta-base [55] SST-2 backbone of Section 5.3) and the reconstruction monitors are released with the harness.
Appendix H. Notation
| Symbol | Meaning |
|---|---|
| deployed classifier | |
| input batch at round t | |
| n | batch size |
| [CLS] embedding of input i () | |
| classifier logits for input i | |
| softmax of | |
| input space | |
| K | number of classes |
| d | embedding dimension ( for bert-base-uncased) |
| y | true class label |
| detector drift score for batch | |
| alert threshold | |
| target false-positive rate (FPR) | |
| m | calibration-set size |
| drift severity (fraction of tokens perturbed) | |
| reconstruction loss of input i (Recon) | |
| R | training reconstruction-loss reference (Recon) |
| training embedding centroid (cosine baseline) | |
| training embedding mean and covariance (Mahalanobis) | |
| Recon data / confidence / latent signals | |
| M | number of constituent detectors combined by Alert |
| clean-standardised score of detector k (Alert) | |
| Alert combiner statistic, (mean-z) | |
| effect size of detector k (location-shift model) | |
| clean correlation matrix of standardised scores | |
| all-ones vector |
Appendix I. Equations Summary
- Equation (A1): Recon batch drift score (reconstruction-loss KS distance).
- Equation (A11): calibrated false-positive control.
- Equation (A2): cosine (SageMaker-style) score.
- Equation (A3): maximum mean discrepancy (MMD).
- Equation (A4): Mahalanobis score.
- Equation (A5): black-box shift detection (BBSD).
- Equation (A6): energy score.
- Equation (A7): maximum-softmax-probability (MSP).
Appendix J. Tables Index
Appendix J.1. Where Alert Is the Point
- Table 9: Alert is best at every SST-2 severity (a) and never significantly worse than any single detector (b, no-regret).
- Table A7: the location-shift model predicts the observed combiner AUC.
- Table 10: label-free selection recovers the right detector family with no hand-tuning (a), and helps most under mismatch (b).
- Table 12: Alert generalises, transferring to a three-class sentence-pair task (a) and to vision and a CLIP backbone (b).
- Table 13: Alert strictly improves on single-modality monitoring under multimodal drift (a), catching drift invisible to either modality alone (b).
Appendix J.2. Comparing Existing Detectors (No Alert Claim), on a Stated Aspect
- Table 1: existing benchmarks/suites versus DecayBench; aspect: which evaluation properties each provides.
- Table 4: detection AUC by task and severity; aspect: no single detector wins at low severity.
- Table A5: three further detectors; aspect: AUC, none is uniformly best either.
- Table A4: supervised versus free detectors; aspect: a supervised detector wins only given drift labels.
- Table 11: natural domain shift; aspect: the industrial cosine check fails while confidence/MMD detect.
- Table A3: graded natural drift; aspect: a severity ramp discriminates detectors that the whole-corpus shift cannot.
- Table 8: aspect: detection delay (timeliness).
- Table 6: aspect: precision/recall/F1 at the calibrated FPR.
- Table A6: aspect: AUPRC.
Appendix J.3. Setup and Validity (Neither Comparison nor Alert Claim)
References
- Socher, R.; Perelygin, A.; Wu, J.; Chuang, J.; Manning, C.D.; Ng, A.Y.; Potts, C. Recursive Deep Models for Semantic Compositionality over a Sentiment Treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing (EMNLP), Seattle, WA, USA, 18–21 October 2013; pp. 1631–1642. [Google Scholar]
- Warstadt, A.; Singh, A.; Bowman, S.R. Neural Network Acceptability Judgments. Trans. Assoc. Comput. Linguist. 2019, 7, 625–641. [Google Scholar] [CrossRef] [Scilit]
- Cuong, H.; Xu, J. Assessing Quality Estimation Models for Sentence-Level Prediction. In Proceedings of the 27th International Conference on Computational Linguistics (COLING), Santa Fe, NM, USA, 20–26 August 2018. [Google Scholar]
- Fomicheva, M.; Sun, S.; Yankovskaya, L.; Blain, F.; Guzmán, F.; Fishel, M.; Aletras, N.; Chaudhary, V.; Specia, L. Unsupervised Quality Estimation for Neural Machine Translation. Trans. Assoc. Comput. Linguist. 2020, 8, 539–555. [Google Scholar] [CrossRef] [Scilit]
- Rabanser, S.; Günnemann, S.; Lipton, Z.C. Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 8–14 December 2019. [Google Scholar]
- Yang, J.; Zhou, K.; Li, Y.; Liu, Z. Generalized Out-of-Distribution Detection: A Survey. Int. J. Comput. Vis. 2024, 132, 5635–5662. [Google Scholar] [CrossRef] [Scilit]
- Lu, J.; Liu, A.; Dong, F.; Gu, F.; Gama, J.; Zhang, G. Learning under Concept Drift: A Review. IEEE Trans. Knowl. Data Eng. 2019, 31, 2346–2363. [Google Scholar] [CrossRef] [Scilit]
- Liu, G.; Borovica-Gajic, R. DriftBench: Defining and Generating Data and Query Workload Drift for Benchmarking. arXiv 2025, arXiv:2510.10858. [Google Scholar]
- Fisher, R.A. Statistical Methods for Research Workers; Oliver and Boyd: Edinburgh, UK, 1925. [Google Scholar]
- Simes, R.J. An improved Bonferroni procedure for multiple tests of significance. Biometrika 1986, 73, 751–754. [Google Scholar] [CrossRef] [Scilit]
- Bonferroni, C.E. Teoria statistica delle classi e calcolo delle probabilità. Pubbl. R Ist. Super. Sci. Econ. E Commer. Firenze 1936, 8, 3–62. [Google Scholar]
- Stouffer, S.A.; Suchman, E.A.; DeVinney, L.C.; Star, S.A.; Williams, R.M. The American Soldier: Adjustment During Army Life; Princeton University Press: Princeton, NJ, USA, 1949; Volume 1. [Google Scholar]
- Hendrycks, D.; Gimpel, K. A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks. In Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France, 24–26 April 2017. [Google Scholar]
- Liang, S.; Li, Y.; Srikant, R. Enhancing the Reliability of Out-of-distribution Image Detection in Neural Networks. In Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Liu, W.; Wang, X.; Owens, J.; Li, Y. Energy-based Out-of-distribution Detection. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 6–12 December 2020. [Google Scholar]
- Sun, Y.; Guo, C.; Li, Y. ReAct: Out-of-distribution Detection with Rectified Activations. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montreal, QC, Canada, 6–14 December 2021. [Google Scholar]
- Wang, H.; Li, Z.; Feng, L.; Zhang, W. ViM: Out-Of-Distribution with Virtual-logit Matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022. [Google Scholar]
- Sun, Y.; Ming, Y.; Zhu, X.; Li, Y. Out-of-Distribution Detection with Deep Nearest Neighbors. In Proceedings of the International Conference on Machine Learning (ICML), Baltimore, MD, USA, 18–23 July 2022. [Google Scholar]
- Lee, K.; Lee, K.; Lee, H.; Shin, J. A Simple Unified Framework for Detecting Out-of-Distribution Samples and Adversarial Attacks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montreal, QC, Canada, 3–8 December 2018. [Google Scholar]
- Gretton, A.; Borgwardt, K.M.; Rasch, M.J.; Schölkopf, B.; Smola, A. A Kernel Two-Sample Test. J. Mach. Learn. Res. 2012, 13, 723–773. [Google Scholar]
- Lopez-Paz, D.; Oquab, M. Revisiting Classifier Two-Sample Tests. In Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France, 24–26 April 2017. [Google Scholar]
- Lipton, Z.C.; Wang, Y.X.; Smola, A. Detecting and Correcting for Label Shift with Black Box Predictors. In Proceedings of the International Conference on Machine Learning (ICML), Stockholm, Sweden, 10–15 July 2018. [Google Scholar]
- Greco, S.; Vacchetti, B.; Apiletti, D.; Cerquitelli, T. Unsupervised Concept Drift Detection from Deep Learning Representations in Real-time. arXiv 2024, arXiv:2406.17813. [Google Scholar]
- An, J.; Cho, S. Variational Autoencoder Based Anomaly Detection Using Reconstruction Probability; Special Lecture on IE; SNU Data Mining Center: Seoul, Republic of Korea, 2015; Volume 2, pp. 1–18. [Google Scholar]
- Nalisnick, E.; Matsukawa, A.; Teh, Y.W.; Gorur, D.; Lakshminarayanan, B. Do Deep Generative Models Know What They Don’t Know? In Proceedings of the International Conference on Learning Representations (ICLR), New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Podolskiy, A.; Lipin, D.; Bout, A.; Artemova, E.; Piontkovskaya, I. Revisiting Mahalanobis Distance for Transformer-Based Out-of-Domain Detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 2–9 February 2021. [Google Scholar]
- Colombo, P.; Gomes, E.D.C.; Staerman, G.; Noiry, N.; Piantanida, P. Beyond Mahalanobis-Based Scores for Textual OOD Detection. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 28 November–9 December 2022. [Google Scholar]
- Uppaal, R.; Hu, J.; Li, Y. Is Fine-tuning Needed? Pre-trained Language Models Are Near Perfect for Out-of-Domain Detection. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), Toronto, ON, Canada, 9–14 July 2023. [Google Scholar]
- Desai, S.; Durrett, G. Calibration of Pre-trained Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Virtual, 16–20 November 2020. [Google Scholar]
- Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K.Q. On Calibration of Modern Neural Networks. In Proceedings of the International Conference on Machine Learning (ICML), Sydney, NSW, Australia, 6–11 August 2017. [Google Scholar]
- Gama, J.; Žliobaitė, I.; Bifet, A.; Pechenizkiy, M.; Bouchachia, A. A Survey on Concept Drift Adaptation. ACM Comput. Surv. 2014, 46, 1–37. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, J.; Wang, P.; Zou, D.; Zhou, Z.; Ding, K.; Peng, W.; Wang, H.; Chen, G.; Li, B.; Sun, Y.; et al. OpenOOD: Benchmarking Generalized Out-of-Distribution Detection. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, New Orleans, LA, USA, 28 November–9 December 2022. [Google Scholar]
- Koh, P.W.; Sagawa, S.; Marklund, H.; Xie, S.M.; Zhang, M.; Balsubramani, A.; Hu, W.; Yasunaga, M.; Phillips, R.L.; Gao, I.; et al. WILDS: A Benchmark of in-the-Wild Distribution Shifts. In Proceedings of the International Conference on Machine Learning (ICML), Virtual, 18–24 July 2021. [Google Scholar]
- Vovk, V.; Petej, I.; Nouretdinov, I.; Ahlberg, E.; Carlsson, L.; Gammerman, A. Retrain or Not Retrain: Conformal Test Martingales for Change-Point Detection. In Proceedings of the Tenth Symposium on Conformal and Probabilistic Prediction and Applications (COPA), Virtual, 29–31 August 2021; Volume 152. [Google Scholar]
- Zhao, Z.; Gao, H.; Xing, N.; Zeng, L.; Zhang, M.; Chen, G.; Rigger, M.; Ooi, B.C. NeurBench: Benchmarking Learned Database Components with Data and Workload Drift Modeling. arXiv 2025, arXiv:2503.13822. [Google Scholar]
- Bao, H.; Zhang, Z.; Jing, P.; Yuan, Z.; Shi, K.; Ye, Y. Drift-Bench: Diagnosing Cooperative Breakdowns in LLM Agents under Input Faults via Multi-Turn Interaction. arXiv 2026, arXiv:2602.02455. [Google Scholar]
- Dadalto, E.; Alberge, F.; Duhamel, P.; Piantanida, P. Combine and Conquer: A Meta-Analysis on Data Shift and Out-of-Distribution Detection. Trans. Mach. Learn. Res. (TMLR). 2024. Available online: https://openreview.net/forum?id=VGNBUS9TrU (accessed on 11 August 2026).
- Bates, S.; Candès, E.; Lei, L.; Romano, Y.; Sesia, M. Testing for Outliers with Conformal p-values. Ann. Stat. 2023, 51, 149–178. [Google Scholar] [CrossRef] [Scilit]
- Fawcett, T. An introduction to ROC analysis. Pattern Recognit. Lett. 2006, 27, 861–874. [Google Scholar] [CrossRef] [Scilit]
- Davis, J.; Goadrich, M. The relationship between Precision-Recall and ROC curves. In Proceedings of the 23rd International Conference on Machine Learning (ICML), Pittsburgh, PA, USA, 25–29 June 2006; pp. 233–240. [Google Scholar]
- van Rijsbergen, C.J. Information Retrieval, 2nd ed.; Butterworths: London, UK, 1979. [Google Scholar]
- Basseville, M.; Nikiforov, I.V. Detection of Abrupt Changes: Theory and Application; Prentice Hall: Englewood Cliffs, NJ, USA, 1993. [Google Scholar]
- Cesa-Bianchi, N.; Lugosi, G. Prediction, Learning, and Games; Cambridge University Press: Cambridge, UK, 2006. [Google Scholar]
- Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; Bowman, S.R. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In Proceedings of the International Conference on Learning Representations (ICLR), New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), Minneapolis, MN, USA, 2–7 June 2019; pp. 4171–4186. [Google Scholar]
- Maas, A.L.; Daly, R.E.; Pham, P.T.; Huang, D.; Ng, A.Y.; Potts, C. Learning Word Vectors for Sentiment Analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics (ACL), Portland, OR, USA, 19–24 June 2011; pp. 142–150. [Google Scholar]
- Zhang, X.; Zhao, J.; LeCun, Y. Character-level Convolutional Networks for Text Classification. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montreal, QC, Canada, 7–12 December 2015. [Google Scholar]
- Hou, Y.; Li, J.; He, Z.; Yan, A.; Chen, X.; McAuley, J. Bridging Language and Items for Retrieval and Recommendation. arXiv 2024, arXiv:2403.03952. [Google Scholar]
- Williams, A.; Nangia, N.; Bowman, S.R. A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT), New Orleans, LA, USA, 1–6 June 2018; pp. 1112–1122. [Google Scholar]
- Krizhevsky, A. Learning Multiple Layers of Features from Tiny Images; Technical Report; University of Toronto: Toronto, ON, Canada, 2009. [Google Scholar]
- Hendrycks, D.; Dietterich, T. Benchmarking Neural Network Robustness to Common Corruptions and Perturbations. In Proceedings of the International Conference on Learning Representations (ICLR), New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Coates, A.; Ng, A.; Lee, H. An Analysis of Single-Layer Networks in Unsupervised Feature Learning. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics (AISTATS), Fort Lauderdale, FL, USA, 11–13 April 2011; pp. 215–223. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning Transferable Visual Models from Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML), Virtual, 18–24 July 2021; Volume 139, pp. 8748–8763. [Google Scholar]
- Hodosh, M.; Young, P.; Hockenmaier, J. Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics. J. Artif. Intell. Res. 2013, 47, 853–899. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; Stoyanov, V. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv 2019, arXiv:1907.11692. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 26 June–1 July 2016; pp. 770–778. [Google Scholar]
- Liu, F.; Xu, W.; Lu, J.; Zhang, G.; Gretton, A.; Sutherland, D.J. Learning Deep Kernels for Non-Parametric Two-Sample Tests. In Proceedings of the International Conference on Machine Learning (ICML), Virtual, 13–18 July 2020. [Google Scholar]

| Benchmark/Suite | Domain | Ref.- Free | Calib. FPR | Signif. | Delay | Typed/ Graded | Multi- Modal | Released |
|---|---|---|---|---|---|---|---|---|
| OpenOOD [32] | vision OOD | ✓ | ∼ | × | × | ∼ | × | ✓ |
| WILDS [33] | in-the-wild | × | × | × | × | × | ∼ | ✓ |
| Failing-Loudly [5] | shift detect. | ✓ | × | × | × | ∼ | × | ∼ |
| Conf. martingale [34] | sequential CP | ✓ | ✓ | ∼ | ✓ | × | × | × |
| DriftBench [8] | DB workload | × | × | × | × | ✓ | × | ✓ |
| NeurBench [35] | DB components | × | × | × | × | ✓ | × | ✓ |
| Drift-Bench [36] | agent dialogue | × | × | × | × | ✓ | × | ✓ |
| DecayBench (ours) | NLP/vision/MM | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Dataset | Modality/Domain | Role in DecayBench | Source |
|---|---|---|---|
| SST-2 | NLP, sentiment (GLUE) | Primary task; graded synthetic drift | [1] |
| CoLA | NLP, acceptability (GLUE) | Primary task; graded synthetic drift | [2] |
| MNLI | NLP, 3-class NLI | Task/label-space transfer | [49] |
| IMDB | NLP, sentiment | Natural domain shift | [46] |
| Yelp | NLP, sentiment | Natural domain shift | [47] |
| Amazon | NLP, sentiment (by year) | Natural/temporal domain shift | [48] |
| CIFAR-10-C | Vision, image | Modality transfer; corruption drift | [50,51] |
| STL-10 | Vision, image | Modality transfer | [52] |
| CIFAR-10 (CLIP pairs) | Multimodal, image-caption | Cross-modal fusion drift | [53] |
| Flickr8k | Multimodal, real pairs | Cross-modal, real-pair check | [54] |
| Detector | Family | Reads | Trained | Source |
|---|---|---|---|---|
| MSP | confidence | softmax | no | [13] |
| ODIN | confidence | logits | no | [14] |
| Energy | confidence | logits | no | [15] |
| MMD | distributional | embeddings | no | [20] |
| BBSD | distributional | softmax | no | [5] |
| Mahalanobis | feature-density | embeddings | no | [19] |
| KNN | feature-density | embeddings | no | [18] |
| DriftLens | per-class emb. | embeddings | no | [23] |
| Cosine baseline | industrial | embeddings | no | SageMaker-style |
| Recon | reconstruction | embeddings | yes | [24] |
| Task | Recon | Cosine | MMD | Mahalanobis | BBSD | Energy | MSP | |
|---|---|---|---|---|---|---|---|---|
| SST-2 | ||||||||
| CoLA | ||||||||
| SST-2 | Median | Fisher | GLRT | Bonferroni | Simes | Alert (Mean-z, Ours) |
|---|---|---|---|---|---|---|
| Task | Detector | Precision | Recall | F1 |
|---|---|---|---|---|
| SST-2 | Recon | |||
| MMD | ||||
| Energy | ||||
| Alert (ours) | ||||
| CoLA | Recon | |||
| MMD | ||||
| Energy | ||||
| Alert (ours) |
| Task | Clean | |||||||
|---|---|---|---|---|---|---|---|---|
| SST-2 | ||||||||
| CoLA |
| Detector | SST-2 | CoLA | ||
|---|---|---|---|---|
| Recon | ||||
| MMD | ||||
| Energy | ||||
| MSP | ||||
| Mahalanobis | ||||
| SageMaker (cosine) | ||||
| Alert (ours) | 2.00 ± 0.00 | |||
| (a) ROC-AUC versus drift severity (SST-2) | |||||
| Detector | |||||
| Recon | |||||
| MMD | |||||
| Energy | |||||
| Alert (ours) | |||||
| (b) Paired-bootstrap no-regret test (, p) at | |||||
| Task | Alert vs. | Batch size 64 | Batch size 32 | ||
| p | p | ||||
| SST-2 | Recon | ||||
| MMD | |||||
| Energy | |||||
| CoLA | Recon | ||||
| MMD | |||||
| Energy | |||||
| (a) Self-configured Alert (ctx) vs. fixed-family Alert | |||||
| Task | Alert variant (selected family) | ||||
| SST-2 | Alert fixed {Energy, MMD, Recon} | ||||
| Alert ctx top-3 {BBSD, Recon, MMD} | |||||
| Alert ctx top-5 | |||||
| CoLA | Alert fixed {Energy, MMD, Recon} | ||||
| Alert ctx top-3 {MSP, ODIN, Energy} | |||||
| Alert ctx top-5 | |||||
| (b) Where selection helps most: mismatch | |||||
| Mismatch setting | Alert fixed | Alert ctx | |||
| CoLA drift, family chosen on SST-2 | |||||
| SST-2 drift, family chosen on CoLA | |||||
| SST-2, roberta-base backbone (BERT family) | |||||
| OOD Set | Acc. | Recon | MMD | Energy | MSP | Mahalanobis | Cosine | Alert (Ours) |
|---|---|---|---|---|---|---|---|---|
| IMDB | ||||||||
| Yelp |
| (a) NLP transfer: three-class NLI (MNLI) | ||||||
| Detector | ||||||
| Classifier acc. | ||||||
| BBSD | ||||||
| ODIN | ||||||
| Energy | ||||||
| MMD | ||||||
| MSP | ||||||
| DriftLens | ||||||
| Recon | ||||||
| Mahalanobis | ||||||
| Cosine | ||||||
| Alert (ours) | ||||||
| (b) Vision/backbone transfer | ||||||
| Detector | ResNet-18 | CLIP | ||||
| sev 1 | 3 | 5 | sev 1 | 3 | 5 | |
| MMD | ||||||
| MSP | ||||||
| ODIN | ||||||
| Energy | ||||||
| KNN | ||||||
| Mahalanobis | ||||||
| Recon | ||||||
| Alert (ours) | ||||||
| Detector (What It Monitors) | Image Drift | Text Drift | Mixed | Mismatch |
|---|---|---|---|---|
| MMD, image only | ||||
| MMD, text only | ||||
| concat-MMD, both (early fusion) | ||||
| alignment, cross-modal | ||||
| Alert, marginals only {img, txt} | ||||
| Alert + alignment (ours) |
| Combiner/Detector (Features) | Mismatch ROC-AUC |
|---|---|
| image-MMD alone | |
| text-MMD alone | |
| concat-MMD (early fusion of marginals) | |
| alignment alone | |
| Alert mean-z, marginals only {img, txt} | |
| Alert mean-z + alignment (ours) | |
| unstandardised mean + alignment | |
| logistic regression + alignment (supervised) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Xu, J.; Tian, Y. DecayBench: A Reference-Free Benchmark for Trustworthy Drift Detection. Mathematics 2026, 14, 3045. https://doi.org/10.3390/math14173045
Xu J, Tian Y. DecayBench: A Reference-Free Benchmark for Trustworthy Drift Detection. Mathematics. 2026; 14(17):3045. https://doi.org/10.3390/math14173045
Chicago/Turabian StyleXu, Jia, and Yingli Tian. 2026. "DecayBench: A Reference-Free Benchmark for Trustworthy Drift Detection" Mathematics 14, no. 17: 3045. https://doi.org/10.3390/math14173045
APA StyleXu, J., & Tian, Y. (2026). DecayBench: A Reference-Free Benchmark for Trustworthy Drift Detection. Mathematics, 14(17), 3045. https://doi.org/10.3390/math14173045




