Next Article in Journal
Algorithm Design and Analysis of a Blockchain-Enabled Reinforcement and Active Learning Framework for Robust IoT Intrusion Detection
Previous Article in Journal
An Open-Source Handheld Mobile LiDAR Pipeline for Tree Detection, Stem Reconstruction and Volume Estimation in Pinus pinaster Stands
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
This is an early access version, the complete PDF, HTML, and XML versions will be available soon.
Article

Martingale Doppelgänger-Eval: A Specification-Driven Interventional Audit Algorithm for Visual Evidence Use in Vision–Language Models

Department of Mathematics and Statistics, Texas Tech University, Lubbock, TX 79409, USA
*
Author to whom correspondence should be addressed.
Algorithms 2026, 19(9), 803; https://doi.org/10.3390/a19090803 (registering DOI)
Submission received: 20 August 2026 / Revised: 17 September 2026 / Accepted: 18 September 2026 / Published: 19 September 2026

Abstract

Assessing chart understanding requires distinguishing responses to local visual evidence from associations with a chart’s overall shape. We present Martingale Doppelgänger-Eval, an interventional audit framework that connects executable edit specifications to benchmark generation, validity checks and response estimation. Matched charts support separate measurements of evidence response and regional sensitivity. Each result carries its uncertainty, response coverage and identification status. We characterize conditions for identifying paired effects and specify a sequential procedure for audits that stop after inspecting accumulating evidence. We directly evaluate nine frozen vision–language models on 12,000 pairs generated with the corrected renderer. Complete-pair evidence scores range from 0.4362 to 0.4990, while supervised pixel controls generalize the learned rules to disjoint held-out symbols. A matched volume-only experiment reveals model-specific responses, and magnitude-matched regional edits quantify oracle-region sensitivity separately from the faithfulness of reported regions. The framework links interpretable visual comparisons to the conditions needed for their statistical evaluation.
Keywords: vision–language models; interventional audit; visual evidence; counterfactual evaluation; shortcut learning; anytime-valid inference; candlestick charts vision–language models; interventional audit; visual evidence; counterfactual evaluation; shortcut learning; anytime-valid inference; candlestick charts

Share and Cite

MDPI and ACS Style

Wang, Z.; Rachev, S.T. Martingale Doppelgänger-Eval: A Specification-Driven Interventional Audit Algorithm for Visual Evidence Use in Vision–Language Models. Algorithms 2026, 19, 803. https://doi.org/10.3390/a19090803

AMA Style

Wang Z, Rachev ST. Martingale Doppelgänger-Eval: A Specification-Driven Interventional Audit Algorithm for Visual Evidence Use in Vision–Language Models. Algorithms. 2026; 19(9):803. https://doi.org/10.3390/a19090803

Chicago/Turabian Style

Wang, Ziyao, and Svetlozar T. Rachev. 2026. "Martingale Doppelgänger-Eval: A Specification-Driven Interventional Audit Algorithm for Visual Evidence Use in Vision–Language Models" Algorithms 19, no. 9: 803. https://doi.org/10.3390/a19090803

APA Style

Wang, Z., & Rachev, S. T. (2026). Martingale Doppelgänger-Eval: A Specification-Driven Interventional Audit Algorithm for Visual Evidence Use in Vision–Language Models. Algorithms, 19(9), 803. https://doi.org/10.3390/a19090803

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Article metric data becomes available approximately 24 hours after publication online.
Back to TopTop