Skip to Content
ComputationComputation
  • This is an early access version, the complete PDF, HTML, and XML versions will be available soon.
  • Article
  • Open Access

9 September 2026

Float32-Induced Distortion in Activation Patching: Precision Floors and Displacement-Dependent Endpoint-Curvature Error

Independent Researcher, Herndon, VA 20171, USA

Abstract

Interaction-level activation patching is used to infer nonlinear cooperation in neural networks, but its numerical validity has not been systematically audited. In two pretrained checkpoints—GPT-2-medium and Pythia-410M, audited on one CUDA backend with a single templated cloze task family—we identified two independent failure modes. First, exact finite pair interactions subtract four O(1) forward evaluations while the target scales as O(α2), creating a precision floor. In matched float32/float64 experiments, at a descriptive 3× floor, 13.0% and 51.2% of full-displacement interactions fell below the threshold—across reasonable 2×10× cutoffs, these fractions ranged from 8.6 to 33.0% and from 39.5 to 82.1%, respectively—with sign disagreements of 2.9% and 13.2%; deterministic repetitions showed zero spread and therefore failed to expose the error. Second, even in float64, endpoint cross-Hessian estimates diverged from finite interactions as displacement increased, reaching 41–68% error at the full-replacement scale commonly used in patching; this displacement relationship is fitted on only these two checkpoints, and the Pythia-410M fit is visibly less stable. The corrupted prompts contain neither candidate answer, and the fixture’s behavioral contrast was not serialized at measurement time. A post hoc audit of the exact deposited fixture now confirms the intended contrast (the answer beats the distractor on all clean prompts in both models; the clean-minus-corrupt contrast is positive on 20/20 and 17/20 prompts). Thus, the quoted constants are properties of this behaviorally supported fixture and backend; we expect the audit procedure, not the constants, to transfer. We introduce an inexpensive α2 scaling audit that detects cancellation floors and hidden mixed-precision bottlenecks; it exposed a hardcoded float32 softmax path inside nominal-float64 GPT-NeoX inference. A motivating negative result—that local logical gate topology does not predict attribution-patching error—was obtained under an author-held internal specification that is not independently timestamped and only in ten small synthetic transformers; it is untested at pretrained scale. These results show that reproducibility alone does not establish measurement validity and that precision error and endpoint approximation error must be audited separately. The proposed checks provide a practical validation standard for interaction-level mechanistic interpretability claims.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.