A Systematic Benchmark of Quantum Support Vector Machines for Interpretable Attribution of AI-Generated Text
Abstract
1. Introduction
2. Materials and Methods
2.1. Dataset
2.2. Problem Formulation
2.3. Research Questions
- RQ1 (Benchmark and ceiling). What is the optimal hyperparameter configuration of QSVM for binary LLM attribution on a corpus of approximately 5800 samples, where does its accuracy ceiling lie, and how does that ceiling compare with a classical SVM operating on identical features as well as on the full TF-IDF representation?
- RQ2 (Kernel concentration). How does quantum kernel concentration manifest as the feature dimension grows from 4 to 16 qubits, and what data-to-qubit ratio is required to avoid it on a problem of this scale and difficulty? In particular, does increasing circuit repetitions or moving to richer Pauli feature maps mitigate this effect, or do they accelerate it?
- RQ3 (Stylometric fingerprinting). What stylometric fingerprint emerges from the similarity structure induced by the quantum kernel for each LLM, and how does this fingerprint evolve as feature dimension increases? Specifically, do additional qubits expose genuinely new stylistic axes, or do they merely sharpen the ones already visible at low dimension?
- RQ4 (Interpretable feature engineering). Can enriched stylometric feature representations, namely, a hybrid character n-gram plus stylometric pipeline and a direct named-feature encoding in which one qubit corresponds to one stylometric feature, surpass the accuracy ceiling of the PCA baseline while preserving, or even sharpening, the interpretability of quantum attribution?
2.4. Feature Extraction Pipeline (Baseline)
2.5. QSVM Architecture
2.6. Classical SVM Baseline
2.7. Evaluation Protocol
3. Results
3.1. Stage 1: Hyperparameter Benchmark
3.1.1. Accuracy Ceiling and Dual Operating Points
3.1.2. Quantum Kernel Concentration and the Data-to-Qubit Ratio
3.1.3. Feature-Map and Circuit-Depth Ablations
3.1.4. Per-Class Behavior at the Joint-Best Configuration
3.2. Stylometric Fingerprint Analysis
3.2.1. Progressive Feature Discovery Across Dimensions
3.2.2. Complete LLM Stylometric Profiles
3.2.3. Classical Versus Quantum Feature Attribution
3.3. Stage 2: Feature Engineering for Quantum Stylometry
3.3.1. Pipeline Design
3.3.2. Extended Stylometric Features
3.3.3. Stage 2 Results
3.3.4. Stylometric Insights from Direct Encoding
3.4. Independent Test-Set Evaluation
3.5. Open-Set Quantum Stylometric Fingerprinting
3.5.1. Quantum Kernel Proximity Scoring
3.5.2. Multi-Model Attribution Workflow
3.5.3. Robustness Considerations
4. Discussion
4.1. The Role of Quantum Kernels in NLP
4.2. Concentration and the Data Bottleneck
4.3. Limitations
4.4. Future Directions
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| CNOT | Controlled-NOT Gate |
| CPU | Central Processing Unit |
| CSVM | Classical Support Vector Machine |
| CV | Cross-Validation |
| F1 | F1 Score (Harmonic Mean of Precision and Recall) |
| GPU | Graphics Processing Unit |
| HPC | High-Performance Computing |
| JSON | JavaScript Object Notation |
| LIME | Local Interpretable Model-Agnostic Explanations |
| LLM | Large Language Model |
| MI | Mutual Information |
| NISQ | Noisy Intermediate-Scale Quantum |
| NLP | Natural Language Processing |
| NLTK | Natural Language Toolkit |
| PCA | Principal Component Analysis |
| POS | Part of Speech |
| pp | Percentage Points |
| QML | Quantum Machine Learning |
| QNLP | Quantum Natural Language Processing |
| QSVC | Quantum Support Vector Classifier |
| QSVM | Quantum Support Vector Machine |
| RAM | Random-Access Memory |
| RBF | Radial Basis Function |
| RQ | Research Question |
| SHAP | Shapley Additive Explanations |
| SVC | Support Vector Classifier |
| SVM | Support Vector Machine |
| TF-IDF | Term Frequency–Inverse Document Frequency |
Appendix A
Appendix A.1. Hardware and Software Environment
| Component | Specification |
|---|---|
| Compute cluster | SLURM HPC, CPU partition |
| Node CPU | 64 physical/64 logical cores |
| Node memory | 248 to 251 GB RAM |
| Python | 3.12.2 (conda-forge) |
| Qiskit | 1.4.3 |
| Qiskit Machine Learning | 0.8.3 |
| Qiskit Aer | 0.17.1 |
| scikit-learn | 1.7.1 |
| NumPy | 1.26.4 (2.2.6 for the replication re-runs; see the replication note below) |
| Simulation backend | FidelityStatevectorKernel (exact, noiseless) |
| Default entanglement | Linear (explicitly set; the library default is ‘full’) |
| Angle normalization | [0, 2], fitted on training data only |
Appendix A.2. Computational Cost Across Experiments
| Experiment | Runs | QKern Time | Pipeline Time |
|---|---|---|---|
| Experiment 1.1 ( to 8, small) | 36 | ∼31 min | ∼1 h 34 min |
| Experiment 1.2 (, larger data) | 4 | ∼11 min | ∼22 min |
| Experiment 1.3 ( to 12, medium) | 19 | ∼1 h 17 min | ∼2 h 09 min |
| Experiment 1.4 (Pauli + reps = 3 ablations) | 9 | ∼28 min | ∼49 min |
| Experiment 1.5 (, large-scale) | 5 | ∼2 h 21 min | ∼2 h 55 min |
| Experiment 1.6 ( to 16, large-scale) | 10 | ∼56 h 02 min | ∼64 h 11 min |
| Total | 83 | ∼60 h 50 min | ∼72 h 00 min |
Appendix B
Appendix B.1. Experiment 1.1: Dimensions 4 to 8, Small-Scale Grid
| Model Tag | d | Train | Val | C | QSVM% | CSVM% | CV% | CV± | QKern | Pipeline |
|---|---|---|---|---|---|---|---|---|---|---|
| 4-600-400-C1 | 4 | 600 | 400 | 1 | 79.75 | 97.75 | 73.50 | ±3.82 | 33 s | 2 min |
| 4-600-400-C10 | 4 | 600 | 400 | 10 | 76.75 | 97.75 | 74.17 | ±4.31 | 33 s | 2 min |
| 4-600-400-C50 | 4 | 600 | 400 | 50 | 78.00 | 97.75 | 70.00 | ±4.80 | 33 s | 2 min |
| 4-600-400-C100 | 4 | 600 | 400 | 100 | 74.00 | 97.75 | 70.50 | ±7.31 | 33 s | 2 min |
| 4-700-300-C1 | 4 | 700 | 300 | 1 | 78.33 | 97.67 | 74.43 | ±3.48 | 41 s | 2 min |
| 4-700-300-C10 | 4 | 700 | 300 | 10 | 74.67 | 97.67 | 73.71 | ±3.54 | 42 s | 2 min |
| 4-700-300-C50 | 4 | 700 | 300 | 50 | 73.33 | 97.67 | 71.57 | ±3.39 | 40 s | 2 min |
| 4-700-300-C100 | 4 | 700 | 300 | 100 | 72.67 | 97.67 | 70.29 | ±4.03 | 41 s | 3 min |
| 4-800-200-C1 | 4 | 800 | 200 | 1 | 77.50 | 97.50 | 76.75 | ±1.95 | 49 s | 2 min |
| 4-800-200-C10 | 4 | 800 | 200 | 10 | 73.50 | 97.50 | 77.25 | ±2.58 | 48 s | 2 min |
| 4-800-200-C50 | 4 | 800 | 200 | 50 | 71.50 | 97.50 | 75.88 | ±4.45 | 51 s | 2 min |
| 4-800-200-C100 | 4 | 800 | 200 | 100 | 66.50 | 97.50 | 74.38 | ±3.47 | 51 s | 2 min |
| 6-600-400-C1 | 6 | 600 | 400 | 1 | 84.25 | 97.75 | 78.00 | ±2.67 | 41 s | 2 min |
| 6-600-400-C10 | 6 | 600 | 400 | 10 | 80.00 | 97.75 | 75.33 | ±1.72 | 42 s | 3 min |
| 6-600-400-C50 | 6 | 600 | 400 | 50 | 80.50 | 97.75 | 75.50 | ±1.13 | 42 s | 3 min |
| 6-600-400-C100 | 6 | 600 | 400 | 100 | 80.50 | 97.75 | 75.50 | ±1.13 | 40 s | 3 min |
| 6-700-300-C1 | 6 | 700 | 300 | 1 | 82.00 | 97.67 | 75.00 | ±2.52 | 51 s | 2 min |
| 6-700-300-C10 | 6 | 700 | 300 | 10 | 78.67 | 97.67 | 73.43 | ±2.19 | 53 s | 3 min |
| 6-700-300-C50 | 6 | 700 | 300 | 50 | 78.67 | 97.67 | 73.43 | ±2.19 | 50 s | 3 min |
| 6-700-300-C100 | 6 | 700 | 300 | 100 | 78.67 | 97.67 | 73.43 | ±2.19 | 52 s | 3 min |
| 6-800-200-C1 | 6 | 800 | 200 | 1 | 83.50 | 97.50 | 80.25 | ±3.50 | 1 min | 3 min |
| 6-800-200-C10 | 6 | 800 | 200 | 10 | 82.00 | 97.50 | 76.25 | ±2.05 | 1 min | 3 min |
| 6-800-200-C50 | 6 | 800 | 200 | 50 | 82.00 | 97.50 | 76.62 | ±1.79 | 1 min | 3 min |
| 6-800-200-C100 | 6 | 800 | 200 | 100 | 82.00 | 97.50 | 76.62 | ±1.79 | 59 s | 3 min |
| 8-600-400-C1 | 8 | 600 | 400 | 1 | 88.00 | 97.75 | 77.67 | ±2.91 | 50 s | 3 min |
| 8-600-400-C10 | 8 | 600 | 400 | 10 | 87.50 | 97.75 | 79.00 | ±2.20 | 52 s | 3 min |
| 8-600-400-C50 | 8 | 600 | 400 | 50 | 87.50 | 97.75 | 79.00 | ±2.20 | 51 s | 3 min |
| 8-600-400-C100 | 8 | 600 | 400 | 100 | 87.50 | 97.75 | 79.00 | ±2.20 | 51 s | 3 min |
| 8-700-300-C1 | 8 | 700 | 300 | 1 | 81.00 | 97.67 | 72.57 | ±3.43 | 1 min | 3 min |
| 8-700-300-C10 | 8 | 700 | 300 | 10 | 82.00 | 97.67 | 73.57 | ±3.75 | 1 min | 3 min |
| 8-700-300-C50 | 8 | 700 | 300 | 50 | 82.00 | 97.67 | 73.57 | ±3.75 | 1 min | 3 min |
| 8-700-300-C100 | 8 | 700 | 300 | 100 | 82.00 | 97.67 | 73.57 | ±3.75 | 1 min | 3 min |
| 8-800-200-C1 | 8 | 800 | 200 | 1 | 86.00 | 97.50 | 79.12 | ±2.76 | 1 min | 3 min |
| 8-800-200-C10 | 8 | 800 | 200 | 10 | 87.50 | 97.50 | 79.38 | ±1.81 | 1 min | 3 min |
| 8-800-200-C50 | 8 | 800 | 200 | 50 | 87.50 | 97.50 | 79.38 | ±1.81 | 1 min | 3 min |
| 8-800-200-C100 | 8 | 800 | 200 | 100 | 87.50 | 97.50 | 79.38 | ±1.81 | 1 min | 3 min |
Appendix B.2. Experiment 1.2: Concentration Onset
| Model Tag | d | Train | Val | C | QSVM% | CSVM% | CV% | CV± | QKern | Pipeline |
|---|---|---|---|---|---|---|---|---|---|---|
| 8-1200-800-C1 | 8 | 1200 | 800 | 1 | 75.00 | 96.88 | 72.17 | ±0.49 | 2 min | 5 min |
| 8-1200-800-C10 | 8 | 1200 | 800 | 10 | 73.25 | 96.88 | 72.00 | ±1.25 | 2 min | 5 min |
| 8-1400-600-C1 | 8 | 1400 | 600 | 1 | 77.33 | 97.33 | 76.86 | ±3.03 | 3 min | 6 min |
| 8-1400-600-C10 | 8 | 1400 | 600 | 10 | 74.33 | 97.33 | 76.43 | ±3.33 | 3 min | 6 min |
Appendix B.3. Experiment 1.3: Dimension 10 and 12 Sweeps
| Model Tag | d | Train | Val | C | QSVM% | CSVM% | CV% | CV± | QKern | Pipeline |
|---|---|---|---|---|---|---|---|---|---|---|
| 10-600-400-C1 | 10 | 600 | 400 | 1 | 86.75 | 97.75 | 76.17 | ±2.96 | 1 min | 3 min |
| 10-600-400-C10 | 10 | 600 | 400 | 10 | 86.50 | 97.75 | 76.17 | ±2.21 | 1 min | 3 min |
| 10-800-200-C1 | 10 | 800 | 200 | 1 | 84.00 | 97.50 | 77.75 | ±2.52 | 2 min | 3 min |
| 10-800-200-C10 | 10 | 800 | 200 | 10 | 83.00 | 97.50 | 78.12 | ±2.40 | 2 min | 3 min |
| 10-1200-800-C1 | 10 | 1200 | 800 | 1 | 79.50 | 96.88 | 76.17 | ±2.48 | 3 min | 6 min |
| 10-1200-800-C10 | 10 | 1200 | 800 | 10 | 79.25 | 96.88 | 75.67 | ±2.86 | 3 min | 6 min |
| 10-1400-600-C1 | 10 | 1400 | 600 | 1 | 81.33 | 97.33 | 76.57 | ±2.23 | 4 min | 7 min |
| 10-1400-600-C10 | 10 | 1400 | 600 | 10 | 79.83 | 97.33 | 76.21 | ±2.09 | 4 min | 7 min |
| 10-1600-600-C1 | 10 | 1600 | 600 | 1 | 83.50 | 98.00 | 78.88 | ±1.51 | 5 min | 8 min |
| 10-2000-600-C1 | 10 | 2000 | 600 | 1 | 82.67 | 97.00 | 77.85 | ±1.83 | 7 min | 11 min |
| 10-2000-800-C1 | 10 | 2000 | 800 | 1 | 80.25 | 97.50 | 74.55 | ±1.58 | 7 min | 11 min |
| 12-600-400-C1 | 12 | 600 | 400 | 1 | 84.50 | 97.75 | 74.00 | ±2.00 | 2 min | 4 min |
| 12-600-400-C10 | 12 | 600 | 400 | 10 | 85.00 | 97.75 | 74.33 | ±1.86 | 2 min | 4 min |
| 12-800-200-C1 | 12 | 800 | 200 | 1 | 84.00 | 97.50 | 76.00 | ±1.09 | 3 min | 5 min |
| 12-800-200-C10 | 12 | 800 | 200 | 10 | 83.50 | 97.50 | 76.38 | ±1.21 | 3 min | 5 min |
| 12-1200-800-C1 | 12 | 1200 | 800 | 1 | 79.12 | 96.88 | 74.25 | ±2.29 | 6 min | 9 min |
| 12-1200-800-C10 | 12 | 1200 | 800 | 10 | 79.25 | 96.88 | 74.17 | ±2.62 | 6 min | 9 min |
| 12-1400-600-C1 | 12 | 1400 | 600 | 1 | 84.00 | 97.33 | 77.50 | ±1.71 | 8 min | 11 min |
| 12-1400-600-C10 | 12 | 1400 | 600 | 10 | 83.33 | 97.33 | 77.79 | ±2.34 | 8 min | 11 min |
Appendix B.4. Experiment 1.4: PauliFeatureMap Comparison
| Model Tag | d | Train | Val | Map | QSVM% | CSVM% | CV% | CV± | QKern | Pipeline |
|---|---|---|---|---|---|---|---|---|---|---|
| 8-600-400-ZZ | 8 | 600 | 400 | ZZFeatureMap | 88.00 | 97.75 | 77.67 | ±2.91 | 50 s | 3 min |
| 8-600-400-PauliXX | 8 | 600 | 400 | Pauli_XX | 84.75 | 97.75 | 77.33 | ±2.76 | 1 min | 3 min |
| 8-600-400-PauliZZXX | 8 | 600 | 400 | Pauli_ZZ_XX | 77.50 | 97.75 | 69.00 | ±2.95 | 1 min | 3 min |
| 8-600-400-Paulifull | 8 | 600 | 400 | Pauli_full | 65.25 | 97.75 | 63.00 | ±3.23 | 4 min | 7 min |
| 8-800-200-ZZ | 8 | 800 | 200 | ZZFeatureMap | 86.00 | 97.50 | 79.12 | ±2.76 | 1 min | 3 min |
| 8-800-200-PauliXX | 8 | 800 | 200 | Pauli_XX | 87.00 | 97.50 | 79.25 | ±2.42 | 2 min | 3 min |
| 8-800-200-PauliZZXX | 8 | 800 | 200 | Pauli_ZZ_XX | 79.00 | 97.50 | 76.88 | ±1.98 | 2 min | 4 min |
| 8-800-200-Paulifull | 8 | 800 | 200 | Pauli_full | 66.00 | 97.50 | 66.88 | ±1.85 | 5 min | 8 min |
Appendix B.5. Experiment 1.4 (Continued): Circuit-Depth Ablation (reps = 2 Versus reps = 3)
| Model Tag | d | Train | Val | reps | QSVM% | CV% | CV± | QKern | Δ vs. reps = 2 |
|---|---|---|---|---|---|---|---|---|---|
| 8-600-400-C1-r2 | 8 | 600 | 400 | 2 | 88.00 | 77.67 | ±2.91 | 50 s | baseline |
| 8-600-400-C1-r3 | 8 | 600 | 400 | 3 | 78.75 | 69.67 | ±2.96 | 1 min | −9.25 pp |
| 8-1200-800-C1-r2 | 8 | 1200 | 800 | 2 | 75.00 | 72.17 | ±0.49 | 2 min | baseline |
| 8-1200-800-C1-r3 | 8 | 1200 | 800 | 3 | 70.00 | 66.00 | ±2.94 | 3 min | −5.00 pp |
| 12-1400-600-C1-r2 | 12 | 1400 | 600 | 2 | 84.00 | 77.50 | ±1.71 | 8 min | baseline |
| 12-1400-600-C1-r3 | 12 | 1400 | 600 | 3 | 78.67 | 72.93 | ±1.87 | 8 min | −5.33 pp |
Appendix B.6. Experiment 1.5: Scaling Runs
| Model Tag | d | Train | Val | QSVM% | CSVM% | CV% | CV± | QKern | Pipeline |
|---|---|---|---|---|---|---|---|---|---|
| 12-1600-600-C1 | 12 | 1600 | 600 | 87.50 | 98.00 | 78.31 | ±1.35 | 9 min | 13 min |
| 12-2000-600-C1 | 12 | 2000 | 600 | 82.50 | 97.00 | 75.95 | ±1.93 | 13 min | 18 min |
| 12-2000-800-C1 | 12 | 2000 | 800 | 86.38 | 97.50 | 78.20 | ±3.01 | 14 min | 19 min |
| 12-3480-1160-C1 | 12 | 3480 | 1160 | 88.02 | 98.02 | 83.19 | ±0.75 | 39 min | 48 min |
| 12-4640-1160-C1 | 12 | 4640 | 1160 | 86.64 | 98.28 | 84.01 | ±0.83 | 1 h 05 min | 1 h 17 min |
Appendix B.7. Experiment 1.6: Dimension 14 and 16 Scaling Runs
| Model Tag | d | Train | Val | QSVM% | CSVM% | CV% | CV± | QKern | Pipeline |
|---|---|---|---|---|---|---|---|---|---|
| 14-2400-1160-C1 | 14 | 2400 | 1160 | 79.91 | 97.16 | 75.88 | ±2.19 | 1 h 30 min | 1 h 47 min |
| 14-2800-1160-C1 | 14 | 2800 | 1160 | 83.28 | 97.33 | 78.46 | ±1.43 | 1 h 55 min | 2 h 13 min |
| 14-3200-800-C1 | 14 | 3200 | 800 | 87.62 | 98.25 | 82.94 | ±0.91 | 2 h 20 min | 2 h 40 min |
| 14-3480-1160-C1 | 14 | 3480 | 1160 | 88.02 | 98.02 | 83.05 | ±0.84 | 2 h 24 min | 2 h 47 min |
| 14-4640-1160-C1 | 14 | 4640 | 1160 | 84.91 | 98.28 | 81.70 | ±1.35 | 4 h 05 min | 4 h 41 min |
| 16-2400-1160-C1 | 16 | 2400 | 1160 | 78.62 | 97.16 | 73.00 | ±2.38 | 4 h 53 min | 5 h 43 min |
| 16-2800-1160-C1 | 16 | 2800 | 1160 | 80.78 | 97.33 | 73.96 | ±2.04 | 6 h 14 min | 7 h 09 min |
| 16-3200-800-C1 | 16 | 3200 | 800 | 84.88 | 98.25 | 80.25 | ±1.74 | 8 h 13 min | 9 h 19 min |
| 16-3480-1160-C1 | 16 | 3480 | 1160 | 86.72 | 98.02 | 80.98 | ±0.84 | 8 h 39 min | 9 h 55 min |
| 16-4640-1160-C1 | 16 | 4640 | 1160 | 82.16 | 98.28 | 78.92 | ±1.34 | 15 h 48 min | 17 h 56 min |
Appendix B.8. Same-Split Feature-Importance Comparability Check
Appendix B.9. Measured Kernel Concentration Statistics
| Kernel | Off-Diagonal Mean | Off-Diagonal Variance |
|---|---|---|
| Quantum ZZ, | ||
| Quantum ZZ, | ||
| Quantum ZZ, | ||
| Quantum ZZ, | ||
| Classical RBF, | 0.2145 |
Appendix C
| Pipeline | reps | QSVM Val. Acc. (%) | CSVM Val. Acc. (%) | |
|---|---|---|---|---|
| hybrid | 1 | 0.1 | 57.41 | 89.74 |
| hybrid | 1 | 1 | 75.17 | 89.74 |
| hybrid | 1 | 10 | 76.90 | 89.74 |
| hybrid | 2 | 0.1 | 53.02 | 89.74 |
| hybrid | 2 | 1 | 76.55 | 89.74 |
| hybrid | 2 | 10 | 76.98 | 89.74 |
| stylometric_direct | 1 | 0.1 | 55.95 | 90.86 |
| stylometric_direct | 1 | 1 | 77.16 | 90.86 |
| stylometric_direct | 1 | 10 | 78.62 | 90.86 |
| stylometric_direct | 2 | 0.1 | 60.00 | 90.86 |
| stylometric_direct | 2 | 1 | 78.19 | 90.86 |
| stylometric_direct | 2 | 10 | 79.14 | 90.86 |
Appendix D
| Feature | Category | Definition |
|---|---|---|
| avg_word_len | Lexical | Mean number of characters per token |
| avg_sent_len | Lexical | Mean number of tokens per sentence |
| unique_ratio | Lexical | Number of unique tokens divided by total tokens (type–token ratio) |
| vocab_richness | Lexical | Guiraud’s richness index: unique tokens divided by the square root of total tokens |
| word_count | Lexical | Total number of tokens in the text |
| char_count | Lexical | Total number of characters in the text |
| noun_ratio | Syntactic | Nouns divided by total POS-tagged tokens |
| verb_ratio | Syntactic | Verbs divided by total POS-tagged tokens |
| adj_ratio | Syntactic | Adjectives divided by total POS-tagged tokens |
| adv_ratio | Syntactic | Adverbs divided by total POS-tagged tokens |
| punct_density | Syntactic | Punctuation characters divided by total characters |
| syntactic_complexity | Syntactic | Mean tokens per sentence, punctuation included, divided by 20 |
| formality | Stylistic | Formal connectives (e.g., therefore, furthermore) divided by formal plus informal markers (e.g., yeah, okay) from fixed keyword lists; a sentence-length proxy is used when neither occurs |
| question_ratio | Stylistic | Question marks divided by sentence count |
| exclaim_ratio | Stylistic | Exclamation marks divided by sentence count |
| capital_ratio | Stylistic | Capitalized characters divided by total characters |
| contraction_ratio | Stylistic | Contraction markers (n’t, ’re, ’ve, ’ll, ’d, ’s, ’m) divided by total tokens |
| list_marker_density | Stylistic | Line-initial list markers (bullets, enumerations) divided by sentence count |
| digit_ratio | Extended | Digit characters divided by total characters |
| sent_len_std | Extended | Standard deviation of sentence lengths in tokens, divided by 20 |
| hedge_ratio | Extended | Distinct hedging phrases present, from a fixed 24-phrase list (e.g., however, I believe, typically), divided by sentence count |
| paragraph_count_ratio | Extended | Paragraphs divided by sentence count |
| repeated_bigram_ratio | Extended | Repeated token bigrams divided by total bigrams |
| avg_word_len_std | Extended | Standard deviation of token lengths in characters, divided by 5 |
References
- Gehrmann, S.; Strobelt, H.; Rush, A. GLTR: Statistical detection and visualization of generated text. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Florence, Italy, 28 July–2 August 2019; pp. 111–116. [Google Scholar] [CrossRef] [Scilit]
- Mitchell, E.; Lee, Y.; Khazatsky, A.; Manning, C.D.; Finn, C. DetectGPT: Zero-shot machine-generated text detection using probability curvature. In Proceedings of the 40th International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July 2023; Volume PMLR 202, pp. 24950–24962. Available online: https://proceedings.mlr.press/v202/mitchell23a.html (accessed on 7 August 2026).
- Miralles-González, P.; Huertas-Tato, J.; Martín, A.; Camacho, D. Not all tokens are created equal: Perplexity attention weighted networks for AI-generated text detection. Inf. Fusion 2026, 125, 103465. [Google Scholar] [CrossRef] [Scilit]
- Crothers, E.N.; Japkowicz, N.; Viktor, H.L. Machine-generated text: A comprehensive survey of threat models and detection methods. IEEE Access 2023, 11, 70977–71002. [Google Scholar] [CrossRef] [Scilit]
- Fariello, S.; Fenza, G.; Forte, F.; Gallo, M.; Marotta, M. Distinguishing human from machine: A review of advances and challenges in AI-generated text detection. Int. J. Interact. Multimed. Artif. Intell. 2025, 9, 6–18. [Google Scholar] [CrossRef] [Scilit]
- Bahad, S.; Bhaskar, Y.; Krishnamurthy, P. NootNoot at SemEval-2024 Task 8: Fine-tuning language models for AI vs human generated text detection. In Proceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024), Mexico City, Mexico, 20–21 June 2024; pp. 918–921. [Google Scholar] [CrossRef] [Scilit]
- Wu, J.; Yang, S.; Zhan, R.; Yuan, Y.; Chao, L.S.; Wong, D.F. A survey on LLM-generated text detection: Necessity, methods, and future directions. Comput. Linguist. 2025, 51, 275–338. [Google Scholar] [CrossRef] [Scilit]
- Mosteller, F.; Wallace, D.L. Inference in an authorship problem: A comparative study of discrimination methods applied to the authorship of the disputed Federalist Papers. J. Am. Stat. Assoc. 1963, 58, 275–309. [Google Scholar] [CrossRef] [Scilit]
- Jeong, S.W.; Ročková, V. From small to large language models: Revisiting the Federalist Papers. arXiv 2025, arXiv:2503.01869. [Google Scholar] [CrossRef] [Scilit]
- Stamatatos, E. A survey of modern authorship attribution methods. J. Am. Soc. Inf. Sci. Technol. 2009, 60, 538–556. [Google Scholar] [CrossRef] [Scilit]
- Gollub, T.; Potthast, M.; Beyer, A.; Busse, M.; Rangel, F.; Rosso, P.; Stamatatos, E.; Stein, B. Recent trends in digital text forensics and its evaluation. In Information Access Evaluation. Multilinguality, Multimodality, and Visualization; Forner, P., Müller, H., Paredes, R., Rosso, P., Stein, B., Eds.; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2013; Volume 8138, pp. 282–302. [Google Scholar] [CrossRef] [Scilit]
- Rosso, P.; Potthast, M.; Stein, B.; Stamatatos, E.; Rangel, F.; Daelemans, W. Evolution of the PAN Lab on digital text forensics. In Information Retrieval Evaluation in a Changing World; Ferro, N., Peters, C., Eds.; The Information Retrieval Series; Springer: Cham, Switzerland, 2019; Volume 41, pp. 461–485. [Google Scholar] [CrossRef] [Scilit]
- Sidorov, G.; Velasquez, F.; Stamatatos, E.; Gelbukh, A.; Chanona-Hernández, L. Syntactic dependency-based n-grams: More evidence of usefulness in classification. In Computational Linguistics and Intelligent Text Processing; Gelbukh, A., Ed.; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2013; Volume 7816, pp. 13–24. [Google Scholar] [CrossRef] [Scilit]
- Huang, B.; Chen, C.; Shu, K. Authorship attribution in the era of LLMs: Problems, methodologies, and challenges. ACM SIGKDD Explor. Newsl. 2025, 26, 21–43. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bitton, Y.; Bitton, E.; Nisan, S. Detecting stylistic fingerprints of large language models. arXiv 2025, arXiv:2503.01659. [Google Scholar] [CrossRef] [Scilit]
- Havlíček, V.; Córcoles, A.D.; Temme, K.; Harrow, A.W.; Kandala, A.; Chow, J.M.; Gambetta, J.M. Supervised learning with quantum-enhanced feature spaces. Nature 2019, 567, 209–212. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Thanasilp, S.; Wang, S.; Cerezo, M.; Holmes, Z. Exponential concentration in quantum kernel methods. Nat. Commun. 2024, 15, 5200. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Schuld, M.; Killoran, N. Quantum machine learning in feature Hilbert spaces. Phys. Rev. Lett. 2019, 122, 040504. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, Y.; Arunachalam, S.; Temme, K. A rigorous and robust quantum speed-up in supervised machine learning. Nat. Phys. 2021, 17, 1013–1017. [Google Scholar] [CrossRef] [Scilit]
- Lorenz, R.; Pearson, A.; Meichanetzidis, K.; Kartsaklis, D.; Coecke, B. QNLP in practice: Running compositional models of meaning on a quantum computer. J. Artif. Intell. Res. 2023, 76, 1305–1342. [Google Scholar] [CrossRef] [Scilit]
- Peral-García, D.; Cruz-Benito, J.; García-Peñalvo, F.J. Using quantum natural language processing for sentiment classification and next-word prediction in sentences without fixed syntactic structure. In Information and Software Technologies; Lopata, A., Gudonienė, D., Butkienė, R., Eds.; Communications in Computer and Information Science; Springer: Cham, Switzerland, 2024; Volume 1979, pp. 235–243. [Google Scholar] [CrossRef] [Scilit]
- Peral-García, D.; Cruz-Benito, J.; García-Peñalvo, F.J. Comparing natural language processing and quantum natural processing approaches in text classification tasks. Expert Syst. Appl. 2024, 254, 124427. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Song, D.; Zhang, P.; Wang, P.; Li, J.; Li, X.; Wang, B. A quantum-inspired multimodal sentiment analysis framework. Theor. Comput. Sci. 2018, 752, 21–40. [Google Scholar] [CrossRef] [Scilit]
- Widdows, D.; Aboumrad, W.; Kim, D.; Ray, S.; Mei, J. Quantum natural language processing. KI—Künstl. Intell. 2024, 38, 293–310. [Google Scholar] [CrossRef] [Scilit]
- Bowles, J.; Ahmed, S.; Schuld, M. Better than classical? The subtle art of benchmarking quantum machine learning models. arXiv 2024, arXiv:2403.07059. [Google Scholar] [CrossRef] [Scilit]
- Dugan, L.; Hwang, A.; Trhlík, F.; Ludan, J.M.; Zhu, A.; Xu, H.; Ippolito, D.; Callison-Burch, C. RAID: A shared benchmark for robust evaluation of machine-generated text detectors. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, 11–16 August 2024; pp. 12463–12492. [Google Scholar] [CrossRef] [Scilit]
- Gemma Team, Google DeepMind. Gemma 3 Technical Report. arXiv 2025, arXiv:2503.19786. [Google Scholar] [CrossRef] [Scilit]
- Qwen Team. Qwen2.5: A Party of Foundation Models! September 2024. Available online: https://qwen.ai/blog?id=qwen2.5 (accessed on 10 May 2026).







| Split | Training Used | Validation | % of Train Pool Used |
|---|---|---|---|
| Full 80/20 | 4640 | 1160 | 100% |
| Large 75/25-equivalent (Experiments 1.5–1.6 at 3480 samples and all Stage 2 experiments) | 3480 | 1160 | 75% |
| Medium 70/30-equivalent | 2800 | 1160 | 60% |
| Small- and medium-scale grids (Experiments 1.1–1.3 and intermediate 1.5 runs) | 600 to 2000 | 200 to 800 | 13 to 43% (hyperparameter search) |
| Category | Features |
|---|---|
| Lexical | avg_word_len, avg_sent_len, unique_ratio, vocab_richness, word_count, char_count |
| Syntactic | noun_ratio, verb_ratio, adj_ratio, adv_ratio, punct_density, syntactic_complexity |
| Stylistic | formality, question_ratio, exclaim_ratio, capital_ratio, contraction_ratio, list_marker_density |
| d | Best QSVM | Best Config | CV Mean | CV Std |
|---|---|---|---|---|
| 4 | 79.75% | 600/400, | 73.50% | ±3.82% |
| 6 | 84.25% | 600/400, | 78.00% | ±2.67% |
| 8 | 88.00% | 600/400, | 77.67% | ±2.91% |
| 10 | 86.75% | 600/400, | 76.17% | ±2.96% |
| 12 | 88.02% | 3480/1160, | 83.19% | ±0.75% |
| 14 | 88.02% | 3480/1160, | 83.05% | ±0.84% |
| 16 | 86.72% | 3480/1160, | 80.98% | ±0.84% |
| d | Optimal Training Range | Samples per Qubit at Optimum | Peak QSVM | Notes |
|---|---|---|---|---|
| 4 | 600 | 150 | 79.75% | Limited expressibility |
| 6 | 600 | 100 | 84.25% | Same |
| 8 | 600 | 75 | 88.00% | Experiment 1.1 sweet spot |
| 10 | 600 to 1600 | 60 to 160 | 86.75% | Concentration onset visible |
| 12 | 1600 and 3480 | 133 and 290 | 88.02% | Right-shifted optimum |
| 14 | 3200 to 3480 | 229 to 249 | 88.02% | Joint peak |
| 16 | 3480 | 218 | 86.72% | Sharp full-data drop |
| 16 | 4640 (full data) | 290 | 82.16% | −4.56 pp concentration penalty |
| Feature Map | 600/400 QSVM | CV (600/400) | 800/200 QSVM | CV (800/200) | Kernel Time (600 Train) |
|---|---|---|---|---|---|
| ZZFeatureMap | 88.00% | 77.67 ± 2.91% | 86.00% | 79.12 ± 2.76% | 50 s |
| Pauli_XX | 84.75% | 77.33 ± 2.76% | 87.00% | 79.25 ± 2.42% | 1 min |
| Pauli_ZZ_XX | 77.50% | 69.00 ± 2.95% | 79.00% | 76.88 ± 1.98% | 1 min |
| Pauli_full | 65.25% | 63.00 ± 3.23% | 66.00% | 66.88 ± 1.85% | 4 min |
| Class | Precision | Recall | F1-Score |
|---|---|---|---|
| Gemma 3 | 91.37% | 83.97% | 87.51% |
| Qwen 2.5 | 85.17% | 92.07% | 88.48% |
| Weighted average | 88.27% | 88.02% | 88.00% |
| Dimension | New Stylistic Signals Entering the Named-Feature Buckets |
|---|---|
| 8 | exclaim_ratio, question_ratio, capital_ratio (all Gemma) |
| 10 | formality (Qwen) first appears; contraction_ratio surfaces marginally (bottom of the top-100; see Figure 4) |
| 12 | vocab_richness (Gemma) appears; all previously named axes persist |
| 14 | Full sixteen-feature fingerprint: twelve Gemma- and four Qwen-aligned features |
| 16 | contraction_ratio (Gemma) reappears in the full-data (4640-sample) run, completing the fifth stylistic axis (configuration-sensitive; see Appendix B.8) |
| Aspect | Linear-Weight Attribution | PCA Representation Salience |
|---|---|---|
| Input space | Full 3018-dimensional vector | 14-dimensional PCA input |
| Source of reported score | Trained LinearSVC weight vector | PCA loadings and explained variance |
| Kernel dependence | Specific to the trained model | None (kernel-agnostic by construction) |
| Profile shape | Concentrated on a few dominant features | Spread across all 18 named features |
| Top-1 feature | vocab_richness | avg_sent_len |
| Feature | Definition | Expected Discriminator |
|---|---|---|
| digit_ratio | count(digits)/len(text) | Qwen uses more numerical references |
| sent_len_std | std of sentence lengths in tokens | Gemma shows higher variance |
| hedge_ratio | distinct hedge phrases/sent_count | Qwen hedges more (e.g., “it is important to note”, “however”) |
| paragraph_count_ratio | num_paragraphs/sent_count | Qwen segments into more paragraphs |
| repeated_bigram_ratio | repeated bigrams/total bigrams | Captures phrase-repetition habits |
| avg_word_len_std | std of word lengths | Lexical variety beyond mean word length |
| Pipeline | QSVM Accuracy | CSVM Accuracy | CV Mean | CV Std |
|---|---|---|---|---|
| Experiment 2.1 (PCA baseline, word TF-IDF) | 88.02% | 98.02% | 83.05% | ±0.84% |
| Experiment 2.2 (char n-gram hybrid) | 76.55% | 89.74% | 74.71% | ±1.18% |
| Experiment 2.3 (direct stylometric, 1 qubit = 1 feature) | 78.19% | 90.86% | 76.84% | ±0.56% |
| Experiment 2.4 (classical, same features as Experiment 2.3) | n/a | 90.86% | n/a | n/a |
| Dimension | Best Classical Kernel (C) | Validation Accuracy | CV Mean ± Std |
|---|---|---|---|
| 8 | RBF () | 83.02% | 82.41 ± 1.70% |
| 10 | RBF () | 95.60% | 94.83 ± 1.31% |
| 12 | RBF () | 97.41% | 97.04 ± 0.96% |
| 14 | RBF () | 97.76% | 96.35 ± 0.55% |
| 16 | RBF () | 97.59% | 96.35 ± 0.85% |
| 14 | LinearSVC () | 96.64% | 96.03 ± 1.06% |
| Frozen Configuration | Input Representation | Training Samples | Validation Accuracy | Test Accuracy |
|---|---|---|---|---|
| QSVM , (Experiment 2.1) | PCA, 14 components | 3480 | 88.02% | 87.30% |
| QSVM , | PCA, 12 components | 3480 | 88.02% | 88.40% |
| QSVM , (Experiment 1.1 sweet spot) | PCA, 8 components | 600 | 88.00% | 79.80% |
| QSVM direct stylometric encoding (Experiment 2.3) | 14 MI-selected features | 3480 | 78.19% | 76.80% |
| Classical LinearSVC ablation (Experiment 2.4) | 14 MI-selected features | 3480 | 90.86% | 90.60% |
| Classical RBF SVM, (Experiment 2.5) | PCA, 14 components | 3480 | 97.76% | 95.60% |
| Global classical SVM | Full 3018-dimensional vector | 3480 | 98.02% | 97.70% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kopanov, K.; Atanasova, T. A Systematic Benchmark of Quantum Support Vector Machines for Interpretable Attribution of AI-Generated Text. Information 2026, 17, 883. https://doi.org/10.3390/info17090883
Kopanov K, Atanasova T. A Systematic Benchmark of Quantum Support Vector Machines for Interpretable Attribution of AI-Generated Text. Information. 2026; 17(9):883. https://doi.org/10.3390/info17090883
Chicago/Turabian StyleKopanov, Kalin, and Tatiana Atanasova. 2026. "A Systematic Benchmark of Quantum Support Vector Machines for Interpretable Attribution of AI-Generated Text" Information 17, no. 9: 883. https://doi.org/10.3390/info17090883
APA StyleKopanov, K., & Atanasova, T. (2026). A Systematic Benchmark of Quantum Support Vector Machines for Interpretable Attribution of AI-Generated Text. Information, 17(9), 883. https://doi.org/10.3390/info17090883

