Next Article in Journal
Production of Aluminum Hydroxide from Aluminate Solution Obtained by Red Mud Sinter Processing via Carbonation
Previous Article in Journal
Experimental and Kinetic Modeling Study on the Autoignition of Ammonia/Propane Mixtures
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Benchmarking Performance Judgment in Open-Weight LLM Controller Tuning: Control Knowledge Does Not Ensure Reliable Gain Ranking

1
Faculty of Humanities and Sciences, Mufu Campus, Jinling Institute of Technology, Nanjing 210038, China
2
Institute of Intelligent Transportation Systems, Zhejiang University, Hangzhou 310058, China
3
College of Water Resources and Civil Engineering, China Agricultural University, Beijing 100083, China
*
Author to whom correspondence should be addressed.
These authors contributed equally to this work.
Processes 2026, 14(18), 2954; https://doi.org/10.3390/pr14182954
Submission received: 10 August 2026 / Revised: 8 September 2026 / Accepted: 11 September 2026 / Published: 16 September 2026
(This article belongs to the Section Process Control, Modeling and Optimization)

Abstract

Large language models (LLMs) are increasingly embedded in control-design workflows, yet their ability to compare candidate controllers remains uncertain. A simulator-grounded performance-judgment benchmark evaluates four open-weight checkpoints across a quadruple-tank process, a recycle reactor, and a synthetic 3-by-3 cyclic plant. Declarative control-theory accuracy is nearly saturated (98–100% pooled by plant), whereas qualitative-description accuracy for ranking PI gain sets is 46.8%, 31.8%, and 53.2%, respectively. Outputs are associated with lower gain magnitudes, but the association varies with scaling, plant, and controller sampling. Across 271 replay-stable pairs, IAE rankings agree with ISE rankings on 95.2%; under a combined simulator perturbation, 88.7% of seed–item labels are preserved. On fixed gain pairs, equations and 14-point response trajectories raise pooled accuracy only from 46.8% to 51.2% and 51.0%. All 12 prespecified trajectory-versus-qualitative intervals include zero, and none survives Holm correction. Swapping displayed trajectory values changes 9.5% of parsed choices. By contrast, a numerical scaffold derived from the same samples raises evidence-consistent accuracy from 51.1% to 88.1%, with 11 of 12 contrasts surviving Holm correction. The main bottleneck under these prompts is extracting and aggregating raw numerical evidence, not the final comparison alone. This result does not identify a unique mechanism or establish a general absence of dynamic reasoning. Simulator-based validation remains necessary before LLM judgments are used for controller tuning.
Keywords: benchmarking; controller tuning; large language models; machine learning; multivariable systems; PID; process control benchmarking; controller tuning; large language models; machine learning; multivariable systems; PID; process control

Share and Cite

MDPI and ACS Style

Chen, J.; Shu, Y.; Li, H. Benchmarking Performance Judgment in Open-Weight LLM Controller Tuning: Control Knowledge Does Not Ensure Reliable Gain Ranking. Processes 2026, 14, 2954. https://doi.org/10.3390/pr14182954

AMA Style

Chen J, Shu Y, Li H. Benchmarking Performance Judgment in Open-Weight LLM Controller Tuning: Control Knowledge Does Not Ensure Reliable Gain Ranking. Processes. 2026; 14(18):2954. https://doi.org/10.3390/pr14182954

Chicago/Turabian Style

Chen, Jiaxuan, Yang Shu, and Haonan Li. 2026. "Benchmarking Performance Judgment in Open-Weight LLM Controller Tuning: Control Knowledge Does Not Ensure Reliable Gain Ranking" Processes 14, no. 18: 2954. https://doi.org/10.3390/pr14182954

APA Style

Chen, J., Shu, Y., & Li, H. (2026). Benchmarking Performance Judgment in Open-Weight LLM Controller Tuning: Control Knowledge Does Not Ensure Reliable Gain Ranking. Processes, 14(18), 2954. https://doi.org/10.3390/pr14182954

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop