Fully Quantized Training vs. Post-Training Quantization for a Small Hyperspectral Transformer Model for Pixel-Level Foreign Plastic Object Classification
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsIt is an interesting work from the perspective of model development for foreign objects detection using hyperspectral imaing in the wavelength range of 1000-1700 nm.
The content of the research is sufficient from the perspective of computer science, but from the perspective of the agricultural food sector, the following information should be added:
- The parameters for hyperspectral imaging should be provided e.g. intergration time, the distance from the lens to the objects.
- The types, sizes and distribution locations of the foreign objects should also be provided.
- If the foreign object is partially obscured or completely hidden, this is a difficult issue in foreign object detection and should also be discussed.
- The types and colors of plastic foreign objects, such as transparent ones with or without color, and the reasons for the differences in reflected light spectra. These were not mentioned in the text.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsThis paper systematically investigates the training and inference performance of low-precision floating-point quantization (FP32, FP16, BF16, FP8, and NVFP4) on a compact transformer-based hyperspectral classification model. Through a rigorously controlled experimental design, the study reveals several key findings: FQT consistently outperforms PTQ in accuracy preservation, FP8 offers a favorable balance between accuracy and compression efficiency, and ultra-low-precision formats exhibit distinctly different batch-size scaling behaviors between training and inference. The topic is novel, and the conclusions provide valuable guidance for model deployment in resource-constrained scenarios. However, the paper still has certain issues that require careful revision. We recommend that the authors thoroughly address the concerns raised before final acceptance. The specific comments are as follows:
- The abstract should be reorganized around a clear logical thread that synthesizes the key findings, rather than listing results in a format-by-format sequence.
- The introduction inadequately contextualizes low-precision quantization in HSI applications, failing to clearly identify that existing studies remain largely confined to FP32/FP16 training and INT8 inference, with no systematic evaluation of FP8/FP4 suitability for HSI Transformers. We recommend strengthening the research gap statement and clarifying the novelty of this work in addressing this void.
- There are extra commas after Equations (1), (4), and (5). We recommend removing them.
- RemoveEquation (3), as the value of σ can be simply stated in the text without loss of clarity.
- The paper states that "quantization is applied to the linear projection layers and layer normalization modules," while "the initial 1×1 convolution, RoPE, softmax, and the final MLP head are retained in FP32." However, no explanation is provided regarding why these specific modules are chosen for quantization and why others are preserved. We suggest adding a paragraph in Section 3.3 or Section 2.2 to elaborate on the selection criteria for sensitive modules—for instance, based on gradient norms, activation distributions, or layer-wise quantization sensitivity analysis—so that the strategy design becomes interpretable rather than merely empirical.
- The strategy is referred to as "mixed-precision FQT" throughout the paper, yet FQT literally stands for "Fully Quantized Training," which may cause confusion among readers: if it is "fully" quantized, why are some modules still retained in FP32? We suggest clarifying the methodological description to resolve this ambiguity.
- It is not specified whether the data split was performed at the pixel level or the image level.If the split was performed at the pixel level, pixels from the same hyperspectral cube could appear in both the training and test sets, potentially leading to overly optimistic evaluation results due to spatial correlation and data leakage.
If the split was performed at the image level—i.e., entire cubes assigned to a single set—then the 85% test proportion would correspond to approximately 44 images for testing and 8 images for training and validation combined. This remains a viable design, but requires explicit confirmation and justification.
- There is an extraneous symbol after the heading of Section 3.2. Please remove it.
- Part of the content in Section 4.2.1 appears to have been placed below Figure 2. Please pay attention to the typesetting and ensure that the text is properly positioned.
- In line 548, the OA improvement of FP16 FQT over PTQ is calculated as 98.33 - 98.02 = 0.31 percentage points, yet the text reports it as "0.30 pp."
The paper contains an excessive number of long and complex sentences, which to some extent impede reading fluency and clarity of expression. In addition, certain passages exhibit redundancy—particularly in the descriptions of FQT and PTQ definitions, the mixed-precision strategy, and the experimental setup, where similar information is reiterated multiple times. We recommend that the authors carefully revise the entire manuscript by breaking down lengthy sentences, streamlining repetitive content, and adopting a more concise and straightforward writing style. This would greatly enhance the overall readability and the scholarly quality of the presentation.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Reviewer 3 Report
Comments and Suggestions for AuthorsFile attached
Comments for author File:
Comments.pdf
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Reviewer 4 Report
Comments and Suggestions for AuthorsThe paper is devoted to FQT vs. PTQ for a Small Hyperspectral Transformer Model for Pixel-Level Foreign Plastic Object Classification. Of course, intelligent data processing, in particular transformer-based HSI, is a key part of modern sensor technologies. Studying the influence of training algorithms and data types on the accuracy and efficiency of foreign object detection and classification is a relevant task for food safety control. Overall, the study is well designed, however I have some comments on the text:
1. It would be helpful to avoid duplicating information and dispersing it across different parts of the paper. For example, the specific features of the FQT, PTQ, and QAT methods are discussed in the introduction (lines 91-103), in sections 2.2.1 and 2.2.2, and in section 3.3. Concentrating similar information could facilitate its comprehension by readers.
2. Line 279, Please explain in more detail the process of obtaining the "dark reference". Namely, by averaging how many frames was it obtained?
3. From a practical perspective, this article focuses on the detection and identification of small (≈5 mm) non-biological plastic foreign objects (12 types of polymers) in food products (poultry fillets). Nowhere in the article are the specific features of the hyperspectral data for this practically important task described. If such features are absent, the article is of a more general nature than clamed in the title. If such features are present, they should be specified: for example, the range of reflectance values in the spectral range of 1000-1700 nm for tissue fillets and polymers, or the difference in reflectance (minimum and maximum values) at a fixed wavelength for tissue fillets and polymers.
4. It would be desirable to explain from a physical point of view the fact of excluding the 600-1000 nm wavelength range from the hyperspectral data set.
5. It would be worthwhile if the conclusion included a forecast for the application of the method developed for identifying foreign plastic objects with sizes greater than and less than 5 mm, i.e. an assessment of the possible range of sizes for detecting and identifying foreign plastic objects in food products.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf

