Next Article in Journal
A Highly Parallel Integrated Process of Unloading, Exchanging, and Collecting for Rail-Changing
Previous Article in Journal
Evidence-Based Assessment of Commercial Fuel Additives Using OBD-Derived Fuel Economy Under Real-World High-Altitude Driving Conditions
Previous Article in Special Issue
Feasibility of Infrared-Based Pedestrian Detectability in Unlit Urban and Rural Road Sections Using Consumer Thermal Cameras
 
 
Article
Peer-Review Record

Thermal-Based Driver Monitoring in an Automotive Environment Using a Mobile Camera: A Feasibility Study

Vehicles 2026, 8(6), 116; https://doi.org/10.3390/vehicles8060116
by Yordan Stoyanov 1,2
Reviewer 1:
Reviewer 2: Anonymous
Reviewer 3: Anonymous
Vehicles 2026, 8(6), 116; https://doi.org/10.3390/vehicles8060116
Submission received: 18 February 2026 / Revised: 22 April 2026 / Accepted: 25 May 2026 / Published: 27 May 2026
(This article belongs to the Special Issue Novel Solutions for Transportation Safety, 2nd Edition)

Round 1

Reviewer 1 Report (New Reviewer)

Comments and Suggestions for Authors

The paper presents a practical feasibility study on using a low-cost LWIR camera for driver monitoring in a real-world automotive environment. The authors are commendably transparent about the scope of their work, explicitly positioning it as a structural stability and temporal consistency evaluation rather than a full drowsiness classification system. The experimental setup is clearly described, and the multi-layer validation framework is logical. However, to better meet the publication standards and improve the robustness of the study, several methodological details and practical limitations require further clarification.

  1. In Section 2.1, it is stated that thermal video was captured at a frame rate of 25 Hz. However, in the temporal evolution graphs (Figures 8-11), data points appear to be plotted at 5-minute intervals. Please clarify how the 25 Hz data was aggregated into these sparse data points. Are these single frames extracted every 5 minutes, or are they averaged values over a certain time window? Providing this detail is essential for understanding the temporal filtering process.
  2. Section 2.3 mentions that regions corresponding to the driver's facial area were identified manually. While acceptable for a baseline feasibility study, manual extraction significantly compromises the "real-world continuous driving" claim. Please add a brief discussion on how future implementation of automatic tracking algorithms might affect the measured spatial noise and uniform index, as algorithmic bounding box jitter could introduce artificial temperature fluctuations.
  3. The experimental sessions were all conducted in winter mornings with an outside temperature -1to 0and a stabilized cabin temperature of 23. Real-world cabin environments are often subjected to HVAC airflow directly hitting the face, or intense solar loading during summer. It is highly recommended to explicitly mention the lack of diverse thermal environmental testing (e.g., summer conditions, active AC blowing) in Section 4.6 (Limitations).
  4. At the end of Section 3.7 (Table 5), the authors attribute the high RMSE and modest Pearson correlation to "baseline temperature offsets." Considering the sessions were all conducted in the morning, could the authors briefly elaborate in the Discussion section on the potential physiological or contextual causes of these offsets? For example, differences in initial metabolic state, clothing insulation, or slight variations in camera calibration at startup.
  5. The results show greater variability in the "Minimum forehead temperature". Given the real driving conditions (60-90 km/h), vehicle vibrations and subtle involuntary head micro-corrections are inevitable. Please briefly discuss whether high-frequency vehicle vibrations contributed to ROI boundary crossing (e.g., catching cooler background pixels) during the manual tracking process, thus explaining this higher variability.
  6. The study relies on exactly two participants (healthy males, aged 37 and 39). While suitable for initial sensor validation, thermal emissivity and distribution can vary significantly with skin type, makeup, facial hair, or the presence of glasses. The authors should explicitly add demographic diversity (gender, age, facial features) to the list of future works/limitations in Section 4.6.

Author Response

Response to Reviewer 1

We sincerely thank the reviewer for the positive overall assessment of the manuscript and for the constructive suggestions. We appreciate the recognition that the work is positioned as a practical feasibility study rather than as a full drowsiness-classification system. The manuscript has been revised accordingly, and each point is addressed below.

Comment 1

“In Section 2.1, it is stated that thermal video was captured at a frame rate of 25 Hz. However, in the temporal evolution graphs (Figures 8–11), data points appear to be plotted at 5-minute intervals. Please clarify how the 25 Hz data was aggregated into these sparse data points.”

Response:
Thank you for this important comment. We clarified in the revised manuscript that the thermal video was acquired continuously at 25 Hz, but the quantitative temporal analysis was not performed frame by frame over the entire dataset. Instead, single representative frames were extracted at 5-minute intervals from each 60-minute session, yielding 12 sampled frames per session and 72 sampled frames in total across all six sessions. This clarification has been added in the revised Data Acquisition and Preprocessing section.

Comment 2

“Section 2.3 mentions that regions corresponding to the driver's facial area were identified manually… Please add a brief discussion on how future implementation of automatic tracking algorithms might affect the measured spatial noise and uniform index…”

Response:
We agree. The manuscript has been revised to explicitly state that the forehead temperature window was positioned manually in each selected frame and that this manually assisted workflow was intentionally adopted to avoid losing the target region in low-resolution LWIR imagery. We also expanded the Limitations section to note that future implementation of automated face/ROI tracking may introduce additional frame-to-frame boundary jitter and should therefore be evaluated separately in future studies.

Comment 3

“The experimental sessions were all conducted in winter mornings… It is highly recommended to explicitly mention the lack of diverse thermal environmental testing…”

Response:
We agree and have revised the manuscript accordingly. The Limitations section now explicitly states that all sessions were conducted during winter mornings under relatively stable conditions, with outside temperatures between approximately −1 °C and 0 °C and a stabilized cabin temperature of approximately 23 °C. We also note that these conditions do not represent summer solar loading, rapidly varying cabin thermal fields, or direct HVAC airflow toward the face.

Comment 4

“At the end of Section 3.7 (Table 5), the authors attribute the high RMSE and modest Pearson correlation to ‘baseline temperature offsets’… could the authors elaborate on the potential physiological or contextual causes of these offsets?”

Response:
Thank you. We have revised the relevant discussion to clarify that the observed inter-session offsets likely reflect day-to-day baseline differences rather than structural instability. We now explicitly mention plausible contributing factors such as pre-drive thermal adaptation, clothing insulation, initial physiological state, and small contextual differences in cabin microclimate.

Comment 5

“The results show greater variability in the ‘Minimum forehead temperature’… Please briefly discuss whether high-frequency vehicle vibrations contributed to ROI boundary crossing…”

Response:
We agree. The revised manuscript now clarifies that the minimum forehead temperature was the most sensitive metric because it was more strongly affected by small ROI boundary shifts, cooler peripheral pixels, and possible residual effects of natural driving vibration and subtle facial micro-movements, even though frames with obvious motion artefacts were excluded during manual selection.

Comment 6

“The study relies on exactly two participants… The authors should explicitly add demographic diversity to the list of future works/limitations…”

Response:
We agree and have revised the manuscript accordingly. The Limitations section now explicitly states that the study involved only two healthy adult male participants of similar age, and that further validation is required across broader driver populations, including differences in sex, age, facial morphology, facial hair, eyewear use, and health condition.

Final response to Reviewer 1

We thank the reviewer again for the helpful and constructive comments. We believe that the revised manuscript now provides a clearer and more transparent description of the sampled-frame workflow, the role of manual ROI placement, the interpretation of inter-session variability, and the limitations associated with subject diversity and environmental conditions.

Reviewer 2 Report (New Reviewer)

Comments and Suggestions for Authors

The concern raised by Reviewer 1 — that the comparison with existing methods is insufficient — has not been addressed in this revision. The paper describes the UTi260M's performance in complete isolation, without benchmarking against any reference thermal system (e.g., a research-grade FLIR or Axis camera) or against alternative driver monitoring approaches (visible-spectrum DMS systems, near-infrared cameras). The authors are required to either: (a) include a direct comparison of key noise metrics (NETD, FPN) between the UTi260M and at least one reference-class thermal sensor with published specifications; or (b) present a structured comparison table situating the current system's metrics within the range of values reported in the published literature for LWIR sensors used in automotive or biomedical thermal monitoring contexts.  Action required: Add a comparative analysis section or table in Section 4 (Discussion) that benchmarks the UTi260M's reported NETD and FPN against at least two published references. This comparison is non-negotiable; without it, the claim that this system is 'suitable' for in-vehicle monitoring has no calibration benchmark.

The addition of a second participant is acknowledged and welcomed. However, n=2 remains extremely small, and both participants are described as healthy adult males of similar age (37 and 39 years old). This demographic homogeneity — identical sex, similar age, comparable health status, and consistent morning recording conditions — is never acknowledged as a limitation on generalizability. The paper's conclusion that the system 'can provide structurally stable and repeatable facial temperature measurements' is stated without caveat.  Action required: The limitations section (Section 4.6) must be expanded to explicitly acknowledge that: (i) the study population consists exclusively of healthy adult males; (ii) the system's performance with drivers of different ages, sexes, body temperatures, health states, or facial morphologies remains unknown; and (iii) winter morning conditions may not represent the full range of in-vehicle thermal environments the system would encounter in deployment.

 

The KSS and drowsiness-detection components have been removed entirely from the revision. This is a legitimate methodological repositioning. However, the scope reduction is never explicitly acknowledged or justified in the manuscript itself. The introduction still frames the study extensively within the drowsiness detection literature (lines 33–90), creating a misleading expectation that the paper will report drowsiness-related findings. Readers unfamiliar with the revision history will reasonably expect behavioral inferences that are not delivered.  Action required: The final paragraph of the introduction should include a clear, explicit statement such as: 'The present study is intentionally scoped to sensor-level validation as a methodological prerequisite for drowsiness detection; behavioral inference is deferred to future work incorporating appropriate ground truth measures.' Additionally, the literature review should be refocused so that drowsiness-detection references (citations [16]–[22]) are cited only as motivating context, not as the methodological precedent for the present study.

 

Section 2.3 (lines 160–164) states that regions corresponding to the driver's facial area 'were identified manually.' No information is provided about how this identification was made consistent across frames, sessions, or participants. At 25 Hz over six 60-minute sessions, the total dataset comprises approximately 540,000 frames. Manual frame-by-frame annotation at this scale is implausible, suggesting either a subset of frames was annotated, or some form of assisted segmentation was used. Neither the frame selection strategy nor the ROI boundary criteria are described.  Action required: Provide the following information in Section 2.3: (i) the total number of frames annotated per session and the selection criterion (e.g., every N-th frame, keyframes only); (ii) the anatomical or geometric criteria used to define forehead ROI boundaries (specific landmarks, fixed pixel window, or proportional head bounding box); (iii) whether more than one annotator was involved, and if so, the inter-rater agreement metric; and (iv) an illustrative figure or supplementary figure showing the ROI superimposed on a representative thermal frame.

 

Section 2.3 now provides a clearer description of the color-to-temperature reconstruction procedure (lines 165–168). This is an improvement over the original submission. However, the accuracy of this reconstruction is never validated. The UTi260M supports both a display mode (color-mapped JPEG/video output) and, potentially, a native radiometric output. If native radiometric output was available for any portion of the recordings, the reconstruction accuracy should be quantified against it. If no native radiometric output was available, this must be stated explicitly, and the reconstruction uncertainty must be bounded.  Action required: Report the reconstruction accuracy (e.g., mean absolute error or RMSE against a known reference or against the few frames where radiometric data were available). If no radiometric ground truth is available, state this explicitly and provide an estimate of the reconstruction precision based on the color scale discretization step size (i.e., the temperature resolution of the color map). Additionally, note that temperature values reported to four decimal places in Table 4 (e.g., 0.0041 °C/min) imply a measurement precision inconsistent with a color-map-based reconstruction unless the color scale resolution is extremely fine — this should be acknowledged and the reported precision adjusted accordingly.

 

Table 2 reports σ_noise, σ_col (column noise), and σ_row (row noise) as numerically identical: 5.1399 for Participant I and 4.6032 for Participant II across all three directional metrics. These quantities are defined by distinct equations (Eqs. 2, 3, and 4): σ_noise is a spatial average of local window-based noise; σ_col is the standard deviation of column-wise mean temperatures; σ_row is the standard deviation of row-wise mean temperatures. It is mathematically near-impossible for all three to yield identical values unless the same variable was inadvertently substituted for all three in the implementation.  The note below Table 2 attempts to justify this as 'uniform spatial fluctuation distribution,' but this explanation is insufficient — even a perfectly noise-uniform image would yield different absolute values for the three distinct formulas applied to the same pixel grid. This raises serious concern about whether the three metrics were computed independently and correctly.  Action required: The authors must either: (a) provide a clear mathematical demonstration of why these three values are expected to be identical for the analyzed frame; or (b) acknowledge that this is a computational implementation error, correct the analysis, and report the independently computed values for each metric. Given that Table 2 is the primary quantitative evidence for sensor characterization, this issue must be resolved with full transparency.

 

All noise metrics in Table 2 (σ_noise, σ_col, σ_row, FPN values, NETD_est, and SNR) are reported without units. The physical interpretation of these values is ambiguous: if expressed in degrees Celsius, σ_noise = 5.14°C would correspond to a thermal sensitivity of 5140 mK — approximately 100 times worse than the manufacturer specification of <50 mK for the UTi260M. This suggests the analysis may have been performed on reconstructed 8-bit pixel intensity values (0–255) rather than calibrated temperature values in °C. If the SNR formula (Eq. 7, SNR = T/σ_global) yields SNR = 20.84 with σ_global = 5.14, then the mean signal T ≈ 107 — consistent with an 8-bit pixel value for a frame containing both cool background and warm facial regions, but inconsistent with a temperature in degrees Celsius.  Action required: The authors must: (i) explicitly state the unit system for all values in Table 2; (ii) clarify whether the noise analysis was performed in pixel space (digital numbers) or in calibrated temperature space (°C); (iii) if performed in pixel space, convert all noise metrics to their equivalent temperature values using the color scale calibration and report both; and (iv) reconcile the reported NETD_est with the manufacturer's specification of <50 mK. The NETD claim is currently unsupported by the numbers as presented.

Table 1 specifies the UTi260M's native sensor resolution as 256×192 pixels. However, Table 2 reports an image resolution of 1440×1080 pixels for both participants — a factor of approximately 5.6× in each dimension, consistent with upscaled/interpolated display output rather than native sensor pixels. This distinction is never acknowledged in the manuscript.  Spatial noise metrics (σ_noise, FPN) computed on bilinearly or bicubically interpolated frames do not accurately represent true sensor pixel noise, because interpolation artificially correlates adjacent pixels and smooths local noise — potentially causing the reported noise values to significantly underestimate true sensor noise.  Action required: Clarify in Section 2.3 or 2.4: (i) whether all noise analyses were performed on the upscaled 1440×1080 display frames or on extracted native 256×192 sensor data; (ii) if performed on upscaled data, quantify the expected impact of interpolation on spatial noise estimates; and (iii) if possible, repeat the noise analysis on the native-resolution frames and report both results. All conclusions about sensor 'suitability' based on the noise metrics depend critically on this clarification.

In Figure 8b (Forehead Mean ROI Temperature — Participant I), Session 2 shows a pronounced temperature decrease to approximately 30°C between approximately 20 and 30 minutes into the session. This represents a 4–6°C drop below the physiologically expected human forehead temperature range and falls below the 35–37°C range cited in the abstract as 'biologically plausible.' The current explanation — that this is a 'localized decrease between approximately 20–30 minutes' attributed to 'localized cooling effects and small ROI boundary variations' (lines 384–386) — is inadequate.  A spontaneous 5°C forehead temperature decrease in a healthy adult male during monotonous highway driving is either: (a) a genuine physiological event (e.g., skin vasoconstriction, perspiration onset, or a transient cabin cooling event) that should be documented with contextual data from that session; or (b) a measurement artifact (ROI tracking failure, partial occlusion, a window reflection, or background thermal intrusion into the ROI). Either interpretation has significant implications for the claimed reliability of the system.  Action required: The authors must investigate and explain this specific event in Session 2 at 20–30 minutes. If it is a measurement artifact, this constitutes a sensor or ROI extraction failure that directly contradicts the paper's claim of stable, reliable measurement — and must be prominently discussed in the limitations section as a known failure mode. If it is a genuine physiological event, contextual documentation from that session (cabin temperature at that time, participant self-report, ambient conditions) must be provided.

The forehead ROI is described only as 'a predefined forehead region of interest (ROI), including center-point, mean, maximum, and minimum temperature metrics.' No spatial specification is provided: neither pixel dimensions, proportion of face bounding box, anatomical landmarks, nor any illustration showing where exactly the ROI is placed relative to the thermal face image. A researcher attempting to replicate this study would have no way to consistently define the same forehead region.  Action required: Provide at minimum: (i) approximate dimensions of the ROI relative to the thermal image (e.g., a W×H pixel window centered on the glabella, occupying approximately X% of the detected face bounding box); (ii) the criteria used to position the ROI per frame; and (iii) at least one illustrative thermal frame figure showing the ROI boundary overlaid on the thermal face image. This can be added as a supplementary figure if space is constrained.

 

Section 2.4 (line 239) states that 'local noise values were computed using small windows (e.g., 5×5 or 7×7).' The phrase 'e.g.' indicates this is illustrative rather than definitive. The actual window size used for all reported calculations is not stated. Noise estimates depend directly on the window size, and using 'e.g.' rather than the actual parameter used is insufficient for reproducibility.  Action required: State unambiguously in Section 2.4 which window size was used for all reported analyses. If both sizes were evaluated, report both results and state the selection criterion.

 

Table 5 reports inter-session Pearson correlation values ranging from −0.02 to 0.40. The authors state that 'despite modest correlation values, the bounded slope magnitudes and low intra-session variability confirm stable sensor behavior' (lines 456–458). This framing conflates two distinct properties: (a) within-session sensor stability (correctly assessed by CV and slope in Table 4) and (b) cross-session temporal signal similarity (assessed by Pearson r in Table 5).  Correlations of −0.02 to 0.40 indicate very weak to negligible linear similarity between temporal thermal trajectories across sessions. This is not evidence of sensor instability per se, but it is also not evidence of sensor stability — it simply means the temporal patterns of temperature evolution differ between sessions, which is physiologically expected if participants' thermoregulatory states differ day-to-day.  The authors should not use low inter-session correlation to support claims of 'stable sensor behavior.' These are orthogonal claims.  Action required: Revise the language in lines 455–459 to clearly distinguish: 'Within-session stability, as assessed by CV (Table 4), confirms that the sensor provides consistent measurements within individual recording sessions. Inter-session comparisons (Table 5) reflect physiological baseline variation between recording days rather than sensor drift, as no progressive increase in within-session variability was observed across sessions.' Remove or reframe any language that implies low inter-session correlation is evidence of sensor stability.

 

The introduction devotes substantial space to the problem of drowsiness detection, fatigued driving, and behavioral indicators of alertness decline (lines 33–90). This framing creates a reasonable expectation in readers that the study will report drowsiness-related findings. However, the study contains no drowsiness induction protocol, no behavioral ground truth, no sleepiness assessment, and no inference about driver state from the thermal measurements. The results section reports only sensor noise metrics and temporal thermal stability.  This disconnect between introduction and results undermines the paper's internal coherence and may lead reviewers or readers to assess the paper as under-delivering on its stated purpose.  Action required: Restructure the introduction so that drowsiness detection is positioned clearly as the downstream motivating application, not the study's direct objective. Insert the following type of bridging statement at the close of the introduction: 'As a prerequisite for thermal-based behavioral inference, the present study evaluates whether a low-cost LWIR sensor can provide the measurement stability and repeatability required as a foundation for future driver state detection algorithms. Behavioral and drowsiness-specific analyses are deferred to subsequent work.' This reframing requires targeted edits to lines 104–111 and the final introduction paragraph.

 

Section 2.5 (lines 191–194) states that 'observable behavioral indicators, such as head flexion and reduced postural micro-corrections, were examined alongside thermal characteristics.' This implies that the study analyzed postural behavior quantitatively. However, no results for these behavioral indicators appear anywhere in the manuscript: not in Section 3 (Results), not in Table 4 or 5, and not in the Discussion.  Action required: If behavioral observation data were collected but not analyzed, remove the claim from Section 2.5. If behavioral indicators were qualitatively observed, this should be stated explicitly as 'observational, not quantified.' The manuscript should not imply that behavioral indicators were 'examined' if no quantitative results are presented.

The Abbreviations section lists 'FTN' as 'Fixed-Pattern Noise.' The correct abbreviation, used consistently and correctly throughout the manuscript body (Sections 2.4, 2.6, Table 2, and the Discussion), is 'FPN' (Fixed-Pattern Noise).  Action required: Correct 'FTN' to 'FPN' in the Abbreviations section.

I draw the authors' attention to the following publications, which are directly relevant to the study's scope and should be considered for citation in the revised manuscript:

Yan, L., Wu, C., Zhu, D., Ran, B., He, Y., Qin, L., & Li, H. (2017). "Driving Mode Decision Making for Intelligent Vehicles in Stressful Traffic Events." Transportation Research Record, 2625(1), 14–23. DOI: 10.3141/2625-02

This paper presents a multi-sensor framework for classifying real-world driver states — integrating physiological indices, vehicle dynamics, and self-reported behavioral records under stressful traffic conditions — which directly parallels the motivating application of the current study. Citing this work would situate the proposed thermal sensor validation within an established multi-sensor driver state monitoring paradigm, strengthening the methodological positioning of thermal imaging as a complementary sensing modality.

 

Comments on the Quality of English Language

P. Mitev 'Development of a Training Station…'

Single quotation marks are used instead of standard double quotation marks. Additionally, the relevance of this reference — which concerns a training station for dice part orientation using machine vision — to a driver thermal monitoring study is not apparent and should be verified or removed.

"Previous works often focus on drowsiness classification…"

"Previous works" is non-standard. Should be "Previous work often focuses…" or "Previous studies often focus…"

"Thermal Image Quality Assessment can by according to the following dependencies"

"can by according to" is ungrammatical. Can be: "Thermal Image Quality Assessment was performed according to the following dependencies" or "…can be conducted according to the following…"

 

"rather than in sensor hardware itself"

Should be: "rather than in the sensor hardware itself."

Unclear Phrase

"without structural amplification between participants"

This phrase is ambiguous. The intended meaning seems to be "without systematic differences in FPN magnitude between participants," which should be stated explicitly.

"at the same time instance" / "at comparable time instances"

"Instance" is incorrect here. Should be "time instant" or "time point" in both occurrences.

"where H is image height and W is image width"

Should read: "where H is the image height and W is the image width."

"Driving speed varied between 60 and 90 km/h in 60 minutes mode."

"in 60 minutes mode" is undefined and unclear. Suggest: "Driving speed varied between 60 and 90 km/h throughout each 60-minute session."

The manuscript uses "artefacts" (UK) in lines 160, 337, and 357, and "nearest-neighbour" (UK) in line 168, while other spellings follow US conventions. The author should choose one convention and apply it consistently throughout. Given MDPI's predominantly international-English style, either is acceptable but must be uniform.

"The system operates under: varying head orientation; natural posture adjustments; realistic cabin environment; prolonged monitoring duration."

A colon followed by semicolons is unconventional. Convert to prose: "The system operates under varying head orientation, natural posture adjustments, a realistic cabin environment, and prolonged monitoring duration."

Awkward Phrasing

"These findings support temporal robustness rather than identical inter-session replication."

Ambiguous. Suggest: "These findings support claims of within-session temporal stability, without implying identical signal replication across sessions."

Run-On Sentence in Conclusions

"…normalized analysis confirmed preservation of relative temporal behavior indicating that variations are more consistent with baseline session offsets rather than structural sensor instability."

This is a run-on. A comma is needed after "behavior," and "rather than structural sensor instability" should read "than with structural sensor instability" for parallel construction: "…preservation of relative temporal behavior, indicating that variations are more consistent with baseline session offsets than with structural sensor instability."

 

Abbreviation List — FTN vs. FPN Already flagged in my previous comment, but worth noting again here: the abbreviation "FTN" in the Abbreviations section is a typographical error for "FPN" (Fixed-Pattern Noise).

Non-Standard Citation Format The National Sleep Foundation reference is formatted as a bracketed web access note rather than MDPI's standard reference style. It should be reformatted to comply with the journal's citation guidelines.

"Conceptualization, Y.S. ;methodology…" and "validation, Y.S.,; formal analysis"

There is a spurious space before the semicolon after the first entry, and an extra comma before the semicolon in the validation entry. These should be corrected to standard MDPI author contributions format. More critically, "software, C.W." attributes software development to a co-author "C.W." who does not appear anywhere else in the manuscript — there is only one listed author (Y.S.). This appears to be a copy-paste error from another submission and must be corrected or explained.

 

Author Response

Response to Reviewer 2

We sincerely thank the reviewer for the detailed and rigorous evaluation of the manuscript. The comments were highly valuable and led to substantial revision and restructuring of the work. In response, we carefully reconsidered the scope, interpretability, and presentation of the study and revised the manuscript accordingly. Most importantly, the manuscript has now been explicitly reframed as a feasibility-level workflow study rather than as a detector-calibration or driver-state-classification paper.

Comment 1

“The concern raised by Reviewer 1 — that the comparison with existing methods is insufficient — has not been addressed…”

Response:
We appreciate this comment. In the revised manuscript, the Introduction was strengthened with additional and more relevant references to driver monitoring, head tracking, NIR-based monitoring, thermal driver-state-related studies, and recent intelligent driver-monitoring systems. Rather than claiming equivalence with specialized automotive or research-grade thermal systems, the revised manuscript now more explicitly positions the present work as a consumer-grade feasibility study. The contribution is therefore framed as signal-level and workflow-level validation rather than as competitive benchmarking against high-end systems.

Comment 2

“The addition of a second participant is acknowledged… demographic homogeneity is never acknowledged as a limitation…”

Response:
We agree and have revised the Limitations section accordingly. The revised manuscript now explicitly states that the study included only two healthy adult male participants of similar age, and that the findings cannot be generalized without further validation across broader driver populations and different physiological and morphological characteristics.

Comment 3

“The KSS and drowsiness-detection components have been removed entirely… The scope reduction is never explicitly acknowledged or justified…”

Response:
We thank the reviewer for this important observation. The Introduction has been thoroughly revised to clarify that the present study is not a behavioral or drowsiness-classification experiment. The final paragraph of the Introduction now explicitly states that the manuscript is intentionally limited to signal-level feasibility validation of a consumer-grade mobile LWIR workflow as a methodological prerequisite for future driver-state inference, rather than a direct drowsiness-detection study.

Comment 4

“Section 2.3 states that regions corresponding to the driver's facial area were identified manually… Please provide the total number of frames annotated, selection criterion, ROI boundary criteria…”

Response:
We agree and have substantially revised Section 2.3. The revised manuscript now states explicitly that:

  • the analysis was performed on single sampled frames extracted at 5-minute intervals rather than exhaustive frame-by-frame annotation;
  • this yielded 12 sampled frames per session and 72 sampled frames in total;
  • the forehead temperature window was positioned manually in the upper frontal facial region above the periorbital area and below the hairline;
  • frames with visually unstable facial capture or obvious motion artefacts were excluded;
  • ROI placement was performed by a single operator to maintain internal consistency.

Comment 5

“The color-to-temperature reconstruction procedure is described more clearly, but the accuracy of this reconstruction is never validated…”

Response:
We appreciate this point and have revised the manuscript accordingly. The revised Methods section now states clearly that the UTi260M workflow did not provide raw detector-level radiometric data. Therefore, the reported temperature-derived values are explicitly interpreted as approximate workflow-level measurements derived from app-displayed thermal information, rather than as fully validated radiometric measurements. We also reduced the scope of the paper accordingly and removed unsupported detector-level interpretations.

Comment 6

“Table 2 reports σ_noise, σ_col, and σ_row as numerically identical…”

Response:
We thank the reviewer for identifying this important issue. Upon reconsideration of the presentation and interpretability of the original noise-related metrics, and given that the workflow is based on app-displayed thermal imagery without raw detector-level export, we concluded that strict detector-style spatial noise characterization was not sufficiently supported in the manuscript. Therefore, the original noise-oriented presentation was removed and replaced with a new workflow-oriented summary table describing the sampled-frame thermal analysis procedure. The manuscript was refocused on workflow-level consistency, geometric feasibility, cross-session repeatability, and temporal stability, which are more appropriate to the available data.

Comment 7

“All noise metrics in Table 2 are reported without units… The NETD claim is currently unsupported…”

Response:
We agree. In response to this comment, we removed unsupported detector-level noise claims, including interpretations related to NETD, and revised the text so that the analysis is explicitly presented as workflow-level thermal frame consistency rather than detector-level thermographic characterization. This change is consistent with the fact that the UTi260M workflow did not provide raw radiometric export.

Comment 8

“Table 1 specifies the native sensor resolution as 256 × 192, while Table 2 reports 1440 × 1080…”

Response:
We agree that this distinction required clarification. The revised manuscript now clearly states that the study relied on smartphone application displayed thermal imagery rather than native detector-level frame export. Accordingly, the analysis is now framed as a practical assessment of the deployed mobile thermal workflow rather than strict native-sensor characterization.

Comment 9

“In Figure 8b, Session 2 shows a pronounced temperature decrease… this must be investigated and explained…”

Response:
We appreciate this comment and revised the interpretation accordingly. The revised manuscript now explicitly states that this isolated decrease is interpreted conservatively as a likely frame-level measurement anomaly related to ROI sensitivity or partial peripheral inclusion, rather than as robust evidence of a physiological forehead-temperature decrease of that magnitude. This interpretation is supported by the overall stability of the remaining trajectories and by the manually assisted, sampled-frame workflow.

Comment 10

“The forehead ROI is described only as a predefined forehead region… no spatial specification is provided…”

Response:
We agree and expanded the methodological description. The revised manuscript now specifies that the manually positioned forehead window was located in the upper frontal facial region above the periorbital area and below the hairline, while avoiding obvious peripheral non-target regions and background inclusion. The manually assisted nature of this ROI definition is also acknowledged more explicitly in the revised Limitations section.

Comment 11

“Section 2.4 states that ‘local noise values were computed using small windows (e.g., 5×5 or 7×7)’…”

Response:
This issue became non-applicable after the manuscript was revised and the unsupported detector-style noise characterization block was removed. The revised manuscript no longer relies on the original local-noise analysis framework.

Comment 12

“Low inter-session correlation should not be used as evidence of sensor stability…”

Response:
We agree and revised the interpretation accordingly. The manuscript now clearly distinguishes:

  • within-session stability, assessed through coefficient of variation and ROI mean slope; and
  • inter-session similarity, assessed through Pearson correlation and RMSE.

The revised wording now states that modest inter-session correlations are more plausibly explained by day-to-day physiological and contextual baseline variation rather than direct evidence of sensor instability.

Comment 13

“The introduction devotes substantial space to drowsiness detection… This disconnect undermines coherence…”

Response:
We agree and completely revised the Introduction. The revised version is shorter, more focused, and more clearly positions drowsiness-related literature as broader motivating context rather than as the direct objective of the present study.

Comment 14

“Section 2.5 states that behavioral indicators were examined…”

Response:
We agree. The relevant section has been revised. The manuscript now states that observable posture changes were considered only qualitatively as contextual observations and were not quantified in the present study.

Comment 15

“FTN should be FPN…”

Response:
Thank you. The abbreviation issue was corrected in the revised manuscript. Subsequently, after restructuring the study and removing the unsupported noise block, the abbreviations list was also simplified to include only the terms that remain relevant to the revised manuscript.

Comment 16

English-language and formatting issues

Response:
We thank the reviewer for the detailed language-related comments. The manuscript has been thoroughly edited for readability and consistency. Numerous sentences were revised, redundant passages were removed, the terminology was simplified and standardized, and formatting issues in the author contributions and references were corrected.

Final response to Reviewer 2

We are grateful for the reviewer’s detailed and rigorous comments. They resulted in substantial improvements to the manuscript. In response, we did not simply revise wording, but fundamentally clarified the scope of the study, removed unsupported interpretations, strengthened the methodological transparency, and reframed the contribution as a feasibility-level validation of a consumer-grade app-based LWIR workflow. We believe the revised manuscript is now significantly clearer, more coherent, and better aligned with the available data.

Reviewer 3 Report (New Reviewer)

Comments and Suggestions for Authors

Overall comments: This is an interesting case study into the feasibility of using LWIR in lieu of RGB cameras for identifying facial features and calculating things like head pose. While this study did not focus on actual facial feature identification / head pose computation, it did begin to lay the basic groundwork for using an LWIR signal from consumer-grade cameras by focusing on the quality of the LWIR signal from driver faces—namely, looking at signal variability in a number of computations. I have a few specific comments to address below, but other than that, it should be noted in the “Limitations” that this testing was done under pretty ideal conditions (extremely consistent external / internal temperatures and lighting), and that a lot of the computations required manual face or facial feature identification and/or bounding box drawing.

Specific comments:

Line 63: Did the investigation you cite here successfully identify drowsiness via thermal imaging? That is a crucial part to report here.

Introduction overall: My reading of the Introduction is that there have been no attempts to use thermal imaging in place of RGB cameras for determining things like head pose or other behavioral or facial indicators of drowsiness that are not specifically heat-related (like body temperature or respiration). Is that true? What you appear to be proposing here is using LWIR in lieu of RGB to have a lighting-independent approach to measures about the head—like head pose, head movement, possibly off-road glancing, and there have not been other approaches to driver monitoring using this signal. It would be helpful if this were explicitly stated, because while you cite a lot of driver monitoring research that uses heat-sensitive cameras to try to identify body temperature or breath, it seems your focus has nothing to do with that and is about laying the groundwork to use LWIR in place of RGB.

Lines 123-127: It appears all tests were done in largely similar lighting / heat conditions (morning). This isn’t a criticism, just something to note—that ostensibly the goal of using LWIR for these purposes is to handle variations in lighting that RGB isn’t good at, but most testing occurred in the environments RGB is good at handling.

Figure 2b: Displaying the camera output to the driver during testing is certainly an odd choice—I can see it being useful for engineering runs, but in the future I’d advise against it for testing on public roadways—it’s certainly likely to draw a lot of glancing. (This isn’t something that has to be addressed in the manuscript, just for future consideration.)

Lines 160-162: Bounding boxes had to be drawn around the participants’ faces on a frame-by-frame basis for 6 hours of video each? That sounds incredibly laborious and may potentially introduce error, but at the very least is a major limitation of the LWIR approach—that there are not current technologies available for automatic face segregation from the surround, as we have with RGB. However, it’s unclear to me whether this is used in any of the analyses—it looks like either the entire sensor field is used, or specific facial features are used as points of measurement.

Author Response

Response to Reviewer 3

We sincerely thank the reviewer for the thoughtful and constructive comments, and for recognizing the value of the study as an initial step toward the use of consumer-grade LWIR imaging in driver monitoring. The manuscript has been revised accordingly, and we respond to each point below.

Comment 1

“This is an interesting case study into the feasibility of using LWIR in lieu of RGB cameras…”

Response:
We sincerely appreciate this positive assessment and the recognition that the work is intended as groundwork for future driver-monitoring use of LWIR imagery. In the revised manuscript, we further clarified this positioning by explicitly defining the present work as a feasibility-level study focused on facial thermal consistency rather than automated feature detection or behavioral inference.

Comment 2

“It should be noted in the ‘Limitations’ that this testing was done under pretty ideal conditions…”

Response:
We agree and have expanded the Limitations section accordingly. The revised manuscript now explicitly states that all sessions were conducted during winter mornings under relatively stable outside and cabin temperatures, and that these conditions do not cover the full thermal variability of real in-vehicle environments.

Comment 3

“A lot of the computations required manual face or facial feature identification and/or bounding box drawing.”

Response:
We agree and clarified this point in the revised manuscript. We now explicitly state that the analysis used a manually positioned forehead temperature window for each sampled frame, that this was performed by a single operator, and that this manually assisted workflow is a limitation of the present study rather than a deployable automated solution.

Comment 4

“Line 63: Did the investigation you cite here successfully identify drowsiness via thermal imaging?”

Response:
Thank you. The Introduction has been revised to better distinguish between broader driver-state-related thermal studies and the specific methodological aim of the present manuscript. Drowsiness-related studies are now used more carefully as contextual motivation rather than as methodological precedent for the current study.

Comment 5

“The Introduction suggests that there have been no attempts to use thermal imaging in place of RGB…”

Response:
We appreciate this comment and revised the Introduction to better position the work in the broader landscape of driver-monitoring research. The revised text now emphasizes that thermal imaging is being considered here as an alternative or complementary non-visible sensing modality, while also making clear that the present manuscript does not present a complete thermal replacement for RGB-based driver monitoring.

Comment 6

“Lines 123–127: It appears all tests were done in largely similar lighting / heat conditions…”

Response:
We agree. This point has now been made explicit in the revised Limitations section, which states that the sessions were carried out under a restricted range of environmental conditions and therefore do not represent the full range of thermal and lighting conditions encountered in practical deployment.

Comment 7

“Lines 160–162: Bounding boxes had to be drawn around the participants’ faces on a frame-by-frame basis for 6 hours of video each?”

Response:
Thank you. We clarified this explicitly in the revised Data Acquisition and Preprocessing section. The revised manuscript now states that the full 25 Hz recordings were not annotated exhaustively frame by frame. Instead, the quantitative analysis used single representative frames extracted at 5-minute intervals, and the manually positioned forehead window was applied only to those selected frames.

Final response to Reviewer 3

We thank the reviewer for the constructive and balanced comments. We believe that the revised manuscript now more clearly presents the study as a feasibility-level thermal monitoring workflow, explicitly acknowledges its restricted environmental and methodological conditions, and avoids implying unsupported behavioral interpretation.

Round 2

Reviewer 1 Report (New Reviewer)

Comments and Suggestions for Authors

I appreciate your revisions.

This manuscript is a resubmission of an earlier submission. The following is a list of the peer review reports and author responses from that submission.


Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

The comparison with other existing methods is not enough.

Author Response

RESPONSE TO REVIEWER 1

We sincerely thank the reviewer for the constructive feedback.

Comment:

The comparison with other existing methods is not enough.

Response:

The Introduction and Discussion sections have been expanded to provide clearer contextualization within existing literature. A distinction is now made between (i) high-cost automotive-grade thermal systems and (ii) consumer-grade LWIR cameras. The contribution of the present study is explicitly defined as a structured validation of radiometric stability under controlled in-vehicle conditions rather than a full driver-monitoring system comparison. Additional references have been integrated to improve positioning within current research.

Comment:

The English could be improved.

Response:

The manuscript has been carefully revised for clarity, consistency, and grammatical correctness. Redundant expressions were removed, technical terminology was standardized, and several sentences were restructured to improve readability.

Reviewer 2 Report

Comments and Suggestions for Authors

In the first version of this article (which was rejected and then resubmitted) I highlighted that there were areas for improvement. In this new version, the authors seem to have integrated everything that had previously emerged. 

Author Response

Comment:

The introduction can be improved.

Response:

The Introduction has been revised to better define the scope of the study and to clarify the distinction between behavioral interpretation and methodological validation. The contribution is now clearly positioned as a radiometric stability assessment.

Comment:

Research design / methods can be improved.

Response:

The research design has been clarified and expanded. A second participant has been included, each completing three independent 60-minute sessions under controlled environmental conditions. A detailed summary of session parameters is now provided in Table 3. The environmental conditions and thermal gradients are explicitly documented.

Comment:

Figures and tables can be improved.

Response:

All figures and tables have been reorganized for logical sequence and consistent citation. Figure numbering has been corrected, and all figures and tables are now referenced explicitly within the text.

Reviewer 3 Report

Comments and Suggestions for Authors

The manuscript by Stoyanov presents a driver tracking method for driver monitoring based on thermal imaging using a long-wave infrared camera. I still like the idea of research. but, despite the quick revision and resubmission, have some concerns that need to be improved before it is ready for re.submission (and ultimately publication). 

Still, in my view, the research design and its description is not sufficient to provide the readers with enough clarity and information for showing the methods feasibility in a field study. While I still believe that the technical data and its evaluation appear to be feasible (section 2.6 thermal image quality assessment), the assessment of its feasibility for detecting drowniness signs from poses is hard to trust in its current form. In particular, the only participant of the study appears to be author himself, who provided self-reported sleepiness values (KSS) for a couple of drives. To add, there is no ground truth data in any form presented for the pose analysis. All together given these shortcomings, it is almot impossible to judge the feasibility of the thermal cam based on the presented data. I recommend to at least invite a few naive participants, let them drive under different drowsiness circumstances and assess some ground truth for the poses (e.g. from manual coding by co-driver or based on webcam that is installed in the cockpit next to thermal cam).

Here are my detailed comments:  

Abstract has some typos

Introduction is still ok

Materials and Methods:

  • The description of the experimental campaign is very misleading, it says it had 3 sessions in line 123 and six sessions in l. 135
  • I do not understand the math: 3 session x ~0,75 h (because each session lasted between 30-60 min) x ~110 km/h (because driving speed was between 90 and 130 km/h) = 247,5 km --> I understand that it could well be 3x0,6x90 (=162 km) or even less (and yes even below 150 km as mentioned) --> but with three sessions, it would be easier if the author would not leave it to the reader to do the math, but just be explicit
  • Worse is that table 3 shows 6 sessions taking exactly 60 mins
  • why only one participant? Is this the author himself? How can the author as single participant guarantee that the KSS scores are not biased by the experimental aim. I really recommend to do the drives again at least with a couple of subjects that do not know the goal of the study beforehand. 
  • l. 147 Figure 1 only shows the camera
  • section 2.3: Please describe how the quality of the manual extraction of the driver's face area was assessed and warranted. A feasible way might be the use of multiple raters while calculating interrater agreement
  • l. 172: in which cases were "radiometric temperature values not directly accessible" --> is this bug? Is there a systematic error? What does this mean? 
  • l.216: I do not understand this sentence
  • section 2.7: 
    • there are some typos in the paragraph
    • this is a very general description of the KSS, please also describe when and how the KSS was applied during the driving sessions 

Results: 

  • l. 275 --> this information on the temperature contradicts table 1 (which says -1° to 0° C)
  • l. 280 ff: in the paragraph potential drowsiness features exhibited by the driver are mentioned, but I am missing a quantitative description/analysis of these as ground truth for the thermal camera.
    • how often and when did these occur? If we do not know this, it this hard to judge the thermal cam's feasibility in detecting these 
    • there needs to be some assessment when these occur,
      • e.g. rated and protocolled by a co-driver during the drive.
      • Alternatively you can record the pose by automated (or manual) analysis of a video feed of non-thermal "ground truth" cam.
      • Another alternative would be to only mimic the poses, e.g. while standin,g to see how well these can be detected by the thermal cam. 
    • Figure 6 and 7 show only one session. Why not all 6 (or 3)? 
    • The KSS values in table 3 look almost too good to be true. Were there no variations between the days? Please show the individual traces of each drives. To add, were all drives really exactly 60 minutes? 

Discussion: 

  • the discussion can in my view only be judged after methods and results have been improved.

Tables & Figures: 

  • Numbering and order is odd, e.g. Table 3 is the first to be mentioned

Ethical concern: 

  • I have some ethical concerns with the sleepiness scores and the driver's traffic participation in the night drives 
  • the Karolinska sleepiness scores were very high for the night drives, so that it may be that the driver shouldn't be driving in order not to be a danger for other road participants. This is even substantiated by the claim in l.285 saying that the driver's eye were predominantly oriented downwards
  • see also Figure 8 which shows a really unattentive driver - this can be very dangerous at 100 km/h
  • I would feel much better if the university's ethic committee would consider the study
Comments on the Quality of English Language

there are some typos in the manuscript.

Author Response

Comment:

Research design not sufficient; only one participant; self-reported sleepiness; lack of ground truth.

Response:

The manuscript has been fundamentally reframed. All sleepiness-related interpretation and KSS-based evaluation have been removed. The study no longer claims behavioral detection but focuses exclusively on radiometric temperature stability. A second participant has been included, each completing three independent sessions. Quantitative temporal stability metrics (slope analysis, coefficient of variation, Pearson correlation, and RMSE) have been introduced to assess reproducibility. The scope is now clearly methodological rather than behavioral.

Comment:

Unclear number of sessions; confusion between 3 and 6 sessions.

Response:

The session structure has been clarified. Six total sessions were conducted (three per participant), each lasting exactly 60 minutes. This is explicitly stated in the revised Methods section and summarized in Table 3.

Comment:

Mathematical consistency regarding duration and distance.

Response:

Ambiguous statements regarding estimated driving distance have been removed. Session duration and contextual parameters are now explicitly documented to avoid interpretative calculations by the reader.

Comment:

Manual face extraction quality unclear.

Response:

The study now focuses on ROI-based temperature stability rather than pose classification accuracy. ROI definitions are consistently applied across sessions, and the analysis is strictly quantitative. Behavioral inference has been removed.

Comment:

Radiometric values not directly accessible.

Response:

The limitation of display-level extraction for UTi260M is now explicitly described in Section 2.3. The methodological implications are acknowledged in the Limitations section.

Comment:

Lack of quantitative assessment of poses.

Response:

A stationary dual-camera validation procedure has been added. Simultaneous RGB and LWIR recordings under controlled conditions verify geometric consistency of head positions. This serves as qualitative geometric validation rather than behavioral ground truth.

Comment:

KSS values and ethical concern regarding sleepiness.

Response:

All KSS-related analysis and sleepiness interpretation have been removed. The study does not induce or assess driver fatigue. Measurements were conducted under controlled and safe conditions.

Comment:

Figure numbering and order inconsistent.

Response:

All figures and tables have been renumbered and reorganized in logical order. Cross-referencing within the text has been verified.

Comment:

Typos.

Response:

The manuscript has undergone careful language revision and formatting correction.

Back to TopTop