Processing of Amplitude-Temporal Acoustic Parameters in the Auditory System During Signal Coding for Image Recognition: Analytical Review
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe manuscript proposes an analytical review of auditory perception mechanisms, with particular attention to the encoding of amplitude-temporal parameters of acoustic signals and their transformation into auditory images. The article addresses heterogeneous fields such as psychophysics, psychophysiology, psychopathology, cognitive neuroscience, and applications in HCI and machine learning. The authors argue that these parameters facilitate discrimination but interfere with recognition, highlighting the role of higher cognitive processes (memory, verbalisation, emotions). Below are the observed findings.
1. How do the authors explain the contradiction between studies that support and deny the preservation of structural properties in auditory signals? How do the authors explain the contradiction between studies that support and deny the preservation of structural properties in auditory signals? Is it possible to propose a unified metric?
2. Why is no clear taxonomy of the analysed cognitive models provided, considering the variety of domains addressed (EEG, fMRI, psychopathology)? Why is no clear taxonomy of the analysed cognitive models provided, considering the variety of domains addressed (EEG, fMRI, psychopathology)?
3. How do the authors justify the absence of a systematic methodology (e.g., PRISMA) for the selection of literature, considering that the work is presented as an "analytical review"? How do the authors justify the absence of a systematic methodology (e.g., PRISMA) for the selection of literature, considering that the work is presented as an "analytical review"? The inclusion of methodologies for analysing complex signals, such as the study doi:10.3390/signals6030038, could strengthen the methodological framework, especially in the section on the coding of amplitude-time parameters.
4. Can the authors clarify the original contribution of the work compared to previous reviews? Currently, the manuscript appears more as a descriptive summary than a critical one.
5. Can the authors clarify how they intend to formally integrate the mentioned machine learning models, given that no architecture, pipeline, or experimental comparison is provided?
6. What is the theoretical justification for the implicit analogy between acoustic coding and other sensory modalities without a shared quantitative framework? The inclusion of integrated AI-signal processing approaches, as discussed in the study doi: 10.3390/eng7010028, could support a more rigorous formalisation of the concept of "auditory image."
7. How is the central claim that amplitude-temporal parameters facilitate discrimination but hinder recognition validated? Are comparative quantitative evidences available?
8. Why are the sections on HCI and BCI marginal and not integrated with the rest of the theoretical discussion?
9. Can the authors clarify the operational distinction between "auditory image," "auditory representation," and "mental imagery," which are used interchangeably in the text?
10. How is inter-individual variability managed in the proposed models? No statistical or computational model emerges to describe it.
11. The section on emotions appears disconnected from the central theme of signal encoding: is it possible to integrate it into a coherent framework or scale it down?
12. How do the authors intend to address the lack of objective behavioural evidence in the formation of auditory images, a problem explicitly acknowledged in the text?
Author Response
Dear Reviewer!
I would like to express my gratitude for the valuable work you did on the manuscript. All 12 comments were assessed, analyzed, and responded to. In most cases, the responses were integrated with the text. A number of points were clarified.
The manuscript proposes an analytical review of auditory perception mechanisms, with particular attention to the encoding of amplitude-temporal parameters of acoustic signals and their transformation into auditory images. The article addresses heterogeneous fields such as psychophysics, psychophysiology, psychopathology, cognitive neuroscience, and applications in HCI and machine learning. The authors argue that these parameters facilitate discrimination but interfere with recognition, highlighting the role of higher cognitive processes (memory, verbalisation, emotions). Below are the observed findings.
1. How do the authors explain the contradiction between studies that support and deny the preservation of structural properties in auditory signals? How do the authors explain the contradiction between studies that support and deny the preservation of structural properties in auditory signals? Is it possible to propose a unified metric?
A: The formation of auditory images is quite subjective, even compared to visual perception. Although the set of amplitude-temporal characteristics of acoustic parameters has been well-studied and described in scientific and practical research, the subsequent transformations into auditory images are quite controversial and often contradictory. This is due both to the lack of auditory detectors (similar to visual detectors) and to the fact that the formation of auditory images in the brain is often viewed through visual images, since auditory images require several seconds to form, while visual images require thousandths and hundredths of a sec.
Why is no clear taxonomy of the analysed cognitive models provided, considering the variety of domains addressed (EEG, fMRI, psychopathology)? Why is no clear taxonomy of the analysed cognitive models provided, considering the variety of domains addressed (EEG, fMRI, psychopathology)?
A: Overall, the arsenal of objective methods for assessing auditory signal processing in the human brain is quite limited. The most commonly used methods (ERPs, EEG, fMRI) are capable of assessing the spatiotemporal characteristics of neural network excitation under such conditions in different regions of the cerebral cortex. Differences among these methods lie in improved performance in assessing either spatial or temporal characteristics [4]. Event-related potentials (ERPs), specifically waves with varying latencies and different psychophysiological significance, occupy an intermediate position. Other methods (PET scanning, MEG of auditory EPs) stand apart. However, given their complexity and expense, these methods provide little information for auditory perception. Intraoperative monitoring of auditory EPs in neurosurgery provides useful information at the precortical signal processing stage, where waves are clearly linked to neurons of the auditory pathway. However, near-field EEPs under such conditions do not reflect signal processing in the cerebral cortex, especially since the patient is unconscious.
How do the authors justify the absence of a systematic methodology (e.g., PRISMA) for the selection of literature, considering that the work is presented as an "analytical review"? How do the authors justify the absence of a systematic methodology (e.g., PRISMA) for the selection of literature, considering that the work is presented as an "analytical review"? The inclusion of methodologies for analysing complex signals, such as the study doi:10.3390/signals6030038, could strengthen the methodological framework, especially in the section on the coding of amplitude-time parameters.
A: The planning and goal of this review were guided by the current classification of scientific reviews. While systematic and meta-analyses using AI exist, narrative reviews based on the authors' experience are also relevant, as are critical reviews that not only summarize existing knowledge but also evaluate it, identifying gaps, contradictions, and limitations. Therefore, the objective was chosen in accordance with the objectives of a critical review, to highlight current controversies and issues in auditory image processing.
We also express our gratitude for the opportunity to review and cite the article from Signals.
Can the authors clarify the original contribution of the work compared to previous reviews? Currently, the manuscript appears more as a descriptive summary than a critical one.
A: Taking this comment into account, it was decided to formulate a conclusion to highlight the main controversial issues and make some changes to the Introduction. Given the diversity of scientific articles on auditory perception, these works, even review papers, are tied either to acoustics, psychoacoustics, psychophysics, or neurophysiology, psychology, clinical ear diseases, or psychiatry. Therefore, given the vast number of works on visual perception, the assessment of auditory perception is clearly more modest, as a number of contradictions have accumulated that need to be classified and analyzed.
Yes, the manuscript may resemble a "descriptive summary," but it emphasizes contemporary controversial issues regarding the transformation of acoustic signals into auditory images at various stages. This information is reflected in the structure of the review.
Can the authors clarify how they intend to formally integrate the mentioned machine learning models, given that no architecture, pipeline, or experimental comparison is provided?
A: Information has been added to the relevant chapter. While EEG analysis is integrated to some extent into machine learning models, the use of ML models for auditory perception research is considered in isolation in most studies, using separate methods. This issue is being explored most fully in BCI systems for creating support systems for individuals with complex auditory impairments, but one of the criteria—the P300 wave of auditory EPs, the alpha rhythm, etc.—is not fully understood.
What is the theoretical justification for the implicit analogy between acoustic coding and other sensory modalities without a shared quantitative framework? The inclusion of integrated AI-signal processing approaches, as discussed in the study doi: 10.3390/eng7010028, could support a more rigorous formalisation of the concept of "auditory image."
A: The analogy between acoustic signals and other sensory modalities, typically visual or verbal imagery, in neurophysiology and psychophysiology is approximate. Using neural activity data obtained in animals, and also partially and very limitedly in neurosurgical patients (e.g., intraoperative monitoring of auditory potentials), as well as data from various psychological studies, this information is interpreted as approaching the auditory processes of healthy humans. Therefore, a common quantitative basis in this case is also a contentious issue.
Also, thanks for the reference Eng, which was reviewed and incorporated into the paper.
How is the central claim that amplitude-temporal parameters facilitate discrimination but hinder recognition validated? Are comparative quantitative evidences available?
A: The thesis has been added to the paper. This is a well-known thesis from sensory physiology: at the receptor level and at the levels of the first neural centers (auditory or other sensory pathways), amplitude-temporal characteristics are clearly encoded as receptor potential, generator potential, and action potential, as demonstrated in numerous animal studies. All controversial signal processing mechanisms are linked to cortical-subcortical relationships in the formation of images, which depend on a number of factors—individual experience, education, personality traits (including emotional sphere), past illnesses, injuries, etc. Comparative quantitative studies primarily involve the use of ML and AI methods.
Why are the sections on HCI and BCI marginal and not integrated with the rest of the theoretical discussion?
A: The sections on HCI and BCI are not secondary; however, the use of these approaches in research in sensory physiology is typically considered separately. For example, long-term participation in the HCII and AHFE conferences also demonstrates the isolated presentation of these research results. This review presents and categorizes such information as a separate chapter. However, the application and interpretation of HCI and BCI in auditory perception research is limited (isolated).
Here I will partially reiterate my response to comment 5. While EEG analysis is integrated to some extent into machine learning models, the use of ML models for auditory (and sensory perception in general) research is considered in most studies in isolation, employing separate methods. This issue is most fully explored in BCI systems for creating assistance systems for individuals with complex visual, auditory, and motor impairments, but one criterion—the P300 wave of auditory EPs, the alpha rhythm, etc.—is considered incomplete.
Can the authors clarify the operational distinction between "auditory image," "auditory representation," and "mental imagery," which are used interchangeably in the text?
A: As the title suggests, the paper's key operational concepts are "acoustic signals" (signals from the environment) and auditory images, which are formed in the secondary projection area of ​​the auditory system and the association cortex. Figure 1 and its description are designed to categorize these concepts. Interchangeability is possible to avoid excessive repetition.
How is inter-individual variability managed in the proposed models? No statistical or computational model emerges to describe it.
A: Interindividual variability is a distinct field in psychology, psychophysiology, and neurophysiology. It is a characteristic of higher-level behavioral acts. For this purpose, this review includes Chapter 2.4, which addresses emotions as an essential characteristic of personality properties.
The section on emotions appears disconnected from the central theme of signal encoding: is it possible to integrate it into a coherent framework or scale it down?
A: On the one hand, emotions are a personal characteristic, and on the other, emotions are considered alongside mental abilities such as perception, memory, mnemonics, coding, and so on. In the review structure, this chapter follows the listed attributes. We believe it would be inappropriate to simplify this information, as such studies are presented in studies of auditory perception.
How do the authors intend to address the lack of objective behavioural evidence in the formation of auditory images, a problem explicitly acknowledged in the text?
A: In sensory physiology, the lack of objective (digital) behavioral evidence is a well-known problem, which has been addressed to a certain extent through animal studies. However, the question of how humans form images of any modality remains.
Reviewer 2 Report
Comments and Suggestions for AuthorsProcessing of Amplitude-Temporal Acoustic Parameters in the Auditory System during Signal Coding for Image Recognition: Analytical Review
1.
The introduction covers several areas — psychophysics, psychophysiology, sychopathology, human-machine interaction and machine learning — but the underlying scientific thread remains somewhat unclear. Could you clarify exactly what gap in the literature your review seeks to fill, and provide a more rigorous justification for why these very different dimensions need to be brought together within a single analytical framework?
2.
You present auditory imagery as a process involving both the encoding of amplitude-temporal parameters, memory, subvocalisation, semantic components and, at times, other perceptual modalities. How do you theoretically articulate these mechanisms within a coherent model of auditory perception, and what, in your view, are the main contradictions or limitations of previous work that your article helps to clarify?
3.
The section Methodology and concept of analytical review presents a neurophysiological conceptual framework for auditory perception, but the methodology of the analytical review remains largely unexplained. Could you specify the criteria used to select, compare and organise the studies, and clearly indicate whether this is a narrative, analytical, systematic or integrative review? As it stands, the objective is stated, but the review protocol is not sufficiently defined to guarantee methodological rigour and the reproducibility of the analysis.
4.
You propose a sequence ranging from the amplitude-temporal parameters of the signal to the formation, recognition and misrecognition of auditory images via various neuroanatomical structures. What is the precise empirical basis for establishing this functional schema as a coherent synthetic model, particularly for distinguishing the respective roles of primary and secondary auditory areas and frontal associative regions? It would be useful to clarify what constitutes an established consensus in the literature and what remains hypothetical or subject to debate.
5.
The section on psychopathology highlights several links between auditory hallucinations, schizophrenia, musical hallucinations, auditory imagery and dysfunction in cortical or subcortical regions. However, the synthesis remains somewhat juxtapositional at times. Could you clarify exactly which unified mechanistic model you advocate to explain the origins of auditory hallucinations, and distinguish more explicitly between clinical associations, neuroimaging correlations and causally proven relationships?
6.
In the sections on the structural and temporal properties of auditory images, as well as on their facilitation or interference effects, you report results that are sometimes contradictory. Could you propose a more structured explanatory framework to reconcile
these discrepancies, for example based on the type of task, the complexity of the
stimulus, the level of attention required, or the brain regions involved? As it stands, the reader understands that there is a controversy, but is less clear on its exact sources and how your review helps to organise them.
Author Response
Dear Reviewer!
I would like to express my gratitude for the valuable work you did on the manuscript. All 6 comments were assessed, analyzed, and responded to. In most cases, the responses were integrated with the text. A number of points were clarified.
Processing of Amplitude-Temporal Acoustic Parameters in the Auditory System during Signal Coding for Image Recognition: Analytical Review
1.
The introduction covers several areas — psychophysics, psychophysiology, sychopathology, human-machine interaction and machine learning — but the underlying scientific thread remains somewhat unclear. Could you clarify exactly what gap in the literature your review seeks to fill, and provide a more rigorous justification for why these very different dimensions need to be brought together within a single analytical framework?
A: The planning and goal of this review were guided by the current classification of scientific reviews. While systematic and meta-analyses using AI exist, narrative reviews based on the authors' experience are also relevant, as are critical reviews that not only summarize existing knowledge but also evaluate it, identifying gaps, contradictions, and limitations. Therefore, the objective was chosen in line with the tasks of a critical review: to identify current controversies and issues in auditory image processing. With the diversity of scientific articles on auditory perception, these works, even review works, are tied either to acoustics, psychoacoustics, psychophysics, neurophysiology, psychology, clinical ear diseases, or psychiatry. Therefore, given the vast number of studies on visual perception, the assessment of auditory perception is clearly more modest.
2.
You present auditory imagery as a process involving both the encoding of amplitude-temporal parameters, memory, subvocalisation, semantic components and, at times, other perceptual modalities. How do you theoretically articulate these mechanisms within a coherent model of auditory perception, and what, in your view, are the main contradictions or limitations of previous work that your article helps to clarify?
A: The digitization and comparative analysis of sensory processes is a well-known thesis in sensory physiology, whereby at the receptor level and at the levels of the first neural centers (auditory or other sensory pathways), amplitude-temporal characteristics are clearly encoded as receptor potential, generator potential, and action potential, as demonstrated in numerous animal studies. All controversial signal processing mechanisms are linked to cortical-subcortical relationships in the formation of images, which depend on a number of factors—individual experience, education, personality traits (including emotional sphere), past illnesses, injuries, etc. Comparative quantitative studies primarily involve the use of ML and AI methods.
3.
The section Methodology and concept of analytical review presents a neurophysiological conceptual framework for auditory perception, but the methodology of the analytical review remains largely unexplained. Could you specify the criteria used to select, compare and organise the studies, and clearly indicate whether this is a narrative, analytical, systematic or integrative review? As it stands, the objective is stated, but the review protocol is not sufficiently defined to guarantee methodological rigour and the reproducibility of the analysis.
A: Despite the diversity of scientific articles on auditory perception, these studies, even review works, are focused either on acoustics, psychoacoustics, psychophysics, or neurophysiology, psychology, ENT clinics, or psychiatry. Therefore, given the vast number of studies on visual perception, the assessment of auditory perception is clearly more modest, but it contains a number of contradictions in the assessment of auditory image formation.
The planning and goal of this review were guided by the current classification of scientific reviews. While systematic and meta-analyses using AI exist, narrative reviews based on the authors' experience are also relevant, as are critical reviews that not only summarize existing knowledge but also evaluate it, identifying gaps, contradictions, and limitations. Therefore, the objective was chosen in accordance with the objectives of a critical review: to identify current contradictions and issues in auditory image processing.
4.
You propose a sequence ranging from the amplitude-temporal parameters of the signal to the formation, recognition and misrecognition of auditory images via various neuroanatomical structures. What is the precise empirical basis for establishing this functional schema as a coherent synthetic model, particularly for distinguishing the respective roles of primary and secondary auditory areas and frontal associative regions? It would be useful to clarify what constitutes an established consensus in the literature and what remains hypothetical or subject to debate.
A: The neuroanatomy and neurophysiology of the auditory pathway are fairly well-researched. Obviously, for this reason, the review did not detail this information, limiting it to the original diagram and the processes associated with neural structures (Figure 1). A well-known thesis from sensory physiology states that at the receptor level and the levels of relay neural centers (of the auditory or other sensory pathways), amplitude-temporal characteristics are clearly encoded as receptor potential, generator potential, and action potential, as demonstrated in numerous animal studies of neural activity. In humans, auditory processing at the precortical level is assessed based on recordings of brainstem auditory evoked potentials, which is possible both in the waking state and during neurosurgery. All controversial signal processing mechanisms are associated with cortical-subcortical relationships in the formation of images, which depend on a number of factors—individual experience, education, personality traits (including emotionality), past illnesses, injuries, etc.
5.
The section on psychopathology highlights several links between auditory hallucinations, schizophrenia, musical hallucinations, auditory imagery and dysfunction in cortical or subcortical regions. However, the synthesis remains somewhat juxtapositional at times. Could you clarify exactly which unified mechanistic model you advocate to explain the origins of auditory hallucinations, and distinguish more explicitly between clinical associations, neuroimaging correlations and causally proven relationships?
A: This is a very interesting question. Given the collection of current controversies in auditory perception research, the information in this section is quite relevant yet compact.
The purpose of studying cognitive processes in brain pathology lies not only in its diagnostic value, although this is undoubtedly a priority. Another valuable aspect is the fact that similar changes in executive functions, under certain conditions, can also be observed in apparently healthy individuals. Based on this, an analysis of current understanding of auditory perception abnormalities such as auditory hallucinations, amusia, and subvocalization in schizophrenia and borderline disorders has been conducted.
Thus, on the one hand, disorganization of neuronal network activity is known, leading, among other things, to increased activity in the auditory tract nuclei and involuntary hallucinations. However, on the other hand, data has been obtained showing a lack of evidence that exceptionally vivid or exceptionally faint auditory images are associated with the presence of auditory hallucinations in schizophrenia. One possible explanation is that auditory hallucinations in schizophrenia arise from disturbances in the analysis of verbalization processes from visible, invisible sources, and inner speech.
6.
In the sections on the structural and temporal properties of auditory images, as well as on their facilitation or interference effects, you report results that are sometimes contradictory. Could you propose a more structured explanatory framework to reconcile
these discrepancies, for example based on the type of task, the complexity of the
stimulus, the level of attention required, or the brain regions involved? As it stands, the reader understands that there is a controversy, but is less clear on its exact sources and how your review helps to organise them.
A: This question also confirms the relevance of the controversies that exist in the study of auditory perception. In visual perception, beginning with amphibians (frogs), the functioning of detector neurons for processing all sorts of small details around us, as well as afferent systems (within the entire visual system) for perceiving a large number of frequencies of different colors, has been proven. Numerous attempts have been made to transfer these mechanisms to the analysis of auditory information. But numerous attempts have only uncovered new contradictions. Facilitation and interference have a specific meaning in neurophysiology. Facilitation is an increase in amplitude (of the auditory EP) and, as a result, an increase in the sound pressure level of the auditory image in the primary projection cortex. Interference, according to the laws of physics, can result in either an increase in amplitude or a reduction. In auditory perception, interference is often used to denote the separation of a signal from "noise," which can ultimately lead to one effect or another.
My answer further corresponds to some extent to the answer to question 4. At the receptor level and the levels of relay neural centers, amplitude-temporal characteristics are clearly encoded as receptor potential, generator potential, and action potential, as demonstrated in numerous animal studies of neural activity. In humans, auditory processing at the precortical level is assessed based on recordings of brainstem auditory EPs, which is possible both in the waking state and during neurosurgery. All conflicting signal processing mechanisms are linked to cortical-subcortical relationships in image formation, which depend on a number of factors.
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsThe authors have given excellent answers to the review comments and have revised the paper well.
