Next Article in Journal
Porosity and Pore-Network Controls on Elastic Properties and Permeability in Porous Ignimbrites
Previous Article in Journal
Exploratory Evaluation of Diagnostic Accuracy and Temporal Reproducibility of Multimodal Large Language Models in the Image-Based Assessment of Oral Mucosal Lesions
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Processing of Amplitude-Temporal Acoustic Parameters in the Auditory System During Signal Coding for Image Recognition: Analytical Review

Department of Normal Physiology, St. Petersburg State Pediatric Medical University, 194100 Saint Petersburg, Russia
Appl. Sci. 2026, 16(8), 4047; https://doi.org/10.3390/app16084047
Submission received: 20 March 2026 / Revised: 17 April 2026 / Accepted: 20 April 2026 / Published: 21 April 2026
(This article belongs to the Special Issue Cognitive, Affective and Behavior Neuroscience)

Abstract

In the study of sensory processes, the visual system has received the most research compared to other sensory systems. The primary difference between visual and auditory perception lies in the nature of the stimuli and the reception processes: vision perceives electromagnetic radiation, while auditory perception perceives acoustic signals of mechanical origin. This review aims to analyze modern approaches and controversies to the mechanisms of auditory perception related to psychophysics, psychophysiology, psychopathology, modern research on hearing in human–computer interaction (HCI) systems, and machine learning methods. Modern studies of acoustic patterns include a comprehensive assessment of the physical characteristics of perception, complex nonverbal auditory cues, verbalization, perception and memory, as well as individual differences in auditory perception. An analysis of the scientific literature allowed us to conclude that acoustic signals transformed in the brain into auditory images retain (encode) a number of amplitude-temporal parameters of acoustic signals that facilitate auditory discrimination (filtering), but interfere with auditory detection (recognition). Signal processing often, but not necessarily, involves brain regions involved in other forms of perception. It depends on subvocalization, includes semantically interpreted information and expectations, pictorial (visual) and descriptive components, functions as a mnemonic, and is linked to individual musical ability and experience (although the mechanisms of this connection are unclear).

1. Psychophysics of Acoustic Signals

1.1. Acoustic Patterns of Mechanical Vibrations from the Environment: Modern Empirical Data

Despite the revival of visual imagery research that began in the late 1960s and early 1970s, patterns of acoustic signals and transformation into auditory images did not attract much scientific interest for a long time. Most research focused on visual and spatial imagery [1,2,3,4,5]. However, this did not mean a complete lack of interest in other (non-visual) forms of imagery, which are essential in everyday life and without which holistic perception is impossible [6,7,8]. Despite the diversity of scientific research on auditory perception, these studies, even review papers, are focused either on acoustics, psychoacoustics, psychophysics, or neurophysiology, psychology, clinical ear diseases, or psychiatry. Therefore, given the vast number of studies on visual perception, the assessment of auditory perception is clearly more modest, and a number of contradictions have accumulated that need to be classified and analyzed.
Contemporary studies of acoustic patterns include a comprehensive assessment of physical characteristics of perception (pitch, timbre, and loudness), complex nonverbal auditory cues (musical contour, melody, harmony, tempo, musical audiology, and environmental sounds), verbalization (speech, text, sleep, and internal monologue), perception and memory (detection, encoding, reminder, mnemonic properties, and phonological loop), and individual differences in auditory perception (vividness, musicality, ability, experience, synesthesia, musical hallucinosis, schizophrenia, and amusia) [9,10].
Acoustic signals are transformed in the brain into auditory images with the storage (coding) of a number of amplitude-temporal parameters of acoustic signals, which facilitate auditory discrimination (filtering), but interfere with auditory detection (recognition), and involve brain areas involved in other forms of perception. Often, but not necessarily, auditory imagery depends on subvocalization, includes semantically interpreted information and expectations, including visual and descriptive components of other modalities, functions as mnemonics, and is associated with individual musical abilities and experience, although the mechanisms of this connection have not been established [11,12,13,14].
To define the terminology of “auditory images”, we turn to the statement that they are considered “the introspective stability of auditory experience, created from components of long-term memory in the absence of direct sensory stimulation of this experience.” It should be noted that this definition is similar to the characteristics used for images of other modalities [3,5]—visual, olfactory, etc.
The perception of auditory images, unlike visual ones, is more subjective. This is because the researcher cannot directly observe the patterns (images) that form in the subject’s brain. The characteristics and properties of images are assessed using indirect measures, which are assumed to be subject to a number of predictable and systematic influences (e.g., selective filtering and parallel and/or subsequent task solving) [15,16,17].
The outputs assessed typically include subjective reports from participants [17,18], comparisons of performance on tasks in which participants are instructed to generate auditory imagery across tasks [19,20] and in which participants are not instructed to generate auditory imagery, brain imaging studies that assess areas of cortical activation during the auditory imagery task [21,22], and clinical data regarding psychopathology or other conditions that involve auditory imagery and cognitive processes [23,24].
The vast majority of studies of acoustic signals focus on the physical structural properties of signals (parameters), including elementary or basic characteristics such as pitch, timbre, and loudness [25]. One classic task is the presentation of relevant and deviant signals, which we will discuss below when considering mismatch negativity (MMN).
One of the main questions is whether loudness is encoded during the generation of auditory images. In one study, participants estimated the typical loudness of sounds associated with each phrase and then generated images of the sounds referenced by the phrases. The time required to generate the images was unrelated to loudness parameters. It has been suggested that the encoding of auditory images differs fundamentally from that of visual images, with loudness encoding not being required for recognition [26,27,28].
Subsequent studies demonstrated loudness encoding, but with the involvement of the visual system in processing, through which two phrases were presented. Subjects were instructed to form an auditory image using the phrase on the left, and then to form an auditory image from the phrase on the right. Subjects were then required to subjectively adjust the loudness of the second sound to match the subjective loudness of the first image. Under these experimental conditions, reaction time increased simultaneously with the difference in the typical loudness levels of the two objects [29,30].
In other conditions, subjects were asked to synthesize auditory images corresponding to pairs of acoustic signals, followed by a subjective loudness assessment. A reduction in reaction time was recorded, while the differences in the base loudness of the auditory images increased. The authors concluded that signals may contain loudness information when comparing and evaluating images. However, in the auditory image generation conditions, they suggested that loudness information is not always processed (or synthesized gradually) during the formation of auditory images. Thus, the results of this study, which involved comparing and assessing subjective loudness equivalents, were also consistent with similar studies [6,7].
In life, we typically encounter complex acoustic patterns that are transformed into auditory images involving multiple sensory systems. Therefore, processing involves combinations of several modalities to synthesize the final recognition version. Such acoustic patterns include, for example, specific musical contours, familiar and unfamiliar melodies, harmony, tempo and duration, musical notation, the sounds of flora and fauna, as well as objects and events in the environment [27,31,32].
An objective method for studying auditory perception is the recording of auditory evoked potentials (AEPs) [31,32], the generation of whose components over a time interval of 10–50 ms is clearly linked to neurons of the auditory pathway (Figure 1, 4–6). Thus, when presenting ascending or descending melodic phrases consisting of eight notes, the melodic phrase was perceived as a whole. Then, the first three or five notes were presented, and the subjects had to imitate the missed notes. The study was carried out with various deviations, with a hint (three notes were presented with five notes) or, conversely, five notes (with the presentation of three notes). Sometimes the conditions were not met at all (subjects expected to hear five notes, but heard only three) [33]. In such conditions, subjects were asked to press a button, expecting to hear the last note of the sequence. Visualization of the continuation, as well as the anticipation of the missed note, were accompanied by waves similar to the auditory evoked potentials of perceived notes. It was concluded that the similarity of patterns during the analysis epoch of the N100 component of the ERP during the anticipation of a note that was missed or actually perceived is consistent with the hypothesis that anticipation in the auditory system and holistic auditory perception activate similar brain mechanisms [31,32].
Overall, the arsenal of objective methods for assessing auditory signal processing in the human brain is quite limited. The most commonly used methods (ERPs, EEG, and fMRI) are capable of assessing the spatiotemporal characteristics of neural network excitation under such conditions in different regions of the cerebral cortex. Differences among these methods lie in improved performance in assessing either spatial or temporal characteristics [4]. Event-related potentials (ERPs), specifically waves with varying latencies and different psychophysiological significance, occupy an intermediate position. Other methods (PET scanning and MEG of auditory EPs) stand apart. However, given their complexity and expense, these methods provide little information for auditory perception. Intraoperative monitoring of auditory EPs in neurosurgery provides useful information at the precortical signal processing stage, where waves are clearly linked to neurons of the auditory pathway. However, near-field EPs under such conditions do not reflect signal processing in the cerebral cortex, especially since the patient is unconscious.
The temporal accuracy of auditory recognition is significantly reduced compared to the temporal precision of holistic auditory perception, which is associated with the involvement of other modalities in holistic perception [32]. Typically, identifying an auditory image requires at least several seconds (sometimes tens of seconds), while holistic auditory perception is activated immediately upon the appearance of signals. In comparison, visual recognition takes thousandths/hundredths of a second [34,35]. Study participants were presented with fragments of well-known melodies [11,12]. They were required to concentrate on the first fragment, after which they had to “mentally reproduce” the melody until the next melody played. In a control study, excluding imagery, participants had to judge whether the second melody matched the first. In the imagery condition, reaction time increased with increasing time between the second and first fragments of the melody. It was concluded that the temporal structure of the melody was preserved in the auditory image. Reaction times also increased as one melody was integrated into another, suggesting that participants began scanning from the beginning of the melody, regardless of whether a second melody was present.
The use of lyrics allows for the possibility that participants may have used a form of representation other than auditory imagery (e.g., abstract verbal or positional representation). To examine this possibility, word pairs were presented, and participants judged whether the pitch of the second verse was higher or lower than the pitch of the first verse. As before, response times increased with increasing interword distance [11,12]. The results of the study suggested a time-stretched generation of auditory images [36], which served as a basis for investigating the mechanisms of sequential presentation of acoustic signals. During perception, participants could adjust the tempo on a computer recording. In conditioned imagery, participants were given the name of a familiar melody, instructed on how to reproduce the melody, and then adjusted a metronome to match the tempo of their image. The tempo settings varied across melodies, suggesting that participants were discriminating between melodies. More importantly, the correlation between perceived tempo and displayed tempo for each melody was high (r 5.63), suggesting that sound images retain tempo information.
In general, the auditory modality is considered a good model for investigating global–local processing without involving the parameters that define visuospatial imagery. In particular, music can be hierarchically organized across spatiotemporal scales and parameters, allowing for the investigation of various abilities such as cognitive control, attention, inhibition, and processing levels [37,38]. As in the visual task, the global–local auditory task is designed such that each stimulus consists of both a global and a local dimension, and these dimensions can be congruent or incongruent.
For example, results from studies in ADHD indicate that global–local signal processing in the auditory domain and the absence of a global information processing bias in individuals with ADHD are not limited to the visuospatial modality and reflect a broader and more general multidimensional information processing bias [39]. The fact that the lack of global information processing bias exists outside the visuospatial domain suggests that this abnormal information processing style may play an important role in the ADHD phenotype.

1.2. Methodology and Concept of Analytical Review

The neuroanatomy and neurophysiology of the auditory pathway are fairly well-researched. Obviously, for this reason, this review did not detail this information, limiting it to the original diagram and the processes associated with neural structures (Figure 1). A well-known thesis from sensory physiology states that at the receptor level and the levels of relay neural centers (of the auditory or other sensory pathways), amplitude-temporal characteristics are clearly encoded as receptor potential, generator potential, and action potential, as demonstrated in numerous animal studies of neural activity. In humans, auditory processing at the precortical level is assessed based on recordings of brainstem auditory evoked potentials, which is possible both in the waking state and during neurosurgery. All controversial signal processing mechanisms are associated with cortical–subcortical relationships in the formation of images, which depend on a number of factors—individual experience, education, personality traits (including emotionality), past illnesses, injuries, etc.
Thus, on the one hand, we consider the amplitude-temporal parameters of acoustic signals, which are mechanical vibrations of the environment, and, on the other hand, the transformation of these signals into auditory images. To further discuss the psychophysiology of auditory perception and the brain structures that support processing, we will examine the general diagram of signal transmission in the auditory sensory system (Figure 1).
The formation of auditory images is quite subjective, even compared to visual perception. Although the set of amplitude-temporal characteristics of acoustic parameters has been well-studied and described in scientific and practical research, the subsequent transformations into auditory images are quite controversial and often contradictory. This is due both to the lack of auditory detectors (similar to visual detectors) and to the fact that the formation of auditory images in the brain is often viewed through visual images, since auditory images require several seconds to form, while visual images require thousandths and hundredths of a second.
Mechanical sound waves, characterized by amplitude-temporal parameters, interact with specialized receptors of the organ of Corti—hair cells. Then the signal is transformed and transmitted to neurons of the auditory ganglion (cranial nerve VIII), from which it passes to the cochlear neurons of the brainstem (with 50% of the nerve fibers crossing to the opposite side) and neurons of the superior olive.
Next, psychophysiological mechanisms for the formation of auditory images are connected through the interaction of the switching nuclei of the thalamus (internal geniculate bodies) and neurons of the primary projection zone of the auditory system in the cerebral cortex, where the first subjective components are formed—auditory sensations (volume, frequency, tonality, etc.).
The analogy between acoustic signals and other sensory modalities, typically visual or verbal imagery, in neurophysiology and psychophysiology is approximate. Using neural activity data obtained in animals, and also partially and very limitedly in neurosurgical patients (e.g., intraoperative monitoring of auditory potentials), as well as data from various psychological studies, this information is interpreted as approaching the auditory processes of healthy humans. Therefore, a common quantitative basis in this case is also a contentious issue.
Auditory images are formed in the secondary projection area of the cerebral cortex, located around the primary area in the temporal lobe. The recognition (or misrecognition) of auditory images occurs with the participation of the frontal association cortex. At the same time, integration with other sensory systems is important for image formation, primarily with the visual system, via the nonspecific nuclei of the thalamus, as well as integration with neurons of the brainstem reticular formation and motor and autonomic nerve centers. Cortico-inhibitory processes on brainstem neurons are of great importance for the control of selective attention processes in the auditory system (Figure 1, green dotted lines).
Based on the above, this analytical (narrative) review aims to highlight current understanding of the transformation of environmental amplitude-temporal parameters of various origins into auditory images as they pass through nerve centers of the auditory pathway. To this end, a comparative analysis of contemporary studies and controversies using various methodological approaches, including those in psychopathology and machine learning, was conducted.

2. Psychophysiology of Auditory Perception

2.1. Environmental Sounds

Typically, people subjectively adjust the tonality [40] and loudness [41] of auditory images generated by common environmental cues in accordance with the pitch or loudness of auditory perception. This is determined by the physical characteristics of environmental stimuli. Following this setup, participants were instructed to form an auditory image of the sound emitted by an object named by a visually presented word. It was found that visualized environmental sounds can influence subsequent perceived cues from the same category (e.g., sounds of traffic, nature, household objects, etc.) [42].
One of the tasks in studying the auditory perception of real objects is to compare the correspondence of different sounds in the auditory image of one object with other sounds of the same object [43]. Subjects were asked to perform this task. Response speed was found to be higher when a sound corresponded to the object being evaluated. In the case of visual perception, the following has been demonstrated. If a visual image evokes the correct auditory image of an object, the hypothesis that a pattern resulting from a visual image can stimulate subsequent auditory perception is supported. The degree to which sound representation is sufficient for a visual image to elicit a holistic auditory image remains controversial. This obviously depends on individual life experience.
Similar tasks were addressed using fMRI. Simultaneous viewing of familiar visual scenes and perception of sounds corresponding to these scenes were assessed. Subsequently, viewing of familiar visual scenes was accompanied by instructions regarding the sounds of the images immediately preceding these scenes. In addition, distorted versions of visual scenes were viewed without any accompanying sounds and without any instructions regarding the auditory images [44]. If subjects perceived and categorized the corresponding sounds, bilateral activation was recorded in the primary (Herschle’s gyrus) and secondary projection areas of the auditory cortex (planum temporale). If subjects were asked to subjectively reproduce the corresponding sounds in the forming auditory images, activation was recorded only in the secondary projection area of the auditory cortex. These results are consistent with the data on the activation of the secondary projection area of the auditory cortex in the virtual absence of activation of the primary projection area of the auditory cortex [45] and with greater activation of the associative regions of the auditory system during periods of complete silence embedded in a familiar melody associated with auditory images [46]. Interpretation of these data suggests the conclusion that the formation of a particular auditory image and auditory perception in general are based on the functioning of overlapping neural networks of the secondary projection area of the auditory cortex. However, auditory images are not associated with the structures of the primary projection area of the auditory cortex (Figure 1, neurons 7), which are involved in the formation of auditory sensations [44].

2.2. Verbal Information—Perception and Generation of Speech

A significant portion of human introspective experience is based on verbal information, realized primarily through “inner monologue”—the transformation of images of various modalities into a subjective auditory image through the mechanism of “self-talk.” It can be argued that auditory verbal images are a fundamental component of consciousness [11,47,48]. These processes are checked during traditional conversations, interviews, interrogations, etc.
EEG and event-related potentials during verbalization of syllables and while listening to instructions for their mental representation as visual images were recorded [19]. Presentation of the auditory signal and formation of the auditory image were synchronized over a period of time with the presentation of the visual stimulus. Subjects were required to press a button simultaneously when the auditory perception or image changed. Verbalized instructions for creating auditory images contributed to the formation of a pronounced component of the auditory evoked potential N1 in the interval of 109–143 ms [3,19], as well as a late positive complex (LPC) with a peak latency of 290–330 ms [49]. The topology of the EEG response was consistent with the hypothesis that the N1 in auditory imagery is associated with activity in the anterior temporal regions, and that the LPC in auditory imagery is associated with activity in the cingulate gyrus, medial frontal regions, and bilateral anterior temporal lobes. However, these findings were not supported by a similar study [19], as no convincing behavioral evidence was found to support the creation of auditory imagery by participants, nor was there any evidence that EEG patterns in the imagery context were involved in reflecting the unique processes of auditory imagery.
PET scans showed similar results to previous studies, showing increased activity in the left inferior frontal gyrus during subjective presentation of auditory images in the brain. Furthermore, auditory image processing is also associated with increased activity in the left premotor and left temporal cortices [50]. Broca’s area in the left hemisphere has been shown to exhibit increased activity during covert speech generation and the application of transcranial magnetic stimulation (TMS) to the left hemisphere, which can disrupt the formation of ordinary speech [51]. With this in mind, the question was raised whether the effects of TMS on the left hemisphere could also affect internal (covert) speech activity [52]. That is, is it possible for TMS to disrupt the formation of verbal auditory images? Participants were asked to count the number of syllables in visually displayed word chains. Reaction time increased and was directly related to the number of syllables. Targeted application of TMS to the anterior or posterior regions (Broca’s area and motor areas, respectively) of the left hemisphere resulted in speech distortion during direct counting and a significant increase in reaction time during explicit or covert counting. Thus, TMS to the posterior/motor areas of the left hemisphere affected explicit speech, but did not affect covert (inner) speech.
An fMRI study was conducted after instruction to use first-person, second-person, or third-person intra-speech auditory imagery [53]. Inner speech resulted in increased activation of the left inferior frontal/insular cortex, left temporoparietal cortex, right cerebellum, superior temporal gyrus, and supplementary motor area. First-person auditory imagery was accompanied by increased activation of the left insular cortex, precentral gyrus, and lingual gyrus. Bilateral activation of the middle temporal gyrus and posterior cerebellar cortex was observed, as well as changes in the right middle frontal gyrus, inferior parietal lobe, hippocampus, and thalamus [53]. In contrast, second- and third-person auditory imagery elicited greater activation of the supplementary motor area, left precentral and middle temporal gyrus with inferior parietal lobe, and right superior temporal gyrus and posterior cerebellar cortex [53]. Compared to control first-person speech perception, second- and third-person images elicited greater activation of the medial parietal lobe and posterior cingulate cortex. However, the degree of internal representation of different voices was not independently assessed. Therefore, the question remains about the articulatory properties of a subject’s inner voice, which must necessarily differ from the properties of first-, second-, or third-person voices [54].
There is evidence of deteriorating changes in frequency resolution and fine temporal structure during vocoding of speech signals with noise. Thus, the original acoustic signal is bandpass filtered by neurons of the auditory pathway, followed by extraction of amplitude envelopes for the output signals. The superimposed bandpass noise of the corresponding passbands is then modulated by these amplitude envelopes, and the modulated bandpass noises are summed over frequencies [55]. At coarser or, conversely, finer spectral resolution, the addition of periodicity did not lead to a significant improvement in intelligibility. Also, smoothing or exaggeration of contours did not affect the intelligibility of vocoded and discontinuous speech stimuli. Improved intelligibility in this study was achieved by stretching the mosaic segments without introducing any external noise, but in the presence of 10 ms breaks. Thus, the key factor in the present study was the facilitation of auditory grouping, and not the masking potential itself [55]. The interpretation of the obtained results concluded that previously inaccessible speech signals became accessible by stretching the mosaic segments and strengthening the auditory grouping. Subvocalization undoubtedly influences the formation of auditory images. Additionally, comparison of test results under distracted conditions demonstrated the influence of articulation on auditory images [56].

2.3. Perception and Memory: Mnemonics and Coding

Assessing the processing of acoustic signals requires a comparison of incoming signals with information stored in memory. For this purpose, not only complete images but also reference characteristics are compared. For example, a standard pitch was compared with a time interval filled with a quiet or distracting signal, after which a (new, different) sound was presented for comparison [57]. Subjects were asked to rate the conformity of the latter signal to the standard. Specifically, recognition results were found to be better if the interval between the standard and comparison sounds was quiet. One possible explanation is that subjects used auditory images to repeat the pitch of the standard sound. Such images were more easily formed or were maintained in the absence of distracting or competing perceptual stimuli.
The mnemonic (memory-enhancing) properties of images in the visual system have been well studied in both animals and humans. Any visual information is better remembered [58,59]. This is the basis for the method of loci and the visual perception of word combinations to increase the likelihood of memorizing verbal material. The perception of acoustic patterns associated with memory has been studied much less frequently. Such studies can only be conducted in humans, since they involve speech. In particular, auditory mnemonics was assessed by presenting printed verbal cues, audio recordings, and visual images of acoustic patterns against a background of free recall [60]. It was found that recall was better when presented with audio recordings or visual images than when presented with verbal cues. However, recall did not differ when presented with only audio recordings or only visual images.
Subjects were then presented with audio recordings, images, or both. Specifically, it was found that memorization with combined perception did not differ significantly from presentation of only visual images or only audio recordings. The authors suggested [60] that auditory images possess mnemonic value similar to visual ones. Both modalities are processed in the same memory system, since different systems will have combined effects. Using combined auditory and visual images may lead to a greater memorization effect than using a single type of image.
This is a well-known thesis from sensory physiology: at the receptor level and at the levels of the first neural centers (auditory or other sensory pathways), amplitude-temporal characteristics are clearly encoded as receptor potential, generator potential, and action potential, as demonstrated in numerous animal studies. All controversial signal processing mechanisms are linked to cortical–subcortical relationships in the formation of images, which depend on a number of factors—individual experience, education, personality traits (including emotional sphere), past illnesses, injuries, etc. Comparative quantitative studies primarily involve the use of ML and AI methods.
Subsequent studies of auditory perception mnemonics were conducted using sets of printed verbal cues, audio recordings, and visual images, similar to previous studies [61,62]. In a series of experiments, cues were either blocked according to specific categories (musical instruments, vehicles, animals, and tools) or presented in a random order. Memorization of audio recordings and visual images did not decrease. Memorization of both modalities was better compared to verbal cues, for which, similar to audio recordings, memorization improved when certain stimuli were blocked. Thus, a category superiority effect was observed, but blocking did not affect memorization of visual images. The authors hypothesized that nonverbal and nonmusical auditory images are intermediate in characteristics between visual images and verbal cues. Moreover, the memorization pattern for audio recordings was similar to the memorization pattern for verbal cues under blocking influences. Thus, auditory perception involves the processing of resources from both the visuospatial component and the phonological reverberation component of working memory. However, the question remains whether auditory images were necessarily used to encode auditory stimuli [62].

2.4. Influence of Emotions

On the one hand, emotions are a personal characteristic, and on the other, emotions are considered alongside mental abilities such as perception, memory, mnemonics, coding, and so on. It is known that emotional characteristics can enhance or, conversely, impair memory function. This applies to all modalities, including auditory stimuli. Various conceptual models have been proposed to explain the mechanism of emotional experience resulting from auditory stimulation. The general idea of these concepts is that the actual emotion is a multiplicative function consisting of several factors. These factors include signal parameters with acoustic and syntactic properties, the listener’s psychophysiological characteristics (musical experience, personality traits as a constant, and mood as a variable), as well as physiological responses to stimuli and the subjective experience of these responses [63].
Early studies of acoustic characteristics found that mood (major vs. minor), harmonic complexity, and rhythmic articulation (e.g., staccato vs. legato) best predicted pleasantness (major mode). Simpler harmony and staccato articulation were positively correlated with greater pleasantness, respectively. Conversely, faster tempos, more pronounced accentuations (e.g., marcato), and more staccato articulation were positively correlated with higher levels of arousal. Positive linear correlations were found between valence and pitch level, as well as tempo. Valence was higher in musical excerpts with higher pitch and faster tempos. Positive linear correlations were also found between arousal and loudness, tempo, timbre, and pitch level—arousal was higher in time periods with higher loudness, faster tempos, higher pitch, and sharper sounds [63]. Lower pitch levels were associated with more negative valence and higher levels of arousal when using piano excerpts whose pitch levels were systematically varied. Furthermore, gender was found to influence the effect of pitch levels on both valence and arousal. The positive relationship between pitch and valence was stronger in women, whereas the negative relationship between pitch and arousal was observed only in men. Finally, the effect of pitch level on valence was found to depend on other musical characteristics, such as tempo and mode [63].
The relationships between the hierarchical levels of neural centers in the auditory system and auditory emotions are of great importance. Higher-level neural structures have been shown to evoke both tension and relaxation. Overall, a complex interaction between the characteristics of the auditory system, listener characteristics, and auditory sensations has been demonstrated, all of which influence the emotional experience evoked by auditory stimuli. One area of research into emotions in auditory perception is the component of aesthetic perception [64]. In particular, aesthetics in humans consists of a number of basic components, the most fundamental of which are directly aesthetic emotions arising from aesthetic experience in response to the aesthetic appeal or merits of sensory objects or works of art [65,66].
For many centuries, the role of art in providing “catharsis,” representing the purification of the soul through the processing of aesthetic experiences that evoke pleasant feelings, has been conceptualized [64]. In line with this paradigm, event-related potential (ERP) recordings have demonstrated that emotional experiences resulting from aesthetics or beauty can be stronger than emotional experiences evoked by control or neutral stimuli. For example, VEP recordings revealed increased amplitudes of the P1 and P3b components during the perception of attractive faces compared to unattractive ones, indicating stronger emotional experiences and the involvement of emotional and reward pathways in the assessment of facial attractiveness. Auditory EP studies have demonstrated that the amplitude of the late positive potential was greater during the assessment of the beauty of chord sequences compared to the assessment of the correctness of chord sequences, especially in inexperienced subjects. The results allowed us to conclude that there is an enhanced affective and motivational component in the assessment of visual or auditory beauty. Moreover, the perception of auditory beauty is more emotionally charged than the perception of visual beauty. Thus, there is an inextricable link between aesthetics or beauty and emotional (internal) feelings [66,67]. However, an unresolved question remains: how do aesthetic emotions differ from basic or life emotions, as well as from reactions associated with affective evaluation in general?
Regarding aesthetic perception and affective evaluation, it is believed that aesthetic emotions differ not only from ordinary life emotions but also from emotions due to their affective evaluation. Thus, aesthetic perception consists of an assessment of the quality or value (analytical/originality, semantic, typical, and affective) of the perceived object or event. Affective evaluation, however, consists of an analysis of the affective content of the object or event [67]. Attention, similar to emotions, is also filled with endogenous mechanisms that support selective attention. There is no doubt that attention functions within the auditory and visual modalities, taking into account spatial characteristics [68]. It has been established that exogenous spatial acoustic patterns influence visual discrimination, but not vice versa. Thus, subjects are able to respond to two given target patterns regardless of modality and hierarchical level. If they showed a priming effect specific to levels from visual to auditory perception or vice versa, this would support the existence of a common (or at least interactive) attentional mechanism for selecting the range of auditory and visual perception.

3. Acoustic Patterns, Psychopathology, Neurogenerators, Feedback Influences

3.1. Psychopathology

The purpose of studying cognitive processes in brain pathology lies not only in its diagnostic value, although this is undoubtedly a priority. Another valuable aspect is the fact that similar changes in executive functions, under certain conditions, can also be observed in apparently healthy individuals. Based on this, an analysis of current understanding of auditory perception abnormalities such as auditory hallucinations, amusia, and subvocalization in schizophrenia and borderline disorders has been conducted.
In psychopathology, mental illness, borderline mental states, and neuropsychiatric instability in general, the generation of auditory images can occur involuntarily, compulsively, and in states of heightened anxiety [9,53,69,70,71]. Excess, deficit, and other deviations in the generation of auditory images are an important component in some forms of psychopathology, for example, in (musical) hallucinosis, schizophrenia, and (potential) amusia [71,72,73]. A condition in which auditory (musical) images involuntarily affect a person is called musical hallucinosis (MH) [72]. Systematic studies of this condition have been conducted retrospectively using case histories [74]. Estimates of the prevalence of MH in patients with psychopathology vary from 0.16% to 27% [75]. Results from neuroimaging studies are consistent with the hypothesis that MG may result from abnormal spontaneous activity of neurogenerators underlying musical perception and figurative thinking [59,71]. Furthermore, signs of bilateral hearing loss [42] have been described in patients with dysfunction of the right auditory cortex [76], desynchronization of the right auditory cortex [77], meningioma of the right occipital region, and brainstem lesions [69].
Auditory hallucinations are a diagnostic criterion and an important characteristic of cognitive impairment in schizophrenia [9,70]. It has been noted that in such patients, auditory hallucinations are more often verbal in origin than figurative [9]. One possible explanation for this phenomenon is the fact that auditory hallucinations result from abnormal neuronal activity in the primary and secondary projection areas of the auditory system [75,78].
Auditory hallucinations in schizophrenia were initially thought to result from increased neural activity in the auditory cortex [79]. Later, it was shown that in patients prone to hallucinations, the psychophysiological characteristics of auditory images were less vivid compared to patients less prone to hallucinations. Later, it was established that there were no differences in the vividness of auditory image perception between patients prone to hallucinations and those less prone to hallucinations [53]. A literature review [59] concluded that there was no evidence that exceptionally vivid or exceptionally faint auditory images were associated with the presence of auditory hallucinations in schizophrenia. On this basis, the hypothesis that auditory hallucinations in schizophrenia were associated with impaired auditory image perception was rejected [54]. One possible explanation remains the fact that auditory hallucinations in schizophrenia arise due to the inability of patients to distinguish speech from visible and invisible sources. Therefore, they visualize their own speech to an external source [80].
An fMRI study of schizophrenia patients with and without auditory hallucinations showed that hallucinations were associated with increased activity in the right and left superior temporal gyrus, left inferior parietal cortex, and left middle frontal gyrus. A link between cerebral blood flow patterns and hallucinations was suggested, along with the hypothesis that auditory hallucinations represent abnormal activation of auditory pathways and connections [81].
In another comparative study [71], all patients were given tasks that activated subvocalization and phonological storage while analyzing letter strings, judging pitch, and interpreting ambiguous auditory signals. The results of patients prone to auditory hallucinations were found to be no different from those of patients without hallucinations, regardless of the task. The authors concluded that a direct link between inner speech and auditory hallucinations is unlikely. A possible explanation is that auditory hallucinations in schizophrenia result from decreased efficiency of inner speech introspection mechanisms.
PET scanning of patients with schizophrenia with varying susceptibility to auditory hallucinations and a control group [81] revealed the following. When participants imagined their own voice, there were no differences between the groups. When participants imagined a sentence spoken in another person’s voice, patients prone to auditory hallucinations showed reduced activity in the left middle temporal gyrus and rostral supplementary motor area [76].
Similar task conditions were assessed using fMRI data. When all participants visualized their own voice, no differences were observed between groups. However, when participants visualized verbal stimuli from another person’s voice, patients showed reduced activity in the posterior cerebellum, hippocampus, bilateral lentiform nuclei, right thalamus, and middle and superior temporal cortex [53,54]. Results from other studies [53,54,80] are consistent with the hypothesis that verbal hallucinations in schizophrenia are associated with a reduced ability to activate brain regions associated with the control of inner speech. Based on fMRI results from patients and control participants [54], it has been suggested that the motor cortex and brainstem may also be involved in the structures generating hallucinations.
Thus, on the one hand, disorganization of neuronal network activity is known, leading, among other things, to increased activity in the auditory tract nuclei and involuntary hallucinations. However, on the other hand, data have been obtained showing a lack of evidence that exceptionally vivid or exceptionally faint auditory images are associated with the presence of auditory hallucinations in schizophrenia. One possible explanation is that auditory hallucinations in schizophrenia arise from disturbances in the analysis of verbalization processes from visible sources, invisible sources, and inner speech.

3.2. Reflection of Structural (Amplitude) and Temporal Properties of Acoustic Signals in Auditory Images

This issue remains controversial to this day. While the encoding of amplitude-temporal characteristics and the functioning of detectors in the visual system have been studied quite comprehensively, from frogs to humans, many gaps remain in the formation of auditory images. It is believed that auditory images preserve the structural properties of acoustic signals, including pitch, loudness, tonality [82], timbre [36], musical contour [83], and melody [84], as well as interval fractions of musical signals [80], musical tempo [84], and speech rate [25]. Furthermore, auditory images can stimulate subsequent perception based on harmonic relationships, timbre, and categories of words and environmental sounds. This is consistent with the hypothesis that the physical amplitude-temporal parameters of acoustic signals are preserved in images [9,85].
The opposing view does not support the hypothesis that the structural properties of acoustic signals are preserved in auditory images. In particular, it has been shown that subjects can experience difficulty detecting embedded melodies [86] or alternative interpretations of an ambiguous stimulus in auditory images. Furthermore, a significant decrease in pitch accuracy and comparative task reflection has been shown in auditory images [84]. Reproduced loudness does not influence subsequent perception, and, combined with the apparent absence of loudness as a necessary part of the auditory image, this indicates that loudness-related structural properties are not part of the basic architecture of auditory images [87].
There are various ways to discuss this mechanism. One possibility is that the structure of relatively simple information, such as pitch, is preserved. In contrast, the structure of complex information, such as embedded melody, is not maintained [85,86]. Another possible explanation is that the structure of isolated or separable information, such as tonality, is preserved, but this structure, in the case of integrated information (embedded melody), is not preserved separately from the larger structure into which this information is integrated [87,88]. A third way to discuss this is that the full structural information is preserved in the auditory image, but it is weaker or more susceptible to interference for some signals than for others [89]. In general, it is believed that auditory images preserve some, but not all, amplitude-temporal properties of acoustic signals. Rather than assessing the preservation of structural properties of acoustic signals in auditory images using the all-or-nothing principle, it is believed that the problem should be approached differently. In particular, it should be studied whether additional structural properties are preserved as an integral pattern or whether the task itself influences the presence of structural properties in the image [90,91].
A line of research supports the hypothesis that the temporal properties of acoustic signals are preserved in auditory images. However, more time is required to transform the displayed pitch. It also takes longer to transform the subjective loudness level of one image to match the subjective loudness level of the second, as the difference between the initial subjective loudness levels increases [92]. It is believed that the generation time of auditory images is unrelated to subjective loudness sensations [93]. However, this conclusion contradicts the assertion that auditory images preserve the temporal properties of acoustic signals and is applicable only to cases of gradual loudness increases. Based on this, it can be generally said that auditory images preserve some temporal properties of auditory stimuli.

3.3. Feedback Influence of Auditory Images on the Modulation of Auditory Perception

The detection threshold for weak acoustic signals increases with the simultaneous perception of multiple auditory patterns, consistent with the hypothesis that auditory patterns themselves interfere with holistic auditory perception. Such interference may limit the processing power and resources of the auditory system [94,95,96]. However, the formation of auditory patterns involves a temporal expectation [3,32], suggesting that the patterns themselves should enhance perception. This occurs if the perceived acoustic signal matches the expectation and should interfere with perception, as well as if the perceived stimulus does not match the expected one [97]. However, even if the perceived acoustic signal precisely matches the content of the pattern (and this match is strong), interference may occur.
There are several possible mechanisms by which auditory images interfere with or, conversely, facilitate (improve) auditory perception. The first possibility is that the signal is converted into an auditory image before the signal’s processing power is exceeded, followed by a back reflection. However, if participants recognize only one signal that does not exceed the auditory system’s processing power, either a back reflection or facilitation of recognition can be detected [98]. The second possibility is that facilitation occurs when an image is formed preceding the acoustic signal, and interference occurs if the auditory image and acoustic pattern are formed simultaneously [99]. A third possible interpretation suggests that facilitation is recorded when the content of the auditory image and the acoustic pattern coincide, and interference occurs when the auditory image and acoustic pattern do not coincide. However, interference can occur even in the case of an exact match with the auditory image [87,99]. A fourth possible explanation, consistent with earlier suggestions [20] regarding the interaction of images and perception, is that auditory images interfere with auditory detection but facilitate auditory differentiation or identification.
The idea that auditory images interfere with the detection of acoustic signals but facilitate their discrimination or identification is supported by a wide range of data. Thus, studies involving the detection of an acoustic signal are usually associated with inhibition of perception [99], whereas studies on the differentiation or identification of auditory stimuli, on the contrary, are usually accompanied by facilitation [79]. This difference arises if other cognitive levels of processing not directly related to sensory recognition are used to detect signals. The detection of an acoustic signal by sensory systems also involves the use of mechanisms of voluntary (selective) attention, which are maximally realized in auditory perception. If some of these mechanisms are blocked, then fewer attentional resources will be available for signal detection. The effects of the reduction of attentional processes are especially noticeable if the signal is weak. Furthermore, if attention mechanisms are modality-specific, this may explain the differences in the effects of same and different modalities on pattern detection [100,101].
Discrimination or identification of an already detected signal involves comparing information about the signal with large amounts of data in memory. An auditory image can lower the threshold for its recognition (retrieval) from memory, so that the completion of the sensory process can be either facilitated or inhibited by incoming detected information [47,99,102].

3.4. Neurogenerators of Auditory Perception and Images

Psychophysiological studies of priming (altered recognition, including set) [68], timbre similarity assessment [36], detection of embedded melodies in the presence of auditory distractors [86], as well as clinical studies of musical hallucinosis [76] in schizophrenia [81] are consistent with the hypothesis that auditory images are formed in brain areas involved in auditory perception, i.e., with primary and secondary auditory projection areas.
For example, patients with damage to the right temporal lobe perform worse on the task of comparing pitch in images and holistic perception than patients with damage to the left temporal lobe or control participants [84]. A similar decrease in performance is observed in healthy participants after applying transcranial magnetic stimulation to the right temporal lobe [36].
Furthermore, the superior temporal gyrus, frontal and parietal lobes, and motor cortex have been shown to be activated by pitch [93,94] and timbre [103,104] comparisons during acoustic pattern perception. The planum temporale (the area around Wernicke’s speech center) is activated by instructions to form an auditory image or by perception of environmental sounds [44]. The level of activation of this area during incoming acoustic signals correlates with ratings of the vividness of auditory images [105]. Auditory and premotor areas of the cortex are activated by playing notes on a musical instrument [103]. The application of TMS to the left hemisphere has been found to disrupt speech [52], while silent articulation and speech images activate the left inferior frontal gyrus and Broca’s motor speech area. The above is consistent with the hypothesis that auditory imagery involves brain regions involved not only in auditory but also in auditory-speech perception [48].
The problem with brain imaging studies and auditory perception assessment is that there is no objective behavioral evidence that auditory imagery has been formed [19,44,45,46,53,54]. It is assumed a priori that “images” are formed because participants have been given a task or the patterns find a plausible verbal interpretation.
It should be noted that the lack of behavioral evidence for auditory image recognition is not limited to brain imaging studies. Some behavioral studies report the evaluation of auditory imagery without evidence [97]. Examples of findings from such studies that could be used to support the above claims include priming [68,79,85], interference with behavioral tasks with specific processes or signal types [86], and measurement of image quality—the auditory image [86].
Answers to previously posed questions regarding the extent to which auditory images preserve the structural (amplitude) or temporal (frequency and timbre) properties of the acoustic signal, and whether images interfere with or facilitate signal perception, may be directly related to the extent to which images engage identical brain regions involved in auditory perception. Even so, the brain regions involved in generating auditory imagery may not be entirely identical across tasks. Differences have been found between the brain regions activated during auditory perception and the brain regions activated when participants receive instructions and/or are presumably generating auditory imagery. For example, the primary auditory cortex is activated by auditory perception but is not activated (or is weakly activated) by instructions to generate or use auditory imagery [36,44,45,84].
Observations that auditory imagery can usually be discerned in auditory perception in the presence of incidental auditory hallucinations [81] or impaired reality monitoring [33,106] suggest from brain imaging that many, but not all, of the brain regions involved in auditory perception are involved.

4. Classification of EEG-ERPs Parameters in Behavioral Studies of the Auditory System

While EEG analysis is integrated to some extent into machine learning models, the use of ML models for auditory perception research is considered in isolation in most studies, using separate methods. This issue is being explored most fully in BCI systems for creating support systems for individuals with complex auditory impairments, but one of the criteria—the P300 wave of auditory ERPs, the alpha rhythm, etc.—is not fully understood. Over the past two decades, dozens of attempts have been made to classify EEG signals for the development of noninvasive brain–computer interface (BCI) systems. According to the National Center for Biotechnology Information (USA), over 2000 results appear in response to a search query with the keywords BCI (brain–computer interaction), EEG, ERPs, and classification [107,108,109,110,111,112].
The classical model for recording mismatch negativity (MMN) during EP recording assumes frequent presentation of a standard signal and rare presentations of a deviant signal that differs only slightly in characteristics [113]. In representative cases, the averaged, typically auditory evoked potential (AEP), deviates negatively from the standard averaged AEP due to changing conditions. This difference is typically recorded approximately 100 ms after the onset of the deviant stimulus and lasts for approximately 100 ms [96,114], i.e., it appears during the 100–200 ms analysis epoch. The MMN model proposes that negativity reveals a mismatch between the incoming auditory deviant stimulus and the template from memory due to the standard stimulus. Importantly, MMN does not occur under conditions of continuously changing stimuli unless the deviant stimulus itself is a repetition. Although MMN is most often modulated by attention in studies, its occurrence is independent of it [113,115]. MMN is often recorded when listeners are actively engaged in an unrelated task [99]. More remarkably, MMN can even be recorded in comatose patients [116], supporting the use of MMN as a powerful tool for investigating automatically processed aspects of auditory input.
The presence of a late MMN, but not an early or both MMNs, suggests something unique about the formation of the auditory template for global processing [113]. At first glance, this appears paradoxical. The late violation elicited an MMN only when it initially matched, whereas no MMN occurred when it violated the standard initially and finally. This finding suggests that the late global MMN actually indexed global structure. This interpretation remains preliminary, as a similar advantage of the late global deviation was not observed in behavior (the late deviation did not differ from the early and both behavioral conditions).
The main challenge here is the classification of EEG-ERP signals with high accuracy in real time [117,118]. In the applied aspect, this is necessary, for example, for the development of rehabilitation complexes for the restoration of motor functions after a stroke [109] and traumatic brain injuries with the involvement of patients in the control of external devices or applications with biofeedback, including in a game form [35] and even the creation of control of information posts on a contactless basis [119,120].
The main challenges in working with bioelectric signals include the low signal-to-noise ratio [34,116], its significant variability, taking into account individual human characteristics [120,121], and even from session to session for a specific subject [26,122]. This dictates the need to make the classifier either robust to signal changes or adaptive (adapting to signal changes, including those of a specific subject), in conditions of a small amount of data (lengthy procedures are tiring for patients).
Among the approaches to classifying EEG signals in auditory system studies, two groups can be distinguished. The first category includes approaches that focus on extracting useful information from the EEG signal, constructing and prospectively selecting the most informative features. The main advantage of the approaches of this group is the ability to interpret the results—to assess the importance of an individual feature, and then to reconstruct which part of the signal this feature corresponds to [26,121].
The second group includes approaches that use automatic feature extraction, such as convolutional neural networks (CNNs) [118]. The main advantages of such approaches are the ability to work with the original signal and high generalization capacity, which can contribute to the robustness of the classifier. In the case of automatic feature extraction, interpretation of the results is difficult unless it is possible to reconstruct the information contained in the extracted features [21,26].
The formal formulation of the problem of classification of EEG patterns is as follows: a set of samples N is given X 1 , , X N , for each of which the class is known y i 1 , , K , where K —number of classes considered. Each sample is a matrix of signal amplitudes of size E × T , where E —number of electrodes used, and T = Δ t · f s —number of time samples during the sample length Δ t and sampling frequency f s . It is necessary to construct and train a classifier using the available data that is capable of determining the class for samples in subsequent training sessions [26,123].
Immediately prior to sample classification, several signal preprocessing methods are applied to improve the final classification accuracy. Typical preprocessing methods include augmentation, decomposition, spatial filtering, and feature extraction and selection.
Any classification method requires a set of numbers, called features, as input. Successful classification requires that these features, taken together, distinguish between samples belonging to different classes. Therefore, feature extraction is often used, in which more stable features are calculated based on the original signal. Feature extraction can consider the time domain, the frequency domain, or both. In the time domain, autoregressive model coefficients [124], covariance matrices [107], and log-variance features [26] are used. To take into account information about the frequency domain, the power in the band and wavelet coefficients [123] are used.
Machine learning methods applicable to EEG pattern classification can be divided into several categories, as reviewed [34,116,124,125,126]. Below is a list of methods divided into categories with references to the studies that used the respective method:
-
Methods for constructing a hyperplane separating classes: linear discriminant analysis (LDA) [125] and support vector machine (SVM) [26,127].
-
Methods based on calculating the proximity between objects: nearest neighbor method [89] and minimum Riemannian distance to mean (MRDM) [128].
-
Probabilistic methods: Bayesian classifier [87] and Markov models [124].
-
Decision trees [128].
-
Deep neural networks [15].
Deep neural networks, depending on the architecture used, can solve both the classification problem and the feature extraction problem. Fully connected feedforward networks [123] are suitable for classification, while autoencoders [101] are suitable for automatic feature extraction by searching for a hidden representation. At the same time, convolutional networks [105] solve both problems by combining convolutional and fully connected layers responsible for feature extraction and classification, respectively. Due to this, the input of the convolutional network is fed with the original signal, which eliminates the need for manual feature extraction. Recurrent networks using LSTM cells [26] allow working with sequences while maintaining context, which is why they are actively used for machine translation and can be applied to signal processing.
Among the studies reviewed, convolutional neural networks (CNNs) are most frequently used. A convolutional neural network is a type of artificial neural network architecture [110] for efficient pattern recognition in images. Unlike a conventional neural network, a convolutional neural network also contains convolutional layers and pooling layers. The idea behind their use stems from an analogy with how the human brain works: some neurons are focused on recognizing simple patterns, such as line angles in images, while another group of cells analyzes the responses of first-level cells [15].
When working with EEG signals, spectrograms can be fed to the input of a convolutional network, in which case the task is reduced to image classification. An alternative input data option is the original signal using an architecture that is an adaptation of the FBCSP method. A similar architecture, ShallowNet, has been described previously [22]. Below are descriptions of all layers used and their purposes:
  • Temporal convolution is performed using a 1 × 25 kernel to isolate characteristic peaks in the signal.
  • Convolution is performed across all electrodes; this step is analogous to spatial filtering in the FBCSP algorithm.
  • All matrix values are squared element-wise.
  • For each 1 × 75 window, temporal pooling is performed: the mean value of the elements in the window is taken.
  • The natural logarithm of each element is taken. The combination of steps 3–5 is equivalent to calculating the log-variance of features in the FBCSP algorithm.
  • The classification problem for the features obtained after pooling is solved by a combination of a fully connected and SoftMax layer.
The main drawback of deep learning methods is their dependence on the amount of available data—the more data, the better the generalization. Augmentation methods, already mentioned earlier [26], are currently used to address this problem.
When discussing methods for analyzing EEG features, it is worth noting an important principle that also underlies EEG generation—the principle of self-similarity. The structure of such processes is based on a special kind of set—fractals—that exhibits scale invariance. Fractal is a term coined by Benoit Mandelbrot in 1975 to describe objects constructed through a repeated action, where some aspects constraining the object are infinite while others are finite, and where at any stage of construction, some parts of the object represent a reduced version of the previous stage. It has now been established that the EEG has a fractal structure [21,26]. Thus, along with oscillatory models of EEG decoding, one can also speak of fractal models.
There are a number of theoretical premises that, if taken into account, can lead to a positive solution to the problem of obtaining information from EEG. Let us consider the main ones.
  • It can be considered a well-established fact that the EEG as a process is characterized, on the one hand, by the manifestation of a deterministic factor—regular brain activity (rhythms), and on the other, by a chaotic factor, reflected in the fractal structure of the process [17,21,26].
  • The overwhelming majority of natural fractals, and especially those generated by living organisms, are in fact not fractals, but multifractals—that is, not regular, but random fractals. The property of exact self-similarity is characteristic only of regular fractals with a deterministic method of their construction. Multifractals exhibit self-similarity only after appropriate averaging over all statistically independent realizations of the object. It can be considered with a high degree of certainty that the EEG is a multifractal. In addition to purely geometric characteristics determined by the value of the fractal dimension D, multifractals also have some statistical properties [17,21].
It is known that any measurement process carries certain information about the phenomenon under study. This information can typically be extracted by knowing the modulation or coding law of the original (base) process. If we consider the base process from an oscillatory perspective, it is believed that useful information is contained in changes in the amplitude, frequency, or phase of oscillations.
The fractal dynamics analysis (FDA) method [17,21] involves dividing the original process into time epochs, the duration of which corresponds to a stationarity or quasi-stationarity interval. Over this interval, the measurable characteristics of the fractal do not change (or change insignificantly). When choosing the stationarity interval, one can also be guided by the duration of the leading rhythm in the system under study: if the system is a living organism, this could be, for example, the heart rate (the duration of the RR interval, which itself can be variable). The choice of the number of time epochs is determined by additional requirements defining the purpose of the study. As a rule, this number does not exceed several dozen; the total time interval should ensure the detection of changes in the dynamic characteristics of the fractal [17,21].
The AFD technology is based on concepts of the structure of electrophysiological processes, including:
-
Deterministic chaos;
-
Fractal fluctuations such as Brownian motion and its variations;
-
A quasi-regular component (biorhythms).
In principle, the higher the dimensionality (and, consequently, the higher the reliability of the correlate of the process being described), the longer the time epochs required [21]. However, for biological signals, this requirement may conflict with stationarity conditions, which require shorter intervals. AFD not only addresses this requirement but also offers a solution. The advantages of the AFD method include the following [17,21].
First, the method is insensitive to short-term interference, which is almost inevitable in functional diagnostics and can sometimes seriously impede the work of even experienced medical specialists. This eliminates the need for careful pre-selection of EEG fragments for analysis. That is, one of the main obstacles to the creation of automatic diagnostic EEG systems is being overcome. Such systems are indispensable for mobile research in conditions remote from medical and preventive institutions, in the event of disasters and natural calamities, etc. [129].
Secondly, relatively short 30 s EEG fragments can be used for analysis, and such a small amount of data is easy to store and transmit over long distances. This allows the AFD method to be used in telemedicine services, during mass screenings, etc. Furthermore, it is very important when time is limited to examine a patient. For example, during intracranial neurosurgery, emergency diagnostics, in pediatrics (when it is difficult to keep a child still for a long time), during monitoring of the functional state of the operator, etc. [17,21].
Thirdly, using the obtained EEG characteristics, it is possible to create effective automatic EEG classification systems. For example, to separate them into normal and pathological, by the nature of the pathology, by the functional state of the brain during a selected 30 s period, etc. The only thing left to do is to train the system on an adequate training sample, in accordance with a specific practical task [17,21].

Funding

This research by Saint Petersburg State Pediatric Medical University was supported by project No. AAAA-A19-119112290090-9.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

Not applicable.

Conflicts of Interest

The author declares no conflicts of interest.

References

  1. Reisberg, D.; Heuer, F. Visuospatial images. In The Cambridge Handbook of Visuospatial Thinking; Shah, P., Miyake, A., Eds.; Cambridge University Press: New York, NY, USA, 2005. [Google Scholar]
  2. Reber, P.J.; Kounios, J. Neural Activity When People Solve Verbal Problems with Insight. PLoS Biol. 2004, 2, 500–510. [Google Scholar]
  3. Janata, P. Brain electrical activity evoked by mental formation of auditory expectations and images. Brain Topogr. 2001, 13, 169–193. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Lytaev, S. Long-latency event-related potentials (300–1000 ms) of the visual insight. Sensors 2022, 22, 1323. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Lytaev, S. Interaction of Sensitivity, Emotions, and Motivations during Visual Perception. Sensors 2024, 24, 7414. [Google Scholar] [CrossRef] [Scilit]
  6. Altman, Y.A.; Vaitulevich, S.F.; Petrovpavlovskaya, E.A.; Shestopalova, L.B. Human discrimination of dynamic changes in the spatial position of sound images (electrophysiological and psychophysical study). Hum. Physiol. 2010, 36, 83–92. [Google Scholar] [CrossRef] [Scilit]
  7. Altman, Y.A.; Vaitulevich, S.F.; Shestopalova, L.B.; Petrovpavlovskaya, E.A.; Nikitin, N.I. Total electrical potentials of the human brain during sound source localization. Uspekhi Fiz. Nauk 2012, 43, 3–18. [Google Scholar]
  8. Eardley, A.F.; Pring, L. Remembering the past and imaging the future: A role for nonvisual imagery in the everyday cognition of blind and sighted people. Memory 2006, 14, 925–936. [Google Scholar] [CrossRef] [Scilit]
  9. Belskaya, K.A.; Lytaev, S.A. Neuropsychological analysis of cognitive deficits in schizophrenia. Hum. Physiol. 2022, 48, 37–45. [Google Scholar] [CrossRef] [Scilit]
  10. Altman, J.A.; Vaitulevich, S.P.; Shestopalova, L.B.; Petropavlovskaia, E.A. How does mismatch negativity reflect auditory motion? Hear. Res. 2010, 268, 194–201. [Google Scholar] [CrossRef] [Scilit]
  11. Bibikov, N.G. Neurophysiological mechanisms of auditory adaptation. I. Adaptation during the stimulus action. Uspekhi Fiz. Nauk 2010, 41, 72–90. (In Russian) [Google Scholar]
  12. Bibikov, N.G. Neurophysiological mechanisms of auditory adaptation. II. Poststimulus effects. Uspekhi Fiz. Nauk 2010, 41, 77–92. (In Russian) [Google Scholar] [PubMed]
  13. Carlile, S. The plastic ear and perceptual relearning in auditory spatial perception. Front. Neurosci. 2014, 8, 237. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Melo, Â.; Mezzomo, C.L.; Garcia, M.V.; Biaggio, E.P.V. Computerized Auditory Training in Students: Electrophysiological and Subjective Analysis of Therapeutic Effectiveness. Int. Arch. Otorhinolar. 2018, 22, 23–32. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Lawhern, V.J.; Solon, A.J.; Waytowich, N.R.; Gordon, S.M.; Hung, C.P.; Lance, B.J. EEGNet: A Compact Convolutional Network for EEG-based Brain-Computer Interfaces. J. Neural Eng. 2018, 15, 056013. [Google Scholar] [CrossRef] [Scilit]
  16. Levi-Aharoni, H.; Shriki, O.; Tishby, N. Surprise response as a probe for compressed memory states. PLoS Comp. Biol. 2020, 16, e1007065. [Google Scholar] [CrossRef] [Scilit]
  17. Lytaev, S.; Surovitskaya, Y. The frustration status and noise proof feature during perception of the auditory images. In Foundations of Augmented Cognition. Directing the Future of Adaptive Systems; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2011; Volume 6780, pp. 186–193. [Google Scholar]
  18. Lytaev, S.A.; Belskaya, K.A. Integration and Disintegration of Auditory Images Perception. In Foundations of Augmented Cognition; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2015; Volume 9183, pp. 470–480. [Google Scholar]
  19. Meyer, M.; Elmer, S.; Baumann, S.; Jancke, L. Short-term plasticity in the auditory system: Differential neural responses to perception and imagery of speech and music. Restor. Neurol. Neurosci. 2007, 25, 411–431. [Google Scholar] [CrossRef] [Scilit]
  20. Milekina, O.N.; Nechaev, D.I.; Supin, A.Y. Estimates of frequency resolving power of humans by different methods: The role of sensory and cognitive factors. Hum. Physiol. 2018, 44, 357–363. [Google Scholar] [CrossRef] [Scilit]
  21. Polonnikov, R.I.; Kartashev, N.K.; Wasserman, E.L. Regular developmental changes in EEG multifractal characteristics. Int. J. Neurosci. 2003, 113, 1615–1639. [Google Scholar] [CrossRef] [Scilit]
  22. Schirrmeister, R.T.; Springenberg, J.T.; Fiederer, L.D.J.; Glasstetter, M.; Eggensperger, K.; Tangermann, M.; Hutter, F.; Burgard, W.; Ball, T. Deep learning with convolutional neural networks for EEG decoding and visualization. Hum. Brain Mapp. 2017, 38, 5391–5420. [Google Scholar] [CrossRef] [Scilit]
  23. Lytaev, S. PET-Neuroimaging and Neuropsychological Study for Early Cognitive Impairment in Parkinson’s Disease. In Proceedings of the Bioinformatics and Biomedical Engineering. IWBBIO 2022, Gran Canaria, Spain, 27–30 June 2022; Rojas, I., Valenzuela, O., Rojas, F., Herrera, L.J., Ortuño, F., Eds.; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2022; Volume 13346, pp. 143–153. [Google Scholar]
  24. Sellers, E.; Donchin, E. A P300-based brain–computer interface: Initial tests by ALS patients. Clin. Neurophysiol. 2006, 117, 538–548. [Google Scholar] [CrossRef] [Scilit]
  25. Lebedeva, N.N.; Karimova, E.D. Acoustic characteristics of speech signal as an indicator of human functional state. Uspekhi Fiz. Nauk 2014, 45, 57–95. (In Russian) [Google Scholar]
  26. Kapralov, N.V.; Nagornova, Z.V.; Shemyakina, N.V. Methods for classifying EEG patterns of imaged movements. Inform. Autom. 2021, 20, 94–132. [Google Scholar]
  27. Luo, H.; Poeppel, D. Cortical oscillations in auditory perception and speech: Evidence for two temporal windows in human auditory cortex. Front. Psychol. 2012, 3, 170. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Luo, H.; Tian, X.; Song, K.; Zhou, K.; Poeppel, D. Neural Response Phase Tracks How Listeners Learn New Acoustic Representations. Curr. Biol. 2013, 23, 968–974. [Google Scholar] [CrossRef] [Scilit]
  29. McDermott, J.H.; Simoncelli, E.P. Sound Texture Perception via Statistics of the Auditory Periphery: Evidence from Sound Synthesis. Neuron 2011, 71, 926–940. [Google Scholar] [CrossRef] [Scilit]
  30. McDermott, J.H.; Schemitsch, M.; Simoncelli, E.P. Summary statistics in auditory perception. Nat. Neurosci. 2013, 16, 493–498. [Google Scholar] [CrossRef] [Scilit]
  31. Zhao, X.; Zhao, J.; Liu, C.; Cai, W. Deep Neural Network with Joint Distribution Matching for Cross-Subject Motor Imagery Brain-Computer Interfaces. BioMed Res. Int. 2020, 2020, 7285057. [Google Scholar] [CrossRef] [Scilit]
  32. Janata, P.; Paroo, K. Acuity of auditory images in pitch and time. Percept. Psychophys. 2006, 68, 829–844. [Google Scholar] [CrossRef] [Scilit]
  33. Jones, S.J.; Vaz Pato, M.; Spraque, L.; Stokes, M.; Munday, R.; Haque, N. Auditory Evoked Potentials to Spectro-Temporal Modulation of Complex Tones in Normal Subjects and Patients with Severe Brain Injury. Brain 2000, 123, 1007–1016. [Google Scholar] [CrossRef] [Scilit]
  34. Lotte, F.; Bougrain, L.; Cichocki, A.; Clerc, M.; Congedo, M.; Rakotomamonjy, A.; Yger, F. A review of classification algorithms for EEG-based brain-computer interfaces: A 10 year update. J. Neural Eng. 2018, 15, 031005. [Google Scholar] [CrossRef] [Scilit]
  35. Lytaev, S. Modeling and Estimation of Physiological, Psychological and Sensory Indicators for Working Capacity. Adv. Intell. Syst. Comput. 2021, 1201, 207–213. [Google Scholar]
  36. Halpern, A.R.; Zatorre, R.J.; Bouffard, M.; Johnson, J.A. Behavioral and neural correlates of perceived and imagined musical timbre. Neuropsychologia 2004, 42, 1281–1292. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Carrer, L.R. Music and sound in time processing of children with ADHD. Front. Psychiatry 2015, 6, 127. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Grinspun, N.; Nijs, L.; Kausel, L.; Onderdijk, K.; Sepúlveda, N.; Rivera-Hutinel, A. Selective attention and inhibitory control of attention are correlated with music audiation. Front. Psychol. 2020, 11, 1109. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Akerman, A.; Etkovitch, A.; Kalanthroff, E. Global-Local Processing in ADHD Is Not Limited to the Visuospatial Domain: Novel Evidence from the Auditory Domain. J. Atten. Disord. 2023, 27, 822–829. [Google Scholar] [CrossRef] [Scilit]
  40. Sanguebuche, T.R.; Peixe, B.P.; Bruno, R.S.; Biaggio, E.P.V.; Garcia, M.V. Speech-evoked Brainstem Auditory Responses and Auditory Processing Skills: A Correlation in Adults with Hearing Loss. Int. Arch. Otorhinolaryngol. 2018, 22, 38–44. [Google Scholar] [CrossRef] [Scilit]
  41. Nechaev, D.I.; Supin, A.Y. Hearing sensitivity to shifts of rippled-spectrum patterns. J. Acoust. Soc. Am. 2013, 134, 2913–2922. [Google Scholar] [CrossRef] [Scilit]
  42. Patterson, R.D. Auditory images: How complex sounds are represented in the auditory system. Acoust. Sci. Technol. 2000, 21, 183–190. [Google Scholar] [CrossRef] [Scilit]
  43. Schneider, T.R.; Engel, A.K.; Debener, S. Multisensory identification of natural objects in a two-way cross modal priming paradigm. Exp. Psychol. 2008, 55, 121–132. [Google Scholar] [CrossRef] [Scilit]
  44. Bunzeck, N.; Wuestenberg, T.; Lutz, K.; Heinze, H.J.; Jancke, L. Scanning silence: Mental imagery of complex sounds. NeuroImage 2005, 26, 1119–1127. [Google Scholar] [CrossRef] [Scilit]
  45. Wu, J.; Mai, X.; Chan, C.C.H.; Zheng, Y.; Luo, Y. Event-related potentials during mental imagery of animal sounds. Psychophysiology 2006, 43, 592–597. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  46. Kraemer, D.J.M.; Macrae, C.N.; Green, A.E.; Kelly, W.M. Sound of silence activates auditory cortex. Nature 2005, 434, 158. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Dmitrieva, E.S.; Gel’man, V.I. Perception of emotional intonation of noisy speech signal with different acoustic parameters by adults of different age and gender. Zh Vyss. Nerv. Deiat Im. I P Pavlov. 2011, 61, 306–316. (In Russian) [Google Scholar] [PubMed]
  48. Carlile, S.; Leung, J. The Perception of Auditory Motion. Trends Hear. 2016, 20, 1–19. [Google Scholar] [CrossRef] [Scilit]
  49. Van Dinteren, R.; Arns, M.; Jongsma, M.L.A.; Kessels, R.P.C. P300 Development across the Lifespan: A Systematic Review and Meta-Analysis. PLoS ONE 2014, 9, 0087347. [Google Scholar] [CrossRef] [Scilit]
  50. Bookheimer, S. Functional MRI of language: New approaches to understanding the cortical organization of semantic processing. Annu. Rev. Neurosci. 2002, 25, 151–188. [Google Scholar] [CrossRef] [Scilit]
  51. Stewart, L.; Walsh, V.; Frith, U.; Rothwell, J.C. TMS produces two dissociable types of speech disruption. NeuroImage 2001, 13, 472–478. [Google Scholar] [CrossRef] [Scilit]
  52. Aziz-Zadeh, L.; Cattaneo, L.; Rochat, M.; Rizzolatti, G. Covert speech arrest induced by rTMS over both motor and nonmotor left hemisphere frontal sites. J. Cogn. Neurosci. 2005, 17, 928–938. [Google Scholar] [CrossRef] [Scilit]
  53. Shergill, S.S.; Murray, R.M.; McGuire, P.K. Auditory hallucinations: A review of psychological treatments. Schizophr. Res. 1998, 32, 137–150. [Google Scholar] [CrossRef] [Scilit]
  54. Shergill, S.S.; Bullmore, E.T.; Brammer, M.J.; Williams, S.C.; Murray, R.M.; McGuire, P.K. A functional study of auditory verbal imagery. Psychol. Med. 2001, 31, 241–253. [Google Scholar] [CrossRef] [Scilit]
  55. Ueda, K.; Takeichi, H.; Wakamiya, K. Auditory grouping is necessary to understand interrupted mosaic speech stimuli. J. Acoust. Soc. Am. 2022, 152, 970–980. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  56. Aleman, A.; Formisano, E.; Koppenhagen, H.; Hagoort, P.; de Haan, E.H.; Kahn, R.S. The functional neuroanatomy of metrical stress evaluation of perceived and imaged spoken words. Cereb. Cortex 2005, 15, 221–228. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  57. Berti, S.; Munzer, S.; Schroger, E.; Pechmann, T. Different interference effects in musicians and a control group. Exp. Psychol. 2006, 53, 111–116. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  58. Maĭorova, L.A.; Martynova, O.V.; Balaban, P.M.; Ivanitskiĭ, A.M.; Shklovskiĭ, V.M. Mismatch negativity and its hemodynamic equivalent (based on fMRI) in research of speech perception in healthy and in speech disorders. Uspekhi Fiz. Nauk 2014, 45, 27–43. (In Russian) [Google Scholar] [PubMed]
  59. Seal, M.L.; Aleman, A.; McGuire, P.K. Compelling imagery, unanticipated speech, and deceptive memory: Neurocognitive models of auditory verbal hallucinations in schizophrenia. Cogn. Neuropsychiatry 2004, 9, 43–72. [Google Scholar] [CrossRef] [Scilit]
  60. Lytaev, S. Short Time Algorithms for Screening Examinations of the Collective and Personal Stress Resilience. In Engineering Psychology and Cognitive Ergonomics. HCII 2023; Harris, D., Li, W.C., Eds.; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2023; Volume 14017, pp. 442–458. [Google Scholar] [CrossRef] [Scilit]
  61. Razavi, B.; O’Neill, W.E.; Paige, G.D. Auditory Spatial Perception Dynamically Realigns with Changing Eye Position. J. Neurosci. 2007, 27, 10249–10258. [Google Scholar] [CrossRef] [Scilit]
  62. Piazza, E.A.; Sweeny, T.D.; Wessel, D.; Silver, M.A.; Whitney, D. Humans use summary statistics to perceive auditory sequences. Psychol. Sci. 2013, 24, 1389–1397. [Google Scholar] [CrossRef] [Scilit]
  63. Minatoya, M.; Daikoku, T.; Kuniyoshi, Y. Emotional responses to auditory hierarchical structures is shaped by bodily sensations and listeners’ sensory traits. Front. Psychol. 2025, 16, 1599430. [Google Scholar] [CrossRef] [Scilit]
  64. Karim, A.K.M.R.; Proulx, M.J.; de Sousa, A.A.; Likova, L.T. Do we enjoy what we sense and perceive? A dissociation between aesthetic appreciation and basic perception of environmental objects or events. Cogn. Affect. Behav. Neurosci. 2022, 5, 904–951. [Google Scholar] [CrossRef] [Scilit]
  65. Menninghaus, W.; Wagner, V.; Wassiliwizky, E.; Schindler, I.; Hanich, J.; Jacobsen, T.; Koelsch, S. What are aesthetic emotions? Psychol. Rev. 2019, 126, 171–195. [Google Scholar] [CrossRef] [Scilit]
  66. Schindler, I.; Hosoya, G.; Menninghaus, W.; Beermann, U.; Wagner, V.; Eid, M.; Scherer, K.R. Measuring aesthetic emotions: A review of the literature and a new assessment tool. PLoS ONE 2017, 12, e0178899. [Google Scholar] [CrossRef] [Scilit]
  67. Egermann, H.; Reuben, F. “Beauty is how you feel inside”: Aesthetic judgments are related to emotional responses to contemporary music. Front. Psychol. 2020, 11, 29. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  68. List, A. Global and local priming in a multi-modal context. Front. Hum. Neurosci. 2023, 16, 1043475. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  69. Lytaev, S.; Belskaya, K. Levels of Consciousness in Psychopathology According to Monitoring of Neural Network Centers Alpha Rhythm rs-EEG. Bull. Electr. Eng. Inform. 2025, 14, 3133–3145. [Google Scholar] [CrossRef] [Scilit]
  70. Lytaev, S.; Belskaya, K. Neural Network of Alpha Rhythm rs-EEG Activity for the Assessment of Consciousness in Psychopathology. OBM Neurobiol. 2025, 9, 299. [Google Scholar] [CrossRef] [Scilit]
  71. Evans, C.L.; McGuire, P.K.; David, A.S. Is auditory imagery defective in patients with auditory hallucinations? Psychol. Med. 2000, 30, 137–148. [Google Scholar] [CrossRef] [Scilit]
  72. Fischer, C.E.; Marchie, A.; Norris, M. Musical and auditory hallucinations: A spectrum. Psychiatry Clin. Neurosci. 2004, 58, 96–98. [Google Scholar] [CrossRef] [Scilit]
  73. Fritzsch, B.; Knipper, M.; Friauf, E. Auditory system: Development, genetics, function, aging and diseases. Cell Tissue Res. 2015, 361, 1–6. [Google Scholar] [CrossRef] [Scilit]
  74. Sacks, O. Musicophilia: Tales of Music and the Brain; Knoft: New York, NY, USA, 2007. [Google Scholar]
  75. Hermesh, H.; Konas, S.; Shiloh, R.; Dar, R.; Marom, S.; Weizman, A.; Gross-Isseroff, R. Musical hallucinations: Prevalence in psychotic and non-psychotic outpatients. J. Clin. Psychiatry 2004, 65, 91–197. [Google Scholar] [CrossRef] [Scilit]
  76. Shinosaki, K.; Yamamoto, M.; Ukai, S.; Kawaguchi, S.; Ogawa, A.; Ishii, R.; Yuko, M.M.; Inouye, T.; Hirabuki, N.; Kaku, Y.; et al. Desynchronization in the right auditory cortex during musical hallucinations: A MEG study. Psychogeriatrics 2003, 3, 88–92. [Google Scholar] [CrossRef] [Scilit]
  77. Lewald, J.; Staedtgen, M.; Sparing, R.; Meister, I.G. Processing of auditory motion in inferior parietal lobule: Evidence from transcranial magnetic stimulation. Neuropsychologia 2011, 49, 209–215. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  78. Ghitza, O.; Greenberg, S. On the Possible Role of Brain Rhythms in Speech Perception: Intelligibility of Time-Compressed Speech with Periodic and Aperiodic Insertions of Silence. Phonetica 2009, 66, 113–126. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  79. Hubbard, T.L. Auditory Imagery: Empirical Findings. Psychol. Bull. 2010, 136, 302–329. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  80. Shoshina, I.; Isajeva, E.; Mukhitova, Y.; Tregubenko, I.; Khan’ko, A.; Limankin, O.; Simon, Y. The internal noise of the visual system and cognitive functions in schizophrenia. Procedia Comput. Sci. 2020, 169, 813–820. [Google Scholar] [CrossRef] [Scilit]
  81. Lennox, B.R.; Park, S.B.G.; Medley, I.; Morris, P.G.; Jones, P.B. The functional anatomy of auditory hallucinations in schizophrenia. Psychiatry Res. Neuroimaging 2000, 100, 13–20. [Google Scholar] [CrossRef] [Scilit]
  82. Brimijoin, W.O.; Akeroyd, M.A. The role of head movements and signal spectrum in an auditory front/back illusion. i-Perception 2012, 3, 179–182. [Google Scholar] [CrossRef] [Scilit]
  83. Nosulenko, V.N. Psychology of Auditory Perception; Nauka: Moscow, Russia, 1988. (In Russian) [Google Scholar]
  84. Giraud, A.-L.; Kleinschmidt, A.; Poeppel, D.; Lund, T.E.; Frackowiak, R.S.; Laufs, H. Endogenous Cortical Rhythms Determine Cerebral Specialization for Speech Perception and Production. Neuron 2007, 56, 1127–1134. [Google Scholar] [CrossRef] [Scilit]
  85. Crowder, R.G. Imagery for musical timbre. J. Exp. Psychol. Hum. Percept. Perform. 1989, 15, 472–478. [Google Scholar] [CrossRef]
  86. Brodsky, W.; Kessler, Y.; Rubenstein, B.S.; Ginsborg, J.; Henik, A. The mental representation of music notation: Notational audition. J. Exp. Psychol. Hum. Percept. Perform. 2008, 34, 427–445. [Google Scholar] [CrossRef] [Scilit]
  87. Frolov, A.A.; Mokienko, O.; Lyukmanov, R.; Biryukova, E.; Kotov, S.; Turbina, L.; Nadareyshvily, G.; Bushkova, Y. Post-stroke Rehabilitation Training with a Motor-Imagery-Based Brain-Computer Interface (BCI)-Controlled Hand Exoskeleton: A Randomized Controlled Multicenter Trial. Front. Neurosci. 2017, 11, 400. [Google Scholar] [CrossRef] [Scilit]
  88. Nechaev, D.I.; Milekhina, O.N.; Supin, A.Y. Hearing sensitivity to gliding rippled spectrum patterns. J. Acoust. Soc. Am. 2018, 143, 2387–2393. [Google Scholar] [CrossRef] [Scilit]
  89. Alink, A.; Euler, F.; Kriegeskorte, N.; Singer, W.; Kohler, A. Auditory motion direction encoding in auditory cortex and high-level visual cortex. Hum. Brain Mapp. 2012, 33, 969–978. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  90. Dietrich, A.; Kanso, R. A review of EEG, ERP, and neuroimaging studies of creativity and insight. Psychol. Bull. 2010, 136, 822–848. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  91. Gordon, M.S.; Russo, F.A.; MacDonald, E. Spectral information for detection of acoustic time to arrival. Atten. Percept. Psychophys. 2013, 75, 738–750. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  92. Chait, M.; Greenberg, S.; Arai, T.; Simon, J.Z.; Poeppel, D. Multi-time resolution analysis of speech: Evidence from psychophysics. Front. Neurosci. 2015, 9, 214. [Google Scholar] [CrossRef] [Scilit]
  93. Lytaev, S. Modern Neurophysiological Research of the Human Brain in Clinic and Psychophysiology. In Bioengineering and Biomedical Signal and Image Processing; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2021; Volume 12940, pp. 231–241. [Google Scholar]
  94. Giraud, A.-L.; Poeppel, D. Cortical oscillations and speech processing: Emerging computational principles and operations. Nat. Neurosci. 2012, 15, 511–517. [Google Scholar] [CrossRef] [Scilit]
  95. Justus, T.; List, A. Auditory attention to frequency and time: An analogy to visual local–global stimuli. Cognition 2005, 98, 31–51. [Google Scholar] [CrossRef] [Scilit]
  96. Näätänen, R.; Paavilainen, P.; Rinne, T.; Alho, K. Mismatch negativity (MMN) in basic research of central auditory processing: A review. Clin. Neurophysiol. 2007, 118, 2544–2590. [Google Scholar] [CrossRef] [Scilit]
  97. Keller, P.E.; Koch, I. Action planning in sequential skills: Relations to music performance. Q. J. Exp. Psychol. 2008, 61, 275–291. [Google Scholar] [CrossRef] [Scilit]
  98. Touge, T.; Gonzalez, D.; Wu, J.; Deguchi, K.; Tsukaguchi, M.; Shimamura, M.; Ikeda, K.; Kuriyama, S. The Interaction between Somatosensory and Auditory Cognitive Processing Assessed with Event-Related Potentials. J. Clin. Neurophysiol. 2008, 25, 90–97. [Google Scholar] [CrossRef] [Scilit]
  99. Altman, Y.A.; Vaitulevich, S.F.; Varfolomeev, A.L.; Petropavlovskaya, E.; Shestopalova, L. Mismatch negativity as an indicator of human discriminative localization ability. Hum. Physiol. 2007, 33, 22–30. [Google Scholar]
  100. Goossens, T.; van de Par, S.; Kohlrausch, A. Gaussian-noise discrimination and its relation to auditory object formation. J. Acoust. Soc. Am. 2009, 125, 3882–3893. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  101. Hinton, G.E.; Salakhutdinov, R.R. Reducing the Dimensionality of Data with Neural Networks. Science 2006, 313, 504–507. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  102. Anderson, E.S.; Oxenham, A.J.; Nelson, P.B.; Nelson, D.A. Assessing the role of spectral and intensity cues in spectral ripple detection and discrimination on cochlear-implant users. J. Acoust. Soc. Am. 2012, 132, 3925–3934. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  103. Akeroyd, M.A. An overview of the major phenomena of the localization of sound sources by normal-hearing, hearing-impaired, and aided listeners. Trends Hear. 2014, 18, 2331216514560442. [Google Scholar] [CrossRef] [Scilit]
  104. Aleman, A.; van’t Wout, M. Subvocalization in auditory-verbal imagery: Just a form of motor imagery? Cogn. Process. 2004, 5, 228–231. [Google Scholar] [CrossRef] [Scilit]
  105. Scheller, M.; Matorres, F.; Little, A.C.; Tompkins, L.; de Sousa, A.A. The role of vision in the emergence of mate preferences. Arch. Sex. Behav. 2021, 50, 3785–3797. [Google Scholar] [CrossRef] [Scilit]
  106. Leung, J.; Wei, V.; Burgess, M.; Carlile, S. Head tracking of auditory, visual and audio-visual targets. Front. Neurosci. 2016, 9, 493. [Google Scholar] [CrossRef] [Scilit]
  107. Barachant, A.; Bonnet, S.; Congedo, M.; Jutten, C. Multiclass Brain–Computer Interface Classification by Riemannian geometry. IEEE Trans. Biomed. Eng. 2012, 59, 920–928. [Google Scholar] [CrossRef] [Scilit]
  108. Bockbrader, M.A.; Francisco, G.; Lee, R.; Olson, J.; Solinsky, R.; Boninger, M.L. Brain Computer Interfaces in Rehabilitation Medicine. PMR 2018, 10, S233–S243. [Google Scholar] [CrossRef] [Scilit]
  109. Cervera, M.A.; Soekadar, S.R.; Ushiba, J.; Millán, J.D.R.; Liu, M.; Birbaumer, N.; Garipelli, G. Brain-computer interfaces for post-stroke motor rehabilitation: A meta-analysis. Ann. Clin. Transl. Neurol. 2018, 5, 651–663. [Google Scholar] [CrossRef] [Scilit]
  110. Donchin, E.; Spencer, K.M.; Wijesinghe, R. The Mental Prosthesis: Assessing the Speed of a P300-Based Brain–Computer Interface. IEEE Trans. Rehabil. Eng. 2000, 8, 174–179. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  111. Gaur, P.; Pachori, R.B.; Wang, H.; Prasad, G. A multi-class EEG-based BCI classification using multivariate empirical mode decomposition based filtering and Riemannian geometry. Expert Syst. Appl. 2018, 95, 201–211. [Google Scholar] [CrossRef] [Scilit]
  112. Haider, A.; Fazel-Rezai, R. Application of P300 Event-Related Potential in Brain-Computer Interface. In Event-Related Potentials and Evoked Potentials; Sittiprapaporn, P., Ed.; IntechOpen: London, UK, 2017. [Google Scholar]
  113. List, A.; Justus, T.; Robertson, L.C.; Bentin, S. A mismatch negativity study of local-global auditory processing. Brain Res. 2007, 1153, 122–133. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  114. Näätänen, R.; Jacobsen, T.; Wincler, I. Memory-based or afferent processes in mismatch negativity (MMN): A review of the evidence. Psychophysiology 2005, 42, 25–32. [Google Scholar] [CrossRef] [Scilit]
  115. List, A.; Grabowecky, M.; Suzuki, S. Local and global level-priming occurs for hierarchical stimuli composed of outlined, but not filled-in, elements. J. Vis. 2013, 13, 23. [Google Scholar] [CrossRef] [Scilit]
  116. Lotte, F.; Condego, M.; Lécuyer, A.; Lamarche, F.; Arnaldi, B. A review of classification algorithms for EEG-based brain-computer interfaces. J. Neural Eng. 2007, 4, R1–R13. [Google Scholar] [CrossRef]
  117. Tang, Z.; Sun, S.; Zhang, S.; Chen, Y.; Li, C.; Chen, S. A Brain-Machine Interface Based on ERD/ERS for an Upper-Limb Exoskeleton Control. Sensors 2016, 16, 2050. [Google Scholar] [CrossRef] [Scilit]
  118. Tayeb, Z.; Fedjaev, J.; Ghaboosi, N.; Richter, C.; Everding, L.; Qu, X.; Wu, Y.; Cheng, G.; Conradt, J. Validating Deep Neural Networks for Online Decoding of Motor Imagery Movements from EEG Signals. Sensors 2019, 19, 210. [Google Scholar] [CrossRef] [Scilit]
  119. Piccione, F.; Giorgi, F.; Tonin, P.; Priftis, K.; Giove, S.; Silvoni, S.; Palmas, G.; Beverina, F. P300-based brain computer interface: Reliability and performance in healthy and paralysed participants. Clin. Neurophysiol. 2006, 117, 531–537. [Google Scholar] [CrossRef] [Scilit]
  120. Xu, L.; Xu, M.; Ke, Y.; An, X.; Liu, S.; Ming, D. Cross-Dataset Variability Problem in EEG Decoding with Deep Learning. Front. Hum. Neurosci. 2020, 14, 103. [Google Scholar] [CrossRef] [Scilit]
  121. Won, J.H.; Humphrey, E.L.; Yeager, K.R.; Martinez, A.A.; Robinson, C.H.; Mills, K.E.; Johnstone, P.M.; Moon, I.J.; Woo, J. Relationship among the physiologic channel interactions, spectral-ripple discrimination, and vowel identification in cochlear implant users. J. Acoust. Soc. Am. 2014, 136, 2714–2725. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  122. Alexander, J.D.; Nygaard, L.C. Reading voices and hearing text: Talker-specific auditory imagery in reading. J. Exp. Psychol. Hum. Percept. Perform. 2008, 34, 446–459. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  123. Lei, B.; Liu, X.; Liang, S.; Hang, W.; Wang, Q.; Choi, K.S.; Qin, J. Walking Imagery Evaluation in Brain Computer Interfaces via a Multi-View Multi-Level Deep Polynomial Network. IEEE Trans. Neural Syst. Rehabil. Eng. 2019, 27, 497–506. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  124. Meng, J.; Zhang, S.; Bekyo, A.; Olsoe, J.; Baxter, B.; He, B. Noninvasive Electroencephalogram Based Control of a Robotic Arm for Reach and Grasp Tasks. Sci. Rep. 2016, 6, 38565. [Google Scholar] [CrossRef] [Scilit]
  125. Pratticò, D.; Laganà, F. Infrared Thermographic Signal Analysis of Bioactive Edible Oils Using CNNsfor Quality Assessment. Signals 2025, 6, 38. [Google Scholar] [CrossRef] [Scilit]
  126. Laganà, F.; Pratticò, D.; Quattrone, M.F.; Pullano, S.A.; Calcagno, S. Hybrid AI–Taguchi–ANOVA Approach for Thermographic Monitoring of Electronic Devices. Eng 2026, 7, 28. [Google Scholar] [CrossRef] [Scilit]
  127. Belwafi, K.; Romain, O.; Gannouni, S.; Ghaffari, F.; Djemal, R.; Ouni, B. An embedded implementation based on adaptive filter bank for brain–computer interface systems. J. Neurosci. Methods 2018, 305, 1–16. [Google Scholar] [CrossRef] [Scilit]
  128. Guan, S.; Zhao, K.; Yang, S. Motor Imagery EEG Classification Based on Decision Tree Framework and Riemannian Geometry. Comput. Intell. Neurosci. 2019, 2019, 5627156. [Google Scholar] [CrossRef] [Scilit]
  129. Hartmann, K.G.; Schirrmeister, R.T.; Ball, T. EEG-GAN: Generative adversarial networks for electroencephalograhic (EEG) brain signals. arXiv 2018, arXiv:1806.01875. [Google Scholar] [CrossRef] [Scilit]
Figure 1. General structure of the auditory sensory system. Original. 1—adequate stimulus (acoustic signals), 2—structures of outer and middle ear, 3—specialized receptors (hair cells of the vestibule (inner ear), 4—peripheral sensory neurons—auditory ganglion (VIII pair of cranial nerves), 5—brainstem cochlear nuclei and superior olive, 6—internal geniculate bodies (switching thalamic nuclei), 7—temporal area (primary projection zone of the brain cortex), 8—thalamic cushion (associative thalamic nuclei), 9—temporal area (secondary projection zone of the brain cortex), 10—associative brain cortex, 11—reticular formation of the brainstem, 12—nonspecific thalamic nuclei, 13—motor (4-hill) and autonomic nervous centers, 14—effectors (muscles). Blue arrows—ascending afferent pathways; green dotted arrows—reverse corticofugal afferentation.
Figure 1. General structure of the auditory sensory system. Original. 1—adequate stimulus (acoustic signals), 2—structures of outer and middle ear, 3—specialized receptors (hair cells of the vestibule (inner ear), 4—peripheral sensory neurons—auditory ganglion (VIII pair of cranial nerves), 5—brainstem cochlear nuclei and superior olive, 6—internal geniculate bodies (switching thalamic nuclei), 7—temporal area (primary projection zone of the brain cortex), 8—thalamic cushion (associative thalamic nuclei), 9—temporal area (secondary projection zone of the brain cortex), 10—associative brain cortex, 11—reticular formation of the brainstem, 12—nonspecific thalamic nuclei, 13—motor (4-hill) and autonomic nervous centers, 14—effectors (muscles). Blue arrows—ascending afferent pathways; green dotted arrows—reverse corticofugal afferentation.
Applsci 16 04047 g001
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Lytaev, S. Processing of Amplitude-Temporal Acoustic Parameters in the Auditory System During Signal Coding for Image Recognition: Analytical Review. Appl. Sci. 2026, 16, 4047. https://doi.org/10.3390/app16084047

AMA Style

Lytaev S. Processing of Amplitude-Temporal Acoustic Parameters in the Auditory System During Signal Coding for Image Recognition: Analytical Review. Applied Sciences. 2026; 16(8):4047. https://doi.org/10.3390/app16084047

Chicago/Turabian Style

Lytaev, Sergey. 2026. "Processing of Amplitude-Temporal Acoustic Parameters in the Auditory System During Signal Coding for Image Recognition: Analytical Review" Applied Sciences 16, no. 8: 4047. https://doi.org/10.3390/app16084047

APA Style

Lytaev, S. (2026). Processing of Amplitude-Temporal Acoustic Parameters in the Auditory System During Signal Coding for Image Recognition: Analytical Review. Applied Sciences, 16(8), 4047. https://doi.org/10.3390/app16084047

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop