Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (11)

Search Parameters:
Keywords = melody contour

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
25 pages, 1115 KB  
Article
Controllable Symbolic Music Generation via Stage-Aware Style Routing and Differentiable Melody Regularization
by Xuanfei Zhou, Yinxuan Huang, Sining Han, Jiangyao Bai, Qianzhen Zhang, Lailong Luo and Chen Wang
Information 2026, 17(6), 568; https://doi.org/10.3390/info17060568 - 8 Jun 2026
Viewed by 323
Abstract
Controllable symbolic music generation must preserve a reference melody while remaining responsive to style prompts. Existing hierarchical diffusion systems typically reuse a shared condition vector across harmony, rhythm, and timbre stages, which can entangle stylistic factors and weaken melody preservation. We present HCDMG++, [...] Read more.
Controllable symbolic music generation must preserve a reference melody while remaining responsive to style prompts. Existing hierarchical diffusion systems typically reuse a shared condition vector across harmony, rhythm, and timbre stages, which can entangle stylistic factors and weaken melody preservation. We present HCDMG++, a hierarchical diffusion framework that addresses these two limitations through stage-aware style routing and differentiable melody regularization. The routing module uses a residual multi-layer perceptron (MLP) with zero-initialized scalar gates to project text-derived style embeddings into harmony-, rhythm-, and timbre-specific subspaces, whereas the regularization branch aligns soft pitch histograms and contour trajectories with the conditioning melody during training without breaking the differentiable computation graph. We evaluate the integrated system on a 384-sample benchmark covering four melodies, eight styles, four random seeds, and three denoising budgets, supplemented by a matched legacy-compatible reference and inference-time component ablation that contrasts legacy behavior, silenced gates, an automated uniform gamma routing sweep, and the full forward pass. HCDMG++ produces valid four-track outputs in all 384 runs, reaches a peak pitch histogram similarity score of 0.508 under a 64-step budget, and improves pitch histogram alignment over Legacy-HCDMG by roughly two orders of magnitude on the matched slice, while attaining a positive Fisher-style style separability score where the legacy benchmark is too sparse to support one. These results indicate that stage-specific conditioning and differentiable structural guidance jointly improve controllability in symbolic music diffusion, while also exposing the remaining limitations in long-form generalization and perceptual validation, which motivate the future work outlined at the end of this paper. Full article
(This article belongs to the Section Information Applications)
Show Figures

Figure 1

18 pages, 1861 KB  
Article
The Interplay between Syllabic Duration and Melody to Indicate Prosodic Functions in Brazilian Portuguese Story Retelling
by Plinio A. Barbosa and Luís H. G. Alvarenga
Languages 2024, 9(8), 268; https://doi.org/10.3390/languages9080268 - 1 Aug 2024
Cited by 2 | Viewed by 2648
Abstract
This paper investigates the relationship between syllabic duration and F0 contours for implementing three prosodic functions. Work on rhythm usually describes the evolution of syllable-sized durations throughout utterances, rarely making reference to melodic events. On the other hand, work on intonation usually describes [...] Read more.
This paper investigates the relationship between syllabic duration and F0 contours for implementing three prosodic functions. Work on rhythm usually describes the evolution of syllable-sized durations throughout utterances, rarely making reference to melodic events. On the other hand, work on intonation usually describes linear sequences of melodic events with indirect references to duration. Although some scholars have explored the relationship between these two parameters for particular functions, to our knowledge, there has been no investigation on the systematic correlation between syllabic duration and F0 values throughout narrative sequences. Based on a corpus of story retelling with nine speakers of Brazilian Portuguese from two regions, our work investigated the interplay between syllabic duration and melody to signal three prosodic functions: terminal and non-terminal boundary marking and prominence. The examination of local syllabic duration maxima and four F0 descriptors revealed that these maxima act as landmarks for particular F0 shapes: for non-terminal boundaries, the great majority of shapes were increasing and increasing–decreasing patterns; for terminal boundaries, almost all shapes were decreasing F0 patterns; and for prominence marking, the great majority of shapes were high tones across the stressed syllable. Time series analyses revealed significant correlations between duration and specific F0 descriptors, pointing to a ruled interplay between F0 and syllabic duration patterns in Brazilian Portuguese story retelling. Full article
(This article belongs to the Special Issue Phonetics and Phonology of Ibero-Romance Languages)
Show Figures

Figure 1

18 pages, 3363 KB  
Article
A Concert-Based Study on Melodic Contour Identification among Varied Hearing Profiles—A Preliminary Report
by Razvan Paisa, Jesper Andersen, Francesco Ganis, Lone M. Percy-Smith and Stefania Serafin
J. Clin. Med. 2024, 13(11), 3142; https://doi.org/10.3390/jcm13113142 - 27 May 2024
Cited by 1 | Viewed by 1559
Abstract
Background: This study investigated how different hearing profiles influenced melodic contour identification (MCI) in a real-world concert setting with a live band including drums, bass, and a lead instrument. We aimed to determine the impact of various auditory assistive technologies on music [...] Read more.
Background: This study investigated how different hearing profiles influenced melodic contour identification (MCI) in a real-world concert setting with a live band including drums, bass, and a lead instrument. We aimed to determine the impact of various auditory assistive technologies on music perception in an ecologically valid environment. Methods: The study involved 43 participants with varying hearing capabilities: normal hearing, bilateral hearing aids, bimodal hearing, single-sided cochlear implants, and bilateral cochlear implants. Participants were exposed to melodies played on a piano or accordion, with and without an electric bass as a masker, accompanied by a basic drum rhythm. Bayesian logistic mixed-effects models were utilized to analyze the data. Results: The introduction of an electric bass as a masker did not significantly affect MCI performance for any hearing group when melodies were played on the piano, contrary to its effect on accordion melodies and previous studies. Greater challenges were observed with accordion melodies, especially when accompanied by an electric bass. Conclusions: MCI performance among hearing aid users was comparable to other hearing-impaired profiles, challenging the hypothesis that they would outperform cochlear implant users. A cohort of short melodies inspired by Western music styles was developed for future contour identification tasks. Full article
(This article belongs to the Special Issue Advances in the Diagnosis, Treatment, and Prognosis of Hearing Loss)
Show Figures

Figure 1

12 pages, 363 KB  
Article
Criterion-Related Validation of a Music-Based Attention Assessment for Individuals with Traumatic Brain Injury
by Eunju Jeong and Susan J. Ireland
Int. J. Environ. Res. Public Health 2022, 19(23), 16285; https://doi.org/10.3390/ijerph192316285 - 5 Dec 2022
Viewed by 3658
Abstract
The music-based attention assessment (MAA) is a melody contour identification task that evaluates different types of attention. Previous studies have examined the psychometric and physiological validity of the MAA across various age groups in clinical and typical populations. The purpose of this study [...] Read more.
The music-based attention assessment (MAA) is a melody contour identification task that evaluates different types of attention. Previous studies have examined the psychometric and physiological validity of the MAA across various age groups in clinical and typical populations. The purpose of this study was to confirm the MAA’s criterion validity in individuals with traumatic brain injury (TBI) and to correlate this with standardized neuropsychological measurements. The MAA and various neurocognitive tests (i.e., the Wechsler adult intelligence scale DST, Delis–Kaplan executive functioning scale color-word interference test, and Conner’s continuous performance test) were administered to 38 patients within two weeks prior to or post to the MAA administration. Significant correlations between MAA and neurocognitive batteries were found, indicating the potential of MAA as a valid measure of different types of attention deficits. An additional multiple regression analysis revealed that MAA was a significant factor in predicting attention ability. Full article
(This article belongs to the Special Issue Music for Health Care and Well-Being)
28 pages, 6632 KB  
Article
Prosodic Transfer in Contact Varieties: Vocative Calls in Metropolitan and Basaá-Cameroonian French
by Fatima Hamlaoui, Marzena Żygis, Jonas Engelmann and Sergio I. Quiroz
Languages 2022, 7(4), 285; https://doi.org/10.3390/languages7040285 - 7 Nov 2022
Cited by 1 | Viewed by 7182
Abstract
This paper examines the production of vocative calls in (Northern) Metropolitan French (MF) and Cameroonian French (CF) as it is spoken by native speakers of a tone language, Basaá. While the results of our Discourse Completion Task confirm previous descriptions of MF, they [...] Read more.
This paper examines the production of vocative calls in (Northern) Metropolitan French (MF) and Cameroonian French (CF) as it is spoken by native speakers of a tone language, Basaá. While the results of our Discourse Completion Task confirm previous descriptions of MF, they also further our understanding of the relationship between pragmatics and prosody across different groups of French speakers. MF favors the vocative chant in routine contexts and a rising-falling contour in urgent contexts. In contrast, context has little influence on the choice of contour in CF. A melody consisting of the surface realization of lexical tones is produced in both contexts. Regarding acoustic parameters, context only exerts a significant effect on the loudness of vocative calls (RMS amplitude) and has little effect on their F0 height, F0 range and duration. A target-use of vocative calls in CF thus does not amount to target-like use of the original standard target language, MF. Our results provide novel evidence for the transfer of lexical tones onto the contact variety of an intonation language. They also corroborate previous studies involving the pragmatics-prosody interface: the more marked a prosodic pattern is (here, the vocative chant), the more difficult it is to acquire. Full article
Show Figures

Figure 1

19 pages, 4204 KB  
Article
Singing Transcription from Polyphonic Music Using Melody Contour Filtering
by Zhuang He and Yin Feng
Appl. Sci. 2021, 11(13), 5913; https://doi.org/10.3390/app11135913 - 25 Jun 2021
Cited by 3 | Viewed by 4556
Abstract
Automatic singing transcription and analysis from polyphonic music records are essential in a number of indexing techniques for computational auditory scenes. To obtain a note-level sequence in this work, we divide the singing transcription task into two subtasks: melody extraction and note transcription. [...] Read more.
Automatic singing transcription and analysis from polyphonic music records are essential in a number of indexing techniques for computational auditory scenes. To obtain a note-level sequence in this work, we divide the singing transcription task into two subtasks: melody extraction and note transcription. We construct a salience function in terms of harmonic and rhythmic similarity and a measurement of spectral balance. Central to our proposed method is the measurement of melody contours, which are calculated using edge searching based on their continuity properties. We calculate the mean contour salience by separating melody analysis from the adjacent breakpoint connective strength matrix, and we select the final melody contour to determine MIDI notes. This unique method, combining audio signals with image edge analysis, provides a more interpretable analysis platform for continuous singing signals. Experimental analysis using Music Information Retrieval Evaluation Exchange (MIREX) datasets shows that our technique achieves promising results both for audio melody extraction and polyphonic singing transcription. Full article
(This article belongs to the Special Issue Advances in Computer Music)
Show Figures

Figure 1

15 pages, 2327 KB  
Article
Is It Speech or Song? Effect of Melody Priming on Pitch Perception of Modified Mandarin Speech
by Chen-Gia Tsai and Chia-Wei Li
Brain Sci. 2019, 9(10), 286; https://doi.org/10.3390/brainsci9100286 - 22 Oct 2019
Cited by 8 | Viewed by 5028
Abstract
Tonal languages make use of pitch variation for distinguishing lexical semantics, and their melodic richness seems comparable to that of music. The present study investigated a novel priming effect of melody on the pitch processing of Mandarin speech. When a spoken Mandarin utterance [...] Read more.
Tonal languages make use of pitch variation for distinguishing lexical semantics, and their melodic richness seems comparable to that of music. The present study investigated a novel priming effect of melody on the pitch processing of Mandarin speech. When a spoken Mandarin utterance is preceded by a musical melody, which mimics the melody of the utterance, the listener is likely to perceive this utterance as song. We used functional magnetic resonance imaging to examine the neural substrates of this speech-to-song transformation. Pitch contours of spoken utterances were modified so that these utterances can be perceived as either speech or song. When modified speech (target) was preceded by a musical melody (prime) that mimics the speech melody, a task of judging the melodic similarity between the target and prime was associated with increased activity in the inferior frontal gyrus (IFG) and superior/middle temporal gyrus (STG/MTG) during target perception. We suggest that the pars triangularis of the right IFG may allocate attentional resources to the multi-modal processing of speech melody, and the STG/MTG may integrate the phonological and musical (melodic) information of this stimulus. These results are discussed in relation to subvocal rehearsal, a speech-to-song illusion, and song perception. Full article
(This article belongs to the Special Issue Advances in the Neurocognition of Music and Language)
Show Figures

Figure 1

13 pages, 784 KB  
Article
Generation of Melodies for the Lost Chant of the Mozarabic Rite
by Darrell Conklin and Geert Maessen
Appl. Sci. 2019, 9(20), 4285; https://doi.org/10.3390/app9204285 - 12 Oct 2019
Cited by 6 | Viewed by 6394
Abstract
Prior to the establishment of the Roman rite with its Gregorian chant, in the Iberian Peninsula and Southern France the Mozarabic rite, with its own tradition of chant, was dominant from the sixth until the eleventh century. Few of these chants are preserved [...] Read more.
Prior to the establishment of the Roman rite with its Gregorian chant, in the Iberian Peninsula and Southern France the Mozarabic rite, with its own tradition of chant, was dominant from the sixth until the eleventh century. Few of these chants are preserved in pitch readable notation and thousands exist only in manuscripts using adiastematic neumes which specify only melodic contour relations and not exact intervals. Though their precise melodies appear to be forever lost it is possible to use computational machine learning and statistical sequence generation methods to produce plausible realizations. Pieces from the León antiphoner, dating from the early tenth century, were encoded into templates then instantiated by sampling from a statistical model trained on pitch-readable Gregorian chants. A concert of ten Mozarabic chant realizations was performed at a music festival in the Netherlands. This study shows that it is possible to construct realizations for incomplete ancient cultural remnants using only partial information compiled into templates, combined with statistical models learned from extant pieces to fill the templates. Full article
(This article belongs to the Special Issue Sound and Music Computing -- Music and Interaction)
Show Figures

Figure 1

17 pages, 1516 KB  
Article
Joint Detection and Classification of Singing Voice Melody Using Convolutional Recurrent Neural Networks
by Sangeun Kum and Juhan Nam
Appl. Sci. 2019, 9(7), 1324; https://doi.org/10.3390/app9071324 - 29 Mar 2019
Cited by 92 | Viewed by 10669
Abstract
Singing melody extraction essentially involves two tasks: one is detecting the activity of a singing voice in polyphonic music, and the other is estimating the pitch of a singing voice in the detected voiced segments. In this paper, we present a joint detection [...] Read more.
Singing melody extraction essentially involves two tasks: one is detecting the activity of a singing voice in polyphonic music, and the other is estimating the pitch of a singing voice in the detected voiced segments. In this paper, we present a joint detection and classification (JDC) network that conducts the singing voice detection and the pitch estimation simultaneously. The JDC network is composed of the main network that predicts the pitch contours of the singing melody and an auxiliary network that facilitates the detection of the singing voice. The main network is built with a convolutional recurrent neural network with residual connections and predicts pitch labels that cover the vocal range with a high resolution, as well as non-voice status. The auxiliary network is trained to detect the singing voice using multi-level features shared from the main network. The two optimization processes are tied with a joint melody loss function. We evaluate the proposed model on multiple melody extraction and vocal detection datasets, including cross-dataset evaluation. The experiments demonstrate how the auxiliary network and the joint melody loss function improve the melody extraction performance. Furthermore, the results show that our method outperforms state-of-the-art algorithms on the datasets. Full article
(This article belongs to the Special Issue Digital Audio and Image Processing with Focus on Music Research)
Show Figures

Figure 1

18 pages, 2192 KB  
Article
Cognitive Load Changes during Music Listening and its Implication in Earcon Design in Public Environments: An fNIRS Study
by Eunju Jeong, Hokyoung Ryu, Geonsang Jo and Jaehyeok Kim
Int. J. Environ. Res. Public Health 2018, 15(10), 2075; https://doi.org/10.3390/ijerph15102075 - 21 Sep 2018
Cited by 10 | Viewed by 6534
Abstract
A key for earcon design in public environments is to incorporate an individual’s perceived level of cognitive load for better communication. This study aimed to examine the cognitive load changes required to perform a melodic contour identification task (CIT). While healthy college students [...] Read more.
A key for earcon design in public environments is to incorporate an individual’s perceived level of cognitive load for better communication. This study aimed to examine the cognitive load changes required to perform a melodic contour identification task (CIT). While healthy college students (N = 16) were presented with five CITs, behavioral (reaction time and accuracy) and cerebral hemodynamic responses were measured using functional near-infrared spectroscopy. Our behavioral findings showed a gradual increase in cognitive load from CIT1 to CIT3 followed by an abrupt increase between CIT4 (i.e., listening to two concurrent melodic contours in an alternating manner and identifying the direction of the target contour, p < 0.001) and CIT5 (i.e., listening to two concurrent melodic contours in a divided manner and identifying the directions of both contours, p < 0.001). Cerebral hemodynamic responses showed a congruent trend with behavioral findings. Specific to the frontopolar area (Brodmann’s area 10), oxygenated hemoglobin increased significantly between CIT4 and CIT5 (p < 0.05) while the level of deoxygenated hemoglobin decreased. Altogether, the findings indicate that the cognitive threshold for young adults (CIT5) and appropriate tuning of the relationship between timbre and pitch contour can lower the perceived cognitive load and, thus, can be an effective design strategy for earcon in a public environment. Full article
Show Figures

Figure 1

21 pages, 2858 KB  
Article
Analyzing Free-Hand Sound-Tracings of Melodic Phrases
by Tejaswinee Kelkar and Alexander Refsum Jensenius
Appl. Sci. 2018, 8(1), 135; https://doi.org/10.3390/app8010135 - 18 Jan 2018
Cited by 22 | Viewed by 11570
Abstract
In this paper, we report on a free-hand motion capture study in which 32 participants ‘traced’ 16 melodic vocal phrases with their hands in the air in two experimental conditions. Melodic contours are often thought of as correlated with vertical movement (up and [...] Read more.
In this paper, we report on a free-hand motion capture study in which 32 participants ‘traced’ 16 melodic vocal phrases with their hands in the air in two experimental conditions. Melodic contours are often thought of as correlated with vertical movement (up and down) in time, and this was also our initial expectation. We did find an arch shape for most of the tracings, although this did not correspond directly to the melodic contours. Furthermore, representation of pitch in the vertical dimension was but one of a diverse range of movement strategies used to trace the melodies. Six different mapping strategies were observed, and these strategies have been quantified and statistically tested. The conclusion is that metaphorical representation is much more common than a ‘graph-like’ rendering for such a melodic sound-tracing task. Other findings include a clear gender difference for some of the tracing strategies and an unexpected representation of melodies in terms of a small object for some of the Hindustani music examples. The data also show a tendency of participants moving within a shared ‘social box’. Full article
(This article belongs to the Special Issue Sound and Music Computing)
Show Figures

Graphical abstract

Back to TopTop