Deep Beats, Deep Thoughts? Predicting General Cognitive Ability from Natural Music-Listening Behavior
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe present manuscript examines associations between generalized cognitive abilities and features of music that people listen to. The data are based on smartphone assessment, thus, providing an ecologically valid insight into participants' music-listening behaviors. Authors analyze a sample of 185 participants. The findings are in line with expectations, showing minor to negligible relationships between the study variables.
I have read the manuscript closely and found it to be well written and fitting into the scope of the Journal of Intelligence and the special issue. The analyses and interpretations are sound and the manuscript contributes to the knowledge of the field. I have only minor comments that should be addressed before publication is recommended. I congratulate authors on their interesting work!
(1) Please examine whether age and gender are related to the music listening features studied. This allows readers to learn more about the role of demographics for the tested associations.
(2) I appreciate the use of LIWC, which allows for a standardized analysis of the lyrics that might offer information on music preferences in relation to cognitive abilities. However, the last years have shown that LIWC analyses of language and broad domains of individual differences (e.g., Big Five) are limited with regard to generalizability and overall show comparatively small associations (Koutsoumpis et al., 2022; Martínez-Huertas et al., 2022). I suggest noting these critical perspectives to the discussion section, as this allows to contextualize the findings and provide readers with a reason why the present findings should be interpreted cautiously.
Koutsoumpis, A., Oostrom, J. K., Holtrop, D., van Breda, W., Ghassemi, S., & de Vries, R. E. (2022). The kernel of truth in text-based personality assessment: A meta-analysis of the relations between the Big Five and the Linguistic Inquiry and Word Count (LIWC). Psychological Bulletin, 148(11-12), 843–868. https://doi.org/10.1037/bul0000381
Martínez-Huertas, J. Á., Moreno, J. D., Olmos, R., Martínez-Mingo, A., & Jorge-Botana, G. (2022). A failed cross-validation study on the relationship between LIWC linguistic indicators and personality: Exemplifying the lack of generalizability of exploratory studies. Psych, 4(4), 803-815.
(3) Please consider changing "females" to "women" and "males" to "men" in line with gender reporting guidelines by the APA.
(4) Since the data for music preferences were mostly derived from streamers such as Spotify, who provide pre-selections of songs (i.e., playlists), it is to some degree unclear how well the song selection reflects an active way of participants choosing the music analyzed. For example, while writing this review a song plays in the background that automatically followed a song I selected. This is based on a Spotify algorithm, but I have no idea what song or band it is that is playing now nor would I have chosen this actively. I find it bearable enough to not skip it. Similarly, people might not even chose music actively but rely on a playlist or pre-selection by others, which would not be necessarily interpreted as an active choice of song selection. This might be added as a note in the paragraph discussing noise in behavioral data on p. 12 (Second, digital behavioral data..."). Future research might replicate the findings, but investigate participants' favorite songs in relation to the study variables.
Author Response
Comment 1: Please examine whether age and gender are related to the music listening features studied. This allows readers to learn more about the role of demographics for the tested associations.
Response 1: We thank the reviewer for this helpful suggestion and agree that associations with age and gender are of interest to readers and help contextualize our findings. We added a table reporting Pearson correlations between music-listening features and demographic variables (age, gender, and education) to our Online Supplementary Materials (OSM; Table S2), which is now referenced in the manuscript (lines 337-340). For the sake of completeness, we also included descriptive analyses of the associations between demographic variables and GCA in the Results section (lines 335–337). Finally, we expanded the Discussion to more thoroughly address the possibility that age may be a confounding variable in our prediction analyses and to clarify its implications for interpretation (lines 612-615).
Comment 2: I appreciate the use of LIWC, which allows for a standardized analysis of the lyrics that might offer information on music preferences in relation to cognitive abilities. However, the last years have shown that LIWC analyses of language and broad domains of individual differences (e.g., Big Five) are limited with regard to generalizability and overall show comparatively small associations (Koutsoumpis et al., 2022; Martínez-Huertas et al., 2022). I suggest noting these critical perspectives to the discussion section, as this allows to contextualize the findings and provide readers with a reason why the present findings should be interpreted cautiously.
Response 2: Thank you for pointing us towards this important limitation of LIWC-based language analysis! We have expanded our discussion of the feature-level interpretations by explicitly noting that prior research has shown LIWC-based analyses of language to exhibit limited generalizability and comparatively small associations when applied in personality predictions (lines 512-514) and cite the suggested references (Koutsoumpis et al., 2022; Martínez-Huertas et al., 2022). This addition is intended to further contextualize our findings and underscore the need for cautious interpretation of associations involving individual LIWC-based features.
Comment 3: Please consider changing "females" to "women" and "males" to "men" in line with gender reporting guidelines by the APA.
Response 3: We thank the reviewer for clarifying the APA guidelines regarding gender reporting! We have implemented the requested changes (line 179).
Comment 4: Since the data for music preferences were mostly derived from streamers such as Spotify, who provide pre-selections of songs (i.e., playlists), it is to some degree unclear how well the song selection reflects an active way of participants choosing the music analyzed. For example, while writing this review a song plays in the background that automatically followed a song I selected. This is based on a Spotify algorithm, but I have no idea what song or band it is that is playing now nor would I have chosen this actively. I find it bearable enough to not skip it. Similarly, people might not even chose music actively but rely on a playlist or pre-selection by others, which would not be necessarily interpreted as an active choice of song selection. This might be added as a note in the paragraph discussing noise in behavioral data on p. 12 (Second, digital behavioral data..."). Future research might replicate the findings, but investigate participants' favorite songs in relation to the study variables.
Response 4: We appreciate this insightful and well-articulated comment. We fully agree that music listening on streaming platforms often involves passive exposure through playlists or algorithmic recommendations, making it difficult to determine the extent to which individual songs reflect active choice. As the reviewer notes, this limitation applies broadly to contemporary music consumption: most streaming services provide playlists and automated follow-ups, and even album-based listening may include songs that listeners tolerate rather than actively select. At present, it is technically not possible to reliably distinguish between actively selected and automatically played songs using smartphone-based sensing data alone as such differentiation would require access to detailed backend data from streaming service providers (see Anderson et al., 2021). Importantly, however, our focus was on what participants listened to in everyday life rather than on their subjective liking or intentional selection of each song. From this perspective, even passively consumed music constitutes a valid representation of real-world listening behavior, as the music was still played and experienced. We agree that designs focusing on favorite or self-selected songs address a related but distinct research question and may yield complementary insights, though such approaches rely on self-reports that are themselves subject to well-known biases. Following the reviewer’s suggestion, we have now explicitly incorporated this limitation as an example of noise in digital behavioral data in the Discussion section (lines 571-574), thereby further contextualizing our findings and their interpretation.
Reviewer 2 Report
Comments and Suggestions for AuthorsSummary
The manuscript “Deep Beats, Deep Thoughts? Predicting General Cognitive Ability from Natural Music-Listening Behavior” uses a smartphone-based app to assess how real-world music listening behaviors relate to general cognitive ability (GCA). Using several features (e.g., listening durations, acoustic features, lyrical content) and two prediction methods (LASSO regression and random forest models), the authors found modest predictive performance, at least using the nonlinear random forest models. These results are proof of concept that digital traces such as music listening behaviors can be used to predict GCA
Evaluation
The methodology and analysis are quite impressive. I also think the authors generally do a nice job of not overstating their results. I have a few questions and comments for the authors to consider in a revision.
- The ability to predict GCA from music listening is impressive (even if prediction is quite modest), but there needs to be a greater attempt at unpacking the theoretical rationale for why these predictors might matter in GCA in the discussion. This is critical in better understanding the generalizability of these findings beyond the present sample. For example, why might “liveness” negatively predict GCA? The authors begin to address this (L473-475), but the explanation they provide seems to apply more to contextual cognitive performance (e.g., speculating that live music might be less suited to focused listening) rather than GCA.
- Why do you think the random forest models performed better than the LASSO regressions? Their nonlinear nature? Something else? What does this tell us (if anything) about the relationship between GCA and music listening?
- L485-486: The authors claim that “much of contemporary German-language music in the 2020s has been dominated by hip hop.” Upon what is this claim based? The authors’ opinion? Examining top songs (e.g., billboard charts)? Streaming data? Something else?
- L487-490: I might also suggest that the predictive relevance of German-language music could relate to individuals’ linguistic experiences outside of their native language, which, in turn, could be associated with GCA.
- L517-519: Although I understand the speculative nature of this section, I would make it quite clear that such an application (e.g., using changes in music listening behaviors to infer something about changes in cognitive ability) would assume that the associations in the present study are causal, which we do not know.
Author Response
Comment 1: The ability to predict GCA from music listening is impressive (even if prediction is quite modest), but there needs to be a greater attempt at unpacking the theoretical rationale for why these predictors might matter in GCA in the discussion. This is critical in better understanding the generalizability of these findings beyond the present sample. For example, why might “liveness” negatively predict GCA? The authors begin to address this (L473-475), but the explanation they provide seems to apply more to contextual cognitive performance (e.g., speculating that live music might be less suited to focused listening) rather than GCA.
Response 1: We thank the reviewer for this thoughtful comment and agree that theoretical interpretation is crucial for understanding the generalizability of predictive findings. At the same time, we deliberately adopted a cautious interpretative stance in this manuscript. Given the modest overall predictive performance, the inherent noise in digital behavioral data, the possibility of nonlinear effects, and the general instability of feature-importance estimates across resampling iterations, we sought to avoid over-interpreting individual predictors or attributing strong theoretical meaning to single features. Moreover, the outcome variable general cognitive ability represents a broad construct derived from multiple subdomains. This makes it difficult to provide precise theoretical explanations for why specific behavioral features (e.g., liveness of recordings) would relate to GCA as a whole rather than to more narrowly defined cognitive processes. As such, we believe that detailed, feature-specific interpretations may risk overstating the implications of exploratory patterns that may as well be sample- and model-dependent. We added an additional sentence to the discussion that explicitly states these concerns regarding feature-level interpretations (lines 509-511). That said, we provide theory-informed interpretations of exemplary features at a more general level, following the uses and gratifications theory that we also introduced in our introduction. We made this interpretation framework more explicit (lines 484-490).
Comment 2: Why do you think the random forest models performed better than the LASSO regressions? Their nonlinear nature? Something else? What does this tell us (if anything) about the relationship between GCA and music listening?
Response 2: Thank you for this important question. There are likely multiple, complementary reasons why the random forest models outperformed the LASSO regressions in our analyses. On the one hand, methodological factors unrelated to the substantive nature of the associations likely played a role. In our cross-validated setting with a relatively small effective sample size per training fold, the LASSO tended to select strong regularization parameters to avoid overfitting. This often resulted in substantial coefficient shrinkage and, in some iterations, intercept-only models. Moreover, when predictors are highly correlated, which is the case for many audio and lyric features, LASSO typically selects one arbitrary feature from a correlated group or shrinks all coefficients toward zero, which can wash out weak distributed signal. Linear models are also particularly sensitive to noise in predictors, which is common in digital behavioral data as described in our discussion. On the other hand, it is also well plausible that the associations detected by the random forest but not by the LASSO were of a form that linear models could not capture. Random forests can model nonlinear, piecewise, and interaction effects without explicit specification, whereas the LASSO assumes additive linear relationships unless interactions or nonlinear terms are manually engineered. Thus, the superior performance of the random forest may indicate that the relationship between music-listening behavior and GCA is weak, conditional, or interaction-based rather than well described by simple linear effects. Taken together, these findings suggest that the predictive signal linking music listening and GCA is both modest and structurally complex, which may explain why it was detectable only by a flexible, ensemble-based model. We have clarified this interpretation in the revised Results (lines 345-351) and Discussion (lines 415-417).
Comment 3: L485-486: The authors claim that “much of contemporary German-language music in the 2020s has been dominated by hip hop.” Upon what is this claim based? The authors’ opinion? Examining top songs (e.g., billboard charts)? Streaming data? Something else?
Response 3: We thank the reviewer for raising this concern. The statement regarding the dominance of German-language hip hop in the 2020s was based on general observations of chart and streaming trends rather than on a single, clearly citable academic source. As we could not identify a suitable reference that would meet the standards of an academic publication, we have removed this sentence from the manuscript. Instead, we now focus on the alternative interpretation approach suggested in the follow-up comment below, which allows us to discuss the finding without relying on the genre argument.
Comment 4: L487-490: I might also suggest that the predictive relevance of German-language music could relate to individuals’ linguistic experiences outside of their native language, which, in turn, could be associated with GCA.
Response 4: We thank the reviewer for pointing us in this direction. We believe this is a very reasonable and more direct explanation of the relevance of German-language music than the one we initially provided, which posited musical genre as a mediator. We rephrased our interpretation attempt in the discussion to account for verbal abilities (lines 500-504).
Comment 5: L517-519: Although I understand the speculative nature of this section, I would make it quite clear that such an application (e.g., using changes in music listening behaviors to infer something about changes in cognitive ability) would assume that the associations in the present study are causal, which we do not know.
Response 5: We appreciate the reviewer’s thoughtful comment. While we had briefly addressed the issue of causality in the original version of the manuscript, we agree that this topic warrants more explicit and prominent discussion. In response, we added a dedicated fifth challenge paragraph that clearly emphasizes the correlational nature of the present findings and explicitly cautions against causal interpretations or prescriptive applications based on music-listening behavior. This new paragraph addresses the assumptions required for causal inference and outlines why such conclusions cannot be drawn from the current study (lines 607- 618).
Reviewer 3 Report
Comments and Suggestions for AuthorsI want to congratulate authors for the intuition and effort to connect everyday music listening with intelligence. As my task is to review your work, I will make some comments to clarify or ask about some aspects:
1) Please, could you provide with more information about the reasons involving the period lapse between the first survey and the fifth survey of Intelligence in september?
2) What is the criterion for the duration of time response to exclude participants due to lack of attention?
3) When your study shows that the Random forest model yields the best performance, I wonder why you did not report any significance related to the Spearman correlations.
4) Did you know why the original study used only the short version of the Cattell Horn Carroll questionnaire? I consider that assessing the fourth dimension of visual processing could enhance the results.
5) Did you find any specific correlation between the three intellectual dimensions and everyday music listening?
6) I consider the Discussion section very well built, with space for confrontation, reflection, implications and limitations.
Author Response
Comment 1: Please, could you provide with more information about the reasons involving the period lapse between the first survey and the fifth survey of Intelligence in september?
Response 1: We thank the reviewer for this question and are happy to clarify the study design. The data used in the present analyses originate from a large-scale longitudinal study that combined six months of continuous mobile sensing with six distinct monthly surveys. To minimize participant burden and avoid extremely long questionnaires at any single measurement occasion, different psychological constructs were distributed across surveys. For example, demographic information was assessed in Survey 1, personality traits in Survey 2, and cognitive ability in Survey 5. Importantly, the intelligence test was administered only once and was included in Survey 5 by random assignment. It was not assessed earlier in the study. Consequently, the time lapse between the assessment of demographics and cognitive ability reflects the overall study design aimed at reducing burden rather than a theoretical or methodological rationale related to intelligence assessment. We have further clarified this point in the manuscript to avoid potential confusion (lines 160-162).
Comment 2: What is the criterion for the duration of time response to exclude participants due to lack of attention?
Response 2: We agree with the reviewer that this issue merits further clarification. Because the mobile version of the INT is still a relatively novel assessment instrument, we could not deduce a cutoff for response-time–based exclusion criteria in this instrument. More generally, we could not identify a consensus in the literature regarding fixed response-time thresholds for ability tests. Given this lack of established benchmarks, we adopted a conservative, two-step exclusion procedure applied separately within each subtest. First, we identified participants who showed more repeated use of the same response option than would be expected based on the number of repeated correct response options in the respective subtest (allowing the expected number of repetitions under perfect performance plus one additional repetition). Only participants flagged by this criterion were then evaluated with respect to response times. Specifically, we defined implausibly short response times as values below one third of the median response time of the fastest item in the respective subtest (calculated across the full sample), and excluded participants who showed such response times on at least three items within that subtest. We included details on the reponse times in the methods sections (lines 171-175). Importantly, this procedure was intentionally non-strict. Our aim was to identify only clear cases of inattentive responding while minimizing the risk of falsely excluding attentive participants, especially given the large variability in response times observed in mobile testing environments. As a result, only a very small number of participants were excluded using this combined criterion (n = 3).
Comment 3: When your study shows that the Random forest model yields the best performance, I wonder why you did not report any significance related to the Spearman correlations.
Response 3: We thank the reviewer for this thoughtful question. While there is a relatively well-established significance test for comparing model performance against the featureless baseline (see Stachl et al., 2020), we deliberately refrained from applying it. We did not report significance tests for the Spearman correlations between predicted and observed GCA scores for several reasons. First, the present study is exploratory in nature and was not preregistered with confirmatory hypotheses regarding specific effect sizes or prediction thresholds. Accordingly, our primary goal was to evaluate out-of-sample predictive performance rather than to conduct null-hypothesis significance testing on individual performance metrics. Second, machine-learning approaches evaluated via cross-validation are generally not designed to test the statistical significance of prediction coefficients or performance measures in the traditional inferential sense. Instead, they focus on estimating generalization performance through resampling-based evaluation (e.g., cross-validation), which yields distributions of performance metrics rather than single estimates. In this framework, descriptive summaries of out-of-sample performance (e.g., median Spearman correlations) are typically reported without accompanying p-values. For these reasons, we focused on reporting descriptive out-of-sample performance estimates and model comparisons where appropriate, rather than inferential significance tests for individual correlation coefficients. We have clarified this rationale in the revised manuscript to avoid confusion. We added a corresponding statement in our methods section (lines 295-298).
Reference: Stachl, C., Au, Q., Schoedel, R., Gosling, S. D., Harari, G. M., Buschek, D., Völkel, S. T., Schuwerk, T., Oldemeier, M., Ullmann, T., Hussmann, H., Bischl, B., & Bühner, M. (2020). Predicting personality from patterns of behavior collected with smartphones. Proceedings of the National Academy of Sciences, 117(30), 17680-17687. https://doi.org/10.1073/pnas.1920484117
Comment 4: Did you know why the original study used only the short version of the Cattell Horn Carroll questionnaire? I consider that assessing the fourth dimension of visual processing could enhance the results.
Response 4: This is a well-justified question. At the time the present study was conducted, the mobile version of the Cattell–Horn–Carroll–based test was still under development, and the fourth dimension (visual processing) had not yet been finalized for smartphone-based administration. Consequently, this dimension was not available for assessment and could not be included in the present analyses. We agree that assessing visual processing would be a valuable extension in this context and may enhance future investigations. We have now made this limitation more explicit in the manuscript (lines 191-193).
Comment 5: Did you find any specific correlation between the three intellectual dimensions and everyday music listening?
Response 5: Thank you for this warranted question. We initially planned to examine associations and prediction performance at the subtest level as well. However, the short mobile version of the cognitive ability test exhibited relatively low internal consistencies for the individual subscales (Cronbach’s α = .56 for fluid reasoning, .70 for comprehension knowledge, and .64 for quantitative knowledge). Given that these reliabilities fall below commonly accepted thresholds, we concluded that the subscale scores were not sufficiently reliable for meaningful predictive analyses. Measurement error at this level would substantially attenuate associations and impair both linear and machine-learning–based prediction models (see, e.g., Jacobucci et al., 2020). In contrast, the combined general cognitive ability (GCA) score demonstrated satisfactory internal consistency (α = .80). We therefore focused all reported analyses on the full GCA measure, which we consider the most psychometrically sound outcome in the present dataset. We agree that examining relations between specific cognitive domains and everyday music listening would be highly informative, and we see this as an important direction for future research using more reliable subscale measures.
Reference: Jacobucci, R., & Grimm, K. J. (2020). Machine learning and psychological research: The unexplored effect of measurement. Perspectives on Psychological Science, 15(3), 809-816. https://doi.org/10.1177/1745691620902467
Comment 6: I consider the Discussion section very well built, with space for confrontation, reflection, implications and limitations.
Response 6: We thank the reviewer for this positive and encouraging feedback on the Discussion section. We are pleased that the balance between interpretation, implications, and limitations was perceived as appropriate.
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsI thank the authors for addressing my comments and I recommend publishing the manuscript in its present form. Congratulations on this interesting study and the contribution to the knowledge of the field!
Reviewer 2 Report
Comments and Suggestions for AuthorsI appreciate the authors' responses to my initial set of comments. I am happy to endorse this version of the manuscript for publication.
Reviewer 3 Report
Comments and Suggestions for AuthorsI am fully satisfied with your answers.

