1. Introduction
News recommendation serves as a core technology for modern content platforms. It directly affects user engagement and platform revenue. Poor recommendations cause users to lose interest, click less, and spend less time on the platform. Over time, this leads to user churn. Biased recommendations also reduce content diversity and fairness. Niche articles and socially important stories struggle to get attention. Therefore, improving both accuracy and fairness in news recommendations is important for platform operators and content providers.
The main research problem is how to model user reading preferences from click history and news content. A good solution must handle two challenges. First, news text contains both stable information and sudden, time-sensitive signals. The stable part includes general topics and writing style. The sudden part includes breaking events, new entities, and changes in public attention. Second, popular news gets most of the clicks. This creates a feedback loop. Popular articles become more dominant in training data. Cold-start or long-tail articles get little exposure, even when they are relevant. A good recommender must be accurate, timely, and fair.
Many methods have been proposed for news recommendation. Early work used attention to extract features from user clicks. Zhu et al. [
1] combined CNNs with attention and RNNs to model user interests. Qi et al. [
2] designed a candidate-aware network that builds user representations for each candidate item. Kim et al. [
3] used a context-aware network with a selection module to enrich headlines with body text. Recent work has explored graph-based models [
4] and multi-interest learning [
5]. These methods have pushed the field forward.
Despite this progress, existing methods have two key limitations. First, none of them leverage frequency-domain analysis for text encoding. As illustrated in
Figure 1, news headlines inherently contain both low-frequency smooth semantics (e.g., general context) and high-frequency abrupt signals (e.g., named entities and event triggers) [
6,
7]. By stacking pretrained word vectors in token order and applying a fast Fourier transform (FFT) along the sequence dimension, our spectral analysis reveals how much each frequency contributes to the overall semantic meaning. Specifically, an empirical analysis of sampled news titles from the MIND-small dataset (
Figure 1b) demonstrates that high-frequency energy accounts for between
and
of the total magnitude, with an average of
. This indicates that high-frequency peaks, driven by sudden keyword changes, are not negligible noise but a substantial part of the semantic content. Yet, existing attention models treat all words equally and inadvertently smooth out these high-frequency signals. This leads to recommendations that are safe but slow to reflect fresh content [
8,
9]. Second, existing methods lack effective calibration for popularity bias in the embedding space. Popular items dominate the training phase, pushing cold-start items to the tail of the distribution. This geometric bias persists in retrieval and ranking, making it difficult for fresh or niche content to surface even when it matches user interests.
To determine whether the observed high-frequency components are discriminative rather than random noise, we further compute the per-frequency variance of magnitude spectra across news categories. Low-frequency components (index ) have coefficients of variation between 0.12 and 0.18, whereas high-frequency components (index ) exhibit 1.8× higher cross-category variance on average and up to 2.3× higher variance at peak frequencies. This contrast indicates that low-frequency components mainly preserve shared structural context, while high-frequency components contain stronger category-specific variation.
To address the above issues, we propose FDDP-RN, a network with frequency-domain denoising and popularity bias correction. The model has three main innovations. First, we apply FFT to word embeddings and design a truncation-and-scaling filter. This filter boosts high-frequency semantic components and reduces low-frequency noise. To the best of our knowledge, this is the first use of frequency-domain text processing in news recommendation. It addresses the problem that existing encoders smooth over sudden, meaningful signals. Second, we propose a click-response clustering method. It models popularity as a temporal pattern, not a fixed number. We extract features from each item’s click curve, such as growth rate, peak time, and decay slope. We then cluster items into popularity modes, such as cold-start or mainstream. Based on cluster membership, we apply a dynamic norm-scaling factor. This adjusts the embedding magnitude of cold-start items. It changes the geometry of the embedding space so that cold-start items are not at a disadvantage. To the best of our knowledge, this is among the first methods to perform popularity debiasing at the embedding level while supporting items with zero interactions. Third, we use an asymmetric routing design. Historically clicked news items go through the popularity calibration module. Candidate news items do not. This keeps candidate embeddings raw and rich. The asymmetry creates an implicit contrastive effect. The attention model learns to match on semantic direction, not embedding size. This separates relevance from popularity in the user model.
Figure 1.
Motivation for frequency-domain denoising. The combination of conceptual semantics (a) and empirical energy distribution (b) confirms that high-frequency abrupt signals constitute a substantial portion (average ) of the overall text features, which are often overlooked by traditional spatial-domain encoders. The colors in the schematic distinguish smooth low-frequency semantics from abrupt high-frequency signals; the plotted colors in the empirical panel distinguish the corresponding frequency components.
We test FDDP-RN on three public datasets: MIND-small, MIND-large, and Adressa. Results show that our model outperforms state-of-the-art baselines on all metrics and all datasets. We also conduct detailed ablations. For FDD, we compare against low-pass only, high-pass only, band-pass only, fixed scaling, learnable frequency-wise filtering, and no FDD. For PBC, we compare against popularity normalization, IPW, popularity regularization, and no PBC. These tests confirm that the gain comes from our specific designs, not from extra parameters.
Our main contributions are:
We introduce a frequency-domain news encoder tailored to news recommendation. Unlike Fourier-based language models that mainly use spectral transforms for efficient token mixing, the proposed encoder performs explicit frequency partitioning on title word embeddings and reconstructs denoised news representations for recommendation.
We design a truncation-and-scaling frequency denoising mechanism that selectively enhances high-frequency abrupt semantic signals while suppressing redundant low-frequency components, thereby improving the representation of timely and event-sensitive news.
We propose an embedding-space popularity bias correction mechanism for cold-start news. Unlike click-count reweighting or popularity prediction methods, it calibrates the magnitude of cold-start news vectors toward a popularity-derived norm benchmark while preserving semantic direction.
We conduct experiments on MIND and Adressa, together with mechanism-level spectral analysis, propagation-state clustering, ablation studies, diversity analysis, and cold-start fairness diagnostics, to verify both recommendation accuracy and the intended bias-correction behavior.
The rest of this paper is organized as follows.
Section 2 reviews related work and gives background on FFT.
Section 3 describes FDDP-RN in detail.
Section 4 covers training and complexity.
Section 5 presents the experimental setup.
Section 6 shows results and analysis.
Section 7 concludes and discusses limitations and future work.
4. Model Training of FDDP-RN
Following standard training principles in news recommendation, we introduce a negative sampling mechanism during the model optimization stage to alleviate the imbalance between positive and negative sample distributions. It is worth noting that while our overall architecture is designed to score any candidate news
, during the training phase (Algorithm 1), we instantiate the candidate set through this negative sampling. Specifically, a batch
containing
independent user sessions is extracted from the training set
. For each user session
, we construct a localized candidate set consisting of one positive clicked item
and
K randomly sampled negative items
from the unclicked candidates. To accurately characterize user interests, we use the negative log-likelihood function as the optimization objective, which is defined as follows:
where
represents the predicted score of the positive sample, and
denote the predicted scores of the
K corresponding negative samples in the
i-th session. We optimize the model parameters by minimizing the batch-wise negative log-likelihood loss, which is formulated as:
where
denotes the training mini-batch, and
is the batch size.
This loss function aims to increase the scores assigned to positive samples and suppress the interference of negative samples by widening the prediction difference between positive and negative samples, thereby optimizing the model’s performance in discriminating click behavior. This allows both news representation and user representation to fully utilize frequency-domain information and popularity-correction information.
The FDDP-RN training procedure is detailed in Algorithm 1. Initially, word embeddings, popularity parameters, and state vectors are initialized to partition news into the cold-start news collection and the popular news collection (line 1). During each epoch (line 2), the shuffled data is optimized in batches. For each session , negative sampling constructs training instances (line 6). The model then computes representations following a structured three-phase process. In Phase 1: shared news encoding (lines 7–13), all involved news items (both historical and candidates) undergo frequency-domain denoising (FDD) and are encoded into initial representations . Subsequently, in Phase 2: user interest modeling (lines 14–18), the representations of historical clicked news items are exclusively routed to the popularity bias correction (PBC) module (line 16) to automatically generate fair representations for both cold and popular news without altering semantic directions. A news-level attention mechanism then aggregates these calibrated vectors into a comprehensive user interest representation (line 18). Candidate news explicitly bypass this PBC module. Finally, in Phase 3: click prediction (lines 19–23), matching scores are calculated via inner products between and candidate representations (lines 20–22). The model is optimized by updating all parameters via backpropagation using the computed batch loss (lines 25–26).
The computational requirements of FDDP-RN can be analyzed through both theoretical complexity bounds and empirical measurements. (1) The encoding of a single news article undergoes a Fourier transform. An FFT is performed along the sequence dimension on the word vector matrix to transform the semantic signal from the spatial domain to the frequency domain. The computational cost is . Scaling is performed through a frequency-domain denoising mechanism that splits high and low frequencies and scales high frequencies. This process involves only element-wise operations and no complex matrix operations. The computational complexity is . The computational cost of IFFT is the same as that of FFT, which is . (2) Popularity bias correction only performs magnitude calibration and distribution normalization on the final embedding vector of historical cold-start news articles. No complex matrix operations or feature reconstruction are required. The calculation mainly includes estimating the average embedding magnitude of popular news, computing the embedding norm of cold-start news, and applying vector scaling based on the correction factor . All operations are scalar operations and element-wise vector operations, with a computational complexity of .
The computational feasibility of FDDP-RN is validated through empirical benchmarks on an NVIDIA GeForce RTX 4070 Ti SUPER GPU. During the training phase, the model achieves a processing throughput of 1.2 s per batch. This efficiency stems from the inherent
complexity of the frequency-domain denoising module, which introduces minimal computational overhead compared to the standard spatial attention mechanisms. Additionally, the popularity bias correction operates as a parameter-free adjustment, ensuring that fairness calibration does not impede training scalability. Consequently, for a dataset with 250,000 samples, the model completes five full training epochs and reaches optimal convergence within approximately 10 h, demonstrating a superior trade-off between structural complexity and training speed. To quantify the additional cost introduced by the two proposed components,
Table 3 reports the measured batch-level overhead under the same hardware and training configuration. Compared with the base news encoder without FDD or PBC, the full FDDP-RN model increases the average training time from 1.09 s to 1.20 s per batch, corresponding to a 10.1% overhead. The FDD module accounts for most of this increase because it performs one rFFT, one lightweight frequency scaling operation, and one inverse real FFT along the title sequence dimension. The PBC module adds only 0.02 s per batch because it uses norm computation and scalar rescaling on historical news embeddings. It introduces no trainable parameters and requires only a one-time offline propagation-state clustering step.
In terms of deployment and inference, if news embeddings are pre-generated and stored in a cache, the computational overhead for encoding candidate articles during a user request is essentially negligible. For scenarios requiring on-the-fly encoding, the complexity for a single news item is governed by the self-attention mechanism at , scaling linearly to for a pool of candidates. Furthermore, the final ranking process, which relies on vector dot products, requires only per item, resulting in a total scoring complexity of .
Author Contributions
Conceptualization, B.M. and X.D.; methodology, B.M. and H.G.; software, Y.D. and X.W.; validation, B.M. and H.G.; formal analysis, B.M. and X.D.; investigation, Y.D. and X.D.; resources, B.M.; data curation, X.D. and Y.D.; writing—original draft preparation, X.D. and Y.D.; writing—review and editing, B.M. and H.G.; visualization, Y.D. and X.W.; supervision, B.M.; project administration, B.M.; funding acquisition, B.M. All authors have read and agreed to the published version of the manuscript.
Funding
This work was supported in part by the Natural Science Foundation of Fujian Province, China (Grant No. 2022J05176), the President’s Fund of Minnan Normal University (No. KJ2022002), and the Minnan Normal University Advanced Cultivation Project (No. MSGJB2023019).
Data Availability Statement
The datasets analyzed in this study are publicly available. The Microsoft News Dataset (MIND) is available at
https://msnews.github.io/ (accessed on 29 July 2026); the Adressa dataset is available from its original public release.
Conflicts of Interest
Author Huifan Gao was employed by Xiamen Airlines Co., Ltd. (Xiamen, China). The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as potential conflicts of interest. All research was conducted at Minnan Normal University under the direction of Dr. Biyang Ma.
References
- Zhu, Q.; Zhou, X.; Song, Z.; Tan, J.; Guo, L. Dan: Deep attention neural network for news recommendation. In Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence (AAAI), Honolulu, HI, USA, 27 January–1 February 2019; pp. 5973–5980. [Google Scholar] [CrossRef] [Scilit]
- Qi, T.; Wu, F.; Wu, C.; Huang, Y. News Recommendation with Candidate-aware User Modeling. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), Madrid, Spain, 11–15 July 2022; pp. 1917–1921. [Google Scholar] [CrossRef] [Scilit]
- Kim, T.; Kim, Y.; Lee, Y.C.; Shin, W.Y.; Kim, S.W. Is It Enough Just Looking at the Title?: Leveraging Body Text To Enrich Title Words Towards Accurate News Recommendation. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management (CIKM), Atlanta, GA, USA, 17–21 October 2022; pp. 4138–4142. [Google Scholar] [CrossRef] [Scilit]
- Jiang, S.; Song, H.; Lu, Y.; Zhang, Z. News Recommendation Method Based on Candidate-Aware Long- and Short-Term Preference Modeling. Appl. Sci. 2025, 15, 300. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Zhu, J.; Bi, Q.; Cai, G.; Shang, L.; Dong, Z.; Jiang, X.; Liu, Q. MINER: Multi-Interest Matching Network for News Recommendation. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2022; Association for Computational Linguistics: Dublin, Ireland, 2022; pp. 343–352. [Google Scholar] [CrossRef] [Scilit]
- Treviso, M.; Lee, J.U.; Ji, T.; Van Aken, B.; Cao, Q.; Ciosici, M.R.; Hassid, M.; Heafield, K.; Hooker, S.; Raffel, C.; et al. Efficient methods for natural language processing: A survey. Trans. Assoc. Comput. Linguist. 2023, 11, 826–860. [Google Scholar] [CrossRef] [Scilit]
- Lee-Thorp, J.; Ainslie, J.; Eckstein, I.; Ontanon, S. FNet: Mixing Tokens with Fourier Transforms. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL); Association for Computational Linguistics: Seattle, WA, USA, 2022; pp. 4296–4313. [Google Scholar] [CrossRef] [Scilit]
- Qi, T.; Wu, F.; Wu, C.; Yang, P.; Yu, Y.; Xie, X.; Huang, Y. HieRec: Hierarchical User Interest Modeling for Personalized News Recommendation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP); Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 5446–5456. [Google Scholar] [CrossRef] [Scilit]
- Ding, Y.; Wang, B.; Cui, X.; Xu, M. Popularity prediction with semantic retrieval for news recommendation. Expert Syst. Appl. 2024, 247, 123308. [Google Scholar] [CrossRef] [Scilit]
- Wu, F.; Qiao, Y.; Chen, J.H.; Wu, C.; Qi, T.; Lian, J.; Liu, D.; Xie, X.; Gao, J.; Wu, W.; et al. MIND: A Large-scale Dataset for News Recommendation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL); Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 3597–3606. [Google Scholar] [CrossRef] [Scilit]
- Wu, C.; Wu, F.; An, M.; Huang, J.; Huang, Y.; Xie, X. Neural News Recommendation with Attentive Multi-View Learning. In Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI), Macao, China, 10–16 August 2019; pp. 3863–3869. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, R.; Wang, S.; Lu, W.; Peng, X. News Recommendation Via Multi-Interest News Sequence Modelling. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore, 23–27 May 2022; pp. 7942–7946. [Google Scholar] [CrossRef] [Scilit]
- Wang, R.; Wang, S.; Lu, W.; Peng, X.; Zhang, W.; Zheng, C.; Qiao, X. Intention-Aware User Modeling for Personalized News Recommendation. In Proceedings of the International Conference on Database Systems for Advanced Applications (DASFAA); Springer: Cham, Switzerland, 2023; pp. 373–389. [Google Scholar] [CrossRef] [Scilit]
- Rao, Y.; Zhao, W.; Zhu, Z.; Lu, J.; Zhou, J. Global Filter Networks for image classification. Adv. Neural Inf. Process. Syst. (NeurIPS) 2021, 34, 980–993. [Google Scholar]
- Zheng, L.; Lu, C.T.; Jiang, F.; Zhang, J.; Yu, P. Spectral Collaborative Filtering. In Proceedings of the 12th ACM Conference on Recommender Systems (RecSys), Vancouver, BC, Canada, 2 October 2018; pp. 311–319. [Google Scholar] [CrossRef] [Scilit]
- Wei, L.; Ni, R.; Wei, J.; Jiang, Y. Frequency-sensitive diffusion model for personalized sequential recommendation. Neurocomputing 2025, 654, 131313. [Google Scholar] [CrossRef] [Scilit]
- Yang, W.; Zhong, R.; Chen, Y.; Li, S.; Ping, H.; Lu, C.; Jiang, P. FITMM: Adaptive Frequency-Aware Multimodal Recommendation via Information-Theoretic Representation Learning. In Proceedings of the 33rd ACM International Conference on Multimedia (MM), Dublin, Ireland, 27–31 October 2025. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Q.; Chen, X.; Zhang, H.; Li, X. Dynamic Hierarchical Attention Network for news recommendation. Expert Syst. Appl. 2024, 255, 124667. [Google Scholar] [CrossRef] [Scilit]
- Qi, T.; Wu, F.; Wu, C.; Huang, Y. PP-Rec: News Recommendation with Personalized User Interest and Time-aware News Popularity. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP), Online, 1–6 August 2021; pp. 5457–5467. [Google Scholar] [CrossRef] [Scilit]
- Meng, L.; Shi, C.; Hao, S.; Su, X. DCAN: Deep Co-Attention Network by Modeling User Preference and News Lifecycle for News Recommendation. In Proceedings of the International Conference on Database Systems for Advanced Applications (DASFAA); Springer: Cham, Switzerland, 2021; pp. 681–696. [Google Scholar] [CrossRef] [Scilit]
- Ma, B.; Deng, Y.; Gao, H. PAD-MPFN: Dynamic Fusion with Popularity Decay for News Recommendation. Electronics 2025, 14, 3057. [Google Scholar] [CrossRef] [Scilit]
- Yang, J. Effects of Popularity-Based News Recommendations (“Most-Viewed”) on Users’ Exposure to Online News. Media Psychol. 2016, 19, 243–271. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Chen, Y.; Wang, Z.; Zhao, W. Popularity-enhanced news recommendation with multi-view interest representation. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management (CIKM), Virtual, 1–5 November 2021; Association for Computing Machinery: New York, NY, USA, 2021; pp. 1949–1958. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Jiang, Y.; Li, H.; Zhao, W. Improving news recommendation with channel-wise dynamic representations and contrastive user modeling. In Proceedings of the 16th ACM International Conference on Web Search and Data Mining (WSDM), Singapore, 27 February–3 March 2023; pp. 562–570. [Google Scholar] [CrossRef] [Scilit]
- Pennington, J.; Socher, R.; Manning, C.D. GloVe: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP); Association for Computational Linguistics: Doha, Qatar, 2014; pp. 1532–1543. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA, 4–9 December 2017; Volume 30. [Google Scholar]
- Gulla, J.A.; Zhang, L.; Liu, P.; Özgübek, Ö.; Su, X. The Adressa Dataset for News Recommendation. In Proceedings of the International Conference on Web Intelligence (WI), Leipzig, Germany, 23–26 August 2017; pp. 1042–1048. [Google Scholar] [CrossRef] [Scilit]
- Yang, B.; Liu, D.; Suzumura, T.; Dong, R.; Li, I. Going Beyond Local: Global Graph-Enhanced Personalized News Recommendations. In Proceedings of the 17th ACM Conference on Recommender Systems (RecSys), Singapore, 18–22 September 2023; Association for Computing Machinery: New York, NY, USA, 2023; pp. 24–34. [Google Scholar] [CrossRef] [Scilit]
- Yang, Z.; Wang, W.; Qi, T.; Zhang, P.; Zhang, T.; Zhang, R.; Liu, J.; Huang, Y. GLoCIM: Global-view Long Chain Interest Modeling for news recommendation. In Proceedings of the 2024 Conference on Computational Linguistics (COLING); Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 5084–5095. [Google Scholar]
- Wang, S.; Guo, S.; Wang, L.; Liu, T.; Xu, H. HDNR: A hyperbolic-based debiased approach for personalized news recommendation. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), Taipei, Taiwan, 23–27 July 2023; pp. 259–268. [Google Scholar] [CrossRef] [Scilit]
- Qi, T.; Wu, F.; Wu, C.; Huang, Y.; Xie, X. Privacy-Preserving News Recommendation Model Learning. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2020; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 1423–1432. [Google Scholar] [CrossRef] [Scilit]
- An, M.; Wu, F.; Wu, C.; Zhang, K.; Liu, Z.; Xie, X. Neural news recommendation with long-and short-term user representations. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL); Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 336–345. [Google Scholar] [CrossRef] [Scilit]
- Pu, X.; Zhang, J.; Chen, X.; Qian, Y.; Zhang, R. News Recommendation with Word-Related Joint Topic Prediction. IEEE Access 2024, 12, 72566–72577. [Google Scholar] [CrossRef] [Scilit]
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |