TGEconomicDataset: A Collection of Russian-Language Economic Telegram Channels and a Synthetic Data Generation Framework for Continuous Authentication
Abstract
1. Summary
2. Related Work
3. Background
3.1. Description of Telegram Infrastructure
- ID (integer)—Message identifier.
- Message (string)—Message text.
- Date (datetime)—Date and time of message creation.
- URLs (array)—Contains all URLs that appear in the text of the message.
- Mentions (array)—Contains explicit references to peers in the form of “@peer_name”.
- fwd_id (long)—An identifier of the peer whose message was forwarded (copied).
- media_group_id (long)—identifier of a message group (media group). A message group is displayed as a single message with multiple media elements and text.
- Entities—Message formatting elements, including text styles and elements highlighted with a special style, such as URLs and email addresses, mentions of user names, bank card numbers, etc.
- Reactions (nested JSON-object)—contains user reactions to the message in the form of special Unicode characters such as
,
,
, etc.:
- Totals (array)—Contains summary information about reactions: how many users reacted to the message with a specific Unicode character.
- Array elements:
- Reaction (Unicode emoji character)—contains one Unicode character from the emoji special character group.
- Count (long)—Number of user reactions.
- recent_reactions (array)—Contains detailed data about reactions: how a specific peer reacted; not supported by all channels and chats.
- Array elements:
- peer_id (long)—Peer identifier.
- Reaction—Contains one Unicode character.
- Replies (array of messages)—User comments (replies) to the message. Each comment is a message represented as a JSON object with the structure shown above.
3.2. Using a Text Dataset in Continuous Authentication Tasks
3.3. Threat Model
3.4. Ethical Concerns
4. Materials and Methods
4.1. Source of Data on Collected Channels
4.2. Description of the Dataset Structure
4.3. Method for Obtaining Labeled Datasets Representing Channel Mixtures Intended for Solving the Problem of Continuous Authentication
- The latest messages from the second channel were added to the messages from the first channel, according to a predefined ratio/percentage. By default, the proportion of added messages is 10%. The choice of a 10% anomaly level is dictated by the key requirement of the continuous authentication task—minimizing detection delay. An anomaly must remain a rare event to preserve its diagnostic value as an indicator of compromise. A high proportion of anomalies transforms them into a new stable pattern, which contradicts the goal of early detection of a subtle deviation in the dataset. This threshold allows the validation of the model’s ability to quickly detect an intrusion before it becomes statistically predominant. Thus, messages added from another channel represent the proportion of “normal behavior” in terms of editorial, stylistic, and content policy, while added messages represent outliers. The two channels were selected so that the messages belonged to the same list of topics and had a similar average publication length. Given that the added publications were written by a different group of authors, they have different stylistic and syntactic characteristics and can thus be identified as outliers relative to “normal” messages. The date of the “injected” messages is taken from the original normal messages, and an additional synthesized date field is added, formed on the basis of the average frequency of publications in the normal part of the collection. In order to conduct experiments on collections of the same size, synthetic datasets of equal size were formed: 5, 10, and 50 thousand original messages.
- A specified percentage of messages (by default, 10% of the messages from the source channel) from the second channel, arranged in chronological order, were added to the messages from the first main channel. The set of messages for these dates was taken from the first channel and replaced with messages from the second channel. The publication dates were also added from the second channel. Unique message identifiers were also added from the second channel. The second method gives a the possibility to take into account the impact of global economic news on the subject of messages in economic channel publications and makes such a dataset more plausible. The number of original messages remains the same as it was initially, with the exception of the inserted messages.
4.4. Empirical Validation of Synthetic Datasets
- Cosine distances between text vector representations allow assessing semantic similarity at the level of individual messages, for which the following are used:
- Inter-class cosine distance—the distance from objects of the opposite class to the centroid.
- Inter-centroid distances—distances between the centroids of opposite classes.
- Fréchet Distance (FD) [44] is a strict measure for comparing multivariate distributions that accounts not only for the position of distribution centers (like cosine distance), but also for their shape and spread. In addition to FD, normalized FD is calculated in the range 0 to 1 for interpretability, enabling comparison of complexity across different datasets independently of absolute FD values. Normalized FD represents an inverted and scaled version of FD in the range 0 to 1, where 1 corresponds to complete closeness of distributions, and 0 to their maximum divergence.
5. Analysis of the Collected Dataset
6. Enriching Collected Data with Financial Indicators and Analyzing Them
- A 200-day moving average (200MA).
- The closing price is above or below the 200-day average (AboveMA).
- The number of consecutive days during which the quote value exceeds the moving average value. The greater the number of days, the longer the upward trend develops without significant corrective movement.
- Relative strength indicator (RSI), reflecting overbought or oversold conditions.
- When trading in the above-mentioned exchange instruments did not take place, holidays and weekends are supplemented with quotes from the last trading day.
7. Conclusions and Discussion
- A resource for stylometry in Social Media Publishing Economic News. We have collected and structured a dataset of Russian-language Telegram channels on economic topics, enriched it with metadata, and synchronized the financial indicators. Its structure allows for the study of linguistic patterns in relation to external context, while maintaining a balance between research utility and ethical requirements, as the data does not allow for author identification.
- A framework for generating realistic test data. The core methodological innovation of this work is our proposed method for generating synthetic labeled datasets designed to simulate an author-change scenario. This framework models a realistic account compromise, where messages from different authors are interwoven while preserving a unified thematic focus. The controlled anomaly level (set at ~10%) is crucial for validating models’ ability to perform early detection of authorship shifts—a fundamental requirement for continuous authentication systems.
- A tool for research reproducibility. We provide a complete set of scripts for dataset formation, enrichment, and the generation of synthetic collections. This ensures full reproducibility of the results and offers researchers a ready-made instrument for creating similar datasets for other subject areas and languages.
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| BTC | bitcoin |
| DXY | dollar index |
| CDF | cumulative distribution function |
| RSI | relative strength indicator |
References
- DemandSage. Telegram Statistics 2025: Worldwide Data. Available online: https://www.demandsage.com/telegram-statistics/ (accessed on 2 December 2025).
- Mehdipour, S.; Jannati, N.; Negarestani, M.; Amirzadeh, S.; Keshvardoost, S.; Zolala, F.; Vaezipour, A.; Hosseinnejad, M.; Fatehi, F. Health pandemic and social media: A content analysis of COVID-related posts on a Telegram channel with more than one million subscribers. Stud. Health Technol. Inform. 2021, 279, 122–129. [Google Scholar] [CrossRef] [Scilit]
- Lou, C.; Tandoc, E.C.; Hong, L.X.; Pong, X.Y.; Lye, W.X.; Sng, N.G. When motivations meet affordances: News consumption on Telegram. J. Stud. 2021, 22, 934–952. [Google Scholar] [CrossRef] [Scilit]
- Griggio, C.F.; Nouwens, M.; Klokmose, C.N. Caught in the network: The impact of WhatsApp’s 2021 privacy policy update on users’ messaging app ecosystems. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ‘22), New York, NY, USA, 30 April–5 May 2022; pp. 1–23. [Google Scholar] [CrossRef] [Scilit]
- Al-Rawi, A. News loopholing: Telegram news as portable alternative media. J. Comput. Soc. Sci. 2022, 5, 949–968. [Google Scholar] [CrossRef] [Scilit]
- Gupta, K.; Oladimeji, D.; Varol, C.; Rasheed, A.; Shahshidhar, N. A comprehensive survey on artifact recovery from social media platforms: Approaches and future research directions. Information 2023, 14, 629. [Google Scholar] [CrossRef] [Scilit]
- Telegram APIs. Available online: https://core.telegram.org/api (accessed on 2 December 2025).
- Herrero-Solana, V.; Castro-Castro, C. Telegram channels and bots: A ranking of media outlets based in Spain. Societies 2022, 12, 164. [Google Scholar] [CrossRef] [Scilit]
- Channels FAQ. Available online: https://telegram.org/faq_channels?setln=en (accessed on 2 December 2025).
- Xiao, C.; Freeman, D.M.; Hwa, T. Detecting clusters of fake accounts in online social networks. In Proceedings of the 8th ACM Workshop on Artificial Intelligence and Security, Denver, CO, USA, 16 October 2015; ACM Press: New York, NY, USA, 2015; pp. 91–101. [Google Scholar]
- La Morgia, M.; Mei, A.; Sassi, F.; Stefa, J. Pump and dumps in the Bitcoin era: Real time detection of cryptocurrency market manipulations. In Proceedings of the 2020 IEEE 29th International Conference on Computer Communications and Networks (ICCCN), Honolulu, HI, USA, 3–6 August 2020; pp. 1–9. [Google Scholar]
- La Morgia, M.; Mei, A.; Sassi, F.; Stefa, J. The Doge of Wall Street: Analysis and detection of pump and dump cryptocurrency manipulations. ACM Trans. Internet Technol. 2023, 23, 1–28. [Google Scholar] [CrossRef] [Scilit]
- Rajaei, M.J.; Mahmoud, Q.H. A survey on pump and dump detection in the cryptocurrency market using machine learning. Future Internet 2023, 15, 267. [Google Scholar] [CrossRef] [Scilit]
- Conti, E.; Salvi, D.; Borrelli, C.; Hosler, B.; Bestagini, P.; Antonacci, F.; Sarti, A.; Stamm, M.C.; Tubaro, S. Deepfake speech detection through emotion recognition: A semantic approach. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2022), Singapore, 22–27 May 2022; pp. 8962–8966. [Google Scholar]
- Rao, S.; Verma, A.K.; Bhatia, T. Review on social spam detection: Challenges, open issues, and future directions. Expert. Syst. Appl. 2021, 186, 115742. [Google Scholar] [CrossRef] [Scilit]
- Ayeswarya, S.; Singh, K.J. A comprehensive review on secure biometric-based continuous authentication and user profiling. IEEE Access 2024, 12, 82996–83021. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Zheng, R.; Chen, H. From Fingerprint to Writeprint. Commun. ACM 2006, 49, 76–82. [Google Scholar] [CrossRef] [Scilit]
- Fedotova, A.; Romanov, A.; Kurtukova, A.; Shelupanov, A. Authorship attribution of social media and literary Russian-language texts using machine learning methods and feature selection. Future Internet 2022, 14, 4. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Zhang, X.; Hu, H. Continuous user authentication on multiple smart devices. Information 2023, 14, 274. [Google Scholar] [CrossRef] [Scilit]
- Telegram Channels and Groups Catalog. Available online: https://tgstat.ru/en (accessed on 2 December 2025).
- Baumgartner, J.; Zannettou, S.; Squire, M.; Blackburn, J. The Pushshift Telegram Dataset. Proc. Int. AAAI Conf. Web Soc. Media 2020, 14, 840–847. [Google Scholar] [CrossRef] [Scilit]
- Hoseini, M.; Melo, P.; Benevenuto, F.; Feldmann, A.; Zannettou, S. On the globalization of the QAnon conspiracy theory through Telegram. arXiv 2021, arXiv:2105.13020. [Google Scholar] [CrossRef] [Scilit]
- Hashemi, A.; Zare Chahooki, M.A. Telegram group quality measurement by user behavior analysis. Soc. Netw. Anal. Min. 2019, 9, 33. [Google Scholar] [CrossRef] [Scilit]
- Sosa, J.; Sharoff, S. Multimodal pipeline for collection of misinformation data from Telegram. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, Marseille, France, 20–25 June 2022; European Language Resources Association: Paris, France, 2022; pp. 1480–1489. [Google Scholar]
- La Morgia, M.; Mei, A.; Mongardini, A.M. TGDataset: Collecting and exploring the largest Telegram channels dataset. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1 (KDD ‘25), New York, NY, USA, 3–7 August 2025; pp. 2325–2334. [Google Scholar] [CrossRef] [Scilit]
- Angermaier, M.; Pinheiro Neto, J.; Höldrich, E.; Lasser, J. Dataset for the “The Schwurbelarchiv: A German Language Telegram dataset for the study of conspiracy theories” paper [Data set]. arXiv 2025, arXiv:2504.06318. [Google Scholar] [CrossRef]
- Kireev, K.; Mykhno, Y.; Troncoso, C.; Overdorf, R. A Telegram dataset of propaganda and its moderation. Proc. Int. AAAI Conf. Web Soc. Media 2025, 19, 2510–2518. [Google Scholar] [CrossRef] [Scilit]
- Gangopadhyay, S.; Dessí, D.; Dimitrov, D.; Dietze, S. TeleScope: A longitudinal dataset for investigating online discourse and information interaction on Telegram. Proc. Int. AAAI Conf. Web Soc. Media 2025, 19, 2423–2433. [Google Scholar] [CrossRef] [Scilit]
- Groups and Channels. Available online: https://telegram.org/faq#groups-and-channels (accessed on 2 December 2025).
- MessageEntity. Available online: https://core.telegram.org/type/MessageEntity (accessed on 2 December 2025).
- Bansal, P.; Ouda, A. Continuous authentication in the digital age: An analysis of reinforcement learning and behavioral biometrics. Computers 2024, 13, 103. [Google Scholar] [CrossRef] [Scilit]
- Shopon, M.; Tumpa, S.N.; Bhatia, Y.; Kumar, K.N.P.; Gavrilova, M.L. Biometric systems de-identification: Current advancements and future directions. J. Cybersecur. Priv. 2021, 1, 470–495. [Google Scholar] [CrossRef] [Scilit]
- Estrela, P.M.A.B.; Albuquerque, R.d.O.; Amaral, D.M.; Giozza, W.F.; Júnior, R.T.d.S. A framework for continuous authentication based on touch dynamics biometrics for mobile banking applications. Sensors 2021, 21, 4212. [Google Scholar] [CrossRef] [Scilit]
- Choi, M.; Lee, S.; Jo, M.; Shin, J.S. Keystroke dynamics-based authentication using unique keypad. Sensors 2021, 21, 2242. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ashibani, Y.; Kauling, D.; Mahmoud, Q.H. Design and implementation of a contextual-based continuous authentication framework for smart homes. Appl. Syst. Innov. 2019, 2, 4. [Google Scholar] [CrossRef] [Scilit]
- Bułat, R.; Ogiela, M.R. Personalized context-aware authentication protocols in IoT. Appl. Sci. 2023, 13, 4216. [Google Scholar] [CrossRef] [Scilit]
- Baig, A.F.; Eskeland, S. Security, privacy, and usability in continuous authentication: A survey. Sensors 2021, 21, 5967. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fedotova, A.; Kurtukova, A.; Romanov, A.; Shelupanov, A. Semantic clustering and transfer learning in social media texts authorship attribution. IEEE Access 2024, 12, 39783–39803. [Google Scholar] [CrossRef] [Scilit]
- Cho, S.H.; Kim, B.; Kwon, H.; Kim, M. Exploring the potential of large language models for author profiling tasks in digital text. In Notebook for PAN at CLEF 2023; Elsevier: Amsterdam, The Netherlands, 2025; pp. 1272–1281. [Google Scholar]
- Brocardo, M.L.; Traore, I.; Woungang, I. Authorship verification of e-mail and tweet messages applied for continuous authentication. J. Comput. Syst. Sci. 2015, 81, 1429–1440. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Zhang, Q.; Huang, M. Author verification of text fragments based on the BERT model. In Notebook for PAN at CLEF 2023; Foshan University: Foshan, China, 2023; pp. 1–4. [Google Scholar]
- Telemetr—Telegram Channel Search and Analytics Platform. Available online: https://telemetr.me/ (accessed on 16 January 2026).
- BERTopic: Neural Topic Modeling with a Class-Based TF-IDF Procedure. Available online: https://arxiv.org/abs/2203.05794 (accessed on 2 December 2025).
- Dowson, D.C.; Landau, B.V. The Fréchet distance between multivariate normal distributions. J. Multivar. Anal. 1982, 12, 450–455. [Google Scholar] [CrossRef] [Scilit]
- Cointegrated/Rubert-Tiny2: Lightweight Russian BERT Model. 2021. Available online: https://huggingface.co/cointegrated/rubert-tiny2 (accessed on 16 January 2026).
- Language Detection Library in Python. Available online: https://github.com/fedelopez77/langdetect (accessed on 2 December 2025).
- Memory Inefficient Algorithm and Getting Error While Saving the Model. Available online: https://github.com/MaartenGr/BERTopic/issues/173 (accessed on 2 December 2025).
- Terragni, S.; Fersini, E. Word embedding-based topic similarity measures. In Proceedings of the International Conference on Applications of Natural Language to Information Systems (ANLIS 2021), Saarbrücken, Germany, 23–25 June 2021; Springer: Berlin/Heidelberg, Germany, 2021; pp. 33–45. [Google Scholar]
- MarketWatch. Download Data. Available online: https://www.marketwatch.com/investing/cryptocurrency/btcusd/download-data? (accessed on 2 December 2025).

















| Attribute Name | Description |
|---|---|
| _id | MongoDB document identifier |
| Title | Channel name exactly as in Telegram |
| posted | Date the document was added to the collection |
| timestamp | Date of data collection in Unix format, i.e., the number of seconds that have elapsed since 00:00:00 UTC on 1 January 1970. |
| Collection | Collection name where messages from the channel are stored |
| Source | Source (if the data is loaded from Telegram, then the tag is “tg”) |
| ch_mentions | Mentioned channels in the channel |
| forwards_from_chat | Channels from which messages are forwarded to the channel |
| forwards_from_user | Users from whom messages are forwarded to the channel |
| tg_urls | URLs mentioned in the channel messages |
| Origin | Source of information about the channel, either TGStat [41] or a mention in another channel |
| members_count | Number of subscribers (participants) |
| Completed | Bit flag (1 if the channel messages are completely downloaded) |
| last_id | Last downloaded message ID |
| updated | Bit flag (1 if all messages of the channel are processed) |
| updated_mg | Bit flag (1 if messages with the same media group are merged) |
| Attribute Name | Description |
|---|---|
| _id | MongoDB document identifier |
| text | Message text |
| timestamp | Date of publication of the message in Unix format, i.e., the number of seconds that have elapsed since 00:00:00 UTC on 1 January 1970. |
| posted_str_date | Date of publication of the message in the format “DD/MM/YYYY” |
| posted | Date of the message creation |
| market_entities | Dictionary with values of financial instruments |
| views | Message view count in a Telegram channel |
| msg_id | Telegram message ID |
| forwards | The number of message forwards/reposts in a Telegram channel |
| media | Type of media |
| reactions | Dictionary of reactions to the message |
| edit_date | Date of the message editing |
| hashtags | Array of hashtags |
| entities | Array of message entities |
| mentions | Array of mentions |
| tg_urls | Array of Telegram URLs (https://t.me/…) in the message |
| symbols | Number of symbols in the message |
| words | Number of words in the message |
| Attribute Name | Description |
|---|---|
| _id | Record identifier |
| avg_word_inj | Average number of words in a channel’s message from which outlier posts are added |
| avg_word_main | Average number of words in posts on the original channel with non-outlier posts. |
| collection | Name of the resulting synthesized collection |
| collection_inj | Name of the collection containing raw data from the Telegram channel with outlier messages |
| collection_main | Name of the collection containing raw data from the main Telegram channel with non-outlier posts |
| injected_count | Number of added outlier messages |
| count_main | Number of messages marked as original non-outlier |
| Percent | Percentage of outliers |
| size | The number of messages; if 0, the channel contains all messages. Otherwise, the number of latest messages that were taken from the source channel. |
| type | collection synthesis method: type 1 corresponds to method 1; type 2 to method 2, described above. |
| Attribute Name | Description |
|---|---|
| _id | Publication ID |
| text | Message text |
| timestamp | Timestamp of the original dataset |
| newtimestamp | For added outlier message, the timestamp is set based on the average frequency of publications in the main channel. |
| outlier | Outlier tag: 0, normal message; 1, outlier. |
| Channel | Message Text | Outlier Tag |
|---|---|---|
| Name of the synthesized collection: PerfecProfit_and_whalebotalerts_10_size_5000_type1 | ||
| PerfecProfit | HI PRICE![]() Tiker—MAGN Price—55.705 06.05.2024—10:03:00 #MAGN” | 0 |
HI PRICE![]() Tiker—YNDX Price—4282.6 06.05.2024—10:03:00 #YNDX” | 0 | |
| whalebotalerts | 117 BTC ($11,321,512) transferred from Kraken to Unknown Sender’s balance: 0 | 1 |
| Crypto Fear/Greed Index Status: Extreme Greed Value: 84/100 Next Update in 19 h | 1 | |
| Metric | Mean ± SD | Range [Min, Max] | Interpretation |
|---|---|---|---|
| Cosine distance between centroids | 0.232 ± 0.076 | [0.069, 0.337] | High semantic similarity between classes |
| Fréchet Distance | 0.607 ± 0.230 | [0.205, 0.864] | Strong distribution overlap (all FD < 1.0) |
| Normalized closeness | 0.940 ± 0.019 | [0.915, 0.979] | Extremely challenging verification scenario |
| Distance from class 0 to centroid 1 | 0.359 ± 0.080 | [0.203, 0.476] | Inter-class proximity (class 0 → 1) |
| Distance from class 1 to centroid 0 | 0.397 ± 0.089 | [0.225, 0.507] | Inter-class proximity (class 1 → 0) |
| Metric | Mean ± SD | Range [Min, Max] | Interpretation |
|---|---|---|---|
| Cosine distance between centroids | 0.135 ± 0.075 | [0.018, 0.279] | High semantic similarity between classes |
| Fréchet Distance (FD) | 0.414 ± 0.218 | [0.094, 0.813] | Substantial distribution overlap |
| Normalized closeness | 0.959 ± 0.022 | [0.919, 0.991] | Challenging verification scenario |
| Distance from class 0 to centroid 1 | 0.359 ± 0.080 | [0.220, 0.490] | Interpretation |
| Distance from class 1 to centroid 0 | 0.322 ± 0.075 | [0.211, 0.491] | High semantic similarity between classes |
| Rank | Hashtag | Frequency of Use in the Dataset | Rank | Hashtag | Frequency of Use in the Dataset |
|---|---|---|---|---|---|
| 1 | #poccия’ | 32,205 | 26 | #Пoлитикa’ | 3781 |
| 2 | #Aкции’ | 22,349 | 27 | #пpoгнoз’ | 3742 |
| 3 | #Импyльc’ | 21,585 | 28 | #инφляция’ | 3716 |
| 4 | #кpиптo’ | 20,152 | 29 | #гaз’ | 3641 |
| 5 | #cшa’ | 13,746 | 30 | #YNDX’ | 3490 |
| 6 | #oтчeтнocть’ | 11,059 | 31 | #ROSN’ | 3452 |
| 7 | #aкции’ | 9588 | 32 | #cot’ | 3438 |
| 8 | #нeφть’ | 8560 | 33 | #бaнки’ | 3359 |
| 9 | #GAZP’ | 7322 | 34 | #LKOH’ | 3307 |
| 10 | #pынки’ | 7084 | 35 | #SMLT’ | 3082 |
| 11 | #BTC’ | 6454 | 36 | #sentiment’ | 2919 |
| 12 | #SBER’ | 6024 | 37 | #дивидeнды’ | 2901 |
| 13 | #fx’ | 5805 | 38 | #OZON’ | 2831 |
| 14 | #китaй’ | 5708 | 39 | #POSI’ | 2829 |
| 15 | #дкп’ | 5659 | 40 | #VKCO’ | 2808 |
| 16 | #экoнoмикa’ | 5515 | 41 | #TCSG’ | 2785 |
| 17 | #гeoпoлитикa’ | 5488 | 42 | #дивидeнд’ | 2736 |
| 18 | #eвpoпa’ | 4854 | 43 | #epзнoвocти’ | 2719 |
| 19 | #мaкpo’ | 4591 | 44 | #нoвocти’ | 2693 |
| 20 | #MOEX’ | 4416 | 45 | #GMKN’ | 2632 |
| 21 | #Ъyзнaл’ | 4211 | 46 | #AFLT’ | 2524 |
| 22 | #caнкции’ | 4100 | 47 | #ALRS’ | 2501 |
| 23 | #VTBR’ | 4024 | 48 | #CHMF’ | 2477 |
| 24 | #Maкpo’ | 3969 | 49 | #Пpeмapкeтинг’ | 2477 |
| 25 | #NVTK’ | 3832 | 50 | #Пoлитикa’ | 3781 |
| № | Mention | Number of Mentions |
|---|---|---|
| 1 | @ifax_go’ | 2655 |
| 2 | @kriptokondrashov’ | 1438 |
| 3 | @selfinvestor’ | 1204 |
| 4 | @investinfopro_mmvb’ | 1137 |
| 5 | @bitkogan’/Bitkogan’ | 953 |
| 6 | @banksta’ | 916 |
| 7 | @AK47pfl’ | 592 |
| 8 | @offcrypt’ | 544 |
| 9 | @wisedruidd’ | 485 |
| 10 | @keepcontact’ | 457 |
| 11 | @voin_dv’ | 452 |
| 12 | @rocket1signals’ | 450 |
| 13 | @interfaxonline’ | 432 |
| 14 | @ntvnews’ | 418 |
| 15 | @fanimani_official’ | 361 |
| 16 | @investi_usa’ | 330 |
| 17 | @rybar’ | 290 |
| 18 | @alfabank’ | 280 |
| 19 | @Dohod’ | 271 |
| 20 | @FiboAlex’ | 264 |
| № | Mention | Number of Mentions |
|---|---|---|
| 1 | @WhaleBotAlerts’ | 148,634 |
| 2 | @banksta’ | 26,161 |
| 3 | @bankrollo’ | 23,360 |
| 4 | @if_market_news’ | 21,955 |
| 5 | @interfaxonline’ | 21,531 |
| 6 | @bankoffo’ | 19,148 |
| 7 | @kommersant’ | 18,549 |
| 8 | @banki_oil’ | 14,939 |
| 9 | @marketsnapshot’ | 14,221 |
| 10 | @ifax_go’ | 13,996 |
| 11 | @voenkorKotenok’ | 13,516 |
| 12 | @PerfecProfit’ | 11,791 |
| 13 | @fatcat18’ | 11,213 |
| 14 | @kstati_p’ | 11,005 |
| 15 | @kfm936’ | 10,473 |
| 16 | @ejdailyru’ | 8586 |
| 17 | @selfinvestor’ | 8461 |
| 18 | @bitkogan’ | 7625 |
| 19 | @ivan_utenkov13’ | 6473 |
| 20 | @milinfolive’ | 6347 |
| № | Original Unigram | Unigram in English | Number of Occurrences |
|---|---|---|---|
| 1 | гoд | year | 1,058,316 |
| 2 | Poccия | Russia | 643,307 |
| 3 | мoчь | be able | 433,262 |
| 4 | кoмпaния | company | 395,978 |
| 5 | pынoк | market | 379,351 |
| 6 | poccийcкий | Russian | 361,870 |
| 7 | cвoй | its | 359,526 |
| 8 | aкция | stock | 330,631 |
| 9 | цeнa | price | 325,888 |
| 10 | нoвый | new | 321,402 |
| 11 | pyбль | ruble | 310,296 |
| 12 | CШA | USA | 309,627 |
| 13 | PΦ | RF | 307,033 |
| 14 | дeнь | day | 263,045 |
| 15 | cтpaнa | country | 261,620 |
| 16 | млpд | billion | 260,943 |
| 17 | Укpaинa | Ukraine | 247,127 |
| 18 | вpeмя | time | 243,380 |
| 19 | чeлoвeк | person | 228,869 |
| 20 | дaть | give/provide | 226,986 |
| № | Original Bigram | Bigram in English | Number of Occurrences |
|---|---|---|---|
| 1 | 2023 гoд | year 2023 | 130,353 |
| 2 | 2022 гoд | year 2022 | 92,893 |
| 3 | 2024 гoд | year 2024 | 85,779 |
| 4 | млpд pyб | billion RUB | 66,443 |
| 5 | млpд pyбль | billion rubles | 61,484 |
| 6 | Bлaдимиp Пyтин | Vladimir Putin | 40,298 |
| 7 | чиcтый пpибыль | net profit | 37,892 |
| 8 | бaнк Poccия | Bank of Russia | 33,731 |
| 9 | binance tx | Binance transaction | 32,004 |
| 10 | ключeвoй cтaвкa | key rate | 30,119 |
| 11 | sender’s balance | sender’s balance | 29,858 |
| 12 | balance tx | balance transaction | 29,858 |
| 13 | PИA нoвocть | RIA Novosti/RIAN | 28,315 |
| 14 | индeкc мocбиpжa | MOEX Index/Moscow Exchange Index | 27,663 |
| 15 | Coinbase tx | Coinbase transaction | 26,726 |
| 16 | млн pyбль | million rubles | 24,845 |
| 17 | BC PΦ | Armed Forces of the Russian Federation (AF RF) | 24,626 |
| 18 | yгoлoвный дeлo | criminal case | 23,973 |
| 19 | coвeт диpeктop | board of directors | 22,764 |
| 20 | oбъëм измeнeниe | volume change | 21,666 |
| 21 | измeнeниe дeнь | daily change | 21,639 |
| 22 | цeнa oткpытиe | open price | 21,593 |
| 23 | oткpытиe cвeчa | opening of the candle | 21,585 |
| Percent | Topic | Keywords in English | Original Keywords |
|---|---|---|---|
| 26 | Russia’s economy, market, banking sector, investments, finance, and events in Russia and around the world | ‘billion’, ‘rubles’, ‘Russia’, ‘USA’, ‘2023’, ‘rub’, ‘million’, ‘companies’, ‘RF’, ‘oil’, ‘CB’, ‘2024’, ‘2022’, ‘market’ | ‘млpд’, ‘pyблeй’, ‘Poccия’, ‘cшa’, ‘2023’, ‘pyб’, ‘млн’, ‘кoмпaнии’, ‘PΦ’, ‘нeφть’, ‘ЦБ’, ‘2024’, ‘2022’, ‘pынoк’ |
| 18 | Russia’s foreign policy and economy. Finance. | ‘Russia’, ‘RF’, ‘said’, ‘USA’, ‘president’, ‘Putin’, ‘Trump’, ‘Ukraine’, ‘ukrainians’, ‘head’, ‘reported’ | ‘Poccия’, ‘PΦ’, ‘зaявил’, CШA’, ‘пpeзидeнтa’, ‘пyтин’, ‘тpaмпa’, ‘Укpaинa’, ‘Tpaмп’, ‘пpeзидeнт’’глaвa’, ‘cooбщил’ |
| 17 | Situation in Russian regions: UAV strikes, air defense operations, statements by authorities | region, UAF, reported, RF, Russia, Defense Ministry, victims, direction, data, result, governor, air defense | ‘oблacти’, ‘BCУ’, ‘cooбщил’, ‘cooбщили’, ‘PΦ’, ‘paйoнe’, ‘Poccия’, ‘минoбopoны’, ‘пocтpaдaвшиx’, ‘нaпpaвлeнии’, ‘дaнным’, ‘peзyльтaтe’, ‘гyбepнaтop’, ‘ПBO’ |
| 16 | Cryptocurrency. Investments | ‘unknown’, ‘tx’, ‘transfered’, ‘transfered unknown’, ‘usdc’, ‘btc’, ‘000’, ‘coinbase’, ‘binance’, ‘sender’, ‘sender balance’, ‘balance tx’, ‘balance’, ‘unknown coinbase’ | ‘unknown’, ‘tx’, ‘transfered’, ‘transfered unknown’, ‘usdc’, ‘btc’, ‘000’, ‘coinbase’, ‘binance’, ‘sender’, ‘sender balance’, ‘balance tx’, ‘balance’, ‘unknown coinbase’ |
| 8 | News. Investment news. Bonds. | More details, bonds, privet Rostov, privet, Rostov, news, offer, rubles, billion, investments, shares, currency, exchange rate | ‘пoдpoбнee’, ‘oблигaций’, ‘privet rostov’, ‘privet’, ‘rostov’, ‘нoвocть’, ‘пpeдлoжить’, ‘pyб’, ‘млpд’, ‘news’, ‘инвecтиции’, ‘aкции’, ‘вaлютa’, ‘кypc’ |
| 4 | Events in Israel | Israel, Hamas, Gaza, USA, Gaza Strip, Iran, Strip, Netanyahu, IDF, stated | ‘Изpaиль’, ‘XAMAC’, ‘Гaзa’, ‘CШA’, ‘ceктope гaзa’, ‘Иpaн’, ‘ceктope’, ‘Heтaньяxy’, ‘ЦAXAЛ’, ‘зaявил’, ‘ceктop Гaзa’ |
| 2 | Telegram-censored publications | infringement, due copyright, message couldn, displayed device, device due, couldn displayed, copyright infringement, copyright, displayed, device | ‘infringement’, ‘due copyright’, ‘message couldn’, ‘displayed device’, ‘device due’, ‘couldn displayed’, ‘copyright infringement’, ‘copyright’, ‘displayed’, ‘device’ |
| 2 | Weather events in Russia. | degrees, temperature, Moscow, expected, snow, wind, heat, abnormal precipitation, ISS, weather, forecasters, at night, rain | ‘гpaдycoв’, ‘тeмпepaтypa’, ‘Mocквa’, ‘oжидaeтcя’, ‘cнeг’, ‘вeтep’, ‘жapa’, ‘aнoмaльный’ ‘ocaдкoв’, ‘мкc’, ‘пoгoдa’, ‘cинoптики’, ‘нoчью’, ‘дoждь’, ‘oжидaютcя’ |
| 2 | China’s economy | China, stocks, USA, alibaba, Xi, People Republic of Chine, dollars, chinese, Hong Kong, chinastocks | ‘Kитaй’, ‘China’, ‘aкции’, ‘CШA’, ‘alibaba’, ‘Cи’, ‘KHP’, ‘дoллapoв’, ‘китaйцы’, ‘китaйcкиx’, ‘гoнкoнгcкиx’, ‘китaйcкиe’, ‘chinastocks’ |
| 1 | Economy of Georgia, South America, Africa, European Union | Georgia, France, Maduro, Tbilisi, Venezuela, protesting, president, protesters, Niger, protests, countries, protest, Argentine, police | ‘Гpyзия’, ‘Φpaнция’, ‘Maдypo’, ‘Tбилиcи’, ‘Beнecyэлa’, ‘пpeзидeнтa’, ‘пpoтecтyющиe’, ‘нигep’, ‘пpoтecты’, ‘cтpaны’, ‘Apгeнтинa’, ‘пoлиция’ |
| 1 | Artificial intelligence, apps, social media, and investments | AI, Apple, Telegram, iPhone, Google, company, Microsoft, VK, Twitter, Russia, users, intelligence, artificial intelligence, artificial | ‘ИИ’, ‘Apple’, ‘Telegram’, ‘iPhone’, ‘Google’, ‘кoмпaнии’, ‘Microsoft’, ‘VK’, ‘Twitter’, ‘Poccия’, ‘пoльзoвaтeлeй’, ‘интeллeктa’, ‘иcкyccтвeннoгo интeллeктa’, ‘иcкyccтвeннoгo’ |
| Exchange-Traded Instrument | Channel | Correlation |
|---|---|---|
| Gold (futures contract) | @CalendarBG | 0.30 |
| USDRUB | @bogdanoffinvest | 0.311 |
| @scalpon | 0.302 | |
| @marketsnapshot | 0.301 | |
| Bitcoin | @whalebotalerts | 0.567 |
| @futures_Utushkin | 0.536 | |
| @PKNCash | 0.339 | |
| @crypto_nwz | 0.335 | |
| @scalpon | 0.333 | |
| @usertrader3 | 0.315 | |
| BRENT crude oil (futures contract) | @tass_agency | 0.382 |
| @cbrstocks | 0.376 | |
| @markettwits | 0.362 | |
| @if_market_news | 0.362 | |
| @trade_system007 | 0.322 | |
| @marketsnapshot | 0.321 | |
| @macroresearch | 0.316 | |
| @prime1 | 0.313 | |
| @oil_capital | 0.307 | |
| Dollar index | @marketsnapshot | 0.37 |
| @finchehov | 0.331 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Luneva, E.; Banokin, P.; Shelupanov, A. TGEconomicDataset: A Collection of Russian-Language Economic Telegram Channels and a Synthetic Data Generation Framework for Continuous Authentication. Data 2026, 11, 25. https://doi.org/10.3390/data11020025
Luneva E, Banokin P, Shelupanov A. TGEconomicDataset: A Collection of Russian-Language Economic Telegram Channels and a Synthetic Data Generation Framework for Continuous Authentication. Data. 2026; 11(2):25. https://doi.org/10.3390/data11020025
Chicago/Turabian StyleLuneva, Elena, Pavel Banokin, and Alexander Shelupanov. 2026. "TGEconomicDataset: A Collection of Russian-Language Economic Telegram Channels and a Synthetic Data Generation Framework for Continuous Authentication" Data 11, no. 2: 25. https://doi.org/10.3390/data11020025
APA StyleLuneva, E., Banokin, P., & Shelupanov, A. (2026). TGEconomicDataset: A Collection of Russian-Language Economic Telegram Channels and a Synthetic Data Generation Framework for Continuous Authentication. Data, 11(2), 25. https://doi.org/10.3390/data11020025






