Next Article in Journal
100 m Resolution Age-Stratified Population Grid Data for China Based on Township-Level in 2020
Previous Article in Journal
Face Typicality–Distinctiveness Norms for the 304 Front-View Faces of the Glasgow Unfamiliar Face Database
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Data Descriptor

TGEconomicDataset: A Collection of Russian-Language Economic Telegram Channels and a Synthetic Data Generation Framework for Continuous Authentication

by
Elena Luneva
*,
Pavel Banokin
and
Alexander Shelupanov
Department of Comprehensive Information Security of Electronic Computer Systems, Tomsk State University of Control Systems and Radioelectronics, 634050 Tomsk, Russia
*
Author to whom correspondence should be addressed.
Data 2026, 11(2), 25; https://doi.org/10.3390/data11020025
Submission received: 7 December 2025 / Revised: 16 January 2026 / Accepted: 25 January 2026 / Published: 28 January 2026
(This article belongs to the Section Information Systems and Data Management)

Abstract

Telegram, along with WhatsApp and Signal, has become very popular due to its hybrid capabilities, including both instant private and public messaging, making it an effective tool for quickly broadcasting content to a wide audience. This article presents TGEconomicDataset, a new dataset containing more than 2.9 million messages from the most popular Russian-language Telegram channels in the field of economics, as well as synthetically generated labeled mixtures of these channels. These mixtures are specifically designed to model authorship change scenarios for testing various methods for solving the problem of continuous authentication, which is of particular interest due to the need for organizations and companies to rely on data posted on social media. The presented dataset is enriched with quotes of important financial instruments such as gold futures, the USD/RUB currency pair, BRENT oil, the dollar index (DXY), and bitcoin (BTC), synchronized with the message timestamps. A detailed joint analysis of the collected data is provided. In addition to the presented dataset, we publish the scripts used to collect the data, integrate the financial indicators, and generate the synthetic mixtures for the continuous authentication task, ensuring full reproducibility of the research.

1. Summary

Today, Telegram, along with WhatsApp and Signal, has become very popular due to its hybrid capabilities, which include both instant private and public messaging. The number of users of this messenger app began to grow during the coronavirus pandemic [1,2,3]. In addition, starting in 2021, there has been a mass migration of users from WhatsApp to other messengers, including Telegram, due to changes in WhatsApp’s privacy policy [4,5,6].
Telegram has proven to be an attractive service due to its support for user privacy, its rich infrastructure that enables cryptocurrency transactions, and its tools for automating channel management, including artificial intelligence bots and an open platform for developers via an application programming interface (API) [7,8]. Along with the ability to privately exchange text, video, and audio messages and files in real time, Telegram has made it possible to interact with a wide audience both publicly and privately through its channel mechanism. A channel is a platform tool that allows messages and other content (video, audio, files) to be broadcast to a large audience [9]. Telegram channel data can be used by organizations in their business processes. Such organizations include mass media, online survey organizations, and financial services.
Telegram’s wide range of features, primarily privacy, also attract malicious actors to the platform for illegal activities such as distributing content (videos, video games, etc.) in violation of copyright laws, spreading radical ideologies, using the messenger to carry out terrorist attacks and using channels for market manipulation, including manipulation on crypto exchanges through pump and dump schemes [10,11,12,13].
The task of identifying typical user behavior becomes relevant here, the solution of which allows the user to be assigned to a particular class based on semantic analysis of the information received from a user. The practical application of methods and software algorithms for solving such a task may be necessary when analyzing data from both messengers and social networks, for example, to determine the fact of theft, sale of accounts, as well as the use of deepfake technologies to commit fraudulent actions in virtual space [14,15]. The methods can also be used to determine the impact of unreliable and knowingly false information on user behavior.
In this paper, we present a dataset of Telegram channels in the field of economics, as well as a mixture of these channels with class labels, formed to test various methods for solving the problem of continuous authentication [16,17,18,19], which is of particular interest due to the need for organizations to rely on data posted on social media. The main contributions of this work are as follows:
A dataset of messages from Russian-language Telegram channels has been created. The channels for downloading were selected using the TGStat [20] analytics system catalog.
A key feature of the proposed dataset is synthetic collections of labeled messages that reflect a scenario where the author changes while the thematic focus of the texts remains the same. These collections can be used to test and validate models that solve the problem of continuous authentication (authorship verification) within a given subject area.
Attributes with values reflecting the state of financial markets at the time of message creation have been added to the dataset messages. These attributes allow us to explore two interrelated aspects: first, the influence of the market context on the content of messages, and second, the suitability of using these quantitative attributes as features to improve the accuracy of continuous authentication models.
The dataset contains public channels whose data does not allow for establishing a link between a message and the user ID of its author. Thus, the dataset strikes a balance between research utility and ethical responsibility, allowing authorship and message style to be studied without revealing private information.
The dataset and scripts for its creation and analysis are publicly available. This allows the results to be reproduced and further research to be conducted in the field of continuous authentication and text analysis in social media.

2. Related Work

The largest dataset from the Telegram messenger is described in [21] and contains 27,801 channels and 317,224,715 messages on various topics, focusing on channels discussing right-wing extremism and cryptocurrency. Other noteworthy datasets are those described in [22], dedicated to the QAnon conspiracy theory and containing 151 channels, as well as those in [23], which contains more than 500,000 Persian channels.
Ref. [24] is devoted to collecting misinformation about COVID-19 in the Telegram messenger. The collected dataset contains almost a million messages collected from 2000 channels and, in addition to text messages, contains images, videos, and documents, mainly in PDF format.
The authors of [25] attempted to provide a more comprehensive picture of the Telegram ecosystem and collected a dataset of 120,979 channels, representing more than 400 million messages, which were then grouped into 13 topics.
Another study [26] was focused on collecting channels in German and represents 6000 groups and channels, about 40 million text messages, and more than 3 million transcribed audio files. The authors suggest that researchers could use this resource to study social dynamics in the areas of disinformation, political extremism, public opinion adaptation, and social network structures on Telegram [26].
Ref. [27] presents a dataset consisting of posts from 13 channels collected in real time, as well as posts collected using an export API that allows all recent messages in a channel to be downloaded. This made it possible to identify posts that had been moderated and deleted. The authors then attempted to estimate the proportion of propaganda posts in the collected data.
The paper [28] presents the largest dataset, TeleScope, containing data from 500,000 public Telegram channels and 120 million messages collected over a period of 2.5 years (until October 2024). The authors suggest that researchers use the collected dataset, including in the context of repeating social media research previously applied to X (formerly Twitter), for analyzing networks and communities, and conducting research on language communities [28].
Previous studies involving Telegram datasets have primarily focused on general content or socio-political themes. Russian-language content, including economic discourse, is represented only fragmentarily and is not treated as a coherent, annotated corpus for specialized research. Moreover, none of the known datasets provide synthetic data specifically constructed for continuous authentication tasks. Our dataset addresses these gaps. It offers a purposefully collected corpus of messages from popular Russian-language economic channels, where each message is enriched with financial attributes (gold and BRENT oil futures prices, USD/RUB exchange rates, bitcoin, and the dollar index) reflecting market conditions at the time of publication. In addition to the original authentic messages, the dataset includes synthetic messages specifically designed for rigorous evaluation of continuous authentication algorithms.

3. Background

3.1. Description of Telegram Infrastructure

Telegram combines the functionality of a social network and a messaging app for making audio and video calls and conferences. The Telegram social network allows you to interact with a large audience through chats: channels and groups. Channels and groups are collections of messages and can be either publicly available or require prior approval of each new member, as well as the ability to join via an invitation link. A channel is a chronologically ordered collection of messages with the option to subscribe to new messages, managed by a single user-administrator or a group of users with administrator or author access rights [29]. Only privileged users can post messages in a channel. The channel administrator can allow comments to be added to messages. Other subscribed users can comment on messages or leave their reactions. When a message is posted in a channel, the channel ID is indicated as the author.
A group is a chronologically ordered collection of messages from participating users with the option to subscribe to new messages, managed by one administrator user or several users with administrator and moderator access rights. Unlike a channel, all group members can create messages and react to them. Comments, as a separate nested collection of messages, are not available in a group. The class diagram (Figure 1) shows the structure of the above-described elements of the Telegram messenger. The activity of non-privileged users may be limited by the number of messages a user can send in an hour. The ability to comment on messages is not provided in the group.
The Telegram data model includes messages, channels, and users subscribed to them. As shown in Figure 1, each message consists of content, reactions, and view statistics. Message content can include text, audio and video data. A separate type of message is a telegram circle video, which does not include other types of data. Messages that include multiple media elements (images, audio recordings, videos, etc.) are represented as a group of simple messages (media_group) with a common media_group_id identifier.
A reaction to a message is a way of evaluating a message in which a user expresses their opinion in the form of one or more Unicode characters from the emoji group. The overall evaluation of a message is represented by an associative array in which the key is a Unicode emoji character and the value is the number of reactions (evaluations) left by users with the specified key.
The Telegram social network API allows developers to download chat information, including its type (channel or group), list of participants, and collection of messages. The content of the downloaded messages includes not only text data, but also design elements: sections of text highlighted with font styles, URL and email addresses, graphic elements, user IDs, hashtags, and other elements. Each message from a channel or chat is characterized by the following attributes [30]:
  • ID (integer)—Message identifier.
  • Message (string)—Message text.
  • Date (datetime)—Date and time of message creation.
  • URLs (array)—Contains all URLs that appear in the text of the message.
  • Mentions (array)—Contains explicit references to peers in the form of “@peer_name”.
  • fwd_id (long)—An identifier of the peer whose message was forwarded (copied).
  • media_group_id (long)—identifier of a message group (media group). A message group is displayed as a single message with multiple media elements and text.
  • Entities—Message formatting elements, including text styles and elements highlighted with a special style, such as URLs and email addresses, mentions of user names, bank card numbers, etc.
  • Reactions (nested JSON-object)—contains user reactions to the message in the form of special Unicode characters such as Data 11 00025 i001, Data 11 00025 i002, Data 11 00025 i003, etc.:
    • Totals (array)—Contains summary information about reactions: how many users reacted to the message with a specific Unicode character.
  • Array elements:
    • Reaction (Unicode emoji character)—contains one Unicode character from the emoji special character group.
    • Count (long)—Number of user reactions.
  • recent_reactions (array)—Contains detailed data about reactions: how a specific peer reacted; not supported by all channels and chats.
  • Array elements:
    • peer_id (long)—Peer identifier.
    • Reaction—Contains one Unicode character.
  • Replies (array of messages)—User comments (replies) to the message. Each comment is a message represented as a JSON object with the structure shown above.
The Telegram network API imposes restrictions on continuous data downloads. Frequent calls to API methods or the use of identifiers for download of non-existent (deleted) messages, users, and channels will result in blocking. The duration of each subsequent block increases and may exceed 20 h. Our experience with API usage has shown that continuous message downloading in batches of 100 items is possible with a pause of at least two seconds between calls.

3.2. Using a Text Dataset in Continuous Authentication Tasks

Continuous authentication is a security process that operates throughout a user session and uses biometric [31,32], behavioral [33,34], or contextual [35,36] (geolocation, network parameters, temporal context) data to determine user authenticity.
Some categories of user data may be unavailable, which limits the possibility of continuous authentication. In this regard, it is interesting to use text data, such as posts in messengers and social networks, as a basis for continuous authentication. Preparing datasets for experimental analysis and developing continuous authentication methods that process text data can allow the use of source text data as an independent source of information available to third-party services without the need to access closed or confidential data. This is especially relevant in the context of implementing business processes for organizations that use user message analysis, for example, for marketing, sociological, and research purposes, where publication texts serve as an independent and valuable resource [16,19].
Machine learning and text preprocessing methods used to establish authorship [37,38,39,40,41] can also be effectively used in continuous authentication tasks. Analysis of stylistic features and linguistic patterns, which are identified using algorithms, allows the creation of a unique user profile based on their written speech. Thus, methods originally aimed at determining text authorship expand the capabilities of continuous authentication, providing more reliable and adaptive protection. It should be noted that group profiles on social networks are often maintained by several authors and contain reposts from news channels, but they often have common requirements for the style of publications and the interpretation of events, which is of particular research interest when solving such problems.
There is a critical shortage of realistic, yet controlled and ethically sound test environments for developing prolonged authentication methods. We fill this gap by offering a formalized method for generating synthetic datasets. This method allows us to create marked sequences based on real public data (Telegram channels) that realistically model a compromise scenario (change in authorship while maintaining the topic), ensuring control over the level of anomalies, and also allow for rigorous benchmarking of different algorithms. Thus, we provide not just another dataset, but a toolkit for creating reproducible benchmarks that allow for the exploration of prolonged authentication methods in realistic but ethically safe conditions, without the need to access private data.

3.3. Threat Model

The dataset is intended for research in the field of continuous authentication of Telegram channels. It simulates three types of threats: (1) author change (another person starts posting messages at a certain point in time, and their style differs from what was previously observed); (2) channel compromise (a small number of messages are added as a result of unauthorized access); and (3) style imitation, hidden advertising (messages from another source are edited to match the style and vocabulary of the original author). Synthetic posts labeled 1 reproduce these threats to evaluate algorithms for detecting changes in authorship, including minor changes in style.

3.4. Ethical Concerns

The dataset consists of messages collected from public Telegram channels. The downloading was not targeted at specific channels or users; the list of channels was selected from TGStat [20] and supplemented with links mentioned within these channels. Although the data is public and not anonymized, no attempts were made to profile or identify individual users. There are two tools in the Telegram messenger: groups and channels. The dataset consists only of channel messages. The channel does not contain explicit indication of the author ID of a message. Thus, the channel does not allow for establishing a correspondence between a message and a user with author privileges.
Synthetic messages added for continuous authentication tasks simulate changes in authorship or style and do not correspond to real users, which allows for safe experimentation. Overall, the dataset strikes a balance between research utility and ethical responsibility, allowing for the study of authorship, style, and message dynamics without revealing private information.

4. Materials and Methods

4.1. Source of Data on Collected Channels

Messages for the time period from 1 September 2022 to 1 December 2024 have been downloaded. The upload process consisted of two stages: initial and additional. The lists of channels for the initial upload were obtained from TGStat [20], Russia’s largest catalog of Telegram channels and chats, under the heading “Economy”. In the alternative analytics service Telemetr [42], the economy category is represented by fewer elements compared to TGStat. During the initial downloading of messages, an additional list of channels was compiled, mentioned in messages in the form of a username, a link to a user, or in the form of a forwarded (quoted) message. Then, messages from channels in the additional list were downloaded. Thus, the most well-known economic channels in the Russian-language segment of the Telegram messenger were covered.
The constructed dataset has several inherent limitations that define its scope. First, it carries a platform selection bias, as it includes only public channels indexed by TGStat [20], potentially omitting private, new, or unlisted sources. Second, it is thematically and linguistically narrow, focused solely on Russian-language economic discourse, which limits its generalizability to other topics or languages. Third, the data reflects a specific temporal snapshot and may not capture long-term linguistic shifts. Therefore, this dataset is best utilized as a purpose-built benchmark for methodological research in stylometry and continuous authentication rather than as a representative sample for broad socio-economic analysis.
To export messages, the generated channel list is saved in a document database and used until the export process is complete. During the export process (Figure 2), an active channel is set, from which messages are downloaded in small batches to prevent blocking. Each channel metadata includes a status and a list of users and other channels mentioned in it. Text data is processed before saving. When each message is saved, log entries are updated, allowing the download to resume after a stop. When the channel is fully downloaded, the next channel in the list is set as active.
After messages are downloaded, they are processed and enriched with economic data (Figure 3). Messages with the same mediagroup_id identifier are combined, as they are displayed as a single message in the Telegram network client. The exchange’s trading day calendar is taken into account when adding economic data. Messages created on weekends have the price and exchange indicators values from the most recent previous trading day.

4.2. Description of the Dataset Structure

TGEconomicDataset dataset was stored in a MongoDB database, which enables fast query execution and the selection of necessary data samples. A document data model was used to create the dataset, which made it possible to preserve the structure of Telegram messages with nested elements and arrays of attribute values. The document model allowed the message schema to be expanded as the experiments progressed. The list of downloaded channels is presented in the “chats” collection. A description of the MongoDB collections, which contain the synthesized labeled collections for research in the field of continuous authentication, is provided in Section 4.3.
Table 1 shows the attributes of the “chats” collection. Channel messages are stored in separate collections. Each message in a channel is described by a set of attributes, which are listed in Table 2. The stock market indicators used to enrich each message are presented in the “market_entities” attribute. A detailed description of these indicators is provided in Section 6 of this study.

4.3. Method for Obtaining Labeled Datasets Representing Channel Mixtures Intended for Solving the Problem of Continuous Authentication

The task of typical user behavior requires evaluation of the sensitivity of the methods being developed on different user groups in different subject areas. Also, it is necessary to assess the possibility of improving the effectiveness of these methods through the use of additional information, such as comments on user messages being analyzed.
To solve such problems, including the problem of continuous authentication based on collected datasets, synthetic datasets with class labels were constructed, simulating a change in editorial policy in connection with the sale or theft of personal pages, using channels as an example. This simulation was carried out by adding messages from a similar channel to the main channel. The channels, on which the mixtures were based, were selected so that they belonged to the same general topic and were similar in terms of the average number of words in a publication (message).
Next, channel mixes were constructed in one of two ways:
  • The latest messages from the second channel were added to the messages from the first channel, according to a predefined ratio/percentage. By default, the proportion of added messages is 10%. The choice of a 10% anomaly level is dictated by the key requirement of the continuous authentication task—minimizing detection delay. An anomaly must remain a rare event to preserve its diagnostic value as an indicator of compromise. A high proportion of anomalies transforms them into a new stable pattern, which contradicts the goal of early detection of a subtle deviation in the dataset. This threshold allows the validation of the model’s ability to quickly detect an intrusion before it becomes statistically predominant. Thus, messages added from another channel represent the proportion of “normal behavior” in terms of editorial, stylistic, and content policy, while added messages represent outliers. The two channels were selected so that the messages belonged to the same list of topics and had a similar average publication length. Given that the added publications were written by a different group of authors, they have different stylistic and syntactic characteristics and can thus be identified as outliers relative to “normal” messages. The date of the “injected” messages is taken from the original normal messages, and an additional synthesized date field is added, formed on the basis of the average frequency of publications in the normal part of the collection. In order to conduct experiments on collections of the same size, synthetic datasets of equal size were formed: 5, 10, and 50 thousand original messages.
  • A specified percentage of messages (by default, 10% of the messages from the source channel) from the second channel, arranged in chronological order, were added to the messages from the first main channel. The set of messages for these dates was taken from the first channel and replaced with messages from the second channel. The publication dates were also added from the second channel. Unique message identifiers were also added from the second channel. The second method gives a the possibility to take into account the impact of global economic news on the subject of messages in economic channel publications and makes such a dataset more plausible. The number of original messages remains the same as it was initially, with the exception of the inserted messages.
The “to_inject” collection stores multiple pairs consisting of two collections that are similar in terms of the average number of words per message. Based on these pairs, the above-mentioned annotated collections are synthesized. Information about mixed messages is presented in a special collection called “injected_chats” with the following set of attributes (Table 3). Each mixed channel message is characterized by a set of added attributes presented in Table 4.
Table 5 shows examples of normal messages and messages considered to be outliers. The data sample shows that it is extremely difficult for humans to visually detect outliers.

4.4. Empirical Validation of Synthetic Datasets

To validate how well the synthetic data reflect real scenarios of channel authorship changes according to the proposed threat model (Section 3.3), we conducted an empirical analysis. We used 30 type 1 datasets and 30 type 2 datasets. Type 1 datasets contained 5000 Telegram messages with a controlled anomaly rate of 10% (label 1 represents anomalous messages). Type 2 datasets had an anomaly percentage of 9–13% and varied in size from approximately 4000 to 34,000 messages.
Due to computational constraints, we applied stratified sampling to type 2 datasets: from each dataset, we extracted a fixed-size subsample of 4000 messages while preserving the original distribution of legitimate and anomalous messages. During stratification, we performed preprocessing to remove messages without meaningful textual content (emoji-only messages, media without text, etc.). This preprocessing slightly adjusted the anomaly proportions, resulting in an actual anomaly rate of 6.2–16.3% (mean 11.7% ± 2.2%) for type 2 datasets.
The experimental design followed the threat model hierarchy: the most challenging scenario (threat 3) combines authorship change with style imitation, while the second-most challenging scenario (threat 1) preserves topic consistency without style imitation (see Section 3.3).
Topic homogeneity was ensured through multiple approaches. First, all Telegram channels were selected from the economics category on TGstat [20] (see Section 4.1). Second, we performed BERTopic [43] topic modeling on a sample of 300,000 messages from the full corpus (~2.9 million messages; Section 5). The modeling confirmed that over 65% of messages belong to economic themes, with the most frequent topics being Russian economy (26%), cryptocurrencies and investments (16%), and foreign policy/economics (18%). Additionally, frequent cross-references between channels further indicate thematic proximity (see Section 5).
Focusing on a close average number of words when creating synthetic datasets, as well as relatively close publication dates, potentially makes the style of publications similar and models threats 1 or 3. In this work, we focus on the most complex cases (threats 1 and 3), as they are of the greatest practical interest for continuous authentication systems. threat 2, although important, is a more trivial task in terms of detection. Thus, the goal of the experimental analysis is to evaluate the degree of similarity and overlap between normal and anomalous messages.
To quantify the distribution overlap between legitimate (label 0) and anomalous (label 1) messages, we employed three complementary metrics:
  • Cosine distances between text vector representations allow assessing semantic similarity at the level of individual messages, for which the following are used:
    • Inter-class cosine distance—the distance from objects of the opposite class to the centroid.
    • Inter-centroid distances—distances between the centroids of opposite classes.
  • Fréchet Distance (FD) [44] is a strict measure for comparing multivariate distributions that accounts not only for the position of distribution centers (like cosine distance), but also for their shape and spread. In addition to FD, normalized FD is calculated in the range 0 to 1 for interpretability, enabling comparison of complexity across different datasets independently of absolute FD values. Normalized FD represents an inverted and scaled version of FD in the range 0 to 1, where 1 corresponds to complete closeness of distributions, and 0 to their maximum divergence.
The choice of metrics is motivated by the need to evaluate both local semantic proximity and global distribution similarity in relation to the threat model. Vector representations for all 270,000 messages were obtained using the pre-trained rubert-tiny2 model [45]. The selection of this compact model is justified by its efficiency for Russian-language texts and computational constraints when processing this volume of data.
Complete numerical results for all 60 datasets are available in the research repository (Data Availability Statement). Analysis of type 1 experimental results (Table 6) via cosine similarity reveals high semantic coherence: 87% of datasets (26 of 30) exhibit inter-centroid distances below 0.30 (mean = 0.232 ± 0.076). Inter-class distance analysis further indicates substantial distribution overlap, with class 0 samples averaging 0.359 ± 0.080 from the class 1 centroid, and class 1 samples averaging 0.397 ± 0.089 from the class 0 centroid.
The Fréchet Distance analysis (Table 6) confirmed strong distribution overlap: all 30 datasets exhibit FD < 1.0 (mean = 0.607 ± 0.230, range [0.205, 0.864]). These values correspond to extremely challenging classification cases. Since FD remained below 1.0 for all datasets, the synthetic datasets are suitable for modeling the most difficult threats—specifically when a new author team attempts precise style imitation.
The normalized closeness score (0.940 ± 0.019) quantitatively confirms that the created datasets constitute a challenging test benchmark for authentication systems.
To visually illustrate the degree of distribution overlap, Figure 4 presents two-dimensional PCA projections of embeddings for two extreme cases from type 1 datasets: the closest case (FD = 0.205) and the farthest case (FD = 0.864). Even for the dataset with the smallest FD value, clear distribution overlap is observable, confirming that all synthetic datasets model complex, non-trivial user verification scenarios.
Table 7 presents the experimental results for type 2 datasets. Cosine similarity analysis reveals high semantic coherence: 90% of datasets (27 of 30) exhibit inter-centroid distances below 0.20 (mean = 0.135 ± 0.075). Inter-class distances further demonstrate distribution proximity: class 0 samples average 0.359 ± 0.080 from the class 1 centroid, while class 1 samples average 0.322 ± 0.075 from the class 0 centroid, indicating substantial class overlap. The mean cosine distance in type 2 datasets is nearly half that of type 1 (0.135 versus 0.232), reflecting stronger semantic relatedness between legitimate and anomalous messages.
Fréchet Distance analysis confirms substantial distribution overlap for type 2 datasets: all 30 datasets exhibit FD < 1.0 (mean = 0.414 ± 0.218, range [0.094, 0.813]). FD values show greater variability compared to type 1, reflecting the impact of stratification and the inherent variability of original anomaly distributions.
The histogram in Figure 5 illustrates the FD distribution, while Figure 6 shows the cosine centroid distance distribution. Both histograms confirm that type 2 datasets model more challenging scenarios associated with the threat of author collective change coupled with message style imitation.
Overall, the experimental results confirm that the synthetic datasets align with threat scenarios 1 and 3. Analysis of Fréchet Distance (FD) and cosine measures revealed substantial overlap between legitimate and anomalous message distributions across all 60 datasets (FD < 1.0), modeling challenging continuous authentication cases.

5. Analysis of the Collected Dataset

The entire dataset is 6.3 GB in size, contains 486 channels, and 2,918,303 posts, mainly in Russian. Only 10 channels contain posts in English. The language of the channels was identified using the langdetect library [46]. Considering that many channels repost messages from news channels and also mention the news title and the news channel itself, these news channels are also included in the dataset described. In order to characterize the channels in this dataset, Figure 7 shows the cumulative distribution function of a discrete random variable—the number of messages per channel. At the same time, with a probability of 0.996, the number of messages in the channel is no more than 100,000.
The cumulative distribution function of the number of subscribers to a channel shows that most channels have fewer than 250,000 subscribers (with a probability of 0.903), which is also confirmed by the histogram shown in Figure 8. The histogram (Figure 9) shows that the number of subscribers to a channel is distributed according to a power distribution function.
Some messages in the dataset are marked with hashtags and also contain references to other channels. Analysis of the dataset showed that only 456,504 posts contain hashtags, which is 15.6% of the total number of posts. The number of hashtags contained in a single post varies from 1 to 136. Figure 10 shows the cumulative distribution function of the discrete random variable of hashtag usage in posts. When constructing this cumulative function, posts in which hashtags are not specified were not taken into account. It should be noted that most publications containing hashtags use fewer than 10 (probability 0.996). Figure 11 shows a fragment of the histogram for 10 or fewer hashtags. Table 8 shows the top 100 most popular hashtags across the entire dataset.
Other channels are mentioned in 18.8% of messages, of which only 1.2% of mentions do not contain mentions of the same channel or mentions of special accounts such as bots. Figure 12 shows the cumulative distribution function of the discrete random variable of mention usage. When constructing this cumulative function, publications in which mentions are not represented or only one mention is used, i.e., a mention of the same channel or a channel bot with the same name, were not taken into account. Most publications contain no more than 10 mentions, with a probability of 0.996. Figure 13 shows a fragment of the histogram for 200 or fewer mentions. The most frequent mentions are listed in Table 9. For comparison, Table 10 shows the most frequent mentions without any filtering.
Analysis of the text of publications in the collected dataset showed that out of the total number of messages, 1690 contain only emojis without a single word, and 174,847, or 6%, contain images without accompanying text. Figure 14 and Figure 15 show the cumulative distribution function of the discrete random variable of the number of characters and words in publications in the collected dataset.
Unigrams and bigrams were also collected across the entire dataset. Each token is a word, with the exception of the stop word list, which also includes channel names. Given the significant computational costs involved in processing all the data, unigrams and bigrams were collected using the following less computationally intensive algorithm. For each channel, the 100 most frequent unigrams and bigrams were found, and then a list of the most frequent unigrams and bigrams in the dataset was compiled from this set of unigrams. Table 11 and Table 12 show the most frequent bigrams and unigrams across the entire dataset.
In addition, an analysis of the message topics was performed based on the Topicbert model [43] in unsupervised mode. Topic modeling across the entire dataset requires significant computational resources, including a large amount of RAM at the stage of clustering the obtained data by topic [47]. Therefore, to reduce costs, a sampling method was used, during which a training dataset of 300,000 publications was formed, preserving the proportions of messages from each channel. This volume is related to the available computing power. The quality criteria [48] for topic modeling were as follows: Coherence Score = 0.77, Topic Diversity = 0.89. The language model “paraphrase-multilingual-MiniLM-L12-v2” [43] was used for topic modeling with the minimum cluster size set to 100 in order to capture more significant topics. The resulting model allows us to identify topics for each publication from the collected dataset. The number of automatically identified topics was 300, with each topic occupying 7% or less of the sample. The model did not recognize 48% of publications, which is a normal result for the given type of model. In order to evaluate the topics globally, all topics were converted by combining similar clusters. Table 13 shows the most frequent topics with a share exceeding 1%, excluding noise publications.
Figure 16 shows the publications obtained, broken down by topic by reducing the dimensionality of their vector representation to two. It can be noted that events occurring in the world are well reflected in economic channels, along with economic topics.
Given that economic channels often quote and refer to news channels, the topic modeling data was compiled taking into account publications from news channels, which were also downloaded as reference channels. It should also be noted that six channels were censored by the Telegram messenger. From a copyright perspective, such publications were placed in a separate cluster, and the texts of their publications were unavailable and replaced with messages from Telegram about censorship.

6. Enriching Collected Data with Financial Indicators and Analyzing Them

In order to understand the reason for the appearance of a particular publication on economic topics, as well as for a deeper semantic understanding of the publications, the collected dataset was enriched with stock market quotes and indicators [49].
The time interval from 1 September 2022, to 1 December 2024, for which messages were uploaded, corresponds to the development of the trend in the gold futures contract for the growth of exchange gold, including correction and resumption of the trend (Figure 17).
Looking at the price of gold as a safe-haven asset in isolation can overlook commodity price cycles, periods of market panic, and declines in the number of lots purchased with borrowed money. Therefore, in addition to gold quotes, quotes for the USD/RUB currency pair, BRENT oil, the dollar index (DXY), and bitcoin (BTC) are also downloaded. The following indicators are calculated for each exchange instrument (ticker):
  • A 200-day moving average (200MA).
  • The closing price is above or below the 200-day average (AboveMA).
  • The number of consecutive days during which the quote value exceeds the moving average value. The greater the number of days, the longer the upward trend develops without significant corrective movement.
  • Relative strength indicator (RSI), reflecting overbought or oversold conditions.
  • When trading in the above-mentioned exchange instruments did not take place, holidays and weekends are supplemented with quotes from the last trading day.
The relationship between intraday volatility (the absolute difference between the maximum and minimum prices) and the number of mentions of a stock exchange instrument was analyzed. The p-value does not exceed 0.001 in all cases. This analysis made it possible to check for the presence of channels whose publications are related to the volatility of the aforementioned exchange instruments (Table 14).
For each of the exchange instruments, channels were identified in which the frequency of mentions of the instrument correlates with volatility. An increase in volatility traditionally attracts the attention of economic analysts and the interested audience. The identified channels confirm the presence of authors and readers who follow price changes of the exchange instrument.

7. Conclusions and Discussion

This paper presents a new open dataset and methodological toolkit developed to address key challenges in text data analysis and continuous authentication. Our main contribution is methodological and consists of three interrelated components:
  • A resource for stylometry in Social Media Publishing Economic News. We have collected and structured a dataset of Russian-language Telegram channels on economic topics, enriched it with metadata, and synchronized the financial indicators. Its structure allows for the study of linguistic patterns in relation to external context, while maintaining a balance between research utility and ethical requirements, as the data does not allow for author identification.
  • A framework for generating realistic test data. The core methodological innovation of this work is our proposed method for generating synthetic labeled datasets designed to simulate an author-change scenario. This framework models a realistic account compromise, where messages from different authors are interwoven while preserving a unified thematic focus. The controlled anomaly level (set at ~10%) is crucial for validating models’ ability to perform early detection of authorship shifts—a fundamental requirement for continuous authentication systems.
  • A tool for research reproducibility. We provide a complete set of scripts for dataset formation, enrichment, and the generation of synthetic collections. This ensures full reproducibility of the results and offers researchers a ready-made instrument for creating similar datasets for other subject areas and languages.
Thus, the primary contribution of this work lies not in economic analysis, but in advancing the methodology for evaluating machine learning models in the field of continuous text-based authentication, while providing a unique testing ground. The resources presented enable the comparison of models under controlled yet realistic conditions, where the core challenge is the detection of stylistic anomalies against a backdrop of thematic homogeneity.
This challenge is further amplified by the specific nature of the data—namely, a substantial volume of texts from professional market analysts. Messages of this type combine a rigorous logical framework with elements of expressiveness; they are logical and structured (akin to reports) on one hand, yet often emotional and persuasive (like forecasts) on the other. This duality allows for a more nuanced validation of algorithms in tasks of stylometry, including sentiment analysis in professional discourse and the extraction of cause-and-effect assertions from text. This opens avenues for further research in the security of text data streams, adaptive authentication methods, and style analysis in social networks.

Author Contributions

Conceptualization, methodology, software, validation, formal analysis, data curation, writing—original draft preparation, writing—editing: E.L. and P.B., writing—review and editing, funding acquisition, supervision: A.S. All authors have read and agreed to the published version of the manuscript.

Funding

This work was carried out with financial support from the Ministry of Science and Higher Education of the Russian Federation as part of the basic part of TUSUR’s state assignment for 2023–2025 (project No. FEWM-2023-0015).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data collection methodology did not aim to collect specific messages or messages from a specific author or channel. The collection was based on the presence of channels on the TGSTAT resource, as well as mentions of channels during data collection. The original data and experimental scripts and results presented in the study are openly available at https://github.com/pavel805/TGEconomicDataset (accessed on 16 January 2026).

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

Abbreviations

The following abbreviations are used in this manuscript:
BTCbitcoin
DXYdollar index
CDFcumulative distribution function
RSIrelative strength indicator

References

  1. DemandSage. Telegram Statistics 2025: Worldwide Data. Available online: https://www.demandsage.com/telegram-statistics/ (accessed on 2 December 2025).
  2. Mehdipour, S.; Jannati, N.; Negarestani, M.; Amirzadeh, S.; Keshvardoost, S.; Zolala, F.; Vaezipour, A.; Hosseinnejad, M.; Fatehi, F. Health pandemic and social media: A content analysis of COVID-related posts on a Telegram channel with more than one million subscribers. Stud. Health Technol. Inform. 2021, 279, 122–129. [Google Scholar] [CrossRef] [Scilit]
  3. Lou, C.; Tandoc, E.C.; Hong, L.X.; Pong, X.Y.; Lye, W.X.; Sng, N.G. When motivations meet affordances: News consumption on Telegram. J. Stud. 2021, 22, 934–952. [Google Scholar] [CrossRef] [Scilit]
  4. Griggio, C.F.; Nouwens, M.; Klokmose, C.N. Caught in the network: The impact of WhatsApp’s 2021 privacy policy update on users’ messaging app ecosystems. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (CHI ‘22), New York, NY, USA, 30 April–5 May 2022; pp. 1–23. [Google Scholar] [CrossRef] [Scilit]
  5. Al-Rawi, A. News loopholing: Telegram news as portable alternative media. J. Comput. Soc. Sci. 2022, 5, 949–968. [Google Scholar] [CrossRef] [Scilit]
  6. Gupta, K.; Oladimeji, D.; Varol, C.; Rasheed, A.; Shahshidhar, N. A comprehensive survey on artifact recovery from social media platforms: Approaches and future research directions. Information 2023, 14, 629. [Google Scholar] [CrossRef] [Scilit]
  7. Telegram APIs. Available online: https://core.telegram.org/api (accessed on 2 December 2025).
  8. Herrero-Solana, V.; Castro-Castro, C. Telegram channels and bots: A ranking of media outlets based in Spain. Societies 2022, 12, 164. [Google Scholar] [CrossRef] [Scilit]
  9. Channels FAQ. Available online: https://telegram.org/faq_channels?setln=en (accessed on 2 December 2025).
  10. Xiao, C.; Freeman, D.M.; Hwa, T. Detecting clusters of fake accounts in online social networks. In Proceedings of the 8th ACM Workshop on Artificial Intelligence and Security, Denver, CO, USA, 16 October 2015; ACM Press: New York, NY, USA, 2015; pp. 91–101. [Google Scholar]
  11. La Morgia, M.; Mei, A.; Sassi, F.; Stefa, J. Pump and dumps in the Bitcoin era: Real time detection of cryptocurrency market manipulations. In Proceedings of the 2020 IEEE 29th International Conference on Computer Communications and Networks (ICCCN), Honolulu, HI, USA, 3–6 August 2020; pp. 1–9. [Google Scholar]
  12. La Morgia, M.; Mei, A.; Sassi, F.; Stefa, J. The Doge of Wall Street: Analysis and detection of pump and dump cryptocurrency manipulations. ACM Trans. Internet Technol. 2023, 23, 1–28. [Google Scholar] [CrossRef] [Scilit]
  13. Rajaei, M.J.; Mahmoud, Q.H. A survey on pump and dump detection in the cryptocurrency market using machine learning. Future Internet 2023, 15, 267. [Google Scholar] [CrossRef] [Scilit]
  14. Conti, E.; Salvi, D.; Borrelli, C.; Hosler, B.; Bestagini, P.; Antonacci, F.; Sarti, A.; Stamm, M.C.; Tubaro, S. Deepfake speech detection through emotion recognition: A semantic approach. In Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2022), Singapore, 22–27 May 2022; pp. 8962–8966. [Google Scholar]
  15. Rao, S.; Verma, A.K.; Bhatia, T. Review on social spam detection: Challenges, open issues, and future directions. Expert. Syst. Appl. 2021, 186, 115742. [Google Scholar] [CrossRef] [Scilit]
  16. Ayeswarya, S.; Singh, K.J. A comprehensive review on secure biometric-based continuous authentication and user profiling. IEEE Access 2024, 12, 82996–83021. [Google Scholar] [CrossRef] [Scilit]
  17. Li, J.; Zheng, R.; Chen, H. From Fingerprint to Writeprint. Commun. ACM 2006, 49, 76–82. [Google Scholar] [CrossRef] [Scilit]
  18. Fedotova, A.; Romanov, A.; Kurtukova, A.; Shelupanov, A. Authorship attribution of social media and literary Russian-language texts using machine learning methods and feature selection. Future Internet 2022, 14, 4. [Google Scholar] [CrossRef] [Scilit]
  19. Wang, Y.; Zhang, X.; Hu, H. Continuous user authentication on multiple smart devices. Information 2023, 14, 274. [Google Scholar] [CrossRef] [Scilit]
  20. Telegram Channels and Groups Catalog. Available online: https://tgstat.ru/en (accessed on 2 December 2025).
  21. Baumgartner, J.; Zannettou, S.; Squire, M.; Blackburn, J. The Pushshift Telegram Dataset. Proc. Int. AAAI Conf. Web Soc. Media 2020, 14, 840–847. [Google Scholar] [CrossRef] [Scilit]
  22. Hoseini, M.; Melo, P.; Benevenuto, F.; Feldmann, A.; Zannettou, S. On the globalization of the QAnon conspiracy theory through Telegram. arXiv 2021, arXiv:2105.13020. [Google Scholar] [CrossRef] [Scilit]
  23. Hashemi, A.; Zare Chahooki, M.A. Telegram group quality measurement by user behavior analysis. Soc. Netw. Anal. Min. 2019, 9, 33. [Google Scholar] [CrossRef] [Scilit]
  24. Sosa, J.; Sharoff, S. Multimodal pipeline for collection of misinformation data from Telegram. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, Marseille, France, 20–25 June 2022; European Language Resources Association: Paris, France, 2022; pp. 1480–1489. [Google Scholar]
  25. La Morgia, M.; Mei, A.; Mongardini, A.M. TGDataset: Collecting and exploring the largest Telegram channels dataset. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1 (KDD ‘25), New York, NY, USA, 3–7 August 2025; pp. 2325–2334. [Google Scholar] [CrossRef] [Scilit]
  26. Angermaier, M.; Pinheiro Neto, J.; Höldrich, E.; Lasser, J. Dataset for the “The Schwurbelarchiv: A German Language Telegram dataset for the study of conspiracy theories” paper [Data set]. arXiv 2025, arXiv:2504.06318. [Google Scholar] [CrossRef]
  27. Kireev, K.; Mykhno, Y.; Troncoso, C.; Overdorf, R. A Telegram dataset of propaganda and its moderation. Proc. Int. AAAI Conf. Web Soc. Media 2025, 19, 2510–2518. [Google Scholar] [CrossRef] [Scilit]
  28. Gangopadhyay, S.; Dessí, D.; Dimitrov, D.; Dietze, S. TeleScope: A longitudinal dataset for investigating online discourse and information interaction on Telegram. Proc. Int. AAAI Conf. Web Soc. Media 2025, 19, 2423–2433. [Google Scholar] [CrossRef] [Scilit]
  29. Groups and Channels. Available online: https://telegram.org/faq#groups-and-channels (accessed on 2 December 2025).
  30. MessageEntity. Available online: https://core.telegram.org/type/MessageEntity (accessed on 2 December 2025).
  31. Bansal, P.; Ouda, A. Continuous authentication in the digital age: An analysis of reinforcement learning and behavioral biometrics. Computers 2024, 13, 103. [Google Scholar] [CrossRef] [Scilit]
  32. Shopon, M.; Tumpa, S.N.; Bhatia, Y.; Kumar, K.N.P.; Gavrilova, M.L. Biometric systems de-identification: Current advancements and future directions. J. Cybersecur. Priv. 2021, 1, 470–495. [Google Scholar] [CrossRef] [Scilit]
  33. Estrela, P.M.A.B.; Albuquerque, R.d.O.; Amaral, D.M.; Giozza, W.F.; Júnior, R.T.d.S. A framework for continuous authentication based on touch dynamics biometrics for mobile banking applications. Sensors 2021, 21, 4212. [Google Scholar] [CrossRef] [Scilit]
  34. Choi, M.; Lee, S.; Jo, M.; Shin, J.S. Keystroke dynamics-based authentication using unique keypad. Sensors 2021, 21, 2242. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Ashibani, Y.; Kauling, D.; Mahmoud, Q.H. Design and implementation of a contextual-based continuous authentication framework for smart homes. Appl. Syst. Innov. 2019, 2, 4. [Google Scholar] [CrossRef] [Scilit]
  36. Bułat, R.; Ogiela, M.R. Personalized context-aware authentication protocols in IoT. Appl. Sci. 2023, 13, 4216. [Google Scholar] [CrossRef] [Scilit]
  37. Baig, A.F.; Eskeland, S. Security, privacy, and usability in continuous authentication: A survey. Sensors 2021, 21, 5967. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Fedotova, A.; Kurtukova, A.; Romanov, A.; Shelupanov, A. Semantic clustering and transfer learning in social media texts authorship attribution. IEEE Access 2024, 12, 39783–39803. [Google Scholar] [CrossRef] [Scilit]
  39. Cho, S.H.; Kim, B.; Kwon, H.; Kim, M. Exploring the potential of large language models for author profiling tasks in digital text. In Notebook for PAN at CLEF 2023; Elsevier: Amsterdam, The Netherlands, 2025; pp. 1272–1281. [Google Scholar]
  40. Brocardo, M.L.; Traore, I.; Woungang, I. Authorship verification of e-mail and tweet messages applied for continuous authentication. J. Comput. Syst. Sci. 2015, 81, 1429–1440. [Google Scholar] [CrossRef] [Scilit]
  41. Li, J.; Zhang, Q.; Huang, M. Author verification of text fragments based on the BERT model. In Notebook for PAN at CLEF 2023; Foshan University: Foshan, China, 2023; pp. 1–4. [Google Scholar]
  42. Telemetr—Telegram Channel Search and Analytics Platform. Available online: https://telemetr.me/ (accessed on 16 January 2026).
  43. BERTopic: Neural Topic Modeling with a Class-Based TF-IDF Procedure. Available online: https://arxiv.org/abs/2203.05794 (accessed on 2 December 2025).
  44. Dowson, D.C.; Landau, B.V. The Fréchet distance between multivariate normal distributions. J. Multivar. Anal. 1982, 12, 450–455. [Google Scholar] [CrossRef] [Scilit]
  45. Cointegrated/Rubert-Tiny2: Lightweight Russian BERT Model. 2021. Available online: https://huggingface.co/cointegrated/rubert-tiny2 (accessed on 16 January 2026).
  46. Language Detection Library in Python. Available online: https://github.com/fedelopez77/langdetect (accessed on 2 December 2025).
  47. Memory Inefficient Algorithm and Getting Error While Saving the Model. Available online: https://github.com/MaartenGr/BERTopic/issues/173 (accessed on 2 December 2025).
  48. Terragni, S.; Fersini, E. Word embedding-based topic similarity measures. In Proceedings of the International Conference on Applications of Natural Language to Information Systems (ANLIS 2021), Saarbrücken, Germany, 23–25 June 2021; Springer: Berlin/Heidelberg, Germany, 2021; pp. 33–45. [Google Scholar]
  49. MarketWatch. Download Data. Available online: https://www.marketwatch.com/investing/cryptocurrency/btcusd/download-data? (accessed on 2 December 2025).
Figure 1. Telegram network message structure and related entities (1–*—one to many).
Figure 1. Telegram network message structure and related entities (1–*—one to many).
Data 11 00025 g001
Figure 2. Message downloading process.
Figure 2. Message downloading process.
Data 11 00025 g002
Figure 3. Message processing and enrichment process.
Figure 3. Message processing and enrichment process.
Data 11 00025 g003
Figure 4. Comparison of PCA projections for closest and farthest cases for type 1 dataset.
Figure 4. Comparison of PCA projections for closest and farthest cases for type 1 dataset.
Data 11 00025 g004
Figure 5. Comparison of Frechet Distance between centroids for type 1 and type 2 datasets.
Figure 5. Comparison of Frechet Distance between centroids for type 1 and type 2 datasets.
Data 11 00025 g005
Figure 6. Comparison of cosine distance between centroids for type 1 and type 2 datasets.
Figure 6. Comparison of cosine distance between centroids for type 1 and type 2 datasets.
Data 11 00025 g006
Figure 7. Cumulative distribution function of the number of messages per channel.
Figure 7. Cumulative distribution function of the number of messages per channel.
Data 11 00025 g007
Figure 8. CDF of the number of registered users per channel.
Figure 8. CDF of the number of registered users per channel.
Data 11 00025 g008
Figure 9. Fragment of a histogram showing the number of subscribers to a channel (for channels with fewer than 100,000 subscribers).
Figure 9. Fragment of a histogram showing the number of subscribers to a channel (for channels with fewer than 100,000 subscribers).
Data 11 00025 g009
Figure 10. CDF of using hashtags in publications.
Figure 10. CDF of using hashtags in publications.
Data 11 00025 g010
Figure 11. Histogram of hashtag number used in publications.
Figure 11. Histogram of hashtag number used in publications.
Data 11 00025 g011
Figure 12. CDF of using mentions in publications.
Figure 12. CDF of using mentions in publications.
Data 11 00025 g012
Figure 13. Histogram fragment of the number of mentions in the dataset publications.
Figure 13. Histogram fragment of the number of mentions in the dataset publications.
Data 11 00025 g013
Figure 14. CDF of number of symbols in publications.
Figure 14. CDF of number of symbols in publications.
Data 11 00025 g014
Figure 15. CDF of number of words in publications.
Figure 15. CDF of number of words in publications.
Data 11 00025 g015
Figure 16. Topics of publications obtained by combining clusters (x and y are dimensions obtained during the dimensionality reduction of the vector representation of messages).
Figure 16. Topics of publications obtained by combining clusters (x and y are dimensions obtained during the dimensionality reduction of the vector representation of messages).
Data 11 00025 g016
Figure 17. Growth of the exchange price of gold futures for the period of data downloaded from the Telegram messenger (orange lines are the trend lines).
Figure 17. Growth of the exchange price of gold futures for the period of data downloaded from the Telegram messenger (orange lines are the trend lines).
Data 11 00025 g017
Table 1. Description of attributes of the “chats” collection.
Table 1. Description of attributes of the “chats” collection.
Attribute NameDescription
_idMongoDB document identifier
TitleChannel name exactly as in Telegram
postedDate the document was added to the collection
timestampDate of data collection in Unix format, i.e., the number of seconds that have elapsed since 00:00:00 UTC on 1 January 1970.
CollectionCollection name where messages from the channel are stored
SourceSource (if the data is loaded from Telegram, then the tag is “tg”)
ch_mentionsMentioned channels in the channel
forwards_from_chatChannels from which messages are forwarded to the channel
forwards_from_userUsers from whom messages are forwarded to the channel
tg_urlsURLs mentioned in the channel messages
OriginSource of information about the channel, either TGStat [41] or a mention in another channel
members_countNumber of subscribers (participants)
CompletedBit flag (1 if the channel messages are completely downloaded)
last_idLast downloaded message ID
updatedBit flag (1 if all messages of the channel are processed)
updated_mgBit flag (1 if messages with the same media group are merged)
Table 2. Description of attributes of collections containing messages from Telegram channels.
Table 2. Description of attributes of collections containing messages from Telegram channels.
Attribute NameDescription
_idMongoDB document identifier
textMessage text
timestampDate of publication of the message in Unix format, i.e., the number of seconds that have elapsed since 00:00:00 UTC on 1 January 1970.
posted_str_dateDate of publication of the message in the format “DD/MM/YYYY”
postedDate of the message creation
market_entitiesDictionary with values of financial instruments
viewsMessage view count in a Telegram channel
msg_idTelegram message ID
forwardsThe number of message forwards/reposts in a Telegram channel
mediaType of media
reactionsDictionary of reactions to the message
edit_dateDate of the message editing
hashtagsArray of hashtags
entitiesArray of message entities
mentionsArray of mentions
tg_urlsArray of Telegram URLs (https://t.me/…) in the message
symbolsNumber of symbols in the message
wordsNumber of words in the message
Table 3. Set of attributes for the collection of injected chats.
Table 3. Set of attributes for the collection of injected chats.
Attribute NameDescription
_idRecord identifier
avg_word_injAverage number of words in a channel’s message from which outlier posts are added
avg_word_mainAverage number of words in posts on the original channel with non-outlier posts.
collectionName of the resulting synthesized collection
collection_injName of the collection containing raw data from the Telegram channel with outlier messages
collection_mainName of the collection containing raw data from the main Telegram channel with non-outlier posts
injected_countNumber of added outlier messages
count_mainNumber of messages marked as original non-outlier
PercentPercentage of outliers
sizeThe number of messages; if 0, the channel contains all messages. Otherwise, the number of latest messages that were taken from the source channel.
typecollection synthesis method: type 1 corresponds to method 1; type 2 to method 2, described above.
Table 4. Metadata for synthesized datasets.
Table 4. Metadata for synthesized datasets.
Attribute NameDescription
_idPublication ID
textMessage text
timestampTimestamp of the original dataset
newtimestampFor added outlier message, the timestamp is set based on the average frequency of publications in the main channel.
outlierOutlier tag: 0, normal message; 1, outlier.
Table 5. Fragments of injected dataset.
Table 5. Fragments of injected dataset.
ChannelMessage TextOutlier Tag
Name of the synthesized collection:
PerfecProfit_and_whalebotalerts_10_size_5000_type1
PerfecProfitData 11 00025 i004HI PRICEData 11 00025 i005
Tiker—MAGN
Price—55.705
06.05.2024—10:03:00
#MAGN”
0
Data 11 00025 i004HI PRICEData 11 00025 i005
Tiker—YNDX
Price—4282.6
06.05.2024—10:03:00
#YNDX”
0
whalebotalerts117 BTC ($11,321,512) transferred from Kraken to Unknown
Sender’s balance: 0
1
Crypto Fear/Greed Index
Status: Extreme Greed
Value: 84/100
Next Update in 19 h
1
Table 6. Validation metrics statistics for 30 type 1 datasets.
Table 6. Validation metrics statistics for 30 type 1 datasets.
MetricMean ± SDRange [Min, Max]Interpretation
Cosine distance between centroids0.232 ± 0.076[0.069, 0.337]High semantic similarity between classes
Fréchet Distance 0.607 ± 0.230[0.205, 0.864]Strong distribution overlap (all FD < 1.0)
Normalized closeness0.940 ± 0.019[0.915, 0.979]Extremely challenging verification scenario
Distance from class 0 to centroid 10.359 ± 0.080[0.203, 0.476]Inter-class proximity (class 0 → 1)
Distance from class 1 to centroid 00.397 ± 0.089[0.225, 0.507]Inter-class proximity (class 1 → 0)
Table 7. Validation metrics statistics for 30 type 2 datasets.
Table 7. Validation metrics statistics for 30 type 2 datasets.
MetricMean ± SDRange [Min, Max]Interpretation
Cosine distance between centroids0.135 ± 0.075[0.018, 0.279]High semantic similarity between classes
Fréchet Distance (FD)0.414 ± 0.218[0.094, 0.813]Substantial distribution overlap
Normalized closeness0.959 ± 0.022[0.919, 0.991]Challenging verification scenario
Distance from class 0 to centroid 10.359 ± 0.080[0.220, 0.490]Interpretation
Distance from class 1 to centroid 00.322 ± 0.075[0.211, 0.491]High semantic similarity between classes
Table 8. Most frequent hashtags in the dataset.
Table 8. Most frequent hashtags in the dataset.
RankHashtagFrequency of Use in the DatasetRankHashtagFrequency of Use in the Dataset
1#poccия’32,20526#Пoлитикa’3781
2#Aкции’22,34927#пpoгнoз’3742
3#Импyльc’21,58528#инφляция’3716
4#кpиптo’20,15229#гaз’3641
5#cшa’13,74630#YNDX’3490
6#oтчeтнocть’11,05931#ROSN’3452
7#aкции’958832#cot’3438
8#нeφть’856033#бaнки’3359
9#GAZP’732234#LKOH’3307
10#pынки’708435#SMLT’3082
11#BTC’645436#sentiment’2919
12#SBER’602437#дивидeнды’2901
13#fx’580538#OZON’2831
14#китaй’570839#POSI’2829
15#дкп’565940#VKCO’2808
16#экoнoмикa’551541#TCSG’2785
17#гeoпoлитикa’548842#дивидeнд’2736
18#eвpoпa’485443#epзнoвocти’2719
19#мaкpo’459144#нoвocти’2693
20#MOEX’441645#GMKN’2632
21#Ъyзнaл’421146#AFLT’2524
22#caнкции’410047#ALRS’2501
23#VTBR’402448#CHMF’2477
24#Maкpo’396949#Пpeмapкeтинг’2477
25#NVTK’383250#Пoлитикa’3781
Table 9. Most frequent mentions.
Table 9. Most frequent mentions.
MentionNumber of Mentions
1@ifax_go’2655
2@kriptokondrashov’1438
3@selfinvestor’1204
4@investinfopro_mmvb’1137
5@bitkogan’/Bitkogan’953
6@banksta’916
7@AK47pfl’592
8@offcrypt’544
9@wisedruidd’485
10@keepcontact’457
11@voin_dv’452
12@rocket1signals’450
13@interfaxonline’432
14@ntvnews’418
15@fanimani_official’361
16@investi_usa’330
17@rybar’290
18@alfabank’280
19@Dohod’271
20@FiboAlex’264
Table 10. Most frequent mentions without filtering.
Table 10. Most frequent mentions without filtering.
MentionNumber of Mentions
1@WhaleBotAlerts’148,634
2@banksta’26,161
3@bankrollo’23,360
4@if_market_news’21,955
5@interfaxonline’21,531
6@bankoffo’19,148
7@kommersant’18,549
8@banki_oil’14,939
9@marketsnapshot’14,221
10@ifax_go’13,996
11@voenkorKotenok’13,516
12@PerfecProfit’11,791
13@fatcat18’11,213
14@kstati_p’11,005
15@kfm936’10,473
16@ejdailyru’8586
17@selfinvestor’8461
18@bitkogan’7625
19@ivan_utenkov13’6473
20@milinfolive’6347
Table 11. Most frequent unigrams.
Table 11. Most frequent unigrams.
Original UnigramUnigram in EnglishNumber of Occurrences
1гoдyear1,058,316
2PoccияRussia643,307
3мoчьbe able433,262
4кoмпaнияcompany395,978
5pынoкmarket379,351
6poccийcкийRussian361,870
7cвoйits359,526
8aкцияstock330,631
9цeнaprice325,888
10нoвыйnew321,402
11pyбльruble310,296
12CШAUSA309,627
13RF307,033
14дeньday263,045
15cтpaнacountry261,620
16млpдbillion260,943
17УкpaинaUkraine247,127
18вpeмяtime243,380
19чeлoвeкperson228,869
20дaтьgive/provide226,986
Table 12. Most frequent bigrams.
Table 12. Most frequent bigrams.
Original BigramBigram in EnglishNumber of Occurrences
12023 гoдyear 2023130,353
22022 гoдyear 202292,893
32024 гoдyear 202485,779
4млpд pyбbillion RUB66,443
5млpд pyбльbillion rubles61,484
6Bлaдимиp ПyтинVladimir Putin40,298
7чиcтый пpибыльnet profit37,892
8бaнк PoccияBank of Russia33,731
9binance txBinance transaction32,004
10ключeвoй cтaвкakey rate30,119
11sender’s balancesender’s balance29,858
12balance txbalance transaction29,858
13PИA нoвocтьRIA Novosti/RIAN28,315
14индeкc мocбиpжaMOEX Index/Moscow Exchange Index27,663
15Coinbase txCoinbase transaction26,726
16млн pyбльmillion rubles24,845
17BC PΦArmed Forces of the Russian Federation (AF RF)24,626
18yгoлoвный дeлocriminal case23,973
19coвeт диpeктopboard of directors22,764
20oбъëм измeнeниevolume change21,666
21измeнeниe дeньdaily change21,639
22цeнa oткpытиeopen price21,593
23oткpытиe cвeчaopening of the candle21,585
Table 13. Message topics obtained by combining small clusters.
Table 13. Message topics obtained by combining small clusters.
PercentTopicKeywords in EnglishOriginal Keywords
26Russia’s economy, market, banking sector, investments, finance, and events in Russia and around the world‘billion’, ‘rubles’, ‘Russia’, ‘USA’, ‘2023’, ‘rub’, ‘million’, ‘companies’, ‘RF’, ‘oil’, ‘CB’, ‘2024’, ‘2022’, ‘market’‘млpд’, ‘pyблeй’, ‘Poccия’, ‘cшa’, ‘2023’, ‘pyб’, ‘млн’, ‘кoмпaнии’, ‘PΦ’, ‘нeφть’, ‘ЦБ’, ‘2024’, ‘2022’, ‘pынoк’
18Russia’s foreign policy and economy. Finance.‘Russia’, ‘RF’, ‘said’, ‘USA’, ‘president’, ‘Putin’, ‘Trump’, ‘Ukraine’, ‘ukrainians’, ‘head’, ‘reported’‘Poccия’, ‘PΦ’, ‘зaявил’, CШA’, ‘пpeзидeнтa’, ‘пyтин’, ‘тpaмпa’, ‘Укpaинa’, ‘Tpaмп’, ‘пpeзидeнт’’глaвa’, ‘cooбщил’
17Situation in Russian regions: UAV strikes, air defense operations, statements by authoritiesregion, UAF, reported, RF, Russia, Defense Ministry,
victims, direction, data, result, governor, air defense
‘oблacти’, ‘BCУ’, ‘cooбщил’, ‘cooбщили’, ‘PΦ’, ‘paйoнe’, ‘Poccия’, ‘минoбopoны’, ‘пocтpaдaвшиx’, ‘нaпpaвлeнии’, ‘дaнным’, ‘peзyльтaтe’, ‘гyбepнaтop’, ‘ПBO’
16Cryptocurrency. Investments‘unknown’, ‘tx’, ‘transfered’, ‘transfered unknown’, ‘usdc’, ‘btc’, ‘000’, ‘coinbase’, ‘binance’, ‘sender’, ‘sender balance’, ‘balance tx’, ‘balance’, ‘unknown coinbase’‘unknown’, ‘tx’, ‘transfered’, ‘transfered unknown’, ‘usdc’, ‘btc’, ‘000’, ‘coinbase’, ‘binance’, ‘sender’, ‘sender balance’, ‘balance tx’, ‘balance’, ‘unknown coinbase’
8News. Investment news. Bonds.More details, bonds, privet Rostov, privet, Rostov, news, offer, rubles, billion, investments, shares, currency, exchange rate‘пoдpoбнee’, ‘oблигaций’, ‘privet rostov’, ‘privet’, ‘rostov’, ‘нoвocть’, ‘пpeдлoжить’, ‘pyб’, ‘млpд’, ‘news’, ‘инвecтиции’, ‘aкции’, ‘вaлютa’, ‘кypc’
4Events in IsraelIsrael, Hamas, Gaza, USA, Gaza Strip, Iran, Strip, Netanyahu, IDF, stated‘Изpaиль’, ‘XAMAC’, ‘Гaзa’, ‘CШA’, ‘ceктope гaзa’, ‘Иpaн’, ‘ceктope’, ‘Heтaньяxy’, ‘ЦAXAЛ’, ‘зaявил’, ‘ceктop Гaзa’
2Telegram-censored publicationsinfringement, due copyright, message couldn, displayed device, device due, couldn displayed, copyright infringement, copyright, displayed, device‘infringement’, ‘due copyright’, ‘message couldn’, ‘displayed device’, ‘device due’, ‘couldn displayed’, ‘copyright infringement’, ‘copyright’, ‘displayed’, ‘device’
2Weather events in Russia.degrees, temperature, Moscow, expected, snow, wind, heat, abnormal precipitation, ISS, weather, forecasters, at night, rain‘гpaдycoв’, ‘тeмпepaтypa’, ‘Mocквa’, ‘oжидaeтcя’, ‘cнeг’, ‘вeтep’, ‘жapa’, ‘aнoмaльный’ ‘ocaдкoв’, ‘мкc’, ‘пoгoдa’, ‘cинoптики’, ‘нoчью’, ‘дoждь’, ‘oжидaютcя’
2China’s economyChina, stocks, USA, alibaba, Xi, People Republic of Chine, dollars, chinese, Hong Kong, chinastocks‘Kитaй’, ‘China’, ‘aкции’, ‘CШA’, ‘alibaba’, ‘Cи’, ‘KHP’, ‘дoллapoв’, ‘китaйцы’, ‘китaйcкиx’, ‘гoнкoнгcкиx’, ‘китaйcкиe’, ‘chinastocks’
1Economy of Georgia, South America, Africa, European UnionGeorgia, France, Maduro, Tbilisi, Venezuela, protesting, president, protesters, Niger, protests, countries, protest, Argentine, police‘Гpyзия’, ‘Φpaнция’, ‘Maдypo’, ‘Tбилиcи’, ‘Beнecyэлa’, ‘пpeзидeнтa’, ‘пpoтecтyющиe’, ‘нигep’, ‘пpoтecты’, ‘cтpaны’, ‘Apгeнтинa’, ‘пoлиция’
1Artificial intelligence, apps, social media, and investmentsAI, Apple, Telegram, iPhone, Google, company, Microsoft, VK, Twitter, Russia, users, intelligence, artificial intelligence, artificial‘ИИ’, ‘Apple’, ‘Telegram’, ‘iPhone’, ‘Google’, ‘кoмпaнии’, ‘Microsoft’, ‘VK’, ‘Twitter’, ‘Poccия’, ‘пoльзoвaтeлeй’, ‘интeллeктa’, ‘иcкyccтвeннoгo интeллeктa’, ‘иcкyccтвeннoгo’
Table 14. Correlation between the volatility of an exchange-traded instrument and the frequency of its mentions in channel messages.
Table 14. Correlation between the volatility of an exchange-traded instrument and the frequency of its mentions in channel messages.
Exchange-Traded InstrumentChannelCorrelation
Gold (futures contract)@CalendarBG0.30
USDRUB@bogdanoffinvest0.311
@scalpon0.302
@marketsnapshot0.301
Bitcoin@whalebotalerts0.567
@futures_Utushkin0.536
@PKNCash0.339
@crypto_nwz0.335
@scalpon0.333
@usertrader30.315
BRENT crude oil (futures contract)@tass_agency0.382
@cbrstocks0.376
@markettwits0.362
@if_market_news0.362
@trade_system0070.322
@marketsnapshot0.321
@macroresearch0.316
@prime10.313
@oil_capital0.307
Dollar index@marketsnapshot0.37
@finchehov0.331
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Luneva, E.; Banokin, P.; Shelupanov, A. TGEconomicDataset: A Collection of Russian-Language Economic Telegram Channels and a Synthetic Data Generation Framework for Continuous Authentication. Data 2026, 11, 25. https://doi.org/10.3390/data11020025

AMA Style

Luneva E, Banokin P, Shelupanov A. TGEconomicDataset: A Collection of Russian-Language Economic Telegram Channels and a Synthetic Data Generation Framework for Continuous Authentication. Data. 2026; 11(2):25. https://doi.org/10.3390/data11020025

Chicago/Turabian Style

Luneva, Elena, Pavel Banokin, and Alexander Shelupanov. 2026. "TGEconomicDataset: A Collection of Russian-Language Economic Telegram Channels and a Synthetic Data Generation Framework for Continuous Authentication" Data 11, no. 2: 25. https://doi.org/10.3390/data11020025

APA Style

Luneva, E., Banokin, P., & Shelupanov, A. (2026). TGEconomicDataset: A Collection of Russian-Language Economic Telegram Channels and a Synthetic Data Generation Framework for Continuous Authentication. Data, 11(2), 25. https://doi.org/10.3390/data11020025

Article Metrics

Back to TopTop