Next Article in Journal
Marginalized Voices, Indigenous Knowledge Systems and Community Media: A Complexity Framework for Community-Engaged Research in Contentious Times
Previous Article in Journal
Examining Small, Medium and Micro-Enterprises’ Contribution to the South African Economy: A Critical Review
Previous Article in Special Issue
Safeguarding the Public Debate from Disinformation: EU Perspectives During the 2024 European Parliament Elections
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Climate-Related Content on Iberian Fact-Checking Platforms: Topics, Entities and Factual Authority in Polígrafo and Maldita

by
Francisco Conrado
1,
Karen Pinto Garzón
2 and
João Pedro Baptista
3,*
1
Communication and Society Research Centre (CECS), Social Sciences Institute, University of Minho, 4710-057 Braga, Portugal
2
Faculty of Information Science, Universidad Complutense Madrid, 28040 Madrid, Spain
3
Centre for Research in Applied Communication, Culture, and New Technologies (CICANT), Department of Letters, Arts and Communication, University of Trás-os-Montes and Alto Douro, 5000-801 Vila Real, Portugal
*
Author to whom correspondence should be addressed.
Soc. Sci. 2026, 15(9), 629; https://doi.org/10.3390/socsci15090629
Submission received: 23 July 2026 / Revised: 11 September 2026 / Accepted: 14 September 2026 / Published: 16 September 2026

Abstract

How is climate-related content produced on fact-checking platforms? This study addresses this question by analysing the climate-related content of two Iberian platforms, Polígrafo (Portugal, n = 26) and Maldita (Spain, n = 277). It combines topic modelling (LDA) and named entity recognition (NER) applied to 303 articles published between 2019 and 2025 with manual coding of the discursive functions of entity mentions. The material includes fact-checks and other editorial genres. The results reveal two contrasting editorial profiles in the analysed corpora that nonetheless share a common epistemic foundation. Both platforms anchor their verdicts and explanations in scientific and meteorological authority, but they diverge in genre and in scale. The Maldita corpus displays an adversarial register devoted to the refutation of climate denialism and to visual fact-checking, with a dense repertoire of named targets, whereas Polígrafo concentrates on a pedagogical mediation of everyday environmental practices, produced largely within an externally funded project. The functional coding also shows that the polarising public figures examined enter verification discourse predominantly as objects of verification, never as sources of authority. The mapping of these differences, which suggest different forms of organising climate-related content on these platforms, opens an avenue for future research on how the impact of verification on disinformation could vary according to the production model that sustains it.

1. Introduction

There has been much discussion about the capacity of fact-checking to keep pace with the constant mutation of climate disinformation (Lewandowsky 2021; Treen et al. 2020). In recent years, the latter has ceased to be confined to explicit denialism and has come to assume subtler forms, whether through conspiracy theories and obstructionist discourses, or by means of manipulated visual content, decontextualised weather maps, greenwashing and multimodal pieces that combine data and images with a technical language that feigns scientific authority (Lamb et al. 2020; O’Neill and Smith 2014; Palau-Sampio et al. 2024; Törnberg and Törnberg 2025). These narratives stem both from campaigns organised by interest groups and political actors and from the dynamics of social media, where the accelerated circulation of content favours decontextualisation and doubt about the scientific consensus (Lewandowsky 2021; Treen et al. 2020; Vu et al. 2023).
In this scenario, verification platforms do more than correct false claims. Insofar as they dismantle disinformative content, they take on the mission of rebuilding the authority of sources before their audiences, since fact-checkers explain how a factual judgement is produced and what evidence allows a public conclusion about truth to be sustained (Graves 2017; Humprecht 2020). The challenge, in the climate case, goes beyond one-off refutation, given that it matters to dismantle narratives that present disinformation as if it were a legitimate controversy, a task for which literacy and preventive inoculation have proved effective (Cook et al. 2017; Lamb et al. 2020; van der Linden 2024).
Despite this role, research on how fact-checkers work on themes such as the climate crisis remains scarce, and comparative research across national contexts is scarcer still (Humprecht 2020; Vu et al. 2023). This gap is particularly sensitive in the Iberian context. Spain has a growing body of work on climate verification, especially around Maldito Clima, EFE Verifica and Newtral (Fernández-Castrillo and Magallón-Rosa 2023; Sendra-Duro 2025), whereas Portugal, even though it has produced studies on political fact-checking and media literacy (Baptista et al. 2022; Oliveira et al. 2024), remains unstudied with regard to climate verification.
The present study seeks to address this gap through a comparative analysis of 303 pieces published between 2019 and 2025 by Maldita (Spain) and Polígrafo (Portugal), with recourse to topic modelling (Latent Dirichlet Allocation, LDA) and named entity recognition (NER), complemented by a manual functional coding of entity mentions. We deliberately place ourselves on the production side, attentive to the platforms’ climate-related content, and we leave for future research its effects on audiences. This strategy allows two complementary dimensions to be analysed, namely which thematic fields structure climate-related content and which people, organisations and locations sustain its authority, and it adds a third, namely what role those entities actually perform in the discourse, whether as sources that anchor verdicts or explanations, as targets of the claims verified, or as mere context. Supported by the epistemologies of digital journalism applied to fact-checking (Ekström and Westlund 2019; Ekström et al. 2020; Graves 2017), we resort to two descriptive lenses that organise this mapping, the referential ecologies of verification and the regimes of factual legitimation, and we understand verification as a practice of epistemic mediation that selects claims for examination and makes explicit the sources and evidence on which its assessments are based. Distinguishing frequently mentioned actors from actual sources can inform newsroom assessments of source use and help future audience studies compare explanatory and corrective genres.

2. Theoretical Framework

2.1. Fact-Checking and Public Knowledge About the Climate Crisis

The media are crucial for the public understanding of the climate crisis (van der Linden 2022). Communicating a global and politically disputed phenomenon with cumulative effects demands translating an abstract, long-term problem into experiences recognisable to audiences, such as heatwaves, wildfires, droughts, energy consumption or public policies (O’Neill and Smith 2014; Vu et al. 2023). As verification initiatives consolidated internationally, especially since the expansion of professional networks such as the International Fact-Checking Network, fact-checking platforms became first-line actors in the fight against disinformation, since they explain where a claim comes from and what kind of evidence allows it to be accepted or rejected (Graves 2017; Humprecht 2020; Vu et al. 2023).
This verification work becomes more complex in the climate case, since it rarely deals with completely false content. What is frequently at stake are ambiguous claims, partial data, manipulated images and interested interpretations of meteorological phenomena (Fernández-Castrillo and Magallón-Rosa 2023; O’Neill and Smith 2014; Törnberg and Törnberg 2025). The effectiveness of climate fact-checking depends, for this reason, less on the debunk itself than on the way the process is explained and the sources are made transparent to the audience (Humprecht 2020; Vu et al. 2023). And the work does not end with publication, given that verification content itself re-enters the circuit of social media, where audiences will accept it or reinterpret it in their own manner (Graves 2017).
As Ekström and Westlund (2019) and Ekström et al. (2020) note, studies on digital journalism and on disinformation should dialogue with each other, since news journalism, data journalism, science communication and fact-checking take part in one and the same dispute over the authority of public knowledge. The epistemologies of digital journalism are plural, and fact-checking can be understood as a specific form of journalistic knowledge production, because it intervenes in controversies in which what is disputed, alongside the facts, are the very criteria that determine what counts as a valid fact in the public sphere (Ekström et al. 2020; Graves 2017). The present study builds on this theoretical tradition.
To operationalise this perspective, we make use of two descriptive lenses. By referential ecology of verification we understand the patterned set of people, organisations and territories that a platform recurrently mobilises to anchor its verifications, that is, the infrastructure of references upon which verdicts rest. The notion takes up what the classical literature on sources described piece by piece and shifts the observation to the aggregate anchoring pattern that emerges from the corpus collected for a platform, measurable by computational means. By regime of factual legitimation we understand, in turn, the dominant principle of authority that organises that ecology, that is, the type of instance, whether scientific, political, technical or observational, to which the platform resorts in order to justify its assessment of a disputed claim through publicly inspectable evidence (Humprecht 2020; Vu et al. 2023). An important caveat follows from these definitions. Frequency of mention does not, by itself, establish a legitimating function, since an entity can appear in a verification as the authority that sustains the verdict, but also as the author or protagonist of the claim under scrutiny, or as mere background. For this reason, the computational mapping is complemented in this study by a manual coding of the function that each mentioned entity performs, without which the passage from mention counts to regimes of legitimation would remain an inference. Far from intending to replace the epistemology of fact-checking described by Graves (2017), which explains how verifiers construct factual judgements at the level of professional routines, these lenses shift the analysis to the structural level and allow platforms to be compared by the systematic pattern of the authorities they mobilise. For example, an IPCC report used as evidence functions as a source, Greta Thunberg as the subject of a debunked rumour as an object, and a European Union funding acknowledgement as context.

2.2. Climate Disinformation, Sources and Politicisation

The politicisation of the climate crisis involves an intricate web of disinformation campaigns oriented towards undermining the scientific consensus and delaying climate action, with the legitimacy of specialised institutions as a recurrent target (Lamb et al. 2020; Treen et al. 2020). Climate disinformation, which rarely stops at denying the phenomenon, displaces responsibilities and exaggerates the costs of measures, and may also discredit renewables or present climate policies as ideological impositions (Lamb et al. 2020; Palau-Sampio et al. 2024). A particularly worrying aspect of this dynamic is the amplification of sceptical positions which, although minoritarian in the face of the scientific consensus, acquire a disproportionate presence in certain communicative spaces, distorting public debate and potentially affecting citizens’ decisions (Palau-Sampio et al. 2024; Treen et al. 2020). This encompasses both old denialism, which rejects warming or its human causes, and newer forms that discredit climate science or the viability of solutions (Coan et al. 2021). False-balance coverage can reinforce doubt by giving contrarian views equal prominence to scientific evidence (Cook et al. 2017).
The actors who produce or amplify this content form a hybrid ecosystem. Vu et al. (2023) find, among the identifiable actors making the climate claims examined, a predominance of politicians, accompanied by businesspeople, organisations, scientists and activists. Palau-Sampio et al. (2024) identify social media, pseudo-media and party-political actors as the main sources of climate disinformation, whereas Fernández-Castrillo and Magallón-Rosa (2023) show, for the Spanish case, that many of the hoaxes verified by Maldito Clima circulate through messaging platforms such as WhatsApp, especially during episodes of high climate visibility such as heatwaves. Törnberg and Törnberg (2025), for their part, document the role of blogs and alternative media linked to the climate countermovement in the diffusion of denialist and delay narratives.
Corrective sources constitute, within this frame, a pillar of verification. The literature shows that climate fact-checkers rely mostly on public authorities and on specialised scientific sources, the category that Vu et al. (2023) identify as the most used in the climate verifications of the United States, the United Kingdom, Germany and Australia. Humprecht (2020) adds that the transparency of sources is a key element of fact-checking credibility, even though it depends on structural factors such as the country, the level of journalistic professionalisation, trust in the media and the national journalistic culture. The themes that are prioritised, much like the sources that are consulted, are ultimately not neutral, given that they show what each platform understands as a verifiable climate problem and what type of authority it mobilises to correct it (Palau-Sampio et al. 2024).

2.3. The Iberian Case and Research Questions

Climate fact-checking reflects the socio-political disputes of each country (Vu et al. 2023), since it is produced from distinct national contexts, each with its own journalistic culture and its own disinformation ecosystem (Humprecht 2020; Palau-Sampio et al. 2024). In the Spanish case, Fernández-Castrillo and Magallón-Rosa (2023) show the relevance of narratives linked to extreme meteorological phenomena, weather maps, viral images and obstructionist discourses, and Sendra-Duro (2025), when comparing EFE Verifica, Maldito Clima and Newtral, identifies a strong national territorial focus and a prominent use of official and specialised sources. Considering that Maldita generates the largest volume of climate verification among Spanish platforms, it can be argued that its model approximates a scientific-observational and visual-forensic approach to verification, whose factual authority is built by combining scientific sources and visual evidence with a strong pedagogical contextualisation (Fernández-Castrillo and Magallón-Rosa 2023; Sendra-Duro 2025).
In Portugal, by contrast, specific studies on climate fact-checking are scarce. There are, even so, contributions on electoral verification, with Baptista et al. (2022) showing that, during the legislative elections, Observador and Polígrafo centred mostly on political statements, and on the mapping of national initiatives against disinformation, in which Oliveira et al. (2024) situate Polígrafo. Previous literature does not offer, however, a map of the climate authorities mobilised by Polígrafo, which makes it relevant to study empirically which institutions and sources sustain its factual authority. From this review, we derive the following research questions.
RQ1. Which thematic fields structure the climate-related coverage of Maldita and Polígrafo between 2019 and 2025?
RQ2. Which people, organisations and locations do Maldita and Polígrafo mobilise to sustain the factual authority of their climate-related content?
RQ3. How do topics and named entities relate to one another in the construction of referential ecologies of verification and regimes of factual legitimation?

3. Materials and Methods

The material analysed in this article consists of climate-related content, including fact-checks and other editorial genres, published by two Iberian platforms, namely Polígrafo (Portugal) and Maldita (Spain). Collection was carried out in June 2025 using keyword searches followed by structured web scraping in R (version 4.5.1). The search expressions were «crise climática», «alterações climáticas» and «emergência climática», in the Portuguese case, and «cambio climático», «crisis climática» and «calentamiento global», in the Spanish case. The Polígrafo search returned two available listing pages, both in the Ambiente section; the Maldita material was located in Maldito Clima. The terms cover the generic designation of climate change and its crisis or warming framings; they do not ensure retrieval of every relevant article, particularly texts using only event-specific vocabulary.
After collection, the Polígrafo corpus was composed of 27 articles published between 2019 and 2024. One of the articles was excluded for containing no textual body for analysis, which reduced the total to 26 documents. In the case of Maldita, 398 records were collected, corresponding to 277 distinct article URLs published between 2019 and 2025; duplicate URLs were removed. The difference in scale between the two corpora (1:11) constitutes in itself a relevant analytical datum, one that describes the material retrieved rather than the total editorial attention devoted to climate issues in each country. This asymmetry was treated as a variable of interest, even though its methodological implications are also acknowledged. The complete article list is provided in Supplementary Materials. Table 1 summarises the characteristics of the two corpora.
The final analysis was conducted in R version 4.5.2. For the purposes of this analysis, we combined two complementary natural language processing techniques with a manual coding stage. Topic modelling via Latent Dirichlet Allocation (LDA) was used to map the latent thematic structure of the articles, that is, what was discussed in the texts. To identify the actors and references mobilised, that is, in which actors and places argumentation was anchored, we resorted to Named Entity Recognition (NER). This double approach is consistent with studies that apply topic modelling to journalistic corpora (Ahmed et al. 2022; Lee 2022; Valdez et al. 2018; Zhang et al. 2013). Because NER registers mentions but not the role those mentions play, a third analytical stage involved functional coding of the extracted entities, described below, which grounds the interpretation of mention frequencies in their discursive roles.

3.1. Textual Pre-Processing

The texts were subjected to a boilerplate cleaning process, that is, the removal of discursive units that appear recurrently in the texts without semantic value for the analysis. In the case of Polígrafo, a funding disclaimer (European Media and Information Fund/Calouste Gulbenkian Foundation) present in the body of 19 of the articles was identified and systematically removed. A second, distinct co-funding disclaimer (European Union/European Commission), present in a single institutional article, was identified at a later stage, during the functional coding, and its implications for the raw entity counts are addressed in the results. In the case of Maldita, the recurrent editorial calls to action (for example, invitations to follow the platform’s WhatsApp channel) were neutralised for topic modelling through the removal of self-references in the stopword list, and their mentions were classified as boilerplate in the functional coding.
In addition, the text of both corpora was converted to lowercase, non-alphabetic characters were eliminated and word-level tokenisation was applied (tidytext::unnest_tokens; tidytext version 0.4.3). Stopwords were removed, along with additional functional terms and editorial self-references (e.g., «Polígrafo», «FactCRICIS», «Maldita», «Maldita Ciencia»), and only tokens with three or more alphabetic characters were retained. The document-term matrix was built with the tidytext package (Silge and Robinson 2017).
We opted for tokenisation without lemmatisation in order to preserve the diversity of forms as they occur in journalistic discourse and to respect the polysemy of words (Brookes and McEnery 2019). The plural form «emissões» (emissions), for example, occurs in discursive contexts different from those of «emissão» (emission). This choice can fragment related forms; the UMass coherence values (Table 2) describe word co-occurrence but do not rule out sensitivity to alternative pre-processing.

3.2. Topic Modelling (LDA)

The present analysis adopts an exploratory and inductive approach, in which thematic categories emerge from the data without recourse to pre-defined coding schemes. LDA is particularly suited to this logic, since it allows latent thematic patterns to be identified in a corpus without supervision, which makes it a recurrent tool in the analysis of journalistic texts (Blei et al. 2003; Zhang et al. 2013). The final models used variational expectation-maximisation, the default method in the topicmodels package version 0.2-17 (Grün and Hornik 2011), with a fixed random seed (seed = 123). The selection of the number of topics (k) was based on a combination of statistical and interpretability criteria, with recourse to a preliminary Gibbs-based search using the ldatuning package version 1.0.3 (Murzintcev 2024), which computes four fit metrics for values of k between 2 and 8.
For both corpora k = 4 was adopted, informed by a sensitivity analysis with k = 3, 4 and 5. In Maldita, temperature-related vocabulary recurs across the tested solutions, while in Polígrafo recycling-related vocabulary recurs but the remaining topics present greater variability, consistent with the limitations of a small corpus (Agrawal et al. 2018). The adoption of k = 4 is thus grounded in the interpretability of the resulting solution, rather than the stability of every topic. Using the same k facilitates presentation but does not make the separately estimated topics equivalent across platforms.
Given the sensitivity of the Polígrafo model, we further computed UMass topic coherence (Mimno et al. 2011) as an indicator of the semantic quality of the extracted topics. The score measures the co-occurrence of each topic’s representative words in the documents: for each word pair, the log ratio uses the number of documents containing both words plus one, divided by the number containing the higher-ranked word. Table 2 reports averages across pairs for the top-10 and top-20 words of the final topics. Higher scores indicate stronger co-occurrence under this implementation, but there is no universal acceptance threshold. Given the small Polígrafo corpus, its LDA results are strictly exploratory and should be read alongside the manual analysis, not as a stable thematic classification.
The LDA results were extracted in the form of a document-topic distribution (θ matrix, γ values), which indicates the proportion of each topic in each document, and of a topic-word distribution (β matrix), which identifies the most representative words of each topic. The attribution of the dominant topic was based on the maximum γ value for each document. Topics were interpreted and labelled on the basis of the 20 words with the highest β values and of the titles of the articles with the highest γ values, in line with the practice recommended by Lee (2022) and Valdez et al. (2018). Labelling was carried out independently by two researchers and divergences were reconciled through discussion until consensus was reached. Results were visualised through LDAvis version 0.3.2 (Sievert and Shirley 2014).

3.3. Named Entity Recognition (NER)

Named entity recognition was used to map the actors and geographic references mobilised in the discourse, which allowed us to identify the people, organisations and locations that structure the argumentation about the climate crisis. As a technique, NER is well established in the literature and allows the automatic extraction of mentions in unstructured text (Ehrmann et al. 2016; Nadeau and Sekine 2007). Extraction was carried out with spaCy version 3.8.11 via reticulate version 1.45.0, using the pre-trained models pt_core_news_lg version 3.8.0 for the Portuguese corpus and es_core_news_lg version 3.8.0 for the Spanish corpus. Post-processing removed identified false positives and editorial self-references, corrected selected entity types (for example, Copernicus to ORG), and combined specified personal-name variants, such as Trump/Donald Trump and Thunberg/Greta Thunberg. Counts refer to retained mentions, not to the number of articles. This normalisation does not resolve every institutional-name variant.

3.4. Functional Coding of Entity Mentions

Because a mention count cannot distinguish an authority that sustains a verdict from the author of the claim under scrutiny, the ORG and PER entities extracted by NER were subjected to manual functional coding. Three mutually exclusive functions were defined. An entity was coded as a source when it anchors the verdict or the explanation, whether by being quoted, by providing data or by being invoked as an authority. It was coded as an object when it is the author, target or protagonist of the claim being verified. It was coded as context when the mention is referential, without an epistemic role in the piece, as happens with institutional affiliations, background references and funding disclaimers. LOC entities were excluded from the functional coding, given that locations rarely perform a legitimating function. A single researcher performed the functional coding; this stage is distinct from the two-researcher interpretation of LDA topic labels.
The Polígrafo corpus was examined manually in full, which its small size made feasible; functional coding records cover 25 of its 26 articles, as the remaining article had no retained ORG or PER entities eligible for this stage. For the larger Maldita corpus, exhaustive manual coding was less practical; LDA and NER mapped all articles, while functional interpretation rested on a manually coded sample. For the latter, the coded random stratum comprised 35 articles selected through stratified random sampling proportional to the yearly distribution of the corpus (fixed seed = 123), a design whose balance against the full corpus was verified through keyword proxies of the four LDA topics. To this random sample was added a purposive complement of the 11 articles that mention Greta Thunberg or Donald Trump, given the centrality of these figures to the analysis of personalisation and their rarity in a random draw. The purposive complement is excluded from the aggregate proportions and reported separately in Section 4.3; the two strata are not pooled. Coding proceeded at the level of the entity–article pair, with the dominant function and the number of mentions registered for each pair, and ambiguous cases resolved through re-reading of the full text. Functional percentages weight each entity–article pair by its recorded mention count. A single dominant label does not retain all secondary roles when an entity performs more than one function in an article. For example, the PSD was coded as object in the Parque das Nações fact-check even though it also supplied documentation. A targeted reconciliation of the original European Union and Zero records against archived cleaned article bodies retained the entity–article unit and dominant functions. It confirmed 15 long-form European Union mentions and corrected Zero’s POL018 and POL023 counts to 3 and 5. These functional counts differ from retained NER strings; stand-alone «UE» abbreviations are outside the reported European Union profile.

3.5. LDA × NER Crossing

The LDA and NER results were combined at the document level, by associating with each text its dominant topic and the respective extracted entities. This operation allowed the referential profile of each topic to be characterised by examining whether certain themes are anchored predominantly in institutional actors (ORG), geographic references (LOC) or individual figures (PER). For each topic, the absolute counts and the relative proportions of each entity type were computed, in line with the analysis proposed by Heidenreich et al. (2019) for the study of media framing.

4. Results

4.1. Thematic Structure (LDA)

The application of LDA with k = 4 to both corpora produced exploratory thematic structures. To assess the co-occurrence of their representative vocabulary, we computed the UMass topic coherence for the top-10 and top-20 words per topic. Figure 1 presents the fit metrics used to inform the selection of the number of topics.
The temporal distribution of the articles reveals, first of all, markedly distinct production patterns. In Polígrafo, 22 of the 26 articles were published in 2022, most of them concentrated between June and August, which indicates that the corpus essentially reflects the editorial agenda of a specific period. In addition, 19 of these articles carry the disclaimer of the same externally funded project (European Media and Information Fund/Calouste Gulbenkian Foundation), a circumstance to which we return in the discussion. Maldita’s production is distributed in a more balanced way, with three articles in 2019, two in 2020, 23 in 2021, 92 in 2022, 88 in 2023 and 61 in 2024, to which eight articles in the first half of 2025 are added (collection truncation). These counts show coverage across several years, with a peak in 2022, coinciding with the launch of the dedicated section Maldito Clima in April of that year.
The LDA model identified four topics in the Polígrafo corpus, with a distribution varying between 19.7% and 33.1% (Table 3). Topic 1 (19.7%, 5 documents) is organised around the global climate crisis, with the keywords «emissões» (emissions), «países» (countries), «mundo» (world), «temperatura» (temperature) and «brasil» (Brazil). Representative articles include pieces on zero-emission targets, banned pesticides and COP28, forming an axis of international climate policy with a strong geopolitical dimension. Topic 2 (33.1%, 9 documents) is the largest topic in the corpus and aggregates texts on recycling, sustainable consumption and European green policies, with the words «água» (water), «verde» (green), «ecoponto» (recycling point), «embalagens» (packaging) and «calor» (heat). The articles include pieces on waste separation, the environmental footprint of events, heatwaves and water saving. This axis was classified as pedagogical mediation, since these articles function here as translators of everyday environmental practices for the citizen. Topic 3 (26.9%, 7 documents) centres on health, oceans and the impacts of climate change, with the terms «amianto» (asbestos), «saúde» (health), «alterações» (changes), «mar» (sea) and «climáticas» (climatic). Topic 4 (20.3%, 5 documents) focuses on the contamination of soils and water in an urban context, with the words «água» (water), «terrenos» (plots of land), «solos» (soils), «lisboa» (Lisbon) and «contaminados» (contaminated), which includes articles on contaminated plots in Parque das Nações, greenwashing and floods.
The Maldita corpus presents, for its part, a thematic structure with one clearly dominant topic. Topic 1 (37.6%, 104 documents) absorbs more than a third of the corpus and is organised around climate denialism and the scientific evidence on temperature, with the words «temperatura» (temperature), «cambio» (change), «climático» (climatic), «calor» (heat) and «calentamiento» (warming). This is the most distinctive finding of the Spanish corpus. Maldita devotes a substantial proportion of its production to the dismantling of denialist arguments, functioning as a discursive trench against climate disinformation. Topic 2 (19.1%, 49 documents) addresses climate policy and the regulation of emissions. Topic 3 (21.7%, 61 documents) constitutes an axis of visual fact-checking, with the words «mapa» (map), «mapas» (maps), «bulo» (hoax), «imagen» (image) and «tiempo» (weather), centred on the dismantling of manipulated maps and decontextualised images. Topic 4 (21.6%, 63 documents) is organised around extreme meteorological phenomena, with the terms «agua» (water), «incendios» (wildfires), «DANA», «lluvias» (rains) and «clima» (climate).
Figure 2 presents the intertopic maps for both corpora. LDAvis renumbers the topics by token-weighted prevalence: circles 1–4 correspond to Table 3 topics 1, 3, 4 and 2 for Polígrafo, and 1, 4, 2 and 3 for Maldita. The distance between circles provides a visual indication of lexical differences, not an independent test of topic stability.

4.2. Entity Ecology (NER)

In Polígrafo, organisations constitute the most represented category (45.1%), followed closely by locations (44.0%) and, at a distance, people (10.9%). In Maldita, locations have the largest share (47.9%), followed by organisations (35.4%) and people (16.7%).
Cramér’s V (0.059) indicates a negligible difference in the pooled distribution of entity types. This is a descriptive measure: repeated mentions within an article are not independent observations and longer texts can contribute more mentions. The two platforms operate, in fact, within a shared referential geometry. Both anchor climate coverage predominantly in organisations and locations, with a lesser mobilisation of individual figures. The differences reside in the identity and, as the functional coding will show, in the role of the mobilised entities, more than in the proportion of types.
The most frequent entities (Table 4) reveal contrasting referential profiles. On the institutional plane (ORG), Polígrafo’s raw counts are led by European and supranational bodies, such as the European Union (14), the European Commission (11) and the European Parliament (6), complemented by national environmental NGOs such as Zero (13) and by the waste-management entity Sociedade Ponto Verde (11). Maldita, in contrast, privileges entities of a scientific and meteorological nature. AEMET (275) and the IPCC (215) are, by far, the most cited entities, followed by NASA (101). On the geographic plane (LOC), Polígrafo presents a globalised perspective, with Brazil as the most frequently mentioned location (27 mentions), followed by the Earth (10), China (9) and the USA (8). Maldita centres strongly on Spain (401), followed by the United States (147) and Europe (131), a predominance that indicates a strong national territorial anchoring. As for personal entities (PER), in Polígrafo the most cited figures are exclusively national experts and technicians. In Maldita, polarising international public figures dominate, such as Greta Thunberg (47) and Donald Trump (33), alongside spokespersons of meteorological institutions, such as Rubén del Campo (45, of AEMET). Whether these raw counts translate into regimes of legitimation is, however, a question that only the functional coding can answer.
Table 4. Most frequent entities by type and platform (top 3).
Table 4. Most frequent entities by type and platform (top 3).
TypePolígrafonMalditan
ORGEuropean Union14AEMET275
Zero13IPCC215
European Commission11NASA101
LOCBrazil27Spain401
Earth10United States147
China9Europe131
PERSusana Fonseca9Greta Thunberg47
J. Poças Martins5Rubén del Campo45
Marta Leandro5Donald Trump33
Note. Absolute retained NER-string counts (n). The European Union and Zero functional denominators in Table 5 also include occurrences within longer expressions. For the percentage distribution by type, cf. Figure 3.
Figure 3. Global distribution of named entities by type and platform. Cramér’s V = 0.059 (descriptive measure).
Figure 3. Global distribution of named entities by type and platform. Cramér’s V = 0.059 (descriptive measure).
Socsci 15 00629 g003

4.3. Functional Profile of Entity Mentions

Table 5 summarises the functional profiles of selected entities by platform and sampling stratum. The functional coding substantially reconfigures the reading of Table 4, in different directions for each platform. In Polígrafo, the apparent primacy of European political institutions dissolves under functional scrutiny. The 15 coded long-form mentions of the European Union comprise one source, two objects and twelve contextual mentions. The latter concern funding disclaimers and regulatory framing (substances «banned in the EU»), while the two object mentions occur in an explainer on climate legislation. Four of the European Parliament’s six mentions come from an explainer about the «Fit for 55» package, in which the institution is the topic of the piece rather than the authority sustaining a verdict. The remaining two were coded as a source in an opinion article by an external expert, where the Parliament is invoked as a normative authority. The European Commission also performs a genuine, though minority, source function (four of 11 mentions, in pieces that cite its studies and documents as evidence).
Table 5. Functional profile of the most frequent entities (share of mentions coded as source/object/context).
Table 5. Functional profile of the most frequent entities (share of mentions coded as source/object/context).
PlatformEntityTypeSourceObjectContext
PolígrafoEuropean Union (n = 15)ORG7%13%80%
European CommissionORG36%36%27%
European ParliamentORG33%67%0%
Zero (n = 14)ORG100%0%0%
Sociedade Ponto VerdeORG100%0%0%
IPCCORG100%0%0%
Susana FonsecaPER100%0%0%
Maldita (random)AEMETORG100%0%0%
IPCCORG93%7%0%
NASAORG92%0%8%
Rubén del CampoPER100%0%0%
Maldita (purposive)Greta ThunbergPER0%100%0%
Donald TrumpPER0%87%13%
Note. Polígrafo values use the functional records for 25 articles. The European Union profile includes the long form «União Europeia», also within longer expressions, but not stand-alone «UE»; the Zero profile excludes the distinct project Do Zero. The two denominators shown were checked against archived cleaned article bodies. Maldita rows are labelled by sampling stratum and are not pooled. The European Parliament’s two source-function mentions invoke normative authority in an opinion article. These selected entity profiles are descriptive illustrations, not population estimates. Percentages rounded.
Once these mentions are set aside, the entities that effectively anchor Polígrafo’s verdicts and explanations are of a scientific, technical and sectoral nature. Counted by source-function mentions, the leading organisations are the environmental NGO Zero (14), the IPCC (13), the UN and its agencies (12), the WHO (11), Sociedade Ponto Verde (11), the Portuguese health authority DGS (8) and the meteorological institute IPMA (5). The three most cited persons (Susana Fonseca, Joaquim Poças Martins, Marta Leandro) perform a source function in 100% of their mentions. The divergence between the raw ranking and the functional ranking is partly attributable to the mechanics of NER itself, since «União Europeia» is a relatively stable name, whereas the IPCC appears under multiple variants («Painel Intergovernamental sobre Mudanças Climáticas», «Painel Intergovernamental para as Alterações Climáticas», «IPCC») that can fragment its count or escape automatic recognition. Scientific authority is in any case dispersed across many distinct institutions while European references concentrate on three.
In Maldita, by contrast, the functional coding is consistent with the raw ranking. In the random stratum, AEMET performs a source function in 53 of 53 mentions, the IPCC in 50 of 54 (the residual mentions belong to an explainer in which the IPCC report is itself the topic) and NASA in 23 of 25. The scientific-observational interpretation of the platform is, therefore, supported by the coded material, and it is complemented by a repertoire of named objects with no equivalent in the Portuguese corpus, composed of politicians (Carlos Mazón, Alberto Garzón, Donald Trump, Republican members of the US Congress), authors of contested claims (Valentina Zharkova, José Ramón Ferrandis), disinformation-carrying outlets (El Toro TV, Mount Vernon News, Alerta Digital), legacy media whose old pieces are corrected (El Mundo, ABC, El Español) and fossil-fuel interest organisations. In the random stratum, the global distribution of coded mentions is 75.8% source, 15.2% object and 9.0% context; in the Polígrafo corpus the corresponding distribution is 69.1%, 14.4% and 16.5% (260, 54 and 62 of 376 coded mentions), with the higher context share driven precisely by disclaimers and institutional pieces.
The purposive stratum clarifies the status of the polarising figures. Greta Thunberg performs an object function in 100% of her coded mentions (35 of 35), invariably as the target of hoaxes that Maldita dismantles in her defence, and never as a source of authority; Donald Trump is an object in 20 of 23 mentions, as the author of claims verified one by one, with the remaining three coded as context. Personalisation in Maldita is, therefore, a personalisation of targets and not of sources, and it is concentrated, in the full corpus, in around a dozen pieces. A final function surfaced by the coding, common to both platforms, deserves mention: the defence of scientific authorities against disinformative appropriation, visible when Maldita corrects distorted readings of NASA or IPCC studies and when Polígrafo restores the context of misused scientific claims, a corrective operation in which the authority appears simultaneously as a misattributed source in the hoax and as a genuine source in the verification.

4.4. LDA × NER Crossing

The crossing between topics and entities in Polígrafo (Figure 4) reveals strongly differentiated anchoring profiles. Topic 1 (global climate crisis) is overwhelmingly dominated by LOC (65.4%), with a residual presence of PER (3.0%). The near absence of PER in this topic indicates that the global climate crisis is discussed in Polígrafo without personalisation, in terms of systems and territories. Topics 2 and 3 present the inverse pattern, with ORG dominating (62.3% and 57.1%). Topic 4 (urban contamination) is the only one with a balance between ORG and LOC (39.1% each) and presents the highest weight of PER (21.8%), which reflects articles on concrete cases with identifiable participants.
In Maldita (Figure 5), LOC is the dominant category in all four topics (42.3–49.7%), which reflects a strong territorialisation of climate discourse in the Maldita corpus. The most relevant difference resides in the weight of PER. Topic 3 (visual fact-checking) has a PER share of 26.6%, the highest value of the whole study, consistent with the nature of visual hoaxes involving public figures. Topic 2 (climate policy) presents the highest weight of ORG (41.0%), as expected given the institutional framing. Topic 1 (denialism), despite being the largest of the corpus, is anchored more in data and institutions than in individual figures (PER = 14.3%).

5. Discussion

The present study comparatively analysed climate-related coverage on two Iberian fact-checking platforms by combining LDA, NER and a functional coding of entity mentions, with a focus on the production side, that is, on what the platforms actually publish. The results reveal contrasting editorial profiles in the analysed corpora and show that their climate-related coverage goes beyond the piecemeal correction of false claims, operating as a form of epistemic mediation that selects the problems worthy of verification and hierarchises the authorities called upon to resolve them in a disputed informational environment (Ekström and Westlund 2019; Ekström et al. 2020). As Graves (2017) argues, the factual judgements of fact-checking are constructed through professional routines and contextual interpretation, with nothing mechanical about their objectivity, and it is this construction that our data allow to be observed at scale.
In response to RQ1, the data suggest that Polígrafo orients itself predominantly towards a register of pedagogical mediation. The most prominent topic of the corpus (T2, 33.1%) centres on the translation of everyday environmental practices for the citizen, an orientation consistent with the model of fact-checking as an educational public service, in which the platform assumes the function of mediator between technical-scientific knowledge and readers’ practices (Graves 2016). This orientation must, however, be read against the material conditions of production. The Portuguese corpus is concentrated in a single period (22 of 26 articles in 2022, most between June and August) and 19 of its pieces belong to a single externally funded project (EMIF/Calouste Gulbenkian Foundation). The pedagogical register of the collected Polígrafo texts may therefore be partly associated with that funded initiative, rather than a settled editorial line; the concentration does not establish what the platform would have published without external funding. The corpus is, moreover, heterogeneous in genre: a close reading distinguishes only six verification pieces in the strict sense, alongside fourteen service explainers, three republications of the Brazilian outlet Agência Pública, an opinion article by an external expert, an institutional press release and a short report. In Maldita, more than a third of the corpus (T1, 37.6%) centres on the dismantling of denialist narratives. This finding aligns with previous studies on the Spanish case, in particular with Fernández-Castrillo and Magallón-Rosa (2023), who showed the role of Maldito Clima during the 2022 heatwave in the verification of narratives such as the supposed chromatic manipulation of weather maps. Sendra-Duro (2025) also identifies in Spanish verification a strong focus on extreme phenomena and viral hoaxes. These patterns can be read in relation to the socio-political contexts discussed by Vu et al. (2023), without treating the two platforms as representative of their countries.
This difference in orientation is reinforced by the presence, in Maldita, of an autonomous topic devoted to visual fact-checking (T3, 21.7%), centred on the dismantling of manipulated maps and decontextualised images. The vocabulary of this topic points to the importance of visual formats in the disinformation addressed by the Maldita corpus. This finding is consistent with the literature on visual communication and climate disinformation, since, as O’Neill and Smith (2014) show, climate images produce meanings and perceptions of truth of their own, instead of serving as simple illustration. Contemporary disinformation rests on an aesthetic of apparent scientific rationality, made of graphs and maps that mimic the visual language of objectivity (Törnberg and Törnberg 2025). In the collected Polígrafo texts, this type of visual verification was not identified, an absence that could be related to the context of creation of each platform and to the media ecosystem of each country. Maldita was born as a project to combat viral hoaxes (Vizoso and Vázquez-Herrero 2019), whereas Polígrafo has a tradition closer to political-parliamentary fact-checking.
As for RQ2 and RQ3, the functional coding imposes a substantial revision of the interpretation that the raw entity counts would invite. Read at the level of mentions, the two platforms seem to embody opposed regimes of legitimation, with Polígrafo anchored in European political authority and Maldita in scientific authority. Read at the level of functions, the opposition weakens in terms of the authority invoked. The European institutions that top Polígrafo’s raw ranking almost never sustain factual assessments: they appear as funders in disclaimers, as the regulatory frame within which claims are situated, or as the very subject matter of pedagogical explainers. Part of their prominence is an artefact of NER mechanics, which can miss or inconsistently classify the multiple variants under which the IPCC surfaces in Portuguese text. When the anchoring function is isolated, Polígrafo’s verdicts and explanations rest on the same type of authority as Maldita’s, namely scientific, meteorological and health institutions (IPCC, WHO, DGS, IPMA), complemented by a national technical layer of environmental NGOs, waste-management entities and individual experts who are sources in the totality of their mentions. The coded material suggests that the two platforms draw on similar kinds of authority.
A visible difference in the analysed material concerns the genre of climate-related content and the object layer that comes with it. Maldita practises adversarial fact-checking populated by named objects, from politicians and hoax authors to disinformation-carrying outlets, legacy media and fossil-fuel interest groups, a repertoire that has almost no equivalent in the Portuguese corpus, where the adversarial function is confined to a handful of pieces. The notion of referential ecology of verification gains precision with this result: the ecologies of the two platforms differ less in the authorities they invoke than in the presence or absence of a structured layer of adversaries against which those authorities are deployed. European references remain distinctive in the Portuguese mention profile, but as a topic and regulatory horizon rather than as a legitimating instance, which is itself coherent with the dependence of Portuguese environmental news on a European regulatory framing (Schmidt and Delicado 2014) and with the tradition of Portuguese verification journalism of anchoring itself in institutional sources as guarantors of credibility (Jerónimo and Sánchez Esparza 2023).
The dimension of personalisation constitutes the third axis of differentiation, and here too the functional coding refines the initial reading. Polígrafo has the lowest proportion of personal entities in the whole study (10.9% overall), mobilising almost exclusively national technical experts, whose three most frequently mentioned representatives function as sources in 100% of their coded mentions. Maldita, with people accounting for 16.7% of mentions, integrates into its discourse polarising international public figures, whose coded appearances are predominantly objects: Thunberg appears solely as the target of hoaxes that the platform dismantles, Trump primarily as the author of claims verified one by one, with some contextual mentions, and both are concentrated in around a dozen pieces of the corpus. The peak of PER in the visual fact-checking topic (26.6%) suggests that the visual hoaxes addressed in this corpus often involve public figures, a difference interpretable in the light of the literature on the personalisation of media discourse (Van Aelst et al. 2012), with the qualification that this is a personalisation of targets, not of sources. The coding also rendered visible a corrective operation shared by both platforms and rarely named in the literature, the defence of scientific authorities against disinformative appropriation, in which an institution such as NASA figures at once as the falsely invoked source of the hoax and as the genuine source of the correction.
We also find it important to address the difference in scale between the two corpora (26 vs. 277 articles). This asymmetry is simultaneously a relevant empirical datum, one that describes the scale of the retrieved material, and a methodological limitation, because it conditions the stability of the LDA and the statistical comparability between the models. The Maldita corpus contains around eleven times as many articles as the Polígrafo corpus, a difference potentially associated with multiple factors (team size, funding model, exposure to viral disinformation, editorial priorities), and its production is sustained over several years and institutionalised in a dedicated vertical, whereas Polígrafo’s is concentrated in a single year and tied to a funded project. Rather than a difference in rigour, this may suggest different forms of editorial organisation in the two platforms; the design does not separate national context from organisational conditions. Comparative studies on fact-checking in Spain and in Latin America report similar structural asymmetries in the scale of production and institutionalisation, with greater development in the Spanish case (Esteban-Navarro et al. 2021; Guallar et al. 2022; Palau-Sampio 2018), although without a specific focus on the climate theme. There are, at the same time, symmetries that the comparison should not obscure: both corpora contain an institutional press release announcing an externally funded project (FactCRICIS in Polígrafo, the launch of Maldito Clima in Maldita), both platforms republish third-party content (Agência Pública in one case, Climática/La Marea in the other), and in one Portuguese piece the verdicts are anchored in experts consulted by Maldita itself, a detail that illustrates the asymmetric circulation of verification authority within the Iberian space.
The differences identified between the two corpora admit, however, alternative explanations that should be pondered. To begin with, the keywords used in the collection differ between languages («crise climática» vs. «cambio climático»), and these lexical differences could capture partially distinct editorial universes, although the Portuguese search also included the generic «alterações climáticas». On the other hand, site architecture and indexing practices could influence the accessibility of content to web scraping, which would make Maldita potentially more accessible than Polígrafo. It should be added that the prevalence of Maldita’s Topic 4 (extreme phenomena, 21.6%) could owe less to an editorial orientation than to the occurrence of meteorological events of great impact in the analysed period, in particular the DANA of October 2024, which generated intensive fact-checking activity. The geographic profile of Polígrafo also deserves a note of caution, since Brazil’s position as the platform’s most frequent location derives essentially from the three republications of Agência Pública, whose content is Brazilian, and should not be read as a geopolitical orientation of the platform’s own production. These alternative explanations suggest that the observed differences may reflect a combination of platform editorial choices, contextual conditions and retrieval effects that the present design cannot disentangle. Climate education and media literacy, public concern, the volume of climate news and agreements with online platforms are further possible influences that were not measured here.
The results have implications for practice and for research on climate fact-checking in the Iberian and European context. The predominance of the pedagogical register in the collected Polígrafo texts can inform newsroom assessments of the balance between explanatory coverage and direct debunking. Its association with project funding also raises questions for funding policy, given that much of the Polígrafo output documented here was linked to a funded project. Audience interest, which was not measured here, could also influence this pattern, since the literature demonstrates that journalistic production, and fact-checking with it, tends to be influenced by consumption metrics and public preferences (Fürst 2020; Graves et al. 2016; Lee et al. 2014; Zamith 2018). The emergence of an autonomous visual fact-checking topic in Maldita signals the growing importance of visual literacy in the verification of climate content, a dimension on which research in science communication has not yet sufficiently dwelt (Lee and Suk 2025; Leßmöllmann 2020; Schäfer 2012, 2020). Systematic visual fact-checking was not identified in the collected Polígrafo texts, a finding that may inform newsroom assessments and research on visual media literacy. A broader collection would be needed to establish whether this reflects a coverage gap. The convergence of epistemic regimes has, finally, its own implication: if both platforms rest their factual assessments on scientific and meteorological authority, the communicative effectiveness of verification may depend less on which authorities are invoked than on the genre through which they reach audiences, whether the service explainer or the adversarial debunk. Confirmation of this hypothesis will demand reception studies that link the production models mapped here to their effects.

Limitations and Future Directions

The main limitation stems from the asymmetry of the corpora (26 vs. 277 documents), which conditions the direct comparability of the LDA models. With only 26 documents, the Polígrafo model is particularly sensitive to corpus composition, and the fit metrics (ldatuning) did not converge on a clear statistical optimum, a pattern known in small corpora (Agrawal et al. 2018). Coherence scores and recurring vocabulary across values of k do not eliminate this uncertainty. A leave-one-out test revealed limited stability of the document-topic assignments in Polígrafo (M = 40.2%), a result consistent with the reduced size of the corpus and one that reinforces the exploratory character of this model. The strong temporal concentration of the Polígrafo corpus means, furthermore, that the identified topic structure reflects the agenda of that specific period, and largely of a specific funded project, and cannot be generalised as a multi-year coverage pattern. The Maldita analysis, with n = 277, is less constrained by corpus size, but neither platform represents its national fact-checking system. Alternative approaches such as the Structural Topic Model (STM), with an appropriate multilingual representation and platform as a covariate, could be explored and constitute a direction for future research.
The functional coding carries its own limitations. In the Portuguese case all 26 articles were examined, with functional coding records for the 25 with retained ORG or PER entities; in the Spanish case, coding rests on a stratified sample (35 random articles plus a purposive complement of 11), stratified by year rather than by topic. Original mention counts constitute lower bounds because the coding used texts condensed to entity-bearing sentences; the European Union and Zero counts were subsequently checked against the archived cleaned article bodies. These constraints, together with single-researcher coding and the absence of an intercoder reliability estimate, recommend reading the functional results as an interpretive complement to the computational mapping rather than as an independently validated classification. Finally, the analytical potential of NER was explored essentially at the level of type counts, entity frequency and functional interpretation. More sophisticated analyses, such as entity co-occurrence networks and centrality measures, could reveal relational patterns that would complement the present analysis. The absence of a manual gold standard for the evaluation of NER precision and recall in the specific corpus constitutes an additional limitation, made salient by the recognition inconsistencies documented here for multi-variant institutional names. A formal evaluation with a manually annotated sample remains a promising direction, in line with the protocol of Ehrmann et al. (2016).

6. Conclusions

To close, we return to what we set out in the introduction. We sought to understand which thematic fields structure climate-related coverage in Polígrafo and Maldita, which entities sustain its factual authority and how the two relate. From the outset, the data show a strong asymmetry between platforms, since Maldita presents a larger corpus spanning several years, institutionalised in a dedicated vertical, whereas Polígrafo reveals a limited production concentrated in 2022 and largely tied to an externally funded project. This pattern suggests different organisational conditions in the two platforms; it does not establish national differences. The exploratory thematic structures also diverge. Maldita’s corpus is organised around denialism, the scientific evidence on temperature, extreme phenomena and visual fact-checking, whereas Polígrafo structures its corpus around the global climate crisis, recycling and green policies, and health and urban contamination. Whereas Maldita addresses visual disinformation in the collected material through an adversarial genre with named targets, Polígrafo orients itself towards an institutional and pedagogical mediation.
Regarding the regimes of factual legitimation, the functional coding delivers the study’s central refinement. The apparent opposition between a Polígrafo anchored in European political authority and a Maldita anchored in science does not survive the analysis of what the mentioned entities actually do in the texts. The coded texts in both platforms support their factual assessments by recourse to scientific, meteorological and health authority, and the European institutions that dominate Polígrafo’s mention rankings operate mainly as topic, regulatory frame and funding context rather than as legitimating instances. The coded material suggests convergence between these regimes; what diverges is the genre of climate-related coverage through which they are exercised, the presence or absence of a structured layer of named adversaries, and the scale and institutionalisation of production. The polarising public figures examined, for their part, enter this economy of authority predominantly as objects, with some contextual mentions of Trump, but never as sources.
From the analytical point of view, the study shows that the climate-related output of fact-checking platforms can be profitably mapped by two descriptive lenses, the referential ecologies of verification and the regimes of factual legitimation, provided that mention frequencies are interpreted in relation to the functions that mentions perform. The coded material suggests that sources constitute the very core of the authority of verification, in such a way that to verify is also to organise a disputed field of knowledge by selecting relevant sources and making explicit the evidence supporting assessments of climate claims. From the empirical point of view, the study offers an Iberian comparison still little developed in the literature, and documents how the same epistemic foundation can sustain very different editorial constructions in two neighbouring national contexts, with implications both for funding policies, given the concentration of funded articles in the Polígrafo corpus, and for the editorial strategies of verification platforms. The next step for future research remains to understand whether and how these production models condition the impact of verification on audiences, since different genres may affect disinformation in different ways.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/socsci15090629/s1. Table S1: List of the 26 Polígrafo articles included in the study; Table S2: List of the 277 Maldita articles included in the study.

Author Contributions

Conceptualisation, F.C. and J.P.B.; methodology, F.C.; software, F.C.; validation, F.C., J.P.B. and K.P.G.; formal analysis, F.C.; investigation, F.C. and J.P.B.; resources, F.C.; data curation, F.C.; writing—original draft preparation, F.C. and K.P.G.; writing—review and editing, J.P.B.; supervision, J.P.B.; project administration, J.P.B.; funding acquisition, J.P.B. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Foundation for Science and Technology (FCT) under Grant Ref. 2023.12581.PEX.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The R analysis scripts, the entity-function coding tables, the reconciliation records and the derived data files that support the results presented are available from the authors upon reasonable request.

Acknowledgments

During the preparation of this manuscript the authors used ChatGPT 5.5 for proofreading, translation and formatting. It also assisted with reconciling entity counts during revision. The authors take full responsibility for the content.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Agrawal, Amritanshu, Wei Fu, and Tim Menzies. 2018. What Is Wrong with Topic Modeling? And How to Fix It Using Search-Based Software Engineering. Information and Software Technology 98: 74–88. [Google Scholar] [CrossRef] [Scilit]
  2. Ahmed, Fasih, Muhammad Nawaz, and Aisha Jadoon. 2022. Topic Modeling of the Pakistani Economy in English Newspapers via Latent Dirichlet Allocation (LDA). SAGE Open 12: 21582440221079931. [Google Scholar] [CrossRef] [Scilit]
  3. Arun, Rajkumar, Venkatasubramaniyan Suresh, C. E. Veni Madhavan, and M. N. Narasimha Murthy. 2010. On Finding the Natural Number of Topics with Latent Dirichlet Allocation: Some Observations. In Advances in Knowledge Discovery and Data Mining. Edited by Mohammed J. Zaki, Jeffrey Xu Yu, B. Ravindran and Vikram Pudi. Berlin and Heidelberg: Springer, pp. 391–402. [Google Scholar] [CrossRef] [Scilit]
  4. Baptista, João-Pedro, Pedro Jerónimo, Valeriano Piñeiro-Naval, and Anabela Gradim. 2022. Elections and Fact-Checking in Portugal: The Case of the 2019 and 2022 Legislative Elections. Profesional de la Información 31: e310611. [Google Scholar] [CrossRef] [Scilit]
  5. Blei, David M., Andrew Y. Ng, and Michael I. Jordan. 2003. Latent Dirichlet Allocation. Journal of Machine Learning Research 3: 993–1022. [Google Scholar]
  6. Brookes, Gavin, and Tony McEnery. 2019. The Utility of Topic Modelling for Discourse Studies: A Critical Evaluation. Discourse Studies 21: 3–21. [Google Scholar] [CrossRef] [Scilit]
  7. Cao, Juan, Tian Xia, Jintao Li, Yongdong Zhang, and Sheng Tang. 2009. A Density-Based Method for Adaptive LDA Model Selection. Neurocomputing 72: 1775–81. [Google Scholar] [CrossRef] [Scilit]
  8. Coan, Travis G., Constantine Boussalis, John Cook, and Mirjam O. Nanko. 2021. Computer-Assisted Classification of Contrarian Claims about Climate Change. Scientific Reports 11: 22320. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Cook, John, Stephan Lewandowsky, and Ullrich K. H. Ecker. 2017. Neutralizing Misinformation through Inoculation: Exposing Misleading Argumentation Techniques Reduces Their Influence. PLoS ONE 12: e0175799. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Deveaud, Romain, Eric SanJuan, and Patrice Bellot. 2014. Accurate and Effective Latent Concept Modeling for Ad Hoc Information Retrieval. Document Numérique 17: 61–84. [Google Scholar] [CrossRef] [Scilit]
  11. Ehrmann, Maud, Giovanni Colavizza, Yannick Rochat, and Frédéric Kaplan. 2016. Diachronic Evaluation of NER Systems on Old Newspapers. In Proceedings of the 13th Conference on Natural Language Processing (KONVENS 2016). Bochum: Bochumer Linguistische Arbeitsberichte, pp. 97–107. [Google Scholar]
  12. Ekström, Mats, and Oscar Westlund. 2019. The Dislocation of News Journalism: A Conceptual Framework for the Study of Epistemologies of Digital Journalism. Media and Communication 7: 259–70. [Google Scholar] [CrossRef] [Scilit]
  13. Ekström, Mats, Seth C. Lewis, and Oscar Westlund. 2020. Epistemologies of Digital Journalism and the Study of Misinformation. New Media & Society 22: 205–12. [Google Scholar] [CrossRef] [Scilit]
  14. Esteban-Navarro, Miguel-Ángel, Antonia-Isabel Nogales-Bocio, Miguel-Ángel García-Madurga, and Tamara Morte-Nadal. 2021. Spanish Fact-Checking Services: An Approach to Their Business Models. Publications 9: 38. [Google Scholar] [CrossRef] [Scilit]
  15. Fernández-Castrillo, Carolina, and Raúl Magallón-Rosa. 2023. El periodismo especializado ante el obstruccionismo climático. El caso de Maldito Clima. Revista Mediterránea de Comunicación 14: 35–52. [Google Scholar] [CrossRef] [Scilit]
  16. Fürst, Silke. 2020. In the Service of Good Journalism and Audience Interests? How Audience Metrics Affect News Quality. Media and Communication 8: 270–80. [Google Scholar] [CrossRef] [Scilit]
  17. Graves, Lucas. 2016. Deciding What’s True: The Rise of Political Fact-Checking in American Journalism. New York: Columbia University Press. [Google Scholar]
  18. Graves, Lucas. 2017. Anatomy of a Fact Check: Objective Practice and the Contested Epistemology of Fact Checking. Communication, Culture & Critique 10: 518–37. [Google Scholar] [CrossRef] [Scilit]
  19. Graves, Lucas, Brendan Nyhan, and Jason Reifler. 2016. Understanding Innovations in Journalistic Practice: A Field Experiment Examining Motivations for Fact-Checking. Journal of Communication 66: 102–38. [Google Scholar] [CrossRef] [Scilit]
  20. Griffiths, Thomas L., and Mark Steyvers. 2004. Finding Scientific Topics. Proceedings of the National Academy of Sciences of the United States of America 101: 5228–35. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Grün, Bettina, and Kurt Hornik. 2011. topicmodels: An R Package for Fitting Topic Models. Journal of Statistical Software 40: 1–30. [Google Scholar] [CrossRef] [Scilit]
  22. Guallar, Javier, Lluís Codina, Pere Freixa, and Mario Pérez-Montoro. 2022. Disinformation, Hoaxes, Curation and Verification: Review of Studies in Ibero-America 2017–2020. Online Media and Global Communication 1: 648–68. [Google Scholar] [CrossRef] [Scilit]
  23. Heidenreich, Tobias, Fabienne Lind, Jakob-Moritz Eberl, and Hajo G. Boomgaarden. 2019. Media Framing Dynamics of the ‘European Refugee Crisis’: A Comparative Topic Modelling Approach. Journal of Refugee Studies 32: i172–82. [Google Scholar] [CrossRef] [Scilit]
  24. Humprecht, Edda. 2020. How Do They Debunk ‘Fake News’? A Cross-National Comparison of Transparency in Fact Checks. Digital Journalism 8: 310–27. [Google Scholar] [CrossRef] [Scilit]
  25. Jerónimo, Pedro, and Marta Sánchez Esparza. 2023. Jornalistas locais e fact-checking: Um estudo exploratório em Portugal e Espanha. Comunicação e Sociedade 44: e023016. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Lamb, William F., Giulio Mattioli, Sebastian Levi, J. Timmons Roberts, Stuart Capstick, Felix Creutzig, Jan C. Minx, Finn Müller-Hansen, Trevor Culhane, and Julia K. Steinberger. 2020. Discourses of Climate Delay. Global Sustainability 3: e17. [Google Scholar] [CrossRef] [Scilit]
  27. Lee, Angela M., Seth C. Lewis, and Matthew Powers. 2014. Audience Clicks and News Placement: A Study of Time-Lagged Influence in Online Journalism. Communication Research 41: 505–30. [Google Scholar] [CrossRef] [Scilit]
  28. Lee, Jiyoung, and Jiyoun Suk. 2025. Navigating the Complexity of Visual Misinformation: Developing the Visual Misinformation Processing Model for Visual-Text Misinformation Dynamics. Communication Theory 35: 238–49. [Google Scholar] [CrossRef] [Scilit]
  29. Lee, So Chung. 2022. Topic Modeling of Korean Newspaper Articles on Aging via Latent Dirichlet Allocation. Asian Journal for Public Opinion Research 10: 4–22. [Google Scholar] [CrossRef]
  30. Leßmöllmann, Annette. 2020. Current Trends and Future Visions of (Research on) Science Communication. In Science Communication. Edited by Annette Leßmöllmann, Marcelo Dascal and Thomas Gloning. Berlin: De Gruyter Mouton, pp. 657–88. [Google Scholar] [CrossRef] [Scilit]
  31. Lewandowsky, Stephan. 2021. Climate Change Disinformation and How to Combat It. Annual Review of Public Health 42: 1–21. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  32. Mimno, David, Hanna M. Wallach, Edmund Talley, Miriam Leenders, and Andrew McCallum. 2011. Optimizing Semantic Coherence in Topic Models. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing. Edinburgh: Association for Computational Linguistics, pp. 262–72. [Google Scholar]
  33. Murzintcev, Nikita. 2024. ldatuning: Tuning of the Latent Dirichlet Allocation Models Parameters. R Package Version 1.0.3. Available online: https://github.com/nikita-moor/ldatuning (accessed on 13 September 2026).
  34. Nadeau, David, and Satoshi Sekine. 2007. A Survey of Named Entity Recognition and Classification. Lingvisticae Investigationes 30: 3–26. [Google Scholar] [CrossRef] [Scilit]
  35. Oliveira, Ana Filipa, Micael Maneta, Maria José Brites, and Vanessa Ribeiro-Rodrigues. 2024. How Is Portugal Addressing Disinformation? Results of a Mapping of Initiatives (2010–2023). Observatorio (OBS*) 18: 158–74. [Google Scholar] [CrossRef] [Scilit]
  36. O’Neill, Saffron J., and Nicholas Smith. 2014. Climate Change and Visual Imagery. WIREs Climate Change 5: 73–87. [Google Scholar] [CrossRef] [Scilit]
  37. Palau-Sampio, Dolors. 2018. Fact-Checking and Scrutiny of Power: Supervision of Public Discourses in New Media Platforms from Latin America. Communication & Society 31: 347–65. [Google Scholar] [CrossRef] [Scilit]
  38. Palau-Sampio, Dolors, Paz Crisóstomo Flores, and Maria Josep Picó Garcés. 2024. Fuelling Climate Change Disinformation: Global Narratives Distorting Environmental Risks in North America, Europe and Latin America. Catalan Journal of Communication & Cultural Studies 16: 217–36. [Google Scholar] [CrossRef] [Scilit]
  39. Schäfer, Mike S. 2012. Online Communication on Climate Change and Climate Politics: A Literature Review. WIREs Climate Change 3: 527–43. [Google Scholar] [CrossRef] [Scilit]
  40. Schäfer, Mike S. 2020. News Media Imagery of Climate Change: Reviewing the Research. In Research Handbook on Communicating Climate Change. Edited by David C. Holmes and Lucy M. Richardson. Cheltenham: Edward Elgar, pp. 131–42. [Google Scholar]
  41. Schmidt, Luísa, and Ana Delicado. 2014. Alterações climáticas na opinião pública portuguesa. In Ambiente, Alterações Climáticas, Alimentação e Energia: A Opinião dos Portugueses. Edited by Luísa Schmidt and Ana Delicado. Lisboa: Imprensa de Ciências Sociais. [Google Scholar]
  42. Sendra-Duro, Enric. 2025. Fact-checking y desinformación climática en España. Tendencias en la cobertura, tematización y gestión de fuentes en la estrategia profesional de EFE Verifica, Maldito Clima y Newtral. Doxa Comunicación 41: 561–87. [Google Scholar] [CrossRef] [Scilit]
  43. Sievert, Carson, and Kenneth Shirley. 2014. LDAvis: A Method for Visualizing and Interpreting Topics. In Proceedings of the Workshop on Interactive Language Learning, Visualization, and Interfaces. Baltimore: Association for Computational Linguistics, pp. 63–70. [Google Scholar] [CrossRef] [Scilit]
  44. Silge, Julia, and David Robinson. 2017. Text Mining with R: A Tidy Approach. Sebastopol: O’Reilly Media. [Google Scholar]
  45. Törnberg, Anton, and Petter Törnberg. 2025. The Aesthetics of Climate Misinformation: Computational Multimodal Framing Analysis with BERTopic and CLIP. Environmental Politics 35: 1070–93. [Google Scholar] [CrossRef] [Scilit]
  46. Treen, Kathie M. d’I., Hywel T. P. Williams, and Saffron J. O’Neill. 2020. Online Misinformation about Climate Change. WIREs Climate Change 11: e665. [Google Scholar] [CrossRef] [Scilit]
  47. Valdez, Danny, Andrew C. Pickett, and Patricia Goodson. 2018. Topic Modeling: Latent Semantic Analysis for the Social Sciences. Social Science Quarterly 99: 1665–79. [Google Scholar] [CrossRef] [Scilit]
  48. Van Aelst, Peter, Tamir Sheafer, and James Stanyer. 2012. The Personalization of Mediated Political Communication: A Review of Concepts, Operationalizations and Key Findings. Journalism 13: 203–20. [Google Scholar] [CrossRef] [Scilit]
  49. van der Linden, Sander. 2022. Misinformation: Susceptibility, Spread, and Interventions to Immunize the Public. Nature Medicine 28: 460–67. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  50. van der Linden, Sander. 2024. Countering Misinformation through Psychological Inoculation. Advances in Experimental Social Psychology 69: 1–58. [Google Scholar] [CrossRef] [Scilit]
  51. Vizoso, Ángel, and Jorge Vázquez-Herrero. 2019. Plataformas de fact-checking en español: Características, organización y método. Communication & Society 32: 127–44. [Google Scholar] [CrossRef] [Scilit]
  52. Vu, Hong Tien, Annalise Baines, and Nhung Nguyen. 2023. Fact-Checking Climate Change: An Analysis of Claims and Verification Practices by Fact-Checkers in Four Countries. Journalism & Mass Communication Quarterly 100: 286–307. [Google Scholar] [CrossRef] [Scilit]
  53. Zamith, Rodrigo. 2018. Quantified Audiences in News Production: A Synthesis and Research Agenda. Digital Journalism 6: 418–35. [Google Scholar] [CrossRef] [Scilit]
  54. Zhang, Zhi, Duoqian Miao, and Can Gao. 2013. Short Text Classification Using Latent Dirichlet Allocation. Journal of Computer Applications 33: 1587–90. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Fit metrics (ldatuning) for the selection of the number of topics (k = 2 to 8). Upper panel (a), Polígrafo (n = 26); lower panel (b), Maldita (n = 277). Metrics: Griffiths and Steyvers (2004), Cao et al. (2009), Arun et al. (2010) and Deveaud et al. (2014).
Figure 1. Fit metrics (ldatuning) for the selection of the number of topics (k = 2 to 8). Upper panel (a), Polígrafo (n = 26); lower panel (b), Maldita (n = 277). Metrics: Griffiths and Steyvers (2004), Cao et al. (2009), Arun et al. (2010) and Deveaud et al. (2014).
Socsci 15 00629 g001
Figure 2. Intertopic maps generated by LDAvis (Sievert and Shirley 2014). (a) Polígrafo; (b) Maldita. Circle size reflects token-weighted topic prevalence and the distance between circles the semantic dissimilarity between topics.
Figure 2. Intertopic maps generated by LDAvis (Sievert and Shirley 2014). (a) Polígrafo; (b) Maldita. Circle size reflects token-weighted topic prevalence and the distance between circles the semantic dissimilarity between topics.
Socsci 15 00629 g002
Figure 4. Named entity profile by topic in Polígrafo (% ORG, LOC, PER per topic, k = 4).
Figure 4. Named entity profile by topic in Polígrafo (% ORG, LOC, PER per topic, k = 4).
Socsci 15 00629 g004
Figure 5. Named entity profile by topic in Maldita (% ORG, LOC, PER per topic, k = 4).
Figure 5. Named entity profile by topic in Maldita (% ORG, LOC, PER per topic, k = 4).
Socsci 15 00629 g005
Table 1. Characterisation of the corpora.
Table 1. Characterisation of the corpora.
PolígrafoMaldita
CountryPortugalSpain
LanguagePortugueseSpanish
n (articles)26277
Period2019–20242019–2025
CollectionJune 2025June 2025
Sourcepoligrafo.sapo.ptmaldita.es
Table 2. UMass topic coherence (top-10 and top-20 words per topic).
Table 2. UMass topic coherence (top-10 and top-20 words per topic).
TopicPOL Top-10POL Top-20MALD Top-10MALD Top-20
1−0.690−0.623−0.493−0.676
2−0.777−0.628−0.928−1.135
3−0.459−0.625−0.898−1.080
4−0.470−0.614−1.182−1.185
Note. Pair-averaged scores for the final reported topics, with add-one smoothing in the co-occurrence numerator. Higher values indicate stronger word co-occurrence; scores do not establish the stability of topic assignments.
Table 3. Proportional distribution of LDA topics and interpretive labels (k = 4).
Table 3. Proportional distribution of LDA topics and interpretive labels (k = 4).
PlatformTLabel% CorpusDocsKeywords
Polígrafo1Global climate crisis19.75emissões, países, temperatura
2Recycling and green policies33.19água, verde, ecoponto
3Health and climate impacts26.97amianto, saúde, mar
4Urban contamination20.35terrenos, solos, lisboa
Maldita1Climate denialism37.6104temperatura, calor, calentamiento
2Climate policy19.149emisiones, COP, medidas
3Visual fact-checking21.761mapa, bulo, imagen
4Extreme weather events21.663agua, incendios, DANA
Note. Keywords are reported in the original language of each corpus. Percentages are mean document-topic weights; Docs counts dominant-topic assignments. The two summaries need not coincide.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Conrado, F.; Garzón, K.P.; Baptista, J.P. Climate-Related Content on Iberian Fact-Checking Platforms: Topics, Entities and Factual Authority in Polígrafo and Maldita. Soc. Sci. 2026, 15, 629. https://doi.org/10.3390/socsci15090629

AMA Style

Conrado F, Garzón KP, Baptista JP. Climate-Related Content on Iberian Fact-Checking Platforms: Topics, Entities and Factual Authority in Polígrafo and Maldita. Social Sciences. 2026; 15(9):629. https://doi.org/10.3390/socsci15090629

Chicago/Turabian Style

Conrado, Francisco, Karen Pinto Garzón, and João Pedro Baptista. 2026. "Climate-Related Content on Iberian Fact-Checking Platforms: Topics, Entities and Factual Authority in Polígrafo and Maldita" Social Sciences 15, no. 9: 629. https://doi.org/10.3390/socsci15090629

APA Style

Conrado, F., Garzón, K. P., & Baptista, J. P. (2026). Climate-Related Content on Iberian Fact-Checking Platforms: Topics, Entities and Factual Authority in Polígrafo and Maldita. Social Sciences, 15(9), 629. https://doi.org/10.3390/socsci15090629

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop