1. Introduction
Article 5 GDPR lies at the core of the EU data protection framework because it sets out the principles governing the lawful processing of personal data, namely lawfulness, fairness and transparency, purpose limitation, data minimisation, accuracy, storage limitation, integrity and confidentiality, and accountability (
Quelle 2018). These principles are not merely introductory statements. They structure the regulation as a whole and inform the interpretation of more specific obligations imposed elsewhere in the GDPR, particularly those concerning lawful basis, transparency, security, and automated decision making (
Bygrave 2017;
Jasmontaite et al. 2018;
Malgieri 2019;
Kaminski 2019). Although GDPR compliance is often discussed primarily as a legal requirement, recurring infringements of Article 5 also reveal how organisations repeatedly fail to justify, limit, explain, secure, and document their processing practices in ways that the regulation requires (
Quelle 2018;
Novelli et al. 2024;
Papagiannidis et al. 2025). This is especially relevant in a period in which artificial intelligence is increasingly embedded in organisational decision making, service delivery, profiling, ranking, and other forms of data-intensive processing, often involving large volumes of personal data and technically complex processing chains (
Ufert 2020;
Laux 2024). While such systems may generate operational efficiencies and new forms of value creation, they also intensify risks relating to opacity, unfairness, security, discrimination, and diminished individual control over personal information (
Ufert 2020;
Hamon et al. 2022;
European Data Protection Board 2018). Recent scholarship also shows that these risks arise not only at the level of outcomes, but already at the level of system design, impact assessment, and privacy protection in machine learning environments, where questions of fairness, proportionality, and privacy preservation are closely interconnected (
Calvi and Kotzinos 2023;
El Mestari et al. 2024). At the same time, the EU digital compliance landscape is often criticised as overly complex and fragmented, creating legal uncertainty and increasing the costs of compliant innovation, particularly for data-intensive and AI-related business models (
Belen-Saglam et al. 2023). The legal relevance of automated decision making in this context has also been clarified in recent CJEU case law, particularly in SCHUFA and subsequent case law on automated assessments, which confirms that certain forms of automated evaluation may fall within the scope of Article 22 GDPR and that data subjects must be able to obtain meaningful information and effective review of such outcomes (
Court of Justice of the European Union 2023,
2025). These concerns are also reflected in the European Commission’s “Digital Omnibus” initiative, which seeks to simplify the practical application of the GDPR and the AI Act, especially for organisations and SMEs, with a view to strengthening legal certainty and supporting innovation (
Skiotytė and Sadauskaitė 2026). Regardless of the direction such simplification efforts may take, Article 5 GDPR remains central to the legal assessment of personal data processing because it contains the principles that shape the lawfulness of processing across the regulation as a whole (
European Union 2016;
Quelle 2018). In practice, weaknesses at the level of Article 5 frequently trigger or coincide with infringements of other GDPR provisions, especially Article 6 on lawful basis, Articles 12 to 15 on transparency and information, Article 22 on automated decision making, and Article 32 on security of processing (
European Union 2016;
Malgieri 2019;
Kaminski 2019;
European Data Protection Board 2018). This broader relationship between the principles in Article 5 and more specific GDPR obligations is also consistent with leading commentary, which treats those principles as central to the structure and application of the regulation as a whole, including in relation to lawful basis, special categories of personal data, and automated decision making (
Kuner et al. 2020). This makes Article 5 a particularly useful entry point for examining how unlawful or deficient processing is reflected in supervisory practice. The broader legal context is also important. In the European legal tradition, personal data protection is rooted in a human rights framework centred on the protection of the individual, human dignity, and autonomy, rather than being treated simply as a matter of market regulation or commercial efficiency (
European Union Agency for Fundamental Rights et al. 2018;
Mantelero 2020). The right to the protection of personal data is distinct from the right to privacy, even though the two are closely connected. Privacy protects the wider personal sphere, whereas data protection provides a more specific legal structure for the collection, use, storage, and disclosure of information relating to identifiable persons (
European Union Agency for Fundamental Rights et al. 2018;
Mantelero 2020). This distinction becomes particularly significant in digital environments, where data that may appear benign at first sight can become highly revealing once they are aggregated, cross referenced, or used for inference and profiling (
Mantelero 2020). Convention 108+ reinforces the same broader orientation by framing data protection as part of a human rights-based legal culture that extends beyond the EU legal order strictly understood (
Council of Europe n.d.;
Mantelero 2020). The relationship between the GDPR and the Artificial Intelligence Act should be understood in that same context. The AI Act does not displace data protection law. Rather, where personal data are processed, the GDPR continues to apply, while the AI Act adds a complementary regulatory layer concerning transparency, human oversight, and risk management (
European Union 2024;
Ufert 2020;
Laux 2024;
Lazcoz and de Hert 2023). Article 2(7) of the AI Act, together with recitals 9 and 10, makes this relationship explicit (
European Union 2024;
Croatian Personal Data Protection Agency 2025). At the same time, this article does not treat the analysed decision set as an AI-enforcement dataset. Only a limited subset of the coded decisions concerns clearly identifiable AI-related or automation- and profiling-related processing. The relevance of that context lies elsewhere. The recurring legal weaknesses visible in general Article 5 enforcement, especially those concerning lawful basis, transparency, security, retention, and accountability, are also the weaknesses that become especially significant where processing is more opaque, scalable, or technologically complex (
Ufert 2020;
Laux 2024;
De Hert and Lazcoz 2021;
Birahim 2025). Against this background, this article addresses a gap in the literature through an empirical analysis of national data protection authority decisions across the European Economic Area. It draws on an initial GDPRhub retrieval pool of 1660 publicly available decisions issued between 25 May 2018 and 15 September 2025, from which a final analytical sample of 790 national DPA decisions involving infringements related to Article 5 GDPR was constructed. Using structured content analysis, the article identifies recurring infringement patterns, examines the co occurrence of Article 5 with other GDPR provisions, analyses selected sectoral and contextual dimensions of supervisory practice, and considers the limited but relevant implications of those findings for automated and data-intensive processing. In that sense, the article contributes in three respects. First, it maps recurring enforcement patterns relating to Article 5 GDPR in a large EEA decision set. Second, it clarifies how Article 5 infringements are connected in supervisory practice with other GDPR provisions, particularly Articles 6, 12 to 15, 22, and 32. Third, it shows why those enforcement patterns remain relevant in more technologically complex processing environments, even though the dataset itself is not predominantly AI-specific (
Papagiannidis et al. 2025;
Pagallo 2019;
Novelli et al. 2024). The analysis is guided by three research questions: What recurrent infringement patterns relating to Article 5 GDPR emerge from national DPA decisions across the EEA? Which GDPR provisions most frequently co-occur with Article 5 infringements in supervisory practice, and what do these patterns reveal about the practical application of Article 5 GDPR? To what extent do sectoral context and automated processing context shape the composition of Article 5-related non-compliance in national DPA decisions?
2. Methods
This study is based on an empirical legal analysis of publicly available supervisory authority decisions and combines structured content analysis with descriptive quantitative examination of the coded dataset. Its purpose is to identify recurrent patterns of non-compliance relating to Article 5 GDPR in national DPA practice across the European Economic Area (EEA), to examine which other GDPR provisions most frequently appear alongside Article 5 and to assess selected contextual features of the published decision record. The study draws on both primary and secondary sources. The primary material consists of supervisory authority decisions available through GDPRhub and, where necessary, through the official websites of national data protection authorities. Secondary sources include academic literature, EU legislation, guidance issued by the European Data Protection Board together with the European Commission, as well as selected materials relevant to the legal interpretation of Article 5 GDPR and related provisions. These materials were used to frame the legal analysis and to support the construction of the coding scheme. In that respect, the coding framework was designed to translate legal requirements into operational analytical categories rather than to treat compliance only as an ex post enforcement outcome, an approach that is consistent with scholarship on data protection by design and the structured evidencing of compliance (
Chhetri et al. 2022;
Koulierakis 2024).
The initial retrieval pool consisted of 1660 decisions extracted from GDPRhub on 16 September 2025. It included decisions from EEA jurisdictions published in the database for the period from 25 May 2018 to 15 September 2025 in which one or more GDPR infringements had been identified. The jurisdictions covered were Austria, Belgium, Bulgaria, Croatia, Cyprus, Czechia, Denmark, Estonia, Finland, France, Germany, Greece, Hungary, Iceland, Ireland, Italy, Latvia, Lithuania, Luxembourg, Malta, the Netherlands, Norway, Poland, Portugal, Romania, Slovenia, Spain, together with Sweden. For dataset construction, the search was limited to the GDPRhub outcome categories complaint upheld, partially upheld and investigation: violation found.
GDPRhub should not be treated as a complete official repository of all national DPA decisions in the EEA. It is a publicly accessible database that collects and summarises GDPR-related decisions, but inclusion depends on publication, availability, together with submission practices. The initial dataset therefore reflects the published decision record available through that source rather than the full underlying universe of supervisory enforcement. This limitation is important and is taken into account in interpreting the findings. The present study was limited to national DPA decisions involving infringements related to Article 5 GDPR. During data preparation, exact duplicate records were removed and the remaining records were screened against the study’s inclusion criteria. A record was retained only if it fell within the defined jurisdictional scope, referred to a unique supervisory authority decision, and contained sufficient publicly available factual and legal information to permit reliable coding of the relevant Article 5 principle or principles, co-occurring GDPR provisions, together with thematic compliance-failure categories. Records that did not meet these conditions were excluded. After screening and cleaning, the final analytical sample comprised 790 national DPA decisions. The unit of analysis was the individual supervisory authority decision. A distinction was therefore made between a decision and an infringement. One decision could include multiple infringements of Article 5, as well as infringements of other GDPR provisions. The number of coded infringements therefore exceeds the number of analysed decisions.
The coding strategy combined deductive and inductive elements. The legal dimension of the coding frame was defined deductively from the structure of the GDPR, with particular attention to Article 5 and to those provisions that frequently appeared alongside it in the decision record, especially Articles 6, 12–15, 22, together with 32. The thematic categories were refined inductively through an initial review of a subsample of decisions and decision summaries. That pilot stage was used to test category boundaries and to stabilise the final codebook before full coding began.
The coding process proceeded in five steps. First, decision records were collected and harmonised into a single working dataset. Second, a pilot subsample was reviewed in order to test the initial coding logic and identify borderline cases. Third, the codebook was revised so that the legal, thematic, contextual, together with outcome variables were accompanied by clearer operational definitions and decision rules. Fourth, the final analytical sample was coded decision by decision using a standardised protocol. Fifth, ambiguous cases were revisited and compared with previously coded decisions together with the underlying summaries in order to improve internal consistency across the dataset.
Coding was carried out manually in Microsoft Excel, which was used for dataset construction, cleaning, variable coding, together with descriptive tabulation. Primary coding was undertaken by one author. The second author reviewed the codebook, checked a subset of coded decisions, and examined ambiguous cases as part of a consistency review. Where disagreements arose, they were resolved through discussion and, where necessary, by refining the coding rules. Because the entire dataset was not independently double-coded, formal inter-coder reliability statistics were not calculated. This should be read as a limitation of the study, but not as a reason to disregard the dataset.
Each decision was coded at four levels: legal, thematic, contextual, together with outcome-related.
At the legal level, coding captured the infringed GDPR provisions identified in the decision record, including the specific Article 5 principle or principles concerned, as well as the presence of co-occurring provisions. Article 5 coding distinguished between Article 5(1)(a) to (f) and Article 5(2). Additional binary flags were used for other GDPR provisions, including Articles 6, 7, 9, 12–15, 17, 21, 22, 24, 25, 28, 32, together with 35. A separate variable recorded whether a decision involved Article 5 alone or whether Article 5 appeared together with other GDPR provisions.
At the thematic level, each decision was coded for its main compliance-failure pattern. The principal categories included lawful-basis deficiencies, transparency failures, inadequate technical or organisational measures, failures relating to data subject rights, employment-related processing, unauthorised disclosure or transfer to third parties, video surveillance, marketing, problematic data collection practices, public disclosure of personal data, children’s data, biometric processing, profiling or automated decision making, together with other or unspecified. Where the decision clearly revealed a second distinct pattern, a secondary thematic category was also recorded.
Two more granular sub-coding layers were added because the general thematic categories were too broad to capture some legally important differences. First, Article 6-related decisions were sub-coded to distinguish between cases in which no lawful basis had been identified, invalid reliance on consent, improper reliance on legitimate interest, improper reliance on contract, improper reliance on legal obligation or public task, together with special-category mismatch. Second, transparency-related decisions were sub-coded to distinguish between failure to inform under Article 13, failure to inform under Article 14, incomplete privacy notice or information, inaccurate or misleading notice, insufficient explanation of automated processing or profiling, retention-period transparency deficiency, together with general transparency deficiency where the available material did not permit a narrower classification.
An additional variable was used to capture the technological context of each decision in relation to AI, automation and profiling. Because publicly available supervisory decisions often contain only limited technical detail, the coding was deliberately conservative. Decisions were grouped into four categories: AI-explicit; automation/profiling-related; digitally intensive but not clearly AI-related; and not AI-related or with no clear indication.
A decision was coded as AI-explicit only where the available material expressly referred to artificial intelligence, machine learning, neural networks, predictive modelling, automated inference or another AI-related technique. References to scoring, ranking, profiling or automated processing were not, by themselves, sufficient for this classification. Where such practices were present, but the material did not indicate the use of AI techniques, the decision was coded as automation/profiling-related.
The digitally intensive category covered cases involving complex digital services, large-scale platforms, extensive data flows or technically intensive processing environments, without a sufficiently clear indication of either AI or automation/profiling. As a result, some cases that may have involved AI in practice were not coded as AI-explicit where the public record did not support that conclusion. This choice was made to avoid overstating the AI dimension of the dataset.
At the contextual level, decisions were coded by jurisdiction, sector, organisation type, together with, where possible, organisation size. Jurisdiction referred to the country of the national supervisory authority issuing the decision. Sector was coded conservatively on the basis of the decision summary and party description into broad categories such as public sector, finance, health and social care, education, technology/media/telecommunications, transport/utilities/industry, retail/services, civil society/professional bodies, together with other or unknown. Organisation type was coded as public or private/other. Organisation size was coded only where size was explicitly stated or could be inferred with sufficient confidence from the available public material. In many decisions, that was not possible. For that reason, size-based analysis is treated in this study as exploratory only.
At the outcome level, each decision was coded according to whether it resulted in an administrative fine or in other corrective measures without a fine. Fine amounts were recorded as reported in the source materials and, where necessary, converted into euro-equivalent values in order to permit aggregate descriptive comparison across the final sample and across jurisdictions.
The study then combined descriptive statistics with cross-tabulations together with selected statistical tests. Descriptive summaries were prepared in Excel. Inferential analyses were conducted in Eviews 12 SV. Cross-tabulations with Pearson chi-square tests and Cramér’s V were used to assess whether the distribution of infringement patterns differed across sectors and, where feasible, across organisation-size categories. Because fine amounts were highly skewed, sectoral differences in sanction levels were examined using the Kruskal–Wallis test rather than parametric comparison. Jurisdictional comparison was treated descriptively through counts of decisions, counts of fine-imposing decisions, together with aggregate euro-equivalent fine totals.
This methodological design does not permit claims about the full prevalence of GDPR non-compliance across the EEA, nor does it establish causal relationships between organisational characteristics and specific infringement outcomes. What it does permit is a structured empirical legal analysis of the published supervisory decision record. In that sense, the study provides evidence of how Article 5 GDPR appears in national DPA practice, which forms of non-compliance recur most often, together with how those findings relate to other GDPR provisions in concrete enforcement settings. A summary of the variables, operational definitions, values/categories and coding rules is provided in
Table 1.
3. Findings
The results presented below are based on the final analytical sample of 790 national DPA decisions involving infringements related to Article 5 GDPR, selected from an initial GDPRhub retrieval pool of 1660 publicly available decisions. Because the unit of analysis was the individual supervisory authority decision, and because one decision could contain more than one infringement, the number of coded Article 5 infringements exceeds the number of analysed decisions. Across the 790 decisions, a total of 1039 Article 5 infringements were identified, covering Article 5(1)(a) to (f) together with Article 5(2). Of the 790 decisions, 65 involved Article 5 alone, while 725 involved Article 5 together with at least one other GDPR provision. The analysis begins with the distribution of co-occurring GDPR provisions, then turns to the internal distribution of Article 5 principle-level infringements, followed by thematic coding, more granular sub-coding of lawful-basis and transparency failures, AI- and automation-related coding, together with selected contextual and outcome-related findings.
3.1. Co-Occurring GDPR Provisions
Figure 1 shows that Article 5-related non-compliance rarely appeared in isolation. After Article 5 itself, the most frequent co-occurring provision was Article 6 on lawful basis, recorded in 342 decisions. This was followed by Article 32 on security of processing in 204 decisions, Article 13 in 200 decisions, together with Article 12 in 193 decisions. Other frequently occurring provisions included Article 25 in 124 decisions, Article 24 in 118, Article 9 in 105, Article 15 in 79, Article 7 in 73, Articles 14 and 17 in 54 each, Article 28 in 40, Article 35 in 38, together with Article 21 in 33 decisions. Article 22 appeared in seven decisions. These figures show that, in published supervisory practice, Article 5 findings most often formed part of broader constellations of GDPR non-compliance rather than stand-alone violations.
3.2. Distribution of Article 5 Infringements
Figure 2 sets out the internal distribution of Article 5 principle-level infringements. Article 5(1)(a) was the most frequent category, appearing in 297 decisions. It was followed by Article 5(1)(c) in 192 decisions, Article 5(1)(f) in 188 decisions, together with Article 5(2) in 166 decisions. Article 5(1)(e) appeared in 77 decisions, Article 5(1)(b) in 75, while Article 5(1)(d) appeared in 44 decisions. In other words, the most recurrent Article 5 findings concerned lawfulness, fairness and transparency, data minimisation, integrity and confidentiality, together with accountability.
3.3. Thematic Categories of Non-Compliance
The thematic coding points in the same direction.
Figure 3 shows that the largest thematic category was lawful-basis deficiencies, identified in 357 decisions. Transparency-related failures were identified in 312 decisions, while inadequate technical or organisational measures appeared in 282. Other recurring categories included failures relating to data subject rights in 124 decisions, employment-related processing in 122, video surveillance in 81, marketing in 64, unauthorised disclosure or transfer in 59, public disclosure of personal data in 32, problematic collection practices in 28, children’s data in 25, biometric processing in 13, together with profiling or automated decision making in 11 decisions. Since more than one thematic pattern could be present in a single case, these counts exceed the total number of decisions.
3.4. Lawful-Basis and Transparency Sub-Coding
A more detailed view of Article 6-related cases showed that the largest subgroup consisted of decisions in which no lawful basis had been properly identified, recorded in 150 cases. This was followed by improper reliance on legal obligation or public task in 60 decisions, improper reliance on legitimate interest in 46, invalid reliance on consent in 41, special-category mismatch in 35, together with improper reliance on contract in 10 decisions. These figures show that lawful-basis problems did not take one uniform form. The published decisions instead reveal several recurrent types of legal deficiency under Article 6.
Transparency-related sub-coding also showed internal variation. The largest subgroup concerned failures to inform under Article 13, recorded in 154 decisions. This was followed by general transparency deficiencies in 75 decisions and failures to inform under Article 14 in 52. Smaller categories included retention-period transparency deficiencies in 12 decisions, insufficient explanation of automated processing or profiling in four, inaccurate or misleading notice in two, together with incomplete privacy notice or information in two decisions. These findings suggest that transparency-related non-compliance in supervisory practice extends beyond formal notice drafting and often concerns more basic failures to provide or structure information in a legally adequate way.
As shown in
Figure 4, the lawful-basis-related cases were not limited to one type of legal deficiency. The largest subgroup consisted of decisions in which no lawful basis had been properly identified, followed by cases involving improper reliance on legal obligation or public task, improper reliance on legitimate interest, invalid reliance on consent, special-category mismatch, and improper reliance on contract.
As shown in
Figure 5, transparency-related failures were not limited to one form of non-compliance. The largest subgroup concerned failures to inform data subjects under Article 13 GDPR, followed by general transparency deficiencies and failures to inform under Article 14 GDPR. Less frequent categories included retention-period transparency deficiencies, insufficient explanation of automated processing or profiling, inaccurate or misleading information and incomplete privacy notices or information.
3.5. AI- and Automation-Related Coding
The technological coding should be read cautiously. Of the 790 analysed decisions, 12 were classified as AI-explicit and nine as automation/profiling-related. A further 133 involved digitally intensive processing but could not be reliably classified as AI-related in the narrower sense. The remaining 636 decisions did not contain a sufficiently clear indication of AI-related or automation-specific processing in the publicly available record. The dataset should therefore not be presented as an AI-enforcement dataset. It is primarily a dataset of Article 5-related supervisory decisions, within which a limited but relevant subset concerns automation or clearly identifiable AI-related processing. The figures therefore reflect only what could be reliably identified from the public decision record, not the possible underlying use of AI in practice.
3.6. Outcome-Related and Contextual Findings
The outcome-related findings are also notable. Of the 790 analysed Article 5-related decisions, 529 resulted in an administrative fine, while 261 resulted in other corrective measures without a fine. The total euro-equivalent value of the fines imposed in these decisions was EUR 1,829,894,914.82. These figures show that Article 5-related findings were involved in high-level corrective action. In a large share of the published decisions, they were associated with pecuniary sanctions, in some cases substantial ones.
The distribution of decisions across jurisdictions was uneven. Italy accounted for 190 decisions, Spain for 112, Belgium for 81, Greece for 58, together with Hungary, Iceland, and Norway for 33 each. Finland contributed 32 decisions, Ireland 28, France 27, while Poland contributed 23. Lower counts were recorded for a number of other jurisdictions, including Croatia, Latvia, Slovenia, Estonia, Germany, Portugal, Lithuania, Malta, the Netherlands, Bulgaria, together with Czechia. These differences should not be read too quickly as differences in enforcement intensity alone. They also reflect differences in publication practices, decision visibility and the structure of the GDPRhub record itself.
Sectoral comparison was carried out on a reduced subset of 439 decisions after excluding cases coded as other/unknown sector and decisions in which the main thematic failure was classified as other/unspecified. The distribution of main thematic failures differed significantly across sectors (chi-square = 253.37, p < 0.001, Cramér’s V = 0.287). The sectoral patterns were not identical. In technology/media/telecommunications, transparency and technical or organisational measures were the most frequent categories. In the public sector, technical and organisational measures, lawful basis, transparency, employment-related processing, together with video surveillance were prominent. In health and social care, technical and organisational measures were again dominant, followed by unauthorised disclosure or transfer and employment-related processing. In finance, technical and organisational measures, lawful basis, data subject rights and transparency featured strongly, while in retail/services marketing clearly stood out. These findings indicate that Article 5-related non-compliance was not distributed evenly across sectors.
A second sectoral test focused on Article 6 subcategories. Using a reduced subset of 208 Article 6-related decisions, the analysis again showed a statistically significant association (chi-square = 134.31, p < 0.001, Cramér’s V = 0.359). Public-sector decisions were especially associated with improper reliance on legal obligation or public task and cases where no lawful basis had been properly identified. In technology/media/telecommunications, the dominant pattern was again the absence of a clearly identified lawful basis, followed by invalid reliance on consent. In health and social care, special-category mismatch was especially prominent. These patterns suggest that lawful-basis problems do not look the same across sectors and that the legal reasoning used to justify processing differs in recurring ways depending on the institutional setting.
Organisation-size analysis was more limited. Size could be coded with sufficient confidence only in a minority of cases, and the inferential comparison was based on a reduced subset of 194 decisions. Within that subset, the chi-square test suggested a statistically significant association between organisation size and main thematic failure (chi-square = 50.26, p = 0.0276, Cramér’s V = 0.294). Large organisations appeared more often in decisions involving technical and organisational measures, transparency, together with data subject rights, whereas medium-sized organisations appeared relatively more often in cases concerning technical and organisational measures, employment-related processing, children’s data, together with video surveillance. These findings should be read as exploratory only, given the limited and selective size-coded subset.
4. Discussion
This study examined how Article 5 GDPR appears in national DPA decisions across the EEA together with what those decisions show about recurring forms of non-compliance in supervisory practice. Three points stand out.
First, Article 5-related infringements rarely appeared on their own. In most cases, they were accompanied by findings under other GDPR provisions, especially Article 6 on lawful basis, Articles 12 to 14 on transparency and information and Article 32 on security of processing. That matters because it shows how Article 5 functions in practice. The principles set out there are not treated by supervisory authorities as abstract background values only. They operate as the legal centre of broader constellations of non-compliance. Where a controller cannot identify a valid legal basis, cannot explain the processing to data subjects, or fails to provide adequate safeguards, the same underlying defect often surfaces simultaneously at the level of Article 5 together with at the level of more specific GDPR obligations (
European Union 2016;
Quelle 2018). This finding should, however, be read in light of the structure of the GDPR. Some co-occurrence patterns are legally expected and should not always be understood as separate or independent failures. The clearest example is the relationship between Article 5(1)(a) and Article 6 GDPR. Since Article 6 sets out the conditions under which processing is lawful, a finding that processing lacked a valid legal basis will often also support a finding that the lawfulness element of Article 5(1)(a) was infringed. A similar relationship can be seen between the transparency element of Article 5(1)(a) and the information duties in Articles 12 to 14, between Article 5(1)(f) and Article 32 on security of processing, and between Article 5(2) and Article 24 on controller responsibility. The relevance of the co-occurrence analysis is therefore not that each provision always represents a fully separate failure, but that supervisory authorities often apply Article 5 principles together with more specific GDPR obligations when assessing the same underlying processing deficiency.
Second, the internal distribution of Article 5 findings is not random. The strongest concentration concerned Article 5(1)(a), followed by Article 5(1)(c), Article 5(1)(f), together with Article 5(2). In other words, published supervisory practice most often pointed to deficiencies in lawfulness, fairness and transparency, data minimisation, integrity and confidentiality, together with accountability. That pattern is legally significant. These are the principles that go to the foundations of lawful processing. When they recur at scale in supervisory decisions, the problem is not a technical irregularity at the margins. It suggests that the basic legal conditions of processing are often not being met or not being documented in a defensible way (
Quelle 2018;
Bygrave 2017;
Jasmontaite et al. 2018).
Third, the thematic findings help explain what these principle-level infringements look like in concrete cases. Lawful-basis deficiencies, transparency failures, together with inadequate technical or organisational measures clearly dominated the published decision set. The more detailed Article 6 coding shows that the difficulty was not limited to one type of legal mistake. Some decisions involved no identifiable lawful basis at all. Others involved misuse of consent, legitimate interest, contract, or legal obligation/public task. That matters because it suggests that Article 6-related non-compliance is not simply a matter of choosing the wrong label. In many cases, the legal reasoning supporting the processing appears to have been absent, incomplete, or badly aligned with the actual processing operation described in the decision.
A similar point applies to transparency. The results do not support the narrow view that transparency failures mainly concern badly drafted privacy notices. Failures to inform under Article 13 together with Article 14 formed the largest identifiable transparency subgroups, with additional cases involving general transparency deficiencies, retention-related omissions and insufficient explanation of automated processing or profiling. Taken together, that points to a more basic problem: in many cases, data subjects were not adequately told what was being done with their data, under what conditions, or for how long. In legal terms, this goes directly to the transparency element of Article 5(1)(a) together with to the practical application of Articles 12 to 14 GDPR (
Kaminski 2019;
Hamon et al. 2022). This is consistent with work on information transparency which emphasises that meaningful compliance depends not only on the formal existence of notices, but also on whether information practices make the actual processing operation visible and intelligible to the data subject (
Alić and Luić 2022).
Security-related findings point in the same direction. Article 5(1)(f) appeared frequently, Article 32 was one of the most common co-occurring provisions, and inadequate technical or organisational measures formed one of the largest thematic categories. These findings suggest that, in published supervisory practice, failures of integrity and confidentiality remain a central part of Article 5 enforcement. This is unsurprising from a legal point of view. Security of processing is not external to the GDPR structure; it is one of the conditions of lawful processing itself (
European Union 2016). The dataset does not allow broader claims about the actual prevalence of security failures across all controllers in the EEA, but it does show that, where published decisions exist, security-related deficiencies recur with considerable frequency. This is particularly important in data-intensive and machine-learning environments, where privacy and security risks may include inference, extraction, adversarial manipulation, and other vulnerabilities that go beyond ordinary unauthorised access to data (
Rigaki and Garcia 2023;
Kalodanis et al. 2024;
Vassilev et al. 2025).
The position of accountability is also worth noting. Article 5(2) was among the most frequent Article 5 findings in the sample. That is important because accountability is often described in general terms, but supervisory decisions show what it means in practice: controllers are expected not only to comply, but to be able to show that they complied, on what basis, together with under what safeguards (
Bygrave 2017;
Jasmontaite et al. 2018;
von Grafenstein 2024). In that respect, the published decision record supports an interpretation of Article 5 according to which failures of documentation, justification and traceability are not merely secondary shortcomings. They are often part of the infringement itself. A similar point has been made in the literature on data protection impact assessments, which treats accountability as requiring structured ex ante assessment, documentation, and review of risks rather than mere ex post record keeping (
Kasirzadeh and Clifford 2021).
The sectoral findings add an important qualification. The distribution of main thematic failures differed significantly across sectors, and the same was true of Article 6 subcategories. Technology/media/telecommunications showed a stronger concentration of transparency and technical-measures cases. Public-sector decisions more often involved legal obligation/public task reasoning, technical measures, together with employment-related processing. Health and social care showed a strong concentration of security-related failures and special-category issues. Retail/services stood out for marketing cases. These patterns do not mean that each sector has one single legal problem. They do, however, suggest that Article 5-related non-compliance takes recurrent forms that are shaped by the legal and operational environment in which processing occurs. The supervisory record is therefore not sector-neutral.
Organisation-size findings are less secure and should remain in the background. The size-coded subset was limited, and a large proportion of the dataset could not be coded for size with sufficient confidence. The statistically significant result is still worth noting, but it does not justify broad claims about whether larger or smaller organisations are generally more compliant or less compliant. At most, it suggests that where published infringements occur and size can be identified, the dominant form of non-compliance may differ by organisational scale. That is a much narrower claim, and it is the one the data can support.
The technological coding requires the same caution. Only a small subset of decisions could be classified as AI-explicit or automation/profiling-related. For that reason, this article should not be read as evidence that Article 5 enforcement is mainly an AI-enforcement story. It is not. The point is narrower. The published supervisory record already reveals a stable set of legal weaknesses, including lawful-basis deficiencies, poor transparency, inadequate security measures, excessive or poorly justified data use and weak accountability. These weaknesses remain directly relevant when processing becomes more automated, data-intensive, or opaque. That relevance is legal, not rhetorical. Automated systems do not suspend Article 5; they make its practical application harder and, in some cases, more consequential (
Ufert 2020;
Laux 2024;
European Data Protection Board 2018).
Taken together, the findings support a straightforward conclusion. Article 5 is not merely a symbolic opening provision of the GDPR. In supervisory practice, it functions as a central legal reference point through which broader patterns of unlawful or insufficiently justified processing become visible. Empirical analysis of Article 5-related decisions is therefore useful not only for describing enforcement outcomes, but also for understanding how the GDPR is applied in practice at the level of principle, rule, together with institutional enforcement.
5. Conclusions
This article examined how Article 5 GDPR is reflected in national supervisory authority practice across the EEA through an empirical legal analysis of 790 DPA decisions involving infringements related to Article 5, drawn from an initial GDPRhub retrieval pool of 1660 publicly available decisions issued between 25 May 2018 and 15 September 2025. The findings show that Article 5-related infringements are neither exceptional nor merely incidental. They recur across a large body of published supervisory decisions, most often in relation to lawfulness, fairness and transparency, data minimisation, integrity and confidentiality, together with accountability. In that sense, Article 5 appears not as a peripheral provision, but as a central legal reference point in supervisory practice.
The analysis further shows that Article 5 findings rarely arise in isolation. In most cases, they appear together with infringements of other GDPR provisions, especially Article 6 on lawful basis, Articles 12 to 14 on transparency and information, and Article 32 on security of processing. These co-occurrence patterns should, however, be read in light of the structure of the GDPR, since some combinations, especially Article 5(1)(a) and Article 6, reflect the relationship between principle-level obligations and their more specific operationalisation in the regulation. This does not diminish the relevance of the co-occurrence finding. Rather, it shows that, in published supervisory practice, Article 5 often functions as the normative core of broader compliance failures. Principle-level infringements under Article 5 frequently reveal the same factual or legal deficiencies that also trigger findings under more specific obligations. From that perspective, empirical examination of Article 5-related decisions contributes to a better understanding of how the GDPR is applied in practice, not only at the level of individual provisions, but across the structure of the regulation as a whole (
European Union 2016;
Quelle 2018). The study contributes to the empirical legal literature in three main respects. First, it maps recurrent infringement patterns relating to Article 5 GDPR in a large set of national DPA decisions across the EEA. Second, it clarifies how Article 5 is operationalised in supervisory practice through recurring co-occurrence with provisions concerning lawful basis, transparency, automated decision making, as well as security. Third, it shows that, although the analysed decision set is not predominantly AI-specific, the legal weaknesses repeatedly visible in general Article 5 enforcement remain highly relevant for automated, data-intensive, or otherwise technologically complex processing environments. In this respect, the article does not claim that Article 5 enforcement is mainly an AI-enforcement story. Its narrower point is that supervisory practice already reveals a stable set of legal weaknesses that become especially important where processing is more opaque, scalable, or partially automated (
Ufert 2020;
European Data Protection Board 2018;
Laux 2024).
The findings also carry broader analytical significance. The recurrence of lawful-basis deficiencies, transparency failures, security-related shortcomings, excessive data use and accountability findings suggests that Article 5 remains one of the main pressure points of GDPR compliance in practice. This is consistent with prior scholarship that treats Article 5, together with accountability, as foundational to the wider architecture of the regulation (
Quelle 2018;
Bygrave 2017;
Jasmontaite et al. 2018;
von Grafenstein 2024). The published decision record therefore supports a reading of Article 5 not as a merely symbolic opening provision, but as a legally operative set of standards through which supervisory authorities assess whether personal data processing is structured, justified, limited, secured, and capable of being defended in legal terms.
The study has clear limitations. It is based on publicly available decisions, so the dataset reflects differences in publication practices, enforcement visibility, reporting pathways, together with the structure of the GDPRhub database. The study is based on published supervisory decisions rather than on the full population of GDPR infringements across the EEA. It therefore identifies patterns visible in the available enforcement record, but it does not measure the total incidence of non-compliance. In addition, some contextual variables, especially organisation size or technological context, could not always be determined with precision from the available public record. The AI-relevance coding was deliberately conservative for that reason. These limitations do not negate the value of the findings, but they do require caution in interpretation.
Further research could develop this line of inquiry in several directions. It could refine the coding of automation-related or AI-related cases, distinguish more clearly between model development and deployment contexts, expand sectoral comparison, or examine whether supervisory practice changes as AI-related enforcement becomes more visible. It would also be useful to develop a more complete, harmonised EEA-wide repository of supervisory decisions, since the current fragmentation of publication practices limits comparability, legal certainty, as well as the ability to draw lessons from recurring enforcement patterns. More broadly, the present findings support the value of empirical legal work on supervisory decisions as a means of understanding how core GDPR principles are interpreted, applied, and then enforced in concrete cases across European jurisdictions (
van Dijk et al. 2016;
Novelli et al. 2024;
Papagiannidis et al. 2025).