CORE PATH Stop 49 / 106

Who Controls the Data About Us?

Data is not merely a record but infrastructure for digital power. This article separates personal from other data, controllers from processors, access from ownership, and examines portability, privacy, concentration and inferred profiles.

“How Do We Measure Whether Decentralisation Is Real?” showed that data is not merely a technical detail but one of the dimensions of real decentralisation. If one organisation holds the only complete copy of critical data, defines its format, controls access and sets the terms of transfer, several apparently independent systems may still depend on the same informational gatekeeper. Whoever controls the data flow often controls not only a record of reality but also the possibilities that can be opened or closed on the basis of that record.

But the phrase “my data” can be misleading. Data is not one single kind of thing, and law does not treat it as a simple object of property. Personal data, data generated by connected devices, business data, public data, anonymised statistics and profiles inferred from behaviour can carry different rules and different rights. This article therefore asks not only who owns the data, but more precisely: who decides why it is collected, who can see it, who can combine it, whether it can be transferred, and who may act on the basis of it.

This is the beginning of THY-REALITY’s digital arc. “How Media Shape What We Think About” examined the selection of information, “How to Evaluate a Source” the evaluation of sources, and “How Do We Measure Whether Decentralisation Is Real?” informational concentration. This article goes one layer lower, into the infrastructure from which later digital systems derive their power: data about people, devices, transactions, locations, habits and relationships. Only then can “Platforms and Algorithms of Attention” ask how platforms and algorithms use that data to shape attention.

Data is not only a record: it is a capacity to observe and act

A row in a database is not power by itself. Power emerges when enough data exists to distinguish between people, predict behaviour, personalise offers, assess risk, filter access or automate decisions. A location history can reveal more than one current location; a sequence of purchases more than one receipt; a network of contacts more than one address book.

That is why context and linkage matter. Two apparently harmless data points can together reveal a sensitive pattern. The European Commission notes that separate pieces of information can constitute personal data when, taken together, they can identify a person. Pseudonymised data remains personal data when re-identification is still possible; only genuinely and irreversibly anonymised data falls outside that regime.

The first lesson of this article is therefore simple: do not ask only what one data point reveals today; ask what it may reveal when combined with other data tomorrow.

Control is broader than legal ownership. In practice it can be divided into several powers: who determines the purpose of collection, who determines the means of processing, who has technical access, who may disclose data to others, who may delete or retain it, who defines the format, and who produces new inferences from it.

The GDPR therefore uses the concept of a controller: the entity that determines why and how personal data is processed. A processor may technically store or process data on the controller’s behalf without independently determining the basic purpose. This legal distinction is also a useful analytical tool: physically hosting the server does not always mean setting the rules, while claiming not to “possess the data” does not eliminate practical power if one actor determines all the critical parameters of the system.

Assessing digital sovereignty therefore requires a map of decision-making and access, not merely a map of servers. Who decides? Who sees? Who can transfer? Who can refuse? Who is accountable?

Personal data is not the same as all data

This article must preserve an important legal boundary. The GDPR protects personal data: information relating to an identified or identifiable living person. It does not govern all industrial, business or machine data in the same way. It would therefore be inaccurate to suggest that an individual has an identical right to every piece of data created near them or through their use of a device.

At the same time, the boundary is not narrow. Device identifiers, location data, online identifiers and combinations of apparently separate information can become personal data when they enable a person to be identified. Pseudonymisation can reduce direct exposure, but it does not by itself make data anonymous.

This avoids two extremes: treating every datum as the private property of an individual, or treating every technical record as non-personal simply because a name is absent. The relevant question is whether the information relates to a person and whether that person can reasonably be identified.

The strongest protection is not always a right to erase information after a system has already collected everything. Often the stronger question comes before collection: is this data necessary at all? The GDPR therefore rests on purpose limitation and data minimisation. An organisation should collect what it needs for a specified and explained purpose rather than accumulating an unlimited reserve for possible future use.

This principle also has systemic value. Every additional field creates another surface for abuse, misinterpretation, breach, accidental disclosure or later purpose drift. A database that does not exist cannot be stolen; a data point the system never needed cannot later become a condition for access.

Digital sovereignty is therefore not only better encryption. It is also a discipline of non-collection: less data, clearer purposes, shorter retention and fewer people or systems with access.

Access and transparency: people need to know what the system knows about them

People cannot easily correct a power imbalance when they do not know which data exists. The GDPR therefore provides rights including information and access: organisations must explain the purposes of processing, categories of data, recipients, retention periods and basic rights. Where data was not collected directly from the person, information about its source can also matter.

But formal transparency is not necessarily practical transparency. Thirty pages of legal language may contain required information while telling an ordinary person almost nothing. This article therefore adds a practical test: can a normal user determine within a few minutes what is collected, why, for how long, with whom it is shared and how they can act?

If the answer requires a specialist, the information asymmetry remains large. Digital sovereignty requires not merely a published policy but an intelligible map of the data relationship.

The GDPR right to data portability allows a person, in defined circumstances, to receive personal data in a machine-readable form and transmit it to another controller. This matters because it can reduce part of the cost of switching. But the right is not a general declaration that every datum about a person is their property or that every internal model or inference must be handed over.

The OECD notes that portability increases user power only when it is operational: the format must be usable, transfer sufficiently easy, and the competing service able to import and use the data effectively. Exporting a file that no other service can meaningfully read is a formal option with little functional value.

Portability can also create new security and privacy risks. More transfer paths mean more opportunities for mistakes and abuse. Well-designed portability therefore reduces lock-in while preserving identity verification, security and clarity about the receiving party.

The biggest gap is often between collected and inferred data

Digital systems store more than what users explicitly tell them. Clicks, locations, purchases, views, typing patterns, social connections and device behaviour can generate inferred profiles: probabilities, segments, recommendations, risk scores and predictions. Such inferences may matter more for decisions than the original raw observations.

A person’s power over their digital position is therefore not complete merely because they can export a list of fields they entered themselves. It also matters to understand what categories and inferences a system produces from behaviour, how long it uses them and in which decisions they have an effect. Where relevant, the GDPR also requires transparency concerning automated decision-making.

This article does not yet open the full question of algorithmic influence—that belongs to “Platforms and Algorithms of Attention”. It establishes the foundation: an algorithm can shape a user’s environment only because it has something from which to infer. Data power is the fuel of algorithmic power.

Data concentration can become competitive advantage and gatekeeping

Data is not always scarce, but its volume, history, combination and feedback loops can create an important advantage. The OECD has noted that data-driven network effects and economies of scale can reinforce market power, especially where a new entrant without historical data struggles to provide an equivalent service.

This connects directly to “Monopoly, Plutocracy and the Concentration of Economic Power” and “How Do We Measure Whether Decentralisation Is Real?”. Even with many applications, the data layer may remain centralised: one provider holds the longest history, one identity system connects all services, one API governs access, or one platform observes both sides of a market. Data concentration is not by itself proof of abuse, but it is a dependency indicator worth measuring.

The practical questions therefore resemble those used for other forms of power: can a new provider enter without privileged data access, can users leave without losing history and functionality, and is there a single point that can cut off access for everyone else?

Connected devices: data is generated even when we do not type it

Cars, smartwatches, industrial machinery, televisions, sensors and other connected products do not produce only a service; they also produce streams of usage and performance data. The EU Data Act, applicable since 12 September 2025, therefore expanded rules concerning access to and use of data generated by connected products and related services.

The practical shift is important: users of a device should not automatically be wholly dependent on the manufacturer for using data generated by that device. The rules expand access and sharing possibilities and also facilitate switching between certain data-processing and cloud services. They do not displace the GDPR; where the data is personal, personal-data protection still applies.

This article takes a broader principle from this: a device you use should not become an invisible one-way data pipe in which everything flows upward while the user has no useful access back.

Privacy, security and availability are not the same thing

Three different goals are often collapsed in data discussions. Privacy asks whether collection and use are appropriate in relation to the person and the purpose. Security asks whether data is protected against unauthorised access, alteration or loss. Availability asks whether an authorised person can obtain and use the data when needed.

A system can be highly secure and still privacy-invasive: it may encrypt data perfectly while collecting far too much. It can be privacy-conscious but badly secured. Or it can protect data so rigidly that users themselves cannot move or reuse it. The NIST Privacy Framework therefore treats privacy as an organisational risk-management problem, not merely a firewall problem.

A sound data architecture must therefore balance these goals: minimum necessary collection, secure processing, intelligible control and enough portability that protection does not become another form of lock-in.

A practical data-sovereignty audit

For an individual, community or organisation, it is useful to examine a data system with the same discipline “How Do We Measure Whether Decentralisation Is Real?” applied to decentralisation. A first audit does not require a complete legal opinion; it requires a map that reveals where power is accumulating.

Use twelve questions: (1) what data is collected, (2) what purpose justifies each category, (3) who is controller and who is processor, (4) who has actual technical access, (5) where data not directly supplied by the person comes from, (6) how long it is retained, (7) to whom it is disclosed, (8) what profiles or inferences are produced, (9) which data the user can receive in a usable format, (10) how much time and cost migration requires, (11) what happens if the largest data provider fails or refuses cooperation, and (12) who can inspect, correct or stop improper processing.

The best outcome is not the slogan “all data local” or “all data open.” Some data should be tightly restricted, some should be portable, and some may benefit from trusted sharing. The more precise test is: is control proportionate to purpose, and are there enough pathways to prevent data from becoming an invisible monopoly over a person’s digital ability to act?

This article began with infrastructure. Before discussing algorithms, recommendations, digital identity or artificial intelligence, we need to know what those systems infer from. Data determines what is visible about a person, what can be compared and which possibilities can be automated.

But data by itself does not determine what a user sees next. Ranking, recommendation and optimisation systems do that. When a platform holds enough behavioural data and simultaneously controls the information flow, it can not only observe preferences but also help shape them.

“Platforms and Algorithms of Attention” will therefore continue with platforms and algorithms of attention. This article leaves behind a foundational rule of digital sovereignty: if we do not understand who collects data, who determines its purpose, who can combine it and how we can exit the system, we do not understand the real distribution of digital power either.

Sources and further reading

  1. European Commission. Principles of personal data processing under the GDPR — lawfulness, transparency, purpose limitation, data minimisation, storage limitation, security and accountability.
  2. European Commission. Information for individuals — GDPR rights including access, rectification, erasure, restriction, portability and objection.
  3. European Commission. Application of the GDPR — personal data, pseudonymisation/anonymisation, controller and processor roles.
  4. European Data Protection Board. Guidelines on the right to data portability under Regulation 2016/679 (WP242 rev.01).
  5. European Data Protection Board. Guidelines 4/2019 on Article 25 — Data Protection by Design and by Default.
  6. European Commission. Data Act — fair access to and use of data; user access to connected-product data and switching between data-processing services; applicable from 12 September 2025.
  7. European Commission. European Data Governance Act — trusted data sharing, data intermediaries and reuse frameworks.
  8. OECD (2024). The impact of data portability on user empowerment, innovation, and competition — portability, interoperability, switching costs, lock-in and privacy/security trade-offs.
  9. OECD. Data governance — mechanisms for user agency, control, trusted data sharing and privacy-enhancing approaches.
  10. OECD (2021). Consumer data and competition — consumer data, privacy, entry barriers and competition in data-driven markets.
  11. OECD (2016). Big Data: Bringing Competition Policy to the Digital Era — data-driven network effects, economies of scale and market power.
  12. NIST. Privacy Framework — voluntary enterprise risk-management framework for identifying and managing privacy risk.