What is ‘personal data’ under EU data protection law?
What counts as "personal data" under GDPR? This guide explains the four building blocks of the definition, why the concept is interpreted broadly, and how identifiers, pseudonymization, biometric data, IP addresses, and context determine whether EU data protection law applies.
Key point: The four “building blocks” of the definition of personal data are: (1) “any information” (2) “relating to” (3) “an identified or indentifiable” (4) “natural person”
The definition of personal data under GDPR is identical to the definition under the 1995 Data Protection Directive.
Article 4 (1) ‘Personal Data’ means any information relating to an identified or identifiable natural person (‘data subject’) […]
The definition contains four closely intertwined building blocks that feed of each other but should be separately analyzed for the sake of clarity. Those blocks are “any information” “relating to” “an identified or indentifiable” “natural person”.
First “building block”: “ANY INFORMATION”

Key points: (1) The concept of personal data includes both objective and subjective statements. (2) For information to be ‘personal data’, it is not necessary that it be true or proven. (3) Personal data covers not only information considered to be “sensitive” but also information that is not sensitive including public information. (4) The concept of personal data includes information available in whatever form. (5) Biometric data can be considered both as content of the information about a particular individual as well as an element to establish a link between one piece of information and the individual.
The term “any information” signals the willingness of the legislator to design a broad concept of personal data and calls for a wide interpretation.
- The concept of personal data includes any sort of statements about a person. It covers not only “objective” information (e.g. the presence of a certain substance in one’s blood) and “subjective” information (e.g. opinions or assessments). Sectors such as banking, insurance and employment process significant volumes of subjective information (e.g. John Doe is “a reliable borrower”, or “not expected to die before he turns 80” or “a good worker who merits promotion”.
- For information to be ‘personal data’, it is not necessary that it be true or proven. In fact, data protection rules already envisage the possibility that information is incorrect and provide for a right of the data subject to access that information and to challenge it through appropriate remedies (1).
- The concept of personal data includes data providing any sort of information. It covers not only information considered to be “sensitive” because of its particularly risky nature, but also more general kinds of information. It includes information touching the individual’s private life “stricto sensu”, but also information regarding whatever types of activity is undertaken by the individual (e.g. working relationships or social behavior) regardless of whether that information is public or not. The European Court of Justice has endorsed this broad approach (2).
Example: Professional habits and practices Drug prescription information (e.g. drug identification number, drug name, drug strength, manufacturer, selling price, new or refill, reasons for use, reasons for no substitution order, prescriber’s first and last name, phone number, etc.), whether in the form of an individual prescription or in the form of patterns discerned from a number of prescriptions, can be considered as personal data about the physician who prescribes this drug, even if the patient is anonymous. Thus, providing information about prescriptions written by identified or identifiable doctors to producers of prescription drugs constitutes a communication of personal data.
Regarding the format or the medium on which the information is contained, the concept of personal data includes information available in whatever form (e.g., alphabetical, numerical, graphical, photographical or acoustic). It includes information kept on paper, as well as information stored in a computer memory by means of binary code, or on a videotape, for instance. This is a logical consequence of covering automatic processing of personal data within its scope. Sound and image data qualify as personal data insofar as they may represent information on an individual.
- It is not necessary for the information to be considered as personal data to be contained in a structured database or file. Also information contained in free text in an electronic document may qualify as personal data, provided the other criteria in the definition of personal data are fulfilled. An e-mail can, for example, contain ‘personal data’.
Example: Telephone Banking: In telephone banking, where the customer’s voice giving instructions to the bank are recorded on tape, those recorded instructions should be considered as personal data.
Example: Videosurveillance: Images of individuals captured by a video surveillance system can be personal data to the extent that the individuals are recognizable.
Example: a child’s drawing: As a result of a neuro-psychiatric test conducted on a girl in the context of a court proceeding about her custody, a drawing made by her representing her family is submitted. The drawing provides information about the girl’s mood and what she feels about different members of her family. As such, it could be considered as being “personal data”. The drawing will indeed reveal information relating to the child (her state of health from a psychiatric point of view) and also about e.g. her father’s or mother’s behaviour. As a result, the parents in that case may be able to exert their right of access on this specific piece of information.
Genetic and biometric data: GDPR defines biometric data and genetic data as follows:
Article 4 (14) : ‘Biometric data” means personal data resulting from specific technical processing relating to the physical, physiological or behavioural characteristics of a natural person, which allow or confirm the unique identification of that natural person, such as facial images or dactyloscopic data.
Article 4 (13) ‘Genetic data” means personal data relating to the inherited or acquired genetic characteristics of a natural person which give unique information about the physiology or the health of that natural person and which result, in particular, from an analysis of a biological sample form the natural person in question.
Typical examples of biometric data include data provided by fingerprints, retinal patterns, facial structure, voices, and also hand geometry, vein patterns or even some deeply ingrained skill or other behavioural characteristic (e.g. handwritten signature, keystrokes, particular way to walk or to speak, etc…).
Biometric data can be considered both as content of the information about a particular individual (e.g. “John Doe has these fingerprints”) as well as an element to establish a link between one piece of information and the individual (e.g. “this object has been touched by someone with these fingerprints and these fingerprints correspond to John Doe; therefore this object has been touched by John Doe”). As such, biometrical elements can work as “identifiers” (e.g. DNA data provides both information about the human body and allows unambiguous and unique identification of a person).
Human tissue samples (like a blood sample) are themselves sources out of which biometric data are extracted, but they are not biometric data themselves (as for instance a pattern for fingerprints is biometric data, but the finger itself is not). Therefore the extraction of information from the samples is collection of personal data, to which the rules of data protection law apply. The collection, storage and use of tissue samples themselves may be subject to separate sets of rules (3).
Second “building block”: “RELATING TO”

Key points: (1) This building block is often overlooked yet it is crucial. (2) Data can “relate” to an individual by virtue of its “content”, “purpose” OR “result” . (3) These three elements (content, purpose, result) must be considered as alternative conditions, and not as cumulative ones. (4) One piece of information may relate to different persons at the same time depending on what element is present
This building block is often overlooked yet is understanding it is crucial to precisely find out which are the relations/links that matter and how to distinguish them. Clearly, information relates to an individual when it is about that individual. In some cases the relationis self evident (e.g. data about an employee in the HR file; the results of a patient’s medical test contained in his medical records, or the image of a person filmed on a video interview of that person). However, where the information is only indirectly connected to the individual determining whether it relates to an individual may require some analysis (e.g. . information concerning objects where those objects belong to someone, or data about processes or events such as information on the functioning of a machine where human intervention is required).
Example: the value of a house: The value of a particular house is information about an object. Data protection rules will clearly not apply when this information will be used solely to illustrate the level of real estate prices in a certain district. However, under certain circumstances such information should also be considered as personal data. Indeed, the house is the asset of an owner, which will hence be used to determine the extent of this person’s obligation to pay some taxes, for instance. In this context, it will be indisputable that such information should be considered as personal data.
Example: car service record: The service register of a car held by a mechanic or garage contains the information about the car, mileage, dates of service checks, technical problems, and material condition. This information is associated in the record with a plate number and an engine number, which in turn can be linked to the owner. Where the garage establishes a connection between the vehicle and the owner, for the purpose of billing, information will “relate” to the owner or to the driver. If the connection is made with the mechanic that worked on the car with the purpose of ascertaining his productivity, this information will also “relate” to the mechanic.
It can be said that data relates to an individual if it refers to the identity, characteristics or behavior of an individual or if such information is used to determine or influence the way in which that person is treated or evaluated. In other words, in order to consider that the data “relate” to an individual, a “content” element OR a “purpose” element OR a “result” element should be present.
- The “content” element is present where, in the light of all circumstances surrounding the case, the information is “about” the person. For example, the results of medical analysis clearly relate to the patient, or the information contained in a company’s folder under the name of a certain client clearly relates to him. In the same way, the information contained in a RFID tag or a bar code incorporated in an identity document of a certain individual relates to that person.
- The “purpose” element exists when, taking into account all of the circumstances surrounding the precise case, the data is used or is likely to be used for the purpose of evaluating, treating in a certain way or influencing the status or behavior of an individual (see, Call log example below).
- A third kind of ‘relating’ to specific persons arises when a “result” element is present. Despite the absence of a “content” or “purpose” element, data can be considered to “relate” to an individual because its use is likely to have an impact on a certain person’s rights and interests, taking into account all the circumstances surrounding the precise case. It is not necessary that the potential result be a major impact. It is sufficient if the individual may be treated differently from other persons as a result of the processing of such data.
Example: call log for a telephone - The call log of a telephone inside a company office provides information about the calls that have been made from that telephone connected to a certain line. That information can be brought into relation with different subjects. On the one hand, the line has been made available to the company, and the company is contractually obliged to pay those calls. The phone set is under the control of a certain employee during working times and calls are supposed to be made by him. The call log may also provide information about the person who was called. The phone can also be used by whatever person is allowed into the premises in the absence of the employee (e.g. by cleaning staff). For different purposes, the information on the use of that phone set can be related to the company, the employee, or the cleaning staff (for instance to check the time cleaning staff leave their workplace, as they are supposed to confirm by phone at what time they leave before locking the premises).
Example: monitoring buses to optimize service having an impact on drivers. A system of satellite location is set up by a transportation agency responsible for buses in a city which makes it possible to determine the position of buses in real time. The purpose is to provide better service by ensuring bus schedules are accurate and assign more or less buses to a route depending on traffic conditions. Strictly speaking the data needed for that system is data relating to the buses, not about the drivers. Yet, the system does allow monitoring the performance of drivers and checking whether they respect speed limits, follow appropriate itineraries, etc. It can therefore have a considerable impact on these individuals, and as such the data may be considered to relate to natural persons and the processing should be subject to data protection rules.
These three elements (content, purpose, result) must be considered as alternative conditions, and not as cumulative ones. In particular, where the content element is present, there is no need for the other elements to be present to consider that the information relates to the individual.
It should be kept in mind that one piece of information may relate to different persons at the same time depending on what element is present with regard to each one. The same information may relate to one individual because of the “content” element (the data is clearly about John Doe), AND to a different individual because of the “purpose” element (it will be used in order to treat Jane Doe in a certain way) AND even to a third individual because of the “result” element (it is likely to have an impact on the rights and interests of a third person).
Example: information contained in the minutes of a meeting. - The minutes of a meeting, typically record the attendance of participants A B and C; the statements made by participants A and B; and a report of proceedings on certain topics as summarized by the author of the minutes C. As personal data relating to A one can only consider the information that he or she attended the meeting at a certain time and place, and that he or she made certain statements. The presence in the meeting of B, his or her statements and the proceedings about an issue as summarized by C are NOT personal data relating to A. This is so even if this information is contained in the same document, and even if it was A who triggered the issue to be discussed at the meeting. It is therefore excluded from A’ right of access to his own personal data. Whether and to what extent that information can be considered as personal data of B and C will have to be determined separately, using the analysis described before.
Third Element: “IDENTIFIED OR IDENTIFIABLE” [NATURAL PERSON]

Key points: (1) Identification is normally achieved through particular pieces of information which we may call “identifiers.” (2) Identification always depends on the circumstances of the case. (3) In order to determine whether a natural person is identifiable, account should be taken of “all the means reasonably likely to be used by the controller of by another person” to achieve identification. (4) EU data protection authorities have taken the position that IP addresses should be treated as personal data. This has been controversial. (5) Under EU data protection law pseudonymised data is personal databut is subject to more flexible data protection rules. (6) Anonymous data (i.e. information relating to a natural person where the person cannot be identified, whether by the data controller or by any other person, taking account of all the means likely reasonably to be used either by the controller or by any other person to identify that individual) is not subject to EU data protection laws.
In general terms:
- a natural person can be considered as “identified” when, within a group of persons, he or she is “distinguished” from all other members of the group
- a natural person is “identifiable” when, although the person has not been identified yet, it is possible to do it (that is the meaning of the suffix “-able”).
GDPR states that:
Article 4(1): “[…] an identifiable natural person is one who can be identified, directly or indirectly, in particular by reference to an identifier such as a name, an identification number, location data, an online identifier or to one or more factors specific to the physical, physiological, genetic, mental, economic, cultural or social identity of that natural person.”
Identification is normally achieved through particular pieces of information which we may call “identifiers” and which hold a particularly privileged and close relationship with the particular individual.
- Examples of identifiers include outward signs of the appearance of this person, like height, hair colour, clothing, etc… or a quality of the person which cannot be immediately perceived, like a profession, a function, a name etc.
The question of whether the individual to whom the information relates is identified or not and the extent to which certain identifiers are sufficient to achieve identification depends on the circumstances of the case:
- A very common family name will not be sufficient to identify someone — i.e. to single someone out — from the whole of a country’s population, while it is likely to achieve identification of a pupil in a classroom.
- Ancillary information, such as “the man wearing a black suit” may identify someone out of the passers-by standing at a traffic light.
A natural person can be “directly or indirectly” identifiable:
- The name of a person is most common direct identifier, and, in practice, the notion of “identified person” implies most often a reference to the person’s name. However, in order to ascertain identity, the name of the person sometimes has to be combined with other pieces of information (date of birth, names of the parents, address or a photograph of the face) to prevent confusion between that person and possible namesakes. For example, the information that a sum of money is owed by John Doe can be considered to relate to an identified individual because it is linked with the name of the person. The name also can be the starting point leading to information about where the person lives or about the person’s family (through the family name).
- The category of indirect identifiers typically relates to the phenomenon of “unique combinations”, whether small or large in size. Some indirect characteristics are so unique that someone can be identified with no effort (“present Prime Minister of Spain”). Others may not allow the identification of an individual in isolation but in combination with other pieces of information (whether the latter is retained by the data controller or not) can allow identification. For example, age category, regional origin, etc may be pretty conclusive in some circumstances, particularly if one has access to additional information of some sort.
Example: Fragmentary information in the press - Information is published about a former criminal case which won much public attention in the past. In the present publication there is none of the traditional identifiers given, especially no name or date of birth of any person involved. It does not seem unreasonably difficult to gain additional information allowing one to find out who the persons mainly involved are, e.g. by looking up newspapers from the relevant time period. Indeed, it can be assumed that it is not completely unlikely that somebody would take such measures (as looking up old newspapers) which would most likely provide names and other identifiers for the persons referred to in the example. It seems therefore justified to consider the information in the given example as being ‘information about identifiable persons’ and as such ‘personal data’.
It is important to note that, while identification through the name is the most common occurrence in practice, a name may itself not be necessary in all cases to identify an individual. Other “identifiers” can be used to single someone out. Indeed, computerized files registering personal data usually assign a unique identifier to the persons registered, in order to avoid confusion between two persons in the file. Also on the Web, web traffic surveillance tools make it easy to identify the behavior of a machine and, behind the machine, that of its user. Thus, the individual’s personality is pieced together in order to attribute certain decisions to him or her. Without even inquiring about the name and address of the individual it is possible to categories this person on the basis of socioeconomic, psychological, philosophical or other criteria and attribute certain decisions to him or her since the individual’s contact point (a computer) no longer necessarily requires the disclosure of his or her identity in the narrow sense. In other words, the possibility of identifying an individual no longer necessarily means the ability to find out his or her name.
The European Court of Justice has spoken in that sense when considering that “referring, on an internet page, to various persons and identifying them by name or by other means, for instance by giving their telephone number or information regarding their working conditions and hobbies, constitutes the processing of personal data […]” (4).
Example: Asylum seekers - Asylum seekers hiding their real names in a sheltering institution have been given a code number for administrative purposes. That number will serve as an identifier, so that different pieces of information concerning the stay of the asylum seeker in the institution will be attached to it, and by means of a photograph or other biometric indicators, the code number will have a close and immediate connection to the physical person, thus allowing him to be distinguished from other asylum seekers and to have attributed to him different pieces of information, which will then refer to an “identified” natural person.
In order to determine whether a natural person is identifiable, account should be taken of “all the means reasonably likely to be used by the controller of by another person” to achieve identification.
Recital 26 of GDPR “[…] To determine whether a natural person is identifiable, account should be taken of all the means reasonably likely to be used, such as singling out, either by the controller or by another person to identify the natural person directly or indirectly. To ascertain whether means are reasonably likely to be used to identify the natural person, account should be taken of all objective factors, such as the costs of and the amount of time required for identification, taking into consideration the available technology at the time of the processing and technological developments.”
The criterion of “all the means likely reasonably to be used either by the controller or by another person” should in particular take into account all the factors at stake.
- The cost of conducting identification is one factor, but not the only one.The way the processing is structured, the advantage expected by the controller, the interests at stake for the individuals, as well as the risk of organisational dysfunctions (e.g. breaches of confidentiality duties) and technical failures should all be taken into account.
- The intended purpose should be considered. A controller cannot successfully argue that data is not personal under EU law because only scattered pieces of information are processed, without reference to a name or any other direct identifiers, when the processing of that information only makes sense if it allows identification of specific individuals and treatment of them in a certain (See Video Surveillance example below). When the purpose of the processing implies the identification of individuals, EU data protection law will assume that the controller or any other person involved have or will have the means “likely reasonably to be used” to identify the data subject. To argue that individuals are not identifiable, where the purpose of the processing is precisely to identify them, is a sheer contradiction in terms under EU data protection law. Where identification of the data subject is not included in the purpose of the processing, putting in place the appropriate state-of-the-art technical and organizational measures to protect the data against identification may make the difference to consider that the persons are not identifiable.
- The test is dynamic and should consider the state of the art in technology at the time of the processing and the possibilities for development during the period for which the data will be processed. Identification may not be possible today with all the means likely reasonably to be used today. If the data are intended to be stored for one month, identification may not be anticipated to be possible during the “lifetime” of the information, and they should not be considered as personal data. However, it they are intended to be kept for 10 years, the controller should consider the possibility of identification that may occur also in the ninth year of their lifetime, and which may make them personal data at that moment.
Example: Publication of X-ray plates together with the patient’s first name - A lady’s X-ray plate had been published in a scientific journal, together with the lady’s first name, which was a very unusual one. The first name of the person, combined by the knowledge by their relatives or acquaintances that she suffered a certain ailment rendered the person identifiable to a number of persons, and the X-ray plate would then be considered as personal data.
Example: Pharmaceutical research data - Hospitals or individual physicians transfer data from medical records of their patients to a company for the purposes of medical research. No names of the patients are used but only serial numbers attributed randomly to each clinical case, in order to ensure coherence and to avoid confusion with information on different patients.. The names of patients stay exclusively in possession of the respective doctors bound by medical secrecy. The data do not contain any additional information which make identification of the patients possible by combining it. In addition, all other measures have been taken to prevent the data subjects from being identified or becoming identifiable, be it legal, technical or organizational. Under these circumstances it is reasonable to consider that no means are present in the processing performed by the pharmaceutical company, which make it likely reasonably to be used to identify the data subjects.
Example: Videosurveillance - Although controllers often argue that identification would only happen in a small percentage of the material collected and that before identification in these few instances actually takes place no personal data is processed, the whole process of video surveillance is considered as processing data about identifiable persons under EU law. This is due to the fact that the purpose of video surveillance is to be able to identify the persons to be seen in the video images in all cases where such identification is deemed necessary by the controller, even if some persons recorded are not identifiable in practice.
Example: damage caused by graffiti - Passenger vehicles owned by a transportation company suffer repeated damage when they are dirtied with graffiti. In order to evaluate the damage and to facilitate the exercise of legal claims against their authors, the company organises a register containing information about the circumstances of the damage, as well as images of the damaged items and of the “tags” or “signature” of the author. At the moment of entering the information into the register, the authors of the damage are not known nor to whom the “signature” corresponds. It may well happen that it will never be known. However, the purpose of the processing is precisely to identify individuals to whom the information relates as the authors of the damage, so as to be able to exercise legal claims against them. Such processing makes sense if the data controller expects as “reasonably likely” that there will one day be means to identify the individual. The information contained in the pictures should be considered as relating to “identifiable” individuals, the information in the register as “personal data”, and the processing should be subject to the data protection rules, which allow such processing as legitimate under certain circumstances and subject to certain safeguards.
IP addresses
EU data protection authorities have taken the position that IP addresses should be treated as data relating to an identifiable person (5). This has been controversial, however:
- Clearly in those cases where the processing of IP addresses is carried out with the purpose of identifying the users of the computer (for instance, by Copyright holders in order to prosecute computer users for violation of intellectual property rights), the controller should anticipate that the “means likely reasonably to be used” to identify the persons will be available (e.g. through the courts), and treat the information as personal data.
- IP addresses under certain circumstances indeed do not allow identification of the user, for various technical and organizational reasons (e.g. addresses attributed to a computer in an internet café, where no identification of the customers is requested). In those cases it is reasonable to take the position that data collected on the use of computer X during a certain timeframe does not allow identification of the user with reasonable means, and therefore it is not personal data. However, because the Internet Service Providers will most probably not know either whether the IP address in question is one allowing identification or not, to be on the safe side, they will likely in practice process it as personal data unless they are in a position to distinguish with absolute certainty that the data correspond to users that cannot be identified.
Pseudonymised data
Pseudonymisation is the process of disguising identities. The aim of such a process is to be able to collect additional data relating to the same individual without having to know his identity. The effectiveness of the pseudonymisation procedure depends on a number of factors (at which stage it is used, how secure it is against reverse tracing, the size of the population in which the individual is concealed, the ability to link individual transactions or records to the same person, etc.). Pseudonymization is specially relevant in the context of research and statistics.
GDPR defines pseudonymization as follows:
Article 4(5) of GDPR ‘pseudonymisation’ means the processing of personal data in such a manner that the personal data can no longer be attributed to a specific data subject without the use of additional information, provided that such additional information is kept separately and is subject to technical and organisational measures to ensure that the personal data are not attributed to an identified or identifiable natural person;
Under EU data protection law, pseudonymised data is considered as information on individuals which are indirectly identifiable. Indeed, using a pseudonym means that it is possible to backtrack to the individual, so that the individual’s identity can be discovered, but then only under predefined circumstances. Although data protection rules apply to pseudonymized data, the risks at stake for the individuals with regard to the processing of such indirectly identifiable information will most often be low, and therefore the application of data protection rules is more flexible.
Key-coded data are a classical example of pseudonimisation. Information relates to individuals that are earmarked by a code, while the key making the correspondence between the code and the common identifiers of the individuals (like name, date of birth, address) is kept separately.
Example: non-aggregated data for statistics - An example to illustrate the importance of taking into account all the circumstances to assess whether the means for identification are “likely reasonably” to be used could be that of personal information processed by the national institute for statistics, where, at a certain stage, the information is kept in non-aggregated form and relates to specific individuals, but these are designated with a code instead of a name (e.g. the individual coded X1234 drinks a glass of wine more than 3 times a week). The institute for statistics keeps separately the key to these codes (the list associating the codes with the names of the persons). That key can be considered to be “likely reasonably to be used” by the institute for statistics, and therefore the set of individual-related information can be considered as personal data and should be subject to the data protection rules by the institute. Now, we can imagine that a list with data about wine drinking habits of consumers is transferred to the national wine-producer organization in order to enable them to back up their public stance by statistical figures.
To determine whether a key coded list of information is still personal data, it should be assessed whether individuals can be identified “taking into account all the means likely reasonably to be used by the controller or another person”.
- If the codes used are unique for each specific person, the risk of identification occurs whenever it is possible to get access to the key used for the encryption. Therefore the risks of an external hack, the likelihood that someone within the sender’s organization — despite his professional secrecy — would provide the key and the feasibility of indirect identification are factors to be taken into account to determine whether the persons can be identified taking into account all the means likely reasonably to be used by the controller or any other person, and therefore whether information should be considered as “personal data”. If they are, the data protection rules will apply. A different question is that those data protection rules could take into account whether risks for the individuals are reduced, and make processing subject to more or less strict conditions, based on the flexibility allowed by the rules of the Directive.
- If, on the contrary, the codes are not unique, but the same code number (e.g. “123”) is used to designate individuals in different towns, and for data from different years (only distinguishing a particular individual within a year and within the sample in the same city), the controller or a third party could only identify a specific individual if they knew to what year and to what town the data refer. If this additional information has disappeared, and it is not likely reasonably to be retrieved, it could be considered that the information does not refer to identifiable individuals and would not be subject to the data protection rules.
This sort of data is commonly used in clinical trials with medicines (6).
Anonymous data
Under EU data protection law anonymous data is information relating to a natural person where the person cannot be identified, whether by the data controller or by any other person, taking account of all the means likely reasonably to be used either by the controller or by any other person to identify that individual.
GDPR does not define “Anonymous data” but it states that:
Recital 26 of GDPR “[…] The principles of data protection should therefore not apply to anonymous information, namely information which does not relate to an identified or identifiable natural person or to personal data rendered anonymous in such a manner that the data subject is not or no longer identifiable. This Regulation does not therefore concern the processing of such anonymous information, including for statistical or research purposes”
The assessment of whether the data allow identification of an individual, and whether the information can be considered as anonymous or not depends on the circumstances, and a case-by-case analysis should be carried out with particular reference to the extent that the means are likely reasonably to be used for identification as described in Recital 26.
This is particularly relevant in the case of statistical information, where despite the fact that the information may be presented as aggregated data, the original sample is not sufficiently large and other pieces of information may enable the identification of individuals.
Example: Statistical surveys and combination of scattered information - Apart from their general obligation to respect data protection rules, in order to ensure anonymity of the statistical surveys, statisticians are subjected to a specific duty of professional secrecy, and under those rules it is forbidden for them to publish non anonymous data. This obliges them to publish aggregated statistical data which cannot possibly be attributed to an identified person behind the statistics. This rule is particularly relevant concerning the publication of census data. In each situation a threshold should be determined under which it is deemed possible to identify the persons concerned. If a criterion appears to lead to identification in a given category of persons, however large (i.e. only one doctor operates in a town of 6000 inhabitants), this “discriminating” criterion should be dropped altogether or other criteria be added to “dilute” the results on a given person so as to allow for statistical secrecy.
Example: Publication of video surveillance - A shopkeeper installs a camera surveillance system in his shop. He publishes, in his shop, the pictures of thieves who have been caught by means of the camera surveillance system. After police intervention, he blanks out the faces of the thieves, by darkening them. However, even after this operation, there still exists a possibility that the persons on the photos can be recognized by their friends, relatives or neighbors, because of the fact that e.g. their figure, haircut and clothes are still recognizable.
Fourth “building block”: “NATURAL PERSON”

Key points: As a general rule, EU data protection law does not apply to: (1) Information related to dead individuals. (2) Information related to unborn babies (unless State Member law has extended protections to them.) (3) Information of legal entities (unless State Member law has extended protections to them) However, if the information in categories (1)-(3) above indirectly identifies natural persons it will be subject to EU data protection laws.
EU data protection law applies to natural persons, that is, to human beings. This means that:
- Information relating to dead individuals is in principle not to be considered as personal data subject to EU data protection law. However, the data of the deceased may still indirectly receive some protection in certain cases. The information on dead individuals may also refer to living persons. For instance, the information that the dead John Doe suffered from hemophilia indicates that his daughter Jane Doe may also suffer from the same disease. Additionally, information on deceased persons may be subject to specific protection granted by sets of rules other than data protection legislation (for example, the obligation of confidentiality of medical staff does not end with the death of the patient. Also, national legislation on the right to one’s own image and honour may grant also protection to the memory of the dead.
- The extent to which data protection rules may apply before birth depends on the general position of national legal systems about the protection of unborn children. To take mainly account of inheritance rights, some Member States acknowledge the principle that children conceived but not yet born are considered as if they were born as far as benefits are concerned (and thus can receive a heritage or accept a donation), subject to the condition that they may effectively be born. In other Member States, specific protection is given by particular legal provisions, also subject to the same condition. To determine whether national data protection provisions protect also information on unborn children, that general approach of the national legal system should be considered, together with the idea that the purpose of data protection rules is to protect the individual (7).
- As a general rule, EU data protection law does not apply to the information of legal persons. However, certain data protection rules may still indirectly apply to information relating to businesses or to legal persons, in a number of circumstances (8). Information about legal persons may also be considered as “relating to” natural persons on their own merits (e.g. where the name of the legal person derives from that of a natural person, or where corporate e-mail is normally used by a certain employee, or information about a small business which may describe the behavior of its owner). In all cases where the criteria of “content”, “purpose” or “result” allow the information on the legal person or on the business to be considered as “relating” to a natural person, under EU data protection law such information should be considered as personal data, and the data protection rules apply. Additionally, as in the case of information on dead people, practical arrangements by the data controller may also result in data on legal person being subject de facto to data protection rules. Where the data controller collects data on natural and legal persons indistinctly and includes them in the same sets of data, the design of the data processing mechanisms and the auditing system may be set up so as to comply with data protection rules. In fact, it may be easier for the controller to apply the data protection rules to all sorts of information in his files than to try to sort out what refers to natural and what to legal persons. Finally, in some cases EU data protection requirements have been extended to legal persons by State Member laws (9).
End notes:
- Rectification could be done by adding contrasting comments or by using the appropriate legal remedies, such as appeal mechanisms of the position or capacity of those persons (as consumer, patient, employee, customer, etc).
- Judgment of the European Court of Justice C-101/2001of 6.11.2003 (Lindqvist), §24: “The term personal data used in Article 3(1) of Directive 95/46 covers, according to the definition in Article 2(a) thereof, any information relating to an identified or identifiable natural person. The term undoubtedly covers the name of a person in conjunction with his telephone coordinates or information about his working conditions or hobbies”.
- See, for example, Council of Europe Recommendation No. Rec (2006) 4 of the Committee of Ministers to Member States on research on biological materials of human origin, of 15.3.2006
- Judgment of the European Court of Justice C-101/2001of 06.11.2003 (Lindqvis), §27
- WP 37/2000 on “Privacy on the Internet — An integrated EU Approach to On-line Data Protection” states that “Internet access providers and managers of local area networks can, using reasonable means, identify Internet users to whom they have attributed IP addresses as they normally systematically “log” in a file the date, time, duration and dynamic IP address given to the Internet user. The same can be said about Internet Service Providers that keep a logbook on the HTTP server. In these cases there is no doubt about the fact that one can talk about personal data in the sense of Article 2 a) of the Directive …)”
- In 2001 Directive 2001/20 of 4 April 2001 on the implementation of good clinical practice and the conduct of clinical trials layed down a legal framework for the pursuit of these activities. The medical professional/researcher (“investigator”) testing the medicines collects the information about clinical results on each patient, earmarking him with a code. The researcher provides the information to the pharmaceutical company or other parties involved (“sponsors”) only in this coded form, as they are only interested in biostatistical information. However, the investigator keeps separately a key associating this code with common information to identify the patients in a separate way. To protect the health of patients in case the medicines turn out to pose dangers, the investigator is obliged to keep this key, so that individual patients may be identified in case of need and receive appropriate treatment. The question here is whether the data used for the clinical trial can be considered to relate to “identifiable” natural persons and thus be subject to the data protection rules. According to the analysis described before, to determine whether a person is identifiable account should be taken of all the means likely reasonably to be used either by the controller or by any other person to identify the said person. In this case, the identification of individuals (to apply the appropriate treatment in case of need) is one of the purposes of the processing of the key-coded data. The pharmaceutical company has construed the means for the processing, included the organizational measures and its relations with the researcher who holds the key in such a way that the identification of individuals is not only something that may happen, but rather as something that must happen under certain circumstances. The identification of patients is thus embedded in the purposes and the means of the processing. In this case, one can conclude that such key-coded data constitutes information relating to identifiable natural persons for all parties that might be involved in the possible identification and should be subject to the rules of data protection legislation. This does not mean, though, that any other data controller processing the same set of coded data would be processing personal data, if within the specific scheme in which those other controllers are operating re-identification is explicitly excluded and appropriate technical measures have been taken in this respect. In other areas of research or of the same project, re-identification of the data subject may have been excluded in the design of protocols and procedure, for instance because there is no therapeutical aspects involved. For technical or other reasons, there may still be a way to find out to what persons correspond what clinical data, but the identification is not supposed or expected to take place under any circumstance, and appropriate technical measures (e.g. cartographic, irreversible hashing) have been put in place to prevent that from happening. In this case, even if identification of certain data subjects may take place despite all those protocols and measures (due to unforeseeable circumstances such as accidental matching of qualities of the data subject that reveal his/her identity ), the information processed by the original controller may not be considered to relate to identified or identifiable individuals taking account of all the means likely reasonably to be used by the controller or by any other person. Its processing may thus not be subject to the provisions of the Directive. A different matter is that for the new controller who has effectively gained access to the identifiable information, it will undoubtedly be considered to be “personal data”.
- The consideration that the legal system’s general response relies on the expectation that the situation of unborn children is typically limited in time to the period of pregnancy. It may not take account of the fact that this situation may actually last considerably longer, as in the case of frozen embryos. Specific legal responses may be found in particular provisions on reproduction techniques, dealing with the use of medical or genetic information about embryos.
- Some provisions of the e-privacy Directive 2002/58/EC extend to legal persons. Article 1 thereof provides that “2. The provisions of this Directive particularize and complement Directive 95/46/EC for the purposes mentioned in paragraph 1. Moreover, they provide for protection of the legitimate interests of subscribers who are legal persons.” Accordingly, Articles 12 and 13 extend the application of some provisions concerning directories of subscribers and unsolicited communication also to legal persons.The European Court of Justice made clear that nothing prevented the Member States from extending the scope of the national legislation implementing the provisions of the 95 Directive to areas not included within the scope thereof, provided that no other provision of community law precludes it (see, Judgment of the European Court of Justice C-101/2001 of 06.11.2003 (Lindqvis), § 98). Accordingly some Member States such as Italy, Austria or Luxembourg extended the application of certain provisions of national law adopted pursuant to the Directive (such as those on security measures) to the processing of data on legal persons.
Resources
Articles and Recitals:
- GDPR Article 4(1) (Concept of Personal Data
- GDPR Recitals 26, 27, and 30 (Concept of Personal Data)
- GDPR Article 4(5), 6(4)(e), 25(1), 32(1)(a) (Pseudo-Anonymization)
- GDPR Recitals 26, 28 and 29 (Pseudo-Anonymization)
- GDPR Article 4(14) (Biometric Data)
- GDPR Article 4(13) (Genetic Data)
Regulator Guides
- Article 29 WP opinion 4/2007 on the concept of Personal Data
- Article 29 WP Document 37/2000: Privacy on the Internet — An integrated EU Approach to On-line Data Protection
Anonymization
- AEPD Guide on K-anonymization (In Spanish)
- Irish DPA guide on anonymization: Anonymisation and Pseudonymisation: Full Guidance Note
