What is an aggregate and anonymised data addendum?
Is using anonymised customer data legal under POPIA in South Africa?
“This Act does not apply to the processing of personal information … that has been de-identified to the extent that it cannot be re-identified again.”
“"de-identify", in relation to personal information of a data subject, means to delete any information that identifies the data subject; can be used or manipulated by a reasonably foreseeable method to identify the data subject; or links to other information that identifies the data subject.”
“Further processing of personal information must be in accordance or compatible with the purpose for which it was collected.”
When you need a Aggregate & Anonymised Data Addendum
- When you run a SaaS or platform business and want to mine usage patterns across all customers to build benchmarks, industry insights, or product analytics — and need the legal right to do so without breaching POPIA.
- When you intend to train or fine-tune machine-learning or AI models on data derived from customer data, and need that training data to fall outside POPIA so consent is not required for every data subject.
- When you sell or publish aggregated reports, trend data, or "state of the market" insights built from data flowing through your service, and must be sure no customer or individual can be identified in them.
- When a data protection addendum or master services agreement is silent on whether you may use de-identified data for your own purposes — leaving you with no clear right, or a customer arguing you have none.
- When investors or an acquirer perform due diligence and ask whether your analytics or AI data assets were lawfully created — a properly drafted addendum is the document that answers that question.
What a Aggregate & Anonymised Data Addendum should contain
Definition of aggregate and de-identified data
Define the dataset precisely: data derived from customer data that has been irreversibly de-identified (matching POPIA's section 1 test) and aggregated across multiple customers, so that neither any individual data subject nor the customer itself can be identified from it. A tight, POPIA-aligned definition is what takes the data outside the Act — a vague "anonymised data" label does not.
Permitted purposes
Spell out exactly what the provider may do with the de-identified data — for example analytics, statistical analysis, benchmarking, service and security improvement, product development, and training of machine-learning or AI models. Listing the purposes keeps the right contained and lets the customer see, and disclose to its own data subjects, what the data will be used for.
Standard of de-identification (irreversibility)
Commit to de-identifying to the POPIA standard: deleting not only direct identifiers but any information that could, by a reasonably foreseeable method, be used or manipulated to re-identify a data subject, or that links to other identifying information. Specify aggregation and minimum-cohort thresholds so small groups cannot be singled out. This clause is the technical heart of the addendum.
No re-identification undertaking
Bar the provider (and anyone it shares the data with) from attempting to re-identify the de-identified data, or to link it back to a data subject or the customer. POPIA does not criminalise re-identification on its own, so this contractual undertaking — treating any re-identification as a material breach — is precisely what keeps the data reliably outside the Act and preserves the section 6(1)(b) exclusion.
Ownership of the aggregate dataset and resulting insights
State who owns the resulting de-identified, aggregated dataset and the insights, benchmarks, models, and trained AI derived from it — typically the provider, since the data is no longer personal information about the customer or its data subjects. Confirm the customer retains its own underlying customer data and grants no broader rights than this addendum sets out.
Transparency and further-processing alignment
Tie the addendum to the customer-facing privacy notice and terms so the use is disclosed up front, satisfying POPIA's section 15 requirement that further processing be compatible with the original collection purpose. The provider should warrant it will not create de-identified data in a way that conflicts with what data subjects were told at the point of collection.
Survival and continued use after termination
Confirm that the provider may keep and continue using de-identified, aggregated data — and any models or insights already built from it — after the agreement ends, precisely because that data is no longer personal information and is not subject to the deletion-on-termination obligations that apply to the customer's personal information.
Security and onward-sharing controls
Require the de-identified data to be held securely and impose conditions on any onward sharing or publication — for instance that recipients are also barred from re-identification — so that the aggregate dataset cannot drift back into being personal information through careless disclosure or combination with other datasets.
De-identified / aggregate data vs personal information under POPIA
| Feature | Genuinely de-identified, aggregated data | Personal information (incl. weak "anonymisation") |
|---|---|---|
| POPIA applies? | No — excluded by s 6(1)(b) | Yes — full Act applies |
| Lawful basis needed to use | None — not personal information | Consent or another s 11 ground required |
| Can be used for AI training freely | Yes, within the permitted purposes | Only on a lawful basis and compatible purpose |
| Re-identification | Barred by contract; reverses the s 6(1)(b) carve-out if done | N/A — data already identifies the subject |
| Survives termination | Yes — provider may keep and use it | Must usually be deleted or returned |
| Key risk | Anonymisation is reversible, so the carve-out fails | Processing without a lawful basis or notice |
Common South African pitfalls
- Weak or reversible "anonymisation": stripping names but leaving data that can be re-combined to single someone out does not meet POPIA's section 1 test. If the data can be re-identified by a reasonably foreseeable method, it is still personal information, POPIA still applies, and the whole addendum collapses — this is the single biggest risk.
- No bar on re-identification: leaving out an express no-re-identification undertaking is dangerous precisely because POPIA does not criminalise re-identification on its own — the contract is the only discipline keeping the data outside the Act. Without the clause, a downstream party could re-identify the data, dragging it back inside POPIA, and the provider would have no remedy and no defence to the resulting unlawful processing.
- Springing the use on data subjects: creating de-identified data for analytics or AI training that was never disclosed at collection can fail POPIA's section 15 compatibility test for further processing. The addendum must sit on top of a privacy notice and customer terms that already tell data subjects this will happen.
- Over-broad permitted purposes: a clause that lets the provider use "all data for any purpose" invites a finding that the arrangement is really about ongoing personal-information processing dressed up as anonymisation. Keep the purposes defined and tied to genuinely de-identified, aggregated data.
- Confusing aggregation with de-identification: aggregating data does not automatically de-identify it — small cohorts or unique outliers can still identify individuals. Without minimum-cohort thresholds and proper de-identification, "aggregate" reports can leak identity.
- Ownership left silent: if the addendum does not state who owns the resulting dataset, models, and insights, the parties can end up in a dispute precisely over the most valuable asset created — and a customer may argue it never gave the provider the right to build or keep it.
Frequently asked questions
Is anonymised data covered by POPIA in South Africa?
No — provided the anonymisation is genuine. POPIA section 6(1)(b) says the Act does not apply to personal information that has been de-identified to the extent that it cannot be re-identified again. If the data can still be re-identified by a reasonably foreseeable method, it remains personal information and POPIA applies in full.
What does "de-identify" mean under POPIA?
Under section 1 of POPIA, to de-identify means to delete any information that identifies the data subject, that can be used or manipulated by a reasonably foreseeable method to identify them, or that links to other information that identifies them. It is a high bar — removing names is not enough if the remaining data can still single someone out.
Can a SaaS provider use customer data to train AI under POPIA?
Yes, if the training data is genuinely de-identified and aggregated so it cannot be re-identified, because then it falls outside POPIA under section 6(1)(b) and no longer counts as personal information. If the data is still identifiable, the provider needs a lawful basis under POPIA and the use must be compatible with the original collection purpose.
Does POPIA stop a provider re-identifying de-identified data?
Not directly. POPIA does not create a stand-alone criminal offence of re-identifying de-identified information — the Chapter 11 offences do not cover it. But the moment data is re-identified it is personal information again and back inside POPIA, so the real protection is contractual: a well-drafted addendum includes an express undertaking by the provider, and anyone it shares the data with, not to attempt re-identification, breach of which is a material breach.
Who owns the aggregated dataset and the AI models trained on it?
Typically the provider, and the addendum should say so. Because genuinely de-identified, aggregated data is no longer personal information about the customer or its data subjects, the provider can be given ownership of that dataset and of the insights, benchmarks, and models built from it, while the customer keeps its own underlying personal information.
Do I need customer consent to use de-identified data?
Not as a matter of POPIA, if the data is truly de-identified and cannot be re-identified, because it then sits outside the Act. As a matter of transparency and section 15, though, the use should be disclosed in the privacy notice and customer terms so that creating the de-identified data is compatible with the purpose for which the data was originally collected.
Can the provider keep using the aggregate data after the contract ends?
Yes, if the addendum provides for it. Because de-identified, aggregated data is not personal information, it is not subject to the deletion-or-return obligations that apply to the customer's personal information on termination, so the provider can keep and continue using the dataset and any models already built from it.
What is the biggest risk with an anonymised data clause?
Weak anonymisation. If the "de-identified" data can in fact be re-identified by a reasonably foreseeable method, it is still personal information, POPIA applies, and the provider is unlawfully processing it without a lawful basis. The clause only works if the de-identification is genuine, irreversible, and backed by a no-re-identification undertaking.
Sources & authority
- Protection of Personal Information Act 4 of 2013 (POPIA), s 6(1)(b) — application of the Act
- Protection of Personal Information Act 4 of 2013 (POPIA), s 1 — definitions of "de-identify" and "re-identify"
- Protection of Personal Information Act 4 of 2013 (POPIA), s 15 — further processing limitation
This guide is general information, not legal advice. It reflects the law as at June 2026.