Search legal guides

Search MJ Kotze Inc legal guides and articles

Data, Privacy & Website

Aggregate & Anonymised Data Addendum in South Africa

The addendum that lets your SaaS or platform use aggregated, de-identified customer data for analytics and AI training — drafted so the data genuinely sits outside POPIA, not just nominally.

Written by

Martin Kotze

Attorney, Conveyancer & Notary Public

Last reviewed:

Quick answer

What is an aggregate and anonymised data addendum?

An aggregate and anonymised data addendum is a clause set, usually bolted onto a SaaS, platform, or services agreement (often alongside the data protection addendum), that gives the provider the right to take the personal information it processes for the customer, strip out everything that identifies any individual, combine it across many customers, and use the resulting aggregated, de-identified dataset for its own purposes — typically analytics, benchmarking, performance reporting, product and feature improvement, and training machine-learning or AI models. The commercial point is straightforward: a provider that serves hundreds of customers sits on enormously valuable patterns, but it can only exploit them lawfully if the data it uses no longer counts as personal information. Under the Protection of Personal Information Act 4 of 2013 (POPIA), information that has been genuinely de-identified — so that it cannot be re-identified again — is carved out of the Act entirely. The addendum is the contractual machinery that captures that right, defines what the provider may do with the de-identified data, fixes who owns it, and (critically) commits the provider not to re-identify it. Done properly, the dataset is no longer about any data subject and POPIA falls away; done badly, the "anonymised" data is still personal information and the entire arrangement is unlawful processing.

Is using anonymised customer data legal under POPIA in South Africa?

Yes — but only if the data is genuinely and irreversibly de-identified. POPIA section 6(1)(b) provides that the Act does not apply to the processing of personal information that "has been de-identified to the extent that it cannot be re-identified again". That single exclusion is the entire legal foundation for an aggregate and anonymised data addendum: once the data sits outside POPIA, the provider can use it for analytics, benchmarking and AI training without needing the data subject's consent or another lawful basis, because the data is no longer personal information at all. The catch is in POPIA's own definition of "de-identify" in section 1 of POPIA: to de-identify means to delete any information that identifies the data subject, that can be used or manipulated by a reasonably foreseeable method to identify the data subject, or that links to other information that identifies the data subject. The bar is therefore high — masking names is not enough if the remaining data can be re-combined to single someone out. POPIA also defines "re-identify" in section 1 as resurrecting information that has been de-identified; but, importantly, POPIA does not create a stand-alone criminal offence of re-identifying de-identified data — the offences in Chapter 11 of the Act (sections 100 to 109) do not include re-identification. The real consequence of weak or reversed de-identification is that the data was never outside the Act in the first place, so the provider is left processing personal information without a lawful basis and exposed to the Information Regulator's enforcement and administrative fines (up to R10 million). That is why the no-re-identification discipline has to be built into the contract: the addendum, not the statute, is what binds the provider not to reverse the anonymisation. And where de-identified data is later created from data originally collected for a different purpose, section 15 requires that any further processing be compatible with the purpose for which it was first collected — so the addendum should sit on top of a privacy notice and customer terms that already disclose this use, rather than springing it on data subjects. In short, the law fully supports using anonymised, aggregate data — provided the anonymisation is real, the provider is contractually barred from reversing it, and the original collection was transparent.
This Act does not apply to the processing of personal information … that has been de-identified to the extent that it cannot be re-identified again.
Protection of Personal Information Act 4 of 2013 (POPIA), s 6(1)(b)
"de-identify", in relation to personal information of a data subject, means to delete any information that identifies the data subject; can be used or manipulated by a reasonably foreseeable method to identify the data subject; or links to other information that identifies the data subject.
Protection of Personal Information Act 4 of 2013 (POPIA), s 1 — definition of "de-identify"
Further processing of personal information must be in accordance or compatible with the purpose for which it was collected.
Protection of Personal Information Act 4 of 2013 (POPIA), s 15 — further processing

When you need a Aggregate & Anonymised Data Addendum

  • When you run a SaaS or platform business and want to mine usage patterns across all customers to build benchmarks, industry insights, or product analytics — and need the legal right to do so without breaching POPIA.
  • When you intend to train or fine-tune machine-learning or AI models on data derived from customer data, and need that training data to fall outside POPIA so consent is not required for every data subject.
  • When you sell or publish aggregated reports, trend data, or "state of the market" insights built from data flowing through your service, and must be sure no customer or individual can be identified in them.
  • When a data protection addendum or master services agreement is silent on whether you may use de-identified data for your own purposes — leaving you with no clear right, or a customer arguing you have none.
  • When investors or an acquirer perform due diligence and ask whether your analytics or AI data assets were lawfully created — a properly drafted addendum is the document that answers that question.

What a Aggregate & Anonymised Data Addendum should contain

1

Definition of aggregate and de-identified data

Define the dataset precisely: data derived from customer data that has been irreversibly de-identified (matching POPIA's section 1 test) and aggregated across multiple customers, so that neither any individual data subject nor the customer itself can be identified from it. A tight, POPIA-aligned definition is what takes the data outside the Act — a vague "anonymised data" label does not.

2

Permitted purposes

Spell out exactly what the provider may do with the de-identified data — for example analytics, statistical analysis, benchmarking, service and security improvement, product development, and training of machine-learning or AI models. Listing the purposes keeps the right contained and lets the customer see, and disclose to its own data subjects, what the data will be used for.

3

Standard of de-identification (irreversibility)

Commit to de-identifying to the POPIA standard: deleting not only direct identifiers but any information that could, by a reasonably foreseeable method, be used or manipulated to re-identify a data subject, or that links to other identifying information. Specify aggregation and minimum-cohort thresholds so small groups cannot be singled out. This clause is the technical heart of the addendum.

4

No re-identification undertaking

Bar the provider (and anyone it shares the data with) from attempting to re-identify the de-identified data, or to link it back to a data subject or the customer. POPIA does not criminalise re-identification on its own, so this contractual undertaking — treating any re-identification as a material breach — is precisely what keeps the data reliably outside the Act and preserves the section 6(1)(b) exclusion.

5

Ownership of the aggregate dataset and resulting insights

State who owns the resulting de-identified, aggregated dataset and the insights, benchmarks, models, and trained AI derived from it — typically the provider, since the data is no longer personal information about the customer or its data subjects. Confirm the customer retains its own underlying customer data and grants no broader rights than this addendum sets out.

6

Transparency and further-processing alignment

Tie the addendum to the customer-facing privacy notice and terms so the use is disclosed up front, satisfying POPIA's section 15 requirement that further processing be compatible with the original collection purpose. The provider should warrant it will not create de-identified data in a way that conflicts with what data subjects were told at the point of collection.

7

Survival and continued use after termination

Confirm that the provider may keep and continue using de-identified, aggregated data — and any models or insights already built from it — after the agreement ends, precisely because that data is no longer personal information and is not subject to the deletion-on-termination obligations that apply to the customer's personal information.

8

Security and onward-sharing controls

Require the de-identified data to be held securely and impose conditions on any onward sharing or publication — for instance that recipients are also barred from re-identification — so that the aggregate dataset cannot drift back into being personal information through careless disclosure or combination with other datasets.

De-identified / aggregate data vs personal information under POPIA

FeatureGenuinely de-identified, aggregated dataPersonal information (incl. weak "anonymisation")
POPIA applies?No — excluded by s 6(1)(b)Yes — full Act applies
Lawful basis needed to useNone — not personal informationConsent or another s 11 ground required
Can be used for AI training freelyYes, within the permitted purposesOnly on a lawful basis and compatible purpose
Re-identificationBarred by contract; reverses the s 6(1)(b) carve-out if doneN/A — data already identifies the subject
Survives terminationYes — provider may keep and use itMust usually be deleted or returned
Key riskAnonymisation is reversible, so the carve-out failsProcessing without a lawful basis or notice

Common South African pitfalls

  • Weak or reversible "anonymisation": stripping names but leaving data that can be re-combined to single someone out does not meet POPIA's section 1 test. If the data can be re-identified by a reasonably foreseeable method, it is still personal information, POPIA still applies, and the whole addendum collapses — this is the single biggest risk.
  • No bar on re-identification: leaving out an express no-re-identification undertaking is dangerous precisely because POPIA does not criminalise re-identification on its own — the contract is the only discipline keeping the data outside the Act. Without the clause, a downstream party could re-identify the data, dragging it back inside POPIA, and the provider would have no remedy and no defence to the resulting unlawful processing.
  • Springing the use on data subjects: creating de-identified data for analytics or AI training that was never disclosed at collection can fail POPIA's section 15 compatibility test for further processing. The addendum must sit on top of a privacy notice and customer terms that already tell data subjects this will happen.
  • Over-broad permitted purposes: a clause that lets the provider use "all data for any purpose" invites a finding that the arrangement is really about ongoing personal-information processing dressed up as anonymisation. Keep the purposes defined and tied to genuinely de-identified, aggregated data.
  • Confusing aggregation with de-identification: aggregating data does not automatically de-identify it — small cohorts or unique outliers can still identify individuals. Without minimum-cohort thresholds and proper de-identification, "aggregate" reports can leak identity.
  • Ownership left silent: if the addendum does not state who owns the resulting dataset, models, and insights, the parties can end up in a dispute precisely over the most valuable asset created — and a customer may argue it never gave the provider the right to build or keep it.

Frequently asked questions

Is anonymised data covered by POPIA in South Africa?

No — provided the anonymisation is genuine. POPIA section 6(1)(b) says the Act does not apply to personal information that has been de-identified to the extent that it cannot be re-identified again. If the data can still be re-identified by a reasonably foreseeable method, it remains personal information and POPIA applies in full.

What does "de-identify" mean under POPIA?

Under section 1 of POPIA, to de-identify means to delete any information that identifies the data subject, that can be used or manipulated by a reasonably foreseeable method to identify them, or that links to other information that identifies them. It is a high bar — removing names is not enough if the remaining data can still single someone out.

Can a SaaS provider use customer data to train AI under POPIA?

Yes, if the training data is genuinely de-identified and aggregated so it cannot be re-identified, because then it falls outside POPIA under section 6(1)(b) and no longer counts as personal information. If the data is still identifiable, the provider needs a lawful basis under POPIA and the use must be compatible with the original collection purpose.

Does POPIA stop a provider re-identifying de-identified data?

Not directly. POPIA does not create a stand-alone criminal offence of re-identifying de-identified information — the Chapter 11 offences do not cover it. But the moment data is re-identified it is personal information again and back inside POPIA, so the real protection is contractual: a well-drafted addendum includes an express undertaking by the provider, and anyone it shares the data with, not to attempt re-identification, breach of which is a material breach.

Who owns the aggregated dataset and the AI models trained on it?

Typically the provider, and the addendum should say so. Because genuinely de-identified, aggregated data is no longer personal information about the customer or its data subjects, the provider can be given ownership of that dataset and of the insights, benchmarks, and models built from it, while the customer keeps its own underlying personal information.

Do I need customer consent to use de-identified data?

Not as a matter of POPIA, if the data is truly de-identified and cannot be re-identified, because it then sits outside the Act. As a matter of transparency and section 15, though, the use should be disclosed in the privacy notice and customer terms so that creating the de-identified data is compatible with the purpose for which the data was originally collected.

Can the provider keep using the aggregate data after the contract ends?

Yes, if the addendum provides for it. Because de-identified, aggregated data is not personal information, it is not subject to the deletion-or-return obligations that apply to the customer's personal information on termination, so the provider can keep and continue using the dataset and any models already built from it.

What is the biggest risk with an anonymised data clause?

Weak anonymisation. If the "de-identified" data can in fact be re-identified by a reasonably foreseeable method, it is still personal information, POPIA applies, and the provider is unlawfully processing it without a lawful basis. The clause only works if the de-identification is genuine, irreversible, and backed by a no-re-identification undertaking.

Sources & authority

This guide is general information, not legal advice. It reflects the law as at June 2026.

Get your Aggregate & Anonymised Data Addendum reviewed or drafted

Upload an existing document for a fixed-fee review, or have a bespoke Aggregate & Anonymised Data Addendum drafted for your business — personally, by a senior corporate and commercial attorney. No obligation to proceed.

Review: Fixed fee from R8 325 (excl. VAT) · 48-hour turnaroundDraft: Fixed fee from R8 175 (excl. VAT)

For the businesses we act for

The Keystone Workspace

The attorney-designed platform the businesses we act for use to run their contracts, e-signatures and company secretarial work in one place.

Why you can trust this: Martin Kotze has been an admitted Attorney of the High Court of South Africa, registered Conveyancer, and Notary Public since 2014, practising from Pretoria. The firm is regulated by the Legal Practice Council under firm registration 17444.

This guide is general information, not legal advice for your specific matter.