What are the three questions?
Which provider, under which agreement, in which region. Every other data question about an AI system resolves into one of those three, and all three are answerable in writing before the first call is made.
They are worth asking in that order, because the answers narrow. The provider determines which agreements exist. The agreement determines what may be done with the data. The region determines which law reaches it.
The consumer tier is a different contract
The same provider, the same model, the same interface, and a different agreement underneath. Consumer tiers commonly train on inputs by default and offer no deletion right. Enterprise tiers commonly do neither, and say so in the terms.
An employee using a personal account is therefore not a variation of approved use. It is data leaving under terms nobody read, and it will not appear on a company card or in the identity provider, which is why exposure assessments start with four discovery sources rather than one.
“Same provider, same model, same screen. A different agreement underneath, and a different answer about training.”
on why the tier matters more than the vendorWhat has to be fixed in writing?
| Term | What to establish | Why it is asked |
|---|---|---|
| Training | Whether inputs or outputs train any model | The single clause most often different between tiers |
| Retention | How long prompts and responses are held, and by whom | Abuse-monitoring retention is separate from training |
| Sub-processors | Who else touches the data, and where they sit | A named region can still route through another |
| Residency | Which region processes and which stores | Processing and storage are frequently not the same |
| Deletion | What can be deleted, how, and how quickly | The PDPA obligation lands on the organisation |
Residency deserves the closest reading. A provider may store in Singapore and process elsewhere, or offer regional processing on some models and not others. The commitment that matters is the one that covers the model the system actually calls.
Private deployment answers it differently
Where the terms cannot be made acceptable, the model runs inside the organisation’s own tenancy and the question changes from what the provider may do to what the organisation has configured. That is a heavier build with its own operating cost, and it is the right answer for a narrower set of cases than it is usually proposed for.
The data class decides it. Where the exposure is regulated personal data or material non-public information, private deployment earns its cost. Where it is internal documentation, an enterprise agreement usually covers it.
What do we do about it?
The three questions are settled in the same session that agrees the error rate, and the answers are written into the build plan rather than assumed from a vendor page. Every model call routes through a gateway, so the provider, the region and the agreement in force are recorded per request rather than remembered.
When a provider changes terms, and they do, the record shows what was in force at the time each request was made.
Article Published 1 August 2026 · By Agile Labs
Agile Labs is a Singapore software engineering company. Since 2016 we have built, taken over, secured and maintained software and AI systems.
Sources
- Personal Data Protection Commission Singapore, Advisory Guidelines on the Use of Personal Data in AI Recommendation and Decision Systems, March 2024.
- OpenAI, Anthropic and Google enterprise terms, training and retention clauses as published.
- MAS, Consultation Paper on Guidelines on Artificial Intelligence Risk Management, third-party and outsourcing sections.
