Singapore - Synthetic Data Generation Guide
Proposed Guide on Synthetic Data Generation (PDPC)
Singapore
RAI-SG-NA-PGSDGXX-2024The Personal Data Protection Commission (PDPC) of Singapore published the Proposed Guide on Synthetic Data Generation on 15 July 2024 to explain the uses, risks and good practices for generating structured synthetic data as a Privacy Enhancing Technology (PET). Jointly developed with A*STAR and supported by IMDA, the Guide provides a five‑step approach, case archetypes, and an annexed handbook of technical and governance controls to mitigate re‑identification risks.
Summary
Read full text ↗Plain English
Overview
The PDPC Proposed Guide on Synthetic Data Generation (15 July 2024) sets out good practices for producing structured synthetic data as a Privacy Enhancing Technology (PET). The Guide was jointly developed with the Agency for Science, Technology and Research (A*STAR) and is intended to be available as a resource within the IMDA PET Sandbox. It frames synthetic data as artificial data generated by purpose‑built mathematical or ML models that can mimic statistical properties of source datasets while warning that synthetic output may still pose re‑identification risks. The Guide focuses on fully synthetic structured data, provides three common use case archetypes (AI training, data sharing/analysis, software testing), and offers a five‑step process and annexed handbook to guide organisations in technical, governance and contractual controls to mitigate residual risks. It emphasises a risk‑based approach, recognising trade‑offs between utility and privacy, and positions the document as a living guide to be updated as the technology and evaluation methods evolve.
Definitions
The Guide defines key terms used throughout: "synthetic data" as artificially generated data created by mathematical or ML models trained on source data; "fully synthetic" (no real records preserved) and "partially synthetic" (some original values retained) distinctions; "re‑identification risk" meaning the probability an individual in source data can be matched to synthetic records; and "Privacy Enhancing Technology (PET)" covering data obfuscation, encrypted computation and federated analytics. The Guide focuses on structured tabular data and clarifies that while high‑quality synthetic data can preserve statistical properties and utility, it is not automatically non‑personal data and still requires privacy risk assessment.
Governance and Institutional Framework
The Guide outlines governance expectations for organisations that generate synthetic data, including role‑based accountability (CIO/CTO/CDO/Data Protection Officer), documented policies, and cross‑functional oversight. It recommends integrating synthetic data governance into existing data protection management systems and PDPA compliance programmes, and using contractual safeguards when engaging external vendors or cloud providers. The PDPC suggests organisations adopt explicit project charters specifying purpose, legal basis, data sources, expected end‑uses and retention timelines. Where appropriate, organisations are encouraged to engage regulators and the IMDA PET Sandbox (IMDA PET Sandbox) for early regulatory feedback. The Guide also recommends periodic board reporting for higher‑risk projects and inclusion of synthetic data projects in enterprise risk registers and privacy impact assessments.
Key Focus Areas
The Guide’s technical and process focal points include: (1) data understanding and minimisation—identify sensitive fields, remove outliers and restrict granularity where feasible; (2) synthetic generation choices—select generation methods (statistical, model‑based, GANs, probabilistic models) appropriate to data types and risk tolerance; (3) privacy controls—apply pseudonymisation on source data, inject calibrated noise, and consider differential privacy for metrics or model training where provable bounds are needed; (4) utility validation—define and measure downstream performance metrics to ensure synthetic data supports intended tasks without overfitting to real individuals; (5) re‑identification testing—use membership, attribute and record linkage tests and threshold‑based risk criteria described in Annex D and Annex E; and (6) contractual and operational controls—access restrictions, purpose limitation clauses, logging, and incident response integration with existing PDPA obligations. The Guide emphasises iterative trade‑off analysis between utility and disclosure risk and recommends conservative deployments in high‑sensitivity domains (e.g., healthcare, finance) unless robust mitigation is in place.
Implementation Framework
The Guide offers a practical five‑step implementation flow: Step 1—Know your data: catalogue fields, provenance and sensitivity; Step 2—Prepare data: pseudonymise/anonymise and remove outliers, apply minimisation; Step 3—Generate synthetic data: choose model architecture and training protocols, incorporate noise or constraints as needed; Step 4—Assess re‑identification risks: apply quantitative and qualitative tests from Annexes D/E and compare to predefined risk thresholds; Step 5—Manage residual risk: apply contractual, organisational and technical safeguards, maintain documentation and monitoring. Each step includes recommended artifacts (data dictionaries, model cards, risk assessment reports) and acceptance criteria for release. The Guide further provides a sample data dictionary format (Annex B) and examples of generation methods (Annex C) to aid implementation teams.
Monitoring and Evaluation
The Guide requires ongoing monitoring of synthetic data projects: scheduled re‑identification testing, validation of downstream model performance, incident detection and reporting, and periodic review of assumptions (e.g., changes in linkable external datasets that might raise re‑identification risk). It recommends maintaining logs of generation runs, model versions and random seeds where feasible for reproducibility and audit. The Guide encourages organisations using synthetic data in production to include synthetic data metrics in their privacy KPI dashboards, and to re‑evaluate risk when dataset composition, use cases or external threat models change. Where possible, results from the IMDA PET Sandbox pilots and public case studies should inform refinements.
Penalties, Liability, and Appeals
Although the Guide itself is non‑binding, organisations are reminded that synthetic data generation activities remain subject to the PDPA and other sectoral rules. PDPC enforcement powers under the PDPA include directions and financial penalties; since 1 October 2022 the PDPA regime allows penalties of up to 10% of an organisation's annual turnover in Singapore (for qualifying thresholds) or SGD 1,000,000 (whichever is higher) depending on circumstances. Organisations are therefore encouraged to document risk assessments, follow the Guide’s practices, and consult PDPC where novel or high‑risk designs are proposed. The Guide also notes that civil liability or sectoral regulatory sanctions may arise from misuse of data or harms traceable to synthetic outputs; it recommends clear allocation of contractual liability and dispute/appeal mechanisms in supplier contracts.
Relationship to Other Instruments
The Guide complements and builds on existing PDPC materials such as advisory guidelines on AI use and prior PDPC technology guides (e.g., anonymisation guidance). It is designed to be used alongside IMDA’s PET Sandbox resources and sectoral rules (e.g., health, finance) administered by agencies like the Ministry of Health and the Monetary Authority of Singapore where more stringent controls may apply. The Guide references OECD/PDP research and international best practices on PETs and situates synthetic data guidance within the PDPA accountability framework, recommending alignment of synthetic data project governance with organisational PDPA accountability obligations.
International Alignment
The Guide acknowledges analogous instruments internationally — for example, PET guidance and synthetic data research emerging from OECD, EU data protection authorities, and standards bodies — and seeks alignment on terminology and evaluation approaches. PDPC suggests interoperability with international standards and encourages use of internationally recognised privacy techniques (e.g., differential privacy) and measurable re‑identification tests to facilitate cross‑border collaborations while respecting local legal constraints. Through IMDA’s sandbox and international engagement, Singapore aims to harmonise practices to support data sharing, research collaboration and trade while maintaining strong privacy safeguards.
Implementation Timeline
| Milestone | Date / Window |
|---|---|
| Guide published (proposed) | 15 July 2024 |
| Inclusion in IMDA PET Sandbox resources | Q3–Q4 2024 (ongoing) |
| PDPC updates / living document reviews | Periodic review — next review indicated 2024–2025 (no fixed date) |
Compliance Checklist
| Action | Acceptance Criteria |
|---|---|
| Purpose & scope documented | Project charter with stated lawful basis and use cases |
| Data mapping and sensitivity assessment | Data dictionary (Annex B) completed |
| Pre‑generation mitigation | Outliers removed / pseudonymisation applied where appropriate |
| Generation method selected | Model card and rationale recorded |
| Re‑identification testing completed | Risk metrics within agreed thresholds |
| Residual risk management | Contractual, access and monitoring controls implemented |
Sources and References
| Source | Type |
|---|---|
| PDPC, Proposed Guide on Synthetic Data Generation (PDF) | Primary Source |
| PDPC, Proposed Guide on Synthetic Data Generation (web summary) | Primary Source |
| IMDA, Privacy Enhancing Technology Sandboxes | Primary Source |
The Singapore Personal Data Protection Commission (PDPC) has issued a Proposed Guide on Synthetic Data Generation, offering best practices for organisations creating artificial data to ensure privacy while maintaining utility. This guide applies to any organisation in Singapore that generates structured synthetic data, whether for AI training, data sharing, analysis, or software testing. While the guide itself is not legally binding, it clarifies how existing data protection laws, specifically the Personal Data Protection Act (PDPA), apply to these activities.
Organisations must adopt a risk-based approach, following a five-step process that includes: - Thoroughly understanding and minimising sensitive data before generation. - Choosing appropriate generation methods and applying privacy controls like pseudonymisation or noise injection. - Rigorously testing for re-identification risks to ensure synthetic data doesn't inadvertently expose real individuals. - Implementing strong governance, including clear accountability, documented policies, and contractual safeguards with vendors. - Continuously monitoring synthetic data projects for ongoing risks and performance.
The guide was published on July 15, 2024, and is currently a proposed document, with ongoing integration into the IMDA Privacy Enhancing Technology Sandbox resources expected in late 2024. It will be periodically reviewed and updated.
Although the guide is advisory, failing to manage privacy risks in synthetic data generation can lead to significant penalties under the PDPA. These can include financial penalties of up to 10% of an organisation's annual turnover in Singapore or S$1,000,000, whichever is higher, in addition to potential civil liability. A crucial takeaway is that synthetic data is *not* automatically considered non-personal data. Organisations must still conduct thorough privacy risk assessments, as even artificially generated data can carry re-identification risks if not properly managed.
Plain-English rewrite by Regulations.ai — not legal advice. Verify against the official text.
What you must do — compliance checklist
0 / 14 marked completePlain-English obligations under Singapore - Synthetic Data Generation Guide. Not legal advice — verify against the official text before relying on it.
- #1CriticalPenalties, Liability, and Appeals⏰ Before placing on market
Applies to: Organisations generating synthetic data.
“Organisations are therefore encouraged to document risk assessments, follow the Guide’s practices...”
- #2CriticalImplementation Framework⏰ Before generation
Applies to: Organisations generating synthetic data.
“Data dictionary (Annex B) completed”
- #3CriticalKey Focus Areas⏰ Before generation
Applies to: Organisations generating synthetic data.
“(1) data understanding and minimisation—identify sensitive fields, remove outliers and restrict granularity where feasible;”
- #4CriticalKey Focus Areas⏰ Before generation
Applies to: Organisations generating synthetic data.
“(3) privacy controls—apply pseudonymisation on source data, inject calibrated noise...”
- #5CriticalKey Focus Areas⏰ Before placing on market
Applies to: Organisations generating synthetic data.
“(5) re‑identification testing—use membership, attribute and record linkage tests and threshold‑based risk criteria...”
- #6CriticalMonitoring and Evaluation⏰ Ongoing
Applies to: Organisations using synthetic data in production.
“The Guide requires ongoing monitoring of synthetic data projects: scheduled re‑identification testing, validation of downstream model performance...”
- #7ImportantGovernance and Institutional Framework⏰ Before project commencement
Applies to: Organisations generating synthetic data.
“The PDPC suggests organisations adopt explicit project charters specifying purpose, legal basis, data sources, expected end‑uses and retention timelines.”
- #8ImportantGovernance and Institutional Framework⏰ Null
Applies to: Organisations generating synthetic data.
“It recommends integrating synthetic data governance into existing data protection management systems and PDPA compliance programmes...”
- #9ImportantGovernance and Institutional Framework⏰ Before vendor engagement
Applies to: Organisations engaging external vendors for synthetic data.
“...and using contractual safeguards when engaging external vendors or cloud providers.”
- #10ImportantGovernance and Institutional Framework⏰ Before project commencement
Applies to: Organisations generating synthetic data.
“...inclusion of synthetic data projects in enterprise risk registers and privacy impact assessments.”
- #11ImportantKey Focus Areas⏰ Before placing on market
Applies to: Organisations generating synthetic data.
“(4) utility validation—define and measure downstream performance metrics to ensure synthetic data supports intended tasks...”
- #12ImportantKey Focus Areas⏰ Before placing on market
Applies to: Organisations generating synthetic data.
“(6) contractual and operational controls—access restrictions, purpose limitation clauses, logging...”
- #13ImportantMonitoring and Evaluation⏰ Ongoing
Applies to: Organisations generating synthetic data.
“It recommends maintaining logs of generation runs, model versions and random seeds where feasible...”
- #14RecommendedGovernance and Institutional Framework⏰ Before project commencement
Applies to: Organisations with novel or high-risk synthetic data projects.
“Where appropriate, organisations are encouraged to engage regulators and the IMDA PET Sandbox for early regulatory feedback.”
Related Regulations
Advisory Guidelines on the Use of Personal Data in AI Recommendation and Decision Systems (PDPC)
Singapore93% similar
Advisory Guidelines on Use of Personal Data in Generative AI
Singapore90% similar
Guide on the Use of Generative Artificial Intelligence Tools by Court Users (Registrar's Circular / Singapore Courts)
Singapore88% similar
AI Verify (AI governance testing framework and toolkit)
Singapore88% similar
Advisory Council on the Ethical Use of AI and Data
Singapore88% similar
© Regulations.AI — created on 13-Jun-2026