Singapore - Synthetic Data Generation Guide

Proposed Guide on Synthetic Data Generation (PDPC)

Singapore

RAI-SG-NA-PGSDGXX-2024
Proposed(Officially filed for action)
GuidelineData Protection and PrivacyRisk ManagementGovernance and Oversight
Export PDF

The Personal Data Protection Commission (PDPC) of Singapore published the Proposed Guide on Synthetic Data Generation on 15 July 2024 to explain the uses, risks and good practices for generating structured synthetic data as a Privacy Enhancing Technology (PET). Jointly developed with A*STAR and supported by IMDA, the Guide provides a five‑step approach, case archetypes, and an annexed handbook of technical and governance controls to mitigate re‑identification risks.

Overview

The PDPC Proposed Guide on Synthetic Data Generation (15 July 2024) sets out good practices for producing structured synthetic data as a Privacy Enhancing Technology (PET). The Guide was jointly developed with the Agency for Science, Technology and Research (A*STAR) and is intended to be available as a resource within the IMDA PET Sandbox. It frames synthetic data as artificial data generated by purpose‑built mathematical or ML models that can mimic statistical properties of source datasets while warning that synthetic output may still pose re‑identification risks. The Guide focuses on fully synthetic structured data, provides three common use case archetypes (AI training, data sharing/analysis, software testing), and offers a five‑step process and annexed handbook to guide organisations in technical, governance and contractual controls to mitigate residual risks. It emphasises a risk‑based approach, recognising trade‑offs between utility and privacy, and positions the document as a living guide to be updated as the technology and evaluation methods evolve.

Definitions

The Guide defines key terms used throughout: "synthetic data" as artificially generated data created by mathematical or ML models trained on source data; "fully synthetic" (no real records preserved) and "partially synthetic" (some original values retained) distinctions; "re‑identification risk" meaning the probability an individual in source data can be matched to synthetic records; and "Privacy Enhancing Technology (PET)" covering data obfuscation, encrypted computation and federated analytics. The Guide focuses on structured tabular data and clarifies that while high‑quality synthetic data can preserve statistical properties and utility, it is not automatically non‑personal data and still requires privacy risk assessment.

Governance and Institutional Framework

The Guide outlines governance expectations for organisations that generate synthetic data, including role‑based accountability (CIO/CTO/CDO/Data Protection Officer), documented policies, and cross‑functional oversight. It recommends integrating synthetic data governance into existing data protection management systems and PDPA compliance programmes, and using contractual safeguards when engaging external vendors or cloud providers. The PDPC suggests organisations adopt explicit project charters specifying purpose, legal basis, data sources, expected end‑uses and retention timelines. Where appropriate, organisations are encouraged to engage regulators and the IMDA PET Sandbox (IMDA PET Sandbox) for early regulatory feedback. The Guide also recommends periodic board reporting for higher‑risk projects and inclusion of synthetic data projects in enterprise risk registers and privacy impact assessments.

Key Focus Areas

The Guide’s technical and process focal points include: (1) data understanding and minimisation—identify sensitive fields, remove outliers and restrict granularity where feasible; (2) synthetic generation choices—select generation methods (statistical, model‑based, GANs, probabilistic models) appropriate to data types and risk tolerance; (3) privacy controls—apply pseudonymisation on source data, inject calibrated noise, and consider differential privacy for metrics or model training where provable bounds are needed; (4) utility validation—define and measure downstream performance metrics to ensure synthetic data supports intended tasks without overfitting to real individuals; (5) re‑identification testing—use membership, attribute and record linkage tests and threshold‑based risk criteria described in Annex D and Annex E; and (6) contractual and operational controls—access restrictions, purpose limitation clauses, logging, and incident response integration with existing PDPA obligations. The Guide emphasises iterative trade‑off analysis between utility and disclosure risk and recommends conservative deployments in high‑sensitivity domains (e.g., healthcare, finance) unless robust mitigation is in place.

Implementation Framework

The Guide offers a practical five‑step implementation flow: Step 1—Know your data: catalogue fields, provenance and sensitivity; Step 2—Prepare data: pseudonymise/anonymise and remove outliers, apply minimisation; Step 3—Generate synthetic data: choose model architecture and training protocols, incorporate noise or constraints as needed; Step 4—Assess re‑identification risks: apply quantitative and qualitative tests from Annexes D/E and compare to predefined risk thresholds; Step 5—Manage residual risk: apply contractual, organisational and technical safeguards, maintain documentation and monitoring. Each step includes recommended artifacts (data dictionaries, model cards, risk assessment reports) and acceptance criteria for release. The Guide further provides a sample data dictionary format (Annex B) and examples of generation methods (Annex C) to aid implementation teams.

Monitoring and Evaluation

The Guide requires ongoing monitoring of synthetic data projects: scheduled re‑identification testing, validation of downstream model performance, incident detection and reporting, and periodic review of assumptions (e.g., changes in linkable external datasets that might raise re‑identification risk). It recommends maintaining logs of generation runs, model versions and random seeds where feasible for reproducibility and audit. The Guide encourages organisations using synthetic data in production to include synthetic data metrics in their privacy KPI dashboards, and to re‑evaluate risk when dataset composition, use cases or external threat models change. Where possible, results from the IMDA PET Sandbox pilots and public case studies should inform refinements.

Penalties, Liability, and Appeals

Although the Guide itself is non‑binding, organisations are reminded that synthetic data generation activities remain subject to the PDPA and other sectoral rules. PDPC enforcement powers under the PDPA include directions and financial penalties; since 1 October 2022 the PDPA regime allows penalties of up to 10% of an organisation's annual turnover in Singapore (for qualifying thresholds) or SGD 1,000,000 (whichever is higher) depending on circumstances. Organisations are therefore encouraged to document risk assessments, follow the Guide’s practices, and consult PDPC where novel or high‑risk designs are proposed. The Guide also notes that civil liability or sectoral regulatory sanctions may arise from misuse of data or harms traceable to synthetic outputs; it recommends clear allocation of contractual liability and dispute/appeal mechanisms in supplier contracts.

Relationship to Other Instruments

The Guide complements and builds on existing PDPC materials such as advisory guidelines on AI use and prior PDPC technology guides (e.g., anonymisation guidance). It is designed to be used alongside IMDA’s PET Sandbox resources and sectoral rules (e.g., health, finance) administered by agencies like the Ministry of Health and the Monetary Authority of Singapore where more stringent controls may apply. The Guide references OECD/PDP research and international best practices on PETs and situates synthetic data guidance within the PDPA accountability framework, recommending alignment of synthetic data project governance with organisational PDPA accountability obligations.

International Alignment

The Guide acknowledges analogous instruments internationally — for example, PET guidance and synthetic data research emerging from OECD, EU data protection authorities, and standards bodies — and seeks alignment on terminology and evaluation approaches. PDPC suggests interoperability with international standards and encourages use of internationally recognised privacy techniques (e.g., differential privacy) and measurable re‑identification tests to facilitate cross‑border collaborations while respecting local legal constraints. Through IMDA’s sandbox and international engagement, Singapore aims to harmonise practices to support data sharing, research collaboration and trade while maintaining strong privacy safeguards.

Implementation Timeline

MilestoneDate / Window
Guide published (proposed)15 July 2024
Inclusion in IMDA PET Sandbox resourcesQ3–Q4 2024 (ongoing)
PDPC updates / living document reviewsPeriodic review — next review indicated 2024–2025 (no fixed date)

Compliance Checklist

ActionAcceptance Criteria
Purpose & scope documentedProject charter with stated lawful basis and use cases
Data mapping and sensitivity assessmentData dictionary (Annex B) completed
Pre‑generation mitigationOutliers removed / pseudonymisation applied where appropriate
Generation method selectedModel card and rationale recorded
Re‑identification testing completedRisk metrics within agreed thresholds
Residual risk managementContractual, access and monitoring controls implemented

Sources and References

SourceType
PDPC, Proposed Guide on Synthetic Data Generation (PDF)Primary Source
PDPC, Proposed Guide on Synthetic Data Generation (web summary)Primary Source
IMDA, Privacy Enhancing Technology SandboxesPrimary Source
Plain English

The Singapore Personal Data Protection Commission (PDPC) has issued a Proposed Guide on Synthetic Data Generation, offering best practices for organisations creating artificial data to ensure privacy while maintaining utility. This guide applies to any organisation in Singapore that generates structured synthetic data, whether for AI training, data sharing, analysis, or software testing. While the guide itself is not legally binding, it clarifies how existing data protection laws, specifically the Personal Data Protection Act (PDPA), apply to these activities.

Organisations must adopt a risk-based approach, following a five-step process that includes: - Thoroughly understanding and minimising sensitive data before generation. - Choosing appropriate generation methods and applying privacy controls like pseudonymisation or noise injection. - Rigorously testing for re-identification risks to ensure synthetic data doesn't inadvertently expose real individuals. - Implementing strong governance, including clear accountability, documented policies, and contractual safeguards with vendors. - Continuously monitoring synthetic data projects for ongoing risks and performance.

The guide was published on July 15, 2024, and is currently a proposed document, with ongoing integration into the IMDA Privacy Enhancing Technology Sandbox resources expected in late 2024. It will be periodically reviewed and updated.

Although the guide is advisory, failing to manage privacy risks in synthetic data generation can lead to significant penalties under the PDPA. These can include financial penalties of up to 10% of an organisation's annual turnover in Singapore or S$1,000,000, whichever is higher, in addition to potential civil liability. A crucial takeaway is that synthetic data is *not* automatically considered non-personal data. Organisations must still conduct thorough privacy risk assessments, as even artificially generated data can carry re-identification risks if not properly managed.

Plain-English rewrite by Regulations.ai — not legal advice. Verify against the official text.

What you must do — compliance checklist

0 / 14 marked complete

Plain-English obligations under Singapore - Synthetic Data Generation Guide. Not legal advice — verify against the official text before relying on it.

  1. #1CriticalPenalties, Liability, and AppealsBefore placing on market

    Applies to: Organisations generating synthetic data.

    Organisations are therefore encouraged to document risk assessments, follow the Guide’s practices...
  2. #2CriticalImplementation FrameworkBefore generation

    Applies to: Organisations generating synthetic data.

    Data dictionary (Annex B) completed
  3. #3CriticalKey Focus AreasBefore generation

    Applies to: Organisations generating synthetic data.

    (1) data understanding and minimisation—identify sensitive fields, remove outliers and restrict granularity where feasible;
  4. #4CriticalKey Focus AreasBefore generation

    Applies to: Organisations generating synthetic data.

    (3) privacy controls—apply pseudonymisation on source data, inject calibrated noise...
  5. #5CriticalKey Focus AreasBefore placing on market

    Applies to: Organisations generating synthetic data.

    (5) re‑identification testing—use membership, attribute and record linkage tests and threshold‑based risk criteria...
  6. #6CriticalMonitoring and EvaluationOngoing

    Applies to: Organisations using synthetic data in production.

    The Guide requires ongoing monitoring of synthetic data projects: scheduled re‑identification testing, validation of downstream model performance...
  7. #7ImportantGovernance and Institutional FrameworkBefore project commencement

    Applies to: Organisations generating synthetic data.

    The PDPC suggests organisations adopt explicit project charters specifying purpose, legal basis, data sources, expected end‑uses and retention timelines.
  8. #8ImportantGovernance and Institutional FrameworkNull

    Applies to: Organisations generating synthetic data.

    It recommends integrating synthetic data governance into existing data protection management systems and PDPA compliance programmes...
  9. #9ImportantGovernance and Institutional FrameworkBefore vendor engagement

    Applies to: Organisations engaging external vendors for synthetic data.

    ...and using contractual safeguards when engaging external vendors or cloud providers.
  10. #10ImportantGovernance and Institutional FrameworkBefore project commencement

    Applies to: Organisations generating synthetic data.

    ...inclusion of synthetic data projects in enterprise risk registers and privacy impact assessments.
  11. #11ImportantKey Focus AreasBefore placing on market

    Applies to: Organisations generating synthetic data.

    (4) utility validation—define and measure downstream performance metrics to ensure synthetic data supports intended tasks...
  12. #12ImportantKey Focus AreasBefore placing on market

    Applies to: Organisations generating synthetic data.

    (6) contractual and operational controls—access restrictions, purpose limitation clauses, logging...
  13. #13ImportantMonitoring and EvaluationOngoing

    Applies to: Organisations generating synthetic data.

    It recommends maintaining logs of generation runs, model versions and random seeds where feasible...
  14. #14RecommendedGovernance and Institutional FrameworkBefore project commencement

    Applies to: Organisations with novel or high-risk synthetic data projects.

    Where appropriate, organisations are encouraged to engage regulators and the IMDA PET Sandbox for early regulatory feedback.

© Regulations.AI — created on 13-Jun-2026