Synthetic Data Terminology in Official Statistics

We define and review application scenarios in official statistics for synthetic data generation and related methodology.

Synthetic data finds its foundation in statistical disclosure control. It was formalized by Rubin in 1993 as a technique to enable safer distribution of microdata to the public. Today, synthetic data has evolved into an interdisciplinary field spanning statistics, computer science, and artificial intelligence. It is no longer used solely for disclosure control but also for a wide range of applications, including data sharing, machine learning development, testing algorithms, and addressing data scarcity.

At the same time, ongoing research continues to address serious challenges related to data utility, disclosure risk, and evaluation methods. In this paper, we aim to create clarity in the terminology surrounding synthetic data, specifically for application in official statistics.

We first give a brief history of synthetic data methodology to place the definition of synthetic data in a historical context.

Subsequently, we describe the bulk of the terminology related to synthetic data, propose definitions and scopes for each term, and reflect on the different methods from the view of different application scenarios in official statistics.

Rosanne J. Turner, Reinoud Stoel and Peter-Paul de Wolf (2026). Synthetic Data Terminology in Official Statistics. Discussion paper, Statistics Netherlands, The Hague/Heerlen.