SIESTA Concepts: Synthetic Data

Some of the most valuable data for research is also the most sensitive. Health records, census microdata and detailed socio-demographic files could power countless studies yet they describe real people, and for good reason they sit behind strict access controls. Researchers who need them can usually reach them only through scientific-use files or controlled access in secure centres if at all.

This is the tension EOSC-SIESTA sets out to resolve. EOSC, the European Open Science Cloud, is the European Union’s initiative for an open, trusted environment where researchers can find, share and reuse data and tools under the FAIR principles. But sensitive data doesn’t fit neatly into an open cloud. That is the gap EOSC-SIESTA fills, and one of its key concepts is synthetic data.

What is synthetic data?

Synthetic data is artificial data. It doesn’t describe real people; it imitates a real dataset. Instead of releasing genuine information about specific individuals, you study in depth how that information behaves (how age relates to gender, household type, municipality or employment status) and then use that knowledge to generate a brand-new dataset, built from scratch, made up of people who never existed. Fictional individuals who, taken together, resemble the real ones statistically: the same age structure, the same mix of household types, the same patterns of migration or income.

More precisely, synthetic data is microdata produced by a statistical model that has learned the patterns of a real, confidential dataset. The model does not copy records, it reproduces them. The result simulates what the real data would look like without being an observation of anyone in particular.

One nuance matters for understanding the concept properly: synthetic data comes in degrees.

  • Fully synthetic data: no single value comes from a real person; the entire dataset is model-generated. This protects privacy the most. It also has a legal consequence: data that can no longer identify anyone, directly or indirectly, may fall outside the definition of “personal data” altogether.
  • Partially synthetic data: only the variables most likely to reveal an identity are replaced by simulated values (exact age, place of birth, nationality, fine geography), while the rest of the record is kept as it is. This usually still counts as personal data and has to be treated as such.
  • Hybrid data: a middle path that mixes real and simulated values within each record according to the risk of each variable. It’s often the most practical option for demographic files, which combine low-risk classification variables with high-risk ones like precise location or rare events.

Good synthetic data has to be two things at once, and they pull in opposite directions. It must be useful, so a researcher gets almost the same result as with the real data. And it must be safe, so it leaves no trail back to any specific person. The whole craft is finding that balance, measuring it, and documenting it openly.

Synthetic data is not only a laboratory idea, official statistical offices are already using it. In the United States, the Census Bureau has published synthetic data for years: one of its tools shows where people work across the country without ever exposing a single worker. Statistical offices in the United Kingdom, Germany, Austria, Denmark and the Netherlands have all built or tested synthetic versions of sensitive datasets, from labour surveys to farm censuses.

This matters because population data is the hard case. Sources like the census and the population register hold rich detail (age, gender, nationality, place of birth, town of residence) and it’s precisely their breakdown into small geographic areas that makes them risky to share. Synthetic data offers a modern alternative to older privacy techniques such as grouping figures together, capping extreme values or hiding small cells.

A legal framework pulls in both directions at once: protect the data, yet open it up. At the European level, Regulation (EC) 223/2009 makes statistical confidentiality a founding principle, while Regulation (EU) 2019/1700 pushes statistical offices to find ways of sharing demographic data without breaching it. The European Statistics Code of Practice sets these two demands: confidentiality and accessibility side by side, and synthetic data is increasingly seen as a way to satisfy both. Crucially, the GDPR shapes the outcome: fully synthetic data that identify no one may fall outside the definition of “personal data” (Article 4.1, Recital 26), whereas partially synthetic data usually remain personal data, a distinction decided case by case. 

EOSC-SIESTA puts the idea to work, across its epidemiology and demography use cases. The goal is not to upload real data and make it fully available, but to use it as the basis for realistic simulations: a proof of concept for a different way of working.

Starting from information that originates in sources like the census, the project builds a synthetic population, an artificial population with the same characteristics as the real one: household composition, age, mobility, socioeconomic traits. That synthetic population then feeds into models. The aim is a simulator useful both for research and for the public management of health crises, paired with an interface that lets users with no technical background configure the models they need for a particular disease and read off the results.