Estonia’s national transport model, LAAM, is being developed to answer a practical question: how will people actually move around the country? Which roads will get congested in ten years? Where is a bus route missing? What happens to traffic if a new road or railway is built? It’s the tool that will let the people planning Estonia’s roads, public transport, and infrastructure investment base decisions on how people really live and travel, instead of guesswork.
The open-source, data-driven model combines conventional transport modeling with an activity-based approach that puts individuals at the center. Rather than treating each trip in isolation, it accounts for people’s daily activities, mobility needs, and the way their travel decisions are shaped by factors such as where they live and who they share a household with. This makes it possible to analyze both the current transport system and future scenarios in much greater detail.
LAAM is a joint effort, led by the University of Tartu’s Mobility Lab together with Positium, Statistics Estonia, Goudappel and STACC. This article is about STACC’s part of it: the synthetic population that the rest of the model runs on.
The problem with modeling a real country
A model like this can’t reason about “Estonia” as an abstraction. It has to reason about actual people: a retired couple in Pärnu, a student commuting into Tartu, a family in Maardu with two children and one car. Multiply that by roughly 1.3 million residents and you hit a wall. Real, address-level records for every person in Estonia obviously can’t be handed to a transport model, or to anyone outside a locked government register system. But if you invent people at random instead, the model just produces confident nonsense.
The way out is a synthetic population: an artificial population of about 1.3 million invented people, grouped into roughly 569,000 households, that collectively match the real Estonia (the same age mix, the same family shapes, the same number of children, the same car ownership, the same commuting patterns) down to the neighborhood level, without any single synthetic person corresponding to a real one.
Problems that sound simple and aren’t
Building the synthetic population meant solving a long list of things that look easy until you try them.
Putting households in the right place
Where does each household actually live? Not “somewhere in this town” but a specific building, because a model that places a family half a kilometer from their real home gets the traffic on that street wrong.
Building households that make sense
Realistic families come in more shapes than a spreadsheet suggests: couples, single parents, people living alone, children, and a single parent sharing a home with a grown child, all assembled so the households actually make sense (no ten-year-olds living by themselves, no household with three heads).
Assigning cars without inventing or losing them
How many cars does a household have? A car can be owned by one person but used by the whole family, or owned by a company and driven by someone who lives at a private address. That matters because one household member’s access to a car can affect the travel options available to other members of the household. Getting this right, without either double-counting cars or losing them, took real work.
Putting jobs where people actually work
Workplaces caused a different problem. Left unchecked, “other” jobs (everything from construction to public administration) pile up wherever there happens to be a large office building, instead of reflecting how jobs are actually spread out.
Handling households that don’t fit the usual pattern
And what about the people who don’t live in a typical household at all: children in institutional care, very large group homes? These edge cases needed their own handling rather than being quietly dropped.
Privacy by design, not by promise
Nothing about a real individual ever leaves the secure government register environment where this is built. Every number produced there passes through a strict rule: no statistic may describe a group smaller than three people. If fewer than three people anywhere share a given combination of traits, that count gets randomized just enough to blur it, so no group is ever small enough to point back to one household or one person.
From privacy-safe numbers to a full population
What comes out of that secure environment is a set of privacy-safe summaries: rounded counts and patterns, nothing individually identifiable. The last step, called disaggregation, rebuilds a full, detailed synthetic population from those summaries outside the secure environment: roughly 1.3 million synthetic people and 569,000 synthetic households that, added together, reproduce the same real statistics, without ever containing a real record.

Why it matters
The result behaves like Estonia’s population without being a copy of it, close enough that a transport model can be trusted to say something true about roads, buses and railways that don’t exist yet.
The wider LAAM project is still under development. If you’d like to follow its progress, you can read the latest news or subscribe to the project newsletter here.



