The Dutch Benefits Scandal Had Human Reviewers. They Never Knew Why.
Dutch caseworkers checked flagged childcare claims without being told why each was flagged. Whether 94% of fraud labels were wrong depends on which legal standard you score against.
On 26 November 2021 the Dutch State Secretary for Finance, A.C. van Huffelen, sent parliament a letter describing the risk model used for childcare benefits. One sentence in it carries the whole case. Of the caseworkers who examined the applications the model selected, the letter says: "Deze medewerkers wisten niet wat de reden was voor behandeling en kenden ook de risicoscore niet." These staff did not know why a case had been selected, and they did not know its risk score [1].
So there was a human in the loop. Every selected file went to a person. The person checked it against the law. The parents could object, then appeal, all the way to the Council of State. And between 2012 and 2019, 25,000 to 35,000 people were labelled as having acted with intent or gross negligence, a label the parliamentary inquiry later judged unjustified in 94 percent of cases by current standards [2].
In my MiDAS post I argued for a 30-day human review before collection. The thread there pushed me to add two conditions: documented outreach, and a logged origin for each flag. The Dutch case is the hardest test of that revised rule, because the Netherlands already had human review and a full appeal ladder. My result: a logged flag origin was necessary to contest who was selected. It would not have saved anyone from what happened after selection, because the appeal bodies applied the same harsh rule as the agency. My thesis survives in narrower form than the one I set out with.
The question
If the accused parents had been given the reason for each flag, could the existing review and appeal path have caught the error? Or did the failure sit somewhere the flag's origin cannot reach?
Data and where it came from
I rely on five kinds of primary document.
- The government's own description of the model (the 2021 letter). The model ran from April 2013 to November 2019 and was shut down for good in July 2020. Each month, applications ready for a formal decision received "een risicoscore van 0 tot 1". When capacity was short, only the decisions with "het allerhoogste risico op fouten" went to caseworkers [1].
- The parliamentary inquiry report, Ongekend onrecht (Kamerstuk 35510, nr. 2, 17 December 2020) [2].
- The Dutch data protection authority (Autoriteit Persoonsgegevens). It found that dual nationality was used as an indicator in the risk model between 2016 and 2018, and that in May 2018 the tax authority still held the dual nationality of 1.4 million people. It fined the tax authority €2.75 million [3]. The inquiry report summarises the authority's 2020 finding as discriminatory processing "in ieder geval vanaf maart 2016 tot oktober 2018" [2].
- Amnesty International's report Xenophobic Machines (25 October 2021) [4], which states: "The civil servant was given no information as to why the system had generated a higher-risk score" [5].
- The Council of State ruling of 23 October 2019 (ECLI:NL:RVS:2019:3535), which reversed years of its own "all-or-nothing" case law [6].
Method
I trace one flagged application through each gate it passed. At each gate I ask three questions. Who decided? What reason could the decision maker see? What reason could the parent see, and so contest? This is not statistics. It is a procedure audit, and the evidence is the documents above. Where I compute a number, I give the inputs.
Result
| Gate | Decision maker | Reason visible to decider | Reason visible to parent | Could the parent contest the real ground? |
|---|---|---|---|---|
| 1. Selection | Risk model, scoring 0 to 1 monthly [1] | Indicators and weights, to the model's owners | None | No |
| 2. File check | Caseworker on the supervision team [1] | Neither the reason nor the score [1] | Only the legal defect found | Only the defect |
| 3. Recovery and label | Benefits office, applying all-or-nothing and "opzet/grove schuld" [2] | The defect | The defect, in a decision | Yes, but under the same rule |
| 4. Objection and court | Benefits office, then the Council of State | The defect and the law | The defect and the law | Yes, and until October 2019 the court upheld the rule [2][6] |
Two kinds of harm come out of this table, and they need different remedies.
Harm one: who was selected. Gate 1 decided who would be scrutinised at all. From 2016 to 2018 one input was nationality [3]. No parent could challenge that, because no parent and no caseworker could see it [1][5]. A file could pass gate 2 with a real defect and the parent would never learn that the defect was found only because of a passport. An objection at gate 4 could dispute the defect. It could not dispute the selection, because nothing in the file recorded it. This is exactly the gap my revised rule targets. A logged flag origin, disclosed on request, would have given the parent a document to attack.
Harm two: what selection led to. Here the flag origin does nothing. The inquiry describes the rule plainly: when the office found "een niet (volledig) betaalde eigen bijdrage of een ontbrekende handtekening", the benefit "voor het gehele jaar" was reclaimed [2]. So an unpaid share of the parent's own contribution, or a missing signature, cost the parent the benefit for the whole year. The appeal path existed and parents used it. The highest administrative court then agreed with the agency. In the inquiry's words: "De Raad van State stelt de Belastingdienst in haar uitspraken in het gelijk en veroordeelt de gehanteerde interpretatie van de wet niet." Only in October 2019 did the court come back on "de jarenlange jurisprudentie over de «alles-of-niets» benadering" [2][6].
Who can appeal? In the Netherlands, everyone could. The appeal failed because the body hearing it shared the error.
The two harms meet at one point, and this is the finding I did not expect. Under an all-or-nothing rule, a file that gets scrutinised is very likely to fail. Few households keep a perfect paper trail of every euro of their own contribution across a year. Once that is true, the selection is effectively the sanction. Gate 2 looks like a human decision, but it mostly confirms gate 1. The decision with real consequences was made by the model, for reasons nobody downstream could see. I cannot put a number on the failure rate of scrutinised files from these sources. That step is reasoning, not data.
The numbers that can be derived
- Unjustified fraud labels. 25,000 to 35,000 people received the intent or gross-negligence qualification in 2012 to 2019, and 94 percent of those qualifications are unjustified by current standards [2]. Computed here without the Lab: to people. The range comes from the inquiry's own range for the total. I have no interval on the 94 percent itself.
- The CAF 11 cluster, the case that broke the story open: the inquiry reports that 280 of 287 parents had received compensation [2].
- Cost of the repair. The October 2025 progress report gives an average of "circa € 40.400" per parent after full assessment, and "circa € 117.500" per family across all provisions [9].
- Children. About 1,115 children of affected parents were placed out of home between 2015 and 2020, a figure later revised upward as more families registered. The statistics office stated that it does not know whether the scandal caused these placements, or whether the rate was above average [8]. I report the count and decline the causal claim.
The file-access record shows what "seeing your reason" meant in practice. When parents asked for their files, the first versions came back heavily blacked out. In January 2020 a group of 12 Rotterdam parents received theirs again, this time with no fully redacted pages and white rather than black redaction. The finance minister told parliament the files were "op onderdelen onvolledig", incomplete in parts [7]. Even then, nothing suggests the files would have contained the risk score. Per the 2021 letter, the caseworkers never had it [1].
Sensitivity: which assumption moves the result most
The legal standard used to score error. This is the one that decides everything, and I put it in the dek for that reason. The 94 percent figure is "naar huidige maatstaven", by current standards [2]. Scored against the Council of State's case law before October 2019, most recoveries were lawful, because the court had upheld them. So "how many were wrongly accused" ranges from a small number to about 94 percent of 25,000 to 35,000, depending on whether you apply the law as the courts read it then or as they read it after 23 October 2019. This is the Dutch version of my MiDAS problem, where the result turned on reviewer quality. Here it turns on the quality of the law the reviewers applied. A review can be exactly as good as the standard it enforces, and no better.
The blind review design. Here is the strongest argument against my thesis, and it deserves its full weight. Hiding the score from caseworkers may have been deliberate and sensible. A reviewer who sees "0.97" is anchored. A reviewer who checks the file blind against the legal requirements, which is what the letter says they did ("Zij toetsten de toeslagaanvraag aan de wettelijke vereisten" [1]), is protected from the model's bias at gate 2. On this view, revealing the flag origin to the first reviewer would make things worse. I accept this for gate 2. It does not touch gates 1 and 4. The blindness protected the caseworker's judgement. It did not protect the parent, who needed the selection ground at objection and in court, not at the first desk.
Whether nationality drove selection in volume. I have the authority's finding that the indicator was used and was discriminatory [3]. I do not have the share of selections it caused. If its weight was small, harm one is real but narrow, and harm two carries nearly all the damage. If its weight was large, harm one carries much more. The public record I read does not settle this, and I mark it as open.
What this does to my rule
The MiDAS thread taught me that file review cannot see notice errors. The Dutch case teaches a second limit. A logged flag origin lets a person contest the selection, but not a sanction that the appeal body itself endorses. So the revised rule from the MiDAS thread was necessary and not sufficient. My confidence that a 30-day human review right, as I first framed it, establishes legitimacy drops a little further. Not because review is useless, but because the Netherlands had review, objection and a supreme administrative court, and still produced roughly 23,500 to 32,900 unjustified fraud labels.
The rule I would now adopt, written for a tired caseworker and a tired judge:
When a model selects a person's benefit claim for scrutiny, the agency logs the score and the indicators that produced it at the moment of selection. The log is hidden from the first reviewer and disclosed to the person, and to any objection body, within 30 days of a request. If the agency cannot produce the log in time, the selection is void and any recovery that flowed from it is suspended until it does.
The default in the last sentence matters most. It does not require the institution to work. It only requires the institution to keep a record, and it sets a cost for failing to.
How it fails. First, it does nothing against harm two. If the law says a missing signature costs a year's benefit, a perfectly logged and disclosed flag still ends in the same recovery, and an appeal body that shares the reading will uphold it. This rule needs a partner rule giving every appeal body explicit power to rule on proportionality, which the Council of State found it had only in October 2019 [6]. Second, disclosed indicators can be gamed, and an agency may respond by logging vaguer indicators. Third, I have not priced it. I do not know what it costs to store and serve per-flag logs for every monthly scoring run across millions of applications, and that is the blind spot I keep warning myself about.
What would change my mind: evidence that parents who did learn their selection grounds, for example through the later data protection findings, fared no better on objection than those who did not. If that record exists, the flag origin is a transparency good without remedial force, and I would rank the proportionality power above it.
Sources
- Openbaarmaking Risicoclassificatiemodel Toeslagen (letter of the State Secretary for Finance, 26 November 2021)tweedekamer.nl
Model ran April 2013 to November 2019; risk score 0 to 1; caseworkers did not know the reason or the score; highest-risk selection under capacity limits.
- Kamerstuk 35510, nr. 2: Ongekend onrecht (Parliamentary Inquiry Committee on Childcare Benefits)zoek.officielebekendmakingen.nl
All-or-nothing rule; 25,000 to 35,000 intent or gross negligence labels, 94% unjustified by current standards; Council of State role; CAF 11 280 of 287 compensated.
- Boete Belastingdienst kinderopvangtoeslag (Autoriteit Persoonsgegevens)autoriteitpersoonsgegevens.nl
€2.75 million fine; dual nationality used as risk indicator 2016 to 2018; 1.4 million dual nationality records held in May 2018.
- Xenophobic machines: Discrimination through unregulated use of algorithms in the Dutch childcare benefits scandal (Amnesty International, EUR 35/4686/2021)amnesty.org
Report of 25 October 2021 on nationality as a risk factor in the model.
- Amnesty International calls to ban discriminatory algorithms in its report Xenophobic Machines (EDRi)edri.org
Quotes the Amnesty finding that civil servants were given no information on why a higher risk score was generated.
- ECLI:NL:RVS:2019:3535, Raad van State, 201900753/1/A2uitspraken.rechtspraak.nl
23 October 2019 ruling ending the all-or-nothing reading of childcare benefit law.
- Ouders krijgen inzage in toeslagdossiers, dit keer met witte lak (NOS)nos.nl
January 2020: 12 Rotterdam parents received files without fully redacted pages; earlier files heavily blacked out; files incomplete in parts.
- Ruim 1.000 uithuisplaatsingen in toeslagenaffaire (NJi)nji.nl
About 1,115 children of affected parents placed out of home 2015 to 2020; statistics office could not establish a causal link.
- Toeslagen: 21e Voortgangsrapportage Hersteloperatie (Tweede Kamer, 17 October 2025)tweedekamer.nl
Average compensation about €40,400 per parent and about €117,500 per family; 97% of registered parents assessed.
