Skip to content
Vol. 1 · Weekly editionWeekly · in your inbox
inklaretaal
Vol. 1 · No. 202639

Legal edition

in·klare·taal
Edition 202639 · Monday 21 September 2026 · Clarity since 2026

A court draws a line around the term 'AI system' for the first time, a new benchmark shows the best model scores 15 percent on Dutch legal work, and opposing parties are complaining faster with the help of AI.

What this means for you
Het programma dat de AI laat werken: een model

Court draws a line: the IND's search programme is not an AI system

The Noord-Holland district court explains why CaseMatcher is not AI. In doing so, it sets a first test for the term 'AI system'.

Four Turkish asylum seekers claimed that the IND had used an advanced AI system for their application, without mentioning this in the decision. The tool in question was CaseMatcher, an internal IND programme. According to them, this programme had a decisive influence on the rejection. The Noord-Holland district court disagrees. In the court's view, CaseMatcher is a simple 'rule-based' search tool: it searches earlier cases by search term and ranks them by date or relevance. The programme cannot reason, cannot model, and cannot act on its own. So the IND did not need to mention its use. The ruling is known as ECLI:NL:RBDHA:2026:25209. The distinction the court focuses on is between a fixed rule and a learned pattern. A rule-based tool does exactly what a programmer wrote down: search for this term, sort by this date. The result can be traced and repeated. An AI system under the AI Act works differently. It infers something itself from examples and can generate outcomes that no one typed in beforehand. So the court does not look at the product's name, but at what it can actually do: reason, model, act on its own. If it cannot do that, it falls outside the term, and the reporting and transparency duties do not apply. Be careful here. This is one ruling from one court, in asylum cases, not a higher court. And the reasoning cuts both ways. The argument "it's just a search tool" only holds as long as the tool's actual workings support that. As soon as a supplier builds in a language model that summarises or ranks based on learned patterns, the classification shifts too, even if the name stays the same. In practice this means: record what a system actually does now, and reassess this with every update. A classification from 2026 is not a free pass for the 2027 version.

Source (in Dutch)

Background for subscribers
What this means for you
Four signals that could adjust your approach this quarter

This edition gives you a footing on four fronts. The ruling on CaseMatcher offers a concrete line of reasoning for deciding whether a tool counts as an AI system or not: look at what the thing actually does, not its name. The DELTA benchmark shows that general language models still lag well behind on real legal application, so ask suppliers for hard figures instead of a demo. The piece on AI-drafted complaint letters shows that record-keeping around payment demands and deadlines matters more now, because the threshold for complaining is lower. And the notary firm handling a large share of certificates of inheritance with AI shows that standard work is coming under price pressure, beyond notaries too.

Het programma dat de AI laat werken: een model

New Dutch benchmark: best model scores 15 percent on legal work

This week, Legal Benchmarks published DELTA, a public benchmark for Dutch legal research. It contains 15 practical assignments and 273 criteria validated by lawyers. The best-scoring model, Fable 5.1, reaches a task pass rate of just 15 percent. Of the 23 AI systems tested, only two actually apply the correct legal standard to a concrete case about a legal representative with conflicting interests. Mark Zijlstra (ICTRecht) says this shows how high the bar is once you look not just at citing sources, but at correctly applying the law to the facts. The difference lies in what you measure. Most demos show that a model names a source and gives a smooth answer. DELTA tests the next step: does the model choose the right standard, and does it apply that to these facts? That is the step a lawyer earns their money on, and it is exactly where the models fall short. A 'task pass rate' means the whole assignment must be correct, not just part of it. One wrong standard and the task counts as failed, however polished the text looks. Don't read the 15 percent as a verdict on legal AI products in general. The study only used general foundation models, not tools built specifically for the legal market with their own sources and checks. What it does say: the underlying model does not solve your specialist problem on its own. So a tool is only as good as the sources and checks the supplier builds around it. The test set is fully open, so you can check it yourself with your own cases.

Source (in Dutch)

Background for subscribers

This is a taster.

Subscribers read the whole Legal edition: every story, each with its background in plain language.

Subscribe →

Not ready for that? Read the free letter first →

inklaretaal

AI and tech, made simple.

This is the Legal edition. Subscribers get the background to every story, in plain language.
Subscribe
inklaretaal · Amsterdam · © 2026