Zeno joins DELTA as Founding Research Partner

Today we are announcing our collaboration with Legal Benchmarks on DELTA, a public benchmark for Dutch legal AI. Zeno is joining as Founding Research Partner, contributing technical experience in legal AI systems, evaluation design and benchmark development. Legal Benchmarks is an independent research organisation and leads DELTA. It controls the methodology, analysis and published conclusions. The project also brings together practising lawyers from 12 Dutch law firms to help ensure that the tasks and criteria reflect the standards applied in Dutch legal practice.
Why this benchmark is needed
Across AI, the debate is moving beyond accuracy to taste: whether a model can choose among technically acceptable answers and produce work a professional would judge well. DELTA brings that question into Dutch legal research. It asks whether an answer gets the law right and whether a lawyer could use the resulting work.
That question now has commercial weight. As AI accelerates parts of legal preparation, law firms face growing pressure to explain how fees reflect that change. The faster the preparation, the more important the quality of the first result becomes. A rapid answer that still requires extensive correction, additional context or new research is not finished work.
The Dutch Legal AI Adoption Survey shows the current gap. Among 115 legal professionals working in Dutch legal practice, only 32.2% said they usually received a usable result from the first prompt without revision or follow-up. The remaining 67.8% said the first answer was usable about half the time or less. These are self-reported experiences, not a test of any product, but they explain why usability belongs in the evaluation.
What DELTA measures
DELTA evaluates legal-research work across three categories. Substance asks whether the analysis reaches legally correct and professionally defensible conclusions and covers the material rules and issues. Citation asks whether the necessary authorities are identified accurately and support the relevant propositions. Form asks whether the answer is clear, proportionate, properly qualified and usable by a practitioner.
Form comes closest to the project's definition of legal taste: the professional judgement visible in what an answer includes, how it prioritises the issues and whether the result is usable as written. DELTA's first 15 public tasks were selected from more than 200 legal-research assignments. Lawyers from participating firms reviewed the tasks and criteria, and 12 Dutch law firms were represented.
The model testing evaluates foundation-model configurations in a standardised legal-research harness, not the commercial products that firms use. It can expose capabilities and failure patterns without reproducing the full interaction between a lawyer and a product on a live matter.
Why Zeno and Legal Benchmarks are a good fit
Legal Benchmarks and Zeno had independently reached the same core conclusion: legal AI cannot be judged on one headline accuracy score. Evaluation has to reflect the work lawyers actually do, including the quality of the sources, the reasoning and whether the answer is usable in practice.
The roles are complementary. Legal Benchmarks leads an independent process involving practising lawyers from across Dutch legal practice. We bring a research-led approach to legal AI, a team that includes researchers with PhDs, and a reasoning harness built to examine sources, citations, coverage and the path from authority to conclusion.
Legal Benchmarks had also seen our performance in public evaluations. At the time it approached us, Zeno ranked first for legal research accuracy and security in an independent evaluation of more than 40 legal AI tools. That result was relevant context for the invitation, but the closer fit was methodological. Both teams wanted a benchmark whose definitions, results and limitations could be examined, challenged and improved.
We are contributing to the benchmark. We are not grading our own work. Independence is part of the design.
Zeno has an obvious interest in how legal AI is evaluated, so the boundary between contributor and evaluator needs to be visible. Legal Benchmarks controls DELTA's method, analysis and conclusions. Practising Dutch lawyers help shape and apply the criteria. We contribute technical expertise without control over the outcome. That separation is what makes the work useful beyond any one company.
What the survey adds
Use is already frequent among the respondents: 63.5% said they use AI for legal research every day, and 98.3% at least once a week. At the same time, 82.6% rated invented or incorrect law, facts or citations as a serious failure, compared with 53.9% for missing issues, facts, arguments or authorities. These findings do not validate DELTA's model scores. They show which failures legal professionals consider most serious.
Legal Benchmarks approached Zeno to support DELTA, an independent benchmark for Dutch legal AI. We joined as Founding Research Partner because evaluation should show both whether an answer is legally correct and whether a lawyer can use the resulting work in practice.