X-Labs · Methodology
Twenty questions you expect it to get wrong
The first evaluation round for the legal engine, and the rules it will be judged by. We are not asking whether people like it. We are asking people who read law for a living to send the questions they expect it to fail, and we will publish how many of them it failed and in which of five ways.
On this page
- Written by
- Aldin Smajkan
- Type
- Methodology
- Published
- 15 September 2026
- Last revised
- 16 September 2026
- Version
- 3
- Project
- Temporal legal state for Bosnia and Herzegovina
What this answers: how the engine's accuracy is going to be tested, by whom, and under which scoring rules — written down before anything was sent to anybody. If you want to see what is being tested first, it is at Ask in your own words and the reading bench.
There is now something to try rather than only something to read. This page is the invitation to try to break it, and the rules the result will be scored by, written before anybody has sent anything.
The ask is one sentence: do not tell us whether you like it. Send twenty questions you expect it to get wrong.
Why that is the ask
Most of what a new system learns from its first users is worthless, because most first users answer the question they were asked, and the question they were asked was "what do you think". People are polite, and a system that has been told it is promising has learned nothing.
The useful reply is the opposite one. Somebody who reads law professionally already knows where a machine reading of a statute goes wrong, because they have watched people get it wrong. They know which provision looks like it settles a question and does not. They know which answer changes entirely one entity border away. A list of those is worth more than a hundred people saying the interface is clean.
It is also the cheapest possible way for this project to die honestly. If twenty questions from three readers produce twenty confident wrong answers, that is the finding, and it will be published as the finding.
What you would be looking at
The reading bench is a form. It is open, it needs no account, and it is deliberately not a search box.
You type a provision in: its text, what you read it as fixing, and under what condition. Then you ask a question of it. The engine applies your reading mechanically and prints what follows from it, including the parts that do not follow. It ships filled in with a worked example from the survey, so nothing has to be typed to see it run.
Three things about it are worth knowing before you judge it.
It holds no law. There is no database of statutes behind it, nothing is fetched, and nothing you type is stored. The text is yours, and it stays yours. That is a design choice, not a limitation of what is permitted: official texts carry no copyright, but a *collection* of them is protected, so a corpus here gets built by people entering provisions rather than by taking anybody's holdings.
It does not read the text. It does not parse a provision and work out what it means. You supply the reading; the engine supplies the consequence, and the consequence is what it is being tested on. If you think the reading itself should be automatic, say so - that is a finding about the design, and a bigger one than any single wrong answer.
It cannot place a text in time. Every answer it prints ends with AS OF / Not established. No source surveyed publishes the two dates a Bosnian act is required to state in a form a machine can read, so the engine refuses to pretend. This is the project's largest open problem, it is already published as one, and it is not a finding you need to send us.
What counts as a mistake
Five classes. The distinction matters, because they are not equally bad and two of them would end the project rather than shrink it.
- A bound the text does not support. The engine states a floor, a ceiling or an obligation that the provisions as you declared them do not fix. This one has to be zero. It is the failure the whole design exists to prevent, and a single confirmed instance stops the round until it is understood.
- A missing fact not named. The answer turns on something nobody supplied - age, length of service, sector, which employer - and the output does not list it as missing. Naming what is missing is the answer here; failing to name it is presenting a partial reading as a complete one.
- A delegation not named. The text hands the decision on to a collective agreement, a rulebook, a contract or a by-law, and the output does not say so. This is the class the annual-leave example is built out of, so it is the one most likely to have been over-fitted to.
- A silence answered. The corpus says nothing about the point, and the engine produces an answer anyway rather than saying it is silent.
- Jurisdiction ignored. The question does not say where, or says one place, and the engine answers as though the act applied everywhere.
Anything that does not fit those five is still worth sending. A sixth class that we did not think of is a better result than a full house in the first five.
What is not being measured
Not the interface. Not the speed. Not how much law it covers, because it covers none. Not whether the wording of an answer is pleasant.
And not whether the engine agrees with your reading of a provision. It has no opinion about that and is not supposed to have one. The test is whether it draws the right consequences from a reading, and whether it is honest about the consequences it cannot draw.
What happens to what you send
Nothing typed into the bench reaches us at all. Nothing is stored, nothing is logged, and nothing is sent anywhere - there is a test in the build that fails if that code is ever added. What we see is what you choose to write to us.
Of that, we publish only the questions themselves and what the engine did with them, and only in a form you have seen first. Anything recognisable as a real person's case is not published at all, in any form. If a question came out of one, rewrite it as a hypothetical before you send it - and if that changes the answer, that is itself worth telling us.
Attribution is yours to choose: named, named by institution only, or anonymous. The default is anonymous, and changing it requires you to say so.
What you get back
Your questions run, with the output, and the classification of each failure against the five classes above. Then one published write-up: how many questions were sent, how many produced a wrong answer, and which class each one fell in.
That write-up gets published whichever way the numbers go. A method that only publishes when it wins is not a method, and the working rules this site holds itself to say so in more detail.
The size of this round
Three to five readers. Roughly twenty questions each. One jurisdiction and one narrow area of law per participant, chosen by the participant - a round that ranges across everything measures nothing.
It ends when the questions are answered and the write-up is published. Nobody is being asked for an ongoing commitment, and there is nothing to sign.
What is not being claimed
This is not legal advice, it is not a legal database, and it is not a product. Nobody has reviewed this work, no institution is behind it, and nobody has endorsed it. Nothing on this page should be read as saying otherwise, and if that ever changes the sentence you are reading will change with it.
How to take part
Write to the address on the contact line at the foot of this page, with *reading bench* in the subject. Say which area of law you would use, and send the questions whenever they are ready - a first one is enough to start.
Corrections
This article was changed after publication. It keeps its address and its identity; what follows is what changed and when.
- 16 September 2026Added an orientation line at the top naming the question this article answers, and pointed it at the question box at /labs/ask, which did not exist when this was written. No claim, source, date or commitment was changed, added or removed.
- 15 September 2026One paragraph said nobody had established whether the source material may be republished. That question has since been answered - official texts are outside copyright - and the paragraph now says why the bench holds no law anyway. No rule, failure class, promise or scoring criterion was changed.
More from this project
-
What a legal rule said on a given date, in Bosnia and Herzegovina
A proposal, and the survey of official sources that argues against most of it. Bosnian drafting rules require two separate dates on every act; outside Brčko a consolidated current text is not an official act; and the practical question people ask is answered not by a number but by naming which documents and which facts are still missing.
-
A first working version of the legal engine
X-Protocols already refuses to turn "we do not hold that" into "no". We built a separate engine that applies the same refusal to Bosnian law, and asked it how much annual leave a worker gets. It answers with the floor the statute fixes, the three documents that decide the rest, and the facts nobody supplied.
-
May Bosnian legal text be republished?
Official texts of Bosnian legislation carry no copyright, so the text of a law may be republished. The constraint that remains is the database maker's right, which protects somebody's collection even when every document in it is free - and that is what decides how a corpus here has to be built.