It gets better by finding out what actually happened — and it refuses to let that be gamed
Most systems that claim to learn get more confident the more they are used.
Corobate does the opposite first: when it starts measuring a source, that source's rating
usually falls, because a handful of results is weak evidence and the system will not
pretend otherwise. It rises slowly after that, and it falls quickly when something turns out
to be wrong.
Every number on this page came from running the actual software. Nothing here is an
illustration. You can reproduce all of it with one command, given at the bottom.
The loop in one paragraph
A decision is made and a record is sealed. Later, somebody independent reports what really
happened. That report changes the rating of the source whose evidence was used. The next
decision that touches that source sees the new rating. Round and round — with three
deliberate brakes, described below, that stop the loop from flattering itself.
Step 1 · Being measured makes a source look worse before it looks better
What we did
We took a record from an accredited laboratory and had independent auditors
confirm it, one at a time. Then one auditor found it wrong.
The rating is on a 0–10,000 scale, where 10,000 would be certainty. Watch what happens at
the fifth confirmation.
Why it drops when the evidence gets better
For the first four confirmations there are too few results to say anything, so the system
ignores them entirely and rates the record on the laboratory's accreditation alone. At the
fifth, it has enough to start counting — and a pessimistic reading of five results is worse
than the benefit of the doubt an accreditation buys. The rating falls from 9,500 to 5,655,
then climbs back as the record of success lengthens. A system that only ever went up when
you fed it good news would be a system telling you what you want to hear.
And one bad result costs more than three good ones bought
Six confirmations took the rating from 5,655 to 6,456. A single auditor finding the record
wrong took it back to 5,291 — below where it stood after four. That asymmetry is deliberate.
Being caught out once is more informative than being confirmed once, and the arithmetic is
built that way rather than tuned that way.
Step 2 · A good record cannot buy a promotion
The same experiment, one thing changed
We ran the identical seven confirmations against a supplier's own document
instead of a laboratory's.
This is the property that stops the loop being farmed
If a good track record could promote a supplier's own paperwork into independently verified
evidence, then anyone who wanted a high rating would simply arrange to be confirmed a lot.
A supplier's word is capped at a supplier's word — permanently, whatever its history. The
track record can only ever pull a rating down from what the source's standing already
allows. It can never push one up past it.
Step 3 · Nobody grades their own homework
Three reports about the same record
We had the supplier who filed the record say it turned out right; then a
laboratory that had backed it up say the same; then an unconnected auditor.
Written down but not counted — and the difference matters
Refusing to count a report is not the same as refusing to record it. All three
are on the permanent record, so anyone reviewing later can see that the supplier said its own
document was fine and that this carried no weight. Deleting it would hide the attempt.
Counting it would let anyone build their own reputation out of their own opinion.
Step 4 · A person who accepts risk gets a scorecard — and it changes nothing about their authority
Five decisions by the same director
Some decisions cannot wait for perfect evidence. A named person can accept a
known gap, in writing, if the possible loss is theirs to bear and is bounded. The system then
tracks how often the risks they accepted actually went wrong.
Read that last column carefully — it is the honest one
After four resolved decisions with one bad outcome, the worst-case reading of this person's
failure rate is —. Not 25%. The system reports the upper bound,
because four decisions genuinely do not tell you much, and a number that looked precise would
be the more dangerous output. It takes a great many resolved decisions before that figure
means anything, and the system says so on every record until it does.
The scorecard cannot widen anybody's authority
A director's limit was — throughout, and the software
contains an explicit check that this limit is identical before and after the scorecard is read.
If a future change ever tried to let a good run buy a bigger limit, that check stops the
program rather than quietly permitting it. Being right is not the same as being entitled.
A system where good luck buys permission is a system that eventually permits the wrong thing
to someone who has been lucky.
Step 5 · A decision can become evidence — and it carries how far it is from anything anyone saw
A sealed decision can be fed back in as evidence for a later one. When that happens the
system records how many steps it stands from a first-hand observation. A conclusion drawn from
a conclusion drawn from a conclusion is still usable — but nobody reading it can mistake it for
something somebody witnessed.
What the whole thing refuses to do
Three brakes, all measured above
A source's standing sets a ceiling that its own track record can never lift. Nobody's report
about their own work counts toward their own rating. And a person's success rate never widens
what they are allowed to approve. Each of those makes the system less flattering to
everyone in it, including its owner. That is the point: an improvement loop worth trusting is
one that can tell you something you would rather not hear.
What it does not claim
Corobate never reads the content of a document to judge whether it is plausible. It knows who
said something, how old it is, whether anyone independent has backed it up, and what happened
afterwards. It does not know whether a claim is true, and it never says it does.