Know whether the change helped

Benchmark scores answer somebody else's questions. Yours is narrower and harder to settle: is this getting better at the work your organisation actually does, judged the way your best people would judge it?

A yardstick in your own terms

An eval is a written account of what good looks like, precise enough that a machine can apply it a thousand times. If your team argues about whether a generated summary is faithful to the meeting it came from, the eval is where that argument gets settled once and in writing, so the next thousand summaries are judged the same way rather than by whoever happens to read them.

Nobody outside your organisation can write that account, because the standard in it is yours. What can be brought in from outside is the method for getting it out of people's heads and onto the page.

The swap test

Satya Nadella stated the test for ownership as being able to switch out a generalist model without losing the company-veteran expertise built into your learning system. Whether you pass it comes down to where that expertise physically sits.

If your rubrics, your labelled examples and your scoring live in your own repository, changing model is a configuration change and a re-run. If they live in a vendor's dashboard, or were never written down because the model seemed to know already, then changing model means starting again. That is why everything here ends up on your side of the line.

Your experts have already decided. Nobody wrote it down

Andrew Hall spent a term teaching this at Stanford, and had his students each pick a criterion they personally cared about: how sycophantic a model becomes in a political argument, whether it gives sound voting advice, whether it holds cultural nuance across languages. Working rubrics and leaderboards, inside a single three-hour session.

The speed is not the interesting part. Those projects worked, on Hall's own reading, because each student already held the standard they were encoding. A company is the same with one complication: the standard is spread across several people who have never had to reconcile theirs with each other. Writing it down is an interviewing problem before it is a technical one.

What we do, what your engineers keep

Your engineers are better than we are at building judges, wiring harnesses and running them in CI, and they keep all of it. We bring the part with no library behind it: deciding what is worth measuring, writing rubrics that survive real disagreement, and adjudicating when two qualified reviewers score the same output differently and both have good reasons.

Everything produced lands in your repository, readable and editable by your own team and still running after the engagement ends. Nothing is hosted, so there is nothing to migrate off later.

Why a template would not do this

The judgements that decide anything are the ones people disagree about. Resolving those means somebody in the room asking why, then writing the answer down in a form a machine can apply. A template gives you the format, but the asking is the work.

Stop being the answer to every question hello@useful.agency