Head-to-Head: Why a New Model Has to Beat the Old One, Not Just Look Good
In my home training rig, a new model doesn't get promoted for clearing some fixed score. It has to beat the current champion, on the same test set, the same day, under the same judge.
For a long time the promotion gate on my training rig was a single number. Finish a fine-tune run, push it through the eval suite, check whether the score cleared a bar I'd picked months back, ship it if yes. Felt like a process. Really it was a number I'd stopped questioning.
The trouble is that bar doesn't hold still. Every so often I'd catch the judge rewarding something dumb: padded, hedge-y answers scoring high on a vague "thoroughness" criterion. I'd go fix the rubric or tweak the judge prompt, reasonable on its own. But the moment I touched the judge, the old bar stopped meaning what it used to mean. A score calibrated against last month's judge tells you nothing about a model graded by this month's judge. I was clearing a bar I'd quietly moved myself, and congratulating each new adapter for it.
Same test set, same day, same judge
So the gate is a straight head-to-head now. A new adapter doesn't get promoted for clearing a threshold. It gets promoted for beating whatever model is currently live, on the identical frozen test set, on the same day, graded by the same judge configuration. If I upgrade the judge itself, I rerun the reigning champion through the new judge in the same sitting, so both models are scored on the same yardstick at the same moment rather than one of them getting graded by a version of the judge that no longer exists. Ties don't count as wins, either. My judge isn't precise enough for me to trust a coin flip, so the challenger has to win outright or it stays on the bench.
What I like about running it this way is that it stops asking the judge to be correct in any absolute sense, only consistent for the length of one match. Whatever quirk the judge happens to reward this week, both models are answering the same questions in the same room on the same afternoon. The gap between their scores is the only number I actually trust, because everything that could make the measurement drift is held fixed for both of them at once.
The harder part is sticking to it when the result is inconvenient. I've had a run finish, checked the standalone score, and felt pretty good about it, every tracked category up a bit from the last checkpoint. Then the head-to-head came back and the new model lost to the champion, on the same conditions, same day. That stings more than it should, given the run chewed through hours on a machine sitting a few feet from my desk. Every instinct says the standalone number is the real story and the match result is noise.
Usually it isn't noise. Digging through the losing transcripts a couple of times, I found the new model had picked up a real regression on one case type the champion still handled cleanly, a weak spot the aggregate score buried under gains everywhere else. The number on the spreadsheet went up. The only question that actually matters, whether this model beats the one currently serving requests, went down. Shipping on the isolated score alone would have quietly made the fleet worse while the report said the opposite.
So the rule now is dull on purpose: nothing goes live without winning its match, no exceptions for a run that ate a whole weekend of compute. Some weeks that means a model I was excited about just doesn't ship, and I put it back in the queue mildly annoyed at myself. It's more overhead than I'd like for a one-person setup. But a fixed number was never really measuring anything for long, it just felt like it was, and I'd rather run a slower gate than a broken one.