← All posts

When the judge can be fooled: a softer way to pick the best answer

A simple summary of our ICLR 2026 paper on “best-of-n”, one of the simplest ways to get better answers out of an AI model.

The simple trick

Suppose you want a better answer from an AI model. One of the easiest methods is this: ask the model the same question several times, say n times, and have a second model, a judge or reward model, score each answer. Then show the user the answer with the highest score.

This is called best-of-n. It needs no retraining, it's easy to run, and it works surprisingly well. That's why it's widely used, both in real applications and as the baseline other methods are measured against.

The catch: the judge (or reward model) isn't perfect

The reward model is itself a model, and it only approximates what we really want. It has blind spots. It might reward answers that sound confident, or long, or polite, even when they are worse.

With a few answers, this rarely matters. But as you generate more and more answers, the top-scoring one becomes more and more likely to be the one that found a blind spot: it looks great to the judge without actually being great. Past a certain point, asking for more answers makes the final result worse. It's Goodhart's law in action: once a measure becomes a target, it stops being a good measure.

A softer rule

Our paper studies a smoother version called soft best-of-n. Instead of always taking the top-scored answer, you pick one at random, giving higher-scored answers a higher chance of being picked. A single parameter controls how strongly the judge’s scores influence your choice:

Best-of-n versus soft best-of-n on five example answers Five answers, A to E, with judge scores 0.62, 0.91, 0.55, 0.88 and 0.70. Answer B has the top score but games the judge; answer D is truly the best. Best-of-n picks B every time. Soft best-of-n picks B about 47% of the time, D about 37%, and the others occasionally. Judge's score Best-of-n Soft best-of-n Chance of being picked: 0.620.910.550.880.70 ABCDE games the judgetruly best 0%100%0%0%0% 5%47%3%37%9%
A made-up example with five answers. Best-of-n always picks B, the answer that fooled the judge. The soft version still favours high scores, but it gives the truly best answer, D, a real chance.

What we showed

Methods like this are often used without a clear picture of when they help or hurt. We worked out mathematical guarantees for both the ordinary and the soft version:

  1. How much the choice changes the model's behaviour. Picking from n answers pushes the model away from its usual output, but only slowly: the effect grows roughly with the logarithm of n. The less strongly you favour high scores, the smaller this shift.
  2. How far the result is from the best possible answer. This gap depends on two things: how wrong the judge is, and how often the model produces a truly great answer in the first place. A weak model or a weak judge both cost you.
  3. When soft beats ordinary. If the judge is accurate, always taking the top score is the right choice. But if the judge has blind spots. The right balance between the scores and allowing randomness can help the soft version perform better.

Does it hold up in practice?

We tested this with a small language model answering prompts designed to provoke harmful replies, and measured how harmless the chosen answers were. We ran it twice: once with a strong judge and once with a weak one.

With the strong judge, taking the top score kept getting better as we generated more answers, as expected. With the weak judge, ordinary best-of-n started to get worse as n grew: the judge was being gamed. The soft version held up noticeably better, just as the theory predicted.

The takeaway

If you fully trust your judge, pick the top answer. In practice you rarely should trust it fully, and then it pays to soften the choice. It's a one-line change.

The same lesson comes up in my current work on agentic benchmarks. Whenever an AI system is pushed hard against an automated checker, whether that's a reward model, a verifier or a test suite, it will eventually learn to please the checker rather than do the task. Knowing how far you can trust the checker matters as much as the checker itself.

← All posts