The simple trick
Suppose you want a better answer from an AI model. One of the easiest methods is this: ask the model the same question several times, say n times, and have a second model, a judge or reward model, score each answer. Then show the user the answer with the highest score.
This is called best-of-n. It needs no retraining, it's easy to run, and it works surprisingly well. That's why it's widely used, both in real applications and as the baseline other methods are measured against.
The catch: the judge (or reward model) isn't perfect
The reward model is itself a model, and it only approximates what we really want. It has blind spots. It might reward answers that sound confident, or long, or polite, even when they are worse.
With a few answers, this rarely matters. But as you generate more and more answers, the top-scoring one becomes more and more likely to be the one that found a blind spot: it looks great to the judge without actually being great. Past a certain point, asking for more answers makes the final result worse. It's Goodhart's law in action: once a measure becomes a target, it stops being a good measure.
A softer rule
Our paper studies a smoother version called soft best-of-n. Instead of always taking the top-scored answer, you pick one at random, giving higher-scored answers a higher chance of being picked. A single parameter controls how strongly the judge’s scores influence your choice:
- Complete reliance: always choose the highest-scoring answer. This is ordinary best-of-n.
- No reliance: ignore the scores and choose uniformly at random. On average, this gives the same quality as asking the model once.
- Partial reliance: favour higher-scoring answers while giving others a chance. The judge guides your choice without dictating it.
What we showed
Methods like this are often used without a clear picture of when they help or hurt. We worked out mathematical guarantees for both the ordinary and the soft version:
- How much the choice changes the model's behaviour. Picking from n answers pushes the model away from its usual output, but only slowly: the effect grows roughly with the logarithm of n. The less strongly you favour high scores, the smaller this shift.
- How far the result is from the best possible answer. This gap depends on two things: how wrong the judge is, and how often the model produces a truly great answer in the first place. A weak model or a weak judge both cost you.
- When soft beats ordinary. If the judge is accurate, always taking the top score is the right choice. But if the judge has blind spots. The right balance between the scores and allowing randomness can help the soft version perform better.
Does it hold up in practice?
We tested this with a small language model answering prompts designed to provoke harmful replies, and measured how harmless the chosen answers were. We ran it twice: once with a strong judge and once with a weak one.
With the strong judge, taking the top score kept getting better as we generated more answers, as expected. With the weak judge, ordinary best-of-n started to get worse as n grew: the judge was being gamed. The soft version held up noticeably better, just as the theory predicted.
The takeaway
If you fully trust your judge, pick the top answer. In practice you rarely should trust it fully, and then it pays to soften the choice. It's a one-line change.
The same lesson comes up in my current work on agentic benchmarks. Whenever an AI system is pushed hard against an automated checker, whether that's a reward model, a verifier or a test suite, it will eventually learn to please the checker rather than do the task. Knowing how far you can trust the checker matters as much as the checker itself.