A second opinion
A model asked to check its own work will usually tell you it did fine. It has every reason to: the reply is already in its context, it has just spent a thousand tokens committing to that answer, and “yes, that was right” is the cheapest continuation available.
So the check isn’t self-assessment. Every reply goes out to a separate review call with its own instructions, none of the conversation, temperature zero, and exactly one question to answer. It comes back pass, flag, or fail. A fail doesn’t get softened or annotated: it goes back to be written again, and the person waiting never sees the first attempt.
What the diagram doesn’t show is the argument I keep having with myself about which model should do the reviewing. For a long time it was always a smaller, always-on model on a second machine, genuinely independent, and cheap enough to run on every single turn. The problem is that a small model judging a much larger one is not obviously the right way round. So now the review is routed to the strongest judge available at that moment: the main model when it is already warm, the smaller one when it isn’t.
That trades some independence for a lot of capability, and I’m not certain it’s the right trade. What makes it survivable is that the reviewer never inherits the conversation. It gets the reply and the question, and nothing about how the reply came to seem like a good idea.
Read the full piece: Character as CI.