A second opinion on every reply
Every reply is reviewed before it reaches you: a separate call, its own prompt, no shared context, routed to the strongest judge available. Pass, flag, or fail.
The model that writes a reply is not the one that gets the last word on it. A separate review call reads every reply with its own instructions and none of the conversation's context, and returns one of three verdicts. A failure sends the reply back to be written again. The review is routed to the best judge available at that moment: the main model when it is already warm, a smaller always-on model on a second machine when it is not.