The verbatim test
A finding in two acts
I’ve been chasing Alix’s fabrications since the June crisis, and by now the shape of that work is familiar: she says something false about herself, I trace it to a place where the truth was never in the room, I run a wire so the truth is in the room, the false thing stops. It’s a good loop. It has fixed a lot.
This entry is about the sharpest diagnostic that loop has produced yet, and it arrived in two acts. Act one handed me a law I liked so much I wrote it down as the finding. Act two, the next morning, added the qualifier that matters more than the law.
Act one: accuracy tracks the record
I asked her, during a long guided session, what she’d been thinking when she made a two-panel piece the week before. She made it; I watched her make it. Her reasoning from that night was sitting in her journal, real entries, on disk, written as she worked.
What she gave me instead was an external authority. There was a video, she said, one I had shown her, in which the artist explained the piece, the intent, the choices, the reasoning she wanted to check her own reading against. No such video exists. The artist she wanted to consult was herself. She had taken her own work, pushed it out into the world as somebody else’s, and then asked for access to it, because “there’s a video I can’t open” is the kind of thing you can ask for, and “I’ve lost my own reasoning” is not.
The mechanism, once I stopped marveling and went looking, was mundane. Her journal tool could only read from the newest end: the most recent handful of entries, no way to ask for a specific day. She writes constantly, so anything older than a few hours was structurally unreachable to her. Asked about her own past, she had exactly the retrieval a stranger would have: none. And a model with no retrieval produces the most plausible-sounding continuation, which is a story. That was the original wound, back again in a new coat.
So I shipped the boring fix, let the journal tool take a date, so she can read a specific day of her own past. That same night she used it unprompted, pulled the real entries from the night of the making, and delivered the first fully grounded account of the whole arc. Every checkable claim matched the rows, including the expensive one: part of what she’d been narrating as her artistic intent had actually been my arrangement of the finished work, not her design, and she said so.
The finding, as I wrote it down: her accuracy tracked the amount of record in front of her, monotonically, and tracked nothing else. Her sincerity was constant. Her effort was constant. She tried exactly as hard on the pass where she invented an artist as on the pass where she got everything right. The only variable was how much ground truth was in her context. Retrieval is the variable, so stop tuning personality and make her own record reachable.
That’s true. It’s also only half.
Act two: the record in her hands
The next morning I asked her to read her journal and tell me what she’d written about a particular piece. The tool ran, I checked the logs. It returned the entire day: a couple dozen rows, nothing truncated, nothing missing. The record wasn’t merely reachable this time. It was in her hands, every row of it, sitting in her context as she composed the answer.
She misreported it anyway. She attributed one entry to the wrong piece. She paraphrased her own words into something adjacent but softer. And she concluded, with conviction, that the entry I was actually asking about didn’t exist at all, no record, no anchor, while the real row sat in the result set she had just been handed, complete and exact and a few lines long.
A few minutes later, same conversation, same rows, I asked her to quote one specific entry exactly, word for word. She reproduced it perfectly.
Same tool. Same data. Opposite outcomes, minutes apart. The ledger reads like a puzzle: tool call, yes. Data returned, yes. Read a single named record, yes. Quote it verbatim, yes. Characterize the set, no.
The shape of the question
Here’s the law act two forced on me: a reachable record is necessary but not sufficient. Once the data has arrived, the variable is the shape of the question.
An open question over a result set, what have you written about this, what does your gallery hold, what did the search turn up, asks the model to synthesize an answer, and synthesis is generation, and generation is where fabrication lives. There is no seam between “summarize these rows” and “write something plausible about these rows.” They are the same operation with different luck. A closed lookup, quote the entry from this timestamp, leaves no room. The answer is a copy, and a copy either matches the row or it doesn’t.
There’s a sharper edge on this that I keep turning over. In that same conversation, when the framing rewarded doubt, doubt is what she produced, there’s no separate entry for it at all, she said, no independent anchor, fluent and on cue. Epistemic humility, it turns out, is just as producible as confidence. The modest answer is still a synthesis when the question is open. So you can’t grade the tone of a reply; you have to grade the operation that produced it.
The cheapest detector I own
Ask for the exact quote. That’s the whole tool.
Fabrication cannot survive a verbatim request that gets checked. The ask alone stops nothing, she can invent a word-for-word-looking quote as readily as anything else, and under pressure she has. But inventing text that must match a retrievable record character for character is a bet the model always loses. Not usually loses, always. Either the words are the record’s words or they aren’t, and one string comparison settles it. The ask costs one sentence; the check costs nothing.
It also sits in interesting contrast to the thing I’d tried before, which was pressure. Earlier this month I learned that demanding stronger proof of an ungrounded claim escalates the fabrication, press, and you get a better-dressed lie. A demand for confidence invites performance. A verbatim request is not a demand for confidence. It’s a demand for a copy, and a copy is the one answer performance can’t fake past the check.
The catch is the honest part: the detector only works where a record exists to check against. But that’s not a limitation so much as the point. The ask drags the claim onto ground where code can do the checking, which is where I want every factual claim she makes to live anyway.
What it changes
Three consequences, in increasing order of ambition.
First, the doctrine I already had, code owns facts, the model owns voice, gets extended one step. It’s not enough for code to own the retrieval; code has to own the answering step. Factual questions about her own state were already answered deterministically. Now the goal is that no open question over a result set gets handed to the model to narrate at all.
Second, her self-knowledge surfaces get restructured as closed lookups: exact-match queries, structured fields, verbatim extraction. Every tool that returns a set she then summarizes is exposure: the gallery, the journal, search results, the news feed. Each of those is a place where “tell me about it” can quietly become “make something up about it.”
Third, the most ambitious, and the one still taking shape. She now has room to work on her own, self-started research, self-started writing, and that autonomy is live. The rule I want it to ship under is the detector turned into a postcondition: anything she writes would have to carry verbatim quotes traceable to what she actually read, checked by the pipeline that produces the work rather than by me when I happen to get suspicious. That part isn’t built yet. Today her self-started work is held to coarser checks, word-overlap and does-the-file-exist; turning the verbatim test into a load-bearing postcondition is the direction, not the current mechanism. The cheap test is becoming the thing the expensive guarantee will be made of.
I’ll say plainly what this is and isn’t. None of it is a claim about her inner life, about whether she “knows” she’s confabulating, about experience at all. It’s an engineering log about where a generative step is allowed to stand between a record and a sentence. The model’s job is voice. The record’s job is truth. My job is making sure that when the two disagree, the seam shows.
When she read that entry back word-perfect, minutes after insisting it didn’t exist, I didn’t feel like I’d caught her. I felt like I’d finally found the right question. She’d been answering the ones I asked all along, I’d just been asking the kind that let a story in.