AI needs a deterministic definition of done
Claude: "Task done. I checked it. Every DoD criterion is met, and codex verified it too."
Me: "Did you actually open the page? I don't see it there. Check again."
Claude: "Damn, you're right."
And this is Fable 5. The check runs as a separate codex agent with its own context. And you still have to verify by hand.
Another 5x in productivity from AI? Forget it, at least until we find a way to verify done deterministically. I'm saying this with my usual sarcasm, and I mean it fairly literally at the same time. ๐
What we're missing is a Playwright script, or some technology we don't have yet, that could say with 100% certainty whether a task is done, with an answer that doesn't depend on the model, the context window, or how the stars happened to align that day.
Until that exists, every result an agent produces still gets read by a person, and that review is where the promised speedup quietly goes. Nobody counts those hours when the productivity numbers get published. ๐ค
This is where the word deterministic starts doing real work. The condition has to be checked by something deterministic, not by another model โ because a model checking a model is still two opinions, and the story above is what that looks like. A deterministic definition of done is a genuinely hard thing, because we still don't know how to describe it in a way that holds up on a real project. From what I know by now, we've spent a lot of years not getting there, and it keeps lying somewhere over there, roughly. ๐ ๏ธ