My multi-agent coding setup, and why GLM-5.2 surprised me
My regular setup:
- Claude Code.
- Opus 5 writes BDD tests (end-to-end, Playwright, CLI, etc.) according to the requirements.
- Sonnet writes the code to fix tests.
- Fable resolves disputes between Sonnet (which sometimes can’t implement a correct tests) and Opus (which sometimes writes impossible tests). It changes architecture in the middle, describes the requirements at the very beginnig and performs smoke testing at the end.
- Codex acts as a second pair of eyes. Each agent is instructed to consult Codex when it gets stuck. Codex crytises solution, finds gaps.
Such setup allows me to implement complex feature that in usual takes weeks or even months within 1-2 days.
But. It token consuming.
After I reached all my limits, I switched to Z.ai, and...
Honestly, GLM-5.2 is fucking good. At one-fifth of the price:
It’s not as good as Fable at orchestrating subagents. It requires explicit instructions and sometimes allows agents to change tests that are already correct.
For most tasks, even creative ones, it’s as good as Opus 5, as I don’t see any difference between them at all.
It writes code better than Sonnet and as well as Opus 5, slower but comparable.
After some improvements to my setup, I stopped seeing any difference compared with the Anthropic stack