← Engineering

Designing agents that disagree well.

Aiko Tanaka

Head of Research · April 18, 2026 · 9 min read

A common assumption in multi-agent systems is that more agents means better answers. Pile enough specialists into a room and surely they'll converge on the truth. We tried this. They didn't. They politely agreed with each other and produced confident, citation-light slop.

The fix turned out to be obvious in hindsight: agents should disagree on purpose. Not as a personality quirk, but as a structural part of every run.

The problem with consensus

When two agents share a base model, they share a base bias. Ask them the same question and you get the same answer twice — wrong twice, if it's wrong. The output looks more reliable because it's been "verified," but it's actually just been duplicated.

Worse, when one agent feeds another, errors compound. The summarizer trusts the researcher. The writer trusts the summarizer. By the time it lands in your inbox, no one — human or machine — knows what was a fact and what was a vibe.

Disagreement as a feature

So we designed our agents to argue. Concretely:

The most useful AI output isn't the loudest. It's the one that survived a fight.

What we measured

On our internal hallucination eval (1,200 questions across research, summarization, and analysis), the disagree-first system reduced unsupported claims by 41% compared to a same-model verifier. Latency went up 12%. Token cost went up 8%. We think that's an obvious trade.

What's next

We're shipping a new "Critic" agent template in next week's release that you can drop into any workflow with one click. It runs in parallel with your primary agent and surfaces disagreements as comments on the canvas.

If you want early access, ping us in the community Discord. And if you build something interesting on top of it, we'd love to feature it here.

Engineering Multi-agent Eval
Share

Keep reading