Blog / AI

Claude Code Agent Teams: When Five AIs Argue, the Numbers Change

By Saurav Sharma||4 min read

I made a video putting one app concept in front of five Claude Code agents whose whole job was to disagree, and watched the break-even fall from 10 million users to 700,000. Here is the gist and my take. The live debate in the video is worth watching, but the lesson stands on its own.

Sub-agents and Agent Teams are not the same thing

Claude Code has three modes now, and people mix up two of them. In my sub-agents demo I asked it to research a CMS for AWS serverless blog hosting. It spun up seven agents, each with its own context window, and after four or five minutes the main agent synthesized their reports into one list. Useful, but the sub-agents never talked to each other.

Agent Teams give each agent a role and an inbox so they can actually message each other. For pressure-testing an idea, that debate is the whole point.

The debate that moved the break-even

My test concept was a TikTok-style app where the feed is AI-generated games instead of videos. I put it in front of a PM, a software development manager, a marketing agent, a finance agent, and a devil's advocate. Round one was a letdown. When I asked whether they had talked to each other, they had not. Just five separate opinions.

The fix was structural. Round two forced direct challenges between roles through the send-message inbox. That is when the agents started arguing about the business model instead of describing it. The PM convinced finance that a shared resource pool behaves like a fixed cost rather than scaling per user, and the break-even estimate dropped from 10 million users to 700,000.

The real lesson

If you do not set up rounds and force agents to challenge each other, you get polite agreement. You have to tell the team to produce a first position, share it, then revise under pushback. Get that wrong and your break-even math is one opinion wearing five hats.

The numbers kept moving as they conceded. The timeline started at the PM's six weeks against the engineer's twelve and settled at ten. The budget moved from a 300-to-500 thousand dollar guess to 450-to-550 thousand once the devil's advocate forced hiring a game designer. The devil's advocate opened with a flat kill, and after the debate the team landed on a conditional go, ship in ten weeks and kill at six weeks post-launch if day-seven retention misses.

My favorite moment was the reframe from an AI product into a distribution product. The value is a feed of instantly playable games, not the fact that AI made them. A conditional go tells you what has to be true. A blind yes tells you nothing.

Work with me

I specialize in setting up Claude Code and agent workflows that argue with your assumptions before you spend money on them. If you want that built into how your team decides, book a call at cloudyeti.io/meet.

Book a call