The Pizza Party is Over
What the Two-Pizza Rule, Brooks’s Law, and a 2,234-Person Field Experiment Say About the Post-Agent Org Chart
In the early days of Amazon, Jeff Bezos set the rule: No team should be bigger than two pizzas can feed, usually six to eight people. Why? Smaller teams “minimize lines of communication and decrease overhead of bureaucracy and decision-making” (AWS Executive Insights).
The Two Pizza Team rule was based on actual research, not folklore. Amazon cites the Ringelmann Effect, individual productivity dropping as a group grows, and a Hackman and Vidmar study finding individual satisfaction drops the same way. It also cites Gallup’s State of the American Workplace data. Organizations under ten employees scored 42 percent engagement or higher, while larger organizations averaged less than 30 percent.
Fifty years earlier, Fred Brooks proposed the formula for the underlying mechanism in his seminal work, The Mythical Man-Month. He wrote that communication channels between team members grow as n(n-1)/2. A nine-person team already carries 36 channels, and fifty developers carry 1,225 (Brooks’s Law explained). Adding people to a late project can make it later, because the new coordination cost outweighs the new hands.
Shifting from functional silos to Two Pizza Teams solved the problem of managing the coordination cost between people who each hold a narrow slice of the problem (the exact cost Brooks’ formula measures). Two Pizza Teams became the reference model, the structure organizations reached for to cut that cost.
Still, many organizations never made the leap to Two Pizza Teams from their traditional functional silos because the transformation required was one of culture and not simple tool or process adoption. Organizations reach for matrix structures and shared-service queues as a structural fix for what’s actually a cultural problem. On paper, a matrix org could look like a group of Two Pizza Teams, but when the communication efficiency wasn’t achieved (because a matrix actually only further complicated the lines of communication, rather than simplifying them), organizations were all too ready to declare failure and go back to the way things were.
Bolting AI Onto the Legacy Org Chart
Fast forward to 2026, companies are rushing toward AI adoption with hopes of driving efficiency and savings.
Enterprises are bolting AI onto their old org charts, as if strapping a rocket pack to the back of each employee would instantaneously make them faster. They never asked “Does this structure still make sense?”
If every specialist gets a coding assistant or an AI tool tuned to their role, the org chart stays exactly the way it is. AI in their minds is a tooling change, nothing more.
The organizations that never made the leap to Two Pizza Teams are the worst off of all. They never fixed the coordination cost in the first place, and bolting an AI tool onto one specialist at a time does nothing to fix it now.
Several governance pieces I’ve read treat the whole problem as sanctioning tools, tiering risk by how sensitive the work is, and cataloging what gets built. All of that is useful hygiene. The guidance assumes every specialist keeps their existing lane and treats AI as one more tool for that lane.
Framing AI this way undercuts the fundamental impact of the technology, and ultimately attempts to solve adoption problems at the wrong layer.
The Agent Absorbs the Coordination Cost
Handoffs, meetings, ticket queues, and status updates exist because no one person can hold everything in their head … except, maybe, an agent with sufficient context.
This context management is what makes agents so useful: they subsume communication overhead. An agent can translate between specialists, draft the boilerplate, and run the non-functional work each specialist used to do by hand. They can also fill in for skill gaps, enabling the ai-pilled “solopreneur.”
Once that overhead drops out, a leader is free to size a team purely for expertise.
That mechanism now has empirical support, not just theory. Harang Ju and Sinan Aral ran a randomized field experiment with 2,234 participants split between human-only and human-AI teams, producing 11,024 ads for a real client and testing them live on X to roughly five million impressions (Ju and Aral, arXiv, 2025/2026). Human-AI teams produced 50 percent more output per worker and higher text quality. Task-oriented messaging rose 25 percent, and interpersonal messaging fell 18 percent, consistent with an agent absorbing the coordination that used to run through people. The same study found delegation also reduces output diversity, what the authors call diversity collapse (more on that later).
Borrowed Terms, Applied Structurally
The idea of pairing two specialists from different fields is not new. Bain’s chairman Orit Gadiesh coined the term expert-generalist for someone who masters multiple disciplines well enough to connect patterns a narrow specialist would miss (Forbes, 2015). Andrew Sobel credits a related term, deep generalist, to leadership scholar Warren Bennis: someone who combines actual depth in one field with genuine range across others (Sobel). While neither piece is particularly data-heavy, both highlight how the pairing model works: Put a person deep in the business next to a person deep in engineering, and the pair behaves like a single expert-generalist, without either person actually having to become one.
Some of my best career successes came through pairing with a business-vertical SME and working together to create software-based solutions aimed at disruption: e-commerce, mortgage banking, video games, even AI infrastructure.
When we started Genvid’s Studio-Grade AI Production app, I structured the team this way. At first, Stephan Bugaj was our subject-matter expert, and I was the deep technologist. We had Opus 4.5 at our disposal, and we converged on solutions fast, with high-bandwidth communication between the two of us and Claude filling in wherever either of us fell short. As we acquired customers, we took on a forward-deployed-engineer (FDE) approach to bake new ideas rapidly.
A pair of expert-generalists moves faster than the six-to-eight person team because there’s almost nothing to coordinate. The business expert can prototype something that works in hours. The engineering expert brings the non-functional rigor and operational discipline so the solution runs cleanly at production scale.
The bottleneck that this exposed surprised me: it moved from dev velocity (classic) to creative input. When ideas are implemented almost at the speed they are conceived, you can only move forward when there are fresh ideas.
Keeping Score
If an organization measures the pair by how many tools got sanctioned or how many builds got cataloged, the same vanity metrics that plague every other AI transformation effort just show up again. Token spend is no different (Token Spend Is the New Lines of Code). An organization has to measure the pair by outcome instead. Did the customer’s problem actually get solved? What did it cost per outcome?
This problem is one of business clarity and outcomes, and if neither exists no amount of AI can help. In the new model, pairs run out of ideas to build long before they run out of coordination capacity. Once execution moves at close to the speed of conception, idea supply becomes the constraint, not headcount.
Lest we think that we can outsource idea generation to AI, Ju and Aral’s caveat rears its ugly head: Delegation improved text quality but also produced diversity collapse. Output became more homogeneous as people handed more work to the AI. If idea supply is the resource the pairing model runs on now, handing the ideation itself to agents translates to a hundred pairs across a company building the same handful of ideas instead of a hundred different ones. Keeping that variety alive once agents generate ideas, not just execute them, is the design problem the next wave of these teams actually has to solve.
