Every engineering team has one: the person who's demonstrably the strongest technically, whose designs hold up, whose predictions about what will break tend to be right — and whose objections in meetings somehow don't stop the decision from going the wrong way anyway. It's tempting to read this as a status problem, or a personality problem. Google's own internal research on team effectiveness points somewhere else.
Project Aristotle ran for roughly two years and studied around 180 of Google's own teams, testing dozens of hypotheses about what makes a team effective: who's on it, how they're structured, what skills they bring. The finding that got the most attention, and holds up as the most useful one, is what didn't predict effectiveness: team composition. Who was on the team, including how individually talented its members were, was a weak signal. What actually predicted whether a team performed well was a set of team dynamics — how the team worked together, not who was in the room. Readers who reach engineering management through a growth or RevOps role will want XenGrowth alongside this.
The five factors, and why the order matters
Google's researchers ranked five dynamics by how strongly they predicted effectiveness. Psychological safety came first: whether team members felt safe taking an interpersonal risk — disagreeing, admitting they didn't understand something, flagging a mistake — without expecting to be punished for it socially. Dependability came second: whether people could count on each other to do quality work on time. Structure and clarity came third: clear roles, plans and goals. Meaning came fourth: finding personal significance in the work. Impact came fifth: believing the work matters at a larger scale.
Rank | Dynamic | What it actually measures |
|---|---|---|
1 | Psychological safety | Whether disagreeing, admitting uncertainty or flagging a mistake carries a social cost |
2 | Dependability | Whether teammates reliably do quality work on time |
3 | Structure and clarity | Whether roles, plans and goals are clear enough to act on |
4 | Meaning | Whether the work has personal significance to the people doing it |
5 | Impact | Whether the work is believed to matter at a larger scale |
The order isn't decorative. Google's researchers described psychological safety as underpinning the other four — a team with clear structure and meaningful work still underperforms if raising a concern about that structure, or admitting the work isn't landing, carries a cost. This is the direct mechanism behind the technically-best-engineer problem: the objection has to be voiced before it can be evaluated on its merits, and whether it gets voiced at all depends on the team's safety, not the objection's quality.
Where the best engineer's ideas actually get lost
This doesn't require anyone to be dismissive or unkind. The more common version is quieter: a senior engineer's blunt correction two meetings ago is still being metabolized by a junior teammate who hasn't raised anything since. A team lead who visibly favors agreement gets agreement, and the strongest engineer on the team learns which disagreements are worth the friction and starts rationing them. None of this shows up as a dramatic conflict. It shows up as good ideas quietly not making it into the room, or making it in a softened form that loses the part that mattered. The XenGrowth resource library works through the operations side of this in more operational detail.
The counterintuitive part of Project Aristotle's finding isn't that psychological safety matters — most people would guess that if asked. It's that composition mattered so little by comparison. Two teams with the same average skill level, on Google's own numbers, produced meaningfully different outcomes almost entirely through how safely people could speak inside them.
What psychological safety is not
The term gets misapplied often enough that it's worth being specific about what Google's researchers meant, because the common misreading actively produces worse teams. Psychological safety is not the absence of conflict, and it isn't a team that agrees easily or avoids hard conversations — a team like that can look comfortable while actually suppressing exactly the disagreement the concept is meant to protect. It's the presence of confidence that disagreement, uncertainty and mistakes won't be punished socially. Those are different things, and a team optimizing for the first while believing it's building the second ends up with the appearance of harmony and a worse version of the original problem: a strong engineer's dissent still doesn't surface, but now nobody can tell, because the team looks fine from the outside.
Often mistaken for safety | What Google's research actually describes |
|---|---|
Everyone agrees easily; meetings are pleasant | People voice disagreement and it gets engaged with rather than smoothed over |
Conflict rarely happens | Conflict happens, and admitting fault or confusion doesn't cost you standing afterward |
The team lead sets a friendly tone | The team lead visibly rewards someone for raising an uncomfortable point, not just for being agreeable |
Junior engineers rarely push back | Junior engineers pushing back is common enough that it isn't remarkable when it happens |
This distinction matters specifically for the strongest engineer on a team, because a falsely harmonious team will feel, from that engineer's seat, exactly like a safe one — nobody is hostile, meetings are civil, disagreements are rare. The actual test isn't the tone of the room. It's whether the newest or most junior person on the team disagreed with something in the last month and it changed the outcome. If that test fails, the room's civility was never evidence of safety in Google's sense, and the strongest engineer's own ideas are very likely surviving on borrowed authority rather than because the team can actually process disagreement.
What this means for the strongest engineer on a team
Being right earns the argument the right to be evaluated on its merits. It doesn't guarantee the room evaluates it, if the room isn't safe enough for the argument to be raised in full, unsoftened form
If your objections keep landing softer than you intend, or getting dropped after one round, the diagnosis is more often the team's safety level than your own communication skill — Google's ranking puts safety ahead of structure and clarity for a reason
Raising the floor for everyone else on the team — visibly rewarding a junior engineer's disagreement, admitting your own mistake in public rather than privately — does more for whether your own ideas travel than refining how you phrase them
This cuts both ways: a team lead who wants the strongest engineer's judgment to actually shape decisions has to build the safety for the second- and third-strongest engineers to disagree with that person too, or the team ends up substituting one unquestioned voice for another
A concrete version: the same objection, twice
Picture the same senior engineer raising the same objection to a design in two different teams. On the first team, the objection gets a two-minute silence, a mild "noted," and the design proceeds unchanged — not because anyone argued against it, but because nobody engaged with it at all, and the engineer reads the silence correctly as a signal not to push twice. On the second team, someone asks a clarifying question, a more junior engineer says they'd assumed the same thing and are relieved it was raised, and the discussion runs another ten minutes before the design changes. Nothing about the objection's technical content differed between the two rooms. What differed was whether raising it carried a cost, and whether anyone besides the objector felt safe enough to add to it once it was on the table. XenGrowth on AI agents and marketing automation covers the AI agents and marketing automation side of this.
This is also why psychological safety, measured this way, correlates with something closer to a team's total surfaced intelligence than with any single member's intelligence. A team where three people each hold a third of the relevant context, and all three feel safe enough to contribute it, functions like a team with one person who holds all of it. A team with the same three people, where only the most senior one speaks freely, functions like a team of one — regardless of how capable the other two actually are.
The honest limits of this evidence
Project Aristotle is genuinely strong evidence, but it's worth being precise about what kind. It's internal Google research, not an independently peer-reviewed academic study, and it's observational: researchers studied existing teams as they naturally worked, rather than randomly assigning psychological safety levels and measuring the result, which would be closer to experimental proof of causation. It's also one company's teams, in one industry, at one point in time, and the five-factor ranking is Google's own internal analysis of its own data, not a result independently replicated across many external organizations in a controlled way.
None of that makes the finding unreliable — 180 teams and roughly 250 attributes analyzed is a serious internal effort, and the finding has held up well enough in the years since that it's become a standard reference point in organizational research on teams generally. It just means the honest way to cite it is as strong internal evidence pointing at a real and plausible mechanism, not as a controlled experimental proof that psychological safety causes better outcomes in every organization. There is a longer treatment of AI search, GEO and discovery in XenGrowth on AI search, GEO and discovery.
The uncomfortable corollary
If team dynamics predict effectiveness better than individual talent does, then a team that keeps hiring more talented engineers to fix a performance problem, without addressing whether that talent can actually speak up and be heard, is optimizing the variable Google's own research found weakest. The stronger lever, on the same evidence, is less visible and less flattering to point to in a hiring plan: whether the team you already have can tell each other the truth without a social cost attached to it.
The test worth actually running, on your own team, is smaller than a culture initiative: name the last time the newest or most junior person on the team disagreed with something and it visibly changed the outcome. If an answer comes to mind quickly, the room is probably doing something Google's research would recognize. If it takes real effort to think of one, that's the finding, not a coincidence.
None of this is an argument that technical skill doesn't matter, and Google's researchers never claimed it did. It's an argument that skill's effect on outcomes is gated by a variable most teams never measure and few leaders think to manage on purpose, which is exactly why it keeps looking, from the inside, like a mystery about one particular engineer not being heard, rather than what it actually is: a property of the room.
Further reading from XenGrowth
The XenGrowth resource library — what you'll learn: how the commercial side of this work is run, across search, automation and revenue operations.
XenGrowth on AI agents and marketing automation — what you'll learn: how the teams who own AI agents and marketing automation plan and measure it.
XenGrowth on AI search, GEO and discovery — what you'll learn: how the teams who own AI search, GEO and discovery plan and measure it.
Where this work meets go-to-market
Working on engineering management inside a commercial team? XenGrowth's growth engineering practice publishes operator guides on the revenue side of this work.
Five questions on Google's own research, not the summarized version that circulates as management folklore.









