Every technical leader knows the feeling: sharp decisions in the morning, mushy ones by 4pm. The explanation everyone reaches for is decision fatigue — willpower as a tank that empties with each choice, running dry by the end of a long day. It's a great story. The evidence behind it has taken a beating that the story's popularity doesn't reflect.
This is worth sorting out properly rather than waving away, because the two candidate explanations point at completely different fixes. If willpower genuinely depletes with each decision, the fix is rationing decisions — batching them, delegating them, deferring the unimportant ones. If the real driver is something else entirely, rationing decisions might miss the actual lever completely while leaving the real cause untouched. I write about decision fatigue from the build side; XenGrowth, who work on the commercial side of this covers what it takes to run it.
The study everyone cites
Nearly every popular treatment of decision fatigue starts in the same place, so it's worth starting there too and following the evidence more carefully than the popular version usually does.
Danziger, Levav and Avnaim-Pesso's 2011 analysis of roughly 1,100 Israeli parole rulings found something striking: the probability of a favorable ruling started each session around 65%, fell steadily to near zero by the session's end, and jumped straight back to about 65% right after a food break. It reads like a perfect natural experiment in mental depletion, and it's been cited that way constantly since — including, often, as the anchor evidence for decision fatigue in general.
The challenge the popular version leaves out
That's the version everyone stops at. There's a second half of the story that the popular retellings almost never mention, and it changes the conclusion substantially.
A later simulation study by Glöckner took the same data and asked a harder question: could this pattern show up without any depletion at all, purely from how the day's cases were scheduled? Two facts turned out to matter. Favorable rulings simply take longer to process than denials, so more of them get processed early in a session before time runs short. And unrepresented prisoners — who are granted parole less often regardless of when they're heard — tend to be scheduled later in the day. Simulate a court session with just those two structural facts, no fatigue mechanism required, and you get a decline that looks a lot like the original finding.
Interpretation | What it requires | Status |
|---|---|---|
Judges deplete a willpower reserve across a session | A mental resource that drains with each decision and refills with food | The underlying theory failed a large, preregistered replication attempt |
The pattern reflects case-scheduling structure, not depletion | Favorable cases take longer + unrepresented prisoners scheduled later | Reproduces a similar-looking decline in simulation without any fatigue mechanism |
Neither of these fully settles the question of what was actually happening in that courtroom. What the challenge does establish is that the parole study is nowhere near the clean, mechanism-proven demonstration of decision fatigue it's usually presented as. On the operations side of this specifically, The XenGrowth resource library is worth reading.
The bigger problem: the theory underneath it didn't replicate
Decision fatigue is usually presented as a specific case of a broader theory called ego depletion — the idea that self-control and willpower draw from one shared, limited resource, so any act of self-control leaves less available for the next one. In 2016, Hagger and colleagues ran a preregistered replication of the core ego-depletion effect across 23 independent labs, with 2,141 participants and a standardized protocol agreed on in advance.
The replication found no significant ego-depletion effect. Not a smaller effect. Not a mixed result. Across 23 labs, using the field's own agreed-upon method, the core mechanism decision fatigue depends on didn't show up.
This is a genuinely awkward finding for a term as widely used as decision fatigue. It doesn't prove judgment never degrades over a day. It does mean the specific, popular mechanism — willpower as a depletable tank — is standing on much shakier ground than the term's ubiquity would suggest.
Why a large, high-profile failure like this matters beyond one study
It's worth pausing on why the 2016 replication failure matters more than a single disappointing result usually would. Ego depletion wasn't a fringe idea — by the mid-2010s it had generated hundreds of published studies, spawned popular books, and become a stock explanation in business writing for everything from why judges rule harshly to why dieters cave in the evening. A 23-lab preregistered replication, agreed on in advance by researchers on both sides of the debate, using a standardized protocol, is about as strong a test as social psychology has the infrastructure to run. Finding nothing under those conditions doesn't just weaken one study — it weakens the entire category of claims that lean on the theory, decision fatigue included, unless those claims can point to independent support that doesn't route through ego depletion.
None of this means every popular claim built on ego depletion is automatically false. It means the burden of proof sits with the claim, not with the skeptic, which is the opposite of how 'decision fatigue' usually gets deployed in practice — as an assumed background fact that other arguments get built on top of without anyone checking the foundation. For the AI agents and marketing automation angle, see XenGrowth on AI agents and marketing automation.
What's actually well-supported instead
Set the shaky mechanism aside and there's still a real phenomenon to explain, and the evidence for a specific alternative is stronger than the evidence for the one everyone defaults to.
The afternoon slump a technical leader feels is very likely real. The better-supported explanation just isn't a draining willpower tank — it's accumulated cognitive load from a day built almost entirely out of context switches. Meyer, Barton, Murphy, Zimmermann and Fritz's instrumented study found collaborative activity — the meetings, escalations, one-off questions and interruptions that fill a leader's day disproportionately — at 24.4% of the measured workday, more than coding itself, with switches happening every 0.3 to 2.0 minutes outside planned meetings.
Layer in Mark, Gudith and Klocke's finding that people compensate for interruption and pressure by working faster and harder, at a real cost in measured stress and effort that showed up within 20 minutes. A leader's day is disproportionately built from exactly the kind of activity that research shows carries a real, accumulating cost — not from a fixed number of 'decisions' draining a shared tank.
Explanation | What it predicts | Evidentiary status |
|---|---|---|
Willpower depletes with each decision made | Judgment quality should track raw decision count, regardless of what filled the time between decisions | Core mechanism failed a 23-lab, 2,141-participant preregistered replication |
Accumulated cognitive load from context-switching | Judgment quality should track how fragmented the day was, regardless of raw decision count | Consistent with instrumented data on collaborative load (24.4%) and switch frequency (0.3-2.0 min) |
The two explanations make genuinely different predictions, which is what makes this worth getting right rather than treating as a semantic quibble. If willpower depletion is the mechanism, a leader who makes ten decisions in a quiet, uninterrupted morning should be just as depleted as one who makes ten decisions across a morning of constant escalations and pings. If accumulated context-switching is the mechanism, those two mornings should look nothing alike by early afternoon, even though the decision count is identical. Anecdotally, most technical leaders report the second pattern, not the first — a calm morning with several real decisions feels nothing like a chaotic one with the same number, which is itself a small piece of evidence favoring the fragmentation account over the depletion account. XenGrowth on AI search, GEO and discovery approaches this from the AI search, GEO and discovery side.
What actually follows from getting the mechanism right
Stop treating 'decision fatigue' as settled science when you plan around it. The specific mechanism it's usually attributed to has a serious replication problem, which matters for how confidently you should generalize it
Address the actual candidate cause instead: reduce the raw number of context switches a leader absorbs in a day, not the raw count of decisions, since it's the switching — not decision-making per se — that has the stronger evidentiary backing
Group similar decisions together rather than scattering them across the day. Nothing in the failed-replication evidence argues against batching reducing cognitive load; it argues against the specific 'depleting resource' story for why batching might help
Treat a food break's apparent effect in the original study with appropriate skepticism as evidence for anything beyond 'a break helps' — which barely needs a depletion theory to explain
If you're building a process around 'protecting decision quality,' design it around measurable fragmentation (meeting density, interruption frequency) rather than an unmeasurable internal resource nobody has been able to reliably detect in a lab
A word on the food-break detail specifically
The original parole study's most quoted image — approval rates snapping back up right after judges ate — deserves its own honest look, because it's usually presented as the clinching detail. A break restoring performance is consistent with willpower depletion, but it's also consistent with almost any other fatigue-adjacent mechanism: simple physical tiredness, low blood sugar affecting mood and patience, or even just a change in the physical environment breaking up monotony. 'A break helped' is compatible with a very large number of explanations, most of which don't require inventing a depletable psychological resource to account for it. Treating the food-break detail as decisive evidence for ego depletion specifically, rather than for 'breaks are generally good,' is exactly the kind of overreach this post is arguing against.
This matters for technical leadership directly. If a leader notices their judgment improves after lunch, the accurate conclusion isn't 'my willpower reserve refilled.' It's one of several plausible, boring explanations — physical fatigue easing, blood sugar stabilizing, or simply having stepped away from the specific fragmented stretch of meetings that preceded it — none of which require believing in a resource that a large registered replication failed to find.
The honest version of this story is less tidy than 'willpower runs out.' It's also more useful: if the real driver is accumulated context-switching rather than a fixed decision budget, the fix isn't rationing how many choices you make. It's rationing how many times your attention gets yanked somewhere new before you're asked to make the next one. That's a scheduling problem, not a metaphysical one, and scheduling problems are the kind you can actually solve.
Further reading from XenGrowth
The XenGrowth resource library — what you'll learn: how the commercial side of this work is run, across search, automation and revenue operations.
XenGrowth on AI agents and marketing automation — what you'll learn: how the teams who own AI agents and marketing automation plan and measure it.
XenGrowth on AI search, GEO and discovery — what you'll learn: how the teams who own AI search, GEO and discovery plan and measure it.
Where this work meets go-to-market
The operational playbooks that sit alongside decision fatigue live with XenGrowth, who work on the commercial side of this.
Four questions on the research behind — and against — the popular 'decision fatigue' story.









