Why Do Technical Leaders Run Out of Good Decisions by Afternoon?
Career

Why Do Technical Leaders Run Out of Good Decisions by Afternoon?

The famous study behind 'decision fatigue' — judges granting parole less often as the day wears on — has been seriously challenged, and the whole 'willpower is a depleting resource' idea failed to replicate across 23 labs. The afternoon slump is real. The explanation everyone reaches for probably isn't.

Published November 23, 20259 min readUpdated Nov 23, 2025

Written by · Full-Stack Agentic AI Software Engineer — AI Agents, Automation & Revenue Systems for GTM/RevOps teams

In brief

Is 'decision fatigue' — running out of willpower to make good calls after a long day — actually real?

The specific evidence usually cited for it is on much shakier ground than the term's popularity suggests. The famous 2011 study finding Israeli parole judges granted parole less often as a session wore on has been directly challenged by a later simulation showing the pattern is consistent with a statistical artifact of case ordering, and the broader 'ego depletion' theory behind decision fatigue — that willpower is a single resource that drains with use — failed to replicate in a 2016 registered report across 23 labs and over 2,100 participants. None of that means a technical leader's afternoon judgment is imaginary; it means the popular explanation is probably wrong. The better-supported mechanism is cumulative cognitive load from fragmented, collaboration-heavy work: Meyer et al.'s instrumented data put collaborative activity at 24.4% of a developer's day, and Mark, Gudith and Klocke's research found that compensating for interruption and pressure carries a real, measurable stress cost that accumulates across a session. A leader's afternoon isn't running out of a willpower tank. It's carrying the compounding cost of a day built almost entirely out of context switches.

  • The famous parole-judge 'decision fatigue' study (Danziger, Levav & Avnaim-Pesso, 2011) found favorable rulings dropping from roughly 65% to near zero across a session, recovering after a food break — a real, striking pattern
  • A later simulation study directly challenged the interpretation, showing the same pattern is consistent with a statistical artifact of case ordering (unrepresented prisoners, less likely to be granted parole regardless, tended to be scheduled later in a session)
  • The broader theory behind 'decision fatigue' — ego depletion, the idea that willpower is a single resource that drains with each decision — failed to replicate in a 2016 preregistered report across 23 independent labs and 2,141 participants
  • This does not mean judgment never degrades across a day — it means the specific mechanism popularly cited for it is on much weaker ground than its ubiquity implies
  • A better-supported explanation for a leader's afternoon decline is accumulated cognitive load from collaborative, fragmented work, not a depleting willpower reserve — the fix that follows from each explanation is different

Evidence notes

Danziger, Levav & Avnaim-Pesso, 'Extraneous factors in judicial decisions' (PNAS, 2011)

Analysis of roughly 1,100 Israeli parole board rulings over 10 months found the probability of a favorable ruling dropping from about 65% at the start of a session to near zero by its end, then jumping back to about 65% after a food break.

Glöckner, 'The irrational hungry judge effect revisited: Simulations reveal that the magnitude of the effect is overestimated' (Judgment and Decision Making, 2016)

A simulation study argued the observed pattern is consistent with a statistical artifact: favorable rulings take longer to process than unfavorable ones, and unrepresented prisoners — who are granted parole less often regardless of timing — tend to be scheduled later in a session. Both factors alone can reproduce a similar-looking decline without invoking depleted willpower.

Hagger et al., 'A Multilab Preregistered Replication of the Ego-Depletion Effect' (Perspectives on Psychological Science, 2016)

23 laboratories, 2,141 participants, using a standardized protocol to test the core ego-depletion effect underlying 'decision fatigue' theory. The registered replication found no significant ego-depletion effect, prompting the replication team to argue the theory needs to specify conditions under which it reliably appears, since this large, preregistered attempt did not find one.

Meyer, Barton, Murphy, Zimmermann & Fritz, 'The Work Life of Developers' (IEEE TSE, 2017)

20 developers, 4 companies, 220 instrumented work days. Collaborative activity — meetings, ad-hoc conversation, email — totalled 24.4% of the measured workday, more than coding itself, with activity switches every 0.3 to 2.0 minutes outside planned meetings.

Every technical leader knows the feeling: sharp decisions in the morning, mushy ones by 4pm. The explanation everyone reaches for is decision fatigue — willpower as a tank that empties with each choice, running dry by the end of a long day. It's a great story. The evidence behind it has taken a beating that the story's popularity doesn't reflect.

This is worth sorting out properly rather than waving away, because the two candidate explanations point at completely different fixes. If willpower genuinely depletes with each decision, the fix is rationing decisions — batching them, delegating them, deferring the unimportant ones. If the real driver is something else entirely, rationing decisions might miss the actual lever completely while leaving the real cause untouched. I write about decision fatigue from the build side; XenGrowth, who work on the commercial side of this covers what it takes to run it.

The study everyone cites

Nearly every popular treatment of decision fatigue starts in the same place, so it's worth starting there too and following the evidence more carefully than the popular version usually does.

Danziger, Levav and Avnaim-Pesso's 2011 analysis of roughly 1,100 Israeli parole rulings found something striking: the probability of a favorable ruling started each session around 65%, fell steadily to near zero by the session's end, and jumped straight back to about 65% right after a food break. It reads like a perfect natural experiment in mental depletion, and it's been cited that way constantly since — including, often, as the anchor evidence for decision fatigue in general.

That's the version everyone stops at. There's a second half of the story that the popular retellings almost never mention, and it changes the conclusion substantially.

A later simulation study by Glöckner took the same data and asked a harder question: could this pattern show up without any depletion at all, purely from how the day's cases were scheduled? Two facts turned out to matter. Favorable rulings simply take longer to process than denials, so more of them get processed early in a session before time runs short. And unrepresented prisoners — who are granted parole less often regardless of when they're heard — tend to be scheduled later in the day. Simulate a court session with just those two structural facts, no fatigue mechanism required, and you get a decline that looks a lot like the original finding.

Interpretation

What it requires

Status

Judges deplete a willpower reserve across a session

A mental resource that drains with each decision and refills with food

The underlying theory failed a large, preregistered replication attempt

The pattern reflects case-scheduling structure, not depletion

Favorable cases take longer + unrepresented prisoners scheduled later

Reproduces a similar-looking decline in simulation without any fatigue mechanism

Neither of these fully settles the question of what was actually happening in that courtroom. What the challenge does establish is that the parole study is nowhere near the clean, mechanism-proven demonstration of decision fatigue it's usually presented as. On the operations side of this specifically, The XenGrowth resource library is worth reading.

The bigger problem: the theory underneath it didn't replicate

Decision fatigue is usually presented as a specific case of a broader theory called ego depletion — the idea that self-control and willpower draw from one shared, limited resource, so any act of self-control leaves less available for the next one. In 2016, Hagger and colleagues ran a preregistered replication of the core ego-depletion effect across 23 independent labs, with 2,141 participants and a standardized protocol agreed on in advance.

The replication found no significant ego-depletion effect. Not a smaller effect. Not a mixed result. Across 23 labs, using the field's own agreed-upon method, the core mechanism decision fatigue depends on didn't show up.

This is a genuinely awkward finding for a term as widely used as decision fatigue. It doesn't prove judgment never degrades over a day. It does mean the specific, popular mechanism — willpower as a depletable tank — is standing on much shakier ground than the term's ubiquity would suggest.

Why a large, high-profile failure like this matters beyond one study

It's worth pausing on why the 2016 replication failure matters more than a single disappointing result usually would. Ego depletion wasn't a fringe idea — by the mid-2010s it had generated hundreds of published studies, spawned popular books, and become a stock explanation in business writing for everything from why judges rule harshly to why dieters cave in the evening. A 23-lab preregistered replication, agreed on in advance by researchers on both sides of the debate, using a standardized protocol, is about as strong a test as social psychology has the infrastructure to run. Finding nothing under those conditions doesn't just weaken one study — it weakens the entire category of claims that lean on the theory, decision fatigue included, unless those claims can point to independent support that doesn't route through ego depletion.

None of this means every popular claim built on ego depletion is automatically false. It means the burden of proof sits with the claim, not with the skeptic, which is the opposite of how 'decision fatigue' usually gets deployed in practice — as an assumed background fact that other arguments get built on top of without anyone checking the foundation. For the AI agents and marketing automation angle, see XenGrowth on AI agents and marketing automation.

What's actually well-supported instead

Set the shaky mechanism aside and there's still a real phenomenon to explain, and the evidence for a specific alternative is stronger than the evidence for the one everyone defaults to.

The afternoon slump a technical leader feels is very likely real. The better-supported explanation just isn't a draining willpower tank — it's accumulated cognitive load from a day built almost entirely out of context switches. Meyer, Barton, Murphy, Zimmermann and Fritz's instrumented study found collaborative activity — the meetings, escalations, one-off questions and interruptions that fill a leader's day disproportionately — at 24.4% of the measured workday, more than coding itself, with switches happening every 0.3 to 2.0 minutes outside planned meetings.

Layer in Mark, Gudith and Klocke's finding that people compensate for interruption and pressure by working faster and harder, at a real cost in measured stress and effort that showed up within 20 minutes. A leader's day is disproportionately built from exactly the kind of activity that research shows carries a real, accumulating cost — not from a fixed number of 'decisions' draining a shared tank.

Explanation

What it predicts

Evidentiary status

Willpower depletes with each decision made

Judgment quality should track raw decision count, regardless of what filled the time between decisions

Core mechanism failed a 23-lab, 2,141-participant preregistered replication

Accumulated cognitive load from context-switching

Judgment quality should track how fragmented the day was, regardless of raw decision count

Consistent with instrumented data on collaborative load (24.4%) and switch frequency (0.3-2.0 min)

The two explanations make genuinely different predictions, which is what makes this worth getting right rather than treating as a semantic quibble. If willpower depletion is the mechanism, a leader who makes ten decisions in a quiet, uninterrupted morning should be just as depleted as one who makes ten decisions across a morning of constant escalations and pings. If accumulated context-switching is the mechanism, those two mornings should look nothing alike by early afternoon, even though the decision count is identical. Anecdotally, most technical leaders report the second pattern, not the first — a calm morning with several real decisions feels nothing like a chaotic one with the same number, which is itself a small piece of evidence favoring the fragmentation account over the depletion account. XenGrowth on AI search, GEO and discovery approaches this from the AI search, GEO and discovery side.

What actually follows from getting the mechanism right

  1. Stop treating 'decision fatigue' as settled science when you plan around it. The specific mechanism it's usually attributed to has a serious replication problem, which matters for how confidently you should generalize it

  2. Address the actual candidate cause instead: reduce the raw number of context switches a leader absorbs in a day, not the raw count of decisions, since it's the switching — not decision-making per se — that has the stronger evidentiary backing

  3. Group similar decisions together rather than scattering them across the day. Nothing in the failed-replication evidence argues against batching reducing cognitive load; it argues against the specific 'depleting resource' story for why batching might help

  4. Treat a food break's apparent effect in the original study with appropriate skepticism as evidence for anything beyond 'a break helps' — which barely needs a depletion theory to explain

  5. If you're building a process around 'protecting decision quality,' design it around measurable fragmentation (meeting density, interruption frequency) rather than an unmeasurable internal resource nobody has been able to reliably detect in a lab

A word on the food-break detail specifically

The original parole study's most quoted image — approval rates snapping back up right after judges ate — deserves its own honest look, because it's usually presented as the clinching detail. A break restoring performance is consistent with willpower depletion, but it's also consistent with almost any other fatigue-adjacent mechanism: simple physical tiredness, low blood sugar affecting mood and patience, or even just a change in the physical environment breaking up monotony. 'A break helped' is compatible with a very large number of explanations, most of which don't require inventing a depletable psychological resource to account for it. Treating the food-break detail as decisive evidence for ego depletion specifically, rather than for 'breaks are generally good,' is exactly the kind of overreach this post is arguing against.

This matters for technical leadership directly. If a leader notices their judgment improves after lunch, the accurate conclusion isn't 'my willpower reserve refilled.' It's one of several plausible, boring explanations — physical fatigue easing, blood sugar stabilizing, or simply having stepped away from the specific fragmented stretch of meetings that preceded it — none of which require believing in a resource that a large registered replication failed to find.

The honest version of this story is less tidy than 'willpower runs out.' It's also more useful: if the real driver is accumulated context-switching rather than a fixed decision budget, the fix isn't rationing how many choices you make. It's rationing how many times your attention gets yanked somewhere new before you're asked to make the next one. That's a scheduling problem, not a metaphysical one, and scheduling problems are the kind you can actually solve.

Further reading from XenGrowth

Where this work meets go-to-market

The operational playbooks that sit alongside decision fatigue live with XenGrowth, who work on the commercial side of this.

What does the decision fatigue evidence actually show?

Four questions on the research behind — and against — the popular 'decision fatigue' story.

1 / 4
What did Glöckner's later simulation study find about the famous parole-judge decision fatigue result?

Apply this article

How to turn insights into execution

A practical sequence for teams turning concepts into production outcomes.

Decision FatigueCognitive LoadLeadershipResearchBurnoutCareerscareer

Audit your current state

Map the bottlenecks and constraints connected to the article’s core problem.

Choose one bounded change

Test the most useful recommendation on one workflow before widening the scope.

Measure what changed

Keep the parts that improve the work, document what failed, and make the next decision from evidence.

Next step

Need help applying this in your stack?

I can translate these patterns into a concrete implementation plan for your team.

Discuss implementationBack to blog

Replies usually within 24 hours.

Next Steps

Continue reading

How Much Does a Single Interruption Really Cost?

The number everyone quotes — multitasking costs you 40% of your productivity — is real, but it isn't from the study everyone cites it from. The study measured something smaller, stranger, and more useful.

Navigate

Why Do Meetings Hit Engineers Harder Than Other Roles?

A 30-minute meeting doesn't cost an engineer 30 minutes. It costs the meeting, the time spent rebuilding the mental model it interrupted, and whatever fraction of that model doesn't come back intact.

Navigate

How Much of Engineering Fatigue Is Just Unclear Requirements?

A lot of what gets labeled 'this project is exhausting' is actually 'this project keeps making me redo work because nobody decided what it should do.' Those have different fixes, and only one of them is about you.

Navigate

Does Being On-Call Cost You Even on Quiet Nights?

A pager that never goes off should be a free night's sleep. Two separate sleep-lab studies say it usually isn't, and the reason has nothing to do with how many alerts actually fired.

Navigate

How Long Should a Deep Work Block Actually Be?

The '90-minute focus cycle' gets quoted as settled science. The research it's built on is real, genuinely interesting, and considerably less precise than the number implies.

Navigate

Is Burnout a Medical Diagnosis?

The WHO added burn-out to the ICD-11 in 2019 and most of the coverage since has gotten the headline backwards. It is not a disease. The actual entry says so in its own second sentence, and almost nobody quotes that part.

Navigate

Why Doesn't Working Fewer Hours Fix Burnout?

Cut the hours and the exhaustion dimension might ease a little. The other two dimensions in the actual clinical model of burnout don't have anything to do with hours, and they're usually the ones that don't move.

Navigate

Am I Tired, Bored, or Actually Burnt Out?

All three feel like 'I don't want to do this today.' They aren't the same problem, and the research says you're a genuinely unreliable judge of which one you're in.

Navigate

How Do Engineers Recover From Burnout Without Quitting the Industry?

Most burnout advice assumes the fix is either 'push through' or 'leave.' The Maslach-Leiter model points at a third option: find which specific area of the job is actually mismatched, and change that one thing.

Navigate