Will AI Cut Engineering Jobs, or Multiply Their Leverage?
Career

Will AI Cut Engineering Jobs, or Multiply Their Leverage?

Both answers are already true, for different people. The payroll data shows a 19% employment gap opening for 22-to-25-year-olds in AI-exposed jobs while experienced workers show no gap at all. That split is the actual story, and it is not the one either side of the argument is telling.

Published July 8, 202611 min readUpdated Jul 8, 2026

Written by · Full-Stack Agentic AI Software Engineer — AI Agents, Automation & Revenue Systems for GTM/RevOps teams

In brief

Is AI reducing the number of software engineering jobs, or making individual engineers more productive?

Both, and the division runs along experience rather than across the profession as a whole. Stanford's Digital Economy Lab, working from ADP payroll records covering millions of US workers, finds employment for 22-to-25-year-olds in AI-exposed occupations sitting about 19% below where it would be had it tracked their less-exposed peers, with no equivalent gap for experienced workers. The mechanism is reduced hiring, not layoffs. Meanwhile the BLS still projects 15% growth for software developers over 2024-2034, roughly 288,000 added roles. Those two figures are not in conflict: the total keeps growing while the entry door narrows. On the leverage side, the evidence is far weaker than the marketing suggests. METR's randomized trial found experienced open-source developers were 19% slower with AI tools while believing they had been 20% faster.

  • The displacement is concentrated in early-career hiring, not spread evenly across the profession, and it works through jobs never posted rather than through layoffs
  • BLS still projects 15% growth and about 129,200 annual openings for software developers through 2034, so the aggregate story is not contraction
  • METR's RCT is the strongest evidence against the leverage claim, and its most useful finding is the perception gap: developers were wrong about their own speed by roughly 39 percentage points
  • DORA's 2025 data shows AI adoption now correlates with higher throughput and, at the same time, lower delivery stability — speed and safety moved in opposite directions
  • Stanford's own data separates occupations where AI substitutes for tasks from ones where it complements them; employment fell in the first group and held or rose in the second, which is the closest thing to career advice in the whole dataset

Evidence notes

Stanford Digital Economy Lab, 'Canaries in the Coal Mine?'

Brynjolfsson, Chandar and Chen use high-frequency ADP payroll data covering millions of US workers through June 2026. Employment for ages 22-25 in the most AI-exposed occupations runs roughly 19% below the counterfactual set by less-exposed peers; experienced workers in the same occupations show no comparable gap. The divergence has widened steadily since August 2025 and operates mainly through reduced hiring rather than increased separations.

BLS Occupational Outlook, software developers

The Bureau of Labor Statistics projects 15% employment growth for software developers, QA analysts and testers from 2024 to 2034 — 'much faster than average' against a 4% all-occupation baseline — adding about 287,900 positions from a 2024 base near 1.7 million, with roughly 129,200 openings each year counting replacements.

METR randomized controlled trial, July 2025

Sixteen experienced open-source developers worked 246 real issues drawn from their own repositories (averaging 22k+ stars, 1M+ lines), each issue randomly assigned to allow or forbid AI tooling. Developers forecast a 24% speedup beforehand and estimated a 20% speedup afterward; measured completion time was 19% slower with AI available.

DORA 2025, State of AI-assisted Software Development

In 2024 DORA measured every 25% rise in AI adoption against roughly a 1.5% drop in throughput and a 7.2% drop in delivery stability. In the 2025 data the throughput sign flipped positive while stability did not recover. DORA's framing is that AI amplifies whatever a team already is rather than fixing it.

Stack Overflow Developer Survey 2025

AI tool use or planned use reached 84%, up from 76% in 2024, while trust in output accuracy fell from 40% to 29% and the share actively distrusting it rose from 31% to 46%. The most-cited frustration was AI solutions that are 'almost right, but not quite,' which respondents reported as making debugging slower rather than faster.

Continue with purpose

The honest answer is that both things are happening, to different people, and almost every version of this argument you'll read picks one group and generalizes it to the whole profession. The displacement camp points at collapsing graduate hiring. The leverage camp points at their own week. Neither is lying. They're describing different ends of the same distribution.

So this post does something narrower and, I think, more useful: it takes the four best datasets we currently have, states what each one actually measured, and shows where they agree. They agree more than the discourse suggests. What they agree on is uncomfortable for both camps. Much of the judgment software engineering demands shows up as process design, which is what XenGrowth's work on go-to-market systems publishes on.

What does the employment data actually show?

The most credible number comes from Stanford's Digital Economy Lab, and it's credible for a boring reason: it isn't a survey. Brynjolfsson, Chandar and Chen used ADP payroll records — actual paychecks for millions of US workers — rather than asking anyone how they felt about AI. Their headline finding is that employment for 22-to-25-year-olds in the most AI-exposed occupations sits about 19% below where it would be had it tracked their less-exposed peers.

The part that gets dropped in the retelling: experienced workers in those same occupations show no comparable gap. None. The effect is almost entirely concentrated in the youngest cohort.

And the mechanism matters as much as the magnitude. This is happening through reduced hiring, not increased separations. Nobody is being marched out. The jobs are simply not being posted, which is a much quieter process and one that produces no news story, no layoff announcement, and no visible cohort of displaced workers to interview. It shows up as a graduate who sends 300 applications and hears back from four.

A hiring freeze aimed at one cohort looks like nothing at all from the inside of a company. It looks like a headcount plan that quietly got shorter.

Then why is the BLS still projecting 15% growth?

Because it's measuring a different thing, over a longer window, and it isn't wrong. The Bureau of Labor Statistics projects software developer employment growing 15% from 2024 to 2034 — about 287,900 added positions against a 2024 base near 1.7 million, with roughly 129,200 openings a year once you count replacement of people retiring or leaving. Against a 4% all-occupation average, that's still filed under 'much faster than average.'

There's no contradiction here once you separate the size of a profession from the shape of its entry. A field can grow in total headcount while its bottom rung gets pulled up. That's the configuration the two datasets jointly describe: more engineers overall, fewer ways in. Google’s Generative AI Report Is Here. Read It Without Inventing a Story. covers the AI search measurement side of this.

Question

What the data says

Source

Is the profession shrinking?

No — 15% projected growth to 2034, ~129,200 openings a year

BLS Occupational Outlook

Is entry-level hiring shrinking?

Yes — ~19% employment gap for ages 22-25 in AI-exposed roles

Stanford Digital Economy Lab

Are experienced engineers affected?

No measurable employment gap in the same occupations

Stanford Digital Economy Lab

Are individual engineers faster with AI?

Experienced devs measured 19% slower on their own repos

METR RCT

Are teams shipping more?

Yes, throughput up — but delivery stability down

DORA 2025

Does the leverage claim survive contact with measurement?

Less well than you'd expect. METR ran the study the industry should have run years earlier: sixteen experienced open-source developers, 246 real issues from their own repositories, each issue randomly assigned to allow or forbid AI tooling. Not a benchmark, not a toy task — actual work on codebases averaging over a million lines that these people already knew intimately.

They were 19% slower with AI available.

I want to be careful about how far that generalizes, because it's a small sample on mature codebases with experts who already had the whole architecture in their heads — close to the worst case for a tool whose main advantage is knowing things you don't. On an unfamiliar codebase, or in a language you're rusty in, the result would likely flip. Treat it as one strong data point, not a law.

But the finding underneath the headline generalizes further than the headline does. Those developers predicted a 24% speedup going in, and estimated a 20% speedup coming out — after doing the work. They were off by roughly 39 percentage points about their own labor, in the direction of flattering the tool. That's the number worth carrying around. It means your sense that AI is making you faster is not evidence that AI is making you faster.

Stack Overflow's 2025 survey suggests why the illusion is so durable. Adoption climbed to 84% while trust in accuracy fell to 29%, and the single most-cited frustration was output that is 'almost right, but not quite.' Almost-right code feels like progress at the moment it appears. The cost lands later, in a debugging session nobody attributes back to the tool that caused it. International Keyword Research Starts Where the Translation Sheet Stops covers the international search research side of this.

What did DORA find, and why does it complicate both stories?

DORA's 2025 report is the most interesting of the four because it changed its own mind. In the 2024 data, every 25% rise in AI adoption tracked with roughly a 1.5% drop in throughput and a 7.2% drop in delivery stability — both bad. A year later the throughput sign flipped positive. Teams really are shipping more.

Stability didn't recover. It's still going the wrong way. Which means the thing AI reliably delivers is volume, and the thing it reliably costs is confidence that what shipped works. DORA's own framing is that AI amplifies whatever a team already is: disciplined teams get faster, undisciplined teams get faster at producing incidents.

GitClear's code analysis points at the mechanism. Across their corpus, copy-pasted lines rose from 8.3% in 2020 to 12.3% in 2024 while refactored — 'moved' — lines fell from 24.1% to 9.5%. 2024 was the first year on record where within-commit copy/paste exceeded moved code. That's a codebase getting larger without getting more organized, which is exactly the shape of output you'd expect from a tool that generates locally-plausible code with no memory of what already exists twelve directories away.

Study

What it actually measured

Where it stops

Stanford / ADP payroll

Real paychecks for millions of US workers through June 2026, by age and AI exposure

Occupation-level exposure scores are a proxy; it cannot tell you which tasks inside a job moved

BLS projections

Modelled headcount for a ten-year window, 2024-2034

A projection, not an observation, and it is silent on the seniority mix inside the total

METR RCT

Completion time on 246 real issues, randomized, experts on their own mature repos

Sixteen developers, early-2025 tooling, and close to the worst case for AI: people who already knew the codebase

DORA 2025

Self-reported delivery metrics across a large practitioner survey

Correlational. AI adoption and delivery instability move together; the report does not establish which causes which

GitClear

Static analysis of commit diffs across a large public corpus

Measures the shape of code churn, not whether the resulting software was any good

So which side of the split are you on?

Stanford's paper contains the closest thing to actionable advice in any of this, and it's buried in a distinction most summaries skip. They separated occupations where AI usage mostly substitutes for human tasks from ones where it mostly complements them. Employment fell in the first group. In the second, it was flat or rising. On AI search measurement specifically, GA4 Has an AI Assistant Channel. It Still Does Not Measure Every AI Search Visit. is worth reading.

That's not a statement about job titles. It's a statement about what your day actually consists of. Two people with the same title can sit on opposite sides of that line.

  • Substitution-shaped work: implementing a spec someone else wrote, translating a ticket into code, producing a known component in a known pattern, writing tests against defined behavior

  • Complement-shaped work: deciding what should be built at all, choosing between architectures whose tradeoffs won't surface for eighteen months, debugging something nobody has a name for yet, negotiating a requirement with the person who wants it

  • The tell is whether your work has a checkable right answer at the moment you start it. Substitution-shaped work does. Complement-shaped work is mostly the process of finding out what the answer even needs to be

  • Most engineering jobs contain both. The question is the ratio, and whether it's moving

This is also why the junior gap is real and not a moral failing of a generation. Entry-level work is defined by being substitution-shaped. That's the whole point of it — you're given bounded problems with known answers precisely so a more experienced person can check your work. AI is very good at bounded problems with known answers. The rung that got removed is the one built out of exactly the tasks the tool does cheapest.

One more caution about all five of these at once. Every dataset here was collected against tooling that has since been replaced, in a period where model capability moved faster than anyone's ability to study it. METR's trial ran on early-2025 assistants. Stanford's payroll window closes in June 2026. Read them as measurements of a moving object, taken from behind, which is the only place measurement can ever be taken from.

That cuts both ways, and I'd be suspicious of anyone who only invokes it in one direction. The people telling you METR is obsolete because the models improved are making the same unfalsifiable move as the people who said in 2023 that displacement was a year away. If the answer to every measurement is that the next version will be different, then no evidence can ever land, and the argument stops being about the world.

What actually follows from all this?

  1. Stop treating your own sense of speed as evidence. METR's participants were experts, motivated, and wrong by 39 points. If you want to know whether a tool helps, measure something — cycle time on comparable tickets, review rounds per PR, incidents per deploy — rather than consulting your impression

  2. Watch stability, not throughput. DORA's finding is that the two decoupled. If your team ships more and your change-failure rate drifts up, the tool is working exactly as measured and you are still losing

  3. Audit your own week against the substitution/complement line. Not your title, your calendar. If most of what you did had a checkable right answer at the moment you started, that is the part of the market under pressure

  4. If you hire, notice that the junior gap is a hiring-plan artifact, not a law of nature. The rung got removed by a thousand individually reasonable headcount decisions, and it can be rebuilt the same way

  5. Do not read BLS growth as reassurance. Aggregate growth in a profession says nothing about whether its entry door is open, and right now those two numbers are pointing in different directions

The framing I'd resist hardest is the one where this is a single question with a single answer. It isn't. There is a version of the next five years where you personally have more leverage than any engineer in history, and a version where the work you're good at is the work that got cheapest, and which one you get is substantially decided by what fraction of your job survives having a checkable right answer.

The 19% gap isn't a prediction. It's already in the payroll data, through June 2026. The argument about whether this will happen is over.

Further reading from XenGrowth

Where this work meets go-to-market

Working on software engineering inside a commercial team? the XenGrowth practice publishes operator guides on the revenue side of this work.

Which side of the line is your work on?

Five questions against the substitution/complement distinction from the Stanford paper. It asks about the shape of your actual week, not your job title, because two people with the same title routinely land on opposite sides.

1 / 5
For most of what you did last week, did a checkable right answer already exist at the moment you started?

Not whether it was easy — whether someone could have graded it against a spec.

Apply this article

How to turn insights into execution

A practical sequence for teams turning concepts into production outcomes.

AICareersSoftware EngineeringLabor MarketProductivityResearchcareer

Audit your current state

Map the bottlenecks and constraints connected to the article’s core problem.

Choose one bounded change

Test the most useful recommendation on one workflow before widening the scope.

Measure what changed

Keep the parts that improve the work, document what failed, and make the next decision from evidence.

Next step

Need help applying this in your stack?

I can translate these patterns into a concrete implementation plan for your team.

Discuss implementationBack to blog

Replies usually within 24 hours.

Next Steps

Continue reading

Coding Is the Smallest Part of Software Engineering

When researchers put monitoring software on 20 professional developers' machines for 220 work days, coding came out at 21% of the day. Not because those developers were slacking — because the other 79% is the job. AI automates a slice of the 21%.

Navigate

Are Junior Developer Jobs Disappearing? What the Data Says

Entry-level hiring at the tech majors is down 65% since 2019 and Stanford measures a 19% employment gap for 22-to-25-year-olds. But an LSE paper covering 243 million hires found that when you control for remote work, the AI effect largely vanishes. The cause matters, because the two have opposite fixes.

Navigate

The Software Engineer of 2030 Will Look Different

Most predictions about this are unfalsifiable, so here are five that aren't. Each one names what would have to be true, and what evidence would prove it wrong — including the two I think are most likely to age badly.

Navigate

Debugging Is Becoming More Valuable Than Writing Code

Stack Overflow's 2025 survey found the top developer frustration wasn't AI being wrong. It was AI being almost right — output that compiles, looks correct, and costs you an afternoon. That failure mode moves work out of writing and into diagnosis, and diagnosis was already the expensive half.

Navigate

How Long Should a Deep Work Block Actually Be?

The '90-minute focus cycle' gets quoted as settled science. The research it's built on is real, genuinely interesting, and considerably less precise than the number implies.

Navigate

How Much Does a Single Interruption Really Cost?

The number everyone quotes — multitasking costs you 40% of your productivity — is real, but it isn't from the study everyone cites it from. The study measured something smaller, stranger, and more useful.

Navigate

What to Learn When AI Can Already Write the Code

The useful question isn't what AI can do — it's what it structurally cannot. Veracode ran 100+ models across 80 tasks and 45% of the output carried an OWASP Top 10 vulnerability, with larger models no better than small ones. That failure has a shape, and the shape tells you what to learn.

Navigate

What Sleep Debt Does to Engineering Judgment

The finding that should worry you isn't that six hours of sleep degrades performance. It's that in the study which established it, subjective sleepiness stopped tracking objective decline — the impaired group did not know they were impaired.

Navigate

Should Software Engineers Become AI Engineers?

Mostly no — and the reason is in the data people cite to argue yes. The Stanford AI Index finds the fastest-growing AI skills are deployment ones: AWS, scalability, workflow management. The market is short of engineers who can ship these systems, not people who understand them.

Navigate
  • What Happens When One Engineer Does the Work of Five?

    The claim gets made constantly and almost never with a number attached. When someone did attach numbers — METR's randomized trial — experienced developers came out 19% slower while believing they were 20% faster. But suppose the claim were true. The consequences are stranger than the people making it seem to expect.