Managers
Field research · 2025–2026

Managers AI Evaluation Study

AI is writing a growing share of production code, which raises a question nobody has fully answered yet: how do you evaluate engineers now? I interviewed 35 software engineering managers about what performance means today, how far they trust the data, and who takes the blame when a review goes wrong.

Stockholm University
Conducted as a bachelor’s thesis at
Stockholm Business School, Stockholm University
35
engineering managers interviewed
14
sectors, from fintech to health tech
17
hours of recorded conversation
300
engineers evaluated by the people I spoke to
01 · The context

Every manager in the study uses AI daily. So do their teams.

This is not an early-adopter story. Copilot, Claude Code and Cursor are standard issue in the organizations I studied, usually rolled out top-down and often with board-level conviction. One company gave every engineer a premium AI subscription and wrote the cost off as table stakes. Another reorganized its engineering function around AI entirely. The managers in this study are not speculating about a future workflow. They evaluate people inside it.

It does not matter if you spend a lot of money on tokens. We think we are going to get that back in the future.
Engineering manager, marketing tech, on his board's company-wide AI mandate
Copilot or equivalent assistants28 of 35
Agentic tools (Claude Code, Cursor)22 of 35
Internal or self-hosted LLMs14 of 35
Custom AI evaluation pipelines6 of 35
Tools managers reported in active use on their teams. Most run several at once.
01
ChatGPT replaces search
Engineers quietly swap Stack Overflow for chat. No policy, no budget line.
02
Org-wide Copilot
Autocomplete becomes standard issue. Security signs off, procurement follows.
03
AI insights teams
Dedicated teams wire AI into delivery metrics, dashboards and reporting.
04
Agents for everyone
Premium AI-agent subscriptions for every engineer, written off as table stakes.
02 · The study

35 conversations with the people who judge engineers

Semi-structured interviews, 25 to 45 minutes each, with managers running teams of 4 to 20 engineers across fintech, SaaS, cloud infrastructure, health tech and ten other sectors. Everyone in the sample carries real evaluative authority: reviews, compensation, promotion. Each conversation covered how they evaluate today, where AI has entered the process, what they do when their judgment disagrees with the data, and who answers when an evaluation goes wrong.

IDRoleSectorTeamYrs managing
M01Engineering ManagerFintech86
M02Senior Engineering ManagerFintech129
M03Engineering ManagerFintech53
M04Engineering ManagerSaaS / CRM87
M05Engineering ManagerSaaS / CRM64
M06Director of EngineeringSaaS / CRM1612
M07Engineering ManagerSaaS / CRM95
M08Engineering ManagerCloud infrastructure98
M09Engineering ManagerCloud infrastructure75
M10Senior Engineering ManagerCloud infrastructure1411
M11Engineering ManagerHealth tech42
M12Engineering ManagerHealth tech129
M13Senior Engineering ManagerHealth tech107
M14Engineering ManagerDeveloper tools64
M15Director of EngineeringDeveloper tools1814
M16Engineering ManagerE-commerce116
M17Engineering ManagerE-commerce73
M18Senior Engineering ManagerE-commerce1510
M19Engineering ManagerAutomotive software1412
M20Engineering ManagerAutomotive software97
M21Engineering ManagerMedtech software65
M22Engineering ManagerMedtech software108
M23Engineering ManagerBanking tech86
M24Senior Engineering ManagerBanking tech1311
M25Engineering ManagerBanking tech52
M26Engineering ManagerLogistics tech94
M27Engineering ManagerLogistics tech128
M28Engineering ManagerGaming106
M29Senior Engineering ManagerGaming1613
M30Engineering ManagerCybersecurity75
M31Engineering ManagerCybersecurity119
M32Engineering ManagerMarketing tech1410
M33Engineering ManagerMarketing tech63
M34Former EM (role abolished)Micromobility7
M35Engineering ManagerMicromobility85
All 35 participants. Anonymized, with roles and sectors lightly generalized to protect identity. Interviews ran 25 to 45 minutes.
03 · Finding one

“Everybody can churn out code now.” Output stopped being the signal.

When AI produces the baseline, raw volume stops telling you anything. 26 of 35 managers said their definition of performance changed materially within the last 18 months. What rose: architectural judgment, risk handling, and what one manager called the multiplier factor, the ability to make everyone around you better. What fell: lines shipped, tickets closed, velocity for its own sake.

The base level of output is no longer sufficient. Everybody can churn out code.
Engineering manager, CRM/SaaS, 8 reports
share of evaluative weight
Raw output volume32%
Delivery speed24%
Code quality & review depth18%
Architectural judgment & risk14%
Multiplier effect: unblocking, mentoring12%
Directional weights, synthesized from how 35 managers described their evaluation criteria before and after AI tools entered the workflow.

The metrics also miss the work that holds teams together. Several managers described engineers who look weak in raw numbers while quietly carrying coordination, mentoring and risk-handling for everyone else. And more than one pointed out that the real bottleneck has moved: AI shifted it from writing code to reviewing and integrating it, which is exactly the work the dashboards weigh least.

26 of 35
said what counts as engineering performance changed materially within 18 months.
04 · Finding two

AI didn't shrink the work of evaluation. It moved it.

The efficiency story goes like this: AI gathers the evidence, the manager saves time. The managers describe something else. 24 of 35 report that evaluation takes more effort since AI entered the workflow, because the hours saved collecting signals now go into validating them. One manager built a fully automated pipeline that pulls delivery metrics, analyzes code and drafts performance summaries. He also described a standing Monday ritual of repairing his agents when they break. His verdict on his own system:

Quite seductive, in the sense that they give you substance that is not there.
Engineering manager, cloud infrastructure, who built his own AI evaluation pipeline
Before AI
45%
30%
25%
With AI
15%
40%
20%
25%
Gathering evidence
Validating AI signals
Interpreting
Conversations
Where managers said evaluation effort goes. The gathering work AI absorbed came back as validation work.
People trust dashboards too quickly because they look clean and rational. But performance is messy.
Engineering manager, health tech
24 of 35
say evaluation now demands more of them, not less
32 of 35
have caught AI being confidently wrong at least once
05 · Finding three

Input, not outcome

31 of 35 managers drew the same line without being asked: AI can inform an evaluation, it cannot make one. They cross-check it, override it, and treat scientific-looking output as the start of a conversation rather than the end of one. Nearly all of them described some routine for verifying AI claims before acting on them. "Trash in, trash out," as one put it.

AI as input only, human makes the call (31)open to AI-led scoring (4)
It feels easy to imagine people being compared just based on this very scientific-looking output that may not be quite so accurate.
Engineering manager, health tech, on an AI-generated report about a colleague's code reviews

The line gets sharper as decisions get more political. One manager split salary from promotion: salary is quantitative enough that AI could plausibly help, but promotions run on company interests that are written down nowhere a model can read them. The most consequential decisions are the least legible ones.

06 · Finding four

The hardest decisions never touched AI

Every manager could recall their hardest evaluation. An engineer unravelling because of troubles at home. A senior hire who could not land a role change. A strong performer who quietly stopped trying. Not one of them reached for an AI tool in those moments, including the heaviest adopters in the study. The decisive inputs were timing, trust and reading a person in the room.

01
The unravelling

Apparent underperformance turned out to be a crisis at home. The fix was a quiet pairing with a trusted colleague and time. No dashboard surfaced it, and none would have.

02
The wrong seat

A respected senior engineer took on a team-lead role and sank. The work was helping them step back without losing face in front of the team.

03
The slow fade

High effort, falling fit. The manager said dragging out a clearly negative review would have been unfair to the person and the team both.

It is not so much what you say, but how you say it. If you catch someone in a very low moment, you can completely destroy the relationship you have with that person.
Engineering manager, automotive software
I am more concerned about people shying away from responsibility than doing things wrong. We can always react.
Engineering manager, CRM/SaaS
07 · The question

When it goes wrong, who answers? Three camps.

Then I asked the question this study was built around. When an AI-shaped evaluation harms someone (a missed promotion, an unfair rating, a termination), who answers for it? Thirty-five managers, three camps. And when asked who answers for a negative review in their org today, the most common reply was one word: "I am."

60%
21 of 35 managers
The signer

Accountability follows the final human signature. Whoever approves the review, or merges the code, answers for it, no matter how much of it AI produced.

Once you push it to GitHub, it is you.

Engineering manager, CRM/SaaS

26%
9 of 35 managers
The deployer

Accountability moves up the chain to whoever mandated the system. If leadership forces tools on managers who cannot fully understand them, leadership owns the outcomes.

The person who decides to use the AI is ultimately accountable.

Engineering manager, industrial software

14%
5 of 35 managers
The system

Accountability dissolves into infrastructure: CI pipelines, AI review agents, engineered safeguards. The question stops being who failed and becomes what the system allowed.

It is not about the people, it is about the systems.

Former engineering manager, AI-native company

One manager grounded his position in a live incident: an AI coding agent had recently wiped a production database. His read was that the failure belonged to the humans who designed insufficient safeguards around it, not to the model.

Behind any computer mistake is a human mistake.

Engineering manager, cloud infrastructure

08 · The core finding

Formal accountability is holding. The conditions for it are eroding.

Put the findings together and a structural picture emerges. Managers remain fully, formally accountable for evaluations. They insist on it. But the three conditions that make accountability meaningful (understanding the basis of a decision, controlling the process, being able to explain it) weaken as AI mediates more of the evidence. Responsibility stays put while control drains out of it. Researchers have a name for where this ends: the moral crumple zone, the point where the nearest human absorbs blame for outcomes an entire system produced.

255075100202220232024202520262027projectedFormal accountabilityPractical controlTHE ACCOUNTABILITY GAP
Index of how managers described their position over time: what they formally answer for (black) against how much of the evaluative process they practically shape and can explain (orange). 2027 projected from their own expectations.
If I cannot explain the basis, I should not rely on it heavily.
Engineering manager, health tech

Every manager in this study is signing off on documents whose evidence base they can audit less and less. They know it. They sign anyway, because the alternative is to stop managing.

09 · The deep end

What the extreme cases look like

Two managers in the study sit at the far end of AI adoption, and what they describe sounds less like the future of performance reviews than their replacement. A third group is worried about a different destination entirely.

Case one

The company that deleted management

One participant, a former engineering manager, works at a company that abolished management outright. The CEO assesses performance by querying the codebase through Claude, and accountability lives in CI pipelines and AI review agents instead of people.

900 200
headcount through the restructure

No managers. No one-on-ones. No review cycle. Performance lives in the codebase.

Case two

Founder mode, all the way down

Another manager hires by watching candidates direct AI agents, and expects individual contributors to operate like founders, each running a small fleet of them. In his world the weak signal is not low output. It is what he calls AI slop: volume without judgment. The strong signal is an engineer whose agents compound.

The warning

The digital prison

A third of managers raised surveillance without being asked. AI makes total monitoring of engineers trivially cheap: every commit, every message, every pause. The people best positioned to deploy that future want no part of it.

I do not care if the employee is taking three coffee breaks in the morning or five in the afternoon, as long as the work is delivered.
Engineering manager, automotive software
AI-driven performance scoring is abdicating some of your managerial duties. Some of your human duties.
Engineering manager, fintech
10 · Takeaways

A few things worth sitting with

None of this comes with instructions. But a few patterns from the conversations seem worth passing on, depending on where you sit.

If you run a company
  • The managers who got the most out of AI were the ones free to override it. Tools that hand down verdicts were trusted least.
  • When leadership mandates the tools, some of the accountability seems to move with that decision. A quarter of the managers already see it that way.
  • The time saved gathering evidence tends to come back as validation work. Worth knowing before counting on efficiency gains.
If you lead engineers
  • Output metrics say less than they used to. Most managers in the study had already shifted their weight toward judgment and the ability to lift others.
  • Almost everyone drew the same line: AI informs the review, a person makes the call. The ones most at ease with their process were the ones who checked the AI's claims.
  • The hardest conversations stayed fully human for every manager I spoke to. Whatever else changes, that part seems likely to stay.
11 · My read

Where I think this is heading

The study ends where the data ends, but the direction feels legible. Within a few years the unit being evaluated will not be the engineer. It will be the engineer plus their agents. Judging the human in isolation already feels artificial to the managers furthest ahead, and the rest of the industry tends to follow them with a lag.

The management role splits in two. One half becomes an editor: cross-examining AI evidence, auditing what the dashboards claim, deciding what deserves trust. The other half is the part that never changed, reading people, timing hard conversations, noticing when a number is hiding a person in trouble. Managers who are only output supervisors sit directly in the path of automation. Managers who own both halves become more valuable, not less.

Accountability will get formalized one way or another. Either organizations write deployment accountability into how they adopt these tools, with named owners, real override rights and time budgeted for validation, or the gap keeps widening quietly until a lawsuit, a works council or a regulator closes it for them. The crumple zone is comfortable for everyone except the person inside it.

Most companies will not delete their managers. But every company will eventually re-answer the question this study kept running into: when the evidence is machine-made and the consequences are human, who signs? The managers who come out ahead will be the ones who keep insisting that the answer is a name. Starting with their own.

12 · About the data

What this is, and what it isn't

Findings draw on roughly 35 semi-structured interviews of 25 to 45 minutes with software engineering managers, conducted in 2025 and 2026 across 14 sectors and analyzed with reflexive thematic analysis. Participants are anonymized and quotes lightly edited for clarity. The percentages summarize qualitative coding of a purposive sample. Read them as directional findings from structured conversations, not survey-grade estimates. The underlying academic study was conducted at Stockholm Business School.

Want the full write-up, the data, or an argument?

aboodalzeno@gmail.com