Managers AI Evaluation Study
AI is writing a growing share of production code, which raises a question nobody has fully answered yet: how do you evaluate engineers now? I interviewed 35 software engineering managers about what performance means today, how far they trust the data, and who takes the blame when a review goes wrong.

Stockholm Business School, Stockholm University
Every manager in the study uses AI daily. So do their teams.
This is not an early-adopter story. Copilot, Claude Code and Cursor are standard issue in the organizations I studied, usually rolled out top-down and often with board-level conviction. One company gave every engineer a premium AI subscription and wrote the cost off as table stakes. Another reorganized its engineering function around AI entirely. The managers in this study are not speculating about a future workflow. They evaluate people inside it.
“It does not matter if you spend a lot of money on tokens. We think we are going to get that back in the future.”
35 conversations with the people who judge engineers
Semi-structured interviews, 25 to 45 minutes each, with managers running teams of 4 to 20 engineers across fintech, SaaS, cloud infrastructure, health tech and ten other sectors. Everyone in the sample carries real evaluative authority: reviews, compensation, promotion. Each conversation covered how they evaluate today, where AI has entered the process, what they do when their judgment disagrees with the data, and who answers when an evaluation goes wrong.
| ID | Role | Sector | Team | Yrs managing |
|---|---|---|---|---|
| M01 | Engineering Manager | Fintech | 8 | 6 |
| M02 | Senior Engineering Manager | Fintech | 12 | 9 |
| M03 | Engineering Manager | Fintech | 5 | 3 |
| M04 | Engineering Manager | SaaS / CRM | 8 | 7 |
| M05 | Engineering Manager | SaaS / CRM | 6 | 4 |
| M06 | Director of Engineering | SaaS / CRM | 16 | 12 |
| M07 | Engineering Manager | SaaS / CRM | 9 | 5 |
| M08 | Engineering Manager | Cloud infrastructure | 9 | 8 |
| M09 | Engineering Manager | Cloud infrastructure | 7 | 5 |
| M10 | Senior Engineering Manager | Cloud infrastructure | 14 | 11 |
| M11 | Engineering Manager | Health tech | 4 | 2 |
| M12 | Engineering Manager | Health tech | 12 | 9 |
| M13 | Senior Engineering Manager | Health tech | 10 | 7 |
| M14 | Engineering Manager | Developer tools | 6 | 4 |
| M15 | Director of Engineering | Developer tools | 18 | 14 |
| M16 | Engineering Manager | E-commerce | 11 | 6 |
| M17 | Engineering Manager | E-commerce | 7 | 3 |
| M18 | Senior Engineering Manager | E-commerce | 15 | 10 |
| M19 | Engineering Manager | Automotive software | 14 | 12 |
| M20 | Engineering Manager | Automotive software | 9 | 7 |
| M21 | Engineering Manager | Medtech software | 6 | 5 |
| M22 | Engineering Manager | Medtech software | 10 | 8 |
| M23 | Engineering Manager | Banking tech | 8 | 6 |
| M24 | Senior Engineering Manager | Banking tech | 13 | 11 |
| M25 | Engineering Manager | Banking tech | 5 | 2 |
| M26 | Engineering Manager | Logistics tech | 9 | 4 |
| M27 | Engineering Manager | Logistics tech | 12 | 8 |
| M28 | Engineering Manager | Gaming | 10 | 6 |
| M29 | Senior Engineering Manager | Gaming | 16 | 13 |
| M30 | Engineering Manager | Cybersecurity | 7 | 5 |
| M31 | Engineering Manager | Cybersecurity | 11 | 9 |
| M32 | Engineering Manager | Marketing tech | 14 | 10 |
| M33 | Engineering Manager | Marketing tech | 6 | 3 |
| M34 | Former EM (role abolished) | Micromobility | — | 7 |
| M35 | Engineering Manager | Micromobility | 8 | 5 |
“Everybody can churn out code now.” Output stopped being the signal.
When AI produces the baseline, raw volume stops telling you anything. 26 of 35 managers said their definition of performance changed materially within the last 18 months. What rose: architectural judgment, risk handling, and what one manager called the multiplier factor, the ability to make everyone around you better. What fell: lines shipped, tickets closed, velocity for its own sake.
“The base level of output is no longer sufficient. Everybody can churn out code.”
The metrics also miss the work that holds teams together. Several managers described engineers who look weak in raw numbers while quietly carrying coordination, mentoring and risk-handling for everyone else. And more than one pointed out that the real bottleneck has moved: AI shifted it from writing code to reviewing and integrating it, which is exactly the work the dashboards weigh least.
AI didn't shrink the work of evaluation. It moved it.
The efficiency story goes like this: AI gathers the evidence, the manager saves time. The managers describe something else. 24 of 35 report that evaluation takes more effort since AI entered the workflow, because the hours saved collecting signals now go into validating them. One manager built a fully automated pipeline that pulls delivery metrics, analyzes code and drafts performance summaries. He also described a standing Monday ritual of repairing his agents when they break. His verdict on his own system:
“Quite seductive, in the sense that they give you substance that is not there.”
“People trust dashboards too quickly because they look clean and rational. But performance is messy.”
Input, not outcome
31 of 35 managers drew the same line without being asked: AI can inform an evaluation, it cannot make one. They cross-check it, override it, and treat scientific-looking output as the start of a conversation rather than the end of one. Nearly all of them described some routine for verifying AI claims before acting on them. "Trash in, trash out," as one put it.
“It feels easy to imagine people being compared just based on this very scientific-looking output that may not be quite so accurate.”
The line gets sharper as decisions get more political. One manager split salary from promotion: salary is quantitative enough that AI could plausibly help, but promotions run on company interests that are written down nowhere a model can read them. The most consequential decisions are the least legible ones.
The hardest decisions never touched AI
Every manager could recall their hardest evaluation. An engineer unravelling because of troubles at home. A senior hire who could not land a role change. A strong performer who quietly stopped trying. Not one of them reached for an AI tool in those moments, including the heaviest adopters in the study. The decisive inputs were timing, trust and reading a person in the room.
Apparent underperformance turned out to be a crisis at home. The fix was a quiet pairing with a trusted colleague and time. No dashboard surfaced it, and none would have.
A respected senior engineer took on a team-lead role and sank. The work was helping them step back without losing face in front of the team.
High effort, falling fit. The manager said dragging out a clearly negative review would have been unfair to the person and the team both.
“It is not so much what you say, but how you say it. If you catch someone in a very low moment, you can completely destroy the relationship you have with that person.”
“I am more concerned about people shying away from responsibility than doing things wrong. We can always react.”
When it goes wrong, who answers? Three camps.
Then I asked the question this study was built around. When an AI-shaped evaluation harms someone (a missed promotion, an unfair rating, a termination), who answers for it? Thirty-five managers, three camps. And when asked who answers for a negative review in their org today, the most common reply was one word: "I am."
Accountability follows the final human signature. Whoever approves the review, or merges the code, answers for it, no matter how much of it AI produced.
“Once you push it to GitHub, it is you.”
Engineering manager, CRM/SaaS
Accountability moves up the chain to whoever mandated the system. If leadership forces tools on managers who cannot fully understand them, leadership owns the outcomes.
“The person who decides to use the AI is ultimately accountable.”
Engineering manager, industrial software
Accountability dissolves into infrastructure: CI pipelines, AI review agents, engineered safeguards. The question stops being who failed and becomes what the system allowed.
“It is not about the people, it is about the systems.”
Former engineering manager, AI-native company
One manager grounded his position in a live incident: an AI coding agent had recently wiped a production database. His read was that the failure belonged to the humans who designed insufficient safeguards around it, not to the model.
“Behind any computer mistake is a human mistake.”
Engineering manager, cloud infrastructure
Formal accountability is holding. The conditions for it are eroding.
Put the findings together and a structural picture emerges. Managers remain fully, formally accountable for evaluations. They insist on it. But the three conditions that make accountability meaningful (understanding the basis of a decision, controlling the process, being able to explain it) weaken as AI mediates more of the evidence. Responsibility stays put while control drains out of it. Researchers have a name for where this ends: the moral crumple zone, the point where the nearest human absorbs blame for outcomes an entire system produced.
“If I cannot explain the basis, I should not rely on it heavily.”
Every manager in this study is signing off on documents whose evidence base they can audit less and less. They know it. They sign anyway, because the alternative is to stop managing.
What the extreme cases look like
Two managers in the study sit at the far end of AI adoption, and what they describe sounds less like the future of performance reviews than their replacement. A third group is worried about a different destination entirely.
The company that deleted management
One participant, a former engineering manager, works at a company that abolished management outright. The CEO assesses performance by querying the codebase through Claude, and accountability lives in CI pipelines and AI review agents instead of people.
No managers. No one-on-ones. No review cycle. Performance lives in the codebase.
Founder mode, all the way down
Another manager hires by watching candidates direct AI agents, and expects individual contributors to operate like founders, each running a small fleet of them. In his world the weak signal is not low output. It is what he calls AI slop: volume without judgment. The strong signal is an engineer whose agents compound.
The digital prison
A third of managers raised surveillance without being asked. AI makes total monitoring of engineers trivially cheap: every commit, every message, every pause. The people best positioned to deploy that future want no part of it.
“I do not care if the employee is taking three coffee breaks in the morning or five in the afternoon, as long as the work is delivered.”
“AI-driven performance scoring is abdicating some of your managerial duties. Some of your human duties.”
A few things worth sitting with
None of this comes with instructions. But a few patterns from the conversations seem worth passing on, depending on where you sit.
- The managers who got the most out of AI were the ones free to override it. Tools that hand down verdicts were trusted least.
- When leadership mandates the tools, some of the accountability seems to move with that decision. A quarter of the managers already see it that way.
- The time saved gathering evidence tends to come back as validation work. Worth knowing before counting on efficiency gains.
- Output metrics say less than they used to. Most managers in the study had already shifted their weight toward judgment and the ability to lift others.
- Almost everyone drew the same line: AI informs the review, a person makes the call. The ones most at ease with their process were the ones who checked the AI's claims.
- The hardest conversations stayed fully human for every manager I spoke to. Whatever else changes, that part seems likely to stay.
Where I think this is heading
The study ends where the data ends, but the direction feels legible. Within a few years the unit being evaluated will not be the engineer. It will be the engineer plus their agents. Judging the human in isolation already feels artificial to the managers furthest ahead, and the rest of the industry tends to follow them with a lag.
The management role splits in two. One half becomes an editor: cross-examining AI evidence, auditing what the dashboards claim, deciding what deserves trust. The other half is the part that never changed, reading people, timing hard conversations, noticing when a number is hiding a person in trouble. Managers who are only output supervisors sit directly in the path of automation. Managers who own both halves become more valuable, not less.
Accountability will get formalized one way or another. Either organizations write deployment accountability into how they adopt these tools, with named owners, real override rights and time budgeted for validation, or the gap keeps widening quietly until a lawsuit, a works council or a regulator closes it for them. The crumple zone is comfortable for everyone except the person inside it.
Most companies will not delete their managers. But every company will eventually re-answer the question this study kept running into: when the evidence is machine-made and the consequences are human, who signs? The managers who come out ahead will be the ones who keep insisting that the answer is a name. Starting with their own.
What this is, and what it isn't
Findings draw on roughly 35 semi-structured interviews of 25 to 45 minutes with software engineering managers, conducted in 2025 and 2026 across 14 sectors and analyzed with reflexive thematic analysis. Participants are anonymized and quotes lightly edited for clarity. The percentages summarize qualitative coding of a purposive sample. Read them as directional findings from structured conversations, not survey-grade estimates. The underlying academic study was conducted at Stockholm Business School.
Want the full write-up, the data, or an argument?
aboodalzeno@gmail.com