The ROI of engineering assessment

THE SHORT ANSWER

The ROI of engineering assessment comes from one equation: the quality of the people you hire multiplied by how much better your best engineers are than your worst ones. For most of the last two decades, that second number moved slowly, and the business case for assessment leaned on cost-of-bad-hire statistics with shaky provenance. That era has ended. In an AI-native environment, the gap between your strongest and weakest engineers is widening, which means every point of selection accuracy is worth two to five times what it was in 2020. The cost of getting hiring wrong has not stayed flat. It has quietly doubled.

This is a long read, and deliberately so. Most articles on this topic open with the same five statistics, all of which fall apart under scrutiny. We want to show our working on where those numbers come from, what the peer-reviewed science actually says, and why the calculation you may have been using to justify assessment investment is probably understating the return by half. We will get to the maths. First, the provenance problem.

SACKETT ET AL. 2022

0.42

VALIDITY OF STRUCTURED INTERVIEWS (VS 0.19 FOR UNSTRUCTURED)

BEHROOZI ET AL. 2020

61%

WHITEBOARD FAILURE RATE UNDER OBSERVATION (VS 36% IN PRIVATE)

DANIOTTI ET AL. SCIENCE 2026

30M

GITHUB COMMITS SHOWING AI WIDENS NOT CLOSES THE SKILL GAP

CODILITY CANDIDATE SURVEY

1.25M

CANDIDATE FEEDBACK RESPONSES SHAPING CODILITY’S ASSESSMENT DESIGN

SECTION 01 · PROVENANCE

ANSWER

There is no trustworthy engineering-specific figure. Every widely-cited cost-of-bad-hire number is a general hiring stat extrapolated to engineering. Most trace back to phantom government citations, self-reported vendor surveys from 2007, or a training-conduct study from 2005. The one source that holds up methodologically (CIPD’s 2024 Resourcing and Talent Planning Survey) reports 41% of UK organisations had new hires resign within the first 12 weeks, though again this is cross-role. General figures still point in the same direction: poor hires cost between 50% and 200% of annual salary. For engineering specifically those costs are almost certainly understated, given higher salaries, longer ramp, and the technical debt trail a bad engineering hire leaves behind.

Why the headline numbers do not stand up

None of the widely-cited cost-of-bad-hire figures are engineering-specific. They are generalising statistics, usually drawn from cross-role HR surveys, then applied to engineering because nobody has produced a defensible engineering-specific alternative. That observation is itself worth pausing on. The category has spent two decades quoting general-hiring averages as if they were engineering numbers, and because senior engineers cost two to three times the average hire in those surveys, ramp longer, and leave deeper technical debt when they do not work out, the general figures almost certainly understate the real engineering cost. With that caveat in mind, here is what happens when you chase the most-repeated numbers to their source.

Phantom Citation

“US DOL: bad hires cost 30% of earnings”

Appears in thousands of recruitment articles. No findable DOL publication. Searching returns nothing.

Actual Origin

Cathy Fyock’s 1993 practitioner book, presented as a rule of thumb. HR vendors began attributing it to the DOL in the mid-2000s. Textbook citogenesis.

Frozen in 2017

“CareerBuilder: $14,900 per bad hire”

Self-estimated costs from a 2017 Harris Poll of 2,257 hiring managers. Never updated.

Status in 2024

CareerBuilder merged with Monster in 2024. Still cited as current. Nine years of cost inflation ignored.

Vendor-Origin

“46% of new hires fail within 18 months”

2005 self-published Leadership IQ study. Never peer-reviewed, never replicated. Sold with the “Hiring for Attitude” product.

Why it is inflated

“Failure” includes termination, voluntary departure under pressure, performance reviews below expectations, and disciplinary action. Twenty years on, nobody has re-run it.

Derivative

“REC: £132,000 per bad mid-manager hire”

2017 REC “Perfect Match” report, modelled calculation co-produced with Indeed.

The core input

Leadership IQ’s US failure rate. The most-cited British number is partially derivative of the most-cited American number.

Anecdote, Not Data

“Zappos: bad hires cost us $100 million”

Tony Hsieh remark in an interview, repeated across a decade of thought leadership.

Methodology

None. No audit, no disaggregation, no breakdown. A founder’s off-the-cuff estimate treated as data ever since.

Stands Up

CIPD: 41% of UK hires resign within 12 weeks

Annual, independent, cross-sector survey of ~1,000 UK HR professionals. Measured consistently each year.

Why this one works

A rare leading indicator, not a modelled cost estimate. Not engineering-specific, but methodologically sound.

What actually stands up

CIPD’s Resourcing and Talent Planning Survey is annual, independent, and draws on roughly a thousand UK HR professionals across all sectors. It does not publish a headline bad-hire figure. What it does publish is a measurable leading indicator: 41% of UK organisations reported new hires resigning within the first 12 weeks in 2024. That is a real number, measured the same way every year, and it tells you the real-world rate of hiring mismatch is high enough to justify serious investment in selection quality, even without a clean engineering-specific figure.

SHRM’s Human Capital Benchmarking Report is a self-reported member survey. Its cost-per-hire figure (around $4,700) and time-to-fill figure (around 42 days) are useful directional benchmarks. The “6 to 9 months of salary” replacement cost you see in SHRM toolkits is practitioner consensus rather than a single landmark study, which is worth knowing when you quote it.

The provenance point matters because it reveals how much of the category’s thought leadership is a house of cards. When an assessment vendor sells you on ROI using any of these statistics, the first question is when the underlying research was done, by whom, and whether it has been re-measured since. The usual answer is: a long time ago, by someone with an interest in the finding, and no.

The deeper point is that bad-hire cost is the wrong anchor for engineering ROI anyway. The utility formula in the next section is the right one, and the evidence it rests on, particularly in an AI-augmented environment, is increasingly engineering-specific. From here on, the research we cite is either directly about software engineers (Behroozi on coding interviews, Daniotti on GitHub commits, METR on open-source developers, DORA on engineering teams, Anthropic on junior engineers) or in meta-analysis validity science that applies across roles including engineering.

SECTION 02 · THE SCIENCE

ANSWER

Yes, when they are designed correctly. Structured assessments that mirror real engineering work (work samples) predict job performance roughly twice as well as unstructured interviews, and many times better than years of experience. The Sackett et al. 2022 meta-analysis, published in the Journal of Applied Psychology, revisited the consensus validity coefficients downward after finding flaws in earlier methodology, but the relative ordering held. Structured methods work. Unstructured ones barely beat chance.

What the research actually shows

For two decades, the academic anchor for “selection quality matters” was Schmidt and Hunter’s 1998 meta-analysis in Psychological Bulletin, synthesizing 85 years of research. It is still the most widely cited paper in personnel selection.

In 2022, Sackett, Zhang, Berry and Lievens– identified systematic flaws in how prior meta-analyses had corrected for range restriction, and re-evidenced the validity coefficients. Most numbers came down. Most recruitment content has not yet caught up.

The revised coefficients, which are the defensible ones to use in 2026.

Validity of selection methods in predicting job performance

Correlation (r) with job performance. Sackett, Zhang, Berry & Lievens (2022). Journal of Applied Psychology. Range 0 to 1.


Structured interview

0.42

Job knowledge tests

0.40

Work sample tests

0.33

General mental ability

0.31

Integrity tests

0.31

Unstructured interview

0.19

Years of experience

0.07


High validity

Moderate

Barely better than chance

0.07

Years of experience predicts job performance at r = 0.07. Explains roughly half a percent of performance variance.

Structured interviews are more than twice as valid as unstructured. The most common hiring method is barely better than chance.

Two numbers deserve to be printed on the wall of every engineering leader’s office. Years of experience predicts job performance at r = 0.07. Unstructured interviews predict it at r = 0.19. Less than half the validity of structured interviews. The gap between scientific best practice and common hiring practice in engineering is significant.

Why work samples have a special claim

The theoretical case for coding assessments comes from Asher and Sciarrino’s 1974 principle of “point-to-point correspondency”: when the behaviour you measure directly mirrors the behaviour you are predicting, validity is improved. A candidate writing code in a realistic IDE is a closer proxy for the job than the same candidate whiteboarding a data-structures puzzle, or answering behavioural questions about past projects.

Work samples can show smaller demographic subgroup differences than cognitive tests alone (Roth, Bobko and McFarland, 2005)
, improving both validity and fairness at the same time. That is rare. Most selection methods trade one for the other.

When the method contaminates the signal: the whiteboard interview

Validity coefficients only hold if the method delivers clean signal. The cleanest illustration of what happens when a method contaminates its own measurement comes from Behroozi, Shirolkar, Barik and Parnin’s controlled experiment at NC State and Microsoft, published at ESEC/FSE 2020.

The researchers randomly assigned 48 computer science students to two conditions. Half completed a whiteboard coding task with an observer present, thinking aloud. Half completed the same task done in private. Same problem. Same difficulty. Same time limit.

Coding interview performance: public vs private conditions

48 computer science students. Randomised assignment. Published at ESEC/FSE 2020.

Condition A

Public whiteboard, observer present, think aloud

61.5%

Failed the task


Measured physiological response

Pupil dilation, fixation duration, self-reported frustration all rose significantly. Median correctness dropped from 3 of 3 test cases to 1 of 3. Zero women in the sample passed.

Condition B

Private, no observer, no think aloud

36.3%

Failed the task


Same content, different delivery

Same task. Same difficulty. Same time limit. When the observer condition was lifted, the candidates measured correctly at the median, not the candidates.

The authors noted that the whiteboard interview “has an uncanny resemblance to the Trier Social Stress Test”, a laboratory procedure designed specifically to induce cortisol spikes. Zero women passed the public whiteboard. All women passed the private one.

The implication for the ROI calculation is direct. Run the utility formula on a whiteboard-style assessment and your Δr collapses. Same content, delivered under conditions that match real engineering work, would give you meaningfully more predictive signal. You are paying for measurement and receiving none. You are also driving away the candidates who would have passed in calmer conditions, with a disproportionate drop-off among women. Ecological validity is not a design preference; it is the difference between an assessment that predicts job performance and an assessment that measures how candidates respond to artificial stress.

SECTION 03 · THE MATHS

ANSWER

The right formula is not “cost of a bad hire divided by number of hires.” It is the Brogden-Cronbach-Gleser utility equation, developed in the 1950s and formalised by Schmidt, Hunter, McKenzie and Muldrow in 1979. In plain English: the dollar return on a hiring method equals the number of people hired, times how long they stay, times how much better your method is at predicting performance, times how much top performers out-earn bottom performers for the business.

The formula, in the form engineering leaders should care about

The Utility Formula

Brogden-Cronbach-Gleser


Return × N × T × Δr × SDy × Z

N

Engineers hired per year

T

Average tenure in years

Δr

Gain in selection validity

SDy

Productivity dispersion in £

Z

Function of selection ratio

The variable that matters most, and that most vendors never discuss, is SDy. It is how much dispersion exists between your best and worst engineers when expressed in business value. Conservative industrial-organisational psychology estimates place it at roughly 40% of mean salary. In knowledge work, most estimates now put it higher.

A worked example

Take a company hiring 100 engineers per year at £140,000 average total compensation, with average tenure of four years, a selection ratio of 0.25 (one in four applicants hired), and SDy estimated conservatively at 40% of salary (£56,000).

Move from the unstructured interview (r = 0.19) to a structured coding work sample (r = 0.33). The validity improvement is 0.14.

Worked example

100

Engineers hired per year

4

Average tenure in years

0.14

Gain in selection validity

£56k

Productivity dispersion in £

1.27

Function of selection ratio

Total return over 4 years≈ £4.0m

That is roughly £40,000 per hire per year, from the single act of switching one selection method for another. Push Δr to 0.20 (unstructured interview to a composite of structured interview plus GMA plus work sample) and the return roughly doubles.

This is the formula that actually answers “what is the ROI of engineering assessment?” It is absent from nearly every piece of marketing in the category, because it does not require you to believe an unverified bad-hire statistic. It requires you to believe that engineers vary in productivity, which every engineering leader knows to be true.

And as of 2024, that SDy variable is doing something it has not done for decades. It is moving.

Section 04 · The AI shift

Answer

The gap between high-performing and low-performing engineers is widening. Research analysing more than 30 million GitHub commits across around 160,000 developers (Daniotti et al., Science 2026) found AI tools boost productivity for senior developers but produce no measurable benefit for early-career ones. Rather than closing skill gaps, AI is widening them. When productivity dispersion widens, the SDy variable in the utility formula moves with it. Every unit of selection accuracy becomes worth two to five times what it was in 2020. This is not a marginal shift in the business case for assessment. It is a structural re-basing.

The productivity dispersion is widening

The most important piece of research for engineering leaders to know is Daniotti et al., published in Science in 2026. The researchers analysed more than 30 million GitHub commits from around 160,000 developers across six countries. Their finding: “Generative AI boosts productivity and accelerates expansion into new technical domains, but only for senior-level developers. Early-career developers, by contrast, despite being the most frequent users of genAI, show no measurable benefits. Rather than closing skill gaps, genAI appears to be widening them.”

Three further studies corroborate the pattern.

Widening

Anthropic · Shen & Tamkin, Jan 2026

AI group scored 17 points lower on comprehension

52 mostly junior engineers taught a new async Python library. Those who asked conceptual questions retained 65%+. Those who delegated code generation retained under 40%.

Widening

DORA 2024 & 2025 · Google Cloud

AI is an amplifier, not a rising tide

Nearly 40,000 respondents. +3.4% code quality, -7.2% delivery stability. Strong teams get stronger. Weak teams ship chaos faster.

Widening

Labour market

Stanford payroll data

Employment for developers aged 22-25 is down 20% since late 2022

Developers 26+ have held steady or grown. The generational productivity gap is already visible in hiring decisions being made today.

Widening

Primary evidence

Daniotti et al. · Science 2026

“AI is widening skill gaps, not closing them”

More than 30 million GitHub commits. Around 160,000 developers. Six countries. Senior developers gain. Early-career developers show no measurable benefit.

The compression evidence, and why it does not hold where the money is

There is a competing narrative, and it is worth addressing directly. The early story on AI coding tools was that they would compress the gap by making everyone more productive. Some evidence supports that on specific task types. GitHub’s Peng et al. randomised trial reported Copilot users finishing a JavaScript HTTP server task 55.8% faster than controls. Cui et al. in Management Science (2025) pooled three field experiments covering nearly 4,900 developers at Microsoft, Accenture and a Fortune 100 electronics firm, and found a 26% increase in completed tasks.

The qualifier matters. These studies measure AI’s impact on contained, well-scoped tasks that resemble training data. Real engineering work looks different. METR’s July 2025 randomised study put 16 experienced open-source developers on 246 real tasks inside mature repositories averaging 22,000 GitHub stars. Allowing AI tools slowed them down by 19%. The developers forecast beforehand that AI would speed them up by 24%. After the study, they still believed AI had sped them up by 20%. A 39-percentage-point gap between how fast they felt they were working and how fast they actually were.

Dell’Acqua et al.’s Harvard and BCG field experiment with 758 consultants found that AI raised performance by 25 to 40% on tasks inside its capability envelope, but made workers 19 percentage points more likely to produce the wrong answer on tasks outside it. They called this the “jagged technological frontier”. Workers cannot tell, in the moment, which side of the frontier they are on.

AI compresses dispersion

On well-scoped greenfield work

Contained tasks that resemble AI training data. Novices gain more than experts. This is where most compression studies measure.

AI compresses dispersion

On real, contextual engineering work

High quality bars, implicit conventions, mature codebases, production operations. Almost all economically valuable software lives here.

These findings reconcile once you separate the task types. AI compresses productivity dispersion on well-scoped greenfield work that resembles training data. It widens dispersion on real, contextual engineering work with high quality bars and implicit conventions. Almost all economically valuable software lives in the second category.

What developers themselves say

Stack Overflow’s 2025 Develop er Survey captured the shift in sentiment in a single year. 84% of developers use or plan to use AI tools. Favourable sentiment dropped from 72% to 60%. Trust dropped from 40% to 29%. Only 3% “highly trust” AI output. 66% reported struggling with “almost right but not quite” AI solutions. 45% said debugging AI-generated code takes longer than writing it themselves. The people closest to the tools are the least confident that they produce net-positive work at the bottom of the distribution.

Charity Majors

Co-founder and CTO, Honeycomb

Section 05 · The new question

Answer

The old question was “what is the cost of a bad engineering hire?” The new question is “what is the cost of hiring an engineer who is a net-negative with AI tools, and you could not tell from their CV or their LeetCode score?” This is a different failure mode. It does not look like a bad hire in the first month. It looks like a bad hire twelve months in, when the technical debt compounds and nobody on the team can reason from first principles about what they have inherited.

The AI-fluency false positive

The new failure mode runs like this. A candidate uses AI tools confidently in an interview, ships fast-looking code, and speaks the right vocabulary. But the AI output contains the “close-but-wrong” category of defect that two-thirds of Stack Overflow respondents report struggling with. The candidate’s own debugging capability has been eroded by over-delegation, the pattern Anthropic’s skill-formation study surfaced. Month three, they are shipping. Month twelve, the code base has accumulated technical debt faster than the team can service it, and when the AI tool gets something subtly wrong, nobody catches it.

This failure mode is not visible in traditional assessments. It does not show up in a LeetCode score. It does not show up in a CV. It only becomes visible once the engineer is inside a real system, at which point the cost of removing them runs into the hundreds of thousands and the cost to the team’s morale and the product’s integrity runs higher still.

What thoughtful engineering leaders are now assessing for

The framework emerging from practitioners who have thought about this seriously looks like this.

Steve Yegge and Gene Kim, in Vibe Coding (IT Revolution, December 2025), argue for a two-part interview structure. One air-gapped round with no AI assistance, to verify the candidate can reason about code unaided. One AI-enabled round to test the actual job: are they thoughtful about framing prompts, adept at managing context, savvy at debugging model misunderstandings, or are they flailing? They write that by late 2025, “it will be a huge red flag” if a candidate cannot code at all without AI.

Kent Beck’s distinction is between “augmented coding” (candidate maintains engineering values like tests, complexity management and coverage while working with AI) and “vibe coding” (accepting whatever the model produces). Both might ship working code in an interview. Only one is hireable.

The practical implications for what engineering leaders should now assess:

1

Ask who designs the assessments and what validation methodology they use. Look for IO psychology credentials and published research.


2

Open the candidate-facing coding environment and try it yourself. Write code, use the terminal, debug something. If it feels like a downgrade from your daily IDE, it will feel that way to candidates too.


3

Ask for the platform’s specific AI position: does it block, embed, or make AI configurable? Ask how it captures and reports AI usage.


4

Verify the platform’s compliance posture (SOC 2 Type II audit, GDPR) and ask specifically about EU AI Act preparation if you operate in the EU.


5

Request actual completion rate data, not marketing estimates. Ask for the methodology behind any published statistics.


6

Test the ATS integration with your specific system. Send a test candidate through the full workflow and check what data flows where.

Frequently asked questions

What is the ROI of engineering assessment?

The ROI of engineering assessment is governed by the Brogden-Cronbach-Gleser utility formula: ΔU = N × T × Δr × SDy × Z, where Δr is the improvement in selection validity and SDy is the standard deviation of engineer productivity in pounds or dollars. For a company hiring 100 engineers at £140,000 average compensation over four years, moving from an unstructured interview (r = 0.19) to a structured coding work sample (r = 0.33) generates roughly £4 million in utility. In an AI-augmented environment, that return is now two to five times what it was in 2020, because productivity dispersion between strong and weak engineers has widened.