Codility AI Capabilities: assess real engineering work in an AI-first world

AI can write code now.
Hire the engineers who can build_

AI has changed what engineering excellence looks like. The engineers who thrived two years ago may not be the engineers who thrive now. You need signal on who can build in this environment, backed by methodology that holds up when questioned.

One methodology across Screen, Interview, and Skills Intelligence

The problem is what candidates do with the output

Every engineer you assess already works with an AI assistant daily. Banning it in an assessment tests a workflow nobody uses on the job. Allowing it without visibility collapses your signal: correctness scores cluster near the top, and you cannot tell who understood what they shipped from who generated it. What a candidate does with the output is the question that matters.

Correctness alone no longer separates candidates

When an AI assistant can produce a working solution for almost anyone, a pass or fail score stops telling you who can build. You need to see the judgment that produced the output, alongside the output itself.

Live interviews still run in an artificial environment

Your engineers work in VS Code with real tooling and an AI copilot beside them every day. A stripped-down browser editor with no AI visibility asks candidates to perform in a setting they never actually work in.

You cannot audit a decision you cannot see

When a hiring outcome gets challenged, a raw score leaves you with nothing to show. You need a record of how the work was produced, reviewed by a human, that holds up when someone asks how the call was made.

From blocking AI to measuring judgment

Codility makes AI collaboration visible, then scores the part of the work that still separates engineers: judgment about what to build, what to trust, and what to fix.

“Everyone’s using AI and it feels sometimes unfair to disregard candidates the access. But then we can monitor it and ensure they’re not using it for the whole thing.”
Engineering leader, consumer analytics
“Within 12 months I think all of the vendors in the market are going to have very similar AI copilot integration. I think the real differentiation is: what are the insights?”
Engineering leaders, European fintech

AI-native realism, built into the work itself

AI-Native Tasks put engineers in a real VS Code environment where an AI assistant is part of the job itself. Two styles cover the two things you actually need to know: can this engineer build with an assistant, and can they catch it when the assistant is wrong.

Build AI

The candidate builds the AI or ML system itself, a retrieval pipeline or an anomaly detector, and is scored on how they design and harden it. Working with an assistant is a requirement here, since the task is often only achievable in the time given when the candidate puts one to real work.

Work with AI

The candidate works alongside an assistant seeded to sometimes be wrong, so they have to catch what it gets wrong, check the rest, and ship a result they would stand behind. Judgment about the assistant’s output is the differentiating skill.

The concrete mechanism, an assistant that is sometimes wrong on purpose, is the proof point. It ships as a fixture of the assessment today.

AI-Native Tasks are live now across Interview, Screen, and Skills Intelligence. Full task-by-task detail is in the AI-Native Tasks one-pager.

Claude Code in the assessment environment

Preview. Open to every customer across Screen, Interview and Skills Intelligence

Anthropic’s Claude Code runs inside the Codility assessment environment, preinstalled and routed through Codility’s own AI infrastructure. A candidate opens the terminal, types claude, and works the problem alongside an agent. Every prompt, edit and correction saves to the candidate report and is covered by Replay, so how an engineer directs an agent becomes reviewable evidence rather than something a reviewer has to reconstruct afterwards.

The same agent in all three

Screen, Interview and Skills Intelligence run the same agent in the same real environment. In Interview the session is shared live, so interviewer and candidate see every prompt and response as it happens, and both sides can interact.

Nothing to set up

Preinstalled and preconfigured, with the trust and onboarding dialogs pre-accepted. No API keys, no logins, and no candidate credentials in play, so a timed session starts on the problem.

Control at two levels

Enabled or disabled per assessment in the same AI Assistant settings. Above that sits one account-wide AI switch, so an organization that cannot permit AI tooling has a single place to say no.

A live Codility Interview: a shared VS Code session with main.py open and a two-person video call between the interviewer and the candidate. In the terminal, Claude Code is running, hosted by Codility, and the interviewer has asked it to create a small Python interview challenge for a candidate.
Claude Code running in the terminal of a shared Codility Interview session, hosted by Codility.
The prompting is the signal. Claude Code helps make it observable: where a candidate trusts the agent, where they push back, and how they turn a rough first answer into code they would stand behind.

Full detail on setup, controls, and what reaches the candidate report is in the Claude Code one-pager.

Controlled AI, reviewable by design

The AI Assistant and Claude Code are enabled or disabled per assessment, with every candidate interaction captured as reviewable AI activity. Reviewers see the prompts, the iterations, and the finished output. Model-level and more granular controls are on the roadmap.

Visible by default

Every prompt and response the candidate exchanges with the assistant is logged. Reviewers can see how a candidate directed the assistant and where they pushed back on it, alongside the code itself.

No automated grading of AI activity

Reviewable AI activity informs human judgment and leaves the decision with the reviewer. The person making the hiring call still makes the call, with more of the picture in front of them.

One place to say no

Account settings carry a single AI switch above every individual AI setting. While it is off, every AI feature is off across the account, whatever the individual settings say, and a feature released later cannot arrive switched on underneath it. Changing it is an admin action.

One methodology, across hiring and the workforce you already have

The same assessment science runs through every product. Hire the right engineers, then verify what your existing team can actually do, on one evaluation model instead of three disconnected tools.

Screen

Standardized early technical signal at scale. Assess AI-enabled work with auditable outcomes before a candidate reaches a live interview.

Interview

Structured technical interviews in a shared VS Code environment with sidecar services, whiteboard, and a full transcript of how the candidate actually worked.

Skills Intelligence

Map technical capability across the engineering org you already have, on the same methodology used to hire them.

The same methodology now extends beyond engineering too. AI for Business is in Preview across Screen and Skills Intelligence, assessing AI collaboration in non-technical roles on the same evidence-first foundation.

AI as a first-class engineering skill

The Engineering Skills Model gives AI collaboration its own category, defined and scored on the same footing as the technical disciplines beside it. Every AI-native task is tagged at the test-case level, so scoring reports against the specific skills it exercises.

Skill model 30

skills across five categories, with AI as a first-class category alongside the technical disciplines engineering teams already assess.

Skill model 180

subskills, giving reviewers granular signal on where a candidate or employee is strong and where they are not.

Validation 210

elements mapped to SWEBOK, SFIA, O*NET, the NIST AI RMF, the EU AI Act, ISO/IEC 42001, and Anthropic’s AI Fluency framework, validated by engineering leaders.

Methodology that holds up when questioned

Codility answers “why should I trust this score” with documented methodology built by occupational psychologists, tested against how engineers actually work worldwide, and mapped to the standards regulators and engineering bodies already recognize.

103-country practitioner surveyIEEE SWECOM and SWEBOK alignmentOccupational psychologists on the methodology teamAdverse impact monitoring at every cut scoreValidated by engineering leaders

Codility’s AI capabilities span Screen, Interview, and Skills Intelligence, all on this one methodology.