Latest from the lab

Content integrity is a format problem · September 2026 update

In short

A task statement leaks because it is text: somebody copies it out of an assessment, posts it, and from then on the task measures who found the post rather than who can do the work. August’s answer to that is structural rather than another signal to watch, because six of the most popular fundamental tasks now arrive as a short narrated, animated video with no statement to lift. Alongside it, two integrity signals moved to where reviewers actually look, one switch now turns every AI feature off across an account, three technologies arrived that Codility could not assess before, and business candidates can open the data behind their task rather than asking for it. The full changelog sits at the end of this page.
  • Device integrity and pattern detection now move the overall Integrity Risk band on their own, rather than waiting for a reviewer to go looking for them.
  • Six of the most popular fundamental tasks now arrive as a short narrated, animated video instead. There is no statement to lift, and each one is a genuine variant of its written original, so the two are interchangeable.
  • Account settings now carry a single AI switch above every individual AI setting, for organizations that cannot permit AI tooling at all.
  • Playwright, Erlang and Clojure all arrived, each with platform support, ready-made tasks and self-serve authoring together, which is now the standard shape for a new technology.
  • The full changelog of everything that shipped in August is at the end of this page.

Where does a business candidate’s time go?

An AI for Business task is built on real data. A team roster, a set of productivity scores, a company brief. That data is what makes a scenario resemble the job rather than a hypothetical, and reading it is part of the work.

Until August it lived in the AI Assistant’s context. The task description named the files, and a candidate got at the contents by asking the assistant for them. That worked, and it spent some of the session on retrieval.

The files now sit under the task description and open in a tab beside it. Anything with rows and columns opens as a table with a fixed header, so a 120-row roster can be read rather than scrolled through as raw text. The assistant also moved to the middle of the workspace, so the layout runs task description, then assistant, then response, which is the order the work actually happens in.

A candidate reads the situation, forms a view, and uses the assistant to pressure-test it. More of the session goes on the judgment the scenario was built to draw out.

One detail that connects back to the rest of this page: the files are served through the same gated route as the task description and carry the same copy and print restrictions. Opening the data did not open a way to take it away.

Why is the task statement the weakest part of an assessment?

Most of the effort in assessment integrity goes on watching the candidate. Proctoring, behavioral signals, similarity checks, risk scoring. All of it useful, and all of it pointed at the same moment: the session where someone is answering the question.

The question itself gets far less attention, and it is the part that leaks. A task statement is text. Text can be copied out of a session, posted on a forum, indexed, and eventually pasted into a model. From that point the task is measuring who found the post rather than who can do the work, and no amount of watching the candidate detects that, because nothing about their session looks unusual.

The industry’s answer has been to rotate content faster and hope the refresh rate stays ahead of the leak rate. That is a treadmill. It buys time without changing the underlying property that makes a task leak, which is that the problem exists as text at all.

This is the same argument as how you prevent cheating in coding assessments, pointed at content rather than at candidates: the durable fix changes the structure of the assessment rather than adding another thing to look for.

What changes when the problem is narrated instead of written?

Six of the most popular fundamental tasks now come as a video. The candidate watches a short lecture where the problem is explained and animated on a board, then solves it as normal.

There is no statement to copy out. Passing the problem around means transcribing it, which is real work rather than a keyboard shortcut, and effort is the only honest lever anyone has here. This does not make a task impossible to leak. Someone determined will sit down and write it out. It moves the cost of doing so from nothing to something.

Two properties matter as much as the format. Each video task is a genuine variant of the written task it came from, so the two can be swapped for each other or mixed inside one randomized assessment. And they sit in the same libraries as their written counterparts, so there is nothing new to learn and nothing new to buy.

Does an integrity signal help if nobody looks at it?

Two of the more telling signals in the platform now sit in a better place.

Device integrity comes from the Codility Desktop App, and is designed to surface tools that sit invisibly over a candidate’s screen and feed them answers without appearing in a screen share or a recording. Pattern detection looks for code that was retyped from a second screen or another device. Both worked. Both required somebody to open a session and go looking.

Most reviewers read one summary reading and move on, which meant those signals existed mainly for the people who already knew to check for them. Either one now raises the overall Integrity Risk band by itself, so a session worth a closer look says so on the summary. Whole batches of candidates can also be filtered by either signal from Test Mission Control or the Candidates page, so triage happens across a pipeline rather than one report at a time.

Raising a band surfaces a session for a person to look at. It is designed to point a reviewer at the right sessions rather than to decide anything on its own, and the underlying evidence is there to read.

What if your organization cannot permit AI at all?

Plenty cannot, and the honest answer used to be a list of settings and a promise that nothing new had been added underneath them.

Account settings now carry a single AI switch that sits above every individual AI setting. While it is off, every AI feature is off across the account, whatever the individual settings say, and a feature shipped later cannot arrive switched on underneath it. While it is on, each feature behaves exactly as it did before, so per-assessment choices work the way a team already uses them.

You cannot put a list of settings in front of a compliance review. You can put one switch in front of it. The point of the control is that a cautious organization can adopt the platform without adopting the AI, and can show which of the two it has done.

What arrives when Codility adds a new technology?

The other half of August was the library getting broader, and the way it got broader matters more than the list.

Playwright, Erlang and Clojure all arrived, and each one arrived complete rather than partial. The platform runs it. A set of ready-made tasks covers the immediate hiring need. Your own engineers can build tasks in it, in the task builder or through the task creation MCP server. That combination is now the standard shape for every new technology rather than a one-off, which means asking for one has a known outcome instead of an open-ended queue.

The reason to care about the process rather than the three names: an assessment library that only grows when a vendor decides to grow it will always lag the stacks people are actually hiring for. Fifty tasks arrived across the three, spread through the Starter, Core and Advanced libraries rather than parked at the premium end.

What else changed in August?

Three more releases worth knowing about, with the full detail in the changelog below.

  • Business candidates can open the data behind their task. An AI for Business task is built on a roster, a set of scores, a company brief. Those files used to live only inside the assistant, so a candidate’s first job was prising the numbers out of it. They now open in a tab beside the brief, with spreadsheets laid out as a readable table.
  • Personalized feedback found the employee. It had been sitting two clicks deep inside a report tab, well enough hidden that almost nobody read it. It now announces itself on the employee’s own profile, marked visible only to them, with a count of what they have not read.
  • API keys can be rotated without downtime. An integration user can now hold several applications and several keys, so rotation runs as create, migrate, verify, revoke, with the integration up throughout.

What does this mean for how you run assessments?

The work of keeping an assessment trustworthy has been framed as a detection problem for a long time: watch harder, score the signals, flag the outliers. August moved a different lever. A narrated problem is harder to pass around than a written one because of what it is, not because of what anyone spotted. A signal that moves the summary reading gets acted on because of where it sits, not because a reviewer got more diligent.

Both are structural changes, and structural changes hold up when nobody is paying attention. That is the property worth optimizing for, because the alternative depends on sustained vigilance from people who have a pipeline to get through.

If you run Codility today, most of what is described here is already on. The AI features are the exception and start switched off, and an admin turns them on in account settings.

Everything that shipped in August

The complete record of the August 2026 releases, including the ones the article above does not cover.

Fundamental tasks delivered as video

Live

Six of the most popular fundamental tasks now have a video counterpart. Instead of reading the problem, the candidate watches a short lecture where it is explained and animated on a board.

Why It Matters

  • A problem that exists only as narration and animation is far more awkward to copy, scrape or paste into a model than a page of text. It raises the effort of passing a task around without changing what the task measures.
  • Each video task is a genuine variant of the task it came from, so the pair can be swapped or mixed inside a randomized assessment.
  • Video tasks sit in the same libraries as their written counterparts and are added to an assessment the same way, so there is nothing new to learn.

What’s New

  • Six fundamental tasks with a narrated, animated video statement
  • One in the Starter library, one in Core, four in Advanced
  • A seventh video task in the coding library, added earlier in August
  • Each one usable interchangeably with the written task it varies

Availability

Live now, available according to your packages. —

Device integrity and pattern detection feed the Integrity Risk band

Live

Integrity Risk is the single reading a reviewer looks at first, and two of the more telling signals were not part of it. Both now move the band, and both can be filtered across a batch of candidates.

Why It Matters

  • A reviewer scanning Integrity Risk now sees the effect of these two signals without opening anything else. Device integrity raises the band to at least high, and pattern detection raises it one level below that.
  • Candidates can be filtered by device integrity and by pattern detection on both Test Mission Control and the Candidates page, so a batch can be triaged rather than opened one at a time.
  • Raising a band surfaces a session worth a human look. It is designed to point a reviewer at the right sessions rather than to decide anything on its own, and the underlying evidence is there to read.

What’s New

  • Device integrity and pattern detection now factored into the Integrity Risk band
  • Candidate filtering by either signal on Test Mission Control and the Candidates page
  • Applies to new assessments created with these features switched on

Availability

Live for new assessments created with device integrity or pattern detection switched on. Device integrity runs through the Codility Desktop App and is in Open preview. —

Playwright as a task environment

Live

Playwright is Microsoft’s browser automation framework. It drives a real browser to test web applications end to end, and it is one of the most frequently requested technologies from quality assurance teams.

Why It Matters

  • Quality assurance hiring has been one of the largest gaps in the library, and Playwright is the framework those teams have moved to.
  • Tasks landed in Core as well as Advanced, so the technology is usable on the packages most teams already hold rather than only at the premium end.
  • Your own engineers can build Playwright tasks, so internal conventions and real repository patterns can go into the assessment.

What’s New

  • Playwright available as a regular task environment
  • 15 ready-made tasks across the Core and Advanced libraries
  • Custom Playwright task authoring in the task builder and through the task creation MCP server
  • Coverage from locators and form controls up to shadow DOM, network diagnostics and stateful checkout

Availability

Live now, available according to your packages. —

Erlang on the platform

Live

Erlang is a functional language built in the 1980s for systems that are not allowed to go down: telecom switches, messaging backbones, payment rails. Lightweight processes, message passing and let-it-crash supervision instead of shared-memory threads.

Why It Matters

  • Teams running Erlang in production had no way to assess for it, which meant hiring on a proxy skill or on a conversation.
  • Tasks cover the distinctive parts of the language rather than generic algorithm work: process-per-node designs, bit syntax, gen-server-style state, supervision under crashing jobs.
  • Starter library coverage is included, unusual for a new technology, so early-stage screening is possible from day one.

What’s New

  • Erlang supported on the platform
  • 20 tasks across the Starter and Core libraries
  • Erlang available for self-serve task authoring, including through the task creation MCP server
  • Coverage from list and map fundamentals up to worker pools, checksummed protocol decoding and nested validation

Availability

Live now, available according to your packages. —

Clojure on the platform

Live

Clojure runs on the same machinery as Java, so companies add it to systems they already have rather than starting over. It suits engineers who would rather describe what should happen to a set of data than write the steps for walking through it.

Why It Matters

  • Clojure sits inside existing Java estates, so the hiring need is often for one team inside a much larger organization, which is exactly the case a general-purpose library serves worst.
  • Tasks are built on the idioms rather than translated from another language: relational operators over sets, threading macros, multimethod dispatch, lazy infinite sequences.
  • Self-serve authoring means a team can encode its own conventions rather than accept a generic interpretation of the language.

What’s New

  • Clojure supported on the platform
  • 15 tasks across the Starter and Core libraries
  • Clojure available for self-serve task authoring
  • Coverage from data joins and username normalization up to multi-tier rate limiters and log aggregation pipelines

Availability

Live now, available according to your packages. —

Every new technology now arrives with three things

Live

Asking for a technology Codility did not support used to mean waiting without much sense of what waiting would produce. Starting with Playwright, each new language or framework arrives as a complete package.

Why It Matters

  • A request now has a known outcome rather than an open-ended queue.
  • The three parts arrive together, so there is no window where the platform runs a technology but nothing can be assessed in it, or where tasks exist but nobody can add their own.
  • Handing over self-serve authoring at the same time means the library stops depending solely on what Codility chooses to build next.

What’s New

  • Platform support for the technology
  • An initial set of ready-made tasks covering the immediate hiring need
  • Self-serve authoring in that technology, in the task builder and through the task creation MCP server
  • Playwright, Erlang and Clojure all followed this shape

Availability

Live. In effect for every new technology from the Playwright release onward. —

Quality assurance and test automation tasks

Live

Quality assurance hiring has been one of the most frequently requested gaps in the library. 30 new Selenium and REST Assured tasks close most of it. Playwright arrived in the same month and has its own block above, because it came with platform support and self-serve authoring rather than tasks alone.

Why It Matters

  • Quality assurance candidates were previously assessed on general coding ability, which measures something adjacent to the job rather than the job.
  • REST Assured gained the easy tier it never had, so the framework is now usable for early-stage screening rather than senior hiring only.
  • Advanced library placement gives teams already covered at Core a way to vary a randomized assessment or refresh content that has been in rotation too long.

What’s New

  • 13 Selenium tasks in Java, from data grids and native dialogs up to accessibility auditing, script injection and waiting that survives a flaky page
  • 17 REST Assured tasks in Java, from pagination and idempotency keys up to contract verification, regional failover and GraphQL query building
  • The easy tier REST Assured never had, so the framework is now usable for early-stage screening rather than senior hiring only
  • Both frameworks sit in the Advanced library, because Core is already covered for each

Availability

Live now, available according to your packages. —

Tasks for teams putting models into production

Live

Retrieval-augmented generation means fetching the most relevant documents from a knowledge base before a model answers, so the answer is grounded in your own data rather than the model’s training. It is now the most common way teams put a model into production, and it has a task family behind it.

Why It Matters

  • Retrieval work is the shape most production model deployments actually take, and it was not assessable as a distinct skill.
  • Building the metrics that judge a model’s output is a separate job from building the model. It is the role most teams are hiring for and least able to test.
  • One brief in six language versions means the same skill can be measured across stacks without maintaining six unrelated tasks.

What’s New

  • Conversational retrieval-augmented generation as one brief in six versions: Python, TypeScript, Go, LangChain, Java, and a version running real embeddings rather than mocked ones
  • A model evaluation task covering answer accuracy, calibration, steadiness across reworded questions and refusal detection, gathered into one scorecard
  • A support email router built on LangChain, a support ticket assignment task, a document clustering task and a linear regression task, all in the VS Code environment

Availability

Live now, available according to your packages. —

More tasks where the candidate picks the language

Live

A task that carries its own language forces a choice between the problem you want to set and the stack you want to test. Three new tasks let the candidate answer in whichever language they are strongest in.

Why It Matters

  • Hiring for engineering judgment rather than for a specific stack becomes possible without maintaining a separate task per language.
  • The same skills are measured whichever language a candidate chooses, so results stay comparable across a mixed pipeline.
  • One of the three is the VS Code version of the most popular REST endpoint task, so a widely used assessment is now stack-agnostic.

What’s New

  • Rebuilding resource state from a webhook stream full of duplicates, disorder and deletes, in one of seven languages
  • Counting phone numbers hidden in free text, in one of nine languages
  • Implementing a REST endpoint with an exact-match filter, in one of nine languages: Python, Go, Java, JavaScript, TypeScript, Ruby, Rust, C++ or C#
  • Candidates may solve in more than one language, and the strongest attempt counts

Availability

Live now, available according to your packages. —

AI for Business candidates can open the task data

Live

An AI for Business task is built on data: a team roster, a set of productivity scores, a company brief. Those files used to live only in the assistant’s context, so the only way to see them was to ask the assistant to read them out.

Why It Matters

  • The candidate reads the data first and uses the assistant to interpret it, rather than spending their first minutes extracting it. Less time guessing what to ask, more time on the judgment being measured.
  • Files with rows and columns open as a table with a fixed header row, so a long roster can be read rather than scrolled through as raw text.
  • Files are served through the same gated route as the task description and carry the same copy and print restrictions, so opening the data does not open a way to take it away.

What’s New

  • Attached files listed under the task description, opening in a tab beside the brief
  • Spreadsheet files rendered as a table with a fixed header row
  • Formatted view for text and document files
  • The assistant repositioned to the middle of the workspace, between the task description and the response
  • Copy and print protection carried across to file previews

Availability

Live across Screen and Skills Intelligence. —

Personalized feedback on the employee profile

Live

A skills program pays off when people act on what they learn, and nobody acts on feedback they cannot find. Personalized feedback on an assessment has existed for a while, two clicks deep inside a report tab. It now comes to the employee.

Why It Matters

  • A personalized feedback banner appears on the profile the moment feedback is waiting, carrying a lock marked visible only to you, so an employee knows their manager cannot read it.
  • Review and Read again buttons on every task open the report at that task’s feedback rather than at the top of it.
  • Past assessments are included, so people find feedback written for them months ago and never seen.
  • Manager views are unchanged and there is nothing to switch on.

What’s New

  • Personalized feedback banner on the employee profile, with a visible-only-to-you lock
  • The Timeline tab is now Assessment feedback, carrying a count of unread items
  • Review and Read again buttons on every task, opening the report at that task’s feedback
  • Past assessments included

Availability

Live for accounts with AI candidate feedback switched on. Nothing new to enable. —

Special arrangements recipients on a program

Live

When a candidate or an employee needs an adjustment to how they sit an assessment, that request has to reach someone who can act on it. The recipient was fixed, so on a large program the request often landed with the wrong owner.

Why It Matters

  • The person who owns a program nominates who receives special arrangements requests for it, rather than accepting a single account-wide destination.
  • The field suggests eligible recipients from the admins and managers on the account, so a request cannot be routed to an address that cannot act on it.
  • The setting can be changed after a program is running, so routing follows a change of owner rather than outliving it.

What’s New

  • Special arrangements recipients field when creating a Skills Intelligence program
  • The same field on the program details view
  • Recipient suggestions drawn from eligible admins and managers

Availability

Live. —

Additional services connect themselves in VS Code

Open preview

A task that needs a database used to need someone to wire the database up. Adding MongoDB, MySQL, Postgres or Redis to a VS Code space left the connection to be made by hand.

Why It Matters

  • Adding one of the four services now connects it to the VS Code workspace automatically, so the environment is ready when it opens.
  • Nobody loses the first stretch of a session working out a connection string, which is setup cost for whoever builds the task and wasted minutes for the candidate.
  • Environments where a database serves no purpose can no longer have one attached, so the option list matches what the environment can actually do.

What’s New

  • Automatic connection between an added service and the VS Code workspace
  • Supported across most VS Code environments, including Bash, C, C++, Dart, Django, .NET, Go, Java, Jupyter Notebook, Kotlin, Next.js, Node.js, Python, R, Ruby, Rust, Spring Boot and Swift
  • Service attachment removed from environments where a database does not apply, namely React, Angular, Terraform and SystemVerilog

Availability

In Open preview, so it is on for every account. —

Manage your API keys in a secure and user friendly manner

Live

An integration user held one application and one API key. Rotating that key meant resetting it, which replaced credentials still in use, so the safe window between the new key working and the old key dying did not exist.

Why It Matters

  • Rotation becomes four steps you control: create the new key, move the integration across, confirm it works, then revoke the old one. The old key stays valid until you choose to revoke it, so nothing is forced to break while you switch.
  • Applications can be renamed and redirect URIs set per application, so a key is recognisable by what it does rather than by when it was made.
  • Individual keys can be revoked as they fall out of use, without touching the rest.
  • A record of what was retired: revoked keys and applications stay listed behind a show and hide toggle, and can be deleted once nobody needs them, so an integration user carries its own history.
  • Eightfold follows the same path, so an Eightfold credential can be moved without a reset step in the middle.

What’s New

  • Multiple applications and API keys per integration user
  • Add, rename and revoke individual applications and keys
  • Redirect URIs set independently per application
  • The existing reset flow kept for applications owned by regular platform users
  • Account admins hold the permission by default and can grant it to other roles. Anyone with integration access can view applications and their values without changing them

Availability

Live for all customers. —

One switch for every AI feature

Live

Some organizations cannot permit AI tooling in an assessment at all. Until now that meant finding several separate settings, turning each one off, and trusting nothing new had been added underneath since somebody last checked. Account settings now carry a single AI switch above all of them.

Why It Matters

  • One place to say no: turn the switch off and every AI feature is off across the account, whatever the individual settings say. A feature added later cannot arrive switched on underneath it.
  • The individual settings are untouched: turn the switch on and each feature behaves exactly as it did before, so per-assessment choices work the way your team already uses them.
  • Built for the review that asks for it: a single account-level control that can be shown to a legal or compliance reviewer is usually what that conversation is actually after.

What’s New

  • A global AI switch in account settings
  • Every AI feature disabled across the account while the switch is off
  • Each feature returned to its own setting when the switch is on
  • Applied to existing accounts, so there is nothing to set up first

Availability

Live in account settings. Changing it is an admin action. —

A new typeface, and a better-oriented AI Copilot

Live

Two platform-wide changes. Inter replaces Roboto across the product, and the AI Copilot carries more context about the workspace it is running in.

Why It Matters

  • Inter was designed for screens, with readability and accessibility in mind, and it brings the product into line with the rest of Codility.
  • The change is subtle and it touches almost every screen, which is the answer if the product looks slightly different than you remember.
  • The AI Copilot staying closer to the task in front of the candidate means fewer answers that are technically correct and beside the point.

What’s New

  • Inter replaces Roboto across the Codility platform
  • The AI Copilot carries more context of the workspace it is connected to

Availability

Live across the Codility platform. — —

Frequently asked questions

Why does Codility deliver some task statements as video?

Because a written statement can be copied out of an assessment and shared, and a narrated, animated problem is far more awkward to move around. Six of the most popular fundamental tasks now have a video counterpart. It raises the effort of passing a task around without changing what the task measures. It does not make a task impossible to leak.