Research Assessment Research

Inside the Engineering Skills Model: what it measures, how it works, and what changed in 2.1

The Engineering Skills Model is Codility’s validated taxonomy of engineering capability: 30 skills and 180 subskills, 210 elements in all, across five categories including a first-class Artificial Intelligence category. Every Codility task and score traces back to it. Version 2.1 is the largest revision since the model began.

Key takeaways

  • The model has five categories: Technical, Artificial Intelligence, Problem-Solving, Business, and Technology Stack.
  • Artificial Intelligence is now a category of its own, with five skills spanning fluency, collaboration, responsibility, governance, and engineering.
  • Programming languages and technologies moved inside the model as a Technology Stack category, 107 entries in all.
  • Depth varies by design. Some categories resolve at skill level because the construct does not support finer observable behavior.
  • The model is a competency framework, not a selection procedure, and not something you buy. It is the methodology underneath the platform.
  • Every task maps to a skill in the model. Not every skill has a task, and that is deliberate.
  • You can walk all 210 elements yourself in the interactive model, linked at the foot of this piece.

Why does an assessment platform need a skills model at all?

Because a score means nothing unless it traces back to a defined capability. A number with nothing behind it invites challenge, and engineering leaders are right to distrust one.

A skills model is the structure that turns a result into evidence. It names what was measured, why that capability matters for the role, and how the task in front of the candidate connects to the work someone will actually do.

This shows up in buying conversations as a blunt question: will this fit our taxonomy? Whether a model can align with how a company already describes its engineers is one of the most common objections raised in evaluation, because a framework that cannot connect to existing internal language creates friction instead of clarity.

There is a second, less discussed reason. A model decides what gets built. Assessment content follows the taxonomy, so the model is also the roadmap for what a platform can measure next.

What is actually inside the model?

Five categories hold 30 skills, which resolve into 180 subskills and a total of 210 addressable elements Codility can help measure across roles and organizations.

CategorySkillsSubskillsWhat it covers
Technical949Core software and systems engineering, from construction and testing to architecture and working with data
Artificial Intelligence524Human and AI competence as a cross-cutting domain, from fluency through to engineering AI systems
Problem-Solving4Resolves at skill levelCognitive skills underpinning innovation and adaptability
Business10Resolves at skill levelInterpersonal, collaborative, and delivery competencies
Technology Stack2107Catalog of programming languages and technologies
Total30180210 elements in all

The structure nests, and one thread makes it concrete. A category sets the domain. A skill within it names a capability. A subskill narrows that capability to something observable. When a candidate completes a task, the resulting signal rolls back up the same path, so a proficiency result ties to the specific capability it reflects.

Each element carries a code, from ESM 01.0 through ESM 30.0 at skill level, with subskills numbered beneath. The codes underpin crosswalks and mapping exercises: they are what makes a mapping exercise between the model and an internal framework a structured job rather than a conversation about synonyms.

Why do two categories have no subskills?

Because depth was added only where a capability resolves into distinct observable behavior. Problem-Solving and Business sit at skill level, and that is a design decision rather than unfinished work.

Systems Thinking, Critical Thinking, Teamwork, and Behaving Ethically are real constructs, and they are in the model because engineering work demands them. Splitting them into six neat subskills each would artificially create a precision the underlying evidence does not support.

This is worth dwelling on if you evaluate assessment frameworks for a living. Taxonomies tend to drift toward uniform depth, because uniform depth looks rigorous in a diagram. Uniform depth is also the fastest route to constructs that exist on a slide and nowhere in the measurement.

The asymmetry runs the other way too. Technology Stack holds 107 of the model’s 180 subskills, so more than half the subskill volume in the model is a catalog of languages and technologies. That is breadth doing a different job from depth, and reading the two as the same quantity would misread the model.

What changed in version 2.1?

Four changes, and most existing structure carries forward. Roles already defined in 2.0 remain valid and have a documented path into 2.1.

ChangeVersion 2.0Version 2.1
AI elevated to a full categoryAI competence scattered through the technical domain as readiness additions: AI Literacy, AI Evaluation, AI Application, AI BuildingA first-class Artificial Intelligence category with five skills: AI Fluency, AI Collaboration, AI Responsibility, AI Governance, AI Engineering
Languages and technologies join the modelHeld outside the model, in standalone reference appendicesA Technology Stack category, two skills, 107 entries in all
Supporting becomes BusinessThe Supporting category, covering collaboration and deliveryThe Business category, plus a new Project Management skill covering delivery and Agile methodologies, scope, timelines, and budgets
Technical refinementsSoftware Architecture and Systems Architecture as two separate skillsOne Software and Systems Architecture skill, all subskills retained, definitions modernized

The AI reorganization is the largest single move, and nothing was dropped in it. Every AI competency from 2.0 has a documented home in 2.1. AI Literacy became AI Fluency. Prompt Engineering became Context Engineering beneath AI Fluency. AI Evaluation split, with output checking landing under AI Collaboration as AI Output Validation and technical evaluation landing under AI Engineering. Responsible AI became AI Responsibility. AI Building became AI Engineering.

AI Governance is entirely new: AI Regulation, AI Policy, Human Oversight, and AI Risk Management, alongside new entries including AI Security, AI Cost Optimization, Context Hygiene, and Agent Orchestration.

Several subskills were also renamed for clarity, with the competency unchanged. Release Management became Revision Control. Leveraging APIs became API Integration. Low-level Programming became Embedded Programming with an embedded-systems focus.

Why did AI become a category of its own?

Because AI matured into a distinct, cross-cutting domain of engineering competence, and the model followed the work. Writing code is no longer the whole job. Directing AI, checking what it produces, and owning what ships are now core to how engineers deliver.

Where AI sits in the model carries an argument. It spans the technical and business halves rather than nesting inside either, because the work it describes draws on judgment and commercial sense as much as technical depth. An engineer who delegates well, spots a plausible and wrong answer, and knows which decision cannot be handed to a model is exercising something broader than a technical subskill.

The five skills split the domain along lines that hold up in practice:

  • AI Fluency. Understanding what AI is, how it works, and where it fails. Covers AI Foundations, AI Failure Modes, Context Engineering, AI Data Literacy, and Context Hygiene.
  • AI Collaboration. Working through the full interaction loop, from delegating a task to refining output, including supervising autonomous agents. Covers AI Task Delegation, AI Iteration, AI Output Validation, and Agent Orchestration.
  • AI Responsibility. Treating ethics as an active competency. Covers Algorithmic Fairness, AI Explainability, AI Accountability, and AI Ethics.
  • AI Governance. Operating inside regulation, policy, and standards. Covers AI Regulation, AI Policy, Human Oversight, and AI Risk Management.
  • AI Engineering. Building and running AI systems. Covers Machine Learning, LLM Development, AI Architecture, MLOps, AI Evaluation, AI Security, and AI Cost Optimization.

The grounding for treating this as measurable capability comes from AI-literacy research, validated measurement instruments, and governance frameworks including the NIST AI Risk Management Framework, the EU AI Act, and ISO/IEC 42001. Effective AI use is a learned competency rather than an automatic byproduct of tool access, which is the case for measuring it directly instead of accepting a self-reported claim.

Why are programming languages part of the model now?

Because tool proficiency is real capability, and keeping it in an appendix implied otherwise. Technology Stack brings it inside the model in the same competency vocabulary as everything else, across two skills: Programming Languages with 27 entries and Technologies with 80.

The catalog itself is largely carried forward from 2.0, so this is a structural change more than a content change. What it buys is a single place to express a role: the capability someone needs and the stack they need it in, described in one model instead of two documents.

It also removes an awkward gap in workforce planning. A team that wants to know who could move onto a Terraform-heavy platform team was previously reading a skills framework and a technology list side by side, and reconciling them by hand.

How does the model actually work in an assessment?

Every task on the platform maps to a skill in the model, and the score a task produces reports against that skill. That mapping is the whole mechanism, and it runs in both directions.

Forwards, it is how a role gets defined. A role profile sets which skills matter and what proficiency each needs, and seniority shifts those expectations rather than changing the skills themselves. Proficiency runs from Novice through Expert, with elements outside the role marked as such instead of scored as zero.

Backwards, it is how a result gets explained. A proficiency reading resolves to a subskill, that subskill to a skill, and that skill to a category, so a candidate result can be discussed in terms of the capability it reflects rather than a single undifferentiated number.

The same taxonomy runs across hiring and the existing workforce. A candidate assessed in Screen and an engineer measured in Skills Intelligence are described in the same vocabulary, which is what makes capability comparable from first touch through to internal redeployment.

One honest limit belongs here. Not every skill in the model has a task today, and it is designed that way. The model describes the competencies that matter now and in the near future, so it runs deliberately broader than current content coverage. Some competencies, including parts of AI Responsibility and Behaving Ethically, are better evaluated through an interview or a work sample than an automated coding task.

How was version 2.1 built and validated?

Through a three-phase, evidence-based methodology, run by Codility’s Assessment Science team and maintained by occupational psychologists.

The first phase was a structured survey of the authoritative skills frameworks, which is what re-anchored the model rather than letting it drift on internal opinion. The second brought in internal and external subject-matter experts. The third was an organizational confirmation survey of engineering subject-matter experts spanning job families and seniority levels.

The confirmation result is the one worth reporting: no respondent proposed changing a single subskill. Skills were retained against agreement thresholds rather than kept because they were already there, which is the discipline that stops a taxonomy from accumulating.

The model’s wider evidence base includes a quantitative criticality survey in which practitioners rated skill importance across job families, alongside four documented iterations of the model with a crosswalk at each revision.

The mapping to external frameworks is deliberate and broad: O*NET, SWEBOK v4 from the IEEE Computer Society, SFIA v9, the NIST AI Risk Management Framework, the EU AI Act, ISO/IEC 42001, Anthropic’s AI Fluency framework, the Great Eight competency model, and Lightcast’s open skills taxonomy, among others.

If your interest is psychometric rather than structural, the boundary matters. The Engineering Skills Model defines and organizes job-relevant constructs, which is the recognized foundation for defensible selection. The reliability and fairness evidence for what the assessments themselves produce lives in Codility’s Technical Manual, structured to the APA Standards for Educational and Psychological Testing and referencing the EEOC Uniform Guidelines and SIOP Principles. Reliability there is reported with McDonald’s omega, the appropriate coefficient for multidimensional composite assessments rather than the more commonly quoted alpha.

Keeping those two documents distinct is the point. A competency framework and a technical validation report answer different questions, and collapsing them is how vendors end up claiming validity they have not evidenced.

How does it compare to SFIA, SWEBOK, and O*NET?

The Engineering Skills Model is purpose-built to measure engineering capability for assessment. SFIA, SWEBOK, and O*NET were built for career frameworks, engineering knowledge, and occupational reference. The model maps to them rather than competing with them.

FrameworkPrimary purposeHow the Engineering Skills Model relates
Engineering Skills ModelMeasures engineering capability for assessment, with AI as a first-class categoryThe native model. Every Codility task and score traces to it
SFIA v9Describes professional skills and levels of responsibility for roles and careersMapped to as an external reference
SWEBOK v4Codifies the recognized body of knowledge for software engineeringMapped to as an external reference
O*NETCatalogs occupations and their associated tasks and skillsMapped to as an external reference

None of the three was designed to produce a scored, traceable assessment result. That is the gap the model fills, and the mapping is what lets a team keep the language it already uses.

What does the model not do?

Four boundaries, stated plainly, because each one is a place where frameworks like this get oversold.

It is not a selection procedure and it makes no hiring decisions. It defines and organizes constructs. The validity of any specific hiring use sits with the implementing organization in its own context.

It is not a product. There is no Engineering Skills Model to buy. It is the methodology underneath the assessments already running on the platform.

It is not the same thing as the AI features inside the product. The Artificial Intelligence category describes the AI skills an engineer needs. The in-product AI Copilot is the controlled assistant inside an assessment, enabled or disabled per assessment, with every interaction captured as reviewable AI activity. AI for Business is the Skills Intelligence offering that measures how people in non-technical and tech-adjacent roles work with AI. The model describes the skills, the products measure them.

It is not a complete map of current content coverage. Model breadth and task coverage are different things, and the model leads.

How do I explore the model myself?

Walk it. The interactive model opens all 210 elements, and reading a taxonomy in a table is a poor substitute for following one thread from a category down to a definition.

Explore the Engineering Skills Model

What it lets you do:

  • Read every definition. All 30 skills and 180 subskills carry their full definition and their ESM code.
  • Change the view. A radial map for structure, a tree for hierarchy, and a grid for scanning breadth.
  • Search the whole model. Jump straight to any element by name.
  • Apply a role and a seniority level. Eight sample engineering roles across five levels, from Junior to Principal, show how proficiency expectations shift while the skills stay fixed.
  • See which roles value a skill. From any element, look across the roles that depend on it.

One thing to read correctly. The role profiles in the interactive model are illustrative, built to demonstrate how role-based expectations work against the taxonomy. They are not published benchmarks, and a real role profile is defined by the implementing team, with Codility Advisory available for taxonomy mapping where a model needs to connect to an internal framework.

That is what a skills model is for. A score you can explain, backed by a capability you can inspect.

If this article was just too light for your needs, you can also read the full Engineering Skills Management Technical Report.

Open the Engineering Skills Model’s Technical Report

Frequently Asked Questions

What is the Engineering Skills Model?

It is Codility’s research-backed competency framework, built and maintained by occupational psychologists on the Assessment Science team. It describes the role-relevant competencies engineering work requires across job families including backend, frontend, data science, quality assurance, mobile, and platform engineering. Customers use it to define roles and hiring profiles, and it guides how Codility designs and prioritizes assessment content.