Article Skills-Based Hiring

How do you run early careers hiring at scale without losing the signal?

In short

Enterprise early careers hiring needs a documented assessment standard across markets, a clear AI policy, an accessible candidate workflow and human review of integrity signals. Keep the assessment to 30 to 45 minutes, then analyze completion, score quality and fairness before the next season.

  • Early careers hiring puts more weight on the assessment because CVs show little job evidence.
  • At enterprise scale, define one assessment and review standard across markets, campus programs and virtual events.
  • Use 30 to 45 minutes across one or two tasks, then watch completion and drop-off.
  • Build each program’s assessment with its engineers, let candidates choose their language and set accessibility up front.
  • State whether AI is allowed, capture reviewable evidence and have a person decide every flagged result.
  • For business and hybrid roles, use AI for Business in Screen to assess a real deliverable made with AI.
  • Close the season with engagement, score distribution, task quality, time use and adverse impact, then connect the results to next season’s design.

Why is early careers hiring a different problem from every other technical hire?

When I think about early careers hiring, I start with the evidence problem. The biggest applicant pools come with the thinnest work history. A graduate scheme, internship or apprenticeship can draw thousands of applications for a few dozen places, while CVs show a course, a project and perhaps a placement. The assessment has to add job-relevant evidence early. Candidates may be applying to several programs at once. A long or confusing process loses them before they show what they can do, and the hiring team still needs a record it can explain when leaders ask about pass rates and fairness.

For many candidates, the assessment is their first interaction with the company. It carries part of the employer brand.

A figure showing thousands of applicants entering one short assessment on the same terms, narrowing to a shortlist, with an end-of-season analysis of completion, fairness at every cut score and score against offers feeding back into the next season's assessment design.

What changes when early careers hiring reaches enterprise scale?

Enterprise scale adds markets, role streams, intake dates, systems and reviewers to the applicant volume. Talent, engineering and legal teams may all need to explain how the decision was made. I would define one documented assessment and review standard for each program before the first invitation goes out, then keep local changes deliberate and recorded.

That standard covers the skills being measured, the scoring criteria, the cut-score rationale, the AI policy, accommodations and the evidence a reviewer needs for an integrity decision. Role-relevant task sets can differ. The decision method should stay stable across the markets running the same program. Codility Screen handles the first pass asynchronously. Candidates take the assessment inside a window you set. Bulk invitations go out through the applicant tracking system (ATS) or graduate portal, and results return there. An intake of 800 and an intake of 8,000 use the same workflow, leaving the recruiting team to focus on the shortlist.

Keep it punchy. Our assessment science team’s guidance is 30 to 45 minutes across one or two tasks. One talent leader put it well: “we don’t want anything that’s going to take two hours, we want it to be quite punchy.” Watch completion and drop-off after launch. Length and unclear instructions show up quickly in that data. Start with a broad pass mark, then refine it over the first few weeks. One HR leader described the approach as “casting a wide net, refining over four weeks“. Compare the score distribution with the interview capacity before you settle the threshold. Randomized task pools reduce the value of a leaked solution. Candidates draw from equivalent tasks, so people in the same intake rarely see the same set.

How do you run an early careers hiring event on campus or online?

Use Codility Screen for campus recruiting events, virtual assessment days and university challenges. Candidates can complete a role-relevant assessment inside a scheduled window, with the results feeding the same review process as the rest of the early careers program.

Screen supports tens of thousands of concurrent sessions. That capacity matters when a national campus campaign, a global virtual event or several university events land in the same window. The talent team can open the event broadly without building a separate manual process for every location. Set the task pool, time limit, AI policy and review rules before the event starts. Randomization limits how often candidates see the same tasks, while reviewable integrity signals give the team a record to inspect after a short, high-volume session.

An event can also widen the entrance to the hiring funnel. Students get a chance to show what they can do before a recruiter has much work history to read, and the strongest results can move into the program’s existing interview workflow.

How do you design the assessment for each program?

I would build the assessment with the engineers who will manage the hires, against the work the program leads to, before the first invitation goes out. A graduate software engineering scheme and a data analyst intake need different evidence. I would protect this design step under time pressure. Codility’s assessment scientists are occupational psychologists. They work with the customer’s subject matter experts to define the skills for the program, then select or build tasks that measure those skills. Standard library tasks have documented content validation. A local validation study is a separate service for the customer’s roles.

Start with language. A cohort may include Python, Java, JavaScript and C++ learners, so language-agnostic tasks let candidates work in the language they know. The same scoring applies across supported languages, keeping the comparison on problem-solving. Then match the task set to the work. Multiple-choice tasks can test reasoning and concepts without code. The library also covers SQL and data work. If a program runs a separate psychometric assessment, Codility can sit alongside it in the same candidate flow. Set the AI policy before the assessment opens. Disable the AI Assistant for a fundamentals check. Enable it when the role calls for AI-assisted work. Every interaction is saved as reviewable AI activity, and the same setting applies to the whole assessment.

How can you assess AI skills in business roles?

This is where AI for Business matters. Early careers programs often run technology and business streams together. AI for Business runs inside Screen and gives business-role candidates a work product to complete with an AI assistant. The result is scored against documented criteria, so the team can review what the candidate produced and how they used the tool. Hybrid roles need both forms of evidence. A forward-deployed engineer, for example, may need to debug a service, understand a customer’s workflow and explain a practical recommendation.

Combine technical tasks with a business-work assessment in AI for Business to examine the full job. The workflow stays in one place. AI for Business uses the same assessment engine, scoring, reporting, integrations and compliance posture as Screen. Set the AI Assistant policy per assessment, and keep every interaction as reviewable AI activity in the candidate record.

What does an accessible assessment mean in practice?

I treat accessibility as part of assessment setup. Every qualified applicant needs a way to complete the test and a clear route to request support. Set the policy before invitations go out. Codility conforms to the Web Content Accessibility Guidelines (WCAG) 2.2 at level AA. The platform supports keyboard operation, screen readers and 400 percent zoom. Candidates can switch on accessibility mode from the test introduction page. A Voluntary Product Accessibility Template is available on request.

A three-column figure of what an accessible early careers assessment means in practice. The platform conforms to WCAG 2.2 AA with keyboard operation, screen reader compatibility and 400 percent zoom. Candidates can turn on accessibility mode themselves and ask the test sponsor for support from inside the test. Common adjustments are extended time set per candidate, the assessment split into shorter sessions, and a live facilitated format.

Inside the test, candidates can tell the test sponsor they need support. The common adjustments are extended time, shorter sessions and a live facilitated format when the screen itself is the barrier. Decide the policy with legal and the talent team, then make it easy to find.

Language affects access too. An independent cApStAn audit found 65 percent of Codility tasks written at or below B1 on the Common European Framework. Plain instructions and program-specific wording help candidates focus on the task.

Candidate feedback gives the team a check on the experience. In a Codility survey, 9 in 10 candidates said the assessment content was fair, and 87 percent rated their experience good or excellent. That is a useful standard for the first interaction with your company.

How do you handle AI cheating in high-volume graduate assessments?

AI-assisted cheating becomes harder to manage when a task reaches thousands of candidates in a few weeks. A leaked prompt travels quickly, and a score alone cannot show how the work was produced. I start with a clear AI rule and a review policy the whole program can follow. The AI Assistant is enabled or disabled per assessment. When it is enabled, every interaction inside the assessment is saved as reviewable AI activity. When it is disabled, the invitation should say what candidates can use and what sits outside the rules.

External AI use rarely leaves one definitive signal. Codility combines secure individual invitations, similarity checks against past submissions, randomized task pools and a replay timeline. The timeline shows paste volume, tab switching and time away from the assessment. Some AI cheating apps use invisible overlays designed to stay hidden during screen sharing. Codility’s cheating apps detection can surface these helper applications and record them in the integrity widget and candidate timeline. A reviewer can inspect that record alongside the other behavior. Optional identity verification adds another layer for programs that need it. Those signals contribute to a deterministic integrity risk rating. The rating directs attention; a person reviews the evidence and decides the outcome. At enterprise scale, write down who reviews a flag, what follow-up happens and how the decision is recorded. I have written more about preventing cheating in coding assessments while keeping human accountability in the process.

What should you analyze when the campaign closes?

At the close of the campaign, I want five answers: did candidates complete the assessment, did scores separate the cohort, did the tasks work, did the time limit fit, and did any group face a lower selection rate at the cut scores we used? The answers tell the team what to change before the next season.

Five cards showing the measures of an end-of-season early careers analysis: engagement from invite to completion, score distribution with a healthy median around 50 to 60 percent, task quality where each task separates stronger from weaker candidates, time use against the window allowed, and adverse impact under the four-fifths rule at every cut score.

Engagement gives you the first signal. Compare invites, starts and completions, then read that against the score distribution. A drop-off after starting points to length or unclear instructions. A ceiling of perfect scores means the task set needs work.

Task quality and time use explain that shape. Each task should separate stronger from weaker candidates, so replace any that everyone passes or everyone fails. Compare completion time with the window too. Candidates using 85 to 90 percent of the time may be carrying too much load.

The fairness check compares selection rates by demographic group at every cut score the program uses. Under the four-fifths rule, a group’s selection rate should be at least 80 percent of the highest group’s rate. The customer’s demographic data powers this analysis.

Use the findings to set the next season’s tasks, time limit and cut-score rationale. Professional Services can support the analysis and document the decisions so the program has a record when leadership asks how the process worked.

How does this fit the systems the team already runs?

Once the assessment is designed, the workflow should disappear into the systems your recruiters already use. Hiring managers should be able to review results without moving between systems. Codility connects with more than 20 applicant tracking systems, including Workday, SAP SuccessFactors, iCIMS, Greenhouse, SmartRecruiters, Lever, Avature and Eightfold. Invitations go out and scores, reports and integrity signals return through the same workflow. Graduate recruitment portals can connect through partners such as Cappfinity, and a public API covers custom flows.

How should you start planning an early careers hiring program?

I would start with the questions leadership will ask when the season ends: who completed the assessment, how scores separated candidates, whether any group was disadvantaged at the cut scores, and what the team will change next time. Then build the assessment with your engineers, keep it short, set the accommodation policy and connect it to the recruiting workflow.

If you are planning a season, talk to us. We will put the assessment science team in the room for the design.

About the author

Christopher Greco is Head of Product Marketing at Codility, where he owns how the platform is positioned across Screen, Interview and Skills Intelligence, and works alongside the Assessment Science team behind the Engineering Skills Model. He came to hiring from the other side of the AI question: before Codility he led product marketing for AI data and model evaluation at Toloka, serving frontier AI labs and large technology companies. He has also built marketing teams from nothing three times, so the hiring problems he writes about are ones he has had himself. He is based in Rome.

Frequently asked questions

What is early careers hiring?

Early careers hiring covers graduate schemes, internships, apprenticeships and entry-level roles for people with little or no professional track record. It often brings thousands of applications for a small number of places, so a structured skills assessment adds job-relevant evidence early in the process.