September 2026 | Improved AI Score Accuracy and Items That Could Not Be Assessed Are Now Shown as "Not Assessed"

Account

We have improved how the AI Score shown in AI Collaboration format reports is evaluated.

The evaluation is now based on the candidate's own messages and actions, and an item with no opportunity to assess it is shown as "Not assessed" instead of receiving a score. The evaluation items and the level criteria for each item are unchanged.

Background

When a candidate works together with AI, the log contains not only the candidate's prompts but also the AI's responses and actions, the task statement, and, in interactive questions, the persona's answers.

The previous evaluation sometimes treated a judgement made by the AI or the persona as the candidate's own behaviour, and sometimes scored items that the session gave no opportunity to assess. As a result, a score could rise because of something the candidate did not actually do, or fall because there was nothing to judge by.

What has changed

Evaluation is based on the candidate's own messages and actions

The evidence for an evaluation is the prompts and messages the candidate wrote and the candidate's own actions in the editor. The AI's responses and actions, the task statement and the persona's answers are used only as context to understand what the candidate responded to, and are not used as evidence.

Any AI output or task statement that the candidate pasted into a prompt is not treated as the candidate's own words either.

The evaluation details quote the candidate's messages and actions that the evaluation rests on.

Only items with an opportunity to assess them are scored

For each evaluation item, the evaluator first decides whether the session gave an opportunity to assess it. Only those items are scored, by how far they meet the level criteria.

An item with no opportunity to assess it receives no score, is shown as "Not assessed", and comes with the reason. Items that are not assessed are left out of the score calculation.

A score of 0 and "Not assessed" mean different things.

Display Meaning
0 There was an opportunity to assess the item, but not even the lowest level was met
Not assessed (-) There was no opportunity to assess the item, or no evidence could be confirmed

Evaluation coverage is shown

"Evaluation coverage" is shown below the AI Score. It indicates how many of the evaluation items the score is based on.

The wider the coverage, the more items the score draws on. Depending on the task and how the candidate proceeded, some items may not be assessable.

The evaluator version is shown

The report shows the version of the evaluation method as "Evaluator version". Reports evaluated with the new method show v0.2.0, and reports evaluated with the previous method show v0.1.0.

Interactive questions evaluate problem solving and communication in the same way

In interactive questions, where the candidate works while talking with a persona, Problem Solving Skills and Communication Skills are evaluated with the same method as AI Collaboration Skills.

Internal validation

Using past assessment data, we evaluated the same submissions with the previous and the new method and had a separate AI model audit the results.

Metric Change from the previous evaluation
Share of evaluations judged valid About 3 times higher
Share of evaluations based on someone other than the candidate About one third

We will keep improving the evaluation method.

* Results of auditing 30 past assessments with an AI model separate from the one used for evaluation. The metric is for comparing methods and does not represent the absolute accuracy of the evaluation.

Impact on existing reports

  • The improvement applies to submissions evaluated after the release. Reports that have already been created are not re-evaluated.
  • Because items with no opportunity to assess them are left out of the score, scores may trend differently from reports evaluated with the previous method. Please take care when comparing reports with different evaluator versions.
  • The evaluation of deliverables and the calculation of the total test score are unchanged.

For the list of evaluation items and how they are evaluated, see About AI Score.