July 2026 | Updated AI Collaboration Skills Evaluation Criteria to Better Assess the Depth of AI Use
AccountWe have updated the evaluation criteria for “AI Collaboration Skills” shown in AI Collaboration format reports.
Rather than assessing only whether a candidate completed a deliverable with AI, the evaluation now uses actual behavior logs to assess whether the candidate independently framed the problem, decided what to delegate to AI, considered its proposals, and led verification and improvement.
Background
As AI use becomes commonplace, hiring teams need to determine not simply whether a candidate can complete a task with AI, but whether the candidate can examine AI proposals, make independent decisions, and take responsibility for the quality of the outcome.
Under the previous evaluation criteria, candidates who could instruct AI to implement and verify their work tended to receive high scores overall. Differences in problem framing, scrutiny of proposals, and depth of verification were therefore not always reflected clearly in scores. We revised the criteria to distinguish standard AI use from skilled AI collaboration.
Five New Evaluation Axes
We reorganized the evaluation criteria into five axes and 13 specific criteria.
| Evaluation Axis | Behaviors Evaluated |
|---|---|
| Environment Engineering | Prepare tools and context so that AI can work appropriately, and establish rules for safe execution |
| Intent Specification | Frame the problem in the candidate’s own words and communicate constraints and completion criteria clearly |
| Co-Reasoning | Elicit alternatives from AI and decide whether to adopt them based on rationale and trade-offs |
| Execution Control | Adjust the scope delegated to AI, steer work during execution, and guide convergence after failures |
| Output Quality Assurance | Design verification methods, critically assess AI output, and further improve quality |
For detailed definitions of each axis and its criteria, see AI Collaboration Skills | SEI.

Reports display the five evaluation axes in a radar chart.
As of July 2026, the five axes comprise 13 evaluation criteria. We will continue to improve and expand these criteria as AI use evolves. Each criterion is evaluated based on observable behaviors in candidate prompts, AI execution logs, test results, and other evidence.
Score Distribution Verified Through an Internal Benchmark
As part of the redesign, we compared score distributions under the previous and new criteria using actual exam reports.
The median full-score rate across evaluation criteria decreased from 88.5% under the previous criteria to 36.7% under the new criteria.
| Comparison Metric | Previous Criteria | New Criteria |
|---|---|---|
| Full-score rate across evaluation criteria (median) | 88.5% | 36.7% |
| Standard deviation by evaluation criterion (median, normalized to a 0–1 scale) | 0.17 | 0.29 |
The concentration of full scores decreased while the standard deviation increased, confirming that differences in candidate behavior are more readily reflected in scores.
The new criteria are not designed to give every candidate a uniformly lower score. Standard behaviors are positioned in the middle of the scale, while skilled behaviors in which candidates lead problem framing, delegation decisions, proposal review, and verification are distinguished at the upper end.
How to Use the Evaluation in Hiring
The updated AI Collaboration Skills evaluation makes it easier to use the report in the following hiring decisions:
- Compare differences in AI collaboration processes between candidates whose deliverables are similar in quality
- Review the behaviors behind each score in Playback to build confidence in the evaluation
- Explore which decisions candidates made themselves and why they chose a particular approach during interviews
This update applies to “AI Collaboration Skills” within AI Score (β) for the AI Collaboration format. There are no changes to deliverable evaluations or the calculation of total test scores.