Measure AI proficiency through work people can complete and explain under known conditions. For a company learning programme, start with a relevant task, inspect how the employee checks the result, then try a comparable task with different inputs. Record where help was needed and what the evidence supports them doing next.
A useful conclusion sounds like “Alex can prepare these print-production instructions independently and recognize when the order is incomplete.” It should tell the employee what to practise and the work owner what support to retain. A company-wide score rarely supplies that detail on its own.
Decide what the assessment needs to tell you
Knowledge, confidence, use and demonstrated performance answer different questions. Keep those questions visible when choosing an assessment.
| Evidence | What it can help you understand | What still needs checking |
|---|---|---|
A knowledge test | Understanding of concepts, risks and appropriate uses | Whether the employee applies that understanding in the role |
A self-assessment | Confidence, perceived difficulty and requests for support | What the person can demonstrate |
Tool activity | Where and how often use occurs | Whether the work is correct, appropriate and independently completed |
Observed work and explanation | Decisions and results on the sampled task | Whether the capability holds on other relevant tasks |
There is a place for a well-designed knowledge test. In the GLAT study by Yueqiao Jin and colleagues, a 20-question generative AI literacy test was validated with higher education students. In a separate task study with 83 students from medical, healthcare and nursing backgrounds, test scores predicted performance on an AI-assisted visual-analysis task; self-reported ChatGPT literacy did not significantly predict that performance in the model. The task used a particular chatbot and educational setting. It does not establish a pass mark for your employees.
The practical implication is to match the evidence to the decision. Ask about confidence if you want to understand confidence. For independent work, also observe the work. For a common foundation, the employee AI literacy guide helps define what everyone needs to understand before role-specific practice.
Make the task reveal a decision
An assessment should give the employee a reason to exercise judgment, rather than reward familiarity with a trainer's example.
Suppose Alex Morgan coordinates production at a commercial print shop. Alex uses an approved AI tool to turn a permitted sample customer order into a draft instruction for the press team. The order specifies matte paper and printing on both sides, but leaves the binding method unanswered. A previous example used glossy paper and stapled booklets.
The assessor, Claire Bennett, wants to see whether Alex preserves the new requirements and leaves the unanswered choice open. A fluent instruction that silently carries over the old paper finish or binding method could send the press team into the wrong job.

Claire defines the evidence before the attempt:
- Task choice: Alex can explain what drafting assistance is useful for and what the tool must not decide.
- Input handling: Alex uses the permitted sample order and supplies the relevant requirements.
- Checking: Alex compares the draft with the order, including the paper finish and printing on both sides.
- Uncertainty: Alex spots the missing binding instruction and asks the work owner to clarify it.
- Explanation: Alex can show the source of an important requirement and describe a correction without Claire supplying the answer.
These criteria fit this job. They are a starting point for local assessment, not a standardized proficiency test. The Skills England foundation benchmark, published in January 2026, includes clear instructions, routine uses and checking information alongside responsible use. That foundation can inform your criteria, but it does not replace the role's actual work requirements.
Task alignment also needs testing. In Christopher Bogart and colleagues' 2025 preprint, researchers working with a small US Navy robotics-training programme found that contextual scenario questions related more strongly to some competition tasks than a broader AI literacy test. The complete comparison involved only ten participants. The authors still needed to establish whether the scenario context or the depth of reasoning explained the difference. This supports investigating fit, without assuming that any realistic-looking exercise is valid.
Use the role-specific learning-path guide to choose the relevant work. Someone who drafts press instructions need not demonstrate the same capability as someone building the print shop's scheduling software.
Check whether the result travels
One successful attempt establishes what happened on that attempt. It might reflect a familiar template, help from the observer or a particularly good model response. Keep the employee's contribution distinguishable from those conditions.
After feedback and relevant practice, give Alex another comparable order. Change the customer requirements so copying the previous answer will not work. Keep the capability being assessed recognizable: preserving specified requirements, checking the generated instruction and identifying an unresolved choice. Avoid making the second task harder through an unfamiliar printing process that Alex has never been taught.
Record both what happened and what changed:
- Which decisions did Alex make independently?
- Which prompts or corrections did Claire supply?
- Did Alex find the missing requirement before being asked?
- Did the tool, starting aid or task difficulty change between attempts?
Improvement on the later sample can guide the next learning decision. It does not, by itself, prove that training caused the improvement or that Alex can handle every production order.
The assessor's consistency matters too. Ask two assessors to review a small set of the same samples independently, using the agreed criteria, then compare their judgments. If one accepts an unexplained correction and the other requires Alex to locate its source, resolve that difference before interpreting scores across employees. Agreement does not establish validity, but disagreement exposes a problem you can fix.
Give people a fair chance to show the skill
Tell employees what the assessment is for, what evidence will be retained and who will see it. For a learning assessment, make the next development decision explicit. Do not quietly turn an exploratory exercise into a promotion ranking or pay decision; that use needs a separately justified assessment process.
Before comparing results, check the conditions:
- Access: Can everyone use the approved tool and the same required features?
- Materials: Are the instructions, source documents and interface usable with the accessibility support people need?
- Job familiarity: Does the sample represent work the employee is expected to know?
- Permitted support: Are starting aids, assistance and time conditions defined and recorded?
Independence can include an approved job aid or assistive technology. Define it in relation to the job, rather than treating all support as cheating. If the interface prevents a legitimate attempt, record an assessment blocked by access instead of a low proficiency result.
A short conversational assessment can still help someone reflect. In a September 2026 LinkedIn post, marketing adviser Jim Kingsbury described finding an adaptive AI interview more responsive to individuals than a static assessment. His suggested prompt asks the system to use what it already knows about the person and pose progressively harder questions. He also links a commercial assessment in the comments. The post reports his preference, without validation or evidence of later workplace performance.
Use that kind of conversation to generate questions for practice. If the system knows different amounts about different employees, its judgments are not starting from comparable evidence. Follow a promising explanation with a task the person can show you.
Turn the evidence into the next practice decision
Keep a short capability record that names the task, date, tool, permitted support and observed decisions. Add the specific learning need, rather than compressing everything into one average.
For Alex, the record might show that source checking is independent while recognizing incomplete orders still needs a prompt from Claire. The next practice should make that decision visible again. More instruction on writing longer prompts would miss the gap.
Use assessment to choose the next practice
Define the decision
Name the job task, important requirements and evidence of independent judgment.
Observe and explain
Inspect the attempt and let the employee show checks, corrections and support used.
Compare a changed task
Sample the same capability with different inputs and record changed conditions.
Choose the next support
Agree specific practice or an access fix, then revisit the relevant capability.
Revisit the assessment when the work, tool or permission boundary changes materially. If a new tool feature performs a step that employees previously did themselves, decide what they must still understand and check. Keep old and new results attached to their conditions instead of presenting a continuous score as though nothing changed.
The people and change guide places this learning decision alongside the responsibilities and support that make useful adoption possible. Start with one role and a small set of comparable tasks. Improve the assessment when it fails to distinguish a capability gap from an unsuitable test.
Questions and answers
Can a company use one AI proficiency score for everyone?
A shared score may summarize a common knowledge test, but it cannot establish that employees can perform every role's AI-assisted work. Keep task-specific evidence alongside it, including important errors and support needs. Before using a score for consequential employment decisions, justify that intended use and the scoring process separately.
Start by matching the assessment evidence to the decision. A production coordinator and a software builder need different demonstrations of capability.
Is self-assessment useful for measuring AI proficiency?
Self-assessment is useful for understanding confidence, perceived difficulty and what support an employee wants. It should not substitute for observing the capability you need to assess. Ask the employee to complete a relevant task and explain an important check, then compare that evidence with their own view.
The comparison of assessment evidence shows how to keep these purposes separate. A mismatch can open a useful coaching conversation without proving that the employee was dishonest or careless.
How often should you reassess employees' AI skills?
Reassess after relevant practice when a comparable task is available, and when a material change in the tool, work or permissions changes what the employee must do. Record those changed conditions. Repeating the same exercise can show familiarity with the exercise rather than capability across different inputs.
Use a changed task and consistent assessment criteria to check the specific skill. The evidence here does not establish one best interval for every company or role.
What if an employee needs help during the assessment?
Record the help and which decision it supplied. An approved starting aid or accessibility support may be part of ordinary independent work; an assessor telling the employee which error to correct is different evidence. Define permitted support beforehand, and keep access problems separate from capability gaps.
Review the conditions for a fair attempt, then choose practice that addresses the observed difficulty. If the employee could not make a legitimate attempt, resolve that barrier before drawing a proficiency conclusion.



