AI use should count in performance reviews when an employee can show how it affected useful work, the judgment they exercised and what they learned. Agree those expectations before the review period. A high prompt count cannot tell you whether a customer received a better answer or a colleague inherited more checking.
For managers and people teams, the practical task is to connect AI learning to the work already being assessed. Keep the quality standard visible, make support available and ask for evidence the employee can explain. This article offers a way to structure that conversation; the broader people and change guide covers the conditions that help employees adopt AI.
Agree what good work looks like before adding an AI goal
Start with the employee's responsibilities. A museum communications officer needs to publish accurate exhibition information that visitors can understand. Using an approved AI tool to draft a listing may help, but the listing must still have the correct dates, opening times, booking conditions and access information.
An objective such as “use AI more” leaves the important decisions unresolved. Before adding it to a review, agree:
- The work. Which recurring task or problem is the employee expected to improve?
- The standard. What makes the finished result acceptable, and who checks it?
- The development. What should the employee learn to do independently?
- The conditions. Which approved tool, information, practice time and support will be available?
- The review point. When will you examine an ordinary attempt and decide what changes next?
CIPD's performance management guidance treats performance as both results and behavior. It also distinguishes core work, contributions beyond the immediate role and adapting to change. For complex tasks, it recommends considering learning outcomes or open-ended objectives rather than rigid targets. That gives managers room to discuss development without pretending that every experiment must immediately produce a saving.
Tell employees how these expectations connect to the existing review process. Do not introduce a new AI criterion at the end of the period or leave different managers to invent incompatible interpretations of it. If the organization wants a formal rating category, HR should establish and test its criteria before it affects employment decisions.
Treat usage counts as a question to investigate
Usage data can help identify access problems, support needs or spending. It needs interpretation. A long conversation with an assistant might represent careful exploration, repeated failure or unnecessary activity. A short interaction might be enough for a familiar task. Neither count establishes the quality of the finished work.
A September 2026 working paper by Gong and colleagues examined a Chinese medical technology company that required 200 AI queries per employee per month, then lowered the target to 100 across branches at different times.
31%
of queries under the 200-query target were repeated or classified as personal or off-task
The study covered about 5,000 nonproduction employees at one company. Repetition meant the same employee submitted an identical query earlier in the month. The measure describes queries, not the correctness of AI outputs.
First use rose after the mandate, but the authors could not establish that the mandate caused that initial increase. Their analysis of the later reduction found that repeated or off-task queries accounted for most of the fall in usage. The study does not establish an ideal quota, and its findings come from one company. It is a reason to examine what a target encourages, not a universal prescription for how much AI people should use.
The concern also appears in employee accounts. In a March 2026 Reddit discussion about AI-related performance reviews, a Reddit user described a former employer's rising thresholds and automatic performance improvement plans, saying team output did not change the rule. The employer and policy were not independently verified. The account illustrates the concern an employee may bring to a review: whether useful work matters more than satisfying a counter.
Keep activity data separate from the judgment about performance. If someone rarely uses an approved tool, ask whether relevant tasks, access and support exist. If usage is high, ask to see the result and the checking it required. Both conversations need more than the dashboard.
Separate results, judgment and learning in the evidence
Bring a small selection of ordinary work to the conversation. The purpose is to understand the employee's contribution, including what they changed or rejected. A polished success story alone may hide the effort needed to make the method dependable.
Use these distinctions to organize the discussion. They are prompts for judgment, not a weighted scoring formula.
| Evidence | What to ask | What it cannot establish alone |
|---|---|---|
Finished work | Did the result meet the agreed standard and help its recipient? | That AI caused an improvement |
Checks and decisions | What did the employee verify, correct, reject or escalate? | That a long checklist was applied well |
Learning from an attempt | What can the employee now explain or do that they could not before? | A production benefit from an unfinished experiment |
Help given to colleagues | Could another person use the guidance, example or feedback? | Team benefit from the number of posts or demonstrations |
Conditions provided | Did access, suitable information and agreed support actually arrive? | An employee's capability when those conditions were missing |
When discussing efficiency, include the work required to check and finish the output. A quicker first draft may still create more revision for an editor. Compare reasonably similar tasks, record important differences and avoid attributing every change to AI. The complete-task measurement guidance explains how to include review and rework.
Learning deserves a specific description too. “Completed the course” says something about participation. “Can identify unsupported claims in an exhibition listing and trace each factual detail to the approved brief” says what the person has learned to do. Ask them to demonstrate that ability on appropriate material.
Review a real task without demanding a success story
Suppose Alex Morgan, a museum communications officer, tries an approved AI assistant to draft an exhibition listing from a curator's brief and the museum's visitor information. The draft is fluent, but it adds free parking that the source material does not promise. Alex removes that statement, checks the dates and booking link, and sends the corrected listing through the normal editorial approval.
The manager can discuss the accepted listing and Alex's checking decisions separately. Alex caught an unsupported claim before publication. That is evidence of judgment. It does not, by itself, prove that drafting with AI was faster or that every future listing will be accurate.

For the next attempt, Alex agrees to test whether supplying a clearer source brief reduces unsupported additions. The learning goal is to explain which changes helped and which checks remain necessary. The manager provides suitable material and editorial feedback, with a review date after the next listing.
If the revised approach still takes longer, record that result. Alex may have completed the learning objective by establishing that the method is unsuitable for this task. The production expectation remains an accurate listing delivered through the agreed process. A learning objective should make that distinction explicit from the start.
Check whether the system supports the behavior it asks for
A review can reward learning on paper while the delivery plan makes it impractical. If employees are expected to practise, coach colleagues or document a useful method, make room for that work. Use the learning-time agreement to name which commitments change and who approves them.
Team contribution needs similar care. Suppose Alex turns the exhibition checks into a short guide and a colleague uses it to catch an incorrect booking date. The evidence is the usable guide and the colleague's application of it. Counting how many guides Alex uploads would miss whether anyone can use them. If helping others becomes a regular responsibility, agree its scope and protected time.
Before managers apply the expectations more widely, compare how they would interpret a few appropriately redacted examples. Include:
- a strong result produced with little AI use;
- a careful experiment that showed AI was unsuitable;
- a fast draft that shifted substantial correction work to someone else;
- an employee whose agreed access or support never arrived.
Ask where the judgments differ and which missing facts would change them. Revise unclear criteria before they shape ratings. Give employees a way to add context and correct inaccurate evidence through the existing review process.
Keep responsibility on both sides of the conversation. The employee explains their work and follows agreed standards. The manager provides the promised conditions, resolves conflicting priorities and assesses the evidence in context. Revisit the agreement as the work or tools change.
Questions about AI use in performance reviews
Should AI adoption be an employee performance goal?
An AI-related goal can be useful when it names a relevant task, a quality standard and a capability the employee is expected to develop. Confirm access, approved information and support before setting it. Avoid making raw usage the goal: it cannot establish whether the work improved. Start with the expectation-setting questions and fit the agreement into the organization's existing review process.
How is an AI learning goal different from a performance goal?
A learning goal describes what the employee will learn to understand or do. A performance goal describes the required work result. A museum communications officer might learn to identify unsupported claims in AI drafts while remaining responsible for accurate, approved exhibition listings. Agree what evidence will demonstrate each goal, and allow an honest learning result that a method is unsuitable. The museum example shows how to discuss both without promising a saving.
What if an employee chooses not to use AI?
Ask about the task and the reason before drawing a conclusion. The tool may be unsuitable, the information may not be approved for it, or promised access and training may be missing. A sound choice to use another method can coexist with good performance. Where developing an AI-related capability is an agreed role expectation, identify the specific gap and arrange appropriate practice and support. Assess demonstrated work and judgment, rather than treating nonuse alone as a rating.
Should managers read every prompt to assess performance?
A complete prompt history is not a substitute for examining finished work and discussing decisions. Start with selected, appropriate work samples and the employee's explanation of checks, corrections and results. Use existing privacy and access rules, and explain what evidence is being collected and why. If usage records are needed for a specific support or governance purpose, keep that purpose distinct from the performance judgment; do not turn the conversation into an unexplained activity ranking.



