Skip to content
Broad slate, terracotta and sage pencil-hatched fields meet along two gently descending irregular edges.

Managers and team practicesArticle

How should managers evaluate AI-assisted work?

Review AI-assisted work against task requirements, check the employee's judgment and decide what to accept, correct or escalate. Includes a venue brief example.

Jump to a section

Evaluate AI-assisted work against the task's requirements, then examine the employee's judgment in producing and checking it. A fluent draft is not proof of a usable result. Equally, knowing that AI helped does not tell you whether the employee was careful, capable or responsible. Start with the deliverable and the evidence behind it.

For a line manager, two decisions need separate answers: can someone rely on this work, and what does it reveal about the employee's contribution? This article focuses on reviewing one work sample. The broader people and change guide covers organizational support, while performance systems for useful AI adoption addresses goals and formal ratings.

Establish what the work must do

Before examining how a draft was produced, identify its intended use. A conference crew's setup brief has to tell people what equipment to prepare and which arrangements the client approved. Elegant wording cannot compensate for an invented recording requirement or an unresolved projector specification.

Make the acceptance conditions visible:

  • Purpose and audience. Who will use the output, and what decision or action will it support?
  • Required accuracy. Which details must match a contract, source record or responsible person's confirmation?
  • Consequences of error. What would happen if someone acted on an incorrect or missing detail?
  • Approval boundary. Who can accept the finished work, and which uncertainties need another owner's decision?

Apply these conditions whether the employee used AI, a template or their own first draft. Follow the organization's tool and data rules too. A good output does not resolve a separate concern about prohibited inputs or an unapproved account.

The Government of Canada's generative AI guide, updated in September 2026, advises checking factual and contextual accuracy against trusted sources or colleagues and considering whether the user can confirm the output's quality. Its rules apply to federal institutions. The useful review principle here is narrower: identify what would establish correctness for this task.

Look behind apparent completeness

Leonid Polupan describes reviewing two initiatives, international expansion and a new product, in a September 2026 LinkedIn account. The documents looked persuasive, but broad market figures and unvalidated customer assumptions did not support the decisions at hand. He reports holding one initiative and returning the other for validation.

This is his account of specific reviews, not an independently verified measure of AI's effect on productivity. Its useful detail is the gap between a completed document and completed investigation. Ask which important statements depend on evidence the employee has actually checked.

Suppose Sam Taylor coordinates an event at a conference venue. Sam uses an approved assistant to prepare the crew brief from permitted client requirements, a room plan and an approved equipment list. The draft says every session must be recorded, although the client has made no such request. It also assumes the venue's projector can meet an unusual presentation requirement.

Sam removes the recording instruction after checking the client requirements, then marks the projector question for the technical lead. The manager should recognize both decisions. Sam has prevented one unsupported commitment and identified a limit to what the available material establishes. The brief still cannot be issued as final while a consequential setup requirement remains open.

An event coordinator points to a room plan while a technical colleague checks source sheets beside him at a conference-room desk.
Use the source material in the review. The employee can show what was checked, while a technical colleague resolves a requirement the draft cannot establish.

Review the event brief before issuing it

  1. Confirm the requirement

    Compare the proposed crew instructions with the client's approved requirements.

  2. Inspect a consequential detail

    Trace the recording instruction and projector assumption to their source.

  3. Hear the checking decisions

    Ask Sam what was removed, verified and left for technical confirmation.

  4. Assign the remaining decision

    Have the technical lead confirm the projector arrangement before final approval.

Issue a brief people can rely on

The approving manager checks that the confirmed arrangement is in the final version and the unsupported recording instruction is absent.

The response follows what you find, rather than the confidence of the prose.

Finding in the venue briefDecision on the workFollow-up with Sam
The required details are verified and the approving owner has confirmed them.
Accept the final brief for its agreed use.
Recognize the checks that made it reliable.
An unsupported recording instruction remains, but the client requirements settle it.
Return the brief for a specific correction, then recheck the final version.
Ask why the instruction survived checking and agree how to catch it next time.
The projector arrangement needs expertise or approval Sam does not have.
Hold final approval and give the technical question an owner.
Credit appropriate escalation; make the answer available before the crew acts.

An unresolved question is not automatically a failure by the employee. Determine whether Sam had the relevant information, authority and time. Nor does appropriate escalation make an unfinished brief ready for use. Keep the employee discussion and the acceptance decision clear.

Interpret the employee's contribution fairly

Tool labels can affect how people judge a worker. In four preregistered online experiments published in PNAS in 2025, Jessica Reif, Richard Larrick and Jack Soll studied social evaluations of AI assistance. In an employee-vignette experiment, evaluators judged an AI-assisted worker less competent and diligent, and lazier. Other results varied with task fit and evaluators' own AI use.

These controlled judgments do not show that the workers produced worse real-world work or that every manager reacts alike. They give a reason to check your inference: would you interpret the same source-checking decision differently if the tool label were absent?

Ask the employee to show their contribution using the materials they normally work with:

  • What did the assistant contribute to this deliverable?
  • Which consequential detail did you verify, and against what?
  • What did you change or reject, and why?
  • What remains uncertain, and who can resolve it?

This should be an explanation of the work, not a memory test. Allow the source sheets, notes and accessible communication methods the person needs. A manager can probe an unsupported assumption without demanding a polished live performance.

There is also a difference between accepting help and transferring the thinking to the reviewer. In a May 2026 Reddit discussion about code review, a developer described receiving changes in technologies the submitters did not understand. Their later clarification focused on that missing understanding, rather than AI use itself. The claimed management response was disputed in replies, so it cannot establish that their approach succeeded.

Another Reddit user reported reviewing such changes together with their authors, exposing gaps and seeing improvement. That is a bounded coaching experience, not a proven intervention. For Sam's brief, the equivalent is to trace one disputed instruction together and have Sam make and explain the correction. Quietly repairing every draft yourself leaves the employee's understanding unresolved.

Give checking enough time and a clear next step

Review effort belongs in the work plan. In an October 2026 X discussion, Noah, whose profile describes protocol engineering, warns that rapid generation can consume coworkers' review capacity. He argues for choosing what deserves to be shipped and matching review to the consequences of failure. This is a practitioner's judgment, not a measured rate of reviewer fatigue.

For the venue, match the scrutiny to what the crew will do. An awkward sentence can receive a quick edit. A recording commitment requires source checking. A projector requirement outside the coordinator's expertise requires technical confirmation. The manager needs to arrange the people and time for those checks, rather than praising drafting speed while treating review as spare-time work.

Use a regular manager review to notice recurring problems. Repeated invented commitments may point to weak checking; repeated unanswered technical questions may point to missing support. Investigate the pattern before attributing it to an employee's motivation. A count of drafts, prompts or requested corrections cannot settle that question by itself.

Questions and answers

Should employees receive less credit when AI helped?

AI assistance alone is not a reason to reduce credit. Evaluate the usable result and the employee's contribution, including selection, verification, correction and appropriate escalation. The standard should reflect the task and the responsibility the employee actually held.

For an event brief, coordinator Sam Taylor's removal of an unsupported recording instruction is a meaningful checking decision. Recognize it while holding final approval until the projector question is resolved. Use the contribution questions to establish what the employee did before drawing a broader conclusion.

What if an AI-assisted document reads well but contains an error?

Treat the error according to its consequence and the evidence needed to correct it. An invented instruction to record an event needs removal and a recheck against the client's requirements. An uncertain technical arrangement needs an expert answer before the crew relies on it.

Discuss how the error passed checking and whether the employee had the necessary information and support. Fluency should not excuse an unsupported detail, but one error does not establish the employee's overall capability. The venue-brief comparison shows separate acceptance, correction and escalation decisions.

Is asking another AI system to check the output enough?

Another AI system can help identify questions or inconsistencies, but its agreement does not establish an external fact. For a conference setup, verify the recording requirement against the client's approved instructions and the projector arrangement with the responsible technical lead. Two generated answers agreeing on a nonexistent requirement still leave it unsupported.

Name the source or person that can settle the consequential detail, and check the corrected final output. Keep that verification responsibility clear even if AI helps with the checking process.

Can one AI-assisted deliverable determine a performance rating?

One deliverable can inform a discussion, but it rarely establishes performance across a role or period by itself. Record what the task required, what was accepted or corrected, the employee's judgment and the conditions under which they worked. Look for comparable evidence over time before interpreting a pattern.

Keep immediate work approval separate from formal appraisal. Use the organization's agreed process and performance-system guidance for useful AI adoption to consider wider responsibilities, support and consistent standards. Neither AI usage frequency nor one polished brief supplies that whole assessment.

Updated

aiready

A home for your company’s AI community.

Share what works and help each other put AI into practice.

  • Real use cases

  • Practical guides

  • Company policies

  • Shared experience

Explore aiready
Explore the blog