A request to make an event invitation “more polished” sounds small. Then the AI rewrites the headline, drops the registration detail, changes the audience, and replaces a warm opening you already approved. By version four, every draft is different, but none is clearly closer to done.

The missing ingredient is rarely a cleverer adjective. It is a visible difference to judge and a boundary around everything that should stay put. “Warmer,” “more professional,” and “more premium” leave tone, length, structure, and vocabulary open at once. Each revision resets the comparison.

A lightweight evaluation habit works better than another paragraph of prompting. OpenAI’s evaluation guidance notes that generative output varies even for the same input and recommends structured tests. It also says comparison, classification, and scoring against specific criteria are more reliable evaluation tasks than open-ended generation. You do not need an eval platform to use that lesson on a document: turn taste into a difference another person can point to.

Freeze the baseline before discussing taste

A useful feedback card has three parts: keep, change, boundary.

Keep names what already works. The date, venue, registration link, and opening sentence might remain untouched. Change describes one observable target for this round: cut the middle from 90 words to 45 and replace three abstract claims with one attendee situation. Boundary prevents an apparently elegant rewrite from inventing an offer, altering the fee, or promoting an unverified speaker credential to fact.

This gives the model both a finish line and a regression test. Anthropic’s evaluation guide describes strong success criteria as specific, measurable, achievable, and relevant, and observes that most uses need more than one dimension. For an everyday revision, that does not mean ten weighted scores. Choose one primary difference, then keep facts, permissions, and commitments as hard constraints.

Make one decision per round

Order matters. If the team edits first and reconstructs the baseline later, every review becomes a memory contest.

Save base → Set target → Pick A/B → Apply A/BSave baseSet targetPick A/BApply A/B

Instead of “B feels friendlier,” write: “Choose B because its first two sentences name what an attendee will leave with and stay under 45 words. Keep A’s date sentence. Reject the promised free consultation in both versions because it is not an approved offer.” The next revision now has evidence, not a mood.

Changing one primary dimension at a time is not a claim that other qualities do not matter. It preserves causality. Settle structure before tone; verify facts before rhythm. If “include all background” conflicts with “halve the length,” a person must rank those goals. Asking AI to satisfy an unresolved contradiction only hides the decision inside the next draft.

Pairwise comparison also creates a legitimate “neither” outcome. If A is concise but drops the audience, while B preserves the audience but invents urgency, do not average them. Keep the baseline, record both failures, and tighten the criterion.

Advertisement

Treat disagreement as information

When two reviewers disagree, teams often ask AI to find a compromise. The result may satisfy neither person because the disagreement was never named.

Google’s People + AI Guidebook advises teams not to discard label disagreements automatically as noise. Differences between raters can reveal problems in tools, workflows, instructions, or the data strategy itself. The same diagnostic move helps with ordinary creative work.

One reviewer may own brand voice while another owns legal disclosure. That is not a midpoint on a style scale. Make the disclosure a hard boundary and let the brand owner decide among compliant phrasings. If the disagreement really is about the degree of enthusiasm, each reviewer can select one acceptable and one unacceptable example, describe the visible distinction, and say where it applies.

Keep a small revision ledger: baseline, criterion, A/B decision, rejection reason, and merged version. NIST AI RMF Playbook Measure 2.1 calls for documenting test sets, metrics, tools, and related details to support repeatability and consistency. A two-page invitation does not need enterprise governance paperwork, but it does need enough history to prevent the same argument next week.

Stop when the criterion is met

If the selected revision meets this round’s target without breaking keep items or boundaries, stop. The existence of a fifth candidate is not evidence that the fourth is unfinished.

Pause for a responsible person when sources conflict, brand and compliance requirements have no priority, or the copy creates a price, contract, health, hiring, or public commitment. AI can surface the conflict; it cannot hold the authority behind an approval. Use only approved tools for confidential material and remove customer or employee data that the revision does not require.

A revision ledger also fits the fields in an AI summary someone else can actually take over: the chosen version, decision basis, and remaining limits travel with the deliverable. If the deeper problem is who owns the work after a fast first draft, use the ownership check for AI’s “finished” output alongside this revision loop.

AI handoff card

Start from the documents you are currently authorized to read and remain read-only. Do not overwrite, rename, move, or delete files. Locate the latest draft and its prior version; if the relationship is uncertain, list candidate paths, modification times, and missing evidence, then stop. Compare the versions under intended changes, unresolved issues, and unintended drift, citing short excerpts or locations. Convert existing vague feedback into one card with keep items, one observable change for this round, and non-negotiable fact, permission, and commitment boundaries. Offer two minimal candidates that edit only the named region. Explain how each meets the criterion and what it trades off. Do not choose, merge, or expand scope. Finish with the A, B, or neither decision I must make and a clear stopping condition; wait for confirmation.

Sort the feedback before revising

A designer first sends a pile of similar-looking feedback slips through an AI and gets a polished but entirely altered invitation; she then sorts the slips into confirmed, unresolved, and protected trays, applies only the confirmed slip, and leaves the other two states visible on the desk

  1. Similar-looking feedback slips lie in one pile, mixing confirmed edits, unresolved questions, and protected constraints.
  2. Treating them all as instructions, the designer feeds the pile to the AI and gets a polished invitation with its color, flower, border, and layout all changed.
  3. She retrieves the slips and sorts them into a green confirmed tray, an amber question tray under a magnifying glass, and a red constraint tray beside a lock.
  4. She sends only the green slip back through the AI; the finished invitation changes one ribbon color while the question and constraint trays remain on the desk.

A tailor pins the one seam being altered. A loose cuff is not permission to replace the collar, fabric, and color. Freeze the garment; move one pin.

Advertisement

Share

Share this mini class

If this lesson helps untangle a work bottleneck, share it with someone deciding how to use AI.

References