Use one repeatable review pack to turn a sample of support conversations into decisions a team lead can act on. This is for support quality leads and managers who need reviewers to judge replies against the same tone, policy and resolution standards.
The result is a calibration sheet, not a score dump. It records the evidence, the decision, the coaching needed and any change required to a macro or help-centre article.
Key point
Review against written standards
Give the model your team’s actual policy and tone rules before you provide conversations. Do not ask it to infer standards from a handful of examples.
1. Build one review pack
Create a document called Support QA calibration pack. Use the same headings every time. This gives the model a stable job and gives team leads a record they can compare across review rounds.
Add these sections:
- Review period and queue: for example,
Returns queue, week ending [date]. - Standards: paste the current tone guide, relevant policy excerpts and resolution requirements.
- Decision rules: define what counts as pass, fail and needs-review.
- Conversation set: paste each conversation with a case ID, issue type, customer messages and agent replies in order.
- Known exceptions: list approved policy exceptions, service incidents or temporary process changes.
Keep the standards short enough to be usable. A useful tone standard names observable behaviour, such as “state the next action and owner before closing”, rather than “be helpful”. A useful resolution standard says what must happen, such as “confirm whether a refund was submitted and state the expected next step”.
If you use uploaded files or features that vary by account, check the current xAI documentation before changing your handling process.
Stop
Remove unnecessary customer data
Do not paste names, full addresses, payment details, account credentials or other unnecessary personal data. Replace them with labels such as [customer] and [order reference].
2. Define the pass or fail rules before review
Write three to five decision rules at the top of the pack. This stops a review becoming a general discussion about whether a reply “feels good”.
Use rules like these, then adjust them to match your operation:
| Standard area | Pass when | Fail when |
|---|---|---|
| Tone | The reply is clear, respectful and takes ownership without making unsupported promises. | It is dismissive, overly casual, defensive or promises an outcome not yet confirmed. |
| Policy | The reply follows the supplied policy and does not add conditions that are not in it. | It gives an incorrect entitlement, misses a required disclosure or invents a policy. |
| Resolution | The customer receives a completed action or a clear next step, owner and timeframe where approved. | The reply closes the conversation without resolving the issue or explaining what happens next. |
| Accuracy | Facts, dates, product details and case status match the conversation. | The reply states a fact that the conversation does not support. |
Use needs-review only for genuine uncertainty, such as a policy conflict or missing case evidence. Do not use it to avoid making a decision on a minor wording issue.
Note
Separate severity from the decision
A reply can fail because of one serious policy error, or pass with several small style improvements. Record severity separately so the fail rate does not hide the risk.
3. Ask for a sheet, not a rewritten queue
Paste the standards first, then the conversation set. Ask for one row per case. Require quotations from the agent reply so every finding can be checked.
Use this instruction:
Review the conversations against the supplied tone, policy, resolution and accuracy standards.
For each case, produce a calibration sheet row with:
- case ID and issue type
- pass, fail or needs-review for tone, policy, resolution and accuracy
- overall decision
- severity: low, medium or high
- exact agent wording that supports each failed or needs-review decision
- the standard that applies
- a specific coaching point, written as an action for the agent
- a proposed macro change, only where the issue is repeatable across cases
- a proposed help-centre change, only where the customer lacked information the article should provide
- an escalation question, only where the supplied standards do not settle the decision
Do not invent case facts, policy rules, customer outcomes or internal actions. If evidence is missing, state “not evidenced in supplied conversation”. Return a Markdown table followed by a short list of recurring patterns.
For a large sample, review one issue type at a time. For example, assess ten delivery-delay cases before moving to ten refund cases. This makes recurring macro problems easier to spot and reduces comparison between unrelated policies.
4. Check the evidence before sharing it
Do not send the first output straight to the team lead. Open the original conversation beside the sheet and check every fail and every high-severity item.
Start with these checks:
- Read the quoted agent wording. Confirm it appears in the case and is not missing a qualifying sentence.
- Find the referenced policy rule in your source material. Confirm the rule actually applies to that issue type and date.
- Check the resolution status. A customer may have received an answer later in the conversation, even if an early reply was weak.
- Test the coaching point. It should tell the agent what to write or do differently on the next similar case.
- Check each proposed macro change against at least two cases. One poor reply is usually a coaching issue, not a macro defect.
Check
A good row can be audited quickly
A team lead should be able to move from the decision, to the quoted wording, to the stated standard, without guessing why the reviewer made the call.
The output is going wrong if you see broad comments such as “show more empathy”, policy citations that do not appear in your pack, or identical macro recommendations for unrelated issues. It is also wrong if it treats an unconfirmed delivery, refund or account action as completed. Change the instruction to require exact evidence, then rerun only the affected cases.
5. Turn findings into a team-lead decision record
End each calibration round with a short section titled Team-lead decisions. Do not make the sheet the final authority. It is a structured draft for review.
For each repeated issue, record:
- Finding: for example, agents omit the next action after explaining a delay.
- Evidence: number the affected case IDs.
- Decision: coach individuals, amend a macro, amend a help-centre article, or clarify policy.
- Owner: name the responsible lead or content owner.
- Due date: use your internal review date.
- Follow-up sample: state which queue and issue type you will check next.
Use a simple threshold agreed by your team, such as requiring a repeated pattern before changing shared wording. This prevents a single unusual case from becoming a permanent macro.
6. Repeat the same cycle each week
Keep the same decision rules for a review period. If policy changes, version the standards section, for example Returns policy, internal revision 3, and state the effective date in the pack. Do not compare older conversations against a later rule without marking the difference.
At the next review, include a small follow-up sample for each macro or coaching action. Your question is not merely whether scores improved. Check whether the specific failure has stopped and whether the new wording caused a new problem.
When the process does not work, reduce the scope. Review fewer cases, use one issue type, and supply only the policy excerpts that govern it. If decisions still vary between reviewers, stop changing prompts and hold a human calibration session: agree the decision rules on three real cases, update the written standards, then start the next sheet from those agreed examples.