A completed spreadsheet can hide a difficult problem: two reviewers may interpret the same request differently. Before you outsource annotation, ask to see how the team handles that disagreement.

Our synthetic data example contains a request that could describe either a payment problem or a page-loading failure. The record stays unresolved because the primary intent depends on which problem happened first. A useful delivery includes that question and preserves the original text.

The demonstration shows a reporting format. It doesn’t establish the quality of an annotation workforce or the accuracy you’ll get on your own data. Your evaluation needs a defined task, representative examples, and an agreed review method.

Start with the label boundaries

Ask for a written rubric that explains each label and the boundaries between similar labels. Include clear examples, ambiguous examples, and instructions for missing context. Name the person who can settle a policy question.

“Billing” and “account access” can both appear in a single support request. Decide whether the task allows multiple labels, requires a primary intent, or permits an unresolved state. An annotation team shouldn’t have to invent that policy while working through the batch.

Version the rubric. When a rule changes, retain enough information to identify which records were labeled under the earlier version. Agree whether those records need another review and include that work in the plan.

Separate completion from quality

Ask the provider to distinguish records processed, records reviewed, and records accepted. Those counts answer different questions. A row that passes a schema check may still contain the wrong label.

HumanSignal’s review documentation describes accepting annotations, correcting and accepting them, or rejecting them. It also supports examining disagreement to prioritize review. Those distinctions are useful when specifying the evidence you want in a delivery. Read HumanSignal’s annotation review documentation.

Ask how reference labels are established and who approves them. Agreement between reviewers shows consistency under a chosen measure. It doesn’t, on its own, establish that the shared answer is correct. Keep schema validity, reviewer agreement, and accuracy against an approved reference separate in reports.

Calibrate on representative work

Before increasing volume, run a small batch that contains the cases your team actually expects to see. Include ordinary records, confusing boundaries, and examples with insufficient context. Choose the batch with the review owner so it doesn’t quietly exclude the hard cases.

Review disagreements together and update the instructions. If an example still needs a product or domain decision, assign that decision to an owner. Record the final rule and why it was chosen, then apply it consistently to affected records.

Agree the sample size and acceptance criteria for your task. A short demonstration can reveal obvious instruction problems. It can’t support a universal accuracy claim across different languages, domains, or input conditions.

Specify what comes back with the labels

A delivery should be inspectable without a meeting. Request stable record identifiers, the final labels, the rubric version, and the review state. Include reasons for material corrections and a separate list of exceptions that still need input.

  • Original records: preserve the input so a reviewer can reconstruct the decision.
  • Review evidence: describe what was checked, by whom or by which method, and what the result means.
  • Unresolved items: identify the missing information and the person responsible for resolving it.
  • Export checks: validate required fields, permitted labels, and identifiers before importing the data into your workflow.

Clarify the use of automated labeling and AI assistance. Identify which stages use automation and which involve independent human review. If a batch has only automated structural checks, the report should say so.

Use the first batch to assess the workflow

Track how many questions return to your team and whether the same question repeats. Repeated ambiguity may point to a gap in the rubric. Repeated errors against a clear rule call for a different response.

Also record the time your reviewer spends accepting the output and the effort required to correct it. A fast initial batch that needs extensive revision can delay the usable result.

Bring a representative sample and your existing instructions to the first discussion. The checklist below helps turn that material into an agreed review plan before you commit to more volume.

PUT IT INTO PRACTICE

A worksheet for your next discussion.

Download an editable text file. No email required.

Download the worksheet ↓