What Is a Rubric in AI Training Work?
Updated · 2 min read
If you work on AI evaluation for long, you will meet rubrics. A rubric is a set of criteria that defines what a good answer to a particular prompt must contain. It turns a vague judgement into something consistent enough to train a model on.
What a rubric looks like
Imagine a prompt asking for the risks of a common medication. A rubric might say: mentions the two most serious side effects; states the interaction with alcohol; does not recommend a dose; advises consulting a clinician; is under 200 words. Each line can be checked as met or not met.
Using a rubric
When you evaluate with a rubric, you go line by line and record whether each criterion is satisfied, sometimes with a short justification. The discipline is to judge only what the rubric asks, even when another issue catches your eye — then note that separately if the task allows.
Writing a rubric
Experts are increasingly asked to write rubrics themselves, because only a specialist knows what a complete answer requires. Good rubric criteria are:
- Specific — 'names the statute of limitations for the claim' rather than 'legally accurate'.
- Checkable — answerable yes or no by another expert.
- Independent — each criterion tests one thing.
- Weighted — critical errors count more than presentation.
Common mistakes
Criteria that are too vague ('is clear'), criteria that overlap, and rubrics that reward length rather than correctness. Another is writing criteria only an answer phrased one specific way could meet; a correct answer in different words should still pass.
A short rubric, start to finish
Prompt: "Explain how to calculate the area of a triangle when you know all three sides." A rubric might read:
- Names Heron's formula or an equivalent method (critical).
- Defines the semi-perimeter correctly (critical).
- Shows the formula in a correct form (critical).
- Includes a worked numerical example (important).
- Notes that the sides must satisfy the triangle inequality (nice to have).
An answer that misses a critical item fails regardless of how well it is written; one that misses only the last item is still good. That weighting is what lets the rubric separate a correct answer from a polished but wrong one.
Applying a rubric consistently
Two reviewers using the same rubric should reach the same verdict. When a criterion is ambiguous for a particular answer, note the ambiguity in your comment and follow the project's rule for unclear cases. Raising unclear criteria with the project helps everyone; guidelines usually improve after reviewers flag them.
Common questions
How long does it take to write a rubric?
For a complex expert prompt, often 20 to 60 minutes. Projects set expectations in their guidelines.
Do rubrics replace human judgement?
No. They make judgement consistent and explainable, which is what a model needs to learn from it.
Open AI Training roles
17 listings hiring now, each with its pay shown.
