Assign traces to reviewers
Queues control which traces go to which reviewers. Findings can push flagged traces straight into a queue. From there the work spreads itself. A queue carries the rubrics to grade against and the reviewers on the team, and it hands traces out round-robin, so no one assigns reviews by hand and the load stays even across the team. Leave a queue’s rubric list empty and it grades against all of the project’s active rubrics.Do the review
Each reviewer sees their My Reviews tab: a queue of traces to grade against the rubric. They can submit corrections on different sections, rationales, and other grades. They work in a grading workbench that sits right next to the trace. They score each dimension, set their confidence and severity, add a short rationale, and can edit the model’s answer side by side to capture what the right response should have been. A built-in copilot drafts the correction and the rationale for each grade, so reviewers check and approve instead of writing from scratch.See where reviewers agreed and disagreed
All Reviews shows completed reviews across all reviewers, so you can see what they found across the board. Because more than one expert can grade the same trace, you can also see where they agree and where they don’t, which is the same signal that tells you how solid your ground truth is. Reviewed traces can be routed onward from here too. Select any subset and use Add to dataset to make them an evaluation dataset or a generation seed dataset — so the cases your experts just corrected become the seeds for the next batch of synthetic data.Feedback gets reused automatically
A reviewer’s correction isn’t a one-off. It’s distilled into memory that speeds up future reviews of similar cases, and feeds directly into judge calibration and regression datasets, so the same mistake, once caught, doesn’t need catching twice.Next steps
Judges & Datasets
Calibrate a judge against this review data

