// ROLE SUMMARY
You'll evaluate prompt-response pairs generated by large language models, scoring them across dimensions like factual accuracy, instruction-following, coherence, and appropriate refusal behavior. Each session involves reviewing batches of 20–50 pairs using a structured rubric inside a web-based annotation platform, flagging edge cases, and writing brief justifications for non-obvious scores.
AI Prompt Quality Rater
// DESCRIPTION
You'll evaluate prompt-response pairs generated by large language models, scoring them across dimensions like factual accuracy, instruction-following, coherence, and appropriate refusal behavior. Each session involves reviewing batches of 20–50 pairs using a structured rubric inside a web-based annotation platform, flagging edge cases, and writing brief justifications for non-obvious scores.
Beyond surface-level ratings, you'll identify subtle failure modes — responses that are technically accurate but misleading, answers that follow the letter of a prompt while violating its intent, and outputs that pass a quick read but contain embedded errors. You'll work asynchronously, with a typical batch taking 60–90 minutes, and you're expected to maintain inter-annotator agreement scores above 0.75 kappa.
Feedback you submit feeds directly into preference datasets used to fine-tune and align production models. Calibration sessions run bi-weekly, where you'll review disagreements with a senior evaluator and update your scoring approach. Strong performance can lead to specialized tracks covering domain-specific evaluation (legal, medical, code).
// SKILLS & REQUIREMENTS
// FREQUENTLY ASKED QUESTIONS
// READY TO GET STARTED?
Apply in minutes
Create your profile, select your areas of expertise, and start working on frontier AI projects.
Apply Now