// ROLE SUMMARY

You'll evaluate prompt-response pairs generated by large language models, scoring them across dimensions like factual accuracy, instruction-following, coherence, and appropriate refusal behavior. Each session involves reviewing batches of 20–50 pairs using a structured rubric inside a web-based annotation platform, flagging edge cases, and writing brief justifications for non-obvious scores.

AI Prompt Quality Rater

RLHF$35–55/hrRemotePosted August 3, 2026

// DESCRIPTION

You'll evaluate prompt-response pairs generated by large language models, scoring them across dimensions like factual accuracy, instruction-following, coherence, and appropriate refusal behavior. Each session involves reviewing batches of 20–50 pairs using a structured rubric inside a web-based annotation platform, flagging edge cases, and writing brief justifications for non-obvious scores.

Beyond surface-level ratings, you'll identify subtle failure modes — responses that are technically accurate but misleading, answers that follow the letter of a prompt while violating its intent, and outputs that pass a quick read but contain embedded errors. You'll work asynchronously, with a typical batch taking 60–90 minutes, and you're expected to maintain inter-annotator agreement scores above 0.75 kappa.

Feedback you submit feeds directly into preference datasets used to fine-tune and align production models. Calibration sessions run bi-weekly, where you'll review disagreements with a senior evaluator and update your scoring approach. Strong performance can lead to specialized tracks covering domain-specific evaluation (legal, medical, code).

// SKILLS & REQUIREMENTS

Prompt evaluationRubric-based scoringLLM output analysisInstruction-following assessmentFactual accuracy verificationAnnotation platform proficiency

// FREQUENTLY ASKED QUESTIONS

// READY TO GET STARTED?

Apply in minutes

Create your profile, select your areas of expertise, and start working on frontier AI projects.

Apply Now