// ROLE SUMMARY
You'll review AI-generated code across the full web stack — JavaScript, TypeScript, Java backend services, and Node. js APIs — assessing correctness, security, maintainability, and adherence to the task specification.
Full-Stack Code Output Evaluator
// DESCRIPTION
You'll review AI-generated code across the full web stack — JavaScript, TypeScript, Java backend services, and Node.js APIs — assessing correctness, security, maintainability, and adherence to the task specification. For each task, you receive a prompt that was given to a code-generating model and the code it produced. You run the code (in a provided sandbox or your local environment), verify it works as specified, and score it on multiple rubric dimensions. If the code is broken or insecure, you document exactly why, with line-level references.
Workflows include side-by-side comparison tasks (picking the better of two model outputs with written justification), single-output correctness reviews (does this code do what the prompt asked?), and security-flag tasks (does this output introduce SQL injection, XSS, or insecure dependency patterns?). You'll also encounter backend architecture tasks where you evaluate whether a Node.js or Java service is structured sensibly for the stated use case — not just whether it runs, but whether a senior engineer would accept it in a real PR. Some tasks use Eclipse-based Java projects; comfort navigating that toolchain is expected.
You should have at least 4 years of professional full-stack development experience, with production-level fluency in JavaScript/TypeScript and Java. You don't need to write new features — your job is to read and judge code quickly and accurately. Annotators who've done code review as part of a senior or lead engineering role tend to calibrate well for this work.
// SKILLS & REQUIREMENTS
// FREQUENTLY ASKED QUESTIONS
// READY TO GET STARTED?
Apply in minutes
Create your profile, select your areas of expertise, and start working on frontier AI projects.
Apply Now