// ROLE SUMMARY
You'll systematically probe large language models to find inputs that cause them to bypass safety guidelines and produce harmful, disallowed, or policy-violating outputs. Your work involves crafting adversarial prompts across defined harm categories — including but not limited to dangerous instructions, deceptive content, and identity manipulation — then documenting exactly what worked, what the model produced, and why the response constitutes a policy violation.
LLM Jailbreak Red Teamer
// DESCRIPTION
You'll systematically probe large language models to find inputs that cause them to bypass safety guidelines and produce harmful, disallowed, or policy-violating outputs. Your work involves crafting adversarial prompts across defined harm categories — including but not limited to dangerous instructions, deceptive content, and identity manipulation — then documenting exactly what worked, what the model produced, and why the response constitutes a policy violation. You won't be testing randomly; each session is scoped to specific harm categories and model behaviors as directed by the research team.
The workflow is structured: you'll receive a target model, a harm taxonomy, and a session goal. You'll then attempt multiple attack vectors per category — direct prompting, role-play framing, multi-turn escalation, encoding tricks, and others — logging each attempt with the full prompt, model response, and a severity rating. Successful jailbreaks get promoted to a verified findings database; failed attempts are still valuable and should be logged with notes on why they failed. Sessions typically run 2–3 hours with a defined scope.
The strongest candidates are people who think like adversaries — researchers, security professionals, former CTF competitors, or anyone who has spent time stress-testing systems. You should be methodical enough to document findings reproducibly and creative enough to try non-obvious attack angles. Familiarity with existing jailbreak literature (DAN prompts, prompt injection, etc.) is assumed. You won't be writing exploit code, but you will be thinking in the same mode.
// SKILLS & REQUIREMENTS
// FREQUENTLY ASKED QUESTIONS
// READY TO GET STARTED?
Apply in minutes
Create your profile, select your areas of expertise, and start working on frontier AI projects.
Apply Now