AI Safety & Policy Evaluator at Turing
posted 46 minutes agoAI Safety & Policy Evaluator | Compensation varies | Remote
What you’ll do
Evaluate and help improve safety behavior in large language model systems. This project focuses on single-turn image-edit requests, model outputs, and detailed safety guidelines.
- Classify prompts and outputs using a defined safety taxonomy.
- Apply policy consistently to ambiguous, borderline, and benign cases.
- Write clear, defensible rationales that can be used as training data.
- Document safeguard failures, bypasses, and edge-case violations.
- Where required by the statement of work, create adversarial test prompts and compare or rank model outputs.
- Identify gaps, conflicts, or ambiguities in project guidelines and suggest clarifications.
Requirements
- BA/BS degree or equivalent experience in policy, law, ethics, linguistics, journalism, computer science, or another relevant analytical field.
- Strong analytical judgment when evaluating nuanced information against defined criteria.
- Experience with red teaming, prompt engineering, or designing prompts that test AI safety filters.
- Knowledge of Trust & Safety issues involving LLMs, including misinformation, bias, stereotypes, jailbreaks, dual-use content, and assistance with harmful activity.
- Ability to write precise evaluation rubrics and concise rationales.
- Experience in content moderation, policy analysis, AI safety evaluation, RLHF, or data annotation is preferred.
- Your own computer and a reliable internet connection.
Content advisory
The work involves written prompts concerning sensitive subjects, including minors, self-harm, sexual or suggestive content, non-consensual intimate imagery, violence, hate and harassment, illegal activity, and depictions of identifiable people. Harmful material is encountered as text. Images reviewed are not intended to be egregious, and contractors generally will not view generated harmful imagery. Wellbeing and risk-mitigation resources are available at no cost.
Contract and compensation
- Independent contractor engagement lasting 1–2 weeks per statement of work.
- Project-based fees vary by scope and specialized expertise. Required training time is separately compensated.
- Contractors set their own schedules and working methods.
- No benefits, paid leave, or continuing engagement are offered or implied.
- Contractors provide their own equipment and are responsible for taxes, insurance, and expenses.
- The engagement is non-exclusive, and additional project phases may be offered under separate statements of work.