Domain Expert – History | Contractor | 8-Week Engagement | Worldwide Remote
Turing is seeking a highly qualified History Domain Expert to help evaluate and improve Large Language Models (LLMs). This is a unique opportunity to apply your academic expertise in shaping the next generation of AI systems — working alongside leading AI researchers to enhance model accuracy, reasoning, and depth of historical knowledge.
About Turing
Turing is one of the world's fastest-growing AI companies, partnering with leading AI labs to advance frontier model capabilities across reasoning, coding, STEM, and specialized domain knowledge. We build real-world AI systems that solve mission-critical challenges for organizations worldwide.
Role Overview
As a History Domain Expert, you will design challenging prompts, assess factual accuracy and reasoning quality, and identify knowledge gaps across a wide range of historical disciplines — including ancient, medieval, modern, political, military history, archaeology, and historiography.
Key Responsibilities
- Create advanced, domain-specific prompts of varying difficulty levels across historical subfields.
- Evaluate AI-generated responses for factual accuracy, reasoning quality, completeness, and nuance.
- Identify hallucinations, logical inconsistencies, outdated information, and edge cases.
- Develop benchmark datasets and adversarial test cases to stress-test model performance.
- Provide evidence-based feedback supported by authoritative references.
- Collaborate with AI researchers to drive measurable improvements in model output.
- Maintain high standards of annotation quality and thorough documentation.
Day-to-Day Tasks
- Write domain-specific prompts across a range of difficulty levels.
- Compare multiple AI-generated responses and rank them with justification.
- Explain why a response is correct or incorrect using credible, authoritative sources.
- Identify ambiguous or poorly framed questions and propose improved alternatives.
Minimum Qualifications
- Master's degree or higher in History or a closely related discipline.
- 3+ years of professional, research, or teaching experience in the field.
- Excellent written English and strong analytical skills.
- High attention to detail and commitment to accuracy.
Preferred Qualifications
- Experience with LLMs, Generative AI, prompt engineering, or AI evaluation workflows.
- Published research, academic recognition, or formal teaching experience.
- Ability to work independently and deliver objective, evidence-backed assessments.
Engagement Details
- Commitment: 40 hours per week with at least 4 hours of PST overlap daily.
- Employment Type: Contractor assignment (no medical benefits or paid leave).
- Duration: 8-week contract.