This job post has expired on December 06, 2025. It is likely that the position has already been filled.

Data Scientist - AI Evaluation & Analysis at Mercor

posted 8 months ago

mercor.com Contractor remote $100/hour 500 views

Data Scientist - AI Evaluation & Analysis | $100–120/hr | Remote Worldwide

We're seeking a data-driven analyst to conduct comprehensive failure analysis on AI agent performance across finance-sector tasks. You'll identify patterns, root causes, and systemic issues in our evaluation framework by analyzing task performance across multiple dimensions.

Key Responsibilities:

Statistical Failure Analysis: Identify patterns in AI agent failures across task components including prompts, rubrics, templates, file types, and tags
Root Cause Analysis: Determine whether failures stem from task design, rubric clarity, file complexity, or agent limitations
Dimension Analysis: Analyze performance variations across finance sub-domains, file types, and task categories
Reporting & Visualization: Create dashboards and reports highlighting failure clusters, edge cases, and improvement opportunities
Quality Framework: Recommend improvements to task design, rubric structure, and evaluation criteria based on statistical findings
Stakeholder Communication: Present insights to data labeling experts and technical teams

Required Qualifications:

Statistical Expertise: Strong foundation in statistical analysis, hypothesis testing, and pattern recognition
Programming: Proficiency in Python (pandas, scipy, matplotlib/seaborn) or R for data analysis
Data Analysis: Experience with exploratory data analysis and creating actionable insights from complex datasets
AI/ML Familiarity: Understanding of LLM evaluation methods and quality metrics
Tools: Comfortable working with Excel, data visualization tools (Tableau/Looker), and SQL

Preferred Qualifications:

Experience with AI/ML model evaluation or quality assurance
Background in finance or willingness to learn finance domain concepts
Experience with multi-dimensional failure analysis
Familiarity with benchmark datasets and evaluation frameworks
2-4 years of relevant experience

Apply on Mercor Go back

Show all jobs of Mercor

How to apply for this role

Upload your resume — keep it up-to-date and in English. Mercor will auto-fill your profile from it.
Complete the AI interview — a 15-minute conversation about your experience. Be ready to discuss specific projects and challenges you've solved.
Submit your application — only about 20% of applicants finish all the steps, so completing yours puts you well ahead.

Benture is an independent job board and is not affiliated with Mercor.

Data Scientist - AI Evaluation & Analysis at Mercor

How to apply for this role

Related Jobs

Mercor

$170/hr remote in US

Mercor

$70/hr remote

Mercor

$180/hr remote in US

Mercor

$150/hr remote in US

Mercor

$120/hr remote in US

Mercor

$50-70/hr remote

Mercor

$210/hr remote

Mercor

$30/hr Remote

Mercor

$38/hr remote

Mercor

$38/hr remote in Germany

Mercor

$38/hr remote

Mercor

$30/hr remote