Benture logo
next job
Turing logo

ML & Data Engineer — Data Quality at Turing

posted 1 hour ago
turing.com Contractor remote TBD 40 views

ML & Data Engineer — Data Quality & PII Compliance | Contract (40 hrs/week) | Worldwide Remote

Turing is seeking an experienced ML/Data Engineer to strengthen the quality, reliability, and compliance of enterprise data pipelines. This role focuses on data quality validation and testing PII/PHI de-identification systems — ensuring sanitized data is accurate, consistent, and free from sensitive information leaks. If you thrive at the intersection of data engineering, ML evaluation, and compliance-focused quality assurance, this role is for you.

About Turing

Turing is a leading AI company accelerating the advancement and deployment of frontier AI systems. We partner with the world's top AI labs and enterprises to push capabilities in reasoning, coding, agentic behavior, multimodality, and other advanced AI domains.

What You'll Do

  • Design and automate validation suites for data pipelines, including schema checks, completeness validation, drift detection, and reconciliation across pipeline stages.
  • Perform deep data quality analysis across enterprise data sources and connectors — covering topic coherence, domain coverage, consistency, and depth.
  • Build adversarial test datasets for PII/PHI de-identification systems, covering edge cases, obfuscated identifiers, multilingual entities, OCR noise, and unusual document formats.
  • Evaluate NER and ML-based de-identification systems using precision, recall, F1 score, leak rates, and false-negative analysis.
  • Identify and investigate PII leakage risks across raw, processed, and sanitized data.
  • Implement automated regression gates in CI/CD to prevent pipeline changes from deploying without passing data quality and privacy checks.
  • Conduct sampling-based human-in-the-loop audits and maintain detailed audit trails for compliance evidence.
  • Partner with data and engineering teams to perform root-cause analysis and resolve data inconsistencies, quality issues, and privacy leaks.
  • Develop monitoring and reporting mechanisms for pipeline quality, de-identification performance, and compliance risks.

What We're Looking For

  • 5+ years of experience in data engineering, ML engineering, data quality, or a related field.
  • Strong Python skills for test automation, data validation, and ML evaluation.
  • Experience with testing frameworks and data-quality tools such as pytest, Great Expectations, Pandera, or similar.
  • Strong SQL skills and experience validating data across multiple pipeline stages.
  • Solid understanding of PII and PHI categories and de-identification concepts.
  • Familiarity with privacy and compliance frameworks such as HIPAA Safe Harbor, GDPR, or LGPD.
  • Experience evaluating NER or other ML-based systems using labeled datasets and precision/recall metrics.
  • Ability to design reliable evaluations for non-deterministic ML or LLM-based systems.
  • Experience integrating automated tests and quality checks into CI/CD pipelines.
  • Familiarity with GCP services such as BigQuery, Google Cloud Storage, and Cloud Run Jobs.
  • Strong analytical, debugging, and communication skills.

Engagement Details

  • Commitment: 40 hours per week with 6 hours/day overlap with PST.
  • Duration: 1 month (adjustable based on engagement).
  • Environment: Fully remote.

Why Work With Turing?

  • Fully remote, flexible work environment.
  • Opportunity to contribute to cutting-edge AI projects with leading LLM companies.

Related Jobs

Benture logo
See All Jobs
Apply Back