Benture logo
next job
Turing logo

DevOps & Cloud Infrastructure Engineer at Turing

posted 54 minutes ago
turing.com Contractor remote TBD 31 views

DevOps & Cloud Infrastructure Engineer | Contract (40 hrs/week) | Fully Remote

Turing is a leading AI company working with top AI labs and enterprises to advance frontier AI systems. We are seeking a hands-on DevOps and Cloud Infrastructure Engineer to build and operate the infrastructure behind Lazarus — a large-scale platform for PII detection, redaction, and human review. This is a 1-month contract (adjustable) requiring 40 hours/week with 6 hours/day overlap with PST.

What You'll Do

  • Design, provision, and operate GCP infrastructure for PII detection and redaction pipelines.
  • Build and manage workloads across Cloud Run, Cloud Run Jobs, Compute Engine, GKE, and GPU-backed infrastructure.
  • Design scalable batch-processing systems capable of handling hundreds of thousands of files.
  • Implement worker parallelism, queues, retries, checkpointing, idempotency, timeouts, and failure recovery.
  • Manage data movement across Google Cloud Storage, Amazon S3, VMs, containers, and external storage systems.
  • Secure sensitive datasets using IAM, service accounts, Secret Manager, private networking, controlled egress, IAP, encryption, and audit logging.
  • Deploy and operate containerized Python and ML workloads using Docker.
  • Support PII and ML services such as Google Sensitive Data Protection, Presidio, OCR/vision systems, NER models, and LLM-based validation pipelines.
  • Provision and manage GPU infrastructure including drivers, CUDA, quotas, autoscaling, and model-serving environments.
  • Support model-serving stacks such as vLLM, Hugging Face Transformers, and Triton.
  • Build and maintain CI/CD pipelines for Cloud Run, VMs, containers, and related services.
  • Establish observability through centralized logging, metrics, alerting, job tracking, and infrastructure dashboards.
  • Optimize throughput and cost through infrastructure selection, concurrency tuning, autoscaling, and API rate-limit management.
  • Develop backup, recovery, migration, and disaster-recovery processes for large datasets and cloud resources.
  • Troubleshoot production issues involving networking, storage, IAM, containers, APIs, compute, and ML infrastructure.

What We're Looking For

  • 5+ years of experience in DevOps, cloud infrastructure, platform engineering, or a related field.
  • Strong hands-on GCP experience including Cloud Run, Compute Engine, GKE, Cloud Storage, IAM, Secret Manager, VPC, Identity-Aware Proxy, Artifact Registry, and Cloud Logging/Monitoring.
  • Working experience with AWS services, particularly Amazon S3 and cross-cloud data movement.
  • Strong Linux administration and shell-scripting skills.
  • Proficient Python skills for infrastructure automation, operational tooling, and data-processing workflows.
  • Strong Docker and containerization experience.
  • Experience designing and operating large-scale batch or data-processing pipelines.
  • Solid understanding of queues, worker parallelism, retries, checkpoints, idempotency, and failure recovery.
  • Experience securely processing and transferring sensitive or regulated data.
  • Experience with CI/CD tools and production deployment workflows.
  • Exposure to ML infrastructure, GPU workloads, OCR, NLP, LLM serving, or model inference systems.
  • Strong problem-solving skills with the ability to independently investigate and resolve production issues.

Perks of Freelancing With Turing

  • Fully remote work environment.
  • Opportunity to contribute to cutting-edge AI projects with leading LLM companies.

Related Jobs

Benture logo
See All Jobs
Apply Back