Site Reliability Engineer (SRE) at Mercor
posted 1 hour agoSite Reliability Engineer | $200/hr | Remote
Mercor, which connects technical talent with organisations working on technology and AI initiatives, is hiring Infrastructure / Site Reliability Engineers for an engagement building and operating complex, enterprise-grade infrastructure. This is a full-time opportunity, and candidates must be able to commit to a full-time engagement.
What You'll Do
- Build, operate and improve highly available, scalable production infrastructure.
- Manage and optimise Kubernetes-based production environments.
- Design and maintain cloud infrastructure, primarily on AWS.
- Build and maintain observability across infrastructure and applications using Datadog or similar platforms.
- Investigate production incidents, perform root-cause analysis and implement durable fixes.
- Improve monitoring, alerting, logging, tracing and overall production visibility.
- Develop automation and internal tooling to reduce manual operational work.
- Partner with software engineering teams on deployments, infrastructure and production reliability.
- Contribute to infrastructure architecture decisions for complex distributed systems.
Requirements
- Professional experience in Infrastructure Engineering, SRE, Platform Engineering, DevOps or Production Engineering.
- Hands-on experience operating complex, enterprise-grade production systems.
- Strong production experience with Kubernetes.
- Strong experience with AWS and cloud-native infrastructure.
- Experience with Datadog, Prometheus, Grafana or comparable observability platforms.
- Experience with Infrastructure as Code (Terraform, Pulumi or equivalent).
- Strong understanding of distributed systems, networking, containers, Linux and cloud architecture.
- Experience building or maintaining CI/CD and production deployment infrastructure.
- Strong debugging, troubleshooting and incident-response skills.
- Proficiency in at least one programming or scripting language, such as Python, Go or Bash.
Strong Signals
- Operating Kubernetes and cloud infrastructure at significant production scale.
- Supporting high-traffic or mission-critical applications.
- Building infrastructure or platform tooling used by large engineering organisations.
- Ownership of production reliability, on-call operations, incident response or capacity planning.
- Demonstrated improvements to SLOs/SLIs, observability, deployment reliability or operational efficiency.
Pay and Schedule
$200/hr. Full-time engagement required; candidates must be able to commit full-time.
Location
Fully remote.
How to apply for this role
- Upload your resume — keep it up-to-date and in English. Mercor will auto-fill your profile from it.
- Complete the AI interview — a 15-minute conversation about your experience. Be ready to discuss specific projects and challenges you've solved.
- Submit your application — only about 20% of applicants finish all the steps, so completing yours puts you well ahead.