Benture logo
next job
Turing logo

Senior AI QA Engineer at Turing

posted 2 hours ago
turing.com Full Time Bengaluru, India TBD 45 views

Senior AI QA Engineer | Full Time | Bengaluru, India (Hybrid) | BFSI Domain

Turing is seeking an experienced Senior AI Quality Assurance Engineer to own the quality, reliability, and production readiness of enterprise-grade AI applications. This is a high-impact role focused on evaluating LLM applications, RAG systems, and AI agents within the Banking, Financial Services & Insurance (BFSI) sector.

About the Role

This role goes well beyond traditional software testing. You will design benchmark datasets, build automated evaluation pipelines, and establish measurable quality standards across AI workflows. Collaborating closely with AI Engineers, Product Managers, and Platform teams, you will ensure every release meets enterprise-grade expectations for accuracy, reliability, performance, and scalability.

Key Responsibilities

  • AI Evaluation & Benchmarking: Design golden datasets and automated evaluation pipelines for LLM applications. Define quality gates and measure hallucination rate, tool selection accuracy, execution accuracy, precision, recall, latency, and cost.
  • RAG & Agent Validation: Validate Retrieval-Augmented Generation pipelines and AI agent workflows, covering retrieval quality, context relevance, tool invocation, reasoning flow, memory, and end-to-end task completion.
  • Python Automation & API Testing: Build and maintain scalable automation frameworks using Python and Pytest for unit, integration, API, regression, and end-to-end testing. Integrate quality checks into CI/CD pipelines.
  • Frontend Automation: Develop automated UI test suites using Playwright to validate AI-powered user journeys, conversational interfaces, and workflow execution across releases.
  • Observability & Root Cause Analysis: Leverage OpenTelemetry and AI observability tools to analyse execution traces, latency, model responses, and workflow behaviour. Identify regressions, hallucinations, and production issues.
  • Performance & Enterprise Readiness: Monitor latency, throughput, reliability, and scalability to ensure production-ready AI applications that meet enterprise quality standards.

Required Skills

  • 7+ years of experience in Software QA, Test Automation, or AI Quality Engineering
  • Strong Python programming skills
  • Hands-on experience with Pytest for unit, integration, API, and regression testing
  • Experience testing REST APIs and backend services
  • Hands-on experience evaluating LLM-powered applications and RAG systems
  • Understanding of AI agent evaluation methodologies
  • Proficiency measuring AI quality metrics: Hallucination Rate, Tool Selection Accuracy, Execution Accuracy, Precision/Recall, Latency (P50/P95/P99), Token Usage
  • Experience with Playwright or similar frontend automation frameworks
  • Experience with OpenTelemetry, tracing, or observability platforms
  • Strong debugging and root cause analysis skills
  • Experience integrating automated tests into CI/CD pipelines

Nice to Have

  • Experience with LangSmith, Langfuse, MLflow, or Arize Phoenix
  • Experience evaluating multi-agent systems (LangGraph, CrewAI, Google ADK, AutoGen)
  • Familiarity with vector databases such as Pinecone, Milvus, Weaviate, or pgvector
  • Exposure to OpenAI, Anthropic, Gemini, or Azure OpenAI
  • Experience with performance and load testing tools
  • Prior experience in BFSI

Education

BE / BTech / MCA / MTech / BSc in Computer Science, Artificial Intelligence, Data Science, Information Technology, or a related field.

What Success Looks Like

You will build a scalable AI quality engineering framework through automated testing, benchmark-driven evaluations, RAG and agent validation, and comprehensive observability — ensuring AI applications consistently deliver accurate, reliable, and production-ready experiences for enterprise users.

Related Jobs

Benture logo
See All Jobs
Apply Back