Sr. AI QA Engineer | Full-Time | Hybrid – Bengaluru, Hyderabad, Mumbai, or Gurugram | BFSI Domain
Turing is seeking an experienced Senior AI Quality Assurance Engineer to own the quality, reliability, and production readiness of enterprise-grade AI applications. This is a high-impact role focused on evaluating LLM applications, RAG systems, and AI agents — going well beyond traditional software testing.
About the Role
You will establish measurable quality standards, build automated evaluation pipelines, and ensure every release meets enterprise expectations for accuracy, reliability, performance, and scalability. You'll collaborate closely with AI Engineers, Product Managers, and Platform teams across the BFSI domain.
Key Responsibilities
- AI Evaluation & Benchmarking: Design golden datasets and automated evaluation pipelines for LLM applications. Define quality gates and measure hallucination rate, tool selection accuracy, execution accuracy, precision, recall, latency, and cost.
- RAG & Agent Quality Validation: Validate Retrieval-Augmented Generation pipelines and AI agent workflows, covering retrieval quality, context relevance, tool invocation, reasoning flow, memory, and end-to-end task completion.
- Python Automation & API Testing: Build and maintain scalable automation frameworks using Python and Pytest for unit, integration, API, regression, and end-to-end testing. Integrate quality checks into CI/CD pipelines.
- Frontend Automation: Develop automated UI test suites using Playwright to validate AI-powered user journeys, conversational interfaces, and workflow execution.
- Observability & Root Cause Analysis: Leverage OpenTelemetry and AI observability tools to analyze execution traces, latency, model responses, and API calls. Identify regressions, hallucinations, and production issues.
- Performance & Enterprise Readiness: Monitor latency, throughput, reliability, and scalability to ensure production-ready AI applications meeting enterprise quality standards.
Required Skills
- 7+ years in Software QA, Test Automation, or AI Quality Engineering
- Strong Python programming skills
- Hands-on experience with Pytest (unit, integration, API, regression testing)
- Experience testing REST APIs and backend services
- Proven experience evaluating LLM-powered applications and RAG systems
- Familiarity with AI quality metrics: Hallucination Rate, Tool Selection Accuracy, Execution Accuracy, Precision/Recall, Latency (P50/P95/P99), Token Usage
- Experience with Playwright or similar frontend automation frameworks
- Experience with OpenTelemetry, tracing, or observability platforms
- Strong debugging and root cause analysis skills
- CI/CD pipeline integration experience
Nice to Have
- Experience with LangSmith, Langfuse, MLflow, or Arize Phoenix
- Familiarity with multi-agent frameworks: LangGraph, CrewAI, Google ADK, or AutoGen
- Experience with vector databases (Pinecone, Milvus, Weaviate, pgvector)
- Exposure to OpenAI, Anthropic, Gemini, or Azure OpenAI
- Performance and load testing experience
- Prior BFSI domain experience
Education
BE / BTech / MCA / MTech / BSc in Computer Science, Artificial Intelligence, Data Science, Information Technology, or a related field.
What Success Looks Like
You will build a scalable AI quality engineering framework — spanning automated testing, benchmark-driven evaluations, RAG and agent validation, frontend and API automation, and comprehensive observability — ensuring AI applications consistently deliver accurate, reliable, and production-ready experiences for enterprise users.