This job post has expired on September 26, 2026. It is likely that the position has already been filled.
Trainium NKI Kernel Expert at Mercor
posted 1 month agoTrainium NKI Kernel Expert | $70–90/hr | Remote (US)
Join a cutting-edge AI evaluation project where your deep expertise in Neuron Kernel Interface (NKI) development will directly shape the quality of training infrastructure for a frontier AI lab. In this role, you will assess, critique, and provide rubric-based written feedback on NKI kernel development tasks — evaluating CUDA→NKI migration fidelity, Trainium-specific performance optimizations, and cross-platform numerical correctness.
What You'll Do
- Evaluate the quality, correctness, and hardware-appropriateness of NKI development tasks used to train and benchmark frontier AI models
- Assess CUDA→NKI migration fidelity, ensuring idiomatic use of Trainium-native patterns
- Review Trainium-specific performance optimization quality, including NeuronCore pipeline utilization and memory-bandwidth efficiency
- Define and apply cross-platform numerical-correctness standards (GPU vs. Trainium accumulation order, rounding behavior, mixed-precision semantics)
- Deliver clear, structured, rubric-based written feedback to guide task improvement
Basic Qualifications
- 2+ years of hands-on experience developing or optimizing kernels using the Neuron Kernel Interface (NKI) targeting AWS Trainium/Inferentia2 hardware
- Strong understanding of NKI-specific development patterns: tile-based computation, SBUF/PSUM/HBM memory-hierarchy management, partition-dimension constraints, and DMA orchestration
- Demonstrated experience assessing CUDA→NKI migration quality
- Familiarity with Trainium-specific performance profiling tools and metrics
- Experience evaluating cross-platform numerical-correctness standards across GPU and Trainium architectures
Preferred Qualifications
- Direct experience with AWS Neuron SDK, Neuron Compiler internals, or contributions to NKI kernel libraries
- Prior CUDA or Triton kernel development experience
- Familiarity with Trainium hardware specifications, including NeuronCore-v2 architecture and supported data types (FP32/BF16/FP8/INT8)
- Experience benchmarking ML training workloads on Trn1/Trn2 instances
This is a fully remote, hourly contractor engagement open to candidates based in the United States.
How to apply for this role
- Upload your resume — keep it up-to-date and in English. Mercor will auto-fill your profile from it.
- Complete the AI interview — a 15-minute conversation about your experience. Be ready to discuss specific projects and challenges you've solved.
- Submit your application — only about 20% of applicants finish all the steps, so completing yours puts you well ahead.