
SWE-Bench Task Auditor | $70–90/hr | Remote (US)
Join a frontier AI lab as a SWE-Bench Task Auditor, where you'll play a critical role in ensuring the quality, correctness, and reproducibility of software engineering benchmark tasks used to train and evaluate cutting-edge AI models. This is a high-impact contract role for experienced engineers with a strong open-source background.
This is a remote contract opportunity open to candidates based in the United States. Ideal for senior engineers passionate about AI evaluation quality and open-source software.