About the role
Help train the world's most advanced AI models. DevFixr is recruiting machine learning engineers for a client that builds training environments for frontier AI labs. You'll write challenging ML engineering coding tasks that are used to train and evaluate cutting-edge AI models. The work covers both CPU and GPU tasks. This is freelance, remote work at around 20 hours a week, paid $30-100 per hour depending on experience. What you'll do- Design realistic, hard ML engineering problems set inside real codebases- Write a working reference solution for each task- Write automated tests that reliably check whether a solution is correct- Make sure every task is clearly specified and can't be passed with shortcuts- Improve tasks based on reviewer feedback What we're looking for- 4+ years of hands-on machine learning engineering in Python- You've trained or fine-tuned models yourself (PyTorch, JAX or similar)- You write thorough tests and are comfortable in large, multi-file codebases- Strong written English, as every task is a written specification- Around 20 hours a week, consistently Nice to have- GPU experience: CUDA, Triton, kernel optimisation, FSDP, DeepSpeed, multi-GPU or multi-node training- Previous AI training or evaluation work (e.g. Turing, Mercor, Outlier, Scale AI)- Experience building benchmarks, evals or RL environments- Kaggle rankings, competitive programming or open-source ML contributions How it works- Freelance contract, paid hourly for the hours you work- Fully remote with flexible hours- A short screening call, then a technical assessment- We never charge candidates any fees Apply with your LinkedIn profile, and include your GitHub or Kaggle link if you have one.
From the employer's public posting. AI Eval HQ isn't affiliated with DevFixr; you apply on their site.
Apply on DevFixr