About the role
Machine Learning Evaluation Specialist (AI Training) About The Role What if your years of research and domain expertise could directly shape how the next generation of AI models think and reason? That's exactly what this role offers. We're looking for ML experts and domain specialists to design fiendishly difficult evaluation problems that push state-of-the-art AI systems to their limits. This isn't about textbook prompts — it's about crafting the kind of challenges that only someone with deep, hard-earned expertise can build and judge. Your work will directly influence how leading AI models are measured, improved, and deployed. Organization: AlignerrType: Hourly ContractLocation: RemoteCommitment: 10–40 hours/week What You'll Do Design complex, original machine learning problems rooted in your specialized domain of expertiseCraft evaluation tasks that go beyond standard ML pipelines and require genuine advanced knowledge to solveDraw from your own research experience to build challenges that would stump a highly capable AIWrite clear, precise problem statements with defined evaluation criteria and gold-standard solutionsAssess AI-generated ML solutions for correctness, creativity, and methodological rigorDocument problem difficulty, required domain knowledge, and expected failure modesCollaborate asynchronously with a global team of researchers and engineers Who You Are You hold graduate-level expertise (MS or PhD preferred) in a scientific or technical domain intersecting with machine learningYou have strong working knowledge of core ML methods — model selection, feature engineering, evaluation metrics, and pipeline designYou're deeply familiar with active, open research problems in your fieldYou can pinpoint exactly where general ML knowledge falls short and specialized domain insight becomes essentialYou have experience conducting or publishing original research (highly valued)You communicate complex ideas clearly and precisely in writingYou're self-motivated and thrive when working independently on intellectually demanding problems Example Domains (Not Exhaustive) We're actively seeking specialists across a wide range of fields, including: Computational biology, genomics, or bioinformaticsClimate science and environmental modelingMedical imaging and healthcare MLMaterials science and computational chemistryAstrophysics and signal processingNLP for low-resource or specialized corporaRobotics, control theory, or reinforcement learningFinancial modeling and quantitative analysis If your domain isn't listed, we still want to hear from you — if your expertise is deep and niche, it's valuable. Why Join Us Work at the frontier — contribute directly to AI evaluation and safety research that mattersMake your expertise count — finally, a role that rewards deep domain knowledge, not just generalist skillsCollaborate globally — work asynchronously alongside top researchers and engineers around the worldFull autonomy — flexible schedule, no micromanagement, work when and how you do your best thinkingOngoing opportunity — strong potential for contract extension and deeper research involvementBuild your profile — be recognized as a contributor to cutting-edge AI development