About the role
We're looking for experienced software engineers to build one AI coding evaluation task each, from a codebase they know well. You pick a real change in a repository you know from the inside, one that a leading AI coding model gets wrong. You write the fix, the tests that prove it, and a short grading rubric. The task is used to measure and train AI coding models. Remote, any location. 15 to 20 hours. Due Tuesday 29 September. $80 to $100 per hour (USD), up to 20 hours. What you'll deliverA one-page proposal, approved by us before you buildThe repo building and running its tests offline in Docker, at a pinned commitYour own fix, written to the standard a maintainer would mergeTests that reject a nearly-right fix3 to 5 review criteria for style, scope and design RequirementsIdeally 3+ years of professional software engineeringCode merged into a public open-source repository, ideally one you know well. A private codebase you own also works (not employer or client code)Strong in at least one of Python, TypeScript, Go, Rust, Java or C++Solid Docker, Linux, Git and testing skillsRegular use of an AI coding toolExperience writing tests, rubrics or evaluation criteria for AI modelsA secure personal or work computer (no public or shared machines)Available now for 15 to 20 hours Nice to have- Deep expertise in one area, such as databases, compilers, distributed systems, security or infrastructure- Experience building agent tasks or RL environments Terms- Paid only for a task that passes our quality review- You sign Askable's contract and accept its terms- The work is used for AI training. The change and your solution must be your own, and you must be free to license them to us- Don't publish the work anywhere: no public pull request or public fork Apply at https://www.askable.com/experts/roles/senior-software-engineer
From the employer's public posting. AI Eval HQ isn't affiliated with Askable; you apply on their site.
Apply on Askable