About the role
About The Role is looking for experienced software engineers to evaluate how AI coding agents handle real engineering work inside realistic codebases: implementing a change to spec, fixing a bug with a regression test, refactoring without breaking behavior, reviewing a pull request, debugging from a trace. Coding agents are strong in many ways and frustratingly bad in others: they miss edge cases, make brittle changes, and need close supervision from an experienced engineer. You bring the judgment you have built over a career shipping production systems. We bring the agent output that judgment is needed to grade. In this role, you will design challenging, realistic tasks drawn from your own practice, such as a bug fix with a regression test, a feature implemented to spec with tests, a refactor with behavior preserved, a pull request review, a debugging write-up from logs and traces, or an API contract with examples, run them through frontier AI agents, and evaluate what comes back against a professional standard. You will work with realistic professional files, the kind a practitioner in your field actually handles, which you assemble yourself. Some tasks are compact, built around a handful of files; others are larger scenarios that take several days to build. In every case the goal is the same: a task a competent professional in your field would complete correctly and a frontier model currently gets wrong. This is not a traditional software engineering role. You will be helping build better AI by putting your knowledge to work in a structured, flexible, fully remote environment. The work is long-form and self-directed, and clear written reasoning matters as much as technical depth. Responsibilities Design challenging, realistic software engineering tasks drawn from your own day-to-day work: the scenario, a spec phrased the way you would brief a trusted teammate, and the codebase or supporting files an engineer would need (repositories you construct or adapt, tests, logs, API contracts, design notes), which you author yourself. Run those tasks through AI coding agents and evaluate the deliverable they produce (the diff, the tests, the design, the review) against the standard you would hold a teammate to. Compare two agent outputs on identical specs and codebases, decide which performed better, and document where each fell short. Write tests, rubrics, and success criteria that specify what a correct deliverable must contain (the right behavior preserved, the right edge cases handled, the right tests added, the right conventions followed), and explain in writing why a submission passes or fails each one. Flag concrete failures with evidence: brittle changes, missed edge cases, tests that pass for the wrong reason, thrashing in the trace, fabricated or ignored files, and off-spec interpretation of the ask. Contribute across your stack and adjacent ones, and review and refine tasks built by other engineers. Qualifications 2+ years of professional software engineering experience preferred, shipping and maintaining production code as part of a team. In progress Bachelor’s degree or higher. Expertise in at least one common stack: TypeScript or JavaScript, Python, Java, Go, Rust, C++, C, Ruby, PHP, Swift, or Kotlin. Specialists are welcome: a strong frontend-only, backend-only, or mobile engineer is a good referral. Comfortable in a terminal with git, tests, and a debugger; able to read an unfamiliar codebase and orient quickly. Docker and CI experience is a plus. Strong judgment about what good code and a real model failure look like, and the ability to explain both in writing. No prior AI or machine learning experience is required. Engineering judgment and attention to detail matter most. Hands-on practitioner: you currently do (or recently did) the work yourself at an individual-contributor level, not solely in a managerial capacity. Full professional or native-level written and spoken English; you can articulate why a result is wrong, not only that it is. General familiarity with AI and LLM tools: you have used models like Claude or ChatGPT in professional work and can tell a well-reasoned answer from a plausible-sounding but incorrect one. Baseline tech literacy: comfortable with cloud file tools (e.g., Google Workspace), managing browser profiles, downloading and installing desktop apps (e.g., Claude), and everyday file handling (e.g., converting between Excel and Google Sheets, zipping files for sharing). A computer science degree is a plus but not required; shipped professional work outweighs credentials.