Back to Jobs
Data Science & Analytics Just now

AI Coding Agent Evaluator

PolandPoland
Contract
Up to $40/hr equivalent
Senior-Level

Job Description

Key Skills Required

Master these to land this role

ReactHighly Demanded
Learn in 26 Hours
PythonBestseller 🔥
Learn in 56 Hours
Data AnalystBestseller 🔥
Learn in 88 Hours
Data ScientistPostgresFastAPIKafkaRedisBusiness IntelligenceDocker

Want to know if you're a match for this job?

Calculate My Match Score

We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.

You'll create challenging tasks and evaluation criteria within realistic simulated environments:

  • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust

Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

How would you rate this job post?

See what other professionals think about this role.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More