Software Engineer, Benchmarking
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
About the role
We're looking for a Software Engineer to build and run the platform behind our AI model evaluations. You'll work closely with our research team to prepare benchmark datasets, build the pipelines and environments our evaluations run in, and turn results into published output.
Our researchers design the methodology. You'll turn it into systems that run consistently and reproducibly, so results stay comparable across models, agent scaffolds, and time. Evaluations draw on the knowledge domains our expert network covers.
What you'll do
Prepare and maintain benchmark datasets: Own the data work behind our benchmarks, including cleaning, preparation, conversion into runnable formats, and ongoing maintenance. Validate that tasks are complete, consistent, and executable, and flag ambiguities that would compromise results.
Build and maintain evaluation pipelines: Build the infrastructure that runs evaluations consistently across model APIs and terminal agents, so results are reproducible and comparable.
Build evaluation environments: Create lightweight, containerized environments and viewers for tasking and for evaluating model performance on tool use.
Support model experiments: Help fine-tune small open-source LLMs and compare baseline against post-training performance.
Develop the scoreboard and leaderboard: Build the published views of our results, including model-level, benchmark-level, task-level, domain-level, and rubric-level performance.
Build analysis tools: Make it easy to identify recurring failure modes, compare models and agent scaffolds fairly, and track capability improvements and regressions over time.
Collaborate: Work closely with researchers and engineers to make sure evaluation data and outputs are accurate, consistent, and well integrated into what we publish.
Build tooling for data creation and review: Support expert annotation and data-generation projects by building lightweight HTML viewers and internal tools for task authoring, review, quality control, and structured data collection.
What you bring
Solid engineering skills: 4+ years of professional experience building and maintaining complex systems, with strong Python. You write robust, maintainable code and are comfortable diving deep into existing codebases and infrastructure.
Data rigor: Experience preparing, cleaning, and maintaining datasets, and the care to make sure two results are genuinely comparable.
Comfort with containers and environments: Experience with Docker and building reproducible execution environments.
Collaborative: You work well alongside researchers and scientists and can translate their methodology into working systems.
Hands-on experience running AI evaluations, or with frameworks like Harbor, Terminal-Bench, or Inspect, is a strong plus.
Tech Stack
Evaluations: Python, model APIs, agent/evaluation frameworks, custom evaluation tooling
Models: APIs from the major AI providers, terminal agents, and open-source models via the Hugging Face ecosystem and PyTorch
Environments: Docker
Publishing: React, Next.js, Tailwind
Workflow: GitHub, Slack, Notion, Linear
Bonus Experience
Experience fine-tuning or post-training open-source LLMs, or other hands-on machine learning work
Experience with agentic, multi-turn, long-context, or tool-use evaluation
Experience validating LLM-as-judge or rubric-based grading setups
Background or strong interest in a scientific or technical domain
Experience building data-heavy dashboards, leaderboards, or visualizations
Open-source contributions or published work related to benchmarks and measurement
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Office Hours
Explore Top Companies in this Space
Degreed
Education Technology
Studyportals
Education Technology
Khan Academy
Education Technology
Study
Education Technology
Office Hours
View Company ProfileOffice Hours is a company that embodies the concept of accessibility and approachability in the professional world. The name itself suggests a platform or service where individuals can connect with experts, mentors, or professionals during their 'office hours', fostering a culture of knowledge sharing, guidance, and collaboration. In today's fast-paced and often remote work environment, Office Hours could be seen as a bridge that connects people across different disciplines and locations. The company's mission might revolve around creating a seamless and intuitive platform for users to find, book, and engage in meaningful conversations with specialists in their desired fields. By leveraging technology, Office Hours aims to break down barriers to knowledge and networking, making it easier for anyone to reach out and learn from others during their designated office hours. This service could be particularly valuable for students seeking academic advice, professionals looking for career guidance, or entrepreneurs needing insights from industry veterans. With a strong focus on user experience and a robust network of knowledgeable individuals, Office Hours has the potential to become a go-to resource for people worldwide, democratizing access to expertise and promoting lifelong learning.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.
