Back to Jobs
Office Hours
AI & Machine Learning 1h ago

Software Engineer, Benchmarking

Office Hours
United StatesUnited States
Full-time
$160,000-$210,000
Mid-Level

Job Description

Key Skills Required

Master these to land this role

React57mFree Trial ✨
Start 10-Day Free Trial
Python2h 41mFree Trial ✨
Start 10-Day Free Trial
Machine Learning41mFree Trial ✨
Start 10-Day Free Trial
DockerAI Engineer

Want to know if you're a match for this job?

Calculate My Match Score

About the role

We're looking for a Software Engineer to build and run the platform behind our AI model evaluations. You'll work closely with our research team to prepare benchmark datasets, build the pipelines and environments our evaluations run in, and turn results into published output.

Our researchers design the methodology. You'll turn it into systems that run consistently and reproducibly, so results stay comparable across models, agent scaffolds, and time. Evaluations draw on the knowledge domains our expert network covers.

What you'll do

  • Prepare and maintain benchmark datasets: Own the data work behind our benchmarks, including cleaning, preparation, conversion into runnable formats, and ongoing maintenance. Validate that tasks are complete, consistent, and executable, and flag ambiguities that would compromise results.

  • Build and maintain evaluation pipelines: Build the infrastructure that runs evaluations consistently across model APIs and terminal agents, so results are reproducible and comparable.

  • Build evaluation environments: Create lightweight, containerized environments and viewers for tasking and for evaluating model performance on tool use.

  • Support model experiments: Help fine-tune small open-source LLMs and compare baseline against post-training performance.

  • Develop the scoreboard and leaderboard: Build the published views of our results, including model-level, benchmark-level, task-level, domain-level, and rubric-level performance.

  • Build analysis tools: Make it easy to identify recurring failure modes, compare models and agent scaffolds fairly, and track capability improvements and regressions over time.

  • Collaborate: Work closely with researchers and engineers to make sure evaluation data and outputs are accurate, consistent, and well integrated into what we publish.

  • Build tooling for data creation and review: Support expert annotation and data-generation projects by building lightweight HTML viewers and internal tools for task authoring, review, quality control, and structured data collection.

What you bring

  • Solid engineering skills: 4+ years of professional experience building and maintaining complex systems, with strong Python. You write robust, maintainable code and are comfortable diving deep into existing codebases and infrastructure.

  • Data rigor: Experience preparing, cleaning, and maintaining datasets, and the care to make sure two results are genuinely comparable.

  • Comfort with containers and environments: Experience with Docker and building reproducible execution environments.

  • Collaborative: You work well alongside researchers and scientists and can translate their methodology into working systems.

Hands-on experience running AI evaluations, or with frameworks like Harbor, Terminal-Bench, or Inspect, is a strong plus.

Tech Stack

  • Evaluations: Python, model APIs, agent/evaluation frameworks, custom evaluation tooling

  • Models: APIs from the major AI providers, terminal agents, and open-source models via the Hugging Face ecosystem and PyTorch

  • Environments: Docker

  • Publishing: React, Next.js, Tailwind

  • Workflow: GitHub, Slack, Notion, Linear

Bonus Experience

  • Experience fine-tuning or post-training open-source LLMs, or other hands-on machine learning work

  • Experience with agentic, multi-turn, long-context, or tool-use evaluation

  • Experience validating LLM-as-judge or rubric-based grading setups

  • Background or strong interest in a scientific or technical domain

  • Experience building data-heavy dashboards, leaderboards, or visualizations

  • Open-source contributions or published work related to benchmarks and measurement

How would you rate this job post?

See what other professionals think about this role.

banner

Office Hours is a company that embodies the concept of accessibility and approachability in the professional world. The name itself suggests a platform or service where individuals can connect with experts, mentors, or professionals during their 'office hours', fostering a culture of knowledge sharing, guidance, and collaboration. In today's fast-paced and often remote work environment, Office Hours could be seen as a bridge that connects people across different disciplines and locations. The company's mission might revolve around creating a seamless and intuitive platform for users to find, book, and engage in meaningful conversations with specialists in their desired fields. By leveraging technology, Office Hours aims to break down barriers to knowledge and networking, making it easier for anyone to reach out and learn from others during their designated office hours. This service could be particularly valuable for students seeking academic advice, professionals looking for career guidance, or entrepreneurs needing insights from industry veterans. With a strong focus on user experience and a robust network of knowledgeable individuals, Office Hours has the potential to become a go-to resource for people worldwide, democratizing access to expertise and promoting lifelong learning.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More