Back to Jobs
AI & Machine Learning 4h ago

Staff AI Scientist

United StatesUnited States
Full-time
$199,000 - $267,950
Senior-Level

Job Description

Key Skills Required

Master these to land this role

Machine LearningBestseller 🔥
Learn in 42 Hours
BackendBestseller 🔥
Learn in 18 Hours
Python ScriptingNLPAI Engineer

Want to know if you're a match for this job?

Calculate My Match Score

About the Role

The Health Intelligence team is at the forefront of integrating modern AI and LLMs into the Oura experience, transforming how members interact with and learn from their data. We are building a next-generation AI-powered platform at the intersection of classical ML and modern GenAI. The serving layer increasingly runs through LLMs, which translates insights from traditional ML into contextually relevant, safe, and personalized insights. Bridging the gap and owning the pipeline of classical ML, backend engineering, and GenAI is one of the defining technical challenges of this role.

As a Staff AI Scientist, you will own the end to end development of critical P1 Health Intelligence initiatives within Oura. You will be hands-on in building, deploying, and iterating on production systems, and you will hold a high bar for the velocity at which the team moves from hypothesis to live experiment to learning. You will work across the full stack of the product development lifecycle — from ideation, research, data engineering, and pipeline generation to Backend API contracts, LLM configuration, fine tuning, retrieval, and evaluation. You will be part of a bespoke versatile high impact team that is the connective tissue between engineering, product, and design. This is a high-visibility role for someone who thinks in systems, ships with urgency, and wants to build something that compounds in value over weeks and months.

This is a US Remote role.

What You Will Do

  • Own end to end development of core intelligent capabilities: Research, build, evaluate, and ship reusable intelligence systems—from scientific signal through product integration, launch, and iteration.
  • Define improvements in personalization tech strategy: Set the agenda for how Oura represents users and delivers relevant content across surfaces. Influence roadmap and technical direction across partner teams.
  • Drive evaluation rigor: Design measurement frameworks that assess the full intelligent Advisor experience. Understanding evaluation only matters if it moves fast enough to inform the next decision — you will build lightweight offline evals and shadow-mode testing infrastructure that let the team iterate quickly without waiting for long A/B cycles. Establish rubrics and tooling others can use and reuse.
  • Support causal and counterfactual model development: Support the causal and counterfactual reasoning necessary to distinguish outcome effects from confounding variables. Design and analyze experiments that measure genuine impact on behavior and health, not just engagement.
  • Mentor and raise the bar: As a Staff scientist, you are expected to grow the people around you by providing technical mentorship to scientists and engineers — shaping team norms around experimentation and evaluation, and helping define what good looks like for personalization science at Oura.
  • Collaborate and communicate across functions: Partner with engineering, science, product, and design across the Health Intelligence team to shape how personalization integrates into the broader member experience. Communicate trade-offs, uncertainty, and modeling assumptions clearly to technical and non-technical stakeholders across the US and EU.

Requirements

We’d love to hear from you if you have:

  • 8+ years of experience in applied AI, AI research, and backend engineering. A graduate degree (MS or PhD) in a relevant quantitative field such as Computer Science, Statistics, or a related discipline is strongly preferred.
  • Deep experience with AI / LLM-backed products and evaluation workflows, such as LLM-as-judge, rubric-based evaluation, safety/red-teaming, and offline vs. online assessment of model quality, latency, and cost. And a track record of shipping these into real production systems in a robust experimentation framework, not just offline analyses or research prototypes.
  • Deep experience with Backend engineering best practices and demonstrated ability to build and own systems that serve millions of users.
  • Hands-on experience across retrieval, ranking, and recommendation system design (including collaborative filtering, embedding-based approaches, graph networks, or related methods), and a track record of shipping these into real production systems in a robust experimentation framework, not just offline analyses or research prototypes.
  • Comfort working closely with server and app engineers on model serving, pipeline architecture, and deployment infrastructure — and an instinct for where to cut scope to ship faster.
  • Practical experience integrating recommendation or retrieval signals with LLM-powered generation, including work on grounding, constrained decoding, prompt design, or evaluation frameworks that assess the efficacy of the generation layer.
  • Demonstrated ability to design lightweight experiments and evaluations that generate signal quickly, such as shadow testing, staged rollouts, and proxy metrics that responsibly accelerate the learning loop without waiting on long A/B cycles.
  • Experience framing personalization problems, modeling user trajectories, and working with stateful or sequential data.
  • Solid exposure to causal methods (uplift modeling, treatment effect estimation, counterfactual evaluation) and experiment design, with the ability to interpret results with appropriate caution and communicate uncertainty clearly.
  • Evidence of operating beyond individual contributions: influencing technical direction, mentoring others, shaping team practices, or leading cross-functional scientific initiatives.
  • Strong ability to explain complex systems, trade-offs, and uncertainty to both technical and non-technical audiences, and to operate effectively in a fast-moving, ambiguous domain.
  • Strong proficiency in Python, including data analysis and modeling, as well as experience with modern data tooling in collaboration with data and engineering partners.

Would be a benefit

These are strong signals of fit:

  • Experience designing personalization specifically for consumer behavior change or health outcomes, where the goal is engagement and longitudinal impact.
  • Familiarity with health, wearables, or digital therapeutics domains, and genuine interest in how personalization compounds over a member's lifetime.
  • Comfort working outside “core” hours across time zones with distributed, cross-functional teams.

What we offer

  • Competitive salary and equity packages
  • Health, dental, vision insurance, and mental health resources
  • An Oura Ring of your own plus employee discounts for friends & family
  • 20 days of paid time off plus 13 paid holidays plus 8 days of flexible wellness time off
  • Paid sick leave and parental leave

Oura takes a market-based approach to pay, which may vary depending on your location. US locations are categorized into tiers based on a cost of labor index for that geographic area. While most offers will be closer to the starting range, successful candidates' pay will be determined based on job-related skills, experience, qualifications, work location, internal peer equity, and market conditions. These ranges may be modified in the future.

  • Region 1: $233,000 - $267,950
  • Region 2: $212,000 - $243,800
  • Region 3: $199,000 - $228,850

How would you rate this job post?

See what other professionals think about this role.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More
Staff AI Scientist at Oura