Back to Jobs
AI & Machine Learning 2h ago

Program Manager, Voice Data Operations

United StatesUnited States
Full-time
Not Disclosed
Senior-Level

Job Description

Key Skills Required

Master these to land this role

Machine LearningBestseller 🔥
Learn in 42 Hours
Python ScriptingNLPData Science & AnalyticsAI Engineer

Want to know if you're a match for this job?

Calculate My Match Score

About the Role

At Deepgram, data isn’t just fuel for our models, it’s a product in its own right. Our vertically integrated voice AI platform depends on strategically created, curated, and labeled audio: from original data collection and augmentation, to structured workflows and evaluation sets. These data pipelines support a range of cutting-edge technologies, including audio intelligence models, conversational voice agents, STT, and TTS.

We’re looking for a hands-on, systems-minded Program Manager to lead the design and execution of various voice data programs. This role is ideal for someone who thrives on building from scratch, someone who can take an abstract modeling goal or product need and turn it into a concrete data strategy and pipeline with tools, guidelines, and quality safeguards in place. This requires ownership to understand frontier research strategies, building custom style guides, prototyping new tools, and directly influencing how data shapes our products.

This is a role for builders, someone who can spot a gap, roll up their sleeves, and design the solution. You’ll be at the center of Deepgram’s model development cycle, working across Research, Engineering, and Product, and you’ll be elbow-deep in both the day-to-day execution and the systems thinking required to scale it.

This role reports to the VP of Data Operations.

What You’ll Own:

  • Design, launch, and own end-to-end data workflows: from raw audio ingestion to production-ready datasets
  • Build and evolve labeling specs, style guides, and instructional documentation for global annotation teams
  • Identify opportunities for better tooling, automation, and workflow optimization, and lead their implementation
  • Translate product goals and model requirements into data creation strategies, deciding what to build, how to build it, and why it matters for product impact
  • Own the full lifecycle for your domains — customer expectations, data acquisition, preparation, scaling, provenance, and evaluation — and be accountable for the model outcome, not just the data hand-off.
  • Prototype and deploy data tools and infrastructure
  • Collaborate with Research and Engineering to align data collection with model training architecture and downstream product impact
  • Track advancements in speech AI research and evolving market use cases to inform labeling approaches and data design priorities
  • Partner with QA and Evaluation leads to deliver high-quality, human-in-the-loop datasets and benchmarks
  • Manage and mentor data vendors, freelancers, and potentially internal ICs as the team grows
  • Track throughput, data quality, and vendor performance
  • Drive continuous improvement in speed, cost-efficiency, and quality across all data operations
  • Curate and refine datasets to align with specific product goals, linguistic coverage, or research hypotheses

What We're Looking For

  • Experience owning data, ML, or operations programs end-to-end in program/project or product management.
  • Fluency working directly with technical teams; you can hold a conversation about data quality, evaluation, and model impact.
  • Systems thinking, understanding how decisions propagate across a system, reason up from fundamentals instead of defaulting to convention, and design solutions that hold up as things scale.
  • A track record of prioritization under constraint — deciding what to fund, what to cut, and how to sequence it.
  • Strong operating instincts: you scope, sequence, assign, and ship, and nothing stalls because someone didn't know the next step.
  • Demonstrated ability to design and build scalable processes, not just manage existing ones

It would be great if you also had:

  • Direct exposure to speech/audio, ASR, or TTS data, and the specific nuances of multilingual, code-switched, low-resource, or domain-specific data.
  • Experience with active-learning or data-selection approaches
  • Startup or high-ambiguity experience

You’ll Love This Role If You

  • Believe that data is a product, not just a resource, and you want to help define what great voice data looks like
  • Enjoy turning experimental ideas into robust, repeatable systems that can scale to production
  • Thrive in ambiguity and take initiative without waiting for perfect specs, preferring action over perfection and iteration over indecision

How would you rate this job post?

See what other professionals think about this role.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More