Senior Engineer - Open Voice-Agent Stack
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
At Hugging Face, we're on a journey to democratize good AI. We are building the fastest growing platform for AI builders with over 11 million users who collectively shared over 3M+ models, 1M+ datasets & 1.47M+ apps. Our open-source libraries have more than 600k+ stars on Github.
About the Role
We are building the open voice-agent stack for Hugging Face, and we are looking for a senior engineer to own a large part of it. Two things sit at the centre of this role. The first is speech-to-speech, our open-source library for realtime voice agents. The second is hf-voice, a new product that will let any developer build and deploy voice agents with their Hugging Face account.
The library already powers the Reachy Mini fleet and there is a public demo running on Spaces, so you won't start from a blank page. But almost everything about how this becomes a product developers rely on is still open, and you will have a direct say in it.
Your missions:
- Own the Open-Source Library:
- Take architectural ownership of large parts of speech-to-speech: pipeline design, latency budget, and the reliability of the realtime loop.
- Integrate new ASR, TTS and end-to-end speech models as they land, and keep the abstractions clean while the model landscape keeps moving.
- Review community PRs, triage issues, cut releases, and grow the group of contributors around the project.
- Ship hf-voice:
- Design the developer API and the streaming protocol: session lifecycle, transport (WebSockets/WebRTC), authentication, error semantics, versioning.
- Build the serving side: realtime inference on GPU, concurrency, autoscaling, observability, and cost per session.
- Work with the Hub and inference teams so that a working voice agent is easy to integrate into products and demos.
- Take the product from demo to production: load testing, SLOs, graceful degradation when a model or a network path misbehaves.
- Work in the Open:
- Write the docs, examples and templates that get a developer from zero to a running agent in minutes.
- Support the deployments already relying on the stack, starting with the Reachy Mini fleet.
- Talk about the work publicly if you enjoy it: blog posts, demos, conference talks. We cover the travel and the prep time.
Requirements
What we're looking for
- Senior engineer, able to own a substantial part of an architecture and drive it forward autonomously.
- Experience building developer-facing infrastructure at an AI or developer-tools company: inference APIs, agent infrastructure, or something comparable.
- Substantial open-source contributions to a Python library. Comfortable with async Python and distributed systems, including their failure modes.
- You have shipped something realtime: streaming, WebSockets or WebRTC, audio or video pipelines, live inference.
- Practical experience with LLMs or multimodal models in production. Clear written communication and a habit of collaborating async and in public.
- Motivated by voice and conversational AI.
Bonus points if you have
- Contributions to a voice-agent framework such as speech-to-speech, pipecat, LiveKit Agents, Vocode or TEN.
- Contributions to llama.cpp or another low-level inference runtime.
- Hands-on work with ASR, TTS or end-to-end speech models, including evaluation of latency and quality trade-offs.
- GPU serving, quantization, or on-device inference experience.
- Audio pipeline knowledge: VAD, echo cancellation, jitter buffers, barge-in and turn detection.
- Experience shipping to embedded or robotics targets.
- A public track record: talks, blog posts, demos.
About You
If you're interested in joining us, but don't tick every box above, we still encourage you to apply! We're building a diverse team whose skills, experiences, and backgrounds complement one another. We're happy to consider where you might be able to make the biggest impact.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Hugging Face
Explore Top Companies in this Space
Tribe AI
Artificial Intelligence / Enterprise Software / AI Consultancy & Services / Machine Learning
Supernal
Artificial Intelligence / Enterprise Software / Machine Learning
Aion
Artificial Intelligence & Machine Learning / Cloud Infrastructure & GPU Compute / Enterprise Software
DigitalGenius
AI / SaaS / Ecommerce / Customer Service
Hugging Face
View Company ProfileHugging Face (operating at huggingface.co) is a machine learning community platform engineered for collaboration. Founded in 2016 by Clément Delangue, Julien Chaumond, and others and headquartered in Manhattan, Hugging Face develops computation tools for building applications using machine learning. Under the hood, the platform leverages a vast repository of pre-trained models, datasets, and metadata, allowing developers and researchers to build and deploy AI applications without manual browsing. This allows users to accelerate their machine learning workflows and focus on innovation. Backed by $395.2 million in funding from investors including Lux Capital, Sequoia, Coatue, and Addition.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.


