Back to Jobs
Neurons Lab
AI & Machine Learning 20h ago

AI Voice Copilot Architect (Real-Time Voice Pipeline)

Neurons Lab
PolandPoland
PortugalPortugal
SloveniaSlovenia
CyprusCyprus
GreeceGreece
SerbiaSerbia
HungaryHungary
SpainSpain
ItalyItaly
SlovakiaSlovakia
LithuaniaLithuania
AlbaniaAlbania
🌍Czechia
MoldovaMoldova
UkraineUkraine
BulgariaBulgaria
EstoniaEstonia
🌍North Macedonia
RomaniaRomania
Contract
Not Disclosed
Senior-Level

Job Description

Key Skills Required

Master these to land this role

Machine Learning41mFree Trial ✨
Start 10-Day Free Trial
Prompt Engineering17mFree Trial ✨
Start 10-Day Free Trial
Python ScriptingNLPAI Engineer

Want to know if you're a match for this job?

Calculate My Match Score

Objective

Own the technical architecture and delivery of the voice copilot from validated PoC to production.

  • Hit the bar this client tests against: latency, accuracy, concurrency, and cost.
  • Keep expectations aligned: production polish is in scope now; protect the team from silent scope creep.
  • Transfer knowledge continuously to the client's team and Neurons Lab engineers.

Areas of Responsibility

Technical architecture & hands-on implementation

  • Own the full pipeline: streaming speech-to-text, LLM field extraction, Chrome-extension delivery, and AWS infrastructure.
  • Drive latency work: cut P95 from ~6s toward ~2s; remove post-processing corner cases (occasional ~1min lag on one field type).
  • Run model A/B tests (current pair: Claude Haiku vs GPT Luna) with golden-set evaluation for phonetic name and email accuracy.
  • Own evaluation and cost: Langfuse traces, accuracy dashboards, real per-call cost from live calls, and an optimization plan.
  • Harden for production: 5–10+ concurrent calls, strict data isolation between users, monitoring, alerting, and safe rollback.
  • Ship epics end to end (example: the SES email briefing service); always keep a demo fallback so a live session never fails.

Working with client stakeholders

  • Front technical discussions with a meticulous client; VCCs test edge cases and expect production quality.
  • Present concrete system behavior, with numbers — this account rewards evidence, not slides.
  • Hold the scope line: tie every feedback item to the SOW; route roadmap items (learning loop, persistent memory) to future phases.
  • Keep internal discussions internal; all client-facing materials pass ADM review before sending.

Team & knowledge

  • Lead the AI Engineer and the pod: set tasks, review output, unblock fast.
  • Absorb the handover from the outgoing architect (0.15–0.2 FTE supervision window) and become independent fast.
  • Run knowledge-transfer sessions; the project must have no single point of failure.
  • Support the production SOW with estimates and architecture options when the account team asks.

Skills

  • Real-time voice pipelines: streaming STT, turn handling, low-latency LLM inference — hands-on.
  • LLM engineering: prompt engineering, structured extraction, guardrails, model A/B evaluation.
  • Observability and evals: Langfuse or similar; golden datasets; latency, accuracy, and cost dashboards.
  • AWS: Bedrock, serverless patterns, SES; token economics and per-call cost engineering.
  • Full-stack pragmatism: strong Python; enough TypeScript / Chrome-extension knowledge to own the integration.
  • Clear spoken and written English for demanding US executives.

Knowledge

  • Contact-center / agent-assist patterns and metrics (handle time, cost per call, concurrency).
  • Production LLM operations: load testing, data isolation, incident handling.
  • Nice to have: empathy-sensitive domains (healthcare, veterinary, insurance) and PE-sponsored rollouts.

Experience

Key characteristics (screen for all four):

  1. Voice AI in production — mandatory. Shipped at least one real-time voice or speech product to real users (agent assist, voice bot, live transcription copilot). Candidates will demo real artifacts at the interview.
  2. 6+ years hands-on AI/ML engineering, with strong recent LLM production practice.
  3. Latency and reliability record. Can show measured P95 reductions and concurrency fixes on a live system.
  4. Consulting / client-facing seniority. Calm and precise under detailed UAT scrutiny; manages expectations well.

Nice to have:

  • Chrome extension delivery; telephony / streaming stacks (Amazon Connect, Twilio, LiveKit).
  • Langfuse in production.
  • US client experience with Eastern-time overlap.

How would you rate this job post?

See what other professionals think about this role.

banner

Neurons Lab (operating at neurons-lab.com) is a leading AI consultancy platform engineered for AI transformation services. Founded by Igor Sydorenko and headquartered in London, United Kingdom, Neurons Lab helps financial institutions move from AI-curious to AI-enabled. Under the hood, the company delivers AI training programs and custom AI agents designed for regulated environments. This allows financial institutions to leverage AI technology effectively. Backed by no publicly disclosed funding information.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More
AI Voice Copilot Architect (Real-Time Voice Pipeline) at Neurons Lab | HireSkys