Back to Jobs
Innodata
AI & Machine Learning 7h ago

Research Scientist in Speech and Audio Data Engineering

Innodata
United StatesUnited States
Full-time
$160,000 - $185,000 p/year
Mid-Level

Job Description

Key Skills Required

Master these to land this role

ASR (Automatic Speech Recognition)Speech-to-SpeechPyTorchText-to-SpeechSpeech and Audio ML

Want to know if you're a match for this job?

Calculate My Match Score

Innodata (Nasdaq: INOD), a global data engineering company specializing in data, Artificial Intelligence (AI), and evaluation frameworks, is seeking a Research Scientist to focus on the science behind speech and audio data for AI models.

Scope of the Role:

Where models currently differ lies in robustness across accents, noise, and code-switching; speaker diarization; the naturalness of generated speech; and latency under streaming. Measuring these honestly, and building the data to train for them, is as much about data and evaluation design as it is about architecture. Innodata builds this data and these evaluations for customers and frontier labs advancing speech and audio models.

You will partner directly with customers and frontier labs building ASR, text-to-speech, speech-to-speech, conversational voice, diarization, and audio-language models. Your work involves making judgments on which conditions and languages a benchmark must cover to be honest, what transcription conventions should be for a given objective, and when an automated metric can be trusted versus when a human ear is required.

What You’ll Own:

  • Define how Innodata designs, structures, and evaluates audio data for speech and audio models, and validate these choices experimentally. Specifically:
  • Translate the requirements of speech and audio models — ASR, text-to-speech, speech generation, speech-to-speech, conversational voice, speaker diarization, and streaming systems — into concrete data specifications: modalities, transcription and annotation schemas, sampling, and evaluation criteria.
  • Build evaluation methodology that goes beyond word error rate, focusing on semantic accuracy, robustness to noise and accent, code-switching, diarization error rate (DER), naturalness and intelligibility of generated speech, and streaming latency. Determine when automated metrics hold and when human evaluation is necessary.
  • Decide how existing and incoming audio should be structured, enriched, and sampled for coverage across languages, accents, acoustic conditions (studio, real-world, telephonic), speaker demographics, emotional and paralinguistic range, scripted versus spontaneous speech, and single- versus multi-speaker settings, including low-resource and code-switched speech.
  • Partner with the transcription and linguistics lead to turn model objectives into transcription specifications, and quantify how transcription conventions and quality affect ASR and speech-model results.
  • Collaborate with the audio solutions and engineering team to ensure the audio collected aligns with model objectives, specifying what good data and evaluation require.
  • Conduct experiments to prove that data decisions matter, fine-tuning and evaluating models on Innodata data, and tying specific data choices to measurable improvements.
  • Design adversarial and stumping evaluations — noisy, accented, and adversarial audio — to identify where speech systems fail and turn those failures into better data.
  • Publish findings, turning insights into benchmarks, methodology, and papers that advance the field and earn the trust of customers and frontier labs.
  • Work with annotation teams, subject-matter experts, and synthetic/augmented audio pipelines to turn specifications into operational plans.

You’ll Thrive in This Role If You Have:

  • Roughly 5+ years of hands-on industry experience in speech or audio ML. Practical experience is weighted over formal credentials; a PhD with a compelling research agenda can offset the lower end.
  • A Bachelor’s degree in computer science, electrical engineering, or a related technical or quantitative field is required; an advanced degree (MS or PhD) in a relevant field is preferred.
  • Experience training and evaluating speech or audio models yourself — ASR, TTS, speech-to-speech, speaker, or audio-language models — with strong PyTorch fundamentals.
  • Fluency in toolchains and metrics for speech work, including ESPnet, NeMo, SpeechBrain, or Kaldi, HuggingFace, forced alignment, and WER/CER, as well as metrics beyond these.
  • Hands-on experience with multilingual, accented, dialectal, low-resource, or code-switched speech, and with synthetic or augmented audio (TTS pipelines, noise and room-response simulation).
  • A way of thinking in datasets: experience building evaluation sets, reasoning about coverage across conditions, and arguing about what makes speech data good for a given objective.
  • A track record recognized by the field, such as first-author publications or strong open-source contributions at venues like Interspeech, ICASSP, ASRU, SLT, or NeurIPS.
  • The ability to work directly with research scientists at partner labs, explaining data and modeling decisions clearly to both expert and non-expert audiences, backed by rigorous, reproducible experiments and documentation.
  • Bonus: Interest or hands-on experience in responsible-AI evaluation and red-teaming, such as spoofing and voice-cloning robustness, or bias across accents and languages.

How would you rate this job post?

See what other professionals think about this role.

banner

Innodata is an elite, global data engineering and AI enablement powerhouse engineered to orchestrate massive-scale high-quality data operations, algorithmic training datasets, and digital transformation workflows for the world’s largest technology companies and enterprises. Operating as a critical "intelligence infrastructure layer" for the modern AI economy, the company eliminates the operational friction of deploying complex LLMs and generative AI applications—which frequently suffer from low-quality training data, biased outputs, and fragmented annotation pipelines—by seamlessly deploying a combination of advanced proprietary data annotation platforms, automated synthetic data generation, and an elite global network of subject matter experts. Moving beyond basic crowdsourced data validation paradigms, Innodata empowers Fortune 500 enterprises, global legal publishers, and leading medical institutions to dynamically scale their core AI foundation models, custom machine learning pipelines, and multi-modal semantic data processing with elite, scalable, and audit-ready precision. Under the hood, their sophisticated operational framework—bolstered by strict data security compliance, persistent programmatic data quality controls, and deep domain expertise across vertical domains like legal, healthcare, and finance—natively manages high-velocity data curation, complex enterprise knowledge graph construction, and high-stakes model evaluation. What sets Innodata apart is its uncompromising dedication to purpose-driven data craftsmanship; by bridging the gap between raw, unstructured enterprise information and high-performance, fine-tuned AI execution, the firm enables global commercial organizations to radically accelerate their AI time-to-market, eliminate systemic data engineering bottlenecks, and build an unassailable foundation for continuous commercial growth in the modern, AI-transformed global marketplace.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More