Research Data Scientist – GenAI/LLM
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Innodata (Nasdaq: INOD), a global data engineering company, specializes in enabling responsible AI advancement through data, evaluation frameworks, and expertise. They provide solutions for Generative AI and AI adopters, with a legacy of delivering high-quality data and outcomes.
Scope of the Role:
We are seeking a highly skilled Research Data Scientist – GenAI/LLM to join our AI/LLM Delivery Unit. This role focuses on research-driven AI/ML initiatives involving Generative AI, Large Language Models (LLMs), NLP, multimodal AI, model evaluation, and AI data.
The role combines strong research and analytical capabilities with hands-on AI/ML expertise, requiring the candidate to design experiments, develop evaluation methodologies, analyze complex datasets, build research prototypes, and translate research findings into practical AI/ML solutions.
The ideal candidate will have a strong research orientation, excellent statistical and analytical skills, and the ability to work collaboratively with researchers, data scientists, AI/ML engineers, domain experts, and client-facing teams.
What You’ll Own:
AI/ML & Generative AI Research:
- Conduct independent and collaborative research in Generative AI, LLMs, NLP, multimodal AI, machine learning, model evaluation, and AI data.
- Formulate research questions and translate complex AI/ML problems into structured research methodologies and experiments.
- Design, execute, and analyze experiments to evaluate and improve AI/ML models and solutions.
- Build analytical models, prototypes, and research pipelines using Python and relevant ML frameworks.
- Stay current with emerging research, methodologies, papers, and developments in GenAI, LLMs, NLP, multimodal models, and AI evaluation.
LLM & Model Evaluation:
- Develop and implement LLM evaluation frameworks, benchmarks, datasets, and evaluation criteria.
- Evaluate models for accuracy, robustness, bias, hallucination, reasoning, relevance, response quality, and other performance dimensions.
- Conduct model benchmarking, error analysis, comparative analysis, and performance evaluation.
- Work on areas such as RAG, SFT, RLHF/DPO, prompt engineering, fine-tuning, embeddings, and LLM optimization, as applicable.
- Identify model and data gaps and recommend improvements to enhance model performance and reliability.
Data Science & Statistical Research:
- Collect, clean, analyze, and interpret large and complex structured and unstructured datasets.
- Perform EDA, statistical analysis, hypothesis testing, significance testing, correlation analysis, sampling, and error analysis.
- Develop data-driven insights and identify patterns, trends, and relationships relevant to AI/ML research.
- Apply appropriate statistical and quantitative methodologies to validate research findings.
AI Data & Dataset Development:
- Develop and evaluate datasets, sampling methodologies, taxonomies, annotation frameworks, data quality frameworks, and evaluation criteria for AI/ML models.
- Analyze data quality and identify issues affecting model performance.
- Collaborate with annotation, data engineering, and AI/ML teams to improve AI training and evaluation data.
- Translate data and research findings into actionable recommendations for improving AI systems.
Research & Innovation:
- Contribute to research papers, technical reports, whitepapers, patents, benchmarks, internal publications, and other research outputs, where applicable.
- Identify opportunities to apply emerging research and technologies to real-world AI and data challenges.
- Explore new methodologies, models, datasets, and evaluation approaches to improve AI capabilities.
- Contribute to capability building and innovation within the AI/LLM practice.
Collaboration & Stakeholder Engagement:
- Work closely with researchers, data scientists, AI/ML engineers, data/annotation teams, domain experts, and delivery teams.
- Present research findings, analytical insights, and technical recommendations to senior technical stakeholders.
- Translate complex research and technical concepts into clear, actionable recommendations.
- Participate in client-facing technical discussions and presentations, translating business requirements into AI/ML solutions.
You’ll Thrive in This Role If You Have:
- A Master’s or PhD in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Statistics, Mathematics, Computational Science, or a related discipline.
- A Bachelor’s/Master’s degree from IITs, NITs, or other premier engineering/research institutions is strongly preferred.
- 4–7 years of hands-on research experience in AI/ML, Data Science, NLP, Generative AI, LLMs, or related areas.
- Strong demonstrated research experience with the ability to independently formulate research questions, design experiments, analyze results, and communicate findings.
- A demonstrated research track record through research publications, patents, conference presentations, open-source contributions, or significant AI/ML research projects.
- Candidates with publications in reputed conferences/journals and a strong academic/research profile will be preferred.
- Strong proficiency in Python and SQL.
- Strong hands-on experience with NumPy, Pandas, Scikit-learn, and preferably PyTorch/TensorFlow.
- Strong understanding of:
- Machine learning algorithms
- Statistics and experimentation
- Data analysis and feature engineering
- Model evaluation and performance metrics
- Hypothesis testing and statistical inference
- Hands-on exposure to LLMs, NLP, Generative AI, and multimodal AI.
- Experience with one or more of RAG, LLM evaluation, prompt engineering, fine-tuning, SFT, RLHF/DPO, embeddings, or model benchmarking.
- Experience working with large-scale structured and unstructured datasets.
- Familiarity with Git and cloud platforms such as AWS, Azure, or GCP is desirable.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Innodata
Technical Training & Quality Manager – AI Data
Innodata
United StatesTechnical Delivery Manager (TDM) for AI/ML Engagements
Innodata
United StatesBusiness Development Representative - Federal Practice
Innodata
United StatesStaff / Lead Data Analyst / Analytics Engineer
Innodata
CanadaExplore Top Companies in this Space
SunnyData
IT Consulting / Data Engineering / AI / Enterprise Software
Mactores
Data Engineering & AI
itD Tech
Digital Transformation Consulting & IT Architecture / Enterprise AI Strategy & Implementation / Advanced Data Engineering & Cloud-Native Development
Bounteous
AI Services & Digital Transformation Consulting / Product Engineering & Experience Design / Enterprise Data, Cloud & Marketing Technology Integration
Innodata
View Company ProfileInnodata is an elite, global data engineering and AI enablement powerhouse engineered to orchestrate massive-scale high-quality data operations, algorithmic training datasets, and digital transformation workflows for the world’s largest technology companies and enterprises. Operating as a critical "intelligence infrastructure layer" for the modern AI economy, the company eliminates the operational friction of deploying complex LLMs and generative AI applications—which frequently suffer from low-quality training data, biased outputs, and fragmented annotation pipelines—by seamlessly deploying a combination of advanced proprietary data annotation platforms, automated synthetic data generation, and an elite global network of subject matter experts. Moving beyond basic crowdsourced data validation paradigms, Innodata empowers Fortune 500 enterprises, global legal publishers, and leading medical institutions to dynamically scale their core AI foundation models, custom machine learning pipelines, and multi-modal semantic data processing with elite, scalable, and audit-ready precision. Under the hood, their sophisticated operational framework—bolstered by strict data security compliance, persistent programmatic data quality controls, and deep domain expertise across vertical domains like legal, healthcare, and finance—natively manages high-velocity data curation, complex enterprise knowledge graph construction, and high-stakes model evaluation. What sets Innodata apart is its uncompromising dedication to purpose-driven data craftsmanship; by bridging the gap between raw, unstructured enterprise information and high-performance, fine-tuned AI execution, the firm enables global commercial organizations to radically accelerate their AI time-to-market, eliminate systemic data engineering bottlenecks, and build an unassailable foundation for continuous commercial growth in the modern, AI-transformed global marketplace.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.