Research Engineer - AI Training Data Quality
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
About the Company
Our partner is a YC-backed company building a new kind of marketplace in the AI training data space. Rather than operating as a labor marketplace, they provide infrastructure that lets data producers transform their existing data into formats AI labs want and sell it directly to those labs. This democratized model unlocks far more high-value data sources, and the team is growing quickly to keep up with demand.
The Opportunity
This is the company's top hiring priority and a genuinely hard research problem. Because data flows through a decentralized marketplace, ensuring quality at scale is the single biggest bottleneck to growth. As a Research Engineer, you will build the automated systems that verify and assure data quality so that suppliers consistently deliver excellent data to buyers.
You will start by digging into the data manually to understand failure modes, then design systems to automate quality checks at scale, combining rule-based approaches with AI for fuzzier cases and human-in-the-loop review where it makes sense. This is fundamentally a research role focused on building automated systems, not manual QA.
Responsibilities
- Identify data quality issues including inconsistencies, formatting problems, and ingestion challenges
- Perform initial manual data quality review to deeply understand failure modes
- Build systems to automate quality checks at scale using rule-based and AI-driven approaches
- Design hybrid systems that balance automation with human-in-the-loop review where appropriate
- Continuously improve verification methods as the data landscape and AI tooling evolve
Requirements
- Deeply technical, with a strong learning slope and the ability to ramp quickly in a fast-moving field
- Background in AI/ML engineering, or software engineering at an AI-focused company with visible data ingestion and processing experience
- Ability to reason about likely data quality problems from first principles
- Comfortable owning ambiguous, open-ended problems end to end
- Comfortable working in person, full-time, in a San Francisco office
- Bonus: experience working with noisy or unstructured data, or judgment on when to use automation versus human-in-the-loop review
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at TalentPluto
TalentPluto
View Company ProfileTalentPluto is an innovative, AI-native recruitment platform and career agent designed specifically to connect elite Go-To-Market (GTM) professionals with high-growth, VC-backed startups. Moving away from the traditional, resume-spamming job board model, the company operates via a proprietary "voice AI" headhunter named Pluto. Under the hood, candidates simply jump on a 5-to-10-minute phone call with the AI to discuss their background, specific experience, and career goals. Pluto then instantly maps this data against a curated network of over 100 top-tier tech companies (ranging from Seed to Series C) actively hiring for roles like Account Executives, SDRs, Growth Marketers, and Customer Success Managers. What sets TalentPluto apart in the crowded HR tech space is its aggressive focus on candidate privacy and the "opt-in" model. Candidates remain entirely anonymous while Pluto pitches their profiles to hiring managers; they only reveal their identities and engage when they receive a match they are genuinely interested in pursuing. By eliminating tedious application forms, ghosting, and cold outreach, TalentPluto provides startups with highly qualified, intent-driven talent while giving professionals a frictionless path to their next big career move.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.


