Back to Jobs
Torc Robotics
AI & Machine Learning 3d ago

Pseudo-Labeling and Data Pipeline Engineer

Torc Robotics
United StatesUnited States
Full-time
$160,800 — $193,000 USD
Senior-Level

Job Description

Key Skills Required

Master these to land this role

Machine LearningBestseller 🔥
Learn in 42 Hours
Data Science & AnalyticsMLflowPyTorchPython Scripting

Want to know if you're a match for this job?

Calculate My Match Score

What You'll Do:

Design, implement, test and deploy tooling and pipelines for internal quality control and issue identification of pseudo-labeled data, ensuring annotations meet the bar required by downstream model training.

  • Provide statistical support and report on pseudo-label quality, coverage, and pipeline health to internal stakeholders and leadership.
  • Support secondary data selection (virtual packaging) to curate and feed on-demand data to online model training.
  • Support the build-out of the ML data delivery system that enables the online perception team to train models on-demand.
  • Drive general pipeline improvement and optimization across the pseudo-labeling and data delivery stack, identifying and resolving bottlenecks in throughput, quality, or reliability.
  • Demonstrate project management skills, serving as project lead guiding less experienced team members in multiple facets of project execution.
  • Stay up to date with the latest developments in offline perception, data pipeline engineering, and ML data infrastructure for autonomous driving.
  • Independently develop tools, services, and algorithms using disciplined software development processes, making recommendations for developing new code or re-using existing code, implementing version control, and maintaining documentation of created applications.
  • Define and implement ingestion, data preparation, curation, and governance of large, multi-faceted data sets supporting analytics and ML training workflows.
  • Proactively assess current capabilities to identify areas for improvement, proposing solutions that align with core strategy and operation.
  • Guide and produce information products, supporting visualization and data accessibility in a customer-centric manner.
  • Evaluate and make recommendations regarding technical advances that improve productivity and quality, reduce flow times, and enhance operational surety.
  • Develop guidelines and standards for data quality control, data delivery systems, and their deployment, and associated processes.
  • Provide technical guidance or business process expertise, technical leadership, coaching and mentoring to team members.

What You'll Need to Succeed:

Considered highly skilled and proficient in discipline; conducts complex, important work under minimal supervision and with wide latitude for independent judgment.

  • Scope of Influence: Expected to drive alignment across team interfaces to the rest of the organization. Designs, maintains and owns team technical solutions and drives consensus. Mentors and guides engineers within the group.
  • Bachelor’s Degree in Computer Science, Robotics, Electrical Engineering or related technical field plus demonstrated competencies and technical proficiencies typically acquired through 6+ years of experience OR;
  • Master’s Degree in Computer Science, Robotics, Electrical Engineering or related technical field plus demonstrated competencies and technical proficiencies typically acquired through 3+ years of experience OR;
  • Required Qualifications (some combination of the following skills):
    • Familiarity with the offline perception stack in general, and knowledge of how pseudo-label data is produced and related best practices.
    • Strong software engineering background building and operating data pipelines and services at scale.
    • Statistical analysis and reporting skills, with the ability to translate data quality findings into actionable insights.
    • Scaled MLOps and ToolingML Frameworks, experiment tracking, model registry, MLflow, Weights & Biases, ML Metrics and Evaluation / Quality.
    • Model Data CurationParquet data processing (PyArrow, Daft, Pandas, etc).
    • Development Tools & Eco-System (at scale) – Proficiency in Python software development. Also, VDI and cloud-based development environments, CI Systems (GitHub Actions), and Docker.
    • Experience with distributed data processing and/or ML frameworksPyTorch, Lightning, Ray, or similar.

Bonus Points!

  • Experience with large-scale data delivery systems and associated quality control.
  • Pseudo-labeling experience in general.

How would you rate this job post?

See what other professionals think about this role.

banner

Torc Robotics (operating at torc.ai) is a leading autonomous vehicle software company engineered for safe, sustained innovation in the trucking industry. Founded in 2005 by Michael Fleming and a group of Virginia Tech students, and headquartered in Blacksburg, Virginia, Torc Robotics offers a complete autonomous software solution for the trucking/freight industry. Under the hood, the physical AI developed at Torc enables self-driving trucks to perceive, understand, and perform complex actions in the real (physical) world. This allows experienced partners to commercialize autonomous solutions. Not specified.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More
Pseudo-Labeling and Data Pipeline Engineer at Torc Robotics