Pseudo-Labeling and Data Pipeline Engineer
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
What You'll Do:
Design, implement, test and deploy tooling and pipelines for internal quality control and issue identification of pseudo-labeled data, ensuring annotations meet the bar required by downstream model training.
- Provide statistical support and report on pseudo-label quality, coverage, and pipeline health to internal stakeholders and leadership.
- Support secondary data selection (virtual packaging) to curate and feed on-demand data to online model training.
- Support the build-out of the ML data delivery system that enables the online perception team to train models on-demand.
- Drive general pipeline improvement and optimization across the pseudo-labeling and data delivery stack, identifying and resolving bottlenecks in throughput, quality, or reliability.
- Demonstrate project management skills, serving as project lead guiding less experienced team members in multiple facets of project execution.
- Stay up to date with the latest developments in offline perception, data pipeline engineering, and ML data infrastructure for autonomous driving.
- Independently develop tools, services, and algorithms using disciplined software development processes, making recommendations for developing new code or re-using existing code, implementing version control, and maintaining documentation of created applications.
- Define and implement ingestion, data preparation, curation, and governance of large, multi-faceted data sets supporting analytics and ML training workflows.
- Proactively assess current capabilities to identify areas for improvement, proposing solutions that align with core strategy and operation.
- Guide and produce information products, supporting visualization and data accessibility in a customer-centric manner.
- Evaluate and make recommendations regarding technical advances that improve productivity and quality, reduce flow times, and enhance operational surety.
- Develop guidelines and standards for data quality control, data delivery systems, and their deployment, and associated processes.
- Provide technical guidance or business process expertise, technical leadership, coaching and mentoring to team members.
What You'll Need to Succeed:
Considered highly skilled and proficient in discipline; conducts complex, important work under minimal supervision and with wide latitude for independent judgment.
- Scope of Influence: Expected to drive alignment across team interfaces to the rest of the organization. Designs, maintains and owns team technical solutions and drives consensus. Mentors and guides engineers within the group.
- Bachelor’s Degree in Computer Science, Robotics, Electrical Engineering or related technical field plus demonstrated competencies and technical proficiencies typically acquired through 6+ years of experience OR;
- Master’s Degree in Computer Science, Robotics, Electrical Engineering or related technical field plus demonstrated competencies and technical proficiencies typically acquired through 3+ years of experience OR;
- Required Qualifications (some combination of the following skills):
- Familiarity with the offline perception stack in general, and knowledge of how pseudo-label data is produced and related best practices.
- Strong software engineering background building and operating data pipelines and services at scale.
- Statistical analysis and reporting skills, with the ability to translate data quality findings into actionable insights.
- Scaled MLOps and Tooling – ML Frameworks, experiment tracking, model registry, MLflow, Weights & Biases, ML Metrics and Evaluation / Quality.
- Model Data Curation – Parquet data processing (PyArrow, Daft, Pandas, etc).
- Development Tools & Eco-System (at scale) – Proficiency in Python software development. Also, VDI and cloud-based development environments, CI Systems (GitHub Actions), and Docker.
- Experience with distributed data processing and/or ML frameworks – PyTorch, Lightning, Ray, or similar.
Bonus Points!
- Experience with large-scale data delivery systems and associated quality control.
- Pseudo-labeling experience in general.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Torc Robotics
Software Engineer - Autonomous Vehicle Data Pipelines
Torc Robotics
United StatesSenior Product Cybersecurity Architect
Torc Robotics
United StatesSenior Autonomous Vehicle Perception and Mapping Engineer
Torc Robotics
United StatesStaff Machine Learning Engineer - Scene Generation
Torc Robotics
United StatesExplore Top Companies in this Space
Pirate Ship
Logistics / E-commerce Shipping / SaaS
Pallet
Logistics and Supply Chain Management
Shipbob
E-commerce Logistics and Supply Chain
Shipbobinc
E-commerce Logistics and Shipping
Torc Robotics
View Company ProfileTorc Robotics (operating at torc.ai) is a leading autonomous vehicle software company engineered for safe, sustained innovation in the trucking industry. Founded in 2005 by Michael Fleming and a group of Virginia Tech students, and headquartered in Blacksburg, Virginia, Torc Robotics offers a complete autonomous software solution for the trucking/freight industry. Under the hood, the physical AI developed at Torc enables self-driving trucks to perceive, understand, and perform complex actions in the real (physical) world. This allows experienced partners to commercialize autonomous solutions. Not specified.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.
