Back to Jobs
Wynd Labs
Data Science & Analytics 2h ago

Data Pipeline Engineer

Wynd Labs
United StatesUnited States
Full-time
Competitive salary, benefits and equity package
Mid-Level

Job Description

Key Skills Required

Master these to land this role

Python2h 41mFree Trial ✨
Start 10-Day Free Trial
Data ScientistBig DataData PipelinesWeb Scraping

Want to know if you're a match for this job?

Calculate My Match Score

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.

We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.

We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.

Who You Are:

  • Bachelor’s degree or equivalent work experience
  • Strong grasp of Python (advanced) — async programming, multiprocessing, and writing production-grade code for long-running data jobs
  • Hands-on experience with web scraping at scale (proxies, rate limiting, anti-bot evasion). Experience with platform APIs and large media/metadata datasets (video platforms, social media)
  • Experience designing and operating distributed data pipelines across many workers/servers using task queues (Celery, Kafka, RabbitMQ, or similar)
  • Practical experience with data warehousing — columnar/analytical warehouses; Databend, ClickHouse, or BigQuery strongly preferred; comfortable with complex analytical queries, partitioning strategies, cost-aware querying on cloud warehouses
  • Containerizing workloads, writing Helm charts/manifests, managing deployments, autoscaling scraping/processing workloads using Docker & Kubernetes
  • Comfortable managing services on Linux servers, debugging performance issues (disk I/O, network, memory) without managed-cloud abstractions (Linux & bare-metal ops)
  • CI/CD for data workflows (GitHub Actions, ArgoCD)
  • Writing Scalable API

What You'll Be Doing:

  • Maintain, optimize, and troubleshoot database queries and related data systems to support efficient data access, processing, and reliability.
  • Assist in creating, maintaining, and improving data pipelines used to collect, process, transform, validate, and deliver large-scale datasets.
  • Support web scraping and data collection initiatives, including developing, testing, and maintaining scripts or tools used to gather publicly available data in accordance with Company requirements.
  • Monitor and troubleshoot data pipeline issues, identify data quality concerns, and assist in implementing timely fixes to maintain data accuracy and operational continuity.
  • Document engineering work, including database queries, pipeline processes, scraping workflows, technical decisions, issues encountered, and resolutions implemented.
  • Participate in research and development projects to improve the Company’s data products and workflows.

Why Work With Us:

  • Opportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.
  • Culture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better. We prioritize low ego and high output. This is a fully remote team.
  • Compensation. You’ll receive a competitive salary, benefits and equity package.

How would you rate this job post?

See what other professionals think about this role.

banner

Wynd Labs is a premier, enterprise-grade data infrastructure platform engineered to orchestrate massive-scale public web data ecosystems and intelligent artificial intelligence training workflows. Operating as a highly integrated decentralized data hub, the company eliminates the operational friction of traditional localized web scraping by seamlessly deploying advanced distributed crawling telemetry, rigorous multimodal data pipelines, and cohesive residential proxy architectures. Moving beyond rigid legacy dataset providers, Wynd Labs empowers frontier AI labs, elite research teams, and data-driven enterprises to dynamically synchronize their machine learning models with instantaneous, internet-scale data ingestion. Under the hood, their sophisticated backend infrastructure natively handles complex high-throughput routing, scalable real-time search extraction, and seamless petabyte-scale multimedia annotation, ensuring frictionless data accessibility and uncompromising model training readiness. What sets Wynd Labs apart is its uncompromising dedication to frictionless data orchestration; by bridging the gap between decentralized bandwidth sharing and rigorous artificial intelligence development, the platform empowers organizations to radically accelerate their algorithmic velocity, optimize data acquisition, and build an unassailable foundation for continuous AI dominance in the modern computational landscape.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More
Data Pipeline Engineer at Wynd Labs | HireSkys