Back to Jobs
Nascent
Data Science & Analytics Just now

Data Acquisition & Web Scraping Engineer

Nascent
CanadaCanada
United StatesUnited States
Full-time
Not Disclosed
Mid-Level

Job Description

Key Skills Required

Master these to land this role

Backend42mFree Trial ✨
Start 10-Day Free Trial
DevOps1h 38mFree Trial ✨
Start 10-Day Free Trial
Web3RustCybersecurity

Want to know if you're a match for this job?

Calculate My Match Score

You'll own the systems that feed Nascent's trading and research operations with the data they need to compete—building and operating the web scraping and data acquisition infrastructure that runs 24/7 across a heterogeneous fleet of machines, providers, and operating systems. This is a roughly 50/50 split between software engineering and operations: half your time you're writing production code (primarily Rust) to build new scrapers, defeat anti-bot countermeasures, and architect resilient pipelines; the other half you're deep in logs, monitoring dashboards, and fleet health—diagnosing failures, tuning proxies, and keeping a distributed system humming across mixed infrastructure.

You'll work closely with analysts and researchers who depend on the data you deliver. The problems are genuinely interesting: every major website is actively trying to stop you, infrastructure spans multiple cloud providers and bare metal, and the surface area of what breaks is enormous. This role is remote-first. Montreal proximity is preferred—but it's not a hard requirement. North American or European time zones work best for team overlap.

Responsibilities

  • Build and maintain production-grade web scraping systems—primarily in Rust—designing scrapers that are resilient to site changes, rate limiting, CAPTCHAs, and evolving anti-bot countermeasures.

  • Operate and monitor a heterogeneous distributed infrastructure spanning multiple cloud providers, bare metal, and mixed operating systems—you own uptime, not just deployments.

  • Diagnose production issues from raw logs and telemetry, building observability into every system you ship so problems surface before they cascade.

  • Design and implement proxy management, rotation strategies, and network-layer evasion techniques to maintain reliable data acquisition at scale.

  • Develop tooling for fleet health monitoring, automated alerting, and self-healing infrastructure across a diverse set of machines and hosting environments.

  • Collaborate with analysts and researchers to understand data requirements, prioritize new source integrations, and ensure data quality and freshness meet trading-grade standards.

  • Reverse-engineer web applications and APIs—inspecting network traffic, deobfuscating JavaScript, and adapting to adversarial changes in target sites.

  • Continuously improve system reliability, throughput, and maintainability—refactoring scraping pipelines, optimizing connection handling, and reducing operational toil through automation.

  • Deploy and maintain agentic workflows for data source discovery, onboarding, and troubleshooting—using LLM-based agents to automate the identification of new sources, accelerate integration, and surface and resolve failures in existing pipelines.

About You

  • You are a builder who also runs what you build—you don't consider a project done when the PR merges; you consider it done when it's been stable in production for weeks.

  • You have a genuine interest in the cat-and-mouse game of web scraping at scale: anti-bot systems, browser fingerprinting, proxy rotation, and the constant adaptation it requires.

  • You are comfortable navigating messy, heterogeneous infrastructure—mixed OS environments, multiple hosting providers, hardware you didn't provision—and you make it better over time.

  • You are energized by operational puzzles: tracing a failure across distributed logs, identifying a subtle network degradation, or figuring out why a scraper that worked yesterday is now blocked.

  • You write clean, production-grade code and care about systems that run unattended and fail gracefully. You are proficient in at least one major programming language (Go, Python, C++, Java, or similar)—the stack is primarily Rust, and prior Rust experience is a strong plus, but if you haven't written it professionally, you're the kind of engineer who picks up new languages fast and takes ownership of the ramp. Python is useful for scripting, prototyping, and analyst-facing tooling.

  • You thrive in less-structured environments where you're trusted to prioritize your own work, and you take ownership of outcomes rather than waiting for tickets.

  • You are fluent with agentic workflows and LLM-based tooling—you know how to design and operate AI agents to discover new data sources, automate onboarding, and triage failures in production pipelines. You don't treat AI as a novelty; you reach for it when it's the right tool and build on top of it.

  • You have strong networking intuition—you think in terms of TCP connections, DNS resolution, HTTP headers, and proxy chains, not just API calls.

Preferred Experience

  • 2–5 years of professional software engineering experience with strong systems or backend fundamentals—production-grade code that runs in anger, not just toy projects or prototypes. Proficiency in at least one major programming language (Go, Python, C++, Java, or similar) is required; Rust experience is a strong plus but not a prerequisite.

  • Deep understanding of networking fundamentals: TCP/IP, DNS, HTTP internals, CDNs, proxies, load balancing, and queue handling.

  • Hands-on experience with web scraping or data acquisition at scale, including familiarity with anti-bot/anti-automation countermeasures and evasion techniques.

  • Demonstrated experience operating heterogeneous distributed infrastructure (mixed OS, hardware, hosting providers)—not just deploying to a single cloud provider.

  • Strong log analysis and monitoring skills: you can diagnose issues from raw logs and build observability into systems you own.

  • Hands-on experience building or operating agentic workflows (LLM-based agents, tool-use pipelines) for automation, data extraction, or system orchestration—this is core to how we discover and onboard new data sources.

  • Nice to have:

  • Experience with Rust—the stack is primarily Rust, and prior production experience accelerates your ramp significantly.

  • Proficiency in Python—useful for scripting, prototyping, and interacting with analyst-facing tooling.

  • Experience with cloud providers (AWS, GCP, DigitalOcean) and managing infrastructure across multiple providers simultaneously.

  • Background in high-throughput HTTP at scale—proxy management, VPN/routing optimization, and connection pooling.

  • Exposure to real-time or latency-sensitive systems (HFT-adjacent experience is a plus).

How would you rate this job post?

See what other professionals think about this role.

banner

Nascent (operating at nascentmarkets.xyz) is a proprietary investment and trading firm engineered for the crypto ecosystem, specializing in early-stage financial innovation. Founded in 2020, Nascent operates as a team of builders and investors, distinct from traditional venture capital by structuring itself as a firm where the only investors are its own team members. Unlike conventional trading firms or venture capitalists, Nascent focuses on backing ‘missionary founders’—individuals and teams pioneering category-defining protocols and products within open financial systems. Under the hood, the firm combines deep research, proprietary trading systems, and direct operational involvement to identify and amplify high-potential opportunities in emerging markets and technologies. This allows early-stage crypto founders to access critical capital, mentorship, and infrastructure while enabling Nascent to capture value through strategic trading and operational expansion. The firm’s approach bridges the gap between speculative innovation and scalable execution, fostering a new paradigm for finance built on decentralization and open systems.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More