Back to Jobs
Benchling
Data Science & Analytics 1d ago

Data Engineer, AI & Data Engineering (AIDE)

Benchling
United StatesUnited States
Full-time
Not Disclosed
Mid-Level

Job Description

Key Skills Required

Master these to land this role

Python2h 41mFree Trial ✨
Start 10-Day Free Trial
SQL15 minFree Trial ✨
Start 10-Day Free Trial
AI EngineerData EngineerSnowflake

Want to know if you're a match for this job?

Calculate My Match Score

ROLE OVERVIEW

Biotechnology is rewriting life as we know it, from the medicines we take, to the crops we grow, the materials we wear, and the household goods that we rely on every day. But moving at the new speed of science requires better technology. Benchling's mission is to unlock the power of biotechnology. The world's most innovative biotech companies use Benchling's R&D Cloud to power the development of breakthrough products and accelerate time to milestone and market. Come help us bring modern software to modern science.

Benchling is building AI & Data Engineering (AIDE), a small, autonomous team within our Security & IT organization. AIDE owns three things: internal AI tooling, adoption, and AI-assisted workflows across the company; cross-functional and company-wide agentic AI applications that no single department owns; and the enterprise data engineering, analytics architecture, and source-of-truth datasets that everything above depends on. AIDE’s data and analytics functions grew out of our former Data, Analytics & Systems (DAS) team, and this role carries forward DAS's original charter: building and running the data pipelines, warehouse, and analytics infrastructure that the entire company relies on for trustworthy answers.

This is a data engineering role — we want someone who builds and operates reliable, production-grade data pipelines and warehouse infrastructure, not a data scientist focused on modeling or analysis.

This role exists because AIDE's data function supports the whole company — GTM, Customer Success, Product, Finance, and beyond — not just one team, and the team needs to grow to support these initiatives as we expand the team’s scope and portfolio. You'll own core pipelines end to end (ingestion, transformation, warehouse, and the BI/analytics layer on top), partner with the rest of the data team on the team's data architecture, and help build the trusted data foundation that AIDE's AI-adoption and agentic AI work increasingly depends on.

Check out our engineering blog for examples of past work across Benchling.

RESPONSIBILITIES

  • Own core data pipelines end to end: Build and operate the ELT pipeline that moves data from Benchling's product, Salesforce, and third-party systems into Snowflake, modeled with dbt, and built to production standards — testing, monitoring, schema versioning — that hold up as usage scales. This is infrastructure the rest of the company builds on, not a one-off project.

  • Build the data foundation for AIDE's AI initiatives: Partner with AIDE's AI engineering side to make governed, trustworthy data available for the agentic AI tooling and internal AI applications the team ships.

  • Own data governance and pipeline health: Maintain Snowflake access controls (RBAC), monitor data quality, uphold PII-handling and data-access policy, and manage warehouse cost and performance as usage grows.

  • Contribute to platform strategy: Weigh in on bigger structural decisions — warehouse architecture, semantic layer/metrics store design— alongside the rest of the data and AI engineering team.

QUALIFICATIONS

  • 3+ years of professional experience building and operating production data pipelines — ingestion, transformation, and modeling data into a cloud data warehouse.

  • Strong SQL and Python skills; hands on experience with data modeling methodologies and tools, preferable with dbt.

  • Experience applying software engineering practices to data systems — version control, code review, CI/CD, automated testing — and comfort working with cloud infrastructure (AWS or similar) supporting production pipelines.

  • Experience with Snowflake or a comparable modern cloud data warehouse in production.

  • Comfort with orchestration tooling (Airflow or similar) for scheduled data jobs.

  • Track record supporting many stakeholders across departments such as Sales, CS, Product, Finance, rather than a single internal customer.

  • Understanding of data privacy, governance, quality, and testing frameworks and best practices.

  • Strong communication skills; comfortable translating ambiguous requests from non-technical stakeholders into a scoped, buildable data solution.

  • Comfortable in a small, fast-moving, still-forming team — AIDE only stood up in its current form in mid-2026 and is actively defining its own processes.

  • Interest in learning more about life science (prior knowledge is not required).

NICE TO HAVE

  • Familiarity with product behavioral data and a modern BI tool (Sigma, Omni, Looker, Tableau) deployed in a self-service model.

  • Experience with product/usage analytics instrumentation and event-taxonomy governance.

  • Familiarity with GTM analytics tools such as Salesforce.

  • Exposure to AI-usage telemetry, LLM observability data, or supporting AI/ML tooling with curated data.

  • Background in enterprise SaaS, life sciences, or biotech.

  • Experience building or maintaining a metrics layer.

How would you rate this job post?

See what other professionals think about this role.

banner

Benchling is a premier, enterprise-grade cloud-based R&D platform engineered to orchestrate massive-scale life-sciences ecosystems and intelligent, frictionless biotechnology-innovation workflows. Operating as the global standard for modern, AI-native R&D, the company eliminates the operational friction of traditional, legacy-disconnected laboratory systems—which frequently trap critical scientific data in fragmented, manual, or paper-based silos—by seamlessly deploying advanced Electronic Lab Notebook (ELN) telemetry, rigorous Laboratory Information Management System (LIMS) architectures, and cohesive cross-platform data-integration frameworks. Moving beyond rigid legacy desktop paradigms, Benchling empowers over 200,000 global scientists, biopharma giants, and academic institutions to dynamically synchronize their molecule-discovery, process-development, and clinical-readiness pipelines with elite, scalable autonomous execution. Under the hood, their sophisticated proprietary operational infrastructure—built by a unique blend of scientists and software engineers—natively manages complex multi-modal biological data ingestion (DNA/protein sequences, cell lines, reagents), instantaneous real-time collaborative-experiment routing, and automated AI-powered insights, providing the necessary foundation for organizations to maintain total data-sovereignty and reproducibility without compromising on performance or functionality. What sets Benchling apart is its uncompromising dedication to frictionless scientific orchestration; by bridging the gap between highly technical, security-intensive biotechnology-enterprise demands and accessible, interoperable lab-management experiences, the firm empowers modern organizations to radically accelerate their discovery velocity, eliminate critical vendor-lock-in and data-silo bottlenecks, and build an unassailable foundation for continuous commercial and institutional independence in the modern, AI-transformed digital biology landscape.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More
Data Engineer, AI & Data Engineering (AIDE) at Benchling | HireSkys