Back to Jobs
Data Science & Analytics 1h ago

Principal Data Engineer

United StatesUnited States
Full-time
Not Disclosed
Lead/Manager

Job Description

Key Skills Required

Master these to land this role

SQLBestseller 🔥
Learn in 9 Hours
PythonBestseller 🔥
Learn in 56 Hours
Data ScientistData EngineeringBig Data

Want to know if you're a match for this job?

Calculate My Match Score

About the Role:

We're looking for a Principal Data Engineer to be the most senior individual contributor on the Data Platform Engineering (DPE) team at Hims & Hers. In this role, you will define and align the architectural vision for our data platform with business goals – working in close partnership with product, engineering, and data science leadership. Your scope is org-wide: you will set technical direction across every surface DPE owns, drive the highest-stakes architectural decisions, and establish the standards that the entire data engineering discipline operates by.

Our platform serves millions of patients across telehealth, prescription, and wellness products. It runs on GCP BigQuery, Airflow on Astronomer/EKS, dbt, Confluent Kafka, Databricks Delta Lake, and Terraform/OpenTofu – and it is in active, consequential evolution: a net-new streaming platform, and a lower environments strategy being built from scratch. You will own those architectural bets.

You Will:

  • Own the long-term technical architecture for DPE across ingestion, orchestration, event streaming, and the platform infrastructure that enables transformation and serving - driving the highest-stakes decisions for the CDC-based streaming platform (Kafka → Flink → BigQuery), orchestration platform evaluation, and lower environment strategy

  • Chair Architecture Review Committee (ARC) decisions; act as the primary technical DRI for cross-team, multi-system, and cost-impacting changes

  • Establish and enforce engineering standards and production readiness criteria across all DPE-owned systems - testing requirements, CI/CD patterns, observability-as-code, logging standards, data contracts, Schema Registry governance, and what 'production-ready' means for emerging streaming and CDC capabilities

  • Own data quality and observability architecture - dbt anomaly detection frameworks, schema validation, data drift alerting, and the platform standards that ensure consumers can trust the data they build on

  • Define the technical strategy for self-service analytics: what platform capabilities enable Analytics Engineering to work independently, what guardrails prevent downstream breakage, and how DPE reduces its bottleneck over time

  • Own evaluation, onboarding, and ongoing governance of DPE-managed tooling - Fivetran, Confluent, and equivalent platforms, including contract management, cost tracking, and deprecation decisions

  • Own data sharing and egress patterns - access provisioning, cross-team data contracts, reverse ETL (Hightouch), and governed consumption paths for internal and external consumers

  • Drive cost governance for platform infrastructure - BigQuery slot reservations, query optimization, partition strategies, orchestration rightsizing, and cloud spend accountability across the full DPE stack

  • Lead incident response for platform-level P1/P2 incidents: act as technical escalation point, facilitate blameless RCAs, and drive systemic fixes that prevent recurrence

  • Produce exemplary technical artifacts - architecture decision records, solution design docs, RFCs - that create alignment and become the team's reference standard

  • Mentor and elevate Staff and Senior Data Engineers; raise the technical ceiling through design reviews, code reviews, and hands-on pairing

  • Partner cross-functionally with ML/Data Science, legal/security/compliance, and DevOps to deliver platform capabilities that are ML-ready, compliant with HIPAA/GDPR, and hardened at the infrastructure layer

  • Contribute hands-on to critical path work

You Have:

  • 15+ years of professional experience designing, building, and owning data platform architecture at company scale

  • Demonstrated ability to define and align architectural vision with business goals across an entire engineering organization

  • Deep expertise in cloud-native data platforms across GCP primary (BigQuery, GCS, Dataflow); AWS operational familiarity required (EKS-based Airflow) - BigQuery strongly preferred; multi-cloud fluency is required, not a plus

  • Hands-on experience with the modern data stack: experience governing dbt at platform scale, Airflow/Astronomer, Kafka/Confluent, Databricks/Spark, Fivetran, and data sharing/activation platforms (Hightouch or equivalent)

  • Experience designing and operating event streaming pipelines at scale - including Schema Registry, data contracts, and consumer lag management

  • Proven track record establishing engineering standards across multiple teams and driving adoption without direct authority

  • Experience owning data quality frameworks - dbt testing, anomaly detection, schema validation, and data observability tooling

  • Experience with data governance and compliance frameworks in a regulated environment - HIPAA/PHI handling, data classification, access controls, audit logging, and GDPR

  • Experience leading incident response for data platform outages - blameless RCA, systemic root cause identification, and operational improvement

  • Infrastructure-as-code fluency - Terraform or equivalent; you treat infrastructure changes like software changes

  • Strong Python and SQL skills; comfortable writing, reviewing, and raising the bar on production-grade pipeline code

  • Clear written communication: you produce design docs and RFCs that build alignment, not confusion. Comfort operating in ambiguity - you define the path, you don't wait for it to be defined

Preferred Qualifications:

  • Databricks, Unity Catalog, and Delta Lake in production at scale

  • Experience with CDC (Change Data Capture) patterns and Flink for real-time data processing

  • PySpark/SparkSQL for large-scale batch and streaming workloads

  • Experience driving a BigQuery → Databricks Lakehouse migration or equivalent cloud data warehouse migration

  • Experience with MLOps - partnering with ML engineers on model training pipelines, feature stores, and experimentation infrastructure

  • Go or Python service development for Kafka producers and consumers

  • Experience at a direct-to-consumer healthcare or telehealth company with HIPAA and GDPR obligations

  • Familiarity with SOX compliance controls in a data engineering context

How would you rate this job post?

See what other professionals think about this role.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More