Back to Jobs
Gravie
Data Science & Analytics 2h ago

Senior Data Engineer

Gravie
United StatesUnited States
Full-time
Not Disclosed
Senior-Level

Job Description

Key Skills Required

Master these to land this role

Data EngineerAWSPySparkApache SparkApache Kafka

Want to know if you're a match for this job?

Calculate My Match Score

We're building a near real-time streaming data platform for operational data across the business. We seek a Senior Data Engineer to both build/extend and operate it: you'll build the infrastructure and you'll own the platform in production - latency and throughput SLOs, backpressure under spiky load, replay and backfill, dead-letter triage, and connector failure recovery. This is a hands-on role for someone who wants to build, extend, and operate a live production platform, not just design it and hand it off.

You are self-driven, calm under production pressure, and comfortable owning streaming systems in a regulated environment.

You Will:

Take full-lifecycle ownership of the streaming platform - architecture, implementation, production operations.

  • Run the platform in production: own latency/throughput SLOs, monitoring and alerting (e.g. Datadog) - including replays, per-source backfills, connector-failure recovery, dead-letter triage, and tuning for spiky, batch-driven claims load.
  • Build and extend streaming pipelines that ingest CDC events from operational databases and SaaS sources into canonical, contract-validated form.
  • Transform and enrich in Spark (Structured Streaming or dbt-on-Spark micro-batch), including cross-stream joins that correlate events into unified lifecycle entities.
  • Enable secure, governed data access over PHI: classification, row/column controls, and access policy applied as data is served to consumers (e.g. ABAC).
  • Make it reliable and observable: idempotent/replayable pipeline design, data-quality validation, runbooks, and observability the broader data team can rely on.
  • Provision as code: define the platform (streaming, processing, storage, and catalog services on AWS) in CDK with CI/CD for data pipelines, and right-size for cost against the latency SLO.
  • Partner across teams: work with upstream producers on source changes and contracts, with downstream consumers on access and data needs, and with stakeholders to turn requirements into what the platform delivers.
  • Demonstrate commitment to our core competencies of being authentic, curious, creative, empathetic and outcome oriented.

You Bring:

  • 6+ years building and operating production data systems, including demonstrated ownership of streaming or event-driven pipelines - on-call, incident response, SLOs, runbooks, and recovery, not just development.
  • Deep, production experience with Apache Kafka - partitioning, consumer groups, consumer-lag and broker-health troubleshooting, exactly-once/idempotent semantics, schema registry, and replay/backfill under load.
  • Strong, hands-on Apache Spark experience (PySpark) for streaming and batch transformation in Production.
  • AWS-native data engineering across streaming, processing, storage, and catalog services (e.g. MSK, EMR, Glue, S3, Athena), with infrastructure-as-code - AWS CDK (preferred) or Terraform - CI/CD for data pipelines, and cost awareness.
  • Comfort debugging distributed data pipelines (consumer lag, data skew, backpressure, late/out-of-order events) with observability tooling (e.g. Datadog/Cloudwatch).
  • Expert-level SQL and Python, and experience building and consuming REST APIs.
  • An AI-forward engineering mindset, with demonstrated hands-on use of AI-assisted and agentic development tools, an opinion on where AI adds value (and where it doesn’t), and an understanding of agentic data consumption patterns and needs to act on trusted operational data—including context management, lineage, provenance, permissions, freshness, and low-latency access for agentic discovery.
  • Change data capture and open table formats - CDC (e.g. Debezium) plus Iceberg or Delta Lake: schema evolution, partitioning, and table maintenance.
  • Data contracts, schema governance, and cataloging - schema registries and compatibility rules with dead-letter handling; and familiarity with a technical metastore (e.g. Glue Data Catalog, Unity Catalog) and a governance/discovery catalog (e.g. Atlan, Alation, Collibra).
  • Degree in Computer Science, Information Systems or another quantitative field, and comfort on the command line / a Unix-based OS (we are 100% Mac+Linux at Gravie).
  • Health insurance domain knowledge - HIPAA Protected Health Information (PHI) and governing access to it.
  • Excellent communication skills and demonstrated success in driving results through influence and collaboration.

Extra Credit:

  • Experience with Apache Flink or other stateful stream processors - helpful context but not required.
  • Knowledge of JVM-based languages like Kotlin or Java.
  • Familiarity with serving data to AI and agentic consumers - exposing canonical data as low-latency context or inputs for automated/agentic workloads.
  • Previous venture-backed start-up company experience.

How would you rate this job post?

See what other professionals think about this role.

banner

Gravie is a health insurance company that aims to simplify the process of finding and enrolling in health insurance plans. By leveraging technology and a consumer-centric approach, Gravie provides individuals, families, and employers with a range of health insurance options and personalized support. The company's platform allows users to compare plans, determine eligibility, and enroll in coverage, making the often complex and confusing process of buying health insurance more accessible and straightforward. Gravie also offers a range of tools and resources to help users navigate the healthcare system, including claims support, provider directories, and wellness programs. Additionally, Gravie's technology-enabled platform enables seamless integration with existing HR systems and benefits administration platforms, streamlining the administration of health benefits for employers. With a focus on innovation, customer satisfaction, and ease of use, Gravie is committed to making high-quality health insurance more affordable and accessible to everyone. The company's approach has resonated with consumers and employers alike, as it continues to expand its reach and build a reputation as a trusted and reliable partner in the health insurance industry. Gravie's mission is to empower individuals and families to take control of their healthcare and make informed decisions about their health and wellness. With its user-friendly platform, comprehensive support services, and commitment to customer satisfaction, Gravie is revolutionizing the way people interact with the healthcare system and making a positive impact on the lives of its users. Gravie's headquarters is located in Minneapolis, Minnesota, and its official website is https://www.gravie.com/careers/. The company operates in the health insurance industry, providing a range of services and support to individuals, families, and employers. Overall, Gravie is a leading provider of health insurance solutions, dedicated to providing high-quality, affordable, and accessible coverage to those who need it most.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More