Staff Software Engineer, Data Engineering
United StatesJob Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
We are seeking a Staff Software Engineer, Data Engineering to lead the design and development of the production data platform that powers machine learning across Omada.
In this role, you will partner closely with Data Scientists, Applied AI Engineers, Product Engineers, and fellow Data Engineers to identify, design and build trusted, reusable datasets foundations that serve as the foundation for feature engineering, model training, experimentation, and production inference.
Rather than building one-off pipelines for individual models, you’ll create scalable data products and feature pipelines that enable multiple machine learning use cases while ensuring consistency, reliability, and governance across the ML lifecycle.
You will own the technical design of feature datasets—from ingesting raw behavioral, clinical, and operational data through transforming, validating, and publishing production-grade datasets that are reusable across modeling teams.
This role is ideal for someone who enjoys solving complex data problems, designing scalable distributed data systems, and enabling machine learning through well-engineered data foundations.
Key Responsibilities:
- Design, build, and maintain reusable feature datasets that support machine learning use cases including personalization, engagement, risk prediction, churn modeling, recommendation systems, and experimentation.
- Establish self-service foundations that streamline and democratize dataset creation across the data organization.
- Partner with Data Scientists to translate modeling requirements into production-ready feature pipelines, supporting the full model lifecycle from exploration to deployment.
- Identify source data, transformations, and historical windows needed for feature engineering. Help define and build shared, reusable feature definitions across models rather than one-off datasets.
- Balance features freshness, correctness, latency, and computational efficiency when designing data pipelines.
- Build datasets that support both historical model training and future production inference.
- Design and implement batch and streaming pipelines that transform raw healthcare, behavioral, product, and operational data into trusted ML-ready datasets.
- Build reliable data processing systems using Python, SQL, Spark, and modern cloud data platforms.
- Optimize large-scale distributed processing for performance, scalability, and cost.
- Design data pipelines that are modular, testable, observable, and easy to evolve as product requirements change.
- Ensure data quality through testing, anomaly detection, schema validation, and pipeline monitoring.
- Partner with platform teams to support near real-time feature generation where appropriate.
- Improve reproducibility by standardizing feature computation across experimentation and production.
- Support rapid experimentation without sacrificing long-term maintainability.
- Ensure data quality through testing, anomaly detection, schema validation, and pipeline monitoring.
- Establish engineering standards for correctness, documentation, and maintainability.
- Familiarity with feature stores or feature management platforms.
- Familiarity with model training pipelines and MLOps workflows.
Technical Leadership:
- Lead architecture and design discussions for large-scale ML data systems, driving adoption of reusable patterns and platform capabilities across Data Engineering.
- Influence technical direction across multiple engineering teams, embedding with Product, Engineering, and business stakeholders (Clinical, Finance, Growth, Enrollment) during early design phases to shape data capture requirements at the source.
- Translate ambiguous business requirements from Business domain SMEs into concrete technical specs, maintaining consistency of business logic and definitions across systems.
- Mentor engineers on distributed data processing, software engineering best practices, and scalable data modeling.
About you:
- 8+ years building large-scale production data platforms and distributed data pipelines.
- Experience designing reusable datasets that power machine learning, experimentation, or advanced analytics.
- Demonstrated experience partnering closely with Data Scientists to productionize feature engineering workflows.
- Experience leading cross-team technical initiatives and influencing engineering direction.
- Strong experience working with cloud-native data platforms such as AWS.
- Experience building production data systems using Databricks, Iceberg, Spark, Redshift, Snowflake, or similar technologies.
- Experience developing reliable batch and streaming data pipelines.
- Experience working with healthcare, behavioral, or other large-scale event data is a plus.
Technical Skills:
- Expert SQL with strong data modeling skills.
- Strong programming skills in Python, Java, or Scala.
- Experience with Apache Spark or similar distributed compute frameworks.
- Experience with Airflow or similar orchestration platforms.
- Experience designing dimensional models, event models, and feature datasets.
- Experience implementing testing, CI/CD, observability, and production monitoring for data pipelines.
- Understanding of Feature Stores and ML data lifecycle concepts.
- Experience with Lakehouse Architecture such as Databricks, Iceberg is a strong plus.
- Experience with streaming technologies such as Kafka, Flink, or Spark Structured Streaming.
- Familiarity with NoSQL Databases (document & graph databases Nepture, Neo4j etc.)
- Understanding of software engineering best practices, distributed systems, and cloud-native architectures.
Communication Skills:
- An exceptional people leader who develops engineers into future technical leaders.
- Comfortable influencing senior executives and cross-functional partners.
- Skilled at balancing business priorities with long-term technical investments.
- Able to communicate complex technical concepts to both technical and non-technical audiences.
- Passionate about building trusted data platforms that enable the business.
Education:
- Bachelor’s degree in Computer Science or a similar discipline preferred.
Technologies we use:
- Ruby on Rails, Redshift, Athena, Postgres, SQL, Python, Apache Airflow, Appflow, S3, SNS, SQS, Kafka, Docker, Kubernetes, AWS infrastructure, Lambda, Serverless, Tableau, Bugsnag, Datadog, GitLabCI, Cursor, OpenMetadata, Databricks
Bonus Points for:
- Experience supporting personalization, recommendation, ranking, or predictive modeling systems.
- Familiarity with model training pipelines and MLOps workflows.
- Experience designing data platforms for experimentation.
- Healthcare industry experience is a plus.
- Experience with Data/AI Governance.
Benefits:
- Competitive salary with generous annual cash bonus
- Equity grants
- Remote first work from home culture
- Flexible Time Off to help you rest, recharge, and connect with loved ones
- Generous parental leave
- Health, dental, and vision insurance (and above market employer contributions)
- 401k retirement savings plan
- Lifestyle Spending Account (LSA)
- Mental Health Support Solutions
- ...and more!
How would you rate this job post?
See what other professionals think about this role.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.