Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
We’re looking for a Staff Data Engineer to join our growing Data Platform team. You’ll play a key role in designing and scaling the infrastructure and pipelines that power analytics, machine learning, and business intelligence across Sonatype. You’ll work closely with stakeholders across product, engineering, and business teams to ensure data is reliable, accessible, and actionable. This role is ideal for someone who thrives on solving complex data challenges at scale and enjoys building high-quality, maintainable systems.
- Use data with purpose: you'll get the chance to work on problems that directly impact how the world builds secure software
- Use modern tooling: you'll get the chance to leverage the best of open-source and cloud-native technologies
- Have a deep collaborative culture: you'll be joining a passionate team that values learning, autonomy, and impact
What you'll do:
Design, build, and maintain scalable data pipelines and ETL/ELT processes
Architect and optimize data models and storage solutions for analytics and operational use
Collaborate with data scientists, analysts, and engineers to deliver trusted, high-quality datasets
Own and evolve parts of our data platform using Databricks and Spark
Implement observability, alerting, and data quality monitoring for critical pipelines
Drive best practices in data engineering, including documentation, testing, and CI/CD
As a Staff Engineer you will help drive long-term architectural vision and mentor the team on engineering best practices, while partnering with stakeholders to ensure data solutions support business outcomes.
Contribute to the design and evolution of our next-generation data lakehouse architecture
What you bring:
8+ years of experience as a Data Engineer or in a similar backend engineering role
Bachelor’s degree in Computer Science, Engineering, or a related technical field
Databricks Optimization: Tune Spark jobs, optimize join performance, and manage Delta Lake architecture for batch and streaming data.
Experience leveraging AI-assisted development tools and AI/ML technologies to improve data engineering workflows, developer productivity, data quality and ops.
Strong programming skills in Python, Scala, or Java
Hands-on experience with distributed data systems like Spark or Kafka
Proficient in writing complex SQL and NoSQL queries and optimizing queries for performance
Experience building and maintaining robust ETL/ELT pipelines in production
Understanding of data modeling techniques (star schema, dimensional modeling, etc.)
It’d be great if you also had:
Familiarity with software supply chain, cybersecurity, or large-scale software ecosystem data
A track record of improving data platform reliability, scalability, performance, and cost efficiency
Familiarity with workflow orchestration tools (Airflow, Dagster, or similar)
Hands-on experience with cloud data platforms, particularly AWS
Familiarity with modern table formats such as Delta Lake, Apache Iceberg, or Apache Hudi
Experience implementing data observability, lineage, governance, and automated data quality frameworks
Experience designing real-time or streaming data architectures using data lake technologies
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Sonatype
Explore Top Companies in this Space
ArmorCode
Application Security Posture Management (ASPM) / Agentic AI Unified Exposure Management / Software Supply Chain & DevSecOps Governance
Apiiro
Application Security Posture Management (ASPM) / Software Supply Chain Security (SSCS) / DevSecOps Cloud Security Automation
Blue Bottle Coffee
Grocery Retail / Retail / Coffee industry
Acquisition.com
Private Equity / Venture Capital / Business Services
Sonatype
View Company ProfileSonatype is a pioneering developer security and software supply chain management platform that empowers engineering and DevSecOps teams to automate open-source governance and secure their software development lifecycles. Founded in 2008 by Jason van Zyl—the creator of Apache Maven—and Brian Fox, the company originated from the development of Maven Central, the world's largest repository of Java components. Sonatype leverages deep intelligence and AI-powered automation to analyze open-source software (OSS) dependencies, identifying code vulnerabilities, license compliance risks, and malicious software before deployment. Its core product suite, led by the Nexus Platform—including Nexus Repository, Nexus Lifecycle, and Nexus Firewall—provides end-to-end visibility and continuous enforcement across CI/CD pipelines. Headquartered in Fulton, Maryland, Sonatype serves thousands of global enterprises, including major banks, technology firms, and government agencies, helping them accelerate software delivery without compromising safety or security.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.

