Senior Data Engineer
United StatesJob Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
We are seeking a seasoned, hands-on Senior Data Engineer to serve as the subject matter expert and technical anchor for our next-generation Databricks Lakehouse platform. In this role, you will design and build scalable ingestion pipelines, establish production best practices, set up CI/CD standards, and elevate the teamβs technical capabilities.
Our current environment is hosted on AWS, but we welcome strong Databricks experts with background in Azure or other major cloud providers, as your core Lakehouse mastery is what matters most. We need a confident, self-directed engineer with deep production experience across Databricks and Apache Spark who can tackle unstructured data sources, optimize distributed workloads, and drive decisive architectural standards in a fast-paced rebuild environment.
Key Responsibilities
Platform Standards & Best Practices: Lead the architecture, design, and implementation of robust Lakehouse ingestion and transformation pipelines using Databricks and PySpark. Establish standards that level up the broader team's engineering maturity.
CI/CD & Operational Excellence: Set up automated delivery pipelines using Databricks Asset Bundles (DABs) and modern orchestration (Airflow or similar) to ensure seamless, reliable deployments.
Complex Data Integration: Architect pipelines to ingest, process, and structure both structured (OLTP) and unstructured/semi-structured data (APIs, logs, cloud storage streams) into high-quality analytical datasets.
Delta Lake & Governance: Enforce Delta Lake principles (CDC, schema evolution, optimization) and integrate data quality frameworks (e.g., Great Expectations) and Unity Catalog governance into daily development.
AI Agent Direction: Direct and evaluate AI agents for rapid code generation, ensuring generated code meets strict quality, security, and performance standards.
Technical Mentorship & Execution: Clearly articulate complex technical trade-offs to stakeholders, mentor peers, and confidently drive pragmatic solutions without needing step-by-step guidance.
Qualifications
Deep Databricks & Spark Expertise: 5+ years of dedicated data engineering experience with proven, large-scale production exposure to Databricks, PySpark performance tuning, and distributed processing architectures.
Cloud Ecosystem Experience: Strong proficiency in major cloud environments (AWS primary, but candidate expertise in Azure or GCP with Databricks is fully welcomed and transferable).
Advanced CI/CD Knowledge: Hands-on experience implementing enterprise CI/CD for Databricks workflows, specifically utilizing Databricks Asset Bundles (DABs) or equivalent infrastructure-as-code patterns.
Unstructured & Diverse Data Processing: Demonstrated history of turning raw, unstructured, or semi-structured data sources into structured, production-ready Lakehouse layers.
Production Discipline: Expert SQL optimization, deep Python proficiency, and practical experience with workflow orchestrators (Airflow) and quality/governance tooling (Unity Catalog, Great Expectations).
Senior Communication: Concise, direct communicator capable of explaining technical concepts clearly and mentoring team members on best practices.
Pragmatic & Autonomous: Comfortable tackling ambiguous, legacy datasets in a fast-paced rebuild environment; excels at guiding AI-assisted coding tools while enforcing high stability standards.
How would you rate this job post?
See what other professionals think about this role.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.