Senior Data QA Automation Engineer
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
v4c.ai was founded with a clear goal: to make data, AI, and machine learning accessible and impactful for every organization. As a Databricks partner, we deliver end-to-end solutions that transform complex data challenges into strategic outcomes.
Key Responsibilities
- Strategic Framework Design: Architect, build, and scale automated test frameworks from scratch natively within Databricks using PySpark, Python, and SQL.
- Lakehouse Quality Engineering: Design robust automated assertions for Delta Lake tables, including checking data drift, schema evolution, and historical data validation via time-travel functions.
- Enterprise Pipeline Testing: Code complex automated scenarios to validate large-scale batch and real-time streaming data pipelines (Structured Streaming), ensuring source-to-target integrity.
- Governance Validation: Programmatically verify data lineage, audit logs, and access controls implemented via Databricks Unity Catalog.
- CI/CD & DevOps Ownership: Lead the integration of automated data quality tests into enterprise CI/CD pipelines (e.g., Azure DevOps, GitHub Actions), leveraging Databricks Workflows, APIs, or Airflow.
- Technical Leadership & Mentorship: Act as the subject matter expert for data quality; mentor junior team members, establish QA standards, and advocate for data quality principles across engineering teams.
- Performance Assessment: Design and execute automated performance and scalability tests on Spark jobs, large clusters, and complex query optimizations.
Required Skills and Qualifications
- Education: Bachelor’s or Master’s degree in Computer Science, Data Engineering, or a related quantitative field.
- Experience: 8+ years of experience in data engineering, data QA, or software development engineering in test (SDET), with at least 2+ years of dedicated experience architecting test automation in Databricks.
- Expert PySpark & Python: Mastery of Python and PySpark (DataFrames and SQL APIs) for processing and profiling large datasets.
- Advanced Spark SQL: Deep expertise in writing advanced SQL queries, optimization techniques, and understanding Spark query execution plans.
- Advanced Testing Tooling: Hands-on mastery of big-data validation libraries (e.g., Great Expectations, pytest, Delta Live Tables expectations).
- Cloud Infrastructure: Strong operational knowledge of Databricks deployment on a major cloud provider (AWS, Azure, or GCP).
Preferred Qualifications
- Certifications: Databricks Certified Data Engineer Professional or Databricks Certified Machine Learning Professional.
- Streaming Expertise: Experience validating real-time event-streaming architectures (Kafka, Event Hubs, Kinesis).
- Data Ops: Solid understanding of DataOps culture, testing infrastructure as code, and data observability principles.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at v4c.ai
Explore Top Companies in this Space
Tiger Analytics
AI & Data Analytics Consulting / Enterprise Software
HyperDev
Artificial Intelligence / Software Development Tools / Enterprise Software / Low-Code Platforms
XNT
Securities & Commodity Contracts Intermediation / Brokerage / Financial Services / Regulated Technology
unybrands
Technology / Information and Internet / E-Commerce / Digital Platforms
v4c.ai
View Company Profilev4c.ai is an IT services consultancy and systems integrator specializing in enterprise data modernization, cloud data platforms, and artificial intelligence solutions. Founded in 2024 and headquartered in Scottsdale, Arizona, the firm delivers end-to-end solutions centered on Databricks Lakehouse architecture and Dataiku, offering expertise across data engineering, MLOps, and generative AI deployment.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.
