Senior Data Engineer (Remote) - Clinical Trial Software & Databricks
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
About Us
At Muttdata, we build innovative Data Products and Machine Learning solutions that help companies solve complex business challenges. As a fast-growing, remote-first startup, we're passionate about technology, collaboration, and continuous learning.
🚀 What We Do
We specialize in:
- Leveraging Machine Learning systems for demand planning and budget forecasting.
- Developing scalable data infrastructures to enhance high-level decision-making for clients.
- Offering Data Engineering and custom AI solutions to optimize cloud-based systems.
- Using Generative AI to help e-commerce platforms and retailers create higher-quality ads, faster.
- Building deep learning models for visual recognition and automation in industries like product categorization, quality control, and information retrieval.
- Developing recommendation models to personalize user experiences in e-commerce, streaming, and digital platforms, driving engagement and conversions.
🌟 Our Partnerships
We collaborate with:
- Amazon Web Services
- Astronomer
- Databricks
🤓 Responsibilities
- Design, build, and optimize enterprise data pipelines, lakehouse storage layers, and data models using Databricks (PySpark, Spark SQL, Delta Lake) to power custom clinical application backends.
- Collaborate with frontend developers, software architects, and clinical research teams to build API-driven endpoints, data ingestion engines, and query layers for proprietary clinical trial software.
- Build performant, standards-compliant data structures to store EDC outputs, audit trails, device telemetry, and patient-reported outcomes, enabling rapid querying and downstream analytics.
- Partner with Clinical QA and Validation teams to ensure compliance with GxP, 21 CFR Part 11, HIPAA, and GDPR.
- Implement real-time and batch ingestion jobs connecting legacy clinical systems, central labs, EHRs, and wearable devices into a unified Databricks Lakehouse architecture.
- Monitor, troubleshoot, and optimize Spark jobs, Delta Lake tables, and query execution times to support high-throughput, low-latency clinical platform workflows.
💻 Required Skills
- 4+ years of hands-on experience building production data pipelines and lakehouse architectures using Databricks, Delta Lake, and Apache Spark (PySpark or Scala).
- Experience building, extending, or maintaining custom software applications for clinical trials (e.g., custom EDC, CTMS, Clinical Data Repositories, or eCOA/ePRO platforms).
- Deep understanding of clinical data standards and regulatory environments, including CDISC (SDTM, ADaM, CDASH), 21 CFR Part 11, GxP validation, and ICH-GCP guidelines.
- Strong experience with relational schema design, dimensional modeling, and unstructured data handling within Delta Lake environments.
- Proficiency in Python, SQL, RESTful API integrations, CI/CD pipelines, Git, and automated testing frameworks.
- Experience working in cloud environments (AWS preferred, Azure or GCP).
- Advanced English to discuss technical requirements and solutions with clients in the United States.
💻 Nice to Have
- Bachelor's or Master's degree in Computer Science, Data Engineering, Bioinformatics, or a related quantitative field.
- Experience with Databricks Workflows, Delta Live Tables (DLT), and Unity Catalog governance.
- Background working in a validated system environment (Computer System Validation / CSV).
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Muttdata
Explore Top Companies in this Space
Datatonic
IT Services / IT Consulting / AI & Machine Learning Operationalization (MLOps) / Enterprise Software
Motia
Fleet Management / Fuel Card Solutions / Telematics / Enterprise Software
apply2day
Human Resources / Recruitment / Job Listings / Talent Acquisition
M+R Strategic Services
Public Relations and Communications Services / Non-Profit & Charitable Organizations / Fundraising / Advocacy
Muttdata
View Company ProfileMuttdata (operating at muttdata.ai) is a modern data platform and AI consulting company engineered for businesses seeking to harness the power of scalable machine learning and data science. Headquartered in Buenos Aires, Argentina, Muttdata specializes in helping startups and enterprises build and deploy AI-driven solutions that transform raw data into actionable business outcomes. Unlike traditional data consultancies, Muttdata focuses on architecting end-to-end AI systems using Databricks and cloud-native frameworks, ensuring seamless integration with existing workflows. This approach allows clients—spanning industries from fintech to retail—to accelerate innovation, improve decision-making, and achieve measurable growth through AI. Backed by a $4 million funding round led by Resultiv Capital and including Databricks Ventures, Muttdata combines technical expertise with a client-centric philosophy to deliver scalable, high-impact AI solutions.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.


