Back to Jobs
Payabl
Data Science & Analytics 1d ago

Data Engineer (Lakehouse & Streaming Data)

Payabl
CyprusCyprus
PolandPoland
PortugalPortugal
Full-time
Not Disclosed
Mid-Level

Job Description

Key Skills Required

Master these to land this role

Python2h 41mFree Trial ✨
Start 10-Day Free Trial
SQL15 minFree Trial ✨
Start 10-Day Free Trial
Apache IcebergAWSPySpark

Want to know if you're a match for this job?

Calculate My Match Score

payabl. empowers businesses to grow through payments innovation and banking services. Our ambition is to expand our strong portfolio of global financial services we provide to businesses and make them all available in one place on our platform we call payabl.one. As a licensed financial company with principal membership with card schemes, we specialize in global payments and providing businesses with multi-currency accounts.

The Role:

You will play a vital role in the Data Team, collecting, analyzing, and interpreting data to support decision-making across the organization. Your tasks include data collection, ensuring quality, and developing databases. Skilled in SQL, Python for data manipulation, and visualization tools, you will collaborate with various teams to understand their needs and provide actionable insights. Passionate about staying updated with new technologies, you will drive innovation and contribute to a culture of continuous improvement.

Key Responsibilities:

Lakehouse Architecture and Data Platform Design

  • Design, build, and maintain scalable data lakehouse solutions on AWS using S3, Apache Iceberg, and AWS Glue Catalog as core platform components.
  • Contribute to the evolution of the medallion architecture, ensuring that bronze, silver, and gold layers are reliable, performant, and aligned with business needs.

CDC and Streaming Data Ingestion

  • Build and support real-time and near-real-time ingestion pipelines that stream operational data from on-premise databases into AWS using Debezium, Kafka, Kafka Connect, and Iceberg sinks.
  • Monitor and troubleshoot streaming pipelines, including connector issues, schema changes, data consistency, and ingestion reliability.

Data Modelling and Business-Ready Layers

  • Design and implement silver and gold layer datasets that transform raw operational data into trusted, business-ready data products.
  • Work closely with analysts, analytics engineers, and business stakeholders to understand requirements and create reusable, well-structured data models.

Batch and Distributed Processing

  • Develop and maintain batch and distributed processing jobs using PySpark on AWS EMR and AWS Glue.
  • Optimize data transformation jobs for performance, scalability, reliability, and cost efficiency.

Data Integration and Orchestration

  • Build and maintain data workflows using Apache Airflow for API ingestion, batch processing, and orchestration across the medallion architecture.
  • Support integrations from external and third-party systems using tools such as Airbyte.

Data Quality, Governance, and Reliability

  • Implement data quality checks, validation processes, and reconciliation logic to ensure trusted and consistent datasets across the platform.
  • Contribute to data governance practices around documentation, ownership, lineage, access, and compliance.

Infrastructure and DevOps Collaboration

  • Collaborate on AWS infrastructure and deployment practices, working with services such as S3, Glue, EMR, IAM, EKS, and related platform components.
  • Support infrastructure-as-code workflows using Terraform and Terragrunt in collaboration with engineering and infrastructure teams.

Requirements:

  • Minimum 3+ years of experience in data engineering or related roles.
  • Strong experience with SQL and data modeling for analytics and reporting use cases.
  • Strong programming experience with Python, ideally including PySpark.
  • Experience designing and maintaining ETL/ELT pipelines in production environments.
  • Experience with real-time or near-real-time data ingestion.
  • Experience working with Kafka or similar streaming technologies.
  • Experience with CDC concepts and tools such as Debezium.
  • Experience with data lake or lakehouse architectures on cloud platforms.
  • Hands-on experience with AWS data services such as: S3, Glue Catalog, Glue Jobs, EMR, IAM, and related AWS services.
  • Experience with Apache Iceberg, Delta Lake, or similar open table formats.
  • Experience designing curated analytics layers, such as silver and gold layers in a medallion architecture.
  • Experience with orchestration tools such as Apache Airflow or Dagster.
  • Experience working with relational databases such as MySQL, MariaDB, or PostgreSQL.
  • Working knowledge of Unix/Linux environments and shell scripting.
  • Understanding of data quality, governance, lineage, and production monitoring concepts.

Good to Have:

  • Experience with Apache Druid, ClickHouse, Snowflake, Databricks, or other analytical databases and warehouse platforms.
  • Experience with Airbyte or similar data integration tools.
  • Experience with dbt or collaboration with analytics engineering teams.
  • Experience with Terraform and Terragrunt.
  • Experience with Docker and Kubernetes.
  • Experience optimizing Spark jobs on EMR or AWS Glue.
  • Experience with data visualization tools such as: Tableau, Power BI, Superset, AWS QuickSight.
  • Experience with data observability, alerting, and monitoring tools.

How would you rate this job post?

See what other professionals think about this role.

banner

Payabl (operating at payabl.com) is a financial services platform engineered for streamlining and automating payment processes within the business-to-business (B2B) sector. Founded in an unspecified year, the company appears to specialize in providing a modern solution to the complexities of invoice management, payment tracking, and cash flow optimization. Under the hood, Payabl likely leverages cloud-based technology and integration APIs to connect with accounting systems, payment gateways, and enterprise resource planning (ERP) tools, ensuring seamless workflows. This allows businesses to reduce manual errors, accelerate payment cycles, and gain real-time visibility into their financial operations. While specific details about funding, investors, or founders remain undisclosed, the platform’s focus on operational efficiency suggests it targets mid-sized to large enterprises seeking to digitize their payment processes.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More