Customer Reliability Engineer
RomaniaJob Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Role Overview:
As a Customer Reliability Engineer, you will own the reliability, availability, and operational excellence of iPiD's production platforms. From monitoring and automation to deployments and incident management, you will build and maintain the practices that keep our services secure, scalable, and running smoothly. Working across Engineering, Product, and Operations, you will support releases, solve complex production challenges, and continuously improve platform resilience.
You will also play a critical role in customer implementations, particularly on-premise deployments, leading technical delivery from installation through to go-live. When issues arise, you will be a trusted technical lead, driving investigation, resolution, and continuous improvement to ensure our customers can depend on iPiD with confidence.
Core Responsibilities:
- Own end-to-end deployment of the product on Kubernetes, including rollout, rolling updates, rollback, scaling, and pause/resume of releases using Kubernetes Deployment primitives that provide declarative updates for Pods and ReplicaSets.
- Ensure product health and reliability through monitoring, logs, traces, dashboards, and SLOs/SLIs with error budgets to guide risk decisions and prioritization.
- Provide customer-facing technical operations — assist customers managing on-prem deployments; deliver technical assistance and health assurance for hosted deployments — acting as the technical lead for customer escalations.
- Maintain per-customer environment knowledge — architecture, configurations, requirements, and customizations.
- Govern configuration consistency across customers and environments using configuration management and GitOps-style single-source-of-truth to reduce drift and enforce compliance automatically.
- Drive continuous operational improvement by codifying resolutions, root causes, and best practices into a knowledge base, automation, and runbooks.
- Partner cross-functionally with Engineering, Product, and Operations teams to support production readiness, release planning, user acceptance testing (UAT), and production go-live sign-off.
Requirements
- 5+ years of experience in infrastructure, DevOps, or Site Reliability Engineering roles, ideally within fintech, financial services, or another regulated environment
- Hands-on experience with Kubernetes and Helm in production.
- Infrastructure as code and configuration management — Terraform, Ansible, or equivalent.
- Experience creating and maintaining CI/CD pipelines.
- Linux and cloud-native security fundamentals.
- Strong coding and scripting skills with a focus on automation and efficiency.
- Customer technical operations — enterprise customer engagement, escalation leadership.
- Must be legally authorized to work in the country where the role is based without requiring current or future visa sponsorship.
Benefits
- Meaningful Impact – Play a key role in shaping the future of trusted cross-border payments and fraud prevention, helping solve critical challenges for financial institutions and businesses worldwide.
- Learn from experienced industry leaders – Join a team with deep expertise across payments, fintech and technology, and gain exposure to global customers and markets .
- Ownership & Growth – Join a fast-growing global fintech where your contributions are visible and valued, with the opportunity to participate in our Employee Stock Option Plan (ESOP) and share in our long-term success.
How would you rate this job post?
See what other professionals think about this role.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.