Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
About The Role
Yuno is looking for a Staff Site Reliability Engineer to set the technical direction for reliability across our infrastructure โ starting with the platform that provisions, deploys, and manages AI agents at scale on AWS, the system powering payments across 190+ countries. The platform is in production and growing, and we need the most senior reliability voice in the room to evolve the architecture and make sure it stays reliable, observable, and ready to scale.
This is not a "maintain what exists" role, and it's not a single-system role. You'll own the reliability strategy โ driving architectural decisions, designing event-driven communication, defining how we measure and defend reliability, and setting the standards other engineering teams build on.
How AI Shows Up in This Role
The platform you own is Yuno's AI agent infrastructure โ provisioning and deploying AI agents at scale, plus the agents that route payments and prevent fraud. Keeping the AI-native layer reliable is the core of the role
AI is our default execution layer: you're encouraged to use AI-assisted tooling across automation, runbooks, incident analysis, and root-cause investigations, and to help define how the wider engineering org adopts it. We care how you use it, not whether you do
Your Contribution Will Be
Reliability strategy and standards โ define the SLO culture, error-budget policy, and incident practices that scale across engineering teams, turning reliability from firefighting into a measurable, org-wide discipline
Platform architecture and evolution โ drive architectural decisions as the platform matures; the deciding voice on choosing technologies, designing systems, and when to evolve the infrastructure
Messaging and event-driven architecture โ design and own the messaging layer for inter-service communication, replacing synchronous patterns with durable, reliable async messaging
Infrastructure and deployment โ own the cloud infrastructure, automate provisioning with IaC, and ensure the platform scales reliably as transaction volume grows
Observability โ build the monitoring, tracing, and alerting that keeps the platform healthy; when something breaks at 3am, your dashboards and alerts should explain why before anyone has to dig
Incident leadership and mentorship โ act as the senior escalation point for the hardest production problems, run blameless postmortems and root-cause analyses that turn into permanent fixes, and raise the reliability bar by mentoring senior and mid-level engineers
Chaos engineering mindset โ continuous fault injection and resilience experiments that surface weaknesses before they turn into incidents, plus identifying and proposing resilience patterns to prevent those failures from reaching production.
What Success Looks Like
Within your first 6โ12 months, you've set the reliability strategy for the platform, driven at least one major architectural evolution (event-driven messaging, streaming reliability, or observability), and engineering teams have adopted the SLO and error-budget framework you defined. You're the person Yuno trusts with the hardest reliability calls.
Skills You Need
Minimum Qualifications
Event-driven architecture and messaging systems โ you've designed and owned systems around message queues (Kafka, NATS, RabbitMQ) and understand at-least-once delivery, consumer groups, dead letters, and backpressure; you've migrated a system from synchronous to async
Deep AWS โ EC2, VPC, IAM, S3, and RDS โ with strong networking fundamentals, since inter-service communication runs over the internal VPC
Infrastructure as Code โ Terraform or Pulumi, reviewed in PRs rather than clicked in consoles
Kubernetes and Docker in production โ container lifecycle, resource limits, health checks, and orchestration at scale
Observability and SLOs โ Datadog fluency or equivalent (dashboards, monitors, APM, distributed tracing), and a track record defining and operating SLOs, SLIs, and error budgets across services
Chaos engineering and resilience testing โ hands-on experience with fault injection, game days, or chaos experiments (Gremlin, Chaos Mesh, AWS FIS, or similar) to harden production systems
Distributed systems debugging โ you've diagnosed async flows and cascading failures in production and can explain what broke and how you fixed it; comfortable coding for automation and tooling (Go, Python, or similar)
Databases โ solid SQL (PostgreSQL) and NoSQL (MongoDB, Redis): when to use each, indexing, replication, and performance tuning
Proven technical leadership โ you've set reliability standards, influenced architecture across teams, and mentored engineers, not just owned your own scope
English โ advanced proficiency, written and spoken
Preferred Qualifications
AI / MLOps infrastructure โ running AI workloads in production (model serving, LLM inference, GPU/resource management, and agent evaluation/observability tools like LangFuse, LangSmith, Braintrust, or MLflow)
Multi-tenant container platforms โ running customer or user workloads in containers (Replit, Railway, Fly.io, or internal PaaS)
Data pipelines and orchestration โ Airflow, Prefect, or similar; data warehouses like Databricks, Snowflake, or BigQuery a plus
Incident management and on-call tooling โ PagerDuty, Opsgenie, or incident.io
Experience in the payments industry
Nice to Have
ECS experience
s6-overlay for container process supervision
Experience with AI agent framework ecosystems
Spanish proficiency
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Yuno
Explore Top Companies in this Space
Shopify
E-commerce
Busha
FinTech / Cryptocurrency / Digital Banking
JustMarkets
FinTech / Online Trading / Financial Services
OSEA Malibu
Cosmetics / Health & Beauty / E-commerce
Yuno (operating at y.uno) is a leading global payment orchestration platform engineered for simplifying payment operations and maximizing revenue. Founded in 2021 by former payments and technology executives from Rappi, Uber, Ali Pay and other leading companies and headquartered in Bogotรก, Yuno does X instead of Y โ it solves the problem of complex payment processes by providing a single platform that integrates every payment method, processor, and fraud prevention system worldwide. Under the hood, Yuno's platform orchestrates the full payment lifecycle โ from checkout to authorization, routing, fraud prevention, settlement, and reconciliation. This allows businesses to simplify their payment operations, maximize revenue, and accelerate global commerce. Backed by a $25 million funding round led by founder Juan Pablo Ortega and a consortium of high-profile investors like DST Global.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.

