Back to Jobs
Development Just now

Application Site Reliability Engineer (SRE)

ArgentinaArgentina
ChileChile
ColombiaColombia
MexicoMexico
PeruPeru
UruguayUruguay
Full-time
Not Disclosed
Mid-Level

Job Description

Key Skills Required

Master these to land this role

PythonBestseller 🔥
Learn in 56 Hours
Grafana.NETPowerShellBashWindowsAurora PostgreSQLPrometheusAWSTerraform

Want to know if you're a match for this job?

Calculate My Match Score

About the Role

Our trading platform powers every customer interaction, making reliability a first-class product concern. You will be responsible for maintaining and improving the operational reliability of our .NET/C# services on Windows, ensuring they remain highly available, observable, and resilient.

You'll collaborate closely with software engineers to improve monitoring, deployment safety, automation, fault isolation, and incident response, while driving continuous improvements in platform reliability and operational excellence.

What You'll Do

  • Participate in the on-call rotation for production trading systems and lead incident response during service disruptions.
  • Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues.
  • Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases.
  • Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact.
  • Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Troubleshoot issues across:
    • .NET/C# applications
    • Windows Server
    • Aurora PostgreSQL databases
    • AWS infrastructure
    • CI/CD pipelines and deployments
  • Improve deployment safety, release automation, and rollback strategies.
  • Partner with developers to improve application operability, resilience, and fault isolation.
  • Automate operational tasks through scripting and infrastructure automation.
  • Create and maintain runbooks, operational documentation, and incident response procedures.
  • Continuously improve monitoring, alert quality, automation, and platform reliability.

Requirements

Required Technical Skills

.NET & Windows

  • Strong experience debugging and supporting .NET/C# applications in production.
  • Hands-on experience with Windows Server environments.

Scripting & Automation

  • Strong PowerShell scripting skills.
  • Experience with Python or Bash.

Observability

  • Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools).
  • Solid understanding of metrics, logging, tracing, and alerting best practices.

CI/CD & DevOps

  • Experience with modern CI/CD pipelines.
  • Knowledge of deployment strategies, release automation, and rollback mechanisms.

Cloud & Infrastructure

  • Experience working with AWS.
  • Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools.

Databases

  • Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms.

Reliability Engineering

  • Practical experience with:
    • SLIs & SLOs
    • Error Budgets
    • Incident Response
    • Root Cause Analysis (RCA)
    • Alert Design
    • Production Operations

How would you rate this job post?

See what other professionals think about this role.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More