Back to Jobs
Development 3h ago

Senior Site Reliability and Infrastructure Engineer

United StatesUnited States
Full-time
$160,000 - $208,000 USD
Senior-Level

Job Description

Key Skills Required

Master these to land this role

Networking FundamentalsLinux System AdministrationContainerizationKubernetesSRECloud Platforms

Want to know if you're a match for this job?

Calculate My Match Score

We are looking for an experienced Site Reliability and Infrastructure Engineer to join our engineering team.

You will support Counterpart Health’s existing technology infrastructure by reviewing and improving processes, developing automation tools to eliminate toil, and troubleshooting issues as they arise.

You will collaborate with technical leads across engineering disciplines, as well as data scientists and technology professionals, to develop and maintain a modern, scalable infrastructure platform that supports domestic and international workloads across various compute, storage, and networking needs.

We are seeking someone with prior experience deploying and maintaining containerized infrastructure and workloads. Kubernetes competency is highly valued.

As a Senior Site Reliability Engineer, you will:

  • Build systems for declarative application and infrastructure lifecycle management, including continuous deployment, continuous integration, Kubernetes cluster management, and service/workload inventory.
  • Prioritize and troubleshoot infrastructure issues, minimizing downtime and responding to alerts efficiently.
  • Contribute to setting the direction of the Site Reliability Engineering (SRE) team, ensuring goals align with Counterpart Health’s company-wide objectives.
  • Foster a collaborative, high-performance culture that promotes motivation, innovation, and cross-disciplinary teamwork.
  • Streamline and automate infrastructure processes, including delivery pipelines and database changes.

You should get in touch if:

  • You have 5+ years of programming experience and are proficient in at least one of the following languages: Python, Go, or Shell Scripting.
  • You have in-depth knowledge of containerization technologies and orchestration, such as Docker, Containerd, and Kubernetes, along with experience with CNCF-based technologies like Helm, gRPC, and Prometheus.
  • You have experience with public cloud platforms such as GCP, Azure, or AWS.
  • You are knowledgeable in networking fundamentals, including TCP/IP, UDP, firewalls, routing, DNS, and load balancing.
  • You have experience with Linux system administration and a solid understanding of Linux design principles.
  • You understand key SRE concepts, such as monitoring, performance tuning, and automation.
  • You can work autonomously with limited guidance, proactively identifying and solving problems.
  • You have excellent communication and collaboration skills, with the ability to work effectively with cross-functional teams and adapt to new challenges and evolving technologies.

How would you rate this job post?

See what other professionals think about this role.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More