Software Engineer III - Site Reliability
United StatesJob Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
At MyFitnessPal, we believe good health starts with what you eat. We provide tools, resources and support to enable users to reach their health goals.
We are looking for a Software Engineer III - Site Reliability to join the MyFitnessPal PEAS team. As a member of the PEAS team, you'll have the opportunity to positively impact MyFitnessPal users through your expertise in the reliability and delivery systems that keep the MyFitnessPal ecosystem fast, available, and safe to change. In addition to technical expertise, you'll find that your teammates value collaboration, mentorship, and inclusive environments.
About the team:
The Productivity Engineering: Automation & Self-Service (PEAS) team is responsible for the automation, CI/CD, and self-service platforms that product teams rely on to build, test, and ship features quickly and safely.
The PEAS team is part of the Technology Operations (TechOps) organization, which includes IT, Infrastructure, Security, Reliability, and DevOps/DevEx disciplines. TechOps seeks to enable MyFitnessPal to "ship with confidence" by delivering a secure, reliable, and low-friction runway to build and operate software at scale.
Essential Duties:
As a Software Engineer III on the PEAS team, you will own reliability across our production services and maintain security in the software delivery pipeline. You'll define how we measure and defend reliability, lead us through incidents, and make our CI/CD pipelines both faster and harder to compromise. You will play a key role in delivering an outstanding experience to MyFitnessPal users.
What you’ll be doing:
- Own and evolve our SLI/SLO and error-budget frameworks, and use them to influence prioritization and product decisions
- Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches
- Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue
- Design and operate resilient, scalable infrastructure using Infrastructure as Code (Terraform)
- Manage production Kubernetes and container workloads, including capacity planning and cloud-cost optimization
- Own CI/CD pipelines and safe deployment strategies (canary, progressive rollout, fast rollback)
- Own the security controls that live inside the delivery pipeline — integrating and tuning SAST, DAST, and SCA scanning (for example, in GitHub Actions) so issues surface while code is still in review
- Implement and maintain policy-as-code (for example, OPA/Rego, Kyverno, or Conftest) to block unsafe infrastructure and Kubernetes changes at admission time
- Drive vulnerability triage and remediation SLAs for pipeline- and infrastructure-level findings, prioritizing by real risk
- Partner with our Security Engineer and the broader Security & Reliability disciplines — you own security in the pipeline and collaborate on the rest, rather than duplicating that function
- Participate in and improve the on-call rotation; build the runbooks and automation that make on-call sustainable
- Coach team members and engineers across the org on reliability patterns and operational best practices
Qualifications to be successful in this role:
- 5+ years in site reliability, platform, or infrastructure engineering, with clear senior-level ownership of production systems
- Strong programming skills for automation and tooling (Go, Python, Typescript or similar) — you have experience building software or custom tooling, not just scripts
- Deep, hands-on experience with a major cloud platform (AWS is a plus), Kubernetes, and Infrastructure as Code (Terraform is a plus)
- Proven track record leading incident response and building SLO-driven reliability practices.
- Working fluency with observability tooling (Datadog is a plus)
- Practical experience integrating security into CI/CD pipelines — SAST/DAST/SCA tooling, dependency scanning, or policy-as-code
- Strong understanding of cloud security fundamentals (identity/IAM, least-privilege patterns, policy/guardrails, secrets management)
- The judgment and communication skills to raise a security or reliability finding with a senior engineer and land it as a shared problem to solve, not a fight to win
- Experience with policy-as-code frameworks (especially Kyverno, but tools like OPA/Rego or Conftest are also relevant) enforced at admission time is a plus
- Exposure to regulated or compliance-driven environments (SOC 2, PCI DSS, HIPAA) is a plus
- Chaos engineering or game-day experience is a plus
- Experience supporting B2C/mobile backend environments with high traffic, rapid iteration, and strong reliability needs is a plus
Values you’ll model
- Be Kind and Care — build with empathy; assume positive intent; support teammates and members.
- Live Good Health — champion healthy habits and balance in how we work and what we ship.
- Be Data-Inspired — ground decisions in research and data; measure what matters.
- Champion Change — lean into ambiguity; iterate, learn, and improve continuously.
- Leave it Better than You Found It — raise quality, clarify systems, and document as you go.
- Make It Happen — bias to action; deliver impact with craft and accountability.
How would you rate this job post?
See what other professionals think about this role.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.