Senior Platform Engineer
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Role Summary
Gravie is seeking a Senior Platform Engineer to build and operate the platform our engineering teams ship on. AWS should be second nature: deep enough to reach for it before documentation, and you should be as comfortable working alongside AI agents as in a terminal.
The ideal candidate has genuinely operational AWS expertise, has worked in or closely with a platform team before, and knows a platform only succeeds when engineers choose it. They will own shared infrastructure, pipelines, and tooling end to end, partner with our architecture team on turning direction into paved paths teams actually adopt, and treat the engineers who use the platform as the customers they are.
At Gravie, engineers work closely with Product and own outcomes end to end, and we practice agentic development to increase the scope and leverage of that ownership: engineers use AI agents to help develop specs and plans, then to execute substantial multi-step engineering work, while setting guardrails and acceptance criteria and owning the quality of everything that ships. As agents take on more of the execution, the bar on problem framing, architectural judgment, system thinking, and risk identification gets higher, not lower.
How Our Platform Team Works
If you have worked on a platform team, this will be familiar. If you have not, read it carefully, because it is the job.
The platform is a product and engineers are its users. Adoption is the measure of success, not tickets closed. If teams route around something we built, that is our problem to solve, not their failure to comply.
Self-service over gatekeeping. If the answer to a common request is a ticket to us, that is a bug in the platform.
We pave roads, we do not police them. Standards belong in pipelines, templates, and defaults, not in reminders and review comments.
Small team, wide surface. Leverage matters more than heroics, which is why AI tooling and automation are not side projects here - they are how the surface stays covered.
We operate what we ship. On-call for platform services is ours. So is the postmortem, and so is fixing the class of problem rather than the instance.
Documentation is a deliverable. A capability nobody can find or adopt without asking us is not finished.
Key Responsibilities
Build and operate the shared platform: CI/CD pipelines, service templates, deployment tooling, and the shared infrastructure services teams depend on.
Own AWS infrastructure as code across multiple accounts, in Terraform, CDK, or both, held to the same standard as application code, including review, testing, and a real rollback story.
Implement and sustain the paved paths teams build on, partnering with architecture on direction while owning the operational reality once a path is live.
Operate what you build: on-call for platform services, incident response, the postmortem, and fixing the class of problem rather than the instance.
Build observability in by default: metrics, logs, traces, SLOs, and alerts that page a human only when a human is actually needed.
Own platform security posture alongside Security: identity and permission boundaries, secrets handling, network and data isolation, and remediation of findings in a HIPAA and SOC 2 environment.
Make cloud spend attributable and defensible through tagging discipline, cost visibility, and removing waste before it compounds.
Use AI agents as a normal part of daily workflow, and build them into the platform where they create leverage, including code review, alert triage, migrations, and repetitive remediation.
Automate toil out of the platform, reducing the number of requests that require a person at all, and measure whether it worked.
Treat documentation and developer onboarding as part of the deliverable, so teams can adopt platform capability without asking first.
Drive adoption of platform capability and standards, including the unglamorous migration work needed to move existing services.
Raise the engineering bar through code review and mentorship, and help other engineers get more out of AI tooling.
Qualifications
Five or more years of software or infrastructure engineering experience, including ownership of production systems that you also operated.
AWS to a depth that is genuinely second nature - reasoning about IAM policy evaluation, VPC routing and DNS, and container or serverless runtime behavior without reaching for documentation first, and having been the person who diagnosed it under real pressure.
Hands-on ownership of a multi-account AWS environment, including identity boundaries, networking, managed data services, and cost and tagging governance.
Infrastructure as code as your default rather than your fallback, in Terraform, CDK, or both, including how you review it, test it, and reason about blast radius before applying it.
Strong CI/CD engineering, including shared or templated pipelines consumed by more than one team.
Experience working in or closely with a platform, infrastructure, or developer experience team, and a clear point of view on what makes a platform get adopted rather than merely tolerated.
Real comfort with AI tooling in daily engineering work - using coding agents such as Claude Code to do meaningful multi-step work, able to describe where they fail and how you catch it, and keeping ownership of what ships.
Production operations experience: on-call, incident response, postmortems, and SLO thinking.
Strong programming ability in at least one language such as Python, Go, TypeScript, or a JVM language, enough to build tooling rather than only configure it.
Sound judgment on security, secrets, access control, and data handling, ideally in a regulated or audited environment.
Clear written communication - on a platform team, documentation and a well-argued change proposal are load-bearing, not overhead.
Preferred Qualifications
Healthcare or another regulated domain, including HIPAA and SOC 2 obligations.
Previous experience in a venture-backed high growth company.
GitLab CI at scale, including templated pipelines with typed inputs consumed across many repositories.
Datadog in depth: monitors, dashboards, SLOs, and keeping observability cost under control.
Container orchestration at production scale, such as ECS or Kubernetes.
Building AI tooling rather than only using it, such as MCP servers, Claude Code plugins or skills, or agent harnesses.
Dependency and currency automation, such as Renovate, and the lifecycle policy around it.
Delivery metrics work, including DORA, and the definitional arguments that come with it.
Familiarity with Team Topologies vocabulary and where platform teams fit alongside stream-aligned and enabling teams.
Experience decommissioning a platform capability, not only launching one.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
Train and Evaluate AI Agents in CAD Environments (Freelance)
Mindrift
United StatesTrain and Evaluate AI Agents in CAD Environments (Freelance)
Mindrift
United StatesSenior Engineer – DoD/U.S. Navy Energetics Facility Design and Construction
Eastern Research Group
United StatesSecurity Operations Specialist
HiddenLayer
United StatesMore Openings at Gravie
Explore Top Companies in this Space
Cigna
Healthcare / Insurance / Health Services
Oscar
Health Insurance
Sana Benefits
InsurTech / Health Insurance / Employee Benefits
OneImaging
Healthcare / Enterprise Software / Digital Health / Insurance Technology
Gravie
View Company ProfileGravie is a health insurance company that aims to simplify the process of finding and enrolling in health insurance plans. By leveraging technology and a consumer-centric approach, Gravie provides individuals, families, and employers with a range of health insurance options and personalized support. The company's platform allows users to compare plans, determine eligibility, and enroll in coverage, making the often complex and confusing process of buying health insurance more accessible and straightforward. Gravie also offers a range of tools and resources to help users navigate the healthcare system, including claims support, provider directories, and wellness programs. Additionally, Gravie's technology-enabled platform enables seamless integration with existing HR systems and benefits administration platforms, streamlining the administration of health benefits for employers. With a focus on innovation, customer satisfaction, and ease of use, Gravie is committed to making high-quality health insurance more affordable and accessible to everyone. The company's approach has resonated with consumers and employers alike, as it continues to expand its reach and build a reputation as a trusted and reliable partner in the health insurance industry. Gravie's mission is to empower individuals and families to take control of their healthcare and make informed decisions about their health and wellness. With its user-friendly platform, comprehensive support services, and commitment to customer satisfaction, Gravie is revolutionizing the way people interact with the healthcare system and making a positive impact on the lives of its users. Gravie's headquarters is located in Minneapolis, Minnesota, and its official website is https://www.gravie.com/careers/. The company operates in the health insurance industry, providing a range of services and support to individuals, families, and employers. Overall, Gravie is a leading provider of health insurance solutions, dedicated to providing high-quality, affordable, and accessible coverage to those who need it most.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.