Staff Platform Engineer
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Role Summary
Gravie is seeking a hands-on Staff Platform Engineer to build the golden roads and paved paths for engineering teams, define the north star architecture for the organization, and design and deliver AI-enabled capabilities. This role serves as the hub for technical direction, ensuring long-term architectural direction and standards converge and radiate outward across teams.
The ideal candidate combines deep distributed systems and AWS experience with real production generative AI experience. They must be equally credible in building paved paths that teams voluntarily adopt and setting a target-state architecture that the entire organization inherits. The role involves establishing standards, reference implementations, and paved roads that engineers actually use—not just architecture documents.
At Gravie, engineers work closely with Product and own outcomes end-to-end, practicing agentic development to increase the scope and leverage of ownership. Engineers use AI agents to develop specs and plans, execute substantial multi-step engineering work, while setting guardrails and owning quality. As agents take on more execution, the bar for problem framing, architectural judgment, system thinking, and risk identification rises.
This role is deliberately split: 40% golden roads and paved paths (hands-on, reference implementations, shared libraries, service templates, scaffolding, pipelines), 30% north star architecture (published decisions adopted and enforced by tooling), and 30% AI solutions (agent systems, retrieval and context design, evaluation, and guardrails). The mix shifts quarterly. The goal is to empower teams to make more decisions independently due to clear direction and paved paths.
Key Responsibilities
- Build paved paths teams build on, including reference implementations, shared libraries, service templates, scaffolding, and pipeline templates, ensuring the recommended way is the easiest.
- Treat paved roads as a product with real users, measuring adoption and time to first success, and fixing parts people route around.
- Implement the first real use case yourself on a real workload with a real team, not just in a demo repository.
- Build distributed systems patterns, including backend services, event-driven workers, job pipelines, APIs, and relational data models.
- Partner with the Platform Engineering team as the primary delivery vehicle, building paved paths, planning into their roadmap, and handing off sustaining ownership.
- Own migration and adoption, including moving existing services onto the path and retiring what it replaces.
- Stay in code review across teams and close to production incidents to identify where the path is failing.
- Define and publish the target-state architecture with clear owners and decision dates, keeping it current as the business changes.
- Own the AWS foundations product, including multi-account structure, identity boundaries, network and data isolation, compute and serverless runtimes, managed data services, and tagging and cost governance.
- Drive standards adoption to completion, from documents into pipelines, with adoption measured across repositories and teams.
- Lead buy, build, and reuse decisions, evaluating models, frameworks, and vendors on accuracy, latency, reliability, privacy, security, and cost.
- Reduce surface area deliberately by consolidating overlapping tools and standards, owning sunset paths for replacements.
- Design and build production-grade AI agent systems, including multi-agent workflows, orchestration layers, and supporting services.
- Design retrieval, context, and memory systems grounding AI outputs, including sensitive data scoping, redaction, and auditing.
- Develop AI-powered decision-support systems meeting accuracy, traceability, and explainability requirements for healthcare and regulated environments.
- Build product-quality user experiences using React and TypeScript where AI capabilities must become understandable and actionable.
- Establish patterns for AI quality assurance, including automated evaluations, regression testing, groundedness checks, and compliance guardrails.
- Establish one standard way to build, run, evaluate, and operate AI systems.
- Act as the connective point across teams, surfacing duplicated effort, picking up decisions with no owner, and aligning teams heading toward the same problem.
- Drive cross-team communication through writing, demos, and working sessions to ensure direction and tradeoffs are understood.
- Build alignment where incentives differ across teams and surface disagreement rather than resolving it quietly.
- Contribute to roadmap and capacity planning, sequencing paved path, foundational, and AI work alongside product commitments.
- Work through the Platform Engineering team to effect change at organization scale rather than standing up competing surfaces.
- Facilitate collaboration between engineering teams, Product, Data, Security, and Infrastructure, including running forums for cross-cutting work planning and unblocking.
- Represent technical direction to leadership and carry business context back to engineering.
- Partner with Product to shape the roadmap and translate ambiguous business problems into architecture.
- Raise the engineering bar through design review, mentorship, and a working decision-record practice, making architectural context available to both engineers and agents.
- Partner with Security, Compliance, and Data on AI governance, access control, and data handling, designed in rather than retrofitted.
Qualifications
- Eight or more years of software engineering experience, including time with architecture scope and an organization-wide remit, and ownership of complex systems from design through production.
- Still shipping: writes and reviews code today, can open a meaningful merge request in an unfamiliar codebase within the first weeks.
- Experience building internal platform capability that other teams voluntarily adopted—can name the paved path or golden road built, how many teams moved onto it, and what they routed around.
- Hands-on experience building and operating production AI applications, not only prototypes, including evaluation, monitoring, failure handling, guardrails, latency, and cost control.
- Deep AWS expertise across a multi-account organization: identity and permission boundaries, networking, compute and serverless runtimes, managed data services, managed model services, and cost and tagging governance. Infrastructure as code in Terraform, CDK, or both.
- A track record of establishing architectural foundations where little existed—specific standards introduced, how they were adopted by teams that did not report to the candidate, and which failed and why.
- Hands-on experience with agentic software development—using AI coding agents to create specs and plans, execute meaningful multi-step work, run and fix tests, validate results, and iterate while owning architecture, quality, and what ships.
- Deep expertise in at least one primary area (Python or JVM backend and distributed systems, or React and TypeScript product engineering), plus ability to work across the full system.
- Experience with retrieval, embeddings, context, memory, evaluation, monitoring, or AI agent architectures and workflows.
- Strong judgment in testing, monitoring, failure handling, privacy, security, access control, and data governance, carried into a regulated or audited environment.
- Experience working with or inside a platform or infrastructure team as the delivery vehicle for organization-wide change, including how ownership was divided and how a parallel platform was avoided.
- Experience serving as a central technical point across multiple teams, with evidence of making those teams faster rather than becoming their approval gate.
- Demonstrated ability to create alignment across disagreeing teams and to contribute to roadmap and capacity planning so foundational work is actually scheduled.
- Clear written communication and effective collaboration across technical and non-technical teams; influence without authority is the primary tool of this role.
Preferred Qualifications
- Healthcare, health insurance, or benefits administration, including ICHRA, claims, eligibility and enrollment, or interoperability standards such as EDI and FHIR.
- Previous experience in a venture-backed high-growth company.
- Internal developer platform or developer experience work, including golden path or paved road programs and adoption measurement.
- Experience with GitLab CI at scale, including shared and templated pipelines.
- Production observability and LLM observability practice, including Datadog.
- Event-driven and streaming architecture, such as Kafka or Flink.
- Practical experience with agent tooling and standards such as Claude Code, the Model Context Protocol, and agent SDKs.
- Experience defining an AI governance or technology radar process alongside a security or compliance partner.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
Train and Evaluate AI Agents in CAD Environments (Freelance)
Mindrift
United StatesTrain and Evaluate AI Agents in CAD Environments (Freelance)
Mindrift
United StatesSenior Engineer – DoD/U.S. Navy Energetics Facility Design and Construction
Eastern Research Group
United StatesSecurity Operations Specialist
HiddenLayer
United StatesMore Openings at Gravie
Explore Top Companies in this Space
SimplyInsured
Health Insurance Technology
BlueCross BlueShield of Tennessee
Health Insurance / Healthcare
Cigna
Healthcare / Insurance / Health Services
OneImaging
Healthcare / Enterprise Software / Digital Health / Insurance Technology
Gravie
View Company ProfileGravie is a health insurance company that aims to simplify the process of finding and enrolling in health insurance plans. By leveraging technology and a consumer-centric approach, Gravie provides individuals, families, and employers with a range of health insurance options and personalized support. The company's platform allows users to compare plans, determine eligibility, and enroll in coverage, making the often complex and confusing process of buying health insurance more accessible and straightforward. Gravie also offers a range of tools and resources to help users navigate the healthcare system, including claims support, provider directories, and wellness programs. Additionally, Gravie's technology-enabled platform enables seamless integration with existing HR systems and benefits administration platforms, streamlining the administration of health benefits for employers. With a focus on innovation, customer satisfaction, and ease of use, Gravie is committed to making high-quality health insurance more affordable and accessible to everyone. The company's approach has resonated with consumers and employers alike, as it continues to expand its reach and build a reputation as a trusted and reliable partner in the health insurance industry. Gravie's mission is to empower individuals and families to take control of their healthcare and make informed decisions about their health and wellness. With its user-friendly platform, comprehensive support services, and commitment to customer satisfaction, Gravie is revolutionizing the way people interact with the healthcare system and making a positive impact on the lives of its users. Gravie's headquarters is located in Minneapolis, Minnesota, and its official website is https://www.gravie.com/careers/. The company operates in the health insurance industry, providing a range of services and support to individuals, families, and employers. Overall, Gravie is a leading provider of health insurance solutions, dedicated to providing high-quality, affordable, and accessible coverage to those who need it most.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.