Back to Jobs
Y CombinatorDevelopment 1d ago

Field Reliability Engineer (FRE)

Honeycomb
United StatesUnited States
Full-time
$200,000 — $240,000 USD
Senior-Level

Job Description

Key Skills Required

Master these to land this role

DevOpsBestseller 🔥
Learn in 63 Hours
AWSOpenTelemetrySREKubernetes

Want to know if you're a match for this job?

Calculate My Match Score

What You’ll Do

Platform Engineering — Architecture & Standards Ownership

  • Define the architecture and operational standards for Refinery as a Service (RaaS) and Honeycomb Private Cloud (HnyPC) — decisions other engineers build within — across multiple AWS accounts and regions.
  • Architect the Terraform modules, Helm charts, and deployment automation that other FREs build on and extend, not just consume.
  • Set the technical direction for how Honeycomb instruments, monitors, and operates its own managed infrastructure — using Honeycomb to monitor Honeycomb.
  • Own capacity planning, scaling strategy, upgrade sequencing, and cost optimization across multi-region AWS environments.
  • Build platforms and automation that change how the FRE team operates at scale — enabling the team to grow without proportional headcount.

Technical Escalation & Unblocking

  • Serve as the final technical escalation point for the most novel, highest-stakes customer situations — problems with no precedent in existing runbooks.
  • Resolve deep infrastructure and observability issues spanning distributed systems, Kubernetes clusters, AWS networking (ALBs, PrivateLink, NLBs, VPCs), and polyglot service meshes, often in real time under revenue-critical pressure.
  • Partner directly with customer SRE, platform, and engineering leadership to navigate multi-week escalations and architecture redesigns tied to the company's largest relationships.
  • Provide senior incident command for managed services (RaaS, HnyPC) — the point of last escalation when Tier 2 support and IC3/IC4 engineers need staff-level judgment.
  • Build the playbooks, diagnostic frameworks, and tooling that let IC3/IC4 engineers operate independently in situations that previously required staff involvement.

Open Source & Ecosystem Leadership

  • Shape Honeycomb's open source strategy in the OpenTelemetry ecosystem — drive initiatives, not just contributions.
  • Represent Honeycomb at the community level in OpenTelemetry SIGs, setting direction on collectors, exporters, and instrumentation libraries the broader ecosystem depends on.
  • Spot whitespace in the OTel ecosystem that creates structural customer friction, and lead the effort to close it — upstream or through Honeycomb tooling.
  • Build reference architectures and integration guides that set the standard for effective instrumentation across common customer environments (Kubernetes, ECS, serverless).
  • Lead feature and architecture contributions to Honeycomb's own open source projects (Refinery, Honeycomb Collector Distro) that support managed service capabilities at scale.

Technical Backstop for the Field

  • Be the final technical authority Solutions Architects call when a deal goes deeper than any existing playbook — join live production troubleshooting, validate architecture decisions, and provide the infrastructure credibility that closes the company's most technically demanding evaluations.
  • Own the infrastructure and data pipeline narrative on the company's most strategic accounts, in partnership with SA leadership.
  • Lead architecture reviews, SLO workshops, and instrumentation deep-dives for the most complex customer environments (multi-cluster Kubernetes, hybrid cloud, high-cardinality workloads) — often advising customer VPs and C-suite directly.
  • Step into the highest-stakes customer-facing POCs and pilots as technical lead, standing up collector pools, configuring Refinery pipelines, and proving out integrations in the customer's actual environment.
  • Drive prioritization at the roadmap level by identifying strategic gaps between Honeycomb's product capabilities and what the company's largest customers need.

Internal Tooling, Mentorship & Cross-Functional Leadership

  • Build internal tools and UIs that change how the FRE function operates — deployment dashboards, rule management interfaces, and monitoring tooling used company-wide.
  • Drive alignment across Solutions Architecture, Customer Success, Support, Product, and Engineering by synthesizing field signal into a coherent picture others can act on.
  • Manage trade-offs across the FRE charter — managed services, technical escalation, open source, and field backstop — making prioritization calls that others follow.
  • Mentor IC3/IC4 engineers on career development, not just technical skills — helping them build the scope, judgment, and stakeholder navigation needed for their next level.
  • Communicate at the executive level with both Honeycomb leadership and customer C-suite when the situation demands it.

What We're Looking For

Required

  • 9+ years in engineering, SRE, infrastructure, DevOps, or equivalent — with demonstrated Staff-level (or equivalent) technical scope and impact, not just tenure.
  • Deep hands-on experience with Kubernetes (EKS strongly preferred) — you've deployed, scaled, and operated production clusters at scale, and set standards for how others do the same.
  • Strong AWS expertise across core services (EC2, EKS, ECS, ALB/NLB, VPC, PrivateLink, IAM, S3, Route53), with fluency in multi-account architecture design, service quotas, and cost optimization strategy.
  • A track record of senior incident command — not just participation in on-call, but owning incident response, triage, and postmortem process improvements at the function level.
  • Infrastructure as Code mastery (Terraform, Helm, Chef, Ansible) — you've architected modules and standards that other engineers build on, not just written your own.
  • Deep observability expertise: structured logging, distributed tracing, metrics, SLOs/SLIs, and the full instrumentation lifecycle, with a record of setting standards for others.
  • Strong command of OpenTelemetry (SDK, Collector architecture, processors, exporters, semantic conventions) or equivalent, including public community contribution and leadership.
  • Proficiency in at least two of: Go, Python, Java, TypeScript/Node.js, .NET — enough to read, instrument, and debug customer code at a deep level.
  • Excellent executive communication skills — equally credible with a customer's staff SRE and their VP or C-suite, and with Honeycomb's own leadership.
  • Demonstrated ability to set direction in ambiguous, high-pressure situations with no precedent — and to build the structure that lets others operate independently afterward.

Nice to Haves

  • Background leading customer-facing engineering functions — solutions architecture, field engineering, or technical consulting — where you've owned the technical answer a deal hinged on.
  • Track record of building platforms and organizational systems, not just tools, that let a team scale without proportional headcount.
  • Public recognition or a leadership role in the CNCF/OpenTelemetry ecosystem — maintainer status, SIG leadership, or conference speaking.
  • Familiarity with Honeycomb or event-based observability approaches.
  • Experience operating telemetry pipelines at scale (OTel Collectors, Refinery/sampling, tail-based sampling, pipeline reliability).
  • Experience with managed SaaS deployments, private cloud offerings, or multi-tenant infrastructure operations.
  • Prior experience at an observability, monitoring, or developer tools vendor.
  • Experience mentoring senior or staff-track engineers on scope and career progression.

How would you rate this job post?

See what other professionals think about this role.

Startup Details

Company Type

YC Backed Startup 🚀

HoneycombHoneycomb
banner

Honeycomb is a cutting-edge technology company that specializes in observability and monitoring solutions for modern software applications. With a strong focus on innovation and customer satisfaction, Honeycomb provides a comprehensive platform that enables developers, engineers, and operators to gain deep insights into their systems, identify performance bottlenecks, and optimize their applications for better reliability, scalability, and user experience. The company's flagship product offers a robust and intuitive interface for exploring, analyzing, and visualizing complex data sets, allowing users to quickly pinpoint issues, troubleshoot problems, and make data-driven decisions to drive business success. By leveraging advanced technologies such as machine learning, artificial intelligence, and cloud computing, Honeycomb empowers organizations to build, deploy, and manage high-quality software applications that meet the evolving needs of their customers, partners, and stakeholders. With a strong commitment to community engagement, open-source collaboration, and customer support, Honeycomb is well-positioned to become a leading player in the rapidly growing observability and monitoring market. As a trusted partner for businesses of all sizes, Honeycomb is dedicated to helping its customers achieve their goals, improve their operations, and stay ahead of the competition in today's fast-paced digital landscape.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More