Lead Observability Engineer (Remote)
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Some of the world’s most innovative global software and technology companies struggle to find engineering partners capable of stepping into complex environments and immediately driving meaningful outcomes. These teams need more than additional hands—they need senior engineers who can quickly understand an environment, identify the path forward, and execute without constant direction.
Enter EverOps – the premier Embedded Service Provider. We partner directly with customer engineering teams to assess and address mission-critical infrastructure, cloud, and delivery challenges.
The Challenge
EverOps is looking for a Lead Observability Engineer with deep hands-on experience across metrics, logs, and traces at very large scale to lead an observability maturity assessment, and the platform consolidation that follows it, for a high-scale consumer mobile platform.
The current estate spans multiple commercial observability vendors alongside self-managed Prometheus, Thanos, Grafana, Vector, and ELK. It carries tens of millions of active time series, tens of thousands of scrape targets, well over a thousand dashboards, and more than ten thousand alert definitions. Much of the log and trace data is sampled or dropped for cost reasons, ownership tagging is sparse, and the self-managed components carry an operational load that crowds out improvement.
The direction under evaluation is consolidation onto AWS-native observability services built on OpenTelemetry. This role requires someone who knows these systems well enough to price that move honestly, prove or disprove performance parity, and then lead the migration if the answer is go.
The Mission
As a Lead Observability Engineer, you will join our U.S.-Based Virtual Operating Center and lead an embedded TechPod as its Pod Leader: a player-coach who sets technical direction, serves as the primary point of contact for the customer’s engineering leadership, and stays hands-on in the work.
Your immediate priority is leading a two-month Observability Maturity Assessment. You’ll validate telemetry volumes, retention, sampling, and cardinality; build a total cost of ownership model; design an AWS-native target architecture; and deliver a migration plan and commercial recommendation that leadership can make a go/no-go decision on.
As discovery closes, your focus shifts to leading the platform migration: standing up the target collection pipeline and storage tiers, porting log transforms to OpenTelemetry, rebuilding dashboards and alerts, and establishing policy-driven retention and ownership tagging that make observability both less expensive and more useful.
The customer’s observability team is capable but stretched thin. You will be expected to add capacity rather than consume it: pull the data yourself, ask targeted questions, keep recommendations unbiased, and leave the team with fewer systems to run and better telemetry than they have today.
What You’ll Do
Estate Assessment: Build a complete picture of the observability estate across vendors, agents, collectors, query surfaces, data volumes, and operating model, grounded in live measurement rather than questionnaires alone.
Telemetry Analysis: Validate metric cardinality, active series, scrape target health, log volumes, trace sampling, and retention, and identify where fidelity is being lost today and why.
Cost Modeling: Build a total cost of ownership comparison between the current state and two to three costed target states, covering vendor spend, self-managed infrastructure, data transfer and egress, and AWS Pricing Calculator estimates.
Target Architecture: Design the AWS-native target across collection (ADOT / OpenTelemetry Collector), pipeline (Amazon Data Firehose), storage tiering (CloudWatch, S3, Amazon Managed Service for Prometheus), query surfaces (Amazon Managed Grafana, CloudWatch, Athena, OpenSearch), and alerting.
Retention & Data Classification: Partner with Security, Legal, and Engineering to classify telemetry data and design policy-driven retention, including long-term compliance archives.
Capability & Gap Analysis: Map current capabilities to AWS-native equivalents and make clear recommendations where no equivalent exists, such as continuous profiling.
Performance Parity: Define and run tests that show whether the target state holds query performance, alert latency, and data fidelity against success criteria agreed with engineering leadership.
Migration Planning & Execution: Size and sequence the migration, then lead it, including pipeline cutover, porting log transforms to OpenTelemetry, rebuilding dashboards, and deduplicating and rebuilding alerts.
OpenTelemetry Standards: Define the OpenTelemetry conventions (semantic conventions, resource attributes, collector topology, sampling strategy) that application teams will adopt as instrumentation moves over.
Ownership & Cost Attribution: Rebuild service ownership tagging so telemetry cost can be attributed back to the teams generating it.
Commercial Analysis: Reconcile platform spend, analyze licensing and commit structures, and inform vendor renewal strategy with a clear, unbiased recommendation.
Technical Leadership: Lead the TechPod as a player-coach, run the working cadence with the customer’s observability leadership, and keep the engagement light-touch on a busy internal team.
Documentation & Readouts: Produce the assessment report, target architecture, cost model, migration plan, and executive summary, and present findings and tradeoffs to engineering leadership.
You Have
Experience: 8+ years in SRE, DevOps, Observability, or Platform Engineering, including 4+ years owning production observability platforms and prior experience in a technical lead, staff, or principal-level role.
Observability at Scale: Deep experience running metrics, logging, and tracing platforms at large scale (millions of active series, multiple terabytes of logs per day) and making them cheaper and more reliable over time.
Prometheus Ecosystem: Advanced production experience with Prometheus and a long-term storage layer such as Thanos, Cortex, or Mimir, including PromQL, recording rules, remote write, and cardinality management.
Commercial Platforms: Hands-on experience with Datadog or a comparable commercial platform, including how its pricing is built across hosts, custom metrics, log indexing, and APM.
AWS Observability: Production experience with CloudWatch, CloudWatch Logs, Container Insights, X-Ray or Application Signals, Amazon Managed Service for Prometheus, and Amazon Managed Grafana.
OpenTelemetry: Production experience with the OpenTelemetry Collector (or ADOT), including pipeline design, processors, tail and head sampling, and multi-backend export.
Log Pipelines: Experience designing and operating log pipelines with Vector, Fluent Bit, Logstash, or Firehose, including transforms, routing, and tiered storage in S3.
Kubernetes: Strong production experience with EKS or Kubernetes, including DaemonSet agent sizing, kube-state-metrics, and monitoring very large clusters.
Infrastructure as Code: Advanced proficiency with Terraform; experience with Terragrunt or Atmos is a plus.
Alerting & Incident Response: Experience designing SLO-based alerting, reducing alert sprawl, and integrating with incident management tooling such as PagerDuty.
Cost Modeling: Ability to build defensible TCO models from usage data and pricing, and to explain the assumptions behind every number.
Automation: Strong scripting ability using Python, Go, or Bash to pull usage data, analyze telemetry, and automate migration work.
Discovery & Assessment: Demonstrated ability to enter an unfamiliar environment, measure it directly, and produce a defensible current-state picture and recommendation in weeks rather than quarters.
Communication: Ability to explain technical, cost, and compliance tradeoffs to engineers, Security and Legal stakeholders, and executive leadership.
Extra Awesome
Vendor Migration: Experience migrating off Datadog, Splunk, New Relic, or similar platforms onto AWS-native or open-source observability stacks.
Dashboards & Alerts as Code: Experience automating large-scale dashboard and alert migrations using Grafana provisioning, Grafonnet, Terraform providers, or conversion tooling.
Grafana Ecosystem: Experience with Grafana Alloy, Beyla, or other eBPF-based instrumentation.
Continuous Profiling: Experience with Pyroscope, Parca, or comparable profiling tools.
Data Governance: Experience designing telemetry retention and data handling to meet SOC 2, GDPR, privacy, or similar compliance requirements.
Analytics Platforms: Familiarity with Athena, OpenSearch, Databricks, or similar platforms used for log analytics and long-term telemetry queries.
Consumer Scale: Experience with high-traffic B2C platforms where telemetry volume tracks tens of millions of users.
Consulting: Experience leading assessment or discovery engagements that end in a committed, funded program of work.
Certifications: Prometheus Certified Associate, OpenTelemetry Certified Associate, AWS Certified DevOps Engineer – Professional, AWS Certified Solutions Architect – Professional, CKA, or similar.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
Train and Evaluate AI Agents in CAD Environments (Freelance)
Mindrift
United StatesTrain and Evaluate AI Agents in CAD Environments (Freelance)
Mindrift
United StatesSenior Engineer – DoD/U.S. Navy Energetics Facility Design and Construction
Eastern Research Group
United StatesSecurity Operations Specialist
HiddenLayer
United StatesExplore Top Companies in this Space
inventYOU
IT Consulting & Strategic Tech Transformation / Custom Software Engineering & Java/.NET SaaS / European Nearshoring & Technical Agile Teams / Cloud Infrastructure & DevOps Operations
Motia
Fleet Management / Fuel Card Solutions / Telematics / Enterprise Software
Salvo Software
Enterprise Software / ERP Systems / Business Automation / IoT
UpSmith
Artificial Intelligence / Enterprise Software / Home Services / SaaS
EverOps
View Company ProfileEverOps (operating via everops.com) is a premier Embedded Service Provider and strategic technology partner engineered to help global enterprise software companies accelerate software delivery, reduce operating risk, and optimize cloud spend. Founded in 2012 and headquartered in San Francisco, California, the enterprise fundamentally disrupts traditional staff augmentation and consulting models by deploying embedded 'TechPods'—elite engineering teams that integrate directly into client environments with proven playbooks and AI-native practices. Moving far beyond advisory roles, EverOps natively unifies Cloud, CI/CD, Observability, Site Reliability Engineering (SRE), Security, and FinOps to deliver guaranteed outcomes on complex, mission-critical initiatives like re-platforming and global network modernization. The company partners closely with industry-leading platform providers, holding prestigious AWS Select Tier and Datadog Advanced partner certifications, to ensure deep architectural expertise and direct vendor support. Under the hood, their execution methodology leverages forward-deployed engineers who co-own high-stakes infrastructure challenges, utilizing infrastructure-as-code (IaC), GitOps delivery models, and comprehensive observability integrations to eliminate technical debt and significantly improve mean time to resolution (MTTR). Capturing massive market acceleration, having supported multiple IPOs and creating billions in combined client market value, EverOps remains a definitive cornerstone for organizations seeking to turn complex operational bottlenecks into highly scalable, automated, and secure digital foundations.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.