Staff Software Engineer, Observability Platform
CanadaJob Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
About the Observability Platform Team
The Observability Platform team exists to make production legible to every engineer and AI agent at Wealthsimple, on their first day and every day after. We treat the platform as a product, with our internal engineering teams as our customers.
Our mission is to provide the best, most consistent logging, metrics, and tracing experience. We do this by brokering tools and patterns into golden paths that accelerate best practices, and by connecting teams to business impact through SLOs, user flows, and customer experience signals. We instrument on open standards so that switching costs are a configuration change rather than a rewrite, we optimize for the questions engineers did not anticipate rather than for pre-built dashboards, and we design for a world where both humans and AI agents investigate production through fast, high-cardinality data.
As the Staff Software Engineer on this team, you are the deepest technical builder and the person who sets the technical bar. You design the foundations, write the code that matters most, and raise the engineering quality of everyone around you.
In this role you'll have the opportunity to:
- Design and build the events-first foundation. Own the architecture of the wide-event data model and the high-throughput ingest, storage, and query pipelines behind it, including the high-cardinality and columnar or streaming systems that make arbitrary questions answerable in production.
- Build the SDKs and golden paths. Design and ship the instrumentation libraries, shared SDKs, and defaults that make rich, consistent telemetry the path of least resistance for every engineering team, and drive their adoption.
- Set the technical standards. Define the instrumentation conventions, naming, tagging, sampling, and context-propagation practices on open standards such as OpenTelemetry, and codify them so humans and machines share one language.
- Make production legible to AI agents. Build the fast query foundation and access patterns, including protocols such as MCP, that let AI agents investigate incidents, verify their own changes, and operate in tight feedback loops alongside engineers.
- Move fast with AI tooling. Use AI coding tools such as Claude Code and modern LLMs fluently to prototype, build, and navigate large systems, and help the team raise its own bar for building with AI.
- Work confidently in ambiguity. Jump into unfamiliar codebases and make significant, well-reasoned changes with high impact.
- Prove value through experiments. Pilot new approaches with one or two teams, measure the results, and scale what works rather than committing everything up front.
- Raise the bar for others. Mentor senior engineers, review complex designs, and lead cross-team technical initiatives through influence rather than authority.
What you'll bring:
- 8+ years of significant software engineering experience, with a software development background rather than a primarily operations or systems-administration one. You build platforms and tools as software, with the design, testing, and engineering rigor that implies.
- A hands-on, current coder in more than one language such as Kotlin and Ruby. You are comfortable across multiple stacks, and you still spend meaningful time writing and shipping production code.
- Direct experience building an observability or telemetry platform and its SDKs before. You have shipped the instrumentation libraries, pipelines, and shared platforms that other engineers build on, and you carry the deep technical expertise this team is founded on.
- Depth in event-based and high-cardinality systems. You understand wide structured events, columnar or streaming backends, cardinality, sampling, tagging, and context propagation, and the tradeoffs of ingesting and querying telemetry at scale.
- Fluency with open standards and instrumentation. Hands-on experience with OpenTelemetry or an equivalent, and with SLI and SLO design, is core to how you work.
- Comfort making significant changes in ambiguous codebases. You can enter a system you did not build, form an accurate model of it quickly, and change it safely and substantially.
- Proficiency with AI coding tools and LLMs. You use tools such as Claude Code fluently in your daily workflow, you have a clear point of view on where they help and where they do not, and you keep quality high while moving faster.
- Strong communication and influence. You partner well with other engineering teams, explain complex tradeoffs clearly, and improve the systems and the people around you.
Nice to have:
- Experience in a regulated or fintech environment, familiarity with Kubernetes and progressive delivery such as Argo Rollouts, experience with columnar stores such as ClickHouse, and experience building observability for LLM and agent-based workloads.
How would you rate this job post?
See what other professionals think about this role.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.