Senior Software Engineer, Backend
United StatesJob Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
About the role:
If you’re passionate about pushing the limits of reliability, fault-tolerance, and operational excellence in complex distributed systems, Camunda offers the perfect stage.
As a Senior Software Engineer, Backend, you’ll take ownership of automated reliability testing and chaos engineering at the core of Camunda 8. This role thrives on curiosity, experimentation, and relentless improvement—you’ll break things (on purpose!) in safe environments to strengthen our platform before customers ever feel the impact. Here, you’ll make a measurable impact and guide product direction, all while learning new technologies and collaborating with talented, friendly colleagues who live our FAITH values (Focus, Ambition, Integrity, Talent, and Humor).
Every day at Camunda, you’ll help us build robust systems, empower teams, and deliver a great user experience—even in failure scenarios.
Curious about the kind of challenges you'll work on at Camunda? Watch this quick 30-minute talk from our engineers to learn more about the new Camunda Exporter and how we’re solving complex problems at scale
What You’ll Be Doing:
Design, implement, and execute automated chaos experiments and reliability tests, validating Camunda's platform under real-world scenarios
You will investigate, root-cause and debug potential failures or performance regressions of our Java based products
Continuously improve existing load testing, chaos engineering and observability infrastructure to keep our system reliable and performant
Introduce new tooling and approaches for reliability testing, driving measurable improvements to operational experience and user outcomes (examples from the past: introducing zbchaos as fault injection tool)
Collaborate closely with QA and engineering (including cross-functional teams), sharing learnings (e.g. with our team blog) and influencing team roadmaps based on experiment findings
Champion a pragmatic, autonomous approach to software design, solving tough problems and learning new principles and technologies (Kubernetes, Helm, Grafana, Java, Go, and more)
Advocate for user-centric reliability by using and understanding Camunda’s products firsthand
Defining success for this role:
Based on all of the above, you have end to end delivered a concrete task, as for example (the concrete goal will be defined during your onboarding):
After 3 months, you have designed and delivered a way to run quick and reproducible load tests, allowing us:
to reduce the feedback loops for engineers
directly fitting into the development lifecycle
Enable engineers to make use of it and extend it
After 3 months, you have designed and implemented a realistic, automated, holistic load testing framework for Optimize, in close collaboration with Senior Engineers and Field teams:
Conducted a deep-dive analysis of performance limits.
Shared results with the relevant core engineering teams and enabled them to re-run the tests.
Documented findings in public-facing documentation to help users better understand performance characteristics and guide their installation design.
What You Bring:
Ability and/or willingness to use our product
5+ years of experience in backend software engineering (Java)
Proven drive to experiment, learn new tech, and conduct automated reliability and chaos testing in distributed environments
Deep enthusiasm for improving performance and fault-tolerance in production systems
Autonomous, pragmatic problem-solving skills with an ability to guide and influence others
Willingness and ability to use Camunda products and approach reliability from a user’s perspective
Nice to Haves:
Hands-on experience with Go or other languages, in a multi-language context
Working as SRE on distributed systems, with a strong software engineering mindset before
Working knowledge of Kubernetes, Helm Charts, Operators, and related production infrastructure
Experience running apps in production, with solid skills in monitoring, troubleshooting, and performance analysis
Background in chaos engineering, automated reliability, load, or performance testing
What We Have to Offer:
Compensation
We offer competitive, fair, and transparent compensation. Salary ranges are location-based, with Standard and Major markets (global tech hubs) reflecting local competition.
The Annual Total Target Cash (base salary + 100% variable target, where applicable) shown below spans from the minimum in a Standard market to the maximum in a Major market. Final offers depend on skills, experience, and location, and we typically hire in the first half of the range to allow room for growth:
United States: $143,800.00 to $231,900.00
United Kingdom: £90,300.00 to £148,500.00
Singapore: S$178,600.00 to S$267,900.00
Canada: C$152,700.00 to C$251,200.00
If you’re based elsewhere, you’ll be hired via Remote.com (our global employer partner), and your Talent Acquisition Partner will provide a personalized Total Rewards Calculator after your first interview.
Equity: We also offer equity (where applicable) through our Virtual Stock Option Plan (VSOP).
How would you rate this job post?
See what other professionals think about this role.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.