Principal Engineer, Commerce Platform
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
About the Role:
We’re building a global commerce platform to power $1T+ in annual transactions for millions of SMEs. We want our Principal Engineers to be the technical force-multiplier who sets the domain modeling, architecture, raises the reliability bar, and multiplies team effectiveness. You’ll steward our core commerce models—subscriptions, payments, catalog, pricing, inventory, fulfillment, reconciliation, and tax—and the services around them, ensuring correctness by design, auditability, and delightful performance at scale. This is an individual contributor (IC) role with organizational-level influence (no direct reports), focused on designing systems, shaping standards, and growing engineers.
Tech Stack:
Backend: Go, ConnectRPC
Databases: MongoDB, Firestore, ClickHouse
Cloud: GCP (GKE), Pub/Sub, Redis, OpenTelemetry
What You’ll Do:
- Architect and ship multi-tenant, planet-scale services (checkout, subscriptions, payments orchestration, invoicing, tax hooks) with clear domain boundaries (Domain-Driven Design) and hard Service Level Objectives (SLOs)
- Be the custodian of API and schema design: own protobuf/ConnectRPC conventions, versioning policy, deprecation playbooks, and Buf breaking-change checks—so our contracts stand the test of time
- Guarantee resilience and availability of core payment paths: timeouts, retries with jitter, circuit breakers, idempotency keys, outbox/Saga patterns, hedged requests, and graceful degradation
- Ensure complete auditability: append-only double-entry ledger, immutable event streams, trace-linked entities (OpenTelemetry trace/span IDs), tamper-evident trails, and reconciliations that tie out to the cent
- Own error boundaries end-to-end: enumerate failure domains (Payment Service Provider, network, data, concurrency, quota, browser, device); design uniform error contracts; implement compensations/backfills and automated replay
- Keep track of every deployed thing: services, workers, triggers, cron, subscriptions—own the service catalog and scorecards (owners, SLOs, runbooks, Pod Disruption Budgets, Horizontal/Vertical Pod Autoscalers, budgets, quotas, timeouts)
- Configuration and limits stewardship: enforce sane defaults across GKE, Pub/Sub, Redis, Firestore/Mongo, ClickHouse—connection pools, acknowledgment deadlines, batch sizes, Time-To-Live, memory/file descriptor limits, and GCP quotas
- Observability as a product: pervasive OpenTelemetry, RED/USE metrics, exemplars, trace sampling, SLO dashboards, and alerting that wakes humans only for user-impacting issues
- Production excellence: canary/blue-green rollouts, automated rollbacks, chaos drills, Disaster Recovery (DR) playbooks (Recovery Point Objective/Recovery Time Objective), multi-region failover strategies, and incident command on rotation
- Security and compliance by design: PCI scope minimization, tokenization/vaulting, secrets/Key Management System (KMS) hygiene, data retention/archival, and privacy controls—embed checks in Continuous Integration/Continuous Deployment (CI/CD)
- Developer acceleration: pave golden paths (service templates, Architecture Decision Records/RFC process, linting/formatting, contract tests, ephemeral environments, load/performance harnesses) to make the right thing the easy thing
What You’ll Lead:
- Core domain evolution: orchestration → ledger → reconciliation flows with crisp invariants and consistency guarantees (read-your-writes where needed, eventual consistency where appropriate)
- Reliability strategy: Service Level Indicators (SLIs)/SLOs, error budgets, capacity planning, cost/FinOps guardrails, multi-region posture, and Disaster Recovery (DR) exercises
- API and data governance: canonical models, schema lifecycle (compatibility matrix, migrations), data lifecycle (retention, archival, compliance)
- Practice leadership for HighLevel: design reviews, postmortems, technical strategy, coding standards, and mentorship across teams—raise the bar for the organization
- Hiring and team growth: help us hire, scale, and train the right team; shape interview loops, rubrics, onboarding, and ongoing learning (brown bags, reviews, pair design)
- Cross-functional partnership: collaborate with Product/Marketing/Support to translate platform capabilities and constraints into roadmaps, Go-To-Market (GTM) narratives, and reliable customer outcomes
- Risk and roadmap: maintain a technical risk register, make build-vs-buy calls, and propose simplifications or deprecations that meaningfully reduce complexity and Mean Time To Recovery (MTTR)
Minimum Qualifications:
- 10+ years building and operating backend systems (at least 5+ years in Go), with 2–3+ years acting as a Staff/Principal-level individual contributor (IC) or Tech Lead for critical paths
- Deep proficiency with protobuf + ConnectRPC/gRPC and API lifecycle management (versioning, compatibility, contract testing, Buf)
- Distributed systems fundamentals: idempotency, exactly-once-ish via deduplication/outbox, ordering, consensus basics, backpressure, concurrency control
- Event-driven architectures on GCP (Pub/Sub), plus Redis for fast paths; strong schema design in MongoDB/Firestore and analytics/reporting patterns on ClickHouse
- Kubernetes/GKE operations at scale: autoscaling (Horizontal/Vertical Pod Autoscalers), Pod Disruption Budgets, resource limits/requests, multi-region topologies, CI/CD, canary/blue-green deployments
- Reliability engineering: SLIs/SLOs, error budgets, capacity and load testing, incident management, Disaster Recovery/Business Continuity Planning (DR/BCP)
- Security and compliance: secrets/KMS best practices, PCI basics (scope reduction, key rotation), and data governance (retention/archival)
- Testing discipline: unit, integration, contract, property-based, performance; test data management and deterministic environments
- Frontend collaboration: solid understanding of Vue.js + TanStack Query to shape clean API surfaces and performance budgets across the boundary
- Exceptional technical writing and communication: design docs, Architecture Decision Records/RFCs, postmortems, and stakeholder updates
Nice to Have:
- Hands-on integrations with major Payment Service Providers (PSPs)/local rails (e.g., UPI, wallets, Buy Now Pay Later (BNPL), cards/3-D Secure 2) and reconciliation at scale
- Experience with active-active or multi-region designs; chaos engineering; traffic management
- Observability leadership with OpenTelemetry at organizational scale (tail-based sampling, exemplars)
- FinOps experience: cost baselining, quotas, budget alarms, and workload right-sizing
- Familiarity with regulatory frameworks (PCI DSS, SOC 2/ISO 27001) and privacy laws relevant to our markets
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at HighLevel
Explore Top Companies in this Space
Promethean
Education Technology / EdTech / Software as a Service (SaaS)
Clipboard
Software as a Service (SaaS), Technology
Procurify
Software as a Service (SaaS), Procurement Technology
Remote
Software as a Service (SaaS), Remote Work Technology
HighLevel
View Company ProfileHighLevel is a comprehensive, all-in-one, AI-powered business operating system and marketing platform designed primarily for digital marketing agencies, freelancers, and small businesses. It provides a centralized suite of tools that replace multiple standalone software subscriptions. Core features include a CRM, sales pipelines, website and funnel builders, email/SMS marketing automation, appointment scheduling, reputation management, and white-labeling capabilities. By consolidating these functions, HighLevel helps businesses capture, nurture, and close leads while streamlining their operations and scaling their growth.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.


