Director of Platform Operations
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
At Arbor, we’re on a mission to transform the way schools work for the better. We believe in a future of work in schools where being challenged doesn’t mean being burnt out and overworked. Where data guides progress without overwhelming staff. And where everyone working in a school is reminded why they got into education every day.
Our MIS and school management tools are already making a difference in over 12,000 schools and trusts. Giving time and power back to staff, turning data into clear, actionable insights, and supporting happier working days.
At the heart of our brand is a recognition that the challenges schools face today aren’t just about efficiency, outputs, and productivity—but about creating happier working lives for the people who drive education every day: the staff. We want to make schools more joyful places to work, as well as learn.
About the Role
We are looking for an experienced and highly knowledgeable Platform Operations Director to join our Engineering team and to own the operational backbone of Arbor’s suite of applications. The remit and focus of the role is to bring together three distinct disciplines — Site Reliability Engineering, Security Engineering, and Developer Experience — under single accountable leadership, and be responsible for their outcomes across every product line and market. It’s a broad and exciting role, so we’re looking for someone up for a challenge—if you’re an effective leader and are highly collaborative, this is the role for you.
Core Responsibilities
Site Reliability Engineering and Availability
Own the availability commitment. Be accountable for meeting the 99.9% availability SLA across the application suite. Define, publish, and govern the SLO and error-budget framework that makes availability measurable, forecastable, and actionable rather than retrospective.
Build the reliability discipline. Lead the SRE pillar to embed observability, capacity planning, performance engineering, resilience testing, and toil reduction as standing practices with clear owners and cadence.
Make reliability visible. Ensure availability, latency, and error-budget consumption are reported transparently to product teams, R&D leadership, and the executive, with credible attribution of loss to cause.
Engineer out recurrence. Drive systemic reliability improvement through problem management, tracking repeat causes, single points of failure, and architectural weak points to closure with named owners and dates.
Security Engineering and Vulnerability Management
Own the vulnerability management framework. Define and operate the end-to-end framework for identifying, triaging, prioritizing, and remediating security vulnerabilities across application code, dependencies, containers, and infrastructure, including agreed severity definitions and remediation SLAs.
Put the reporting in place. Build the reporting and dashboards that allow the organization — R&D leadership, the executive, and where relevant the board and customers — to monitor vulnerability burn-down and resolution against SLA, by severity, age, and owning team.
Hold the line on remediation. Ensure vulnerabilities are resolved rather than merely recorded: drive burn-down of the existing backlog, prevent aging, and escalate credibly where remediation is not being prioritized.
Shift security left. Embed security into the engineering workflow through automated scanning, secure-by-default platform patterns, dependency hygiene, and practical enablement for product teams. Partner closely with security leadership and the CISO on posture, roadmap, compliance obligations, and audit evidence.
Developer Experience and the Build-It, Run-It Transformation
Lead the transformation. Own the program of work that moves product engineering teams to “build it, run it” — taking genuine production ownership of their services, including on-call, alerting, and operational health — via a staged, evidenced adoption path rather than a mandate.
Make the right thing the easy thing. Lead the DevX pillar to deliver the internal developer platform, golden paths, self-service tooling, CI/CD, and environment provisioning that make ownership viable for teams and reduce cognitive load.
Treat platform as a product. Run the platform with product discipline: known internal customers, articulated service levels, adoption metrics, feedback loops, and a roadmap prioritized on developer impact.
Measure and improve engineering throughput. Own the engineering productivity metrics — deployment frequency, lead time for change, change failure rate, and time to restore — and use them to target investment where it demonstrably lifts delivery.
Requirements
About You
Proven senior leadership of SRE, platform, or infrastructure functions at scale, with accountability for availability against a defined SLA in a customer-facing SaaS environment.
Deep, practical command of reliability engineering: SLIs, SLOs, error budgets, observability, capacity planning, and problem management.
Demonstrated ownership of a major incident management process, including incident command models, on-call design, and blameless post-incident review, with evidence of materially improving MTTR.
Track record of owning or closely partnering on security posture, including vulnerability management at scale, remediation SLAs, and reporting to executive or board level.
Experience leading a shift to distributed production ownership (build it, run it) in an organization that previously centralized operations.
Experience of internal developer platform or platform-as-a-product models, and of using engineering productivity metrics to drive investment.
Credible judgment on where and how to apply AI to operational workflows, with a clear-eyed view of what to automate, what to augment, and where human accountability must remain.
Strong track record of building and developing engineering leaders, including handling performance robustly.
Ability to deliver outcomes through influence across teams outside direct reporting lines, and to hold peers to account constructively.
Strong technical credibility across modern cloud platforms (AWS) and infrastructure as code (Terraform), with the judgment to hold teams to a high engineering standard.
Excellent communication and influencing skills, able to operate confidently at executive level and convey technical and risk concepts to non-technical audiences, including during live incidents.
Desirable
Experience in enterprise SaaS at scale, ideally in EdTech or another data-sensitive or regulated domain.
Familiarity with security and compliance frameworks relevant to education data, for example ISO 27001, SOC 2, Cyber Essentials Plus, and UK GDPR obligations.
Hands-on experience of AIOps or AI-assisted incident tooling, and of evaluating such tooling responsibly.
Experience running a multi-product, multi-market estate on shared platform foundations.
Familiarity with PHP-based estates, Docker and containerization, and Kanban and agile delivery.
FinOps experience and accountability for cloud cost efficiency.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Arbor Education
Explore Top Companies in this Space
Nova Pioneer
Primary & Secondary Education (K-12) / Independent School Networks / Pan-African EdTech & Curricular Development
vironix.ai
AI / Manufacturing / Logistics / Enterprise Software
Pinewood
Technology / Artificial Intelligence / Media & Entertainment / Enterprise Software
OxDynamics
Chemical Engineering / Industrial Automation / Environmental Technology / Process Optimization
Arbor Education
View Company ProfileArbor Education (operating under arbor-education.com) is the premier, enterprise-grade cloud Management Information System (MIS) pioneer, school operations innovator, and educational data powerhouse engineered to operate as the definitive, high-velocity administration, reporting, and institutional governance layer for schools and Multi-Academy Trusts (MATs) across the United Kingdom. The company completely eliminates the severe systemic friction of traditional educational administration—where headteachers, data managers, and teachers face fragmented legacy databases, high on-premise server maintenance overhead, manual statutory census formatting, and disconnected parent communication loops—by deploying a unified, cloud-native operational ecosystem. Moving far beyond traditional, passive spreadsheet logging or rigid legacy school software, Arbor natively unifies full-lifecycle statutory record keeping, formative and summative progress tracking, real-time behavior management pipelines, integrated communication modules (SMS, email, and app notifications), and cashless payment gateways into a single high-availability system of record. Trusted by over 12,500 schools and 930 Multi-Academy Trusts, the platform provides schools with the largest live national benchmarking dataset, ensuring data accuracy within 0.1% of official Department for Education (DfE) drops. Empowered by its integrated, context-aware artificial intelligence core, Arbor AI, the system introduces automated workflow builders, custom data quality dashboards, and its signature Ofsted Inspection Companion to significantly reduce non-teaching workloads and reclaim valuable instruction time with production-hardened precision. Under the hood, its sophisticated technical architecture utilizes single-tenanted private firewalled databases, high-memory Amazon EC2 worker instances, and high-concurrency GraphQL and REST APIs to provide seamless third-party app integrations and real-time parallel data analytics processing. What sets Arbor Education apart is its uncompromising dedication to replacing administrative educational logjams with absolute operational momentum and cloud-first collaboration; by bridging the gap between complex multi-tier institutional telemetry and intuitive, mobile-first parent and teacher portals, the corporation remains the absolute, undisputed category cornerstone of modern school management and scalable algorithmic education tech.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.

