Jobbie
← Discover jobs
Arbor Education

Director of Platform Operations

Industry Other industries

Remote, United KingdomPosted 8h ago

Job description

Location: Remote

Salary: £150,000 - £160,000

About us

At Arbor, we’re on a mission to transform the way schools work for the better.

We believe in a future of work in schools where being challenged doesn’t mean being burnt out and overworked. Where data guides progress without overwhelming staff. And where everyone working in a school is reminded why they got into education every day.

Our MIS and school management tools are already making a difference in over 12,000 schools and trusts. Giving time and power back to staff, turning data into clear, actionable insights, and supporting happier working days.

At the heart of our brand is a recognition that the challenges schools face today aren’t just about efficiency, outputs and productivity - but about creating happier working lives for the people who drive education everyday: the staff. We want to make schools more joyful places to work, as well as learn.

About the role

We are looking for an experienced and highly knowledgeable Platform Operations Director to join our Engineering team and to own the operational backbone of Arbor’s suite of applications. The remit and focus of the role is to bring together three distinct disciplines — Site Reliability Engineering, Security Engineering, and Developer Experience — under single accountable leadership, and be responsible for their outcomes across every product line and market. It’s a broad and exciting role, so we’re looking for someone up for a challenge - if you’re an effective leader and are highly collaborative, this is the role for you.

Core responsibilities

Site Reliability Engineering and availability - Own the availability commitment. Be accountable for meeting the 99.9% availability SLA across the application suite. Define, publish, and govern the SLO and error-budget framework that makes availability measurable, forecastable, and actionable rather than retrospective. - Build the reliability discipline. Lead the SRE pillar to embed observability, capacity planning, performance engineering, resilience testing, and toil reduction as standing practices with clear owners and cadence. - Make reliability visible. Ensure availability, latency, and error-budget consumption are reported transparently to product teams, R&D leadership, and the executive, with credible attribution of loss to cause. - Engineer out recurrence. Drive systemic reliability improvement through problem management, tracking repeat causes, single points of failure, and architectural weak points to closure with named owners and dates.

Security engineering and vulnerability management - Own the vulnerability management framework. Define and operate the end-to-end framework for identifying, triaging, prioritising, and remediating security vulnerabilities across application code, dependencies, containers, and infrastructure, including agreed severity definitions and remediation SLAs. - Put the reporting in place. Build the reporting and dashboards that allow the organisation — R&D leadership, the executive, and where relevant the board and customers — to monitor vulnerability burn-down and resolution against SLA, by severity, age, and owning team. - Hold the line on remediation. Ensure vulnerabilities are resolved rather than merely recorded: drive burn-down of the existing backlog, prevent ageing, and escalate credibly where remediation is not being prioritised. - Shift security left. Embed security into the engineering workflow through automated scanning, secure-by-default platform patterns, dependency hygiene, and practical enablement for product teams. Partner closely with security leadership and the CISO on posture, roadmap, compliance obligations, and audit evidence.

Developer Experience and the build-it, run-it transformation - Lead the transformation. Own the programme of work that moves product engineering teams to “build it, run it” — taking genuine production ownership of their services, including on-call, alerting, and operational health — via a staged, evidenced adoption path rather than a mandate. - Make the right thing the easy thing. Lead the DevX pillar to deliver the internal developer platform, golden paths, self-service tooling, CI/CD, and environment provisioning that make ownership viable for teams and reduce cognitive load. - Treat platform as a product. Run the platform with product discipline: known internal customers, articulated service levels, adoption metrics, feedback loops, and a roadmap prioritised on developer impact. - Measure and improve engineering throughput. Own the engineering productivity metrics — deployment frequency, lead time for change, change failure rate, and time to restore — and use them to target investment where it demonstrably lifts delivery.

Requirements

About you - Proven senior leadership of SRE, platform, or infrastructure functions at scale, with accountability for availability against a defined SLA in a customer-facing SaaS environment. - Deep, practical command of reliability engineering: SLIs, SLOs, error budgets, observability, capacity planning, and problem management. - Demonstrated ownership of a major incident management process, including incident command models, on-call design, and blameless post-incident review, with evidence of materially improving MTTR. - Track record of owning or closely partnering on security posture, including vulnerability management at scale, remediation SLAs, and reporting to executive or board level. - Experience leading a shift to distributed production ownership (build it, run it) in an organisation that previously centralised operations. - Experience of internal developer platform or platform-as-a-product models, and of using engineering productivity metrics to drive investment. - Credible judgement on where and how to apply AI to operational workflows, with a clear-eyed view of what to automate, what to augment, and where human accountability must remain. - Strong track record of building and developing engineering leaders, including handling performance robustly. - Ability to deliver outcomes through influence across teams outside direct reporting lines, and to hold peers to account constructively. - Strong technical credibility across modern cloud platforms (AWS) and infrastructure as code (Terraform), with the judgement to hold teams to a high engineering standard. - Excellent communication and influencing skills, able to operate confidently at executive level and convey technical and risk concepts to non-technical audiences, including during live incidents.

Desirable - Experience in enterprise SaaS at scale, ideally in EdTech or another data-sensitive or regulated domain. - Familiarity with security and compliance frameworks relevant to education data, for example ISO 27001, SOC 2, Cyber Essentials Plus, and UK GDPR obligations. - Hands-on experience of AIOps or AI-assisted incident tooling, and of evaluating such tooling responsibly. - Experience running a multi-product, multi-market estate on shared platform foundations. - Familiarity with PHP-based estates, Docker and containerisation, and Kanban and agile delivery. - FinOps experience and accountability for cloud cost efficiency.

Benefits

What we offer

The chance to work alongside a team of hard-working, passionate people in a role where you’ll see the impact of your work everyday. We also offer: - A dedicated wellbeing team who champion initiatives such as mindfulness, lunch n learns, manager training, mental health first aid training and much more! - 32 days holiday (plus Bank Holidays). This is made up of 25 days annual leave plus 7 extra company wide days given over Easter, Summer & Christmas - Life Assurance paid out at 3x annual salary - Comprehensive wellness benefit provided by AIG Smart Health, which provides a 24/7 virtual GP service, Mental health support, Counselling, and personalised Health Checks  - Private Dental Insurance with Bupa  - Salary sacrifice Pension provided by Scottish Widows - Enhanced maternity and adoption leave (20 weeks full pay) and paternity (6 weeks full pay) pay - 5 free return to work maternity coaching sessions, helping you adapt to this new exciting time of life! - Access to services such as Calm and Bippit (financial wellbeing coaching)  - All of our roles champion flexible working and we are happy to discuss what this means to you - Social committees that plan team, office and company wide events to bring people together and celebrate success - Dedicated professional development training budget (CPD courses, upskilling resources, professional memberships etc) - Volunteer with a charity of your choice for a day each year - Dog friendly offices!

Interview process - Phone screen - 1st stage - 2nd stage

We are committed to a fair and comfortable recruitment process, so if you require any reasonable adjustments during your application or interview process, please reach out to a member of the team at careers@arbor-education.com .

Our commitment is also backed by our partnership with Neurodiversity Consultancy, Lexxic who provide us with training, support and advice.

Arbor Education is an equal opportunities organisation

Our goal is for Arbor to be a workplace which represents, celebrates and supports people from all backgrounds, and which gives them the tools they need to thrive - whatever their ambitions may be so we support and promote diversity and equality, and actively  encourage applications from people of all backgrounds.

Refer a friend

Know someone else who would be good for this role? You can refer a friend, family member or colleague, if they are offered a role with Arbor, we will say thank you with a voucher valued up to £200! Simply email: careers@arbor-education.com

Please note: We are unable to provide visa sponsorship at this time.