Europe Arbeitnow ·
full-timeLead Site Reliability Engineer
Zego · London
About Zego 🚀 At Zego, we're on a mission to do the good thing, not the insurance thing. Insurance hasn't changed much in over a century. The way we live, work and travel has. We're building the real-time, AI-driven infrastructure that powers innovative insurance, so good drivers get cover that works the way they actually drive. We're not just updating insurance; We're leading the AI evolution in insurance 🤖 For us, AI isn't a line on a roadmap or a buzzword on a slide. It's our operating reality, and it's how we build products that back drivers instead of the old insurance playbook. We don't do things slowly, and we don't do bureaucracy. We back high-performance builders who want ownership, early responsibility and the chance to do the most career-defining work of their lives. You'll get the space to try things, the tools to move fast, and the room to see your ideas reach millions of drivers. Do not take our word for it. Read what Zegons say about us on Glassdoor . If you're ready to build the future of insurance, we're hiring. Overview of the Role You will build the Site Reliability Engineering function at Zego, embedding reliability, observability and operational excellence as core engineering concerns - AI is the primary lever for doing that at scale, not a bolt on. You will be our dedicated SRE, working alongside Systems Engineering and embedded with Product and Engineering across roughly 120 engineers. Teams own their services and their own on-call. You own the framework, the instrumentation and the AI tooling that makes that ownership work. You will make teams good at running what they build: alert quality over alert volume, runbooks an agent can execute rather than prose that describes, and diagnostics good enough that an engineer or an agent reaches resolution without an SRE in the room. You will be the primary advocate for reliability with Product and Engineering, making the case with data rather than assertion, so platform health is prioritised alongside delivery commitments. Key Responsibilities Define and land the reliability standard: SLIs, SLOs, error budgets and production readiness criteria that teams apply to their own services, with adoption measured rather than assumed. Encode standards into automation rather than enforcing them by hand. Guardrails in CI, agents that check reliability and observability posture on pull requests, and safe defaults in shared infrastructure so the reliable path is the easy path. Build SRE capability on Zego's AI platform, extending our MCP servers and agents, and making the estate legible to them through machine readable runbooks, structured telemetry and diagnostics an agent can act on. Own observability as a practice, including instrumentation design, signal quality, cardinality and cost. Raise the standard of incident response end-to-end, from detection and triage through to retrospectives that produce change. Put AI to work where it pays off most, stripping toil out of incidents so responders can focus on judgement Measure toil, publish it, and remove it through tooling that others can run without you. What you will need to be successful in the Role We are looking for an engineer who lives SRE and DevOps culture, and treats AI tooling as a default part of how the work gets done, not a side experiment. You will engage and empower teams through decisions grounded in data, and you will be as comfortable building with AI as consuming it: extending MCP servers, writing agents, and automating the operational work that would otherwise fill your week. We treat reliability at Zego as a platform capability, so we care more about what you can make repeatable for others than what you can fix yourself. What you’ll bring to the Team Deep SRE experience: SLI and SLO definition, error budget management, and ownership of incident response through to blameless retrospectives. Fluency with AI as an engineering tool rather than a chat window: building agents and automation, working with MCP servers or equivalent, and judgement about where an agent can act and where a human stays in the loop. Strong software engineering in Python, writing tested, maintainable code that other engineers can run and extend. Observability depth with OpenTelemetry and Datadog. AWS at scale, and Kubernetes based platforms managed through Infrastructure as Code. Desirable Istio, Crossplane and ArgoCD. We run all three and will teach the specifics to the right person. Running ML or AI workloads in production, such as model serving or feature pipelines. We do this at Zego and you would help support it. Having been a sole delivery owner somewhere, and knowing what that does and does not mean. The Zego ways of working 🏡 Teams work better with time to collaborate and space to get things done. We call it Zego Hybrid: some of us are in our office (central London or Halifax) weekly, others monthly or quarterly. It's about finding the balance between face time and focus that produces great work and a healthy
| Role | Lead Site Reliability Engineer |
|---|---|
| Company | Zego |
| Location | London |
| Compensation | Not disclosed |
| Deadline | Rolling |
Typical process for this type of role
A general guide — the exact steps for this specific listing may vary; check the original posting for details.
- 1ApplicationSubmit your resume through the apply link.
- 2ScreeningRecruiter reviews your background against the role.
- 3AssessmentA technical test, assignment, or coding round, depending on the role.
- 4Interview(s)One or more rounds with the hiring team.
- 5OfferOffer letter with compensation and start date.
Before you apply
0/4Explore Related Placements
// similar opportunities
You might also like
Site Reliability Engineer
Deepjudge
About Us: DeepJudge is Switzerland’s leading AI and ICT scale-up, transforming how law firms and legal departments access and leverage their knowledge. Founded ...
Technical Site Reliability Engineer
Andurilindustries
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the e...
Engineering Site Lead
Perplexity
Perplexity is revolutionizing how people discover and interact with information through AI-powered search and knowledge tools. As we expand our global footprint...
Senior Site Reliability Engineer
Camunda
Camunda is the enterprise platform for agentic orchestration , enabling organizations to coordinate AI agents, people, and systems across complex, end-to-end bu...
Site Reliability Engineers (SRE)
Xm
The Role: You will join a team working with Observability, Escalations, Post-mortems, Correction of Errors, and other practices that will contribute to the comp...
Site Reliability Engineer
Tokyodev
Site Reliability Engineer...
0-1 year experience eligible.