Greenhouse ·
full-timePrincipal Site Reliability Engineer, Platform Engineering
Gitlab · Remote, Canada; Remote, United Kingdom; Remote, United States
GitLab provides an intelligent orchestration platform for DevSecOps, helping organizations boost developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. The Principal Site Reliability Engineer, Platform Engineering role focuses on shaping the architecture and strategy for GitLab Dedicated, the fully managed single‑tenant SaaS offering. Day‑to‑day responsibilities include setting technical direction for scaling isolated customer environments, leading platform transformations in resilience, failover, tenant orchestration, change management, self‑service tooling, and integrations, and driving a modular, cell‑based architecture that aligns with GitLab’s broader Cells strategy. The role also involves strengthening service ownership, establishing reusable patterns and automation to reduce toil, identifying systemic reliability and scalability risks using production signals, and influencing complex technical decisions across teams while mentoring senior engineers and advancing engineering excellence. Ideal candidates have deep expertise in site reliability, platform, infrastructure, or backend engineering, with hands‑on experience designing and operating large‑scale production systems, working with cloud infrastructure, automation, observability, infrastructure as code, and modern production engineering practices.
Education
- Level: Bachelor's degree
Key Skills
About the Role GitLab is the intelligent orchestration platform for DevSecOps. It enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. The company embraces AI as a core productivity multiplier and values high‑performance culture, continuous knowledge exchange, and inclusive collaboration. Responsibilities Set technical direction for GitLab Dedicated, shaping architecture and platform strategy as we scale a growing fleet of isolated, single‑tenant environments. Lead platform transformations across resilience, failover, tenant orchestration, change management, self‑service tooling, and platform integrations. Drive scalable, modular architecture aligned with GitLab’s Cells strategy while preserving security, isolation, and compliance. Strengthen service ownership and operational maturity, helping engineering teams build, operate, and improve production systems. Identify and address systemic reliability and scalability risks using production signals, incident patterns, and architectural insight. Establish reusable platform patterns and automation to reduce operational toil and enable efficient scaling. Lead complex technical decisions balancing reliability, security, cost, maintainability, and customer needs. Advance engineering excellence through architectural leadership, mentorship, and influence. Requirements Deep expertise in Site Reliability, Platform, Infrastructure, or Backend Engineering with experience designing and operating large‑scale production systems. Hands‑on experience with cloud infrastructure, automation, observability, infrastructure as code, and modern production engineering practices. Strong software engineering fundamentals and experience building production systems or infrastructure tooling.
Skills for this role
Required skills
Also mentioned in the listing
| Role | Principal Site Reliability Engineer, Platform Engineering: Dedicated |
|---|---|
| Company | Gitlab |
| Location | Remote, Canada; Remote, United Kingdom; Remote, United States |
| Compensation | Not disclosed |
| Posted | 2026-09-29 |
| Deadline | Rolling |
Typical process for this type of role
A general guide — the exact steps for this specific listing may vary; check the original posting for details.
- 1ApplicationSubmit your resume through the apply link.
- 2ScreeningRecruiter reviews your background against the role.
- 3AssessmentA technical test, assignment, or coding round, depending on the role.
- 4Interview(s)One or more rounds with the hiring team.
- 5OfferOffer letter with compensation and start date.
Before you apply
0/4Principal Site Reliability Engineer, Platform Engineering at Gitlab: frequently asked questions
- Who can apply for the Principal Site Reliability Engineer, Platform Engineering role at Gitlab?
- The listing asks for Bachelor's degree; freshers are welcome.
- What skills does the Principal Site Reliability Engineer, Platform Engineering role require?
- The listing highlights Site Reliability Engineering, Platform Engineering, Backend Engineering, Cloud Infrastructure, Automation, Observability, Infrastructure as Code, Production Engineering Practices, Software Engineering Fundamentals, Production Systems. Show each of these in a project or past role on your resume.
- What is the salary for this role?
- Gitlab has not stated compensation in the listing. Check the original posting or ask during the application process.
- Is the Principal Site Reliability Engineer, Platform Engineering position remote, hybrid or onsite?
- The listing marks this role as remote, with Remote, Canada; Remote, United Kingdom; Remote, United States as the location.
- What is the application deadline?
- Gitlab has not listed a fixed deadline, so apply early in case the opening is filled.
- How do I apply for the Principal Site Reliability Engineer, Platform Engineering role?
- Use the Apply button on this page. It opens the original listing on job-boards.greenhouse.io, where you submit your application with the company.
More at Gitlab
Other jobs at Gitlab
- Engineering Manager, Dedicated Infrastructure · Remote, Canada; Remote, United States
- Manager, Assigned Support Engineering (emea) · Remote, United Kingdom
- Support Engineer, U.S. Government Support · Remote, United States
- Senior Product Manager, CRM & GTM Systems · Remote, United States
- Major Account Executive · Remote, Germany
Explore Related Placements
// similar opportunities
You might also like
Site Reliability Engineer
Deepjudge
About Us: DeepJudge is Switzerland’s leading AI and ICT scale-up, transforming how law firms and legal departments access and leverage their knowledge. Founded ...
Lead Site Reliability Engineer
Zego
About Zego 🚀 At Zego, we're on a mission to do the good thing, not the insurance thing. Insurance hasn't changed much in over a century. The way we live, work ...
Senior Site Reliability Engineer
Camunda
Camunda is the enterprise platform for agentic orchestration , enabling organizations to coordinate AI agents, people, and systems across complex, end-to-end bu...
Technical Site Reliability Engineer
Andurilindustries
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the e...
Site Reliability Engineers (SRE)
Xm
The Role: You will join a team working with Observability, Escalations, Post-mortems, Correction of Errors, and other practices that will contribute to the comp...
Principal Presales Engineer
Twilio
Who we are At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of...