Europe Arbeitnow ·
full-timeSenior Software Engineer, Chaos Engineering
Datadog · Paris
Education
- Level: Bachelor's degree
Key Skills
Notes
0-1 year experience eligible.
Datadog’s Chaos Engineering team builds systems that surface reliability weaknesses before they become outages. As a Senior Software Engineer, you will initially focus on zonal resilience, building automation that helps services safely evacuate and recover from zonal failures, while also contributing to fault injection, incident replay, gameday orchestration, and reliability tooling. You will work across engineering teams to design systems that safely exercise production failure modes and turn findings into verified remediation. You will also help advance the use of AI and automation to identify, test, and close resilience gaps as Datadog’s software and infrastructure evolve. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Build zonal-resilience automation that coordinates safe workload evacuations, switchovers, and recovery in partnership with the teams that own affected services. Design and build fault-injection systems for production environments, including infrastructure- and application-level testing, incident replay, and controlled resilience experiments. Develop safeguards such as blast-radius controls, kill switches, validation mechanisms, and rollback paths that keep production experiments contained and reversible. Build agents and automation that help propose failure scenarios, triage experiment results, and connect reliability findings to tracked remediation and verification. Lead gamedays from hypothesis and scenario design through execution, documented findings, remediation tracking, and validation of completed fixes. Design and implement reliable distributed systems, including gRPC services, Kubernetes controllers, and shared platform components, while contributing to technical design and mentoring other engineers. Who You Are: You have strong distributed systems fundamentals and can reason about consistency, failure modes, backpressure, idempotency, quorum, retries, and failure recovery. You understand Kubernetes workload lifecycles, including how pods, controllers, scheduling, draining, and eviction interact with resilient system design. You have experience designing, building, or operating production systems where safety, availability, and controlled failure handling are important. You communicate complex technical decisions clearly through design documents, runbooks, postmortems, and cross-functional technical discussions. You are comfortable collaborating across engineering teams to understand unfamiliar systems, identify failure modes, and drive resilience improvements. Experience with reliability engineering, chaos engineering, zonal failover, AI-assisted operational workflows, traffic interception, or large-scale observability systems is beneficial but not required. Datadog values people from all walks of life. We know not everyone will meet all the above qualifications on day one. That’s okay. If you’re passionate about technology and want to grow your experience, we encourage you to apply. Benefits and Growth: Develop deep expertise in distributed systems, production resilience, Kubernetes, and large-scale infrastructure. Work on reliability systems that operate across Datadog’s production environment and influence how engineering teams design for failure. Grow your experience designing safe, automated approaches to fault injection, zonal resilience, and incident reproduction. Explore practical applications of AI and automation to reliability engineering and operational workflows. Collaborate with engineers across infrastructure, databases, observability, and service teams on complex systems challenges. Mentor other engineers and contribute to technical designs, engineering practices, and platform strategy. Benefits and Growth listed above may vary based on the country of your employment and the nature of your employment with Datadog. #LI-Hybrid About Datadog: Datadog is the leading observability and security platform for the AI era, providing
Skills for this role
| Role | Senior Software Engineer, Chaos Engineering |
|---|---|
| Company | Datadog |
| Location | Paris |
| Type | full-time |
| Compensation | Not disclosed |
| Deadline | Rolling |
Typical process for this type of role
A general guide — the exact steps for this specific listing may vary; check the original posting for details.
- 1ApplicationSubmit your resume through the apply link.
- 2ScreeningRecruiter reviews your background against the role.
- 3AssessmentA technical test, assignment, or coding round, depending on the role.
- 4Interview(s)One or more rounds with the hiring team.
- 5OfferOffer letter with compensation and start date.
Before you apply
0/4About Datadog
Datadog provides a modern monitoring and security platform designed for developers, IT operations teams, and business users operating in the cloud. The company's platform offers a wide range of capabilities, including infrastructure, network, container, and serverless monitoring, as well as application performance monitoring, log management, and cloud security. These tools are utilized by organizations across various industries, such as financial services, healthcare, retail, and technology, to enable digital transformation and drive collaboration across teams. Datadog employs over 8,100 people globally and continues to expand its presence to support customers across diverse markets and regions.
More at Datadog
Other jobs at Datadog
- Senior Group Manager, Experience Design · New York, New York, USA
- Event Marketing Manager - Global Sponsorships (namer/latam) · New York, New York, USA
- Manager, Revenue Accounting · New York, New York, USA
- Field Enablement Manager (emea) · Amsterdam, The Netherlands; Munich, Germany
- Senior Technology Partner Engineer · New York, New York, USA
Internships at Datadog
- Applied Science Intern · Paris, France
- Software Engineering Intern · Paris, France
- IT Support Technician Intern · Paris, France
- Product Management Intern · New York, New York, USA
- Research Science Intern (phd) · New York, New York, USA; Pittsburgh, Pennsylvania, USA
Explore Related Placements
// similar opportunities
You might also like
Senior Software Engineer - SRE
Mercury
When the Tarr Steps, a footbridge assembled of heavy stones in Exmoor National Park in England, washed away in a flood in 1942, the Royal Engineers rebuilt it. ...
Senior Software Engineer
Chaosindustries
CHAOS Industries is redefining modern defense with a multi-product portfolio that gives the ultimate advantage—domain dominance. The company's products are powe...
Senior Software Engineer
Gymshark
OVERVIEW: We're looking for a Senior Backend Software Engineer to help build the technology that support Gymshark's global growth. Working with industry best to...
Senior Software Engineer
Deliveroo
Our Global Structure Deliveroo is now part of DoorDash, bringing together teams with even greater reach, scale, and ambition. Depending on your role, you may co...
Senior Software Engineer
Kraken
Help us use technology to make a big green dent in the universe! Kraken powers some of the most innovative global developments in energy. We create the technolo...
Senior Software Engineer
Catapult Sports
SENIOR SOFTWARE ENGINEER Catapult is building the future of sports performance technology, with a mission to Unleash the Potential of every athlete and team on ...