Greenhouse ·
full-timeSenior Software Engineer - Incident Insights & Readiness
Datadog · Boston, Massachusetts, USA; New York, New York, USA
Datadog provides a monitoring and observability platform that processes trillions of data points daily, offering alerting, metrics, logs and tracing to thousands of companies. The Incident Insights & Readiness SRE team works inside Datadog to turn incidents into learning opportunities, building tooling and operational frameworks that help engineers prepare for, respond to, and learn from failures. In this role you would own the company‑wide on‑call experience, defining best practices and building platforms that support rotations and compensation. You would design software to streamline incident response, collaborate with product teams to improve those tools, and contribute to the post‑mortem process by guiding teams on writing effective reviews and identifying friction points. The position also involves technical leadership: coaching teammates through design reviews, training on‑callers in incident management, and leading cross‑functional initiatives that embed reliability practices across engineering groups. Ideal candidates have at least five years of software engineering experience, strong skills in Go, Python or TypeScript, familiarity with distributed systems and Kubernetes, a track record of analyzing incidents and driving improvements, and experience mentoring engineers and influencing technical direction without formal authority.
We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way The Incident Insights & Readiness SRE team at Datadog fosters a resilient culture by using incidents as learning opportunities and catalysts for growth. Our users are Datadog engineers, and we build the software, tooling, and operational frameworks that help them prepare for, respond to, and learn from incidents. We work closely with engineering teams across Datadog to analyze incidents and turn those insights into better tools, stronger incident response, and organizational learning. Our efforts empower Datadog to navigate unexpected failures confidently, efficiently, and with a commitment to continuous learning and systems improvement. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own and improve the on-call experience for the company by establishing best practices and building platforms to support on-call rotations and compensation. Define how we respond to incidents, lead the design and implementation of software to streamline the process, and collaborate with product teams to improve incident response across Datadog. Our aim is to fully support our incident responders in dealing with complexity. Contribute to the post-mortem process for the company, collaborating with teams on writing them, and identifying opportunities to reduce friction and enhance learning value for the organization. Our team also runs a weekly postmortem reading group. Support various teams in facilitating incident reviews that emphasize learning and blamelessness. Help them share their learnings across the organization to improve the resilience of our people. Provide technical leadership and day-to-day coaching to team members, accelerating their growth through design reviews, collaborative problem-solving and operational excellence best practices. Train our on-callers in incident and post-mortem processes, sharing expertise in incident management best practices. This involves both introducing newcomers to on-call responsibilities and refreshing the knowledge of existing engineers. Lead cross-functional initiatives in engineering organizations across Datadog, embedding with teams to understand their challenges and drive lasting improvements to reliability and operational excellence. Who You Are: At least 5 years of experience building software that solves real user problems. Experience designing new features and collaborating on code and technical design reviews. We primarily develop in Go and Python, with a bit of TypeScript. Experience building or operating distributed systems, with familiarity with Kubernetes and an understanding of complex failure modes. Demonstrated ability to independently own ambiguous technical problems from design through delivery while balancing long-term engineering quality with pragmatic execution. Experience analyzing incidents, identifying systemic risks, and driving engineering improvements informed by operational learnings. Experience participating in on-call rotations and improving incident response processes. Experience serving as an incident commander or incident coordinator is a plus. Empathy, collaboration, and communication skills in English to cultivate strong relationships across various teams in the organization Experience mentoring engineers, driving cross-functional initiatives, and influencing technical direction without relying on organizational authority. We welcome candidates from a variety of backgrounds, including software engineering, site reliability engineering, production engineering, infrastructure, and other roles focused on building reliable systems or improving incident response. Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications on day one. That's okay. If you’re passionate about technology and want to grow your skills, we encourage you to apply. Benefits and Growth: New hire stock equity (RSUs) and employee stock purchase plan (ESPP) Continuous professional development, product training, and career pathing Intradepartmental mentor and buddy program for in-house networking An inclusive company culture, ability to join our Community Guilds (Datadog employee resource groups) Access to Inclusion Talks, our internal panel discussions Free, global mental health benefits for employees and dependents age 6+ Competitive global benefits Benefits and Growth listed above may vary based on the country
| Role | Senior Software Engineer - Incident Insights & Readiness |
|---|---|
| Company | Datadog |
| Location | Boston, Massachusetts, USA; New York, New York, USA |
| Type | internship |
| Compensation | Not disclosed |
| Posted | 2026-10-01 |
| Deadline | Rolling |
Typical process for this type of role
A general guide — the exact steps for this specific listing may vary; check the original posting for details.
- 1ApplicationSubmit your resume through the apply link.
- 2ScreeningRecruiter reviews your background against the role.
- 3AssessmentA technical test, assignment, or coding round, depending on the role.
- 4Interview(s)One or more rounds with the hiring team.
- 5OfferOffer letter with compensation and start date.
Before you apply
0/4About Datadog
Datadog provides a modern monitoring and security platform designed for developers, IT operations teams, and business users operating in the cloud. The company's platform offers a wide range of capabilities, including infrastructure, network, container, and serverless monitoring, as well as application performance monitoring, log management, and cloud security. These tools are utilized by organizations across various industries, such as financial services, healthcare, retail, and technology, to enable digital transformation and drive collaboration across teams. Datadog employs over 8,100 people globally and continues to expand its presence to support customers across diverse markets and regions.
More at Datadog
Other jobs at Datadog
- Senior Group Manager, Experience Design · New York, New York, USA
- Event Marketing Manager - Global Sponsorships (namer/latam) · New York, New York, USA
- Senior Software Engineer, Chaos Engineering · Paris
- Manager, Revenue Accounting · New York, New York, USA
- Field Enablement Manager (emea) · Amsterdam, The Netherlands; Munich, Germany
Internships at Datadog
- Applied Science Intern · Paris, France
- Software Engineering Intern · Paris, France
- IT Support Technician Intern · Paris, France
- Product Management Intern · New York, New York, USA
- Research Science Intern (phd) · New York, New York, USA; Pittsburgh, Pennsylvania, USA
Explore Related Placements
// similar opportunities
You might also like
Senior Software Engineer
Datadog
About Datadog: We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We opera...
Senior Software Engineer
Spotify
The Platform team creates the technology that enables Spotify to learn quickly and scale easily, enabling rapid growth in our users and our business around the ...
Senior Software Engineer
Servicenow
No description available.
Senior Software Engineer
Okta
Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure th...
Senior Software Engineer, Engagement
Discord
Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly ever...
Senior Software Engineer - Identity & Access Management
Brex
Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corpo...