InternFlow

Greenhouse ·

full-time

Senior Software Engineer - Incident Insights & Readiness

Datadog · Boston, Massachusetts, USA; New York, New York, USA

Datadog provides a monitoring and observability platform that processes trillions of data points daily, offering alerting, metrics, logs and tracing to thousands of companies. The Incident Insights & Readiness SRE team works inside Datadog to turn incidents into learning opportunities, building tooling and operational frameworks that help engineers prepare for, respond to, and learn from failures. In this role you would own the company‑wide on‑call experience, defining best practices and building platforms that support rotations and compensation. You would design software to streamline incident response, collaborate with product teams to improve those tools, and contribute to the post‑mortem process by guiding teams on writing effective reviews and identifying friction points. The position also involves technical leadership: coaching teammates through design reviews, training on‑callers in incident management, and leading cross‑functional initiatives that embed reliability practices across engineering groups. Ideal candidates have at least five years of software engineering experience, strong skills in Go, Python or TypeScript, familiarity with distributed systems and Kubernetes, a track record of analyzing incidents and driving improvements, and experience mentoring engineers and influencing technical direction without formal authority.

senior software engineerincident responseon-call managementpostmortem analysisdistributed systemsgo programmingpython developmentkubernetestechnical mentorshipcross-functional collaboration

We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way The Incident Insights & Readiness SRE team at Datadog fosters a resilient culture by using incidents as learning opportunities and catalysts for growth. Our users are Datadog engineers, and we build the software, tooling, and operational frameworks that help them prepare for, respond to, and learn from incidents. We work closely with engineering teams across Datadog to analyze incidents and turn those insights into better tools, stronger incident response, and organizational learning. Our efforts empower Datadog to navigate unexpected failures confidently, efficiently, and with a commitment to continuous learning and systems improvement. At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them. What You’ll Do: Own and improve the on-call experience for the company by establishing best practices and building platforms to support on-call rotations and compensation. Define how we respond to incidents, lead the design and implementation of software to streamline the process, and collaborate with product teams to improve incident response across Datadog. Our aim is to fully support our incident responders in dealing with complexity. Contribute to the post-mortem process for the company, collaborating with teams on writing them, and identifying opportunities to reduce friction and enhance learning value for the organization. Our team also runs a weekly postmortem reading group. Support various teams in facilitating incident reviews that emphasize learning and blamelessness. Help them share their learnings across the organization to improve the resilience of our people. Provide technical leadership and day-to-day coaching to team members, accelerating their growth through design reviews, collaborative problem-solving and operational excellence best practices. Train our on-callers in incident and post-mortem processes, sharing expertise in incident management best practices. This involves both introducing newcomers to on-call responsibilities and refreshing the knowledge of existing engineers. Lead cross-functional initiatives in engineering organizations across Datadog, embedding with teams to understand their challenges and drive lasting improvements to reliability and operational excellence. Who You Are: At least 5 years of experience building software that solves real user problems. Experience designing new features and collaborating on code and technical design reviews. We primarily develop in Go and Python, with a bit of TypeScript. Experience building or operating distributed systems, with familiarity with Kubernetes and an understanding of complex failure modes. Demonstrated ability to independently own ambiguous technical problems from design through delivery while balancing long-term engineering quality with pragmatic execution. Experience analyzing incidents, identifying systemic risks, and driving engineering improvements informed by operational learnings. Experience participating in on-call rotations and improving incident response processes. Experience serving as an incident commander or incident coordinator is a plus. Empathy, collaboration, and communication skills in English to cultivate strong relationships across various teams in the organization Experience mentoring engineers, driving cross-functional initiatives, and influencing technical direction without relying on organizational authority. We welcome candidates from a variety of backgrounds, including software engineering, site reliability engineering, production engineering, infrastructure, and other roles focused on building reliable systems or improving incident response. Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications on day one. That's okay. If you’re passionate about technology and want to grow your skills, we encourage you to apply. Benefits and Growth: New hire stock equity (RSUs) and employee stock purchase plan (ESPP) Continuous professional development, product training, and career pathing Intradepartmental mentor and buddy program for in-house networking An inclusive company culture, ability to join our Community Guilds (Datadog employee resource groups) Access to Inclusion Talks, our internal panel discussions Free, global mental health benefits for employees and dependents age 6+ Competitive global benefits Benefits and Growth listed above may vary based on the country

RoleSenior Software Engineer - Incident Insights & Readiness
CompanyDatadog
LocationBoston, Massachusetts, USA; New York, New York, USA
Typeinternship
CompensationNot disclosed
Posted2026-10-01
DeadlineRolling

Typical process for this type of role

A general guide — the exact steps for this specific listing may vary; check the original posting for details.

  1. 1ApplicationSubmit your resume through the apply link.
  2. 2ScreeningRecruiter reviews your background against the role.
  3. 3AssessmentA technical test, assignment, or coding round, depending on the role.
  4. 4Interview(s)One or more rounds with the hiring team.
  5. 5OfferOffer letter with compensation and start date.

Before you apply

0/4

About Datadog

Datadog provides a modern monitoring and security platform designed for developers, IT operations teams, and business users operating in the cloud. The company's platform offers a wide range of capabilities, including infrastructure, network, container, and serverless monitoring, as well as application performance monitoring, log management, and cloud security. These tools are utilized by organizations across various industries, such as financial services, healthcare, retail, and technology, to enable digital transformation and drive collaboration across teams. Datadog employs over 8,100 people globally and continues to expand its presence to support customers across diverse markets and regions.

All jobs and hiring details at Datadogcareers.datadoghq.com

More at Datadog

Other jobs at Datadog

Internships at Datadog

See all 250 openings at Datadog

Explore Related Placements


// similar opportunities

You might also like

Senior Software Engineer

Datadog

Datadog logo
Boston, Massachusetts, USA; New York, New York, USAfull-time
Greenhouse

About Datadog: We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We opera...

Apply now →

Senior Software Engineer

Spotify

Torontofull-time
Any GraduateLeverSoftware Development

The Platform team creates the technology that enables Spotify to learn quickly and scale easily, enabling rapid growth in our users and our business around the ...

Apply now →

Senior Software Engineer

Servicenow

Santa Clara, usfull-time
Top Company
Any GraduateSmartRecruitersSoftware Development

No description available.

Apply now →

Senior Software Engineer

Okta

Okta logo
Bengalurufull-time
Greenhouse

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure th...

Apply now →

Senior Software Engineer, Engagement

Discord

San Francisco Bay Areafull-time
B.EGreenhouse# React

Discord has a highly engaged community of millions of daily active users who use the platform for many different reasons, but there’s one thing that nearly ever...

Apply now →

Senior Software Engineer - Identity & Access Management

Brex

Brex logo
San Francisco, California, United Statesfull-time
Greenhouse

Why join us Brex is the intelligent finance platform that enables companies to spend smarter and move faster in more than 200 markets. By combining global corpo...

Apply now →