Greenhouse ·
full-timeSenior Applied Scientist - AI Platform
Datadog · Paris, France
AI Platform builds the foundations of Datadog's AI efforts. The org is 70+ people organised in three pillars: training and serving (GPU clusters, distributed training, low-level infrastructure), agents (agent harnesses, memory systems, the internal AI gateway that routes every LLM request at Datadog), and evaluation and experimentation. This role sits in the evaluation and experimentation pillar, which owns Datadog's shared annotation and evaluation infrastructure — including the evaluation scenario store and the telemetry archival systems used across the Bits org. Together they let an agent travel back in time and query what Datadog looked like at the exact moment an incident happened, so scenarios can be replayed and agent performance tracked over time. Specifically, you'll be the first applied scientist on GenSim (Generative Simulations), the team that builds the environments Datadog's agents learn in. GenSim doesn't replay sampled telemetry — it stands up real, fully instrumented applications that talk to Datadog, drives them with representative traffic, injects controlled failures, and records what happens. Because GenSim injected the failure, it knows the ground truth. That corpus — hundreds of postmortem-derived scenarios and thousands of runnable applications — is today the primary source of post-training data for Datadog's own SRE model, and the substrate that Bits AI SRE and our other agents are trained and evaluated against. The team has no applied science support today and is learning post-training data methodology on the fly. That's the gap this role fills, and the open questions are the interesting part. How do you tell whether a generated environment is actually representative of the messy, incomplete telemetry real customers run — rather than a suspiciously clean one where every monitor exists and every service emits complete logs? How do you make injected problems genuinely hard, and how do you even measure difficulty? How do you define and control the quality of post-training data when correctness, representativeness and difficulty pull in different directions? How do you evaluate an agent end to end when the trajectory is non-deterministic? Creating simulated agent environments for monitoring and SRE work is not well solved in open source or in published research, and Datadog is the leading company in this field. If those are the problems you want to spend your time on, come build this with us. What You'll Do: Own the applied science direction for GenSim: set the methodology and the forward-looking technical calls on how simulated environments and post-training data should be built, on a team where that decision-making does not exist yet. Define, measure and raise the quality of post-training data — basic correctness, representativeness against the real distribution of customer systems and production telemetry, and difficulty — and make those measures something the team can act on release over release. Close the realism gap. Simulated environments today are too clean and the injected problems are not yet hard enough; you'll drive the research and the engineering that make them look like real, imperfect production systems. Build scalable, production-grade systems rather than research scripts. The output is not just a dataset — it is a system of synthetic environments that must be reliable and invokable inside a training loop. Determine how this data is best applied, in LLM post-training and in evaluation, and own the agent and LLM application evaluation approaches for these environments. Work cross-functionally with the engineers and applied scientists on adjacent teams — Bits AI SRE, the model training effort, and the wider evaluation and experimentation pillar — so that what you learn moves freely in both directions. Who You Are: You have a PhD, MS or equivalent research experience in a scientific field, with strong applied mathematics grounding. 6+ years of relevant applied science or ML engineering experience, including setting technical direction for others. You have hands-on experience with LLM and agent post-training data: how it is created, managed, and how training-data quality is controlled. This is the requirement that matters most. You have real domain expertise in LLMs and agentic applications — not classical ML fine-tuning. Fine-tuning classifiers or traditional models is a different problem from the one this team is solving. You have evaluated agents or LLM applications, and can define what 'good' means before you measure it. You are a strong programmer and production software engineer. Python at minimum, plus the ability to ship scalable production systems and work with distributed systems. You collaborate well across engineering and science teams, and you're comfortable being the domain expert who decides what comes next. You thrive in ambiguity and can make sound technical calls when the path isn't yet defined. Bonus Points: Hands-on LLM fine-tuning, post-training or mode
| Role | Senior Applied Scientist - AI Platform |
|---|---|
| Company | Datadog |
| Location | Paris, France |
| Type | internship |
| Compensation | Not disclosed |
| Posted | 2026-09-29 |
| Deadline | Rolling |
Typical process for this type of role
A general guide — the exact steps for this specific listing may vary; check the original posting for details.
- 1ApplicationSubmit your resume through the apply link.
- 2ScreeningRecruiter reviews your background against the role.
- 3AssessmentA technical test, assignment, or coding round, depending on the role.
- 4Interview(s)One or more rounds with the hiring team.
- 5OfferOffer letter with compensation and start date.
Before you apply
0/4About Datadog
Datadog provides a modern monitoring and security platform designed for developers, IT operations teams, and business users operating in the cloud. The company's platform offers a wide range of capabilities, including infrastructure, network, container, and serverless monitoring, as well as application performance monitoring, log management, and cloud security. These tools are utilized by organizations across various industries, such as financial services, healthcare, retail, and technology, to enable digital transformation and drive collaboration across teams. Datadog employs over 8,100 people globally and continues to expand its presence to support customers across diverse markets and regions.
More at Datadog
Other jobs at Datadog
- Senior Group Manager, Experience Design · New York, New York, USA
- Event Marketing Manager - Global Sponsorships (namer/latam) · New York, New York, USA
- Senior Software Engineer, Chaos Engineering · Paris
- Manager, Revenue Accounting · New York, New York, USA
- Field Enablement Manager (emea) · Amsterdam, The Netherlands; Munich, Germany
Internships at Datadog
- Applied Science Intern · Paris, France
- Software Engineering Intern · Paris, France
- IT Support Technician Intern · Paris, France
- Product Management Intern · New York, New York, USA
- Research Science Intern (phd) · New York, New York, USA; Pittsburgh, Pennsylvania, USA
Explore Related Placements
// similar opportunities
You might also like
Senior Applied Scientist
Datadog
The Applied AI team designs and builds algorithmically driven features in the Datadog app. We work across a range of applications, primarily focusing on analysi...
Senior Data Scientist - Applied ML
Justworks
Who We Are At Justworks, you’ll enjoy a welcoming and casual environment, great benefits, wellness program offerings, company retreats, and the ability to inter...
Senior Applied AI/ML Scientist - Search Ranking
Faire
About Faire Faire is a technology wholesale platform built on the belief that the future is local. Independent retailers around the globe collectively represent...
Senior Applied AI/ML Scientist - Brand Growth
Faire
About Faire Faire is a technology wholesale platform built on the belief that the future is local. Independent retailers around the globe collectively represent...
Senior Software Engineer - Action Platform
Datadog
We are looking for a strong technical leader to join the Private Action Runner team, part of the larger Action Platform group and help shape one of the core exe...
Senior Applied AI/ML Scientist - Listing Quality
Faire
About Faire Faire is a technology wholesale platform built on the belief that the future is local. Independent retailers around the globe collectively represent...