InternFlow

Greenhouse ·

full-time

Staff + Sr. Software Engineer, Scaling

Anthropic · New York City, NY; San Francisco, CA; Seattle, WA

Anthropic is an AI safety and research company focused on developing reliable, interpretable, and steerable artificial intelligence systems. This role within the Inference team centers on building and scaling the infrastructure required to serve the Claude model to a global user base. Day-to-day responsibilities involve designing and maintaining high-performance distributed systems, developing intelligent request routing and load balancing mechanisms, and managing fleet-wide orchestration across diverse AI accelerators and cloud platforms. Engineers in this position will work to maximize compute efficiency while providing the necessary infrastructure to support research breakthroughs. The role requires a strong background in distributed systems and a willingness to tackle complex challenges related to networking, autoscaling, and deployment pipelines. It is well-suited for experienced software engineers who are comfortable working in performance-sensitive environments and have a desire to optimize large-scale machine learning infrastructure. Candidates should be prepared to operate across the entire stack, from hardware-level integration to high-level traffic management, ensuring that production systems remain resilient and scalable as the company continues to grow.

distributed systemsinference infrastructuremachine learning systemssoftware engineeringcloud infrastructureload balancingllm optimizationkubernetes

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role Our Inference team is responsible for building and scaling the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry’s largest compute-agnostic inference deployments. We are responsible for the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators. The team has a dual mandate: maximizing compute efficiency to reliably serve our explosive customer growth, while enabling breakthrough research by giving our scientists the high-performance inference infrastructure they need to develop next-generation models. We tackle complex, distributed systems challenges across multiple accelerator families and emerging AI hardware running in multiple cloud platforms. Inference systems are highly performance sensitive distributed systems. Inference serves hundreds of thousands of customers every day, and the size & span of the inference fleet requires sophisticated routing, scaling, and networking systems. Key responsibilities Design, build, and maintain the distributed systems that serve Claude to millions of users worldwide Develop resilient, flexible systems that adapt in real time to real world events Develop intelligent request routing, load balancing, and traffic management systems across thousands of accelerators and multiple cloud providers Maximize compute efficiency and optimize cost across the fleet by autoscaling and orchestrating production, research, and experimental workloads across multiple cloud providers Build and operate production-grade deployment pipelines for releasing new models to users Provide high-performance inference infrastructure that enables researchers to develop next-generation models Integrate new AI accelerator platforms and support inference for new model architectures Minimum qualifications Significant software engineering experience, particularly with distributed systems Results-oriented, with a bias towards flexibility and impact Willingness to pick up slack, even if it goes outside your job description Desire to learn more about machine learning systems and infrastructure Thrive in environments where technical excellence directly drives both business results and research breakthroughs Care about the societal impacts of your work Preferred qualifications Experience with high-performance, large-scale distributed systems Experience implementing and deploying machine learning systems at scale Experience with load balancing, request routing, or traffic management systems Familiarity with LLM inference optimization, batching, and caching strategies Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure) Proficiency in Python or Rust Representative projects Designing intelligent routing algorithms that optimize request distribution across many accelerators in different environments Autoscaling our compute fleet to dynamically match supply with demand across production, research, and experimental workloads Building production-grade deployment pipelines for releasing new models to millions of users reliably Contributing to new inference features Supporting inference for new model architectures Analyzing observability data to tune performance based on real-world production workloads Managing multi-region deployments and geographic routing for global customers The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary: $320,000 — $485,000 USD Logistics Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as l

RoleStaff + Sr. Software Engineer, Scaling
CompanyAnthropic
LocationNew York City, NY; San Francisco, CA; Seattle, WA
CompensationNot disclosed
Posted2026-09-28
DeadlineRolling

Typical process for this type of role

A general guide — the exact steps for this specific listing may vary; check the original posting for details.

  1. 1ApplicationSubmit your resume through the apply link.
  2. 2ScreeningRecruiter reviews your background against the role.
  3. 3AssessmentA technical test, assignment, or coding round, depending on the role.
  4. 4Interview(s)One or more rounds with the hiring team.
  5. 5OfferOffer letter with compensation and start date.

Before you apply

0/4

Staff + Sr. Software Engineer, Scaling at Anthropic: frequently asked questions

What is the salary for this role?
Anthropic has not stated compensation in the listing. Check the original posting or ask during the application process.
Is the Staff + Sr. Software Engineer, Scaling position remote, hybrid or onsite?
The listing gives New York City, NY; San Francisco, CA; Seattle, WA as the location and does not state a work mode.
What is the application deadline?
Anthropic has not listed a fixed deadline, so apply early in case the opening is filled.
How do I apply for the Staff + Sr. Software Engineer, Scaling role?
Use the Apply button on this page. It opens the original listing on job-boards.greenhouse.io, where you submit your application with the company.

More at Anthropic

Other jobs at Anthropic

See all 122 openings at Anthropic

Explore Related Placements


// similar opportunities

You might also like

Sr Staff Software Engineer

Servicenow

Hyderabad, infull-time
Top Company
SmartRecruiters

No description available.

Apply now →

Staff Software Engineer

Okta

Okta logo
Bengalurufull-time
Greenhouse

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure th...

Apply now →

Staff Software Engineer

Okta

Okta logo
Bengalurufull-time
Greenhouse

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure th...

Apply now →

Staff Software Engineer

Okta

Okta logo
Toronto, Ontario, Canadafull-time
Greenhouse

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure th...

Apply now →

Staff Software Engineer

Amplitude

Vancouver, BC, Canadafull-time
Any GraduateGreenhouseSoftware Development# Query Planning# Columnar Storage# Encoding# Compression+8 more

Amplitude is the leading AI analytics platform, helping over 4,700 customers—including Atlassian, Burger King, NBCUniversal, and Square—build better products an...

Apply now →

Staff Software Engineer

Okta

Okta logo
Bengalurufull-time
Any GraduateGreenhouseSoftware Development# Java# C## Object-Oriented Programming# Software Architecture+5 more

Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure th...

Apply now →