InternFlow

Greenhouse ·

contract

Staff Data Infrastructure Engineer

Faire · New York City, NY; San Francisco, CA

Faire operates a wholesale marketplace platform designed to connect independent retailers with brands globally through the use of data and machine learning. This Staff Data Infrastructure Engineer role focuses on the architecture and reliability of the company's data pipelines. The successful candidate will lead the technical direction for moving data from production systems like CockroachDB and MySQL into analytical stores, specifically focusing on building robust CDC and streaming ingestion layers. Daily responsibilities include designing scalable infrastructure, implementing data contracts and quality checks, managing Airflow and Fivetran workflows, and ensuring data reliability through SLOs and incident management. This is a hands-on position that requires deep technical expertise in lakehouse architectures, particularly Apache Iceberg, Spark, and Kafka. The role is well-suited for an experienced engineer who has previously owned data infrastructure at scale, possesses strong opinions on build-versus-buy decisions, and is capable of mentoring team members while driving complex migrations without disrupting existing business operations.

data infrastructuredata engineeringstaff engineerapache icebergstreaming ingestioncdcdata platformsparkkafkadata reliability

About Faire Faire is a technology wholesale platform built on the belief that the future is local. Independent retailers around the globe collectively represent a multi-hundred-billion-dollar wholesale market that has historically been fragmented and offline. At Faire, we're using the power of tech, data, and machine learning to connect this thriving community of entrepreneurs across the globe. Picture your favorite boutique in town — we help them discover the best products from around the world to sell in their stores. With the right tools and insights, we believe that we can level the playing field so businesses can grow and local communities can thrive. We’re looking for smart, resourceful and passionate people to join us as we power the shop local movement. If you believe in community, come join ours. About this role: Our Engineering organization owns the software that makes our marketplace work. The Data Platform group supports everyone at Faire who depends on data: Product Engineering, Data Science, Machine Learning, Analytics, Strategy, Finance, and Product. Our job is to make sure the data is there, it's right, and people can find it and query it without having to think about the plumbing underneath. We are hiring a Staff Engineer to own that plumbing. Concretely, this means the path data takes out of our production databases (CockroachDB and MySQL) and into a place where analysts and data scientists can query it. Today that involves Fivetran, Kafka, Spark, and Airflow landing data in Snowflake and Databricks. It works, but it grew up over time and it shows. We want someone who can design the next version of it, build the hard parts personally, and bring the rest of the company along. This is a hands-on role. You'll also be the person other teams come to when they need to know how data should move at Faire. What you'll do: Set the technical direction for how data moves from production systems into our analytical stores, and own the roadmap to get there over the next couple of years. Build the CDC and streaming ingestion layer: CockroachDB changefeeds and MySQL binlogs into Kafka, then into Iceberg tables on S3. You'll be responsible for the hard details like ordering, deduplication, late data, schema changes, and backfills. Implement data contracts and quality checks throughout our platform Put real ownership and SLAs on the datasets the business runs on, and wire quality checks into the platform with tools like Anomalo and Monte Carlo so we hear about broken data before a dashboard or a model does. Run Airflow and Fivetran well, and have an opinion about what we should keep buying versus what we should build. Own reliability for the platform: SLOs, on-call, incident reviews, and the follow-through so the same thing doesn't break twice. Work with the senior engineers, data scientists, and analysts who depend on this platform, and lead the migration of existing pipelines onto the new one without breaking what they rely on. Mentor the engineers around you. We want the team's data engineering practice to be better because you were here. Qualifications: You've built and run data infrastructure that other teams depended on, at meaningful scale, and you've been the person setting direction for it, not just working on it. Deep experience with change data capture and streaming ingestion from operational databases through Kafka. You know what goes wrong with ordering, duplicates, snapshots, and schema evolution because you've dealt with it. Hands-on experience with lakehouse architectures on an open table format. Iceberg on S3 is what we use, so that's especially valuable. You should be comfortable talking about partitioning, compaction, catalogs, and copy-on-write versus merge-on-read. Strong Spark skills, and experience running Databricks and Snowflake against shared storage. Experience with data quality and observability in practice, including data contracts, SLAs, and tools like Anomalo or Monte Carlo. Experience operating Airflow at scale and working with managed ingestion like Fivetran. Strong SQL, and good instincts for how to model data so analysts and data scientists can actually use it. Solid Python plus at least one of Kotlin, Java, Scala, or Go. Experience shipping infrastructure on AWS with Terraform. A working understanding of data governance: access control, PII, retention and deletion, lineage, and audit. A track record of leading cross-team data initiatives and migrations, and of mentoring senior engineers. You can explain a technical tradeoff to a leadership team and to a new grad, and you can get people who disagree with each other to a decision. You take ownership of things that are broken or unowned, and you're willing to be on call for the systems you build. Experience in a marketplace, e-commerce, or other transaction-heavy business is a plus. Technologies we use and teach: Python, Kotlin, SQL Kafka, Fivetran, Airflow S3, Apache Iceberg, Snowflake, Databricks, Apache Spark AWS, Terrafo

RoleStaff Data Infrastructure Engineer
CompanyFaire
LocationNew York City, NY; San Francisco, CA
Typecontract
CompensationNot disclosed
Posted2026-10-01
DeadlineRolling

Typical process for this type of role

A general guide — the exact steps for this specific listing may vary; check the original posting for details.

  1. 1ApplicationSubmit your resume through the apply link.
  2. 2ScreeningRecruiter reviews your background against the role.
  3. 3AssessmentA technical test, assignment, or coding round, depending on the role.
  4. 4Interview(s)One or more rounds with the hiring team.
  5. 5OfferOffer letter with compensation and start date.

Before you apply

0/4

Staff Data Infrastructure Engineer at Faire: frequently asked questions

What is the salary for this role?
Faire has not stated compensation in the listing. Check the original posting or ask during the application process.
Is the Staff Data Infrastructure Engineer position remote, hybrid or onsite?
The listing gives New York City, NY; San Francisco, CA as the location and does not state a work mode.
What is the application deadline?
Faire has not listed a fixed deadline, so apply early in case the opening is filled.
How do I apply for the Staff Data Infrastructure Engineer role?
Use the Apply button on this page. It opens the original listing on boards.greenhouse.io, where you submit your application with the company.

More at Faire

Other jobs at Faire

See all 31 openings at Faire

Explore Related Placements


// similar opportunities

You might also like

Data Center Engineer, Reliability & Infrastructure Management – Compute Supply

Anthropic

San Francisco, CA | New York City, NYcontract
Greenhouse

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for s...

Apply now →

Staff Detection Engineer

Asana

Asana logo
Warsawcontract
Greenhouse

Our Security team keeps Asana's employees, users, and customers safe by proactively addressing threats and fostering a culture of security across our product an...

Apply now →

Software Engineer, Data Loading Infrastructure

Asana

Asana logo
San Franciscofull-time
Greenhouse

The Data Loading Infrastructure ( LunaDb ) team establishes core foundational architecture required to power web, mobile, and external API platforms, helping de...

Apply now →

Staff Software Engineer - NLP

Gitlab

Bangalorecontract
RemoteAny GraduateGreenhouse# AI# NLP

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency,...

0-1 year experience eligible.

Apply now →

Staff Frontend Engineer

Gitlab

Bangalorecontract
Remote12th / InterGreenhouse# JavaScript# TypeScript# React# Angular+7 more

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency,...

0-1 year experience eligible.

Apply now →

Senior Software Engineer, Data Foundation

Mixpanel

San Francisco, US (Hybrid)contract
Greenhouse

About Mixpanel Mixpanel is the leading product intelligence and analytics platform, trusted by more than 29,000 companies to help understand how people use the ...

Apply now →