Forward Deployed Engineer
Inference performance depends on how efficiently models use the underlying hardware. Techniques across kernels, compilers, runtimes, and serving systems can dramatically improve latency and throughput. None of that matters until it is running in a customer's production traffic.
At Wafer, we are building AI systems that automatically optimize inference workloads across silicon. The goal is fungible token capacity. Any accelerator optimized toward serving inference most efficiently.
Wafer is well funded and serves trillions of tokens a day for mission critical workloads. We serve the highest performance inference to fast-growing AI startups.
Forward Deployed Engineers serve as the link between Wafer's core product and customers' challenges, managing the relationship from start to finish.
What you'll do
Own customer accounts end to end. Sales gets the first meeting. From there you are both the customer's engineer and their point of contact: you decide what to prove, you build it, you keep it running in production, and you carry the relationship.
Win the technical evaluation. Prove Wafer on the customer's own workload rather than a synthetic benchmark, and be the person who can explain the result to their engineers.
Ship the integration. Stand up dedicated deployments and optimize the full stack for the customer's workload.
Own production for your accounts. When latency moves or error rates climb, you find it, you fix it or route it, and you are who the customer hears from.
Inform product roadmap. You sit closer than anyone to how Wafer behaves under real load, and the engineering team builds against what you report.
What we look for
You are an exceptional engineer. You have shipped and operated production systems, and you can still open a profiler, read a trace, and find the problem yourself. Inference experience helps, but it matters more that you can be handed an unfamiliar system and own it inside a week.
You have owned a customer relationship. You can name an account that was yours, say what was going wrong, and describe what you personally did about it.
You are credible in a room full of engineers. You can take a skeptical ML team through a benchmark, defend the methodology, and concede the point when they are right.
You work without a spec. The problem arrives half-defined from a customer who does not yet know what they need, and you come back with a scoped answer rather than a list of questions.
How we evaluate
We score every candidate on seven values:
Infinitely Resourceful
Exceptionalism
Unreasonable Standards
Company Over Self
High EQ
Learns Quickly
First Principles Thinker
Compensation and benefits
$180-240K base salary + generous equity.
Fully covered medical, dental, and vision insurance.
Daily lunch and dinner, unlimited PTO, and parental leave.
$1K/month housing stipend (post-tax) if you live within walking distance (0.5 miles) from the office.
Covered Uber/Waymo from/to office.
Visa sponsorship available.
How we work
On-site in San Francisco, five days a week. Small team with massive surface area and ownership. You operate with complete autonomy of how to solve problems. We don't see engineers as code writers, but as problem solvers.
Forward deployed here means deployed into the customer's stack, their Slack, and their problems. Most of that happens from San Francisco but sometimes that means working with the customer on site.