Machine Learning Engineer - Evals

Location
New York
Workplace
On-site
Compensation
$220k – $300k + equity

About this role

You'll own evaluation of Honcho end to end, the instrument that tells a research-driven team whether their identity representations are actually getting better. Plastic Labs builds the memory and identity layer for the agentic world, and the product works. It's a complex multi-agent harness with real black-box behavior, so the job is defining what "better" means for representations that change over time, without ground truth, and building the machinery that measures it.

If you love constructing measurements from nothing, reading traces to find what's actually broken, and shipping the fix yourself, you'll feel at home here.

What You'll Do

Design the evals. Define the scores and decide what "better" means for an entity representation that changes over time, and keep that definition current as the product and methods move.

Build the pipelines and harnesses. Data in, labels, versions, reruns, judges: the machinery that lets the team ask a new question this week and get an answer this week.

Run them and harvest insights. Read the traces and results, find what's actually broken, not what's easy to measure, and propose fixes that produce higher-fidelity representations.

Build simulation agents at scale. The ceiling on iteration speed is how many entities can be modeled and measured at once. You raise it.

Own the loop end to end. The question, the pipeline, the rerun, the writeup. Nothing gets scoped and handed off, and nothing waits on someone else's sprint.

What happens next

Skip the application pile. I get you in front of the people who decide.

Confirm the fit

A few questions to make sure this role is the right shape for you. Two minutes.

I pitch you to the company

I write the intro, send it to the founder, and handle the back-and-forth.

A meeting lands on your calendar

When the company wants to meet, I get the call on your calendar. You just show up.

Culture & values

Engineering-driven AI lab culture

High-touch, collaborative culture

Five days a week in-person work at Williamsburg office in Domino Refinery, Brooklyn

Flat organizational structure with everyone reporting directly to the CTO

Culture prizes high-agency autodidacts who thrive with broad freedom and significant responsibility

Values intellectual diversity and interdisciplinary curiosity spanning cognitive science, linguistics, neuroscience, and philosophy

Techno-optimistic mindset

Engineers own end-to-end projects

Team builds in the open

Daily teaching and learning among team members

Engineers engage with users in Discord and act as thought partners on product direction

Generalist mindset where everyone wears multiple hats

Know someone who'd be great for this?