Research Engineer, Benchmarks

Location
San Francisco, Singapore
Workplace
On-site
Compensation
$150k – $250k
Visa
Visa Sponsorship Available

About this role

HUD builds high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. Joining a small, technical team of researchers and engineers—including International Olympiad medalists and published AI researchers—you’ll own the design and implementation of evaluations that frontier labs and customers trust. This role is critical to ensuring HUD’s benchmarks are rigorous, credible, and aligned with real-world agent performance.

What you'll do

  • Design, implement, and own the quality of HUD’s internal benchmarks for evaluating frontier agents on domain-specific tasks.
  • Partner with subject-matter experts to define realistic workflows and tasks for domain-specific evaluations.
  • Build reliable infrastructure to run models and agents against benchmark tasks at scale.
  • Develop metrics and analyses that measure benchmark difficulty, reliability, and failure modes.
  • Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations.
  • Write clear documentation and benchmark reports that make results legible and credible to technical audiences.

What HUD is looking for

  • 2–4 years of experience in software engineering, ML engineering, or research roles.
  • Strong proficiency in Python, Docker, and Linux environments.
  • Experience building environments, evaluations, or benchmarks for AI systems.
  • Published research or technical writing on topics such as public benchmarks, model failure modes, or evaluation methodology.
  • Deep understanding of what makes a benchmark realistic, reliable, and practically useful.
  • Curiosity and ability to truly understand how workflows operate across various domains.
  • Strong attention to detail with a habit of spotting subtle inconsistencies and edge cases in tasks.
  • Ability to reason from first principles about task design, scoring, and failure modes.
  • Comfort thriving in unstructured problem spaces and working independently in fast-paced, early-stage startup environments.
  • Excellent communication skills for collaborating across time zones and technical teams.

What happens next

Skip the application pile. I get you in front of the people who decide.

Confirm the fit

A few questions to make sure this role is the right shape for you. Two minutes.

I pitch you to the company

I write the intro, send it to the founder, and handle the back-and-forth.

A meeting lands on your calendar

When the company wants to meet, I get the call on your calendar. You just show up.

Know someone who'd be great for this?

Top Benefits

  • 100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA
  • Lunch and dinner when you’re in the office
  • Company-wide holiday break (Christmas Eve to New Year’s Day) on top of PTO and paid holidays
  • Other perks including an Equinox membership, 401k, and commuter benefits
  • Unlimited* access to tokens for ChatGPT, Claude Code, Cursor, etc. *By unlimited, we mean no one on our token usage leaderboard has ever hit a limit. So we have no idea what the limit is.
  • We provide support for relocation and visas for strong full-time candidates to the US.