Clera home
·Dashboard

Jobs at Nous Research (Now Hiring) — 7 open

Nous Research

Machine Learning Engineer, Evals

United States · Remote OK

Mid level

The Role You'll work across the lab on agent capability evals, benchmark design, LLM-as-judge systems, failure analysis, and the infrastructure that ties it together. This is a high-growth, high-ownership role on a small…

Skills: Machine Learning Engineering, LLM Evaluation, Python, Prompt Engineering, Failure Analysis

Nous Research

Finance & Ops Lead

New York, New York, United States · Hybrid

Mid level

The Role As our first Finance & Operations hire, you'll work directly with company leadership to help scale the business behind Nous Research. You'll own financial planning, operational execution, and many of the interna…

Skills: Financial Modeling, Strategic Planning, Business Operations, Budgeting, Forecasting

Nous Research

Product Analytics Engineer

United States · Remote OK

Senior

The Role As a Product Analytics Engineer at Nous Research, you'll own the measurement systems that help us understand how users interact with Hermes Agent. You'll build the analytics foundation across our products—from i…

Skills: SQL, Python, TypeScript, JavaScript, Product Analytics

Nous Research

Forward Deployed Engineer

United States · Remote OK

Senior

The Role As a Forward Deployed Engineer at Nous Research, you'll deploy and adapt Hermes Agent Enterprise inside complex customer environments. You'll partner directly with enterprise customers to understand their workfl…

Skills: Software Engineering, Solutions Engineering, Distributed Systems, Cloud Infrastructure, API Integration

Nous Research

UI/UX Designer

United States · Remote OK

Mid level

The Role As a UI/UX Designer at Nous Research, you'll help define the future of human-AI interaction across Hermes Agent and Nous Portal. You'll design intuitive experiences for complex agent workflows, explore new inter…

Skills: UI/UX Design, Product Design, Interaction Design, Visual Design, Figma

Nous Research

Software Engineer, GUI/Product

United States · Remote OK

Senior

The Role Join the Product team building the customer-facing surfaces of Nous Portal, Hermes Agent, and the products around them across mobile, desktop, and web. This role centers on the everyday Hermes Agent experience p…

Skills: TypeScript, Node.js, React, Python, Rust

Nous Research

Software Engineer, Hermes Cloud

United States · Remote OK

Mid level

The Role As a Software Engineer on the Hermes Cloud team, you'll own full-stack engineering across the cloud platform that powers Hermes Agent and Nous Portal. You'll design and build the infrastructure, deployment syste…

Skills: Full-stack Engineering, Cloud Infrastructure, TypeScript, Node.js, Python

Machine Learning Engineer, Evals

Nous Research

United States • Remote OK

Apply
Mid level

Tired of cold applications?

Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.

  • Full-time
  • Posted 13d ago
  • ~40 hrs/week
  • Remote in United States

Responsibilities

Develop and maintain evaluation infrastructure, including benchmark design and LLM-as-judge systems, to assess agent capabilities. Conduct failure analysis on model outputs and manage recurring evaluation workflows to support research efforts.

Requirements

Requires 3+ years of experience in software or ML engineering with a strong background in LLM evaluation frameworks and Python. Candidates must be proficient in basic evaluation statistics and have hands-on experience with prompting and agent benchmarks.

Full job description

The Role

You'll work across the lab on agent capability evals, benchmark design, LLM-as-judge systems, failure analysis, and the infrastructure that ties it together. This is a high-growth, high-ownership role on a small team, and you'll ship evaluation infrastructure that researchers depend on from day one.

Responsibilities

  • Run the full eval pipeline end to end and reproduce known results during onboarding, pairing with a senior engineer on your first task

  • Build a judge calibration protocol: sample human-labeled decisions, measure agreement (κ, per-class P/R), identify drift zones, and document it so anyone can re-run it

  • Extend an existing benchmark (GAIA, τ-Bench, SWE-bench slice, etc.) with new tasks targeting known capability gaps, including the prompt, environment, rubric, automated grader, and QA

  • Run failure analysis on model outputs: categorize failure modes, quantify prevalence, and write up findings with recommendations for training data, judge prompts, or benchmark changes

  • Own a recurring eval workflow (weekly regression suite, judge drift dashboard, red-team evaluation for a new capability) and ship tooling researchers actually use

Qualifications

  • 3+ years in software engineering, ML engineering, data science, or a research-adjacent role, with concrete evaluation experience from coursework, an internship, a side project, open source work, or a job

  • Experience with at least one LLM evaluation framework (Harbor, Nemo Evaluator, etc.), with real opinions on what it does well and where it falls short

  • Hands-on experience with LLMs: prompting, few-shot design, and ideally fine-tuning or RAG; regular use of coding agents

  • Solid Python. You write clean, tested, version-controlled code that a colleague could run without you babysitting it

  • Comfort with Git, CI/CD basics, Docker, and the Linux command line (SSH, tmux, debugging a remote job)

  • Understanding of basic eval statistics: why accuracy misleads on imbalanced judges, what Cohen's κ measures, how to think about confidence intervals on a metric

  • At least 3 of the following: you can explain why LLM-as-judge needs calibration; you've done failure analysis and can tell model bugs apart from prompt, grader, or retrieval issues; you know at least two agent benchmarks (GAIA, AgentBench, τ-Bench, MINT, SWE-bench, WebShop, ALFWorld) and a limitation of each; you've designed or extended an eval dataset with happy paths, edge cases, and adversarial examples; you've thought about non-determinism in eval, how you sample, how many runs, how you report variance

  • You communicate clearly to both researchers and engineers, in the right language for each

  • You're comfortable with ambiguity, can turn a half-formed request into a plan, and know when to ask for help

Preferred

  • RLVR / RLHF pipeline experience

  • Training data curation experience

  • Distributed eval orchestration experience

  • Benchmark design from scratch

  • Red teaming and adversarial eval experience

  • Familiarity with psychometrics or measurement theory

Related keywords

Machine LearningLLMEvalsGAIATau-BenchSWE-benchPythonDockerCI/CDGitRAGRLHFRLVRPsychometricsMeasurement TheoryAgent Capability

About Nous Research

LinkedInVisit site

The AI Accelerator Company https://discord.gg/nousresearch

Industry
Blockchain Services
Company size
51-200 employees
View all jobs at Nous Research

About Nous Research

LinkedInVisit site

The AI Accelerator Company https://discord.gg/nousresearch

Industry
Blockchain Services
Company size
51-200 employees
View all jobs at Nous Research

Similar companies hiring

OneBullEx (11)Somnia (11)TRG | Technology Research Group Ltd. (8)Contango (7)Coinspaid Solutions (7)Tatum (6)Creatify Labs (3)Halliday (3)Quicknode (3)Sui Foundation (3)Cardano Foundation (2)Reown (2)
Clera home

Your AI-talent agent. Connecting talents with dream jobs.

Earn $5,000

Tools

  • Salary Calculator
  • Resume Review
  • Startup Map

Explore

  • Jobs
  • Discover Jobs
  • Companies
  • Referral

Platform

  • Pricing
  • Integrations
  • Partners
  • Acquihire

Clera

  • Manifesto
  • Engineering
  • We are hiring!
  • FAQs
  • Blog
  • Press

Tools

  • Salary Calculator
  • Resume Review
  • Startup Map

Explore

  • Jobs
  • Discover Jobs
  • Companies
  • Referral

Platform

  • Pricing
  • Integrations
  • Partners
  • Acquihire

Clera

  • Manifesto
  • Engineering
  • We are hiring!
  • FAQs
  • Blog
  • Press

© 2026 Clera Labs, Inc.

PrivacyTermsBug Bounty