Clinical Research Scientist, Mental Health AI

San Francisco · On-site$140k – $185k + Equity

About this role

About the Role

We’re looking for a Clinical Research Scientist to help lead and expand our work evaluating AI systems in mental health and other clinically sensitive settings.

As people increasingly use AI for emotional support, health information, and guidance during periods of distress, we need better ways to determine whether these systems respond safely and appropriately. Many important risks emerge over the course of a conversation, including missed signs of escalating distress, reinforcement of harmful beliefs, diagnostic overreach, and unhealthy emotional dependence.

You’ll work closely with our existing research team and engineers to build clinically grounded benchmarks, realistic multi-turn scenarios, scoring criteria, and validation studies. You’ll help define what safe and unsafe model behavior looks like and ensure our evaluations reflect meaningful clinical risks.

This role is a strong fit for a clinical psychologist, psychiatrist, or clinical scientist who wants to help shape a growing research program in mental health AI evaluation while maintaining active academic and clinical collaborations.

What You’ll Do

Lead the clinical design of evaluations for AI systems used in mental health and other sensitive domains.

Identify important clinical failure modes and translate them into realistic scenarios and clear scoring criteria.

Design studies to validate evaluation methods, including clinician review, inter-rater reliability, and comparisons with real-world interaction data.

Work with psychologists, psychiatrists, researchers, and academic collaborators to develop rigorous benchmarks.

Partner with technical researchers and engineers to implement evaluations at scale.

Analyze model behavior, publish findings, and help shape our broader research agenda.

What We Offer

Significant ownership over the direction of our mental health AI research

Close collaboration with clinical, technical, and research colleagues

Strong technical support and resources for large-scale studies

Competitive salary, meaningful equity, and title flexibility based on experience

Relocation and transportation support

Health and dental insurance

Lunch and dinner provided, plus snacks, coffee, and drinks

Unlimited PTO

Further Reading:

Hugging Face blog on evaluation

Anthropic’s blog on challenges in evaluation

New York Times article on issues in benchmarking

Stanford HAI report showing hallucinations in legal tech tools

Company at a glance

Vals AI provides high-quality benchmarks and large-scale evaluations for assessing large language model performance, trusted by foundation model labs and enterprises worldwide. The company combines Stanford-grounded NLP research with expertise from NVIDIA, Meta, Microsoft, and other leading tech firms.

Founded2024
Team Size1-10
WorkspaceOn-site
StagePre-seed
IndustryAI/ML
Location
San Francisco, California, United States
Websitevals.ai
LinkedInLinkedIn

What happens next

Skip the application pile. I get you in front of the people who decide.

Confirm the fit

A few questions to make sure this role is the right shape for you. Two minutes.

I pitch you to the company

I write the intro, send it to the founder, and handle the back-and-forth.

A meeting lands on your calendar

When the company wants to meet, I get the call on your calendar. You just show up.

Know someone who'd be great for this?