Clera home
·Dashboard

Jobs at Vals AI (Now Hiring) — 1 open

Vals AI logoVals AI

Clinical Research Scientist, Mental Health AI

San Francisco, California, United States · On-site

$140k–$185k/yr

Senior

About the Role We're looking for a Clinical Research Scientist to help build the next generation of evaluations for AI systems used in mental health and other clinically sensitive settings. As large language models becom…

Skills: Clinical research, AI evaluation, Behavioral science, Study design, Data collection

Vals AI logo

Clinical Research Scientist, Mental Health AI

Vals AI

San Francisco, California, United States • On-site

Apply
Senior

Tired of cold applications?

Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.

  • $140k–$185k/yr
  • Full-time
  • postgraduate degree
  • Competitive salary, Equity, Relocation support, Transportation support, Health insurance, Dental insurance
  • Posted 2d ago
  • ~40 hrs/week

Responsibilities

You will design clinician-informed benchmarks and evaluation methodologies to assess AI systems in mental health and sensitive clinical domains. Additionally, you will collaborate with researchers and engineers to conduct validation studies and publish research findings.

Requirements

Candidates must hold a PhD, PsyD, MD, or equivalent research degree in a field such as Clinical Psychology, Psychiatry, or Behavioral Science. You must demonstrate a strong track record of independent research, scientific publication, and collaborative project leadership.

Full job description

About the Role

We're looking for a Clinical Research Scientist to help build the next generation of evaluations for AI systems used in mental health and other clinically sensitive settings. As large language models become part of how people seek advice, emotional support, and health information, there is a growing need for rigorous ways to understand how these systems behave in real-world conversations. Many of the most important questions including how models respond to psychological distress, uncertainty, or vulnerable users, can't be answered with traditional AI benchmarks alone. They require clinical expertise, careful study design, and realistic evaluations grounded in human behavior.

You'll work with researchers and engineers to design clinician-informed benchmarks, develop new evaluation methodologies, and build datasets that measure model behavior in realistic, multi-turn interactions. The role combines clinical research, behavioral science, and AI evaluation, with opportunities to publish, collaborate with leading universities and help shape emerging standards for evaluating AI.

We welcome applicants from academia, hospitals, nonprofit research institutes, and digital health organizations who are excited to bring their research into industry while continuing to publish and collaborate with the broader research community.

What You'll Do

  • Design clinician-informed benchmarks, datasets, and evaluation methods for AI systems used in mental health and other clinically sensitive domains.

  • Design and conduct validation studies to ensure benchmark performance reflects real-world model behavior.

  • Partner with psychologists, psychiatrists, researchers, and academic collaborators to identify important evaluation problems and translate them into rigorous benchmarks.

  • Build collaborative research projects with universities, hospitals, and nonprofit organizations.

  • Analyze model behavior, publish research findings, and communicate results through technical reports and presentations.

  • Collaborate with research engineers to implement large-scale evaluation pipelines and benchmark infrastructure.

  • Help shape our research agenda in mental health AI evaluation and identify emerging research directions.

Requirements

Research background: PhD, PsyD, MD, or equivalent research experience in Clinical Psychology, Psychiatry, Behavioral Science, Public Health, or a related field.

Research experience: Demonstrated experience designing and leading research projects, including study design, data collection, statistical analysis, and scientific writing.

Publications: Track record of publishing independent research in peer-reviewed journals or conferences.

Collaborative research: Demonstrated ability to initiate and lead collaborative research with external partners, including universities, hospitals, or other research organizations.

Research methods: Strong understanding of behavioral research methods, human subjects research, survey design, psychometrics, qualitative or quantitative methods, or clinical study design.

Communication: Excellent written and verbal communication skills, including experience presenting research to diverse audiences.

Nice to Haves

  • Research focused on adolescent mental health, suicide prevention, psychotherapy, digital mental health, clinical decision making, or related areas

  • Experience studying how people interact with AI or other digital technologies

  • Familiarity with large language models or AI evaluation

  • Experience with longitudinal studies, conversation analysis, or real-world behavioral datasets

  • Experience leading IRBs, multi-site studies, or collaborations across institutions

  • Existing collaborations within academia or healthcare that you'd like to continue growing

What We Offer

  • Highly competitive salary and meaningful ownership. Excellence is well rewarded.

  • Relocation and transportation support

  • Health/dental insurance coverage

  • Lunch and dinner provided, free snacks/coffee/drinks

  • Unlimited PTO

  • Opportunity to publish and present your work

About Us

Founding team: The core methodology behind this platform comes from NLP evaluation research we had done at Stanford. We raised a $5M seed from some of the top institutional and angel investors in the valley. Our team has prior work experience at NVIDIA, Meta, Microsoft, Palantir and HRT. Collectively, we have over 300 citations in our published work. Our early team include Stanford PhDs, ex-Jane Street quants, and the first designer at Snorkel.

Tech stack: We use Python for most things at Vals. Our platform is built on Django, with a React frontend. All of the infra is on AWS using CDK for IaC.

What We're Looking For

  • Learning velocity: The role encompasses a wide variety of tasks. Rather than expecting you to be an expert on Day 1, we are looking for someone who can learn new skills and technologies extremely quickly.

  • Ownership: Working in a small, talent-dense team, we expect everyone to show initiative to build where it's needed, not where it's asked. We strive for autonomy over consensus. This is especially true for this role.

  • Intensity: The LLM landscape is constantly changing. Foundation model labs are continuously pushing the frontier. The unicorn companies that will emerge from this technology shift are being built now. Those that win will have an incredibly high speed of execution.

  • Solution-oriented mindset: We're looking for people who see opportunities to craft solutions at each juncture, not those who pass hard problems to others or admit defeat.

Further Reading:

  • Hugging Face blog on evaluation

  • Anthropic’s blog on challenges in evaluation

  • New York Times article on issues in benchmarking

  • Stanford HAI report showing hallucinations in legal tech tools

Related keywords

Clinical Research ScientistMental Health AILarge Language ModelsAI EvaluationClinical PsychologyPsychiatryBehavioral SciencePublic HealthStudy DesignStatistical AnalysisPsychometricsPythonDjangoReactAWSInfrastructure as Code

About Vals AI

LinkedInVisit site

Evaluating Large Language Models.

Industry
Software Development
Company size
11-50 employees
Headquarters
San Francisco
LinkedIn followers
6,935

Model benchmarks are seriously lacking. With Vals AI, we report how language models perform on the industry-specific tasks where they will be used.

Offices: San Francisco, US

View all jobs at Vals AI

About Vals AI

LinkedInVisit site

Evaluating Large Language Models.

Industry
Software Development
Company size
11-50 employees
Headquarters
San Francisco
LinkedIn followers
6,935

Model benchmarks are seriously lacking. With Vals AI, we report how language models perform on the industry-specific tasks where they will be used.

Offices: San Francisco, US

View all jobs at Vals AI

Similar companies hiring

Amazon (12278)Bosch (3775)Google (3629)Prolific (3468)AgileEngine (2660)Transport AI (2088)Microsoft (1841)Booz Allen Hamilton (1726)Speechify (1503)BJAK (1216)Salesforce (1130)Meta (1061)
Clera home

Your AI-talent agent. Connecting talents with dream jobs.

Earn $5,000

Tools

  • Salary Calculator
  • Resume Review
  • Startup Map

Explore

  • Jobs
  • Discover Jobs
  • Companies
  • Case Studies
  • Referral

Platform

  • Pricing
  • Integrations
  • Partners
  • Acquihire

Clera

  • Manifesto
  • Engineering
  • We are hiring!
  • FAQs
  • Blog
  • Press

Tools

  • Salary Calculator
  • Resume Review
  • Startup Map

Explore

  • Jobs
  • Discover Jobs
  • Companies
  • Case Studies
  • Referral

Platform

  • Pricing
  • Integrations
  • Partners
  • Acquihire

Clera

  • Manifesto
  • Engineering
  • We are hiring!
  • FAQs
  • Blog
  • Press

© 2026 Clera Labs, Inc.

PrivacyTermsBug Bounty