About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like fin…
Skills: Kubernetes, Python, Go, Distributed systems, ML infrastructure
About us At Vinci, we are building the operator intelligence infrastructure that modern hardware programs rely on daily. We have already proven that a single foundation model works out of the box across physics on realis…
Skills: Backend Engineering, Infrastructure as Code, Terraform, CI/CD, Jenkins
If you're ready to be part of our legacy of hope and innovation, we encourage you to take the first step and explore our current job openings. Your best is waiting to be discovered. Day - 08 Hour (United States of Americ…
Valar Labs is a well funded, fast-growing AI diagnostics startup backed by leading VCs revolutionizing oncology and urology through cutting-edge technology and innovation. We are dedicated to improving patient outcomes t…
About Arc Institute Arc Institute is an independent nonprofit research organization at the interface of artificial intelligence and biology, working to accelerate scientific progress and understand the root causes of com…
Skills: Human immunology, Neurobiology, Computational biology, Functional genomics, Protein display technologies
Research Design and Data Analysis Student Consultants (RDDA Student Consultant) Job Description Palo Alto University (PAU), a private, non-profit university, founded in 1975 and located in the heart of Northern Californi…
Skills: Research Design, Data Analysis, Statistical Methods, Methodology, Data Collection
ABOUT QUINCE Quince is a destination for builders, creators, innovators, and operators who want to come together and challenge the status quo. Our mission is simple: make really high quality essentials for really low pri…
Skills: Demand Forecasting, Inventory Placement Optimization, Operations Research, Machine Learning, Large Language Models
ABOUT QUINCE Quince is a destination for builders, creators, innovators, and operators who want to come together and challenge the status quo. Our mission is simple: make really high quality essentials for really low pri…
About Clockwork Systems Clockwork.io – Software Driven Fabrics to increase GPU cluster utilization Clockwork Systems was founded by Stanford researchers and veteran systems engineers who share a vision for redefining the…
Latitude AI (lat.ai) is building the future of Ford’s autonomy roadmap to make travel safer, less stressful, and more enjoyable for everyone. Bringing this vision to scale, our fully in-house developed hands-free ADAS pl…
Skills: Modern C++, C++17, Parsing Software, Python, Entity Component Systems
Ascendis Pharma is a dynamic, fast-growing global biopharmaceutical company with locations in Denmark, Europe, and the United States. Today, we're advancing programs in Endocrinology Rare Disease and Oncology. Here at As…
We are so glad you are interested in joining Sutter Health! Organization: PAMF-Palo Alto Medical Foundation PAD Position Overview: Competently performs routine and specialized nuclear medicine procedures including perfor…
Skills: Nuclear medicine procedures, PET CT, Patient assessment, Time management, Clinical skills
Job Summary We're continuously building a pipeline of talented maintenance, instrumentation, controls, and equipment professionals interested in future career opportunities at our Palo Alto biotechnology site. At IFF Hea…
Skills: Process instrumentation, Calibration, Troubleshooting, Control systems, PLC
About Arc Institute Arc Institute is an independent nonprofit research organization at the interface of artificial intelligence and biology, working to accelerate scientific progress and understand the root causes of com…
Skills: Legal strategy, Corporate law, Transactional law, Technology transfer, Life sciences law
SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this poss…
Skills: Electrical Design, PCB Layout, Mixed-Signal Circuits, RF Front Ends, Signal Integrity
SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this poss…
SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this poss…
Skills: RFIC Design, Analog Circuit Design, Mixed-Signal Design, SiGe, CMOS
SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this poss…
Skills: Hardware Design, PCB Layout, Mixed-Signal Circuits, RF Front Ends, Signal Integrity
SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this poss…
SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this poss…
Skills: C, C++, Wireless Communications, Signal Processing, Network Protocols
You will build and operate the ML platform that powers large-scale training, evaluation, and batch inference at Mistral AI. This involves developing infrastructure for distributed GPU workloads, managing compute capacity, and ensuring system reliability through observability and production operations.
Requirements
Candidates must have 4+ years of experience in ML infrastructure, distributed systems, or Kubernetes platform engineering. Proficiency in Python or Go and deep knowledge of GPU infrastructure and scheduling technologies are required.
Full job description
About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.
We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.
The Role
This role focuses on building and operating the ML platform that powers large-scale training, evaluation, and batch inference at Mistral AI. You will develop the infrastructure that enables researchers and engineers to run distributed GPU workloads reliably across clusters, hardware types, and regions.
You will work across the full ML lifecycle, from workload scheduling and capacity management to platform APIs, observability, and production operations. You will take ownership of critical systems and help turn complex infrastructure into reliable, self-service capabilities.
What You Will Do
Build the ML Platform: Develop services, APIs, controllers, and tooling for training, evaluation, fine-tuning, and batch inference.
Orchestrate GPU Workloads: Build systems for queueing, admission control, quotas, priorities, preemption, and topology-aware placement.
Manage Compute Capacity: Improve how heterogeneous GPU resources are provisioned, allocated, and utilized across clusters.
Enable Multi-Cluster Execution: Place workloads based on capacity, data locality, hardware requirements, and organizational priorities.
Improve Researcher Experience: Create self-service workflows that make distributed workloads easy to launch, observe, debug, and reproduce.
Optimize Performance: Improve GPU utilization, scheduling latency, workload startup time, throughput, and infrastructure efficiency.
Build for Reliability: Develop observability, failure recovery, capacity planning, and operational tooling for critical ML workloads.
Operate What You Build: Participate in on-call rotations and troubleshoot issues across applications, schedulers, networking, storage, and GPU infrastructure.
What We're Looking For
Have 4+ years of experience in ML infrastructure, distributed systems, Kubernetes platform engineering, or a related field.
Are proficient in Python or Go and comfortable working with production-grade distributed systems.
Have strong Kubernetes knowledge, including controllers, operators, CRDs, scheduling, networking, storage, and resource management.
Understand technologies such as Kueue, Karpenter, Volcano, and Kyverno, and the problems they address in workload scheduling, provisioning, and policy enforcement.
Understand distributed ML workloads, including training, fine-tuning, evaluation, checkpointing, and batch inference.
Are familiar with GPU infrastructure and technologies such as PyTorch, CUDA, NCCL, and high-performance networking.
Understand concepts such as quotas, priorities, preemption, gang scheduling, topology awareness, and workload admission.
Can diagnose performance and reliability problems across software, orchestration, networking, storage, and hardware.
Care about developer experience and enjoy turning complex infrastructure into simple, reliable interfaces.
Thrive in an ambiguous, fast-moving environment shaped by frontier AI research.
What We Offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.
For the most up-to-date details on benefits available in your location, please refer to our Benefits page.
Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.
Related keywords
ML PlatformKubernetesPythonGoDistributed SystemsGPUPyTorchCUDANCCLKueueKarpenterVolcanoKyvernoInfrastructureAPIObservability
Frontier AI. In your hands.
We believe in a future where AI is abundant and accessible. We aspire to empower the world to build with—and benefit from—the most significant technology of our time.
Join us: mistral.ai/careers
How many Science & Research jobs are open in Palo Alto, CA right now?
There are currently 496 open science & research positions in Palo Alto, CA listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Science & Research roles in Palo Alto, CA?
Companies currently hiring include Stanford Health Care, CONCEPT Continuing & Professional Studies Division, Palo Alto University, Stanford Medicine Children's Health, Amazon, Summit Therapeutics, Inc., among others. Browse the listings above to see every active employer.
Are there remote or hybrid Science & Research jobs in Palo Alto, CA?
Yes — 137 of the 496 open science & research positions offer remote or hybrid work (21 remote, 116 hybrid).
How do I apply for Science & Research jobs in Palo Alto, CA?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.