About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like fin…
Skills: Kubernetes, Python, Go, Distributed systems, ML infrastructure
Software Development Engineer , Adaptive Search Relevance
Palo Alto, California, United States · On-site
$144k–$224k/yr
Mid level$67M raised
Build the multi-agent orchestration platform that autonomously executes ML research workflows — from hypothesis generation through production deployment — accelerating how we deliver relevant sponsored products to hundre…
Skills: Software development, System architecture, Machine learning, Generative AI, LLM
Principal Technical Program Manager, Sponsored Products and Brands Agent
Palo Alto, California, United States · On-site
$177k–$275k/yr
Senior+$67M raised
At Amazon Ads, we’re re-imagining the advertising landscape through advanced generative AI technologies and AI agents, revolutionizing how millions of customers discover products and engage with brands online. We are at …
Skills: Technical program management, Generative AI, AI agents, Software development, System architecture
A World-Changing Company Palantir builds the world’s leading software for data-driven decisions and operations. By bringing the right data to the people who need it, our platforms empower our partners to develop lifesavi…
Skills: Legal operations, Project management, Process optimization, Data analysis, Technical tooling
Why Join GEICO? At GEICO, we offer a rewarding career where your ambitions are met with endless possibilities. Every day we honor our iconic brand by offering quality coverage to millions of customers and being there whe…
About us At Vinci, we are building the operator intelligence infrastructure that modern hardware programs rely on daily. We have already proven that a single foundation model works out of the box across physics on realis…
Skills: Backend Engineering, Infrastructure as Code, Terraform, CI/CD, Jenkins
Our Mission As humans, there are few things more exciting than meeting someone new. At Tinder, we’re inspired by the challenge of keeping the magic of human connection alive. With tens of millions of users, hundreds of m…
Solutions Architect, Customer Success - US (Remote)
Palo Alto, California, United States · Hybrid
$160k–$230k/yr
Senior$94M raised
Our Purpose At Fiddler, we understand the implications of AI and the impact that it has on human lives. Our company was born with the mission of building trust into AI. The rise of Generative AI and Agents has unlocked g…
Skills: Solutions architecture, Customer success, Machine learning, AI observability, Data science
Member of Technical Staff - Full-Stack Software Engineer
Palo Alto, California, United States · Hybrid
$180k–$220k/yr
Senior$46M raised
About Us Vinci combines a foundation model for physics with GPU-native solvers to deliver unprecedented simulation speed and accuracy. There’s no meshing, no approximations, and customer data is not required to train the…
About Us 🚀 We're building an AI assistant for hardware designers. Our mission is to enable millions of hardware designers and engineers to iterate through designs 1000x faster. We are building our geometry + physics dri…
ABOUT QUINCE Quince is a destination for builders, creators, innovators, and operators who want to come together and challenge the status quo. Our mission is simple: make really high quality essentials for really low pri…
Skills: Python, SQL, Machine Learning, Recommender Systems, Generative AI
ABOUT QUINCE Quince is a destination for builders, creators, innovators, and operators who want to come together and challenge the status quo. Our mission is simple: make really high quality essentials for really low pri…
Skills: Full-stack Development, Backend Services, RESTful APIs, Java, Spring Boot
Research Design and Data Analysis Student Consultants (RDDA Student Consultant) Job Description Palo Alto University (PAU), a private, non-profit university, founded in 1975 and located in the heart of Northern Californi…
Skills: Research Design, Data Analysis, Statistical Methods, Methodology, Data Collection
ABOUT QUINCE Quince is a destination for builders, creators, innovators, and operators who want to come together and challenge the status quo. Our mission is simple: make really high quality essentials for really low pri…
Skills: Demand Forecasting, Inventory Placement Optimization, Operations Research, Machine Learning, Large Language Models
ABOUT QUINCE Quince is a destination for builders, creators, innovators, and operators who want to come together and challenge the status quo. Our mission is simple: make really high quality essentials for really low pri…
Who We Are The world is moving towards instant digital payments and TabaPay is leading the way. We help thousands of Fintechs in the US and Canada instantly move money in and out of accounts and we are actively expanding…
About Clockwork Systems Clockwork.io – Software Driven Fabrics to increase GPU cluster utilization Clockwork Systems was founded by Stanford researchers and veteran systems engineers who share a vision for redefining the…
Senior Manager - Lifecycle Marketing, North America
Palo Alto, California, United States · Remote OK
Senior
About tonies: tonies is the globally leading interactive audio platform for children with more than 10 million Tonieboxes and 125 million Tonies sold. The intuitive and award-winning audio system has changed the way youn…
Senior Software Engineering Manager - Multi-Object Tracking & State Estimation
Palo Alto, California, United States · On-site
$253k–$380k/yr
Senior+Visa sponsorship
Latitude AI (lat.ai) is building the future of Ford’s autonomy roadmap to make travel safer, less stressful, and more enjoyable for everyone. Bringing this vision to scale, our fully in-house developed hands-free ADAS pl…
Skills: Multi-object tracking, State estimation, Machine learning, Computer vision, Python
Triage Associate I, Platform Triage - Second Shift (Contract)
Palo Alto, California, United States · On-site
$26/hr–$39/hr
Mid level
Latitude AI (lat.ai) is building the future of Ford’s autonomy roadmap to make travel safer, less stressful, and more enjoyable for everyone. Bringing this vision to scale, our fully in-house developed hands-free ADAS pl…
Skills: Data Analysis, Troubleshooting, ADAS Testing, Sensor Data Analysis, Technical Documentation
You will build and operate the ML platform that powers large-scale training, evaluation, and batch inference at Mistral AI. This involves developing infrastructure for distributed GPU workloads, managing compute capacity, and ensuring system reliability through observability and production operations.
Requirements
Candidates must have 4+ years of experience in ML infrastructure, distributed systems, or Kubernetes platform engineering. Proficiency in Python or Go and deep knowledge of GPU infrastructure and scheduling technologies are required.
Full job description
About Mistral
Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms.
We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.
The Role
This role focuses on building and operating the ML platform that powers large-scale training, evaluation, and batch inference at Mistral AI. You will develop the infrastructure that enables researchers and engineers to run distributed GPU workloads reliably across clusters, hardware types, and regions.
You will work across the full ML lifecycle, from workload scheduling and capacity management to platform APIs, observability, and production operations. You will take ownership of critical systems and help turn complex infrastructure into reliable, self-service capabilities.
What You Will Do
Build the ML Platform: Develop services, APIs, controllers, and tooling for training, evaluation, fine-tuning, and batch inference.
Orchestrate GPU Workloads: Build systems for queueing, admission control, quotas, priorities, preemption, and topology-aware placement.
Manage Compute Capacity: Improve how heterogeneous GPU resources are provisioned, allocated, and utilized across clusters.
Enable Multi-Cluster Execution: Place workloads based on capacity, data locality, hardware requirements, and organizational priorities.
Improve Researcher Experience: Create self-service workflows that make distributed workloads easy to launch, observe, debug, and reproduce.
Optimize Performance: Improve GPU utilization, scheduling latency, workload startup time, throughput, and infrastructure efficiency.
Build for Reliability: Develop observability, failure recovery, capacity planning, and operational tooling for critical ML workloads.
Operate What You Build: Participate in on-call rotations and troubleshoot issues across applications, schedulers, networking, storage, and GPU infrastructure.
What We're Looking For
Have 4+ years of experience in ML infrastructure, distributed systems, Kubernetes platform engineering, or a related field.
Are proficient in Python or Go and comfortable working with production-grade distributed systems.
Have strong Kubernetes knowledge, including controllers, operators, CRDs, scheduling, networking, storage, and resource management.
Understand technologies such as Kueue, Karpenter, Volcano, and Kyverno, and the problems they address in workload scheduling, provisioning, and policy enforcement.
Understand distributed ML workloads, including training, fine-tuning, evaluation, checkpointing, and batch inference.
Are familiar with GPU infrastructure and technologies such as PyTorch, CUDA, NCCL, and high-performance networking.
Understand concepts such as quotas, priorities, preemption, gang scheduling, topology awareness, and workload admission.
Can diagnose performance and reliability problems across software, orchestration, networking, storage, and hardware.
Care about developer experience and enjoy turning complex infrastructure into simple, reliable interfaces.
Thrive in an ambiguous, fast-moving environment shaped by frontier AI research.
What We Offer
We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.
For the most up-to-date details on benefits available in your location, please refer to our Benefits page.
Privacy Policy
Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy.
Related keywords
ML PlatformKubernetesPythonGoDistributed SystemsGPUPyTorchCUDANCCLKueueKarpenterVolcanoKyvernoInfrastructureAPIObservability
Frontier AI. In your hands.
We believe in a future where AI is abundant and accessible. We aspire to empower the world to build with—and benefit from—the most significant technology of our time.
Join us: mistral.ai/careers
How much do Data & Analytics jobs in Palo Alto, CA pay?
Based on 498 listings with disclosed salaries, most data & analytics jobs in Palo Alto, CA pay between $114k–$270k per year. Individual offers vary with seniority, company size, and specialization.
How many Data & Analytics jobs are open in Palo Alto, CA right now?
There are currently 608 open data & analytics positions in Palo Alto, CA listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Data & Analytics roles in Palo Alto, CA?
Companies currently hiring include Amazon, Rivian, JPMorganChase, SpaceXAIA, SpaceX, among others. Browse the listings above to see every active employer.
Are there remote or hybrid Data & Analytics jobs in Palo Alto, CA?
Yes — 278 of the 608 open data & analytics positions offer remote or hybrid work (59 remote, 219 hybrid).
How do I apply for Data & Analytics jobs in Palo Alto, CA?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.