About the Role As a Member of Technical Staff, Training, you will design, build, and operate the distributed systems behind large-scale model post-training — spanning training, inference, and orchestration, with a focus …
Skills: ML systems, Large-scale training infrastructure, LLMs, Generative models, Distributed training
Company Description IFS is a billion-dollar revenue company with 7000+ employees on all continents. Our leading AI technology is the backbone of our award-winning enterprise software solutions, enabling our customers to …
Skills: Python, TypeScript, JavaScript, LLMs, Generative AI
Company Description IFS is a billion-dollar revenue company with 7000+ employees on all continents. Our leading AI technology is the backbone of our award-winning enterprise software solutions, enabling our customers to …
PsiQuantum's mission is to build the first useful quantum computers: machines capable of delivering the breakthroughs the field has long promised. Since our founding in 2016, our singular focus has been to build and depl…
PsiQuantum's mission is to build the first useful quantum computers: machines capable of delivering the breakthroughs the field has long promised. Since our founding in 2016, our singular focus has been to build and depl…
Skills: Computational chemistry, GPU acceleration, CUDA, HIP/ROCm, C++
SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this poss…
SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this poss…
DESCRIPTION: Duties: Design and develop automation-based systems for data analysis, document management, and client intelligence. Research machine learning methods for data processing and analysis. Develop machine learni…
Skills: Machine learning, Natural language processing, Speech recognition, Python, PyTorch
Pivotal is the leader in the emerging market of electric Vertical Takeoff and Landing (eVTOL) aircraft. We design, develop, and manufacture light eVTOL aircraft and are renowned for the BlackFly aircraft the world's firs…
Skills: Embedded systems, C++, C, Python, Firmware development
Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active develo…
Skills: Developer Relations, LLMs, Technical Writing, Public Speaking, Software Engineering
About Us Rivian and Volkswagen Group Technologies is a joint venture between two industry leaders with a clear vision for automotive’s next chapter. From operating systems to zonal controllers to cloud and connectivity s…
About Us Rivian and Volkswagen Group Technologies is a joint venture between two industry leaders with a clear vision for automotive’s next chapter. From operating systems to zonal controllers to cloud and connectivity s…
Skills: C/C++, Embedded software development, Linux, QNX, Android Automotive OS
About Us Rivian and Volkswagen Group Technologies is a joint venture between two industry leaders with a clear vision for automotive’s next chapter. From operating systems to zonal controllers to cloud and connectivity s…
Skills: Technical program management, Systems engineering, Hardware development lifecycle, Software development lifecycle, Test and validation
Medical Scribe Intern, Clinical AI Safety & Evaluation
Palo Alto, California, United States · Hybrid
Entry level$155M raised
Transform healthcare with us. At Qualified Health, we're redefining what's possible with Generative AI in healthcare. Our infrastructure provides the guardrails for safe AI governance, healthcare-specific agent creation,…
Skills: Medical terminology, Clinical documentation, Data labeling, Quality assurance, Attention to detail
About Nu Nu is the leading digital bank in Latin America, serving 135 million customers across Brazil, Mexico, and Colombia. The company has been leading an industry transformation by leveraging data and proprietary tech…
Skills: Product Design, Mobile Design, User Experience, Fintech, Blockchain
G'day! We are ServiceRocket🚀, a global tech-enabled services company headquartered in Palo Alto, California. Our purpose is to be the single most reliable partner in the acceleration of your growth. At ServiceRocket, we…
If you love variety, enjoy making digital tools work harmoniously behind the scenes, and have a meticulous eye for detail, we want to hear from you! Pay Range: $72,500-$76,000 /annual This is an on-site position in Palo …
About Rivian Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. As a company, we con…
Senior Product Manager, Consumer AI Assistant — Mobile
Palo Alto, California, United States · Hybrid
$115k–$163k/yr
Senior$16B raised
Senior Product Manager, Consumer AI Assistant — Mobile About the Role We are looking for a Senior Product Manager to lead the strategy and execution of our consumer AI assistant on mobile. You will own two critical outco…
Senior Product Manager, Consumer AI Assistant — Mobile
Palo Alto, California, United States · Hybrid
$100k–$163k/yr
Senior$16B raised
Senior Product Manager, Consumer AI Assistant — Mobile About the Role We are looking for a Senior Product Manager to lead the strategy and execution of our consumer AI assistant on mobile. You will own two critical outco…
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Full-time
Competitive compensation, Meaningful equity, Comprehensive benefits, Flexible work arrangements
Posted 15d ago
~40 hrs/week
Responsibilities
Contribute to open-source large-scale post-training infrastructure and optimize throughput, scalability, and hardware efficiency. Develop training frameworks, collaborate with model researchers, and build observability systems for training performance.
Requirements
Requires 3+ years of experience in ML systems or large-scale training infrastructure. Candidates must have experience with training/inference correctness, performance debugging, and building or operating agentic post-training systems.
Full job description
About the Role
As a Member of Technical Staff, Training, you will design, build, and operate the distributed systems behind large-scale model post-training — spanning training, inference, and orchestration, with a focus on the performance, correctness, scalability, and reliability of workloads running across large GPU clusters.
This role suits engineers who move fluidly across modeling recipes, complex infrastructure, and low-level systems, identify bottlenecks in distributed workloads, and translate experimental requirements into robust software.
In This Role, You Will
Design, build, and operate distributed training, rollout, and orchestration systems for large-scale LLM and multimodal post-training across multi-GPU, multi-node environments.
Profile and optimize performance across the full-stack — model implementation, parallelism strategies, communication libraries, and GPU kernels — to improve throughput, latency, memory efficiency, hardware utilization, and cost.
Investigate numerical correctness and low-precision issues in distributed training and inference, including train–inference consistency for reinforcement learning.
Improve the reliability of long-running workloads through checkpointing, fault recovery, observability, and operational tooling.
Build supporting infrastructure for reinforcement learning and agentic post-training, including asynchronous rollout, trajectory collection, sandboxed execution, evaluation harnesses, and data pipelines.
Contribute to open-source training and inference systems, including Miles and SGLang, and partner with researchers to turn experimental requirements into production systems.
Minimum Qualifications
3+ years of experience building or operating distributed machine learning systems, large-scale training infrastructure, or high-performance inference systems.
Hands-on experience with post-training systems, training backends, or inference systems for large language models (e.g., Megatron-LM, FSDP, SGLang, TensorRT-LLM, vLLM).
Experience in at least two of the following areas:
Performance, efficiency, and scalability of multi-GPU, multi-node workloads
Numerical correctness or low precision
Stability, reliability, or fault tolerance
Post-training algorithm recipes and orchestration infrastructure for large training runs
Multimodal training or inference, including vision-language models and multimodal generation
Agent infrastructure, including sandboxes, harnesses, and eval systems
Building and maintaining open-source projects widely adopted in industry and academia
Preferred Qualifications
Familiarity with RL algorithms such as PPO, GRPO, and their variants, and experience applying them in large-scale post-training.
Experience with modern post-training frameworks (e.g., Miles, slime, AReaL, verl, Prime-RL).
Key open-source contributions to training or inference frameworks (e.g., SGLang, vLLM, Megatron-LM).
GPU kernel development (e.g., CUDA, Triton, CUTLASS) or communication-layer optimization (e.g., NCCL, RDMA, NVLink/NVSwitch).
Experience training or serving models at very large scale (e.g., Mixture-of-Experts models on clusters of thousands of GPUs).
Top-tier publications in ML systems or other systems fields.
Even if you don't meet every qualification above, we encourage you to apply — we care most about demonstrated ability to build and reason about large-scale systems.
About RadixArk
RadixArk builds open-source and production infrastructure for large language models and multimodal post-training. Our systems — including Miles, an enterprise-grade reinforcement learning training framework, and SGLang, a widely deployed high-performance LLM inference engine — power distributed post-training across clusters of 10k–100k+ GPUs.
Compensation
Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 to $400,000, plus equity.
Benefits include a 401(k) plan and unlimited PTO.
RadixArk sponsors employment visas (e.g., H-1B, O-1) for eligible candidates.
Equal Opportunity
RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Related keywords
ML systemsLarge-scale trainingLLMGenerative modelsDistributed trainingPerformance engineeringMegatron-LMFSDPTorchtitanSGLangvLLMRDMAInfiniBandNVLinkNCCLRCCL
Based on 729 listings with disclosed salaries, most software jobs in Palo Alto, CA pay between $120k–$275k per year. Individual offers vary with seniority, company size, and specialization.
How many Software jobs are open in Palo Alto, CA right now?
There are currently 952 open software positions in Palo Alto, CA listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Software roles in Palo Alto, CA?
Companies currently hiring include Rivian, Amazon, Rivian and Volkswagen Group Technologies, Ford Motor Company, JPMorganChase, among others. Browse the listings above to see every active employer.
Are there remote or hybrid Software jobs in Palo Alto, CA?
Yes — 470 of the 952 open software positions offer remote or hybrid work (102 remote, 368 hybrid).
How do I apply for Software jobs in Palo Alto, CA?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.