Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale …
Applied Intuition, Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion, the Silicon Valley company is creating the digital infrastructure needed to bring intelligence to every moving…
Senior Software Development Engineer in Test (Senior SDET / SETI)
Sunnyvale, California, United States · On-site
$140k–$210k/yr
Senior$3.6B raised
Calling all innovators - find your future at Fiserv. We're Fiserv, a global leader in Fintech and payments, and we move money and information in a way that moves the world. We connect financial institutions, corporations…
Your Impact Join a small team building the next generation of Cyber Security products from the ground up. Led by industry veterans with a proven track record of success, you will play a critical role in bridging engineer…
Your Impact Join a small team building the next generation of Cyber Security products from the ground up. Led by industry veterans with a proven track record of success - you will get to architect, build, and deliver the…
Skills: Cyber Security, Test automation, Python, Go, Rust
Amazon Music is an immersive audio entertainment service that deepens connections between fans, artists, and creators. From personalized music playlists to exclusive podcasts, concert livestreams to artist merch, Amazon …
Skills: Machine learning, Deep learning, LLMs, Agentic AI, Java
Your Impact Join a small team building the next generation of Cyber Security products from the ground up. Led by industry veterans with a proven track record of success - you will get to architect, build, and deliver hug…
Skills: Cyber security, Malware analysis, Python, C, C++
The Role Wayve is developing driver-out Robotaxi systems that bring together the AI Driver, vehicle platform, sensors and compute, safety mechanisms, Remote Assistance, fleet services, and cloud operations. We are lookin…
Job Description About the Team The AV Launch team builds the software foundation that brings GM’s autonomous driving capabilities to life on the vehicle. We configure, deploy, launch, and monitor the applications that ma…
Senior Software Systems Engineer, Autonomous Systems Validation Confidence
Sunnyvale, California, United States · Remote OK
$153k–$234k/yr
Senior$8.5B raised
Job Description About the role We are looking for a Senior Software Systems Engineer to develop the methods, software, and quantitative evidence used to measure confidence in autonomous vehicle validation results. This r…
Skills: Python, C++, Systems engineering, Data science, Statistics
ML Systems Engineer, Data Labeling Engineering - Early Career
Sunnyvale, California, United States · Hybrid
$125k–$165k/yr
Entry level$8.5B raised
Job Description About the Team Help teach our self-driving vehicles how to see and understand the world. The Data Labeling Engineering team designs, builds, and operates hybrid human/machine data labeling tools and pipel…
Job Description Work Arrangement: This role is categorized as Remote/hybrid. Remote: This role is based remotely but if you live within a 50-mile radius of [Austin, Detroit, Warren, Milford, Sunnyvale, CA ], you are expe…
Skills: Autonomous driving, Robotics, Python, Data analysis, Machine learning
At Sonatus, we’re driving the transformation to AI-enabled software-defined vehicles. Traditional automotive software methods can’t keep pace with consumer expectations shaped by the mobile industry—where features evolve…
Skills: AI Testing, Test Automation, Embedded Software, Python, Linux
At GFiber, we believe that great internet has the power to drive innovation, strengthen communities, enable the impossible, and do all the everyday things that make all of our world go round. And the job of creating bett…
Are you passionate about mentoring, coding, writing documentation, and helping developers solve engineering challenges? Developer Relations is looking for a talented software engineer to help developers create outstandin…
Senior Technical Program Manager II, Gemini Enterprise, Cloud AI
Sunnyvale, California, United States · On-site
$240k–$333k/yr
Senior+$26M raised
Minimum qualifications: Bachelor's degree in a a technical field, or equivalent practical experience. 10 years of experience in program management. Experience with AI model training, testing, evaluation, and tuning proce…
Skills: Program Management, AI Model Training, Large Language Models, KPI Development, Strategic Planning
Apple is where individual imaginations gather together, committing to the values that lead to great work. Every new product we build or service we create is the result of us making each other’s ideas stronger. That happe…
Skills: C++, Modern C++, Computer Vision, Machine Learning, 3D Perception
Technical Release Manager, Retail and Marcom Engineering
Sunnyvale, California, United States · On-site
Mid level$11B raised
Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind…
Position Summary Join Metis Technology Solutions in supporting Lockheed Martin Space’s Combined Orbital Operations Logistics and Resiliency (COOLR) program. This role delivers IT and cybersystems security support f…
Skills: System Administration, Linux, UNIX, Windows, Embedded Hardware
This role is pivotal in shaping our cloud security strategy and safeguarding data privacy. As our Principal Security Engineer - Cloud, you will be a key leader in our cloud security journey, working closely with the Kube…
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Full-time
Posted 2d ago
~40 hrs/week
Responsibilities
You will design and build automated test frameworks and release qualification pipelines for the GPU inference stack. Additionally, you will lead validation efforts for multi-node cluster bring-up, performance modeling, and fault-injection testing to ensure production-grade reliability.
Requirements
Candidates must have 8+ years of experience in software engineering as an SDET or Systems Test Engineer with deep expertise in GPU cluster infrastructure. Proficiency in Python, container orchestration, and distributed LLM serving engines is required.
Full job description
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
About the Role As a Staff GPU Inference SDET, you will be the founding quality, reliability, and validation lead for a new GPU Inference Development team. Working closely with engineering leads and cross-functional systems infrastructure teams, you will design, build, and scale the end-to-end release qualification and automated test ecosystem for our GPU inference stack and rack-scale accelerated compute fleets. In this high-impact role, you will be responsible for building automated test suites to validate multi-node GPU cluster bring-up, verifying prefill worker optimizations, testing open-source and custom serving engines, and ensuring numerical correctness and performance stability under real-world streaming workloads. You will be the primary technical anchor ensuring production-grade reliability, fault isolation, and peak inference performance across accelerated GPU infrastructure.
WHAT YOU’LL DO
Build GPU Release Qualification Systems:
Design and implement automated test automation frameworks, regression gates, and release qualification pipelines for the complete GPU inference stack—spanning custom API services, model-serving workers, container runtimes, serving engines, driver stacks, and firmware.
Inference Serving & Workload Validation:
Benchmark and stress-test distributed LLM serving frameworks, focusing on prefill vs. decode worker performance, continuous batching, prefix caching, KV-cache efficiency, and tensor/expert parallelism.
Performance & Performance Modeling Verification:
Build automated workload replay and benchmarking tools to validate GPU performance models. Track critical serving metrics including Time-to-First-Token (TTFT), Inter-Token Latency (ITL), request throughput, tail latency (P99), and capacity efficiency.
Numerical Correctness & Quality Gates:
Build validation infrastructure to ensure model accuracy, precision stability (FP16/FP8/quantization), determinism, and output correctness across software updates, kernel fusions, and hardware revisions.
Fault Injection & Fleet Resilience:
Engineer chaos engineering and fault-injection suites to simulate node failures, inter-node network degradation, GPU memory leaks, driver/firmware mismatches, and automated recovery paths for multi-node GPU clusters.
Observability & CI/CD Integration:
Integrate automated test pipelines with telemetry tools (e.g., Prometheus, Grafana) to turn one-off investigations into repeatable engineering gates and continuous performance monitoring.
REQUIREMENTS:
8+ years of software engineering experience as an SDET, Infrastructure Quality Lead, or Systems Test Engineer.
GPU & Cluster Infrastructure Expertise:
Hands-on experience bringing up, provisioning, and validating multi-node GPU clusters (NVIDIA or AMD ecosystem) across public cloud infrastructure or enterprise data center environments.
Inference Stack Knowledge:
Deep understanding of LLM serving engines and distributed runtimes, including prefill vs. decode disaggregation, KV-cache management, and dynamic batching.
Automation & Scripting:
Expert-level Python programming skills with extensive experience designing custom test automation frameworks, diagnostic tooling, and CI/CD integration.
Orchestration & Networking:
Strong proficiency with container orchestration tools (e.g., Kubernetes, Slurm, Ray) and high-performance cluster interconnects (e.g., InfiniBand, RoCE, NCCL).
Failure Analysis & Debugging:
Proven background in root-cause analysis across software/hardware boundaries, stress testing, and node failure simulation in distributed systems.
NICE TO HAVES:
Direct experience with either AMD (ROCm / HIP) or NVIDIA software stacks.
Experience building workload replay tools, ML evaluation pipelines, or MLPerf Inference benchmark suites.
Familiarity with low-level kernel profiling tools (PyTorch Profiler, NVTX, ROCm profilers) or C++
Why Join Cerebras
People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:
Build a breakthrough AI platform beyond the constraints of the GPU.
Publish and open source their cutting-edge AI research.
Work on one of the fastest AI supercomputers in the world.
Enjoy job stability with startup vitality.
Our simple, non-corporate work culture that respects individual beliefs.
Find out more about what it's like to work at Cerebras here!
Apply today and become part of the forefront of groundbreaking advancements in AI!
Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.
This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.
Related keywords
GPUInferenceSDETLLMPythonKubernetesSlurmRayInfiniBandRoCENCCLCI/CDPrometheusGrafanaFault InjectionDistributed Systems
Cerebras Systems builds the world's fastest AI inference. We are powering the future of generative AI. We’re a team of pioneering computer architects, deep learning researchers, and engineers building a new class of AI supercomputers from the ground up.
From sub-second inference speeds to breakthrough training performance, Cerebras makes it easier to build and deploy state-of-the-art AI—from proprietary enterprise models to open-source projects downloaded millions of times.
Here’s what makes our platform different:
🔦 Sub-second reasoning – Instant intelligence and real-time responsiveness, even at massive scale
⚡ Blazing-fast inference – Up to 30x faster than GPUs
🧠 Agentic AI in action – Models that can plan, act, and adapt autonomously
🌍 Scalable infrastructure – Built to move from prototype to global deployment without friction
Cerebras solutions are available in the Cerebras Cloud or on-prem, serving leading enterprises, research labs, and government agencies worldwide.
👉 Learn more: https://www.cerebras.ai
Join us: https://cerebras.net/careers/
Offices: 1237 E Arques Ave, Sunnyvale, California 94085, US · 150 King St W, Toronto, Ontario M5H 1J9, CA · Tokyo, JP · Bangalore, IN
artificial intelligencedeep learningnatural language processinginferencemachine learningllmAIenterprise AIand fast inferenceSemiconductor
Based on 1502 listings with disclosed salaries, most software jobs in Sunnyvale, CA pay between $140k–$286k per year. Individual offers vary with seniority, company size, and specialization.
How many Software jobs are open in Sunnyvale, CA right now?
There are currently 1,772 open software positions in Sunnyvale, CA listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Software roles in Sunnyvale, CA?
Companies currently hiring include Google, Walmart, Apple, Wayve, Applied Intuition, among others. Browse the listings above to see every active employer.
Are there remote or hybrid Software jobs in Sunnyvale, CA?
Yes — 571 of the 1772 open software positions offer remote or hybrid work (40 remote, 531 hybrid).
How do I apply for Software jobs in Sunnyvale, CA?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.