Staff Site Reliability Engineer - AI Platform Runtime
Santa Clara, California, United States · Hybrid
$168k–$334k/yr
Senior+$29B raised
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems e…
Skills: Site Reliability Engineering, Kubernetes, Python, Typescript, JavaScript
Principal Software Engineer - Networking - DGX Cloud
Santa Clara, California, United States · On-site
$272k–$431k/yr
Senior+$29B raised
At NVIDIA, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing…
Skills: Software-defined networking, EVPN, BGP, IP subnetting, Overlay networking
We are now looking for a Senior Firmware Verification and Bringup Engineer to join our Memory Subsystem Team! Widely considered to be one of the technology world’s most desirable employers, NVIDIA is an industry leader w…
The NVIDIA Supply Chain Business Transformation organization is seeking an experienced data and Cloud technology professional for the position of a Data Systems Analyst who performs complex, advanced work, designing and …
Skills: Data warehousing, Supply chain planning, Power BI, Tableau, Databricks
Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing wha…
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Driving continued aggr…
Role Overview Velaura AI is seeking an experienced Director of IT Infrastructure & Operations to own IT and infrastructure end to end: the hybrid compute environment our designers depend on, the security program that pro…
Skills: IT infrastructure, Operations management, Linux, HPC environments, Cloud architecture
2027 Software Engineering Intern (Masters - Santa Clara, CA)
Santa Clara, California, United States · On-site
Entry level$2.2B raised
Who We Are Applied Materials is the global leader in materials science and engineering solutions that are at the foundation of virtually every new semiconductor chip and advanced display in the world. The equipment that …
Skills: C, C++, Software design, Real-time control, Motion control
Principal Applied Research Engineer, Content Authenticity
Santa Clara, California, United States · On-site
$272k–$431k/yr
Senior+$29B raised
NVIDIA has been transforming computer graphics and accelerated computing for more than 25 years. In the AI era, it’s a unique legacy of innovation that’s fueled by great technology and amazing people. NVIDIA AI for Media…
Skills: Deep Learning, Computer Vision, PyTorch, TensorFlow, ONNX
Senior Business Systems Analyst - SAP IBP Planning
Santa Clara, California, United States · On-site
$144k–$270k/yr
Senior+$29B raised
The Supply Chain Information Technology team is looking to fill a Senior Business Systems Analyst — SAP IBP Planning position. This job is accountable for crafting sophisticated planning solutions, optimization models, a…
Skills: SAP IBP, SAP APO, Supply Chain Planning, S&OP, Demand Planning
Today, NVIDIA is tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing…
Skills: Enterprise Integration, Microservices, SAP Integration Suite, API Management, Agentic AI
Senior Software Engineer, Cosmos Infrastructure and End to End Performance
Santa Clara, California, United States · On-site
$184k–$288k/yr
Senior$29B raised
We're now looking for a Senior Software Engineer, Cosmos Infrastructure and End to End Performance! NVIDIA Cosmos is an open omni-model platform of generative world foundation models (WFMs) designed to accelerate physica…
Software Engineer - Hardware Diagnostics - New Grad
Santa Clara, California, United States · Hybrid
Entry level
Company Description: As a cutting-edge AI infrastructure startup, we are building high-performance systems that power next-generation, large-scale AI deployments. Our team brings together industry veterans and passionate…
Skills: Python, C, Systems software, Embedded systems, Distributed systems
Senior Software Engineer, Agentic AI and Observability
Santa Clara, California, United States · On-site
$168k–$270k/yr
Senior$29B raised
Ready to develop the future of AI at NVIDIA? Join our BizApps SRE team to advance the Agentic AI Factory model, a critical initiative aimed at fast development, deployment, and operation of AI-powered applications across…
At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who …
Skills: Python, Test automation, Software verification, Image processing, Pattern recognition
Senior System Software Engineer, ML and Vector Search
Santa Clara, California, United States · On-site
$184k–$288k/yr
Senior+$29B raised
It’s an exciting time for NVIDIA as we expand our capabilities into the world of data science, data processing, and database acceleration. We're looking for an outstanding software engineer to apply their skills in the d…
Are you seeking an outstanding opportunity? We are looking for a Senior Photonic Layout Design Engineer – someone who is excited to join a growing group of diverse individuals responsible for handling high-speed mixed-si…
NVIDIA is looking for a talented Machine Learning Engineer to drive the development, evaluation, deployment and end-to-end lifecycle management of our AI-powered systems. This role bridges advanced AI application develop…
NVIDIA is looking for a Senior Software Engineer in Object Storage to design, implement, and extend the capabilities of our internal object storage system. This system is a core service that is critical to NVIDIA AI/ML r…
Skills: Object storage, Distributed systems, Python, Go, C++
Santa Clara, California, United States · Remote OK
$221k–$387k/yr
Senior+$4.1B raised
Company Description It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for …
Skills: Product design, Leadership, IT Service Management, AI-native UX, Design strategy
Staff Site Reliability Engineer - AI Platform Runtime
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
$168k–$334k/yr
Full-time
bachelor degree
Equity, Health Insurance
Posted 4d ago
~40 hrs/week
Responsibilities
Lead the technical strategy and roadmap for large-scale SRE initiatives to improve reliability, scalability, and developer productivity. Design and build resilient distributed systems while driving automation and observability improvements across enterprise AI platforms.
Requirements
Requires 10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architecture. Candidates must possess a bachelor degree in Computer Science or a related field and strong proficiency in programming languages like Python, Go, or JavaScript.
Full job description
Site Reliability Engineering (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination of software and systems engineering practices. This is a highly specialized discipline which demands knowledge across different systems, networking, coding, database, capacity management, continuous delivery and deployment, open source cloud enabling technologies like Kubernetes and Public Cloud. SRE at NVIDIA ensures that our internal and external facing services run maximum reliability and uptime as promised to the users and at the same time enabling developers to make changes to the existing system through careful preparation and planning while keeping an eye on capacity, latency and performance. SRE is also a mindset and a set of engineering approaches to running better production systems and optimizations. Much of our software development focuses on building components to eliminate manual work through automation, performance tuning and growing efficiency of production systems.
As SREs are responsible for the big picture of how our systems relate to each other, we use a breadth of tools and approaches to tackle a broad spectrum of problems. Practices such as limiting time spent on reactive operational work, blameless postmortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and interesting dynamic day-to-day work. SRE's culture of diversity, intellectual curiosity, problem solving and openness is important to our success. Our organization brings together people with a wide variety of backgrounds, experiences and perspectives. We encourage them to collaborate, think big and take risks in a blame-free environment. We promote self-direction to work on meaningful projects, while we also strive to build an environment that provides the support and mentorship needed to learn and grow.
What you’ll be doing:
Lead the technical strategy and roadmap for large-scale, cross-functional SRE initiatives that improve reliability, scalability, and developer productivity across enterprise systems.
Design, and build resilient distributed systems that power NVIDIA’s next-generation AI-driven enterprise products and services.
Architect and develop AI Agents, AI Skills to accelerate platform operations
Drive automation and observability improvements, using metrics and analytics to enhance performance, reliability, and efficiency.
Collaborate across Cloud, Platform, Security, and AI/ML teams to implement modern SRE components that ensure high availability and secure operations.
Analyze and troubleshoot complex systems, championing best practices in system design, incident management, and postmortem analysis.
Mentor and influence engineers across teams, fostering technical excellence and a culture of reliability engineering.
What we need to see:
10+ years of experience in Site Reliability Engineering, Platform Engineering, or Cloud Architect roles.
BS degree in Computer Science or a related technical field involving coding (e.g., physics or mathematics), or equivalent experience
Strong proficiency in programming languages such as Python, Typescript, JavaScript, or Go, with a focus on automation and infrastructure-as-code.
Experience with infrastructure-as-code such as AWS CDK, AWS CloudFormation, Terraform or CrossPlane
Solid understanding of OpenTelemetry or other Observability implementation at scale.
Deep expertise in systems architecture, networking, Kubernetes, and public cloud services (AWS, Azure, or GCP).
Outstanding problem-solving, communication, and teamwork skills, with the ability to influence across technical and interpersonal boundaries.
Ways to stand out from the crowd:
Passion for and experience with Public Cloud or large-scale automation systems.
Demonstrated ability to drive technical strategy and deliver measurable reliability outcomes in complex environments.
A strong sense of ownership, curiosity, and innovation, you thrive in ambiguity and turn challenges into opportunities.
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables outstanding creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence.
#LI-Hybrid
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until September 12, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Related keywords
Site Reliability EngineeringAI PlatformKubernetesPythonTypescriptJavaScriptGoAWSAzureGCPTerraformAWS CDKAWS CloudFormationCrossPlaneOpenTelemetryInfrastructure-as-code
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
Offices: 2701 San Tomas Expressway, Santa Clara, CA 95050, US · No. 8, Ji Hu Rd., Taipei City, Taipei City 114, TW · Nanakramguda, Serilingampally Mandal, Plot # 6A&B, IT Park Layout, RR District, Hyderabad, Telangana 500046, IN · No. 127 Andheri Kurla Road, CNB Square, Mumbai, Village Chakala, Andheri East 400 093, IN · Survey No. 144/145, Samrat Ashok Path, Off Airport Road, Pune, Yerwada 411 006, IN
Based on 1388 listings with disclosed salaries, most software jobs in Santa Clara, CA pay between $131k–$328k per year. Individual offers vary with seniority, company size, and specialization.
How many Software jobs are open in Santa Clara, CA right now?
There are currently 1,521 open software positions in Santa Clara, CA listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Software roles in Santa Clara, CA?
Companies currently hiring include NVIDIA, ServiceNow, Qualcomm, AMD, Applied Materials, among others. Browse the listings above to see every active employer.
Are there remote or hybrid Software jobs in Santa Clara, CA?
Yes — 450 of the 1521 open software positions offer remote or hybrid work (106 remote, 344 hybrid).
How do I apply for Software jobs in Santa Clara, CA?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.