Principal Engineer, Cloud Site Reliability Engineering
Santa Clara, California, United States · On-site
$272k–$431k/yr
Senior+$29B raised
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with va…
We are looking for an experienced Firmware Engineer to join the NIC Firmware team at the Yokneam site. The Firmware team develops cutting edge networking features for cloud, HPC and storage. We drive the data growth of t…
Skills: C, C++, Firmware Development, Networking Protocols, Data Structures
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping i…
Today, NVIDIA is tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing…
Skills: Large Language Models, Agentic AI, Python, Java, Go
Tech Lead Ethernet Networking Verification Engineer
Austin, Texas, United States · On-site
$184k–$288k/yr
Senior+$29B raised
NVIDIA wants a highly qualified engineer for its Ethernet Switch Networking Development team based in Austin. You will work within a rapidly growing team to better serve customers and show results. This role blends leade…
Principal Linux Kernel Engineer – NVIDIA CPU Server Platforms
Santa Clara, California, United States · On-site
$272k–$431k/yr
Senior+$29B raised
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping i…
Skills: Linux Kernel, CPU Architecture, Memory Management, Scheduling, Virtualization
Senior Technical Program Manager, DGX Cloud Software - Product and Services
Santa Clara, California, United States · Hybrid
$200k–$322k/yr
Senior+$29B raised
NVIDIA's DGX Cloud (DGXC) powers AI for strategic research and product workloads. The company seeks a Senior Technical Program Manager (TPM) to lead complex, cross-functional programs powering NVIDIA’s next-generation AI…
Skills: Technical Program Management, Software Engineering, Cloud Infrastructure, System Integration, AI Workloads
NVIDIA's Performance Lab (PerfLab) builds the systems and automation used to evaluate the performance and quality of accelerated computing and AI workloads. We turn complex benchmark experiments into reliable, scalable, …
Skills: Python, Linux, System Software, Infrastructure Engineering, Distributed Systems
Developer Technology Engineer, Public Sector - New College Grad 2026
Santa Clara, California, United States · On-site
$124k–$242k/yr
Entry level$29B raised
Our work at NVIDIA is dedicated towards a computing model focused on visual and AI computing. For two decades, NVIDIA has pioneered visual computing, the art and science of computer graphics, with our invention of the GP…
Skills: C++, Fortran, CUDA, OpenACC, Linear Algebra
At NVIDIA, we pride ourselves in having energy-efficient products. We believe that continuing to maintain our products' energy efficiency compared to the competition is key to our continued success. Our team researches a…
Skills: Python, C++, Artificial Intelligence, Machine Learning, Deep Learning
We are now looking for an ASIC Top Floorplan Design Engineer. NVIDIA is seeking a talented ASIC Floorplan Engineer to design and implement the world’s leading SoC's and GPU's. This position offers you a unique opportunit…
Skills: ASIC Floorplanning, Physical Design, Verilog, System Verilog, Python
Senior Systems Software Engineer, Windows and Linux Enablement - DGX Station
Santa Clara, California, United States · On-site
$224k–$357k/yr
Senior+$29B raised
DGX Station is NVIDIA’s next-generation personal AI supercomputer—a deskside workstation built on the NVIDIA Grace Blackwell GB300 Superchip with massive coherent CPU+GPU memory, designed to bring data-center-class AI ca…
Skills: Windows Internals, Linux Kernel, Driver Development, Firmware, UEFI
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping i…
Skills: C++, System software, 3D graphics, Direct3D, Vulkan
Research Scientist, AI for Clinical Translatability
Santa Clara, California, United States · Remote OK
$168k–$265k/yr
Senior$29B raised
For decades, NVIDIA has been at the forefront of accelerated computing, driving innovation in fields ranging from computer graphics to artificial intelligence. Today, we're leveraging this legacy to revolutionize digital…
Skills: Machine learning, Computational biology, Drug discovery, PyTorch, JAX
Senior Developer Relations Manager, Physical AI Metropolis
Germany · Remote OK
Senior$29B raised
Are you a visionary at the intersection of AI innovation and executive influence? Do you thrive on shaping the future of visual understanding and AI-driven operational intelligence? If you’re a bold, strategic problem so…
Senior Manager, Engineering - Data Center Firmware
Santa Clara, California, United States · On-site
$272k–$431k/yr
Senior+$29B raised
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning -…
Skills: Firmware Engineering, OpenBMC, MCU Firmware, Data Center Architecture, System Software
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. Doing what’s never…
Senior Embedded System Software Engineer - Hypervisor and Virtualization
Taipei, Taiwan · On-site
Senior$29B raised
NVIDIA has continuously reinvented itself over two decades. NVIDIA’s invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More…
Skills: Embedded systems, Hypervisor, Linux, C, ARM processor
Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, an…
Skills: Kubernetes, Go, Python, GitOps, Infrastructure As Code
For over 25 years, NVIDIA has been revolutionizing computer graphics, PC gaming, and accelerated computing. It’s an outstanding legacy of innovation that’s motivated by great technology—and amazing people. Today, we’re t…
Skills: AI Agents, LLM, Prompt Engineering, RAG, Tool Calling
Principal Engineer, Cloud Site Reliability Engineering
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
$272k–$431k/yr
Full-time
bachelor degree, postgraduate degree
Equity, Health Insurance
Posted 3d ago
~40 hrs/week
Responsibilities
The SRE Architect will lead the GPU Private Cloud team to optimize software development workflows and CI/CD systems for internal NVIDIA teams. They are responsible for identifying performance bottlenecks, architecting scalable solutions, and technically directing a team of engineers.
Requirements
Candidates must have 15+ years of systems software development experience, including expertise in cloud infrastructure and AI. A bachelor's or master's degree in Electrical Engineering, Computer Science, or a relevant field is required.
Full job description
NVIDIA is looking for a Cloud Site Reliability Engineering Architect to work in IPP's (Infrastructure, Planning and Process) Cloud Infrastructure Team. IPP is a global organization within NVIDIA. This group works with various other groups within NVIDIA such as Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence and Autonomous Vehicles to cater to their infrastructure needs. These cloud services provide almost half a million automated jobs per day on thousands of servers helping with the efficiency of thousands of NVIDIA's software engineers worldwide. The cloud hosts various machines and devices with operating systems like Windows, Linux, and Android. It supports hardware platforms including NVIDIA GPUs and Tegra Processors. It delivers unified CI/CD solutions and cloud-based software development. Are you passionate about distributed infrastructure and looking for sophisticated, critical issues, ready to build the next generation of cloud services, design creative solutions, mine through data to uncover real problems and fix them?
What you'll be doing:
Serve as an SRE Architect part of GPU Private Cloud team used by thousands of NVIDIANs globally for interactive development, centralized CI/CD, and QA testing.
Evaluating, identifying and developing software solutions to optimize critical software development workflows across various organizations within NVIDIA.
Architecting, implementing, and supporting end-to-end CI/CD system using open-source and NVIDIA proprietary software.
Customer (NVIDIA Internal development teams) onboarding to Private cloud infrastructure with a good discovery of the use case and available solutions within the cloud.
Identify performance bottlenecks and optimize the speed and cost efficiency of AI development and testing systems.
Leading software development projects and technically direct a team of brilliant engineers and guide them to provide efficient and impactful solutions.
Looking for problems within software systems and resolving the issues
Craft and implement critical metrics using various analytics methods and dashboards.
What we need to see:
BS or MS in Electrical Engineering, Computer Science, or relevant field (or equivalent experience).
15+ years of systems software development including at least 1 year dedicated to developing/exploring AI.
Experience of maintaining cloud infrastructure and highly available production environment.
Strong programming and software development skills in JAVA, Python, Shell-script along with good understanding of distributed systems and REST APIs.
Experience in working with SQL/NoSQL database systems such as MySQL, Cassandra, MongoDB or Elasticsearch.
Excellent knowledge and working experience with Docker containers and Virtual Machines.
Good background of Cloud technologies like: OpenStack, Docker, Kubernetes, Chef/Puppet, Hadoop/Ceph/SwiftStack, LXC, Git, Perforce, JFrog, Kafka.
Ability to work across organizational boundaries effectively to improve alignment and productivity between teams in a multi-national, multi-time-zone corporate environment.
Ways to stand out from the crowd:
Depth in AI, Machine Learning and Deep Learning algorithms and techniques.
Strong collaborative and interpersonal skills, with a consistent record of guiding and influencing others in dynamic environments.
Experience developing large-scale software systems using modular architecture under real-time performance requirements.
Background in designing high-performance, scalable software systems with a strong focus on hardware cost optimization.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until August 9, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Related keywords
Cloud Site Reliability EngineeringGPU Private CloudCI/CDJavaPythonShell-scriptDistributed systemsREST APIsMySQLCassandraMongoDBElasticsearchDockerKubernetesOpenStackChef
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
Offices: 2701 San Tomas Expressway, Santa Clara, CA 95050, US · No. 8, Ji Hu Rd., Taipei City, Taipei City 114, TW · Nanakramguda, Serilingampally Mandal, Plot # 6A&B, IT Park Layout, RR District, Hyderabad, Telangana 500046, IN · No. 127 Andheri Kurla Road, CNB Square, Mumbai, Village Chakala, Andheri East 400 093, IN · Survey No. 144/145, Samrat Ashok Path, Off Airport Road, Pune, Yerwada 411 006, IN
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
Offices: 2701 San Tomas Expressway, Santa Clara, CA 95050, US · No. 8, Ji Hu Rd., Taipei City, Taipei City 114, TW · Nanakramguda, Serilingampally Mandal, Plot # 6A&B, IT Park Layout, RR District, Hyderabad, Telangana 500046, IN · No. 127 Andheri Kurla Road, CNB Square, Mumbai, Village Chakala, Andheri East 400 093, IN · Survey No. 144/145, Samrat Ashok Path, Off Airport Road, Pune, Yerwada 411 006, IN