Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, an…
Skills: Kubernetes, Go, Python, GitOps, Infrastructure As Code
NVIDIA is the world leader in accelerated computing and artificial intelligence and has been powering Deep Learning, AI, accelerated data analytics, autonomous systems and robots around the world. NVIDIA's customers incl…
Skills: CUDA programming, GPU platforms, Deep Learning, Machine Learning, AI agents
We are now looking for an ASIC Top Floorplan Design Engineer. NVIDIA is seeking a talented ASIC Floorplan Engineer to design and implement the world’s leading SoC's and GPU's. This position offers you a unique opportunit…
Skills: ASIC Floorplanning, Physical Design, Verilog, System Verilog, Python
NVIDIA is looking for an experienced software engineer with infrastructure experience to become a senior member of the Cloud Foundations Automation - Development Team. We build and manage the automation ecosystem support…
Principal Product Manager, NVIDIA Trust Services – DGX Cloud
Santa Clara, California, United States · Hybrid
$240k–$380k/yr
Senior+$29B raised
Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing wha…
Senior Systems Software Engineer, Windows and Linux Enablement - DGX Station
Santa Clara, California, United States · On-site
$224k–$357k/yr
Senior+$29B raised
DGX Station is NVIDIA’s next-generation personal AI supercomputer—a deskside workstation built on the NVIDIA Grace Blackwell GB300 Superchip with massive coherent CPU+GPU memory, designed to bring data-center-class AI ca…
Skills: Windows Internals, Linux Kernel, Driver Development, Firmware, UEFI
NVIDIA’s Networking Silicon Engineering group is looking for a Local Product Engineer. NVIDIA's GPUs,DPUs and SOCs are the world leaders in performance and efficiency, and we are continually innovating in creative and ou…
NVIDIA's Object Storage Platform team builds and operates the company's internal S3-compatible distributed object storage service — a critical piece of infrastructure that stores, manages, and serves exabytes of data acr…
Santa Clara, California, United States · Remote OK
$168k–$322k/yr
Senior$29B raised
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping i…
Senior System Software Engineer, SoC Power and Performance - O-RAN Infrastructure
Santa Clara, California, United States · On-site
$152k–$242k/yr
Senior$29B raised
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping i…
Skills: System software, SoC architecture, Firmware, C, C++
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping i…
Skills: Python, Linux, CI/CD, Test automation, Infrastructure engineering
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping i…
Skills: Infrastructure Engineering, CI/CD, Test Automation, Python, Infrastructure-as-code
Senior Technical Program Manager, DGX Cloud - Trust Services
Santa Clara, California, United States · On-site
$200k–$322k/yr
Senior+$29B raised
NVIDIA is seeking a Senior Technical Program Manager to lead Trust Services programs for DGX Cloud. DGX Cloud powers large-scale AI infrastructure across NVIDIA, cloud service providers, and NVIDIA Cloud Partners, making…
Senior AI Tools Engineer, SRE Operations - GeForce NOW
Santa Clara, California, United States · Remote OK
$144k–$230k/yr
Senior$29B raised
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping i…
NVIDIA is seeking a highly skilled Senior Performance Engineer to join our Performance and R&D organizations. In this role, you will help build and evolve systems that support performance analysis, telemetry, and optimiz…
Skills: Performance analysis, Systems engineering, HPC, AI infrastructure, RDMA
Developer Technology Engineer, AI - New College Grad 2026
Santa Clara, California, United States · On-site
$124k–$242k/yr
Entry level$29B raised
Today, NVIDIA is tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing…
Skills: C, C++, Software design, AI algorithms, Parallel programming
Senior Software Engineer - Manufacturing and Factory
Santa Clara, California, United States · On-site
$184k–$357k/yr
Senior+$29B raised
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping i…
Skills: System Architecture, Server Systems Design, Networking Protocols, Firmware Flashing, Security Provisioning
NVIDIA is looking for a Senior Network Engineer to develop a cloud network infrastructure. The goal is to craft a reliable, scalable and efficient network to support NVIDIA software development workflows and tools, inclu…
We are looking for a Storage Software Architect to join the architecture group. You will be part of a team that shapes the next generation of storage for AI: training, inferencing, KV cache, RAG and more. Utilizing the b…
Skills: Storage Architecture, AI Workloads, DPUs, NICs, Software Stack Design
Senior Manager, Engineering - Data Center Firmware
Santa Clara, California, United States · On-site
$272k–$431k/yr
Senior+$29B raised
NVIDIA's invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern deep learning -…
Skills: Firmware Engineering, OpenBMC, MCU Firmware, Data Center Architecture, System Software
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
$168k–$334k/yr
Full-time
bachelor degree
Equity, Benefits
Posted 4d ago
~40 hrs/week
Responsibilities
Design, build, and operate the Kubernetes platform supporting NVIDIA's global network infrastructure and automation services. Own the full lifecycle of cluster provisioning, upgrades, and production support, including on-call incident response.
Requirements
Requires a Bachelor's degree and 8+ years of experience operating production Kubernetes platforms or distributed systems. Proficiency in Go or Python and deep expertise in GitOps and network telemetry services are essential.
Full job description
Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, and operate NVIDIA’s global network across data centers, colocation facilities, and cloud environments. The team owns the architecture and lifecycle of this platform, including cluster provisioning and upgrades, GitOps delivery, observability, capacity, and service enablement. We build software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments.
We are looking for a hands-on senior engineer to own the lifecycle and automation of the Kubernetes platform supporting GNI network systems. You will also provide production support for network services running on the platform, partnering with their engineering owners when issues or changes cross the platform boundary. You will take complex problems from design through production and remain accountable for the outcome. You will bring deep Kubernetes expertise and help establish consistent engineering practices across the US and Bangalore teams. This is a senior individual contributor role with end-to-end ownership and production responsibility.
What You’ll Be Doing:
Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments.
Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.
Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps.
Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features.
Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. Drive issues from initial signal through verified resolution.
Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery.
Participate in CFR’s production on-call rotation, including scheduled after-hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion.
What We Need to See:
Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems.
Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.
Proficiency in at least one general-purpose programming language, such as Go or Python.
Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.
Experience deploying and supporting network automation or telemetry services on Kubernetes.
Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion.
Ways to Stand Out From the Crowd:
Strong knowledge of IP routing, data center fabrics, and cloud networking is a great plus.
Experience designing and operating large, multi-region Kubernetes fleets, including fleet-wide upgrades and recovery.
Hands-on experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle, machine remediation, and upgrades.
Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns.Experience designing or operating network automation and telemetry services on Kubernetes at global scale.
Contributions to Cluster API, Metal3, or other open-source Kubernetes infrastructure projects.
NVIDIA’s deep learning platforms have made major impact to various fields is broadly used across leading academic institutions, start-ups, and industry, including the world’s largest Internet companies. We need passionate, hard-working and creative people to help us take on more of these unique opportunities in deep learning cloud solutions. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD - 276,000 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until August 8, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Related keywords
KubernetesDGX CloudNetwork InfrastructureGitOpsCluster APIMetal3GoPythonCI/CDInfrastructure As CodeTelemetryDistributed SystemsIP RoutingData Center FabricsCloud NetworkingCustom Resources
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
Offices: 2701 San Tomas Expressway, Santa Clara, CA 95050, US · No. 8, Ji Hu Rd., Taipei City, Taipei City 114, TW · Nanakramguda, Serilingampally Mandal, Plot # 6A&B, IT Park Layout, RR District, Hyderabad, Telangana 500046, IN · No. 127 Andheri Kurla Road, CNB Square, Mumbai, Village Chakala, Andheri East 400 093, IN · Survey No. 144/145, Samrat Ashok Path, Off Airport Road, Pune, Yerwada 411 006, IN
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
Offices: 2701 San Tomas Expressway, Santa Clara, CA 95050, US · No. 8, Ji Hu Rd., Taipei City, Taipei City 114, TW · Nanakramguda, Serilingampally Mandal, Plot # 6A&B, IT Park Layout, RR District, Hyderabad, Telangana 500046, IN · No. 127 Andheri Kurla Road, CNB Square, Mumbai, Village Chakala, Andheri East 400 093, IN · Survey No. 144/145, Samrat Ashok Path, Off Airport Road, Pune, Yerwada 411 006, IN