Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, an…
Skills: Kubernetes, Go, Python, GitOps, Infrastructure as code
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. We are now seeking a h…
Skills: Software verification, SONiC NOS, Python, CI/CD, Jenkins
NVIDIA’s System-on-Chip (SOC) group is seeking a motivated Padring Verification Engineer to develop scalable verification methodologies and deliver high-quality padring verification results for complex GPU and Tegra SOCs…
NVIDIA is seeking a Senior SoC Design Engineer to design the next-generation SoCs. We are looking for special individuals to deliver innovative products. Together, we will build the next generation of life-changing SoCs.…
Silicon Performance, Power and Binning Tools Engineer
Shanghai, Shanghai, China · On-site
Mid level$29B raised
Every CPU, GPU, and Tegra SoC NVIDIA has shipped in the past four years passed through our toolchain on its way to production. Over 200 product SKUs were optimized during the Blackwell generation alone! Now we're looking…
Skills: Python, Perl, LLMs, AI Agents, Computer Architecture
NVIDIA is growing its SDK Verification team and looking for engineers early in their careers who want to go deep on verification, automation, and networking technology. You'll work alongside SDK development and architect…
Hardware Program Manager – HW Fulfillment and R&D Procurement
Rawabi, West Bank, Palestinian Territory · On-site
Mid level$29B raised
NVIDIA is looking for a Hardware Program Manager with technical expertise to join our HW R&D organization, focusing on Hardware Fulfillment, HW Component Orders, and Prototyping Operations for our advanced Network produc…
Skills: Hardware Program Management, Procurement, Supply Chain Management, NPI, Hardware Development Cycles
Technical Product Marketing Engineer, Metropolis - New College Grad 2026
Santa Clara, California, United States · On-site
$92k–$184k/yr
Entry level$29B raised
NVIDIA has been redefining computer graphics, accelerated computing, and AI for more than 25 years - an outstanding legacy of innovation fueled by groundbreaking technology and phenomenal people. Today, we are tapping in…
Skills: Vision AI, Agentic frameworks, Model training, Model fine-tuning, Agent orchestration
NVIDIA's GPUs and SOCs are the world leaders in power, performance and efficiency. We are continually innovating to deliver new and creative, unusual solutions to extraordinary problems in a wide range of sectors. To thi…
NVIDIA’s GPUs and SOCs are the world leaders in power, performance, and efficiency. We are continually innovating to deliver new and creative, unusual solutions to extraordinary problems in a wide range of sectors. To th…
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by phenomenal technology—and amazing people. Today, we’re tapp…
Skills: Account Management, Game Development, Technical Partnerships, Business Development, GPU Acceleration
For more than 25 years, NVIDIA has redefined computer graphics, video game development, and enhanced processing capabilities. Today, we're leading the next era of AI, building the platforms that power intelligent applica…
Skills: Technical program management, Software engineering, Data science, Data processing, Accelerated computing
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping i…
Skills: C, C++, System Architecture, Embedded Systems, Multithreading
We are now seeking a Senior Infrastructure Software Engineer for NVIDIA TensorRT Edge-LLM! NVIDIA's TensorRT Infrastructure group is seeking excellent software engineers to enable the next generation of edge AI. This is …
Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing wha…
Senior Deep Learning Sofware Infrastructure Engineer
Canada · Remote OK
$224k–$431k/yr
Senior+$29B raised
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping i…
Skills: Deep Learning, Distributed Systems, PyTorch, GPU Clusters, Python
NVIDIA is at the forefront of the AI revolution, and our research is shaping the future of large language models. We are looking for a Senior Scientist to join our team and help advance our capabilities in synthetic data…
Skills: Synthetic data generation, Large language models, Generative modeling, Multimodal machine learning, Software engineering
Nvidia DRIVE platform offers solutions to build safe, scalable, AI-enabled Autonomous Vehicles. The end‑to‑end full stack platform spans from in‑car supercomputers to cloud‑scale training and simulation. Our goal is to s…
NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More rece…
Senior Manufacturing and System Co-Design Workflow Engineer
Santa Clara, California, United States · On-site
$196k–$311k/yr
Senior+$29B raised
Build the infrastructure that keeps every NVIDIA chip aligned from first spec to final shipment. NVIDIA's Silicon Co-Design Group sits at the convergence of architecture, silicon, systems, and manufacturing. The System–M…
Skills: Python, Systems Engineering, Silicon Bring-up, Productization Engineering, Data Pipelines
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Full-time
bachelor degree
Posted 11h ago
~40 hrs/week
Responsibilities
You will design, build, and operate the Kubernetes platform that powers NVIDIA's global network automation, telemetry, and operations. Additionally, you will manage the lifecycle of Kubernetes environments and provide production support for network services.
Requirements
Candidates must have 8+ years of experience in building or operating production Kubernetes platforms and distributed systems. Proficiency in Go or Python, along with deep expertise in Kubernetes networking, storage, and cluster lifecycle management, is required.
Full job description
Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, and operate NVIDIA’s global network across data centers, colocation facilities, and cloud environments. The team owns the architecture and lifecycle of this platform, including cluster provisioning and upgrades, GitOps delivery, observability, capacity, and service enablement. We build software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments.
We are looking for a hands-on senior engineer to own the lifecycle and automation of the Kubernetes platform supporting GNI network systems. You will also provide production support for network services running on the platform, partnering with their engineering owners when issues or changes cross the platform boundary. You will take complex problems from design through production and remain accountable for the outcome. You will bring deep Kubernetes expertise and help establish consistent engineering practices across the US and Bangalore teams. This is a senior individual contributor role with end-to-end ownership and production responsibility.
What You’ll Be Doing:
Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments.
Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.
Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps.
Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features.
Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. Drive issues from initial signal through verified resolution.
Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery.
Participate in CFR’s production on-call rotation, including scheduled after-hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion.
What We Need to See:
Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems.
Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.
Proficiency in at least one general-purpose programming language, such as Go or Python.
Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.
Experience deploying and supporting network automation or telemetry services on Kubernetes.
Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion.
Ways to Stand Out From the Crowd:
Strong knowledge of IP routing, data center fabrics, and cloud networking is a great plus.
Experience designing and operating large, multi-region Kubernetes fleets, including fleet-wide upgrades and recovery.
Hands-on experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle, machine remediation, and upgrades.
Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns.Experience designing or operating network automation and telemetry services on Kubernetes at global scale.
Contributions to Cluster API, Metal3, or other open-source Kubernetes infrastructure projects.
NVIDIA’s deep learning platforms have made major impact to various fields is broadly used across leading academic institutions, start-ups, and industry, including the world’s largest Internet companies. We need passionate, hard-working and creative people to help us take on more of these unique opportunities in deep learning cloud solutions. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you.
Related keywords
KubernetesNetwork InfrastructureCloud Foundations ReliabilityGitOpsAutomationObservabilityGoPythonCluster APIMetal3CI/CDInfrastructure as CodeData CenterCloud EnvironmentsTelemetryIP Routing
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
Offices: 2701 San Tomas Expressway, Santa Clara, CA 95050, US · No. 8, Ji Hu Rd., Taipei City, Taipei City 114, TW · Nanakramguda, Serilingampally Mandal, Plot # 6A&B, IT Park Layout, RR District, Hyderabad, Telangana 500046, IN · No. 127 Andheri Kurla Road, CNB Square, Mumbai, Village Chakala, Andheri East 400 093, IN · Survey No. 144/145, Samrat Ashok Path, Off Airport Road, Pune, Yerwada 411 006, IN
Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.
Offices: 2701 San Tomas Expressway, Santa Clara, CA 95050, US · No. 8, Ji Hu Rd., Taipei City, Taipei City 114, TW · Nanakramguda, Serilingampally Mandal, Plot # 6A&B, IT Park Layout, RR District, Hyderabad, Telangana 500046, IN · No. 127 Andheri Kurla Road, CNB Square, Mumbai, Village Chakala, Andheri East 400 093, IN · Survey No. 144/145, Samrat Ashok Path, Off Airport Road, Pune, Yerwada 411 006, IN