You'll own the reliability of Luma's 10k+ GPU fleet: the scheduling, efficiency, and resilience that research and products depend on. As a Staff AI Infrastructure Engineer, you'll be a technical authority who turns deep …
Forward Deployed Engineers turn what Luma's models can do into systems customers actually rely on. You embed with a customer, learn their workflow, define the problem with them, and build the production system that solve…
Skills: Software Engineering, Python, TypeScript, Go, System Architecture
Team: Infra Reliability · SF Bay Area / Remote (US) You'll own the GPU infrastructure Luma's research and product run on — thousands of NVIDIA and AMD GPUs across on-prem and multi-cloud (AWS and OCI). As a Senior SRE, y…
You'll set technical direction for how products get built at Luma, directing AI coding agents to ship production systems fast while making the disciplined calls on scope, structure, and trade-offs. This is a senior role …
Skills: System design, AI coding agents, Technical leadership, Product intuition, Distributed systems
To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts. Job Category Software Engineering Job Details About Salesforce Salesforc…
Skills: ReactJS, TypeScript, Java, Spring Framework, AWS
You'll own the infrastructure that tells Luma whether its models are getting better. As a Research Engineer on Evaluations, you'll build the pipelines, metrics, and automated systems that close the loop between model out…
At some point in any digital investigation, an analyst needs to step beyond the perimeter and engage threats at the source. Authentic8 Silo places any type of digital analyst in region-specific, multi-application workspa…
Skills: Jamf Pro, Microsoft Intune, Endpoint Management, Google Workspace, Microsoft 365
To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts. Job Category User Experience Job Details About Salesforce Salesforce is …
You'll own how Luma's models get served — integrating new architectures into the inference engine, scaling deployments across thousands of machines, and keeping expensive GPU fleets busy while meeting internal SLOs. This…
Skills: Python, System architecture, PyTorch, Hugging Face, vLLM
You'll lead the team that owns Luma's entire inference serving stack — routing, scheduling, and fleet-wide orchestration across thousands of GPUs, multiple clouds, and hardware vendors — where all of Luma's compute meets…
Skills: Distributed systems, ML infrastructure, Model serving, Inference platforms, LLM
You'll turn Luma's industry-leading generative video models into world models: interactive, controllable, physically faithful, and useful as a substrate for embodied reasoning. This is the role at the center of the thesi…
Skills: Generative modeling, World models, PyTorch, Computer vision, Robotics
You'll own how Luma brings its models to market for creators, brands, and studios: the positioning, the launches, and the first strategic partnerships that turn frontier research into tools people actually adopt. This is…
As Luma's founding Robotics Engineer, you'll bring up commercial robot platforms — humanoids, arms, mobile bases — wire up the sensors and data pipelines our world models need, and run the experiments that tell us whethe…
Skills: Robotics, Systems integration, ROS, ROS2, Python
Research Scientist / Engineer — Foundation Model (Agent)
Redwood City, California, United States · Hybrid
Senior$1.1B raised
You'll build and train large-scale multimodal agentic models — systems that reason, plan, code, and call tools to do complex, multi-step work over pixels. This is core research shaping how users interact with what Luma's…
Skills: Machine Learning, Foundation Models, Agentic Systems, Multimodal Models, PyTorch
Research Scientist / Engineer – Training Infrastructure
Redwood City, California, United States · Hybrid
Senior$1.1B raised
You'll build the distributed systems that train Luma's large-scale multimodal models across thousands of GPUs, so researchers can focus on innovation on top of reliable, efficient, scalable infrastructure. This is hard P…
You'll own features end to end, from a fuzzy concept to a shipped, polished experience, and you'll design the interaction paradigms for a product where AI is the medium, not a feature bolted on. This is a hands-on genera…
Skills: Product Design, Interaction Design, Visual Design, Prototyping, AI Tools
Research Scientist / Engineer – Reinforcement Learning Infrastructure
Redwood City, California, United States · Hybrid
Senior$1.1B raised
You'll build the systems that make reinforcement learning work at frontier scale — coupling policy optimization with large fleets of inference workers, agentic environments, and the reward and verification systems that t…
You'll help define the simulation substrate Luma uses to train general-purpose robot policies — a faithful, controllable simulation of the world built on our generative video and 3D models. You'll sit at the boundary bet…
Research Scientist / Engineer – Performance Optimization
Redwood City, California, United States · Hybrid
Senior$1.1B raised
You'll make Luma's multimodal models fast — profiling and optimizing GPU, CPU, and accelerator code so they train efficiently and deploy at scale without sacrificing quality. You'll write the kernels and operations that …
You'll own the bridge between Luma's frontier research and its products — Canvas, Agents, and the model platform — making sure what researchers build is shaped by what customers need, and what ships takes full advantage …
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
$230k–$360k/yr
Full-time
Equity
Posted 19d ago
~40 hrs/week
Responsibilities
You will own the reliability, efficiency, and scaling of a large-scale GPU infrastructure fleet while building and leading a team of systems engineers. You will also partner with research teams to architect infrastructure that supports new model capabilities and high-performance inference.
Requirements
The role requires deep expertise in Linux, distributed systems, and operating GPU clusters within production environments. Candidates must possess strong fluency in Kubernetes and the ability to debug complex issues across hardware, kernels, and orchestration layers.
Full job description
You'll own the reliability of Luma's 10k+ GPU fleet: the scheduling, efficiency, and resilience that research and products depend on. As a Staff AI Infrastructure Engineer, you'll be a technical authority who turns deep systems knowledge into repeatable, company-wide reliability, and a leader other strong engineers want to work with.
This is close-to-the-metal work — kernels, containers, schedulers, networking, storage, GPU behavior — under demand hard enough that yesterday's solutions break regularly. It's also a technical-leadership role: you'll set the bar and grow the team. If most of your experience has been inside highly abstracted internal platforms where others owned the underlying machinery, this likely isn't a match.
What You'll Own
Architect and operate large, heterogeneous GPU environments under extreme demand, improving utilization and performance where small gains change company outcomes.
Resolve failures spanning hardware, OS, runtimes, and orchestration, and eliminate whole classes of instability.
Define how infrastructure and workloads evolve as cluster size and concurrency grow — scheduling, placement, resource management.
Work directly with research to build the systems new model capabilities require, and scale inference without sacrificing reliability or latency.
Hire and develop exceptional systems and reliability engineers, and set the bar for depth, judgment, and production ownership.
Shape product and research architecture early through strong partnerships.
First 90 Days
One way the first 90 could unfold.
Days 1–30 — Immerse & Diagnose: Learn the fleet, its failure modes, and the biggest reliability and utilization gaps.
Days 30–60 — Ship & Validate: Eliminate a recurring class of instability or land a utilization or performance win that moves company outcomes.
Days 60–90 — Scale & Systemize: Set the reliability direction, redesign ahead of where today's abstractions will fail, and begin building the team.
What You Bring
Deep expertise in Linux and distributed systems.
Experience operating GPU or accelerator clusters in real production environments.
Strong fluency in Kubernetes and modern open-source infrastructure.
Comfort debugging across hardware, kernel, runtime, and orchestration, and understanding how systems behave under contention and at scale.
You write code and build automation, and think in bottlenecks, failure modes, and trade-offs.
Judgment engineers trust, especially when things break.
Nice to Have
You raise reliability standards company-wide and influence product and research architecture early.
You build partnerships rather than ticket queues, and attract and level up strong engineers.
Curiosity for how models use infrastructure, because improving systems expands what becomes possible.
About Luma: Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.
Related keywords
Staff AI Infrastructure EngineerGPULinuxDistributed SystemsKubernetesInfrastructureReliabilityKernelContainersNetworkingStorageOrchestrationAutomationInferenceLatencySystems Engineering
Luma AI’s mission is to build Multimodal AGI: AI that can generate, understand, and operate in the physical world.
We develop multimodal models across video, 3D, and generative media, and ship them in products like Dream Machine to help creators and teams turn ideas into compelling visuals—fast.
Offices: San Francisco Bay Area, CA, US
Machine LearningGenerative MediaGenerative AIand AI VideoGraphic DesignMedia and EntertainmentVirtual RealityArtificial IntelligenceAugmented RealityFoundational AI
How much do Software jobs in Redwood City, CA pay?
Based on 206 listings with disclosed salaries, most software jobs in Redwood City, CA pay between $131k–$261k per year. Individual offers vary with seniority, company size, and specialization.
How many Software jobs are open in Redwood City, CA right now?
There are currently 257 open software positions in Redwood City, CA listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Software roles in Redwood City, CA?
Companies currently hiring include Delinea, Luma, Box, C3 AI, Retell, among others. Browse the listings above to see every active employer.
Are there remote or hybrid Software jobs in Redwood City, CA?
Yes — 185 of the 257 open software positions offer remote or hybrid work (54 remote, 131 hybrid).
How do I apply for Software jobs in Redwood City, CA?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.