San Francisco, California, United States · On-site
Senior+$1.1B raised
The Role Inference is where all of Luma’s compute meets all of Luma’s users. The inference platform team owns the entire serving stack — from request routing, scheduling, and queueing to fleet-wide orchestration across t…
Skills: Distributed Systems, ML Infrastructure, Model Serving, vLLM, SGLang
Software Engineer - Site Controller, Energy Storage
San Francisco, California, United States · On-site
$180k–$238k/yr
Mid level$4.2B raised
About Redwood Materials Redwood is localizing a global battery supply chain that seamlessly integrates recovery, reuse, and recycling — keeping critical minerals in circulation and driving the energy transition. Founded …
Skills: Rust, Python, Modbus TCP, CAN, Linux System Administration
San Francisco, California, United States · Remote OK
$150k–$190k/yr
Mid level$487M raised
About us PhysicsX is a deep-tech company with roots in numerical physics and Formula One, dedicated to accelerating hardware innovation at the speed of software. We are building an AI-driven simulation software stack for…
Skills: Machine Learning, Python, 3D Point Cloud Processing, Mesh Data Manipulation, MLOps
San Francisco, California, United States · On-site
$130k–$160k/yr
Senior+$30M raised
Hornblower Group is a global leader in experience and transportation. Spanning a 100-year history, Hornblower Group’s portfolio of international offerings includes water- and land-based experiences and ferry and tr…
HOK is a collective of future-forward thinkers and designers who are driven to face the critical challenges of our time. We are dedicated to improving people's lives, serving our clients and healing the planet. Together,…
About Faire Faire is a technology wholesale platform built on the belief that the future is local. Independent retailers around the globe collectively represent a multi-hundred-billion-dollar wholesale market that has hi…
San Francisco, California, United States · On-site
$190k–$270k/yr
Senior+$534M raised
About the Role As a Senior Network Engineer at Together, you are responsible for designing, implementing, and maintaining our network infrastructure to ensure seamless connectivity and optimal performance for all user-fa…
San Francisco, California, United States · On-site
$250k–$450k/yr
Mid level$31M raised
About AfterQuery AfterQuery is an applied research lab curating data solutions for foundation model development. We serve every frontier AI lab with the mission of delivering the best data to power the best models. In do…
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make com…
Skills: Site Reliability Engineering, Kubernetes, SDN, Linux Networking, Python
San Francisco, California, United States · On-site
$200k–$300k/yr
Senior$122M raised
Who We Are Serval is an AI-native automation platform transforming how enterprises operate. We build intelligent agents that understand real-world workflows and execute them end-to-end — replacing manual processes and ri…
About Artos: At Artos, we build tools that help biopharma companies create and manage their R&D documentation in a fraction of the time. If you’re looking to join a team whose mission is to fundamentally change the way t…
San Francisco, California, United States · On-site
$180k–$350k/yr
Mid levelVisa sponsorship
Exa is an applied AI lab building a search engine unlike the world has ever seen. We build massive-scale infra to crawl the entire web, train state-of-the-art embedding models to process it, and design super high perform…
Skills: Crawling, Parsing, ML Performance, Retrieval Algorithms, Vector Databases
San Francisco, California, United States · On-site
$300k–$405k/yr
Senior+$1.5B raised
Perplexity API Platform Perplexity innovates at the frontier of AI infrastructure, search, and orchestration to serve the world's most discerning users. The Perplexity API Platform brings our technology to the world's mo…
Skills: Technical Leadership, API Design, Python, Distributed Systems, Performance Optimization
San Francisco, California, United States · On-site
Mid level
Omnifold’s Mission Every bad forecast has a physical consequence. Unnecessary goods are manufactured, shipped, and stored. Emergency air freight is needed for misallocated products. Poor production planning means workers…
San Francisco, California, United States · On-site
$120k–$200k/yr
Entry level
About Simple Simple AI is building state of the art voice AI agents for enterprise. Our realistic agents help iconic businesses like Doordash, xAI, and Omaha Steaks handle all kinds of phone operations, from customer sup…
Skills: Problem solving, Communication, Product design, Prioritization, Prompt engineering
San Francisco, California, United States · Remote OK
$170k–$230k/yr
Senior$5M raised
About the Role Roger is an AI platform that frees home health clinicians from paperwork so they can focus on what matters: delivering life-changing care to our most vulnerable elderly patients in the comfort of their hom…
Certain terms and conditions of employment for this position, including the rate of pay, benefits, etc., are currently subject to negotiation with the appropriate union. As part of the People Analytics team, the Workforc…
ABOUT ABUNDANT As the need for scaling data becomes the core bottleneck to progress, moving from general knowledge to domain expertise, and from chatbots to agents. Abundant is solving this by designing and operating sim…
Skills: Model Capability Design, Literature Review, Experimentation, Benchmarking, Data Pipeline Management
About Browserbase Browserbase is the complete platform to build and deploy agents that browse and interact with the web like humans. We provide one API key for agents to search and fetch information, fill out forms, and …
Skills: Product Design, Systems Thinking, Prototyping, Visual Craft, Typography
San Francisco, California, United States · Remote OK
$127k–$190k/yr
Senior$150M raised
About Collective: Collective is on a mission to redefine the way businesses-of-one work. Our technology and team of trusted advisors help members achieve financial independence by taking care of everything from business …
Skills: SQL, dbt, BigQuery, Metabase, Data Modeling
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Full-time
Posted 29d ago
~40 hrs/week
Responsibilities
Lead the inference engineering team by balancing hands-on technical contributions with team management and growth. Architect and optimize the serving stack to maximize efficiency, reliability, and unit economics for millions of users.
Requirements
Requires 8+ years of experience in large-scale distributed systems or ML infrastructure, specifically operating GPU fleets at scale. Must have deep expertise in LLM serving engines and a strong command of Python, PyTorch, and Kubernetes.
Full job description
The Role
Inference is where all of Luma’s compute meets all of Luma’s users. The inference platform team owns the entire serving stack — from request routing, scheduling, and queueing to fleet-wide orchestration across thousands of GPUs spanning multiple clusters, clouds, and hardware vendors. The team has a dual mandate: maximize the efficiency, reliability, and unit economics of production inference for millions of users, and enable research to move fast — new model architectures should go from research checkpoint to production in days, and our serving stack increasingly powers training itself through online reinforcement learning.
We are hiring a Tech Lead Manager to lead this team through its next phase of growth. This is a hands-on leadership role, not a pure management position: we expect you to spend at least 50% of your time as an individual contributor — designing, building, and debugging in the serving stack — alongside hiring and growing the team, setting technical direction, and partnering across research, product, and infrastructure. You lead by shipping, and you set the technical bar for the team through your own work.
What You’ll Do
Spend at least half your time hands-on in the serving stack: architect and build core platform components, own the hardest design decisions, and debug the toughest production incidents yourself
Lead, grow, and develop the inference engineering team: own hiring, coaching, and career growth, and build the team’s operational culture — on-call, incident response, capacity planning, and postmortems
Set the technical roadmap for the serving platform: model serving engines, request routing and scheduling, autoscaling, caching, observability, and deployment pipelines
Own the platform’s SLOs and economics: latency and availability targets, GPU utilization, and cost per generation across every model we serve
Partner closely with research to ship new model architectures into production on day zero, and to integrate serving into online RL and evaluation loops
Manage and optimize inference workloads across heterogeneous fleets — multiple clusters, clouds, and GPU vendors — including capacity planning and hardware bring-up
Build sophisticated scheduling and queueing systems that optimally leverage expensive GPU resources against live traffic patterns, cluster availability, and user priority
Representative Projects
Design intelligent routing and scheduling that optimizes request distribution across thousands of GPUs in multiple regions and clouds
Stand up disaggregated prefill/decode serving with tiered KV-cache reuse across GPU memory, DRAM, NVMe, and network storage
Autoscale and hot-swap models across the fleet to dynamically match GPU supply with live demand across production, research, and experimental workloads
Take a new multimodal architecture from research checkpoint to a production deployment serving millions of users, including quantization, speculative decoding, and precision/regression validation across hardware platforms
Build end-to-end tracing that follows any inference request through its full lifetime — queueing, routing, prefill, decode, and delivery
Integrate the inference stack into an online reinforcement learning pipeline where serving throughput directly gates training progress
Background
8+ years of engineering experience in large-scale distributed systems or ML infrastructure, with several years building and operating model-serving or inference platforms in production
Experience running inference platforms at scale — you have operated fleets on the order of thousands of GPUs across multiple clusters or clouds, and you understand what breaks at that scale
Technical leadership experience, including managing or leading engineers through periods of rapid growth — and a genuine desire to keep at least half your time in hands-on technical work rather than move into pure management
Deep, practical expertise in LLM and foundation-model serving engines (vLLM, SGLang, TensorRT-LLM, or equivalent) — ideally you’ve modified engine internals, debugged edge cases under load, and contributed improvements back
Strong command of the serving-performance toolkit: continuous batching, KV-cache management, quantization, speculative decoding, and parallelism strategies (TP/EP/pipeline)
Strong Python and PyTorch; experience operating services on Kubernetes at scale
Experience with queues, scheduling, traffic control, and fleet management at scale
Bonus Points
Experience serving diffusion, video, or other multimodal generative models (not just text), and with FFmpeg/multimedia processing
Experience with modern networking stacks — RDMA (RoCE, InfiniBand), NVLink — including KV-cache transfer and multi-node serving topologies
Experience across heterogeneous accelerator platforms (NVIDIA, AMD, TPU, Trainium) and the porting/validation work that comes with them
Contributions to open-source serving infrastructure (vLLM, SGLang, Ray, Kubernetes ecosystem)
Systems-language depth (Rust, C++, CUDA/HIP) for kernel- and runtime-level optimization
Luma AI’s mission is to build Multimodal AGI: AI that can generate, understand, and operate in the physical world.
We develop multimodal models across video, 3D, and generative media, and ship them in products like Dream Machine to help creators and teams turn ideas into compelling visuals—fast.
Offices: San Francisco Bay Area, CA, US
Machine LearningGenerative MediaGenerative AIand AI VideoGraphic DesignVirtual RealityArtificial IntelligenceAugmented RealityFoundational AIGenerative AI
How much do Engineering jobs in San Francisco, CA pay?
Based on 2822 listings with disclosed salaries, most engineering jobs in San Francisco, CA pay between $130k–$284k per year. Individual offers vary with seniority, company size, and specialization.
How many Engineering jobs are open in San Francisco, CA right now?
There are currently 3,686 open engineering positions in San Francisco, CA listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Engineering roles in San Francisco, CA?
Companies currently hiring include OpenAI, Anthropic, San Francisco Department of Public Health, Crusoe, Pinterest, among others. Browse the listings above to see every active employer.
Are there remote or hybrid Engineering jobs in San Francisco, CA?
Yes — 1815 of the 3686 open engineering positions offer remote or hybrid work (370 remote, 1445 hybrid).
How do I apply for Engineering jobs in San Francisco, CA?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.