Member of Technical Staff, AMD GPU Performance Engineering
San Francisco, California, United States · Remote OK
$200k–$400k/yr
SeniorVisa sponsorship$150M raised
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of…
Senior Software Engineer - Research Platform, Consumer Devices
San Francisco, California, United States · Hybrid
$293k–$325k/yr
Senior$201B raised
About the Team The Future of Computing Research team is an Applied Research team within the Consumer Devices group focused on developing new methods and models as we advance forward in our mission of building AGI that be…
Skills: Full-stack development, Backend services, API design, Data modeling, Generative AI
San Francisco, California, United States · Remote OK
$180k–$250k/yr
Senior+$41M raised
Role Engineer focusing on providing optimal user and agent experiences on the LanceDB platform. Takes high-level direction (e.g., "identify where UX is significant ahead of AX and drive more parity") and drives to result…
San Francisco, California, United States · On-site
$177k–$239k/yr
Senior+$35B raised
At eero, our mission is to serve as the central nervous system of the home. While we began by revolutionizing home WiFi, we aim to create comprehensive solutions that serve both wireless and wired connectivity needs for …
San Francisco, California, United States · On-site
Mid level
Physical Intelligence is bringing general-purpose AI into the physical world. We are a group of engineers, scientists, roboticists, and company builders developing foundation models and learning algorithms to power the r…
Senior Solutions Engineer About ArangoDB Arango makes your business data AI-ready, giving agents, apps, and assistants trusted context at scale. Every answer is traceable. Every decision is governed. No more stitching to…
San Francisco, California, United States · Remote Solely
$200k–$220k/yr
Senior+$47M raised
Senior Solutions Engineer About ArangoDB Arango makes your business data AI-ready, giving agents, apps, and assistants trusted context at scale. Every answer is traceable. Every decision is governed. No more stitching to…
Skills: Technical Sales, Consulting, AI, Databases, Big Data Systems
San Francisco, California, United States · On-site
$110k–$130k/yr
Senior
We’re hiring a Senior Project Engineer to work closely with the superintendent and project manager on estimating and scheduling efforts, subcontractor buy-out, and project documentation/ tracking. Senior project engineer…
Skills: Construction Management, Estimating, Scheduling, Subcontractor Buy-out, Project Documentation
San Francisco, California, United States · On-site
$295k–$445k/yr
Senior$201B raised
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proacti…
Artificial intelligence is rewriting every industry it touches — but construction, the second largest sector in the global economy, has barely felt it yet. That's not a problem. That's an opening. ZeroRFI is the AI compa…
Skills: System Architecture, React.js, Python, Cloud-native Systems, Microservices
Agent Post-Training, Frontier Evals and Environments Research
San Francisco, California, United States · On-site
$295k–$445k/yr
Senior$201B raised
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proacti…
Skills: Machine Learning, Software Engineering, Statistics, Large Language Models, Reinforcement Learning
San Francisco, California, United States · On-site
$140k–$200k/yr
Senior$85M raised
About Mariana Minerals Mariana Minerals is a software-first, vertically integrated minerals company on a mission to supply the critical minerals powering modern energy, AI, and defense technologies. We’re reimagining the…
San Francisco, California, United States · On-site
$295k–$445k/yr
Senior$201B raised
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proacti…
Skills: Machine Learning, Software Engineering, Statistics, Large Language Models, Reinforcement Learning
San Francisco, California, United States · On-site
$295k–$445k/yr
Senior$201B raised
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proacti…
San Francisco, California, United States · On-site
$140k–$200k/yr
Senior$85M raised
About Mariana Minerals Mariana Minerals is a software-first, vertically integrated minerals company on a mission to supply the critical minerals powering modern energy, AI, and defense technologies. We’re reimagining the…
Skills: 3D CAD, Mechatronics Design, Robotics Hardware, Drive-by-wire Actuation, Electromechanical Systems
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make com…
Skills: Storage Systems Engineering, Distributed Systems, C, C++, Rust
San Francisco, California, United States · On-site
$295k–$445k/yr
Senior$201B raised
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proacti…
San Francisco, California, United States · On-site
$295k–$445k/yr
Senior$201B raised
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proacti…
San Francisco, California, United States · On-site
$295k–$445k/yr
Senior$201B raised
About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proacti…
San Francisco, California, United States · Remote OK
$179k–$273k/yr
Senior$773M raised
Thumbtack helps millions of people confidently care for their homes. Thumbtack is the one app you need to take care of and improve your home — from personalized guidance to AI tools and a best-in-class hiring experience.…
Skills: Backend Engineering, API Design, System Integration, AI Workflows, Technical Design
Member of Technical Staff, AMD GPU Performance Engineering
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
$200k–$400k/yr
Full-time
bachelor degree
Health Insurance, Dental Insurance, Vision Insurance, 401(k) Company Match, Equity
Visa sponsorship available
Posted 41d ago
~40 hrs/week
Remote in United States
Responsibilities
Build and optimize AMD GPU backends, kernels, and runtime paths to make vLLM a first-class inference engine. Improve performance-critical paths including attention, GEMM, and communication-heavy operations using ROCm and related tooling.
Requirements
Requires a Bachelor's degree in CS or a related field with hands-on experience optimizing AMD GPU workloads and ML kernels. Candidates should possess deep knowledge of AMD GPU execution, memory behavior, and performance profiling skills.
Full job description
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.
About the Role
We're looking for an AMD GPU performance engineer to make vLLM a first-class inference engine across the AMD accelerator ecosystem. You'll build and optimize AMD GPU backends, kernels, runtime paths, and benchmarking infrastructure using ROCm, HIP, Triton, CK, AITER, and related tooling so vLLM can deliver frontier inference performance on AMD GPUs.
You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving performance-critical paths such as attention, GEMM, sampling, KV cache, and communication-heavy operations. Your work will help make AMD GPU support in vLLM usable, fast, benchmarked, and maintainable.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar.
Hands-on experience optimizing AMD GPU workloads using ROCm, HIP, Triton, CK, AITER, or similar AMD ecosystem tools.
Deep understanding of AMD GPU execution, memory behavior, toolchains, kernel performance, and backend-specific performance constraints.
Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or communication-heavy runtime paths.
Strong performance profiling and benchmarking skills, with the ability to use measurements, hardware counters, correctness tests, and reproducible benchmarks to guide optimization work.
Preferred qualifications:
Experience with vLLM, SGLang, TensorRT-LLM, ROCm-based serving, or other LLM inference systems.
Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems.
Experience with compiler and kernel technologies such as Triton, MLIR, LLVM, CK, AITER, HIP, or other kernel DSLs and backend libraries.
Knowledge of quantization methods such as INT8, FP8, mixed precision, or AMD hardware-specific numeric formats, including accuracy and performance tradeoffs.
Bonus points if you have:
Contributed to vLLM, ROCm, HIP, Triton, CK, AITER, PyTorch, compiler projects, or other open-source ML infrastructure.
Built AMD GPU benchmarking infrastructure or automated performance regression detection for accelerator workloads.
Worked directly with AMD, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Inferact is a startup founded by creators and core maintainers of vLLM, the most popular open-source LLM inference engine.
Our mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.
Offices: San Francisco, CA, US
Information TechnologySoftwareArtificial Intelligence (AI)
How much do Engineering jobs in San Francisco, CA pay?
Based on 2907 listings with disclosed salaries, most engineering jobs in San Francisco, CA pay between $130k–$286k per year. Individual offers vary with seniority, company size, and specialization.
How many Engineering jobs are open in San Francisco, CA right now?
There are currently 3,788 open engineering positions in San Francisco, CA listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Engineering roles in San Francisco, CA?
Companies currently hiring include OpenAI, Anthropic, San Francisco Department of Public Health, Crusoe, Pinterest, among others. Browse the listings above to see every active employer.
Are there remote or hybrid Engineering jobs in San Francisco, CA?
Yes — 1850 of the 3788 open engineering positions offer remote or hybrid work (371 remote, 1479 hybrid).
How do I apply for Engineering jobs in San Francisco, CA?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.