Jobs at Inferact (Now Hiring) — 14 open

Founding Product Designer

San Francisco, California, United States · Remote OK

Mid levelVisa sponsorship$150M raised

Overview Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the inters…

Skills: Product design, UI/UX design, Visual identity, Brand systems, Figma

Head of Engineering

San Francisco, California, United States · On-site

Senior+Visa sponsorship$150M raised

Overview Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the inters…

Skills: LLM inference, ML systems, GPU optimization, Distributed systems, Hardware-software co-design

Product Marketing Manager

San Francisco, California, United States · Remote OK

Mid levelVisa sponsorship$150M raised

Overview Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the inters…

Skills: Product Marketing, Go-To-Market Strategy, Event Management, Positioning And Messaging, Partner Marketing

Member of Technical Staff, CI/CD Infrastructure

San Francisco, California, United States · Remote OK

$200k–$400k/yr

SeniorVisa sponsorship$150M raised

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference efficient and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection …

Skills: Docker, Kubernetes, GitHub Actions, Buildkite, Python

Member of Technical Staff, TPU Performance Engineering

Singapore, Singapore · On-site

$200k–$400k/yr

SeniorVisa sponsorship$150M raised

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of…

Skills: TPU Performance Engineering, JAX, XLA, Pallas, ML Kernel Optimization

Member of Technical Staff, AMD GPU Performance Engineering

San Francisco, California, United States · Remote OK

$200k–$400k/yr

SeniorVisa sponsorship$150M raised

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of…

Skills: ROCm, HIP, Triton, CK, AITER

Member of Technical Staff, AMD GPU Performance Engineering

Singapore, Singapore · On-site

$200k–$400k/yr

SeniorVisa sponsorship$150M raised

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of…

Skills: AMD GPU Optimization, TPU Performance Engineering, ROCm, HIP, Triton

Member of Technical Staff, Inference

Singapore, Singapore · On-site

$200k–$400k/yr

SeniorVisa sponsorship$150M raised

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of…

Skills: Python, PyTorch, vLLM, TensorRT-LLM, SGLang

Member of Technical Staff, Cloud Orchestration

Singapore, Singapore · On-site

$200k–$400k/yr

SeniorVisa sponsorship$150M raised

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of…

Skills: Kubernetes, Container Orchestration, Kubernetes Operators, Python, Rust

Member of Technical Staff, Performance and Scale

Singapore, Singapore · On-site

$200k–$400k/yr

SeniorVisa sponsorship$150M raised

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of…

Skills: Rust, Go, C++, Distributed Systems, Network Protocols

Member of Technical Staff, Kernel Engineering

Singapore, Singapore · On-site

$200k–$400k/yr

SeniorVisa sponsorship$150M raised

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of…

Skills: CUDA, C++, Python, GPU Architecture, Kernel Optimization

Member of Technical Staff, Inference

San Francisco, California, United States · Remote OK

$200k–$400k/yr

SeniorVisa sponsorship$150M raised

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of…

Skills: Python, PyTorch, vLLM, TensorRT-LLM, SGLang

Member of Technical Staff, TPU Performance Engineering

San Francisco, California, United States · Remote OK

$200k–$400k/yr

SeniorVisa sponsorship$150M raised

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of…

Skills: AMD GPU Optimization, TPU Performance Engineering, ROCm, HIP, Triton

Member of Technical Staff, Developer Relations

San Francisco, California, United States · Remote OK

$200k–$400k/yr

SeniorVisa sponsorship$150M raised

Overview Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the inters…

Skills: LLM Inference Systems, Developer Relations, GPU Serving, Technical Writing, Model Serving

About Inferact

Building the future of inference

Industry
Software Development
Company size
11-50 employees
Founded
2025
Headquarters
San Francisco, CA
LinkedIn followers
3,730
Total funding
$150M

Inferact is a startup founded by creators and core maintainers of vLLM, the most popular open-source LLM inference engine. Our mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.

Offices: San Francisco, CA, US

Information TechnologySoftwareArtificial Intelligence (AI)AI Infrastructure
View all jobs at Inferact