San Francisco, California, United States · Remote OK
$200k–$400k/yr
SeniorVisa sponsorship$150M raised
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of…
Construction is the 2nd largest industry in the world (4x the size of SaaS!). But unlike software (with observability platforms such as AppDynamics and Datadog), construction teams lack automated feedback loops to help p…
San Francisco, California, United States · Remote OK
Senior
Site Reliability Engineer Platform and software · shared across customers Reports to: Director, Site Reliability Location: Remote (US) Department: Cloud Platform Engineering / SRE/Reliability Position summary The Site Re…
San Francisco, California, United States · Remote OK
$195k–$258k/yr
Senior+$2.7B raised
Circle (NYSE: CRCL) is one of the world’s leading internet financial platform companies, building the foundation of a more open, global economy through digital assets, payment applications, and programmable blockchain in…
San Francisco, California, United States · Remote OK
$99k–$150k/yr
Mid level$773M raised
Thumbtack helps millions of people confidently care for their homes. Thumbtack is the one app you need to take care of and improve your home — from personalized guidance to AI tools and a best-in-class hiring experience.…
Skills: Salesforce Administration, AI Agent Development, Workflow Automation, API Integration, Middleware Management
San Francisco, California, United States · On-site
$185k–$210k/yr
Senior+$57M raised
Company Credit Genie is a mobile-first financial wellness platform designed to help individuals take control of their financial future. We leverage artificial intelligence to provide personalized insights and are buildin…
San Francisco, California, United States · On-site
Mid level$50M raised
About Us At Rox, we believe in empowering people to do their best work. Our platform supercharges sellers with autonomous revenue agents to do the manual work so they can focus on what they do best: selling. Just as codi…
Skills: SQL, Python, Dbt, Data Modeling, Dashboarding
San Francisco, California, United States · On-site
Senior
About Engram Today’s AI is a brilliant stranger: it can solve the world’s hardest math problems, but it knows next to nothing about you and your work. It rereads your files to answer even basic questions, burns an enormo…
San Francisco, California, United States · On-site
$285k–$335k/yr
Senior+$244M raised
About the Company: Tools for Humanity (TFH) designs and builds technology behind World. World is building a real human network designed to accelerate people in the age of AI. As bots and autonomous agents reshape the int…
Skills: Product Management, Stakeholder Management, Web3, Digital Identity, Privacy Standards
Tabula Bio
Research Scientist
San Francisco, California, United States · On-site
Senior
About Tabula Tabula is building an AI-first therapeutics company. We are starting with bacteriophages, natural predators of bacteria, and building the models and experimental systems needed to design better therapies for…
About Distyl AI Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in…
Skills: AI System Architecture, Technical Leadership, Solutions Architecture, Cloud Systems, API Design
About Us At Hayden AI, we are on a mission to harness the power of computer vision to transform the way transit systems and other government agencies address real-world challenges. From bus lane and bus stop enforcement …
San Francisco, California, United States · On-site
Senior
We are so glad you are interested in joining Sutter Health! Organization: CPMC-California Pacific Med Center Position Overview: This position replacing r-130399 and now has PD added to job code Job Description: Performan…
Our Mission Rebuild how the world works, to make institutions work better for the people they serve. About Brain Co. Brain Co. builds AI-native operating systems for large, regulated institutions. Each system is built fo…
Why Harvey At Harvey, we’re transforming how legal and professional services operate. By combining frontier agentic AI, an enterprise-grade platform, and deep domain expertise, we’re reshaping how critical knowledge work…
Staff Machine Learning Engineer - Computer Vision & Multi-Modal AI
San Francisco, California, United States · On-site
Senior$2.9B raised
The opportunity We are building the next generation of AI-driven game experiences — generative world models, neural rendering, and multi-modal understanding that turn images, text, and 3D primitives into interactive worl…
About Us Beast Industries is a multifaceted media and entertainment company founded by Jimmy Donaldson, popularly known as MrBeast, the most watched person in the world. Renowned for revolutionizing digital content creat…
Skills: Data Modeling, ETL/ELT, Distributed Data Processing, Batch Processing, Streaming Pipelines
About Us Beast Industries is a multifaceted media and entertainment company founded by Jimmy Donaldson, popularly known as MrBeast, the most watched person in the world. Renowned for revolutionizing digital content creat…
Skills: Machine Learning Engineering, Python, MLOps, Distributed Systems, Data Pipelines
About Us Beast Industries is a multifaceted media and entertainment company founded by Jimmy Donaldson, popularly known as MrBeast, the most watched person in the world. Renowned for revolutionizing digital content creat…
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
$200k–$400k/yr
Full-time
bachelor degree
Health Insurance, Dental Insurance, Vision Insurance, 401(k) Company Match, Equity
Visa sponsorship available
Posted 48d ago
~40 hrs/week
Remote in United States
Responsibilities
Optimize the vLLM inference engine to improve the speed and cost of running LLMs and diffusion models. Develop innovations for diverse hardware and architectures, including mixture-of-experts and multimodal models.
Requirements
Requires a bachelor's degree in computer science or equivalent with deep expertise in transformer architectures and PyTorch internals. Candidates must have experience with LLM inference systems and the ability to implement research papers into performant code.
Full job description
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.
About the Role
We're looking for an inference runtime engineer to push the boundaries of what's possible in LLM and diffusion model serving. Models grow larger. Architectures shift: mixture-of-experts, multimodal, agentic. Every breakthrough demands innovations on the inference engine itself. You'll work at the core of vLLM, optimizing how models execute across diverse hardware and architectures. Your work will directly impact how the world runs AI inference.
Skills and Qualifications
Minimum qualifications:
Bachelor's degree or equivalent experience in computer science, engineering, or similar.
Deep understanding of transformer architectures and their variants.
Strong programming skills in Python with experience in PyTorch internals.
Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).
Ability to read and implement model architectures and inference techniques from research papers.
Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.
Preferred qualifications:
Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving.
Familiarity with RL frameworks and algorithms for LLMs.
Experience with multimodal inference (audio/image/video/text).
Contributions to open-source ML or system infrastructure projects.
Bonus points if you have:
Implemented core features in vLLM or other inference engine projects.
Contributed to vLLM integrations (verl, OpenRLHF, Unsloth, LlamaFactory, etc).
Written widely-shared technical blogs or side projects on vLLM or LLM inference.
Logistics
Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
Visa sponsorship: We sponsor visas on a case-by-case basis.
Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.
Inferact is a startup founded by creators and core maintainers of vLLM, the most popular open-source LLM inference engine.
Our mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster.
Offices: San Francisco, CA, US
Information TechnologySoftwareArtificial Intelligence (AI)
How much do Data & Analytics jobs in San Francisco, CA pay?
Based on 2332 listings with disclosed salaries, most data & analytics jobs in San Francisco, CA pay between $125k–$278k per year. Individual offers vary with seniority, company size, and specialization.
How many Data & Analytics jobs are open in San Francisco, CA right now?
There are currently 2,910 open data & analytics positions in San Francisco, CA listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Data & Analytics roles in San Francisco, CA?
Companies currently hiring include OpenAI, UCSF Department of Anesthesia and Perioperative Care, Salesforce, Anthropic, Pinterest, among others. Browse the listings above to see every active employer.
Are there remote or hybrid Data & Analytics jobs in San Francisco, CA?
Yes — 1794 of the 2910 open data & analytics positions offer remote or hybrid work (381 remote, 1413 hybrid).
How do I apply for Data & Analytics jobs in San Francisco, CA?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.