San Francisco, California, United States · Remote OK
Senior
Locations: San Francisco or Remote About The Role The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient…
San Francisco, California, United States · Remote OK
Senior
About Near AI Near AI is building the world's best user-owned open-source AI assistant – and the models, agent network, and privacy-preserving infrastructure powering it. Our mission centers on democratizing access to AI…
Skills: Python, TypeScript, Rust, Fullstack Development, AI Systems
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Full-time
Posted 22d ago
~40 hrs/week
Remote in San Francisco, California, United States
Responsibilities
Architect and maintain production high-traffic LLM serving systems. Focus on optimizing throughput, latency, and cost for leading open-source LLMs.
Requirements
Requires deep expertise in GPU architectures and hands-on experience optimizing inference engines like vLLM or TensorRT. A proven track record of designing end-to-end high-traffic serving systems is essential.
Full job description
Locations: San Francisco or Remote
About The Role
The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient infrastructure for open-source AI at a global scale.
We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served.
What You'll Be Doing
Architect and maintain production high-traffic LLM serving systems.
Optimize throughput, latency, and cost for leading open-source LLMs.
What We're Looking For
Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.
Deep knowledge of state-of-the-art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc.
Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems.
Strong problem-solving skills and ability to communicate technical ideas clearly.
We'd Love If You Have
Experience with Trusted Execution Environments (TEE).
Active contributor to open-source LLM inference engines.
Please let us know if you require any special requirements for your interview and we'll do our best to accommodate.
NEAR AI is an artificial intelligence research, engineering, and product development company committed to building an AI future owned by everyone. Founded by AI pioneer and former Google Deepmind researcher Illia Polosukhin, NEAR AI’s verifiable private inference infrastructure empowers developers and enterprises to deploy AI models with full control over their data.
With hardware-backed private inference via a simple API, NEAR AI Cloud runs sensitive AI workloads securely and at scale, from privacy-critical consumer interactions to autonomous systems and critical infrastructure. NEAR AI Private Chat brings the same guarantees to users’ everyday questions and research. Serving over 100 million users across platforms such as Brave Nightly and OpenMind, NEAR AI is proven infrastructure for transforming sensitive data into safe intelligence and advancing a user-owned AI future. Learn more at https://near.ai/.
Offices: 535 Mission St, San Francisco, California, US
Machine LearningArtificial Intelligenceand Natural Language Processing
NEAR AI is an artificial intelligence research, engineering, and product development company committed to building an AI future owned by everyone. Founded by AI pioneer and former Google Deepmind researcher Illia Polosukhin, NEAR AI’s verifiable private inference infrastructure empowers developers and enterprises to deploy AI models with full control over their data.
With hardware-backed private inference via a simple API, NEAR AI Cloud runs sensitive AI workloads securely and at scale, from privacy-critical consumer interactions to autonomous systems and critical infrastructure. NEAR AI Private Chat brings the same guarantees to users’ everyday questions and research. Serving over 100 million users across platforms such as Brave Nightly and OpenMind, NEAR AI is proven infrastructure for transforming sensitive data into safe intelligence and advancing a user-owned AI future. Learn more at https://near.ai/.
Offices: 535 Mission St, San Francisco, California, US
Machine LearningArtificial Intelligenceand Natural Language Processing