About the Role We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inferen…
Amazon Search is building a first-of-its-kind AI-powered visual search experience that lets customers describe products they're imagining, instantly see AI-generated images in response, and tap those images to search for…
Skills: Generative AI, Multimodal Retrieval, Diffusion Models, Large Language Models, Text-to-Image Generation
About Rivian Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. As a company, we con…
About Rivian Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. As a company, we con…
About Subsense Subsense is a deep-tech company developing the world’s first non-surgical, bidirectional brain-computer interface powered by plasmonic and magnetoelectric nanoparticles. Our mission is to unlock direct com…
Skills: Python, Instrument Control, PyVISA/SCPI, PyQt, Data Acquisition
Who: You! And the rest of the IT department & their cross-functional partners across Sales, Marketing, CX, Finance, and Engineering What: An Enterprise Applications Staff Architect position responsible for defining the o…
About Nclusion Nclusion is on a mission to provide traditional financial services to 1.4 billion people worldwide without access today. Without a secure way to save, invest, or transfer money, individuals are not empower…
Skills: Finance Transformation, Systems Implementation, Process Optimization, Cross-Functional Leadership, NetSuite
Product Communications Lead, Core Print Solutions Description - Reporting to the Global Head of Print Communications, this role leads communications for HP’s Core Print business, spanning Office, Consumer, and Supplies. …
ABOUT ALLOCATE Allocate is building the intelligent private markets operating system for the wealth channel. We give RIAs, family offices, and institutional allocators modern infrastructure for discovering, accessing, an…
Skills: Product Management, AI Workflow Automation, Internal Tools Development, Operational Platforms, GTM Systems
About the Role The seasonal Operations Technical Specialist supports the operational readiness of H&R Block’s seasonal tax office network. This role ensures offices are fully equipped, functional, and aligned with brand …
Director, Applied Science, Alexa for Shopping (Rufus)
Palo Alto, California, United States · On-site
$263k–$350k/yr
Senior+$35B raised
Alexa for Shopping (Rufus) is Amazon's new AI-powered shopping assistant that combines the capabilities of Rufus and Alexa+ to provide a more personalized and intelligent shopping experience. We are building the future o…
Skills: Large Language Models, Multi-agent Systems, Reinforcement Learning, RLHF, DPO
Aptos is a people-first blockchain on a mission to help billions of people achieve universal and fair access to decentralized assets in a safe and scalable way. Founded by some of the original creators and maintainers th…
About Rivian Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. As a company, we con…
Skills: Autonomous vehicle systems, Data engineering, Machine learning, Triage infrastructure, Root cause analysis
About Rivian Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. As a company, we con…
Skills: Quantized deep learning, Hardware acceleration, Autonomous systems, Perception model design, Embedded compute platforms
Model AI
Founding Machine Learning Infrastructure Engineer
Palo Alto, California, United States · On-site
Senior
Founding Machine Learning Infrastructure Engineer Location: Onsite in Palo Alto Compensation: Competitive Salary + Equity About Model AI Model AI is building the infrastructure and application stack for the next generati…
Skills: ML Systems, Distributed Systems, High-Performance Computing, LLM Inference, CUDA
Model AI
Founding Agent Harness Engineer
Palo Alto, California, United States · On-site
Mid level
Founding Agent Systems Engineer Location: Onsite in Palo Alto Compensation: Competitive Salary + Equity About Model AI Model AI is building the infrastructure and application stack for the next generation of agentic AI s…
Skills: Python, Systems Engineering, LLM Agents, Evaluation Harnesses, Developer Tools
Enterprise Account Executive Enterprise Sales • Remote (US) • Full-Time About Neo.Tax Enterprises waste millions of dollars trying to calculate and substantiate R&D tax credits and capitalization of software costs. Neo.T…
Skills: Enterprise SaaS Sales, New Business Development, Pipeline Management, Executive Presence, Strategic Account Planning
About Rivian Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. As a company, we con…
About Rivian Rivian is on a mission to keep the world adventurous forever. This goes for the emissions-free Electric Adventure Vehicles we build, and the curious, courageous souls we seek to attract. As a company, we con…
Skills: Camera systems, Optics, Image sensor, Image quality, Image signal processor
WindBorne Systems is supercharging weather forecasts with a unique proprietary data source: a global constellation of next-generation smart weather balloons targeting the most critical atmospheric data. We design, manufa…
Skills: Ruby on Rails, Postgres, Full-stack Development, Infrastructure Architecture, Data Pipeline Management
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
$185k–$300k/yr
Full-time
Competitive salary, Equity, Comprehensive health benefits, Monthly stipends, Company retreats
Posted 82d ago
~40 hrs/week
Responsibilities
Lead the implementation of advanced inference acceleration and GPU parallelism strategies to optimize Pika's AI-driven video and language models. Collaborate with research and engineering teams to deploy high-performance computing kernels and scalable production pipelines.
Requirements
Requires 5+ years of engineering experience with deep expertise in CUDA, NCCL, and distributed inference techniques like TP, SP, and PP. Candidates should have a proven track record in model quantization and familiarity with video generation or LLMs.
Full job description
About the Role
We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.
You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.
What You’ll Do
Accelerate Inference: Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.
Maximize GPU Parallelism: Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.
Programming for Performance: Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.
Advance AI Deployment: Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.
Improve Training Efficiency: (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.
Technical Excellence: Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.
What We’re Looking For
Experience: 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale.
Inference Mastery: Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.
GPU & Parallelism: Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.
AI Domain Knowledge: Familiarity with video generation (videogen) models and large language models (LLMs).
Collaboration: Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions.
Ownership Mindset: Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment.
Bonus: Experience in enhancing training efficiency, stability, or resource optimization for large models.
Nice to Have
Experience with high-throughput video or real-time streaming model deployment
Familiarity with distributed training and optimization toolkits
Contributions to open source projects in AI infrastructure or deep learning compilers
Startup or rapid prototyping experience
What We Offer
Competitive salary in the AI industry
Equity in a fast-growing startup shaping the future of AI
Comprehensive health benefits, monthly stipends, company retreats
A supportive and collaborative office culture—we’re all building and launching together
About Pika
At Pika, we're crafting a future where video creation is seamless, intuitive, and universally accessible. Our mission is to empower creativity by breaking down technical barriers using the transformative power of AI. We’re a tight-knit, energetic team based in Palo Alto, CA, valuing efficiency, curiosity, and the ambition to make a meaningful impact on the world.
We work from our Palo Alto office 3–5 days a week and welcome applicants who are eager to contribute onsite.
Based on 985 listings with disclosed salaries, most technology jobs in Palo Alto, CA pay between $110k–$262k per year. Individual offers vary with seniority, company size, and specialization.
How many Technology jobs are open in Palo Alto, CA right now?
There are currently 1,289 open technology positions in Palo Alto, CA listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Technology roles in Palo Alto, CA?
Companies currently hiring include Rivian, SpaceX, Ford Motor Company, Amazon, Rivian and Volkswagen Group Technologies, among others. Browse the listings above to see every active employer.
Are there remote or hybrid Technology jobs in Palo Alto, CA?
Yes — 583 of the 1289 open technology positions offer remote or hybrid work (112 remote, 471 hybrid).
How do I apply for Technology jobs in Palo Alto, CA?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.