The Video Computer Vision (VCV) organization is an applied research and engineering team developing real-time, on-device Computer Vision and Machine Perception technologies across Apple products. Within VCV, our team bui…
Skills: Multimodal Reasoning, Computer Vision, Machine Learning, Deep Learning, LLM
Minimum qualifications: Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, or a related field, or equivalent practical experience. 8 years of experience in ASIC design. Experience with S…
Staff Software Engineer, Data Center Network Switches
Sunnyvale, California, United States · On-site
$207k–$300k/yr
Senior+$26M raised
Minimum qualifications: Bachelor's degree or equivalent practical experience. 8 years of experience programming in C++. 5 years of experience testing, and launching software products. 3 years of experience with software …
Skills: C++, Software design, Network architecture, Ethernet switching, Data structures
Apple Cloud Networking team builds and operates large-scale, software-defined networking platforms that enable secure, resilient, and highly available multi-cloud connectivity with a global footprint. Our infrastructure …
Skills: Software engineering, Systems engineering, Infrastructure engineering, Technical leadership, People management
Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind…
The Role We are looking for a Research Scientist to join the Multi-Embodiment Generalist Agent (MEGA) team within Wayve Science as a founding member. MEGA is building foundation models for general-purpose robots beyond n…
Skills: Machine learning, Vision-language models, Video models, Robot policies, Foundation models
Principal Research Scientist, Robot Foundation Model
Sunnyvale, California, United States · Hybrid
$407k–$512k/yr
Senior$2.5B raised
About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing …
Skills: Machine learning, Vision-language models, Video models, Robot policies, Foundation models
San Francisco, California, United States · On-site
$127k–$184k/yr
Senior$26M raised
Minimum qualifications: Bachelor's degree or equivalent practical experience. 6 years of experience with cloud native architecture in a customer-facing or support role. Experience with cloud engineering, on-premise engin…
Skills: Cloud Architecture, Infrastructure Modernization, Application Modernization, Data Analytics, Cloud AI
Staff Software Engineer, Internet Egress Traffic Engineering
Sunnyvale, California, United States · On-site
$207k–$300k/yr
Senior+$26M raised
Minimum qualifications: Bachelor's degree or equivalent practical experience. 8 years of experience programming in C++. 5 years of experience building and developing large-scale infrastructure, distributed systems and ne…
About the role: We’re looking for an experienced GPU Driver Developer. You will focus the design, development, implementation, and toolchain of a high-performance host-side driver stack and API for our proprietary RISC-V…
About the role: We’re looking for a highly experienced GPU Standards Developer. You will focus on the design, development, implementation, and toolchain of a high-performance host-side driver stack layering industry stan…
Apple's Cellular Software team is seeking talented, highly motivated and disciplined engineers to work across layers on groundbreaking cellular technologies. The position involves identifying and/or developing core cellu…
Your Impact Join a small team building the next generation of cybersecurity products from the ground up. Led by industry veterans with a proven track record of success - you will get to architect, build, and deliver huge…
Your Impact Join a small team building the next generation of cybersecurity products from the ground up. Led by industry veterans with a proven track record of success - you will get to architect, build, and deliver huge…
Skills: Artificial Intelligence, Cybersecurity, Machine Learning, Deep Learning, Large Language Models
Job Description At General Motors, our product teams are redefining mobility. Through a human-centered design process, we create vehicles and experiences that are designed not just to be seen, but to be felt. We’re turni…
Are you ready to be at the forefront of Agentic AI innovation and redefine the future of communication? Join our team as a Sr TPM and lead futuristic initiatives that will shape the next generation of intelligent, conver…
Skills: Technical program management, Software development, Generative AI, Agentic AI, Real-time AI
At Coram AI, we’re reimagining video security for the modern world. Our cloud-native platform uses computer vision and AI to help businesses stay safe, make smarter decisions, and move faster; from real-time alerts to se…
Skills: Sales, Prospecting, Pipeline management, Product demonstration, Communication
Meta is seeking highly skilled Design Engineers to join our team. In this role, you will contribute to the development of advanced technology solutions, including machine learning and network acceleration. You will colla…
Position Summary...As the Director of Product Management for Seller Risk, you will define the vision, strategy, and execution for the products that establish trust across the Walmart Marketplace seller lifecycle—from onb…
Position Summary... What you'll do...Role summary: As a Senior Software Engineer at Walmart, you will lead the delivery of scalable software solutions by managing feature implementation, testing, and ongoing support with…
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Full-time
bachelor degree, postgraduate degree
Posted 38d ago
~40 hrs/week
Responsibilities
The Applied Researcher will design and optimize compact vision-language models for on-device reasoning under strict compute and memory constraints. They will collaborate with hardware and software teams to integrate these models into the Apple ecosystem, ensuring efficient performance for real-time multimodal tasks.
Requirements
Candidates must hold at least an MS in Computer Science, AI, or a related field with a strong foundation in deep learning and multimodal model training. Proficiency in Python, PyTorch, and experience with model compression or inference optimization in resource-constrained environments is required.
Full job description
The Video Computer Vision (VCV) organization is an applied research and engineering team developing real-time, on-device Computer Vision and Machine Perception technologies across Apple products. Within VCV, our team builds next-generation multimodal AI systems that combine on-device multimodal encoders, large language models, and foundation models to create intelligent systems capable of understanding, reasoning, and acting across language, vision, audio, and tools. Our work is deeply integrated into the Apple ecosystem, partnering across hardware, software, and ML teams to deliver real-time, scalable, and privacy-preserving experiences reaching millions of users.
Description
We are seeking an Applied Researcher with deep expertise in multimodal reasoning at small model scale — making vision-language models in the smaller regime (under ~10B parameters, down to sub-1B) think, plan, and act reliably under strict compute, memory, and latency constraints. In this role you will own the reasoning side of the on-device multimodal stack: designing compact VLMs that reason over images, video, and 3D scene content; compressing and distilling the reasoning process itself; and engineering the decoding and inference path that makes multi-step reasoning affordable on an Apple device. This role offers the unique opportunity to define what on-device intelligence looks like for hundreds of millions of users. You'll push the boundaries of what small models can achieve — enabling real-time multimodal understanding and multi-step reasoning without reliance on cloud connectivity. You'll collaborate with hardware teams, compiler engineers, and ML researchers to unlock capabilities that few organizations can deliver at Apple's scale and quality bar.
This role spans multiple dimensions of efficient on-device reasoning — including VLM architecture and connector design, reasoning post-training (SFT/RL), chain-of-thought compression, speculative and structured decoding, visual token reduction, quantization and distillation, and hardware-aware inference optimization.
A core focus of this role is efficient reasoning: compressed and latent chain-of-thought, reasoning distillation from frontier teachers, adaptive test-time compute (knowing when — and how long — to think), speculative and structured decoding, KV-cache compression, and visual-token efficiency. A second focus is reasoning over real-time visual perception experts. Rather than consuming pixels alone, the VLM should be able to invoke and reason over the outputs of specialist on-device vision models — feed-forward 3D scene and geometry estimators (VGGT-style reconstruction, depth, camera pose), human body and hand mesh/pose recovery, object detectors, localizers, and trackers — and fuse those structured, metric outputs into its reasoning about the scene. This raises real research questions: how to represent geometry, body parameters, and detections compactly in a token-budgeted context; how to schedule which experts run at which frame rate within a real-time budget; and how to train a small model to invoke, trust, and cross-check them. Efficiency is treated as a first-class metric here: reasoning quality is measured at a fixed latency, memory, and power budget.
Minimum Qualifications
MS in Computer Science, Machine Learning, AI, Computer Vision, or a related field (or equivalent practical experience)
Strong foundation in deep learning, with specific experience in LLM or VLM training, post-training, or inference optimization
Demonstrated experience working with multimodal models (vision-language models, multimodal LLMs) in resource-constrained environments, including hands-on work with reasoning quality, decoding, or model compression
Proficiency in Python and modern deep learning frameworks (PyTorch preferred), with familiarity with inference and optimization toolchains (quantization, distillation, pruning — e.g., vLLM/SGLang, llama.cpp, MLX, CoreML)
Preferred Qualifications
PhD with research in efficient multimodal reasoning, LLM reasoning, model compression, efficient inference/decoding, or lightweight VLM architectures
Experience training or post-training vision-language models end-to-end — connector/projector design, visual instruction tuning, resolution and token-budget trade-offs, small-model recipes
Hands-on experience with reasoning techniques: chain-of-thought distillation and compression, latent/implicit reasoning, reward-guided decoding, RL for reasoning (GRPO, RLVR, STaR), or test-time compute allocation
Expertise in decoding and serving optimizations: speculative decoding, structured/grammar-constrained generation, KV-cache quantization and eviction, continuous batching, long-context inference
Experience combining LLMs with real-time perception models — 3D reconstruction and geometry (VGGT, DUSt3R/MASt3R-style, SLAM, monocular depth), human pose and body/hand mesh recovery (SMPL-family), detection, segmentation, or tracking — and with spatial or 3D-grounded reasoning and embodied/spatial VQA
Experience deploying LLM or multimodal models on mobile or edge hardware (CoreML, MLX, TensorRT-LLM, or equivalent), with attention to ANE/GPU kernel and memory constraints
Experience with quantization-aware training, mixed-precision inference, and knowledge distillation for vision-language models
Familiarity with efficient vision encoders and self-supervised/joint-embedding pretraining (V-JEPA, I-JEPA, MAE, DINO/DINOv2, SigLIP, CLIP), Mamba/SSM vision backbones, or streaming architectures for real-time video with fixed memory budgets
Interest in Video-LLMs, long-video reasoning, and world models for prediction and planning
Publication record in top-tier venues is a plus (NeurIPS, ICML, ICLR, CVPR, ECCV, ACL, MLSys, ICRA, etc.)
We’re a diverse collective of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. And the same innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it. This is where your work can make a difference in people’s lives. Including your own.
Apple is an equal opportunity employer that is committed to inclusion and diversity. Visit apple.com/careers to learn more.
Offices: 1 Apple Park Way, Cupertino, California 95014, US
Based on 1493 listings with disclosed salaries, most software jobs in Sunnyvale, CA pay between $140k–$286k per year. Individual offers vary with seniority, company size, and specialization.
How many Software jobs are open in Sunnyvale, CA right now?
There are currently 1,773 open software positions in Sunnyvale, CA listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Software roles in Sunnyvale, CA?
Companies currently hiring include Google, Walmart, Apple, Wayve, Applied Intuition, among others. Browse the listings above to see every active employer.
Are there remote or hybrid Software jobs in Sunnyvale, CA?
Yes — 569 of the 1773 open software positions offer remote or hybrid work (42 remote, 527 hybrid).
How do I apply for Software jobs in Sunnyvale, CA?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.