Your Impact at LILA The Staff/Principal DevOps Engineer - AI Inference will drive the design, implementation, and optimization of infrastructure purpose-built for serving machine learning models at scale. This role bridg…
Your Impact at LILA Lila Sciences is hiring a Software Engineer to develop the next generation of Lab Instrument Integration Software, which is a foundational component of our AI-enabled laboratory. We are looking for se…
Your Impact at LILA Lila Sciences is seeking a highly skilled, hands-on Automated Systems Engineer to design, develop, and implement automation solutions across a broad range of laboratory workflows. Reporting to the Aut…
Skills: Automated workflows, Python, API integration, CAD design, Rapid prototyping
Overview: Draper is an independent, nonprofit research and development company headquartered in Cambridge, MA. The 2,000+ employees of Draper tackle important national challenges with a promise of delivering successful a…
Overview: Draper is an independent, nonprofit research and development company headquartered in Cambridge, MA. The 2,000+ employees of Draper tackle important national challenges with a promise of delivering successful a…
Skills: UVM, System Verilog, Digital verification, FPGA, ASIC
Overview: Draper is an independent, nonprofit research and development company headquartered in Cambridge, MA. The 2,000+ employees of Draper tackle important national challenges with a promise of delivering successful a…
Skills: Integrated circuit design, Computer architecture, System Verilog, Verilog, VHDL
ROLE SUMMARY: The technical engineering counterpart to AIDE’s applied workflow roles, embedded within the Research Unit to convert promising AI workflow concepts into durable, evaluated, and supportable systems that acce…
ROLE SUMMARY: Drives the practical application of large language models (LLMs) and agentic artificial intelligence (AI) across the Inflammation & Immunology (I&I) portfolio by partnering closely with clinical and scienti…
Skills: Large language models, Agentic AI, Clinical development, Workflow automation, Clinical data review
Overview: Draper is an independent, nonprofit research and development company headquartered in Cambridge, MA. The 2,000+ employees of Draper tackle important national challenges with a promise of delivering successful a…
Skills: Cadence EDA, Siemens EDA, Gitlab, CI/CD, SOS
Please Note: To provide the best candidate experience amidst our high application volumes, each candidate is limited to 10 applications across all open jobs within a 6-month period. Advancing the World’s Technology Toget…
Cambridge, Massachusetts, United States · Remote OK
$139k–$251k/yr
Senior+$5.2B raised
Would you like to use your hands on technical expertise to help influence Akamai's Cloud evolution? Would you like to work cross-functionally with product management and sales engineers? Join our Competitive Intelligence…
Cambridge, Massachusetts, United States · Remote OK
$95k–$171k/yr
Mid level$5.2B raised
Do you have a passion for cutting edge technologies and tackling system problems? Are you a self-starting professional who thrives in a dynamic environment? Join our Site Reliability team Our Team builds and delivers hig…
Skills: Site Reliability Engineering, DevOps, Python, Go, Shell
Cambridge, Massachusetts, United States · Remote Solely
Mid level
Company Description By working at Harvard University, you join a vibrant community that advances Harvard's world-changing mission in meaningful ways, inspires innovation and collaboration, and builds skills and expertise…
Skills: Python, Test automation, Playwright, Selenium, Pytest
The Customer Success group is the relationship management partner for our clients during the course of their partnership with CMT. As a member of Customer Success in the Sales organization, you will partner closely with …
Software Engineer III, Infrastructure, Google Cloud Networking
Cambridge, Massachusetts, United States · On-site
$147k–$210k/yr
Mid level$26M raised
Minimum qualifications: Bachelor’s degree or equivalent practical experience. 2 years of experience with software development with C++, Python, or Go programming language. 2 years of experience with data structures and a…
Skills: C++, Python, Go, Data structures, Algorithms
Co-op – Design Release Engineer - Cambridge, MA - Jan - Aug 2027
Cambridge, Massachusetts, United States · Hybrid
$25/hr–$45/hr
Entry level
Job TitleCo-op – Design Release Engineer - Cambridge, MA - Jan - Aug 2027 Job Description Co-op – Design Release Engineer - Cambridge, MA - June- December 2026 Are you interested in an Internship opportunity with Philips…
Skills: Software design, Design control, Process improvement, Continuous value delivery, Cross-functional collaboration
Overview: Draper is an independent, nonprofit research and development company headquartered in Cambridge, MA. The 2,000+ employees of Draper tackle important national challenges with a promise of delivering successful a…
Co-op – Software Engineering (APM) - Cambridge, MA - Jan - Aug 2027
Cambridge, Massachusetts, United States · Hybrid
$25/hr–$45/hr
Entry level
Job TitleCo-op – Software Engineering (APM) - Cambridge, MA - Jan - Aug 2027 Job Description Are you interested in an Internship opportunity with Philips? We welcome individuals who are currently pursuing an undergraduat…
Co-op - Software Development Engineer - Cambridge, MA - Jan-Aug 2027
Cambridge, Massachusetts, United States · On-site
$26/hr–$45/hr
Entry level
Job TitleCo-op - Software Development Engineer - Cambridge, MA - Jan-Aug 2027 Job Description Are you interested in an co-op opportunity with Philips? We welcome individuals who are currently pursuing an undergraduate (B…
Job Description SummaryAs a Senior Staff AI Engineer, you’ll lead enterprise full stack and AI platform architecture at scale for multiple initiatives across the company. You’ll design and ship microservices and REST API…
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
$192k–$272k/yr
Full-time
Medical coverage, Dental coverage, Vision coverage, Life insurance, Disability insurance, Flexible time off
Posted 9d ago
~40 hrs/week
Responsibilities
You will design and optimize infrastructure for serving machine learning models at scale, focusing on low-latency and high-throughput inference. This involves managing GPU clusters, building model serving platforms, and ensuring reliable production deployments.
Requirements
The role requires extensive experience in DevOps, SRE, or Platform Engineering with a focus on GPU/accelerator infrastructure and Kubernetes. Proficiency in AWS, infrastructure-as-code, and Python is essential for managing high-performance ML workloads.
Full job description
Your Impact at LILA
The Staff/Principal DevOps Engineer - AI Inference will drive the design, implementation, and optimization of infrastructure purpose-built for serving machine learning models at scale. This role bridges platform engineering, site reliability, and ML infrastructure, building the systems that power low-latency, high-throughput inference across GPU clusters and cloud accelerators. You will collaborate with ML engineers, research scientists, and software engineers to build inference platforms that serve models reliably to production users while maximizing compute efficiency.
What You'll Be Building
GPU/accelerator infrastructure on Kubernetes: scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement for inference workloads
Model serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing
Intelligent request routing and load balancing across heterogeneous accelerator fleets (NVIDIA GPUs, AWS Inferentia/Trainium) to maximize utilization and minimize latency
Autoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads
Production-grade deployment pipelines for ML models: canary rollouts, A/B testing, model versioning, and safe rollback across multi-region deployments
Infrastructure-as-code with Terraform and Helm for GPU-accelerated EKS clusters, including node pools, spot/on-demand strategies, and accelerator-specific networking
Observability and performance optimization: GPU utilization monitoring, inference latency profiling, token throughput dashboards, and SLO/SLI tracking for model endpoints
CI/CD pipelines for model artifacts: container image builds with CUDA/driver dependencies, model registry integration, and automated inference benchmarking in CI
AWS cloud infrastructure for ML: EKS with GPU node groups, EC2 accelerated instances (P4/P5, Inf2, Trn1), S3 model storage, EFA/high-bandwidth networking, and IAM least privilege
Cost optimization and capacity planning: right-sizing accelerator instances, spot instance strategies for inference, and fleet-wide efficiency reporting
What You'll Need to Succeed
Expertise in DevOps, SRE, or Platform Engineering with significant experience operating GPU/accelerator infrastructure at scale
Deep experience with Kubernetes for ML workloads: GPU scheduling, resource quotas, node affinity, and accelerator device management
Strong proficiency deploying to AWS using infrastructure-as-code (Terraform, Helm) with hands-on experience managing GPU-based compute (EKS, EC2 P-series/Inf/Trn instances)
Experience with model serving infrastructure: inference servers, request batching, KV-cache optimization, or LLM serving frameworks
Strong understanding of networking for distributed inference: high-bandwidth interconnects, NCCL, VPC/PrivateLink, and load balancing at L4/L7
Strong proficiency in Python for automation, tooling, and integration with ML frameworks
Bonus Points For
Experience with LLM inference optimization: continuous batching, speculative decoding, quantization (GPTQ, AWQ, FP8), tensor parallelism, and pipeline parallelism
Hands-on experience with multiple accelerator families (NVIDIA A100/H100, AWS Inferentia2, Trainium, AMD MI300X) and maintaining hardware-agnostic serving infrastructure
Multi-region deployment experience with geographic routing and failover for latency-sensitive inference endpoints
Proficiency in Rust or Go for performance-critical infrastructure components
SRE practices for ML systems: chaos engineering on GPU workloads, incident management, capacity modeling for bursty inference traffic
Experience with model registries, artifact versioning, and ML supply chain security
Observability platform expertise: building custom metrics for token-level throughput, time-to-first-token, and per-request GPU memory profiling
Prior startup/high-growth experience balancing velocity with reliability in rapidly scaling AI systems
Compensation
We offer competitive base compensation with bonus potential and generous early-stage equity. Your final offer will reflect your background, expertise, and expected impact.
U.S. Benefits. Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and vision coverage; employer-paid life and disability insurance; flexible time off with generous company wide holidays; paid parental leave; an educational assistance program; commuter benefits, including bike share memberships for office based employees; and a company subsidized lunch program.
International Benefits. Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region. USD salary ranges apply only to U.S.-based positions; international salaries are set to local market.
Expected Base Salary Range
$192,000—$272,000 USD
About LILA
Lila Sciences is building Scientific Superintelligence™ to solve humankind's greatest challenges. We believe science is the most inspiring frontier for AI. Rather than hard-coding expert knowledge into tools, LILA builds systems that can learn for themselves.
LILA combines advanced AI models with proprietary AI Science Factory™ instruments into an operating system for science that executes the entire scientific method autonomously, accelerating discovery at unprecedented speed, scale, and impact across medicine, materials, and energy. Learn more at www.lila.ai.
Guided by our core values of truth, trust, curiosity, grit, and velocity, we move with startup speed while tackling problems of historic importance. If this sounds like an environment you'd love to work in, even if you don't meet every qualification listed above, we encourage you to apply.
We’re All In
Lila Sciences is committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status.
Information you provide during your application process will be handled in accordance with our Candidate Privacy Policy.
A Note to Agencies
Lila Sciences does not accept unsolicited resumes from any source other than candidates. The submission of unsolicited resumes by recruitment or staffing agencies to Lila Sciences or its employees is strictly prohibited unless contacted directly by Lila Science’s internal Talent Acquisition team. Any resume submitted by an agency in the absence of a signed agreement will automatically become the property of Lila Sciences, and Lila Sciences will not owe any referral or other fees with respect thereto.
Lila Sciences is the world’s first scientific superintelligence platform and autonomous lab for life, chemistry, and materials science. We are building the foundation to apply AI to every aspect of the scientific method, enabling scientists to bring forth solutions in human health and sustainability at a pace and scale never experienced before.
Based on 358 listings with disclosed salaries, most software jobs in Cambridge, MA pay between $100k–$251k per year. Individual offers vary with seniority, company size, and specialization.
How many Software jobs are open in Cambridge, MA right now?
There are currently 433 open software positions in Cambridge, MA listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Software roles in Cambridge, MA?
Companies currently hiring include Akamai Technologies, Lila Sciences, Google, MORSE Corp, Draper, among others. Browse the listings above to see every active employer.
Are there remote or hybrid Software jobs in Cambridge, MA?
Yes — 210 of the 433 open software positions offer remote or hybrid work (68 remote, 142 hybrid).
How do I apply for Software jobs in Cambridge, MA?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.