Sr. Engineer, Cloud HPC Platform

San Jose, CA · On-site$120k – $150k

About this role

Senior Engineer, Cloud HPC Platform

Location: San Jose, CA (Headquarters)
 
Ayar Labs is shattering AI data bottlenecks by moving data at the speed of light. As pioneers of co-packaged optics (CPO), we are using light instead of electricity to move data faster, further, and with a fraction of the energy needed to fuel the explosive growth of AI models.
 
Backed by industry giants like NVIDIA, AMD and Intel and manufactured in partnership with the world’s leading semiconductor ecosystem, Ayar Labs’ co-packaged optics solution is key to unleashing next-generation AI scale-up architectures.

Ayar Labs is moving silicon engineering compute workloads into AWS. The Senior Cloud HPC Platform Engineer will design, build, and operate the cloud HPC platform that supports EDA, simulation, verification, physical design, AMS, and related engineering workflows.

You will lead the migration from the current RHEL-based compute environment to AWS while protecting engineering productivity, design data, license availability, and output correctness. You will work closely with IT, TFM, ASIC, AMS, verification, physical design, security, finance, and EDA vendors.

This is a hands-on infrastructure role. You will own the platform from workload discovery and architecture through migration, production operations, cost management, and on-call support.
 

Key Responsibilities
  • Lead the AWS migration: Inventory engineering workloads, dependencies, data, licenses, and performance requirements; define migration waves, cutover plans, rollback procedures, and acceptance criteria.
  • Build the cloud HPC platform: Design and operate scalable AWS compute using appropriate EC2 instance families, accelerated networking, autoscaling, placement strategies, and workload isolation.
  • Own scheduling and job execution: Make Slurm the primary scheduler and operational control plane for interactive, batch, regression, and multi-day simulation workloads. Deploy and operate Slurm directly and/or through AWS ParallelCluster where appropriate; configure partitions, QoS, priorities, fair-share, reservations, preemption, accounting, dependencies, job arrays, and policy-based autoscaling.
  • Plan engineering run capacity: Partner with design and verification teams before major regressions, simulations, and tapeout milestones to translate run manifests and workload forecasts into CPU/core, memory, GPU, wall-time, scratch and capacity I/O, network, license-token, Slurm partition/reservation, and budget requirements; publish capacity scenarios, reservations, and readiness risks.
  • Engineer storage and data movement: Design high-performance storage and tiering across services such as Amazon FSx, EFS, S3, and on-prem systems. Establish backup, lifecycle, replication, and recovery controls.
  • Enable EDA workloads: Build reproducible RHEL-compatible environments for Cadence, Synopsys, Ansys, and other engineering tools. Support PDKs, third-party IP, shared flows, and controlled releases.
  • Manage licenses: Design reliable FlexNet/FlexLM access across hybrid and cloud environments, monitor utilization, and prevent licensing from becoming a scaling bottleneck.
  • Automate the environment: Define infrastructure through Terraform or OpenTofu and automate images, configuration, patching, and application deployment with tools such as Packer and Ansible.
  • Prove performance and correctness: Benchmark representative workloads before and after migration. Validate runtime, queue time, storage performance, reliability, cost, and quality-of-results with engineering owners.
  • Deliver self-service access: Architect and operate secure virtual desktop infrastructure (VDI/DVI) for engineering workflows, including Citrix Virtual Apps and Desktops and/or NICE DCV/VNC. Own application publishing, golden images, patching, SSO/MFA, session brokering and policies, profile and storage integration, GPU/graphics support, clipboard and file-transfer controls, monitoring, capacity, high availability, and performance troubleshooting; provide documented self-service paths for launching jobs and remote sessions.
  • Own security and connectivity: Implement least-privilege IAM, network segmentation, encryption, secrets management, logging, vendor access controls, and secure connectivity through VPN and/or Direct Connect.
  • Operate for reliability: Establish observability, service objectives, incident response, runbooks, change controls, and disaster-recovery testing. Partner with engineering on run-demand forecasting and capacity planning for major regressions, simulations, and tapeout milestones.
  • Control cloud cost: Implement tagging, budgets, chargeback/showback, scheduling policies, idle-resource controls, and workload-specific cost/performance optimization. Use Slurm accounting and workload forecasts to provide engineering with resource and cost estimates before large runs.
  • Reduce operational fragility: Replace undocumented manual steps and one-off scripts with versioned, tested, supportable automation and clear documentation.

Basic Qualifications
  • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience.
  • 7+ years building and operating Linux infrastructure, including 3+ years in AWS or a comparable cloud environment.
  • Deep hands-on experience with Enterprise Linux in production as a System Administrator
  • Experience designing or operating HPC, batch compute, large-scale simulation, or similarly compute-intensive platforms.
  • Deep production experience administering Slurm as the primary HPC scheduler, including partitions, QoS, priorities, fair-share, reservations, preemption, accounting, job arrays and dependencies, failure recovery, upgrades, and integration with AWS ParallelCluster or equivalent cloud capacity.
  • Strong AWS experience across EC2, IAM, VPC, S3, CloudWatch, Systems Manager, KMS, and high-performance storage services.
  • Strong infrastructure-as-code skills using Terraform or OpenTofu, including reusable modules, state management, review, and testing.
  • Experience automating Linux images and configuration with Packer, Ansible, Python, and/or Bash.
  • Strong knowledge of high-performance and shared storage, Linux file systems, data transfer, backup, and recovery.
  • Production experience operating secure virtual desktop infrastructure (VDI/DVI) for engineering workloads, preferably Citrix Virtual Apps and Desktops, including application publishing, image and patch lifecycle, SSO/MFA, session brokering and policy, profile and storage integration, GPU/graphics, clipboard and file-transfer controls, monitoring, capacity and high availability, and performance troubleshooting.
  • Strong knowledge of cloud networking, DNS, routing, firewalls, VPN, and hybrid connectivity.
  • Experience supporting FlexNet/FlexLM or another network-license system.
  • Experience establishing monitoring, alerting, incident response, capacity management, and cost controls for production infrastructure.
  • Ability to partner directly with engineers, translate run manifests and workload forecasts into CPU/core, memory, GPU, wall-time, storage I/O and capacity, network, license-token, Slurm partition/reservation, and budget requirements, and communicate capacity and migration risks clearly.
  • Clear written documentation, design proposals, operating procedures, and post-incident reviews.

Preferred Qualifications
  • Experience supporting semiconductor EDA environments, including Cadence, Synopsys, Ansys, Siemens EDA, PDKs, and IP libraries.
  • Experience migrating EDA, HPC, simulation, or verification workloads from on-prem infrastructure to AWS.
  • Experience with Amazon FSx for Lustre, FSx for OpenZFS, EFA, AWS Batch, ParallelCluster, or equivalent HPC services.
  • Experience designing and operating Citrix Virtual Apps and Desktops or comparable virtual desktop infrastructure (VDI/DVI) for engineering workloads, including application publishing, image and patch lifecycle, SSO/MFA, session brokering and policies, GPU/graphics, profile and storage integration, high availability, monitoring, and performance troubleshooting.
  • Experience benchmarking workload runtime, queue time, storage I/O, scaling efficiency, quality-of-results, and cost.
  • Familiarity with GitLab CI, artifact repositories, observability platforms, and controlled release processes.
  • AWS Professional or Specialty certification.

Salary range:  $120,000 - $150,000
 
 
NOTE TO RECRUITERS:
Principals only. We are not accepting resumes from recruiters for this position. Remuneration for recruiting activities is only applicable subject to a signed and executed agreement between the parties. Please don’t send candidates to Ayar Labs, and do not contact our managers.

Ayar Labs is an Equal Opportunity Employer and is strongly committed to all policies which will afford equal opportunity employment to all qualified persons without regard to age, sex, national origin, race, color, ethnicity, creed, religion, gender identity, sexual orientation, disability, veteran status, or any other characteristic protected by law. It is the policy of Ayar Labs to provide reasonable accommodation when requested by a qualified applicant or employee with a disability, unless such accommodation would cause an undue hardship. Veterans are more than welcome and encouraged to apply.

Company at a glance

Ayar Labs is transforming AI infrastructure by accelerating data movement. Recognizing that the complexity and size of AI models are increasing at a rate that traditional interconnect technology cannot handle, the company has developed the industry’s first optical I/O solution that enables customers to maximize the compute efficiency and performance of growing AI infrastructure, while reducing costs, latency and power consumption. Based on open standards and optimized for both AI training and inference, Ayar Labs’ optical I/O solution is backed by a robust ecosystem that enables it to integrate smoothly into AI systems at scale.

Founded2015
Team Size51-200 employees
WorkspaceOn-site
IndustryComputer Hardware Manufacturing
Location
San Jose, California, United States
LinkedInLinkedIn

Tired of cold applications?

Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.

Know someone who'd be great for this?