Who We Are Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with…
Skills: Kubernetes, Linux, PyTorch, CUDA, NCCL
mouhamedzz
Wargaming Research & Systems Analyst — White Cell
Abu Dhabi, Abu Dhabi Emirate, United Arab Emirates · On-site
Mid level
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century…
Skills: Operations Research, Systems Analysis, Wargaming, Modeling and Simulation, Python
mouhamedzz
Wargaming Research & Systems Analyst — Red Cell
Abu Dhabi, Abu Dhabi Emirate, United Arab Emirates · On-site
Mid level
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century…
Skills: Operations Research, Systems Analysis, Wargaming, Modeling and Simulation, Python
mouhamedzz
Wargaming Research & Systems Analyst — Red Cell Lead
Abu Dhabi, Abu Dhabi Emirate, United Arab Emirates · On-site
Senior
Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century…
Skills: Wargaming, Operations Research, Systems Analysis, Python, C++
About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing …
Skills: Security Assurance, GRC, Control Testing, Risk Management, TISAX
Software Engineer, Science and Strategic Initiatives, DeepMind
London, England, United Kingdom · On-site
$174k–$252k/yr
Senior$26M raised
Minimum qualifications: Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience. 5 years of experience in software design and development using Python, distributed systems, or…
Skills: Python, Distributed systems, Cloud infrastructure, AI systems, Machine learning
Company Description Wise is a global technology company, building the best way to move and manage the world’s money. Min fees. Max ease. Full speed. Whether people and businesses are sending money to another country, spe…
Skills: Java, Spring Framework, React, Javascript, Typescript
About Arondite Arondite enables organisations to wield autonomy, data and AI at scale. We equip operators with the foundational software to orchestrate their mix of platforms and capabilities, and then take mission-criti…
Location: 1-2 days per week in the London office Salary: Up to £75,000 (+ £1,800 wellbeing allowance + up to 10% bonus) Who are we? We’re the original pioneers in connected commerce marketing. Since 2008, we’ve been part…
Location: 1-2 days per week in the London office Salary: Up to £65,000 (+ £1,800 wellbeing allowance + up to 10% bonus) Who are we? We’re the original pioneers in connected commerce marketing. Since 2008, we’ve been part…
Lead Data Engineer - 12 Month FTC Who We Are: AND Digital are a tech company focused on accelerating digital delivery and dedicated to closing the digital skills gap. We’ve been helping organisations build better digital…
Skills: SQL, Python, Data engineering, ETL/ELT, Data architecture
About Neo4j: Neo4j is the graph intelligence platform that transforms data into knowledge to power the next generation of intelligent applications and AI systems. It includes enterprise-ready knowledge graphs for accurat…
About Eucalyptus We're on a mission to make good health last a lifetime. More than 1 billion people live with obesity worldwide, driving preventable chronic conditions. We're here to build better long-term care. Eucalypt…
MongoDB Professional Services (PS) works with customers of all shapes and sizes, in all verticals, from tier-1 banks to small web startups, on a variety of exciting use cases. This role solves technically sophisticated p…
Note for Recruitment Agencies: We prefer to hire directly and we will be in touch with our PSL Agencies if this role is eligible for release. We do not accept speculative CVs from agencies. If speculative CVs are sent, n…
Gibson Dunn is a leading global law firm, advising clients on significant transactions and disputes. Our exceptional teams craft and deploy creative legal strategies that are meticulously tailored to every matter, howeve…
Company Description Are you looking to join an organization that is growing and dynamic? What about a high-energy, collaborative environment that rewards hard work? J.S. Held is a global consulting firm that combines tec…
Company Description The Product Security team at NielsenIQ is dedicated to enhancing business application security and making product teams accountable throughout the software development lifecycle. It not only protects …
Lead Mobile Maintenance Engineer (Capacity Engineer) Location: Central London (Multiple Sites) About the Role We are looking for a Lead Mobile Maintenance Engineer to provide technical expertise across a portfolio of com…
Skills: Electrical maintenance, Mechanical maintenance, Fault-finding, Capacity planning, System optimisation
Shift Engineer Shift Pattern: Continental shifts 12 hour day / 12 hour nights Location: Central London Purpose of Job To ensure all environmental conditions are maintained at all times with regard to critical building sy…
Skills: Air conditioning, UPS, Generators, LV systems, Fault finding
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
$75k–$95k/yr
Full-time
Comprehensive Health Coverage, Meaningful Equity, Pension Contributions, Unlimited PTO, Company-Wide Winter Break, Paid Parental & Family Leave
Posted 3h ago
~40 hrs/week
Responsibilities
You will act as a technical partner to ML engineering teams, diagnosing complex distributed systems issues and improving platform reliability. Responsibilities include investigating failures in training and inference workloads, analyzing system performance, and contributing to internal tooling and documentation.
Requirements
Candidates must have a strong background in software engineering, Linux systems, and Kubernetes, along with hands-on experience operating machine learning workloads. Proficiency in distributed systems, GPU infrastructure, and observability tools is essential for this technical support role.
Full job description
Who We Are
Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems—designed to take ideas from research to production with less friction.
Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in.
We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.
The Way We Work
The people who thrive here are builders who move fast, communicate openly, take ownership, and continuously improve themselves, their teams, and our company. Here's what that looks like in practice:
Move with Urgency: We move quickly, make thoughtful decisions, and keep momentum. We value action over perfection and learn by shipping.
Take Ownership: We own outcomes, not just our individual work. We make decisions that move the company forward and follow through.
Communicate Openly: We communicate directly, seek to understand, and create clarity for others. Honest conversations help us move faster together.
Build Great Teams: We lead by example, empower others, and create healthy teams where people can do their best work.
Raise the Bar: We're always improving ourselves. We learn from feedback, consistently challenge ourselves to grow, and focus on the work that matters most.
Think Long-Term: We design for what's next. We create scalable systems, simplify complexity, and use AI and automation to amplify our impact.
What We’re Looking For
Lightning AI is looking to hire AIPlatform Support Engineers to join our EMEA Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production environments.
This role is not a ticket router or traditional support engineer. You are a technical partner to ML teams - helping diagnose failures, improve reliability, and guide customers through complex distributed systems problems.The problems range from Kubernetes scheduling and GPU orchestration to distributed PyTorch failures, inference latency, networking bottlenecks, storage performance, and platform reliability. You’ll gain exposure to a wide variety of real world AI workloads across industries and help shape the infrastructure powering the next generation of ML applications.
We are currently hiring for two EMEA shifts (9AM–7PM CET/CEST):
Saturday–Tuesday
Thursday–Sunday
This role is hybrid out of our London office, with an in-office requirement of at least 2 days per week and occasional team and company offsites. We are not able to provide visa sponsorship for this role at this time.
What You'll Do
Work Directly With ML Engineers
Partner directly with customer engineering teams running training and inference workloads in production
Help customers diagnose and resolve complex distributed systems and ML infrastructure issues
Act as a technical advisor during high impact incidents and platform degradation events
Translate infrastructure level issues into actionable guidance for ML engineers
Build credibility with customers through strong technical reasoning and clear communication
Debug ML Infrastructure & Distributed Workloads
Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems
Troubleshoot PyTorch, CUDA, NCCL, and inference serving related issues
Analyze logs, metrics, traces, and system behavior to isolate root causes
Debug containerized workloads running across Kubernetes and bare metal GPU environments
Support customers scaling workloads across multi node GPU systems
Diagnose performance bottlenecks involving compute, memory, networking, or storage
Improve Reliability & Platform Operations
Identify recurring patterns across customer issues and drive long term reliability improvements
Contribute to post incident reviews and operational improvements
Build internal tooling, automation, documentation, and runbooks
Partner closely with infrastructure, networking, and platform engineering teams
Help improve observability, operational visibility, and troubleshooting workflows
Improve the customer experience through better processes and technical guidance
What This Role Is Not
To set clear expectations:
This is not a traditional help desk or ticket routing support role
This is not purely customer success or account management
This is not a backend engineering role
This is not a passive escalation position
This role is for engineers who enjoy solving difficult technical problems while working closely with other engineers.
What You’ll Need
Required Qualifications
Infrastructure & Systems
Strong software engineering and systems troubleshooting background
Experience with Kubernetes and containerized environments
Linux systems knowledge, including networking, storage, process management, and performance tuning
Experience with cloud infrastructure and distributed systems
Experience with observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry
ML Infrastructure Experience
Hands on experience operating machine learning workloads in production or research environments
Experience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL
Familiarity with GPU infrastructure and orchestration
Experience troubleshooting performance, reliability, or scaling issues in ML infrastructure
Understanding of the operational challenges involved in running ML systems at scale
Collaboration
Strong communication skills and ability to work directly with highly technical customers and engineering teams
Comfortable operating in fast moving, highly ambiguous environments
Experience with large scale model training or distributed inference systems
Familiarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms
Experience with InfiniBand, RDMA, or high-performance networking
Experience operating bare metal infrastructure
Familiarity with storage systems commonly used in ML environments
Experience working at an AI infrastructure, cloud, MLOps, or developer tooling company
Contributions to platform engineering, developer infrastructure, or operational tooling projects
Experience writing automation, tooling, or scripts in Python or similar languages
Compensation
We are committed to offering competitive compensation that reflects the value each team member brings to our mission. Final offers are based on factors such as experience, skills, geographic location, and role expectations. In addition to base salary, our total rewards package for eligible roles includes a discretionary bonus, a meaningful equity component, and comprehensive benefits.
The anticipated annual base salary range for this role is:
£75,000—£95,000 GBP
Benefits and Perks
We offer a comprehensive and competitive benefits package designed to support our employees’ health, well-being, and long-term success:
Comprehensive Health Coverage: Medical, dental, and vision coverage for employees and eligible dependents.
Meaningful Equity: RSUs that give employees a stake in the company's long-term success.
Retirement Savings: 401(k) matching (U.S.) and pension contributions (U.K.).
Flexible Time Off: Unlimited PTO, company holidays, and floating holidays to support work-life balance.
Company-Wide Winter Break: Two weeks of company closure each winter to disconnect and recharge.
Paid Parental & Family Leave: Paid leave to support you and your family through life's important moments.
Professional Development: Annual learning and development allowance to support your professional growth.
Wellness Benefits: Wellness and work-from-home stipends to support your physical and mental well-being.
Sabbatical Program: Four weeks of paid sabbatical leave after four years of service.
Flexible Work: Flexible schedules and a hybrid work model for our office-based teams.
In-Office Meals: Complimentary meals at our office hubs.
Benefits may vary by location, team, and role.
At Lightning AI, we are committed to fostering an inclusive and diverse workplace. We believe that diverse teams drive innovation and create better products. We provide equal employment opportunities to all employees and applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic. We are dedicated to building a culture where everyone can thrive and contribute to their fullest potential.
Related keywords
AI PlatformMachine LearningKubernetesPyTorchCUDANCCLGPUDistributed SystemsCloud InfrastructureObservabilityPrometheusGrafanaOpenTelemetryNetworkingStoragePython
The full-stack AI infrastructure platform. From GPU to endpoint.
Industry
Software Development
Company size
51-200 employees
Headquarters
New York, NY
LinkedIn followers
102,195
Total funding
$109M
The AI development platform - From idea to AI, Lightning fast ⚡️. Code together. Prototype. Train on GPUs. Scale. Serve.
From your browser - with zero setup.
AI Studio is your laptop on the cloud. Zero setup. Always ready. Persistent storage and environments. Code on CPU. Debug on GPU. Scale to multi-node. Run sweeps, jobs and more.
Scale models with PyTorch Lightning, Fabric, Lit-GPT, torchmetrics and more.
Offices: New York, NY, US · San Francisco, US · Seattle, US · London, GB
Artificial IntelligenceMachine LearningInfrastructuredeep learningdata scienceand open sourceSoftware EngineeringInformation TechnologyArtificial IntelligenceSoftware
How many Engineering jobs are open in London, United Kingdom right now?
There are currently 5,707 open engineering positions in London, United Kingdom listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Engineering roles in London, United Kingdom?
Companies currently hiring include Turner & Townsend, Amazon, Wise, AECOM, JPMorganChase, among others. Browse the listings above to see every active employer.
Are there remote or hybrid Engineering jobs in London, United Kingdom?
Yes — 3118 of the 5707 open engineering positions offer remote or hybrid work (334 remote, 2784 hybrid).
How do I apply for Engineering jobs in London, United Kingdom?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.