Senior Platform Reliability Engineer

Location
Melbourne, Victoria, Australia
Workplace
Hybrid

About this role

Firmus Technologies

Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.  

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability. 

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally. 

 

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers. 

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale. 

 

AI FactoryOS Operations 

AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system. 

AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against. 

The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability, and builds and operates the shared services the estate's own operation depends on. The function works closely with the engineering teams that build the platform, supplying the production evidence that shapes what they fix and what they build next. 

 

Role Summary  

The Senior Platform Reliability Engineer is part of the team that operates Firmus AI FactoryOS in production: the GPU compute fleet, and the platform services it depends on, including exabyte-scale storage, the shared core services, the virtualisation hosting the management plane, and the observability infrastructure the estate is measured through. This is state-of-the-art AI infrastructure, among the largest deployments in Asia Pacific, built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation. 

This is a hands-on senior role with deep technical expertise, working in a team that shares accountability for the compute fleet and the platform services it depends on. The team runs those to a declared service level and sets the acceptance requirements each service has to meet before it goes live. The team also builds the shared administrative infrastructure the estate is run from, in consultation with the AI Infrastructure team, and operates it as a shared service. Automation is a first-class part of this role: the team builds and maintains the guarded automation and remediation tooling that turns manual response into a self-healing capability. 

 

Key Responsibilities 

  • Operate the multi-tenant control and management plane that Firmus' AI and infrastructure services depend on, including the tenancy, quota and access controls that support separation between tenant workloads and data. 
  • Operate Firmus' exabyte-scale distributed and high-performance filesystems (for example VAST, WEKA, Ceph) and S3-compatible object storage to their declared service levels, own the operational automation around them, and drive continuous improvement in how they are operated, feeding platform improvement requirements to AI Infrastructure with evidence. 
  • Operate the GPU compute fleet in production: node health and readiness, GPU and node fault detection and handling, firmware and driver currency to the supported baselines, remediation and return-to-service, and the hardware fault and replacement workflow with vendors and site operations. 
  • Build and maintain the guarded automation and remediation tooling that turns manual response into a self-healing capability, delivered as controlled code and reviewed by AI Infrastructure where it affects service behaviour. 
  • Diagnose and tune performance across the full data path, applying a deep understanding of operating systems, computer networks and software-defined storage, and working with technologies including RDMA, GPU Direct Storage, RoCE and InfiniBand. 
  • Deliver change as code, and drive continuous improvement in cluster validation, CI/CD automation, and provisioning and testing frameworks. 
  • Operate the observability infrastructure as a shared service across metrics, logs, traces, alerting and retention, and work with service owners so that telemetry becomes alerting that is actionable and tied to a runbook. 
  • Operate each service in the portfolio to its published service level, carry new services through production readiness review, and execute the monthly patching cycle and urgent vulnerability remediation. 
  • Provide the deepest technical expertise for these platforms, taking on and diagnosing the faults that require internals-level knowledge to root cause, and driving the permanent fix to closure through AI Infrastructure, and manage vendor escalations at engineering level. 
  • Lead technical recovery during major incidents, drive the changes that remove repeat causes, share the follow-the-sun on-call roster, and mentor the engineers who carry frontline diagnosis, documenting operational procedures, runbooks and performance results. 

 

Skills & Experience 

Required Skills 

  • Strong skills in infrastructure, systems and platform engineering, with 8+ years of experience including substantial ownership of production storage and shared infrastructure services in a 24/7 environment. 
  • Extensive experience with scale-out, parallel or enterprise storage supporting demanding workloads (for example VAST, WEKA, Ceph, Lustre, GPFS or NetApp), including diagnosis of performance and capacity problems across the full data path using evidence rather than assumption. 
  • Substantial experience operating a virtualisation platform (for example Proxmox, VMware or KVM) and building shared services such as databases and object storage to a defined service level. 
  • Expert-level knowledge of Linux systems, including storage and file system internals, kernel and driver behaviour, networking, memory and I/O subsystems, and systematic performance analysis. 
  • Extensive experience operating observability infrastructure as a service across metrics, logs, traces and alerting (for example Prometheus, Grafana, OpenTelemetry, Loki or Elasticsearch), and designing alerting that is actionable. 
  • Strong skills in infrastructure automation, infrastructure-as-code and GitOps practices (for example OpenTofu or Terraform, Ansible, Argo CD, CircleCI), with change delivered through peer review, automated testing and progressive rollout. 
  • Practical experience with scripting or programming for operational automation and tooling, such as Python, Go or Bash. 
  • Working competence with Kubernetes, sufficient to diagnose faults where shared services meet the container estate. 
  • Proven ability to act as a senior escalation point in production, including major incident response, on-call participation, post-incident review, vendor escalation, and the production of runbooks that others can execute successfully. 
  • Practical application of least privilege, secrets management, certificate-based access and audited privileged operations, and an understanding of how operational practice produces audit evidence. 
  • Clear technical judgement and communication skills, with the ability to produce incident updates, design notes and escalations that non-specialists can act on, and a working style suited to a small team that shares accountability for a broad portfolio. 

 

Preferred Experience 

  • Experience in storage and infrastructure operations for GPU or HPC environments, including the demands of large-scale distributed training and inference. 
  • Experience in multi-tenant service provider, cloud or colocation environments with contractual response commitments and customer-facing reporting. 
  • Familiarity with bare-metal lifecycle and provisioning automation (for example NVIDIA Infra Controller, MAAS, Ironic or Metal3), and standard operating environments with known-good-state enforcement. 
  • Experience with multi-tenant Kubernetes patterns and GPU-enabled Kubernetes. 
  • Knowledge of data centre and hardware fundamentals, including firmware management, DPU or SmartNIC-based networking, and high-performance fabrics such as InfiniBand or RoCE. 
  • Familiarity with policy-as-code and software-supply-chain controls, such as image signing, software bills of materials and vulnerability gating. 
  • Experience preparing for or operating under ISO 27001, or SOC 2, and responding to enterprise customer due diligence. 
  • A Bachelor's degree in computer science, engineering or a related discipline, or an equivalent combination of relevant experience and training.  

 

Location & Reporting 

LocationBased in Australia or Singapore, with travel to Australian AI Factory sites as required. 

On-call: The function runs 24/7. First line monitoring and first response sit with the operations centre. This role shares the after-hours escalation roster for its domain with the other senior engineers in the function. 

Reporting to:  Reports to the Head of AI FactoryOS Operations while the function is being established, working under broad direction with a high degree of autonomy and direct access to the decision makers. As the function reaches its planned structure, the role will report to the Service Reliability Manager, with the Head of AI FactoryOS Operations remaining accountable for the function. The scope, level and remit of the role do not change under either arrangement. 

 

Employment Basis

Permanent full-time

 

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering. Apply now to be part of shaping the future of sustainable AI infrastructure.

Tired of cold applications?

Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.

Know someone who'd be great for this?