Site Reliability Engineer

Antananarivo · Remote ok

About this role

Our client's Cloud Operations team is expanding its SRE function. Site Reliability Engineers keep all user-facing services and production systems running smoothly. SREs here are a blend of pragmatic operators and software craftspeople who apply sound engineering principles, operational discipline and mature automation to the environment and the codebase. The team specialises in systems — networking, the Linux kernel, and scaling, algorithms and distributed systems.

As an SRE you will

  • Be on an on-call rotation responding to production availability incidents, and support service engineers with customer incidents
  • Use your on-call shift to prevent incidents from ever happening
  • Run infrastructure with Ansible, Puppet, Terraform and Kubernetes
  • Make monitoring and alerting alert on symptoms, not outages
  • Document every action, so findings turn into repeatable actions — and then into automation
  • Improve the deployment process to make it as boring as possible
  • Design, build and maintain core infrastructure that scales to hundreds of thousands of concurrent users
  • Debug production issues across services and levels of the stack
  • Plan the growth of the infrastructure

You may be a fit if you

  • Think cloud-first, regardless of the flavour of public cloud
  • Think security-first
  • Think about systems — edge cases, failure modes, behaviours, specific implementations
  • Know your way around Linux and Windows
  • Know the use of config-management systems like Ansible or Puppet
  • Have strong programming skills — Python, Java, Golang, Node.js
  • Collaborate and communicate asynchronously, and document so nothing is learned twice
  • Have a go-for-it attitude: when you see something broken, you fix it
  • Have experience with Nginx, HAProxy, Docker, Kubernetes, Terraform or similar technologies

Projects you could work on

  • Coding infrastructure automation with Ansible and Terraform
  • Improving Prometheus monitoring or building new metrics
  • Helping release managers deploy and fix new versions of application software
  • Planning and executing the migration from AWS virtual machines to cloud-native, container-based deployments on Kubernetes (EKS)
  • Developing a relationship with a product group and defining their SRE KPIs — the SRE practice here is early in its journey

Company at a glance

Ontrac Solutions helps organizations adopt emerging technologies to scale smarter.
We build GenAI platforms, predictive analytics solutions, and drive cloud adoption.
We're also a HubSpot partner, supporting landing page design, website development, CRM integration, workflows, and automation.

From infrastructure to marketing ops, we deliver strategy and execution that drives growth.

Founded2010
Team Size11-50 employees
WorkspaceRemote ok
IndustryIT Services and IT Consulting
Location
Antananarivo, Analamanga, Madagascar

Tired of cold applications?

Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.

Know someone who'd be great for this?