Site Reliability Engineer (m/f/d)

Location
Köln
Workplace
Hybrid

About this role

At our company, it’s all about #OneTeam! Join gridscale and help shape the future of the cloud together with OVH.

As a leading tech company, we’ve been working for over two decades to reduce our environmental footprint - with innovative solutions and an open cloud designed to be sustainable from the ground up: #SustainableByDesign.

Our Tech Stack 🚀

  • OpenStack • Kubernetes • GitOps • FluxCD / ArgoCD • Ansible • Terraform
    • Linux • Bare Metal • Prometheus • Grafana • Python • Go

Your Role💻

As a Senior Site Reliability Engineer you'll be part of a small and experienced team responsible for building, operating, and industrializing OVHcloud's on-premise cloud platform. You'll work on an infrastructure based on OpenStack and a Kubernetes / GitOps stack, with a strong focus on automation, compute lifecycle management, and AI-assisted engineering.

You'll help drive the OPCP roadmap, contribute to architectural and automation decisions, and continuously improve the platform in a security-oriented and highly automated environment. As a senior team member, you'll also share your expertise and support your peers on Platform Engineering, automation, and AI tooling.

Your Tasks

  • Design and build an on-premise cloud infrastructure based on OpenStack that can be deployed autonomously

  • Develop and maintain our Infrastructure as Code with Ansible and Terraform, including AI-assisted and specification-driven engineering workflows

  • Drive the ongoing development of our Kubernetes stack and GitOps workflows with FluxCD / ArgoCD

  • Take ownership of the compute infrastructure lifecycle, from bare metal and hypervisors to virtual nodes, and automate its ongoing operation

  • Develop and improve the team’s AI engineering capabilities, including knowledge bases and agent-based workflows for incident handling and capacity planning

  • Contribute to self-healing capabilities by transforming manual operational procedures into automated workflows

  • Design and implement test suites covering functional and technical specifications, including regression, performance, and security testing

  • Document and package the solution to ensure smooth deployment and operation for users

  • Continuously improve the platform based on telemetry and operational feedback

  • Act as a technical point of contact and support your peers on automation, Platform Engineering, and AI tooling

What we offer you💼

  • Exceptional team spirit across all departments and national borders; we live #OneTeam

  • Exciting work in a highly innovative and international environment with cutting-edge technologies

  • 32 vacation days, increasing with length of service

  • Flexible working hours, home office options, and a secure permanent position with market- and performance-based compensation

  • Employer-funded pension plan and an attractive insurance package

  • OVHcloud covers 50% of public transportation costs

  • Up to €400 annual financial contribution from OVHcloud towards sports activities (gym membership, sports classes, etc.)

  • Through Corporate Benefits, you receive attractive discounts at numerous shops and companies

  • We contribute to the leasing of your cargo bike

  • Regular company events and free cold and hot beverages



  • Several years of hands-on experience as an SRE, Platform Engineer, or DevOps Engineer in production environments, with strong expertise in Linux system administration and bare-metal infrastructure

  • Strong hands-on experience with OpenStack as well as end-to-end compute infrastructure management, from bare metal and firmware to hypervisors and virtualized environments

  • Extensive experience with Kubernetes and the cloud-native ecosystem, combined with strong knowledge of Infrastructure as Code and GitOps using technologies such as Ansible, Terraform, FluxCD, or ArgoCD

  • Practical experience with AI-assisted engineering, LLMs, and agent-based workflows, with a pragmatic understanding of their value in production environments

  • Good programming skills in Python and/or Go; knowledge of observability, networking concepts such as VLAN/BGP, or advanced compute optimization is a plus

  • Strong problem-solving and communication skills, ownership mentality, and fluent English skills

Tired of cold applications?

Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.

Know someone who'd be great for this?

Top Benefits

  • 32 vacation days
  • Flexible working hours
  • Home office options
  • Employer-funded pension plan
  • Insurance package
  • Public transportation subsidy
  • Sports activities contribution
  • Corporate discounts
  • Cargo bike leasing
  • Company events
  • Free beverages