Clera home
·Dashboard

Jobs at Runloop (Now Hiring) — 1 open

Runloop logoRunloop

Site Reliability Engineer

San Francisco, California, United States · Hybrid

Senior

About Runloop Runloop.ai is pioneering the next generation of infrastructure and orchestration to power the Agentic Web/age of AI Agents. Our platform empowers developers to deploy agents that write code, browse the web,…

Skills: Python, Go, Docker, Kubernetes, Terraform

Runloop logo

Site Reliability Engineer

Runloop

San Francisco, California, United States • Hybrid

Apply
SeniorHybrid · 4 days in office

Tired of cold applications?

Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.

  • Full-time
  • bachelor degree
  • Competitive Salary, Equity, Health Insurance, Dental Insurance, Vision Insurance, Catered Lunch
  • Posted 15d ago
  • ~40 hrs/week

Responsibilities

The SRE will design and maintain production infrastructure on cloud platforms while ensuring high availability and security for code sandboxes. They will automate deployments, manage observability frameworks, and lead incident response and root-cause analysis.

Requirements

Candidates need 5+ years of software engineering experience, with at least 3 years focused on SRE or DevOps. Proficiency in Python or Go, containerization tools like Kubernetes, and cloud infrastructure management is required.

Full job description

About Runloop

Runloop.ai is pioneering the next generation of infrastructure and orchestration to power the Agentic Web/age of AI Agents. Our platform empowers developers to deploy agents that write code, browse the web, and use computers the way a human would. We're a small team of former Google and Stripe engineers, including the co-founder of Google Wallet and 100% of its founding team, dedicated to solving the complex challenges of productionizing AI for software engineering at scale.

The Role

We're looking for a skilled and passionate Site Reliability Engineer to join our team. As an SRE, you'll be responsible for the reliability, observability, performance, and security of our core platform, the foundation our users build their work on. You'll work closely with our engineering team to develop and maintain the systems that power our code sandboxes, ensuring a seamless and stable experience for our customers. This is a critical role that blends a deep understanding of operations with a software engineering mindset.

Responsibilities

  • Design and maintain our production infrastructure on cloud platforms like AWS, GCP, or Azure

  • Monitor and respond to system alerts and incidents using Grafana and Prometheus, ensuring high availability and a secure environment for our users' code

  • Collaborate with developers to ensure new features and services are designed with scalability and reliability in mind

  • Troubleshoot and resolve complex issues related to our infrastructure, networking, and the sandbox environment

  • Participate in an on-call rotation to support our production systems

  • Define and track SLIs/SLOs, manage error budgets, and proactively monitor distributed systems with logging and tracing

  • Automate deployments, scaling, provisioning, and recovery tasks to reduce toil and build self-healing systems

  • Lead incident response, conduct root-cause analysis, and facilitate blameless post-mortems to drive continual improvement

  • Collaborate cross-functionally with product, engineering, and developer relations to ensure reliable releases and an outstanding developer experience

  • Plan for capacity growth, forecast system usage, and contribute to safe release and change management processes

Qualifications

  • Strong computer science fundamentals, backed by a degree from a top-tier CS/EE program, or equivalent experience

  • 5+ years of experience in software engineering, with at least 3 years focused explicitly on site reliability, DevOps, or infrastructure operations

  • Strong programming skills in languages like Python or Go

  • Deep expertise in containerization technologies such as Docker and Kubernetes

  • Experience with cloud infrastructure and tools like Terraform and/or Pulumi

  • Familiarity with monitoring and alerting tools like Prometheus, Grafana, or Datadog

  • A solid understanding of networking, security, and Linux systems administration

  • Experience designing, scaling, and maintaining distributed systems (backend platforms, APIs, or front-end infrastructure)

  • Proficiency in implementing observability frameworks (metrics, logging, tracing) and aligning reliability goals with developer velocity

  • Hands-on experience managing incidents, running on-call operations, and producing actionable post-mortems

  • Ability to mentor engineers and influence reliability practices across teams, especially for front-end infrastructure and performance

Bonus Points

  • Experience with chaos engineering techniques, front-end observability tools (e.g., Sentry, RUM, synthetic monitoring), or building CI/CD pipelines for front-end delivery

Benefits

  • Competitive salary and equity

  • Comprehensive health, dental, and vision insurance for employee and dependents

  • Opportunity to work on cutting-edge technology and make a real impact on the future of software engineering

  • Daily catered lunch for all employees and a fridge full of your favorite snacks and drinks

Location:

  • Onsite 4 days a week in San Francisco; Optional 1 day a week remote

Join Us! If you're excited about shaping the future of AI-driven software engineering and empowering developers to build the next generation of AI powered coding tools, we want to hear from you. Join the Runloop team and be at the forefront of the AI revolution in software development.

Runloop AI is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, sexual orientation, gender identity, or any other characteristic protected by law.

Related keywords

AWSGCPAzureGrafanaPrometheusSLIsSLOsError BudgetsDistributed SystemsTerraformPulumiDockerKubernetesDatadogLinuxChaos Engineering

About Runloop

LinkedInVisit site

We build autonomous agents for DeFi.

Industry
Software Development
Company size
2-10 employees
Founded
2021
LinkedIn followers
145

Runloop builds simulation environments for DeFi and trains autonomous agents to provide liquidity, manage risk, and discover trading opportunities at scale.

View all jobs at Runloop

About Runloop

LinkedInVisit site

We build autonomous agents for DeFi.

Industry
Software Development
Company size
2-10 employees
Founded
2021
LinkedIn followers
145

Runloop builds simulation environments for DeFi and trains autonomous agents to provide liquidity, manage risk, and discover trading opportunities at scale.

View all jobs at Runloop

Similar companies hiring

Amazon (11177)Bosch (3616)Google (3535)Prolific (3434)AgileEngine (3057)Transport AI (1791)Booz Allen Hamilton (1584)Microsoft (1582)Speechify (1529)BJAK (1317)Salesforce (1031)Cisco (970)
Clera home

Your AI-talent agent. Connecting talents with dream jobs.

Earn $5,000

Tools

  • Salary Calculator
  • Resume Review
  • Startup Map

Explore

  • Jobs
  • Discover Jobs
  • Companies
  • Referral

Platform

  • Pricing
  • Integrations
  • Partners
  • Acquihire

Clera

  • Manifesto
  • Engineering
  • We are hiring!
  • FAQs
  • Blog
  • Press

Tools

  • Salary Calculator
  • Resume Review
  • Startup Map

Explore

  • Jobs
  • Discover Jobs
  • Companies
  • Referral

Platform

  • Pricing
  • Integrations
  • Partners
  • Acquihire

Clera

  • Manifesto
  • Engineering
  • We are hiring!
  • FAQs
  • Blog
  • Press

© 2026 Clera Labs, Inc.

PrivacyTermsBug Bounty