About this role
Senior Platform Engineer / SRE
Level: Senior
Location: Remote in LATAM Only (Brazil preferred)
Working hours: EST
Type: Full-Time (Payment in USD)
Reports to: Engineering Manager, Platform
Company Overview
ReadyOn.AI is an AI native Labor Operating System redefining how the world’s largest enterprises manage frontline labor. Born out of a Stanford AI Lab, ReadyOn applies advanced AI and market design principles to one of the world’s hardest optimization problems: matching 2.7 billion frontline workers to the right shifts, in real time.
Frontline workers increasingly expect the flexibility and autonomy of gig platforms, while large employers face relentless pressure to control labor costs. ReadyOn bridges that gap with a system of action that predicts workforce demand, dynamically matches it to available employees, and automates thousands of staffing decisions across complex, multi site operations.
The platform is already proven at global scale, powering labor operations for some of the world’s largest enterprises, including F250 and F500 organizations spanning hundreds of thousands of employees and billions of dollars in annual labor spend. ReadyOn has demonstrated that scheduling was never the real problem. It was a symptom. The real challenge is dynamically matching people and work at scale.
Headquartered in San Francisco with over 100 employees, ReadyOn grew revenue 8x year over year in 2025, driven by multiple seven figure Fortune 250 deployments and a rapidly expanding pipeline.
Transform How Frontline Work Runs
Frontline labor can represent up to 40% of a company’s P&L, yet the systems managing this multi trillion dollar market were built around static schedules and manual processes.
ReadyOn is rejecting that paradigm. Staffing is not a scheduling problem. It is a real time supply and demand orchestration problem. Our AI native Labor Operating System is built from the ground up for AI agents to optimize labor in real time, much like ridesharing platforms match drivers and riders, but applied to frontline labor instead of fixed, one-size-fits-all schedules.
The Role
We're looking for a Senior Platform Engineer / SRE who can set the standard for how we automate and operate systems at scale. You'll tackle hard problems — multi-tenant isolation, self-service infrastructure, reliability engineering — and have the scope to solve them properly.
This is not a ticket-processing role. Seniors here identify problems before they're asked, and raise the ceiling on what the platform can do.
What You'll Work On
Shape how we do infrastructure-as-code — Terraform patterns, multi-account design, and the standards that hold it together across teams.
Operate GitOps at scale — ArgoCD configuration, managing promotion workflows, and ensuring deployment reliability across multiple environments and tenants.
Operate multi-tenant Kubernetes infrastructure on AWS EKS — managing tenant isolation, workload placement, cluster topology, and maintaining scalability.
Maintain self-service infrastructure automation — managing provisioning pipelines and configuration management.
Use agentic coding tools for infrastructure work — scaffolding new environments, generating and reviewing IaC, and accelerating automation.
Own reliability — managing SLOs, monitoring error budgets, ensuring incident response quality, and driving the feedback loop that turns incidents into platform improvements.
Maintain observability standards — ensuring trace coverage, alert quality, on-call ergonomics, and runbook culture.
Maintain security posture — focusing on secrets management at scale and infrastructure hardening.
Must Have
5+ years in platform engineering, SRE, or infrastructure — with meaningful time operating production systems at scale.
Deep IaC expertise — you actively manage complex Terraform state and multi-account configurations in production.
Strong GitOps background — you understand declarative infrastructure management at depth and have opinions on how to do it well.
Deep Kubernetes knowledge — you've operated clusters in production, dealt with real failure modes, and understand the system at the control plane level.
Strong AWS background — networking, compute, IAM, storage, multi-account design
Automation-first thinking at a senior level — you implement systems that eliminate entire categories of manual work.
Hands-on experience building and operating CI pipelines — GitHub Actions, CircleCI, GitLab CI, or equivalent.
Active user of agentic coding tools — you know how to direct them effectively, review their output critically, and use them to multiply your output.
Reliability engineering track record — SLOs defined and measured, post-mortems run, measurable improvements driven.
Strong communicator — you can articulate operational decisions and incident summaries clearly to engineers and leadership alike.
Nice to Have
Experience with Keycloak or other IdP
Experience with Argo: ArgoCD, Argo Workflows, Rollout
Experience with Karpenter and node lifecycle management in production
Background in FinOps — cost attribution, reserved capacity planning, workload right-sizing
Familiarity with data infrastructure — object storage, CDC pipelines, or lakehouse patterns
Experience with multi-tenant infrastructure — isolation patterns, noisy neighbor mitigation, and tenant lifecycle management.
Experience supporting AI/ML inference workloads or GPU-based compute in production
Prior experience scaling platform infrastructure at a startup moving toward enterprise-grade requirements
What You Won't Find Here
A platform team that maintains the status quo. We're actively building: new scale requirements, new architectural domains, and an ML/AI footprint that's growing fast. Senior engineers here shape how the platform evolves, and the tools available to do it are better than they've ever been.
Tired of cold applications?
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Know someone who'd be great for this?