San Francisco, California, United States · Remote OK
Senior
Site Reliability Engineer Platform and software · shared across customers Reports to: Director, Site Reliability Location: Remote (US) Department: Cloud Platform Engineering / SRE/Reliability Position summary The Site Re…
Data Center Technician/ Field Engineer Implementation and delivery · shared across sites Reports to: Director, Field Operations Location: (Odessa, TX); travel up to 50% Department: Infrastructure & DC Operations / Data C…
Hardware Engineer Infrastructure operations · shared across sites Reports to: Director, Hardware Engineering Location: Pleasanton, CA (hybrid) or assigned site; travel up to 25% Department: Infrastructure & DC Operations…
Skills: Hardware Engineering, GPU Systems, x86 Server Architecture, Linux, Python
Sales Operations AdministratorLocation: Onsite – Pleasanton, California Reporting to: Sales Ops At STN, we don't just adapt to the digital future, we engineer it. Our mission is to help organizations thrive in a rapidly …
Orangeburg, South Carolina, United States · On-site
$52k–$72k/yr
Mid level
Logistics CoordinatorLocation: Onsite – Orangeburg, SC Reporting to: DCO At STN, we don't just adapt to the digital future, we engineer it. Our mission is to help organizations thrive in a rapidly evolving technology lan…
Skills: Inventory Management, Asset Lifecycle Management, Shipping and Receiving, Supply Chain Management, Microsoft Excel
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Full-time
Posted 55d ago
~40 hrs/week
Remote in United States
Responsibilities
The SRE owns reliability, observability, and incident response for the GPUaaS platform. Key duties include defining SLOs, building the observability stack, and leading major incident resolution.
Requirements
Requires 5+ years of experience in SRE or DevOps with strong programming skills in Go or Python. Candidates must have hands-on experience with Kubernetes and observability tooling at scale.
The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution.
Key responsibilities
Define and operate Service Level Objectives (SLOs) aligned with customer SLAs
Build and maintain the observability stack including metrics, logs, traces, and alerting
Lead incident response and chair post-incident reviews
Drive automation to reduce toil and improve mean-time-to-recover (MTTR)
Author and maintain operational runbooks alongside the NOC
Manage on-call rotation, escalation paths, and incident-management tooling
Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering
Drive chaos engineering, game days, and reliability testing programs
Produce SLA performance reports in coordination with the SLA Manager
Mentor junior engineers and contribute to engineering culture
Required qualifications
5+ years in SRE, DevOps, or production engineering roles
Strong programming skills in Go, Python, or both
Hands-on experience operating Kubernetes-based platforms at scale
Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)
Strong incident management experience including major-incident command
Preferred qualifications
GPU or HPC platform operational experience
Familiarity with SLA-driven customer environments and credit calculations
Experience with chaos engineering tools (Gremlin, Litmus, or similar)
Strategy, Innovation & Consulting.
Human Element to the Best-in-class Technology.
Industry
IT Services and IT Consulting
Company size
11-50 employees
Founded
2016
Headquarters
Pleasanton, California
LinkedIn followers
4,997
At STN, we don’t just deliver technology, we build the foundation that modern organizations run on. From enterprise IT to AI infrastructure, we design, operate, and support systems that are reliable, secure, and built for performance. But what sets us apart isn’t just our stack it’s how we show up.
We don’t believe in one-size-fits-all. We don’t drop in hardware and disappear. We work side by side with our customers to understand what they actually need and build solutions that fit, flex, and scale as they grow.Whether you're running business-critical systems, deploying AI models, or training large-scale workloads on NVIDIA GPUs, we’re here to make sure your infrastructure isn’t just running, it’s working for you.
Our team brings deep technical expertise, hands-on support, and a people-first mindset to everything we do. Because we believe technology should unlock potential, not create more problems. STN exists to help teams thrive in complex environments with custom engineering, real partnership, and a clear plan forward.
Offices: 4464 Willow Rd, 102, Pleasanton, California 94588, US
Strategy, Innovation & Consulting.
Human Element to the Best-in-class Technology.
Industry
IT Services and IT Consulting
Company size
11-50 employees
Founded
2016
Headquarters
Pleasanton, California
LinkedIn followers
4,997
At STN, we don’t just deliver technology, we build the foundation that modern organizations run on. From enterprise IT to AI infrastructure, we design, operate, and support systems that are reliable, secure, and built for performance. But what sets us apart isn’t just our stack it’s how we show up.
We don’t believe in one-size-fits-all. We don’t drop in hardware and disappear. We work side by side with our customers to understand what they actually need and build solutions that fit, flex, and scale as they grow.Whether you're running business-critical systems, deploying AI models, or training large-scale workloads on NVIDIA GPUs, we’re here to make sure your infrastructure isn’t just running, it’s working for you.
Our team brings deep technical expertise, hands-on support, and a people-first mindset to everything we do. Because we believe technology should unlock potential, not create more problems. STN exists to help teams thrive in complex environments with custom engineering, real partnership, and a clear plan forward.
Offices: 4464 Willow Rd, 102, Pleasanton, California 94588, US