Principal Software Engineer - DevOps / Site Reliability Engineer
Singapore, Singapore · On-site
SeniorVisa sponsorship$21M raised
Riot Games was established in 2006 by entrepreneurial gamers who believe that player-focused game development can result in great games. In 2009, Riot released its debut title League of Legends to critical and player acc…
Skills: Site Reliability Engineering, DevOps, Python, Go, JavaScript
Job Responsibilities The Executive will manage the following areas: Clinical Quality and Patient Safety Coordinate and conduct initiatives and activities related to clinical quality and patient safety, including performi…
Job Responsibilities The Executive will manage the following areas: Clinical Quality and Patient Safety Coordinate and conduct initiatives and activities related to clinical quality and patient safety, including performi…
Job Title: A2AD Routing Analyst, Senior Job Category: Logistics Time Type: Full time Minimum Clearance Required to Start: TS/SCI Employee Type: Regular-Long Term Assignment Percentage of Travel Required: Up to 25% Type o…
Summary As a Safety intern, you will be providing support to the safety team, assisting with editorial support for department manuals and supporting activities associated with workplace safety and health. You will also b…
Our Mission At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology an…
Skills: Environmental Health and Safety, Risk Management, Compliance, Strategic Planning, Incident Investigation
Key Responsibilities • Lead the planning, execution, and delivery of Vulnerability Assessment and Penetration Testing (VAPT) engagements across network, application, endpoint, cloud, and hybrid environments. • Execute an…
Skills: Vulnerability Assessment, Penetration Testing, Red Team Operations, Adversary Emulation, Threat Simulation
Cushman & Wakefield
Senior Facilities Engineer
Singapore, Singapore · On-site
Senior$735M raised
Job Title Senior Facilities Engineer Job Description Summary Job Description About the Role · Responsible for the management and maintenance of the M&E facilities in the Building. · Oversee functions and activities of da…
Skills: Facilities management, M&E systems maintenance, Building services, Troubleshooting, Air conditioning systems
Job Title: APL/ASL Manager, Journeyman Job Category: Logistics Time Type: Full time Minimum Clearance Required to Start: Secret Employee Type: Regular-Long Term Assignment Percentage of Travel Required: Up to 10% Type of…
Location: Singapore, Singapore Thales is a global technology leader trusted by governments, institutions, and enterprises to tackle their most demanding challenges. From quantum applications and artificial intelligence t…
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deplo…
We invite suitable candidate to join a strong research team in Temasek Laboratories@NTU with the following requisites: Key Responsibilities: The candidate is expected to conduct and guide research in Wireless and Cellula…
Skills: Wireless Security, Cellular Security, Cybersecurity, Deep Learning, Software Engineering
At Temasek Laboratories@NTU, our mission is to undertake cutting-edge research in defence science and technology, driving new solutions and advancements across various research areas, including Microsystems Technology, H…
Skills: Statistical signal processing, Estimation theory, Detection and classification, Waveform-based geolocation, Machine learning
GIC is one of the world’s largest sovereign wealth funds. With over 2,000 employees across 11 locations around the world, we invest in more than 40 countries globally across asset classes and businesses. Working at GIC g…
Join DayOne – Shaping the Future of Data Infrastructure DayOne is a global leader in the development and operation of high-performance data centers. As one of the fastest-growing companies in the industry, we’ve built a …
Skills: EHS Management, ISO 14001, ISO 45001, Regulatory Compliance, Contractor Governance
Cushman & Wakefield
Facilities Manager
Singapore, Singapore · On-site
Senior$735M raised
Job Title Facilities Manager Job Description Summary Job Description About the Role: Oversee the facilities management, operations, maintenance, repair, and minor improvement works across the Site assigned.Manage and coo…
Skills: Facilities management, Building maintenance, M&E systems, Contractor management, Fire safety management
Job Title: Air Mobility/Airlift Planner, Journeyman Job Category: Logistics Time Type: Full time Minimum Clearance Required to Start: Secret Employee Type: Regular-Long Term Assignment Percentage of Travel Required: Up t…
About Mimecast Mimecast is a global cybersecurity leader redefining how organisations secure human risk. Our AI-powered, API-enabled Human Risk Management platform is purpose-built to protect organisations from the full …
Skills: Cybersecurity, CISO, Risk management, Strategic advisory, Executive communication
The job profile for this position is Cyber Security Lead Analyst, which is a Band 3 Senior Contributor Career Track Role. Excited to grow your career? We value our talented employees, and whenever possible strive to help…
Cohesity is the leader in AI-powered data security. Over 13,600 enterprise customers, including over 85 of the Fortune 100 and nearly 70% of the Global 500, rely on Cohesity to strengthen their resilience while providing…
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Full-time
bachelor degree
Full relocation support, Comprehensive health insurance, Open paid time off, Retirement benefits with company matching, Life insurance, Parental leave
Visa sponsorship available
Posted 12d ago
~40 hrs/week
Responsibilities
You will own and evolve the operational foundations for the AI Efficiency team's platform to ensure reliability, scalability, and security in production. This includes designing infrastructure, improving CI/CD pipelines, and establishing standards for observability, incident management, and service resilience.
Requirements
Candidates must have at least 5 years of experience in SRE, DevOps, or platform engineering with strong programming skills in languages like Python or Go. Proficiency in cloud environments, container orchestration, and infrastructure-as-code tools is essential for this role.
Full job description
Riot Games was established in 2006 by entrepreneurial gamers who believe that player-focused game development can result in great games. In 2009, Riot released its debut title League of Legends to critical and player acclaim. As the most played PC game in the world, over 100 million play every month. Players form the foundation of our community and it’s for them that we continue to evolve and improve the League of Legends experience.
We’re looking for humble but ambitious, razor-sharp professionals who can teach us a thing or two. We promise to return the favor. Like us, you take play seriously; you’re passionate about games. We embrace those who see things differently, aren’t afraid to experiment, and who have a healthy disregard for constraints.
That's where you come in.
The AI Efficiency team at Riot Games builds the platforms, tools, and technical foundations that help Rioters safely and effectively use AI to accelerate how we work. As these systems become increasingly important to creative, product, and development workflows across Riot, we need dedicated engineering leadership to ensure these systems remain stable, scalable, secure, and dependable in production.
As a Principal DevOps / Site Reliability Engineer on the AI Efficiency team, you will own and evolve the operational foundations that allow the AI Efficiency team’s tech platform and the tools deployed within it to run reliably at growing scale. You will establish the systems, standards, automation, and support practices required to move quickly without compromising availability, deployment safety, maintainability, or user trust.
You will partner closely with software engineers, ML platform engineers, technical artists, data scientists, and Riot’s infrastructure and security teams to improve developer experience, production readiness, observability, incident response, capacity planning, and service resilience. You will also help evaluate and operationalize AI-native engineering workflows such as agent-assisted code review, automated bug triage, AI-driven performance and security analysis, and browser-based UI validation. This role ensures the broader platform and its services are safely operated, supported, and continuously improved in production.
You’re right for this role if you enjoy making complex systems reliable, reducing operational toil, improving how engineers build and ship software, and anticipating how systems will fail before those failures affect users. You are comfortable taking ownership of production health, leading through incidents, building sustainable operational practices, and creating paved roads that help teams move quickly and safely. You are also energized by the opportunity to responsibly bring new AI-native automation patterns into real engineering workflows, thoughtfully applying emerging capabilities to reduce friction, improve reliability, and enhance how engineers interact with production systems without compromising safety or control.
Responsibilities:
Own and continuously improve the reliability, availability, scalability, performance, and operational health of the Efficiency team’s (web) platform and the tools deployed within it
Design, build, and maintain the infrastructure, deployment systems, and operational foundations required to support a growing portfolio of production AI services and internal tools
Improve CI/CD pipelines, release engineering practices, environment management, and deployment automation so software can be shipped safely, quickly, and consistently
Establish production-readiness standards and ensure new utilities have appropriate monitoring, alerting, ownership, documentation, rollback strategies, and support plans before launch
Define and operationalize service health indicators, SLIs, SLOs, error budgets, and reliability metrics that guide engineering priorities and tradeoffs between reliability, velocity, cost, and complexity
Build comprehensive observability across applications, infrastructure, service dependencies, and user workflows using metrics, logs, traces, dashboards, synthetic monitoring, and actionable alerts
Establish sustainable incident-management and on-call practices, including escalation paths, runbooks, severity definitions, communication protocols, and clear service ownership
Lead or contribute to the diagnosis and resolution of production incidents, coordinating across teams and driving blameless post-incident reviews and durable corrective actions
Build automation that reduces operational toil, improves mean time to detect and recover, and eliminates recurring sources of failure or manual intervention
Implement safe deployment patterns such as automated validation, progressive delivery, canary releases, feature flags, health checks, rollback mechanisms, and controlled environment promotion
Perform capacity planning, load testing, performance analysis, and resource forecasting to ensure the Toolkit can support increasing adoption and usage across Riot
Design and validate resilience, backup, recovery, failover, and disaster-recovery strategies for critical services, data, configurations, and infrastructure
Identify single points of failure and systemic risks across applications, cloud infrastructure, networking, databases, queues, caches, third-party dependencies, and operational workflows
Improve developer experience by building self-service workflows, reusable infrastructure components, local development environments, test environments, deployment tooling, and clear operational documentation
Establish and maintain infrastructure-as-code, configuration-management, secrets-management, and environment-governance practices that make infrastructure changes safe, repeatable, and auditable
Partner with engineers throughout the software development lifecycle to embed reliability, operability, security, and maintainability into system design rather than addressing them only after launch
Troubleshoot complex production issues across web applications, APIs, distributed services, containerized workloads, cloud infrastructure, network boundaries, authentication systems, and external service dependencies
Partner with ML Platform Engineers to ensure model-serving and inference systems integrate cleanly with the team’s broader observability, deployment, incident-management, and reliability standards
Collaborate with Riot infrastructure, information security, IT, developer-platform, and compliance teams to ensure the team follows appropriate operational and security requirements
Evaluate and implement AI-assisted operational workflows such as automated anomaly investigation, log analysis, remediation recommendations, regression detection, and runbook automation
Define guardrails, approval requirements, auditability, and escalation paths for agentic or automated operational systems that can interact with production environments
Champion operational excellence through technical leadership, mentoring, documentation, standards, architecture reviews, and tooling that raise the reliability bar across the team
Required Qualifications:
Bachelor’s degree in Computer Science or a related field, or equivalent professional experience
5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, Platform Engineering, Production Engineering, Developer Experience, or a similar role supporting production systems
Strong programming and automation skills in one or more languages such as Python, Go, JavaScript, or TypeScript
Experience designing, operating, and improving cloud-based production systems in AWS, GCP, Azure, or comparable environments
Experience building and maintaining CI/CD pipelines, release systems, deployment automation, and environment-management workflows
Strong understanding of observability practices, including metrics, logging, distributed tracing, dashboards, synthetic monitoring, and alert design
Experience participating in or leading incident response, on-call support, root-cause analysis, and post-incident improvement work
Experience improving the reliability, availability, scalability, and performance of distributed systems, service-oriented architectures, APIs, or web platforms
Strong understanding of containerized environments and orchestration technologies such as ECS, Docker, Kubernetes, or comparable systems
Experience with infrastructure-as-code and configuration-management tools such as Terraform, Pulumi, CloudFormation, or similar technologies
Working knowledge of Linux systems, networking, DNS, load balancing, service discovery, authentication, secrets management, and cloud security fundamentals
Ability to identify systemic operational risks and drive durable improvements across systems owned by multiple engineers or teams
Ability to collaborate across organizational boundaries, influence technical direction, and communicate clearly during both planned work and high-pressure incidents
Experience providing technical leadership, mentoring engineers, and establishing engineering standards across a team or organization
Desired Qualifications:
Experience supporting AI/ML platforms, inference services, model-serving systems, GPU-backed workloads, data pipelines, or other compute-intensive services
Experience defining and using SLOs, error budgets, and reliability metrics to guide prioritization and engineering decisions
Experience designing or improving internal developer platforms, self-service infrastructure, paved roads, golden paths, or shared engineering services
Experience building sustainable on-call rotations and operational support models for services used by multiple teams
Experience with progressive delivery, canary deployments, blue-green deployments, feature-flag systems, and automated rollback strategies
Experience with performance testing, capacity modeling, chaos engineering, fault injection, resilience testing, or failure-mode analysis
Experience designing backup, disaster-recovery, business-continuity, and regional failover strategies
Experience operating databases, caches, message queues, object storage, service meshes, API gateways, and other common distributed-system components
Experience improving security posture through access controls, secrets management, dependency management, vulnerability remediation, network segmentation, and infrastructure hardening
Experience balancing availability, latency, engineering velocity, infrastructure efficiency, and cost in systems operating at scale
Familiarity with browser automation and end-to-end testing frameworks such as Playwright for validating critical user workflows and detecting production regressions
Experience evaluating or integrating AI-assisted tools for incident investigation, anomaly detection, operational diagnostics, code review, test generation, or automated remediation
Familiarity with the risks and operational controls required when AI agents interact with source control, CI/CD pipelines, cloud infrastructure, or production systems
Experience establishing governance, approval workflows, audit trails, and quality controls for automated operational systems
Experience working in environments where experimental tools must be transitioned into reliable, supported, and maintainable production services
For this role, you'll find success through craft expertise, a collaborative spirit, and decision-making that prioritizes your fellow Rioters, who are the customers of your work. Being a dedicated fan of games is not necessary for this position!
Our Perks:
Full relocation support
Comprehensive health insurance for you, your spouse, and children
Open paid time off
Retirement benefits with company matching
Life insurance, parental leave, plus short-term and long-term disability
Play Fund so you can deepen your knowledge of our players and community through games
We’ll double down on your donations of time and money to non-profits
Related keywords
Site Reliability EngineeringDevOpsInfrastructure EngineeringPlatform EngineeringPythonGoJavaScriptTypeScriptAWSGCPAzureCI/CDKubernetesDockerTerraformPulumi
Rioters wanted: we’re looking for humble, but ambitious, razor-sharp pros who take play seriously.
Industry
Computer Games
Company size
1,001-5,000 employees
Founded
2006
Headquarters
Los Angeles, CA
LinkedIn followers
1,528,602
Total funding
$21M
Since 2006, Riot Games has stayed committed to changing the way video games are developed, published, and supported for players. From our first title, League of Legends, to 2020’s VALORANT; we have strived to evolve the community with growth in Esports, and expansion from games into entertainment. Players are the foundation of Riot's community and because of them, we’re able to reach new heights.
Founded by Brandon Beck and Marc Merrill, Riot is headquartered in Los Angeles, California, and has 4,500+ Rioters in 20+ offices worldwide. Riot has been featured on numerous lists including Fortune’s “100 Best Companies to Work For,” “25 Best Companies to Work in Technology,” “100 Best Workplaces for Millennials,” and “50 Best Workplaces for Flexibility.”
Riot Games recruiters will never ask for money or request sensitive information, and they'll always reach out from an @riotgames.com email address. You can learn more about Riot’s interview process here: https://www.riotgames.com/en/work-with-us/interviewing-at-riot/interview-process
Offices: Los Angeles, CA 90064, US · Dublin, IE · St. Louis, MO, US · Seoul, KR · Barcelona, ES
How many Security & Safety jobs are open in Singapore, Singapore right now?
There are currently 1,359 open security & safety positions in Singapore, Singapore listed on Clera. New openings are added daily as companies post roles.
Which companies are hiring for Security & Safety roles in Singapore, Singapore?
Companies currently hiring include NCS Group, Thales, UOB, SMRT Corporation Ltd, Cushman & Wakefield, among others. Browse the listings above to see every active employer.
Are there remote or hybrid Security & Safety jobs in Singapore, Singapore?
Yes — 257 of the 1359 open security & safety positions offer remote or hybrid work (38 remote, 219 hybrid).
How do I apply for Security & Safety jobs in Singapore, Singapore?
Each listing links directly to the employer's application page. Apply early — fresh listings get the most recruiter attention in the first two weeks.