Head of Data Center Operations

San Francisco · On-site$250k – $350k

About this role

fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.

As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.

Own how fal runs its data centers. You'll lead data-center operations end to end, strategy, capacity, build-out, and steady-state, for the GPU fleet that powers every fal inference. You'll operate through our on-site lead and technicians while staying close to engineering and leadership, and help shape the compute organization as we scale.

WHAT YOU'LL OWN

  • Own the full DC operations lifecycle — planning, design, build-out, deployment, and steady-state operations across our sites.

  • Set infrastructure operations strategy — capacity, cost, reliability, and scale, with clear OKRs/KPIs, in lockstep with engineering, finance, and leadership.

  • Run operations through the on-site team — direct our site lead and technicians, own incident and escalation management, and drive uptime and MTTR across sites.

  • Partner with Engineering, Network, and Capacity Planning to execute infrastructure expansion while optimizing power, cooling, rack density, and deployment schedules.

  • Develop operational processes, documentation, and vendor governance as fal continues to scale its global infrastructure footprint.

WHAT YOU BRING

  • 10+ years in data center / infrastructure operations including senior leadership (Director level) owning multi-site or global operations.

  • 10MW+ facility experience — you've owned operations for a large-scale, mission-critical data center.

  • Full-lifecycle experience — planning and build-out through steady-state — ideally for high-density GPU / accelerator (AI inference) environments.

  • Strategic and cross-functional range — you partner naturally with engineering, finance, and business leadership on capacity, cost, and scale.

  • Command of commercial / vendor relationships and colocation / leased-space operations.

BONUS POINTS

  • You've stood up new data-center capacity from the ground up.

  • Experience operating high-density LPU/GPU or liquid-cooled environments.

  • You've built and scaled DC operations teams and the pipeline that feeds them.

Company at a glance

fal is a generative media platform that provides developers with access to the world's best generative image, video, and audio models through a unified API. Trusted by over 2.5 million developers and leading companies, fal offers the fastest inference engine for diffusion models, on-demand serverless GPUs, and dedicated compute clusters for frontier research.

For press inquiries: [email protected]
For customer support: [email protected]

Team Size51-200 employees
WorkspaceOn-site
IndustryTechnology, Information and Internet
Location
San Francisco, California, United States
Websitefal.ai
LinkedInLinkedIn

Top Benefits

  • Equity

Tired of cold applications?

Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.

Know someone who'd be great for this?