San Francisco, California, United States · Remote OK
$200k–$250k/yr
Senior$1.3B raised
About The Role Together AI is growing its compute footprint, and making sure new capacity meets our technical standards is an important priority for the company. Every new cluster has to clear a technical bar before it c…
Skills: Technical Program Management, Infrastructure Qualification, GPU Hardware, High-Performance Networking, Data Center Operations
About the role As a Forward Deployed Engineer (FDE) focused on Inference & Post-Training, you will be a hands-on technical partner to our most strategic customers — production AI teams looking to leverage high quality mo…
Technical Support Engineer (GPU Clusters) - US Weekends
Remote OK
$160k–$230k/yr
Mid level$1.3B raised
About the role As a Technical Support Engineer at a pioneering AI company, you'll be the first line of defense to support customers as they build out training, fine tuning, and inference solutions with Together AI. You'l…
Technical Support Engineer (Inference) - US Weekends
Remote OK
$160k–$230k/yr
Senior$1.3B raised
About the role As a Technical Support Engineer at a pioneering AI company, you'll be the first line of defense to support customers as they build out training, fine tuning, and inference solutions with Together AI. You'l…
AI Infrastructure Systems Engineer Build the infrastructure powering the next generation of AI. At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inferenc…
About the Role As an IT Engineer, you will be working across identity, devices, SaaS platforms, and supporting end users. You should be organized, collaborative, and eager to learn while working within established system…
Skills: Okta, Google Workspace, MacOS, Linux, FleetDM
San Francisco, California, United States · On-site
$200k–$290k/yr
Senior$1.3B raised
About the Role The Model Shaping team at Together AI works on products and research for tailoring open foundation models to downstream applications. We build services that allow machine learning developers to choose the …
Skills: Python, PyTorch, Distributed training, GPU architecture, Mixed-precision training
San Francisco, California, United States · On-site
$150k–$200k/yr
Mid level$1.3B raised
About the Role Together.ai is looking for a Software Engineer to join the Customer Insights team — a great role for a full-stack or backend engineer who wants to grow into event-driven systems, analytics, and the custome…
San Francisco, California, United States · On-site
$160k–$230k/yr
Senior$534M raised
About the Role Together AI's product is what developers touch every day — Playground, Model Garden, Fine-Tuning, File Management, Batch Processing, Model Evals. These surfaces are how users experience the platform, and t…
San Francisco, California, United States · On-site
$240k–$280k/yr
Senior+$1.3B raised
About the Role We're looking for a Software Engineer to build the systems that treat infrastructure as software. This role owns the software state machines that provision hardware, bring it into service, and manage its f…
About the Role We’re hiring a Strategic Account Executive to own a portfolio of high-priority AI-native and enterprise accounts. You’ll be responsible for generating pipeline, closing new business, and expanding customer…
Senior Software Engineer Full Stack Engineer (TypeScript) Together.ai is looking for a Senior Software Engineer to take a leading role in the authentication, authorization, and collaboration systems that every Together p…
San Francisco, California, United States · On-site
$300k–$350k/yr
Senior+$534M raised
About the Role Together AI is building the AI-native cloud — the fastest inference infrastructure on the planet, paired with a model ecosystem, orchestration layer, and data center footprint that enterprises and frontier…
Skills: Strategic Partnerships, Business Development, Alliance Management, Commercial Negotiation, Cloud Distribution
San Francisco, California, United States · On-site
$240k–$280k/yr
Senior$534M raised
About the Role Together AI is hiring a Staff Platform Engineer to join the Product Foundations engineering organization and drive its service infrastructure strategy. Product Foundations builds and operates Together’s mi…
San Francisco, California, United States · On-site
$200k–$230k/yr
Senior$534M raised
About the Role Together AI is looking for a Commercial Counsel to support both the agreements that secure our infrastructure and the deals that drive our growth. You'll work alongside the Infrastructure, Sales, Finance, …
About the Role As a Senior Network Engineer at Together, you are responsible for designing, implementing, and maintaining our network infrastructure to ensure seamless connectivity and optimal performance for all user-fa…
San Francisco, California, United States · On-site
$175k–$220k/yr
Mid level$534M raised
About the Role Our product surface is expanding fast - GPU clusters, managed storage, networking, and observability - and we're adding a Product Manager to the Together Cloud team to own the day-to-day product work that …
Skills: Product Management, Data Analytics, Cloud Infrastructure, AI Infrastructure, GPU Clusters
Research Intern RL & Post-Training Systems, Turbo (Fall 2026)
San Francisco, California, United States · On-site
$58/hr–$63/hr
Entry level$534M raised
About the Role The Turbo Research team investigates how to make post-training and reinforcement learning for large language models efficient, scalable, and reliable. Our work sits at the intersection of RL algorithms, in…
Skills: Reinforcement Learning, Post-Training, ML Systems, Python, C++
San Francisco, California, United States · On-site
$150k–$200k/yr
Mid level$534M raised
About the Role We’re looking for a detail-oriented Data Center Operations professional to manage and track all break/fix activities across multiple data center locations. This role acts as the central point of coordinati…
Skills: Data Center Operations, Incident Management, SLA Compliance, Vendor Management, Asset Tracking
San Francisco, California, United States · On-site
$58/hr–$63/hr
Entry level$534M raised
About The Role As a Research Intern in the Model Shaping team, you will work on one or more of the following areas: Advanced post-training methods across supervised learning, preference optimization, and reinforcement le…
Skills: Machine Learning, Deep Learning, PyTorch, JAX, Transformer Architecture
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
$200k–$250k/yr
Full-time
Equity, Health Insurance, Remote Work Flexibility
Posted 1d ago
~40 hrs/week
Remote in United States
Responsibilities
Own and improve the end-to-end qualification process for new compute capacity to ensure it meets technical standards. Coordinate with engineering teams to validate hardware, networking, and operational resilience before recommending go/no-go decisions.
Requirements
Requires 5+ years of experience in technical program management within large-scale compute or data center environments. Must possess technical fluency in GPU hardware and networking, along with the ability to analyze data using Python or SQL.
Full job description
About The Role
Together AI is growing its compute footprint, and making sure new capacity meets our technical standards is an important priority for the company. Every new cluster has to clear a technical bar before it carries customer workloads, and this role owns that bar. As Technical Program Manager, Compute Qualification, you will run the process that screens and qualifies prospective compute providers, taking each prospective deployment through a structured evaluation across compute, networking, storage, power, cooling, and operations.
You will coordinate various engineering partners through validation, review provider specifications and test results, and produce clear go/no-go recommendations on whether new capacity meets our standards. It is a high-impact, process-driven role for someone technical enough to know when a spec sheet does not add up, and additional diligence needs to be completed, and organized enough to drive many evaluations to closure in parallel. "You will deep-dive into critical hardware performance metrics, proactively identifying potential bottlenecks in cluster architecture before they impact our end customers training or inference workloads." Conduct diligence and work with engineering teams to make assessments regarding technical and operational resilience.
Responsibilities
Own and continuously improve the end-to-end qualification process for new compute capacity, from initial provider intake through final go/no-go recommendation.
Run multiple provider evaluations in parallel, setting timelines, tracking status, and keeping every stakeholder aligned on what is needed and by when.
Partner with infrastructure engineering, network engineering, data center engineering, and SRE teams to plan and coordinate technical validation, then translate their findings into clear decisions for leadership.
Review provider technical specifications and questionnaire responses for completeness and accuracy, flagging gaps, inconsistencies, and risks that warrant follow-up.
Conduct first-pass analysis of provider data yourself: compare specifications across suppliers , sanity-check performance claims, and surface issues before deeper engineering review.
Maintain the standards, templates, and documentation that define what meets spec across compute, networking, storage, power, cooling, and operational support.
Build a structured, auditable record of evaluation outcomes that informs sourcing decisions and scales the qualification function as the team grows.
Requirements
5+ years in technical program or project management, infrastructure program management, or a comparable technical operations role, ideally involving hardware, data center, or large-scale compute environments.
Proven ability to run multiple complex, cross-functional workstreams to deadline, with strong organization and stakeholder management.
Working technical fluency across data center infrastructure: server and GPU hardware, high-performance networking (InfiniBand or Ethernet fabrics), storage, and power and cooling fundamentals; enough depth to read a detailed technical specification and know what to question.
Hands-on comfort with data: able to write scripts or queries (for example, Python or SQL) to compare, validate, and analyze provider specifications and test results independently.
Excellent written and verbal communication; able to turn dense technical detail into clear recommendations for both engineers and executives.
Willingness to travel to provider and data center sites as needed.
Nice to Have
Experience qualifying, commissioning, or accepting GPU clusters or HPC infrastructure against defined performance and reliability standards.
Familiarity with AI training and inference infrastructure, including interconnect topologies, cluster bring-up, and acceptance testing.
Experience in AI/HPC cluster design.
Background working directly with hardware vendors, colocation providers, or cloud capacity providers.
About Together AI
Together AI is an AI-native cloud company building the infrastructure to make AI faster, cheaper, and more accessible. We’re rapidly scaling our GPU footprint: signing our own data center leases, building large-scale clusters, and expanding toward a global owned-infrastructure presence. Our research team has contributed to breakthroughs like FlashAttention, Hyena, and RedPajama, and we co-design across software, hardware, and algorithms to push the frontier of AI efficiency.
Compensation
We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $200-250K + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.
Equal Opportunity
Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more. Please see our Privacy Policy at https://www.together.ai/privacy
Accelerate inference, model shaping, and pre-training on a research-optimized platform.
Industry
Software Development
Company size
201-500 employees
Founded
2022
Headquarters
San Francisco, California
LinkedIn followers
96,747
Total funding
$1.3B
Together AI is the AI Native Cloud, purpose-built for AI engineers and researchers with a full suite of tooling across inference, model shaping, and pre-training. AI natives can use Together AI as a full-stack AI platform — from a high- performance inference engine built for reliable and fast scaling to on-demand GPU clusters and massive-scale AI factories.
Together AI continuously pushes the frontier forward by productizing cutting-edge research from our world-leading AI systems research team. By combining research velocity with production-grade infrastructure, we enable companies to reliably scale AI-native applications as fast as the field evolves.
Trusted by leading AI natives like Cursor, Decagon, Eleven Labs, AI21, Hedra, and Cartesia, as well as SaaS innovators such as Salesforce, Zoom, and Zomato, Together AI powers the next generation of AI-native applications.
Offices: 251 Rhode Island St, Suite 205, San Francisco, California 94103, US
Accelerate inference, model shaping, and pre-training on a research-optimized platform.
Industry
Software Development
Company size
201-500 employees
Founded
2022
Headquarters
San Francisco, California
LinkedIn followers
96,747
Total funding
$1.3B
Together AI is the AI Native Cloud, purpose-built for AI engineers and researchers with a full suite of tooling across inference, model shaping, and pre-training. AI natives can use Together AI as a full-stack AI platform — from a high- performance inference engine built for reliable and fast scaling to on-demand GPU clusters and massive-scale AI factories.
Together AI continuously pushes the frontier forward by productizing cutting-edge research from our world-leading AI systems research team. By combining research velocity with production-grade infrastructure, we enable companies to reliably scale AI-native applications as fast as the field evolves.
Trusted by leading AI natives like Cursor, Decagon, Eleven Labs, AI21, Hedra, and Cartesia, as well as SaaS innovators such as Salesforce, Zoom, and Zomato, Together AI powers the next generation of AI-native applications.
Offices: 251 Rhode Island St, Suite 205, San Francisco, California 94103, US