About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Skills: Data Pipeline Design, Machine Learning Infrastructure, Spark, Beam, Flink
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Skills: Backend Engineering, API Design, OAuth Flows, Webhooks, Integration Frameworks
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Skills: Systems Engineering, API Design, Infrastructure, Concurrency, LLMs
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Skills: React, State Management, Websockets, Real-time UI, Frontend Engineering
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Skills: Swift, UIKit, SwiftUI, iOS System Design, Concurrency Patterns
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Skills: Reinforcement Learning, PyTorch, Python, Large Language Models, Reward Modeling
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Skills: Machine Learning, Reinforcement Learning, PyTorch, Python, Distributed Training
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Skills: Multimodal Learning, Computer Vision, Video Modeling, Generative AI, Distributed Training
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Skills: Large-scale Pretraining, Multimodal Models, Distributed Training, Data Curation, Model Architecture
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Skills: Speech Recognition, Speech Synthesis, Multimodal Foundation Models, ASR, TTS
Member of Technical Staff, Multimodal Post-train/RL
San Jose, California, United States · On-site
$180k–$450k/yr
Senior
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Skills: RL Algorithms, PPO, GRPO, RLHF, Multimodal Foundation Models
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Skills: Multimodal AI, Neural Networks, Distributed Machine Learning, Data Curation, Synthetic Data Generation
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
Skills: Data Collection, Vendor Management, Program Management, Data Operations, Quality Assurance
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
About Hark Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persisten…
About Hark Hark is an artificial intelligence laboratory, building the most advanced, personal intelligence in the world. We believe that artificial intelligence can be used to help offload our mental workload through th…
Skills: Brand Building, Brand Design, Creative Direction, Visual Design, Film Production
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
$170k–$450k/yr
Full-time
Posted 88d ago
~40 hrs/week
Responsibilities
Design and build scalable end-to-end data infrastructure to process multimodal training data for AI models. Collaborate with researchers to ensure data quality, reproducibility, and efficient delivery to training systems.
Requirements
Requires 5+ years of data engineering experience with proficiency in the modern data stack and ML-specific pipelines. Candidates should possess strong systems thinking and experience with large-scale batch and streaming processing.
Full job description
About Hark
Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.
We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.
To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.
About the Role
You'll build the data infrastructure that turns raw signals into the training data Hark's models learn from, and the pipelines that keep it flowing at scale.
That means owning the full data engineering stack: ingestion, transformation, quality filtering, and delivery to training and evaluation systems. The models we ship are only as good as the data behind them, and this role owns that foundation.
This is a high-ownership role on a small team. You'll work directly with model researchers, data collection leads, and infrastructure engineers, and the systems you build will directly shape the quality and pace of model development.
Responsibilities
Design and build scalable data pipelines that ingest, process, and deliver training data across multiple modalities: text, audio, vision, and structured feedback signals.
Own the data infrastructure stack end-to-end: ingestion, transformation, deduplication, quality filtering, versioning, and delivery to model training and evaluation systems.
Collaborate closely with model researchers and data collection leads to understand data requirements and translate them into reliable, auditable pipelines.
Build tooling and frameworks that make it easy for the team to inspect, evaluate, and iterate on data quality. The insights surfaced should feed back into collection and curation decisions.
Define and enforce data quality standards. Instrument pipelines for correctness, freshness, and coverage. Catch regressions before they reach training.
Design data systems for reproducibility and scale. The pipelines you build need to handle growing volumes across modalities without becoming a bottleneck.
Identify gaps in the current stack and drive concrete improvements to throughput, quality, and reliability.
Requirements
Strong data engineering fundamentals. You are comfortable designing and operating large-scale batch and streaming pipelines, and you care about correctness and reliability.
Experience building data systems for machine learning. You understand the difference between a data pipeline for analytics and one that feeds model training, and you know what it takes to get the latter right.
Fluency with the modern data stack. You've worked with tools like Spark, Beam, or Flink, and you know how to make tradeoffs between them. Experience with data versioning systems (e.g., DVC, Delta Lake, Iceberg) is a strong plus.
Systems thinking. You reason about schema evolution, backfills, and failure modes before they become production incidents. You build for the day-2 case, not just the demo.
A quality instinct. You don't just move data. You understand what's in it, catch problems early, and close the feedback loop with the people who need clean data.
Strong communication. You can work closely with model researchers and engineers, explain data tradeoffs clearly, and make good decisions across team boundaries.
5+ years of relevant data engineering experience. Experience at a fast-growing AI or research-driven company is a strong plus.
Bonus Qualifications
Experience building data infrastructure for large language model or multimodal model training.
Familiarity with multimodal data formats and processing pipelines (audio, video, image).
Experience with human feedback or preference data pipelines (RLHF, DPO, or similar).
Hands-on experience with data quality evaluation frameworks or annotation tooling.
Background in distributed systems, stream processing, or large-scale ETL.
Experience at a fast-moving AI lab or research-driven company.
Compensation
The US base salary range for this full-time position is between $170,000 - $450,000 annually.
The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.
Building the Intelligence Behind the Next Generation of Products.
Industry
Business Consulting and Services
Company size
2-10 employees
Hark Labs is a GenAI and product innovation studio helping startups and enterprises harness the power of artificial intelligence. We design smarter products, automate complex workflows, and turn data into dynamic intelligence — bridging the gap between cutting-edge AI research and real-world business impact.
Building the Intelligence Behind the Next Generation of Products.
Industry
Business Consulting and Services
Company size
2-10 employees
Hark Labs is a GenAI and product innovation studio helping startups and enterprises harness the power of artificial intelligence. We design smarter products, automate complex workflows, and turn data into dynamic intelligence — bridging the gap between cutting-edge AI research and real-world business impact.