About this role
- Own ML/AI systems end-to-end: data pipelines, model
training, serving infrastructure, monitoring, and iteration
- Build LLM-powered applications with custom pipelines,
prompt management, evaluation, and optimization
- Implement multi-agent orchestration systems using
LangGraph, CrewAI, or AutoGen for autonomous workflows
- Build and optimize RAG pipelines using LlamaIndex with
chunking strategies, embedding selection, re-ranking, and evaluation
- Deploy and manage LLM inference infrastructure using vLLM
or Ollama for on-premise sovereign deployments
- Build traditional ML scoring models: churn prediction,
propensity scoring, LTV estimation, next-best-action
- Design and build feature pipelines using Apache Flink
(streaming) and Spark (batch) for real-time and batch ML
- Implement MLOps practices: model versioning, registry,
drift monitoring, A/B testing, and staged rollouts
- Design and implement AI operators for visual low-code
canvas (LLM Gateway, RAG Pipeline, Intent Classifier)
- Optimize ML inference for latency and throughput at scale
(10K+ QPS)
- Collaborate with Data Engineering and Platform teams to
integrate ML systems with data infrastructure
Requirements
- 3+ years of hands-on ML/AI engineering with demonstrated
end-to-end system ownership
- Production experience building LLM-powered applications
(not just API consumption)
- Hands-on experience with agent orchestration: LangGraph,
CrewAI, or AutoGen in production
- Production RAG experience with evaluation metrics, hybrid
search, and re-ranking strategies
- Experience building ML models: churn, propensity, LTV,
segmentation, recommendation systems
- Hands-on experience with data pipelines: Spark for batch,
Flink or Kafka Streams for real-time
- Strong Python proficiency: production code structure,
async, multiprocessing, profiling, optimization
- Experience with vector databases at scale: OpenSearch
k-NN, Qdrant, or Milvus
- Production MLOps experience: MLflow, experiment tracking,
model registry, drift monitoring
- Real-time ML inference experience at 1,000+ QPS
Good to Have:
- Experience at AI-first companies or building AI/ML
platforms from scratch
- Telco or enterprise data platform background
- Experience with LLM fine-tuning: LoRA, QLoRA, PEFT
techniques
- Experience with embedding models: sentence-transformers,
fine-tuning for domain
- Kubernetes for ML workload orchestration and GPU
scheduling
- Knowledge of PII detection (Presidio) and LLM guardrails
(NeMo Guardrails)
Tired of cold applications?
Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.
Know someone who'd be great for this?