AGENTIC WORKFLOWS (LANGCHAIN / LANGGRAPH)

Location
Zagreb
Workplace
Remote ok
Compensation
$50 – $100

About this role

ABOUT THE ROLE

Tight Line LLC is a software consultancy that takes on hard problems for enterprise clients. We're looking for a contract AI engineer to work on short-term client projects, typically a few weeks to a few months each, across multiple enterprise engagements. The work is Python-first for agent backends, with JavaScript/TypeScript for UIs and the occasional all-TS build: building AI agents and agentic workflows that run in production, integrating with real CRMs, ERPs, databases, and legacy systems, not proof-of-concept demos.

You'll often be embedded directly with client teams, so this is as much a consulting role as an engineering role. You should be comfortable joining a project mid-stream, earning trust quickly, and leaving the client's team better than you found it.


WHAT YOU'LL DO

Design and build agentic workflows using LangChain, LangGraph, DeepAgents, and LangSmith

Build MCP servers and register them with agents (the MCP client side) so tools and data sources plug in cleanly

Build and harden RAG pipelines end to end: ETL and ingestion, chunking, embedding, vector stores, hybrid retrieval, and evaluation

Instrument and evaluate agent behavior (tracing, evals, regression suites) so "it works" is something you can prove, not just claim

Integrate agents with client systems and data sources, including legacy ones

Pair with client engineers, explain your decisions, and transfer knowledge so clients can own the system after you roll off

Participate in client meetings: demos, status updates, technical discussions with both engineers and stakeholders



WHAT WE'RE LOOKING FOR

Strong Python engineering skills plus solid JavaScript/TypeScript. The LangChain SDK supports both; Python tends to lead for agents and JS/TS for UIs. Either way, you'll be writing production code: typed, tested, and packaged for someone else's team to maintain

Solid CS fundamentals: data structures and algorithms, and a real grasp of concurrency with async/await in both Python and JS/TS

Comfortable building HTTP clients against whatever a client system speaks: REST, JSON-RPC, gRPC, GraphQL

Testing discipline: TDD where it fits, and at minimum fluent with unit, integration, and system tests, including mocking and stubbing

Demonstrated, shippable work building AI agentic workflows. You can show us something you built and talk in depth about the design decisions, failure modes, and how you evaluated it

Strong hands-on experience with the LangChain ecosystem: LangChain, LangGraph, DeepAgents, and LangSmith in particular

LLM fundamentals: clear, direct prompting with well-structured system messages and context, and tool calling with error handling and retries

MCP experience: building MCP servers and registering them with an agent as an MCP client

RAG fundamentals: chunking, embedding, vector stores (Qdrant, pgvector), hybrid queries (BM25), and reciprocal rank fusion, plus the ETL pipelines that feed them

Agent observability and evals: tracing, single- and multi-turn evals, benchmarking, graders in code, LLM-as-judge, trajectory-based evals, and the data flywheels that connect them (offline and online evals, automations that feed an annotation queue). Nobody has evals fully figured out; we want people who treat them as a first-class problem

Comfortable with Docker and Docker Compose for local development, and with at least one cloud provider: object storage, managed databases, and deploying a container

Fluency with AI-assisted development tools (Claude Code, Cursor, Copilot, or similar) as a core part of your workflow, plus good judgment about when to trust them

A consulting mindset: you listen before you prescribe, you scope realistically, you communicate tradeoffs honestly, and you're comfortable with ambiguity and context-switching across engagements

Professional-level English (C1 or better), written and spoken. You'll be communicating directly with clients, including non-technical stakeholders, in meetings and in writing

Working hours that overlap substantially with US Central (Chicago), US Mountain (Denver), US Pacific (San Francisco), or UK (London) business hours


NICE TO HAVE

Experience with CrewAI, AutoGen, LlamaIndex, or other agent frameworks beyond the LangChain ecosystem

Production RAG experience at enterprise scale

Kubernetes (e.g. EKS) and queueing infrastructure (SQS, ElastiCache)

Experience building agents that support skills, in the coding-agent sense

Deep eval-driven development experience beyond the basics (custom harnesses, LangSmith evals, EDD as a working style)

Prior consulting or client-services background


DETAILS

Rate: $50-90/hr USD depending on experience

Engagement: contract, project-based, with potential for ongoing work across multiple client engagements

Location: work from anywhere, as long as your working hours substantially overlap one of the time zones above
Possible travel to client sites if needed

Tired of cold applications?

Sign up with Clera and we'll reach out the moment a role actually fits you — no more spraying applications into the void.

Know someone who'd be great for this?