Careers AI & Agents
AI Engineer, Agents & Retrieval
Build agent loops that can be audited and retrieval that holds up on somebody else’s documents, not on the demo set.
Full-time Hybrid Pokhara, Nepal
Posted 22 Aug 2026
The job
The demo is never the hard part. An agent that calls three tools in a notebook takes an afternoon; the same agent running against a client’s real data, with a budget, a permission boundary and a written answer for what happens when it does something wrong, is the actual job.
You would work on the parts that decide whether an agent is allowed near production: what it is permitted to touch, what is logged, what is reversible, and how anyone would tell afterwards what it did. Retrieval sits alongside that — chunking, embedding and ranking that survives contact with documents nobody cleaned first.
What you would be doing
- Design tool boundaries before capabilities: what an agent may call, with what budget, and what the failure path is.
- Build retrieval that is evaluated rather than eyeballed — a fixed question set, scored, before and after every change.
- Make every run auditable. If nobody can reconstruct what the system did and why, it does not go near a client’s data.
- Write the evaluation harness as part of the feature, not after somebody asks for numbers.
- Say clearly when a problem does not need a model. Refusing to use one is often the deliverable.
What we need from you
Four or five things, not fifteen. A long list filters out the people who read it honestly and keeps the ones who do not.
- You have shipped something using an LLM that real people used, and can describe what broke.
- Strong Python or TypeScript, and enough of the other to read it.
- You understand embeddings and vector search as engineering rather than as an API call — chunking strategy, recall against a fixed set, and why the naive version disappoints.
- You are sceptical of your own demos and instrument them accordingly.
What would help, and is not required
- You have worked with MCP servers, or built tool interfaces for a model against an existing API.
- You have run pgvector or a dedicated vector store on a corpus large enough for the choice to matter.
- You have written an evaluation set that later caught a regression.
What you would work with
- Python
- TypeScript
- MCP
- PostgreSQL
- pgvector
- OpenAI / Anthropic APIs
It takes about ten minutes and goes to a person, not to an applicant tracking system. You will hear back either way.
Also open
Other roles
Engineering
Engineering Internship
Six months on real work with a named mentor, real review, and your name on commits that ship.
Internship On-site Pokhara, Nepal
Engineering
Mobile Engineer (Flutter)
Ship apps to two stores from one codebase, including the offline case that nobody specifies until it fails.
Full-time Hybrid Pokhara, Nepal
Engineering
Senior Full-Stack Engineer
Own a product surface end to end — schema, API, interface and the deploy — inside a client’s own repository and to their standards.
Full-time Hybrid Pokhara, Nepal