Skip to content

Careers AI & Agents

AI Engineer, Agents & Retrieval

Build agent loops that can be audited and retrieval that holds up on somebody else’s documents, not on the demo set.

Full-time Hybrid Pokhara, Nepal

Posted 22 Aug 2026

The job

The demo is never the hard part. An agent that calls three tools in a notebook takes an afternoon; the same agent running against a client’s real data, with a budget, a permission boundary and a written answer for what happens when it does something wrong, is the actual job.

You would work on the parts that decide whether an agent is allowed near production: what it is permitted to touch, what is logged, what is reversible, and how anyone would tell afterwards what it did. Retrieval sits alongside that — chunking, embedding and ranking that survives contact with documents nobody cleaned first.

What you would be doing

  1. Design tool boundaries before capabilities: what an agent may call, with what budget, and what the failure path is.
  2. Build retrieval that is evaluated rather than eyeballed — a fixed question set, scored, before and after every change.
  3. Make every run auditable. If nobody can reconstruct what the system did and why, it does not go near a client’s data.
  4. Write the evaluation harness as part of the feature, not after somebody asks for numbers.
  5. Say clearly when a problem does not need a model. Refusing to use one is often the deliverable.

What we need from you

Four or five things, not fifteen. A long list filters out the people who read it honestly and keeps the ones who do not.

  • You have shipped something using an LLM that real people used, and can describe what broke.
  • Strong Python or TypeScript, and enough of the other to read it.
  • You understand embeddings and vector search as engineering rather than as an API call — chunking strategy, recall against a fixed set, and why the naive version disappoints.
  • You are sceptical of your own demos and instrument them accordingly.

What would help, and is not required

  • You have worked with MCP servers, or built tool interfaces for a model against an existing API.
  • You have run pgvector or a dedicated vector store on a corpus large enough for the choice to matter.
  • You have written an evaluation set that later caught a regression.

What you would work with

  • Python
  • TypeScript
  • MCP
  • PostgreSQL
  • pgvector
  • OpenAI / Anthropic APIs
Apply for this role

It takes about ten minutes and goes to a person, not to an applicant tracking system. You will hear back either way.

Also open

Other roles

  1. Engineering

    Engineering Internship

    Six months on real work with a named mentor, real review, and your name on commits that ship.

    Internship On-site Pokhara, Nepal

  2. Engineering

    Mobile Engineer (Flutter)

    Ship apps to two stores from one codebase, including the offline case that nobody specifies until it fails.

    Full-time Hybrid Pokhara, Nepal

  3. Engineering

    Senior Full-Stack Engineer

    Own a product surface end to end — schema, API, interface and the deploy — inside a client’s own repository and to their standards.

    Full-time Hybrid Pokhara, Nepal