What we build, and what it looks like in a real system

Each section below states the work, the design decisions that matter and the kinds of task it covers. Nothing here is a case study: the studio publishes no client names, logos or numbers, because none have been supplied as verifiable.

01

LLM integration into the systems you already run

A model is the smallest part of the work. The rest is the seam into your ERP, CRM, ticketing or document store.

Most of the effort goes into the part that is not the model: where the answer appears in the tool people already use, what permissions they keep, what happens when the model is wrong, and who is accountable for the output.

We build against your existing systems through their APIs rather than replacing them. If a step can be solved with a rule, a query or a form, it stays a rule — a model is used where language or unstructured documents genuinely are the problem.

Typical tasks

  • Answer drafting inside a case-management tool, with the source documents attached to every draft
  • Classification and routing of inbound tickets, claims or applications, with a review queue for low confidence
  • Summaries of long records, each statement linked back to the page it came from

02

Retrieval over internal documents

Answers grounded in your own corpus — policies, contracts, manuals, tickets — with retrieval quality measured separately from writing quality.

Retrieval decides whether the answer is possible at all; generation only decides how it reads. We build ingestion, chunking, indexing and permission-aware filtering, then measure whether the right passage was found before looking at the wording.

Access rules are applied at retrieval time, so a user cannot receive a passage from a document they are not allowed to open. Where a source is scanned or inconsistent, that is handled during ingestion and reported rather than hidden.

Typical tasks

  • Policy assistant for staff, restricted by department and seniority
  • Contract review support with clause-level references and a checklist for the reviewer
  • Technical search across drawings, PDFs and spreadsheets that were never written to be searched

03

Agents and process automation

Multi-step work with tools: reading a queue, calling internal APIs, filling forms, escalating to a person.

An agent is a sequence of decisions with side effects, so the design question is which decisions may be made without a person. We keep irreversible steps behind explicit approval and make every step visible in a log.

Automation is delivered with the failure path first: what happens when an upstream system is down, returns nonsense, or the business rule is ambiguous. Those paths are specified before the happy path is built.

Typical tasks

  • First-line triage that drafts a reply and hands it over with the full context
  • Reconciliation between two systems, with differences queued for a human decision
  • Scheduled reporting assembled from several sources, with the source of every figure recorded

04

Evaluation and hallucination control

A model that answers fluently is not a model that answers correctly. Quality is measured on your own hard cases, release after release.

We build a test set from your operation: real questions, real documents, the cases your staff find difficult. Retrieval and answer quality are scored separately so a regression points at the right layer.

Refusal and escalation rules are part of the product, not an afterthought: the system must know what it does not know and route that case to a person with the context attached.

Typical tasks

  • Regression suite over a labelled set of questions, run on every change
  • Citation checks that block an answer whose sources do not support it
  • Regular reporting of failure cases by category, with the worst examples named

05

Data privacy and regulatory requirements

Where the data lives, what leaves the perimeter, what is retained and what is logged — written so compliance can review it.

For corporate deployments the constraint usually decides the architecture, not the other way round. We document the data flow, the retention rules and the sub-processors before writing code, and choose hosting accordingly.

Where a requirement forbids external model calls, open-weight models run inside your perimeter. That choice has an accuracy cost, and it is stated in writing rather than discovered later.

Typical tasks

  • Processing inside your cloud tenant, with no vendor-side retention
  • Redaction or pseudonymisation before any external call
  • Per-request audit trail of prompts, retrieved passages and outputs

06

Support after launch

Models and documents change. Either a clean hand-over with documentation and monitoring, or a monthly engagement that keeps quality measured.

A deployment is not finished when it works on the day of the demo: documents are re-issued, providers deprecate model versions, and the questions people ask drift. Someone has to notice.

Hand-over includes runbooks, monitoring and a re-evaluation procedure your own team can run. If you would rather not run it, the same work continues as a retainer with a named engineer.

Typical tasks

  • Monitoring for retrieval drift and refusal spikes
  • Re-indexing when the document set changes
  • Quarterly evaluation report with the failure cases named