1. Home
  2. Services
  3. AI development

AI development company in the UK: LLM features, document assistants and agents that do real work

Fixology is a UK AI development company that builds practical AI into business software: LLM features that read and draft documents, assistants that answer from your own files, and AI agents that complete tasks with a person approving the risky steps. Every build is tested against your real data and owned by you, to the standard of the UK's leading bespoke software house.

AI development screenshot
The short answer

Fixology builds AI into the systems you already use: document extraction, drafting, assistants that answer from your own files with sources, and agents that act with a person approving the risky steps. A focused AI feature takes 2 to 4 weeks, an assistant over your documents 6 to 10 weeks, and an agent across several systems 8 to 14 weeks. Every build is scored against your real examples, runs under UK GDPR and belongs to you.

AI features we build for UK businesses

Most useful business AI is not a chatbot on the homepage. It is a large language model (LLM) doing one well-defined job inside a system your team already uses, with the output checked before it matters.

  • Extraction from documents. Reading supplier invoices, delivery notes, CVs, contracts or claim forms and turning them into structured data for Xero, Sage, your CRM or an ERP, with low-confidence fields flagged for a person.
  • Drafting. First drafts of proposals, tender answers, customer replies or case notes, written from your own templates and data, for a member of staff to edit and send.
  • Classification and routing. Sorting a shared inbox or ticket queue by type, urgency and team, and pulling out order numbers and account details on the way.
  • Assistants over your documents. A tool that answers staff questions from your policies, product manuals, past proposals or knowledge base, and shows the source for every answer.
  • Customer service cover. Answering repeated questions on your website or WhatsApp out of hours, and handing anything complex to a named person.

This is AI integration in the practical sense: we build these as part of real software, connected through proper API integrations to the systems that hold your data, not as a separate tool someone has to remember to open.

Assistants that answer from your own documents (RAG)

Retrieval-augmented generation, usually shortened to RAG, is the standard way to make an LLM answer from your information instead of its general training. When someone asks a question, the system first searches your documents for the most relevant passages, then gives only those passages to the model and asks it to answer from them, citing where each point came from.

The model is the easy part. The work that decides whether staff trust the answers is everything around it:

  • Getting the documents in. SharePoint, Google Drive, a document management system, PDFs and scanned files all need different handling. Tables and scanned forms are where weak builds fall over.
  • Splitting and indexing. Documents are broken into sections and stored with embeddings (numeric fingerprints of meaning) in a search index, usually PostgreSQL with pgvector, so there is no extra database to run.
  • Permissions. If a junior member of staff cannot open the HR folder, the assistant must not cite it either. Access rules are applied at search time, before the model sees anything.
  • Freshness. When a policy changes, the index updates within minutes, and old versions stop being cited.
  • Saying "I do not know". If nothing relevant is found, the assistant says so and points to a person, rather than filling the gap with a plausible guess.

AI agents with a human approval step

An AI agent is software that decides which steps to take to finish a task, calling tools such as your calendar, CRM or order system along the way. That is useful and also where the risk sits, because an agent that can act can also act wrongly.

We design every agent around a clear line between what it may do alone and what needs a person. Reading data, drafting and preparing are usually fine to automate. Anything that moves funds, makes a commitment to a customer, sends a new kind of email, deletes records or changes a contract waits in an approval queue, where a member of staff sees exactly what the agent intends to do and why, then approves, edits or rejects it.

Each agent also gets hard limits enforced in ordinary code, not in the prompt: the records it may change, the systems it can touch, the number of actions in an hour. Every step is written to an audit log you can read. As the approval history builds, you can choose to let reliable actions run unassisted, on evidence rather than by default.

How we test AI before it goes live

LLM output varies, so clicking through once proves nothing. We build an evaluation set at the start of every AI project and use it to decide when the feature is good enough.

  1. Collect real examples. Typically 50 to 200 real questions, documents or tickets from your business, with the right answer agreed by someone who knows the work.
  2. Agree the bar. For example: 95 percent of invoice fields extracted correctly, every answer citing a real source, zero answers on topics outside scope.
  3. Score every change. Each new prompt, model or retrieval tweak runs against the whole set, so improvements are measured and regressions caught before release.
  4. Keep testing after launch. Real conversations are reviewed and the difficult ones added to the set.

You see the scores at each two-week demo. If the feature cannot reach the bar you agreed, you find out in week three, not after launch.

Prompt injection, outages and model changes in live AI

An AI feature that passes its tests still has to survive the real world. Three risks are specific to LLM systems, and we design for each of them from the first sprint.

Prompt injection

A document, email or web page can contain text written to hijack the model, such as an instruction to ignore its rules or reveal other data. We treat everything the model reads as untrusted. The model never holds credentials, tools check permissions in ordinary code before acting, and output that would trigger an action is validated against a strict format before anything happens.

Provider outages and slow responses

Model APIs have bad days. Every call has a timeout and a retry, long jobs run on a queue rather than making staff wait, and where it matters a second provider is configured as a fallback. If AI is unavailable, the system degrades to the manual process instead of failing.

Model changes

Providers retire and update model versions regularly. We pin the exact version in use, keep prompts under version control alongside the code, and rerun the evaluation set before any switch. Rate limits stop a runaway loop flooding the provider.

UK GDPR, client data and model training

The consumer chat apps and the business APIs treat your data differently. The business API terms and data processing addenda of the main providers state that what you send is not used to train their models by default. We check the current terms for the provider and model you use, record them, and do not use consumer accounts for client work.

The rest is ordinary UK GDPR done properly:

  • A data protection impact assessment where the system processes personal data in a new way, which most AI projects do. We draft the technical sections for you.
  • UK or EU processing where the provider offers it, for example through Microsoft Azure or AWS regions, and a record of every sub-processor.
  • Personal data removed or masked before it reaches the model when the task does not need it.
  • A person involved in any decision with legal or similarly significant effects on someone, such as credit, hiring or access to a service.

The ICO publishes detailed guidance on AI and data protection, and we have written a plain-English summary in AI and UK GDPR for small businesses. Our wider approach is on the security page.

When AI is the wrong tool

We will tell you when a different approach would do the job better. AI is usually the wrong choice when:

  • The rules are already known. If a person could write the logic as a flowchart (VAT rules, approval limits, discount tables), ordinary code is faster to run and right every time.
  • The answer must be exact every time. Payroll, stock levels and financial reporting need deterministic calculations. An LLM can help explain a figure, not produce it.
  • The real problem is the data. If customer records live in four systems that disagree, an AI layer on top will answer confidently from the wrong one. Fix the data flow first, which is the argument in our piece on data plumbing before agents.
  • Volume is low. If the task happens ten times a week, a better form or template is often the more sensible fix.

In those cases we will usually point you at business process automation instead, which often removes the work that made AI look necessary.

Models and tools we use

We are not tied to one provider. Most work runs on the commercial APIs from Anthropic (Claude), OpenAI and Google (Gemini), chosen per task on accuracy, speed and where the data is processed. For clients who need data to stay entirely within their own infrastructure, open-weight models such as Llama or Mistral can run on private servers, with a trade-off in capability that we spell out before you commit.

We keep the model behind a thin layer in the code, so changing provider is a configuration change plus an evaluation run, not a rewrite. Around it we use the tools you would expect in any well-built system: TypeScript or Python, PostgreSQL, queues for long-running jobs, and n8n where a visual workflow suits the people who will maintain it.

How an AI project runs and what you own

Every AI project starts small, and you speak directly to the AI developers doing the work, not an account manager. A one to two week discovery tests the idea on your real data, agrees the evaluation bar and fixes the scope. We then build in two-week sprints with a demo on a test link and evaluation scores each time, starting with the narrowest useful version and widening it once it proves itself.

You own all of it: the code, prompts, evaluation set, integrations and the provider accounts, which sit in your company's name from day one. Nothing is a black box rented from us. If you later bring the work in-house or move to another supplier, the evaluation set and logs go with it, so whoever comes next can see exactly how well the system performs.

AI development: questions we get asked

How long does it take to build an AI assistant for my business?

An assistant that answers staff questions from your own documents typically takes 6 to 10 weeks, including document ingestion, permissions, an evaluation set and launch. A single AI feature added to an existing system, such as extracting invoice data, usually takes 2 to 4 weeks. An agent working across two or three systems with approval steps takes 8 to 14 weeks.

Will our data be used to train ChatGPT or other AI models?

Not if the system is built properly on business API terms. The main providers state in their business terms that API inputs are not used to train their models by default, which is different from the consumer chat apps. We confirm the current terms for the provider you use, put the account in your company's name, and keep a record of it for your data protection file.

What is the difference between a chatbot and an AI agent?

A chatbot answers questions in a conversation. An agent completes a task: it can check your calendar, look up a customer in the CRM, prepare a booking or order and update records. Because agents can act, we build them with hard limits in code and an approval step for anything that commits funds, makes promises to customers or sends them messages, so a member of staff stays in control.

Can AI be connected to our existing systems like Microsoft 365 or HubSpot?

Yes. Most of our AI work is connected to systems clients already use, including Microsoft 365 and SharePoint, Google Workspace, HubSpot, Salesforce, Xero, Sage and bespoke databases. The AI reads and writes through the same APIs as any other integration, using an account with only the permissions it needs, and every action it takes is logged.

How do you stop AI from making things up?

Three ways. First, the model answers only from documents or data we retrieve for it and must cite its source. Second, it is told to say it does not know when nothing relevant is found, and we test that it does. Third, an evaluation set of real examples is scored on every change, so we can show you the accuracy rate rather than promise it.

Do we need our own AI model trained on our data?

Almost never. Training or fine-tuning a model takes a lot of effort and goes out of date as your information changes. For most businesses, retrieving the right documents at question time (RAG) gives better, more current and more traceable answers. Fine-tuning can make sense for very high volumes of a narrow task, and we would show you the evidence before recommending it.

Is it legal to use AI with customer data in the UK?

Yes, with the same obligations as any other processing of personal data under UK GDPR: a lawful basis, transparency, minimising what you share, security and a data processing agreement with the provider. Many AI projects need a data protection impact assessment, and decisions with significant effects on people need meaningful human involvement. We build these in from the start.

Tell us about your project

Three short steps. You will hear back from a developer, not a salesperson, within one working day. We are happy to sign an NDA first.

Or call 020 7096 2842, Monday to Friday, 9am to 6pm.

Tell us what you want to build

Three quick steps. A senior developer reads every brief and replies within one working day with how we would approach it and a realistic timeline.

What do you want to build?
Call us Start your project
Chat with a developerUsually replies in minutes