Skip to main content
Fugen Services logo

AI & Automation

Search across everything your company knows

Organisational knowledge tends to live in the worst possible places: a shared drive, a decade of email, PDFs nobody has opened since signing. Retrieval-augmented generation makes that searchable by meaning rather than by filename, and returns answers that cite their sources.

Indicative

From £9,000

Fixed price agreed in writing before any build starts.

Get a quote+44 7488 265083

Background photo by Eric Lozaga on Pexels

The problem this solves

Keyword search fails because people do not remember the words a document used. So the same questions get asked in Slack, experienced staff become bottlenecks, and decisions get made on half-remembered detail. The knowledge exists; it is just unreachable.

What you get

Document ingestion pipeline

PDFs, Word files, spreadsheets, wikis, tickets and email processed with layout-aware parsing, so tables and structure survive rather than collapsing into noise.

Chunking and embedding strategy

Chunk sizes and overlap tuned to your document types and evaluated — this single decision drives most of the accuracy difference between a good RAG system and a useless one.

Hybrid retrieval

Semantic vector search combined with keyword matching and re-ranking, because pure vector search reliably misses exact identifiers like part numbers and clause references.

Permission-aware results

Retrieval filtered by the user’s access rights, so an AI search layer cannot become an accidental data-leak channel.

Cited answers

Responses quote and link the source passage, making verification possible and building the trust the system needs to actually get used.

Freshness and re-indexing

Scheduled and event-driven re-indexing with change detection, so superseded policy documents stop being quoted as current.

How we work

  1. Corpus survey

    We inventory what documents exist, in what formats, and where the authoritative versions live — often the hardest part of the project.

  2. Retrieval evaluation set

    Real questions with known correct source documents, so retrieval quality is measured rather than assumed.

  3. Pipeline build and tuning

    Ingestion, chunking, embedding and re-ranking, iterated against the evaluation set until precision is acceptable.

  4. Interface

    A search interface, an internal assistant, or an API your existing tools call — whichever puts answers where the work happens.

  5. Rollout with feedback capture

    Thumbs-up/down on answers feeding a review queue, so retrieval keeps improving on real usage.

What you should expect

  • Measured retrieval precision on your own question set
  • Answers cite and link their source passage
  • Access control respected at retrieval time
  • Superseded documents stop surfacing as current guidance

Built with

  • pgvector
  • Qdrant
  • Mistral
  • Claude
  • Python
  • TypeScript
  • PostgreSQL
  • Redis
  • Docker
  • Unstructured

Mainstream, well-supported technology — chosen so you can hire for it and so another team could take the project over.

RAG & Knowledge Bases — your questions

Including the ones about cost, which most agencies leave off the page.

A general assistant knows nothing about your contracts, your pricing or your internal policies, and will happily improvise when asked. RAG retrieves the relevant passages from your own documents first and constrains the answer to them, with citations. The difference is verifiability.

Yes. We connect to SharePoint, Google Drive, Confluence, Notion and plain file shares, and we honour the source system’s permissions so users only retrieve what they were already allowed to see.

That depends on your documents, which is why we build an evaluation set before quoting on the full system. Well-structured corpora commonly reach 85–95% retrieval precision on realistic questions. Poorly scanned PDFs and inconsistent terminology bring that down, and we would rather tell you that during assessment than after delivery.

Wherever you require. We can keep embeddings and documents entirely within UK or EU regions, or fully on your own infrastructure. Only the retrieved passage is sent to the model at query time, and with enterprise terms that content is excluded from training.

Talk to someone who has built this before

A short call is usually enough to tell you whether this is the right service for your situation — including when it is not.