Build · Knowledge assistants

Ask your documents. Get cited answers.

A secure internal assistant that answers staff questions from your own policies, procedures, research and archives, with every answer linked to the document and page it came from, and a plain refusal when the answer is not in your material.

From R45,000Timeline 6–10 weeksAnswers Cited to sourceData Stays yours
Why it matters

The knowledge exists. Finding it doesn’t.

Recall

A repository is not an answer

Storing documents solves storage. It does not help a new staff member who does not know what the report was called, who wrote it, or that it exists at all.

Trust

Generic chatbots guess

A general-purpose AI will answer confidently whether or not your documents support it. In a regulated or scientific setting, an invented citation is worse than no answer.

Security

Permissions have to hold

Access rules are mirrored from your repository and enforced when content is retrieved, with anything unclear withheld by default. Restricted content gets its own test before go-live. That is designed in from the start, not bolted on.

What we build

Concrete deliverables, agreed upfront.

Everything below is written into the engagement letter before work starts, so you know exactly what you are paying for.

  • A connection to where your documents live: SharePoint, Google Drive, file shares, institutional repositories such as DSpace, or regular exports
  • Extraction that copes with real files, including scanned PDFs read optically and, where needed, transcripts of video and audio
  • Answers with citations to the specific document and page, so anyone can check the source in one click
  • Permission-aware retrieval: your repository's access rules enforced at retrieval, withheld by default where a permission is missing
  • An evaluation harness: a test set written by your own staff and an accuracy threshold agreed in writing before go-live
  • An admin view showing what is indexed, what failed to process and why, and what staff ask that your documents cannot answer
  • Deployment where you need it: a South African data centre, your own cloud tenant, or your own servers
Process

How it runs

Week 1

Discovery

Inventory the content, test extraction on your worst files rather than your best, map the permission model, and agree the test questions with your staff.

Weeks 2–4

Ingest and index

Connect the sources, extract and clean the text, and index it for both exact-term and meaning-based search. Technical vocabulary and codes need both.

Weeks 4–6

Answer and measure

Build the answering layer with citations and refusals, then measure it against the test set until it clears the agreed threshold.

Weeks 6+

Pilot and go-live

A small group uses it alongside the old way, feedback is fed back in, then we roll out, train your people and hand over.

A good fit
  • Research institutes and knowledge-heavy organisations
  • Professional firms with hundreds of procedures and precedents
  • Manufacturers with manuals, SOPs and safety documents
  • Any organisation that is constantly onboarding staff
Probably not for you
  • A handful of documents: a well-organised shared folder is cheaper
  • Content that changes by the minute, which needs a different architecture
  • Anyone who wants the system to answer beyond what the documents say
Investment
R45,000from

Typical internal assistants run R45,000 to R120,000. Large or multimedia archives and fully on-premise deployments are quoted after a short, paid discovery, so the price rests on evidence rather than guesswork.

Timeline
Six to ten weeks for most builds
Running costs
Model and hosting usage billed at cost, with the usage report attached
Support
Optional monthly maintenance and re-indexing
VAT
Prices exclude VAT where applicable
Questions

Things people ask.

Will it make things up?

It is built not to. Answers are composed only from passages retrieved from your documents, every claim carries a citation, and when the material does not support an answer the system says so. We measure this before launch rather than hoping.

Which AI model does it use?

It depends on your data residency and cost requirements. We work with commercial models such as Claude and GPT under terms that prohibit training on your data, or open-weight models running entirely on your own hardware.

Is our content used to train anyone's model?

We never authorise it. We only use model providers whose terms prohibit training on submitted data, we name them in the agreement, and we attach their terms.

What about POPIA?

It applies, and we treat it seriously. Even a research or policy collection contains personal information: author names, details of customers or growers in records, recordings, and the logs of who asked what. We identify it in discovery, remove it from passages before anything is sent to a model outside South Africa, never identify speakers in recordings, keep usage logs for a defined purpose and period, and work under a written operator agreement.

Tell us what you're trying to fix.

A 30-minute call is enough to tell whether this is the right build, and whether you should start with an audit instead.