A repository is not an answer
Storing documents solves storage. It does not help a new staff member who does not know what the report was called, who wrote it, or that it exists at all.
A secure internal assistant that answers staff questions from your own policies, procedures, research and archives, with every answer linked to the document and page it came from, and a plain refusal when the answer is not in your material.
Storing documents solves storage. It does not help a new staff member who does not know what the report was called, who wrote it, or that it exists at all.
A general-purpose AI will answer confidently whether or not your documents support it. In a regulated or scientific setting, an invented citation is worse than no answer.
Access rules are mirrored from your repository and enforced when content is retrieved, with anything unclear withheld by default. Restricted content gets its own test before go-live. That is designed in from the start, not bolted on.
Everything below is written into the engagement letter before work starts, so you know exactly what you are paying for.
Inventory the content, test extraction on your worst files rather than your best, map the permission model, and agree the test questions with your staff.
Connect the sources, extract and clean the text, and index it for both exact-term and meaning-based search. Technical vocabulary and codes need both.
Build the answering layer with citations and refusals, then measure it against the test set until it clears the agreed threshold.
A small group uses it alongside the old way, feedback is fed back in, then we roll out, train your people and hand over.
Typical internal assistants run R45,000 to R120,000. Large or multimedia archives and fully on-premise deployments are quoted after a short, paid discovery, so the price rests on evidence rather than guesswork.
It is built not to. Answers are composed only from passages retrieved from your documents, every claim carries a citation, and when the material does not support an answer the system says so. We measure this before launch rather than hoping.
It depends on your data residency and cost requirements. We work with commercial models such as Claude and GPT under terms that prohibit training on your data, or open-weight models running entirely on your own hardware.
We never authorise it. We only use model providers whose terms prohibit training on submitted data, we name them in the agreement, and we attach their terms.
It applies, and we treat it seriously. Even a research or policy collection contains personal information: author names, details of customers or growers in records, recordings, and the logs of who asked what. We identify it in discovery, remove it from passages before anything is sent to a model outside South Africa, never identify speakers in recordings, keep usage logs for a defined purpose and period, and work under a written operator agreement.
A 30-minute call is enough to tell whether this is the right build, and whether you should start with an audit instead.