When helping transition multiple services into a support team, I’ve repeatedly encountered the same requirement: support engineers need to get up to speed quickly. They need to understand how each service works, who owns it and what to do when something goes wrong.
Across the projects I’ve worked on, the information they need has often been scattered across Confluence pages, tickets, databases, PDFs and Word documents. Each serves a purpose. The difficulty comes when an engineer needs to pull that information together under pressure.
Imagine an alert arriving at 3 am. A critical service is affected, people are waiting for an update, and the engineer investigating it is also searching for the right runbook. Finding a document is only part of the job. They still need to establish whether it is current and whether its advice applies to this incident.
That pressure is one of the reasons we started exploring knowledge bases and governed retrieval. The same information that helps an engineer respond during an incident can also help a new colleague understand a service during onboarding.
These experiences give us a practical starting point:
Start with the work. Then identify the knowledge required to perform it reliably.
What constitutes a knowledge base?
A knowledge base is a collection of information organised so that people or software can use it to answer questions and complete tasks. For a support team, that might include service ownership, troubleshooting procedures and known issues. For onboarding, it could include the explanations, decisions and working practices a new colleague needs.
Three layers help make this concrete.
Scroll horizontally to read all columns.
| Layer | What it provides | Example |
|---|---|---|
| Source knowledge | The information needed to understand or perform the work | A recovery procedure, architecture decision or service ownership record |
| Storage and organisation | A place to maintain that information and its relationships | Workspace pages, Markdown files, repositories or database records |
| Knowledge service | A way for someone to find and use it | Search, guided onboarding or a question-answering interface |
These layers can exist within one product or span several systems. Their usefulness depends on how well they support the task.
An organised documentation site may be sufficient. Where an AI-assisted interface helps, retrieval-augmented generation, usually shortened to RAG, is one pattern for connecting a model to relevant material. It searches for content and supplies that content to the model as context for its response. Microsoft’s RAG documentation describes that retrieve, augment and generate sequence.
For a buyer, the more useful question is what happens when someone asks for help. Can they find relevant information, understand the answer and check the source?
Start with the questions, then choose the sources
For the engineer in our opening example, the starting questions are straightforward: Who owns this service? Which checks should I run first? What conditions require escalation? Is there an approved procedure for this symptom?
Those questions define an initial source set. Service records, approved runbooks, escalation procedures and selected incident reports may be enough to test the workflow. Start with the smallest source set that supports the workflow. Expand it when unanswered questions or observed failures justify additional knowledge.
Scroll horizontally to read all columns.
| Need | Candidate sources | What to establish before use |
|---|---|---|
| Onboard engineers across services | Service overviews, architecture explanations, ownership records | Audience, owner, current version and required context |
| Find troubleshooting guidance | Runbooks, known issues, selected incident reviews | Applicability, review status and escalation conditions |
| Answer internal policy questions | Approved policies and procedures | Effective date, authority and conflicting versions |
| Help support teams understand a product | Product guides and selected resolved tickets | Accuracy, sensitive content and suitability for reuse |
| Answer questions about current system state | Authoritative operational records or APIs | Access permissions, freshness and the limits of the query |
Current operational questions introduce a separate requirement. “What does the runbook say about this alert?” can be answered from documentation. “Is the service healthy now?” requires current operational data.
This distinction helps contain the first implementation. A service that finds approved procedures can be useful on its own. Connecting it to live operational data is a further design decision, with additional access and reliability requirements.
The same applies to information held in someone’s head. Undocumented decisions first need to be captured before retrieval can expose them. Discovering that gap is valuable: it tells the team which knowledge needs to be captured and who can supply it.
Four ways we are turning information into usable knowledge
The source material already exists in different forms. In work I have built with colleagues, and in a document-processing project I inherited, we have approached the problem through several routes. Each illustrates a different part of the journey from stored information to something a person or an agent can use.
AI Ops: from Confluence to reviewed operational knowledge
Confluence is a workspace where teams create and maintain documentation. In the AI Ops platform I have been developing with a colleague, we have built a path for capturing selected Confluence pages and attachments with information about their source and version. Capture, review and publication into searchable operational knowledge are separate steps. An imported page can remain transition evidence until review approves it as a published knowledge source. The onboarding and question-answering journeys can then use published service knowledge to help engineers understand an unfamiliar service. That separation gives us places to check provenance, sensitive content and readiness. Source-permission and content-change propagation require explicit design and verification.
Uploaded PDFs: creating a searchable evidence collection
In an inherited document-processing project, users can upload PDFs into a pipeline that extracts passages, creates a searchable index and returns matching text with source information. This is a different starting point from a team maintaining a wiki: the knowledge collection grows through documents people submit. The search path returns source passages and metadata for direct inspection. This can be useful in its own right when the user needs to inspect the evidence. File handling, access, extraction quality and traceability all matter before an agent is introduced.
GoldenPath: Markdown with identity and relationships
In GoldenPath, we use Markdown files in Git as part of the knowledge foundation. The documents carry metadata such as an identifier, owner, status and links to related documents. The loader retains that metadata and the source file path; the graph-ingestion code can turn declared relationships into links between documents. That makes it possible to trace a piece of guidance back to a file and follow its relationship to a decision or another document. Git also provides a record of changes. The metadata creates a structure for traceability, while people and validation still need to keep those relationships meaningful. Our knowledge graph and RAG implementation journey explores that metadata foundation in more detail.
Scalet: using the knowledge through an agent
Scalet shows how an agent can consume knowledge held in several places. Its source registry distinguishes repository snapshots from sources queried live, and records what each source covers and how freshness should be understood. Tools search the copied knowledge and query operational records separately. This lets us distinguish an answer drawn from a document snapshot from one based on a current tracker read. Scalet currently provides internal implementation evidence for this approach. Customer outcome evidence will come from future deployments and measurement, while each organisation's context will determine the appropriate implementation.

Figure 1. Scalet’s internal workspace, showing a source query, response and run timeline.
Tools such as Obsidian, Notion, SQL or NoSQL databases and platforms such as Supabase can support different parts of these patterns. The choice follows the source material, access requirements and work to be done. Existing systems may remain the authoritative home of the information. Any searchable copy then needs a defined process for changes, deletions and permission updates. These examples provide inspectable implementation patterns. Measured time savings and production-security evidence require dedicated evaluation.

Figure 2. Three source patterns and the decisions needed to make their content usable. This is a conceptual synthesis across several implementation patterns. Each implementation requires its own designed and verified review and access controls.
What makes knowledge retrieval governed?
Governed knowledge retrieval means finding and using organisational information with explicit rules for access, source authority, maintenance and accountability.
Consider a support engineer asking for a recovery procedure. The service should retrieve material they are permitted to use, identify the procedure it relied on and make relevant review information visible. If documents conflict, the response should surface that conflict. When supporting information has gaps, the response should identify them and give the person a useful next step.
Access needs to be checked before restricted material is supplied to the model. Microsoft’s guidance explicitly calls for access control at retrieval time and treating retrieved content as untrusted input. It also notes that grounded responses can still be inaccurate. A source link gives the reader evidence to inspect, while correctness still requires evaluation. Microsoft RAG security and limitations.
For a first implementation, I would turn those principles into practical checks:
-
Does retrieval give authorised users the procedure and enforce the appropriate access boundary for every other request?
-
Does the cited passage actually support the answer?
-
How does the service respond when the source changes, expires or is removed?
-
Can the service identify gaps or conflicting information?
-
Who receives feedback and owns the correction?
Ownership matters after the initial build. Someone needs to decide whether a procedure remains valid, resolve competing versions and approve changes. A retrieval service makes knowledge easier to reach; a maintenance process keeps it worth reaching for.
A useful first implementation for a team supporting several services
Start with a named group of engineers and one bounded workflow, such as onboarding into service support. Select the service pages and runbooks that group actually needs. Agree on the questions they should be able to answer using them.
A response should explain the relevant guidance, link to the supporting material and make its limits clear. When information is incomplete, it should point to the appropriate owner or escalation route. Keep the engineer in charge of the operational decision.
The AI Ops work has given us concrete ingestion, onboarding and retrieval paths to examine. The next evidence comes from engineers using them: which questions they can answer, where the sources need strengthening and which retrieval paths need improvement. The implementation gives us something to test; user feedback and measured outcomes will tell us how well it supports the job.
That evidence should shape any later move into agent actions. Reading a recovery procedure and restarting a service are different capabilities. Each needs its own permission boundary and tests. OWASP’s guidance on excessive agency recommends limiting available functionality and permissions, and requiring approval for high-impact actions. OWASP excessive agency guidance. We explore the action boundary further in our article on governed agent runtimes.
After organising and publishing the knowledge, the next step is making it useful through search or governed retrieval. A person can use that interface directly. An agent can also use retrieval as a tool when the work calls for it. The knowledge service provides the information foundation; permission to take an operational action is a separate design decision.
If you are considering that next step, try our Agent Test. It helps you examine whether the proposed work needs an agent or could be served by retrieval or a conventional workflow, and surfaces questions about identity, sensitive information and actions. It provides a starting point for the design conversation; evaluation of the actual system establishes the evidence required for implementation.
Test usefulness before expanding
A useful test measures whether someone who knows the work considers the response accurate, supported and actionable.
For a practical first check, choose one service and ask an engineer new to it to establish what it does, who owns it, where to look first during an incident, which recovery guidance is current and when to escalate. Record where they need additional support and who owns the required information. That gives the improvement work a concrete starting point.
Build a small evaluation set from the questions engineers actually ask. For each question, record the relevant source, the expected answer or outcome, and who is qualified to assess it. Include cases involving source gaps, conflicting documents and requests that sit outside the user's access boundary.
Then compare the proposed service with the current way of finding the information. Relevant measures include whether the person reached a usable answer, whether the supporting material was correct, how long the task took and whether another colleague had to intervene.
Keep the comparison fair. An engineer answering a familiar question is different from a new starter encountering it for the first time. Recording that context keeps conclusions proportionate to the evidence available.
Feedback should also explain adoption. When someone returns to the old process, ask why. Perhaps search found the right document more quickly. Perhaps the answer needs greater precision. Perhaps the source set needs additional material. Each observation points to a different improvement.
The result may justify extending the knowledge service. It may show that better documentation and ordinary search are enough. It may reveal that the chosen workflow is too poorly documented to support reliable answers yet. All three outcomes help make the next investment more deliberate.
Where the Knowledge Audit fits
If your team is considering an internal knowledge service, bring one workflow you want to improve and the places people currently search for answers. That gives us a concrete starting point for assessing the knowledge, integrations, controls and evaluation needs involved.
Scaletific’s Knowledge Audit assesses your proposed use case, its knowledge sources and the controls needed to make it useful. It helps establish whether the next step should be better-organised information, governed retrieval or an agent with a clearly scoped role. Where a build is justified, we can scope that work and the ongoing maintenance and assurance separately. The audit can also recommend a simpler approach or establish that the current workflow is best served at its present level of complexity.
The immediate decision is whether the available knowledge can support useful, appropriately controlled work, and what needs to change before it can.
Bring us one workflow your team wants to improve and the places they currently search for answers.