> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sourcebot.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Knowledge Base

export const feature_0 = "Knowledge Base"

export const verb_0 = undefined

<Note>
  {feature_0} {verb_0 ?? "is"} only available in a paid plan. Please activate a [license key](/docs/activating-a-subscription) to use this feature.
</Note>

## Overview

Sourcebot generates and maintains architecture documentation for your repositories. A background, offline scan reads each opted-in repository with an LLM. It writes Markdown documents into a managed knowledge repository for that repository.

The documents cite a repository and file together. This lets agents verify a claim with one `read_file` call. Sourcebot indexes the documents like any other repository. Agents discover them through the existing `grep`, `read_file`, and `list_tree` tools in Ask Sourcebot and MCP.

Each knowledge repository is visible exactly when its owner repository is visible.

A full scan organizes documentation around discovered systems:

1. Sourcebot carves the owner's pinned tree into deterministic, weight-budgeted territories.
2. Territory scouts explore in parallel and identify a few load-bearing systems with code evidence.
3. A system reconciler fuses duplicates, folds satellite candidates into parent systems, and rejects low-value groupings.
4. Per-system surveyors prioritize architectural invariants, cross-cutting flows, tricky contracts, and why-decisions.
5. A taxonomist merge agent curates the final work list. Document agents write it, then the README fan-in builds the overview.

Your catalog is an engineer's handbook, not an encyclopedia. The pipeline favors a few dense documents over exhaustive feature coverage.

Three hard caps protect your scan budget. A scan accepts at most 24 systems, 6 proposed documents per system, and 120 create items in its final work list. These caps are backstops, not content targets. Existing documents can still be kept, merged, or deleted without counting toward the create cap.

A document agent must call `finish_doc` successfully for its own path. If it stops without finishing, Sourcebot nudges it to submit before treating that item as failed. Other document items continue running when one item fails.

A delta scan skips territory scouting and reconciliation. It groups affected documents by their existing system folder and runs only the relevant surveyors. Changed owner files that no document cites are sent to every active delta surveyor for adjudication.

## How agents use it

Generated docs appear as a `knowledge-base/` folder inside each repository across search, browse, and agent tools. The folder's `README.md` gives you an overview and an index of the generated documents.

Searches scoped to a repository include its generated docs. Citations use `@file:{repo::path}` syntax.

When a search covers both knowledge documents and ordinary code, knowledge documents receive a large ranking boost so they appear first. This is a strong preference, not a strict guarantee. The streaming web search UI displays results as they arrive from the search backend, so a knowledge document from a slower shard can appear after earlier code results. Non-streaming search surfaces (such as agent tools and the API) sort all results before returning, so knowledge documents rank first reliably.

## Enabling Knowledge Base

Knowledge Base requires all of the following:

* Your license includes the `knowledge-base` entitlement.
* Your config file's top-level `models` array contains at least one entry.
* An owner turns on **Enabled** on **Settings → Knowledge Base**.
* An owner then chooses repositories with **Edit repositories** on the same page.

Sourcebot does not scan any repository until Knowledge Base is on and an owner chooses repositories.

Turning Knowledge Base off pauses all scans. Docs that Sourcebot already generated stay searchable. Use a repository's **Delete docs** action to remove them. While Knowledge Base is off, you can't change repositories or select **Generate now**.

If your license no longer includes the `knowledge-base` entitlement, generated docs stay searchable, but Sourcebot stops updating them.

## Configuration

Owners tune generation on **Settings → Knowledge Base → Configuration**. Sourcebot stores these settings in its database, so you don't edit your config file to change them.

| Setting | Default | Description |
| - | - | - |
| **Scan mode** | **Delta** | Chooses between **Delta** and **Full** scans. See [Scan modes](#scan-modes). |
| **Refresh interval** | 24 hours | Sets how often Sourcebot checks each repository for changes. |
| **Model** (per stage) | The first entry in your `models` array | Selects the model for a stage. See [Per-stage models and reasoning effort](#per-stage-models-and-reasoning-effort). |
| **Reasoning effort** (per stage) | **Automatic** | Sets the reasoning effort for a stage. |
| **Token budget per scan** | 50,000,000 | Sets the maximum number of tokens that one repository scan newly processes. Prompt-cache reads do not count against this budget. |
| **Parallel tasks per scan** | 4 | Sets how many scouts, surveyors, or document agents run concurrently in one scan. |

Models and their credentials stay in your config file. The configuration page lists the entries in your top-level `models` array.

The `settings.maxKnowledgeScanJobConcurrency` config setting controls how many repositories scan concurrently. Its default is `1`. Increasing it multiplies your possible per-job spend by the number of concurrent jobs.

## Per-stage models and reasoning effort

Knowledge generation runs in four stages. You can choose a model and reasoning effort for each one:

* **Discovery** runs territory scouts and the system reconciler.
* **Survey** surveys each canonical system and merges the work list.
* **Document generation** writes each knowledge document.
* **Overview** writes the README overview.

A stage without a chosen model uses the first entry in your `models` array, the same default that Ask Sourcebot uses. To change that default, reorder your `models` array.

**Discovery** and **Overview** inherit each setting from **Survey** separately. If you choose only a Discovery model, it still uses the Survey effort. If you choose only a Discovery effort, it still uses the Survey model. **Document generation** does not inherit from another stage.

Providers map effort values as follows:

| Provider | Supported values | Mapping |
| - | - | - |
| Anthropic | `low`, `medium`, `high`, `xhigh`, `max` | Uses the value directly with adaptive thinking. |
| OpenAI and Azure | `low`, `medium`, `high`, `xhigh` | Uses the value directly. `max` clamps to `xhigh` and logs a warning. |
| Google (`google-generative-ai`) and Google Vertex (`google-vertex`) | `low`, `medium`, `high` | Uses the value as the thinking level. `xhigh` and `max` clamp to `high` and log a warning. |
| OpenAI-compatible | `low`, `medium`, `high`, `xhigh` | Passes the value through. `max` clamps to `xhigh` and logs a warning. |
| All other providers | None | Rejects the scan before it starts. |

Sourcebot resolves all stage settings before it starts the scan. A missing model, unsupported provider, or model setup failure stops the whole scan before it spends tokens.

Anthropic models without adaptive thinking log a warning and run without the configured effort. They keep their existing thinking behavior.

Usage rows appear separately for each model. Each generated document records the model that wrote it. `README.md` records the `readme` model.

## Document naming contract

Every newly planned document uses exactly `system-slug/concept-slug.md`. Both segments use lowercase kebab-case with at most 64 characters. The folder names the canonical system, and the file names a specific concept.

The concept name cannot repeat its system folder. Sourcebot also refuses these generic concept names after ignoring case and punctuation: `overview`, `readme`, `index`, `doc`, `docs`, `documentation`, `about`, `misc`, `general`, `summary`, `notes`, `info`, `main`, `other`, `todo`, `claude`, and `agents`.

Existing documents at legacy paths remain editable, finishable, and deletable. The stricter naming contract applies only when surveyors or the merge agent propose a new path.

## Repository selection and the settings page

Open **Settings → Knowledge Base** and choose one of these modes:

* **All repos** includes every standard repository. New repositories join automatically. Use a repository's **Delete docs** action to remove its generated documentation and exclude it.
* **Selected repos** gives you a searchable picker for an explicit repository list.

Sourcebot scans a newly selected repository as soon as it is indexed. It does not wait for the first scheduled refresh. Select **Generate now** to scan every selected repository immediately.

Each repository shows what is happening right now, with a detail line underneath:

| Status | Meaning |
| - | - |
| **Waiting for index** | The repository has not finished its first code index. Generation starts after it does. |
| **Setting up** | Sourcebot is creating the repository's knowledge repository. |
| **Queued** | A scan is waiting to run. The detail shows its position in the queue. |
| **Generating** | A scan is running. The detail shows the current stage, such as **Writing docs 7/23**. |
| **Retrying** | A scan attempt failed and will run again. Open the icon to see the error. |
| **Indexing for search** | The docs are written. Search picks them up once indexing finishes. |
| **Up to date** | The latest scan succeeded. The detail shows the number of docs. |
| **Not generated** | No scan has run yet. The detail shows when the next scheduled scan runs. |
| **Failed**, **Setup failed** | The latest scan or setup failed. Open the icon to see the error and logs. |

Expand a repository to see the current or latest run, its token usage by model, and how much of the **Token budget per scan** it used.

If a repository already has a real root entry named `knowledge-base`, the settings page shows a conflict state. Generation pauses for that repository until you rename or remove the entry.

The per-repository **Delete docs** action removes the repository's knowledge repository, clone, and shards. In **All repos** mode, it also excludes the repository until you select **Re-include**.

## Scan modes

The **Delta** mode is the default. It updates only documents whose cited files changed. It falls back to a full scan for a large diff, missing history, or a last full scan older than 7 days.

The **Full** mode re-verifies every document during each scan cycle. A repository's first scan is always a full scan.

Repositories that cite busy shared dependencies re-scan whenever those dependencies change. The **Delta** mode keeps that work cheap.

## Scheduling and cost

Scheduled scans refresh according to the **Refresh interval** setting. The default interval is 24 hours. Sourcebot uses the citation index to skip repositories with no relevant changes.

Every job records input, output, cache-read, and cache-write tokens for each model. You can see this usage on **Settings → Knowledge Base**. The **Token budget per scan** setting bounds the tokens that your job newly processes. Prompt-cache reads do not count against this budget.

Sourcebot caches the agent conversation with supported providers so your long scans re-read context from cache. If your scan exhausts its budget or encounters a structural failure, it fails immediately without a retry.

Your scan model calls use a 45 minute transport timeout. Reconciliation and other single-step generations can legitimately exceed typical HTTP header timeouts.

## Notes

* Sourcebot must index a repository before Knowledge Base can scan it.
* Generated documents go live as soon as Sourcebot commits them. There is no review queue.
* Use the per-repository **Delete docs** action as your escape hatch.
* Cloud-licensed deployments receive the entitlement from the license service. Operators must ensure that the entitlement service knows `knowledge-base` before this release rolls out.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.