> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pavoai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Index knowledge

> How Pavo reads your sources and extracts typed, cited facts: the raw material for everything else.

Indexing is the first thing System Knowledge does. Once you connect a source, Pavo reads all of it - every code file, warehouse table, dashboard, and past query - and pulls out **facts**: atomic, typed, cited units of knowledge. Everything downstream (the book, the bench, the hub) is built from these facts.

## What indexing does

For each source you connect, Pavo works through every entity it contains and asks the same question: *what did I just learn?* A warehouse table yields its columns, owner, and a semantic description of what it holds. A code file yields how a component works. A past query yields how a metric is actually computed. Each answer is stored as one or more facts, with a citation back to where it came from.

At scale this is a large amount of knowledge: a single connected system commonly produces tens of thousands of facts across its sources.

## Facts: the atomic unit

A fact is the smallest unit of knowledge Pavo stores. Every fact is **typed**, which is what lets the book and bench reason about the right things in the right places.

| Fact type            | Captures                                      | Example                                                        |
| -------------------- | --------------------------------------------- | -------------------------------------------------------------- |
| **Metric fact**      | How a metric is defined, computed, or bounded | "Activation rate excludes users who churned within 24h."       |
| **Data insight**     | Something true about the data itself          | "`events.session_id` is null for \~3% of rows before 2024-06." |
| **Business context** | Why the system or a choice exists             | "Autoplay was launched to lift session depth for new users."   |
| **Platform feature** | How a component or capability works           | "The ranker recently added Thompson sampling for exploration." |

## Every fact is grounded

No fact stands on its own. Each links back, by **citation**, to the exact code, query, table, or record it was extracted from. This is what keeps the knowledge trustworthy: when an agent or a teammate relies on a fact, they can trace it to its source and see whether it still holds.

<Info>
  Grounding is the difference between a knowledge base and a plausible summary. A fact you can trace, you can trust; a fact you cannot, you have to re-verify.
</Info>

## Using the index

The facts Pavo extracts are exposed as a searchable index (a retrieval index over your knowledge). Two ways to use it:

* **In Pavo**: the book, the bench, and every agent draw on the index directly.

<img src="https://mintcdn.com/pavo/1gMFgORUyJjR73jb/images/CleanShot-2026-08-24-at-12.43.48.gif?s=7e8bdbb4d4055c6bbbacdb12cbfe9293" alt="Clean Shot 2026 08 24 At 12 43 48" width="960" height="723" data-path="images/CleanShot-2026-08-24-at-12.43.48.gif" />

* **In your own tools**: query the same index from your coding agents (for example Cursor or Claude Code) so they answer with your system's real facts instead of guesses.

<img src="https://mintcdn.com/pavo/1gMFgORUyJjR73jb/images/image-1.png?fit=max&auto=format&n=1gMFgORUyJjR73jb&q=85&s=aa9929cf1bc1054f25718c8edb86838d" alt="Image" width="2508" height="1344" data-path="images/image-1.png" />

## Coverage and freshness

Indexing is not a one-time event. As your system changes and you do more work, the index needs to stay current and complete.

* **Re-scan** to pull in new sources, or to pick up findings from tasks you have run since the last pass.
* **Coverage** reflects what has been read versus what is connected, use it to spot a source that was connected but never fully ingested.
* What Pavo indexes is scoped to what you grant. A source you did not connect, or files you excluded, are simply absent, not silently guessed at.

<Note>
  If a connected source shows little or no extracted knowledge, treat it as a finding, not a given: it usually means a permission gap or an unreadable format upstream. See [Connectors](/connectors/overview).
</Note>

## Next steps

<CardGroup cols={2}>
  <Card title="Tribal book" icon="book" href="/system-knowledge/tribal-book">
    How facts are distilled into a reviewable account of your system.
  </Card>

  <Card title="Connectors" icon="plug" href="/connectors/overview">
    What each source contributes to the index.
  </Card>
</CardGroup>
