> ## Documentation Index
> Fetch the complete documentation index at: https://docs.pavoai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Knowledge bench

> A benchmark that scores how correct and complete your system knowledge is, before you rely on it.

A book can read well but it's not useful until it's factually correct. <br />The **knowledge bench** is how you find out. It is a benchmark of pointed questions about your system, scored by an independent judge, that tells you how correct and complete the current knowledge is, before an agent acts on it.

## Anatomy of the bench

The bench has four parts. Understanding them makes the scores easy to read.

| Part                  | What it is                                                                                                                          |
| --------------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| **Questions**         | Specific, checkable questions about your system, for example, "at what exact granularity is the canonical activation rate counted?" |
| **System under test** | The current compiled knowledge, the facts and the book, that has to answer them.                                                    |
| **Judge**             | An independent model that scores each answer against the question.                                                                  |
| **Escalation**        | The path for questions the judge cannot settle on its own, routed to a human.                                                       |

## What it measures

Knowledge can fail in two different ways, and the bench measures both:

* **Coverage**: *can* the knowledge answer the question at all? This is a matter of degree: you might answer 40% of the questions today and 90% next week. Coverage is scored as a number.
* **Correctness**: is the answer *right*? This is a pass/fail judgment on the answers you do have.

<Info>
  Keep the two apart. High coverage with low correctness is dangerous: confident, wrong answers. High correctness with low coverage is safe but incomplete. You want both climbing together.
</Info>

## The independent judge

Each question is scored by a judge that is separate from whatever produced the knowledge, so the system is not grading its own work. The judge reads the question, checks the compiled knowledge, and returns a verdict with its reasoning attached, so a score is never a black box. You can open any verdict and see why it landed where it did.

## Human escalation

The judge does not guess when it cannot tell. When a question falls outside what the judge can confidently resolve: the knowledge is ambiguous, or the question's rubric does not cleanly apply, it escalates to a human rather than inventing a score.

This keeps the bench score honest: a number you can trust, because the uncertain cases were adjudicated by a person, not auto-scored.

<Note>
  Escalation is a feature. The questions the judge routes to you are exactly the ones where your judgment adds the most, and answering them sharpens both the knowledge and the bench.
</Note>

## Managing questions

Pavo generates a first version of the bench so you have something to score against on day one. But you own the question set, it should encode the gotchas *your* team cares about.

<Steps>
  <Step title="Review Pavo's questions">
    Read each question and the judge's answer. Confirm the answer is right, or mark it wrong.
  </Step>

  <Step title="Add the gotchas you know">
    Add questions for the mistakes you have seen people make, the wrong join, the metric variant, the edge case. You can seed these from files you upload.
  </Step>

  <Step title="Edit or remove the misses">
    Delete questions that do not matter and refine ones that are imprecise, or open the task that generated them to steer it.
  </Step>
</Steps>

<Tip>
  In most engagements, the highest-leverage thing a team does in the first couple of weeks is sit together and make the bench collectively represent the gotchas everyone knows. Correct knowledge is the bottleneck to useful agents, and the bench is how you enforce it.
</Tip>

## Reading bench scores

A bench score is not just a grade: it is a to-do list. When knowledge scores 87 out of 100, the interesting question is *which* 13 it missed and what would close them. Pavo and its agents use the failed questions to go back and improve the knowledge: fill a coverage gap, correct a wrong fact, resolve an ambiguity in the book. Re-score, and watch the number move for a reason you can name.

## Next steps

<CardGroup cols={2}>
  <Card title="Tribal book" icon="book" href="/system-knowledge/tribal-book">
    Fix what the bench surfaces as wrong or missing.
  </Card>

  <Card title="Best practices" icon="star" href="/system-knowledge/best-practices">
    How to keep vetting the bench as part of your routine.
  </Card>
</CardGroup>
