Skip to main content
A book can read well but it’s not useful until it’s factually correct.
The knowledge bench is how you find out. It is a benchmark of pointed questions about your system, scored by an independent judge, that tells you how correct and complete the current knowledge is, before an agent acts on it.

Anatomy of the bench

The bench has four parts. Understanding them makes the scores easy to read.

What it measures

Knowledge can fail in two different ways, and the bench measures both:
  • Coverage: can the knowledge answer the question at all? This is a matter of degree: you might answer 40% of the questions today and 90% next week. Coverage is scored as a number.
  • Correctness: is the answer right? This is a pass/fail judgment on the answers you do have.
Keep the two apart. High coverage with low correctness is dangerous: confident, wrong answers. High correctness with low coverage is safe but incomplete. You want both climbing together.

The independent judge

Each question is scored by a judge that is separate from whatever produced the knowledge, so the system is not grading its own work. The judge reads the question, checks the compiled knowledge, and returns a verdict with its reasoning attached, so a score is never a black box. You can open any verdict and see why it landed where it did.

Human escalation

The judge does not guess when it cannot tell. When a question falls outside what the judge can confidently resolve: the knowledge is ambiguous, or the question’s rubric does not cleanly apply, it escalates to a human rather than inventing a score. This keeps the bench score honest: a number you can trust, because the uncertain cases were adjudicated by a person, not auto-scored.
Escalation is a feature. The questions the judge routes to you are exactly the ones where your judgment adds the most, and answering them sharpens both the knowledge and the bench.

Managing questions

Pavo generates a first version of the bench so you have something to score against on day one. But you own the question set, it should encode the gotchas your team cares about.
1

Review Pavo's questions

Read each question and the judge’s answer. Confirm the answer is right, or mark it wrong.
2

Add the gotchas you know

Add questions for the mistakes you have seen people make, the wrong join, the metric variant, the edge case. You can seed these from files you upload.
3

Edit or remove the misses

Delete questions that do not matter and refine ones that are imprecise, or open the task that generated them to steer it.
In most engagements, the highest-leverage thing a team does in the first couple of weeks is sit together and make the bench collectively represent the gotchas everyone knows. Correct knowledge is the bottleneck to useful agents, and the bench is how you enforce it.

Reading bench scores

A bench score is not just a grade: it is a to-do list. When knowledge scores 87 out of 100, the interesting question is which 13 it missed and what would close them. Pavo and its agents use the failed questions to go back and improve the knowledge: fill a coverage gap, correct a wrong fact, resolve an ambiguity in the book. Re-score, and watch the number move for a reason you can name.

Next steps

Tribal book

Fix what the bench surfaces as wrong or missing.

Best practices

How to keep vetting the bench as part of your routine.