Skip to main content
An agent that hands you a confident number is easy. An agent that hands you a number you can act on is the hard part, and it is the whole point of Agent Studio. Every result the fleet produces carries a decision-safe contract: it says what it is, where it came from, what it was checked against, and what it does not prove. This page is that contract, and where the work lives.

What a result is

Not a claim: an account. When a Task Studio finishes, it returns a typed result, not free text. In the LTV program, one task returned a MetricDefinitionSet, another a MeasurementContract, another an ExperimentSpecification. The type tells you what kind of answer it is, and every result of a type carries the same fields, so you always know how to read it. Each result declares four things:
That last field, the honest statement of what a result does not prove, is the one most tools skip and the one that matters most. A measurement of learner value is not permission to change a price. Agent Studio keeps the answer and its authority separate, on purpose.

Keep the three kinds of value apart

The sharpest discipline in the LTV program was a rule the plan set and every result kept: accounting value, predictive value, and causal value are different things, and one may never stand in for another.
  • Accounting: money already earned. A fact about the past.
  • Predictive: a forecast of future behavior. Useful for planning, calibrated and bounded.
  • Causal: evidence that an action changes an outcome. Only a controlled experiment earns this.
The program’s final answer said it plainly: use 730-day discounted contribution as the economic account, cohort-calibrated renewal risk as a planning signal, and randomized experiments for policy, and never use the forecast itself as a treatment rule. Collapsing these is the most common way applied-science work quietly goes wrong; the contract makes it hard to do by accident.
A result only travels once you accept it, and downstream modules may consume only an accepted, versioned result. This is why a long program doesn’t drift: every stage stands on a frozen, signed-off answer, never on another agent’s work-in-progress.

Where the work lives

Everything a program produces stays in its studio, across a few views: Because the files are real, they connect to how you already work: a notebook is a notebook, a dataset is a dataset, and they can be traced or exported. The Output tab: the program's headline, key output, and accepted reports The Files tab: the studio's real tree of plans, notebooks, and datasets

How knowledge compounds

This is the payoff. Because the whole program lives in one studio, work multiplies instead of repeating:
  • Delta, not rerun. Add a new proxy or model and ask the Director to redo the analysis. It already knows the analyses, the findings, and the prior comparisons, so it hands back the same plots with one new row, not a from-scratch pass.
  • Write-back to the hub. Any result worth keeping is pushed to the Knowledge Hub with its evidence chain, shared at the system level. The next program, and the next engineer, starts from it.
  • Lower cost over time. Because consolidated knowledge is reused instead of rediscovered, later agent runs don’t redo settled work, which cuts token spend as much as it saves time.
This is the loop closing. Agent Studio spends the understanding System Knowledge built, and writes what it learns back to it. Each program leaves the system knowing more about itself than the last one found it.

Next steps

Best practices

Habits that turn program volume into compounding knowledge.

Knowledge hub

Where the findings you keep go to live.