The same saved recipe crossed harness boundaries.
A pantry recipe is a portable artifact. The service returns exact saved code with provenance, and each caller chooses its own isolate. This proof page points at a dogfood note where one recipe was fetched and run from Pi and a plain client process, with the orchestrator leg recorded separately.
Cross-harness dogfood
The field note at .context/loops/recipe-rounds/crossharness-dogfood.md follows one recipe, TitleCaseDemo, through multiple callers. Pi used its pantry tool to list, get, and run the recipe. A plain Bun process used src/client.ts to fetch the same source and execute it in a separate process. The note records the exact outputs and the provenance fields visible during fetch.
Measured token note, bounded to the old eval: pantry can reduce generated output for repeated procedures, while total tokens can rise for small procedures because discovery and tool calls add input overhead. That is not the product claim; the product claim here is artifact custody across harnesses.
What pantry proves at the boundary
The registry boundary is small: list returns metadata without code; get returns the full saved source; run is outside pantry. Runtime extensions live inside one agent. A pantry recipe is a portable artifact any harness can fetch.
Think extensions are runtime-local capabilities for a Think agent; pantry is a runtime-neutral code registry where any harness fetches the exact saved source and chooses its own execution authority.
Eval aside: payload sizes and local execution
This table reports tokenizer-estimated payload sizing for discovery and local eval scaffolding, plus measured local execution for pantry reuse. Provider billing appears only when live mode records provider usage.
| Technique | Tokenizer-estimated payload tokens | Basis | Measured ms | Correctness |
|---|---|---|---|---|
| prompt-from-scratch | 270 | estimate unless live mode records provider usage | not measured unless live | not measured unless live |
| inline-tool-def-each-time | 2,268 | estimate unless live mode records provider usage | not measured unless live | not measured unless live |
| pantry discovery + reuse | 396 | tokenizer-estimated discovery payload only; local run is deterministic code execution | 0 | 4 / 4 measured pass |
Reproduce with bun run evals. Try live usage with LIVE_MODEL=1 bun run evals. The machine-readable file is evals/results.json in the repo.
Limits
- This applies to recurring patterns with a saved recipe. Novel work still needs reasoning.
- There is always discovery cost. The agent must inspect recipe names, descriptions, schemas, and capabilities.
- Tokenizer counts are payload sizes, not provider billing, unless live mode records provider usage.
- The live-provider observation, run through an orchestrator, is exploratory: one prod Kimi K2.7 sample reduced output tokens and raised total tokens for the tiny procedure tested.
- Local deterministic wall-clock is not network latency and not a live model timing result.
- pantry stores and hands back code. Running fetched code remains the caller's trust decision.