Public evidence archive / Prototype series 01

Recorded research runs

This report examines how bounded AI systems plan research, route work across specialist tools, preserve evidence lineage, divide authority, and remain inspectable when something fails. These static recordings expose the execution—not merely the final answer.

Everything here is prototype maturity. These two recordings are signed fixture runs: canned inputs, a rule-based planner, and a demo signing key so the files cannot be silently swapped. That is integrity checking, not a production signature, and the numbers are not live market data. Local Qwen or FRED runs stay on the operator's machine; they are not published here. Read the research context.

Run index

Read the completed baseline first, then compare the degraded recording. The pair shows both the intended workflow and the system's behavior when one specialist fails.

Baseline workflow

success replay

The full evidence-to-scenario path: planning, specialist routing, deterministic calculation, evaluation, and a completed record.

  • completed
  • 14 events
  • 5 evidence items

Inspect baseline workflow →

Resilience test

degraded replay

A specialist fails mid-run. The system records the fault, preserves the partial evidence, and completes with an explicit degraded outcome.

  • degraded
  • 17 events
  • 5 evidence items

Inspect resilience test →

What you can inspect

ResearchTR-2026-001Repository