Required facts are stored beside observed production answers; this is not universal correctness.
Production evidence labThe benchmark is
The benchmark is
part of the product.
Replay the same hard questions against the live app. Expected facts, real answers, sources, speed, and cost stay together in one receipt.
Loading evaluation results…
LIVE CASES—production cross-source runs
P50 LATENCY—p95 not recorded
PROVIDER PROOF—not recorded
CLAIM RELEVANCE—zero-claim check not recorded
Quick versus Investigate
Not yet comparableUse more reasoning only when it earns its cost.
The comparison has not been claimed.
Run the same frozen questions in both modes against one verified release.
npm run benchmark:live -- --mode fastnpm run benchmark:live -- --mode thinking Judge lens
EVIDENCE BUILDINGEvery scoring claim has a receipt.
Distinct providers appear in the replayable benchmark receipt.
P50 and p95 are measured on the deployed public target.
Average HydraDB calls per answer—not a vague efficiency claim.
The exact benchmark command is published with the results.
The same proof packet is available in web, API, and MCP.
Replay the receipt
Not availablenpm run benchmark:live -- --url https://queueproof.vercel.app