QueueProof
Production evidence lab

The benchmark is
part of the product.

Replay the same hard questions against the live app. Expected facts, real answers, sources, speed, and cost stay together in one receipt.

Loading evaluation results…

LIVE CASESproduction cross-source runs
P50 LATENCYp95 not recorded
PROVIDER PROOFnot recorded
CLAIM RELEVANCEzero-claim check not recorded
Quick versus Investigate

Use more reasoning only when it earns its cost.

Not yet comparable
The comparison has not been claimed.

Run the same frozen questions in both modes against one verified release.

npm run benchmark:live -- --mode fastnpm run benchmark:live -- --mode thinking
Judge lens

Every scoring claim has a receipt.

EVIDENCE BUILDING
01 · FACTS + RELEVANCE

Required facts are stored beside observed production answers; this is not universal correctness.

02 · CROSS-SOURCE

Distinct providers appear in the replayable benchmark receipt.

03 · LATENCY

P50 and p95 are measured on the deployed public target.

04 · COST

Average HydraDB calls per answer—not a vague efficiency claim.

05 · REPRODUCIBILITY1 CMD

The exact benchmark command is published with the results.

06 · DEVELOPER EXPERIENCE3 SURFACES

The same proof packet is available in web, API, and MCP.

Replay the receiptnpm run benchmark:live -- --url https://queueproof.vercel.app
Not available