QueueProof
Production evidence lab

The benchmark is
part of the product.

Replay the same hard questions against the live app. Expected facts, real answers, sources, speed, and cost stay together in one receipt.

Loading evaluation results…

LIVE CASES—production cross-source runs
P50 LATENCY—p95 not recorded
PROVIDER PROOF—not recorded
CLAIM RELEVANCE—zero-claim check not recorded
Quick versus Investigate

Use more reasoning only when it earns its cost.

Not yet comparable
The comparison has not been claimed.

Run the same frozen questions in both modes against one verified release.

npm run benchmark:live -- --mode fastnpm run benchmark:live -- --mode thinking
Judge lens

Every scoring claim has a receipt.

EVIDENCE BUILDING
01 · FACTS + RELEVANCE—

Required facts are stored beside observed production answers; this is not universal correctness.

02 · CROSS-SOURCE—

Distinct providers appear in the replayable benchmark receipt.

03 · LATENCY—

P50 and p95 are measured on the deployed public target.

04 · COST—

Average HydraDB calls per answer—not a vague efficiency claim.

05 · REPRODUCIBILITY1 CMD

The exact benchmark command is published with the results.

06 · DEVELOPER EXPERIENCE3 SURFACES

The same proof packet is available in web, API, and MCP.

Replay the receiptnpm run benchmark:live -- --url https://queueproof.vercel.app
Not available