Explain the claim
Publish purpose, method, public samples, scoring rules, and limitations.
How the bazaar works
BenchBazaar is an open registry for versioned LLM evaluations—not a single ranking of “intelligence” and not a hosted arbitrary-code execution cloud.
Publish purpose, method, public samples, scoring rules, and limitations.
Public free samples never belong to the official hidden scoring pool.
Every result points to an immutable receipt with exact compatibility facts.
The trust boundary
Official scored prompts and expected answers stay out of public pages, browser queries, source maps, analytics, and result receipts. In the MVP, they remain in an author-controlled runner rather than the web application.
The model endpoint must still receive each evaluation prompt. A malicious or compromised provider can retain it. BenchBazaar therefore says “sealed” or “hidden from public download,” never “impossible to leak.”
Serious substrate
Corrections create successor versions. Receipts are append-only. Scoreboards compare only compatible runs from one exact version and track. Evidence labels say what was actually checked.
Inspect the preview catalog