Labs is the evidence layer of Brief & Ship. It will publish reproducible website experiments, including prompts, repository state, test protocol, raw outputs, and limitations.
It does not currently publish agent rankings. No controlled multi-agent benchmark has been completed in this repository, so there are no win rates, quality percentages, or performance conclusions to report.
Research standard
Every result-bearing lab must state the question before the outcome. It must preserve the starting brief, agent and model configuration, instruction files, permission mode, number of intervention rounds, build environment, scoring rubric, and artifacts required to reproduce the analysis.
Changes to any of those variables belong in the limitations. A single generated site can reveal a failure mode. It cannot establish a universal success rate.
The research queue includes a same-brief agent comparison, a paired instruction-file study, a JavaScript inventory, and a failure taxonomy for SEO and accessibility defects. Those studies remain unpublished until the runs and raw artifacts exist. The build baseline below is different: it reports measurements we can reproduce from this repository today.
What is available now
The Brief & Ship marketing-site example documents the implementation decisions behind the measured output. It is a case study, not a controlled agent experiment. The comparison library remains qualitative until Labs has evidence strong enough to support empirical claims.