EFFECT SETS · TWO DOCUMENTS, TWO RESULTS
The first scored run of sixteen frozen scenarios against the v1 manifest, hand-labelled before the runner existed. Five failed. The result is never replaced.
First scored run →Run once against v2, a separately versioned label correction in which one label moved: S12’s hold, because started kitchen work is never held. Taken once on the current release, 740a062.
v2 release condition →The original benchmark did not become 16/16. The other four v1 failures were implementation defects, fixed under published SHAs. Both manifests are developer-authored, finite and public, not an independent benchmark. The effect-set CI workflow stays red on purpose, because it judges v1.