SecondRunStanding open challenge · release 1.4.0

Does the rating predict the job?

ClusterMAX 3.0 is out. Test whether its medals improve predictions of completed, accepted customer work within cost and latency limits, beyond specifications and price alone.

Customer-outcome test: awaiting real submissions.

The scorer works now. Run the labeled synthetic example, or bring a design frozen before the jobs, its rating binding and complete outcomes. The separate historical study remains inconclusive.

Submit job recordsReview the 3.0 evidence

Disclosure: Hot Aisle provided $200 in compute credits for my own testing.

Test a frozen prediction

Load three records, in order: the design you locked before the jobs, the released medals bound to that design, and every terminal outcome. Public release of a medal does not prevent a prospective test of future jobs. Files stay on this device.

Real campaign evidence

Run 3 has three scored arms across two providers. The existing calculator recomputed their accepted work and exported a sealed comparison. Whole-allocation invoices remain unresolved; this intake is separate from the prospective medal study.

Inspect the real calculated report and intake · Follow the claim / rating / behavior research

Other work on this challenge

What must survive the test?
  1. Freeze: exact plan hash, rating/rubric version, medal-to-probability transform, predictor mapping, cohort and effect threshold.
  2. Bind, don't backfill: the design and predictions must precede the test jobs. A separate binding record supplies the released medals, hashes the exact plan it binds to, and must be timestamped after the freeze and before every job started -- so a medal can't be quietly chosen once outcomes are in.
  3. Complete cohort: every planned job has one terminal record, including errors and timeouts.
  4. Customer scope: ordinary tenant permissions and standard support; no silent reviewer upgrade.
  5. Prospective prediction: model and prediction timestamps precede the work.
  6. Held-out providers: no test provider appears in either training set; the rating is the only additional declared input.
  7. Coverage floor: at least 10 trials, 5 distinct sessions and 3 distinct UTC calendar dates per provider.
  8. Ranking value: a medal-permutation significance test -- does giving each provider its own bound medal's anchor beat giving every provider the cohort-mean anchor, by more than a preregistered minimum, more often than random medal-to-provider reshuffles would? The v1.3 comparisons against a plain baseline and a fixed 0.50 reference are still reported, but as descriptive context only -- they include a level shift and do not isolate medal-specific information.
Disclosure

Hot Aisle provided $200 in compute credits for my own testing.

What can this still get wrong?

The scorer checks declarations and arithmetic. It does not authenticate billing records, establish an independent timestamp, detect hidden model inputs or prove that a submitted baseline was competently fitted. Freeze and independently timestamp the plan; inspect the evidence behind the hashes.

Each provider receives equal weight after averaging its jobs. The 95% percentile bootstrap resamples providers 5,000 times. The eight-provider minimum and the per-provider coverage floor are declared demonstration gates, not a power calculation. Shared hardware, geography and time shocks can still create dependence.

The medal-to-probability transform (design/transform.json) is SecondRun's uncalibrated research assumption, not a probability asserted by SemiAnalysis; it must canonically hash to a transform registered in design/registry.json. Results are reported as a pilot descriptive signal from the medal permutation test — positive, negative, not_testable or inconclusive — never as a pass/fail certification.

The primary endpoint is a completed job meeting all frozen accepted-work, p95 latency and total-cost gates. This tests predictive value for that endpoint. A claim about ClusterMAX 3.0 requires service-matched managed-cluster evidence; standalone bare metal or token APIs cannot substitute. It does not establish rare-event security, long-run reliability, creditworthiness, or what special access caused. Reviewer-versus-customer treatment requires a separate paired or randomized study.

Why this test exists

ClusterMAX 2.0 describes a relative rating, providers preparing after receiving its checklist, and limits on what roughly five days of testing establish about reliability. Those are reasons to test incremental customer-outcome value rather than count endorsements. They are not findings that all its measurements are wrong.

Published 2.0 methodology · Source ledger

Archive: earlier public-code probes

An earlier release ran seven controlled probes of the exact public audit_runner.py at commit 97865af001e5bbf2f0ea5672ec0f8c0d7fb123f4, with mocked collector, reporter and target detection — no cloud or host audit ran. These are narrow API observations about return-code control flow, not findings about the CLI: the public CLI exits 2 on failed security checks. They are archived, not part of the standing challenge above.

Exact probe record · Reproduce the probes

ClusterMAX 3.0: release ledger and scope

The complete 77-provider transcription and 2.1-to-3.0 transitions are preserved. The ledger identifies exactly which particulars were reviewed.

The post-release ledger records 17 selected assignments and distinguishes supplied facts, partial evidence and unreviewed particulars. The article contains real performance and recovery testing. Our separate question is the medal's added predictive value for customer jobs.

3.0 covers managed clusters. Bare-metal-only and token-endpoint results cannot stand in for that service. Participation Ribbon is recorded but has no anchor in the original five-tier transform; such bindings HOLD. Unavailable is not a quality score. Protocol and publication history.