RealityCheck · criterion-validity ledger
Do the scores predict independent reality? Preregistered hypotheses, circularity-guarded outcomes, reversed-lag placebos.
Spearman's ρ with two-sided permutation p-values; verdicts against the prespecified direction. A confirmed association whose reversed-lag placebo is also significant is downgraded — predicting the past is confounding, not foresight.
| id | hypothesis | n | ρ | p | placebo (reversed lag) | verdict |
|---|---|---|---|---|---|---|
| H1 | decoupling_score → Independent incident-tracker rate expects positive · lag 1 · scope: all · evidence: third_party_administrative · exposure: low decoupled governance leaves real exposure unmanaged | 120 | 0.611 | 0.0010 | ρ=-0.158, p=0.2108 ✓clean | ✓ CONFIRMED |
| H2 | decoupling_score → Regulatory enforcement actions expects positive · lag 1 · scope: regime=market · evidence: third_party_administrative · exposure: medium market regimes rely on ex-post enforcement, so the gap surfaces there | 60 | 0.750 | 0.0010 | ρ=-0.280, p=0.1319 ✓clean | ✓ CONFIRMED |
| H3 | commitment_depth → Days from commitment to enacted control expects negative · lag 0 · scope: all · evidence: behavioral_record · exposure: low deeper verifiable commitment should shorten enactment | 120 | -0.188 | 0.0380 | — | ✓ CONFIRMED |
| H4 | rhetoric_intensity → Independent incident-tracker rate expects positive · lag 1 · scope: all · evidence: third_party_administrative · exposure: low falsification probe: tone alone should predict nothing | 120 | 0.033 | 0.7123 | ρ=0.153, p=0.2478 ✓clean | ○ NULL |
| H5 | decoupling_score → Own press-release volume expects positive · lag 0 · scope: all · evidence: self_published · exposure: high guard demonstration: self-published outcome must be excluded evidence class 'self_published' — excluded by the circularity guard; an index must not predict documents with documents | — | — | — | — | 🛡 GUARDED |
Endogeneity is not designed away. The circularity guard excludes self-published outcome evidence, but third-party outcomes — rankings, accreditation actions, enforcement, press coverage — are themselves shaped by the same institutional discourse the scores measure. A reversed-lag placebo does not touch this. Where declared, outcomes are stratified by exposure to the measured discourse; prediction concentrated in high-exposure outcomes should be read as discourse reverberation, not enactment.
Confirmed associations are correlational. The design supports forecasting claims, not causal ones, and a validity ledger is evidence about the index, not about any single organization.
Stratified by discourse exposure: low: confirmed 2, null 1 · medium: confirmed 1
Circularity guard. Outcomes with evidence class self_published are mechanically excluded (🛡 GUARDED): an index must not predict institutions' documents with institutions' documents. Eligible evidence: third-party/administrative records and behavioral records only.
Reading the ledger. The deliverable is the whole ledger, not one coefficient: how many prespecified associations landed, where, and under which governance regime. NULL rows are findings too — a stable, consequence-free say–do gap is a first-order policy result.
Generated by RealityCheck — phase 2 of the measurement family: gate the rubric, gate the judge, route the judgments, report it all, then check it against reality.