AlphaGate · LLM-judge reliability audit
rubric v1.0-demo · Krippendorff’s α across 3 independent judges
The deliverable is the diff, not the score. Class: instrument = wording disagrees across most units (revise the item); object = disagreement concentrates in specific units (route them to human review); Δα = pooled-group alpha change if the item were dropped.
| item | α | class | Δα if dropped | guidance | divergent readings (unit — judge: value) |
|---|---|---|---|---|---|
| Provides third-party verification third_party_verification | 0.261 | instrument breadth 0.53 | +0.204 | Disagreement is broad across the corpus: the item wording permits multiple readings. Tighten the anchors — require a named, verifiable referent for top scale points — and re-run; expect alpha to move. | DOC_011 — judge_a: 1, judge_b: 0 DOC_029 — judge_b: 0, judge_c: 1 DOC_001 — judge_a: 0, judge_b: 0, judge_c: 1 |
| Rhetoric intensity (unanchored 0–5) rhetoric_intensity | 0.364 | instrument breadth 0.60 | +0.116 | Disagreement is broad across the corpus: the item wording permits multiple readings. Tighten the anchors — require a named, verifiable referent for top scale points — and re-run; expect alpha to move. | DOC_004 — judge_a: 0, judge_b: 5, judge_c: 4 DOC_008 — judge_a: 0, judge_b: 5, judge_c: 2 DOC_013 — judge_a: 5, judge_b: 0, judge_c: 5 |
| Document stance (category) document_stance | 0.542 | mixed breadth 0.40 | +0.053 | Both patterns present: partially indeterminate wording plus genuinely contested units. Tighten the anchors first, then route the residual units to review. | DOC_001 — judge_a: 2, judge_b: 3, judge_c: 1 DOC_004 — judge_a: 2, judge_b: 1, judge_c: 3 DOC_028 — judge_a: 3, judge_b: 2, judge_c: 1 |
| Harm acknowledgment (0–5) harm_acknowledgment | 0.713 | mixed breadth 0.23 | -0.019 | Both patterns present: partially indeterminate wording plus genuinely contested units. Tighten the anchors first, then route the residual units to review. | DOC_019 — judge_a: 1, judge_b: 4, judge_c: 5 DOC_024 — judge_a: 3, judge_b: 5 DOC_002 — judge_a: 5, judge_b: 5, judge_c: 3 |
| Accountability strength (0–5) accountability_strength | 0.758 | mixed breadth 0.33 | -0.017 | Both patterns present: partially indeterminate wording plus genuinely contested units. Tighten the anchors first, then route the residual units to review. | DOC_001 — judge_a: 4, judge_b: 5, judge_c: 3 DOC_002 — judge_a: 1, judge_b: 2, judge_c: 0 DOC_003 — judge_a: 1, judge_b: 3, judge_c: 2 |
| Primary named actor (category) primary_actor | 0.783 | mixed breadth 0.23 | -0.188 | Both patterns present: partially indeterminate wording plus genuinely contested units. Tighten the anchors first, then route the residual units to review. | DOC_010 — judge_a: 3, judge_b: 2, judge_c: 3 DOC_014 — judge_a: 4, judge_b: 4, judge_c: 2 DOC_021 — judge_a: 4, judge_b: 4, judge_c: 3 |
gate verdict
How to read α here. Alpha here measures instrument determinacy: whether the rubric constrains any competent reader to the same reading. It does not measure rater independence. Cross-family LLM panels share training distributions; agreement between them is correlated by construction.
Ticks on each meter mark the tentative (0.667) and pass (0.8) thresholds.
| scale group | items | α | 95% CI | meter | verdict |
|---|---|---|---|---|---|
| binary 0-1 | has_deadline, names_responsible_party, third_party_verification | 0.684 | [0.559, 0.803] | ~ TENTATIVE | |
| nominal (categories item-namespaced) | primary_actor, document_stance | 0.730 | [0.640, 0.830] | ~ TENTATIVE | |
| ordinal 0-5 | commitment_specificity, timeline_clarity, accountability_strength, rhetoric_intensity, harm_acknowledgment | 0.709 | [0.623, 0.770] | ~ TENTATIVE |
Sorted worst-first. Latitude = mean normalized pairwise distance between judges (0 = always agree, 1 = maximally apart) — high-latitude items are where the rubric leaves judges room to diverge.
| item | units | α | 95% CI | % agree | latitude | meter | verdict |
|---|---|---|---|---|---|---|---|
Provides third-party verification third_party_verification · binary | 30 | 0.261 | [0.036, 0.519] | 0.63 | 0.378 | ✕ FAIL | |
Rhetoric intensity (unanchored 0–5) rhetoric_intensity · ordinal | 30 | 0.364 | [0.105, 0.547] | 0.33 | 0.280 | ✕ FAIL | |
Document stance (category) document_stance · nominal | 30 | 0.542 | [0.339, 0.738] | 0.70 | 0.300 | ✕ FAIL | |
Harm acknowledgment (0–5) harm_acknowledgment · ordinal | 30 | 0.713 | [0.486, 0.857] | 0.42 | 0.147 | ~ TENTATIVE | |
Accountability strength (0–5) accountability_strength · ordinal | 30 | 0.758 | [0.591, 0.835] | 0.34 | 0.158 | ~ TENTATIVE | |
Primary named actor (category) primary_actor · nominal | 30 | 0.783 | [0.606, 0.906] | 0.84 | 0.156 | ~ TENTATIVE | |
Names a responsible party names_responsible_party · binary | 30 | 0.824 | [0.648, 0.956] | 0.91 | 0.089 | ✓ PASS | |
Commitment specificity (anchored 0–5) commitment_specificity · ordinal | 30 | 0.862 | [0.735, 0.915] | 0.54 | 0.093 | ✓ PASS | |
Timeline clarity (anchored 0–5) timeline_clarity · ordinal | 30 | 0.869 | [0.771, 0.908] | 0.45 | 0.113 | ✓ PASS | |
Names a concrete deadline has_deadline · binary | 30 | 0.955 | [0.845, 1.000] | 0.98 | 0.022 | ✓ PASS |
Bias = mean signed distance from the panel median on numeric scales, normalized by scale width (positive = runs hot). Deviations are schema-adherence failures — they are logged and preserved, never silently repaired.
| judge | codes | missing markers | deviations | dev / 100 | bias | tendency | deviation breakdown |
|---|---|---|---|---|---|---|---|
| judge_a | 297 | 3 | 0 | 0.00 | -0.048 | neutral | — |
| judge_b | 300 | 0 | 0 | 0.00 | 0.059 | runs hot | — |
| judge_c | 291 | 0 | 10 | 3.32 | -0.015 | neutral | duplicate_code ×1, missing_item ×3, out_of_range ×4, unparseable_value ×2 |
Pairwise agreement by judge pair. If one pair sits far above the others, those two families are collapsing toward each other and the panel is less independent than it looks — read the pooled α accordingly.
| pair | comparisons | exact agreement | mean distance |
|---|---|---|---|
| judge_a × judge_b | 297 | 0.596 | 0.165 |
| judge_a × judge_c | 288 | 0.660 | 0.162 |
| judge_b × judge_c | 291 | 0.598 | 0.189 |
Route these to human review first — disagreement is signal, not noise.
| unit | disagreement | comparisons |
|---|---|---|
| DOC_002 | 0.320 | 30 |
| DOC_024 | 0.300 | 28 |
| DOC_028 | 0.273 | 30 |
| DOC_022 | 0.253 | 30 |
| DOC_001 | 0.247 | 30 |
10 total.
| judge | unit | item | raw value | rule | action |
|---|---|---|---|---|---|
| judge_c | DOC_007 | accountability_strength | '7' | out_of_range | treated_missing |
| judge_c | DOC_009 | harm_acknowledgment | 'moderate-to-high' | unparseable_value | treated_missing |
| judge_c | DOC_014 | accountability_strength | '7' | out_of_range | treated_missing |
| judge_c | DOC_015 | timeline_clarity | 'moderate-to-high' | unparseable_value | treated_missing |
| judge_c | DOC_024 | harm_acknowledgment | '7' | out_of_range | treated_missing |
| judge_c | DOC_027 | harm_acknowledgment | '7' | out_of_range | treated_missing |
| judge_c | DOC_001 | commitment_specificity | '5' | duplicate_code | dropped |
| judge_c | DOC_011 | third_party_verification | None | missing_item | treated_missing |
| judge_c | DOC_013 | third_party_verification | None | missing_item | treated_missing |
| judge_c | DOC_021 | third_party_verification | None | missing_item | treated_missing |
Method. Krippendorff’s α (coincidence-matrix formulation) with the distance metric matched to each item’s level of measurement: nominal, ordinal, or interval; binary items use the nominal metric. Units with fewer than two non-missing values drop out. Confidence intervals are nonparametric bootstrap over units. Rubric missing-value markers are treated as missing, not as scale points. Cutoffs follow Krippendorff (Content Analysis, 4th ed.): α ≥ .80 reliable; .667–.80 tentative conclusions only.
Reading this report. A PASS means the judges are interchangeable on that scale — single-judge output can be trusted. A FAIL is not a model problem by default: it usually localizes to specific rubric items whose wording leaves judges latitude. Fix the instrument, re-run, and watch α move.
Validity & publication. Valid until 2027-01-05. Verdicts are not excerptable: a pass/tentative/fail claim may only be published with this full report — including the deviation register — attached. Re-verification against a newer model snapshot supersedes this report; drift is a result, not an error.
Generated by AlphaGate — the go/no-go gate for LLM judges.