AlphaGate · LLM-judge reliability audit

Institutional Commitment Rubric (demo)

rubric v1.0-demo · Krippendorff’s α across 3 independent judges

judges judge_a, judge_b, judge_cunits 30codes analyzed 888deviations logged 10generated 2026-07-09

Remediation — fix these first

The deliverable is the diff, not the score. Class: instrument = wording disagrees across most units (revise the item); object = disagreement concentrates in specific units (route them to human review); Δα = pooled-group alpha change if the item were dropped.

itemαclassΔα if droppedguidancedivergent readings (unit — judge: value)
Provides third-party verification
third_party_verification
0.261instrument
breadth 0.53
+0.204Disagreement is broad across the corpus: the item wording permits multiple readings. Tighten the anchors — require a named, verifiable referent for top scale points — and re-run; expect alpha to move.DOC_011 — judge_a: 1, judge_b: 0
DOC_029 — judge_b: 0, judge_c: 1
DOC_001 — judge_a: 0, judge_b: 0, judge_c: 1
Rhetoric intensity (unanchored 0–5)
rhetoric_intensity
0.364instrument
breadth 0.60
+0.116Disagreement is broad across the corpus: the item wording permits multiple readings. Tighten the anchors — require a named, verifiable referent for top scale points — and re-run; expect alpha to move.DOC_004 — judge_a: 0, judge_b: 5, judge_c: 4
DOC_008 — judge_a: 0, judge_b: 5, judge_c: 2
DOC_013 — judge_a: 5, judge_b: 0, judge_c: 5
Document stance (category)
document_stance
0.542mixed
breadth 0.40
+0.053Both patterns present: partially indeterminate wording plus genuinely contested units. Tighten the anchors first, then route the residual units to review.DOC_001 — judge_a: 2, judge_b: 3, judge_c: 1
DOC_004 — judge_a: 2, judge_b: 1, judge_c: 3
DOC_028 — judge_a: 3, judge_b: 2, judge_c: 1
Harm acknowledgment (0–5)
harm_acknowledgment
0.713mixed
breadth 0.23
-0.019Both patterns present: partially indeterminate wording plus genuinely contested units. Tighten the anchors first, then route the residual units to review.DOC_019 — judge_a: 1, judge_b: 4, judge_c: 5
DOC_024 — judge_a: 3, judge_b: 5
DOC_002 — judge_a: 5, judge_b: 5, judge_c: 3
Accountability strength (0–5)
accountability_strength
0.758mixed
breadth 0.33
-0.017Both patterns present: partially indeterminate wording plus genuinely contested units. Tighten the anchors first, then route the residual units to review.DOC_001 — judge_a: 4, judge_b: 5, judge_c: 3
DOC_002 — judge_a: 1, judge_b: 2, judge_c: 0
DOC_003 — judge_a: 1, judge_b: 3, judge_c: 2
Primary named actor (category)
primary_actor
0.783mixed
breadth 0.23
-0.188Both patterns present: partially indeterminate wording plus genuinely contested units. Tighten the anchors first, then route the residual units to review.DOC_010 — judge_a: 3, judge_b: 2, judge_c: 3
DOC_014 — judge_a: 4, judge_b: 4, judge_c: 2
DOC_021 — judge_a: 4, judge_b: 4, judge_c: 3

gate verdict

~ TENTATIVE
acceptable only for tentative conclusions — human review recommended
4 pass · 3 tentative · 3 fail of 10 items
0.684
α · binary 0-1 (3 items)
~ TENTATIVE
0.730
α · nominal (categories item-namespaced) (2 items)
~ TENTATIVE
0.709
α · ordinal 0-5 (5 items)
~ TENTATIVE

How to read α here. Alpha here measures instrument determinacy: whether the rubric constrains any competent reader to the same reading. It does not measure rater independence. Cross-family LLM panels share training distributions; agreement between them is correlated by construction.

Headline reliability — pooled by shared scale

Ticks on each meter mark the tentative (0.667) and pass (0.8) thresholds.

scale groupitemsα95% CImeterverdict
binary 0-1has_deadline, names_responsible_party, third_party_verification0.684[0.559, 0.803]~ TENTATIVE
nominal (categories item-namespaced)primary_actor, document_stance0.730[0.640, 0.830]~ TENTATIVE
ordinal 0-5commitment_specificity, timeline_clarity, accountability_strength, rhetoric_intensity, harm_acknowledgment0.709[0.623, 0.770]~ TENTATIVE

Per-item reliability

Sorted worst-first. Latitude = mean normalized pairwise distance between judges (0 = always agree, 1 = maximally apart) — high-latitude items are where the rubric leaves judges room to diverge.

itemunitsα95% CI% agreelatitudemeterverdict
Provides third-party verification
third_party_verification · binary
300.261[0.036, 0.519]0.630.378✕ FAIL
Rhetoric intensity (unanchored 0–5)
rhetoric_intensity · ordinal
300.364[0.105, 0.547]0.330.280✕ FAIL
Document stance (category)
document_stance · nominal
300.542[0.339, 0.738]0.700.300✕ FAIL
Harm acknowledgment (0–5)
harm_acknowledgment · ordinal
300.713[0.486, 0.857]0.420.147~ TENTATIVE
Accountability strength (0–5)
accountability_strength · ordinal
300.758[0.591, 0.835]0.340.158~ TENTATIVE
Primary named actor (category)
primary_actor · nominal
300.783[0.606, 0.906]0.840.156~ TENTATIVE
Names a responsible party
names_responsible_party · binary
300.824[0.648, 0.956]0.910.089✓ PASS
Commitment specificity (anchored 0–5)
commitment_specificity · ordinal
300.862[0.735, 0.915]0.540.093✓ PASS
Timeline clarity (anchored 0–5)
timeline_clarity · ordinal
300.869[0.771, 0.908]0.450.113✓ PASS
Names a concrete deadline
has_deadline · binary
300.955[0.845, 1.000]0.980.022✓ PASS

Judge adherence profiles

Bias = mean signed distance from the panel median on numeric scales, normalized by scale width (positive = runs hot). Deviations are schema-adherence failures — they are logged and preserved, never silently repaired.

judgecodesmissing markersdeviationsdev / 100biastendencydeviation breakdown
judge_a297300.00-0.048neutral
judge_b300000.000.059runs hot
judge_c2910103.32-0.015neutralduplicate_code ×1, missing_item ×3, out_of_range ×4, unparseable_value ×2

Judge dependence diagnostic

Pairwise agreement by judge pair. If one pair sits far above the others, those two families are collapsing toward each other and the panel is less independent than it looks — read the pooled α accordingly.

paircomparisonsexact agreementmean distance
judge_a × judge_b2970.5960.165
judge_a × judge_c2880.6600.162
judge_b × judge_c2910.5980.189

Highest-disagreement units

Route these to human review first — disagreement is signal, not noise.

unitdisagreementcomparisons
DOC_0020.32030
DOC_0240.30028
DOC_0280.27330
DOC_0220.25330
DOC_0010.24730

Deviation register

10 total.

judgeunititemraw valueruleaction
judge_cDOC_007accountability_strength'7'out_of_rangetreated_missing
judge_cDOC_009harm_acknowledgment'moderate-to-high'unparseable_valuetreated_missing
judge_cDOC_014accountability_strength'7'out_of_rangetreated_missing
judge_cDOC_015timeline_clarity'moderate-to-high'unparseable_valuetreated_missing
judge_cDOC_024harm_acknowledgment'7'out_of_rangetreated_missing
judge_cDOC_027harm_acknowledgment'7'out_of_rangetreated_missing
judge_cDOC_001commitment_specificity'5'duplicate_codedropped
judge_cDOC_011third_party_verificationNonemissing_itemtreated_missing
judge_cDOC_013third_party_verificationNonemissing_itemtreated_missing
judge_cDOC_021third_party_verificationNonemissing_itemtreated_missing

Method. Krippendorff’s α (coincidence-matrix formulation) with the distance metric matched to each item’s level of measurement: nominal, ordinal, or interval; binary items use the nominal metric. Units with fewer than two non-missing values drop out. Confidence intervals are nonparametric bootstrap over units. Rubric missing-value markers are treated as missing, not as scale points. Cutoffs follow Krippendorff (Content Analysis, 4th ed.): α ≥ .80 reliable; .667–.80 tentative conclusions only.

Reading this report. A PASS means the judges are interchangeable on that scale — single-judge output can be trusted. A FAIL is not a model problem by default: it usually localizes to specific rubric items whose wording leaves judges latitude. Fix the instrument, re-run, and watch α move.

Validity & publication. Valid until 2027-01-05. Verdicts are not excerptable: a pass/tentative/fail claim may only be published with this full report — including the deviation register — attached. Re-verification against a newer model snapshot supersedes this report; drift is a result, not an error.

Generated by AlphaGate — the go/no-go gate for LLM judges.