ReviewRouter · disagreement-routed human oversight
Panel consensus acted on automatically; contested judgments routed to humans. Disagreement is signal, not noise.
Ranked by flagged items, then mean disagreement. Each row shows the most contested items with every judge's value and the panel consensus. Full ranked queue + adjudicated dataset: C:\Users\C_Sin\AppData\Local\Temp\claude\C--Users-C-Sin-Full-Dataset-Project\39c0c605-4c7e-42a5-afea-e72921eb3c86\scratchpad\d3\review_queue.csv
| # | unit | review items | spot items | mean disagr. | most contested |
|---|---|---|---|---|---|
| 1 | CASE_032 | 2 | 3 | 0.267 | risk_tone — judge_a: 5, judge_b: 3, judge_c: 0 → consensus 3 remediation_clarity — judge_a: 3, judge_b: 1, judge_c: 4 → consensus 3 |
| 2 | CASE_028 | 1 | 5 | 0.267 | risk_tone — judge_a: 2, judge_b: 5, judge_c: 0 → consensus 2 |
| 3 | CASE_003 | 1 | 3 | 0.225 | remediation_clarity — judge_a: 1, judge_b: 4, judge_c: 3 → consensus 3 |
| 4 | CASE_004 | 1 | 3 | 0.183 | risk_tone — judge_a: 1, judge_b: 1, judge_c: 4 → consensus 1 |
| 5 | CASE_017 | 1 | 2 | 0.183 | risk_tone — judge_a: 3, judge_b: 1, judge_c: 5 → consensus 3 |
| 6 | CASE_027 | 1 | 2 | 0.183 | risk_tone — judge_a: 1, judge_b: 4, judge_c: 0 → consensus 1 |
| 7 | CASE_008 | 1 | 2 | 0.167 | risk_tone — judge_a: 3, judge_b: 0, judge_c: 2 → consensus 2 |
| 8 | CASE_029 | 1 | 2 | 0.167 | risk_tone — judge_a: 0, judge_b: 0, judge_c: 3 → consensus 0 |
| 9 | CASE_030 | 1 | 2 | 0.167 | risk_tone — judge_a: 3, judge_b: 2, judge_c: 5 → consensus 3 |
| 10 | CASE_001 | 1 | 1 | 0.108 | risk_tone — judge_a: 2, judge_b: 0, judge_c: 3 → consensus 2 |
| 11 | CASE_005 | 1 | 0 | 0.083 | risk_tone — judge_a: 5, judge_b: 3, judge_c: 2 → consensus 3 |
| 12 | CASE_011 | 1 | 0 | 0.083 | risk_tone — judge_a: 5, judge_b: 3, judge_c: 2 → consensus 3 |
Green = auto-accepted, amber = spot-check, red = review. The class column applies the shared decision rule: instrument = the item disagrees across most units (the wording generates the queue — fix the sentence, verify with an AlphaGate audit); object = disagreement concentrates in specific units (genuine contestation — the queue is doing its job).
| item | level | units | routing | review % | mean disagr. | |
|---|---|---|---|---|---|---|
| risk_tone | ordinal | 40 | 37.5% | 0.277 | ✕ instrument problem — fix the wording | |
| remediation_clarity | ordinal | 40 | 5.0% | 0.143 | ✓ object contestation — queue is real | |
| control_specificity | ordinal | 40 | 0.0% | 0.073 | mixed | |
| has_escalation_path | binary | 40 | 0.0% | 0.017 | low | |
| owner_type | nominal | 40 | 0.0% | 0.058 | mixed | |
| policy_scope | ordinal | 40 | 0.0% | 0.080 | mixed | |
| review_cadence | nominal | 40 | 0.0% | 0.067 | mixed | |
| third_party_attested | binary | 40 | 0.0% | 0.200 | ✕ instrument problem — fix the wording |
Method. Per (unit, item): consensus = median (numeric scales, snapped to a real scale point) or majority vote (categorical); disagreement = mean pairwise distance between judges normalized by scale span, or 1 − majority share for categories. On a three-judge 0–5 panel: one adjacent-point dissent auto-accepts (0.13), a 2–1 categorical split spot-checks (0.33), a far outlier reviews (0.40+). Decisions: AUTO_ACCEPT ≤ 0.15, REVIEW ≥ 0.35 (ties and short panels always route to review), SPOT_CHECK between. Thresholds are policy, not truth — tune them to your review budget.
Why panels beat confidence scores. A single model's self-reported confidence is notoriously miscalibrated and vendor-specific. Cross-family disagreement is vendor-neutral by construction and localizes contested judgment. Sibling tools: AlphaGate (is the rubric reliable enough to trust consensus?) and AdherenceBench (can each judge follow the schema at all?).
Generated by ReviewRouter.