ReviewRouter · disagreement-routed human oversight

Compliance Screening Rubric (demo)

Panel consensus acted on automatically; contested judgments routed to humans. Disagreement is signal, not noise.

judges judge_a, judge_b, judge_cunits 40decisions 320policy auto ≤ 0.15 · review ≥ 0.35generated 2026-07-09
76.6%
decisions auto-accepted on panel consensus
18.1%
spot-check band (sample these)
5.3%
routed to human review
16
units in the review queue

Review queue — work in this order

Ranked by flagged items, then mean disagreement. Each row shows the most contested items with every judge's value and the panel consensus. Full ranked queue + adjudicated dataset: C:\Users\C_Sin\AppData\Local\Temp\claude\C--Users-C-Sin-Full-Dataset-Project\39c0c605-4c7e-42a5-afea-e72921eb3c86\scratchpad\d3\review_queue.csv

#unitreview itemsspot itemsmean disagr.most contested
1CASE_032230.267risk_tone — judge_a: 5, judge_b: 3, judge_c: 0 → consensus 3
remediation_clarity — judge_a: 3, judge_b: 1, judge_c: 4 → consensus 3
2CASE_028150.267risk_tone — judge_a: 2, judge_b: 5, judge_c: 0 → consensus 2
3CASE_003130.225remediation_clarity — judge_a: 1, judge_b: 4, judge_c: 3 → consensus 3
4CASE_004130.183risk_tone — judge_a: 1, judge_b: 1, judge_c: 4 → consensus 1
5CASE_017120.183risk_tone — judge_a: 3, judge_b: 1, judge_c: 5 → consensus 3
6CASE_027120.183risk_tone — judge_a: 1, judge_b: 4, judge_c: 0 → consensus 1
7CASE_008120.167risk_tone — judge_a: 3, judge_b: 0, judge_c: 2 → consensus 2
8CASE_029120.167risk_tone — judge_a: 0, judge_b: 0, judge_c: 3 → consensus 0
9CASE_030120.167risk_tone — judge_a: 3, judge_b: 2, judge_c: 5 → consensus 3
10CASE_001110.108risk_tone — judge_a: 2, judge_b: 0, judge_c: 3 → consensus 2
11CASE_005100.083risk_tone — judge_a: 5, judge_b: 3, judge_c: 2 → consensus 3
12CASE_011100.083risk_tone — judge_a: 5, judge_b: 3, judge_c: 2 → consensus 3

Item routing profile

Green = auto-accepted, amber = spot-check, red = review. The class column applies the shared decision rule: instrument = the item disagrees across most units (the wording generates the queue — fix the sentence, verify with an AlphaGate audit); object = disagreement concentrates in specific units (genuine contestation — the queue is doing its job).

itemlevelunitsroutingreview %mean disagr.
risk_toneordinal4037.5%0.277✕ instrument problem — fix the wording
remediation_clarityordinal405.0%0.143✓ object contestation — queue is real
control_specificityordinal400.0%0.073mixed
has_escalation_pathbinary400.0%0.017low
owner_typenominal400.0%0.058mixed
policy_scopeordinal400.0%0.080mixed
review_cadencenominal400.0%0.067mixed
third_party_attestedbinary400.0%0.200✕ instrument problem — fix the wording

Method. Per (unit, item): consensus = median (numeric scales, snapped to a real scale point) or majority vote (categorical); disagreement = mean pairwise distance between judges normalized by scale span, or 1 − majority share for categories. On a three-judge 0–5 panel: one adjacent-point dissent auto-accepts (0.13), a 2–1 categorical split spot-checks (0.33), a far outlier reviews (0.40+). Decisions: AUTO_ACCEPT ≤ 0.15, REVIEW ≥ 0.35 (ties and short panels always route to review), SPOT_CHECK between. Thresholds are policy, not truth — tune them to your review budget.

Why panels beat confidence scores. A single model's self-reported confidence is notoriously miscalibrated and vendor-specific. Cross-family disagreement is vendor-neutral by construction and localizes contested judgment. Sibling tools: AlphaGate (is the rubric reliable enough to trust consensus?) and AdherenceBench (can each judge follow the schema at all?).

Generated by ReviewRouter.