The reference instrument — most mature of the five
AlphaGate — working title
Is your scoring rubric reliable enough to trust a machine on?
Organizations increasingly let language models grade, score, and screen — usually validated by spot checks. Content analysis solved this problem decades ago for human raters: independent raters must demonstrably agree, measured with chance-corrected statistics. AlphaGate applies that standard to machine judges: it runs any rubric across multiple independent models, computes Krippendorff's alpha per item, and returns a pass / tentative / fail verdict against published cutoffs — led by a remediation section: the exact items to fix, the divergent readings observed, and the expected alpha change. Its central lesson: low reliability usually localizes to specific rubric wording — you fix the sentence, not the model. Interpretation is stated, not assumed: with cross-family LLM panels, alpha measures instrument determinacy (does the rubric constrain any competent reader to the same reading?), not rater independence — model families share training distributions, and the report includes a pairwise dependence diagnostic so a reader can check the panel rather than trust it.