The Ethics Observatory

Future & Related Projects

Five instruments grown from one methodology.

The Observatory's core question — where does ethical commitment couple to verifiable practice? — required building measurement discipline for AI coders that didn't exist yet. That discipline turned out to be worth more than one study. Each project below is a working instrument with a live demo, spun out of the DPPRA/ICST research program and applicable far beyond it: anywhere a machine is trusted to judge.

gate the rubric · gate the judge · route the judgments · report it all · check it against reality

STATUS — working prototypes with reproducible demos · public release follows the methodology paper (2026) · demo data below is synthetic
α

The reference instrument — most mature of the five

AlphaGate — working title

Is your scoring rubric reliable enough to trust a machine on?

Organizations increasingly let language models grade, score, and screen — usually validated by spot checks. Content analysis solved this problem decades ago for human raters: independent raters must demonstrably agree, measured with chance-corrected statistics. AlphaGate applies that standard to machine judges: it runs any rubric across multiple independent models, computes Krippendorff's alpha per item, and returns a pass / tentative / fail verdict against published cutoffs — led by a remediation section: the exact items to fix, the divergent readings observed, and the expected alpha change. Its central lesson: low reliability usually localizes to specific rubric wording — you fix the sentence, not the model. Interpretation is stated, not assumed: with cross-family LLM panels, alpha measures instrument determinacy (does the rubric constrain any competent reader to the same reading?), not rater independence — model families share training distributions, and the report includes a pairwise dependence diagnostic so a reader can check the panel rather than trust it.

keylesszero dependencies math validated against reference implementation

What it makes possible — the roadmap

The four instruments below extend the same discipline outward: from the rubric to the judge, from the judge to the oversight loop, from the loop to the reporting standard, and finally to the question every index owes an answer — does any of it predict reality?

7 ∉ [0..5]

Adherence · the second gate

AdherenceBench — working title

Can the model actually follow your rules?

Public benchmarks measure whether models are smart. Deployments fail on a different axis: whether the model stays inside the instrument — the schema, the bounded scales, the citation rules. AdherenceBench is a deterministic battery of structured-coding tasks whose ground truth is known by construction, so it catches what schema validation can't: fabrication (confidently coding a question the source never answers), paraphrased "verbatim" quotes, invented scale points, and instructions followed from inside the documents being analyzed. Models fail characteristically, not randomly — the failure-mode profile, not the headline score, is the procurement signal.

5-dimension adherence score 91-field marathon tiernothing silently repaired
auto human

Oversight · the routing layer

ReviewRouter — working title

Which machine judgments need a human, right now?

Every AI-governance framework demands "human oversight," and none of them says which judgments the scarce humans should look at. ReviewRouter uses a better signal than any single model's self-reported confidence: disagreement between independent judges. Where a cross-family panel agrees, the consensus is acted on automatically — with provenance. Where it disagrees, the case lands in a ranked review queue, because contested readings are exactly where one model would silently pick a side. Disagreement is signal, not noise — and rubric items that generate the queue by themselves get flagged as instrument problems.

ranked review queue adjudicated dataset with provenancetunable to review budget

Reporting · the standard

RELIANCE — v1.0 draft standard

Can a reader verify any of it?

Research and industry are adopting LLM coding at speed, and most of it would not survive the scrutiny routinely applied to a single human coder. RELIANCE (REporting LLM-based content ANalysis: Checklist and Evidence) is a CONSORT-style 12-item checklist — versioned instruments, cross-family panels, chance-corrected reliability, verbatim deviation registers, human anchors, audit availability — with a companion tool that grades a study's declaration into compliance tiers and generates the methods-section text from the declaration, so a paper can only claim what it declared. Twelve rows, one screen: faster than excavating a methods section, comparable across studies.

12 items · 4 critical compliance tiersmethods text auto-drafted

Validity · phase two

RealityCheck — working title

Do the scores predict anything real?

The question every index eventually faces, and most never answer. An index that only describes documents is a mirror; one whose scores forecast independent outcomes is an instrument. RealityCheck is a criterion-validity harness with the discipline enforced in code: hypotheses are hash-locked before outcome data is joined (preregistration as software), a circularity guard mechanically excludes self-published outcome evidence — an index must not predict institutions' documents with institutions' documents — and every lagged association is tested against a reversed-lag placebo. Null results are published as findings. This is the planned validation phase of the Decoupling Index.

preregistration as code circularity guardplacebo-tested