RELIANCE v1.1 · reporting standard for LLM-coded data

Institutional AI-policy coding pilot (example declaration)

Example Research Group

RELIANCE-Substantial

11 met · 1 partial · 0 not met · 0 n/a (justified) — of 12 items · 2026-07-09

valid until 2026-12-31

itemrequirement & evidencestatus
Instrument
R1 ★ CRITICALInstrument versioning
Instrument (codebook + prompt text) is version-identified and frozen per run; any change is a new, reported version.
Evidence: instrument v2.3, SHA-256 hash-locked; every prompt or field change bumps the version and triggers a re-code
✓ met
R2Level of measurement
Every field declares binary/nominal/ordinal/interval, and statistics match the level.
Evidence: per-field level declared in the codebook (binary/nominal/ordinal); reliability statistics matched per level
✓ met
R3Missing-data convention
'Unclear'/'N/A' codes are declared, distinct from scale points, and treated as missing in reliability.
Evidence: codes 8 (unclear) and 9 (not applicable) declared as missing markers, excluded from reliability computation
✓ met
Panel
R4 ★ CRITICALPanel composition
>=2 independent automated coders per unit (>=3 recommended), independence declared; single-model designs justified and labeled.
Evidence: three independent coders from three model families (Anthropic, OpenAI, Google), selected for training independence
✓ met
R5Model identity pinning
Exact model IDs, versions/dates, settings, and access route reported per coder.
Evidence: exact model IDs, versions, run dates, and access routes logged per coder in the run manifest
✓ met
R6Packet archival
The exact material each coder received is archived and available.
Evidence: full packets (instrument + document) archived per coder per unit in the project repository
✓ met
Reliability
R7 ★ CRITICALChance-corrected reliability
Chance-corrected statistic matched to measurement level, with uncertainty, per item/family — not pooled-only.
Evidence: Krippendorff's alpha per item family (ordinal/nominal/binary) with bootstrap CIs; per-item table in the appendix
✓ met
R8A-priori reliability gate
Acceptance threshold and failure consequence declared before the run.
Evidence: alpha >= .667 gate declared before the run; failing families trigger instrument revision and re-code, not post-hoc exclusion
✓ met
Transparency
R9 ★ CRITICALDeviation register
Every off-schema output logged verbatim with disposition; register available; nothing silently repaired.
Evidence: public deviation register (JSONL): every off-schema output preserved verbatim with disposition; no silent repair
✓ met
R10Disagreement handling
Final-code rule (consensus/adjudication/majority) declared; extent and location of disagreement reported.
Evidence: final codes by panel median/majority; disagreement reported per item and per document, contested units listed
✓ met
Validity
R11Human anchor
Human-coded subsample anchors machine coding (or absence justified and claims bounded).
Evidence: human anchor limited to author adjudication of contested units; stratified human-coded subsample planned for the scale-up
~ partial
R12Audit availability
Coded outputs, register, reliability computation, and instrument available to readers.
Evidence: coded outputs, register, reliability code, and instrument available in the project repository
✓ met

Required disclosures

What did this instrument fail to detect?
The instrument reads documents, not conduct: it cannot detect commitments enacted but never documented, and tonal items remained below the reliability gate through two revisions before stabilizing

What null or negative results were produced?
Rhetorical-intensity items showed no reliable association with any verifiable-action item; one full instrument revision produced a reliability collapse on previously stable items and was rolled back — both are reported in the run record

A tier is a property of a report, not a team: it states what a reader can verify. Critical items (★): instrument versioning, panel composition, chance-corrected reliability, deviation register. Generate methods-section text from this declaration with reliance methods.

The limit, stated plainly. This standard makes discourse-practice gaps legible, costly to maintain, and awkward to explain. It cannot supply the will to close them: no instrument closes a gap an organization has an interest in keeping open.

A tier claim may only be published with this full declaration attached — there is no badge, and there never will be. Declarations expire and are revocable; re-verification against a newer model snapshot supersedes this document.

RELIANCE — the reporting checklist for machine-coded data. Companion tools: AlphaGate (reliability gating), AdherenceBench (schema adherence), ReviewRouter (disagreement routing).