RELIANCE v1.1 · reporting standard for LLM-coded data
Example Research Group
11 met · 1 partial · 0 not met · 0 n/a (justified) — of 12 items · 2026-07-09
valid until 2026-12-31
| item | requirement & evidence | status |
|---|---|---|
| Instrument | ||
| R1 ★ CRITICAL | Instrument versioning Instrument (codebook + prompt text) is version-identified and frozen per run; any change is a new, reported version. Evidence: instrument v2.3, SHA-256 hash-locked; every prompt or field change bumps the version and triggers a re-code | ✓ met |
| R2 | Level of measurement Every field declares binary/nominal/ordinal/interval, and statistics match the level. Evidence: per-field level declared in the codebook (binary/nominal/ordinal); reliability statistics matched per level | ✓ met |
| R3 | Missing-data convention 'Unclear'/'N/A' codes are declared, distinct from scale points, and treated as missing in reliability. Evidence: codes 8 (unclear) and 9 (not applicable) declared as missing markers, excluded from reliability computation | ✓ met |
| Panel | ||
| R4 ★ CRITICAL | Panel composition >=2 independent automated coders per unit (>=3 recommended), independence declared; single-model designs justified and labeled. Evidence: three independent coders from three model families (Anthropic, OpenAI, Google), selected for training independence | ✓ met |
| R5 | Model identity pinning Exact model IDs, versions/dates, settings, and access route reported per coder. Evidence: exact model IDs, versions, run dates, and access routes logged per coder in the run manifest | ✓ met |
| R6 | Packet archival The exact material each coder received is archived and available. Evidence: full packets (instrument + document) archived per coder per unit in the project repository | ✓ met |
| Reliability | ||
| R7 ★ CRITICAL | Chance-corrected reliability Chance-corrected statistic matched to measurement level, with uncertainty, per item/family — not pooled-only. Evidence: Krippendorff's alpha per item family (ordinal/nominal/binary) with bootstrap CIs; per-item table in the appendix | ✓ met |
| R8 | A-priori reliability gate Acceptance threshold and failure consequence declared before the run. Evidence: alpha >= .667 gate declared before the run; failing families trigger instrument revision and re-code, not post-hoc exclusion | ✓ met |
| Transparency | ||
| R9 ★ CRITICAL | Deviation register Every off-schema output logged verbatim with disposition; register available; nothing silently repaired. Evidence: public deviation register (JSONL): every off-schema output preserved verbatim with disposition; no silent repair | ✓ met |
| R10 | Disagreement handling Final-code rule (consensus/adjudication/majority) declared; extent and location of disagreement reported. Evidence: final codes by panel median/majority; disagreement reported per item and per document, contested units listed | ✓ met |
| Validity | ||
| R11 | Human anchor Human-coded subsample anchors machine coding (or absence justified and claims bounded). Evidence: human anchor limited to author adjudication of contested units; stratified human-coded subsample planned for the scale-up | ~ partial |
| R12 | Audit availability Coded outputs, register, reliability computation, and instrument available to readers. Evidence: coded outputs, register, reliability code, and instrument available in the project repository | ✓ met |
What did this instrument fail to detect?
The instrument reads documents, not conduct: it cannot detect commitments enacted but never documented, and tonal items remained below the reliability gate through two revisions before stabilizing
What null or negative results were produced?
Rhetorical-intensity items showed no reliable association with any verifiable-action item; one full instrument revision produced a reliability collapse on previously stable items and was rolled back — both are reported in the run record
A tier is a property of a report, not a team: it states what a reader can verify. Critical items (★): instrument versioning, panel composition, chance-corrected reliability, deviation register. Generate methods-section text from this declaration with reliance methods.
The limit, stated plainly. This standard makes discourse-practice gaps legible, costly to maintain, and awkward to explain. It cannot supply the will to close them: no instrument closes a gap an organization has an interest in keeping open.
A tier claim may only be published with this full declaration attached — there is no badge, and there never will be. Declarations expire and are revocable; re-verification against a newer model snapshot supersedes this document.
RELIANCE — the reporting checklist for machine-coded data. Companion tools: AlphaGate (reliability gating), AdherenceBench (schema adherence), ReviewRouter (disagreement routing).