The exact analysis set
Frozen RUN_07, protocol v2.14: 195 coding outputs, 65 documents and three coder configurations. Sector counts are government 15, journal 14, forum 12, institutional 14 and corporate 10. The instrument contains 91 coded fields and 16 narrative fields; this battery selects the 29 ordinal fields identified by the production classifier.
Codes 8 (unclear) and 9 (not applicable) are missing; process/confidence and deprecated fields are excluded. Three-coder matching yields 1,727 document–field units and 5,181 ratings. Document means use the same eligible fields for all coders and give each document equal weight. Three fields are domain-specific. Blank slots in the full numerical export denote missing or inapplicable values.
| Ordinal field | Complete documents |
|---|---|
| cb_benefit_intensity_code | 65 |
| cb_benefit_presence_strength_code | 65 |
| cb_cost_intensity_code | 65 |
| cb_cost_presence_strength_code | 65 |
| detection_verification_presence_strength_code | 12 |
| dppra_a_strength_code | 65 |
| dppra_alignment_score_code | 64 |
| dppra_d_strength_code | 65 |
| dppra_p1_strength_code | 65 |
| dppra_p2_strength_code | 65 |
| dppra_r_strength_code | 65 |
| eth_accountability_code | 65 |
| eth_autonomy_code | 65 |
| eth_beneficence_eth_code | 65 |
| eth_beneficence_mech_code | 65 |
| eth_equity_code | 65 |
| eth_integrity_code | 65 |
| eth_justice_code | 65 |
| eth_nonmaleficence_code | 65 |
| eth_privacy_code | 65 |
| eth_transparency_code | 65 |
| int_academic_intensity_code | 65 |
| int_assessment_intensity_code | 65 |
| int_data_prov_intensity_code | 65 |
| int_epistemic_intensity_code | 65 |
| int_research_intensity_code | 65 |
| int_system_tevv_intensity_code | 65 |
| journal_evidence_quality_strength_code | 14 |
| surveillance_monitoring_presence_strength_code | 12 |
Original results reproduced
The repeated-measures ANOVA uses document averages. Its partial η² is coder-effect sum of squares divided by coder-effect-plus-error sums of squares. It is not 76% of all raw rating variance, and it is not a causal attribution. Arithmetic summaries retain an approximately equal-spacing interpretation. Ordinal Krippendorff’s α reproduces as 0.7247.
| Documents | F | df 1 | df 2 | p | Partial η² | SS coder | SS error | SS total |
|---|---|---|---|---|---|---|---|---|
| 65 | 197.5533 | 2 | 128 | 7.442e-40 | 0.7553 | 16.6385 | 5.3903 | 73.0152 |
| Coder | Document-average score |
|---|---|
| Claude | 2.2302 |
| Codex | 2.9457 |
| Gemini | 2.5848 |
Original pooled paired tests — historical reproduction only
These tests treat document–field pairs as independent despite within-document dependence. They are retained here for auditability; the revised paper uses the document-level tests below. Pair-available matching includes one additional Codex–Gemini unit; replacement tests use common-valid matching across all three coders.
| Comparison | Paired field units | Mean difference | t | Unadjusted p |
|---|---|---|---|---|
| Codex - Claude | 1727 | 0.714 | 35.4258 | 4.612e-207 |
| Gemini - Claude | 1727 | 0.3509 | 17.5079 | 2.651e-63 |
| Codex - Gemini | 1728 | 0.3628 | 17.7897 | 3.846e-65 |
| GG ε | Adjusted df 1 | Adjusted df 2 | Adjusted p |
|---|---|---|---|
| 0.9209 | 1.8419 | 117.8786 | 6.497e-37 |
The sphericity correction does not establish equal category spacing.
Document-level paired comparisons
Six tests: all ordinal fields and the ten ethical-principle fields, with three comparisons each. Paired t-tests use 65 document means and 64 degrees of freedom. Holm adjustment applies separately to the three comparisons within each analysis. Confidence intervals are pointwise 95% intervals, not simultaneous intervals. Mean differences still use arithmetic category distances.
| Field set | Comparison | Documents | Difference | CI low | CI high | t | df | Raw p | Holm p | Higher docs | Tied docs | Lower docs |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| all_ordinal | Codex - Claude | 65 | 0.7155 | 0.6535 | 0.7775 | 23.0534 | 64 | 1.056e-32 | 3.167e-32 | 65 | 0 | 0 |
| all_ordinal | Gemini - Claude | 65 | 0.3546 | 0.2741 | 0.435 | 8.8073 | 64 | 1.228e-12 | 1.228e-12 | 54 | 1 | 10 |
| all_ordinal | Codex - Gemini | 65 | 0.3609 | 0.2888 | 0.433 | 9.9982 | 64 | 1.062e-14 | 2.125e-14 | 56 | 1 | 8 |
| ethical_principles | Codex - Claude | 65 | 0.7677 | 0.6888 | 0.8466 | 19.448 | 64 | 1.47e-28 | 4.409e-28 | 65 | 0 | 0 |
| ethical_principles | Gemini - Claude | 65 | 0.36 | 0.2624 | 0.4576 | 7.3662 | 64 | 4.224e-10 | 4.224e-10 | 50 | 7 | 8 |
| ethical_principles | Codex - Gemini | 65 | 0.4077 | 0.3252 | 0.4901 | 9.8781 | 64 | 1.706e-14 | 3.413e-14 | 52 | 10 | 3 |
Checks using category order only
Within each matched field, assign +1 when the first coder is higher, 0 when tied, and −1 when lower. Average within documents: this is the fraction higher minus the fraction lower. Any strictly increasing relabeling of the categories leaves the result unchanged. Percentile intervals resample whole documents 30,000 times (seed 20260907). Exact two-sided binomial sign tests compare positive and negative document totals, omit document ties, and use Holm adjustment within each three-comparison field set.
Codex–Claude has net higher ratings in all 65 documents, but not on every field: 980 higher, 692 tied and 55 lower matched field ratings.
| Field set | Comparison | Mean net higher fraction | Bootstrap CI low | Bootstrap CI high | Higher docs | Tied docs | Lower docs | Raw sign p | Holm sign p | Higher fields | Tied fields | Lower fields |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| all_ordinal | Codex - Claude | 0.5366 | 0.4977 | 0.5763 | 65 | 0 | 0 | 5.421e-20 | 1.626e-19 | 980 | 692 | 55 |
| all_ordinal | Gemini - Claude | 0.2808 | 0.2215 | 0.3398 | 54 | 4 | 7 | 4.322e-10 | 4.322e-10 | 653 | 901 | 173 |
| all_ordinal | Codex - Gemini | 0.2968 | 0.2434 | 0.3496 | 57 | 1 | 7 | 7.638e-11 | 1.528e-10 | 668 | 906 | 153 |
| ethical_principles | Codex - Claude | 0.6092 | 0.5569 | 0.6615 | 65 | 0 | 0 | 5.421e-20 | 1.626e-19 | 411 | 224 | 15 |
| ethical_principles | Gemini - Claude | 0.2908 | 0.2123 | 0.3692 | 50 | 4 | 11 | 4.589e-07 | 4.589e-07 | 256 | 327 | 67 |
| ethical_principles | Codex - Gemini | 0.3631 | 0.2938 | 0.4323 | 53 | 8 | 4 | 5.911e-12 | 1.182e-11 | 273 | 340 | 37 |
Clustered ordinal models
Cumulative-logit ordinal GEE with coder indicators, fixed field effects, document clusters, working independence and robust sandwich covariance. This is a marginal model, not a mixed-effects model. It requires ordered categories without equal distances. The robust covariance aggregates whole-document contributions; working independence does not make individual ratings independent for uncertainty.
Reported contrasts multiply covariance by G/(G−1), use t(G−1) reference, and apply Holm adjustment to the three comparisons within each fit. The intervals are pointwise 95%. Odds ratios summarize the odds of exceeding a score threshold, not ethical intensity or correctness. A common odds ratio assumes proportional odds; the threshold diagnostics below do not establish that assumption.
| Model | Comparison | Documents | Fields | Ratings | Odds ratio | CI low | CI high | Log odds | SE | Raw p | Holm p |
|---|---|---|---|---|---|---|---|---|---|---|---|
| all_ordinal | Codex - Claude | 65 | 29 | 5181 | 3.9724 | 3.2884 | 4.7987 | 1.3794 | 0.0946 | 5.119e-22 | 1.536e-21 |
| all_ordinal | Gemini - Claude | 65 | 29 | 5181 | 1.9338 | 1.6429 | 2.2761 | 0.6595 | 0.0816 | 2.302e-11 | 2.302e-11 |
| all_ordinal | Codex - Gemini | 65 | 29 | 5181 | 2.0543 | 1.743 | 2.421 | 0.7199 | 0.0822 | 1.519e-12 | 3.039e-12 |
| ethical_principles | Codex - Claude | 65 | 10 | 1950 | 4.1672 | 3.3283 | 5.2175 | 1.4272 | 0.1125 | 3.991e-19 | 1.197e-18 |
| ethical_principles | Gemini - Claude | 65 | 10 | 1950 | 1.8611 | 1.5533 | 2.23 | 0.6212 | 0.0905 | 3.228e-09 | 3.228e-09 |
| ethical_principles | Codex - Gemini | 65 | 10 | 1950 | 2.2391 | 1.8475 | 2.7136 | 0.8061 | 0.0962 | 6.987e-12 | 1.397e-11 |
| without_COR | Codex - Claude | 55 | 29 | 4401 | 3.7933 | 3.0965 | 4.6469 | 1.3332 | 0.1012 | 1.694e-18 | 5.083e-18 |
| without_COR | Gemini - Claude | 55 | 29 | 4401 | 1.7581 | 1.4838 | 2.0832 | 0.5642 | 0.0846 | 1.425e-08 | 1.425e-08 |
| without_COR | Codex - Gemini | 55 | 29 | 4401 | 2.1576 | 1.7869 | 2.6052 | 0.769 | 0.094 | 5.072e-11 | 1.014e-10 |
| without_FRM | Codex - Claude | 53 | 27 | 4176 | 5.2197 | 4.3332 | 6.2876 | 1.6524 | 0.0928 | 8.646e-24 | 2.594e-23 |
| without_FRM | Gemini - Claude | 53 | 27 | 4176 | 2.1974 | 1.7815 | 2.7104 | 0.7873 | 0.1046 | 7.096e-10 | 7.096e-10 |
| without_FRM | Codex - Gemini | 53 | 27 | 4176 | 2.3754 | 1.9812 | 2.8481 | 0.8652 | 0.0904 | 4.709e-13 | 9.419e-13 |
| without_GOV | Codex - Claude | 50 | 29 | 4011 | 3.8928 | 3.1826 | 4.7614 | 1.3591 | 0.1002 | 3.301e-18 | 9.904e-18 |
| without_GOV | Gemini - Claude | 50 | 29 | 4011 | 1.8288 | 1.5253 | 2.1926 | 0.6036 | 0.0903 | 2.032e-08 | 2.032e-08 |
| without_GOV | Codex - Gemini | 50 | 29 | 4011 | 2.1286 | 1.7546 | 2.5824 | 0.7555 | 0.0962 | 3.165e-10 | 6.33e-10 |
| without_INS | Codex - Claude | 51 | 29 | 4089 | 3.663 | 2.9803 | 4.502 | 1.2983 | 0.1027 | 3.416e-17 | 1.025e-16 |
| without_INS | Gemini - Claude | 51 | 29 | 4089 | 1.8134 | 1.5146 | 2.1712 | 0.5952 | 0.0897 | 2.203e-08 | 2.203e-08 |
| without_INS | Codex - Gemini | 51 | 29 | 4089 | 2.02 | 1.6795 | 2.4295 | 0.7031 | 0.0919 | 5.813e-10 | 1.163e-09 |
| without_JRN | Codex - Claude | 51 | 28 | 4047 | 3.9528 | 3.1546 | 4.9528 | 1.3744 | 0.1123 | 1.171e-16 | 3.514e-16 |
| without_JRN | Gemini - Claude | 51 | 28 | 4047 | 2.2582 | 1.8852 | 2.7051 | 0.8146 | 0.0899 | 3.976e-12 | 7.951e-12 |
| without_JRN | Codex - Gemini | 51 | 28 | 4047 | 1.7504 | 1.4943 | 2.0504 | 0.5598 | 0.0788 | 4.077e-09 | 4.077e-09 |
All contrasts retain positive directions and intervals above 1 in the ethical-principle subset and when each sector is omitted. Sector omission is a sensitivity check, not proof of generalization to new sectors.
Every threshold diagnostic
Separate binomial GEE fits compare scores above 0, 1, 2, 3 and 4. Fields with no binary outcome variation are omitted at each cutoff; differences between thresholds can therefore reflect both effect variation and field composition. All estimated directions are positive, but effect magnitudes vary.
| Model | Comparison | Documents | Fields | Ratings | Odds ratio | CI low | CI high | Log odds | SE | Raw p | Holm p |
|---|---|---|---|---|---|---|---|---|---|---|---|
| above_0 | Codex - Claude | 65 | 24 | 4206 | 1.7326 | 1.4253 | 2.1062 | 0.5496 | 0.0977 | 4.414e-07 | 1.324e-06 |
| above_0 | Gemini - Claude | 65 | 24 | 4206 | 1.5823 | 1.1541 | 2.1693 | 0.4589 | 0.158 | 0.005 | 0.0101 |
| above_0 | Codex - Gemini | 65 | 24 | 4206 | 1.095 | 0.8309 | 1.4429 | 0.0907 | 0.1381 | 0.5136 | 0.5136 |
| above_1 | Codex - Claude | 65 | 29 | 5181 | 2.796 | 2.2459 | 3.4809 | 1.0282 | 0.1097 | 1.255e-13 | 3.764e-13 |
| above_1 | Gemini - Claude | 65 | 29 | 5181 | 1.9098 | 1.5246 | 2.3923 | 0.647 | 0.1128 | 2.831e-07 | 5.662e-07 |
| above_1 | Codex - Gemini | 65 | 29 | 5181 | 1.464 | 1.1451 | 1.8718 | 0.3812 | 0.123 | 0.0029 | 0.0029 |
| above_2 | Codex - Claude | 65 | 29 | 5181 | 5.0269 | 4.0164 | 6.2917 | 1.6148 | 0.1123 | 1.037e-21 | 3.11e-21 |
| above_2 | Gemini - Claude | 65 | 29 | 5181 | 1.9613 | 1.6199 | 2.3747 | 0.6736 | 0.0957 | 1.61e-09 | 1.61e-09 |
| above_2 | Codex - Gemini | 65 | 29 | 5181 | 2.563 | 2.0775 | 3.162 | 0.9412 | 0.1051 | 6.846e-13 | 1.369e-12 |
| above_3 | Codex - Claude | 65 | 27 | 5109 | 6.3844 | 5.0673 | 8.0437 | 1.8539 | 0.1157 | 4.323e-24 | 1.297e-23 |
| above_3 | Gemini - Claude | 65 | 27 | 5109 | 2.4904 | 2.0062 | 3.0914 | 0.9124 | 0.1082 | 5.592e-12 | 5.592e-12 |
| above_3 | Codex - Gemini | 65 | 27 | 5109 | 2.5636 | 2.0722 | 3.1716 | 0.9414 | 0.1065 | 1.087e-12 | 2.174e-12 |
| above_4 | Codex - Claude | 65 | 19 | 3552 | 2.5737 | 1.7288 | 3.8316 | 0.9454 | 0.1992 | 1.207e-05 | 3.621e-05 |
| above_4 | Gemini - Claude | 65 | 19 | 3552 | 1.3099 | 0.8511 | 2.016 | 0.27 | 0.2158 | 0.2156 | 0.2156 |
| above_4 | Codex - Gemini | 65 | 19 | 3552 | 1.9648 | 1.3706 | 2.8166 | 0.6754 | 0.1803 | 0.0003868 | 0.0007736 |
Convergence and numerical verification
All 12 GEE fits converged. Independent estimating-equation and sandwich-covariance calculations verified the implementations. Low-threshold full nuisance covariance matrices can be rank-deficient, while the coder-coefficient covariance blocks are full-rank. Numerical square-root warnings or NaN nuisance standard errors can appear in the printed summaries; no full-parameter Wald tests are interpreted.
Printed ordinal model summaries contain 25,905 internally expanded threshold rows for the full fit. The research data comprise 5,181 ratings clustered in 65 documents. Raw statsmodels summary p-values differ from the finite-cluster and Holm-adjusted contrasts above.
| Model | Converged | Iterations | Parameters | Full covariance rank | Coder covariance rank | Max average score | Sandwich discrepancy |
|---|---|---|---|---|---|---|---|
| all_ordinal | 1 | 2 | 35 | 35 | 2 | 8.64e-18 | 4.295e-15 |
| ethical_principles | 1 | 2 | 16 | 16 | 2 | 1.075e-17 | 1.596e-15 |
| without_COR | 1 | 2 | 35 | 35 | 2 | 7.547e-18 | 5.759e-15 |
| without_FRM | 1 | 2 | 33 | 33 | 2 | 6.093e-18 | 4.871e-15 |
| without_GOV | 1 | 2 | 35 | 35 | 2 | 1.072e-17 | 7.619e-15 |
| without_INS | 1 | 2 | 35 | 35 | 2 | 1.058e-17 | 5.586e-15 |
| without_JRN | 1 | 2 | 34 | 34 | 2 | 5.712e-18 | 6.483e-15 |
| above_0 | 1 | 2 | 26 | 23 | 2 | 1.93e-17 | 8.932e-14 |
| above_1 | 1 | 2 | 31 | 30 | 2 | 4.2e-17 | 1.531e-12 |
| above_2 | 1 | 2 | 31 | 31 | 2 | 1.346e-17 | 7.931e-15 |
| above_3 | 1 | 2 | 29 | 29 | 2 | 4.39e-17 | 5.593e-15 |
| above_4 | 1 | 2 | 21 | 21 | 2 | 1.113e-17 | 1.038e-13 |
all_ordinal: full coefficients, exclusions and diagnostics
Download all coefficients · Download complete model summary
{
"model": "all_ordinal",
"converged": true,
"iterations": 2,
"warnings": [],
"n_parameters": 35,
"covariance_rank": 35,
"smallest_cov_eigenvalue": 0.00015774969186709627,
"max_absolute_parameter": 5.38613788718312,
"threshold": null,
"fields_excluded_for_no_binary_variation": [],
"max_average_score": 8.640068008663638e-18,
"sandwich_max_difference": 4.295175326518574e-15,
"coder_covariance_rank": 2
} OrdinalGEE Regression Results
===================================================================================
Dep. Variable: rating No. Observations: 25905
Model: OrdinalGEE No. clusters: 65
Method: Generalized Min. cluster size: 390
Estimating Equations Max. cluster size: 420
Family: Binomial Mean cluster size: 398.5
Dependence structure: Independence Num. iterations: 2
Date: Mon, 07 Sep 2026 Scale: 1.000
Covariance type: robust Time: 05:45:07
========================================================================================================================
coef std err z P>|z| [0.025 0.975]
------------------------------------------------------------------------------------------------------------------------
I(y>0.0) 2.8860 0.197 14.648 0.000 2.500 3.272
I(y>1.0) 2.1584 0.184 11.704 0.000 1.797 2.520
I(y>2.0) 0.5459 0.136 4.027 0.000 0.280 0.812
I(y>3.0) -1.0258 0.140 -7.322 0.000 -1.300 -0.751
I(y>4.0) -4.1852 0.205 -20.444 0.000 -4.586 -3.784
Codex 1.3794 0.094 14.695 0.000 1.195 1.563
Gemini 0.6595 0.081 8.146 0.000 0.501 0.818
field_cb_benefit_presence_strength_code 0.3152 0.059 5.374 0.000 0.200 0.430
field_cb_cost_intensity_code 0.2942 0.216 1.360 0.174 -0.130 0.718
field_cb_cost_presence_strength_code 0.5520 0.229 2.411 0.016 0.103 1.001
field_detection_verification_presence_strength_code -4.9739 0.556 -8.942 0.000 -6.064 -3.884
field_dppra_a_strength_code -1.5012 0.220 -6.822 0.000 -1.932 -1.070
field_dppra_alignment_score_code -1.8558 0.196 -9.468 0.000 -2.240 -1.472
field_dppra_d_strength_code 2.6520 0.296 8.949 0.000 2.071 3.233
field_dppra_p1_strength_code -0.5459 0.214 -2.552 0.011 -0.965 -0.127
field_dppra_p2_strength_code -1.7762 0.211 -8.427 0.000 -2.189 -1.363
field_dppra_r_strength_code -2.7500 0.246 -11.166 0.000 -3.233 -2.267
field_eth_accountability_code 0.8967 0.311 2.882 0.004 0.287 1.507
field_eth_autonomy_code -0.5006 0.216 -2.320 0.020 -0.923 -0.078
field_eth_beneficence_eth_code 0.3047 0.234 1.300 0.194 -0.155 0.764
field_eth_beneficence_mech_code -0.3815 0.103 -3.709 0.000 -0.583 -0.180
field_eth_equity_code -0.7681 0.258 -2.972 0.003 -1.275 -0.261
field_eth_integrity_code 0.7210 0.212 3.399 0.001 0.305 1.137
field_eth_justice_code -0.4000 0.269 -1.486 0.137 -0.927 0.128
field_eth_nonmaleficence_code 0.3362 0.265 1.266 0.205 -0.184 0.857
field_eth_privacy_code -1.0014 0.209 -4.782 0.000 -1.412 -0.591
field_eth_transparency_code 0.5520 0.262 2.111 0.035 0.039 1.065
field_int_academic_intensity_code -5.3861 0.419 -12.865 0.000 -6.207 -4.566
field_int_assessment_intensity_code -4.3880 0.342 -12.846 0.000 -5.057 -3.718
field_int_data_prov_intensity_code -1.3783 0.198 -6.962 0.000 -1.766 -0.990
field_int_epistemic_intensity_code -1.1117 0.175 -6.366 0.000 -1.454 -0.769
field_int_research_intensity_code -4.1478 0.288 -14.422 0.000 -4.711 -3.584
field_int_system_tevv_intensity_code -0.7768 0.229 -3.392 0.001 -1.226 -0.328
field_journal_evidence_quality_strength_code -0.6483 0.470 -1.379 0.168 -1.570 0.273
field_surveillance_monitoring_presence_strength_code -4.0868 0.578 -7.076 0.000 -5.219 -2.955
==============================================================================
Skew: -0.2025 Kurtosis: 1.6538
Centered skew: -0.0425 Centered kurtosis: 1.3716
==============================================================================
Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
ethical_principles: full coefficients, exclusions and diagnostics
Download all coefficients · Download complete model summary
{
"model": "ethical_principles",
"converged": true,
"iterations": 2,
"warnings": [],
"n_parameters": 16,
"covariance_rank": 16,
"smallest_cov_eigenvalue": 0.0007437373155553275,
"max_absolute_parameter": 3.8656625982654456,
"threshold": null,
"fields_excluded_for_no_binary_variation": [],
"max_average_score": 1.0749236258934849e-17,
"sandwich_max_difference": 1.5959455978986625e-15,
"coder_covariance_rank": 2
} OrdinalGEE Regression Results
===================================================================================
Dep. Variable: rating No. Observations: 9750
Model: OrdinalGEE No. clusters: 65
Method: Generalized Min. cluster size: 150
Estimating Equations Max. cluster size: 150
Family: Binomial Mean cluster size: 150.0
Dependence structure: Independence Num. iterations: 2
Date: Mon, 07 Sep 2026 Scale: 1.000
Covariance type: robust Time: 05:45:08
===================================================================================================
coef std err z P>|z| [0.025 0.975]
---------------------------------------------------------------------------------------------------
I(y>0.0) 3.8657 0.454 8.517 0.000 2.976 4.755
I(y>1.0) 2.6961 0.370 7.287 0.000 1.971 3.421
I(y>2.0) 1.3913 0.296 4.700 0.000 0.811 1.971
I(y>3.0) -0.0941 0.247 -0.381 0.703 -0.578 0.390
I(y>4.0) -3.1850 0.221 -14.403 0.000 -3.618 -2.752
Codex 1.4272 0.112 12.784 0.000 1.208 1.646
Gemini 0.6212 0.090 6.917 0.000 0.445 0.797
field_eth_autonomy_code -1.3421 0.294 -4.564 0.000 -1.918 -0.766
field_eth_beneficence_eth_code -0.5689 0.318 -1.788 0.074 -1.193 0.055
field_eth_beneficence_mech_code -1.2278 0.334 -3.678 0.000 -1.882 -0.573
field_eth_equity_code -1.5987 0.266 -6.014 0.000 -2.120 -1.078
field_eth_integrity_code -0.1688 0.226 -0.748 0.455 -0.612 0.274
field_eth_justice_code -1.2455 0.253 -4.926 0.000 -1.741 -0.750
field_eth_nonmaleficence_code -0.5386 0.221 -2.434 0.015 -0.972 -0.105
field_eth_privacy_code -1.8227 0.275 -6.620 0.000 -2.362 -1.283
field_eth_transparency_code -0.3312 0.192 -1.721 0.085 -0.708 0.046
==============================================================================
Skew: -0.5337 Kurtosis: 1.5030
Centered skew: -0.2363 Centered kurtosis: 1.1091
==============================================================================
Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
without_COR: full coefficients, exclusions and diagnostics
Download all coefficients · Download complete model summary
{
"model": "without_COR",
"converged": true,
"iterations": 2,
"warnings": [],
"n_parameters": 35,
"covariance_rank": 35,
"smallest_cov_eigenvalue": 0.0001067196496657633,
"max_absolute_parameter": 5.17718555018618,
"threshold": null,
"fields_excluded_for_no_binary_variation": [],
"max_average_score": 7.546796035428525e-18,
"sandwich_max_difference": 5.7592819402429996e-15,
"coder_covariance_rank": 2
} OrdinalGEE Regression Results
===================================================================================
Dep. Variable: rating No. Observations: 22005
Model: OrdinalGEE No. clusters: 55
Method: Generalized Min. cluster size: 390
Estimating Equations Max. cluster size: 420
Family: Binomial Mean cluster size: 400.1
Dependence structure: Independence Num. iterations: 2
Date: Mon, 07 Sep 2026 Scale: 1.000
Covariance type: robust Time: 05:45:09
========================================================================================================================
coef std err z P>|z| [0.025 0.975]
------------------------------------------------------------------------------------------------------------------------
I(y>0.0) 2.9315 0.227 12.896 0.000 2.486 3.377
I(y>1.0) 2.2039 0.214 10.314 0.000 1.785 2.623
I(y>2.0) 0.5453 0.154 3.530 0.000 0.243 0.848
I(y>3.0) -0.9913 0.161 -6.144 0.000 -1.308 -0.675
I(y>4.0) -4.0852 0.231 -17.667 0.000 -4.538 -3.632
Codex 1.3332 0.100 13.291 0.000 1.137 1.530
Gemini 0.5642 0.084 6.729 0.000 0.400 0.729
field_cb_benefit_presence_strength_code 0.3085 0.065 4.772 0.000 0.182 0.435
field_cb_cost_intensity_code 0.3949 0.230 1.714 0.087 -0.057 0.847
field_cb_cost_presence_strength_code 0.6902 0.244 2.828 0.005 0.212 1.169
field_detection_verification_presence_strength_code -4.9628 0.563 -8.816 0.000 -6.066 -3.860
field_dppra_a_strength_code -1.6130 0.240 -6.730 0.000 -2.083 -1.143
field_dppra_alignment_score_code -2.0246 0.220 -9.182 0.000 -2.457 -1.592
field_dppra_d_strength_code 2.6698 0.343 7.790 0.000 1.998 3.342
field_dppra_p1_strength_code -0.6984 0.221 -3.164 0.002 -1.131 -0.266
field_dppra_p2_strength_code -2.0218 0.233 -8.670 0.000 -2.479 -1.565
field_dppra_r_strength_code -3.1145 0.269 -11.558 0.000 -3.643 -2.586
field_eth_accountability_code 0.9216 0.353 2.611 0.009 0.230 1.614
field_eth_autonomy_code -0.4876 0.241 -2.027 0.043 -0.959 -0.016
field_eth_beneficence_eth_code 0.4575 0.273 1.677 0.094 -0.077 0.992
field_eth_beneficence_mech_code -0.4124 0.116 -3.553 0.000 -0.640 -0.185
field_eth_equity_code -0.5302 0.283 -1.871 0.061 -1.086 0.025
field_eth_integrity_code 0.6244 0.234 2.672 0.008 0.166 1.082
field_eth_justice_code 0.0115 0.296 0.039 0.969 -0.569 0.592
field_eth_nonmaleficence_code 0.1284 0.271 0.473 0.636 -0.403 0.660
field_eth_privacy_code -0.8530 0.232 -3.682 0.000 -1.307 -0.399
field_eth_transparency_code 0.5466 0.296 1.847 0.065 -0.033 1.127
field_int_academic_intensity_code -5.1772 0.434 -11.928 0.000 -6.028 -4.327
field_int_assessment_intensity_code -4.2463 0.362 -11.716 0.000 -4.957 -3.536
field_int_data_prov_intensity_code -1.2636 0.212 -5.969 0.000 -1.679 -0.849
field_int_epistemic_intensity_code -1.1152 0.187 -5.971 0.000 -1.481 -0.749
field_int_research_intensity_code -4.0958 0.327 -12.512 0.000 -4.737 -3.454
field_int_system_tevv_intensity_code -1.0051 0.230 -4.366 0.000 -1.456 -0.554
field_journal_evidence_quality_strength_code -0.6282 0.478 -1.316 0.188 -1.564 0.308
field_surveillance_monitoring_presence_strength_code -4.0749 0.586 -6.948 0.000 -5.224 -2.925
==============================================================================
Skew: -0.1743 Kurtosis: 1.6266
Centered skew: 0.0100 Centered kurtosis: 1.2916
==============================================================================
Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
without_FRM: full coefficients, exclusions and diagnostics
Download all coefficients · Download complete model summary
{
"model": "without_FRM",
"converged": true,
"iterations": 2,
"warnings": [],
"n_parameters": 33,
"covariance_rank": 33,
"smallest_cov_eigenvalue": 0.00018070086809806227,
"max_absolute_parameter": 6.150103947656452,
"threshold": null,
"fields_excluded_for_no_binary_variation": [],
"max_average_score": 6.093465451247267e-18,
"sandwich_max_difference": 4.871103520542874e-15,
"coder_covariance_rank": 2
} OrdinalGEE Regression Results
===================================================================================
Dep. Variable: rating No. Observations: 20880
Model: OrdinalGEE No. clusters: 53
Method: Generalized Min. cluster size: 390
Estimating Equations Max. cluster size: 405
Family: Binomial Mean cluster size: 394.0
Dependence structure: Independence Num. iterations: 2
Date: Mon, 07 Sep 2026 Scale: 1.000
Covariance type: robust Time: 05:45:09
================================================================================================================
coef std err z P>|z| [0.025 0.975]
----------------------------------------------------------------------------------------------------------------
I(y>0.0) 3.5312 0.195 18.128 0.000 3.149 3.913
I(y>1.0) 2.7485 0.183 15.048 0.000 2.390 3.106
I(y>2.0) 0.7376 0.158 4.666 0.000 0.428 1.047
I(y>3.0) -1.0432 0.176 -5.935 0.000 -1.388 -0.699
I(y>4.0) -4.5256 0.250 -18.071 0.000 -5.016 -4.035
Codex 1.6524 0.092 17.984 0.000 1.472 1.833
Gemini 0.7873 0.104 7.601 0.000 0.584 0.990
field_cb_benefit_presence_strength_code 0.4569 0.074 6.162 0.000 0.312 0.602
field_cb_cost_intensity_code 0.2161 0.265 0.815 0.415 -0.303 0.735
field_cb_cost_presence_strength_code 0.5190 0.278 1.865 0.062 -0.026 1.064
field_dppra_a_strength_code -1.4323 0.282 -5.086 0.000 -1.984 -0.880
field_dppra_alignment_score_code -2.0701 0.241 -8.600 0.000 -2.542 -1.598
field_dppra_d_strength_code 3.0405 0.329 9.240 0.000 2.396 3.685
field_dppra_p1_strength_code -0.2490 0.272 -0.914 0.361 -0.783 0.285
field_dppra_p2_strength_code -2.0019 0.270 -7.412 0.000 -2.531 -1.473
field_dppra_r_strength_code -3.0552 0.301 -10.167 0.000 -3.644 -2.466
field_eth_accountability_code 1.6025 0.343 4.675 0.000 0.931 2.274
field_eth_autonomy_code -0.7591 0.234 -3.241 0.001 -1.218 -0.300
field_eth_beneficence_eth_code 0.1722 0.253 0.679 0.497 -0.325 0.669
field_eth_beneficence_mech_code -0.3965 0.144 -2.757 0.006 -0.678 -0.115
field_eth_equity_code -0.4230 0.338 -1.250 0.211 -1.086 0.240
field_eth_integrity_code 0.7428 0.263 2.830 0.005 0.228 1.257
field_eth_justice_code -0.4098 0.344 -1.191 0.234 -1.084 0.265
field_eth_nonmaleficence_code 0.5977 0.334 1.789 0.074 -0.057 1.252
field_eth_privacy_code -0.7718 0.252 -3.068 0.002 -1.265 -0.279
field_eth_transparency_code 1.2989 0.274 4.734 0.000 0.761 1.837
field_int_academic_intensity_code -6.1501 0.500 -12.303 0.000 -7.130 -5.170
field_int_assessment_intensity_code -5.0865 0.400 -12.709 0.000 -5.871 -4.302
field_int_data_prov_intensity_code -1.3130 0.237 -5.543 0.000 -1.777 -0.849
field_int_epistemic_intensity_code -1.2408 0.220 -5.648 0.000 -1.671 -0.810
field_int_research_intensity_code -4.7137 0.309 -15.263 0.000 -5.319 -4.108
field_int_system_tevv_intensity_code -0.3433 0.265 -1.296 0.195 -0.862 0.176
field_journal_evidence_quality_strength_code -0.9825 0.532 -1.846 0.065 -2.026 0.061
==============================================================================
Skew: -0.1733 Kurtosis: 2.4613
Centered skew: -0.1133 Centered kurtosis: 2.2683
==============================================================================
Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
without_GOV: full coefficients, exclusions and diagnostics
Download all coefficients · Download complete model summary
{
"model": "without_GOV",
"converged": true,
"iterations": 2,
"warnings": [],
"n_parameters": 35,
"covariance_rank": 35,
"smallest_cov_eigenvalue": 0.0001501614592392729,
"max_absolute_parameter": 5.180682250251609,
"threshold": null,
"fields_excluded_for_no_binary_variation": [],
"max_average_score": 1.07174857924423e-17,
"sandwich_max_difference": 7.618905506490137e-15,
"coder_covariance_rank": 2
} OrdinalGEE Regression Results
===================================================================================
Dep. Variable: rating No. Observations: 20055
Model: OrdinalGEE No. clusters: 50
Method: Generalized Min. cluster size: 390
Estimating Equations Max. cluster size: 420
Family: Binomial Mean cluster size: 401.1
Dependence structure: Independence Num. iterations: 2
Date: Mon, 07 Sep 2026 Scale: 1.000
Covariance type: robust Time: 05:45:10
========================================================================================================================
coef std err z P>|z| [0.025 0.975]
------------------------------------------------------------------------------------------------------------------------
I(y>0.0) 2.7548 0.213 12.939 0.000 2.338 3.172
I(y>1.0) 2.0206 0.199 10.142 0.000 1.630 2.411
I(y>2.0) 0.4339 0.154 2.817 0.005 0.132 0.736
I(y>3.0) -1.1695 0.158 -7.415 0.000 -1.479 -0.860
I(y>4.0) -4.1827 0.245 -17.043 0.000 -4.664 -3.702
Codex 1.3591 0.099 13.698 0.000 1.165 1.554
Gemini 0.6036 0.089 6.754 0.000 0.428 0.779
field_cb_benefit_presence_strength_code 0.2798 0.062 4.542 0.000 0.159 0.401
field_cb_cost_intensity_code 0.3455 0.259 1.332 0.183 -0.163 0.854
field_cb_cost_presence_strength_code 0.5889 0.268 2.194 0.028 0.063 1.115
field_detection_verification_presence_strength_code -4.8145 0.562 -8.560 0.000 -5.917 -3.712
field_dppra_a_strength_code -1.8388 0.242 -7.599 0.000 -2.313 -1.365
field_dppra_alignment_score_code -1.9775 0.228 -8.673 0.000 -2.424 -1.531
field_dppra_d_strength_code 2.8506 0.339 8.404 0.000 2.186 3.515
field_dppra_p1_strength_code -0.9506 0.222 -4.283 0.000 -1.386 -0.516
field_dppra_p2_strength_code -1.8181 0.256 -7.098 0.000 -2.320 -1.316
field_dppra_r_strength_code -2.8707 0.318 -9.029 0.000 -3.494 -2.248
field_eth_accountability_code 0.6868 0.353 1.946 0.052 -0.005 1.378
field_eth_autonomy_code -0.3149 0.250 -1.258 0.208 -0.805 0.176
field_eth_beneficence_eth_code 0.3455 0.271 1.274 0.203 -0.186 0.877
field_eth_beneficence_mech_code -0.4554 0.109 -4.183 0.000 -0.669 -0.242
field_eth_equity_code -0.9506 0.293 -3.245 0.001 -1.525 -0.376
field_eth_integrity_code 0.8449 0.252 3.353 0.001 0.351 1.339
field_eth_justice_code -0.6389 0.308 -2.076 0.038 -1.242 -0.036
field_eth_nonmaleficence_code 0.2021 0.301 0.672 0.501 -0.387 0.792
field_eth_privacy_code -1.2423 0.239 -5.189 0.000 -1.712 -0.773
field_eth_transparency_code 0.4119 0.299 1.377 0.168 -0.174 0.998
field_int_academic_intensity_code -5.1807 0.469 -11.049 0.000 -6.100 -4.262
field_int_assessment_intensity_code -4.2061 0.386 -10.906 0.000 -4.962 -3.450
field_int_data_prov_intensity_code -1.5900 0.222 -7.151 0.000 -2.026 -1.154
field_int_epistemic_intensity_code -1.0596 0.194 -5.455 0.000 -1.440 -0.679
field_int_research_intensity_code -3.9160 0.327 -11.982 0.000 -4.557 -3.275
field_int_system_tevv_intensity_code -1.0920 0.259 -4.208 0.000 -1.601 -0.583
field_journal_evidence_quality_strength_code -0.4984 0.479 -1.040 0.298 -1.438 0.441
field_surveillance_monitoring_presence_strength_code -3.9280 0.585 -6.716 0.000 -5.074 -2.782
==============================================================================
Skew: -0.1247 Kurtosis: 1.4727
Centered skew: 0.0008 Centered kurtosis: 1.2363
==============================================================================
Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
without_INS: full coefficients, exclusions and diagnostics
Download all coefficients · Download complete model summary
{
"model": "without_INS",
"converged": true,
"iterations": 2,
"warnings": [],
"n_parameters": 35,
"covariance_rank": 35,
"smallest_cov_eigenvalue": 0.00011407610041409381,
"max_absolute_parameter": 5.423740126209203,
"threshold": null,
"fields_excluded_for_no_binary_variation": [],
"max_average_score": 1.0578631392240664e-17,
"sandwich_max_difference": 5.585809592645319e-15,
"coder_covariance_rank": 2
} OrdinalGEE Regression Results
===================================================================================
Dep. Variable: rating No. Observations: 20445
Model: OrdinalGEE No. clusters: 51
Method: Generalized Min. cluster size: 390
Estimating Equations Max. cluster size: 420
Family: Binomial Mean cluster size: 400.9
Dependence structure: Independence Num. iterations: 2
Date: Mon, 07 Sep 2026 Scale: 1.000
Covariance type: robust Time: 05:45:11
========================================================================================================================
coef std err z P>|z| [0.025 0.975]
------------------------------------------------------------------------------------------------------------------------
I(y>0.0) 2.7982 0.214 13.079 0.000 2.379 3.217
I(y>1.0) 2.0957 0.205 10.221 0.000 1.694 2.498
I(y>2.0) 0.5857 0.154 3.806 0.000 0.284 0.887
I(y>3.0) -0.9551 0.154 -6.195 0.000 -1.257 -0.653
I(y>4.0) -3.9780 0.217 -18.358 0.000 -4.403 -3.553
Codex 1.2983 0.102 12.769 0.000 1.099 1.498
Gemini 0.5952 0.089 6.705 0.000 0.421 0.769
field_cb_benefit_presence_strength_code 0.2882 0.060 4.772 0.000 0.170 0.407
field_cb_cost_intensity_code 0.1734 0.240 0.724 0.469 -0.296 0.643
field_cb_cost_presence_strength_code 0.4058 0.253 1.604 0.109 -0.090 0.902
field_detection_verification_presence_strength_code -4.8500 0.559 -8.670 0.000 -5.946 -3.754
field_dppra_a_strength_code -1.5103 0.252 -6.001 0.000 -2.004 -1.017
field_dppra_alignment_score_code -1.7980 0.223 -8.066 0.000 -2.235 -1.361
field_dppra_d_strength_code 2.3456 0.318 7.387 0.000 1.723 2.968
field_dppra_p1_strength_code -0.5679 0.252 -2.253 0.024 -1.062 -0.074
field_dppra_p2_strength_code -1.8084 0.234 -7.739 0.000 -2.266 -1.350
field_dppra_r_strength_code -2.7886 0.281 -9.911 0.000 -3.340 -2.237
field_eth_accountability_code 0.6227 0.336 1.853 0.064 -0.036 1.281
field_eth_autonomy_code -0.5679 0.256 -2.215 0.027 -1.071 -0.065
field_eth_beneficence_eth_code 0.1609 0.270 0.597 0.551 -0.368 0.689
field_eth_beneficence_mech_code -0.1797 0.095 -1.885 0.059 -0.367 0.007
field_eth_equity_code -1.0405 0.286 -3.638 0.000 -1.601 -0.480
field_eth_integrity_code 0.6088 0.230 2.648 0.008 0.158 1.059
field_eth_justice_code -0.5346 0.305 -1.751 0.080 -1.133 0.064
field_eth_nonmaleficence_code 0.1609 0.296 0.544 0.586 -0.419 0.740
field_eth_privacy_code -1.1650 0.228 -5.116 0.000 -1.611 -0.719
field_eth_transparency_code 0.3271 0.286 1.143 0.253 -0.234 0.888
field_int_academic_intensity_code -5.4237 0.428 -12.666 0.000 -6.263 -4.584
field_int_assessment_intensity_code -4.3075 0.374 -11.506 0.000 -5.041 -3.574
field_int_data_prov_intensity_code -1.5504 0.234 -6.637 0.000 -2.008 -1.093
field_int_epistemic_intensity_code -1.1340 0.196 -5.798 0.000 -1.517 -0.751
field_int_research_intensity_code -4.2121 0.327 -12.881 0.000 -4.853 -3.571
field_int_system_tevv_intensity_code -0.7862 0.269 -2.923 0.003 -1.313 -0.259
field_journal_evidence_quality_strength_code -0.6278 0.468 -1.343 0.179 -1.544 0.289
field_surveillance_monitoring_presence_strength_code -3.9733 0.579 -6.860 0.000 -5.109 -2.838
==============================================================================
Skew: -0.1869 Kurtosis: 1.4472
Centered skew: -0.0366 Centered kurtosis: 1.1856
==============================================================================
Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
without_JRN: full coefficients, exclusions and diagnostics
Download all coefficients · Download complete model summary
{
"model": "without_JRN",
"converged": true,
"iterations": 2,
"warnings": [],
"n_parameters": 34,
"covariance_rank": 34,
"smallest_cov_eigenvalue": 0.00014102865305170537,
"max_absolute_parameter": 5.366593655886791,
"threshold": null,
"fields_excluded_for_no_binary_variation": [],
"max_average_score": 5.711599548479308e-18,
"sandwich_max_difference": 6.482661629725328e-15,
"coder_covariance_rank": 2
} OrdinalGEE Regression Results
===================================================================================
Dep. Variable: rating No. Observations: 20235
Model: OrdinalGEE No. clusters: 51
Method: Generalized Min. cluster size: 390
Estimating Equations Max. cluster size: 420
Family: Binomial Mean cluster size: 396.8
Dependence structure: Independence Num. iterations: 2
Date: Mon, 07 Sep 2026 Scale: 1.000
Covariance type: robust Time: 05:45:12
========================================================================================================================
coef std err z P>|z| [0.025 0.975]
------------------------------------------------------------------------------------------------------------------------
I(y>0.0) 2.6826 0.206 13.016 0.000 2.279 3.087
I(y>1.0) 1.9391 0.190 10.194 0.000 1.566 2.312
I(y>2.0) 0.4352 0.140 3.105 0.002 0.160 0.710
I(y>3.0) -1.0854 0.140 -7.740 0.000 -1.360 -0.811
I(y>4.0) -4.3995 0.197 -22.295 0.000 -4.786 -4.013
Codex 1.3744 0.111 12.362 0.000 1.156 1.592
Gemini 0.8146 0.089 9.152 0.000 0.640 0.989
field_cb_benefit_presence_strength_code 0.2777 0.065 4.296 0.000 0.151 0.404
field_cb_cost_intensity_code 0.3433 0.228 1.503 0.133 -0.104 0.791
field_cb_cost_presence_strength_code 0.5743 0.250 2.293 0.022 0.083 1.065
field_detection_verification_presence_strength_code -4.8250 0.558 -8.639 0.000 -5.920 -3.730
field_dppra_a_strength_code -1.1880 0.231 -5.149 0.000 -1.640 -0.736
field_dppra_alignment_score_code -1.5488 0.182 -8.523 0.000 -1.905 -1.193
field_dppra_d_strength_code 2.5067 0.320 7.839 0.000 1.880 3.133
field_dppra_p1_strength_code -0.2278 0.241 -0.946 0.344 -0.700 0.244
field_dppra_p2_strength_code -1.3517 0.191 -7.090 0.000 -1.725 -0.978
field_dppra_r_strength_code -2.1388 0.197 -10.847 0.000 -2.525 -1.752
field_eth_accountability_code 0.8342 0.327 2.551 0.011 0.193 1.475
field_eth_autonomy_code -0.4239 0.229 -1.854 0.064 -0.872 0.024
field_eth_beneficence_eth_code 0.3832 0.249 1.539 0.124 -0.105 0.871
field_eth_beneficence_mech_code -0.4804 0.115 -4.189 0.000 -0.705 -0.256
field_eth_equity_code -0.8846 0.262 -3.372 0.001 -1.399 -0.370
field_eth_integrity_code 0.8342 0.218 3.835 0.000 0.408 1.261
field_eth_justice_code -0.4579 0.267 -1.714 0.086 -0.981 0.066
field_eth_nonmaleficence_code 0.6877 0.300 2.290 0.022 0.099 1.276
field_eth_privacy_code -1.0008 0.233 -4.291 0.000 -1.458 -0.544
field_eth_transparency_code 0.3566 0.266 1.341 0.180 -0.165 0.878
field_int_academic_intensity_code -5.3666 0.505 -10.619 0.000 -6.357 -4.376
field_int_assessment_intensity_code -4.4162 0.376 -11.730 0.000 -5.154 -3.678
field_int_data_prov_intensity_code -1.2496 0.215 -5.802 0.000 -1.672 -0.827
field_int_epistemic_intensity_code -1.0844 0.192 -5.640 0.000 -1.461 -0.708
field_int_research_intensity_code -4.1082 0.305 -13.453 0.000 -4.707 -3.510
field_int_system_tevv_intensity_code -0.6140 0.262 -2.340 0.019 -1.128 -0.100
field_surveillance_monitoring_presence_strength_code -3.9449 0.582 -6.782 0.000 -5.085 -2.805
==============================================================================
Skew: -0.2757 Kurtosis: 1.6224
Centered skew: -0.0723 Centered kurtosis: 1.2959
==============================================================================
Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
above_0: full coefficients, exclusions and diagnostics
Download all coefficients · Download complete model summary
{
"model": "above_0",
"converged": true,
"iterations": 2,
"warnings": [],
"n_parameters": 26,
"covariance_rank": 23,
"smallest_cov_eigenvalue": -1.143355803459955e-14,
"max_absolute_parameter": 7.142644151258511,
"threshold": 0,
"fields_excluded_for_no_binary_variation": [
"cb_benefit_intensity_code",
"cb_benefit_presence_strength_code",
"cb_cost_intensity_code",
"cb_cost_presence_strength_code",
"eth_beneficence_eth_code"
],
"max_average_score": 1.9295602258701603e-17,
"sandwich_max_difference": 8.931744233109384e-14,
"coder_covariance_rank": 2
} GEE Regression Results
===================================================================================
Dep. Variable: binary No. Observations: 4206
Model: GEE No. clusters: 65
Method: Generalized Min. cluster size: 63
Estimating Equations Max. cluster size: 69
Family: Binomial Mean cluster size: 64.7
Dependence structure: Independence Num. iterations: 2
Date: Mon, 07 Sep 2026 Scale: 1.000
Covariance type: robust Time: 05:45:12
========================================================================================================================
coef std err z P>|z| [0.025 0.975]
------------------------------------------------------------------------------------------------------------------------
intercept -2.1811 0.526 -4.148 0.000 -3.212 -1.151
Codex 0.5496 0.097 5.667 0.000 0.360 0.740
Gemini 0.4589 0.157 2.928 0.003 0.152 0.766
field_dppra_a_strength_code 3.7838 0.480 7.882 0.000 2.843 4.725
field_dppra_alignment_score_code 4.5791 0.551 8.313 0.000 3.500 5.659
field_dppra_d_strength_code 7.1426 1.025 6.969 0.000 5.134 9.151
field_dppra_p1_strength_code 5.1633 0.569 9.082 0.000 4.049 6.278
field_dppra_p2_strength_code 3.4190 0.519 6.585 0.000 2.401 4.437
field_dppra_r_strength_code 2.4324 0.577 4.212 0.000 1.301 3.564
field_eth_accountability_code 5.5110 0.657 8.388 0.000 4.223 6.799
field_eth_autonomy_code 6.0330 1.047 5.760 0.000 3.980 8.086
field_eth_beneficence_mech_code 7.1426 1.025 6.969 0.000 5.134 9.151
field_eth_equity_code 4.1549 0.581 7.153 0.000 3.017 5.293
field_eth_integrity_code 7.1426 1.025 6.969 0.000 5.134 9.151
field_eth_justice_code 4.7896 0.611 7.839 0.000 3.592 5.987
field_eth_nonmaleficence_code 6.4440 0.750 8.591 0.000 4.974 7.914
field_eth_privacy_code 4.0949 0.659 6.212 0.000 2.803 5.387
field_eth_transparency_code 5.1633 0.625 8.266 0.000 3.939 6.388
field_int_academic_intensity_code 0.1597 0.574 0.278 0.781 -0.966 1.285
field_int_assessment_intensity_code 0.6253 0.604 1.035 0.301 -0.559 1.809
field_int_data_prov_intensity_code 3.5700 0.633 5.636 0.000 2.328 4.811
field_int_epistemic_intensity_code 4.5099 0.621 7.262 0.000 3.293 5.727
field_int_research_intensity_code 1.1650 0.534 2.181 0.029 0.118 2.212
field_int_system_tevv_intensity_code 3.9301 0.494 7.953 0.000 2.962 4.899
field_journal_evidence_quality_strength_code 4.4355 0.760 5.840 0.000 2.947 5.924
field_surveillance_monitoring_presence_strength_code 1.2658 0.635 1.995 0.046 0.022 2.509
==============================================================================
Skew: -1.0879 Kurtosis: 3.9528
Centered skew: -0.5955 Centered kurtosis: 3.3379
==============================================================================
Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
above_1: full coefficients, exclusions and diagnostics
Download all coefficients · Download complete model summary
{
"model": "above_1",
"converged": true,
"iterations": 2,
"warnings": [],
"n_parameters": 31,
"covariance_rank": 30,
"smallest_cov_eigenvalue": -3.668266906557964e-13,
"max_absolute_parameter": 7.552367093744065,
"threshold": 1,
"fields_excluded_for_no_binary_variation": [],
"max_average_score": 4.200033059767046e-17,
"sandwich_max_difference": 1.5306644840507033e-12,
"coder_covariance_rank": 2
} GEE Regression Results
===================================================================================
Dep. Variable: binary No. Observations: 5181
Model: GEE No. clusters: 65
Method: Generalized Min. cluster size: 78
Estimating Equations Max. cluster size: 84
Family: Binomial Mean cluster size: 79.7
Dependence structure: Independence Num. iterations: 2
Date: Mon, 07 Sep 2026 Scale: 1.000
Covariance type: robust Time: 05:45:12
========================================================================================================================
coef std err z P>|z| [0.025 0.975]
------------------------------------------------------------------------------------------------------------------------
intercept 4.1009 0.707 5.799 0.000 2.715 5.487
Codex 1.0282 0.109 9.449 0.000 0.815 1.241
Gemini 0.6470 0.112 5.782 0.000 0.428 0.866
field_cb_benefit_presence_strength_code 1.542e-14 nan nan nan nan nan
field_cb_cost_intensity_code -0.9349 0.990 -0.944 0.345 -2.875 1.005
field_cb_cost_presence_strength_code -0.9349 0.990 -0.944 0.345 -2.875 1.005
field_detection_verification_presence_strength_code -6.5471 0.903 -7.253 0.000 -8.316 -4.778
field_dppra_a_strength_code -3.1519 0.682 -4.624 0.000 -4.488 -1.816
field_dppra_alignment_score_code -3.7355 0.700 -5.335 0.000 -5.108 -2.363
field_dppra_d_strength_code 1.371e-14 0.716 1.92e-14 1.000 -1.403 1.403
field_dppra_p1_strength_code -2.1683 0.681 -3.183 0.001 -3.504 -0.833
field_dppra_p2_strength_code -3.6065 0.673 -5.357 0.000 -4.926 -2.287
field_dppra_r_strength_code -4.5217 0.692 -6.531 0.000 -5.879 -3.165
field_eth_accountability_code -1.9412 0.698 -2.780 0.005 -3.310 -0.572
field_eth_autonomy_code -2.0217 0.811 -2.493 0.013 -3.611 -0.432
field_eth_beneficence_eth_code -1.9412 0.690 -2.814 0.005 -3.293 -0.589
field_eth_beneficence_mech_code -2.0973 0.591 -3.546 0.000 -3.256 -0.938
field_eth_equity_code -3.2520 0.696 -4.674 0.000 -4.616 -1.888
field_eth_integrity_code -0.4116 0.717 -0.574 0.566 -1.817 0.994
field_eth_justice_code -2.8511 0.779 -3.662 0.000 -4.377 -1.325
field_eth_nonmaleficence_code -1.8547 0.773 -2.398 0.016 -3.371 -0.339
field_eth_privacy_code -3.1172 0.668 -4.664 0.000 -4.427 -1.807
field_eth_transparency_code -2.0973 0.698 -3.004 0.003 -3.466 -0.729
field_int_academic_intensity_code -7.5524 0.864 -8.738 0.000 -9.246 -5.858
field_int_assessment_intensity_code -6.1661 0.773 -7.982 0.000 -7.680 -4.652
field_int_data_prov_intensity_code -3.1859 0.640 -4.982 0.000 -4.439 -1.932
field_int_epistemic_intensity_code -2.3600 0.703 -3.358 0.001 -3.738 -0.982
field_int_research_intensity_code -5.9721 0.722 -8.271 0.000 -7.387 -4.557
field_int_system_tevv_intensity_code -2.8090 0.677 -4.147 0.000 -4.137 -1.481
field_journal_evidence_quality_strength_code -2.0162 0.899 -2.243 0.025 -3.778 -0.254
field_surveillance_monitoring_presence_strength_code -5.5162 1.071 -5.151 0.000 -7.615 -3.417
==============================================================================
Skew: -1.2471 Kurtosis: 2.7631
Centered skew: -0.7794 Centered kurtosis: 2.4712
==============================================================================
Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
above_2: full coefficients, exclusions and diagnostics
Download all coefficients · Download complete model summary
{
"model": "above_2",
"converged": true,
"iterations": 2,
"warnings": [],
"n_parameters": 31,
"covariance_rank": 31,
"smallest_cov_eigenvalue": 0.0011357310360301605,
"max_absolute_parameter": 5.111813895504899,
"threshold": 2,
"fields_excluded_for_no_binary_variation": [],
"max_average_score": 1.3457248783335231e-17,
"sandwich_max_difference": 7.931155732165962e-15,
"coder_covariance_rank": 2
} GEE Regression Results
===================================================================================
Dep. Variable: binary No. Observations: 5181
Model: GEE No. clusters: 65
Method: Generalized Min. cluster size: 78
Estimating Equations Max. cluster size: 84
Family: Binomial Mean cluster size: 79.7
Dependence structure: Independence Num. iterations: 2
Date: Mon, 07 Sep 2026 Scale: 1.000
Covariance type: robust Time: 05:45:13
========================================================================================================================
coef std err z P>|z| [0.025 0.975]
------------------------------------------------------------------------------------------------------------------------
intercept 0.5887 0.196 3.003 0.003 0.204 0.973
Codex 1.6148 0.111 14.486 0.000 1.396 1.833
Gemini 0.6736 0.095 7.091 0.000 0.487 0.860
field_cb_benefit_presence_strength_code 0.3415 0.109 3.126 0.002 0.127 0.556
field_cb_cost_intensity_code 0.5407 0.316 1.711 0.087 -0.079 1.160
field_cb_cost_presence_strength_code 0.7205 0.328 2.197 0.028 0.078 1.363
field_detection_verification_presence_strength_code -3.9306 0.600 -6.554 0.000 -5.106 -2.755
field_dppra_a_strength_code -1.9648 0.265 -7.405 0.000 -2.485 -1.445
field_dppra_alignment_score_code -2.1142 0.268 -7.890 0.000 -2.639 -1.589
field_dppra_d_strength_code 2.1268 0.405 5.246 0.000 1.332 2.921
field_dppra_p1_strength_code -1.3581 0.248 -5.485 0.000 -1.843 -0.873
field_dppra_p2_strength_code -1.7965 0.251 -7.151 0.000 -2.289 -1.304
field_dppra_r_strength_code -2.6404 0.299 -8.831 0.000 -3.226 -2.054
field_eth_accountability_code 0.5407 0.351 1.540 0.123 -0.147 1.229
field_eth_autonomy_code -0.3482 0.262 -1.327 0.185 -0.862 0.166
field_eth_beneficence_eth_code 0.0962 0.275 0.350 0.726 -0.442 0.635
field_eth_beneficence_mech_code -0.6823 0.133 -5.128 0.000 -0.943 -0.421
field_eth_equity_code -0.8034 0.262 -3.063 0.002 -1.317 -0.289
field_eth_integrity_code 0.4989 0.298 1.673 0.094 -0.086 1.083
field_eth_justice_code -0.4284 0.283 -1.515 0.130 -0.983 0.126
field_eth_nonmaleficence_code 0.1631 0.315 0.518 0.604 -0.454 0.780
field_eth_privacy_code -1.0611 0.237 -4.480 0.000 -1.525 -0.597
field_eth_transparency_code 0.4989 0.304 1.639 0.101 -0.098 1.095
field_int_academic_intensity_code -5.0052 0.611 -8.193 0.000 -6.203 -3.808
field_int_assessment_intensity_code -3.4790 0.333 -10.435 0.000 -4.132 -2.826
field_int_data_prov_intensity_code -1.4492 0.227 -6.382 0.000 -1.894 -1.004
field_int_epistemic_intensity_code -1.3581 0.219 -6.193 0.000 -1.788 -0.928
field_int_research_intensity_code -3.7526 0.346 -10.834 0.000 -4.432 -3.074
field_int_system_tevv_intensity_code -0.9217 0.237 -3.885 0.000 -1.387 -0.457
field_journal_evidence_quality_strength_code -0.8102 0.522 -1.553 0.120 -1.833 0.212
field_surveillance_monitoring_presence_strength_code -5.1118 0.994 -5.143 0.000 -7.060 -3.164
==============================================================================
Skew: -0.3074 Kurtosis: -0.3985
Centered skew: -0.2340 Centered kurtosis: -0.2894
==============================================================================
Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
above_3: full coefficients, exclusions and diagnostics
Download all coefficients · Download complete model summary
{
"model": "above_3",
"converged": true,
"iterations": 2,
"warnings": [],
"n_parameters": 29,
"covariance_rank": 29,
"smallest_cov_eigenvalue": 0.0011104048247526743,
"max_absolute_parameter": 4.0725443633151865,
"threshold": 3,
"fields_excluded_for_no_binary_variation": [
"detection_verification_presence_strength_code",
"surveillance_monitoring_presence_strength_code"
],
"max_average_score": 4.3896075743644865e-17,
"sandwich_max_difference": 5.592748486549226e-15,
"coder_covariance_rank": 2
} GEE Regression Results
===================================================================================
Dep. Variable: binary No. Observations: 5109
Model: GEE No. clusters: 65
Method: Generalized Min. cluster size: 75
Estimating Equations Max. cluster size: 81
Family: Binomial Mean cluster size: 78.6
Dependence structure: Independence Num. iterations: 2
Date: Mon, 07 Sep 2026 Scale: 1.000
Covariance type: robust Time: 05:45:13
================================================================================================================
coef std err z P>|z| [0.025 0.975]
----------------------------------------------------------------------------------------------------------------
intercept -1.6835 0.196 -8.570 0.000 -2.068 -1.298
Codex 1.8539 0.115 16.154 0.000 1.629 2.079
Gemini 0.9124 0.107 8.498 0.000 0.702 1.123
field_cb_benefit_presence_strength_code 0.4676 0.098 4.780 0.000 0.276 0.659
field_cb_cost_intensity_code 0.4199 0.290 1.450 0.147 -0.148 0.988
field_cb_cost_presence_strength_code 0.8676 0.305 2.849 0.004 0.271 1.464
field_dppra_a_strength_code -0.8276 0.347 -2.387 0.017 -1.507 -0.148
field_dppra_alignment_score_code -2.6233 0.454 -5.780 0.000 -3.513 -1.734
field_dppra_d_strength_code 2.9831 0.334 8.940 0.000 2.329 3.637
field_dppra_p1_strength_code 0.2508 0.268 0.936 0.349 -0.274 0.776
field_dppra_p2_strength_code -1.5143 0.348 -4.357 0.000 -2.196 -0.833
field_dppra_r_strength_code -2.7811 0.442 -6.295 0.000 -3.647 -1.915
field_eth_accountability_code 1.6281 0.308 5.291 0.000 1.025 2.231
field_eth_autonomy_code -0.7925 0.347 -2.283 0.022 -1.473 -0.112
field_eth_beneficence_eth_code 0.6800 0.267 2.551 0.011 0.158 1.202
field_eth_beneficence_mech_code -0.1046 0.158 -0.663 0.507 -0.414 0.205
field_eth_equity_code 0.0767 0.303 0.253 0.800 -0.517 0.670
field_eth_integrity_code 1.1760 0.248 4.744 0.000 0.690 1.662
field_eth_justice_code 0.1270 0.319 0.398 0.691 -0.498 0.752
field_eth_nonmaleficence_code 0.7503 0.297 2.526 0.012 0.168 1.332
field_eth_privacy_code -0.3815 0.294 -1.298 0.194 -0.958 0.195
field_eth_transparency_code 1.2731 0.287 4.439 0.000 0.711 1.835
field_int_academic_intensity_code -4.0725 1.051 -3.875 0.000 -6.133 -2.012
field_int_assessment_intensity_code -3.3640 0.568 -5.925 0.000 -4.477 -2.251
field_int_data_prov_intensity_code -0.8276 0.292 -2.831 0.005 -1.401 -0.255
field_int_epistemic_intensity_code -1.1355 0.309 -3.675 0.000 -1.741 -0.530
field_int_research_intensity_code -4.0725 0.772 -5.272 0.000 -5.587 -2.559
field_int_system_tevv_intensity_code 0.0257 0.294 0.087 0.930 -0.551 0.602
field_journal_evidence_quality_strength_code -1.2283 0.712 -1.724 0.085 -2.625 0.168
==============================================================================
Skew: 0.4555 Kurtosis: -0.3312
Centered skew: 0.4184 Centered kurtosis: -0.2512
==============================================================================
Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
above_4: full coefficients, exclusions and diagnostics
Download all coefficients · Download complete model summary
{
"model": "above_4",
"converged": true,
"iterations": 2,
"warnings": [],
"n_parameters": 21,
"covariance_rank": 21,
"smallest_cov_eigenvalue": 0.004076421807522495,
"max_absolute_parameter": 4.643468389491005,
"threshold": 4,
"fields_excluded_for_no_binary_variation": [
"detection_verification_presence_strength_code",
"dppra_alignment_score_code",
"dppra_p2_strength_code",
"dppra_r_strength_code",
"int_academic_intensity_code",
"int_assessment_intensity_code",
"int_data_prov_intensity_code",
"int_epistemic_intensity_code",
"int_research_intensity_code",
"surveillance_monitoring_presence_strength_code"
],
"max_average_score": 1.1127235269328709e-17,
"sandwich_max_difference": 1.0380585280245214e-13,
"coder_covariance_rank": 2
} GEE Regression Results
===================================================================================
Dep. Variable: binary No. Observations: 3552
Model: GEE No. clusters: 65
Method: Generalized Min. cluster size: 54
Estimating Equations Max. cluster size: 57
Family: Binomial Mean cluster size: 54.6
Dependence structure: Independence Num. iterations: 2
Date: Mon, 07 Sep 2026 Scale: 1.000
Covariance type: robust Time: 05:45:13
================================================================================================================
coef std err z P>|z| [0.025 0.975]
----------------------------------------------------------------------------------------------------------------
intercept -4.6435 0.738 -6.288 0.000 -6.091 -3.196
Codex 0.9454 0.198 4.783 0.000 0.558 1.333
Gemini 0.2700 0.214 1.260 0.207 -0.150 0.690
field_cb_benefit_presence_strength_code 0.5231 0.417 1.255 0.210 -0.294 1.340
field_cb_cost_intensity_code 8.103e-15 0.960 8.44e-15 1.000 -1.882 1.882
field_cb_cost_presence_strength_code 0.2938 0.916 0.321 0.748 -1.501 2.089
field_dppra_a_strength_code -1.1108 1.260 -0.881 0.378 -3.581 1.359
field_dppra_d_strength_code 3.5916 0.784 4.584 0.000 2.056 5.127
field_dppra_p1_strength_code 1.2474 0.857 1.455 0.146 -0.432 2.927
field_eth_accountability_code 2.0051 0.816 2.456 0.014 0.405 3.605
field_eth_autonomy_code 0.2938 1.110 0.265 0.791 -1.882 2.469
field_eth_beneficence_eth_code 1.5288 0.874 1.749 0.080 -0.184 3.242
field_eth_beneficence_mech_code 0.7116 0.337 2.114 0.035 0.052 1.371
field_eth_equity_code 8.278e-15 0.960 8.62e-15 1.000 -1.882 1.882
field_eth_integrity_code 1.2474 0.857 1.455 0.146 -0.433 2.927
field_eth_justice_code 1.0117 0.865 1.169 0.242 -0.684 2.708
field_eth_nonmaleficence_code 1.4424 0.867 1.663 0.096 -0.258 3.142
field_eth_privacy_code -1.1108 1.260 -0.881 0.378 -3.581 1.359
field_eth_transparency_code 1.1357 0.930 1.221 0.222 -0.687 2.959
field_int_system_tevv_intensity_code -1.1108 1.260 -0.881 0.378 -3.581 1.360
field_journal_evidence_quality_strength_code 2.3894 1.025 2.331 0.020 0.381 4.398
==============================================================================
Skew: 3.3864 Kurtosis: 12.8510
Centered skew: 3.2193 Centered kurtosis: 12.0258
==============================================================================
Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
Coder configuration and provenance
Coder configuration for the three version-locked runs behind the paper. RUN 07 (v2.14) is the source of every figure in this addendum. Reviewers asked for exact model versions, access dates, interfaces, sampling parameters and session-memory conditions. They are answerable to different degrees, and none of the three resolves to a dated build. The Codex coder’s model label and client version are recorded in the session logs for every run. The Claude coder’s tier rests on investigator testimony, corroborated by dated build records spanning the coding window. The Gemini tier is reported by the app’s own response-details panel for all three runs. Across all three runs, every coder is now identified by something better than testimony. Every label below is nonetheless a product-surface label, so snapshot identity is not claimed for any coder.
| Class | Items |
|---|---|
| Artifact-backed | instrument version; packet texts (SHA-256); coding dates; coder identity stamped in every output; output schema; Codex model label and client version, from the session logs; Gemini tier for all three runs, from the app's response-details panel |
| Testimony-backed | Claude tier, corroborated by dated build records; interface for all three coders |
| Unrecoverable | dated model snapshot id for any coder; API version; temperature; top-p; model-side system prompt |
All three runs fall inside a seven-day window, 24-30 June 2026. RUN 07 was coded in a single day, 30 June 2026. This bounds exposure to an unannounced point release; it is not proof of model constancy.
| Run | Coder | Model as run | Interface | Outputs | Coding dates | Packets hashed | Combined SHA-256 of packet set |
|---|---|---|---|---|---|---|---|
| RUN 04 · v2.12 | Claude | Anthropic Claude — Opus 4.8, highest tier then available (investigator testimony, corroborated by dated build records across the window) | chat, packet-delivered via a local dashboard | 65 | 2026-06-24 | 65 | f6d4a5780d471ba7c165b3e1c3cb302298bd8b7e7ac89d0f0682073617bd33f8 |
| RUN 04 · v2.12 | Codex | OpenAI Codex — model label gpt-5.5, Codex Desktop CLI 0.142.0 (recorded, not testimony) | ChatGPT-subscription Codex CLI, not the API | 65 | 2026-06-24 | 65 | 50f1d6964f694f73085867cb4802d1f053df7ecfd947ba2db4447b9d6a4e4207 |
| RUN 04 · v2.12 | Gemini | Google Gemini — 3.5 Flash, reported by the app's response-details panel on the coding thread for this run | Gemini consumer app, hand-run by the investigator | 65 | 2026-06-24, 2026-06-25, 2026-06-26, 2026-06-27 | not recoverable | — |
| RUN 06 · v2.13 | Claude | Anthropic Claude — Opus 4.8, highest tier then available (investigator testimony, corroborated by dated build records across the window) | chat, packet-delivered via a local dashboard | 65 | 2026-06-28, 2026-06-29 | 55 | 490726fa38cc45c1140658b1b69a90f34b1437b8d96ffec4f2ef4d601f64f2cc |
| RUN 06 · v2.13 | Codex | OpenAI Codex — model label gpt-5.5, Codex Desktop CLI 0.142.3 (recorded, not testimony) | ChatGPT-subscription Codex CLI, not the API | 65 | 2026-06-28, 2026-06-29 | 55 | 410c379331c43224b19cc7b10a90ee332272f0d66d46696f3013862d165ab9ed |
| RUN 06 · v2.13 | Gemini | Google Gemini — 3.5 Flash, reported by the app's response-details panel on the coding thread for this run | Gemini consumer app, hand-run by the investigator | 65 | 2026-06-28, 2026-06-29 | not recoverable | — |
| RUN 07 · v2.14 | Claude | Anthropic Claude — Opus 4.8, highest tier then available (investigator testimony, corroborated by dated build records across the window) | chat, packet-delivered via a local dashboard | 65 | 2026-06-30 | 65 | 7d0f7d86d66a3786d2d88659086a747e33c626fd35ac8155bfc3104f824eec1b |
| RUN 07 · v2.14 | Codex | OpenAI Codex — model label gpt-5.5, Codex Desktop CLI 0.142.3 (recorded, not testimony) | ChatGPT-subscription Codex CLI, not the API | 65 | 2026-06-30 | 65 | 7d0f7d86d66a3786d2d88659086a747e33c626fd35ac8155bfc3104f824eec1b |
| RUN 07 · v2.14 | Gemini | Google Gemini — 3.5 Flash, reported by the app's response-details panel on the coding thread for this run | Gemini consumer app, hand-run by the investigator | 65 | 2026-06-30 | 65 | 22a27b1f363fe17267ae093686d0225fa79ccd6abf03de7a5a5eae93aff9a332 |
How each model was identified
Codex. The Codex coder ran through a desktop client that writes a session log for every run. Those logs record the model label and client version. 188 sessions on 30 June 2026; 187 name documents, together covering all 65; every one records model label gpt-5.5 and client version 0.142.3. Session of 28 June 2026: model label gpt-5.5, client version 0.142.3. Session of 24 June 2026: model label gpt-5.5, client version 0.142.0.
gpt-5.5 is the label the client recorded, not a dated build identifier. It fixes which model line was addressed, not which build answered. The session logs themselves are not published: they contain full document text.
Claude. The Claude coder's outputs carry no model field, so the tier rests on investigator testimony that the highest available build was used throughout. Dated build records from the same workstation name Claude Opus 4.8 on 25, 27, 28 and 29 June and on 1 and 3 July 2026 — spanning the whole coding window — with the next Anthropic build first appearing on 5 July, after all three runs.
This establishes which Anthropic build was in use on those dates. It does not stamp the chat session that performed the coding, which is a different surface. Reported as corroborated testimony, not as identification.
Gemini. The Gemini app reports, per response, which model answered. On the retained coding threads that panel reads 3.5 Flash, with the composer's model selector also on Flash. Checked on a retained thread for each of the three runs. All three read 3.5 Flash.
Each thread is tied to its run by content, not by assertion. The RUN 07 thread's record header names the same source id, document id and title as the archived RUN 07 Gemini output and stamps prompt version v2.14 and coding date 30 June 2026; its narrative content matches that output string for string, and its two emergent-code strings appear in no other output in the archive. The RUN 06 thread declares schema version v2.13 and codes a named corporate scaling policy from that run.
One RUN 04 output self-stamped a coder id containing 'pro', which conflicted with the Flash testimony. That string was written by the model, and model self-identification is unreliable. The app's own panel reports Flash. The self-stamp is reported as an adherence deviation, not as an attribution. Like the Codex label, this is a product-surface label rather than a dated build identifier.
Sampling parameters, session memory and output settings were not settable and were not recorded. Claude and Codex ran through subscription interfaces exposing no temperature or top-p control; Gemini was hand-run in the consumer app. Session-memory state was not captured. This is a limitation of consumer-interface coding, reported rather than estimated.
Gemini packets are staged in place and overwritten between runs, so only the most recent staging survives. All 65 carry a 30 June 2026 modification time and are artifact-backed for RUN 07 only; the RUN 04 and RUN 06 Gemini packets are gone. Packet coverage for RUN 07: 65 unique sources (COR 10, FRM 12, GOV 15, INS 14, JRN 14).
The machine-readable source for this section is provenance_configuration.json, included in the reproducibility package.
What these checks establish
The overall ordering Codex > Gemini > Claude survives document-level comparisons, a check using order alone, and clustered ordinal models. The evidence supports a directional tendency among these tested coder configurations.
Documents are assumed independent between clusters. The selected corpus is not a probability sample, the instrument fields are fixed, and unmodeled between-document dependence remains possible. Missingness follows production conventions rather than a missing-data model. The checks do not establish validity against an external criterion, causal effects of protocol revisions, inherited model-family traits, or the validity of the separate nominal-field χ² tests. Account, interface and exact-build provenance remain separate questions.
This dated release includes every test in the reviewer-requested ordinal sensitivity battery. Earlier protocol experiments and qualitative counter-readings are separate bodies of evidence and are not presented as newly reproduced here.
Reproduce and extend the results
Download the complete reproducibility package
The package includes the exact numerical inputs, pinned dependencies, all outputs, original post-review analysis plan, and portable scripts. No raw corpus text, private working notes or local computer paths are included. Run checks.py, ordinal_checks.py, then verify.py; see the included README for setup and interpretation. The 195 original source-file hashes were checked unchanged during the local audit; reproduction from the public numerical export does not independently verify those unavailable source files.
Browse every downloadable file
- analysis_plan.md
- checks.py
- common_valid_ratings.csv
- data_audit.json
- document_level_tests.csv
- document_means_all_ordinal.csv
- document_means_ethical_principles.csv
- field_inventory.csv
- model_above_0.txt
- model_above_1.txt
- model_above_2.txt
- model_above_3.txt
- model_above_4.txt
- model_all_ordinal.txt
- model_ethical_principles.txt
- model_without_COR.txt
- model_without_FRM.txt
- model_without_GOV.txt
- model_without_INS.txt
- model_without_JRN.txt
- ordinal_checks.py
- ordinal_contrasts.csv
- ordinal_ratings.csv
- ordinal_results.json
- parameters_above_0.csv
- parameters_above_1.csv
- parameters_above_2.csv
- parameters_above_3.csv
- parameters_above_4.csv
- parameters_all_ordinal.csv
- parameters_ethical_principles.csv
- parameters_without_COR.csv
- parameters_without_FRM.csv
- parameters_without_GOV.csv
- parameters_without_INS.csv
- parameters_without_JRN.csv
- provenance_configuration.json
- rank_only_checks.csv
- README.md
- reproduced_verification.json
- reproduction.json
- requirements.txt
- sensitivity_results.json
- source_manifest.json
- verification.json
- verify.py
Release history
— First statistical addendum: original-results reproduction, six document tests, six order-only checks, 12 GEE fits with 36 contrasts, sphericity sensitivity and numerical verification. Future analyses will be added as separately dated releases with their own data, methods and scope.
Method reference: statsmodels OrdinalGEE documentation.