← Ethics Observatory

SAMYRAD 2026 · DPPRA / ICST · RUN 07 · 7 September 2026

Statistical addendum

The overall coder ordering holds when we account for document dependence and use category order alone. Its strength varies across the scale.

Complete reviewer-requested checks for Coders as Standpoint Instruments, including inconclusive results, model diagnostics and the files needed to reproduce every calculation in this release.

65document clusters
29ordinal fields
12GEE fits, all converged

The exact analysis set

Frozen RUN_07, protocol v2.14: 195 coding outputs, 65 documents and three coder configurations. Sector counts are government 15, journal 14, forum 12, institutional 14 and corporate 10. The instrument contains 91 coded fields and 16 narrative fields; this battery selects the 29 ordinal fields identified by the production classifier.

Codes 8 (unclear) and 9 (not applicable) are missing; process/confidence and deprecated fields are excluded. Three-coder matching yields 1,727 document–field units and 5,181 ratings. Document means use the same eligible fields for all coders and give each document equal weight. Three fields are domain-specific. Blank slots in the full numerical export denote missing or inapplicable values.

All 29 field denominators
Ordinal fieldComplete documents
cb_benefit_intensity_code65
cb_benefit_presence_strength_code65
cb_cost_intensity_code65
cb_cost_presence_strength_code65
detection_verification_presence_strength_code12
dppra_a_strength_code65
dppra_alignment_score_code64
dppra_d_strength_code65
dppra_p1_strength_code65
dppra_p2_strength_code65
dppra_r_strength_code65
eth_accountability_code65
eth_autonomy_code65
eth_beneficence_eth_code65
eth_beneficence_mech_code65
eth_equity_code65
eth_integrity_code65
eth_justice_code65
eth_nonmaleficence_code65
eth_privacy_code65
eth_transparency_code65
int_academic_intensity_code65
int_assessment_intensity_code65
int_data_prov_intensity_code65
int_epistemic_intensity_code65
int_research_intensity_code65
int_system_tevv_intensity_code65
journal_evidence_quality_strength_code14
surveillance_monitoring_presence_strength_code12

Original results reproduced

The repeated-measures ANOVA uses document averages. Its partial η² is coder-effect sum of squares divided by coder-effect-plus-error sums of squares. It is not 76% of all raw rating variance, and it is not a causal attribution. Arithmetic summaries retain an approximately equal-spacing interpretation. Ordinal Krippendorff’s α reproduces as 0.7247.

Repeated-measures ANOVA
DocumentsFdf 1df 2pPartial η²SS coderSS errorSS total
65197.553321287.442e-400.755316.63855.390373.0152
Coder means
CoderDocument-average score
Claude2.2302
Codex2.9457
Gemini2.5848

Original pooled paired tests — historical reproduction only

These tests treat document–field pairs as independent despite within-document dependence. They are retained here for auditability; the revised paper uses the document-level tests below. Pair-available matching includes one additional Codex–Gemini unit; replacement tests use common-valid matching across all three coders.

All three original pooled paired tests
ComparisonPaired field unitsMean differencetUnadjusted p
Codex - Claude17270.71435.42584.612e-207
Gemini - Claude17270.350917.50792.651e-63
Codex - Gemini17280.362817.78973.846e-65
Greenhouse–Geisser sphericity sensitivity
GG εAdjusted df 1Adjusted df 2Adjusted p
0.92091.8419117.87866.497e-37

The sphericity correction does not establish equal category spacing.

Document-level paired comparisons

Six tests: all ordinal fields and the ten ethical-principle fields, with three comparisons each. Paired t-tests use 65 document means and 64 degrees of freedom. Holm adjustment applies separately to the three comparisons within each analysis. Confidence intervals are pointwise 95% intervals, not simultaneous intervals. Mean differences still use arithmetic category distances.

All six document-level tests
Field setComparisonDocumentsDifferenceCI lowCI hightdfRaw pHolm pHigher docsTied docsLower docs
all_ordinalCodex - Claude650.71550.65350.777523.0534641.056e-323.167e-326500
all_ordinalGemini - Claude650.35460.27410.4358.8073641.228e-121.228e-1254110
all_ordinalCodex - Gemini650.36090.28880.4339.9982641.062e-142.125e-145618
ethical_principlesCodex - Claude650.76770.68880.846619.448641.47e-284.409e-286500
ethical_principlesGemini - Claude650.360.26240.45767.3662644.224e-104.224e-105078
ethical_principlesCodex - Gemini650.40770.32520.49019.8781641.706e-143.413e-1452103

Checks using category order only

Within each matched field, assign +1 when the first coder is higher, 0 when tied, and −1 when lower. Average within documents: this is the fraction higher minus the fraction lower. Any strictly increasing relabeling of the categories leaves the result unchanged. Percentile intervals resample whole documents 30,000 times (seed 20260907). Exact two-sided binomial sign tests compare positive and negative document totals, omit document ties, and use Holm adjustment within each three-comparison field set.

Codex–Claude has net higher ratings in all 65 documents, but not on every field: 980 higher, 692 tied and 55 lower matched field ratings.

All six order-only checks
Field setComparisonMean net higher fractionBootstrap CI lowBootstrap CI highHigher docsTied docsLower docsRaw sign pHolm sign pHigher fieldsTied fieldsLower fields
all_ordinalCodex - Claude0.53660.49770.576365005.421e-201.626e-1998069255
all_ordinalGemini - Claude0.28080.22150.339854474.322e-104.322e-10653901173
all_ordinalCodex - Gemini0.29680.24340.349657177.638e-111.528e-10668906153
ethical_principlesCodex - Claude0.60920.55690.661565005.421e-201.626e-1941122415
ethical_principlesGemini - Claude0.29080.21230.3692504114.589e-074.589e-0725632767
ethical_principlesCodex - Gemini0.36310.29380.432353845.911e-121.182e-1127334037

Clustered ordinal models

Cumulative-logit ordinal GEE with coder indicators, fixed field effects, document clusters, working independence and robust sandwich covariance. This is a marginal model, not a mixed-effects model. It requires ordered categories without equal distances. The robust covariance aggregates whole-document contributions; working independence does not make individual ratings independent for uncertainty.

Reported contrasts multiply covariance by G/(G−1), use t(G−1) reference, and apply Holm adjustment to the three comparisons within each fit. The intervals are pointwise 95%. Odds ratios summarize the odds of exceeding a score threshold, not ethical intensity or correctness. A common odds ratio assumes proportional odds; the threshold diagnostics below do not establish that assumption.

All 21 ordinal contrasts: pooled, ethical-principle and five leave-sector-out models
ModelComparisonDocumentsFieldsRatingsOdds ratioCI lowCI highLog oddsSERaw pHolm p
all_ordinalCodex - Claude652951813.97243.28844.79871.37940.09465.119e-221.536e-21
all_ordinalGemini - Claude652951811.93381.64292.27610.65950.08162.302e-112.302e-11
all_ordinalCodex - Gemini652951812.05431.7432.4210.71990.08221.519e-123.039e-12
ethical_principlesCodex - Claude651019504.16723.32835.21751.42720.11253.991e-191.197e-18
ethical_principlesGemini - Claude651019501.86111.55332.230.62120.09053.228e-093.228e-09
ethical_principlesCodex - Gemini651019502.23911.84752.71360.80610.09626.987e-121.397e-11
without_CORCodex - Claude552944013.79333.09654.64691.33320.10121.694e-185.083e-18
without_CORGemini - Claude552944011.75811.48382.08320.56420.08461.425e-081.425e-08
without_CORCodex - Gemini552944012.15761.78692.60520.7690.0945.072e-111.014e-10
without_FRMCodex - Claude532741765.21974.33326.28761.65240.09288.646e-242.594e-23
without_FRMGemini - Claude532741762.19741.78152.71040.78730.10467.096e-107.096e-10
without_FRMCodex - Gemini532741762.37541.98122.84810.86520.09044.709e-139.419e-13
without_GOVCodex - Claude502940113.89283.18264.76141.35910.10023.301e-189.904e-18
without_GOVGemini - Claude502940111.82881.52532.19260.60360.09032.032e-082.032e-08
without_GOVCodex - Gemini502940112.12861.75462.58240.75550.09623.165e-106.33e-10
without_INSCodex - Claude512940893.6632.98034.5021.29830.10273.416e-171.025e-16
without_INSGemini - Claude512940891.81341.51462.17120.59520.08972.203e-082.203e-08
without_INSCodex - Gemini512940892.021.67952.42950.70310.09195.813e-101.163e-09
without_JRNCodex - Claude512840473.95283.15464.95281.37440.11231.171e-163.514e-16
without_JRNGemini - Claude512840472.25821.88522.70510.81460.08993.976e-127.951e-12
without_JRNCodex - Gemini512840471.75041.49432.05040.55980.07884.077e-094.077e-09

All contrasts retain positive directions and intervals above 1 in the ethical-principle subset and when each sector is omitted. Sector omission is a sensitivity check, not proof of generalization to new sectors.

Every threshold diagnostic

Separate binomial GEE fits compare scores above 0, 1, 2, 3 and 4. Fields with no binary outcome variation are omitted at each cutoff; differences between thresholds can therefore reflect both effect variation and field composition. All estimated directions are positive, but effect magnitudes vary.

Two comparisons remain unresolved. Codex–Gemini above 0 has Holm p = .514; Gemini–Claude above 4 has Holm p = .216. These results neither prove equality nor support uniform separation at every boundary. The common ordinal odds ratios are pooled summaries.
All 15 threshold-specific contrasts
ModelComparisonDocumentsFieldsRatingsOdds ratioCI lowCI highLog oddsSERaw pHolm p
above_0Codex - Claude652442061.73261.42532.10620.54960.09774.414e-071.324e-06
above_0Gemini - Claude652442061.58231.15412.16930.45890.1580.0050.0101
above_0Codex - Gemini652442061.0950.83091.44290.09070.13810.51360.5136
above_1Codex - Claude652951812.7962.24593.48091.02820.10971.255e-133.764e-13
above_1Gemini - Claude652951811.90981.52462.39230.6470.11282.831e-075.662e-07
above_1Codex - Gemini652951811.4641.14511.87180.38120.1230.00290.0029
above_2Codex - Claude652951815.02694.01646.29171.61480.11231.037e-213.11e-21
above_2Gemini - Claude652951811.96131.61992.37470.67360.09571.61e-091.61e-09
above_2Codex - Gemini652951812.5632.07753.1620.94120.10516.846e-131.369e-12
above_3Codex - Claude652751096.38445.06738.04371.85390.11574.323e-241.297e-23
above_3Gemini - Claude652751092.49042.00623.09140.91240.10825.592e-125.592e-12
above_3Codex - Gemini652751092.56362.07223.17160.94140.10651.087e-122.174e-12
above_4Codex - Claude651935522.57371.72883.83160.94540.19921.207e-053.621e-05
above_4Gemini - Claude651935521.30990.85112.0160.270.21580.21560.2156
above_4Codex - Gemini651935521.96481.37062.81660.67540.18030.00038680.0007736

Convergence and numerical verification

All 12 GEE fits converged. Independent estimating-equation and sandwich-covariance calculations verified the implementations. Low-threshold full nuisance covariance matrices can be rank-deficient, while the coder-coefficient covariance blocks are full-rank. Numerical square-root warnings or NaN nuisance standard errors can appear in the printed summaries; no full-parameter Wald tests are interpreted.

Printed ordinal model summaries contain 25,905 internally expanded threshold rows for the full fit. The research data comprise 5,181 ratings clustered in 65 documents. Raw statsmodels summary p-values differ from the finite-cluster and Holm-adjusted contrasts above.

Diagnostics for all 12 fits
ModelConvergedIterationsParametersFull covariance rankCoder covariance rankMax average scoreSandwich discrepancy
all_ordinal12353528.64e-184.295e-15
ethical_principles12161621.075e-171.596e-15
without_COR12353527.547e-185.759e-15
without_FRM12333326.093e-184.871e-15
without_GOV12353521.072e-177.619e-15
without_INS12353521.058e-175.586e-15
without_JRN12343425.712e-186.483e-15
above_012262321.93e-178.932e-14
above_112313024.2e-171.531e-12
above_212313121.346e-177.931e-15
above_312292924.39e-175.593e-15
above_412212121.113e-171.038e-13
all_ordinal: full coefficients, exclusions and diagnostics

Download all coefficients · Download complete model summary

{
  "model": "all_ordinal",
  "converged": true,
  "iterations": 2,
  "warnings": [],
  "n_parameters": 35,
  "covariance_rank": 35,
  "smallest_cov_eigenvalue": 0.00015774969186709627,
  "max_absolute_parameter": 5.38613788718312,
  "threshold": null,
  "fields_excluded_for_no_binary_variation": [],
  "max_average_score": 8.640068008663638e-18,
  "sandwich_max_difference": 4.295175326518574e-15,
  "coder_covariance_rank": 2
}
                           OrdinalGEE Regression Results                           
===================================================================================
Dep. Variable:                      rating   No. Observations:                25905
Model:                          OrdinalGEE   No. clusters:                       65
Method:                        Generalized   Min. cluster size:                 390
                      Estimating Equations   Max. cluster size:                 420
Family:                           Binomial   Mean cluster size:               398.5
Dependence structure:         Independence   Num. iterations:                     2
Date:                     Mon, 07 Sep 2026   Scale:                           1.000
Covariance type:                    robust   Time:                         05:45:07
========================================================================================================================
                                                           coef    std err          z      P>|z|      [0.025      0.975]
------------------------------------------------------------------------------------------------------------------------
I(y>0.0)                                                 2.8860      0.197     14.648      0.000       2.500       3.272
I(y>1.0)                                                 2.1584      0.184     11.704      0.000       1.797       2.520
I(y>2.0)                                                 0.5459      0.136      4.027      0.000       0.280       0.812
I(y>3.0)                                                -1.0258      0.140     -7.322      0.000      -1.300      -0.751
I(y>4.0)                                                -4.1852      0.205    -20.444      0.000      -4.586      -3.784
Codex                                                    1.3794      0.094     14.695      0.000       1.195       1.563
Gemini                                                   0.6595      0.081      8.146      0.000       0.501       0.818
field_cb_benefit_presence_strength_code                  0.3152      0.059      5.374      0.000       0.200       0.430
field_cb_cost_intensity_code                             0.2942      0.216      1.360      0.174      -0.130       0.718
field_cb_cost_presence_strength_code                     0.5520      0.229      2.411      0.016       0.103       1.001
field_detection_verification_presence_strength_code     -4.9739      0.556     -8.942      0.000      -6.064      -3.884
field_dppra_a_strength_code                             -1.5012      0.220     -6.822      0.000      -1.932      -1.070
field_dppra_alignment_score_code                        -1.8558      0.196     -9.468      0.000      -2.240      -1.472
field_dppra_d_strength_code                              2.6520      0.296      8.949      0.000       2.071       3.233
field_dppra_p1_strength_code                            -0.5459      0.214     -2.552      0.011      -0.965      -0.127
field_dppra_p2_strength_code                            -1.7762      0.211     -8.427      0.000      -2.189      -1.363
field_dppra_r_strength_code                             -2.7500      0.246    -11.166      0.000      -3.233      -2.267
field_eth_accountability_code                            0.8967      0.311      2.882      0.004       0.287       1.507
field_eth_autonomy_code                                 -0.5006      0.216     -2.320      0.020      -0.923      -0.078
field_eth_beneficence_eth_code                           0.3047      0.234      1.300      0.194      -0.155       0.764
field_eth_beneficence_mech_code                         -0.3815      0.103     -3.709      0.000      -0.583      -0.180
field_eth_equity_code                                   -0.7681      0.258     -2.972      0.003      -1.275      -0.261
field_eth_integrity_code                                 0.7210      0.212      3.399      0.001       0.305       1.137
field_eth_justice_code                                  -0.4000      0.269     -1.486      0.137      -0.927       0.128
field_eth_nonmaleficence_code                            0.3362      0.265      1.266      0.205      -0.184       0.857
field_eth_privacy_code                                  -1.0014      0.209     -4.782      0.000      -1.412      -0.591
field_eth_transparency_code                              0.5520      0.262      2.111      0.035       0.039       1.065
field_int_academic_intensity_code                       -5.3861      0.419    -12.865      0.000      -6.207      -4.566
field_int_assessment_intensity_code                     -4.3880      0.342    -12.846      0.000      -5.057      -3.718
field_int_data_prov_intensity_code                      -1.3783      0.198     -6.962      0.000      -1.766      -0.990
field_int_epistemic_intensity_code                      -1.1117      0.175     -6.366      0.000      -1.454      -0.769
field_int_research_intensity_code                       -4.1478      0.288    -14.422      0.000      -4.711      -3.584
field_int_system_tevv_intensity_code                    -0.7768      0.229     -3.392      0.001      -1.226      -0.328
field_journal_evidence_quality_strength_code            -0.6483      0.470     -1.379      0.168      -1.570       0.273
field_surveillance_monitoring_presence_strength_code    -4.0868      0.578     -7.076      0.000      -5.219      -2.955
==============================================================================
Skew:                         -0.2025   Kurtosis:                       1.6538
Centered skew:                -0.0425   Centered kurtosis:              1.3716
==============================================================================

Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
ethical_principles: full coefficients, exclusions and diagnostics

Download all coefficients · Download complete model summary

{
  "model": "ethical_principles",
  "converged": true,
  "iterations": 2,
  "warnings": [],
  "n_parameters": 16,
  "covariance_rank": 16,
  "smallest_cov_eigenvalue": 0.0007437373155553275,
  "max_absolute_parameter": 3.8656625982654456,
  "threshold": null,
  "fields_excluded_for_no_binary_variation": [],
  "max_average_score": 1.0749236258934849e-17,
  "sandwich_max_difference": 1.5959455978986625e-15,
  "coder_covariance_rank": 2
}
                           OrdinalGEE Regression Results                           
===================================================================================
Dep. Variable:                      rating   No. Observations:                 9750
Model:                          OrdinalGEE   No. clusters:                       65
Method:                        Generalized   Min. cluster size:                 150
                      Estimating Equations   Max. cluster size:                 150
Family:                           Binomial   Mean cluster size:               150.0
Dependence structure:         Independence   Num. iterations:                     2
Date:                     Mon, 07 Sep 2026   Scale:                           1.000
Covariance type:                    robust   Time:                         05:45:08
===================================================================================================
                                      coef    std err          z      P>|z|      [0.025      0.975]
---------------------------------------------------------------------------------------------------
I(y>0.0)                            3.8657      0.454      8.517      0.000       2.976       4.755
I(y>1.0)                            2.6961      0.370      7.287      0.000       1.971       3.421
I(y>2.0)                            1.3913      0.296      4.700      0.000       0.811       1.971
I(y>3.0)                           -0.0941      0.247     -0.381      0.703      -0.578       0.390
I(y>4.0)                           -3.1850      0.221    -14.403      0.000      -3.618      -2.752
Codex                               1.4272      0.112     12.784      0.000       1.208       1.646
Gemini                              0.6212      0.090      6.917      0.000       0.445       0.797
field_eth_autonomy_code            -1.3421      0.294     -4.564      0.000      -1.918      -0.766
field_eth_beneficence_eth_code     -0.5689      0.318     -1.788      0.074      -1.193       0.055
field_eth_beneficence_mech_code    -1.2278      0.334     -3.678      0.000      -1.882      -0.573
field_eth_equity_code              -1.5987      0.266     -6.014      0.000      -2.120      -1.078
field_eth_integrity_code           -0.1688      0.226     -0.748      0.455      -0.612       0.274
field_eth_justice_code             -1.2455      0.253     -4.926      0.000      -1.741      -0.750
field_eth_nonmaleficence_code      -0.5386      0.221     -2.434      0.015      -0.972      -0.105
field_eth_privacy_code             -1.8227      0.275     -6.620      0.000      -2.362      -1.283
field_eth_transparency_code        -0.3312      0.192     -1.721      0.085      -0.708       0.046
==============================================================================
Skew:                         -0.5337   Kurtosis:                       1.5030
Centered skew:                -0.2363   Centered kurtosis:              1.1091
==============================================================================

Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
without_COR: full coefficients, exclusions and diagnostics

Download all coefficients · Download complete model summary

{
  "model": "without_COR",
  "converged": true,
  "iterations": 2,
  "warnings": [],
  "n_parameters": 35,
  "covariance_rank": 35,
  "smallest_cov_eigenvalue": 0.0001067196496657633,
  "max_absolute_parameter": 5.17718555018618,
  "threshold": null,
  "fields_excluded_for_no_binary_variation": [],
  "max_average_score": 7.546796035428525e-18,
  "sandwich_max_difference": 5.7592819402429996e-15,
  "coder_covariance_rank": 2
}
                           OrdinalGEE Regression Results                           
===================================================================================
Dep. Variable:                      rating   No. Observations:                22005
Model:                          OrdinalGEE   No. clusters:                       55
Method:                        Generalized   Min. cluster size:                 390
                      Estimating Equations   Max. cluster size:                 420
Family:                           Binomial   Mean cluster size:               400.1
Dependence structure:         Independence   Num. iterations:                     2
Date:                     Mon, 07 Sep 2026   Scale:                           1.000
Covariance type:                    robust   Time:                         05:45:09
========================================================================================================================
                                                           coef    std err          z      P>|z|      [0.025      0.975]
------------------------------------------------------------------------------------------------------------------------
I(y>0.0)                                                 2.9315      0.227     12.896      0.000       2.486       3.377
I(y>1.0)                                                 2.2039      0.214     10.314      0.000       1.785       2.623
I(y>2.0)                                                 0.5453      0.154      3.530      0.000       0.243       0.848
I(y>3.0)                                                -0.9913      0.161     -6.144      0.000      -1.308      -0.675
I(y>4.0)                                                -4.0852      0.231    -17.667      0.000      -4.538      -3.632
Codex                                                    1.3332      0.100     13.291      0.000       1.137       1.530
Gemini                                                   0.5642      0.084      6.729      0.000       0.400       0.729
field_cb_benefit_presence_strength_code                  0.3085      0.065      4.772      0.000       0.182       0.435
field_cb_cost_intensity_code                             0.3949      0.230      1.714      0.087      -0.057       0.847
field_cb_cost_presence_strength_code                     0.6902      0.244      2.828      0.005       0.212       1.169
field_detection_verification_presence_strength_code     -4.9628      0.563     -8.816      0.000      -6.066      -3.860
field_dppra_a_strength_code                             -1.6130      0.240     -6.730      0.000      -2.083      -1.143
field_dppra_alignment_score_code                        -2.0246      0.220     -9.182      0.000      -2.457      -1.592
field_dppra_d_strength_code                              2.6698      0.343      7.790      0.000       1.998       3.342
field_dppra_p1_strength_code                            -0.6984      0.221     -3.164      0.002      -1.131      -0.266
field_dppra_p2_strength_code                            -2.0218      0.233     -8.670      0.000      -2.479      -1.565
field_dppra_r_strength_code                             -3.1145      0.269    -11.558      0.000      -3.643      -2.586
field_eth_accountability_code                            0.9216      0.353      2.611      0.009       0.230       1.614
field_eth_autonomy_code                                 -0.4876      0.241     -2.027      0.043      -0.959      -0.016
field_eth_beneficence_eth_code                           0.4575      0.273      1.677      0.094      -0.077       0.992
field_eth_beneficence_mech_code                         -0.4124      0.116     -3.553      0.000      -0.640      -0.185
field_eth_equity_code                                   -0.5302      0.283     -1.871      0.061      -1.086       0.025
field_eth_integrity_code                                 0.6244      0.234      2.672      0.008       0.166       1.082
field_eth_justice_code                                   0.0115      0.296      0.039      0.969      -0.569       0.592
field_eth_nonmaleficence_code                            0.1284      0.271      0.473      0.636      -0.403       0.660
field_eth_privacy_code                                  -0.8530      0.232     -3.682      0.000      -1.307      -0.399
field_eth_transparency_code                              0.5466      0.296      1.847      0.065      -0.033       1.127
field_int_academic_intensity_code                       -5.1772      0.434    -11.928      0.000      -6.028      -4.327
field_int_assessment_intensity_code                     -4.2463      0.362    -11.716      0.000      -4.957      -3.536
field_int_data_prov_intensity_code                      -1.2636      0.212     -5.969      0.000      -1.679      -0.849
field_int_epistemic_intensity_code                      -1.1152      0.187     -5.971      0.000      -1.481      -0.749
field_int_research_intensity_code                       -4.0958      0.327    -12.512      0.000      -4.737      -3.454
field_int_system_tevv_intensity_code                    -1.0051      0.230     -4.366      0.000      -1.456      -0.554
field_journal_evidence_quality_strength_code            -0.6282      0.478     -1.316      0.188      -1.564       0.308
field_surveillance_monitoring_presence_strength_code    -4.0749      0.586     -6.948      0.000      -5.224      -2.925
==============================================================================
Skew:                         -0.1743   Kurtosis:                       1.6266
Centered skew:                 0.0100   Centered kurtosis:              1.2916
==============================================================================

Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
without_FRM: full coefficients, exclusions and diagnostics

Download all coefficients · Download complete model summary

{
  "model": "without_FRM",
  "converged": true,
  "iterations": 2,
  "warnings": [],
  "n_parameters": 33,
  "covariance_rank": 33,
  "smallest_cov_eigenvalue": 0.00018070086809806227,
  "max_absolute_parameter": 6.150103947656452,
  "threshold": null,
  "fields_excluded_for_no_binary_variation": [],
  "max_average_score": 6.093465451247267e-18,
  "sandwich_max_difference": 4.871103520542874e-15,
  "coder_covariance_rank": 2
}
                           OrdinalGEE Regression Results                           
===================================================================================
Dep. Variable:                      rating   No. Observations:                20880
Model:                          OrdinalGEE   No. clusters:                       53
Method:                        Generalized   Min. cluster size:                 390
                      Estimating Equations   Max. cluster size:                 405
Family:                           Binomial   Mean cluster size:               394.0
Dependence structure:         Independence   Num. iterations:                     2
Date:                     Mon, 07 Sep 2026   Scale:                           1.000
Covariance type:                    robust   Time:                         05:45:09
================================================================================================================
                                                   coef    std err          z      P>|z|      [0.025      0.975]
----------------------------------------------------------------------------------------------------------------
I(y>0.0)                                         3.5312      0.195     18.128      0.000       3.149       3.913
I(y>1.0)                                         2.7485      0.183     15.048      0.000       2.390       3.106
I(y>2.0)                                         0.7376      0.158      4.666      0.000       0.428       1.047
I(y>3.0)                                        -1.0432      0.176     -5.935      0.000      -1.388      -0.699
I(y>4.0)                                        -4.5256      0.250    -18.071      0.000      -5.016      -4.035
Codex                                            1.6524      0.092     17.984      0.000       1.472       1.833
Gemini                                           0.7873      0.104      7.601      0.000       0.584       0.990
field_cb_benefit_presence_strength_code          0.4569      0.074      6.162      0.000       0.312       0.602
field_cb_cost_intensity_code                     0.2161      0.265      0.815      0.415      -0.303       0.735
field_cb_cost_presence_strength_code             0.5190      0.278      1.865      0.062      -0.026       1.064
field_dppra_a_strength_code                     -1.4323      0.282     -5.086      0.000      -1.984      -0.880
field_dppra_alignment_score_code                -2.0701      0.241     -8.600      0.000      -2.542      -1.598
field_dppra_d_strength_code                      3.0405      0.329      9.240      0.000       2.396       3.685
field_dppra_p1_strength_code                    -0.2490      0.272     -0.914      0.361      -0.783       0.285
field_dppra_p2_strength_code                    -2.0019      0.270     -7.412      0.000      -2.531      -1.473
field_dppra_r_strength_code                     -3.0552      0.301    -10.167      0.000      -3.644      -2.466
field_eth_accountability_code                    1.6025      0.343      4.675      0.000       0.931       2.274
field_eth_autonomy_code                         -0.7591      0.234     -3.241      0.001      -1.218      -0.300
field_eth_beneficence_eth_code                   0.1722      0.253      0.679      0.497      -0.325       0.669
field_eth_beneficence_mech_code                 -0.3965      0.144     -2.757      0.006      -0.678      -0.115
field_eth_equity_code                           -0.4230      0.338     -1.250      0.211      -1.086       0.240
field_eth_integrity_code                         0.7428      0.263      2.830      0.005       0.228       1.257
field_eth_justice_code                          -0.4098      0.344     -1.191      0.234      -1.084       0.265
field_eth_nonmaleficence_code                    0.5977      0.334      1.789      0.074      -0.057       1.252
field_eth_privacy_code                          -0.7718      0.252     -3.068      0.002      -1.265      -0.279
field_eth_transparency_code                      1.2989      0.274      4.734      0.000       0.761       1.837
field_int_academic_intensity_code               -6.1501      0.500    -12.303      0.000      -7.130      -5.170
field_int_assessment_intensity_code             -5.0865      0.400    -12.709      0.000      -5.871      -4.302
field_int_data_prov_intensity_code              -1.3130      0.237     -5.543      0.000      -1.777      -0.849
field_int_epistemic_intensity_code              -1.2408      0.220     -5.648      0.000      -1.671      -0.810
field_int_research_intensity_code               -4.7137      0.309    -15.263      0.000      -5.319      -4.108
field_int_system_tevv_intensity_code            -0.3433      0.265     -1.296      0.195      -0.862       0.176
field_journal_evidence_quality_strength_code    -0.9825      0.532     -1.846      0.065      -2.026       0.061
==============================================================================
Skew:                         -0.1733   Kurtosis:                       2.4613
Centered skew:                -0.1133   Centered kurtosis:              2.2683
==============================================================================

Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
without_GOV: full coefficients, exclusions and diagnostics

Download all coefficients · Download complete model summary

{
  "model": "without_GOV",
  "converged": true,
  "iterations": 2,
  "warnings": [],
  "n_parameters": 35,
  "covariance_rank": 35,
  "smallest_cov_eigenvalue": 0.0001501614592392729,
  "max_absolute_parameter": 5.180682250251609,
  "threshold": null,
  "fields_excluded_for_no_binary_variation": [],
  "max_average_score": 1.07174857924423e-17,
  "sandwich_max_difference": 7.618905506490137e-15,
  "coder_covariance_rank": 2
}
                           OrdinalGEE Regression Results                           
===================================================================================
Dep. Variable:                      rating   No. Observations:                20055
Model:                          OrdinalGEE   No. clusters:                       50
Method:                        Generalized   Min. cluster size:                 390
                      Estimating Equations   Max. cluster size:                 420
Family:                           Binomial   Mean cluster size:               401.1
Dependence structure:         Independence   Num. iterations:                     2
Date:                     Mon, 07 Sep 2026   Scale:                           1.000
Covariance type:                    robust   Time:                         05:45:10
========================================================================================================================
                                                           coef    std err          z      P>|z|      [0.025      0.975]
------------------------------------------------------------------------------------------------------------------------
I(y>0.0)                                                 2.7548      0.213     12.939      0.000       2.338       3.172
I(y>1.0)                                                 2.0206      0.199     10.142      0.000       1.630       2.411
I(y>2.0)                                                 0.4339      0.154      2.817      0.005       0.132       0.736
I(y>3.0)                                                -1.1695      0.158     -7.415      0.000      -1.479      -0.860
I(y>4.0)                                                -4.1827      0.245    -17.043      0.000      -4.664      -3.702
Codex                                                    1.3591      0.099     13.698      0.000       1.165       1.554
Gemini                                                   0.6036      0.089      6.754      0.000       0.428       0.779
field_cb_benefit_presence_strength_code                  0.2798      0.062      4.542      0.000       0.159       0.401
field_cb_cost_intensity_code                             0.3455      0.259      1.332      0.183      -0.163       0.854
field_cb_cost_presence_strength_code                     0.5889      0.268      2.194      0.028       0.063       1.115
field_detection_verification_presence_strength_code     -4.8145      0.562     -8.560      0.000      -5.917      -3.712
field_dppra_a_strength_code                             -1.8388      0.242     -7.599      0.000      -2.313      -1.365
field_dppra_alignment_score_code                        -1.9775      0.228     -8.673      0.000      -2.424      -1.531
field_dppra_d_strength_code                              2.8506      0.339      8.404      0.000       2.186       3.515
field_dppra_p1_strength_code                            -0.9506      0.222     -4.283      0.000      -1.386      -0.516
field_dppra_p2_strength_code                            -1.8181      0.256     -7.098      0.000      -2.320      -1.316
field_dppra_r_strength_code                             -2.8707      0.318     -9.029      0.000      -3.494      -2.248
field_eth_accountability_code                            0.6868      0.353      1.946      0.052      -0.005       1.378
field_eth_autonomy_code                                 -0.3149      0.250     -1.258      0.208      -0.805       0.176
field_eth_beneficence_eth_code                           0.3455      0.271      1.274      0.203      -0.186       0.877
field_eth_beneficence_mech_code                         -0.4554      0.109     -4.183      0.000      -0.669      -0.242
field_eth_equity_code                                   -0.9506      0.293     -3.245      0.001      -1.525      -0.376
field_eth_integrity_code                                 0.8449      0.252      3.353      0.001       0.351       1.339
field_eth_justice_code                                  -0.6389      0.308     -2.076      0.038      -1.242      -0.036
field_eth_nonmaleficence_code                            0.2021      0.301      0.672      0.501      -0.387       0.792
field_eth_privacy_code                                  -1.2423      0.239     -5.189      0.000      -1.712      -0.773
field_eth_transparency_code                              0.4119      0.299      1.377      0.168      -0.174       0.998
field_int_academic_intensity_code                       -5.1807      0.469    -11.049      0.000      -6.100      -4.262
field_int_assessment_intensity_code                     -4.2061      0.386    -10.906      0.000      -4.962      -3.450
field_int_data_prov_intensity_code                      -1.5900      0.222     -7.151      0.000      -2.026      -1.154
field_int_epistemic_intensity_code                      -1.0596      0.194     -5.455      0.000      -1.440      -0.679
field_int_research_intensity_code                       -3.9160      0.327    -11.982      0.000      -4.557      -3.275
field_int_system_tevv_intensity_code                    -1.0920      0.259     -4.208      0.000      -1.601      -0.583
field_journal_evidence_quality_strength_code            -0.4984      0.479     -1.040      0.298      -1.438       0.441
field_surveillance_monitoring_presence_strength_code    -3.9280      0.585     -6.716      0.000      -5.074      -2.782
==============================================================================
Skew:                         -0.1247   Kurtosis:                       1.4727
Centered skew:                 0.0008   Centered kurtosis:              1.2363
==============================================================================

Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
without_INS: full coefficients, exclusions and diagnostics

Download all coefficients · Download complete model summary

{
  "model": "without_INS",
  "converged": true,
  "iterations": 2,
  "warnings": [],
  "n_parameters": 35,
  "covariance_rank": 35,
  "smallest_cov_eigenvalue": 0.00011407610041409381,
  "max_absolute_parameter": 5.423740126209203,
  "threshold": null,
  "fields_excluded_for_no_binary_variation": [],
  "max_average_score": 1.0578631392240664e-17,
  "sandwich_max_difference": 5.585809592645319e-15,
  "coder_covariance_rank": 2
}
                           OrdinalGEE Regression Results                           
===================================================================================
Dep. Variable:                      rating   No. Observations:                20445
Model:                          OrdinalGEE   No. clusters:                       51
Method:                        Generalized   Min. cluster size:                 390
                      Estimating Equations   Max. cluster size:                 420
Family:                           Binomial   Mean cluster size:               400.9
Dependence structure:         Independence   Num. iterations:                     2
Date:                     Mon, 07 Sep 2026   Scale:                           1.000
Covariance type:                    robust   Time:                         05:45:11
========================================================================================================================
                                                           coef    std err          z      P>|z|      [0.025      0.975]
------------------------------------------------------------------------------------------------------------------------
I(y>0.0)                                                 2.7982      0.214     13.079      0.000       2.379       3.217
I(y>1.0)                                                 2.0957      0.205     10.221      0.000       1.694       2.498
I(y>2.0)                                                 0.5857      0.154      3.806      0.000       0.284       0.887
I(y>3.0)                                                -0.9551      0.154     -6.195      0.000      -1.257      -0.653
I(y>4.0)                                                -3.9780      0.217    -18.358      0.000      -4.403      -3.553
Codex                                                    1.2983      0.102     12.769      0.000       1.099       1.498
Gemini                                                   0.5952      0.089      6.705      0.000       0.421       0.769
field_cb_benefit_presence_strength_code                  0.2882      0.060      4.772      0.000       0.170       0.407
field_cb_cost_intensity_code                             0.1734      0.240      0.724      0.469      -0.296       0.643
field_cb_cost_presence_strength_code                     0.4058      0.253      1.604      0.109      -0.090       0.902
field_detection_verification_presence_strength_code     -4.8500      0.559     -8.670      0.000      -5.946      -3.754
field_dppra_a_strength_code                             -1.5103      0.252     -6.001      0.000      -2.004      -1.017
field_dppra_alignment_score_code                        -1.7980      0.223     -8.066      0.000      -2.235      -1.361
field_dppra_d_strength_code                              2.3456      0.318      7.387      0.000       1.723       2.968
field_dppra_p1_strength_code                            -0.5679      0.252     -2.253      0.024      -1.062      -0.074
field_dppra_p2_strength_code                            -1.8084      0.234     -7.739      0.000      -2.266      -1.350
field_dppra_r_strength_code                             -2.7886      0.281     -9.911      0.000      -3.340      -2.237
field_eth_accountability_code                            0.6227      0.336      1.853      0.064      -0.036       1.281
field_eth_autonomy_code                                 -0.5679      0.256     -2.215      0.027      -1.071      -0.065
field_eth_beneficence_eth_code                           0.1609      0.270      0.597      0.551      -0.368       0.689
field_eth_beneficence_mech_code                         -0.1797      0.095     -1.885      0.059      -0.367       0.007
field_eth_equity_code                                   -1.0405      0.286     -3.638      0.000      -1.601      -0.480
field_eth_integrity_code                                 0.6088      0.230      2.648      0.008       0.158       1.059
field_eth_justice_code                                  -0.5346      0.305     -1.751      0.080      -1.133       0.064
field_eth_nonmaleficence_code                            0.1609      0.296      0.544      0.586      -0.419       0.740
field_eth_privacy_code                                  -1.1650      0.228     -5.116      0.000      -1.611      -0.719
field_eth_transparency_code                              0.3271      0.286      1.143      0.253      -0.234       0.888
field_int_academic_intensity_code                       -5.4237      0.428    -12.666      0.000      -6.263      -4.584
field_int_assessment_intensity_code                     -4.3075      0.374    -11.506      0.000      -5.041      -3.574
field_int_data_prov_intensity_code                      -1.5504      0.234     -6.637      0.000      -2.008      -1.093
field_int_epistemic_intensity_code                      -1.1340      0.196     -5.798      0.000      -1.517      -0.751
field_int_research_intensity_code                       -4.2121      0.327    -12.881      0.000      -4.853      -3.571
field_int_system_tevv_intensity_code                    -0.7862      0.269     -2.923      0.003      -1.313      -0.259
field_journal_evidence_quality_strength_code            -0.6278      0.468     -1.343      0.179      -1.544       0.289
field_surveillance_monitoring_presence_strength_code    -3.9733      0.579     -6.860      0.000      -5.109      -2.838
==============================================================================
Skew:                         -0.1869   Kurtosis:                       1.4472
Centered skew:                -0.0366   Centered kurtosis:              1.1856
==============================================================================

Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
without_JRN: full coefficients, exclusions and diagnostics

Download all coefficients · Download complete model summary

{
  "model": "without_JRN",
  "converged": true,
  "iterations": 2,
  "warnings": [],
  "n_parameters": 34,
  "covariance_rank": 34,
  "smallest_cov_eigenvalue": 0.00014102865305170537,
  "max_absolute_parameter": 5.366593655886791,
  "threshold": null,
  "fields_excluded_for_no_binary_variation": [],
  "max_average_score": 5.711599548479308e-18,
  "sandwich_max_difference": 6.482661629725328e-15,
  "coder_covariance_rank": 2
}
                           OrdinalGEE Regression Results                           
===================================================================================
Dep. Variable:                      rating   No. Observations:                20235
Model:                          OrdinalGEE   No. clusters:                       51
Method:                        Generalized   Min. cluster size:                 390
                      Estimating Equations   Max. cluster size:                 420
Family:                           Binomial   Mean cluster size:               396.8
Dependence structure:         Independence   Num. iterations:                     2
Date:                     Mon, 07 Sep 2026   Scale:                           1.000
Covariance type:                    robust   Time:                         05:45:12
========================================================================================================================
                                                           coef    std err          z      P>|z|      [0.025      0.975]
------------------------------------------------------------------------------------------------------------------------
I(y>0.0)                                                 2.6826      0.206     13.016      0.000       2.279       3.087
I(y>1.0)                                                 1.9391      0.190     10.194      0.000       1.566       2.312
I(y>2.0)                                                 0.4352      0.140      3.105      0.002       0.160       0.710
I(y>3.0)                                                -1.0854      0.140     -7.740      0.000      -1.360      -0.811
I(y>4.0)                                                -4.3995      0.197    -22.295      0.000      -4.786      -4.013
Codex                                                    1.3744      0.111     12.362      0.000       1.156       1.592
Gemini                                                   0.8146      0.089      9.152      0.000       0.640       0.989
field_cb_benefit_presence_strength_code                  0.2777      0.065      4.296      0.000       0.151       0.404
field_cb_cost_intensity_code                             0.3433      0.228      1.503      0.133      -0.104       0.791
field_cb_cost_presence_strength_code                     0.5743      0.250      2.293      0.022       0.083       1.065
field_detection_verification_presence_strength_code     -4.8250      0.558     -8.639      0.000      -5.920      -3.730
field_dppra_a_strength_code                             -1.1880      0.231     -5.149      0.000      -1.640      -0.736
field_dppra_alignment_score_code                        -1.5488      0.182     -8.523      0.000      -1.905      -1.193
field_dppra_d_strength_code                              2.5067      0.320      7.839      0.000       1.880       3.133
field_dppra_p1_strength_code                            -0.2278      0.241     -0.946      0.344      -0.700       0.244
field_dppra_p2_strength_code                            -1.3517      0.191     -7.090      0.000      -1.725      -0.978
field_dppra_r_strength_code                             -2.1388      0.197    -10.847      0.000      -2.525      -1.752
field_eth_accountability_code                            0.8342      0.327      2.551      0.011       0.193       1.475
field_eth_autonomy_code                                 -0.4239      0.229     -1.854      0.064      -0.872       0.024
field_eth_beneficence_eth_code                           0.3832      0.249      1.539      0.124      -0.105       0.871
field_eth_beneficence_mech_code                         -0.4804      0.115     -4.189      0.000      -0.705      -0.256
field_eth_equity_code                                   -0.8846      0.262     -3.372      0.001      -1.399      -0.370
field_eth_integrity_code                                 0.8342      0.218      3.835      0.000       0.408       1.261
field_eth_justice_code                                  -0.4579      0.267     -1.714      0.086      -0.981       0.066
field_eth_nonmaleficence_code                            0.6877      0.300      2.290      0.022       0.099       1.276
field_eth_privacy_code                                  -1.0008      0.233     -4.291      0.000      -1.458      -0.544
field_eth_transparency_code                              0.3566      0.266      1.341      0.180      -0.165       0.878
field_int_academic_intensity_code                       -5.3666      0.505    -10.619      0.000      -6.357      -4.376
field_int_assessment_intensity_code                     -4.4162      0.376    -11.730      0.000      -5.154      -3.678
field_int_data_prov_intensity_code                      -1.2496      0.215     -5.802      0.000      -1.672      -0.827
field_int_epistemic_intensity_code                      -1.0844      0.192     -5.640      0.000      -1.461      -0.708
field_int_research_intensity_code                       -4.1082      0.305    -13.453      0.000      -4.707      -3.510
field_int_system_tevv_intensity_code                    -0.6140      0.262     -2.340      0.019      -1.128      -0.100
field_surveillance_monitoring_presence_strength_code    -3.9449      0.582     -6.782      0.000      -5.085      -2.805
==============================================================================
Skew:                         -0.2757   Kurtosis:                       1.6224
Centered skew:                -0.0723   Centered kurtosis:              1.2959
==============================================================================

Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
above_0: full coefficients, exclusions and diagnostics

Download all coefficients · Download complete model summary

{
  "model": "above_0",
  "converged": true,
  "iterations": 2,
  "warnings": [],
  "n_parameters": 26,
  "covariance_rank": 23,
  "smallest_cov_eigenvalue": -1.143355803459955e-14,
  "max_absolute_parameter": 7.142644151258511,
  "threshold": 0,
  "fields_excluded_for_no_binary_variation": [
    "cb_benefit_intensity_code",
    "cb_benefit_presence_strength_code",
    "cb_cost_intensity_code",
    "cb_cost_presence_strength_code",
    "eth_beneficence_eth_code"
  ],
  "max_average_score": 1.9295602258701603e-17,
  "sandwich_max_difference": 8.931744233109384e-14,
  "coder_covariance_rank": 2
}
                               GEE Regression Results                              
===================================================================================
Dep. Variable:                      binary   No. Observations:                 4206
Model:                                 GEE   No. clusters:                       65
Method:                        Generalized   Min. cluster size:                  63
                      Estimating Equations   Max. cluster size:                  69
Family:                           Binomial   Mean cluster size:                64.7
Dependence structure:         Independence   Num. iterations:                     2
Date:                     Mon, 07 Sep 2026   Scale:                           1.000
Covariance type:                    robust   Time:                         05:45:12
========================================================================================================================
                                                           coef    std err          z      P>|z|      [0.025      0.975]
------------------------------------------------------------------------------------------------------------------------
intercept                                               -2.1811      0.526     -4.148      0.000      -3.212      -1.151
Codex                                                    0.5496      0.097      5.667      0.000       0.360       0.740
Gemini                                                   0.4589      0.157      2.928      0.003       0.152       0.766
field_dppra_a_strength_code                              3.7838      0.480      7.882      0.000       2.843       4.725
field_dppra_alignment_score_code                         4.5791      0.551      8.313      0.000       3.500       5.659
field_dppra_d_strength_code                              7.1426      1.025      6.969      0.000       5.134       9.151
field_dppra_p1_strength_code                             5.1633      0.569      9.082      0.000       4.049       6.278
field_dppra_p2_strength_code                             3.4190      0.519      6.585      0.000       2.401       4.437
field_dppra_r_strength_code                              2.4324      0.577      4.212      0.000       1.301       3.564
field_eth_accountability_code                            5.5110      0.657      8.388      0.000       4.223       6.799
field_eth_autonomy_code                                  6.0330      1.047      5.760      0.000       3.980       8.086
field_eth_beneficence_mech_code                          7.1426      1.025      6.969      0.000       5.134       9.151
field_eth_equity_code                                    4.1549      0.581      7.153      0.000       3.017       5.293
field_eth_integrity_code                                 7.1426      1.025      6.969      0.000       5.134       9.151
field_eth_justice_code                                   4.7896      0.611      7.839      0.000       3.592       5.987
field_eth_nonmaleficence_code                            6.4440      0.750      8.591      0.000       4.974       7.914
field_eth_privacy_code                                   4.0949      0.659      6.212      0.000       2.803       5.387
field_eth_transparency_code                              5.1633      0.625      8.266      0.000       3.939       6.388
field_int_academic_intensity_code                        0.1597      0.574      0.278      0.781      -0.966       1.285
field_int_assessment_intensity_code                      0.6253      0.604      1.035      0.301      -0.559       1.809
field_int_data_prov_intensity_code                       3.5700      0.633      5.636      0.000       2.328       4.811
field_int_epistemic_intensity_code                       4.5099      0.621      7.262      0.000       3.293       5.727
field_int_research_intensity_code                        1.1650      0.534      2.181      0.029       0.118       2.212
field_int_system_tevv_intensity_code                     3.9301      0.494      7.953      0.000       2.962       4.899
field_journal_evidence_quality_strength_code             4.4355      0.760      5.840      0.000       2.947       5.924
field_surveillance_monitoring_presence_strength_code     1.2658      0.635      1.995      0.046       0.022       2.509
==============================================================================
Skew:                         -1.0879   Kurtosis:                       3.9528
Centered skew:                -0.5955   Centered kurtosis:              3.3379
==============================================================================

Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
above_1: full coefficients, exclusions and diagnostics

Download all coefficients · Download complete model summary

{
  "model": "above_1",
  "converged": true,
  "iterations": 2,
  "warnings": [],
  "n_parameters": 31,
  "covariance_rank": 30,
  "smallest_cov_eigenvalue": -3.668266906557964e-13,
  "max_absolute_parameter": 7.552367093744065,
  "threshold": 1,
  "fields_excluded_for_no_binary_variation": [],
  "max_average_score": 4.200033059767046e-17,
  "sandwich_max_difference": 1.5306644840507033e-12,
  "coder_covariance_rank": 2
}
                               GEE Regression Results                              
===================================================================================
Dep. Variable:                      binary   No. Observations:                 5181
Model:                                 GEE   No. clusters:                       65
Method:                        Generalized   Min. cluster size:                  78
                      Estimating Equations   Max. cluster size:                  84
Family:                           Binomial   Mean cluster size:                79.7
Dependence structure:         Independence   Num. iterations:                     2
Date:                     Mon, 07 Sep 2026   Scale:                           1.000
Covariance type:                    robust   Time:                         05:45:12
========================================================================================================================
                                                           coef    std err          z      P>|z|      [0.025      0.975]
------------------------------------------------------------------------------------------------------------------------
intercept                                                4.1009      0.707      5.799      0.000       2.715       5.487
Codex                                                    1.0282      0.109      9.449      0.000       0.815       1.241
Gemini                                                   0.6470      0.112      5.782      0.000       0.428       0.866
field_cb_benefit_presence_strength_code               1.542e-14        nan        nan        nan         nan         nan
field_cb_cost_intensity_code                            -0.9349      0.990     -0.944      0.345      -2.875       1.005
field_cb_cost_presence_strength_code                    -0.9349      0.990     -0.944      0.345      -2.875       1.005
field_detection_verification_presence_strength_code     -6.5471      0.903     -7.253      0.000      -8.316      -4.778
field_dppra_a_strength_code                             -3.1519      0.682     -4.624      0.000      -4.488      -1.816
field_dppra_alignment_score_code                        -3.7355      0.700     -5.335      0.000      -5.108      -2.363
field_dppra_d_strength_code                           1.371e-14      0.716   1.92e-14      1.000      -1.403       1.403
field_dppra_p1_strength_code                            -2.1683      0.681     -3.183      0.001      -3.504      -0.833
field_dppra_p2_strength_code                            -3.6065      0.673     -5.357      0.000      -4.926      -2.287
field_dppra_r_strength_code                             -4.5217      0.692     -6.531      0.000      -5.879      -3.165
field_eth_accountability_code                           -1.9412      0.698     -2.780      0.005      -3.310      -0.572
field_eth_autonomy_code                                 -2.0217      0.811     -2.493      0.013      -3.611      -0.432
field_eth_beneficence_eth_code                          -1.9412      0.690     -2.814      0.005      -3.293      -0.589
field_eth_beneficence_mech_code                         -2.0973      0.591     -3.546      0.000      -3.256      -0.938
field_eth_equity_code                                   -3.2520      0.696     -4.674      0.000      -4.616      -1.888
field_eth_integrity_code                                -0.4116      0.717     -0.574      0.566      -1.817       0.994
field_eth_justice_code                                  -2.8511      0.779     -3.662      0.000      -4.377      -1.325
field_eth_nonmaleficence_code                           -1.8547      0.773     -2.398      0.016      -3.371      -0.339
field_eth_privacy_code                                  -3.1172      0.668     -4.664      0.000      -4.427      -1.807
field_eth_transparency_code                             -2.0973      0.698     -3.004      0.003      -3.466      -0.729
field_int_academic_intensity_code                       -7.5524      0.864     -8.738      0.000      -9.246      -5.858
field_int_assessment_intensity_code                     -6.1661      0.773     -7.982      0.000      -7.680      -4.652
field_int_data_prov_intensity_code                      -3.1859      0.640     -4.982      0.000      -4.439      -1.932
field_int_epistemic_intensity_code                      -2.3600      0.703     -3.358      0.001      -3.738      -0.982
field_int_research_intensity_code                       -5.9721      0.722     -8.271      0.000      -7.387      -4.557
field_int_system_tevv_intensity_code                    -2.8090      0.677     -4.147      0.000      -4.137      -1.481
field_journal_evidence_quality_strength_code            -2.0162      0.899     -2.243      0.025      -3.778      -0.254
field_surveillance_monitoring_presence_strength_code    -5.5162      1.071     -5.151      0.000      -7.615      -3.417
==============================================================================
Skew:                         -1.2471   Kurtosis:                       2.7631
Centered skew:                -0.7794   Centered kurtosis:              2.4712
==============================================================================

Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
above_2: full coefficients, exclusions and diagnostics

Download all coefficients · Download complete model summary

{
  "model": "above_2",
  "converged": true,
  "iterations": 2,
  "warnings": [],
  "n_parameters": 31,
  "covariance_rank": 31,
  "smallest_cov_eigenvalue": 0.0011357310360301605,
  "max_absolute_parameter": 5.111813895504899,
  "threshold": 2,
  "fields_excluded_for_no_binary_variation": [],
  "max_average_score": 1.3457248783335231e-17,
  "sandwich_max_difference": 7.931155732165962e-15,
  "coder_covariance_rank": 2
}
                               GEE Regression Results                              
===================================================================================
Dep. Variable:                      binary   No. Observations:                 5181
Model:                                 GEE   No. clusters:                       65
Method:                        Generalized   Min. cluster size:                  78
                      Estimating Equations   Max. cluster size:                  84
Family:                           Binomial   Mean cluster size:                79.7
Dependence structure:         Independence   Num. iterations:                     2
Date:                     Mon, 07 Sep 2026   Scale:                           1.000
Covariance type:                    robust   Time:                         05:45:13
========================================================================================================================
                                                           coef    std err          z      P>|z|      [0.025      0.975]
------------------------------------------------------------------------------------------------------------------------
intercept                                                0.5887      0.196      3.003      0.003       0.204       0.973
Codex                                                    1.6148      0.111     14.486      0.000       1.396       1.833
Gemini                                                   0.6736      0.095      7.091      0.000       0.487       0.860
field_cb_benefit_presence_strength_code                  0.3415      0.109      3.126      0.002       0.127       0.556
field_cb_cost_intensity_code                             0.5407      0.316      1.711      0.087      -0.079       1.160
field_cb_cost_presence_strength_code                     0.7205      0.328      2.197      0.028       0.078       1.363
field_detection_verification_presence_strength_code     -3.9306      0.600     -6.554      0.000      -5.106      -2.755
field_dppra_a_strength_code                             -1.9648      0.265     -7.405      0.000      -2.485      -1.445
field_dppra_alignment_score_code                        -2.1142      0.268     -7.890      0.000      -2.639      -1.589
field_dppra_d_strength_code                              2.1268      0.405      5.246      0.000       1.332       2.921
field_dppra_p1_strength_code                            -1.3581      0.248     -5.485      0.000      -1.843      -0.873
field_dppra_p2_strength_code                            -1.7965      0.251     -7.151      0.000      -2.289      -1.304
field_dppra_r_strength_code                             -2.6404      0.299     -8.831      0.000      -3.226      -2.054
field_eth_accountability_code                            0.5407      0.351      1.540      0.123      -0.147       1.229
field_eth_autonomy_code                                 -0.3482      0.262     -1.327      0.185      -0.862       0.166
field_eth_beneficence_eth_code                           0.0962      0.275      0.350      0.726      -0.442       0.635
field_eth_beneficence_mech_code                         -0.6823      0.133     -5.128      0.000      -0.943      -0.421
field_eth_equity_code                                   -0.8034      0.262     -3.063      0.002      -1.317      -0.289
field_eth_integrity_code                                 0.4989      0.298      1.673      0.094      -0.086       1.083
field_eth_justice_code                                  -0.4284      0.283     -1.515      0.130      -0.983       0.126
field_eth_nonmaleficence_code                            0.1631      0.315      0.518      0.604      -0.454       0.780
field_eth_privacy_code                                  -1.0611      0.237     -4.480      0.000      -1.525      -0.597
field_eth_transparency_code                              0.4989      0.304      1.639      0.101      -0.098       1.095
field_int_academic_intensity_code                       -5.0052      0.611     -8.193      0.000      -6.203      -3.808
field_int_assessment_intensity_code                     -3.4790      0.333    -10.435      0.000      -4.132      -2.826
field_int_data_prov_intensity_code                      -1.4492      0.227     -6.382      0.000      -1.894      -1.004
field_int_epistemic_intensity_code                      -1.3581      0.219     -6.193      0.000      -1.788      -0.928
field_int_research_intensity_code                       -3.7526      0.346    -10.834      0.000      -4.432      -3.074
field_int_system_tevv_intensity_code                    -0.9217      0.237     -3.885      0.000      -1.387      -0.457
field_journal_evidence_quality_strength_code            -0.8102      0.522     -1.553      0.120      -1.833       0.212
field_surveillance_monitoring_presence_strength_code    -5.1118      0.994     -5.143      0.000      -7.060      -3.164
==============================================================================
Skew:                         -0.3074   Kurtosis:                      -0.3985
Centered skew:                -0.2340   Centered kurtosis:             -0.2894
==============================================================================

Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
above_3: full coefficients, exclusions and diagnostics

Download all coefficients · Download complete model summary

{
  "model": "above_3",
  "converged": true,
  "iterations": 2,
  "warnings": [],
  "n_parameters": 29,
  "covariance_rank": 29,
  "smallest_cov_eigenvalue": 0.0011104048247526743,
  "max_absolute_parameter": 4.0725443633151865,
  "threshold": 3,
  "fields_excluded_for_no_binary_variation": [
    "detection_verification_presence_strength_code",
    "surveillance_monitoring_presence_strength_code"
  ],
  "max_average_score": 4.3896075743644865e-17,
  "sandwich_max_difference": 5.592748486549226e-15,
  "coder_covariance_rank": 2
}
                               GEE Regression Results                              
===================================================================================
Dep. Variable:                      binary   No. Observations:                 5109
Model:                                 GEE   No. clusters:                       65
Method:                        Generalized   Min. cluster size:                  75
                      Estimating Equations   Max. cluster size:                  81
Family:                           Binomial   Mean cluster size:                78.6
Dependence structure:         Independence   Num. iterations:                     2
Date:                     Mon, 07 Sep 2026   Scale:                           1.000
Covariance type:                    robust   Time:                         05:45:13
================================================================================================================
                                                   coef    std err          z      P>|z|      [0.025      0.975]
----------------------------------------------------------------------------------------------------------------
intercept                                       -1.6835      0.196     -8.570      0.000      -2.068      -1.298
Codex                                            1.8539      0.115     16.154      0.000       1.629       2.079
Gemini                                           0.9124      0.107      8.498      0.000       0.702       1.123
field_cb_benefit_presence_strength_code          0.4676      0.098      4.780      0.000       0.276       0.659
field_cb_cost_intensity_code                     0.4199      0.290      1.450      0.147      -0.148       0.988
field_cb_cost_presence_strength_code             0.8676      0.305      2.849      0.004       0.271       1.464
field_dppra_a_strength_code                     -0.8276      0.347     -2.387      0.017      -1.507      -0.148
field_dppra_alignment_score_code                -2.6233      0.454     -5.780      0.000      -3.513      -1.734
field_dppra_d_strength_code                      2.9831      0.334      8.940      0.000       2.329       3.637
field_dppra_p1_strength_code                     0.2508      0.268      0.936      0.349      -0.274       0.776
field_dppra_p2_strength_code                    -1.5143      0.348     -4.357      0.000      -2.196      -0.833
field_dppra_r_strength_code                     -2.7811      0.442     -6.295      0.000      -3.647      -1.915
field_eth_accountability_code                    1.6281      0.308      5.291      0.000       1.025       2.231
field_eth_autonomy_code                         -0.7925      0.347     -2.283      0.022      -1.473      -0.112
field_eth_beneficence_eth_code                   0.6800      0.267      2.551      0.011       0.158       1.202
field_eth_beneficence_mech_code                 -0.1046      0.158     -0.663      0.507      -0.414       0.205
field_eth_equity_code                            0.0767      0.303      0.253      0.800      -0.517       0.670
field_eth_integrity_code                         1.1760      0.248      4.744      0.000       0.690       1.662
field_eth_justice_code                           0.1270      0.319      0.398      0.691      -0.498       0.752
field_eth_nonmaleficence_code                    0.7503      0.297      2.526      0.012       0.168       1.332
field_eth_privacy_code                          -0.3815      0.294     -1.298      0.194      -0.958       0.195
field_eth_transparency_code                      1.2731      0.287      4.439      0.000       0.711       1.835
field_int_academic_intensity_code               -4.0725      1.051     -3.875      0.000      -6.133      -2.012
field_int_assessment_intensity_code             -3.3640      0.568     -5.925      0.000      -4.477      -2.251
field_int_data_prov_intensity_code              -0.8276      0.292     -2.831      0.005      -1.401      -0.255
field_int_epistemic_intensity_code              -1.1355      0.309     -3.675      0.000      -1.741      -0.530
field_int_research_intensity_code               -4.0725      0.772     -5.272      0.000      -5.587      -2.559
field_int_system_tevv_intensity_code             0.0257      0.294      0.087      0.930      -0.551       0.602
field_journal_evidence_quality_strength_code    -1.2283      0.712     -1.724      0.085      -2.625       0.168
==============================================================================
Skew:                          0.4555   Kurtosis:                      -0.3312
Centered skew:                 0.4184   Centered kurtosis:             -0.2512
==============================================================================

Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.
above_4: full coefficients, exclusions and diagnostics

Download all coefficients · Download complete model summary

{
  "model": "above_4",
  "converged": true,
  "iterations": 2,
  "warnings": [],
  "n_parameters": 21,
  "covariance_rank": 21,
  "smallest_cov_eigenvalue": 0.004076421807522495,
  "max_absolute_parameter": 4.643468389491005,
  "threshold": 4,
  "fields_excluded_for_no_binary_variation": [
    "detection_verification_presence_strength_code",
    "dppra_alignment_score_code",
    "dppra_p2_strength_code",
    "dppra_r_strength_code",
    "int_academic_intensity_code",
    "int_assessment_intensity_code",
    "int_data_prov_intensity_code",
    "int_epistemic_intensity_code",
    "int_research_intensity_code",
    "surveillance_monitoring_presence_strength_code"
  ],
  "max_average_score": 1.1127235269328709e-17,
  "sandwich_max_difference": 1.0380585280245214e-13,
  "coder_covariance_rank": 2
}
                               GEE Regression Results                              
===================================================================================
Dep. Variable:                      binary   No. Observations:                 3552
Model:                                 GEE   No. clusters:                       65
Method:                        Generalized   Min. cluster size:                  54
                      Estimating Equations   Max. cluster size:                  57
Family:                           Binomial   Mean cluster size:                54.6
Dependence structure:         Independence   Num. iterations:                     2
Date:                     Mon, 07 Sep 2026   Scale:                           1.000
Covariance type:                    robust   Time:                         05:45:13
================================================================================================================
                                                   coef    std err          z      P>|z|      [0.025      0.975]
----------------------------------------------------------------------------------------------------------------
intercept                                       -4.6435      0.738     -6.288      0.000      -6.091      -3.196
Codex                                            0.9454      0.198      4.783      0.000       0.558       1.333
Gemini                                           0.2700      0.214      1.260      0.207      -0.150       0.690
field_cb_benefit_presence_strength_code          0.5231      0.417      1.255      0.210      -0.294       1.340
field_cb_cost_intensity_code                  8.103e-15      0.960   8.44e-15      1.000      -1.882       1.882
field_cb_cost_presence_strength_code             0.2938      0.916      0.321      0.748      -1.501       2.089
field_dppra_a_strength_code                     -1.1108      1.260     -0.881      0.378      -3.581       1.359
field_dppra_d_strength_code                      3.5916      0.784      4.584      0.000       2.056       5.127
field_dppra_p1_strength_code                     1.2474      0.857      1.455      0.146      -0.432       2.927
field_eth_accountability_code                    2.0051      0.816      2.456      0.014       0.405       3.605
field_eth_autonomy_code                          0.2938      1.110      0.265      0.791      -1.882       2.469
field_eth_beneficence_eth_code                   1.5288      0.874      1.749      0.080      -0.184       3.242
field_eth_beneficence_mech_code                  0.7116      0.337      2.114      0.035       0.052       1.371
field_eth_equity_code                         8.278e-15      0.960   8.62e-15      1.000      -1.882       1.882
field_eth_integrity_code                         1.2474      0.857      1.455      0.146      -0.433       2.927
field_eth_justice_code                           1.0117      0.865      1.169      0.242      -0.684       2.708
field_eth_nonmaleficence_code                    1.4424      0.867      1.663      0.096      -0.258       3.142
field_eth_privacy_code                          -1.1108      1.260     -0.881      0.378      -3.581       1.359
field_eth_transparency_code                      1.1357      0.930      1.221      0.222      -0.687       2.959
field_int_system_tevv_intensity_code            -1.1108      1.260     -0.881      0.378      -3.581       1.360
field_journal_evidence_quality_strength_code     2.3894      1.025      2.331      0.020       0.381       4.398
==============================================================================
Skew:                          3.3864   Kurtosis:                      12.8510
Centered skew:                 3.2193   Centered kurtosis:             12.0258
==============================================================================

Reported pairwise contrasts use robust covariance multiplied by G/(G-1) and t(G-1) reference, with Holm adjustment.

Coder configuration and provenance

Coder configuration for the three version-locked runs behind the paper. RUN 07 (v2.14) is the source of every figure in this addendum. Reviewers asked for exact model versions, access dates, interfaces, sampling parameters and session-memory conditions. They are answerable to different degrees, and none of the three resolves to a dated build. The Codex coder’s model label and client version are recorded in the session logs for every run. The Claude coder’s tier rests on investigator testimony, corroborated by dated build records spanning the coding window. The Gemini tier is reported by the app’s own response-details panel for all three runs. Across all three runs, every coder is now identified by something better than testimony. Every label below is nonetheless a product-surface label, so snapshot identity is not claimed for any coder.

What is and is not recoverable
ClassItems
Artifact-backedinstrument version; packet texts (SHA-256); coding dates; coder identity stamped in every output; output schema; Codex model label and client version, from the session logs; Gemini tier for all three runs, from the app's response-details panel
Testimony-backedClaude tier, corroborated by dated build records; interface for all three coders
Unrecoverabledated model snapshot id for any coder; API version; temperature; top-p; model-side system prompt

All three runs fall inside a seven-day window, 24-30 June 2026. RUN 07 was coded in a single day, 30 June 2026. This bounds exposure to an unannounced point release; it is not proof of model constancy.

Coder configuration for all three version-locked runs
RunCoderModel as runInterfaceOutputsCoding datesPackets hashedCombined SHA-256 of packet set
RUN 04 · v2.12ClaudeAnthropic Claude — Opus 4.8, highest tier then available (investigator testimony, corroborated by dated build records across the window)chat, packet-delivered via a local dashboard652026-06-2465f6d4a5780d471ba7c165b3e1c3cb302298bd8b7e7ac89d0f0682073617bd33f8
RUN 04 · v2.12CodexOpenAI Codex — model label gpt-5.5, Codex Desktop CLI 0.142.0 (recorded, not testimony)ChatGPT-subscription Codex CLI, not the API652026-06-246550f1d6964f694f73085867cb4802d1f053df7ecfd947ba2db4447b9d6a4e4207
RUN 04 · v2.12GeminiGoogle Gemini — 3.5 Flash, reported by the app's response-details panel on the coding thread for this runGemini consumer app, hand-run by the investigator652026-06-24, 2026-06-25, 2026-06-26, 2026-06-27not recoverable
RUN 06 · v2.13ClaudeAnthropic Claude — Opus 4.8, highest tier then available (investigator testimony, corroborated by dated build records across the window)chat, packet-delivered via a local dashboard652026-06-28, 2026-06-2955490726fa38cc45c1140658b1b69a90f34b1437b8d96ffec4f2ef4d601f64f2cc
RUN 06 · v2.13CodexOpenAI Codex — model label gpt-5.5, Codex Desktop CLI 0.142.3 (recorded, not testimony)ChatGPT-subscription Codex CLI, not the API652026-06-28, 2026-06-2955410c379331c43224b19cc7b10a90ee332272f0d66d46696f3013862d165ab9ed
RUN 06 · v2.13GeminiGoogle Gemini — 3.5 Flash, reported by the app's response-details panel on the coding thread for this runGemini consumer app, hand-run by the investigator652026-06-28, 2026-06-29not recoverable
RUN 07 · v2.14ClaudeAnthropic Claude — Opus 4.8, highest tier then available (investigator testimony, corroborated by dated build records across the window)chat, packet-delivered via a local dashboard652026-06-30657d0f7d86d66a3786d2d88659086a747e33c626fd35ac8155bfc3104f824eec1b
RUN 07 · v2.14CodexOpenAI Codex — model label gpt-5.5, Codex Desktop CLI 0.142.3 (recorded, not testimony)ChatGPT-subscription Codex CLI, not the API652026-06-30657d0f7d86d66a3786d2d88659086a747e33c626fd35ac8155bfc3104f824eec1b
RUN 07 · v2.14GeminiGoogle Gemini — 3.5 Flash, reported by the app's response-details panel on the coding thread for this runGemini consumer app, hand-run by the investigator652026-06-306522a27b1f363fe17267ae093686d0225fa79ccd6abf03de7a5a5eae93aff9a332

How each model was identified

Codex. The Codex coder ran through a desktop client that writes a session log for every run. Those logs record the model label and client version. 188 sessions on 30 June 2026; 187 name documents, together covering all 65; every one records model label gpt-5.5 and client version 0.142.3. Session of 28 June 2026: model label gpt-5.5, client version 0.142.3. Session of 24 June 2026: model label gpt-5.5, client version 0.142.0.

gpt-5.5 is the label the client recorded, not a dated build identifier. It fixes which model line was addressed, not which build answered. The session logs themselves are not published: they contain full document text.

Claude. The Claude coder's outputs carry no model field, so the tier rests on investigator testimony that the highest available build was used throughout. Dated build records from the same workstation name Claude Opus 4.8 on 25, 27, 28 and 29 June and on 1 and 3 July 2026 — spanning the whole coding window — with the next Anthropic build first appearing on 5 July, after all three runs.

This establishes which Anthropic build was in use on those dates. It does not stamp the chat session that performed the coding, which is a different surface. Reported as corroborated testimony, not as identification.

Gemini. The Gemini app reports, per response, which model answered. On the retained coding threads that panel reads 3.5 Flash, with the composer's model selector also on Flash. Checked on a retained thread for each of the three runs. All three read 3.5 Flash.

Each thread is tied to its run by content, not by assertion. The RUN 07 thread's record header names the same source id, document id and title as the archived RUN 07 Gemini output and stamps prompt version v2.14 and coding date 30 June 2026; its narrative content matches that output string for string, and its two emergent-code strings appear in no other output in the archive. The RUN 06 thread declares schema version v2.13 and codes a named corporate scaling policy from that run.

One RUN 04 output self-stamped a coder id containing 'pro', which conflicted with the Flash testimony. That string was written by the model, and model self-identification is unreliable. The app's own panel reports Flash. The self-stamp is reported as an adherence deviation, not as an attribution. Like the Codex label, this is a product-surface label rather than a dated build identifier.

Sampling parameters, session memory and output settings were not settable and were not recorded. Claude and Codex ran through subscription interfaces exposing no temperature or top-p control; Gemini was hand-run in the consumer app. Session-memory state was not captured. This is a limitation of consumer-interface coding, reported rather than estimated.

Gemini packets are staged in place and overwritten between runs, so only the most recent staging survives. All 65 carry a 30 June 2026 modification time and are artifact-backed for RUN 07 only; the RUN 04 and RUN 06 Gemini packets are gone. Packet coverage for RUN 07: 65 unique sources (COR 10, FRM 12, GOV 15, INS 14, JRN 14).

Stale version string in the watch-out block. Every v2.14 packet declares prompt_version v2.14 in the body and v2.13 inside the coder watch-out block, because v2.14 was additive and did not touch that block. It did not propagate: all 195 RUN 07 outputs stamped v2.14.
RUN 07 Codex packets are byte-identical copies of the Claude packets. In RUN 06 the two packet sets differ in exactly three lines - the addressing line, the block header and the coder instruction. In RUN 07 the Codex folder holds byte-identical copies of the Claude files, still addressed to Claude. Coder identity for that run was carried by the run sheet instead, and Codex stamped its own coder name on 65 of 65 outputs. The watch-out block is coder-neutral by design from v2.13 onward, so the substance Codex received matches a Codex-addressed packet. The residual risk is an addressing line reading as though the block applied to another coder. Reported here rather than left for a reader to find.

The machine-readable source for this section is provenance_configuration.json, included in the reproducibility package.

What these checks establish

The overall ordering Codex > Gemini > Claude survives document-level comparisons, a check using order alone, and clustered ordinal models. The evidence supports a directional tendency among these tested coder configurations.

Documents are assumed independent between clusters. The selected corpus is not a probability sample, the instrument fields are fixed, and unmodeled between-document dependence remains possible. Missingness follows production conventions rather than a missing-data model. The checks do not establish validity against an external criterion, causal effects of protocol revisions, inherited model-family traits, or the validity of the separate nominal-field χ² tests. Account, interface and exact-build provenance remain separate questions.

This dated release includes every test in the reviewer-requested ordinal sensitivity battery. Earlier protocol experiments and qualitative counter-readings are separate bodies of evidence and are not presented as newly reproduced here.

Reproduce and extend the results

Download the complete reproducibility package

The package includes the exact numerical inputs, pinned dependencies, all outputs, original post-review analysis plan, and portable scripts. No raw corpus text, private working notes or local computer paths are included. Run checks.py, ordinal_checks.py, then verify.py; see the included README for setup and interpretation. The 195 original source-file hashes were checked unchanged during the local audit; reproduction from the public numerical export does not independently verify those unavailable source files.

Browse every downloadable file

Release history

— First statistical addendum: original-results reproduction, six document tests, six order-only checks, 12 GEE fits with 36 contrasts, sphericity sensitivity and numerical verification. Future analyses will be added as separately dated releases with their own data, methods and scope.

Method reference: statsmodels OrdinalGEE documentation.