Discussion
GPT-5 Pro shows positive rank agreement with human expert evaluations. The comparison with human inter-rater agreement remains uncertain: the paired bootstrap uses an approximate reliability adjustment, and failure to detect a difference does not establish equivalence to an additional expert. Central compression—the tendency to pull extreme ratings toward the middle of the scale—is the most consistent pattern, whose mechanism this observational comparison cannot identify. Journal-tier predictions are the dimension where human and model judgments align most closely, and on the small set of papers whose publication outcomes have already resolved, both predict realized venues with broadly similar accuracy. Qualitative coverage varies widely across papers: on some, the model captures nearly all consensus human concerns; on others, it misses key critiques or raises issues absent from the expert consensus. Appendix results describe five additional models, but differences in paper availability and run configurations limit comparisons across models.
Limitations. Several caveats temper these conclusions. Our sample comprises 60 social-science papers evaluated by the model, of which 47 have published human evaluation packages for direct comparison; all were specifically selected by The Unjournal for evaluation, not a random draw from the research literature; performance may differ in other fields or on less polished manuscripts. Human evaluations are themselves a noisy reference signal rather than ground truth, with substantial inter-rater variation; agreement with a panel mean need not be bounded by pairwise human agreement. We cannot fully rule out knowledge contamination: while we instruct the model to ignore prior knowledge about authors, institutions, or publication history in the system prompt, the models’ training data may include fragments of these papers or related discussions. Robustness checks with models whose training cutoffs predate the papers, or out-of-time validation on papers entering The Unjournal’s pipeline after model training, would help address this concern. Prompt wording, reference-group interpretation, training, and shrinkage toward typical scores are possible explanations for rating patterns; this design does not separate them. Self-reported interval width alone does not establish calibration. All LLM evaluations are single-run; aggregating across multiple runs or temperature settings could change the picture. Finally, the qualitative human-issue coverage and LLM overlap-rate estimates are themselves LLM-assessed (GPT-5.2 Pro as judge), introducing a further layer of model dependence.
Implications. Even a reasoning-capable model costing several dollars per paper is orders of magnitude cheaper than human expert review. The qualitative gaps we observe—missed critiques, generic issues, and central compression of ratings—argue against full automation of peer review. AI evaluation appears most promising as a supplement: providing fast structured feedback, flagging potential concerns for human reviewers, and enabling systematic comparison across large paper sets that would be infeasible with human effort alone.
Governance and attack surface. As AI review tools move from research prototypes to deployed products, the attack surface expands. Prompt-injection techniques—embedding hidden instructions in a manuscript’s metadata, footnotes, or even white-on-white text—could steer model outputs toward inflated ratings or suppressed critiques. Because our pipeline (and similar commercial services) routes unpublished manuscripts through third-party APIs, confidentiality depends on provider access, retention and training policies, contractual protections, and deployment configuration; encryption in transit alone does not prevent access during inference. Over-reliance on AI scores introduces a further governance risk: if editorial decisions weight model ratings, authors may optimise papers for the model rather than for scientific rigour, creating a Goodhart dynamic. Finally, current evaluations reflect a single model checkpoint; model updates, alignment changes, or fine-tuning can shift ratings in ways that are invisible to users. We recommend that any operational deployment include adversarial red-teaming of prompts, formal confidentiality agreements with API providers, transparent disclosure of AI involvement in review, and periodic re-calibration against fresh human evaluations.
Future directions. All LLM evaluations reported here are single-run; multi-run robustness and prompt-sensitivity analysis would directly address the key open methodological question. Prospective validation should archive the input PDF, model identifier, prompt, and prediction timestamp before human evaluations and publication outcomes become available. Entry into The Unjournal’s pipeline after a training cutoff is insufficient if an earlier manuscript or related review was already public. The qualitative human-issue coverage and LLM overlap-rate estimates are themselves LLM-assessed (GPT-5.2 Pro as judge); human validation of these scores is needed to close the model-dependence loop. Finally, the tier-prediction comparison in Results is a first step toward externally verifiable accuracy measurement: roughly thirty papers with registered predictions from both evaluator types have not yet resolved, and scoring those predictions as publication outcomes accrue can strengthen external validation. Eligibility requires checking that each prediction predates any publicly known outcome or venue cue. A fixed follow-up window and a rule for unresolved papers are needed to avoid selectively scoring early publications.