Contemporary AI lacks the imagination to diverge or negate in science

🔘 About this critique

This is not peer review. No proof has been checked, no experiment reproduced, no result validated. The audited version is an arXiv preprint and is not presented here as a peer-reviewed publication.

What is audited is correspondence: the distance between what the article claims and what it presents as grounds for it. Where something could not be checked, that is stated.

The recommendation is to read the paper.

Critique

Audited versionarXiv:2606.08251v2 · submitted 9 June 2026, 03:31 UTC · licence CC BY 4.0
Incorporated materialgithub.com/listar2000/science-reward-model · commit a64f62e90c9aecfab8085b4889a5939d7d4c1bbc · consulted 3 August 2026 solely to verify the Data and Code Availability statement. No code was executed.

Unit 1 — What the abstract asserts and what the body reports

Article data. The abstract states that no model class spontaneously proposes null hypotheses, a move humans make more freely. The body, introducing Figure 2d, states that humans formulate such hypotheses infrequently, that every LLM in the panel formulates them less often, and that even agentic deep-research systems with live web search rarely articulate one.

Reading. Absolute absence and relative deficit against an already low human baseline are not the same finding. The second is what the measurement supports; the first is what the abstract states.

Critical conclusion. The distance is internal to the document. The abstract converts a measured relative deficit into a categorical formulation. The title does not repeat the zero-frequency claim verbatim, but amplifies the same direction of interpretation.

Claim typeOwn result
ExposureHigh. The abstract’s categorical wording underwrites the title’s broader claim about negation.

Unit 2 — The scope of the claim and the scope of the corpus

Article data. The corpus is 121,640 post-2023 preprints from six platforms: bioRxiv 68 per cent, medRxiv 20 per cent, ChemRxiv 3 per cent, and PsyArXiv, EdArXiv and SocArXiv jointly 9 per cent. arXiv was excluded deliberately, and the exclusion is argued in the text.

Reading. Biology and medicine account for 88 per cent of the material. Physics, mathematics and engineering are not sampled as declared corpus domains, and the formal sciences fall outside the study’s stated empirical-paper design. The measured domain is hypothesis generation from recent empirical preprints, predominantly in biology and medicine, with substantially smaller chemistry and social-science strata.

Critical conclusion. The title speaks of science without carrying the corpus restriction. The composition is published, so a reader can establish this; the formulation does not.

Claim typeOwn result
ExposureHigh. The scope of the title exceeds the scope of the sample.

Unit 3 — Participation reported without a denominator

Article data. Five hypotheses per paper were sent to the original author, in the wording of the supplementary material, as long as the authors could access their email. The number of authors contacted is not reported. The main text states that 6,749 scientists participated and returned 25,139 rating sets. The supplementary material states that 6,749 scientists participated, some contributing only partially completed responses — for example, before the understanding test — and that the 25,139 valid evaluations ultimately came from 5,259 participants. It also states that data were excluded from authors who did not fully recall their paper or who consented only partially.

Reading. Without the number of invitations, no participation rate can be computed by the reader. The figure that appears in the abstract is the larger one; the figure that carries the analysis appears only in the supplement.

Critical conclusion. The article reports no non-response analysis, respondent–non-respondent comparison or participation weighting. This absence does not by itself invalidate comparisons within the respondent sample. What the article does not demonstrate is that this self-selected sample represents scientists at large.

Claim typeOwn result
ExposureHigh. All human-side inferences depend on this respondent sample.

Unit 4 — The negation detector was trained on synthetic examples generated by a model in the evaluated panel

Article data. The supplementary material describes the null-hypothesis detector: Treebank tokenisation, stopwords deliberately retained so that negations such as no effect and not significant survive, TF-IDF vectorisation, and an ensemble of four classifiers averaged at a 0.5 threshold. The training set is 1,000 null and 1,000 non-null instances generated by o3-mini. Two annotators independently checked 100 of each, with full agreement reported. Cross-validated accuracy is reported at 99.5 per cent. o3-mini is listed among the reasoning models of the evaluated panel.

Reading. One evaluated model generated the synthetic examples used to train the detector. The reported accuracy comes from cross-validation within that synthetic distribution, and the 200-item human check is not an independent test on the scientific hypotheses the panel actually produced. The instrument is explicitly designed around lexical signals of negation; its sensitivity to equivalence, non-inferiority, conditional absence of effect or implicit scientific negation is not established.

Critical conclusion. The analysis provides evidence about explicit lexical null formulation. It does not validate the detector against scientific negation in its full logical and disciplinary range.

Claim typeOwn result
ExposureHigh. The detector carries the paper’s central finding on negation.

Unit 5 — The verdict is recorded; the reason is not

Article data. The full questionnaire appears in the supplementary material. Each hypothesis is rated on three 1-to-9 scales and adoption on a 1-to-7 scale. The rubric is published with verbal anchors at 9, 5 and 1, and participants must pass a comprehension check before rating. No free-text field is attached to any rating. A single optional box appears at the end of the form, after the administrative questions, unlinked to any hypothesis; no analysis of this open-response field is reported in the article. The discussion states that what experts contribute is not only a verdict on individual ideas but the collective construction of the criteria by which ideas come to be judged.

Reading. The rating criteria are traceable: a reader knows what raters were asked to measure. The score is also recorded; what is not recorded is the hypothesis-specific reasoning by which the rater reached it. The reward model is therefore trained on numerical ratings and derived pairwise preferences that do not encode the reason for each judgment.

Critical conclusion. The discussion identifies collective criterion-building as the distinctive contribution of expert communities, but the survey does not elicit the hypothesis-specific reasoning through which each score was reached.

Claim typeOwn result
ExposureHigh. The survey ratings supply the human preference signal used to train the reward model.

Unit 6 — A ground truth the article itself shows to be situated

Article data. The article treats the 25,139 expert ratings as ground truth for testing automated proxies. The same analysis reports that raters prefer ideas resembling their own, that probability weighs more than novelty and feasibility, that seniority reduces adoption, and that fields differ. The reward model is benchmarked against a floor of 61.0 ± 0.1 per cent, obtained from 26,731 OpenReview submissions to 46 conferences by constructing within-conference comparisons between different papers and measuring how often a reviewer in one position agreed with the direction of another position’s preference.

Reading. A model trained on these preferences learns the preference structure of this population, including the biases the article reports. The 61 per cent reference is reconstructed cross-paper directional agreement, not inter-rater reliability on the same hypotheses under the same rubric. Exceeding it shows that the reward model predicts held-out recorded preferences more accurately than that reconstructed floor. It does not establish task-equivalent inter-rater reliability or shared scientific reasoning.

Critical conclusion. The article is explicit that distilling current taste is not distilling the practice by which taste evolves. The framing of the ratings as ground truth sits at some distance from that acknowledgement.

Claim typeOwn result
ExposureMedium–High. The floor supports the paper’s claim that the reward model closes the gap with independent reviewer consistency.

Unit 7 — Declared as released, but only partly available

Article data. The Data and Code Availability statement says that aggregated, de-identified ratings, trained reward-model checkpoints and pipeline code have been released at the project repository, and that individual identifiers will not be released under the governing protocol. The words naming the project repository carry an active link to the authors’ GitHub repository.

Incorporated material. At the audited snapshot, the dataset-building, training and analysis code is present. The data documentation states that the training file link is redacted for now and will be made available upon manuscript acceptance. The checkpoint section carries a Hugging Face badge resolving to the site root, an inline placeholder marked as a link to be released, and two literal editorial notes asking for the placeholder to be replaced with the real link.

Reading. The repository is locatable and the code is available. The documentary gap lies between a release statement written in the completed tense and the present accessibility of the data and the trained checkpoints.

Critical conclusion. At the audited snapshot, two of the three resource classes described as already released were not accessible. This limits independent reproduction of the reward-model results. It does not demonstrate any intention to withhold them.

Claim typeQuotation / incorporated material
ExposureHigh. The discrepancy is directly checkable and bears on reproducibility.

What this leaves

The paper documents not one isolated failure of AI, but a coupled circuit: published literature shaped by selection, models that recombine it, authors whose preferences are situated, and a reward model trained on the resulting scores. Some stages widen the candidate space and others filter or compress it. The study makes this circuit visible at unusual scale. What it leaves unresolved is that the ratings used to train the reward model preserve the decision, but not the hypothesis-specific reason by which that decision was reached.

This article is recommended here because its scale, methodological disclosure and publication of the full survey make it possible to examine not only what it concludes about scientific AI, but also what its own apparatus can and cannot establish.


r0: 10f953633b3db3b4dc07e63a493e9b66

🔘 Paper page: https://doi.org/10.48550/arXiv.2606.08251

Abstract

Bold projections that artificial intelligence will accelerate scientific discovery have raced ahead of evidence from working scientists, and the field still lacks large-scale, scientist-in-the-loop tests of these claims. Here we mount the largest such evaluation to date and map what AI cannot yet do for science. We invited authors of 121,640 recent preprints across biology, medicine, chemistry, and the social sciences to judge ideas that large language models (LLMs) generated from the context and puzzles of their own papers. 6,749 scientists returned 25,139 sets of ratings on novelty, empirical feasibility, probability of being true, and favorability of adoption. Three patterns emerge. First, non-reasoning LLMs collapse into a narrow «hivemind» of similar ideas; reasoning models roam a wider hypothesis space, yet no model class spontaneously proposes null hypotheses — a move humans make more freely. Second, scientists reward ideas that resemble their own and prize probability over novelty, though social scientists tolerate risk more readily than life scientists. Senior social scientists are the harshest critics, and their skepticism is well-earned: LLMs falter most in pluralistic fields like the social sciences that demand context-aware interpretation and evolving theories. Third, automated evaluators on which the community currently relies — LLM-as-a-judge, artificial metrics, and even state-of-the-art (SOTA) models — agree only weakly with expert judgment, and retrieval augmentation and scientist persona prompting yield only marginal gains. A Qwen3-14B reward model we post-trained on human ratings captures field taste nuances, beats SOTA models by up to 27%, and closes the gap to the inter-rater consistency of independent peer reviewers. For all the hype, today’s scientific AI still represents a collaborator whose imagination, outputs and judgment benefit from human grounding.

Authors

Bao, H., Wu, S., Liu, X. et al

Liked this post? Follow this blog to get more.