A new automated tool exposes how easily human editors miss shifted endpoints in published medical research.
How do we know a clinical trial actually worked? Often, we do not, because researchers quietly swap their primary goals mid-study to make disappointing data look like a win. For years, detecting this “outcome switching” required tedious, manual audits by human experts.
Now, a validation study of an AI workflow called RegCheck suggests we can automate this policing for the price of a cup of coffee. This is not just about saving time. It challenges the assumption that peer review at elite medical journals is a reliable shield against spin. If a cheap algorithm can spot hidden discrepancies better than human editors, the traditional publishing model has a major quality control problem.
The audit results
Researchers tested RegCheck on 62 clinical trials originally audited by the COMPare Trials project, sampled from five high-impact general medical journals. The tool compared published trial reports against their original registrations. The results show that algorithms can match, and sometimes exceed, human diligence.
The system’s performance metrics include:
- A 91.2% outcome extraction recall relative to human experts.
- An 83.6% accuracy rate in classifying outcomes as primary, secondary, or non-prespecified.
- An initial misreporting detection accuracy of 85.6%.
- A revised misreporting detection accuracy of 94.8% after resolving discrepancies.
- A mean operating cost of just $5.94 per paper.
During the discrepancy resolution, researchers found that the AI’s judgment was frequently superior to the original human auditors. The tool caught valid reporting issues that human experts had overlooked.
The cost of silence
This cheap automation could disrupt a long-standing crisis in medical publishing. Selective reporting and outcome switching are not new; they are well-documented drivers of publication bias that distort the clinical evidence base. A classic review on dissemination and publication biases highlighted how these practices skew what treatments doctors believe are effective. Yet, journals rarely dedicate the staff needed to check these registrations before publication.
By lowering the cost of auditing to under six dollars, this tool removes the financial excuse for lazy peer review. It shifts the burden of proof back to the authors and editors before a paper ever goes to print.
Not a perfect shield
We must view these results with some caution. The validation cohort was small, limited to just 62 papers from top-tier journals. These journals generally have better formatting and more complete registries than lower-tier publications, meaning the tool might struggle with messier data. Furthermore, the tool still relies on human adjudication to resolve its trickiest discrepancies and achieve that high 94.8% accuracy rate. It is an assistant, not an independent judge.
Even with these limits, the implications are clear. Automated compliance checking should become a mandatory gatekeeper in scientific publishing.
Source: medRxiv
