Select Page


    Key Takeaways

    A BERT-based language model is now trawling the cancer literature the way email filters comb for junk. In a study published in The BMJ, Queensland University of Technology biostatistician Adrian Barnett and collaborators trained the system on 2,202 retracted paper-mill papers, then ran it across 2.6 million cancer studies from 1999 to 2024, flagging 261,245 papers, or 9.87%, for writing patterns that resemble suspected fabrications. The share of flagged papers climbed from roughly 1% in the early 2000s to more than 16% by 2022, a surge that turns peer review into an arms race between industrialized fakery and what Barnett calls a “scientific spam filter.” Three journals are already testing the screening tech, as editors brace for what gets through when the templates evolve.

    A staggering problem uncovered by AI

    One of the quiet anxieties in US science right now is that the literature itself might be getting harder to trust at scale. A new paper in The BMJ suggests that worry is not paranoia. Researchers screened an ocean of cancer papers and found a surprisingly large slice that “looks like” work that later got pulled for suspected fabrication.

    In the study titled “Machine learning based screening of potential paper mill publications in cancer research,” an international team led by Professor Adrian Barnett of Queensland University of Technology (QUT) analyzed 2.6 million cancer papers published from 1999 through 2024. The model flagged 261,245 papers, or 9.87%, for writing patterns similar to already retracted work linked to suspected fabrication.

    How “paper mills” became an industrial problem

    Barnett’s framing is blunt: “Paper mills are companies that sell fake or low-quality scientific studies. They are producing ‘research’ on an industrial scale, and our findings suggest the problem in cancer research is far larger than most people realised,” he said.

    The pattern also appears to be worsening over time. The proportion of flagged papers rose from roughly 1% of annual cancer research output in the early 2000s to more than 16% by 2022. And the concentration was not uniform: gastric cancer papers were flagged at about 22%, bone cancer at about 21%, and liver cancer at about 20%.

    A “scientific spam filter” built on BERT

    To do the screening, the team trained a BERT-based language model on 2,202 retracted papers cataloged in the Retraction Watch database. After validation against independent expert datasets, the system reached 91% accuracy in identifying suspicious papers that matched the known “template” style.

    Barnett also urged caution on interpretation. “If it’s actually ten percent, we don’t really know. It could actually be more because we’re just detecting one particular kind of template,” he said, warning that more sophisticated templates could slip through.

    Why this matters in the real world

    The stakes are not abstract. “Cancer research influences clinical trials, drug development and patient care. If fabricated studies make their way into the evidence base, they can mislead real scientists and ultimately slow progress for patients,” Barnett said.

    There are early signs of operational adoption: 3 scientific journals are already testing the BERT screening technology in their editorial process. That is a small start, but it hints at where peer review may be heading: routine, automated triage before human experts ever see a manuscript.



    Source link

    Translate »