How has LLM use impacted the language of scientific writing?

In recent years on this website, we have given a pretty critical eye on the use and prevalence of large language models (LLMs), more often realised as generative artificial intelligence (GAI), in academic literature that I tend to surround myself with at work and on here. The main critical points I have made in the two essays on its use in academia largely revolve around its undeclared use, the apparently lax regulations surrounding its use, and the apparent lack of efficacy in detecting and identifying where LLMs or GAI have been used in the writing of a manuscript. I have also covered the possible implications of the use of LLMs and GAI in the academic world, such as the potential undermining of trust in scientific institutions, the erosion of ethics in academia, and potential spread of misinformation and disinformation through scientific journals. Today, however, I want to take a look at how LLMs have impacted the language used in scientific writing, due to its prevalent use, misuse, or even abuse, in the early 2020s.

The primary article of interest is the article by Kobak et al. published in Scientific Advances on 2 July 2025. Here, the authors aimed to assess if there is a footprint left by LLM use in scientific writing, and compare these changes, if there were any, with existing changes made in the academic world, such as the trends in research topics, with the COVID-19 pandemic being a rather recent memory (and still is relevant to this day), and other global events. The approach the researchers used did not require two separate labeled corpora, one for human-written text, and one for LLM-generated text. One salient limitation they highlighted with this methodology was the potential for biases to be introduced, as assumptions about the LLMs used by authors of LLM-labeled text would be required.

In their study, they focused on using PubMed articles in their corpus, assessing millions of article abstracts in English-language publications over 2010 to 2024. Next, they had to define what exactly an ‘excess word’ is. Generally, these are words that contribute to verbosity, or just fluff that plays no significant function in scientific writing. Perhaps a more colloquial term is ‘waffle’ or ‘fluff’. These words would have shown some extent of excess usage in articles published within that time period, and were picked up by the researchers. Some 900 or so of them. These words would be sorted into parts of speech, and into style words and content words. Using this list, frequencies of excess words within the abstracts were computed and were compared between years.

(Kobak et al., 2025)

Perhaps the most salient finding was presented in Figure 1 of their publication, as it showed the trend of frequencies of some example words from 2010 to 2024. They noted that some words exploded in usage in just 2023 to 2024 alone (I am pretty sure you know why), such as the words ‘delves’, ‘crucial’, ‘significant’ and ‘these’, though the frequency for the word ‘important’ plummeted during that period. Comparing these with the word ‘pandemic’ which surged in use in 2020 to 2022, and ‘Ebola’, which surged in 2016 before decreasing over time, these changes attributed to LLM used were likened to how COVID-19 research influenced the language of biomedical reporting at the time. These frequencies exhibited sizable frequency gaps and excess frequency ratios, leading to the finding that there is indeed an influence on the language of scientific writing. In contrast, from the years of 2013 to 2019, the pre-LLM and pre-COVID period assessed, the authors found no words with a significant excess frequency gap, and the word ‘Ebola’ only had an excess frequency ratio of 9.9 in 2015, compared to ‘delves’ with 28.0, though this was surpassed by mentions of ‘zika’ in 2017 with an excess frequency ratio of 40.4.

Separating these excess words into content and style words, it appeared that most excess words with an increased frequency in 2024 alone were style words. Things like ‘delve’, ‘underscore’, ‘crucial’, and ‘showcase’, words which convey style, but pretty much lack substance. This contrasted with the pattern of excess words during the COVID-era of scientific literature, which mainly consisted of content words like ‘respiratory’. The contrast in distribution of excess words by parts of speech is also pretty interesting as well. The authors found that most pre-2024 excess words consisted of nouns, but in 2024, two-thirds of excess words consisted of verbs.

(Kobak et al., 2025)

Putting the frequency gaps based on common and rare words together, the authors estimated that at least 13.5% of abstracts in 2024 contained excess words (frequency gap 0.135) that could be attributed to some form of LLM usage. This estimate was only a lower bound, since some abstracts that have been through some LLM usage lack some excess words captured by the researchers in this study, and hence this would be an underestimation. Comparing this with the frequency gap of 0.069 for excess words in abstracts published during the COVID-era, this is pretty concerning.

Further findings concerned the distribution of excess word prevalence, by frequency gaps based on common and rare words, across fields of research covered by PubMed, countries of origin of these articles, and journals that publish these articles. For instance, non-anglosphere countries, most notably Taiwan, Iran, and Thailand tended to have a higher frequency gap compared to anglosphere countries, and abstracts related to healthcare, environment, and computation tended to have substantially higher frequency gaps compared to abstracts related to ecology. More concerningly, while abstracts in more reputable major journals like Nature, Science, and Cell tended to have lower frequency gaps (0,07), more controversial journals such as MDPI had drastically higher frequency gaps (0.21).

To address this heterogeneity in excess word usage across research fields, countries of origin, and journals, the authors proposed attitudes and general reception of LLMs in individual research fields, and the use of LLMs by non-native English speakers to write something that would sound more ‘natural’. And I am sure that I have criticised this before, but some journals tend to have a more relaxed review process than others, and hence more lax journals would allow low-effort articles churned out by LLMs to pass through the review process than journals with a more stringent protocol targeting low-effort or AI-generated scientific writing.

Now, of course, there are some limitations which might suggest how much of an underestimation LLM prevalence is in abstracts written in 2024 and beyond. Some might be more aware of LLM signatures left in the generated text, and would try to censor out these signature words, thus, it could be the case that anglophone and non-anglophone countries used LLMs to similar extents in the writing process, but native English speakers are able to hide this involvement in their writing than their non-native English speaking counterparts.

Another limitation is the lack of distinction between direct LLM word usage, and word usage that human writers adopted because of LLMs. This one is a pretty difficult thing to dissect, but it could be very well the case that a shift in human writing styles is happening, no matter how slowly, due to LLM influence. However, there is some evidence to suggest that LLM use has influenced how we speak, highlighted in a separate study by a different research group which we will cover soon.

So, having summarised this article, what do I think about it?

The study by Kobak et al. (2025) is one of the few, and one of the first studies providing empirical evidence for some form of shift or change in the distribution of excess words, and the use of excess words in scientific writing, more precisely, in fields related to the life sciences, medicine, and healthcare. Verbs overtook nouns as the dominant part of speech found in excess words, and style words dominated the type of excess words in 2024 by a landslide.

I am concerned with the implications that such findings have. In the subgroup analyses by journals, I would question if this is necessarily an indictment on the review processes in the journals that feature higher frequency gaps, or if they were slower or softer to act on LLM usage compared to their more reputable counterparts. Some journals have imposed bans on LLM use in the writing process, which is great, but other journals might prioritise publication output more than anything else (like a paper mill), which could end up publishing low-effort generated articles that undermine or mislead the scientific community. After all, LLMs can hallucinate, plagiarise, and generate fake references or evidence in their output. And so, paper mills could very well poison scientific writing by allowing these articles to pass.

Additionally, there is the future of scientific writing to think about. While we generally can pick out generated articles today, given the prevalence of style words that LLMs absolutely love to populate in abstracts and full texts, there is the question if these style words would eventually die out so LLM usage is less detectable, or if human writing would be influenced by LLM excess words, and excess words written by humans would adopt similar patterns to those of LLMs today.

All in all, this article has shown the current situation in PubMed articles, and I would be interested in seeing if such an approach could be extended to other forms of media, such as news and social media, to track the patterns of change in how LLMs affect the linguistic landscape of written and spoken media. To conclude, the authors put out a rather tongue in cheek closing statement which almost mirrors the way LLMs write, and honestly, there is perhaps no better irony than this:

Leave a comment