Over the course of human history, several paradigm shifts have occurred to change how we communicate with one another. From the invention of writing to the printing press, radio, and later, the Internet, the patterns in which we communicate have developed in punctuated spurts, such as the development of Internet and messaging lingo, with the latter driven by character limits in text messages back then. This decade, we are dealing with yet another possible paradigm shift, with the advent of the large language model (LLM). Trained on an extensive wealth of corpora and data, which ethics remain pretty controversial, LLMs often display a certain pattern in how they phrase responses, to which humans, through growing and prolonged interaction with LLMs, could grow receptive to how LLMs ‘speak’. Previously, we have taken a look at the impact of LLMs in academic and scientific writing, but it is time to broaden the scope a little bit, and ask ourselves if LLMs has impacted human speech patterns. Today, we will take a look at this preprint paper by Yakura et al. (2025) found on arXiv, which goes into this very topic.
To clarify, arXiv is a pre-print catalogue by Cornell University containing academic papers, abstracts, and conference papers covering various research topics. The papers hosted on arXiv are not peer reviewed, but go through another moderation process. As such, we cannot say that these papers are journal publications, as arXiv is not a journal, and neither do they publish material, but more rather, archive them.
With this out of the way, let us look at how they assessed for the presence of such linguistic influences.
The premise of the study centers around word preferences, and how these patterns occur in LLMs compared to humans. LLMs like ChatGPT are trained on multiple media types spanning numerous genres and contexts, but the reinforcement learning process and fine-tuning are not quite open nor transparent, as they are typically proprietary methods. OpenAI, the company behind ChatGPT, has asserted that the LLM generally prefers politeness, neutrality, and avoidance of conflict, portraying itself as being ‘professional’. Yet, this behaviour generally is distinguishable from normal human communication. Hence, we would expect ChatGPT to prefer words that reflect these traits, shaped by statistical learning of human discourse. However, inferring from the findings of the study we covered last week, ChatGPT tends to favour fluff words like ‘delve’, which sheds more light on its lexicographic biases. Humans also display preferences for certain words, but these are shaped by culture, geography, social contexts, and even historical contexts. For instance, the word ‘wee’ may be preferred in some anglophone communities to mean ‘urination’, while in Scotland, it would be preferred when referring something in the diminutive, like a wee lad (young boy).
To assess these word preferences quantitatively, word frequencies could be examined. It has been previously used to assess trends in word preferences in media, such as news discourse, as well as how colloquialisms have changed over time. It is this comparison which could allow us to understand the extent to which interaction with LLMs could impact human word preferences.
A couple of data sources were used in the study. The first was an audio transcript of academic videos from educational entities, where nearly 3 million YouTube videos were compiled. The second was a collection of podcasts pertaining to business, education, religion, spirituality, science and technology, and sports, from 2017 to 2024. This totaled 4 million podcast series, from which the researchers randomly selected 6000 podcasts from each quarter and from each category to include in their sample.
With this abundance of data, the main limitation here would be the computational resources the researchers had at hand to process all of these podcasts and videos. As such, videos and podcasts shorter than 20 minutes and 15 minutes respectively, and longer than 3 hours and 5.5 hours respectively were excluded from the study. Furthermore, these materials were further cut to include only conversations, as the main objective of the study primarily concerned spontaneous language use, something that is done through conversing. Of course, this exclusion would affect each content category differently, as religion and spirituality was reported to be more likely to contain monologues, and hence would demonstrate a sharp drop-off in included content for the study. The audio from these podcasts and videos included in the study would be transcribed by the researchers to avoid potential bias from automated transcription features from platforms such as YouTube.
Using this corpus, the words from the transcripts were filtered to exclude words that do not carry semantic meaning, such as some grammatical words like a and the, and filler words that do not carry meaningful information. Variations in word forms such as by conjugation (like underscore, underscores, underscored) were reduced to the main root form to facilitate subsequent analyses more efficiently.
The main method of analyses, like the previous study we covered, compared word frequencies between human-generated and LLM-generated texts. To do this, contrastive datasets were constructed that made use of prompting LLMs to improve, polish, or rephrase texts drawn from various publicly available sources like arXiv abstracts. This process involved the use of different LLM versions, to better understand trends in word preferences by LLMs over their version history. Comparison would be done using log-odds ratios, with positive log-odds ratios for a word indicating that there was a higher usage by LLMs, and negative log-odds ratios indicating that a given word tended to be used by humans.
Now, this analysis method would tell us a correlation, or a tendency, but does not imply a causative effect that LLMs have or have not affected how humans converse. As such, the researchers aimed to estimate the causative effect through the use of a synthetic control, based on the assumption that words sharing similar pre-release usage patterns would have continued exhibiting comparable patterns in the absence of the release.
It was found that style words like delve exhibited the strongest preferences by LLMs like ChatGPT, with other words like boast, comprehend, meticulous, and swift showing strong preferences by LLMs from data obtained from YouTube talks. This finding is generally not surprising, as it corroborates with the findings of the previous study we discussed, though that used a corpus of published studies in the PubMed database. However, this study has also revealed that YouTube videos and podcasts demonstrated an increase in the use of LLM-preferred words. Possible explanations for this increase may include video presenters reading off LLM-generated scripts, or some other machine-to-human interaction.
The main point of interest is the increase in LLM-preferred word usage in podcasts, which usually feature unscripted speech. The increase did not affect all disciplines uniformly though, as non-statistically significant effects were reported in podcasts related to sports, religion, and spirituality, but significant effects were found in STEM-related topics, business, and education. It is apparent that more academically-aligned topics tended to display a stronger adoption of LLM-preferred words, likely due to the increased likelihood of exposure to LLM-generated content in academia. These effects would have bled into informal discourse, such as those talking about sports.
Overall, this study illustrates some salient patterns of linguistic shifts as a result of LLM use. These shifts are not affected by whether or not content was scripted, but more rather the speech topics, and the likely pattern of how LLM-preferred words enter and get embedded into human speech. The precise mechanisms of these linguistic shifts are not really known or agreed upon though, as several possible processes have been brought up as mechanisms. This includes direct imitation, cognitive ease, and the integration of LLM-preferred words into human thought processes. However, these mechanisms would entail the involvement of an internalisation of LLM-driven linguistic patterns into the human thought process, lending credibility that it could be a mix of these processes that contribute to this LLM-driven linguistic shift in humans.
There are several concerning implications in how human culture would evolve under the influence of LLMs and generative AI. As LLM-preferred words enter human use, it would eventually be possible that, LLM-preferred words would be embedded in the word preferences in humans who might never have interacted with an LLM before. LLMs would thus hold an immense leverage over how human culture evolves, and how we humans may communicate in everyday discourse.
A dangerous implication that the researchers raised was the possibility of cultural homogenization by the use of LLMs, as LLMs may exhibit preferences to some certain cultural traits, and would prompt the erosion of existing cultural traits already held by humans. As humans and LLMs interact, this would prompt a self-perpetuating cycle of increasing homogenization, as some patterns in speech and writing become monopolized. This would also carry concerning implications regarding the social benefits of LLMs, how social identity and social dynamics may change.
Though in a pre-print state as of the time of writing, I think that this study contributes a lot to our understanding of how we humans communicate as interactions between LLMs and humans become increasingly common. With an extensive corpus encompassing multiple domains of human speech, this study presented large-scale empirical evidence that the influence of LLMs have penetrated beyond written content, and into human speech, adding onto the existing knowledge that LLM use has been occurring in scientific writing. I think that this study, alongside the previous one we looked at, has triggered me to do a retrospective on how I have been communicating with others, and perhaps even how I write here on The Language Closet. My last interaction with an LLM may have occurred several years ago, but I have been interacting with people who might have interacted with an LLM more recently, and may have incorporated LLM-preferred words into their speech. Have I started communicating like an LLM, or am I still human? That is the question I asked myself from time to time.
Further reading
Yakura, H., Lopez-Lopez, E., Brinkmann, L., Serna, I., Gupta, P., Soraperra, I., and Rahwan, I. (2025) ‘Empirical evidence of Large Language Model’s influence on human spoken communication’, arXiv, 2409.01754. https://doi.org/10.48550/arXiv.2409.01754.