In late-2025, denizens of the Internet noticed a pattern in how punctuations are used by large language models (LLMs) when generating their responses to a given prompt. Putting aside the verbosity, pervasive use of style words, and very ironically here, tendency to use the rule of three, users have pointed out that the ‘quickest’ way to identify an AI-generated text is to look out for this punctuation mark called the em-dash. And so today, for a light intermission in the series of AI use in languages and linguistics, we will take a deeper look into this punctuation mark that has been proposed to be a fingerprint of AI-generated text.
The em-dash, not to be confused with the shorter en-dash and hyphen, is a punctuation mark that is rendered as the length of the font’s height. As a result, the length of the em-dash varies by the font used, though the proportion of the em-dash length to the font height is uniform. The en-dash in contrast, is about half the length of the em-dash, while the hyphen is a little shorter, and is positioned differently in height from the dashes.
Some of the most common functions of the em-dash include the expression of interruption, in which there is an abrupt change or cutoff of thought or speech. Long pauses or breaks may also be expressed using the em-dash, and it gives a different literary interpretation from the ellipsis. It also serves to replace other punctuation marks in written texts, though it seems to be a stylistic preference, such as the parentheses and the colon. For em-dashes used like a colon, this also allows the inversion of the order in equivalence clauses. Examples include the one below:
- The food, which was delicious, reminded me of home.
- The food—which was delicious—reminded me of home.
- The food (which was delicious) reminded me of home.
- These are the main spices used in curries: turmeric, cumin, garam masala, cardamom, mustard seeds.
- Turmeric, cumin, garam masala, cardamom, mustard seeds—these are the main spices used in curries.
Sometimes, it even takes the place of a semicolon, something that may be used to link two statements, where a period to separate them is too strong to drive a certain point, but a comma might be too weak or inappropriate.
Quotations, redactions, and repetition may also be functions served by the em-dash as well, though perhaps not as particularly well-known or well-appreciated compared to the expression of interruption.
So, given these human-attributed functions of the em-dash for the past centuries, how does this link to AI-generated texts today?
LLMs are trained on how we write and express ourselves, and try to replicate these patterns in response to a given prompt. It is not inherently intelligent, it just gives us something it guesses is acceptable, given the patterns of speech and prompt. As a result, if an LLM is prompted to produce a list, it would form its response by observing patterns of elements of lists, the semantics and connotations desired by the list, and so on. If it sees an em-dash preferentially used in the training data, it will think that given this preference of punctuation mark, it should use the em-dash there as well. Past printed literature is full of examples where em-dashes would have been used to express ideas, and would have been captured, ethically or unethically speaking, by LLMs for training.
LLMs such as ChatGPT also render em-dashes without spaces flanking the punctuation mark, but humans do put spaces when they want to write dashes. Flanking hyphens or double hyphens with spaces can create dashes in some text editors, which might render them as the shorter en-dash instead. Perhaps this difference is what people have been pointing out as the “ChatGPT dash”, and it ultimately boils down to this very orthographical difference.
Does this mean that the em-dash is a strong indicator of AI-generated text today?
Well, not necessarily. As noted earlier, the em-dash still serves several literary and orthographical functions in human-written texts. Interrupted speech being one of — see what I did there? Additionally, prompts may affect what the rendered response would be, orthography and usage of certain words alike. These can mask, albeit not completely, some fingerprints left behind which may indicate AI-generated text. Thus, the em-dash alone is not particularly a strong indicator of AI-generated text, and would therefore entail further scrutiny of the text’s contents.
The em-dash, however, is slightly more troublesome to input into various text editors compared to the shorter en-dash. This discrepancy in user burden could have contributed to the user preference for the en-dash to express the aforementioned functions, leaving the em-dash a rather telling signature of generated texts. For instance, the en-dash is typically formed from two hyphens (–), while the em-dash is formed using three hyphens (—), at least on most text editors. User intuition would be inclined to just use two hyphens, creating the en-dash. Additionally, the em-dash may be treated as a special character, either requiring pulling up a selection menu for it, or using the Alt+0151 shortcut, which is substantially more burdensome than the en-dash. Some text editors might also render the triple hyphen sequence as an en-dash and a hyphen rather than an em-dash. This might substantiate some arguments that the frequent feature of the em-dash is a hallmark of AI-generated text.
Nevertheless, we also cannot rule out the fact that LLMs are trained on how we write, speak, and express ourselves through various means, our preference and use of the em-dash being one of them. This would create preferences for which LLMs might want to churn out their responses, and hence producing something that it thinks a human would expect. Then again, preferences have resulted in AI-generated responses featuring more style words, as well as signatures like ‘delve’. Who knows, perhaps over time, should these preferences continue, the em-dash could very well develop into a preferred punctuation mark by LLMs instead of several alternatives which serve the same function. Additionally, as we have seen with human word preferences under the influence of LLMs, an aversion from using orthographical and lexical stuff preferred by LLMs could develop to try to make human-written texts more distinct from what an LLM churns out. Alternatively, there could be a shift in the opposite direction, in which human-written text, under the influence and prolonged exposure to LLM usage, would subconsciously converge with the styles and preferences adopted by LLMs in AI-generated text.
As of now, much of the discourse surrounding the em-dash and LLMs and AI appear to center around which users have experienced and seen; there have been no empirical studies surrounding the use of em-dash by humans and generative AI, but who knows, over time, we would have built large enough corpora to answer this very question.