Read enough AI-generated text and you start spotting the same habits. Neat headings. Three-item lists. Phrases like “the key is.” And, of course, em dashes everywhere.
The em dash is this mark:
The answer sounds simple — until you look at how the model was trained.
Humans have used it for a long time, so it is not an invention of ChatGPT or any other AI system. But models do seem unusually fond of it. The reason is less mysterious than it looks: em dashes were already common in the kind of polished writing models learned from, and later training encouraged the same smooth, explanatory style.
In other words, AI did not decide that em dashes were cool. It picked up a human writing habit and then repeated it at a scale humans never could.
Models learned from polished writing
During pre-training, a language model reads a huge amount of text and learns which tokens tend to follow other tokens. It does not memorize a rule saying, “Add an em dash when the paragraph needs some flair.” It learns patterns from examples.
Em dashes show up often in edited prose: essays, journalism, books, reference material, and carefully written explanations. They are useful because they can introduce a contrast, add a qualification, or squeeze an aside into a sentence without stopping the flow.
Those are also things an AI assistant does constantly.
A typical answer might begin with a broad point and then add a caveat:
The tool is easy to use — but it still needs careful review.
Or it might set up a simple idea and then make it sound more precise:
The model is not checking the truth — it is predicting likely text.
These sentence shapes appear again and again in explanatory writing. If training data contains plenty of polished examples that use em dashes this way, the model learns that the mark is a natural continuation.
The exact mix of data used to train major models is usually private, so nobody outside the companies can give a perfect breakdown. Still, the basic mechanism is clear. Models reproduce patterns from the text selected for training, and high-quality edited text contains more typographic punctuation than a random collection of texts, comments, and hurried emails.
Post-training made the style stronger
Pre-training gives a model many possible ways to write. Post-training pushes it toward the ways people prefer.
Chat models are shown examples of good answers and receive feedback about which responses are more helpful, clear, professional, or complete. Reviewers are not necessarily voting for em dashes directly. They are rewarding the larger style that often includes them.
Compare these two versions:
The model predicts likely text. It does not verify every claim.
The model predicts likely text — it does not verify every claim.
The first is plain and direct. The second connects the warning to the main point and sounds a little more polished. Neither is automatically better, but the second has the kind of smooth rhythm that often does well in examples of “good assistant writing.”
Repeat that preference across thousands of training examples and ratings, and the model starts reaching for the same structure more often. The punctuation comes along for the ride.
This can also turn into a feedback loop. A tuned model produces polished answers. Some of those answers may help create synthetic training examples, evaluation sets, or demonstrations for newer models. If the same writing habits keep appearing in those materials, later systems learn them too.
That does not mean every model copied one original em-dash-loving chatbot. It means the industry often trains assistants toward a similar idea of a good answer, so the same habits can show up across different systems.
Assistants have a narrow default voice
Human writing is all over the place. Some people write in fragments. Some love semicolons. Some use five-word sentences. Others take half a page to reach the period.
A general-purpose assistant cannot begin with that much personality. Its default voice has to work for a student asking about photosynthesis, a developer debugging code, and someone drafting an email to their landlord.
The result is usually a safe middle: friendly, organized, confident, and slightly formal. It explains the main point, adds a caveat, and tries to keep everything flowing.
The em dash is almost perfect for that voice. It is less formal than a semicolon, more dramatic than a comma, and smoother than starting a new sentence. It lets the assistant sound conversational while still looking edited.
Because millions of answers use roughly the same default voice, small habits become very visible. One human writer using three em dashes in an essay does not look unusual. Thousands of chatbot answers using the same sentence rhythm make it feel like an AI trademark.
It is not really a trademark, though. It is a human punctuation pattern that models learned from polished prose, post-training rewarded as part of a clear explanatory style, and the default assistant voice repeated until everyone noticed.