Sitemap

When LLMs Write Our Papers

Four writing issues I notice as a reviewer — and how to fix them

--

Press enter or click to view image in full size
A 2D flat-shaded illustration of a paper ship with text on it sitting dangerously unstable on top of a rolling wave.
The current wave of LLM use can send our writing off course, from the swirl of hyped language to the cliff of rocky citations.

While reviewing for the CHI conference, I’ve recently found myself commenting on issues I had rarely mentioned in reviews before. They appeared again in the next paper in my AC stack, and the next one after that. When I shared these observations with colleagues, they described strikingly similar experiences.

Although anecdotal, these patterns motivated me to write this article on what I call “LLM writing smells.“ These problems likely arise from using AI for literature research and drafting.

Let’s look at these issues and how to fix them.

Marketing Language

The problem

One new writing issue I’ve noticed is the increased use of feel-good adjectives and adverbs, especially when framing the work and reporting and discussing results. Words like “seamless”, “elegant”, “vividly”, and “dramatically” often appear meaningless and likely slip in when an LLM is used for drafting or polish. For example, what’s the real difference between designing a “seamless workflow” and a “workflow” in your prototype?

Why this is problematic

Such words colour the writing subjective and invite critical questions because they assign qualities to the work that were never evaluated or even intended. For example, if you describe your interaction design as “elegant”, as a reviewer I expect you to provide evidence or at least a rationale for that claim. Yet spelling out a justification will most likely just bloat the paper because it is rarely relevant to research questions in HCI whether a system is “elegant” or not.

How to fix it

Don’t put in adjectives and adverbs that “hype up” your work — or any at all. Let the work speak for itself by evaluating what is relevant. More broadly, don’t think that your paper will only be convincing if your system wins in every way possible. Being better at A but not B is also a fine insight, especially if you discuss why that is the case.

Performative Related Work Section

The problem

Yes, that looks like a related work section that covers relevant areas. It’s not good though. It cites something, but far from the most relevant or fitting work, likely because an LLM was asked for sources. For example, it cites a 25-year-old paper that once made a point in passing. Indeed, that work includes a short string of letters that supports your argument. However, it is a far-fetched choice for support, when more fitting work exists, focused directly on your technology or research question. Another case is citing work so distant from your own domain that even a non-expert reviewer can tell it’s off. For example, referencing an insight about perceived control from a smart home study while your research concerns touchscreen keyboards.

Why this is problematic

Related work provides a foundation for your work and for the reader’s understanding of it. If your references come from “deep research” with an LLM on the day of the deadline, your related work section signals to reviewers that you do not know the literature well — or worse, that you do not care about the work of the community to which you are submitting your manuscript. Reviewers in that field easily spot this, yet still have to spend time explaining why the citations don’t fit. This is a waste of time for everyone involved.

How to fix it

This is not about memorising the entire literature but about how you think about this part of the paper: Do you see related work as a way to motivate, inform, and contextualise your work, or as a tedious writing requirement? One way to address this is to read related work already at the start and continuously throughout your project, making notes. When it comes to writing, use these to draft the related work. Maybe use LLMs to revise this draft but not unguided. And when using AI to find papers, read them and question whether they are indeed related (enough) and whether their context is fitting for yours.

Incorrect Representation of Cited Work

The problem

This issue is a more focused version of the previous one. Here, a cited paper is relevant but misrepresented or misinterpreted, likely because an LLM was asked to summarise it or draw the connection to fit it in. For example, I first noticed this in a manuscript that cited one of our own papers, claiming we had shown that when people do X, it leads to effect E. In reality, participants did X in all study conditions, as this was part of our task and not a comparison. The submission thus misrepresented our findings.

Why this is problematic

Misrepresenting related work clearly demonstrates that you haven’t properly read that source and risks building your argument on findings that do not exist. It can also make you miss opportunities, such as overlooking a research gap because you incorrectly assume something has already been investigated, as in my example above.

How to fix it

I suspect this is a result of relying on an LLM’s representation of related work, either through “deep research” features or generated summaries. Nuanced understanding matters. Read papers yourself, or at the very least double-check any AI-provided summaries and sources before relying on them or writing about them in your manuscript.

Stretched Result Summaries

The problem

I’ve noticed that result summaries in my reviewed papers stretch interpretations, likely because they were generated by prompting an LLM with the results. This is different from overclaiming: The summary may not exaggerate but rather claim something that does not quite represent the results reported earlier in the paper. This may also involve metaphors. For example, a summary might state that participants experienced a new feature as an “unobtrusive guardrail”, even though none of the reported quotes or interview questions relate to “unobtrusiveness” or mention a “guardrail”.

Why this is problematic

These interpretations come across as surprising, unfounded, or far-fetched. They leave readers wondering whether the authors know more relevant information than they reported, or whether the interpretation is simply imprecise or incorrect, reflecting poorly on the manuscript overall. Moreover, metaphors can help paint a bigger picture, but only if they fit the reader’s understanding at that point in the paper. It also needs to be clear that the metaphor reflects the authors’ interpretation, not something participants said (unless they did, of course).

How to fix it

Summarise your findings in your own words first, then maybe polish it with an LLM if you feel that is needed. Watch out for unintended nuances or metaphors that AI might introduce, and add constraints in your prompt (e.g. “avoid metaphors”). Double-check AI text against your reported results or ask a colleague to read it and flag anything that feels off.

Conclusion

These four problematic patterns all concern more than writing style: maintaining scientific rigor and integrity. While there are likely more than these four, my hope is that this collection helps raise awareness and provides a useful resource to share with an AI-eager co-author, colleague or student.

All concrete examples are paraphrased or fictionalised to avoid revealing any details about the reviewed work that originally inspired this article.

--

--

Daniel Buschek
Daniel Buschek

Written by Daniel Buschek

Human-Computer Interaction Prof at U. of Bayreuth, Germany. Building tools for creative people. Critically evaluating AI's impact on users, workflows, outcomes.