Sitemap

Don’t Review with an LLM (Laundry List Method)

The problem of asking AI for problems

--

Press enter or click to view image in full size
Illustration of a sheet of paper with a stylised bullet list in front of a fiery red background.
The “Laundry List Method” of reviewing lists what could have been done instead, easily generated with LLMs.

Not just at CHI 2026, I’ve noticed many people now review-to-reject. All studies have tradeoffs, all papers have limited scope. Sure, reviewing is fast, even with “passing knowledge”, if we list what isn’t there and conclude it’s lacking. But we should review what has been done.

Unfortunately, Large Language Models (LLMs) have made it easy to produce laundry lists of stereotypical, unspecific, and unfitting “issues”. I suspect this has made this kind of bad review more prevalent — and motivated me to run a little experiment:

In this article, I “review” one of our own papers with ChatGPT to show how this leads to bad review quality.

The “Laundry List Method” (LLM) of reviewing

We look at the use of AI to generate a list of comments, which reviewers might then expand on, also in own words: For example, a reviewer might ask a chatbot to “List critical points for a review of the attached paper.” This saves the time and effort required to deeply engage with the paper.

I tried it out with ChatGPT and one of our own papers. I categorise the responses below, following common elements of HCI papers: design, method, evaluation, interpretation.

Press enter or click to view image in full size
Screenshot of prompt and response in ChatGPT. The concrete paper name is redacted for the screenshot, while a line in the response is highlighted. It says, referring to the generated list of issues: “You can pick those that fit best depending on your overall evaluation goal.”
Example prompt and output in ChatGPT. Note how the highlighted response also suggests that, as a reviewer, I should have a fixed “evaluation goal” in mind upfront, to support with the generated issues.

1. Design: “What about these alternatives?”

Example: ChatGPT’s output raised many points on our design and prototype. It “may suffer scalability issues” and it “assumes that users are willing and able to meaningfully [do X, which] may under-address usability burden”. It also “does not support [alternative]”, and finally, “the paper does not engage deeply with these trade-offs”.

The problem: Such responses do not meaningfully evaluate the design/prototype. A prototype need not be on a product scale. Every design makes assumptions. And as long as a paper has a prototype, its design could have been different. However, that does not mean the presented prototype is not adequate, valuable, and useful for how it is employed in this specific research.

A better review would assess, for example, how the design assumptions, rationale, and goals are communicated and embedded into literature, how the resulting design addresses these goals, which features are novel or interesting recombinations, and how appropriate the prototype is for its use in the chosen empirical method.

2. Method: “It does not generalise!”

Example: ChatGPT pointed out the “limited scope of tasks and context”, highlighting that there are many other tasks and so it is “unclear how well the approach generalizes”. Further points in this category were “short-term study duration” and sample size.

The problem: This is exactly what one might expect to read in reviews “on average”. It misses that all studies have to make decisions. For any study, it could have been longer and larger and included other tasks. While it can be fair to question choices related to these aspects, a reviewer who mainly lists such generic points likely has not read the paper in detail. The generic points also may not fit the paper’s specific method.

A better review would consider the study choices and assumptions in light of the research context, methodological approach, and goals. Potentially missing pieces can be assessed for the paper’s method, without envisioning an entirely different paper.

3. Evaluation: “More objective metrics!”

Example: ChatGPT produced the argument that “The evaluation appears to rely mostly on interaction logs and interviews. While qualitative insights are valuable, the paper might benefit from more rigorous metrics”, plus a long list of examples.

The problem: It can’t get more stereotypical. It also put interaction log analysis into the qualitative bucket. Yet crucially, it again covers what alternative work could have been done instead of reviewing ours.

A better review would assess, for example, if the analysis choices are adequately motivated and reported, and how they fit the method and data and serve the research questions.

4. Interpretation: “Not deep enough!”

Example: ChatGPT generated text saying “While the paper reports [X] the analysis may be superficial”. It then listed aspects mentioned in our results and discussion, claiming we “underexplore” them by giving examples of further questions.

The problem: While shallow discussions are indeed a common issue in HCI in my experience, it is also always possible to generate questions that are unanswered (yet). For any analysis, one can claim it should do even more. If this is the main way a review engages with that part of a paper, this misses the potential value of the insight that is indeed in the paper.

A better review would assess the value of what has been analysed and discussed, for example, considering how it is embedded into theory and reflected on in light of related studies, how this informs future design and what to study and clarify next, and what the potential big-picture impact is for the research community and beyond.

Takeaways

In summary, asking LLMs to list problems of a paper leads to unsuitable reviews because most points will miss the point:

Such reviews list what could have been done instead of assessing the value of what has been done.

I hope this article can raise awareness and helps researchers, in particular junior colleagues, to spot and contextualise such bad (AI) reviews. After all, the worst outcome of receiving such a review might not be the initial frustration but rather listening to it and changing your paper for the worse.

Towards more constructive use

Finally, I don’t think the examples generated above can never be useful. For instance, it might be constructive to question a design making certain assumptions. Yet making this point in a review requires substantiating it by engaging with what’s in the paper, instead of simply listing it generically.

Moreover, I don’t think that LLMs can never be used constructively in the review context more broadly. For example, my “experiment” here might indeed be valuable from an author perspective: I get to ignore ChatGPT’s unfitting points yet some might stimulate valuable reflection to improve my paper, before sending it to peer review. We explored a concrete concept in this direction at CHI’24. There, crucially, the power over engagement with AI-generated text rests in the hands of the paper authors, not the evaluaters, which I think makes this use of AI more resilient against AI errors and “slop”.

What to do now?

As a community, the issue of reviewing-to-reject via the “Laundry List Method” (or LLMs in their other meaning) needs to be discussed in the context of systems and pressures that make engaged reviewing difficult. There are likely no quick and easy answers. That said, here’s a list of concrete actions you can take right now:

  • Clarify to reviewers that ACM forbids reviewers to upload papers to chatbots. In a response letter, you could perhaps combine this with a diplomatically worded hope for human review.
  • Also dare to respond clearly to unfitting expectations, in particular regarding positivist vs interpretivist approaches (resourses below).
  • As a PhD student or advisor, run a similar experiment with your own paper, or include it in your teaching, to sensitise colleagues and students to the downsides of this review approach.
  • As an AC (or fellow reviewer), remind reviewers that they are expected to engage with the paper, and ask them to update their review if it reads like a generated list of generic concerns or alternative visions for the work.

Resources

  • Excellence in reviewing and championing papers (Ken Hinckley)
  • ACM policy on LLMs and reviews (acm)
  • On qualitative HCI work and reviewing it (arxiv)
  • How to respond to reviewer critiques for qualitative work (google doc)
  • Extreme example of a review raising 80 questions (open review)

--

--

Daniel Buschek
Daniel Buschek

Written by Daniel Buschek

Human-Computer Interaction Prof at U. of Bayreuth, Germany. Building tools for creative people. Critically evaluating AI's impact on users, workflows, outcomes.