We talk a lot about how AI can make requirements clearer and more complete - but what about the level below user stories, the acceptance criteria? For his Master’s thesis, Patrick Kainer built a prototype that generates acceptance criteria straight from a product’s own documentation, and then had real developers compare them, blind, against acceptance criteria written by hand. Let’s look at what he found.
AI-generated vs. human-written acceptance criteria – the numbers at a glance:
- Readability: +20 %
- Understandability: +6 %
- Definability (how clearly scoped and testable): +19 %
- Technical accuracy: +18 %
→ Overall quality rated 21 % higher for AI-generated acceptance criteria
(Participants rated the criteria on a -2 to +2 agreement scale, from “strongly disagree” to “strongly agree”. The percentages above show the difference between the AI-generated and human-written averages, expressed as a share of that scale’s full range - so +20 % in readability means participants agreed noticeably more strongly that the AI-generated criteria were easy to read.)
In his thesis, Patrick Kainer investigated whether AI, using a product’s own documentation as a knowledge base, can generate acceptance criteria that hold up against ones written by hand. The prototype he built takes a user story together with relevant product documentation, finds the passages that actually relate to that story, and feeds both into a language model to draft acceptance criteria - a similar idea to how AI-supported tools like storywise pull context into the writing process. Developers at an Austrian software company were then shown acceptance criteria for the same user stories - some written manually, some generated by the prototype - without knowing which was which. They rated each set on four dimensions: readability, understandability, definability (how clearly the criteria are scoped and testable), and technical accuracy. And the results were pretty clear-cut. Let’s dive into the most interesting findings of the study.
Finding #1: AI is strongest exactly where precision matters most
“AI can automate routine tasks in requirements engineering, make requirements more traceable, and support a more objective, more reproducible way of writing acceptance criteria.”
The gaps were largest in readability, definability, and technical accuracy - readability came out on top at +20 %, with definability and technical accuracy close behind at +19 % and +18 %. Understandability was the clear outlier, improving by only +6 %: human-written criteria were already reasonably easy to grasp, just less precise and less consistently scoped.
This lines up with a pattern we’ve seen before: AI support tends to act as a kind of formal quality control. It’s less about making requirements sound nicer, and more about making them unambiguous, verifiable, and consistently structured - exactly the qualities that acceptance criteria live or die by.
Finding #2: developers said they’d actually use it
“The AI results weren’t just seen as linguistically convincing - participants judged them as genuinely ready to work with in practice.”
It’s one thing for criteria to score well on a survey, and another for developers to say they’d actually put them into their backlog. Asked whether they’d add a user story with the given acceptance criteria to their backlog as-is, 80% said yes for the AI-generated criteria, compared to 54 % for the human-written ones. That’s a strong signal that the improvement wasn’t just measurable - it was also felt as practically useful by the people who’d have to work with these criteria day to day.
Finding #3: AI captures most of the substance, but not all of it
“AI should be seen as a supporting tool that speeds up and standardises the writing of acceptance criteria - it doesn’t replace the professional responsibility of the developers who own them.”
Here’s the honest caveat, and it’s an important one: alongside the developer survey, the thesis also ran a machine-based comparison (BERTScore) to check how much of the content of the human-written criteria actually made it into the AI-generated ones. The overlap came out moderate - roughly 62–66 % semantic similarity. In other words, the AI reliably picks up on the general topic and structure, but it doesn’t always catch every detail or edge case a human author would have included.
So while the AI-generated criteria were consistently rated as clearer, better scoped, and more technically sound, the thesis is careful to point out that they shouldn’t go straight into a sprint unchecked. Human review still matters - to catch the details that get lost, and to judge whether a criterion is actually relevant to the business context. AI support here works best as a strong first draft, not a replacement for that final check.
Study details – the hard facts
If you care about the details of the study but don’t want to read the entire thesis, here’s a breakdown of how the study was conducted.
| Research question | How can product documentation, combined with AI, be used to raise the quality of acceptance criteria for user stories? |
|---|---|
| Method | A prototype converts a user story and its related documentation into vectors, clusters them to find the relevant passages, and passes both to a language model (ChatGPT-5-mini) to generate acceptance criteria. |
| Participants | Developers at an Austrian software company, all working with acceptance criteria as part of their regular role. |
| Task | Participants blind-rated sets of acceptance criteria for the same user stories - some written manually, some generated by the prototype - without knowing which was which. |
| Evaluation criteria |
|
| Findings | AI-generated acceptance criteria were rated more highly than manually written ones across all four dimensions, with the largest gains in readability, definability, and technical accuracy, and a much smaller gain in understandability. A separate machine-based content comparison (BERTScore) showed moderate semantic overlap (62–66 %) between AI-generated and human-written criteria, indicating that AI support boosts clarity and structure but doesn’t yet guarantee full content coverage - human review remains necessary. |
If you do want to check out the entire thesis (only available in German), you can do so here.
