Can AI write better acceptance criteria than we can?

Can AI write better acceptance criteria than we can?

Requirements Engineering·09/17/2026

We talk a lot about how AI can make requirements clearer and more complete - but what about the level below user stories, the acceptance criteria? For his Master’s thesis, Patrick Kainer built a prototype that generates acceptance criteria straight from a product’s own documentation, and then had real developers compare them, blind, against acceptance criteria written by hand. Let’s look at what he found.

AI-generated vs. human-written acceptance criteria – the numbers at a glance:

  • Readability: +20 %
  • Understandability: +6 %
  • Definability (how clearly scoped and testable): +19 %
  • Technical accuracy: +18 %

→ Overall quality rated 21 % higher for AI-generated acceptance criteria

(Participants rated the criteria on a -2 to +2 agreement scale, from “strongly disagree” to “strongly agree”. The percentages above show the difference between the AI-generated and human-written averages, expressed as a share of that scale’s full range - so +20 % in readability means participants agreed noticeably more strongly that the AI-generated criteria were easy to read.)

In his thesis, Patrick Kainer investigated whether AI, using a product’s own documentation as a knowledge base, can generate acceptance criteria that hold up against ones written by hand. The prototype he built takes a user story together with relevant product documentation, finds the passages that actually relate to that story, and feeds both into a language model to draft acceptance criteria - a similar idea to how AI-supported tools like storywise pull context into the writing process. Developers at an Austrian software company were then shown acceptance criteria for the same user stories - some written manually, some generated by the prototype - without knowing which was which. They rated each set on four dimensions: readability, understandability, definability (how clearly the criteria are scoped and testable), and technical accuracy. And the results were pretty clear-cut. Let’s dive into the most interesting findings of the study.

Finding #1: AI is strongest exactly where precision matters most

“AI can automate routine tasks in requirements engineering, make requirements more traceable, and support a more objective, more reproducible way of writing acceptance criteria.”

— Patrick Kainer

The gaps were largest in readability, definability, and technical accuracy - readability came out on top at +20 %, with definability and technical accuracy close behind at +19 % and +18 %. Understandability was the clear outlier, improving by only +6 %: human-written criteria were already reasonably easy to grasp, just less precise and less consistently scoped.

This lines up with a pattern we’ve seen before: AI support tends to act as a kind of formal quality control. It’s less about making requirements sound nicer, and more about making them unambiguous, verifiable, and consistently structured - exactly the qualities that acceptance criteria live or die by.

Finding #2: developers said they’d actually use it

“The AI results weren’t just seen as linguistically convincing - participants judged them as genuinely ready to work with in practice.”

— Patrick Kainer

It’s one thing for criteria to score well on a survey, and another for developers to say they’d actually put them into their backlog. Asked whether they’d add a user story with the given acceptance criteria to their backlog as-is, 80% said yes for the AI-generated criteria, compared to 54 % for the human-written ones. That’s a strong signal that the improvement wasn’t just measurable - it was also felt as practically useful by the people who’d have to work with these criteria day to day.

Finding #3: AI captures most of the substance, but not all of it

“AI should be seen as a supporting tool that speeds up and standardises the writing of acceptance criteria - it doesn’t replace the professional responsibility of the developers who own them.”

— Patrick Kainer

Here’s the honest caveat, and it’s an important one: alongside the developer survey, the thesis also ran a machine-based comparison (BERTScore) to check how much of the content of the human-written criteria actually made it into the AI-generated ones. The overlap came out moderate - roughly 62–66 % semantic similarity. In other words, the AI reliably picks up on the general topic and structure, but it doesn’t always catch every detail or edge case a human author would have included.

So while the AI-generated criteria were consistently rated as clearer, better scoped, and more technically sound, the thesis is careful to point out that they shouldn’t go straight into a sprint unchecked. Human review still matters - to catch the details that get lost, and to judge whether a criterion is actually relevant to the business context. AI support here works best as a strong first draft, not a replacement for that final check.

Study details – the hard facts

If you care about the details of the study but don’t want to read the entire thesis, here’s a breakdown of how the study was conducted.


Research question

How can product documentation, combined with AI, be used to raise the quality of acceptance criteria for user stories?

Method

A prototype converts a user story and its related documentation into vectors, clusters them to find the relevant passages, and passes both to a language model (ChatGPT-5-mini) to generate acceptance criteria.

Participants

Developers at an Austrian software company, all working with acceptance criteria as part of their regular role.

Task

Participants blind-rated sets of acceptance criteria for the same user stories - some written manually, some generated by the prototype - without knowing which was which.

Evaluation criteria
  • Readability: How easy is the wording to read?
  • Understandability: How easy is the criterion to grasp and interpret correctly?
  • Definability: Is the criterion clearly scoped and testable?
  • Technical accuracy: Is the criterion logically and technically consistent?
Findings

AI-generated acceptance criteria were rated more highly than manually written ones across all four dimensions, with the largest gains in readability, definability, and technical accuracy, and a much smaller gain in understandability. A separate machine-based content comparison (BERTScore) showed moderate semantic overlap (62–66 %) between AI-generated and human-written criteria, indicating that AI support boosts clarity and structure but doesn’t yet guarantee full content coverage - human review remains necessary.


If you do want to check out the entire thesis (only available in German), you can do so here.

Are you ready for better requirements engineering?

Getting started with storywise just takes a few minutes.
Yes, it's THAT intuitive!
Explore all features -->