We evaluated our chatbot using a cognitive walkthrough. I'm writing this article to share what we did, and to hear how others are evaluating their own conversational AI chatbots.
A traveler asked for a fall getaway for two. The bot answered with a full paragraph on each option. Then it repeated three of those same names under a new heading called "top picks." Then it listed all of them again under "available now." Finally it offered to check availability it had already checked.
There was no button in the wrong place, no color with poor contrast, no label that misled anyone. It was however, by any reasonable measure, unusable. The traveler had to read the same six names three times to find one thing to click.
That is a new kind of bug. It does not live in a layout. It lives in a sequence of sentences a model generated, one token after another, with no designer ever laying out a single pixel. My first instinct was to file it under data science or engineering: it is not a screen, so it is not design. But it is the experience a person is having, so that makes it design.
The same tools that catch a broken checkout flow caught this. The same principles that trim a cluttered dashboard fixed it. Only the canvas changed.
Testing
A cognitive walkthrough is normally run against a screen. Someone steps through a task one click at a time and asks the same four questions at each step:
- Will the person know what to do here?
- Will they see that this control does it?
- Will they understand the feedback once they act?
- Will they know they're making progress?
Nothing about that method requires a screen. It requires a sequence of steps a person moves through toward a goal. Conversational AI is just that. So the walkthrough ran the same way it would against a checkout flow. The findings provided a locatable failure at a specific step: the person could not tell, after step three, whether step three had already happened. That is the same failure mode as a form that submits silently, with no confirmation. The walkthrough did not know it was looking at an LLM. It was looking for a missing signal, and it found one.
That distinction matters for what happens next. A vague complaint gets a vague fix, usually "make the tone friendlier." A located failure gets a specific one: this step needs a visible outcome, this step needs to stop repeating itself.
Density
For screen design, density means how much visual weight sits in a given space. In a chat reply it means something almost identical: how much information does a reader have to wade through before they are confident to act.
The fall-getaway reply is exactly this failure: six properties, named across three headings, none attached to a single clear next step.
| Before — roughly 400 words | After — under 100 words |
|---|---|
| Full blurb for every option, in prose | One headline: count, dates, party size |
| A second list of "top picks," repeating names already shown | Top three only, one line of reasoning each, one action per item |
| A third list under a new heading, repeating them again | Remaining options named once, in a single line |
| An offer to check availability the reply already contained | One optional follow-up, only if not already answered |
| No consistent action attached to any single item | — |
Same underlying information. A quarter of the length. Three established laws explain why that mattered more than it might seem. Hick's Law says more options take longer to decide between. Choice overload, closely related, says a large set of options does not just slow the decision, it makes people feel worse about the choice they eventually make. Six properties restated three times is not six choices, it behaves like eighteen, and it costs both time and confidence. Miller's Law says people hold information more easily in small chunks than as one continuous stream. The fix applied all three: fewer things to choose from, grouped under a headline, a shortlist, and a single closing line.
Similar to a visual designer cutting a dashboard from twelve visible metrics to four. The screen did not get smaller. The content that mattered got easier to find.
Interaction
Interaction design is about more than what a screen contains. It is about order: what happens first, and what a person has to do at each step before moving to the next. A chat reply has the same problem, one turn at a time instead of one screen at a time.
The walkthrough found two separate interaction failures, both about sequence rather than content.
| Before | After |
|---|---|
| Party size asked again after the traveler already gave it | When, Where, Who persist for the whole conversation once known |
| Itinerary link buried mid-paragraph after several dense exchanges | Draft link surfaced early, as the primary thing in the reply |
| Bot offers to check availability it already checked | No offer to repeat a step already completed in this same reply |
| Only ever offers to add more to the plan | Reply ends with a clear Affirm, Change, or Remove |
This is progressive disclosure, applied to a conversation instead of a screen. The principle is usually described as showing only what a person needs at the current step, and revealing more only once they ask for it. A reply that re-asks a known answer, or dumps the next three steps into one turn "just in case," is overload. One job per turn is the same restraint, paced across exchanges instead of across a form.
The link to the actual plan the AI built is the clearest case of Fitts's Law at work: how long it takes someone to reach a target depends on how far away it is and how big it is. A link sitting mid-paragraph, after several exchanges, is a small target buried at the end of a long reach. Moving it to the top of the reply and treating it as the primary artifact, not a line inside a sentence, is the same fix in different material.
The guidelines
Before this walkthrough, we'd already put together a set of conversational AI guidelines as a foundation for the chatbot. The document is principle-level, not feature-level: a five-phase model for how the conversation should move, then a numbered set of guidelines, each with a plain-language principle and a Do/Avoid pair.
It is written for developers, data scientists, and interaction designers together, not for designers alone. A guidelines document that only designers read stays a design opinion. One written for the people who own the model becomes a shared reference both sides can point to.
That shared reference is also what decided the shape of the actual handoff. Normally a fix like this leaves a designer's hands as a redline: a mockup, spacing values, a spec an engineer can build against. This one left as something closer to a style guide than a screen:
- A short list of shared rules
- A named shape for each type of reply
- Pairs of examples, a bad version and a good version of the same reply, side by side
A style guide teaches a human writer what "on-brand" sounds like the same way, through paired examples rather than abstract rules alone. This is that same convention, aimed at a model instead of a person.
The walkthrough did not have to argue that repetition was bad, in the abstract, to a new audience. It had to show that a specific reply violated a principle that already existed, in terms the guidelines already used. Where a genuine gap turned up, the fix was to add a new principle to the same numbered list, not invent a new kind of document. The next finding, whatever it turns out to be, has the same place to go.
That is the real reason this belongs in the same toolkit as testing, density, and interaction. It is not one more clever fix. It is the standing artifact that makes every future fix cheaper.
Isn't this just prompt engineering?
Someone will call this prompt engineering with a UX label taped on. They're half right. The fix does live in a system prompt. But prompt engineering did not find the problem, and it did not decide what the right shape was.
A cognitive walkthrough found it, design guidelines decided the fix: fewer items to compare, a clear hierarchy of headline then shortlist then one action per item.
Prompt engineering carried the fix into production. It is the medium. It is not the method.
Constraints
This is not a case study with a lift number attached to it. There is no conversion chart here, no before-and-after engagement metric, because the point was never to make people chat longer. If anything, a shorter, clearer reply should mean the conversation ends sooner, which is the goal, not the failure.
It is also not a guarantee. A system prompt strongly steers a model. It does not enforce behavior the way code enforces a validation rule. The same reply can drift back toward the old pattern if the prompt is edited later by someone who never saw the walkthrough, and nothing stops that automatically. The fix has to be re-checked the way any design pattern has to be re-checked once other people start touching it.
And it depends on a partner that has nothing to do with design. Asking a model to emit a clean bulleted shortlist with a link only matters if the interface actually renders bullets and links as something other than plain text. The prompt can ask for structure. Somebody else still has to paint it.
Takeaways
- Run a cognitive walkthrough on a chat transcript the same way you would run one on a screen flow. The four questions do not require pixels to ask.
- Treat word count as a design signal, not just a writing one. A reply that repeats itself is a density problem before it is a tone problem.
- Write the shape down as rules plus paired bad and good examples, not as a paragraph of guidance. Models learn from contrast the same way people do.
- Decide, explicitly, what a dead end looks like. If failure and success share a format, someone will mistake one for the other.
- Expect the deliverable to change form. The judgment stays the same. The artifact might be a spec instead of a mockup.
Closing
The ticket that started this did not mention design. It mentioned a chatbot reply that felt too long. Nobody asked for a cognitive walkthrough, and nobody asked for a response-shape spec with paired examples. Both showed up anyway, because the problem underneath the complaint was a design problem wearing a different coat.
The canvas will keep changing. It already has, more than once, in the years most of us have been doing this work. The skills underneath it, watching a person move through a task and noticing exactly where they lose the thread, have not changed nearly as much.