We evaluated our chatbot using a cognitive walkthrough. I'm writing this article to share what we did, and to hear how others are evaluating their own conversational AI chatbots.

A traveler asked for a fall getaway for two. The bot answered with a full paragraph on each option. Then it repeated three of those same names under a new heading called "top picks." Then it listed all of them again under "available now." Finally it offered to check availability it had already checked.

There was no button in the wrong place, no color with poor contrast, no label that misled anyone. It was however, by any reasonable measure, unusable. The traveler had to read the same six names three times to find one thing to click.

That is a new kind of bug. It does not live in a layout. It lives in a sequence of sentences a model generated, one token after another, with no designer ever laying out a single pixel. My first instinct was to file it under data science or engineering: it is not a screen, so it is not design. But it is the experience a person is having, so that makes it design.

The same tools that catch a broken checkout flow caught this. The same principles that trim a cluttered dashboard fixed it. Only the canvas changed.

Testing

A cognitive walkthrough is normally run against a screen. Someone steps through a task one click at a time and asks the same four questions at each step:

  • Will the person know what to do here?
  • Will they see that this control does it?
  • Will they understand the feedback once they act?
  • Will they know they're making progress?

Nothing about that method requires a screen. It requires a sequence of steps a person moves through toward a goal. Conversational AI is just that. So the walkthrough ran the same way it would against a checkout flow. The findings provided a locatable failure at a specific step: the person could not tell, after step three, whether step three had already happened. That is the same failure mode as a form that submits silently, with no confirmation. The walkthrough did not know it was looking at an LLM. It was looking for a missing signal, and it found one.

That distinction matters for what happens next. A vague complaint gets a vague fix, usually "make the tone friendlier." A located failure gets a specific one: this step needs a visible outcome, this step needs to stop repeating itself.

Density

For screen design, density means how much visual weight sits in a given space. In a chat reply it means something almost identical: how much information does a reader have to wade through before they are confident to act.

The fall-getaway reply is exactly this failure: six properties, named across three headings, none attached to a single clear next step.

Word count and structure, current reply versus proposed reply
Before — roughly 400 words After — under 100 words
Full blurb for every option, in prose One headline: count, dates, party size
A second list of "top picks," repeating names already shown Top three only, one line of reasoning each, one action per item
A third list under a new heading, repeating them again Remaining options named once, in a single line
An offer to check availability the reply already contained One optional follow-up, only if not already answered
No consistent action attached to any single item —

Same underlying information. A quarter of the length. Three established laws explain why that mattered more than it might seem. Hick's Law says more options take longer to decide between. Choice overload, closely related, says a large set of options does not just slow the decision, it makes people feel worse about the choice they eventually make. Six properties restated three times is not six choices, it behaves like eighteen, and it costs both time and confidence. Miller's Law says people hold information more easily in small chunks than as one continuous stream. The fix applied all three: fewer things to choose from, grouped under a headline, a shortlist, and a single closing line.

Similar to a visual designer cutting a dashboard from twelve visible metrics to four. The screen did not get smaller. The content that mattered got easier to find.

Interaction

Interaction design is about more than what a screen contains. It is about order: what happens first, and what a person has to do at each step before moving to the next. A chat reply has the same problem, one turn at a time instead of one screen at a time.

The walkthrough found two separate interaction failures, both about sequence rather than content.

Turn sequencing, current reply versus proposed reply
Before After
Party size asked again after the traveler already gave it When, Where, Who persist for the whole conversation once known
Itinerary link buried mid-paragraph after several dense exchanges Draft link surfaced early, as the primary thing in the reply
Bot offers to check availability it already checked No offer to repeat a step already completed in this same reply
Only ever offers to add more to the plan Reply ends with a clear Affirm, Change, or Remove

This is progressive disclosure, applied to a conversation instead of a screen. The principle is usually described as showing only what a person needs at the current step, and revealing more only once they ask for it. A reply that re-asks a known answer, or dumps the next three steps into one turn "just in case," is overload. One job per turn is the same restraint, paced across exchanges instead of across a form.

The link to the actual plan the AI built is the clearest case of Fitts's Law at work: how long it takes someone to reach a target depends on how far away it is and how big it is. A link sitting mid-paragraph, after several exchanges, is a small target buried at the end of a long reach. Moving it to the top of the reply and treating it as the primary artifact, not a line inside a sentence, is the same fix in different material.

Current Experience
6 bot turns~440 words0 button-style CTAs
⋮ ✛
✥ ⌄
Powered by AI ?
✦

Hi! Welcome to Solstice Vacation Club. I'm your virtual assistant, here to help you discover personalized travel ideas, check availability and make the most of your travel planning. What can I assist you with today?

Today
Plan a trip
5m ago
✦

Where would you like to go?

Florida
4m ago
✦

What dates should I plan for your Florida trip? If you know your party size too, include that as well so I can tailor the itinerary.

Oct 23-27 family of 4
4m ago
✦

What part of Florida would you like to visit? Orlando, Miami, Tampa Bay, the Gulf Coast, the Keys, or another area all work.

Orlando
3m ago
✦

Great choice. For your October 23rd, 2026 to October 27th, 2026 Orlando trip, here's a solid family-friendly plan to start with:

Solstice Vacation Club at Coral Cove or Solstice Vacation Club at Palmetto Grove
A strong home base for a family of 4, with easy access to the parks and plenty to do nearby.

Day 1 — Easy arrival day
Check in, enjoy the pool, and have a relaxed dinner nearby. Good first-night options include Trattoria Bellina or a casual meal along the main tourist strip.

Day 2 — Starlight Kingdom
Best if you want the classic Orlando experience. In late October, Starlight Kingdom's After Hours Halloween event may also be an option on select nights.

Day 3 — Reef Line Aquarium + Skyline Commons
A lighter day that works well between big park days. Reef Line Aquarium has seasonal family programming, and Skyline Commons gives you several easy attractions in one area.

Day 4 — Apex Studios Orlando
Good for a full day of rides and themed attractions. If your group prefers a less intense day, this could be swapped for Coral Cove Marine Park.

Day 5 — Slow morning before departure
Resort time, outlet shopping at the Lakeside Outlet Collection, or one last meal before heading out.

If you want, I can turn this into a saved trip and start by adding your first confirmed items. A simple starting set would be:

  • your Orlando stay
  • Starlight Kingdom
  • Reef Line Aquarium + Skyline Commons
  • and one Universal day

Just say "add those," or tell me which items you want on the trip.

add those
2m ago
✦

Done — your trip is saved.

View your trip: Orlando

Right now it includes:

  • Oct 23, 2026, 3:00 PM — Solstice Vacation Club at Palmetto Grove
  • Oct 24, 2026, 9:00 AM — Starlight Kingdom
  • Oct 25, 2026, 10:00 AM — Reef Line Aquarium
  • Oct 25, 2026, 2:00 PM — Skyline Commons
  • Oct 26, 2026, 9:00 AM — Apex Studios Orlando

A good next step for Oct 23 or the morning of Oct 27 would be a relaxed family meal or a low-key activity near the resort. I can add one of those next if you want.

Type a message...
Proposed Experience
5 bot turns~120 words1 button-style CTA
⋮ ✛
✥ ⌄
Powered by AI ?
✦

Hi! Welcome to Solstice Vacation Club. I'm your virtual assistant, here to help you discover personalized travel ideas, check availability and make the most of your travel planning. What can I assist you with today?

Today
Plan a trip
3m ago
✦

Where would you like to go?

Florida
3m ago
✦

What dates, and how many people?

Oct 23-27 family of 4
2m ago
✦

Which area — Orlando, Miami, Tampa Bay, the Gulf Coast, or the Keys?

Orlando
2m ago
✦

Orlando trip ready — Oct 23–27, family of 4.

I've drafted a Coral Cove / Palmetto Grove stay with Starlight Kingdom, Skyline Commons, and Apex Studios days — suggested because parks and a nearby resort work well for a family of 4.

View Plan →

What did I get right? What would you swap or cut?

Type a message...

Open the standalone prototype in a new tab →

The guidelines

Before this walkthrough, we'd already put together a set of conversational AI guidelines as a foundation for the chatbot. The document is principle-level, not feature-level: a five-phase model for how the conversation should move, then a numbered set of guidelines, each with a plain-language principle and a Do/Avoid pair.

It is written for developers, data scientists, and interaction designers together, not for designers alone. A guidelines document that only designers read stays a design opinion. One written for the people who own the model becomes a shared reference both sides can point to.

That shared reference is also what decided the shape of the actual handoff. Normally a fix like this leaves a designer's hands as a redline: a mockup, spacing values, a spec an engineer can build against. This one left as something closer to a style guide than a screen:

  • A short list of shared rules
  • A named shape for each type of reply
  • Pairs of examples, a bad version and a good version of the same reply, side by side

A style guide teaches a human writer what "on-brand" sounds like the same way, through paired examples rather than abstract rules alone. This is that same convention, aimed at a model instead of a person.

The walkthrough did not have to argue that repetition was bad, in the abstract, to a new audience. It had to show that a specific reply violated a principle that already existed, in terms the guidelines already used. Where a genuine gap turned up, the fix was to add a new principle to the same numbered list, not invent a new kind of document. The next finding, whatever it turns out to be, has the same place to go.

That is the real reason this belongs in the same toolkit as testing, density, and interaction. It is not one more clever fix. It is the standing artifact that makes every future fix cheaper.

Isn't this just prompt engineering?

Someone will call this prompt engineering with a UX label taped on. They're half right. The fix does live in a system prompt. But prompt engineering did not find the problem, and it did not decide what the right shape was.

A cognitive walkthrough found it, design guidelines decided the fix: fewer items to compare, a clear hierarchy of headline then shortlist then one action per item.

Prompt engineering carried the fix into production. It is the medium. It is not the method.

Constraints

This is not a case study with a lift number attached to it. There is no conversion chart here, no before-and-after engagement metric, because the point was never to make people chat longer. If anything, a shorter, clearer reply should mean the conversation ends sooner, which is the goal, not the failure.

It is also not a guarantee. A system prompt strongly steers a model. It does not enforce behavior the way code enforces a validation rule. The same reply can drift back toward the old pattern if the prompt is edited later by someone who never saw the walkthrough, and nothing stops that automatically. The fix has to be re-checked the way any design pattern has to be re-checked once other people start touching it.

And it depends on a partner that has nothing to do with design. Asking a model to emit a clean bulleted shortlist with a link only matters if the interface actually renders bullets and links as something other than plain text. The prompt can ask for structure. Somebody else still has to paint it.

Takeaways

  • Run a cognitive walkthrough on a chat transcript the same way you would run one on a screen flow. The four questions do not require pixels to ask.
  • Treat word count as a design signal, not just a writing one. A reply that repeats itself is a density problem before it is a tone problem.
  • Write the shape down as rules plus paired bad and good examples, not as a paragraph of guidance. Models learn from contrast the same way people do.
  • Decide, explicitly, what a dead end looks like. If failure and success share a format, someone will mistake one for the other.
  • Expect the deliverable to change form. The judgment stays the same. The artifact might be a spec instead of a mockup.

Closing

The ticket that started this did not mention design. It mentioned a chatbot reply that felt too long. Nobody asked for a cognitive walkthrough, and nobody asked for a response-shape spec with paired examples. Both showed up anyway, because the problem underneath the complaint was a design problem wearing a different coat.

The canvas will keep changing. It already has, more than once, in the years most of us have been doing this work. The skills underneath it, watching a person move through a task and noticing exactly where they lose the thread, have not changed nearly as much.