Home About me Contact
Experiment · UX audit with AI

18 findings from AI, 2 of them worth keeping

I ran a heuristic UX audit of a live onboarding flow with Claude Design — to see whether AI could catch what a tested, refined flow had missed

Category
Real estate · AI experiment
My role
UX audit, AI-assisted analysis
Method
Nielsen's heuristics
Tool
Claude Design ↗
Audit subject
My own 3-year-old onboarding
Claude open on a laptop

Context

I took the onboarding flow from my previous project, PerCheck. It shipped about 3 years ago, was tested with real users and refined after that — and I asked Claude Design to audit it against Nielsen's heuristics.

The point was to see whether AI could spot something I had overlooked, or argue with decisions I had already made at the time

Limitations

On a PRO plan the usage limits meant I could not feed the model all the context of the product. I shared onboarding screenshots and a short description of the product and the flow; the links I gave it were not reachable from its environment. So the AI only used what fit inside the limit — isolated screens, not a working product.

The screenshots I gave Claude as context
{{ ctxCounter }}
Welcome modal of the onboarding First tooltip pointing at the order button Tooltip on the add item dropdown Tour step explaining what can be checked Tour step about passport recognition

What came back

Claude produced 18 findings. Without the full flow and the business logic behind it, the model reasoned from screenshots alone, so it raised «problems» that were never problems in the product. For example:

  • it flagged the missing progress bar — which was left out on purpose, because the flow is non-linear and has no fixed order of steps;
  • it proposed preset configurations — impossible here, since there is no way to predict which entities, and how many, a user will want to check.

I went through every finding and kept only what held up against the product.

16
findings were irrelevant or needed context the model did not have
2
gave a useful direction, which I then refined against the product logic

The 2 findings worth acting on

1.
An ambiguous CTA on the entry modal. If a user skips the modal text — which happens often — «Let's go» does not say whether they are starting a guided tour or jumping straight into the product. I shortened the modal to one sentence and made the button name the action, so the screen reads at a glance.
Original modal with the Let's go button
Revised modal with a shorter text and an explicit button label
2.
Tooltips repeating the buttons they pointed at. We knew about this while building the feature: those steps had to be highlighted, but there was little to say about them. In a couple of places the tooltip partly repeated the text on the button, and Claude pointed that out. The step still had to be highlighted, so the copy had to either add real information or get as short as possible. The tooltips still start the flow, but now they add something the button does not.
Original tooltips repeating the button labels
Refined tooltips carrying more meaningful content
What the experiment showed

AI is usable for a UX audit, but its output tracks the context it is given. Without enough of it, the model:

  • writes generic, textbook conclusions;
  • misreads how the flow actually works;
  • suggests solutions that ignore the business logic behind the product.

At the same time, even with limited data, AI can highlight certain growth points that can be used as a starting point for further analysis. Even so, it pointed at growth areas worth a second look.

11%
of the AI findings turned out to be genuinely useful — 2 out of 18

What I took away

I would use it as a supporting tool: collect hypotheses, judge each one, drop what does not apply, refine what does. A good assistant — but it cannot fully replace a designer or a researcher.

Alena Slastnikova · Product Designer Next case: toi_b →