A participant pauses before pressing a button. Writing “the button needs more contrast” skips the most important part of the analysis. The pause is observable. Its cause is still uncertain. The participant might be checking the price, looking for reassurance or wondering whether the action is reversible.
View full image
That distinction matters because the same visible hesitation can lead to very different changes. A brighter button will not answer a question about what happens after payment.
Keep a short evidence trail
For each issue, record the task, what happened, the consequence and where to find the moment in the session. Then add an interpretation separately. This makes it possible for another person to disagree with your explanation without losing the original observation.
Consider a hypothetical checkout test: a participant submits the form twice while the first request is still processing. The observation is the repeated submission. Possible causes include missing progress feedback or uncertainty about whether the first click registered. Disabling repeat submission and showing a processing state is one candidate response; changing the label alone might not address it.
Frequency is only one part of priority
A minor annoyance repeated in every session can matter less than a single failure that risks losing work. Ask whether the issue blocks the task, whether recovery is possible and how often the affected situation occurs in the product.
Small qualitative samples help expose mechanisms. They do not reliably estimate the percentage of all customers who will encounter a problem. Report “four of six participants in these sessions” rather than presenting the fraction as a population rate.
A result from my RDR2 study
In my independent target-selection study, the design question was who would receive an interaction. Making the available commands readable was insufficient if the player associated them with the wrong character. That directed the exploration toward target identification and anticipation.
The final controller prototype was tested with seven participants who had already seen the proposal during the visual evaluation. Their performance supported the clarity of the tested flow, but familiarity limits what the result says about first-time use. Moving characters and combat were outside that prototype. Those are conditions for a next test, not details to hide beneath a success claim.
Turn the finding into a bounded next step
A useful recommendation names the change, the behavior it should affect and the next check. For example: “Show the potential recipient before opening the action menu; check whether unfamiliar players correctly anticipate the target when two characters overlap.”
AI can help organize notes, but check every proposed cluster against the recordings or transcripts. A fluent summary can merge different causes or lose a contradiction. Prioritization still needs product context and human review; an automatically assigned impact score does not supply either.
Research is ready to influence delivery when the team can trace a proposed change back to evidence and understands what remains untested.
