Scaling UX Research Through AI-Assisted Evaluation
Automating the repetitive layer of heuristic evaluation while keeping research judgment in the loop.
The Challenge
Heuristic evaluation is a valuable method for identifying usability issues early in the design process. However, conducting the evaluation manually can become repetitive and time-consuming, especially when researchers need to review multiple screens and flows based on Nielsen’s 10 usability heuristics.
In our existing workflow, researchers were responsible for reviewing the entire design flow, identifying potential violations, documenting findings, and determining their relevance. While the final interpretation still requires research expertise, the initial screening process created a significant amount of repetitive work.
“Can we automate the first layer of heuristic evaluation without removing researchers from the decision-making process?”
The Opportunity
Instead of using AI to replace heuristic evaluation, I explored how it could be used as a researcher’s first-pass screening assistant.
The goal was to shift the workflow from:
This approach allows automation to handle the repetitive part of the process while keeping human judgment where it matters most.
The Solution
I developed an AI-assisted heuristic evaluation workflow that evaluates a design flow based on Nielsen’s 10 usability heuristics using three key inputs.
1. Design Context
Information about what the product or feature is, who it is designed for, and the situation in which it will be used.
2. User & Product Goals
The intended user goal and what the flow is expected to help users accomplish.
3. Design Flow
Screenshots representing the sequence of the user journey or interaction flow.
How It Works
Define the Context
The researcher provides the AI with the context needed to understand the design. This prevents the AI from evaluating the interface purely based on visual characteristics without understanding its intended use.
- Who is the user?
- What are they trying to accomplish?
- What is the purpose of this flow?
Upload the Design Flow
The researcher provides screenshots of the complete design flow. The screenshots give the evaluator visual and contextual information about the interaction sequence, rather than evaluating individual screens in isolation.
AI First-Pass Evaluation
The AI evaluates the flow based on Nielsen’s 10 usability heuristics. The output is intended to function as a screening layer, not a final research finding.
Researcher Validation
This is where the researcher remains critical. Researchers review the issues identified by the AI and determine relevance, accuracy, severity, and context.
This creates a human-in-the-loop workflow, where AI accelerates the screening process while researchers retain ownership of interpretation and decision-making.
Actionable Research Findings
After validation, relevant issues can be synthesized into actionable findings and recommendations for the design or product team. The result is a researcher-validated evaluation produced through a more efficient workflow.