OverviewProblemMethodsFindingsRedesignResultsReflection
Back to all work
Case Study · 06 of 06

WhatsApp, enhanced.

Small interventions, outsized friction. A 12-week self-directed usability study that diagnosed low predictability as the through-line — and used Figma redesigns to prove the fix.

Type
Self-directed
Methods
Performance + think-aloud
Validation
Paired t-test
Duration
12 weeks
Outcome
4 of 5 tasks improved (sig.)
Heuristic evaluationPerformance testingThink-aloud protocolTask analysisFigma prototypingPaired t-testNielsen heuristics
My Role

Sole researcher and designer — study design, moderation, statistical analysis, and Figma prototyping. End-to-end ownership of a 12-week project.

Approach

Mixed-methods: heuristic evaluation surfaced candidate issues; performance testing quantified them; think-aloud protocol explained why they happened; redesigns fixed them; t-tests confirmed the improvement.

Platform

WhatsApp iOS. Six tasks covering core messaging interactions: group creation, message reactions, media sharing, status updates, voice messages, and contact search.

Problem statement and objective for the WhatsApp study, alongside a quote from an international student in Korea who struggles to find hidden features in the app.
Results

Statistically significant gains.

4/5
Tasks improved
(p < 0.05)
↓31%
Average task
completion time
↑22%
Subjective satisfaction
(1–10 scale)
The Problem

WhatsApp is widely used. That doesn't mean it's easy to use.

With over 2 billion users, WhatsApp is one of the most-used apps in the world. But ubiquity doesn't equal usability. Users adapt to friction — they work around it, forget it, or assume it's normal.

The research question: are there specific interactions where WhatsApp consistently imposes unnecessary cognitive load or time cost on users — and if so, can targeted redesigns measurably fix them?

I chose this as a self-directed study because it offered a contained, testable product — ideal for demonstrating end-to-end mixed-methods capability including quantitative validation.

WhatsApp overview: founded in 2009, acquired by Facebook in 2014, with over 2 billion users — yet held back by poor ratings tied to user experience issues.
Methods

Three-stage mixed-methods design.

Qualitative discovery → quantitative measurement → Figma redesign → statistical validation. Each stage informed the next.

Three research questions guiding the study: how WhatsApp can improve feature discoverability for international users, which hidden features users most need surfaced, and how the group chat admission process affects engagement.
01

Heuristic evaluation

I evaluated WhatsApp iOS against Nielsen's 10 heuristics to identify candidate usability issues. This gave the study a structured starting point and surfaced six tasks worth testing.

02

Performance testing · Baseline

Six participants completed six tasks on the original WhatsApp. I recorded task completion time (seconds), error rate (number of incorrect actions), and subjective satisfaction (1–10 scale). Task order was randomized to control for learning effects.

03

Think-aloud protocol

Concurrent think-aloud during performance testing. Participants narrated their decision-making in real time. This surfaced the "why" behind the performance data — revealing that users were slowed not by technical failure but by unpredictable outcomes from familiar gestures.

04

Thematic synthesis

"Low predictability" emerged as the through-line. Across all six tasks, users hesitated when they couldn't anticipate what would happen next — whether from inconsistent gesture affordances, buried navigation, or feedback that didn't match their mental model.

05

Figma redesigns

Five targeted redesigns addressing the highest-severity issues: improved gesture affordance signifiers, surfaced group creation flow, enhanced message reaction discoverability, clearer media attachment indicators, and visible status update entry points.

06

Post-redesign performance testing + paired t-test

The same six participants completed the same six tasks using prototype flows on the Figma redesigns. Paired t-tests compared baseline vs. redesign performance. Four of five tested tasks showed statistically significant improvement (p < 0.05).

Key Findings

"Low predictability" as the through-line.

Across all six tasks, the performance data and think-aloud transcripts told the same story: users weren't confused by complex features. They were slowed by small, recurring moments of unpredictability — interactions that should have been obvious but weren't.

The pattern mapped cleanly to two of Nielsen's heuristics: consistency and standards and visibility of system status. WhatsApp's gesture vocabulary wasn't broken — it was inconsistently applied.

Task-level findings

Redesign

Five targeted interventions.

Each redesign addressed a specific predictability failure — making affordances explicit, surfacing buried entry points, and aligning system feedback with user mental models.

01

Group creation — surfaced entry point

Added "New Group" as a direct action in the main chat list header (alongside the search icon), eliminating the need to navigate through New Chat → New Group.

02

Message reactions — visible affordance

Added a subtle emoji icon on hover/long-press preview that communicates the gesture affordance before the user commits — reducing failed short-presses.

03

Media attachment — labeled sub-menu

Replaced the generic + icon with labeled icons (Photo, Document, Contact) visible in the attachment tray on first tap — no sub-menu navigation required.

04

Status update — profile-proximate entry

Added a status ring affordance to a contact's profile photo (consistent with Instagram/Snapchat mental model), with a tap-to-add-status CTA in your own profile.

05

Voice messages — onboarding tooltip

Added a one-time contextual tooltip explaining the lock-to-record gesture on first voice message attempt — dismissed after acknowledgment, visible to new users.

Statistical Results

The numbers confirmed the diagnosis.

Paired t-test results

Compared pre-redesign vs. post-redesign performance across all participants for each task. Four tasks reached significance at p < 0.05. The fifth (voice messages) showed improvement but did not reach significance — likely due to the small sample (n=6) and the confound of prior familiarity.

What statistical significance means here

With only 6 participants, the effect sizes needed to be large to clear the threshold — and they were. That's a meaningful signal: the usability failures were severe enough, and the redesigns targeted enough, that even a small sample returned confident results.

Impact summary slide: the finalized prototype enhanced user satisfaction by 21%, reduced task completion time and errors by 47%, and improved usability, visibility, and consistency.

A "significant improvement" without a p-value is just a feeling.

Reflection

What self-directed research teaches.

This project was a deliberate demonstration of full-cycle mixed-methods capability — from qualitative discovery to statistical validation.

The most important thing I learned wasn't about WhatsApp. It was about how the methods compound: the heuristic evaluation told me where to look; the performance testing told me how bad the problem was; the think-aloud told me why it happened; and the t-test told me whether the fix was real.

That sequence — discovery → measurement → diagnosis → intervention → validation — is the same one I bring to every client engagement. The platform changes. The discipline doesn't.