Chris Robeck · Research Practice
Virtual Portfolio Review

ChrisRobeck

Senior UX Researcher

San Francisco · Available for Senior / Staff UXR Roles

Originally From All Over

"My parents said 'be whatever you want.'
So I went to clown school."

2009–2014 · UMD
Clown, Director, Videographer. Certified. Actually.
HorrorQuest
Built a live-action Halloween game. It got too successful. Only questionably allowed.
Karen Holtzblatt
A professor caught me. "Expelled… or work for me." Took the deal. Connected to the inventor of Contextual Design.
Ensu → Meta
Mental health startup, $0→$1M. Then the big ones called.
Off the Clock in SF

📷 Photography

Volunteer body-image work

🎲 D&D Campaign

Homebrew world-building with an obscene amount of AI. Built my own Agentic Harness just for my campaign.

🎭 LARP / Event Design

SAGA Court of Seasons — immersive events for SF clubs

Today's Agenda

Two Case Studies · One Research Practice

Case Study 1
Meta
Creator Identity: 2% D30 retention, 11% sharing
Case Study 2
Ensu
Retention Crisis: 7% → 23% D30 retention
Case Study 1 · Meta

Creator Identity: From contradictory signals to a validated strategic framework

A mixed-methods approach to finding the missing variable behind creator-tool adoption.

Demonstrating: Deep, grounded research craft and technical skill

Situation

Creation volume rose while native tool use fell; sampled content revealed repost inflation.

Task

Identify a segmentation model that explained creator needs and could guide product investment.

Action

Natural-context research, journal study, cross-functional synthesis, NLP clustering, and N500 validation.

Result

20 hypotheses became 12 validated identities, a 4/6/2 strategy, and 2% young-adult retention.

Case Study 1 Roadmap

How the story moves from a misleading metric to a validated strategy

Situation
Behavioral Signal

Creation rose while native tool use fell; sampled content exposed repost inflation.

Task
Strategic Questions

Find the missing variable and make it useful for segmentation and investment.

Action
Mixed Methods

Natural-context research, AI-assisted diary synthesis, human audit, and N500 validation design.

Result
Framework → Strategy

Validation produced 12 identities, revealed how they stack, and unified three research streams.

Situation · Behavioral Signal

The dashboard measured output, not creator intent.

Creation rose. Native tool adoption fell. The top-line metric looked healthy; the behaviors underneath pointed in the opposite direction.

+12%
Weekly posts increased across the eight-week trend
−2.5–11.8%
All five tracked native tools declined (trim, text, voiceover, remix, templates)

Why it mattered: More content did not mean stronger ownership of creation. Reels risked becoming the distribution endpoint — not the tool of choice.

8-week weekly trend · Sep–Oct 2024

Situation

The volume metric hid what people were actually doing

~95%
of sampled videos were edited or posted on another platform first

Creation was rising — but Reels was often the distribution endpoint, not the creation tool.

Sampled content review

Situation · Research Activation

I created awareness and empathy…
(made a nuisance of myself)

A structured Data Science pull made the pattern credible; a gamified stand-up ritual built shared intuition.

01 · Validate

Frame "Ours — or theirs?" Lock the eligible Reels population and observation window before viewing.

02 · Engage

Watch one Reel and lock in two guesses before the reveal. Closest calls win small prizes.

Guess 1 · Origin: Created with our tools — or first made on another platform?

Guess 2 · Popularity: More popular on Reels — or on the original platform?

Guessing exposed assumptions; the reveal gave the team a shared pattern language.

Task

We had parity — so why didn't we have stickiness?

The platform was broad but underspecialized. Strategic goal: Tailor native tools to outcompete on a small number of strategically valuable creator segments.

Question 1

Which segments have the highest strategic value?

Inputs: Viewership analytics, platform consumption. Analysis: Cross-referencing output against audience demand.

Question 2

Which segments can we realistically convert?

Inputs: Switching behavior, competitive use vs. TikTok/Shorts. Analysis: Friction and habit strength.

Question 3

Which tools matter to each winnable segment?

Inputs: Feature requirements for filters, effects, cuts, audio syncing. Analysis: Tool adoption to creator intent.

Action

Conventional segments failed to predict tool choice

Existing variables like demographics and content genre showed almost no correlation with feature preference. A 20-year-old comedy creator and a 40-year-old comedy creator often needed entirely different tools.

R² = 0.04
Age vs. Tool
No predictive power
R² = 0.07
Region vs. Tool
No predictive power
R² = 0.11
Genre vs. Tool
Weak correlation
R² = 0.09
Followers vs. Tool
No predictive power

The pivot: We needed to find the missing variable that actually drove creation behavior.

Action · Method & Execution

Diary Study → Contextual Inquiry

Instead of assigning a lab task, the protocol followed naturally occurring creation moments and reconstructed the decisions behind them.

Purposeful Sample · N=10

✦ Active creators — at least 1 post per week
✦ Convertible maturity — mostly beginning–mid stage
✦ Platform spread — varied creation ecosystems
✦ Domain diversity — mixed creator categories
✦ Demographic mix — women, men, varied profiles

Study Cadence

Week 1
7-Day Diary — Capture natural creation
Participants documented real, self-initiated creation on whichever platform they chose.
Week 2
Contextual Interviews
Interviews used each participant's own Week 1 artifacts as concrete prompts. Replay, probe, reconstruct.

Longitudinal capture reduced recall bias and preserved the context surrounding each creation decision.

Action · Analysis

AI proposed the structure; researchers retained judgment

The workflow accelerated a quote-level affinity analysis while keeping interpretation, category boundaries, and final synthesis under researcher control.

01
Build the corpus — Combined diary entries, post-creation videos, and interview transcripts. Kept participant and source context attached.
02
NotebookLM assist — First-pass synthesis — Generated participant-level summaries. Surfaced recurring concepts and candidate patterns for researcher review.
03
Agentic Miro plugin — Quote-level placement — Converted quotes into digital post-its. Proposed placement into the team's predefined research categories.
04
Human audit — Interpretation & synthesis — Researchers reviewed every cluster, re-sorted misplaced evidence, challenged weak groupings. Merged, split, and refined into the final affinity structure.

Decision rule: AI accelerated organization; researchers owned interpretation, exceptions, and the final category model.

Action · Findings

Creator identity — not genre — explained tool needs

The natural creation record showed that creators organized their work around what they contributed, not which tool they used or category they occupied. We had to decouple the creator from their creation.

01 · Contribution came before production

Creators described presenting knowledge, talent, perspective, or a subject; tool operation was secondary to that intent. "Presenting," not "making."

02 · Genre concealed shared motivations

A gaming creator and a beauty creator could share the same presentation identity — even when their content categories looked unrelated.

03 · Segmentation needed a new unit

We proposed decoupling who the creator believed they were from the genre of the content they produced.

Action · Survey Design

Planned Missingness

N=500 active creators. Each person answers 40 questions out of 100 — planned missingness lets us build the complete picture from incomplete individual views.

500
Active creators · ≥3 Reels in last 30 days
40
Questions per person · 8 assigned elements, ~10 min
60%
Planned missingness — by design, not attrition

All profiling from platform metadata. All tool usage from behavioral logs. Zero self-report beyond identity items. Selection stratified by genre, follower tier, frequency, and platform.

Result · Survey Analysis

From Partial Data to Validated Framework

STEP 1
Recover Covariance
FIML method: partial data from 500 people → complete 100×100 correlation matrix
STEP 2
Blind EFA
Promax rotation (oblique) — creators hold avg 2.4 identities. Conservative loading threshold 0.50 (standard is 0.40)
STEP 3
Result Funnel
8 didn't survive — proof the test was real. Remaining 12 form the final framework.
GENRE → TOOL ADOPTION
R² = 0.03–0.07
✗ Does not predict
IDENTITY ELEMENTS → TOOL ADOPTION
Significant
✓ This is why identity replaced genre
Result · The 12 Validated Identities

The validation pipeline turned identity into a measurable framework

ë
My Knowledge
Text persistence, data cards, chapter markers
💬
My Opinions
Direct-to-camera, green screen overlays
My Skill
Speed ramps, zoom, multi-clip sequencing
🐾
My Pet
Audio dubbing, trending sound sync
🏠
My Lifestyle
Fast editing capture, minimal unpolished
👥
My Community
Duet, stitch, reply-to-comment
My Look
Beauty filters, lighting adjustments
✈️
My Travel
Stabilization, location tagging
📹
My Subject
Object tracking, high-res export
🛤️
My Journey
Before/after templates, transitions
🎨
My Art
Time-lapse, precise color grading
✍️
My Writing
Structure, argument, text overlays
"I'm presenting what I know about how to do this specific thing. The video is just the delivery method."
Skincare / beauty creator · 47K — Participant P3
"I need a stat to stay on screen for eight seconds while I explain it. So I'm in CapCut building text layers like it's a PowerPoint — for a 60-second Reel."
Personal finance creator · 112K — Participant P7

Identities don't exist in isolation — creators mix and match. Avg 2.4 identities per creator. Tool adoption isn't driven by a single label.

Result · Strategic Decision

Leadership used the synthesis to make a strategic prioritization decision

4
Focus
Knowledge · Opinions · Skill · Pet
Actively recruit and build dedicated tools
6
Neutral
Look · Subject · Lifestyle · Community · Writing · Art
Maintain support, don't prioritize new development
2
Discourage
Travel · Journey
Deliberately deprioritize; stop building edge-case tools
Result · Measurable Outcomes

The framework broke down silos and drove measurable original creation

+2%
D30 Retention among young adults
+11%
Sharing & Following across key segments
>90% → ~40%
Educational & Skill repost drop

🗣️ Common Strategic Language — Smoothed conversations between Research, Marketing, and Outreach teams that historically didn't collaborate closely.

🔄 Cross-Functional Model — The marketing partnership became the internal model for cross-functional work.

📈 Creator Recruitment & Retention — By building the exact tools these identities needed, repost drops were massive.

Result · Reflection

Rigor means challenging your own method

A framework is only as strong as the blind spots you actively hunt for.

📊

Leaner Survey Design

500 was generous and 8 elements per person meant 40 questions when 5 would have worked. Next time: leaner on both axes, put savings toward qual follow-up.

🤝

Force Co-Ownership

I would have brought Data Science into the qualitative coding phases so they understood the nuance behind the variables they were factoring.

⚖️

Trust the Qualitative Signal

When the dashboard contradicted behavior, we spent too long trying to fix the query before going out to talk to users.

Case Study 2 · Ensu

Retention Crisis: From a collapsing metric to a new onboarding architecture

A triangulated approach to diagnosing and fixing a product failure under extreme time pressure.

Demonstrating: Systems thinking, data triangulation, and intelligent decision-making frameworks

Situation

D30 retention collapsed from 20% to 7% while four variables changed at once.

Task

Isolate the root cause quickly and rigorously without lowering the evidence standard.

Action

Cohort tracing, rapid survey, and guided activity interviews with churned and new users.

Result

Replaced clinical forced-choice with guided discovery, recovering retention to 23%.

Case Study 2 Roadmap

How the story moves from a collapsing metric to a user-designed onboarding architecture

Situation
The Retention Cliff

D30 retention collapsed to 7% while four variables changed at once.

Task
12 Research Questions

Isolate the root cause quickly and rigorously without lowering the evidence standard.

Action
4-Phase Triangulation

Cohort tracing, rapid survey, guided activity interviews with churned and new users, co-design sessions.

Result
Recovery → Infrastructure

Recovered retention to 23%, then turned the crisis response into always-on research infrastructure.

Situation

The Retention Cliff

Ensu — ML-powered mood tracking, linking music to mental health. Six months in, the metric that determined whether the company could keep building collapsed.

20% → 7%
D30 retention · 65% decline

⚠ If we cannot diagnose and fix this drop, we lose funding.

📢
Marketing
New TikTok and Discord campaigns drove a large user spike
📅
Seasonality
Finals season; most users were college students
New Features
Two major product features launched simultaneously
🏴
Competitor
A direct competitor had started pushing hard

Startup resources were limited. I could not chase all four — I needed to narrow fast.

Task · Research Architecture

Stakeholder timelines shaped the research architecture

Twelve research questions, sequenced so each phase could only be designed after the previous phase answered its questions.

ENGINEERING: FINDINGS NEEDED IN 16 DAYS
MARKETING: STRATEGY DECISION NEEDED FIRST
PHASE 01
Understanding the Drop
Q1–Q3: Temporal + behavioral scope. When does divergence first appear?
PHASE 02
Diagnosing the Cause
Q4–Q6: External vs. internal. Is churn a failure of discovery, satisfaction, or relevance?
PHASE 03
Understanding the Mechanism
Q7–Q9: Mental model + behavior. How does directed vs. self-directed flow shape exploration?
PHASE 04
Designing the Intervention
Q10–Q12: Language + shippable model. What can ship in one engineering week?

Each phase answered a subset of questions — and determined whether the next phase was needed.

Action · Phase 1: Mixpanel Cohort Analysis

Three Signals in the Data

Are users taking different actions during week 1?
NO
Are the distributions of those actions different?
YES
Is the order they take actions different?
YES
Are any orders/actions correlated to long-term retention?
YES

Cohort-Specific

Retention decline isolated to users who onboarded after a specific mid-month date. Earlier cohorts completely unaffected — same product, same features, different outcome.

Same Actions, Different Sequence

New-cohort users were doing the same things — but engagement concentrated heavily in one feature instead of spread across four. The signal wasn't in what they did — it was the pattern.

Correlated with Retention

The order users engaged with features predicted whether they stayed. Certain paths led to retention; others led to churn. Specific sequences were reliably associated.

"I could see the pattern shift. I couldn't see why it shifted. Was this because marketing brought in different people? Or because the product was changing the experience?"
That question — external vs. internal — became the trigger for Phase 2.
Action · Phase 2: Rapid Churn Survey

The Right Users, the Wrong Assumptions

~600
Users targeted · In-app survey triggered as they exhibited churn behavior
137
Complete responses · 22.8% response rate

The Audience Did Shift

Survey demographics revealed a measurably different user profile from the new marketing channels. Not the "wrong" users — but less mental health vocabulary and domain familiarity.

Near-Zero Feature Awareness

When asked "Which features are you aware of?" — churned users averaged 1.3 vs. retained users at 3.4. They weren't rejecting features. They didn't know they existed.

External Causes Eliminated

✗ Competitor: 5% mentioned alternatives. ✗ Seasonality: No academic calendar correlation. ✗ Marketing quality: Same intent and interest level. The problem was confirmed internal.

We knew WHO had shifted. We knew they weren't finding features. We knew it was the product, not the market. But we didn't know the mechanism — HOW was the product failing them?

Action · Phase 3: Guided Activity Interviews

How the Product Was Failing

12 users, 60 minutes each. Two groups, two protocols — because churned users and new users held different kinds of evidence.

GROUP A · 6 PARTICIPANTS
Churned Users
10m — Warm-up & Context
20m — Retrospective Walkthrough
20m — Usage Reconstruction
10m — Moment of Decline
GROUP B · 6 PARTICIPANTS
New Users
10m — Warm-up & Context
20m — Peer Translation Exercise
20m — Live Onboarding (think-aloud)
10m — Reflection & Comparison

Why two protocols: churned users could tell me what happened. New users could show me what happens. Together, they triangulated the mechanism.

Action · Phase 3 Analysis

The Silent Failure Loop

Four findings from observation and retrospection that surveys and analytics couldn't reach.

1. The Preference Paradox

Every user said they preferred the new lateral layout's freedom. But behavior showed the opposite: they chose one feature and never explored further. The directed flow had created a habit loop. The lateral flow broke it.

You can't get this from a survey (stated preference) or analytics (behavior). You need both side by side.

2. Clinical Intimidation

Watching live onboarding revealed the exact moment language became a barrier. Terms like "sentiment" and "cognitive patterns" caused visible hesitation. The audience shift + clinical language = users retreating to whatever felt safest.

3. First-Feature Anchoring

The first feature didn't just become their favorite — it became their concept of the entire app. "I thought it was a mood tracker." The mental model narrowed at first contact and never expanded.

4. The Satisfied Exit

Churned users consistently described feeling finished, not frustrated. High satisfaction with their primary feature. Zero awareness there was more. They weren't leaving the app — they were leaving what they thought the app was.

The most dangerous kind of churn: no complaints, no friction, no signal in support tickets.

The product had created a silent failure loop: unfamiliar language pushed new users toward one safe feature , the lateral layout let them stay there , and they left satisfied they'd seen it all . No complaints. Just a quiet exit.

Action · Phase 4: Co-Design Sessions

What Users Designed

5 sessions · 2 participants each · 90 minutes. Each session paired one mental-health-aware user with one user without prior vocabulary — the same audience split the product now needed to serve.

The Vocabulary

Clinical (Before)User Language (After)
"Toolbox""Compass"
"Sentiment""Mood Predictions"
"Cognitive Patterns""Your Patterns"
"Therapeutic Playlist""Your Mood Mix"
"Mental Health Assessment""Quick Check-in"

Design Philosophy

Guided Intent: Replaced the lateral dashboard with a step-by-step emotional sequence.

🔗 Integrated Outcome: Connected mood-tracking directly to the playlist generation.

"More horoscope than medical chart."
Co-design participant's design philosophy
Result

The Third Model

12 days
Problem identification → engineering handoff
1 week
Engineering implementation

Not a revert. Not a tweak. A new onboarding architecture built from user-generated design.

OLD MODEL ✗
Directed · Rigid · Clinical
CURRENT MODEL ✗
Open · Unstructured · Overwhelming
THIRD MODEL ✓
Guided Discovery · Structured Agency

Guided Discovery — Start with one clear proposition, then reveal paths as confidence grows.

User-Generated Language — Clinical terminology replaced with phrases participants used to explain the product to each other.

Progressive Disclosure — Features unlock as users engage, reducing first-session overload while ensuring ecosystem discovery.

Only one week of engineering effort — because the research was precise about the cause and users had already designed the solution.

Result · Impact

Impact & Reflection

D30 RETENTION RECOVERY
7% → 23%
Exceeded the original 20% baseline by 3 points

Rapid Execution: 12 days of research, 1 week of engineering.

💰 Business Goal: Made the investor deadline with a clear recovery narrative.

🎯 Strategic Shift: The "Unaware" persona became a permanent part of the user model.

🔄 Lasting Practice: Usability test battery became a standing weekly practice.

What I'd Do Differently
📊

Better demographic data earlier

I would have identified the new "Unaware" persona sooner if we had captured better context during the initial acquisition spike.

📹

Proactive onboarding recordings

We should have been watching session recordings before the crisis hit, rather than using them only as a post-mortem diagnostic tool.

⚖️

Less binary elimination

Under extreme time pressure, I closed off some research threads too quickly when early data seemed clear. I needed to maintain a wider lens.

Result · Democratizing Research

Making Sure It Never Happens Again

"The crisis was solved in 12 days. The systems I built after lasted the life of the company."

Automated Retention Monitoring

Built automated Mixpanel dashboards tracking cohort retention curves in real-time. Configured threshold alerts: if D7 or D14 dropped below baseline, the team was notified immediately.

Standing Churn Survey

Took the Phase 2 rapid churn survey and made it permanent with an automated trigger. Responses flow into a live dashboard. Quarterly review of questions to keep them relevant.

Weekly Usability Test Battery

Repeatable battery tied to the 8 key conversion points from the Mixpanel funnel. Ran weekly with rotating participants from the 2,000-member community.

Self-Serve Research Dashboard

Aggregated retention metrics, churn survey, usability battery, and community feedback. Designed for non-researchers: plain language, visual summaries, trend indicators.

Before the CrisisAfter Operationalization
Research was project-based — someone had to ask for itResearch was always-on — data flowed continuously
Problems were discovered reactivelyProblems were caught proactively — alerts fired before crisis
Only the researcher could interpret user healthAnyone on the team could check the dashboard
Usability was tested when scheduledUsability was tested every week, automatically
Churn reasons required a new studyChurn reasons were continuously collected

⏱ Time from signal to team awareness: Weeks → Hours

Automated systems were still running when I left Ensu.

Two Stories. One Belief.

Great research doesn't just answer questions — it accelerates decisions.

"I believe research is only as good as the decisions it changes. But the best research changes how decisions get made — permanently."

Ready for the Next Chapter

🔍

The Discover

Building robust, user-trusted AI experiences at unprecedented scale.

DynaGuard

Merging distinct user mental models and architectures for next-gen guardrails.

📊

Continuous Integration

Empowering cross-functional teams to act on democratized, rigorous insights.

Download Full Case Study PDF ↓

chrisrobeck.com · San Francisco