Back to blog

The confirmation bias trap in user research

Kalle·

Richard Feynman used his 1974 Caltech commencement address to make a blunt point about scientific integrity: the easiest person to fool is yourself. He was talking about physics. He could have been talking about your last five customer interviews.

Here is the uncomfortable version for founders. You had the idea. You wrote the interview questions. You chose the participants. You ran the calls. You took the notes. You decided which quotes mattered. At every step, the person filtering the evidence was also the person who desperately wants the idea to be good.

That is how user research turns into a machine for manufacturing agreement.

This article is about how that machine works, why it runs hardest when the stakes are highest, and what you can do about it. Because “just be objective” is not a process. It is a hope.

What confirmation bias means in user research

In his widely cited 1998 review, psychologist Raymond Nickerson defines confirmation bias as seeking or interpreting evidence in a way that favors an existing belief, expectation, or hypothesis. The two verbs matter.

Seeking: you look for evidence in places where your belief is likely to survive.

Interpreting: when mixed evidence arrives anyway, you read it in the most friendly possible way.

Most founders notice the first problem and underestimate the second. They know a leading question can contaminate a call. They do not always notice what happens later, when a lukewarm answer becomes a positive signal in the notes, or when one enthusiastic quote starts carrying the whole product decision.

The classic demonstration is Peter Wason’s 1960 2-4-6 task. Participants saw a number sequence, 2, 4, 6, and were told it fit a rule the experimenter had in mind. They could propose more triples, get yes-or-no feedback, and announce the rule when ready. The actual rule was broad: any ascending sequence. But many participants formed a narrower hypothesis, such as even numbers increasing by two, and then tested examples like 8, 10, 12 or 20, 22, 24. Those examples passed, so confidence rose. Only 6 of 29 subjects reached the correct conclusion without first announcing an incorrect one.

The failed participants were not lazy. They ran tests. The problem was that their tests were designed to confirm the rule they already had in mind, not to break it.

Klayman and Ha later reframed this as a positive test strategy: people often test cases they expect to fit. In ordinary life, that can be a useful shortcut. In early product research, it can be disastrous, because many explanations fit the same transcript. “Customers have this problem and would pay for my product” and “customers are being polite to a founder” can produce nearly identical notes unless your interview process is designed to separate them.

That is why confirmation bias is so expensive in customer discovery. It does not stop you from collecting evidence. It makes the wrong evidence feel rigorous.

Why founders get the strongest version

Everyone has confirmation bias. Founders doing their own research have the compounding version, because the belief being tested is not abstract. It has your time, reputation, identity, and runway attached to it.

Nielsen Norman Group’s guide to confirmation bias in UX makes the point plainly: the more invested you are in an assumption about a design or user, the stronger the bias becomes. Their example is a designer who spent months on a design versus one who spent two days on a paper prototype. The two-day designer can hear bad news. The two-month designer has more to protect.

Now scale that up. You did not spend two months on a design. You quit a job. You told your partner this would work. You recruited a co-founder. You have a runway number in a spreadsheet. You are not evaluating evidence from neutral ground.

Jeanette Mellinger, Head of UX Research at BetterUp and former Head of UXR for Uber Eats, calls this the founder’s “happy ears” problem in First Round Review. The more you have at stake, the easier it is to hear the answer you want. With the right framing, almost anything can be interpreted as a green flag.

The hardest part is that more interviews do not automatically fix this. In the 1979 Lord, Ross, and Lepper study on biased assimilation, supporters and opponents of capital punishment saw the same mixed body of evidence. Both sides found the evidence that agreed with them more convincing, scrutinized the opposing evidence more harshly, and became more polarized.

Translate that to product work: twenty contaminated interviews can be worse than five, because now the original belief has a larger quote bank behind it.

The stakes are not theoretical. CB Insights’ March 2026 analysis of 431 VC-backed startup shutdowns found that poor product-market fit appeared in 43% of identifiable failure reasons. That does not mean better interviews would have saved every company. It does mean that talking to customers is not enough. The learning process has to be capable of telling you that your idea is wrong.

What the trap looks like from the inside

Confirmation bias rarely feels like bias. From the inside it feels like diligence. These are the five forms that show up most often in founder interviews.

The question already contains the answer

Nielsen Norman Group gives a simple example: an ecommerce team believes the checkout button is the problem, then asks whether the red checkout button was difficult to locate. The question plants the button, plants the difficulty, and closes off every other explanation.

The founder version sounds like this:

Biased questionBetter question
“Would an AI interviewer help you get more honest feedback?”“Tell me about the last time you needed feedback and worried people were holding back.”
“Is onboarding too complicated?”“Walk me through the last time you invited a teammate.”
“Would you pay for this?”“What have you paid for recently to solve this problem?”
“Do you think the dashboard is useful?”“What did you do the last time you needed this information?”

The better questions are not magic. They simply move the conversation away from your preferred answer and back into the participant’s actual life.

For more examples, use Maren’s guide to avoiding leading questions.

Compliments become evidence

Rob Fitzpatrick’s The Mom Test is useful because it treats compliments, hypotheticals, and feature wishlists as weak data. That is exactly the failure mode founders want to ignore. A participant saying “this sounds great” may be kind, interested, or mildly curious. None of those prove they will change behaviour.

The research mistake is not receiving the compliment. It is converting the compliment into roadmap confidence.

When praise appears, slow down. Ask what specific past event made the idea feel relevant. Ask what they currently do instead. Ask what would keep them from switching. If the answer cannot be attached to a real situation, treat it as social lubrication, not evidence.

The notes warm up the transcript

One participant says, “I might look at it if we had the budget.” The summary says, “Interested in paid plan.”

Three people give lukewarm non-answers. The synthesis says, “Mixed but generally positive.”

Nobody has to be dishonest for this to happen. Nickerson’s review emphasizes that confirmation bias is often unwitting. The person doing the analysis is not inventing evidence. They are rounding ambiguity toward the belief they already hold.

This is why raw transcripts matter. If you cannot trace a theme back to the exact participant language that supports it, you are no longer synthesizing. You are narrating.

Ambiguity gets a friendly interpretation

“I’d probably use this” is not a yes. It is a sentence about an imagined future, said to a person who wants the idea to work.

Teresa Torres argues that customer interviews should be grounded in specific past behaviour because general questions invite summaries and speculation. Ask about the last time the problem happened. Ask what the person did, what it cost, who got involved, and what changed afterwards. Past behaviour can contradict your hypothesis. Polite speculation rarely does.

This is the same reason Maren’s guide to observed versus reported behaviour treats opinions as leads, not conclusions.

The calendar says “validation”

If your calendar invite says “validation calls”, the conclusion is already leaning forward.

Nielsen Norman Group’s article on why teams should stop saying they “validate” designs is not just vocabulary policing. The word primes the team to seek proof and primes participants to treat the thing as nearly finished. The healthier frame is test, study, examine, or learn.

There is nothing wrong with validation interviews when you have a specific artefact to test. The problem is a validation mindset with no chance of invalidation. If the interview cannot produce an answer that changes your plan, it is not research. It is ceremony.

For the distinction, see Maren’s guide to discovery vs validation interviews.

How to design interviews that can prove you wrong

You cannot remove confirmation bias by deciding to be unbiased. You reduce its surface area by changing the process.

Write the belief down before the call

Before you recruit anyone, write three sentences:

  1. “We believe...”
  2. “We would be more confident if...”
  3. “We would be less confident if...”

The third sentence is the important one. If you cannot write it, you are not ready to interview. You have built a process that can only return encouragement.

Steve Portigal’s Interviewing Users treats research planning, interview technique, documentation, and synthesis as one system. That matters here because bias can enter at any stage. A neutral interview guide does not help if the analysis later upgrades every ambiguous answer.

Erika Hall’s Just Enough Research frames good research as asking better questions, thinking critically about answers, reducing unknowns, and spotting blind spots. In founder terms, that means aiming research at the assumption that would hurt most if it were false, not the assumption you are most excited to confirm.

Ask for stories, not opinions

Your research question might be “do founders trust AI-led interviews?” Do not ask that directly.

Ask:

  • “Tell me about the last time you needed feedback from users.”
  • “Who did you ask?”
  • “What did you avoid asking?”
  • “Where did you suspect people were being polite?”
  • “What did you do with the answers afterwards?”

The answer to your research question lives inside the story. A specific incident gives you sequence, context, stakes, workarounds, and consequences. An opinion gives you a tidy self-image.

This is also the practical overlap between confirmation-bias prevention and The Mom Test: keep the conversation about the participant’s life, not your idea.

Probe praise and criticism with the same energy

When a participant says something positive, do not relax. Treat it as the moment to get more precise.

Ask:

  • “What specifically makes that feel useful?”
  • “When would you have used it last month?”
  • “What would stop you from trying it?”
  • “What would you keep doing the old way?”

When a participant says something negative, do not defend. Ask the same kind of follow-up:

  • “Tell me more about that.”
  • “Where did that happen last time?”
  • “What did you do next?”
  • “How big a problem was it compared with everything else that week?”

Mellinger’s “ask the counter” tactic is useful because it makes balance operational. If you ask how something was useful, also ask how it was frustrating. If you ask what worked, ask what did not. Confirmation bias thrives in asymmetric follow-up.

Triangulate before you conclude

One flattering evidence source is easy to bend. Three independent sources are harder to bend in the same direction.

Nielsen Norman Group recommends triangulating research findings with sources such as analytics, support logs, usability testing, and customer-service data. For a founder, that might look like:

  • Interviewees say onboarding is fine, but 60% of invited teammates never activate.
  • Customers say the feature is important, but usage logs show it is touched once and abandoned.
  • A buyer praises the concept, but will not introduce the actual admin who owns the workflow.
  • A churned user says the price was the issue, but support history shows unresolved setup failures.

The point is not that behavioural data is always right and interviews are always wrong. The point is that disagreement between sources is where learning starts.

Take yourself out of the room when the answer matters

The cleanest way to reduce founder confirmation bias is to remove the founder from the interview.

That can mean a colleague with no stake in the feature. It can mean an outside researcher. It can mean an asynchronous interview setup where participants do not have to soften their answer to the person who built the thing.

This is not a moral judgement about founders. It is an instrument problem. If the instrument changes the measurement, you either calibrate it hard or use another instrument.

A ten-minute audit of your last five interviews

If you want to know whether this applies to you, do not introspect. Audit.

Open the transcripts or notes from your last five customer conversations and check four things.

1. Count questions that contained your answer. Watch for “do you”, “would you”, “is it easy”, “was it frustrating”, and any adjective doing your arguing for you. A yes-or-no answer is often a sign that you narrowed the evidence too early.

2. Count minutes spent on your idea versus their life. If the participant reacted to your product for 25 of 30 minutes, you ran a demo with a survey attached. Discovery evidence lives in their workflow, last incident, workaround, budget, and decision process.

3. Find the most negative thing each person said. Then check what happened next. Did you probe it, or did you explain, clarify, and move on? The move-on is the tell.

4. Compare the transcript temperature to the summary temperature. If “I might look at it” became “interested”, your analysis is warming the data after the call.

This audit is small, but it exposes the central question: did the interview process give the truth a path to reach you?

Where Maren fits

Maren is built for founders who need disciplined customer learning without putting founder hope in every conversation.

Maren runs structured, asynchronous interviews that ask for recent stories, probe vague answers, and keep pushing for the concrete event underneath the opinion. When a participant says they would use something, Maren does not round that up to a yes. Maren asks what happened last time, what they did instead, and what would make the new behaviour worth the switch.

The output is not a pile of compliments. It is a set of themes tied back to what participants actually said, including the friction, hesitation, and contradictions a founder might unconsciously filter out.

That does not replace judgement. You still choose the segment, define the research goal, and decide what evidence is strong enough. Maren’s job is narrower and more valuable: run the conversation in a way that gives disconfirming evidence a fair chance to appear.

However you run the research, the standard is the same. A good interview process should not merely help you feel more confident. It should help you find out you are wrong while it is still cheap to be wrong.

The 2-4-6 rule was not “even numbers increasing by two.” The market’s rule is probably not your first hypothesis either. The only way to find out is to test the triple that might break it.

Tell Maren what you want to learn

Try Pro for 30 days. Your first interview can be live in five minutes. No credit card required.