There is no defensible universal interview target. The right sample depends on what the team needs to learn, how heterogeneous the market is, how much depth the decision requires, and whether the research has seriously searched for evidence that could overturn the emerging view.
A commercial-diligence team is three days from investment committee.
It has completed eight expert interviews. Seven respondents broadly support the market-growth thesis. The eighth is an outlier. The engagement manager wants four more calls. The partner asks the obvious question:
“Why four?”
Answers to that question are often less rigorous than the research itself.
Twelve is standard. Ten feels thin. Twenty should be enough. We usually do five per geography. The themes are starting to repeat.
Each answer offers the comfort of a number without establishing what the number is supposed to accomplish.
That is a problem because expert interviews can play very different roles in an engagement. A team may be trying to identify the major forces shaping a market, understand why customers behave as they do, test a specific causal explanation, discover exceptions to an apparent rule, estimate an uncertain quantity, or determine whether a finding travels across customer segments and geographies.
Those objectives do not require the same evidence.
Qualitative-methods research makes this point more sharply than consulting practice often does. In a study of 25 in-depth interviews, Hennink, Kaiser, and Marconi distinguished between code saturation, having identified the range of relevant issues, and meaning saturation, having developed a richer understanding of what those issues mean and how they operate. In their dataset, code saturation occurred after nine interviews; meaning saturation required 16 to 24.
That is not a prescription to conduct 24 calls.
It is a warning that “we are hearing the same themes” and “we understand the mechanism well enough to advise the client” are different claims.
For consulting teams working under severe time constraints, the right question is therefore not how many interviews constitute rigor.
It is:
What does this decision require the interviews to establish, and what observable condition would tell us that another interview is unlikely to change the answer?
The false precision of an interview target
The idea that there is a generally correct number of qualitative interviews persists partly because methodological studies have produced memorable numbers.
Guest, Bunce, and Johnson analyzed 60 interviews with female sex workers in Ghana and Nigeria and found that, in their particular study, most thematic saturation occurred within the first 12 interviews, with many high-level themes visible earlier. After 12 interviews, 92% of the codes ultimately identified in the 30 Ghana interviews had emerged.
The headline number has traveled much further than the qualification.
The authors explicitly cautioned against assuming that six to 12 interviews would always suffice. Their population was relatively homogeneous and their research objectives narrow. They noted that heterogeneous populations, diffuse questions, lower-quality data, and objectives involving variation between groups require more sampling.
A later systematic review by Hennink and Kaiser similarly found relatively modest interview counts could reach saturation in empirical studies, particularly when populations were homogeneous and study objectives narrowly defined. The authors' conclusion was not that researchers had discovered a universal number, but that sample-size guidance has to be interpreted alongside the characteristics of the study.
This matters enormously for consulting.
Twelve procurement managers buying the same product in one country may produce substantial repetition.
Twelve respondents split among four countries, three customer segments, two channel models, and regulators may represent almost no replication within any relevant subgroup.
The denominator is therefore not simply “interviews.”
It is interviews within the parts of the problem where the recommendation depends on variation.
A global study with 30 calls can be empirically thinner than a tightly framed study with eight.
First decide what the interviews are allowed to prove
Before determining the sample, the team should define the evidentiary job of the interviews.
This is where many diligence programs become confused.
Suppose eight of ten distributors say a competitor's product is gaining traction.
That observation can support a legitimate statement:
Most of the distributors we interviewed described increasing competitive traction.
It does not automatically support:
80% of distributors in the market believe the competitor is gaining share.
The interviews were not a probability sample. The respondents may have been selected because they were accessible, knowledgeable, recently active, known to the expert network, or especially relevant to the hypothesis. The sample may overrepresent large firms, former executives, urban customers, English speakers, successful operators, or people motivated to take expert calls.
Purposeful sampling is designed precisely to find information-rich cases, not to reproduce the statistical characteristics of a population. Palinkas and colleagues describe purposeful sampling as selecting cases because of their relevance to the phenomenon under study, with different strategies appropriate to different research objectives.
That makes expert interviews extremely useful for questions such as:
Why is this happening?
What mechanism might explain the data?
Which variables are insiders watching?
Where does the conventional explanation fail?
What differences exist across customer types?
Which hypothesis should we investigate next?
It makes them much less suited, by themselves, to estimating population prevalence.
This distinction should appear in the engagement team's evidence language.
Instead of saying “70% of experts believe X,” a qualitative study may be better summarized as: X was a recurring view across these respondent strata; Y was the principal alternative explanation; and Z remains unresolved.
The objective is not weaker evidence.
It is a more accurate description of what the evidence can support.
What kind of saturation does the decision require?
Once the evidentiary purpose is clear, saturation becomes more useful.
Consider a consulting team investigating why adoption of an industrial technology remains slow.
After seven interviews, it has heard the same five barriers repeatedly: capital cost, integration difficulty, shortage of technical staff, uncertain ROI, and procurement conservatism.
If the engagement question is simply “What are the main barriers?”, the team may be approaching code saturation.
But suppose the client needs to decide which barrier to attack with a new go-to-market model.
Now the research must establish considerably more.
Does “uncertain ROI” mean customers lack evidence of savings, cannot obtain budget approval, distrust vendors' calculations, or have payback requirements the product cannot satisfy?
Is integration difficulty a genuine technical limitation or a purchasing objection used to conceal another concern?
Does the answer change between new facilities and retrofit environments?
Who inside the customer organization can override procurement?
Under what circumstances have buyers adopted despite these barriers?
Those questions seek meaning, causal structure, and boundary conditions rather than another mention of the same theme.
Hennink and colleagues' distinction is useful because it prevents a common diligence failure: stopping when the labels are repeating even though the implications for the decision are not understood. Their empirical results were specific to one study, but the conceptual distinction is broadly applicable.
Malterud, Siersma, and Guassora offer another useful way to frame the problem. Their concept of information power proposes that the necessary sample becomes smaller when participants contain more information relevant to the study. They identify five factors affecting that information power: the specificity of the study aim, specificity of the sample, use of established theory, quality of interview dialogue, and analysis strategy.
Translated into consulting practice, five interviews with precisely selected respondents discussing a tightly specified question in depth may provide more decision-relevant information than 20 loosely matched calls built around a generic discussion guide.
The target should therefore move with the question.
Narrow question + homogeneous respondent universe + high- quality interviews = potentially smaller sample.
Broad question + heterogeneous market + multiple mechanisms + high consequence of missing an exception = larger and more deliberately structured sample.
The second case does not necessarily need dozens more respondents. It needs a better sampling design.
Design the sample before counting responses
The fastest route to superficial rigor is to pick a target, say 20 interviews, and fill it with whoever is easiest to recruit.
The more defensible approach is to construct the sample architecture first.
Suppose a consultancy is testing the attractiveness of a B2B software market across France, Germany, and the United Kingdom.
The thesis depends on three propositions:
- customers face the same underlying pain point across markets;
- buying authority is moving toward a central functional owner;
- willingness to pay is high enough to support the proposed model.
A generic target of 15 interviews says little. The sampling question is which perspectives could cause the team to reach materially different conclusions. The team might distinguish:
- existing customers versus prospects who declined to buy;
- enterprise versus mid-market buyers;
- technical users versus economic buyers;
- mature adopters versus recent adopters;
- countries with meaningfully different buying or regulatory contexts;
- companies that implemented successfully versus those where adoption stalled.
Maximum-variation sampling formalizes part of this intuition: select cases across dimensions on which meaningful differences are expected so the study can identify both variation and patterns that persist despite it. Methodological work on purposeful sampling contrasts this with homogeneous sampling, which intentionally narrows variation when the research question permits it.
The key is not to maximize diversity indiscriminately.
Every stratum consumes interviews. Split the sample into too many cells and the team ends with one respondent representing each supposed market segment: a taxonomy rather than evidence.
A useful rule is to create a respondent stratum only when the team can finish the sentence:
If this group sees the issue differently, our recommendation might change because…
If no meaningful decision consequence follows, the distinction probably does not deserve its own quota.
Independence matters before consensus
Sampling design also has a temporal dimension.
Teams often adapt interview guides as research progresses. That is sensible. Early interviews expose better questions.
But adaptation introduces another danger: later respondents can be recruited and questioned inside a frame created by the early respondents.
By interview eight, the team “knows” the market has three principal problems. Interviews nine through 15 then explore those problems in depth. Unsurprisingly, the synthesis becomes increasingly coherent.
The apparent saturation may partly reflect the research process narrowing itself.
Experimental work on collective judgment illustrates why independence is valuable. Lorenz and colleagues had 144 participants make quantitative estimates and found that exposure to others' estimates reduced diversity without reliably improving collective accuracy; it could also increase participants' confidence as their answers converged. The experiment involved simple estimation tasks rather than expert interviews, so it should not be treated as direct evidence about commercial diligence. Its narrower implication is useful: convergence created after exposure is not equivalent to independent convergence.
For important questions, consulting teams should therefore consider independent rounds.
Run an initial set of interviews against the same broad question before exposing respondents, or recruiters, to an emerging consensus where practical.
Then synthesize.
Use the next round deliberately to test gaps, alternative explanations, and respondent groups missing from the first.
RAND's review of expert-judgment elicitation makes a related recommendation in a very different technical context: use multiple experts where possible and include independent experts alongside people closely associated with the project being assessed. It also recommends explicitly pressing experts to consider reasons their uncertainty range may be wider than first stated.
The broader principle is valuable for diligence:
Do not let the first plausible explanation determine the entire remaining sample.
Search for negative cases before declaring saturation
A second danger arises when teams measure saturation by repetition alone.
Imagine 14 experts agree that adoption of a medical technology depends primarily on physician preference.
Interview 15 is a procurement executive from a large hospital system who explains that centralized contracting has removed most physicians' practical choice.
Has the 15th interview added only one minority view?
Or has it revealed that the original conclusion fails in the exact customer segment responsible for most of the client's expected growth?
Frequency alone cannot answer that question.
A negative case is valuable not because dissent is automatically correct but because it tests the boundaries of the emerging explanation.
The research team should ask:
What respondent would we most expect to disagree with our current synthesis?
Which customer tried the product and rejected it?
Which operator failed while peers succeeded?
Which country does not fit the apparent regional pattern?
Which former executive competed unsuccessfully under the strategy we are recommending?
Which member of the value chain loses if the apparent trend continues?
Purposeful-sampling methodologies explicitly include approaches built around extreme, deviant, and maximum-variation cases because unusual cases can reveal mechanisms or conditions hidden by typical ones.
This is particularly important in consulting because many recommendations are not statements about average experience.
They are conditional claims.
A market-entry recommendation may depend on a regulatory exception.
A commercial-diligence case may depend on the defensibility of one customer segment.
A pricing strategy may fail if an important buyer archetype has a fundamentally different procurement process.
The higher the cost of missing such an exception, the less comfortable the team should be with simple thematic repetition.
“Expert” is not a quality control
The composition problem does not disappear when every respondent has impressive credentials.
Expertise is multidimensional.
Someone may have tremendous industry seniority but limited visibility into current purchasing behavior. A former CEO may understand market structure but not today's technical workflow. A current operator may have excellent local knowledge but unusual incentives. An adviser may have broad exposure while receiving much of their information from the same industry narrative as every other adviser.
Structured-expert-judgment research developed in risk analysis takes this problem seriously enough to evaluate not merely whether people are labeled experts but, where possible, how well their probabilistic judgments are calibrated. Cooke's Classical Model, for example, explicitly distinguishes statistical accuracy from informativeness rather than assuming credentials guarantee forecasting quality.
Most consulting interviews cannot calibrate experts formally in this manner.
But teams can borrow the underlying skepticism.
Record why each respondent is in a position to know the answer.
Separate firsthand experience from secondhand opinion.
Ask what period their knowledge covers.
Look for commercial or professional incentives that could systematically affect the perspective.
Do not turn five people repeating the same industry talking point into five independent observations.
Interview quality is partly a sampling problem and partly a provenance problem.
A defensible stop rule for fast research
Formal saturation is difficult to establish inside a five-day diligence sprint.
The response should not be to abandon stopping rules. It should be to use one appropriate to the decision.
Francis and colleagues proposed one operational method for theory-based qualitative research: establish an initial analysis sample in advance, then specify how many further interviews must occur without new ideas before stopping. In their studies they used an initial set of ten and a stopping criterion of three subsequent interviews. Those specific numbers came from their research design, not a general law. The useful innovation was making the rule prospective and observable.
Consulting teams can adopt the principle without pretending to conduct an academic qualitative study.
Before interviews begin, define seven things.
1. Decision purpose
What decision is the interviews supposed to improve?
“Understand the market” is too vague. “Determine whether low adoption reflects implementation friction or weak customer economics” can guide a sample.
2. Respondent strata
Which groups have structurally different information relevant to the decision?
Do not add strata solely for cosmetic coverage.
3. Minimum diversity
Which perspectives must be represented before the evidence can be considered adequate?
This is a floor, not the final sample size.
4. Independent first round
Gather enough responses before narrowing the inquiry around the emerging consensus.
The objective is to preserve some independent discovery.
5. Saturation test
After each round, distinguish three questions:
- Are genuinely new themes still appearing?
- Are interviews still changing the team's understanding of the mechanisms behind important themes?
- Are material differences between strata still unexplained?
The first tests code saturation. The latter two move closer to meaning saturation.
6. Negative-case search
Before stopping, conduct a deliberate search for respondents and evidence that should be least compatible with the emerging conclusion. A research program that has only sampled likely confirmers has not earned saturation merely because they agree.
7. Decision stop rule
Stop when the required respondent strata have been covered; another independent round is producing little new decision-relevant information; the mechanisms material to the recommendation are understood to the required depth; credible negative cases have been actively sought; and remaining disagreements or sampling limitations are explicit enough for the decision maker to judge.
That last condition is important.
“Enough” does not mean uncertainty has disappeared.
It means the expected decision value of another interview has become lower than the cost, in time, money, or delayed action, of obtaining it.
For a reversible market-sizing discussion, that point may arrive quickly.
For a large acquisition whose thesis depends on subtle changes in customer behavior across countries, it should arrive later.
The downside of being wrong belongs in the sample-size decision even if no qualitative-methods paper can supply a formula for it.
More interviews can make the research worse
There is a final reason not to maximize sample size.
Every interview adds analysis burden.
A team conducting 40 calls in a week may have less time to interrogate contradictions, trace claims, compare respondent types, or change the next interview guide than a team conducting 15 carefully sequenced conversations.
Scale can produce false confidence if the synthesis process collapses nuanced evidence into theme counts.
This becomes particularly relevant as AI lowers the coordination and synthesis costs of interviewing.
Being able to collect 50 responses does not establish that 50 is methodologically superior to 15. It may allow the team to cover more strata, run an independent second round, or investigate minority cases that previously would have been cut for time. Those are genuine improvements.
Simply adding more similar respondents after the central themes are already understood is not.
The scarce resource shifts from access to interviews toward sampling judgment.
Where Enquire fits: use scale to improve sample design, not to substitute for it
Enquire's current product offering includes an AI-led interview engine designed to conduct expert interviews asynchronously, as well as text-based input from multiple vetted experts. Its broader research environment is designed to synthesize signals, identify gaps, and make differences across perspectives visible.
For consulting teams, the most defensible use of that additional capacity is not “we can interview more people.”
It is we can design a better evidence sequence within the same deadline.
A first round can cover independently selected respondent strata. Structured synthesis can show which themes repeat, which perspectives differ, and where important gaps remain. A second round can then target missing geographies, contradictory cases, or mechanisms that have not reached sufficient depth.
Asynchronous interviewing can also make it economically easier to include perspectives that a traditional schedule might exclude simply because arranging another call is inconvenient.
But the technology does not resolve the methodological questions.
A larger biased sample remains biased.
Repeated expert opinion does not become market prevalence.
AI synthesis cannot infer a missing respondent category that the research design never considered with complete reliability.
And a polished summary can obscure weak sampling as easily as a slide deck can.
Enquire's role is therefore most useful when additional interview capacity expands variation, independence, and negative-case search, while structured synthesis helps the team decide what it still does not know. Its current positioning explicitly emphasizes combining expert perspectives with structured research and surfacing where views align or diverge.
The methodological standard must still come from the engagement team.
Defend the sample with logic, not arithmetic
When the partner asks why the team needs four more calls, there should be a better answer than “because 12 is the target.”
Perhaps the first eight interviews have identified the major themes, but only existing customers have been sampled. Four former prospects are needed to test whether the team's explanation of lost sales survives outside the customer base.
Perhaps the research has reached theme saturation in Germany, but the recommendation assumes that the purchasing mechanism transfers to France and no French economic buyer has yet been interviewed.
Perhaps no more calls are needed. The remaining uncertainty concerns market prevalence, which should be answered with transaction data or a survey rather than another qualitative interview.
Or perhaps one interview could be more valuable than ten: the team has found a credible negative case capable of disproving the causal assumption that drives the recommendation.
Those are defensible sampling decisions.
A universal count is not.
The best expert-interview program is therefore not the one with the largest N or the neatest claim of saturation. It is the one that can show a clear chain from decision → research objective → respondent design → evidence gained → disconfirmation attempted → stopping rule.
Then “How many interviews are enough?” becomes answerable.
Enough to know what the major explanations are. Enough to understand the ones that matter. Enough variation to detect where they fail. And no more than the decision requires.
Sources and further reading
- Monique Hennink, Bonnie Kaiser, and Vincent Marconi, “Code Saturation Versus Meaning Saturation,” Qualitative Health Research journals.sagepub.com
- Greg Guest, Arwen Bunce, and Laura Johnson, “How Many Interviews Are Enough?,” Field Methods sfu.ca
- Monique Hennink and Bonnie Kaiser, “Sample Sizes for Saturation in Qualitative Research,” Social Science & Medicine sciencedirect.com
- Kirsti Malterud, Volkert Siersma, and Ann Dorrit Guassora, “Sample Size in Qualitative Interview Studies: Guided by Information Power” pubmed.ncbi.nlm.nih.gov
- Jill Francis et al., “What Is an Adequate Sample Size? Operationalising Data Saturation for Theory-Based Interview Studies” tandfonline.com
- Lawrence Palinkas et al., “Purposeful Sampling for Qualitative Data Collection and Analysis in Mixed Method Implementation Research” pubmed.ncbi.nlm.nih.gov
- Jan Lorenz et al., “How Social Influence Can Undermine the Wisdom of Crowd Effect,” PNAS pnas.org
- RAND, Subjective Probability Distribution Elicitation in Cost Risk Analysis: A Review rand.org
- Roger M. Cooke, “Supplementary Information for Structured Expert Judgment” rogermcooke.net
- Enquire, current product overview enquire.ai