← Back to blog

When the First Draft Is Cheap, Evidence Becomes Premium

ConsultingBy Enquire Team · February 12, 2026

AI is compressing retrieval and synthesis faster than it is eliminating uncertainty. Consulting firms should redesign delivery around verifiable evidence, field validation, and named accountability, not output volume.

A consulting team can now move from a broad question to a respectable first synthesis in minutes. Market structure, competitor positions, regulatory themes, likely strategic options: an AI system can assemble the analytical map before an analyst would once have finished collecting sources.

That is a genuine productivity gain. It also creates a less comfortable question.

If every credible firm can produce a competent first draft quickly, what is the scarce part of the work?

Increasingly, it is not the draft. It is establishing which parts of the draft deserve to be believed.

For consulting firms, that changes the logic of AI adoption. The objective should not be to automate the largest possible share of an engagement. It should be to automate the parts where output is inexpensive to check, then redirect scarce human effort toward the parts that determine confidence in the recommendation: framing the problem correctly, establishing provenance, finding disconfirming evidence, obtaining operating context that is absent from published material, and taking responsibility for the judgment that remains.

In other words, AI is making synthesis abundant. Evidence is becoming the bottleneck.

The first-draft advantage is disappearing

The productivity case for AI in knowledge work is already substantial. In a preregistered experiment developed with Boston Consulting Group, 758 knowledge workers using GPT-4 completed 12.2% more tasks and worked 25.1% faster on tasks judged to lie within the model's capability frontier, while producing higher-quality outputs. On a more complex task deliberately chosen to sit outside that frontier, however, participants using AI were less likely to reach the correct answer.

More recent evidence suggests the productivity effect can reach beyond individual drafting. A 2026 Organization Science field experiment with 791 Procter & Gamble professionals found that individuals using AI produced work of comparable quality to two-person teams without AI in the product-development tasks studied. AI also helped participants span functional expertise more effectively.

These findings should make consulting leaders optimistic about automating a large portion of the work that historically consumed junior capacity: searching, summarizing, structuring, generating alternatives, and producing a first analytical narrative.

But the BCG experiment contains a more consequential finding for delivery design. On the task where AI reduced accuracy, AI-assisted participants still produced recommendations that evaluators found more coherent and persuasive, even when the underlying answer was wrong.

That matters because consulting has traditionally relied on visible friction as one informal quality-control mechanism. Weak analysis often looked weak: gaps appeared in the logic, a junior team member struggled to defend a claim, or an incomplete workstream arrived visibly unfinished.

Generative AI can remove some of those warning signals. A questionable conclusion can now arrive polished, internally consistent, and boardroom-ready.

The cost of producing plausible analysis has therefore fallen faster than the cost of establishing its validity. That is the source of the evidence premium.

The jagged frontier is a workflow problem, not a job-description problem

It is tempting to translate current AI limitations into a reassuring division of labor: machines draft; humans judge.

That distinction is unlikely to survive.

The more useful lesson from current evaluations is that AI capability remains jagged. Performance varies materially across tasks that may look similar to the user, which makes occupational labels such as “consultant,” “researcher,” or “analyst” poor units for automation decisions.

The 2026 APEX-Agents benchmark, for example, tested agents on 480 long-horizon tasks across investment banking, management consulting, and corporate law. The highest overall first-attempt completion rate in the paper was 24%; the strongest reported result on management-consulting tasks was 22.7%. Multiple attempts improved performance, but inconsistency remained substantial.

Those percentages should not be treated as a forecast of how much consulting work AI can automate. APEX is deliberately constructed: 22 of its 33 work environments are entirely fictional, web search is disabled for reproducibility, and agents work without the accumulated context of a real engagement team. Its value is different. It demonstrates how quickly performance can degrade when work requires sustained planning, navigation across tools and files, and correct execution over many linked steps.

At the same time, capability is moving quickly.METR's task-horizon evaluations continue to find an exponential trend in how long a well-specified task frontier models can reliably complete, although METR stresses that its tests concentrate on software, machine learning, and cybersecurity and should not be read as direct measures of real occupational automation. It also notes that real jobs contain tacit context, interpersonal interaction, and success criteria that cannot be cleanly scored.

The implication is not to protect a permanent category of “human work.” It is to redesign engagements at the level of individual workflow steps.

A practical distinction is whether a step is observable, reversible, and consequential.

When the correct output can be cheaply checked, errors are easy to reverse, and the context can be made explicit, aggressive automation makes sense. When quality is difficult to observe, evidence is incomplete, or an error could redirect a major recommendation, the workflow needs stronger verification, additional evidence, or named human ownership.

That boundary should move as models improve.

Evidence becomes more valuable precisely because plausible synthesis becomes abundant

Most client decisions are not constrained by a total absence of information. They are constrained by uncertainty about what information means, whether it applies to the situation at hand, and which assumptions will survive contact with operating reality.

AI can summarize ten reports about a new market. That is different from establishing whether distributors actually behave as those reports imply.

It can identify the consensus view of an acquisition target's competitive advantage. That is different from finding the former operator who can explain the operational exception that invalidates the consensus.

It can generate six reasons a transformation program may fail. That is different from determining which failure mechanism is active inside the client's organization.

The highest-value evidence in these situations tends to possess one or more characteristics that generic synthesis cannot manufacture: provenance, specificity to the decision, independence from the prevailing narrative, or access to information that has not already been encoded into the public corpus.

Primary research is one way to obtain such evidence, but consulting firms should resist an equally simplistic conclusion: an expert interview is not automatically evidence merely because a human said something.

Qualitative-research methodology is instructive here. In one study of 25 in-depth interviews, Hennink, Kaiser, and Marconi distinguished between code saturation, having identified the range of issues being discussed, and meaning saturation, understanding the nuances and dimensions of those issues. In their particular sample, the broad themes emerged considerably earlier than the deeper conceptual understanding. The authors explicitly caution that saturation depends on study purpose, population, sampling strategy, data quality, and the type of issue being explored.

The consulting parallel is important. Three calls with well-chosen operators may be sufficient to expose the basic structure of a problem. They may be nowhere near sufficient to understand an ambiguous incentive, detect an edge case, or distinguish a sector-wide pattern from one executive's experience.

The goal of fieldwork therefore should not be “get more expert calls.” It should be to design an evidence portfolio around the decision: Who is in a position to know? Whose incentives or experience might systematically distort the answer? Which perspective is missing? What observation would disprove the working hypothesis?

That last question deserves more attention as AI increases analytical speed. MIT CISR's research on organizational learning makes a related argument: teams learn faster when they articulate and test explicit hypotheses rather than merely generating solutions and hoping to learn from failure.

AI makes hypothesis generation inexpensive. Firms should spend some of the saved effort making hypothesis testing more demanding.

A four-part division of labor for AI-enabled engagements

For most consulting workflows today, the useful split is neither “AI does analysis” nor “humans stay in the loop.” It is a four-part operating model.

1. Automate retrieval and the first analytical map.

AI should take an increasingly large role in searching authorized sources, extracting relevant information, comparing documents, summarizing established positions, building issue trees, generating preliminary hypotheses, and identifying obvious gaps.

The output should be treated as a map of what to investigate, not as accumulated proof.

This distinction prevents an engagement team from wasting expensive human time reconstructing information that a machine can assemble efficiently, while avoiding the equally costly mistake of treating synthetic fluency as confidence.

2. Instrument verification instead of adding generic human review.

“Have a manager check it” is not a scalable control system. As the volume of AI-generated work expands, undifferentiated review simply moves the bottleneck upward.

Verification should instead be embedded in the workflow. Material claims should carry source provenance. High-consequence assertions should be separated from low-consequence descriptive material. Contradictory evidence should be surfaced rather than averaged away. Assertions that cannot be tied to an adequate source should remain explicitly unresolved.

This is consistent with the NIST Generative AI Risk Management Profile, which recommends documenting upstream data sources and content provenance and assessing output accuracy and reliability against known ground truth using multiple evaluation methods. BCG similarly argues that organizations redesigning work around agentic AI need testing, monitoring, evaluation, and auditability built into the operating model rather than added after deployment.

The standard should become: the more effortlessly a claim was produced, the easier its evidentiary trail should be to inspect.

3. Invest field research where it can change the answer.

AI-generated research should make primary research more targeted, not eliminate it.

Before an expert conversation, the team should already know what it believes, why it believes it, and which unresolved assumptions matter to the recommendation. Interviews can then be designed to discriminate between competing explanations.

That changes recruitment, too. The most useful respondent may not be the person most likely to confirm the current synthesis. It may be the operator from the failed implementation, the former buyer who rejected the category, the regulator with a different interpretation, or the executive whose market behaves differently from the apparent norm.

Field research becomes valuable when it introduces new information or disconfirmation, not when it merely adds quotations to a conclusion already reached.

4. Preserve named human accountability for framing and tradeoffs.

A recommendation is more than the sum of verified facts.

Someone has to decide what problem the engagement is actually solving, which uncertainties matter, what level of evidence is sufficient, which stakeholder tradeoffs are acceptable, and what action should follow when the evidence remains incomplete.

For high-stakes work, that responsibility should be explicit. The relevant partner or engagement leader should be able to answer not only “What do we recommend?” but “Which assumptions does this recommendation depend on, which evidence could overturn it, and why are we comfortable acting before the remaining uncertainty is resolved?”

Human accountability matters here not because machines are constitutionally incapable of judgment. It matters because advice creates consequences, and governance requires an identifiable owner of the decision standard.

Measure confidence gained, not content produced

Consulting firms can easily measure the wrong benefits from AI.

Tokens generated, research hours saved, slides produced per analyst, and time to first draft all capture activity. They say relatively little about whether an engagement reached a better-supported answer.

That measurement gap is visible in adjacent professional services. Thomson Reuters'2026 AI in Professional Services Report, based on more than 1,500 professionals in legal, tax, accounting, risk, fraud, and government, not management consulting, found that 40% said their organizations were using generative AI, while only 18% said their organizations tracked AI return on investment.

Consulting firms should supplement productivity metrics with a small set of confidence metrics:

  • Material-claim support: What share of claims that could change the recommendation have adequate, inspectable evidence?
  • Critical-assumption coverage: Which high-impact assumptions have been tested against independent or primary evidence rather than repeated across secondary sources?
  • Disconfirmation yield: Did the research process actively surface evidence against the preferred hypothesis, and what changed because of it?
  • Late-stage correction: How much work is being reopened during partner review or client challenge because an earlier factual or interpretive error survived the workflow?

The target is not zero uncertainty. That would make consulting impossibly slow.

The target is to know where uncertainty remains and whether the decision justifies taking it.

Where Enquire fits: use synthesis to decide what deserves field evidence

A research system can support this model when it connects synthetic analysis to evidence collection rather than treating them as substitutes.

Enquire's current positioning combines structured AI research with expert perspective, and its product offering includes text-based input from vetted experts, AI-led asynchronous interviews, and research context that can persist across successive inquiries.

The useful workflow is straightforward. AI research produces the initial analytical landscape: major explanations, known evidence, disagreements, and missing information. Those gaps then determine where operator or specialist perspective is worth acquiring. New field evidence is fed back into the synthesis, where it can confirm, qualify, or overturn the starting view.

Neither side of that loop is self-validating. AI-generated synthesis still requires provenance and verification. Expert statements still require appropriate sampling, interpretation, and triangulation.

The value lies in making those two forms of research inform each other faster.

The boundary will keep moving

Any workflow designed in 2026 should assume that today's allocation of work will become obsolete.

The evidence already points in both directions. Current systems remain inconsistent on some long-horizon professional tasks. They can also match or exceed human performance on increasingly consequential bounded tasks, and evaluations such as METR's indicate continued rapid capability improvement.

Verification will become more automatable. Primary research will become easier to conduct and analyze automatically. Models will gain more organizational context. Some judgments that seem irreducibly human today will become routine machine work.

That does not weaken the operating principle. It strengthens it.

Consulting firms should continuously move automation toward tasks whose correctness can be observed and whose failure can be contained, while moving scarce professional attention toward whatever remains difficult to verify, expensive to misunderstand, and consequential to the client.

The strategic question is therefore not how many analyst hours AI can remove from an engagement.

It is more demanding:

Where does the confidence behind our recommendation actually come from, and who owns it?

Sources and further reading

  1. Dell'Acqua et al., “Navigating the Jagged Technological Frontier,” Organization Science hbs.edu
  2. Dell'Acqua et al., “The Cybernetic Teammate,” Organization Science pubsonline.informs.org
  3. Vidgen et al., “APEX-Agents” arxiv.org
  4. METR, “Task-Completion Time Horizons of Frontier AI Models” metr.org
  5. Vaccaro, Almaatouq, and Malone, “When combinations of humans and AI are useful,” Nature Human Behaviour nature.com
  6. Hennink, Kaiser, and Marconi, “Code Saturation Versus Meaning Saturation” journals.sagepub.com
  7. NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile nist.gov
  8. BCG, “Reinventing the Operating System of Work with AI” bcg.com
  9. Thomson Reuters Institute, 2026 AI in Professional Services Report thomsonreuters.com
  10. Enquire, current product overview enquire.ai

See how Enquire fits your workflow.