AI in ecommerce operations

How AI changes D2C research without replacing judgement

AI can widen retrieval, organise evidence and expose unanswered questions. An experienced operator still owns the question, the sources and the decision.

A human hand sorts evidence cards while a subtle machine-analysis grid remains in the background.

Key takeaways

Delegate the gather: competitor and concept research, transcripts, first-pass variants and option scoring. Keep test selection, numerical truth, customer language and publication approval with the accountable operator.

AI changes D2C research by making retrieval and organisation cheaper. It does not make judgement optional.

A research model can search across more material than one operator will read in an afternoon. It can cluster repeated claims, compare definitions, pull candidate quotations and show where an outline lacks evidence. It can also state a false fact confidently, misread a source, flatten disagreement or attach a citation that does not support the sentence beside it.

Split the work by accountability. Give the model tasks that can be inspected and repeated. Keep the choices that determine what the business believes and does with people who can defend them.

For ecommerce research, that means a person owns the commercial question, source rules, claim boundaries, first-party context and final recommendation. AI helps build the evidence surface underneath those decisions.

What do I delegate to AI?

Delegate the gather:

  • Competitor and concept research.
  • Transcripts and extraction from source material.
  • First-pass variants.
  • Scoring a stack of options against an approved rubric.

These tasks supply raw concepts. Strategy decides which concept deserves the first test. A clean score or confident summary is not the decision; it is material for the person accountable for the result.

Delegate the gather. I sign off what goes out under my name.
Eddie Cheng, Ecommerce operator

Two boundaries are absolute in this workflow. AI does not introduce a new figure. Every number must come from an approved source, a supplied company record or a calculation a reviewer can reproduce. AI-generated customer language is hypothesis language. It remains test-only until a real person has said it.

That distinction protects the customer voice from being quietly manufactured. A model can propose a plausible objection or phrase, but the research record must label it as an AI hypothesis. Interviews, reviews, support conversations or another genuine customer source are what move language into the evidence set.

Which parts of research become faster?

Retrieval is the obvious gain. A model with web or document access can search several formulations of a question, follow related terms and assemble a working source list. That is useful when the question crosses media, commerce, finance and operations.

It can also help with document triage. Instead of reading every source in full before knowing whether it is relevant, the researcher can ask for the sections that define a metric, document a product behaviour or state a limitation. The original source still needs inspection before its claim enters an article or decision.

Other suitable tasks include:

  • Turning a long document into narrow, linked claim notes.
  • Comparing how two sources define the same term.
  • Finding dates, markets and populations that make two figures incomparable.
  • Grouping search results by the question they answer.
  • Listing claims in a draft that lack a source.
  • Generating alternate queries for an unresolved part of the brief.
  • Converting an approved source pack into an outline.

OpenAI's deep research documentation describes a plan, source selection and multi-source synthesis process with links in the final report. Those links make review faster because the reader can move from the summary to the material behind it.

The links make the report traceable, but the source still has to support the claim.

Which decisions should the operator keep?

Keep any decision where a plausible but wrong answer would change what the business does.

That includes the research question itself. "How do we scale Meta?" is too broad to produce a defensible answer. "Can this business increase Meta spend next month without breaching its contribution and stock constraints?" points towards the required evidence and the people who own it.

The operator should also decide:

  • Which sources are acceptable for each kind of claim.
  • Whether a platform document, independent test or internal record has the right scope.
  • What uncertainty is material.
  • Which company data may be used and how it is protected.
  • Whether the evidence is current enough for the decision.
  • What first-party experience changes the interpretation.
  • Which recommendation follows, if any.

NIST's AI Risk Management Framework Core calls for explicit task scope, defined roles and documented human oversight in human-AI configurations. Those ideas translate cleanly into a small research desk. Name the job the model performs, the evidence it may use, the person who reviews it and the condition that blocks the output.

Why is a cited answer still not a verified answer?

A citation answers "where might this have come from?" Verification asks a longer set of questions.

Open the source and check:

  1. Does it contain the claimed fact?
  2. Does the surrounding passage change the meaning?
  3. Is the source describing the same product, market and period?
  4. Is it the original source or another summary?
  5. Does it separate observation from recommendation?
  6. Has a later update replaced the guidance?

OpenAI's own accuracy guidance says language models can produce incorrect facts, fabricated quotations or references and overconfident answers. It tells readers to verify important information against reliable sources.

The NIST Generative AI Profile treats confident false content and false citations as confabulation risks. A polished report can therefore require more scrutiny, not less, because weak claims are easier to accept when the presentation feels complete.

What does a dependable workflow look like?

Separate discovery, evidence and judgement. Do not ask one prompt to complete all three invisibly.

  1. Write the commercial question and the decision it will inform.
  2. Set source rules, date range, market and known exclusions.
  3. Use AI to discover and triage candidate material.
  4. Open the original sources and record narrow claim notes.
  5. Build an outline in which every factual section names its evidence.
  6. Add approved first-party context and state what remains uncertain.
  7. Run a claim audit, then require a named human approval.

Each stage should leave an artefact. The search stage leaves queries and candidate URLs. The evidence stage leaves claim notes. The outline maps sections to sources. The review records which claims changed or were removed.

This makes correction cheap. If a source is withdrawn, the editor can find the claims that depend on it. If the question changes, the team can see which evidence still applies. If an answer cannot be supported, it disappears instead of being softened into vague copy.

The source pack behind this article follows that pattern. It records what each source can support, what it cannot support and which operating claims came directly from Eddie rather than being inferred from public material.

How should you write the research brief?

The brief should constrain the work enough to make failure visible.

Include:

  • The decision and audience.
  • The exact question.
  • Geographic and time scope.
  • Approved source types and named exclusions.
  • Definitions that must remain consistent.
  • Claims the model may not make.
  • Required first-party input.
  • The output format and review owner.
  • The condition for returning an error instead of an answer.

That last condition matters. A model is usually rewarded for producing something. The workflow needs permission to stop when evidence is missing.

For example, a brief about increasing Meta spend can ask for current Meta delivery guidance, commerce inventory definitions and a pre-scale framework. It cannot ask public sources to reveal the brand's contribution limit or stock appetite. Those inputs belong to the operator.

Similarly, a brief about ROAS and profit can map how platforms define conversion value and how commerce reports define net sales. The company still has to choose its own contribution-profit contract.

How do you judge source quality?

Source quality depends on the claim.

A platform's documentation is usually the best source for how its own setting currently behaves. It is not independent evidence that the setting will improve every advertiser's result. A commerce help centre can define a report field. It cannot set the company's commercial target. A vendor case study may describe a real outcome, but the vendor selected the customer, measure and presentation.

Use a source ladder rather than one blanket rule:

  • Primary product or regulatory documentation for current behaviour and rules.
  • Original datasets or research papers for empirical claims.
  • Company financial and operational records for this business.
  • Reputable analysis for interpretation, with its incentives visible.
  • Community discussion for questions and failure reports, not settled facts.

The model can label these categories. A person decides whether the source is suitable for the sentence and the consequence attached to it.

Do not confuse recency with authority. A recent post may repeat an old error. Do not confuse authority with relevance either. An official document for another market can be accurate and still wrong for the decision in front of you.

Where does first-party experience enter?

It enters after public evidence has defined the common ground and before the article settles its recommendation.

The operator can add:

  • The constraint that public guides usually omit.
  • A decision sequence used in practice.
  • An anonymised example with approved details.
  • A result that contradicted the expected explanation.
  • A boundary where the advice stops working.

This contribution should be recorded before it is polished. Keep the original answer, the edited version and where it appears. Do not turn hesitation into certainty or attribute generic advice to Eddie simply because it suits the article.

First-party experience has its own limits. One account or brand example does not establish a market rule. Present it as experience, not a benchmark.

What failure modes should the review catch?

The most obvious failure is a false claim. Several quieter failures can do as much damage.

The model may use a source that mentions the topic but does not support the sentence. It may combine two definitions into one. It may omit a date that makes a figure look current. It may treat platform attribution as incremental revenue. It may summarise a qualified recommendation as a rule. It may find five sources that all repeat the same original claim and present them as independent agreement.

A final claim audit should mark each material statement as:

  • Directly supported by a named source.
  • A calculation from supplied data.
  • An inference, labelled as such.
  • Approved first-party experience.
  • Unsupported and removed.

The audit should also search for invented clients, precise numbers without provenance, quotation marks, stale product instructions and links that redirect to a different page.

The review should preserve disagreement too. If two credible sources define the measure differently, do not ask the model to blend them into one smooth answer. Record both definitions, decide which one applies to the current work and explain the choice. If the evidence remains mixed, the brief should say so.

Uncertainty needs an owner. The researcher decides whether a gap blocks publication, narrows the claim or becomes a question for the operator. The model can identify the gap, but it cannot decide how much commercial risk the company should accept because a source is incomplete.

Keep rejected claims in the working log with a short reason. That prevents the same weak statement returning during a later rewrite and gives the reviewer a view of what the research could not establish.

How do you measure whether AI improved the work?

Do not measure success by the amount of text produced.

Useful operating measures include research time to an approved brief, percentage of material claims with verified sources, unsupported claims caught before review, source diversity, correction time and the number of briefs returned for missing evidence.

Track human effort too. A workflow that creates a long report and transfers all verification to an editor may be slower than a focused manual search. Faster discovery only matters if the evidence becomes easier to trust and reuse.

The review log should record what the model did well and where it failed. Those notes improve the next brief more than saving a giant prompt with no explanation.

How would this work on a live operator question?

Take a question such as whether an ecommerce brand should increase paid-media spend before a product launch.

A person first defines the decision: the proposed increase, market, period and commercial boundary. The model can then search current platform guidance, inventory-report definitions and measurement documentation. It can organise those sources under delivery, stock, event quality and profit, then list the facts still missing from the business.

The missing facts are likely to include available stock by important variant, product contribution, cash timing, current creative supply and the acceptable downside. Public research cannot answer them. The operator or the company's records must.

After those inputs arrive, the model can draft scenarios and identify where two assumptions conflict. A person checks the calculations and decides whether the scenario is plausible. If the recommendation changes spend, the decision record keeps the inputs, source links and named approver.

This split avoids two bad extremes. One is asking the model for a confident budget answer with no company data. The other is refusing assistance for the mechanical work of locating platform rules and comparing report definitions.

AI can gather and arrange the public material while the operator supplies private context and makes the call. If a source changes or the stock position moves, the team can update the affected part without rerunning an opaque conversation from memory.

What should never be delegated silently?

Do not silently delegate the decision, the source threshold or the final claim.

AI can widen the desk. It can find terms an operator did not think to search, compare documents and turn a messy folder into a usable starting point. The operator still decides which question deserves time and whether the answer changes spend, stock, pricing or customer communication.

The publish gate belongs to a person for the same reason. I sign off what goes out under my name. That approval covers the figures, the status of customer language, the chosen test and the uncertainty that remains.

Use the model where inspection is cheap. Keep judgement where the cost of being wrong belongs to the business.

Useful answers

Questions operators ask

Can AI do ecommerce market research?
AI can help retrieve, compare and organise public material, especially when it has access to current sources. A person must still define the question, inspect the sources and judge whether the evidence supports the decision.
Are citations from an AI research tool reliable?
Citations make a claim traceable, but they do not make it correct. Open the source, confirm that it says what the report claims and check whether the source is suitable for the question.
Which D2C research tasks are suitable for AI?
Delegate competitor and concept research, transcripts, first-pass variants and scoring a stack of options. Treat the outputs as raw material that remains open to inspection.
Which research decisions should remain human?
Keep judgement over which concept deserves a test, whether a number is real, whether language came from a customer or from AI, and whether the work ships.
How do you stop AI research from inventing facts?
Restrict it to approved sources, require claim-level links and reject unsupported material. AI does not introduce a new figure, and AI-generated customer language remains a test hypothesis until a real person has said it.

About the author

Eddie Cheng

Eddie Cheng founded Penang Media and co-owns VIBAe. He writes from the agency and brand sides of ecommerce growth, connecting paid acquisition with stock, margins, cash flow and contribution profit.

More from Eddie Cheng

The operating context

Growth from the agency and brand sides.

Eddie Cheng writes about profit-first ecommerce growth from both sides of the work: Penang Media, the performance agency he founded, and VIBAe, the footwear brand he co-owns. His articles connect paid acquisition with stock, margins, cash flow and the decisions that determine profitable growth.

About the publication