Skip to content
All posts
PLAYBOOKJul 27, 2026·6 min read

AI in Ecommerce: What's Actually Working vs. What's Just Hype

Most AI pilots show zero measurable return within six months. One category is a real exception. Here's how to tell the difference before you invest.

The Suggesto team
Title card reading Hype vs. Reality, beside two panels labelled mostly noise and one job, done well.

Open any ecommerce newsletter this year and you'll find the same word doing a lot of work: AI. AI-powered pricing, AI product descriptions, AI customer service, AI everything. It's hard to tell which of it actually moves revenue and which is just a feature label bolted onto something that already existed.

A study published this year gives an unusually honest answer, and it isn't the one most vendors are selling.

The uncomfortable number

MIT's Media Lab, through its Project NANDA initiative, published a report in July 2025 called The GenAI Divide: State of AI in Business 2025. After analysing around 300 real AI deployments and surveying hundreds of executives and employees, the researchers found that 95% of enterprise generative AI pilots showed no measurable impact on profit or loss within six months. Only about 5% delivered real, rapid value.

That's not "AI doesn't work." It's a specific, narrower finding: most pilots, measured against a strict six-month P&L bar, aren't proving their worth yet.

Worth knowing before you quote it: the study's own methodology has been publicly challenged. Paul Roetzer, founder of the Marketing AI Institute, argued the six-month, P&L-only definition of success ignores real value AI can deliver in other forms, efficiency gains, cost reduction, faster workflows, and pointed out the "zero return" finding rests on a relatively small sample of interviews rather than official company reporting. Both things can be true at once: a lot of AI pilots really are underperforming, and a strict, narrow measurement window makes the failure rate look starker than it might otherwise be.

Chart showing 95 percent of AI pilots deliver no measurable return versus 5 percent that succeed
MIT's 2025 research found 95% of enterprise AI pilots showed no measurable P&L impact within six months.

The category that's the exception

Here's what makes the picture more interesting than a flat "AI is overhyped" headline. Not all AI is the same bet.

Personalization, a narrower and more mature category than generative AI broadly, has real, repeated, independently documented evidence behind it. McKinsey's own research has consistently found that 71% of consumers expect personalized interactions from the brands they buy from, and 76% get frustrated when that personalization is missing. Where it's implemented well, McKinsey found it typically drives a 10 to 15% revenue lift. Separately, McKinsey's research on promotions found 65% of customers name targeted, relevant offers as a top reason they actually make a purchase.

That's a meaningfully different picture from a generic AI chatbot pilot with no clear job to do. Personalization succeeds where it does because it solves one specific, measurable problem: matching the right thing to the right person. It isn't a strategy. It's a narrow tool aimed at a narrow job, and narrow tools aimed at narrow jobs are exactly the kind of AI that tends to show up in the working 5%, not the stalled 95%.

Diagram comparing broad unfocused AI pilots with a narrow single-purpose AI use case
A narrow, measurable job tends to succeed where a vague, broad AI initiative stalls.

How to tell which side of the divide you're looking at

Before investing time or budget in an "AI-powered" anything, it's worth asking a few blunt questions:

  • What specific job is this doing? "AI-powered insights" is vague. "Recommends the right product from a customer's answers to three questions" is specific. Specific jobs get measured. Vague ones don't.
  • Can you measure it within weeks, not quarters? The pilots that stall tend to be the ones with no clear, near-term signal of whether they're working.
  • Is this replacing a task you already understand, or inventing a new one? Tools that augment something you already do well (recommending products, writing tighter subject lines) tend to outperform tools built around a capability nobody asked for.

Full disclosure, since it's directly relevant here: this is the category Suggesto sits in. The app uses AI to auto-tag a merchant's catalogue and help build product-recommendation quizzes, a narrow job with a clear measurable output, not a general-purpose AI layer bolted onto everything. That's not a coincidence. It's exactly the kind of scoped, single-job use case the evidence actually supports.

Key takeaways

The honest story isn't "AI works" or "AI is hype." It's that AI is not one bet. Broad, ambitious generative AI pilots are mostly struggling to prove measurable value within a short window, while narrower, more mature applications like personalization have real, independently verified revenue evidence behind them.

Before adopting anything with "AI" in the name, ask what specific job it's doing and how soon you'll know if it worked. The tools solving one clear problem tend to be the ones actually paying off.

Frequently Asked Questions

Is AI actually helping ecommerce businesses right now?

It depends heavily on what kind of AI. Broad generative AI pilots have mostly struggled to show measurable financial return within six months, according to MIT research. Narrower applications like personalized recommendations have much stronger, independently documented evidence behind them.

Why do so many AI pilots fail to show ROI?

MIT's research points to a mix of factors: budgets often go toward flashy sales and marketing pilots rather than higher-ROI back-office uses, and many tools are built to look good in a demo rather than adapt to a business's actual workflow and data.

Should small Shopify merchants invest in AI tools?

Selectively, yes. The evidence favours narrow, well-scoped tools that do one measurable job, like personalized product recommendations, over broad "AI-powered" platforms with no specific, near-term outcome to point to.

What's the difference between AI hype and AI that works?

Hype tends to be vague about the specific job being done and hard to measure quickly. Tools that work tend to solve one clear, narrow problem with a result you can measure within weeks.

References

  • MIT Media Lab, Project NANDA. The GenAI Divide: State of AI in Business 2025 (July 2025).
  • Marketing AI Institute. Critique of MIT GenAI Divide methodology.
  • McKinsey & Company. The Value of Getting Personalization Right, or Wrong, Is Multiplying.
  • McKinsey & Company. Unlocking the Next Frontier of Personalized Marketing.