AI in retail: the product data was never clean
Retail has the clearest gap in this set between the use case that gets discussed and the one that works, and the gap is entirely about data. Personalisation and recommendation attract the budget and the board attention, and both sit directly on top of product data that is fragmented across hundreds of suppliers, several systems and twenty years of inconsistent entry. The same attribute is expressed six different ways, the same product appears under three identifiers, and nobody in the organisation holds the full picture. A recommendation engine built on that foundation produces confident nonsense, and the failure is reported upward as a model problem. The project that works first is the one nobody wants to fund: normalising and enriching the catalogue itself, which is unglamorous, measurable field by field, and unblocks every other item on the list behind it.
Retail, in short
| The constraint | Product data is fragmented across suppliers, systems and years of inconsistent entry, and the quality problem is upstream of anything a model can fix. |
|---|---|
| The distraction | Personalisation, which is the loudest use case and depends on the data quality nobody has fixed yet. |
| A first project that works | Product data normalisation and enrichment: unglamorous, measurable, and it unblocks everything else on the list. |
| Where ground truth lives | Catalogue records already corrected by merchandising teams, which shows exactly what good looks like. |
What shows up most often, not a description of any particular organisation. No named clients and no case studies: see editorial policy.
Why catalogue work is a good fit for current technology
It is the most reliable category of task available: reading unstructured text written by a human and producing structured fields. Supplier descriptions, specification sheets and product copy go in; normalised attributes come out.
The properties that make it work are the same ones that make any first project work. High volume, so the saving is real. Checkable output, so quality is a matter of looking rather than of opinion. And a merchandising team already doing it by hand, which means the ground truth exists and the people who can label are already employed.
It also degrades gracefully. A system unsure about an attribute can leave it blank and flag it, which is exactly what a human would do, and a blank field costs far less than a wrong one. Few tasks in other sectors have such a forgiving failure mode.
The heterogeneity problem, and who actually understands it
A catalogue assembled over two decades from hundreds of suppliers contains every convention anyone ever used, including ones the current team has never seen. Sizes in three unit systems. Colours as names, as codes and as supplier-specific strings. Categories that were reorganised twice and never backfilled.
The people who understand this are merchandisers, not technologists, and they hold it informally. They know that this supplier always puts the pack quantity in the description, that this category's weights were entered in grams until 2019, that these three brands share an identifier because of an acquisition.
None of it is documented. Extracting it is the substance of stage one on a retail project, and the technique is the same as everywhere: ask to see the last fifty items somebody fixed by hand, rather than asking how the data works.
Seasonality is a constraint on the calendar
Retail has periods when nothing may change. Peak trading, whether it falls in November or around a regional festival, freezes deployment entirely, and the freeze is usually longer than outsiders expect because it includes the run-up.
This is not an obstacle so much as a scheduling fact that determines project shape. A deployment that misses its window waits for the next one, which can be months. Projects here should be scoped to finish well before a freeze rather than up against it, and the freeze dates belong in the plan from week one.
There is an upside. The quiet period after peak is when merchandising teams have time to label, review and give the project attention, which is the scarce input. Aligning the evaluation work with that window rather than fighting for attention during peak is worth more than any amount of extra engineering capacity.
The returns and customer service layer
The second project that reliably works in retail sits in customer operations: classifying returns reasons, routing complaints, and extracting structure from the free text customers write. High volume, recoverable when wrong, and currently handled by people reading the same twenty patterns over and over.
It has a property the catalogue work does not: the output feeds directly back into merchandising. Knowing that a fifth of returns on a product line cite the same sizing issue is worth more to the business than the handling time saved, and that finding is invisible while the reasons sit in free text nobody aggregates.
What to do with the catalogue once it is fixed
The point of the unglamorous first project is that it makes the second one possible. Search that works, because the attributes are consistent. Recommendation that means something, because products can be compared. Supplier onboarding that takes hours rather than weeks.
Sequencing it this way also produces something politically useful: a visible, measurable win early, on a problem the merchandising team already knew they had. That buys the credibility to attempt the more interesting work, which is what a first deployment is actually for.