AI in retail: the product data was never clean

Retail has the clearest gap in this set between the use case that gets discussed and the one that works, and the gap is entirely about data. Personalisation and recommendation attract the budget and the board attention, and both sit directly on top of product data that is fragmented across hundreds of suppliers, several systems and twenty years of inconsistent entry. The same attribute is expressed six different ways, the same product appears under three identifiers, and nobody in the organisation holds the full picture. A recommendation engine built on that foundation produces confident nonsense, and the failure is reported upward as a model problem. The project that works first is the one nobody wants to fund: normalising and enriching the catalogue itself, which is unglamorous, measurable field by field, and unblocks every other item on the list behind it.

Retail, in short

Retail : the constraint, the distraction and where to start
The constraint Product data is fragmented across suppliers, systems and years of inconsistent entry, and the quality problem is upstream of anything a model can fix.
The distraction Personalisation, which is the loudest use case and depends on the data quality nobody has fixed yet.
A first project that works Product data normalisation and enrichment: unglamorous, measurable, and it unblocks everything else on the list.
Where ground truth lives Catalogue records already corrected by merchandising teams, which shows exactly what good looks like.

What shows up most often, not a description of any particular organisation. No named clients and no case studies: see editorial policy.

Why catalogue work is a good fit for current technology

It is the most reliable category of task available: reading unstructured text written by a human and producing structured fields. Supplier descriptions, specification sheets and product copy go in; normalised attributes come out.

The properties that make it work are the same ones that make any first project work. High volume, so the saving is real. Checkable output, so quality is a matter of looking rather than of opinion. And a merchandising team already doing it by hand, which means the ground truth exists and the people who can label are already employed.

It also degrades gracefully. A system unsure about an attribute can leave it blank and flag it, which is exactly what a human would do, and a blank field costs far less than a wrong one. Few tasks in other sectors have such a forgiving failure mode.

The heterogeneity problem, and who actually understands it

A catalogue assembled over two decades from hundreds of suppliers contains every convention anyone ever used, including ones the current team has never seen. Sizes in three unit systems. Colours as names, as codes and as supplier-specific strings. Categories that were reorganised twice and never backfilled.

The people who understand this are merchandisers, not technologists, and they hold it informally. They know that this supplier always puts the pack quantity in the description, that this category's weights were entered in grams until 2019, that these three brands share an identifier because of an acquisition.

None of it is documented. Extracting it is the substance of stage one on a retail project, and the technique is the same as everywhere: ask to see the last fifty items somebody fixed by hand, rather than asking how the data works.

Seasonality is a constraint on the calendar

Retail has periods when nothing may change. Peak trading, whether it falls in November or around a regional festival, freezes deployment entirely, and the freeze is usually longer than outsiders expect because it includes the run-up.

This is not an obstacle so much as a scheduling fact that determines project shape. A deployment that misses its window waits for the next one, which can be months. Projects here should be scoped to finish well before a freeze rather than up against it, and the freeze dates belong in the plan from week one.

There is an upside. The quiet period after peak is when merchandising teams have time to label, review and give the project attention, which is the scarce input. Aligning the evaluation work with that window rather than fighting for attention during peak is worth more than any amount of extra engineering capacity.

The returns and customer service layer

The second project that reliably works in retail sits in customer operations: classifying returns reasons, routing complaints, and extracting structure from the free text customers write. High volume, recoverable when wrong, and currently handled by people reading the same twenty patterns over and over.

It has a property the catalogue work does not: the output feeds directly back into merchandising. Knowing that a fifth of returns on a product line cite the same sizing issue is worth more to the business than the handling time saved, and that finding is invisible while the reasons sit in free text nobody aggregates.

What to do with the catalogue once it is fixed

The point of the unglamorous first project is that it makes the second one possible. Search that works, because the attributes are consistent. Recommendation that means something, because products can be compared. Supplier onboarding that takes hours rather than weeks.

Sequencing it this way also produces something politically useful: a visible, measurable win early, on a problem the merchandising team already knew they had. That buys the credibility to attempt the more interesting work, which is what a first deployment is actually for.

Questions people actually ask

Why is personalisation usually the wrong first project?

Because it sits on top of product data that is fragmented across suppliers, systems and years of inconsistent entry. A recommendation engine built on a catalogue where the same attribute is expressed six ways will produce confident nonsense, and the failure will be attributed to the model rather than to the data underneath it.

Is catalogue normalisation really an AI project?

It is one of the better fits available, because the work is reading unstructured supplier text and producing structured fields, which is the most reliable category of task with current technology. It is also measurable field by field, which means quality arguments are settled by looking rather than by opinion.

What makes retail data harder than it looks?

Scale multiplied by heterogeneity. A catalogue of hundreds of thousands of items assembled from hundreds of suppliers over two decades contains every convention anyone ever used, and nobody holds the full picture. The people who know are usually merchandisers, not technologists.

Where is the ground truth?

In catalogue records that merchandising teams have already corrected. Every retailer has a set of items that were fixed by hand because they mattered, and that set is exactly what good looks like. It is rarely thought of as a dataset and it is sitting in the product information system.

Read next

Sources

Radif Partners

Written and maintained by Radif Partners

Applied AI deployment practice · Forward deployed engineering

Covers 2026, · last reviewed 2026-09-24