Every marketplace operator eventually confronts the same uncomfortable fact: you are judged on the quality of a catalog you do not control.
Your search relevance, your conversion rate, your return rate and your paid acquisition efficiency all depend on product data. That data arrives from hundreds or thousands of independent sellers, each with their own systems, their own formats, their own definition of what a complete product record looks like, and no particular incentive to care about yours.
This is not a data problem. It is a structural one. And most of the standard solutions to it are trade-offs dressed up as strategies.
Dispersion is the starting condition
A single retailer with a bad catalog has a project. A marketplace with a bad catalog has an operating model.
The retailer can standardise. They control the source, they can pick a taxonomy, and they can push a cleanup through their own team. A marketplace can do none of that unilaterally. The same product arrives from one seller as a spreadsheet with eleven populated fields, from another as an XML feed built for a different channel entirely, and from a third as a folder of photographs and a price list.
Everything downstream — search, filtering, recommendations, comparison, buy-box logic — assumes a homogeneous catalog. Nothing upstream produces one.
Then multiply by taxonomy
The dispersion problem would be manageable if there were one target format. There isn't.
A marketplace spanning sixty categories isn't running one normalisation problem. It's running sixty of them. Footwear needs size systems, materials and fit. Power tools need voltage, torque, battery compatibility and safety certifications. Wine needs vintage, region, grape and alcohol content. These aren't variations on a schema — they're different schemas, each with its own required attributes, its own units, its own controlled vocabularies and its own idea of what makes a listing complete.
Expanding into a new category is therefore never just a commercial decision. It's a data modelling project, and it lands on whoever owns catalog operations, usually late and usually without additional headcount.
The trade-off everyone makes
Faced with this, marketplaces choose a point on a spectrum.
At one end: push the work to sellers. Strict templates, mandatory attributes, validation on upload, rejection of anything non-conforming. The catalog stays clean and the marketplace's costs stay low.
The cost shows up somewhere else. Onboarding friction is a seller acquisition problem. A merchant with genuinely interesting inventory and no digital product data — which describes a large share of physical retail — simply cannot complete your onboarding. They don't complain. They just don't list. Every hour you add to seller onboarding is an hour spent on your marketplace instead of a competitor's, and sellers are acutely aware of the difference.
At the other end: absorb the work internally. A merchandising team receives whatever sellers send and turns it into catalog. Sellers love you. Your operating costs scale linearly with your seller count, which is the opposite of what a marketplace business model is supposed to do.
Most operators end up somewhere in the middle and feel like they've compromised badly in both directions. They have.
Templates don't eliminate the internal team
Here's the part that surprises people who haven't run marketplace operations: even at the strictest end of the spectrum, you still need the merchandising team.
Template compliance is not data correctness. A seller can populate every mandatory field, pass every validation rule, and still produce a listing that is wrong. The material says "leather" because that was the first dropdown option. The category is close enough to pass but wrong enough to break filtering. The dimensions are in the wrong unit. The images technically meet the specification and show the product from an angle that tells a buyer nothing.
Validation catches structure. It does not catch meaning. So the internal team stays, doing QA on data that already passed automated checks — which means you've imposed the friction cost on sellers and kept the headcount cost internally.
The cost nobody puts on the spreadsheet
Onboarding delay is usually measured in operational terms: tickets, throughput, backlog. The real cost is commercial, and it's larger.
Products have selling windows. Seasonal goods, trend-driven categories, new model releases, promotional periods. A product that takes six weeks to list into a twelve-week season has lost half its revenue potential before it ever appears in search. A product listed after the launch buzz has passed enters a market where the price has already moved against it.
That loss doesn't appear in an operations report. It appears as inventory that didn't sell, sellers who churned, and categories that underperformed for reasons that get attributed to demand.
Deduplication is a marketplace-only problem
Independent retailers never face this one. Marketplaces face it constantly.
Five sellers list the same product. You now have five titles, five descriptions, five attribute sets with conflicting values, and five product pages competing with each other in your own search results. Reviews fragment. Ranking signals split. The buyer sees what looks like five different products and has no way to compare offers on the same item — which is the single most valuable thing a marketplace can offer.
Getting this right requires matching records across sellers, resolving conflicts between them, and building a canonical product that sits above all the offers. Then doing the same for variants — the same shoe in nine sizes and three colours, arriving from four sellers who each model variants differently, or don't model them at all.
This is genuinely hard, it requires seeing across the whole catalog at once, and it cannot be delegated to sellers. No individual seller has the visibility to solve it.
And then it changes every day
Product data is relatively stable. Price and availability are not.
Conflating the two is one of the more common architectural mistakes in marketplace operations. They have completely different cadences: descriptions and attributes change rarely, prices and stock change hourly. If your pipeline treats them as one dataset, you either process everything constantly at enormous cost, or you update slowly and accumulate the two failure modes that damage marketplaces most — selling something you don't have, and showing a price you won't honour.
What AI enrichment actually changes
Against this backdrop, AI-based product data enrichment is worth taking seriously, because it changes something structural rather than incremental.
Historically, enriching a product required human attention. Research the product, find the specifications, write the description, assign the category, extract the attributes, check the images. That's minutes per SKU at best, which means catalog throughput is limited by headcount.
AI enrichment breaks that bound. Missing data can be found and extracted from supplier sites, manufacturer pages, PDFs and images. Content can be generated in a consistent voice, at length, in multiple languages. Attributes can be extracted from unstructured text and mapped to a controlled vocabulary. Categorisation can be applied consistently across a catalog rather than inconsistently across a team.
Two things improve simultaneously that used to be in tension: cost and quality. Enrichment gets cheaper per SKU, and the output is frequently more complete than what a rushed human would have produced, because the machine doesn't get bored on product 4,000 and doesn't skip the attribute fields at the bottom of the form.
The shift is measurable, not theoretical. One retailer we work with cut a six-month manual localization process to two weeks. Another onboarded 3,000 products across 60+ marketplace categories — each with its own attribute template — and doubled marketplace orders after listing.
That's the genuine advance. But it moves the hard question rather than answering it.
The hard question: who runs it, and who pays?
Enrichment costs something. Not much per SKU, but a marketplace catalog is a large number multiplied by that small cost, and someone has to bear it. More importantly, someone has to operate it. There are three workable answers.
Model 1 — The marketplace merchandising team runs it
Enrichment sits inside marketplace operations, integrated into the catalog systems, running on a schedule.
This is the right model where control matters most: strategic categories where data quality is a competitive differentiator, bulk onboarding when a large seller signs and delivers 50,000 SKUs at once, international expansion where an existing catalog needs translating and re-localising, and any own-inventory the marketplace holds directly.
It's also the only model that can handle deduplication, variant merging and canonical product construction — because those require cross-seller visibility that no seller has.
The limitation is economic. The marketplace pays for everything, and cost still scales with catalog size. It's better than manual enrichment by a wide margin, but it doesn't change who's funding the work.
Model 2 — Sellers run it themselves
Enrichment is offered to merchants as a self-service capability. They bring whatever data they have, produce listings that meet marketplace requirements, and cover their own usage.
The economics here are much more attractive to a marketplace, because cost scales with the party that benefits from the listing. It also solves the acquisition problem directly: a seller with inventory and no product data is no longer disqualified, because the gap between what they have and what you require is now something they can close in an afternoon rather than a barrier they can't cross.
The risk is obvious. Quality becomes variable again, distributed across thousands of independent actors with differing levels of care and capability. Whether this model works depends entirely on how it's implemented — more on that below.
Model 3 — Both, deliberately divided
In practice this is where most marketplaces should end up, and the division isn't arbitrary. It follows from what each party can actually see.
Sellers handle intake. Turning their own messy source data into structured, complete, marketplace-conformant records. They know their products. They have the source material. They benefit directly from the listing. This is the long tail, and it's the part that scales with seller count.
The marketplace handles what requires cross-catalog visibility. Deduplication. Variant consolidation. Canonical product records. Taxonomy governance. Category-level quality standards. Bulk operations. Anything where the correct answer depends on seeing more than one seller's data at once.
The clean version of this: sellers produce offers, the marketplace produces the catalog. Enrichment serves both, from different ends, with different economics.
What has to be true for seller-run enrichment to work
Model 2 is the one that determines whether Model 3 is viable, and it fails easily. Six things separate implementations that work from implementations that quietly get abandoned.
Self-registration. If a seller needs a sales call, a provisioning step, or an approval from marketplace operations before they can start, the long tail is excluded by definition. You cannot onboard three thousand merchants through a funnel with humans in it. The seller has to be able to go from interest to first enriched product without anyone at the marketplace being involved.
No assumptions about incoming data. This is the one most implementations get wrong. The moment the tool says "upload a CSV with these columns," you have rebuilt the template problem with extra steps and added a tool to it. The seller's data arrives as it arrives — a spreadsheet from 2019, an XML feed built for a different channel, a supplier PDF, a list of part numbers, photographs. Any format constraint on the input reproduces exactly the friction you were trying to remove.
AI assistance for non-technical sellers. Your typical marketplace merchant is a shop owner, not a data engineer. They cannot build a field mapping, define a transformation, or write a rule. What they can do is describe what they sell and what they need. The system has to close that gap — turning a conversation into a working enrichment process, and asking clarifying questions when the requirements are ambiguous.
Seller experience is the whole game. Sellers compare effort across marketplaces, continuously, without telling you. If enrichment takes longer than filling in your template, they'll fill in the template badly or list somewhere else. The bar isn't "better than nothing" — it's "faster and less painful than the alternative they already have."
Validation rules stay independent of the enrichment process. This matters more than it sounds. The marketplace owns the rules; the seller owns the process for meeting them. If the two are entangled — if the tool defines what "valid" means — sellers optimize for passing rather than for accuracy, and you're back to compliant-but-wrong listings.
Human review before publication. On both sides. The seller should see enriched output and be able to correct it before it goes live, because they know things about their products that no amount of research will surface. The marketplace should sample and spot-check, particularly in new categories and for new sellers. Automation handles the volume. Judgement handles the exceptions. Any implementation that skips the review step is trading a quality problem you can see for one you can't.
These six aren't theoretical. They're the pattern I've seen in separate seller-facing enrichment programs that get adopted from ones that get quietly abandoned after the pilot.
Where this leaves us
The marketplace catalog problem has always been a distribution question rather than a technology question: someone has to do the work, and every option for who that someone is has been unattractive.
What's changed is that the work itself got dramatically cheaper, and — more importantly — became something a non-specialist can direct. That makes it feasible, for the first time, to put real enrichment capability in the hands of sellers rather than concentrating it in an internal team that scales with your seller count.
That's a meaningful shift. But it only pays off if the implementation removes friction rather than relocating it, and if the marketplace retains ownership of the things only it can see: the taxonomy, the validation rules, the canonical catalog, and the final quality bar.
The technology is no longer the constraint. The operating model is.

