AI product manager: deciding what accuracy is good enough

Managing an AI product is ordinary product management until the first time the system is confidently wrong, and then it is a different job. A conventional product either does what it says or has a defect somebody can fix. This one produces a plausible, well-formed, incorrect answer as part of normal operation, and no amount of engineering removes that property. The decision that follows belongs to product rather than to engineering, and it is the one most teams never make explicitly: which errors are acceptable, on which inputs, for which users, and what the system does when it is unsure. A team without a written answer to that has delegated it by default to whoever tunes the prompt, and will discover its position on acceptable error from a customer complaint rather than from a decision.

AI product manager, in short

Emerging role
AI product manager : what the role is and how it is judged
In one sentence Decides what an AI product should do, with the unusual complication that its behaviour is probabilistic.
Judged on Whether the product is used and whether its failures are the ones the team chose to accept.
Fails when It is run like ordinary product management, with a roadmap of features and no position on acceptable error.
Most confused with Ordinary product management, which it resembles until the first time the system is confidently wrong.

Real work, but the scope differs enough between employers that the title alone tells you little. No compensation figures: see methodology.

The acceptable error position, and why nobody asks for it

No stakeholder requests this document. Customers assume the system is right, executives assume quality is an engineering concern, and engineers assume someone has decided what good means. So it does not get written, and the absence is invisible until it is expensive.

What it contains is short. The categories of input the system handles, the kind of mistake that is tolerable in each, the kind that is not, and what happens on the intolerable ones. Half a page for most products.

Writing it forces three useful arguments. Whether being wrong quietly is worse than declining to answer, which for most business uses it is. Whether all users are alike, which they are not, since an expert can catch an error a novice acts on. And whether the system should ever act rather than suggest, which is a different product entirely and is usually decided by drift rather than by choice.

Why the roadmap looks wrong to everyone else

On a conventional product, progress is features. Here, a large share of the most valuable work is making something the product already nominally does actually work often enough to rely on, and that does not appear on a roadmap in any satisfying form.

The pressure is therefore constant to add rather than to raise reliability, because additions are visible and legible to people outside the team. Products that give in to it end up wide and untrusted: nominally capable of many things, reliable at none, and quietly abandoned by users who tried three features and were let down by two.

The defence is measurement. A reliability figure per category, reported alongside the feature list, makes invisible work visible and turns an argument about priorities into an argument about numbers. Without it, the product manager is defending an intuition against a roadmap, which is not a fight anyone wins repeatedly.

Owning the evaluation set is the job, not a delegation

Engineers build the harness. What belongs in it is a product decision, because it encodes which cases matter, and a set assembled purely by engineering measures what is convenient to measure.

The symptom is a system whose numbers improve while users complain. That happens when the evaluation set over-represents easy, well-formed inputs and under-represents the messy minority that generates most of the frustration. Nobody did anything wrong; the set was built from what was available rather than from what matters.

The correction is a standing obligation: sample real production inputs, on a schedule, and feed them back. It is unglamorous, it takes an hour a week, and it is the single practice that most reliably separates products that improve from products whose metrics improve.

Pricing a product whose cost varies per request

A conventional software product has a marginal cost close to zero. This one does not: each request costs something, the amount depends on how much work the system did, and a heavy user can cost a multiple of a light one on the same plan.

That makes pricing a product decision with an engineering dependency, which is unusual. Flat pricing on a product with variable cost invites the customers who cost most to use it most, and teams discover this when a single account turns a healthy margin negative.

What the role has to say no to

Demos that set an expectation the system cannot hold. A curated demonstration creates a belief about reliability that the product then has to live with, and the person who pays is the user who tries the same thing with their own data.

Agentic behaviour introduced because it is impressive. Moving from suggesting to acting changes the approval model, the logging obligation and the accountability question, and it is worth doing only when the suggestion genuinely is not enough. This is covered on our agent and chatbot comparison, and it is a product decision rather than an architectural one.

And confident output on inputs the system handles badly. Surfacing uncertainty costs perceived capability and buys trust, and trust is what determines whether anybody is still using the thing in six months.

Questions people actually ask

How is this different from ordinary product management?

Everything is the same until the system is wrong. A conventional product either works or has a bug; this one produces a plausible, well-formed, incorrect answer as normal operation. Deciding which errors are acceptable, for whom, and what happens when they occur is a product decision that has no equivalent elsewhere.

Does the role need technical depth?

Enough to know what cannot be fixed by asking. A product manager who believes any quality problem can be solved with a better prompt will set commitments the team cannot meet. The useful floor is understanding why a system fails on a class of inputs, which is diagnostic rather than mathematical.

What does a roadmap look like here?

Less feature-shaped than elsewhere. A meaningful share of the work is raising reliability on things the product already nominally does, which is invisible on a roadmap and is what users actually notice. Teams that only ship features end up with a wide product nobody trusts.

Who owns the evaluation set?

In practice the product manager should, even when engineers build it, because what belongs in it is a product judgement about which cases matter. Handing that decision entirely to engineering produces a set that measures what is easy to measure rather than what determines whether the thing is used.

Read next

Sources

Radif Partners

Written and maintained by Radif Partners

Applied AI deployment practice · Forward deployed engineering

Covers 2026, · last reviewed 2026-09-24