← All posts

AI names the brand it expects, not the one it saw

· 7 min read

A magnifying glass leaning against a phone

Ask an AI what brand that is and you will get a brand name. You will get one whether or not a logo was ever in frame. The name is the model's best completion of a sentence, not something it read off a label, and from the outside those two are impossible to tell apart because they arrive in the same even, unhesitating voice.

So before you spend anything, ask one question: where does that name come from? If nobody said it out loud, it was not written on screen, and the brand's own site does not carry the item, you are looking at a guess with good posture.

The picture is the weaker half of the model

A multimodal model has two parts doing very different jobs. A vision encoder turns the frame into something the language model can work with. The language model then writes the answer. The language half is enormously larger and far better trained than the vision half, and when what it sees and what it expects disagree, the expectation usually wins. Researchers call it modality imbalance. You can call it the model finishing your sentence.

Picture what it is actually given. A creator films a beige ribbed tank in a bathroom mirror. Two seconds, mixed light, half the garment out of frame, no logo, no on-screen text, nothing named in the audio. There is no brand information in that clip. The model answers anyway, because “beige ribbed tank” has a most likely neighbour in its training data, and that neighbour is a brand name.

It is worth being precise about what went wrong there. The model did not misread a logo. It never had one to read, and it produced a brand name regardless, because producing text is the only thing it does.

It reaches for the brand that usually stands there

Work on vision-language models has found the failure is patterned rather than random. These models hallucinate most for objects that are frequent in their training data, and for objects that commonly co-occur with what is genuinely in the frame. The benchmarks that test this deliberately ask about the plausible absent thing, because that is the one models fall for. Show a plate, get a fork.

Shopping is that problem with money attached. The brands most likely to be named are the ones that were already famous, because fame in the world is frequency in the training data. Meanwhile:

In none of those cases does the model return nothing. It returns the nearest famous neighbour. That is why the answer is so often a household name and so rarely the small brand that actually made the thing.

Fluent and unsure look identical

Humans hedge without being asked. We slow down, we say probably, our face does half the work. A model has no separate channel for doubt. A guess and a certainty come out of the same process, in the same sentence shape, at the same speed.

Hedging has to be deliberately built in, and it tests badly, because the honest answer reads as a failure. “I can see a ribbed knit tank in oat and no brand marking anywhere” is the most useful sentence such a tool could give you, and it is the sentence most likely to be engineered away for sounding unhelpful.

This applies to the model inside our app exactly as much as any other, and pretending otherwise would be the same failure in a different costume. It applies to visual search too, in its own way: Google Lens returns matches, not authentications, so two products sharing a mould or a silhouette look like one product to it. That is genuinely useful for finding an item and its lookalikes, and it is the wrong instrument for deciding which of them you are looking at. Lens has its own set of failure modes on video, and they are worth knowing separately.

What a wrong brand actually costs

The damage is not that you learned something false. It is that a specific name sends you somewhere specific.

What counts as evidence for a brand

A brand claim is only as good as what sits behind it. In rough order of strength:

What does not count, however confident it sounds:

A thirty-second test for any AI answer

Why we would rather come back with less

We build against this on purpose, and the honest way to describe it is as a trade rather than a solution. Stealsy holds every brand claim against the post's own words, what was said out loud, what was on screen, what the caption wrote, or against a verified link on that brand's own site. No evidence, no brand: the product ships honestly generic instead of confidently mislabelled. The rest of that pipeline is written up here.

The cost is real and we should say it. Sometimes you get “ribbed knit tank in oat” when there was a genuine brand behind it that we could not prove, and a generic answer is less satisfying than a name. We take that trade every time, because a description that is merely incomplete leaves you searching from something true. A brand name that is wrong sends you to a checkout.

Get Stealsy for iPhone — share a video, get the products, prices and verified links.

Common questions

Why does AI get the brand wrong when identifying a product from a video?
Because naming a brand and reading a logo are two different things, and only one of them is happening. In a multimodal model, a vision encoder converts the frame into something the language model can work with, and the language model then writes the answer. The language half is far larger and far better trained than the vision half, so when what the model sees and what the model expects disagree, the expectation usually wins. Researchers call this modality imbalance: ingrained language priors override the actual visual evidence. On top of that, vision-language models hallucinate most often for things that are frequent in their training data or that commonly co-occur with what is genuinely in the frame. Applied to shopping, that means the wrong brand is not random. It skews toward whichever brand is most famous in that product category.
How do I check whether an AI's brand answer is right?
Ask where the name came from, then verify it in one click. A brand claim is only as good as its source: someone saying the name out loud in the audio, the name written in on-screen text or the caption, a logo legible enough to read rather than infer, the creator's own storefront listing the item, or the brand's own website carrying that exact product. Open the brand's site and search for the item. If the brand does not make it, you have your answer immediately. Also check that the shade or variant name exists in the brand's real lineup, because invented variant names are one of the loudest signs an answer was generated rather than read.
Can Google Lens or an AI tool tell a real brand from a dupe?
Not reliably, and it is not built to. Visual search returns visual matches, not authentication. It compares shape, colour, texture, pattern and packaging, so two products that share a mould or a silhouette look like the same thing to it even when one carries a brand and the other does not. That is genuinely useful for finding the item or its lookalikes, and it is the wrong tool for deciding which one you are holding. Authentication needs evidence a picture does not contain: the seller, the listing history, the brand's own catalogue, the serial or batch markings.
Is a generic product description better than a confident brand guess?
For anything you are about to pay for, yes. A description like ribbed knit tank in oat tells you exactly what is known and lets you keep searching from a true starting point. A wrong brand name sends you somewhere specific and wrong: onto a brand's site looking for a product they have never made, or to a checkout for a more expensive item that was never the one in the video. The wrong name also spreads, because you will repeat it in a comment or to a friend, and a guess with two sources starts looking like a consensus.