upsunday

Your Brand Is a Dataset Now: How AI Models Form an Opinion of Your Company

Language models describe your company from everything written about it. What model reputation is, why consistency beats volume, how to measure AI visibility when the answers keep changing, and the Four Cs we use to shape what AI says about you.

Jake Young8 min read
Cover art for “Your Brand Is a Dataset Now: How AI Models Form an Opinion of Your Company”

Right now someone is asking an AI assistant about your company. Is it any good? Who is it for? How does it compare with the one they already know? The assistant answers in a confident paragraph, and nobody from your company was in the room. That paragraph is your brand as the model understands it, assembled from everything the web has said about you.

We call that paragraph model reputation. It forms like an average rather than an impression, it turns inconsistency into vagueness, and your naming, copy and design decisions are now part of its training material. It can also be shaped, deliberately, with four layers of work.

Model Reputation Is a Consensus

One great experience can change a person’s mind about a brand. A model’s impression works more like an average, formed from training data and, increasingly, from live retrieval of pages, reviews, forum threads and videos. The inputs that matter most appear to sit off your site. Ahrefs studied 75,000 brands and found branded web mentions correlated with visibility in Google’s AI Overviews at 0.664, against 0.218 for backlinks, and a follow up in May 2026 found YouTube mentions were the strongest single signal it measured.

The sources models favour also shift without warning. Semrush tracked more than 100 million citations across ChatGPT, Google AI Mode and Perplexity over thirteen weeks in 2025 and watched Reddit and Wikipedia citations in ChatGPT fall, in mid-September, from roughly 60% and 55% of responses to somewhere between 10% and 20%. No brand planned for that. Brands described the same way across many kinds of sources barely felt it.

Why Inconsistency Becomes Vagueness

Consider what a model does with contradictory inputs. If your site calls you a platform, your founder’s interviews say agency, LinkedIn says consultancy and reviews describe a software tool, the model doesn’t pick the right one. It blurs them. You get described in the most generic terms that fit everything, or you get left out of answers that need a confident match.

Blur becomes expensive when the model is also wrong. The EBU and BBC reviewed more than 3,000 answers from ChatGPT, Copilot, Gemini and Perplexity across 18 countries and found almost half had at least one significant issue, and a fifth had major accuracy problems such as outdated or invented details. That study covered news, not brands, but the mechanism transfers. Where sources are thin, stale or contradictory, the answer fills the gap.

Big brands aren’t safe either. A June 2026 preprint by Chu and Hou tested product recommendations in GPT-4o-mini, Claude Sonnet and Gemini, and found established brands won nearly every recommendation when products looked identical, then lost that edge to a rival with a rating advantage of less than a tenth of a star. Familiarity gets you considered. Evidence decides the answer.

The Four Cs of Model Reputation

When we audit how models see a brand, we work through four layers in order. Each one makes the next easier.

Claim: what you say about yourself

Write one sentence that states what you are, who you serve and why you’re different, plainly enough that a model can repeat it. Then use nearly the same words on your homepage, About page, structured data description, LinkedIn, press boilerplate and founder bios. Repetition bores a marketing team. To a model it’s the difference between a fact and a guess. An atmospheric headline can stay; the plain sentence needs to sit close to it, in text.

Corroboration: what others say about you

A claim only you make is an advertisement. The same claim from a trade publication, a customer review, a podcast host and a forum thread starts to look like a fact. The original GEO paper by Aggarwal and colleagues, published at KDD 2024, found that adding citations, quotations and statistics could lift a source’s visibility in generative answers by up to 40%. Models favour content that shows its evidence. So earn coverage that restates your claim in the writer’s own words, make reviews easy to leave where models read them, and publish in the formats they cite, video included.

Category: what you’re filed under

Models organise the world into entities. Yours needs to be unambiguous, and it needs to sit in the right category next to the right neighbours. This is where most of the technical work lives, so we cover it in detail below.

Comparison: how you stack up when asked

Many of the questions that matter are comparisons. In Seer Interactive’s 2026 dataset, 95.4% of comparison queries triggered an AI Overview, the highest rate of any query type. When a model compares you with rivals, it reaches for attributes it can state with confidence: specialisms, ratings, policies, public customers. Publish your differentiators as checkable facts, or the model compares you on whatever it can find.

Engineering Entity Clarity

Fix the name first. One legal name, one trading name, one spelling, one logo, one primary domain. If you share a name with another company, add a consistent disambiguator everywhere, such as the category or city. If a product has been renamed, redirect the old URLs and state the history once, plainly, on an official page, so models can connect the two names instead of treating them as rivals.

Then describe the entity in structured data. Put Organization JSON-LD on the homepage with a stable @id, your canonical name, logo, description, founding date and sameAs links to the profiles that are unmistakably yours: LinkedIn, Crunchbase, official social accounts and a Wikidata item if one exists. Google’s documentation describes this markup as helping it disambiguate your organisation. Give products and sub-brands their own markup that points back to the parent with brand or parentOrganization, so the family tree is explicit.

Then align the places models read. Update directory and marketplace profiles, review platforms, app store listings and old press releases that still carry outdated descriptions. On Wikipedia, don’t write or edit your own article. Its conflict of interest rules are strict, and a deleted page does more harm than no page. If you’re notable, independent coverage will get you there.

The trade-off is creative freedom. Campaign copy can play around the canonical description but shouldn’t contradict it, and product teams lose the ability to rename things casually. Brands that rename every season pay for it in how clearly models can describe them.

Citation-Worthy Content Architecture

If corroboration drives reputation, your site’s job is to be the most quotable primary source about you. We structure it as fact pages: one canonical URL per important claim, with a plain answer first, the evidence underneath, an author and a date. Add a methodology page for any number you publish and a policies page that states terms once. Original data, even modest, gives other writers a reason to cite you, and every citation carries your claim further.

Measuring AI Visibility, and Why It’s Hard

Most AI visibility dashboards report a rank. That number is close to meaningless. SparkToro and Gumshoe had 600 volunteers run the same 12 recommendation prompts through ChatGPT, Claude and Google’s AI nearly 3,000 times, and the chance of getting the same list of brands twice was under one in a hundred. Answers also vary with memory, location, language and model version, and much AI usage never produces a click you can see.

What holds steady across runs is how often you appear and how you’re described, so that’s what we measure. Build a fixed panel of prompts that real customers ask: about you by name, about your category and head-to-head against named rivals. Run each prompt many times on each major assistant, logged out and in a clean session, and record four things: appearance rate, accuracy of the description, the words used about you, and which sources are cited. Report the rate with its range, and compare month to month on the same panel.

Then triangulate with signals you already own. ChatGPT tags the links it sends with utm_source=chatgpt.com, and other assistants show up in analytics by referrer. Server logs show which AI crawlers fetch which pages, and Google has begun adding generative AI reports to Search Console. Branded search volume and direct traffic remain the slowest and most honest indicators of whether people are asking for you.

Design and Copy Decisions Now Train the Answer

Every naming choice, tagline, page title and piece of alt text is now input to a system that summarises you for strangers. A product renamed for a campaign creates two entities. A case study with its results baked into an image says nothing to a crawler. Multimodal models read screenshots, product photos and video, so a brand that looks the same everywhere is easier to recognise and connect across sources. Consistency was always good brand practice. Now it compounds.

A 3D rubber stamp pressing down, its face carrying a sun emblem
Every page you publish leaves an impression that models learn from.

There’s a trap here too. Chu and Hou also tested what happens when every brand in a category optimises its claims for the model at once: the recommendation rate for any single brand fell from 0.802 to 0.007 on their measure. Optimisation anyone can copy cancels out. What lasts is substance other people will repeat.

The Next Two to Five Years

These are predictions. Share of answer becomes a board metric beside share of search, reported as ranges because positions don’t hold still. PR becomes infrastructure, since earned media, reviews and community presence are inputs to your most influential recommender. Brand guidelines grow a machine chapter covering the canonical description, entity names and published facts. And errors get slower to fix: once a wrong fact is widely repeated, correcting it takes many consistent sources and model updates you don’t control.

What to Do This Quarter

Build the prompt panel. Twenty prompts across your name, your category and head-to-head comparisons, run at least ten times each on ChatGPT, Claude, Gemini and Perplexity. Record appearance rate, description and cited sources. That’s your baseline.

Write the canonical sentence, get it signed off and put it everywhere you control. Then fix the entity: one name and spelling, Organization JSON-LD with a stable @id and honest sameAs links, and explicit parent relationships for sub-brands.

Publish your differentiators as fact pages with evidence, authors and dates, and trace every wrong answer to its likely source, such as an old directory listing or a stale press release, and fix it there.

Your brand has always lived in other people’s heads. Now it also lives in a model’s, and the model reads everything. If you want to see how AI describes your company today, and build a brand consistent enough to be described well, we’d love to talk. We usually start with the prompt panel.

  • AI Search
  • Brand
  • GEO
Share

More insights

Keep reading

Contact

Let's talkabout something yours