All AI recommendations
    Methodology

    How we measure what AI recommends

    Every page in this section reports what four major AI models said when asked a specific buyer question. This page documents how the measurement is performed, the sample size, and the limits of the method.

    The procedure

    For each category, we run one fixed question against four AI models: ChatGPT, Gemini, Claude, Perplexity. The question is identical across models. We attach one instruction, the same for every model, asking it to return a ranked shortlist of three to six brands as JSON, ordered by how strongly it would recommend each, along with any source URLs and a one-line rationale. That is a formatting instruction only. We do not add a persona, a preference, or any context that would steer which brands come back.

    We read the brands straight from each model's structured answer, in the order it returned them, together with the citations. The brand names are the model's own. We do not re-parse free text or match against a brand dictionary, so a brand is counted only when the model itself names it.

    Measurements are refreshed quarterly by re-running the set, which replaces the previous data. The "measured on" date on each page reflects the most recent run. This is a manual cadence, not an automated schedule.

    Honest limitations

    Sample size is small

    One question per category per quarter. AI responses vary substantially across runs: SparkToro's research found less than a 1-in-100 chance that two identical queries return the same brand list. A single run is a snapshot, not a definitive ranking.

    Citations are model-dependent

    Perplexity, Gemini and Claude return their citation sources reliably. ChatGPT returns them inconsistently. When citations are missing, the model did not return them on that run; it does not mean no sources exist.

    Brands are the model's own words

    We take each brand exactly as the model returned it. If a model spells a name differently or leaves one out, that is reflected as-is, and case or spelling variants can split counts when we tally agreement across models.

    Models are not pinned to one version

    Different measurement waves used different versions within each model family. Read a run as a snapshot of what that family said at the time, not a fixed, reproducible model ID.

    This is an observatory, not a buyer's guide

    These pages show what AI says about each category at one point in time. They are not buying recommendations. Use them to understand AI recommendation patterns, not to choose vendors.

    Scope

    We currently track 41 category questions across SaaS, SEO, content, marketing, PR, social, e-commerce, creator, and agency tooling. Questions are picked because they are common buyer queries in Honeyb's audience.

    The SEO categories also power our SEO tool directory, where each tool carries the share of buyer questions in which AI names it.

    We expand the question set when a category reaches a clear gap. Suggest a category by emailing us, or run your own brand against the same models using our free AI visibility checker.

    Why publish this at all

    Two reasons.

    First, AI recommendation has become a real discovery channel. 58 percent of consumers have replaced traditional search with AI for product research according to Capgemini's 2025 data. Buyers are reading these answers and making decisions from them. Showing what AI actually says, transparently, helps everyone understand the channel better.

    Second, Honeyb runs this kind of measurement professionally for customers, at much higher sample sizes, across full prompt sets, with proper sentiment validation. The pages here are an honest, public-facing version of the same discipline.

    See your own brand the same way

    Run the same kind of measurement on your brand across every major AI model. Free to start.