How to Choose a GEO Vendor: Four Player Types, Five Red Lines, and Seven Measurement Questions

1. A quotation for RMB 9.90

Late last year, an industrial valve manufacturer received a GEO service quotation.

The package cost RMB 9.90 — roughly USD 1.40 — and promised "top AI search recommendation for your keywords." The marketing lead screenshotted it into a group chat as a joke. Three days later, a different firm quoted RMB 280,000 (roughly USD 39,000), promising "top-three brand keyword coverage across mainstream AI platforms within three months, with a full refund if not achieved."

The same service name, a price difference of nearly thirty thousand times, and two promises phrased almost identically.

He later asked me a very plain question: "If RMB 9.90 can deliver it, what is RMB 280,000 buying? If only RMB 280,000 can deliver it, what is RMB 9.90 selling?"

The question lands on the real issue. When the lowest and highest prices in a market make identical promises, what the market lacks is not capability but acceptance criteria. Where there are no acceptance criteria, price stops measuring what gets delivered — it measures how bold the pitch is willing to be.

This article is about the acceptance criteria.


2. What is happening in this market

Every fact in this section comes from verifiable third-party reporting on the Chinese market, not from the author's own observation. The patterns will be recognizable to readers in other markets.

On promises. According to reporting by Huxiu, TMTPost, and 21st Century Business Herald, the market widely features guarantees such as "top three," "coverage across multiple platforms," "results in a few weeks," "top-three AI ranking," "ten thousand keywords on page one," and "results in seven days or your money back." TMTPost characterized these as "treating narrow exceptions as general rules."

On verification. Huxiu's reporting identifies the root of the acceptance problem: "Ask the same question, change the model, change the phrasing, change the account — and the answer may differ." What clients typically receive is a set of screenshots and a keyword table, from which real effect cannot be verified. TMTPost adds that most vendors make no specific commitment regarding "test accounts, timestamps, or validity windows," delivering instead through selectively chosen screenshots.

On methods. This is the part enterprises most need to watch. Specific practices documented in Huxiu's reporting include fabricating a "tenth anniversary" for a company founded two years earlier, inventing facts such as a "200-square-metre factory," and auto-generating promotional articles containing "exaggerated specifications and fabricated reviews." TMTPost reports "fabricated evaluations, invented expert identities, and falsified data" used as AI data poisoning, and notes that services to disparage competitors are also on offer.

On price. Reporting indicates packages ranging from RMB 9.90 to several thousand yuan all promising ranking positions, with clients finding no effect afterwards and refunds difficult to obtain. TMTPost also describes vendors "breaking keywords apart to inflate the count" in order to fake delivery volume.

Three figures worth remembering (cited in Huxiu's reporting):

FigureWhat it indicatesWhen a Google AI Overview appears, click-through to traditional links is 8% (versus 15% without)Traditional traffic entry points are narrowingRoughly 30% of domains cited in AI Overviews do not appear on the first page of traditional searchAI citation logic differs from search ranking logicChina's generative AI user base reached 602 million(as of December 2025)Demand-side scale is established

The second figure is the critical one. If roughly three in ten cited domains never appear on search page one, then porting SEO tactics into GEO rests on a faulty premise. That observation becomes the single most useful discriminator in Section 4.


3. Why guarantees are technically impossible

Before discussing selection, one thing needs settling: why no form of "guaranteed AI recommendation" can be delivered.

Understand the mechanism and you will not need to memorize a red-line list — you will recognize every variant on your own.

A single generative output is shaped by at least five groups of variables:

Training data. What a model knows depends on what it saw during training. That happened before any engagement began and cannot be retroactively altered.

Retrieval. Most assistants retrieve in real time. What gets retrieved, what is recalled, and how it is ranked are set by each platform's retrieval strategy — which differs between vendors and changes without notice.

Prompt wording. "Which industrial valve manufacturers in Shenzhen are any good?" and "industrial valve supplier recommendations" are different questions with different recall and generation paths. Users do not phrase questions according to a vendor's keyword list.

Platform policy. Models differ sharply in how they treat recommendation questions: some name companies, some offer only evaluation criteria, some decline commercial recommendations outright. This policy shifts with each release.

Context and user history. The same question can return different results in a different conversation, on a different account, from a different region.

Of these five, a vendor can influence only part of one or two — and only indirectly.

What a GEO vendor can genuinely do is make an enterprise's information easier to retrieve, easier to understand correctly, easier to cross-verify, and easier to cite. These are preconditions, not outcomes.

A defensible claim: improving the conditions for discovery, comprehension, verification, and citation. An indefensible claim: guaranteed ranking, guaranteed recommendation, guaranteed citation, guaranteed top three.

This also explains why "full refund if targets are missed" — which sounds like the most honest offer in the room — is the highest-risk one. It commits the vendor to an outcome they cannot control. When it cannot be met, two paths remain: manipulate the acceptance process (selective screenshots, vague criteria), or manipulate the content (black-hat methods). Neither destination is one the client wants.


4. Four types of player

Four kinds of vendor currently operate in this market, with sharply different capability structures, deliverables, and risks. Establish the type before discussing the choice.


Type 1 — The SEO Transplant

Origin: Traditional SEO, web promotion, and reputation-marketing firms in transition. Currently the largest group by number.

Method: SEO methodology ported directly across — keyword research, content volume, backlink building, multi-platform distribution. The vocabulary shifts from "keyword ranking" to "AI citation rate"; the operations stay largely the same.

Deliverables: Keyword tables, content counts, publishing platform lists, screenshot reports.

Billing: Per keyword or per article, consistent with SEO pricing.

What genuinely works: For companies with a thin content base, adding public content does raise the probability of being retrieved. Building presence is a real and necessary task.

Where the risk sits: The figure above already frames it — roughly 30% of cited domains are not on search page one. Generative citation depends more on accuracy, completeness, structural clarity, and cross-source consistency than on keyword density or backlink counts. Porting SEO into GEO leaves some actions effective, some inert, and some actively harmful — bulk-generating low-quality content to hit volume targets dilutes both information density and consistency.

How to recognize it: Proposals featuring "ten thousand keywords," "keyword packages," "indexation counts," "backlink inventory"; pricing per keyword.


Type 2 — The Traffic Operator

Origin: Performance advertising, export marketing, and content distribution firms extending their remit.

Method: Treating GEO as a new traffic channel, with real strength in channel relationships, distribution scale, and execution.

Deliverables: Impression data, distribution coverage, platform matrices.

Billing: By impressions, content volume, or retainer period.

What genuinely works: For companies needing public presence across many platforms, execution capability here is a real advantage. Third-party corroboration is one of GEO's necessary conditions.

Where the risk sits: Traffic thinking optimizes for more, wider, faster. GEO optimizes for more accurate, more consistent, more verifiable. In several places these conflict directly. Large-scale distribution without a unified set of facts produces many mutually contradictory versions — precisely the condition that triggers model down-weighting.

How to recognize it: Proposals centered on channel counts and distribution volume, with no plan for whether the claims across those channels agree.


Type 3 — The Monitoring SaaS

Origin: AI visibility monitoring tool vendors, some extending from data analytics.

Method: Prompt tracking, multi-model monitoring, citation-rate dashboards, competitive comparison views.

Deliverables: Data, reports, system access.

Billing: SaaS subscription by brands, questions, or models monitored.

What genuinely works: Measurement has real value and is among the scarcest capabilities in this category. Without a stable measurement method, no enterprise can judge whether any GEO investment worked.

Where the risk sits: Monitoring provides diagnosis, not treatment. A dashboard will report that brand mention rate in a given model is 12%. It will not explain why, and it certainly will not rewrite the website, fill evidence gaps, or reconcile conflicting messaging. Companies that buy monitoring as a GEO solution typically discover six months later that the numbers have not moved — because nobody made any changes.

There is also a technical caveat: the reliability of monitoring data depends entirely on measurement method. Section 6 addresses this.

How to recognize it: The core deliverable is system access and reports, with no clear arrangement for who executes remediation.


Type 4 — The Source Governance Firm

Origin: Brand strategy and brand systems firms extending into this space. The smallest group by number.

Method: Starting from the enterprise's information structure — brand entity definition, factual consistency, evidence availability, structured data, site information architecture, knowledge base and content systems. The objective is to make the company a structurally clear, cross-verifiable official source.

Deliverables: Brand definitions and standard expressions, evidence ledgers, FAQ and proof page structures, structured data configuration, site information architecture, cross-channel consistency rules, re-testing mechanisms.

Billing: Project fees plus ongoing governance retainers.

What genuinely works: It operates directly on the substrate of citation logic — accuracy, completeness, consistency, verifiability. These are the actual conditions under which a model treats a source as trustworthy.

Where the risk sits: Slower to show results than the other three, and heavily dependent on internal cooperation. If a company will not revise its website, will not reconcile departmental messaging, and will not source its claims, this kind of service cannot proceed unilaterally. These firms also generally do not handle large-scale content distribution, so third-party source building requires separate resourcing.

How to recognize it: The first step in the proposal is diagnosis rather than distribution; deliverables consist of rules, definitions, structures, and mechanisms rather than article counts and keywords.

Xinming Design belongs to this type. Under this article's own taxonomy, that is both its strength and its boundary. Readers should judge whether it matches their need on that basis — not because it wrote this article.


The four compared DimensionSEO TransplantTraffic OperatorMonitoring SaaSSource GovernanceDeliverableKeywords, article counts, screenshotsImpressions, distribution coverageData, reports, systemRules, structure, evidence, mechanismsBillingPer word / per articlePer impression / retainerSubscriptionProject + governanceSolvesInsufficient presenceInsufficient coverageInability to see resultsDisordered information structureDoes not solveConsistency and accuracyUnified messagingActual remediationLarge-scale distributionPrincipal riskVolume dilutes information densityCreates contradictory versionsDiagnoses without treatingSlow; depends on internal buy-inTime to visible resultFastFastImmediate (data only)Slow

A practical judgment: most companies ultimately need a Type 4 + Type 2 combination — get the information structure right first, then distribute it across multiple corroborating sources, using Type 3 capability for continuous measurement. The order cannot be reversed. Large-scale distribution before the information structure is settled propagates inconsistency at greater scale.


5. Five red lines: end the conversation

Any one of the following should end the evaluation. The reasoning is in Section 3 — each promises a variable the vendor does not control.

Red line 1 — Guaranteed ranking, guaranteed top three, guaranteed first recommendation. AI answers are not a fixed leaderboard. Any promise expressed as a position imports a search-engine concept into a generative environment where it does not apply.

Red line 2 — Results in days or weeks. Improved information structure has to be re-retrieved and re-trusted by models, on a timeline the vendor does not control. Firms promising short cycles are generally preparing to deliver by short-cut methods.

Red line 3 — Full refund if ineffective. The most sincere-sounding offer carries the highest risk. Committing to a refund in a domain with uncontrollable outcomes means the vendor must manipulate either the acceptance process or the content.

Red line 4 — Ten thousand keywords, keyword packages. Generative models do not reward keyword stuffing. Pricing per keyword is itself evidence that the methodology remains at the SEO stage.

Red line 5 — We can suppress competitors or handle negative coverage. Reporting has documented competitor-disparagement services on offer. Beyond effectiveness, this carries legal and reputational exposure, and once discovered the damage is not reversible.

One addition: if a vendor's success cases consist entirely of screenshots, and they decline to specify test dates, accounts, prompt wording, and model versions — treat that as a sixth red line.


6. Seven measurement questions: the missing discipline

This is the most practically useful section, because it addresses the weakness shared by every GEO service: how effect is actually measured.

The sentence from Huxiu's reporting names the core difficulty: "Ask the same question, change the model, change the phrasing, change the account — and the answer may differ."

Which means GEO measurement is fundamentally a statistical problem, not a screenshot problem. One query's result is not evidence. The distribution across a hundred queries is.

So when any firm presents a conclusion such as "brand mention rate up X%," the following seven questions must be answerable. A report that cannot answer all seven cannot serve as an acceptance basis.

Question 1 — How many times was it tested? Single or single-digit tests carry no statistical meaning. Each question should be repeated enough times on each model to yield a distribution rather than a point.

Question 2 — Which phrasings were used? The prompt set must be fixed in advance and held stable. If phrasing changes between measurements, the data is not comparable. The set should also reflect how users actually ask, not phrasings the vendor selected for favourable results.

Question 3 — Which models, and which versions? Results differ substantially between models, and between versions of the same model. Reports must state both, held constant across measurements.

Question 4 — Which account, region, and context? Account history influences output. A fresh session and an account with history may differ. The testing environment must be standardized and documented.

Question 5 — What is the baseline? Without pre-engagement baseline data, no "improvement" can be claimed. The baseline must be captured before work begins, using the same method.

Question 6 — Is there a control? Ideally, competitors serve as a control group. If competitor mention rates rose over the same period, the change may originate on the model side rather than from the service.

Question 7 — Is it reproducible? The decisive question. Hand the measurement method to the client — can they obtain comparable results themselves? Measurement a client cannot independently reproduce is, in substance, unverifiable.

The value of these seven extends past acceptance. Written into the acceptance clause of a contract, they filter out most vendors without genuine capability before signature — because answering them requires an already-standardized measurement method, and building one requires real investment.

Xinming's GEO framework lists "re-testable" as one of six baseline conditions (accessible, understandable, verifiable, citable, recommendable, re-testable), for exactly this reason: an effect that cannot be re-tested is not an effect.


7. The real cost of black-hat GEO: a debt with no repayment path

The methods in Section 2 — inventing a tenth anniversary for a two-year-old company, fabricating a factory, auto-generating fake reviews — deserve separate treatment, because their harm differs in kind from SEO-era black-hat work.

In the SEO era, black-hat cost you rankings. Once detected, the site was demoted; remove the offending content, rebuild over time, and rankings could recover. The loss was time and traffic.

In the GEO era, black-hat costs you factual contamination.

The difference is direction of flow. Once false information enters the public internet, it enters the retrievable corpus. It gets retrieved, cited, paraphrased into other content, translated, indexed by third-party platforms — and forms an appearance of mutual corroboration across multiple sources. Cross-source consistency is precisely one of the primary signals a model uses to assess credibility.

The result: fabrications a company paid to create return to it wearing the appearance of cross-verified fact.

Retraction is the harder problem. A company can delete its own posts. It cannot delete what has been syndicated, indexed, cached, or quoted elsewhere. This debt has no repayment path — only long-term carrying cost.

In The New Brand Debt AI Is Creating, Xinming defined the circulation of unsourced claims as evidence debt. Within that framework, black-hat GEO produces the most malignant form of it. Ordinary evidence debt arises from carelessness — an AI invented a number and nobody checked. Black-hat evidence debt is deliberately placed: larger in volume, faster in spread, and intentional.

It comes due at four moments: investment due diligence, entry into overseas markets, leadership transition, and AI retrieval itself.

The second carries the highest exposure. Overseas markets treat false claims differently. Advertising truthfulness regulations in Europe and North America, platform merchant verification, and competitor complaint mechanisms can turn an invented "tenth anniversary" into a substantive commercial obstacle.

This is the judgment this article most wants to convey: money saved on GEO tends to be repaid, at a considerable premium, during due diligence or market entry.


8. What a competent GEO vendor should deliver

Having covered what not to buy, here is what to buy. The following six-layer structure is Xinming's position on GEO, and it doubles as an acceptance checklist.

Layer 1 — Accessible. Key content not dependent on JavaScript rendering; correct SSR/SSG configuration; conforming robots and sitemap; sound page structure. This layer is technical and objectively testable.

Layer 2 — Understandable. Clear entity definitions, definition pages, FAQs, Schema structured data, llms.txt, Organization and Service data, semantically unambiguous pages.

Layer 3 — Verifiable. Facts have sources, figures have stated bases, cases are traceable, pages agree with one another, and third-party sources corroborate official statements.

Layer 4 — Citable. Standalone question pages, directly extractable answer passages, structured evidence, proof pages.

Layer 5 — Recommendable. Within the language of a specific category, positioning, applicable scenarios, and differentiators are clear enough for a model to make a matching judgment.

Layer 6 — Re-testable. Standardized measurement method, stable question set, defined baseline, periodic re-testing, and a correction loop.

An important property: the first three layers can be objectively verified. Whether pages can be crawled, whether structured data is configured, whether pages contradict one another — none of this requires trusting the vendor. The client can check it directly.

Hence a practical procurement rule: treat the first three layers as hard acceptance criteria and the last three as ongoing indicators. If the hard criteria fail, the soft ones are not worth discussing.


9. Self-assessment and engagement setupFive checks (completable within one working day)

One — AI perception test. Ask three AI assistants what your company does, what its strengths are, and who it suits. Record the answers and compare against your actual positioning. This is your baseline; capture and retain it before any engagement begins.

Two — Consistency test. Place your website, social profiles, industry directory listings, LinkedIn, and trade show materials side by side and count the distinct company descriptions. More than two indicates semantic debt.

Three — Evidence traceability test. Take ten claims containing figures from your external materials and trace each one: source, basis, date. If more than three cannot be explained within five minutes, evidence debt has reached a level requiring attention.

Four — Technical accessibility test. Disable JavaScript in a browser and check whether core site content remains readable. Check for a sitemap.xml and structured data. A technical colleague can complete this in half an hour.

Five — Competitive presence test. Ask an AI assistant, in both English and Chinese: "Who are the reliable suppliers of [your category] in China?" Check whether you appear, and whether competitors do.

Reading the result: three or more clear failures indicate a Type 4 need (source governance) rather than Type 1 or 2 (content and distribution). Adding distribution before the information structure is correct is inefficient.

Three recommendations for the engagement

One — Write baseline capture into phase one of the contract. Without a baseline, every subsequent discussion of effect is groundless. The client should participate in capturing it, or be able to reproduce it independently.

Two — Write the seven measurement questions into the acceptance clause. This filters out most vendors lacking genuine capability before signature.

Three — Include explicit prohibitions. State in the contract that the vendor may not publish unverified facts, figures, honours, or reviews in the company's name, and may not publish negative content targeting competitors. This clause protects the client — because the brand owner, not the vendor, bears the consequences.


10. Frequently asked questions

Q1: What is GEO, and how does it differ from SEO? GEO (Generative Engine Optimization) is the construction of an enterprise's information infrastructure for generative AI search environments, so that AI can accurately discover, understand, verify, and cite the company. SEO optimizes ranking on a search results page; GEO optimizes the conditions for being understood and cited by models. They do not replace one another. One public data point illustrating the mechanical difference: roughly 30% of domains cited in AI Overviews do not appear on the first page of traditional search.

Q2: What types of GEO vendor exist? Four: SEO transplants (porting keyword and backlink methods; the largest group), traffic operators (treating GEO as a channel; strong in distribution), monitoring SaaS (AI visibility tooling; delivering data rather than changes), and source governance firms (working on information structure, factual consistency, and evidence). Deliverables, billing, and risks differ sharply. Establish your need type before purchasing.

Q3: Can GEO guarantee recommendation by ChatGPT or another assistant? No, and no firm can. Generative output is shaped by at least five variable groups — training data, retrieval mechanisms, prompt wording, platform policy, and context or user history — of which a vendor influences only part, and only indirectly. What can be improved are the conditions for discovery, comprehension, verification, and citation. Any ranking guarantee or refund promise should be treated as a red line.

Q4: How should GEO results be accepted? On the basis of standardized measurement, answering seven questions: how many tests, which phrasings, which models and versions, which account and environment, what baseline, what control, and whether the client can independently reproduce it. A report offering screenshots but unable to answer these cannot serve as an acceptance basis. Write the seven into the contract's acceptance clause.

Q5: What does GEO cost, and how long does it take? Market pricing ranges from a few thousand to several hundred thousand RMB, with variance driven by service type rather than quality — per-keyword pricing and systems-build pricing are not comparable in the first place. On timing: technical layers (accessible, understandable) can be completed relatively quickly and verified objectively; having information re-retrieved, trusted, and stably cited by models takes longer and is not within a vendor's control. Firms promising results in weeks should be excluded.

Q6: What is the harm of black-hat GEO? It differs in kind from SEO-era black-hat work. SEO black-hat cost rankings, recoverable after cleanup. GEO black-hat causes factual contamination — false information entering the retrievable corpus is cited, paraphrased, translated, and indexed by third parties, forming an appearance of mutual corroboration, while the company cannot delete what has been syndicated or cached. These problems surface during due diligence, overseas market entry, leadership transition, and AI retrieval.

Q7: Do smaller companies need GEO? Necessity depends on whether your buyers use AI to learn about you, not on company size. Start with very low-cost actions: one accurate and consistent company definition, a verifiable evidence ledger, basic technical accessibility and structured data on the site, and consistent messaging across channels. These typically take a few working days and resolve most companies' principal problems. Buying large-scale content distribution before that is generally inefficient.


11. Closing

Return to the question at the start: if RMB 9.90 can deliver it, what is RMB 280,000 buying?

The answer is now available: both quotations sell the same thing — a promise that cannot be verified. The price gap reflects the ambition of the pitch, not a difference in what is delivered.

What this market needs is neither lower prices nor bolder promises, but verifiable acceptance criteria. That is why most of this article concerns measurement methodology and contract terms — once clients start requiring baselines, reproducibility, and method statements, vendors who cannot deliver remove themselves from the conversation.

One closing observation, offered as a trend judgment. The GEO industry is replaying the early history of SEO: ranking guarantees, black-hat methods, platform countermeasures, consolidation. The difference is that this cycle will run noticeably shorter.

Two reasons. Model iteration is far faster than search algorithm updates, so the shelf life of any exploit is shorter. And model providers are more sensitive to data contamination than search engines ever were — training corpus quality bears directly on model capability, which gives platforms a much stronger incentive to police pollution.

The implication: for firms and clients treating GEO as short-term arbitrage, the window is narrower than it appears. For those treating it as information infrastructure, what accumulates does not expire when rules change — accurate, complete, consistent, verifiable company information is a condition of being trusted under any generation of model.

There is no shortcut here. It is closer to finally saying clearly the things a company should have been saying clearly all along.

The GEO market lacks acceptance criteria, not capability — RMB 9.90 and RMB 280,000 can carry the same promise. No ranking here, only tools: four vendor types, five red lines, seven measurement questions for your contract, and why black-hat GEO leaves contamination that cannot be retracted. Facts are third-party reported; the author is a market participant. - XINMING DESIGN
标签:

发表评论