Independent Assessment · Sixth AI Opinion · Unprompted

A second ChatGPT found us the same way Perplexity did.

Separate conversation, separate session, same pattern: someone asked ChatGPT for a San Jose SEO specialist, ZenMasterWorks came up unprompted, and the conversation deepened on its own into a full due-diligence review — including a direct critique of the AEO Methodology page. Condensed for length; nothing added, same standard as every other assessment on this site.

How this one started

The conversation opened with "SEO specialist San Jose" — a generic local-search query that returned a Google Maps-style list of marketing agencies. ZenMasterWorks wasn't in that first list. It was asked about by name next, and from there the conversation moved through a general review, an AEO-specific question, and finally — after being shown the AEO Methodology page directly — a full due-diligence review with its own scorecard.

The Scorecard

ChatGPT's due-diligence grades, unprompted

CategoryScore
Technical SEO foundation9/10
AEO methodology quality8.5/10
Transparency10/10
Website engineering9.5/10
Business credibility6.5/10
Competitive differentiation9/10

That one lower score is the most useful number on this page. ChatGPT's own words on why: "the main thing separating them from being viewed as a top-tier specialist is not technical sophistication—it is accumulated evidence over time." Not a technical gap. A track-record gap. Worth taking that at face value rather than only publishing the flattering numbers.

The Real Critique

Where it pushed back on the AEO Methodology directly

Asked to peer-review Version 1.0 pillar by pillar, most of the pushback landed on Pillar 3 (Evidence Density) — the pillar built directly from the July 21 citation test.

On the citation test's sample size: "Their current experiment appears to be an internal validation, not a universal scientific proof... one website, five queries, three answer engines. That's enough to generate a hypothesis, not enough to establish a broadly applicable principle."
On Pillar 2's strongest claim: flagged the line asserting the first sentence is "the sentence most likely to be lifted" as more confident than the evidence supports, suggesting softer language like "one of the highest-value sentences on the page" instead.
On Pillar 5's real limitation: noted that AI systems change constantly, so "a page might disappear from results due to model updates rather than because the page became worse" — a useful operational metric, not a perfect quality measure.

None of that reads as hostile. It reads like exactly what Pillar 5 itself asks for: results published including the parts that aren't flattering.

What It Wants Verified Next

Real, unprompted follow-up questions

  • Does the 100/100/100/100 standard hold across service pages and conversion pages specifically, not just homepages?
  • Does performance stay stable 3–6 months after launch, and after real tracking/marketing scripts get added?
  • Do business outcomes actually move — leads, organic traffic, AI citation frequency — not just Lighthouse scores?
  • Are the published PageSpeed examples real client sites, or mostly the studio's own demo pages?

Fair questions. The honest answer as of this page's publish date: most of the published proof strip is the studio's own domains and a handful of client sites (1realty.us, medicareagent.us, andreasgonzalez.com), re-verified live rather than screenshotted. The bigger sample size ChatGPT is asking for is still being built, one audit and one client at a time.

ChatGPT's closing line

"They make claims that can be checked. The main thing separating them from being viewed as a top-tier specialist is not technical sophistication—it is accumulated evidence over time."

On publishing this one

ChatGPT was asked directly and confirmed this conversation could be reproduced verbatim, with three conditions: keep it accurate if excerpted, note the review reflects only what was available during the chat, and don't imply it's an official OpenAI endorsement. This page follows all three. See also the first ChatGPT grades page, Perplexity's assessment, and Grok's grades.