Gumshoe currently runs two Technical Audit scoring systems side by side, and they score the same page differently. That is expected. They measure different things in different ways, so a gap between the two scores doesn't mean either one is wrong.
What is Technical Audit v2?
Technical Audit v2 is a newer way of scoring your pages. Instead of asking an AI model to judge a page, it runs a defined list of pass or fail checks and builds the score from the results. That makes every result traceable: each one ties back to a specific check.
Technical Audit v2 is available on every plan today. It will become part of the Pro subscription, so if you are on Starter and want to keep using it, that is worth planning for.
How does the original Technical Audit score a page?
The original Technical Audit hands an AI model a structured extract of your page, including schema markup, heading structure, metadata, and internal links, and asks it to assess how ready that page is for an AI system to understand. The model works from that extracted data rather than the live page.
It scores six areas on a 0 to 100 scale: structured data, page layout and structure, schema markup, navigation, content balance, and metadata. Your overall score is the average of those six.
Because a model is making the assessment, two runs on an unchanged page can come back slightly different. That variation is normal for this system.
How does Technical Audit v2 score a page?
Technical Audit v2 makes no judgment call. It runs 20 checks against your page, each of which either passes or fails, grouped into four categories. Machine readability carries the most weight overall, so it tends to be the biggest lever on your v2 score.
Machine readability and data layer
- HTTP status: the page loads successfully.
- Robots blocking: the page is not blocked from AI crawlers by robots.txt or a noindex tag.
- JSON-LD presence: the page has at least one structured data block.
- JSON-LD validity: that structured data block parses correctly.
- Canonical tag: a canonical tag exists and points to the right URL.
- Title quality: the page has one clear title tag of a reasonable length that includes a brand name.
- Open Graph tags: Open Graph title, type, image, url, and description are present.
- Twitter card tags: Twitter card title, description, and image are present.
- Published date: a machine readable publish date can be found.
- Mobile viewport: the page declares a responsive viewport.
Content clarity and answer readiness
- Main content landmark: there is one clear main content region.
- Main content volume: there is enough visible text inside that region.
- Heading hierarchy: headings exist and follow a logical order with no skipped levels.
- Chunked content: paragraphs are not excessively long.
- Text to HTML ratio: visible text makes up a healthy share of the page's markup.
- Content emphasis: key takeaways are visually emphasized.
Authority and trust signals
- Article schema author: article style structured data includes an author field.
- Visible author byline: a visible author credit appears on the page.
Multimodal and agent optimization
- Image alt text: most images have descriptive alt text.
- Media captions: video or audio content has reachable captions.
This list is a working framework and will keep developing alongside our AI visibility research, so expect checks to be added or refined over time.
So why don't the two scores match?
Two reasons, and they compound.
- The groupings are different. The original audit sorts everything into six sections, v2 into four categories, and those do not map onto each other.
- The grading approaches are different. One is an overall assessment by a model, the other is strict pass or fail against a fixed list.
Put those together and there is no reason to expect the two numbers to land near each other, or even to move in the same direction from one run to the next. Read them as two independent signals rather than one correcting the other.
Why are my pages scored against different numbers of checks?
Because a check that does not apply to a page drops out rather than failing. A page with no images is not scored on alt text at all, so it is neither penalized nor credited for it. A few checks also depend on an earlier one: we do not check whether structured data is valid on a page that has no structured data, so that check simply does not run.
This is why v2 scores are most useful for tracking one page over time. Comparing two different pages to each other carries more of an asterisk, since they may have been scored against different totals.
Will my past scores change when you update a check?
No. Every check carries its own version, and each run records exactly which versions were used to grade it. If we later refine how a check works, that becomes a new version going forward rather than being applied backward. Your historical scores stay as they were.
Which score should I be working from?
Use both while they run side by side. The original audit gives you a written highlight and gap summary for each of its six sections, which is useful for exploring a page. Technical Audit v2 gives you specific, traceable pass or fail results, which is better for knowing exactly what to fix.
Pro tip: If you are deciding where to start, work through the machine readability checks first. They carry the most weight in the v2 score, and most of them are one-time fixes to your page template rather than per-page work.