Was it found
Whether the product, its terms and its entry point can be retrieved as facts, and whether your access policy lets a permitted agent in at all.
Beyond AI visibility
Most tools that claim to measure how ready your website is for AI answer one question: can a model retrieve and read this page. It is a fair question and it is not the one your business runs on. The question that matters is whether an agent your customer sent can finish the job, safely, and at a cost that still makes the product worth selling.

Visibility measurement grew out of search. It looks at whether your content is crawlable, whether it is cited by assistants, whether the markup is clean, and whether the site publishes the files a model expects to find. All of that is worth doing. None of it tells you what happened after the agent arrived.
The gap shows up as soon as you follow a real journey. An agent can be served a perfectly indexed rates page and still be unable to say which rate applies to a nine month tenure for a customer of that profile. It can find the application form and still lose everything it entered at the identity step. It can submit and receive nothing it can hand back. A visibility score reads all three of those as success.
Readiness is not a property of your content. It is a property of your journey.
Crawlability, citations, markup, published policy files. Measured on the page.
Completion, recovery, consent, evidence and cost. Measured on the journey.
Invisible means the agent never arrives. Visible without ready means it arrives and leaves empty.
An agent has to clear each layer to reach the next, so the score is not an average of independent factors. A property can be excellent at the first two and return nothing, which is exactly the pattern we keep finding in banking and insurance.
Whether the product, its terms and its entry point can be retrieved as facts, and whether your access policy lets a permitted agent in at all.
Whether the agent can tell which rate, fee, tenure or exclusion governs its case, rather than collecting every number on the page.
Identity, conditional forms, one-time passwords, step-up authentication and vendor hand-offs, each measured for whether accumulated state survives it.
The share of benchmark tasks completed, the errors hit, and how many of those errors the agent could recover from without human rescue.
Whether a verifiable result came back: a reference, a status, a receipt the caller can hand to the customer who sent it.
The five roll up into one reading per journey, and the journey readings roll up into one reading for the property. Both are reported with the trajectory evidence attached, because a score without the run behind it is an opinion with a number on it.
This is the whole methodological argument, and it is why the measurement is built the way it is. Agents vary: the same model, given the same instruction twice, can take different paths, spend different tokens and fail in different places. A measurement that moves with that variance is useless for a regulated release process.
A benchmark fleet and task library, frozen between runs. When the number rises, the property changed, not the probe. It also means the improvement survives the next model release.
Each task runs repeatedly. A journey that completes sometimes is a different risk from one that completes reliably, and a spread that wide is itself a finding for risk and operations.
Runs happen in an environment you nominate, rate-bound and logged, stopping before anything irreversible. If your access policy refuses the agent, that refusal is recorded as a finding.
What comes back is not a grade. It is a reading per layer, the failure point for every incomplete task, the token and step cost of the ones that completed, and a remediation list ranked by score movement per unit of engineering effort.
The reason to compress five layers into one number is not simplicity, it is shared language. Risk needs to know whether a delegated journey is safe. Engineering needs to know which surface to fix first. Marketing needs to know whether the traffic it is buying can convert when the visitor is software. The same reading answers all three, as long as the evidence travels with it.
Consent scope, confirmation gates, audit trail and the failure modes an agent can reach. Refusals are recorded rather than routed around.
A backlog ranked by score movement per unit of effort, with the failing trajectory attached to each item so it can be reproduced locally.
The step that ends most attempts, and what it costs in tokens and retries to reach it. Usually not the step teams expect.
The same frozen suite, re-run. One number per journey and one for the property, comparable across releases because the probe did not change.
Two limits worth stating plainly. The reading decays: agentic readiness is not a certification you hold, and a release that reworks a form or swaps a payments vendor can move it materially in one sprint. And it is not a league table. We do not publish client readings, journeys or trajectories, including on our own site.
Two weeks, a permitted environment, synthetic identities, and a w0 score with the trajectory evidence behind it. No change to your stack.