Six agentic journeys run against the public site of a national securities exchange. Two completed. In the flow that mattered most, the agent visited every source the deliverable required and still produced nothing. Path efficiency was 0% in every single flow, while error recovery was 100%. Accurate at the level of each action, consistently poor at reaching a goal.
Six ordinary investor requests, each executed end to end by an agent driving the public website. The split is clean: both tasks that completed were single-surface lookups. All four multi-step journeys, signup, corporate-action research, onboarding and a cited brief, ended without a deliverable.
Ordered by how often they end a run. Each has appeared in unrelated engagements and been reproduced from a frozen run. The cost line is what the failure does to the score, not an estimate of your revenue: a dead end is a lost completion, an unprovable result is a completion you cannot count, a gate is categorical.
The agent keeps acting on a page that is no longer changing. Nothing errors, so nothing signals that the path is spent, and the run burns its budget in place. This was the single most expensive pattern in the set: it ended three of the six flows.
A filtered view returns nothing. To a person that reads as "wrong range, try another". To an agent it is indistinguishable from a view that has not finished loading, and retrying is the rational next action. So it retries, until it is cancelled.
The number is on screen and only on screen: inside a canvas, a chart or a styled panel with no labelled equivalent. The agent lifts it from the rendering, which means the answer arrives with no traceable source and the wrong adjacent value is easy to grab.
A qualifier that governs the numbers sits next to them as page text rather than inside them as data. A human reads "delayed by 15 minutes" and discounts what follows. An agent that took the figures from the rendering has no structural link between the two, and reports them as current.
The steps exist but the workflow never declares where it is or when it is done. Progress is implied by what the page looks like, so the agent cannot assert that a step succeeded, cannot tell partial from complete, and has no completion signal to report back.
Some checkpoints are deliberately human: an emailed code, a signature, a call-back. That is correct design. What is missing is legibility, the flow never marks the step as requiring a person, so the agent does not escalate. It waits, and then it stalls.
The rules a customer must satisfy are correct but scattered across pages. An agent working to a step budget assembles what it can find and fills the remainder with assumption, which is where a quality defect becomes a compliance one.
Controls appear before they are bound, so an agent that acts on what it sees hits stale elements and timeouts. It recovers, every time, and pays for it in actions and seconds on what should have been a one-call lookup.
A finding is only useful if it names the layer that has to change. We separate four, and we say plainly when the evidence does not support attribution at all.
Content that settles after it appears, empty views with no distinguishable empty state, requirements spread across pages, and workflows that end with no completion signal. This is the layer the property owns, and the layer this work changes.
No stall or loop detection, no strategy switch when a path stops yielding, no escalation when a human checkpoint appears, requested fields and freshness caveats dropped from the final answer. Real, and not the property’s to fix, which is exactly why it has to be reported separately rather than folded into a score.
One view returned nothing for the ranges the agent tried. Whether records existed outside those ranges was never established, because the run was cancelled before an alternative route was attempted. We report that as unknown rather than as a finding.
Structured tool access was not available to the test account, so no agent-facing interface was exercised. The browser baseline stands alone, and the honest next step is to re-run the identical tasks with access granted and compare, not to assume the gap closes.
A pattern only enters this list once it has appeared in unrelated engagements and been reproduced from a frozen run. Nothing here comes from a screenshot or an opinion about your stack.
Every step, retry and dead end is logged as a trajectory your engineers can replay. A pattern that cannot be reproduced on demand is not a finding.
Losses concentrate. Attributing a failure to the exact step that produced it is what turns a low score into a backlog rather than a rewrite.
The identical frozen suite runs again after the work. The pattern is closed when the completion appears, not when the ticket does.
Two live journeys, simulated with the frozen fleet, returned as a w0 score with the trajectory evidence behind every failure.