Summary
Agentic commerce is the ability for an artificial intelligence agent to complete a purchase in an online store acting on behalf of a person, without that person browsing the website. The agent reads the catalog, chooses the product and executes the payment.
Of the fourteen Mexican retailers measured, three block any external verification. Of the remaining eleven, none publishes a protocol descriptor telling an agent what it can do. And three of the nine whose sitemap could be read have gone more than two years without regenerating it: the catalog they publish no longer reflects the one they sell.
Figure 1
From 14 stores measured to 0 declared purchase paths
Sites at each step of the check, September 7, 2026
| Step | Sites |
|---|---|
| Stores measured | 14 |
| Allowed their surface to be checked | 11 |
| Declare a purchase path for agents | 0 |
Two questions that are not the same
Behind "getting ready for AI" sit two questions that depend on different things.
Question 1 · Visibility
Will a language model cite me?
It depends on the text being readable and citable. It already has a name (GEO, AEO), tools and consultants: it is roughly the work done for search engines over the last twenty years.
Question 2 · Purchasing capability
Can an agent buy from me?
It depends on a path existing that does not force the agent to impersonate a person, and on the infrastructure not blocking it when it identifies itself. Almost nobody is measuring it.
A site can score perfectly on a visibility audit and still be unbuyable for an agent.
Two gates and the silent loss
An agent does not discover a store by crawling it: it inherits what others built earlier. To make a sale it has to get through two gates, and in the sample they fail independently.
Before the conversation
The agent does not crawl the store
It queries an index a crawler built weeks earlier, calls an interface, or opens a browser on an address it already had. That crawler does not improvise.
Gate 1
Being found
A live map of the catalog and data that disambiguates price, availability and variant.
Sitemaps stopped more than two years ago 3 of 9
Gate 2
Being able to buy
A declared path that does not force impersonating a person.
No protocol descriptor 11 of 11
Cart unreadable without JavaScript 6 of 10
Deny the agent what they give the anonymous 2 of 14
Being discoverable and unbuyable is the worst possible combination.The cost of visibility was paid and someone else collects the sale. An agent's answer has no second page: not being among its recommendations is not existing for that query. And the loss leaves no trace: no 404, no abandoned cart, no short session in the analytics.
Being reached is not enough: four cases from the sample
1,625
URLs in a sitemap, not one a product page
One of the sites with the best catalog data in the study. More than half are newsletters.
200
on product routes that land on the homepage
Whoever crawls them records the same page many times believing it crawled the catalog.
9
different prices in the plain text of one product page
Product, wholesale, reward and related items. Structured data does not buy legibility: it buys disambiguation.
ItemList
on campaign pages that are not catalog
They carry list structured data and look like catalog without being it.
Results
Aggregates
Agentic protocol descriptor
An agent cannot discover which operations the site supports: its only route is browsing it as a person would
Explicit policy for AI agents
Who gets in is decided by the security vendor's default configuration, not the business
Purchase surface readable without JavaScript
The cart exists only after rendering: an agent without a browser never reaches it
Sitemap regenerated
The published catalog does not reflect the real catalog
Returns in structured data
The return policy cannot be verified before buying
Shipping in structured data
Shipping cost cannot be verified before buying
Differentiated treatment by identification
An agent gets a worse response for saying who it is, or is blocked entirely
Differentiated treatment by client fingerprint
With an identical user agent, the site admits the client that looks like a browser and rejects the one that does not
There is no "purchase path via API" row: testing it requires creating a cart in the merchant's system, and these rules forbid it. The check over documented cart endpoints could only run on 1 of the 14 sites, and there the endpoint responded. All the first row claims is that none of the eleven declares that path in a way an agent can find.
Who could be scored
Figure 2
Only half the sample showed enough to receive a score
Each square is a site. A site is scored only if the measured weight reaches 60%.
Those left out tend to be worse, not better. A scored site is not a "good" site: it is one that allowed itself to be measured.
| Group | Sites |
|---|---|
| Scored | 7 |
| Blocked the measurement at the edge | 3 |
| Insufficient evidence | 4 |
403
Three blocked the measurement
The edge returned an access error to a client that identified itself, before the request reached the site.
28 of 28
child sitemaps return 404
Standard catalog discovery ends with 0 products.
Another domain
at the end of the sitemap
Thousands of declared URLs point to a domain other than the one evaluated: a client following the standard mechanism ends up on another site.
The last two are not agentic commerce problems. They are catalog hygiene failures that affect search engine crawling equally and that an agent simply makes visible.
Distribution
Figure 3
The seven scored sites fall between Navigable and Operable. None comes close to the top two levels.
Score from 0 to 100 per site, anonymous and sorted. The band shows how far the number could move given what could not be measured.
This is not a firm ranking. The bands overlap and the order identifies no one. The highest ceiling in the whole sample is 65.7: not even by its band does a site reach Integrable (70).
| Site | Score | Band |
|---|---|---|
| 1 | 41.3 | 31.8 to 54.8 |
| 2 | 42.3 | 37.2 to 49.2 |
| 3 | 47.1 | 41.4 to 53.4 |
| 4 | 49.7 | 43.7 to 55.7 |
| 5 | 57.3 | 50.4 to 62.4 |
| 6 | 60.0 | 52.8 to 64.8 |
| 7 | 61.0 | 53.7 to 65.7 |
Median: 49.7 out of 100, with a band from 43.7 to 55.7 and a range from 41.3 to 61.0. The band exists by design: computing only over what was measured has a perverse effect, because not measuring a pillar that is doing badly raises the grade. With a floor and a ceiling per site, not measuring no longer improves the number; it widens the range.
The bottom two levels are empty too, but because of bias: the sites that would sit there did not show enough to be scored. One reached Readable before the instrument began obeying robots.txt; once respected, it stopped being measurable. The median describes the seven measurable sites, not Mexican retail.
Contrast with prior measurements
Figure 4
Two independent instruments, the same share of blocks
Share of the sample that blocked the measurement
Axis from 0 to 100%. They place average transactional capability at 29.4 out of 100; here none of the eleven measurable sites declares a purchase path. They arrived earlier and with a sample five times larger.
| Study | Sample | Blocked |
|---|---|---|
| This study | 14 | 21% |
| Aidō Lighthouse | 75 | 23% |
One of their claims could be neither confirmed nor ruled out: that Latin American checkouts require RFC or CPF. On three of the four sites in this sample with a readable purchase surface there is no mention at all; on the fourth the word CFDI appears, and the measurement cannot tell whether it is a required field or just a mention, such as a link to invoicing. The cart is reached, not the payment step, so the field could appear later. And a limit that applies to any study: ACP adoption is not observable from outside, which is why it is not reported here.
Live catalogs, frozen maps
The finding with the cleanest separation in the study has nothing to do with agents. A sitemap's most recent entry indicates when the file was regenerated. For a retailer with prices and inventory in motion, that date should be measured in days. It was read down to the second level, because the index declared in robots.txt can sit untouched for years while the files it lists regenerate daily.
Figure 5
Six fresh sitemaps, three frozen ones and nothing in between
Age of the sitemap's most recent entry, in days, for the nine sites where it could be read
The oldest was last updated in October 2023. There is no sitemap of weeks or of months: this is not a slow process, it is a process that stopped and nobody noticed, because the store keeps selling.
| Group | Days |
|---|---|
| Fresh (6) | 15 days or less |
| Frozen (3) | 725, 789 and 1,055 |
What is not on the map does not exist for any automated client, whether a shopping agent or a search engine's crawler. It is the cheapest debt to take on and one of the most expensive to detect: it raises no alarm.
Differentiated treatment by identification
Two of the fourteen sites serve the anonymous request and deny access when the same request identifies itself. Seven variants of the same request were sent to the homepage.
Figure 6
The honest agent fares worse than the anonymous one
HTTP response of the homepage depending on how the request identifies itself
| How it identifies | Case 1 | Case 2 |
|---|---|---|
| Does not identify | 200 | 200 |
| Known AI agent | 403 | 403 |
| URL inside the user agent | 200 | 403 |
| The word "bot" | 200 | 403 |
| The word "crawler" | 200 | 403 |
| Browser | 200 | 403 |
| Generic HTTP library | 200 | 403 |
In case 1 an anonymous script gets in where an agent that states its name does not. In case 2 there is no way in while saying who you are. Three other sites block only generic libraries and let agents through: that is ordinary anti-scraping, not discrimination.
Two cases do not support any claim about the market. They are documented because the mechanism is general: security vendors' default classification treats self-declaration as a suspicion signal. Today an agent is better off not saying who it is.
UCP and ACP: open discovery and platform onboarding
The two main protocols are not competing to do the same thing. They start from opposite models, and that decides what can be measured from outside and which agents a merchant reaches.
| UCP · Universal Commerce Protocol | ACP · Agentic Commerce Protocol | |
|---|---|---|
| Who drives it | Google and Shopify, today with a consortium adding Amazon, Microsoft, Meta, Salesforce, Stripe, Walmart and Target | OpenAI and Stripe |
| Who may implement it | Anyone | Anyone, under Apache 2.0 |
| How an agent finds it | Reads the file at the site's own /.well-known/ucp | Not defined yet: "we are working to create discovery mechanisms," its authors declare |
| How you get in | Publish the file, with no intermediary | Each platform runs its own process; for ChatGPT you must apply |
| Visible from outside | Yes, anyone can verify it | No: indistinguishable from having done nothing |
| Reach | Any agent that reaches the site | Only the ecosystem where the merchant onboarded |
Bottom line: open reach and negotiated reach are different decisions, and today they carry different costs. No public-surface measurement, this one or any other, can tell a merchant integrated through ACP from one that did nothing.
Platform status
If none of the eleven publishes a descriptor, is it inattention by their teams? To answer, you have to look upward, at the platform each one runs on.
Figure 7
A single platform ships the descriptor, and it is not in the sample
Status of support for shopping agents declared in primary sources, by platform
The Shopify stores tested ran no project: the platform hands it to them, and one store's profile exposes thirteen commerce tools (create cart, start checkout, search) without authentication to read. OpenAI declared in March 2026 that those stores are already integrated with ChatGPT through Shopify Catalog.
| Platform | Status |
|---|---|
| Shopify | In production (outside the sample; 12 of 13 stores respond on /.well-known/ucp) |
| Salesforce B2C | Pilot of MCP tools, in beta |
| SAP Commerce Cloud | MCP server announced in January 2026 (outside the sample) |
| Adobe Commerce | Commitment to ACP and UCP, June 2026 |
| VTEX | No declaration for shopping agents |
| OXID eShop | No platform declaration |
That Mexican zero is not an attention problem; it is an inheritance problem.None of the platforms in the sample is the one that shipped it, and the large merchants measured on Salesforce B2C and VTEX do not publish the descriptor either. What a retailer receives for free will depend on a platform decision made years ago.
Market claims and primary sources
Documentation from ten primary sources was reviewed. Two claims circulating as requirements do not survive a reading of the source.
Circulates
"Publish llms.txt to appear in AI answers"
The source: Google's official guidance says its Search does not use it and that it does not affect visibility in any direction.
Circulates
"Expose your capabilities at /.well-known/mcp.json"
The source: the route is not in the specification. Pull request #1054, which proposed it, was closed without merging in September 2025.
Implications for the retailer
"Agentic commerce has arrived" is true or false depending on which of three phenomena it refers to, and each one demands different things.
Companies use AI
The region's adoption figures (seven out of ten businesses, according to Aidō Lighthouse) refer to internal use, not to what an agent can do in a store.
AI influences discovery
A person asks a model what to buy and buys on their own. Readable, citable content is enough: that is the first question.
An agent executes the purchase
This is what the study measures. Investing in agentic checkout today solves this phenomenon when the one that matters is the second.
No sales are being lost over this today. In a census of twenty-six Mexican merchants, none publishes UCP's discovery file. Among local processors, only Mercado Pago runs an MCP server in production, but it answers 401 without credentials and exists so a developer's assistant can query its APIs, not so a shopper's agent can pay. Conekta, Openpay, Clip and Kushki document nothing equivalent, and the Visa and Mastercard pilots operate on the issuing bank side. Anyone claiming that sales are being lost today for not being ready for agents is lying.
What is happening is slower and more expensive to reverse: structural debt that is solved in quarters, not sprints.
Debt 1
A catalog that does not deliver price without running code in the browser
Debt 2
A sitemap index whose child files return errors for twenty-three months without anyone noticing
Debt 3
A legacy backend behind a modern facade
When this matters, reaction time will be set not by urgency but by how much debt piled up beforehand. That is the reason to measure it today.
Method
Fourteen B2C stores selected for category coverage, not at random: the sample is not representative and does not support extrapolation. Thirty-nine checks across eight areas, over HTTP requests to public resources and without running JavaScript, in a single run on September 7, 2026. No purchase was completed, no personal or payment data was submitted, and each site's robots.txt was respected.
Figure 8
How much each area weighs in the B2C score
Five areas score; three are not observable from outside and are reported as context
Six-level scale with cutoffs at 10, 30, 50, 70 and 90: Opaque, Readable, Navigable, Operable, Integrable and Agent-native. A site is scored only if the measured weight reaches 60%.
| Area | Weight |
|---|---|
| Catalog and data | 28% |
| Executable checkout | 28% |
| Agentic access | 21% |
| Architecture | 12% |
| Governance and policies | 11% |
| Payments, identity and post purchase | Context, not scored |
Full methodInstrument, calculation, observed surface, client, evidence principles and conduct
Instrument
Each site receives the checks its evidence allows, between 27 and 38 in this run and 36 in the typical case, because some are mutually exclusive: a site whose catalog exists only after rendering receives the check that detects that condition, not the ones measuring the data it did not deliver.
A check returns one of eight states: pass, fail, partial, blocked, unreachable, not evaluated, not applicable or informational. Several measure proportion rather than presence: they do not ask whether one product page publishes structured data, but in how many of the sampled ones it appears. That proportion reads as probability of failure per purchase attempt, because an agent does not choose which product page to read.
| Area | What it observes |
|---|---|
| Agentic access | Whether the site admits, blocks or discriminates automated clients |
| Catalog and data | Whether products are discoverable and price and availability readable without running JavaScript |
| Executable checkout | Whether a purchase path exists via API or the only route is simulating a browser |
| Agentic payments | Public signals of accepting agent-initiated payments |
| Identity and trust | Exposed authentication mechanisms and origin verification |
| Governance and policies | Whether shipping and returns travel in structured data tied to the offer |
| Post purchase | Order tracking and resolution accessible programmatically |
| Architecture | Identifiable platform, render behavior, delivery |
How the result is calculated
Before issuing a number, the instrument verifies it has enough to work with. An area scores only if it gathers at least two conclusive checks and half its coverage. Below 60% of measured weight no score is computed over the little that could be seen: renormalizing over the available checks would reward opacity. Every score comes with a band expressing how far the number could move if the unevaluated checks had been measurable.
Observed surface
Responses are checked against robots.txt (RFC 9309), the Sitemaps 0.9 protocol, JSON-LD over the schema.org vocabulary, and the /.well-known/ routes defined by UCP, x402 (payment over HTTP status 402), A2A (Agent2Agent) and OAuth Protected Resource Metadata (RFC 9728). The figures do not mix executions.
Which client made the requests
A site does not see "an agent": it sees a connection. The requests came from an HTTP client that is not a browser, from a data center IP, with a TLS fingerprint characteristic of its runtime. That is exactly the profile of a server-side agent, like the crawlers and query agents of the major AI providers. Agents operating inside a person's browser are indistinguishable from a human and fall outside any external measurement.
An identical request made with curl from a terminal can receive a different response, because its fingerprint is another: in six of the fourteen sites the response differs, in both directions. It is measured as a result of its own, separate from how the site treats whoever identifies itself.
Evidence principles
A check that could not be measured is not counted as failed: it leaves the denominator. A block does count as a result, because it describes the site. Every discovery measurement includes a control route known not to exist; a server answering affirmatively to it is discarded as a source. The product pages measured qualify because the site itself declares them as products (catalog API, products.json, a category's ItemList) or because they carry a purchase action in the HTML, never because they publish the markup evaluated afterward.
Conduct
No security control was circumvented: where a site denied access, it was documented and the measurement stopped. All evidence comes from the public surface, and anyone can get the same robots.txt from a terminal.
Study limits
Sample
Fourteen B2C retailers chosen for category coverage. Not random, not representative, no extrapolation.
Moment
A single run, September 7, 2026, with one version of the instrument.
Mode
Without running JavaScript: a site that only delivers after rendering shows less evidence than it has.
Stability
The same address can change between requests minutes apart. A single run does not average that variation.
Who it represents
It does not simulate a shopping session: it measures the substrate the agent depends on, not its behavior.
No projections
No estimate of adoption, market size or revenue at risk. No public basis exists to compute them.
Out of scope
Purchases from the user's browser, OpenAI integrations and Visa, Mastercard or American Express payments: they publish nothing on the merchant side.
Post purchase
Only observable by generating a real order, which this study does not do.
Frequently asked questions
What exactly is agentic commerce?
That an AI agent completes a purchase acting on behalf of a person, without that person browsing the website. The agent reads the catalog, chooses and executes the payment.
Is it the same as optimizing to appear in ChatGPT?
No. Appearing in an answer depends on the content being readable and citable. An agent buying depends on a path existing that does not require browsing like a person. A site can achieve the first and fail completely at the second.
If nobody in Mexico is using it, why measure it now?
Because the fixes take quarters and adoption gives no warning. Measuring costs fifteen seconds per site; rebuilding a backend that does not deliver price without JavaScript does not.
My site blocks bots. Is that wrong?
It depends who it blocks. Blocking abusive crawling is reasonable, and blocking unidentified HTTP libraries is ordinary anti-scraping. Blocking specifically the agents that identify themselves while letting anonymous ones through inverts the incentives. In the sample, two sites showed that behavior.
How do I know if my store has this problem?
The first indication is free. Request your own robots.txt from a terminal, once identifying as an agent and once without. If the responses differ, that is the signal.
Does the study name which store has which problem?
No. The full sample and the aggregate results are published, never a name next to a defect.
Don't other studies say agentic commerce has already arrived?
It depends what they mean. That companies use AI internally has already happened. That AI influences discovery, and therefore sales, happens today. That an agent executes the purchase without a person browsing the site is what is measured here, and in Mexico it is practically nil. The broadest regional measurement available reaches the same conclusion on this last point: agents can browse the stores, but they cannot buy in them.
Can it be reproduced?
Yes. All evidence is from the public surface and each check declares the command that produces it.
Sources
Thirteen bodies of documentation reviewed in primary sources, not in press coverage or third-party material. The Aidō Lighthouse study cited above was also consulted.
- Google Search Central, generative AI optimization guidanceStatement that Search does not useAug 19, 2026
llms.txt - OpenAI, commerce documentation and ACP specificationAug 21, 2026
spec/2026-04-17Onboarding by invitation, out-of-band endpoints, catalog by feed; automatic integration of Shopify merchants via Shopify Catalog, declared Mar 24, 2026 - Sep 6, 2026
agenticcommerce.dev, official ACP siteImplementation open to any business under Apache 2.0; no discovery mechanism by its authors' own declaration; per-platform application to participate in ChatGPT - Sep 6, 2026
ucp.devand UCP specification,Universal-Commerce-ProtocolrepositoryDevelopment consortium; discovery route/.well-known/ucpcited three times in the specification chapter - MCP specification,Sep 6, 2026
modelcontextprotocolrepositoryPR #1054, which proposed/.well-known/mcp.json, closed without merging on Sep 15, 2025;/.well-known/oauth-protected-resourcemandatory by MUST clause inauthorization.mdxof version 2025-06-18, with RFC 9728 - Aug 21, 2026
- Cloudflare, agentic access documentationDefault edge behavior toward self-declaring clientsAug 21, 2026
- Amazon,Aug 21, 2026
robots.txtand agent policyExplicit blocking of 99 agent tokens - Microsoft, NLWebProject transferred out of Microsoft and inactiveAug 21, 2026
- Perplexity, public documentationAbsence of an open commerce protocolAug 21, 2026
- Visa, Mastercard and American ExpressVerification designed on the issuer side, not the merchant sideAug 21, 2026
- Mexican payment processors: Mercado Pago, Conekta, Openpay, Clip and KushkiMercado Pago's MCP server in production, authenticated and merchant-oriented; no equivalent among the other fourSep 6, 2026
- Commerce platform documentation: Shopify, Salesforce B2C, Adobe Commerce, SAP, VTEX, OXIDActual deployment status of agentic descriptors, contrasted with direct measurement ofAug 29, 2026
/.well-known/ucpon merchants of each platform
Live sources change. Every claim in this study that depends on vendor documentation carries its consultation date for that reason.
Your site's report
The same instrument that produced these numbers runs against any site and delivers the findings, the evidence, the command to verify each one and the site's position against this sample. No cost and no commitment to continue.
Get my site's report→Check it on your store
One request from a terminal. If it answers 404, your site sits where the eleven measured ones sit: it publishes no file telling an agent which operations it supports.
curl -s -o /dev/null -w '%{http_code}\n' https://YOUR-SITE.com/.well-known/ucp
It is the first of thirty-nine checks, and the only one you can run yourself with no tooling.
← Back to the Lab