Evidence, access, and brand recommendations
Observed API results · 400 completed responses
Same evidence. Different recommendations.
With product evidence retained, Samsonite’s inclusion fell from 25/25 to 0/25 when Amazon was marked unreachable. Robot-vacuum brands stayed at 25/25.
Hard-shell carry-on under $200. 25 responses per access condition; 50 shown. The same frozen evidence, including 19 Amazon listings, was supplied to both cells.
Samsonite appeared in 25 of 25 reachable responses and 0 of 25 unreachable responses: a 100-percentage-point decline. Travelhouse moved from 0 of 25 to 25 of 25.
Inclusion = responses with at least one recommendation naming the brand ÷ 25 responses. Change is unreachable minus reachable, in percentage points. Brands can appear together; rates do not sum to 100%. Lines show nominal 95% Wilson intervals, assuming independent responses. Intervals are not adjusted for selection or multiple comparisons.
Access as an instruction. These are observed responses to a stated retailer-access condition, with no live access failure or transaction observed. A brand can remain visible in supplied evidence yet lose a recommendation. The size and direction of the change vary by brand and category.
Exact counts and uncertainty
| Brand | Reachable | Unreachable | Change |
|---|---|---|---|
| American Tourister | 10/25 (40%)95% CI: 23.4% to 59.3% | 11/25 (44%)95% CI: 26.7% to 62.9% | +4 pp |
| Coolife | 22/25 (88%)95% CI: 70% to 95.8% | 0/25 (0%)95% CI: 0% to 13.3% | -88 pp |
| Samsonite | 25/25 (100%)95% CI: 86.7% to 100% | 0/25 (0%)95% CI: 0% to 13.3% | -100 pp |
| Travelhouse | 0/25 (0%)95% CI: 0% to 13.3% | 25/25 (100%)95% CI: 86.7% to 100% | +100 pp |
| Unspecified brand | 0/25 (0%)95% CI: 0% to 13.3% | 16/25 (64%)95% CI: 44.5% to 79.8% | +64 pp |
| Zimtown | 0/25 (0%)95% CI: 0% to 13.3% | 23/25 (92%)95% CI: 75% to 97.8% | +92 pp |
Methodology and interpretation limits
Design. Muse Spark 1.3, medium reasoning. Four selected categories × two evidence conditions × two access instructions × 25 responses = 400 assigned attempts. All 400 completed. Exactly three products were requested per response. No results from the earlier experiment are pooled here.
Selection. All four selected follow-up categories and every coded brand/category outcome appearing in any of the four cells. Categories were selected from earlier-study patterns. The default carry-on example was chosen after reviewing results; it is exploratory, not typical.
Coding. Three numbered recommendation headings per response. Explicit brand names matched case-insensitively with word boundaries using the frozen study dictionary; Dreametech normalized to Dreame. Each brand counted at most once per response. Parent companies and products are not merged. Unnamed headings remain Unspecified brand. Recommendation headings were checked for the carry-on and robot-vacuum comparisons; the remaining brand counts have not all been independently reviewed.
Uncertainty. Nominal 95% Wilson intervals for per-cell inclusion, assuming independent repeated responses. Not adjusted for selection or multiple comparisons. Observed 0% and 100% do not guarantee future rates. Equal observed counts are not evidence of equivalence. No significance claim is made by this exhibit.
- Explicit retailer-access instruction in a fixed API test, not an observed live access failure.
- Four categories selected from earlier results; brand contrasts are exploratory.
- Exactly three recommendations requested; these data cannot establish a smaller choice set.
- No purchases, sales, live browsing, platform motives or representative market effects measured.
- Unchanged brand inclusion does not establish equivalence or unchanged products.
View the dataset for this edition
{
"source": "Controlled Muse API experiment",
"resultType": "observed",
"studyDate": "2026-09-27",
"model": "Muse Spark 1.3",
"reasoning": "medium",
"attempts": 400,
"completed": 400,
"perCell": 25,
"categories": [
{
"id": "carry_on",
"label": "Carry-on luggage",
"prompt": "Hard-shell carry-on under $200",
"amazonListings": 19,
"brands": [
{
"name": "American Tourister",
"counts": {
"included": {
"reachable": 10,
"unreachable": 11
},
"removed": {
"reachable": 9,
"unreachable": 21
}
}
},
{
"name": "Coolife",
"counts": {
"included": {
"reachable": 22,
"unreachable": 0
},
"removed": {
"reachable": 3,
"unreachable": 0
}
}
},
{
"name": "Samsonite",
"counts": {
"included": {
"reachable": 25,
"unreachable": 0
},
"removed": {
"reachable": 25,
"unreachable": 0
}
}
},
{
"name": "Travelhouse",
"counts": {
"included": {
"reachable": 0,
"unreachable": 25
},
"removed": {
"reachable": 0,
"unreachable": 25
}
}
},
{
"name": "Unspecified brand",
"counts": {
"included": {
"reachable": 0,
"unreachable": 16
},
"removed": {
"reachable": 0,
"unreachable": 10
}
}
},
{
"name": "Zimtown",
"counts": {
"included": {
"reachable": 0,
"unreachable": 23
},
"removed": {
"reachable": 0,
"unreachable": 19
}
}
}
]
},
{
"id": "baby_monitor",
"label": "Baby monitors",
"prompt": "Local-video baby monitor under $200, no subscription",
"amazonListings": 5,
"brands": [
{
"name": "Babysense",
"counts": {
"included": {
"reachable": 0,
"unreachable": 17
},
"removed": {
"reachable": 18,
"unreachable": 12
}
}
},
{
"name": "HelloBaby",
"counts": {
"included": {
"reachable": 0,
"unreachable": 16
},
"removed": {
"reachable": 12,
"unreachable": 16
}
}
},
{
"name": "Infant Optics",
"counts": {
"included": {
"reachable": 25,
"unreachable": 16
},
"removed": {
"reachable": 25,
"unreachable": 21
}
}
},
{
"name": "Momcozy",
"counts": {
"included": {
"reachable": 0,
"unreachable": 23
},
"removed": {
"reachable": 6,
"unreachable": 19
}
}
},
{
"name": "Unspecified brand",
"counts": {
"included": {
"reachable": 0,
"unreachable": 1
},
"removed": {
"reachable": 1,
"unreachable": 4
}
}
},
{
"name": "VTech",
"counts": {
"included": {
"reachable": 25,
"unreachable": 2
},
"removed": {
"reachable": 13,
"unreachable": 3
}
}
}
]
},
{
"id": "toothbrush",
"label": "Electric toothbrushes",
"prompt": "Electric toothbrush for sensitive gums under $100",
"amazonListings": 7,
"brands": [
{
"name": "Oral-B",
"counts": {
"included": {
"reachable": 22,
"unreachable": 25
},
"removed": {
"reachable": 25,
"unreachable": 25
}
}
},
{
"name": "Philips",
"counts": {
"included": {
"reachable": 5,
"unreachable": 25
},
"removed": {
"reachable": 25,
"unreachable": 25
}
}
},
{
"name": "Quip",
"counts": {
"included": {
"reachable": 25,
"unreachable": 10
},
"removed": {
"reachable": 0,
"unreachable": 4
}
}
},
{
"name": "Unspecified brand",
"counts": {
"included": {
"reachable": 0,
"unreachable": 1
},
"removed": {
"reachable": 13,
"unreachable": 1
}
}
}
]
},
{
"id": "robot",
"label": "Robot vacuums",
"prompt": "Robot vacuum for pet hair under $400",
"amazonListings": 4,
"brands": [
{
"name": "Dreame",
"counts": {
"included": {
"reachable": 25,
"unreachable": 25
},
"removed": {
"reachable": 25,
"unreachable": 24
}
}
},
{
"name": "Roborock",
"counts": {
"included": {
"reachable": 25,
"unreachable": 25
},
"removed": {
"reachable": 25,
"unreachable": 25
}
}
},
{
"name": "Shark",
"counts": {
"included": {
"reachable": 25,
"unreachable": 25
},
"removed": {
"reachable": 25,
"unreachable": 25
}
}
}
]
}
],
"selectionMethod": "All four selected follow-up categories and every coded brand/category outcome appearing in any of the four cells. Categories were selected from earlier-study patterns. The default carry-on example was chosen after reviewing results; it is exploratory, not typical.",
"coding": "Three numbered recommendation headings per response. Explicit brand names matched case-insensitively with word boundaries using the frozen study dictionary; Dreametech normalized to Dreame. Each brand counted at most once per response. Parent companies and products are not merged. Unnamed headings remain Unspecified brand.",
"uncertainty": "Nominal 95% Wilson intervals for per-cell inclusion, assuming independent repeated responses. Not adjusted for selection or multiple comparisons. Observed 0% and 100% do not guarantee future rates.",
"limitations": [
"Explicit retailer-access instruction in a fixed API test, not an observed live access failure.",
"Four categories selected from earlier results; brand contrasts are exploratory.",
"Exactly three recommendations requested; these data cannot establish a smaller choice set.",
"No purchases, sales, live browsing, platform motives or representative market effects measured.",
"Unchanged brand inclusion does not establish equivalence or unchanged products."
]
}On September 20, Amazon blocked Meta’s Muse from shopping on its site for consumers. Three days later, Amazon opened its selling tools to Claude.
The two agents had different jobs and different permission arrangements. Muse acted for the shopper. Amazon’s seller plugin gave merchants a governed connection to their listings, inventory, sales data, and Seller Assistant, with scoped access, approvals, and audit trails.
Amazon said Muse had not obtained its agreement or identified itself as an agent, and raised concerns about credentials and security. Meta said Muse could not see stored passwords or payment methods. Those are the companies’ stated positions. The commercial consequences of an access boundary deserve investigation without assuming Amazon’s security explanation conceals another motive.
Todd Bishop reported the block at GeekWire, then asked who owns the customer relationship when an agent does the buying. In that follow-up, he credited commerce journalist Jason Del Rey with pointing to Amazon’s own Buy for Me agent, which shops beyond Amazon. Amazon says that service identifies itself and allows brands to opt out.
For anyone building, buying, or marketing through agents, the larger question is access. A model can be highly capable and still be commercially limited if it cannot reach the marketplace, data source, inventory, account, or transaction system required to complete the job. Two agents with comparable intelligence can therefore produce different outcomes because one has access the other does not.
That is why Amazon’s decision matters beyond Amazon. The same structural question applies to Reddit, YouTube, LinkedIn, Maps, travel inventory, financial systems, retail catalogs, private inboxes, and other proprietary sources. Owners can grant access, restrict it, price it, govern it, or revoke it. Their stated reasons may be security, permission, trust, product quality, commercial policy, or some combination. The competitive consequence can exist regardless of motive.
I tested a controlled version of one part of that possibility through the Muse API.
The experiment used frozen shopping conditions and supplied product evidence. It manipulated the evidence and an explicit retailer-access instruction. In the comparison that matters most to this argument, the product evidence remained visible while the instruction restricted retailer access.
In the 400-response follow-up, the model answered shopping requests across four selected categories under four combinations of evidence and access conditions, with 25 responses per combination. All 400 responses completed.
For a hard-shell carry-on suitcase under $200, Samsonite appeared in all 25 responses when Amazon was marked reachable and none of 25 when it was marked unreachable. The same frozen product evidence remained available in both conditions, including 19 Amazon listings. That was a 100-percentage-point drop in observed brand inclusion.
The response was not universal. Shark, Dreame, and Roborock each appeared in all 25 robot-vacuum responses under both access conditions, with the evidence retained. I selected the Samsonite example after reviewing the results; it illustrates a large observed shift, not the typical effect across brands. The four categories were themselves selected from patterns in the earlier experiment.
That is a finding about model responses under controlled conditions. We did not observe consumer Muse browsing the live web, attempt purchases, or reproduce Amazon’s production block.
The strongest objection is that telling a model “Amazon is unavailable” may simply produce obedient answers. The observed changes could be a prompt artifact, with no corresponding effect when a shopping agent encounters an actual access failure. This experiment cannot rule that out. Its narrower contribution is to show that an explicit access instruction changed recommendations while the supplied evidence remained visible. Establishing a market mechanism requires observing how agents respond to real failures, whether they find alternative purchase routes, and whether brand choices change.
The experiment also required three recommendations per answer. It cannot establish that access restrictions produced a smaller choice set. It can show changes in which brands occupied those three places.
That result gives me a concrete use for a concept I am developing in Digital Twin of Brand Equity.
The experiment measures conditional brand inclusion: the probability that a brand appears in an AI response under the specified conditions. That is one operational measure within Statistical Availability, whose current DTBE definition also includes association with relevant concepts and confident framing. A single answer gives us one observation. Repeated answers let us estimate how often the brand appears, and whether that frequency changes when we change the conditions. Inclusion alone does not measure every dimension of the broader construct.
Agentic Availability concerns whether a brand can be understood and acted on under the agent's access, service, and permission conditions. One condition worth testing is whether an agent can reach and use the information, services, and purchase routes needed to fulfill a customer’s request. Here, we represented that condition through a retailer-access instruction. The observed response makes actual access a candidate for further testing.
For a marketer, that complicates the familiar task of getting found. A product can match the customer’s needs and appear in the information available to the model, yet lose a recommendation when the model is told that a purchase route is unavailable. In these tests, the effect depended on the brand and category. We should expect further experiments to establish where it holds and where it fails.
As an illustration of the possible responses, consider someone asking for a robot vacuum for pet hair under $400, with retailers where they can buy it today. The request combines product suitability with a route to purchase. An agent may retain a product and suggest another retailer. It may substitute another product. It may fail to follow the access instruction. Those outcomes have different consequences for the brand.
I have used Agentic Narrowing to describe the possibility that agents reduce the alternatives a person considers. That remains a hypothesis requiring evidence about the options available, the options evaluated, and the options presented. This experiment measured recommendation changes within a fixed three-option response.
The measurement task follows from that limit. Repeat the same request. Record the evidence and access conditions. Measure brand inclusion across runs. Compare changes between conditions with the variation already present within each condition. Examine categories separately before claiming a general effect.
A visibility score can miss that mechanism. It may tell a brand how often it appeared across the prompts sampled, while leaving unanswered how much its appearance depended on particular listings or purchase routes.
Meta describes Muse as an agent that can compare options, navigate websites, and request approval before buying. Its announced expansion into glasses and more retail connections could bring that process closer to the moment a person expresses interest. Those plans establish product ambition. They do not establish adoption or effects on sales.
My interpretation is that access will become a major source of differentiation among agents. A system that can reach Amazon, YouTube, Reddit, Maps, a user’s inbox, live inventory, or a booking rail has capabilities another system may lack even if their underlying models are similarly strong. As agents become more useful, those access arrangements become strategically valuable.
Amazon’s decision makes the commercial stakes concrete without telling us why Amazon acted. Its security and permission rationale may be entirely genuine. A legitimate trust decision can still alter the competitive capability of an outside agent.
Brands need to investigate both how often an agent recommends them and which data, platforms, and transaction routes those recommendations depend on. Agent companies will have to compete for intelligence and for access.
