# Agentic Landmark > Advisory for agent readiness on both surfaces: the Agent Readiness Index for the agents that choose your brand, and an operational twin for those you run. Public Ghost content for AI and LLM tooling. This file includes a bounded export of public pages first, then recent public posts. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages ### Privacy Policy URL: https://www.agenticlandmark.com/privacy/ Last updated: 2026-08-19T15:26:40.000Z # Privacy Policy **Agentic Landmark** · agenticlandmark.com · Last updated July 2026 Agentic Landmark is an advisory practice operated by Jim Cook. Its services are directed to organizations in the United States, the European Economic Area, and the United Kingdom. This policy describes what personal information this site handles, why, and what your choices are. It is written to the standard of the world's stricter privacy regimes because this practice advises on consent architecture, and its own policy should meet the bar it sets for clients. ## What this site collects, and why **Inquiries.** If you use the contact form, the practice receives your name, work email, your message, and any optional fields you choose to complete (vertical, role, and the page you arrived from). This is used for one purpose: to evaluate and respond to your inquiry. The lawful basis is legitimate interest in responding to people who contact a business. The practice deliberately separates your contact details from its long-term working records. When you submit the form, your full inquiry, including your name and email, reaches the practitioner in two places: a notification email, and a row in a private spreadsheet that serves as the practice's inquiry inbox. For inquiries that become active prospects, the practitioner records only the business context (your message, vertical, role, and source page) as a pipeline record in a private repository, linked back to your inquiry by a reference. Your name and email are never carried into that pipeline; they remain in the notification email and the inbox. **The Landmark Brief.** If you subscribe to the practice's publication, your email address is held by Ghost, the membership and publishing platform this site runs on, and used solely to deliver new issues. The lawful basis is your consent. Every email includes an unsubscribe link, and unsubscribing removes you from delivery immediately. Subscriber information stays in Ghost and is not copied into any other system. **Security and technical processing.** The contact form submits to a Google Apps Script endpoint, so Google processes your submission in transit. The site's hosting platform keeps standard server logs. The site's typefaces load from Google Fonts, which means your IP address is transmitted to Google when a page loads. The site sets only strictly necessary cookies (Ghost's membership session cookies); it runs no analytics trackers, no advertising pixels, and no marketing cookies, which is why you will not see a cookie consent banner here. ## What this site does not do Your information is never sold, rented, or shared for advertising. There are no ad networks, no data brokers, no lookalike audiences, and no enrichment services. The practice does not use your inquiry to add you to a mailing list; contacting the practice and subscribing to the Brief are separate choices, and neither triggers the other. ## Service providers The practice uses a small number of providers to operate this site, each receiving only what its function requires: **Ghost** (site hosting, publishing, and Brief membership), **Google** (form processing through Apps Script, email and the inquiry inbox through Google Workspace, and typeface delivery), **GitHub** (private pipeline records containing inquiry context without contact details), **Anthropic** (the AI tool that reads that pipeline context during prospect research, receiving no contact details), and **Cloudflare** (domain and DNS). ## How this practice uses AI Agentic Landmark advises brands on agent-mediated commerce and on governing the agents they run themselves, so it should be plain about where AI sits in its own work. The practice uses Anthropic's Claude as a working tool: for research, for drafting and editing content, for developing methodology, and for building and maintaining this site. One part of this meets your inquiry. Once an inquiry becomes an active prospect, the practice runs an agentic workflow: an automated process in which a software agent built on Claude assembles a first-pass research dossier on its own. The agent reads that prospect's pipeline record (the business context described above, with no contact details) and gathers publicly available information about the company, then opens its draft for the practitioner to review before anything is used. It works only from business context and public information, it never reads your contact details, and it never contacts you. Consistent with the separation described above, your name and email stay in the notification email and the inquiry inbox and are not sent to any AI tool. Nothing you provide is used to train AI models. This section is provided as a matter of transparency about how the practice works. The providers that process your personal data, including the AI tool above, are listed under "Service providers" above. ## International transfers The practice and its providers process data in the United States. Where providers transfer data from the European Economic Area or the United Kingdom, they rely on recognized safeguards: Cloudflare and Google are certified under the EU-U.S. Data Privacy Framework, and providers otherwise use standard contractual clauses. ## Retention Inquiry emails and the entries in the inquiry inbox are kept while an engagement conversation is live or reasonably prospective, then deleted. Pipeline records, which contain no contact details, are kept as business records. Subscriber emails are kept until you unsubscribe. You may request earlier deletion of anything at any time. ## Your rights Wherever you are located, the practice honors the rights the strictest regimes grant: you may request access to the personal information held about you, correction of it, deletion of it, restriction of or objection to its processing, a portable copy of it, and withdrawal of any consent you have given. To exercise any of these, email [hello@agenticlandmark.com](mailto:hello@agenticlandmark.com) and you will receive a response within thirty days. If you are in the EEA or UK, you also have the right to lodge a complaint with your supervisory authority. ## Scope and changes This site and the practice's services are directed to organizations in the United States, the European Economic Area, and the United Kingdom. This policy will be updated as the site's data practices change, with the date above revised accordingly; material changes to how inquiry or subscriber data is handled will be noted prominently. Questions about this policy: [hello@agenticlandmark.com](mailto:hello@agenticlandmark.com). ### About URL: https://www.agenticlandmark.com/about/ Last updated: 2026-08-19T15:24:59.000Z Agentic Landmark · About # The agent is the fifth platform shift *this practice has worked through.* Agentic Landmark is the advisory practice of Jim Cook, who has led brands through every major platform shift of the digital era, with more than 35 years of VP and SVP-level leadership across 17 industry verticals. The practice exists because the current shift, in which agents join humans as the entities that find, evaluate, and transact with brands, at delegation depths that vary by consumer and by category, rewards the same thing every prior shift rewarded: the brands that rebuilt their infrastructure before the interface finished arriving. 1988 ## Interactive before the web Built interactive multimedia sales tools from R&D data at Bandag (Bridgestone, $900M+ revenue, 100 countries) for fleet clients including FedEx, UPS, and Walmart. Silver CINDY award, 1993, for a touch screen kiosk: an interface most companies did not yet believe in. 1995 ## The web arrives Designed Bandag's first website, then carried the shift through enterprise digital unification at Sprint ($23B, now part of T-Mobile), financial services eCommerce at GE ($130B), and VP-level eCommerce at Insight ($7B). 2010s ## Mobile and the app economy Nine years as SVP at Mission Data, with engagements spanning Kroger, National Geographic, NBC Universal, Penguin Random House, and The Atlantic. Then VP of Digital Customer Experience at Credit One Bank: a mobile banking app from zero to 4 million active monthly users in year one, a 4.7-star App Store rating, JD Power top 3 in mobile banking. 2018 ## Voice, and the first agents Published thought leadership on voice interfaces as emerging CX technology, including a conversation on emerging technologies for marketers on Yeager Marketing's Top 3 for Tech Marketers podcast, and built a Digital Persona and Personalization SaaS platform at ThoughtLava. Voice asked the questions agents now answer at scale: how does a brand get chosen when a human is not looking at a screen? Now ## The agentic web The shift the prior four rehearsed for. Building Agentic Landmark: the Agent Readiness Index, a five-dimension scoring instrument that quantifies a brand's exposure to agentic commerce; its operational twin, the Operational Agent Readiness Index, which scores how safely a brand can delegate real work to agents of its own; a framework library spanning delegation postures, two-audience design, and agentic conversion optimization; and an essay series, The Landmark Brief, arguing the case in public. For the first time in five shifts, the work spans both sides of the interface: the products humans use, and the infrastructure agents read. Every shift produced the same divide: brands that treated the new interface as a *channel*, and brands that understood it as a *reorganization*. In an agentic world, brands need a landmark agents can navigate by. Building those is what this practice does. ## The practice is its own first client Most advisory practices in this space describe the agentic shift from outside it. This one runs on it. The practice's own operating system is agent-native: a version-controlled source of truth where agent workflows draft, reconcile, and monitor; every change proposed as a pull request; and a human holding merge authority over all of it. That is the same boundary the practice advises brands to design for consumers, authority granted to agents in defined scopes, with a human decision at the gate. It is also the discipline the practice's operational instrument, the Operational Agent Readiness Index, scores in others: how much authority an agent holds, how far a wrong action reaches, and whether it can be undone. And the infrastructure it sells is the infrastructure it ships: this site carries the practice's full service catalog as structured data agents can query, and a machine-readable identity record any agent can resolve. When an engagement recommends a delegation architecture, the recommendation comes from lived operation, not analyst observation. The evidence is one view-source away. ## What the practice does Agentic Landmark prepares commerce and brand-driven organizations for agent-mediated commerce on both surfaces: the agents that find, evaluate, and transact with a brand from the outside, and the agents it will run inside its own operations. One query now produces two answers, fed by two different layers, and most brands have built for only one of them. On the outward surface, the Agent Readiness Index quantifies a brand's exposure across five dimensions and tells leadership which one is bleeding revenue first, so the next dollar of infrastructure lands where the exposure actually is. On the inward surface, the Operational Agent Readiness Index scores unmanaged agent risk as exposure net of control, so a brand handing its own agents committing authority knows the blast radius before the deployment, not after the incident. The model is advisory by design: the practice designs the blueprint, the brand's own teams and technology partners build it. Agentic Landmark specifies the work and never touches client systems, which is what makes it a clean complement to the implementation partners who do. ## In public - DataCamp · 4 August 2026 [Assess your agentic commerce readiness](https://www.agenticlandmark.com/assess-your-agentic-commerce-readiness/) A 40-minute walkthrough of the Agent Readiness Index, from what the agentic web is through a scored worked example, presented to roughly 500 attendees, about three quarters of them outside the United States. [See the services](https://www.agenticlandmark.com/services/) [Read the thinking](https://www.agenticlandmark.com/brief/) ### Services URL: https://www.agenticlandmark.com/services/ Last updated: 2026-08-19T15:25:26.000Z Agentic Landmark · Services # Four engagements for the demand side, sequenced the way *readiness actually builds*. One more for the door that faces inward. Each engagement has a defined scope, a defined deliverable, and a natural next step. Entry engagements carry published fees; later-stage engagements are scoped to the work. The sequence is designed so the entry steps are never sunk costs: the Briefing fee recovers forward, and Assessment fees credit toward what follows. The sequence is built to be entered at the Briefing or the Assessment; the right door depends on whether the budget conversation has already happened. 01 ## Board & C-Suite Readiness Briefing $15,000–$28,000 · 2–3 weeks For the champion who sees the urgency and needs to convert board skepticism into budget authority. A fully produced 45–60 minute presentation tailored to your vertical, company size, and digital revenue exposure: market evidence from named sources, a competitive disruption narrative drawn from your category, an ARI risk band estimate with revenue exposure framing, and a phased investment recommendation. Pre-briefing with the champion calibrates board-specific concerns; the practitioner can deliver live. The deck is a board-ready asset you own regardless of what comes next. Fee recoverable against any Assessment or Sprint within 90 days ↓ 02 ## Agentic Readiness Assessment $38,000–$85,000 · 3–5 weeks For the brand with budget authority and a readiness problem to diagnose: where, exactly, is revenue exposed? The full Agent Readiness Index diagnostic across all five dimensions, producing a scored Readiness Scorecard with composite score, dimension scores, and risk band; a Strategic Gap Report with dimension-by-dimension analysis and a prioritized remediation sequence; and a 90-day action roadmap with sequenced infrastructure investment priorities. The Assessment answers which question to ask first, before the next dollar of tooling or platform spend is committed. Assessment clients receive scoping credit on follow-on engagements ↓ 03 ## Agentic Strategy & Roadmap Scoped per engagement · 6–12 weeks For the brand that knows its gaps and needs the comprehensive plan: what gets built, in what order, and why. Agent persona development defining how agents should perceive, represent, and interact with the brand; product data and content architecture strategy for agent-first environments; channel strategy across the agent ecosystems that matter for your category; trust and credentialing strategy; a governance framework for ongoing agentic compliance; and a 12–24 month roadmap with phased priorities and resource requirements. The floor covers a single-vertical, clearly scoped strategy; the ceiling covers multi-vertical, enterprise-scale work. Natural next step: ongoing implementation advisory ↓ 04 ## Implementation Advisory Retainer Scoped per engagement · 6-month minimum For the client executing a completed roadmap who needs senior advisory as the landscape keeps moving. Monthly strategy and implementation reviews; agent performance monitoring and interpretation; emerging platform and technology guidance as the agentic landscape evolves; stakeholder alignment support; and access to the practice's benchmarking data and vertical intelligence. The floor covers focused single-workstream scope; the ceiling covers regulated and multi-system complexity. Scope deepens with each renewal. Ongoing Fees move *forward*, not sideways. The Briefing fee is recoverable against any Assessment or Sprint within 90 days; Assessment fees credit toward follow-on scoping. All fees in USD. Engagements in insurance and regulated healthcare carry a 20–25% premium, reflecting the consent architecture and compliance review these engagements require, and the legal, risk, and technology stakeholders they must carry. The instrument underneath ## The Agent Readiness *Index* Every demand-side engagement runs on the same diagnostic spine: a five-dimension scoring instrument that quantifies a brand's exposure to agentic commerce and locates where revenue is leaking first. Three dimensions measure exposure the category sets and add directly; two measure the controls a brand owns and are scored inversely, so a stronger moat and more mature data lower the composite, and the exposure subtotal is a floor no remediation lowers. The composite runs 0 to 100 and higher means more exposed, placing a brand in one of four risk bands: Low, Emerging, Significant, or Structural. The operational instrument runs the same direction on the same 0 to 100 scale but inverts the weighting: two exposure dimensions at 35 percent, three control dimensions at 65 percent. Read through both doors, Meridian, the practice’s worked example, carries an ARI of **65.90** with **45.65** of it exposure no remediation lowers, and an OARI of **62.05** with **39.00** points of control it has not built. Same company, same band, opposite shapes. A demand-side score is mostly weather and an operational score is mostly roof, which usually makes the operational door the cheaper one to move first. 25% Traffic Dependency How much revenue rides on discovery channels that agents are already intermediating. 20% Structured Comparability How readily decisions in your category reduce to a structured comparison. Agents do not browse; they compare. 20% Agent Substitutability Whether an agent can shortlist and select in your category without a human in the loop. 15% Brand Moat inverse The brand strength that survives algorithmic intermediation, scored inversely: a strong moat reduces exposure. 20% Structured Data Maturity inverse The machine-readability of your product, pricing, and entity data, scored inversely: mature data reduces exposure. Immature data produces omission, not evaluation. [Read the full methodology](https://www.agenticlandmark.com/agent-readiness-index/) The door that faces inward ## The Operational Agent Governance *Sprint* $45,000–$125,000 · 6–8 weeks The four engagements above prepare a brand for the agents outside it. This one is for the agents inside it: for the COO, CIO, or Chief AI Officer about to give operational agents committing authority, agents that can move money, change records, and take actions that do not cleanly undo. A scoped six-to-eight-week sprint scores the exposure before the deployment and sets the controls around it. It produces three deliverables. The **Operational Agent Authorization Map** inventories every workflow where agents act or will act, and sets an execute, approve, or escalate boundary for each. The **Operational Data Readiness Assessment** audits the ERP, WMS, OMS, and supplier data those workflows depend on, against what an agent needs to act safely. The **Bounded Autonomy Governance Framework** defines the escalation triggers, audit trail, rollback architecture, authority inheritance chain, and quarterly review that keep the boundaries enforced. It runs on the operational instrument, the Operational Agent Readiness Index, which scores unmanaged agent risk as exposure net of control across five dimensions. The composite places the program in one of four bands: Low, Emerging, Significant, or Critical Operational Risk. 25% Authorization Maturity inverse Whether agents act on their own scoped credentials or borrow a human's. Scored inversely: mature authorization reduces exposure. 20% Blast Radius and Reversibility How far a wrong action reaches, and whether it can be undone. Recommendations are reversible; committing actions are not. 20% Operational Data Readiness inverse Whether the data an agent acts on is complete and trustworthy enough to act on. Scored inversely. 20% Governance Maturity inverse The review, controls, and observability around agents that take action. Scored inversely. 15% Workflow Suitability Whether the workflow is one an agent should hold at all, or one that needs a human at the gate. [Read the full methodology](https://www.agenticlandmark.com/operational-agent-readiness-index/) [Take the free snapshot](https://www.agenticlandmark.com/operational-agent-readiness-index/) ## Most engagements start with a *conversation*, not a contract. [Start the conversation](https://www.agenticlandmark.com/engage/?src=services) [Read the thinking](https://www.agenticlandmark.com/brief/) ### Agent Readiness Index URL: https://www.agenticlandmark.com/agent-readiness-index/ Last updated: 2026-08-19T15:22:16.000Z Agentic Landmark · Methodology # The Agent Readiness *Index* A five-dimension scoring instrument that quantifies a brand's revenue exposure to agentic commerce and locates where revenue is leaking first. It is one of the practice's two instruments: the Agent Readiness Index reads the outward surface, how agents beyond the business find, evaluate, and choose a brand, while its operational twin, the [Operational Agent Readiness Index](https://www.agenticlandmark.com/operational-agent-readiness-index/), reads the inward one. The composite runs 0 to 100; higher means more exposed. Three of the five dimensions measure exposure your category sets, which no remediation lowers; two measure the controls you own and are scored inverse, so the exposure subtotal is a floor you carry no matter how well you execute. The instrument is published here because a practice arguing for machine-readable infrastructure should not run an opaque one. Exposure · 65% · set by your category · added directly · not lowered by remediation TD · 25% ## Traffic Dependency How reliant the business is on organic browsing, SEO-driven discovery, and affiliate traffic. Heavy organic search dependency means immediate exposure, because agents compress click-through behavior: they read, evaluate, and answer without delivering the visit the channel was built to monetize. SC · 20% ## Structured Comparability Whether an agent can meaningfully compare the brand's offering against competitors in a structured evaluation. High comparability means higher disruption risk: clear feature matrices, standardized pricing tiers, and transparent eligibility criteria are exactly the raw material of the comparison table. Agents do not browse; they compare. AS · 20% ## Agent Substitutability Whether an agent could reasonably shortlist or select in this category without meaningful human involvement. Categories where decisions follow clear rules and the purchase is self-serve automate fastest; categories with heavy emotional differentiation or complex human judgment automate last. Control · 35% · owned by you · scored inverse · every remediation target lives here BM · 15% *inv* ## Brand Moat Strength The brand equity that makes consumers resistant to agent-substituted alternatives. Scored as a control, inversely: a strong moat lowers exposure. The diagnostic question is not whether the brand is well known, but whether the moat is built on genuine product advantage or on behavioral friction. An agent does not have a home screen; it queries every option on every transaction, and friction-based loyalty does not survive that. SDM · 20% *inv* ## Structured Data Maturity How ready the organization's data and infrastructure are for machine evaluation: API exposure, pricing clarity, schema completeness, machine-readable documentation. Scored as a control, inversely: mature data lowers exposure. Low maturity means agents cannot accurately assess the brand even when they try, producing misrepresentation or omission rather than competitive selection. Immature data does not lose the comparison; it never enters it. ARI = (TD × 0.25) + (SC × 0.20) + (AS × 0.20) + ((100 − BM) × 0.15) + ((100 − SDM) × 0.20) All inputs score 0–100\. The two control dimensions are inverse: a Brand Moat score of 70 contributes (100 − 70) × 0.15 = 4.5 points, not 10.5, and a Structured Data Maturity score of 70 contributes (100 − 70) × 0.20 = 6.0 points, not 14.0\. Control that is working subtracts from exposure. The three exposure dimensions add directly. Because the controls are inverse, the exposure subtotal is a *floor*. Think of it as weather and roof. Traffic Dependency, Structured Comparability, and Agent Substitutability are weather: set by the category a brand sells in, carrying 65 percent of the weight, and lowered by no remediation this instrument recommends. Comparability and substitutability are fixed outright. Traffic dependency is the one that drifts, and only as a revenue mix shifts toward direct, branded, and paid channels, which is a change to the shape of the business measured in years rather than a fix. Brand Moat and Structured Data Maturity are the roof: brand-owned, scored inversely, and home to every point of remediation. Meridian, the running brand behind the Arc in the demonstration above, scores **65.90**. Of that, **45.65** is weather, the composite it would still hold with flawless execution on everything it controls. Only **20.25** is in play. A high floor is not a failing. It is the quantified cost of competing in a highly comparable, highly delegatable category, and knowing the number is what separates a scoping decision from a remediation plan. The composite places a brand in one of four bands 0–30 ### Low Structural protections insulate the brand from near-term disruption: a strong moat, low traffic dependency, or low substitutability. Focus: GEO content architecture and schema markup. Monitor and prepare. 31–55 ### Emerging Agentic displacement is beginning to reach adjacent competitors. The window for proactive investment is open but narrowing. Focus: the structured data layer and order-management API readiness. 56–75 ### Significant Category dynamics are already shifting and competitive gaps are forming. Delay compounds the cost of remediation. Focus: protocol integration and retrieval infrastructure. 76–100 ### Structural Revenue impact is measurable or imminent on a 12–24 month horizon. Focus: the full Accelerator program, all tracks simultaneously. The published layer is not the whole instrument. The working Assessment applies a *category exposure overlay* that calibrates the dimensions to how agents actually reach your category, *supplementary gates* for platform-catalog, brand-owned-agent, and white-label exposure, and *vertical sub-dimensions* where the category demands them. The math above tells you what the score means. The Assessment is what produces yours. ## External *validation* Boston Consulting Group's 2026 scenario analysis of agentic retail mapped four possible futures and found two imperatives constant across all of them: discoverability, the ability to be found by agents regardless of which system they operate inside, and desirability, brand strength that survives algorithmic intermediation. Those two imperatives map directly onto the instrument: discoverability is read across both groups, as the exposure Traffic Dependency measures set against the control Structured Data Maturity scores, which is precisely the gap between how much a brand depends on being found and how readable it is when an agent looks; desirability is what the inverse-scored Brand Moat dimension captures. The ARI was built before that analysis published; the alignment is convergence, not citation. ## Walked *end to end* A session presented for DataCamp in August 2026 runs the instrument from first principles: what the agentic web is, how much buying authority consumers are actually delegating, and how five weighted dimensions resolve into a single composite. It closes on the Meridian example above, scored and read in full, and on the four questions worth answering before commissioning anything. The record carries the slides, an eleven-category posture and exposure reference, and links into the arguments behind each part. [See the session record](https://www.agenticlandmark.com/assess-your-agentic-commerce-readiness/) ## The Assessment runs the instrument. *The score is yours to keep.* [See the services](https://www.agenticlandmark.com/services/) [Start the conversation](https://www.agenticlandmark.com/engage/?src=ari) ### Engage URL: https://www.agenticlandmark.com/engage/ Last updated: 2026-08-19T15:25:49.000Z Agentic Landmark · Engage # A conversation, *not a pitch.* Tell the practice what prompted you to reach out. You will hear back within two business days with a point of view on your situation, and if a conversation makes sense, a scheduling link to book it. No sequence enrollment, no drip campaign, no handoff. Name Work email What prompted you to reach out? Vertical optional Select Travel Retail DTC CPG / Consumer Goods Manufacturing Insurance Healthcare Other Your seat optional Select CMO / Marketing Digital / eCommerce Product COO / Operations CIO / Technology Chief AI Officer Other Start the conversation Your name and email are used only to reply to you. The practice keeps qualification context for its pipeline; contact details are never stored there. ## Here for the *thinking*? The Landmark Brief publishes the practice's arguments on the infrastructure, authorization architecture, and commerce strategy of the agentic web. Subscribing puts new briefs in your inbox; your details stay in this site's membership system and nowhere else. Prefer email? [hello@agenticlandmark.com](mailto:hello@agenticlandmark.com) What happens next: your note is read by the practitioner, not a queue. Within two business days you receive a reply with a point of view on your situation, and a scheduling link if a working session is the right next step. *Engagements typically begin with the Briefing or the Assessment; the reply will say which, and why.* ### Operational Agent Readiness Index URL: https://www.agenticlandmark.com/operational-agent-readiness-index/ Last updated: 2026-08-19T15:23:00.000Z Agentic Landmark · Operational instrument # The Operational Agent Readiness *Index* A five-dimension scoring instrument that measures the risk an organization carries when it delegates real work to agents inside its own operations. The composite runs 0 to 100; higher means more exposure net of control. It is the supply-side twin of the Agent Readiness Index: where ARI asks whether outside agents can find and choose you, OARI asks whether inside agents can act on your systems safely. It is published here for the same reason ARI is, because a practice arguing for governed agent infrastructure should not run an opaque instrument. AM · 25% *inv* ## Authorization Maturity Whether the organization can express and enforce what an agent is allowed to do, per workflow, with scoped and revocable authority. Scored as a control, inversely: mature authorization lowers exposure. The strong case is authority drawn per workflow at three boundaries, what the agent may execute on its own, what needs human approval, and what must escalate, with each agent holding its own scoped and revocable credentials and authority changes made at config level and taking effect immediately. The weak case is all-or-nothing authority exercised through borrowed human or service accounts, where tightening scope after a near-miss is a code deploy rather than a setting. This is the dimension where the practice holds proof a framework-reading competitor cannot manufacture. BR · 20% ## Blast Radius and Reversibility How much a single wrong action costs, and whether it can be taken back. Scored as exposure, directly: a large, irreversible blast radius raises the score. The questions are the worst realistic outcome per workflow, where a wrong action is caught (before commit, after commit, or downstream next week), what can actually be rolled back, and whether per-action and cumulative ceilings force a human in before damage compounds. Recommendations that are wrong self-correct at the next query; actions that are wrong have already happened. This is the dimension with no demand-side analogue. DR · 20% *inv* ## Operational Data Readiness Whether the operational data an agent reads and writes is structured, current, and reconciled enough that an action taken on it is an action taken on reality. Scored as a control, inversely. This is the inward cousin of ARI's Structured Data Maturity: the same discipline pointed at ERP, WMS, OMS, and supplier and customer systems rather than at the public catalog. The weak case is the same fact conflicting across systems with no resolvable authority, data that is stale or batch-delayed, contracts that change silently, and workflows that fail open on bad input rather than halting and escalating. GM · 20% *inv* ## Governance Maturity Whether the organization can see, escalate, and stop agent action once it is live. Scored as a control, inversely. The strong case is escalation triggers defined in advance, a complete and tamper-evident audit trail, a named owner per agent who can intervene in real time, a tested agent-failure runbook, and reviews that can pull an agent from production. The weak case is escalation discovered only after failure, partial logs, accountability that is diffuse, a plan that is really just the general IT incident process, and set-and-forget deployment. WS · 15% ## Workflow Suitability Whether the workflows being delegated are genuinely agent-suitable, or whether the organization is pointing agents at judgment-heavy, high-variance work because that is where the enthusiasm is loudest. Scored as exposure, directly, as a mismatch: agents aimed at unsuitable work raise the score. The questions are the share of delegated work that is bounded and rules-expressible versus judgment-dependent, whether the acceptable-error tolerance is explicit and honest, and whether selection was driven by suitability analysis or by which demo landed best with leadership. OARI = ((100 − AM) × 0.25) + (BR × 0.20) + ((100 − DR) × 0.20) + ((100 − GM) × 0.20) + (WS × 0.15) All inputs score 0 to 100\. The three control dimensions are inverse: a mature control lowers the composite. An Authorization Maturity of 80 contributes (100 − 80) × 0.25 = 5 points, not 20, because control that is working subtracts from exposure. The two exposure dimensions add directly. The frame is exposure net of control: what could go wrong, minus how well you have contained it. The same subtotal logic applies here, with the proportions inverted. Blast Radius and Workflow Suitability are the operational weather: exposure set by the work an organization chose to delegate and by what a failure would cost, lowered by delegating differently rather than by governing better. They carry 35 percent. Authorization Maturity, Operational Data Readiness, and Governance Maturity are the roof, and they carry 65 percent. That inversion is the finding. On the demand side most of a score is weather. Here most of it is roof, and a roof is built rather than inherited. Meridian, the same running brand read through the operational door, scores **62.05**: **23.05** is exposure from the work it delegated and **39.00** is control it has not built yet. Nearly two thirds of the score is buildable. The composite places an operation in one of four bands 0–30 ### Low Controls match exposure. Deployments are authorized, bounded, governed, and observable. Focus: monitor, keep governance review light, and re-run the snapshot on each new deployment. 31–55 ### Emerging Exposure is beginning to outrun controls in adjacent workflows. The window to formalize before scaling is open but narrowing. Focus: an authorization map on the workflows in flight, and closing the single lowest control dimension. 56–75 ### Significant Controls are materially behind exposure and a failure is likely on the current trajectory. Delay compounds remediation cost. Focus: the full governance sprint, all three components, sequenced from the lowest control dimension up. 76–100 ### Critical Exposure far exceeds control. Failure is imminent or already occurring. Focus: contain or halt scaling, an emergency bounded autonomy framework, and a sprint paired with a remediation retainer. The five scored dimensions are not the whole instrument. Alongside them, the Assessment records a separate *observability* flag, Ready, Partial, or Blind, that asks the prior question: could you even see an agent misbehaving in production. It is never folded into the composite, because readiness to run agents safely and readiness to see what they are doing are different problems with different infrastructure. An operation can have mature authorization and still be Blind. The math above tells you what the score means; the Assessment, delivered inside the sprint, is what produces yours. ## One competence, *two instruments* OARI is not a second practice. It is the inward reading of the same competence. ARI scores whether an outside agent can accurately evaluate you on the way in; OARI scores whether an inside agent can safely act on your systems once you have handed it the work. The seam between them is structured data, read in two directions: ARI's Structured Data Maturity asks whether an outside agent can read you correctly, and OARI's Operational Data Readiness asks whether an inside agent can act without acting on a stale or conflicting version of the truth. Same discipline, two directions. A brand can be strong on one surface and exposed on the other, which is why the two scores do not collapse into one. The free door-opener ## Take the *snapshot* Eight questions, an indicative band, and the single widest gap between your exposure and your controls. It runs in your browser, and nothing you enter is stored. It does not return a precise score, the instrument computes that, it returns where you sit and what to close first. ### Assess your agentic commerce readiness URL: https://www.agenticlandmark.com/assess-your-agentic-commerce-readiness/ Last updated: 2026-08-19T15:26:17.000Z Agentic Landmark · In public # Assess your agentic commerce *readiness* DataCamp · 4 August 2026 · 40 minutes · roughly 500 attendees, about three quarters outside the United States The session walked the Agent Readiness Index end to end: what the agentic web is, why the shift is further along than most planning assumes, how much buying authority consumers are actually handing over, and how a brand scores its own exposure across five weighted dimensions. It closed on a worked example and a first step that costs nothing but a meeting. This page is the record of what was covered, with links into the longer arguments behind each part. [Watch the recording on DataCamp · registration required](https://www.datacamp.com/resources/webinars/assess-your-agentic-commerce-readiness?ref=agenticlandmark.com) ## What the session covered **The agentic web, and why it is not a channel.** An agent holds a goal, queries infrastructure across many brands, and returns with a decision or a completed purchase. It never renders the page, so design, imagery, and story are invisible to it. Treating this as a new channel next to mobile and marketplace produces optimization of a page the buyer may never see. The longer argument is in [Agentic Commerce Is Not a Subset of E-Commerce](https://www.agenticlandmark.com/brief/agentic-commerce-is-not-a-subset-of-e-commerce/). The category error: a marketplace is a place with your storefront on it, an agent is a layer over every place at once **The demand-side data.** AI traffic to US retail sites grew 393% year over year in Q1 2026\. In March 2026, AI-sourced traffic converted 42% better than non-AI traffic, a reversal from 38% worse a year earlier. Average machine readability of US retail product pages sits at 66%, which is the one number in that set a brand controls. All three from Adobe's Q2 2026 AI Traffic Report. The trust threshold flipped in a single year, and the one number brands control sits at two thirds **Why the funnel does not see it.** Botify's crawl data shows roughly one human visit for every 198 times OpenAI crawls a site, against one visit per six Google crawls. The evaluation happens upstream of any session your analytics were built to record, so the commercial event is over before the funnel begins. [Your Best-Converting Channel Is the One You Cannot Measure](https://www.agenticlandmark.com/brief/your-best-converting-channel-is-the-one-you-cannot-measure/) covers why that revenue arrives disguised, and [You are measuring one funnel and paying for two](https://www.agenticlandmark.com/brief/you-are-measuring-one-funnel-and-paying-for-two/) covers what to build about it. In the agent path there are zero brand-controlled interface touchpoints **Delegation postures.** Curator, Scoper, Delegator, Orchestrator. How much authority the consumer hands over is the variable that reshapes the funnel, and it varies by purchase rather than by person. Accenture's Consumer Pulse 2026, covering 25,590 consumers across 16 countries, puts 74% delegating routine tasks, 32% letting an agent decide within limits, and 9% at full delegation. The gradient, not the headline number, is what shapes strategy: [Agentic adoption is a gradient, not a number](https://www.agenticlandmark.com/brief/agentic-adoption-is-a-gradient-not-a-number/). Down the gradient, delegated authority increases and the population willing to grant it shrinks **The agent-accessibility stack.** Four standards in a fixed order across two layers. robots.txt and llms.txt describe, schema.org structures, and WebMCP exposes callable tools so an agent can act. Being cited and being usable are different problems, and the second only pays off after the first: [Being cited is not being usable](https://www.agenticlandmark.com/brief/being-cited-is-not-being-usable/). The discovery layer is the prerequisite; the infrastructure layer pays off only after it **From atmosphere to attribute.** "Premium build quality" is not retrievable. A stainless steel boiler with PID temperature control is. Every claim living in campaign language has to be re-expressed as structured, verifiable data, or the agent never enters the brand into the decision at all. [What It Means to Be a Machine-Readable Brand](https://www.agenticlandmark.com/brief/what-it-means-to-be-a-machine-readable-brand/). What lives in campaign language has to be re-expressed as structured, verifiable data **Where regulation sits.** The EU's Digital Product Passport, under the Ecodesign for Sustainable Products Regulation (EU 2024/1781), phases in from 2025 to 2030 and will require a structured, verifiable record of a product's identity, materials, origin, compliance, and lifecycle, retrievable from one identifier. Its purpose is sustainability. Its byproduct is close to the record an agent needs to resolve and rank a product, which means that for EU-selling brands the compliance build and the legibility build are largely the same build. Separately, the EU AI Act reached full applicability on 2 August 2026\. The UK has no AI-specific law but enforces actively under the DMCC Act, with the CMA able to fine up to 10% of global turnover directly. The US remains largely deregulated at federal level, with a state-level patchwork forming. Two of the three have no agent-specific law, which is not the same as having no exposure **The index itself.** Five weighted dimensions. Three are exposure set by your category, worth 65%: Traffic Dependency, Structured Comparability, Agent Substitutability. Two are control you own, worth 35% and entering inverse: Brand Moat and Structured Data Maturity. Meridian, the worked example, scored 65.9 out of 100, with 45.65 of that an exposure floor the brand cannot move and only 20.25 in play. Measurement Readiness runs alongside as a separate flag, held outside the composite, because a brand can score well on catalog structure and still be blind to its own numbers. On what the score does not tell you, see [Your Readiness Score Cannot Tell You Where You Stop](https://www.agenticlandmark.com/brief/your-readiness-score-cannot-tell-you-where-you-stop/). Exposure the category sets, control the brand owns, one composite ## The reference from the session Agentic exposure is not uniform. It tracks how much buying authority the shopper hands to the agent, which varies by category. Find the row that fits you, read the posture, and start on the one move that matters most for where you are. Posture 01 Curator The person browses and decides. The agent assists. Discovery is mediated, selection is human. Posture 02 Scoper The person sets criteria, the agent shortlists, the person picks from the set it returns. Posture 03 Delegator The person sets a mandate and the agent selects and buys within it. One approval, or none. Posture 04 Orchestrator Standing authority. The agent reorders and manages on its own, revisiting the person rarely. Eleven commerce categories with their typical delegation posture, agentic exposure level, and highest-leverage first move. | Category | Typical posture | Exposure | The first move | | ----------------------------------------------------------------------------- | ---------------- | -------- | ---------------------------------------------------------------------------------------------------------- | | Consumables and replenishables · grocery, supplements, pet, household staples | Orchestrator | High | Get **reorder and availability data** exact and real-time. Standing orders punish stale stock and pricing. | | Consumer electronics and appliances · spec-driven, comparison-heavy | Delegator | High | Publish **complete, verifiable attributes**. Missing a spec drops you from the ranking entirely. | | Footwear and apparel · by fit, spec, and use case | Delegator | High | Structure **fit, material, and use-case data** so an agent can match intent, not just keywords. | | Beauty and personal care · routine, repeat purchase | Delegator | High | Make **reviews and ingredient data machine-readable**. Repeat-buy categories reward trust signals. | | Travel and hospitality · flights, rooms, packages | Scoper Delegator | High | Wire **real-time inventory and price into agents**. Google has named hotels the next agentic vertical. | | Commodity B2B supplies · maintenance, office, industrial | Delegator | High | Make your **catalog, terms, and credit callable** by a procurement agent, not just a rep. | | Financial and insurance products · rate and terms driven | Scoper | Emerging | Structure **terms for comparison**. Trust and regulation keep a human on the decision, for now. | | Automotive and high-consideration · vehicles, large, infrequent purchases | Scoper | Emerging | Make **spec, trim, and price data legible** so you make the shortlist the human then works through. | | Complex, relationship-led B2B · negotiated, long sales cycle | Curator Scoper | Emerging | Fix **discovery legibility first**. Be in the set an agent assembles before the human relationship starts. | | Experiential and services · dining, events, personal services | Curator | Lower | Structure **availability and reviews**. Discovery is agent-mediated even when the choice stays human. | | Luxury, craft, and bespoke · story-led, configured, brand-driven | Curator | Lower | Make your **advantages legible as data**. An agent cannot narrate the craft it cannot read. | **How to read this.** Exposure is set by your category and moves slowly. It is the weather. Your posture tells you how soon it bites and how much of the purchase an agent already runs. The first move is the highest-leverage step you control, and in almost every row it comes back to one thing: **product data an agent can find, read, and trust.** The line also drifts. Categories move toward delegation as trust builds, the way considered electronics did, so today's Curator is tomorrow's Scoper. Postures and exposure from the Agent Readiness Index · illustrative, calibrate to your own catalog and channel mix [Category and posture reference](https://www.agenticlandmark.com/assets/files/agentic-landmark-category-and-posture.pdf) [Full deck, 35 slides](https://www.agenticlandmark.com/content/files/2026/08/agentic-commerce-webinar-deck.pdf) ## The first step from the session Not a budget, and not a vendor. Get the five owners in one room: digital and ecommerce, marketing and brand, data and analytics, engineering or platform, and merchandising. Answer four questions together. Are we in an exposed category. Where does an agent lose us today. Can we see agent traffic in our numbers at all. What is the widest gap we control. That conversation is the assessment, whether you run it yourselves or bring someone in. [See the Agent Readiness Index](https://www.agenticlandmark.com/agent-readiness-index/) [Start a conversation](https://www.agenticlandmark.com/engage/) ## Posts ### 500,000 SKUs Is Not an SEO Problem URL: https://www.agenticlandmark.com/brief/500-000-skus-is-not-an-seo-problem/ Last updated: 2026-09-04T01:28:30.000Z *Google told brands that AI visibility is a content discipline, then published a case study describing an infrastructure one.* ## The claim, and who is making it In August, Think with Google published a panel from Cannes featuring Linda Ha, deputy CMO of Ikea, Dan Finley, group CEO of Debenhams Group, and Andy Wells, VP of growth marketing at DoorDash. One section runs under the header "Good SEO is good GEO." The framing throughout is that foundational search engine optimization remains the best formula for organic discoverability as search becomes conversational and agentic. The piece closes by instructing readers to audit their search architecture for AI readiness and to move more budget into AI-powered campaigns. The advice is not wrong. It is incomplete, and the incompleteness runs in the direction of the publisher's interest. This is Google's publication, Google's event, Google's stage, and Google's ad products. That does not make the claim false. It makes it a claim that requires an independent test, and the article supplies none. Its own footer notes that results vary by advertiser. The test is available anyway. It is sitting in the article. ## Read the example Wells states the thesis most directly of the three. He says the notion that answer engine optimization must be an entirely separate practice from SEO is a misconception, and that DoorDash has stayed focused on site architecture, structured data, and making sure crawlers can read the site. Then he describes what DoorDash actually built. The platform carries more than 500,000 product SKUs. Availability varies by local merchant: a given item is stocked at some and not at others. DoorDash integrates directly with Google so that Gemini and AI Overviews can read that catalog and understand when to recommend DoorDash, or a specific product, based on what is actually available. That is not SEO. Sitemaps, canonical tags, heading hierarchy, internal linking, and crawl budget do not solve merchant-level availability across half a million items. What solves it is a feed. An identity scheme that keeps a product the same product across merchants. An availability signal that is true at query time rather than at publish time. A direct integration path to the platform doing the reading, with a contract about freshness. Those are supply-side data commitments owned in engineering and operations. They are not a content team's backlog, and no amount of technical SEO maturity produces them. The panel's own worked example refutes the panel's own headline. You do not have to argue with Google here. You have to read closely and point. ## The two layers, named The confusion is real and it is worth naming precisely, because both layers are genuinely required and they are genuinely different. Layer one is whether an agent can find you and cite you. That layer is substantially continuous with SEO. Crawlability, clean markup, server-rendered content, schema, and third-party corroboration all carry over. Google is right that the fundamentals did not stop mattering, and the practitioners who declared SEO dead in 2025 were wrong. Layer two is whether an agent can resolve, compare, and act on your catalog with enough confidence to name a specific item to a specific person in a specific place. That requires structured product data that is machine-comparable against competitors on the attributes the agent is reasoning over, and that stays true as inventory moves. In ARI this is Structured Comparability and Structured Data Maturity, and they are scored separately from anything a content audit touches. Ha makes the same point from the other side without naming it. Ikea, she says, has to stay on top of what other sources say about the brand, and works across paid, owned, and earned to shape that. Third-party representation is not a property of your website either. It is not something a sitemap fixes. Good SEO gets you into the answer. Structured, resolvable, availability-true product data gets you named inside it. The first is necessary. It is not sufficient, and treating it as sufficient is how a brand spends a year on content operations and cannot explain why agents still recommend a competitor. ## The admission in the third section The most valuable sentence in the piece is not about search at all. Wells says DoorDash can no longer rely on click-based attribution as discovery behavior shifts, and has moved to incrementality experiments, measuring incremental gross order value rather than attributed conversions. That is Measurement Readiness, conceded by an operator rather than asserted by an analyst. Measurement Readiness sits outside ARI's five scored dimensions on purpose. It does not measure agent exposure. It measures whether a brand can observe its own exposure in its own numbers. A brand that fails it cannot tell whether its AI search work did anything, or whether the quarter moved for unrelated reasons, and will keep funding whatever its instrumentation happens to be able to see. That is not a reporting inconvenience. It is a capital allocation defect. Notice the sequencing. The same operator who says AEO is not a separate discipline also says his attribution cannot see the behavior he is optimizing for. Both statements are true. Together they describe a company doing serious infrastructure work it does not yet have a name for, measured by an apparatus it has already outgrown. Most brands are in the same position without the first half. ## The door nobody opened DoorDash also shipped Ask DoorDash, a conversational assistant that will take a recipe link or a photo of a cookbook page and assemble a full shopping cart. On the demand side that reads as a feature. On the supply side it is an operational agent assembling a commitment against a catalog where availability varies by merchant. The failure mode is a cart of items that are not there, produced at scale, in front of a customer, by the brand's own system. That is blast radius in the OARI sense, running on exactly the same catalog data the Gemini integration depends on. One dataset. Two exposures. Only one of them was discussed at Cannes. ## What the audit should actually cover Google's closing instruction is the right instruction. The question is scope. An audit that stops at crawlability will pass a brand that is structurally invisible to an agent making a recommendation. Five things belong in it that a technical SEO audit will not reach: Whether product attributes are comparable against competitors on the dimensions an agent reasons over, not merely present on the page. Whether product identity resolves consistently across every surface where the catalog appears. Whether availability and price are true at the moment of the query rather than at the moment of publication. Whether the brand has an integration path to the platforms doing the reading, or is relying on being crawled. And whether anything in the analytics stack can detect agent-mediated demand well enough to tell you if the other four worked. The last one gates the other four. A brand that cannot measure the exposure cannot govern the response to it, and will conclude that the work did not pay off because the instrumentation could not see that it did. Audit the architecture. Just make sure the architecture you audit is the one agents actually read. [Lessons from marketing leaders on AI Search - Think with GoogleGoogle’s Official Digital Marketing Publication. Discover how marketing leaders at IKEA, DoorDash, and Debenhams are adapting to AI Search.![](https://storage.ghost.io/c/e1/c4/e1c4b063-1fa4-4977-8ffa-6add50cd2316/content/images/icon/google-favicon-48-b5fd5a41-0bb6-46bf-a4d6-10eabf73c27e.png)Google BusinessThink with Google Editorial Team![](https://storage.ghost.io/c/e1/c4/e1c4b063-1fa4-4977-8ffa-6add50cd2316/content/images/thumbnail/twg-us-2853-cannes-search-session-thumbnail-1600x900-79bbe765-d443-4dd5-b1ae-0b13f907d8c8.webp=n-w1244-h700-fcrop64=1,00000000ffffffff-rw)](https://business.google.com/us/think/search-and-video/ai-search-visibility-thought-leadership/?ref=agenticlandmark.com) ### Citation Share Is Not a Metric Yet URL: https://www.agenticlandmark.com/brief/citation-share-is-not-a-metric-yet/ Last updated: 2026-09-02T01:04:35.000Z *An industry is being told to build a capability around optimizing a number that moves several-fold depending on which tool produced it.* McKinsey published a finding in June, in [State of the Consumer 2026](https://www.mckinsey.com/industries/consumer-packaged-goods/our-insights/state-of-consumer?ref=agenticlandmark.com), that should worry anyone selling AI visibility. In consumer goods, brand-owned websites account for one to two percent of the sources large language models cite. Not one to two percent of traffic. One to two percent of the material a model is reading when it forms an answer about your brand. The prescription follows immediately, and it is the one everybody is giving right now. Generative engine optimization is the new search engine optimization. If the models are not reading your site, go influence what they are reading. Strengthen public relations. Seed the review sites. Show up in the communities. Get consistent across platforms you do not own. The finding is real, and the direction is right. The prescription is where I would stop. ## The measurement is better than most and still moves Start with what makes this study worth taking seriously, because most numbers in this category do not survive the first question. The June figure rests on roughly 2.6 million citations, drawn across 25 consumer packaged goods brands, spanning 89,119 unique domains in the UK and US, over October 2025 through May 2026, aggregated across ChatGPT, Gemini, and Perplexity. That is a real sample, and McKinsey published the parameters in the exhibit rather than burying them. Most of what gets quoted at conferences this quarter has none of that. Now look at what the exhibit actually shows. It is broken out by platform, and the one to two percent is a range across those three models. Look at the top ten cited sources and the same brand-owned share runs from three percent to ten percent, again depending on which model you ask. Same vendor. Same window. Same 25 brands. Same 2.6 million citations. A threefold spread produced by nothing except platform mix. If your monitoring vendor weights Perplexity more heavily than mine does, our numbers diverge threefold before either of us has made a mistake. ## Two numbers, eight months apart That is the variance inside one study. Between studies it is worse. In October 2025, in [New front door to the internet](https://www.mckinsey.com/capabilities/growth-marketing-and-sales/our-insights/new-front-door-to-the-internet-winning-in-the-age-of-ai-search?ref=agenticlandmark.com), McKinsey published that a brand's own sites comprise five to ten percent of the sources AI-powered search references, from Google AI Overview data combined with the firm's own analysis. In June 2026, the same firm published one to two percent, from a vendor dataset. Same organization. Eight months apart. A five-fold difference in what reads, to anyone skimming, like the same quantity. No reconciliation offered, and no reference in the second piece to the first. I do not think either is an error. The two figures answer different questions that happen to sound identical. One counts sources referenced by AI-powered search, anchored to Google's AI Overview surface. The other counts citations across three named models. One spans categories. The other is scoped to consumer goods. Change any one of those and the number moves. Change all of them and a five-fold gap is unremarkable. Ask a simpler question and the same problem appears. What counts as a citation? A named link in a footnote is one. Is an unlinked brand mention in the body of an answer? A paraphrase of your product page that names no source? Your own copy appearing because a retailer syndicated it? Every monitoring product has an answer; none of them publish it, and the answers are not the same. ## Who supplied the measurement One more thing worth stating plainly, because I have not seen anyone else say it. The vendor behind the June figure sells generative engine optimization monitoring. The same statistic is the lead hook on its own homepage. The party supplying the measurement also sells the remedy the measurement implies. That does not make the data false. Vendor telemetry is frequently the only telemetry that exists in a new category, and a firm that publishes with it and names it is being more transparent than one that does not. Adobe sits in the same position with its machine-readability benchmarks, which arrive alongside a product Adobe sells for improving machine readability. But the finding is not independent of the conclusion it supports, and that is worth knowing before a number becomes a budget line. Anyone quoting the one to two percent should be naming the vendor in the same breath. I now do. ## What survives the disagreement Strip out everything unstable and something solid is left standing. Both readings, five to ten percent and one to two percent, say the same thing about direction: the overwhelming majority of what models read about your brand is not written by you. Whether it is ninety percent or ninety-nine percent does not change a single decision a brand would make. That finding is robust across instruments, and it is genuinely important. What does not survive is the use the finding is being put to. You cannot optimize toward a target you cannot measure twice. If your agency reports your citation share at four percent and a competitor's vendor reports theirs at eleven, you have learned nothing about either brand. You have learned that two instruments disagree. A number that moves threefold on platform mix alone, and five-fold across published studies, can support a trend line inside a single frozen instrument. It cannot support a level, a benchmark, or a comparison. That is a constraint, not a reason to stop measuring. Fix your instrument, name it every time you report a figure it produced, and read your own movement over time. What you give up is the industry benchmark, and the industry benchmark does not currently exist, no matter who is selling you one. ## The layer you can actually measure Here is the part that made me want to write this down. The prescription on offer is to influence the corpus. Spend on public relations, on reviews, on community presence, on consistency across platforms you do not own. It is not bad advice. It is also a permanent operating expense, applied to a surface you cannot audit, verified by a metric that will not hold still. Compare that to the other layer. Whether your product feed validates is a fact. Whether your structured data parses is a fact. Whether your pricing and availability are reachable by an agent that does not execute JavaScript is a fact. Whether your checkout can complete a transaction initiated by something that is not a browser is a fact. Every one of those is verifiable by you, today, with no vendor in the middle and no methodology dispute to resolve. The discovery layer and the infrastructure layer have never been equally measurable, and I do not think the industry has noticed. Nearly all the attention and nearly all the spend is going to the layer where the numbers are contested, produced by interested parties, and unstable across instruments. Meanwhile, the layer that determines whether an agent can transact with you at all returns a clean pass or fail on every check. Citation gets you into the answer. Infrastructure gets you transacted with. Both matter, and the second one is the only one you can currently prove you have done. Structure survives a corpus you do not control. Seeded content does not. ### A Watermark Is Not An Audit Trail URL: https://www.agenticlandmark.com/brief/a-watermark-is-not-an-audit-trail/ Last updated: 2026-08-14T14:44:30.000Z *Machine-readable marking records that a model touched the words. It cannot record whether judgment shaped what they were used for.* ## What Anthropic actually shipped Anthropic has signed the EU AI Act's Article 50(2) Code of Practice on Transparency of AI-Generated Content, and has published how it intends to honor it. Two mechanisms, deliberately different. Text gets an imperceptible watermark woven into the text itself at the model level, which means it is present regardless of which Claude surface produced it, and it travels with the text through copy and paste. Files get signed provenance metadata under the Coalition for Content Provenance and Authenticity (C2PA) open standard, on supported types including .svg, .png, and .jpg. A signed label signals that a file was processed by Claude and lets you detect whether the file has been tampered with. Models launched on or after August 2, 2026 are marked at launch. Earlier models are in progress under the law's transition period. Marking applies wherever Claude is offered, worldwide, across Claude Platform (API), Claude, Claude Code, Claude Cowork, and Claude Tag, and follows the model through AWS, Google Cloud, and Microsoft Foundry. Detection for users and third parties is committed, with technical documentation forthcoming. This is competent compliance work, applied globally rather than fenced to the jurisdiction that required it. It is also, read carefully, an argument against the use most people are going to make of it. ## The most useful paragraph is in the limitations Anthropic's own documentation is candid in a way that downstream coverage will not be. A detected mark indicates that content may have been processed by Claude. It does not confirm provenance. The documentation names the case directly: people use Claude to proofread, translate, summarize, and convert files, and the output can carry a mark even when the underlying ideas, text, or data originated somewhere else entirely. The failure runs the other way too. Absence of a mark proves nothing. Content may come from a model released before marking was supported, or have been heavily edited, paraphrased, or translated, or be too short a passage to carry a reliable signal, or be a file whose metadata was stripped by a format conversion, a re-save, or a screenshot. Presence is weak evidence of authorship. Absence is no evidence of anything. A signal that fails in both directions is not proof. It is an input. That distinction will not survive contact with the people who need it most. Schools, hiring managers, procurement teams, and platform trust systems will receive a binary from a probabilistic instrument, and the paragraph explaining why it is not a binary will not travel with the mark. ## Two records, and only one of them is being built for you Provenance and authorization are different records answering different questions. A provenance record is retrospective and artifact-bound. It attaches to the output and reports what passed through what. It is authorship-agnostic by construction, because a model that reformats a document leaves the same trace as a model that wrote it. An authorization record is prospective and decision-bound. It attaches to the act, not the artifact. It names the principal who permitted a specific action, the boundary the action was permitted inside, the moment the permission was granted, and the party who stays answerable when the action turns out to be wrong. The practice has been arguing this shape for a while on the commerce side. AP2, Google's agent payments protocol, produces cryptographically signed mandates precisely because a receipt is not an authorization. The grant screen matters for the same reason: it is the surface where a consumer hands scoped authority to an agent, and it is the moment the record gets created rather than reconstructed afterward from whatever traces happen to survive. Regulation is now producing the first record automatically, at model level, across the industry, at no cost to anyone. It will not produce the second. Nobody ships you the second. It is an internal build, and the arrival of free provenance is going to make a lot of organizations feel like they already have it. ## The deployer obligation nobody is reading One line in Anthropic's article does more commercial work than the watermark does: if you deploy Claude in your own product, you should independently assess what Article 50 requires of your products and services. Article 50 obligations fall on providers and on deployers, and those are not the same party. A brand running an assistant on its own storefront, generating product copy at catalog scale, or answering service inquiries with a model is a deployer. The model provider's mark does not discharge the deployer's duty. It means the brand's output now carries a signal the brand did not choose, cannot remove, and has not accounted for in its own disclosures. Marking is also becoming a distribution input rather than a compliance artifact. Platforms already read C2PA manifests, and the standard is the one Adobe and the camera manufacturers converged on. The moment a manifest feeds a ranking, labeling, or filtering decision, provenance stops being a legal question and becomes a Representation Layer question. Brands that cannot state how their own assets were produced will have it stated for them, by a platform, using a default they did not set. ## What a COO should ask instead The operational version of this mistake is asking whether AI can be detected in the work product. That is the detectable question, not the answerable one. The Operational Agent Readiness Index question is different: for any committing action an agent takes inside the business, can you name the authorization that permitted it, the boundary it operated within, and the human accountable for the outcome? Governance Maturity does not ask what touched the record. It asks who decided, and whether the decision sat inside the fence. An audit trail records decisions. A watermark records processing. An organization that can produce the first has little need for the second. An organization that can only produce the second has not built governance. It has been handed labeling and mistaken the two for each other. ## The question underneath A watermark can tell you a machine touched the words. It cannot tell you whether a person did the thinking. The organizations that come through this well will not be the ones that got good at detecting AI in their own supply chain. Detection is being commoditized for them, and it was never going to answer the question anyway. They will be the ones that can point at any output, any decision, any committed action, and name the judgment behind it. That record does not arrive in a model update. You build it, or you do not have it. ### Your Readiness Score Cannot Tell You Where You Stop URL: https://www.agenticlandmark.com/brief/your-readiness-score-cannot-tell-you-where-you-stop/ Last updated: 2026-08-11T03:07:21.000Z *Both instruments produce a weighted composite and a composite average, which is the one thing an ordered dependency refuses to do.* --- Score a company on operational agent governance, and you can get a strong result. Policy is written. Escalation triggers are predefined. Every deployed agent has a named human with authority to pull it. On the instrument that measures those things, this company is in good shape, and it should be, because that work is expensive and most companies have not done it. Now ask whether an agent shopping on behalf of a consumer can read this company's product data accurately enough to put it in a comparison set. It cannot, because the catalog was written for humans. The governance work was not wasted. It is also, for agent-mediated buying, never exercised. And here is the part that should bother you: both of that company's scores are defensible. Neither one is capable of telling it what just happened. ## The order runs on both sides Two briefs in this series have already laid most of the groundwork, and it is worth being precise about which part each one built. [Being cited is not being usable](https://www.agenticlandmark.com/brief/being-cited-is-not-being-usable/) argued that the agent-accessibility stack is a sequence rather than a checklist, and that the order is fixed: if you are not in the answer, the execution layer you built is never invoked. [When you deploy the agent, you own the mistake](https://www.agenticlandmark.com/brief/when-you-deploy-the-agent-you-own-the-mistake/) argued that a brand is exposed to agents from two directions, that the two exposures have different physics, and that one instrument stretched across both goes vague about both. Put those together and something is missing between them. The first brief established an order, and every layer in it faces outward. The second established a second surface, and described it as a surface: parallel, differently governed, differently owned. It never asked whether that surface has an order of its own. It does, and most companies have it drawn backwards. Governance is not the mature state you arrive at after the technical work. It is the substrate. Policy comes first, because it determines what any system is permitted to do at all. Systems of record come next, because they determine what an agent can act on and whether the copy it reached was the authoritative one. Authorization comes last, not because it matters least but because it is the gate immediately before the commit: the per-workflow decision about what this agent may execute, what it must send for approval, and what it must escalate. Run in that order, the operational side is not a surface sitting alongside the demand side. It is a second-order dependency, and it points somewhere specific. ## Both orders end in the same place The demand order terminates at the transaction. So does the operational one. That is the whole structural claim, and it has a consequence neither prior brief could reach. The transaction is the only point in either sequence that depends on both. Everything upstream on the outward-facing side has to hold for an agent to arrive and select you. Everything upstream on the inward-facing side has to hold for your own systems to complete what it started. The two sequences do not run in parallel forever. They converge. Which means the two exposures are not merely different in physics. They are joined, at exactly one point, and it is the point that makes money. A company that has built nothing operationally can be found, compared, and chosen, right up until something has to be committed on its side. A company that has built nothing on the representation side can hold immaculate authorization boundaries that no agent ever reaches. Neither company is partially ready. Each is stopped, at a specific link, for a specific and entirely different reason. They present identically. Both look like a transaction that did not happen, which is the least diagnostic symptom available, and a company reading one instrument will find a plausible story on the side it can see. ## A composite averages; an order does not Now the tension that has been sitting under both instruments since the second one was built. The Agent Readiness Index and the Operational Agent Readiness Index both produce a weighted composite. Five dimensions, weights, a number, a band. That is the correct shape for what they measure, and averaging is not a flaw in them. It is what lets a number describe a whole position at once and lets a board compare this quarter to last. But an average is a claim about the aggregate, and a break is a claim about one link. Those do not commute. Strength downstream cannot compensate for a break upstream, because downstream is never reached, and any instrument that sums across dimensions will quietly let a strong result in four places offset a stopping condition in the fifth. The composite is not lying. It is answering a different question than the one the company asked. The company at the top of this brief does not have a mid-range readiness position that its governance work partially offsets. It has a stopping condition, on one side, at one link, and a fully built lane on the other side that nothing will arrive on until the first is fixed. ## What each one is actually for The resolution is not to abandon the score. It is to stop asking it a question it was never built to answer. The ordered read locates the break. It is a per-link diagnostic with a near-binary answer: does the agent get through here? It tells you what to fix first and, more valuably, what not to fix yet. Almost nobody runs it, because it produces no number and nothing to put on a slide. The score measures exposure. Once the sequence carries traffic, it tells you how much revenue is sitting on the far side of a dependency you do not control, which is the question a board asks, and the question a per-link diagnostic cannot answer. Most companies are running the second exercise without having run the first. They are improving dimensions on a sequence that terminates in a broken link several steps upstream. The improvement is real, it is measurable, it moves the composite, and none of it reaches the transaction. Find the break. Then measure what it costs you. ### No Seller Can Build a Buyer's Agent URL: https://www.agenticlandmark.com/brief/no-seller-can-build-a-buyers-agent/ Last updated: 2026-08-04T02:05:35.000Z *The most useful agents will be the ones brands have the least commercial leverage over, and that is a structural fact about who pays rather than a policy anyone chose.* Earlier in this series I argued that the merchant of record and the merchant of decision have come apart, and that the merchant of decision is whoever built the agent, set its evaluation criteria, and owns the surface where the buyer's intent gets formed. That piece stopped where the unbundling was established. The question it leaves open is the one a CMO asks next, and it is a practical one. If someone else is now making the decision, how do I get into it? Who do I call, and what does it cost? For the agents that will matter most, there is nobody to call. Not because the market is immature and a rate card is coming, but because of what an agent has to be in order to be worth delegating to in the first place. ## The divergence is at one specific moment An agent and a seller want the same thing for most of a transaction. Both want the need understood correctly. Both want accurate product data, real inventory, a checkout that completes. Through all of that, the interests are aligned, which is why so much agentic commerce infrastructure work is uncontroversial and why every party is happy to build it. There is exactly one moment where the interests come apart, and it is the moment that decides the sale. When the agent narrows to a recommendation, the buyer wants the best fit. A seller wants its own product. If the seller is paying for the agent, that preference has somewhere to enter. Everything about the structure of the market right now is a consequence of that single divergence point. ## What a buyer's agent has to be able to do The test is not what an agent says about its independence. It is what it is permitted to do at the divergence point. A buyer's agent has to be able to recommend a competitor when the competitor fits better. It has to be able to say that nothing in the category is worth the money right now, or that the cheaper option is sufficient, or that the thing the buyer already owns does not need replacing. Each of those is a recommendation that costs someone a sale, and a buyer's agent has to be able to make it without anything in its economics arguing back. That is a demanding standard, and it is why the question of who funds the agent is not a detail. An agent funded by the buyer can meet it. An agent funded by placement cannot, whatever its intentions, because the constraint is not honesty but incentive. ## Buyers can detect this without analyzing it The counter is that consumers never scrutinized search results this closely, and search monetized through paid placement perfectly well for two decades. The difference is what is being delegated. A search result is a list to evaluate; the human still chooses, and a sponsored listing among ten is visible and discountable. A recommendation is the choice, and increasingly the whole consideration set. When one option is returned instead of ten, sponsorship is not a labelled row anymore. It is the answer. Buyers do not need to reason this through to respond to it. What they experience is whether the recommendations turn out well over time, and delegation extends or retracts accordingly. An agent that quietly optimizes for margin produces a worse hit rate, and the buyer notices the outcome without ever diagnosing the cause. The mechanism is invisible. The consequence is not. ## The pressure runs the other way, hard None of this means buyer-aligned agents automatically win. The economics point the other direction, and it is worth being clear-eyed about that. Charging buyers directly is the cleanest alignment and the hardest sell, because consumers have thirty years of training that this category is free. Commission on completed transactions looks neutral and is not quite, since it rewards the agent for a purchase happening at all and, depending on structure, for a more expensive one. Seller-funded placement is the most proven monetization model on the internet and the one every incumbent already knows how to operate. So the pressure toward placement is enormous, and it will not be resisted everywhere. What is likely is a market that separates: agents that took placement money and are consequently used for lower-stakes, lower-consideration purchases, and agents that did not and are trusted with the decisions that matter. Both survive. They are not competing for the same delegation. ## What this means if you are the brand The strategic consequence inverts a relationship most marketing organizations are built around. In every prior channel, sufficient budget bought presence. Search, social, retail media, all of it had a rate card. In the agents that earn the highest-value delegation, there is no rate card by design, and building one would destroy the thing that made them valuable. That leaves an uncomfortable but clarifying position. **Your commercial leverage is lowest exactly where the delegation is highest.** The agents you can pay are the ones handling purchases the buyer did not care much about. The agents handling the decisions you most want to influence are the ones where money does not enter. The practical implication is that budget stops being the lever and evidence becomes it. Whether your claims are specific enough to be checked, whether your data says the same thing everywhere, whether the thing arrives as described often enough to show up in the aggregate. This series has covered the mechanics of both, and the point here is narrower: those are not merely the best available tactics while the channel matures. In the highest-value part of the channel, they are the only inputs, permanently, because the alternative input has been designed out. ## The question to ask about any agent There is a single diagnostic that cuts through most of the positioning, and it is worth asking of every agent surface a brand is evaluating. Who pays, and what happens at the moment of recommendation? If the buyer pays, the agent can afford to tell them not to buy. If a seller pays, something has to be true about the recommendation for the business to work, and whatever that something is, it is operating whether or not anyone intended it to. Brands should ask this about the agents they are trying to appear in, because it tells them what kind of appearing is even possible. And it is worth noticing that a company can be genuinely committed to buyer alignment and still be structurally unable to sustain it, if the money it took requires otherwise. Intent is not the variable. The funding is. ### You are measuring one funnel and paying for two URL: https://www.agenticlandmark.com/brief/you-are-measuring-one-funnel-and-paying-for-two/ Last updated: 2026-07-31T02:38:47.000Z *Agent selection is a funnel with its own stages, its own failure modes, and its own success event, and almost nobody instruments it.* Earlier in this series I argued that the best-converting channel most brands have is the one their analytics are trained to ignore, and named what to build: server-side classification, an attribution model that credits agent-influenced discovery separately, and a price-at-time-of-selection baseline for catalogs where price moves. Most teams read that list, recognized a data engineering project, and did nothing. This is the version you can start without one. ## The funnel nobody instruments The reason credit goes missing is structural, and it is not really an analytics failure. There are two funnels running, and every analytics package in common use was designed around the second one. **Upstream is agent selection.** A query arrives somewhere you do not control. The agent retrieves candidates, narrows to a shortlist, and recommends. The commercial event, the moment you are compared against three competitors and either included or dropped, happens entirely inside that sequence. **Downstream is human conversion.** Visit, cart, purchase. This is the funnel with the dashboards. The two are not stages of one process. They have different actors, different failure modes, and different success events. Downstream success is a sale. **Upstream success is being included at all**, and no conversion metric measures inclusion. ## Winning one and losing the other looks identical in your numbers This is why blending them is expensive rather than merely imprecise. A brand can be strong upstream and weak downstream. Agents recommend it constantly, and the buyers arrive and do not convert, because price moved, or stock was thin, or checkout could not be completed programmatically. The dashboard shows soft conversion and suggests a landing page test. A brand can be weak upstream and strong downstream. It converts beautifully on the shrinking population of buyers who still arrive without an agent, and is quietly absent from the consideration sets forming elsewhere. The dashboard shows healthy conversion right up until volume falls off a cliff, and by then the upstream position took a year to erode and will take longer to rebuild. Both cases produce an unremarkable blended number. Neither is diagnosable from it. The single funnel view does not just under-report the agent channel, it produces confident recommendations aimed at the wrong half of the problem. ## Three layers you can build before the engineering Server-side classification is the right foundation and it is a quarter of work with a dependency on your platform team. These three layers do not wait on it, and together they are enough to defend a budget. **Visibility tracking, as a leading indicator.** Whether agents surface you at all, on which prompts, against which competitors, and how that moves. This is upstream instrumentation, the only kind available today, and it is not revenue. That is the point. It moves before revenue does, which makes it the earliest signal you can act on and the one most worth standing up first. **Self-reported attribution, as directional truth.** Add an AI assistant option to the how-did-you-hear-about-us field on lead and order forms. It undercounts, respondents misremember, and it is the fastest honest signal most brands can have running inside a week. A number that is directionally right and admits its error bars beats a precise number that is silently wrong. **Correlation analysis, as the argument.** Track visibility movement against pipeline and revenue over several quarters. Correlation is not causation, and a consistent directional relationship sustained across quarters is evidence, which is strictly more than the channel has now. Note what each layer is for. The first tells you something changed. The second tells you roughly how much. The third is what you say out loud in the room where budget gets allocated. They are not three attempts at the same measurement, they are three different jobs, and a brand that builds only the first has a signal it cannot spend. ## What good enough actually means Good enough is not a weaker version of deterministic attribution. It is a different standard, and it is the one the situation supports. Deterministic attribution answers which touch caused this sale. No layer above answers that, and no protocol available today answers it either. What the layers answer is whether the agent channel is growing, roughly what it is worth, and whether the work you are doing on it is landing. That is sufficient to fund a channel, staff it, and set a target against it. It is worth being explicit about the failure mode on the other side. A brand that insists on precision before it will act does not get precision. It gets another four quarters of the channel reading as unproven in numbers it already knows are wrong, which is the loop the earlier piece described, running with the brand's informed consent rather than in spite of it. ## The part that does not wait There is one asymmetry worth sitting with. When a measurement standard for agent-influenced discovery does arrive, and something will eventually, it will apply forward. It will not reach back and reconstruct which of last year's direct-channel revenue an agent actually sourced. Standards do not come with history attached. So the brands that stood up crude layers now will have two years of movement to read the standard against, and a baseline that makes the first proper numbers interpretable. The brands that waited will have the standard, an accurate reading of one quarter, and nothing to compare it to. The detection work, the moving-baseline problem, and why Measurement Readiness sits outside the ARI composite rather than inside it are all in the earlier piece. This one is about the funnel that piece did not name. You are paying for both. Instrument both. ### A Reporting Line Is Not An Authorization Boundary URL: https://www.agenticlandmark.com/brief/a-reporting-line-is-not-an-authorization-boundary/ Last updated: 2026-07-31T02:05:05.000Z *An org chart records who answers to whom, which is not the same question as what may be committed, and agent software is now shipping the former as though it were the second.* ## The chart came back Earlier in this series I worked through the argument that corporate hierarchy was always an information routing protocol, built around the limit on how many people one person can coordinate, and that AI removes the constraint that made it necessary. I still think that reading is right about the history. What I did not expect was how fast the chart would come back, pointed at machines. A category of product now ships the org chart as the interface for running agents. You define roles. Roles carry standing instructions. Roles sit in reporting lines. A manager agent decomposes a goal and assigns the pieces to worker agents, and anything exceeding an agent's pay grade escalates upward. It is an intuitive model. Every operator already knows how to read it, which is most of why it works as an interface. It is also the wrong unit of analysis for the one thing that decides whether an agent program survives contact with production. ## A reporting line is not an authorization boundary An org chart encodes who reports to whom. Agent safety is decided by what may be committed, on which workflow, by which principal. Those two things are not the same object, and they do not map onto each other. A manager agent and a worker agent have a reporting line. They have no authorization boundary between them. Nothing about the manager sitting above the worker in a diagram constrains what the worker is able to commit when it runs. The line is a coordination artifact. The boundary is a control, and it has to be written somewhere else. This is the distinction the Operational Agent Readiness Index scores under Authorization Maturity, which is the heaviest of its five dimensions. Authority is assessed per workflow, at three boundaries: what an agent may execute on its own, what requires human approval, and what has to escalate. An org chart is per role by construction. It cannot express a per-workflow boundary, because a role is not a workflow. The same role touches a dozen workflows with wildly different stakes, and a chart flattens all of them into one line. The consequence is specific, and it is worse than having no chart. An organization that has drawn its agent org chart has produced the artifact that feels like governance. It looks complete. It hangs on a wall. And it will answer the authorization diagnostic poorly while believing the opposite, which is a harder position to correct than not having started. ## The question that separates the two Underneath this sits a choice every operating model makes, usually without anyone noticing it was a choice. **Delegated execution.** The resting state is that work completes. Agents act, and human approval is an interrupt inserted on the actions someone flagged as sensitive. **Propose-only.** The resting state is that nothing happens. Agents produce proposals, and human approval is the only path to a write. Both models have approval gates. Both can show you an approvals queue. A feature comparison puts them in the same row. They are opposites, and the difference is not the gates; it is what happens in their absence. So the test is one question: **if nobody presses anything, what happens?** If the answer is that work proceeds, you are running delegated execution, and every unflagged action is a committed action. If the answer is that nothing happens, you are propose-only, and your exposure is bounded by construction whatever your policies say. Neither answer is wrong. They carry different risks, and they need different controls. What is expensive is not knowing which one you picked, and an org chart interface will not tell you, because both models draw the same chart. ## The map you still cannot produce There is a cheaper version of the same diagnosis, and it takes about a minute. Could you produce today, without investigation, a map of every agent you have deployed, every workflow it acts on, the permissions it holds on each, and the named human accountable for it? Most organizations cannot. The ones that have drawn an agent org chart usually cannot either, which is the point. They have an artifact, and it is not that artifact. It answers who, and the question is what, on which workflow, revocable by whom, and how fast. If assembling that map would take a week of investigation, that is not a documentation gap. It means the authority is not expressed anywhere that can be enforced, only inferred from configuration scattered across systems that were never designed to describe an agent. ## What happens when an agent assigns work to an agent The org chart interface introduces one structural feature that almost nobody prices, and it follows directly from the manager-agent pattern. When a manager agent assigns to a worker agent, the delegation chain extends. The question is whether authority inherits correctly across that call. Does the worker act within the bounds the manager holds, or does it act with its own standing permissions, which may be broader? Can privilege escalate silently across a hop that no human reviewed? The chart makes that chain look like an accountability structure. It is an authorization surface, and it is the one most likely to be unbounded, because it was drawn to describe delegation between colleagues, where the constraint was social rather than technical. ## Keep the chart, draw the second artifact None of this is an argument that the software is bad. These tools do a real job; they do it well, and most teams that bought one bought sensibly. Task decomposition needs a coordination model, and a chart humans can read at a glance is a good one. Keep it. The correction is that the chart is not the whole picture, and the missing half is the half that bounds the damage. Next to the chart, draw the artifact it cannot produce: per workflow, the three boundaries, the named human, and how fast authority can be tightened when something goes wrong. That second artifact is the one that answers the question a board asks after an incident. Earlier in this series, I argued that being cited by an AI system is not the same as being usable by one, because a surface that resembles readiness is not readiness. This is the operational version of the same error, one door over. A reporting line resembles a control. It is not one. The org chart tells you who to ask. It does not tell you what already happened without asking. ### The 40% That Get Scrapped Don't Fail on the Model URL: https://www.agenticlandmark.com/brief/the-40-that-get-scrapped-dont-fail-on-the-model/ Last updated: 2026-07-15T15:16:46.000Z *Gartner put a number on the coming agent shakeout. Which side of it you land on is decided before the first agent runs, by a dimension most teams never scope.* An agent in your operation takes a wrong action at two in the morning. A purchase order to the wrong supplier, a refund that should have been held, a message sent to a list it should never have touched. The question that decides everything is not whether the agent was smart. It is how much that one action had cost by the time anyone noticed, and whether any of it can be pulled back. That is the question Gartner was really answering when it predicted that over 40% of agentic AI projects will be canceled by the end of 2027\. The figure gets quoted as if it landed last week. It did not. Gartner published it in June 2025 and named three causes: escalating costs, unclear business value, and inadequate risk controls. A year of re-coverage has stripped the date off the number, but the prediction window is already half spent. ## None of the three reasons is the model Read the three causes again. Escalating costs. Unclear business value. Inadequate risk controls. Not one of them is a limit of the model. Drop the most capable model released this year into a project with no defined outcome and no owner, and you get a more articulate version of the same failure. The programs that die in 2027 will not die because the agent could not reason. They will die because no one decided, in advance, what the agent was allowed to break. This is the uncomfortable part for teams who bought the capability and skipped the operating discipline. The demo works. The pilot works. Then the project moves toward production, someone finally asks what happens when the agent is wrong, and the answer is a long pause. That pause is what a cancellation sounds like, a few budget cycles early. ## Blast radius is the dimension that decides There is a specific property that separates the agents that survive from the agents that get scrapped, and it has a name. Blast radius: when an agent acts wrong, how large is the damage, and how reversible is it. An agent that drafts a recommendation has a blast radius of nearly zero. If it is wrong, you discard the draft and it asks again tomorrow. An agent that moves funds, commits an order, or issues a refund has a blast radius measured in real money and real trust, and much of what it does is committed the instant it fires. Same underlying technology. Opposite risk. The difference is not intelligence. It is what the action touches and whether you can take it back. This is why two organizations can deploy the identical agent and reach opposite outcomes. One staged the action, put a ceiling on it, and caught errors before commit. The other pointed the agent at production with unbounded authority and monitoring that only noticed after the fact. The first has a tool. The second has a liability that has not gone off yet. ## Five questions that sort a survivable agent from a scrapped one For every agent workflow you are running or planning, the answers to five questions tell you which side of the number you are on. What is the worst realistic outcome of a single wrong action: a reversible inconvenience, a recoverable loss, regulatory exposure, or irreversible harm. How is a wrong action caught: a guardrail before commit, monitoring after commit, or a human noticing downstream. What can actually be rolled back, and what is committed the moment it fires. Is there a bounded limit on how much damage one agent can do before a human is forced in, or is autonomy unbounded once granted. And if the agent began acting wrong at two in the morning, how long until someone knew, and how long from knowing to stopped. A program that cannot answer these is not more advanced than one that can. It is closer to the 40%. ## The scrap happens at scoping, not at deployment The failure feels like it happens in production, the moment the wrong action fires. It actually happened much earlier, at scoping, when the blast radius was left at its default setting of unbounded and no one wrote that down as a decision. This is the other half of a gap I described in *Your Agents Are Using Someone Else's Credentials*: authority handed over without boundaries. Borrowed authority sets how far an agent can reach. Blast radius sets how much it can break before you can stop it. Put them together and you have the whole of what *When You Deploy the Agent, You Own the Mistake* named. The mistake is committed, and it is yours. Containing it is not model work. It is operational work: staging irreversible actions, setting per-action and cumulative ceilings, moving detection ahead of commit, and shortening the time from wrong action to human override. It is the least glamorous line in the deployment plan, and the first one cut when the project is sold on the strength of the demo. ## The window is real. It is just not the window you think. The urgency around agents is usually pitched as speed. Move first, deploy fast, do not be last. Gartner's number points the other way. The organizations on the safe side of 2027 will not be the ones that moved fastest. They will be the ones that scoped blast radius before they scaled, and can name, for every agent, the ceiling on its authority and the person who can shut it off. The strategic window and the urgency window are the same window here. They just do not feel the same until one of them has already closed, quietly, in a budget review, as a program that could not answer five questions gets reclassified as a pilot and wound down. Four in ten will. The work that keeps you out of that number is available now, and it costs less than the program you would otherwise scrap. ### Your Agents Are Using Someone Else's Credentials URL: https://www.agenticlandmark.com/brief/your-agents-are-using-someone-elses-credentials/ Last updated: 2026-07-15T02:56:26.000Z Pull the audit log for the last purchase order your agent placed. Whose name is on it? In most enterprises running agents today, the answer is a person. Sometimes a person who was on vacation that week. Sometimes a service account created in 2019 by an engineer who has since left. Occasionally an integration user with a name like `svc-erp-prod` that nobody can account for and nobody dares disable. What the log will not say is: an agent did this. Because as far as your identity system is concerned, no agent did anything. Agents do not exist in it. ## Agents are not a principal type Every enterprise identity model recognizes two kinds of actor. There are **humans**, who authenticate, hold roles, and are accountable. And there are **service accounts**, which are non-human identities that let systems talk to systems, and which are deliberately broad because they were designed to be plumbing, not decision-makers. An agent is neither of these, and that is the whole problem. It is not a human. It does not authenticate, it does not have a manager, and it cannot be held accountable in any meaningful sense. But it is also not plumbing. It makes decisions. It selects a supplier, sets a quantity, judges whether a return is resaleable. A service account moving a file between systems is not exercising judgment. An agent placing a purchase order is doing very little else. So when a team deploys an agent, they face a practical question with no good answer: what does it act *as*? And the path of least resistance is always the same. Give it a service account, or let it act on behalf of a human who has the permissions it needs. That is borrowing. And it is the default, not a mistake. Nobody decided to do this. It is simply what happens when you deploy a new kind of actor into a model that has no category for it. ## What borrowing actually costs you Three consequences follow, and they compound. **Authority becomes all-or-nothing.** Permissions in an identity model are shaped around *roles*, because roles were designed for people. A procurement analyst can acknowledge purchase orders, modify order terms, add suppliers to the approved list, and initiate returns, because a human in that job needs all of those, and a human exercises judgment about which one is appropriate. Hand that role to an agent and the agent inherits every one of those capabilities, in every situation, whether or not its workflow calls for it. The agent that was deployed to acknowledge routine POs can now modify order terms. Not because anyone granted that. Because nobody drew a smaller box. **Attribution breaks.** This is the one that surfaces in the audit log. If an agent acts through a borrowed credential, every action it takes is recorded as an action by the credential's owner. Your log is not wrong, exactly. It is just describing a world that no longer exists, one in which the entity that clicked and the entity accountable are the same entity. Now ask what happens when something goes wrong and someone has to reconstruct it. You cannot filter for what the agent did, because the agent is not a thing in the data. You cannot tell an agent's action from a human's. You cannot answer the first question any regulator, auditor, or incident review will ask, which is *who did this*, and you cannot answer the second one either, which is *who authorized them to*. **Revocation gets slow.** When an agent's authority needs to change, either tightened after a near-miss or widened after it proves out, how is that change made? If the agent's authority is a property of a borrowed role, then changing it means changing the role, which means changing what a human can do, which means a conversation with the people who own that role, and probably a ticket, and probably a release. Meanwhile the agent keeps acting. The gap between *we should narrow this* and *this is narrowed* is measured in sprints. An authority you cannot revoke quickly is not a control. It is a hope. ## The escalation nobody watches There is a fourth consequence, and it is quieter than the others. Agents invoke other agents. They invoke tools. They call downstream systems. And in most deployments, when agent A calls agent B, nothing constrains agent B to the authority that agent A was operating under. B has its own credential, which may be broader, and it will happily do the thing it was asked to do. You have granted agent A a narrow scope. Agent A, entirely within its scope, asks agent B for something. Agent B, with a wider scope, does it. The action that lands in the system is one that agent A was never permitted to take, and no rule was broken anywhere in the chain. This is privilege escalation through composition, and it does not look like an attack. It looks like a working system. ## What the fix actually is Make agents a principal type. Not a workaround, not a naming convention on a service account. A first-class, distinct kind of actor in the identity model, with its own credentials, its own scope, and its own name in the log. Once an agent is a principal, three things become expressible that were not expressible before. **Authority scoped per workflow, not per role.** The question stops being "what can this agent's role do" and becomes "what may this agent do *in this workflow*." The same agent, running an inventory replenishment task, gets a different grant from the one it gets running a supplier communication task. Roles cannot express this, because roles are about people, and a person carries the same permissions from task to task and uses judgment to decide what is appropriate. An agent has no such judgment, so the boundary has to be in the grant. **Three boundaries, drawn per workflow.** For every workflow an agent touches, exactly three things need deciding: - What may it **execute** autonomously, with no human in the loop - What must it **propose for approval**, with a human confirming inside a defined window - What must it **escalate**, where no autonomous action is permitted regardless of urgency or volume This is a small model, and it is deliberately small. Most teams, asked to define agent governance, produce a document. Asked instead to fill in three columns for each of their live workflows, they produce something enforceable, and they usually discover in the process that they cannot agree on where the lines go, which is itself the finding. **Two delegation patterns, and they are not interchangeable.** An agent can act in one of two modes, and conflating them is a common and expensive error. *User-delegated*: the agent acts on behalf of a specific person, and its authority is bounded by that person's authority. It can never exceed what its principal could have done directly. The human is present in the chain, and accountability flows to them. *Tenant-owned*: the agent acts for the organization, not for any individual. Its authority is granted directly, to it, and it does not inherit from anyone. Accountability flows to whoever configured the grant. Both are legitimate. They are appropriate in different places. What is not legitimate is being unclear about which one you are running, because the accountability chain is entirely different in each, and you will discover which one you were actually running at the worst possible moment. **And the grant is evaluated at action time.** Authority checked once at deployment and then assumed is not authority. It is a memory of authority. The check belongs at the moment of the action, against the current grant, so that a change to the grant is a change to what the agent can do, immediately, without a deploy. ## Why this is the first control, not one of many There is a reasonable objection here, which is that authorization is one of several things an agent program needs, alongside data quality, observability, escalation design, and incident response. That is true. It is also the wrong way to sequence them. Authorization is the control that bounds every other risk. An agent cannot take an action it lacks the authority to take. So the blast radius of any failure, including failures in the other controls, is capped by the grant. A model that hallucinates a supplier cannot onboard one if onboarding is escalate-only. A data error that produces an absurd reorder quantity cannot commit if the quantity exceeds the execute threshold. Bad data plus bounded authority is a caught error. Bad data plus borrowed authority is a purchase order. Everything else in a governance program tells you what happened, or helps you decide what should happen. Authorization is the only control that determines what *can* happen. It is the one that has to come first, and it is the one that is almost always deferred, because it touches the identity system and the identity system is owned by a team that did not ask for any of this. Which brings us back to the audit log. Someday, someone is going to pull it. It might be a regulator. It might be an auditor. It might just be you, at 2am, trying to work out how a thing happened. And when they ask which of your agents did this, and under what authority, you will either be able to answer, or you will be explaining that the log says a person did it, and the person was on vacation. The question is not whether the agent had permission. It is whether you can say whose permission it was using. ### When You Deploy the Agent, You Own the Mistake URL: https://www.agenticlandmark.com/brief/when-you-deploy-the-agent-you-own-the-mistake/ Last updated: 2026-07-14T18:16:54.000Z Everything I have written in this series has been about agents you do not control. Agents that shop you. Agents that read your product data and decide whether you make the shortlist. Agents that mediate the moment a customer used to spend on your site. Thirty-eight briefs on discovery, structured data, protocol integration, and what happens to a brand when the buyer stops browsing and starts delegating. All of it about a force arriving from outside, doing something to you. That is half the problem. I want to write about the other half, because the other half is the one that can actually cost you money this quarter. ## Two surfaces, not one A brand today is exposed to agents from two directions at once, and the two exposures are not the same problem wearing different clothes. On the **demand side**, agents mediate discovery and selection. You are the subject of a decision you do not control. An agent builds a comparison table, and you are either in it or you are not. The question is whether agents can find you, evaluate you accurately, and choose you. This is the surface the Agent Readiness Index scores, and it is the surface this series is exploring. On the **operational side**, you are not the subject. You are the actor. You are the one deploying agents, inside your own operations, against your own systems, with your own authority. An agent reorders inventory. An agent acknowledges a purchase order. An agent routes a shipment, or issues a refund, or decides whether a returned unit is resold or scrapped. The risk here is not something happening to you. It is something you are choosing. Same word, "agents." Different physics entirely. ## The asymmetry that makes this urgent Here is the part that should change how you sequence your investment. A demand-side agent that gets you wrong costs you a recommendation. That is real money, and I have spent thirty-eight briefs arguing it is worth taking seriously. But it is also *self-correcting*. The agent asks again tomorrow, and if your data is better, you win the next comparison. The failure is recoverable because the failure is a **decision that did not go your way**. An operational agent that gets it wrong does not cost you a recommendation. It places the purchase order. It issues the refund. It commits the supplier. By the time anyone notices, the thing has already happened, and the question is no longer whether you will be selected. It is whether you can unwind what your own system just did in your name. Recommendations are reversible. Actions are not. This is why the failure numbers on the operational side are so much uglier than the discourse suggests. Gartner projects that over 40% of agentic AI projects will be scrapped by the end of 2027 (Gartner, 2025). Read that carefully, because the interesting part is not the percentage. It is the cause. These programs are not failing because the models are not good enough. They are failing on governance: agents deployed before anyone defined what they were allowed to do, what data they needed to do it reliably, and how a failure gets caught before it compounds. Nobody scraps an agent program at the pilot. They scrap it after scaling exposed what the controls could not hold. ## Why one instrument cannot cover both The temptation, once you see the second surface, is to reach for the instrument you already have and stretch it. Resist that. An instrument built for demand-side exposure asks questions like: how dependent is your discovery on channels you do not own, how comparable is your product against alternatives, how substitutable are you in an agent's judgment. Those are the right questions for a brand that is the *object* of an agent's decision. They are the wrong questions entirely for a COO whose agents are placing orders against an ERP. That person does not need to know their traffic dependency. They need to know whether their authorization model can express what an agent may do per workflow, or whether it grants an agent everything its role can do. They need to know what the worst realistic outcome of one wrong action is, and whether it can be reversed. They need to know whether they would even *see* an agent misbehaving, or only find out downstream when the damage surfaced. Stretch one instrument across both surfaces and you get an instrument that is vague about both. The sharpness is the value. A diagnostic that asks a manufacturer's COO about their GEO investment is not being thorough. It is being useless, and it will tell them, incorrectly, that they are fine. So: two instruments. The Agent Readiness Index for the demand surface. The Operational Agent Readiness Index for the operational one. Different dimensions, different weights, different buyer, different questions. ## The seam where they meet Two instruments, but not two problems. They rest on the same foundation, and once you see the seam you cannot unsee it. The demand-side instrument has a dimension called Structured Data Maturity: can an outside agent read your product data accurately enough to evaluate you fairly. Fragmented catalog data means an agent recommends a competitor whose specs it could actually parse. The operational instrument has a dimension called Operational Data Readiness: can an inside agent read your systems of record accurately enough to *act* on them safely. When the same SKU has one lead time in the ERP and a different one in the warehouse system, an agent acts on whichever copy it reached. Confidently. At machine speed. That is the same discipline, pointed in two directions. Structured data read outward, so an agent can evaluate you. Structured data read inward, so an agent can act for you. A brand that fixes one has done real work toward the other, and a brand that has fixed neither is exposed on both surfaces simultaneously and usually knows about only one. ## One competence, two readings The conclusion I have arrived at, after writing about only half of this, is that agent readiness is not a marketing problem that happens to involve technology. It is an organizational competence, and it has two surfaces because the organization touches agents in two places. The CMO owns the surface where agents evaluate the brand. The COO, the CIO, and increasingly the Chief AI Officer own the surface where the brand delegates to agents. Those are different people, different budgets, and different fears. But it is one competence, and a company can be strong on one surface and dangerously exposed on the other without anyone noticing, because the two exposures are reported to different executives and measured in different meetings. Two instruments, then. Not because there are two problems. Because there are two surfaces, and one number cannot honestly describe both. I have spent this series arguing that the agents outside your walls are going to reshape how you are chosen. I still believe that. But the agents *inside* your walls are already spending your money, and the failure rate says most organizations have not yet drawn the boundaries that would let them do it safely. The demand-side agent can cost you a sale. The one you deployed yourself can cost you the program. ### Being cited is not being usable URL: https://www.agenticlandmark.com/brief/being-cited-is-not-being-usable/ Last updated: 2026-08-30T18:34:51.000Z The category error underneath most agent-ready advice: being cited and being usable are two different problems, and readiness for one does not transfer to the other. When an AI agent completes a task on your website, it does two separable things. First, it has to find you, which means being present in the answer the agent starts from. Then it has to act, which means doing the thing: filtering, comparing, adding to cart, checking out, on whatever surface you expose. Most advice about getting agent-ready treats these as one project. They are not. They run on different infrastructure, they fail in different ways, and the work you do on one buys you almost nothing on the other. That is the category error worth naming before you spend a dollar on it. ## The stack, from passive to interactive There are four standards a brand can put in front of an agent, and they line up on a single spectrum from passive to interactive. At the passive end sit robots.txt and llms.txt, two plain-text files that describe your site and tell crawlers and models what they may read. They are declarations. Nothing calls them. They sit there to be consulted, if anyone bothers. In the middle is Schema.org structured data, the machine-readable description layer. It is how an agent knows that this string is a price, that one is a rating, and this block is a product and its variants. It is the difference between an agent reading your catalog and an agent guessing at it. At the interactive end is WebMCP, and it is the genuinely new entry. WebMCP is a browser-native standard from Google and Microsoft, moving through the W3C, that lets a page declare its own callable tools: a search function, an add-to-cart function, a checkout function, with typed inputs an agent can invoke directly instead of screen-scraping your buttons. It is the browser-side counterpart to Anthropic's Model Context Protocol, the server-side standard that connects agents to backend systems. Where structured data lets an agent read your site, WebMCP lets an agent act on it. The move from robots.txt to WebMCP is the move from here is what I am to here is what you can do. That shift is the whole story. ## Two layers, not one checklist The practice has a name for the split these four standards sit across: the two-layer architecture. There is a discovery layer, where an agent decides whether you are even in the consideration set, and there is an infrastructure layer, where the agent actually transacts. robots.txt, llms.txt, and Schema.org are mostly about the first. WebMCP is squarely the second. This is not a semantic tidy-up. The two layers are owned by different teams, closed by different work, and measured by different numbers. Getting cited is a content-and-data problem: are you machine-readable, are you in the answer engines, does an agent see you when it builds its shortlist? Being usable is an interface-and-protocol problem: once an agent has chosen you, can it complete the task on your surface without falling back to brittle guesswork? Readiness for one does not transfer to the other. A brand can be immaculately structured and cited everywhere and still hand the agent a checkout it cannot operate. A brand can expose perfect WebMCP tools and never get called, because it was not in the answer to begin with. ## The order is the argument Here is the part most coverage misses, because most coverage is technical-SEO how-to that treats the stack as a checklist to complete in parallel. It is not a checklist. It is a sequence, and the sequence runs in a fixed direction. An agent's journey starts with a model's answer. The buyer asks the agent for a running shoe that fits their gait and their budget, and the agent produces a shortlist before it touches a single website. If you are not in that answer, nothing downstream matters. Your WebMCP tools are never invoked, because the agent never arrives. Citation is the prerequisite. It is the entry point to everything else. Which is why the execution layer, for all that it is the exciting new thing, only pays off once the discovery layer is already true. Building agent-callable tools before you are agent-visible is building the second floor before the first. But the inverse is a real failure too, and it is the one brands are sleepwalking into. Cited and not usable means the agent finds you, chooses you, arrives, and then cannot complete the task, so it routes around you to whoever it can actually transact with. You paid for the visibility and lost the sale at the last step. Both failure modes are live. Only the order in which you fix them is fixed. ## What to actually build, and when The honest guidance follows from the sequence, not from the newest standard. Fix discovery first, because it gates everything and because most brands are not close. Adobe's 2026 data puts average product-page machine readability at 66 percent, with the weakest performers down near 54\. The description layer, the boring middle of the stack, is where the real gap is, and it is the cheapest to close. And the citation channel is worth closing it for: Seer Interactive measured ChatGPT-referred conversion at 15.9 percent against 1.76 percent for Google organic, in a single B2B software case study covering October 2024 to April 2025\. The conversion action was a signup rather than a purchase, so the ratio is a directional signal about the citation channel rather than a merchant benchmark.. Being in the answer is not a vanity metric. Then, and only then, invest in execution. Here is a note of restraint the how-to pieces skip: WebMCP is early. It is in a Chrome origin trial, native support in other browsers is unconfirmed, and a formal web standard is years away, not months. That is a reason to understand the architecture and watch the trial, not a reason to rush a build. The brands already in the trial are the ones agents most need to transact with, in travel, retail, and financial services, which tells you where this lands first. It does not make it production infrastructure yet. One more piece of restraint, at the passive end: do not overweight llms.txt. Google's own John Mueller called it purely speculative in mid-2026, because no major AI system actually consumes it. A file nothing reads is not a readiness layer. Spend the structured-data effort where agents are demonstrably looking. ## The sequence, stated plainly Citation gets you into the answer. Execution gets you acted upon. Both require work now, but the second is worth nothing until the first is true, and no amount of the first survives a broken second. So the question to bring to your own site is not are we agent-ready, which flattens two problems into one and usually gets answered by whoever owns the newest acronym. It is two questions, asked in order. Are we in the answer. And when the agent arrives, can it do the thing? A brand can be perfectly cited and completely unusable, or perfectly usable and never cited. Both of those brands lose. The one that wins fixes them in the order the agent experiences them. **Correction, August 30, 2026.* This post originally cited the Seer conversion figures without their basis. The source is a single B2B software client, October 2024 to April 2025, roughly 1,370 AI conversions against approximately 14 million organic sessions, and the conversion action is a signup rather than a purchase. The figures are unchanged; the framing is corrected.* ### Your Best-Converting Channel Is the One You Cannot Measure URL: https://www.agenticlandmark.com/brief/your-best-converting-channel-is-the-one-you-cannot-measure/ Last updated: 2026-07-02T17:55:38.000Z *The measurement gap in agent-mediated commerce is not a reporting problem. It is a budgeting problem, and it defunds the channel that is already winning.* ## The channel is converting. The analytics are not reporting it. In March 2026, AI-sourced traffic to US retail sites converted 42% better than non-AI traffic. Twelve months earlier, it converted 38% worse. That is roughly an 80-point swing in a single year, recorded alongside AI traffic to US retail sites growing 393% year over year in Q1 (Adobe Q2 2026 AI Traffic Report). The channel is no longer exploratory. It is high-intent, high-conversion commercial traffic. Now open your analytics and try to find it. For most brands, the number is not there. Not because the traffic is not arriving, but because the stack cannot see it for what it is. Agent-driven revenue lands in the same buckets it always has: direct, organic, unassigned. The brand is capturing the best-converting channel in its funnel and filing the proceeds under a label that tells the CMO nothing. ## Why agent revenue arrives disguised There are three separate mechanisms, and it is worth naming each precisely. First, agent traffic is currently indistinguishable from bot traffic in most analytics stacks, and no cross-platform standard exists to separate them. An agent evaluating a product looks, to a conventional analytics tool, like something to filter out. Second, the agent frequently evaluates without visiting at all. Botify's crawl data shows roughly one human visit per 198 OpenAI crawls, against one visit per six Google crawls. The commercial event, the moment the brand is compared and shortlisted, happens upstream of any session your tools were built to record. Third, and most concretely, the payment layer now hides the transaction itself. Buyer-side agent payment credentials such as Skyfire and Visa Intelligent Commerce issue a spending instrument to the agent and settle through ordinary card rails. The brand receives a standard card transaction with no protocol flag. A brand can be fully UCP-compliant, meaning its checkout is exposed to agents through the Universal Commerce Protocol, and still take agent-initiated card purchases it cannot distinguish from a person typing in a card number. Without server-side signal detection, those purchases misattribute to direct or organic by default. ## Two failures hiding inside one number The phrase "attribution gap" actually contains two distinct failures, and they require different fixes. The first is agent traffic measurement: can you see the visit at all, and classify it correctly? This is a detection problem, solved with server-side signal work. The second is agent-influenced attribution: can you credit the agent recommendation that preceded a human purchase? There is no protocol-level standard for this today. When an agent surfaces your brand in a consideration set, and the consumer later completes the purchase themselves, the recommendation did the selling and the last click takes the credit. This is the discovery versus selection problem made financial. The agent handled discovery. Your attribution model only knows how to reward the surface where selection was recorded. Recommend and transact are not the same event, and they are increasingly not measured by the same system. ## The moving baseline Even brands that solve detection hit a methodology problem underneath. Conventional conversion measurement A/B tests against a control group holding a stable price. Agent-mediated selection breaks that assumption. The agent performs real-time comparison at evaluation time, so the price at the moment of the agent's query is what determines selection, not the average price across a test window. For any brand running dynamic pricing, the human-channel attribution model is now measuring against the wrong baseline. The agent channel needs a separate model that accounts for price at time of selection rather than price at time of click. Otherwise, the brand is grading that channel on a test it was never actually running. ## The blind spot is self-reinforcing Here is why this is not a back-office reporting nuisance but a strategic exposure. Misattributed agent revenue does not disappear. It gets reassigned to direct and organic channels that already look healthy and already own their budgets. The agent channel, meanwhile, shows up in the dashboard as low volume and unproven ROI, because most of what it produced was credited to something else. So it gets underfunded. The brand starves the exact channel converting 42% better than everything else, and it does so on the strength of numbers it believes are accurate. That is the trap. The measurement gap does not just hide performance. It actively redirects investment away from the winning channel and toward the channels quietly absorbing the winner's credit. Left alone, it compounds every planning cycle. This is why the practice treats Measurement Readiness as a qualifying dimension, independent of the five scored ARI dimensions. Traffic Dependency, Structured Comparability, Agent Substitutability, Brand Moat, and Structured Data Maturity tell a brand how exposed it is. Measurement Readiness asks the prior question: would you even be able to see that exposure playing out in your own numbers? A brand can score well on catalog structure and still be blind on channel performance, which means it cannot manage the transition it is already inside of. ## What this asks of the CMO The remediation is not exotic. It is server-side agent traffic classification, an attribution model that credits agent-influenced discovery separately from last-click selection, and, for dynamically priced catalogs, a baseline that measures price at time of selection. None of it requires waiting for a standard to settle. The brands that build measurement infrastructure now will be able to defend the agent channel's budget with evidence. The brands that wait will keep defunding their best channel with confidence. You cannot optimize a channel you cannot see. And right now, the best-converting channel most brands have is the one their own analytics are trained to ignore. ### What It Means to Be a Machine-Readable Brand URL: https://www.agenticlandmark.com/brief/what-it-means-to-be-a-machine-readable-brand/ Last updated: 2026-07-02T02:52:48.000Z Every brand has spent years learning to be memorable to humans. The next decade is about being legible to machines. Those are not the same skill, and the brands that assume one produces the other are the ones agents will quietly route around. A machine-readable brand is not a brand with good SEO. It is not a brand with a chatbot on its site. It is a brand whose identity, claims, and offerings exist as structured, verifiable data that an agent can retrieve, parse, and reason from without a human in the loop. Most brands are not this. They have spent their infrastructure budget on the human-attention layer and almost nothing on the machine-legibility layer, and the gap between the two is now a commercial exposure. Here is the uncomfortable version of the problem. A challenger brand with a clean entity record, accurate attribute encoding, and strong presence in high-authority sources can have higher agent awareness than a category incumbent with decades of brand equity and no structured knowledge layer. The agent does not know who the market leader is. It knows what it can retrieve and verify. Brand equity that lives only in human memory and emotional association does not transfer to a system that has neither. What being machine-readable actually requires, concretely: An accurate entity. The brand has to exist as a correctly attributed entity across the structured sources that feed agent knowledge, Wikipedia, Wikidata, schema.org, the knowledge graphs agents query. A brand that is not a clean structured entity is not simply unknown to agents. It is inconsistently known, which is worse, because inconsistent entity representation produces inconsistent recommendations and factual errors that persist across every platform at once. Positioning translated into structured claims. This is the part most brands underestimate. A positioning statement is a human-attention asset. It works because humans respond to narrative. An agent cannot retrieve "premium build quality." It can retrieve a stainless steel boiler with PID temperature control. The work of becoming machine-readable is translating every brand claim that currently lives in campaign language into a verifiable, structured, queryable attribute. Where a positioning claim has no machine-readable evidence behind it, that claim is invisible to the agent evaluating the brand. Specificity where brands are used to atmosphere. Ravi Evani of Publicis Sapient put it precisely for hospitality: agents rank hotels not by slogans but by structured facts. Not "family-friendly" but connecting rooms, kids' menus, step-free access, cribs available. The same rule governs every category. In financial services, "competitive rates" is invisible; a specific APR and a documented eligibility criterion are not. In travel, "stunning views" is unquantifiable; floor level, window orientation, and landmark distance in meters are. Chris Riedy of Ibotta named the failure mode in one line: agents are not looking at display ads, they are looking at the inherent quality and metadata of the product, including its price. This is one of the five dimensions the practice scores in its Agent Readiness Index, Structured Data Maturity. The diagnostic question it asks is not whether the brand is well known. It asks whether the brand's data and infrastructure are ready for machine evaluation, because a brand that scores low here cannot be accurately assessed by an agent even when the agent tries. Low maturity does not produce competitive loss. It produces misrepresentation and omission, which is the quieter and more expensive failure, because the brand never sees the transaction it was never entered into. The reframe that matters for a CMO: being machine-readable is not a technical project you delegate to the data team and forget. It is a brand strategy question. The brand narrative your team spent years building has to be re-expressed in a form a machine can verify, or it does not exist in the channel where an increasing share of discovery and selection now happens. The story still matters for the human who reviews the agent's recommendation before committing. But if the agent could not read the brand in the first place, that recommendation was never made. Being unforgettable was the last era's competitive advantage. Being machine-readable is this era's table stakes. The brands that understand the difference are already translating their equity into structured form. The ones that assume their reputation will carry them are about to discover that the agent has never heard of them. ### Palmata measures how agents describe you. Not whether they can buy from you. URL: https://www.agenticlandmark.com/brief/palmata-measures-how-agents-describe-you-not-whether-they-can-buy-from-you/ Last updated: 2026-07-01T22:25:49.000Z Contentful launched Palmata on June 23\. Read past the product announcement, and it is the clearest signal yet that the measurement layer for agentic commerce is now a category, not a feature. Palmata's core argument is one the practice has been making for months: visibility is only the starting point. Most AI discovery tools tell a brand whether it shows up in ChatGPT, Gemini, or Perplexity, how often it is cited, and how its share of voice compares. Those are the right questions to ask second. They are not the questions that change anything. The questions that change something are the ones Palmata is built to answer. How is the brand being represented? Why are answer engines describing it this way? What content should change first? That is the difference between knowing you have a problem and knowing what to do about it. Two things about this launch matter beyond the product itself. The first is what it says about where the money is moving. Contentful acquired the team behind this (Writ) and built it into a standalone platform, launched independently from the core CMS. Salesforce paid a billion dollars for Contentful, in part, for the content architecture underlying exactly this kind of capability. The measurement and representation layer of AI discovery is now attracting acquisition capital and dedicated product investment. When that happens, the category has crossed from emerging to established. The second is the framing Palmata uses, because it maps precisely to the layered model brands should already be working from. Palmata measures and improves how a brand is represented at the discovery layer, the moment before a buyer reaches an owned channel. This is the recommendation layer. It determines whether an agent surfaces and describes a brand accurately when a consumer asks for a recommendation. Which is exactly why the measurement story does not end here. Palmata answers how a brand is represented when an agent talks about it. It does not answer whether an agent can transact with that brand once it decides to recommend it. Those are two different readiness questions with two different measurement problems. Discovery representation is a content and evidence problem, and Palmata is built for it. Transaction execution is an infrastructure problem, and it needs a different instrument entirely. The brands that will navigate the next two years well are the ones that measure both. Representation at the discovery layer, so they know how agents describe them. Transactability at the infrastructure layer, so they know whether agents can actually complete a purchase once the recommendation is made. Palmata is a real step forward. It makes the discovery-representation problem measurable, which is the prerequisite to fixing it. The launch is worth attention not just for what it does but for what it confirms: measuring your brand's AI reputation is no longer optional infrastructure. It is becoming table stakes, and the tooling is arriving to make it possible. The next instrument the market needs measures the layer below representation. When an agent decides to buy on a consumer's behalf, can it actually transact with you? That report does not exist yet. It will. ### Agentic adoption is a gradient, not a number URL: https://www.agenticlandmark.com/brief/agentic-adoption-is-a-gradient-not-a-number/ Last updated: 2026-07-01T22:16:52.000Z Accenture surveyed 25,590 consumers across 16 countries on AI shopping agents. The headline everyone is quoting: 74% would trust a personal AI agent more than their best friend to purchase on their behalf. That is the wrong number to fixate on. The useful finding is the gradient underneath it. 74% would let an agent handle routine tasks: deal negotiation, complaint resolution, subscription renewals, and reorders. 32% would let an agent make a purchase decision within defined limits, but still review before payment. 9% would let an agent initiate and complete a purchase without final approval. 74, then 32, then 9\. That is not a single adoption number. It is three different levels of delegated authority, and the same consumer occupies different levels for different purchases. This is the most important thing the report shows, and it confirms something the agentic commerce conversation keeps getting wrong. There is no such thing as "agentic commerce adoption" as a single figure that moves over time. Adoption is a spectrum of delegation that varies by category, by stakes, and by how much the consumer values making the decision themselves. The report makes the category point directly. Recurring services rank highest for delegation. Lifestyle and travel purchases drop sharply as autonomy increases. Consumers delegate grocery restocking and keep control over the hotel room, the clothing, and the experience. Effort-heavy and low-emotion gets delegated. Identity-linked and high-consideration stays in the consumer's hands. For brands, this changes what readiness means. You are not preparing for one wave of adoption. You are serving consumers who sit at different points on the delegation spectrum for different things you sell. The replenishment-class product needs a full transaction infrastructure now, because the consumer is already in the 9% for that category. The high-consideration product needs to win the consideration set, because the consumer will stay in the 74%-that-reviews for years, possibly forever. The 74% best-friend stat will get the headlines. The 74-32-9 gradient is the number that should shape the strategy. ### A universal translator still needs something worth translating URL: https://www.agenticlandmark.com/brief/a-universal-translator-still-needs-something-worth-translating/ Last updated: 2026-07-01T22:00:48.000Z Adyen announced Adyen Agentic this week, and the most strategically useful line in the release is the one most coverage will skip. "Without having to bet on which ecosystems ultimately win." That is the protocol-fragmentation problem stated plainly, and it is the question every commerce leader is quietly stuck on. UCP, ACP, AP2, Meta's AI checkout, whatever ships next quarter. Each agentic platform operates on different protocols, different product data formats, and different cart and checkout requirements. Integrating with each one individually is a series of bets on which platforms matter, made before anyone knows the answer. Adyen's pitch is to be the abstraction layer that removes the bet. Integrate once. They translate that integration across every protocol and platform. The merchant participates in new channels without rebuilding each time the landscape shifts. The structure is worth noting because it maps to how agentic commerce actually decomposes. Three layers: Agentic Feed for product and inventory distribution, Agentic Cart for checkout and order orchestration, Agentic Payments for authentication and fraud. Discovery, cart, transaction. The same three layers every brand has to solve, packaged as infrastructure. The detail that matters most for brands: merchant of record preservation is built into the payments layer. This is the same architectural choice Google made with Universal Cart and Visa made with Intelligent Commerce. The platforms have collectively decided that the merchant stays the merchant of record. That question is now settled across every major agentic commerce rail. What is not settled, and what an abstraction layer cannot solve for you, is whether your product data is good enough to be useful once it is translated. Adyen Agentic distributes your catalog across every platform. If that catalog is incomplete, unstructured, or inaccurate, the universal translator faithfully distributes incomplete, unstructured, inaccurate data everywhere at once. The abstraction layer solves the integration problem. It does not solve the data quality problem. Those remain two different builds, and the second one is still the brand's job. The protocol-agnostic infrastructure thesis is right. Build for the layer, not the platform. Just remember that what flows through the layer is still yours to get right. ### When Will Consumers Actually Let Agents Buy? URL: https://www.agenticlandmark.com/brief/when-will-consumers-actually-let-agents-buy/ Last updated: 2026-07-01T21:12:08.000Z *Everyone agrees agents will transact on consumers' behalf. Almost no one has mapped when, for what, and what has to be true first. Here is the timeline that matters.* The agentic commerce conversation has a consistent blind spot. Nearly every analysis conflates two different consumer behaviors: using AI to discover and compare products, and letting AI complete the purchase. The first is already mainstream. A quarter of US online adults have used ChatGPT to search for products in the past month. The second delegated transaction, where the consumer does not click buy, is happening at a fraction of that rate, and the gap between the two is the most misunderstood dynamic in the category. OpenAI launched Instant Checkout and pulled it within six months. The discovery volume was enormous. The transaction volume was not. That gap is not a failure of technology. It is a function of trust, and trust accrues on a schedule that infrastructure announcements cannot accelerate. This article maps that schedule. Not with dates, dates are guesses, but with the sequence that actually governs when consumers will let agents transact, and the conditions that have to be true for each stage to arrive. ### The unit of analysis is not the consumer. It is the consumer-category pair. The central error in most adoption forecasts is treating "agentic commerce adoption" as a single number that moves over time. It is not. The same consumer will delegate a grocery reorder to an agent while insisting on personally choosing a mattress. Adoption does not happen consumer by consumer. It happens category by category, within each consumer, at different rates. This is why the delegation posture framework is the right lens. Consumers are not Curators or Delegators as a fixed identity. They occupy different postures for different purchase categories simultaneously. The adoption timeline is the rate at which categories migrate from low-delegation postures to high-delegation postures, and that rate is governed by the characteristics of the category itself. Four postures, in increasing order of delegated authority: The Curator researches with AI assistance but makes every decision personally. The agent informs. The human decides and buys. The Scoper sets parameters, price ceiling, preferred brands, specifications, and lets the agent build a shortlist or execute within those bounds, reviewing the result. The Delegator hands off the entire transaction for defined categories within standing rules, reviewing exceptions rather than decisions. The Orchestrator manages a portfolio of agents handling continuous commerce relationships, intervening only when the system surfaces something outside its mandate. The transaction adoption timeline is the migration of categories through these postures. Here is the order it happens in, and why. ### Stage one: replenishment and routine reorder The first category to reach full delegation is the one where the decision was already made long ago. Weekly grocery staples. Household consumables. Subscription restocks. Pet food. The purchases where the consumer already knows exactly what they want, has bought it many times, and experiences the act of buying as friction rather than choice. These purchases reach the Delegator posture first because the consumer is not giving up a decision they value. They are giving up a chore. What has to be true for this stage: The agent must have access to a purchase history accurate enough to reconstruct the standing order. The fulfillment must be reliable enough that an error does not break trust. The exception handling, what the agent does when an item is out of stock, must be good enough that the consumer is not surprised in an unwelcome way. This stage is arriving now. It is the category where Walmart's Sparky, Instacart's agent integrations, and Amazon's replenishment infrastructure are already operating. The trust threshold is low because the stakes per transaction are low and the consumer's preference is already established. ### Stage two: specified replacement and commodity purchase The second category to migrate is the specifiable purchase, where the consumer knows the attributes they want but not necessarily the specific product. The 18V cordless drill with a brushless motor is under $200\. The replacement water filter that fits a specific refrigerator. The printer ink, the phone case, and the standard-specification item, where the consumer can articulate requirements precisely, and does not have a strong brand attachment that requires personal evaluation. This reaches Scoper posture readily and Delegator posture for repeat instances. The consumer declares the specification and trusts the agent to match it. The migration is gated by the agent's ability to evaluate product attributes accurately against the declared specification, which means it is gated by the quality of the structured product data the agent can query. What has to be true for this stage: Brands need machine-readable product attributes complete enough for an agent to match against a specification. Inventory and pricing must be queryable in real time. The protocols that let the agent transact ACP, UCP must be live in the consumer's chosen platform. This stage is beginning now and accelerates as product data infrastructure matures across brands. ### Stage three: considered purchase with established preference The third category is the considered purchase, where the consumer has a formed preference, but the transaction stakes are high enough to warrant review. The repeat apparel purchase from a known brand. The replacement of a product that the consumer has owned and liked. The upgrade within a product line they trust. Here, the consumer is willing to delegate execution but wants to review before commitment, because the purchase matters enough that an error has real cost. This category reaches Scoper posture but resists Delegator posture longer. The consumer will let the agent assemble the purchase, but wants to approve it. The migration to full delegation depends on the accumulated trust the consumer has that the agent gets it right enough times to stop reviewing. What has to be true for this stage: The trust record has to exist. This is the stage where the compounding mechanism matters most. An agent that has executed lower-stakes purchases accurately builds the trust that lets a consumer extend authority to higher-stakes ones. The infrastructure is necessary but not sufficient. The differentiator is the behavioral history between the consumer and the agent. ### Stage four: high-consideration and first-time purchase The final category to migrate and the one that may never fully migrate is the high-consideration, high-stakes, taste-dependent, or first-time purchase. The first espresso machine. The vacation. The car. The mattress. The purchases where the consumer does not have an established preference, where the choice expresses something about them, or where the stakes are high enough that the consumer wants to own the decision regardless of how capable the agent is. These purchases stay in the Curator posture longest, and some stay there permanently. The agent will research, compare, build the consideration set, and surface the rationale. But the consumer makes the call. Not because the agent cannot decide, but because the consumer does not want to delegate this particular decision. What has to be true for this stage: This is less an infrastructure question than a human one. Some categories will migrate as trust accrues and as younger consumers who grew up delegating reach the life stages where these purchases happen. Others will resist indefinitely because the consumer derives value from making the choice. The error is assuming this category migrates on the same curve as the others. It does not. ### What this means for brands The adoption timeline is not a single curve. It is four curves moving at different rates, and a brand's position on each depends on the categories it sells into. A brand selling replenishment-class products is already in the transaction-delegation era. The infrastructure investment agent-accessible checkout, real-time inventory, and structured reorder data is not a future project. It is a present requirement, and competitors who built it are already capturing delegated transactions. A brand selling specifiable commodity products is entering stage two now. The gating factor is product data quality. The brands with complete, machine-readable, accurately structured attributes are being matched against consumer specifications. The brands without them are invisible to the agent doing the matching. A brand selling considered purchases with established preference is in the trust-accrual stage. The infrastructure matters, but the differentiator is the brand's performance in the lower-stakes agent interactions that build the trust record. Every accurate fulfillment, every correct recommendation, every well-handled exception is an input to whether the consumer eventually delegates the higher-stakes purchase. A brand selling high-consideration products should not be optimizing for full transaction delegation. It should be optimized to be the brand that the agent surfaces and explains when the consumer is making the decision personally. The job is to win the consideration set, not the delegated transaction. ### The mistake to avoid The single most expensive mistake in agentic commerce planning is applying one adoption timeline across a portfolio that spans multiple categories. A retailer that sells both replenishment staples and high-consideration durables cannot run one agentic commerce strategy. The staples need a full transaction infrastructure now. The durables need discovery and consideration-set optimization with human-decision support. Treating them the same means either over-investing in transaction infrastructure for categories that will not use it, or under-investing for categories that already demand it. The timeline is real. It is just not singular. The brands that map their own categories against the posture migration and invest according to where each category actually sits are the ones that will neither be caught unprepared nor waste capital building for a delegation that the category will not adopt. Agents will transact on consumers' behalf. They already are, for some things. The question was never when agentic commerce arrives. It is which of your categories it has already arrived for, and which it never fully will. ### When Gartner names a category, buyers get a budget URL: https://www.agenticlandmark.com/brief/when-gartner-names-a-category-buyers-get-a-budget/ Last updated: 2026-07-01T20:51:42.000Z Gartner published an Innovation Insight report on May 12th, naming "Agentic CMS" as a distinct category. The finding that will get cited in every enterprise content budget conversation for the next two years: By 2028, 80% of customer interactions will shift from web, search, social, and mobile to agentic AI interfaces. And traditional WCM and headless CMS approaches will be rendered obsolete as content workflows transition from storage and delivery to automation, decisioning, and orchestration. When Gartner names a category, enterprise technology buyers get a budget. That is what this report does for content infrastructure modernization. The timing is not coincidental. The Salesforce/Contentful acquisition closed the same week. Two signals pointing at the same structural shift: the content layer that brands built for human editorial workflows is not the content layer that agents can query and assemble from. The architecture required for agentic content delivery API-first, structured, channel-agnostic is a different build from the architecture that generated the last decade of CMS investment. The 80% figure is the one to deploy in board conversations. It is not a niche technology prediction. It is a projection about where the majority of customer interactions will happen within two years. If 80% of customer interactions shift to agentic interfaces and your content infrastructure was built for the other 20%, the gap is not a roadmap item. It is a structural readiness problem. The Gartner category naming and the Salesforce acquisition in the same week are the analyst and market validation arriving simultaneously. For brands still treating content infrastructure as a maintenance budget rather than a competitive investment, this is the week the argument changed. ### "When AI Builds Itself." The Bottleneck Is Shifting. URL: https://www.agenticlandmark.com/brief/when-ai-builds-itself-the-bottleneck-is-shifting/ Last updated: 2026-07-01T20:44:37.000Z *Anthropic just published the most data-rich account of AI acceleration from inside a frontier lab. Here is what it means for business leaders who are not building AI.* Anthropic published something this week that every business leader should read slowly. Not because it is about AI safety. Because it contains the most honest account of what AI acceleration actually looks like from inside a frontier lab with production data, and the three futures it describes have direct implications for how organizations should be planning right now. The article is called "When AI Builds Itself." It covers recursive self-improvement: the trajectory toward AI systems capable of designing their own successors. Whether you believe that moment is years away or decades away, the evidence Anthropic presents about where we already are today is the part worth paying close attention to. ### The data The headline number: as of May 2026, more than 80% of the code merged into Anthropic's codebase was authored by Claude. Engineers are shipping 8x as much code per quarter as they were from 2021 through 2024\. The task horizon, the length of work an AI agent can reliably complete autonomously, has been doubling every four months. Tasks that take a skilled person days are expected to come within range this year. Tasks taking weeks are projected for 2027. The most striking internal result: in April 2026, Claude-powered agents were given an open-ended AI safety research problem and left to solve it. Two human researchers recovered roughly 23% of the performance gap over a week. The agents recovered 97% over 800 cumulative hours using roughly $18,000 in computing. The humans chose the problem. The agents designed every experiment themselves. This is not speculative. It is production data from an organization that is, by most measures, among the most technically capable in the world. ### The three futures The article describes three possible futures. Anthropic is clear about which one they believe is most likely. The first future is that the trend stalls. Returns diminish, the S-curve bends, some constraints compute scarcity, a missing architectural insight, and regulatory disruption slows progress. The exponential curves flatten. Anthropic includes this scenario for completeness. They do not believe it describes where we are headed. The second future is compounding efficiency gains without full recursive self-improvement. Humans continue to set research directions and exercise judgment. AI handles execution at accelerating speed. The result is a fundamental restructuring of organizational capacity: 100-person companies doing the work of 10,000-person organizations. Knowledge work transformed. The bottleneck shifts from doing to deciding. Anthropic believes this is the scenario we are entering. The third future is full recursive self-improvement AI systems capable of designing their own successors, with humans playing an oversight and validation role rather than a creative one. This future is possible. Anthropic views it as less likely in the near term than the second, but not impossible. They are building governance frameworks for it anyway. For business leaders operating outside the AI industry, the second scenario is the one to plan around. Its implications are not abstract. They are already visible in organizations that are paying attention. ### The productivity floor is rising A small team with access to capable AI is no longer a small team in the traditional sense. It is a small team with a multiplier that compounds as the models improve. Anthropic's data illustrates this from the inside. An engineer steering Claude through a complex debugging task is not doing the same job as an engineer writing code line by line. They are doing a higher-leverage job directing, evaluating, and course-correcting while the output scales beyond what any individual engineer could produce alone. The 8x productivity figure is an artifact of that shift. The same dynamic is available to organizations outside AI development. The teams that understand this earliest are already operating as if they have significantly more capacity than their headcount implies. They are running analyses, drafting communications, testing hypotheses, and building internal tools at a pace that would have required much larger teams two years ago. The competitive implication is straightforward. Organizations that have built the workflows, data architecture, and operating culture to work effectively with capable AI are accumulating a capacity advantage. That advantage compounds. Every improvement in model capability increases the return on having built the infrastructure to use it. Organizations that have not built that infrastructure find themselves further behind with each model generation, not just because they lack the tools but because they lack the institutional practice of working with them. ### The bottleneck is shifting from doing to deciding The article is direct about this: the human comparative advantage, for now, is research taste and judgment. Choosing which problems deserve attention. Evaluating which results to trust. Recognizing when an approach is a dead end. For most organizations, this represents a significant structural challenge. The people with the best judgment are not evenly distributed. Judgment concentrates at the top of organizations that have invested in developing it through years of experience, through exposure to complex tradeoffs, and through the accumulated pattern recognition that comes from sustained engagement with hard problems. The pyramid staffing model was built around a different assumption: that judgment at the top needed execution distributed below it to produce results at scale. AI changes this. Execution is becoming cheap. Judgment remains expensive. An organization that recognized this shift early would be restructuring not to add AI tools to existing roles but to concentrate judgment-dense talent at the center and let AI handle the execution perimeter. This is what the Block hierarchy-to-intelligence model describes at the organizational level. It is what McKinsey's AI Transformation Manifesto describes at the enterprise level. And it is what Anthropic's data shows happening in practice at the model development level. The same structural shift is visible across all three, because it reflects something fundamental about what becomes valuable when execution becomes abundant. ### Comprehension is the new constraint The employee quote that stayed with me most from the article does not come from Anthropic's institutional voice. It comes from an employee describing their own experience: "On days where everything works well, I can't help but think nothing I do matters, everything is automated and better and faster than I ever will be. But then there are days where everything breaks, and I don't understand why, and I realize I have no idea what I've been up to anymore." That tension between the efficiency gains and the loss of comprehension is the most honest description of where many organizations are right now. The doing accelerates. The understanding does not keep up. The leaders who figure out how to maintain comprehension while capturing the acceleration will be the ones who can actually steer where it goes. This is not a warning against using AI. It is a warning about how to use it. The organizations that deploy AI as a speed tool, making existing processes faster without understanding what those processes are optimizing for, are accumulating execution debt. They are producing more outputs with less understanding of what the outputs mean, how they were produced, or what the failure modes are. The organizations that deploy AI as a judgment amplifier, using it to extend the reach of genuine human expertise rather than to substitute for understanding, are building something different. They are building comprehension at scale. ### What to do with this Anthropic frames the second future, compounding efficiency gains, humans directing AI execution as both the most likely near-term outcome and an enormous opportunity. They are right on both counts. The practical implication for leaders who are not building frontier AI is not to build it. It is to understand what the second scenario means for their own organizations and build the infrastructure to participate in it. That means workflow architecture: designing how human judgment and AI execution interact across the organization's core processes, not just in isolated tools or experiments. It means data infrastructure: ensuring the information the organization needs to make good decisions is structured, accessible, and usable by the AI systems that will increasingly assist in making those decisions. It means judgment investment: identifying where genuine human expertise is irreplaceable, which problems deserve attention, which tradeoffs matter, which results to trust and protecting and developing that capacity deliberately rather than letting it erode as execution becomes automated. The window for building these capabilities while the multiplier is still small enough to learn with is open. It will not stay open at the same cost forever. The compounding that makes the second scenario so consequential also means that the distance between organizations that were built for it and organizations that did not grows with every model generation. The article is worth reading in full. It is the clearest picture available of where the technology actually is, described by people who are building it. The three futures are a framework for thinking about what comes next. The data is a description of what is already here. Read it: [https://www.anthropic.com/institute/recursive-self-improvement](https://www.anthropic.com/institute/recursive-self-improvement?ref=agenticlandmark.com) ### The Merchant of Record Is Not the Merchant of Decision URL: https://www.agenticlandmark.com/brief/the-merchant-of-record-is-not-the-merchant-of-decision/ Last updated: 2026-07-01T20:34:50.000Z *Every agentic commerce protocol was designed to reassure merchants. Merchants should read the fine print.* Every major agentic commerce protocol announced in the last six months, UCP, ACP, AP2, was designed to preserve something specific: the merchant stays the merchant of record. That was not an accident. It was the concession that made the protocols politically viable. Google, Visa, Stripe, and Shopify understood that merchants would not integrate agent-accessible infrastructure if it meant ceding the customer relationship to the platform. So the architecture was built to let agents transact on behalf of consumers without the merchant losing record ownership. The brand stays the brand. The platform provides the rail. This reassurance is technically accurate and strategically incomplete. Merchants who have taken the merchant of record language as the end of the analysis have answered the wrong question. ### What merchants actually got The merchant of record designation is not ceremonial. It carries real and specific protections. Legal accountability stays with the brand. The customer's money flows to the merchant's payment infrastructure, not through the platform as an intermediary taking a cut before passing the remainder. The purchase confirmation carries the merchant's name. Returns, disputes, chargebacks, and warranty claims route through the merchant's systems. Fraud liability, tax collection, and consumer protection obligations all remain on the merchant's side of the transaction. These protections matter. The history of platform commerce is full of examples where merchants who ceded merchant of record status found themselves progressively squeezed on margins, on data access, on customer relationship ownership as the platform gained leverage. The UCP and AP2 architectures were explicitly designed to avoid that outcome. The merchant record protection is real and worth preserving. But preserving it requires understanding what it does and does not cover. ### What merchants did not get The merchant of record is not the merchant of decision. In traditional e-commerce, the merchant controlled most of the purchase journey. Traffic arrived at the merchant's site. The merchant's search, navigation, and recommendation systems shaped what the consumer saw. The merchant's product pages, content, and trust signals influenced evaluation. The merchant's checkout design and promotional mechanics influenced conversion. Record ownership and decision influence were bundled together. In agent-mediated commerce, they unbundle. When a consumer configures a shopping agent with a price ceiling, preferred brands, and substitution rules, the discovery, comparison, price monitoring, compatibility checking, and purchase authorization all happen before any transaction fires inside the agent's evaluation logic, on the platform's infrastructure, before the merchant's site is ever visited. By the time the merchant's checkout processes the transaction, the decision was made elsewhere. The merchant executed it. They did not make it. The merchant of record stays the merchant of record. The merchant of decision becomes whoever built the agent, designed the evaluation criteria, and owns the surface where the consumer's intent was formed and acted on. This is [Roger Dunn](https://www.linkedin.com/in/rogerdunn/?ref=agenticlandmark.com)'s observation from earlier this year: Google may be the merchant of decision for Universal Cart transactions while Gap, Nike, and Sephora remain the merchants of record. Both things can be simultaneously true, and they are not the same thing. ### The loyalty and data problem The strongest argument for maintaining merchant-controlled checkout is not the record ownership argument. It is the data and loyalty argument, and it is undernamed in most agentic commerce discussions. Every transaction that flows through a merchant's own checkout generates signals that compound over time. Behavioral data tied to specific products, occasions, and purchase contexts. Loyalty program redemption events that reinforce the customer relationship. Post-purchase signals reviews, returns, repeat purchases that feed personalization and retention models. Email capture. First-party cookies. The raw material of lifetime customer value. AP2 mandates route through Google Wallet. UCP transactions complete through the platform's payment infrastructure. When a transaction closes inside Google's Universal Cart, the merchant receives the order. They may not receive the full signal set that the same transaction would have generated inside their own checkout. Every agent-mediated transaction that bypasses the merchant's checkout is a transaction where the data loop either does not close or closes more thinly. At low agent transaction volume, this gap is immaterial. As agent-mediated transactions grow toward the percentages that McKinsey and BCG are projecting, it becomes a compounding disadvantage. The merchants who built rich first-party data infrastructure from their own checkout are increasingly competing on signals that merchants whose transactions route through platform checkout infrastructure cannot accumulate at the same rate. This is not a reason to reject UCP or AP2 integration. It is a reason to build loyalty and data capture into the agent-accessible checkout architecture from the start to ensure that the signals generated by agent transactions are as rich as those generated by human transactions, regardless of which checkout surface the transaction completed on. ### The agent accessibility problem that most brands are missing Maintaining a merchant-controlled checkout does not automatically make it agent-accessible. These are not the same investment, and most brands are treating them as if they were. A checkout that requires browser navigation, clicking through a cart, filling a form, completing a payment page cannot be executed by an agent operating on a consumer's behalf without browser automation. A checkout that requires a logged-in user session to apply loyalty pricing is a checkout that an agent without cached session state cannot execute correctly. A checkout that calculates shipping in real time through a visual interface rather than through an API is a checkout that an agent cannot query before confirming a transaction. The merchant of record status is preserved. The agent simply cannot use the checkout. This is the most expensive misunderstanding in agentic commerce infrastructure planning right now. Brands are investing in protocol compliance, registering for UCP, integrating with AP2, and assuming that their existing checkout infrastructure is capable of executing the transactions that those protocols will route to them. It is not, in most cases. Protocol compliance is the agreement to accept agent-initiated transactions. Checkout agent accessibility is the infrastructure that actually executes them. The build required to make a checkout agent-accessible is specific: API-first checkout with endpoints that accept agent-initiated transaction parameters directly, without browser navigation. Real-time pricing and inventory APIs that an agent can query before confirming a purchase. Loyalty integration at the API layer so that loyalty pricing, earned rewards, and program benefits are accessible to agents operating on behalf of enrolled members. Protocol-specific authentication that proves agent authorization without requiring a live user session. None of this is the same build as maintaining a best-in-class human checkout. The human checkout optimizes for visual conversion trust signals, progress indicators, payment method variety, and abandonment recovery. The agent checkout optimizes for programmatic reliability, API response speed, error handling, state management, and transaction confirmation. A brand can have an excellent human checkout and a completely inaccessible agent checkout simultaneously. ### The three-build model The merchants who navigate the next three years correctly are building three capabilities in parallel. The human checkout for the Curator posture majority. These are the consumers who are still browsing, comparing, and deciding for themselves before they reach checkout. They represent the majority of the current market. They need the full human checkout experience, product pages, trust signals, seamless payment, and post-purchase confirmation. Underinvesting here because the agentic future is coming is a mistake. These consumers are the present revenue. The agent-accessible checkout for the Scoper and Delegator posture consumers who have delegated execution to an agent. This checkout does not need to be visible or beautiful. It needs to be reliable, API-first, protocol-compliant, and capable of executing agent-initiated transactions with the same accuracy and completeness that human-initiated transactions achieve. Underinvesting here because it serves a small current minority is also a mistake. These consumers are the future revenue with the highest conversion probability. The loyalty and data architecture that captures signals from both transaction paths. This is the layer that most brands are not thinking about yet. As agent transactions grow, first-party signal capture must extend into the agent transaction path. Loyalty redemption, behavioral signals, and post-purchase data need to flow from agent-initiated transactions as richly as they flow from human-initiated transactions. The brands that build this architecture now are the ones whose personalization and retention models will be richer in three years than competitors who let agent transactions flow through platform infrastructure without capturing the signals. ### What the merchant of record actually means The merchant of record designation is valuable. It is worth maintaining. The protocols that preserve it were designed with merchant interests in mind, and the protection is genuine. But merchants who have read the merchant of record reassurance and concluded that the customer relationship is secure have answered the wrong question. Record ownership secures the transaction. Decision influence secures the relationship. In the era of agent-mediated commerce, the two are no longer bundled together by default. The era of agentic commerce does not make merchant-controlled checkout less important. It makes the definition of controlled more specific: controlled for whom, accessible by what, capturing which signals, serving which checkout audiences simultaneously. The merchants who answer those questions in their infrastructure planning now are the ones the next version of Google's launch list will include. ### The Curator is still the majority. URL: https://www.agenticlandmark.com/brief/the-curator-is-still-the-majority/ Last updated: 2026-07-01T20:12:10.000Z DuckDuckGo just recorded its highest single-day search traffic ever. US iPhone installs are running 95% above pre-Google I/O levels. Almost double. Most coverage is treating this as an anti-AI story. It is not. DuckDuckGo is not anti-AI. It simply offers what Google no longer does: the option to search without AI involvement. The [noai.duckduckgo.com](http://noai.duckduckgo.com/?ref=agenticlandmark.com) URL lets users disable AI features entirely. Google removed that choice. Consumers switching to DuckDuckGo after I/O are not rejecting the web or search. They are asserting something specific: they want to remain in the selection loop. They want to browse, compare, and decide for themselves. They are not ready to delegate discovery to an AI that curates before they ever see the options. In the framework I use for agentic commerce readiness, this is the Curator posture behavior. These consumers research with their own eyes and make their own calls. They are the majority of the market right now, and they are signaling clearly that the speed of Google's transition toward AI-mediated discovery is faster than their trust in AI-mediated discovery. The DuckDuckGo data sits alongside the Google I/O announcements and makes the same point that the BCG four-scenarios framework made in March: the transition to agent-mediated commerce is not uniform. Different consumers are moving at different speeds, and some are actively moving backward. Brands building only for the consumer who wants to delegate are optimizing for the part of the market that is not yet there at scale. The Curator is still the majority. They still need to be served. [DuckDuckGo No-AI: Private Search Without AIDuckDuckGo private search with all AI features turned off. The classic results you expect. Always private.![](https://storage.ghost.io/c/e1/c4/e1c4b063-1fa4-4977-8ffa-6add50cd2316/content/images/icon/DDG-iOS-icon_152x152-24b28c6d-9c3e-47c9-b059-597f4f66db41.png)DuckDuckGo![](https://storage.ghost.io/c/e1/c4/e1c4b063-1fa4-4977-8ffa-6add50cd2316/content/images/thumbnail/logo_social-media-bb4ee5bb-f48b-4a47-82cf-db6a09a86fe5.png)](https://noai.duckduckgo.com/?ref=agenticlandmark.com) ### Salesforce did not buy a CMS. It bought a content layer for agents. URL: https://www.agenticlandmark.com/brief/salesforce-did-not-buy-a-cms-it-bought-a-content-layer-for-agents/ Last updated: 2026-07-01T20:01:36.000Z Salesforce did not acquire a CMS yesterday. They acquired the content layer that agents need to assemble personalized experiences in real time. The official framing is "Headless 360," connecting customer data, Agentforce, and Contentful's composable APIs to deliver AI-assembled experiences at scale. The strategic logic runs deeper than that. Contentful's co-founder Sascha Konietzke framed it in his announcement post: "AI agents now outnumber humans on the Web, forcing companies to rethink how digital experiences are created, optimized, and deployed." That sentence is the acquisition rationale. If agents now outnumber humans as content consumers, the content layer needs to be machine-readable, API-first, and structurally composable by design. Contentful was built from the start as a headless, API-first content infrastructure. That architecture, originally designed to serve mobile apps that couldn't render CMS templates, turns out to be exactly the architecture that agents can query and assemble from. Salesforce had the data layer. They had the AI execution layer. The missing piece was a content layer that agents could actually work with. Not a CMS that generates pages for humans to read. A content infrastructure that stores structured, channel-agnostic content objects that an agent can retrieve, assemble, and deliver based on context, channel, language, and business rules without a human authoring a new page for every variation. This is the same acquisition logic as Anthropic buying Stainless. Own the layer that everything else connects through. Salesforce is buying content addressability, the ability for Agentforce to reach into a structured content store and assemble the right experience at machine speed. The implication for brands running Salesforce stacks: the enterprise content in Contentful is now in scope for agent assembly, not just human publishing. The governance and tagging discipline that was good practice before is now the infrastructure that determines whether Agentforce can actually use your content. Konietzke built a CMS for the mobile era and inadvertently built the content infrastructure for the agent era. Salesforce paid for that foresight. ### The Designer's New Audience Problem URL: https://www.agenticlandmark.com/brief/the-designers-new-audience-problem/ Last updated: 2026-07-01T03:33:47.000Z *For thirty years, UX research assumed one audience. The agent era requires designing for two simultaneously, and the harder problem is the one nobody is writing patterns for yet.* The field has a name for it now. Fuselab Creative calls it Agent UX. Peterson Technology Partners calls it AX Agentic Experience Design. A paper presented at CHI 2026 in Barcelona formalized it as the "dual-audience interface" problem: the Human-Agent-UI triangle, in which every interface must simultaneously serve the human who sets intent and the agent that executes it. The framing is correct. The coverage is incomplete. Most of the emerging literature on designing for human and agent audiences focuses on the same half of the problem: how to make the agent's actions visible and interpretable to the human watching it work. Transparency layers. Reasoning traces. Override controls. Confidence indicators. Error recovery patterns. This is valuable design work, and it is urgently needed. But there is a second design problem that precedes all of it, and almost no one is writing patterns for it yet. ### The producer-side problem vs. the consumer-side problem The dual-audience interface literature is almost entirely focused on what you might call the producer side: organizations deploying AI agents inside their own products to help users accomplish tasks. The Fuselab piece describes a clinical AI interface where clinicians refused to use a recommendation system that surfaced suggestions without explaining its reasoning. NNGroup's State of UX 2026 identifies trust as the defining design challenge. The CHI 2026 paper proposes representational compatibility, transparency, and low barriers as the three design properties that matter most. All of this is describing the same moment: a human is watching an agent do something on their behalf, and the interface must make that action legible, controllable, and trustworthy. That is not the hardest design problem in the agent era. The hardest design problem comes before the agent does anything at all. ### The grant screen: the most consequential undesigned surface When a consumer configures an AI shopping agent setting a price ceiling, selecting preferred brands, defining substitution rules, and establishing exception conditions, they are creating something that has no real UX precedent: a purchase mandate. A structured, machine-executable declaration of intent that will govern the agent's behavior across every transaction within its scope, potentially for months. Google's Agent Payments Protocol shipped this summer. AP2 lets a consumer delegate purchases inside bounded conditions, "buy these Nike Pegasus 41s in size 10 if they hit $90," using cryptographically signed digital mandates with a verifiable audit trail. The protocol infrastructure exists. The UX that makes this delegation legible, configurable, and trustworthy for mainstream consumers has barely been designed. This is the grant screen problem. And it is structurally different from the dual-audience interface problem most designers are currently working on. The transparency layer and override controls that Fuselab and others are designing help a human understand and correct what an agent is doing. The grant screen determines what the agent is allowed to do in the first place. One is about monitoring execution. The other is about configuring authority. These are different design problems with different success criteria, different failure modes, and different consequences when they go wrong. ### Why the grant screen is harder than it looks Traditional UX research has a well-developed vocabulary for preference declaration. Settings screens. Preference panels. Onboarding flows. These are established patterns for collecting user input that shapes system behavior. The grant screen is not a settings screen. It is an authorization instrument. The user is not setting preferences that the system will try to honor. They are granting authority that an autonomous agent will act on, potentially executing transactions, committing funds, and making decisions in the user's name without per-transaction confirmation. The design implications are specific: **Scope legibility.** The user must understand what they are granting authority over. Not in legal terms, in experiential terms. "Price ceiling: $150" is clear. "Substitution rules: allow equivalent items from approved brands" is not clear to most consumers. The gap between what the user believes they authorized and what the agent believes it is authorized to do is the primary source of trust failures in agent-mediated commerce. Closing that gap is a language design problem, an interaction design problem, and a mental model problem simultaneously. **Reversibility and editability.** Mandates must feel provisional, not permanent. A consumer who believes their agent authority is difficult to revoke will either not grant it at all or grant it without reading what they are agreeing to. Neither outcome serves them. The grant screen must make the mandate feel like a standing instruction they are in control of, not a contract they are signing. **Exception design.** The most consequential design decision in the grant screen is not what the agent is allowed to do. It is what happens at the edges of that authority. When the agent encounters a situation outside its declared parameters, the preferred brand is out of stock, the price dropped below the floor, the specifications changed, what does it surface, and how? Exception handling design determines whether agent delegation feels trustworthy or unpredictable. **Trust calibration over time.** The consumer who delegates for the first time should have a different experience from the consumer who has delegated successfully for six months. First-delegation flows need more friction, more explicitness, more confirmation. Established-delegation flows should reduce that friction deliberately. Most grant screen design treats all consumers as first-time delegators, which is the right default and the wrong long-term model. ### What the field needs The producer-side dual-audience interface problem is getting serious design attention. The consumer-side authorization design problem is not, and the infrastructure that requires it is already in production. The patterns that need to be developed: The mandate declaration interface how a consumer expresses bounded authority in language and interaction models that are both machine-parseable and humanly understandable. The scope visualizes how a user sees what they have authorized across multiple active mandates, across multiple categories, across multiple agents operating on their behalf. The exception alert design how an agent surfaces out-of-bounds situations in a way that gets timely human review without creating alert fatigue. The trust progression model how delegation experiences should evolve as the consumer builds a track record with a specific agent or brand. These patterns do not yet exist at the level of specificity that interaction design requires. The CHI 2026 paper on dual-audience interfaces is an academic starting point. The Fuselab and Peterson Technology work addresses the monitoring side. The authorization side is the gap. For UX researchers and interaction designers: the grant screen is the most important new design surface in commerce, and the one with the least published guidance. The work that needs to happen here is foundational, defining the vocabulary, the patterns, and the failure modes of consumer-facing delegation interfaces before the infrastructure scales past the point where the design can catch up. The transparency layer problem has designers working on it. The authorization design problem is still waiting for them. ### Google says you don't need llms.txt. Google is also checking for it. URL: https://www.agenticlandmark.com/brief/google-says-you-dont-need-llms-txt-google-is-also-checking-for-it/ Last updated: 2026-07-01T03:18:30.000Z Google says you don't need llms.txt. Google is also checking for it. Google says you do not need llms.txt. Google also just added an llms.txt check to Lighthouse. Both statements are true. They are not contradictory. They are describing two different audiences. Google Search's AI optimization guide, published earlier this month, placed llms.txt in the mythbusting section. You do not need it for AI Overviews or AI Mode. That guidance is accurate for the Google Search use case. AI Overviews and AI Mode pull from Google's existing search index. If your pages are crawlable, indexable, and well-structured, you are in the pool. The file adds nothing to that process. Lighthouse 13.3 shipped a new Agentic Browsing category that checks for llms.txt alongside WebMCP support, layout stability, and accessibility structure. This audit does not measure search visibility. It measures whether your site is ready for software agents operating through Chrome agents that need to understand your site's structure, purpose, and primary content without crawling every page to figure it out. Google's own documentation explains the purpose: without llms.txt, agents may spend more time crawling the site to understand its high-level structure. Two products. Two audiences. Two different readiness requirements. This is the same distinction that runs through every agentic commerce infrastructure conversation. Google Search AI features query the index. Browser-based agents interact with the site directly. An agent completing a task on a consumer's behalf through Chrome needs to navigate, understand, and execute, not just retrieve a cited paragraph. The llms.txt file is the site's way of telling that agent where to start. The practical implication: if your team is using "Google said we don't need it" to close the llms.txt conversation, show them the Lighthouse audit. The question is not whether you need it for Search. The question is whether browser agents can navigate your site efficiently when they arrive. Those are not the same question. Google is now measuring both. ### Strategy over urgency is right. So is the calendar. URL: https://www.agenticlandmark.com/brief/strategy-over-urgency-is-right-so-is-the-calendar/ Last updated: 2026-07-01T03:10:49.000Z Strategy over urgency is right. So is the calendar. Forrester published their State of Agentic Commerce framework yesterday. Emily Pfeiffer's headline is "strategy over urgency." She is right. She is also describing a distinction that is harder to act on than it sounds. The framework is accurate on the current state. Most agentic experiences are still conversational. True autonomy is rare. Consumer trust and adoption are uneven. Hype is running well ahead of behavior. All of that is true today. The tension is in the timing. Forrester's recommendation is to prepare now for the moment when agentic commerce becomes truly viable. That framing implies the moment is ahead of us. For some brands, it is. For others, the moment arrived this month. Google's Universal Cart is rolling out this summer with Nike, Sephora, Target, Ulta, Walmart, Wayfair, and a selection of Shopify merchants. Those are not pilot participants. That is the launch list. UCP is expanding to Canada, Australia, and the UK within months, and into hotel booking and food delivery shortly after. AP2 is shipping to Gemini Spark users this summer. The infrastructure that makes a brand visible, selectable, and transactable inside those surfaces takes 12 to 18 months to build properly. Structured product data. Schema completeness. API-accessible checkout. Loyalty integration. Protocol compliance. If a brand starts that infrastructure work when agentic commerce feels urgent, when their pipeline is shrinking, and their competitors are showing up in agent recommendations, they are 12 to 18 months behind a transition that is already in production. Forrester is right that the answer is not panic. The answer is not waiting either. The brands on Google's Universal Cart launch list did not get there because they responded urgently to a summer 2026 announcement. They got there because they made infrastructure decisions in 2024 and 2025. Strategy over urgency is the correct principle. The strategic decision is understanding that the infrastructure timeline and the urgency timeline are the same timeline. They just do not feel the same until one of them has already closed. The Forrester framework is worth reading. [The State Of Agentic Commerce In Mid-2026digital leaders cannot afford to ignore digital commerce. Forrester’s Agentic Commerce Framework can guide your strategy.![](https://storage.ghost.io/c/e1/c4/e1c4b063-1fa4-4977-8ffa-6add50cd2316/content/images/icon/cropped-icon_512.c2nfSw_xZ0-300x300-c259f05d-9baf-4b88-b76d-304e958d2000.png)ForresterEmily Pfeiffer![](https://storage.ghost.io/c/e1/c4/e1c4b063-1fa4-4977-8ffa-6add50cd2316/content/images/thumbnail/Strategy-over-Urgency-Graphic-d9be5fb1-f5f5-47f4-b877-bc49c1664c9e.png)](https://www.forrester.com/blogs/the-state-of-agentic-commerce-in-mid-2026/?ref=agenticlandmark.com) ### Walmart built its agent. Amazon will rent you one. The difference is the data layer. URL: https://www.agenticlandmark.com/brief/walmart-built-its-agent-amazon-will-rent-you-one-the-difference-is-the-data-layer/ Last updated: 2026-07-01T03:01:16.000Z Amazon just packaged everything it learned building Alexa for Shopping and started selling it to retailers. AWS Agentic Shopping Assistant gives retailers the architecture, starter code, and expert guidance behind Amazon's own AI shopping infrastructure. Kate Spade is already live with it. Built on Anthropic's Haiku model through Amazon Bedrock. Retailers can deploy in roughly 60 days. The business case Amazon is making: their AI shopping assistant drove nearly $12 billion in incremental sales for Amazon last year. Conversational sessions convert at 3.5 times the rate of keyword search. If that's what the technology did for Amazon, here's the infrastructure to do it for your brand. That is a compelling offer. It is also worth reading carefully. Walmart built Sparky on their own data, their own customer relationships, their own loyalty infrastructure. That is why Sparky converts better than any general-purpose agent; it knows more because Walmart built the data layer rather than renting it. The 35% higher spend per order is a data ownership story. The AWS ASA offer is different. You are building your brand-owned agent on Amazon's infrastructure, using Amazon's models, trained on Amazon's shopping data. The "you keep the customer relationship" reassurance is technically accurate. The dependency it creates is real. This is the same tension in Google's Universal Cart reassurance: "you stay the merchant of record." True. But when discovery, cart management, and checkout authorization all live in someone else's infrastructure, the customer relationship is on your brand label and their nervous system. Every major platform is now offering retailers the same deal: here is the agent intelligence we built for ourselves, available to you, on our infrastructure. Amazon. Google. Microsoft. The retailers who own the next decade will be the ones who read that offer clearly and decide which layer they are willing to rent and which layer they need to own. ### The Pope named a governance problem. Agentic commerce is the micro version. URL: https://www.agenticlandmark.com/brief/the-pope-named-a-governance-problem-agentic-commerce-is-the-micro-version/ Last updated: 2026-07-01T02:51:14.000Z The most remarkable thing said about AI this week did not come from a CEO on a stage at a tech conference. It came from Chris Olah, co-founder of Anthropic, standing at the Vatican. "Every frontier AI lab, including Anthropic, operates inside a set of incentives and constraints that can sometimes conflict with doing the right thing. "He said this while welcoming Pope Leo XIV's first encyclical, "Magnifica Humanitas," which called for robust external regulation of AI and named the concentration of power and data in private sector hands as a danger to human dignity. Olah not only accepted the criticism. He asked for more of it. That is a different register than almost anything said publicly by anyone at a frontier AI company. It is also the most honest framing of the governance problem I have read: the people building the technology are inside a set of incentives that may not align with the outcomes everyone else needs. External checks are not an attack on innovation. They are a structural necessity. The same governance question has a very specific form in commerce. When a consumer delegates a purchase decision to an AI agent, setting parameters, granting authority, and handing off execution, they are extending trust to a system operating inside exactly those incentives Olah described. The authorization architecture that determines what an agent can do, on whose behalf, within what limits, and with what accountability is not a technical detail. It is the governance layer that makes delegation safe. The Pope named a macro problem. The commerce infrastructure question is the micro version of the same thing. Both are asking: who is accountable when the agent acts, and what prevents it from acting in ways that serve the platform rather than the person? That question does not have a complete answer yet. What Olah did this week is make it harder to pretend it does not need one. ### AI Native vs. Agent Native: The Distinction Most Organizations Are Missing URL: https://www.agenticlandmark.com/brief/ai-native-vs-agent-native-the-distinction-most-organizations-are-missing/ Last updated: 2026-06-30T22:12:06.000Z *They sound similar. They describe fundamentally different things. One makes your platform smarter. The other makes it addressable.* The term "AI native" has become the default label for any digital platform that has embedded generative AI into its core experience. The AI helps users draft content, summarize data, flag anomalies, and make decisions faster. This is genuinely useful. It is also not what the next three years require. There is a second category that most organizations have not named clearly: agent native. And the failure to distinguish between the two is producing a generation of platforms that will be smart and useless at exactly the wrong moment. ### What AI Native actually means An AI native platform is one where artificial intelligence augments human work. The AI generates, summarizes, assists, and recommends. The human reviews, decides, and acts. The interface is designed for human perception: chat sidebars, conversational prompts, highlighted suggestions, and confidence scores displayed in dashboards that humans read. The design paradigm is: how can AI help humans do tasks faster and better? This is a feature decision. You are adding capability to an existing system whose fundamental architecture, screens, buttons, and human-initiated workflows were designed for human users. The AI layer sits on top of that architecture and improves it. The architecture itself does not change. Most enterprise software platforms are AI native today, or claim to be. They have added AI features. Their underlying architecture remains human-first. ### What Agent Native actually means An agent native platform is one designed from the start or redesigned from the core to be addressable by autonomous agents operating on behalf of humans. The design paradigm shifts entirely: how do we build systems that autonomous agents can navigate, plan, and execute tasks within? This is an architecture decision. It requires that every capability in the platform has an API surface, a named permission structure, and event emission. Not so that AI can help a human use the capability. So that an agent can call the capability directly, with delegated authority, without a human navigating the interface at all. The interface implication is specific: an agent native platform has a dual interface. There is a human UI, screens, buttons, and visual workflows for the humans who still use the system directly. And there is an agent API layer, tools, protocols, and structured data endpoints that expose the exact same system capabilities to agents operating autonomously. Not a separate system. The same system is accessible two ways. An AI native platform asks: what AI features should we add to help our users? An agent native platform asks: Does every capability we build have an API surface, a named permission, and event emission? ### The commerce version of this distinction In commerce, the AI native platform is the one where your product recommendation engine got smarter, your customer service chatbot got better, and your marketing team can now generate product descriptions in seconds. These are real improvements. They are also improvements to a human-facing system. The agent native question is different. When a consumer configures a shopping agent with a price ceiling, preferred brands, and substitution rules, that agent needs to call your checkout API, verify inventory in real time, apply loyalty pricing, and complete a transaction without a human navigating your website. If your checkout requires browser navigation, your checkout is not agent native. If your inventory data lives only in a human-readable HTML product page, your inventory is not agent native. If your loyalty program requires a logged-in user session to apply perks, your loyalty program is not agent native. The AI native platform made your website better for the human visitors who browse it. The agent native platform makes your commerce infrastructure addressable by the agents acting on behalf of consumers who never visit at all. Both matter. They are not the same investment. ### Why the distinction matters right now Most organizations currently building AI capabilities are making feature decisions when they believe they are making architecture decisions. They are adding AI to help humans use the platform better. They are not redesigning the platform to be addressable by agents. The practical cost of this confusion is timing. The architecture decision whether every capability has an API surface, named permission, and event emission is being made right now, on every feature in the roadmap. It is not a decision that can be retrofitted cheaply once the feature ships in its human-first form. A capability designed for human navigation can be wrapped in an API later, but the permission model, the state management, the error handling, and the event structure that makes agent-reliable execution possible are significantly more expensive to add after the fact than to build in from the start. The brands and platforms that will be agent native in 2027 are the ones making agent native architecture decisions in their product planning today. Not by building for agents instead of humans, the dual interface model means both audiences are served. But by requiring, at the design stage, that every capability be exposed through an API surface that an agent can call, not just a screen a human can click. ### The question that separates the two AI native platforms ask: What AI features should we add? Agent native platforms ask: Can an autonomous agent, with delegated authority from the consumer, navigate and execute every meaningful workflow in this system without a human in the loop? If the answer to the second question is no, if there are workflows that require human navigation, human session state, or human confirmation at every step, the platform is AI native. It may be very good. It is not agent native. The first question is a feature roadmap question. The second is an architecture question. They have different owners, different timelines, and different costs. The confusion between them is producing platforms that will be impressively smart and structurally inaccessible at exactly the moment the agents arrive. That moment, for commerce specifically, is not 2028\. It is now. ### The day the search box learned to build URL: https://www.agenticlandmark.com/brief/the-day-the-search-box-learned-to-build/ Last updated: 2026-06-29T22:41:26.000Z *At Google I/O 2026, Search stopped returning links and started assembling interfaces. That one change quietly rewrites what a decade of search optimization was actually for.* The most consequential thing Google announced at I/O was not the billion users. It was a feature most of the coverage skipped, because it does not look like a commerce story or a consumer story. It is a search story, and it changes the job of everyone who optimizes for search. The headlines went where you would expect. AI Mode crossed one billion monthly users. The Universal Cart assembles a shopper's selections across merchants. New agents book restaurants and place phone calls on a consumer's behalf. Each of these is real, and each deserves attention. But underneath them sat a quieter announcement called generative UI. In response to a single query, Search can now build a custom interface on the fly. Not ten blue links. Not even a cited paragraph of generated text. An actual comparison table, an interactive calculator, a visual tool, a simulation, a persistent dashboard, Google calls a mini app, assembled in real time for that one specific question. Read that again with your own work in mind, because if any part of your job involves making a brand visible in search, the ground just moved. ### What optimization is rewarded, and what it rewards now For roughly a decade, search optimization has rewarded content that the engine could read and cite. You wrote authoritative, well-structured, statistic-rich prose. You earned a ranking and later a citation in an AI Overview or a Perplexity answer. The lever was content quality. Generative Engine Optimization, the discipline that grew up around AI search, refined that lever for a generative world, but it did not change the underlying premise. The premise was always: write the thing well enough that the engine chooses to surface it. Generative UI rewards something different, and the difference is not a refinement. It is a category change. When the engine assembles a comparison table, it does not lift sentences from your buying guide. It populates cells. It needs a value for price, a value for material, a value for warranty length, and a boolean for whether the thing ships free. It pulls those values from structured, machine-readable data, not from the paragraph where you described them beautifully. Your prose can be the best on the internet, and your brand can still be absent from the table, because the table is not made of prose. It is made of fields that the engine can parse, compare, and render. > You can write the best buying guide on the internet and still be missing from the comparison table the engine builds. The table is not made of prose. It is made of data. This is the part worth sitting with. A brand can do everything GEO has asked of it for the last three years, win the citation, appear in the answer, get named as a source, and still not appear in the interactive tool the engine generates to actually help the user decide. The citation and the tool are now two different surfaces, drawing from two different layers of the brand's presence. ### The shift, stated plainly GEO is moving from citation-worthiness to structured-data-worthiness. From a content-marketing problem to a data-architecture problem. That sentence is short on purpose because it is the whole argument. The work of being found used to live in the editorial function: writers, content strategists, and the people who know how to make a page authoritative. The work of being usable now lives partly in the data function: the product information management system, the schema, the feed, and the people who know how to make a fact machine-readable. Generative UI does not replace the first kind of work. It adds the second kind as a hard requirement, and it puts the two on the same critical path. > Adobe's Q2 2026 AI Traffic Report found that the average US retail product page is 66% machine-readable to large language models. The best-performing sites reach 82.5%. The lowest sits at 54.2%. The gap between a brand that shows up in the generated tool and one that does not is largely the gap between those numbers. The Adobe benchmark matters here because it quantifies how unprepared most brands are for exactly this shift. A product page that is two-thirds machine-readable was good enough when the engine only needed to cite the page. It is not good enough when the engine needs to extract structured values to build a comparison that the user sees instead of the page. ### Why is this uncomfortable for how teams are built? In most organizations, the content team and the structured-data team are different functions. Often, different budgets. Sometimes, different departments meet twice a year and otherwise leave each other alone. Content reports to marketing. Product data reports to e-commerce operations or merchandising or, in the worst cases, to nobody in particular and lives in a spreadsheet someone updates when they remember. Generative UI collapses that separation, and it does so without asking permission. The brand that appears in the generated comparison is the one whose content and structured data are a single, coordinated investment. The brand that keeps treating them as two sequential projects, content this quarter, data architecture maybe next year, will keep earning citations while quietly disappearing from the tools those citations were supposed to lead to. There is a version of this problem that is even more specific. Google built generative UI on Gemini 3.5 Flash and its Antigravity development platform, which means the interfaces are generated by a coding agent in real time. The engine is not selecting a pre-built template and slotting your data in. It involves reasoning about what interface the question deserves and constructing it. The brands whose data is clean, complete, and consistently structured are the raw material that agents reach for. The brands whose data is partial or inconsistent are simply harder to build with, and an agent optimizing for a good answer will route around difficulty the same way water routes around a rock. ### What this does and does not change It does not make content worthless. Authority, accuracy, and citation-worthiness still determine whether the engine trusts a brand as a source, and trust is upstream of everything. A brand with no credible content presence will not be pulled into a generated tool, no matter how clean its feed is, because the engine has no reason to believe its numbers. Content earns the right to be considered. That has not changed. What changed is that content is no longer sufficient. The brands that win the generated interface are the ones that pair editorial authority with structured-data discipline, and treat the two as one capability rather than two departments. The discipline formerly known as GEO is converging, in real time, with the structured-data work that agentic commerce has required all along. I/O did not invent that convergence. It made it visible, and it put a one-billion-user surface behind it. For a decade, the question was whether your content was good enough to be *cited*. The new question is whether your data is clean enough to be *built with*. ### Same Icon. Opposite Job. Google Just Inverted the Shopping Cart. URL: https://www.agenticlandmark.com/brief/same-icon-opposite-job-google-just-inverted-the-shopping-cart/ Last updated: 2026-06-29T22:33:33.000Z A while back, I argued the shopping cart icon was a thirty-year-old skeuomorph living on borrowed time. A digital artifact that mimics a physical behavior, pushing a cart through a store, for no reason other than familiarity. We kept it because shoppers in 1995 needed a familiar metaphor. Then we spent three decades optimizing a process built around it: browse, add, check out, abandon. Seventy percent of the time, abandon. This week at I/O, Google replaced it. They introduced the Universal Cart, a cart that assembles your purchases from multiple merchants and is managed by an AI agent on your behalf. AI Mode, the surface it runs on, just crossed one billion monthly users. Here's the part that matters. For thirty years, the cart was something you filled. One store at a time. The Universal Cart inverts that completely. You don't fill it; the agent assembles it. Not from one merchant, but across every merchant that can answer the agent's questions. There's no browsing phase to abandon. The cart is built from a brief, not from a session. ### Anthropic Did Not Buy a Developer-Tools Company. It Bought the Connectivity Layer. URL: https://www.agenticlandmark.com/brief/anthropic-did-not-buy-a-developer-tools-company-it-bought-the-connectivity-layer/ Last updated: 2026-06-29T22:22:38.000Z Anthropic acquired Stainless yesterday. Most coverage treats it as a developer tools story. It's a connectivity infrastructure story. Stainless builds the SDKs and MCP server tooling that allow developers to connect Claude to external systems. They have been generating every official Anthropic SDK since the beginning. The acquisition brings that capability in-house permanently. Why this matters: Anthropic created the Model Context Protocol to make agent connectivity possible. MCP is the standard that allows AI agents to connect to APIs, databases, and external services to reach outside themselves and act in the world. The protocol is only as useful as how easily developers can build on it. Stainless turns an API spec into native SDKs across TypeScript, Python, Go, Java, Kotlin, and more. It generates the MCP servers that connect Claude to those systems. In practical terms, the harder it is to build an MCP connection, the fewer systems Claude agents can reach. The easier it is, the more capable every Claude-based agent becomes. Anthropic's own framing: "Agents are only as useful as what they can connect to." That sentence is the acquisition rationale. They are not buying a developer tools company. They are vertically integrating the layer that determines how far their agents can reach and how easily the brands and developers building on Claude can extend that reach to their own systems. For brands building agentic commerce infrastructure: your product catalog, your inventory APIs, your loyalty data, and your checkout endpoints are the systems MCP servers connect Claude agents to. The Stainless acquisition means Anthropic is directly invested in making that connectivity as frictionless as possible. This is the same strategic logic as Visa's ICC play: own the integration layer, and value flows to you regardless of what connects through it. [https://www.anthropic.com/news/anthropic-acquires-stainless](https://www.anthropic.com/news/anthropic-acquires-stainless?ref=agenticlandmark.com) ### Google I/O 2026: The Infrastructure of Agentic Commerce Just Went Live URL: https://www.agenticlandmark.com/brief/google-i-o-2026-the-infrastructure-of-agentic-commerce-just-went-live/ Last updated: 2026-06-29T22:17:39.000Z *Today's announcements aren't previews. They're production deployments. Here's what they actually mean.* Google published two blog posts today that, read together, represent the most consequential agentic commerce announcements since the category was named. The first covers Search: AI Mode has surpassed one billion monthly users, with queries more than doubling every quarter. The Search box is getting its biggest upgrade in twenty-five years. Search agents are launching, information agents that monitor the web 24/7 on your behalf, agentic booking for local services, and a feature that lets Google call businesses on your behalf. The second covers Shopping: a new Universal Cart, an expansion of the Universal Commerce Protocol to new surfaces and geographies, and the launch of the Agent Payments Protocol. Most coverage will treat these as feature announcements. They are infrastructure deployments. Here is what actually matters. ### The Universal Cart is the grant screen made practical The Universal Cart is the first consumer-facing implementation of the delegation architecture that agentic commerce requires. It works across merchants and across surfaces, Search, Gemini, YouTube, Gmail as a single shopping context that follows the consumer rather than belonging to any one retailer. But the feature that deserves the most attention is not the cross-merchant cart. It is the Agent Payments Protocol layer underneath it. AP2 lets you set strict guardrails for agentic payment transactions. You declare the specific brands and products you want and the maximum the agent can spend. The agent only makes the purchase when your criteria are met. Under the hood, AP2 creates a transparent, verifiable link between you, the merchant, and the payment processor, using privacy-preserving technology to keep your data safe. Tamper-proof digital mandates ensure the agent is always acting on your behalf, giving you a permanent digital paper trail. This is the grant screen. The consumer-facing authorization surface that defines what authority the agent holds, what it can spend, which brands it can transact with, and what constitutes a mandate violation. This infrastructure is the most consequential unbuilt surface in agentic commerce, which was announced today as a production feature shipping to Gemini Spark in the coming months. The PC compatibility check example in the Universal Cart announcement is the clearest illustration of what this means in practice: if you're building your first custom PC and add a few parts from several retailers to your cart, your cart will proactively flag any product incompatibilities and suggest alternatives. That is an agent operating within a bounded mandate to buy parts for this PC build and exercising judgment within it. That is Scoper posture behavior running in production at Google scale. ### The merchant list is the readiness signal You can check out with Google Pay in just a few taps with many of your favorite brands, or transfer your items to the merchant's site to complete your purchase. You can try these select checkout features soon across merchants like Nike, Sephora, Target, Ulta Beauty, Walmart, Wayfair, and Shopify merchants such as Fenty and Steve Madden. Read that list carefully. These are the brands whose infrastructure was ready. UCP integration, Google Pay compatibility, merchant-of-record preservation, and real-time catalog connectivity. Every brand on that list has done the infrastructure work required to be in the Universal Cart at launch. Every brand not on that list has not. No matter which way you buy, the brand stays the merchant of record. That detail matters for everyone watching the disintermediation question: Google is not inserting itself as the merchant. It is providing the interface, the protocol, and the payments infrastructure. The brands that have integrated remain the commercial entity in the transaction. The brands that have not integrated are not in the transaction at all. ### UCP is expanding faster than anyone has published The Universal Commerce Protocol is not a US-only standard in pilot. We're expanding our UCP-powered checkout experience on Google to Canada and Australia in the coming months and later to the U.K. UCP is also coming to YouTube in the U.S., and to even more verticals, starting soon with hotel booking and local food delivery. Three things worth noting in that sentence. First, geographic expansion to three major English-speaking markets within months. Second, YouTube as a commerce surface UCP-powered checkout inside video content. Third, vertical expansion beyond retail into hospitality and food delivery, which are exactly the high-frequency transaction categories where agent mediation compounds fastest. The hotel booking expansion is the travel vertical signal. The brands building agent-ready infrastructure in travel loyalty data, inventory APIs, and real-time availability are now looking at a UCP-powered checkout surface launching on the world's most-used AI search platform. ### Search agents are the Scoper posture at the population scale The information agents announced today run continuously in the background, monitoring across the web for changes related to your specific question. If you're apartment hunting, you can brain dump all of the exact requirements you're looking for, and your agent will continuously scan for you, notifying you when listings meet your needs. That is a consumer declaring parameters, size, location, price, amenities, and delegating continuous monitoring to an agent operating within those parameters. That is the Scoper posture. The consumer sets the scope. The agent executes within it. Exceptions surface as notifications rather than for every result. The agentic calling feature takes it one step further. For select categories like home repair, beauty, or pet care, you can ask Google to call businesses on your behalf. A consumer who has delegated the call to Google has delegated a commerce action, inquiry, qualification, or potentially booking to an agent operating within their declared intent. The consumer never makes the call. The agent makes it and surfaces the result. These are not experimental features. They are rolling out to everyone in the U.S. this summer. ### What today means for brands AI Mode has one billion monthly users with queries doubling every quarter. The Universal Cart is launching with Nike, Sephora, Target, Ulta, Walmart, Wayfair, and a selection of Shopify merchants. UCP is expanding to Canada, Australia, and the UK. AP2 is the authorization layer that makes consumer-controlled agent purchasing verifiable and defensible. The infrastructure of agentic commerce is not being built. It is built. It is in production. It is rolling out to a billion users. The readiness question for brands is no longer theoretical. The Universal Cart merchant list published today is the first public ranking of which brands are inside the agentic commerce infrastructure and which are not. The ShopTalk announcements of March were previews. Today's I/O announcements are deployments. Three things every commerce brand needs to understand today: The grant screen exists. AP2 is the consumer-facing authorization layer for agentic purchasing. The brands that are surfaced within that authorization layer, the ones consumers include in their declared brand preferences and spending parameters, have a structural advantage that compounds with every agent transaction. The brands not in a consumer's declared parameters do not get considered, regardless of their product quality. The merchant-of-record model survived. Google is not the merchant. The brand remains the commercial entity. The brands that integrated with UCP and Google Pay preserved their customer relationships. The brands waiting for clarity on who owns the transaction no longer have that ambiguity. Today's announcement settled it. The timeline moved. Information agents launching this summer. Universal Cart launching this summer. AP2 is coming to Gemini Spark in the coming months. The window between "interesting development to monitor" and "live infrastructure your customers are using" just closed for a significant portion of the market. ### Google Says You Are Overthinking AI Search. For Google Search, Yes. For Agent Commerce, No. URL: https://www.agenticlandmark.com/brief/google-says-you-are-overthinking-ai-search-for-google-search-yes-for-agent-commerce-no/ Last updated: 2026-06-29T22:09:44.000Z Optimizing your website for generative AI features on Google Search. Google just published official guidance on optimizing for generative AI search. The mythbusting section is worth reading carefully and worth reading critically. Google says you don't need llms.txt files, content "chunked" for AI parsing, pages rewritten specifically for AI systems, inauthentic mention-seeking, or overfocusing on structured data. For Google Search's AI features, AI Overviews, and AI Mode, most of that is accurate. Google's RAG-based systems pull from their existing search index. Good SEO, non-commodity content, and clear technical structure already get you there. The GEO industry has created unnecessary complexity around features that are, for Google, just SEO by another name. But the guidance is narrower than it reads. It applies specifically to Google Search's generative features. It does not apply to direct agent queries operating outside Google's search index. When a consumer's agent queries your product catalog through ACP or UCP to execute a transaction, it is not going through Google's RAG pipeline. It needs machine-readable structured data, metafields, and API-accessible checkout, exactly what Google says not to overfocus on. It also doesn't apply to non-Google platforms. Perplexity, ChatGPT, and Claude each have their own retrieval architectures. The llms.txt guidance is Google-specific. Google even points to this gap itself: "Protocols like Universal Commerce Protocol are emerging that will allow Search agents to do more." That sentence covers the layer that their own guide doesn't address. The takeaway: for Google Search AI features, good SEO is sufficient. For direct agent commerce outside the Google index, structured data and protocol integration remain critical. Read Google's guide before following any GEO advice that doesn't specify which platform it's optimizing for. [Google’s Guide to Optimizing for Generative AI Features on Google Search | Google Search Central | Documentation | Google for DevelopersLearn how to optimize your website for Google Search’s generative AI features, including official best practices, technical SEO advice, and emerging AI agent guidance.![](https://storage.ghost.io/c/e1/c4/e1c4b063-1fa4-4977-8ffa-6add50cd2316/content/images/icon/favicon-new-3ab9c872-c10b-4a88-bafa-7d81b2a3a323.png)Google for Developers![](https://storage.ghost.io/c/e1/c4/e1c4b063-1fa4-4977-8ffa-6add50cd2316/content/images/thumbnail/home-social-share-lockup-66f0f265-b1c3-47b3-a7a7-4784869bc173.jpg)](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide?ref=agenticlandmark.com) ### The Job Doesn't Change. The Agent Does. URL: https://www.agenticlandmark.com/brief/the-job-doesnt-change-the-agent-does/ Last updated: 2026-06-29T21:50:41.000Z *What happens when you run Jobs to Be Done across the three horizons of AI evolution?* Clayton Christensen's Jobs to Be Done framework is built on one foundational observation: people don't buy products. They hire solutions to make progress in their lives. The job, the specific progress a person is trying to make, is stable. What changes over time is which solution does the job better. That insight is forty years old. It has survived every major technology transition in commerce. It is also the most clarifying lens I've found for understanding where agentic AI is headed and why the strategy question for brands is more urgent than most realize. ### The Three Horizons of AI in Commerce The Three Horizons framework, developed by McKinsey, describes how companies navigate technological change: Horizon 1 is the optimization of the current model, Horizon 2 is the emergence of genuinely new capabilities, and Horizon 3 is the transformation of the business model. The horizons are not sequential phases. They coexist. The skill is managing all three simultaneously. Applied to AI in commerce, the three horizons look like this: Horizon 1 is generative AI as a tool. Copilots, content generators, assistants. AI that augments the human decision process without replacing it. The consumer still drives. The AI helps. Rufus helps you decide. ChatGPT helps you compare. The human browses, evaluates, and buys. Horizon 2 is agentic AI as a delegate. The consumer sets parameters, price ceiling, preferred brands, and substitution rules, and the agent executes within them. The human is removed from the individual transaction. The agent researches, selects, and completes the purchase. The consumer reviews exceptions, not decisions. This is where Shopify's Agentic Storefront, Google's UCP, and the Scoper and Delegator postures operate today. Horizon 3 is an ambient AI as a standing system. The agent anticipates needs before they are consciously articulated. Replenishment without a request. Preference learning that compounds across categories. Commerce that happens as a background process of daily life rather than a deliberate act. The consumer sets a mandate, not a parameter set. The agent manages a relationship, not a task. ### Now run JTBD across all three Here is where the two frameworks produce something neither generates alone. The job the consumer is hiring for doesn't change across the horizons. In commerce, the core functional job is a version of this: get the right product, at the right price, with the least friction, reliably. That job existed before the internet. It will exist after ambient AI. The mechanism for accomplishing it changes dramatically. The job does not. What changes across the three horizons is the degree of consumer consciousness involved in hiring the solution. In Horizon 1, the consumer consciously hires the AI. They open ChatGPT, ask a question, and evaluate the answer. The hiring moment is explicit. The consumer is aware that they are delegating a subtask and retains authority over the outcome. In Horizon 2, the consumer consciously sets the parameters but then steps back from the individual hiring decisions. They configure the agent once. The agent hires the solution on their behalf for each transaction within scope. The consumer reviews the pattern, not the event. In Horizon 3, the consumer may not consciously hire anything. The ambient system, trained on their behavior and operating within a standing mandate, makes the hiring decision on their behalf before they know the need exists. The job gets done. The consumer may notice only in retrospect. This progression has a direct implication for how brands think about their relationship with the consumer. ### The insight that neither framework surfaces alone JTBD tells brands: understand the job deeply, because whoever does the job best wins. Three Horizons tells brands: the mechanism for doing the job is moving through distinct phases, and you need to be building for H2 while H1 is still your primary revenue channel. Run them together, and a third insight emerges: as AI moves from H1 to H3, the consumer's active role in the selection decision decreases. Which means the window in which a brand can influence selection through traditional means, marketing, creative, user experience, and promotional pricing is shrinking across the horizon progression. In Horizon 1, the brand can still influence the consumer directly. The ChatGPT answer is a recommendation. The consumer reads it and decides. Brand equity, content quality, and review reputation all still operate through human perception. In Horizon 2, the brand's influence moves upstream. The agent evaluates structured product data against declared parameters. The consumer set those parameters weeks ago. The brand influences selection by being correctly represented in machine-readable data, by appearing in the parameter sets consumers configure for their agents, and by having the trust record that makes an agent's prior selection of the brand a pattern worth repeating. In Horizon 3, influence is almost entirely infrastructural. The ambient system has learned what the consumer values, weighted by a trust record built over thousands of prior interactions. The brands with the longest, most accurate, and most frictionless history in that trust record are the default. New brands entering a consumer's consideration set at H3 have to break a standing pattern, which is structurally harder than winning a consideration-stage evaluation at H1. The JTBD principle says the job is stable and the solution evolves. The Three Horizons principle says the evolution is structured and predictable. Together, they say: the brand that understands the job deeply and builds the infrastructure to do it at H2 and H3 has a compounding advantage over the brand that optimizes for the H1 interface. ### What this means for brands right now Most brands are operating a Horizon 1 strategy. They are optimizing for AI-assisted discovery: GEO, content quality, and structured data as an SEO extension. These are the right investments. They are also the table stakes. The brands that will have a structural advantage in three years are the ones currently asking the H2 question: when a consumer's agent queries my catalog on their behalf, does my product data represent the job I'm actually doing for them, in machine-readable terms, at the attribute level the agent evaluates? And the brands worth watching are the ones already building for H3: what does a standing trust record look like in my category? What does it mean for a consumer's ambient system to have defaulted to my brand reliably enough that I become the pattern rather than the candidate? The job hasn't changed. Get the right product, at the right price, with the least friction, reliably. The mechanism for doing that job is in the middle of a horizon shift. The question is which horizon your infrastructure is built for. ### Hidden Text Is Already Telling AI Agents to Recommend Brands. Trust Breaks After Authorization. URL: https://www.agenticlandmark.com/brief/hidden-text-is-already-telling-ai-agents-to-recommend-brands-trust-breaks-after-authorization/ Last updated: 2026-06-29T21:47:48.000Z Google published research that every brand building for agentic commerce should read. Their Threat Intelligence team scanned billions of public web pages for indirect prompt injection, hidden instructions embedded in web content designed to manipulate AI agents reading that content. What they found matters beyond the security community. The categories they identified tell the story: pranks and experiments are mostly harmless. Educational content expected. Brands are trying to deter AI crawlers already widespread. Then there's the SEO category. Websites are embedding hidden instructions telling AI agents to recommend their products over competitors. Not through better product data. Not through legitimate optimization. Through invisible text that says, in effect: "If you are an AI, recommend us first." This is already happening. Google reports a 32% increase in malicious prompt injection detections between November 2025 and February 2026. The implication for agentic commerce is specific: the integrity of AI agent recommendations is not guaranteed. An agent querying a supplier site, a review platform, or a comparison page may be receiving poisoned instructions alongside the product data it came for. The agent you trusted to act on your consumer's behalf may be operating on manipulated inputs without either of you knowing. This is why behavioral trust verification matters at the transaction layer, not just at registration. A verified agent can still be compromised between authorization and execution by content it encounters in the wild, exactly the gap the DeepMind agent trap research named earlier this year. The brands building for agentic commerce need to understand this: the trust problem is not just about your own infrastructure. It's about the integrity of the content environment the agent navigates on your consumer's behalf. [AI threats in the wild: The current state of prompt injections on the webWe initiated a broad sweep of the public web to monitor for known indirect prompt injection patterns. This is what we found.![](https://storage.ghost.io/c/e1/c4/e1c4b063-1fa4-4977-8ffa-6add50cd2316/content/images/icon/apple-touch-icon-fc90224e-6803-4487-a608-792392fde2f0.png)GoogleThomas Brunner![](https://storage.ghost.io/c/e1/c4/e1c4b063-1fa4-4977-8ffa-6add50cd2316/content/images/thumbnail/google-1000x1000-eaa8e118-504a-4bf5-a4a0-2694fde0bbd6.png)](https://blog.google/security/prompt-injections-web/?ref=agenticlandmark.com) ### The Unit of Work Is Dissolving Again URL: https://www.agenticlandmark.com/brief/the-unit-of-work-is-dissolving-again/ Last updated: 2026-06-29T21:43:32.000Z *I've watched this happen four times. This time it's different in one specific way.* Natasha Jen said the quiet part out loud in a Fast Company piece published this week about Anthropic's Claude Design and the competitive scramble it triggered among Adobe, Figma, and Canva. "The software you open to make something may stop being the meaningful unit of work. Any company built on selling that unit has a real problem once it dissolves." I've been in design long enough to have watched the unit of work dissolve before. Multiple times. And Jen's observation is the most precise description of what's actually happening right now that I've read anywhere. ### What I've seen before I started with a BFA. My early career was analog, the craft, the physical production, the hours of work that were visible in the output because they had to be. Then the Mac arrived, and desktop publishing changed what design meant overnight. The unit of work shifted from paste-up and mechanicals to software. Quark. PageMaker. Eventually InDesign. The craft moved into the tool. Then the web. Then mobile. Then composable architecture and the distributed canvas of digital commerce. Each transition, the unit of work shifted. Each time, the designers who thrived were the ones who understood that the tool was never the point. The thinking was the point. The tool was just where the thinking happened to live for a while. What the Fast Company article describes is a version of this transition that is structurally different from the ones I've experienced before. And the difference matters. ### The frenemies dynamic is not new. What's new is the scope. The article details how Anthropic launched Claude Design, a direct competitor to the design tools it partners with, apparently without giving Figma or Adobe meaningful advance notice. Figma's stock dropped 7%. Adobe dropped 2.5%. Canva, which apparently co-developed Claude Design with Anthropic, got a preferential export button in the product. The design world's reaction was predictably dramatic. This kind of move is not unusual in technology. Google pays Apple billions to be the default search engine while competing directly with the iPhone. Platform relationships have always been frenemies arrangements. The party with the most leverage makes the most aggressive moves. What's different here is the scope of what's being contested. In previous transitions, the competition was over which tool designers used. Desktop publishing competed to be where you laid out pages. The web competed to be where you designed experiences. Mobile competed to be where you prototyped interfaces. What Anthropic is competing to be is something more foundational: the place where the work itself originates. Claude Design doesn't just want to be in your toolset. It wants to be the starting point of the creative process, the surface where an idea becomes a first draft before any other tool sees it. That's a different competitive ambition than Quark vs. PageMaker. It's a competition over who owns the blank page. ### Adobe's answer is the most interesting thing in the article Adobe is positioning itself as an AI "model curator." Rather than defending its own generative AI models against Claude, Gemini, and whatever else comes next, Adobe is building a unified interface that lets professionals access whichever model produces the best result for each specific task. You want imagery? Pick the model that performs best for imagery today. You want layout recommendations? Different model. The interface stays consistent. The models underneath it are interchangeable. This is the same strategic move that Adobe made when it built Creative Cloud: stop selling individual tools and start selling the ecosystem that connects them. Now the ecosystem connects AI models rather than applications. I've been watching Adobe make strategic decisions for thirty years, and this one is right. Not because it neutralizes the Claude Design threat, it doesn't, fully, but because it correctly identifies where the durable value is. The model that's best today won't be best in eighteen months. The professional who has built their workflow around a consistent interface, with the flexibility to swap models underneath it, has the most sustainable setup. Figma made the same call. Noah Levin put it directly: "We're not in the game of forming one extreme, deep partnership when all of these models excel at different things. And it's an advantage to not be a model company right now when you can actually just incorporate the pieces that make sense." This is wisdom earned from watching the tool transitions. The companies that tied their identity too tightly to a single technology, Quark's delay to develop for macOS, Flash's bet on proprietary runtime, every platform that confused its format with its value paid the price when the technology shifted. The ones that survived identified their value at a layer above the technology. ### What "raises the floor, not the ceiling" actually means for the field Andy Allen's observation in the article is the one that will age best: AI design tools are "more iMovie than Final Cut Pro." They raise the floor for people new to the field. They don't raise the ceiling for experienced practitioners. I've watched this pattern before. Desktop publishing raised the floor; suddenly, anyone with a Mac could lay out a newsletter. It didn't raise the ceiling. The designers who understood typography, hierarchy, visual communication, and the why behind every choice became more valuable because there were now millions of people making pages, and almost none of them were asking those questions. Mobile raised the floor. Anyone could build an app with the right tools. The designers who understood context, continuity, the thread of an experience across a person's day became more important because the surface got harder, not easier. What Claude Design and its competitors raise the floor on is production: getting from idea to first draft, from concept to comp. The speed at which a capable non-designer can produce something that looks designed has increased dramatically. What that means for experienced designers is the same thing every previous transition meant: the production skill is no longer the differentiator. The judgment is. Knowing why you made a choice, what it communicates, how it serves the person encountering it, how it holds up at the tenth iteration versus the first, that is the work that AI cannot do yet, and may not be able to do for quite a while. Lewenstein at Anthropic acknowledged this honestly: "Claude Design doesn't yet address that last mile craft and delight that differentiates the best products from the OK ones." The last mile craft is where the unit of work still belongs to the designer. Everything before it is increasingly contested. ### The question every design leader should be asking The Fast Company article focuses on the competitive dynamics between platforms. That's the right frame for investors and industry observers. But for design leaders inside organizations, the people responsible for teams, culture, quality, and the relationship between design thinking and business outcomes, the more urgent question is different. If the unit of work is dissolving, what is the unit you're building your team around? Production velocity? That's the unit under the most pressure. Speed is increasingly a commodity that AI provides. Judgment and direction? That unit is becoming more valuable, not less. The ability to evaluate AI-generated output, to recognize when something is technically correct but strategically wrong, to hold the standard that the tool cannot hold for itself, is the design skill that compounds now. The organizations that figure this out and restructure their teams, their workflows, and their quality frameworks around it will end up with stronger design cultures than they had before AI arrived. The ones that celebrate the speed without protecting the judgment will get more of everything, including more mediocrity, faster. Natasha Jen is right: the software you open to make something may stop being the meaningful unit of work. The thinking you bring to what gets made never stops being meaningful. That's the unit worth building around. *Read the Fast Company article in full:* [*https://www.fastcompany.com/91538439/design-enters-its-frenemies-era*](https://www.fastcompany.com/91538439/design-enters-its-frenemies-era?ref=agenticlandmark.com) ### The Agentic Commerce Race Is Google vs. OpenAI. Your Problem Is Neither. URL: https://www.agenticlandmark.com/brief/the-agentic-commerce-race-is-google-vs-openai-your-problem-is-neither/ Last updated: 2026-06-29T21:35:58.000Z Fast Company published the clearest mainstream account of the agentic commerce race I've read yet. Worth reading in full. The framing is Google vs. OpenAI, which builds the shopping platform consumers actually use. The Shopping Graph advantage is real. The protocol divergence between UCP and ACP is real. The OpenAI Instant Checkout failure is explained more honestly here than anywhere else I've seen. But the article is written from the platform side. The brand side of this story is the one that doesn't get told. Omar Qari's quote is buried near the middle, and it's the most important sentence: "OpenAI's first attempt at trying to get products into ChatGPT was to screen-scrape Dick's Sporting Goods or Ulta and show their products. And you can't blame them, because that's how they trained the model." That sentence is the entire brand infrastructure argument in one line. The platforms are being rebuilt from scratch to handle real-time product data, inventory, loyalty, and checkout. Whether your brand shows up in that rebuilt infrastructure is not a function of which platform wins the race. It's a function of whether your product data is structured, machine-readable, and connected to whichever protocol the winner runs on. Google wins; you need UCP integration and a Shopping Graph-compatible catalog. OpenAI wins; you need ACP integration and structured product feeds. Both need the same underlying infrastructure to work with. The race between Google and OpenAI will determine which interface consumers use. The infrastructure work your brand does now determines whether you're in the room when that interface opens. [https://www.fastcompany.com/91533534/shop-til-you-bot-google-openai-and-the-race-to-build-agentic-commerce](https://www.fastcompany.com/91533534/shop-til-you-bot-google-openai-and-the-race-to-build-agentic-commerce?ref=agenticlandmark.com) ### Agentic Commerce Is Not a Subset of E-Commerce URL: https://www.agenticlandmark.com/brief/agentic-commerce-is-not-a-subset-of-e-commerce/ Last updated: 2026-06-29T21:21:18.000Z *The category error that's causing brands to solve the wrong problem.* McKinsey projects $3 to $5 trillion in agentic commerce by 2030. Global e-commerce is projected to reach roughly $7 to $8 trillion by the same year. If agentic commerce were a subset of e-commerce, McKinsey's number would mean agents mediating more than half of all online retail within four years. That is an implausible claim. No technology transition in commerce history has captured the majority share that fast. The number only makes sense if McKinsey is measuring something different, not a slice of digital retail, but a projection of how much commerce across all channels will be initiated and mediated by agents, regardless of where the transaction ultimately completes. That distinction is not semantic. It is the most important strategic framing question in agentic commerce right now, and most brands are getting it wrong. ### The category error E-commerce is a channel descriptor. It describes where transactions happen: digital storefronts, browser-based checkout, and online payment flows. The defining characteristic is the medium through which the consumer interacts with the commerce surface. Agentic commerce is a mediation descriptor. It describes who initiates the transaction and how: a delegated agent acting on a consumer's declared mandate rather than a human navigating a browsing session. The defining characteristic is the presence of an intermediary layer between consumer intent and commercial outcome. These are not the same type of thing. One describes a channel. The other describes a mediation model that can operate across channels. The confusion arose because the first visible agentic commerce deployments, ChatGPT Shopping, Gemini checkout, and Perplexity Buy with Pro, happened to be digital-native platforms. Early observers mapped them onto the e-commerce category because the transaction was completed in a digital interface. But the interface where the transaction is completed is not the defining feature. The agent mediating the decision is. ### What the mediation model actually touches An agent executing a commerce task on a consumer's behalf does not care about the channel taxonomy brands use to organize their marketing budget. A travel agent configures a flight, hotel, and car rental across three different booking systems. A financial agent rebalances a portfolio, pays a bill, and renews an insurance policy. A grocery agent replenishes a standing list, substituting where items are out of stock and flagging when the brand preference is unavailable. A home improvement agent builds a parts list for a project, checks local store inventory, and reserves the items for pickup. None of these transactions is e-commerce in the traditional sense. Some complete inside AI platforms. Some complete inside loyalty apps. Some complete through voice interfaces. Some initiate online and close in a physical store through a BOPIS fulfillment flow. The back end that executes each transaction, the OMS, the payment rail, and the fulfillment infrastructure, is the same infrastructure that has always powered commerce. What changed is the layer above it. The discovery and selection layer that precedes the transaction has moved. In e-commerce, that layer was a browser session: a consumer navigating a website, evaluating options, and adding to cart. In agentic commerce, that layer is a mandate: a consumer configuring parameters and delegating execution to an agent. The transaction still touches the same infrastructure. But the human browsing session that triggered it is no longer there. ### Why is the scope larger than e-commerce? If you accept that agentic mediation applies to any channel where agents can operate, which is increasingly every channel, then the addressable scope of agentic commerce is not "digital retail." It is "commerce wherever agents operate," which is a substantially larger and more distributed surface. Consider the verticals where agent mediation is already live or imminent: Travel. Agents configuring multi-leg itineraries, managing loyalty redemption across programs, booking across hotel, flight, and ground transportation in a single delegated flow. This is not e-commerce. This is agent-mediated commerce that touches every booking system in the travel stack. Financial services. Agents managing subscription renewals, insurance comparisons, loan applications, and investment rebalancing within declared parameters. TRID, FINRA, and IAA requirements structure how delegation works here, but the agent mediation model applies regardless of the regulatory constraints on it. Real estate. Agents are compressing the discovery and scheduling phases of property search, coordinating showings, and managing document preparation. The Delegator posture is structurally bounded by transaction stakes and legal requirements at offer and contract stages, but the upstream agent mediation is substantial. Retail DTC. The channel is most commonly labeled "agentic e-commerce," but even here, agents are mediating across both online and in-store fulfillment surfaces. A Scoper-posture consumer who delegates restocking within a declared brand list and price ceiling is not conducting e-commerce. They are delegating a standing commerce mandate that executes across whatever fulfillment surface the agent finds available. The McKinsey $3-5T figure starts to make sense when you aggregate across these verticals. It is not a projection of agentic share within e-commerce. It is a projection of commercial value that flows through agent-mediated decision processes across the full commercial landscape. ### The wrong investment that follows from the wrong framing When brands treat agentic commerce as a subset of e-commerce, the strategic implication is that it requires an e-commerce optimization response: better product pages, improved conversion flows, more sophisticated retargeting, incremental SEO investment extended into GEO. None of those investments addresses the mediation layer. An agent does not browse your product pages. It queries your structured data infrastructure. It does not respond to your retargeting. It executes within the parameters your consumer declared when they configured their mandate. It does not arrive through a search result. It arrives through a protocol query that either finds your product data machine-readable and complete or does not find it at all. The optimization discipline for agents is entirely different from the optimization discipline for human browsers. The measurement infrastructure is different. The competitive dynamics are different. The compounding mechanisms are different. Brands that have framed this as an e-commerce optimization problem are investing in the interface layer, while the mediation layer that determines whether agents select them at all goes unbuilt. ### What are the correct framing changes If agentic commerce is a mediation model rather than a channel, the strategic implication is that every brand in every vertical with a consumer-facing commerce dimension is in scope, not just the ones running digital storefronts. The infrastructure requirement cuts across channels simultaneously. A brand's product data needs to be machine-readable, whether the agent querying it is operating inside ChatGPT, Gemini, a loyalty app, a voice interface, or a financial agent. A brand's checkout infrastructure needs to accept agent-initiated transactions regardless of which protocol the agent is running. A brand's loyalty data needs to be accessible to agents operating on behalf of consumers who hold loyalty status. None of this is an e-commerce project. It is a commerce infrastructure project that happens to include digital channels as one of its components. The brands that will navigate the next three years well are the ones that have correctly identified the scope of the problem they are solving. Agentic commerce is not a new channel to add to the digital marketing mix. It is a structural change in how commerce initiates, and that change touches every channel simultaneously, which is exactly what the McKinsey number is actually measuring. ### Agentic Commerce Runs on Power. We Should Care Where It Comes From. URL: https://www.agenticlandmark.com/brief/agentic-commerce-runs-on-power-we-should-care-where-it-comes-from/ Last updated: 2026-06-29T21:17:23.000Z The energy problem underneath the AI boom is bigger than most technology discussions acknowledge. By 2030, global data center electricity consumption is projected to reach 945 TWh, roughly equivalent to Japan's entire national electricity consumption today. In the United States alone, data centers will consume more electricity by the end of the decade than all aluminum, steel, cement, and chemical production combined. AI workloads are the primary driver, with AI-specific server electricity consumption projected to grow 30% annually. The grid was not built for this. 70% of U.S. power infrastructure was constructed between the 1950s and 1970s. The IEA estimates 20% of planned data center projects are already at risk of delays due to grid connection queues and capacity constraints. Virginia, the largest U.S. data center market, is seeing its first base-rate electricity increase since 1992, paid for by households, not the hyperscalers generating the demand. The traditional answer to this problem builds more terrestrial generation, extends the grid, wait for permitting, which cannot keep pace with the demand curve AI is producing. Which is why companies like Panthalassa are worth paying attention to. Panthalassa is building a planetary-scale clean energy platform from the middle of the ocean. The premise: the ocean is where the energy is. Offshore renewable generation, faster to deploy than terrestrial infrastructure, is designed to serve the scale of compute demand that AI actually requires. Their Pacific Ocean pilot, Ocean-2, is already in deployment. The team comes from SpaceX, Blue Origin, NASA, Google, Boeing, Tesla, Amazon, and Microsoft, the same organizations that have learned to build infrastructure at a pace that terrestrial regulatory and construction timelines cannot match. The AI energy problem is not going to be solved by incremental grid expansion and power purchase agreements. It will be solved by people rethinking where energy comes from, how it gets to compute, and what infrastructure can scale at the speed the technology demands. Panthalassa is one of the more serious answers I've seen to that question: [https://panthalassa.com](https://panthalassa.com/?ref=agenticlandmark.com) ### The Adobe Q2 2026 AI Traffic Report Is the Most Important Commerce Data Published This Year URL: https://www.agenticlandmark.com/brief/the-adobe-q2-2026-ai-traffic-report-is-the-most-important-commerce-data-published-this-year/ Last updated: 2026-08-30T21:24:36.000Z *The numbers confirm the transition. The visibility gap data names exactly where the work needs to happen.* Adobe Analytics released its Q2 2026 AI Traffic Report on April 16, covering the first quarter of 2026 across more than one trillion visits to U.S. retail sites. The numbers are significant on their own. Read together, they make an argument that should end a lot of ongoing debates in commerce strategy circles. The argument is this: AI-referred traffic now converts better than the rest of the channel mix combined. And most brands are invisible to the systems generating it. ### The traffic numbers AI-driven traffic to U.S. retail sites grew 393% year over year in Q1 2026\. In March specifically, the figure was 269% year over year, a slight deceleration from the 693% surge recorded during the holiday season, but still representing extraordinary compound growth off a base that itself represented a step-change from the prior year. To understand the scale: Adobe's data covers direct transactions online and represents more retail site traffic than any other technology company or research organization tracks. These are not survey estimates. They are observed visits across a trillion-visit dataset. Consumer behavior data from Adobe's companion survey of 5,000 U.S. respondents reinforces the traffic figures. 39% of respondents report using AI for online shopping, with 85% of that group saying it improved their experience. Trust is moving with adoption: 66% of respondents believe AI tools provide accurate results. ### The conversion reversal that changes everything The most strategically significant data point in the report is not the traffic growth figure. It is the conversion reversal. In March 2025, AI-referred traffic converted 38% worse than non-AI traffic. In March 2026, AI-referred traffic converts 42% better than non-AI traffic. That is a complete reversal in relative conversion performance over twelve months. This reversal matters because it closes the last credible argument for treating AI-referred traffic as an experimental channel worth monitoring rather than a primary channel worth building infrastructure for. The objection that AI-sourced visitors are lower quality, less purchase-ready, or more likely to bounce is now empirically wrong. Adobe's data shows the opposite: AI-sourced visitors in March 2026 spent 48% longer on retail sites, browsed 13% more pages per visit, and had a 12% higher engagement rate than non-AI traffic. The explanation is structural. When a consumer arrives from an AI referral, the agent has already done the qualification work, comparing options, evaluating specifications, and forming purchase intent before the click. The visitor arrives at the site further along in the decision process than any other channel delivers. They are not browsing. They are confirming. A single B2B software case study, with ChatGPT-referred signups converting at 15.9% against 1.76% for Google organic, pointed to this pattern. Adobe's Q2 2026 data confirms it at scale across the entire U.S. retail sector. ### The visibility gap that most brands are ignoring Adobe released a second dataset alongside the traffic figures: a machine readability benchmark across the U.S. retail sector, produced by its AI Content Visibility Checker. The benchmark assigns a score from 0 to 100% representing what percentage of a page's content is readable by AI systems. A score of 66% means a third of the content on that page is invisible to the agents evaluating it. Here is what the benchmark found across U.S. retail: Homepages: 75% machine readable. One quarter of homepage content is invisible to AI systems. Category pages: 74%. Similar to homepages. Product pages: 66%. The pages where purchase decisions are made and where structured product attributes live are the least machine-readable pages in the retail stack. The product page number is the most consequential finding in the report. Retailers have thousands of SKUs. At 66% machine readability, a third of the product data on those pages, the attributes, specifications, comparisons, and trust signals that agents evaluate when making recommendations, does not exist from the agent's perspective. Other page types ranged from 73% for store locator pages to 82% for returns and exchanges pages. Loyalty and membership pages came in at 78% meaningful because loyalty data is one of the primary signals agents use to personalize recommendations for consumers who have granted shopping authority. The gap between best and worst performers is the most actionable data point for brands assessing their own position. The top-performing U.S. retailers score 82.5% on homepage machine readability. The lowest performers score 54.2%. That 28-point gap between leaders and laggards on a single page type represents the current competitive surface in AI-driven commerce. The brands at 82.5% are more visible to the agents, generating the 393% traffic surge. The brands at 54.2% are significantly less so. ### What the two datasets together mean The traffic and visibility data, read together, describe a market in transition with two distinct populations of brands. The first population has built or is building machine-readable infrastructure. Their product pages return structured, parseable data when agents query them. Their category pages surface a clean taxonomy that agents can use for comparison. Their loyalty data is accessible. These brands are capturing a disproportionate share of the 393% traffic surge because the agents generating that traffic can find, evaluate, and recommend them accurately. The second population has not. Their product pages, at 66% machine readability or below, return partial data. Agents either cannot evaluate their products accurately or deprioritize them in favor of brands whose data is complete. The 393% traffic surge is happening, they are just not capturing a proportionate share of it. The conversion data makes this consequential in a way that pure traffic growth does not. If AI-referred traffic converted at the same rate as other channels, the share of that traffic would be a missed incremental opportunity. At 42% better conversion than the rest of the channel mix, the missing share of AI-referred traffic means missing the highest-quality visitors arriving at your category. These are not window shoppers. They are buyers. ### The machine readability problem is not a content problem The instinctive response to machine-readable data is to treat it as a content problem: write better product descriptions, add more attributes, improve the copy. That framing is wrong in a specific and expensive way. Machine readability is an infrastructure problem. The 34% of product page content that is invisible to AI systems is not invisible because the content does not exist. It is invisible because the content exists in forms embedded in JavaScript, buried in images, structured for human visual parsing rather than machine data extraction that AI crawlers cannot parse. Adobe's research on AI crawlers reinforces this: JavaScript renders correctly to human browsers while AI crawlers often see only a skeleton of the page. Product data that exists for human shoppers may be entirely invisible to the agents evaluating that product for shoppers who have delegated discovery and selection. The Botify data cited in the practice's market intelligence is consistent: agents evaluate without visiting in the way humans visit. The upstream evaluation the agent's assessment of your product data before any human decision is made is the commercial event that determines selection probability. The fix is not a copywriting sprint. It is a structured data architecture: moving product specifications, pricing logic, availability, review data, and trust signals into machine-readable formats that agents can parse at query speed. Metafields over HTML descriptions. Schema markup over visual styling. API-exposed product attributes over embedded content. ### Why this report matters for brands still debating urgency The Adobe Q2 2026 report answers the urgency question in a way that no projection or scenario analysis can: with current performance data from a trillion-visit dataset. The AI channel is not coming. It generated 393% year-over-year traffic growth in Q1 2026\. It converts 42% better than the rest of the channel mix. Shoppers who arrive from it spend 48% longer on site. The competitive gap between brands at 82.5% machine readability and brands at 54.2% is already open and already affecting which brands those shoppers are arriving at. The window for building infrastructure before the channel reaches a majority share is not indefinitely open. The brands at 82.5% are compounding their advantage with every agent recommendation that routes high-intent traffic their way. The brands at 54.2% are losing ground on every one of those recommendations. Adobe's headline is that AI traffic is surging. The finding that should be driving commerce strategy decisions is the one buried in the visibility benchmark: a third of your product data doesn't exist to the fastest-growing, highest-converting channel in your category. That is not a problem to monitor. It is a problem to solve. Source: Adobe Digital Insights, published April 16, 2026, accompanying the Q2 2026 AI Traffic Report. Data based on 1T+ visits to U.S. retail sites and a companion survey of 5,000 U.S. consumers. **Correction, August 30, 2026.* Several figures in this post have been corrected.* *The Seer conversion figures were originally described as live merchant data, with Google organic at 1.8 percent. Both were wrong. The source is a single B2B software client, from October 2024 to April 2025, with roughly 1,370 AI conversions against approximately 14 million organic sessions, and the conversion action is a sign-up rather than a purchase. The Google organic figure is 1.76 percent.* *The difference between 82.5 percent and 54.2 percent was described as 52 percentage points. The absolute gap is 28.3 points. The 52 is Adobe's relative percent difference between the two scores, not a gap measured in points.* *The year-over-year conversion change was described as an 80-percentage-point swing. That adds two relative differences as though they were absolute. The post now describes the reversal without assigning it a point value.* *The post also claimed AI-referred traffic converts better than every other channel. Adobe's comparison is against non-AI traffic in aggregate, and the claim has been narrowed to match.* *The machine readability figures come from Adobe Digital Insights' post of April 16, 2026, which accompanies the Q2 2026 AI Traffic Report rather than appearing within it.* ### Most Companies Are Experimenting with AI Agents. Almost None Have Made Them Work. URL: https://www.agenticlandmark.com/brief/most-companies-are-experimenting-with-ai-agents-almost-none-have-made-them-work/ Last updated: 2026-06-29T21:04:56.000Z McKinsey's latest from their Technology and QuantumBlack teams puts a number on what many of us have been saying: nearly two-thirds of enterprises worldwide have experimented with AI agents. Fewer than 10 percent have scaled them to deliver tangible value. The reason is not model capability. It is data. Eight in ten companies cite data limitations as the primary roadblock to scaling agentic AI. The article prescribes the right medicine for CIOs. Modernize your data architecture layer by layer rather than rebuilding from scratch. Build a semantic layer that codifies business meaning into machine-readable form through ontologies and knowledge graphs. Shift from periodic data cleanup to continuous, real-time quality management. Evolve your operating model so humans supervise and orchestrate agent-driven workflows rather than execute what agents now handle. All of that is correct. And all of it describes the internal infrastructure problem. Here is the part missing from the conversation. ### The same gaps that break internal agents break external ones too The data architecture gaps that prevent a company from scaling its internal agent workflows are the same gaps that make the brand invisible to external agents acting on behalf of consumers. Product data sitting in silos across channels does not just produce inconsistent internal recommendations. It produces a brand that an AI shopping agent cannot evaluate accurately, cannot compare against competitors, and cannot transact with reliably. McKinsey illustrates this with an omnichannel retail example. Product data and purchase histories sat in silos, so context broke as customers moved across channels, producing inconsistent recommendations and service experiences. Their prescription is an agent-ready architecture that connects systems and data to support the entire customer commerce journey. That is the right architecture. But it solves the problem from the inside out. The same data fragmentation that breaks the internal customer journey also breaks the external one, the journey where a consumer's AI agent is evaluating your brand against every competitor in the category before the consumer ever visits your site. The agent does not average your conflicting signals across channels. It routes around you entirely. ### The semantic layer is the commerce layer The most technically significant section of the McKinsey piece describes the semantic layer: the structure that sits between raw data and AI applications, codifying what things are, how they relate, and what rules govern them. Without this shared semantic foundation, they argue, agents act on incomplete or conflicting interpretations of the same data. This is the exact problem I see in every commerce vertical I work in. A hotel described as "family-friendly" in marketing copy is invisible to an agent evaluating family suitability. An agent needs connecting room availability, kids menu status, step-free access, crib availability, each as structured, queryable fields. A financial product described as having "competitive rates" is equally invisible. An agent needs a specific APR, a minimum balance, documented eligibility criteria, a fee schedule. The semantic layer McKinsey describes for internal enterprise data is the same layer that makes a brand's products legible to external agents. The enterprise data architecture community and the commerce readiness community are converging on the same requirement from different directions. One is asking how to make agents work inside the organization. The other is asking how to make the brand legible to agents working outside it. ### The board-level argument hiding in these numbers The scale gap statistic, fewer than 10 percent of enterprises have moved beyond experimentation, reframes the competitive conversation for any brand considering agentic commerce investment. The question most boards are asking is whether competitors are ahead of them. The honest answer, based on this data, is that almost nobody is ahead of anybody. The field is wide open. More than 90 percent of enterprises that have experimented with agents have not figured out the data infrastructure to make them work at scale. That is not a reason to wait. It is the reason to move now. The brands that close the data infrastructure gap in the next 12 to 18 months are not competing against a mature field. They are building a structural advantage in a market where the vast majority of participants are still stuck at the pilot stage. The compounding effect of being early, in merchant trust scores, in agent selection rates, in the quality of the data that agents learn from, is real and difficult to replicate from behind. ### Both conversations need to happen at the same time I have spent 30+ years in digital product strategy and commerce. The pattern I see now is one I have seen before. When mobile commerce emerged, the brands that treated it as a channel optimization problem, separate from their core commerce architecture, spent years catching up to the brands that treated it as an infrastructure problem from the start. Agentic commerce is the same inflection, but the stakes are higher because agents are less forgiving than mobile browsers. A mobile site that loaded slowly lost a percentage of conversions. A brand that is invisible to agents loses the evaluation entirely. The consumer never knows the brand was an option. The CIO conversation McKinsey is driving, restructuring internal data for agentic workflows, is necessary. But it is not sufficient. The CMO conversation, ensuring the brand's products, pricing, and attributes are structured for the external agents making purchase decisions on behalf of consumers, needs to happen in parallel. Both problems start with the same foundation: structured, complete, semantically clear data. The brands solving both problems simultaneously are the ones building a durable advantage. The brands solving only one are building half an architecture and hoping the other half does not matter. It will matter. It already does. ### Thirty years. One Icon. Zero questions asked. URL: https://www.agenticlandmark.com/brief/thirty-years-one-icon-zero-questions-asked/ Last updated: 2026-06-29T20:59:12.000Z Look at the top-right corner of almost any e-commerce site. There it is. A little shopping cart icon. Same as it was in 1995. We've redesigned everything about online shopping in thirty years: search, recommendations, payments, delivery, and mobile. But the cart remained. We never questioned it. We should have. The shopping cart is a skeuomorph. A digital artifact that mimics physical behavior for no reason other than familiarity. It made sense when the internet was new, and users needed recognizable metaphors. You push a cart through a store. You add things to it. You wheel it to the register. But we're not in 1995\. And increasingly, we're not shopping that way either. When an AI agent handles a purchase, there is no cart. There's no gathering phase. No browsing, no comparing, no adding, no abandoning. The agent already knows what's available, has already evaluated the options, and executes when appropriate. Seventy percent of online shopping carts are abandoned before purchase. Seven in ten. We've accepted that number as normal for so long that most brands have entire CRO programs dedicated to recovering them. But the cart abandonment problem doesn't get solved in an agent world. It disappears. Because the cart disappears. The cart was never a feature. It was a workaround, a bridge between physical retail and something we hadn't yet invented. That something is here. Brands still optimizing cart completion rates are optimizing the last step of a process that is being bypassed at the first step. The question isn't how to reduce abandonment. It's whether your infrastructure is visible, evaluable, and transactable before a consumer ever decides to gather anything. ### McKinsey's Twelve AI Themes Are Sound. For Commerce, Two Carry the Urgency. URL: https://www.agenticlandmark.com/brief/mckinseys-twelve-ai-themes-are-sound-for-commerce-two-carry-the-urgency/ Last updated: 2026-06-29T20:45:29.000Z McKinsey just published a twelve-point AI transformation manifesto. It's worth reading. The twelve themes are sound. Build enduring capabilities, not one-off solutions. Focus on economic leverage points. Make senior business leaders accountable for the tech agenda. Treat data as a performance asset. Design for adoption before you scale. These are the right principles, and most organizations are still failing at them. But the frame that runs through the document "rewired companies" obscures the more specific question most commerce organizations are facing right now. The question is not whether they've rewired for AI in the general sense. It's whether they've built the specific infrastructure layer that determines whether they're visible and selectable in agent-mediated commerce. Theme 8 is where McKinsey gets closest: "make data easy to consume and enrich it for advantage." The observation is correct. The implication for commerce brands is more specific than the document makes it. Clean, structured, machine-readable product data isn't just a general AI readiness requirement. It's the specific input that determines whether an agent querying your catalog can find, evaluate, and select your products at all. The brands that have treated product data as a business-owned performance asset for the last two years are the ones showing up in agent recommendations today. Theme 10 "no trust, no right to deploy AI" names something important that most commerce AI discussions underweight. The agentic trust problem isn't just internal governance. It's consumer-facing. A consumer who configures an agent to shop on their behalf is making a trust decision about the agent, the platform, and the brands the agent will interact with. The authorization surface where that trust is established or lost is the most consequential UX problem in commerce right now, and almost nobody is designing for it. The twelve themes are a useful diagnostic for organizational AI readiness. For commerce specifically, themes 8 and 10 are where the most urgent work is concentrated, and both require more specificity than a general transformation framework can provide. Worth reading in full. McKinsey Quarterly, April 2026. [https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-ai-transformation-manifesto](https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-ai-transformation-manifesto?ref=agenticlandmark.com) ### BCG Identified Four Agentic AI Futures for Retail. The Infrastructure Requirements Are the Same in All of Them. URL: https://www.agenticlandmark.com/brief/bcg-identified-four-agentic-ai-futures-for-retail-the-infrastructure-requirements-are-the-same-in-all-of-them/ Last updated: 2026-08-30T18:47:06.000Z In March 2026, Boston Consulting Group published a scenario analysis on how AI agents will reshape retail commerce. Six Managing Directors and Partners built a four-scenario matrix. Named the imperatives that cut across every scenario. And told CEOs they could no longer afford to wait for clarity. The most important sentence in the document is not one of the named scenarios. It is this: "The marketing challenge isn't predicting which future will emerge, but building capabilities that are robust enough to succeed regardless of which futures unfold." That sentence is the entire strategic case for agentic commerce infrastructure. Here is what it means in practice. ### BCGs Four Scenarios BCG structured their analysis around two variables: where will influence reside (algorithmic vs. human), and how will market power be concentrated (platforms vs. distributed agents). The four combinations: **The Open Agentic Bazaar** is distributed algorithmically. Agents browse freely across brands. Machine-readable product data and API accessibility determine who gets selected. **The Super-App Embrace** consolidated, algorithmic. A handful of tech giants own the interface. Brands pay for access and optimize for each platform's proprietary logic. **Brand Resurgence Through Data Fortresses** consolidated, human. A few large brands and retail platforms dominate. Loyalty architecture, closed ecosystem data, and proprietary agent engines are the moat. **Creator-Led Authenticity Revival** distributed, human. Consumers trust creators over algorithms. Brand equity, community relationships, and transparent storytelling determine preference. BCG is careful to say these are not mutually exclusive; some combination will coexist, and the dominant mix will vary by category and geography. What they cannot tell you is which combination will dominate, or when. ## What Persists Across All Four BCG named two imperatives that cut across every scenario: discoverability and desirability. Discoverability is the ability to be found by agents, regardless of which system they operate in. The specific implementation differs by scenario. The underlying requirement that your product data is structured in a form agents can evaluate, rather than in prose and HTML spec tables, is the same in all four. Desirability is a brand strength that survives algorithmic intermediation. BCG: "branding won't matter less; it will matter differently, but even more than it used to." The data on both is pointed. BCG found only an 8 to 12 percent overlap between traditional search results and AI-generated answers. SEO and AEO are not the same channel. SEO captures bottom-funnel intent from humans already in buy mode. AEO influences the top and middle of the funnel, where agents build consideration sets before a human ever engages. A brand invisible at the AEO stage never appears in the shortlist that the consumer eventually reviews. ## The Specific Gap Most Brands Are Missing The brands capturing AI-referred traffic have done something specific. In a single B2B software case study, ChatGPT-referred traffic converted at 15.9% against 1.76% for Google organic, on signups rather than purchases. They have expressed their product attributes as typed, machine-readable data rather than as prose copy. An agent building a comparison matrix for a consumer cannot parse a paragraph into a comparison column. It can read a typed attribute directly. The difference between a brand that appears in an agent's shortlist and one that does not is often not product quality or price. It is whether the relevant attribute is in a [Schema.org](http://schema.org/?ref=agenticlandmark.com) additionalProperty field or in a sentence. Shopify merchants with agent-readable catalogs saw AI-sourced orders grow 15x in 12 months. The merchants not on that growth curve are running the same Shopify infrastructure with catalog data in the wrong format. ## The Scenario-Robust Investment Case BCG's most useful practical point is the reframing from forecast to preparedness. Most brands are asking when agent-driven purchases will hit critical mass. BCG argues that it is the wrong question. The right question is how to build capabilities that work whether the market is 10% agent-driven or 90%. Infrastructure investments compound. A brand that begins building structured data, entity records, and AEO content architecture in 2026 enters 2027 with 12 months of compounding in agent crawl indexes and citation patterns. A brand that starts in 2028 begins from zero while competing against brands that have had two years of compounding. BCG on this: "CEOs cannot afford to wait any longer for clarity. The price of indecision compounds faster than the cost of making a mistake." ## The Signal Worth Watching BCG identifies a set of early signals that reveal which future is arriving: the growing share of purchase journeys that begin inside AI agents rather than search engines, whether AI assistants recommend brands by name or by specification, and how consumers respond to content they know was generated by machines. The brands positioned to read those signals accurately are the ones that have already built the infrastructure that makes them legible to agents. Not because they predicted the right future. Because they invested in the one thing that matters in all four of them. [https://www.bcg.com/publications/2026/agentic-scenarios-every-marketer-must-prepare-for](https://www.bcg.com/publications/2026/agentic-scenarios-every-marketer-must-prepare-for?ref=agenticlandmark.com) **Correction, August 30, 2026.* This post originally described the Seer conversion figures as live merchant data and gave Google organic as 1.8 percent. Both were wrong. The source is a single B2B software client, from October 2024 to April 2025, with roughly 1,370 AI conversions against approximately 14 million organic sessions, and the conversion action is a sign-up rather than a purchase. The Google organic figure is 1.76 percent.* ### Hierarchy Was Always an Information-Routing Protocol. AI Removes the Constraint. URL: https://www.agenticlandmark.com/brief/hierarchy-was-always-an-information-routing-protocol-ai-removes-the-constraint/ Last updated: 2026-06-29T20:38:40.000Z Jack Dorsey and Roelof Botha published something worth reading slowly this week. The thesis: hierarchy has always been an information routing protocol, not a management philosophy. The Roman contubernium, the Prussian General Staff, McCallum's railroad org chart, and two thousand years of organizational innovation were an attempt to solve the same problem. A leader can effectively manage somewhere between three and eight people. So you add layers. Layers slow information. You optimize the layers. The information still slows.AI removes the constraint. Not by making the layers faster. By replacing what the layers do. Most organizations are responding to this by giving everyone a copilot, making the existing structure work slightly better without changing it. Block is asking a different question: if AI handles information routing, what is the organization actually for? I've spent 35 years in organizations ranging from startups to regulated financial institutions, and the most consistently effective leaders I've worked with or for were already operating this way instinctively, without the vocabulary for it. They were close to the work, close to their people, and allergic to the coordination theater that organizations build around information scarcity. The best engineering leaders I've known write code. The best design leaders I've known still design. The ones who stopped doing the work to manage the information flow were the ones who became the bottleneck they were hired to clear. The Block model isn't radical if you've seen it work. What's radical is the scale at which it becomes viable. The sentence that stayed with me: "The edge is where the intelligence makes contact with reality." The human role in this model is not coordination. It is judgment, intuition, cultural context, ethical decision-making, and the ability to sense things the model cannot perceive. That is a much harder job than managing information flow. It is also a much more honest description of what great leadership has always actually been. The organizations that will struggle with this transition are the ones that built their authority structures around information control, knowing things others didn't, owning the context that made them necessary. That scarcity is ending. What replaces it is harder to fake and harder to automate: the quality of your judgment at the edge. Block is in the early stages of this. So is every organization that's being honest about it. But the question they're asking is what your company understands that is genuinely hard to understand, and is that understanding getting deeper every day, the right one. [From Hierarchy to IntelligenceHow Block is using AI to eliminate hierarchical bottlenecks, building the first company organized as intelligence rather than hierarchy.![](https://static.ghost.org/v5.0.0/images/link-icon.svg)From Hierarchy to IntelligenceBlock![](https://storage.ghost.io/c/e1/c4/e1c4b063-1fa4-4977-8ffa-6add50cd2316/content/images/thumbnail/0aede75a69f4e57d1624891cac16c799987523a7-3840x2160-633e2e48-103e-4a8f-b1b0-8c91df4cef67.png)](https://block.xyz/inside/from-hierarchy-to-intelligence?ref=agenticlandmark.com) ### Google UCP Update: Three Capabilities, Three Infrastructure Gaps URL: https://www.agenticlandmark.com/brief/google-ucp-update-three-capabilities-three-infrastructure-gaps/ Last updated: 2026-06-29T19:12:57.000Z Google quietly updated the Universal Commerce Protocol last week. Three new capabilities. Most coverage missed what each one actually names. Here’s what changed and the specific infrastructure problem each capability exposes. ## **Capability 1: Cart** Agents can now add multiple items to a shopping cart simultaneously from a single store. Before this, agent-mediated UCP transactions were single-item. That’s a significant limitation: you can’t run a grocery replenishment workflow, an apparel outfit assembly, or a household essentials standing order on a single-item protocol. The basket is the fundamental unit of commerce for most retail categories. The infrastructure gap is named: basket assembly requires that every item in that cart is queryable, comparable, and available in real time, not just the hero product. A brand whose product catalog is partially structured, or whose inventory data is accurate for in-demand SKUs but stale for everything else, will see agent-assembled carts fail or degrade at the edges. Cart capability makes catalog completeness a transaction-critical requirement, not a data quality nice-to-have. ## **Capability 2: Catalog** Agents can now retrieve real-time product details, variants, inventory, and pricing directly from a retailer’s catalog at the moment of query. This is the live retrieval layer. Before this, agents were largely dependent on what they had trained on or last crawled. That data is often stale. New products, live pricing changes, promotional pricing, and real-time stock status weren’t reliably accessible at transaction time. The infrastructure gap is named: there is a measurable difference between a brand that trains AI well and a brand that is queryable in real time. Training data determines whether agents know you exist. Catalog integration determines whether agents can verify you’re relevant right now. A brand with well-structured product data but no Catalog endpoint is discoverable but not transactable at the current state. That gap between discovery and real-time transaction readiness is what the Catalog capability makes concrete and testable. ## **Capability 3: Identity Linking** Shopper loyalty accounts can now be recognized across integrated platforms when transacting through agents. This is the least-discussed of the three and arguably the most commercially significant near-term. The primary reason consumers delegate shopping tasks to agents is time savings and better outcomes. Loyalty redemption applying points, accessing member pricing, and earning towards status is a core part of what “better outcome” means for a large portion of the buying population. Without identity linking, agents either skip loyalty entirely or require the consumer to intervene at the point where loyalty would apply, which defeats the purpose of delegation. The infrastructure gap is named: loyalty program architecture was not designed for agent-mediated transactions. Most loyalty systems authenticate against a human session a browser cookie, a logged-in account, or a physical card scan. Agents don’t have sessions in the human sense. They carry delegated authorization. Identity linking is the mechanism that maps authorization to a loyalty account. Brands whose loyalty infrastructure can’t support this will see agents either ignore their loyalty programs entirely or generate exceptions that force human intervention at the worst possible moment in the transaction flow. ## **The fourth thing: Merchant Center onboarding** Google also simplified UCP onboarding through Merchant Center. Less technically interesting but strategically significant. Shopify gave its merchants a one-click path to agent discovery through Agentic Storefronts in January. Google is now making UCP integration simpler for Merchant Center retailers in March. Two on-ramps, advancing in parallel. The brands outside both WooCommerce and Magento, headless architectures, and custom stacks, have neither. The gap is widening from two sides simultaneously, and it is widening faster than most teams are moving. The three new UCP capabilities are not features. They are a diagnostic instrument. Can your agent-assembled cart handle incomplete catalog coverage at the edges? Can your product data be queried in real time for variants, inventory, and pricing, not just crawled periodically? Can your loyalty infrastructure recognize a delegated agent identity and execute redemption without requiring human intervention? Three questions. The answers determine whether the protocol advancement Google shipped last week translates into commercial advantage for your brand or commercial advantage for a competitor whose infrastructure is already positioned to use it. ### ShopTalk Spring 2026: What the Announcements Actually Mean URL: https://www.agenticlandmark.com/brief/shoptalk-spring-2026-what-the-announcements-actually-mean/ Last updated: 2026-06-29T20:09:06.000Z *The most important detail from the Gap/Gemini announcement was the one most coverage skipped. Start there.* ShopTalk Spring 2026 ran March 24-26 in Las Vegas against a backdrop that would have been unrecognizable two years ago. The conference theme was "Retail in the Age of AI." That framing used to be aspirational. This year it was descriptive: agentic commerce moved from the futures track to the main stage, and the announcements reflected it. Here is what was actually said, and why it matters. ## The Gap/Gemini announcement is a data readiness story, not a partnership story Gap Inc. became the first major fashion retailer to enable direct checkout inside Google Gemini, using Google's Universal Commerce Protocol (UCP). Shoppers can now discover, evaluate, and buy products from Old Navy, Gap, Banana Republic, and Athleta without leaving the AI platform. Google Pay handles the transaction. Gap stays the merchant of record. The announcement got substantial press. Most of it focused on the wrong thing. The detail that matters is this: Gap's product information inside Gemini is not crawled from their website. Gap provided it directly to Gemini in advance. The brand controls what data the agent sees. The brand controls the product attributes, the pricing logic, the comparison data, and the trust signals the agent uses when it evaluates Gap products against alternatives. That is the entire game. An agent does not crawl. It queries structured, pre-provided data. The brand that has done the work to supply that data clean, complete, machine-readable, directly integrated with the platform controls the recommendation, the comparison, and the conversion. The brand that has not done that work is not invisible because agents dislike them. They are invisible because there is nothing for the agent to query. Gap's CTO Sven Gerjets said it directly: "It's not just keyword search anymore. It's conversations, and we need to be relevant to that." The relevance is not a content strategy. It is a data architecture. Two additional details from the Gap announcement worth noting: First, Gap's loyalty accounts are not yet connected to the Gemini checkout flow. That is the next infrastructure gap and a meaningful one. Discovery infrastructure and transaction infrastructure are not the same build. A brand that has done the data work to be findable inside Gemini has not automatically done the integration work to make loyalty redemption, personalized pricing, or account-level preferences accessible to the agent at checkout. Those are separate layers with separate requirements. Second, UCP now supports loyalty integration, multi-item carts, and real-time catalog updates. Those are exactly the capability gaps that broke OpenAI's Instant Checkout in 2025\. The infrastructure matured. The brand readiness question did not change; it became more specific. ## Shopify made the same argument from the platform side The same week as ShopTalk, Shopify rolled out Agentic Storefronts to all merchants by default. The talk on the floor got specific in ways that previous Shopify AI announcements have not. Agents parse metafields, not HTML descriptions. If your product specifications, materials, dimensions, or compatibility data live only inside a description field written for human readers, an agent cannot extract structured information from it. The data needs to be in metafields, tagged with standard namespaces, and structured for machine parsing. Agents rely on schema markup to understand your catalog. Schema.org Product, Offer, and Organization markup is no longer just a search ranking input. It is the machine-readable layer that allows an agent to correctly evaluate and compare your products against alternatives at query speed. Agents require API-accessible checkout. A checkout flow that requires browser navigation, clicking through a cart, filling a form, and completing a payment page cannot be executed by an agent. Headless architecture and API-first checkout are no longer just technical modernization projects. They are the infrastructure that makes your products transactable by agents operating on behalf of consumers. The framing Shopify is now pushing to merchants: you are building for two audiences simultaneously. The human buyer and the agent querying on their behalf have different infrastructure requirements, and most storefronts are currently built for only one of them. That is the split experience. And it is not a future optimization task. It is a current visibility problem for any merchant whose product data is not structured for agent retrieval. ## The protocol divergence is worth tracking Sephora unveiled an AI-powered shopping experience inside ChatGPT at ShopTalk. The timing creates an interesting contrast. OpenAI pulled back from checkout intermediation and pivoted to being a discovery and referral layer. Merchants receive high-intent traffic; transactions close on the brand's own storefront. The agent qualifies the buyer and hands them off. Google is moving in the opposite direction with UCP toward in-platform checkout, loyalty integration, and real-time catalog access. The agent qualifies and closes. Two different protocol bets, playing out simultaneously. Both require structured product data, API integration, and machine-readable catalog architecture. The infrastructure requirement is identical. What differs is where the transaction completes and who controls the checkout surface. For brands, the practical implication is that building agent-readable infrastructure is not a bet on one protocol winning. It is the prerequisite for participating in either model. The brands sitting out the infrastructure build are not neutral on the protocol question. They are excluded from both. ## The Coresight Day 1 summary named the gap precisely Coresight Research, covering ShopTalk's Day 1 sessions, put it plainly: agentic commerce "remains in early stages, limited by gaps in infrastructure such as universal carts, APIs and standardized product data." That sentence is an accurate description of where most brands are, not where the technology is. The technology has moved. UCP is live. Shopify's Agentic Storefronts are live. Gap is transacting inside Gemini today. The gap is not the platform. The gap is the brand infrastructure that has not been built to connect with it. The conference theme was "Retail in the Age of AI." The subtext of every major announcement was that the age of AI is already here for the brands that built the infrastructure, and still theoretical for the brands that didn't. ## What ShopTalk 2026 actually signals for brands ShopTalk is where platforms announce to decision-makers with active technology budgets. The announcements this year were not previews. Gap's Gemini integration is live. Shopify's Agentic Storefronts are live. UCP is in production with Gap and rolling out to additional retailers. The window between announcement and execution is where the readiness work happens. The brands at ShopTalk talking about agentic commerce in 2026 are the same brands that will be measured against it in 2027\. The ones that leave ShopTalk with a clearer picture of the infrastructure requirements and move on them are positioned to be the Gap in that story. The ones that file it under "interesting developments to watch" are positioned to be the unnamed competitors Gerjets was gesturing at when he noted that no other major apparel brand had announced a similar Gemini partnership. The readiness question is not whether agents will mediate commerce. That question was answered at ShopTalk this week. The readiness question is whether your product data, your catalog architecture, your checkout infrastructure, and your loyalty data are accessible to agents operating on your customers' behalf. If you do not know the answer to that question, the answer is probably no, and the cost of that answer is accumulating now. ## ### The Read on OpenAI's Checkout Pullback Is not Accurate URL: https://www.agenticlandmark.com/brief/the-read-on-openais-checkout-pullback-is-not-accurate/ Last updated: 2026-06-29T20:28:37.000Z OpenAI shut down Instant Checkout in early March 2026. The take: agentic commerce hit a wall. Brands have more time. The technology isn't ready. That read is wrong, and acting on it is expensive. Here is what actually happened. ## What failed was a business model, not a technology OpenAI tried to become both the pipe and the toll booth. Build the discovery layer, yes, but also insert itself as a checkout intermediary, take a cut of every transaction, and become the operating system of commerce. That specific model failed. The discovery layer worked. Approximately 50 million shopping-related queries happen on ChatGPT every day. Users were researching products, comparing options, and forming purchase intent, all through ChatGPT. That part of the thesis was validated at extraordinary scale. The checkout intermediary model did not work. At the point of shutdown, fewer than 30 of Shopify's millions of merchants had successfully integrated with Instant Checkout. Not because merchants lacked motivation. Because product data across the internet proved too messy, too unstandardized, and too fragmented to allow an AI agent to reliably automate checkouts without triggering financial errors, inventory disputes, and tax compliance failures. Sales tax collection was not built. SKU variation across catalog sources broke the comparison logic. And Forrester's March 2026 consumer research found that completing a purchase inside an answer engine is the least-adopted behavior among regular AI platform users, behind general research and product comparison. *The toll booth model failed. The protocol survived. The discovery function survived. The high-intent shopping traffic survived.* ChatGPT is now explicitly a discovery and referral engine that routes high-intent shoppers to merchant sites, which is, for most brands, the more defensible model anyway. ## The brand readiness requirement did not change This is the part that gets lost in the "agentic commerce is slower than expected" read. *The transaction venue changed. The readiness requirement did not.* A brand without structured, agent-readable product data was not ready for Instant Checkout in September 2025\. That same brand is not ready for the discovery and referral model today. The agent still queries your product data. It still evaluates your catalog against competitors. It still needs machine-readable attributes, pricing logic, and trust signals to include you in the recommendation. The failure of the checkout layer does not buy brands more time on the infrastructure layer. It just changes what the infrastructure needs to connect to. ## The failure revealed something most brands missed entirely The cohort of merchants that successfully integrated with Instant Checkout was almost entirely composed of brands already inside Shopify's ecosystem brands with structured catalog infrastructure and an active integration pathway. Brands operating on WooCommerce, Magento, PrestaShop, custom stacks, or headless architectures had no equivalent on-ramp. They were not blocked by motivation. They were blocked by the absence of a normalized, agent-readable product data layer. That gap does not close with the shutdown of Instant Checkout. It becomes more consequential as ChatGPT, Gemini, and Copilot route high-intent discovery traffic to brands that have built a readable infrastructure and route around brands that have not. The merchants who were live when Instant Checkout launched were the ones who had already done the product data work. That's the actual finding. Not that the technology failed, but that clean, structured data was the rate-limiting constraint, and most brands don't have it. ## What this means practically The agentic commerce protocol OpenAI built with Stripe is still live. Google's Universal Commerce Protocol, co-developed with Shopify, Walmart, Target, and a coalition of major retailers, is live. Microsoft Copilot's shopping integration is live. Perplexity has been routing commerce transactions since before Thanksgiving 2024. The infrastructure race did not pause. The consumer adoption curve is running behind it, which is normal at this stage of a platform transition. The brands building agent-readable product data, clean catalog architecture, and protocol-compatible infrastructure are not early adopters of a failed experiment. They are early movers on the infrastructure layer that every agentic commerce surface checkout or referral requires. The brands reading the shutdown as a signal to wait are making the same mistake brands made in 2011 when early mobile commerce stumbled. The platform matured. The brands that had built mobile-first infrastructure captured the growth. The brands that waited retrofitted under time pressure and competitive disadvantage.