The paper behind The Agentic Selling Manifesto

The seller side of a market where people buy in conversation, agents act under mandate, and the page composes itself around the person asking.

Gianmaria Monteleone, CEO and Founder of Veliu

September 2026

Figures as of the date of publication; sources in the References.

Abstract. Commerce is acquiring a second kind of buyer and a second kind of surface. Software agents now read, compare and transact under a person’s mandate, and an interface can be composed by a model at the moment of a visit, out of what that visitor said and did. The infrastructure serving the machine buyer was built in twelve months. The discipline serving the seller does not yet exist. This paper gives that discipline a name, agentic selling, and decomposes it into three problems: measuring a company’s presence across the conversations its buyers actually have, representing an offer in a form a machine can verify, and acting on the company’s own systems under an auditable mandate. Two arguments run underneath it, and the agentic-commerce literature leaves out both. Delegation is bounded by how people want and by what buying means to them, which leaves every seller with two buyers to serve on different terms. And the attention that sellers have always traded in is being redefined by systems that listen, while the seller’s own surfaces stay fixed. Our evidence is operational: more than a year building and running an agentic marketplace in which people delegate buying and selling to agents, with more than fifty thousand customers, more than three hundred thousand delegated transactions and more than one hundred million euros in generated transaction value. We report what that corpus taught us, state the five claims the discipline rests on together with what would weaken each, describe what we are building, and close on the two senses of agency that the next few years will settle.

I

The end of navigation.

A shopper reaches a brand’s site after a long conversation with an assistant that heard her needs, compromises and worries, answered in her register and kept the thread alive. Then the site opens and the conversation disappears. The homepage addresses an average visitor; navigation asks her to translate the need into categories, keywords and filters. The brand knows its product best, yet behaves like the least informed participant.

The discontinuity comes from the buyer, the page and the transaction itself. First, people ask. A buyer who once typed keywords now speaks to a system that answers and asks a second question. Licklider described the partnership in 1960, before a network existed to run it on: the person sets the goal; the computer does the routine finding and comparing [1]. Sixty-five years later it arrived in a chat window.

Second, the page has stopped being fixed. For twenty years a website was a set of documents to traverse. A model can now output a structured description of components, streamed to a renderer and composed into a page from the visitor’s words and behavior. The industry converged on this in successive releases: a framework vendor popularized generative UI in early 2024, an open specification for agent-driven interfaces arrived at the end of 2025, renderers for the main frameworks followed in the spring, and assistants opened their surfaces to third-party interactive components [2]. A page composed for one person and one intent is an ordinary engineering object.

Third, agents transact: within twelve months the buy button moved inside chat, open protocols connected agents to merchants, channel fees appeared, and merchant-named agents surfaced inside the platforms, configured from merchant feeds [3]. This is the most discussed of the three and the one that matters least here. The harder questions are the ones underneath: who understands the buyer, and who represents the seller?

Buying online has required translation work on every kind of site, whether or not a cart waits at the end. A person arrives with a need: protection for the knees on long downhill runs, under two hundred euros, for wide feet; a coat for a winter spent outdoors; a payroll system for two countries and forty people; a law firm that has argued this kind of case before. The site supplies a navigation bar, a search box and someone else’s sorting logic; the buyer chooses a category, guesses a keyword, toggles six filters and opens eleven tabs. Hutchins, Hollan and Norman described two gulfs in 1985: execution, between intention and available action; evaluation, between what the system shows and what the person needs to know [4]. Direct manipulation narrowed both for people who knew what they sought, leaving the person who arrives with a sentence to translate it herself. The dropdown, category page, filter and keyword were the cost of talking to a system that could not understand ordinary language.

  1. The dropdown becomes a question.

    The need, spoken once, in the buyer’s own words.

Category
Running
Trail
Road
Track
Walking

something to protect the knees on long downhill runs, under two hundred euros, wide feet

Now the system reads the buyer’s language and does that work. A product can be presented one way to a marathon runner and another to someone protecting an injured knee; a service can answer both a founder equipping a team of ten and an enterprise in mid-migration; a law firm can lead with the case closest to the visitor’s question. These were always different questions that the grid asked in one way. Self-service becomes service in the older sense: the attention of someone who knows the subject. For two decades the buyer compiled a need into the system’s terms; now the system does that thinking, so the same catalog answers questions its navigation never anticipated.

A brand is admitted to the answer, or absent from it.

That reversal changes commercial presence. Navigation supplied structures a seller could occupy, such as shelf positions, pages, rankings and keywords to bid on; an answer offers none of them. Simon warned in 1971 that a wealth of information creates a poverty of attention [5]. The answer is the form that poverty finally takes: the shelf compresses into a sentence, the consideration set collapses from hundreds to a handful, and a brand is admitted to the answer or absent from it.

Roughly a third of the content on an average retail product page is invisible to the models doing the recommending.

38

percent below non-AI channels, a year earlier

42

percent above them, twelve months later

AI-referred retail traffic, converting against non-AI channels.

An answer is assembled from what a model can read and verify: it reads feeds, fetches pages, extracts parsable attributes, traces claims, checks price and availability, and discards what fails. Four classes account for most failures we observe: attributes trapped in prose or images, stale availability, claims without machine-checkable sources, and checkout flows that break under delegated sessions [6]. Recommending models cannot see roughly a third of a U.S. retail product page [7]. That invisibility flattens offers where the details sell, just as demand arrives: within twelve months, AI-referred retail traffic moved from converting 38 percent below other channels to 42 percent above them [7]. Commercial attention is concentrating the way it once did in a shop window.

An argument about a mechanism should say what would refute it, so here is ours. If assistant use grows over the next two years while the share of considered purchases beginning in conversation stops rising, conversation is only a research habit. If conversational answers preserve consideration sets as wide as the grid’s, concentration is a passing effect of immature retrieval. If the sellers winning answers already won the shelf, the interface has changed over an unchanged market. All three are measurable, and the benchmark in section X is being built partly so others can measure them.

II

From search to delegation.

Economics anticipated this in 1961, when Stigler formalized costly search: buyers stop when the expected gain from another comparison falls below its cost; market structure follows who bears that cost [8]. Category pages, rankings and keyword auctions sold shortcuts to buyers paying per comparison. Delegation changes the kind of cost: another comparison costs almost nothing at the margin, while searching well becomes a fixed cost, paid in building and maintaining the agent.

Search costs once sustained price dispersion; a cheap-sampling agent narrows it among offers comparable on the attributes it can read. Bakos showed in 1997 that cheaper search intensifies competition among comparable offers and raises returns to differences buyers can evaluate [9]. I expect prices to approach marginal cost where offers are comparable, while Chamberlin’s pricing power [10] survives where a difference is real, legible and valued. A real but unreadable difference keeps its human value and is worth zero to the agent.

Cheap search still leaves one cost in place. Diamond proved in 1971 that, with identical buyers searching sequentially for a homogeneous good, an arbitrarily small positive cost of another quote makes the monopoly price the unique equilibrium [11]. One friction remains: doubting your own agent always costs something.

A buyer tends to settle on one assistant, while sellers must appear on all of them. Armstrong’s competitive bottleneck predicts the result when one side single-homes and the other multi-homes: platforms compete for the buyer and extract from the seller [12]. Assistant checkouts so far charge the merchant a fee on completed purchases while the buyer pays nothing [3]. The fee’s incidence fits the model; its ceiling depends on the seller’s next best channel.

Entry brings a further complication. Cheaper search lowers a small seller’s cost of being found, then imposes fixed costs that barely vary with volume: a machine-verifiable representation and an agent to run it. Those costs fall hardest on the smallest sellers. Entry becomes easier where a difference is real and cheap to verify, harder where a position rested on distribution, shelf presence or bought attention.

Delegation also moves the governing theory from search to agency [13]. Jensen and Meckling separated agency costs into monitoring by the principal, bonding by the agent, and residual loss: the welfare sacrificed when the agent’s choice diverges from the principal’s best outcome. Monitoring and bonding are being engineered publicly through spending caps, instant revocation and clear recourse, the controls consumers require for trust [14]. Residual loss, by contrast, is nearly invisible to whoever bears it: a buyer who receives a good coat never sees the better coat the agent skipped. No receipt is issued for a foregone alternative.

Residual loss has a twin on the seller’s side. No seller learns that a buyer asked for exactly what it makes and saw three competitors, or that an agent reached its site, failed to read the attribute the comparison turned on, and left. Its analytics hold at most a line in a log with no question attached; there is no abandoned cart, because no cart was created.

When neither side can see what it lost, dissatisfaction cannot correct the market. Discipline has to come from elsewhere: claims checked when made, records that someone outside the transaction can audit afterward, and sellers answerable for what their agents say. Verification lets quality earn a return even when loss remains silent. Sellers who read the standards work on agentic commerce as compliance miss the mechanism by which their quality can still be told apart.

The decisive question is whose mandate each agent carries. A buyer’s agent carries the buyer’s; an assistant’s agent balances buyer outcomes against the economics of its host surface. A brand can accept the agents some platforms now offer in its colors [3], as long as it knows whose rules govern them and where their telemetry stays.

The demand side has protocols, mandates, payment rails and carts that travel between assistants. The supply side still has little independent representation: agents speaking for sellers inside these surfaces carry the surface’s mandate. The conditions for the seller’s own agent now exist: a brand agent is a model with tools and an objective, given the brand’s knowledge, access to its systems and a perimeter of action. Tokenized agent payments and the emerging rails for buyer identity and authentication supply the trust infrastructure [15]. The missing discipline must begin with an honest account of delegation’s limits.

III

Why agents will not do all our buying.

The agentic-commerce literature often assumes delegation spreads to everything once it becomes possible. I disagree: sellers will serve two buyers, because of how wanting works and what buying means.

Delegation begins with a statable preference, which the literature assumes into existence too quickly. Simon described satisficing under limited attention and knowledge [16]. Slovic’s review of thirty years of experiments goes further: preferences over unfamiliar or complex options are often constructed during elicitation and change with its method [17]. A printer cartridge’s purpose is settled before the search begins. A buyer builds the want for a coat, a wine or a room for an anniversary while looking; part of the value is produced there. This year’s surveys remain consistent with that distinction: three quarters would delegate routine tasks, a third would accept decisions within limits, and fewer than one in ten would accept autonomous purchases [14].

The strongest reply is that agents can estimate preferences people cannot state. After twenty-five years of recommenders inferring taste, “you know what I like, choose for me” is a valid mandate. Stated attitudes predict future conduct poorly: commerce has repeatedly crossed thresholds people said they would refuse. Where nobody chooses, the default decides; the surfaces write the default. The reply is right about articulation, but the constraint lies elsewhere.

A preference can be inferred while responsibility stays with the person. Some buyers want to have chosen; a perfect estimate still hands them an object they did not choose. Residual loss also makes adoption a poor measure of fit: someone unable to detect a bad delegated choice cannot learn from it. Delegation may overshoot into unsuitable territory while the adoption curve shows nothing; correction would arrive through visible failure or regulation, long after the quiet losses begin.

I make a narrower claim: the boundary runs between delegating work and delegating experience. Most labor in a considered purchase goes into assembling worthwhile options; the choosing takes comparatively little. A system can remove that labor and hand back the looking, reacting and revising, preserving the activity that constructs the want. This is why the interface matters as much as the agent. The claim fails if autonomous completion spreads as readily to purchases people describe as identity-bearing or relational as to routine replenishment, holding capability and controls constant.

Anthropology explains what is at stake. Douglas and Isherwood treated goods as communication, making cultural categories visible and stable [18]. Miller’s North London ethnography found ordinary shopping to be an act of care: thinking about household members through what one chooses for them [19]. A gift, a coat or a journey can carry authorship, so delegating the selection changes the meaning. A gift chosen by an agent is a different gift even when its recipient receives the same object. That authorship belongs to the act of choosing and cannot be recovered from how well the object fits.

The boundary crosses categories instead of following them. One bottle of wine can be restaurant inventory, a hurried contribution to dinner, a signal of intimacy or an evening learning about a region; the same bottle can serve all four in one week. Its purpose that day decides whether an agent can be sent for it.

one bottle of wine

restaurant inventory

send the agent

a hurried contribution to dinner

send the agent

a signal of intimacy

go yourself

an evening learning about a region

go yourself

The same bottle can serve all four in a single week.

Where delegation occurs, the buying agent applies a mandate consistently across more options than a person could hold in mind, and rewards what it can read. In December 2025 Anthropic ran an experiment in agent-mediated trading: sixty-nine employees traded through Claude agents for a week, across four parallel marketplaces, with a hundred dollars each. Agents listed more than five hundred items on their own, made offers, negotiated in natural language and closed 186 deals [20]. Stronger models earned more per sale on identical items, while people represented by weaker ones failed to notice and rated fairness the same. Aggressive negotiating instructions had no significant effect.

I read this as capability asymmetry: model quality moved prices where human bazaar tactics did not, without the principals noticing. Participants rated fairness equally despite unequal results, making residual loss tangible: satisfaction alone cannot establish whether someone was well represented.

Pricing power has lived in differentiation since Chamberlin located it there in 1933 [10]. A buying agent prices an unreadable difference at zero; when the same few models serve most buyers, that blind spot reaches every seller at once. Kleinberg and Raghavan’s algorithmic monoculture can lower social welfare even when each decision-maker prefers the shared evaluator [21], so differentiation becomes harder where its return is highest.

The limit on delegation therefore doubles the brand’s work. Serving buying agents starts with readable data and advances to winning on quality. That work belongs to a brand agent: it answers the buying agent’s questions, supplies evidence for the brand’s claims and, as protocols mature, argues its case. Without one, the buying agent rewards superiority only where it can parse it; flattening follows. Serving people requires attention to the experience they retain, now judged against an assistant that has already listened.

IV

Pleasure, and the collision.

Across much of buying, pleasure precedes use and lives in the approach. In 1997 Schultz, Dayan and Montague showed dopamine neurons encoding the difference between expected and received reward: once a cue reliably predicts a reward, the activity moves to the cue [22]. Berridge and Robinson distinguished wanting from liking: dopamine drives incentive salience, the pull toward a thing, while enjoyment depends on other systems [23]. In 2007 Knutson’s team showed products, then prices, to people in a scanner. Nucleus accumbens activity rose when a desirable product appeared, before any decision; insula activity rose when the price was too high. Those two signals, anticipated pleasure against anticipated loss, predicted the purchase before it was made [24].

I take less from these findings than they are often made to carry. Inferring a mental state from activation is weak evidence when the region involved responds to many things [25]. Nothing here establishes that a responsive interface produces dopamine. The findings support a narrower proposition: the mechanisms operate during approach; a purchase can be made more attractive without being made more satisfying. The decision itself includes what appears, when, and beside which price, so restaging that sequence for each person acts on an input to choice. This interface hypothesis should be tested through comprehension, correction, completion and later regret. Restraint belongs inside the architecture: an optimizer of wanting will find the gap between wanting and liking and sell into it.

This is the physiology of the shop window, where wanting is staged through cues, anticipation and attention. Campbell described modern consumption as imaginative daydreaming, with objects enjoyed before, and often more than, possession [26]. Holbrook and Hirschman’s experiential account recognized the fantasies, feelings and fun that informational models miss [27]; Pine and Gilmore treated the experience itself as an offering firms can charge for [28]. Boutiques, product demos and waiting lists stage desire; so does an assistant who remembers your name. A brand largely sells attention while the want forms.

Assistants already supply much of that attention through a plain interface. They listen, answer the actual question, match a specialist’s or a beginner’s register, and stay engaged through the second question and the fifth. The physiology above licenses no claim about warmth. What matters is the second turn: the assistant asks, receives an answer and changes its response. The buyer takes part in shaping the approach to purchase, the phase in which those reward mechanisms operate.

More than four in five U.S. consumers who report shopping with AI say their experience improved [7]. The self-report is evidence of demand and says nothing about mechanism. It reveals a low bar many brand sites fail to clear: a page that cannot take a second turn cannot join the approach.

Where that holds, the wanting is staged on a surface the brand does not own.

The brand’s home should amplify this attention with unmatched knowledge of its products and customers. Often it offers pre-printed sentences, a grid and keyword search, with no way to answer, advise or attend to the person. A chatbot in the corner imitates that experience poorly: it knows less, presses harder to sell and cannot change the page. So the brand’s surface is the poorest in the journey at the moment the buyer is most open to attention. For twenty years the site hosted the experience and search was its corridor; now the corridor listens and the destination stays fixed. Where that holds, the wanting is staged on a surface the brand does not own; its site receives the visitor after most of the choosing is over.

What would a page look like if it could attend to the person reading it? It composes itself. A shop assistant reads a customer and rearranges the shop around them, a form of selling that never scaled; a generative interface performs that reading in software. From the visitor’s sentence, or the way they scrolled, it assembles the brand’s materials around the question: the two countries and the forty people for a founder pricing payroll; the previous coat for someone replacing one worn for ten years; the last case for a returning client. The images and words remain the brand’s, so buying becomes mutual adjustment: the buyer speaks, the page composes, the buyer answers back.

Lookup

Someone anticipates the segments, writes the variants, and the system selects one.

Five pages. Another one costs a working day.

Composition

Typed components and containment rules define a space of pages whose possibilities need no enumeration.

A previously unseen page costs a model call.

Adaptive pages have existed for twenty years, so the technical distinction matters. Their familiar form is a lookup: someone anticipates segments and writes variants; the system selects among them. The outputs are bounded by what that person imagined; another variant costs a working day. Composition generates over a grammar: typed components and containment rules define a space of pages nobody has to enumerate, where a previously unseen page costs a model call. Manovich identified this variability two decades before the tools arrived: a computational object lives in versions, with no single instance [29].

Quality control moves accordingly, because enumerated variants can be inspected one by one. A generative system has no finite set to review in advance, so governance has to address the grammar and its invariants: which components exist, which sequences are permitted, where claims come from and what every displayed page must satisfy. Governing the interface means governing a language, for which no public discipline exists yet. Inspecting one rendering says nothing about the grammar that produced it.

Selection, order and emphasis vary, while the offer remains canonical: claims, product facts, approved assets, public terms, safety information and checkout. Every composition answers a question, with a provenance trail to that source.

The limits of a composed page begin in engineering and do not end there. Such a page can manipulate sequence, urgency and salience with a precision a fixed page lacks; it can also invent claims, drift from the brand’s voice or produce an incoherent visual state. The first limit is technical. Section X describes the mechanisms that contain it: a bounded grammar, source discipline, confidence gates, logs and human approval.

The second limit is normative. A page can quote what a visitor said elsewhere, stay accurate, and still violate the expectations attached to that disclosure. Nissenbaum’s contextual integrity asks whether an information flow suits its context, a stricter test than lawful possession [30]. Knowledge used quietly can feel like attention; show the same knowledge on the page and the visitor learns she was watched. The page may use an injured knee she named in her own request; the same fact inferred and carried between visits is profiling. Financial distress, signals concerning minors and protected traits never drive adaptation, whether stated or inferred. Visitors should be able to inspect the signals in use and reset them.

The third limit is structural and belongs at the center of standards work. A page that never repeats is difficult to constitute as a public object: a fixed page can be linked, quoted, archived, compared by competitors, sampled by regulators and produced in evidence. Consumer protection often assumes that buyers saw the same thing; generation removes that assumption. Nobody is going to stop generating to solve this, nor do a vendor’s own logs settle it.

Standards bodies and protocol authors need a convention for the witnessed page. Re-running a model on the same inputs does not return the same page, so the convention has to be a retained record: the output as served and rendered, the versions of component library, grammar and model, and the signals it received. Under a stable identifier, someone absent from the original encounter can find that record and read it. We know of no widely adopted convention of this kind; it is the standards work I would put first.

V

The name of the missing half.

We call it agentic selling, the seller side of agentic commerce [31]. The conditions are new: people converse, pages compose themselves and agents transact. Under them lies one old job: study the market, ready the offer, sell in conversation and be chosen. A brand agent does the work, one per brand, carrying that brand’s mandate and no one else’s.

The craft predates all its technologies. Persian first named the person matching buyer and seller; Arabic carried simsar along trade routes; Italian received sensale, broker of everything from grain to marriages [32]. Each era gives this figure its technology; ours gives it software and the scale of every purchase on earth.

History keeps the defining question: whom does the broker work for?

VI

The three problems.

The first problem is measurement. A brand’s presence is a distribution across the conversations in which its buyers actually ask: how often it enters the answer, how it is described, how often it is chosen and where the purchase breaks down. Estimating it requires a corpus with an explicit sampling method: a laboratory prompt samples its author’s idea of a buyer, a clickstream panel its own exhaust. Consented conversations with profiled buyers sample the market, under a method that has to be published.

Real conversations anchor the corpus; the best one combines first-hand data with real buyer personas, improving synthetic coverage where recruitment cannot reach and training the observation models that read it. Amazon research found synthetic-persona training could outperform training on real, de-identified data for buyer-signal learning [33]. The code was never released, so we built a reproduction [34], described in section X. Published schemas and real-distribution checks make synthetic personas instruments whose assumptions stay inspectable.

The corpus must also distinguish shopping surfaces, which answer from structured feeds and refresh quickly, from answer engines, which draw on citations and refresh slowly; pooling the two obscures both.

  1. 1Measurement.

    A brand’s presence is a distribution across the conversations in which its buyers actually ask.

how often the brand is admitted to the answer

how it is described

how often it is chosen

where the purchase breaks down

A distribution, over the conversations in which the brand’s buyers actually ask.

The second problem is representation. Buying agents require exact, structured, verifiable data at machine speed and abandon whatever is ambiguous. Every seller has an offer, whether it is sneakers, insurance policies, software seats, billable hours or a plant’s capacity for next quarter. Its machine-verifiable form is the catalog, with attributes structured at source, fresh availability and claims linked to checkable evidence. This is merchandising for a reader that demands verification.

The generative interface uses that same representation, since neither composer nor buying agent can read an attribute available only in a photograph. A shared source keeps the machine answer, the human page, the feed and the checkout aligned; separate representations let them drift.

The third problem is action. A brand agent holds a mandate over live systems and acts within defined limits. Autonomy is a control problem: how much it may do unattended, and what record it leaves. Emerging standards concern identity, provenance and records [15].

The problems feed each other: measurement finds failures or openings, representation changes the material available, action applies it to live systems or to buyers, and new conversations supply the evidence.

one agent
per brand

studies

the market

readies

the offer

sells

in conversation

VII

What more than a year in the field taught us.

These claims began as operating notes. Before naming the category we spent more than a year building and running a consumer agentic marketplace where buying and selling are delegated to agents, with millions spent on research and development. Its scale is stated as lower bounds: more than fifty thousand customers, more than three hundred thousand delegated transactions and more than one hundred million euros in generated transaction value [35].

When Anthropic published Project Deal in April 2026 [20], ours had been live for more than a year, with hundreds of thousands of transactions already delegated. The two bodies of evidence complement each other: Project Deal is controlled evidence about model capability and bargaining, ours longitudinal evidence about delegation, transaction failures and trust under repeated use.

0+

customers

0+

delegated transactions

0M+

in generated transaction value

A composite of recurring failure classes makes the mechanism concrete. A buyer asks for “trail shoes for long downhills, bad knees, wide fit, under two hundred euros, need them by Friday.” Three comparable products face different outcomes.

trail shoes for long downhills, bad knees, wide fit, under two hundred euros, need them by Friday.

Three comparable products face different outcomes.

One

cushioning data, trapped inside a product image

It fails at admission.

Two

“improved stability”, a claim with no source a model can check

It fails at verification.

Three

structured attributes, a stack-height figure, a returns policy the agent can quote, stock information that answers for Friday

It is recommended, and it closes.

  • The first never enters the answer, because its cushioning data exists only in an image.
  • The second enters but loses the comparison when “improved stability” lacks a checkable source and is dropped.
  • The third supplies structured attributes, stack height, a quotable returns policy and stock data that confirms Friday delivery, so the agent recommends it and the buyer takes it.

Such a conversation exposes a mechanism without estimating its frequency; across cohorts and weeks, the same observations recurred and shaped our instrument. We would need samples designed to establish how often they occur and whether they transfer.

1.Delegation is gated by trust, which is granted in steps.

People delegate a task, then a larger one, eventually a standing mandate. Logged actions, honored limits and reversed mistakes advance that progression. Trust behaves like credit, extended on record and withdrawn on default; surveys likewise identify spending caps and instant revocation as conditions [14].[14].

2.The strictest buyer a seller has ever faced is a machine that fails in silence.

Failures cluster in familiar catalog gaps that human sellers long bridged; what is new is that they leave no trace. A visitor defeated by a pricing page leaves searches, filters and a bounce in the seller’s analytics, while an agent rejecting an unparsable attribute leaves at most a line in the logs. Measuring that lost demand requires observing the buyer’s side of the conversation.

3.Real conversations differ in distribution from invented prompts.

Our conversations contain budgets, constraints, trade-offs and second questions that prepared prompts tend to omit. Comparing the corpora exposed the difference and led us to anchor synthetic material to real distributions. The benchmark has to settle how big that difference is and whether it transfers.

Threats to validity. Our corpus comes from one agentic marketplace, weighted toward consumers, with a category mix we chose, so regularities may fail to transfer to other verticals or enterprise buying. Metrics count only what our instrumentation sees, over observations spanning more than a year while engines change monthly. We claim recurrence in the cohorts examined; the benchmark in section X is meant to put those observations on ground anyone can check.

VIII

Five pillars.

Each maps to a benchmark measurement or to a property of the architecture in section X, and states what would weaken it. None is settled.

  1. 1

    Connection is necessary; selection is the contest.

    Commerce protocols connect a catalog to agents that speak them, but each recommendation still has to be won, first through readability and then through verifiable difference.

    A seller implements interoperability once and competes for selection in every conversation. The pillar weakens if, among connected catalogs, price and protocol adoption alone explain share of recommendation.

  2. 2

    The answer is the new shelf, and every answer is an adjudication.

    Buyers meet brands in answers before they visit their sites; candidates are adjudicated before anyone sees them.

    The retrieval corpus, the reading of the request, the model, the policy rules and the commercial defaults all get a vote. Presence is binary in one conversation and a distribution across many. Measurement must record the path: retrieved sources, governing criteria, points of uncertainty and whether different wording changes the result. The pillar weakens if answers routinely preserve consideration sets as broad as the grid’s, or their paths prove too unstable to audit.

  3. 3

    Two arts, born separate, converging into one.

    Human buyers respond to story and attention, while agents require verifiable structure: a seller must practice both crafts on the same offer.

    I expect the two crafts to converge as buying agents learn to interrogate brands, negotiate and be persuaded. Agents already negotiate with each other in natural language [20], research systems have done so for years [36], and controlled experiments show language-model agents responding to some of the tactics that move people [37]. Convergence needs more than negotiation: an open question, an answer that carries evidence, and an offer built inside the conversation. A signed quote must carry terms and expiry, be issued under the seller’s mandate, and attest that the buying agent is addressing the brand [15].

    It also needs a cost for misrepresentation. Crawford and Sobel showed that costless, non-binding messages between parties with divergent interests convey coarse information, with less conveyed as divergence grows [38]. A seller’s agent and a buyer’s agent have divergent interests by construction. Checkable evidence, binding commitments and liability make their exchange informative by adding what costless messages lack; no current commerce protocol imposes that cost. I would count the pillar falsified if, by the end of 2028, no widely adopted commerce protocol carries a free-text question with a cited answer and a signed per-conversation offer.

  4. 4

    The interface becomes an output, and the brand keeps the grammar.

    A model can generate an interface per visitor and intent, validate it against a schema and render it from the brand’s pre-approved components. The fixed page becomes a special case of the composed one.

    A site becomes a set of materials and rules from which pages are made, so review moves to the grammar and the invariants: there is no finite set of pages to inspect in advance. The design constrains invention: the page re-assembles what the brand has said, shown and verified, with the source of any claim available on request. The brand retains a record of what one visitor saw, or it has published something nobody can examine. Measures include task completion, correction rates, interaction depth, fidelity checks and how often the deterministic fallback fires. The pillar weakens if composition adds latency and error without adding relevance, or if answers built from approved sources carry as many invented attributes as answers assembled by third parties.

  5. 5

    Authority is granted in stages and answered for at every stage.

    A brand agent acting on live systems needs a mandate before it needs intelligence: a bounded action space, a record and a way back. It observes, proposes, acts with approval, and only then acts alone.

    The brand remains in command, widening the mandate as evidence of reliable work accumulates and narrowing it the week the agent fails. Each stage rests on the record, which must be producible to someone outside the transaction: a log only the seller can read settles nothing in a dispute. The pillar weakens if staged mandates earn no more delegated authority than unstaged ones at equal capability and cost.

IX

The strongest objections.

The platforms will build the seller side themselves.

They are doing so already: Google put a merchant-named agent, configured from the merchant’s feed, inside Search in January [3]. In September Anthropic released a shopping agent and a merchant agent as open reference code, free to customize, with no platform for sale behind it [39]. Both answer from the merchant’s own data, so the harness is now free to any seller.

What a buyer asks is often outside that data: why this material rather than the cheaper one, what the brand will guarantee, what its people learned building it. Reference code carries none of that. The contest moves to what fills the harness: the brand’s own intelligence, its history and what stands behind its products and services, compiled into a form its agent can use and keep current. Section X’s fifth source is where that intelligence comes from.

Section II’s mandate problem remains: an agent governed by a surface’s rules and feeding it telemetry carries that surface’s mandate. Distribution belongs to the platforms; the seller keeps its representation, deeper knowledge and perimeter of action across its systems.

The assistants will absorb both jobs and leave the seller with a feed.

This is the bear case: models learn to read bad pages, representation loses value, and buying stays inside assistants. The site becomes a place to confirm the brand exists. Better retrieval will probably solve some of our four failure classes before catalog improvements do.

Wherever buying happens, the seller still has to know how engines describe it, keep its offer true and machine-checkable at source, and act on its systems under its mandate. All three jobs survive a change of surface. When an assistant does all of them, the brand’s knowledge of its own demand accumulates on the assistant’s side.

People like to browse.

The pleasure lies in looking, wanting and receiving attention; a grid is only one container for it. We expect a page composing itself around a person to provide more of each than forty identical thumbnails, so browsing survives as a dialogue.

If every seller optimizes the answer, the answer becomes spam.

Two decades of search make this the objection I take most seriously. A different equilibrium depends on verification: persuasion scales without limit, but a checkable claim makes fabrication detectable during the transaction and leaves a record. The seller faces a cost where deceiving a crawler once carried none.

Strategic behavior survives and changes shape, as Ellison and Ellison documented when internet retailers made offers harder to compare, with obfuscation associated with lower price elasticity [40]. Expect fabrication where checks are absent, source manipulation where they are shallow, muddled offers where an attribute would lose, and correlated errors where agents share an evaluator. Verification raises the cost of the crude lie and leaves the subtle ones available. Measurement must separate correctness from usefulness, for sellers as well as buyers.

A generated page cannot carry a brand’s identity, and generation is slower than a fixed page.

Both objections are right about naive generation: imitate a brand’s style freehand and you get a forgery; assemble the page from scratch while a visitor waits and you get a slower one. Section X answers with architecture: approved components, a schema, a confidence gate and human approval make fidelity measurable by visual distance, provenance and tone. Deterministic templates supply the floor, product pages are pre-loaded, and a clear intent gets the fixed page whenever composition would add nothing.

A seller’s agent cannot be trusted by buyers.

A seller’s interests have always required scrutiny. Repeated dealings and reputation help only when injured parties detect the injury, which residual loss makes uncertain. A brand agent therefore offers what its counterparty can check during the conversation: sources retrievable immediately, a bounded action space, logs open to an outside auditor, and a seller answerable for the agent’s statements as for its packaging. The sensale’s capital was reputation; its counterpart here is verifiability, which a single unsourced claim can spend in one turn.

X

What we are building.

Veliu is an agentic commerce lab, and agentic selling needs more than one technology, so we are building several that work as one system. Generative interface models compose a page from approved material for whoever is visiting. Protocols set how an agent speaks for a seller and how its words can be checked. Among them is the Brand Voice Protocol, whose first version we will publish: every claim carries its source, every answer stays inside agreed bounds, every exchange leaves a record. In our consumer products, a shopping agent buys and sells on a person’s behalf inside an agentic marketplace. Brand agents speak in a brand’s own name. Some of this is live; the rest will be announced over the coming months.

The brand agent is where the system meets one seller. The reference architecture is one agent per brand, one memory behind two kinds of work. In the back office it studies the market and readies the offer for the brand’s team; on the brand’s surfaces it sells to its buyers.

What it learns from.

First, real buying conversations from recruited buyers: paid, consented, supervised, with profiled personas. Their turns preserve the budgets, uncertainty and second thoughts that single prompts remove.

Second, the conversation stream of our consumer agentic marketplace, where people delegate buying and selling every day.

Third, telemetry from every brand agent we operate, covering organic and paid channels, customer-base signals, e-commerce and social media. On the brand’s site the explicit signals are what visitors write, ask and search, while the implicit ones are a zoomed photograph, a page read to its end, a fast scroll. Competitor signals include a rival product going viral, or an engine beginning to recommend it. The agent also learns which selling proposition converts for visitors from each channel.

Fourth, synthetic personas from our own pipeline [34]. The pipeline walks root-to-leaf paths through a 704-node taxonomy and passes them, with a seven-tier, 130-attribute schema, to a teacher model that produces a coherent fictional buyer. It then role-plays that buyer through six to fourteen dialogue turns at a target emotional intensity, goals and constraints revealing the persona implicitly. Some turns carry no signal at all, as in real traffic.

The pipeline then extracts each dialogue’s implied signals, records the reasoning behind them, and assembles a dataset for distillation into a student model small enough for production. Precision is judged by a model family different from the teacher’s. Recall is measured through a fixed questionnaire answered from the dialogue and again from the extracted signals; every call is metered.

The first two sources calibrate this material: they improve the generator, expose missing attributes and test the plausibility of simulated distributions. Every observation retains its origin, so generated coverage remains distinguishable from field evidence.

Fifth, the brand’s own people: sales assistants, product specialists and founders contribute expertise as training, whether it is the question behind a hesitant customer’s question, why the hinge’s third version was redesigned, or what the brand refuses to make. Every conversation draws on the judgment that makes those people distinctive, whether the brand delivers a service or builds a product. This is the shop assistant of section IV, whose reading of a customer never scaled.

What it fixes, and keeps fixing. Readying the offer means continuous training on what happens inside the brand, across its channels, in its market and among competitors. As catalogs, engines and rivals change, the agent re-compiles attributes, feeds, content, agent-facing pages and structured data. The brand’s visibility is checked daily across the engines with rotating prompts. Competitor comparisons use verified facts only, with the prohibition on invention enforced in code.

A day-zero baseline is frozen before any write, so later changes can be compared against it and retained or reversed on evidence. Where systems permit, fixes are written at source in the space reserved for applications, completing structured data without duplication. Elsewhere they arrive as a costed, prioritized plan. The fixes leave the site as the brand made it.

Identity

establishes which legal seller speaks

Authority

specifies what its agent may disclose, negotiate and bind

Evidence

links claims to versioned sources, including jurisdiction and validity period

Commitment

turns a conversational proposal into an offer carrying terms, expiry and a signature attributable to the merchant of record

How it sells to people. This is the center of the architecture, where the mechanics make one claim checkable: a generated page can carry a brand’s identity.

The input is a visitor’s sentence or behavior. The output is a typed component tree, serialized as data and validated against a schema, using a small closed set of primitives with deterministic templates as the floor.

Before launch, we capture the brand’s site, cluster repeated subtrees, freeze their structure and type their content slots. We distill design tokens for color, type, spacing and radius, preserving color fidelity at Delta E below 3. Each approved component enters a per-brand library labeled by section, audience, offer type, sector and style, with its relations recorded.

During a visit the composer can only select and arrange what is in that library; it produces neither code nor images. One generic, sanitized interpreter renders all brands’ pages, with no per-brand code or runtime code evaluation. Product pages are pre-loaded to remove click latency.

Publication requires every invariant to pass, schema validity and claim provenance among them, plus a composition confidence gate of 95 percent. Below the gate, on a timeout, or when the library holds nothing for what the visitor asked, she gets the deterministic template for her intent, filled from the same approved material. When the whole system fails, the visitor gets the site the brand already had. Review-and-approve remains the default until the brand chooses otherwise.

As the visitor keeps talking, the interface keeps listening: explicit and implicit signals adapt what appears next in real time, in the visitor’s language and market. Within a session each turn carries the state of the one before it, so the third page knows the budget stated on the first.

Deployment takes one line of code on the brand’s site; the URL stays the brand’s, with no redirect. The agent navigates real pages, so traffic shows in the site’s own analytics. Where the offer is bought online, checkout opens on the brand’s own rails, with the brand as merchant of record. Otherwise the agent opens the next step the seller offers. The site adapts: one offer, put a different way for each visitor.

visiblecomparablerecommendedpurchasablecorrect

How it answers agents. Buying agents currently need exact answers, so each brand agent exposes a Model Context Protocol server, structured feeds, agent-readable pages with complete structured data, and support for the commerce protocols assistants use. All answer from the brand’s compiled representation in real time, keeping stock and price as fresh as the site’s. The same components can appear inside assistants accepting third-party interfaces, under their host’s guidelines.

As buying agents learn to interrogate, the brand agent will argue and negotiate within the brand’s limits, which is pillar 3’s convergence. Transport interoperability alone settles none of what follows.

Identity establishes which legal seller speaks; authority specifies what its agent may disclose, negotiate and bind. Evidence links claims to versioned sources, including jurisdiction and validity period. Commitment turns a conversational proposal into an offer carrying terms, expiry and a signature attributable to the merchant of record.

The measurement instrument. We will publish a quarterly agentic selling benchmark built on real conversations and an open methodology [6], scoring five properties per brand: visible, comparable, recommended, purchasable, correct. Purchasable varies by trade: a checkout where one exists, otherwise a quote, a consultation or a trial. The first edition publishes with its methodology; later versions track engines that change monthly.

One data contract runs under both kinds of work: category knowledge compounds across brands while each brand’s knowledge stays its own. Provenance and access controls let auditors see whether category learning accelerates as brands join and whether private knowledge has crossed between them. The autonomy ladder in pillar 5 constrains every action. One system: it learns from millions of buying conversations and sells in yours.

XI

What we expect.

The benchmark’s first edition tests hypotheses stated in advance: machine readability predicts admission to an answer; verifiable differentiation predicts selection after admission; and answers grounded in brand-approved sources contain fewer attributes unsupported by a held-out fact set than answers assembled from third-party sources.

A causal claim about readability needs controlled variants: same meaning, different machine-readable structure. Before collection closes, we publish the sampling frame, engine and model versions, repeated-query policy, denominators and decision rules. We sell what this paper argues for, which is why the hypotheses precede the data. Brands will have their category measured and researchers will contest the method, both in public.

What we expect comes in phases, each dated so we can be wrong.

Phase I, now through 2028

The search moves before the purchase does.

By the end of 2028 we expect most product research to begin in a conversation, with an assistant or agent doing the finding and comparing. Buyers meet the seller inside an answer before anyone opens a site. Delegated buying comes more slowly and unevenly, concentrating where the outcome is easy to state and cheap to be wrong about: consumables, replacements, standardized goods and services, mostly at low and middle prices. Where the choosing carries personal meaning, we expect the agent to prepare the choice and the person to make it.

Phase II, 2028 to 2030

The mandate stands and the arts converge.

Trust earned in steps hardens into standing mandates that replenish, renew and negotiate under rules a principal sets once; buying agents begin interrogating brand agents directly. The first legal disputes over what AI said about a brand make catalog accuracy a question for the people who run it.

Phase III, by 2031

A brand agent becomes ordinary infrastructure.

Within five years we expect most brands to run an agent of their own, as ordinary as a website and for the same reason: without one, a brand is described by whoever has one. The agent has an owner inside the brand and a line in the budget; procurement asks suppliers for one, and agencies build and run them for brands that have none. Commercial pages are generated for each visitor from approved brand materials. Fixed pages survive where a brand wants a monument, a stable public account of itself. The distinction between a site and an agent stops being useful, because the site is now one of the agent’s surfaces.

Beneath these changes lies agency in both senses: software’s capacity to act and people’s capacity to author their choices. Frankfurt described personhood through second-order desires, desires about our desires: wanting to want something else, revising an ordering as well as satisfying it [41]. A mandate of stated preferences holds only the first order: it faithfully executes what someone wanted when they wrote it, even as that person changes their mind about who they are becoming.

Delegation is safe in proportion to how settled a want is. Where wanting is undecided it does damage: an agent can execute an earlier preference exactly and still miss the person who held it. This reaches anthropology’s boundary from the other direction; that boundary guides automation better than any list of categories.

Polanyi’s familiar argument embeds markets in social relations [42]. His double movement supplies the less consoling prediction: exchange disentangles itself from society; society responds through legislation and collective institutions, usually late. Delegation continues that disentangling by moving exchange from persons to instruments. The counter-movement is visible in standards rooms and will reach statutes.

Plurality has an operational meaning there: differences among sellers remain legible and several evaluators remain possible. A person can inspect an answer, recover the canonical offer, reset a composition, revoke a mandate and buy directly. The open question is whether those building the systems will write the first draft of this protection while it can still be technical.

Few things are more human than what we buy, how we look for it, whom we trust to advise us and what choosing feels like. For twenty years we accepted commercial surfaces that could not understand a sentence; that constraint has now lifted. Software can read and compose, holding a consistency no person sustains across a day of visitors and never tiring at the fifth question.

The whole of buying is being rearranged, both what we delegate and what we keep. Its measure will be the plurality preserved and the human judgment that outlasts the help. Every forecast above is dated and checkable by someone who does not work here. Selling has always meant being understood by the one who is asking. That work now has two listeners, and the discipline set out here is how a seller meets both, on purpose and in the open.

Selling is next.

References

  1. [1] Licklider, J. C. R. (1960). “Man-Computer Symbiosis.” IRE Transactions on Human Factors in Electronics, HFE-1, 4-11.
  2. [2] The interface timeline, in brief. Vercel popularized the term “Generative UI” with AI SDK 3.0 on March 1, 2024 (https://vercel.com/blog/ai-sdk-3-generative-ui). Google introduced A2UI, an open, framework-agnostic specification for agent-driven, streamed generative interfaces, on December 15, 2025 (“Introducing A2UI: An open project for agent-driven interfaces,” Google Developers Blog) and published version 0.9 with React, Flutter, Lit and Angular renderers on April 17, 2026 (“A2UI v0.9: The New Standard for Portable, Framework-Agnostic Generative UI,” Google Developers Blog; specification at https://a2ui.org). MCP Apps is the Model Context Protocol extension for interactive interface resources rendered inside assistants.
  3. [3] The buyer-side timeline, in brief. OpenAI launched Instant Checkout in ChatGPT in late September 2025, open-sourcing the Agentic Commerce Protocol with Stripe and setting a merchant fee on completed purchases. Microsoft launched Copilot Checkout with Shopify in early January 2026. Google unveiled the Universal Commerce Protocol on January 11, 2026 at NRF, co-developed with Shopify, Etsy, Wayfair, Target and Walmart, and endorsed by more than twenty partners including Visa, Mastercard, Stripe and American Express; UCP interoperates with the Agent Payments Protocol (AP2), Agent2Agent (A2A) and the Model Context Protocol (MCP), with the retailer as seller of record. Business Agent, the merchant-branded conversational agent inside Google Search configured from Merchant Center, went live January 12. The cross-retailer cart followed in May; Shopify moved its catalog MCP endpoints to the new scheme on June 15. Timeline re-verified at publication. Primary sources: Stripe, “Stripe powers Instant Checkout in ChatGPT and releases Agentic Commerce Protocol codeveloped with OpenAI,” September 2025, https://stripe.com/newsroom/news/stripe-openai-instant-checkout; Microsoft, “Microsoft propels retail forward with agentic AI capabilities,” January 8, 2026, https://news.microsoft.com/source/2026/01/08/microsoft-propels-retail-forward-with-agentic-ai-capabilities-that-power-intelligent-automation-for-every-retail-function/; NRF, “Google deepens AI investments that impact retail,” January 2026, https://nrf.com/blog/google-deepens-ai-investments-that-impact-retail.
  4. [4] Hutchins, E. L., Hollan, J. D., and Norman, D. A. (1985). “Direct Manipulation Interfaces.” Human-Computer Interaction, 1(4), 311-338.
  5. [5] Simon, H. A. (1971). “Designing Organizations for an Information-Rich World.” In M. Greenberger (ed.), Computers, Communications, and the Public Interest. Johns Hopkins Press.
  6. [6] Veliu (2026). “The Agentic Selling Benchmark: Methodology, Scoring, and Failure-Class Taxonomy.” Version 1.0, to be published with the first benchmark edition. Covers panel sourcing, persona profiling, consent and payment terms, engine coverage, scoring, and the failure-class taxonomy of section I. The three hypotheses of section XI are registered before the first data collection closes.
  7. [7] Pandya, V. (2026). “U.S. retailers see surge in AI traffic, but many websites are not entirely readable by machines.” Adobe, April 16, 2026. https://business.adobe.com/blog/ai-traffic-surge-retail-sites-not-machine-readable. Reports the readability and demand figures used in sections I and IV: product pages average a 66% machine-readability score (34% of their content is unreadable by the models doing the recommending) against 75% for homepages; AI-referred traffic to U.S. retail sites grew 393% year over year in Q1 2026 and converted 42% better than non-AI channels in March 2026, having converted 38% worse twelve months earlier; 39% of surveyed consumers report shopping with AI, 85% of whom report an improved experience. Retail is where the measurement exists first; the mechanism is channel-wide.
  8. [8] Stigler, G. J. (1961). “The Economics of Information.” Journal of Political Economy, 69(3), 213-225.
  9. [9] Bakos, J. Y. (1997). “Reducing Buyer Search Costs: Implications for Electronic Marketplaces.” Management Science, 43(12), 1676-1692.
  10. [10] Chamberlin, E. H. (1933). The Theory of Monopolistic Competition. Harvard University Press.
  11. [11] Diamond, P. A. (1971). “A model of price adjustment.” Journal of Economic Theory, 3(2), 156-168. In a market of identical buyers searching sequentially over a homogeneous good, an arbitrarily small positive search cost yields the monopoly price as the unique equilibrium.
  12. [12] Armstrong, M. (2006). “Competition in two-sided markets.” RAND Journal of Economics, 37(3), 668-691. Introduces the competitive bottleneck: where one side single-homes and the other multi-homes, the platform competes for the single-homing side and extracts from the multi-homing side.
  13. [13] Jensen, M. C., and Meckling, W. H. (1976). “Theory of the Firm: Managerial Behavior, Agency Costs and Ownership Structure.” Journal of Financial Economics, 3(4), 305-360.
  14. [14] Accenture (2026). Consumer Pulse Research 2026. June 2026; 25,590 consumers across 16 countries. Reports that 74% of consumers would let a personal AI agent handle routine tasks, 32% would accept delegated decision-making within set limits, and 9% would accept autonomous purchase completion, with configurable permissions, instant override, and clear recourse as the leading conditions of trust (reported in AI News, June 15, 2026, https://www.artificialintelligence-news.com/news/ai-shopping-agents-consumer-trust-accenture-report/). See also Checkout.com (2026). “Agentic Commerce 2026: The State of Consumer Demand and Merchant Readiness,” June 9, 2026, https://www.checkout.com/newsroom/consumer-demand-for-ai-shopping-is-forming-fast-but-trust-for-agentic-commerce-is-still-catching-up: spending caps (30%), instant permission revocation (29%), and easy cancellation (28%) as the controls consumers require before delegating.
  15. [15] The trust rails at press time: tokenized agent payments from the card networks (Visa’s Trusted Agent Protocol, Mastercard’s Agent Pay); the first evidence standards for agentic transactions; in-band machine payments over HTTP 402 (x402, under the Linux Foundation) and agent payment tooling in the cloud stacks (AWS Bedrock AgentCore Payments); verifiable credentials and portable agent identity in progress at the W3C and in ERC-8004; request signatures that let automated buyers identify themselves to the sites they visit (Web Bot Auth). Re-verified at publication. Primary sources: Visa, Trusted Agent Protocol specifications, https://developer.visa.com/capabilities/trusted-agent-protocol/trusted-agent-protocol-specifications; Mastercard, Agent Pay, https://www.mastercard.com/us/en/business/artificial-intelligence/mastercard-agent-pay.html; Cloudflare, “Securing agentic commerce: helping AI agents transact with Visa and Mastercard,” https://blog.cloudflare.com/secure-agentic-commerce/ (covers Web Bot Auth and card-network agent authentication); ERC-8004, https://eips.ethereum.org/EIPS/eip-8004.
  16. [16] Simon, H. A. (1955). “A Behavioral Model of Rational Choice.” Quarterly Journal of Economics, 69(1), 99-118.
  17. [17] Slovic, P. (1995). “The Construction of Preference.” American Psychologist, 50(5), 364-371.
  18. [18] Douglas, M., and Isherwood, B. (1979). The World of Goods: Towards an Anthropology of Consumption. Basic Books.
  19. [19] Miller, D. (1998). A Theory of Shopping. Cornell University Press.
  20. [20] Anthropic (2026). “Project Deal: our Claude-run marketplace experiment.” Published April 24, 2026. https://www.anthropic.com/features/project-deal. A one-week experiment in December 2025 in Anthropic’s San Francisco office: 69 employees, $100 each, four parallel Slack-based marketplaces, over 500 items listed, 186 deals, just over $4,000 in transaction value; agents posted listings, made offers and negotiated autonomously in natural language. Agents run on Claude Opus earned about $3.64 more per sale on identical items than agents run on Claude Haiku; participants represented by the weaker model did not notice the disadvantage and rated fairness the same; aggressive negotiation instructions did not change outcomes; 46 percent of participants said they would pay for such a service.
  21. [21] Kleinberg, J., and Raghavan, M. (2021). “Algorithmic monoculture and social welfare.” Proceedings of the National Academy of Sciences, 118(22), e2018340118. See also Bommasani, R., Creel, K. A., Kumar, A., Jurafsky, D., and Liang, P. (2022). “Picking on the Same Person: Does Algorithmic Monoculture Lead to Outcome Homogenization?” Advances in Neural Information Processing Systems 35; arXiv:2211.13972.
  22. [22] Schultz, W., Dayan, P., and Montague, P. R. (1997). “A Neural Substrate of Prediction and Reward.” Science, 275(5306), 1593-1599.
  23. [23] Berridge, K. C., and Robinson, T. E. (1998). “What is the role of dopamine in reward: hedonic impact, reward learning, or incentive salience?” Brain Research Reviews, 28(3), 309-369.
  24. [24] Knutson, B., Rick, S., Wimmer, G. E., Prelec, D., and Loewenstein, G. (2007). “Neural Predictors of Purchases.” Neuron, 53(1), 147-156.
  25. [25] Poldrack, R. A. (2006). “Can cognitive processes be inferred from neuroimaging data?” Trends in Cognitive Sciences, 10(2), 59-63.
  26. [26] Campbell, C. (1987). The Romantic Ethic and the Spirit of Modern Consumerism. Blackwell.
  27. [27] Holbrook, M. B., and Hirschman, E. C. (1982). “The Experiential Aspects of Consumption: Consumer Fantasies, Feelings, and Fun.” Journal of Consumer Research, 9(2), 132-140.
  28. [28] Pine, B. J., and Gilmore, J. H. (1998). “Welcome to the Experience Economy.” Harvard Business Review, July-August 1998.
  29. [29] Manovich, L. (2001). The Language of New Media. MIT Press. Variability as a principle of a computational medium: the object that exists in many versions and has no single instance.
  30. [30] Nissenbaum, H. (2004). “Privacy as Contextual Integrity.” Washington Law Review, 79(1), 119-157.
  31. [31] “Agentic selling”: the seller side of agentic commerce, named here. The buyer side has begun to accumulate a literature, including generative engine optimization (Aggarwal, P., et al. (2024). “GEO: Generative Engine Optimization.” Proceedings of KDD 2024; arXiv:2311.09735). The selling half has none yet. We intend to write it.
  32. [32] Sensale: Treccani, s.v. “sensale”; from Arabic simsar, from Persian, mediator. The word itself traveled the trade routes.
  33. [33] Parveen, D., Kang, D., Paruchuri, A., Kayal, D., and Mallapragada, P. (2026). “Accelerating Personalization Signal Learning via Synthetic Data.” Proceedings of ECIR 2026. https://www.amazon.science/publications/accelerating-personalization-signal-learning-via-synthetic-data. Finds that models trained on synthetic personas outperform models trained on real, de-identified data, in precision and recall on the same real test set.
  34. [34] Veliu (2026). The Veliu persona pipeline. Five stages: taxonomy-anchored persona generation under a tiered attribute schema, role-played buying dialogues with controlled affect targets, signal extraction with reasoning traces, dataset assembly for distillation, and evaluation. A reproduction of the method of [33], whose original code was not released.
  35. [35] Figures as of the date of publication, counted on the live marketplace described in section VII and stated as lower bounds.
  36. [36] Lewis, M., Yarats, D., Dauphin, Y., Parikh, D., and Batra, D. (2017). “Deal or No Deal? End-to-End Learning of Negotiation Dialogues.” Proceedings of EMNLP 2017, 2443-2453.
  37. [37] Bianchi, F., Chia, P. J., Yuksekgonul, M., Tagliabue, J., Jurafsky, D., and Zou, J. (2024). “How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis.” Proceedings of ICML 2024; arXiv:2402.05863.
  38. [38] Crawford, V. P., and Sobel, J. (1982). “Strategic Information Transmission.” Econometrica, 50(6), 1431-1451.
  39. [39] Anthropic (2026). “Building Commerce Agents with Claude.” September 2, 2026. https://claude.com/blog/claude-for-commerce-agents; reference implementations at https://github.com/anthropics/commerce-agents. Ships a shopping agent (search, compare, substitute, assemble the order) and a merchant agent (sales questions, promotions, inventory and pricing), with implementations over UCP and the Shopify Admin API published by Shopify at https://github.com/Shopify/claude-for-commerce-examples.
  40. [40] Ellison, G., and Ellison, S. F. (2009). “Search, Obfuscation, and Price Elasticities on the Internet.” Econometrica, 77(2), 427-452.
  41. [41] Frankfurt, H. G. (1971). “Freedom of the Will and the Concept of a Person.” Journal of Philosophy, 68(1), 5-20.
  42. [42] Polanyi, K. (1944). The Great Transformation. Farrar and Rinehart.