Free to read6 min read

Who Pays When an AI Agent Does the Work?

From the The Economics of Personal Agents: Costs, Revenue, and the Business of AI That Works for You collection

A web search costs its provider a tiny fraction of a cent. So does a refreshed social feed. That near-zero marginal cost was the quiet foundation of the consumer internet for two decades: when serving one more person costs almost nothing, advertising can pay for everything, and the product can be free.

Personal AI agents change that foundation. An agent asked to plan a weekend trip may read dozens of web pages, compare prices, call booking tools, draft an itinerary, and check its own work before reporting back. Each step consumes compute, and the bill arrives every time the agent acts.

Software that books, buys, researches, writes, and operates other software on a person's behalf moved from demonstrations to products in 2025 and 2026. The products arrived faster than any shared account of how they make money.

The useful way to read this business is as a ledger. On one side sits the cost of a completed task; on the other, the value that task delivers and the share of that value someone will pay for. Every pricing decision, revenue experiment, and competitive bet in the agent market is an attempt to keep the second column larger than the first. Four ideas make the ledger legible.

A cost that rises with use

Search and social networks monetized attention. They spent almost nothing per user and earned by placing ads in front of eyes that were already there. Agents run the logic in the opposite direction: they spend compute in order to save their user's attention. The person who once scrolled through ten results now receives one answer or one completed booking.

That inversion gives agent businesses a different shape. Revenue has to cover a cost that rises with use, and the heaviest users are the most expensive ones to serve. An ad-funded service welcomes its most engaged users; an agent provider charging a flat price has reason to watch them closely.

The inversion also reaches beyond the agent companies. Much of the web was built for human visitors: ads beside articles, product placements, store layouts designed for browsing. When software does the browsing, publishers and retailers see fewer people arrive.

Content licensing deals, pay-per-crawl arrangements, and agent checkout fees are early attempts to reprice access to the web for an audience made partly of software. These arrangements sit alongside the older ones. Advertising and storefronts continue to fund much of what people read and buy.

The token, the task, and cheaper intelligence

The basic unit of cost is the token. Models charge separately for input, output, and reasoning tokens, and previously cached context costs less than fresh context. Choosing a small model or a frontier model for the same job can move the price by one or two orders of magnitude.

Agents multiply tokens. Planning, tool calls, long documents, and self-checking loops can make a single task consume many times the compute of a chat reply. Then come the costs outside the model: browsing and computer-use infrastructure, sandboxes, memory storage, third-party API fees, payment handling, and human review when something goes wrong.

Failed attempts and retries count too, which makes reliability an economic variable as much as a technical one. A more reliable agent is a cheaper agent per completed task.

Against this sits a steep decline in the price per token for a given level of capability, a trend that ran through 2023, 2024, and 2025. Optimists read it as the path to affordable agents for everyone. Skeptics point to an older observation: in 1865 William Stanley Jevons argued that more efficient use of coal would raise total coal consumption, because cheaper energy finds more uses.

The Jevons effect appears in AI as well. Falling unit prices widen the range of tasks worth delegating, longer and more ambitious tasks follow, and total spending keeps growing. Both readings fit the evidence, and which one dominates for a given product depends on how its users respond to lower prices.

The same tension runs through the capital stack underneath. Hyperscaler capital spending on AI reached the hundreds of billions of dollars a year in 2025. The bull case holds that agent workloads will turn that capacity into revenue; the bubble case holds that revenue will trail the spending for longer than investors expect.

Agents tilt the economics toward inference, the cost of running models. The answer to the capital question therefore depends heavily on how many tasks people delegate and what they pay for each.

Heavy users and the limits of a flat price

Consumer AI assistants converged on a subscription of roughly $20 per month. In December 2024 OpenAI added a tier at $200 per month. A heavy tier priced at ten times the standard one signals that some users want, and consume, far more compute than the standard price covers.

This is the power-user problem. Most everyday requests are cheap to serve, while a small share of long-running tasks drives most of the compute bill. An agent that works for hours on a project can turn a profitable subscriber into a loss under a flat monthly price.

Providers respond with the tools of rationing and pricing: rate limits, credits, higher tiers, and a gradual shift toward usage-based and outcome-based charges. Model routing is the quieter response. Sending simple work to small or on-device models and hard work to frontier models lowers the average cost per task, so the routing policy becomes a margin policy.

Bundling spreads the cost across a larger relationship. Folding agents into existing subscriptions for storage, office software, devices, or shopping lets one payment carry several services, with the agent as one of them.

Other ways to earn, and the question of trust

Subscriptions are one revenue stream among several. Agents that search, compare, and complete purchases create a place to earn a share of each transaction, the take rate, through merchant fees, affiliate arrangements, in-chat checkout, and payment programs built for agents.

Enterprises may pay per seat, per agent, or per task, and they may pay more than consumers do, because work done inside a business often has a clear cost attached. Devices shift part of the bill onto the user's own hardware when models run locally.

Each stream carries its own trade-off between earnings and trust. Search advertising faced a version of the question: whether a ranked list paid for partly by sellers still serves the person searching. An agent that chooses and buys on someone's behalf raises the same question with higher stakes, because the user sees the outcome and rarely sees the options that were passed over.

Incentives that stay visible and aligned with the user may prove to be an economic asset in their own right. Trust, in this market, has a price on both sides of the ledger.

Where the value settles

The stack has many layers competing for margin: chipmakers, cloud providers, model labs, platform owners, and application builders. Capable models are multiplying, and open-weight releases put steady pressure on model prices.

That pressure moves attention toward control of distribution: the phone, the operating system, the browser, the default assistant, and the accumulated memory and integrations that make one agent more useful to a particular person than any rival.

Three scenarios remain open. A few general-purpose agents could hold most users, much as a few search engines and social networks did. Many specialized agents could divide the market by task and industry. Or agents could become a utility layer, metered and capital-heavy, earning thin but durable margins in the manner of cloud computing or electricity.

Each scenario leaves its own evidence trail. Rising switching costs from memory and integrations would favor the first. Strong pricing power in specific trades and industries would favor the second. Falling model prices alongside steady margins at the infrastructure layer would favor the third.

A ledger that keeps moving

Prices in this market change month to month, so any fixed figure ages quickly. The durable tool is the method: estimate the cost of a completed task from tokens, tools, and retries; set it against revenue per user; read the margin; and ask what would change the answer.

Those four steps apply to any agent product or forecast. Who pays, and for what, is the question each era of the internet eventually settled for itself: advertisers paid for attention, subscribers paid for access, and merchants paid for traffic.

Personal agents add a new entry to that list, a payment for work done on someone's behalf. The arithmetic behind it favors whoever keeps the cost of each completed task below its value to the person who asked.