Notes from an intimate Baseten conversation with Parag Agrawal, founder of Parallel Web Systems and former CEO of Twitter, interviewed by Charles O’Neill, co-head of model training at Baseten_._
Last night at Baseten, Parag Agrawal led an intimate group through a fascinating journey into how search economics break when agents show up. Generous with time, unhurried with questions, including the awkward and malformed one (mine), answering at a level of depth that assumed we could keep up.
Parag opened with how he founded Parallel on a prognosis that agents would use the web a 1,000 times more than humans do. Asked what he had gotten wrong since, he went the other way.
“I think today it feels that I was wrong, but in the other direction…I think 1000X is a huge underestimate of what’s going to happen, and my imagination just wasn’t big enough at the time.”
That is the premise for everything that followed. If you accept it, a set of economics that looked stable for twenty-five years stops working.

Agents will use the web way more than 1,000x than humans. Parag says his original estimate was too small.
The bill flipped, and it flipped recently
The most concrete thing he said had like nothing to do with vision. It was about invoices. Model inference got cheap enough, fast enough, that when he looked at where his own spend was going, search had quietly become the dominant line item.
Check the arithmetic Parag laid out. Human search monetizes somewhere in the range of $50 to $250 dollars per thousand searches. Agent web search, purchased through the incumbent APIs, has been running $10 to $20 per thousand. Meanwhile the model call underneath it has collapsed toward free.

Inference costs collapsed. Search didn’t. And now’s dominant line item in agent work.
So the ratio inverted. Search used to be a rounding error next to inference. Now, in his words, whatever you’re doing, cost ends up dominated by web search.
Old search pricing doesn’t work at that volume. You can’t pay $50 per thousand queries and then run a thousand times more queries. The math breaks.
“We need to deliver better search than Google at one hundredth the cost.”
Perhaps ten times faster too. That’s the goal.
Parallel says it recently shipped in that range, around a dollar per thousand requests, roughly a tenth of prevailing API pricing. Published rate for turbo and fast search is $1 per 1,000 requests. Numbers in the appendix.

Why the incumbents don’t chase this
Charles pushed him on the obvious alternative:
why not make search free and fund it with advertising, the way every search business before this one has worked?
The sharpest structural claim of the night was about incentives, not tech. Once a search business monetizes attention well, he argued, it has no real incentive to chase a hundredfold cost reduction. The pressure goes into growing the audience and tightening monetization instead. A business clearing fifty dollars per thousand queries is rationally allocating its engineering effort elsewhere. That’s rational, not lazy. imho.

An ad buys a slot in human attention. Agents have no attention to sell.
Agrawal is unusually generous about advertising, which makes the argument stronger rather than weaker. He worked on ads. Advertising, he said, is the most efficient engine for differential pricing that exists, so good at it that a business can lose money on ninety percent of its queries and ninety percent of its users and still post extraordinary margin.
But he rejects it for agents on principle. Advertising is the wrong business model here, full stop. Parallel is in the business of giving customers the best possible data, not influencing it.
The reason is mechanical. An ad works by buying a slot in human attention. An agent has no attention to sell. Put a sponsored result into an agent’s context window and you haven’t monetized anything, you haven’t monetized anything, you’ve broken the answer.

What replaces the old bargain
His framing: every web business model so far assumes you’re either monetizing attention or selling to a human. Agents do neither.
So Parallel built the alternative. He described Index as, in effect, the AI-era version of AdSense. They’ve trained models to estimate, in a principled way, how much marginal value a piece of content added each time an agent used it to do work on the web, and to pay accordingly. Partners let Parallel use their data, and Parallel pays them based on what the model predicts they’re owed.

Shapley analogy: every stone gets paid. The one holding the most gets the most. Contribution measured, not guessed.
The estimator is Shapley value, borrowed from cooperative game theory. Given a set of sources that jointly produced an answer, a source’s Shapley value is its average marginal contribution across every possible ordering of the sources. Drop a redundant source and the answer barely moves, so it earns little. Drop a load-bearing one and the answer collapses, so it earns more. Uniqueness gets paid.
Worth noting alongside that: by his estimate, agents read on the order of a hundred times more content than a human does to accomplish roughly the same task, and Cloudflare has reported that close to half of web traffic today is agents, not people.
The cognitive core
Charles had his own thesis he wanted tested: that models are heading toward smaller and more purely reasoning-focused over time, with facts pulled live from outside sources rather than memorized during training. He put the target state simply, wanting a model brought down to “just its cognitive core,” stripped of the job of memorizing things it doesn’t need to know.
Agrawal agreed with the direction. Parameters spent memorizing facts are parameters not spent reasoning, and it’s precisely why models are unreliable on the long tail of what they were trained on, the obscure date, the minor fact, the thing that shows up once in the corpus. His read on where this goes: smaller, sharper models, with less invested in memorization, paired with live access to your personal data, your company’s data, and the web through something like Parallel.

The best audience question of the night
Someone asked:
If Parallel pays out based on how unique your contribution is, what stops people from manufacturing fake, artificially unique information to game the payout?
They framed it as a verification-layer question, nothing about scoring mechanics.
His first answer downplayed the risk: ’they’re not big enough yet to be worth gaming, and if your SEO strategy today is built around outsmarting Parallel specifically, you’re chasing the wrong target.'
The second answer was the real one. He conceded the adverse incentives are coming and that Parallel will have to solve for them, then put a number on it: something like a year or two before this shows up as a genuine vector, before rational actors start actively gaming the score.
That’s a bigger admission than it sounds like in the moment. It means the scoring mechanism is unhardened today, by his own account, with a self-imposed clock already running.

402 handshake: bridging agentic demand with instant payouts for long-tail creators (draft)
That is the whole case for machine payments in one sentence, made by someone who is not selling a payments protocol. Bilateral licensing works for The Atlantic. It does not work yet for ten thousand independent writers owed forty cents each. Invoicing does not scale down. A standard rail is the only thing that does, and that is where machine payments can open the door for independent creators. The buy side is already there, and already small by his own account, a handful of customers paying through machine payment rails today, with room to grow.

The question I left with
If the pricing model is the defensible thing and the rails are commodity, then the rails will be built by whoever benefits most from them existing. Parallel benefits. The payment rail providers benefit. The infrastructure layer benefits.
Here’s what I’d ask next: as this rail gets built, what would it take to make sure it works as well for the smallest publisher as it does for the biggest platform.

The rail gets built by the big players. The long tail may be for who needs it most.
PS Thanks to Baseten, Devin Fuller & team for making the evening so memorable!
Sources:
Parallel pricing documentation
Agentic payments integration (MPP/x402 setup for agents)
docs.parallel.ai/integrations/agentic-payments
twitter.com/schwentker/status/2090711103823388687
bsky.app/profile/schwentker.sandboxlabs.ai/post/3mtslmewhtc2c
