Dwarkesh Patel warns compute prices will surge 10x
- Falling token prices reflect smart software routing, not dying demand for artificial intelligence.
- Compute bottlenecks pushed spot GPU prices up 40 percent as top labs outbid rivals.
- Goldman Sachs projects AI infrastructure buildout costs will hit $1.1 trillion by 2027.
Wall Street is misinterpreting the dip in token prices. A Citadel Securities chart showing downward-trending token costs convinced critics that the AI bubble is popping, but they are reading the data backward.
On The AI Daily Brief, host Nathaniel Whittemore pointed out that falling unit prices reflect third-party routing software finding cheaper suppliers, not shrinking adoption. Total volume is exploding as enterprises shift from simple prompts to autonomous agents that consume massive amounts of compute. Guest Nufar Gaspar noted that token pricing is misleading anyway, especially when reasoning models spend internal monologue tokens that cost up to 20 times more than visible output. A 27% cost bump from tokenizer updates can easily get masked by nominal unit prices, making cost per successful task the true metric.
"The most expensive token is the one your best person is afraid to spend."
- Nufar Gaspar, The AI Daily Brief
That explosion in agentic volume is running straight into a physical wall. Discussing the mechanics on the Dwarkesh Podcast, host Dwarkesh Patel explained that while AI lab revenues are growing tenfold annually, physical compute capacity expands only three times a year. Spot prices for compute have already jumped 40% since February. Google is reportedly paying SpaceX double the spot rate - roughly $900 million a month - to secure 110,000 GPUs.
Patel argued that hardware builders cannot simply manufacture their way out of this crunch. Moore’s Law now delivers just a 1.4x performance gain per cycle, and expanding silicon output requires scarce lithography machines from ASML. Meanwhile, AI hardware has already swallowed 86% of TSMC’s N3 wafer node allocation, leaving almost no consumer chip capacity left to cannibalize. The result is a brutal selection pressure where elite labs outbid minor startups for scarce silicon.
"Using a weaker model wastes expensive tokens."
- Dwarkesh Patel, Dwarkesh Podcast
The next day on The AI Daily Brief, Whittemore highlighted how Wall Street is adjusting to this scarcity era. Goldman Sachs strategist Ryan Hammond now projects AI infrastructure spending will reach $1.1 trillion by 2027, driven by an expected 24-fold rise in token consumption by 2030. Tech giants are rewriting their operational playbooks to secure capacity: SpaceX is evolving into a space-based data center host, while Jeff Bezos’s new startup, Prometheus, raised $12 billion from JP Morgan, Goldman Sachs, and BlackRock to build industrial AI.
The bottleneck has moved from software design to the physical power grid.
This desperate search for hardware coincides with a sharp geopolitical split. Whittemore detailed how Meta was forced by Beijing to dismantle its $2 billion acquisition of Chinese startup Manas, as Chinese authorities seize passports of key AI researchers to prevent talent flight. As Chinese firms re-incorporate inside the mainland, Western platforms are left fighting for whatever domestic compute and talent remain inside their own borders.
Cheap tokens are an illusion of market efficiency. The underlying compute powering them is becoming the scarcest asset on Earth.
Source Intelligence
- Deep dive into what was said in the episodes

Nathaniel Whittemore
What Happens When AI Breakthroughs Outrun Human Understanding • Aug 3
- Goldman Sachs, simultaneously managing the SpaceX IPO, forecast the company could achieve $474 billion in revenue by 2030, with its AI division growing a hundredfold.
- Jeff Bezos' AI startup, Prometheus, closed a $12 billion funding round, valuing the company at $41 billion, with participation from JP Morgan, Goldman Sachs, and BlackRock.
- Prometheus aims to build an "artificial general engineer" (AGE) capable of designing and manufacturing complex equipment, having already hired 150 people across global offices.
- Bezos dismisses fears of an AI jobs apocalypse, asserting AI will create a labor shortage by generating tenfold more opportunities, even while reducing the number of people needed by ten times.
- Prometheus is also exploring a $100 billion fund for industrial buyouts, intending to apply private equity roll-up models and proprietary AI to enhance productivity in the manufacturing sector.
- Meta completed an operational split from Manas, its AI acquisition, following Chinese government orders in April to unwind the $2 billion deal.
- Eugene Wang of Wintell & Co. states that dismantling "red-chip" corporate structures is no longer optional for Chinese tech firms, with the focus now on efficient restructuring. Chinese officials are also seizing passports of key AI researchers and executives.
- Goldman Sachs strategists, led by Ryan Hammond, forecast 2027 AI spending at $1.1 trillion in a baseline scenario and $1.4 trillion in a bullish scenario, significantly exceeding Wall Street's median $920 billion estimate for next year.
Also from this episode: (12)
Startups (1)
- Nathaniel Whittemore reports SpaceX is conducting the largest IPO in history, priced at $135 per share, implying a valuation just under $1.8 trillion. This makes SpaceX the world's seventh-largest company upon debut.
Markets (3)
- Bloomberg reports retail investors submitted over $100 billion in orders for the SpaceX IPO, which offered $75 billion worth of stock, reducing retail allocation from 30% to 20%. The retail portion was nearly 7x oversubscribed.
- A Reuters opinion piece warns of significant risk for retail investors in the SpaceX IPO, citing the company's relative lack of revenue: a $5 billion loss on $18.7 billion revenue in 2025.
- Nathaniel Whittemore notes the index indicates that in mid-June, the average price paid for a million tokens had decreased from its early June peak, returning to early May levels.
Models (4)
- Nathaniel Whittemore believes the SpaceX IPO will not serve as a referendum on AI models, arguing Elon Musk's unique market position and SpaceX's recharacterization as a "Neocloud" infrastructure company separate it from frontier AI labs.
- Nathaniel Whittemore highlights a "token panic" on Wall Street following Citadel Securities' "Silicon Valley LLM Token Expenditure Index," which social media misinterprets as declining token demand or volume.
- Silicon Data clarified their "LLM Token Expenditure Index" is a "usage weighted average token price index," measuring the average amount the market pays for a million LLM tokens, not total volume, demand, or expenditure.
- Analyst Max Weinbach estimates high inference-intensive API token margins, suggesting OpenAI could cut prices by 60% and remain profitable, potentially to drive higher adoption volumes amidst competition.
China (2)
- The Chinese government investigated Meta's acquisition of Manas in March and later barred Manas founders from leaving the country, despite Manas attempting to circumvent tech export controls by relocating to Singapore.
- The unwinding of the Manas deal has created a chilling effect in China, prompting numerous startups, including StepFun, Kimi Creator Moonshot, and Kling, to consider unwinding foreign corporate structures to reincorporate domestically.
Big Tech (2)
- Google is evaluating Samsung's 2-nanometer process for components of its 10th-generation TPUs, Icefish, due to years-long backlogs and capacity limitations at its exclusive manufacturer, TSMC.
- Google has also placed orders with Intel for 2028 production for advanced packaging services, indicating a strategy to diversify its chip supply chain beyond TSMC for less sensitive components.
Everything You Need to Know About AI Tokens • Aug 2
Also from this episode: (22)
Agents (4)
- Nufar Gaspar and Nathaniel Whittemore observe that companies frequently grapple with AI token economics, requiring strategies to maximize AI value efficiently without excessive cost in the agentic era.
- Nufar Gaspar states that agentic AI tasks consume 5 to 30 times more tokens than simple chats, with McKinsey estimating 60% of an agentic task's cost tied to checking, refining, and regenerating answers.
- Nathaniel Whittemore describes an agent designed for perpetual research into AI data sources, which, despite functioning as intended, generated "spin" tokens because its continuous internet crawling was not valuable enough to justify its cost.
- Nufar Gaspar lists common silent token spenders, including idle agents, unused automations, a "pre-prompt tax" from always-on rules, immortal conversations, unfiltered data retrieval, and rework loops.
AI Infrastructure (7)
- Nufar Gaspar notes a "new anxiety" around AI token usage, as practitioners feel watched, users fear high costs, and leadership questions the value of growing AI expenses.
- Nathaniel Whittemore expresses concern that rising AI costs lead to a retreat to low-stakes use cases, arguing that "token maxing" can accelerate progress more than underspending.
- Nufar Gaspar describes the "token maximizing" era, citing Meta's reported 60-74 trillion monthly token usage and an unnamed company's $500 million Claude bill without usage limits.
- Nufar Gaspar illustrates AI token costs: drafting an email uses 500-700 tokens (around half a cent), deep research can consume hundreds of thousands, and an accidental deep research query cost one learner over 4 million tokens.
- Nufar Gaspar categorizes token spending into three types: "tokens that teach" (experimentation), "tokens that produce" (deliverables), and "tokens that spin" (wasteful, like idle agents or unused automations) that should be eliminated first.
- Nufar Gaspar recounts spending $1,500 in two weeks on an unused OpenClaw agent, which consumed almost 400 million input tokens with near-zero output (a 2600:1 ratio) due to continuous internal maintenance jobs.
- Nufar Gaspar suggests identifying token "spin" by conducting a "weekend test" for unexpected bills, monitoring extreme input-to-output ratios, or for regular users, auditing automations that lack recent business value.
Enterprise (2)
- Nufar Gaspar highlights the OpenAI CFO's proposed "Useful Intelligence Per Dollar" scorecard, emphasizing the need to measure the actual cost of each successful AI task.
- Nufar Gaspar identifies the current "token anxious" era, where employees self-censor to avoid costs, with Uber capping employee spending at $1,500 and Meta constraining AI usage after previous token-maximizing initiatives.
Models (8)
- Nufar Gaspar defines an AI token as a text chunk typically larger than a character but smaller than a word, noting that a page of English text is roughly 1,000 tokens.
- Nufar Gaspar explains that non-Latin languages can require 2-5 times more tokens for the same content, incurring a "language tax" due to per-token billing, while code's indentation and whitespace also become tokens.
- Nufar Gaspar explains that tokens are not equal across providers; the April Opus 4.7 update changed its tokenizer to produce 30-45% more tokens for the same text, causing real-world bills to grow by 12-27% despite an identical price sheet.
- Nufar Gaspar explains AI requests have three token layers: input (cheapest), reasoning (model's internal thinking, 4-20x cost of input), and output (3-5x input cost), with the reasoning layer often surprisingly expensive.
- Nufar Gaspar highlights Databricks' experiment where Sonnet 5, though 1.7 times cheaper per token than Opus 4.8, ultimately cost more per task ($2.09 vs. $1.94) due to increased iterations and reasoning, underscoring that "cost per accepted task" is the crucial metric.
- Nufar Gaspar recommends habits for efficient token use, including starting new sessions for new tasks, intentionally choosing models, right-sizing context, building reusable capabilities, and filtering data to guide AI effectively.
- Nufar Gaspar advocates protecting "tokens that teach" - for experimentation and learning AI context - as they improve overall return, citing a study where 20,000 developers found the heaviest AI users were twice as productive.
- Nufar Gaspar advises individuals to audit for "spin" tokens, practice token-smart habits, and protect their learning budget, while organizations should make usage visible, budget by workload, and tier spending allowances.
Coding (1)
- Nufar Gaspar advises using CloudCode's `/doctor` command for system audits and acknowledges model routing is an early, complex area, with Nathaniel Whittemore noting that personal model preferences will remain crucial.
Why smarter AI models could drive up compute prices 10x • Aug 3
- Mercury's Command AI uses vendor, team, and memo data to auto-categorize business transactions and provide rationales, automating financial bookkeeping tasks.
Also from this episode: (15)
Startups (1)
- Dwarkesh notes Anthropic's revenue 10xed annually for three consecutive years, reaching $9 billion last year, with projections of $100-$150 billion this year. If this trend continues, next year's revenue would hit $1 trillion.
AI Infrastructure (7)
- This rapid revenue growth contrasts with lab compute capacity, which only increases 3x year-over-year. Bridging this gap requires higher lab margins, increased compute prices, or a greater share of compute allocated to inference.
- Anthropic's inference margins reportedly rose from 40% mid-last year to over 80% currently. Additionally, spot prices for compute have increased over 40% since their February trough.
- OpenAI's compute allocation for inference rose from 25% in 2024 to an estimated 50% or higher, according to Epoch. Labs generally prefer to dedicate compute to training new models, viewing inference revenue primarily as a means to fund further development.
- Dwarkesh highlights that frontier labs need specialized compute at scale for efficiency, flexibility, and security, not just spot instances. Google, for example, pays SpaceX $900 million monthly for 110,000 GB200 and GB300 GPUs.
- Dwarkesh argues that, based on standard economics and the 'lump of labor fallacy,' the marginal value of AI compute could remain astonishingly high despite increased supply, similar to how immigration doesn't always decrease wages long-term.
- As top labs improve compute monetization and costs rise, competition will become harder for others who cannot utilize resources as efficiently. The Alchian-Allen effect suggests models that economize scarce compute inputs will command higher margins.
- Dwarkesh believes the 3x year-over-year compute capacity growth is challenging to sustain due to bottlenecks in its three components.
Chips (3)
- Google's payment for these GPUs is double the hourly spot price, which itself is over 40% higher than February's prices.
- Moore's Law contributes 1.4x, but maintaining it is uncertain. New fabs contribute 1.2x, bottlenecked by ASML EUV machines until at least 2030.
- Wafer allocation shifts contribute 1.8x, but AI has already absorbed 60% of TSMC's N3 node capacity, projected to hit 86% by next year, indicating this source will soon be exhausted.
Models (3)
- Smarter AI models monetize compute more effectively; an H100 equivalent running a human-level software engineer would theoretically rent for over $250,000 annually. This is more than 15 times the current spot price, even without accounting for continuous operation.
- Many current popular AI applications may be priced out as leading companies prioritize compute for automating AI research. These companies will outbid general users for scarce tokens as AI capabilities advance.
- Dwarkesh states that Anthropic's 10x revenue growth versus 3x compute growth demonstrates strong economies of scale in the AI model business, which he worries could lead to power concentration.
Robotics (1)
- Compute prices will eventually fall when robots can autonomously produce chips from raw materials, reflecting only input costs. However, the current 'pre-singularity regime' faces compute scarcity.
Related Stories
BUSINESS
Citadel snaps up AI assets after fund collapse
What Bitcoin Did · Forward Guidance · This Week in Startups · All-In with Chamath, Jason, Sacks & Friedberg · Simon Dixon Hard Talk · Simon Dixon Hard Talk · Bankless · TFTC: A Bitcoin Podcast
BUSINESS
Morpho unbundles DeFi credit to capture institutional debt
Bankless · Bankless · Bankless
AI & TECH
Jensen Huang rallies 100 tech firms against OpenAI rules
Hard Fork · The Daily · Rabbit Hole Recap · This Week in AI · Moonshots with Peter Diamandis
