This website uses cookies

Read our Privacy policy and Terms of use for more information.

Hi {{first_name|Investor}}

Memory is where an agent's bill shows up first, and the buyers are already paying it. Google Cloud told the SEMICON audience on September 2 that memory is now more than 75% of the hardware bill for an AI server. Amazon raised its 2026 capital spending from about $200 billion to $220 billion in July, and Jassy gave one reason: the higher cost of memory. NVIDIA's CFO said on August 26 that gross margin will bottom near 71% on "extreme pricing conditions in memory." The day after OpenAI shipped its computer-use model Astra on September 3, the memory and storage names led the tape: SanDisk ($SNDK) closed up 11.9%, Micron ($MU) 6%, SK hynix 8%, with NVIDIA's Hugging Face deal and Dell's "DRAM, DRAM, DRAM" comment landing the same day.

The memory an agent hoards is being pushed out of the GPU into a ladder of cheaper, slower tiers, and each rung is a product somebody sells. This issue walks the ladder.

How an agent remembers

A language model keeps a working copy of the context it's processing, in a structure called the KV cache. Every token it has read or written in the session adds to it, it lives in the fastest memory available (the HBM on the GPU), and it's temporary: intermediate state for the task in front of it, not a permanent record.

Size is the problem. Llama 3.1 70B at a 128K context in 16-bit precision holds about 40 gigabytes of cache for one user. A million-token context, which Astra supports, is 8 times that on the same math. Run 100 agents at once and the cache leaves the GPU in seconds.

So the cache spills. First into the server's DRAM, which tops out at 1 to 2 terabytes per node. WEKA's June analysis put the share of context that's actually still in memory when an agent needs it in the 70s, against a theoretical 90-plus. Providers sell cache windows of 5 minutes or an hour, and hardware memory pressure evicts on its own schedule; either way, once the context is gone it gets rebuilt from scratch.

That rebuild has a price. Anthropic's rate for Fable 5.1 is $10 per million tokens of fresh input and $0.25 for a cache read, plus a premium when the cache is written. Remembering costs a small fraction of re-reading, and the memory layer exists to keep it that way.

Think of a chef's station. The cutting board is HBM: tiny, instant, holding only the dish in front of you. The counter behind is DRAM, the walk-in fridge is SSD, the warehouse across town is disk. A restaurant serving 100 tables lives or dies by how well it moves ingredients up and down that chain, and agents are 100 tables that never leave.

The rest of this issue is for Insiders.

Below the line: the 5-rung ladder itself, with the company on each rung and what it reported. Micron's data center SSD line more than doubled in a quarter and its CEO named the reason. SanDisk grew data center revenue 103% and disclosed that two-thirds of its growth came from one thing that reverses. Plus the newest rung, where the chips ship to 2 hyperscalers in 2027, the slowest rung, which is also the one least exposed to efficiency gains, which names are positions, what would tell me I'm wrong, and what I'm reading first on September 30.

logo

This is where it gets interesting.

Insiders get the full research — the emerging patterns, the second-order effects, and the trends I think matter before they're obvious

Unlock the full post

A subscription gets you:

  • Weekly deep dives on trends before they're consensus
  • Operator-level research and frameworks behind every thesis
  • Full archive of every past issue and report