Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Why AI Infrastructure Bottlenecks Are Moving Beyond GPUs

Дата публикации: 25-06-2026 20:05:31

The variable most organizations are missing isn’t compute — it’s storage purpose-built for AI context, not just data capacity.

Основное содержимое страницы с новостью.

Getty

Token costs are falling. Yet, enterprise AI bills are rising. The variable most organizations are missing isn’t compute — it’s storage purpose-built for AI context, not just data capacity.

The Agentic Cost Paradox

Enterprise AI economics contain a contradiction most budget models never saw coming.

The cost to generate a token has collapsed — down roughly 60 to 70 percent per year, by Goldman Sachs’ estimate.1 Leaner models, better hardware, and fierce competition have all pushed the same way. Generating intelligence has never been cheaper.

And yet AI bills are climbing. The FinOps Foundation found nearly three-quarters of enterprises exceeded their AI cost projections last year.2 The arithmetic is unforgiving: spend equals price times volume. The price is falling. The volume is exploding faster than any forecast assumed.

Agentic AI is why. Agents don’t ask one question and stop — they reason in chains, hold context across long sessions, and run continuously. Goldman Sachs projects token consumption will grow 24 times by 2030.3 But the deeper issue isn’t just that agents consume more. It’s that long-context workloads inflate the context window, and the cost of holding that context scales disproportionately as it grows. More context doesn’t cost a little more. It costs a lot more.

That’s what the cost-per-token conversation misses. What drives AI economics now isn’t just the price of a token — it’s also whether the infrastructure underneath can hold and serve exploding volumes of context without breaking down at scale. That’s not a compute question. It’s an infrastructure question. And the piece most AI strategies haven’t accounted for – is storage.

Storage Is the Missing Link

Here is what most AI infrastructure conversations miss: storage is not a passive component. It is an active performance variable — one that directly determines how productively GPU accelerators operate, how efficiently inference runs, and what it costs to serve every token at scale.

In production AI, GPU memory is being asked to do two jobs at once: computation and context storage. Those are not the same job. Computation is what GPUs were built for. Holding large volumes of reusable context across active inference sessions is work that purpose-built storage should be doing in coordination with GPU memory. When GPUs are forced to do both, the result is predictable — accelerators stall waiting on memory, inference degrades under concurrent load, and the GPUs you are paying for sit idle instead of generating output. The missing link in most AI strategies isn’t more compute. Its storage built to carry the load GPU memory was never designed to hold.

The Payoff: What KV Cache Offloading Delivers

The mechanism at the center of this problem has a name: the KV cache. During inference, large language models store computed token relationships in this cache rather than recalculating them with every output. It is what makes AI fast and coherent at scale. It is also what breaks first under agentic AI.

Agents maintain context across extended, multi-turn sessions — histories, long documents, chains of reasoning that compound with every interaction. Each exchange expands the KV cache. At scale, across many concurrent agents and users, it consumes gigabytes, then terabytes, of GPU memory. When it overflows, older context is evicted, and recomputing it is expensive and slow — users feel it immediately as degraded response time. This is the two-jobs problem made concrete: GPU memory forced to store context when it should be free to compute. More GPU memory doesn't fix it. It just raises the cost of getting the architecture wrong.

GPU memory is for computation. Context storage is a job for storage infrastructure. Asking your most expensive resource to do both is what breaks the token economics — and it is a problem that scales with your ambition, not against it.

The fix is simpler than most organizations expect. Move the KV cache off the GPU and onto storage purpose-built to hold and serve it at inference speed. The GPU gets its memory back. The context doesn’t disappear — it lives where it belongs. In testing on four NVIDIA H100 GPUs, Dell AI storage delivered a 1-second Time to First Token at a 131K-token context window. The same configuration without KV cache offloading took 17 seconds.4 Seventeen seconds is a failed user experience. One second is a viable product. Multi-turn throughput nearly tripled. Nothing changed except where the context lived.

It’s important to note that not all shared AI storage performs equally — and at agentic concurrency levels, that gap compounds fast. Independent testing on Qwen3-32B showed Dell's storage running at nearly half the response latency of the nearest competitor, and up to 14× faster than inference without KV cache offloading. At scale, that difference isn't a benchmark stat. It's the distance between an AI product that performs under load and one that degrades.

Time to Compute Matters

Performance is only half the equation. The other half is how fast you can put it to work.

Most organizations discover their storage wasn’t built for their GPU environment six months into deployment. The integration work, the tuning, the gap between infrastructure that is provisioned and infrastructure that is productive — these costs rarely surface in procurement but show up reliably in the ROI analysis. And in a market where deployment speed has become its own competitive variable, the lag between capital deployed and capacity earning is rarely recoverable.

This is where sourcing matters more than most procurement models assume. When compute, storage, and networking are designed, validated, and delivered as one system rather than assembled from parts that were never built to work together, the entire stack reaches production on a single timeline. There’s no integration project waiting on the other side of the purchase order, no multi-vendor finger-pointing when performance falls short. AI performance at scale isn’t defined by any single component — it’s defined by how well the whole system works together, and how quickly it can start working at all.

The benefit compounds over time. Because Dell engineers its storage and compute in alignment — on coordinated roadmaps rather than as separate product lines chasing each other — a new GPU generation doesn’t trigger a storage re-architecture. The storage is already designed for what’s coming. That is only possible when a single vendor is accountable for the whole stack: when one company owns the roadmap across compute, storage, and networking, the pieces evolve together instead of leaving customers to absorb the gaps between them. Time to compute — the distance between capital deployed and capacity earning — is where that advantage is won.

The Chapter Most Enterprises Haven’t Written Yet

Agentic AI is not a future state. It is the operating model every serious enterprise AI program is moving toward: continuous inference, long context, multi-agent concurrency, persistent session state. These are precisely the demands that overwhelm GPU memory alone and reward a storage architecture purpose-built to carry the load. The organizations that lead won’t simply have the most compute. They’ll have built for the agentic era rather than adapting to it after the fact and that distinction will compound for years.

Storage is no longer a footnote in that story. It is the chapter most enterprises haven’t written yet.

1, 3 Goldman Sachs Research, “AI Agents Forecast to Boost Tech Cash Flow as Usage Soars” (May 2026): per-token inference costs declining 60–70% per year; global token consumption projected to grow 24x between 2026 and 2030, reaching ~120 quadrillion tokens per month. goldmansachs.com

2 FinOps Foundation, State of FinOps 2026: nearly three-quarters of surveyed enterprises reported AI costs exceeding original projections. finops.org

4 Dell Technologies, “Breaking the AI Inference Context Memory Barrier with Dell Storage,” Dell Technologies Info Hub: benchmark testing of KV cache offloading on NVIDIA H100 GPUs (LLaMA-3.3-70B and Qwen3-32B), including 19x Time to First Token improvement at 131K-token context and comparative results vs. competing storage. infohub.delltechnologies.com

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1The Next AI Advantage Will Be Built, Not Bought05.7307-07-2026
2From Procurement To Production: The Real Bottleneck In The AI Infrastructure Buildout08.4927-07-2026
3Assess your storage strategy for the AI era09.6306-08-2026
4Why Research Infrastructure Is A Strategic Priority For University Leaders07.7909-07-2026
5Government’s AI Problem Isn’t The Model. It’s The Silos.07.1223-07-2026
6Data And Storage At The Center Of The AI Stack08.129-05-2026
7PCIe Gen6 and Gen5 Will Both Matter for AI Storage011.5301-08-2026
8Why Europe’s AI Strategy Must Start with Wireless Infrastructure011.6302-04-2026
9AI without the hype: What AML transformation means09.0706-10-2026
10Stop automating inefficiency and scale AI the right way 0525-06-2026

Классификация: Мнения. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 12.07. Источник: www.forbes.com.