{"id":6554,"date":"2026-07-24T23:56:21","date_gmt":"2026-07-25T03:56:21","guid":{"rendered":"https:\/\/workai.tv\/news\/2026\/07\/ai-data\/vast-data-expands-collaboration-with-amd-to-advance-ai\/"},"modified":"2026-07-24T23:56:21","modified_gmt":"2026-07-25T03:56:21","slug":"vast-data-expands-collaboration-with-amd-to-advance-ai","status":"publish","type":"post","link":"https:\/\/workai.tv\/news\/2026\/07\/ai-data\/vast-data-expands-collaboration-with-amd-to-advance-ai\/","title":{"rendered":"VAST Data Expands Collaboration with AMD to Advance AI"},"content":{"rendered":"<h2>Share with your CTO<\/h2>\n<p>VAST Data is betting that inference, not training, is where AI infrastructure money gets made next, and it&#8217;s pulling AMD into that argument with hardware and a reference architecture to back it. The <a href=\"https:\/\/www.globenewswire.com\/news-release\/2026\/07\/24\/3333077\/0\/en\/VAST-Data-Expands-Collaboration-with-AMD-to-Advance-AI-Infrastructure-for-the-Inference-Era.html\" target=\"_blank\" rel=\"noopener nofollow\">expanded VAST-AMD collaboration<\/a> pairs AMD&#8217;s 6th Gen EPYC CPUs and Instinct GPUs with VAST&#8217;s AI Operating System, with early testing on an MI355X GPU showing 9x faster time-to-first-token and 9.7x more token throughput through KV cache offloading. A joint reference architecture with DriveNets targets AI cloud providers building for training, inference, and agentic workloads at scale.<\/p>\n<h2>What this means for your business<\/h2>\n<p>The organizations most exposed to this announcement are those currently running NVIDIA-centric AI infrastructure stacks and treating data storage as a commodity layer beneath their GPU fleet. VAST is arguing, loudly, that context management, KV cache (the stored memory of an ongoing AI conversation that lets models avoid recomputing prior context), and model loading speed are now first-order performance variables. If your inference architecture doesn&#8217;t treat storage as a peer to compute, you are already accruing a cost and latency disadvantage as workloads shift from one-shot queries to multi-turn agentic sessions.<\/p>\n<p>The 9x TTFT and 9.7x throughput figures deserve scrutiny. VAST itself notes in the release that these speedups are relative to a GPU baseline, and that a lower-capacity GPU baseline could inflate the multiple further. That disclosure, buried in a bullet point, is doing real work. The numbers are not audited by AMD. What is credible is the architectural direction: as agents run longer sessions with larger context windows, the bottleneck migrates from raw GPU flops to how fast you can read and write state. VAST&#8217;s DASE architecture, which treats storage, database, and streaming as one unified namespace rather than separate layers, is a plausible structural answer to that problem, even if the exact multiples are vendor-reported.<\/p>\n<p>The AMD angle matters beyond the benchmarks. Vultr, Crusoe, Core42, and TensorWave are all quoted as AMD-aligned cloud providers validating this stack. That&#8217;s a real signal that a credible alternative supply chain to NVIDIA H100 clusters is cohering around AMD Instinct GPUs plus a purpose-built data layer. The vendor-supplied testimonials are predictable, but the roster of cloud names is not trivial. CTOs evaluating GPU cloud procurement for inference workloads in 2026 and 2027 now have a reference architecture with named deployment partners, not just a whitepaper. The falsification condition here is whether KV cache hit rates in production agentic deployments actually move total cost meaningfully; if agents remain mostly stateless, the entire data-layer argument deflates.<\/p>\n<h2>Concept deep-dive: KV Cache Offloading<\/h2>\n<p>In large language model inference, the KV cache holds the intermediate computations representing everything a model has already &#8220;read&#8221; in a conversation, so it doesn&#8217;t reprocess prior tokens on each new response. Normally this lives in expensive, limited GPU memory. KV cache offloading moves that stored context to fast NVMe storage, freeing GPU memory for active computation. Think of it as the difference between keeping every document open on your desktop versus retrieving them from a fast local drive as needed. For agentic AI running long, multi-turn sessions, this swap is the difference between a system that scales and one that stalls.<\/p>\n<p><em>Based on reporting from <a href=\"https:\/\/www.globenewswire.com\/news-release\/2026\/07\/24\/3333077\/0\/en\/VAST-Data-Expands-Collaboration-with-AMD-to-Advance-AI-Infrastructure-for-the-Inference-Era.html\" target=\"_blank\" rel=\"noopener nofollow\">VAST Data Expands Collaboration with AMD to Advance AI<\/a>, originally published 2026-07-24 15:22:00.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Share with your CTO VAST Data is betting that inference, not training, is where AI infrastructure money gets made next, and it&#8217;s pulling AMD into that argument with hardware and a reference architecture to back it. The expanded VAST-AMD collaboration pairs AMD&#8217;s 6th Gen EPYC CPUs and Instinct GPUs with VAST&#8217;s AI Operating System, with [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":6555,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[146],"tags":[207],"tmauthors":[],"class_list":["post-6554","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai-data","tag-cto"],"_links":{"self":[{"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/posts\/6554","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/comments?post=6554"}],"version-history":[{"count":0,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/posts\/6554\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/media\/6555"}],"wp:attachment":[{"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/media?parent=6554"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/categories?post=6554"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/tags?post=6554"},{"taxonomy":"tmauthors","embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/tmauthors?post=6554"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}