{"id":6444,"date":"2026-07-23T23:39:05","date_gmt":"2026-07-24T03:39:05","guid":{"rendered":"https:\/\/workai.tv\/news\/2026\/07\/ai-data\/instinctive-and-epyc-vast-data-sets-up-widespread-amd-cpu-gpu-collaboration\/"},"modified":"2026-07-23T23:39:05","modified_gmt":"2026-07-24T03:39:05","slug":"instinctive-and-epyc-vast-data-sets-up-widespread-amd-cpu-gpu-collaboration","status":"publish","type":"post","link":"https:\/\/workai.tv\/news\/2026\/07\/ai-data\/instinctive-and-epyc-vast-data-sets-up-widespread-amd-cpu-gpu-collaboration\/","title":{"rendered":"Instinctive and EPYC; VAST Data sets up widespread AMD CPU\/GPU collaboration"},"content":{"rendered":"<h2>Share with your CTO<\/h2>\n<p>VAST Data is formalizing a broad partnership with AMD, committing its AI Operating System to run on AMD EPYC CPUs and Instinct GPUs across training, inference, and agentic workloads. The headline technical move is adopting 6th-gen EPYC processors for VAST&#8217;s storage hardware platforms, gaining PCIe Gen 6 bandwidth that doubles I\/O throughput versus the prior generation. Early testing on an AMD Instinct MI355X showed a 9x speedup in time-to-first-token and 9.7x more token throughput when offloading KV cache to <a href=\"https:\/\/www.blocksandfiles.com\/flash\/2026\/07\/23\/instinctive-and-epyc-vast-data-sets-up-widespread-amd-cpu\/gpu-collaboration\/5276173\" target=\"_blank\" rel=\"noopener nofollow\">VAST&#8217;s storage layer<\/a> over NFS\/RDMA, compared to local host RAM.<\/p>\n<h2>What this means for your business<\/h2>\n<p>If your AI infrastructure is currently Nvidia-only, this announcement is less about AMD winning and more about VAST signaling that it intends to be the data layer for whatever GPU fleet you run. Organizations already evaluating AMD Instinct as a cost alternative to Nvidia H100 or B200 clusters now have a credible, named-customer answer to the software integration question. Crusoe, Vultr, and Core42 are already running AMD-powered services on VAST, which means the reference architecture is not purely theoretical.<\/p>\n<p>The KV cache story is the part worth examining closely. KV cache, the temporary memory store that lets a model recall earlier context in a long conversation or agentic task, has historically lived in GPU HBM or host RAM because latency requirements were too tight for network storage. VAST&#8217;s claim is that NFS over RDMA, combined with AMD&#8217;s Pensando Pollara 400 NIC, has closed that gap enough to make remote storage viable. If true, it shifts the economics of large-scale inference materially: GPU HBM is expensive and finite, while NVMe-backed storage is cheap and elastic. The 9x TTFT speedup is real but framed relative to a specific baseline, and Blocks and Files, whose editorial model depends on storage vendors remaining strategically relevant, has an interest in the storage-as-inference-tier thesis landing well. That doesn&#8217;t make the benchmark wrong, but it does mean your team should validate the comparison conditions before treating 9x as a procurement input.<\/p>\n<p>The automated KV cache lifecycle management feature, which ties cache expiration to data governance policies, is the detail most CTOs will under-weight and most CISOs will eventually demand. Agentic workloads cache sensitive user context by design, and regulatory pressure on AI data retention is arriving faster than most infrastructure teams have planned for. Whether you are on AMD or Nvidia today, the vendor that owns your data layer will own your compliance posture for inference. That is the renewal conversation worth preparing for now.<\/p>\n<h2>Concept deep-dive: KV Cache Offloading<\/h2>\n<p>A KV cache stores the intermediate computations a language model generates while processing a long prompt or a multi-turn conversation, so the model does not have to recompute that context with every new token it generates. Think of it as the model&#8217;s working memory for a single conversation. Normally it lives on the GPU itself. Offloading moves it to external storage over a fast network, freeing GPU memory for more concurrent users, at the cost of retrieval latency. The business case is more users per GPU dollar.<\/p>\n<p><em>Based on reporting from <a href=\"https:\/\/www.blocksandfiles.com\/flash\/2026\/07\/23\/instinctive-and-epyc-vast-data-sets-up-widespread-amd-cpu\/gpu-collaboration\/5276173\" target=\"_blank\" rel=\"noopener nofollow\">Instinctive and EPYC; VAST Data sets up widespread AMD CPU\/GPU collaboration<\/a>, originally published 2026-07-23 13:30:00.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Share with your CTO VAST Data is formalizing a broad partnership with AMD, committing its AI Operating System to run on AMD EPYC CPUs and Instinct GPUs across training, inference, and agentic workloads. The headline technical move is adopting 6th-gen EPYC processors for VAST&#8217;s storage hardware platforms, gaining PCIe Gen 6 bandwidth that doubles I\/O [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":6445,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[146],"tags":[207],"tmauthors":[],"class_list":["post-6444","post","type-post","status-publish","format-standard","has-post-thumbnail","category-ai-data","tag-cto"],"_links":{"self":[{"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/posts\/6444","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/comments?post=6444"}],"version-history":[{"count":0,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/posts\/6444\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/media\/6445"}],"wp:attachment":[{"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/media?parent=6444"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/categories?post=6444"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/tags?post=6444"},{"taxonomy":"tmauthors","embeddable":true,"href":"https:\/\/workai.tv\/news\/wp-json\/wp\/v2\/tmauthors?post=6444"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}