- Nvidia’s SCADA framework and open-source cuFile APIs aim to ease AI inference bottlenecks by letting GPUs access storage directly
- The system could reduce DRAM pressure and improve AI serving economics, but analysts warn it won’t fix the broader memory shortage
- By pushing AI storage interfaces toward open industry standards, Nvidia could extend its influence across the AI infrastructure stack
Nvidia hasn’t magically solved the memory crunch, but it’s doing the next best thing: buying time for existing systems via some nifty tricks involving storage. And unlike a magician, it’s revealing its software secrets via an open-source release.
In traditional systems, CPUs and Dynamic Random-Access Memory or DRAM have acted as the middlemen between the GPU and storage. But in a technical blog, Nvidia detailed how its scaled, accelerated data access (SCADA) framework and cuFile APIs can allow GPUs to read and write from storage directly.
It turns out this tweak – allowing the GPU to talk directly to storage without memory in the middle – is a big deal for two reasons.
First, the old way, while functional, is “nowhere near fast enough to matter in terms of serving valuable tokens today,” WEKA Chief AI Officer Val Bercovici told Fierce. Getting rid of the intermediary makes the whole process faster and more efficient.
“It’s not necessarily a magic bullet. It’s a strategy that says we can make mass memory access faster because we can eliminate the path through the CPU,” J. Gold Associates Founder and Principal Analyst Jack Gold said.
Second, AI inferencing is highly memory dependent, Bercovici said. And given memory is both extremely costly and in short supply at the moment, finding ways to reduce the amount of memory needed for certain tasks can help.
“It’s not really expanding memory; it’s bypassing,” Moor Insights and Strategy VP and Principal Analyst Matt Kimball told Fierce.
Kimball said Nvidia is essentially tackling the same problem AMD did with its MEXT acquisition. AMD’s approach (via MEXT) is to have software predict what data an app will need and proactively move it from flash to DRAM. “So, the operating system sees what looks like a bigger memory pool without any code changes. Nvidia is solving the problem from the other direction,” he said.
But Kimball stressed neither approach is actually adding memory to the market. Rather, “they reduce the amount of DRAM needed per unit of work, which certainly helps. But in the aggregate, the impact is minimal, and we will not see significant price relief as a result.”
In a nutshell, Kimball said moves like these will “definitely buy some headroom” in a market where every little bit counts and could lead to a “real improvement in inference serving economics.” But they’re not “going to impact cost in the way that vendors would like us to believe.”
Open avenues
Gold told Fierce that what Nvidia is doing isn’t necessarily new. But what is new is the fact that the company has decided to release its cuFile APIs in open source, with Google, Intel, NVIDIA and Meta named as inaugural maintainers.
Nvidia is also launching the Storage-Next program, which aims to define storage architectures and software interfaces for AI workloads and turn them into open industry standards.
Open sourcing its technology means others will be able to tap into the benefits of the system Nvidia developed. But Gold noted the move likely not altruistic.
“Politically it’s good. It makes them look like they want to play fairly with everyone,” he said. “Number two, it gets their stuff that they’ve invented to become an industry standard if it really takes hold, and that’s a benefit to them.”
Read more about memory and storage here:
SambaNova targets AI inference boom with chips built for existing data centers
Broadband groups warn White House memory prices could upend supply chain
Memory shortage puts the squeeze on telco modernization efforts
AI growth is about to hit a (memory) wall
Podcast: Data Center 2.0—Decoding modern data storage
VAST Data raised another $1B. Here’s where its CEO says the money’s going