Kioxia's new PCIe 6.0 SSDs want to be the cheap alternative to expensive AI memory

Alfonso Maruccia

Posts: 2,662   +1,009
Staff
Gone In A Flash: Generative AI and other LLM-based tools are facing a significant data bandwidth challenge. Computational power continues to grow exponentially, but the massive amounts of data required for AI training and inference still has to pass through interfaces and storage technologies with limited bandwidth capabilities. Kioxia is proposing a novel solution by bringing NAND flash memory closer to the GPU.

Japanese memory manufacturer Kioxia is bringing a new class of super-fast SSDs to the enterprise market. The Kioxia GP1 Series is based on the PCIe 6.0 interface and can leverage the NVMe 2.2 protocol to serve as a middle ground between "pure" NVMe storage solutions and high-bandwidth memory VRAM soldered directly onto a GPU or AI accelerator's PCB.

The GP1 Super High IOPS SSDs can achieve 10 million input/output operations per second, Kioxia explains, thanks to the company's second-generation XL-FLASH memory. Unlike traditional TLC memory chips used in consumer products, XL-FLASH is designed to deliver higher IOPS performance with 512-byte block sizes. Furthermore, Kioxia is already working on future XL-FLASH generations that could increase data rates to 100 million IOPS.

The GP1 SSDs are optimized for direct GPU access through the PCIe interface. The drives should improve computational efficiency by allowing AI accelerators to quickly access larger datasets stored on flash memory, without requiring the massive expense of expanding the HBM memory pool attached to the GPU's PCB.

The new SSDs are available in E3.S and E1.S 9.5mm/15mm form factors, with support for either traditional air-cooling solutions or cold-plate liquid cooling (available only for E3.S and E1.S 9.5mm models). The drives offer storage endurance of up to 50 drive writes per day.

According to Neville Ichhaporia, senior vice president and general manager of the SSD unit at Kioxia America, XL-FLASH memory can deliver very high performance and low-latency operation, allowing it to serve as an effective memory extension tier for server GPUs. The cost per gigabyte is significantly lower than HBM or even DRAM-based memory expansions, which could be a major advantage for hyperscalers building new data centers in the current AI boom.

Chipmakers such as SanDisk and SK Hynix are attempting to overcome AI's memory wall by developing the novel High Bandwidth Flash memory solution. Meanwhile, Kioxia continues to invest heavily in traditional, albeit massively capable, storage formats such as its new 245TB LC9 enterprise SSDs.

Earlier this year, the Japanese company said that affordable consumer SSDs costing $50 per terabyte are now a thing of the past. However, the manufacturer is still working to bring solid-state storage solutions to budget-friendly PC builds with its recently introduced EG7 Series, which is based on quadruple-level cell (QLC) NAND technology.

Permalink to story:

 
It would be nice if AI companies started to perfect computing power usage.
It would be nice if they could make much cheaper hardware that only cost
a fraction of what they need today. Then maybe, we could get back our
freaking computers and phones without costing 3 times the fair price.
 
It would be nice if AI companies started to perfect computing power usage.
It would be nice if they could make much cheaper hardware that only cost
a fraction of what they need today. Then maybe, we could get back our
freaking computers and phones without costing 3 times the fair price.
So there are some interesting I'm seeing specifically with power usage coming on the next generation of AI hardware. I think the major problem with that is that the current generation of AI hardware has yet to pay for itself.

I buy lots of used server hardware and there are Chinese companies that will remove chips sever chips and put them on something to a desktop/workstation format. This was really big with Xeons a few years ago.

So what I think we will see lots of Nvidia chips being resoldered and making there way to the secondary market. All this stuff is water-cooled and the only mod available is basically taking a car radiator and a sump pump then using a bunch of plumbing adapters.

But I really don't see this current gen making its way to the used market quickly when it's still yet to pay for itself. I also think we are going to see more realistic business models surrounding AI about 3 to 4 years for now. But a concerning trend that I've seen is that there are now digital market places where datacenters are trying to sell their unused capacity at an alarming rate. The cost per token is going down in these market places when the hardware costs are still on the rise.

The real interesting stuff to me right now are these mindblowing pcie6.0 network cards and drives. I even saw a prototype GPU(not in person, I don't think it's a real product yet) but it was a blackwell GPU that had 48PCie lanes on it where the extra lanes were used for low latency connectioms to directly to other GPUs
 
Are these "massive amounts of data" in the room with us now?

Fake Artificial Stupids already work and run on local systems, phones, and even tiny watches. There is no "massive amounts of data" to be processed. They aren't real. The demand isn't real, and there are exactly zero AIs in existence today.

Just like "quantum computers" and "qubits", it's all literally vaporware.
 
Back