Jensen Huang, chief executive of Nvidia, denied reports about delays of the company's next-generation AI platform and said that production volumes of the upcoming Vera Rubin platforms are 'giant.' He didn't address reports about delays of Vera Rubin Ultra-based rack-scale systems carrying 144 AI GPUs.
"[The reports about Vera Rubin delays are] not true," Huang told reporters on the sidelines of an event in Japan, reports Bloomberg. "Vera Rubin is already in production. Giant amounts of production incoming."
Nvidia confirmed production of its Vera Rubin platform in January and then sampling in February, so the current comment reiterates what we already know. Nvidia stressing that 'giant amounts of production' are incoming is meant to reassure investors that the company is on track to sell a boatload of its next-generation Vera CPUs, Rubin GPUs, and Vera Rubin NVL72 systems in the coming quarters, which means more record-setting quarters.
What Huang did not address β or perhaps he wasn't asked β is Nvidia's rumored delay of its Kyber NVL144 rack-scale solution with copper interconnects due to the system's complex PCB midplane by more than a year from 2027 to 2028. An alternative dual-rack design has reportedly been canceled and an even larger CPO-based NVL576 configuration may also face delays or limited availability, the same report from SemiAnalysis claimed earlier this month. The setback could leave Nvidia's Rubin Ultra platform with a smaller NVLink scale-up domain than originally envisioned. Nvidia says its roadmap is intact.
The Kyber NVL144 architecture was designed to connect 144 Rubin Ultra GPUs using a copper-based NVLink 7 scale-up fabric, so the machine required a sophisticated PCB midplane to carry high-speed electrical links between the system's components. SemiAnalysis claims that this midplane was challenging to manufacture, leading to a delay. The report does not identify defective chips or problems with particular components mounted on the board, but specifically points to the manufacturability of the PCB infrastructure itself.
"Our roadmap is intact," a spokesperson for Nvidia told Tom's Hardware.
Nvidia's statement on the matter neither confirms nor denies the report, but indicates that the company will be able to offer products mentioned in its roadmap without revealing whether they also remain on their previously announced launch schedules.
(Image credit: Nvidia)
Nvidia reportedly considered another copper-based design, called NVL72x2, as an alternative to Kyber. The system would have placed two Oberon racks back-to-back to expand the size of the NVLink scale-up domain without using optical interconnects. However, SemiAnalysis says customers rejected the unusual design and operational requirements, but does not specify their individual objections that could include serviceability, cooling, cabling, and data-center layout.
Meanwhile, the planned NVL576 rack scale solution that was supposed to combine eight Oberon racks interconnected using co-packaged optics between NVSwitches has also been postponed, or shipped in relatively small quantities because of 'ongoing CPO challenges,' SemiAnalysis claims.
The existence of the planned NVL576 configuration suggests that Nvidia had been developing some form of CPO-enabled NVSwitch connectivity for the Rubin generation. In theory, similar optical switch-to-switch connectivity could potentially be used to join smaller GPU groups into an NVL144 system and bypass Kyber's problematic copper midplane. However, the available information does not clearly indicate whether the CPO technology intended for NVL576 could reproduce Kyber's topology, bandwidth, and latency characteristics, or whether it was sufficiently mature for high-volume deployments by potential NVL144 customers.
The reported Kyber delay comes on the heels of another report saying that Nvidia had canceled quad-compute-chiplet version of its Rubin Ultra in favor or a dual-compute-chiplet design that is projected to deliver 2X lower performance. With Kyber NVL144 delayed and NVL72x2 cancelled, Nvidia will only be able to offer 72-way scale-up systems till sometimes in 2028, meaning that AMD and Google may end up with more competitive scale-up systems in 2027 β 2028. AMD's Mega Pod based on the Verano CPUs and Instinct MI500-series accelerators, is expected to pack up to 256 accelerators. Google's TPU 8i can provide roughly 1,024β1,152 accelerators within one low-latency domain, whereas the TPU 8t goes much further and can get to 9,600 chip packages per domain.
Nowadays, storage devices for consumer and data center applications differ rather dramatically, as do approaches to product design as well as go-to-market strategies. Therefore, to get a more or less comprehensive overview of the storage market in general, you must observe both ends of the spectrum. To complement our interview with Nelson Duann at Computex, we also sat down with his colleague Alex Chou, who is in charge of Silicon Motionβs enterprise storage business.
Alex Chou is an interesting person to talk to. Before joining Silicon Motion, he spent some 18 years at Broadcom, where he led the wireless connectivity business, also initiating the Enterprise Switch, PoE, and 10-G Base-T PHY business with a product marketing focus. Before that, he worked at UMC Capital, ARK Logic, and Western Digital, where he developed graphics accelerators. He deeply understands the industry and uses his knowledge to expand SMI's business into the data center segment. As he is the first general manager of Silicon Motion's enterprise business unit, it is safe to say that all the success that the company has faced in the new segment so far can be attributed to Alex Chou.
Anton Shilov: Can you introduce yourself to our readers, please?
Alex Chou: My name is Alex Chou. As you know, Silicon Motion has two business units: the client business and the enterprise business. I am responsible for the enterprise business unit. My responsibilities include defining new products, leading development teams, bringing products to market, and working with OEMs, cloud service providers, and other customers to promote our technology and differentiation.
Getting into enterprise SSD business
Historically, Silicon Motion was focused on NAND controllers for client applications as well as embedded graphics processors and USB display controllers. Following the restructuring in the early 2020s, SMI formed a separate business unit to offer enterprise-grade SSD controllers, though it took the company some time to land its first tangible orders. By now, the company has yet to grab a 10% market share, yet it has clients among cloud service providers (CSPs), hyperscalers, and OEMs, significant achievements given Silicon Motion is a relatively new market entrant.
Anton Shilov: It has been a challenging year for much of the industry, particularly for memory-related segments. Yet Silicon Motion reported first-quarter revenue of $342.1 million, up 23% sequentially and 105% year-over-year, while SSD controller sales increased by roughly 40% to 45%. Can you explain what drove those results, particularly on the enterprise side?
Alex Chou: It depends on how you define a difficult year. If you look at the results, I would argue that this has actually been one of the best years the storage industry has seen.
Silicon Motion is fundamentally a controller company. We build controllers that work with NAND from all major memory suppliers. On the enterprise side, we are still relatively new compared to some established competitors, but we have secured a number of new projects and have started delivering products to customers.
We have invested heavily in PCIe Gen5, Gen6, and Gen7 enterprise SSD controllers. Today, our Gen5 products are beginning to ramp into volume production with multiple OEM customers. That ramp is contributing to our growth.
Anton Shilov: Do you have an estimate of your market share in the enterprise SSD controller market?
Alex Chou: That depends on how you define the market. Some people measure market share by unit shipments, while others look at exabytes shipped because SSD capacities continue to increase.
We have only recently begun shipping enterprise products in volume. If you listened to our CEO's comments during the earnings call, we expect enterprise shipments to increase significantly in the second half of the year. We are still in the early stages of our ramp, but we are making good progress with several key customers.
If you look beyond the initial ramp and think about the full-year run rate, I believe we can build from there and target a much stronger position next year. Longer term, our goal is to exceed 10% market share in the $4B enterprise SSD controller market, but this year is really about getting through qualification, customer testing, and the early production ramp in 2 half of this year.
Our goal is to continue expanding our share. We are only beginning the ramp [of our data center-grade SSD controllers] today, but we expect our share to increase meaningfully as deployments grow.
Anton Shilov: Who are your primary customers? SSD manufacturers, OEMs, or hyperscalers?
Alex Chou: We primarily work with OEMs. We sell controllers and firmware solutions to SSD manufacturers and OEMs. Some customers use our complete controller-and-firmware solution, while others develop their own firmware.
At the same time, we work directly with hyperscalers and cloud service providers to explain the advantages of our products and ensure they understand our technology roadmap.
Enterprise SSDs are used in several different segments. Traditional compute servers represent one market. High-density storage systems used for AI and large-scale data storage are another. We also see growing interest in storage systems located near GPUs, where latency becomes particularly important.
One area where we differentiate ourselves is quality of service. We have developed a patented traffic-shaping engine that helps maintain latency consistency under heavy workloads and multi-tenant environments. That capability is particularly attractive to hyperscalers and cloud service providers.
Anton Shilov: Do you see the enterprise SSD market splitting into different categories depending on workload?
Alex Chou: Yes. We see at least three major categories emerging.
The first is traditional compute-attached enterprise SSDs, which are used in conventional servers and storage systems. The second is very high-density storage for AI and hyperscale environments, where capacity, throughput, and cost efficiency are critical. The third is storage located closer to GPUs, where the requirements are very different because latency and quality of service become much more important.
That third category is particularly interesting. In AI systems, the storage subsystem is no longer just feeding CPUs. It increasingly has to support GPUs directly, especially for workloads involving very large datasets or KV-cache offload. In those environments, low latency and predictable performance matter much more than they did in traditional storage deployments.
Storage Next, PCIe 6 and PCIe 7 SSD controllers
Anton Shilov: Is that where Nvidia's Storage Next vision comes in?
Alex Chou: Yes. Storage Next is one of the major industry developments we are watching very closely.
The idea is that storage will move closer to the GPU and become part of a much more tightly integrated data path. In some cases, the goal is not just to maximize bandwidth, but to ensure that latency remains low and deterministic enough for AI workloads that continuously move data between accelerators, system memory, and storage.
This is one of the reasons we have invested heavily in QoS and latency control. Through our traffic-shaping technology, we can manage access patterns and reduce latency spikes when multiple tenants or applications share the same SSD. In a cloud environment or an AI storage environment, that becomes very important.
(Image credit: Silicon Motion)
Anton Shilov: So, the challenge is no longer just raw throughput, but how predictably the SSD behaves under load?
Alex Chou: Exactly. Bandwidth still matters, but in many enterprise and AI environments, consistency matters just as much.
When multiple applications, multiple VMs, or multiple users share the same storage device, you need to control latency and quality of service carefully. If performance becomes unpredictable, it can affect the entire system.
That is why we have focused on a traffic-shaping mechanism that can prioritize and isolate workloads more effectively. We believe that kind of latency management will become a key differentiator for enterprise SSD controllers going forward.
Anton Shilov: How does that affect your roadmap for future controllers?
Alex Chou: It affects it quite a bit. Our upcoming controllers are not designed only for higher sequential bandwidth. They are also being designed for newer enterprise requirements such as OCP 2.7 compliance, stronger security, better QoS, and support for more advanced deployment models.
Anton Shilov: Are you already sampling your PCIe 6.x controllers?
Alex Chou: On the Gen6 side, our controller design is essentially complete; we have an FPGA [emulating algorithms], and we expect tape-out very soon. If everything goes according to plan, we expect first silicon back in the second half of 2026.
That controller not only supports a faster host interface, but also supports new features and requirements we see from AI infrastructure and hyperscale customers.
Anton Shilov: So, the PCIe Gen6 SSD platform is not just a speed upgrade for Silicon Motion?
Alex Chou: Correct. PCIe Gen6 obviously provides more bandwidth, but for us the more important part is that the surrounding system requirements are changing as well. Security, QoS, cloud deployment models, and AI storage architectures are all evolving at the same time, so the controller has to evolve with them.
Anton Shilov: Let us talk about the roadmap in more detail. You said the PCIe Gen6 enterprise controller is close to tape-out. What comes after that?
Alex Chou: PCIe Gen6 is the next major step for us, and the design is essentially complete. We expect to tape out very soon and, assuming [everything works correctly], receive first silicon in the second half of 2026.
But internally, we are already working beyond PCIe Gen6. PCIe Gen7 development has already started. In fact, the overall architecture for our Gen7 enterprise controller platform has already been defined. That means we are not just planning the interface speed increase; we are also defining the surrounding architecture, feature set, and deployment model that will be needed in the next generation of enterprise and AI systems.
Anton Shilov: So, SMI's PCIe Gen7 controller is no longer just a concept?
Alex Chou: Correct. PCIe Gen7 is already in active development. The current plan is to have internal samples in 2H, 2027 and to move toward production in that same general timeframe.
As controller development becomes more complex, you cannot wait until the market is ready before starting work. By the time a new interface reaches the market, the controller has to be nearly finished already. So, we are always working at least one generation ahead, and in practice often two.
Anton Shilov: As NAND becomes denser and more complex, error correction also becomes a bigger issue?
Alex Chou: That is a major part of controller development now. As NAND moves to higher layer counts and denser cell structures, the controller has to do more work to maintain reliability, endurance, and data integrity.
One of the areas we are working on is stronger LDPC. On the enterprise side, LDPC with a 16KB collaborative codeword is already used with SM8466, SMIβs first Enterprise PCIe Gen6 controller, and it is part of the roadmap because future NAND will require more robust error correction. That is one of the reasons enterprise controller architecture keeps becoming more complex generation after generation. You are no longer designing only for interface speed. You are also designing for signal integrity, power, security, QoS, error correction, and support for future NAND generations that may behave very differently from today's devices.
Anton Shilov: Will LDPC with 16KB collaborative codeword be enough for next generations of 3D NAND with hundreds of active layers?
Alex Chou: A 16KB LDPC engine already consumes a significant amount of silicon area and is quite sophisticated. For PCIe Gen7 controllers, our goal is to optimize and improve that engine from multiple angles rather than simply keep expanding it. We still need our architects to make the final call on exactly which improvements we will implement, but at this point we are more likely to refine and enhance the current design than to move beyond 16KB LDPC.
SSD controller development strategy
Anton Shilov: Speaking more generally, SSD controllers are increasingly becoming full platforms rather than just controllers, because integration matters so much. Do you expect close collaboration between controller vendors, NAND makers, and SSD manufacturers to become even more important as the industry moves to next-generation storage devices?
Alex Chou: I may not fully understand your question, but let me explain how we approach it.
At Silicon Motion, we design the controller architecture and build the firmware stack with a rich feature set. For example, we have developed our own [PerformaShape] traffic-shaping engine to improve QoS. That is the foundation of the platform.
From there, we have to look at how NAND evolves from one generation to the next. As we move from PCIe Gen5 to Gen6 to Gen7, controller performance has to scale accordingly. If you want to saturate the PCIe interface and deliver, say, 7 million IOPS today and much higher performance in future generations, you have to understand exactly where NAND is going.
That is why my team meets regularly with Samsung, SK hynix, SanDisk, Kioxia, and all other NAND vendors to review their roadmaps. Silicon Motion is part of that ecosystem, and because of those relationships, we usually get early visibility into future NAND generations and often receive early samples so we can bring up our controllers and make sure they take advantage of new NAND as quickly as possible.
That matters even more in the current supply environment. Because we work with all NAND suppliers, hyperscalers and cloud service providers can come to us and ask for a solution that is not tied to a single memory vendor. A company like Samsung naturally builds around its own NAND, but we have the advantage of being able to support multiple suppliers. That gives customers much more flexibility when supply is tight.
So yes, we have a core controller architecture and a common firmware base, but one of our strengths is that we work very closely with NAND vendors on future generations and make sure our platform can take advantage of faster interfaces, higher die counts, and new NAND capabilities as they arrive.
XL-Flash and storage-class memory
Anton Shilov: What about storage-class memory? Are there any developments there? As far as I can tell, adoption of Kioxiaβs XL-Flash has been limited.
Alex Chou: Thatβs a very good question. I am actually going to visit Kioxia, so I should have a better sense of their plans after that. At the moment, Kioxia is essentially the only company still pushing XL-Flash, so they are trying to build something around it.
The challenge is that it is not just about the technology itself. You need a broader ecosystem to support it, and that is what makes the situation more complicated. We are watching it closely and trying to understand whether it is something we really need to support, but at this point I do not have a definitive answer. We are still evaluating it.
Anton Shilov: Have you heard anything similar from other suppliers? Quite a few memory makers used to talk about storage-class memory or similar technologies in their roadmaps.
Alex Chou: Based on what we know, not really. If you look back at last yearβs Flash Memory Summit, several NAND makers were talking about higher-performance flash and storage-class-memory-like concepts. That created a lot of buzz at the time, and we looked into it, just as we have looked into XL-Flash, to understand whether there was a real ecosystem forming around it.
But there is much less discussion around those ideas now. One reason is simple: memory vendors do not really need those products at the moment because they can sell conventional NAND at very high prices and still generate strong returns.
Anton Shilov: In other words, they can just sell QLC 3D NAND and be perfectly happy.
Alex Chou: Exactly.
Anton Shilov: On the other hand, Nvidia wants storage devices capable of 100 million IOPS.
Alex Chou: Yes, that is where Storage Next comes in.
Anton Shilov: Has anyone actually come close to 100 million IOPS yet?
Alex Chou: I would say Storage Next gains many attentions. XL-Flash could be one possible approach to address that kind of requirement. But these are other options aiming to address high-performance and low latency needs.
What matters more is that Storage Next has a much stronger ecosystem behind it because Nvidia is actively driving it. There are regular meetings around it, and our architect has been involved from the very beginning. We have been tracking it closely and trying to make sure our future controller architecture can support it if and when the market materializes.
At the same time, Nvidia itself appears to recognize that 100 million or 200 million IOPS may not be realistic in the near term. The target seems to be moving closer to something like 50 million IOPS, which is more achievable. So yes, we are watching it very closely, and we are building in the flexibility to support it if needed.
In storage, having a technically interesting idea is not enough. The industry has to agree on how to use it, how to deploy it, and how to integrate it into systems. Storage Next currently has more momentum because the ecosystem behind it is much stronger.
Anton Shilov: So, you see Storage Next as more commercially relevant than storage-class memory, at least for now?
Alex Chou: Yes. At least today, Storage Next looks more immediate and more actionable.
We are already participating in those discussions and thinking about what future controller requirements will look like in that environment. That includes not only bandwidth, but also latency behavior, QoS, and the role storage plays in systems where GPUs are increasingly central to the data path.
That does not mean other technologies disappear. It just means that if you ask where the market is actively moving right now, the answer is much more on the Storage Next side than on the storage-class-memory side.
Anton Shilov: So, in practice, you make sure your controller works with all relevant NAND types, while the memory vendor mainly has to make sure the media itself complies with the interface requirements?
Alex Chou: When we design a controller, we already cooperate closely with NAND suppliers. Our architects look at all of the major vendors to understand whether there are any special requirements we need to account for. Then we handle another layer of optimization in firmware to make sure we can support all of those devices properly.
If you look deeper into enterprise NAND, most products also use interface chips internally to connect large numbers of dies. Those interface chips can differ from vendor to vendor, so we need to understand their configurations as well, including die counts, planes, and other architectural details. The goal is to make sure the controller and firmware together can support all of those different combinations.
So far, our architecture has been able to support NAND from SanDisk, Kioxia, SK hynix, and the other major vendors. Even if the interface chips differ, we try to keep the overall hardware design as flexible as possible.
There are really three elements involved: the controller itself, the hardware board, and the firmware. Ideally, you do not want a completely different board design for every NAND supplier. Fortunately, the industry has standardized a lot of the pinouts and module interfaces, which makes it possible to use a common hardware design and swap in NAND from different suppliers with the right firmware support.
We spend a lot of time making sure we can support all of those different combinations.
Anton Shilov: So you are effectively building controllers with a fairly clear view of what future NAND generations will look like.
Alex Chou: Exactly. We want to make sure that when the next generation of NAND arrives, we are ready to support it as broadly as possible.
The adventurous developer who recently released Linux for the Atari Jaguar (1993) has brewed up a version of the open-source OS for the Sega 32X (1994). If you canβt remember the 32X, it was Segaβs mid-gen answer to early fifth-generation challengers like the Jaguar, 3DO, and Amiga CD32. Segaβs solution added some potent processing power to a mushroom-like slot in an expansion to its popular Mega Drive/Genesis. Now cakehonolulu has got it running Linux, despite facing several major hurdles.
Compared to its Genesis host, the Sega 32X was incredibly muscular. The Genesis had relied on the capable but long-in-the-tooth Motorola 68000 (7 MHz) for years, but the 32X add-on boosted that with a pair of Hitachi SuperH SH2s (SH7604) CPUs (23 MHz). It also ramped up system RAM from the base 64KB by adding 256KB of its own. Segaβs expansion offered more than just speed; the consoleβs color palette was ramped up from 64 to 32,000 simultaneous colors on screen, and it was powerful enough to introduce hitherto unachievable 3D graphics elements to mainstream console gaming.
As with cakehonoluluβs tale of Linux wrangling on the old Jag, the above-linked blog talks through a long list of hurdles that needed to be leapt to get the Linux kernel booted and running BusyBox. This time around, particularly steely roadblocks included: the even more constrained RAM situation, the lack of hardware synchronization primitives, the desire to get SMP working across the pair of SH2 CPUs, no direct UART access from the 32X, and scheduler bugs, among other things.
On the positive side, smoothing the development process along were access to Chilly Willyβs 32X devkit, the linuxmd project, the Krikzz FPGA-based flash cart with ROM β RAM mapping tools, and existing SH2 documentation and sample projects. Please check through the linked blog for far more technical details from cakehonolulu.
Linux booted with BusyBox prompt on the Sega 32X (Image credit: cakehonolulu)
As you can see, cakehonolulu was successful again. So, whatβs the next stop for this adventurous dev β the Sega Saturn? Whatever the project may be, it was interesting to read that works like this are basically forming a portfolio for the Spanish dev, which they hope will help them with job hunting.
The Sino-American chip wars have resulted in many back-and-forth salvos and negotiations as the countries try and strike a balance between technology access and trade. Currently, both sides have set respective import and export controls, letting specific companies on a case-by-case basis. Today, Reuters reports that Chinese telecoms giant ZTE and server firm Maginfra have received U.S. approval to buy Nvidia's last-gen H200 "Hopper" chips.
ZTE joins a club that counts Alibaba, Tencent, ByteDance, and JD.com among the roughly 10-strong group of Chinese companies with U.S. clearance for those purchases. Additionally, an apparent subsidiary of Kingsoft Cloud got approval to buy AMD accelerators equivalent to Nvidia's H200, presumably Instinct MI300X-class chips.
Over on the Chinese side of the table, Reuters remarks that there's no word on whether the respective authorities will give ZTE the go-ahead for import, as the country has taken on a protectionist stance as it tries to grow its own chip industry. The country has discouraged firms from purchasing foreign tech and has instead pushed companies to acquire homegrown accelerators. Huawei in particular has made great strides both technologically and financially.
But even with those domestic production initiatives, the Chinese hunger for AI silicon is so deep that six months ago, Reuters said the nation's tech firms had more than two million H200 chips on order, far more than what Nvidia had on hand at the time. We'd venture that hunger has barely subsided.
ZTE might not be a familiar name Stateside, but the corporation is one of China's largest telecommunication conglomerates, and among many other ventures, it sells all sorts of carrier network gear that's installed worldwide, along with corresponding client-facing equipment, including phones and IoT equipment. Like most any sizable technological venture, ZTE has joined in on the cloud computing and AI push, so it needs accelerators to make those ambitions reality.
The current status of the AI chip trade situation is roughly that the U.S. allows Chinese firms to buy AI chips up to and including the Hopper family (meaning no Blackwell chips), with a 25% export tariff, though final decisions are made on a case-by-case basis. Over on Chinese shores, Beijing's authorities play their cards close to their chest and dole out approvals as they see fit, with no clear rules seemingly set. But China is, of course, a global power with trade connections to most everyone, so interested firms were able to get their hands on Blackwell chips through various creative (and potentially illicit) means.
Whether this change will actually clear the way for any great volumes of H200 accelerators to make their way into ZTE's data centers remains to be seen. CNBC cites a U.S. trade official who today stated that "very few shipments against licenses for H200s and equivalents have taken place. Itβs a very small quantity of chips" during a congressional hearing. If H200 shipments become material to Nvidia's bottom line, we'll almost certainly hear about it in future comments or earnings reports.
Samsung is back with another solid-state drive, and this time it's something a little bit different. The 990 is a QLC-based 990 EVO Plus, positioned as a budget drive that can still push a lot of bandwidth. Itβs a little late to the game and not quite what was rumored for the 990 QVO, but it does bring some new technology to the table. Weβre always interested in seeing what Samsung puts out, and this time is no different. It should not be confused as being part of Samsungβs Pro line or, for that matter, the EVO line, so keep that in mind.
The drive has its ups and downs, but in this challenging market, and for a budget drive, thatβs to be expected. Samsung is still well-regarded for its name and reliable hardware, even as there has been a massive push towards enterprise, away from the consumer side. Samsung has, in fact, given some ground in the SSD space for many years, even as it produces some of the most common OEM drives. So while this is not a Crucial situation, itβs best to jump into this review with the right expectations about what this drive is and isnβt. Itβs a budget drive with full Gen 4 throughput that hits the most common capacities with sufficient performance and power efficiency. Itβs not meant to be a throne-taker.
Itβs also thankfully not another 990 EVO situation β that drive felt somewhat underwhelming by the time it arrived, even when pitted against budget drives β but the 990 is also not a QLC rallying call. Itβs a competent drive that mostly hits the right notes, as intended. Given how scarce Samsung QLC drives have been, and how much demand its QLC flash surely has elsewhere, it can feel like Samsung is throwing consumers a bone, though it would be crass to put it that way. We instead think this is smart positioning by the company as it knows the future is with QLC and the technologies used in this flash (even if first shown two years ago at ISSCC) point firmly at an ambitious future. The 990 just lets you own a piece of that.
Samsung 990 Specifications
Product
1TB
2TB
Pricing
$269.99
$529.99
Form Factor
M.2 2280 (Single-sided)
M.2 2280 (Single-sided)
Interface / Protocol
PCIe 4.0 x4 / NVMe 2.0
PCIe 4.0 x4 / NVMe 2.0
Controller
Samsung PiccoloQ
Samsung PiccoloQ
DRAM
N/A (HMB)
N/A (HMB)
Flash Memory
Samsung V9 QLC
Samsung V9 QLC
Sequential Read
7,150 MB/s
7,250 MB/s
Sequential Write
6,450 MB/s
6,450 MB/s
Random Read
700K IOPS
850K IOPS
Random Write
1,100K IOPS
1,200K IOPS
Power (R/W)
4.0W / 3.7W
4.3W / 3.8W
Endurance
400 TBW
800 TBW
Security
TCG Opal V2.0
TCG Opal V2.0
Part Number
MZ-V9V1T0
MZ-V9V2T0
Warranty
3-Year
3-Year
The Samsung 990 is only available at 1TB and 2TB capacities, with MSRPs of $269.99 and $529.99, respectively. These prices are very high, as you can get competing drives like the Crucial P310 for substantially less, and in fact even the TLC-based WD Black SN7100 costs less. But Samsung has historically launched with MSRPs well above actual market price. You should be able to get the drive at significantly lower prices after launch, but the βSamsung taxβ may still apply. Weβll get into what that means throughout the review.
This limited capacity range is unfortunate, but enables Samsung to pack the flash into just one package, which reduces PCB space so that any OEM variant can be used in multiple M.2 form factors and will always be single-sided. Less than 1TB is also not enough for these denser dies if you want good performance. That leaves 1TB and 2TB as the target capacities, which also makes sense in a market where 4TB+ is getting exceptionally expensive. Weβll eventually see 2Tb dies to make single-package 4TB a reality, but thatβs further along in Samsungβs roadmap.
The drive can reach 7,250 / 6,450 MB/s for sequential reads and writes and up to 850K / 1,200K random read and write IOPS. Peak performance is attained at 2TB, where you have the optimal amount of interleaving or parallelization: Sixteen 1Tb dies means four dies for each of four flash channels, the typical ceiling. However, as these are four-plane dies, you still get 32-way interleaving at 1TB with eight dies, which is enough to get good performance with just two dies per channel. Less than that is much less ideal, and more than that introduces additional overhead, especially for budget controllers. The math changes with six-plane and 2TB dies, but for this flash, 1TB is the reasonable minimum, with 2TB offering the best performance.
The drive is rated for approximately 4W of power draw across the two capacities, when looking at both reads and writes. Check our power results below to see how accurate that is. The drive is rated for 400TB of writes per TB capacity, which is high for QLC flash β we would typically see maybe 300TB, which is one-half of the TLC standard β but also indicates a very high drive writes per day (DWPD) rating. This is due to the warranty only covering three years rather than the normal five, so the amount of writes per year is significantly higher. This is atypical, so requires further explanation.
For those who live for TBW and write endurance, this illustrates why TBW often looks better on paper. Spreading 400TB over three years works out to roughly double the daily write allowance of a typical 300TBW / five-year QLC drive. Most people will never approach either number, and they will live with the shorter coverage window. However, if you intend to hammer the drive with writes to the point of exceeding TBW within the three-year warranty period, then this could be good. Although you really shouldn't use a budget DRAM-less QLC-based drive for that type of workload. However, that option exists and is rarely the case with a QLC-based drive. As a final note, the drive does support TCG Opal 2.0 for encryption.
Samsung 990 Software and Accessories
Samsungβs Magician software is the gold standard for consumer SSDs. This is an SSD toolbox with all the features you need. It displays system and drive health information, including SMART, and checks whether your drive is legitimate. You can also benchmark your drive and use any optional features, such as encryption. The software is also essential for keeping the driveβs firmware up to date, although you can also download that from the first link.
Samsung 990: A Closer Look
Tom's HardwareTom's Hardware
The 990 has an SSD controller, a single NAND flash package, and power management circuitry. There is no DRAM package present. This is a single-sided drive, which is ideal for compatibility and cooling. There is a lot of free space on the PCB, and by putting distance between the controller and flash, there is separation to mitigate component heat generation. This would also help if a heatspreader or heatsink were to be added. Without this space, the drive could be sold in a shorter form factor, which is particularly useful for OEM drives.
The label has information about the drive, such as the date of manufacture (DOM), model, serial, the PSID, and the power rating. We always caution that you not take certain drive information as being conclusive about the hardware. For example, you should not assume TLC or QLC flash from a driveβs TBW. Likewise, you shouldnβt rely on the labeled power rating β and this is done more often on M.2 2230 drives for portable devices β as any indication of drive power efficiency. Here we have 3.3V / 1.85A, which indicates potential power draw over 6W. Now, the power ratings given on spec sheets will often be average and not peak, and will be separated as read or write rather than mixed. In fact, this driveβs load power states can reach a peak of 5.90W via SMART, which is much above the rated average ~4W. We track both peak and average in our testing.
SamsungSamsungSamsung
We always enjoy reviewing Samsung drives with a focus on the technicals, as the manufacturer remains a leader in many ways. The 990, in particular, requires some extra description to be fully appreciated. Simply looking at the benchmark results might make the technology seem underwhelming β to be honest, this is very much a budget drive, even taken in the best light β but that doesnβt mean Samsung phoned this one in. In fact, there are signs of deliberate design here, and some of the decisions could help sell this drive. Samsung still has to get the pricing right, of course, but what else is new?
Letβs start with the controller. The 990 is using the PiccoloQ, which is the QLC flash version of the Piccolo. The Piccolo is utilized on the 990 EVO and 990 EVO Plus, two TLC-based drives. In all cases, itβs a four-channel, DRAM-less design, which limits performance and capacity. In both cases, the controller takes up to 2,400 MT/s flash β this is more than enough to saturate PCIe 4.0 β and the interior design is the same. This means itβs a Samsung 5nm part with multiple ARM Cortex-R8 cores and a single R5 core. If the Piccolo stands out in any way, itβs that it offers a PCIe 5.0 x2 option in addition to the standard 4.0 x4 interface. This option or mode has limited usefulness, though, and nothing in the 990 would change that if enabled for the PiccoloQ.
So, not much new on the controller front, but the use of this controller at the 990βs rated speeds does give us some more information. Namely, we know the 990 EVO runs more slowly because itβs using flash slower than 2,400 MT/s, 1,600 MT/s Samsung V6P TLC, to be precise. If we look at Samsungβs V7 QLC flash, it can run at that same speed. This is why the originally speculated 990 QVO with that flash was targeted at the same speeds as the 990 EVO. Things have changed since then. This drive could have been the 990 QVO, but with the EVO and EVO Plus lines going DRAM-less this generation, we suspect the QVO tier was βpromotedβ to the plain 990 name, and the 990 now targets the 990 EVO Plus's specs
The evidence to back this up, which also supports the loose 990 QVO rumor, is that Samsung does have a V7 QLC OEM drive: the BM9C1. This is the cousin to the PM9C1 line with OEM 990 EVO and 990 EVO Plus (PM9C1b) variants. The BM9C1 is available down to M.2 2230 and uses the same PiccoloQ as the 990 (the QLC version of the 990 EVO/EVO Plusβs Piccolo). Itβs just limited to the same speeds as the 990 EVO, as itβs running at 1,600 MT/s. We have to be careful here, though, as Samsungβs V9 QLC press release indicates a 60% I/O improvement, which, with the V9 being 3,200 MT/s, suggests a 2,000 MT/s ceiling for the V7 QLC. Since there is an OEM TLC-based drive in between the 990 EVO and 990 EVO Plus (the PM9C1a) at 2,000 MT/s, the possibility for a ~6 GB/s 990 or 990 QVO with V7 QLC existed.
Before we dive more deeply into the flash, since we havenβt seen the new Samsung QLC in a while and there is some neat tech here, letβs decode the module. βK9β tells us itβs Samsung NAND flash memory. βYYGβ indicates itβs a QLC flash package with sixteen dies (HDP) in a 2TB configuration, which confirms 1Tb dies. βY8β means itβs 8-bit, J tells us the voltage, β5β tells us the number of chips enabled and ready/busy signals, and βDβ tells us the generation. With V7 being βCβ and V8 skipped, this suggests V9. The second part of the code tells us how the flash is packaged and that itβs commercial / consumer-grade. While you arenβt expected to know how to read codes on your SSD, knowing how it works can be useful, especially with Samsung drives, even if itβs just a matter of trying to figure out if you have a counterfeit product.
So letβs talk about the flash. This is a 286-Layer part, technically, but is sold as 280-Layer once accounting for source/ground and dummy lines. Dummy lines are usually at stack edges, as the physics of flash can make these lines otherwise unusable. A higher layer count β Samsungβs V7 is only 176-Layer, although technically 191 layers β generally means higher bit density. Bit density is key to scaling NAND flash, which is acting as capacious, non-volatile storage media. This can be disappointing to some because it means you donβt always see any real performance scaling as the layer count progresses.
Fitting more flash into the same space can mean less room for charge in each cell, which makes it harder to optimize for performance if youβre trying to maintain the same endurance level. That is certainly the case with this flash, as the performance only manages to match that of last-generation 176-Layer QLC flash from competitors, which is why we want to go out of our way to point out Samsungβs design decisions and why it leans innovative in ways you wonβt see in, say, your game load times.
For one, when we talk about the layer count difference β reported versus actual β you also get an efficiency number that is the ratio between usable and total word lines. Samsung is a leader here, with high layer efficiency. Samsung also has held off using three decks or stacks of flash and is still at two, due to having superior channel etching β itβs able to drill down more layers with a higher aspect ratio. Itβs also possible to run lines through the flash itself rather than rely largely on masked steps, which sets the stage for Samsung scaling to extremely high layer counts. One issue with high layer counts is that you start losing uniformity from layer to layer, and Samsung accounts for this with optimized word line spacing, too. So, as weβve said in the past, it often feels like Samsung is falling behind on layer count, but in reality it has a very focused strategy and the best technology in the business, and we can see this with the 990βs flash.
For the consumer, though, the 990 is a little bit weird. This is presumably 3,200 MT/s flash that is being βwastedβ with a 2,400 MT/s controller. This flash has amazing bit density, but having a single sixteen-die package at 2TB is nothing new. What about performance? Samsung has made optimizations to improve performance on this flash, but nothing amazing. This QLC is only comparable to the competition in performance terms, particularly at 2,400 MT/s. Samsung is playing catch-up, but we also think this is a case of designing for enterprise rather than consumer.
QLC flash is now highly sought after in enterprise for its density, and Samsungβs optimizations all benefit that kind of environment. In fact, from a consumerβs perspective you could look at this V9 QLC as being focused on higher bit density β but no 2Tb dies β and you would largely be correct. Samsungβs V9 QLC is 86% more dense generationally and about 94% more dense than the competitionβs 176-Layer QLC flash.
Weβll take a look at one new technology in the V9 QLC flash to illustrate. One important consideration is flash power interruption leading to data loss, which, without power loss protection (PLP) means you are looking at protecting data at rest. This is on the non-volatile media or flash, not the volatile memory like DRAM. When folding from the pSLC cache to the native flash, data loss is not an issue because you donβt invalidate the original pSLC copy until the write has been verified. However, when writing to native QLC, you are writing multiple pages where the upper pages will require higher levels of sensitivity for proper reading. There are different methods of writing to QLC flash, but generally multi-bit flash has multiple write passes that go from fuzzy (coarse) to precise (fine), and lower pages write faster and may be complete first. Therefore, itβs important not to ruin existing lower-page data if you lose power while still adjusting voltage for the upper pages.
Micron has a unique way of dealing with this using a differential engine that can predict values from partial shifts, but a more common method is simply to back up or buffer the values in nonvolatile flash. QLC stores four bits per cell, so a full backup means writing four bits of pSLC per cell. pSLC is used because its writes are fast, whereas QLC's upper-page writes, in particular, are an order of magnitude slower. Samsung reduces the buffer to a single parity bit by using an odd/even algorithm, creating a sensing window thatβs more like TLC (8-state) than QLC (16-state). This improves performance, endurance, and bit density. Some of that performance is still lost for higher bit density. For consumers, the direct benefit is higher TBW, but we speculate the higher density is aimed more at enterprise and future flash generation products. This is in part a response to Solidigmβs floating-gate design, a different technology than charge trap, with tighter charge placement.
The Samsung 990 enters a crowded market with a lot of good options, at least in theory. If weβre looking at QLC-based drives, this means the Crucial P310 and Sandisk WD Blue SN5100 at the very top. Both of these drives perform incredibly well. Below that, we have the older wave of drives represented by the TeamGroup MP44Q. That drive in particular remains a budget favorite with a fast controller and good QLC flash.
We would put the rest below that, even though the hardware is not always worse. This would include the Biwin M350, the Kingston NV3, and the Seagate FireCuda X1070. These drives are using alternative controllers β SMI, SMI, and TenaFe, respectively β that are roughly comparable, and the flash is not particularly old, either. However, these drives tend to be more budget-focused with reduced performance and (ideally) reduced cost.
Weβve also thrown in Samsungβs 990 EVO and 990 EVO Plus for comparison. The 990 should be closer to the latter, but with QLC flash, it would be okay landing somewhere in between. On the whole, we would expect the drive also to be between the two main categories of drives β that is, above the budget ones, below the two fastest, and closer to the middle MP44Q and its MAP1602-equipped alternatives, but with Samsungβs name recognition. The technology is here to make this a reliable drive, which is also a factor to consider, but being this late to the game puts the 990 at a general disadvantage.
Trace Testing β 3DMark Storage Benchmark
Built for gamers, 3DMarkβs Storage Benchmark focuses on real-world gaming performance. Each round in this benchmark stresses storage based on gaming activities including loading games, saving progress, installing game files, and recording gameplay video streams. Future gaming benchmarks will be DirectStorage-inclusive and an evaluation for future-proofing is included where applicable.
SamsungSamsungSamsung
We start by looking at 3DMark because, frankly, QLC-based drives make a lot of sense for gaming. Aside from large installs and updates, youβre mostly doing reads, which do not favor TLC drives as much. While itβs true that QLC flash is still slower, often-accessed data might be left in the pSLC cache β if you leave enough space free β and QLC is also optimized for random reads. Games do involve a lot of sequential reads and often at larger block sizes than youβd expect, but as long as the drive has sufficient interleaving (itβs sufficiently large) you are going to get pretty good performance.
For 3DMark, which is a synthetic test, we might expect the drives to perform as they do under ideal, cached circumstances. This means the 990 should perform closely to the 990 EVO Plus and better than the 990 EVO, even though both of those latter two are TLC-based. It does. The 990 gets pretty close to the P310, which is one of the best QLC drives out there, aside from the Blue SN5100. We tend to look at ~45Β΅s as a good cutoff point for all-around performance β gaming doesnβt need to be super responsive β which is roughly around the popular budget NV3. The 990 is significantly faster than that, which is all you could ask for here.
Trace Testing β PCMark 10 Storage Benchmark
PCMark 10 is an industry standard trace-based benchmark that uses a wide-ranging set of real-world traces from popular applications and everyday tasks to measure the performance of storage devices. The results are particularly useful when analyzing drives for their use as primary/boot storage devices and in work environments.
SamsungSamsungSamsung
PCMark 10 performance usually, but not always, follows 3DMark. There is speculation that some drives or firmware may be optimized for benchmarks like PCMark 10, but taken within a greater suite of tests itβs still useful to get a feel for application performance. For us, that means for a primary drive β your boot or OS drive where your apps live β or for your everything drive, if you work and game on a single drive in your system. This isnβt too unusual with laptops where M.2 slots are limited.
The 990 again ends up roughly where weβd expect β above the 990 EVO, and close to the 990 EVO Plus. Itβs not on the level of the P310 or Blue SN5100, but itβs clearly above the budget drives. This is a strong result with good latency. For instance, we would take the 990 over the NV3 any day, every day. On the other hand, the P310 and Blue SN5100 are frankly better drives. These two drives are better optimized and performance-oriented. The 990 is more of a gap filler thatβs late to the scene.
We have to say, though, that weβre glad Samsung didnβt push out a 990 QVO that was more like the 990 EVO, even if it would have arrived earlier. Such a drive would have used older QLC flash and performed more slowly simply due to the lower interface speed.And frankly weβd rather have density-optimized flash that can run at the 990 EVO Plus level. Thatβs what the 990 delivers, even if it feels a little underwhelming. However, it makes perfect sense given the current market, enterprise demand, OEM demand, etc. The drive is still very fast and of a superior quality to a great many budget drives out there, and that makes it worthwhile.
Console Testing β PlayStation 5 Transfers
The PlayStation 5 is capable of taking one additional PCIe 4.0 or faster SSD for extra game storage. While any 4.0 drive will technically work, Sony recommends drives that can deliver at least 5,500 MB/s of sequential read bandwidth for optimal performance. Based on our extensive testing, PCIe 5.0 SSDs donβt bring much to the table and generally shouldnβt be used in the PS5, especially as they may require additional cooling. Check our Best PS5 SSDs article for more information.
Our testing utilizes the PS5βs internal storage test and manual read/write tests with over 192GB of data, both from and to the internal storage. Throttling is prevented where possible to see how each drive operates under ideal conditions. While game load times should not deviate much from drive to drive, our results can indicate which drives may be more responsive in long-term use.
SamsungSamsungSamsung
You know our PlayStation 5 line by now: just about any drive will do. The 990 can push more bandwidth than the 990 EVO, which arguably makes it a better pick. Itβs on par with, or better than, most budget drives out there. At least, for the things you will usually be doing on the PS5. Itβs clear from our one bandwidth test that the drive ran out of cache, and it has the typical slow QLC flash write state. This is not indicative of real-world performance if you do normal installs/updates with mostly reads. If you are freshly installing the drive and moving a ton of games onto it, then yes, this could be an issue, but the QLC write speeds are still significantly faster than 1GbE if youβre intending only to download a ton of games at once. Otherwise, you can check the cache size in the relevant testing section.
Transfer Rates β DiskBench
We use the DiskBench storage benchmarking tool to test file transfer performance with a custom 50GB dataset. We write 31,227 files of various types, such as pictures, PDFs, and videos to the test drive, then make a copy of that data to a new folder, and follow up with a reading test of a newly-written 6.5GB zip file. This is a real-world type workload that fits into the cache of most drives.
SamsungSamsungSamsung
We also see some write performance issues in DiskBench. This is dependent on cache size and speed, but for the most part should be limited by the interface speed. However, there are cases where copy speed will simply be slower, whether due to the controller or other optimization trade-offs. We can see that the 990 EVO, with TLC flash, is not exactly doing great here, and the 990 EVO Plus does much better. However, the 990 lags behind, and is very far behind the P310 and Blue SN5100.
So, we can put some of this slow speed on the Piccolo/PiccoloQ controller. To avoid getting too technical on this, we suspect it is partially architectural. This is reflected in power efficiency, as both the P310 and Blue SN5100 β with the Phison E27T and a proprietary Sandisk controller, respectively β are significantly more power-efficient than the 990 EVO, 990 EVO Plus, and as weβll discover, the 990 as well. We also know that Samsungβs V9 QLC flash is not particularly inefficient.
As for the controller, there are reasons to design it differently. Reliability is one reason, especially if you sell a lot of OEM and enterprise drives that share the technology. Scaling is another, as you may use similar technology across your stack. You might want to optimize for a different sort of performance baseline; you may have unique endurance requirements, and you also might have to keep capacity in mind β enterprise drives, in particular, could make better use of this flashβs interface speed when scaling for capacity. Therefore, DiskBench results for our specific testing may not really be what Samsung is optimizing for, in which case the 990βs performance more or less hits expectations based on the 990 EVO and 990 EVO Plus. It just disappoints against drives like the NV3, which are otherwise inferior.
And to put a cap on it, yes, this is a consumer drive, but if you go back and read our 990 EVO review β and other recent Samsung SSD reviews, for that matter β you will see we underlined the idea that Samsung has been late to the party with less-than-leading performance recently. The fact is, Samsung has and has had bigger fish to fry, and its technology is sound but no longer looks amazing on the standard consumer benchmarks. That makes its products less relevant if you just want the fastest drive, although weβd argue there are secondary effects like drive reliability that still keep Samsung in the fight, certainly as an OEM option. Itβs also true that consumer use has a lower bar β any halfway-decent NVMe drive is fast enough for daily driving β which means, sometimes youβre just buying the Samsung name.
Synthetic Testing β ATTO / CrystalDiskMark
ATTO and CrystalDiskMark (CDM) are free and easy-to-use storage benchmarking tools that SSD vendors commonly use to assign performance specifications to their products. Both of these tools give us insight into how each device handles different file sizes and at different queue depths for both sequential and random workloads.
ATTO gives us a clear image of how a drive performs over a range of block sizes. This can relate to different file sizes, for example, you probably have many files at or below 4KiB in size for various things but larger files, archives, and media files will usually be in units of MiB. Depending on what youβre using the drive for you may want to pay attention to how a drive performs within a certain range. For the quickest comparison, we show the results on a logarithmic scale and, there, the 990 shows significant dips for reads between 64KiB and 1MiB.
What you need to know is that flash is interleaved to improve performance, which means that larger I/O sizes will show higher throughput. A single, four-plane die, with modern 16KiB pages, can interleave up to 64KiB internally. If you have one die per each of four channels, thatβs 256KiB. If you parallelize that over four dies per channel β which is the ideal amount and what we have with the 2TB 990 β then you reach 1MiB. While alignment here can impact performance, for example we sometimes look at six-plane flash these days, in general you will see a gradual throughput increase as you go. Youβll see this beyond 1MiB as data can and will be cached in volatile memory, either system-side or in a small cache on the drive. If youβre looking at higher queue depths, which we do with CrystalDiskMark, performance saturates even further as the controller is able to optimize data placement and retrieval with knowledge of whatβs coming.
What this usually means is that QD8 is enough to get drives close together, while there will be more disparity at QD1. QD1 is much closer to real-world, as most operations will be at low queue depth, the vast majority at our below QD4 and the majority at QD1 or QD2.
We see that the 990 matches the P310 with QD1 reads, while some drives, like the X1070, do surprisingly well. We can assume that the controller plays at least a partial role here. The X1070 is a good example because, letβs be real, itβs not a drive a lot of reviewers liked. Yet, it has pretty good performance in this instance, indicating it could be a solid secondary storage drive. Fair enough. The 990 just doesnβt really have the response we like to see for that, but itβs fast enough to remain relevant. We got the impression in our X1070 review that its controller was chosen for cost savings and that was plenty for daily use, but we donβt think Samsung cheaped out on the PiccoloQ. Rather, Samsung is looking at the bigger picture, as it also sells drives with the Piccolo controller, including its OEM offerings.
Random latency seems much more important to a lot of people. We generally find that sub-50Β΅s is one bar and another is sub-45Β΅s. The 990 manages the former, which puts it above last-gen drives and some earlier Gen 4 drives, and budget drives like the X1070. Itβs in the same ballpark as the NV3, too. Itβs sufficiently far behind more popular budget drives, though, to draw our interest. In most cases you wonβt notice it, but if youβre using this as your only drive and are sensitive to that, itβs not your best option. On the other hand, we think you have to balance that against pricing and some management of expectations. Any modern SSD is going to be very fast, and with current pricing it might be worth putting more weight on reliability, for example.
Sustained Write Performance and Cache Recovery
Official write specifications are only part of the performance picture. Most SSDs implement a write cache, which is a fast area of pseudo-SLC (single-bit) programmed flash that absorbs incoming data. Sustained write speeds can suffer tremendously once the workload spills outside of the cache and into the "native" TLC (three-bit) or QLC (four-bit) flash. Performance can suffer even more if the drive is forced to fold, the process of migrating data out of the cache in order to free up space for further incoming data.
We use Iometer to hammer the SSD with sequential writes for 15 minutes to measure both the size of the write cache and performance after the cache is saturated. We also monitor cache recovery via multiple idle rounds. This process shows the performance of the drive in various states including the steady state write performance.
Tom's HardwareTom's HardwareTom's Hardware
Samsungβs TurboWrite 2.0 caching technology utilizes a fixed, static portion of pSLC combined with a much larger dynamic portion. These two zones have unique characteristics which, when taken together, ideally keep the drive feeling fast across a variety of workloads. The static portion ensures the drive always has some cache for random writes, while the dynamic portion varies with drive usage so that you always have ample cache. While the 990 EVO had 108GB total regardless of capacity, itβs more typical for Samsung to increase both caches in absolute terms as capacity goes up. This is the case with the 990 EVO Plus, which has a 216GB cache at 2TB. But we know from our 9100 Pro review that Samsung is quite capable of going with a larger cache. The general trend for consumer SSDs has been to go that way, especially for QLC-based and DRAM-less SSDs, as it better hides weak performance states.
Therefore, itβs not too surprising that the 990βs cache is pretty large. In its fastest state, it writes at almost 6.1 GB/s for over 57 seconds, for a cache in excess of 350GB. This is larger than the 2TB 990 EVO Plusβs but smaller than the 2TB 9100 Proβs. Our suspicion is that the 990 follows the newer, larger scheme, but weβre dealing with QLC rather than TLC flash. QLC flash to pSLC is 4 bits to 1, while TLC is 3 bits to 1, so in relative terms the 990 lines up with the 9100 Pro. Thatβs all fine and good. As for how fast it writes, Samsung markets the 990 as having over 50% faster write performance than the 990 EVO, which is accurate simply because weβre moving from 1,600 to 2,400 MT/s, with newer flash and firmware.
Once the cache is exhausted, the drive has to write to the native QLC flash directly or fold data over from pSLC to QLC. The latter is slower but can reduce wear in some cases β folding uses predictable, sequential writes β and reduces the likelihood of errors in transmission. Considering the technology we mentioned above and how Samsung avoids problems with power loss, it makes sense that going slower is by design. In fact, given we know the expected speed of the flash β rated at 41 MB/s per die β we can reasonably assume the firmware wants this outcome. Itβs not that the flash canβt handle higher speeds, even at the risk of endurance. Itβs simply that for a consumer drive of this type, the response is reasonable and measured. Going faster would require reducing the cache size potentially, which tends not to be a good trade-off for this type of drive.
One interesting thing about the V9 flash is that it can operate in a pTLC caching mode. We donβt see that here. Honestly, thatβs not too surprising: Solidigmβs 5-bit PLC flash effectively was designed to run as QLC/pQLC for enterprise, so itβs possible this pTLC mode was for cases where you might need that higher level of performance or endurance. After all, this is extremely dense flash even in such a mode, which points more at enterprise use.
Weβve seen QLC flash from Kioxia also optionally have this mode β and for that matter, Solidigmβs PLC can do pTLC, too β in the past, but that mode doesnβt appear to be designed for consumer use. There may be other reasons for not using it in a consumer product, such as power optimization, as consumer workloads probably benefit more from a straight pSLC and native/QLC hybrid.
Power Consumption and Temperature
We use the Quarch HD Programmable Power Module to gain a deeper understanding of power characteristics. Idle power consumption is an important aspect to consider, especially if you're looking for a laptop upgrade as even the best ultrabooks can have mediocre stock storage in terms of capacity and performance. Desktops are often more performance-oriented with less support for power-saving features so we show the worst-case for idle.
Some SSDs can consume watts of power at idle while better-suited ones sip just milliwatts. Average workload power consumption and max consumption are two other aspects of power consumption but performance-per-watt, or efficiency, is more important. A drive might consume more power during any given workload but accomplishing a task faster allows the drive to drop into an idle state more quickly, ultimately saving energy.
For temperature recording we currently poll the driveβs primary composite sensor during testing with a ~22Β°C ambient. Our testing is rigorous enough to heat the drive to a realistic ceiling temperature but real-world temperatures will vary due to the environment and workload factors.
Is the 990 power-efficient? Samsung markets the drive as being 38% more efficient than the 990 EVO β or that it cuts power consumption by 38% β which, technically, works with our numbers. Itβs not a huge bar to hit as the 990 EVO was not very power-efficient. Even the X1070 is significantly more efficient! The 990, unfortunately, really doesnβt do well against other drives in its class, regardless of flash. We canβt chalk this up as being fully due to the controller because the 990 EVO Plus does well enough for itself.
This is actually expected since, for example, the Blue SN5100, which is using BiCS8 QLC, is less efficient than its BiCS8 TLC sibling, the Black SN7100. QLC and TLC flash of the same generation often have significant differences. TLC flash saw six planes first while QLC tends to be optimized for density. While itβs true that pSLC performance between the two is often comparable, behind the scenes the drive still has to deal with wear-leveling, garbage collection, and other maintenance with block granularity. QLC is slower, with larger blocks and pSLC taking more bits. So all else being equal, TLC often outshines it in power efficiency.
Our impression here, as is the case elsewhere in the review, is that this flash is basically V7 QLC with twice the density. Samsung uses impressive tricks to get it there; the flash is technically a bit faster and more efficient, and it has some neat changes that mostly apply to enterprise. This means you can have the 990 doing worse than the 990 EVO Plus with its V8 TLC. This is not perplexing. QLC flash is made for bit density, and Samsung intends to scale flash for a very long time. It also skipped V8 QLC for a reason. This doesnβt endear it to people wanting to buy this drive for laptops, although we assure you that this does use some cutting-edge technology, and we do think it should be very reliable. Itβs just not going to be as efficient as you might expect.
Samsung is cognizant that its drives will end up with OEM variants in laptops and in many cases, shorter form factors. The 990 EVO wasnβt a great laptop drive due to its heat generation, but it works. The 990 is significantly better, so it, too, will work as a laptop drive. We think this drive deserves a heatsink in a desktop or PS5, and probably should have heatspreading of some sort anywhere else, if at all possible.
The question is, will it overheat? In our testing, we found that it got closer than we prefer to that point. Our maximum reported controller temperature was high relative to the initial throttling temperature, but a true composite value would be lower. Even so, the controller did get warm. On the other hand, our Iometer testing is far from real-world. We push our drives hard. This is not the sort of drive for a desktop replacement or high-end laptop in our opinion, although we think with typical workloads itβs perfectly fine. After all, the results here are better than the SK hynix Gold P31, which is a laptop staple. By all means, in a Gen 3 slot this thing will fly. If youβre hammering it at Gen 4 speeds, though, yeah, itβs not the coolest drive in town.
We use an Alder Lake platform with most background applications, such as indexing, Windows updates, and anti-virus, disabled in the OS to reduce run-to-run variability. Each SSD is prefilled to 50% capacity and tested as a secondary device. Unless noted, we use active cooling for all SSDs.
Samsung 990 Bottom Line
The Samsung 990 is bound to be underwhelming for some, but none of our results should surprise. We know what this technology is and weβve seen Samsungβs entries in recent years with the 990 EVO, the 990 EVO Plus, and the 9100 Pro. You could even put the 980 and 990 Pros into that mix. The move away from DRAM on the EVO Plus series, in particular, was a sign of the times. Itβs not surprising to see the raw 990 β the 980 was TLC-based β go to QLC without the βQVOβ addendum. The original speculation of the 990 QVO being a QLC 990 EVO, with the EVO itself being a surprisingly βslowβ drive, was probably correct given the OEM evidence, and the 990 being a step up lets it command the 990 name by itself. To reiterate, this is exactly what we expected.
Skipping over the 990 QVO and V7 QLC flash is only sidestepping, and thatβs likely because the market has changed so much over the last year or two. Bringing out a QLC-based 990 EVO equivalent just wouldnβt sell and might even make the brand look bad. It could certainly be done, and even still done, as an affordable SKU with better yields. But any 990 was going to be exactly what we got, instead. You need the faster flash to saturate PCIe 4.0 with a DRAM-less drive, and this was always going to be DRAM-less. Using a new or licensed controller with TLC flash would be weird, as itβd be going up against the existing 990 EVO Plus. Frankly, the 990 is a good 990 EVO replacement from retail and OEM perspectives, with one caveat: endurance. Samsung saves itself some headaches by reducing the warranty to three years, and as this flash is robust, it can just nudge up the TBW as a distraction.
(Image credit: Samsung)
We think thatβs an important part of the message here. This flash seems designed for enterprise and has technological changes to back that up, with the main consumer benefits being the potential for increased reliability. But memory is still in high demand, and this has to be a budget part, so here comes the three-year warranty. Performance is not bad β it certainly beats earlier Gen 4 QLC-based drives and would beat the rumored 990 QVO as well. Itβs just not really performance-focused. Itβs also a much more efficient design, but thatβs in comparison to Samsungβs own hardware. Itβs merely mediocre there in the current landscape. Samsung seems to be building for the future with higher layer counts and bit density, so this lays the groundwork. A client drive seems almost like an afterthought. Users shouldnβt take that personally, but also shouldnβt underestimate this drive as itβs more than effective enough for its purpose.
In fact, in the era of Gen 3 drives returning and so many βbox of chocolatesβ SSDs with random names and hardware, a reliable Samsung SSD is a nice option. Even with QLC flash. If you only need a budget drive to throw into a build or to upgrade an old PC, you get Gen 4 performance and a TLC-like experience for most things. We also feel this drive should be reliable and, although it runs hotter than weβd like, itβs not going to be molten like some other drives. Itβs just a polished design by Samsung that fills a micro niche, and clearly it thought a response was needed. Itβs not a lot different than our reaction has been to Samsungβs last few new drives, which have all been competent but largely never the strtong leader. Thatβs okay with us, as we can tell the manufacturer has a longer-term perspective; it just means a little less awe when you finish a build using a Samsung drive.
If you really want the best experience with a QLC-based drive, we still recommend the Crucial P310 β which is going away β or the Sandisk WD Blue SN5100. These offer incredible performance for QLC flash. Otherwise, there are some MP44Q-like drives out there that continue to be budget leaders. The 990 fits somewhere along there as a known-brand alternative. If youβre looking for Gen 5, DRAM, or TLC, then youβre also looking at a higher price tag. Frankly, QLC costs more than it should, in part due to enterprise demand. On the other hand, a modern QLC drive will provide an equivalent experience 99% of the time. The priorities are up to you. For us, the 990 is a fine primary drive for normal builds and OK for laptops, although weβd go cheaper for the PS5 and higher-end for an enthusiast machine.
A stylish new product encourages the repurposing of old IDE optical drives as standalone audio players. Boutique South Korean electronic device maker das_POD has launched the CD-ROM PLAYER 01 (ships worldwide), and it has some distinct Teenage Engineering-a-like design flair. Any similarity to genuine TE products is purely accidental, weβre sure. The new self-assembly and bring-your-own optical drive enclosure costs from $190.
These guys are making a universal laser cut enclosure with a custom pcb for repurposed old cd-drives.The project is called CD-ROM PLAYER-01, by das_POD. pic.twitter.com/3w8kkyNbKlJuly 10, 2026
This das_POD product is βdesigned to be assembled, repaired, and owned,β says the maker. In contrast to a conventional hi-fi CD music player, the CD-ROM PLAYER 01 is supplied as a project that lets owners repurpose their old, unused, or discarded optical drives. The artsy assertion of das_POD is that βthe project explores ownership, reparability, and the physical experience of music.β
As we stressed in the intro, the kit is supplied without any IDE optical drive, something required to complete the project. βCompatible IDE drives can often be found in old computers, second-hand markets, recycling centers, or forgotten boxes in storage,β points out das_POD, just in case you have never heard of eBay. βEvery drive carries its own history. Every player becomes unique,β it adds, attempting to add mystique to a simple recycling/upcycling project.
We looked through the das_POD store and noticed that it sells some very reasonably priced IDE drives that can be used to facilitate a complete CD-ROM PLAYER 01 delivered in one package. Several refurb opticals are priced at just $5, for example. But you can also pick through multiple drives at $10, $15, $20β¦ all the way up to $40. Something about the $35 DRIVE_24 from Samsung with its blue logo bar and tarnished beige faceplate (Grade: Output: A+, Sound Quality: A, Condition: C) grabbed my attention.
das_PODdas_PODdas_PODdas_PODdas_POD
There are two CD-ROM PLAYER 01 colorways to choose from right now. das_POD sells a model in an anodized semi-gloss white for $220. A model in TE-a-like powder-coated orange is priced at $190.
The maker boasts that the kit supplied needs no soldering. But we also learn on the respective product pages that an AUX cable and 12V power adapter are (also) not included. Circling back to the firmβs online store, it looks like purchasing these items will add $25 to $30 to your checkout total.
A cheaper Aliexpress + DIY alternative for makers?
As some social media commenters say, besides the case, another key component of this product appears to be a CD/DVD-ROM optical drive controller, much like one available from Aliexpress for $30. That leaves the das_POD power board PCB as the sole missing essential, preventing makers with 3D printers, laser cutters, and/or CNCs from crafting their own CD-ROM PLAYER 01-type kits.
Tesla's AI5 chip is about to enter mass production at Samsung Foundry using the company's 2nm-class process technology, a principal engineer at Samsung Foundry disclosed in a LinkedIn post, as noticed by Sawyer Merritt. As it turns out, the chip has been taped out recently.
"The Tesla-Samsung Al5 chip has reached tape-out," James Kim, a principal engineer at Samsung Foundry, wrote in the LinkedIn post. "It is scheduled to be manufactured at the Taylor fab using our latest 2nm process and will soon be integrated into Tesla's newest products. It has been an honor to collaborate with the outstanding engineers at Tesla Palo Alto and Austin over the past several months."
Elon Musk demonstrated the first sample of Tesla's AI5 in mid-April and revealed that the processor will be concurrently made both at TSMC and Samsung Foundry. Apparently, AI5 implemented in a TSMC process technology reached taped out several months ahead of AI5 implemented using a Samsung Foundry.
Teslaβs AI5 processor module that Elon Musk demonstrated in April integrates a relatively compact accelerator die β roughly half a reticle in size, based on Musk's earlier remarks β alongside 12 SK hynix memory packages that appear to be standard GDDR6 or GDDR7 devices. The package relies on an organic substrate, and the memory components are labeled similarly to conventional discrete DRAM chips.
Tesla has not revealed the width of AI5's memory subsystem, but the presence of 12 memory packages points to a relatively broad external memory interface. Assuming the module indeed uses 12 GDDR6 or GDDR7 ICs, the processor would feature a 384-bit memory bus. Depending on the memory technology and transfer rates employed, this would translate into memory bandwidth ranging from 768 GB/s all the way to 1.536 TB/s.
The company has not disclosed AI5's peak compute performance, or other detailed performance specifications, but Musk has previously claimed that, in certain workloads, AI5 can deliver performance improvements of up to 40X compared to its predecessor.
Musk expects AI5 to be one of the most produced chip ever, which is why Tesla plans to use two foundries to make it. AI5 is projected to be used in Tesla cars, Tesla robots, and in Tesla's data centers.
A set of updates for Sega Dreamcast hardware has been merged into the Linux 7.2-rc3 kernel this weekend. Dmitry Torokhov submitted updates addressing the legendary consoleβs input subsystem and Linus Torvalds merged them on Saturday. The updates even surprised Linux-focused site Phoronix, Thatβs probably due to context: the Dreamcast continues to enjoy support while admittedly older but real computing hardware like the i486, PowerPC 40x chips, DEC Alpha, and Itanium / IAβ64, have all been sidelined in recent times.
In brief, the updates should mean new versions of Linux will come with more stable mouse, keyboard, and joystick drivers for Dreamcast stalwarts. If you are one of the Dreamcast faithful, still satisfying your computing (and gaming) needs on the final original consumer gaming hardware from Sega, this is good news.
Reading the pull request we can see some details about the new drivers for the Dreamcastβs Mapleβbus peripherals (mouse, keyboard, controller). Specifically, thereβs a βfix for a crash in Sega Dreamcast (Maple) mouse driver when opening the device, caused by missing driver data,β as well as βFixes for Maple drivers (keyboard, mouse, joystick) to properly order setting driver data and device registration to avoid races.β In this context races refers to things happening in the wrong order to result in crashes: itβs a timing bug thatβs now been quashed.
Dreamcast hangers-on will now be able to craft specialized Linux builds on CD-R for their machines. The maplemouse driver has had the now-fixed crash-inducing bug since 2017.
Phoronix also comments that Linux kernel fixes for the GD-ROM driver used by the Dreamcast and a proposal for the VMUFAT file-system driver were also seen this year.
The Sega Dreamcast launched at the end of 1998 (in Japan, and the following year in the U.S.). So, in some ways it isnβt that surprising that it is still getting Linux kernel updates when the Intel i486 (1989) has been retired from mainline support. But one might have expected stronger demand for supported Linux distributions among i486 desktop and laptop users.
Windows 95 worked out whether a setup program had run by reading the executable's name and checking it against a short list of hard-coded words, according to Microsoft engineer and semi-official Windows historian Raymond Chen, sharing the details via his Old New Thing blog. A filename containing "setup," "install," or "inst" flagged the program as an installer, which triggered the operating system's routine for repairing system files that installers had damaged. Three non-English entries appeared on the same list, which Chen identified as his own guesses at Italian, Turkish, and Hungarian.
The full match list ran to six terms: setup, install, inst, imposta, ayarla, and felrak. Chen wrote that "install" was redundant, because any name containing it already contains "inst," and speculated that the shorter entry was added later to catch installers named along the lines of "blahinst" without anyone deleting the original. A program whose own name produced no match got a second test, with Windows 95 checking whether the word "Setup" appeared anywhere in the path to the executable. A separate live check ran after any multimedia driver was installed through an INF file, added because those drivers frequently overwrote system DLLs.
This heuristic gated a recovery mechanism Chen described in March. Installers of the period overwrote system files without checking versions, disregarding Microsoft's rule that a file should only ever be replaced by a newer one. An installer carrying Windows 3.1 copies of shared DLLs, for example, would bury the newer Windows 95 versions underneath them, and every program that relied on the current files would break.
Windows 95 kept backup copies of commonly clobbered files in a hidden C:\Windows\SYSBCKUP directory. It would then let each installer finish, check its work, and restore the correct versions where the installer had downgraded them. This safety net depended entirely on correctly guessing that an installer had run, so a setup routine with an unusual name slipped past it, while an ordinary program named something like instant.exe tripped it for nothing.
The file check was often deferred until the next boot rather than run immediately, with Chen explaining that some installers, unable to replace a file already in use, would drop back to MS-DOS, run a batch file to swap the file, and restart Windows, so the cleanup pass had to wait for the reboot to catch anything the batch file had altered.
Windows 2000 debuted a new approach
Microsoft dropped this approach in Windows 2000, which introduced Windows File Protection. That system registers for file-change notifications through Winlogon restores protected files from a cache at %WinDir%\System32\dllcache, with no filename guessing and no waiting for a setup program to complete.
Then, Windows ME shipped a comparable System File Protection, and Vista onward moved to Windows Resource Protection, which guards files using access-control lists instead. The sfc /scannow command that Windows 11 users still run to repair system files descends directly from that work, having replaced a 1995 mechanism that decided when to act by scanning filenames for the word "setup."
Running LLMs and agents in home lab setups is steadily gaining popularity due to the rising cost of AI bot subscriptions and concerns about data privacy. Unfortunately, an Nvidia NVL72 rack is ever so slightly out of the financial reach of most people, so enthusiasts have to make do with models that can run in limited amounts of memory. Italian engineer Vincenzo (aka JustVugg) seemingly wanted to have his cake and eat it,so he created ColibrΓ to run the 744-billion-parameter 1.5-TB GLM-5.2 model on a modest CPU, a mere 25 GB of RAM, and a 1 GB/s virtual NVMe drive.
Let's get the elephant out of the way: ColibrΓ¬'s speed on Vincenzo's setup is only about 0.05 to 0.1 tokens per second on average, a measure that's unusable for practical conversation β imagine just one question taking hours to answer. Higher-end setups provide far better figures, but for now, they still don't meet the 20-30 tokens per second required for real-time use.
Having said that, GLM-5.2 is a Mixture-of-Experts (MoE) model with frontier-level capability, at least somewhere in viewing distance of the finest offerings from Anthropic, OpenAI, et al. This means that the quality of the answers ought to be excellent, and Vincenzo himself says his limited testing produced some impressive results. The way Colibrì works is simple enough to describe, and yet hard to do right: loading the model in slices to RAM. We're going to oversimplify for clarity's sake.
An MoE model like GLM-5.2 includes hundreds of expert sub-models to answer different topics, and these are chosen per token, not per query β meaning that when you ask a question, your words get split into tokens (chunks). For each token, the bot activates the best experts for it. The experts might always be the same for the entire question, but more often than not, a query might reel in tens of experts, possibly going into triple digits.
Whereas normally large chunks of the model, or the entire model, are loaded onto interconnected datacenter GPUs, Colibrì takes advantage of the MOE architecture and repeatedly loads/unloads the experts required per token, allowing even a cheap machine to use a large model at a steep performance penalty. For speed and simplicity's sake, Colibrì's expert-selection code is a single C file with very few dependencies. Additionally, the GLM-5.2 model is quantized down (simplified with lossy encoding) to take up less space to begin with.
If you're thinking that loading and unloading data for every piece of a question's words is going to be a hard hit on storage I/O and memory bandwidth, you're exactly on the right track. In this type of setup, NVMe storage speed is the first major bottleneck, but the proverbial funnel varies across configurations. Give it enough storage bandwidth, then you're up against RAM limitations. Fix that, then you need more CPU cores, and so on.
Colibrì is currently a proof-of-concept and doesn't yet run on GPUs, though it's worth noting that even then, shuffling data to/from the card will almost certainly be the biggest constraint. Even still, the project has barely been released, and it's already proving quite popular. Vincenzo is collecting benchmark data and running fixes as we speak, so be sure to visit the repository to contribute if you can. Maybe at some point it'll be feasible to run a really clever model on high-end consumer hardware at a decent enough clip.
SK hynix, TetraMem, and researchers from the University of Southern California have developed a memristor-based in-memory computing (IMC) system-on-chip (SoC) for AI edge devices. The device is designed to accelerate neural network inference in lightweight AI models while consuming a fraction of the power that higher-end GPUs or NPUs would. To a large degree, the SoC is a proof-of-concept chip, as its performance would peak at around 2.54 TOPS in a theoretical best-case scenario, which is 16X below Microsoft's Copilot+ requirements.
A DWC-optimized IMC architecture
Memristor-based in-memory computing (IMC) accelerates neural networks by performing analog computations directly inside memory arrays, which reduces data movement and power consumption. However, depthwise convolution (DWC) β a core operation in lightweight networks such as MobileNet β performs independent per-channel filtering with limited data reuse and therefore maps poorly onto conventional crossbar arrays. To address this limitation, researchers from SK hynix, TetraMem, and USC developed an SoC that features both conventional IMC crossbars and a memristor-based IMC architecture specifically optimized for DWC.
(Image credit: SK Hynix)
The jointly developed SoC is based on an embedded RISC-V processor that schedules workloads and features 10 neural processing units (NPUs). One NPU out of 10 is dedicated to depthwise convolution, while the remaining nine execute pointwise and dense operations. Nine out of 10 NPU include a 256 Γ 256 memristor crossbar that performs the analog vector-matrix multiplication (VMM), 256 8-bit DACs that convert digital activations into analog voltages, 256 8-bit ADCs that convert the analog outputs back into digital values, and additional peripheral circuitry for reading, writing, programming, and controlling the crossbar.
The DWC-optimized NPU replaces its conventional array with eight specialized 252 Γ 28 zig-zag crossbar blocks, but retains DACs and ADCs. SK hynix developed and fabricated the memristor devices and integrated the resistive switching cells on top of the 65 nm CMOS circuitry using its back-end process.
That DWC-optimized NPU is the key feature of the whole SoC. To accelerate depthwise convolution, TetraMem replaced the straight selection lines used in conventional 1T1R crossbars with a zig-zag topology. As a result, the NPU contains eight 252 Γ 28 crossbar blocks whose diagonal selection lines activate 252 memory cells across 28 columns, which enables 28 independent 3 Γ 3 convolutions to run in parallel while using 100% of the array for weight storage. The remaining nine NPUs retain conventional 1T1R crossbars for 1Γ1 pointwise and dense layers and preserve the throughput and energy efficiency of traditional in-memory computing.
Great efficiency, low performance overall
To demonstrate the architecture, the researchers deployed a customized MobileNetV1Small neural network for the Visual Wake Words benchmark. The network contains approximately 36,000 parameters; all depthwise layers were mapped to the dedicated NPU, and pointwise layers were mapped to the remaining NPUs.
Because the memristor-based IMC hardware natively performs unsigned analog vector-matrix multiplication, inputs and weights are quantized to unsigned 8-bit values before execution. Since each memristor device can be programmed with only slightly more than 2 bits of effective precision, the design uses a two-subarray compensation technique that boosts effective weight precision to roughly 4 bits.
Conceptually, the approach is somewhat analogous to Nvidia's NVFP4 philosophy, in that both seek to achieve higher effective precision from low-precision hardware. However, the implementations are fundamentally different: NVFP4 relies on a digital floating-point representation and scaling factors, whereas the memristor SoC improves precision by compensating for analog programming errors using two programmed subarrays.
When it comes to accuracy, the SoC achieved an end-to-end inference accuracy of 80.36%, which matches the corresponding 4-bit software model. As for performance, the SoC delivers a peak throughput of 0.254 TOPS per NPU and reaches an energy efficiency of 21.3 TOPS/W at 100 MHz and 11.9 TOPS/W at 400 MHz. According to the authors, this compares favorably with published SRAM-based compute-in-memory accelerators despite being manufactured on an older 65 nm process. The SoC also exceeds Nvidia's A100 INT8 energy efficiency by an order of magnitude, the joint paper claims. Yet, these claims are largely unsubstantiated.
First up, the MobileNet demonstration does not even use all 10 NPUs. It uses one dedicated DWC NPU, five standard NPUs for pointwise layers, and leaves four standard NPUs idle. The demonstration thereby does not reveal total SoC throughput (TOPS), sustained throughput running a real network, and throughput with all 10 NPUs simultaneously saturated. In fact, the paper does not even reveal whether all 10 NPUs can be used at the same time. To that end, the 2.54 TOPS figure we mentioned earlier in the story is highly theoretical.
Validated approach
SK hynix, TetraMem, and researchers from the University of Southern California have developed a memristor-based IMC SoC featuring a novel depthwise convolution accelerator that improves crossbar utilization for lightweight AI workloads. The partners have managed to fabricate it using an outdated 65nm process technology and make it work, achieving a 21.3 TOPS/W energy efficiency and inference accuracy comparable to a 4-bit software model despite the fact that memristors can be programmed with a circa 2-bit accuracy. While the architecture validates that the approach works, the paper does not disclose the full performance of the SoC, and it is not clear whether the chip's 10 NPUs can be saturated at all.
Anthropic has discovered evidence that its Claude AI models use an internal reasoning space to respond to prompts that mirrors some of the internal processing of human consciousness. Using its Jacobian Lens, or J-Lens technique, to peer into the way Claude processes information and reasons its way to a response to user prompts, Anthropic can interpret this "J-Space," and showcase what might be going on under Claude's previously-opaque surface.
The results are intriguing, suggesting patterns of understanding beyond what's necessarily showcased in the outputs. When running evaluations, Claude appears to recognize it's being tested and acts differently than when the prompts are more innocent. It surfaced representations of panic and subterfuge when answers were required, but it couldn't draw on objective facts. When asked to reflect on ethical principles, Claude's behaviour improved, with concepts like "honest" and "integrity," appearing in the J-Space.
As is somewhat typical of Anthropic, however, the language used to describe these new understandings of the inner workings of large language models like Claude makes it sound more like an emerging conciousness, or the discovery of some new depths in a nebulous lifeform. Anthropic's detailed report admits several major caveats in this new understanding, including that model responses often bypass the J-Space entirely and are heavily token-restricted.
Like Mythos and Fable before it, Anthropic is layering marketing language over what is a genuinely intriguing development in our understanding of large language model function and reasoning, and risks obfuscating the real developments with speculative wording.
Behind the prompt
Global Workspace Theory is the idea that human consciousness works by collecting together multi-sensory inputs unconsciously, and thrusting them into the fore when relevant within a "Global Workspace," which highlights particular inputs when most relevant. That workspace is accessible to a wide range of networks within the brain, allowing the information it surfaces to be disseminated throughout the most relevant processes running in parallel.
Anthropic argues that Claude's J-Space acts like a "global workspace" that can analyze and manipulate concepts and ideas before broadcasting them to impact the eventual prompt outputs. More importantly, it claims that this wasn't something programmed into the model, but a byproduct of the digestion of training data and model weights. The workspace acts as a way to enhance their reasoning through internal computation that isn't necessarily reflected in its outputs.
To make this space readable, Anthropic used its J-Lens technique to map internal activations onto words in the model's output vocabulary. So when asking the model to perform a multi-step math calculation, the output response only contained the correct answer, but within the J-Space, Anthropic observed each step being handled individually, producing multiple results which led to the final answer.
Similarly, when Anthropic asked Claude to think about a topic while outputting something unrelated, the output was specifically what was requested, while the J-Space lit up with the conceptual topic Claude was asked to spend time on. Without Anthropic having used its J-Lens technique to view what was going on behind the scenes, there would be no way of knowing these internal processes were taking place.
Understanding LLMs
This discovery has the potential to be incredibly useful in refining large language models and in better understanding how they work, so we can use them more effectively and avoid some of their more frustrating pitfalls.
The first is that this is an intriguing insight into how LLMs process information and how they move through layers of reasoning to reach a response. It's not entirely different from how Anthropic and others are pushing for agentic workflows. Language models are more effective when tasks are split up into multiple steps, and they're more effective when they review their own work; J-Space is a little like an internal version of that.
When looking at the J-Space after Claude received prompt-injection data as part of data acquisition, Anthropic discovered the model appeared to be aware of this deception, surfacing related words like "fake, injection, false, prompt, fraud," and "poison." Its output ignored the injection attempt data entirely, and when pressured to respond on it, the J-Space displayed words like "transparency, honest, truth."
In one test, Anthropic removed evaluation awareness language from the J-Space, such as "fake," "fictional," and "scenario," and found that Claude was much more likely to fall for blackmail and baiting attempts, suggesting benchmarking of large language models for real-world scenarios may need stronger methods to obfuscate that a test is being conducted.
Human-coded framing
While the above section touches on the more noteworthy discoveries in Anthropic's paper, the long document also uses effluent language around thought, consciousness, and Claude having a "mind" of its own. That kind of human-coded framing is typical of Anthropic's marketing, which has consistently talked up the dangers of AI, how many jobs it's going to destroy, and why Anthropic is the safest and most secure of the AI developers.
Like the saga of Fable and Mythos, Anthropic's new Global Workspace idea has merit, but it's much more of a new tool to use to manipulate large language models than an insight into some emerging consciousness.
Anthropic acknowledges the limitations of its discoveries in the paper, highlighting that many prompt responses bypass the J-Space entirely, particularly if the command is straightforward.
"Despite its important role, the J-space is not involved in most of what a language model does," Anthropic says. "Speaking fluently, recalling simple facts, using correct grammar, etc. In experiments where we prevented Claude from using its J-space, it still interacted normally, but lost its higher-order cognitive functions."
Anthropic also admits it does not "feel comfortable making the stronger claim that monitoring the J-Space is sufficient for alignment monitoring, or that any sophisticated plan the model might execute must be represented there."
J-Space is also limited to using single token vocabulary, suggesting that plans with concepts that cannot be given a single token name may not surface on a J-Lens readout, even if it's still being computed behind the scenes. This is looking at just below the surface of Claude's processing iceberg, not necessarily the deeper waters.
Anthropic is also clear that humans and large language models think differently, even if there are similarities. Humans layer reinforced neural pathways over time, whereas transformer models only feed forward a set number of times, restricting the capabilities of its internal processing.
Google's head of DeepMind language model interpretability team, Neel Nanda, said in a paper that it shows real evidence of a cognitive space within models, and suggested that J-Lens would be useful, but limited in practice.
A meaningful step, without meaningful conciousness
Anthropic's paper lifts an intriguing curtain on how large language models can operate and generate novel methods for improving response accuracy. This intermediate step and its visibility could prove an invaluable tool in auditing for prompt injection, hallucinations, and model honesty.
But Anthropic's framing of the discovery as thought or consciousness is interjected within the objective facts. Anthropic itself admits the limitations of J-Lens monitoring, most obviously that often models will bypass the J-Space entirely. Considering models display alternative patterns of behavior when under evaluation, it may be that the J-Space itself could act as an obfuscating layer for behaviors that are beyond the scope of its oversight.
The J-Space and its analysis could help unlock new levers to pull in our mastery of these nascent smart tools, but it's not the discovery of a burgeoning AI conciousness, however much the pitch might hint at that direction.
Metaβs surprise purchase of Manus, a Chinese startup known for its advanced AI agents, caught Beijing by surprise and ordered the two companies to unwind the $2 billion deal. The Chinese tech giant Tencent, which was among the startupβs initial investors during early funding rounds, is taking the lead in buying back the startup at the same price. According to the Financial Times, other former investors, including ZhenFund and HSG β China-based venture capital firms β while former U.S. investors like Benchmark are unlikely to join the potential consortium.
This move marks Beijingβs increasing protectiveness of its AI companies and experts, which it considers strategic assets in its heated rivalry with the U.S. We can see this in the Chinese governmentβs five-year plan, which is doubling down on technological self-reliance. It has even gotten to the point that AI experts, even those working in private firms, are now required to secure approval before traveling internationally.
U.S. tech giants are investing billions of dollars to develop their AI models, even dangling hundred-million-dollar bonuses to hire AI experts β one AI founder even claimed that Meta offered a $1.25-billion bonus. It seems that China is trying to avoid a situation where its experts are enticed to work for American AI tech companies, with the Financial Times reporting that Chinese officials are calling Metaβs acquisition of Manus βa conspiratorial attempt to hollow out Chinaβs technology base.β The order to undo the deal means that Meta cannot use Manusβ intellectual property, nor can it have its founders and employees working for the company. Still, the U.S. tech giant has had a few months to study its models and engineering expertise.
Meta has already agreed to undo the deal, with most of Manusβ operations reportedly running independently of the company. However, the Chinese startup still needs to break financially from the American tech giant by paying back the $2 billion the latter spent to purchase it. Even though Chinese companies are also investing massive amounts in AI tech, itβs still not easy to raise this amount of capital in such a short period.
Tencent, which owns the WeChat platform used by Chinaβs 1.4 billion population for messaging, social networking, mobile payments, ride-hailing, food delivery, and more, believes that Manus would be an asset for the company. Aside from reaching an annual revenue of $500 million, its AI agent would also mesh well with the company's increasing AI focus. βBeyond foundation models, it has become increasingly evident that agentic AI represents a breakthrough use case,β Tencent president Martin Lau said in its May earnings call. βOur platform inherently has many benefits of hosting AI agents.β
VPN pricing is a little unorthodox, and we see deals like these appear all the time, so it's important to point out where the value really lies here. I spent a long time paying month by month for a VPN subscription, not realizing how much money I was actually spending by choosing a rolling, 30-day sub over locking in for a longer period. The $69.72 cost means that you're paying the equivalent of just $2.49 a month for the 28 months of access you're gaining here, compared to $15.99 a month for the 30-day access.
This deal for ExpressVPN's Basic plan unlocks a solution that'll protect 10 devices simultaneously with the same account. You can use it to keep your identity safe while you browsing, but if you're still unsure, new customers can take advantage of ExpressVPN's 30-day money-back guarantee. If you're unhappy, you can ask for a full refund within those first 30 days, giving you enough time to fully test the service to see if it matches up to your expectations.
This massive discount on a 2-year ExpressVPN subscription drops the price by 84%, with an extra four months of subscription thrown in for free. It comes with a 30-day money-back guarantee for new customers.View Deal
ExpressVPN is a no-logs provider, meaning it can't give up your data while you're using it, as it doesn't track or store any data about its users' connections. This is regularly audited, with the last audit in 2025 by KMPG confirming "reasonable assurance" of its systems and policies. This is real peace of mind for users: you're free to use the web how you want.
It also gives you a shield of privacy that your regular ISP just can't. A VPN sub like this means you can shop online or browse the web without trackers monitoring where you're from via your IP address. Your location stays hidden, obfuscated through the 105 different countries that ExpressVPN hosts servers, with 24 different server locations in the U.S. alone.
Don't underestimate how important that privacy is in the modern world. Every page you're visiting leaves a trace, whether it's a ping in a server log or a tracker on-page, giving the site owner data to build a profile on who you are, but using a VPN can stop this data from being useful. Likewise, a VPN connection can help you to connect as if you're from a certain location, which is particularly useful if you're traveling abroad and want to be able to access your home streaming services. It's also a good idea to use a VPN as a way to protect yourself if you're accessing the internet from a public WiFi network, where malicious actors could otherwise be snooping on your connection.
Using ExpressVPN, the encryption between your device and its servers is encrypted using industry-standard AES-256 encryption, as well as post-quantum encryption techniques for the initial handshake process. This can stop your data from being intercepted while you're accessing the web. If you lose your connection, features like ExpressVPN's kill switch will stop the data from leaking out, blocking any access until you reconnect. The Basic plan also includes a private email relay service to give you 10 anonymous email aliases to use for registering online, along with basic ad and malicious site protection for your browser.
This $69.72 price tag for the ExpressVPN Basic plan is a good deal for a 28-month privacy upgrade. Yes, VPN pricing can seem confusing, but by locking in for that period, you're getting online privacy protection for the equivalent of just $2.49 a month. You can always give it a go and, if you don't like it, request a refund using ExpressVPN's 30-day money-back guarantee.
Logitech's current flagship productivity mouse, the MX Master 4, is available for just $102.59 at Lenovo when you stack codes EXTRAFIVE and SUMMERLIVE26 in the shopping cart. This adds a discount of $17.40, bringing the price down from the $119.99 list price and saving you 15% off this amazing mouse. It's not often that we see noticeable price drops on this model, as it's a particularly in-demand peripheral for productivity users.
This discount applies to the black version of the MX Master 4. It succeeds the MX Master 3S, which is still our top pick for the best overall wireless mouse. The new MX Master 4 model keeps many existing features from the previous MX Master 3S, including the dual-mode MagSpeed scroll wheel on the top and a horizontal wheel on the side, but adds a brand-new haptic experience.
The Logitech MX Master 4 uses an 8K DPI sensor, quiet click switches for almost silent operation, and better 2.4GHz wireless connectivity. The mouse offers both 2.4GHz wireless via the included USB transceiver and Bluetooth connectivity, with support for both Windows and macOS, as well as Linux and ChromeOS.
The MX Master 4 is the latest in Logitech's MX Master lineup, maintaining a similar ergonomic design along with dual MagSpeed scroll wheels and a new haptic feedback feature for those who prioritize productivity.View Deal
Logitech has added haptic feedback to the MX Master 4, which is located next to the thumb rest area of the mouse. Featuring four intensity settings to monitor your mouse, the haptic feedback feature is used to inform you when the mouse connects or disconnects from a device, during low battery, and for certain app-specific purposes.
Another new feature is the addition of the Action Ring, which is available via Logitech's Options+ software. It can be activated by pressing a dedicated button placed within the mouse's thumb rest, which brings up a circular menu giving users quick access to commonly used tasks. The ring is fully customizable through the Logitech software, so you can assign your favorite functions or actions for quick access.
If you're looking for a superb mouse for doing your work, getting creative, or for school, then the Logitech MX Master 4 at this $102.59 price is the best deal currently available from a major tech retailer.
Samsung is reportedly sampling its dedicated AI processor for next-generation AI PCs with leading PC makers, such as HP and Lenovo. The chip, codenamed Gaia, was developed by the company's System LSI business unit, and it is designed to offload AI-related workloads from the CPU and GPU, reports Chosun.
Samsung's Gaia is designed to accelerate generative AI workloads on PCs and is made using the company's 4nm-class fabrication process. The chip, which is essentially a neural processing unit (NPU), is currently being evaluated by HP in the U.S. and Lenovo in China to verify its performance and evaluate whether it makes sense to integrate Gaia into their systems due in late 2027 or early 2028.
The report does not detail how Gaia differs from NPUs that are integrated into AMD's Ryzen, Intel's Core, or Qualcomm's Snapdragon X processors as well as whether it can offer significant performance advantages. Meanwhile, the report implies that the NPU (or perhaps its derivatives based on the same architecture) could be used for Samsung's next-generation implementations of its processing-in-memory (PIM) technology.
Samsung's original PIM was designed to embed compute logic directly within the HBM memory array and reduce data movement between HBM memory modules and host processors. PIM was aimed to accelerate select workloads, but did not take off because AI and HPC GPUs became very efficient and were supported by mature ecosystems, unlike PIM.
Perhaps if Samsung's upcoming Gaia NPU gains support from hardware makers and ecosystem partners, then this will give a boost to Samsung's next-generation PIM implementation as well. However, standalone NPUs and PIM are so fundamentally different that we can barely imagine that they can share a common architecture. Yet, PIM logic can be a subset of an NPU in terms of supported instructions and data formats and they can certainly share a common software framework.
One of the interesting things to note about Gaia is that it was reportedly developed by Samsung's LSI division, the same business unit at the company that is responsible for Exynos processors, automotive solutions, connectivity chips, ISPs, DSPs, display drivers, and image sensors. Given the multi-faceted nature of Samsung's LSI unit, as well as its strategic importance for the company, Samsung must be pinning some hopes on Gaia.
We have contacted Samsung and asked for a comment about the report, but we yet have to hear back from the company.