Normal view

There are new articles available, click to refresh the page.
Before yesterdayMain stream

Dell Pro 7 14 AMD Review: Ryzen AI 9 HX PRO 470 in a 2.8-Pound Business Laptop

22 July 2026 at 21:13

Dell sent us four Pro laptops this cycle, covering two sizes, two product tiers, and both AMD and Intel platforms. The two Pro 5 models approach the business-laptop formula from different directions: the Pro 5 16 AMD uses a larger chassis, replaceable DDR5 memory, and the Ryzen AI 9 HX PRO 470, while the compact Pro 5 14 Intel combines a Core Ultra X7 368H with fast LPCAMM2 memory and Arc B390 graphics. Dell then moves up to the thinner Pro 7 14 chassis with a choice of Intel or AMD processors. The AMD model reviewed here features the same 12-core, 24-thread Ryzen chip and 55 TOPS NPU as the Pro 5 16, but pairs it with 64GB of LPDDR5x-8533 memory, Radeon 890M graphics, a 1TB SSD, and a 14-inch WUXGA display.
Dell Pro 7 14 AMD review

Putting the same processor into two very different laptops gives us a useful look at how chassis size and memory choice affect performance. The Pro 7 14’s faster memory gives the Radeon 890M and local AI workloads an advantage over the Pro 5 16’s DDR5-5600, allowing the smaller laptop to lead the Blender tests, the AI Computer Vision GPU test, and all four AI text-generation workloads. Longer all-core workloads favor the larger Pro 5 16, which finishes Cinebench R23 multi-core 24% ahead with the same processor. Battery life reached 19 hours and 28 minutes, several hours beyond the Pro 5 16 but well short of the two Intel 14-inch systems. Our unit also had the slowest SSD of the four, with its Samsung BM9C1a trailing the Gen5 drives in the Intel laptops by a wide margin.

Beyond those performance differences, the Pro 7 14 AMD is built primarily for businesses deploying and managing a fleet of notebooks. AMD PRO provides DASH and AIM-T management features, while Dell adds SafeBIOS, SafeID, quantum-resistant BIOS verification, and 36 months of ProSupport Next Business Day Onsite service. Our configuration carries a $5,377 single-unit price on Dell.com, although volume buyers will typically pay less. Its $729 premium over the Pro 5 16 buys a thinner and lighter design, much faster memory, stronger integrated graphics, and better AI text-generation performance than the larger AMD model.

Design and Build

The Dell Pro 7 14 uses a compact, dark gray chassis that looks appropriate for a business laptop without feeling plain. At 315.5 × 226 mm, it takes up little space in a bag. The front measures 10.3 mm thick and the maximum thickness reaches 16.45 mm. Dell lists a starting weight of 2.80 pounds, though the final weight varies by configuration. The lid and palmrest have a smooth finish, and the magnesium bottom door keeps the system light while giving the chassis a solid feel.

Dell Pro 7 14 AMD review lidRemoving the bottom panel exposes a single cooling fan, the heat-pipe assembly, wireless card, M.2 SSD, and the large 70Wh battery. The LPDDR5x memory is soldered to the motherboard and cannot be upgraded, but the SSD, wireless card, and battery are accessible for service or replacement. Dell also uses modular components in several areas, making common repairs less involved than on many thin systems. The customer-replaceable battery is especially useful for companies planning to keep these laptops deployed for several years.

Dell Pro 7 14 AMD review insides

Display and Input

The 14-inch WUXGA display in our review build has a 1920 x 1200 resolution, giving it a taller 16:10 aspect ratio that provides more vertical space for documents, spreadsheets, and web pages. It is a non-touch panel with variable refresh rate support, 500-nit brightness, full sRGB coverage, an anti-glare finish, and Dell’s ComfortView Plus (Low Blue Light) certification. The super-low-power panel also helps battery life. Our unit lasted 19 hours and 28 minutes in the PCMark 10 Modern Office test.

Dell pairs the display with an 8MP infrared camera that supports Windows Hello facial recognition, presence detection, temporal noise reduction, and a physical camera shutter. The Mini-LED backlit keyboard uses the available width well, with full-sized primary keys and a familiar layout that does not require much adjustment. There is no numeric keypad on this 14-inch model, but the centered keyboard and large precision touchpad leave plenty of room for everyday work. The display bezels are reasonably narrow along the sides, although the camera hardware requires a slightly thicker section along the top.

Ports and Connectivity

Port selection is good for a laptop this thin, with our review build including three USB Type-C connections. The left side has two Thunderbolt 4 ports with USB4, DisplayPort 2.1, and Power Delivery, plus HDMI 2.1 and a USB 3.2 Gen 1 Type-A port. The right side adds a third USB Type-C port with USB 3.2 Gen2x2 (20 Gbps), DisplayPort 1.4, and Power Delivery, along with a headset jack and wedge-shaped security slot. Dell offers that third Type-C connection as an alternative to a second USB Type-A port with PowerShare, so the exact layout depends on the configuration.

It also includes MediaTek Wi-Fi 7 MT7925 and Bluetooth for wireless connections. Charging is done via USB-C using the included 65W adapter, which can be connected on either side of the laptop. The 70Wh battery supports ExpressCharge and ExpressCharge Boost when paired with a 100W adapter, which Dell offers as an option. With the included 65W adapter, standard charging applies.

Security and Manageability

The Dell Pro 7 14 is built for managed business environments, and much of its value comes from features not shown in benchmark charts. AMD PRO manageability includes DASH and AMD Integrated Management Technology, or AIM-T, giving IT departments tools to monitor, configure, and support systems remotely. Dell Management Portal can also work alongside Microsoft Intune, allowing administrators to manage Dell-specific settings through an existing cloud-based device-management setup.

Dell adds several layers of hardware and firmware protection through SafeBIOS, SafeID, and Trusted Device. SafeBIOS monitors BIOS settings and detects unexpected changes, while SafeID keeps credentials in dedicated hardware away from the operating system. Quantum-resistant BIOS verification protects firmware updates against current and emerging cryptographic threats. Our review build also has a fingerprint reader, and the infrared camera provides another Windows Hello sign-in option.

Our configuration comes with Dell ProSupport and next-business-day onsite service for 36 months following remote diagnosis. If Dell determines that hardware needs to be replaced, a technician can be sent to the customer’s location, reducing the time an employee is left without their primary laptop.

Dell Pro 7 14 (AMD) Specifications

Specification Dell Pro 7 14 (P714265)
Platform Overview
Processor AMD Ryzen AI 9 HX PRO 470
12 cores / 24 threads, up to 5.2 GHz
55 TOPS NPU (Copilot+ PC)
Graphics AMD Radeon 890M (integrated)
Operating System Windows 11 Pro (Copilot+ PC)
Memory and Storage
Memory 64 GB LPDDR5x-8533, dual-channel, onboard
Storage 1 TB SSD (Samsung BM9C1a)
Display and Camera
Display 14″ WUXGA (1920 x 1200), non-touch, VRR
500 nits, 100% sRGB, anti-glare, Low Blue Light, super-low-power
Camera 8 MP + IR (Windows Hello)
Connectivity and Input
Wireless MediaTek Wi-Fi 7 MT7925, Bluetooth
Keyboard English (US) Mini-LED backlit
Ports 2x Thunderbolt 4 (USB4, DisplayPort 2.1, Power Delivery); 1x USB-C 3.2 Gen 2×2 (DisplayPort 1.4, Power Delivery); 1x USB 3.2 Gen 1 Type-A; HDMI 2.1; headset jack; wedge-shaped lock slot. Configurations without the third USB-C port include a second USB Type-A port with PowerShare instead.
Security and Manageability
Security Fingerprint reader
Dell SafeBIOS, SafeID, Trusted Device
Quantum-resistant BIOS verification
Manageability AMD PRO manageability, AMD DASH, AMD Integrated Management Technology
Dell Management Portal, Microsoft Intune
Power and Physical
Battery 3-cell, 70 Wh, Long Life Cycle, ExpressCharge / ExpressCharge Boost
Power Adapter 65 W USB-C
Chassis Aluminum (Top Cover, Palmrest), Lightweight Magnesium (Bottom Cover)
Weight / Dimensions From 2.80 lb; 315.5 x 226 mm; 10.3 to 16.45 mm thick.
Warranty and Pricing
Service Dell ProSupport, Next Business Day Onsite, 36 months
Base Price $2,279
Price as Tested $5,377 (Dell.com single-unit, no discount)

Performance Testing

To see how the Dell Pro 7 14 AMD compares with Dell’s current commercial lineup, we tested it alongside the larger Pro 5 16 AMD and both 14-inch Intel models. All four laptops were tested using our standard power profile across general productivity, CPU rendering, GPU compute, professional visualization, storage, AI, and battery workloads.

The Pro 7 14 and Pro 5 16 use the same Ryzen AI 9 HX PRO 470 and Radeon 890M, but the smaller model has faster LPDDR5x-8533 memory, giving it an advantage in several integrated graphics, memory-heavy, and AI tests. Its thinner chassis limits sustained multi-core work, giving the Pro 5 16 more room. Comparisons with the Intel laptops vary by application. AMD performs better across many SPECviewperf CAD viewsets, while Intel’s Arc graphics lead in Blender GPU rendering.

Test Systems

Specification Dell Pro 5 16 (AMD) Dell Pro 7 14 (AMD) Dell Pro 7 14 (Intel) Dell Pro 5 14 (Intel)
CPU Ryzen AI 9 HX PRO 470 (12C/24T) Ryzen AI 9 HX PRO 470 (12C/24T) Core Ultra 7 366H (16C) Core Ultra X7 368H (16C)
GPU Radeon 890M Radeon 890M Intel Graphics Intel Arc B390
Memory 64 GB DDR5-5600 64 GB LPDDR5x-8533 64 GB LPDDR5x-8533 64 GB LPCAMM2-8533
Storage SanDisk PC SN5100S 1 TB Samsung BM9C1a 1 TB SK hynix PCB01 1 TB SK hynix PCB01 1 TB
Display 16″ WQXGA 14″ WUXGA 14″ WUXGA 14″ WUXGA
Price as Tested $4,648 $5,377 $5,600 $5,492

UL Procyon: AI Computer Vision

The Procyon AI Computer Vision Benchmark measures AI inference performance across CPUs, GPUs, and dedicated accelerators using a range of state-of-the-art neural networks. It evaluates tasks such as image classification, object detection, segmentation, and super-resolution using models including MobileNet V3, Inception V4, YOLO V3, DeepLab V3, Real ESRGAN, and ResNet 50. Tests are run on multiple inference engines, including NVIDIA TensorRT, Intel OpenVINO, Qualcomm SNPE, Microsoft Windows ML, and Apple Core ML, providing a broad view of hardware and software efficiency. Results are reported for float- and integer-optimized models, providing a consistent, practical measure of machine vision performance for professional workloads.

The Dell Pro 7 14 AMD scored 243 overall in the UL Procyon AI Computer Vision GPU benchmark, making it the second-fastest system in the comparison. It outperformed the Dell Pro 5 16 AMD’s score of 214 by roughly 14% and held an 18.5% lead over the Dell Pro 7 14 Intel, which posted 205. The only system ahead was the Dell Pro 5 14 Intel at 398, a substantial 64% advantage. Despite not taking the top position, the Pro 7 14 AMD demonstrated a strong showing for an ultraportable business notebook, particularly given its balanced performance across the suite’s diverse machine vision workloads.

On the CPU side, the Dell Pro 7 14 AMD recorded an overall score of 77, trailing the Dell Pro 5 16 AMD’s 91 and both Intel configurations, which scored 121 and 119, respectively. This placed the system approximately 15% behind the larger AMD notebook and roughly 36% behind the leading Intel Pro 7 14. While Intel’s software optimizations continue to provide an advantage in CPU-based inference, AMD’s GPU-accelerated results in the Pro 7 14 paint a much more competitive picture, highlighting the importance of leveraging modern AI accelerators for professional computer vision applications.

CPU Results (average time in ms) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
AI Computer Vision Overall Score 91 77 121 119
MobileNet V3 1.57 ms 1.79 ms 1.24 ms 1.09 ms
ResNet 50 13.64 ms 16.46 ms 11.57 ms 10.11 ms
Inception V4 40.43 ms 51.10 ms 34.01 ms 29.77 ms
DeepLab V3 69.73 ms 74.74 ms 37.98 ms 50.18 ms
YOLO V3 99.38 ms 122.35 ms 81.22 ms 112.64 ms
REAL-ESRGAN 4,504.35 ms 5,383.21ms 3,163.45 ms 2,884.51 ms
GPU Results (average time in ms) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
AI Computer Vision Overall Score 214 243 205 398
MobileNet V3 1.56 ms 1.17 ms 1.04 ms 0.84 ms
ResNet 50 8.76 ms 7.93 ms 6.52 ms 2.94 ms
Inception V4 22.46 ms 20.06 ms 20.72 ms 8.38 ms
DeepLab V3 38.40 ms 34.04 ms 26.53 ms 17.54 ms
YOLO V3 23.56 ms 21.92 ms 43.07 ms 19.14 ms
REAL-ESRGAN 570.35 ms 536.81 ms 1,309.49 ms 562.87 ms

UL Procyon: AI Text Generation

The Procyon AI Text Generation Benchmark streamlines LLM performance testing by providing a concise, consistent evaluation method. It enables repeated testing across multiple LLM models while minimizing the complexity of large models and the number of variables. Developed with AI hardware leaders, it optimizes the use of local AI accelerators to deliver more reliable, efficient performance assessments.

The Dell Pro 7 14 AMD consistently improved upon the larger Dell Pro 5 16 AMD across all UL Procyon AI Text Generation workloads, highlighting the benefits of its tuning and implementation of the Ryzen AI 9 HX PRO 470 platform. In the Phi test, the Pro 7 14 AMD achieved an overall score of 427, a 15% improvement over the Pro 5 16’s 371. Similar gains were observed in Mistral (389 versus 346, +12%), Llama3 (345 versus 306, +13%), and Llama2 (367 versus 329, +12%). It also delivered lower time-to-first-token metrics and higher token generation rates across the board, making it the stronger of the two AMD-based systems for local LLM inference.

Compared to its Intel counterparts, however, the Dell Pro 7 14 AMD trailed in every workload. The Intel-based Dell Pro 7 14 posted scores of 689, 525, 509, and 530 in Phi, Mistral, Llama3, and Llama2, respectively, representing advantages ranging from 35% to 61% over the AMD model. The Dell Pro 5 14 widened the gap further, leading the group with scores of 904 in Phi, 716 in Mistral, 708 in Llama3, and 641 in Llama2. Intel’s systems also substantially reduced time-to-first-token, with the Pro 5 14 reaching just 1.158 seconds in Phi compared to 4.286 seconds on the Pro 7 14 AMD, while simultaneously delivering higher sustained token throughput.

UL Procyon: AI Text Generation Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Phi
Phi Overall Score 371 427 689 904
Phi Output Time To First Token 5.052 s 4.286 s 1.843 s 1.158 s
Phi Output Tokens Per Second 27.163 tokens/s 30.492 tokens/s 34.172 tokens/s 36.952 tokens/s
Phi Overall Duration 135.343 s 118 s 92.273 s 81.304 s
Mistral
Mistral Overall Score 346 389 525 716
Mistral Output Time To First Token 6.929 s 6.060 s 3.778 s 2.195 s
Mistral Output Tokens Per Second 18.174 tokens/s 20.125 tokens/s 22.843 tokens/s 24.724 tokens/s
Mistral Overall Duration 196.185 s 176 s 143.831 s 123.773 s
Llama3
Llama3 Overall Score 306 345 509 708
Llama3 Output Time To First Token 6.518 s 5.707s 2.992 s 1.657 s
Llama3 Output Tokens Per Second 15.050 tokens/s 16.725 tokens/s 19.104 tokens/s 20.421 tokens/s
Llama3 Overall Duration 224.176 s 199 s 161.290 s 142.814 s
Llama2
Llama2 Overall Score 329 367 530 641
Llama2 Output Time To First Token 11.089 s 10.307 s 5.412 s 4.082 s
Llama2 Output Tokens Per Second 8.878 tokens/s 10.260 tokens/s 11.230 tokens/s 12.397 tokens/s
Llama2 Overall Duration 381.464 s 334 s 278.109 s 245.572 s

UL Procyon: AI Image Generation

The Procyon AI Image Generation Benchmark provides a consistent and accurate method for measuring AI inference performance across a range of hardware, from low-power NPUs to high-end GPUs. It includes three tests: Stable Diffusion XL (FP16) for high-end GPUs, Stable Diffusion 1.5 (FP16) for moderately powerful GPUs, and Stable Diffusion 1.5 (INT8) for low-power devices. The benchmark uses the optimal inference engine for each system, ensuring fair and comparable results.

The Dell Pro 7 14 AMD delivered largely middle-of-the-pack results in the UL Procyon AI Image Generation benchmark, though it remained highly competitive with the other non-leading systems. In Stable Diffusion 1.5 (FP16), it posted an overall score of 247 with an image generation speed of 25.3 seconds per image, placing it about 4% behind the Dell Pro 7 14 Intel (258) and roughly 3% behind the Dell Pro 5 16 AMD (255). The Dell Pro 5 14 was the clear outlier, however, producing a score of 635 and generating images approximately 2.6 times faster at 9.8 seconds per image.

Stable Diffusion 1.5 (INT8) told a similar story. The Dell Pro 7 14 AMD scored 3,521, essentially tying the Dell Pro 5 16 AMD (3,598) and Dell Pro 7 14 Intel (3,575), with less than a 3% spread separating the three systems. Image generation speeds were nearly identical as well, ranging from 8.69 to 8.87 seconds per image. Once again, the Dell Pro 5 14 established a substantial lead, posting a score of 7,693 and cutting generation times to just 4.06 seconds per image, more than twice as fast as the rest of the field.

UL Procyon: AI Image Generation Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Stable Diffusion 1.5 (FP16)
Stable Diffusion 1.5 (FP16) – Overall Score 255 247 258 635
Stable Diffusion 1.5 (FP16) – Overall Time 391.577 s 404.455 s 386.614 s 157.296 s
Stable Diffusion 1.5 (FP16) – Image Generation Speed 24.474 s/image 25.278 s/image 24.163 s/image 9.831 s/image
Stable Diffusion 1.5 (INT8)
Stable Diffusion 1.5 (INT8) – Overall Score 3,598 3,521 3,575 7,693
Stable Diffusion 1.5 (INT8) – Overall Time 69.478 s 70.985 s 69.911 s 32.495 s
Stable Diffusion 1.5 (INT8) – Image Generation Speed 8.685 s/image 8.873 s/image 8.739 s/image 4.062 s/image
Stable Diffusion XL (FP16)
Stable Diffusion XL (FP16) – Overall Score 173 177 268 646
Stable Diffusion XL (FP16) – Overall Time 3,448.478 s 3,379.388 s 2,230.563 s 928.747 s
Stable Diffusion XL (FP16) – Image Generation Speed 215.530 s/image 211.212 s/image 139.410 s/image 58.047 s/image

PCMark 10

PCMark 10 measures general system performance across everyday work such as web browsing, video conferencing, spreadsheets, writing, photo editing, and rendering. Higher scores are better.

The Dell Pro 7 14 AMD scored 8,237 overall in PCMark 10, putting it within 2.5% of the Pro 7 14 Intel’s leading score of 8,438. The four laptops were close in Essentials, where the AMD model scored 10,783, but it moved into first place in Productivity with 14,366. Digital Content Creation reached 9,792, only 60 points behind the larger Pro 5 16 but 818 points behind the Pro 7 14 Intel. For common office work, the Pro 7 14 AMD performed much like the other Dell models and had the best Productivity result of the group.

PCMark 10 (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Overall 8,268 8,237 8,438 7,945
Essentials 10,870 10,783 10,981 10,751
Productivity 14,322 14,366 13,992 13,821
Digital Content Creation 9,852 9,792 10,610 9,158

Geekbench 6

Geekbench 6 measures processor performance using a mix of common tasks, with separate scores for single-core and multi-core workloads. Higher scores are better.

The Dell Pro 7 14 AMD scored 2,888 in Geekbench 6 single-core, keeping it within a relatively narrow range of 127 points across all four laptops. Its multi-core score of 14,768 was 420 points higher than the larger Pro 5 16, despite both systems using the Ryzen AI 9 HX PRO 470. The two 16-core Intel models were faster in this portion of the test, scoring just under 17,000, but the Pro 7 14 AMD still performed well for a thin 14-inch laptop.

Geekbench 6 (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Single-Core 2,989 2,888 2,862 2,968
Multi-Core 14,348 14,768 16,787 16,874

Cinebench R23 and 2024

Cinebench measures how quickly the processor can render a complex scene, with separate tests for single-core and multi-core performance. Higher scores are better.

The Dell Pro 7 14 AMD scored 1,946 in Cinebench R23 single-core and 15,173 in multi-core, while the larger Pro 5 16 finished about 24% ahead in the multi-core test. Since both laptops use the same Ryzen AI 9 HX PRO 470, the additional thermal room available in the 16-inch model likely contributed to its higher score. Cinebench 2024 followed a similar pattern, with the Pro 7 14 AMD scoring 105 in single-core and 847 in multi-core, compared with 119 and 1,055 for the Pro 5 16. Even with that difference, its 2024 multi-core result beat both Intel laptops.

Cinebench (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
R23 Single-Core 2,029 1,946 2,043 2,010
R23 Multi-Core 18,764 15,173 14,640 16,915
2024 Single-Core 119 105 116 122
2024 Multi-Core 1,055 847 683 807

7-Zip Compression

The 7-Zip benchmark measures how quickly the processor can compress and decompress data using multiple threads. Higher GIPS scores are better.

The Dell Pro 7 14 AMD recorded 90.2 GIPS in 7-Zip, placing second behind the Pro 5 16 at 103.9 GIPS. It narrowly passed the Pro 5 14 Intel’s 89.4 GIPS and finished 9 GIPS ahead of the Pro 7 14 Intel. The larger AMD laptop had an advantage during sustained compression, but the Pro 7 14 still produced the best result among the three 14-inch systems.

7-Zip (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Total Rating (GIPS) 103.9 90.2 81.2 89.4

y-cruncher

y-cruncher measures how quickly the processor can calculate large numbers of Pi digits, placing a heavy load on the CPU and memory. Results are measured in seconds, so lower times are better.

The Dell Pro 7 14 AMD completed the 1-billion-digit y-cruncher test in 25.199 seconds, only 0.039 seconds behind the Pro 5 16. The larger AMD model gained more distance as the calculation increased, finishing the 2.5-billion test in 73.320 seconds compared with 79.393 seconds for the Pro 7 14. At 5 billion digits, the Pro 7 14 took 177.776 seconds, about 14 seconds longer than the Pro 5 16 but roughly 27 seconds faster than either Intel system.

y-cruncher — seconds (lower is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
1 Billion 25.160 25.199 34.775 29.405
2.5 Billion 73.320 79.393 104.586 90.719
5 Billion 163.768 177.776 240.554 204.685

Blender 5.1.1 (GPU)

The Blender benchmark measures GPU rendering performance using three different 3D scenes: Monster, Junkshop, and Classroom. Results are reported in samples per minute, so higher scores are better.

The Radeon 890M in the Dell Pro 7 14 AMD reached 140.9 samples per minute in Monster, 122.0 in Junkshop, and 105.6 in Classroom. Those results were 9% to 18% faster than the Pro 5 16, even though both laptops use the same integrated GPU. The Pro 7 14’s faster LPDDR5x-8533 memory likely helped here, since the Radeon 890M shares system memory. Both Intel laptops were much faster in Blender, however, with the Arc B390-equipped Pro 5 14 leading all three scenes.

Blender 5.1.1 GPU — samples/min (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Monster 129.8 140.9 253.6 366.6
Junkshop 103.1 122.0 189.3 312.6
Classroom 91.1 105.6 153.0 244.4

V-Ray and LuxMark

LuxMark measures GPU compute performance by rendering complex scenes through OpenCL, while V-Ray GPU measures how quickly the graphics processor can render a scene using the V-Ray engine. Higher scores are better.

The Dell Pro 7 14 AMD scored 2,042 in LuxMark Hall, 1,050 in LuxMark Food, and 784 vpaths in V-Ray GPU. LuxMark Hall placed it slightly behind the Pro 5 16 and Pro 7 14 Intel, while its Food result beat both of those systems and trailed only the Arc B390-equipped Pro 5 14. V-Ray was close across the group, although the Pro 7 14 AMD finished ahead of only the Pro 7 14 Intel. The Radeon 890M performed reasonably well in these tests, but the Arc B390 had a large advantage in both LuxMark scenes.

GPU Compute (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
LuxMark — Hall 2,125 2,042 2,168 3,287
LuxMark — Food 982 1,050 880 1,585
V-Ray GPU (vpaths) 861 784 775 919

SPECviewperf 15

SPECviewperf 15 measures graphics performance using viewsets based on professional applications for CAD, 3D modeling, rendering, engineering, and medical visualization. Higher scores are better, although performance can vary considerably between applications and graphics architectures.

The Dell Pro 7 14 AMD performed particularly well in the CAD-focused portions of SPECviewperf 15, leading the group in 3ds Max (24.69), CATIA (21.21), Creo (45.70), and Siemens NX (51.84). It also led the two AMD systems in Enscape (8.29) and Maya (53.64), although the Arc B390-equipped Pro 5 14 posted the highest scores in both tests at 14.28 and 82.42. The Pro 5 14 Intel also led Blender (21.19) and Unreal Engine (38.77). The Pro 7 14 AMD remained close to the Pro 5 16 in Energy (25.08), Medical (60.50), and SolidWorks (32.02), while both Radeon 890M systems were much faster than the Pro 7 14 Intel across the engineering viewsets.

SPECviewperf 15 (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
3dsmax-08 23.21 24.69 9.58 20.21
blender-01 19.89 20.85 9.11 21.19
catia-07 20.16 21.21 5.27 10.27
creo-04 44.27 45.70 18.27 32.53
energy-04 25.39 25.08 3.68 10.78
enscape-01 8.02 8.29 6.22 14.28
maya-07 48.67 53.64 49.54 82.42
medical-04 65.45 60.50 9.91 22.96
snx-05 51.67 51.84 37.74 46.22
solidworks-08 33.12 32.02 11.86 23.60
unreal_engine-01 26.61 26.29 21.36 38.77

SPECworkstation 4

SPECworkstation 4 measures workstation performance across CPU, graphics, storage, AI, product design, engineering, financial services, and other professional workloads. Higher scores are better, while DNF means the system did not complete every workload required for that category.

The Dell Pro 7 14 AMD led the Graphics subsystem (2.41), narrowly passing the Pro 5 16 and finishing far ahead of both Intel laptops. Its CPU score (1.00) was the lowest of the four, while AI and Machine Learning (1.29), Energy (1.13), Financial Services (0.88), and Life Sciences (1.08) placed it around the middle of the group. Product Design (1.18) and Productivity and Development (0.72) were also behind the other systems. Storage (0.55) was the weakest result, reflecting the slower Samsung SSD in this system, while Media and Entertainment did not finish.

SPECworkstation 4 (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
CPU (subsystem) 1.09 1.00 1.08 1.19
Graphics (subsystem) 2.36 2.41 0.82 1.68
Storage (subsystem) 0.89 0.55 1.68 1.76
AI & Machine Learning 1.37 1.29 1.18 1.36
Energy 1.24 1.13 0.90 1.18
Financial Services 0.98 0.88 0.78 0.78
Life Sciences 1.34 1.08 1.04 1.34
Media & Entertainment DNF DNF 1.16 DNF
Product Design 1.34 1.18 1.41 1.64
Productivity & Development 0.78 0.72 1.04 1.10

Storage Performance

3DMark Storage measures how an SSD performs during gaming-related tasks such as loading games, installing software, saving progress, and moving game files. Blackmagic Disk Speed Test measures an SSD’s sequential read and write speeds using large media files.

The 1TB Samsung BM9C1a in the Dell Pro 7 14 AMD scored 894 in 3DMark Storage, well behind the other three drives in the comparison. Blackmagic Disk measured 3,103.4MB/s read and 4,034.3MB/s write, compared with more than 8,000MB/s from the SK hynix Gen5 drives installed in both Intel laptops. Much of that gap comes down to the specific SSD in our review build, but Dell’s spec sheet notes that Gen5 SSDs run at Gen4 speed on the AMD version of the Pro 7, so even upgraded configurations will not match the sequential numbers of the Intel units. Either way, storage is one of the weaker areas of this configuration, especially considering its $5,377 price.

Storage (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
3DMark Storage (score) 2,477 894 3,259 3,144
Blackmagic Disk — Read (MB/s) 4,758.0 3,103.4 8,398.6 8,609.6
Blackmagic Disk — Write (MB/s) 5,166.5 4,034.3 8,934.5 8,747.3

Battery Life

The PCMark 10 Modern Office battery test repeatedly runs common office tasks until the battery reaches the test’s cutoff point. Longer runtimes are better.

The Dell Pro 7 14 AMD lasted 19 hours and 28 minutes in the PCMark 10 Modern Office battery test, which was run in Balanced mode at 50% display brightness. That was more than four hours longer than the 16-inch Pro 5, but roughly seven hours behind both 14-inch Intel systems. All four laptops use a 70Wh battery, so the comparison also shows the efficiency advantage of the Intel configurations during lighter office workloads. Even with that gap, the Pro 7 14 AMD provided enough runtime for a long day away from an outlet.

Battery — PCMark 10 Modern Office (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Runtime 15h 22m 19h 28m 26h 18m 26h 48m

Conclusion

The Dell Pro 7 14 AMD is a decent choice for business users who want a portable 14-inch laptop with excellent integrated graphics and local AI performance. Its Ryzen AI 9 HX PRO 470 and 64GB of LPDDR5x-8533 memory worked especially well together, helping it lead the AMD pair in Blender, several SPECviewperf viewsets, the Procyon AI Computer Vision GPU test, and all four AI text-generation workloads. It also led PCMark 10 Productivity and produced competitive single-core performance, giving it plenty of speed for office work, heavier multitasking, CAD applications, and supported local AI tools.

The smaller chassis does place limits on sustained multi-core performance, with the larger Pro 5 16 finishing Cinebench R23 multi-core about 24% faster despite using the same processor. Our review unit’s Samsung BM9C1a SSD was also the slowest of the four Dell drives, trailing the Gen5 SSDs in the Intel systems by a wide margin. Battery life reached 19 hours and 28 minutes, which is excellent for a full day of work, but the two 14-inch Intel models lasted close to 27 hours. Buyers who prioritize long rendering workloads, faster storage, or maximum battery life have better options among the other three Dell configurations.

Pricing may be the largest concern, since our Pro 7 14 AMD review build costs $5,377 before commercial discounts, which is $729 more than the Pro 5 16. However, that premium pays for a thinner and more portable design, faster memory, stronger integrated graphics, and better AI text-generation performance than its larger AMD counterpart. Dell also backs it with AMD PRO manageability, its commercial security tools, and three years of next-business-day onsite support. For companies that need those features in a compact AMD laptop, the Pro 7 14 has a strong case, but buyers focused mainly on sustained CPU performance will get more for their money from the Pro 5 16.

Product Page: Dell Pro 7 14

The post Dell Pro 7 14 AMD Review: Ryzen AI 9 HX PRO 470 in a 2.8-Pound Business Laptop appeared first on StorageReview.com.

Dell Pro 5 16 (AMD) Review: The Desk-First AMD Option in Dell’s Pro Laptop Line

21 July 2026 at 20:30
Dell Pro 5 16 inside Dell Pro 5 16 inside

The Dell Pro 5 16 is the large-screen entry in Dell’s commercial Pro laptop line, and the one that leads on value rather than portability. Our review unit pairs the AMD Ryzen AI 9 HX PRO 470 (12 cores, 24 threads, 55 TOPS NPU) with AMD Radeon 890M integrated graphics, 64 GB of DDR5-5600 RAM, a 1 TB SSD, and a 16-inch WQXGA display. At $4,648 as configured (Dell.com single-unit, before the volume discounts most fleets actually pay), it comes in below the rest of the Pro family we tested.

Dell Pro 5 16 review hero

To be fair, this is a commercial fleet machine. The questions that matter are sustained productivity, manageability, serviceability, battery, and total cost of ownership. It ships with AMD PRO manageability, Dell SafeBIOS and SafeID, quantum-resistant BIOS verification, and ProSupport Next Business Day Onsite coverage; the details IT buyers weigh more heavily than a spec-sheet clock speed.

What makes the Pro 5 16 interesting is that it runs the same silicon as the 14-inch Pro 7 AMD but changes two variables: it moves to slower DDR5-5600 memory, versus LPDDR5x-8533 in the thinner units, and it puts that chip in a larger 16-inch chassis. Those two choices define its performance character, and not in the same direction. The larger chassis lets the Ryzen AI 9 HX PRO 470 maintain higher sustained clocks; the Pro 5 16 posts the strongest multi-core Cinebench result among the four Dell units we tested. The slower memory is the ceiling for its integrated graphics, and the 16-inch panel is why it finishes last in the group on battery life. Paired with the full numeric keypad, this is a desk-first productivity machine: strong at sustained CPU work and everyday multitasking, weaker where memory bandwidth or all-day unplugged runtime matter most.

Design and Build

With its 16-inch screen and full-size keyboard, the Dell Pro 5 16 is the largest laptop in this group, measuring 14.12 inches wide, 9.98 inches deep, and up to 0.75 inches thick, with a starting weight of 4.02 pounds. That is still a noticeable jump from the 14-inch models, especially once the charger is added to a bag, but the larger footprint pays off with more screen space and a roomier keyboard.

Dell Pro 5 16 review closed

The Magnetite aluminum chassis has a dark finish with a light texture across the lid, keyboard deck, and bottom cover. It has the usual plain, business-oriented appearance with little beyond the silver Dell logo. The finish does a good job of hiding fingerprints, too.

Dell Pro 5 16 review inside

Once the bottom cover is removed, the two DDR5 SODIMM slots, M.2 SSD, wireless card, cooling fan, and 70Wh battery are all accessible. Both memory modules can be replaced, and the battery is also designed for customer replacement, giving businesses more options for repairs and upgrades as the laptop ages. For cooling, it features a single large fan and heat pipe, with a wide intake grille covering much of the underside and an exhaust vent running along the rear.

Display and Input

Our review build has the upgraded 16-inch WQXGA display, with a 2560 x 1600 resolution, variable refresh rate support, 500-nit brightness, and full sRGB coverage. The combination gives Windows, photos, and video a sharper and more colorful appearance than Dell’s lower-resolution display options. Its anti-glare coating also helps reduce reflections under office lighting, while Low Blue Light technology is included for longer work sessions.

Dell Pro 5 16 front

Having a 16:10 panel gave us plenty of vertical room for documents, spreadsheets, web pages, and applications with crowded interfaces. The extra width is also useful when working with two windows side by side, particularly when the laptop is being used away from an external monitor. Our panel does not support touch, but that likely isn’t too big of a deal for the business users this configuration is built for.

Dell uses the wider keyboard deck to include a dedicated numeric keypad, which is a useful difference from the three 14-inch laptops in this review group. Anyone who regularly works with spreadsheets, accounting software, or large sets of numerical data should find it much quicker than relying on the number row. The Mini-LED backlighting provides even illumination around the keys, and the keyboard also includes a Copilot key and a fingerprint reader built into the power button.

Below the keyboard is a large clickpad that provides plenty of room for navigation and multi-finger gestures. Since the number pad shifts the main typing area to the left, the clickpad is also positioned left of the laptop’s centerline. An 8MP HDR camera is located above the display with infrared support, presence detection, Windows Hello facial recognition, and a physical privacy shutter.

Ports and Connectivity

The Dell Pro 5 16 offers a good selection of ports across both sides of the chassis. On the left are HDMI 2.1, one USB 3.2 Gen 1 Type-A port, and two 40Gbps Thunderbolt 4 Type-C ports with DisplayPort 2.1 and USB Power Delivery. Either Type-C port can be used with the included 65W charger, leaving some flexibility when deciding where to route the cable on a desk.

Along the right side are a second USB 3.2 Gen 1 Type-A port with PowerShare, a global headset jack, Gigabit Ethernet, and a Kensington wedge-shaped lock slot. Smart-card and nano-SIM slots are available as optional additions, depending on the selected configuration. The built-in RJ45 connection is particularly useful in an office or test environment, where a wired network connection may be preferable to carrying a USB Ethernet adapter.

Dell Pro 5 16 left side

Wireless connectivity includes a MediaTek MT7925 adapter supporting Wi-Fi 7 and Bluetooth 5.4. The 70Wh battery supports Dell ExpressCharge and is ExpressCharge Boost capable.

Security and Manageability

Business buyers will find most of the security and fleet-management features they are likely to need already included in this configuration. The Ryzen AI 9 HX PRO 470 brings AMD PRO management and security features, while Dell Management Portal can work with Microsoft Intune to help IT teams configure systems, distribute updates, and oversee a larger device fleet.

Dell Pro 5 16 (AMD) Specifications

Specification Dell Pro 5 16 (P516265)
Platform Overview
Processor AMD Ryzen AI 9 HX PRO 470
12 cores / 24 threads, up to 5.2 GHz
55 TOPS NPU (Copilot+ PC)
Graphics AMD Radeon 890M (integrated)
Operating System Windows 11 Pro (Copilot+ PC)
Memory and Storage
Memory 64 GB DDR5-5600 (2 × 32 GB), dual-channel
Storage 1 TB SSD (SanDisk PC SN5100S)
Display and Camera
Display 16″ WQXGA (2560 × 1600), non-touch, VRR
500 nits, 100% sRGB, anti-glare, Low Blue Light
Camera 8 MP + IR (Windows Hello)
Connectivity and Input
Wireless MediaTek Wi-Fi 7 MT7925, Bluetooth 5.4
Keyboard US English mini-LED backlit with numeric keypad
Ports 2x Thunderbolt 4 40Gbps Type-C (DisplayPort 2.1, Power Delivery); 2x USB 3.2 Gen 1 Type-A (one with PowerShare); HDMI 2.1; Gigabit Ethernet (RJ45); global headset jack; wedge-shaped lock slot. Optional smart-card reader and nano-SIM slot.
Security and Manageability
Security Fingerprint reader
Dell SafeBIOS, SafeID, Trusted Device
Quantum-resistant BIOS verification
Manageability AMD PRO manageability
Dell Management Portal, Microsoft Intune
Power and Physical
Battery 3-cell, 70 Wh, ExpressCharge / ExpressCharge Boost
Power Adapter 65 W USB-C
Weight / Dimensions From 4.02 lb; 14.12 x 9.98 x up to 0.75 in.
Warranty and Pricing
Service Dell ProSupport, Next Business Day Onsite, 36 months
Base Price $2,089
Price as Tested $4,648 (Dell.com single-unit, no discount)

Performance Testing

Test Systems

Specification Dell Pro 5 16 (AMD) Dell Pro 7 14 (AMD) Dell Pro 7 14 (Intel) Dell Pro 5 14 (Intel)
CPU Ryzen AI 9 HX PRO 470 (12C/24T) Ryzen AI 9 HX PRO 470 (12C/24T) Core Ultra 7 366H (16C) Core Ultra X7 368H (16C)
GPU Radeon 890M Radeon 890M Intel Graphics Intel Arc B390
Memory 64 GB DDR5-5600 64 GB LPDDR5x-8533 64 GB LPDDR5x-8533 64 GB LPCAMM2-8533
Storage SanDisk PC SN5100S 1 TB Samsung BM9C1a 1 TB SK hynix PCB01 1 TB SK hynix PCB01 1 TB
Display 16″ WQXGA 14″ WUXGA 14″ WUXGA 14″ WUXGA
Price as Tested $4,648 $5,377 $5,600 $5,492

UL Procyon: AI Computer Vision

The Procyon AI Computer Vision Benchmark measures AI inference performance across CPUs, GPUs, and dedicated accelerators using a range of state-of-the-art neural networks. It evaluates tasks such as image classification, object detection, segmentation, and super-resolution using models including MobileNet V3, Inception V4, YOLO V3, DeepLab V3, Real ESRGAN, and ResNet 50. Tests are run on multiple inference engines, including NVIDIA TensorRT, Intel OpenVINO, Qualcomm SNPE, Microsoft Windows ML, and Apple Core ML, providing a broad view of hardware and software efficiency. Results are reported for float- and integer-optimized models, providing a consistent, practical measure of machine vision performance for professional workloads.

The Dell Pro 5 16 scored 91 in the CPU test, placing it ahead of the Pro 7 14 AMD at 77 but behind both Intel systems. Its Ryzen processor completed every model faster than the same chip in the Pro 7 14, including YOLO V3 at 99.38 ms versus 122.35 ms and REAL-ESRGAN at 4,504.35 ms versus 5,383.21 ms. Its Radeon 890M raised the GPU score to 214, edging past the Pro 7 14 Intel at 205 but trailing the Pro 7 14 AMD at 243 and the Pro 5 14 Intel at 398. The smaller AMD laptop’s faster LPDDR5x-8533 memory appears to help its integrated graphics, while the Arc B390 gives the Pro 5 14 Intel a large advantage in this test.

CPU Results (average time in ms) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
AI Computer Vision Overall Score 91 77 121 119
MobileNet V3 1.57 ms 1.79 ms 1.24 ms 1.09 ms
ResNet 50 13.64 ms 16.46 ms 11.57 ms 10.11 ms
Inception V4 40.43 ms 51.10 ms 34.01 ms 29.77 ms
DeepLab V3 69.73 ms 74.74 ms 37.98 ms 50.18 ms
YOLO V3 99.38 ms 122.35 ms 81.22 ms 112.64 ms
REAL-ESRGAN 4,504.35 ms 5,383.21ms 3,163.45 ms 2,884.51 ms
GPU Results (average time in ms) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
AI Computer Vision Overall Score 214 243 205 398
MobileNet V3 1.56 ms 1.17 ms 1.04 ms 0.84 ms
ResNet 50 8.76 ms 7.93 ms 6.52 ms 2.94 ms
Inception V4 22.46 ms 20.06 ms 20.72 ms 8.38 ms
DeepLab V3 38.40 ms 34.04 ms 26.53 ms 17.54 ms
YOLO V3 23.56 ms 21.92 ms 43.07 ms 19.14 ms
REAL-ESRGAN 570.35 ms 536.81 ms 1,309.49 ms 562.87 ms

UL Procyon: AI Text Generation

The Procyon AI Text Generation Benchmark streamlines LLM performance testing by providing a concise, consistent evaluation method. It enables repeated testing across multiple LLM models while minimizing the complexity of large models and the number of variables. Developed with AI hardware leaders, it optimizes the use of local AI accelerators to deliver more reliable, efficient performance assessments.

Local text generation was a weak area for the Dell Pro 5 16, as it finished last with all four language models. Scores ranged from 306 with Llama3 to 371 with Phi, while the Pro 7 14 AMD reached 345 and 427 with the same processor. The difference also appeared in generation speed, with the Dell Pro 5 16 producing 27.163 tokens per second in Phi and 15.050 tokens per second in Llama3, compared with 30.492 and 16.725 tokens per second from the Pro 7 14 AMD. Faster memory gave the other systems an advantage in these bandwidth-heavy local AI workloads, with Intel’s Arc B390 producing the best results by a wide margin.

UL Procyon: AI Text Generation Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Phi
Phi Overall Score 371 427 689 904
Phi Output Time To First Token 5.052 s 4.286 s 1.843 s 1.158 s
Phi Output Tokens Per Second 27.163 tokens/s 30.492 tokens/s 34.172 tokens/s 36.952 tokens/s
Phi Overall Duration 135.343 s 118 s 92.273 s 81.304 s
Mistral
Mistral Overall Score 346 389 525 716
Mistral Output Time To First Token 6.929 s 6.060 s 3.778 s 2.195 s
Mistral Output Tokens Per Second 18.174 tokens/s 20.125 tokens/s 22.843 tokens/s 24.724 tokens/s
Mistral Overall Duration 196.185 s 176 s 143.831 s 123.773 s
Llama3
Llama3 Overall Score 306 345 509 708
Llama3 Output Time To First Token 6.518 s 5.707s 2.992 s 1.657 s
Llama3 Output Tokens Per Second 15.050 tokens/s 16.725 tokens/s 19.104 tokens/s 20.421 tokens/s
Llama3 Overall Duration 224.176 s 199 s 161.290 s 142.814 s
Llama2
Llama2 Overall Score 329 367 530 641
Llama2 Output Time To First Token 11.089 s 10.307 s 5.412 s 4.082 s
Llama2 Output Tokens Per Second 8.878 tokens/s 10.260 tokens/s 11.230 tokens/s 12.397 tokens/s
Llama2 Overall Duration 381.464 s 334 s 278.109 s 245.572 s

UL Procyon: AI Image Generation

The Procyon AI Image Generation Benchmark provides a consistent and accurate method for measuring AI inference performance across a range of hardware, from low-power NPUs to high-end GPUs. It includes three tests: Stable Diffusion XL (FP16) for high-end GPUs, Stable Diffusion 1.5 (FP16) for moderately powerful GPUs, and Stable Diffusion 1.5 (INT8) for low-power devices. The benchmark uses the optimal inference engine for each system, ensuring fair and comparable results.

The Dell Pro 5 16 stayed close to the Pro 7 14 AMD and Pro 7 14 Intel in the two Stable Diffusion 1.5 tests. It scored 255 in FP16 and generated each image in 24.474 seconds, while its INT8 score of 3,598 was the best of those three laptops by a narrow margin. Stable Diffusion XL was far more demanding, requiring 215.530 seconds per image for a score of 173, which placed it just behind the Pro 7 14 AMD and well behind both Intel systems. The Radeon 890M works reasonably well with Stable Diffusion 1.5, but generation times become lengthy with the larger SDXL model.

UL Procyon: AI Image Generation Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Stable Diffusion 1.5 (FP16)
Stable Diffusion 1.5 (FP16) – Overall Score 255 247 258 635
Stable Diffusion 1.5 (FP16) – Overall Time 391.577 s 404.455 s 386.614 s 157.296 s
Stable Diffusion 1.5 (FP16) – Image Generation Speed 24.474 s/image 25.278 s/image 24.163 s/image 9.831 s/image
Stable Diffusion 1.5 (INT8)
Stable Diffusion 1.5 (INT8) – Overall Score 3,598 3,521 3,575 7,693
Stable Diffusion 1.5 (INT8) – Overall Time 69.478 s 70.985 s 69.911 s 32.495 s
Stable Diffusion 1.5 (INT8) – Image Generation Speed 8.685 s/image 8.873 s/image 8.739 s/image 4.062 s/image
Stable Diffusion XL (FP16)
Stable Diffusion XL (FP16) – Overall Score 173 177 268 646
Stable Diffusion XL (FP16) – Overall Time 3,448.478 s 3,379.388 s 2,230.563 s 928.747 s
Stable Diffusion XL (FP16) – Image Generation Speed 215.530 s/image 211.212 s/image 139.410 s/image 58.047 s/image

PCMark 10

PCMark 10 measures general system performance across everyday work such as web browsing, video conferencing, spreadsheets, writing, photo editing, and rendering. Higher scores are better.

The Dell Pro 5 16 posted an overall PCMark 10 score of 8,268, finishing only 170 points behind the Pro 7 14 Intel and 31 points ahead of the Pro 7 14 AMD. Its Productivity score of 14,322 beat both Intel laptops and fell only 44 points behind the smaller AMD system, while Digital Content Creation reached 9,852 and placed second. This was a strong showing across the office, communication, and creative applications represented in PCMark 10.

PCMark 10 (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Overall 8,268 8,237 8,438 7,945
Essentials 10,870 10,783 10,981 10,751
Productivity 14,322 14,366 13,992 13,821
Digital Content Creation 9,852 9,792 10,610 9,158

Geekbench 6

Geekbench 6 measures processor performance using a mix of common tasks, with separate scores for single-core and multi-core workloads. Higher scores are better.

The Dell Pro 5 16 recorded the highest Geekbench 6 single-core score in the group at 2,989, narrowly beating the Pro 5 14 Intel at 2,968. Its multi-core score of 14,348 placed it last, although it was only about 3% behind the Pro 7 14 AMD at 14,768. Both Intel systems finished above 16,700, giving them a stronger result in Geekbench’s collection of relatively short multi-core workloads.

Geekbench 6 (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Single-Core 2,989 2,888 2,862 2,968
Multi-Core 14,348 14,768 16,787 16,874

Cinebench R23 and 2024

Cinebench measures how quickly the processor can render a complex scene, with separate tests for single-core and multi-core performance. Higher scores are better.

Longer CPU rendering tests allowed the Dell Pro 5 16 to take better advantage of its larger cooling system. It led Cinebench R23 multi-core with 18,764, beating the Pro 5 14 Intel by roughly 11% and the Pro 7 14 AMD by nearly 24%, while its single-core score of 2,029 was only 14 points behind the leader. Cinebench 2024 widened the multi-core gap, with the Dell Pro 5 16 scoring 1,055 compared with 847 for the Pro 7 14 AMD and 807 for the Pro 5 14 Intel. These results show how much better the Ryzen AI 9 HX PRO 470 performs during sustained rendering when installed in the larger 16-inch chassis.

Cinebench (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
R23 Single-Core 2,029 1,946 2,043 2,010
R23 Multi-Core 18,764 15,173 14,640 16,915
2024 Single-Core 119 105 116 122
2024 Multi-Core 1,055 847 683 807

7-Zip Compression

The 7-Zip benchmark measures how quickly the processor can compress and decompress data using multiple threads. Higher GIPS scores are better.

The Dell Pro 5 16 led the 7-Zip benchmark with a total rating of 103.9 GIPS. That placed it roughly 15% ahead of the Pro 7 14 AMD, 16% ahead of the Pro 5 14 Intel, and 28% ahead of the Pro 7 14 Intel. Users regularly compressing or extracting large archives should see a useful reduction in processing time compared with the three smaller laptops.

7-Zip (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Total Rating (GIPS) 103.9 90.2 81.2 89.4

y-cruncher

y-cruncher measures how quickly the processor can calculate large numbers of Pi digits, placing a heavy load on the CPU and memory. Results are measured in seconds, so lower times are better.

The Dell Pro 5 16 completed all three y-cruncher calculations in the shortest time, beginning with 25.160 seconds in the one-billion-digit test. That was nearly identical to the Pro 7 14 AMD at 25.199 seconds, but the larger system opened a wider gap as the workload increased. Its five-billion-digit result of 163.768 seconds was 14 seconds faster than the smaller AMD laptop, 41 seconds faster than the Pro 5 14 Intel, and nearly 77 seconds faster than the Pro 7 14 Intel.

y-cruncher — seconds (lower is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
1 Billion 25.160 25.199 34.775 29.405
2.5 Billion 73.320 79.393 104.586 90.719
5 Billion 163.768 177.776 240.554 204.685

Blender 5.1.1 (GPU)

The Blender benchmark measures GPU rendering performance using three different 3D scenes: Monster, Junkshop, and Classroom. Results are reported in samples per minute, so higher scores are better.

Blender GPU rendering favored Intel graphics, leaving the Dell Pro 5 16 at the bottom of all three scenes. The Radeon 890M produced 129.8 samples per minute in Monster, 103.1 in Junkshop, and 91.1 in Classroom. The Pro 7 14 AMD was between nine and 18% faster, while the Arc B390 in the Pro 5 14 Intel delivered close to three times the performance in Junkshop and Classroom.

Blender 5.1.1 GPU — samples/min (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Monster 129.8 140.9 253.6 366.6
Junkshop 103.1 122.0 189.3 312.6
Classroom 91.1 105.6 153.0 244.4

V-Ray and LuxMark

LuxMark measures GPU compute performance by rendering complex scenes through OpenCL while V-Ray GPU measures how quickly the graphics processor can render a scene using the V-Ray engine. Higher scores are better.

GPU compute performance was stronger than the Blender results, with the Dell Pro 5 16 scoring 2,125 in LuxMark Hall, 982 in LuxMark Food, and 861 vpaths in V-Ray GPU. The Hall result placed it just behind the Pro 7 14 Intel at 2,168, while its V-Ray score was second only to the Pro 5 14 Intel at 919. Although the Radeon 890M did not perform particularly well with Blender’s renderer, it was far more competitive in the OpenCL and V-Ray workloads tested here.

GPU Compute (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
LuxMark — Hall 2,125 2,042 2,168 3,287
LuxMark — Food 982 1,050 880 1,585
V-Ray GPU (vpaths) 861 784 775 919

SPECviewperf 15

SPECviewperf 15 measures graphics performance using viewsets based on professional applications for CAD, 3D modeling, rendering, engineering, and medical visualization. Higher scores are better, although performance can vary considerably between applications and graphics architectures.

The Dell Pro 5 16 led several engineering and scientific viewsets, including Energy (25.39), Medical (65.45), and SolidWorks (33.12), while staying close to the Pro 7 14 AMD in 3ds Max (23.21), CATIA (20.16), Creo (44.27), and SNX (51.67). Results shifted in the visualization tests, where the Arc B390 led Blender, Enscape, Maya, and Unreal Engine, with the Dell Pro 5 16 scoring 19.89, 8.02, 48.67, and 26.61, respectively. The Radeon 890M performed best in many of the CAD, engineering, and medical workloads represented here, while Intel’s Arc B390 had the advantage in several rendering applications.

SPECviewperf 15 (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
3dsmax-08 23.21 24.69 9.58 20.21
blender-01 19.89 20.85 9.11 21.19
catia-07 20.16 21.21 5.27 10.27
creo-04 44.27 45.70 18.27 32.53
energy-04 25.39 25.08 3.68 10.78
enscape-01 8.02 8.29 6.22 14.28
maya-07 48.67 53.64 49.54 82.42
medical-04 65.45 60.50 9.91 22.96
snx-05 51.67 51.84 37.74 46.22
solidworks-08 33.12 32.02 11.86 23.60
unreal_engine-01 26.61 26.29 21.36 38.77

SPECworkstation 4

SPECworkstation 4 measures workstation performance across CPU, graphics, storage, AI, product design, engineering, financial services, and other professional workloads. Higher scores are better, while DNF means the system did not complete every workload required for that category.

The Dell Pro 5 16 recorded the highest scores in AI and Machine Learning at 1.37, Energy at 1.24, and Financial Services at 0.98, while tying the Pro 5 14 Intel in Life Sciences at 1.34. Its Graphics score of 2.36 was just behind the Pro 7 14 AMD at 2.41 and well ahead of both Intel systems, while the CPU subsystem reached 1.09. Storage was weaker at 0.89, and the laptop also trailed both Intel systems in Product Design and Productivity and Development. Media and Entertainment did not finish just like the other AMD system.

SPECworkstation 4 (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
CPU (subsystem) 1.09 1.00 1.08 1.19
Graphics (subsystem) 2.36 2.41 0.82 1.68
Storage (subsystem) 0.89 0.55 1.68 1.76
AI & Machine Learning 1.37 1.29 1.18 1.36
Energy 1.24 1.13 0.90 1.18
Financial Services 0.98 0.88 0.78 0.78
Life Sciences 1.34 1.08 1.04 1.34
Media & Entertainment DNF DNF 1.16 DNF
Product Design 1.34 1.18 1.41 1.64
Productivity & Development 0.78 0.72 1.04 1.10

Storage Performance

3DMark Storage measures how an SSD performs during gaming-related tasks such as loading games, installing software, saving progress, and moving game files. Blackmagic Disk Speed Test measures an SSD’s sequential read and write speeds using large media files.

The 1TB SanDisk PC SN5100S delivered a 3DMark Storage score of 2,477, placing the Dell Pro 5 16 behind the two Intel systems but far ahead of the Pro 7 14 AMD at 894. Blackmagic measured sequential read and write speeds of 4,758.0 MB/s and 5,166.5 MB/s, respectively. Those speeds should be plenty for office work, large file transfers, and creative applications, although the SK hynix drives in the Intel laptops approached or exceeded 8,400 MB/s in both directions.

Storage (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
3DMark Storage (score) 2,477 894 3,259 3,144
Blackmagic Disk — Read (MB/s) 4,758.0 3,103.4 8,398.6 8,609.6
Blackmagic Disk — Write (MB/s) 5,166.5 4,034.3 8,934.5 8,747.3

Battery Life

The PCMark 10 Modern Office battery test repeatedly runs common office tasks until the battery reaches the test’s cutoff point. Longer runtimes are better.

The Dell Pro 5 16 lasted 15 hours and 22 minutes in the PCMark 10 Modern Office battery test, giving it the shortest runtime of the four laptops. The Pro 7 14 AMD ran for four hours and six minutes longer, while both Intel systems passed 26 hours. The 70Wh battery can cover a normal workday, but its runtime falls well short of the three 14-inch models.

Battery — PCMark 10 Modern Office (higher is better) Dell Pro 5 16 (AMD Ryzen AI 9 HX PRO 470 12C)  Dell Pro 7 14 (AMD Ryzen AI 9 HX PRO 470 12C) Dell Pro 7 14 (Intel Core Ultra 7 366H 16C) Dell Pro 5 14 (Intel Core Ultra X7 368H 16C)
Runtime 15h 22m 19h 28m 26h 18m 26h 48m

Conclusion

The Dell Pro 5 16 is the best fit among the tested laptops for buyers who want a larger screen, a dedicated numeric keypad, and stronger performance during long CPU-heavy workloads. Its Ryzen AI 9 HX PRO 470 led both Cinebench multi-core tests, including an 18,764 in R23 that beat every 14-inch system by 11% or more, topped 7-Zip at 103.9 GIPS, and swept all three y-cruncher calculations, showing that the larger cooling system gives the processor room to sustain performance the thinner chassis can’t. It even posted the group’s best Geekbench single-core score at 2,989. The 16-inch WQXGA display also provides a better workspace for spreadsheets, creative applications, and side-by-side windows than the three 14-inch alternatives.

Our review build costs $4,648, making it the least expensive of the four Dell configurations tested despite including 64GB of replaceable memory, a 1TB SSD, Wi-Fi 7, a 70Wh battery, and three years of next-business-day onsite support. That is still a considerable investment, but within a managed Dell fleet, the Pro 5 16 offers the best combination of sustained CPU performance, serviceability, screen size, and price. It works especially well as a desk-replacement system that may travel occasionally but will spend much of its time connected to office equipment.

Anyone who needs the fastest integrated graphics should look elsewhere in the family, since the DDR5-5600 memory limits what the Radeon 890M can do in Blender, local AI, and other bandwidth-heavy workloads. The Pro 7 14 AMD delivers better Radeon performance with its LPDDR5x-8533 memory, while the Arc B390 in the Pro 5 14 Intel is considerably faster in Blender and AI generation. The Dell Pro 5 16 also has the shortest battery life and highest starting weight in this group, so one of the 14-inch models will be a better choice for users who spend most of the day away from a desk.

For buyers who prioritize screen space, a numeric keypad, and sustained CPU performance over graphics speed and portability, the Dell Pro 5 16 is the strongest desk-replacement option in this group.

Dell Pro 5 16 Product Page

The post Dell Pro 5 16 (AMD) Review: The Desk-First AMD Option in Dell’s Pro Laptop Line appeared first on StorageReview.com.

Supermicro Expands DCBBS Liquid Cooling With Ten Rear Door Heat Exchangers Up to 120kW

15 July 2026 at 19:12
Supermicro rear door heat exchanger family Supermicro rear door heat exchanger family

Supermicro has expanded its Data Center Building Block Solutions liquid cooling lineup with a ten-model Rear Door Heat Exchanger (RDHx) portfolio, spanning 10kW to 120kW of cooling capacity at the door level and up to 240kW at the rack level. The expanded portfolio is aimed at data center operators who need a fast, low-disruption path to liquid cooling, whether they are standing up new AI infrastructure or trying to squeeze denser racks into facilities that were never designed for them.

Supermicro rear door heat exchanger family

The doors mount to standard EIA, ORv3, and MGX racks and can serve as the primary liquid cooling solution or run alongside Supermicro’s direct-to-chip (D2C) cold plates as part of a complete DCBBS deployment. Because the doors neutralize exhaust heat at the rack, they also eliminate the need for dedicated hot- and cold-aisle containment, which is a meaningful chunk of most retrofit budgets. In its video summary, Supermicro says the lineup is sized for rack platforms spanning NVIDIA’s GB300 and B300 systems through Vera Rubin and AMD’s upcoming Helios, which explains the 240kW rack-level ceiling. Supermicro offers DC-powered models that tie into rack busbars and AC-powered models for broader infrastructure compatibility, with intelligent fan control, N+1 fan redundancy, and anti-condensation protection across the line. Monitoring runs through Redfish, SNMP, a web interface, or Supermicro’s SuperCloud Composer, and the doors can be deployed without dedicated facility chilled water, which removes one of the classic blockers for retrofits.

As with the rest of DCBBS, the doors can be delivered as part of a validated rack-scale package with accelerated systems, rack integration, facility power and cooling, management software, and deployment services, with the usual pitch of simplified procurement and reduced time-to-online.

Why Rear Doors, and Why They’re Back

Rear door heat exchangers are among the older ideas in data center cooling, and the AI buildout has made them relevant all over again. IBM shipped the first commercial version in 2005 as the “Cool Blue” Rear Door Heat eXchanger, a passive, water-cooled coil that replaced a rack’s rear door and absorbed roughly 15kW, enough to neutralize more than half the heat from the densest racks of that era. The problem it solved then is the same one it solves now: once rack power climbs past what hot-aisle/cold-aisle air distribution can absorb, you either capture the heat at the rack or spend enormous energy chasing it around the room.

The mechanics are simple; server exhaust passes through a liquid-filled coil on its way out of the rack, transferring heat to the facility water loop before the air ever reaches the room. Passive doors rely on the servers’ own fans to push air through the coil; active designs like Supermicro’s add fan arrays to handle deeper coils and higher capacities. Because the heat never enters the data center, the room stays thermally neutral, which means no aisle containment gymnastics and far less load on AC units. Most importantly for operators with existing facilities, none of it requires touching the servers themselves, which is why rear doors have become the standard on-ramp to liquid cooling for retrofits.

Supermicro rear door heat exchanger 2026

In the AI era, rear doors have taken on a second job: finishing what direct-to-chip cooling starts. Cold plates on CPUs and GPUs typically capture 70 to 80% of a server’s heat into liquid, but the remainder, from memory, NICs, drives, and power supplies, still exits as hot air. Pairing D2C with a rear door catches that residual heat too, pushing all of the rack’s thermal output effectively into the facility loop where it can be rejected efficiently or, increasingly, reused. That is the same math Dell was chasing at Dell Technologies World last year with its PowerCool Enclosed Rear Door Heat Exchanger, a sealed-airflow design Dell claims captures 100% of IT heat while cutting cooling energy costs by up to 60%, part of a broader industry push we covered in [Liquid Cooling is Coming to Your Data Center: Supermicro’s approach is more modular than Dell’s enclosed door, but the destination is the same: get every possible watt of heat into water, because water is where heat becomes manageable, and potentially valuable, again.

Availability

The expanded RDHx portfolio is available now as part of Supermicro’s DCBBS offerings, deployable from individual racks to full data center buildouts.

The post Supermicro Expands DCBBS Liquid Cooling With Ten Rear Door Heat Exchangers Up to 120kW appeared first on StorageReview.com.

Samsung PM1763 PCIe Gen6 SSD Enters Mass Production With 28.4 GB/s Reads

8 July 2026 at 18:57

Samsung has started mass production of the PM1763, its first PCIe Gen6 enterprise SSD, pairing 9th-generation V-NAND with a newly developed 4nm controller. The 15.36TB flagship is rated for sequential reads up to 28,400 MB/s and writes up to 21,000 MB/s, with reads landing at 1.96x the 14,500 MB/s ceiling of the Gen5 PM1753 it replaces. Samsung says the PM1763 has completed validation for next-generation AI platforms, and the timing aligns with the first wave of Gen6-capable servers expected in the coming year.

samsung pm1763

Why PCIe Gen6 Matters

PCIe Gen6 doubles the per-lane signaling rate to 64 GT/s using PAM4, which puts roughly 32 GB/s of bandwidth each direction on the standard x4 SSD link, up from about 16 GB/s on Gen5. Gen5 drives have been bumping against that ceiling for a while, with the fastest models rated in the 14-14.5 GB/s range the interface allows.

The PM1763’s 28.4 GB/s rating uses most of the new headroom, and for AI infrastructure the practical effect is that fewer drives are needed to saturate a GPU server’s storage path, while checkpoint and model-load operations spend less time blocking accelerators. Samsung frames it in model terms, claiming a 40GB LLM can move from drive to memory in roughly 1.4 seconds, versus about 2.7 seconds on the PM1753. That 1.4-second figure is the sequential read rating for a 40GB file, not a measured transfer, so treat it as a best-case illustration of the interface math.

Random performance scales alongside the sequential numbers. Samsung’s product materials list up to 6.8 million random read IOPS and 950,000 random write IOPS for the 16TB-class configuration; the SCADA white paper’s measured 6.92 million IOPS per drive (covered below) suggests the read figure is conservative.

Power efficiency improves up to 1.8x over the prior generation, which matters as much as the raw speed in dense deployments, since Gen6 controllers run hot and every watt saved per drive multiplies across a dense chassis. To that point, Samsung has optimized the PM1763 for liquid-cooled servers with direct-to-chip (D2C) cooling, targeting sustained peak performance under extended load rather than burst figures.

samsung pm1763 liquid cooling design

PM1763 Specifications

Specification Samsung PM1763
Platform Overview
Interface PCIe Gen6 x4, NVMe 2.1, OCP 2.6
NAND Samsung 9th-generation V-NAND
Controller New Samsung 4nm controller
Capacities 4TB, 8TB, 16TB-class (15.36TB formatted) at launch
Product page lists 30.72TB and 61.44TB
Form Factors E1.S
E3.S
U.2 (PCIe Gen5 only)
Performance (16TB)
Sequential Read Up to 28,400 MB/s
Sequential Write Up to 21,000 MB/s
Random Read Up to 6,800,000 IOPS
Random Write Up to 950,000 IOPS
Power and Cooling
Power Efficiency Up to 1.8x improvement over PM1753
Cooling Optimized for liquid-cooled servers, direct-to-chip (D2C)
Security
Features Post-quantum cryptography (PQC)
TEE Device Interface Security Protocol (TDISP)

GPU-Driven I/O: 281 Million IOPS Across 42 Drives

Samsung has also published a white paper pairing the PM1763 with SCADA (Scaled Accelerated Data Access), the NVIDIA-developed framework that lets GPU threads submit NVMe commands directly to the drives, bypassing the CPU and kernel storage stack entirely. The argument for GPU-initiated I/O is concurrency: a CPU can keep roughly 45 million IOPS in flight across its thread pool, while a GPU dispatching from around 100,000 threads pushed past 95 million IOPS per GPU in Samsung’s testing.

The numbers come from a 512-byte random read workload on an H3 Falcon 6048 Gen6 server with one H100, two H200s, and 42 PM1763 E1.S 15.36TB drives behind three Broadcom PEX90144 Gen6 switches. A single PM1763 processed roughly 6.92 million GPU-issued IOPS, an 86% gain over the 3.72 million the Gen5 PM1753 managed in the same setup. With 14 drives per GPU, the system held near-linear scaling at about 96 million IOPS per GPU group, and the full 42-drive configuration aggregated 281 million IOPS, with per-drive results staying within 5% of the single-drive peak. We did not conduct this testing, nor did we independently audit the results, but the scaling behavior is interesting: latency consistency across drives, not peak IOPS, determined how well the aggregate held up.

What Gen6 Does for Dense Storage Servers

The question is what a shelf of these drives does inside a single box. We recently pushed the Dell PowerEdge R7725xd past 300 GB/s of local throughput with 24 Gen5 SSDs, each drive on dedicated x4 lanes from the CPU complex, and served 160 GB/s over the network with PEAK:AIO’s software keeping the queues saturated. That system rivals multi-node storage clusters from a single 2U chassis. Swap in Gen6 drives at the PM1763’s rated speeds and the same 24-bay topology carries a theoretical ceiling of roughly 681 GB/s of raw read bandwidth, though CPU lane budgets and network egress become the binding constraints well before the drives do.

Capacity density moves in tandem. PEAK:AIO’s 2U AI Data Server already packs 1.5PB using 61.44TB QLC drives while delivering 120 GB/s over RDMA. When Samsung ships the 61.44TB PM1763, that class of system gets Gen6 bandwidth and petabyte-plus density in a similar footprint (next-gen servers are likely to add a rack unit for cooling). For AI shops trying to keep GPU clusters fed without building out a parallel file system across a dozen nodes, the single-server storage argument keeps getting stronger.

Security and Availability

Samsung has also extended the drive’s security stack for multi-tenant AI environments. The PM1763 supports post-quantum cryptography algorithms ahead of anticipated quantum attacks on classical encryption, as well as TDISP (TEE Device Interface Security Protocol), which secures the data path between confidential VMs and the device in virtualized deployments.

“Built on industry-leading performance, PM1763 has successfully completed validation for next-generation AI platforms and is well positioned to support evolving AI infrastructure requirements,” said Jangseok Choi, Vice President and Head of Memory Product Planning at Samsung Electronics. “As AI models continue to grow in size and complexity, PM1763 will serve as a key solution that enables customers to efficiently scale memory capacity and optimize AI operations.”

The PM1763 is in mass production now in 4TB, 8TB, and 16TB capacities. Samsung has not announced ship dates for the larger capacity points.

Samsung PM1763 Product Page

The post Samsung PM1763 PCIe Gen6 SSD Enters Mass Production With 28.4 GB/s Reads appeared first on StorageReview.com.

Canonical LXD 6.9 Adds Dell PowerStore Driver and Fibre Channel Support

6 July 2026 at 17:05

Canonical has released LXD 6.9 with updates aimed squarely at storage teams: a native driver for Dell PowerStore arrays, a new Fibre Channel connector for remote storage generally, and support for Dell PowerFlex 5. For a platform that started life as a container manager, that is a notable amount of enterprise SAN plumbing in a single release.

For readers who have not tracked it, LXD is Canonical’s open-source virtualization platform that manages both system containers and full KVM virtual machines across clustered hosts. It has gained attention as organizations reassess their hypervisor options in the wake of Broadcom’s VMware licensing changes, and Canonical has been steadily building out the enterprise features (clustering, live migration, a Kubernetes CSI driver, disaster-recovery replication) that a VMware alternative needs. The missing piece for many has been the storage they already have, and this release addresses much of it.

The reason a native driver matters comes down to where instance data lives. Without one, LXD typically puts VM and container volumes on host-local ZFS, LVM, or Btrfs, or on a Ceph cluster, which means an existing array is reduced to serving LUNs that the host then carves up itself. With a native driver, LXD provisions each instance volume directly on the array, so snapshots, clones, and thin provisioning are handled by the array’s own data services, and volumes are reachable from any cluster member. PowerStore now joins Dell PowerFlex, Pure Storage, and HPE Alletra on that list, with both iSCSI and Fibre Channel connectivity supported at launch.

The Fibre Channel connector is arguably the bigger long-term change. Until now, LXD’s remote storage drivers supported only iSCSI or NVMe/TCP, which excluded the large installed base of FC fabrics that dominate legacy SAN estates. The connector is a general transport layer, so drivers beyond PowerStore can adopt it. In related housekeeping, the NVMe/TCP pool mode has been renamed from nvme to nvme/tcp, with existing pools migrated automatically on upgrade.

On the PowerFlex side, the driver now supports PowerFlex 5, including thin clone support, and automatically detects the array’s software version while remaining compatible with PowerFlex 4. The ZFS driver also gains a practical speedup: LXD now caches image variants matching an instance’s configuration, so repeated deployments from the same image no longer rebuild the clone.

Beyond storage, 6.9 adds load balancer pools for OVN networks with health checking, OWASP-compliant security event logging that can be routed to Grafana Loki, and quorum protection for cluster evacuations. The release also lands fixes for eleven CVEs, several of which involved symlink attacks in crafted images or project restriction bypasses.

One caveat: 6.9 is a feature release, which Canonical explicitly does not recommend for production use. Shops that want these capabilities on a supported footing will be waiting for the next LTS. For everyone else, the release is available now via snap install lxd --channel=6/stable, with the snap base moving from core24 to core26.

The full release notes are available on Canonical’s LXD documentation site.

The post Canonical LXD 6.9 Adds Dell PowerStore Driver and Fibre Channel Support appeared first on StorageReview.com.

AMD Ryzen AI Halo Review: A Dual-OS, 200B-Parameter Desktop Takes On the DGX Spark

6 July 2026 at 14:59

AMD silicon arguably got here first: Strix Halo mini PCs and laptops were shipping with 128GB of unified memory well before NVIDIA entered the picture. But the local-AI desktop as a category is one NVIDIA effectively created when it put a Grace Blackwell superchip in a one-liter box and called it DGX Spark. The pitch was simple: a developer-class machine with enough unified memory to hold capable models, sitting on a desk instead of metered in the cloud. The AMD Ryzen AI Halo is AMD’s answer to that machine. AMD announced it alongside the Ryzen AI Max PRO 400 Series in May 2026; pre-orders opened in June exclusively through Micro Center, with in-store availability July 10th. Ryzen AI Halo arrives with a similar footprint, 128GB of unified memory, a $3,999 price, and a short list of decisions that make it a significantly different proposition than the Spark.

AMD Ryzen AI Halo front view.

AMD bills the Halo as its first AI developer platform, giving developers a fast, low-friction path to build and run AI locally. Under the hood is the Ryzen AI Max+ 395, a 16-core, 32-thread “Zen 5” part with Radeon 8060S integrated graphics (40 RDNA 3.5 compute units) and an XDNA 2 NPU rated at 50 TOPS, all within a platform AMD markets at up to 126 TOPS of combined AI throughput. The 128GB of LPDDR5x runs at 8000 MT/s, delivering 256 GB/s of bandwidth, and the whole platform draws power from a single USB-C input rated at 120W. AMD says the memory pool is sufficient to hold models with up to 200 billion parameters in device memory. It is a complete x86 mini-workstation, which is the root of the difference that matters most.

That difference is Windows. The Spark runs NVIDIA’s Linux-based DGX OS and nothing else; the Halo boots Windows 11 or AMD’s Linux developer image on the same hardware. Native Windows support was the single most common request we heard from people eyeing a Spark, and it reshapes who the box is for. AMD’s own positioning makes the same point: Windows and Linux versus Spark’s Linux-only (as of today, anyway).

Three more decisions separate the Halo from the Spark, and each addresses a complaint we have heard about the incumbent. The Halo uses a standard M.2 2280 SSD rather than the less common 2242 drive the Spark fits, which opens up a much larger pool of aftermarket options, including 8TB capacities, for anyone who wants to replace the drive. It ships with Variable Graphics Memory pre-set to its maximum allocation on both operating systems, so large models load without manual tuning. It also wraps the chassis in a lit status ring, a small correction to the dark, lightless Spark units some buyers received, depending on the OEM. The cost of AMD’s approach shows up at the back of the box, where there is no high-speed fabric, only 10GbE, which caps what you can do with multi-node clustering. Those are the trade-offs, and they define the workloads this machine is built for before any benchmark runs.

AMD Ryzen AI Halo disassembled.

On price, the nuance matters. The Halo lists at $3,999 with a 2TB SSD, placing it right at the going rate for a base Spark-class system. The comparison depends on which Spark you mean. NVIDIA’s own DGX Spark with the larger drive now runs closer to $4,700, a move tied to the LPDDR5X and NAND supply crunch. Base Grace Blackwell systems from ASUS and others still sell around $4,000 and are listed on Amazon today (affiliate link). Measured against the category baseline, the Halo is at parity, and its case rests on the decisions above rather than on undercutting the field.

We ran the Halo through StorageReview’s local AI suite on both Windows and Linux, and the results are presented alongside the design and software analysis throughout this review.

Key Takeaways

  • The dual-OS x86 alternative: Ryzen AI Halo is the only box in the Spark’s category that boots Windows 11 or Linux on a full x86 platform, at $3,999 with a 2TB SSD against the DGX Spark Founders Edition’s $4,699.
  • Memory to hold big models: 128GB of LPDDR5x-8000 unified memory at 256GB/s supports models up to 200 billion parameters locally, with Variable Graphics Memory pre-tuned so large models load without manual configuration.
  • The strongest Ryzen AI Max+ 395 we’ve tested: The Halo topped the HP Z2 Mini G1a and ZBook Ultra G1a in nearly every Windows workload, including 37,316 in Cinebench R23 multi-core, 184.2 GIPS in 7-Zip, and the leading Procyon AI text generation (Phi 1,192) and image generation (SD 1.5 FP16 937) scores.
  • CPU wins, inference losses vs. Spark: On Linux the Halo beat the DGX Spark outright in CPU work, compressing 11% faster and decompressing 38% faster in 7-Zip and finishing the LLVM compile 14% sooner, but trailed 2x to 4x in most vLLM serving scenarios at higher concurrency, stretching to 8.8x in prefill-heavy GPT OSS 120B work.
  • Serviceable storage, modest networking: A standard M.2 2280 bay opens aftermarket upgrades to 8TB, though the platform negotiates the included Gen5 Micron 4600 down to Gen4 by design, and the single 10GbE port with no high-speed fabric rules out the multi-node clustering the Spark’s 200G ConnectX-7 enables.

Specifications

Specification AMD Ryzen AI Halo System
Processor
CPU AMD Ryzen™ AI Max+ 395 Processor
16 Cores / 32 Threads (Zen 5 Architecture)
GPU AMD Radeon™ 8060S Integrated Graphics
40 Compute Units (RDNA™ 3.5 Architecture)
NPU AMD XDNA™ 2 NPU
Memory
Memory Type LPDDR5x
Memory Capacity 128GB
Memory Speed 8000MT/s
Memory Bandwidth 256GB/s
Storage
Storage 2TB M.2 NVMe SSD (SED)
Networking & Connectivity
Ethernet 1 × 10GbE
Wi-Fi Wi-Fi 7
Bluetooth Bluetooth 5.4
I/O
USB 3 × USB-C, 1 × USB-C (Power Input)
Display Output 1 × HDMI 2.1b
System
TDP 120W
Operating System Linux or Windows 11
Dimensions 150 × 150 × 45.4 mm (5.9 × 5.9 × 1.79 in)
Weight Less than 1.2 kg (2.65 lbs)

AMD Ryzen AI Halo Build and Design

The AMD Ryzen AI Halo system uses a compact, “NVIDIA Spark”-like form factor, measuring just 150 × 150 × 45.4mm and weighing less than 1.2kg (2.65lbs). The aluminum chassis features an aggressive geometric ventilation pattern across the top and front, giving the system a distinctive appearance while maximizing airflow into the cooling system. Despite its small footprint, the platform is designed as a full desktop AI workstation capable of handling workstation applications, local LLM inference, and AI development workloads.

Front

The front of the system is intentionally clean, consisting almost entirely of a large mesh ventilation grille that spans the chassis width. Rather than placing ports on the front, AMD dedicates this area to airflow, allowing cool air to enter the system through the large patterned intake while keeping the front uncluttered. A silver accent along the bottom of the chassis provides subtle visual contrast to the otherwise matte black enclosure.

Rear I/O

AMD Ryzen AI Halo rear connectivity.

All external connectivity is located on the rear of the system. From left to right, the rear panel includes the following:

  • USB-C power input
  • USB-C port with DisplayPort Alt Mode
  • Two additional USB-C ports
  • HDMI 2.1b output
  • 10GbE RJ45 Ethernet
  • Kensington security lock slot

The system provides three USB-C data ports alongside a dedicated USB-C power connector, enabling multiple high-speed peripherals and displays to connect simultaneously. HDMI 2.1b provides native display output, while the integrated 10GbE Ethernet interface makes the platform well suited for high-speed NAS connectivity, AI dataset transfers, and local development environments.

Internal Design

The underside of the Ryzen AI Halo features a large perforated vent that draws in air to cool the components mounted along the bottom of the board, and four rubber feet at the corners keep the unit stable on a desk or shelf.

AMD Ryzen AI Halo bottom.

Removing the four screws securing the bottom panel reveals the M.2 SSD bay and a few of the system’s lower-board components, including the Wi-Fi card connector. In our review unit, that slot holds a Micron 4600 2TB Gen5 x4 SSD, while a black adhesive sheet shields the rest of the board from the exposed underside.

AMD Ryzen AI Halo bottom lid removed.

For cooling, AMD took a similar approach to NVIDIA’s DGX Spark, using dual fans that pull air across a finned heatsink and exhaust it out the rear of the chassis.

AMD Ryzen AI Halo heatsink and fans.

With the cooling assembly lifted away, the main board comes into view, revealing the APU die at the center, flanked by memory packages on either side and VRM circuitry running down the left edge. The silver square visible on the die isn’t liquid metal or a paste-based compound; it’s the residue of a solid thermal pad or coating AMD applied as the die-to-heatsink interface, which explains its uniform, dry appearance rather than the wet, smeared look paste or liquid metal typically leaves behind.

AMD Ryzen AI Halo main board.

Flipping the heatsink over reveals the underside of its cold plate, where a mirror-polished section makes direct contact with the die. At the same time, the surrounding memory and power-delivery zones are covered with pre-applied thermal pads of varying thickness.

AMD Ryzen AI Halo cooler underside view.

AMD Ryzen AI Halo Performance Testing

We evaluated the AMD Ryzen AI Halo platform on Windows and Linux to assess its performance in workstation applications and AI-focused workloads. For Windows testing, the AMD Ryzen AI Halo system was compared with two commercially available systems powered by the Ryzen AI Max+ PRO 395: the HP Z2 Mini G1a, a compact desktop workstation, and the HP ZBook Ultra G1a 14-inch, a mobile workstation. Because all three systems share the same underlying processor architecture and Radeon 8060S integrated graphics, these comparisons highlight how the Halo platform performs across different thermal envelopes and system designs.

AMD Ryzen AI Halo top cover off rear view.

For Linux testing, we shifted to AI development and storage workloads, comparing the Ryzen AI Halo system with the NVIDIA DGX Spark. These tests focused on FIO storage benchmarking and vLLM inference performance, comparing AMD’s Ryzen AI Halo-based developer platform with NVIDIA’s purpose-built AI development system for local large-language-model inference.

Tested Units

UL Procyon: AI Computer Vision

The Procyon AI Computer Vision Benchmark provides detailed insights into how AI inference engines perform at a professional level. By incorporating engines from multiple vendors, it delivers performance scores that accurately reflect a device’s capabilities. The benchmark evaluates state-of-the-art neural network models by comparing their AI acceleration performance across hardware types—including CPU, GPU, and NPU—enabling users to assess relative efficiency across a range of workloads and conditions.

To reflect real-world AI workloads, the benchmark uses six diverse neural network models, each selected for its relevance to modern computer vision tasks. MobileNet V3 is a compact, mobile-focused model designed for subject identification in images, whereas Inception V4 performs the same task with a deeper, more complex architecture.

YOLO V3 (You Only Look Once) specializes in real-time object detection by estimating object probabilities. DeepLab V3, built on MobileNet V2, focuses on semantic image segmentation and pixel clustering. Real-ESRGAN, the most computationally demanding test, upscales images from 250×250 to 1,000×1,000 resolution. Finally, ResNet 50 is a robust classification model that enables more effective training of deep neural networks.

The Ryzen AI Halo platform delivered consistently strong AI inference performance across CPU and GPU workloads. On the CPU tests, it posted an overall score of 216, narrowly trailing the HP Z2 Mini’s 227 and comfortably outperforming the HP ZBook Ultra’s 186. GPU inference was even more impressive, with the Halo system earning the highest overall score at 553, ahead of the ZBook Ultra (528) and the Z2 Mini (528). It also recorded the fastest MobileNet V3 inference at 0.38ms and completed the demanding REAL-ESRGAN workload in 185.08ms, compared with 211.76ms on the Z2 Mini and 200.40ms on the ZBook Ultra.

UL Procyon: AI Computer Vision Inference (Lower is better) AMD Ryzen AI Halo (Ryzen AI Max+ 395 | Radeon 8060S) HP Z2 Mini G1a (Ryzen AI Max+ PRO 395 | Radeon 8060S) HP ZBook Ultra G1a 14″ (Ryzen AI Max+ PRO 395 | Radeon 8060S)
CPU Times
AI Computer Vision Overall Score (higher is better) 216 227 186
MobileNet V3 0.73 ms 0.75 ms 1.09 ms
ResNet 50 6.02 ms 5.99 ms 6.84 ms
Inception V4 17.18 ms 17.12 ms 19.80 ms
DeepLab V3 29.21 ms 21.20 ms 28.27 ms
YOLO V3 35.60 ms 36.58 ms 41.64 ms
REAL-ESRGAN 1,931.13 ms 1,892.10 ms 2,138.97 ms
GPU Times
AI Computer Vision Overall Score (higher is better) 553 528 583
MobileNet V3 0.38 ms 0.42 ms 0.46 ms
ResNet 50 4.04 ms 3.85 ms 3.27 ms
Inception V4 13.26 ms 15.15 ms 11.62 ms
DeepLab V3 12.72 ms 10.98 ms 10.72 ms
YOLO V3 11.17 ms 12.64 ms 10.57 ms
REAL-ESRGAN 185.08 ms 211.76 ms 200.40 ms

UL Procyon: AI Text Generation

The Procyon AI Text Generation Benchmark streamlines AI LLM performance testing by providing a concise, consistent evaluation method. It enables repeated testing across multiple LLM models while minimizing the complexity of large model sizes and variable factors. Developed with AI hardware leaders, it optimizes the use of local AI accelerators for more reliable and efficient performance assessments. The results below were measured using TensorRT.

The Ryzen AI Halo reference platform led all tested language models in Procyon AI Text Generation. It achieved the highest overall Phi score at 1,192, compared with 965 for the HP Z2 Mini and 922 for the HP ZBook Ultra. Mistral followed a similar trend with a score of 998, ahead of 850 and 829, while Llama3 finished at 847, outperforming the competing systems at 766 and 756, respectively. The Halo platform also reduced time-to-first-token across every model, producing the first Phi token in just 0.996 seconds, nearly half the latency of the HP systems. Although token generation throughput occasionally favored the Z2 Mini, the Halo platform’s significantly lower startup latency resulted in the best overall benchmark scores.

UL Procyon: AI Text Generation AMD Ryzen AI Halo
(Ryzen AI Max+ 395 | Radeon 8060S)
HP Z2 Mini G1a (Ryzen AI Max+ PRO 395 | Radeon 8060S) HP ZBook Ultra G1a 14″ (Ryzen AI Max+ PRO 395 | Radeon 8060S)
Phi Overall Score 1,192 965 922
Phi Output Time To First Token 0.996 seconds 1.898 seconds 1.956 seconds
Phi Output Tokens Per Second 55.271 tokens/s 68.967 tokens/s 64.986 tokens/s
Phi Overall Duration 56.237 seconds 52.666 seconds 55.501 seconds
Mistral Overall Score 998 850 829
Mistral Output Time To First Token 1.603 seconds 2.734 seconds 2.783 seconds
Mistral Output Tokens Per Second 35.027 tokens/s 43.358 tokens/s 41.992 tokens/s
Mistral Overall Duration 88.582 seconds 81.716 seconds 84.065 seconds
Llama3 Overall Score 847 766 756
Llama3 Output Time To First Token 1.963 seconds 2.545 seconds 2.578 seconds
Llama3 Output Tokens Per Second 34.630 tokens/s 36.752 tokens/s 36.243 tokens/s
Llama3 Overall Duration 92.026 seconds 91.987 seconds 93.200 seconds
Llama2 Overall Score N/A 936 929
Llama2 Output Time To First Token N/A seconds 3.813 seconds 3.860 seconds
Llama2 Output Tokens Per Second N/A tokens/s 24.685 tokens/s 24.619 tokens/s
Llama2 Overall Duration N/A seconds 136.077 seconds 136.720 seconds

UL Procyon: AI Image Generation

The Procyon AI Image Generation Benchmark offers a consistent, accurate way to measure AI inference performance across hardware ranging from low-power NPUs to high-end GPUs. It includes three tests: Stable Diffusion XL (FP16) for high-end GPUs, Stable Diffusion 1.5 (FP16) for moderately powerful GPUs, and Stable Diffusion 1.5 (INT8) for low-power devices. The benchmark uses the optimal inference engine for each system, ensuring fair and comparable results.

Image generation proved to be one of Ryzen AI Halo’s strongest workloads. On Stable Diffusion 1.5 FP16, the Halo platform achieved an overall score of 937, compared with 725 for the Z2 Mini and 648 for the ZBook Ultra, while reducing generation time to 106.7 seconds, compared with 137.8 and 154.2 seconds, respectively. Stable Diffusion XL showed an equally strong lead, with the Halo system finishing in 878.5 seconds, approximately 174 seconds faster than the Z2 Mini and more than 450 seconds faster than the ZBook Ultra. The platform also completed the Stable Diffusion 1.5 INT8 benchmark with an overall score of 9,158, a workload unavailable on either HP comparison system.

UL Procyon: AI Image Generation AMD Ryzen AI Halo
(Ryzen AI Max+ 395 | Radeon 8060S)
HP Z2 Mini G1a (Ryzen AI Max+ PRO 395 | Radeon 8060S) HP ZBook Ultra G1a 14″ (Ryzen AI Max+ PRO 395 | Radeon 8060S)
Stable Diffusion 1.5 (FP16) – Overall Score 937 725 648
Stable Diffusion 1.5 (FP16) – Overall Time 106.7 seconds 137.815 seconds 154.203 seconds
Stable Diffusion 1.5 (FP16) – Image Generation Speed 6.670 s/image 8.613 s/image 9.638 s/image
Stable Diffusion 1.5 (INT8) – Overall Score 9,158 N/A N/A
Stable Diffusion 1.5 (INT8) – Overall Time 27.298 seconds N/A N/A
Stable Diffusion 1.5 (INT8) – Image Generation Speed 3.412 s/image N/A N/A
Stable Diffusion XL (FP16) – Overall Score 682 570 451
Stable Diffusion XL (FP16) – Overall Time 878.493 seconds 1,052.468 seconds 1,329.592 seconds
Stable Diffusion XL (FP16) – Image Generation Speed 54.906 s/image 65.779 s/image 83.100 s/image

SPECworkstation 4

The SPECworkstation 4.0 benchmark is a comprehensive tool for evaluating all key aspects of workstation performance. It provides a real-world measure of CPU, graphics, accelerator, and disk performance, giving professionals the data needed to make informed decisions about their hardware investments. The benchmark includes a dedicated set of tests focused on AI and ML workloads, such as data science tasks and ONNX Runtime-based inference tests, reflecting the growing importance of AI/ML in workstation environments. It covers seven industry verticals and four hardware subsystems, providing a detailed and relevant measure of today’s workstations’ performance.

SPECworkstation 4 highlighted the balanced workstation capabilities of the Ryzen AI Halo platform. It posted the highest scores in Financial Services (2.92), Media & Entertainment (2.97), Product Design (2.22), and Productivity & Development (1.27), outperforming both the HP Z2 Mini and HP ZBook Ultra in those categories. The Z2 Mini held a slight lead in Energy (2.50 vs. 2.35) and Life Sciences (2.60 vs. 2.33). Overall, the Halo reference platform demonstrated strong performance across the benchmark’s professional workloads and remained competitive in every category tested.

SPECworkstation 4.0.0 (Higher is better) AMD Ryzen AI Halo
(Ryzen AI Max+ 395 | Radeon 8060S)
 

HP Z2 Mini G1a (Ryzen AI Max+ PRO 395 | Radeon 8060S)

 

HP ZBook Ultra G1a 14″ (Ryzen AI Max+ PRO 395 | Radeon 8060S)
Energy 2.35 2.50 2.20
Financial Services 2.92 2.35 1.60
Life Sciences 2.33 2.60 2.20
Media & Entertainment 2.97 2.22 1.90
Product Design 2.22 2.00 1.74
Productivity & Development 1.27 1.00 1.03

Luxmark

Luxmark is a GPU benchmark that uses LuxRender, an open-source ray-tracing renderer, to evaluate a system’s performance with highly detailed 3D scenes. This benchmark is useful for assessing the graphical rendering capabilities of servers and workstations, especially for visual effects and architectural visualization applications, where accurate light simulation is crucial.

Luxmark results showed minimal separation among the three Ryzen AI Max+ 395 platforms. The Halo reference system posted the highest Food score at 4,158, edging out the Z2 Mini (3,943) and ZBook Ultra (3,915). In the Hallbench workload, it scored 8,014, slightly behind the Z2 Mini’s 8,477 but ahead of the ZBook Ultra’s 7,833. Overall, the results suggest that systems built around the Radeon 8060S deliver very similar ray-tracing performance regardless of form factor.

Luxmark (Higher is better) AMD Ryzen AI Halo
(Ryzen AI Max+ 395 | Radeon 8060S)
HP Z2 Mini G1a (Ryzen AI Max+ PRO 395 | Radeon 8060S) HP ZBook Ultra G1a 14″ (Ryzen AI Max+ PRO 395 | Radeon 8060S)
Hallbench 8,014 8,477 7,833
Food 4,158 3,943 3,915

7-Zip Compression

The 7-Zip Compression Benchmark evaluates CPU performance during compression and decompression, measuring performance in GIPS (Giga Instructions Per Second) and CPU usage. Higher GIPS and efficient CPU usage indicate superior performance.

The Ryzen AI Halo platform led every major category in the 7-Zip benchmark. It achieved a compression rating of 176.7 GIPS, compared to 139.3 GIPS on the HP Z2 Mini and 139.6 GIPS on the ZBook Ultra. Decompression performance remained equally strong at 191.6 GIPS, exceeding the Z2 Mini’s 164.0 GIPS and the ZBook Ultra’s 174.0 GIPS. Combined, the Halo platform achieved the highest overall rating of 184.2 GIPS, outperforming competing systems by roughly 20% and demonstrating excellent integer throughput for archive creation and extraction workloads.

7-Zip Compression Benchmark (Higher is Better) AMD Ryzen AI Halo
(Ryzen AI Max+ 395 | Radeon 8060S)
HP Z2 Mini G1a (Ryzen AI Max+ PRO 395 | Radeon 8060S) HP ZBook Ultra G1a 14″ (Ryzen AI Max+ PRO 395 | Radeon 8060S)
Compressing
Current CPU Usage 2,741% 2,734% 2,868%
Current Rating/Usage 6.370 GIPS 5.136 GIPS 4.883 GIPS
Current Rating 174.589 GIPS 140.405 GIPS 140.061 GIPS
Resulting CPU Usage 2,751% 2,718% 2,855%
Resulting Rating/Usage 6.424 GIPS 5.126 GIPS 4.890 GIPS
Resulting Rating 176.742 GIPS 139.298 GIPS 139.617 GIPS
Decompressing
Current CPU Usage 2,568% 2,343% 2,904%
Current Rating/Usage 7.670 GIPS 6.805 GIPS 6.029 GIPS
Current Rating 196.986 GIPS 159.451 GIPS 175.104 GIPS
Resulting CPU Usage 2,469% 2,414% 2,887%
Resulting Rating/Usage 7.766 GIPS 6.793 GIPS 6.028 GIPS
Resulting Rating 191.645 GIPS 163.969 GIPS 174.046 GIPS
Total Rating
Total CPU Usage 2,610% 2,566% 2,871%
Total Rating/Usage 7.095 GIPS 5.959 GIPS 5.459 GIPS
Total Rating 184.194 GIPS 151.634 GIPS 156.832 GIPS

Blender Benchmark

Blender is an open-source 3D modeling application. This benchmark was run with the Blender Benchmark utility. The score is measured in samples per minute, with higher values indicating better performance.

CPU rendering was another area where Ryzen AI Halo excelled. In the Monster scene, it achieved 244.7 samples per minute, ahead of the HP Z2 Mini (224.3) and HP ZBook Ultra (189.3). Junkshop followed with 159.4 samples per minute, compared with 149.5 and 129.4, while Classroom completed at 131.6 samples per minute, outperforming the Z2 Mini (116.3) by roughly 13% and the ZBook Ultra (94.1) by nearly 40%. These results demonstrate that the Ryzen AI Max+ 395 delivers excellent multithreaded rendering performance despite its compact workstation footprint.

Blender Benchmark CPU (Samples per minute, Higher is better) AMD Ryzen AI Halo
(Ryzen AI Max+ 395 | Radeon 8060S)
HP Z2 Mini G1a (Ryzen AI Max+ PRO 395 | Radeon 8060S) HP ZBook Ultra G1a 14″ (Ryzen AI Max+ PRO 395 | Radeon 8060S)
Monster 244.7 samples/m 224.3 samples/m 189.29 samples/m
Junkshop 159.4 samples/m 149.5 samples/m 129.42 samples/m
Classroom 131.6 samples/m 116.3 samples/m 94.14 samples/m

GPU rendering results were much closer across systems. The Halo reference platform rendered the Monster scene at 704.2 samples per minute, trailing the Z2 Mini (745.6) but ahead of the ZBook Ultra (661.5). Junkshop was effectively tied between the Halo platform (366.6) and the Z2 Mini (366.5). At the same time, the Halo system posted the highest Classroom score at 361.5 samples per minute, narrowly exceeding the Z2 Mini (359.0) and the ZBook Ultra (333.3). The results indicate that the Radeon 8060S delivers remarkably consistent GPU rendering performance across implementations.

Blender Benchmark GPU (Samples per minute, Higher is better) AMD Ryzen AI Halo
(Ryzen AI Max+ 395 | Radeon 8060S)
HP Z2 Mini G1a (Ryzen AI Max+ PRO 395 | Radeon 8060S) HP ZBook Ultra G1a 14″ (Ryzen AI Max+ PRO 395 | Radeon 8060S)
Monster 704.2 samples/m 745.55 samples/m 661.50 samples/m
Junkshop 366.6 samples/m 366.54 samples/m 341.92 samples/m
Classroom 361.51 samples/m 359.01 samples/m 333.26 samples/m

y-cruncher

y-cruncher is a multithreaded, scalable program capable of computing Pi and other mathematical constants to trillions of digits. Since its launch in 2009, it has become a popular benchmarking and stress-testing tool for overclockers and hardware enthusiasts.

The y-cruncher results split along computation size. At the 1-billion- and 2.5-billion-digit runs, all three systems finished within a few tenths of a second of one another, with the HP systems fractionally ahead. As the workload scaled, the Halo pulled away, completing the 5-billion-digit computation in 71.948 seconds, compared with 75.021 seconds for the Z2 Mini and 78.19 seconds for the ZBook Ultra. At 10 billion digits, the Halo finished in 151.409 seconds, roughly 6% ahead of the Z2 Mini and 12% ahead of the ZBook Ultra, suggesting the desktop chassis sustains heavy multithreaded load better as run times stretch out.

Y-Cruncher (Total Computation Time) AMD Ryzen AI Halo
(Ryzen AI Max+ 395 | Radeon 8060S)
HP Z2 Mini G1a (Ryzen AI Max+ PRO 395 | Radeon 8060S) HP ZBook Ultra G1a 14″ (Ryzen AI Max+ PRO 395 | Radeon 8060S)
1 Billion 13.193 seconds 12.965 seconds 12.93 seconds
2.5 Billion 34.578 seconds 34.533 seconds 34.91 seconds
5 Billion 71.948 seconds 75.021 seconds 78.19 seconds
10 Billion 151.409 seconds 160.252 seconds 171.72 seconds

Geekbench 6

Geekbench 6 is a cross-platform benchmark measuring overall system performance.

Geekbench 6 reinforced the Halo platform’s balanced performance profile. It achieved the highest single-core score of 2,986, ahead of the HP Z2 Mini (2,862) and ZBook Ultra (2,825). Multi-core performance also led the comparison with 18,068. The Radeon 8060S recorded the highest OpenCL GPU score at 92,883, narrowly exceeding the Z2 Mini (91,591) and comfortably outperforming the ZBook Ultra (85,337). The consistent leads across CPU and GPU testing illustrate the well-rounded performance of the Ryzen AI Max+ 395 platform.

Geekbench 6 (Higher is better) AMD Ryzen AI Halo
(Ryzen AI Max+ 395 | Radeon 8060S)
HP Z2 Mini G1a (Ryzen AI Max+ PRO 395 | Radeon 8060S) HP ZBook Ultra G1a 14″ (Ryzen AI Max+ PRO 395 | Radeon 8060S)
CPU Single-Core 2,986 2,862 2,825
CPU Multi-Core 18,068 17,210 17,562
GPU OpenCL 92,883 91,591 85,337

Cinebench R23

Cinebench R23 is a widely recognized benchmark for evaluating CPU performance in 3D rendering workloads. Powered by the Cinema 4D engine, it measures how well a processor handles single-threaded and multithreaded tasks, offering insight into overall responsiveness and parallel processing capabilities.

Cinebench R23 showed a very close race between the two Ryzen AI Max+ 395 desktop implementations. The Halo reference platform posted a multi-core score of 37,316, edging out the HP Z2 Mini’s 37,156, while both comfortably surpassed the HP ZBook Ultra’s 29,112. Single-core performance followed the same pattern, with the Halo platform scoring 2,047, compared to 2,020 on the Z2 Mini and 1,984 on the ZBook Ultra. Although the margins over the Z2 Mini were small, the Halo system consistently finished at the top of the benchmark.

Cinebench R23 (Higher is better) AMD Ryzen AI Halo
(Ryzen AI Max+ 395 | Radeon 8060S)
HP Z2 Mini G1a (Ryzen AI Max+ PRO 395 | Radeon 8060S) HP ZBook Ultra G1a 14″ (Ryzen AI Max+ PRO 395 | Radeon 8060S)
Multi-Core 37,316 37,156 29,112
Single-Core 2,047 2,020 1,984

Cinebench 2024

Cinebench 2024 builds on the foundation of R23 by introducing GPU-based rendering tests while maintaining its focus on CPU performance. For this segment, we examine only the CPU scores, which offer updated insight into how well each system handles modern 3D rendering tasks.

The newer Cinebench 2024 benchmark mirrored the R23 results. Ryzen AI Halo recorded the highest multi-core score at 1,916, narrowly ahead of the HP Z2 Mini (1,906) and maintaining a sizable lead over the HP ZBook Ultra (1,579). Single-core performance also favored the Halo platform, with 116 points compared with 112 for the Z2 Mini and 111 for the ZBook Ultra. While the differences between the desktop systems remained small, the results reinforce the Ryzen AI Max+ 395’s ability to deliver consistently top-tier CPU rendering performance.

Cinebench 2024 (Higher is better) AMD Ryzen AI Halo
(Ryzen AI Max+ 395 | Radeon 8060S)
HP Z2 Mini G1a (Ryzen AI Max+ PRO 395 | Radeon 8060S) HP ZBook Ultra G1a 14″ (Ryzen AI Max+ PRO 395 | Radeon 8060S)
Multi-Core 1,916 1,906 1,579
Single-Core 116 112 111

Phoronix Benchmarks

Phoronix Test Suite is an open-source, automated benchmarking platform that supports over 450 test profiles and more than 100 test suites via OpenBenchmarking.org. It handles everything from installing dependencies to running tests and collecting results, making it ideal for performance comparisons, hardware validation, and continuous integration. We will look at performance in Stream, 7-Zip, and LLVM tests.

Looking at the Phoronix benchmark suite, we observed a fairly even split between the AMD Ryzen AI Halo platform and NVIDIA DGX Spark. The Ryzen AI Halo system consistently led CPU-centric workloads, outperforming the DGX Spark by 11% in 7-Zip compression, 38% in 7-Zip decompression, and completing the LLVM compile test 14% faster. The DGX Spark, on the other hand, showed stronger sustained memory bandwidth, leading the STREAM Scale, Triad, and Add tests by 17%, 13%, and 15%, respectively. Halo retained an 18% advantage in the STREAM Copy benchmark, illustrating that while both systems are highly capable, the Ryzen AI Halo platform favors general-purpose compute performance, whereas the DGX Spark demonstrates higher throughput in memory-bandwidth-focused workloads.

Test AMD Ryzen AI Halo NVIDIA DGX Spark Winner
7-Zip Compression (MIPS) 186,921 169,052 Halo (+11%)
7-Zip Decompression (MIPS) 146,109 106,084 Halo (+38%)
STREAM Copy (MB/s) 145,363 123,730 Halo (+18%)
STREAM Scale (MB/s) 110,611 128,970 Spark (+17%)
STREAM Triad (MB/s) 107,105 121,433 Spark (+13%)
STREAM Add (MB/s) 107,022 122,815 Spark (+15%)
LLVM Compile (Make, Seconds) 431.8 504.3 Halo (14% Faster)

FIO Performance Benchmark

To measure storage performance across common industry metrics, we use fio. Our traditional SSD test preconditions each drive with two full drive fills as a secondary drive; here, we tested each drive in-system, through the filesystem. We previously tested the NVIDIA systems with GDSIO, but AMD doesn’t offer a comparable GPU-direct storage test, so running filesystem-level fio on both platforms puts them on a level playing field, which matters more than benchmarking the drives discretely.

Both the NVIDIA DGX Spark (Founders Edition) and the AMD Ryzen AI Halo ship with a PCIe Gen5 drive, so on paper the two platforms have identical ceilings for raw drive bandwidth. In practice, though, the Halo’s drive runs at PCIe Gen4 link speeds in this system, which caps the available bandwidth well below what the drive itself is capable of. That link-speed limitation is worth keeping in mind throughout this section, as it accounts for a meaningful portion of the bandwidth gap between the two platforms.

In this section, we focus on the following FIO benchmarks:

  • 128K Sequential
  • 64K Random
  • 16K Sequential
  • 16K Random
  • 4K Random

128K Sequential Read

At a single-threaded, deep-queue 128K sequential read (64/1), the NVIDIA DGX Spark posted 13,399.6 MB/s, nearly double the AMD Ryzen AI Halo’s 6,897.5 MB/s. Latency followed the same pattern: the Spark’s mean latency at this depth was 0.597 ms, while the Halo trailed at 1.159 ms, almost twice as slow per request despite moving half the data. This is the clearest gap across the entire sweep, suggesting the Spark’s storage stack is built for higher sustained sequential throughput at this block size.

128K Sequential Write

The write side tells a similar story. At IODepth 16/1, Spark reached 2,984.4 MB/s, compared with the Halo’s 1,392.4 MB/s, a 114% advantage for NVIDIA. Latency again favored Spark, at 0.67 ms versus the Halo’s 1.436 ms. Combined with the read results, 128K sequential is comfortably Spark’s strongest showing relative to the Halo.

64K Random Read

This is where the Halo’s low-queue-depth latency advantage really shows. At 1/1, the Halo answered in just 0.051 ms, versus Spark’s 0.275 ms, over 5x faster for a single outstanding request. That gap holds across low-to-mid queue depths, with the Halo tracking well under 0.2 ms while the Spark hovers in the 0.2–0.45 ms range for most of the sweep.

The story flips at scale, though. As queue depth and thread count climb past roughly 16/4, the Spark’s bandwidth breaks away and plateaus around 9.8 GB/s from 8/16 onward, topping out at 10,019.1 MB/s (160.3K IOPS) at 32/8. The Halo never reaches that ceiling: it saturates in the 6.0–6.9 GB/s range, peaking at 6,926.5 MB/s (110.8K IOPS) at 4/8. Its bandwidth curve is noticeably more erratic, sawtoothing between roughly 2 and 6.8 GB/s depending on the depth/thread combination rather than climbing smoothly.

Latency at full saturation also swings in Spark’s favor: at 32/16, the Halo’s mean latency balloons to 5.185 ms, compared with Spark’s 3.204 ms, meaning the Halo pays for its low-queue-depth responsiveness with worse tail behavior once the queue is fully loaded.

64K Random Write

Random 64K write bandwidth is closer between the two platforms than read bandwidth, though the shapes of the curves are very different. The Halo’s bandwidth is spikier: it repeatedly jumps above 2.5 GB/s (peaking at 2,929.1 MB/s / 46.9K IOPS at 32/4) before dropping back to the 1.1–1.5 GB/s range on adjacent combinations. The Spark is comparatively steady, settling into a 1.7–2.0 GB/s band for most of the sweep and peaking at 2,106.6 MB/s (33.7K IOPS) at 2/16.

Latency mirrors the 64K read pattern: the Halo is dramatically faster at low queue depth (0.054 ms at 1/1 vs. 0.217 ms for the Spark), and both drives degrade sharply as queue depth and thread count climb into double digits. At full saturation (32/16), the Halo’s latency reaches 19.611 ms, compared with the Spark’s 17.857 ms. Both drives are clearly under heavy write-amplification stress at this point, with the Halo again slightly worse at the very top of the curve.

16K Sequential Read

Bandwidth remains high throughout most of the sweep, with both drives sawtoothing between roughly 1–8 GB/s depending on the depth/thread combination. The Spark edges out the higher peak, reaching 8,069.3 MB/s (516.4K IOPS) at 32/1, while the Halo tops out at 6,715.8 MB/s (429.8K IOPS) at 8/4.

Latency again favors the Halo everywhere except at the very top of the queue. At 1/1, the Halo answers in 0.027 ms versus the Spark’s 0.214 ms, and the Halo maintains lower latency through most of the sweep. Only at the deepest combinations does the gap close. At 32/16, the Halo’s mean latency of 1.441 ms lands just under the Spark’s 1.599 ms, making this one of the few points where the Halo’s latency curve doesn’t blow past the Spark’s at saturation.

16K Sequential Write

Bandwidth is close here as well: the Spark peaks at 2,141.6 MB/s (137.1K IOPS) at 16/1, and the Halo isn’t far behind at 1,970.9 MB/s (126.1K IOPS) at 1/8.

Latency is where the two diverge sharply. The Halo starts far ahead at low queue depth (0.022 ms at 1/1 versus the Spark’s 0.231 ms), but its latency curve spikes violently as the queue fills. By 32/16, the Halo’s mean latency has climbed to 6.967 ms, well past the Spark’s 4.306 ms at the same point. The Halo’s curve is also far less predictable along the way, with sharp spikes at several mid-range combinations where the Spark remains comparatively flat.

16K Random Read

This is one of Halo’s better showings. It actually posts a higher peak bandwidth of 6,609.4 MB/s (423.0K IOPS) at 32/16, versus the Spark’s 6,082.6 MB/s (389.3K IOPS) at the same combination.

Latency again starts heavily in the Halo’s favor (0.047 ms at 1/1 vs. 0.222 ms for the Spark). Still, the Halo’s latency curve is far more volatile across the sweep, with sharp spikes at 8/1, 16/1, 8/4, and 8/8 that shoot well above the Spark’s comparatively smooth (if higher-baseline) curve. At the very top of the queue, the Halo’s peak latency of 1.971 ms (at 16/16) exceeds the Spark’s peak of 1.315 ms (at 32/16), so the Halo trades consistency for raw low-queue-depth speed.

16K Random Write

Bandwidth favors the Spark here, which reaches 2,074.1 MB/s (132.7K IOPS) at 32/1, compared with the Halo’s 1,503.5 MB/s (96.2K IOPS) at 8/4. The Spark’s bandwidth curve is also considerably more consistent, holding in the 1.6–2.0 GB/s range for most of the sweep after the initial ramp. In comparison, the Halo swings wildly between roughly 0.1 and 1.5 GB/s from one combination to the next.

Latency is where this test stands out: the Halo starts lower at 1/1 (0.154 ms vs. 0.224 ms), but at 8/16 its mean latency spikes to an extreme 20.698 ms, nearly 4.5x Spark’s worst-case 4.664 ms anywhere in the sweep. Aside from that single spike, the Halo’s latency is otherwise reasonable, but it’s a significant outlier worth flagging for any workload that might land on that specific depth/thread combination.

4K Random Read

At low queue depth, the Halo responds dramatically faster (0.04 ms at 1/1 versus the Spark’s 0.215 ms), and it holds a latency advantage through most of the low-to-mid range of the sweep. But the Spark pulls ahead decisively in throughput at scale: its peak IOPS reaches 1,750.7K (6,838.8 MB/s) at 32/16, well clear of the Halo’s peak of 1,050.4K IOPS (4,103.2 MB/s) at 32/8.

Interestingly, the latency picture inverts at high queue depth. Spark’s latency curve remains relatively contained even as depth and threads climb, peaking at just 0.292 ms. The Halo, by contrast, spikes sharply at 16/16 (0.513 ms) and again at 32/16 (1.009 ms), exceeding its steady-state baseline by over 3x, meaning its excellent low-queue-depth responsiveness doesn’t carry through to full saturation.

4K Random Write

This is the widest IOPS gap in the sweep. The Spark’s random 4K write performance climbs steeply as queue depth and thread count increase, reaching 490.4K IOPS (1,915.6 MB/s) at 8/16. The Halo, meanwhile, plateaus much earlier and at a much lower level, topping out at 125.6K IOPS (490.7 MB/s) at 4/4 and never exceeding roughly 110–115K IOPS for the remainder of the sweep. The Spark is running at nearly 4x the Halo’s ceiling here.

Latency again starts in the Halo’s favor at low queue depth (0.079 ms vs. 0.235 ms at 1/1). By 32/16, the Halo’s mean latency has climbed to 4.546 ms, while the Spark’s remains comparatively controlled at 1.792 ms. Combined with the IOPS gap, this is the test where the Spark’s advantage is most pronounced and most consistent across the full depth/thread sweep.

vLLM Online Serving – LLM Inference Performance

vLLM is one of the most popular high-throughput inference and serving engines for LLMs. The vLLM online serving benchmark evaluates the real-world serving performance of this inference engine under concurrent requests. It simulates production workloads by sending requests to a running vLLM server, with configurable parameters such as request rate, input and output lengths, and the number of concurrent clients. The benchmark measures key metrics, including throughput (tokens per second), time to first token, and time per output token (TPOT), helping users understand how vLLM performs under different load conditions.

We tested inference performance across a comprehensive suite of models spanning various architectures, parameter scales, and quantization strategies to evaluate throughput across different concurrency profiles.

GPT OSS 120B

Equal ISL/OSL (256/256): The Halo scaled to 222 tok/s at batch size 64, while the Spark reached 701 (about 3.2x behind).

Prefill Heavy (8k/1k): The Halo’s throughput peaked at batch 32 (427 tok/s), then dipped to 314 at batch 64, while the Spark surged to 2,760 (about 8.8x) — the widest gap in the set.

Decode Heavy (1k/8k): The Halo reached 127 tok/s, compared with the Spark’s 305 at batch size 64 (about 2.4x).

GPT OSS 20B

Equal ISL/OSL (256/256): Spark led in every batch, scaling to 1,917 vs 617 tok/s at batch 64 (Spark about 3.1x ahead).

Prefill Heavy (8k/1k): Spark’s strongest lead, climbing to 3,672 vs 881 tok/s at batch size 64 (Spark about 4.2x ahead).

Decode Heavy (1k/8k): Spark is ahead throughout, reaching 728 vs 330 tok/s at batch 64 (Spark about 2.2x ahead).

Qwen3 Coder 30B A3B Instruct

Equal ISL/OSL (256/256): The Halo scaled to 376 tok/s at batch size 64, roughly half of Spark’s 729 (about 1.9x) — one of the tighter Equal ISL/OSL results.

Prefill Heavy (8k/1k): The Halo peaked at batch 32 (408 tok/s), then dipped to 362 at batch 64, trailing Spark’s 1,663 (about 4.6x).

Decode Heavy (1k/8k): The Halo reached 188 tok/s, compared with Spark’s 357 at batch size 64 (about 1.9x).

Mistral Small 3.1 24B Instruct

Equal ISL/OSL (256/256): The Halo actually led in batch 1 (9 vs 8 tok/s), then settled to 202 tok/s at batch 64, while the Spark was at 498 (about 2.5x behind).

Prefill Heavy (8k/1k): The Halo peaked at batch 32 (188 tok/s) and dipped slightly to 164 at batch 64, trailing the Spark’s 540 (about 3.3x).

Decode Heavy (1k/8k): Nearly tied throughout, the Halo reached 119 tok/s against the Spark’s 132 at batch 64 (within about 11%).

Llama 3.1 8B Instruct

Equal ISL/OSL (256/256): The Halo scaled cleanly from 25 tok/s at batch 1 to 407 tok/s at batch 64, tracking the Spark closely through batch 4, after which the Spark pulled ahead to 1,330 (Halo trailing by about 3.3x at peak).

Prefill Heavy (8k/1k): The Halo reached 546 tok/s at batch size 64, holding roughly half of Spark’s 1,059.

Decode Heavy (1k/8k): The Halo’s most competitive scenario, reaching 235 tok/s against the Spark’s 263 at a batch size of 64 (within about 12%).

Llama 3.1 8B Instruct FP4

Equal ISL/OSL (256/256): The Halo peaked at 267 tok/s at batch size 64, far behind the Spark’s 3,573 — FP4 shows the widest gap (about 13.4x).

Prefill Heavy (8k/1k): The Halo reached 457 tok/s at batch size 64, compared with the Spark’s 2,713 (about 5.9x).

Decode Heavy (1k/8k): The Halo scaled to 146 tok/s, compared with the Spark’s 588 tok/s at a batch size of 64 (about 4x).

Conclusion

The AMD Ryzen AI Halo enters a category NVIDIA effectively created with the DGX Spark, and it does not try to beat the Spark at its own game. In our vLLM sweeps, the Spark held a 2x to 4x throughput advantage in most scenarios at higher concurrency, stretching to 8.8x in prefill-heavy GPT OSS 120B work and narrowing to roughly 10% only in a pair of decode-heavy runs, and its storage subsystem won nearly every fio test at saturation. What the Halo offers instead is the same 128GB of unified memory and 200-billion-parameter ceiling in a similar footprint, at $3,999 against the Spark Founders Edition’s $4,699 (though some OEMs have them under $4,000), in a full x86 machine that boots Windows 11 or Linux rather than DGX OS alone.

AMD Ryzen AI Halo main board top view.

Set the Spark aside, and the Halo is the strongest Ryzen AI Max+ 395 implementation we’ve tested. It posted the top marks against the HP Z2 Mini G1a and ZBook Ultra G1a in nearly every Windows workload we ran: 37,316 in Cinebench R23 multi-core, 184.2 GIPS in 7-Zip against 151.6 and 156.8 for the HPs, the highest Procyon AI text and image generation scores, and the fastest y-cruncher times at 5 and 10 billion digits. On Linux, against the Spark, it won the CPU-bound Phoronix tests outright, compressing 11% faster in 7-Zip and finishing the LLVM compile 14% sooner. This is a compact workstation first, with the AI developer role layered on top rather than replacing it.

Developers who need maximum local tokens per second, or who plan to cluster nodes over high-speed fabric, should still buy the Spark; the Halo’s 10GbE and its vLLM ceilings are not close. For developers building against ROCm, teams that need Windows in the loop, or anyone who wants one box to cover professional workloads and local model work, the Halo is the better fit, and the standard M.2 2280 bay and pre-tuned Variable Graphics Memory remove two of the most common complaints Spark owners have raised. AMD has also said the platform will pick up Ryzen AI Max PRO 400 Series silicon with up to 192GB of unified memory in the third quarter, so buyers chasing larger models have a clear path forward without changing platforms.

Product Page – AMD Ryzen AI Halo

The post AMD Ryzen AI Halo Review: A Dual-OS, 200B-Parameter Desktop Takes On the DGX Spark appeared first on StorageReview.com.

NVIDIA Quietly Makes Omniverse Free for Production Use

3 July 2026 at 18:47
nvidia omniverse free nvidia omniverse free

NVIDIA has dropped the subscription requirement for Omniverse. The platform is now free for development, production, and redistribution, with no NVIDIA AI Enterprise subscription required. For a product that carried a $4,500-per-GPU-per-year list price under NVIDIA AI Enterprise, and whose Nucleus server pricing prompted a developer forum thread titled “Pricing: $25,000 for the nucleus server!?”, this is one of the larger licensing resets NVIDIA has made. It’s great news for the community, but it was announced very quietly via a forum sticky.

nvidia omniverse free

Licensing documentation was updated to read “As of May 2026, Omniverse is freely available for both development and production use,” and a short announcement was posted to the NVIDIA Developer Forums on July 1. Two days later, the announcement threads had zero replies, which says less about interest than about how few people know about them. The timing is intuitive though: SIGGRAPH 2026 runs July 19 to 23 in Los Angeles, and NVIDIA has a full slate of OpenUSD, neural rendering, and physical AI sessions on the schedule. Expect the licensing change to get a more formal introduction there.

What Changed?

Under the old model, Omniverse development was free, but production deployments required an NVIDIA AI Enterprise subscription, which lists at $4,500 per GPU per year, $22,500 per GPU perpetual, or $1 per GPU-hour on CSP marketplaces. The new terms remove that requirement entirely; anything built with Omniverse can also be redistributed under the same free terms, which is a substantive impact for ISVs and integrators who embed Omniverse applications inside their own products.

The catch, such as it is, sits in support. Without a subscription, support is limited to community channels, meaning the Developer Forums and Discord. Organizations that want enterprise support SLAs still buy NVIDIA AI Enterprise through a reseller or CSP marketplace. Partners who redistribute Omniverse within commercial products can use an embedded licensing model in which the partner handles front-line support, and NVIDIA provides back-line support. In practice, NVIDIA has moved Omniverse from a licensed product to a free platform with a paid support add-on, the same model as most open-source infrastructure businesses, minus the open source.

Omniverse Is Not Just a Rendering Tool

Omniverse is typically thought of as a CAD visualization and content creation platform, a way to import Rhino, Revit, and Maya scenes into a single physically based renderer. That framing undersells what NVIDIA has spent the last several years building. Omniverse paired with OpenUSD is the foundation for digital twin simulation, and digital twin simulation has become critical infrastructure for training physical AI: the world models that let robots and vehicles reason about environments they have never seen.

A great example is Alpamayo, the family of open autonomous vehicle models NVIDIA introduced at CES 2026. Alpamayo 1 is a 10-billion-parameter vision-language-action model, built on Cosmos, that generates driving trajectories along with chain-of-causation reasoning traces explaining why it acted. Models like that cannot be trained or validated on road miles alone; rare and dangerous edge cases are better off simulated than on busy streets. That is what the AlpaSim blueprint and NVIDIA’s neural reconstruction pipelines do, and the Omniverse with OpenUSD is the underlying simulation environment. The first production deployment, the Mercedes-Benz CLA on the DRIVE platform, reached the US market this year. Every developer who wants to work in that stack now gets the simulation layer without a license conversation.

Digital Twins, Up to and Including the AI Factory

A digital twin in this context is not a 3D model; it is a physically accurate, continuously updated simulation of a real facility, built from SimReady assets that carry mass, friction, and material properties, described in OpenUSD so that data from CAD, PLM, and simulation tools stays interoperable. The point is to run the facility in software before and during operation in the physical world: test a robot fleet’s routing on a virtual factory floor, validate camera placement, rehearse a production line change, and only then commit steel and concrete. BMW’s virtual factory work remains the canonical manufacturing example, and NVIDIA’s stated ambition is for every factory to have a digital twin before ground breaks.

NVIDIA Omniverse free - AI Data Center Digital Twin

The most self-referential version of that ambition is the AI factory itself. In March, NVIDIA released the Omniverse DSX Blueprint alongside the Vera Rubin DSX AI factory reference design, a framework for building digital twins of gigawatt-scale GPU data centers. The blueprint unifies power, cooling, networking, and operations into a single simulated environment, with Cadence, Schneider Electric, Siemens, Vertiv, Eaton, Jacobs, and others contributing SimReady assets and integrating their design tools. The logic is the same as the factory floor, with larger numbers attached. These buildouts run into the billions of dollars; mistakes in power and cooling design surface too late, after the concrete is poured, and simulating the facility first is cheaper than discovering a thermal problem in production. NVIDIA is using digital twins of the factories that build tokens to sell the chips that fill them, and Omniverse is the tooling for it all.

Why Free, and Why Now?

Omniverse software subscriptions were never going to register next to a data center business now measured in the tens of billions per quarter, but as a licensed product, Omniverse was friction in front of the workloads NVIDIA most wants to exist. Digital twins, robot training, AV simulation, and AI factory design all consume RTX and data center GPUs at scale, and all of them start with a developer standing up Omniverse. Removing the license removes the reason to prototype on something else. The support-attach model keeps an enterprise revenue path open while the platform grows adoption.

The post NVIDIA Quietly Makes Omniverse Free for Production Use appeared first on StorageReview.com.

HP Z8 Fury G6i Review: One Xeon, up to Four Blackwell GPUs

23 June 2026 at 18:02

For most of the last decade, the high-core workstation conversation has been largely led by AMD. Threadripper PRO pushed core counts, cache, and PCIe lanes past what Intel’s Xeon-W line could offer; the prior Xeon-W flagship topped out at 60 cores, while AMD kept climbing. Intel’s Xeon 600 series, launched in February of this year, is the first CPU in years built to take that argument back, reaching 86 cores on the Granite Rapids-WS platform with 128 PCIe 5.0 lanes. The HP Z8 Fury G6i is the system HP built around it.

HP Z8 Fury G6i Front view.

HP frames the Z8 Fury G6i as an AI workstation rather than a CAD tower, which aligns with current industry trends. One Xeon 600 processor pairs with up to four NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition GPUs (300-watt) for 384 GB of aggregate VRAM, or a single 600W RTX PRO 6000 Workstation Edition for peak single-card performance. The platform supports up to 2 TB of DDR5-6400 ECC memory across eight channels, four M.2 drive slots (with options for more), and up to nine PCIe slots in a roughly 54-liter tower. HP also offers a rack mounting kit (5RU), so a team can centralize it, should the use case dictate.

Our review unit sits one rung below the top of the stack. It runs the 64-core Xeon 696X with two RTX PRO 6000 Max-Q cards for 192 GB of combined VRAM, 128 GB of DDR5-6400, and a four-drive NVMe layout that separates a fast Gen5 boot volume from three Gen4 data drives. That configuration frames the question this review works through: how a single high-core Xeon and a dense Blackwell GPU config balance against each other across professional graphics, GPU rendering, and AI inference, and where the platform gives ground.

The Z8 Fury G6i is configurable on the HP website and starts at roughly $7900 at the time of this review. Our configuration comes in at $74,878. It’s worth noting that most of these systems are bought through corporate acquisition, and volume pricing will be better.

Specifications

Specification HP Z8 Fury G6i
Processor Options (Intel W890 Chipset)
Flagship Model Intel Xeon 698X: 86 cores, 172 threads, 2.0GHz base, up to 4.8GHz Turbo Boost, 336MB L3 cache, 350W
High-Core Options Intel Xeon 696X: 64 cores, 128 threads, 2.4GHz base, up to 4.8GHz Turbo, 336MB L3 cache, 350W
Intel Xeon 678X: 48 cores, 96 threads, 2.4GHz base, up to 4.9GHz Turbo, 192MB L3 cache, 300W
Performance Options Intel Xeon 676X: 32 cores, 64 threads, 2.8GHz base, up to 4.9GHz Turbo, 144MB L3 cache, 275W
Intel Xeon 674X: 28 cores, 56 threads, 3.0GHz base, up to 4.9GHz Turbo, 144MB L3 cache, 270W
Intel Xeon 658X: 24 cores, 48 threads, 3.0GHz base, up to 4.9GHz Turbo, 144MB L3 cache, 250W
Intel Xeon 656: 20 cores, 40 threads, 2.9GHz base, up to 4.8GHz Turbo, 72MB L3 cache, 210W
Intel Xeon 654: 18 cores, 36 threads, 3.1GHz base, up to 4.8GHz Turbo, 72MB L3 cache, 200W
Memory & Storage
System Memory 16 DIMM slots; Up to 2TB DDR5-6400 ECC Registered Memory
Total Storage Capacity Up to 104TB total storage
Internal NVMe Slots Supports up to eight PCIe M.2 SSD devices
Front Accessible Storage Up to four front-accessible hot-swappable NVMe drives with LED indicators and email notifications
SATA Support 4TB-12TB 7200RPM SATA Enterprise HDD support; optional slim DVD-ROM/DVD-Writer
Available Graphics
Ultra High-End NVIDIA A800 (40GB GDDR6)
NVIDIA RTX PRO 6000 Blackwell Generation (96GB GDDR7)
High-End NVIDIA RTX PRO 5000 Blackwell Generation (48GB GDDR7)
NVIDIA RTX PRO 4500 Blackwell Generation (32GB GDDR7)
Mid-Range NVIDIA RTX PRO 4000 Blackwell Generation (24GB GDDR7)
NVIDIA RTX PRO 2000 Blackwell Generation (16GB GDDR7)
Entry NVIDIA RTX A1000 (8GB GDDR6)
I/O & Networking
Front Ports 4x USB Type-A 5Gbps (1 charging)
Optional premium front I/O with 2x USB-C 20Gbps
1x headphone/microphone combo jack
Rear Ports 1x USB Type-C 10Gbps
5x USB Type-A 5Gbps
Optional dual Thunderbolt 5 USB-C 40Gbps ports
Networking Integrated Intel I219-LM PCIe GbE
Optional 10GbE / 25GbE networking modules and NICs
Optional Wi-Fi 7 and Bluetooth 5.4
Certifications & Software
ISV Certifications Certified for professional applications and advanced workstation workflows
HP Software Suite HP Anyware Pro
HP Z Remote Graphics Software (RGS)
HP Support Assistant
HP Smart Sense
Security & Management HP Wolf Security
HP Sure Start
HP Sure Click
HP Sure Sense
TPM 2.0
HP BIOSphere
Sustainability & Efficiency EPEAT Gold certified
ENERGY STAR configurations available
60% recycled plastics
20% recycled steel
80 Plus Platinum power supplies
Physical Specifications
Dimensions (H x W x D) 17.5 x 8.6 x 22 in (44.5 x 21.95 x 55.9 cm) up to 17.5 x 10 x 22 in (44.5 x 25.35 x 55.9 cm) with max side panel
Weight Starting at 48.9 lb (22.2 kg)
Power Supply 1350W, 1700W, or 2700W PSU options
Redundant and aggregate power configurations are available

Design and Build

The Z8 Fury G6i carries over the look HP has settled on across its latest workstation lineup, with a uniform matte black finish from the chassis to the front fascia. The face is dominated by a plastic diamond-mesh grille that runs the full height of the tower for airflow, broken only by the front I/O strip near the top and the metallic HP logo lower down. The tower itself is substantial: it starts at 48.9 lb and measures 17.5 x 8.6 x 22 inches without the rear handle, growing to 17.5 x 10 x 22 inches with the maximum side panel fitted. That heft is a function of the dual-PSU, multi-GPU support built inside, but the result is a rigid, well-damped enclosure that feels every bit the professional-grade workstation it is.

Storage

For boot and high-speed flash storage, the Z8 Fury G6i provides four onboard PCIe Gen5 M.2 slots (labeled SSD0 through SSD3), each fitted with a finned heatsink and a tool-free blue latch for retention. HP sells the drives in 1, 2, 4, and 8TB capacities, so the four slots can be populated to suit anything from a single boot drive to a high-capacity NVMe array.

For bulk storage, the Z8 Fury G6i includes two internal 3.5-inch drive bays with tool-free carriers, letting you slot in high-capacity HDDs without a screwdriver. The bays sit on a backplane with the SATA/power connectors fixed in place, so drives seat directly as they slide in.

I/O and Expansion

The Z8 Fury G6i has a fairly standard set of I/O ports for workstations, with four USB Type-A 5 Gbps ports (one charging) and a single 3.5 mm headphone/microphone combo port. HP, however, offers a premium front I/O setup with two USB Type-A 5 Gbps ports (one charging), two USB Type-C 20 Gbps ports (both charging), and a 3.5 mm headphone/microphone combo port. Both front I/O configurations also offer an SD card reader alongside the I/O ports. Our review unit included the premium front I/O package.

HP Z8 Fury G6i front ports.

On the rear of the Z8, we again see a fairly standard I/O setup with an integrated GbE Ethernet port, five USB Type-A 5 Gbps ports, a single USB Type-C 10 Gbps port, and another headphone/microphone combination port. Near our standard I/O ports, we also see the Flex I/O port, which offers up to 10GBASE-T or 2 Thunderbolt 5 ports. Also on the rear is a collapsible handle that folds flush against the chassis when not in use, but pops out to provide a sturdy hold point for lifting or repositioning the workstation.

HP Z8 Fury G6i rear.

When it comes to expansion slots, the Z8 Fury G6i has a total of 9 PCIe slots: 4 PCIe 5 x16, 3 PCIe 5 x8, 1 PCIe 5 x4, and 1 PCIe 4 x4. These four PCIe 5 x16 slots are what allow the Z8 to house the quad NVIDIA RTX PRO 6000 Blackwell Max-Q cards in a fully loaded configuration.

HP Z8 Fury G6i internal view with side panel removed.

The board also exposes a dedicated network MCIO connector that accepts an optional add-in module for 2x 10GbE or 2x 25GbE LANs. Our review unit shipped without the module, leaving the slot open, but it’s the path HP provides for adding higher-speed networking without consuming a standard PCIe expansion slot.

HP pairs a tool-free PCIe retention latch at the top slot with a pivoting bar that swings to release the card, so the top GPU disengages cleanly without fighting the PCIe slot lock.

We can also see to the left of the chassis the dual removable power supplies that can be configured in either a redundant or cumulative configuration, totaling up to 2700W, to feed configurations like the highest spec buildout that contains a 350W TDP CPU and up to 1200W of GPUs, being either a single RTX PRO 6000 Blackwell card or 4x RTX Pro 6000 Blackwell Max-Q (300W). The power supply setup is unique in that it can support high-end configurations that would exceed what a 15-A or 20-A 120V circuit can handle on its own, before requiring a move to a 240V circuit. For areas that could supply two discrete 120V circuits, you can run the hardware in that environment off this dual-PSU configuration

HP Z8 Fury G6i removeable power supplies.

Performance Testing

HP Z8 Fury G6i Side panel.

Review Unit Specifications

Our HP Z8 Fury G6i review unit arrived at the lab with the following specifications:

  • CPU: Intel Xeon 696X (64c/128t)
  • GPU: 2x NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
  • RAM: 128 GB DDR5-6400 ECC (4×32 GB)
  • Storage:
    • Boot Drive: 1x 2 TB HP Z Turbo Drive PCIe 5×4 M.2 SSD
    • Data Drives: 3x 2 TB HP Z Turbo Drive PCIe 4×4 TLC M.2 SSD

Comparison Specifications

For comparative results, we have lined up the HP Z8 Fury G6i against our Intel Xeon 658x test platform, which is loaded with the same 128 GB of DDR5 RAM and an NVIDIA RTX 4090. This platform has a Xeon 6-series workstation CPU in the same class as the Z8, but with a lower core count of 24 cores/48 threads. We have also set the Z8 against our previously reviewed Dell Precision 7875, which features the 96-core/192-thread AMD Threadripper 9995WX, 512 GB of DDR5-5200 ECC RAM, and dual NVIDIA RTX PRO 6000 Blackwell GPUs.

Procyon AI Computer Vision

The Procyon AI Computer Vision Benchmark measures AI inference performance across CPUs, GPUs, and dedicated accelerators using a range of state-of-the-art neural networks. It evaluates tasks such as image classification, object detection, segmentation, and super-resolution using models that include MobileNet V3, Inception V4, YOLO V3, DeepLab V3, Real ESRGAN, and ResNet 50. Tests are run on multiple inference engines, including NVIDIA TensorRT, Intel OpenVINO, Qualcomm SNPE, Microsoft Windows ML, and Apple Core ML, providing a broad view of hardware and software efficiency. Results are reported for float- and integer-optimized models, providing a consistent, practical measure of machine vision performance for professional workloads.

CPU AI Computer Vision Overall Score

In the Procyon AI Computer Vision CPU benchmark, the HP Z8 Fury G6i achieved an overall score of 207, placing it between the Intel Xeon 658x platform (248) and the Dell Precision 7875 (157). The HP system trailed the Xeon platform by 16.5%, while outperforming the AMD-based Precision workstation by 31.8%, demonstrating strong CPU inference performance across a range of computer vision models.

GPU AI Computer Vision Overall Score

Using its dual NVIDIA RTX PRO 6000 Max-Q GPUs, the HP Z8 Fury G6i posted a Procyon AI Computer Vision GPU score of 1,151. While this was lower than the Dell Precision 7875’s 1,619 score, the HP platform still delivered substantial AI acceleration, finishing approximately 29% behind the Dell workstation in overall GPU-based computer vision performance.

CPU Results HP Z8 Fury G6i (Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q) Intel Xeon 658x Test Platform (128 GB RAM | NVIDIA RTX 4090) Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
CPU Results
AI Computer Vision Overall Score 207 248 157
MobileNet V3 3.74 ms 1.26 ms 5.74 ms
ResNet 50 5.34 ms 4.58 ms 6.52 ms
Inception V4 17.16 ms 14.86 ms 20.42 ms
DeepLab V3 29.18 ms 26.29 ms 47.75 ms
YOLO V3 23.23 ms 26.20 ms 21.97 ms
REAL-ESRGAN 837.83 ms 1113.34 ms 1288.54 ms
GPU Results
AI Computer Vision Overall Score 1,151 N/A 1,619
MobileNet V3 0.61 ms N/A 0.45 ms
ResNet 50 0.96 ms N/A 0.82 ms
Inception V4 2.31 ms N/A 2.16 ms
DeepLab V3 21.00 ms N/A 6.60 ms
YOLO V3 4.74 ms N/A 3.48 ms
REAL-ESRGAN 49.81 ms N/A 47.33 ms

Blender 4.5 CPU

Blender is an open-source 3D modeling application. This benchmark was run using the Blender Benchmark utility across CPU and GPU. The score is measured in samples per minute, with higher values indicating better performance.

In the Blender CPU benchmark, the HP Z8 Fury G6i delivered mixed but competitive results. Compared to the Intel Xeon 658x platform, the HP system was substantially faster, posting gains of approximately 90% across all three scenes. However, the AMD-powered Dell Precision 7875 remained the performance leader, outperforming the HP by roughly 50–66% depending on the workload. Even so, the Z8 Fury established itself as a strong CPU rendering platform, comfortably outperforming the Xeon comparison system while narrowing the gap to the high-core-count Threadripper Pro workstation.

Blender CPU (samples per minute; higher is better) HP Z8 Fury G6i (Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q) Intel Xeon 658x Test Platform (128 GB RAM | NVIDIA RTX 4090) Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
Monster 694.751 365.315 1039.121
Junkshop 449.356 234.081 744.601
Classroom 347.940 186.163 574.705

Blender 4.5 GPU

When rendering on the GPU, the HP Z8 Fury G6i and Dell Precision 7875 were effectively neck-and-neck. The HP system held a slight advantage in the Monster (+1.3%) and Junkshop (+1.2%) scenes, while the Dell workstation edged ahead by less than 0.5% in Classroom. Overall, GPU rendering performance between the two dual RTX PRO 6000 platforms was essentially identical, with differences small enough to fall within normal benchmark variance.

Blender GPU (samples per minute; higher is better) HP Z8 Fury G6i (Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q) Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
Monster 7351.497 7259.413
Junkshop 3992.506 3943.343
Classroom 3648.637 3665.272

PCMark 10

PCMark 10 is an industry-standard benchmark that measures overall system performance in modern office environments. It features updated workloads for Windows 10 or 11 and evaluates everyday tasks such as productivity, web browsing, video conferencing, and content creation. The benchmark is easy to run, delivers multi-level scoring (from high-level overall scores to detailed workload scores), and includes dedicated battery-life and storage tests. While UL Solutions now recommends Procyon for newer application-based testing, PCMark 10 remains a reliable and widely used tool for assessing overall PC performance.

In PCMark 10, which measures overall system responsiveness across common productivity, content creation, and office workloads, the HP Z8 Fury G6i posted a score of 7,742. This placed it behind both comparison systems, trailing the Intel Xeon 658x platform (9,657) by approximately 20% and the AMD-based Dell Precision 7875 (11,433) by roughly 32%. While the Z8 Fury is clearly optimized for professional workstations and accelerated compute workloads, the PCMark 10 results show that competing platforms deliver stronger performance across a broader mix of desktop-oriented tasks.

PCMark10 (higher is better) HP Z8 Fury G6i (Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q) Intel Xeon 658x Test Platform (128 GB RAM | NVIDIA RTX 4090) Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
Overall Score 7,742 9,657 11,433

Blackmagic RAW Speed Test

The Blackmagic RAW Speed Test is a performance benchmarking tool that measures a system’s ability to handle video playback and editing with the Blackmagic RAW codec. It evaluates how well a system can decode and play back high-resolution video files, providing frame rates for both CPU- and GPU-based processing.

The HP Z8 Fury G6i turned in an impressive showing in the Blackmagic RAW Speed Test, leading both comparison systems in CPU and GPU decoding performance. In the 8K CPU test, the HP reached 311 FPS, outperforming the Intel Xeon 658x platform (205 FPS) by approximately 52% and nearly doubling the performance of the Dell Precision 7875 (158 FPS). The gap widened even further in the 8K GPU test, where the HP delivered 650 FPS, compared to 181 FPS from the Xeon platform and 276 FPS from the Precision 7875. These results highlight the Z8 Fury’s exceptional capability for high-resolution Blackmagic RAW playback and editing workflows.

Blackmagic RAW (higher FPS is better) HP Z8 Fury G6i (Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q) Intel Xeon 658x Test Platform (128 GB RAM | NVIDIA RTX 4090) Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
8K CPU
311 205 158
8k GPU 650 181 276

Blackmagic Disk Speed Test

The Blackmagic Disk Speed Test evaluates storage performance by measuring read and write speeds, providing insights into a system’s ability to handle data-intensive tasks, such as video editing and large file transfers.

Storage performance was another area where the HP Z8 Fury G6i remained highly competitive. Equipped with a 2TB PCIe Gen5 HP Z Turbo boot drive and three 2TB PCIe Gen4 HP Z Turbo SSDs for data storage, the system delivered 8,911.5 MB/s read and 8,166.5 MB/s write performance on its primary drive. Compared to the Dell Precision 7875, which posted 9,111.4 MB/s read and 9,292.0 MB/s write, the HP trailed by just 2.2% in read throughput and by approximately 12.1% in write throughput.

The HP system also included a secondary storage volume that achieved 4,136.8 MB/s read and 5,149.3 MB/s write, providing ample bandwidth for active project data, scratch disks, and large media workloads. While the Dell workstation held a modest advantage on the primary drive benchmark, the Z8 Fury still delivered more than enough storage performance for demanding content creation, AI, and professional visualization workflows.

DiskSpeedTest (higher is better) HP Z8 Fury G6i (Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q) Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
Read 8,911.5 MB/s 9,111.4 MB/s
Write 8,166.5 MB/s 9,292.0 MB/s
Secondary Disk Test
Read 4,136.8 MB/s N/A
Write 5,149.3 MB/s N/A

3DMark CPU

The 3DMark CPU Profile evaluates processor performance across six threading levels: 1, 2, 4, 8, 16, and max threads. Each test runs the same boid-based simulation workload to assess how well the CPU scales under different thread counts, with minimal GPU involvement. The benchmark helps identify single-threaded efficiency and multithreaded potential for tasks such as gaming, content creation, and rendering. Scores on 8 threads often align with modern DirectX 12 gaming performance, while 1–4-thread results reflect older or esports scenarios.

The 3DMark CPU Profile benchmark showed the HP Z8 Fury G6i delivering solid scaling across thread counts, though it trailed both comparison platforms throughout the test suite. At Max Threads, the HP scored 15,792, finishing about 6.5% behind the Intel Xeon 658x platform (16,890) and 43% behind the AMD-based Dell Precision 7875 (27,670). This trend continued in the lower-thread-count tests, where the HP generally landed within 5–10% of the Xeon system but was further behind the high-core-count Threadripper Pro workstation.

3DMark CPU (higher is better) HP Z8 Fury G6i (Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q) Intel Xeon 658x Test Platform (128 GB RAM | NVIDIA RTX 4090)
Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
Max Threads 15,792 16,890 27,670
16 Threads 11,241 13,022 15,378
8 Threads 6,635 7,185 8,477
4 Threads 3,594 3,751 4,701
2 Threads 1,816 1,944 2,378
1 Threads 895 990 1,237

3DMark Storage

The 3DMark Storage Benchmark tests your SSD’s gaming performance by measuring tasks like loading games, saving progress, installing game files, and recording gameplay. It evaluates how well your storage performs in real-world gaming and supports the latest storage technologies, providing accurate performance insights.

In the 3DMark Storage Benchmark, the HP Z8 Fury G6i achieved an overall score of 2,944, placing it close to the Dell Precision 7875’s 3,221 result. This left the HP system approximately 8.6% behind the Dell workstation, indicating comparable storage responsiveness for game loading, file transfers, and other storage-intensive workloads. The HP system’s secondary drive also posted a respectable score of 2,067.

3DMark Storage (higher is better) HP Z8 Fury G6i (Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q) Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
Overall Score 2,944 3,221
Overall Score (Secondary Drive) 2,067 N/A

LuxMark

LuxMark is a GPU benchmark that uses LuxRender, an open-source ray-tracing renderer, to evaluate a system’s performance on highly detailed 3D scenes. This benchmark is relevant for assessing the graphical rendering capabilities of servers and workstations, especially for visual effects and architectural visualization applications, where accurate light simulation is crucial.

In LuxMark, the HP Z8 Fury G6i delivered performance very close to that of the Dell Precision 7875. In the Food scene, the HP scored 41,476, trailing Dell’s 41,981 by just 1.2%. The gap widened slightly in the more demanding Hall workload, where the HP reached 95,414 compared to 101,808 from the Precision 7875, a difference of roughly 6.3%. Overall, the results show that the Z8 Fury provides GPU rendering performance nearly equivalent to that of the Dell workstation in ray-traced rendering workloads.

LuxMark (higher is better) HP Z8 Fury G6i (Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q) Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
Food 41,476 41,981
Hall 95,414 101,808

Geekbench 6

Geekbench 6 is a cross-platform benchmark that measures overall system performance.

In Geekbench 6, the HP Z8 Fury G6i delivered performance very similar to the Intel Xeon 658x platform while trailing the AMD-powered Dell Precision 7875. In the CPU tests, the HP scored 2,333 in single-core and 21,110 in multi-core performance, placing it within 2% of the Xeon system while trailing the Precision 7875 by approximately 28% in single-core and 26% in multi-core performance.

On the GPU side, the HP posted 291,727 in OpenCL and 276,201 in Vulkan. Compared to the Dell Precision 7875, which scored 330,765 and 309,146, respectively, the HP trailed by roughly 12% in OpenCL and 11% in Vulkan.

GeekBench (Higher is better) HP Z8 Fury G6i (Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q) Intel Xeon 658x Test Platform (128 GB RAM | NVIDIA RTX 4090) Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
CPU Single Core 2,333 2,383 3,240
CPU Multi-Core 21,110 21,447 28,618
GPU OpenCL 291,727 N/A 330,765
GPU Vulkan 276,201 N/A 309,146

Y-Cruncher

y-cruncher is a multithreaded and scalable program that can compute Pi and other mathematical constants to trillions of digits. Since its launch in 2009, it has become a popular benchmarking and stress-testing application for overclockers and hardware enthusiasts.

The HP Z8 Fury G6i performed exceptionally well in the Y-Cruncher benchmark, consistently outperforming the Intel Xeon 658x platform and remaining highly competitive with the AMD-based Dell Precision 7875. In the smaller datasets, the HP was the fastest system tested, completing the 250 million-digit run in 1.203 seconds, approximately 30% faster than the Xeon platform and nearly 50% faster than the Precision 7875. This trend continued through the 500-million- and 1-billion-digit tests, with HP maintaining the lead.

As workload sizes increased, the Dell Precision 7875’s larger memory capacity and higher core count began to show their advantage. At 2.5 billion digits, the HP completed the run in 17.0 seconds, trailing the Dell by about 12% while dramatically outperforming the Xeon platform. The gap widened in the larger datasets, with the Precision 7875 leading the 5-, 10-, and 25-billion-digit tests by approximately 21–27%.

Y-Cruncher (lower duration is better) HP Z8 Fury G6i (Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q) Intel Xeon 658x Test Platform (128 GB RAM | NVIDIA RTX 4090) Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
250 Million 1.203 s 1.711 s 2.369 s
500 Million 2.667 s 3.758 s 4.281 s
1 Billion 6.042 s 8.312 s 7.617 s
2.5 Billion 17.047 s 23.065 s 15.188 s
5 Billion 37.890 s 54.207 s 29.795 s
10 Billion 82.983 s 118.898 s 61.572 s
25 Billion 232.832 s 326.454 s 169.289 s
50 Billion N/A N/A 371.039 s
100 Billion N/A N/A 844.503 s

7-Zip Compression

The 7-Zip Compression Benchmark evaluates CPU performance during compression and decompression, measuring GIPS (Giga Instructions Per Second) and CPU usage. Higher GIPS and efficient CPU usage indicate superior performance.

In the 7-Zip Compression Benchmark, the HP Z8 Fury G6i delivered the strongest overall result among the systems tested. Looking at the Total Rating, the HP achieved 344.964 GIPS, outperforming the Intel Xeon 658x platform’s 250.083 GIPS by approximately 38% and finishing well ahead of the Dell Precision 7875. The HP also led the Resulting Compression Rating, posting 340.875 GIPS compared to 233.557 GIPS from the Xeon platform, a margin of roughly 46%.

The decompression results followed a similar pattern, with the HP reaching a Resulting Decompression Rating of 349.054 GIPS, compared to 266.608 GIPS on the Xeon platform.

7-Zip Compression Benchmark (higher is better) HP Z8 Fury G6i (Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q) Intel Xeon 658x Test Platform (128 GB RAM | NVIDIA RTX 4090) Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
Compression
Current CPU Usage 5,483% 4,038% 6,445%
Current Rating/Usage 6.247 GIPS 5.786 GIPS 6.949 GIPS
Current Rating 342.522 GIPS 233.641 GIPS 48.392 GIPS
Resulting CPU Usage 5,461% 4,037% 701%
Resulting Rating/Usage 6.242 GIPS 5.705 GIPS 7.010 GIPS
Resulting Rating 340.875 GIPS 233.557 GIPS 49.108 GIPS
Decompression
Current CPU Usage 6,029% 4,688% 728%
Current Rating/Usage 5.839 GIPS 5.705 GIPS 6.801 GIPS
Current Rating 352.023 GIPS 267.475 GIPS 49.526 GIPS
Resulting CPU Usage 5,990% 4,657% 749%
Resulting Rating/Usage 5.827 GIPS 5.725 GIPS 6.832 GIPS
Resulting Rating 349.054 GIPS 266.608 GIPS 51.181 GIPS
Total Rating
Total CPU Usage 5,726% 4,347% 725%
Total Rating/Usage 6.034 GIPS 5.755 GIPS 6.921 GIPS
Total Rating 344.964 GIPS 250.083 GIPS 50.145 GIPS

V-Ray

The V-Ray Benchmark measures rendering performance on CPUs, NVIDIA GPUs, or both, using the advanced V-Ray 6 engines. It uses quick tests and a simple scoring system to help users evaluate and compare their systems’ rendering capabilities. It’s an essential tool for professionals seeking efficient performance insights.

In the V-Ray benchmark, the HP Z8 Fury G6i delivered a score of 28,237, placing it close to the Dell Precision 7875’s 30,356 result. The HP trailed by approximately 7%, indicating that both systems offer similar rendering capabilities for professional visualization and content creation workloads.

V-Ray (higher is better) HP Z8 Fury G6i (Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q) Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
Score 28,237 30,356

SPECworkstation 4.0 Results

The SPECworkstation 4.0 benchmark is a comprehensive tool that evaluates all key aspects of workstation performance. It offers a real-world measure of CPU, graphics, accelerator, and disk performance, ensuring professionals have the data to make informed decisions about their hardware investments. The benchmark includes a dedicated set of tests focused on AI and ML workloads, such as data science tasks and ONNX Runtime-based inference tests, reflecting the growing importance of AI/ML in workstation environments. It encompasses seven industry verticals and four hardware subsystems, providing a detailed and relevant measure of today’s workstations’ performance.

The HP Z8 Fury G6i turned in a strong overall showing in SPECworkstation 4.0, consistently outperforming the Intel Xeon 658x test platform across most workloads while remaining competitive with the AMD-powered Dell Precision 7875. In the industry vertical tests, the HP led the Xeon system in AI & Machine Learning (+11%), Energy (+71%), Financial Services (+90%), Life Sciences (+46%), Media & Entertainment (+9%), and Product Design (+8%), highlighting the benefits of its higher-end workstation configuration and dual professional GPUs.

Within the hardware subsystem scores, the HP maintained advantages over the Xeon platform in CPU performance (+27%), Accelerator performance (+2%), and Graphics performance (+43%), while trailing only in storage performance. Compared to the Dell Precision 7875, the HP generally ranked second, though it remained relatively close in AI & Machine Learning (3.82 vs 4.42), Life Sciences (5.01 vs 5.34), and Accelerator performance (6.27 vs 7.51).

SPECworkstation 4.0 HP Z8 Fury G6i
Intel Xeon 696X | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q
Intel Xeon 658X Test Platform
128 GB RAM | NVIDIA RTX 4090
Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
Industry Vertical Scores
AI & Machine Learning 3.82 3.44 4.42
Energy 6.11 3.57 10.11
Financial Services 5.48 2.89 8.03
Life Sciences 5.01 3.44 5.34
Media & Entertainment 3.61 3.31 4.82
Product Design 2.97 2.75 3.97
Productivity & Development 1.05 1.35 1.72
Hardware Subsystem Scores
CPU 3.21 2.52 4.58
Accelerator 6.27 6.16 7.51
Graphics 11.17 7.82 15.07
Storage 1.03 1.52 1.29

SPECviewperf 15 Results

SPECviewperf 15 is the industry-standard benchmark for evaluating 3D graphics performance across OpenGL, DirectX, and Vulkan APIs. It introduces new workloads, including blender-01 (Blender 3.6), unreal_engine-01 (Unreal Engine 5.4, DirectX 12), and enscape-01 (Enscape 4.0, Vulkan ray tracing), along with updated traces for 3ds Max, CATIA, Creo, Maya, and SolidWorks. With its redesigned GUI, modern application support, and advanced rendering workloads, SPECviewperf 15 provides consistent, real-world insights into professional graphics performance.

In SPECviewperf 15, the HP Z8 Fury G6i delivered strong professional graphics performance and remained competitive with the Dell Precision 7875 despite both systems utilizing dual RTX PRO 6000 GPUs. The HP was particularly strong in Energy and Medical workloads, scoring 116.75 vs. 114.42 and 136.95 vs. 136.06, respectively, giving it a slight advantage in those tests. The two systems were also effectively tied in Blender (90.38 vs. 90.83) and Enscape (52.22 vs. 52.53), with less than a 1% difference between them.

The Dell workstation maintained larger leads in several engineering and CAD-focused workloads, including Creo (+46%), Unreal Engine (+42%), CATIA (+19%), SolidWorks (+16%), and Maya (+17%). However, the HP remained highly competitive across the benchmark suite and demonstrated particularly strong performance in visualization, rendering, and simulation-oriented workloads.

Workload HP Z8 Fury G6i
Intel Xeon 696X | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q
Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
Composite Scores
3dsmax-08 86.06 92.80
blender-01 90.38 90.83
catia-07 94.77 112.41
creo-04 155.15 227.18
energy-04 116.75 114.42
enscape-01 52.22 52.53
maya-07 137.90 162.04
medical-04 136.95 136.06
snx-05 N/A 93.87
solidworks-08 154.40 178.68
unreal_engine-01 68.35 96.86

Topaz Video AI

Topaz Video AI is a professional application for enhancing and restoring video using advanced AI models. It supports tasks such as upscaling footage to 4K or 8K, sharpening blurry content, reducing noise, improving facial details, colorizing black-and-white footage, and interpolating frames for smoother motion. The suite includes an onboard benchmark that measures system performance across its various video-enhancing algorithms, providing a clear view of how well hardware platforms handle demanding AI video-processing workloads.

In Topaz Video AI, the HP Z8 Fury G6i delivered strong performance across the suite’s video enhancement and upscaling models. However, the Dell Precision 7875 generally maintained the lead in the most demanding AI workloads. In the commonly used 1X enhancement models, the HP reached 37.5 FPS in Artemis, 37.4 FPS in Iris, and 38.5 FPS in Proteus, while the Dell workstation achieved roughly 25–40% higher performance in those same tests. However, the HP did post a notable win in the Gaia model, achieving 16.3 FPS compared to 14.7 FPS on the Dell system.

The trend continued in the heavier 2X and 4X upscaling workloads, where the Dell platform generally delivered higher throughput, reflecting the advantage of its higher-end CPU and larger memory configuration. That said, the HP remained competitive in several motion interpolation tests, outperforming the Dell in 4X Slowmo APFast (52.9 FPS vs. 34.8 FPS) and 4X Slowmo Chronos (37.0 FPS vs. 33.0 FPS).

Overall, the results show the HP Z8 Fury G6i as a capable AI video-processing workstation that performs well across Topaz Video AI’s broad range of enhancement models. At the same time, the Dell Precision 7875 generally leads in the most computationally intensive upscaling and restoration workloads.

Test / Model HP Z8 Fury G6i
(Intel Xeon 696x | 128 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
Dell Precision 7875 (AMD 9995WX 96C | 512 GB RAM | 2x NVIDIA RTX PRO 6000 Max-Q)
Benchmark Results – 1X
Artemis 37.49 fps 46.87 fps
Iris 37.35 fps 52.63 fps
Proteus 38.49 fps 48.19 fps
Gaia 16.30 fps 14.71 fps
Nyx 15.35 fps 22.90 fps
Nyx Fast 34.83 fps 50.42 fps
Nyx XL N/A 3.66 fps
Hyperion HDR 21.55 fps 28.57 fps
Benchmark Results – 2X
Artemis 13.45 fps 22.99 fps
Iris 13.24 fps 20.24 fps
Proteus 13.28 fps 23.67 fps
Gaia 11.83 fps 10.36 fps
Nyx 11.89 fps 18.81 fps
Benchmark Results – 4X
Artemis 3.49 fps 6.53 fps
Iris 3.71 fps 6.53 fps
Proteus 3.71 fps 6.40 fps
Gaia 3.61 fps 6.04 fps
Rhea 3.63 fps 5.71 fps
RXL 3.64 fps 5.84 fps
Slow Motion Benchmarks
4X Slowmo – Apollo 30.15 fps 37.51 fps
4X Slowmo – APFast 52.91 fps 34.78 fps
4X Slowmo – Chronos 36.99 fps 33.00 fps
4X Slowmo – CHFast 28.66 fps 38.86 fps
16X Slowmo – Aion 24.69 fps 37.29 fps

HP Z8 Fury G6i vLLM Performance Testing

To evaluate the HP Z8 Fury G6i, we tested configurations using the vLLM Online Serving benchmark, one of the most widely adopted high-throughput inference and serving engines for large language models. The vLLM online serving benchmark simulates real-world production workloads by sending concurrent requests to a running vLLM server and measuring key metrics, including total token throughput (tokens per second), time to first token, and time per output token, under varying load conditions.

Our testing spanned a range of models, from dense architectures to micro-scaling data types. The tests evaluated performance across three workload scenarios: Equal ISL/OSL, Prefill Heavy, and Decode Heavy. These scenarios represent distinct real-world serving patterns, from balanced input and output loads to compute-intensive prompt processing and memory-bandwidth-bound token generation.

To benchmark the HP Z8 Fury G6i, we tested a dual-GPU configuration (2x NVIDIA RTX PRO 6000 Blackwell). Because the system was tested with the same NVIDIA RTX PRO 6000 Blackwell cards used in our Dell Precision 7875 review, the results provide a direct platform-to-platform comparison between the two workstations.

GPT-OSS-120B

Equal ISL/OSL (256/256): The Dell led marginally through batch 32 (4,409 vs 4,341), then the HP edged ahead to 12,604 vs 11,848 tok/s at batch 256 (about 6%).

Prefill Heavy (8k/1k): Nearly identical, HP slightly ahead, climbing to 20,347 vs 18,954 tok/s at batch 256 (about 7%).

Decode Heavy (1k/8k): Virtually tied the entire way, HP peaking at 5,445 vs 5,275 tok/s at batch 256 (about 3%).

GPT-OSS-20B

Equal ISL/OSL (256/256): The Dell led slightly through batch 8 (3,085 vs 3,071), then the HP took over, peaking at 24,131 vs 22,034 tok/s at batch 256 (about 10%).

Prefill Heavy (8k/1k): HP ahead throughout, reaching 35,193 vs 31,982 tok/s at batch 256, the highest absolute throughput in the suite, about 10% over the Dell.

Decode Heavy (1k/8k): Essentially matched through batch 16, HP pulling slightly ahead at scale to 10,461 vs 9,985 tok/s at batch 256 (about 5%).

Qwen3 Coder 30B FP8

Equal ISL/OSL (256/256): The Dell led through the HP’s batch-8 dip (1,783 vs 539), but the HP recovered at batch 16 and pulled steadily away to 18,435 tok/s vs 13,577 at batch 256, a 36% HP advantage.

Prefill Heavy (8k/1k): HP ahead throughout, peaking at 14,780 vs 13,661 tok/s at batch 128 (about 8%); both tapered at batch 256.

Decode Heavy (1k/8k): HP led the full curve, finishing at 3,908 vs 3,464 tok/s at batch 256 (about 13%).

Qwen3 Coder 30B BF16

Equal ISL/OSL (256/256): Even early (the HP took a sharp batch-8 dip to 457), then the HP climbed steeply to 14,282 tok/s vs the Dell’s 10,171 at batch 256, a 40% advantage.

Prefill Heavy (8k/1k): Closely tracked through batch 32, then the HP separated, peaking at 13,485 vs 11,789 tok/s. Both rolled off at batch 256 (HP 12,011 vs 9,381), with HP holding a 28% lead.

Decode Heavy (1k/8k): HP ahead for most of the curve, peaking at 3,409 vs 3,019 tok/s at batch 128 (about 13%).

Mistral Small 24B

Equal ISL/OSL (256/256): Essentially overlapping through batch 64 (4,745 vs 4,730), with the HP edging ahead at the top to 8,766 vs 8,261 tok/s at batch 256 (about 6%).

Prefill Heavy (8k/1k): Tightly matched, both peaking at batch 64 (HP 6,789 vs Dell 6,627) before falling off sharply at higher batches, HP just 2% ahead at peak.

Decode Heavy (1k/8k): Near-identical curves; the Dell even nudged ahead at batch 64, with the HP peaking at 1,894 vs 1,831 tok/s at batch 128 (about 3%). The closest-matched model overall.

Llama 3.1 8B (FP8)

Equal ISL/OSL (256/256): The Dell led early, 427 vs 342 tok/s at batch 1, and held the lead through batch 16 (5,050 vs 2,310), where the HP run dipped at batch 8. The HP recovered hard at batch 32 (8,945 vs 8,341) and pulled away from there, peaking at 23,004 tok/s vs the Dell’s 16,833 at batch 256, a 37% HP advantage.

Prefill Heavy (8k/1k): The two tracked closely, with the Dell marginally ahead at batch 1 (1,803 vs 1,682). The HP took over by batch 4 and stayed ahead, reaching 21,693 tok/s at batch 256 vs 18,822 tok/s, about 15% higher.

Decode Heavy (1k/8k): The HP led throughout by a steady margin, finishing at 6,287 tok/s vs the Dell’s 5,429 at batch 256, a 16% gain.

Llama 3.1 8B BF16

Equal ISL/OSL (256/256): Roughly even through batch 4 (about 1,109 each), with the Dell briefly ahead at batches 8 to 16 during an HP dip. From batch 32 on the HP led, peaking at 17,921 tok/s vs 13,789 at batch 256, 30% higher.

Prefill Heavy (8k/1k): Nearly identical curves, HP slightly ahead the whole way. Both peaked at batch 128 (HP 12,542 vs Dell 11,639), then tapered at batch 256.

Decode Heavy (1k/8k): HP led by a small, consistent margin, peaking at 3,435 vs 3,225 tok/s at batch 128 (about 7%).

Platform Differences: Why the HP Pulls Ahead

Both workstations run identical 2x RTX PRO 6000 Blackwell GPUs, and neither platform supports NVLink, so all inter-GPU communication during tensor-parallel inference travels over PCIe. The Dell Precision 7875 has one PCIe Gen 5 x16 slot with the second GPU running at Gen 4 x16, while the HP Z8 Fury G6i provides full PCIe Gen 5 x16 to all GPU slots.

The host CPU also plays a role since vLLM’s scheduler and token processing run on the CPU rather than the GPU, and the Intel Xeon 6 platform in the HP offers architectural advantages that the Threadripper PRO does not. These platform differences are most pronounced on smaller and quantized models at high concurrency, where the HP leads by 30-37%, and shrink to the low single digits on large MoE models like GPT-OSS-120B, where GPU compute time is the dominant factor.

Conclusion

The HP Z8 Fury G6i is HP’s statement that Intel is back in the high-core workstation conversation. Built around the Granite Rapids-WS Xeon 600 series, our review unit paired the 64-core Xeon 696X with two NVIDIA RTX PRO 6000 Blackwell Max-Q cards, 128 GB of DDR5-6400, and a four-drive NVMe layout, positioning it squarely as an AI workstation rather than a traditional CAD tower. The build reflects HP’s current design language: a uniform matte black chassis with a full-height diamond-mesh grille, paired with a genuinely serviceable interior featuring tool-free NVMe latches, removable dual power supplies, a collapsible rear handle, and a nine-slot, four-PCIe-Gen5-x16 layout that scales to four Max-Q cards.

In our benchmarks, the Z8 Fury staked out clear wins where its platform strengths matter. It dominated the Blackmagic RAW Speed Test at 311 FPS (8K CPU) and 650 FPS (GPU), topped the field in 7-Zip compression at 344.96 GIPS, led the smaller Y-Cruncher datasets, and posted strong SPECworkstation 4.0 results that outpaced the Intel Xeon 658x platform across nearly every vertical. GPU rendering in Blender and LuxMark was effectively a tie with the Dell Precision 7875, as expected given the shared RTX PRO 6000 silicon.

HP Z8 Fury G6i inside with gpus removed.

Where it gives ground, the pattern is consistent. The 96-core Threadripper PRO in the Dell still leads in heavily multithreaded and memory-bound workloads, including Blender CPU, the larger Y-Cruncher runs, 3DMark CPU, and several Topaz Video AI models, reflecting its higher core count and 512 GB of memory. The vLLM inference results are where the platform argument gets interesting: running identical dual RTX PRO 6000 cards, the HP pulled ahead by 30 to 40 percent on smaller and quantized models at high concurrency, narrowing to low single digits on large MoE models like GPT-OSS-120B, where GPU compute dominates. With neither platform supporting NVLink, that gap traces back to HP’s full PCIe Gen5 x16 to every GPU slot and the Xeon 6 host handling vLLM’s CPU-side scheduling more effectively than the Threadripper PRO.

At $52,139 as configured, the Z8 Fury G6i is a serious investment, though most buyers will see better volume pricing through corporate channels. What you get is a thoroughly engineered, highly serviceable AI workstation that, depending on workload, trades blows with the best Threadripper PRO towers and pulls clearly ahead in PCIe-bound multi-GPU inference. For organizations standardizing on Intel and prioritizing GPU-accelerated AI work, the HP Z8 Fury G6i makes a strong case for itself.

HP Z8 Fury G6i Product Page

The post HP Z8 Fury G6i Review: One Xeon, up to Four Blackwell GPUs appeared first on StorageReview.com.

HighPoint Rocket 1604L Review: Four Gen5 M.2 SSDs, One Slot, 55.6GB/s

19 June 2026 at 16:28

The HighPoint Rocket 1604L is a $399 PCIe Gen5 x16 add-in card that carries four M.2 NVMe SSDs, each on a dedicated Gen5 x4 connection. In our testing with four Samsung 9100 PRO 4TB drives installed, the card sustained 55.6GB/s of 128K sequential read bandwidth and 10.1 million 4K random write IOPS, numbers that are within a few percent of what the four drives are rated to deliver on native motherboard slots. That is the entire pitch of this card: it adds drive bays without subtracting performance.

HighPoint Rocket 1604L front view.

The 1604L takes a different architectural path than most quad-M.2 cards we have looked at. It is not a passive bifurcation riser, nor is it a PCIe switch card. Instead, it is built around an Astera Labs PT5161LRS retimer, which sits in the data path at the physical layer, re-clocking and regenerating the Gen5 signal between the host slot and each M.2 connector. At Gen4 speeds, passive cards that simply route traces from the slot to the connectors are usually fine. At Gen5’s 32GT/s signaling rate, trace length and connector transitions start eating into the signal budget, and marginal links train down to Gen4 or throw correctable errors under load. The retimer approach addresses that without the cost, power, and latency of a full PCIe switch. The trade-off is that the host platform must support x4/x4/x4/x4 bifurcation on the slot, since the retimer does not perform any lane virtualization of its own.

This card joins a HighPoint Gen5 family we have covered previously, which includes the switch-based Rocket 1604A, which works in any x16 slot regardless of bifurcation support, and the Rocket 7604A, which adds bootable RAID on top. The 1604L is the leanest of the three. There is no RAID stack and no driver; the operating system simply enumerates four native NVMe devices, and anything beyond that (mdadm, Storage Spaces, ZFS) is up to the user. HighPoint positions the card heavily toward servers hosting M.2 accelerator modules like the Hailo-8 series, but for our purposes, the storage use case is the more universal one. The card is a full-height, half-length design that HighPoint claims is roughly 40% shorter than typical four-bay M.2 cards, with a full-length anodized aluminum heatsink, thermal padding for the drives, an integrated low-decibel fan, and a ventilated bracket. Firmware-level monitoring exposes per-port lane allocation, power draw, and board health, with present and activity LEDs for each SSD.

The bifurcation requirement is the caveat to settle before buying. Most mainstream consumer boards either cannot split a x16 slot four ways or steal those lanes from the primary GPU slot. Where the 1604L makes immediate sense is on platforms with PCIe lanes to spare: Threadripper TRX50 and WRX90, Xeon W, and EPYC or Xeon server boards, where x4/x4/x4/x4 is a BIOS toggle and a spare x16 slot is not a sacrifice. That describes our test rig, so the fit was natural.

HighPoint Rocket 1604L Specifications

Specification Rocket 1604L (R1604L)
Bus Interface PCIe 5.0 x16
Chipset Astera Labs PT5161LRS retimer
Working Mode 4 x 4-lane (host bifurcation x4/x4/x4/x4 required)
Ports 4x M.2 NVMe (dedicated PCIe 5.0 x4 per port)
Device Support M.2 NVMe SSDs or M.2 PCIe accelerator modules
SSD Form Factors M.2 2242, 2260, 2280
Data Transfer Rate Up to 64GB/s
RAID Support None (OS-level software RAID optional)
Form Factor Full-height, half-length
Cooling Full-length aluminum heatsink, integrated fan, thermal pads, ventilated bracket
Monitoring Per-port lane allocation, power, and health via smart firmware; present and activity LEDs
OS Support Native NVMe support in mainstream operating systems, x86 Intel/AMD and ARM
Price $399 (HighPoint eStore)

Build and Design

HighPoint Rocket 1604L top heatsink removed with 4 M.2 drives installed.

The 1604L’s compact footprint is the visible difference from the sprawling four-bay cards of the Gen4 era. Drive installation is conventional: heatsink off, drives into the four sockets, thermal pads aligned, heatsink back on. The single fan exhausts through the ventilated bracket, which matters in workstation towers where slot airflow is unpredictable. We did not observe thermal throttling from any of the four drives during sustained 60-second test runs.

HighPoint Rocket 1604L heatsink removed from card.

Testing Setup

We tested the Rocket 1604L in our consumer Threadripper platform, the same water-cooled rig that has handled our recent high-end GPU and HEDT CPU reviews. The card was installed in a Gen5 x16 slot configured for x4/x4/x4/x4 bifurcation.

StorageReview Threadripper Test Platform

  • CPU: AMD Ryzen Threadripper 7980X (64C/128T)
  • Motherboard: ASUS Pro WS TRX50-SAGE WIFI
  • RAM: 128GB DDR5-6400
  • Storage: 1TB Gen4 Boot SSD, 4x Samsung 9100 PRO 4TB (FW 0B2QNXH7) on the Rocket 1604L
  • OS: Ubuntu Server 24.04

The four Samsung 9100 PRO drives are each rated at 14,800MB/s sequential read, 13,400MB/s sequential write, 2,200K random read IOPS, and 2,600K random write IOPS, which puts the theoretical aggregate at 59.2GB/s read and 8.8 million random read IOPS. Since a Gen5 x16 slot tops out at roughly 63GB/s of usable bandwidth, the drives, not the slot, are the ceiling in this configuration. That is the right way around; a card like this should never be the bottleneck.

All workloads were run with FIO 3.36 using the io_uring engine against the raw block devices, with a 5% LBA span per drive, 60-second runtimes with a 5-second ramp, and one job per drive at QD64 for sequential transfers or 16 jobs per drive at QD32 (64 total) for 4K random. These are burst-oriented consumer test parameters rather than enterprise steady-state methodology, consistent with how we evaluate client platform accessories.

HighPoint Rocket 1604L Performance

Sequential Bandwidth

Workload (4 drives aggregate) IOPS Bandwidth Avg Latency 99th % Latency
128K Sequential Read, QD64 424K 55.6GB/s 604µs 906µs
128K Sequential Write, QD64 279K 36.5GB/s 918µs 1,303µs
64K Sequential Read, QD64 668K 43.8GB/s 383µs 570µs
64K Sequential Write, QD64 462K 30.3GB/s 553µs 914µs

The headline number is the 128K sequential read result of 55.6GB/s, which works out to 13.9GB/s per drive, or about 94% of Samsung’s 14,800MB/s rating for the 9100 PRO. Getting four Gen5 drives to within striking distance of their individual spec sheets, simultaneously, through a single add-in card is the result that validates the retimer architecture. Average latency held at 604µs with the 99th percentile at 906µs, and per-drive utilization stayed pinned above 99% for the duration of the run. The 64K read result of 43.8GB/s trails the 128K figure as expected, since larger transfers amortize protocol overhead more efficiently.

Sequential writes landed at 36.5GB/s at 128K and 30.3GB/s at 64K. That is below the four drives’ combined 53.6GB/s write rating, which is a drive behavior rather than a card limitation: vendor write specs reflect short bursts into pSLC cache, while our 60-second sustained runs push past that window. The write latency profile stayed orderly, with the 128K test averaging 918µs and holding 1,303µs at the 99th percentile.

4K Random Performance

Workload (4 drives aggregate) IOPS Bandwidth Avg Latency 99th % Latency
4K Random Read, QD32 x 64 jobs 8.83M 36.2GB/s 231µs 553µs
4K Random Write, QD32 x 64 jobs 10.1M 41.5GB/s 202µs 461µs

The random results are the cleanest evidence that the 1604L’s data path is transparent. Samsung rates the 9100 PRO 4TB at 2,200K random read IOPS, and four of them behind the 1604L produced 8.83 million, which is the rated aggregate almost to the decimal. Random write reached 10.1 million IOPS against a theoretical ceiling of 10.4 million, about 97% of spec. Writes-outrunning-reads looks odd at first glance but matches the drives’ own ratings, helped along by the 5% working set, which keeps the controllers operating in their happiest caching range.

Latency under these loads stayed tight, averaging 231µs for reads and 202µs for writes, with 99th percentile figures of 553µs and 461µs, respectively. The other observation worth passing along is host cost: driving nearly 10 million IOPS through 64 FIO jobs consumed roughly 60% of the system CPU time over the run. The card will hand a workstation more storage performance than most applications can absorb, and feeding it is a workload in its own right.

Conclusion

The Rocket 1604L does one job, and our test data shows it doing that job with effectively no overhead. Four Samsung 9100 PRO 4TB drives delivered 55.6GB/s of sequential read bandwidth, 8.83 million random read IOPS, and 10.1 million random write IOPS through the card, figures that sit at 94 to 100% of the drives’ combined ratings. For a device whose value proposition is invisibility, that is a clean sweep.

HighPoint Rocket 1604L rear view.

The buyer’s question is whether the $399 ask is justified, given that passive bifurcation cards sell for a fraction of that price. At Gen4 and below, it often is not. At Gen5, the signal integrity margin is thin enough that the retimer earns its keep, particularly for users planning to load the card with drives that each move 14GB/s. Worked out per bay, $100 per Gen5 M.2 slot with cooling and monitoring included is reasonable against the alternative of unstable link training on a passive card, and it undercuts switch-based options while preserving the full bandwidth of every port.

Who should buy it: TRX50, WRX90, Xeon W, and server platform owners who want 16TB or more of Gen5 flash in a single slot for media work, AI dataset staging, or scratch space, and who are comfortable with OS-level RAID or none at all. Who should not: anyone on a platform without x4/x4/x4/x4 bifurcation support, who should look at the switch-based Rocket 1604A instead, and anyone needing bootable hardware RAID, which is the Rocket 7604A’s territory. Buyers running Gen4 drives can also save money with simpler cards, since the retimer’s advantages are largely wasted below 32GT/s.

HighPoint Rocket 1604L Product Page

The post HighPoint Rocket 1604L Review: Four Gen5 M.2 SSDs, One Slot, 55.6GB/s appeared first on StorageReview.com.

Kioxia CD9P-R Review: Read-Intensive Gen5 Up to 61.44TB

18 June 2026 at 17:36

The Kioxia CD9P-R is the read-intensive arm of the company’s new data center NVMe SSD generation, and the first CD-series drive built on BiCS FLASH generation 8 TLC. The series pairs Kioxia’s own controller and firmware with PCIe 5.0 and NVMe 2.0, with rated performance reaching 14,800 MB/s sequential read and 2.6 million random read IOPS depending on capacity. E3.S models run from 1.92TB to 30.72TB, while the 2.5-inch variant extends the stack to 61.44TB. Our review unit is the 7.68TB E3.S model in SED trim (KCD9DPJE7T68).

KIOXIA CD9P-R E3.S SSD front view.

The capacity stack has a wrinkle worth understanding before the charts. Kioxia’s published specs peak in the middle of the family, not at the top: the 7.68TB and 15.36TB models carry the line’s best sequential read rating at 14,800 MB/s and its best random write rating at 450K IOPS, while the 30.72TB flagship steps down to 13,500 MB/s and 270K IOPS. The two smallest capacities also ship on the prior BiCS generation 5 NAND rather than generation 8. In other words, the 7.68TB drive on our bench is the configuration where this platform shows its full hand, and buyers chasing maximum density at 30.72TB give some of that back, which is not uncommon.

The generational step from the CD8P is where Kioxia is making its case, and the published spec tables back it up. At 7.68TB, the outgoing CD8P-R E3.S rated at 200K random write IOPS; the CD9P-R lifts that to 450K, a 2.25x improvement. Sequential read climbs 23% from 12,000 MB/s to 14,800 MB/s, and random read moves 30% from 2 million to 2.6 million IOPS. The active power rating rises slightly, from 21W typical to 23W at this capacity, putting the CD9P-R in line with the rest of the Gen5 read-intensive class. The efficiency argument here concerns what the drive delivers within that envelope.

KIOXIA CD9P-R rear side view.

The fine print that defines where this drive belongs: the CD9P-R features a single-port, 1 DWPD design. There is no dual-port path for traditional enterprise storage arrays, and write-heavy workloads belong on the CD9P-V mixed-use sibling. This is a drive built for hyperscale and cloud server fleets, OLTP read tiers, content delivery, and virtualized environments where the access pattern is read-dominated, and the power budget is fixed. It checks the expected platform boxes along the way, with OCP Datacenter NVMe SSD v2.5 support (not all requirements), power loss protection, end-to-end data protection, SIE and SED security options, a 2.5 million-hour MTTF at 50°C, and a five-year warranty.

KIOXIA CD9P-R Specifications

The table below outlines the KIOXIA CD9P-R Series in the E3.S form factor across its capacity points, highlighting performance metrics, endurance ratings, power, and reliability specifications.

KIOXIA CD9P-R Series Specifications (E3.S)
30.72TB 15.36TB 7.68TB 3.84TB 1.92TB
Model Numbers
SIE Model Number KCD9XPJE30T7 KCD9XPJE15T3 KCD9XPJE7T68 KCD9XPJE3T84 KCD9XPJE1T92
SED Model Number KCD9DPJE30T7 KCD9DPJE15T3 KCD9DPJE7T68 KCD9DPJE3T84 KCD9DPJE1T92
Basic Specifications
Use Case Read Intensive (1 Drive Write Per Day)
Form Factor E3.S, 7.5mm thickness
Interface / Protocol PCIe 5.0 x4, NVMe 2.0
Maximum Interface Speed 128 GT/s (PCIe Gen5 x4)
NAND KIOXIA BiCS FLASH 3D TLC (Gen 8 for 7.68TB-30.72TB; Gen 5 for 1.92TB-3.84TB)
OCP Compliance OCP Datacenter NVMe SSD Specification v2.5 (partial)
Security SIE (Sanitize Instant Erase), SED (TCG Opal & Ruby SSC)
Performance (Up To)
Sequential Read (128KiB, MB/s) 13,500 14,800 14,800 14,500 14,500
Sequential Write (128KiB, MB/s) 7,000 7,000 7,000 7,000 3,600
Random Read (4KiB, K IOPS) 2,600 2,600 2,600 2,600 2,000
Random Write (4KiB, K IOPS) 270 450 450 320 160
Power Requirements
Supply Voltage 12V ±10%, 3.3V ±15%
Power (Active) 23W typ.
Power (Ready/Idle) 5W typ.
Reliability
MTTF 2,500,000 hours @ 0–50°C | 2,000,000 hours @ 0–55°C
UBER < 1 sector per 1017 bits read
DWPD 1
Warranty 5 Years
Data Protection Power Loss Protection (PLP), End-to-End Data Protection
Dimensions
Thickness 7.5mm +0.2 / -0.5mm
Width 76mm ±0.25mm
Length 112.75mm ±0.4mm
Weight 110g max
Environmental
Temperature (Operating) 0°C to 75°C
Temperature (Non-operating) -40°C to 85°C
Humidity (Operating) 5% to 95% RH
Vibration (Operating) 21.27 m/s² { 2.17 Grms } (5–800 Hz)
Shock (Operating) 9.8 km/s² { 1,000 G } (0.5 ms)

KIOXIA CD9P-R Performance

KIOXIA CD9P-R E3S connector side view

Drive Testing Platform

We use a Dell PowerEdge R760 running Ubuntu 22.04.2 LTS as our test platform for all workloads in this review. Equipped with a Serial Cables Gen5 JBOF, it offers wide compatibility with U.2, E1.S, E3.S, and M.2 SSDs. Our system configuration is outlined below:

  • 2 x Intel Xeon Gold 6430 (32-Core, 2.1GHz)
  • 16 x 64GB DDR5-4400
  • 480GB Dell BOSS SSD
  • Serial Cables Gen5 JBOF
  • NVIDIA L4

Drives Compared

DLIO Checkpointing Benchmark

To evaluate SSD real-world performance in AI training environments, we utilized the Data and Learning Input/Output (DLIO) benchmark tool. Developed by Argonne National Laboratory, DLIO is specifically designed to test I/O patterns in deep learning workloads. It provides insights into how storage systems handle challenges such as checkpointing, data ingestion, and model training.

The chart below illustrates how the drives handle the process across 18 checkpoints. When training machine learning models, checkpoints are essential for periodically saving the model’s state, preventing loss of progress during interruptions or power failures. This storage demand requires robust performance, especially under sustained or intensive workloads. We used DLIO benchmark version 2.0 from the August 13, 2024, release.

To ensure our benchmarking reflected real-world scenarios, we based our testing on the LLAMA 3.1 405B model architecture. We implemented checkpointing using torch.save() to capture model parameters, optimizer states, and layer states. Our setup simulated an eight-GPU system, implementing a hybrid parallelism strategy with 4-way tensor parallelism and 2-way pipeline parallel processing distributed across the eight GPUs. This configuration yielded a checkpoint size of 1,636GB, reflecting the requirements of training modern large language models.

Looking at the pass averages, the KIOXIA CD9P-R started at 464.7 seconds in Pass 1 before increasing to 575.6 seconds in Pass 2 and settling at 572.2 seconds in Pass 3. This behavior closely mirrored the majority of the comparison group, which clustered between roughly 553 and 590 seconds by the final pass. The standout outlier was the Pascari X200P, which finished substantially higher at 674.5 seconds. Overall, the CD9P-R demonstrated predictable scaling across repeated checkpoint operations and remained competitive with the mainstream enterprise Gen5 SSDs in the test.

For the DLIO Checkpoint Benchmark through checkpoint 12, the KIOXIA CD9P-R 7.68TB remained one of the more consistent drives in the comparison group. After starting at 471.4 seconds at the first checkpoint, it settled into a relatively narrow operating range of roughly 560 to 580 seconds for the remainder of the test, finishing checkpoint 12 at 569.7 seconds. The KIOXIA drive closely tracked the Solidigm PS1010, Micron 7600 MAX, and Kingston DC3000ME throughout most of the workload.

The Pascari X200P was the clear outlier, jumping sharply after checkpoint 4 and remaining well above the field, reaching nearly 690 seconds by checkpoint 12. The Micron 9550 MAX showed the lowest sustained checkpoint times during the latter half of the run, dipping as low as 531.3 seconds before ending at 569.1 seconds. While the CD9P-R was not the fastest drive at any individual checkpoint, it avoided the large swings seen from several competitors and delivered stable checkpoint performance across the full test window.

FIO Performance Benchmark

To measure the storage performance of each SSD across common industry metrics, we leverage FIO. Each drive undergoes the same testing process, which includes a preconditioning step of two full drive fills with a sequential write workload, followed by steady-state performance measurement. As each workload type being measured changes, we run another preconditioning fill of that new transfer size.

In this section, we focus on the following FIO benchmarks:

  • 128K Sequential
  • 64K Random
  • 16K Random
  • 4K Random

128K Sequential Write (IODepth 16 / NumJobs 1)

Moving to the steady-state 128K Sequential Write test at a lower IODepth of 16, the overall group ranking remained largely unchanged compared to preconditioning. The Micron 9550 Max (12.8TB) continued to lead at 10,957.9 MB/s, with the Micron 9550 Pro (7.68TB) close behind at 10,354.6 MB/s. The Kingston DC3000ME (7.68TB) held third at 8,477.4 MB/s, and the Pascari X200P (7.68TB) was right behind at 8,369.7 MB/s.

The KIOXIA CD9P-R (7.68TB) delivered 6,912.4 MB/s, landing at the back of the field. The Solidigm PS1010 (7.68TB) at 7,126.5 MB/s and the SanDisk DC SN861 (7.68TB) at 7,116.5 MB/s both trailed the mid-pack drives, but still edged out the CD9P-R and the Micron 7600 Max (6.4TB) at 6,960.6 MB/s. The KIOXIA result is consistent and predictable for a read-optimized NVMe drive.

128K Sequential Write Latency (IODepth 16 / NumJobs 1)

At an IODepth of 16 for the steady-state write test, latency dropped substantially across all drives compared to preconditioning conditions. The Micron 9550 Max (12.8TB) again led with the lowest mean latency at 182.2 µs, comfortably ahead of the Micron 9550 Pro (7.68TB) at 192.9 µs, with both drives benefiting from their higher write throughput to service IOs more efficiently.

The KIOXIA CD9P-R (7.68TB) posted 289.0 µs, the highest write latency in the group at this queue depth. The Solidigm PS1010 (7.68TB) at 280.3 µs and the SanDisk DC SN861 (7.68TB) at 280.7 µs were just below the CD9P-R, while the Micron 7600 Max (6.4TB) came in at 287.1 µs. The Kingston DC3000ME (7.68TB) and Pascari X200P (7.68TB) occupied the middle tier at 235.6 µs and 238.6 µs, respectively.

128K Sequential Read (IODepth 64 / NumJobs 1)

The 128K Sequential Read test produced a complete reversal of the write rankings, and the KIOXIA CD9P-R (7.68TB) came out as one of the top performers in the group. The CD9P-R delivered 14,235.9 MB/s, effectively tying with the Pascari X200P (7.68TB) at 14,242.1 MB/s at the top of the chart. The Solidigm PS1010 (7.68TB) at 14,163.3 MB/s, the Micron 9550 Pro (7.68TB) at 14,050.1 MB/s, and the Micron 9550 Max (12.8TB) at 14,047.5 MB/s all clustered tightly in a 200 MB/s band at the top.

The Kingston DC3000ME (7.68TB) trailed the leaders at 13,513.8 MB/s, and the SanDisk DC SN861 (7.68TB) came in at 12,631.2 MB/s. The Micron 7600 Max (6.4TB) at 11,240.5 MB/s was the only drive to fall below the 12 GB/s threshold.

128K Sequential Read latency (IODepth 64 / NumJobs 1)

The 128K Sequential Read latency results closely mirror the bandwidth outcome. The Pascari X200P (7.68TB) led with 561.4 µs, with the KIOXIA CD9P-R (7.68TB) essentially matched at 561.7 µs. The Solidigm PS1010 (7.68TB) at 564.5 µs, Micron 9550 Pro (7.68TB) at 569.0 µs, and Micron 9550 Max (12.8TB) at 569.1 µs all fell within an 8 µs window of the leader, confirming that this tier of drives is constrained by Gen5 interface bandwidth rather than internal latency.

The Kingston DC3000ME (7.68TB) followed at 591.6 µs and the SanDisk DC SN861 (7.68TB) at 633.0 µs, while the Micron 7600 Max (6.4TB) at 711.4 µs posted latency that was 26% higher than the top performers, consistent with its lower sequential read throughput.

 

64K Random Write

Across the full 64K Random Write sweep, the KIOXIA CD9P-R (7.68TB) delivered a consistently respectable bandwidth profile, averaging in the 3-6 GB/s range and reaching a peak of 6,906 MB/s at the highest queue depths tested (IODepth 32 / NumJobs 8). This positioned the CD9P-R in the middle of the field for 64K write throughput, clearly behind the Micron 9550 Max (12.8TB), which scaled to 10+ GB/s peaks, but ahead of the Solidigm PS1010 (7.68TB) and SanDisk DC SN861 (7.68TB), which lagged in the lower half of the chart. The Micron 7600 Max (6.4TB) tracked closely, reaching a similar ceiling and ending just above the CD9P-R.

64K Random Write Latency

The 64K Random Write latency sweep for the KIOXIA CD9P-R (7.68TB) showed us a fairly balanced drive. At low queue depths, latency was well-controlled, starting in the sub-100 µs range at IODepth 1 / NumJobs 1. As concurrency increased, latency rose gradually across the 300-700 µs range for most of the mid-depth sweep, then climbed further at peak queue depths, reaching into the low 2,000 µs range. This placed the CD9P-R in the middle of the group for most of the sweep, performing more predictably than the Solidigm PS1010 (7.68TB) and Pascari X200P (7.68TB), which had sharper spikes in the 4,000-6,000 µs range at high concurrency.

The Micron 9550 Max (12.8TB) maintained the most consistent latency across the sweep, rarely exceeding 1,700 µs even at peak depths, while the Micron 7600 Max (6.4TB) and Micron 9550 Pro (7.68TB) tracked nearby.

64K Random Read

The 64K Random Read sweep revealed one of the more distinctive benefits of the KIOXIA CD9P-R (7.68TB): exceptional performance at low queue depths. At the low end of the sweep, IODepth 1 / NumJobs 1, the CD9P-R started at approximately 1,334 MB/s, leading by a wide margin thanks to its extremely low per-IO read latency. This advantage persisted across the lower queue-depth range, as it consistently ran at or near the top.

As queue depth increased into the 8/4-32/4 range, other drives caught up and surpassed the CD9P-R. At the highest concurrency levels, the CD9P-R stabilized around 11.0-12.0 GB/s, placing it behind the Pascari X200P (7.68TB), Micron 9550 Pro (7.68TB), and Micron 9550 Max (12.8TB), which reached 13.5-14.2 GB/s.

64K Random Read Latency

The 64K Random Read latency sweep confirmed the KIOXIA CD9P-R’s (7.68TB) read latency advantage at low queue depths. From IODepth 1 / NumJobs 1 through the low-concurrency range, the CD9P-R consistently posted some of the tightest latency in the group, tracking well below most competitors and within a narrow band through mid-queue depths.

As queue depth increased beyond the 32/1-to-16/4 range, the group converged, and by the end of the sweep, all drives had climbed into the 600 µs to 1,400 µs range.

16K Random Write

Across the 16K Random Write IOPS sweep, the KIOXIA CD9P-R (7.68TB) delivered a consistent performance profile through the mid-range of the sweep. The Micron 9550 Max (12.8TB) again dominated, sustaining a highly elevated IOPS trajectory well above the field, often approaching 600-690K IOPS at higher queue depths, while the Micron 7600 Max (6.4TB) maintained strong throughput in the 400-450K range.

The CD9P-R tracked alongside the Kingston DC3000ME (7.68TB), Micron 9550 Pro (7.68TB), and SanDisk DC SN861 (7.68TB) through much of the sweep, generally ranging between 200-250K IOPS at moderate depths and scaling toward 365-440K IOPS at the highest tested concurrency (IODepth 32 / NumJobs 16). The Solidigm PS1010 (7.68TB) was the weakest performer in 16K random write, frequently trailing all other drives.

16K Random Write Latency

In the 16K Random Write latency sweep, the KIOXIA CD9P-R (7.68TB) consistently tracked in the lower-to-middle tier across most of the queue depth range. At low queue depths, the drive began under 50 µs and maintained stable, well-controlled latency through the moderate portion of the sweep. As queue depth increased into the 8/8-32/8 range, latency climbed more steeply across all drives, and the CD9P-R moved through the 500-750 µs range before peaking at approximately 1,165 µs.

The Micron 9550 Max (12.8TB) posted the most stable latency across the sweep, holding below the field at most data points. The Solidigm PS1010 (7.68TB) and Pascari X200P (7.68TB) showed the most pronounced latency spikes at high queue depths, peaking at 3,300 µs and 2,050 µs, respectively. In contrast, the CD9P-R, Micron 7600 Max (6.4TB), and Kingston DC3000ME (7.68TB) tracked more predictably through the upper concurrency range. The CD9P-R’s 16K write latency is one of the most consistent in the group under heavy mixed-parallelism loads.

 

 

16K Random Read

In the 16K Random Read IOPS sweep, the Pascari X200P (7.68TB) and Micron 9550 Max (12.8TB) posted the highest sustained read IOPS, approaching 900K at saturation. The Micron 9550 Pro (7.68TB) followed closely with similar scaling, while the Solidigm PS1010 (7.68TB) rounded out the top tier.

The KIOXIA CD9P-R (7.68TB) delivered 734.3K IOPS at its peak measured point (IODepth 32 / NumJobs 8), placing it in the upper-middle tier, ahead of the Micron 7600 Max (6.4TB) at 719.3K, Kingston DC3000ME (7.68TB) at 665.6K, and SanDisk DC SN861 (7.68TB) at 661.1K IOPS.

16K Random Read Latency

In the 16K Random Read latency sweep, the KIOXIA CD9P-R (7.68TB) showed excellent read latency across the majority of the tested queue depth range. From the low end of the sweep, the CD9P-R began at approximately 33 µs at IODepth 1 / NumJobs 1. It remained tightly grouped with the top performers through the moderate-concurrency portion of the sweep, generally staying below 100 µs.

As queue depth increased into the highest concurrency ranges, all drives saw latency climb sharply. The CD9P-R scaled to approximately 713 µs at the peak, landing behind the SanDisk DC SN861 (7.68TB) and Kingston DC3000ME (7.68TB), which exceeded 820-845 µs, while the Micron 9550 Max (12.8TB) and Micron 9550 Pro (7.68TB) maintained a lower peak latency.

4K Random Write

In the 4K Random Write IOPS sweep, the Micron 7600 Max (6.4TB) and Micron 9550 Max (12.8TB) held the top positions through most of the sweep, sustaining 1,500-1,780K IOPS at peak concurrency. The Micron 9550 Pro (7.68TB) and Pascari X200P (7.68TB) followed in the upper tier, while the Solidigm PS1010 (7.68TB) and SanDisk DC SN861 (7.68TB) traded positions in the mid-upper range.

The KIOXIA CD9P-R (7.68TB) reached 1,263.1K IOPS at IODepth 32 / NumJobs 16, placing it in the lower-middle tier for 4K random write performance.

4K Random Write Latency

The 4K Random Write latency sweep produced one of the more variable pictures in the group, and the KIOXIA CD9P-R (7.68TB) again occupied the middle ground. At low queue depths, the CD9P-R’s write latency started in the 9-13 µs range at IODepth 1, competitive but slightly above the best performers at those settings. Through the moderate portion of the sweep, the drive held in the 20-120 µs range alongside most of the group.

At the highest queue depths, the CD9P-R climbed to 200-410 µs, staying well below the Solidigm PS1010 (7.68TB), which peaked near 740 µs and showed significant volatility, and the Pascari X200P (7.68TB) at 544 µs. The Micron 9550 Max (12.8TB), Micron 7600 Max (6.4TB), and Micron 9550 Pro (7.68TB) posted the most controlled latency profiles across the full sweep.

4K Random Read

The 4K Random Read sweep delivered one of the more interesting results in our FIO tests for the KIOXIA CD9P-R (7.68TB). At the lowest queue depths, IODepth 1 / NumJobs 1 and IODepth 2 / NumJobs 1, the CD9P-R separated itself clearly from the field, producing approximately 32.3K IOPS (126 MB/s). As concurrency increased, the rest of the group caught up, and by the peak of the sweep, the rankings shifted. The SanDisk DC SN861 (7.68TB) led in concurrency with 2,555.6K IOPS. At the same time, the KIOXIA CD9P-R reached 2,165.0K IOPS at IODepth 32 / NumJobs 16, effectively tying with the Micron 9550 Max (12.8TB) at 2,164.8K IOPS for second place in the group. The Solidigm PS1010 (7.68TB) and Pascari X200P (7.68TB) followed, with the Micron 7600 Max (6.4TB) and Kingston DC3000ME (7.68TB) at the lower end.

4K Random Read Latency

The 4K Random Read latency sweep confirmed what the IOPS chart suggested: the KIOXIA CD9P-R (7.68TB) has the lowest measured read latency at low queue depths among the drives we tested in the comparison group. Starting at approximately 30 µs with IODepth 1 / NumJobs 1, the CD9P-R had nearly half the latency of most competing drives, clustered in the 60-90 µs range at the same data point. Through the moderate portion of the sweep, the CD9P-R maintained its latency lead, tracking below the group through IODepth 8 and 16 before converging with the pack as concurrency climbed. At the highest tested queue depths, all drives moved into the 120-275 µs range, with the SanDisk DC SN861 (7.68TB) posting the highest latency at roughly 200 µs. The KIOXIA CD9P-R ended the sweep at approximately 236 µs, putting it midway in the group.

GPU Direct Storage

One of the tests we conducted on this testbench was the Magnum IO GPU Direct Storage (GDS) test. GDS is a feature developed by NVIDIA that allows GPUs to bypass the CPU when accessing data stored on NVMe drives or other high-speed storage devices. Instead of routing data through the CPU and system memory, GDS enables direct communication between the GPU and the storage device, significantly reducing latency and improving data throughput.

How GPU Direct Storage Works

Traditionally, when a GPU processes data stored on an NVMe drive, the data must first travel through the CPU and system memory before reaching the GPU. This process introduces bottlenecks, as the CPU acts as an intermediary, adding latency and consuming valuable system resources. GPU Direct Storage eliminates this inefficiency by enabling the GPU to access data directly from the storage device via the PCIe bus. This direct path reduces data-movement overhead, enabling faster, more efficient data transfers.

AI workloads, especially those involving deep learning, are highly data-intensive. Training large neural networks requires processing terabytes of data, and any delay in data transfer can lead to underutilized GPUs and longer training times. GPU Direct Storage addresses this challenge by ensuring that data is delivered to the GPU as quickly as possible, minimizing idle time and maximizing computational efficiency.

In addition, GDS is particularly beneficial for workloads that involve streaming large datasets, such as video processing, natural language processing, or real-time inference. By reducing the reliance on the CPU, GDS accelerates data movement and frees up CPU resources for other tasks, further enhancing overall system performance.

GDSIO Sequential Read Throughput

The GDSIO Sequential Read Throughput sweep highlighted the strong read performance of the KIOXIA CD9P-R (7.68TB). Across all three block size segments, the CD9P-R emerged as the highest-throughput drive at most thread counts.

In the 16K block-size segment, the CD9P-R opened at approximately 0.3 GiB/s on 16K/1 thread, starting behind most of the field. Drives like the Pascari X200P (7.68TB) and Micron 7600 Max (6.4TB) entered the segment at roughly 0.55-0.6 GiB/s on a single thread. As the thread count increased, the CD9P-R scaled aggressively, pulling ahead of the group through the mid-range. By 16K/16 it had overtaken most competitors, and by 16K/64 it reached approximately 2.0 GiB/s, the highest throughput in the group at that point. At 16K/128, the CD9P-R held at roughly 1.9 GiB/s and remained among the top performers as several other drives plateaued or declined. The CD9P-R’s 16K read throughput advantage is a product of how efficiently it scales with thread count, not single-stream dominance.

Moving into the 128K block-size segment, the CD9P-R opened at approximately 1.6 GiB/s at 128K/1, near the top of the group at that thread count, alongside the SanDisk DC SN861 (7.68TB). It scaled cleanly through the moderate thread range, reaching approximately 3.9 GiB/s at 128K/8, the highest throughput in the group at that point, where most competitors were still in the 1.7 to 2.0 GiB/s range. Through 128K/16 and 128K/32, the CD9P-R continued leading the field as other drives scaled up. By 128K/64 and 128K/128, the Micron 9550 Pro (7.68TB) and Micron 9550 Max (12.8TB) caught and passed the CD9P-R, reaching approximately 5.2 to 5.3 GiB/s, while the CD9P-R peaked at approximately 5.1 GiB/s and then pulled back to roughly 4.8 GiB/s at 128K/128.

In the 1M block-size segment, the CD9P-R opened at approximately 3.9 GiB/s at 1M/1 and scaled steadily with each subsequent thread count. At 1M/128, it reached approximately 6.2 GiB/s, the highest peak read throughput in the comparison group. The Pascari X200P (7.68TB) and Micron 9550 Max (12.8TB) were close behind at approximately 6.1 GiB/s each, followed by the Solidigm PS1010 (7.68TB) and Micron 9550 Pro (7.68TB) at roughly 6.0 GiB/s. The Kingston DC3000ME (7.68TB) reached approximately 5.9 GiB/s, while the Micron 7600 Max (6.4TB) was the clear laggard at approximately 5.6 GiB/s.

GDSIO Sequential Read Latency

The GDSIO Sequential Read Latency chart showed that all eight drives followed nearly identical trajectories throughout the full sweep. Latency values are dominated by block size and thread count rather than drive-specific characteristics, making this chart less differentiating than the throughput and IOPS results.

At the single-thread baseline (16K block, 1 thread), the CD9P-R posted approximately 44 µs, placing it in the middle of the group. Faster drives at this configuration included the Pascari X200P (7.68TB) at approximately 26 µs, the SanDisk DC SN861 (7.68TB) at approximately 26 µs, and the Micron 7600 Max (6.4TB) at approximately 27 µs. At the other end, the Kingston DC3000ME (7.68TB) posted approximately 83 µs, and the Solidigm PS1010 (7.68TB) approximately 71 µs, by far the widest single-thread read latency outliers in the group.

Through the rest of the sweep, the lines converged. All drives reached the 20,000-22,000 µs range by 1M/128, with the CD9P-R at approximately 20.3 ms landing in the lower portion of the group, consistent with its higher throughput at that configuration.

GDSIO Sequential Write Throughput

The GDSIO Sequential Write Throughput sweep revealed notable differences across the group.

The Solidigm PS1010 (7.68TB) exhibited a dramatic collapse in write throughput at high thread counts in the 128K block-size segment. After delivering a competitive 3.9 GiB/s at 128K/32, the drive fell sharply to approximately 2.5 GiB/s at 128K/64 and then to just 1.6 GiB/s at 128K/128, less than half the throughput of most competing drives at that configuration. The degradation continued into the 1M segment, where the drive’s peak reached only 4.2 GiB/s with moderate thread counts, then declined again at the high end. This sustained collapse under heavy multi-thread write load stands out as the most significant anomaly in the dataset. The Micron 9550 Max (12.8TB) also showed a sharp throughput drop in the 1M segment, falling from approximately 4.7 GiB/s at 1M/32 to roughly 2.2 GiB/s at 1M/64 before a partial recovery at 1M/128.

In the 16K block-size segment, all drives tracked closely within the 0.6 to 1.5 GiB/s range, with differences of only a few tenths of GiB/s across most thread counts. The CD9P-R opened at approximately 0.7 GiB/s at 16K/1, scaled to approximately 1.5 GiB/s at 16K/32, and then leveled off. The Micron 9550 Max (12.8TB) was the outlier, dropping abruptly to approximately 0.7 GiB/s at 16K/128 while the rest of the group held in the 1.2 to 1.4 GiB/s range.

In the 128K segment, the CD9P-R was among the strongest performers at low thread counts. At 128K/1, it entered at approximately 2.7 GiB/s near the top of the group, and scaled to approximately 4.7 GiB/s at 128K/8, the highest in the group at that point. From 128K/16 onward, the Micron 9550 Pro (7.68TB) and Micron 9550 Max (12.8TB) scaled past the CD9P-R, reaching approximately 5.1 to 5.3 GiB/s at their respective peaks. The CD9P-R held at roughly 4.1-4.5 GiB/s for the remainder of the 128K segment.

In the 1M-block-size segment, the CD9P-R entered the segment at approximately 4.9 GiB/s at 1M/1, the highest throughput in the group at that single-thread configuration. Still, it remained essentially flat at approximately 4.3-4.4 GiB/s across all subsequent thread counts, rather than continuing to scale. Most other drives used additional threads to climb further: the Micron 9550 Pro (7.68TB), Micron 7600 Max (6.4TB), and Pascari X200P (7.68TB) all scaled past the CD9P-R at 1M/4 and higher, reaching peak throughput of approximately 5.2 to 5.7 GiB/s. The CD9P-R’s peak write throughput of 4.9 GiB/s ranked fifth among the eight drives, behind the Micron 9550 Max at 5.69 GiB/s, Micron 9550 Pro at 5.54 GiB/s, Micron 7600 Max at 5.44 GiB/s, and Pascari X200P at 5.41 GiB/s.

GDSIO Sequential Write Latency

In the 16K segment, all drives tracked closely at low thread counts. At 16K/1, the CD9P-R posted approximately 21 µs, the lowest write latency in the group at that configuration, just below the Solidigm PS1010 (7.68TB) at approximately 22 µs. The Pascari X200P (7.68TB) followed at approximately 25 µs, while the Micron 9550 Max (12.8TB) and Micron 7600 Max (6.4TB) both came in around 30 µs. At 16K/128, the Micron 9550 Max spiked to approximately 2,500 µs, the highest write latency in the group at that configuration, while the CD9P-R held at approximately 1,500 µs and remained among the lower values in the group.

In the 128K segment, the Solidigm PS1010’s throughput collapse appeared in the latency chart as a sharp spike to approximately 9,700 µs at 128K/64, the most visible anomaly in that portion of the sweep. The CD9P-R tracked with the main group throughout the 128K segment, remaining in the 1,900 to 3,800 µs range across 128K/64 and 128K/128.

In the 1M segment, two drives showed elevated latency spikes at 1M/64 consistent with their throughput degradation at that configuration. The Micron 9550 Max (12.8TB) and Micron 9550 Pro (7.68TB) each reached approximately 27,000-28,000 µs at 1M/64, then climbed further at 1M/128. The CD9P-R reached approximately 14,500 µs at 1M/64, reflecting its more stable write throughput in that configuration, and ended the sweep at approximately 29,100 µs at 1M/128. The Pascari X200P (7.68TB) delivered the lowest write latency at 1M/128 at approximately 25,000 µs. The highest values at 1M/128 were observed for the Micron 9550 Pro at approximately 44,900 µs and the Solidigm PS1010 at approximately 40,700 µs, with the Micron 7600 Max following at approximately 36,000 µs.

Conclusion

The KIOXIA CD9P-R E3.S 7.68TB does exactly what a read-intensive drive should, and the test data tracks that brief from start to finish. Sequential reads landed near the top of the group at 14,235.9 MB/s in our 128K test, effectively tying the Pascari X200P, and the drive’s real signature showed up at low queue depths: roughly 30 µs of 4K random read latency at QD1, close to half the 60 to 90 µs posted by most of the field. That low-latency read behavior carried over to GPU Direct Storage, where the CD9P-R scaled cleanly with thread count and achieved the group’s highest 1M read throughput of approximately 6.2 GiB/s.

KIOXIA CD9P-R top view.

The trade-offs are just as clear and align with the drive’s 1 DWPD design rather than working against it. Write performance sat mid-pack, with 128K sequential writes of 6,912.4 MB/s, landing at the back of the group, and GDSIO sequential write throughput peaking at 4.9 GiB/s, placing fifth among the eight drives. The Micron 9550 family led those write workloads in both throughput and latency. None of that is a knock on the CD9P-R; it is a read-intensive SSD, and buyers who need sustained write performance should be looking at the mixed-use CD9P-V instead.

Consistency under load is the other part of the story. In DLIO checkpointing against the LLAMA 3.1 405B profile, the CD9P-R settled around 572 seconds per pass and stayed inside the 553 to 590 second band that defined the mainstream Gen5 field, avoiding the swings that pushed the Pascari X200P out to 674.5 seconds. The platform also scales well beyond our sample: the CD9P-R family runs to 30.72TB in E3.S and 61.44TB in the 2.5-inch form factor on the same architecture, which lets a single qualification cover everything from performance compute nodes to high-density read tiers.

For cloud fleets, AI data pipelines, content delivery, and virtualized read tiers with read-dominated access patterns, the CD9P-R is an easy drive to recommend. It pairs top-tier Gen5 read throughput with the lowest low-QD read latency in our comparison group and a deep capacity stack, and it holds steady under sustained load.

Product Page – KIOXIA CD9P-R 7.68TB

The post Kioxia CD9P-R Review: Read-Intensive Gen5 Up to 61.44TB appeared first on StorageReview.com.

HPE Alletra Storage MP B10000 and NIST CSF 2.0: A Full-Stack Cyber Resilience Architecture

12 June 2026 at 15:53

HPE has built a coordinated cyber resilience architecture around the Alletra Storage MP B10000. It extends the platform’s native security capabilities through an integrated stack that includes virtualization with Morpheus and VM Essentials, continuous data protection with Zerto, long-term backup retention with StoreOnce, and observability via vendor-agnostic security information and event management (SIEM) integration. Taken together, the architecture is designed to align directly with the NIST Cybersecurity Framework 2.0, with the B10000 serving as the operational center across every function of the framework. The whole design exists to answer the two questions a security practitioner will ultimately ask about any piece of infrastructure: Is it still under our control, and is the data on it still being protected?

HPE B10000 cyber resilience B10000 stack

Enterprise storage historically sits outside the broader security conversation. In most organizations, a chief information security officer who is worried about a misconfigured router in a branch office or an unpatched laptop on the corporate network pays little attention to the storage array in the data center. The array is behind layers of perimeter security, accessible to a small handful of administrators, and largely invisible to the broader IT organization. On paper, it is one of the safest assets in the building.

That assumption no longer holds. Modern ransomware operations have learned that the array is the highest-value target in the data center. Endpoints and servers can be reimaged. Primary storage is where the data actually lives, and an attacker who gains administrative control of the array can encrypt the data, delete the snapshots meant to recover it, and destroy the backups in a single coordinated motion. At that point, the organization is not dealing with an inconvenience; it is negotiating for its survival. The tools available to attackers, including AI-assisted variants that adapt faster than signature-based defenses can keep up with, are increasingly capable of reaching that target.

The stakes are no longer only operational. Regulators in the United States, the European Union, and the United Kingdom have moved infrastructure security from a best practice toward a legal obligation. Frameworks such as the EU’s Digital Operational Resilience Act and the NIS2 Directive require demonstrable controls for detection, recovery, and incident reporting, with accountability that extends to the executive level. The organizations carrying these obligations are exactly the ones running enterprise storage at scale: banks and financial services firms, hospitals and healthcare networks, utilities and critical infrastructure operators, government agencies, and the cloud and service providers that host all of the above. For them, failing to secure the infrastructure layer is not only a risk to the business but also a compliance failure with serious consequences.

The work of cyber resilience sits in the gap between storage and security teams that report to different leaders, measure success differently, and rarely share the operational language needed to coordinate during an incident. Storage administrators understand throughput, capacity, and recovery objectives. Security teams understand kill chains, attack vectors, and posture management. Most organizations discover the gap only after they have been forced to operate in it. The architecture HPE has assembled is designed to close that gap, and the sections that follow work through each NIST function to test how well it does, returning throughout to the two questions of control and protection.

Key Takeaways

  • Full-stack resilience: HPE pairs the Alletra Storage MP B10000 with VM Essentials, Zerto, StoreOnce, and vendor-agnostic SIEM integration, mapping the coordinated stack to every NIST CSF 2.0 function.
  • Fast block-level detection: The B10000’s built-in entropy-based ransomware engine, validated against 100+ strains, flagged a simulated encryption run in 4-5 minutes and auto-captured a forensic snapshot at the moment of detection.
  • Immutability that survives admin compromise: Array-enforced Virtual Lock snapshots stay read-only through retention; even a compromised admin account can’t delete them, and any attempt becomes a logged, SIEM-visible event.
  • Tiered recovery: Four tiers (Zerto for lowest RPO, VME snapshots, Virtual Lock promote, and StoreOnce Catalyst for long-term retention) were exercised in the lab, with clean restores from both Virtual Lock and Catalyst.
  • Compliance-ready telemetry: Structured security syslog normalized to Elastic Common Schema feeds Elastic Security, CrowdStrike Falcon, or any modern SIEM, supporting DORA/NIS2 evidence and automated SOAR playbooks.

 Architecture Overview

To provide a practical demonstration environment rather than an oversized enterprise deployment, the HPE team built out a compact but functional recovery and backup stack in their Fort Collins lab. At the center of the environment is a three-node HPE VM Essentials (VME) cluster running on HPE ProLiant DL325 Gen 11 Servers, which provides the compute layer for the demo infrastructure. These hosts are interconnected through the local IP network, which also provides connectivity to the NAS layer used within the environment.

Running inside the VME cluster are the HPE Zerto virtual machines, configured similarly to how many organizations currently deploy Zerto in a traditional VMware environment. While the underlying infrastructure here uses HPE VME, the operational flow and recovery functionality remain familiar for administrators experienced with VMware-based disaster recovery workflows.

HPE B10000 cyber resilience ProLiant servers

A WAN-connected secondary Zerto environment is available in the Bristol lab in the UK, with firewall rules and network paths in place to support replication from Fort Collins. The intent of this topology is multi-site recovery: continuous replication and disaster-recovery orchestration across geographically separated sites. This would allow workloads protected in Fort Collins to fail over to Bristol if the primary site became unavailable.

During this engagement, HPE ran Zerto replication locally in Fort Collins due to bandwidth and distance constraints, and because a second parallel VME cluster was not available to exercise the full cross-site flow. The Bristol leg is documented here as configured and available infrastructure, not as a cross-site failover performed in this session.

HPE B10000 cyber resilience stack rear

Storage connectivity in the environment is split between IP and Fibre Channel networking, depending on the workload. The FC SAN fabric connects the VME hosts directly to the HPE Alletra MP B10000 platform, providing shared access to enterprise storage across the cluster. The B10000 also presents Catalyst over Fibre Channel to the StoreOnce Virtual Storage Appliance (VSA), which is the path used for application-consistent backups between the array and the backup target. Virtual Lock snapshots play a central role in the ransomware resilience and immutability workflows on the platform and are exercised in detail in the Protect and Recover sections.

For backup infrastructure, HPE deployed a StoreOnce Gen5 VSA running on a single HPE server in the lab. The VSA exposes the same core StoreOnce functionality as the dedicated appliance. Deduplication, backup target presentation, replication behavior, and integration with the broader HPE data protection stack remain fundamentally the same whether deployed as a VSA or on dedicated hardware. The VSA is well-suited for labs, branch offices, testing environments, and smaller production use cases, making this compact demonstration practical to deploy.

The physical StoreOnce appliance carries an advantage that goes beyond scale and throughput, and it matters specifically in the ransomware context this paper examines. A VSA runs as a guest on a hypervisor. If an attacker compromises the virtualization layer, every workload on it, including a VSA backup target, is within reach. A physical StoreOnce appliance has no such dependency. It runs on dedicated hardware outside the hypervisor that an attacker would have to traverse, keeping the backup of last resort on infrastructure the attacker has not already breached. For production deployments where StoreOnce is the long-term retention tier in a cyber resilience design, the physical appliance is the stronger choice for exactly this reason.

Both the HPE Alletra MP system and the StoreOnce VSA were onboarded to HPE’s Data Services Cloud Console, providing centralized, cloud-based management and visibility across the environment. Through Data Services Cloud Console, the infrastructure can be monitored, managed, and integrated into broader HPE data services workflows from a single interface, tying together storage, backup, and recovery operations.

Govern: Implementing Your Strategy

HPE does not write your governance plan. Your organization, your insurer, and your legal team arrive at the policy. HPE provides the interfaces (APIs, CLIs, Data Services Cloud Console, and the published hardening guides) that let you implement that policy, evidence it, and audit against it.

That distinction matters because most storage vendors quietly assume the opposite. They ship opinionated defaults dressed up as best practices and treat configuration drift as a support problem rather than a governance signal. The B10000 inverts that posture. Every administrative action is logged. Every policy-relevant setting is exposed through the API. The security syslog stream is structured to feed the same SIEM where the rest of the organization’s governance evidence already lives. When an auditor asks who changed the password policy on a specific date, the answer sits in the same pane of glass as every other governed system in the estate.

Two governance disciplines deserve specific call-outs. The first is regulatory alignment. The B10000’s audit logging and SIEM integration architecture supports the operational controls required by the EU’s Digital Operational Resilience Act and the NIS2 Directive, covering detection, response, and recovery activities, incident reporting workflows, and the executive-oversight visibility expected of leadership. Centralizing the audit and security telemetry in the SIEM operationalizes those obligations. The controls the auditor expects to see are not bolted on. They are the same controls that the security operations center (SOC) already uses to operate the platform day-to-day.

The second is executive oversight. Regulatory frameworks increasingly require that leadership demonstrate timely, consolidated visibility into information and communication technology (ICT) risk. A storage array that hides its audit trail behind a vendor-only support portal cannot satisfy that requirement. One that streams audit and security events directly into the customer’s SIEM can. The B10000 was designed for the second pattern.

The platform provides several concrete capabilities. Documented hardening guides define the starting posture. Role-based access controls and dual-authorization mechanisms enforce separation of duties. Multi-factor authentication is built in. Defense Information Systems Agency Security Technical Implementation Guide (DISA STIG) hardening guidance is published for environments that require it. The security syslog stream turns every governance-relevant action into a queryable event. The plan is the customer’s. Instrumentation is the platform’s responsibility.

Identify: Trust in the Platform

Identify, in NIST CSF 2.0 framing, is about understanding what the organization has, where it came from, and what risks it carries. For a storage platform, that translates to supply chain integrity, secure development practices, and a verifiable starting posture.

HPE’s hardware and software supply chain is the foundation. Components are sourced through audited channels, firmware is signed, and platforms ship with tamper-evident packaging and seals that the receiving team can verify against shipment manifests. The supply chain assessment is a documented process that the HPE Cybersecurity Center of Excellence uses to verify that a B10000 leaving the factory in one week has the same provenance as a B10000 leaving in another week.

Secure development practices extend that chain into the software. The B10000 operating system undergoes structured penetration testing, including internal red-team exercises and third-party engagements, with findings fed back into the release cycle. Identify is not a one-time activity completed at install. It is the ongoing verification that the platform trusted on day one is still the platform present on day 400. The B10000 provides the artifacts (signed firmware, hash-verified downloads, hardening guides, secure development documentation, and the audit log stream) that enable that verification. What the security team does with those artifacts is, again, the security team’s decision.

Protect: Maintaining Posture

Protection on the B10000 rests on four mechanisms: immutable Virtual Lock snapshot scheduling, a deliberate separation between array snapshots and hypervisor snapshots, replication to a second site through both array-native Remote Copy and Zerto, and continuous drift detection through the audit log stream.

Immutable Virtual Lock snapshot scheduling

Virtual Lock snapshots are created on a schedule defined at the array and enforced directly by the B10000 storage platform. The snapshots remain immutable and read-only throughout their configured retention window, preventing modification or deletion until expiration. Administrative actions against protected snapshots are recorded in the audit log. They can be forwarded to external SIEM platforms, providing visibility into attempted policy violations, including deletion attempts against protected recovery points.

B10000 snapshots versus VME snapshots

Snapshots come from two places in this environment, and policy design depends on keeping them distinct.

B10000-scheduled snapshots are immutable and read-only and are governed by Virtual Lock retention. VME snapshots, taken by HPE Morpheus VM Essentials at the hypervisor tier, are mutable and read-write for VM-level operational workflows.

If an attacker gains access to the hypervisor, VME snapshots become accessible. Virtual Lock snapshots are not. The Virtual Lock schedule defines the retention floor. VME snapshots fill in the operational layer above it.

Replication: array-native and application-level

A Virtual Lock snapshot at the primary site does not protect against site-level events, and the architecture provides two distinct replication paths to address that. The first is native to the array: B10000 Remote Copy asynchronously replicates volumes to a second B10000 at another site, and, when combined with Virtual Lock snapshot schedules at the target, results in an immutable, physically separate vault that does not depend on the hypervisor or any application-layer software. The second is Zerto, operating at the application level with near-synchronous replication and continuous journal logging, providing the granular, low recovery point objective (RPO) recovery path described in the Recover section.

The two are complementary rather than redundant. Remote Copy preserves the array-enforced immutability guarantees across sites, while Zerto provides checkpoint granularity and orchestrated VM-level failover. In this engagement, the topology was configured to support Zerto replication from Fort Collins to HPE’s Bristol, UK, facility, but during the test session, the replication ran locally in Fort Collins due to bandwidth and parallel-cluster constraints. The Bristol leg is in place as an available topology and a follow-up activity, not as a demonstrated cross-site recovery.

Configuration drift detection

The hardening guide defines a baseline. The continuous audit log stream lets a SIEM verify, in near real time, that the baseline has not drifted. Password policy changes, lockout policy changes, new accounts, role elevations, schedule modifications, and replication target changes all surface as queryable events. The lab environment did not include a SIEM, so drift detection in this engagement was observed at the B10000 audit log level rather than in a centralized correlation tool. In a customer environment with a SIEM in place, the same audit events are fed to the SIEM via the syslog integration described in the Detect section.

Detect: Visibility and Early Warning

The B10000 ships with a built-in ransomware detection engine that monitors storage I/O patterns in real time and raises an alert when the statistical signature of large-scale encryption appears in writes to the array. The HPE Cybersecurity Center of Excellence has validated the engine against more than 100 leading ransomware strains.

Detection is enabled on a per-data-store basis from the VME side. When a data store is created or edited in VME against a B10000 storage server, the Ransomware Detection toggle is enabled by default for every volume provisioned in that data store, with per-volume customization available after creation. The capability requires B10000 OS version 10.5.0 or later. In the Fort Collins lab, the ftc-vme-mg1 data store was configured to connect to the FTC-AMP-5 B10000 over Fibre Channel, with ransomware detection enabled, which is the configuration used for the simulated ransomware run later in this section.

VME data store configuration showing the ftc-vme-mg1 data store on the FTC-AMP-5 B10000 over Fibre Channel, with Ransomware Detection enabled.

The B10000 also supports system-wide ransomware detection policies directly on the storage array, allowing administrators to define global detection sensitivity and automated snapshot retention behavior across protected volumes. As shown below, the Fort Collins environment had the array-level ransomware detection policy enabled with a low-confidence sensitivity profile and automatic snapshot retention configured. This provides an additional layer of protection beyond the per-data-store enablement in VME, ensuring the detection engine remains consistently enforced at the platform level regardless of how individual volumes are provisioned.

The detection logic for both is entropy-based. Random data, which is what encrypted data looks like at the block level, has high entropy. Legitimate workload data, even when compressed, has a structure the engine recognizes. The engine uses a CuSum entropy calculation to measure the disorder in incoming writes relative to an established baseline. The baseline is learned during a training window when a new volume is first protected and is then maintained by a rolling one-hour I/O profile, so that legitimate workload shifts do not generate false positives. When entropy departs from the baseline by a statistically significant margin, the engine flags the volume.

For the SOC, the key framing is that this is an encryption-detection engine serving as a last line of defense for the environment, not a replacement for endpoint protection. The endpoint is where ransomware should be caught first. The B10000 catches what the endpoint missed.

SIEM integration

The integration is deliberately vendor-agnostic. HPE’s role is to provide the groundwork and tools so the B10000 can feed audit and security log data into whatever SIEM an organization already runs, rather than steering customers toward a single HPE-preferred tool. Detection alerts and audit events are exported through the security syslog stream. The Fluentd-based aggregator HPE publishes on GitHub takes the raw syslog message, normalizes it to Elastic Common Schema, and tags the event with fields such as “event.kind: alert” and “event.severity: high” so the receiving platform’s existing rule engine picks it up without custom parsing. The important distinction is that this data is not merely compatible with modern SIEMs; it is active. It arrives structured to trigger automated alerts and downstream responses rather than sitting in a log archive waiting to be queried after the fact. Elastic Security and CrowdStrike Falcon Next-Gen SIEM are platforms HPE has explicitly validated, and the open syslog format means Microsoft Sentinel and other modern SIEMs can be onboarded with comparable effort.

Simulated ransomware on a Windows host

During remote testing against the HPE Fort Collins lab, a simulated ransomware workload was executed against a Windows host attached to a Virtual Lock-protected volume on the B10000. The simulation ran a controlled encryption script that produced the high-entropy write pattern characteristic of a ransomware encryption event, without using actual malware.

The B10000 detection engine raised the alert within four to five minutes of the encryption script beginning to write to the volume. The alert surfaced on the B10000 management console with the affected volume identified.

The four-to-five-minute window is short enough that immutable snapshots taken at or near the moment of detection capture a clean point-in-time before the attacker has finished traversing the volume. It is faster than most behavioral analytics running at the application layer because the B10000 sees writes at the block level as soon as they occur.

Respond: From Alert to Action

When the B10000 ransomware detection engine fires, the platform takes protective action on its own, before any external system is involved. At the moment of detection, the array captures an immutable snapshot of the affected volume and marks its status as degraded. This alert snapshot is not a clean recovery point and should not be used for restoration, as it may contain encrypted data. It is intended for forensic analysis, preserving the state of the volume at the time the attack was detected. Notably, the array does not cut off access to the volume. The data remains reachable, which is a deliberate design choice. The array raises the alarm and preserves the evidence. Still, it leaves the decision of how to proceed to the operators, who are better positioned to judge whether the event is a genuine attack or a false positive, such as a large, legitimate encryption or compression workload.

This autonomous behavior is the floor, not the ceiling, of the response capability. Because it runs directly on the array, it does not depend on a SIEM rule firing, a syslog stream remaining healthy, or an external automation tool being online. If every other layer fails, the moment-in-time snapshot still exists. That independence is the point: the response that matters most happens whether or not the rest of the stack is functioning.

Above that floor, the active SIEM integration described in the Detect section enables the orchestrated response. Because the B10000’s alerts arrive structured to trigger action rather than to be archived, a customer can build automated playbooks in the SIEM or security orchestration, automation, and response (SOAR) platform they already operate, whether that means extending snapshot retention on the affected volumes, isolating the host, opening an incident ticket, or notifying the response team. HPE provides the active, structured signal. The orchestration logic resides in the customer’s chosen tooling, keeping the response aligned with the runbooks and approval workflows the security team has already built.

One clarification matters here because the interface language can be layered. The B10000 alert offers an option to analyze the data, but the array itself does not scan for or identify specific malware. That option assumes the customer has their own scanning and forensic tooling to point at the affected data. The B10000’s job is detection and preservation, not remediation.

Recover: Tiered Restoration and Validation

Recovery is a tiered set of decisions: which recovery point, from which tier, restored to which environment, validated by which process. The architecture built in the Fort Collins lab supports a four-tier recovery model, where each tier exists because the others do not solve the same problem.

Tier 1: Zerto for continuous protection. Near-synchronous replication with continuous journal logging gives the lowest RPO in the stack. Granular restore lets the incident response team roll back a single VM or file to a checkpoint from moments before the encryption event. This is the tier to reach for when the alert fires quickly and the attacker’s window is narrow.

Tier 2: VME-native protection for daily VM-level restore. The HPE Morpheus VM Essentials snapshot stream provides operational recovery at the VM granularity. This is the tier that handles routine restore work and serves as the fallback if Tier 1 is unavailable.

Tier 3: B10000 immutable Virtual Lock snapshots for volume-level rollback. This is the tier that matters most for the ransomware case. Virtual Lock snapshots are protected against deletion and modification, and the promote operation is B10000-native, with no third-party ISV in the dependency chain.

Tier 4: StoreOnce Catalyst restore for long-term retention and cross-array recovery. Tiers 1 through 3 protect against events measured in hours or days. Tier 4 protects against events measured in weeks or months, in cases where the attacker established persistence long before the encryption fired, and when the only clean recovery point is older than any primary storage snapshot retention period. Catalyst stores are deduplicated, immutable, and can be replicated across StoreOnce systems or detached to Cloud Bank object storage. Catalyst Copy and Cloud Bank Detach together provide the “1” in the 3-2-1-1 model. This is also the tier where the physical-versus-VSA distinction noted in the architecture overview carries the most weight. As the last-resort recovery option, Tier 4 must be able to survive a hypervisor compromise. That is the argument for deploying StoreOnce on a dedicated physical appliance rather than as a VSA in production.

Closing Assessment

The two questions a security practitioner asks about a storage platform are whether the infrastructure remains under the organization’s control and whether the data remains protected. After working through the architecture in the Fort Collins lab, the B10000 answers both more completely than a storage array alone usually can, because the answer does not stop at the array.

The platform logs every administrative action, exposes every policy-relevant setting through its API, and streams the resulting audit and security telemetry to whatever SIEM the organization already runs. The governance plan remains the customer’s to write, but the instrumentation needed to implement it, evidence it, and prove it to an auditor is present and exposed rather than hidden behind a vendor portal. Immutable Virtual Lock snapshots mean that even a compromised administrator account cannot quietly destroy the recovery points, and any attempt to do so becomes a logged, queryable event.

In terms of protection, the detect-and-recover loop at the center of the architecture is the part that held up most convincingly. The B10000 flagged the simulated encryption workload within four to five minutes at the block level and captured a forensic snapshot at the moment of detection without waiting on an external system, and the affected volume was recovered cleanly from both a scheduled Virtual Lock snapshot promote and a StoreOnce Catalyst restore. That is the sequence that matters most during an incident, and it ran without a third-party product in the dependency chain.

What sets the architecture apart is not that the B10000 is dramatically more secure than competing arrays when considered in isolation. It is that HPE is one of the few vendors positioned to deliver storage, virtualization, replication, backup, and observability as a coordinated system, and to test that system as a whole against live ransomware in its own lab rather than validating each component in isolation. For an enterprise that owns the infrastructure but answers to a separate security team and an external regulator, that coordination is the difference between a collection of capable products and a resilience strategy that can be operated under pressure.

The architecture is also advancing quickly. HPE is on a steady release cadence across the stack, with tighter integration between VM Essentials and the B10000, ransomware detection expanding to additional data types beyond block volumes, and a published full-stack reference architecture, all on the near-term roadmap. The direction points toward more of the coordination handled natively by the platform and a clearer blueprint for assembling the kind of resilience design this paper examines.

For years, the storage array was treated as one of the safest assets in the building, largely because no one was paying attention to it. The case this architecture makes is that the array should be the opposite: not the overlooked box in the corner, but an active participant in detecting and surviving an attack, provided the organization does the work of turning the capabilities into a plan.

HPE B10000 Product Page

This report is sponsored by HPE. All views and opinions expressed in this report are based on our unbiased view of the product(s) under consideration.

The post HPE Alletra Storage MP B10000 and NIST CSF 2.0: A Full-Stack Cyber Resilience Architecture appeared first on StorageReview.com.

From Database and Virtualized Workloads to Backup: Dell PowerEdge R4715 and R5715 for SMB Realities

5 June 2026 at 11:30

Although Dell’s PowerEdge R4715 and R5715 are two separate products, they should be viewed as a configurable matrix. The matrix includes two chassis, four AMD EPYC 9005 Series CPU options, a wide range of storage configurations, and the full Dell management and support ecosystem. The combination is purpose-built for SMB organizations that need to match infrastructure investment to actual workload requirements, and for the channel partners who help them get there.

Both platforms launched in March 2026. We covered each one individually in our R4715 and R5715 reviews. This piece is different. Rather than evaluating either server in isolation, we examine how the two platforms and the four CPU options perform across the workloads SMB buyers actually run and where configuration choices make the most difference.

SMB AMD EPYC Workload Guide on Dell PowerEdge

Hypervisor flexibility is a key reason these platforms make sense right now. The virtualization market is evolving, and organizations of all sizes are reevaluating the assumptions underlying their infrastructure. Some are sticking with their established stack; others are migrating, and many are running two or more hypervisors in parallel for the foreseeable future. The R4715 and R5715 support the full range of common options, including VMware ESXi, Microsoft Hyper-V, Proxmox VE, and the major Linux KVM distributions, with consistent management and provisioning regardless of choice. For SMB customers who do not have the luxury of standardizing on a single platform, that flexibility is part of the value and one of the reasons we tested across multiple hypervisors for this piece.

The Dell Ecosystem Advantage for SMB and the Channel

The conversation around server platforms often centers on silicon, which makes sense at the spec sheet level. But for SMB buyers and the VARs and system integrators who support them, the operational experience with the silicon is often the deciding factor. Dell’s PowerEdge ecosystem is mature and well understood, and it manifests in ways that disproportionately benefit organizations with lean IT teams.

iDRAC10 and OpenManage Enterprise are the most visible components. The same management plane spans the entire 17th Generation PowerEdge family, so an SMB that buys an R4715 today can later add an R7725 or any other PowerEdge model without needing to learn a new toolset. For VARs and SIs supporting scores of customers, that consistency is even more valuable. A technician who knows iDRAC understands every PowerEdge customer’s infrastructure. The platform supports remote console access, firmware management, hardware health monitoring, and full Redfish API access for automation. For customers without dedicated infrastructure staff, that capability often means the difference between a phone call to a partner and a site visit.

Beneath the management layer, Dell brings a security and supply chain story that is hard to match. Silicon root of trust, cryptographically signed firmware, secured component verification, and TPM 2.0 with FIPS certification are all standard. ProSupport and ProDeploy services are available with global coverage, which matters for distributed SMBs and for partners selling into multiple geographies. Dell’s supply chain is one of the few in the industry that can deliver predictable lead times at scale. For VARs trying to close deals against backorder uncertainty, that is a competitive asset in itself.

For an SMB IT team of two or three people, or for a channel partner supporting dozens of customers with limited bench depth, the Dell ecosystem significantly reduces the operational surface area. The R4715 and R5715 inherit all of that.

R4715 and R5715 at a Glance

Here’s a brief recap for readers who have not seen our individual reviews. The R4715 is a 1U single-socket server optimized for compute density. It supports up to 24 DDR5 RDIMMs, three PCIe Gen5 slots, and a range of storage options, including 2.5-inch and 3.5-inch SAS/SATA configurations and an 8-bay 2.5-inch U.2 NVMe configuration. This is the right choice when rack density and per-rack-unit compute matter more than drive count.

Dell PowerEdge R4715 top down view

Dell PowerEdge R4715 Chassis View

The R5715 is a 2U single-socket server optimized for storage capacity and I/O expandability. It supports the same 24 DDR5 RDIMMs and the four CPU options. However, the R5715 adds a fourth PCIe Gen5 slot and steps up to either 12 bays of 3.5-inch SAS/SATA drives or 16 bays of 2.5-inch SAS/SATA drives. The 3.5-inch configuration can reach 288TB of raw capacity in a single node, which is the configuration we built our R5715 around for this article.

Dell PowerEdge R5715 top down view

Dell PowerEdge R5715 Chassis View

Both platforms are air-cooled, ship with iDRAC10, and support 800W and 1100W power supplies in Platinum or Titanium efficiency grades. These PowerEdge servers do not support GPUs, DPUs, or Fibre Channel, which aligns with Dell’s positioning of these servers as right-sized platforms with a specific target rather than maximum-flexibility platforms.

It is worth understanding where these two servers sit within Dell’s broader AMD-based PowerEdge lineup. The R4715 and R5715 are the value-optimized entry points, purpose-built for the SMB workloads covered in this article. Customers who need accelerators, higher core counts, or greater expansion have a clear path up the stack to the R6715 and R7715, which add GPU and DPU support, processors scaling well beyond 32 cores, and additional PCIe capacity for accelerator-driven and performance-intensive workloads. This tiering is an advantage for the channel; a VAR can place an R4715 or R5715 with a customer today and scale that customer up to the R6715 or R7715 as requirements grow, all within the same management plane, deployment workflow, and support model.

Platform Specifications

Specification Dell PowerEdge R4715 Dell PowerEdge R5715
Processor
Processor One 5th Generation AMD EPYC 9005 Series processor, up to 32 cores
Form Factor 1U rack server 2U rack server
Memory
DIMM Slots 24 DDR5 DIMM slots
Maximum Memory 1.5 TB (up to 64 GB per DIMM)
Memory Speed Up to 5200 MT/s
Memory Type Registered ECC DDR5 RDIMMs only
Storage
Internal Controllers (RAID) PERC H365i, H965i
Internal Boot BOSS-N1 DC-MHS
External HBAs N/A
Front Drive Bays 4x 3.5-inch SAS
8x 2.5-inch SAS/SATA
8x U.2 NVMe Gen4
12x 3.5-inch SAS/SATA
16x 2.5-inch SAS/SATA
Power
Power Supplies Platinum 800W, 1100W
Titanium 800W, 1100W
FTR supported
Cooling & Fans
Cooling Options Air cooling
Fans Up to four sets (dual fan module) hot-plug fans Up to six hot-plug fans
Dimensions
Height 42.8 mm (1.68 inches) 86.8 mm (3.41 inches)
Width 482.0 mm (18.97 inches)
Depth (with bezel) 816.921 mm (32.16 inches) 802.4 mm (31.59 inches)
Depth (without bezel) 815.141 mm (32.09 inches) 801.51 mm (31.55 inches)
Bezel Optional metal bezel

Four AMD EPYC 9005 Series CPU Options

Dell offers four specific CPU SKUs across both servers. The selection is deliberate, covering the range of SMB workloads without overlap or unnecessary complexity. Each CPU is built on AMD’s Zen 5 microarchitecture and shares the same platform-level memory and PCIe characteristics.

CPU Cores Default TDP cTDP Range Base Clock Max Boost L3 Cache
EPYC 9335 32 210W 200-240W 3.0 GHz 4.4 GHz 128 MB
EPYC 9255 24 200W 200-240W 3.2 GHz 4.3 GHz 128 MB
EPYC 9135 16 200W 200-240W 3.65 GHz 4.3 GHz 64 MB
EPYC 9015 8 125W 120-155W 3.6 GHz 4.1 GHz 64 MB

 

The 32-core 9335 is the top of the stack and the most flexible for compute-bound workloads. The 24-core 9255 is the closest match to the price-performance sweet spot we observed in our testing, particularly for database workloads, where marginal gains in core count slow after 24 cores. The 16-core 9135 and 8-core 9015 are where the value story is strongest. Because much of the software SMBs run is licensed per core, including Windows Server, many relational databases, and some hypervisor and backup platforms, the core count chosen at purchase carries forward as a recurring cost for the life of the deployment.

Selecting an 8-core or 16-core SKU that comfortably covers the workload, rather than over-provisioning cores that sit idle, reduces both acquisition cost and ongoing licensing exposure. The 16-core 9135 is especially relevant here, as it aligns with the minimum core count for Windows Server licensing, making it a natural starting point for Windows-centric environments that want to right-size the CPU to the license floor without leaving performance on the table. The 8-core 9015 is the lowest-power and lowest-cost option, ideal for storage-forward workloads or roles where the CPU is not the constraint. Each CPU runs at the same memory speeds and supports the same PCIe Gen5 lane count, so configuration decisions higher up the stack are not constrained by SKU choice.

Performance Testing

Testing Configurations Dell PowerEdge R4715 Dell PowerEdge R5715
Tested CPU’s AMD EPYC 9335, 9255, 9135, 9015 AMD EPYC 9015
Memory 384GB DDR5 384GB DDR5
Boot Storage BOSS RAID1 BOSS RAID1
Front Storage Configuration 8x Samsung PM9D3a RI U.2 Gen5 NVMe SSDs (1.92TB) Raid 10 x 6 12x 20TB HDDs in RAID6

Database Performance: HammerDB MariaDB TPC-C

The headline workload for this evaluation is HammerDB running TPC-C against MariaDB 12.3.1. TPC-C is a long-established OLTP benchmark that produces measurable, comparable results across CPU and storage configurations and represents the kind of transactional database workload at the core of most SMB application stacks. We tested two distinct profiles: a CPU-intensive profile that stresses transaction processing and an I/O-intensive profile that places greater load on the storage subsystem. Both profiles were run across all four CPU options on the R4715 flash configuration to produce a clean CPU scaling curve, and then on the R5715 HDD configuration to show what changes when the storage substrate shifts.

Dell PowerEdge R4715

The HammerDB results show a clear scaling trend as core counts increase across the R4715 platform. Starting with the 8-core EPYC 9015, the system reached 480,818 NOPM in the CPU-intensive profile and 296,105 NOPM in the I/O-intensive profile before leveling off as the number of virtual users increased. Moving to the 16-core EPYC 9135 brought a substantial jump in throughput, pushing CPU-intensive performance to 737,445 NOPM and I/O-intensive performance to 493,093 NOPM, while also allowing the system to sustain higher virtual user counts before saturation.

The jump to the 24-core EPYC 9255 pushed the platform past the one million NOPM mark in the CPU-intensive profile, peaking at 1,017,429 NOPM, while the I/O-intensive profile climbed to 740,574 NOPM. At this point, the additional cores continued to translate directly into usable transactional throughput, while the NVMe storage subsystem kept pace with the growing database load.

At the top end, the 32-core EPYC 9335 delivered the highest results across both profiles, reaching 1,133,714 NOPM in the CPU-intensive workload and 910,321 NOPM in the I/O-intensive workload. The scaling curve remained relatively smooth even at high virtual user counts, indicating that the R4715 flash configuration effectively utilized the larger CPU configurations without storage bottlenecks prematurely limiting performance.

Dell PowerEdge R5715

Then we tested the Dell PowerEdge R5715, configured with 12×20TB HDDs in RAID 6, alongside the 8-core AMD EPYC 9015. In the CPU-intensive profile, the platform reached a peak of 484,715 NOPM at 16 virtual users, with throughput scaling cleanly as additional users were introduced before leveling off near saturation.

The I/O-intensive profile peaked at 308,012 NOPM with 24 virtual users, indicating solid transactional performance of the high-capacity HDD array under moderate concurrency. As the workload ramped, the scaling curve flattened as the spinning-disk subsystem approached its practical performance ceiling under sustained concurrent database activity.

Windows Server Shared Storage

The second practical use case targets a different SMB workload pattern: Windows-based shared storage. For organizations running file shares, departmental applications, or general-purpose Windows Server roles, these platforms can meet modern SMB performance expectations. We ran FIO on Windows Server to characterize sequential and random performance across two storage configurations: a RAID 6 HDD array on the R5715 as a baseline, and an 8-drive SSD JBOD array on the R4715 representing a high-performance storage configuration. The comparison illustrates the operational gap between the two storage tiers in this platform class.

The gap between the two storage configurations becomes immediately apparent in the FIO results. While the RAID6 HDD array in the R5715 delivered respectable sequential throughput for a high-capacity spinning-disk platform, the SSD-equipped R4715 operated in a completely different performance class, particularly in random workloads, where SMB environments tend to feel storage latency most.

Sequential performance on the HDD array reached up to 3.7GB/s writes and 2.2GB/s reads in the 4-thread tests, which is more than adequate for traditional file serving, backups, and bulk storage tasks. However, the SSD configuration pushed sequential throughput into the tens of gigabytes per second, exceeding 56GB/s reads and 26GB/s writes while maintaining dramatically lower latency.

The separation widened even further in 4K random workloads. The HDD array peaked at under 1,300 IOPS in random-write testing, with latency climbing above 100ms, whereas the SSD configuration delivered over 4 million IOPS with sub-millisecond latency. In practical terms, this translates directly into application responsiveness, multi-user file share performance, VM storage behavior, and the ability to sustain concurrent SMB workloads without the storage layer becoming a bottleneck.

FIO Workload R5715 RAID6 HDD 1T R5715 RAID6 HDD 4T R4715 8x SSD 1T R4715 8x SSD 4T

Sequential Read (128K)

Bandwidth 1,475.89 MB/s 2,198.89 MB/s 56,861.09 MB/s 56,866.76 MB/s
IOPS 11,807 17,589 454,885 454,918
Latency 2.70ms 7.28ms 0.56ms 2.25ms

Sequential Write (128K)

Bandwidth 2,665.31 MB/s 3,726.63 MB/s 26,739.52 MB/s 26,753.48 MB/s
IOPS 21,322 29,811 213,912 214,011
Latency 1.49ms 4.39ms 1.20ms 4.78ms

Random Read (4K)

Bandwidth 1.00 MB/s 3.60 MB/s 7,268.78 MB/s 16,143.19 MB/s
IOPS 256 919 1,860,803 4,132,645
Latency 125.11ms 139.00ms 0.13ms 0.17ms

Random Write (4K)

Bandwidth 4.88 MB/s 4.67 MB/s 7,555.53 MB/s 16,010.24 MB/s
IOPS 1,248 1,195 1,934,214 4,098,613
Latency 25.63ms 106.98ms 0.07ms 0.13ms

Proxmox Backup Server

Beyond raw benchmarks, the R5715 with HDD storage is exactly the kind of platform that suits a virtualized backup target workload. To validate that fit, we deployed Proxmox Backup Server on the R5715 configured with the 8-core EPYC 9015 and the same 12-bay 3.5-inch HDD array. Proxmox is a representative example of the broader open-source hypervisor and infrastructure ecosystem that has gained significant traction in SMB environments, and Proxmox Backup Server, in particular, is well-suited to this server’s storage profile.

We deployed Proxmox Backup Server 4.2.0 and used it to back up the virtual machines that power our Proxmox community Discord server environment.

In our specific configuration, backup and restore operations were somewhat limited by the system’s 1GbE networking connection, which became the primary bottleneck during larger transfers. However, the platform supports straightforward networking upgrades via OCP expansion cards, making it easy to migrate to 10GbE or even 25GbE connectivity. With faster networking in place, the R5715 would be able to handle significantly higher backup throughput and restore performance, especially in environments with larger VM datasets or more demanding backup windows.

Conclusion

The Dell PowerEdge R4715 and R5715 succeed by getting the configuration matrix right. Two chassis with clearly differentiated form factors, four CPU options that cover the practical range of SMB workloads without overlap, and a storage menu broad enough to support everything from low-cost bulk capacity to all-flash performance. The configuration flexibility is not theoretical. Across the workloads we tested, the right answer changed with each one. The 24-core 9255 hit the sweet spot for transactional database performance on flash. The 8-core 9015 provided exactly enough CPU for a Proxmox Backup Server deployment with bulk HDDs. The R4715 with flash storage was the appropriate choice for Windows shared storage, while the R5715 with HDDs was the appropriate choice for capacity-focused workloads where peak I/O is not the constraint.

For SMB customers and the partners who serve them, the value of these platforms lies in the ability to match the build to the workload, rather than buying more capacity than the workload requires. The Dell ecosystem, including iDRAC10 management, ProSupport, the security stack, and supply chain predictability, amplifies that value at the operational level. Lean IT teams and channel partners both benefit when the platform behind the customer’s workload is a known quantity.

Dell’s positioning of these servers as a way to consolidate legacy infrastructure and reduce per-socket and per-core licensing exposure holds up against the data we collected. With four CPU options spanning 8 to 32 cores at the same platform level, customers and partners have the flexibility to right-size acquisition, licensing, and operational costs to match actual requirements. That is the value proposition, and the R4715 and R5715 deliver against it.

Dell PowerEdge R5715 Product Page

Dell PowerEdge R4715 Product Page

The post From Database and Virtualized Workloads to Backup: Dell PowerEdge R4715 and R5715 for SMB Realities appeared first on StorageReview.com.

Micron 6600 ION 245TB SSD Review: A Quarter Petabyte Per Drive Bay

2 June 2026 at 19:16

Micron’s 6600 ION NVMe SSD has reached 245.76TB, pushing the company’s capacity-focused PCIe Gen5 QLC line into quarter-petabyte territory. Micron began shipping the 245TB model on May 5, 2026, and is positioning it as the highest-capacity commercially available SSD. The drive sits at the top of the 6600 ION family, uses Micron’s ninth-generation G9 QLC NAND, and is available at this capacity in E3.L 9.5mm and U.2 15mm form factors. The 30.72TB, 61.44TB, and 122.88TB models also support E3.S.

Micron’s G9 QLC uses a six-plane architecture and pushes NAND I/O to 3.6 GB/s, which the company says makes it the fastest QLC currently shipping in a data center SSD. That gives Micron a useful platform for a drive that is clearly aimed at hyperscale object storage, AI data lakes, analytics, content repositories, and other environments where capacity per rack and watts per terabyte matter more than small-block write intensity.

Under Micron’s rack assumptions, a 720-drive E3.L configuration reaches 176.9PB of raw capacity with 245.76TB 6600 ION drives, compared with 31.7PB for the same number of bays populated with 44TB HDDs. Power density moves in the same direction. At Micron’s 30W peak rating, the 245TB 6600 ION delivers roughly 8.2TB per watt. The HDD Micron uses for comparison is 44TB at 10W, or about 4.4TB per watt. That 10W HDD number is an estimate because Seagate has not publicly disclosed full power specifications for its 44TB Mozaic 4+ drives, but it is consistent with the broader power envelope of current high-capacity enterprise HDDs.

That comparison is relevant because the macro picture of data centers is shifting. The IEA estimates global data center electricity consumption at roughly 415 TWh in 2024 and projects it will more than double to about 945 TWh by 2030, with AI identified as the largest driver of the increase. In the U.S., data centers are projected to account for nearly half of electricity demand growth through 2030. For large operators, grid access, floor space, cooling capacity, and watts per terabyte are no longer secondary planning details; they are constraints that shape what can actually be deployed.

There are trade-offs, and they are important to consider before getting into the test data. A 245TB QLC SSD is built for density and read-heavy access, not for replacing every high-write TLC workload. The 245.76TB 6600 ION is rated for up to 13.7 GB/s sequential read and 1.78M random read IOPS, which puts it near the top of the QLC field, on paper. Writes are more restrained, with sequential write rated at 3.0 GB/s and random write at 42,000 IOPS for both 4K and 16K transfers. The top-capacity model also uses a 16K indirection unit rather than the 4K IU used in the 30.72TB version, which matters for workloads that issue sub-16K random writes. That shows up in endurance as well, with 4K RDWPD dropping to 0.075 while 16K RDWPD holds at 0.3.

On the enterprise feature side, the 6600 ION checks the expected boxes. The drive supports OCP 2.6, NVMe 2.0d, NVMe-MI 1.2d, SPDM 1.2, CNSA 2.0 firmware verification with dual-signed updates, and SED options. It is TAA-compliant and FIPS 140-3 L2 certifiable. Micron also rates the drive at 2.5 million hours MTTF at 50°C per OCP 2.5 REL-1, which is important given the role these drives are expected to play in dense, always-on storage infrastructure.

Micron 6600 ION 245TB Specifications

Specification Micron 6600 ION 245.76TB
Platform Overview
Capacity 245.76TB
Form Factors U.2 (15mm)
E3.L
Interface PCIe Gen5 x4 NVMe (v2.0b)
NAND Micron G9 QLC NAND
Performance
Sequential Read 13,700 MB/s
Sequential Write 3,000 MB/s
Random Read 1,780,000 IOPS
Random Write (4K) 42,000 IOPS
Random Write (16K) 42,000 IOPS
Read Latency 100µs (QD1, Typical)
Write Latency 20µs (QD1, Typical)
Power and Endurance
Maximum Power Consumption ≤30W
Idle Power ≤5W
Endurance 1.0 SDWPD (128KB sequential write)
0.3 RDWPD (16KB random write)
0.075 RDWPD (4KB random write)
MTTF 2.5 million device hours
UBER <1 sector per 1017 bits read
Features and Security
Compliance OCP 2.6
NVMe 2.0d
NVMe-MI 1.2d
TAA-compliant
FIPS 140-3 L2 certifiable
Security Features CNSA 2.0
SPDM 1.2
Micron SEE
SED options
Additional Features SGLs
SRIS
PCIe lane reversals

Micron 6600 ION 245TB Performance

Drive Testing Platform

We use a Dell PowerEdge R760 running Ubuntu 22.04.2 LTS as our test platform for all workloads in this review. Equipped with a Serial Cables Gen5 JBOF, it offers wide compatibility with U.2, E1.S, E3.S, and M.2 SSDs. Our system configuration is outlined below:

  • 2 x Intel Xeon Gold 6430 (32-Core, 2.1GHz)
  • 16 x 64GB DDR5-4400
  • 480GB Dell BOSS SSD
  • Serial Cables Gen5 JBOF
  • NVIDIA L4

Drives Compared

FIO Performance Benchmark

To measure the storage performance of each SSD across common industry metrics, we leverage FIO. Prior to this review, each drive undergoes the same testing process, which includes a preconditioning step of two full drive fills with a sequential write workload, followed by steady-state performance measurement. With the Micron 6600 ION pushing to nearly 1/4 petabyte, we leveraged the sprandom preconditioning process for the random workloads on this SSD. This was to speed up the preconditioning process, which would otherwise have taken days or weeks per drive fill. As each workload type being measured changes, we run another preconditioning fill of that new transfer size.

In this section, we focus on the following FIO benchmarks:

  • 128K Sequential
  • 64K Random
  • 16K Random
  • 4K Random

128K Sequential Write (IODepth 16 / NumJobs 1)

For 128K Sequential Write, the Micron 6600 ION delivered 2,838.0MB/s. While that placed it behind the Micron 6550 ION at 10,456.4MB/s and the DapuStor R6060 at 3,920.6MB/s, it remained competitive with the rest of the comparison group. The Solidigm P5336 122.88TB reached 3,152.5MB/s, while the DapuStor J5060 and Solidigm P5336 61.44TB posted 2,883.1MB/s and 2,503.5MB/s, respectively. The result highlights the 6600 ION’s read-centric design, where write performance remains adequate but is not the primary focus.

128K Sequential Write Latency (IODepth 16 / NumJobs 1)

For 128K Sequential Write latency, the Micron 6600 ION measured 704.0µs. The Micron 6550 ION led the workload with just 191.0µs, while the DapuStor R6060 followed at 510.0µs. The Solidigm P5336 122.88TB delivered 634.0µs, narrowly outperforming the 6600 ION, while the Solidigm P5336 61.44TB recorded the highest latency at 798.0µs. Despite landing in the middle of the pack, the 6600 ION remained comfortably below the 1ms threshold.

128K Sequential Read (IODepth 64 / NumJobs 1)

For 128K Sequential Read, the Micron 6600 ION 245TB reached 12,729.8MB/s, finishing second overall behind the Micron 6550 ION 61.44TB, which led the group at 13,979.7MB/s. The DapuStor R6060 122.88TB was close behind at 11,554.0MB/s, while the remainder of the comparison group trailed significantly, with both Solidigm P5336 drives and the DapuStor J5060 clustered around the 7.1GB/s mark. The 6600 ION clearly separated itself from most of the field, with only the 6550 maintaining a measurable advantage.

128K Sequential Read Latency (IODepth 64 / NumJobs 1)

For 128K Sequential Read latency, the Micron 6600 ION recorded 628.0µs, giving it the second-best result in the comparison. The Micron 6550 ION posted the lowest latency at 571.9µs, while the DapuStor R6060 followed at 692.1µs. The remaining drives all exceeded 1ms of latency, with the Solidigm P5336 122.88TB posting the highest figure at 1123.0µs. The 6600 ION maintained a substantial latency advantage over the majority of the field, trailing only its smaller Micron sibling.

64K Random Write

For 64K Random Write, the Micron 6600 ION achieved 2,999.6MB/s. The Micron 6550 ION dominated this workload at 10,516.5MB/s, while the DapuStor R6060 followed at 3,916.9MB/s. The Solidigm P5336 122.88TB, DapuStor J5060, and Solidigm P5336 61.44TB delivered 3,182.7MB/s, 2,883.7MB/s, and 2,721.4MB/s, respectively. The 6600 ION landed in the middle of the comparison group, ahead of several competing QLC offerings.

64K Random Write Latency

For 64K Random Write latency, the Micron 6600 ION recorded 83.0µs, one of the strongest results in the group. Only the DapuStor R6060 posted a lower figure at 63.0µs. The Micron 6550 ION followed at 95.0µs, while the DapuStor J5060 recorded the highest latency at 1386.0µs. This workload showcased the 6600 ION’s ability to maintain extremely responsive write behavior despite its massive capacity.

64K Random Read

For 64K Random Read, the Micron 6600 ION reached 11,946.5MB/s, placing third overall. Both the DapuStor R6060 and Micron 6550 ION edged ahead at 13,274.8MB/s and 13,204.6MB/s, respectively. The rest of the field remained well behind, with the Solidigm P5336 models and DapuStor J5060 all hovering near 7.1GB/s. While not the outright leader, the 6600 ION remained firmly within the top tier of performers.

64K Random Read Latency

For 64K Random Read latency, the Micron 6600 ION posted 669.0µs. The Solidigm P5336 122.88TB surprisingly delivered the lowest latency at 563.0µs, followed by the Micron 6550 ION at 607.0µs. The DapuStor R6060 recorded the highest latency among the leading performers at 1285.0µs. The 6600 ION landed near the front of the group and maintained a healthy balance between throughput and response time.

16K Random Read

The Micron 6600 ION reached approximately 808K IOPS in the 16K Random Read workload, placing it among the top performers in the comparison group. The Micron 6550 ION led the field at roughly 858K IOPS, while the DapuStor R6060 followed closely behind at approximately 818K IOPS. The rest of the comparison group trailed significantly, with the Solidigm P5336, Solidigm P5336, and DapuStor J5060 all landing near the 450K IOPS mark. The result demonstrates the 6600 ION’s ability to deliver strong random read performance despite its massive capacity.

16K Random Read Latency

The Micron 6600 ION maintained one of the more consistent latency curves in the 16K Random Read workload, staying below 150µs through most of the test before finishing around 640µs at peak load. The Micron 6550 ION ended with the lowest latency at roughly 300µs, while the DapuStor R6060 finished near 650µs. The DapuStor J5060 showed the most volatility, with several large latency spikes and a peak exceeding 1ms.

16K Random Write

The Micron 6600 ION delivered approximately 192K IOPS in the 16K Random Write workload. The Micron 6550 ION remained the clear leader at roughly 661K IOPS, with the DapuStor R6060 reaching about 224K IOPS. The Solidigm P5336 61.44TB and 122.88TB followed closely at approximately 201K IOPS and 195K IOPS, while the DapuStor J5060 landed just below the 6600 ION. Although the 6600 ION was not designed as a write-focused drive, it remained competitive against the other high-capacity QLC SSDs in the comparison.

16K Random Write Latency

In the 16K Random Write workload, the Micron 6600 ION maintained low latency through much of the test before rising sharply at the highest queue depths, finishing around 11.6ms. The Micron 6550 ION delivered the most stable behavior, remaining under 500µs throughout, while the DapuStor R6060 peaked near 18.7ms and the Solidigm P5336 ended around 9.3ms. The 6600 ION remained competitive until saturation, where latency increased rapidly under maximum load.

4K Random Read

The Micron 6600 ION 245TB reached approximately 1.75 million IOPS in the 4K Random Read workload, delivering one of the strongest results in the comparison group. Only the Micron 6550 ION and DapuStor R6060 finished ahead, while the remaining drives trailed by a substantial margin. The result reinforces the 6600 ION’s ability to handle highly transactional read-heavy workloads despite being optimized for extreme storage density.

4K Random Read Latency

The Micron 6600 ION recorded 289.0µs average latency in the 4K Random Read test. This placed it among the better-performing drives in the comparison, balancing high IOPS output with responsive access times. While several competitors posted lower latency figures, the 6600 ION remained close within the leading tier of drives tested.

GPU Direct Storage

One of the tests we conducted on this testbench was the Magnum IO GPU Direct Storage (GDS) test. GDS is a feature developed by NVIDIA that allows GPUs to bypass the CPU when accessing data stored on NVMe drives or other high-speed storage devices. Instead of routing data through the CPU and system memory, GDS enables direct communication between the GPU and the storage device, significantly reducing latency and improving data throughput.

How GPU Direct Storage Works

Traditionally, when a GPU processes data stored on an NVMe drive, the data must first travel through the CPU and system memory before reaching the GPU. This process introduces bottlenecks, as the CPU acts as an intermediary, adding latency and consuming valuable system resources. GPU Direct Storage eliminates this inefficiency by enabling the GPU to access data directly from the storage device via the PCIe bus. This direct path reduces data-movement overhead, enabling faster, more efficient data transfers.

AI workloads, especially those involving deep learning, are highly data-intensive. Training large neural networks requires processing terabytes of data, and any delay in data transfer can lead to underutilized GPUs and longer training times. GPU Direct Storage addresses this challenge by ensuring that data is delivered to the GPU as quickly as possible, minimizing idle time and maximizing computational efficiency.

In addition, GDS is particularly beneficial for workloads that involve streaming large datasets, such as video processing, natural language processing, or real-time inference. By reducing the reliance on the CPU, GDS accelerates data movement and frees up CPU resources for other tasks, further enhancing overall system performance.

GDSIO Sequential Read Throughput

In our GDSIO sequential read throughput test, the Micron 6600 ION delivered the strongest small-block showing of the group and stayed competitive all the way through the larger transfer sizes. At 16K, it opened at 409.6MiB/s with a single thread, climbed steadily to 532.5MiB/s at four threads, 614.4MiB/s at eight, and 921.6MiB/s at 16, then pushed to 1.4GiB/s at 32 threads, 2.0GiB/s at 64, and settled at 1.9GiB/s by 128 threads. That gave it a clear lead in the 16K section over the DapuStor J5060, DapuStor J6060, Micron 6550 ION, and Solidigm D5-P5336, with the gap widening considerably as thread counts moved into the upper half of the test.

At 128K, the 6600 ION continued to look strong. It posted 921.6MiB/s at one thread, jumped to 1.9GiB/s at four, 3.0GiB/s at eight, and 3.8GiB/s at 16, then climbed to 4.5GiB/s at 32 threads, 5.0GiB/s at 64, and finished at 5.2GiB/s at 128. That put it at the top of the comparison group across nearly every thread count in the 128K runs, with only the DapuStor J6060 closing the gap at the very heaviest concurrency. The drive’s ability to scale smoothly from low to high thread counts at this block size was one of its standout traits.

The 1M results were a bit more uneven at the low end but finished on a high note. The 6600 ION opened at 2.9GiB/s with a single thread, dipped to 2.2GiB/s at four threads, then recovered to 3.2GiB/s at eight, 3.7GiB/s at 16, and 4.2GiB/s at 32. From there, it climbed sharply to 5.0GiB/s at 64 threads and topped out at 5.9GiB/s at 128, matching the DapuStor J6060 for the highest 1M result in the group. So while the 6600 ION had a small soft spot in the mid-thread range at 1M, its overall sequential read profile was excellent, particularly in the 16K and 128K sections where it consistently led the comparables.

GDSIO Sequential Read IOPS

In the GDSIO sequential read IOPS test, the Micron 6600 ION delivered the strongest small-block performance in the group by a wide margin. At 16K, it opened at 24.9K IOPS with a single thread, climbed to 34.5K at four threads, 41.3K at eight, and 61.5K at 16, then surged to 92.4K at 32 threads and peaked at 130.4K at 64 before settling slightly to 126.7K at 128. That peak was well ahead of the rest of the comparables, with the Solidigm D5-P5336 122.88TB the next closest at roughly 93K and the DapuStor J5060, DapuStor J6060, and Micron 6550 ION 61.44TB all topping out in the 55K to 63K range. The 6600 ION’s separation from the field grew sharply from 16 threads onward, making it the clear leader through the entire small-block portion of the test.

At 128K, the 6600 ION continued to lead the group across nearly every thread count. It posted 7.6K IOPS at one thread, jumped to 15.6K at four, 24.5K at eight, and 30.9K at 16, then climbed to 36.6K at 32 threads, 41.2K at 64, and finished at 43.0K at 128. That kept it ahead of the DapuStor J6060, which trailed by a few thousand IOPS at the heaviest thread counts, and well ahead of the DapuStor J5060, Solidigm D5-P5336, and Micron 6550 ION. The drive’s scaling from low to high concurrency at this block size was smooth and consistent, with no dips or plateaus.

In the 1M runs, the IOPS spread between drives narrowed considerably, as expected with larger transfer sizes. The 6600 ION opened at 3.0K IOPS at one thread, dipped to 2.3K at four threads, then recovered to 3.2K at eight, 3.7K at 16, and 4.3K at 32. From there, it climbed to 5.2K at 64 threads and topped out at 6.0K at 128, finishing right alongside the DapuStor J6060 at the top of the group. So while the 1M section was a closer race, the 6600 ION’s 16K and 128K results were dominant, giving it the best overall sequential read IOPS profile in the comparison.

GDSIO Sequential Read Latency

Finally, in our GDSIO sequential read latency test, the Micron 6600 ION posted some of the lowest average latencies in the group, particularly at the smaller block sizes and across the higher thread counts, where the rest of the comparables started to fall off. At 16K, it opened at 39µs with a single thread, climbed to 115µs at four threads, 192µs at eight, and 259µs at 16, then moved to 344µs at 32 threads, 487µs at 64, and 1.0ms at 128. Those results kept it ahead of the DapuStor J5060, DapuStor J6060, Micron 6550 ION, and Solidigm D5-P5336 across most of the 16K runs, with the gap widening at higher thread counts, where the 6600 ION held its latency much tighter than the rest of the field.

At 128K, the 6600 ION continued to look strong. It posted 131µs at one thread, 256µs at four, 325µs at eight, and 516µs at 16, then climbed to 873µs at 32 threads, 1.6ms at 64, and 3.0ms at 128. That kept it near the bottom of the latency stack across the entire 128K section, trailing only the DapuStor J6060 by a narrow margin at the heaviest thread counts and pulling well ahead of the Micron 6550 ION and Solidigm D5-P5336, which began to climb sharply once concurrency moved past 32 threads.

The 1M runs showed the steepest latency growth, as expected with the largest transfer size. The 6600 ION opened at 332µs with a single thread, then climbed to 1.8ms at four threads, 2.5ms at eight, 4.3ms at 16, 7.4ms at 32, 12.4ms at 64, and topped out at 21.3ms at 128. Even with that increase, the 6600 ION still held the second-lowest 1M latencies in the group, trailing only the DapuStor J6060, while the Micron 6550 ION and Solidigm D5-P5336 climbed to 47.5ms and 28.9ms, respectively, at 128 threads.

GDSIO Sequential Write Throughput

In our GDSIO sequential write throughput test, the Micron 6600 ION 245.76TB delivered a more modest showing than on the read side, particularly at larger transfer sizes, where the rest of the comparables pulled ahead. At 16K, it opened at 512.0MiB/s with a single thread, climbed to 1.1GiB/s at four threads, 1.3GiB/s at eight, and 1.4GiB/s at 16, then peaked at 1.5GiB/s at both 32 and 64 threads before easing back to 1.2GiB/s at 128. That placed it in a tight cluster with the DapuStor J5060, DapuStor J6060, Micron 6550 ION, and Solidigm D5-P5336 across the 16K runs, with no drive meaningfully separating itself at this block size.

At 128K, the 6600 ION lagged the rest of the field. It posted 2.2GiB/s at one thread, climbed to 2.9GiB/s at four threads, and held that level through eight, then settled into a slow decline at 2.8GiB/s at 16, 2.7GiB/s at 32, 2.8GiB/s at 64, and 2.5GiB/s at 128. That kept it at the bottom of the group for most of the 128K runs, with the DapuStor J6060 leading at roughly 3.7 to 3.8GiB/s, the Micron 6550 ION in second through the mid-thread counts, and the Solidigm D5-P5336 and DapuStor J5060 both sitting comfortably above the 6600 ION.

The 1M runs followed a similar pattern. The 6600 ION started at 2.9GiB/s with a single thread, held 2.8GiB/s at four, returned to 2.9GiB/s at eight, then dropped to 2.6GiB/s at 16, 2.4GiB/s at 32, and finished at 2.3GiB/s at both 64 and 128 threads. That negative scaling left it at or near the bottom of the comparison group across the entire 1M section, with the DapuStor J6060 leading at the lower thread counts and the Micron 6550 ION peaking at 3.9GiB/s at 32 threads. So while the 6600 ION had a competitive 16K showing, its 128K and 1M sequential write results trailed the rest of the comparables, especially as thread counts climbed.

GDSIO Sequential Write IOPS

In our GDSIO sequential write IOPS test, the Micron 6600 ION delivered a competitive 16K showing but trailed the rest of the field once block sizes increased. At 16K, it opened at 31.8K IOPS with a single thread, climbed to 70.2K at four threads, 85.7K at eight, and 92.9K at 16, then peaked at 98.0K at 32 threads before easing back to 95.5K at 64 and 81.9K at 128. That kept it in a tight pack with the DapuStor J6060, Micron 6550 ION, and Solidigm D5-P5336 across the 16K runs, with the Micron 6550 ION nudging slightly ahead at the 32-thread peak and the DapuStor J5060 sitting a bit lower through the mid-thread counts.

At 128K, the 6600 ION fell to the back of the comparison group. It posted 18.4K IOPS at one thread, climbed to 24.0K at four threads, and held that through eight, then settled at 23.1K at 16, 22.2K at 32, 22.8K at 64, and finished at 20.6K at 128. That left it trailing the DapuStor J6060 and Micron 6550 ION by a wide margin at the lower and middle thread counts, and only narrowly ahead of the DapuStor J5060 across most of the 128K section. The drive’s negative scaling pattern beyond four threads stood out, since the leaders maintained a flatter profile.

The 1M runs followed the same trend. The 6600 ION opened at 3.0K IOPS at one thread, dipped to 2.9K at both four and eight threads, then continued declining to 2.7K at 16, 2.5K at 32, and 2.4K at both 64 and 128 threads. That placed it at or near the bottom of the comparison group throughout the 1M section, well behind the DapuStor J6060, which led at 3.9K IOPS at the low-thread counts, and the Micron 6550 ION, which peaked at 4.0K at 32 threads. So while the 6600 ION held its own at 16K, its 128K and 1M sequential write IOPS results were the weakest of the comparables, particularly as concurrency increased.

GDSIO Sequential Write Latency

In our GDSIO sequential write latency test, the Micron 6600 ION delivered competitive numbers at the smaller block sizes but climbed to the highest latencies in the group as transfer size and thread count scaled up. At 16K, it opened at 31µs with a single thread, moved to 56µs at four threads, 92µs at eight, and 170µs at 16, then continued to 324µs at 32 threads, 666µs at 64, and 1.6ms at 128. Those results placed it in a tight band with the DapuStor J5060, DapuStor J6060, Micron 6550 ION, and Solidigm D5-P5336 across the 16K runs, with no drive meaningfully separating itself at this block size.

At 128K, the 6600 ION started to fall behind the leaders. It posted 53µs at one thread, climbed to 165µs at four, 331µs at eight, and 690µs at 16, then moved to 1.4ms at 32 threads, 2.8ms at 64, and 6.2ms at 128. That put it near the upper end of the latency stack across the 128K runs, with the DapuStor J6060 delivering the lowest latency and the Micron 6550 ION close behind, while the 6600 ION and Solidigm D5-P5336 trailed at the highest thread counts.

The 1M runs were where the 6600 ION fell furthest behind. It opened at 330µs with a single thread, then climbed sharply to 1.4ms at four threads, 2.7ms at eight, 6.0ms at 16, 12.8ms at 32, 26.8ms at 64, and topped out at 53.7ms at 128. That 128-thread result was the highest 1M write latency in the comparison group, ahead of the DapuStor J5060 at roughly 45ms, the DapuStor J6060 at 43ms, the Solidigm D5-P5336 at 41ms, and the Micron 6550 ION at 39ms. So while the 6600 ION held its own at 16K, its sequential write latencies climbed faster than the comparables as block size and concurrency increased, leaving it at the bottom of the group on the heaviest 1M workloads.

DLIO Checkpointing Benchmark

To evaluate SSD real-world performance in AI training environments, we utilized the Data and Learning Input/Output (DLIO) benchmark tool. Developed by Argonne National Laboratory, DLIO is specifically designed to test I/O patterns in deep learning workloads. It provides insights into how storage systems handle challenges such as checkpointing, data ingestion, and model training. The test is designed so that each drive is fully populated with full checkpoints; larger SSDs fit more checkpoints. The chart below illustrates how both drives handle the process across 99 checkpoints (198 for the 122TB). When training machine learning models, checkpoints are essential for periodically saving the model’s state, preventing loss of progress during interruptions or power failures. This storage demand requires robust performance, especially under sustained or intensive workloads. We used DLIO benchmark version 2.0 from the August 13, 2024, release.

To ensure our benchmarking reflected real-world scenarios, we based our testing on the LLAMA 3.1 405B model architecture. We implemented checkpointing using torch.save() to capture model parameters, optimizer states, and layer states. Our setup simulated an eight-GPU system, implementing a hybrid parallelism strategy with 4-way tensor parallelism and 2-way pipeline parallelism distributed across the eight GPUs. This configuration resulted in checkpoint sizes of 1,636GB, representative of modern large language model training requirements.

In our DLIO checkpoint benchmark, which measures the (Pass Average) time required to complete checkpoint operations across three sequential passes, the Micron 6600 ION 245.76TB started in the middle of the group and slowed considerably as additional passes accumulated. On pass one, it finished in 580.98 seconds, placing it third in the comparison behind the DapuStor R6060 122TB at 465.33 seconds and the Micron P5336 ION 61.44TB at 484.30 seconds, and ahead of the Solidigm P5336 122.88TB at 530.75 seconds and the Solidigm P5336 61.44TB at 662.68 seconds.

On pass two, the 6600 ION climbed sharply to 926.70 seconds, a jump of roughly 60 percent from its first pass. That moved it to the back of the group alongside the DapuStor R6060, which posted 934.50 seconds. The Solidigm P5336 122.88TB climbed more moderately to 746.14 seconds, while the Solidigm P5336 61.44TB and Micron P5336 ION 61.44TB stayed comparatively flat at 640.71 and 570.23 seconds, respectively. That contrast was the key takeaway from the second pass: the high-capacity QLC drives saw a much steeper rise in checkpoint time once the workload was no longer running against a fresh state.

Pass three showed the same pattern continuing. The 6600 ION finished at 968.18 seconds, a small step up from pass two and the slowest result in the group, just behind the DapuStor R6060 at 965.27 seconds. The Solidigm P5336 122.88TB landed at 757.31 seconds, the Solidigm P5336 61.44TB at 639.63 seconds, and the Micron P5336 ION 61.44TB at 585.03 seconds.

Looking at the full per-checkpoint trace rather than the three-pass averages, the Micron 6600 ION 245.76TB held a steady band in the mid-560-second range for the first 24 checkpoints, then briefly spiked to around 990 seconds for checkpoints 26 through 30 before dropping back to roughly 560 seconds and holding that level out to checkpoint 130. From checkpoint 131 onward, it stepped up to the 750 to 850-second range, then climbed again around checkpoint 200 and settled into a sustained band between 880 and 1,030 seconds for the remainder of the run, finishing 390 checkpoints in total, more than any other drive in the comparison. The DapuStor R6060 showed a similar two-stage step pattern, while the Solidigm and Micron 6550 drives held tighter, lower bands, but did not run as many checkpoints overall.

Conclusion

The Micron 6600 ION 245.76TB is a capacity play first, and it executes that brief without apology. A single drive that puts nearly a quarter petabyte into a single drive slot changes the rack math in a way that is hard to argue with, and our testing supports the read-centric positioning Micron built the drive around. Sequential and random read throughput landed at or near the top of the comparison group; GDS read performance scaled cleanly across high thread counts at 16K and 128K; and 4K random read approached 1.75M IOPS. For object stores, AI data lakes, analytics, and content repositories where the access pattern is read-dominated and the constraint is capacity per watt, this drive does what it is supposed to do.

The trade-offs are equally clear, and buyers should size for them rather than around them. Sequential write throughput sat in the middle of the field, GDS write performance and latency fell behind the comparables as block size and concurrency climbed, and the DLIO checkpoint runs showed the 6600 ION slowing meaningfully once the workload moved off a fresh state, finishing at the back of the group on passes two and three. The 16K indirection unit and the 0.075 4K RDWPD rating are the relevant fine print for anyone considering sub-16K random write traffic. None of this is a surprise for a QLC drive built for density, but it does define where the 6600 ION belongs and where it does not.

The broader case rests on power and footprint, not peak performance. At 8.2TB per watt against roughly 4.4TB per watt for the high-capacity HDDs Micron compares to, and with a 720-drive rack reaching 176.9PB versus 31.7PB for an equivalent HDD count, the efficiency argument is the product’s point. For operators managing storage growth inside fixed power and floor-space budgets, the 6600 ION 245TB is a credible answer, provided the workload is matched to its read-heavy design.

Micron 6600 ION Product Page

The post Micron 6600 ION 245TB SSD Review: A Quarter Petabyte Per Drive Bay appeared first on StorageReview.com.

Dell PowerStore Gen 3: Inside the Most Aggressive Enterprise Storage Reset in Years

19 May 2026 at 17:01

Storage refreshes usually come in two flavors. There’s the quiet uplift, where a vendor rolls in a new CPU, claims a few percentage points of performance, and ships the same chassis with a different sticker. And then there’s the generational reset, where the chassis, drives, interconnect, cache architecture, and management plane all move at once. Dell PowerStore Gen 3, branded at launch as PowerStore Elite, is the second kind, and it’s not even close. Every major subsystem in the platform has changed, and most of them have done so in ways that materially redefine what a unified array is supposed to look like in 2026.

Dell PowerStore Gen 3 hero

Dell PowerStore Gen3 Fully Populated w/ Bezel

We’ve been watching PowerStore since the beginning. In our view, this is the most consequential release since the original launch in 2020, and arguably the most aggressive storage platform reset any major vendor has shipped in years. Dell didn’t simply iterate on Gen 2. They rebuilt the platform from the chassis up, made bets on form factors and architectural choices that most of the industry hasn’t yet committed to, and engineered the result for a ten-year service life with multiple in-place controller upgrades. To explore these new models, Dell invited us out to Hopkinton, MA.

Storage density is a key selling point of the new Dell PowerStore, and Dell doesn’t disappoint on that front. The platform supports up to 40 E3.S NVMe drives in a 3U chassis, with planned support for E3.L drives, while keeping every bay user-addressable for data instead of reserving slots for cache SSDs. Dell has also significantly modernized the underlying hardware platform, moving to next-gen Intel processors, alongside DDR5 memory, end-to-end PCIe Gen5 connectivity, and OCP 3.0 modules that replace the older Dell-specific SLIC carrier design.

Connectivity between controllers scales up to 200GbE RDMA today, with a path to even faster speeds via a future I/O card upgrade. On the software side, PowerStoreOS 5.0 introduces new autonomous data path intelligence and log-structured metadata to optimize performance and endurance for high-capacity QLC flash, I/O-level telemetry to lay the groundwork for future inline ransomware detection, and dynamic resource sharing between block and file services. All PowerStore appliances are now unified out of the box, providing enhanced support for file- and block-scale-out and non-disruptive data mobility across clusters.

Dell has also added unaligned deduplication and enhanced compression offloads, a key factor behind the company’s decision to increase its data reduction guarantee from 5:1 in previous generations to 6:1 with the new platform.

Dell PowerStore Gen 3 - 9500 top view with airflow baffle

Dell PowerStore Gen3 Internal Controller View with Fan Shroud

The hardware changes and software updates don’t tell a complete story, though. The architectural decisions Dell made beneath them are material because they determine whether the platform will hold up over the next decade or look dated in a few years. The shift to E3.S/L drives, the larger chassis, the move to Software-Defined Persistent Memory, the cable-free midplane interconnect, and the decoupling of the inter-node fabric from CPU generation are all decisions that look forward rather than sideways. Taken together, they’re what make the Gen 3 chassis a credible foundation for the multi-generational upgrade story Dell is telling with Lifecycle Extension.

There are three new appliance models in the family. The PowerStore 1500 is the single-socket platform with 24 drive bays at launch and a 100 GbE RDMA inter-node fabric. The 5500 and 9500 are dual-socket on the same 3U chassis, with 40 drive bays and 200 GbE RDMA inter-node connectivity; the 9500 offers twice the memory and higher core counts than the 5500. A future controller-swap upgrade will enable the 1500 to scale to 40 drives and 200GbE RDMA. All three run the same PowerStoreOS 5.0 image, share the same OCP 3.0 I/O architecture, and support TLC or QLC media in the same model, with no performance penalty for switching to QLC. This eliminates the need to choose between separate model tiers for different use cases or performance requirements.

The Gen 2 platforms remain available, and existing PowerStore customers have a clear path forward with intelligent clustering. Gen 1, Gen 2, and Gen 3 appliances can coexist in a single cluster with non-disruptive workload mobility, which is the right answer for an install base that doesn’t want to shift platforms every few years.

Dell PowerStore Gen3 Specifications

All three Gen 3 models share a 3U dual-node chassis, the same OCP 3.0 I/O architecture, and a unified PowerStoreOS that supports scale-out block or file out of the box. They differ in CPU socket count, DRAM capacity, drive count, and inter-node bandwidth.

Specification PowerStore 1500 PowerStore 5500 PowerStore 9500
Overview
Positioning Mid Mid / high-end Flagship high-end
Chassis 3U, dual-node
Compute & Memory
CPU platform Intel single-socket Intel dual-socket Intel dual-socket
CPU per appliance 2× 24-core @ 1.9 GHz 4× 24-core @ 1.8 GHz 4× 32-core @ 2.2 GHz
Memory per appliance 512 GB (16 × 32 GB) 1,024 GB (32 × 32 GB) 2,048 GB (64 × 32 GB)
Storage
Drives per base appliance Up to 24 EDSFF Up to 40 EDSFF Up to 40 EDSFF
Drives per expansion (post-RTS) 44 EDSFF
Drive support TLC: 3.84 / 7.68 / 15.36 TB · QLC: 30.72 TB
Min. drive config TLC 6× 3.84 TB · QLC 7× 30.72 TB TLC 6× 3.84 TB · QLC 11× 30.72 TB TLC 6× 3.84 TB · QLC 11× 30.72 TB
Max raw capacity (base) ~737 TB (24 × 30.72 TB) ~1.2 PB (40 × 30.72 TB) ~1.2 PB (40 × 30.72 TB)
I/O & Networking
OCP 3.0 line cards per node 3 (+1 reserved) 5 (+1 reserved) 5 (+1 reserved)
Cross-node interconnect 100 GbE RDMA 200 GbE RDMA 200 GbE RDMA

Dell PowerStore Gen 2 vs. Gen 3

Every major hardware subsystem in Gen 3 series PowerStore appliances has advanced by at least one generation. The most consequential changes for capacity planning are the 3U chassis, the E3.S drive bays, and the jump from 2× 10 GbE to up to 200 GbE on the inter-node fabric.

Dell PowerStore Gen 3 front drive bays

Dell PowerStore Gen3 Fully Populated Front View

The base 3U chassis has 40 drive slots on the 5500 and 9500. The 1500 ships with 24 populated bays, and an additional 16 bays unlocked via a future Data-In-Place upgrade. That works out to 13.3 drives per rack unit, up from 10.5 on Gen 2, or roughly 40% more drives per RU in the base appliance and 83% greater density on the post-RTS 44-drive expansion shelf. The drives themselves are standard E3.S 1T NVMe SSDs from multiple vendors, with no proprietary carrier, which reduces exposure to supply chain constraints or single-source dependencies. Capacities run 3.84 TB, 7.68 TB, and 15.36 TB in TLC, plus 30.72 TB in QLC, and the same set of options is supported across all three models (1500, 5500, and 9500) with no performance trade-off between media types.

Dell PowerStore Elite E3.S drive sled with drive installed

Dell PowerStore Gen3 30TB SSD

The choice between TLC and QLC comes down to $/GB and density rather than performance tier. Minimum drive configurations vary slightly by model: TLC starts at 6× 3.84 TB across the lineup, while QLC starts at 7× 30.72 TB on the 1500 and 11× 30.72 TB on the 5500 and 9500. All drives are SED or FIPS-validated, and the platform supports +1 and +2 Data Resiliency Engine (DRE) configurations. Thermals improve as well, with Dell citing roughly 50% lower cooling requirement than an equivalent 2.5″ deployment, helped by the EDSFF airflow geometry. All 40 bays are also user-addressable for data, since the cache is now handled via Software-Defined Persistent Memory rather than by NVRAM drives that consume front bays, as on Gen 2.

PowerStore Gen 2 PowerStore Elite (Gen 3)
Chassis & Platform
Base chassis 2U, 2-node 3U, 2-node
I/O Bus PCIe Gen 3 PCIe Gen 5
Memory DDR4 DDR5
Storage
Drive form factor Up to 25× 2.5″ U.2 NVMe (dual-ported) Up to 40× E3.S/L 1T NVMe (dual-ported)
Expansion shelf Up to 24× 2.5″ U.2 NVMe Up to 44× E3.S/L NVMe (post-RTS)
Cache strategy U.2 NVRAM drives Cache to local flash (SDPM)
I/O & Networking
I/O slots 3× SLIC / OCP 2.0 (PCIe Gen 3 x16) Up to 5× OCP 3.0 (PCIe Gen 4/5 x16)
Cross-node connect 2× 10 GbE, RDMA Up to 200 GbE, RDMA (400 GbE-ready)
Management & Power
Out-of-band mgmt (BMC) EMC GEM iDRAC (storage version)
Power EMC BBUs / custom PSUs PowerEdge BBU / PSUs

Inside the PowerStore Gen 3 Hardware Platform

Looking at the front of the new 3U PowerStore Gen 3 unit compared to the previous Gen 2, the updated design carries the same look and feel as the 17th-generation PowerEdge servers we have previously reviewed. It ships in Dell’s newer grey colorway with the matching honeycomb bezel that has become a visual signature across the company’s current enterprise hardware lineup.

Dell PowerStore 9500 front bays with drive removed

Dell PowerStore Gen3 With Drive Partially Ejected

Compute & Memory

On the compute side, Intel CPUs power the system, with per-CPU TDP ranging from 165W to 270W and up to 32 cores per CPU on the 9500. Dell quotes up to 50% more cores per node. Memory moves to DDR5 throughout, with the 9500 shipping with 2 TB per appliance (64 × 32 GB DIMMs), the 5500 with 1 TB, and the 1500 with 512 GB. Dell’s shift to DDR5 also significantly increases memory throughput, offering 2x higher bandwidth than previous DDR4 generations.

Dell PowerStore Gen 3 - 1500 cpu and memory

Dell PowerStore Gen3 1500 Controller CPU Heatsink

The internal fabric jumps to PCIe Gen 5, four times the per-lane bandwidth of the Gen 3 fabric used in Gen 2 PowerStore. Socket layout is the main differentiator across the lineup: the 1500 is single-socket, while the 5500 and 9500 are dual-socket, all running the same chassis and the same PowerStore OS software image, with core count and memory as the levers for right-sizing.

Chassis Cooling

Cooling on the 9500, pictured below, is an impressive design. The two CPU coolers on a single controller are connected via heat pipes and extended into a wide, conjoined fin stack, which is a smart move given how much heat a dual-socket controller has to dissipate. An earlier design choice comes into play for cooling: the 3U chassis. Each controller is now 1.5U tall, giving air a longer path to flow through than a 1U tray. This gives each model plenty of cooling headroom throughout the platform’s lifespan. The 1500 uses a more traditional-looking CPU cooler but retains all the airflow benefits of the 5500 and 9500 models. Regardless of the model, cooling across the family is handled by six fans per controller, giving the unit a total of 12 fans that keep the drives and the rest of the unit within their thermal envelope.

Dell PowerStore Gen 3 - 9500 top down view

Dell PowerStore Gen3 9500 Controller Open

Networking and Storage

Compared to previous generations, the new models move to a modular, standardized I/O plane built on OCP 3.0 slots across the family, with up to 5 slots per node on the 5500 and 9500 (+1 reserved for expansion) and 3 per node on the 1500 (+1 reserved). The modules are tool-less hot-swap and replace the proprietary SLIC carrier used in Gen 2. Just as importantly, the I/O fabric jumps to PCIe Gen 5 on the newer models, substantially increasing the per-lane bandwidth available to each card and dramatically raising the networking ceiling. At launch, card options include 4× 32/64 Gb FC, 4× 1/10 GbE, 4× 10/25 GbE, and 2× 100 GbE, with 200/400 GbE Ethernet and 128 Gb FC listed as future I/O card releases. Dell is also strengthening network security, as all Fibre Channel cards will support EDIF (Encrypted Data-in-Flight) through a forthcoming non-disruptive software release. The net result is up to 40 network ports per appliance, double the previous generation and roughly 11% more port density than the Gen 2 controller layout.

Dell PowerStore 9500 rear with OCP cards removed

OCP Gen 2 and Gen 3 Cards on the Rear of the 9500 model with a Quick Latch for Removal.

The controller layout has also been rethought. On Gen 2, the top controller was inverted relative to the bottom one, which made servicing awkward and increased the risk of pulling the wrong controller or grabbing the wrong component during a hot swap. On Gen 3, both controllers are oriented the same way, and either controller can be released using the two handles and levers at the sides of the chassis, sliding out as a single 1.5U tray. The change is small on paper but meaningful for serviceability, particularly in racks where access from above is constrained or where a tech is working under pressure.

Dell PowerStore Gen 3 - 9500 rear with OCP cards installed

Dell PowerStore Gen3 9500 Rear View With FC and Ethernet Cabling

PCIe signaling plays an important role in the new PowerStore Elite, drawing on design elements from current-generation PowerEdge servers. When Dell started offering Gen5 E3.S support on platforms such as the PowerEdge R770 or PowerEdge R7725, it was decided to discontinue the use of PCIe switches on the drive backplane. Older models, such as the PowerEdge R760 with a 24-drive backplane, used a PCIe switch to enable all drive lanes, which increased complexity and cost and reduced drive performance. On models using E3.S, configurations with up to 16 SSDs allocated 4 PCIe lanes per drive, while configurations with up to 40 bays allocated 2 PCIe lanes per drive. Dell applies the same philosophy to the new PowerStore Gen 3 models, with each controller using just 2 PCIe lanes to communicate with each dual-ported E3.S SSD. The remaining PCIe Gen5 lanes are then used for inter-node communication and rear I/O connectivity. Nearly every PCIe lane inside the chassis is utilized, down to the remaining PCIe Gen4 lanes off the CPU chipset, which feed OCP slots that don’t require high bandwidth.

Inter-Node Interconnect

The inter-node fabric is one of the bigger jumps from Gen 2. The 5500 and 9500 run up to 200 GbE RDMA between controllers (the 1500 uses 100 GbE RDMA), compared to just 2× 10 GbE on Gen 2. The links are cable-free, midplane-routed point-to-point connections dedicated to write ingest. Pictured below is one of the 200GbE internal interconnect modules on the 9500.

Dell PowerStore 9500 controller interconnect

Dell PowerStore Gen3 9500 Front Interconnect Card

Importantly, the new interconnect layer is CPU-agnostic and decoupled from processor generations and vendors, positioning the chassis to accept upgrades to future CPU platforms across the unit’s lifecycle without disturbing the drives or re-engineering the fabric.

Dell PowerStore 9500 controller interconnect removed

200 GbE RDMA Controller Interconnect Card Removed

Cache and persistent memory

PowerStore Gen 2 uses front U.2 NVRAM drives for write cache persistence, reducing storage capacity by 4 slots. Gen 3 introduces Software-Defined Persistent Memory (SDPM) instead: standard DDR5 DRAM is presented to the OS as an ACPI-compliant NVDIMM-N, and on power loss, a BIOS SMI handler de-stages the volatile state to the M.2 SSD drive, backed by a hold-up battery subsystem. Each PowerStore Gen 3 array offers substantial lithium-backed power. On the 5500 and 9500 models, each controller has two 54Wh batteries (216Wh total per chassis) while the 1500 model has one per controller (108Wh total per chassis).

Pictured below is the PowerStore 9500 controller with its M.2 boot drive and two battery packs that keep the system powered long enough to commit the state to disk in the event of a power loss.

Dell PowerStore 9500 Hold Battery packs

Dell PowerStore Gen3 Battery Packs

On the PowerStore 1500, we can see the single-socket layout and a single battery pack to match. The same SDPM architecture is in play, just sized appropriately for the smaller controller.

Dell PowerStore Gen3 1500 Internal Controller View

Dell PowerStore Gen3 1500 Internal Controller View

Power and Management

Power and management have been re-platformed onto Dell’s standard server components. The BMC moves from the legacy EMC GEM controller to a storage-tuned iDRAC, aligning PowerStore with the rest of the Dell server portfolio in terms of serviceability. PSUs and battery backup units are now standard PowerEdge parts rather than custom EMC hardware, which should mean better support operations, less downtime, and broader supply. The hold-up power subsystem covers CPUs, DIMMs, drives, fans, and the iDRAC itself, keeping them powered long enough to de-stage the volatile state to disk on a power loss.

Dell PowerStore Gen3 Cooling Fan

Dell PowerStore Gen3 Cooling Fan

PowerStore Gen 3 Performance

Dell shared preliminary numbers comparing Gen 3 to Gen 2. Treat them as directional, but they line up with what the hardware uplift would suggest, given the move to the latest Intel x86, DDR5, PCIe Gen 5, and the 200GbE RDMA fabric between controllers. Compared to the prior generation, Dell targets up to 3x higher IOPS on 8K mixed workloads, up to 3x higher throughput on 1MB sequential reads, and up to 3x higher throughput on 1MB sequential writes, with meaningful latency improvements on both reads and writes. For an independent comparison, Principled Technologies ran the 9500 against a comparable all-NVMe competitor (unnamed in the report) using Vdbench. On an enterprise OLTP-with-analytics workload, the 9500 delivered 834,558 IOPS, compared to the competitor’s 357,427, a 2.33x advantage.

Dell PowerStore Gen3 OLTP workload IOPS Chart

The latency picture is the same shape. At a 310,000 IOPS target for a mixed small-block database workload, the 9500 held at 0.44ms while the competitor sat at 1.22ms, a roughly 64% reduction.Dell PowerStore Gen3 OLTP workload latency Chart

Data efficiency rounds out the comparison and is the one that matters most as NAND pricing climbs. On a dataset built for 2:1 compression and 2.5:1 deduplication at an 8KB dedupe unit, the 9500 hit 6.6:1 overall reduction. The competitor came in at 2.76:1 on the same data.Dell PowerStore Gen3 data reduction Chart

Conclusion

PowerStore Gen 3 is the most significant release since the platform debuted in 2020, and it’s the kind of generational reset that will force the rest of the enterprise storage array market to respond. Dell rebuilt the chassis and modernized the drive form factor, cache architecture, inter-node fabric, I/O plane, and management stack in one massive step. Every one of those choices is forward-looking, which is what makes the ten-year Lifecycle Extension story credible rather than aspirational.

Dell PowerStore Gen3 9500 Fully Populated Front View

Dell PowerStore Gen3 9500 Front View

The design wins are easy to enumerate. Forty E3.S bays in 3U with no slots burned on cache. A 200 GbE RDMA midplane that’s CPU-agnostic and ready for whatever silicon comes next. SDPM in place of NVRAM drives, which both frees front bays for data and removes the dependency on a single persistent memory technology. iDRAC, PowerEdge PSUs, and Dell battery backup units in place of legacy EMC parts, which means PowerStore now benefits from the same serviceability and supply chain as the rest of the Dell enterprise portfolio. And the unified-by-default approach, paired with the ability to support TLC or QLC media in a single model without a performance penalty, simplifies the buying decision in ways that matter for platforms that must serve diverse workloads in the same stack.

We’re looking forward to spending more time with the platform in our lab and seeing how the Gen 3 hardware comes together with everything PowerStoreOS 5.0 brings on the software side. The autonomous data path work, log-structured metadata for high-capacity QLC, I/O-level telemetry, and dynamic block-and-file resource sharing are all the types of enhancements that compound over time on a chassis with this much headroom. It’s that combination, hardware and software moving forward together, that earns this release the Elite name.

PowerStore Product Page

This report is sponsored by Dell Technologies. All views and opinions expressed in this report are based on our unbiased view of the product(s) under consideration.

The post Dell PowerStore Gen 3: Inside the Most Aggressive Enterprise Storage Reset in Years appeared first on StorageReview.com.

Dell PowerProtect Data Domain All-Flash Appliance: The Intel Powered All-Flash Foundation for Cyber Resilience

19 May 2026 at 16:59

Infrastructure for cyber resilience occupies a position in the enterprise stack that primary storage does not. When a system fails, data becomes corrupted, or ransomware locks an organization out of its environment, recovery ultimately depends on the backup platform. That reality has sustained strong demand for purpose-built backup appliances even as cloud-based alternatives have expanded. According to IDC’s 4Q25 Purpose-Built Backup Appliance (PBBA) Tracker, Dell holds the top revenue position in the category, a standing built on the strength of the Dell PowerProtect Data Domain portfolio (based on the Intel® Xeon® processor) and its more than 15,000 active customer deployments worldwide.

dell powerprotect dd9910F hero

We covered the Dell PowerProtect Data Domain DD9410 and DD9910, both utilizing the Intel Xeon Scalable processor, in depth when those systems launched, examining the Data-Less Head architecture, the DDOS software stack, the Hardware Root of Trust and Secure Boot chain, and the performance improvements over the prior Data Domain DD9400/DD9900 generation. The Dell PowerProtect Data Domain DD9910F All-Flash appliance builds on that same foundation without redesigning it. The software architecture, deduplication engine, Data Domain Boost (DD Boost) ecosystem, and security capabilities remain consistent across the family. What changes is the storage medium: the All-Flash appliance replaces spinning disk with flash across the entire appliance, targeting the workflows where that transition has the most direct operational impact.

The case for flash in a platform for cyber resilience differs from that in primary storage. Where flash changes the equation is in restore and replication throughput, and the speed of analytics-driven integrity validation in isolated cyber-recovery vaults. Those are the areas where enterprise recovery SLAs are tightening most aggressively, and they are the focus of this analysis.

We examine the All-Flash appliance, starting with the hardware build, including how Intel’s QAT enables hardware-accelerated compression, then turning to how flash concretely changes restore, replication, and cyber recovery workflows, how the appliance’s data reduction reduces storage footprint and TCO, and where the appliance sits within Dell’s current PowerProtect cyber resilience portfolio.

Inside the Data Domain DD9910F All-Flash Appliance

The PowerProtect Data Domain All-Flash appliance follows the architectural model introduced with the current PowerProtect Data Domain generation, pairing a 2U controller with external storage shelves. Dell refers to this design as a Data-Less Head architecture, in which the controller handles compute, metadata services, and system orchestration while backup data resides on externally attached storage. Separating the compute layer from the storage capacity layer allows the system to scale while maintaining consistent performance characteristics across the Data Domain portfolio.

dell powerprotect dd9910F ssd ejected

The controller occupies a 2U chassis powered by dual 5th Gen Intel Xeon Scalable processors, which drive the Data Domain file system and inline deduplication engine. Large DDR5 memory pools support the platform’s metadata and data reduction workloads. Intel Quick Assist Technology (QAT) is integrated directly into the 5th Gen Xeon silicon, enabling hardware-accelerated compression within DDOS without consuming a PCIe slot or requiring a dedicated add-on card. By offloading compression to Intel QAT’s on-die accelerators, the Intel Xeon processor platform keeps cores available for deduplication metadata processing and data path operations that directly affect restore throughput and replication speed. It can incur increasingly high compute costs as throughput scales. PCIe Gen5 connectivity in the rear expansion slots provides the bandwidth headroom needed to support high-density networking configurations at the top end of the portfolio.

From the rear of the chassis, the system exposes the networking and expansion capabilities required for large enterprise backup environments. Multiple PCIe slots support high-bandwidth networking adapters, including configurations supporting 10GbE, 25GbE, and 100GbE connectivity. These options allow the appliance to scale network throughput depending on deployment requirements, while dedicated connectivity links the controller to the external storage shelves. Power is delivered through dual redundant power supplies configured to maintain operation if a single PSU fails.

dell powerprotect dd9910F rear

Backup capacity for the All-Flash appliance is provided by the FS240 flash enclosure, a 2U storage shelf populated with enterprise SSDs. Each shelf provides high flash capacity in a form factor that is consistent with the controller chassis. The platform supports up to four FS240 shelves, allowing the system to scale from 272TB to 1.1PB of usable capacity depending on configuration. Each enclosure includes redundant power and cooling components consistent with enterprise availability expectations.

dell powerprotect dd9910F drive bays

At a 544TB  configuration, the platform consists of the 2U controller and two 2U flash shelves, for a total footprint of 6U. An equivalent HDD-based Data Domain DD9910 deployment occupies roughly 10U of rack space. This difference helps explain Dell’s claims around improved rack efficiency and reduced power consumption with the all-flash design.

Flash Changes the Recovery Equation

Restore operations, replication throughput, and analytics-driven integrity validation in cyber recovery vaults all rely heavily on read performance and rapid access to deduplication metadata. These are the use cases where performance matters most in practice. The proprietary architecture of the All-Flash appliance for the file system, along with performance tuning, optimizes drive performance.

Dell cites up to 4x faster restore performance for the All-Flash appliance compared with the disk-based DD9910 when configured at equivalent capacity. This performance advantage is a system-level outcome, driven by both the flash storage tier and Intel Xeon scalable processor technology, which manages the metadata-intensive workload required by recovery operations.

Since both arrays share the same controller hardware, all gains are directly attainable in the all-flash configuration of the All-Flash appliance. In testing presented during the platform briefing, the restore workload consisted of multiple data streams and successive backup generations with realistic change rates applied between cycles.

In the test scenario shared by Dell and Intel, restore throughput climbs rapidly, then stabilizes at roughly 60 TB per hour and remains at that level throughout the restore operation. The HDD-based DD9910 exhibits a more gradual performance ramp and lower sustained throughput under equivalent workloads. The difference reflects how flash handles read-intensive operations compared with a spinning disk, particularly when a restore workload requires rapid access to many deduplicated data segments across the storage pool. Data Domain systems store unique data segments once and reference them through metadata pointers for subsequent backups. During a restore, the system must locate and reassemble those segments to reconstruct the requested dataset. Because this process involves numerous read operations and metadata lookups, lower-latency flash storage can significantly accelerate the reconstruction of large backup images.

Replication throughput also benefits from faster reads on the flash storage tier. Dell reports up to 2x faster replication for the All-Flash appliance compared with the DD9910 at similar capacity levels. Replication in Data Domain environments typically involves reading data from the source appliance and transferring it to a secondary system, often located at a disaster recovery site or within a cyber recovery vault.

Reducing replication time carries operational implications beyond raw throughput. In cyber recovery architectures, backup data is replicated into an isolated vault environment designed to protect against ransomware. The vault is exposed to production systems only during tightly controlled replication windows. Shorter replication windows reduce the period during which the vault must be accessible to the production environment, narrowing the potential exposure window.

Dell also reports improvements in analytics-driven validation workloads performed within cyber recovery vaults. Using CyberSense analytics in internal testing, the company cites up to 2.8x faster analytics performance on the All-Flash appliance compared with the disk-based DD9910.

CyberSense analytics workflows scan backup data to verify integrity and detect potential signs of corruption or ransomware activity before recovery operations begin. Because the validation process involves scanning large volumes of backup data, the improved read performance of flash media reduces the time required to confirm that a recovery point is clean before initiating a restore.

Power, Space, and the Operational Case for Flash

Performance improvements are the most visible benefit of replacing disks with flash in a cyber resilience platform. Still, the efficiency gains carry independent weight for enterprise buyers managing infrastructure at scale. The Dell internal testing comparing the Intel-powered DD9910F against the disk-based, Intel-powered DD9910 at equivalent capacity indicates up to 80% lower power consumption and a 40% reduction in rack space. In isolation, those figures read as spec sheet claims. Across a multi-site data protection deployment, they translate into a different kind of conversation.

A large enterprise running Data Domain appliances at a primary site, a disaster recovery site, and an isolated cyber recovery vault operates three distinct infrastructure footprints. Power and cooling costs at each of those locations are real line items, and rack density affects what fits in a given facility without additional build-out. The shift from roughly 10U to 6U per fully configured system, combined with significantly lower power draw, changes the infrastructure math for organizations that are either constrained on data center capacity or actively managing energy costs as part of their total cost of ownership calculations.

The cyber recovery vault use case is worth calling out specifically. Vault infrastructure is, by design, isolated and purpose-built, often deployed in a dedicated cage or colocation environment, where power and space are billed directly. Reducing the physical and power footprint of vault infrastructure without sacrificing recovery performance is a meaningful operational benefit in that context, and one that compounds as retention requirements grow and vault capacity scales over time.

Data Reduction and Effective Capacity

While the introduction of flash changes the performance profile of the All-Flash appliance, the underlying efficiency model remains rooted in the Data Domain deduplication engine. Dell now cites data reduction ratios of up to 75:1 for the current Intel Xeon-based PowerProtect Data Domain generation, extending the platform’s long-standing focus on capacity efficiency.

This figure is not purely theoretical. In testing conducted by Prowess Consulting using a representative VMware backup workload, the Data Domain All-Flash appliance achieved approximately 75:1 data reduction, while competing systems required up to 3.8× more physical capacity to protect the same dataset. Over time, as additional backup copies accumulated and redundancy increased, the effective reduction ratio continued to improve, reaching 78:1 after seven days and exceeding 100:1 after two weeks in that test environment.

The behavior reflects how deduplication operates in backup environments. Initial backup copies contain a higher proportion of unique data, while subsequent backups introduce incremental changes that can be efficiently deduplicated against existing data sets. As a result, effective data reduction improves over time as retention periods extend.

From a practical standpoint, this reduction capability is central to how the All-Flash appliance balances flash economics with large-scale backup requirements. While the All-Flash appliance’s raw usable capacity ranges from 272TB to 1.1PB, the disk-based Data Domain DD9910 scales to higher raw capacity levels, up to 2.1PB. Effective capacity in both cases can be significantly higher depending on data characteristics, change rates, and retention policies. Dell positions this efficiency as a key factor in reducing overall infrastructure footprint, power consumption, and long-term storage costs.

Management and the Data Domain System Manager Interface

Management of the All-Flash appliance is handled through the Data Domain System Manager, the same web-based interface used across the Data Domain family. DDOS 8.7 runs identically across the portfolio, so the management experience carries over without retraining, and the All-Flash appliance drops into existing workflows without introducing a new operational model.

 

The dashboard surfaces what administrators reach for most often without requiring navigation: filesystem capacity, used and available space, compression factor, last write time, active alerts, and licensed services status.

Real-time performance charts are one click away, covering CPU utilization, DD Boost throughput, active connections, filesystem operations, and network activity. In active environments receiving backup and replication traffic simultaneously, those charts give administrators an immediate read on whether the platform is performing within expected parameters.

The Data Management section organizes the operational details below the dashboard. The filesystem view covers capacity, usage, and compression at the system level, while M-trees provide a granular layer. Every backup policy from each connected application creates a unique M tree, logically separating workloads within the system. Per-M tree views show space utilization, daily write patterns, and pre- and post-compression breakdowns over configurable time windows, which matter when diagnosing whether a specific workload is compressing as expected or consuming capacity faster than anticipated.

Replication, protocols, hardware, and administration each have dedicated sections in a layout consistent with prior Data Domain generations. The Replication section consolidates pair status, sync state, and job health across connected systems, with both scheduled and on-demand replication managed from the same screen. Protocol configuration covers DD Boost, CIFS, NFS, and VTL in a single section, with DD Boost serving as the primary path for enterprise backup applications and exposing active connections, client lists, plugin versions, and authentication settings in a single view.

The interface reflects a platform designed for use by people who are not exclusively backup specialists. The backup administrator role has contracted meaningfully over the past decade as infrastructure teams consolidate and generalists take on broader responsibilities. The DD System Manager has been refined to remain accessible in that environment, with logical structure, clear real-time reporting, and short paths from alert to diagnosis.

Where the Data Domain All-Flash Appliance Fits in the PowerProtect Portfolio

The current Data Domain lineup spans five hardware models. The Data Domain DD3410 serves small and medium businesses and the Data Domain DD6410 serves medium businesses. At the same time, the Data Domain DD9410 and DD9910 address larger environments, where capacity and throughput requirements scale accordingly. The All-Flash appliance and the All-Flash Ready Node bring flash into the portfolio at different points and for different deployment scenarios.

The All-Flash appliance targets large enterprise environments where recovery SLAs are aggressive, and the operational cost of slow restores or extended replication windows is measurable. It sits alongside the disk-based Data Domain DD9910 rather than replacing it, with the two systems sharing the same software stack and integration ecosystem while differing primarily in storage medium and the performance and efficiency characteristics that follow from that choice. Organizations with large retention requirements and cost-sensitive capacity economics will likely continue to find the Data Domain DD9910 a better fit. Those prioritizing restore speed, vault analytics performance, and power and space efficiency at the high end of the portfolio have a direct path to the All-Flash appliance.

The All-Flash Ready Node is a separate product category for organizations building software-defined or hyperconverged environments, in which Data Domain capabilities are delivered through a customer-supplied server platform rather than a purpose-built appliance. It is not a like-for-like alternative to the DD9910F.

For most large enterprise evaluations, the practical decision point is between the Data Domain DD9910, which uses Intel Xeon Scalable processors, and the All-Flash appliance at an equivalent capacity. The all-flash model carries a higher acquisition cost. Still, Dell positions the reductions in power, cooling, and rack space as meaningful offsets when evaluated against the total cost of ownership over a multi-year deployment.

Conclusion

The Intel-powered Dell PowerProtect Data Domain All-Flash appliance makes a clear case for where flash belongs in data protection. Backup ingestion is largely sequential and throughput-bound, and flash does not fundamentally change that equation. Recovery is where flash matters, and the All-Flash appliance applies it precisely there: up to 4x faster restores, up to 2x faster replication, and up to 2.8x faster CyberSense analytics versus the disk-based DD9910 at equivalent capacity. Each of those translates into a concrete operational outcome, shorter recovery windows, narrower cyber vault exposure periods, and faster integrity validation before data returns to production.

The efficiency story is the second half of the argument. A 6U all-flash footprint replacing roughly 10U of disk-based infrastructure, combined with up to 80% lower power consumption, changes the math for organizations running data protection across primary, DR, and cyber vault sites. In vault deployments specifically, where power and rack space are billed directly, and isolation requirements make every rack unit count, those reductions compound into deep operational savings over a multi-year deployment. The data reduction engine reinforces the economics from the other direction, with effective ratios reaching 75:1 and continuing to improve as retention increases.

What makes the All-Flash appliance work well is that none of this comes at the cost of operational continuity. The architecture has defined Data Domain for years. The DDOS software stack, the DD Boost ecosystem, the security model, and the System Manager interface all carry over unchanged. For the more than 15,000 organizations already running Data Domain, the All-Flash appliance is an easy decision for mission-critical workloads.

For enterprise environments evaluating where flash fits in their data protection strategy, the All-Flash appliance answers that question directly. Dell did not redesign Data Domain to accommodate flash. It applied flash where the architecture benefits most, delivered measurable gains in the workflows where recovery SLAs are tightening fastest, and held the operational model steady.

Dell Technologies Cyber Resilience

This content was produced in partnership with Dell Technologies and Intel. All analyses and conclusions are based on StorageReview’s independent evaluation.


Intel® Xeon® is a trademark of Intel® Corporation or its subsidiaries

The post Dell PowerProtect Data Domain All-Flash Appliance: The Intel Powered All-Flash Foundation for Cyber Resilience appeared first on StorageReview.com.

LaCie 8big Pro5 Review: 256TB of HAMR-Powered Thunderbolt 5 DAS

23 April 2026 at 20:53

LaCie has been a fixture in our lab for well over a decade. From the 8big Rack Thunderbolt 2 we covered in 2014 through the many generations of 5big, 6big, 8big, and Rugged devices that have followed, the formula has been consistent: premium Neil Poulton-designed enclosures, Seagate drives inside, Mac-centric polish, a solid warranty, and a clear focus on creative professionals. The new LaCie 8big Pro5 carries that pedigree forward in build quality, design, and purpose, and arrives at a notable inflection point for high-capacity direct-attached storage.

With eight 32TB HAMR-based Seagate IronWolf Pro drives on board, the 8big Pro5 tops out at 256TB of raw capacity. As far as turnkey desktop DAS products go, nothing else on the market ships at that capacity today. Competing 8-bay Thunderbolt enclosures from OWC, Sabrent, and others cap out at around 192 TB with the previous-generation PMR drives. While it is technically possible to roll your own by pairing a bare enclosure with eight 32TB IronWolf Pros, that DIY route leaves you stitching together the warranties across vendors. Seagate backs the complete LaCie kit end-to-end, including the drives, which is an advantage at this capacity point and for the value of the workloads involved.

Heat-assisted magnetic recording has been more than two decades in the making, and it has finally moved from hyperscale sampling to a product that a creative professional can put on a desk. For teams working with multi-stream 4K and 8K RAW footage, large photogrammetry or virtual production asset libraries, or AI-assisted content pipelines that consume storage faster than any prior generation, the jump from 24TB-era PMR drives to 32TB HAMR in the same eight bays is a meaningful change. We walked through the technical foundations of HAMR with Seagate’s Colin Presly on Podcast #124: The Path to 50TB HDDs with Frickin Lasers. The roadmap Colin laid out then is now shipping as product, with Mozaic 3+ drives at 30TB and up, Mozaic 4+ pushing to 44TB, and a longer arc toward 100TB drives as platter density continues to climb.

Around that storage core, LaCie delivers the rest of the package you would expect. The 8big Pro5 connects via Thunderbolt 5, which Seagate quotes at up to 80Gbps bidirectional for data, with additional headroom when combined with display traffic. In practice, the ceiling for a hard-drive array is set by the drives themselves. The IronWolf Pro 32TB is rated for up to 285 MB/s sustained, so eight drives in parallel have a theoretical maximum of about 2.2 GB/s before caching effects are taken into account.

The host port delivers up to 140W of power to a connected laptop, with two downstream Thunderbolt 5 ports rated at 30W each and a USB 20Gbps port rated at 15W for daisy-chained peripherals and displays. The LaCie 8big Pro5 ships preconfigured as a single RAID 5 array for 224TB of usable capacity, with RAID 0, 1, 6, 10, 50, and 60 available through LaCie RAID Manager. Build quality, thermals, and design are vintage LaCie, which we will cover in detail throughout the rest of this review. Pricing starts at $5,979 for the 32TB base configuration, with SKUs available up to 64TB, 128TB, 192TB, and 256TB.

LaCie 8big Pro5 – Build and Design

At the front of the LaCie 8big Pro5, the unit features a clean, minimal industrial design that aligns with its professional focus. It measures 11.69 inches in length, 9.13 inches in width, and 8.46 inches in height, giving it a compact yet substantial footprint for an eight-bay system.

Our review unit shipped fully populated with eight of Seagate’s new IronWolf Pro 32TB drives, for a total raw capacity of 256TB. With all drives installed, the system weighs just over 29 pounds, underscoring both its density and solid construction.

The enclosure itself is crafted from a single-piece aluminum chassis finished in metallic gray, giving it a premium, durable feel. Up front, each drive bay is tool-less, allowing quick, easy access to swap or service drives. Each tray is paired with an individual status LED, providing clear, at-a-glance visibility into drive activity and health without requiring interaction with the software.

At the rear, the LaCie 8big Pro5 maintains the same clean, functional design, with heavy perforations across the back panel to support airflow in a fully populated chassis. Power is handled via a standard C19 input and a physical power switch, confirming that the power supply is fully integrated into the unit rather than relying on an external brick.

Connectivity centers on four USB-C ports, each clearly labeled for its role. The leftmost port serves as the primary host connection, operating over Thunderbolt 5 with up to 80Gbps bandwidth and delivering up to 140W of power, making it well-suited for powering and connecting a laptop with a single cable.

Next to it are two additional Thunderbolt 5 downstream ports. These ports enable expansion beyond the enclosure, supporting external storage devices or displays while also delivering up to 30W of power to connected peripherals. This makes the unit function as both a high-capacity storage array and a compact docking hub.

The final USB-C port supports a 20 Gbps connection, intended primarily for additional storage expansion. It also provides up to 15W of power, which is sufficient for bus-powered drives and similar accessories.

To round things out, there is a Kensington lock slot for physically securing the device, a practical addition for shared workspaces or studio environments where the unit may not always be in a controlled rack or locked room.

From a wider rear view, the airflow design becomes much more apparent. The majority of the back panel is perforated, allowing the system to move a significant amount of air across all eight drives. Cooling is handled by a three-fan setup, with two larger fans serving the primary drive bay area and a smaller fan dedicated to the lower section housing the controller and power components. This separation helps ensure consistent airflow across both the storage and internal electronics. This is especially important in a fully populated 256TB configuration where thermal buildup can become a limiting factor over sustained workloads.

You can also see the subtle branding here, with “LaCie – design by Neil Poulton” centered along the upper portion of the rear panel, reinforcing the industrial design heritage that has been a hallmark of LaCie systems for years.

Up top, LaCie adds a simple yet practical touch with the integrated handle cutouts. Machined directly into the aluminum, these recessed grips provide a secure way to lift and move the unit without compromising the clean design language.

Given that the system weighs just over 29 pounds when fully populated, a built-in grip like this makes a noticeable difference during deployment or repositioning. It is a small detail, but one that reflects an understanding that this is not a lightweight desktop accessory and will occasionally need to be handled with a bit more care.

LaCie 8big Pro5 – LaCie RAID Manager software

To manage the 8big Pro5’s storage configuration, LaCie requires its RAID Manager software. This utility is available for Windows and macOS and is necessary to configure the array in RAID modes or switch the unit to JBOD, depending on your deployment needs.

Through RAID Manager, users can choose from a full range of RAID levels, including RAID 0, RAID 1, RAID 5, RAID 6, RAID 10, RAID 50, and RAID 60. This flexibility allows the unit to be tailored for everything from maximum performance to high levels of redundancy and fault tolerance. As shown here, a RAID 5 configuration using all eight 32TB drives yields 224TB of usable capacity and provides single-drive fault tolerance through parity.

In addition to RAID configuration, the software also allows you to format the array in either APFS for macOS environments or NTFS for Windows deployments, making it easy to integrate into mixed or platform-specific workflows. The interface itself is straightforward, providing visibility into drive status, serial numbers, and overall array health, while also confirming valid configurations before deployment.

LaCie 8big Pro5 – Performance

For Windows testing, we leveraged a Dell Pro Max 14 with the following configuration:

  • Intel Core Ultra 9 285H
  • NVIDIA RTX PRO 2000 8GB GDDR7
  • 64GB LPDDR5X-8400
  • 1TB SSD

For macOS testing, we used an M4 MacBook Air.

To evaluate the performance of the 8big Pro5, we began testing in a Windows environment with ExFat, configuring the array in RAID 5. This setup reflects a common balance of capacity, performance, and redundancy for general-purpose use. In this configuration, we ran a series of benchmarks, including IOMeter for synthetic workload analysis, Blackmagic Disk Speed Test for media-focused throughput, and PCMark 10 Disk Benchmark to capture more real-world application behavior.

After completing Windows testing, we switched to a macOS environment using RAID 5 and ExFAT. This allowed us to measure the performance of the same configuration across Windows and Mac environments. In this configuration, we reran Blackmagic Disk Speed Test to compare results in a macOS-native workflow and added ATTO Disk Benchmark to analyze performance across varying transfer sizes.

Blackmagic Disk Speed Test

The Blackmagic Disk Speed Test benchmarks a drive’s read and write speeds to estimate its performance, especially for video editing tasks. It helps users ensure their storage is fast enough for high-resolution content, such as 4K or 8K video.

The Blackmagic results show clear, real-world performance gains across RAID configurations. In RAID 5 in Windows, the 8big Pro5 delivers 1,418.4 MB/s read and 2,061.5 MB/s write speeds, offering a strong balance of performance and data protection. When moved to macOS, read performance remains nearly identical at 1,414.9 MB/s, while write speeds are 1,751.3 MB/s, reflecting some platform differences rather than a limitation of the array itself.

Looking at the Blackmagic workload breakdown, RAID 5 still proves more than capable for high-resolution media workflows. At these speeds, the array comfortably supports formats up through 8K, including 8K DCI and even 12K playback in several codecs, with consistent results across ProRes 422 HQ and H.265. This reinforces that RAID 5 is not just a safe option, but a practical one for professional video editing where both performance and redundancy matter.

In practice, RAID 5 delivers more than enough performance for demanding video workflows while maintaining data protection.

Blackmagic (higher is better) LaCie 8big Pro5 – Windows Raid 5 ExFat LaCie 8big Pro5 – macOS Raid 5 ExFat
Read 1,418.4 MB/s 1,414.9 MB/s
Write 2,061.5 MB/s 1,751.3 MB/s

PCmark 10 Storage

PCMark 10 Storage Benchmarks evaluate real-world storage performance using application-based traces. They test the system and data drives, measuring bandwidth, access times, and consistency under load. These benchmarks offer practical insights beyond synthetic tests, enabling users to compare modern storage solutions effectively.

The PCMark 10 result of 717 gives a useful look at how the 8big Pro5 behaves under real-world workloads rather than pure synthetic throughput. This benchmark incorporates traces from everyday applications, which tend to be more sensitive to latency and mixed I/O patterns than large sequential transfers.

PCmark 10 Storage (higher is better) LaCie 8big Pro5 – Windows Raid 5 ExFat
Overall Score 717

IOMeter

We also ran the LaCie 8big Pro5 array through IOMeter. This lets us dig deeper into workloads, including random and sequential performance. We tested the 8big with a single queue to simulate lighter use and with four queue to see how the DAS handles heavier, more demanding scenarios.

At 1 queue, sequential performance is 1,752.2 MB/s read and 1,851.5 MB/s write, showing strong throughput even under a lighter load. Random 2MB performance lands at 233.8 MB/s read, and 654.1 MB/s write, while small-block 4K operations reach 297 IOPS read and 5,482 IOPS write.

IOMeter (1  queue) LaCie 8big Pro5 – Windows Raid 5 Raw
Seq 2MB Read 1,752.2 MB/s
Seq 2MB Write 1,851.5 MB/s
Random 2MB Read 233.8 MB/s
Random 2MB Write 654.1 MB/s
Random 4K Read 297 IOPS
Random 4K Write 5,482 IOPS

Scaling to 4 queue, sequential reads increase to 1,949.1 MB/s, while writes remain steady at 1,873.6 MB/s, indicating the array is already near its write ceiling. Random 2MB performance improves more noticeably, with reads rising to 391.1 MB/s and writes to 980.5 MB/s. For 4K workloads, reads scale to 1,103 IOPS, while writes settle at 4,458 IOPS.

IOMeter (4 queue) LaCie 8big Pro5 – Windows Raid 5 Raw
Seq 2MB Read 1,949.1 MB/s
Seq 2MB Write 1,873.6 MB/s
Random 2MB Read 391.1 MB/s
Random 2MB Write 980.5 MB/s
Random 4K Read 1,103 IOPS
Random 4K Write 4,458 IOPS

ATTO Disk Benchmark Summary (LaCie 8big Pro5 – macOS RAID 5, ExFat)

The ATTO results provide a clear picture of how the 8big Pro5 behaves in macOS when pushed to maximum throughput across a wide range of transfer sizes in a RAID 5 configuration.

At lower transfer sizes, performance ramps up gradually, as expected for an HDD-based array. Small-block operations (under 16KB) remain relatively modest, but once you move to larger transfer sizes, the system scales more effectively.

From around 64KB onward, throughput stabilizes and becomes a far more representative measure of real-world performance. Peak read speeds reach approximately 3.4 GB/s, while write performance settles slightly lower in the 2.7-3.1 GB/s range across larger block sizes.

Overall, the results show strong sequential performance, with the array delivering high read throughput and slightly lower, but still consistent, write speeds under sustained workloads.

Conclusion

The LaCie 8big Pro5 marks a meaningful leap forward for the line. At 256TB raw over Thunderbolt 5, with eight HAMR-based IronWolf Pro drives housed in a well-designed Neil Poulton enclosure, it is the first turnkey desktop DAS to deliver both a massive capacity jump and next-generation interface bandwidth to creative pros in a single box. The 8big formula is all here: premium build, thoughtful thermals, quiet operation, mature RAID management through LaCie RAID Manager, and a clear focus on the video, photo, and 3D asset workflows that have consistently outpaced the storage they rely on.

Performance lands where a well-tuned eight-bay array should. In RAID 5, the array comfortably handles multi-stream 4K and 8K editing with room to spare. Small-block random performance is modest, as expected for any HDD-based array, but that is not the workload profile this product is built for. For bulk sequential transfers, active project storage, and long-form media ingest, the array delivers the throughput that modern creative workflows need. The Thunderbolt 5 host port with 140W of power delivery, plus the two downstream TB5 ports and the 20Gbps USB-C, also make the unit a legitimate one-cable docking solution for a laptop-based edit bay, not just a storage target.

Pricing starts at $5,979 for the 32TB base configuration and scales up through 64TB, 128TB, 192TB, and 256TB tiers. That is a meaningful investment, but a 5-year warranty that covers both the enclosure and the drives end-to-end, Rescue Data Recovery Services, and the operational simplicity of a single-box deployment distinguish it from a DIY build using bare IronWolf Pros and a third-party enclosure. For creative professionals, production teams, and studios working at 4K, 8K, and beyond, and for anyone whose project data has outgrown what previous-generation PMR arrays could deliver in the same footprint, the 8big Pro5 is the most capable turnkey desktop DAS available today and earns the shortlist spot for high-end workflows that need both the capacity and the interface to match.

Product Page – LaCie 8big Pro5

The post LaCie 8big Pro5 Review: 256TB of HAMR-Powered Thunderbolt 5 DAS appeared first on StorageReview.com.

How Metrum AI and Oregon State University Are Building the New Standard for Academic Assessment

23 April 2026 at 18:22

When we published our story on Oregon State University’s plankton imaging research last November, the headline was the science: AI-accelerated infrastructure aboard research vessels, processing terabytes of ocean data in near real-time before the ship ever reached port. But something else happened quietly in the weeks that followed. Word spread across campus about what a single Dell PowerEdge XE7745 with eight Solidigm D5-P5336 E3.S SSDs and NVIDIA RTX PRO 6000 GPUs could actually accomplish. Other departments started asking questions. Then they started making calls. Christopher Sullivan, Director of Research and Academic Computing at OSU’s College of Earth, Ocean, and Atmospheric Sciences, now wants a rack of these servers to meet the growing AI demand across the university, and the story driving that ambition goes well beyond plankton.

Oregon State has established itself as one of the most forward-thinking universities in the nation in its adoption of AI for both research and academic use. The infrastructure decisions being made on campus today, along with the formation of partnerships with companies such as Metrum AI, Dell, NVIDIA, and Solidigm, are not just academic experiments. They lay the groundwork for a new way for universities to deliver education, assess learning, and protect their students. This is the story of how that model was developed.

The Problem That Generative AI Made Worse

For decades, written assignments were central to academic evaluation. Submit a paper, show understanding, and get a grade. Generative AI has fundamentally altered that system. Now, a student can craft a polished, well-structured essay with little real engagement with the material, and even seasoned faculty can’t reliably tell if it is authentic work. The evidence of genuine understanding that universities relied on for generations has weakened.

The obvious alternative is oral evaluation. Ask students to explain their reasoning out loud, walk through their analysis, and defend their conclusions. That is hard to fake. The problem is scale. A professor teaching 200 students cannot sit across from each one and conduct a substantive oral exam. In the modern university, that constraint has effectively shelved oral assessment as a primary evaluation tool. Metrum AI was built to change that equation.

What Metrum AI Built

Metrum AI, co-founded by CEO Steen Graham and CTO Chetan Gadgil, was built around a simple conviction: AI should do real operational work, not just demonstrate potential. The company deploys multimodal AI agents that reason across video, audio, documents, and structured data for customers in industries from insurance to manufacturing. Metrum has developed a close partnership with Dell Technologies, validating its platforms against Dell’s enterprise server infrastructure across a range of GPU configurations. The academic evaluation system at Oregon State is not a pivot for Metrum; it is the same underlying capability applied to a new problem domain, with the same on-premises, human-in-the-loop design philosophy that runs through everything the company builds.

metrum ai oregon state workflow

Applied to academic assessment, the platform processes recorded student video presentations using multimodal AI and returns rubric-aligned draft evaluations for faculty review. The pitch is specific: give instructors an AI partner that handles the repetitive, time-consuming extraction work so they can focus on the judgment calls that truly require human adjudication.

At a functional level, the platform carries out three operations. It extracts multimodal artifacts from submitted videos, generating timestamped audio transcripts with OpenAI Whisper and capturing slide content via visual analysis powered by Qwen3-VL-30B. It then applies instructor-designed rubrics to the extracted content, using Qwen3-30B-A3B reasoning models that run on vLLM. Finally, it presents draft evaluations with evidence pointers, linking each score to a specific transcript timestamp or slide identifier for faculty review and approval before anything reaches a student.

metrum ai oregon state screen

That final step is crucial. No score, comment, or piece of feedback is visible to students until an instructor has reviewed it, made any necessary modifications, and explicitly approved it. The system is built around faculty authority. Additionally, the platform operates entirely on-premises. This decision influences everything about how the system functions, who trusts it, and what hardware it requires.

metrum ai oregon state screen 2

From a Professor’s Side Project to a Provost Mandate

Jonathan Kalodimos is an Associate Professor of Finance and the Harley and Brigitte Smith Fellow in the College of Business at Oregon State University. His background is not what you might expect from someone at the center of an AI infrastructure story. Before joining OSU, he was a financial economist at the U.S. Securities and Exchange Commission, where he served as lead economist on Dodd-Frank Act Section 954, which established rules around executive compensation clawbacks. His research on corporate governance and financial regulation has been cited in The Wall Street Journal, The New York Times, Bloomberg, and the Harvard Business Review. He also, it turns out, has a physicist’s instinct to measure things precisely.

metrum ai oregon state Jon

About a year ago, Kalodimos coded a simple tool for his MBA class: an AI agent to evaluate the oral component of case study presentations. The students were impressed with the quality of feedback. He presented the project during AI Week at Oregon State. Dell took notice, connected him to Metrum AI, and a classroom experiment became something much larger.

“Once you have the tool, you can refine your teaching style and your teaching methods to leverage the strength of the tool to provide a better educational experience.”

— Jonathan Kalodimos, Associate Professor of Finance and Harley & Brigitte Smith Fellow, Oregon State University

What Kalodimos is building toward is what he calls evidence-based extraction, underpinned by what he describes as rubric engineering. This encompasses determining which features are extractable from a student presentation, aggregating those features into learning outcomes, and providing faculty with a structured view of where each student demonstrated understanding and where they fell short. “The way I explain this to skeptical students,” he said, “is if I had a very detailed checklist, and I went through your presentation checking off things you did, that’s what the system is doing. Obviously way more sophisticated than that, but it’s allowing me to see all the opportunities for the student to demonstrate that they know this material.”

He offered two examples that illustrate what the system changes in practice. In the first, a student condensed a ten-minute presentation into five minutes, spoke in a monotone, and spoke at a pace as if English were a second language. His delivery obscured his comprehension entirely. “Even though I was listening carefully,” Kalodimos said, “I just couldn’t, or wouldn’t, break it down into that level of granularity to overcome the delivery element so I could focus on the actual evidence.” When he later walked through the AI-generated evidence breakdown in fifteen-second increments, it became clear the student understood the material. The delivery had been graded, not the knowledge.

The second case involved a student who built a presentation slowly, with what seemed like disjointed slides, and only pulled the argument together on the final slide. Watching live, Kalodimos had already formed a low opinion of the presentation. The system evaluated the work as a complete arc and scored it well. “I didn’t even think that would be a benefit of this type of evaluation,” he said. “It’s getting away from the time element of evaluation.”

Christopher Sullivan, Director of Research and Academic Computing, is the infrastructure lead on the deployment. His involvement sharpened when a critical compliance gap emerged in the original Metrum and Dell design.

FERPA, Data Sovereignty, and Why the Cloud Is Not the Answer

When Sullivan stepped in to build OSU’s on-premises implementation of the Metrum platform, the first thing he identified was a problem nobody had fully solved: The Family Educational Rights and Privacy Act (FERPA).

FERPA is the federal law governing student education records. It establishes strict requirements around who can access student data, under what conditions, and how it must be protected. For a system like Metrum’s, one that ingests student video submissions, generates transcripts, produces evaluations, and stores the complete history of every grading decision, FERPA compliance is not a checkbox; it’s an architectural constraint.

“We needed to be able to bring something on-premises that would meet all of my FERPA conditions,” Sullivan said, “but also have a large amount of storage space.” Cloud processing was not compatible with that requirement. Routing student video files, audio transcripts, and evaluation records through external AI APIs would mean transmitting personally identifiable student information to third-party systems outside the university’s direct control. The contractual and technical complexity of maintaining FERPA compliance in that environment, across every vendor in the chain, made it a non-starter.

There is also a practical student experience dimension. Students submitting recorded presentations are offering something personal: their voice, their face, their reasoning under pressure, sometimes in their non-native language. When they understand that their video is stored on an OSU server, processed by a model running on OSU hardware, and governed by OSU’s own data policies, the dynamic changes. Kalodimos saw this play out directly during the pilot. “The idea that this was a local model with local storage and OSU has the student’s back,” he said, “was palpable. We really need to use the institutional trust that OSU has built, protecting our students, leveraging these on-prem solutions.”

Cloud AI platforms are easy and quick to deploy, but they require institutions to place trust in a contract rather than in their architecture. For students who are already wary of how their data is managed, that distinction can significantly influence their willingness to adopt. On-premises deployment isn’t just about compliance; it establishes a foundation of trust.

The Pilot, the Provost, and What Comes Next at OSU

The pilot is underway. Approximately 500 students across multiple sections are submitting final projects at the close of finals week. Graded evaluations must be returned within 4 days of the last submission. The AI-generated reports have to be ready before faculty begin grading. “There’s a human component running in parallel,” Kalodimos noted, “but they need the report first. If it takes two days to process all of these, then the human element is even more compressed.” The pressure is real, the deadline is fixed, and the infrastructure is doing its job.

The pilot has surfaced something else worth noting. When a professor was recently promoted to an administrative role mid-term, an instructor had to step in on short notice and finish out the course. Having a consistent AI evaluation framework already in place, with defined rubrics and an established review workflow, gave that instructor a thread of continuity that otherwise would not have existed. “Having a consistent AI evaluation companion,” Kalodimos said, “is going to definitely improve the student experience” in exactly those situations when continuity of human instruction cannot be guaranteed.

The story eventually reached OSU’s Provost. Kalodimos presented the full stack: the Dell system, Solidigm storage performance, developer capabilities, and infrastructure benchmarks, in what was supposed to be a ten-minute meeting. It ran for forty minutes. The Provost followed up with an email to the CIO, the CTO, and Sullivan. OSU is now planning a university-wide deployment to make the resource available to faculty starting the spring term, managed by a newly defined Research Computing office that sits under the Provost and the research office.

Sullivan is thinking about that deployment in terms of rack-scale infrastructure. The same XE7745 platform that anchored the plankton imaging work and is now powering the Metrum AI evaluation pipeline is the foundation he wants to scale. The goal is a rack of these servers, available to float between academic compute and research compute workloads as demand shifts. Ideally, the servers would be dedicated to the Metrum evaluation pipeline during midterm and final submission surges, and redeployed to research workloads during quieter periods of the academic calendar. “We can take machines from that set and shove them into the academic compute side for a period of time, and then bring them back and leverage them for the research compute,” Sullivan said. “We want to be able to redeploy them on the fly.”

The organic adoption is already underway. Faculty from the College of Health and the College of Engineering have independently approached Kalodimos after hearing about the project through informal channels. The platform has not been formally announced beyond the pilot. It found its audience anyway.

The Capacity Problem Behind the Grading Problem

There is a version of this story that is only about speeding up grading. That is an incomplete vision.

The larger version is about class capacity. Sullivan described a 100-level geology course, a class that OSU treats as part of its core educational mission and that every student is meant to take. It currently runs two sections of 300 students each, for a total of 600 per quarter. The instructors are at their limit. Adding sections is not feasible given current teaching loads. “I can’t have the teachers do more work,” Sullivan said. “I need to create pathways for us to either create more sections by reducing the load, or put more students in the sections we’ve got.”

Kalodimos similarly framed the College of Business dimension. Professors with sections capped at 45 students for fire code reasons would have the option to explore large-lecture formats with breakout-room support once individualized evaluation can scale. “It’s not just about packing bodies,” he said. “It’s about maintaining quality while exploring different delivery modes.” The AI evaluation layer is what makes individualized assessment at a lecture-hall scale operationally possible.

“The AI is helping us increase the numbers without changing the impact or the message or what’s being learned.”

— Christopher Sullivan, Director of Research and Academic Computing, Oregon State University

Storage Was the Missing Piece

When Sullivan assessed what it would take to bring the Metrum system on-premises at OSU with full FERPA compliance, the GPU side of the equation was already established. The Metrum and Dell reference architecture had demonstrated that the XE7745 with NVIDIA RTX PRO 6000 GPUs could handle the inference workload at scale. What remained unsolved was storage.

The XE7745 is a 4U air-cooled platform optimized for GPU density. That design is its strength, but it comes with a real constraint: drive bay count is limited. “I needed to put a lot of space into a single piece of equipment without compromising speed,” Sullivan said, “because I didn’t want to lose all the value of the GPUs and everything that XE7745 was worth. And there really weren’t a lot of large-capacity SSD solutions out there to do that in the box.”

The storage layer in a system like this carries more than the headline AI workload. Video files arrive from the student portal and need somewhere to buffer immediately. Extracted audio tracks and timestamped transcripts are stored as discrete artifacts for faculty review. Slide images and OCR output occupy their own tier. The Supabase database tracking submission metadata, draft evaluations, faculty edits, and approval records runs continuously. Model weights for Whisper, Qwen3-VL, and the reasoning model need to load quickly enough to avoid inference bottlenecks. And the full audit trail for every AI-generated draft, every faculty override, and every approval action must be retained as a queryable record for accreditation reviews, academic integrity investigations, and administrative reporting.

Every one of those workloads lives on storage. The GPU gets the credit for the AI output. Storage keeps the GPU fed and working continuously.

Sullivan’s team selected the Solidigm D5-P5336 in the E3.S form factor. The XE7745 holds eight of these drives. At 30.72 TB per drive, that is over 245 TB of flash storage in a single 4U chassis. The D5-P5336 uses QLC NAND with enterprise firmware tuned for sustained write performance and data integrity, which matters here because the system is not handling occasional bursts. During peak submission windows around finals, it simultaneously ingests videos, writes transcripts, logs evaluation output, and updates the database.

As we documented in our ocean research story covering this same hardware configuration, the Solidigm drives in RAID 10 delivered sustained read and write performance without falling behind the processing pipeline. Storage was not the bottleneck. The architecture exposed the actual workload constraints, so the team could tune them where they mattered. That validated conclusion carries directly into the academic evaluation deployment.

OSU as a Blueprint for AI-Ready Higher Education

Oregon State’s approach to AI infrastructure is deliberate and worth examining as a model for other educational institutions. Rather than deploying AI tools opportunistically through cloud APIs, the university has made a series of architectural decisions that treat AI as a durable institutional capability rather than a vendor service. Hardware is standardized around platforms that span research and academic compute workloads. Storage is on-premises, high-density, and compliant by design. Faculty retain final authority over every evaluation the system produces.

Kalodimos is explicit about wanting the platform to travel beyond OSU. “Not every university is going to be as well-resourced as us,” he said. “I want to make sure this technology is available to all universities.” That foundation is what makes the broader argument for educational equity credible. An instructor at a smaller institution with a heavier teaching load and fewer resources arguably needs this tool the most.

“I need storage, and I need that storage to be fast. AI is a dead technology without storage. It’s a data-driven system. We had the algorithms back in the 1960s and 70s. We didn’t have any data to do it because we didn’t have any storage to actually hold that data.”

— Christopher Sullivan, Director of Research and Academic Computing, Oregon State University

Sullivan frames the hardware planning question the same way he frames every infrastructure decision at OSU. Models will change. The types of input students submit will evolve. The evaluation techniques faculty want to run will grow more sophisticated. “I’m going to get a bigger fork, a bigger knife, or a bigger spoon,” he said, “but it’s still a fork, knife, and spoon on the hardware side. I’m going to be changing the models and the inputs dramatically in the years to come, and it’s more important to me right now that the hardware keeps ahead of whatever those are going to be.”

Every processed transcript, every extracted slide, every draft evaluation, every approved grade, and every audit record must all live somewhere. In this system, there is 245 TB of Solidigm QLC flash on-premises within OSU’s infrastructure, doing the quiet work that makes the visible AI possible. The rack Sullivan is planning will not be the last one. The university is watching what this pilot produces, and other institutions will be watching what OSU does. That is what it means to lead in AI.

The post How Metrum AI and Oregon State University Are Building the New Standard for Academic Assessment appeared first on StorageReview.com.

Supermicro JumpStart Review: H14 with AMD Instinct MI350X

13 April 2026 at 19:50

Supermicro’s JumpStart program has established itself as one of the more useful tools in the pre-purchase evaluation toolkit for AI infrastructure. Rather than a scripted demo in a shared environment, JumpStart gives qualified users free, time-boxed, bare-metal access to real production servers via SSH, IPMI, and VNC, enabling them to run workloads on actual hardware. We covered the program in depth last November using an X14 system with an NVIDIA HGX B200, and came away with a clear picture of what a week of focused access can and cannot tell you. This time, Supermicro provided access to an H14 8U system with a very different accelerator story.

We tested the AS-8126GS-TNMR system, an 8U air-cooled platform built around dual AMD EPYC 9575F processors and eight AMD Instinct MI350X GPUs. The MI350X is AMD’s current flagship data center accelerator, built on the 4th Gen CDNA architecture at TSMC’s 3nm node and featuring 288GB of HBM3e per GPU. Across eight GPUs interconnected via AMD Infinity Fabric, the server offers 2.3 TB of total GPU memory in a single node, with an aggregate bandwidth of 1,024 GB/s. The full system uses six 5,250W Titanium-level power supplies in a 3+3 redundant configuration, and Supermicro has provisioned dedicated 400 Gbps networking per GPU for scale-out deployments.

GPU A+ Server AS -8126GS-TNMR Front.

AMD’s position in the data center GPU market has shifted meaningfully in the past two years, and the MI350X generation represents a more serious competitive challenge to NVIDIA than any prior Instinct product. ROCm 7, released in September 2025 and now at version 7.2, brought native MI350X support alongside dramatically improved inference performance, HIP API updates that close the CUDA compatibility gap, and broadened framework support, including PyTorch, JAX, TensorFlow, ONNX Runtime, vLLM, and SGLang.

The vLLM project added a dedicated AMD ROCm CI pipeline in late December 2025, making AMD hardware a first-class platform in that inference stack rather than a downstream port. The ecosystem’s adoption is also hard to ignore: AMD and Meta announced a multi-year, multi-generation 6-gigawatt GPU deployment agreement in February 2026, building on Meta’s existing production deployments of MI300 and MI350 series hardware. That level of commitment from one of the world’s largest AI infrastructure operators is no marketing footnote.

For organizations currently evaluating AI accelerator infrastructure, the lead time for NVIDIA hardware remains a concern. The question is whether AMD is a credible alternative rather than a fallback. Based on a week of testing with ROCm 7.2.0 and the current vLLM, the answer is meaningfully different from what it was 18 months ago.

GPU A+ Server AS -8126GS-TNMR side profile

Our testing covered a selection of popular models; the 2.3TB of HBM3e across a single node enabled single-server inference on large-parameter models, including Moonshot’s Kimi K2.5 and MiniMax M2.5.

AMD Instinct MI350X: Architecture and Generational Improvements

The MI350X represents AMD’s most architecturally ambitious generational leap in the Instinct product line to date. Understanding the engineering decisions behind it provides important context for interpreting the subsequent performance results.

CDNA 4 Architecture and Process Node Transition

The foundational shift from the MI300 series to the MI350 series centers on adopting TSMC’s N3P process node for the Accelerator Compute Chiplets (XCDs), moving from the 5nm fabrication used in the prior generation. The total transistor count reaches approximately 185 billion, a roughly 21% increase over the MI300 generation, achieved without a corresponding increase in power consumption.

The MI350X retains AMD’s proven multi-chiplet packaging strategy. At its core, the GPU package features eight Accelerator Compute Chiplets (XCDs) as the primary computational engines. Each XCD houses four shader engines, each with eight active CDNA 4 compute units, yielding 32 CUs per XCD and a total of 256 CUs for the full accelerator.

The I/O Die layer was also consolidated from four tiles to two in the CDNA 4 package design. This reorganization enabled AMD to double the Infinity Fabric bus width, improving bi-sectional bandwidth while lowering the bus frequency and operating voltage to reduce power consumption.

Redesigned Compute Units and Expanded Precision Support

With the CDNA 4 compute unit matrix math capabilities, there is a substantial boost: MI350 CUs deliver a 2x increase in throughput per CU for 16-bit (BF16, FP16) and 8-bit (FP8, INT8) operations compared to their MI300 counterparts.

Beyond raw throughput gains, CDNA 4 introduces hardware support for lower-precision data types absent from the MI300 series, specifically FP6 and FP4, alongside the existing FP8 support carried forward from the prior generation.

In addition to these standard formats, the MI350X adds native hardware support for the OCP microscaling variants: MXFP4, MXFP6, and MXFP8. Microscaling formats are designed to deliver the throughput advantages of lower-precision compute while maintaining output quality closer to higher-precision baselines than standard quantization typically allows. This is not an AMD-specific development. NVIDIA’s NVFP4 format operates on the same microscaling principles and has seen broad adoption across frontier model deployments, with the GPT-OSS family from OpenAI as one of the most prominent examples built around these formats. The MI350X’s native MXFP4 support allows it to serve these and similar quantized model families without falling back to software emulation or precision promotion.

The MI350X delivers 9.2 PFLOPs at MXFP4 and MXFP6, compared with 4.6 PFLOPs at OCP-FP8, with FP16 at 2.3 PFLOPs and a peak engine clock of 2,200 MHz. For inference-optimized deployments where microscaling quantization is viable, the compute headroom effectively doubles relative to FP8 workloads. A new vector ALU has also been added to the CDNA 4 compute unit, supporting 2-bit operations and capable of accumulating BF16 results into FP32, providing additional flexibility for low-precision vector workloads outside the primary matrix compute path.

Memory Subsystem: HBM3e, Infinity Cache, and Bandwidth Efficiency

The MI350 series features a substantially upgraded memory subsystem with eight HBM3e memory stacks, providing a total capacity of 288GB per GPU. Each 36GB stack, composed of 12-high 24Gbit devices, operates at the full HBM3e pin speed of 8Gbps per pin. The architecture retains AMD’s Infinity Cache, a memory-side cache positioned between the HBM and the Infinity Fabric/L2 caches. It comprises 128 channels, each backed by 2 MB of cache, for a total of 256 MB per GPU. AMD has widened the on-die network buses within the IODs and operates them at a reduced voltage, enabling approximately 1.3x higher memory bandwidth per watt compared to the MI300 series.

The increase in memory capacity from the MI300X’s 192GB to 288GB extends AMD’s lead in per-GPU memory headroom, with direct implications for large-model inference. Each MI350X GPU can independently host models with more than 500 billion parameters. Across an eight-GPU server, the aggregate 2.3TB of HBM3e eliminates the multi-node distribution requirements that complicate trillion-parameter deployments, as the Kimi K2.5 and MiniMax M2.5 results in this review demonstrate.

Flexible Partitioning and Deployment Architecture

The MI350 series supports flexible GPU partitioning per socket, with memory split into two separate clusters. This flexibility also applies to the XCDs, where the quad XCD cluster can be split into dual or single blocks, enabling the chip to support configurations such as 8 instances of 70B models in CPX+NPS2. For organizations running heterogeneous inference workloads across shared infrastructure, this partitioning capability reduces the need for dedicated hardware per model tier and improves utilization economics across mixed deployment environments.

The MI350 series also maintains drop-in compatibility with the UBB (Universal Base Board) infrastructure used in MI300 Series systems. Existing server chassis, power delivery, and cooling infrastructure carry forward without modification, reducing upgrade friction for organizations with active MI300 deployments.

MI355X: The Liquid-Cooled Sibling

The MI350 series ships in two variants built on identical underlying silicon and optimized for different thermal operating envelopes. The MI350X tested here is the air-cooled variant, while the MI355X is its liquid-cooled counterpart, designed for higher-density deployments where direct liquid-cooling infrastructure is available.

While both variants are built on the same fundamental hardware, the MI355X’s higher operational power envelope enables higher sustained clock frequencies, resulting in an approximate 20% performance advantage in real-world, end-to-end workloads compared to the MI350X. The MI355X carries a TBP ceiling of 1,400W versus the MI350X’s 1,000W, with clock speed topping out at 2.4 GHz compared to 2.2 GHz on the air-cooled variant.

In generational terms, the MI355X platform delivers up to 4x peak theoretical performance improvement over the MI300X, with real-world inference gains of approximately 4.2x in agentic and chatbot workloads and about 3x in content generation scenarios. For organizations evaluating MI350X deployments, the 20% performance differential between the two variants represents a clear ceiling. Facilities with DLC infrastructure should evaluate the MI355X to determine whether the thermal investment yields sufficient throughput uplift for their specific workload profile before committing to air-cooled configurations at scale.

Accessing the AMD Instinct MI350X via Supermicro JumpStart Program

Getting started with JumpStart requires registering on Supermicro’s portal, where qualified users can browse available systems and schedule a reservation window. Once approved, the portal provides SSH credentials, IPMI access, and a web-based remote console for the duration of the booking. The system arrives preinstalled with Ubuntu and ready to use. There is no provisioning delay and no support interaction required to get started. Our reservation ran from March 23 through March 27, 2026, giving us a full week on the platform, consistent with our prior JumpStart engagement on the HGX B200.

The screenshot below shows our terminal output from jumpstart for the H14 system, with the AMD-SMI tool displaying the eight AMD Instinct MI350X GPUs and their running software versions.

AMD Instinct MI350X Performance Testing Results

System Configuration

  • Chassis: Supermicro H14
  • CPU: Dual AMD EPYC 9575F
  • Memory: 3TB DDR5
  • GPU: eight AMD Instinct MI350X
  • Storage: 2x 3.8TB PCIe 4.0 M.2 NVMe SSD and 1 x 1.92TB NVMe M.2

Summary of Results

Model Precision Equal (256/256) Prefill-Heavy (8k/1k) Decode-Heavy (1k/8k)
GPT-OSS 20B NVFP4 62,247 123,714 32,468
GPT-OSS 120B NVFP4 33,538 84,018 20,602
Llama 3.1 8B Instruct BF16 51,467 77,658* 19,326
Mistral Small 3.1 24B FP8 40,742 56,093 14,557
Mistral Small 3.1 24B BF16 30,530 53,740 13,559
Qwen3 Coder 30B A3B BF16 34,980 51,550 11,782
Qwen3 Coder 30B A3B FP8 25,928 47,179 11,014
MiniMax M2.5 Block-Scaled FP8 14,391 23,689 6,068
Kimi K2.5 INT4 QAT + BF16 6,527 11,256 2,513
All values in tok/s, peak throughput at BS=256. *Llama 3.1 8B prefill-heavy peaked at BS=128 (77,658 tok/s); BS=256 was 76,893 tok/s.

Claude Code Serving – MiniMax M2.5

Beyond traditional raw LLM inference benchmarks, we wanted to evaluate how well this hardware performs in an agentic coding workflow, specifically serving multiple concurrent Claude Code sessions with a locally hosted model. This use case maps directly to development team productivity: how many engineers can simultaneously use an AI coding assistant served from a single node before the experience degrades?

To test this, we built a benchmark harness that generates a dataset of moderately difficult coding problems (tasks like implementing an LRU cache, building a CLI todo application, writing a markdown converter, and constructing a REST API) and runs each Claude Code session in its own Docker container against the local vLLM server. A transparent proxy sits between the sessions and the inference endpoint, capturing per-request metrics for each Claude code instance. The model used was MiniMax M2.5, served via vLLM on the eight MI350X GPUs. While not the top-ranked coding model on public leaderboards, M2.5 is a capable model that many users run locally, including many of our developer friends.

For a baseline reference point, we use Anthropic’s Claude Opus 4.6 average output throughput via OpenRouter.ai, one of the most popular routing services for production API access. That baseline comes in at approximately 37 tokens per second per API request.

We measured two key metrics: the average output tokens per second per Claude Code session (what each developer experiences) and the aggregate output tokens per second across all sessions (the total work the server is producing).

Looking at the results, a single concurrent session delivers 38.8 tok/s per user and 38 tok/s aggregate, slightly above the OpenRouter cloud baseline. At two sessions, the system edges up to 39.5 tok/s per user as vLLM’s batching begins to amortize overhead, with aggregate throughput climbing to 63 tok/s. Four concurrent sessions are held at 37.3 tok/s per user, matching the cloud baseline while serving four developers simultaneously, with aggregate throughput reaching 128 tok/s. From eight sessions onward, per-instance throughput begins to decline: 34.6 tok/s per user at eight sessions, 31.4 tok/s at sixteen with an aggregate of 190 tok/s, and settles around 23 tok/s per user at 32 and 64 sessions, while aggregate throughput climbs to 578 tok/s and 986 tok/s, respectively. This is the classic batching-throughput-versus-interactivity trade-off: the system can achieve significantly higher total throughput by batching more requests, but each user experiences slower response times. Even at 64 concurrent users, each developer still sees a usable interactive experience, though noticeably slower than the cloud baseline.

For organizations weighing the cost of dozens of simultaneous commercial API subscriptions against self-hosted infrastructure, the tradeoff is clear: a single MI350X node can serve a development team of 16 to 32 engineers, maintaining per-user response speeds within 60-85% of the cloud baseline while delivering aggregate output of 600 to 1,000 tok/s, with added benefits of data locality, no per-token API charges, and full control over model selection.

vLLM Online Serving – LLM Inference Performance

vLLM is one of the most popular high-throughput inference and serving engines for LLMs. The vLLM online serving benchmark evaluates the real-world serving performance of this inference engine under concurrent requests. It simulates production workloads by sending requests to a running vLLM server, with configurable parameters such as request rate, input/output lengths, and the number of concurrent clients. The benchmark measures key metrics, including throughput (tokens per second), time to first token, and time per output token (TPOT), helping users understand how vLLM performs under different load conditions.

We tested inference performance across a comprehensive suite of models spanning various architectures, parameter scales, and quantization strategies to evaluate throughput under different concurrency profiles.

GPT-OSS 120B and 20B

The GPT-OSS model family was tested in both 120B and 20B configurations on the Supermicro H14.

GPT-OSS 120B

The 120B model under an equal workload (256/256) delivers 313.42 tok/s at BS=1, reaches 11,261.72 tok/s at BS=64, and peaks at 33,538.23 tok/s at BS=256. Prefill-heavy (8k/1k) starts at 1,724.84 tok/s, climbs to 36,156.80 tok/s at BS=32 and 79,247.76 tok/s at BS=128, peaking at 84,018.79 tok/s at BS=256. Decode-heavy (1k/8k) grows from 288.90 tok/s at BS=1 to 20,602.52 tok/s at BS=256, with latency remaining well-controlled at lower concurrency levels.

GPT-OSS 20B

The 20B model delivers 485.17 tok/s at BS=1 under the equal workload, reaching 17,986.36 tok/s at BS=64 and peaking at 62,247.52 tok/s at BS=256. Prefill-heavy starts at 3,120.72 tok/s, climbs to 48,132.52 tok/s at BS=32 and 83,968.71 tok/s at BS=64, peaking at 123,714.50 tok/s at BS=256—the highest absolute prefill throughput recorded across both model sizes. Decode-heavy grows from 378.20 tok/s at BS=1 to 32,468.67 tok/s at BS=256, delivering roughly 1.6× the decode throughput of the 120B at peak concurrency while maintaining tighter latency characteristics throughout.

Qwen3 Coder 30B A3B Instruct and FP8 Instruct

The Qwen3-Coder-30B-A3B-Instruct on the Supermicro H14 was tested at both standard (BF16) and FP8 precisions.

Qwen3-Coder-30B-A3B-Instruct (BF16)

At BF16, the equal workload (256/256) delivers 240.53 tok/s at BS=1, reaching 13,312.70 tok/s at BS=64 and 21,333.79 tok/s at BS=128, with peak throughput of 34,980.97 tok/s at BS=256. Prefill-heavy (8k/1k) starts at 1,276.76 tok/s, climbs to 25,069.32 tok/s at BS=32 and 50,198.94 tok/s at BS=128, peaking at 51,550.66 tok/s at BS=256. Decode-heavy (1k/8k) grows steadily from approximately 188 tok/s at BS=1 to 11,782 tok/s at BS=256, maintaining the tightest latency profile of the three scenarios.

Qwen3-Coder-30B-A3B-Instruct (FP8)

The FP8 variant delivers 188.92 tok/s at BS=1 under the equal workload, reaching 10,866.27 tok/s at BS=64 and 17,617.60 tok/s at BS=128, peaking at 25,928.77 tok/s at BS=256—running slightly behind BF16 results across the full range. Prefill-heavy starts at 860.07 tok/s, climbs to 20,513.77 tok/s at BS=32 and 44,205.46 tok/s at BS=128, peaking at 47,179.15 tok/s at BS=256. Decode-heavy grows from 133.79 tok/s at BS=1 to 11,014.95 tok/s at BS=256, scaling consistently and remaining close to BF16 throughout.

Mistral Small 3.1 24B Instruct 2503

The Mistral-Small-3.1-24B-Instruct-2503 on the H14 was tested with both standard and FP8-dynamic precision, showing consistent scaling across all three workload profiles.

Mistral-Small-3.1-24B-Instruct-2503 (BF16)

With BF16 precision, the equal workload (256/256) delivers 236.15 tok/s at BS=1, reaching 15,494.56 tok/s at BS=64, 24,216.52 tok/s at BS=128, and peaking at 30,530.54 tok/s at BS=256. Prefill-heavy (8k/1k) starts at 1,429.41 tok/s, climbs to 29,631.68 tok/s at BS=32 and 54,871.74 tok/s at BS=128, peaking at 53,740.04 tok/s at BS=256. Decode-heavy (1k/8k) grows from 242.66 tok/s at BS=1 to 13,559.19 tok/s at BS=256, scaling steadily across the full range.

Mistral-Small-3.1-24B-Instruct-2503 (FP8-dynamic)

The FP8-dynamic variant delivers 184.25 tok/s at BS=1 under the equal workload, reaching 16,113.95 tok/s at BS=64 and 26,409.01 tok/s at BS=128, peaking at 40,742.04 tok/s at BS=256. Prefill-heavy starts at 1,210.06 tok/s, climbs to 28,773.52 tok/s at BS=32 and 57,765.02 tok/s at BS=128, peaking at 56,093.09 tok/s at BS=256, leading the standard precision result from BS=64 onward. Decode-heavy grows from 183.94 tok/s at BS=1 to 14,557.94 tok/s at BS=256, tracking closely through mid-range before pulling slightly ahead at BS=128 and BS=256.

Llama 3.1 8B Instruct

For the Llama-3.1-8B-Instruct, we saw the equal workload (256/256) delivers 373.26 tok/s at BS=1, reaching 19,363.33 tok/s at BS=64, 34,155.70 tok/s at BS=128, and peaking at 51,467.30 tok/s at BS=256. Prefill-heavy (8k/1k) starts at 1,959.04 tok/s, climbs to 37,227.63 tok/s at BS=32 and 60,062.40 tok/s at BS=64, peaking at 77,658.50 tok/s at BS=128 before tailing off slightly to 76,893.77 tok/s at BS=256. Decode-heavy (1k/8k) starts at 326.48 tok/s, reaching 17,877.52 tok/s at BS=128 and 19,326.35 tok/s at BS=256, maintaining lower per-token latency further into the concurrency range than any of the larger models tested.

MiniMax M2.5

The MiniMax-M2.5 on the H14 rounds out the model lineup, sitting between the Kimi K2.5 and the mid-sized models in terms of throughput profile, with characteristics that reflect its mixture-of-experts architecture. The equal workload (256/256) delivers 79.31 tok/s at BS=1, reaching 5,029.76 tok/s at BS=64, 7,801.10 tok/s at BS=128, and 14,391.98 tok/s at BS=256. Prefill-heavy (8k/1k) shows the strongest scaling of the three scenarios, starting at 424.41 tok/s and climbing to 10,376.75 tok/s at BS=32 and 20,658.57 tok/s at BS=128, peaking at 23,689.18 tok/s at BS=256. Decode-heavy (1k/8k) scales steadily to 4,257.68 tok/s at BS=128 and 6,068.70 tok/s at BS=256, offering the most consistent latency growth across the full concurrency range.

Kimi K2.5

The Kimi K2.5 1-trillion-parameter model on the H14 is the largest and smartest model tested in this review, and its throughput reflects that weight.

The equal workload (256/256) delivers 72.06 tok/s at BS=1, reaching 2,693.07 tok/s at BS=64, 4,244.27 tok/s at BS=128, and peaking at 6,527.62 tok/s at BS=256. Prefill-heavy (8k/1k) scales more aggressively, starting at 185.29 tok/s and reaching 3,798.85 tok/s at BS=32 and 9,153.12 tok/s at BS=128, with peak throughput of 11,256.69 tok/s at BS=256. The step increase from BS=128 to BS=256 carries a significant latency cost, indicating the system is approaching its memory and compute limits at full batch depth for this model size. Decode-heavy (1k/8k) grows from 29.88 tok/s at BS=1 to 2,513.85 tok/s at BS=256, delivering the tightest scaling curve of the three scenarios while demonstrating consistent throughput gains across the full range.

Conclusion

The AMD Instinct MI350X delivers very competitive inference performance across the workload profiles tested here, and the Supermicro AS-8126GS-TNMR provides a well-engineered platform to take advantage of it. With 288GB of HBM3e per accelerator and eight GPUs interconnected over Infinity Fabric, the 2.3TB of aggregate GPU memory available in a single node is sufficient to serve trillion-parameter models like Kimi K2.5 and MiniMax M2.5 without requiring multi-node distribution or model-partitioning workarounds. This capability materially simplifies deployment architecture for large-scale inference.

Smaller models also delivered strong results. Llama 3.1 8B exceeded 77,000 tok/s under prefill-heavy workloads, and mid-range architectures such as Mistral Small 3.1 24B and Qwen3 Coder 30B sustained high throughput with well-controlled latency across the concurrency range. Across the board, the results indicate a hardware platform that scales predictably under load rather than falling off a cliff at higher batch depths.

GPU A+ Server AS -8126GS-TNMR rear

ROCm 7.2 brings significant improvements to the AMD inference software stack, particularly when paired with vLLM 0.18. This pairing delivers a noticeably more stable and higher-performing serving experience than prior ROCm generations, with broader framework support and fewer of the rough edges that characterized earlier Instinct deployments. The ecosystem momentum around AMD hardware is also worth noting: upstream vLLM now maintains a dedicated AMD ROCm CI pipeline, and Meta’s multi-generation deployment commitment at the 6-gigawatt scale reinforces that production validation extends well beyond controlled benchmarking environments.

The Claude Code serving evaluation adds a practical lens to the raw throughput numbers. A single MI350X node sustained near-cloud-baseline response speeds for up to 16 concurrent coding sessions and remained interactive with up to 64 simultaneous users while producing nearly 1,000 tok/s of aggregate output. For organizations weighing the cost of commercial API subscriptions against self-hosted infrastructure, the economics become straightforward at that density, with additional advantages in data locality, elimination of per-token costs, and unrestricted model selection.

Supermicro’s JumpStart program continues to earn its place in the infrastructure evaluation process. Bare-metal access to production hardware, with no provisioning overhead, allowed us to run real workloads under real-world conditions throughout the full test window. For teams conducting accelerator procurement evaluations, this level of hands-on access remains far more informative than spec sheet comparisons or curated vendor demonstrations.

Supermicro JumpStart Program

Product Page – GPU A+ Server AS -8126GS-TNMR

The post Supermicro JumpStart Review: H14 with AMD Instinct MI350X appeared first on StorageReview.com.

Emulex SecureHBA Enables Autonomous Fibre Channel SAN Encryption with Everpure FlashArray

19 March 2026 at 13:00

For the first time, end-to-end, hardware-based in-flight encryption is available across the entire Fibre Channel data path—from server to storage array. Everpure (formerly Pure Storage) is now shipping Fibre Channel-capable FlashArray systems with Emulex SecureHBA technology integrated directly into the array, completing the encrypted transport path the industry has been working toward. Future FlashArray models will ship with SecureHBA as the standard Fibre Channel adapter option. In our evaluation, we assessed the FlashArray//XL130 R5, the first primary storage array to natively bring this capability to the platform.

FlashArray//XL130 R5 Everpure - emulex secureHBA

When SecureHBA-equipped servers connect to the Everpure array, encrypted sessions are established automatically in hardware. There are no additional agents to install and no external encryption infrastructure required. The result is an end-to-end encrypted Fibre Channel connection from server to storage, implemented transparently and without requiring additional software or external encryption infrastructure.

While encryption at rest is standard practice across enterprise storage, encrypting data in-flight, especially within Fibre Channel SAN environments, has been difficult to implement without adding complexity and impacting application performance. Historically, organizations have relied on software-based methods with external key management systems, which introduce additional management overhead and potential risks. Delivering seamless, hardware-level encryption across the storage network has historically been difficult.

Emulex set out to solve that challenge with its SecureHBA. Rather than depending on host software or external systems, SecureHBA performs encryption in hardware, ensuring that application performance and workflows remain unaffected. Since that early coverage, the technology has moved decisively into production. In February, the Fibre Channel Industry Association (FCIA) announced completion of the INCITS Fibre Channel, Security Protocols, Third Edition (FC-SP-3) specification. Emulex SecureHBA is fully compliant with the new standard. SecureHBAs are now widely shipped by major server OEMs, including Dell, HPE, and Lenovo, and more than 120,000 adapters have already been deployed in enterprise environments. In numerous instances, organizations have installed encryption-capable Fibre Channel endpoints during routine server refresh cycles, but haven’t fully realized their benefits because the storage arrays lacked native support.

This technical analysis reviews how SecureHBA functions with the tested FlashArray//XL130 R5, examines how Emulex SAN Manager 3.0 provides operational visibility and reporting, and confirms the performance of hardware-based encryption in a real-world setting. A key distinction throughout is where encryption occurs in the data path: unlike application-layer encryption, which renders data incompressible before it reaches the array, SecureHBA encrypts at the Fibre Channel transport layer, preserving array-level compression, deduplication, and ransomware detection capabilities.

Why In-Flight SAN Encryption Matters

Fibre Channel networks have traditionally operated inside a presumed trust boundary. Storage fabrics were physically segmented, tightly managed, and considered insulated from broader network threats. That assumption is increasingly difficult to defend. Ransomware activity remains persistent and adaptive. The Symantec and Carbon Black Threat Hunter Team’s Ransomware 2026 report notes that threat actor groups continue to evolve rapidly, with disrupted operations quickly replaced by new affiliates, resulting in sustained year-over-year attack volume. The 2025 Verizon Data Breach Investigations Report corroborates this trend, finding ransomware present in 44% of all breaches analyzed, a 37% increase from the prior year.

These campaigns are infrastructure-aware. Once an attacker gains privileged access, internal lateral visibility becomes possible. While Fibre Channel traffic is not routed like traditional IP traffic, it still operates at the transport layer, moving highly sensitive data between servers and storage. Treating the SAN as inherently secure assumes that perimeter defenses remain intact and that internal compromise never occurs, assumptions that breach patterns increasingly challenge.

The more realistic risk scenario is not an external attacker targeting the Fibre Channel fabric directly, but rather a compromised host or a privileged insider already operating within the environment. When administrative control of a server or virtualization host is obtained, storage traffic moving between the host and the array becomes observable within the infrastructure boundary. Encrypting the transport path ensures that even trusted internal networks do not expose sensitive data during lateral movement or post-compromise activity.

Regulatory expectations are evolving alongside the threat landscape. Frameworks such as CNSA 2.0, NIS2, and DORA place greater emphasis on cryptographic resilience and on protecting data in motion, not just data at rest. Security models are increasingly built around zero-trust principles, where internal network segments are no longer implicitly trusted. As a result, organizations are being pushed to evaluate encryption across all layers of infrastructure, including storage transport paths that were previously assumed to be secure.

There is also an operational dimension. Many modern storage platforms rely on identifying anomalous behavior (such as sudden changes in compression ratios or write patterns) to detect ransomware early. If encryption is applied at the application layer before data reaches the array, those detection mechanisms can be impaired. Encrypting traffic at the transport layer instead preserves the array’s visibility into data characteristics while still protecting it in motion. That architectural distinction becomes increasingly relevant as organizations refine recovery strategies and seek to shorten dwell time during attacks.

Encryption placement within the data path also affects storage efficiency. When data is encrypted at the application or host level before it reaches the storage system, it typically becomes incompressible because the encryption process removes recognizable data patterns. That can significantly reduce the effectiveness of array-level compression and data reduction services. By contrast, transport-layer encryption within the Fibre Channel stack occurs after data leaves the host but before it traverses the SAN, and is decrypted upon entering the storage array. This approach allows the array to process the original data stream normally, preserving compression, deduplication, and other data services while still protecting the data in flight. The distinction between application-layer and transport-layer encryption is particularly important for organizations that rely on modern storage arrays for efficiency and ransomware detection.

Infrastructure refresh cycles amplify the significance of this moment. Server and storage platforms frequently remain in production for five years or more, meaning the architectural decisions made today shape the security posture for much of the next decade. Encryption capabilities embedded at the hardware level are rarely retrofitted; they are typically selected as part of the foundational design. For organizations evaluating new server deployments or modernizing storage arrays, transport-layer encryption within the SAN should no longer be treated as optional. Hardware-based Fibre Channel PQC-secure encryption represents a durable security control that aligns with evolving compliance expectations, zero-trust models, and long-term risk management.

In-flight SAN encryption is not a reaction to a specific Fibre Channel weakness. It reflects a broader shift in proactive enterprise security thinking: no internal network layer is implicitly trusted, and resilience must be engineered into infrastructure from the outset at every possible level.

SecureHBA Architecture Overview

The Emulex SecureHBA ecosystem is built on the INCITS Fibre Channel Security Protocols, Third Edition (FC-SP-3) standard, providing backward and forward compatibility with existing Fibre Channel networks and adapters. SecureHBAs can be introduced into an existing fabric without requiring switch hardware or software upgrades. When both endpoints support hardware-based encryption, encryption capabilities are negotiated automatically during the standard Fibre Channel login process, and an encrypted session is established without manual configuration.

Because PQC-safe key negotiation is managed entirely by the SecureHBAs, no additional key management infrastructure or operating system configuration is required. PQC-safe, in this context, refers to cryptographic key exchange mechanisms designed to remain resistant to future quantum-computing attacks on traditional public-key algorithms. Session keys are generated and renewed autonomously, allowing encryption to operate transparently within the transport layer.

Administrators can change the default behavior, but it’s not necessary for full compatibility. An admin may want to change settings to allow only encrypted connections; this will prevent servers that do not support them from connecting. The admin could also disable the feature, allowing all servers to connect without encryption if so desired.

From an architectural standpoint, SecureHBA embeds encryption directly into the Fibre Channel transport path rather than layering it on top of applications or host software. This preserves existing fabric design while enabling in-flight encryption to be deployed without introducing operational complexity.

Modern server platforms increasingly include hardware-based trust mechanisms that verify the authenticity and integrity of system components before they engage in production workloads. Technologies such as the silicon root of trust, secure boot chains, and device attestation protocols, including the DMTF Security Protocol and Data Model (SPDM), enable systems to cryptographically verify firmware, adapters, and other devices at startup. These mechanisms ensure that the hardware involved in a workload is genuine and runs with verified firmware before any data transfer begins.

SecureHBA extends the hardware trust model into the storage transport layer itself. Once trusted devices are established at the server and storage endpoints, the encrypted Fibre Channel session protects data in transit between them. In this architecture, identity verification and data protection work together: trusted infrastructure components communicate across an encrypted SAN path without needing additional software layers or external security infrastructure. The result is a security posture that aligns with modern zero-trust principles while maintaining the operational simplicity typical of Fibre Channel environments.

Performance Validation: Oracle TPROC-H Workload

Broadcom and Everpure representatives met with us for a virtual lab session to demonstrate the performance characteristics of unencrypted and encrypted connections using the Emulex SecureHBA and Everpure FlashArray//XL130 R5. The setup consisted of a server with a SecureHBA installed, running Oracle 26ai Database software, with eight 200-gigabyte volumes provisioned on the FlashArray//XL130 R5 for a HammerDB TPROC-H data warehousing test. Designed for maximum throughput, this test allowed us to push the array and SecureHBA to their absolute limits and review their performance with hardware encryption enabled and disabled.

Emulex Secure HBA Oracle Layout

Because software-based encryption solutions impose their penalty directly on the host CPU, the host CPU is the right place to look when evaluating whether hardware offload delivers on its promise. Since all encryption processes occur directly on the SecureHBA’s silicon, CPU utilization on the server remained virtually identical across the unencrypted and encrypted test runs. By offloading encryption and decryption to the HBA, data transmission can still occur at maximum throughput without any CPU bottleneck.

Zooming in a bit further, we saw a nearly indistinguishable difference in maximum CPU usage under the TPROC-H workload: 30.46% with an encrypted connection and 31.68% with an insecure connection. Averages for the CPU utilization data were also extremely tight here, with encryption on at 24.35% compared to encryption off at 24.28%. These results are so close that they are well within run-to-run variance and demonstrate that the SecureHBA’s onboard encryption processes do not affect host CPU usage.

When using solutions that depend on key manager servers or services, both the active and passive sides of the connection must be configured and tested to ensure proper authentication and encryption, as well as successful failover. Since Emulex performs this automatically in hardware, it isn’t a concern for administrators. For alternative solutions that depend on external key manager services, both active and standby paths must be individually configured, authenticated, and validated to ensure proper failover behavior.

In large SAN environments with thousands of connections, the time and operational complexity required to verify those paths can increase significantly. This process could involve thousands of connections per port and might significantly extend the overall duration. By handling encryption negotiation directly within the SecureHBA hardware, these additional authentication steps are eliminated, minimizing the risk of misconfiguration or delayed failover.  Additionally, solutions that rely on external key managers risk misconfiguration or out-of-sync certificates across all connections, including the passive path. This may not be immediately obvious, since most connectivity issues are only noticed when the active path encounters a problem.

With Everpure’s FlashArray//XL130 R5 and SecureHBAs, customers can expect a rapid, secure cutover during critical failures, thanks to its resilient dual-controller architecture and autonomous encrypted connection re-establishment. When storage fabrics fail, Everpure and Emulex allow administrators to focus on the underlying issue, with SAN failover handled automatically in the background.

Operational Validation: How Encryption Is Observed in Practice

During our lab session, Emulex gave us a quick demonstration of their zero-touch provisioning feature using the Emulex SecureHBA. This lab’s hardware topology consisted of an initiator host server with two HBAs: one 64G Fibre Channel SecureHBA and one 64G Fibre Channel HBA without hardware encryption features. Both initiators were connected to an Everpure FlashArray//XL130 R5 with two SecureHBAs configured as targets.

Emulex SecureHBA zero touch layout

Using the command-line interface on the XL130 R5, Everpure showcased seamless, automatically encrypted connections and showed exactly how the encryption status of each port is displayed, as pictured below.

The Everpure CLI can also reveal more detailed encryption information from SecureHBAs, including whether the connection uses post-quantum cryptography and the HBA’s re-keying deadline.

By default, encrypted connections were immediately established between the SecureHBAs on the initiator host server and the Everpure FlashArray//XL130 R5. An unencrypted connection was established between the standard Fibre Channel HBA and the same array port, demonstrating ease of use while maintaining broad compatibility.

Emulex SAN Manager 3.0: Fabric-Level Visibility and Compliance

Highlights and New Features

  • Quick access to the encryption status of every connection in a SAN
  • Built-in reporting tools for auditing and evaluating SAN-wide security compliance
  • Scripting-friendly platform with a Python-based CLI and automatic event notification via SMTP

For customers using Emulex Fibre Channel HBAs (including SecureHBAs), Broadcom offers Emulex SAN Manager 3.0 software at no additional cost. Since its inception in 2020, the software has been designed to give storage administrators deeper insight into their SAN and to provide tools for monitoring, maintenance, and reporting across multiple Fibre Channel environments. Upon login, clear doughnut-style charts provide an overview of SAN-wide performance and health data using widgets that can be customized and rearranged to suit the user’s preferences. Each one of these widgets can be clicked on to reveal the data behind each chart.

Released alongside the Emulex SecureHBA are new features in SAN Manager 3.0. Users can now use the “Encryption Policy” widget on the dashboard to see exactly how many Emulex Fibre Channel nodes are capable of encryption, have encryption required, or have encryption capabilities turned off. Clicking on the “Encryption Policy” widget or the “Encryption Info” header in SAN Manager 3.0 provides a detailed list of all connected ports, which can be filtered and sorted by attributes such as WWPN, Host Name, Policy, and Key Interval. This makes it easy to find and manage ports quickly, and it can provide even more information when a specific port is selected.

In addition to viewing configured encryption policies in the dashboard, Emulex SAN Manager 3.0 now includes built-in reporting tools for auditing and compliance verification. CSV-formatted security reports can now be generated and exported (via download or email) directly in the application, saving busy IT managers the hassle of gathering information from potentially thousands of Fibre Channel nodes and manually verifying compliance.

The Emulex SAN Manager also provides users with visibility into their SAN health, with tools to moderate congestion, identify failing transceivers, and view link speeds. In addition to being an excellent GUI-based tool, Broadcom has created opportunities for automation-minded IT folks with a Python-based command-line interface for scripting. It plans to add built-in webhooks that analysts can use to export and review performance and reliability data from the SAN Manager to other applications and tools for review. Even email alerts (via SMTP) that warn administrators of critical events can be configured using a powerful task scheduler. Driven mainly by customer requests, the development team is working hard to make the application an extensible, all-in-one platform. Emulex SAN Manager 3.0 will be a welcome addition to any storage administrator’s arsenal when an Emulex SAN’s encryption compliance, health, multipathing, or performance needs to be reviewed.

Final Thoughts

The Everpure FlashArray with Emulex SecureHBA represents the first production deployment of transport-layer, PQC-safe encryption across the Fibre Channel data path.

The introduction of hardware-based, transport-layer encryption into the Fibre Channel stack marks a meaningful shift in how enterprises think about storage security. For years, Fibre Channel fabrics operated inside an implicit trust boundary. Emulex SecureHBA challenges that assumption by embedding encryption directly into the transport layer, without altering fabric design or introducing operational overhead. Our evaluation of the Everpure FlashArray with Emulex SecureHBAs demonstrates that end-to-end, in-flight encryption can be deployed without measurable performance degradation or added complexity. Encryption was negotiated automatically, scaled across more than one thousand sessions, and preserved the array’s data services and host performance characteristics.

Emulex has already established a significant footprint, with more than 120,000 SecureHBAs shipped across major server OEM platforms. The addition of a storage array endpoint closes the loop, enabling true end-to-end encryption across the SAN. While Everpure is the first to integrate SecureHBA at the array level, the architectural approach is standards-based and positioned to extend across future Fibre Channel platforms as hardware refresh cycles continue.

Because this solution is standards-based, it ensures interoperability. Although Emulex and Everpure are the first to introduce it, it still leaves room for interoperability between other brands’ HBAs and storage arrays. This solution is fabric-agnostic and can seamlessly extend to fabric extension solutions (FC-IP). Since everything is managed in hardware, the solution does not require specific software versions or additional OS version-specific feature packages.

Infrastructure decisions made today often define enterprise risk posture for the next five to seven years. As organizations modernize servers and storage arrays, embedding PQC-safe encryption at the transport layer provides a durable security control that aligns with evolving compliance expectations and zero-trust principles. SecureHBA positions in-flight SAN encryption as a foundational element of modern Fibre Channel design, aligning security, performance, and operational simplicity within the transport layer itself.

Find out more about Emulex LPe38102 SecureHBAs, Emulex LPe37102 SecureHBAs, and Emulex SAN Manager.

Learn more about the Everpure FlashArray//XL series.

This report is sponsored by Broadcom. All views and opinions expressed in this report are based on our unbiased view of the product(s) under consideration.

The post Emulex SecureHBA Enables Autonomous Fibre Channel SAN Encryption with Everpure FlashArray appeared first on StorageReview.com.

❌
❌