Normal view

There are new articles available, click to refresh the page.
Today — 15 September 2026StorageReview

Fujitsu MONAKA Server Brings 2nm 144-Core CPUs to Air-Cooled AI Inference, On Sale in November

14 September 2026 at 18:03
Fujitsu MONAKA Server Fujitsu MONAKA Server

Fujitsu is bringing its 2nm FUJITSU-MONAKA processor to AI infrastructure with a new server family designed to run AI inference in air-cooled data centers without requiring specialized liquid cooling. The MONAKA Server is designed, developed, and manufactured in Japan, with component and manufacturing traceability for sovereign AI deployments.

FUJITSU-MONAKA CPU package render with the Fujitsu logo, the 2nm 3D-stacked processor at the heart of the Fujitsu MONAKA Server

The first MONAKA Servers will come in 1U and 2U configurations for AI inference and AI agents, with a separate 2U multi-node system planned for academic and HPC environments. Fujitsu is also making the MONAKA processor available separately for cloud and data center operators and other server vendors.

MONAKA runs at up to 3.8GHz and supports memory transfer speeds up to 8800MT/s, with dedicated matrix instructions and SVE2 vector processing accelerating AI inference directly on the CPU. Fujitsu rates MONAKA at twice the AI inference throughput of other CPUs, although it hasn’t identified the processors used for that comparison or provided benchmark results.

MONAKA Brings AI Inference to Air-Cooled Servers

The 1U MONAKA Server supports air cooling at ambient temperatures up to 40 degrees Celsius, while its liquid-cooled configuration supports water temperatures up to 45 degrees Celsius.

Fujitsu rates its server cooling technology for up to an 80% reduction in cooling power consumption, although specific system configurations and test conditions weren’t provided. The air-cooled configuration is particularly interesting for data centers that want to add AI inference capacity without changing their existing cooling infrastructure.

Fujitsu MONAKA Server 2U chassis with the lid off, showing two FUJITSU-MONAKA CPUs under their heatsinks, the DIMM banks around them, and the front drive bays

The processor itself uses a 3D-stacked design that combines CPU cores manufactured on a 2nm process with cache and I/O produced at 5nm. Fujitsu pairs the design with its ultra-low-voltage operation technology, which is where it says the power efficiency comes from.

Confidential computing is implemented at the hardware level through Arm CCA, encrypting applications and data while they are running in memory, including workloads operating in multi-tenant cloud environments.

MONAKA is also being developed for integration with accelerated platforms through Fujitsu’s work with NVIDIA to connect MONAKA CPUs and NVIDIA GPUs through NVLink Fusion.

MONAKA Server Keeps Development and Manufacturing in Japan

Sovereign infrastructure is the other major focus for MONAKA Server, with Fujitsu keeping the design, development, and manufacturing of the systems in Japan. Production will take place at the company’s Kasashima Plant, where component origins and manufacturing histories can be tracked through the production process.

That traceability takes the sovereign AI pitch down to the physical hardware, past the usual question of where data is stored or processed. Fujitsu is pairing the domestically produced hardware with its own AI software and management technologies for organizations that need greater control over their infrastructure, data, and supply chain.

The initial server lineup includes a 2U system with two MONAKA processors and 1U configurations with either one or two CPUs. Both air- and water-cooled deployments are supported, with the systems targeting AI inference and AI agent workloads.

Specification 2U Rackmount Model 1U Rackmount Model
Features All-in-one CPU/GPU/NW model optimized for existing DC/Edge environments Scalable model with flexible configuration and expansion according to environmental changes
Applications Digital Twin, Physical AI, Agentic AI, AI Inference Optimized for facility environments; flexible expansion from small to large scale
CPU Type / Quantity FUJITSU-MONAKA
2 CPUs
FUJITSU-MONAKA
1 to 2 CPUs
Base Frequency / Cores 2.1GHz × 144 cores 2.1GHz × 144 cores (air-cooled)
2.9GHz × 144 cores (liquid-cooled)
Memory Type / Slots RDIMM
24
RDIMM
12 (1-CPU configuration) / 24 (2-CPU configuration)
Storage Type / Slots E3.S SSD × 4
M.2 SSD × 2
E3.S SSD × 8
M.2 SSD × 2
Expansion Slots PCIe Gen6 (GPU support available) PCIe Gen6
Chassis Size 2U height 1U height
Cooling Method Air-cooled Air-cooled / Liquid-cooled

 

The 1U MONAKA Server also uses CDI/CXL technology to pool memory and accelerators, allowing those resources to be allocated across systems.

For academic and HPC environments, Fujitsu is preparing a 2U multi-node MONAKA Server with four nodes and two CPUs per node. The configuration is intended for computational workloads including surrogate models and fluid analysis.

Specification 2U Multi-node Model (4 nodes per chassis)
Features Multi-node model maximizing processing power within limited power and space
Applications AI data centers, large-scale simulations
CPU Type / Quantity FUJITSU-MONAKA
2 CPUs per node / 8 CPUs per chassis
Base Frequency / Cores 2.9GHz × 144 cores
Memory Type / Slots RDIMM
24 per node / 96 per chassis
Storage Type / Slots E1.S SSD × 2 per node
M.2 SSD × 2 per node
Expansion Slots PCIe Gen6
Chassis Size 2U height
Cooling Method Liquid-cooled

Fujitsu Connects MONAKA With Its AI Software

Fujitsu is extending the sovereign infrastructure approach into software through the Fujitsu Kozuchi AI platform, Takane enterprise generative AI, and industry-specific AI models. The combination brings together the processor, server hardware, AI platform, and generative AI software while supporting domestic data management, governance, traceability, security, and operational autonomy.

Future MONAKA hardware will include higher-density systems for AI data centers and rack-scale servers with autonomous operation features. Fujitsu is also developing robotic maintenance technology for future rack-scale systems, extending automation into physical server operations and maintenance.

Fujitsu’s math on the twice-the-throughput claim is that customers need half the servers and half the power for the same inference load, which is the figure the air-cooled pitch rests on. With no named comparison CPU or published benchmarks behind it, that’s the number to test when systems ship.

FUJITSU-MONAKA Availability

Standalone FUJITSU-MONAKA processors are scheduled to go on sale in November 2026, with availability planned for cloud and data center operators and server vendors in Japan, the US, APAC, and other markets during the fourth quarter of Fujitsu’s fiscal 2026.

Fujitsu also plans to begin sales of its 1U and 2U MONAKA Servers in Japan and Europe in November 2026 to data center operators, enterprises, academic and HPC buyers, and the defense sector. Select customers in the financial, telecommunications, and manufacturing sectors are expected to receive systems during the second half of fiscal 2026, with broader shipments in Japan and Europe scheduled to begin sequentially in April 2027.

FUJITSU-MONAKA

The post Fujitsu MONAKA Server Brings 2nm 144-Core CPUs to Air-Cooled AI Inference, On Sale in November appeared first on StorageReview.com.

NASA-IBM Lunar Foundation Model Goes Open Source With a 2M-Tile Dataset and 22% Lower Ice-Mapping Error

14 September 2026 at 16:43

IBM and NASA have released the NASA-IBM Lunar Foundation Model as open source, one of the first publicly available foundation models built for scientific study of the Moon. The weights, a technical report, and the machine-learning-ready dataset it was trained on are up on Hugging Face under the Prithvi family, which already covers Earth observation, weather, and heliophysics. The pitch is that decades of multi-instrument lunar data have outgrown hand-sifted maps and narrow, task-specific models, and a single pretrained backbone can be adapted to crater mapping, volcanic-feature detection, and ice prospecting without starting over each time. In IBM and NASA’s own benchmarks, the model cut root-mean-square error in flagging likely ice deposits by up to 22% against a SwinV2-B baseline.

NASA-IBM Lunar Foundation Model crater detection output: two grayscale lunar surface tiles with blue bounding boxes drawn around detected craters

A TerraMind Backbone on 30 Layers From Nine Instruments

The model is built on a version of TerraMind, the Earth-observation model IBM developed with the European Space Agency, chosen for how it handles mixed data types and resolutions and learns cross-modal correlations that can fill in missing or noisy values. Fine-tuning for each lunar task uses low-rank adapters that keep 90% of the base weights frozen, which is what keeps the adaptation cost low. The training corpus is the part NASA says didn’t exist before: a unified, spatially aligned dataset of more than 30 layers from nine instruments across four missions, on the order of 2 million image tiles, with more than 1 million 1-meter camera images and close to 964,000 multispectral images at 100-meter resolution. Sources include imagery from NASA’s Lunar Reconnaissance Orbiter, gravity-field maps from the GRAIL mission at 20 kilometers per pixel, Lunar Prospector data, and complementary observations from JAXA’s SELENE/Kaguya. The benchmark collection, called SOMBench, ships alongside the model.

NASA-IBM Lunar Foundation Model ice prospectivity comparison: north and south polar maps beside label, ConvNeXt, and lunar foundation model prediction tiles on a blue-to-yellow prospectivity scale with a 10 km bar

Ice, Volcanic Patches, and Craters at Two Scales

The 22% figure comes from the ice task, where the model combines multimodal, multi-resolution observations to predict where ice may sit below the surface of permanently shadowed regions, the places a future Moon base would mine for water, oxygen, and rocket propellant. For volcanic history, it maps Irregular Mare Patches with 3% better coverage than SwinV2-B while training on imperfect labels, which IBM frames as comparable accuracy at lower fine-tuning cost. Crater detection splits by scale: at meter-scale resolution the model matches SwinV2-B, and at roughly 100-meter context scale it outperforms it by nearly 19% with half the training data. Michael Barker, the NASA lunar topography expert who co-led the project, said the payoff is in features whose ages are still contested: “The ages of these features remain a matter of great debate, so the more we can understand their distribution and properties, the better chance we have of resolving this mystery.”

“NASA has spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job. We also have to make data easier for scientists to explore and use,” said Kevin Murphy, chief science data officer and acting chief data and AI officer at NASA Headquarters. Juan Bernabe-Moreno, director of IBM Research Europe, UK and Ireland, said the model “gives scientists a foundation to explore the Moon at scale, connecting observations across instruments, revealing patterns that are difficult to see in isolation, and providing an open platform the global research community can build on.”

Everything is available now under an open-source license, and IBM’s research blog walks through the three initial NASA use cases in more depth. It’s the same open-weights playbook IBM has run on the enterprise side with Granite 4.0, applied here to a science domain where the scarce resource is labeled data and the compute to fine-tune against it, and the immediate consumers are landing-site hazard analysis and resource prospecting for a sustained lunar presence.

NASA-IBM Lunar Foundation Model

The post NASA-IBM Lunar Foundation Model Goes Open Source With a 2M-Tile Dataset and 22% Lower Ice-Mapping Error appeared first on StorageReview.com.

Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts

14 September 2026 at 16:23

Lightbits Labs, the company that invented NVMe over TCP, is moving into inference software with Inferra, a KV cache orchestration engine that makes its public debut tomorrow, September 15, at the AI Infra Summit in Santa Clara. The software virtualizes GPU memory across DRAM and NVMe storage tiers and turns the KV cache into a persistent data layer, so attention states for long-context and multi-session workloads don’t have to fit in HBM or be recomputed when they fall out. Lightbits claims up to 16x more concurrent inference sessions on existing GPUs, a greater than 100x latency improvement from prefetching attention states in place of recomputing them, and context windows up to 10 million tokens, and says the engine has already run in customer beta programs, including one with OVH.

Lightbits Inferra architecture diagram with the KV cache engine sitting between GPU HBM, DRAM, and storage tiers, showing prefetch, compression, quantization, scheduling, encrypted transfer, and tenant isolation functions

Prefetch Instead of Recompute

The bottleneck Inferra targets is the one we walked through in our KV cache offload to flash piece: as context lengths grow and sessions multiply, the KV cache outgrows GPU memory, and the serving stack either evicts it and recomputes the prefill later or caps how many sessions a GPU can hold. Both paths burn GPU cycles on work that has already been done. Inferra keeps the cache alive across memory and storage tiers, and a predictive prefetcher pulls attention states back toward the GPU ahead of when the model needs them. Lightbits says that approach drives down both Time-to-First-Token (TTFT) and Time Per Output Token (TPOT), and the 100x figure it quotes is the latency improvement against recomputing those states from scratch.

Lightbits Inferra block diagram showing its intelligent prefetcher, tiering manager, log-structured KV store, and security and isolation engines between LLM serving frameworks (vLLM, TensorRT, SGLang) and a standard NVMe SSD pool

Lightbits’ block diagram lays out the four pieces: an intelligent prefetcher, a tiering manager, a log-structured KV store, and security and isolation engines, sitting between the serving framework above and a standard NVMe SSD pool below. The company lists vLLM, TensorRT, and SGLang as supported serving frameworks and says the engine is GPU-and SSD-agnostic. The log-structured store is what makes a flash tier practical for a cache that changes constantly, since it appends updates sequentially and keeps small random writes off the drives, and we looked at how a 3 DWPD drive holds up in that role in our Solidigm D7-PS1030 review. On the multi-tenant side, Lightbits describes secure tenant isolation and intelligent cache management across shared inference infrastructure, with encrypted transfer between tiers, which is what lets a neocloud hand one KV pool to many customers with consistent SLAs.

Avigdor Willenz, co-founder and chairman of Lightbits Labs, framed the launch as a break from storage built for training pipelines. Inference “is a completely different paradigm,” he said, and “retrofitting legacy training and storage systems to solve the KV cache bottleneck simply doesn’t work. It was built for a different era. We engineered Inferra from the ground up specifically to solve the GPU efficiency problem, and it has already proven successful in customer beta programs.”

OVH, Solidigm, and the Booth 219 Demos

OVH is the named beta customer. “With Inferra’s intelligent KV cache tiering, we were able to demonstrate substantial GPU utilization gains, paving the way to providing our customers a more scalable and cost-effective foundation for their AI agent and RAG workloads,” said Yaniv Fdida, chief product and technology officer at OVH. Solidigm ran the engine in its AI Central Lab on D7-PS1010 drives, and Avi Shetty, the company’s VP of ecosystem, solutions, and market enablement, said the results “showcased how network-attached storage with intelligent software layer virtualization can break through the memory wall for large context, agentic AI.” That pairing of a network-attached NVMe pool with a software tier is Lightbits’ home turf; its LightOS block storage has been built on NVMe/TCP since 2019.

Ramesh Chettuvetty, senior vice president of AI product and business at Lightbits, called Inferra “the world’s first KV cache acceleration engine that eliminates idle GPUs and power context windows of up to 10M tokens on commodity hardware,” and said it “delivers instant payback and generates net positive savings from day one.” Those are vendor claims ahead of any independent testing, and the release doesn’t give pricing, packaging, or a general availability date; the engine is in customer beta today. Lightbits is running live demos of the session-density and latency results at Booth 219 during the summit, and has posted a Pod Efficiency Analyzer for teams that want to model the GPU savings against their own cluster before booking one.

Inferra by Lightbits

The post Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts appeared first on StorageReview.com.

Before yesterdayStorageReview

LTM Builds a Lightwell Remediation Services Practice Around IBM and Red Hat’s $5B Open-Source Program

11 September 2026 at 16:35
IBM and Red Hat Project Lightwell graphic IBM and Red Hat Project Lightwell graphic

LTM, the Larsen & Toubro Group services company that was LTIMindtree until its February rebrand, is building a Lightwell remediation services practice around the $5 billion IBM and Red Hat program for securing open-source software with AI-generated, vendor-validated fixes. IBM’s clearinghouse produces validated, production-ready patches for open-source dependencies; LTM’s job is getting them into customer environments. The company says its planned portfolio spans remediation strategy, dependency analysis, risk-based prioritization, remediation program management, DevSecOps integration, testing and validation, and large-scale deployment support.

IBM and Red Hat Lightwell key art with blue data streams converging on a central patch symbol, the program LTM is now building Lightwell remediation services around

Where LTM Fits in the Lightwell Model

Lightwell’s premise, as IBM and Red Hat laid it out in May, is that AI is accelerating vulnerability discovery faster than enterprises can remediate, so a clearinghouse staffed by more than 20,000 engineers ingests vulnerability data from live deployments, validates fixes, and ships them as production-ready patches through subscription services. In July, the two companies added a Lightwell Network tier, which is generally available with a library of remediations spanning current and legacy libraries, and a Lightwell Clearinghouse Premier tier in limited-availability commercial onboarding, and named a bench of deployment partners that included LTM alongside Accenture, Deloitte, HCLTech, Infosys, Kyndryl, TCS, and others. This week’s announcement turns LTM’s spot on that list into a defined offering, backed by its standing as an IBM Platinum Partner.

“As AI accelerates software development and vulnerability discovery, enterprises need a faster and more scalable approach to remediation,” said Chandan Pani, chief information security officer at LTM. “Lightwell represents a significant advancement in securing the open-source software supply chain by bringing AI-driven remediation and trusted software maintenance into the enterprise. Through our collaboration with IBM, LTM will help organisations strengthen their cyber resilience at scale.” Sandip Patel, managing director of IBM India and South Asia, framed it as collective defense: “As AI accelerates vulnerability discovery, collaboration across the ecosystem becomes increasingly important. Bringing LTM’s engineering and transformation expertise to Lightwell can help enterprises mitigate risk and build more resilient software supply chains.”

Why IBM and Red Hat Want the Integrators

Red Hat’s Ryan King, vice president of AI and infrastructure partners, was the most direct about why the program needs a partner channel at all: “AI-driven discovery has pressed the demand for speed and patch delivery far beyond the means of any one single vendor.” A validated patch from the clearinghouse still has to be prioritized against a customer’s actual dependency graph, tested in that customer’s pipelines, and rolled out across an estate without taking down the applications it protects, and that’s the work LTM says it’ll take on, with the stated goal of heading off zero-day exploits and production downtime.

LTM describes the portfolio as planned, and the release carries no pricing, availability dates, or named customers, so this is a go-to-market announcement for services that are being built around Lightwell’s existing tiers. What it does signal is that Lightwell’s financial-services early adopters, which included Bank of America, Citi, Goldman Sachs, JPMorganChase, and Visa at launch, will be followed by a wider enterprise base that reaches the program through integrators like LTM, and we’ll be watching for the first deployment specifics.

IBM Lightwell

The post LTM Builds a Lightwell Remediation Services Practice Around IBM and Red Hat’s $5B Open-Source Program appeared first on StorageReview.com.

IBM Quantum System Two Heads to Switzerland: 120-Qubit Nighthawk r2 at CSCS by End of 2026

11 September 2026 at 16:25
IBM Quantum System Two in IBM's Poughkeepsie quantum data center, with its mirrored hexagonal enclosure beside cooling manifolds and rows of black compute racks IBM Quantum System Two in IBM's Poughkeepsie quantum data center, with its mirrored hexagonal enclosure beside cooling manifolds and rows of black compute racks

IBM and Lockheed Martin are setting up a quantum innovation hub at ETH Zurich, and its core is Switzerland’s first IBM Quantum System Two, to be installed at the Swiss National Supercomputing Centre (CSCS) in Lugano by the end of 2026. The hub comes out of an offset agreement with armasuisse, Switzerland’s Federal Office for Defence Procurement. IBM will operate the machine, which runs on an IBM Quantum Nighthawk processor, while ETH Zurich supplies the expertise and coordinates access for Swiss universities, startups, and companies.

IBM Quantum System Two in IBM's Poughkeepsie quantum data center, with its mirrored hexagonal enclosure beside cooling manifolds and rows of black compute racks

A 120-Qubit Nighthawk r2 Next to Alps

The Swiss system will run IBM’s Nighthawk r2, which IBM calls the fastest and most advanced processor it has made available to users. It carries 120 programmable qubits and a new high-speed qubit reset architecture that lets it execute more than 100,000 circuits per second, which IBM puts at up to 25x the circuit throughput of the Heron processors it succeeds. IBM also says the part has already demonstrated accurate computations on quantum circuits containing 7,500 gates. Nighthawk is the processor family IBM has been building its near-term roadmap around, most recently when it linked two cryogenic modules below 15 millikelvin on the way to its 2029 fault-tolerant machine.

CSCS will house the System Two in the same facility as its Alps supercomputer and provide the power, cooling, and security, so the quantum node sits directly alongside national HPC infrastructure for the chemistry, materials science, optimization, and financial services work the hub is targeting. That’s the same quantum-centric supercomputing pattern IBM has been describing with partners like AMD, where a quantum processor handles the parts of a problem classical hardware can’t simulate efficiently, and the supercomputer handles the rest. The IBM and ETH Zurich agreement to install and operate the system runs for an initial three years, through 2029. Until the hardware arrives, organizations joining through ETH Zurich get access to IBM’s cloud-based quantum fleet, and the whole arrangement builds on a 10-year collaboration between IBM and ETH Zurich on the next generation of algorithms for AI and quantum computing.

Lockheed Martin’s Two Projects and the Training Pipeline

The hub formalizes two joint projects between IBM and Lockheed Martin: quantum sensing for navigation, and quantum simulation aimed at improving the additive manufacturing of metallic alloys. “This project extends the strategic relationship between Lockheed Martin and IBM in quantum and AI while positioning Switzerland as a leader in these critical, cutting-edge fields,” said Dr. Craig Martell, vice president and chief technology officer at Lockheed Martin. Alessandro Curioni, vice president of algorithms and applications at IBM Research and director of the Zurich Research Laboratory, said the goal is “to put world-class quantum hardware, software and expertise directly into the hands of Switzerland’s thriving academic institutions and industries.”

“Switzerland will gain access to essential research infrastructure that will enable us to further advance the exploration of this emerging technology,” said ETH Zurich President Joël Mesot. Alongside the hardware, the Swiss ecosystem gets IBM Quantum Network and IBM Quantum Platform learning offerings, including coursework, certifications, and workshops, plus support for hackathons, conferences, and partner forums. IBM’s Starling roadmap puts a fault-tolerant system in Poughkeepsie by 2029, and installations like the one at CSCS are where the algorithms for that machine get worked out on utility-scale hardware in the meantime.

IBM Quantum Hardware

The post IBM Quantum System Two Heads to Switzerland: 120-Qubit Nighthawk r2 at CSCS by End of 2026 appeared first on StorageReview.com.

Palantir and NVIDIA Deploy a Sovereign Nemotron Supply Chain Stack, Starting With the 1.3 Million Parts in Every Vera Rubin Rack

10 September 2026 at 20:56
Palantir and NVIDIA logos side by side on a black background Palantir and NVIDIA logos side by side on a black background

Palantir and NVIDIA have built a sovereign AI stack for supply chain operations and are running it first inside NVIDIA’s own supply chain, the one that has to line up 1.3 million parts for every Vera Rubin rack. The stack brings NVIDIA Nemotron open models into Palantir Foundry and its Artificial Intelligence Platform (AIP), grounded in the Palantir Ontology, with the aim of giving planners visibility across the chain, surfacing constraints early, and codifying the operational judgment that, to this point, lives in people’s heads. NVIDIA’s deployment starts with materials allocation, the decisions about which parts go where that set how fast a rack moves from wafer to first token.

Palantir and NVIDIA logos side by side on a black background

NVIDIA’s Supply Chain as the First Customer

A rack-scale AI system needs compute, memory, networking, power, cooling, and mechanical parts to arrive together, across thousands of suppliers and a global manufacturing network, and NVIDIA says the Vera Rubin supply chain is twice the size of Grace Blackwell’s. The new stack gives NVIDIA’s supply chain teams what the companies call a shared command center, beginning with allocation decisions, so they can identify constraints earlier, evaluate alternatives faster, and allocate materials based on end-to-end production impact.

“Supply chains are the operating system of the physical economy, and AI factories are among the most complex systems ever built,” said Jensen Huang, founder and CEO of NVIDIA. “From wafers and components to manufacturing, systems and customer delivery, hundreds of companies and trillions of dollars of global economic activity come together to deliver AI infrastructure.” Palantir cofounder and CEO Alex Karp put it more bluntly: “NVIDIA has arguably the most valuable, intricate and complex supply chain in the world.” The two companies first announced their operational AI work together at GTC DC last October; this is that partnership producing a deployed system.

Post-Trained Nemotron, cuOpt, and a Human in the Loop

Palantir customers post-train Nemotron open models on their own operational data inside Foundry and AIP, using NVIDIA NeMo Data Libraries to prepare and augment it. Within AIP, NVIDIA cuOpt handles optimization and scenario planning, so teams can model supply constraints, weigh tradeoffs, and see the operational impact of an allocation decision before making it. The post-trained model recommends actions, explains the tradeoffs, and flags emerging risks, while the supply chain experts keep the final call. Palantir Autopilot, integrated with the NeMo AutoModel and NeMo RL libraries, closes the loop by feeding each recommendation, planner action, and production outcome back into model improvement, which is how the companies say operational knowledge gets preserved instead of lost when people move on.

NVIDIA’s technical write-up of its own deployment gives a sense of how light the model work is. The production model is Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model with roughly 3 billion parameters active per pass, chosen over the larger Nemotron 3 Ultra for agentic workflows. Post-training used LoRA adapters with the base weights frozen and finished on two B200 GPUs in minutes. On NVIDIA’s internal allocation-decision benchmark, the post-trained Lightning model scored 86.7 percent accuracy, 31.2 points ahead of the untuned Nemotron 3 Ultra and 69.2 points ahead of its own base weights. Those are NVIDIA’s numbers on NVIDIA’s data, but they make the case for the approach: a small open model with the organization’s own decisions trained into it beats a much larger general one on that organization’s problem.

Sovereign by Design

Because the data is NVIDIA’s supply chain, the deployment runs on NVIDIA reference architectures and the jointly developed Palantir Sovereign AI Operating System Reference Architecture, or SAIOS, which Dell Technologies and Cisco support. Proprietary data, model weights, and inference stay inside a single governed environment, and the stack can be deployed on premises with Cisco or Dell, or in colocation and cloud with Rackspace and Nebius. Palantir and NVIDIA say they intend to extend what they learn from NVIDIA’s deployment to customers in agriculture, manufacturing, pharmaceuticals, retail, energy, healthcare, automotive, aerospace, and government, and will show the stack and its industry applications at Palantir’s AIPCon 11.

Palantir AIP

NVIDIA Nemotron

The post Palantir and NVIDIA Deploy a Sovereign Nemotron Supply Chain Stack, Starting With the 1.3 Million Parts in Every Vera Rubin Rack appeared first on StorageReview.com.

Backblaze B2 and WEKA NeuralMesh Validated as a Two-Tier AI Storage Pipeline, With Snap-to-Object Checkpoints in B2

10 September 2026 at 20:40
WEKA and Backblaze logos side by side on a dark red and purple gradient background WEKA and Backblaze logos side by side on a dark red and purple gradient background

Backblaze and WEKA have validated their two platforms together for AI pipelines, pairing WEKA NeuralMesh as the performance tier that feeds GPUs with Backblaze B2 Cloud Storage as the capacity tier that holds everything else. The integration, sizing, tuning, and testing are already done, so an AI infrastructure team can deploy a proven two-tier layout without building and qualifying its own. Certification of B2 for NeuralMesh is underway, and both companies say customers can contact either one to get started now.

WEKA and Backblaze logos side by side on a dark red and purple gradient background

Two Tiers, One Data Lifecycle

Raw, unstructured data, meaning the training sets, media libraries, and source files, lives in B2. When a dataset becomes part of a performance-sensitive job, it’s made available to NeuralMesh and served to the accelerators from there. Once a checkpoint, an output, or any other asset no longer needs high-performance access, it goes back to B2, where it can be reused in a later run or pulled back if a job has to recover to an earlier stage. The companies frame it as speed where the GPUs are and capacity everywhere else, with the data moving between the two as its access pattern changes.

“AI teams need their GPUs fed and an infrastructure with the performance and capacity to support the full AI data workflow. WEKA has mastered the performance tier. We’ve spent nearly two decades doing the same for capacity storage,” said Gleb Budman, CEO of Backblaze. Nilesh Patel, Chief Strategy Officer at WEKA, described the same pressure from the other direction: “AI workloads are stretching storage in two directions at once. GPUs need microsecond access to data to stay fed, while datasets and checkpoints are growing to exabyte scale. Our collaboration with Backblaze gives customers a validated path to both, without the cost of building and testing that integration themselves.”

Snap-to-Object Tested Against B2

Another interesting piece is NeuralMesh’s Snap-to-Object, which the companies say they’ve tested with Backblaze. Snap-to-Object writes a consistent snapshot of a NeuralMesh file system out to an object store, and with B2 as the target, that store is the same capacity tier the raw data and retired checkpoints already end up in. For a training run, that means a team can revert to a checkpoint or recover saved inference data from B2 without improvising a fix in the middle of the job, and without maintaining a separate destination for snapshots. The economic effect is a cloud object tier behind NeuralMesh that’s priced as capacity, holding the snapshots alongside the data they protect.

Backblaze has been positioning B2 as the capacity layer for AI throughout the year, from the B2 Neo offering for neocloud platforms to its performance benchmarking program, and WEKA gives it a performance-tier partner on the GPU side of that.

Backblaze B2 Cloud Storage

WEKA NeuralMesh

The post Backblaze B2 and WEKA NeuralMesh Validated as a Two-Tier AI Storage Pipeline, With Snap-to-Object Checkpoints in B2 appeared first on StorageReview.com.

d-Matrix Raptor XPUs Join NVIDIA MGX Racks Through NVLink Fusion, With First Systems Due Q4 2027

10 September 2026 at 20:07
NVIDIA rendering of a d-Matrix Raptor compute tray with NVLink Fusion, an NVIDIA NVLink switch tray, and the full d-Matrix MGX rack NVIDIA rendering of a d-Matrix Raptor compute tray with NVLink Fusion, an NVIDIA NVLink switch tray, and the full d-Matrix MGX rack

d-Matrix will put its next-generation Raptor inference XPUs into NVIDIA’s MGX rack architecture using NVLink Fusion, under a collaboration with NVIDIA that the company describes as a multi-year product roadmap. The first product is a d-Matrix rack built on the MGX reference design with NVIDIA Vera CPUs, NVLink switches, BlueField-4 DPUs, ConnectX-9 SuperNICs, and Spectrum-X Ethernet, aimed at AI labs, hyperscalers, and neoclouds selling what d-Matrix calls premium, ultra-low-latency token services. Initial availability of Raptor XPUs in the MGX rack is expected in the fourth quarter of 2027.

NVIDIA rendering of a d-Matrix Raptor compute tray with NVLink Fusion, an NVIDIA NVLink switch tray, and the full d-Matrix MGX rack

What NVLink Fusion Gives d-Matrix

The argument for the deal is centered around time and risk. Getting a custom accelerator into production at AI factory scale means sourcing and validating a scale-up interconnect, a rack design, power delivery, liquid cooling, scale-out networking, and a supply chain. NVLink Fusion is NVIDIA’s program for letting third-party XPU and CPU designers plug into a stack NVIDIA has already built. For d-Matrix, that means Raptor gets the same NVLink scale-up domain, MGX rack, cooling, and supply chain that NVIDIA’s own systems use, and data center operators can stand up one-rack architecture that carries GPUs, CPUs, and XPUs.

“Demand for inference is soaring, but capital, time and energy remain finite,” said Sid Sheth, cofounder and CEO of d-Matrix. “With NVLink Fusion and MGX, we can integrate our Raptor XPUs into a broadly deployed, liquid-cooled architecture, giving customers a faster, lower-risk path to deploy and scale ultralow-latency inference.” NVIDIA CEO Jensen Huang framed it from the other side: “With NVIDIA AI infrastructure deployed across cloud and on-premises data centers worldwide, NVLink Fusion gives partners like d-Matrix a path to integrate seamlessly with NVIDIA compute platforms, expanding accelerator choice for customers building the next generation of AI factories.”

The interconnect itself is sixth-generation NVLink, which NVIDIA rates at 3.6TB/s per GPU or XPU and 260TB/s of aggregate bandwidth across a 72-accelerator all-to-all domain, more than 14 times the bandwidth of PCIe Gen6 by NVIDIA’s comparison. d-Matrix plans to use it to connect Raptor XPUs into a single high-bandwidth scale-up domain, with the racks built from modular, cable-free MGX trays. Astera Labs is also part of the design, supplying connectivity for the system as an existing member of the NVLink Fusion ecosystem, which also includes Arm, Intel, Fujitsu, SiFive, Marvell, MediaTek, Samsung, Alchip, GUC, Cadence, Synopsys, Ayar Labs, and Lightmatter. We covered MediaTek’s NVLink Fusion XPU work and AWS bringing NVLink Fusion to Trainium in the last two weeks; d-Matrix is the latest to take the same route.

Disaggregated Inference: GPUs for Prefill, Raptor for Decode

d-Matrix is pitching what it calls heterogeneous disaggregation, where a Raptor rack sits alongside a Vera Rubin NVL72 and takes the phase of inference it’s built for. For AI coding assistants, the example d-Matrix uses, the GPUs handle the compute-heavy prefill phase while the Raptor XPUs run the latency-sensitive decode phase, where interactivity is what the customer is paying for. The same split applies to real-time chatbots and voice agents, and it’s why d-Matrix keeps describing the target as a premium token economy: workloads where buyers pay more for speed.

Raptor is the follow-on to d-Matrix’s Corsair XPU, which is in production today. It extends the company’s memory-centric design with what d-Matrix calls a first-of-its-kind 3D DRAM stacking approach, pairing a DRAM chip with an SRAM compute chip in a single two-story package. Co-founder and CTO Sudeep Bhoja previewed the technology at Hot Chips 2026, and the technical details have been published through IEEE. d-Matrix says Raptor was designed from the start for NVLink Fusion and MGX integration, is expected to tape out before the end of this year, is under evaluation at hyperscalers and frontier labs, and is backed by more than 100 patents. The company is demonstrating the design at the AI Infra Summit next week.

d-Matrix

The post d-Matrix Raptor XPUs Join NVIDIA MGX Racks Through NVLink Fusion, With First Systems Due Q4 2027 appeared first on StorageReview.com.

Samsung High NA EUV DRAM Set for 2028, With 12-Inch Photomasks and a Mistral AI Series D Lead

9 September 2026 at 18:23
ASML TWINSCAN EXE:5000 High NA EUV lithography system with its panels open, the platform Samsung plans to use for DRAM production by 2028 ASML TWINSCAN EXE:5000 High NA EUV lithography system with its panels open, the platform Samsung plans to use for DRAM production by 2028

Samsung Electronics made two moves in two days that both point at the same problem: keeping its memory and foundry lines on the leading edge as AI demand outruns conventional scaling. On September 8, Samsung and ASML announced an expanded partnership under which Samsung will put ASML’s High NA extreme ultraviolet (EUV) lithography into high-volume DRAM manufacturing by 2028 and join the industry effort to move from 6-inch to 12-inch photomasks. A day later, at the South Korea and France state summit in Paris, Samsung announced a strategic partnership with Mistral AI to run the French developer’s models on-premises across its chip operations, along with a lead investment in Mistral’s Series D round.

Samsung High NA EUV for DRAM by 2028 and the 12-Inch Photomask Transition

Samsung says it plans to introduce ASML’s High NA EUV systems into future DRAM high-volume manufacturing by 2028, which the two companies describe as the first such deployment in the industry. High NA raises the numerical aperture of the projection optics from 0.33 to 0.55, and the tighter resolution is meant to extend the DRAM scaling roadmap by simplifying process steps that today require multiple patterning passes. Samsung laid out where that scaling is headed in its 3D memory roadmap at FMS 2026, and High NA is the lithography piece that has to be completed first.

ASML TWINSCAN EXE:5000 High NA EUV lithography system with its panels open, the platform Samsung plans to use for DRAM production by 2028

In parallel, Samsung is joining the industry initiative to develop a 12-inch photomask platform for High NA EUV. The 6-inch mask has been the standard for decades, but High NA systems use anamorphic optics that halve the exposure field, so printing a full-size die on a 6-inch reticle means stitching two exposures together. ASML and Samsung say the larger 12-inch format is expected to raise fab productivity, lower chipmaking costs, and remove those stitching constraints, letting manufacturers use High NA without the throughput penalty. Samsung says it will work with industry partners to develop the mask technologies and supporting infrastructure the new format requires, drawing on its experience running EUV in both memory and foundry production.

ASML TWINSCAN EXE:5000 High NA EUV lithography system with two technicians at the front panel

“The AI era is transforming the semiconductor industry and increasing the importance of technological innovation across the entire value chain,” said Young Hyun Jun, Vice Chairman and CEO of Samsung Electronics. “By further strengthening our collaboration with ASML, we are helping lay the foundation for the next generation of AI and semiconductor innovation.” ASML President and CEO Christophe Fouquet called Samsung “one of ASML’s most important innovation partners for many years,” and both companies said they expect to keep exploring collaboration in advanced memory and new manufacturing approaches.

Mistral AI Models Running On-Premises Inside Samsung’s Fabs

The second announcement puts Mistral’s models, including its flagship Mistral Large, to work inside Samsung’s semiconductor operations. Samsung says it will integrate Mistral’s AI services and solutions to develop customized on-premises models for what it calls intelligence-driven infrastructure, with the on-premises deployment chosen so that process technology and operational data stay entirely within Samsung’s own semiconductor infrastructure.

Mistral AI graphic with icons for regional control, third-party models, and data security on a blue grid

The stated targets are practical fab problems rather than chip design in the abstract. Samsung says it will apply targeted models to defect detection and equipment optimization, with the goal of accelerating development cycles, manufacturing precision, and yield stabilization across its advanced memory and logic chips. As process nodes get more complex, the company argues, fast data analysis inside the fab becomes essential, and Mistral’s stack is meant to change how its chips are designed and manufactured.

“Increasing complexities involved in AI chip design and manufacturing requires continuous innovation in semiconductor technologies,” Jun said, this time bylined as head of Samsung’s Device Solutions Division. Mistral co-founder and CEO Arthur Mensch said the company is “proud to support Samsung Electronics with our expertise in electronics and semiconductors, helping to improve how chips are designed and manufactured.”

Samsung also led Mistral’s Series D funding round, taking what it describes as a strategic equity stake to support long-term technology collaboration. Samsung did not disclose the round size or its share. The investment puts Samsung alongside its own lithography supplier on Mistral’s cap table: ASML led Mistral’s Series C exactly one year earlier, in September 2025, with a 1.3 billion euro investment for roughly an 11 percent stake. Mistral has been building out its own infrastructure as well, with Mistral Compute AI factories on NVIDIA GB300 NVL72 and a deepening collaboration with Dell on enterprise deployments.

Taken together, the two announcements show Samsung buying capability on both ends of the fab: the lithography that determines how small the next DRAM cell can get, and the models it wants watching the line once those tools are running.

Samsung Semiconductor

The post Samsung High NA EUV DRAM Set for 2028, With 12-Inch Photomasks and a Mistral AI Series D Lead appeared first on StorageReview.com.

Qualcomm and Amazon Sign Multi-Generation Deal for Custom AI Inference Silicon and 1.6T Optical Interconnects

8 September 2026 at 17:17

Qualcomm Technologies and Amazon have entered into a multi-generation collaboration to deliver customized silicon at scale for AWS’s AI data centers, with AI inference as the primary target. The agreement pairs Qualcomm’s power-efficient processing, silicon design, and system-level integration with Amazon’s AI infrastructure, and is aimed at the compute, memory bandwidth, networking, and energy constraints that come with running inference at hyperscale volume. Neither company named specific parts, delivery dates, or financial terms.

Qualcomm Dragonfly AI300 data center render with rows of racks and the Dragonfly AI300 badge

Custom Silicon and 1.6T Optical Connectivity

Beyond the compute silicon, the two companies are co-developing high-performance optical connectivity for Amazon’s data center networks, with solutions reaching 1.6T and future generations on the roadmap. That work draws on Qualcomm’s high-speed SerDes and optical DSP technologies, the same connectivity portfolio the company laid out at its June Investor Day alongside the Dragonfly C1000 CPU and AI300 accelerator. The intent is to keep network bandwidth from becoming the bottleneck as inference clusters scale out across racks.

The relationship also runs in the other direction. Qualcomm plans to deepen its use of AWS AI infrastructure, including Amazon Bedrock, for its electronic design automation (EDA) workloads, with the stated goal of shortening chip design cycles. That puts AWS compute behind the modeling and verification work for the very silicon Qualcomm will be building for Amazon.

What the Deal Signals

Executives on both sides framed the agreement around compute and connectivity advancing together. “As AI demand accelerates, data center infrastructure will require advances in both computing and connectivity to deliver greater performance with more efficiency,” said Cristiano Amon, President and CEO of Qualcomm Incorporated. Prasad Kalyanaraman, Vice President at AWS, said the collaboration “builds on a strong foundation of partnership” and that “by working together on customized silicon and advanced connectivity, we’re delivering more performant, efficient, and cost-effective infrastructure for our customers.”

The announcement lands as Qualcomm pushes hard into the data center. The June Investor Day put a CPU, a roadmap of AI200, AI250, and AI300 inference accelerators, and a connectivity portfolio on the table, and the company closed its acquisition of Modular in July to give that silicon a software stack that does not depend on CUDA. AWS already designs its own Trainium and Inferentia accelerators through Annapurna Labs, so the open question is where Qualcomm-built silicon fits alongside those parts. The release does not say, but a multi-generation commitment from the largest cloud provider is the strongest customer signal Qualcomm’s data center business has shown so far.

The post Qualcomm and Amazon Sign Multi-Generation Deal for Custom AI Inference Silicon and 1.6T Optical Interconnects appeared first on StorageReview.com.

VMware Explore 2026: Broadcom Doubles Down on Private Cloud Economics and Agentic AI

5 September 2026 at 17:10
Crowd entering The Hub welcome reception at VMware Explore 2026 with mirrored costumed performers Crowd entering The Hub welcome reception at VMware Explore 2026 with mirrored costumed performers

VMware Explore kicked off Monday morning, though the schedule felt unusual compared to past VMworld Explore events, as Broadcom moved the traditional opening keynote to the afternoon.

Attendees in front of the VMware Explore 2026 venue map and schedule wall in Las Vegas

That was not necessarily a bad thing, as it gave us time to catch a few breakout sessions and catch up with vendors on the show floor. In this piece, we’ll dive into the keynote announcements, along with notes from the one-on-one interviews and roundtable panels we attended throughout the conference. Our earlier on-the-ground report covered Lenovo’s turnkey AI and memory-crunch conversation.

Keynote: Shaping the Future of Private AI Cloud and Agentic Innovation

We covered VMware’s morning news releases in an earlier post, but the general session keynote added technical and economic context to those announcements. The morning press releases focused on VMware Private AI Cloud, validated models, and deny-by-default runtime agent security; the keynote focused on operational realities of IT. Broadcom leadership spent their time on stage addressing the economics facing modern IT departments.

Ram Velaga, president of Broadcom’s Infrastructure Software Group, set the tone for the keynotes by emphasizing that hardware cost inflation, particularly pressure on memory pricing, has become a significant infrastructure-planning concern rather than a temporary supply-chain hiccup. The takeaway wasn’t a call to forklift everything to the public cloud, but rather a call for pragmatic workload repatriation. Velaga hammered home VMware’s foundational argument: enterprise IT succeeds by decoupling complex compute, storage, and networking layers from the applications running on top of them.

A Broadcom presenter opens the VMware Explore 2026 general session keynote

Next up, Paul Turner and his engineering team demonstrated the VCF AI Assistant and VMware’s live-patching capabilities. This demonstration addressed a familiar operational challenge: remediating critical vulnerabilities without introducing unnecessary downtime. During the live demo, the system detected administrative configuration drift (in this case, a cluster in which DRS had been inadvertently disabled weeks earlier), ran automated pre-checks, and applied ESXi live patches to resolve performance bottlenecks. This was all accomplished without bouncing virtual machines, which could have caused expensive and unnecessary downtime.

Keynote presenter on the VMware Explore 2026 stage during the VCF AI Assistant and ESXi live patching demo

One of the more insightful portions of the keynote was the customer perspective from Hemant Deshpande of Standard Chartered Bank, who walked the audience through their standardized, factory-integrated rack build.

For anyone who has managed multi-rack deployments, the appeal is obvious: pre-cabling, validation, and configuration of compute, networking, and vSAN storage at the supplier facility cut deployment timelines from months to days while establishing clean, isolated failure domains. As a bonus, moving to high-density virtualized clusters dropped their data center power consumption by 40 to 45 percent. We loved how they showed that a standardized software-defined footprint directly curbs capital and operating expenses.

Two presenters on the VMware Explore 2026 keynote stage during the Standard Chartered factory-integrated rack segment

Continuing the customer-focus theme, United Airlines took the stage to address the platform engineering side, detailing its adoption of VMware’s Kubernetes capabilities on VCF. Rather than running separate environments for containers and legacy VMs, they integrated infrastructure provisioning directly into their CI/CD pipelines via Terraform and Avi Global Server Load Balancing. This gives application teams active-active microservices across multiple on-prem sites and hybrid clouds without submitting help desk tickets, offering developers a public-cloud-style experience while keeping infrastructure tightly managed on-premises.

United Airlines presenters on the VMware Explore 2026 keynote stage discussing Kubernetes on VCF

Finally, Purnima Padmanabhan, VP and General Manager of the Tanzu Division, stepped up to address something that we have been pondering lately: vulnerabilities in the AI software supply chain. To address these issues, Broadcom’s Tanzu division announced TrueSource on Monday. Autonomous and agentic AI pipelines frequently pull dependencies for Python, Java, Node.js, and Spring from public repositories, exposing a significant security risk. Their new product, TrueSource, serves as a curated, cryptographically signed, and continuously patched upstream repository for both runtime frameworks and core data stores including PostgreSQL, MySQL, RabbitMQ, and Valkey. This can be layered alongside Tanzu agent sandboxing and hypervisor-level distributed firewalls. The message was clear: enterprise AI governance requires locking down everything from bare metal to the model runtime before agents touch production environments.

They ended the Keynote with the announcement that Explore 2027 will still be in Las Vegas but moves to Resorts World and to May 3 through 6, 2027.

VMware Explore 2027 save-the-date slide: Las Vegas, Resorts World, May 3 to 6, 2027

Once the keynote was wrapped, we headed over to The Hub for the Welcome Reception.

Crowd entering The Hub welcome reception at VMware Explore 2026 with mirrored costumed performers

In the sections below, we break down what we learned from our discussions, interviews, and small-group panels across the rest of day one.

What We Heard from VMware Leadership

During our discussion, Dilpreet Bindra, Senior Director of Engineering at VMware by Broadcom, reflected on how the definition of an enterprise “workload” has expanded from traditional virtual machines to encompass cloud-native containers and advanced AI applications.

Both Bindra and I have been attending KubeCon religiously, so Kubernetes and container technology seemed like a good place to start our conversation.

Bindra began our conversation by discussing his involvement in the Cloud Native Computing Foundation (CNCF) and his experiences navigating the CNCF and the energy at KubeCon.

Tom Fenton with Dilpreet Bindra, Senior Director of Engineering at VMware by Broadcom, at VMware Explore 2026

He highlighted the shift in VMware’s engineering culture toward open-source contributions, prioritizing upstream collaboration to ensure innovations don’t remain locked inside proprietary products. Bindra noted that VMware has accelerated Kubernetes innovation by actively engaging with the CNCF community.  My feeling was that for him, Kubernetes on VCF is a critical arena for aligning enterprise-grade orchestration with the incredibly fast-paced, community-driven ecosystem that modern developers rely on.

He then drew on his deep background in core VMware technologies, including ESXi and vMotion.  He explained that foundational principles like resource isolation and mobility remain as critical today as ever. However, they must be adapted to avoid the bottlenecks inherent in legacy infrastructure. He emphasized that the true promise of VMware Cloud Foundation (VCF) 9.1 lies in its ability to converge distinct workloads, such as VMs, containers, and AI, onto a unified platform. Bindra sees VCF 9.1 as a way to bring VMs, containers, and AI workloads onto a common infrastructure platform. The goal is to give IT teams the controls they need without forcing developers to navigate separate infrastructure stacks for each workload type.

Addressing the rapid evolution of Private AI and autonomous infrastructure, Bindra offered pragmatic advice for architects designing tomorrow’s data centers. He cautioned against rigid architectural lock-in, given how quickly GPU capabilities, model sizes, and network demands are changing in the AI space. When we touched on the AIOps workflows showcased at VMware Explore 2026, he drew a distinct line between AI-assisted operations and full autonomy. Bindra argued that while AI is becoming exceptionally good at recommending optimizations, establishing true trust in autonomous infrastructure takes time. He recommended that platform teams start by delegating routine, low-risk operational tasks to AI agents, keeping critical production decisions firmly under human supervision until technology and organizational trust mature.

Later, we sat down with Sabina Anja, Chief Technologist and Executive Advisor for VMware Cloud Foundation, to discuss how modern infrastructure is rapidly evolving under the weight of AI workloads and distributed systems. But before diving into that, we touched on my fascination with agentic AI troubleshooting tools, which led to a deeper conversation about the role of AI assistants.

Tom Fenton with Sabina Anja, Chief Technologist for VMware Cloud Foundation, in front of the Explore backdrop

Sabina argued that the growing complexity of modern environments is increasingly overwhelming the people responsible for operating them. With environments growing increasingly complex and distributed, an AI assistant shouldn’t just be an automated script; it needs to synthesize tribal best practices, correlate logs with historical telemetry, and surface critical insights so engineers don’t spend hours manually correlating data across disconnected logging and telemetry tools.

From there, our conversation turned toward the physical realities of the modern data center, particularly resource allocation, traffic patterns, and memory tiering. Sabina, drawing on her deep networking background, pointed out that as AI agents constantly communicate, we will see consistent, saturated traffic that demands smarter workload localization, quality of service, and traffic engineering.

Wrapping up our discussion, we touched on the economic and physical constraints facing infrastructure today, from volatile DRAM costs driving the adoption of memory tiering using NVMe to the staggering power demands of new accelerators. Ultimately, Sabina’s advice to architects was straightforward: look ahead, understand the true power and cooling requirements of your workloads, and design for where data demand is headed rather than relying on legacy assumptions.

Sabina also delivered a standout keynote segment detailing how distributed intelligence must map directly to the realities of physical data centers.

Sabina Anja, Chief Technologist, VMware Cloud Foundation, on screen during her VMware Explore 2026 keynote segment

We had a chance to sit down with Chris Wolf, Global Head of AI and Advanced Services at Broadcom’s VMware Cloud Foundation Division, and reflect on VMware’s private AI journey, which he initiated three years ago.

Tom Fenton with Chris Wolf, Global Head of AI and Advanced Services for the VMware Cloud Foundation Division, at VMware Explore 2026

We first asked him how his early assumptions about private AI held up against reality. He explained that while foundational bets on open-source runtimes like vLLM (an open-source large language model (LLM) serving runtime) and Harbor container registries proved spot-on, they were initially ahead of their time, and it took the market a couple of years to catch up. Wolf said VMware initially placed greater emphasis on NVIDIA vGPU technology, but customer demand for practical edge deployments pushed the team toward simpler, more cost-effective passthrough configurations. He then went on to say that they were very much in touch with their customer base and, in response to customer demand for practical edge deployments, they adapted VMware’s virtualization stack to accommodate simpler, cost-effective passthrough configurations.

As our conversation turned to the broader data center landscape, we pressed him on whether organizations were beginning to silo their infrastructure again around traditional virtual machines, containers, and AI workloads. Chris pointed out that the opposite is happening: he argued that rising power demand and cooling requirements, along with the operating costs of cloud-based AI consumption, are prompting organizations to reconsider where they run AI workloads. Wolf argued that modern virtualization overhead is sufficiently low for many AI workloads that organizations can consolidate infrastructure without giving up meaningful performance, and, as such, enterprises can greatly reduce their physical footprints, rein in software licensing costs, and deploy modest GPU clusters that maximize efficiency without incurring high public cloud expenses.

Looking toward the future, we asked Chris to forecast the next evolution of enterprise infrastructure and where agentic AI fits into the equation. While he is optimistic that specialized smaller models and autonomous agents will relieve human operators of tedious operational toil, he emphasized the critical need for strict governance, human-in-the-loop approvals, and audit trails to prevent unmonitored scripts from destabilizing mission-critical environments. Wolf’s view was that enterprise AI will increasingly rely on distributed architectures and smaller, specialized models rather than depending exclusively on massive centralized models.

Q & A with Ram Velaga

We, along with a few other journalists, were invited to a question-and-answer session with Ram Velaga, president of Broadcom’s Infrastructure Software Group.

Title slide for Ram Velaga, President of the Infrastructure Software Group at Broadcom, at the VMware Explore 2026 press Q and A

The session started with a question about how Broadcom is prioritizing its engineering resources to support enterprise customers as they transition agentic AI workloads into production. He was quick to put the numbers in perspective, revealing that Broadcom is investing well over $5 billion annually in software engineering.  However, he made it clear that they are not pouring every dollar into AI features and that between 50% and 60% of their engineering capacity is strictly devoted to core operational foundations: compute, networking, and storage, to provide rock-solid reliability and seamless, zero-downtime fleet updates that VMware customers have come to expect.

Ram then went on to say that rather than buying into dense, single-vendor architectures, he outlined a future driven by heterogeneous pools of compute connected over high-speed network fabrics with shared memory. He pushed back firmly against industry buzzwords like “control plane.” He argued that VMware’s role remains the same: abstracting underlying infrastructure so customers can run workloads across different processors, accelerators, and storage technologies.

Culturally, he described VMware as an efficiency-first business. While public hyperscalers profit when customers consume excess storage and compute resources, VMware proves its value by reducing resource footprints and optimizing on-premises hardware utilization. To reinforce this across its portfolio, Broadcom has been tearing down internal silos, for example by mobilizing Symantec’s engineering team to deliver EDR-like hypervisor defense and virtual patching directly within VMware Cloud Foundation.

He broke down the broader market forces driving enterprise workloads back to private clouds. Ram pointed out that the rush to on-premises infrastructure is not just about avoiding unpredictable public cloud token bills but also about preserving corporate costs. Velaga felt that data sovereignty and intellectual property concerns are becoming important factors in private AI adoption. For organizations in regulated industries, keeping sensitive data and AI workflows under tighter control can be as important as managing the cost of public AI services.

On the partner front, he defended VMware’s aggressive overhaul of the Cloud Service Provider (CSP) ecosystem. Likening the previous setup to an uncurated franchise in which low-value resellers undercut quality operators, he explained that Broadcom is deliberately trimming out transactional intermediaries to elevate true partners who invest in hardware, delivery, and the customer experience. Moving forward, his priority is to eliminate surprises, deliver seamless monthly maintenance releases, and leverage AI internally to write better, more resilient code.

The Partner Perspective

VMware has become more focused and reliant on its partners. VMware brought Stephen Ayoub, Co-Founder and President at AHEAD, and Ryan Sheehan, Senior Vice President of Advanced Solutions at SHI International Corp.

Partner roundtable panel at VMware Explore 2026 with AHEAD and SHI executives on the Explore stage

The RoundTable opened with an overview of how customer conversations have shifted over the past 12 to 18 months. The panelists agreed that infrastructure conversations have shifted away from purely operational spending toward investments that must demonstrate measurable business outcomes. Instead of infrastructure discussions being limited to technical managers, modern purchases now require alignment among the CEO, CFO, CIO, and business stakeholders who demand measurable outcomes in revenue, security, and productivity before greenlighting multi-million-dollar investments. Stephen from AHEAD noted that while AI experimentation has rapidly matured, organizations are deeply concerned about “shadow AI,” data leakage, and unpredictable tokenomics. Consequently, enterprises are prioritizing architectural foundations such as VMware Cloud Foundation (VCF 9) to secure their data, govern access, and rein in runaway public cloud costs.

The panel was pressed on how Broadcom’s restructuring, specifically slashing thousands of complex SKUs down to core product bundles, is playing out with customers. While acknowledging that pricing adjustments initially triggered emotional pushback, the panelists agreed that their customers have started to accept the changes.

Ryan from SHI and Stephen agreed that the traditional transactional reseller model is becoming less viable as customers increasingly expect partners to take responsibility for architecture, implementation, and operational outcomes. Today, clients demand technical systems integrators who take full accountability for deploying the entire stack.

To wrap up the RoundTable, Broadcom discussed how they are evolving their go-to-market strategy and partner enablement to ensure customers actually turn on and adopt what they purchase. They reiterated that Broadcom is strictly a product and innovation company with no internal professional services wing, meaning the business is completely dependent on elite channel partners to drive deployment. To support this, Broadcom is heavily funding deployment services and expanding rigorous, hands-on training initiatives, which give partner engineers direct access to VMware’s core product teams.

vSphere Standard in the News

Paul Turner told reporters at the show that Broadcom plans to update vSphere Standard, a move first reported by The Register. If confirmed, the move could be significant for smaller organizations that do not require the full VMware Cloud Foundation stack.

Final Thoughts: The Economics Behind the AI Message

In reflection, VMware Explore 2026 made it clear that Broadcom’s private AI strategy is about far more than running large language models. Across the keynote, executive interviews, and partner discussions, the recurring themes were infrastructure economics, operational efficiency, and the practical challenges of deploying AI at scale.

Broadcom executives discussed everything from live patching and workload consolidation to AI governance, power consumption, cooling constraints, and the growing pressure on memory and accelerator resources. Conversations with VMware leaders reinforced a common message: the next generation of enterprise infrastructure must support traditional VMs, containers, and AI workloads on a common platform while giving organizations greater control over costs, security, and data location.

The post VMware Explore 2026: Broadcom Doubles Down on Private Cloud Economics and Agentic AI appeared first on StorageReview.com.

OpenAI GPT-6 Astra Hits GA in Microsoft Foundry: Computer Use, Agentic Execution, and $10 to $75 per Million Tokens

5 September 2026 at 17:09
GPT-6 Astra key visual from Microsoft Foundry: a spiral galaxy of stars on black with GPT and Astra lettering GPT-6 Astra key visual from Microsoft Foundry: a spiral galaxy of stars on black with GPT and Astra lettering

Microsoft announced the general availability of GPT-6 Astra within Microsoft Foundry on Azure. The frontier model is engineered to transition enterprise generative AI from interactive chat interfaces toward autonomous agentic workflows, providing multi-step planning, deliberate decision support, and cross-application tool execution. OpenAI calls Astra its most aligned model to date, and Microsoft says it is designed for token efficiency on complex work.

GPT-6 Astra key visual from Microsoft Foundry: a spiral galaxy of stars on black with GPT and Astra lettering

The release focuses on solving deployment friction around operational fundamentals such as identity management, secure networking, compliance, and governance. By integrating GPT-6 Astra directly into Microsoft Foundry, enterprise IT teams can configure and deploy agentic pipelines while maintaining enterprise boundary controls across their cloud infrastructure.

Multi-Step Reasoning and Application Execution

GPT-6 Astra is designed to handle open-ended objectives by evaluating options, building structured plans, weighing execution trade-offs, and dynamically incorporating new parameters as tasks proceed. The model generates structured artifacts, including technical reports, analysis spreadsheets, and slide decks, that align with pre-existing enterprise templates and business standards for human verification.

A central architectural feature is the model’s computer-use capability, which enables it to operate software interfaces on a user’s behalf. Astra can parse on-screen visual information and manipulate approved user interfaces, allowing it to navigate legacy applications and tools that lack dedicated REST APIs. Key deployment scenarios include:

Software engineering teams can use the model to reproduce reported defects, trace root causes across code repositories, generate proposed fixes, and prepare pull requests for developer review. For business intelligence, Astra automates the construction and iterative refinement of Power BI dashboards. In broader operations, the engine supports automated record updates, form processing, and end-to-end interface testing across enterprise software suites.

Governance, Security, and Containment Controls

Because computer use and autonomous agency present unique security challenges, Microsoft Foundry integrates containment safeguards around Astra. Workflows can be constructed with scoped credentials, role-based access control (RBAC), and required human-approval checkpoints before executing high-impact actions.

Foundry wraps the model with Microsoft Entra identity and access management, pervasive encryption in transit and at rest, private networking routing, automated content filtering, and comprehensive activity logging. Microsoft states that prompts and outputs within Foundry are not used to train the models.

Industry partners evaluating the platform highlighted the balance of development velocity and operational discipline. Replit CTO Luis Hector Chavez said Astra moves AI tooling “beyond code generation to active software creation.” Anirban Nandi, VP of Data and AI at Albertsons Companies, said Azure OpenAI on Microsoft Foundry gives his teams “that balance of speed and control,” with consistent security, governance, and operational controls as they scale.

Deployment Options and Availability

Deployment Context Length Input Cached Input Cached Writes Output
Standard Global (USD $/million tokens)
Standard Global Short context $10.00 $1.00 $12.50 $50.00
Standard Global Long context $20.00 $2.00 $25.00 $75.00
Standard Data Zone (US) (USD $/million tokens)
Standard Data Zone (US) Short context $11.00 $1.10 $13.75 $55.00
Standard Data Zone (US) Long context $22.00 $2.20 $27.50 $82.50
Provisioned Throughput pricing varies by deployment type. U.S. Data Zone Provisioned Throughput is priced at a 10% premium to Global Provisioned Throughput. For current rates and terms, see the Azure OpenAI pricing page.

 

GPT-6 Astra is available immediately across Global and U.S. Data Zone geographies with two distinct infrastructure deployment tiers. The Standard deployment tier provides consumption-based, pay-as-you-go provisioning suited for variable enterprise demand. For workloads requiring deterministic latency baselines and dedicated compute capacity, the Provisioned Throughput tier offers guaranteed model-processing throughput. Microsoft has been expanding Azure’s inference capacity with its own Maia 200 accelerators alongside NVIDIA and AMD hardware. In terms of infrastructure pricing, U.S. Data Zone Provisioned Throughput carries a 10 percent premium over standard Global Provisioned Throughput rates.

The post OpenAI GPT-6 Astra Hits GA in Microsoft Foundry: Computer Use, Agentic Execution, and $10 to $75 per Million Tokens appeared first on StorageReview.com.

NVIDIA to Acquire Hugging Face for $12.93B, Pledges the Platform Stays Open and Hardware Neutral

4 September 2026 at 18:01

NVIDIA has agreed to acquire Hugging Face for $12.93 billion, a transaction that would extend the company’s position from accelerated compute and AI infrastructure into one of the industry’s most widely used platforms for open models, datasets, and application development.

NVIDIA and Hugging Face logos joined by a heart, the lockup NVIDIA used to announce the Hugging Face acquisition

In an announcement published on the NVIDIA website, CEO Jensen Huang said the company plans to scale Hugging Face’s platform, strengthen its underlying infrastructure, and broaden access to AI development and deployment capabilities. The company said Hugging Face will retain its brand and continue operating as an open platform after the transaction.

Hugging Face has developed into a major distribution and collaboration layer for the open-model ecosystem. NVIDIA stated that the platform serves more than 18 million developers, researchers, and creators, hosting over 3 million models, 500,000 datasets, and 1 million applications. More than 200,000 companies use the platform to discover, evaluate, customize, and deploy AI capabilities.

The Neutrality Question

The central issue for existing Hugging Face users will be platform neutrality. NVIDIA said developers will continue to select their preferred models, frameworks, clouds, inference providers, and compute platforms. The company also said that NVIDIA hardware will not be required to build or deploy through Hugging Face, and that the platform will continue to support multi-cloud and multi-accelerator development.

That commitment matters because Hugging Face acts as a common layer across a fragmented AI stack. Enterprises use the platform to assess models, manage artifacts, access datasets, and move AI workloads between development and deployment environments. Maintaining support for non-NVIDIA hardware, alternative cloud providers, and models from other developers would preserve the flexibility that has helped make the platform foundational to the adoption of open AI.

Doubling Down on Open-Weight AI

The deal also reinforces NVIDIA’s investment in open-weight AI. Unlike proprietary AI services, open-weight models can be downloaded, adapted, and operated by organizations using their own infrastructure. NVIDIA positioned the model as a way for enterprises, public institutions, startups, and universities to apply advanced AI capabilities without training each model from scratch or relying exclusively on frontier-model APIs.

Chart of Hugging Face repositories by organization from January 2025 to August 2026, with NVIDIA leading at more than 800 repos

NVIDIA has already been an active contributor to Hugging Face. According to Huang’s post, NVIDIA has published more than 500 models and over 250 open datasets through the platform. NVIDIA characterized itself as the platform’s largest contributor of open models and data, and it was already an investor, having participated in Hugging Face’s $235 million funding round in 2023.

Deal Context and What Comes Next

Huang framed the deal in his post as a continuation rather than a takeover, crediting “Clem, Julien, Thomas and the team at Hugging Face” with having “built something remarkable,” and describing the platform as “a vibrant home for the open model developer community.” NVIDIA’s stated plan is to “scale Hugging Face’s platform, strengthen its infrastructure and expand access to AI,” with the Hugging Face team continuing on under its own brand. Huang closed the post on the ecosystem pitch: “Together, we will make AI more open, more capable and more accessible to people and institutions around the world.”

For NVIDIA, the acquisition provides control of a strategically significant developer channel at a time when AI infrastructure competition is moving beyond GPUs. The company’s 2020 acquisition of Mellanox helped establish its position in data center networking and full-stack infrastructure. Hugging Face would add a widely used platform at the model and developer layer, connecting NVIDIA more directly with organizations deciding which models, tools, datasets, clouds, and accelerators to use.

Neither company has given a closing timeline, and a transaction of this size will need regulatory sign-off. NVIDIA has said Hugging Face will remain open and hardware-neutral, a position that will be examined closely by the developer community and the enterprise customers that depend on the platform’s multi-vendor model.

The post NVIDIA to Acquire Hugging Face for $12.93B, Pledges the Platform Stays Open and Hardware Neutral appeared first on StorageReview.com.

ASUS Lays Out a Full AI Factory Platform: Vera Rubin NVL72 Racks, STX Storage, and a Governance Layer

4 September 2026 at 17:54
ASUS ESC8000A E13P ASUS ESC8000A E13P

ASUS is expanding its role in AI infrastructure, moving from individual AI servers to platforms for building, deploying, and operating entire AI factories. At AI Tech 2026 in Seoul, the company laid out a broader strategy that brings accelerated computing, networking, storage, deployment software, infrastructure management, and AI governance together under one platform.

That puts ASUS more directly alongside Supermicro and GIGABYTE, which are also expanding from server hardware into complete AI infrastructure platforms. The difference from a conventional server launch is that ASUS is extending its role into operations, with tools for infrastructure planning, resource management, MLOps, and governance covering the deployment lifecycle before and after the hardware is installed.

Storage is also an important part of their broader strategy, with new AI-native and object storage systems and an ecosystem that includes Samsung and WD alongside NVIDIA, AMD, Intel, IBM, Schneider Electric, Crusoe, Foxlink, and Aleria.

ASUS Takes AI Factory Planning Into Operations

ASUS is extending its involvement into the planning stage through the NVIDIA DSX Sim Blueprint, which can be used to create digital twins of proposed AI factories before the physical infrastructure is installed. Working with Schneider Electric, AVEVA, and IBM, the environment can model compute, networking, storage, power, cooling, and facility infrastructure during the design process.

Once the infrastructure is deployed, ASUS Control Center and ASUS Infrastructure Deployment Center handle deployment and infrastructure management. The ASUS AI software platform extends into day-to-day AI operations with Quota & Billing tools for resource management and an MLOps Portal for AI development and deployment workflows.

ASUS is also adding a governance layer that connects enterprise policies with AI services and autonomous agents. This extends the platform beyond monitoring the physical infrastructure and into how organizations manage access, resources, and AI workloads running on it.

At the rack level, the ASUS AI POD XA VR721-E3 is based on NVIDIA Vera Rubin NVL72. ASUS claims the system delivers 10 times the performance per watt of the previous generation. The company is also introducing the XA NR1I-E12LR and XA NR1I-E12L based on NVIDIA HGX Rubin NVL8 for AI training, inference, and post-training workloads.

NVIDIA CES 2026 slide of a Vera Rubin POD with its six chips: Vera, Rubin, NVLink6 Switch, CX9, BF4, and Spectrum-X CPO

For agentic AI workloads, the 2U ASUS XA P2N-E2 uses the NVIDIA MGX architecture with two NVIDIA Vera CPUs and support for up to two dual-slot NVIDIA GPUs. ASUS lists agentic reasoning, data processing, and orchestration among the intended workloads for the system.

The ESC8000-E12P supports NVIDIA RTX PRO 6000 and RTX PRO 4500 Blackwell Server Edition GPUs for enterprise inference, vision AI, and visual computing. ASUS is also taking the NVIDIA platform to the edge with the PE3000N, which uses the NVIDIA Jetson Thor T5000 module for real-time inference, sensor fusion, robotics, and industrial automation.

Storage and Partners Broaden the AI Factory Platform

Storage is being integrated directly into ASUS’s AI factory platform. The UF920-E3-RS24 is based on the NVIDIA STX modular foundation for AI-native storage and is joined by the ASUS OJ340A-RS60 object storage system and VS320D-RS26N storage system.

ASUS UF920-E3-RS24 AI-native storage server built on the NVIDIA STX foundation, part of the ASUS AI factory platform

The systems are intended to address the capacity, availability, and data management requirements of AI pipelines, spanning AI-native storage, object storage, and broader storage infrastructure. This also brings storage into the same infrastructure strategy as ASUS’s compute, networking, deployment, and management products.

The partner roster also shows how far ASUS is extending its infrastructure strategy. Samsung and WD are both named storage partners, while NVIDIA, AMD, and Intel cover major compute platforms. IBM, Schneider Electric, Crusoe, Foxlink, Aleria, and others are involved across software, facilities, infrastructure, and deployment.

ASUS is positioning itself closer to large infrastructure providers that can supply more than compute servers, including storage, software, deployment tools, and operational management.

Intel and AMD Expand the Server Portfolio

NVIDIA isn’t the only compute platform involved in ASUS’s AI infrastructure plans. The company also showed new Intel and AMD systems covering AI, HPC, and more traditional enterprise workloads.

The RS700-E12-RS4U and RS720-E12-RS12U use Intel Xeon 6 processors, while the 6U XA P8I-E13A uses Intel’s next-generation Xeon processor platform with support for GPU acceleration.

On the AMD side, the RS720A-E14B-R32U and RS500A-E14B-R12U use AMD EPYC 9006 Series processors. ASUS recently expanded its EPYC 9006 server portfolio around AMD’s efficiency-focused SP8 socket, and these systems extend that platform into its broader AI infrastructure lineup.

The ESC8000A-E13P pairs AMD server hardware with AMD Instinct MI350P PCIe accelerators. ASUS is positioning the system for inference, agentic AI, and HPC workloads.

ASUS ESC8000A-E13P server with AMD Instinct MI350P PCIe accelerators for inference and agentic AI

ASUS Extends AI Infrastructure to the Edge

ASUS also introduced the RUC-2000 series for industrial edge AI. The systems use Intel Core Ultra Series 3 processors and offer up to 180 AI TOPS, with ASUS targeting machine vision, video analytics, industrial automation, and in-vehicle applications. The rugged, fanless design is intended for environments where hardware may need to operate across wider temperature and voltage ranges and tolerate electrical noise.

ASUS RUC-2000 series fanless industrial edge AI computer with Intel Core Ultra Series 3

ASUS’s industrial edge portfolio extends the AI factory strategy to manufacturing, transportation, public safety, healthcare, and other environments where data is generated outside centralized infrastructure.

ASUS RUC-2000H rugged edge AI system from the ASUS AI factory portfolio

ASUS is effectively broadening its role from supplying the servers that run AI workloads to supplying more of the surrounding infrastructure. With Vera Rubin rack-scale systems, Intel and AMD servers, dedicated AI storage, edge hardware, digital-twin planning, management software, and governance tools, ASUS is covering more of the infrastructure required to build and operate an AI factory.

ASUS AI Infrastructure Availability

ASUS servers are available worldwide. Availability of individual systems and other products varies by region and local regulatory requirements, and ASUS directs customers to regional representatives for specific availability information.

 

The post ASUS Lays Out a Full AI Factory Platform: Vera Rubin NVL72 Racks, STX Storage, and a Governance Layer appeared first on StorageReview.com.

Equinix Inference Exchange Brings NVIDIA Compute and 200+ Open Models Closer to Enterprise Data

3 September 2026 at 16:22

Equinix has expanded its partnership with NVIDIA and entered a new collaboration with Together AI to launch Equinix Inference Exchange. Designed as a distributed AI inference architecture for enterprise deployments, the platform aims to shift compute workloads closer to core data repositories, end users, and operational applications. Announced alongside Equinix Fabric One at the Equinix Horizon event, the initiative combines NVIDIA enterprise hardware architectures, Together AI’s open-source model platform, and Equinix’s global interconnection infrastructure to address latency, data sovereignty, and networking bottlenecks.

Equinix Inference Exchange graphic showing interconnected Equinix data centers across the globe, bringing inference closer to where data, users, and applications live

Addressing Distributed Inference and Data Gravity

As enterprise AI initiatives transition from prototype validation to high-throughput production, inference placement dictates overall cost, response latency, and compliance posture. Moving model execution closer to data sources reduces the cost of backhauling data across public clouds while providing deterministic latency profiles. Managing distributed infrastructure across disparate environments, however, introduces operational overhead and networking complexity.

Equinix addresses this operational hurdle by leveraging its footprint of more than 280 data centers across 77 metros, 230 cloud on-ramps, and a network of over 10,500 interconnected enterprises. With eight of the top ten AI model providers and nine of the top ten AI clouds already operating within Equinix facilities, the platform functions as a neutral aggregation layer for distributed AI pipelines.

Architecture and Stack Integration

The Equinix Inference Exchange stack is structured into three distinct layers across compute, software, and physical infrastructure:

Equinix manages the base physical and interconnect foundation. This includes power provisioning, advanced liquid cooling capabilities, day-two facility operations, and private interconnectivity to networks, public clouds, and SaaS platforms via Equinix Fabric.

NVIDIA provides the compute and reference validation layer, supplying enterprise AI infrastructure engineered for high token throughput and reduced inference costs per query.

Together AI provides the inference-serving software layer that supports over 200 open-source models. The platform supports both shared multitenant environments for general compute efficiency and dedicated single-tenant topologies for workloads that require isolated capacity.

Together AI logo; the company supplies the inference-serving software layer of Equinix Inference Exchange

By routing traffic over Equinix Fabric, the platform reduces time-to-first-token (TTFT) metrics for regional users while establishing private peering paths between on-premises storage and model endpoints.

Enterprise Deployment Targets and Availability

The platform targets three primary deployment topologies. For metro-edge inference, organizations can run models in local metro hubs, delivering low-latency API responses while maintaining centralized management and private interconnects. For open-source model migration, the architecture provides a low-friction path for shifting from closed, proprietary APIs to open-source foundation models, reducing vendor lock-in while leveraging existing network fabrics. For sovereign AI and compliance mandates, the distributed layout enables enterprises in regulated sectors to keep inference execution and data processing within specific geographic boundaries, helping meet strict data residency and governance requirements.

Equinix Fabric One diagram showing a connectivity request declared through an MCP client or API and composed into a network spanning Chicago and Ashburn metros

Leadership from Equinix, NVIDIA, and Together AI noted that the integration establishes a vendor-neutral, high-throughput inference fabric that balances compute performance with operational agility.

Equinix Inference Exchange is scheduled to be generally available to enterprises in the first quarter of 2027.

The post Equinix Inference Exchange Brings NVIDIA Compute and 200+ Open Models Closer to Enterprise Data appeared first on StorageReview.com.

On the Ground at VMware Explore 2026: How Lenovo Is Partnering with VMware for Turnkey AI and Taming the Memory Crunch

1 September 2026 at 18:28
Open Lenovo server chassis showing GPU bays and internal cabling at VMware Explore 2026 Open Lenovo server chassis showing GPU bays and internal cabling at VMware Explore 2026

At VMware Explore 2026, we had the chance to sit down with Stuart McRae, Executive Director of Lenovo’s Enterprise Storage, Software, and Solutions Offering Group.

Lenovo booth at VMware Explore 2026 with AMD, Intel, and NVIDIA partner signage

Rather than focusing on a specific product or technology, as we often do, this conversation took a broader view. We covered everything from how AI is reshaping buyer options and the ongoing RAM shortage to a spirited debate over whether Texas or North Carolina has the better BBQ. Below are our takeaways from that discussion.

Lenovo server hardware on display at VMware Explore 2026 with two attendees beside the table

The High Stakes Return to On-Prem

In the early years of the data center, system architects took pride in the “build-it-yourself” ethos. Each infrastructure build had its own local flavor, carefully assembled from preferred “Lego pieces” of compute, storage, and networking. We were the architects of bespoke data centers.

However, the enterprise is now undergoing a fundamental shift. The high-cost gamble of generative AI, combined with the escalating and often unpredictable cost of public cloud, has accelerated the return to on-premises infrastructure. Today’s environment demands immediate, predictable results rather than architectural tinkering.

In our conversation with McRae, it became clear that the mission has changed: We are moving away from piecemeal assembly and toward integrated, turnkey “factories,” where the infrastructure becomes invisible, reliable, and supportable.

The “RAM Apocalypse” is a Catalyst for Innovation

The industry is navigating what many have called the “RAM apocalypse.” A volatile mix of DRAM and flash supply constraints has driven commodity prices sharply higher, turning memory into a growing bottleneck for data center expansion. For memory architecture, this is a now-or-never moment for innovation.

Open Lenovo server chassis showing GPU bays and internal cabling at VMware Explore 2026

Memory tiering is not new, but previous approaches were often technically viable yet operationally fragile, depending on complex configurations at the operating system or application layer. That made them difficult to deploy consistently across a diverse enterprise environment. What is changing now, led by Lenovo and VMware, is the move to bring this intelligence into the hypervisor layer within VMware Cloud Foundation (VCF). By tiering data down to NVMe drives at the software-defined layer, organizations can begin to decouple performance from DRAM cost.

  • Broad Adoptability: Because this happens at the hypervisor level, it requires no “rip and replace” of existing applications or OS rewrites.
  • Economic Impact: Organizations are realizing up to 45% savings on DRAM costs by optimized tiering.

The old proverb asks, ‘ When’s the best time to plant a tree? 20 years ago. When’s the next best time? Tomorrow. Well, the best time to buy a server would have been a year ago… the next best time is today,” McRae told us.

Moving from “Lego Pieces” to Turnkey AI Factories

We are witnessing the inevitable collision of enterprise density and the necessity for “just give me something that works.” The psychological shift in IT is profound: the joy of tinkering has been replaced by the necessity of the “appliance approach.” For organizations looking to stand up AI capabilities, the complexity of sizing and validation is too high to manage on their own.

Row of open Lenovo server chassis and trays on a display table

This has necessitated the rise of the Hybrid AI Factory. By integrating Lenovo’s hardware with the VCF Private AI stack, Lenovo provides a pre-validated environment in which SSDs, network cards, and firmware revisions are guaranteed to work together.

But this is more than hardware; it is about life-cycle management and peace of mind. In an era where security and resilience are non-negotiable, a validated stack ensures that the “Lego pieces” not only fit together on day one but also stay together through every subsequent update and patch. We are moving from being component builders to supervisors of a unified, automated production line.

The Enterprise is Finally Embracing Liquid Cooling

For decades, liquid cooling was a niche luxury, the province of High-Performance Computing (HPC) and specialized research labs. Today, we are witnessing a sea change as water cooling becomes a mainstream requirement in the enterprise. This shift is driven by the immutable laws of thermal physics: as we pack high-density CPUs and power-hungry GPUs into the same rack, traditional air cooling has reached its limit.

Liquid-cooled Lenovo server tray with copper cold plates and quick-disconnect hoses

To manage this thermal load, Lenovo has three distinct architectures for the modern data center:

  • Warm Water Cooling (Lenovo Neptune): A sophisticated direct-to-node approach that manages heat with incredible efficiency, often eliminating the need for power-intensive chillers.
  • Closed-Loop Cooling: A vital “middle ground” for facilities without existing water infrastructure, utilizing liquid-cooled heat sinks that circulate entirely within the chassis.
  • Air Cooling: Still the baseline for standard workloads but increasingly relegated to low-density zones as liquid takes over the heavy lifting.

The 10-Day Delivery: Logistics as a Competitive Feature

In a volatile market, logistics and purchasing are no longer boring back-office functions; Lenovo has turned them into a competitive advantage. Speed of deployment has become a primary metric of infrastructure efficiency. Through programs like Lenovo’s “Top Choice Express” and “AI Express,” Lenovo is offering 10-day delivery cycles for VCF-ready and AI-optimized servers.

Perhaps more importantly, Lenovo has a 30-day price guarantee. In a market where quotes often expire within 72 hours due to component volatility, this stability is a strategic lifeline, particularly for public-sector entities bound by tight and rigid budgeting processes.

The hunger for this hardware-centric stability was palpable at recent industry events, even at software-focused conferences like VMware Explore, where sessions on infrastructure cost control filled up early.

VMware Explore welcome reception schedule on a large conference display

We’re the only server-focused vendor here at Explore. We are seeing massive traffic at our booth and intense interest in hardware… usually at software conferences, those are not the sessions that fill up fast,” McRae said.

AI as the “Quantum Step” in Systems Administration

One of the most transformative shifts we at Storage Review anticipate is not AI as a workload, but AI as the interface for the hypervisor itself. We covered VMware’s new AI Assistant in our write-up on VMware announcements made at Explore this year.

VMware Private AI Cloud platform slide covering agent platform, security, model as a service, and AI infrastructure

We are moving toward a vision of “Natural Language administration,” where the friction between an administrator and the hardware is virtually eliminated. We got a taste of this when we reviewed Aipex, a product that does this for Windows, Linux, and macOS systems, and we were excited when we saw that VMware will be deploying an AI Assistant in VCF, and even better, VMware builds on this AI integration by supporting 3rd-party Management Packs for it.

It’s easy to imagine a future where a kernel fault or a system crash no longer triggers a weekend of manual log analysis. Instead, an administrator asks the system, “Why did this machine crash?” and the AI assistant, deeply integrated into the VCF stack, identifies a specific driver fault, locates the patch, and requests permission to deploy it.

VMware AI Assistant for VCF slide listing conversational diagnostics, health visibility, and management packs builder

This represents a quantum leap forward, harkening back to the early days of virtualization, when provisioning required a month-long dance among storage, network, and server teams. By reducing these complex workflows to simple dialogue analyzed and implemented by AI, we finally remove a barrier that has plagued IT for decades.

“It’s the way it should be,” as McRae put it.

Conclusion: The Blueprint for the Next Three Years

To wrap up our conversation with McRae, we hit him with an important, yet impossible-to-answer question: What will the data center look like in the next three years?

He said that the blueprint for the private cloud over the next three years is toward total integration. We can expect a shift in which VCF and vSAN are the default standards, and GPU resources are integrated directly into clusters as peers rather than as specialized outliers. The “integration-first” philosophy will dominate as enterprises prioritize data sovereignty and operational efficiency over the public cloud’s complexity.

Ultimately, AI is the tool that will finally make our infrastructure “invisible.” By automating the manual labor of troubleshooting, remediation, and provisioning, we are reaching the point where technology works for the business rather than the business working for technology. The goal is no longer to build the best “Lego” tower; it is to run the most efficient factory.

Finally, when pushed to choose a BBQ style, McRae diplomatically said he liked the vinegary tang of North Carolina barbecue just as much as the slow-smoked brisket of Houston and Austin. He had no comment on Cincinnati BBQ.

The post On the Ground at VMware Explore 2026: How Lenovo Is Partnering with VMware for Turnkey AI and Taming the Memory Crunch appeared first on StorageReview.com.

VAST Data CrowdStrike Integration Goes Live: Native Falcon Sensor Now, Next-Gen SIEM and AIDR in Preview

1 September 2026 at 15:37
VAST Data Platform diagram showing objects, tables, files, triggers, and functions spanning edge, core, and cloud VAST Data Platform diagram showing objects, tables, files, triggers, and functions spanning edge, core, and cloud

VAST Data and CrowdStrike are turning the AI security partnership they announced at VAST Forward in February into a shipping product, detailing a multi-layered integration that embeds enterprise-grade cybersecurity directly into AI storage infrastructure, data pipelines, and production AI workloads. By combining the VAST AI Operating System with the CrowdStrike Falcon platform, the collaboration addresses the security challenges in modern AI environments, where proprietary data flows across complex pipelines, models, and autonomous agents.

Securing AI Infrastructure and Ingestion Pipelines

The partnership begins at the hardware and operating system level, where VAST now natively supports the CrowdStrike Falcon sensor. This integration brings Falcon endpoint protection and workload visibility directly to the infrastructure nodes hosting mission-critical AI datasets.

CrowdStrike Falcon console image assessment dashboard showing vulnerability counts, severity breakdown, and top CVEs

Beyond core infrastructure protection, VAST audit telemetry will feed directly into CrowdStrike Falcon Next-Gen SIEM. This telemetry pipeline gives security operations teams deep visibility into how internal users, automated services, and compute nodes interact with storage assets. Correlating low-level storage access patterns with CrowdStrike threat intelligence enables rapid detection of anomalous data access, unauthorized exfiltration attempts, and compromised service accounts across the AI environment.

VAST Data Platform diagram showing objects, tables, files, triggers, and functions spanning edge, core, and cloud

At the data pipeline layer, CrowdStrike Falcon AI Detection and Response (AIDR) integrates directly with the VAST InsightEngine, building on the NVIDIA AI Data Platform reference architecture. As InsightEngine processes, enriches, and vectorizes unstructured data for model training and retrieval-augmented generation (RAG) knowledge bases, Falcon AIDR inspects the data in flight. This allows organizations to classify sensitive information, including personally identifiable information (PII), before it enters downstream AI workflows.

Runtime AI Protection and AI Factory Validation

The integrated platform also extends protection to real-time model interaction. As users and AI agents query models and knowledge bases, CrowdStrike Falcon AIDR detects and flags adversarial threats, such as prompt injection and model jailbreak attempts. This dual-layer approach secures both the foundational training data and the operational inferencing surface.

NVIDIA AI data platform illustration with stylized server racks and data pipelines

To validate the integration in high-throughput enterprise deployments, VAST and CrowdStrike will demonstrate the joint architecture within an NVIDIA-powered AI factory environment. The setup showcases CrowdStrike real-time protection running inline with high-performance VAST data pipelines accelerating GPU-intensive workloads.

According to leadership from both companies, securing AI requires moving beyond basic perimeter defense. Integrating security directly into data movement, ingest engines, and model access ensures that security controls are foundational rather than reactive additions to modern AI architecture.

Availability

Native CrowdStrike Falcon sensor support for the VAST platform is certified and available immediately. Integrations that connect VAST telemetry to CrowdStrike Falcon Next-Gen SIEM and Falcon AIDR are currently available in private preview.

The post VAST Data CrowdStrike Integration Goes Live: Native Falcon Sensor Now, Next-Gen SIEM and AIDR in Preview appeared first on StorageReview.com.

MLPerf Storage v3.0: 877 GiB/s Checkpoints, a Cloud First, and a Leaderboard Turned Over

1 September 2026 at 15:00
Bar chart of MLPerf Storage v3.0 checkpoint write scaling for Everpure FlashBlade//EXA, from 327.6 GiB/s at 10 data nodes to 877.5 GiB/s at 30 Bar chart of MLPerf Storage v3.0 checkpoint write scaling for Everpure FlashBlade//EXA, from 327.6 GiB/s at 10 data nodes to 877.5 GiB/s at 30

MLCommons published the MLPerf Storage v3.0 results today, and the round changes not only what the benchmark is but also who leads it. For the first time, the only audited AI storage benchmark covers the full pipeline: training throughput, checkpointing, and two new inference workloads, a vector database test and a KV cache test. Nineteen organizations submitted 143 results, eleven of them first-time submitters, including Microsoft Azure, NVIDIA, Everpure, and a wave of AI storage startups most readers will not recognize. Just as notable is that DDN, Huawei, Hammerspace, and Lightbits, the names that defined the v2.0 leaderboard, did not submit this round.

One comparability note up front. v3.0 moved its simulated accelerators from H100 to B200, which raises the per-accelerator bandwidth bar, so v3.0 numbers cannot be compared against v2.0 results. These are MLCommons-audited vendor submissions; StorageReview did not conduct or independently verify this testing.

Checkpointing Is the Headline Event

MLCommons added checkpointing as a first-class workload because it has become a verified bottleneck: frontier-scale training jobs write enormous amounts of state at regular intervals, and every second spent flushing a checkpoint is a second of idle accelerator time. The largest number in the round belongs here. Everpure (formerly Pure Storage), in its first MLPerf Storage appearance, submitted FlashBlade//EXA results that scale almost linearly with data node count: 327.6 GiB/s of checkpoint write at 10 data nodes, 484.1 GiB/s at 15, 655.6 GiB/s at 20, and 877.5 GiB/s write with 833.0 GiB/s read at 30 data nodes. Everpure has spent the past year making large throughput claims for //EXA; this is the first time an audited result backs the direction of those claims.

Bar chart of MLPerf Storage v3.0 checkpoint write scaling for Everpure FlashBlade//EXA, from 327.6 GiB/s at 10 data nodes to 877.5 GiB/s at 30

The second-largest checkpoint number is arguably the bigger industry story. Azure Managed Lustre, the first hyperscale cloud storage service ever submitted to the benchmark, posted 642.2 GiB/s of checkpoint write from a 4,096 TiB deployment with 128 clients. A managed cloud file service putting up numbers in the same conversation as purpose-built AI storage arrays would have been hard to picture two rounds ago.

Checkpointing (Top Submissions) System Write B/W (GiB/s)
Everpure FlashBlade//EXA, 30 data nodes 877.5
Azure Managed Lustre, 4,096TiB, 128 clients 642.2
YanRongTech F9000X, 8 clients 313.7
Suzhou Zishan Longlin Multiple configurations 313.7
TuringData F9200, 8 clients 307.1
UBIX UbiPower18000, 3 storage nodes 293.8

Training: Names You Likely Don’t Know

The 3D U-Net training leaderboard belongs to newcomers. YanRongTech’s F9000X sustained 543.9 GiB/s of read bandwidth feeding 99 simulated B200 accelerators, the highest training throughput of the round. TuringData, a Singapore AI infrastructure company making its first submission, landed 4 GiB/s behind at 541.5 GiB/s while feeding 96 simulated B200s from just three storage nodes, and the same three-node F9200 cluster also submitted Checkpoint-70B at 539.9 GiB/s read and 307.1 GiB/s write. UBIX’s UbiPower18000, three storage nodes with 16 x 15.36TB NVMe drives and quad NDR400 networking, hit 454.1 GiB/s. HPE was the only established enterprise array vendor to submit training results, with the K3000 at 345.8 GiB/s.

3D U-Net Training (Top Submissions) Read B/W (GiB/s) Simulated B200 Accelerators
YanRongTech F9000X 543.9 99
TuringData F9200 541.5 96
UBIX UbiPower18000 454.1 80
Suzhou Zishan Longlin 433.4 80
FarmGPU 406.7 75
Azure Managed Lustre 379.1 70

Inference Joins the Benchmark

The two new workloads acknowledge where storage demand is actually growing. The KV cache test measures how quickly a storage system can page the attention state in and out for long-context LLM serving, a pattern we examined in depth in our Dell and Solidigm KV cache offload work. Everpure again posted the round’s largest number, 1,623.4 GiB/s of KV cache read on a 51-host FlashBlade//EXA configuration, with UBIX (443.7 GiB/s) and FarmGPU (443.6 GiB/s) leading the standard-scale submissions. On the vector database side, TTA’s Seahorse system topped the query throughput chart at 57,620 queries per second.

The other structural change is protocol. v3.0 added S3 object storage support, and roughly one-sixth of submissions used it. NVIDIA accounted for most of that with 20 AIStore submissions spanning bare-metal configurations on OCI, AWS, and GCP, topping out at 133.3 GiB/s of checkpoint write. The absolute numbers trail the parallel file systems, but object storage running training and checkpoint workloads within the benchmark marks territory that POSIX file systems had to themselves a round ago.

MLCommons also began tracking power and rack-space efficiency as first-class metrics this round.

Everpure FlashBlade//EXA

Everpure FlashBlade//EXA system held the lead across the 405B and 1.25T parameter model checkpointing categories as well as Key-Value (KV) cache retrieval benchmarks, highlighting the platform’s ability to maintain high throughput during data-intensive artificial intelligence training and inference workloads.

Everpure FlashBlade//EXA architecture diagram showing a compute cluster connected to a metadata core and data nodes over RDMA

Everpure’s submitted configuration paired 30 FlashBlade//EXA blades (120 DirectFlash Modules handling metadata) with 30 Linux/NVMe data nodes for roughly 866 TB of usable capacity in a single file system, splitting metadata over pNFS/TCP and bulk data over NFSv3 on RDMA. In the 1.25T parameter checkpointing simulation with 1,024 client accelerators, the 30-node configuration sustained 877.52 GiB/s of write bandwidth over a 17.74-second operation, alongside 588.28 GiB/s of read bandwidth completed in 28.99 seconds. On the KV cache inference test it reported 85,736 tokens/sec on a storage-only Llama 3.1 8B run, 67,642 tokens/sec on an 8B run combining storage and system memory, and 33,403 tokens/sec on a 70B storage-only run.

What the No-Shows Mean

A benchmark round is evidence of two things: what the submitters can do and what the absentees chose not to show. VAST Data, WEKA, and DDN, the vendors winning the largest neocloud and AI factory deployments, all sat v3.0 out, as did v2.0 submitters Huawei, Hammerspace, and Lightbits. Vendors skip rounds for mundane reasons, engineering cycles, release timing, and benchmark politics; a skipped round is not evidence that a system is slow. But it does mean the audited record and the deployment record now point to different vendor lists, and buyers have to weigh both. Our Best Storage Arrays page tracks exactly that split, labeling what is audited and what is deployment evidence, and has been updated with the full v3.0 results.

The v3.0 results are public today on the MLCommons storage benchmark page.

The post MLPerf Storage v3.0: 877 GiB/s Checkpoints, a Cloud First, and a Leaderboard Turned Over appeared first on StorageReview.com.

NVIDIA MediaTek Partnership Deepens With $3.5 Billion Investment, NVLink Fusion XPUs, and RTX Spark PCs

31 August 2026 at 20:58
NVIDIA render of a custom XPU package, featured image for the NVIDIA MediaTek partnership news NVIDIA render of a custom XPU package, featured image for the NVIDIA MediaTek partnership news

NVIDIA and MediaTek announced a significant expansion of their strategic collaboration, covering enterprise AI infrastructure, edge computing, client PCs, and automotive platforms. Backing the initiative, NVIDIA has invested $3.5 billion in MediaTek-issued convertible bonds. The technical partnership centers on MediaTek integrating the NVIDIA NVLink Fusion platform to provide hyperscalers, cloud service providers, and frontier model developers with a pre-validated framework for designing custom accelerators (XPUs) that can drop directly into NVIDIA rack-scale AI environments.

MediaTek and NVIDIA logos side by side

Leadership from both companies framed the expanded alliance around convergence across compute domains. NVIDIA founder and CEO Jensen Huang pointed to the requirement for accelerated computing to scale from massive data centers down to personal systems and software-defined vehicles, citing MediaTek’s expertise in system-on-chip design, connectivity, and performance per watt. MediaTek Vice Chairman and CEO Rick Tsai noted that combining MediaTek’s custom silicon design capabilities with NVIDIA’s accelerated compute stack and software ecosystem will accelerate innovation for customers across cloud AI infrastructure, local AI computing, and automotive.

Custom Silicon Design on the NVLink Fusion Platform

Developing semi-custom AI accelerators for rack-scale deployments introduces significant engineering hurdles around high-speed SerDes, interconnect topologies, High Bandwidth Memory (HBM) integration, multi-die advanced packaging, and system-level thermal and mechanical qualification. Rather than requiring customers to engineer these underlying interconnects and packaging layers from scratch, MediaTek is adopting NVLink Fusion as a modular design blueprint.

NVIDIA render of an NVLink Fusion board with two custom XPU packages and a central interconnect chip

The NVLink Fusion architecture provides a standardized, prequalified multi-die foundation that includes several primary subsystem technologies. The NVIDIA NVLink Fusion chiplet establishes direct links between custom XPUs and the broader NVLink scale-up fabric using NVIDIA photonics or electrical interconnects. To support coherent, low-latency inter-chip communications, NVIDIA NVLink-C2C provides high-bandwidth, energy-efficient die-to-die connectivity between customer XPUs, NVIDIA Rosa CPUs, and other compatible processing units. On the memory side, NVIDIA NVHBM integrates customized memory interfaces to expand available bandwidth and improve power efficiency while reserving more die area for active compute engines.

NVIDIA render of a custom XPU package with NVHBM memory stacks flanking the compute dies

Through this framework, enterprise customers can bring proprietary processing architectures to MediaTek to tailor specific performance profiles, memory densities, and power targets to their target workloads. In turn, MediaTek manages the complex packaging, physical implementation, and supply-chain logistics required to manufacture multi-die packages. These custom XPUs can then interface with NVIDIA MGX modular rack-scale architectures and scale-out networking fabrics, significantly reducing development risk and shortening time-to-market for specialized AI factories.

Scaling Client Compute and Automotive Platforms

Beyond data center infrastructure, the partnership addresses local client compute for generative and agentic AI workloads. The companies previously collaborated on the GB10 Grace Blackwell Superchip, which links an NVIDIA Blackwell GPU with an ARM-based Grace CPU via NVLink-C2C to power the NVIDIA DGX Spark platform for developer and edge environments. That work is extending into client systems through NVIDIA RTX Spark, which the companies say will power the next generation of consumer PCs built for the AI era, alongside the GB10-based DGX Spark that pairs a Blackwell GPU and Grace CPU over NVLink-C2C.

Bare mainboard of the NVIDIA DGX Spark on the StorageReview bench from our teardown

In the automotive sector, the collaboration spans multiple product generations of AI-powered, software-defined vehicles. MediaTek’s Dimensity Auto platforms integrate NVIDIA AI and RTX graphics for intelligent vehicle cockpits and can operate alongside NVIDIA DRIVE AGX, with both companies building on that foundation toward a scalable architecture for increasingly intelligent, AI-defined vehicles.

The post NVIDIA MediaTek Partnership Deepens With $3.5 Billion Investment, NVLink Fusion XPUs, and RTX Spark PCs appeared first on StorageReview.com.

❌
❌