Reading view

There are new articles available, click to refresh the page.

Anthropic And OpenAI Can Sidestep AGI-Related Scaremongering By Opening Up The Traces Of Their Most Advanced Models, But Their Real Goal For Pacing AI Might Have To Do With Managing Chronic Power Shortages

A group of people in a control room examining maps and charts on screens with the text 'US Power Grid Status: Critical Energy Shortage' and 'Increased grid strain attributed to AI data center energy consumption (OpenAI, Anthropic).'

When you have Anthropic's Jacob Coxon literally boasting about the fact that he collaborated with others to make his doom-and-gloom-style resignation post go viral, while continuing to tout the supposedly noble intentions of his former employer in one podcast after another, one is hard-pressed not to adopt a predominantly cynical bent on the coordinated push towards pacing the progress on AI, made all the more incredulous by the fact that OpenAI echoed that call within hours when Dario Amodei and Sam Altman rarely see eye to eye on any topic under the sun. To a skeptic, an assumption of obfuscation […]

Read full article at https://wccftech.com/anthropic-and-openai-can-sidestep-agi-related-scaremongering-by-opening-up-the-traces-of-their-most-advanced-models-but-their-real-goal-for-pacing-ai-might-have-to-do-with-managing-chronic-power-shor/

Chinese Chip Firm Empyrean Says It Slashed Circuit Design Time By 75% Using Agentic AI Amidst Race To Agentic AI Semiconductor Design

Chinese electronic design automation (EDA) firm Empyrean Technology is using agentic artificial intelligence to accelerate its chip design process. The firm's chairman, Liu Weiping, shared the details at a conference in Shenzen, China, last week. His remarks come as Chinese firms continue to manage around the constraints levied on them through chip equipment sanctions. According to Weiping, his firm has managed to cut down design times for some tasks by as much as 75% through using agentic AI, and his comments came days before China's Ministry of Industry and Information Technology issued a new development plan for the information and […]

Read full article at https://wccftech.com/chinese-chip-firm-empyrean-slashes-circuit-design-time-by-75-using-agentic-ai-amidst-race-to-agentic-ai-semiconductor-design/

NVIDIA Intros RTX PRO 5500 Blackwell Graphics Card, Features Same Core Count As The 5090 But With 2.6x The VRAM

NVIDIA RTX Pro 5500 GPU on a textured black background.

NVIDIA's latest addition to the Blackwell PRO family is the RTX PRO 5500 Workstation Edition graphics card with 84 GB of memory. NVIDIA RTX PRO 5500 Blackwell Workstation Edition Graphics Card Packs 21,760 Cores & 84 GB Memory The NVIDIA RTX PRO "Blackwell" family is one of the most extensive PRO lineups we have seen so far. It features various SKUs, and the latest one is the RTX PRO 5500 Blackwell, which targets the Pro and Workstation segment. The NVIDIA RTX PRO 5500 Blackwell Workstation Edition accelerates agentic AI, physical simulation, and graphics workloads to teams that need workstation power […]

Read full article at https://wccftech.com/nvidia-intros-rtx-pro-5500-blackwell-graphics-card-21760-cores-84-gb-memory/

Qwen3.8-27B Was Allegedly Used To Develop A First-Person Zombie Shooter In About 5 Hours; It’s Far From a Visual Masterpiece, But Just Shows How Far AI Can Go

Qwen3.8-27B was used to develop a first-person zombie shooter

A stage has been reached where you can seemingly develop games using local AI models that don’t require an internet connection or employ any token usage. Based on one individual’s claim, Qwen3.8-27B was leveraged to make a first-person zombie shooter in around five hours, and while it doesn’t even come near those visually breathtaking visuals of AAA titles, this is an excellent example of how far AI has evolved. An older RTX 3090 gaming PC was used to develop the game from scratch, concluding that serious hardware is still necessary to venture into this category The exact version of Qwen3.8-27B […]

Read full article at https://wccftech.com/qwen-38-27b-local-llm-first-person-zombie-shooter-game-5-hours/

Fujitsu MONAKA Server Brings 2nm 144-Core CPUs to Air-Cooled AI Inference, On Sale in November

Fujitsu MONAKA Server Fujitsu MONAKA Server

Fujitsu is bringing its 2nm FUJITSU-MONAKA processor to AI infrastructure with a new server family designed to run AI inference in air-cooled data centers without requiring specialized liquid cooling. The MONAKA Server is designed, developed, and manufactured in Japan, with component and manufacturing traceability for sovereign AI deployments.

FUJITSU-MONAKA CPU package render with the Fujitsu logo, the 2nm 3D-stacked processor at the heart of the Fujitsu MONAKA Server

The first MONAKA Servers will come in 1U and 2U configurations for AI inference and AI agents, with a separate 2U multi-node system planned for academic and HPC environments. Fujitsu is also making the MONAKA processor available separately for cloud and data center operators and other server vendors.

MONAKA runs at up to 3.8GHz and supports memory transfer speeds up to 8800MT/s, with dedicated matrix instructions and SVE2 vector processing accelerating AI inference directly on the CPU. Fujitsu rates MONAKA at twice the AI inference throughput of other CPUs, although it hasn’t identified the processors used for that comparison or provided benchmark results.

MONAKA Brings AI Inference to Air-Cooled Servers

The 1U MONAKA Server supports air cooling at ambient temperatures up to 40 degrees Celsius, while its liquid-cooled configuration supports water temperatures up to 45 degrees Celsius.

Fujitsu rates its server cooling technology for up to an 80% reduction in cooling power consumption, although specific system configurations and test conditions weren’t provided. The air-cooled configuration is particularly interesting for data centers that want to add AI inference capacity without changing their existing cooling infrastructure.

Fujitsu MONAKA Server 2U chassis with the lid off, showing two FUJITSU-MONAKA CPUs under their heatsinks, the DIMM banks around them, and the front drive bays

The processor itself uses a 3D-stacked design that combines CPU cores manufactured on a 2nm process with cache and I/O produced at 5nm. Fujitsu pairs the design with its ultra-low-voltage operation technology, which is where it says the power efficiency comes from.

Confidential computing is implemented at the hardware level through Arm CCA, encrypting applications and data while they are running in memory, including workloads operating in multi-tenant cloud environments.

MONAKA is also being developed for integration with accelerated platforms through Fujitsu’s work with NVIDIA to connect MONAKA CPUs and NVIDIA GPUs through NVLink Fusion.

MONAKA Server Keeps Development and Manufacturing in Japan

Sovereign infrastructure is the other major focus for MONAKA Server, with Fujitsu keeping the design, development, and manufacturing of the systems in Japan. Production will take place at the company’s Kasashima Plant, where component origins and manufacturing histories can be tracked through the production process.

That traceability takes the sovereign AI pitch down to the physical hardware, past the usual question of where data is stored or processed. Fujitsu is pairing the domestically produced hardware with its own AI software and management technologies for organizations that need greater control over their infrastructure, data, and supply chain.

The initial server lineup includes a 2U system with two MONAKA processors and 1U configurations with either one or two CPUs. Both air- and water-cooled deployments are supported, with the systems targeting AI inference and AI agent workloads.

Specification 2U Rackmount Model 1U Rackmount Model
Features All-in-one CPU/GPU/NW model optimized for existing DC/Edge environments Scalable model with flexible configuration and expansion according to environmental changes
Applications Digital Twin, Physical AI, Agentic AI, AI Inference Optimized for facility environments; flexible expansion from small to large scale
CPU Type / Quantity FUJITSU-MONAKA
2 CPUs
FUJITSU-MONAKA
1 to 2 CPUs
Base Frequency / Cores 2.1GHz × 144 cores 2.1GHz × 144 cores (air-cooled)
2.9GHz × 144 cores (liquid-cooled)
Memory Type / Slots RDIMM
24
RDIMM
12 (1-CPU configuration) / 24 (2-CPU configuration)
Storage Type / Slots E3.S SSD × 4
M.2 SSD × 2
E3.S SSD × 8
M.2 SSD × 2
Expansion Slots PCIe Gen6 (GPU support available) PCIe Gen6
Chassis Size 2U height 1U height
Cooling Method Air-cooled Air-cooled / Liquid-cooled

 

The 1U MONAKA Server also uses CDI/CXL technology to pool memory and accelerators, allowing those resources to be allocated across systems.

For academic and HPC environments, Fujitsu is preparing a 2U multi-node MONAKA Server with four nodes and two CPUs per node. The configuration is intended for computational workloads including surrogate models and fluid analysis.

Specification 2U Multi-node Model (4 nodes per chassis)
Features Multi-node model maximizing processing power within limited power and space
Applications AI data centers, large-scale simulations
CPU Type / Quantity FUJITSU-MONAKA
2 CPUs per node / 8 CPUs per chassis
Base Frequency / Cores 2.9GHz × 144 cores
Memory Type / Slots RDIMM
24 per node / 96 per chassis
Storage Type / Slots E1.S SSD × 2 per node
M.2 SSD × 2 per node
Expansion Slots PCIe Gen6
Chassis Size 2U height
Cooling Method Liquid-cooled

Fujitsu Connects MONAKA With Its AI Software

Fujitsu is extending the sovereign infrastructure approach into software through the Fujitsu Kozuchi AI platform, Takane enterprise generative AI, and industry-specific AI models. The combination brings together the processor, server hardware, AI platform, and generative AI software while supporting domestic data management, governance, traceability, security, and operational autonomy.

Future MONAKA hardware will include higher-density systems for AI data centers and rack-scale servers with autonomous operation features. Fujitsu is also developing robotic maintenance technology for future rack-scale systems, extending automation into physical server operations and maintenance.

Fujitsu’s math on the twice-the-throughput claim is that customers need half the servers and half the power for the same inference load, which is the figure the air-cooled pitch rests on. With no named comparison CPU or published benchmarks behind it, that’s the number to test when systems ship.

FUJITSU-MONAKA Availability

Standalone FUJITSU-MONAKA processors are scheduled to go on sale in November 2026, with availability planned for cloud and data center operators and server vendors in Japan, the US, APAC, and other markets during the fourth quarter of Fujitsu’s fiscal 2026.

Fujitsu also plans to begin sales of its 1U and 2U MONAKA Servers in Japan and Europe in November 2026 to data center operators, enterprises, academic and HPC buyers, and the defense sector. Select customers in the financial, telecommunications, and manufacturing sectors are expected to receive systems during the second half of fiscal 2026, with broader shipments in Japan and Europe scheduled to begin sequentially in April 2027.

FUJITSU-MONAKA

The post Fujitsu MONAKA Server Brings 2nm 144-Core CPUs to Air-Cooled AI Inference, On Sale in November appeared first on StorageReview.com.

NASA-IBM Lunar Foundation Model Goes Open Source With a 2M-Tile Dataset and 22% Lower Ice-Mapping Error

IBM and NASA have released the NASA-IBM Lunar Foundation Model as open source, one of the first publicly available foundation models built for scientific study of the Moon. The weights, a technical report, and the machine-learning-ready dataset it was trained on are up on Hugging Face under the Prithvi family, which already covers Earth observation, weather, and heliophysics. The pitch is that decades of multi-instrument lunar data have outgrown hand-sifted maps and narrow, task-specific models, and a single pretrained backbone can be adapted to crater mapping, volcanic-feature detection, and ice prospecting without starting over each time. In IBM and NASA’s own benchmarks, the model cut root-mean-square error in flagging likely ice deposits by up to 22% against a SwinV2-B baseline.

NASA-IBM Lunar Foundation Model crater detection output: two grayscale lunar surface tiles with blue bounding boxes drawn around detected craters

A TerraMind Backbone on 30 Layers From Nine Instruments

The model is built on a version of TerraMind, the Earth-observation model IBM developed with the European Space Agency, chosen for how it handles mixed data types and resolutions and learns cross-modal correlations that can fill in missing or noisy values. Fine-tuning for each lunar task uses low-rank adapters that keep 90% of the base weights frozen, which is what keeps the adaptation cost low. The training corpus is the part NASA says didn’t exist before: a unified, spatially aligned dataset of more than 30 layers from nine instruments across four missions, on the order of 2 million image tiles, with more than 1 million 1-meter camera images and close to 964,000 multispectral images at 100-meter resolution. Sources include imagery from NASA’s Lunar Reconnaissance Orbiter, gravity-field maps from the GRAIL mission at 20 kilometers per pixel, Lunar Prospector data, and complementary observations from JAXA’s SELENE/Kaguya. The benchmark collection, called SOMBench, ships alongside the model.

NASA-IBM Lunar Foundation Model ice prospectivity comparison: north and south polar maps beside label, ConvNeXt, and lunar foundation model prediction tiles on a blue-to-yellow prospectivity scale with a 10 km bar

Ice, Volcanic Patches, and Craters at Two Scales

The 22% figure comes from the ice task, where the model combines multimodal, multi-resolution observations to predict where ice may sit below the surface of permanently shadowed regions, the places a future Moon base would mine for water, oxygen, and rocket propellant. For volcanic history, it maps Irregular Mare Patches with 3% better coverage than SwinV2-B while training on imperfect labels, which IBM frames as comparable accuracy at lower fine-tuning cost. Crater detection splits by scale: at meter-scale resolution the model matches SwinV2-B, and at roughly 100-meter context scale it outperforms it by nearly 19% with half the training data. Michael Barker, the NASA lunar topography expert who co-led the project, said the payoff is in features whose ages are still contested: “The ages of these features remain a matter of great debate, so the more we can understand their distribution and properties, the better chance we have of resolving this mystery.”

“NASA has spent decades building an extraordinary scientific record of the Moon, but collecting data is only part of the job. We also have to make data easier for scientists to explore and use,” said Kevin Murphy, chief science data officer and acting chief data and AI officer at NASA Headquarters. Juan Bernabe-Moreno, director of IBM Research Europe, UK and Ireland, said the model “gives scientists a foundation to explore the Moon at scale, connecting observations across instruments, revealing patterns that are difficult to see in isolation, and providing an open platform the global research community can build on.”

Everything is available now under an open-source license, and IBM’s research blog walks through the three initial NASA use cases in more depth. It’s the same open-weights playbook IBM has run on the enterprise side with Granite 4.0, applied here to a science domain where the scarce resource is labeled data and the compute to fine-tune against it, and the immediate consumers are landing-site hazard analysis and resource prospecting for a sustained lunar presence.

NASA-IBM Lunar Foundation Model

The post NASA-IBM Lunar Foundation Model Goes Open Source With a 2M-Tile Dataset and 22% Lower Ice-Mapping Error appeared first on StorageReview.com.

Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts

Lightbits Labs, the company that invented NVMe over TCP, is moving into inference software with Inferra, a KV cache orchestration engine that makes its public debut tomorrow, September 15, at the AI Infra Summit in Santa Clara. The software virtualizes GPU memory across DRAM and NVMe storage tiers and turns the KV cache into a persistent data layer, so attention states for long-context and multi-session workloads don’t have to fit in HBM or be recomputed when they fall out. Lightbits claims up to 16x more concurrent inference sessions on existing GPUs, a greater than 100x latency improvement from prefetching attention states in place of recomputing them, and context windows up to 10 million tokens, and says the engine has already run in customer beta programs, including one with OVH.

Lightbits Inferra architecture diagram with the KV cache engine sitting between GPU HBM, DRAM, and storage tiers, showing prefetch, compression, quantization, scheduling, encrypted transfer, and tenant isolation functions

Prefetch Instead of Recompute

The bottleneck Inferra targets is the one we walked through in our KV cache offload to flash piece: as context lengths grow and sessions multiply, the KV cache outgrows GPU memory, and the serving stack either evicts it and recomputes the prefill later or caps how many sessions a GPU can hold. Both paths burn GPU cycles on work that has already been done. Inferra keeps the cache alive across memory and storage tiers, and a predictive prefetcher pulls attention states back toward the GPU ahead of when the model needs them. Lightbits says that approach drives down both Time-to-First-Token (TTFT) and Time Per Output Token (TPOT), and the 100x figure it quotes is the latency improvement against recomputing those states from scratch.

Lightbits Inferra block diagram showing its intelligent prefetcher, tiering manager, log-structured KV store, and security and isolation engines between LLM serving frameworks (vLLM, TensorRT, SGLang) and a standard NVMe SSD pool

Lightbits’ block diagram lays out the four pieces: an intelligent prefetcher, a tiering manager, a log-structured KV store, and security and isolation engines, sitting between the serving framework above and a standard NVMe SSD pool below. The company lists vLLM, TensorRT, and SGLang as supported serving frameworks and says the engine is GPU-and SSD-agnostic. The log-structured store is what makes a flash tier practical for a cache that changes constantly, since it appends updates sequentially and keeps small random writes off the drives, and we looked at how a 3 DWPD drive holds up in that role in our Solidigm D7-PS1030 review. On the multi-tenant side, Lightbits describes secure tenant isolation and intelligent cache management across shared inference infrastructure, with encrypted transfer between tiers, which is what lets a neocloud hand one KV pool to many customers with consistent SLAs.

Avigdor Willenz, co-founder and chairman of Lightbits Labs, framed the launch as a break from storage built for training pipelines. Inference “is a completely different paradigm,” he said, and “retrofitting legacy training and storage systems to solve the KV cache bottleneck simply doesn’t work. It was built for a different era. We engineered Inferra from the ground up specifically to solve the GPU efficiency problem, and it has already proven successful in customer beta programs.”

OVH, Solidigm, and the Booth 219 Demos

OVH is the named beta customer. “With Inferra’s intelligent KV cache tiering, we were able to demonstrate substantial GPU utilization gains, paving the way to providing our customers a more scalable and cost-effective foundation for their AI agent and RAG workloads,” said Yaniv Fdida, chief product and technology officer at OVH. Solidigm ran the engine in its AI Central Lab on D7-PS1010 drives, and Avi Shetty, the company’s VP of ecosystem, solutions, and market enablement, said the results “showcased how network-attached storage with intelligent software layer virtualization can break through the memory wall for large context, agentic AI.” That pairing of a network-attached NVMe pool with a software tier is Lightbits’ home turf; its LightOS block storage has been built on NVMe/TCP since 2019.

Ramesh Chettuvetty, senior vice president of AI product and business at Lightbits, called Inferra “the world’s first KV cache acceleration engine that eliminates idle GPUs and power context windows of up to 10M tokens on commodity hardware,” and said it “delivers instant payback and generates net positive savings from day one.” Those are vendor claims ahead of any independent testing, and the release doesn’t give pricing, packaging, or a general availability date; the engine is in customer beta today. Lightbits is running live demos of the session-density and latency results at Booth 219 during the summit, and has posted a Pod Efficiency Analyzer for teams that want to model the GPU savings against their own cluster before booking one.

Inferra by Lightbits

The post Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts appeared first on StorageReview.com.

❌