Normal view

There are new articles available, click to refresh the page.
Before yesterdayMain stream

NVIDIA Enters Full Production of Groq 3 LPX AI Inference Accelerator Chips, Supercharging Vera Rubin With The Fastest Token Generation Speeds Ever Recorded

24 August 2026 at 15:00

A close-up image of the AMD MI300X GPU, featuring its intricate circuit design against a black background.

NVIDIA's Groq 3 LPX is now in full production, offering big token generation speedups for Vera Rubin platforms in the Agentic AI space. NVIDIA Dials Up Vera Rubin NVL72 Token Generation Capabilities, Recording 3,400 TPS With Groq 3 LPX AI Inference Accelerators As part of its Hot Chips 2026 announcements, NVIDIA today announced that its Groq 3 LPX AI inference accelerator chip is in full production. This announcement follows the mass production announcements of Vera CPUs and Vera Rubin servers, marking the robust execution of NVIDIA's AI roadmap. The Groq 3 LPX racks serve as an extension to the NVIDIA […]

Read full article at https://wccftech.com/nvidia-groq-3-lpx-ai-inference-accelerator-full-production-supercharging-vera-rubin/

Cerebras CS-4 Generates In 1 Second What A GPU Rack Needs 30 Seconds For, Powered By 4-Trillion-Transistor WSE-3 Turbo

19 August 2026 at 20:45

A close-up of a hexagonal-patterned front panel with a Ceres logo and the text 'Introducing CS-4' beside it.

Cerebras, the creators of the wafer-scale engine chip, have released their latest CS-4 solution that packs its brand-new WSE-3 Turbo chip. Cerebras CS-4 Is Powered By Wafer-Scale Chips, Packing Up To 250 PFLOPS Per Wafer & Double The Bandwidth of 43.2 PB/s Using 44 GB SRAM Back in 2024, Cerebras unveiled its 3rd Gen Wafer Scale Engine, or WSE-3. This wafer acts as a single chip and offers lots of compute capabilities. Cerebras has been doing Wafer-Scale chips since their advent, and today, the company unveiled the next chapter in its wafer-scale journey. “In AI, speed is productivity,” said Andrew […]

Read full article at https://wccftech.com/cerebras-cs-4-generates-in-1-second-what-a-gpu-rack-needs-30-seconds-for-powered-by-4-trillion-transistor-wse-3-turbo/

Jim Keller Says Cerebras IPO Was Helpful As Tenstorrent Set To “Beat Them on Everything”, Confirms Meeting With Intel & Qualcomm CEOs “Hoping To Get A Big Deal”

27 June 2026 at 21:45

Samsung To Build Next-Gen Tenstorrent AI Chiplet Leveraging RISC-V Architecture 1

Jim Keller isn't bothered by Cerebras's recent IPO and says that he welcomes it, but Tenstorrent will still beat them on everything. Tenstorrent CEO, Jim Keller, Signals Deal With Intel or Qualcomm While Promising To Beat Cerebras "on everything" Tenstorrent recently introduced its latest BlackHole Galaxy server, a system with which it can disrupt the entire AI segment, with performance levels that crush the competition. We covered the announcement last month when the company demoed its Blackhole server undercutting a NVIDIA GB300 with up to five times better TCO. Keller Accepts The Challenge To Beat NVIDIA, Cerebras & Others At […]

Read full article at https://wccftech.com/jim-keller-cerebras-ipo-was-helpful-tenstorrent-to-beat-them-on-everything/

Tensordyne’s 3nm Napier AI Chip Promises 13x Higher Token Throughput Than Blackwell & Blazes Past Rubin With 1000 Tokens/s In Multi-Trillion Parameter Models

15 June 2026 at 17:40

A person wearing gloves holds a TensorDyne TDN AIP chip, with visible circuitry and labeled sections.

US-based AI company, Tensordyne, has announced the successful tape-out of its Napier chip, which it claims to demolish NVIDIA's Blackwell & Rubin chips with leading token throughput and efficiency. Tensordyne’s new Napier AI Chip arrives with one clear mission: to make NVIDIA’s Blackwell and Rubin chips look considerably less impressive The Napier chip will be the core component of the Tensordyne Napier TDN system, which is designed in collaboration with Broadcom and HPE Juniper Networks. The Napier platform has one goal: to unify AI through novel logarithmic AI math, a tightly integrated memory architecture, and a high-performance scale-up interconnect that […]

Read full article at https://wccftech.com/tensordyne-3nm-napier-ai-chip-13x-higher-token-throughput-blackwell-blazes-past-rubin/

AMD Finally Overtakes Intel in Q1 Data Center Revenue as Agentic AI Forces Hyperscalers to Hoard CPUs Over GPUs

8 May 2026 at 16:40

Global Client CPU Shipments Were Up 2.7% In Q4 2025 While Server CPUs Saw 6.5% Growth 1

AMD's EPYC CPU has boosted the company's data center revenue to new heights, surpassing Intel for the first time. As Agentic AI Rages On & CPUs Become The New Favorite, Both AMD & Intel Are Seeing Big Successes Due To Overwhelming Demand In the race for Data Center CPU revenue, there could be only one winner. While both Intel and AMD are enjoying their reignited demand thanks to Agentic AI, the Red Team has managed to grab the largest share, surpassing its rival for the first time in the first quarter of the year. As per DigiTimes, AMD saw its […]

Read full article at https://wccftech.com/amd-overtakes-intel-in-q1-data-center-revenue/

Anthropic Eyes UK Startup’s Fusion Tech Promising 100x Faster AI Inference at One-Tenth the Cost of NVIDIA’s Groq

3 May 2026 at 20:10

Anthropic Eyes UK Startup's Fusion Tech Promising 100x Faster AI Inference at One-Tenth the Cost of NVIDIA's Groq

Anthropic, the creators of Claude AI, are reportedly in early talks with a UK startup whose SRAM tech can boost AI inference by 100x & reduce costs by 10x. Anthropic Reportedly In Early Talks With Fractile, A UK-based Startup Working on the fusion architecture as an AI Inference Booster Currently, Anthropic sources its chips from various companies, including NVIDIA, Google, and Amazon. This trio allows the company to keep running its AI infrastructure without major concerns that are often associated with relying on a single chipmaker. But as compute demand intensifies in the AI space, many AI firms are now […]

Read full article at https://wccftech.com/anthropic-sets-eyes-on-uk-startup-tech-speeds-up-ai-inference-100x-reduces-costs-10x/

Decoding the Future of Inference At NVIDIA: Groq LPUs Join Vera Rubin Platform For Low-Latency Inference

17 March 2026 at 16:00

With its upcoming Vera Rubin rackscale architecture, NVIDIA is going to be integrating LPUs from acquihire Groq, marking a major expansion beyond using GPUs alone for AI inference

The post Decoding the Future of Inference At NVIDIA: Groq LPUs Join Vera Rubin Platform For Low-Latency Inference appeared first on ServeTheHome.

NVIDIA Unveils Vera Rubin With Groq’s LPX to Break Into Inference, a Market Where It Has Never Been First

16 March 2026 at 19:48

A presenter on stage with three open computer servers, showcasing internal components against a black background.

NVIDIA's Groq partnership is now formalizing, as Jensen unveils a hybrid compute tray featuring Groq's third-generation LPU units in a Rubin rack. NVIDIA's Idea With Groq Is to Target 'High-Speed' Workloads, Hoping to Crack the Inference Competition The debate over what NVIDIA would do with Groq has been ongoing for quite some time, and we have maintained a key lead on developments. At GTC 2026, NVIDIA unveiled a new Vera Rubin hybrid compute tray, the Groq 3 LPX, which features eight of the 'unannounced' Groq3 units, which we'll discuss ahead. According to NVIDIA, LPX and Rubin together deliver unprecedented inference […]

Read full article at https://wccftech.com/nvidia-unveils-vera-rubin-with-groq-lpx-to-break-into-inference/

❌
❌