News

Nvidia Begins Mass Production of Groq 3 Inference Accelerator Produced by Samsung, Secures Clients

Nvidia Begins Mass Production of Groq 3 Inference Accelerator Produced by Samsung, Secures Clients
안내

We only offer this video
to viewers located within Korea
(해당 영상은 해외에서 재생이 불가합니다)

▲ The Groq chip produced by Samsung Foundry

Nvidia's inference-only chips, manufactured by Samsung Electronics, have secured clients and entered full-scale mass production.

Nvidia announced on the 24th local time that it has begun mass production of the Groq 3 LPX, which is specialized for running AI agents and accelerates inference speeds.

The Groq 3 LPX is a system composed of a cabinet-sized rack bundling 256 Language Processing Units (LPUs), which are inference-only chips.

The Groq LPU is a chip for which Nvidia acquired a license through an indirect acquisition late last year. It is characterized by adopting SRAM instead of conventional DRAM memory to enable rapid AI responses.

While somewhat unsuited for AI training, it can complete agent-based tasks such as coding in a matter of minutes, a process that would take hours with conventional AI chips.

Nvidia designed an AI system in which conventional Graphics Processing Units (GPUs) handle general-purpose tasks including model training, while the high-speed Groq 3 LPX handles the inference process where the trained AI derives answers.

In line with the mass production, clients have also been secured.

Nvidia introduced Nebius, an AI cloud operator, as the first client for the Groq 3 LPX.

Nvidia CEO Jensen Huang emphasized, "Inference is the growth engine of AI," adding, "This will bring another giant leap in AI throughput, efficiency, and responsiveness right at a time when global demand for AI computation is accelerating."

Nvidia is mass-producing the Groq 3 LPX entirely through Samsung Electronics' foundry (semiconductor contract manufacturing) division.

Nvidia headquarters building (Photo: Getty Images)

Nvidia has also secured clients for "Vera," the agent-dedicated central processing unit (CPU) that is another core product of the "Vera Rubin" AI system.

Nvidia announced on the same day that Space XAI, an AI company founded by Elon Musk and a subsidiary of SpaceX, has decided to adopt the Vera CPU on a large scale.

The Vera CPU is a product tasked with overseeing various AI agents, allowing other AI chips such as GPUs to focus solely on model computations.

Nvidia explained that the Vera CPU achieves 1.8 times the processing speed for such agent tasks compared to x86-based CPUs produced by companies like Intel and AMD.

In particular, SpaceXAI plans to expand this type of AI system not only to ground data centers but also to data centers in space orbit.

SpaceXAI also presented a blueprint to equip its first-generation "Starmind" AI satellites with the Nvidia Vera Rubin system, which has been optimized for the space environment.

However, Nvidia did not disclose transaction sizes or supply schedules with Nebius and SpaceXAI.

Nvidia's successive announcements of securing major clients for the Groq 3 LPX and Vera CPU to defend its market share come as competition among big tech companies to secure low-latency inference computing resources becomes increasingly fierce.

In particular, this is seen as a move to keep competitor AMD in check, as AMD has unveiled the Helios rack system to counter the Vera Rubin system and has also partnered with low-latency chip startup Cerebras.

(Photo courtesy of Samsung Electronics, Yonhap News)
※ Please note: This article was translated by AI and may contain errors.
Copyright Ⓒ SBS. All rights reserved. 무단 전재, 재배포 및 AI학습 이용 금지

Most Read