Nvidia Commercializes Groq Chip Technology With New LPX Rack

Prime Highlights

  • Nvidia’s Groq 3 LPX rack has entered full production and will deploy at Nebius later this year alongside Vera and Rubin processors.
  • Each rack can deliver 3,400 tokens per second, based on a third-party benchmark, targeting low-latency AI inference demand.

Key Facts

  • Nvidia acquired Groq’s assets in December for $20 billion, its largest acquisition on record.
  • Nvidia is scheduled to report quarterly earnings on Wednesday.

Background

Nvidia announced this week that its Groq 3 LPX rack has entered full production, commercializing technology gained through the company’s largest acquisition to date. The Groq rack will be deployed alongside Nvidia’s Vera processors and Rubin graphics processors at cloud provider Nebius, with deployment expected later this year, according to Nvidia senior director Dion Harris.

The move reflects growing demand for low-latency inference, which helps artificial intelligence agents respond quickly without noticeable delays, particularly for coding tasks. Nvidia said cloud companies can charge premium prices for such fast-response services. Harris said the technology allows companies serving tokens to offer premium service tiers to customers who need the fastest response times.

Nvidia acquired assets from chip startup Groq in December for $20 billion, marking its largest purchase on record. The Groq architecture includes 500 megabytes of fast on-chip memory to reduce bottlenecks. Samsung manufactures Groq chips, while Taiwan Semiconductor Manufacturing Co. produces Nvidia’s graphics processors. Each Groq 3 LPX rack packages 256 individual chips and can deliver 3,400 tokens per second, based on a benchmark from Artificial Analysis.

Competition in the space is intensifying. Advanced Micro Devices said earlier this year it would integrate its systems with chips from Cerebras, which recently went public. OpenAI’s newly announced Ultrafast mode currently offers 750 tokens per second using Cerebras technology.

Harris clarified that low-latency chips do not replace GPUs, which remain central to training and broader AI workloads. He said the goal is matching the right processor to the right part of each workload. Nvidia is set to report earnings on Wednesday.