NVIDIA Groq 3 LPX Achieves Full Production for Agentic AI
Caroline Bishop Aug 24, 2026 16:01 NVIDIAs Groq 3 LPX and Vera Rubin NVL72 enhance AI inference with faster token generation and lower costs, reshaping agentic AI infrastructure. NVIDIA announced on August 24 that its Groq 3 LPX inference accelerator is now in full production, marking a significant step in supporting the next generation of agentic AI systems. Groq 3 LPX, designed to complement NVIDIAs Vera Rubin NVL72 AI platform, delivers industry-leading token generation speeds—crucial for applications requiring real-time responsiveness. In a benchmark test using Gemma 4 31B, an open-source agentic AI model, Groq 3 LPX produced 3,400 output tokens per second for long-context use cases, a fourfold improvement over competing platforms. This performance positions NVIDIA to address the growing demand for faster AI inference as the industry shifts from model training to large-scale reasoning and multi-agent systems. AI Factories Demand Infrastructure Evolution Agentic AI systems are reshaping the AI market by requiring extensive context windows, high token throughput, and low latency for real-time decision-making. NVIDIAs Vera Rubin NVL72 platform integrates CPUs, GPUs, and NVLink networking into a unified system, offering a scalable solution for these workloads. By enabling lower token costs and higher efficiency, the platform is designed to meet the needs of hyperscale AI









