Key Takeaways
- NVIDIA has moved Groq 3 LPX into full production as an inference extension for Vera Rubin NVL72 systems.
- NVIDIA's press release reports 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context in Artificial Analysis benchmarking.
- The accelerator addresses token-generation latency, one component of a broader agentic pipeline that also includes context processing, networking, storage, tool calls and orchestration.
- Nebius is the first AI cloud adopting Groq 3 LPX, with access planned through Nebius Token Factory; Groq plans to be among the platform's earliest adopters.
- Buyers should test complete agent workflows under representative loads rather than treating tokens per second as a substitute for end-to-end latency.
The AI infrastructure market is moving from a training-first era toward one in which inference performance determines whether an application is commercially usable. That change becomes especially pronounced with agentic AI: A single user request can trigger hundreds or thousands of inference steps as an agent examines context, invokes tools, generates code and checks its own work.
The investment reflects that shift. IDC reported global AI infrastructure spending of $318B in 2025, up from $153B in 2024, with servers representing approximately 98% of AI-centric spending and accelerated servers accounting for 91.8% of AI server spending, according to AI Infrastructure Spending Reached a Record $86B in Q3 2025, According to IDC. IDC now forecasts $497B in AI infrastructure spending in 2026 and more than $1T by 2029, as detailed in AI Infrastructure Spending Hits $89.7B as ARM Passes x86.
Yet more compute does not automatically create a responsive agent. Akamai-commissioned research finds that 64% of organizations require end-to-end AI response times below 250 ms for their most critical use cases, while 50% fail to meet those targets at peak load. Industry benchmarking places the perceived real-time threshold for generative media and agent workflows at sub-300 ms, making system architecture as important as model capability.
Generation Speed Becomes a Systems Metric
NVIDIA Groq 3 LPX is designed to accelerate the generation phase of inference, or the rate at which a model produces output tokens for an individual user. NVIDIA positions it as an extension of Vera Rubin NVL72 rather than a replacement for the rack-scale platform's broader training and inference capabilities.
According to NVIDIA's press release, Groq 3 LPX produced a record 3,400 output tokens per second in Artificial Analysis benchmarking with Gemma 4 31B, using a 100,000-token context. The company also describes agentic tasks such as coding in minutes versus hours, providing up to 4x faster responsiveness for agents and latency-sensitive workloads compared to certain alternative platforms.
Those are consequential results, but buyers will need details about the alternatives, service configuration, concurrency, accuracy constraints and behavior at peak utilization. A high single-request token rate can shorten individual agent turns, while production responsiveness still depends on queueing, prompt processing, retrieval, network transit and external tool execution.
"Inference is the growth engine of AI. NVIDIA Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency. Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation. This transforms how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness, just as demand for AI computation is accelerating worldwide." Jensen Huang, founder and CEO, NVIDIA
Codesign Extends Beyond the Accelerator
Groq 3 LPX sits within Vera Rubin's larger AI-factory architecture. NVIDIA's press release describes codesign across seven chips and five purpose-built racks, with Vera Rubin NVL72 and Groq 3 LPX addressing different workload requirements for frontier model developers and open-model service providers.
The surrounding systems include NVIDIA BlueField®-4 DPUs, NVIDIA Vera CPU racks, NVIDIA Vera BlueField-4 STX storage and NVIDIA Spectrum™-6 SPX Ethernet. That design matters because agentic latency can migrate from the accelerator to storage, networking or orchestration as token generation becomes faster. Gartner's enterprise storage analysis identifies ultra-low, sub-millisecond data-access latency as a requirement for scalable AI and real-time inference.
Deployment location is another variable. IDC's Worldwide Edge Spending Guide Forecast 2025 reports that worldwide edge computing spend reached $265B in 2025 and is expected to nearly double to around $450B by 2029, driven largely by AI workloads requiring low-latency inference and real-time context. IDC also predicts that by 2027, 80% of CIOs will rely on edge services from cloud providers to meet AI inferencing performance and compliance requirements.
Benchmarks Must Reflect Agent Behavior
Traditional throughput measures remain useful for capacity planning, but they do not fully describe interactive agents. MLCommons expanded the MLPerf Inference interactive scenario in 2025 to better represent agentic and other large-language-model applications, measuring time to first token and time per output token under latency and accuracy constraints.
Its 2026 Edge Agentic Inference work goes further by reporting p50, p90 and p99 end-to-end latency alongside time to first token and time per output token in single-accelerator, single-stream runs. The organization outlined that approach in its Call for Submission: Edge Agentic Inference Benchmark ....
That multidimensional approach should guide enterprise evaluations of Groq 3 LPX. A coding agent, for example, may alternate among context ingestion, token generation, file inspection, compilation and testing. The useful benchmark is therefore not only how quickly the model writes tokens, but how quickly the complete loop reaches a verified result.
Nebius Provides the First Cloud Route
Nebius plans to offer NVIDIA Groq 3 LPX through Nebius Token Factory, its production inference platform. According to NVIDIA's press release, Nebius is the first AI cloud adopting the accelerator, while Groq plans to be among its earliest adopters after Nebius.
"Generation is the phase of inference that determines how responsive an AI system actually is, and that's exactly what NVIDIA Groq 3 LPX is built to accelerate. As the first AI cloud bringing it to production via Nebius Token Factory, we're making sure every step of an agent's loop feels instant — through the same API developers are already using, with no migration to a new stack." Danila Shtan, chief technology officer, Nebius
The API continuity described by Nebius could reduce application migration work, but enterprises will still need to assess pricing, regional availability, model support, service-level objectives and performance under shared-cloud load. Cloud access will also determine whether Groq 3 LPX becomes a broadly consumable inference tier or remains concentrated in specialized deployments.
Common Questions
Is 3,400 output tokens per second equivalent to end-to-end application latency?
No. It measures token-generation performance for Gemma 4 31B with a 100,000-token context in the cited Artificial Analysis benchmark. End-to-end latency also includes prompt processing, networking, storage access, queueing, tool calls and application orchestration.
What should enterprises benchmark before adopting Groq 3 LPX?
They should test representative agent workflows at expected concurrency, including time to first token, time per output token and p50, p90 and p99 completion latency. Evaluations should also maintain the required model accuracy and include retrieval, tool execution and verification steps.
Does Groq 3 LPX replace Vera Rubin NVL72?
No. NVIDIA presents Groq 3 LPX as an extension that increases token-generation performance for Vera Rubin NVL72 systems. Vera Rubin remains the broader training and inference platform, while LPX targets interactive generation for latency-sensitive workloads.
The Next Inference Contest
Groq 3 LPX indicates that inference competition is becoming more specialized: context processing, generation, storage and networking can be optimized as distinct parts of an AI factory. Its commercial importance will ultimately depend on whether the benchmarked token rate translates into predictable end-to-end gains across real agent workloads, concurrency levels and cloud environments.
⬇️