AI Infrastructure
9 min read

NVIDIA Put Groq 3 Into Production as Intel Split Agentic AI

NVIDIA moved Groq 3 LPX into full production while Intel separated agentic AI across CPU, GPU and edge roles—and power became the constraint.

A monumental aged-brass speaking horn turns densely folded cream paper into a swift ribbon of blank ivory tiles.
AI Infrastructure / 9 min read
AIENGINE

9 min read

Share

The AI infrastructure announcements in the 24 hours ending 25 August 2026 at 09:01 in Tehran converged on one operating reality: an agent is not one accelerator workload. NVIDIA moved its Groq 3 LPX decode system into full production, while Intel assigned orchestration, inference and edge work to different architectures. Infineon separately agreed to acquire digital power-control expertise for AI servers.

The strongest new evidence is about specialization, not a general claim that every agent will become faster or cheaper. NVIDIA published one third-party speed result on a particular 31-billion-parameter model and named Nebius as the first adopting cloud. Intel disclosed substantial hardware specifications without prices or application-level benchmarks. Infineon did not disclose its purchase price. The day made the infrastructure stack more concrete while leaving its economics incomplete.

The 24-hour brief

DevelopmentConfirmed evidenceOperational meaningImportant limit
NVIDIA Groq 3 LPXFull production; 3,431 output tokens per second in a 100,000-token-context Gemma 4 31B test; Nebius named first AI-cloud adopterDecode can be assigned to a low-latency architecture while other hardware handles context and broader inferenceOne model and serving point do not establish end-to-end task time, price, power or capacity under mixed production load
Intel Hot Chips disclosuresDiamond Rapids up to 256 cores; Crescent Island up to 480GB LPDDR5X at 350W; Wildcat Lake up to 17 NPU TOPSCPU orchestration, data-centre inference and client/edge work have different memory, I/O and deployment constraintsIntel did not publish comparable agent-task results, prices or delivery schedules in the announcement
Infineon–C2i transactionC2i adds software-defined multiphase control and smart power-stage expertise; expected close in Q3 2026Rapid accelerator load changes make voltage regulation part of usable compute capacityConsideration, revenue, customer commitments and integration milestones were not disclosed
X Financial quarterLoan facilitation and origination fell 70.2% year over year; 91–180-day delinquency was 9.09%Automated underwriting remains bounded by credit vintages, funding and regulationThe company says recent tightening helped sequentially, but the historical comparison remains weak

Reports of prospective NVIDIA investments and licensing arrangements based on unnamed sources are excluded. This brief uses disclosed production status, scheduled conference presentations, a signed acquisition announcement and filed financial results.

NVIDIA made decode a separate production tier

NVIDIA’s 24 August production announcement says Groq 3 LPX is now in full production as an extension to Vera Rubin NVL72. NVIDIA describes two separate inference jobs: processing the prompt and accumulated context, then generating output tokens sequentially. LPX is optimized for the second job, where small delays repeat across every token and every agent turn.

That is a narrower and more useful claim than “faster AI”. NVIDIA’s technical account dated 24 August says Artificial Analysis ran Gemma 4 31B on an NVIDIA-hosted LPX system using 100,000 input tokens and measured a median 3,431 output tokens per second. The same post reports 3,382 at 10,000 tokens. It attributes the result to compiler-scheduled execution, low first-bit latency, fine-grained overlap of computation and chip-to-chip communication, and tensor parallelism at small batch sizes.

The important architecture is the hand-off. NVIDIA describes configurations in which Vera Rubin handles prefill and transfers the key-value cache to LPX for decode; splits attention and feed-forward layers across the platforms; or uses LPX as an external draft model whose tokens Rubin verifies. An operator is buying a serving topology and software path, not just a fast component.

The benchmark still has hard boundaries. It measures output speed on one dense FP8 model, not answer quality, retrieval, tool latency, queueing, concurrency, rack power or cost per completed task. NVIDIA disclosed that the test system was in its own data centre, but not the exact chip count or a price. The Register’s 24 August analysis also notes that Gemma 4 31B is a favourable fit and that larger mixture-of-experts models create different placement and communication costs. “Full production” establishes a manufacturing stage; Nebius’s planned adoption does not yet provide public customer traffic, service pricing or utilisation.

Intel split the workload across three architectures

Intel used Hot Chips to make a similar point with a different portfolio. The official 24 August Intel disclosure positions Diamond Rapids for high-performance orchestration, Crescent Island for data-centre inference and Wildcat Lake for client and edge systems. The Hot Chips programme confirms that all three were presented on 24 August in California, within the locked window.

Diamond Rapids, Intel’s next Xeon architecture, is specified with up to 256 cores, 1.28GB of last-level cache, 16 memory channels at 12,800 MT/s, and 128 PCIe Gen6 lanes with CXL 3.0. The operational interest is not only core count. Agents repeatedly move between model calls, code execution, data processing, storage and tools. High memory bandwidth, I/O and a unified fabric can determine whether accelerators remain fed and whether orchestration becomes the bottleneck.

Crescent Island targets another constraint. Intel specifies 32 Xe3P cores, 256 XMX engines, as much as 480GB of LPDDR5X memory and a 350W air-cooled PCIe form factor. The large non-HBM memory pool and air-cooling target could make deployment easier in existing racks, but Intel published no tokens-per-second, concurrency, energy-per-token or price comparison. “Up to” capacity is a design ceiling, not evidence of useful throughput.

Wildcat Lake is already launched as Intel Core Series 3. Its six-core configuration, integrated Xe3 graphics and NPU of up to 17 TOPS are aimed at price-sensitive laptops and edge systems. That is not a substitute for a data-centre GPU. It is evidence that some classification, preprocessing, privacy-sensitive work or local control may sit closer to the user—provided the actual model fits and its quality is evaluated.

The lesson is to partition from measured workload traces. Our context working-set guide explains why advertised memory or context is not permission to fill it, while the backpressure guide shows why fast hardware still needs bounded admission, queue age and retry budgets.

Infineon bought capability in the power path

Infineon’s 24 August acquisition notice says it will acquire Bangalore-based C2i Semiconductors, with closing expected in the third quarter of 2026. C2i develops software-defined multiphase controllers, smart power stages and system-level designs for vertical power delivery, including work toward substrate-integrated voltage regulators.

The rationale is physical. An accelerator does not draw a steady average load: prompt processing, decode, communication and idle gaps produce rapid changes. Voltage regulation must follow those changes without instability or excess conversion loss. Moving regulation closer to the compute package can reduce path losses and support higher density, but it also raises thermal, packaging and control complexity.

Infineon contributes silicon, silicon-carbide and gallium-nitride power devices plus manufacturing and customer access; C2i contributes digital control and architecture expertise. That combination may shorten development, but the notice provides no acquisition consideration, C2i revenue, customer list, product qualification dates or quantified efficiency gain. Until the transaction closes and products ship, it is an R&D and integration thesis rather than booked AI-infrastructure growth.

This is the lower layer of the capital problem examined in our review of VNET’s one-gigawatt build-out: purchased compute only becomes useful capacity after power delivery, cooling, interconnect, commissioning and billable utilisation all work together.

X Financial showed the limit outside the data centre

The day’s most consequential fintech filing came from X Financial. Its 6-K exhibit accepted on 24 August reports RMB11.63 billion of facilitated and originated loans in the June quarter, down 20.5% sequentially and 70.2% year over year. Active borrowers fell 74.8% year over year to 720,258, while the outstanding balance fell 61.5% to RMB24.97 billion.

Credit signals improved from the March quarter but remained materially worse than a year earlier. The 31–60-day delinquency rate fell sequentially from 2.61% to 1.73%, versus 1.16% a year earlier. The 91–180-day rate eased from 9.95% to 9.09%, versus 2.91% a year earlier. The denominator also needs care: the filing excludes most loans more than 60 days delinquent from outstanding balance, and excludes balances beyond 180 days from the 91–180-day calculation.

Revenue fell 56.3% year over year to RMB993.6 million and GAAP net income fell 91.1% to RMB47.0 million. Loan-facilitation fees dropped 85.5%, while guarantee income rose 119.3% as the older guaranteed portfolio continued to run off. Management says tighter borrower selection and conservative capital deployment drove the contraction, and warns that emerging internet-lending requirements could materially hurt results. Its platform uses proprietary big-data technology, but the release does not isolate an AI model’s approval, loss or fairness performance. Automation cannot be credited for a portfolio outcome that the filing does not measure.

What operators should take from the day

  • Separate prefill, decode, orchestration, retrieval, networking and tool execution in traces before selecting hardware.
  • Compare completed tasks per unit of power and cost, not a vendor’s best token rate in isolation.
  • Require model, precision, context, batch size, concurrency, chip count, power envelope and quality checks with every benchmark.
  • Treat “full production”, “first adopter”, “planned deployment” and generally available paid service as different milestones.
  • Model voltage regulation, cooling and rack integration as capacity dependencies, not facilities footnotes.
  • For automated credit, reconcile approval policy, borrower mix, delinquency definitions, vintage losses, funding and regulation before attributing results to technology.

What remains uncertain

NVIDIA has not disclosed LPX pricing, rack power, shipment volume or Nebius service availability. Intel has not supplied comparable application benchmarks, prices or firm delivery dates for the future data-centre products in its announcement. Infineon has not disclosed the C2i purchase price or product roadmap. X Financial has limited visibility into the Chinese regulatory implementation it says may affect future results.

The common uncertainty is economic: specialized systems can improve one constrained stage while moving cost or delay elsewhere. A fast decoder can wait on context, retrieval or tools; a large-memory accelerator can remain underused; a denser rack can be power-limited; a stricter credit model can improve recent vintages while shrinking revenue.

What to watch next

  • Nebius’s Groq 3 LPX service date, price, supported models and production latency under concurrency.
  • Independent end-to-end tests that include prefill, decode, retrieval, tools, queueing, power and answer quality.
  • Intel’s Diamond Rapids and Crescent Island availability, system pricing and workload-level results.
  • C2i transaction closing, engineering retention and a quantified power-conversion or density milestone.
  • X Financial’s later loan vintages, 91–180-day delinquency, guarantee exposure and final regulatory requirements.

The day did not establish one winning agentic-AI machine. It established a more demanding procurement question: which part of the real workflow is constrained, and does specialized hardware remove that constraint without creating a larger one in memory, power, software or finance?

TaggedNVIDIA Groq 3 LPXIntelAI InferenceAgentic AIData Centre PowerFintech
Work With Us

Interested in implementing this for your business?

We help UK businesses put these ideas into practice. Book a call to discuss your specific situation.