EN

Nvidia Expands Agent Infrastructure as OpenAI Brings GPT-5.6 to Kiro

Nicole Jeffrey

Key Points

  1. Nvidia put Groq 3 LPX into full production, with Nebius adopting the inference chip for its Token Factory service.
  2. Amazon and OpenAI expanded agent models, discovery tools and payment infrastructure across AWS platforms.
  3. Alibaba’s $10.3 billion placement shows the scale of capital being directed toward AI infrastructure.

The latest

Nvidia put its Groq 3 LPX inference chip into full production and unveiled plans to deploy Vera processors with SpaceXAI, accelerating its push into hardware optimized for autonomous agents. The announcements came as Amazon and OpenAI broadened agentic software and model offerings, while Alibaba priced a HK$80 billion share placement whose net proceeds will be committed entirely to full-stack AI capabilities.

Details

  • Inference chip: Built on LPU technology licensed from Groq, the chip extends Vera Rubin for interactive inference. Nvidia says it can generate up to 3,400 output tokens per second on a 31-billion-parameter Gemma 4 model with a 100,000-token context window, with responsiveness up to four times faster. Nebius is the first customer.
  • Efficiency claims: Nvidia says Vera Rubin NVL72 delivers 30 times more throughput per megawatt and 35 times lower cost per token than GB300 NVL72, citing SemiAnalysis’ AgentX benchmark. Jensen Huang called inference “the growth engine of AI” and tied the configuration to rising agentic workloads.
  • Orbital deployment: SpaceXAI will use Vera CPUs for agentic workloads, scale Grok infrastructure through Vera Rubin toward gigawatt-scale capacity and place NVL72 systems aboard the first-generation Starmind satellite. Vera has 88 Olympus cores and LPDDR5X memory delivering up to 1.2 terabytes per second of bandwidth. Nvidia claims up to 1.8 times x86 performance on targeted workloads.
  • Coding models: OpenAI made GPT-5.6 variants Sol, Terra and Luna available in AWS’s Kiro coding platform, spanning planning, building, review and testing. It cited an approximately 82% cost reduction for Terra on Terminal-Bench 2.1 within Kiro. Bedrock also added GPT-5.6 and xAI’s Grok 4.6 with cross-region inference.
  • Agent discovery: AWS introduced a registry for managing agents, MCP servers, tools and skills with approval workflows. Its Apache 2.0-licensed Agentic Resource Discovery specification is designed to work like DNS across clouds and organizations, exposing resources without migration while keeping access and governance local. AgentCore payments also reached general availability with Coinbase integration and Machine Payment Protocol support.
  • Capital and defense: Alibaba priced 710 million new shares at HK$112.70 each for offshore investors outside the United States under Regulation S. Separately, Sakana AI disclosed a Japanese Defense Ministry contract to research AI for intelligence collection, analysis and information management, extending a March 2026 tactical command-and-control contract into strategic security-policy work.

Background

Nvidia cites OpenRouter data indicating agentic workloads consume far more tokens than simple chat requests, driving hardware optimization around inference speed, energy use and token cost.

What’s next

Alibaba’s placement is scheduled to close on August 26, 2026; its net proceeds are designated for full-stack AI investment, including infrastructure expansion and enhancement.

Source

 

What to read next