2026 AI Infra Summit
Welcome to our AI Infra Summit 2026 Video Showcase! The conference, organized by Kisaco Research, brought roughly 8,000 attendees to the Santa Clara Convention Center, September 15–17, for a full-stack view of AI infrastructure — chips, interconnects, networking, memory, storage, power, and the data centers that tie them together.
We were on site filming interviews with the executives and companies below. Videos are in production and will appear here as they are published, so check back often. Meanwhile, grab our latest research brief right from this page.
Thanks to our sponsors for supporting our site!
AI Cooling Tech Evolution
Performance at Scale, Energy Efficiency & Cybersecurity
Astera Labs Taurus 200G Smart Signal Conditioners
Smart Memory Controllers for AI Infrastructure
Power Efficiency: 10-30% More Performance Per Watt
Scaling Beyond Power Constraints with CPO
Making AI Work with Enterprise Systems
Scaling AI GPU Clusters with Ethernet
Explaining Open Compute Interconnect (OCI)
Cognichip's AI Cuts Chip Design Time
Cornelis Unveils Active Compute Fabric for AI
Credo's 1.6T AEC with 200G SerDes Powers 288-GPU Cluster
Credo's 1.6T ZeroFlap Optics
Credo's ZeroFlap Optics: Proactive Link Monitoring
d-Matrix's 3D Stacked XPU & NVIDIA Partnership
AI Infrastructure Beyond GPUs: Networking Scale & Multi-Vendor Challenges
Silicon Photonics: 20x Denser Optical Switching
Bidirectional Optics for Massive GPU Clusters
Think Optical for Efficiency & Memory Scaling
AI Infrastructure Bottleneck: Solving Connectivity at 100 Terabit Scale
NextSilicon: HPC, AI & RISC-V CPU Integration
Open-Source: Cutting Costs and Keeping Control of Your Data
NVIDIA's Vera Rubin Performance Leap Explained
CXL Solutions for Data Centers
Managing Heat at Die Level for Maximum GPU Output
From Milliwatts to Gigawatts Across Devices
Making Inferencing Affordable at Scale
Building Trust in Physical AI
UALink Consortium Shows Working Models

Latest Research Brief
Download Our Latest Report
Data Center Networking in 2026 — Scaling to New Heights
More about this report →
No personal info needed to download
By downloading this report, you agree to our Terms of Use and Privacy Policy.
What we learned at AI Infra Summit 2026
Interviews from the Santa Clara show floor converge on one constraint: power is fixed, so every layer of the stack is judged by the tokens it extracts from each watt.
Tokens per watt is the new scoreboard
Aegis tracks power compute effectiveness, Phononic controls heat at the die, Axiado claims 10–30% more tokens per watt from idle XPU and CPU cycles, and Rebellions builds 4-kilowatt inference servers. Many vendors squeezing more out of constrained power.
The network decides how much of the GPU you get
Marvell puts AI cluster efficiency at only 25–50% and calls connectivity the primary bottleneck. DriveNets links data centers past single-site power ceilings, Cornelis moves compute into NICs and switches, and Credo cuts optical fabric bring-up from 4–6 weeks to about one.
Optics moves toward the silicon
Ayar Labs has raised $650M+ for co-packaged optics; Broadcom's OCI carries 200 Gbps as multiple 50 Gbps wavelengths on one fiber; Lightmatter halves fiber count with bidirectional optics; iPronics delivers optical switching at 10–20x MEMS density.
Scale-up is no longer one vendor's domain
UALink showed working models from Cadence, NetForward, and Synopsys, with 3.0 in progress. The OCP ESUN initiative, seeded by Broadcom, has 75 members and a 1.0 spec. Even NVIDIA opens up: NVLink Fusion brings d-Matrix's Raptor XPU into its racks.
Inference drives innovation — in memory and beyond
NVIDIA shows Vera Rubin at 67x Blackwell on the AgentX benchmark. To feed agentic workloads, Marvell proposes optically connected shared memory for KV cache, Astera Labs puts memory controllers on the fabric, Penguin's CXL cards enable an 11TB server, and d-Matrix stacks DRAM on compute.
Beyond the hyperscaler: control, diversity, and cost
Nutanix sends only 10–20% of workloads to frontier models, keeping the rest on open models it controls at lower cost. DriveNets sees AMD's Helios running alongside NVIDIA, and Qualcomm scales from milliwatts to gigawatts. Buyers want choice, not lock-in.
Highlights from industry thought leaders

























Video interviews
AI Cooling Tech Evolution
- Every watt a GPU draws creates a matching thermal obligation, so power and cooling have to be solved as one problem as AI compute density climbs.
- Cooling is moving from air to single-phase and two-phase direct-to-chip, which works only when the system is engineered as a whole rather than assembled from standalone components.
- Power compute effectiveness is the metric that matters — the share of each watt that reaches the GPU rather than the supporting infrastructure.
Abstract
Performance at Scale, Energy Efficiency & Cybersecurity
- Performance at scale dominates the summit conversation, spanning compute, data centers, edge, and physical AI, with chiplet architectures drawing the most excitement.
- Energy efficiency has to be architected from the silicon up through racks to the whole data center, not bolted on afterwards.
- The attack surface has widened since Spectre and Meltdown to CPUs, NPUs, XPUs and whole semiconductor subsystems, so security must run from silicon through firmware to software.
Abstract
Astera Labs Taurus 200G Smart Signal Conditioners
- Taurus 200G-per-lane signal conditioners cover Ethernet, UALink, and PCIe in the industry's first OCP-compatible footprint spanning both retimers and redrivers.
- Smart swap lets customers switch between retimer and redriver at any point in the design cycle, with one Cosmos software stack across both.
- The Taurus 4 3.2T retimer and redriver were shown driving PRBS patterns across a backplane mockup, trading reach against power for scale-up links.
Abstract
Smart Memory Controllers for AI Infrastructure
- Three new LEO family members land at AI Infra 2026: the fabric-attached LEO X series, plus LEO 2 E and P series for expansion and pooling.
- LEO X scales memory controllers alongside GPUs to add KV cache to the fabric, pairing with the Scorpio 2 320-lane switch for faster time to first token.
- LEO 2 CXL PCIe 6 hangs 12 DDR4 DIMMs at 3DPC or 8 DDR5 at 2DPC off one controller — double the bandwidth and capacity for cloud compute, KV cache, and in-memory databases.
Abstract
Power Efficiency: 10-30% More Performance Per Watt
- With 160 neoclouds starting up globally, the pitch is to optimize the power a data center already has rather than build more generation capacity.
- The management layer boots first on the platform, collects forensic telemetry on every component, and load-balances the idle periods in XPU and CPU operation.
- Treating racks and pods as single units with no human intervention is claimed to yield 10–30% more tokens per dollar or per watt than cooling-only approaches.
Abstract
Scaling Beyond Power Constraints with CPO
- Co-packaged optics is Ayar Labs' answer to rising token demand under hard power constraints, backed by more than $650M raised this year.
- The ecosystem on show spans TSMC silicon photonics wafers, multi-chip packages with Alchip and GUC, external laser sources, and fiber connectors from Broadcom, Foscy, and Senko.
- Multi-source supply is the precondition for high-volume production, with Wiwynn as strategic partner and investor for hyperscale deployment.
Abstract
Making AI Work with Enterprise Systems
- The hard part of enterprise AI sits downstream of the model — connecting agentic workflows to the systems that already run the business.
- Azul's Java enterprise platform is that bridge, letting major global brands pull data and knowledge out of existing infrastructure into AI workflows.
- Sellers' background in 3D graphics silicon frames today's GPU build-out as a familiar hardware cycle with a new integration problem attached.
Abstract
Scaling AI GPU Clusters with Ethernet
- Ethernet now spans scale-up, scale-out, and scale-across, with the scale-up networking initiative at 75 members and a 1.0 specification published.
- Tomahawk Ultra is in production at multiple hyperscalers at 250ns latency and 77 billion packets per second; Tomahawk 6, the first 100TB switch in volume, enables 128,000-GPU clusters in two tiers.
- Thor Ultra is the first 800G NIC in production, supporting UAC and — through work with Microsoft and NVIDIA — MRC.
Abstract
Explaining Open Compute Interconnect (OCI)
- OCI replaces copper SerDes with an optical equivalent, using fiber's multi-wavelength capacity instead of ever-faster electrical lanes.
- Gen 1 carries eight wavelengths per fiber — four each way — dropping the required line rate from 200 Gbps to 50 Gbps while running bidirectionally on a single fiber.
- 50 Gbps NRZ is the energy-efficiency sweet spot, and Broadcom expects to reveal OCI-based hardware later this year.
Abstract
Suresh Vasudevan, Clockwork.io (coming soon)
Cognichip's AI Cuts Chip Design Time
- Cognichip's artificial chip intelligence platform aims to compress chip development from years to weeks.
- Physics-informed foundation models underpin AI-native workflows running from specification to formally verified RTL with zero-touch UVM verification.
- Subsecond PPA optimization across hundreds of implementation points makes designs Pareto-aware; an idea-to-FPGA demo for the Altera ecosystem cut 20-plus weeks to days.
Abstract
Cornelis Unveils Active Compute Fabric for AI
- The second-generation CN6000 applies Ethernet and Ultra Ethernet standards to Cornelis' native architecture, giving AI system builders interoperability.
- A third-generation roadmap extends the company's scale-out heritage into scale-up switching via UALink and EON, targeting new accelerators, GPUs, and XPUs.
- The active compute fabric moves compute into NICs and switches, turning the network from a passive packet-forwarding layer into an intelligent one.
Abstract
Credo's 1.6T AEC with 200G SerDes Powers 288-GPU Cluster
- Credo's first 1.6T-to-1.6T active electrical cable is built on the company's first 200G SerDes, with thin cables reaching up to 6 meters plus Y cables with two 800G ends.
- PILOT software captures runtime telemetry — FEC rate, histograms, temperature, and SNR — from both the host and remote ends of every cable.
- A GB200 NVL72 deployment wires 288 GPUs across four racks plus a network rack with Credo's 6-meter cables to build a zero-flap scale-out fabric.
Abstract
Credo's 1.6T ZeroFlap Optics
- Credo's 1.6T ZeroFlap optics extend its ecosystem with far higher reliability than commodity optics, plus built-in telemetry and inband communication.
- The optics flag the one to two percent of problem links at installation and raise intelligent service tickets, removing the need for fiber plant characterization and manual cleaning.
- Optical fabric bring-up drops from four to six weeks to about one week, with the Rubin-generation parts ramping for most customers in 2027.
Abstract
Credo's ZeroFlap Optics: Proactive Link Monitoring
- A live 51T switch cluster with ConnectX-8 NICs streams optic telemetry across 64 ports, rolled up into a single green, yellow, or red link score.
- The link score combines pre-FEC bit error rate, multipath interference (a telltale of dust), and signal-to-noise ratio.
- For neocloud bare-metal GPU fleets, telemetry is carried inband to the switch, alerts fire only on change with recommended actions, and optics pull themselves out of service before they start flapping.
Abstract
d-Matrix's 3D Stacked XPU & NVIDIA Partnership
- d-Matrix builds silicon, subsystems, and software for ultra-low-latency inference; two acquisitions in its first year reflect that the product is now the rack, not the chip.
- The Raptor XPU is the first 3D stacked DRAM-based XPU, co-packaging DRAM directly on compute.
- An NVIDIA NVLink Fusion partnership puts Raptor inside NVIDIA's rack ecosystem, letting 144 XPUs communicate in one large scale-up domain.
Abstract
AI Infrastructure Beyond GPUs: Networking Scale & Multi-Vendor Challenges
- Idle GPUs are the most expensive wasted resource, making the network that feeds them as critical as the accelerators themselves.
- Scale-across networking links multiple data centers so AI clusters can keep growing past a single facility's power ceiling — already deployed in the field.
- Multi-vendor GPU environments, with NVIDIA alongside AMD's Helios, need networking tuned to heterogeneous traffic patterns, especially for inference.
Abstract
Silicon Photonics: 20x Denser Optical Switching
- iPronics builds optical circuit switching directly into silicon photonic chips, aimed at the scale-up portion of the data center network.
- The chip-integrated approach runs 10–20x denser than MEMS-based switching and reconfigures the network roughly 1,000 times faster.
- A recent $125M round funds supply chain, production volume, and the qualification work needed to ship a reliable product.
Abstract
Bidirectional Optics for Massive GPU Clusters
- Passage L20 brings bidirectional optics to scale-up networks, linking giant GPU clusters with half the optical fiber.
- Halving fiber count directly addresses an anticipated global shortage in optical fiber supply.
- With power constrained, the optics let a gigawatt data center perform like a three-gigawatt one — the same compute and GPU count, faster tokens.
Abstract
Think Optical for Efficiency & Memory Scaling
- As workloads shift from training to inference-heavy agentic AI, CPU-to-GPU ratios move back toward 1:1 and the traditional memory pyramid has to be broken apart.
- Optically connected shared memory adds a tier that absorbs the exponential KV cache growth of agentic AI, pooling capacity across racks at low latency and high bandwidth.
- Optics keeps moving closer to compute — from scale-out rack links to scale-up inside the rack, and eventually onto the die itself as electrical signal integrity and beachfront limits bite.
Abstract
AI Infrastructure Bottleneck: Solving Connectivity at 100 Terabit Scale
- Connectivity — not compute or memory — is the primary bottleneck, and AI cluster efficiency today runs at only 25–50%.
- Refresh cycles have compressed from 3–4 years to 12–18 months, pushing photonic fabric and optics closer to silicon as copper runs out of reach.
- Marvell's 100Tbps switch on 3nm with 200G SerDes backs open standards (UALink, Ethernet for AI) while staying neutral across ecosystems, NVLink Fusion included.
Abstract
NextSilicon: HPC, AI & RISC-V CPU Integration
- Eight years of platform work has production chips accelerating HPC workloads at major U.S. national laboratories.
- The Maverick chip runs the latest open-source models, with an improved successor already in development.
- The new Arbel RISC-V CPU completes a single platform spanning HPC acceleration, AI workloads, and agentic frameworks.
Abstract
Open-Source: Cutting Costs and Keeping Control of Your Data
- Open-source models are closing on closed frontier models, making near-frontier capability runnable on your own infrastructure.
- Sovereign full-stack infrastructure gives organizations control over data access and how models operate.
- A hybrid split — frontier models for 10–20% of workloads, open models on private infrastructure for the rest — improves both governance and tokenomics.
Abstract
NVIDIA's Vera Rubin Performance Leap Explained
- Agentic AI is the most complex workload data centers have faced, requiring coordination across CPUs, GPUs, networking, and storage.
- Vera Rubin shows 67x more performance than Blackwell on the AgentX benchmark with variable sequence lengths, with further gains expected.
- NVLink Fusion opens NVIDIA's scale-up technology to partners such as d-Matrix, keeping the platform vertically integrated but horizontally open.
Abstract
CXL Solutions for Data Centers
- Memory scaling is the bottleneck: expanding capacity affordably while cutting the data movement that drives up power and bandwidth.
- CXL add-in cards are the answer shipping today — and what made Penguin's own 11TB AI server possible.
- RDMA Ethernet-connected boxes extend memory across networked servers, with silicon photonics for GPU memory scaling still ahead.
Abstract
Managing Heat at Die Level for Maximum GPU Output
- Phononic applies thermal control at the die level for GPU and networking partners, rather than at the rack or the room.
- The cooling infrastructure layers onto existing liquid cooling instead of replacing it.
- Software control over that layer is what converts cooling headroom into more tokens per megawatt.
Abstract
From Milliwatts to Gigawatts Across Devices
- Qualcomm's AI portfolio runs from doorbells to data centers — a milliwatt to a gigawatt — and it claims to be the only vendor spanning that whole range.
- High bandwidth compute and token acceleration target sustained token generation at lower power.
- A company built on endpoints is now scaling into rack-scale deployment and a push into physical AI, humanoid robotics included.
Abstract
Making Inferencing Affordable at Scale
- Inferencing buyers have moved past benchmark tokens to production concerns: security, enterprise readiness, and economics.
- Rebellions' 4-kilowatt servers cut operating expense enough to support lower-cost service tiers and push inferencing unit economics toward zero.
- The newly released Rebel product ships alongside a partnership with AI&, an inference provider serving the Japanese market.
Abstract
Building Trust in Physical AI
- As AI moves into drones, vehicles, and household machines, trust becomes the constraint — predictions need ±1% accuracy when the decision is life-or-death.
- Using physics as the source of truth in synthetic data and simulation is what makes a digital twin validated rather than merely plausible.
- A partnership with NVIDIA brings accelerated computing and agentic orchestration to pre-processing, simulation, and solving, lowering the barrier to 99% engineering accuracy.
Abstract
UALink Consortium Shows Working Models
- Three vendors showed working UALink models: Cadence linking two FPGAs directly, NetForward's FPGA-based switch, and Synopsys proving the 224G Ethernet PHY layer over cable.
- The demonstrations run on the 1.0 specification, with 2.0 released in April and 3.0 now in progress.
- Working IP from multiple vendors marks the shift from specification to silicon for scale-up interconnect.
Abstract
By downloading this report, you agree to our Terms of Use and Privacy Policy.