Tensor Processing with Large-scale Homodyne Photonic Crossbar
arXiv preprint arXiv:2604.18496, 2026
Presented in the HOT CHIPS 2026 poster session by Opticore.
Abstract
High-performance computing underpins modern artificial intelligence, enabling foundation models, real-time inference and perception in autonomous systems, and data-intensive scientific simulations. Recent advances in quantization techniques using low-precision computation without degrading model accuracy create new opportunities for analog photonic computing characterized by ultra-high clock rates and low energy consumption. Here we propose and demonstrate a coherent homodyne integrated circuit capable of general matrix multiplication with aggregate throughput exceeding 1,000 tera-operations per second, enabled by massive on-chip optical fanout and parallelism. By leveraging time multiplexing, the required modulator count is reduced from quadratic to linear scaling, allowing dense integration of record-scale 256 by 256 homodyne units within a single reticle. The system achieves up to 7-bit computational accuracy across 8 by 8 parallel channels at a record computing clock rate of 120 Gbaud/s, and 6-bit statistical accuracy across 256 by 100 channels at 20–128 Gbaud/s, representing total throughput of 1,000–6,000 TOPS. Massive parallelism amortizes optoelectronic conversion to allow 330 TOPS/W efficiency using foundry-available packaging technology. The system throughput is benchmarked with Qwen2.5-0.5B models that generate accurate tokens. High throughput and energy efficiency establish a near-term pathway toward light-based accelerators for large-scale training and low-latency inference from data centers to edge devices.
