Topology design for distributed training
Tree, torus, dragonfly — and what comes next when wavelength multiplexing and reconfigurable switches change the cost model.
Wiring up thousands of accelerators with light.
Training a frontier AI model requires thousands of GPUs working together. The bottleneck is no longer compute — it's communication. Electrical interconnects between accelerators burn energy and limit bandwidth; as model parallelism deepens, the network becomes the dominant cost.
We design photonic interconnects and network topologies for warehouse-scale AI training and inference, asking how the physical layer should change when the workload is distributed transformer training, mixture-of-experts inference, or large graph neural networks. The interesting questions aren't just "more bandwidth, less energy" — they're about co-designing the network and the algorithm so the whole system gets faster and greener at once.
Tree, torus, dragonfly — and what comes next when wavelength multiplexing and reconfigurable switches change the cost model.
Bringing the optical I/O into the accelerator package itself: latency, density, and the system-level energy budget.
SUSTAINPHOT — modeling and minimizing the carbon footprint of training runs across geographically-distributed photonic fabrics.
How collective-communication primitives (all-reduce, all-to-all) should be redesigned for a fabric where bandwidth is plentiful but topology is sparse.