Researchers at Stanford have developed Onyx, a coarse-grained reconfigurable array (CGRA) chip that is the first hardware accelerator to natively support both sparse and dense AI computations. Most AI model parameters are zero or near-zero (sparsity), but current CPUs and GPUs waste energy computing with zeros. Onyx skips these unnecessary calculations, achieving up to 565x better energy-delay product over a 12-core Intel Xeon CPU. Unlike partial solutions from Cerebras or Meta's MTIA, Onyx handles both structured and unstructured sparsity across all operation types including matrix, vector, and tensor math. The team is working on next-generation chips with full ML operation support and better dense-sparse integration.

14m read timeFrom spectrum.ieee.org
Post cover image
Table of contents
What is sparsity?The case for sparsityThe trouble with GPUs and CPUsOnyxThe future with sparsity
311 Impressions