상세 보기
초록
Systolic array architectures have become the dominant solution for accelerating deep neural network (DNN) computations, yet designing optimal configurations remains challenging due to the vast design space spanning array dimensions, dataflow strategies, SRAM allocations, and tiling sizes. Existing tool-aided optimization frameworks constrain this design space by treating SRAM sizes or tiling sizes as fixed input parameters, while focusing only on conventional systolic arrays, limiting their ability to discover a truly optimal design. This article presents a comprehensive framework for systolic tensor array (STA) design space exploration that jointly optimizes array configurations, SRAM sizes, tiling sizes, and dataflow types. We introduce TensorSim, a cycle-accurate simulator with an ML-based synthesis prediction model, and TensorOptimizer, an automated multiobjective optimization framework. Evaluation with TensorOptimizer across 12 representative DNNs-from edge-level convolutional neural networks (CNNs) to billion-parameter large language models (LLMs)-reveals that output stationary (OS) dataflows outperform weight stationary (WS) dataflows by on edge workloads, while WS achieves higher efficiency over OS on server-scale computations. Based on these insights, we propose Fused-STA, a reconfigurable architecture that dynamically switches between OS and WS modes within a unified hardware substrate. Fused-STA achieves over 91% of the oracle single-dataflow performance, specifically designed for a single network, providing and improvements over a tensor processing unit (TPU)-like baseline on edge-level and server-level benchmarks, respectively.
키워드
- 제목
- Fused-STA: Automated Design Space Exploration of a Fused Systolic Tensor Array for Universal Deep Learning Acceleration
- 저자
- Lee, Jooyeon; Jung, Sangwoo; Kung, Jaeha
- 발행일
- 2026-05-27
- 유형
- Article; Early Access