Cross-Filter Structured Pruning for Efficient Sparse CNN Acceleration

  • Pham, Ngoc-Son; 
  • Shin, Sangwon; 
  • Xu, Lei; 
  • Shi, Weidong; 
  • Suh, Taeweon
Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

Convolutional Neural Networks (CNNs) are widely used in vision tasks for resource-constrained environments due to their computational efficiency and strong generalization. However, the dominance of 1 x 1 convolutions in modern CNN architectures introduces challenges for sparsity-aware hardware accelerators, particularly in processing element (PE) load balancing, which limits speedup in sparse inference. To address this issue, this paper proposes a cross-filter structured pruning method, enforcing a uniform sparsity pattern across multiple filters to ensure balanced workload distribution among PEs. This approach is further extended to k x k convolutions by decomposing them into 1 x 1 filters, improving applicability across various CNN layers. This paper also proposes an intra-kernel parallelism technique which significantly reduces the size of PE's local buffers, a critical bottleneck in sparse CNN accelerators. Experimental results show that the proposed approach maintains accuracy comparable to globally unstructured pruning while significantly enhancing inference speed. FPGA implementation and cycle-accurate simulations confirm improvements in processing speed, energy efficiency, and hardware utilization, making this method well-suited for edge and mobile AI applications. Specifically, the proposed architecture achieves 1.14x to 1.6x speedup over Sparten for various CNN models and delivers 7.6x and 1.9x higher energy efficiency compared to Sparten and StarSPA, respectively. In terms of area efficiency, synthesis results show a 1.73x-10.95x reduction in required hardware primitives compared to Sparten and StarSPA.

키워드

Filters; Convolutional neural networks; Computer architecture; Computational efficiency; Accuracy; Parallel processing; Feature extraction; Computational modeling; Residual neural networks; Energy efficiency; AI accelerator; convolutional neural networks (CNNs); sparsity exploitation; data compression; dataflow; network on a chip (NoC); CONVOLUTIONAL NEURAL-NETWORK
제목
Cross-Filter Structured Pruning for Efficient Sparse CNN Acceleration
저자
Pham, Ngoc-Son; Shin, Sangwon; Xu, Lei; Shi, Weidong; Suh, Taeweon
DOI
10.1109/ACCESS.2025.3587027
발행일
2025
유형
Article
저널명
IEEE Access
권
13
페이지
129461 ~ 129475