Efficient Non-fused Winograd on GPUs

Hui Wei¹⁶,
Enjie Liu¹⁶,
Youbing Zhao¹⁷ &
…
Hongqing Yu¹⁶

Part of the book series: Lecture Notes in Computer Science ((LNIP,volume 12221))

Included in the following conference series:

Computer Graphics International Conference

2160 Accesses
3 Citations

Abstract

This paper presents an optimized implementation for Winograd non-fused convolution. Our optimizations comprise application-independent grouped producer-consumer chains and a set of Winograd-specific software techniques, including specialized interface-kernels data format which enhances memory access efficiency; warp specialization and double buffer prefetching which effectively exploit computational resources and memory bandwidth; utilizing “shuffle” instruction which conserves hardware resources. The paper also provides supplementary explanation of Winograds’ tile extraction, which saves memory and computing resources.

The proposed techniques has been evaluated head to head by kernel level in GTX 980 GPU, CUDA 9.2 with a wide range of parameters which meet CNN layers benchmark. Compared with the state-of-the-art Winograd Non-fused convolution in CuDnn 7.6.4 (released in Sept, 2019), our implementation achieves a total speedup of 1.64x.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Subscribe and save

Springer+ Basic

£29.99 /Month

Get 10 units per month
Download Article/Chapter or eBook
1 Unit = 1 Article or 1 Chapter
Cancel anytime

Buy Now

Chapter: GBP 19.95; Price includes VAT (United Kingdom)

eBook: GBP 79.50; Price includes VAT (United Kingdom)

Softcover Book: GBP 99.99; Price includes VAT (United Kingdom)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Efficient and portable Winograd convolutions for multi-core processors

Article Open access 12 February 2023

Im2win: An Efficient Convolution Paradigm on GPU

Computing large 2D convolutions on GPU efficiently with the im2tensor algorithm

Article 23 August 2022

References

Lavin, A., Gray, S.: Fast algorithms for convolutional neural networks. In: Proceedings of the CVPR 2016, pp. 4013–4021 (2016)
Google Scholar
Xygkis, A., Soudris, D., Papadopoulos, L., Yous, S., Moloney, D.: Efficient winograd-based convolution kernel implementation on edge devices. In: 55th DAC, pp. 1–6 (2018)
Google Scholar
Jordà, M., Valero-Lara, P., Peña, A.J.: Performance evaluation of cuDNN convolution algorithms on NVIDIA Volta GPUs. IEEE Access 7, 70461–70473 (2019)
Article Google Scholar
Xiao, Q., Liang, Y., Lu, L., Yan, S., Tai, Y.W.: Exploring heterogeneous algorithms for accelerating deep convolutional neural networks on FPGAs. In: Proceedings of the 54th Annual Design Automation Conference, p. 62. ACM (2017)
Google Scholar
Abdelfattah, A., Haidar, A., Tomov, S., Dongarra, J.: Performance, design, and autotuning of batched GEMM for GPUs. In: Kunkel, J.M., Balaji, P., Dongarra, J. (eds.) ISC High Performance 2016. LNCS, vol. 9697, pp. 21–38. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-41321-1_2
Chapter Google Scholar
Tan, G., Li, L., Triechle, S., Phillips, E., Bao, Y., Sun, N.: Fast implementation of DGEMM on Fermi GPU. In: HiPC, Networking, Storage and Analysis, pp. 1–11 (2011)
Google Scholar
Jia, L., Liang, Y., Li, X., Lu, L., Yan, S.: Enabling efficient fast convolution algorithms on GPUs via MegaKernels. IEEE Trans. Comput. 69(7), 986–997 (2020)
MathSciNet MATH Google Scholar
Yan, D., Wang, W., Chu, X.: Optimizing batched winograd convolution on GPUs. In: PPoPP, Main Conference (2020)
Google Scholar
Michael, B., Sean, T., Alex, A.: Singe: leveraging warp specialization for high performance on GPUs. In: ACM SIGPLAN Notices, pp. 119–130 (2014)
Google Scholar
Bauer, M., et al.: CudaDMA: optimizing GPU memory bandwidth via warp specialization. In: HiPC, Networking Storage and Analysis (2011)
Google Scholar

Download references

Author information

Authors and Affiliations

University of Bedfordshire, Luton, UK
Hui Wei, Enjie Liu & Hongqing Yu
Communication University of Zhejiang, Hangzhou, China
Youbing Zhao

Authors

Hui Wei
View author publications
You can also search for this author in PubMed Google Scholar
Enjie Liu
View author publications
You can also search for this author in PubMed Google Scholar
Youbing Zhao
View author publications
You can also search for this author in PubMed Google Scholar
Hongqing Yu
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

University of Geneva, Geneva, Switzerland
Nadia Magnenat-Thalmann
University of Crete, Heraklion, Greece
Constantine Stephanidis
University of Macau, Macau, China
Enhua Wu
Swiss Federal Institute of Technology, Lausanne, Switzerland
Daniel Thalmann
Shanghai Jiao Tong University, Shanghai, China
Bin Sheng
University of Sydney, Sydney, Australia
Jinman Kim
University of Crete, Heraklion, Greece
George Papagiannakis
University of Calgary, Calgary, AB, Canada
Marina Gavrilova

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Wei, H., Liu, E., Zhao, Y., Yu, H. (2020). Efficient Non-fused Winograd on GPUs. In: Magnenat-Thalmann, N., et al. Advances in Computer Graphics. CGI 2020. Lecture Notes in Computer Science(), vol 12221. Springer, Cham. https://doi.org/10.1007/978-3-030-61864-3_35

Download citation

DOI: https://doi.org/10.1007/978-3-030-61864-3_35
Published: 18 October 2020
Publisher Name: Springer, Cham
Print ISBN: 978-3-030-61863-6
Online ISBN: 978-3-030-61864-3
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics

Efficient Non-fused Winograd on GPUs

Abstract

Access this chapter

Subscribe and save

Buy Now

Similar content being viewed by others

Efficient and portable Winograd convolutions for multi-core processors

Im2win: An Efficient Convolution Paradigm on GPU

Computing large 2D convolutions on GPU efficiently with the im2tensor algorithm

References

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Publish with us

Subscribe and save

Buy Now

Navigation

Efficient Non-fused Winograd on GPUs

Abstract

Access this chapter

Subscribe and save

Buy Now

Similar content being viewed by others

Efficient and portable Winograd convolutions for multi-core processors

Im2win: An Efficient Convolution Paradigm on GPU

Computing large 2D convolutions on GPU efficiently with the im2tensor algorithm

References

Author information

Authors and Affiliations

Editor information

Editors and Affiliations

Rights and permissions

Copyright information

About this paper

Cite this paper

Download citation

Share this paper

Publish with us

Search

Navigation