More Web Proxy on the site http://driver.im/

research-article

Open access

EECache: A Comprehensive Study on the Architectural Design for Energy-Efficient Last-Level Caches in Chip Multiprocessors

Authors:

Hsiang-Yun Cheng,

Narges Shahidi,

Mary Jane Irwin,

Mahmut Kandemir,

Yuan XieAuthors Info & Claims

ACM Transactions on Architecture and Code Optimization (TACO), Volume 12, Issue 2

Article No.: 17, Pages 1 - 22

https://doi.org/10.1145/2756552

Published: 08 July 2015 Publication History

Abstract

Power management for large last-level caches (LLCs) is important in chip multiprocessors (CMPs), as the leakage power of LLCs accounts for a significant fraction of the limited on-chip power budget. Since not all workloads running on CMPs need the entire cache, portions of a large, shared LLC can be disabled to save energy. In this article, we explore different design choices, from circuit-level cache organization to microarchitectural management policies, to propose a low-overhead runtime mechanism for energy reduction in the large, shared LLC. We first introduce a slice-based cache organization that can shut down parts of the shared LLC with minimal circuit overhead. Based on this slice-based organization, part of the shared LLC can be turned off according to the spatial and temporal cache access behavior captured by low-overhead sampling-based hardware. In order to eliminate the performance penalties caused by flushing data before powering off a cache slice, we propose data migration policies to prevent the loss of useful data in the LLC. Results show that our energy-efficient cache design (EECache) provides 14.1% energy savings at only 1.2% performance degradation and consumes negligible hardware overhead compared to prior work.

References

[1]

B. Ahsan, L. Ndreu, I Sideris, Y. Sazeides, S. Idgunji, and E. Ozer. 2011. Eliminating energy of same-content- cell-columns of on-chip SRAM arrays. In Proceedings of the 2011 International Symposium on Low Power Electronics and Design (ISLPED). 181--186.

Digital Library

[2]

D. H. Albonesi. 1999. Selective cache ways: On-demand cache resource allocation. In Proceedings of the 32nd Annual ACM/IEEE International Symposium on Microarchitecture (MICRO). 248--259.

Digital Library

[3]

J. Allred, S. Roy, and K. Chakraborty. 2012. Designing for dark silicon: A methodological perspective on energy efficient systems. In Proceedings of the 2012 International Symposium on Low Power Electronics and Design (ISLPED). 255--260.

Digital Library

[4]

R. Balasubramonian, N. P. Jouppi, and N. Muralimanohar. 2011. Multi-Core Cache Hierarchies. Morgan and Claypool.

Digital Library

[5]

A. Bardine, M. Comparetti, P. Foglia, G. Gabrielli, C. A. Prete, and Per Stenström. 2008. Leveraging data promotion for low power D-NUCA caches. In Proceedings of the 11th EUROMICRO Conference on Digital System Design Architectures, Methods and Tools (DSD). 307--316.

Digital Library

[6]

A. Bardine, M. Comparetti, P. Foglia, G. Gabrielli, and C. A. Prete. 2010. Way adaptable D-NUCA caches. Int. J. High Perform. Syst. Archit. 2, 3/4 (Aug. 2010), 215--228.

Digital Library

[7]

A. Bardine, M. Comparetti, P. Foglia, and C. A. Prete. 2014. Evaluation of leakage reduction alternatives for deep submicron dynamic nonuniform cache architecture caches. IEEE Trans. Very Large Scale Integr. Syst. 22, 1 (Jan. 2014), 185--190.

Digital Library

[8]

A. Basu, D. R. Hower, M. D. Hill, and M. M. Swift. 2013. FreshCache: Statically and dynamically exploiting dataless ways. In Proceedings of the 2013 IEEE 31st International Conference on Computer Design (ICCD). 286--293.

[9]

N. Binkert, B. Beckmann, G. Black, S. K. Reinhardt, A. Saidi, A. Basu, J. Hestness, D. R. Hower, T. Krishna, S. Sardashti, R. Sen, K. Sewell, M. Shoaib, N. Vaish, M. D. Hill, and D. A. Wood. 2011. The gem5 simulator. SIGARCH Comput. Archit. News 39, 2 (Aug. 2011), 1--7.

Digital Library

[10]

K. Chandrasekar, C. Weis, Y. Li, S. Goossens, M. Jung, O. Naji, B. Akesson, N. Wehn, and K. Goossens. 2014. DRAMPower: Open-Source DRAM Power and Energy Estimation Tool. Retrieved from http://www.drampower.info/.

[11]

Y.-T. Chen, J. Cong, H. Huang, B. Liu, C. Liu, M. Potkonjak, and G. Reinman. 2012. Dynamically reconfigurable hybrid cache: An energy-efficient last-level cache design. In Design, Automation & Test in Europe Conference & Exhibition (DATE). 45--50.

Digital Library

[12]

H.-Y. Cheng, M. Poremba, N. Shahidi, I. Stalev, M. J. Irwin, M. Kandemir, J. Sampson, and Y. Xie. 2014. EECache: Exploiting design choices in energy-efficient last-level caches for chip multiprocessors. In Proceedings of the 2014 International Symposium on Low Power Electronics and Design (ISLPED). 303--306.

Digital Library

[13]

J. Cong, K. Gururaj, H. Huang, C. Liu, G. Reinman, and Yi Zou. 2011. An energy-efficient adaptive hybrid cache. In Proceedings of the 2011 International Symposium on Low Power Electronics and Design (ISLPED). 67--72.

Digital Library

[14]

H. Esmaeilzadeh, E. Blem, R. St. Amant, K. Sankaralingam, and D. Burger. 2011. Dark silicon and the end of multicore scaling. In Proceedings of the 38th Annual International Symposium on Computer Architecture (ISCA). 365--376.

Digital Library

[15]

K. Flautner, N. S. Kim, S. Martin, D. Blaauw, and T. Mudge. 2002. Drowsy caches: Simple techniques for reducing leakage power. In Proceedings of the 29th Annual International Symposium on Computer Architecture (ISCA). 148--157.

Digital Library

[16]

P. Foglia and M. Comparetti. 2014. A workload independent energy reduction strategy for D-NUCA caches. J. Supercomput. 68, 1 (April 2014), 157--182.

Digital Library

[17]

M. Gebhart, J. Hestness, E. Fatehi, P. Gratz, and S. W. Keckler. 2009. Running PARSEC 2.1 on M5. Technical Report. The University of Texas at Austin, Department of Computer Science.

[18]

V. George, S. Jahagirdar, C. Tong, K. Smits, S. Damaraju, S. Siers, V. Naydenov, T. Khondker, S. Sarkar, and P. Singh. 2007. Penryn: 45-nm next generation Intel core 2 processor. In Proceedings of the 2007 IEEE Asian Solid-State Circuits Conference (ASSCC). 14--17.

[19]

H. R. Ghasemi, S. C. Draper, and N. S. Kim. 2011. Low-voltage on-chip cache architecture using heterogeneous cell sizes for high-performance processors. In Proceedings of the 2011 IEEE 17th International Symposium on High Performance Computer Architecture (HPCA). 38--49.

Digital Library

[20]

M. Ghosh, E. Ozer, S. Ford, S. Biles, and H.-H. S. Lee. 2009. Way guard: A segmented counting bloom filter approach to reducing energy for set-associative caches. In Proceedings of the 14th ACM/IEEE International Symposium on Low Power Electronics and Design (ISLPED). 165--170.

Digital Library

[21]

N. Goulding-Hotta, J. Sampson, Qiaoshi Zheng, V. Bhatt, J. Auricchio, S. Swanson, and M. B. Taylor. 2012. GreenDroid: An architecture for the Dark Silicon Age. In Proceedings of the 17th Asia and South Pacific Design Automation Conference (ASP-DAC). 100--105.

[22]

J. Hu, J. K. John, and S. Wang. 2008. Thermal-aware subarrayed data cache microarchitectures. Int. J. Intelligent Control Syst. 13, 4 (Dec. 2008), 251--263.

[23]

A. Jadidi, M. Arjomand, and H. Sarbazi-Azad. 2011. High-endurance and performance-efficient design of hybrid cache architectures through adaptive line replacement. In Proceedings of the 2011 International Symposium on Low Power Electronics and Design (ISLPED). 79--84.

Digital Library

[24]

J. K. John, J. S. Hu, and S. G. Ziavras. 2005. Optimizing the thermal behavior of subarrayed data caches. In Proceedings of the 2005 IEEE International Conference on Computer Design: VLSI in Computers and Processors (ICCD). 625--630.

Digital Library

[25]

D. Kadjo, H. Kim, P. Gratz, J. Hu, and R. Ayoub. 2013. Power gating with block migration in chip-multiprocessor last-level caches. In Proceedings of the 2013 IEEE International Conference on Computer Design (ICCD). 93--99.

[26]

S. Kaxiras, Z. Hu, and M. Martonosi. 2001. Cache decay: Exploiting generational behavior to reduce cache leakage power. In Proceedings of the 28th Annual International Symposium on Computer Architecture (ISCA). 240--251.

Digital Library

[27]

C. Kim, D. Burger, and S.W. Keckler. 2003. Nonuniform cache architectures for wire-delay dominated on-chip caches. IEEE Micro 23, 6 (Nov. 2003), 99--107.

Digital Library

[28]

H. Kim, J.-H. Ahn, and J. Kim. 2010. Replication-aware leakage management in chip multi-processors with private L2 caches. In Proceedings of the 2010 ACM/IEEE International Symposium on Low-Power Electronics and Design (ISLPED). 135--140.

Digital Library

[29]

J. Kim, N. Hardavellas, K. Mai, B. Falsafi, and J. C. Hoe. 2007. Multi-bit error tolerant caches using two-dimensional error coding. In Proceedings of the 40th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). 197--209.

Digital Library

[30]

N. A. Kurd, S. Bhamidipati, C. Mozak, J. L. Miller, T. M. Wilson, M. Nemani, and M. Chowdhury. 2010. Westmere: A family of 32nm IA processors. In Proceedings of the 2010 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC). 96--97.

[31]

S. Li, J.-H. Ahn, R. D. Strong, J. B. Brockman, D. M. Tullsen, and N. P. Jouppi. 2009. McPAT: An integrated power, area, and timing modeling framework for multicore and manycore architectures. In Proceedings of the 42nd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). 469--480.

Digital Library

[32]

S. Mittal, Z. Zhang, and J. S. Vetter. 2013. FlexiWay: A cache energy saving technique using fine-grained cache reconfiguration. In Proceedings of the 2013 IEEE International Conference on Computer Design (ICCD). 100--107.

[33]

N. Muralimanohar, R. Balasubramonia, and N. P. Jouppi. 2009. CACTI 6.0: A Tool to Model Large Caches. Technical Report. HP Lab.

[34]

S. Naffziger, B. Stackhouse, T. Grutkowski, D. Josephson, J. Desai, E. Alon, and M. Horowitz. 2006. The implementation of a 2-core, multi-threaded itanium family processor. IEEE J. Solid-State Circuits 41, 1, 197--209.

[35]

H. Park, S. Yoo, and S. Lee. 2011. A novel tag access scheme for low power L2 cache. In Proceedings of the 2011 Design, Automation Test in Europe Conference Exhibition (DATE). 1--6.

[36]

H. Park, S. Yoo, and S. Lee. 2012. A multistep tag comparison method for a low-power L2 cache. IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 31, 4, 559--572.

Digital Library

[37]

M. Powell, S.-H. Yang, B. Falsafi, K. Roy, and T. N. Vijaykumar. 2000. Gated-Vdd: A circuit technique to reduce leakage in deep-submicron cache memories. In Proceedings of the 2000 International Symposium on Low Power Electronics and Design (ISLPED). 90--95.

Digital Library

[38]

M. K. Qureshi, D. N. Lynch, O. Mutlu, and Y. N. Patt. 2006. A case for MLP-aware cache replacement. In Proceedings of the 33rd Annual International Symposium on Computer Architecture (ISCA). 167--178.

Digital Library

[39]

S. Rusu, S. Tam, H. Muljono, D. Ayers, and J. Chang. 2006. A dual-core multi-threaded Xeon processor with 16MB L3 cache. In Proceedings of the 2006 IEEE International Solid-State Circuits Conference - Digest of Technical Papers (ISSCC). 315--324.

[40]

S. Rusu, S. Tam, H. Muljono, J. Stinson, D. Ayers, J. Chang, R. Varada, M. Ratta, and S. Kottapalli. 2009. A 45nm 8-core enterprise Xeon processor. In Proceedings of the 2009 IEEE International Solid-State Circuits Conference - Digest of Technical Papers (ISSCC). 56--57.

[41]

K. T. Sundararajan, V. Porpodas, T. M. Jones, N. P. Topham, and B. Franke. 2012. Cooperative partitioning: Energy-efficient cache partitioning for high-performance CMPs. In Proceedings of the 2012 IEEE 18th International Symposium on High Performance Computer Architecture (HPCA). 1--12.

Digital Library

[42]

K. Swaminathan, E. Kultursay, V. Saripalli, V. Narayanan, and M. Kandemir. 2012. Design space exploration of workload-specific last-level caches. In Proceedings of the 2012 ACM/IEEE International Symposium on Low Power Electronics and Design (ISLPED). 243--248.

Digital Library

[43]

M. B. Taylor. 2012. Is dark silicon useful? Harnessing the four horsemen of the coming dark silicon apocalypse. In Proceedings of the 49th ACM/EDAC/IEEE Design Automation Conference (DAC). 1131--1136.

Digital Library

[44]

S. Thoziyoor, N. Muralimanohar, and N. P. Jouppi. 2007. CACTI 5.0. Technical Report. HP Lab.

[45]

H.-J. Tsai, C.-C. Chen, K.-H. Yang, T.-C. Yang, L.-Y. Huang, C.-H. Chung, M.-F. Chang, and T.-F. Chen. 2014. Leveraging data lifetime for energy-aware last level non-volatile SRAM caches using redundant store elimination. In Proceedings of the 51st ACM/EDAC/IEEE Design Automation Conference (DAC). 1--6.

Digital Library

[46]

G. Venkatesh, J. Sampson, N. Goulding, S. Garcia, V. Bryksin, J. Lugo-Martinez, S. Swanson, and M. B. Taylor. 2010. Conservation cores: Reducing the energy of mature computations. In Proceedings of the Fifteenth Edition of ASPLOS on Architectural Support for Programming Languages and Operating Systems (ASPLOS). 205--218.

Digital Library

[47]

Y. Wang, S. Roy, and N. Ranganathan. 2012. Run-time power-gating in caches of GPUs for leakage energy savings. In 2012 Design, Automation Test in Europe Conference Exhibition (DATE). 300--303.

Digital Library

[48]

D. Wendel, R. Kalla, R. Cargoni, J. Clables, J. Friedrich, R. Frech, J. Kahle, B. Sinharoy, W. Starke, S. Taylor, S. Weitzel, S. G. Chu, S. Islam, and V. Zyuban. 2010. The implementation of POWER7: A highly parallel and scalable multi-core high-end server processor. In Proceedings of the 2010 IEEE International Solid-State Circuits Conference - Digest of Technical Papers (ISSCC). 102--103.

[49]

C. Wilkerson, A. R. Alameldeen, Z. Chishti, W. Wu, D. Somasekhar, and S.-L. Lu. 2010. Reducing cache power with low-cost, multi-bit error-correcting codes. In Proceedings of the 37th Annual International Symposium on Computer Architecture (ISCA). 83--93.

Digital Library

[50]

S.-H. Yang, M. D. Powell, B. Falsafi, and T. N. Vijaykumar. 2002. Exploiting choice in resizable cache design to optimize deep-submicron processor energy-delay. In Proceedings of the 8th International Symposium on High Performance Computer Architecture (HPCA). 151--161.

Digital Library

Cited By

Ravipati DGoel RSanten VAmrouch HPanda P(2024)CAPE: Criticality-Aware Performance and Energy Optimization Policy for NCFET-Based CachesIEEE Transactions on Computers10.1109/TC.2024.345773473:12(2830-2843)Online publication date: Dec-2024
https://doi.org/10.1109/TC.2024.3457734
Biswas ATyagi A(2023)Huffman Cache Trails2023 IEEE International Symposium on Smart Electronic Systems (iSES)10.1109/iSES58672.2023.00063(277-282)Online publication date: 18-Dec-2023
https://doi.org/10.1109/iSES58672.2023.00063
Ravipati Dvan Santen VSalamin SAmrouch HPanda P(2023)Performance and Energy Studies on NC-FinFET Cache-Based Systems With FN-McPATIEEE Transactions on Very Large Scale Integration (VLSI) Systems10.1109/TVLSI.2023.328510531:9(1280-1293)Online publication date: 1-Sep-2023
https://dl.acm.org/doi/10.1109/TVLSI.2023.3285105
Show More Cited By

Index Terms

EECache: A Comprehensive Study on the Architectural Design for Energy-Efficient Last-Level Caches in Chip Multiprocessors
1. Hardware
  1. Integrated circuits
    1. Semiconductor memory

Recommendations

EECache: exploiting design choices in energy-efficient last-level caches for chip multiprocessors
ISLPED '14: Proceedings of the 2014 international symposium on Low power electronics and design

Power management for large last-level caches (LLCs) is important in chip-multiprocessors (CMPs), as the leakage power of LLCs accounts for a significant fraction of the limited on-chip power budget. Since not all workloads need the entire cache, ...
The ZCache: Decoupling Ways and Associativity
MICRO '43: Proceedings of the 2010 43rd Annual IEEE/ACM International Symposium on Microarchitecture

The ever-increasing importance of main memory latency and bandwidth is pushing CMPs towards caches with higher capacity and associativity. Associativity is typically improved by increasing the number of ways. This reduces conflict misses, but increases ...
Increasing hardware data prefetching performance using the second-level cache

Techniques to reduce or tolerate large memory latencies are critical for achieving high processor performance. Hardware data prefetching is one of the most heavily studied solutions, but it is essentially applied to first-level caches where it can ...

Comments

Please enable JavaScript to view thecomments powered by Disqus.

Information & Contributors

Information

Published In

cover image ACM Transactions on Architecture and Code Optimization

ACM Transactions on Architecture and Code Optimization Volume 12, Issue 2

July 2015

410 pages

ISSN:1544-3566

EISSN:1544-3973

DOI:10.1145/2775085

Editor:
Koen De Bosschere
Ghent University

Issue’s Table of Contents

Copyright © 2015 ACM.

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]

Publisher

Association for Computing Machinery

New York, NY, United States

Publication History

Published: 08 July 2015

Accepted: 01 April 2015

Revised: 01 January 2015

Received: 01 October 2014

Published in TACO Volume 12, Issue 2

Permissions

Request permissions for this article.

Request Permissions

Check for updates

Author Tags

Qualifiers

Research-article
Research
Refereed

Funding Sources

Contributors

Other Metrics

View Article Metrics

Bibliometrics & Citations

Bibliometrics

Article Metrics

9
Total Citations
View Citations
945
Total Downloads

Downloads (Last 12 months)94
Downloads (Last 6 weeks)10

Reflects downloads up to 03 Mar 2025

Other Metrics

View Author Metrics

Citations

Cited By

Ravipati DGoel RSanten VAmrouch HPanda P(2024)CAPE: Criticality-Aware Performance and Energy Optimization Policy for NCFET-Based CachesIEEE Transactions on Computers10.1109/TC.2024.345773473:12(2830-2843)Online publication date: Dec-2024
https://doi.org/10.1109/TC.2024.3457734
Biswas ATyagi A(2023)Huffman Cache Trails2023 IEEE International Symposium on Smart Electronic Systems (iSES)10.1109/iSES58672.2023.00063(277-282)Online publication date: 18-Dec-2023
https://doi.org/10.1109/iSES58672.2023.00063
Ravipati Dvan Santen VSalamin SAmrouch HPanda P(2023)Performance and Energy Studies on NC-FinFET Cache-Based Systems With FN-McPATIEEE Transactions on Very Large Scale Integration (VLSI) Systems10.1109/TVLSI.2023.328510531:9(1280-1293)Online publication date: 1-Sep-2023
https://dl.acm.org/doi/10.1109/TVLSI.2023.3285105
Ge FWang LWu NZhou F(2019)A Cache Fill and Migration Policy for STT-RAM-Based Multi-Level Hybrid Cache in 3D CMPsElectronics10.3390/electronics80606398:6(639)Online publication date: 6-Jun-2019
https://doi.org/10.3390/electronics8060639
Ghaemi SAhmadpour IArdebili MFarbeh H(2019)Sleepy-LRUThe Journal of Supercomputing10.1007/s11227-019-02758-075:7(3945-3974)Online publication date: 1-Jul-2019
https://dl.acm.org/doi/10.1007/s11227-019-02758-0
Asad AOzturk OFathy MJahed-Motlagh M(2017)Optimization-based power and thermal management for dark silicon aware 3D chip multiprocessors using heterogeneous cache hierarchyMicroprocessors and Microsystems10.1016/j.micpro.2017.03.01151(76-98)Online publication date: Jun-2017
https://doi.org/10.1016/j.micpro.2017.03.011
Cataldo RKorol GFernandes RMatos DMarcon Cde Lima Monteiro DTorres FIndrusiak L(2016)Architectural exploration of last-level caches targeting homogeneous multicore systemsProceedings of the 29th Symposium on Integrated Circuits and Systems Design: Chip on the Mountains10.5555/3145862.3145876(1-6)Online publication date: 29-Aug-2016
https://dl.acm.org/doi/10.5555/3145862.3145876
Zhang ZJu LJia ZFanucci LTeich J(2016)Unified DRAM and NVM hybrid buffer cache architecture for reducing journaling overheadProceedings of the 2016 Conference on Design, Automation & Test in Europe10.5555/2971808.2972025(942-947)Online publication date: 14-Mar-2016
https://dl.acm.org/doi/10.5555/2971808.2972025
Cataldo RKorol GFernandes RMatos DMarcon C(2016)Architectural exploration of Last-Level Caches targeting homogeneous multicore systems2016 29th Symposium on Integrated Circuits and Systems Design (SBCCI)10.1109/SBCCI.2016.7724050(1-6)Online publication date: Aug-2016
https://doi.org/10.1109/SBCCI.2016.7724050

View Options

View options

PDF

View or Download as a PDF file.

eReader

View online with eReader.

Login options

Check if you have access through your login credentials or your institution to get full access on this article.

Full Access

Get this Article

Figures

Tables

Media

View Issue’s Table of Contents