RISC-V Ratified Specifications Library ==================== Welcome to the RISC-V Ratified Specifications Library. This site provides access to all ratified RISC-V specifications. ## [](#about-riscv)About RISC-V RISC-V (pronounced "risk-five") is an open standard instruction set architecture (ISA) based on established reduced instruction set computer (RISC) principles. These specifications are maintained by RISC-V International and represent the ratified, stable specifications for the RISC-V ecosystem. For more information, visit [riscv.org](https://riscv.org). Core Architecture ##### Unprivileged ISA **Version:** v20260120 January 2026 User-level instruction set and standard extensions. [HTML](../isa/unpriv/unpriv-index.html)[PDF](../isa/%5Fattachments/riscv-unprivileged.pdf) [More](https://riscv.atlassian.net/wiki/external/MzM1NWFmYmFjNzlkNGMyMDg0NThjM2ZlODVkZTA3MDE) ##### Privileged ISA **Version:** v20260120 January 2026 Privileged architecture, execution modes, and system control. [HTML](../isa/priv/priv-index.html)[PDF](../isa/%5Fattachments/riscv-privileged.pdf) [More](https://riscv.atlassian.net/wiki/external/NDdlNzQ3YTRiNWZkNDI2ZjlhY2M2YzFkMmExNjI5NTg) Profiles ##### RISC-V Profiles **Version:** v1.0 April 2023 Base profile definitions and guidance. [HTML](../rva20-rvi20-rva22/index.html)[PDF](../rva20-rvi20-rva22/%5Fattachments/RISC-V%5FProfiles.pdf) [More](https://riscv.atlassian.net/wiki/external/ODAwNTY1YzNmZDQ3NDg5N2FkMGQ5MTRjZWY1ZGY4Yjc) ##### RVA23 Profile **Version:** v1.0 October 2024 Application-class profile requirements. [HTML](../rva23/index.html)[PDF](../rva23/%5Fattachments/rva23-profile.pdf) [More](https://riscv.atlassian.net/wiki/external/MWRiOWY4ZjBhZjkzNDI2MTk4ZGMxZDhjNjZjOTMyNDk) ##### RVB23 Profile **Version:** v1.0 October 2024 Embedded and edge profile requirements. [HTML](../rvb23/index.html)[PDF](../rvb23/%5Fattachments/rvb23-profile.pdf) [More](https://riscv.atlassian.net/wiki/external/OTQ3N2Q3MmQwNDFmNDUxMDlkMzY0NTgyMjk5YTI3MzE) Platforms ##### Server Platform **Version:** v1.0 May 2026 Hardware and sofware capabilities that portable system software can rely on being present in a server platform. [HTML](../server-platform/index.html)[PDF](../server-platform/%5Fattachments/riscv-server-platform.pdf) [More](https://riscv.atlassian.net/wiki/external/YjgzMGRlYmQzZTEzNGQwZjhjZjBiZjUyMzlhODBiNDQ) Hardware ##### Advanced Interrupt Architecture **Version:** v1.0 June 2023 Interrupt architecture and related interfaces. [HTML](../aia/index.html)[PDF](../aia/%5Fattachments/riscv-interrupts.pdf) [More](https://riscv.atlassian.net/wiki/external/MTc1MTUwOTg5NTRkNDk0YjlkOGY1MDk5MTljZGU2MjA) ##### IOMMU **Version:** v20260222 February 2026 IOMMU architecture, registers, queues, and integration guidance. [HTML](../iommu/index.html)[PDF](../iommu/%5Fattachments/riscv-iommu.pdf) [More](https://riscv.atlassian.net/wiki/external/ZGY0YjgxZTZhOTY3NDhiZGJhMGZiNzAwMzk3ZmNkNDg) ##### Platform-Level Interrupt Controller **Version:** v1.0.0 February 2023 Interrupt controller behavior and programming model. [HTML](../plic/index.html)[PDF](../plic/%5Fattachments/riscv-plic.pdf) [More](https://riscv.atlassian.net/wiki/external/Mzg2YjQ2ZTBiODYyNDBkNDg2N2JmZmQ1ZTUwOWNmZWI) ##### Server SoC **Version:** v1.0 February 2025 Server-class SoC requirements and conventions. [HTML](../server-soc/index.html)[PDF](../server-soc/%5Fattachments/riscv-server-soc.pdf) [More](https://riscv.atlassian.net/wiki/external/NTg4MWJjYjJmZDkxNDdmNWFmMTdkZjYxYTk2MzFjOGI) Debug, Trace, and RAS ##### Efficient Trace for RISC-V **Version:** v2.0 June 2025 Trace architecture for efficient execution visibility. [HTML](../e-trace/index.html)[PDF](../e-trace/%5Fattachments/riscv-trace-spec.pdf) [More](https://riscv.atlassian.net/wiki/external/NjI1Y2M5ODdiYmUxNGE5MzkzYzdkNzNjNGE3NWFjNDA) ##### QoS Register Interface **Version:** v1.0 June 2024 QoS register interface for capacity and bandwidth control. [HTML](../cbqri/index.html)[PDF](../cbqri/%5Fattachments/riscv-cbqri.pdf) [More](https://riscv.atlassian.net/wiki/external/YmYxMjUzMmYyOTIxNGIzMmJiZjNiYWY0MGVhNjBkYTc) ##### RERI Architecture **Version:** v1.0 May 2024 Reliability, availability, and serviceability error records. [HTML](../ras-eri/index.html)[PDF](../ras-eri/%5Fattachments/riscv-reri.pdf) [More](https://riscv.atlassian.net/wiki/external/ODE1N2Q2OTJkY2U5NDk2MmI5Zjk2ZmJmYzFiNmI1OTg) ##### Debug Specification **Version:** v1.0 February 2025 Defines the interfaces for external debugging of RISC-V processors. [HTML](../debug/index.html)[PDF](../debug/%5Fattachments/riscv-debug-specification.pdf) [More](https://riscv.atlassian.net/wiki/external/YTljNWI1NDgxMzE2NDIzOGEyYjMzMWRmZmRjOGEwNmE) ##### Unformatted Trace & Data Encapsulation **Version:** v1.0 June 2024 Encapsulation format for emitted trace data. [HTML](../trace-encap/index.html)[PDF](../trace-encap/%5Fattachments/e-trace-encap.pdf) [More](https://riscv.atlassian.net/wiki/external/NDZjYTFmZjQ1YTRhNDIxMjljODViYWU1YzQwMjRlNjk) ##### N-Trace **Version:** v1.0 November 2024 Nexus-based trace data formatting and transport. [HTML](../nexus-trace/index.html)[PDF](../nexus-trace/%5Fattachments/RISC-V-N-Trace.pdf) [More](https://riscv.atlassian.net/wiki/external/ZTUxYTJmNWRkYzhmNDM4ZDg5OWNlMmQxODMyYjY3YjI) ##### Trace Connectors **Version:** v1.0 November 2024 Connector definitions for trace interoperability. [HTML](../trace-connectors/index.html)[PDF](../trace-connectors/%5Fattachments/RISC-V-Trace-Connectors.pdf) [More](https://riscv.atlassian.net/wiki/external/N2E0Mjk1ZDkzZTM0NDQ2YmJmM2IzZDIxNjBlZTkxNTI) ##### Trace Control Interface **Version:** v1.0 November 2024 Control and configuration interface for trace features. [HTML](../trace-control-interface/index.html)[PDF](../trace-control-interface/%5Fattachments/RISC-V-Trace-Control-Interface.pdf) [More](https://riscv.atlassian.net/wiki/external/ZmZkYWMzYmU5NzZjNDk0M2FkNmJmOWFhOWVjODk4MzE) Platform Software ##### Semihosting **Version:** v1.0 February 2025 Semihosting interface for development and debugging. [HTML](../semihosting/index.html)[PDF](../semihosting/%5Fattachments/riscv-semihosting.pdf) [More](https://riscv.atlassian.net/wiki/external/YjY0ZDBkZjg2M2E4NDYxM2FmMmMwMWM0YTgxZjRmYTc) ##### Boot and Runtime Services **Version:** v1.0 August 2025 Boot-time and runtime software interface requirements. [HTML](../brs/index.html)[PDF](../brs/%5Fattachments/riscv-brs-spec.pdf) [More](https://riscv.atlassian.net/wiki/external/YTc0N2FlZDJmNTdiNDJlYzhiMGIyODFlOGQ0YWQyZDg) ##### Functional Fixed Hardware **Version:** v1.0.1 October 2024 FFH interface definitions for platform integration. [HTML](../acpi-ffh/index.html)[PDF](../acpi-ffh/%5Fattachments/riscv-ffh.pdf) [More](https://riscv.atlassian.net/wiki/external/NGVhZDQ2MzY3MzZkNGE1MWIxYTYyMWM5N2Y2NTFjZjA) ##### IO Mapping Table **Version:** v1.0 March 2025 Standardized interrupt mapping table format. [HTML](../acpi-rimt/index.html)[PDF](../acpi-rimt/%5Fattachments/rimt-spec.pdf) [More](https://riscv.atlassian.net/wiki/external/ODdiOTYxZjIwOTJlNGM1NWJlM2Y4MDllMzZiODIyNTU) ##### Platform Management Interface **Version:** v1.0 July 2025 Platform management services and messaging definitions. [HTML](../rpmi/index.html)[PDF](../rpmi/%5Fattachments/riscv-rpmi.pdf) [More](https://riscv.atlassian.net/wiki/external/OWIyYzBlNTUyNWM5NDI4ZmIzMDliZjcxOGY1ZDYwZWY) ##### Supervisor Binary Interface **Version:** v3.0 July 2025 Standard interface between supervisor software and firmware. [HTML](../sbi/index.html)[PDF](../sbi/%5Fattachments/riscv-sbi.pdf) [More](https://riscv.atlassian.net/wiki/external/NDZmZTk3MDI0ZTA0NDVhOTg5Y2ZlZWI3NGNjZTlmYTQ) ##### UEFI PROTOCOL **Version:** v1.0.0 May 2022 Defines a software interface between an operating system and platform firmware. [HTML](../uefi/index.html)[PDF](../uefi/%5Fattachments/RISCV%5FUEFI%5FPROTOCOL-spec.pdf) [More](https://riscv.atlassian.net/wiki/external/NDEzOThmNmRjZmZhNDdhNzgxYmYwOWRmMTRkOTBlMTE) ##### QoS Controllers Table **Version:** v1.0 June 2026 Defines the ACPI interface for OSPM configuration and control of QoS controller features. [HTML](../rqsc/index.html)[PDF](../rqsc/%5Fattachments/riscv-rqsc.pdf) [More](https://riscv.atlassian.net/wiki/external/YmYyNWFkZWFkMTc1NDZjNjhkNjc5YmIwMmJkMzAxNTU) Application Enablement ##### Application Binary Interface **Version:** v1.0 November 2022 Application binary interface definitions for ELF tooling. [HTML](../abi/index.html)[PDF](../abi/%5Fattachments/riscv-abi.pdf) [More](https://riscv.atlassian.net/wiki/external/NWMwYmRhMDBkZGI5NDdiNjllZDRkODI2YTRhZDE4ZDc) ##### Vector C Intrinsic **Version:** v1.0 April 2025 Compiler intrinsics for vector extension programming. [HTML](../vector-c-intrinsics/index.html)[PDF](../vector-c-intrinsics/%5Fattachments/v-intrinsic-spec.pdf) [More](https://riscv.atlassian.net/wiki/external/YTMyZTE5NDkyNTM5NDZjMjliYTVmMjBjOTBhZWM4ODY) Hardware ==================== ##### Advanced Interrupt Architecture **Version:** v1.0 June 2023 Interrupt architecture and related interfaces. [HTML](../aia/index.html)[PDF](../aia/%5Fattachments/riscv-interrupts.pdf) [More](https://riscv.atlassian.net/wiki/external/MTc1MTUwOTg5NTRkNDk0YjlkOGY1MDk5MTljZGU2MjA) ##### IOMMU **Version:** v20260222 February 2026 IOMMU architecture, registers, queues, and integration guidance. [HTML](../iommu/index.html)[PDF](../iommu/%5Fattachments/riscv-iommu.pdf) [More](https://riscv.atlassian.net/wiki/external/ZGY0YjgxZTZhOTY3NDhiZGJhMGZiNzAwMzk3ZmNkNDg) ##### Platform-Level Interrupt Controller **Version:** v1.0.0 February 2023 Interrupt controller behavior and programming model. [HTML](../plic/index.html)[PDF](../plic/%5Fattachments/riscv-plic.pdf) [More](https://riscv.atlassian.net/wiki/external/Mzg2YjQ2ZTBiODYyNDBkNDg2N2JmZmQ1ZTUwOWNmZWI) ##### Server SoC **Version:** v1.0 February 2025 Server-class SoC requirements and conventions. [HTML](../server-soc/index.html)[PDF](../server-soc/%5Fattachments/riscv-server-soc.pdf) [More](https://riscv.atlassian.net/wiki/external/NTg4MWJjYjJmZDkxNDdmNWFmMTdkZjYxYTk2MzFjOGI) Profiles ==================== ##### RISC-V Profiles **Version:** v1.0 April 2023 Base profile definitions and guidance. [HTML](../rva20-rvi20-rva22/index.html)[PDF](../rva20-rvi20-rva22/%5Fattachments/RISC-V%5FProfiles.pdf) [More](https://riscv.atlassian.net/wiki/external/ODAwNTY1YzNmZDQ3NDg5N2FkMGQ5MTRjZWY1ZGY4Yjc) ##### RVA23 Profile **Version:** v1.0 October 2024 Application-class profile requirements. [HTML](../rva23/index.html)[PDF](../rva23/%5Fattachments/rva23-profile.pdf) [More](https://riscv.atlassian.net/wiki/external/MWRiOWY4ZjBhZjkzNDI2MTk4ZGMxZDhjNjZjOTMyNDk) ##### RVB23 Profile **Version:** v1.0 October 2024 Embedded and edge profile requirements. [HTML](../rvb23/index.html)[PDF](../rvb23/%5Fattachments/rvb23-profile.pdf) [More](https://riscv.atlassian.net/wiki/external/OTQ3N2Q3MmQwNDFmNDUxMDlkMzY0NTgyMjk5YTI3MzE) Platform Software ==================== ##### Semihosting **Version:** v1.0 February 2025 Semihosting interface for development and debugging. [HTML](../semihosting/index.html)[PDF](../semihosting/%5Fattachments/riscv-semihosting.pdf) [More](https://riscv.atlassian.net/wiki/external/YjY0ZDBkZjg2M2E4NDYxM2FmMmMwMWM0YTgxZjRmYTc) ##### Boot and Runtime Services **Version:** v1.0 August 2025 Boot-time and runtime software interface requirements. [HTML](../brs/index.html)[PDF](../brs/%5Fattachments/riscv-brs-spec.pdf) [More](https://riscv.atlassian.net/wiki/external/YTc0N2FlZDJmNTdiNDJlYzhiMGIyODFlOGQ0YWQyZDg) ##### Functional Fixed Hardware **Version:** v1.0.1 October 2024 FFH interface definitions for platform integration. [HTML](../acpi-ffh/index.html)[PDF](../acpi-ffh/%5Fattachments/riscv-ffh.pdf) [More](https://riscv.atlassian.net/wiki/external/NGVhZDQ2MzY3MzZkNGE1MWIxYTYyMWM5N2Y2NTFjZjA) ##### IO Mapping Table **Version:** v1.0 March 2025 Standardized interrupt mapping table format. [HTML](../acpi-rimt/index.html)[PDF](../acpi-rimt/%5Fattachments/rimt-spec.pdf) [More](https://riscv.atlassian.net/wiki/external/ODdiOTYxZjIwOTJlNGM1NWJlM2Y4MDllMzZiODIyNTU) ##### Platform Management Interface **Version:** v1.0 July 2025 Platform management services and messaging definitions. [HTML](../rpmi/index.html)[PDF](../rpmi/%5Fattachments/riscv-rpmi.pdf) [More](https://riscv.atlassian.net/wiki/external/OWIyYzBlNTUyNWM5NDI4ZmIzMDliZjcxOGY1ZDYwZWY) ##### Supervisor Binary Interface **Version:** v3.0 July 2025 Standard interface between supervisor software and firmware. [HTML](../sbi/index.html)[PDF](../sbi/%5Fattachments/riscv-sbi.pdf) [More](https://riscv.atlassian.net/wiki/external/NDZmZTk3MDI0ZTA0NDVhOTg5Y2ZlZWI3NGNjZTlmYTQ) ##### UEFI PROTOCOL **Version:** v1.0.0 May 2022 Defines a software interface between an operating system and platform firmware. [HTML](../uefi/index.html)[PDF](../uefi/%5Fattachments/RISCV%5FUEFI%5FPROTOCOL-spec.pdf) [More](https://riscv.atlassian.net/wiki/external/NDEzOThmNmRjZmZhNDdhNzgxYmYwOWRmMTRkOTBlMTE) ##### QoS Controllers Table **Version:** v1.0 June 2026 Defines the ACPI interface for OSPM configuration and control of QoS controller features. [HTML](../rqsc/index.html)[PDF](../rqsc/%5Fattachments/riscv-rqsc.pdf) [More](https://riscv.atlassian.net/wiki/external/YmYyNWFkZWFkMTc1NDZjNjhkNjc5YmIwMmJkMzAxNTU) Application Enablement ==================== ##### Application Binary Interface **Version:** v1.0 November 2022 Application binary interface definitions for ELF tooling. [HTML](../abi/index.html)[PDF](../abi/%5Fattachments/riscv-abi.pdf) [More](https://riscv.atlassian.net/wiki/external/NWMwYmRhMDBkZGI5NDdiNjllZDRkODI2YTRhZDE4ZDc) ##### Vector C Intrinsic **Version:** v1.0 April 2025 Compiler intrinsics for vector extension programming. [HTML](../vector-c-intrinsics/index.html)[PDF](../vector-c-intrinsics/%5Fattachments/v-intrinsic-spec.pdf) [More](https://riscv.atlassian.net/wiki/external/YTMyZTE5NDkyNTM5NDZjMjliYTVmMjBjOTBhZWM4ODY) Debug, Trace, and RAS ==================== ##### Efficient Trace for RISC-V **Version:** v2.0 June 2025 Trace architecture for efficient execution visibility. [HTML](../e-trace/index.html)[PDF](../e-trace/%5Fattachments/riscv-trace-spec.pdf) [More](https://riscv.atlassian.net/wiki/external/NjI1Y2M5ODdiYmUxNGE5MzkzYzdkNzNjNGE3NWFjNDA) ##### QoS Register Interface **Version:** v1.0 June 2024 QoS register interface for capacity and bandwidth control. [HTML](../cbqri/index.html)[PDF](../cbqri/%5Fattachments/riscv-cbqri.pdf) [More](https://riscv.atlassian.net/wiki/external/YmYxMjUzMmYyOTIxNGIzMmJiZjNiYWY0MGVhNjBkYTc) ##### RERI Architecture **Version:** v1.0 May 2024 Reliability, availability, and serviceability error records. [HTML](../ras-eri/index.html)[PDF](../ras-eri/%5Fattachments/riscv-reri.pdf) [More](https://riscv.atlassian.net/wiki/external/ODE1N2Q2OTJkY2U5NDk2MmI5Zjk2ZmJmYzFiNmI1OTg) ##### Debug Specification **Version:** v1.0 February 2025 Defines the interfaces for external debugging of RISC-V processors. [HTML](../debug/index.html)[PDF](../debug/%5Fattachments/riscv-debug-specification.pdf) [More](https://riscv.atlassian.net/wiki/external/YTljNWI1NDgxMzE2NDIzOGEyYjMzMWRmZmRjOGEwNmE) ##### Unformatted Trace & Data Encapsulation **Version:** v1.0 June 2024 Encapsulation format for emitted trace data. [HTML](../trace-encap/index.html)[PDF](../trace-encap/%5Fattachments/e-trace-encap.pdf) [More](https://riscv.atlassian.net/wiki/external/NDZjYTFmZjQ1YTRhNDIxMjljODViYWU1YzQwMjRlNjk) ##### N-Trace **Version:** v1.0 November 2024 Nexus-based trace data formatting and transport. [HTML](../nexus-trace/index.html)[PDF](../nexus-trace/%5Fattachments/RISC-V-N-Trace.pdf) [More](https://riscv.atlassian.net/wiki/external/ZTUxYTJmNWRkYzhmNDM4ZDg5OWNlMmQxODMyYjY3YjI) ##### Trace Connectors **Version:** v1.0 November 2024 Connector definitions for trace interoperability. [HTML](../trace-connectors/index.html)[PDF](../trace-connectors/%5Fattachments/RISC-V-Trace-Connectors.pdf) [More](https://riscv.atlassian.net/wiki/external/N2E0Mjk1ZDkzZTM0NDQ2YmJmM2IzZDIxNjBlZTkxNTI) ##### Trace Control Interface **Version:** v1.0 November 2024 Control and configuration interface for trace features. [HTML](../trace-control-interface/index.html)[PDF](../trace-control-interface/%5Fattachments/RISC-V-Trace-Control-Interface.pdf) [More](https://riscv.atlassian.net/wiki/external/ZmZkYWMzYmU5NzZjNDk0M2FkNmJmOWFhOWVjODk4MzE) Platforms ==================== ##### Server Platform **Version:** v1.0 May 2026 Hardware and sofware capabilities that portable system software can rely on being present in a server platform. [HTML](../server-platform/index.html)[PDF](../server-platform/%5Fattachments/riscv-server-platform.pdf) [More](https://riscv.atlassian.net/wiki/external/YjgzMGRlYmQzZTEzNGQwZjhjZjBiZjUyMzlhODBiNDQ) Specifications Under Development ==================== RISC-V Instruction Set Architecture (ISA) Manuals ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-instruction-set-architecture-isa-manuals)RISC-V Instruction Set Architecture (ISA) Manuals RISC-V ISA Manuals * [Volume I: Unprivileged Architecture](unpriv/unpriv-index.html) * [Volume II: Privileged Architecture](priv/priv-index.html) * [Bibliography](biblio/bibliography.html) Preamble ==================== _This document is released under a Creative Commons Attribution 4.0 International License._ _This document is a derivative of the RISC-V privileged specification version 1.9.1 released under following license: ©2010-2017 Andrew Waterman, Yunsup Lee, Rimas Avižienis, David Patterson, Krste Asanović. Creative Commons Attribution 4.0 International License._ ## [](#contributors)Contributors _Contributors to all versions of the spec in alphabetical order (please contact editors to suggest corrections): Krste Asanović, Peter Ashenden, Rimas Avižienis, Jacob Bachmeyer, Allen J. Baum, Jonathan Behrens, Paolo Bonzini, Ruslan Bukin, Christopher Celio, Chuanhua Chang, David Chisnall, Anthony Coulter, Palmer Dabbelt, Monte Dalrymple, Paul Donahue, Greg Favor, Dennis Ferguson, Marc Gauthier, Andy Glew, Gary Guo, Mike Frysinger, John Hauser, David Horner, Olof Johansson, David Kruckemyer, Yunsup Lee, Daniel Lustig, Andrew Lutomirski, Martin Maas, Prashanth Mundkur, Jonathan Neuschäfer, Rishiyur Nikhil, Stefan O’Rear, Albert Ou, John Ousterhout, David Patterson, Dmitri Pavlov, Kade Phillips, Josh Scheid, Colin Schmidt, Michael Taylor, Wesley Terpstra, Matt Thomas, Tommy Thorn, Ray VanDeWalker, Megan Wachs, Steve Wallach, Andrew Waterman, Claire Wolf, Adam Zabrocki, and Reinoud Zandijk._ Please contact RISC-V International to suggest corrections. The RISC-V Instruction Set Manual, Volume II ==================== ![RISCV](../../../common/_images/risc-v_logo.svg) ## [](#the-risc-v-instruction-set-manual-volume-ii)The RISC-V Instruction Set Manual, Volume II ## [](#privileged-architecture)Privileged Architecture Version 20260120: Official Release Preamble ==================== _This document is a derivative of “The RISC-V Instruction Set Manual, Volume I: User-Level ISA Version 2.1” released under the following license: ©2010-2017 Andrew Waterman, Yunsup Lee, David Patterson, Krste Asanović. Creative Commons Attribution 4.0 International License. Please cite as: “The RISC-V Instruction Set Manual, Volume I: User-Level ISA, Document Version 20191214-draft”, Editors Andrew Waterman and Krste Asanović, RISC-V Foundation, December 2019._ _This document is released under a Creative Commons Attribution 4.0 International License._ ## [](#contributors)Contributors _Contributors to all versions of the spec in alphabetical order (please contact editors to suggest corrections): Arvind, Krste Asanović, Derek Atkins, Rimas Avižienis, Jacob Bachmeyer, Christopher F. Batten, Allen J. Baum, Scott Beamer, Abel Bernabeu, Hans Boehm, Alex Bradbury, Preston Briggs, Christopher Celio, Chuanhua Chang, David Chisnall, Paul Clayton, Palmer Dabbelt, L Peter Deutsch, Ken Dockser, Paul Donahue, Aaron Durbin, Roger Espasa, Greg Favor, Shaked Flur, Stefan Freudenberger, Marc Gauthier, Andy Glew, Jan Gray, Gianluca Guida, Michael Hamburg, John Hauser, Christian Herber, David Horner, Bruce Hoult, Bill Huffman, John Ingalls, Alexandre Joannou, Olof Johansson, Ben Keller, David Kruckemyer, Tariq Kurd, Yunsup Lee, Paul Loewenstein, Daniel Lustig, Yatin Manerkar, Luc Maranget, Ben Marshall, Margaret Martonosi, Phil McCoy, Nathan Menhorn, Christoph Müllner, Joseph Myers, Vijayanand Nagarajan, Torbjørn Viem Ness, Rishiyur Nikhil, Jonas Oberhauser, Stefan O’Rear, Markku-Juhani O. Saarinen, Albert Ou, John Ousterhout, Daniel Page, David Patterson, Christopher Pulte, Jose Renau, Susmit Sarkar, Josh Scheid, Colin Schmidt, Peter Sewell, Ved Shanbhogue, Brent Spinney, Brendan Sweeney, Michael Taylor, Wesley Terpstra, Matt Thomas, Tommy Thorn, Philipp Tomsich, Caroline Trippel, Ray VanDeWalker, Muralidaran Vijayaraghavan, Megan Wachs, Paul Wamsley, Andrew Waterman, Robert Watson, David Weaver, Derek Williams, Claire Wolf, Andrew Wright, Reinoud Zandijk, and Sizhuo Zhang._ Please contact RISC-V International to suggest corrections. The RISC-V Instruction Set Manual, Volume I ==================== ![RISCV](../../../common/_images/risc-v_logo.svg) ## [](#the-risc-v-instruction-set-manual-volume-i)The RISC-V Instruction Set Manual, Volume I ## [](#unprivileged-architecture)Unprivileged Architecture Version 20260120: Official Release Bibliography ==================== ## [](#bibliography)Bibliography \[1\] _RISC-V ELF psABI Specification_, \\urlhttps://github.com/riscv/riscv-elf-psabi-doc/. \[2\] _The RISC-V Instruction Set Manual, Volume I: Base User-Level ISA Version 2.0_, UCB/EECS-2014-54, EECS Department, University of California, Berkeley, May 2014. \[3\] _The RISC-V Instruction Set Manual, Volume I: Base User-Level ISA_, UCB/EECS-2011-62, EECS Department, University of California, Berkeley, May 2011. \[4\] _ANSI/IEEE Std 754-2008, IEEE standard for floating-point arithmetic_, 2008. \[5\] D. A. Patterson and C. H. S'equin, "RISC I: A Reduced Instruction Set VLSI Computer" in _ISCA_. 1981, pp. 443-458. \[6\] K. M. G.H. and S. R. W. and P. D. A. and S. C. H., "The RISC II micro-architecture" in _Proceedings VLSI 83 Conference_. August 1983. \[7\] D. Ungar and R. Blau and P. Foley and D. Samples and D. Patterson, "Architecture of SOAR: Smalltalk on a RISC" in _ISCA_. Ann Arbor, MI:, 1984, pp. 188—​197. \[8\] D. D. Lee and S. I. Kong and M. D. H. a. . . . . . . . . . . . . . . . . . G. S. Taylor and D. A. Hodges and R. . . . . . . . . . . . . . . . . . H. Katz and D. A. Patterson, "A VLSI Chip Set for a Multiprocessor Workstation—​Part I: An RISC Microprocessor with Coprocessor Interface and Support for Symbolic Processing", _IEEE JSSC_, vol. 24, no. 6, December 1989\. pp. 1688—​1698. \[9\] H. Pan and B. Hindman and K. Asanovi'c, "Lithe: Enabling Efficient Composition of Parallel Libraries" in _Proceedings of the 1st USENIX Workshop on Hot Topics in Parallelism (HotPar\~'09)_. Berkeley, CA:, March 2009. \[10\] H. Pan and B. Hindman and K. Asanovi'c, "Composing Parallel Software Efficiently with Lithe" in _31st Conference on Programming Language Design and Implementation_. Toronto, Canada:, June 2010. \[11\] _RISC-V Assembly Programmer’s Manual_, \\urlhttps://github.com/riscv/riscv-asm-manual. \[12\] J. Tseng and K. Asanovi'c, "Energy-Efficient Register Access" in _Proc. of the 13th Symposium on Integrated Circuits and Systems Design_. Manaus, Brazil:, September 2000, pp. 377—​384. \[13\] _Selective Dual Path Execution_, University of Wisconsin - Madison, November 1996. \[14\] K. A. and A. T. and G. D. and C. B., "Dynamic Hammock Predication for Non-Predicated Instruction Set Architectures" in _Proceedings of the 1998 International Conference on Parallel Architectures and Compilation Techniques_, PACT '98\. Washington, DC, USA:, 1998. \[15\] K. Hyesoon and M. Onur and S. Jared and P. Y. N., "Wish Branches: Combining Conditional Branching and Predication for Adaptive Predicated Execution" in _Proceedings of the 38th annual IEEE/ACM International Symposium on Microarchitecture_, MICRO 38\. 2005, pp. 43—​54. \[16\] S. Balaram et al., "IBM POWER7 multicore server processor", _IBM Journal of Research and Development_, vol. 55, no. 3, 2011\. pp. 1—​1. \[17\] T. Marc and C. Jeffrey and C. Shailender and C. A. W. and T. S. Sheung, "The MAJC Architecture: A Synthesis of Parallelism and Scalability", _IEEE Micro_, vol. 20, no. 6, November 2000\. pp. 12—​25. \[18\] K. Gharachorloo and D. Lenoski and J. Laudon and P. Gibbons and A. Gupta and J. . . . . . . . . . . . . . . . . . Hennessy, "Memory Consistency and Event Ordering in Scalable Shared-Memory Multiprocessors" in _In Proceedings of the 17th Annual International Symposium on Computer Architecture_. 1990, pp. 15—​26. \[19\] R. Ravi and G. J. R., "Speculative lock elision: enabling highly concurrent multithreaded execution" in _Proceedings of the 34th annual ACM/IEEE International Symposium on Microarchitecture_, MICRO 34\. IEEE Computer Society, 2001, pp. 294—​305. \[20\] M. M. M. and S. M. L., "Simple, Fast, and Practical Non-Blocking and Blocking Concurrent Queue Algorithms" in _Proceedings of the Fifteenth Annual ACM Symposium on Principles of Distributed Computing_, PODC '96\. New York, NY, USA:, Association for Computing Machinery, 1996, pp. 267–275, Available: . \[21\] Roux and Pierre, "Innocuous Double Rounding of Basic Arithmetic Operations", _Journal of Formalized Reasoning_, vol. 7, no. 1, Nov 2014\. pp. 131-142, \[Online\]. Available: . \[22\] W. Buchholz, _Planning a computer system: Project Stretch_. McGraw-Hill Book Company, 1962. \[23\] G. M. Amdahl and G. A. Blaauw and F. P. B. Jr., "Architecture of the IBM System/360", _IBM Journal of R. \\& D._, vol. 8, no. 2, 1964. \[24\] Thornton and J. E., "Parallel Operation in the Control Data 6600" in _Proceedings of the October 27-29, 1964, Fall Joint Computer Conference, Part II: Very High Speed Computer Systems_, AFIPS '64 (Fall, part II). 1965, pp. 33—​40. \[25\] . \[26\] . \[27\] _SAIL ISA Specification Language_. \[Online\]. Available: \[28\] L. R. B and S. ZJ and Y. Y. Lisa and R. R. L and R. M. JB, "On permutation operations in cipher design" in _International Conference on Information Technology: Coding and Computing, 2004\. Proceedings. ITCC 2004._, vol. 2\. IEEE, 2004, pp. 569—​577. \[29\] NIST, _Secure Hash Standard (SHS)_, Federal Information Processing Standards Publication FIPS 180-4, August 2015\. \[Online\]. Available: \[30\] NIST, _Advanced Encryption Standard (AES)_, Federal Information Processing Standards Publication FIPS 197, November 2001\. \[Online\]. Available: \[31\] _GBT 32905-2016: SM3 Cryptographic Hash Algorithm_, Also GM/T 0004-2012\. Standardization Administration of China, August 2016\. \[Online\]. Available: \[32\] ISO/IEC, _IT Security techniques — Hash-functions — Part 3: Dedicated hash-functions_, ISO/IEC Standard 10118-3:2018, 2018. \[33\] M. O. Saarinen, _Lightweight SHA ISA_, \\urlhttps://github.com/mjosaarinen/lwsha\_isa, 03 2020. \[34\] _GB/T 32907-2016: SM4 Block Cipher Algorithm_, Also GM/T 0002-2012\. Standardization Administration of China, August 2016\. \[Online\]. Available: \[35\] ISO/IEC, _Information technology — Security techniques — Encryption algorithms — Part 3: Block ciphers. Amendment 2: SM4_, ISO/IEC Standard 18033-3:2010/DAmd 2 (en), 2018. \[36\] M. S. Turan and E. Barker and J. K. a. . . . K. A. McKay and M. L. Baish and M. Boyle, _Recommendation for the Entropy Sources Used for Random Bit Generation_, NIST Special Publication SP 800-90B, January 2018. \[37\] W. Killmann and W. Schindler, _A Proposal for: Functionality classes for random number generators_, AIS 20 / AIS 31, Version 2.0, English Translation, BSI, September 2011\. \[Online\]. Available: \[38\] E. Barker and J. Kelsey, _Recommendation for Random Number Generation Using Deterministic Random Bit Generators_, NIST Special Publication SP 800-90A Revision 1, June 2015. \[39\] E. Barker and J. Kelsey and A. R. a. . . . M. S. Turan and D. Buller and A. Kaufer, _Recommendation for Random Bit Generator (RBG) Constructions_, Draft NIST Special Publication SP 800-90C, March 2021. \[40\] NIST, _Submission Requirements and Evaluation Criteria for the Post-Quantum Cryptography Standardization Process_, Official Call for Proposals, National Institute for Standards and Technology, December 2016\. \[Online\]. Available: \[41\] _Information technology — Security techniques — Testing methods for the mitigation of non-invasive attack classes against cryptographic modules_, ISO/IEC 17825:2016, International Organization for Standardization, 2016. \[42\] M. O. Saarinen, _Lightweight AES ISA_, \\urlhttps://github.com/mjosaarinen/lwaes\_isa, 01 2020. \[43\] M. Ben and N. G. Richard and P. Dan and S. M. O. and W. Claire, "The design of scalar AES Instruction Set Extensions for RISC-V", _IACR Transactions on Cryptographic Hardware and Embedded Systems_, vol. 2021, no. 1, Dec. 2020\. pp. 109-136, \[Online\]. Available: . \[44\] _XCrypto: a cryptographic ISE for RISC-V_, 1.0.0, 2019\. \[Online\]. Available: \[45\] M. Dworkin, _Recommendation for Block Cipher Modes of Operation: Galois/Counter Mode (GCM) and GMAC_, NIST Special Publication SP 800-38D, November 2007\. \[Online\]. Available: \[46\] NIST, _SHA-3 Standard: Permutation-Based Hash and Extendable-Output Functions_, Federal Information Processing Standards Publication FIPS 202, August 2015\. \[Online\]. Available: \[47\] B. Andrey et al., "PRESENT: An ultra-lightweight block cipher" in _International workshop on cryptographic hardware and embedded systems_. Springer, 2007, pp. 450—​466. \[48\] Z. Wentao and B. Zhenzhen and L. Dongdai and R. Vincent and Y. Bohan and V. Ingrid, "RECTANGLE: a bit-slice lightweight block cipher suitable for multiple platforms", _Science China Information Sciences_, vol. 58, no. 12, 2015\. pp. 1—​15. \[49\] B. Subhadeep and P. S. Kumar and P. Thomas and S. Yu and S. S. Meng and T. Yosuke, "GIFT: a small present" in _International Conference on Cryptographic Hardware and Embedded Systems_. Springer, 2017, pp. 321—​345. \[50\] S. Tomoyasu and M. Kazuhiko and M. Sumio and K. Eita, "TWINE: A Lightweight Block Cipher for Multiple Platforms" in _International Conference on Selected Areas in Cryptography_. Springer, 2012, pp. 339—​354. \[51\] B. Christof et al., "The SKINNY family of block ciphers and its low-latency variant MANTIS" in _Annual International Cryptology Conference_. Springer, 2016, pp. 123—​153. \[52\] B. Subhadeep et al., "Midori: A block cipher for low energy" in _International Conference on the Theory and Application of Cryptology and Information Security_. Springer, 2015, pp. 411—​436. \[53\] A. Kazumaro et al., "Camellia: A 128-bit block cipher suitable for multiple platforms—design andanalysis" in _International Workshop on Selected Areas in Cryptography_. Springer, 2000, pp. 39—​56. \[54\] K. Daesung et al., "New block cipher: ARIA" in _International Conference on Information Security and Cryptology_. Springer, 2003, pp. 432—​445. \[55\] M. O. Saarinen, _On Entropy and Bit Patterns of Ring Oscillator Jitter_, Preprint, February 2021\. \[Online\]. Available: \[56\] NIST and CCCS, _Implementation Guidance for FIPS 140-3 and the Cryptographic Module Validation Program_, CMVP, May 2021\. \[Online\]. Available: \[57\] NIST, _Security Requirements for Cryptographic Modules_, Federal Information Processing Standards Publication FIPS 140-3, March 2019\. \[Online\]. Available: \[58\] C. Criteria, _Common Methodology for Information Technology Security Evaluation: Evaluation methodology_, Specification: Version 3.1 Revision 5, April 2017\. \[Online\]. Available: \[59\] W. Killmann and W. Schindler, _A Proposal for: Functionality classes and evaluation methodology for true (physical) random number generators_, AIS 31, Version 3.1, English Translation, BSI, September 2001\. \[Online\]. Available: \[60\] \[61\] NSA/CSS, _Commercial National Security Algorithm Suite_, August 2015\. \[Online\]. Available: \[62\] R. Bardou and R. Focardi and Y. K. a. . . . L. Simionato and G. Steel and J. Tsay, "Efficient Padding Oracle Attacks on Cryptographic Hardware" in _Advances in Cryptology - CRYPTO 2012 - 32nd Annual Cryptology Conference, Santa Barbara, CA, USA, August 19-23, 2012\. Proceedings_. 2012, pp. 608—​625. \[63\] D. Moghimi and B. Sunar and T. E. a. . . . N. Heninger, "TPM-FAIL: TPM meets Timing and Lattice Attacks" in _29th USENIX Security Symposium (USENIX Security 20)_. USENIX Association, August 2020, pp. To appear, Available: \[64\] R. J. Anderson, _Security engineering - a guide to building dependable distributed systems (3\. ed.)_. Wiley, December 2020, Available: \[65\] D. Karaklajic and J. Schmidt and I. . . . Verbauwhede, "Hardware Designer’s Guide to Fault Attacks", _IEEE Trans. Very Large Scale Integr. Syst._, vol. 21, no. 12, 2013\. pp. 2295—​2306. \[66\] D. Evtyushkin and D. V. Ponomarev, "Covert Channels through Random Number Generator: Mechanisms, Capacity Estimation and Mitigations" in _Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, Vienna, Austria, October 24-28, 2016_. 2016, pp. 843—​857. \[67\] M. Baudet and D. Lubicz and J. M. a. . . . A. Tassiaux, "On the Security of Oscillator-Based Random Number Generators", _J. Cryptology_, vol. 24, no. 2, 2011\. pp. 398—​425. \[68\] AMD, _AMD Random Number Generator_, AMD TechDocs, June 2017\. \[Online\]. Available: \[69\] ARM, _ARM TrustZone True Random Number Generator: Technical Reference Manual_, ARM 100976\\\_0000\\\_00\\\_en (rev. r0p0), May 2017\. \[Online\]. Available: \[70\] J. S. Liberty et al., "True hardware random number generation implemented in the 32-nm SOI POWER7+ processor", _IBM J. Res. Dev._, vol. 57, no. 6, 2013. \[71\] M. Varchola and M. Drutarovsk'y, "New High Entropy Element for FPGA Based True Random Number Generators" in _Cryptographic Hardware and Embedded Systems, CHES 2010, 12th International Workshop, Santa Barbara, CA, USA, August 17-20, 2010\. Proceedings_. 2010, pp. 351—​365. \[72\] M. Hamburg and P. Kocher and M. E. Marson, _Analysis of Intel’s Ivy Bridge Digital Random Number Generator_, Technical Report, Cryptography Research (Prepared for Intel), March 2012. \[73\] B. Valtchanov and V. Fischer and A. A. a. . . . F. Bernard, "Characterization of randomness sources in ring oscillator-based true random number generators in FPGAs" in _13th IEEE International Symposium on Design and Diagnostics of Electronic Circuits and Systems, DDECS 2010, Vienna, Austria, April 14-16, 2010_. 2010, pp. 48—​53. \[74\] A. Hajimiri and T. H. Lee, "A general theory of phase noise in electrical oscillators", _IEEE Journal of Solid-State Circuits_, vol. 33, no. 2, 1998\. pp. 179—​194. \[75\] A. Hajimiri and S. Limotyrakis and T. H. Lee, "Jitter and phase noise in ring oscillators", _IEEE Journal of Solid-State Circuits_, vol. 34, no. 6, June 1999\. pp. 790—​804, \[Online\]. Available: . \[76\] P. Bak, "The Devil’s Staircase", _Phys. Today_, vol. 39, no. 12, December 1986\. pp. 38—​45. \[77\] A. T. Markettos and S. W. Moore, "The Frequency Injection Attack on Ring-Oscillator-Based True Random Number Generators" in _Cryptographic Hardware and Embedded Systems - CHES 2009, 11th International Workshop, Lausanne, Switzerland, September 6-9, 2009, Proceedings_. 2009, pp. 317—​331. \[78\] Rambus, _TRNG-IP-76 / EIP-76 Family of FIPS Approved True Random Generators_, Commercial Crypto IP. Formerly (2017) available from Inside Secure., 2020\. \[Online\]. Available: \[79\] M. Blum, "Independent unbiased coin flips from a correlated biased source — A finite state Markov chain", _Combinatorica_, vol. 6, no. 2, 1986\. pp. 97—​108. \[80\] P. Lacharme, "Post-Processing Functions for a Biased Physical Random Number Generator" in _Fast Software Encryption, 15th International Workshop, FSE 2008, Lausanne, Switzerland, February 10-13, 2008, Revised Selected Papers_. 2008, pp. 334—​342. \[81\] J. P. Mechalas, _Intel Digital Random Number Generator (DRNG) Software Implementation Guide_, Intel Technical Report, Version 2.1, October 2018\. \[Online\]. Available: \[82\] S. M\\"uller, _Documentation and Analysis of the Linux Random Number Generator, Version 3.6_, Prepared for BSI by atsec information security GmbH, April 2020\. \[Online\]. Available: \[83\] ITU, _Quantum noise random number generator architecture_, Recommendation ITU-T X.1702, November 2019\. \[Online\]. Available: \[84\] D. Hurley-Smith and J. C. Hern'andez-Castro, "Quantum Leap and Crash: Searching and Finding Bias in Quantum Random Number Generators", _ACM Transactions on Privacy and Security_, vol. 23, no. 3, June 2020\. pp. 1—​25. \[85\] P. W. Shor, "Algorithms for quantum computation: Discrete logarithms and factoring" in _35th Annual Symposium on Foundations of Computer Science, Santa Fe, New Mexico, USA, 20-22 November 1994_. IEEE, 1994, pp. 124—​134, Available: . \[86\] L. Blum and M. Blum and M. Shub, "A Simple Unpredictable Pseudo-Random Number Generator", _SIAM J. Comput._, vol. 15, no. 2, 1986\. pp. 364—​383. \[87\] L. K. Grover, "A Fast Quantum Mechanical Algorithm for Database Search" in _Proceedings of the Twenty-eighth Annual ACM Symposium on Theory of Computing_, STOC '96\. ACM, 1996, pp. 212—​219, Available: . \[88\] S. Jaques and M. Naehrig and M. R. a. . . . F. Virdia, "Implementing Grover Oracles for Quantum Key Search on AES and LowMC" in _Advances in Cryptology - EUROCRYPT 2020 - 39th Annual International Conference on the Theory and Applications of Cryptographic Techniques, Zagreb, Croatia, May 10-14, 2020, Proceedings, Part II_. 2020, pp. 280—​310, Available: . \[89\] NIST, _Digital Signature Standard (DSS)_, Federal Information Processing Standards Publication FIPS 186-4, July 2013\. \[Online\]. Available: \[90\] Goldberg and R. P., "Survey of virtual machine research", _Computer_, vol. 7, no. 6, June 1974\. pp. 34-45. \[91\] \[92\] N. Juan and I. Sitaram and D. Peter and C. Alan, "Practical, Transparent Operating System Support for Superpages", _SIGOPS Oper. Syst. Rev._, vol. 36, no. SI, dec 2002\. pp. 89—​104, \[Online\]. Available: . \[93\] K. S. a. . . . . . . . . . . . . . . . . . E. S. a. . . . . . . . . . . . . . . . . . A. S. a. . . . . . . . . . . . . . . . . . V. T. a. . . . . . . . . . . . . . . . . . D. Vyukov, "Memory Tagging and how it improves C/C++ memory safety", _CoRR_, vol. abs/1802.09517, 2018\. \[Online\]. Available: 15.1. "H" Extension for Hypervisor Support, Version 1.0 ==================== ## [](#hypervisor)15.1\. "H" Extension for Hypervisor Support, Version 1.0 This chapter describes the RISC-V hypervisor extension, which virtualizes the supervisor-level architecture to support the efficient hosting of guest operating systems atop a type-1 or type-2 hypervisor. The hypervisor extension changes supervisor mode into_hypervisor-extended supervisor mode_ (HS-mode, or _hypervisor mode_ for short), where a hypervisor or a hosting-capable operating system runs. The hypervisor extension also adds another stage of address translation, from _guest physical addresses_ to supervisor physical addresses, to virtualize the memory and memory-mapped I/O subsystems for a guest operating system. HS-mode acts the same as S-mode, but with additional instructions and CSRs that control the new stage of address translation and support hosting a guest OS in virtual S-mode (VS-mode). Regular S-mode operating systems can execute without modification either in HS-mode or as VS-mode guests. In HS-mode, an OS or hypervisor interacts with the machine through the same SBI as an OS normally does from S-mode. An HS-mode hypervisor is expected to implement the SBI for its VS-mode guest. The hypervisor extension depends on an "I" base integer ISA with 32`x` registers (RV32I or RV64I), not RV32E or RV64E, which have only 16 `x`registers. CSR `mtval` must not be read-only zero, andstandard page-based address translation must be supported, either Sv32 for RV32, or a minimum of Sv39 for RV64. The hypervisor extension is enabled by setting bit 7 in the `misa` CSR, which corresponds to the letter H. RISC-V harts that implement the hypervisor extension are encouraged not to hardwire `misa`\[7\], so that the extension may be disabled. | | The baseline privileged architecture is designed to simplify the use of classic virtualization techniques, where a guest OS is run at user-level, as the few privileged instructions can be easily detected and trapped. The hypervisor extension improves virtualization performance by reducing the frequency of these traps. The hypervisor extension has been designed to be efficiently emulable on platforms that do not implement the extension, by running the hypervisor in S-mode and trapping into M-mode for hypervisor CSR accesses and to maintain shadow page tables. The majority of CSR accesses for type-2 hypervisors are valid S-mode accesses so need not be trapped. Hypervisors can support nested virtualization analogously. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#15-1-1-privilege-modes)15.1.1\. Privilege Modes The current _virtualization mode_, denoted V, indicates whether the hart is currently executing in a guest. When V=1, the hart is either in virtual S-mode (VS-mode), or in virtual U-mode (VU-mode) atop a guest OS running in VS-mode. When V=0, the hart is either in M-mode, in HS-mode, or in U-mode atop an OS running in HS-mode. The virtualization mode also indicates whether two-stage address translation is active (V=1) or inactive (V=0). [Table 1](#HPrivModes) lists the possible privilege modes of a RISC-V hart with the hypervisor extension. __Table 1\. Privilege modes with the hypervisor extension.__ | VirtualizationMode (V) | Nominal Privilege | Abbreviation | Name | Two-Stage Translation | | ---------------------- | ----------------- | ------------------- | -------------------------------------------------------- | --------------------- | | 000 | USM | U-modeHS-modeM-mode | User modeHypervisor-extended supervisor modeMachine mode | OffOffOff | | 11 | US | VU-modeVS-mode | Virtual user modeVirtual supervisor mode | OnOn | For privilege modes U and VU, the _nominal privilege mode_ is U, and for privilege modes HS and VS, the nominal privilege mode is S. HS-mode is more privileged than VS-mode, and VS-mode is more privileged than VU-mode. VS-mode interrupts are globally disabled when executing in U-mode. | | This description does not consider the possibility of U-mode or VU-mode interrupts and will be revised if an extension for user-level interrupts is adopted. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#15-1-2-hypervisor-and-virtual-supervisor-csrs)15.1.2\. Hypervisor and Virtual Supervisor CSRs An OS or hypervisor running in HS-mode uses the supervisor CSRs to interact with the exception, interrupt, and address-translation subsystems. Additional CSRs are provided to HS-mode, but not to VS-mode, to manage two-stage address translation and to control the behavior of a VS-mode guest: `hstatus`, `hedeleg`, `hideleg`, `hvip`, `hip`, `hie`,`hgeip`, `hgeie`, `henvcfg`, `henvcfgh`, `hcounteren`, `htimedelta`,`htimedeltah`, `htval`, `htinst`, and `hgatp`. Furthermore, several _virtual supervisor_ CSRs (VS CSRs) are replicas of the normal supervisor CSRs. For example, `vsstatus` is the VS CSR that duplicates the usual `sstatus` CSR. When V=1, the VS CSRs substitute for the corresponding supervisor CSRs, taking over all functions of the usual supervisor CSRs except as specified otherwise. Instructions that normally read or modify a supervisor CSR shall instead access the corresponding VS CSR. When V=1, an attempt to read or write a VS CSR directly by its own separate CSR address causes a virtual-instruction exception.(Attempts from U-mode cause an illegal-instruction exception as usual.)The VS CSRs can be accessed as themselves only from M-mode or HS-mode. While V=1, the normal HS-level supervisor CSRs that are replaced by VS CSRs retain their values but do not affect the behavior of the machine unless specifically documented to do so.Conversely, when V=0, the VS CSRs do not ordinarily affect the behavior of the machine other than being readable and writable by CSR instructions. Some standard supervisor CSRs (`senvcfg`, `scounteren`, and `scontext`, possibly others) have no matching VS CSR. These supervisor CSRs continue to have their usual function and accessibility even when V=1, except with VS-mode and VU-mode substituting for HS-mode and U-mode. Hypervisor software is expected to manually swap the contents of these registers as needed. | | Matching VS CSRs exist only for the supervisor CSRs that must be duplicated, which are mainly those that get automatically written by traps or that impact instruction execution immediately after trap entry and/or right before SRET, when software alone is unable to swap a CSR at exactly the right moment. Currently, most supervisor CSRs fall into this category, but future ones might not. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | In this chapter, we use the term _HSXLEN_ to refer to the effective XLEN when executing in HS-mode, and _VSXLEN_ to refer to the effective XLEN when executing in VS-mode. #### [](#sec:hstatus)15.1.2.1\. Hypervisor Status (`hstatus`) Register The `hstatus` register is an HSXLEN-bit read/write register formatted as shown in [Figure 1](#hstatusreg-rv32) when HSXLEN=32 and [Figure 2](#hstatusreg) when HSXLEN=64\. The `hstatus`register provides facilities analogous to the `mstatus` register for tracking and controlling the exception behavior of a VS-mode guest. ![Hypervisor status register (`hstatus`) when HSXLEN=32](_images/svg-7dc63b24f7dc0b69402ba43d617673231e6e203b.svg) Figure 1\. Hypervisor status register (`hstatus`) when HSXLEN=32 ![Hypervisor status register (`hstatus`) when HSXLEN=64.](_images/svg-b97ad0e41da0d1183b32af4abb0c8004872ca9ef.svg) Figure 2\. Hypervisor status register (`hstatus`) when HSXLEN=64. The VSXL field controls the effective XLEN for VS-mode (known as VSXLEN), which may differ from the XLEN for HS-mode (HSXLEN). When HSXLEN=32, the VSXL field does not exist, and VSXLEN=32. When HSXLEN=64, VSXL is a **WARL** field that is encoded the same as the MXL field of `misa`, shown in [Encoding of MXL field in misa](machine.html#norm:misa%5Fmxl%5Fenc). In particular, an implementation may make VSXL be a read-only field whose value always ensures that VSXLEN=HSXLEN. If HSXLEN is changed from 32 to a wider width, and if field VSXL is not restricted to a single value, it gets the value corresponding to the widest supported width not wider than the new HSXLEN. The `hstatus` fields VTSR, VTW, and VTVM are defined analogously to the`mstatus` fields TSR, TW, and TVM, but affect execution only in VS-mode, and cause virtual-instruction exceptions instead of illegal-instruction exceptions.When VTSR=1, an attempt in VS-mode to execute SRET raises a virtual-instruction exception. When VTW=1 (and assuming `mstatus`.TW=0), an attempt in VS-mode to execute WFI raises a virtual-instruction exception if the WFI does not complete within an implementation-specific, bounded time limit. An implementation may have WFI always raise a virtual-instruction exception in VS-mode when VTW=1 (and `mstatus`.TW=0), even if there are pending globally-disabled interrupts when the instruction is executed. When VTVM=1, an attempt in VS-mode to execute SFENCE.VMA or SINVAL.VMA or to access CSR `satp`raises a virtual-instruction exception. The VGEIN (Virtual Guest External Interrupt Number) field selects a guest external interrupt source for VS-level external interrupts. VGEIN is a **WLRL** field that must be able to hold values between zero and the maximum guest external interrupt number (known as GEILEN), inclusive. When VGEIN=0, no guest external interrupt source is selected for VS-level external interrupts. GEILEN may be zero, in which case VGEIN may be read-only zero. Guest external interrupts are explained in[15.1.2.4\. Hypervisor Guest External Interrupt Registers (hgeip and hgeie)](#hgeinterruptregs), and the use of VGEIN is covered further in [15.1.2.3\. Hypervisor Interrupt (hvip, hip, and hie) Registers](#hinterruptregs). Field HU (Hypervisor in U-mode) controls whether the virtual-machine load/store instructions, HLV, HLVX, and HSV, can be used also in U-mode. When HU=1, these instructions can be executed in U-mode the same as in HS-mode. When HU=0, all hypervisor instructions cause an illegal-instruction exception in U-mode. | | The HU bit allows a portion of a hypervisor to be run in U-mode for greater protection against software bugs, while still retaining access to a virtual machine’s memory. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When Ssnpm extension is implemented, the `HUPMM` field enables or disables pointer masking (see [Pointer Masking Extensions](zpm.html)) for `HLV.*` and `HSV.*` instructions in U-mode, according to the values in [Table 4](#henvcfg-pmm-values), when their explicit memory access is performed as though in VU-mode. In HS- and M-modes, pointer masking for these instructions is enabled or disabled by `senvcfg.PMM`, when their explicit memory access is performed as though in VU-mode. Setting `henvcfg.PMM`enables or disables pointer masking for `HLV.*` and `HSV.*` when their explicit memory access is performed as though in VS-mode. When the Ssnpm extension is not implemented, the `HUPMM` field is read-only zero. The `HUPMM`field is read-only zero for RV32. | | The hypervisor should copy the value written to senvcfg.PMM by the guest to the hstatus.HUPMM field prior to invoking HLV.\* or HSV.\* instructions in U-mode. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | The SPV bit (Supervisor Previous Virtualization mode) is written by the implementation whenever a trap is taken into HS-mode. Just as the SPP bit in `sstatus` is set to the (nominal) privilege mode at the time of the trap, the SPV bit in `hstatus` is set to the value of the virtualization mode V at the time of the trap. When an SRET instruction is executed when V=0, V is set to SPV. When V=1 and a trap is taken into HS-mode, bit SPVP (Supervisor Previous Virtual Privilege) is set to the nominal privilege mode at the time of the trap, the same as `sstatus`.SPP. But if V=0 before a trap, SPVP is left unchanged on trap entry. SPVP controls the effective privilege of explicit memory accesses made by the virtual-machine load/store instructions, HLV, HLVX, and HSV. | | Without SPVP, if instructions HLV, HLVX, and HSV looked instead tosstatus.SPP for the effective privilege of their memory accesses, then, even with HU=1, U-mode could not access virtual machine memory at VS-level, because to enter U-mode using SRET always leaves SPP=0\. Unlike SPP, field SPVP is untouched by transitions back-and-forth between HS-mode and U-mode. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Field GVA (Guest Virtual Address) is written by the implementation whenever a trap is taken into HS-mode. For any trap (breakpoint, address misaligned, access fault, page fault, or guest-page fault) that writes a guest virtual address to `stval`, GVA is set to 1\. For any other trap into HS-mode, GVA is set to 0. | | For breakpoint and memory access traps that write a nonzero value tostval, GVA is redundant with field SPV (the two bits are set the same) except when the explicit memory access of an HLV, HLVX, or HSV instruction causes a fault. In that case, SPV=0 but GVA=1. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The VSBE bit is a **WARL** field that controls the endianness of explicit memory accesses made from VS-mode. If VSBE=0, explicit load and store memory accesses made from VS-mode are little-endian, and if VSBE=1, they are big-endian. VSBE also controls the endianness of all implicit accesses to VS-level memory management data structures, such as page tables. An implementation may make VSBE a read-only field that always specifies the same endianness as HS-mode. #### [](#15-1-2-2-hypervisor-trap-delegation-hedeleg-and-hideleg-registers)15.1.2.2\. Hypervisor Trap Delegation (`hedeleg` and `hideleg`) Registers Register `hedeleg` is a 64-bit read/write register, formatted as shown in[Figure 3](#hedelegreg). Register `hideleg` is an HSXLEN-bit read/write register, formatted as shown in[Figure 4](#hidelegreg).By default, all traps at any privilege level are handled in M-mode, though M-mode usually uses the `medeleg` and `mideleg` CSRs to delegate some traps to HS-mode. The`hedeleg` and `hideleg` CSRs allow these traps to be further delegated to a VS-mode guest; their layout is the same as `medeleg` and `mideleg`. ![Hypervisor exception delegation register (`hedeleg`).](_images/diag-deac15a1a1746a73f03b800db076cacc5c37d42e.svg) Figure 3\. Hypervisor exception delegation register (`hedeleg`). ![Hypervisor interrupt delegation register (`hideleg`).](_images/diag-98109afb52752a7e03303ca03c793475670b3448.svg) Figure 4\. Hypervisor interrupt delegation register (`hideleg`). A synchronous trap that has been delegated to HS-mode (using `medeleg`) is further delegated to VS-mode if V=1 before the trap and the corresponding `hedeleg` bit is set. Each bit of `hedeleg` shall be either writable or read-only zero. Many bits of `hedeleg` are required specifically to be writable or zero, as enumerated in[Table 2](#hedeleg-bits). Bit 0, corresponding to instruction address-misaligned exceptions, must be writable if IALIGN=32. | | Requiring that certain bits of hedeleg be writable reduces some of the burden on a hypervisor to handle variations of implementation. | | ---------------------------------------------------------------------------------------------------------------------------------------- | When XLEN=32, `hedelegh` is a 32-bit read/write register that aliases bits 63:32 of `hedeleg`. Register `hedelegh` does not exist when XLEN=64. An interrupt that has been delegated to HS-mode (using `mideleg`) is further delegated to VS-mode if the corresponding `hideleg` bit is set. Among bits 15:0 of `hideleg`, bits 10, 6, and 2 (corresponding to the standard VS-level interrupts) are writable, and bits 12, 9, 5, and 1 (corresponding to the standard S-level interrupts) are read-only zeros. When a virtual supervisor external interrupt (code 10) is delegated to VS-mode, it is automatically translated by the machine into a supervisor external interrupt (code 9) for VS-mode, including the value written to`vscause` on an interrupt trap. Likewise, a virtual supervisor timer interrupt (6) is translated into a supervisor timer interrupt (5) for VS-mode, and a virtual supervisor software interrupt (2) is translated into a supervisor software interrupt (1) for VS-mode. Similar translations may or may not be done for platform interrupt causes (codes 16 and above). __Table 2\. Bits of hedeleg that must be writable or must be read-only zero.__ | Bit | Attribute | Corresponding Exception | | ---------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0123456789101112131516181920212223 | (See text)WritableWritableWritableWritableWritableWritableWritableWritableRead-only 0Read-only 0Read-only 0WritableWritableWritableRead-only 0WritableWritableRead-only 0Read-only 0Read-only 0Read-only 0 | Instruction address misalignedInstruction access faultIllegal instructionBreakpointLoad address misalignedLoad access faultStore/AMO address misalignedStore/AMO access faultEnvironment call from U-mode or VU-modeEnvironment call from HS-modeEnvironment call from VS-modeEnvironment call from M-modeInstruction page faultLoad page faultStore/AMO page faultDouble trapSoftware checkHardware errorInstruction guest-page faultLoad guest-page faultVirtual instructionStore/AMO guest-page fault | #### [](#hinterruptregs)15.1.2.3\. Hypervisor Interrupt (`hvip`, `hip`, and `hie`) Registers Register `hvip` is an HSXLEN-bit read/write register that a hypervisor can write to indicate virtual interrupts intended for VS-mode. Bits of`hvip` that are not writable are read-only zeros. ![Hypervisor virtual-interrupt-pending register(`hvip`).](_images/diag-a3691683650140652ed21233e737654bdd4af31b.svg) Figure 5\. Hypervisor virtual-interrupt-pending register(`hvip`). The standard portion (bits 15:0) of `hvip` is formatted as shown in[Figure 6](#hvipreg-standard). Bits VSEIP, VSTIP, and VSSIP of `hvip` are writable. Setting VSEIP=1 in `hvip` asserts a VS-level external interrupt; setting VSTIP asserts a VS-level timer interrupt; and setting VSSIP asserts a VS-level software interrupt. ![Standard portion (bits 15:0) of `hvip`.](_images/diag-4ae17791884f10d1782d02edffc5ca41743c6cde.svg) Figure 6\. Standard portion (bits 15:0) of `hvip`. Registers `hip` and `hie` are HSXLEN-bit read/write registers that supplement HS-level’s `sip` and `sie` respectively. The `hip` register indicates pending VS-level and hypervisor-specific interrupts, while`hie` contains enable bits for the same interrupts. ![Hypervisor interrupt-pending register (`hip`).](_images/diag-aa69c49675c41963c5ea659bbe297daced24999e.svg) Figure 7\. Hypervisor interrupt-pending register (`hip`). ![Hypervisor interrupt-enable register (`hie`).](_images/diag-0065999a63f1408d93c8460292a9ef1fcebeeef7.svg) Figure 8\. Hypervisor interrupt-enable register (`hie`). For each writable bit in `sie`, the corresponding bit shall be read-only zero in both `hip` and `hie`. Hence, the nonzero bits in `sie` and `hie`are always mutually exclusive, and likewise for `sip` and `hip`. | | The active bits of hip and hie cannot be placed in HS-level’s sipand sie because doing so would make it impossible for software to emulate the hypervisor extension on platforms that do not implement it in hardware. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | An interrupt _i_ will trap to HS-mode whenever all of the following are true: (a) either the current operating mode is HS-mode and the SIE bit in the `sstatus` register is set, or the current operating mode has less privilege than HS-mode; (b) bit _i_ is set in both `sip` and `sie`, or in both `hip` and `hie`; and (c) bit _i_ is not set in `hideleg`. If bit _i_ of `sie` is read-only zero, the same bit in register `hip`may be writable or may be read-only. When bit _i_ in `hip` is writable, a pending interrupt _i_ can be cleared by writing 0 to this bit. If interrupt _i_ can become pending in `hip` but bit _i_ in `hip` is read-only, then either the interrupt can be cleared by clearing bit _i_of `hvip`, or the implementation must provide some other mechanism for clearing the pending interrupt (which may involve a call to the execution environment). A bit in `hie` shall be writable if the corresponding interrupt can ever become pending in `hip`. Bits of `hie` that are not writable shall be read-only zero. The standard portions (bits 15:0) of registers `hip` and `hie` are formatted as shown in [Figure 9](#hipreg-standard) and [Figure 10](#hiereg-standard) respectively. ![Standard portion (bits 15:0) of `hip`.](_images/diag-a76c6d2e116c6c17d81d98ae789e5068cf929a8a.svg) Figure 9\. Standard portion (bits 15:0) of `hip`. ![Standard portion (bits 15:0) of `hie`.](_images/diag-3455796d1882f91823efdaf9358ddc29821f08a4.svg) Figure 10\. Standard portion (bits 15:0) of `hie`. Bits `hip`.SGEIP and `hie`.SGEIE are the interrupt-pending and interrupt-enable bits for guest external interrupts at supervisor level (HS-level). SGEIP is read-only in `hip`, and is 1 if and only if the bitwise logical-AND of CSRs `hgeip` and `hgeie` is nonzero in any bit. (See [15.1.2.4\. Hypervisor Guest External Interrupt Registers (hgeip and hgeie)](#hgeinterruptregs).) Bits `hip`.VSEIP and `hie`.VSEIE are the interrupt-pending and interrupt-enable bits for VS-level external interrupts. VSEIP is read-only in `hip`, and is the logical-OR of these interrupt sources: * bit VSEIP of `hvip`; * the bit of `hgeip` selected by `hstatus`.VGEIN; and * any other platform-specific external interrupt signal directed to VS-level. Bits `hip`.VSTIP and `hie`.VSTIE are the interrupt-pending and interrupt-enable bits for VS-level timer interrupts. VSTIP is read-only in `hip`, and is the logical-OR of `hvip`.VSTIP and, when the Sstc extension is implemented, the timer interrupt signal resulting from `vstimecmp`. The `hip`.VSTIP bit, in response to timer interrupts generated by `vstimecmp`, is set by writing`vstimecmp` with a value that is less than or equal to the sum of `time` and`htimedelta`, truncated to 64 bits; it is cleared by writing `vstimecmp` with a greater value. The `hip`.VSTIP bit remains defined while V=0 as well as V=1. Bits `hip`.VSSIP and `hie`.VSSIE are the interrupt-pending and interrupt-enable bits for VS-level software interrupts. VSSIP in `hip`is an alias (writable) of the same bit in `hvip`. Multiple simultaneous interrupts destined for HS-mode are handled in the following decreasing priority order: SEI, SSI, STI, SGEI, VSEI, VSSI, VSTI, LCOFI. #### [](#hgeinterruptregs)15.1.2.4\. Hypervisor Guest External Interrupt Registers (`hgeip` and `hgeie`) The `hgeip` register is an HSXLEN-bit read-only register, formatted as shown in [Figure 11](#hgeipreg), that indicates pending guest external interrupts for this hart. The `hgeie` register is an HSXLEN-bit read/write register, formatted as shown in[Figure 12](#hgeiereg), that contains enable bits for the guest external interrupts at this hart. Guest external interrupt number_i_ corresponds with bit _i_ in both `hgeip` and `hgeie`. ![Hypervisor guest external interrupt-pending register (`hgeip`).](_images/diag-637b6d6a881df7d48a3f70eecd37bcc39fa7294f.svg) Figure 11\. Hypervisor guest external interrupt-pending register (`hgeip`). ![Hypervisor guest external interrupt-enable register (`hgeie`).](_images/diag-ae5ebabd5533390af23281514ea6a01b12f4459a.svg) Figure 12\. Hypervisor guest external interrupt-enable register (`hgeie`). Guest external interrupts represent interrupts directed to individual virtual machines at VS-level. If a RISC-V platform supports placing a physical device under the direct control of a guest OS with minimal hypervisor intervention (known as _pass-through_ or _direct assignment_between a virtual machine and the physical device), then, in such circumstance, interrupts from the device are intended for a specific virtual machine. Each bit of `hgeip` summarizes _all_ pending interrupts directed to one virtual hart, as collected and reported by an interrupt controller. To distinguish specific pending interrupts from multiple devices, software must query the interrupt controller. | | Support for guest external interrupts requires an interrupt controller that can collect virtual-machine-directed interrupts separately from other interrupts. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | The number of bits implemented in `hgeip` and `hgeie` for guest external interrupts is UNSPECIFIED and may be zero. This number is known as _GEILEN_. The least-significant bits are implemented first, apart from bit 0\. Hence, if GEILEN is nonzero, bits GEILEN:1 shall be writable in `hgeie`, and all other bit positions shall be read-only zeros in both `hgeip` and`hgeie`. | | The set of guest external interrupts received and handled at one physical hart may differ from those received at other harts. Guest external interrupt number _i_ at one physical hart is typically expected not to be the same as guest external interrupt _i_ at any other hart. For any one physical hart, the maximum number of virtual harts that may directly receive guest external interrupts is limited by GEILEN. The maximum this number can be for any implementation is 31 for RV32 and 63 for RV64, per physical hart. A hypervisor is always free to _emulate_ devices for any number of virtual harts without being limited by GEILEN. Only direct pass-through (direct assignment) of interrupts is affected by the GEILEN limit, and the limit is on the number of virtual harts receiving such interrupts, not the number of distinct interrupts received. The number of distinct interrupts a single virtual hart may receive is determined by the interrupt controller. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Register `hgeie` selects the subset of guest external interrupts that cause a supervisor-level (HS-level) guest external interrupt. The enable bits in `hgeie` do not affect the VS-level external interrupt signal selected from `hgeip` by `hstatus`.VGEIN. #### [](#sec:henvcfg)15.1.2.5\. Hypervisor Environment Configuration Register (`henvcfg`) The `henvcfg` CSR is a 64-bit read/write register, formatted as shown in [Figure 13](#henvcfg), that controls certain characteristics of the execution environment when virtualization mode V=1. ![Hypervisor environment configuration register (`henvcfg`).](_images/svg-6c1054cbba95551da58f4ef98ea985c3e7dd592a.svg) Figure 13\. Hypervisor environment configuration register (`henvcfg`). If bit FIOM (Fence of I/O implies Memory) is set to one in `henvcfg`, FENCE instructions executed when V=1 are modified so the requirement to order accesses to device I/O implies also the requirement to order main memory accesses. [Table 3](#henvcfg-FIOM) details the modified interpretation of FENCE instruction bits PI, PO, SI, and SO when FIOM=1 and V=1. Similarly, when FIOM=1 and V=1, if an atomic instruction that accesses a region ordered as device I/O has its _aq_ and/or _rl_ bit set, then that instruction is ordered as though it accesses both device I/O and memory. __Table 3\. Modified interpretation of FENCE predecessor and successor sets when FIOM=1 and virtualization mode V=1.__ | Instruction bit | Meaning when set | | --------------- | -------------------------------------------------------------------------------------------------------------- | | PIPO | Predecessor device input and memory reads (PR implied)Predecessor device output and memory writes (PW implied) | | SISO | Successor device input and memory reads (SR implied)Successor device output and memory writes (SW implied) | The PBMTE bit controls whether the Svpbmt extension is available for use in VS-stage address translation. When PBMTE=1, Svpbmt is available for VS-stage address translation. When PBMTE=0, the implementation behaves as though Svpbmt were not implemented for VS-stage address translation. If Svpbmt is not implemented, PBMTE is read-only zero. If the Svadu extension is implemented, the ADUE bit controls whether hardware updating of PTE A/D bits is enabled for VS-stage address translation. When ADUE=1, hardware updating of PTE A/D bits is enabled during VS-stage address translation, and the implementation behaves as though the Svade extension were not implemented for VS-mode address translation. When ADUE=0, the implementation behaves as though Svade were implemented for VS-stage address translation. If Svadu is not implemented, ADUE is read-only zero. The Sstc extension adds the `STCE` (STimecmp Enable) bit to `henvcfg` CSR. When the Sstc extension is not implemented, `STCE` is read-only zero. The `STCE` bit enables `vstimecmp` for VS-mode when set to one. When `STCE` bit is `henvcfg` is zero, an attempt to access `stimecmp` (really `vstimecmp`) when V=1 raises a virtual-instruction exception, and `VSTIP` in `hip` reverts to its defined behavior as if this extension is not implemented. The Zicboz extension adds the `CBZE` (Cache Block Zero instruction enable) field to `henvcfg`. The `CBZE` field applies to execution of the cache block zero instruction (`CBO.ZERO`) in privilege modes VS and VU, and only when the instruction is HS-qualified. If the instruction is not HS-qualified, it raises an illegal-instruction exception. If the instruction is HS-qualified and the`CBZE` field is set to 1, the instruction is enabled for execution; otherwise, if the `CBZE` field is set to 0, it raises a virtual-instruction exception. When the Zicboz extension is not implemented, `CBZE` is read-only zero. The Zicbom extension adds the `CBCFE` (Cache Block Clean and Flush instruction Enable) field to `henvcfg`. When V=1, if the `CBO.CLEAN` and `CBO.FLUSH`instructions are not HS-qualified, they raise an illegal-instruction exception. If the instructions are HS-qualified and the `CBCFE` field is set to 1, the instructions are enabled for execution; otherwise, if the `CBCFE` field is set to 0, they raise a virtual-instruction exception. When the Zicbom extension is not implemented, `CBCFE` is read-only zero. The Zicbom extension adds the `CBIE` (Cache Block Invalidate instruction Enable) WARL field to `henvcfg`. The `CBIE` field controls execution of the cache block invalidate instruction (`CBO.INVAL`) in privilege modes VS and VU. The encoding`10b` is reserved. When the Zicbom extension is not implemented, `CBIE` is read-only zero. When V=1, if the `CBO.INVAL` instruction is not HS-qualified, it raises an illegal-instruction exception. If the instruction is HS-qualified and the`CBIE` field is set to `01b` or `11b`, the instruction is enabled for execution; otherwise, it raises a virtual-instruction exception. If `CBO.INVAL` is enabled in HS-mode to perform a flush operation, then when the instruction is enabled in VS- or VU-mode it performs a flush operation, even if`CBIE` is set to `11b`. Otherwise, when the instruction is enabled for execution, its behavior depends on the `CBIE` encoding, as follows: * `01b` — The instruction is executed and performs a flush operation, even if configured by VS-mode to perform an invalidate operation. * `11b` — The instruction is executed and performs an invalidate operation, unless configured by VS-mode to perform a flush operation. If the Ssnpm extension is implemented, the `PMM` field enables or disables pointer masking (see [Pointer Masking Extensions](zpm.html)) for VS-mode, according to the values in[Table 4](#henvcfg-pmm-values). When the Ssnpm extension is not implemented, the `PMM`field is read-only zero. The `PMM` field is read-only zero for RV32. __Table 4\. Legal values of PMM WARL field__ | Value | Description | | ----- | ---------------------------------------------------------------------- | | 00 | Pointer masking is disabled (PMLEN = 0) | | 01 | Reserved | | 10 | Pointer masking is enabled with PMLEN = XLEN - 57 (PMLEN = 7 on RV64) | | 11 | Pointer masking is enabled with PMLEN = XLEN - 48 (PMLEN = 16 on RV64) | The Zicfilp extension adds the `LPE` field in `henvcfg`. When the `LPE` field is set to 1, the Zicfilp extension is enabled in VS-mode. When the `LPE` field is 0, the Zicfilp extension is not enabled in VS-mode and the following rules apply to VS-mode: * The hart does not update the `ELP` state; it remains as `NO_LP_EXPECTED`. * The `LPAD` instruction operates as a no-op. The Zicfiss extension adds the `SSE` field in `henvcfg`. If the `SSE` field is set to 1, the Zicfiss extension is activated in VS-mode. When the `SSE` field is 0, the Zicfiss extension remains inactive in VS-mode, and the following rules apply when `V=1`: * 32-bit Zicfiss instructions will revert to their behavior as defined by Zimop. * 16-bit Zicfiss instructions will revert to their behavior as defined by Zcmop. * The `pte.xwr=010b` encoding in VS-stage page tables becomes reserved. * The `senvcfg.SSE` field will read as zero and is read-only. * When `menvcfg.SSE` is one, `SSAMOSWAP.W/D` raises a virtual-instruction exception. The Ssdbltrp extension adds the double-trap-enable (`DTE`) field in `henvcfg`. When `henvcfg.DTE` is zero, the implementation behaves as though Ssdbltrp is not implemented for VS-mode and the `vsstatus.SDT` bit is read-only zero. When XLEN=32, `henvcfgh` is a 32-bit read/write register that aliases bits 63:32 of `henvcfg`. Register `henvcfgh` does not exist when XLEN=64. #### [](#15-1-2-6-hypervisor-counter-enable-hcounteren-register)15.1.2.6\. Hypervisor Counter-Enable (`hcounteren`) Register The counter-enable register `hcounteren` is a 32-bit register that controls the availability of the hardware performance monitoring counters to the guest virtual machine. ![Hypervisor counter-enable register (`hcounteren`).](_images/diag-a39c492e0a4c6b6d4563223c564a9c49b862c9c3.svg) Figure 14\. Hypervisor counter-enable register (`hcounteren`). When the CY, TM, IR, or HPM_n_ bit in the `hcounteren` register is clear, attempts to read the `cycle`, `time`, `instret`, or`hpmcounter` _n_ register while V=1 will cause a virtual-instruction exception if the same bit in `mcounteren` is 1\. When one of these bits is set, access to the corresponding register is permitted when V=1, unless prevented for some other reason. In VU-mode, a counter is not readable unless the applicable bits are set in both `hcounteren` and`scounteren`. In addition, when the TM bit in the `hcounteren` register is clear, attempts to access the `vstimecmp` register (via `stimecmp`) while executing in VS-mode will cause a virtual-instruction exception if the same bit in `mcounteren` is set. When this bit and the same bit in `mcounteren` are both set, access to the`vstimecmp` register (if implemented) is permitted in VS-mode. `hcounteren` must be implemented. However, any of the bits may be read-only zero, indicating reads to the corresponding counter will cause an exception when V=1\. Hence, they are effectively **WARL** fields. #### [](#15-1-2-7-hypervisor-time-delta-htimedelta-register)15.1.2.7\. Hypervisor Time Delta (`htimedelta`) Register The `htimedelta` CSR is a 64-bit read/write register that contains the delta between the value of the `time` CSR and the value returned in VS-mode or VU-mode. That is, reading the `time` CSR in VS or VU mode returns the sum of the contents of `htimedelta` and the actual value of `time`. | | Because overflow is ignored when summing htimedelta and time, large values of htimedelta may be used to represent negative time offsets. | | ------------------------------------------------------------------------------------------------------------------------------------------- | ![Hypervisor time delta register.](_images/diag-dfe31c16b4dd81aa2d34f985f502022839192c0c.svg) Figure 15\. Hypervisor time delta register. When XLEN=32, `htimedeltah` is a 32-bit read/write register that aliases bits 63:32 of `htimedelta`. Register `htimedeltah` does not exist when XLEN=64. If the `time` CSR is implemented, `htimedelta` (and `htimedeltah` for XLEN=32) must be implemented. #### [](#15-1-2-8-hypervisor-trap-value-htval-register)15.1.2.8\. Hypervisor Trap Value (`htval`) Register The `htval` register is an HSXLEN-bit read/write register formatted as shown in [Figure 16](#htvalreg). When a trap is taken into HS-mode, `htval` is written with additional exception-specific information, alongside `stval`, to assist software in handling the trap. ![Hypervisor trap value register (`htval`).](_images/diag-cd3108085a557f813388c94d049bc7a9e4c16a0c.svg) Figure 16\. Hypervisor trap value register (`htval`). When a guest-page-fault trap is taken into HS-mode, `htval` is written with either zero or the guest physical address that faulted, shifted right by 2 bits. For other traps, `htval` is set to zero, but a future standard or extension may redefine `htval’s` setting for other traps. A guest-page fault may arise due to an implicit memory access during first-stage (VS-stage) address translation, in which case a guest physical address written to `htval` is that of the implicit memory access that faulted—for example, the address of a VS-level page table entry that could not be read. (The guest physical address corresponding to the original virtual address is unknown when VS-stage translation fails to complete.) Additional information is provided in CSR `htinst`to disambiguate such situations. Otherwise, for misaligned loads and stores that cause guest-page faults, a nonzero guest physical address in `htval` corresponds to the faulting portion of the access as indicated by the virtual address in `stval`. For instruction guest-page faults on systems with variable-length instructions, a nonzero `htval` corresponds to the faulting portion of the instruction as indicated by the virtual address in `stval`. | | A guest physical address written to htval is shifted right by 2 bits to accommodate addresses wider than the current XLEN. For RV32, the hypervisor extension permits guest physical addresses as wide as 34 bits, and htval reports bits 33:2 of the address. This shift-by-2 encoding of guest physical addresses matches the encoding of physical addresses in PMP address registers ([Physical Memory Protection](machine.html#pmp)) and in page table entries ([Sv32: Page-Based 32-bit Virtual-Memory Systems](supervisor.html#sv32), [Sv32: Page-Based 39-bit Virtual-Memory Systems](supervisor.html#sv39), [Sv48: Page-Based 48-bit Virtual-Memory Systems](supervisor.html#sv48), and xref:supervisor.adoc#sv57\["Sv57: Page-Based 57-bit Virtual-Memory System). If the least-significant two bits of a faulting guest physical address are needed, these bits are ordinarily the same as the least-significant two bits of the faulting virtual address in stval. For faults due to implicit memory accesses for VS-stage address translation, the least-significant two bits are instead zeros. These cases can be distinguished using the value provided in register htinst. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | `htval` is a **WARL** register that must be able to hold zero and may be capable of holding only an arbitrary subset of other 2-bit-shifted guest physical addresses, if any. | | Unless it has reason to assume otherwise (such as a platform standard), software that writes a value to htval should read back from htval to confirm the stored value. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#15-1-2-9-hypervisor-trap-instruction-htinst-register)15.1.2.9\. Hypervisor Trap Instruction (`htinst`) Register The `htinst` register is an HSXLEN-bit read/write register formatted as shown in [Figure 17](#htinstreg). When a trap is taken into HS-mode, `htinst` is written with a value that, if nonzero, provides information about the instruction that trapped, to assist software in handling the trap. The values that may be written to `htinst` on a trap are documented in [15.1.6.3\. Transformed Instruction or Pseudoinstruction for mtinst or htinst](#tinst-vals). ![Hypervisor trap instruction (`htinst`) register.](_images/diag-88556cab04dcc9a8872c722d170d6a807e792be0.svg) Figure 17\. Hypervisor trap instruction (`htinst`) register. `htinst` is a **WARL** register that need only be able to hold the values that the implementation may automatically write to it on a trap. #### [](#hgatp)15.1.2.10\. Hypervisor Guest Address Translation and Protection (`hgatp`) Register The `hgatp` register is an HSXLEN-bit read/write register, formatted as shown in [Figure 18](#rv32hgatp) for HSXLEN=32 and[Figure 19](#rv64hgatp) for HSXLEN=64, which controls G-stage address translation and protection, the second stage of two-stage translation for guest virtual addresses (see[15.1.5\. Two-Stage Address Translation](#two-stage-translation)). Similar to CSR `satp`, this register holds the physical page number (PPN) of the guest-physical root page table; a virtual machine identifier (VMID), which facilitates address-translation fences on a per-virtual-machine basis; and the MODE field, which selects the address-translation scheme for guest physical addresses. When `mstatus`.TVM=1, attempts to read or write `hgatp` while executing in HS-mode will raise an illegal-instruction exception. ![Hypervisor guest address translation and protection register `hgatp` when HSXLEN=32.](_images/diag-05a282c2790c56f3693eb70780f67ba18be1f408.svg) Figure 18\. Hypervisor guest address translation and protection register `hgatp` when HSXLEN=32. ![Hypervisor guest address translation and protection register `hgatp` when HSXLEN=64 for MODE values Bare, Sv39x4, Sv48x4, and Sv57x4.](_images/diag-e014dc8af24dd841d90f7703dd4cb854399548e8.svg) Figure 19\. Hypervisor guest address translation and protection register `hgatp` when HSXLEN=64 for MODE values Bare, Sv39x4, Sv48x4, and Sv57x4. [Table 5](#hgatp-mode) shows the encodings of the MODE field when HSXLEN=32 and HSXLEN=64\. When MODE=Bare, guest physical addresses are equal to supervisor physical addresses, and there is no further memory protection for a guest virtual machine beyond the physical memory protection scheme described in [Physical Memory Protection](machine.html#pmp). In this case, software must write zero to the remaining fields in `hgatp`.Attempting to select MODE=Bare with a nonzero pattern in the remaining fields has an UNSPECIFIED effect on the value that the remaining fields assume and an UNSPECIFIED effect on G-stage address translation and protection behavior. When HSXLEN=32, the only other valid setting for MODE is Sv32x4, which is a modification of the usual Sv32 paged virtual-memory scheme, extended to support 34-bit guest physical addresses. When HSXLEN=64, modes Sv39x4, Sv48x4, and Sv57x4 are defined as modifications of the Sv39, Sv48, and Sv57 paged virtual-memory schemes. All of these paged virtual-memory schemes are described in[15.1.5.1\. Guest Physical Address Translation](#guest-addr-translation). The remaining MODE settings when HSXLEN=64 are reserved for future use and may define different interpretations of the other fields in `hgatp`. __Table 5\. Encoding of hgatp MODE field.__ | HSXLEN=32 | | | | ------------- | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Value | Name | Description | | 01 | BareSv32x4 | No translation or protection.Page-based 34-bit virtual addressing (2-bit extension of Sv32). | | **HSXLEN=64** | | | | Value | Name | Description | | 01-7891011-15 | Bare—Sv39x4Sv48x4Sv57x4— | No translation or protection. _Reserved_Page-based 41-bit virtual addressing (2-bit extension of Sv39).Page-based 50-bit virtual addressing (2-bit extension of Sv48).Page-based 59-bit virtual addressing (2-bit extension of Sv57). _Reserved_ | Implementations are not required to support all defined MODE settings when HSXLEN=64. A write to `hgatp` with an unsupported MODE value is not ignored as it is for `satp`. Instead, the fields of `hgatp` are **WARL** in the normal way, when so indicated. As explained in [15.1.5.1\. Guest Physical Address Translation](#guest-addr-translation), for the paged virtual-memory schemes (Sv32x4, Sv39x4, Sv48x4, and Sv57x4), the root page table is 16 KiB and must be aligned to a 16-KiB boundary. In these modes, the lowest two bits of the physical page number (PPN) in`hgatp` always read as zeros. An implementation that supports only the defined paged virtual-memory schemes and/or Bare may make PPN\[1:0\] read-only zero. The number of VMID bits is UNSPECIFIED and may be zero. The number of implemented VMID bits, termed _VMIDLEN_, may be determined by writing one to every bit position in the VMID field, then reading back the value in `hgatp`to see which bit positions in the VMID field hold a one. The least-significant bits of VMID are implemented first: that is, if VMIDLEN > 0, VMID\[VMIDLEN-1:0\] is writable. The maximal value of VMIDLEN, termed VMIDMAX, is 7 for Sv32x4 or 14 for Sv39x4, Sv48x4, and Sv57x4. The `hgatp` register is considered _active_ for the purposes of the address-translation algorithm _unless_ the effective privilege mode is U and `hstatus`.HU=0. | | This definition simplifies the implementation of speculative execution of HLV, HLVX, and HSV instructions. | | ------------------------------------------------------------------------------------------------------------- | Note that writing `hgatp` does not imply any ordering constraints between page-table updates and subsequent G-stage address translations. If the new virtual machine’s guest physical page tables have been modified, or if a VMID is reused, it may be necessary to execute an HFENCE.GVMA instruction (see [15.1.3.2\. Hypervisor Memory-Management Fence Instructions](#hfence.vma)) before or after writing `hgatp`. #### [](#vsstatus)15.1.2.11\. Virtual Supervisor Status (`vsstatus`) Register The `vsstatus` register is a VSXLEN-bit read/write register that is VS-mode’s version of supervisor register `sstatus`, formatted as shown in [Figure 20](#vsstatusreg-rv32) when VSXLEN=32 and[Figure 21](#vsstatusreg) when VSXLEN=64\. When V=1,`vsstatus` substitutes for the usual `sstatus`, so instructions that normally read or modify `sstatus` actually access `vsstatus` instead. ![Virtual supervisor status (`vsstatus`) register when VSXLEN=32.](_images/svg-2f9763230ddaad31f92e8cb2ea0e22627ce19655.svg) Figure 20\. Virtual supervisor status (`vsstatus`) register when VSXLEN=32. ![Virtual supervisor status (`vsstatus`) register when VSXLEN=64.](_images/svg-0b37713d9486d387f633475513bb0fb1a2119399.svg) Figure 21\. Virtual supervisor status (`vsstatus`) register when VSXLEN=64. The UXL field controls the effective XLEN for VU-mode, which may differ from the XLEN for VS-mode (VSXLEN). When VSXLEN=32, the UXL field does not exist, and VU-mode XLEN=32\. When VSXLEN=64, UXL is a **WARL** field that is encoded the same as the MXL field of `misa`, shown in [Encoding of MXL field in misa](machine.html#norm:misa%5Fmxl%5Fenc). In particular, an implementation may make UXL be a read-only copy of field VSXL of `hstatus`, forcing VU-mode XLEN=VSXLEN. If VSXLEN is changed from 32 to a wider width, and if field UXL is not restricted to a single value, it gets the value corresponding to the widest supported width not wider than the new VSXLEN. When V=1, both `vsstatus`.FS and the HS-level `sstatus`.FS are in effect. Attempts to execute a floating-point instruction when either field is 0 (Off) raise an illegal-instruction exception. Modifying the floating-point state when V=1 causes both fields to be set to 3 (Dirty). | | For a hypervisor to benefit from the extension context status, it must have its own copy in the HS-level sstatus, maintained independently of a guest OS running in VS-mode. While a version of the extension context status obviously must exist in vsstatus for VS-mode, a hypervisor cannot rely on this version being maintained correctly, given that VS-level software can change vsstatus.FS arbitrarily. If the HS-levelsstatus.FS were not independently active and maintained by the hardware in parallel with vsstatus.FS while V=1, hypervisors would always be forced to conservatively swap all floating-point state when context-switching between virtual machines. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Similarly, when V=1, both `vsstatus`.VS and the HS-level `sstatus`.VS are in effect. Attempts to execute a vector instruction when either field is 0 (Off) raise an illegal-instruction exception. Modifying the vector state when V=1 causes both fields to be set to 3 (Dirty). Read-only fields SD and XS summarize the extension context status as it is visible to VS-mode only. For example, the value of the HS-level`sstatus`.FS does not affect `vsstatus`.SD. An implementation may make field UBE be a read-only copy of`hstatus`.VSBE. When V=0, `vsstatus` does not directly affect the behavior of the machine, unless a virtual-machine load/store (HLV, HLVX, or HSV) or the MPRV feature in the `mstatus` register is used to execute a load or store _as though_ V=1. The Zicfilp extension adds the `SPELP` field that holds the previous `ELP`, and is updated as specified in [Preserving Expected Landing Pad State on Traps](priv-cfi.html#ZICFILP%5FFORWARD%5FTRAPS). The `SPELP` field is encoded as follows: * 0 - `NO_LP_EXPECTED` \- no landing pad instruction expected. * 1 - `LP_EXPECTED` \- a landing pad instruction is expected. The Ssdbltrp adds an S-mode-disable-trap (`SDT`) field extension to address double trap (See [Double Trap Control in sstatus Register](supervisor.html#supv-double-trap)) in VS-mode. #### [](#15-1-2-12-virtual-supervisor-interrupt-vsip-and-vsie-registers)15.1.2.12\. Virtual Supervisor Interrupt (`vsip` and `vsie`) Registers The `vsip` and `vsie` registers are VSXLEN-bit read/write registers that are VS-mode’s versions of supervisor CSRs `sip` and `sie`, formatted as shown in [Figure 22](#vsipreg) and [Figure 23](#vsiereg)respectively. When V=1, `vsip` and `vsie` substitute for the usual `sip`and `sie`, so instructions that normally read or modify `sip`/`sie`actually access `vsip`/`vsie` instead. However, interrupts directed to HS-level continue to be indicated in the HS-level `sip` register, not in`vsip`, when V=1. ![Virtual supervisor interrupt-pending register (`vsip`).](_images/diag-5b61c51e4ca8c647ea13e7a030ff04d640a5674c.svg) Figure 22\. Virtual supervisor interrupt-pending register (`vsip`). ![Virtual supervisor interrupt-enable register (`vsie`).](_images/diag-4531c6c4895d71b9995adf155cd2313ccf2829bb.svg) Figure 23\. Virtual supervisor interrupt-enable register (`vsie`). The standard portions (bits 15:0) of registers `vsip` and `vsie` are formatted as shown in [Figure 24](#vsipreg-standard)and [Figure 25](#vsiereg-standard) respectively. ![Standard portion (bits 15:0) of `vsip`.](_images/diag-48307345e8ac533bf1871bbbc334000a3e1efc5b.svg) Figure 24\. Standard portion (bits 15:0) of `vsip`. ![Standard portion (bits 15:0) of `vsie`.](_images/diag-cdd5475343af7c489bbb26c81488bbca678c317d.svg) Figure 25\. Standard portion (bits 15:0) of `vsie`. Extension Shlcofideleg supports delegating LCOFI interrupts to VS-mode. If the Shlcofideleg extension is implemented, `hideleg` bit 13 is writable; otherwise, it is read-only zero. When bit 13 of `hideleg` is zero, `vsip`.LCOFIP and `vsie`.LCOFIE are read-only zeros. Else, `vsip`.LCOFIP and `vsie`.LCOFIE are aliases of `sip`.LCOFIP and `sie`.LCOFIE. When bit 10 of `hideleg` is zero, `vsip`.SEIP and `vsie`.SEIE are read-only zeros. Else, `vsip`.SEIP and `vsie`.SEIE are aliases of`hip`.VSEIP and `hie`.VSEIE. When bit 6 of `hideleg` is zero, `vsip`.STIP and `vsie`.STIE are read-only zeros. Else, `vsip`.STIP and `vsie`.STIE are aliases of`hip`.VSTIP and `hie`.VSTIE. When bit 2 of `hideleg` is zero, `vsip`.SSIP and `vsie`.SSIE are read-only zeros. Else, `vsip`.SSIP and `vsie`.SSIE are aliases of`hip`.VSSIP and `hie`.VSSIE. #### [](#15-1-2-13-virtual-supervisor-trap-vector-base-address-vstvec-register)15.1.2.13\. Virtual Supervisor Trap Vector Base Address (`vstvec`) Register The `vstvec` register is a VSXLEN-bit read/write register that is VS-mode’s version of supervisor register `stvec`, formatted as shown in[Figure 26](#vstvecreg). When V=1, `vstvec` substitutes for the usual `stvec`, so instructions that normally read or modify `stvec`actually access `vstvec` instead. When V=0, `vstvec` does not directly affect the behavior of the machine. ![Virtual supervisor trap vector base address register `vstvec`.](_images/diag-994bb7d5d228376ece4a919a6aaac59d30314364.svg) Figure 26\. Virtual supervisor trap vector base address register `vstvec`. #### [](#15-1-2-14-virtual-supervisor-scratch-vsscratch-register)15.1.2.14\. Virtual Supervisor Scratch (`vsscratch`) Register The `vsscratch` register is a VSXLEN-bit read/write register that is VS-mode’s version of supervisor register `sscratch`, formatted as shown in [Figure 27](#vsscratchreg). When V=1, `vsscratch`substitutes for the usual `sscratch`, so instructions that normally read or modify `sscratch` actually access `vsscratch` instead. The contents of `vsscratch` never directly affect the behavior of the machine. ![Virtual supervisor scratch register `vsscratch`.](_images/diag-fc2aa2eb6a237d948166043f929f70745f2d99f8.svg) Figure 27\. Virtual supervisor scratch register `vsscratch`. #### [](#15-1-2-15-virtual-supervisor-exception-program-counter-vsepc-register)15.1.2.15\. Virtual Supervisor Exception Program Counter (`vsepc`) Register The `vsepc` register is a VSXLEN-bit read/write register that is VS-mode’s version of supervisor register `sepc`, formatted as shown in[Figure 28](#vsepcreg). When V=1, `vsepc` substitutes for the usual `sepc`, so instructions that normally read or modify `sepc`actually access `vsepc` instead. When V=0, `vsepc` does not directly affect the behavior of the machine. `vsepc` is a **WARL** register that must be able to hold the same set of values that `sepc` can hold. ![Virtual supervisor exception program counter (`vsepc`).](_images/diag-2ad4f8b2843e94416a6c6f2ceab50368e772c2b0.svg) Figure 28\. Virtual supervisor exception program counter (`vsepc`). #### [](#15-1-2-16-virtual-supervisor-cause-vscause-register)15.1.2.16\. Virtual Supervisor Cause (`vscause`) Register The `vscause` register is a VSXLEN-bit read/write register that is VS-mode’s version of supervisor register `scause`, formatted as shown in[Figure 29](#vscausereg). When V=1, `vscause` substitutes for the usual `scause`, so instructions that normally read or modify`scause` actually access `vscause` instead. When V=0, `vscause` does not directly affect the behavior of the machine. `vscause` is a **WLRL** register that must be able to hold the same set of values that `scause` can hold. ![Virtual supervisor cause register (`vscause`).](_images/diag-f15ebf1570e4a40adecd89e52c26a826a0ee6105.svg) Figure 29\. Virtual supervisor cause register (`vscause`). #### [](#15-1-2-17-virtual-supervisor-trap-value-vstval-register)15.1.2.17\. Virtual Supervisor Trap Value (`vstval`) Register The `vstval` register is a VSXLEN-bit read/write register that is VS-mode’s version of supervisor register `stval`, formatted as shown in[Figure 30](#vstvalreg). When V=1, `vstval` substitutes for the usual `stval`, so instructions that normally read or modify `stval`actually access `vstval` instead. When V=0, `vstval` does not directly affect the behavior of the machine. `vstval` is a **WARL** register that must be able to hold the same set of values that `stval` can hold. ![Virtual supervisor trap value register (`vstval`).](_images/diag-be52639561522792111869e9915d4cd58a4b9fc3.svg) Figure 30\. Virtual supervisor trap value register (`vstval`). #### [](#15-1-2-18-virtual-supervisor-address-translation-and-protection-vsatp-register)15.1.2.18\. Virtual Supervisor Address Translation and Protection (`vsatp`) Register The `vsatp` register is a VSXLEN-bit read/write register that is VS-mode’s version of supervisor register `satp`, formatted as shown in[Figure 31](#rv32vsatpreg) for VSXLEN=32 and [Figure 32](#rv64vsatpreg) for VSXLEN=64\. When V=1,`vsatp` substitutes for the usual `satp`, so instructions that normally read or modify `satp` actually access `vsatp` instead. `vsatp` controls VS-stage address translation, the first stage of two-stage translation for guest virtual addresses (see[15.1.5\. Two-Stage Address Translation](#two-stage-translation)). ![Virtual supervisor address translation and protection `vsatp` register when VSXLEN=32.](_images/diag-3e32007a34cf997898a5e47dd1321277556b7a71.svg) Figure 31\. Virtual supervisor address translation and protection `vsatp` register when VSXLEN=32. ![Virtual supervisor address translation and protection `vsatp` register when VSXLEN=64.](_images/diag-bb70793b73e937a69e853ee794441dee0fb4560d.svg) Figure 32\. Virtual supervisor address translation and protection `vsatp` register when VSXLEN=64. The `vsatp` register is considered _active_ for the purposes of the address-translation algorithm _unless_ the effective privilege mode is U and `hstatus`.HU=0\. However, even when `vsatp` is active, VS-stage page-table entries’ A bits must not be set as a result of speculative execution, unless the effective privilege mode is VS or VU. | | In particular, virtual-machine load/store (HLV, HLVX, or HSV) instructions that are mispredicted must not cause VS-stage A bits to be set. | | --------------------------------------------------------------------------------------------------------------------------------------------- | When V=0, a write to `vsatp` with an unsupported MODE value is either ignored as it is for `satp`, or the fields of `vsatp` are treated as **WARL** in the normal way. However, when V=1, a write to `satp` with an unsupported MODE value _is_ ignored and no write to `vsatp` is effected. When V=0, `vsatp` does not directly affect the behavior of the machine, unless a virtual-machine load/store (HLV, HLVX, or HSV) or the MPRV feature in the `mstatus` register is used to execute a load or store _as though_ V=1. #### [](#vstimecmp)15.1.2.19\. Virtual Supervisor Timer (`vstimecmp`) Register The `vstimecmp` CSR is a 64-bit register and has 64-bit precision on all RV32 and RV64 systems. In RV32 only, accesses to the `vstimecmp` CSR access the low 32 bits, while accesses to the `vstimecmph` CSR access the high 32 bits of`vstimecmp`. A virtual supervisor timer interrupt becomes pending, as reflected in the VSTIP bit in the `hip` register, whenever (`time` \+ `htimedelta`), truncated to 64 bits, contains a value greater than or equal to `vstimecmp`, treating the values as unsigned integers. If the result of this comparison changes, it is guaranteed to be reflected in VSTIP eventually, but not necessarily immediately. The interrupt remains posted until `vstimecmp` becomes greater than (`time`\+ `htimedelta`), typically as a result of writing `vstimecmp`. The interrupt will be taken based on the standard interrupt enable and delegation rules while V=1. | | In systems in which a supervisor execution environment (SEE) implemented by an HS-mode hypervisor provides timer facilities via an SBI function call, this SBI call will continue to support requests to schedule a timer interrupt. The SEE will simply make use of vstimecmp, changing its value as appropriate. This ensures compatibility with existing guest VS-mode software that uses this SEE facility, while new VS-mode software takes advantage of vstimecmp directly.) | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#15-1-3-hypervisor-instructions)15.1.3\. Hypervisor Instructions The hypervisor extension adds virtual-machine load and store instructions and two privileged fence instructions. #### [](#15-1-3-1-hypervisor-virtual-machine-load-and-store-instructions)15.1.3.1\. Hypervisor Virtual-Machine Load and Store Instructions ![svg](_images/svg-7ee5c36eedcec558e882f8ea3274cf450cfd841a.svg) The hypervisor virtual-machine load and store instructions are valid only in M-mode or HS-mode, or in U-mode when `hstatus`.HU=1. Each instruction performs an explicit memory access with an effective privilege mode of VS or VU. The effective privilege mode of the explicit memory access is VU when `hstatus`.SPVP=0, and VS when `hstatus`.SPVP=1. As usual for VS-mode and VU-mode, two-stage address translation is applied, and the HS-level `sstatus`.SUM is ignored. HS-level `sstatus`.MXR makes execute-only pages readable by explicit loads for both stages of address translation (VS-stage and G-stage), whereas`vsstatus`.MXR affects only the first translation stage (VS-stage). For every RV32I or RV64I load instruction, LB, LBU, LH, LHU, LW, LWU, and LD, there is a corresponding virtual-machine load instruction: HLV.B, HLV.BU, HLV.H, HLV.HU, HLV.W, HLV.WU, and HLV.D. For every RV32I or RV64I store instruction, SB, SH, SW, and SD, there is a corresponding virtual-machine store instruction: HSV.B, HSV.H, HSV.W, and HSV.D. Instructions HLV.WU, HLV.D, and HSV.D are not valid for RV32, of course. Instructions HLVX.HU and HLVX.WU are the same as HLV.HU and HLV.WU, except that _execute_ permission takes the place of _read_ permission during address translation. That is, the memory being read must be executable in both stages of address translation, but read permission is not required. For the supervisor physical address that results from address translation, the supervisor physical memory attributes must grant both _execute_ and _read_ permissions. (The _supervisor physical memory attributes_ are the machine’s physical memory attributes as modified by physical memory protection, [Physical Memory Protection](machine.html#pmp), for supervisor level.) | | HLVX cannot override machine-level physical memory protection (PMP), so attempting to read memory that PMP designates as execute-only still results in an access-fault exception. Although HLVX instructions’ explicit memory accesses require execute permissions, they still raise the same exceptions as other load instructions, rather than raising fetch exceptions instead. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | HLVX.WU is valid for RV32, even though LWU and HLV.WU are not. (For RV32, HLVX.WU can be considered a variant of HLV.W, as sign extension is irrelevant for 32-bit values.) The memory accesses performed by the `HLVX.*` instructions are not subject to pointer masking (see [Pointer Masking Extensions](zpm.html)). | | HLVX.\* instructions, designed for emulating implicit access to fetch instructions from guest memory, perform memory accesses that are exempt from pointer masking to facilitate this emulation. For the same reason, pointer masking does not apply when MXR is set. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Attempts to execute a virtual-machine load/store instruction (HLV, HLVX, or HSV) when V=1 cause a virtual-instruction exception. Attempts to execute one of these same instructions from U-mode when `hstatus`.HU=0 cause an illegal-instruction exception. #### [](#hfence.vma)15.1.3.2\. Hypervisor Memory-Management Fence Instructions ![svg](_images/svg-18753613efe684d2f90b312d0abb8044757cc6e8.svg) The hypervisor memory-management fence instructions, HFENCE.VVMA and HFENCE.GVMA, perform a function similar to SFENCE.VMA ([Supervisor Memory-Management Fence Instruction](supervisor.html#sfence.vma)), except applying to the VS-level memory-management data structures controlled by CSR `vsatp`(HFENCE.VVMA) or the guest-physical memory-management data structures controlled by CSR `hgatp` (HFENCE.GVMA). Instruction SFENCE.VMA applies only to the memory-management data structures controlled by the current`satp` (either the HS-level `satp` when V=0 or `vsatp` when V=1). HFENCE.VVMA is valid only in M-mode or HS-mode. Its effect is much the same as temporarily entering VS-mode and executing SFENCE.VMA. Executing an HFENCE.VVMA guarantees that any previous stores already visible to the current hart are ordered before all implicit reads by that hart done for VS-stage address translation for instructions that * are subsequent to the HFENCE.VVMA, and * execute when `hgatp`.VMID has the same setting as it did when HFENCE.VVMA executed. Implicit reads need not be ordered when `hgatp`.VMID is different than at the time HFENCE.VVMA executed. If operand _rs1_≠`x0`, it specifies a single guest virtual address, and if operand _rs2_≠`x0`, it specifies a single guest address-space identifier (ASID). | | An HFENCE.VVMA instruction applies only to a single virtual machine, identified by the setting of hgatp.VMID when HFENCE.VVMA executes. | | ------------------------------------------------------------------------------------------------------------------------------------------ | When _rs2_≠`x0`, bits XLEN-1:ASIDMAX of the value held in _rs2_ are reserved for future standard use. Until their use is defined by a standard extension, they should be zeroed by software and ignored by current implementations. Furthermore, if ASIDLEN < ASIDMAX, the implementation shall ignore bits ASIDMAX-1:ASIDLEN of the value held in _rs2_. | | Simpler implementations of HFENCE.VVMA can ignore the guest virtual address in _rs1_ and the guest ASID value in _rs2_, as well ashgatp.VMID, and always perform a global fence for the VS-level memory management of all virtual machines, or even a global fence for all memory-management data structures. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Neither `mstatus`.TVM nor `hstatus`.VTVM causes HFENCE.VVMA to trap. HFENCE.GVMA is valid only in HS-mode when `mstatus`.TVM=0, or in M-mode (irrespective of `mstatus`.TVM). Executing an HFENCE.GVMA instruction guarantees that any previous stores already visible to the current hart are ordered before all implicit reads by that hart done for G-stage address translation for instructions that follow the HFENCE.GVMA. If operand _rs1_≠`x0`, it specifies a single guest physical address, shifted right by 2 bits, and if operand_rs2_≠`x0`, it specifies a single virtual machine identifier (VMID). | | Conceptually, an implementation might contain two address-translation caches: one that maps guest virtual addresses to guest physical addresses, and another that maps guest physical addresses to supervisor physical addresses. HFENCE.GVMA need not flush the former cache, but it must flush entries from the latter cache that match the HFENCE.GVMA’s address and VMID arguments. More commonly, implementations contain address-translation caches that map guest virtual addresses directly to supervisor physical addresses, removing a level of indirection. For such implementations, any entry whose guest virtual address maps to a guest physical address that matches the HFENCE.GVMA’s address and VMID arguments must be flushed. Selectively flushing entries in this fashion requires tagging them with the guest physical address, which is costly, and so a common technique is to flush all entries that match the HFENCE.GVMA’s VMID argument, regardless of the address argument. Like for a guest physical address written to htval on a trap, a guest physical address specified in _rs1_ is shifted right by 2 bits to accommodate addresses wider than the current XLEN. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When _rs2_≠`x0`, bits XLEN-1:VMIDMAX of the value held in _rs2_ are reserved for future standard use. Until their use is defined by a standard extension, they should be zeroed by software and ignored by current implementations. Furthermore, if VMIDLEN < VMIDMAX, the implementation shall ignore bits VMIDMAX-1:VMIDLEN of the value held in _rs2_. | | Simpler implementations of HFENCE.GVMA can ignore the guest physical address in _rs1_ and the VMID value in _rs2_ and always perform a global fence for the guest-physical memory management of all virtual machines, or even a global fence for all memory-management data structures. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | If `hgatp`.MODE is changed for a given VMID, an HFENCE.GVMA with_rs1_\=`x0` (and _rs2_ set to either `x0` or the VMID) must be executed to order subsequent guest translations with the MODE change—even if the old MODE or new MODE is Bare. Attempts to execute HFENCE.VVMA or HFENCE.GVMA when V=1 cause a virtual-instruction exception, while attempts to do the same in U-mode cause an illegal-instruction exception. Attempting to execute HFENCE.GVMA in HS-mode when `mstatus`.TVM=1 also causes an illegal-instruction exception. ### [](#15-1-4-machine-level-csrs)15.1.4\. Machine-Level CSRs The hypervisor extension augments or modifies machine CSRs `mstatus`,`mstatush`, `mideleg`, `mip`, and `mie`, and adds CSRs `mtval2` and`mtinst`. #### [](#15-1-4-1-machine-status-mstatus-and-mstatush-registers)15.1.4.1\. Machine Status (`mstatus` and `mstatush`) Registers The hypervisor extension adds two fields, MPV and GVA, to the machine-level `mstatus` or `mstatush` CSR, and modifies the behavior of several existing `mstatus` fields.[Figure 33](#hypervisor-mstatus) shows the modified`mstatus` register when the hypervisor extension is implemented and MXLEN=64\. When MXLEN=32, the hypervisor extension adds MPV and GVA not to `mstatus` but to `mstatush`.[Figure 34](#hypervisor-mstatush) shows the`mstatush` register when the hypervisor extension is implemented and MXLEN=32. ![Machine status (`mstatus`) register for RV64 when the hypervisor extension is implemented.](_images/diag-dca39abcf45a1ea53c5d4d6f8f624b8b551d99f4.svg) Figure 33\. Machine status (`mstatus`) register for RV64 when the hypervisor extension is implemented. ![Additional machine status (`mstatush`) register for RV32 when the hypervisor extension is implemented. The format of `mstatus` is unchanged for RV32.](_images/diag-659e8497152f4217fa29db72bbf70c7848c0c46d.svg) Figure 34\. Additional machine status (`mstatush`) register for RV32 when the hypervisor extension is implemented. The format of `mstatus` is unchanged for RV32. The MPV bit (Machine Previous Virtualization Mode) is written by the implementation whenever a trap is taken into M-mode. Just as the MPP field is set to the (nominal) privilege mode at the time of the trap, the MPV bit is set to the value of the virtualization mode V at the time of the trap. When an MRET instruction is executed, the virtualization mode V is set to MPV, unless MPP=3, in which case V remains 0. Field GVA (Guest Virtual Address) is written by the implementation whenever a trap is taken into M-mode. For any trap (breakpoint, address misaligned, access fault, page fault, or guest-page fault) that writes a guest virtual address to `mtval`, GVA is set to 1\. For any other trap into M-mode, GVA is set to 0. The TSR and TVM fields of `mstatus` affect execution only in HS-mode, not in VS-mode. The TW field affects execution in all modes except M-mode. Setting TVM=1 prevents HS-mode from accessing `hgatp` or executing HFENCE.GVMA or HINVAL.GVMA, but has no effect on accesses to `vsatp` or instructions HFENCE.VVMA or HINVAL.VVMA. | | TVM exists in mstatus to allow machine-level software to modify the address translations managed by a supervisor-level OS, usually for the purpose of inserting another stage of address translation below that controlled by the OS. The instruction traps enabled by TVM=1 permit machine level to co-opt both satp and hgatp and substitute _shadow page tables_ that merge the OS’s chosen page translations with M-level’s lower-stage translations, all without the OS being aware. M-level software needs this ability not only to emulate the hypervisor extension if not already supported, but also to emulate any future RISC-V extensions that may modify or add address translation stages, perhaps, for example, to improve support for nested hypervisors, i.e., running hypervisors atop other hypervisors. However, setting TVM=1 does not cause traps for accesses to vsatp or instructions HFENCE.VVMA or HINVAL.VVMA, or for any actions taken in VS-mode, because M-level software is not expected to need to involve itself in VS-stage address translation. For virtual machines, it should be sufficient, and in all likelihood faster as well, to leave VS-stage address translation alone and merge all other translation stages into G-stage shadow page tables controlled by hgatp. This assumption does place some constraints on possible future RISC-V extensions that current machines will be able to emulate efficiently. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The hypervisor extension changes the behavior of the Modify Privilege field, MPRV, of `mstatus`. When MPRV=0, translation and protection behave as normal. When MPRV=1, explicit memory accesses are translated and protected, and endianness is applied, as though the current virtualization mode were set to MPV and the current nominal privilege mode were set to MPP. [Table 6](#h-mprv) enumerates the cases. __Table 6\. Effect of MPRV on the translation and protection of explicit memory accesses.__ | MPRV | MPV | MPP | Effect | | ---- | --- | --- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | \- | \- | Normal access; current privilege mode applies. | | 1 | 0 | 0 | U-level access with HS-level translation and protection only. | | 1 | 0 | 1 | HS-level access with HS-level translation and protection only. | | 1 | \- | 3 | M-level access with no translation. | | 1 | 1 | 0 | VU-level access with two-stage translation and protection. The HS-level MXR bit makes any executable page readable. vsstatus.MXR makes readable those pages marked executable at the VS translation stage, but only if readable at the guest-physical translation stage. | | 1 | 1 | 1 | VS-level access with two-stage translation and protection. The HS-level MXR bit makes any executable page readable. vsstatus.MXR makes readable those pages marked executable at the VS translation stage, but only if readable at the guest-physical translation stage.vsstatus.SUM applies instead of the HS-level SUM bit. | MPRV does not affect the virtual-machine load/store instructions, HLV, HLVX, and HSV. The explicit loads and stores of these instructions always act as though V=1 and the nominal privilege mode were`hstatus`.SPVP, overriding MPRV. The `mstatus` register is a superset of the HS-level `sstatus` register but is not a superset of `vsstatus`. #### [](#15-1-4-2-machine-interrupt-delegation-mideleg-register)15.1.4.2\. Machine Interrupt Delegation (`mideleg`) Register When the hypervisor extension is implemented, bits 10, 6, and 2 of`mideleg` (corresponding to the standard VS-level interrupts) are each read-only one. Furthermore, if any guest external interrupts are implemented (GEILEN is nonzero), bit 12 of `mideleg` (corresponding to supervisor-level guest external interrupts) is also read-only one. VS-level interrupts and guest external interrupts are always delegated past M-mode to HS-mode. For bits of `mideleg` that are zero, the corresponding bits in`hideleg`, `hip`, and `hie` are read-only zeros. #### [](#15-1-4-3-machine-interrupt-mip-and-mie-registers)15.1.4.3\. Machine Interrupt (`mip` and `mie`) Registers The hypervisor extension gives registers `mip` and `mie` additional active bits for the hypervisor-added interrupts. [Figure 35](#hypervisor-mipreg-standard) and [Figure 36](#hypervisor-miereg-standard) show the standard portions (bits 15:0) of registers `mip` and `mie` when the hypervisor extension is implemented. ![Standard portion (bits 15:0) of `mip`.](_images/diag-07326d5d93f0744523dadc60c93323573d46d402.svg) Figure 35\. Standard portion (bits 15:0) of `mip`. ![Standard portion (bits 15:0) of `mie`.](_images/diag-e5d65520c28944a61d575b540066004fcfa39fda.svg) Figure 36\. Standard portion (bits 15:0) of `mie`. Bits SGEIP, VSEIP, VSTIP, and VSSIP in `mip` are aliases for the same bits in hypervisor CSR `hip`, while SGEIE, VSEIE, VSTIE, and VSSIE in`mie` are aliases for the same bits in `hie`. #### [](#15-1-4-4-machine-second-trap-value-mtval2-register)15.1.4.4\. Machine Second Trap Value (`mtval2`) Register The `mtval2` register is an MXLEN-bit read/write register formatted as shown in [Figure 37](#mtval2reg). When a trap is taken into M-mode, `mtval2` is written with additional exception-specific information, alongside `mtval`, to assist software in handling the trap. ![Machine second trap value register (`mtval2`).](_images/diag-5634e190eb0777099692d3d4eeafb7a1bccec6ef.svg) Figure 37\. Machine second trap value register (`mtval2`). When a guest-page-fault trap is taken into M-mode, `mtval2` is written with either zero or the guest physical address that faulted, shifted right by 2 bits. For other traps, `mtval2` is set to zero, but a future standard or extension may redefine `mtval2’s` setting for other traps. If a guest-page fault is due to an implicit memory access during first-stage (VS-stage) address translation, a guest physical address written to `mtval2` is that of the implicit memory access that faulted. Additional information is provided in CSR `mtinst` to disambiguate such situations. Otherwise, for misaligned loads and stores that cause guest-page faults, a nonzero guest physical address in `mtval2` corresponds to the faulting portion of the access as indicated by the virtual address in `mtval`. For instruction guest-page faults on systems with variable-length instructions, a nonzero `mtval2` corresponds to the faulting portion of the instruction as indicated by the virtual address in `mtval`. `mtval2` is a **WARL** register that must be able to hold zero and may be capable of holding only an arbitrary subset of other 2-bit-shifted guest physical addresses, if any. The Ssdbltrap extension (See ["Ssdbltrp" Double Trap Extension](ssdbltrp.html#ssdbltrp)) requires the implementation of the `mtval2` CSR. #### [](#15-1-4-5-machine-trap-instruction-mtinst-register)15.1.4.5\. Machine Trap Instruction (`mtinst`) Register The `mtinst` register is an MXLEN-bit read/write register formatted as shown in [Figure 38](#mtinstreg). When a trap is taken into M-mode, `mtinst` is written with a value that, if nonzero, provides information about the instruction that trapped, to assist software in handling the trap. The values that may be written to `mtinst` on a trap are documented in [15.1.6.3\. Transformed Instruction or Pseudoinstruction for mtinst or htinst](#tinst-vals). ![Machine trap instruction (`mtinst`) register.](_images/diag-b102d8b6a52595dc6d0dcb538762ae933ee465f5.svg) Figure 38\. Machine trap instruction (`mtinst`) register. `mtinst` is a **WARL** register that need only be able to hold the values that the implementation may automatically write to it on a trap. ### [](#two-stage-translation)15.1.5\. Two-Stage Address Translation Whenever the current virtualization mode V is 1, two-stage address translation and protection is in effect. For any virtual memory access, the original virtual address is converted in the first stage by VS-level address translation, as controlled by the `vsatp` register, into a_guest physical address_. The guest physical address is then converted in the second stage by guest physical address translation, as controlled by the `hgatp` register, into a supervisor physical address. The two stages are known also as VS-stage and G-stage translation. Although there is no option to disable two-stage address translation when V=1, either stage of translation can be effectively disabled by zeroing the corresponding `vsatp` or `hgatp` register. The `vsstatus` field MXR, which makes execute-only pages readable by explicit loads, only overrides VS-stage page protection. Setting MXR at VS-level does not override guest-physical page protections. Setting MXR at HS-level, however, overrides both VS-stage and G-stage execute-only permissions. When V=1, memory accesses that would normally bypass address translation are subject to G-stage address translation alone. This includes memory accesses made in support of VS-stage address translation, such as reads and writes of VS-level page tables. Machine-level physical memory protection applies to supervisor physical addresses and is in effect regardless of virtualization mode. #### [](#guest-addr-translation)15.1.5.1\. Guest Physical Address Translation The mapping of guest physical addresses to supervisor physical addresses is controlled by CSR `hgatp` ([15.1.2.10\. Hypervisor Guest Address Translation and Protection (hgatp) Register](#hgatp)). When the address translation scheme selected by the MODE field of`hgatp` is Bare, guest physical addresses are equal to supervisor physical addresses without modification, and no memory protection applies in the trivial translation of guest physical addresses to supervisor physical addresses. When `hgatp`.MODE specifies a translation scheme of Sv32x4, Sv39x4, Sv48x4, or Sv57x4, G-stage address translation is a variation on the usual page-based virtual address translation scheme of Sv32, Sv39, Sv48, or Sv57, respectively. In each case, the size of the incoming address is widened by 2 bits (to 34, 41, 50, or 59 bits). To accommodate the 2 extra bits, the root page table (only) is expanded by a factor of four to be 16 KiB instead of the usual 4 KiB. Matching its larger size, the root page table also must be aligned to a 16 KiB boundary instead of the usual 4 KiB page boundary. Except as noted, all other aspects of Sv32, Sv39, Sv48, or Sv57 are adopted unchanged for G-stage translation. Non-root page tables and all page table entries (PTEs) have the same formats as documented in [Sv32: Page-Based 32-bit Virtual-Memory Systems](supervisor.html#sv32), [Sv39: Page-Based 39-bit Virtual-Memory Systems](supervisor.html#sv39), [Sv48: Page-Based 48-bit Virtual-Memory Systems](supervisor.html#sv48), and ["Sv57: Page-Based 57-bit Virtual-Memory System](supervisor.html#sv57). For Sv32x4, an incoming guest physical address is partitioned into a virtual page number (VPN) and page offset as shown in[Figure 39](#sv32x4va). This partitioning is identical to that for an Sv32 virtual address as depicted in[Sv32 virtual address](supervisor.html#sv32va), except with 2 more bits at the high end in VPN\[1\]. (Note that the fields of a partitioned guest physical address also correspond one-for-one with the structure that Sv32 assigns to a physical address, depicted in[Sv32 virtual address](supervisor.html#sv32va).) ![Sv32x4 virtual address (guest physical address).](_images/diag-04304fa0cf5cff107257250c235d34cca62e2d58.svg) Figure 39\. Sv32x4 virtual address (guest physical address). For Sv39x4, an incoming guest physical address is partitioned as shown in [Figure 40](#sv39x4va). This partitioning is identical to that for an Sv39 virtual address as depicted in ["Sv39 virtual address](supervisor.html#sv39va), except with 2 more bits at the high end in VPN\[2\]. Address bits 63:41 must all be zeros, or else a guest-page-fault exception occurs. ![Sv39x4 virtual address (guest physical address).](_images/diag-c00743ef7bd9ba1e327bd5f46a2a9e72b17ae9dd.svg) Figure 40\. Sv39x4 virtual address (guest physical address). For Sv48x4, an incoming guest physical address is partitioned as shown in [Figure 41](#sv48x4va). This partitioning is identical to that for an Sv48 virtual address as depicted in["Sv48 virtual address](supervisor.html#sv48va), except with 2 more bits at the high end in VPN\[3\]. Address bits 63:50 must all be zeros, or else a guest-page-fault exception occurs. ![Sv48x4 virtual address (guest physical address).](_images/diag-b1c5cfb755c43cf32a006470f5c8d54e9902114a.svg) Figure 41\. Sv48x4 virtual address (guest physical address). For Sv57x4, an incoming guest physical address is partitioned as shown in [Figure 42](#sv57x4va). This partitioning is identical to that for an Sv57 virtual address as depicted in[Sv57 virtual address](supervisor.html#sv57va), except with 2 more bits at the high end in VPN\[4\]. Address bits 63:59 must all be zeros, or else a guest-page-fault exception occurs. ![Sv57x4 virtual address (guest physical address).](_images/diag-265cdaed432bfedc182b55792f50784d70ca0591.svg) Figure 42\. Sv57x4 virtual address (guest physical address). | | The page-based G-stage address translation scheme for RV32, Sv32x4, is defined to support a 34-bit guest physical address so that an RV32 hypervisor need not be limited in its ability to virtualize real 32-bit RISC-V machines, even those with 33-bit or 34-bit physical addresses. This may include the possibility of a machine virtualizing itself, if it happens to use 33-bit or 34-bit physical addresses. Multiplying the size and alignment of the root page table by a factor of four is the cheapest way to extend Sv32 to cover a 34-bit address. The possible wastage of 12 KiB for an unnecessarily large root page table is expected to be of negligible consequence for most (maybe all) real uses. A consistent ability to virtualize machines having as much as four times the physical address space as virtual address space is believed to be of some utility also for RV64\. For a machine implementing 39-bit virtual addresses (Sv39), for example, this allows the hypervisor extension to support up to a 41-bit guest physical address space without either necessitating hardware support for 48-bit virtual addresses (Sv48) or falling back to emulating the larger address space using shadow page tables. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The conversion of an Sv32x4, Sv39x4, Sv48x4, or Sv57x4 guest physical address is accomplished with the same algorithm used for Sv32, Sv39, Sv48, or Sv57, as presented in[Virtual Address Translation Process](supervisor.html#sv32algorithm), except that: * `hgatp` substitutes for the usual `satp`; * for the translation to begin, the effective privilege mode must be VS-mode or VU-mode; * when checking the U bit, the current privilege mode is always taken to be U-mode; and * guest-page-fault exceptions are raised instead of regular page-fault exceptions. For G-stage address translation, all memory accesses (including those made to access data structures for VS-stage address translation) are considered to be user-level accesses, as though executed in U-mode. Access type permissions—readable, writable, or executable—are checked during G-stage translation the same as for VS-stage translation. For a memory access made to support VS-stage address translation (such as to read/write a VS-level page table), permissions and the need to set A and/or D bits at the G-stage level are checked as though for an implicit load or store, not for the original access type. However, any exception is always reported for the original access type (instruction, load, or store/AMO). The G bit in all G-stage PTEs is currently not used. Until its use is defined by a standard extension, it should be cleared by software for forward compatibility, and must be ignored by hardware. | | G-stage address translation uses the identical format for PTEs as regular address translation, even including the U bit, due to the possibility of sharing some (or all) page tables between G-stage translation and regular HS-level address translation. Regardless of whether this usage will ever become common, we chose not to preclude it. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#15-1-5-2-guest-page-faults)15.1.5.2\. Guest-Page Faults Guest-page-fault traps may be delegated from M-mode to HS-mode under the control of CSR `medeleg`, but cannot be delegated to other privilege modes. On a guest-page fault, CSR `mtval` or `stval` is written with the faulting guest virtual address as usual, and `mtval2` or `htval` is written either with zero or with the faulting guest physical address, shifted right by 2 bits. CSR `mtinst` or `htinst` may also be written with information about the faulting instruction or other reason for the access, as explained in [15.1.6.3\. Transformed Instruction or Pseudoinstruction for mtinst or htinst](#tinst-vals). When an instruction fetch or a misaligned memory access straddles a page boundary, two different address translations are involved. When a guest-page fault occurs in such a circumstance, the faulting virtual address written to `mtval`/`stval` is the same as would be required for a regular page fault. Thus, the faulting virtual address may be a page-boundary address that is higher than the instruction’s original virtual address, if the byte at that page boundary is among the accessed bytes. When a guest-page fault is not due to an implicit memory access for VS-stage address translation, a nonzero guest physical address written to `mtval2`/`htval` shall correspond to the exact virtual address written to `mtval`/`stval`. #### [](#hyp-mm-fences)15.1.5.3\. Memory-Management Fences The behavior of the SFENCE.VMA instruction is affected by the current virtualization mode V. When V=0, the virtual-address argument is an HS-level virtual address, and the ASID argument is an HS-level ASID. The instruction orders stores only to HS-level address-translation structures with subsequent HS-level address translations. When V=1, the virtual-address argument to SFENCE.VMA is a guest virtual address within the current virtual machine, and the ASID argument is a VS-level ASID within the current virtual machine. The current virtual machine is identified by the VMID field of CSR `hgatp`, and the effective ASID can be considered to be the combination of this VMID with the VS-level ASID. The SFENCE.VMA instruction orders stores only to the VS-level address-translation structures with subsequent VS-stage address translations for the same virtual machine, i.e., only when `hgatp`.VMID is the same as when the SFENCE.VMA executed. Hypervisor instructions HFENCE.VVMA and HFENCE.GVMA provide additional memory-management fences to complement SFENCE.VMA. These instructions are described in [15.1.3.2\. Hypervisor Memory-Management Fence Instructions](#hfence.vma). [Physical Memory Protection and Paging](machine.html#pmp-vmem) discusses the intersection between physical memory protection (PMP) and page-based address translation. It is noted there that, when PMP settings are modified in a manner that affects either the physical memory that holds page tables or the physical memory to which page tables point, M-mode software must synchronize the PMP settings with the virtual memory system. For HS-level address translation, this is accomplished by executing in M-mode an SFENCE.VMA instruction with _rs1_\=`x0` and _rs2_\=`x0`, after the PMP CSRs are written. Synchronization with G-stage and VS-stage data structures is also needed. Executing an HFENCE.GVMA instruction with_rs1_\=`x0` and _rs2_\=`x0` suffices to flush all G-stage or VS-stage address-translation cache entries that have cached PMP settings corresponding to the final translated supervisor physical address. An HFENCE.VVMA instruction is not required. Similarly, if the setting of the PBMTE or ADUE bits in `menvcfg` are changed, an HFENCE.GVMA instruction with _rs1_\=`x0` and _rs2_\=`x0` suffices to synchronize with respect to the altered interpretation of G-stage and VS-stage PTEs' PBMT and A/D bit fields, respectively. By contrast, if the PBMTE or ADUE bits in `henvcfg` are changed, executing an HFENCE.VVMA with _rs1_\=`x0` and _rs2_\=`x0` suffices to synchronize with respect to the altered interpretation of VS-stage PTEs' PBMT and A/D bit fields for the currently active VMID. | | No mechanism is provided to atomically change vsatp and hgatptogether. Hence, to prevent speculative execution causing one guest’s VS-stage translations to be cached under another guest’s VMID, world-switch code should zero vsatp, then swap hgatp, then finally write the newvsatp value. Similarly, if henvcfg.PBMTE/ADUE need be world-switched, they should be switched after zeroing vsatp but before writing the new vsatpvalue, obviating the need to execute an HFENCE.VVMA instruction. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#pm-two-stage)15.1.5.4\. Interaction with Pointer Masking Guest physical addresses (GPAs) are 2 bits wider than the corresponding virtual address translation modes, resulting in additional address translation schemes Sv32x4, Sv39x4, Sv48x4, and Sv57x4 for translating guest physical addresses to supervisor physical addresses. When running with virtualization in VS/VU mode with `vsatp.MODE` \= Bare, this means that those two bits may be subject to pointer masking, depending on `hgatp.MODE` and `senvcfg.PMM`/`henvcfg.PMM` (for VU/VS mode). If `vsatp.MODE` != BARE, this issue does **not** apply. | | An implementation could mask those two bits on the TLB access path, but this can have a significant timing impact. Alternatively, an implementation may choose to "waste" TLB capacity by having up to 4 duplicate entries for each page. In this case, the pointer masking operation can be applied on the TLB refill path, where it is unlikely to affect timing. To support this approach, some TLB entries need to be flushed when PMLEN changes in a way that may affect these duplicate entries. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | To support implementations where (XLEN-PMLEN) can be less than the GPA width supported by `hgatp.MODE`, hypervisors should execute an `HFENCE.GVMA` with_rs1_\=`x0` if the `henvcfg.PMM` is changed from or to a value where (XLEN-PMLEN) is less than GPA width supported by the `hgatp` translation mode of that guest. Specifically, these cases are: * `PMLEN=7` and `hgatp.MODE=sv57x4` * `PMLEN=16` and `hgatp.MODE=sv57x4` * `PMLEN=16` and `hgatp.MODE=sv48x4` Implementation of an address-specific `HFENCE.GVMA` should either ignore the address argument, or should ignore the top masked GPA bits of entries when comparing for an address match. ### [](#15-1-6-traps)15.1.6\. Traps #### [](#sec:hcauses)15.1.6.1\. Trap Cause Codes The hypervisor extension augments the trap cause encoding.[Table 7](#hcauses) lists the possible M-mode and HS-mode trap cause codes when the hypervisor extension is implemented. Codes are added for VS-level interrupts (interrupts 2, 6, 10), for supervisor-level guest external interrupts (interrupt 12), for virtual-instruction exceptions (exception 22), and for guest-page faults (exceptions 20, 21, 23). Furthermore, environment calls from VS-mode are assigned cause 10, whereas those from HS-mode or S-mode use cause 9 as usual. __Table 7\. Machine and supervisor cause register (mcause and scause) values when the hypervisor extension is implemented.__ | Interrupt | Exception Code | Description | | ---------------------------- | -------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 1111 | 0123 | _Reserved_Supervisor software interruptVirtual supervisor software interruptMachine software interrupt | | 1111 | 4567 | _Reserved_Supervisor timer interruptVirtual supervisor timer interruptMachine timer interrupt | | 1111 | 891011 | _Reserved_Supervisor external interruptVirtual supervisor external interruptMachine external interrupt | | 1111 | 121314-15≥16 | Supervisor guest external interruptCounter-overflow interrupt _Reserved_ _Designated for platform use_ | | 0000000000000000000000000000 | 0123456789101112131415161718192021222324-3132-4748-63≥64 | Instruction address misalignedInstruction access fault Illegal instruction Breakpoint Load address misalignedLoad access faultStore/AMO address misalignedStore/AMO access faultEnvironment call from U-mode or VU-modeEnvironment call from HS-modeEnvironment call from VS-modeEnvironment call from M-modeInstruction page faultLoad page fault _Reserved_Store/AMO page faultDouble trap _Reserved_ Software checkHardware errorInstruction guest-page faultLoad guest-page faultVirtual instructionStore/AMO guest-page fault _Designated for custom use_ _Reserved_ _Designated for custom use_ _Reserved_ | HS-mode and VS-mode ECALLs use different cause values so they can be delegated separately. When V=1, a virtual-instruction exception (code 22) is normally raised instead of an illegal-instruction exception if the attempted instruction is _HS-qualified_ but is prevented from executing when V=1 either due to insufficient privilege or because the instruction is expressly disabled by a supervisor or hypervisor CSR such as `scounteren` or `hcounteren`. An instruction is _HS-qualified_ if it would be valid to execute in HS-mode (for some values of the instruction’s register operands), assuming fields TSR and TVM of CSR `mstatus` are both zero. A special rule applies for CSR instructions that access 32-bit high-half CSRs such as `cycleh` and `htimedeltah`. When V=1 and XLEN=32, an invalid attempt to access a high-half CSR raises a virtual-instruction exception instead of an illegal-instruction exception if the same CSR instruction for the corresponding _low-half_ CSR (e.g.`cycle` or`htimedelta`) is HS-qualified. | | When XLEN>32, an attempt to access a high-half CSR always raises an illegal-instruction exception. | | ----------------------------------------------------------------------------------------------------- | Specifically, a virtual-instruction exception is raised for the following cases: * in VS-mode, attempts to access a non-high-half counter CSR when the corresponding bit in `hcounteren` is 0 and the same bit in `mcounteren`is 1; * in VS-mode, if XLEN=32, attempts to access a high-half counter CSR when the corresponding bit in `hcounteren` is 0 and the same bit in`mcounteren` is 1; * in VU-mode, attempts to access a non-high-half counter CSR when the corresponding bit in either `hcounteren` or `scounteren` is 0 and the same bit in `mcounteren` is 1; * in VU-mode, if XLEN=32, attempts to access a high-half counter CSR when the corresponding bit in either `hcounteren` or `scounteren` is 0 and the same bit in `mcounteren` is 1; * in VS-mode or VU-mode, attempts to execute a hypervisor instruction (HLV, HLVX, HSV, or HFENCE); * in VS-mode or VU-mode, attempts to access an implemented non-high-half hypervisor CSR or VS CSR when the same access (read/write) would be allowed in HS-mode, assuming `mstatus`.TVM=0; * in VS-mode or VU-mode, if XLEN=32, attempts to access an implemented high-half hypervisor CSR or high-half VS CSR when the same access (read/write) to the CSR"s low-half partner would be allowed in HS-mode, assuming `mstatus`.TVM=0; * in VU-mode, attempts to execute WFI when `mstatus`.TW=0,or to execute a supervisor instruction (SRET or SFENCE); * in VU-mode, attempts to access an implemented non-high-half supervisor CSR when the same access (read/write) would be allowed in HS-mode, assuming `mstatus`.TVM=0; * in VU-mode, if XLEN=32, attempts to access an implemented high-half supervisor CSR when the same access to the CSR’s low-half partner would be allowed in HS-mode, assuming `mstatus`.TVM=0; * in VS-mode, attempts to execute WFI when `hstatus`.VTW=1 and`mstatus`.TW=0, unless the instruction completes within an implementation-specific, bounded time; * in VS-mode, attempts to execute SRET when `hstatus`.VTSR=1; and * in VS-mode, attempts to execute an SFENCE.VMA or SINVAL.VMA instruction or to access `satp`, when `hstatus`.VTVM=1. Other extensions to the RISC-V Privileged Architecture may add to the set of circumstances that cause a virtual-instruction exception when V=1. On a virtual-instruction trap, `mtval` or `stval` is written the same as for an illegal-instruction trap. | | It is not unusual that hypervisors must emulate the instructions that raise virtual-instruction exceptions, to support nested hypervisors or for other reasons. Machine level is expected ordinarily to delegate virtual-instruction traps directly to HS-level, whereas illegal-instruction traps are likely to be processed first in M-mode before being conditionally delegated (by software) to HS-level. Consequently, virtual-instruction traps are expected typically to be handled faster than illegal-instruction traps. When not emulating the trapping instruction, a hypervisor should convert a virtual-instruction trap into an illegal-instruction exception for the guest virtual machine. Because TSR and TVM in mstatus are intended to impact only S-mode (HS-mode), they are ignored for determining exceptions in VS-mode. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Fields FS and VS in registers `sstatus` and `vsstatus` deviate from the usual_HS-qualified_ rule. If an instruction is prevented from executing because FS or VS is zero in either `sstatus` or `vsstatus`, the exception raised is always an illegal-instruction exception, never a virtual-instruction exception. | | Early implementations of the H extension treated FS and VS in sstatus andvsstatus specially this way, and the behavior has been codified to maintain compatibility for software. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 8\. Synchronous exception priority when the hypervisor extension is implemented.__ | Priority | Exc.Code | Description | | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------- | | _Highest_ | 3 | Instruction address breakpoint | | 12, 20, 1 | During instruction address translation: First encountered page fault, guest-page fault, or access fault | | | 1 | With physical address for instruction: Instruction access fault | | | 22208, 9, 10, 1133 | Illegal instructionVirtual instructionInstruction address misalignedEnvironment callEnvironment break Load/store/AMO address breakpoint | | | 4,6 | Optionally: Load/store/AMO address misaligned | | | 13, 15, 21, 23, 5, 7 | During address translation for an explicit memory access: First encountered page fault, guest-page fault, or access fault | | | 5, 7 | With physical address for an explicit memory access: Load/store/AMO access fault | | | _Lowest_ | 4, 6 | If not higher priority: Load/store/AMO address misaligned | If an instruction may raise multiple synchronous exceptions, the decreasing priority order of [Table 8](#HSyncExcPrio)indicates which exception is taken and reported in `mcause` or `scause`. #### [](#15-1-6-2-trap-entry)15.1.6.2\. Trap Entry When a trap occurs in HS-mode or U-mode, it goes to M-mode, unless delegated by `medeleg` or `mideleg`, in which case it goes to HS-mode. When a trap occurs in VS-mode or VU-mode, it goes to M-mode, unless delegated by `medeleg` or `mideleg`, in which case it goes to HS-mode, unless further delegated by `hedeleg` or `hideleg`, in which case it goes to VS-mode. When a trap is taken into M-mode, virtualization mode V gets set to 0, and fields MPV and MPP in `mstatus` (or `mstatush`) are set according to[Table 9](#h-mpp). A trap into M-mode also writes fields GVA, MPIE, and MIE in `mstatus`/`mstatush` and writes CSRs `mepc`, `mcause`,`mtval`, `mtval2`, and `mtinst`. __Table 9\. Value of mstatus/mstatush fields MPV and MPP after a trap into M-mode. Upon trap return, MPV is ignored when MPP=3.__ | Previous Mode | MPV | MPP | | ------------------- | --- | --- | | U-modeHS-modeM-mode | 000 | 013 | | VU-modeVS-mode | 11 | 01 | When a trap is taken into HS-mode, virtualization mode V is set to 0, and `hstatus`.SPV and `sstatus`.SPP are set according to[Table 10](#h-spp). If V was 1 before the trap, field SPVP in`hstatus` is set the same as `sstatus`.SPP; otherwise, SPVP is left unchanged. A trap into HS-mode also writes field GVA in `hstatus`, fields SPIE and SIE in `sstatus`, and CSRs `sepc`, `scause`, `stval`,`htval`, and `htinst`. __Table 10\. Value of hstatus field SPV and sstatus field SPP after a trap into HS-mode.__ | Previous Mode | SPV | SPP | | -------------- | --- | --- | | U-modeHS-mode | 00 | 01 | | VU-modeVS-mode | 11 | 01 | When a trap is taken into VS-mode, `vsstatus`.SPP is set according to[Table 11](#h-vspp). Register `hstatus` and the HS-level`sstatus` are not modified, and the virtualization mode V remains 1\. A trap into VS-mode also writes fields SPIE and SIE in `vsstatus` and writes CSRs `vsepc`, `vscause`, and `vstval`. __Table 11\. Value of vsstatus field SPP after a trap into VS-mode.__ | Previous Mode | SPP | | -------------- | --- | | VU-modeVS-mode | 01 | #### [](#tinst-vals)15.1.6.3\. Transformed Instruction or Pseudoinstruction for `mtinst` or `htinst` On any trap into M-mode or HS-mode, one of these values is written automatically into the appropriate trap instruction CSR, `mtinst` or`htinst`: * zero; * a transformation of the trapping instruction; * a custom value (allowed only if the trapping instruction is non-standard); or * a special pseudoinstruction. Except when a pseudoinstruction value is required (described later), the value written to `mtinst` or `htinst` may always be zero, indicating that the hardware is providing no information in the register for this particular trap. | | The value written to the trap instruction CSR serves two purposes. The first is to improve the speed of instruction emulation in a trap handler, partly by allowing the handler to skip loading the trapping instruction from memory, and partly by obviating some of the work of decoding and executing the instruction. The second purpose is to supply, via pseudoinstructions, additional information about guest-page-fault exceptions caused by implicit memory accesses done for VS-stage address translation. A _transformation_ of the trapping instruction is written instead of simply a copy of the original instruction in order to minimize the burden for hardware yet still provide to a trap handler the information needed to emulate the instruction. An implementation may at any time reduce its effort by substituting zero in place of the transformed instruction. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | On an interrupt, the value written to the trap instruction register is always zero. On a synchronous exception, if a nonzero value is written, one of the following shall be true about the value: * Bit 0 is `1`, and replacing bit 1 with `1` makes the value into a valid encoding of a standard instruction. In this case, the instruction that trapped is the same kind as indicated by the register value, and the register value is the transformation of the trapping instruction, as defined later. For example, if bits 1:0 are binary `11` and the register value is the encoding of a standard LW (load word) instruction, then the trapping instruction is LW, and the register value is the transformation of the trapping LW instruction. * Bit 0 is `1`, and replacing bit 1 with `1` makes the value into an instruction encoding that is explicitly designated for a custom instruction (_not_ an unused reserved encoding). This is a _custom value_. The instruction that trapped is a non-standard instruction. The interpretation of a custom value is not otherwise specified by this standard. * The value is one of the special pseudoinstructions defined later, all of which have bits 1:0 equal to `00`. These three cases exclude a large number of other possible values, such as all those having bits 1:0 equal to binary `10`. A future standard or extension may define additional cases, thus allowing values that are currently excluded. Software may safely treat an unrecognized value in a trap instruction register the same as zero. | | To be forward-compatible with future revisions of this standard, software that interprets a nonzero value from mtinst or htinst must fully verify that the value conforms to one of the cases listed above. For instance, for RV64, discovering that bits 6:0 of mtinst are0000011 and bits 14:12 are 010 is not sufficient to establish that the first case applies and the trapping instruction is a standard LW instruction; rather, software must also confirm that bits 63:32 ofmtinst are all zeros. A future standard might define new values for 64-bit mtinst that are nonzero in bits 63:32 yet may coincidentally have in bits 31:0 the same bit patterns as standard RV64 instructions. Unlike for standard instructions, there is no requirement that the instruction encoding of a custom value be of the same \`\`kind'' as the instruction that trapped (or even have any correlation with the trapping instruction). | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | [Table 12](#tinst-values) shows the values that may be automatically written to the trap instruction register for each standard exception cause. For exceptions that prevent the fetching of an instruction, only zero or a pseudoinstruction value may be written. A custom value may be automatically written only if the instruction that traps is non-standard. A future standard or extension may permit other values to be written, chosen from the set of allowed values established earlier. __Table 12\. Values that may be automatically written to the trap instruction (mtinst or htinst) register on an exception trap.__ | Exception | Zero | TransformedStandardInstruction | Custom Value | Pseudoinstruction Value | | ------------------------------------------------------------------------------------------ | ------------ | ------------------------------ | ------------ | ----------------------- | | Instruction address misaligned | Yes | No | Yes | No | | Instruction access faultIllegal instructionBreakpointVirtual instruction | YesYesYesYes | NoNoNoNo | NoNoYesYes | NoNoNoNo | | Load address misalignedLoad access faultStore/AMO address misalignedStore/AMO access fault | YesYesYesYes | YesYesYesYes | YesYesYesYes | NoNoNoNo | | Environment call | Yes | No | Yes | No | | Instruction page faultLoad page faultStore/AMO page fault | YesYesYes | NoYesYes | NoYesYes | NoNoNo | | Instruction guest-page faultLoad guest-page faultStore/AMO guest-page fault | YesYesYes | NoYesYes | NoYesYes | YesYesYes | As enumerated in the table, a synchronous exception may write to the trap instruction register a standard transformation of the trapping instruction only for exceptions that arise from explicit memory accesses (from loads, stores, and AMO instructions). Accordingly, standard transformations are currently defined only for these memory-access instructions. If a synchronous trap occurs for a standard instruction for which no transformation has been defined, the trap instruction register shall be written with zero (or, under certain circumstances, with a special pseudoinstruction value). For a standard load instruction that is not a compressed instruction and is one of LB, LBU, LH, LHU, LW, LWU, LD, FLW, FLD, FLQ, or FLH, the transformed instruction has the format shown in[Figure 43](#transformedloadinst). ![Transformed load instruction (LB, LBU, LH, LHU, LW, LWU, LD, FLW, FLD, FLQ, or FLH). Fields funct3, rd, and opcode are the same as the trapping load instruction.](_images/svg-cf1ef4657bdfb0a4e1db1aa780b0db12900cc5ad.svg) Figure 43\. Transformed load instruction (LB, LBU, LH, LHU, LW, LWU, LD, FLW, FLD, FLQ, or FLH). Fields funct3, rd, and opcode are the same as the trapping load instruction. For a standard store instruction that is not a compressed instruction and is one of SB, SH, SW, SD, FSW, FSD, FSQ, or FSH, the transformed instruction has the format shown in[Figure 44](#transformedstoreinst). ![Transformed store instruction (SB, SH, SW, SD, FSW, FSD, FSQ, or FSH). Fields rs2, funct3, and opcode are the same as the trapping store instruction.](_images/svg-5260a745bf0aac5aab2066c278df0238deffc9d2.svg) Figure 44\. Transformed store instruction (SB, SH, SW, SD, FSW, FSD, FSQ, or FSH). Fields rs2, funct3, and opcode are the same as the trapping store instruction. For a standard atomic instruction (load-reserved, store-conditional, or AMO instruction), the transformed instruction has the format shown in [Figure 45](#transformedatomicinst). ![Transformed atomic instruction (load-reserved, store-conditional, or AMO instruction). All fields are the same as the trapping instruction except bits 19:15, Addr. Offset.](_images/svg-4ae05637ae8bdd401ab34c16eea4842b69ad5c0d.svg) Figure 45\. Transformed atomic instruction (load-reserved, store-conditional, or AMO instruction). All fields are the same as the trapping instruction except bits 19:15, Addr. Offset. For a standard virtual-machine load/store instruction (HLV, HLVX, or HSV), the transformed instruction has the format shown in [Figure 46](#transformedvmaccessinst). ![Transformed virtual-machine load/store instruction (HLV, HLVX, HSV). All fields are the same as the trapping instruction except bits 19:15, Addr. Offset](_images/svg-125b4667eb298016120eabea5e564734d12c3d5c.svg) Figure 46\. Transformed virtual-machine load/store instruction (HLV, HLVX, HSV). All fields are the same as the trapping instruction except bits 19:15, Addr. Offset In all the transformed instructions above, the Addr. Offset field that replaces the instruction’s rs1 field in bits 19:15 is the positive difference between the faulting virtual address (written to `mtval` or`stval`) and the original virtual address. This difference can be nonzero only for a misaligned memory access. Note also that, for basic loads and stores, the transformations replace the instruction’s immediate offset fields with zero. For a standard compressed instruction (16-bit size), the transformed instruction is found as follows: 1. Expand the compressed instruction to its 32-bit equivalent. 2. Transform the 32-bit equivalent instruction. 3. Replace bit 1 with a `0`. Bits 1:0 of a transformed standard instruction will be binary `01` if the trapping instruction is compressed and `11` if not. | | In decoding the contents of mtinst or htinst, once software has determined that the register contains the encoding of a standard basic load (LB, LBU, LH, LHU, LW, LWU, LD, FLW, FLD, FLQ, or FLH) or basic store (SB, SH, SW, SD, FSW, FSD, FSQ, or FSH), it is not necessary to confirm also that the immediate offset fields (31:25, and 24:20 or 11:7) are zeros. The knowledge that the register’s value is the encoding of a basic load/store is sufficient to prove that the trapping instruction is of the same kind. A future version of this standard may add information to the fields that are currently zeros. However, for backwards compatibility, any such information will be for performance purposes only and can safely be ignored. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For guest-page faults, the trap instruction register is written with a special pseudoinstruction value if: (a) the fault is caused by an implicit memory access for VS-stage address translation, and (b) a nonzero value (the faulting guest physical address) is written to`mtval2` or `htval`. If both conditions are met, the value written to`mtinst` or `htinst` must be taken from[Table 13](#pseudoinsts); zero is not allowed. __Table 13\. Special pseudoinstruction values for guest-page faults. The RV32 values are used when VSXLEN=32, and the RV64 values when VSXLEN=64.__ | Value | Meaning | | --------------------- | ------------------------------------------------------------------------------------------------------- | | 0x00002000 0x00002020 | 32-bit read for VS-stage address translation (RV32)32-bit write for VS-stage address translation (RV32) | | 0x00003000 0x00003020 | 64-bit read for VS-stage address translation (RV64)64-bit write for VS-stage address translation (RV64) | The defined pseudoinstruction values are designed to correspond closely with the encodings of basic loads and stores, as illustrated by[Table 14](#pseudoinsts-basis). __Table 14\. Standard instructions corresponding to the special pseudoinstructions of [Table 13](#pseudoinsts).__ | Encoding | Instruction | | --------------------- | ----------------------- | | 0x00002003 0x00002023 | lw x0,0(x0) sw x0,0(x0) | | 0x00003003 0x00003023 | ld x0,0(x0) sd x0,0(x0) | A _write_ pseudoinstruction (`0x00002020` or `0x00003020`) is used for the case that the machine is attempting automatically to update bits A and/or D in VS-level page tables. All other implicit memory accesses for VS-stage address translation will be reads. If a machine never automatically updates bits A or D in VS-level page tables (leaving this to software), the _write_ case will never arise. The fact that such a page table update must actually be atomic, not just a simple write, is ignored for the pseudoinstruction. | | If the conditions that necessitate a pseudoinstruction value can ever occur for M-mode, then mtinst cannot be entirely read-only zero; and likewise for HS-mode and htinst. However, in that case, the trap instruction registers may minimally support only values 0 and0x00002000 or 0x00003000, and possibly 0x00002020 or 0x00003020, requiring as few as one or two flip-flops in hardware, per register. There is no harm here in ignoring the atomicity requirement for page table updates, because a hypervisor is not expected in these circumstances to emulate an implicit memory access that fails. Rather, the hypervisor is given enough information about the faulting access to be able to make the memory accessible (e.g. by restoring a missing page of virtual memory) before resuming execution by retrying the faulting instruction. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#15-1-6-4-trap-return)15.1.6.4\. Trap Return The MRET instruction is used to return from a trap taken into M-mode. MRET first determines what the new privilege mode will be according to the values of MPP and MPV in `mstatus` or `mstatush`, as encoded in[Table 9](#h-mpp). MRET then in `mstatus`/`mstatush` sets MPV=0, MPP=0, MIE=MPIE, and MPIE=1\. Lastly, MRET sets the privilege mode as previously determined, and sets `pc`\=`mepc`. The SRET instruction is used to return from a trap taken into HS-mode or VS-mode. Its behavior depends on the current virtualization mode. When executed in M-mode or HS-mode (i.e., V=0), SRET first determines what the new privilege mode will be according to the values in`hstatus`.SPV and `sstatus`.SPP, as encoded in[Table 10](#h-spp). SRET then sets `hstatus`.SPV=0, and in`sstatus` sets SPP=0, SIE=SPIE, and SPIE=1\. Lastly, SRET sets the privilege mode as previously determined, and sets `pc`\=`sepc`. When executed in VS-mode (i.e., V=1), SRET sets the privilege mode according to [Table 11](#h-vspp), in `vsstatus` sets SPP=0, SIE=SPIE, and SPIE=1, and lastly sets `pc`\=`vsepc`. If the Ssdbltrp extension is implemented, when `SRET` is executed in HS-mode, if the new privilege mode is VU, the `SRET` instruction sets `vsstatus.SDT`to 0\. When executed in VS-mode, `vsstatus.SDT` is set to 0. 5.1. "Smcsrind/Sscsrind" Indirect CSR Access, Version 1.0 ==================== ## [](#indirect-csr)5.1\. "Smcsrind/Sscsrind" Indirect CSR Access, Version 1.0 ### [](#5-1-1-introduction)5.1.1\. Introduction Smcsrind/Sscsrind is an ISA extension that extends the indirect CSR access mechanism originally defined as part of the[Smaia/Ssaia extensions](https://github.com/riscv/riscv-aia), in order to make it available for use by other extensions without creating an unnecessary dependence on Smaia/Ssaia. This extension confers two benefits: 1. It provides a means to access an array of registers via CSRs without requiring allocation of large chunks of the limited CSR address space. 2. It enables software to access each of an array of registers by index, without requiring a switch statement with a case for each register. | | CSRs are accessed indirectly via this extension using select values, in contrast to being accessed directly using standard CSR numbers. A CSR accessible via one method may or may not be accessible via the other method. Select values are a separate address space from CSR numbers, and from tselect values in the Sdtrig extension. If a CSR is both directly and indirectly accessible, the CSR’s select value is unrelated to its CSR number. Further, Machine-level and Supervisor-level select values are separate address spaces from each other; however, Machine-level and Supervisor-level CSRs with the same select value may be defined by an extension as partial or full aliases with respect to each other. This typically would be done for CSRs that can be delegated from Machine-level to Supervisor-level. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The machine-level extension **Smcsrind** encompasses all added CSRs and all behavior modifications for a hart, over all privilege levels. For a supervisor-level environment, extension **Sscsrind** is essentially the same as Smcsrind except excluding the machine-level CSRs and behavior not directly visible to supervisor level. ### [](#body)5.1.2\. Machine-level CSRs | **Number** | **Privilege** | **Width** | **Name** | **Description** | | ---------- | ------------- | --------- | -------- | --------------------------------- | | 0x350 | MRW | XLEN | miselect | Machine indirect register select | | 0x351 | MRW | XLEN | mireg | Machine indirect register alias | | 0x352 | MRW | XLEN | mireg2 | Machine indirect register alias 2 | | 0x353 | MRW | XLEN | mireg3 | Machine indirect register alias 3 | | 0x355 | MRW | XLEN | mireg4 | Machine indirect register alias 4 | | 0x356 | MRW | XLEN | mireg5 | Machine indirect register alias 5 | | 0x357 | MRW | XLEN | mireg6 | Machine indirect register alias 6 | | | The mireg\* CSR numbers are not consecutive because miph is CSR number 0x354. | | -------------------------------------------------------------------------------- | The CSRs listed in the table above provide a window for accessing register state indirectly. The value of `miselect` determines which register is accessed upon read or write of each of the machine indirect alias CSRs (`mireg*`). `miselect` value ranges are allocated to dependent extensions, which specify the register state accessible via each`mireg_i_` register, for each `miselect` value. `miselect` is a WARL register. The `miselect` register implements at least enough bits to support all implemented `miselect` values (corresponding to the implemented extensions that utilize `miselect`/`mireg*` to indirectly access register state). The`miselect` register may be read-only zero if there are no extensions implemented that utilize it. Values of `miselect` with the most-significant bit set (bit XLEN - 1 = 1) are designated only for custom use, presumably for accessing custom registers through the alias CSRs. Values of `miselect` with the most-significant bit clear are designated only for standard use and are reserved until allocated to a standard architecture extension. If XLEN is changed, the most-significant bit of `miselect` moves to the new position, retaining its value from before. | | An implementation is not required to support any custom values formiselect. | | ------------------------------------------------------------------------------ | The behavior upon accessing `mireg*` from M-mode, while `miselect` holds a value that is not implemented, is UNSPECIFIED. | | It is expected that implementations will typically raise an illegal-instruction exception for such accesses, so that, for example, they can be identified as software bugs. Platform specs, profile specs, and/or the Privileged ISA spec may place more restrictions on behavior for such accesses. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Attempts to access `mireg*` while `miselect` holds a number in an allocated and implemented range results in a specific behavior that, for each combination of `miselect` and `mireg_i_`, is defined by the extension to which the `miselect` value is allocated. | | Ordinarily, each mireg**_i_** will access register state, access read-only 0 state, or raise an illegal-instruction exception. For RV32, if an extension defines an indirectly accessed register as 64 bits wide, it is recommended that the lower 32 bits of the register are accessed through one of mireg, mireg2, or mireg3, while the upper 32 bits are accessed through mireg4, mireg5, or mireg6, respectively. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Six \*ireg\* registers are defined in order to ensure that the needs of extensions in development are covered, with some room for growth. For example, for an siselect value associated with counter X, sireg/sireg2 could be used to access mhpmcounterX/mhpmeventX, while sireg4/sireg5 could access mhpmcounterXh/mhpmeventXh. Six \*ireg\* registers allows for accessing up to 3 CSR arrays per index (\*iselect) with RV32-only CSRs, or up to 6 CSR arrays per index value without RV32-only CSRs. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#5-1-3-supervisor-level-csrs)5.1.3\. Supervisor-level CSRs | **Number** | **Privilege** | **Width** | **Name** | **Description** | | ---------- | ------------- | --------- | -------- | ------------------------------------ | | 0x150 | SRW | XLEN | siselect | Supervisor indirect register select | | 0x151 | SRW | XLEN | sireg | Supervisor indirect register alias | | 0x152 | SRW | XLEN | sireg2 | Supervisor indirect register alias 2 | | 0x153 | SRW | XLEN | sireg3 | Supervisor indirect register alias 3 | | 0x155 | SRW | XLEN | sireg4 | Supervisor indirect register alias 4 | | 0x156 | SRW | XLEN | sireg5 | Supervisor indirect register alias 5 | | 0x157 | SRW | XLEN | sireg6 | Supervisor indirect register alias 6 | The CSRs in the table above are required if S-mode is implemented. The `siselect` register will support the value range 0..0xFFF at a minimum. A future extension may define a value range outside of this minimum range. Only if such an extension is implemented will `siselect` be required to support larger values. | | Requiring a range of 0–0xFFF for siselect, even though most or all of the space may be reserved or inaccessible, permits M-mode to emulate indirectly accessed registers in this implemented range, including registers that may be standardized in the future. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Values of `siselect` with the most-significant bit set (bit XLEN - 1 = 1) are designated only for custom use, presumably for accessing custom registers through the alias CSRs. Values of `siselect` with the most-significant bit clear are designated only for standard use and are reserved until allocated to a standard architecture extension. If XLEN is changed, the most-significant bit of `siselect` moves to the new position, retaining its value from before. The behavior upon accessing `sireg*` from M-mode or S-mode, while `siselect`holds a value that is not implemented at supervisor level, is UNSPECIFIED. | | It is recommended that implementations raise an illegal-instruction exception for such accesses, to facilitate possible emulation (by M-mode) of these accesses. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | An extension is considered not to be implemented at supervisor level if machine level has disabled the extension for S-mode, such as by the settings of certain fields in CSR menvcfg, for example. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Otherwise, attempts to access `sireg*` from M-mode or S-mode while`siselect` holds a number in a standard-defined and implemented range result in specific behavior that, for each combination of `siselect` and`sireg_i_`, is defined by the extension to which the `siselect` value is allocated. | | Ordinarily, each sireg**_i_** will access register state, access read-only 0 state, or, unless executing in a virtual machine (covered in the next section), raise an illegal-instruction exception. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Note that the widths of `siselect` and `sireg*` are always the current XLEN rather than SXLEN. Hence, for example, if MXLEN = 64 and SXLEN = 32, then these registers are 64 bits when the current privilege mode is M (running RV64 code) but 32 bits when the privilege mode is S (RV32 code). ### [](#5-1-4-virtual-supervisor-level-csrs)5.1.4\. Virtual Supervisor-level CSRs | **Number** | **Privilege** | **Width** | **Name** | **Description** | | ---------- | ------------- | --------- | --------- | -------------------------------------------- | | 0x250 | HRW | XLEN | vsiselect | Virtual supervisor indirect register select | | 0x251 | HRW | XLEN | vsireg | Virtual supervisor indirect register alias | | 0x252 | HRW | XLEN | vsireg2 | Virtual supervisor indirect register alias 2 | | 0x253 | HRW | XLEN | vsireg3 | Virtual supervisor indirect register alias 3 | | 0x255 | HRW | XLEN | vsireg4 | Virtual supervisor indirect register alias 4 | | 0x256 | HRW | XLEN | vsireg5 | Virtual supervisor indirect register alias 5 | | 0x257 | HRW | XLEN | vsireg6 | Virtual supervisor indirect register alias 6 | The CSRs in the table above are required if the hypervisor extension is implemented. These VS CSRs all match supervisor CSRs, and substitute for those supervisor CSRs when executing in a virtual machine (in VS-mode or VU-mode). The `vsiselect` register will support the value range 0..0xFFF at a minimum. A future extension may define a value range outside of this minimum range. Only if such an extension is implemented will `vsiselect`be required to support larger values. | | Requiring a range of 0–0xFFF for vsiselect, even though most or all of the space may be reserved or inaccessible, permits a hypervisor to emulate indirectly accessed registers in this implemented range, including registers that may be standardized in the future. More generally it is recommended that vsiselect and siselect be implemented with the same number of bits. This also avoids creation of a virtualization hole due to observable differences between vsiselect andsiselect widths. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Values of `vsiselect` with the most-significant bit set (bit XLEN - 1 = 1) are designated only for custom use, presumably for accessing custom registers through the alias CSRs. Values of `vsiselect` with the most-significant bit clear are designated only for standard use and are reserved until allocated to a standard architecture extension. If XLEN is changed, the most-significant bit of `vsiselect` moves to the new position, retaining its value from before. For alias CSRs `sireg*` and `vsireg*`, the hypervisor extension’s usual rules for when to raise a virtual-instruction exception (based on whether an instruction is HS-qualified) are not applicable. The rules given in this section for `sireg` and `vsireg` apply instead, unless overridden by the requirements specified in the section below, which take precedence over this section when extension Smstateen is also implemented. A virtual-instruction exception is raised for attempts from VS-mode or VU-mode to directly access `vsiselect` or `vsireg*`, or attempts from VU-mode to access `siselect` or `sireg*`. The behavior upon accessing `vsireg*` from M-mode or HS-mode, or accessing `sireg*` (really `vsireg*`) from VS-mode, while `vsiselect` holds a value that is not implemented at HS level, is UNSPECIFIED. | | It is recommended that implementations raise an illegal-instruction exception for such accesses, to facilitate possible emulation (by M-mode) of these accesses. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Otherwise, while `vsiselect` holds a number in a standard-defined and implemented range, attempts to access `vsireg*` from a sufficiently privileged mode, or to access `sireg*` (really `vsireg*`) from VS-mode, result in specific behavior that, for each combination of `vsiselect` and`vsireg_i_`, is defined by the extension to which the `vsiselect` value is allocated. | | Ordinarily, each vsireg**_i_** will access register state, access read-only 0 state, or raise an exception (either an illegal-instruction exception or, for select accesses from VS-mode, a virtual-instruction exception). When vsiselect holds a value that is implemented at HS level but not at VS level, attempts to access sireg\* (really vsireg\*) from VS-mode will typically raise a virtual-instruction exception. But there may be cases specific to an extension where different behavior is more appropriate. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Like `siselect` and `sireg*`, the widths of `vsiselect` and `vsireg*` are always the current XLEN rather than VSXLEN. Hence, for example, if HSXLEN = 64 and VSXLEN = 32, then these registers are 64 bits when accessed by a hypervisor in HS-mode (running RV64 code) but 32 bits for a guest OS in VS-mode (RV32 code). ### [](#5-1-5-access-control-by-the-state-enable-csrs)5.1.5\. Access control by the state-enable CSRs If extension Smstateen is implemented together with Smcsrind, bit 60 of state-enable register `mstateen0` controls access to `siselect`, `sireg*`,`vsiselect`, and `vsireg*`. When `mstateen0`\[60\]=0, an attempt to access one of these CSRs from a privilege mode less privileged than M-mode results in an illegal-instruction exception. As always, the state-enable CSRs do not affect the accessibility of any state when in M-mode, only in less privileged modes. For more explanation, see the documentation for extension [Smstateen](smstateen.html#]). Other extensions may specify that certain mstateen bits control access to registers accessed indirectly through `siselect` \+ `sireg*`, and/or`vsiselect` \+ `vsireg*`. However, regardless of any other mstateen bits, if`mstateen0`\[60\] = 1, a virtual-instruction exception is raised as described in the previous section for all attempts from VS-mode or VU-mode to directly access `vsiselect` or `vsireg*`, and for all attempts from VU-mode to access `siselect` or `sireg*`. If the hypervisor extension is implemented, the same bit is defined also in hypervisor CSR `hstateen0`, but controls access to only `siselect` and `sireg*`(really `vsiselect` and `vsireg*`), which is the state potentially accessible to a virtual machine executing in VS or VU-mode. When`hstateen0`\[60\]=0 and `mstateen0`\[60\]=1, all attempts from VS or VU-mode to access `siselect` or `sireg*` raise a virtual-instruction exception, not an illegal-instruction exception, regardless of the value of `vsiselect` or any other mstateen bit. Extension Ssstateen is defined as the supervisor-level view of Smstateen. Therefore, the combination of Sscsrind and Ssstateen incorporates the bit defined above for `hstateen0` but not that for`mstateen0`, since machine-level CSRs are not visible to supervisor level. | | CSR address space is reserved for a possible future "Sucsrind" extension that extends indirect CSR access to user mode. | | -------------------------------------------------------------------------------------------------------------------------- | 3.1. Machine-Level ISA, Version 1.13 ==================== ## [](#machine)3.1\. Machine-Level ISA, Version 1.13 This chapter describes the machine-level operations available inmachine-mode (M-mode), which is the highest privilege mode in a RISC-V hart. M-mode is used for low-level access to a hardware platform and is the first mode entered at reset. M-mode can also be used to implement features that are too difficult or expensive to implement in hardware directly. The RISC-V machine-level ISA contains a common core that is extended depending on which other privilege levels are supported and other details of the hardware implementation. ### [](#3-1-1-machine-level-csrs)3.1.1\. Machine-Level CSRs In addition to the machine-level CSRs described in this section,M-mode code can access all CSRs at lower privilege levels. #### [](#misa)3.1.1.1\. Machine ISA (`misa`) Register The `misa` CSR is a **WARL** read-write register reporting the ISA supported by the hart.This register must be readable in any implementation, but a value of zero can be returned to indicate the `misa` register has not been implemented, requiring that CPU capabilities be determined through a separate non-standard mechanism. ![Machine ISA register (misa)](_images/diag-c0894d1a6dba146f15f3229ab4f03434805a1b9e.svg) Figure 1\. Machine ISA register (misa) The MXL (Machine XLEN) field encodes the native base integer ISA width as shown in [Table 1](#norm:misa%5Fmxl%5Fenc). The MXL field is read-only. If `misa` is nonzero, the MXL field indicates the effective XLEN in M-mode, a constant termed _MXLEN_. XLEN is never greater than MXLEN, but XLEN might be smaller than MXLEN in less-privileged modes. __Table 1\. Encoding of MXL field in misa__ | MXL | XLEN | | --- | --------------- | | 123 | 3264 _Reserved_ | The `misa` CSR is MXLEN bits wide. | | The base width can be quickly ascertained using branches on the sign of the returned misa value, and possibly a shift left by one and a second branch on the sign. These checks can be written in assembly code without knowing the register width (MXLEN) of the hart. The base width is given by _MXLEN=2MXL+4_. The base width can also be found if misa is zero, by placing the immediate 2 in a register, then shifting the register left by 31 bits. If zero, the hart is RV32, else it is RV64. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The Extensions field encodes the presence of the standard extensions, with a single bit per letter of the alphabet (bit 0 encodes presence of extension "A" , bit 1 encodes presence of extension "B", through to bit 25 which encodes "Z"). The "I" bit will be set for the RV32I and RV64I base ISAs, and the "E" bit will be set for RV32E and RV64E. The Extensions field is a **WARL** field that can contain writable bits where the implementation allows the supported ISA to be modified. At reset, the Extensions field shall contain the maximal set of supported extensions, and "I" shall be selected over "E" if both are available. When a standard extension is disabled by clearing its bit in `misa`, the instructions and CSRs defined or modified by the extension revert to their defined or reserved behaviors as if the extension is not implemented. | | For a given RISC-V execution environment, an instruction, extension, or other feature of the RISC-V ISA is ordinarily judged to be _implemented_or not by the observable execution behavior in that environment. For example, the F extension is said to be implemented for an execution environment if and only if the instructions that the RISC-V Unprivileged ISA defines for F execute as specified. With this definition of _implemented_, disabling an extension by clearing its bit in misa results in the extension being considered_not implemented_ in M-mode. For example, setting misa.F=0 results in the F extension being not implemented for M-mode, because the F extension’s instructions will not act as the Unprivileged ISA requires but may instead raise an illegal-instruction exception. Defining the term _implemented_ based strictly on the observable behavior might conflict with other common understandings of the same word. In particular, although common usage may allow for the combination "implemented but disabled," in this document it is considered a contradiction of terms, because _disabled_ implies execution will not behave as required for the feature to be considered _implemented_. In the same vein, "implemented and enabled" is redundant here; "implemented" suffices. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | All bits that are reserved for future use must return zero when read. __Table 2\. Encoding of Extensions field in misa.__ | Bit | Character | Description | | ------------------------------------------ | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 012345678910111213141516171819202122232425 | ABCDEFGHIJKLMNOPQRSTUVWXYZ | Atomic extensionB extensionCompressed extensionDouble-precision floating-point extensionRV32E/64E base ISASingle-precision floating-point extension _Reserved_Hypervisor extensionRV32I/64I base ISA _Reserved_ _Reserved_ _Reserved_Integer Multiply/Divide extension _Tentatively reserved for User-Level Interrupts extension_ _Reserved_ _Tentatively reserved for Packed-SIMD extension_Quad-precision floating-point extension _Reserved_Supervisor mode implemented _Reserved_User mode implementedVector extension _Reserved_Non-standard extensions present _Reserved_ _Reserved_ | The "X" bit will be set if there are any non-standard extensions. When the "B" bit is 1, the implementation supports the instructions provided by the Zba, Zbb, and Zbs extensions. When the "B" bit is 0, it indicates that the implementation might not support one or more of the Zba, Zbb, or Zbs extensions. When the "M" bit is 1, the implementation supports all multiply and division instructions defined by the M extension. When the "M" bit is 0, it indicates that the implementation might not support those instructions. However if the Zmmul extension is supported then the multiply instructions it specifies are supported irrespective of the value of the "M" bit. When the "S" bit is 1, the implementation supports supervisor mode.When the "S" bit is 0, the implementation might not support supervisor mode. When the "U" bit is 1, the implementation supports user mode.When the "U" bit is 0, the implementation might not support user mode. | | The misa CSR exposes a rudimentary catalog of CPU features to machine-mode code. More extensive information can be obtained in machine mode by probing other machine registers, and examining other ROM storage in the system as part of the boot process. We require that lower privilege levels execute environment calls instead of reading CPU registers to determine features available at each privilege level. This enables virtualization layers to alter the ISA observed at any level, and supports a much richer command interface without burdening hardware designs. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The "E" bit is read-only. Unless `misa` is all read-only zero, the "E" bit always reads as the complement of the "I" bit.If an execution environment supports both RV32E and RV32I, software can select RV32E by clearing the "I" bit. If an ISA feature _x_ depends on an ISA feature _y_, then attempting to enable feature _x_ but disable feature _y_ results in both features being disabled.For example, setting "F"=0 and "D"=1 results in both "F" and "D" being cleared. Similarly, setting "U"=0 and "S"=1" results in both "U" and "S" being cleared. An implementation may impose additional constraints on the collective setting of two or more `misa` fields, in which case they function collectively as a single **WARL** field. An attempt to write an unsupported combination causes those bits to be set to some supported combination. Writing `misa` may increase IALIGN, e.g., by disabling the "C" extension. If an instruction that would write `misa` increases IALIGN, and the subsequent instruction’s address is not IALIGN-bit aligned, the write to `misa` is suppressed, leaving `misa` unchanged. When software enables an extension that was previously disabled, then all state uniquely associated with that extension is UNSPECIFIED, unless otherwise specified by that extension. | | Although one of the bits 25—​0 in misa being set to 1 implies that the corresponding feature is implemented, the inverse is not necessarily true: one of these bits being clear does not necessarily imply that the corresponding feature is not implemented. This follows from the fact that, when a feature is not implemented, the corresponding opcodes and CSRs become reserved, not necessarily illegal. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-1-2-machine-vendor-id-mvendorid-register)3.1.1.2\. Machine Vendor ID (`mvendorid`) Register The `mvendorid` CSR is a 32-bit read-only register providing the JEDEC manufacturer ID of the provider of the core. This register must be readable in any implementation, but a value of 0 can be returned to indicate the field is not implemented or that this is a non-commercial implementation. ![Vendor ID register (`mvendorid`)](_images/diag-ddfc4bbe6f46125fbb9cbae8cfbe1e26894a17e2.svg) Figure 2\. Vendor ID register (`mvendorid`) JEDEC manufacturer IDs are ordinarily encoded as a sequence of one-byte continuation codes `0x7f`, terminated by a one-byte ID not equal to`0x7f`, with an odd parity bit in the most-significant bit of each byte.`mvendorid` encodes the number of one-byte continuation codes in the Bank field, and encodes the final byte in the Offset field, discarding the parity bit. For example, the JEDEC manufacturer ID`0x7f 0x7f 0x7f 0x7f 0x7f 0x7f 0x7f 0x7f 0x7f 0x7f 0x7f 0x7f 0x8a`(twelve continuation codes followed by `0x8a`) would be encoded in the`mvendorid` CSR as `0x60a`. | | In JEDEC’s parlance, the bank number is one greater than the number of continuation codes; hence, the mvendorid Bank field encodes a value that is one less than the JEDEC bank number. Previously the vendor ID was to be a number allocated by RISC-V International, but this duplicates the work of JEDEC in maintaining a manufacturer ID standard. At time of writing, registering a manufacturer ID with JEDEC has a one-time cost of $500. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-1-3-machine-architecture-id-marchid-register)3.1.1.3\. Machine Architecture ID (`marchid`) Register The `marchid` CSR is an MXLEN-bit read-only register encoding the base microarchitecture of the hart. This register must be readable in any implementation, but a value of 0 can be returned to indicate the field is not implemented.The combination of `mvendorid` and `marchid` should uniquely identify the type of hart microarchitecture that is implemented. ![Machine Architecture ID (`marchid`) register](_images/diag-0172e586c48e82ee6d49172653ff84884762f2b3.svg) Figure 3\. Machine Architecture ID (`marchid`) register Open-source project architecture IDs are allocated globally by RISC-V International, and have non-zero architecture IDs with a zero most-significant-bit (MSB). Commercial architecture IDs are allocated by each commercial vendor independently, but must have the MSB set and cannot contain zero in the remaining MXLEN-1 bits. | | The intent is for the architecture ID to represent the microarchitecture associated with the project around which development occurs rather than a particular organization. Commercial fabrications of open-source designs should (and might be required by the license to) retain the original architecture ID. This will aid in reducing fragmentation and tool support costs, as well as provide attribution. Open-source architecture IDs are administered by RISC-V International and should only be allocated to released, functioning open-source projects. Commercial architecture IDs can be managed independently by any registered vendor but are required to have IDs disjoint from the open-source architecture IDs (MSB set) to prevent collisions if a vendor wishes to use both closed-source and open-source microarchitectures. The convention adopted within the following Implementation field can be used to segregate branches of the same architecture design, including by organization. The misa register also helps distinguish different variants of a design. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-1-4-machine-implementation-id-mimpid-register)3.1.1.4\. Machine Implementation ID (`mimpid`) Register The `mimpid` CSR provides a unique encoding of the version of the processor implementation. This register must be readable in any implementation, but a value of 0 can be returned to indicate that the field is not implemented.The Implementation value should reflect the design of the RISC-V processor itself and not any surrounding system. ![Machine Implementation ID (`mimpid`) register](_images/diag-e422f9ad698aaf51925b6573abeffd24b6f79b3c.svg) Figure 4\. Machine Implementation ID (`mimpid`) register | | The format of this field is left to the provider of the architecture source code, but will often be printed by standard tools as a hexadecimal string without any leading or trailing zeros, so the Implementation value can be left-justified (i.e., filled in from most-significant nibble down) with subfields aligned on nibble boundaries to ease human readability. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-1-5-hart-id-mhartid-register)3.1.1.5\. Hart ID (`mhartid`) Register The `mhartid` CSR is an MXLEN-bit read-only register containing the integer ID of the hardware thread running the code. This register must be readable in any implementation.Hart IDs might not necessarily be numbered contiguously in a multiprocessor system, butone hart must have a hart ID of zero. Hart IDs must be unique within the execution environment. ![Hart ID (`mhartid`) register](_images/diag-9cfb7f99d6bf6243afe25ba0f8194061441c7584.svg) Figure 5\. Hart ID (`mhartid`) register | | In certain cases, we must ensure exactly one hart runs some code (e.g., at reset), and so require one hart to have a known hart ID of zero. For efficiency, system implementers should aim to reduce the magnitude of the largest hart ID used in a system. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-1-6-machine-status-mstatus-and-mstatush-registers)3.1.1.6\. Machine Status (`mstatus` and `mstatush`) Registers The `mstatus` register is an MXLEN-bit read/write register formatted as shown in [Figure 6](#mstatusreg-rv32) for RV32 and [Figure 7](#mstatusreg) for RV64.The `mstatus` register keeps track of and controls the hart’s current operating state. A restricted view of `mstatus` appears as the `sstatus` register in the S-level ISA. ![Machine-mode status (`mstatus`) register for RV32](_images/svg-f9f82fa39c51771f2754bd4db2c720b4e72901c5.svg) Figure 6\. Machine-mode status (`mstatus`) register for RV32 ![Machine-mode status (`mstatus`) register for RV64](_images/svg-c4c6c230ca7f061db767de8401b8409554823c5c.svg) Figure 7\. Machine-mode status (`mstatus`) register for RV64 For RV32 only, `mstatush` is a 32-bit read/write register formatted as shown in [Figure 8](#mstatushreg). Bits 30:4 of `mstatush` generally contain the same fields found in bits 62:36 of `mstatus` for RV64\. Fields SD, SXL, and UXL do not exist in `mstatush`. ![Additional machine-mode status (`mstatush`) register for RV32.](_images/svg-7c2313b1d75ab45bd84f0e6a0926363eb8fe1668.svg) Figure 8\. Additional machine-mode status (`mstatush`) register for RV32. ##### [](#privstack)3.1.1.6.1\. Privilege and Global Interrupt-Enable Stack in `mstatus` register Global interrupt-enable bits, MIE and SIE, are provided for M-mode and S-mode respectively. These bits are primarily used to guarantee atomicity with respect to interrupt handlers in the current privilege mode. | | The global _x_IE bits are located in the low-order bits of mstatus, allowing them to be atomically set or cleared with a single CSR instruction. | | --------------------------------------------------------------------------------------------------------------------------------------------------- | When a hart is executing in privilege mode _x_, interrupts are globally enabled when _x_IE=1 and globally disabled when _x_IE=0. Interrupts for lower-privilege modes, _w_<_x_, are always globally disabled regardless of the setting of any global _w_IE bit for the lower-privilege mode. Interrupts for higher-privilege modes, _y_\>_x_, are always globally enabled regardless of the setting of the global _y_IE bit for the higher-privilege mode.Higher-privilege-level code can use separate per-interrupt enable bits to disable selected higher-privilege-mode interrupts before ceding control to a lower-privilege mode.If supervisor mode is not implemented, then SIE and SPIE are read-only 0. | | A higher-privilege mode _y_ could disable all of its interrupts before ceding control to a lower-privilege mode but this would be unusual as it would leave only a synchronous trap, non-maskable interrupt, or reset as means to regain control of the hart. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | To support nested traps, each privilege mode _x_ that can respond to interrupts has a two-level stack of interrupt-enable bits and privilege modes._x_PIE holds the value of the interrupt-enable bit active prior to the trap, and _x_PP holds the previous privilege mode. The _x_PP fields can only hold privilege modes up to _x_, soMPP is two bits wide and SPP is one bit wide. When a trap is taken from privilege mode _y_into privilege mode _x_, _x_PIE is set to the value of _x_IE; _x_IE is set to 0; and _x_PP is set to _y_. | | For lower privilege modes, any trap (synchronous or asynchronous) is usually taken at a higher privilege mode with interrupts disabled upon entry. The higher-level trap handler will either service the trap and return using the stacked information, or, if not returning immediately to the interrupted context, will save the privilege stack before re-enabling interrupts, so only one entry per stack is required. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | An MRET or SRET instruction is used to return from a trap in M-mode or S-mode respectively.When executing an _x_RET instruction, supposing_x_PP holds the value _y_, _x_IE is set to _x_PIE; the privilege mode is changed to _y_; _x_PIE is set to 1; and _x_PP is set to the least-privileged supported mode (U if U-mode is implemented, else M). If_y_≠M, _x_RET also sets MPRV=0. | | Setting _x_PP to the least-privileged supported mode on an _x_RET helps identify software bugs in the management of the two-level privilege-mode stack. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Trap handlers must be designed to neither enable interrupts nor cause exceptions during the phase of handling where the trap handler preserves the critical state information required to handle and resume from the trap. An exception or interrupt in this critical phase of trap handling may lead to a trap that can overwrite such critical state. This could result in the loss of data needed to recover from the initial trap. Further, if an exception occurs in the code path needed to handle traps, then such a situation may lead to an infinite loop of traps. To prevent this, trap handlers must be meticulously designed to identify and safely manage exceptions within their operational flow. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | _x_PP fields are **WARL** fields that can hold only privilege mode _x_ and any implemented privilege mode lower than _x_. If privilege mode _x_ is not implemented, then _x_PP must be read-only 0. | | M-mode software can determine whether a privilege mode is implemented by writing that mode to MPP then reading it back. If the machine provides only U and M modes, then only a single hardware storage bit is required to represent either 00 or 11 in MPP. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#machine-double-trap)3.1.1.6.2\. Double Trap Control in `mstatus` Register A double trap typically arises during a sensitive phase in trap handling operations — when an exception or interrupt occurs while the trap handler (the component responsible for managing these events) is in a non-reentrant state. This non-reentrancy usually occurs in the early phase of trap handling, wherein the trap handler has not yet preserved the necessary state to handle and resume from the trap. The occurrence of a trap during this phase can lead to an overwrite of critical state information, resulting in the loss of data needed to recover from the initial trap. The trap that caused this critical error condition is henceforth called the _unexpected trap_. Trap handlers are designed to neither enable interrupts nor cause exceptions during this phase of handling. However, managing Hardware-Error exceptions, which may occur unpredictably, presents significant challenges in trap handler implementation due to the potential risk of a double trap. The M-mode-disable-trap (`MDT`) bit is a WARL field introduced by the Smdbltrp extension. Upon reset, the `MDT` field is set to 1. When the `MDT` bit is set to 1 by an explicit CSR write, the `MIE` (Machine Interrupt Enable) bit is cleared to 0. For RV64, this clearing occurs regardless of the value written, if any, to the `MIE` bit by the same write. The `MIE` bit can only be set to 1 by an explicit CSR write if the `MDT` bit is already 0 or, for RV64, is being set to 0 by the same write(For RV32, the `MDT` bit is in `mstatush` and the `MIE` bit in `mstatus` register). When a trap is to be taken into M-mode, if the `MDT` bit is currently 0, it is then set to 1, and the trap is delivered as expected. However, if `MDT` is already set to 1, then this is an _unexpected trap_. When the Smrnmi extension is implemented, a trap caused by an RNMI is not considered an _unexpected trap_irrespective of the state of the `MDT` bit. A trap caused by an RNMI does not set the `MDT` bit. However,a trap that occurs when executing in M-mode with`mnstatus.NMIE` set to 0 is an _unexpected trap_. In the event of a _unexpected trap_, the handling is as follows: * When the Smrnmi extension is implemented and `mnstatus.NMIE` is 1, the hart traps to the RNMI handler. To deliver this trap, the `mnepc` and `mncause`registers are written with the values that the _unexpected trap_ would have written to the `mepc` and `mcause` registers respectively. The privilege mode information fields in the `mnstatus` register are written to indicate M-mode and its `NMIE` field is set to 0. | | The consequence of this specification is that on occurrence of double trap the RNMI handler is not provided with information that a trap reports in themtval and the mtval2 registers. This information, if needed, can be obtained by the RNMI handler by decoding the instruction at the address in mnepc and examining its source register contents. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * When the Smrnmi extension is not implemented, or if the Smrnmi extension is implemented and `mnstatus.NMIE` is 0, the hart enters a critical-error state without updating any architectural state, including the `pc`. This state involves ceasing execution, disabling all interrupts (including NMIs), and asserting a `critical-error` signal to the platform. Whether performance counters and timers are updated in the critical-error state is UNSPECIFIED. | | The actions performed by the platform when a hart asserts a critical-error signal are platform-specific. The range of possible actions include restarting the affected hart or restarting the entire platform, among others. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `MRET` and `SRET` instructions, when executed in M-mode, set the `MDT` bit to 0. If the new privilege mode is U, VS, or VU, then `sstatus.SDT` is also set to 0. Additionally, if it is VU, then `vsstatus.SDT` is also set to 0. The `MNRET` instruction, provided by the Smrnmi extension, sets the `MDT` bit to 0 if the new privilege mode is not M. If it is U, VS, or VU, then `sstatus.SDT` is also set to 0. Additionally, if it is VU, then `vsstatus.SDT` is also set to 0. ##### [](#xlen-control)3.1.1.6.3\. Base ISA Control in `mstatus` Register For RV64 harts, the SXL and UXL fields are **WARL** fields that control the value of XLEN for S-mode and U-mode, respectively. The encoding of these fields is the same as the MXL field of `misa`, shown in [Table 1](#norm:misa%5Fmxl%5Fenc). The effective XLEN in S-mode and U-mode are termed _SXLEN_ and _UXLEN_, respectively. When MXLEN=32, the SXL and UXL fields do not exist, and SXLEN=32 and UXLEN=32. When MXLEN=64, if S-mode is not supported, then SXL is read-only zero. Otherwise, it is a **WARL** field that encodes the current value of SXLEN.In particular, an implementation may make SXL be a read-only field whose value always ensures that SXLEN=MXLEN. When MXLEN=64, if U-mode is not supported, then UXL is read-only zero. Otherwise, it is a **WARL** field that encodes the current value of UXLEN.In particular, an implementation may make UXL be a read-only field whose value always ensures that UXLEN=MXLEN or UXLEN=SXLEN. If S-mode is implemented, the set of legal values that the UXL field may assume excludes those that would cause UXLEN to be greater than SXLEN. Whenever XLEN in any mode is set to a value less than the widest supported XLEN, all operations must ignore source operand register bits above the configured XLEN, and must sign-extend results to fill the entire widest supported XLEN in the destination register. Similarly, `pc` bits above XLEN are ignored, and when the `pc` is written, it is sign-extended to fill the widest supported XLEN. | | We require that operations always fill the entire underlying hardware registers with defined values to avoid implementation-defined behavior. To reduce hardware complexity, the architecture imposes no checks that lower-privilege modes have XLEN settings less than or equal to the next-higher privilege mode. In practice, such settings would almost always be a software bug, but machine operation is well-defined even in this case. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Some HINT instructions are encoded as integer computational instructions that overwrite their destination register with its current value, e.g.,`c.addi x8, 0`. When such a HINT is executed with XLEN < MXLEN and bits MXLEN..XLEN of the destination register not all equal to bit XLEN-1, it is implementation-defined whether bits MXLEN..XLEN of the destination register are unchanged or are overwritten with copies of bit XLEN-1. | | This definition allows implementations to elide register write-back for some HINTs, while allowing them to execute other HINTs in the same manner as other integer computational instructions. The implementation choice is observable only by privilege modes with an XLEN setting greater than the current XLEN; it is invisible to the current privilege mode. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#3-1-1-6-4-memory-privilege-in-mstatus-register)3.1.1.6.4\. Memory Privilege in `mstatus` Register The MPRV (Modify PRiVilege) bit modifies the _effective privilege mode_, i.e., the privilege level at which explicit memory accesses execute.When MPRV=0, explicit memory accesses behave as normal, using the translation and protection mechanisms of the current privilege mode. When MPRV=1, load and store memory addresses are translated and protected, and endianness is applied, as though the current privilege mode were set to MPP. Instruction address-translation and protection are unaffected by the setting of MPRV. MPRV is read-only 0 if U-mode is not supported. An MRET or SRET instruction that changes the privilege mode to a mode less privileged than M also sets MPRV=0. The MXR (Make eXecutable Readable) bit modifies the privilege with which loads access virtual memory.When MXR=0, only loads from pages marked readable (R=1 in [Sv32 page table entry](supervisor.html#sv32pte)) will succeed. When MXR=1, loads from pages marked either readable or executable (R=1 or X=1) will succeed. MXR has no effect when page-based virtual memory is not in effect. MXR is read-only 0 if S-mode is not supported. | | The MPRV and MXR mechanisms were conceived to improve the efficiency of M-mode routines that emulate missing hardware features, e.g., misaligned loads and stores. MPRV obviates the need to perform address translation in software. MXR allows instruction words to be loaded from pages marked execute-only. The current privilege mode and the privilege mode specified by MPP might have different XLEN settings. When MPRV=1, load and store memory addresses are treated as though the current XLEN were set to MPP’s XLEN, following the rules in [3.1.1.6.3\. Base ISA Control in mstatus Register](#xlen-control). | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The SUM (permit Supervisor User Memory access) bit modifies the privilege with which S-mode loads and stores access virtual memory.When SUM=0, S-mode memory accesses to pages that are accessible by U-mode (U=1 in [Sv32 page table entry](supervisor.html#sv32pte)) will fault. When SUM=1, these accesses are permitted. SUM has no effect when page-based virtual memory is not in effect.Note that, while SUM is ordinarily ignored when not executing in S-mode, it _is_ in effect when MPRV=1 and MPP=S. SUM is read-only 0 if S-mode is not supported or if `satp`.MODE is read-only 0. The MXR and SUM mechanisms only affect the interpretation of permissions encoded in page-table entries. In particular, they have no impact on whether access-fault exceptions are raised due to PMAs or PMP. ##### [](#3-1-1-6-5-endianness-control-in-mstatus-and-mstatush-registers)3.1.1.6.5\. Endianness Control in `mstatus` and `mstatush` Registers The MBE, SBE, and UBE bits in `mstatus` and `mstatush` are **WARL** fields that control the endianness of memory accesses other than instruction fetches.Instruction fetches are always little-endian. MBE controls whether non-instruction-fetch memory accesses made from M-mode (assuming `mstatus`.MPRV=0) are little-endian (MBE=0) or big-endian (MBE=1). If S-mode is not supported, SBE is read-only 0\. Otherwise, SBE controls whether explicit load and store memory accesses made from S-mode are little-endian (SBE=0) or big-endian (SBE=1). If U-mode is not supported, UBE is read-only 0\. Otherwise, UBE controls whether explicit load and store memory accesses made from U-mode are little-endian (UBE=0) or big-endian (UBE=1). For _implicit_ accesses to supervisor-level memory management data structures, such as page tables, endianness is always controlled by SBE. Since changing SBE alters the implementation’s interpretation of these data structures, if any such data structures remain in use across a change to SBE, M-mode software must follow such a change to SBE by executing an SFENCE.VMA instruction with _rs1_\=`x0` and _rs2_\=`x0`. | | Only in contrived scenarios will a given memory-management data structure be interpreted as both little-endian and big-endian. In practice, SBE will only be changed at runtime on world switches, in which case neither the old nor new memory-management data structure will be reinterpreted in a different endianness. In this case, no additional SFENCE.VMA is necessary, beyond what would ordinarily be required for a world switch. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If S-mode is supported, an implementation may make SBE be a read-only copy of MBE. If U-mode is supported, an implementation may make UBE be a read-only copy of either MBE or SBE. | | An implementation supports only little-endian memory accesses if fields MBE, SBE, and UBE are all read-only 0\. An implementation supports only big-endian memory accesses (aside from instruction fetches) if MBE is read-only 1 and SBE and UBE are each read-only 1 when S-mode and U-mode are supported. Volume I defines a hart’s address space as a circular sequence of 2XLEN bytes at consecutive addresses. The correspondence between addresses and byte locations is fixed and not affected by any endianness mode. Rather, the applicable endianness mode determines the order of mapping between memory bytes and a multibyte quantity (halfword, word, etc.). Standard RISC-V ABIs are expected to be purely little-endian-only or big-endian-only, with no accommodation for mixing endianness. Nevertheless, endianness control has been defined so as to permit, for instance, an OS of one endianness to execute user-mode programs of the opposite endianness. Consideration has been given also to the possibility of non-standard usages whereby software flips the endianness of memory accesses as needed. RISC-V instructions are uniformly little-endian to decouple instruction encoding from the current endianness settings, for the benefit of both hardware and software. Otherwise, for instance, a RISC-V assembler or disassembler would always need to know the intended active endianness, despite that the endianness mode might change dynamically during execution. In contrast, by giving instructions a fixed endianness, it is sometimes possible for carefully written software to be endianness-agnostic even in binary form, much like position-independent code. The choice to have instructions be only little-endian does have consequences, however, for RISC-V software that encodes or decodes machine instructions. In big-endian mode, such software must account for the fact that explicit loads and stores have endianness opposite that of instructions, for example by swapping byte order after loads and before stores. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#virt-control)3.1.1.6.6\. Virtualization Support in `mstatus` Register The TVM (Trap Virtual Memory) bit is a **WARL** field that supports intercepting supervisor virtual-memory management operations. When TVM=1, attempts to read or write the `satp` CSR or execute an SFENCE.VMA or SINVAL.VMA instruction while executing in S-mode will raise an illegal-instruction exception. When TVM=0, these operations are permitted in S-mode. TVM is read-only 0 when S-mode is not supported. | | The TVM mechanism improves virtualization efficiency by permitting guest operating systems to execute in S-mode, rather than classically virtualizing them in U-mode. This approach obviates the need to trap accesses to most S-mode CSRs. Trapping satp accesses and the SFENCE.VMA and SINVAL.VMA instructions provides the hooks necessary to lazily populate shadow page tables. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The TW (Timeout Wait) bit is a **WARL** field that supports intercepting the WFI instruction (see [3.1.3.3\. Wait for Interrupt](#wfi)). When TW=0, the WFI instruction may execute in modes less privileged than M when not prevented for some other reason. When TW=1, then if WFI is executed in any less-privileged mode, and it does not complete within an implementation-specific, bounded time limit, the WFI instruction causes an illegal-instruction exception. An implementation may have WFI always raise an illegal-instruction exception in modes less privileged than M when TW=1, even if there are pending globally-disabled interrupts when the instruction is executed. TW is read-only 0 when there are no modes less privileged than M. | | Trapping the WFI instruction can trigger a world switch to another guest OS, rather than wastefully idling in the current guest. | | ----------------------------------------------------------------------------------------------------------------------------------- | When S-mode is implemented, then executing WFI in U-mode causes an illegal-instruction exception, regardless of the value of the TW bit, unless the instruction completes within an implementation-specific, bounded time limit. The TSR (Trap SRET) bit is a **WARL** field that supports intercepting the supervisor exception return instruction, SRET. When TSR=1, attempts to execute SRET while executing in S-mode will raise an illegal-instruction exception. When TSR=0, this operation is permitted in S-mode. TSR is read-only 0 when S-mode is not supported. | | Trapping SRET is necessary to emulate the hypervisor extension (see["H" Extension for Hypervisor Support](hypervisor.html#hypervisor)) on implementations that do not provide it. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ##### [](#3-1-1-6-7-extension-context-status-in-mstatus-register)3.1.1.6.7\. Extension Context Status in `mstatus` Register Supporting substantial extensions is one of the primary goals of RISC-V, and hence we define a standard interface to allow unchanged privileged-mode code, particularly a supervisor-level OS, to support arbitrary user-mode state extensions. | | To date, the V extension is the only standard extension that defines additional state beyond the floating-point CSR and data registers. | | ------------------------------------------------------------------------------------------------------------------------------------------ | The FS\[1:0\] and VS\[1:0\] **WARL** fieldsand the XS\[1:0\] read-only field are used to reduce the cost of context save and restore by setting and tracking the current state of the floating-point unit and any other user-mode extensions respectively.The FS field encodes the status of the floating-point unit state, including the floating-point registers`f0`–`f31` and the CSRs `fcsr`, `frm`, and `fflags`. The VS field encodes the status of the vector extension state, including the vector registers `v0`–`v31` and the CSRs `vcsr`, `vxrm`, `vxsat`, `vstart`,`vl`, `vtype`, and `vlenb`. The XS field encodes the status of additional user-mode extensions and associated state.These fields can be checked by a context switch routine to quickly determine whether a state save or restore is required. If a save or restore is required, additional instructions and CSRs are typically required to effect and optimize the process. | | The design anticipates that most context switches will not need to save/restore state in either or both of the floating-point unit or other extensions, so provides a fast check via the SD bit. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The FS, VS, and XS fields use the same status encoding as shown in[Table 3](#norm:mstatus%5Ffs%5Fvs%5Fxs%5Fenc), with the four possible status values being Off, Initial, Clean, and Dirty. __Table 3\. Encoding of FS\[1:0\], VS\[1:0\], and XS\[1:0\] status fields__ | Status | FS and VS Meaning | XS Meaning | | ------ | -------------------- | ------------------------------------------------------------------- | | 0123 | OffInitialCleanDirty | All offNone dirty or clean, some onNone dirty, some cleanSome dirty | If the F extension is implemented, the FS field shall not be read-only zero. If neither the F extension nor S-mode is implemented, then FS is read-only zero. If S-mode is implemented but the F extension is not, FS may optionally be read-only zero. | | Implementations with S-mode but without the F extension are permitted, but not required, to make the FS field be read-only zero. Some such implementations will choose _not_ to have the FS field be read-only zero, so as to enable emulation of the F extension for both S-mode and U-mode via invisible traps into M-mode. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If the `v` registers are implemented, the VS field shall not be read-only zero. If neither the `v` registers nor S-mode is implemented, then VS is read-only zero. If S-mode is implemented but the `v` registers are not, VS may optionally be read-only zero. In harts without additional user extensions requiring new state, the XS field is read-only zero. Every additional extension with state provides a CSR field that encodes the equivalent of the XS states. The XS field represents a summary of all extensions' status as shown in[Table 3](#norm:mstatus%5Ffs%5Fvs%5Fxs%5Fenc). | | The XS field effectively reports the maximum status value across all user-extension status fields, though individual extensions can use a different encoding than XS. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The SD bit is a read-only bit thatsummarizes whether either the FS, VS, or XS fields signal the presence of some dirty state that will require saving extended user context to memory. If FS, XS, and VS are all read-only zero, then SD is also always zero. When an extension’s status is set to Off, any instruction that attempts to read or write the corresponding state will cause an illegal-instruction exception. When the status is Initial, the corresponding state should have an initial constant value. When the status is Clean, the corresponding state is potentially different from the initial value, but matches the last value stored on a context swap. When the status is Dirty, the corresponding state has potentially been modified since the last context save. During a context save, the responsible privileged code need only write out the corresponding state if its status is Dirty, and can then reset the extension’s status to Clean. During a context restore, the context need only be loaded from memory if the status is Clean (it should never be Dirty at restore). If the status is Initial, the context must be set to an initial constant value on context restore to avoid a security hole, but this can be done without accessing memory. For example, the floating-point registers can all be initialized to the immediate value 0. The FS and XS fields are read by the privileged code before saving the context. The FS field is set directly by privileged code when resuming a user context, while the XS field is set indirectly by writing to the status register of the individual extensions.The status fields will also be updated during execution of instructions, regardless of privilege mode. Extensions to the user-mode ISA often include additional user-mode state, and this state can be considerably larger than the base integer registers. The extensions might only be used for some applications, or might only be needed for short phases within a single application. To improve performance, the user-mode extension can define additional instructions to allow user-mode software to return the unit to an initial state or even to turn off the unit. For example, a coprocessor might require to be configured before use and can be "unconfigured" after use. The unconfigured state would be represented as the Initial state for context save. If the same application remains running between the unconfigure and the next configure (which would set status to Dirty), there is no need to actually reinitialize the state at the unconfigure instruction, as all state is local to the user process, i.e., the Initial state may only cause the coprocessor state to be initialized to a constant value at context restore, not at every unconfigure. Executing a user-mode instruction to disable a unit and place it into the Off state will cause an illegal-instruction exception to be raised if any subsequent instruction tries to use the unit before it is turned back on. A user-mode instruction to turn a unit on must also ensure the unit’s state is properly initialized, as the unit might have been used by another context meantime. Changing the setting of FS has no effect on the contents of the floating-point register state. In particular, setting FS=Off does not destroy the state, nor does setting FS=Initial clear the contents. Similarly,the setting of VS has no effect on the contents of the vector register state.Other extensions, however, might not preserve state when set to Off. Implementations may choose to track the dirtiness of the floating-point register state imprecisely by reporting the state to be dirty even when it has not been modified. On some implementations, some instructions that do not mutate the floating-point state may cause the state to transition from Initial or Clean to Dirty.On other implementations,dirtiness might not be tracked at all, in which case the valid FS states are Off and Dirty, and an attempt to set FS to Initial or Clean causes it to be set to Dirty. | | This definition of FS does not disallow setting FS to Dirty as a result of errant speculation. Some platforms may choose to disallow speculatively writing FS to close a potential side channel. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If an instruction explicitly or implicitly writes a floating-point register or the `fcsr` but does not alter its contents, and FS=Initial or FS=Clean, it is implementation-defined whether FS transitions to Dirty. Implementations may choose to track the dirtiness of the vector register state in an analogous imprecise fashion, including possibly setting VS to Dirty when software attempts to set VS=Initial or VS=Clean. When VS=Initial or VS=Clean, it is implementation-defined whether an instruction that writes a vector register or vector CSR but does not alter its contents causes VS to transition to Dirty. [Table 4](#fsxsstates) shows all the possible state transitions for the FS, VS, or XS status bits. Note that the standard floating-point and vector extensions do not support user-mode unconfigure or disable/enable instructions. __Table 4\. FS, VS, and XS state transitions.__ | Current StateAction | Off | Initial | Clean | Dirty | | ------------------- | --- | ------- | ----- | ----- | | At context save in privileged code | | | | | | ---------------------------------- | ----- | --------- | ------- | -------- | | Save state?Next state | NoOff | NoInitial | NoClean | YesClean | | At context restore in privileged code | | | | | | ------------------------------------- | ----- | ---------------------- | --------------------- | ------ | | Restore state?Next state | NoOff | Yes, to initialInitial | Yes, from memoryClean | N/AN/A | | Execute instruction to read state | | | | | | --------------------------------- | ------------ | -------------- | ------------ | ------------ | | Action?Next state | ExceptionOff | ExecuteInitial | ExecuteClean | ExecuteDirty | | Execute instruction that possibly modifies state, including configuration | | | | | | ------------------------------------------------------------------------- | ------------ | ------------ | ------------ | ------------ | | Action?Next state | ExceptionOff | ExecuteDirty | ExecuteDirty | ExecuteDirty | | Execute instruction to unconfigure unit | | | | | | --------------------------------------- | ------------ | -------------- | -------------- | -------------- | | Action?Next state | ExceptionOff | ExecuteInitial | ExecuteInitial | ExecuteInitial | | Execute instruction to disable unit | | | | | | ----------------------------------- | ---------- | ---------- | ---------- | ---------- | | Action?Next state | ExecuteOff | ExecuteOff | ExecuteOff | ExecuteOff | | Execute instruction to enable unit | | | | | | ---------------------------------- | -------------- | -------------- | -------------- | -------------- | | Action?Next state | ExecuteInitial | ExecuteInitial | ExecuteInitial | ExecuteInitial | Standard privileged instructions to initialize, save, and restore extension state are provided to insulate privileged code from details of the added extension state by treating the state as an opaque object. | | Many coprocessor extensions are only used in limited contexts that allows software to safely unconfigure or even disable units when done. This reduces the context-switch overhead of large stateful coprocessors. We separate out floating-point state from other extension state, as when a floating-point unit is present the floating-point registers are part of the standard calling convention, and so user-mode software cannot know when it is safe to disable the floating-point unit. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The XS field provides a summary of all added extension state, but additional microarchitectural bits might be maintained in the extension to further reduce context save and restore overhead. The SD bit is read-only and is set when either the FS, VS, or XS bits encode a Dirty state (i.e., `SD=(FS==0b11 OR XS==0b11 OR VS==0b11)`). This allows privileged code to quickly determine when no additional context save is required beyond the integer register set and `pc`. The floating-point unit state is always initialized, saved, and restored using standard instructions (F, D, and/or Q), and privileged code must be aware of FLEN to determine the appropriate space to reserve for each`f` register. Machine and Supervisor modes share a single copy of the FS, VS, and XS bits. Supervisor-level software normally uses the FS, VS, and XS bits directly to record the status with respect to the supervisor-level saved context. Machine-level software must be more conservative in saving and restoring the extension state in their corresponding version of the context. | | In any reasonable use case, the number of context switches between user and supervisor level should far outweigh the number of context switches to other privilege levels. Note that coprocessors should not require their context to be saved and restored to service asynchronous interrupts, unless the interrupt results in a user-level context swap. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#3-1-1-6-8-previous-expected-landing-pad-elp-state-in-mstatus-register)3.1.1.6.8\. Previous Expected Landing Pad (ELP) State in `mstatus` Register The Zicfilp extension adds the `SPELP` and `MPELP` fields that hold the previous`ELP`, and are updated as specified in [Preserving Expected Landing Pad State on Traps](priv-cfi.html#ZICFILP%5FFORWARD%5FTRAPS). The **_x_**`PELP` fields are encoded as follows: * 0 - `NO_LP_EXPECTED` \- no landing pad instruction expected. * 1 - `LP_EXPECTED` \- a landing pad instruction is expected. #### [](#3-1-1-7-machine-trap-vector-base-address-mtvec-register)3.1.1.7\. Machine Trap-Vector Base-Address (`mtvec`) Register The `mtvec` register is an MXLEN-bit **WARL** read/write register that holds trap vector configuration, consisting of a vector base address (BASE) and a vector mode (MODE). ![Encoding of mtvec MODE field.](_images/diag-589a37244b7049ccd446c0e663599da666607981.svg) Figure 9\. Encoding of mtvec MODE field. The `mtvec` register must always be implemented, butcan contain a read-only value.If `mtvec` is writable, the set of values the register may hold can vary by implementation.The value in the BASE field must always be aligned on a 4-byte boundary, andthe MODE setting may impose additional alignment constraints on the value in the BASE field.Note that the CSR contains only bits XLEN-1 through 2 of the address BASE. When used as an address, the lower two bits are filled with zeroes to obtain an XLEN-bit address that is always aligned on a 4-byte boundary. | | We allow for considerable flexibility in implementation of the trap vector base address. On the one hand, we do not wish to burden low-end implementations with a large number of state bits, but on the other hand, we wish to allow flexibility for larger systems. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | __Table 5\. Encoding of mtvec MODE field.__ | Value | Name | Description | | ----- | ------------------ | ----------------------------------------------------------------------------------- | | 01≥2 | DirectVectored\--- | All traps set pc to BASE.Asynchronous interrupts set pc to BASE+4×cause. _Reserved_ | The encoding of the MODE field is shown in [Table 5](#norm:mtvec%5Fmode%5Fenc).When MODE=Direct, all traps into machine mode cause the `pc` to be set to the address in the BASE field. When MODE=Vectored, all synchronous exceptions into machine mode cause the `pc` to be set to the address in the BASE field, whereas interrupts cause the `pc` to be set to the address in the BASE field plus four times the interrupt cause number.For example, a machine-mode timer interrupt (see [Table 6](#norm:mcause%5Fexccode%5Fenc%5Fimg)) causes the `pc` to be set to BASE+`0x1c`. An implementation may have different alignment constraints for different modes. In particular, MODE=Vectored may have stricter alignment constraints than MODE=Direct. | | Allowing coarser alignments in Vectored mode enables vectoring to be implemented without a hardware adder circuit. Reset and NMI vector locations are given in a platform specification. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-1-8-machine-trap-delegation-medeleg-and-mideleg-registers)3.1.1.8\. Machine Trap Delegation (`medeleg` and `mideleg`) Registers By default, all traps at any privilege level are handled in machine mode, though a machine-mode handler can redirect traps back to the appropriate level with the MRET instruction ([3.1.3.2\. Trap-Return Instructions](#otherpriv)). To increase performance,implementations can provide individual read/write bits within `medeleg`and `mideleg` to indicate that certain exceptions and interrupts should be processed directly by a lower privilege level. The machine exception delegation register (`medeleg`) is a 64-bit read/write register. The machine interrupt delegation (`mideleg`) register is an MXLEN-bit read/write register. In harts with S-mode, the `medeleg` and `mideleg` registers must exist, and setting a bit in `medeleg` or `mideleg` will delegate the corresponding trap, when occurring in S-mode or U-mode, to the S-mode trap handler. In harts without S-mode, the `medeleg` and `mideleg` registers should not exist. | | In versions 1.9.1 and earlier , these registers existed but were hardwired to zero in M-mode only, or M/U without N harts. There is no reason to require they return zero in those cases, as the misaregister indicates whether they exist. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When a trap is delegated to S-mode,the `scause` register is written with the trap cause; the `sepc` register is written with the virtual address of the instruction that took the trap; the `stval` register is written with an exception-specific datum; the SPP field of `mstatus` is written with the active privilege mode at the time of the trap; the SPIE field of `mstatus` is written with the value of the SIE field at the time of the trap; and the SIE field of `mstatus` is cleared. The `mcause`, `mepc`, and `mtval` registers and the MPP and MPIE fields of`mstatus` are not written. An implementation can choose to subset the delegatable traps, with the supported delegatable bits found by writing one to every bit location, then reading back the value in `medeleg` or `mideleg` to see which bit positions hold a one. An implementation shall not have any bits of `medeleg` be read-only one, i.e., any synchronous trap that can be delegated must support not being delegated. Similarly,an implementation shall not fix as read-only one any bits of `mideleg` corresponding to machine-level interrupts(but may do so for lower-level interrupts). | | Version 1.11 and earlier prohibited having any bits of mideleg be read-only one. Platform standards may always add such restrictions. | | ---------------------------------------------------------------------------------------------------------------------------------------- | Traps never transition from a more-privileged mode to a less-privileged mode. For example, if M-mode has delegated illegal-instruction exceptions to S-mode, and M-mode software later executes an illegal instruction, the trap is taken in M-mode, rather than being delegated to S-mode.By contrast, traps may be taken horizontally. Using the same example, if M-mode has delegated illegal-instruction exceptions to S-mode, and S-mode software later executes an illegal instruction, the trap is taken in S-mode. Delegated interrupts result in the interrupt being masked at the delegator privilege level. For example, if the supervisor timer interrupt (STI) is delegated to S-mode by setting `mideleg`\[5\], STIs will not be taken when executing in M-mode. By contrast, if `mideleg`\[5\] is clear, STIs can be taken in any mode and regardless of current mode will transfer control to M-mode. ![Machine Exception Delegation (`medeleg`) register.](_images/diag-12b16eb0283c18b822b893c13363978243d6b368.svg) Figure 10\. Machine Exception Delegation (`medeleg`) register. `medeleg` has a bit position allocated for every synchronous exception shown in [Table 6](#norm:mcause%5Fexccode%5Fenc%5Fimg), with the index of the bit position equal to the value returned in the `mcause` register(i.e., setting bit 8 allows user-mode environment calls to be delegated to a lower-privilege trap handler). When XLEN=32, `medelegh` is a 32-bit read/write register that aliases bits 63:32 of `medeleg`. The `medelegh` register does not exist when XLEN=64. ![Machine Interrupt Delegation (`mideleg`) Register.](_images/diag-1faf03db279c406c2313714b6af38f73bb8e98d2.svg) Figure 11\. Machine Interrupt Delegation (`mideleg`) Register. `mideleg` holds trap delegation bits for individual interrupts, with the layout of bits matching those in the `mip` register (i.e., STIP interrupt delegation control is located in bit 5). For exceptions that cannot occur in less privileged modes, the corresponding `medeleg` bits should be read-only zero. In particular,`medeleg`\[11\] is read-only zero. The `medeleg`\[16\] is read-only zero as double trap is not delegatable. #### [](#3-1-1-9-machine-interrupt-mip-and-mie-registers)3.1.1.9\. Machine Interrupt (`mip` and `mie`) Registers The `mip` register is an MXLEN-bit read/write register containing information on pending interrupts, while `mie` is the corresponding MXLEN-bit read/write register containing interrupt enable bits. Interrupt cause number _i_ (as reported in CSR `mcause`,[3.1.1.15\. Machine Cause (mcause) Register](#mcause)) corresponds with bit _i_ in both `mip` and `mie`. Bits 15:0 are allocated to standard interrupt causes only, while bits 16 and above are designated for platform use. | | Interrupts designated for platform use may be designated for custom use at the platform’s discretion. | | -------------------------------------------------------------------------------------------------------- | ![Machine Interrupt-Pending (`mip`) register.](_images/diag-1d262afabf81c4bae6cf5ab46cb757a331aa464a.svg) Figure 12\. Machine Interrupt-Pending (`mip`) register. ![Machine Interrupt-Enable (`mie`) register](_images/diag-456ec1fe13b910a55f76c883703e2fcf7a6e3190.svg) Figure 13\. Machine Interrupt-Enable (`mie`) register An interrupt _i_ will trap to M-mode (causing the privilege mode to change to M-mode) if all of the following are true: (a) either the current privilege mode is M and the MIE bit in the `mstatus` register is set, or the current privilege mode has less privilege than M-mode; (b) bit _i_ is set in both `mip` and `mie`; and (c) if register`mideleg` exists, bit _i_ is not set in `mideleg`. These conditions for an interrupt trap to occur must be evaluated in a bounded amount of time from when an interrupt becomes, or ceases to be, pending in `mip`, and must also be evaluated immediately following the execution of an _x_RET instruction or an explicit write to a CSR on which these interrupt trap conditions expressly depend (including `mip`,`mie`, `mstatus`, and `mideleg`). Interrupts to M-mode take priority over any interrupts to lower privilege modes. Each individual bit in register `mip` may be writable or may be read-only. When bit _i_ in `mip` is writable, a pending interrupt _i_can be cleared by writing 0 to this bit. If interrupt _i_ can become pending but bit _i_ in `mip` is read-only, the implementation must provide some other mechanism for clearing the pending interrupt. A bit in `mie` must be writable if the corresponding interrupt can ever become pending. Bits of `mie` that are not writable must be read-only zero. The standard portions (bits 15:0) of the `mip` and `mie` registers are formatted as shown in [Figure 14](#norm:mip%5Fstd%5Fenc%5Fimg) and [Figure 15](#norm:mie%5Fstd%5Fenc%5Fimg) respectively. ![Standard portion (bits 15:0) of `mip`.](_images/diag-d9bf50577ae7c1e6834ddee7189f1a68498562af.svg) Figure 14\. Standard portion (bits 15:0) of `mip`. ![Standard portion (bits 15:0) of `mie`.](_images/diag-a13587a597b21dd19c8aea2b0306eacb055c911c.svg) Figure 15\. Standard portion (bits 15:0) of `mie`. | | The machine-level interrupt registers handle a few root interrupt sources which are assigned a fixed service priority for simplicity, while separate external interrupt controllers can implement a more complex prioritization scheme over a much larger set of interrupts that are then multiplexed into the machine-level interrupt sources. The non-maskable interrupt is not made visible via the mip register as its presence is implicitly known when executing the NMI trap handler. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Bits `mip`.MEIP and `mie`.MEIE are the interrupt-pending and interrupt-enable bits for machine-level external interrupts. MEIP is read-only in `mip`, and is set and cleared by a platform-specific interrupt controller. Bits `mip`.MTIP and `mie`.MTIE are the interrupt-pending and interrupt-enable bits for machine timer interrupts. MTIP is read-only in the `mip` register, and is cleared by writing to the memory-mapped machine-mode timer compare register. Bits `mip`.MSIP and `mie`.MSIE are the interrupt-pending and interrupt-enable bits for machine-level software interrupts. MSIP is read-only in `mip`, and is written by accesses to memory-mapped control registers, which are used to provide machine-level interprocessor interrupts. A hart’s memory-mapped `msip` register is a 32-bit read/write register, where bits 31—​1 read as zero and bit 0 contains the MSIP bit. When the memory-mapped `msip` register changes, it is guaranteed to be reflected in `mip`.MSIP eventually, but not necessarily immediately. If a system has only one hart, or if a platform standard supports the delivery of machine-level interprocessor interrupts through external interrupts (MEI) instead, then `mip`.MSIP and `mie`.MSIE may both be read-only zeros. If supervisor mode is not implemented, bits SEIP, STIP, and SSIP of`mip` and SEIE, STIE, and SSIE of `mie` are read-only zeros. If supervisor mode is implemented, bits `mip`.SEIP and `mie`.SEIE are the interrupt-pending and interrupt-enable bits for supervisor-level external interrupts. SEIP is writable in `mip`, and may be written by M-mode software to indicate to S-mode that an external interrupt is pending. Additionally, the platform-level interrupt controller may generate supervisor-level external interrupts.Supervisor-level external interrupts are made pending based on the logical-OR of the software-writable SEIP bit and the signal from the external interrupt controller. When `mip` is read with a CSR instruction, the value of the SEIP bit returned in the `rd` destination register is the logical-OR of the software-writable bit and the interrupt signal from the interrupt controller, but the signal from the interrupt controller is not used to calculate the value written to SEIP. Only the software-writable SEIP bit participates in the read-modify-write sequence of a CSRRS or CSRRC instruction. | | For example, if we name the software-writable SEIP bit B and the signal from the external interrupt controller E, then ifcsrrs t0, mip, t1 is executed, t0\[9\] is written with B \|| E, thenB is written with B | | t1\[9\]. If csrrw t0, mip, t1 is executed, then t0\[9\] is written with B | | E, and B is simply written witht1\[9\]. In neither case does B depend upon E. The SEIP field behavior is designed to allow a higher privilege layer to mimic external interrupts cleanly, without losing any real external interrupts. The behavior of the CSR instructions is slightly modified from regular CSR accesses as a result. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If supervisor mode is implemented, its `mip`.STIP and `mie`.STIE are the interrupt-pending and interrupt-enable bits for supervisor-level timer interrupts. If the stimecmp register is not implemented, STIP is writable in mip, and may be written by M-mode software to deliver timer interrupts to S-mode. If the `stimecmp` (supervisor-mode timer compare) register is implemented, STIP is read-only in mip and reflects the supervisor-level timer interrupt signal resulting from stimecmp. This timer interrupt signal is cleared by writing `stimecmp` with a value greater than the current time value. If supervisor mode is implemented, bits `mip`.SSIP and `mie`.SSIE are the interrupt-pending and interrupt-enable bits for supervisor-level software interrupts. SSIP is writable in `mip` and may also be set to 1 by a platform-specific interrupt controller. If the Sscofpmf extension is implemented, bits `mip`.LCOFIP and `mie`.LCOFIE are the interrupt-pending and interrupt-enable bits for local-counter-overflow interrupts. LCOFIP is read-write in `mip` and reflects the occurrence of a local counter-overflow overflow interrupt request resulting from any of the `mhpmevent_n_`.OF bits being set. If the Sscofpmf extension is not implemented, `mip`.LCOFIP and `mie`.LCOFIE are read-only zeros. Multiple simultaneous interrupts destined for M-mode are handled in the following decreasing priority order: MEI, MSI, MTI, SEI, SSI, STI, LCOFI. | | The machine-level interrupt fixed-priority ordering rules were developed with the following rationale. Interrupts for higher privilege modes must be serviced before interrupts for lower privilege modes to support preemption. The platform-specific machine-level interrupt sources in bits 16 and above have platform-specific priority, but are typically chosen to have the highest service priority to support very fast local vectored interrupts. External interrupts are handled before internal (timer/software) interrupts as external interrupts are usually generated by devices that might require low interrupt service times. Software interrupts are handled before internal timer interrupts, because internal timer interrupts are usually intended for time slicing, where time precision is less important, whereas software interrupts are used for inter-processor messaging. Software interrupts can be avoided when high-precision timing is required, or high-precision timer interrupts can be routed via a different interrupt path. Software interrupts are located in the lowest four bits of mip as these are often written by software, and this position allows the use of a single CSR instruction with a five-bit immediate. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Restricted views of the `mip` and `mie` registers appear as the `sip`and `sie` registers for supervisor level. If an interrupt is delegated to S-mode by setting a bit in the `mideleg` register, it becomes visible in the `sip` register and is maskable using the `sie` register. Otherwise, the corresponding bits in `sip` and `sie` are read-only zero. #### [](#3-1-1-10-hardware-performance-monitor)3.1.1.10\. Hardware Performance Monitor M-mode includes a basic hardware performance-monitoring facility. The `mcycle` CSR counts the number of clock cycles executed by the processor core on which the hart is running. The `minstret` CSR counts the number of instructions the hart has retired. The `mcycle` and `minstret` registers have 64-bit precision on all RV32 and RV64 harts. The counter registers have an arbitrary value after the hart is reset, and can be written with a given value. Any CSR write takes effect after the writing instruction has otherwise completed. The `mcycle` CSR may be shared between harts on the same core, in which case writes to `mcycle`will be visible to those harts. The platform should provide a mechanism to indicate which harts share an `mcycle` CSR. The hardware performance monitor includes 29 additional 64-bit event counters, `mhpmcounter3`\-`mhpmcounter31`. The event selector CSRs,`mhpmevent3`\-`mhpmevent31`, are 64-bit **WARL** registers that control which event causes the corresponding counter to increment. The meaning of these events is defined by the platform, but event 0 is defined to mean "no event." All counters should be implemented, buta legal implementation is to make both the counter and its corresponding event selector be read-only 0. ![Hardware performance monitor counters.](_images/diag-d9fc092ea45f5e72ef2104d5deaba43e644b7f89.svg) Figure 16\. Hardware performance monitor counters. The `mhpmcounters` are **WARL** registers thatsupport up to 64 bits of precision on RV32 and RV64. When XLEN=32, reads of the `mcycle`, `minstret`, `mhpmcounter_n_`, and `mhpmevent_n_`CSRs return bitj 31-0 of the corresponding register, and writes change only bits 31-0;reads of the `mcycleh`, `minstreth`, `mhpmcounter_n_h`, and `mhpmevent_n_h`CSRs return bits 63-32 of the corresponding register, and writes change only bits 63-32. The `mhpmevent_n_h` CSRs are provided only if the Sscofpmf extension is implemented. #### [](#mcounteren)3.1.1.11\. Machine Counter-Enable (`mcounteren`) Register The counter-enable `mcounteren` register is a 32-bit register thatcontrols the availability of the hardware performance-monitoring counters to the next-lower privileged mode. ![Counter-enable (`mcounteren`) register.](_images/diag-4239aeb3047e5b5faab786795bcb031aada45ba3.svg) Figure 17\. Counter-enable (`mcounteren`) register. The settings in this register only control accessibility. The act of reading or writing this register does not affect the underlying counters, which continue to increment even when not accessible. When the CY, TM, IR, or HPM_n_ bit in the `mcounteren` register is clear, attempts to read the `cycle`, `time`, `instret`, or`hpmcountern` register while executing in S-mode or U-mode will cause an illegal-instruction exception. When one of these bits is set, access to the corresponding register is permitted in the next implemented privilege mode (S-mode if implemented, otherwise U-mode). | | The counter-enable bits support two common use cases with minimal hardware. For harts that do not need high-performance timers and counters, machine-mode software can trap accesses and implement all features in software. For harts that need high-performance timers and counters but are not concerned with obfuscating the underlying hardware counters, the counters can be directly exposed to lower privilege modes. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | In addition, when the TM bit in the `mcounteren` register is clear, attempts to access the `stimecmp` or `vstimecmp` register while executing in a mode less privileged than M will cause an illegal-instruction exception. When this bit is set, access to the `stimecmp` or `vstimecmp` register is permitted in S-mode if implemented, and access to the `vstimecmp` register (via `stimecmp`) is permitted in VS-mode if implemented and not otherwise prevented by the TM bit in`hcounteren`. The `cycle`, `instret`, and `hpmcountern` CSRs are read-only shadows of `mcycle`, `minstret`, and `mhpmcounter n`, respectively. The `time` CSR is a read-only shadow of the memory-mapped `mtime` register.Analogously, when XLEN=32, the `cycleh`, `instreth` and `hpmcounternh` CSRs are read-only shadows of `mcycleh`, `minstreth` and `mhpmcounternh`, respectively. When XLEN=32, the `timeh` CSR is a read-only shadow of the upper 32 bits of the memory-mapped `mtime` register, while `time`shadows only the lower 32 bits of `mtime`. | | Implementations can convert reads of the time and timeh CSRs into loads to the memory-mapped mtime register, or emulate this functionality on behalf of less-privileged modes in M-mode software. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | In harts with U-mode, the `mcounteren` must be implemented, but all fields are **WARL** andmay be read-only zero, indicating reads to the corresponding counter will cause an illegal-instruction exception when executing in a less-privileged mode. In harts without U-mode, the `mcounteren` register should not exist. #### [](#3-1-1-12-machine-counter-inhibit-mcountinhibit-register)3.1.1.12\. Machine Counter-Inhibit (`mcountinhibit`) Register ![Counter-inhibit `mcountinhibit` register](_images/diag-4f012ecae5b40f8e88b7b98df2ea6fd27dabd1c8.svg) Figure 18\. Counter-inhibit `mcountinhibit` register The counter-inhibit register `mcountinhibit` is a 32-bit**WARL** register that controls which of the hardware performance-monitoring counters increment. The settings in this register only control whether the counters increment; their accessibility is not affected by the setting of this register. When the CY, IR, or HPM_n_ bit in the `mcountinhibit` register is clear, the `mcycle`, `minstret`, or `mhpmcountern` register increments as usual. When the CY, IR, or HPM_n_ bit is set, the corresponding counter does not increment. The `mcycle` CSR may be shared between harts on the same core, in which case the `mcountinhibit.CY` field is also shared between those harts, and so writes to `mcountinhibit.CY` will be visible to those harts. If the `mcountinhibit` register is not implemented, the implementation behaves as though the register were set to zero. | | When the mcycle and minstret counters are not needed, it is desirable to conditionally inhibit them to reduce energy consumption. Providing a single CSR to inhibit all counters also allows the counters to be atomically sampled. Because the mtime counter can be shared between multiple cores, it cannot be inhibited with the mcountinhibit mechanism. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-1-13-machine-scratch-mscratch-register)3.1.1.13\. Machine Scratch (`mscratch`) Register The `mscratch` register is an MXLEN-bit read/write register dedicated for use by machine mode. Typically, it is used to hold a pointer to a machine-mode hart-local context space and swapped with a user register upon entry to an M-mode trap handler. ![Machine-mode scratch register.](_images/diag-375dcad7e1c9a5d683ef8a0fcefd101ae5b4b245.svg) Figure 19\. Machine-mode scratch register. | | The MIPS ISA allocated two user registers (k0/k1) for use by the operating system. Although the MIPS scheme provides a fast and simple implementation, it also reduces available user registers, and does not scale to further privilege levels, or nested traps. It can also require both registers are cleared before returning to user level to avoid a potential security hole and to provide deterministic debugging behavior. The RISC-V user ISA was designed to support many possible privileged system environments and so we did not want to infect the user-level ISA with any OS-dependent features. The RISC-V CSR swap instructions can quickly save/restore values to the mscratch register. Unlike the MIPS design, the OS can rely on holding a value in the mscratch register while the user context is running. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-1-14-machine-exception-program-counter-mepc-register)3.1.1.14\. Machine Exception Program Counter (`mepc`) Register `mepc` is an MXLEN-bit read/write register formatted as shown in [Figure 20](#norm:mepc%5Fenc%5Fimg).The low bit of `mepc` (`mepc[0]`) is always zero. On implementations that support only IALIGN=32, the two low bits (`mepc[1:0]`) are always zero. If an implementation allows IALIGN to be either 16 or 32 (by changing CSR `misa`, for example), then, whenever IALIGN=32, bit `mepc[1]` is masked on reads so that it appears to be 0\. This masking occurs also for the implicit read by the MRET instruction. Though masked, `mepc[1]`remains writable when IALIGN=32. `mepc` is a **WARL** register that must be able to hold all valid virtual addresses. It need not be capable of holding all possible invalid addresses.Prior to writing `mepc`, implementations may convert an invalid address into some other invalid address that `mepc` is capable of holding. | | When address translation is not in effect, virtual addresses and physical addresses are equal. Hence, the set of addresses mepc must be able to represent includes the set of physical addresses that can be used as a valid pc or effective address. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When a trap is taken into M-mode, `mepc` is written with the virtual address of the instruction that was interrupted or that encountered the exception. Otherwise, `mepc` is never written by the implementation, though it may be explicitly written by software. ![Machine exception program counter register.](_images/diag-4f8212aedadd4ccdd1c08f7cb4985be93d1a2102.svg) Figure 20\. Machine exception program counter register. #### [](#mcause)3.1.1.15\. Machine Cause (`mcause`) Register The `mcause` register is an MXLEN-bit read-write register formatted as shown in [Figure 21](#norm:mcause%5Fenc%5Fimg).When a trap is taken into M-mode, `mcause` is written with a code indicating the event that caused the trap. Otherwise, `mcause` is never written by the implementation, though it may be explicitly written by software. The Interrupt bit in the `mcause` register is set if the trap was caused by an interrupt. The Exception Code field contains a code identifying the last exception or interrupt. [Table 6](#norm:mcause%5Fexccode%5Fenc%5Fimg) lists the possible machine-level exception codes.#The Exception Code is a **WLRL** field, so is only guaranteed to hold supported exception codes. ![Machine Cause (`mcause`) register.](_images/diag-781747e6db819b3cadebdde0acb5ae75cad09eb6.svg) Figure 21\. Machine Cause (`mcause`) register. Note that load and load-reserved instructions generate load exceptions, whereasstore, store-conditional, and AMO instructions generate store/AMO exceptions. | | Interrupts can be separated from other traps with a single branch on the sign of the mcause register value. A shift left can remove the interrupt bit and scale the exception codes to index into a trap vector table. We do not distinguish privileged instruction exceptions from illegal-instruction exceptions. This simplifies the architecture and also hides details of which higher-privilege instructions are supported by an implementation. The privilege level servicing the trap can implement a policy on whether these need to be distinguished, and if so, whether a given opcode should be treated as illegal or privileged. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | If an instruction may raise multiple synchronous exceptions, the decreasing priority order of [Table 7](#norm:exc%5Fpriority) indicates which exception is taken and reported in `mcause`.The priority of any custom synchronous exceptions is implementation-defined. __Table 6\. Machine cause (mcause) register values after trap.__ | Interrupt | Exception Code | Description | | ------------------------- | ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 1111 | 0123 | _Reserved_Supervisor software interrupt _Reserved_Machine software interrupt | | 1111 | 4567 | _Reserved_Supervisor timer interrupt _Reserved_Machine timer interrupt | | 1111 | 891011 | _Reserved_Supervisor external interrupt _Reserved_Machine external interrupt | | 1111 | 121314-15≥16 | _Reserved_Counter-overflow interrupt _Reserved_ _Designated for platform use_ | | 0000000000000000000000000 | 01234567891011121314151617181920-2324-3132-4748-63≥64 | Instruction address misalignedInstruction access faultIllegal instructionBreakpointLoad address misalignedLoad access faultStore/AMO address misalignedStore/AMO access faultEnvironment call from U-modeEnvironment call from S-mode _Reserved_Environment call from M-modeInstruction page faultLoad page fault _Reserved_Store/AMO page faultDouble trap _Reserved_Software checkHardware error _Reserved_ _Designated for custom use_ _Reserved_ _Designated for custom use_ _Reserved_ | __Table 7\. Synchronous exception priority in decreasing priority order.__ | Priority | Exc.Code | Description | | ------------ | ------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------- | | _Highest_ | 3 | Instruction address breakpoint | | 12, 1 | During instruction address translation:First encountered page fault or access fault | | | 1 | With physical address for instruction:Instruction access fault | | | 208,9,1133 | Illegal instructionInstruction address misalignedEnvironment callEnvironment breakLoad/store/AMO address breakpoint | | | 4,6 | Optionally:Load/store/AMO address misaligned | | | 13, 15, 5, 7 | During address translation for an explicit memory access:First encountered page fault or access fault | | | 5,7 | With physical address for an explicit memory access:Load/store/AMO access fault | | | _Lowest_ | 4,6 | If not higher priority:Load/store/AMO address misaligned | When a virtual address is translated into a physical address, the address translation algorithm determines what specific exception may be raised. Load/store/AMO address-misaligned exceptions may have either higher or lower priority than load/store/AMO page-fault and access-fault exceptions. | | The relative priority of load/store/AMO address-misaligned and page-fault exceptions is implementation-defined to flexibly cater to two design points. Implementations that never support misaligned accesses can unconditionally raise the misaligned-address exception without performing address translation or protection checks. Implementations that support misaligned accesses only to some physical addresses must translate and check the address before determining whether the misaligned access may proceed, in which case raising the page-fault exception or access is more appropriate. Instruction address breakpoints have the same cause value as, but different priority than, data address breakpoints (a.k.a. watchpoints) and environment break exceptions (which are raised by the EBREAK instruction). Instruction address-misaligned exceptions are raised by control-flow instructions with misaligned targets, rather than by the act of fetching an instruction. Therefore, these exceptions have lower priority than other instruction address exceptions. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | A software-check exception is a synchronous exception that is triggered when there are violations of checks and assertions defined by ISA extensions that aim to safeguard the integrity of software assets, including e.g. control-flow and memory-access constraints. When this exception is raised, the _x_tvalregister is set either to 0 or to an informative value defined by the extension that stipulated the exception be raised. The priority of this exception, relative to other synchronous exceptions, depends on the cause of this exception and is defined by the extension that stipulated the exception be raised. A hardware-error exception is a synchronous exception triggered when corrupted or uncorrectable data is accessed explicitly or implicitly by an instruction. In this context, "data" encompasses all types of information used within a RISC-V hart. Upon a hardware-error exception, the _x_epc register is set to the address of the instruction that attempted to access corrupted data, while the_x_tval register is set either to 0 or to the virtual address of an instruction fetch, load, or store that attempted to access corrupted data. The priority of hardware-error exception is implementation-defined, but any given occurrence is generally expected to be recognized at the point in the overall priority order at which the hardware error is discovered. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-1-16-machine-trap-value-mtval-register)3.1.1.16\. Machine Trap Value (`mtval`) Register The `mtval` register is an MXLEN-bit read-write register formatted as shown in [Figure 22](#norm:mtval%5Fenc%5Fimg).When a trap is taken into M-mode,`mtval` is either set to zero or written with exception-specific information to assist software in handling the trap. Otherwise, `mtval`is never written by the implementation, though it may be explicitly written by software. The hardware platform will specify which exceptions must set `mtval` informatively, which may unconditionally set it to zero, and which may exhibit either behavior, depending on the underlying event that caused the exception. If the hardware platform specifies that no exceptions set `mtval`to a nonzero value, then `mtval` is read-only zero. If `mtval` is written with a nonzero value when a breakpoint, address-misaligned, access-fault, page-fault, or hardware-error exception occurs on an instruction fetch, load, or store, then `mtval` will contain the faulting virtual address. On a breakpoint exception raised by an EBREAK or C.EBREAK instruction, `mtval`is written with either zero or the virtual address of the instruction. | | For breakpoint exceptions raised by \[C.\]EBREAK, the virtual address of the instruction is already recorded in mepc. Recording the same address in mtval is redundant; the option is provided for backwards compatibility. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | When page-based virtual memory is enabled, `mtval` is written with the faulting virtual address, even for physical-memory access-fault exceptions.This design reduces datapath cost for most implementations, particularly those with hardware page-table walkers. ![Machine Trap Value (`mtval`) register.](_images/diag-d5b13353e79d09b718c8ebe5165017d4b18af01e.svg) Figure 22\. Machine Trap Value (`mtval`) register. If `mtval` is written with a nonzero value when a misaligned load or store causes an access-fault, page-fault, or hardware-error exception, then `mtval` will contain the virtual address of the portion of the access that caused the fault. If `mtval` is written with a nonzero value when an instruction access-fault, page-fault, or hardware-error exception occurs on a hart with variable-length instructions, then `mtval` will contain the virtual address of the portion of the instruction that caused the fault, while`mepc` will point to the beginning of the instruction. The `mtval` register can optionally also be used to return the faulting instruction bits on an illegal-instruction exception (`mepc` points to the faulting instruction in memory). If `mtval` is written with a nonzero value when an illegal-instruction exception occurs, then `mtval`will contain the shortest of:# * the actual faulting instruction * the first ILEN bits of the faulting instruction * the first MXLEN bits of the faulting instruction The value loaded into `mtval` on an illegal-instruction exception is right-justified and all unused upper bits are cleared to zero. | | Capturing the faulting instruction in mtval reduces the overhead of instruction emulation, potentially avoiding several partial instruction loads if the instruction is misaligned, and likely data cache misses or slow uncached accesses when loads are used to fetch the instruction into a data register. There is also a problem of atomicity if another agent is manipulating the instruction memory, as might occur in a dynamic translation system. A requirement is that the entire instruction (or at least the first MXLEN bits) are fetched into mtval before taking the trap. This should not constrain implementations, which would typically fetch the entire instruction before attempting to decode the instruction, and avoids complicating software handlers. A value of zero in mtval signifies either that the feature is not supported, or an illegal zero instruction was fetched. A load from the instruction memory pointed to by mepc can be used to distinguish these two cases (or alternatively, the system configuration information can be interrogated to install the appropriate trap handling before runtime). | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | On a trap caused by a software-check exception, the `mtval` register holds the cause for the exception. The following encodings are defined: * 0 - No information provided. * 2 - Landing Pad Fault. Defined by the Zicfilp extension ([Landing Pad (Zicfilp)](priv-cfi.html#priv-forward)). * 3 - Shadow Stack Fault. Defined by the Zicfiss extension ([Shadow Stack (Zicfiss)](priv-cfi.html#priv-backward)). For other traps, `mtval` is set to zero, but a future standard may redefine `mtval`’s setting for other traps. If `mtval` is not read-only zero, it is a **WARL** register thatmust be able to hold all valid virtual addresses and the value zero.It need not be capable of holding all possible invalid addresses.Prior to writing `mtval`, implementations may convert an invalid address into some other invalid address that `mtval` is capable of holding. If the feature to return the faulting instruction bits is implemented,`mtval` must also be able to hold all values less than 2_N_, where_N_ is the smaller of MXLEN and ILEN. #### [](#3-1-1-17-machine-configuration-pointer-mconfigptr-register)3.1.1.17\. Machine Configuration Pointer (`mconfigptr`) Register The `mconfigptr` register is an MXLEN-bit read-only CSR formatted as shown in [Figure 23](#norm:mconfigptr%5Fenc%5Fimg), thatholds the physical address of a configuration data structure.Software can traverse this data structure to discover information about the harts, the platform, and their configuration. ![Machine Configuration Pointer (`mconfigptr`) register.](_images/diag-bbcca0b223d8541a702e6ff514c8b145f9b34b9f.svg) Figure 23\. Machine Configuration Pointer (`mconfigptr`) register. The pointer alignment in bits must be no smaller than MXLEN: i.e., if MXLEN is 8×_n_, then `mconfigptr`\[log2n\-1:0\] must be zero. The `mconfigptr` register must be implemented, butit may be zero to indicate the configuration data structure does not exist or that an alternative mechanism must be used to locate it. | | The format and schema of the configuration data structure have yet to be standardized. While the mconfigptr register will simply be hardwired in some implementations, other implementations may provide a means to configure the value returned on CSR reads. For example, mconfigptr might present the value of a memory-mapped register that is programmed by the platform or by M-mode software towards the beginning of the boot process. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#sec:menvcfg)3.1.1.18\. Machine Environment Configuration (`menvcfg`) Register The `menvcfg` CSR is a 64-bit read/write register, formatted as shown in [Figure 24](#norm:menvcfg%5Fenc%5Fimg), thatcontrols certain characteristics of the execution environment for modes less privileged than M. ![Machine environment configuration (`menvcfg`) register.](_images/svg-25bdc5e38b0e0804c5b53a2fa176913bbef4f635.svg) Figure 24\. Machine environment configuration (`menvcfg`) register. If bit FIOM (Fence of I/O implies Memory) is set to one in `menvcfg`, FENCE instructions executed in modes less privileged than M are modified so the requirement to order accesses to device I/O implies also the requirement to order main memory accesses. [Table 8](#norm:menvcfg%5Ffiom%5Ffence%5Fpresuc%5Fop)details the modified interpretation of FENCE instruction bits PI, PO, SI, and SO for modes less privileged than M when FIOM=1. Similarly, for modes less privileged than M when FIOM=1, if an atomic instruction that accesses a region ordered as device I/O has its _aq_and/or _rl_ bit set, then that instruction is ordered as though it accesses both device I/O and memory. If S-mode is not supported, or if `satp`.MODE is read-only zero (always Bare), the implementation may make FIOM read-only zero. __Table 8\. Modified interpretation of FENCE predecessor and successor sets for modes less privileged than M when FIOM=1.__ | Instruction bit | Meaning when set | | --------------- | -------------------------------------------------------------------------------------------------------------- | | PIPO | Predecessor device input and memory reads (PR implied)Predecessor device output and memory writes (PW implied) | | SISO | Successor device input and memory reads (SR implied)Successor device output and memory writes (SW implied) | | | Bit FIOM is needed in menvcfg so M-mode can emulate the hypervisor extension of ["H" Extension for Hypervisor Support](hypervisor.html#hypervisor), which has an equivalent FIOM bit in the hypervisor CSR henvcfg. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The PBMTE bit controls whether the Svpbmt extension is available for use in S-mode and G-stage address translation (i.e., for page tables pointed to by `satp` or `hgatp`). When PBMTE=1, Svpbmt is available for S-mode and G-stage address translation. When PBMTE=0, the implementation behaves as though Svpbmt were not implemented. If Svpbmt is not implemented, PBMTE is read-only zero.Furthermore,for implementations with the hypervisor extension,`henvcfg`.PBMTE is read-only zero if `menvcfg`.PBMTE is zero. After changing `menvcfg`.PBMTE, executing an SFENCE.VMA instruction with_rs1_\=`x0` and _rs2_\=`x0` suffices to synchronize address-translation caches with respect to the altered interpretation of page-table entries' PBMT fields.See [Memory-Management Fences](hypervisor.html#hyp-mm-fences) for additional synchronization requirements when the hypervisor extension is implemented. If the Svadu extension is implemented, the ADUE bit controls whether hardware updating of PTE A/D bits is enabled for S-mode and G-stage address translations. When ADUE=1, hardware updating of PTE A/D bits is enabled during S-mode address translation, and the implementation behaves as though the Svade extension were not implemented for S-mode address translation. When the hypervisor extension is implemented, if ADUE=1, hardware updating of PTE A/D bits is enabled during G-stage address translation, and the implementation behaves as though the Svade extension were not implemented for G-stage address translation. When ADUE=0, the implementation behaves as though Svade were implemented for S-mode and G-stage address translation. If Svadu is not implemented, ADUE is read-only zero.Furthermore, for implementations with the hypervisor extension,`henvcfg`.ADUE is read-only zero if `menvcfg`.ADUE is zero. After changing `menvcfg`.ADUE, executing an SFENCE.VMA instruction with_rs1_\=`x0` and _rs2_\=`x0` suffices to synchronize address-translation caches with respect to the altered interpretation of page-table entries' A/D bits.See [Memory-Management Fences](hypervisor.html#hyp-mm-fences) for additional synchronization requirements when the hypervisor extension is implemented. | | The Svade extension requires page-fault exceptions be raised when PTE A/D bits need be set, hence Svade is implemented when ADUE=0. | | -------------------------------------------------------------------------------------------------------------------------------------- | If the Smcdeleg extension is implemented, the CDE (Counter Delegation Enable) bit controls whether Zicntr and Zihpm counters can be delegated to S-mode. When CDE=1, the Smcdeleg extension is enabled, see ["Smcdeleg/Ssccfg" Counter Delegation Extensions](smcdeleg.html#smcdeleg). When CDE=0, the Smcdeleg and Ssccfg extensions appear to be not implemented. If Smcdeleg is not implemented, CDE is read-only zero. The Sstc extension adds the `STCE` (STimecmp Enable) bit to `menvcfg` CSR. When the Sstc extension is not implemented, `STCE` is read-only zero. The `STCE` bit enables `stimecmp` for S-mode when set to one. When this extension is implemented and `STCE` in `menvcfg` is zero, an attempt to access `stimecmp`in a mode other than M-mode raises an illegal-instruction exception, `STCE` in`henvcfg` is read-only zero, and `STIP` in `mip` and `sip` reverts to its defined behavior as if this extension is not implemented. Further, if the H extension is implemented, then `hip`.VSTIP also reverts its defined behavior as if this extension is not implemented. The Zicboz extension adds the `CBZE` (Cache Block Zero instruction enable) field to `menvcfg`. When the `CBZE` field is set to 1, it enables execution of the cache block zero instruction, `CBO.ZERO`, in modes less privileged than M. Otherwise, the instruction raises an illegal-instruction exception in modes less privileged than M. When the Zicboz extension is not implemented, `CBZE` is read-only zero. The Zicbom extension adds the `CBCFE` (Cache Block Clean and Flush instruction Enable) field to `menvcfg`. When the `CBCFE` field is set to 1, it enables execution of the cache block clean instruction (`CBO.CLEAN`) and the cache block flush instruction (`CBO.FLUSH`) in modes less privileged than M. Otherwise, these instructions raise an illegal-instruction exception in modes less privileged than M. When the Zicbom extension is not implemented, `CBCFE` is read-only zero. The Zicbom extension adds the `CBIE` (Cache Block Invalidate instruction Enable) WARL field to `menvcfg` to control execution of the cache block invalidate instruction (`CBO.INVAL`) in modes less privileged than M. When `CBIE` is set to`00b`, the instruction raises an illegal-instruction exception in modes less privileged than M. When the Zicbom extension is not implemented, `CBIE` is read-only zero.The encoding `10b` is reserved.When `CBIE` is set to `01b`or `11b`, and when enabled for execution in modes less privileged than M, it behaves as follows: * `01b` — The instruction is executed and performs a flush operation, even if configured by a mode less privileged than M to perform an invalidate operation. * `11b` — The instruction is executed and performs an invalidate operation, unless configured by a mode less privileged than M to perform a flush operation. If the Smnpm extension is implemented, the `PMM` field enables or disables pointer masking (see [Pointer Masking Extensions](zpm.html)) for the next-lower privilege mode (S-/HS-mode if S-mode is implemented, or U-mode otherwise), according to the values in[Table 9](#norm:menvcfg%5Fpmm%5Fenc). If Smnpm is not implemented, `PMM` is read-only zero. The `PMM` field is read-only zero for RV32. __Table 9\. Legal values of PMM WARL field__ | Value | Description | | ----- | ---------------------------------------------------------------------- | | 00 | Pointer masking is disabled (PMLEN = 0) | | 01 | Reserved | | 10 | Pointer masking is enabled with PMLEN = XLEN - 57 (PMLEN = 7 on RV64) | | 11 | Pointer masking is enabled with PMLEN = XLEN - 48 (PMLEN = 16 on RV64) | The Zicfilp extension adds the `LPE` field in `menvcfg`. When the `LPE` field is set to 1 and S-mode is implemented, the Zicfilp extension is enabled in S-mode. If `LPE` field is set to 1 and S-mode is not implemented, the Zicfilp extension is enabled in U-mode. When the `LPE` field is 0, the Zicfilp extension is not enabled in S-mode, and the following rules apply to S-mode. If the `LPE` field is 0 and S-mode is not implemented, then the same rules apply to U-mode. * The hart does not update the `ELP` state; it remains as `NO_LP_EXPECTED`. * The `LPAD` instruction operates as a no-op. The Zicfiss extension adds the `SSE` field to `menvcfg`. When the `SSE` field is set to 1 the Zicfiss extension is activated in S-mode. When `SSE` field is 0, the following rules apply to privilege modes that are less than M: * 32-bit Zicfiss instructions will revert to their behavior as defined by Zimop. * 16-bit Zicfiss instructions will revert to their behavior as defined by Zcmop. * The `pte.xwr=010b` encoding in VS/S-stage page tables becomes reserved. * `SSAMOSWAP.W/D` raises an illegal-instruction exception. When `menvcfg.SSE` is 0, the `henvcfg.SSE` and `senvcfg.SSE` fields are read-only zero. The Ssdbltrp extension adds the double-trap-enable (`DTE`) field in `menvcfg`. When `menvcfg.DTE` is zero, the implementation behaves as though Ssdbltrp is not implemented. When Ssdbltrp is not implemented `sstatus.SDT`, `vsstatus.SDT`, and`henvcfg.DTE` bits are read-only zero. When XLEN=32, `menvcfgh` is a 32-bit read/write register that aliases bits 63:32 of `menvcfg`. The `menvcfgh` register does not exist when XLEN=64. If U-mode is not supported, then registers `menvcfg` and `menvcfgh` do not exist. #### [](#sec:mseccfg)3.1.1.19\. Machine Security Configuration (`mseccfg`) Register `mseccfg` is a 64-bit read/write register, formatted as shown in [Figure 25](#norm:mseccfg%5Fenc%5Fimg), that controls security features. It exists if any extension that adds a field to `mseccfg` is implemented. Otherwise, it is reserved. ![Machine security configuration (`mseccfg`) register.](_images/svg-25df12cdde1a753b4a3ab2f85f370cc8a277e286.svg) Figure 25\. Machine security configuration (`mseccfg`) register. The Zkr extension adds the `SSEED` and `USEED` fields to the `mseccfg` CSR to control access to the `seed` CSR from modes less privileged than M. When `USEED` is 0, access to the `seed` CSR in U-mode raises an illegal-instruction exception. When `USEED` is 1, read-write access to the`seed` CSR from U-mode is allowed; all other types of accesses raise an illegal-instruction exception. If Zkr or U-mode is not implemented, `USEED` is read-only zero. When `SSEED` is 0, access to the `seed` CSR from S-/HS-mode raises an illegal-instruction exception. When `SSEED` is 1, read-write access to the`seed` CSR from S-/HS-mode is allowed; all other types of accesses raise an illegal-instruction exception. If Zkr or S-mode is not implemented, `SSEED` is read-only zero. When the H extension is also implemented, access to the `seed` CSR from an HS-qualified instruction leads to a virtual-instruction exception in VS and VU modes; all other types of accesses raise an illegal-instruction exception. __Table 10\. Entropy Source Access Control.__ | Mode | SSEED | USEED | Description | | ----- | ----- | ----- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | M | \- | \- | The seed CSR is always available in machine mode as normal (with a CSR read-write instruction.) Attempted read without a write raises an illegal-instruction exception regardless of mode and access control bits. | | U | \- | 0 | Any seed CSR access raises an illegal-instruction exception. | | U | \- | 1 | The seed CSR is accessible as normal. No exception is raised for read-write. | | S/HS | 0 | \- | Any seed CSR access raises an illegal-instruction exception. | | S/HS | 1 | \- | The seed CSR is accessible as normal. No exception is raised for read-write. | | VS/VU | 0 | \- | Any seed CSR access raises an illegal-instruction exception. | | VS/VU | 1 | \- | A read-write seed access raises a virtual-instruction exception, while other access conditions raise an illegal-instruction exception. | The Smepmp extension adds the `RLB`, `MMWP`, and the `MML` fields in`mseccfg`. When `mseccfg.RLB` (Rule Locking Bypass) a WARL field that provides a mechanism to temporarily modify **Locked** PMP rules. When `mseccfg.RLB` is 1, locked PMP rules may be removed or modified and locked PMP rules may be edited. When`mseccfg.RLB` is 0 and `pmpcfg.L` is 1 in any rule or entry (including disabled entries), then `mseccfg.RLB` remains 0 and any further modifications to`mseccfg.RLB` are ignored until a **PMP reset**. | | This feature is intended to be used as a debug mechanism, or as a temporary workaround during the boot process for simplifying software, and optimizing the allocation of memory and PMP rules. Using this functionality under normal operation, after the boot process is completed, should be avoided since it weakens the protection of _M-mode-only_ rules. Vendors who don’t need this functionality may hardwire this field to 0. The terminology used to specify the fields introduced by the Smepmp extension is listed in ["Smepmp" Extension for PMP Enhancements for memory access and execution prevention in Machine mode](smepmp.html). | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `mseccfg.MMWP` (Machine-Mode Allowlist Policy) is a WARL field. This field changes the default PMP policy for Machine mode when accessing memory regions that don’t have a matching PMP rule. This is a sticky bit, meaning that once set it cannot be unset until a **PMP reset**. When set it changes the default PMP policy for M-mode when accessing memory regions that don’t have a matching**PMP rule**, to **denied** instead of **ignored**. The `mseccfg.MML` (Machine Mode Lockdown) is a WARL field.The `MML` bit changes the interpretation of the `pmpcfg.L` bit defined in [3.1.7.1.2\. Locking and Privilege Mode](#pmp-locking).This is a sticky bit, meaning that once set it cannot be unset until a **PMP reset**. When `mseccfg.MML` is set the system’s behavior changes in the following way: 1. The meaning of `pmpcfg.L` changes: Instead of marking a rule as **locked** and**enforced** in all modes, it now marks a rule as **M-mode-only** when set and**S/U-mode-only** when unset. The formerly reserved encoding of `pmpcfg.RW=01`, and the encoding `pmpcfg.LRWX=1111`, now encode a **Shared-Region**. An _M-mode-only_ rule is **enforced** on Machine mode and **denied** in Supervisor or User mode. It also remains **locked** so that any further modifications to its associated configuration or address registers are ignored until a **PMP reset**, unless `mseccfg.RLB` is set. An _S/U-mode-only_ rule is **enforced** on Supervisor and User modes and**denied** on Machine mode. A _Shared-Region_ rule is **enforced** on all modes, with restrictions depending on the `pmpcfg.L` and `pmpcfg.X` bits: * A _Shared-Region_ rule where `pmpcfg.L` is not set can be used for sharing data between M-mode and S/U-mode, so is not executable. M-mode has read/write access to that region, and S/U-mode has read access if `pmpcfg.X`is not set, or read/write access if `pmpcfg.X` is set. * A _Shared-Region_ rule where `pmpcfg.L` is set can be used for sharing code between M-mode and S/U-mode, so is not writable. Both M-mode and S/U-mode have execute access on the region, and M-mode also has read access if`pmpcfg.X` is set. The rule remains **locked** so that any further modifications to its associated configuration or address registers are ignored until a **PMP reset**, unless `mseccfg.RLB` is set. * The encoding `pmpcfg.LRWX=1111` can be used for sharing data between M-mode and S/U mode, where both modes only have read-only access to the region. The rule remains **locked** so that any further modifications to its associated configuration or address registers are ignored until a **PMP reset**, unless`mseccfg.RLB` is set. 2. Adding a rule with executable privileges that either is **M-mode-only** or a**locked** **Shared-Region** is not possible and such `pmpcfg` writes are ignored, leaving `pmpcfg` unchanged. This restriction can be temporarily lifted by setting `mseccfg.RLB` e.g. during the boot process. 3. Executing code with Machine mode privileges is only possible from memory regions with a matching **M-mode-only** rule or a **locked** **Shared-Region** rule with executable privileges. Executing code from a region without a matching rule or with a matching _S/U-mode-only_ rule is **denied**. 4. If `mseccfg.MML` is not set, the combination of `pmpcfg.RW=01` remains reserved for future standard use. If the Smmpm extension is implemented, the `PMM` field enables or disables pointer masking (see [ZPM](zpm.html)) for M-mode according to the values in [Table 11](#norm:mseccfg%5Fpmm%5Fenc). If Smmpm is not implemented, `PMM` is read-only zero. The `PMM` field is read-only zero for RV32. __Table 11\. Legal values of PMM WARL field__ | Value | Description | | ----- | ---------------------------------------------------------------------- | | 00 | Pointer masking is disabled (PMLEN = 0) | | 01 | Reserved | | 10 | Pointer masking is enabled with PMLEN = XLEN - 57 (PMLEN = 7 on RV64) | | 11 | Pointer masking is enabled with PMLEN = XLEN - 48 (PMLEN = 16 on RV64) | | | Smmpm implementations need to satisfy max(largest supported virtual address size, largest supported supervisor physical address size) ⇐ (XLEN - PMLEN) bits to avoid any masking logic on the TLB access path. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The Zicfilp extension adds the `MLPE` field in `mseccfg`. When `MLPE` field is 1, Zicfilp extension is enabled in M-mode. When the `MLPE` field is 0, the Zicfilp extension is not enabled in M-mode and the following rules apply to M-mode. * The hart does not update the `ELP` state; it remains as `NO_LP_EXPECTED`. * The `LPAD` instruction operates as a no-op. When XLEN=32 only, `mseccfgh` is a 32-bit read/write register that aliases bits 63:32 of `mseccfg`. Register `mseccfgh` exists when XLEN=32 and `mseccfg` is implemented; it does not exist when XLEN=64. ### [](#3-1-2-machine-level-memory-mapped-registers)3.1.2\. Machine-Level Memory-Mapped Registers #### [](#3-1-2-1-machine-timer-mtime-and-mtimecmp-registers)3.1.2.1\. Machine Timer (`mtime` and `mtimecmp`) Registers Platforms provide a real-time counter, exposed as a memory-mapped machine-mode read-write register, `mtime`. `mtime` must increment at constant frequency, andthe platform must provide a mechanism for determining the period of an `mtime` tick. The `mtime` register will wrap around if the count overflows. The `mtime` register has a 64-bit precision on all RV32 and RV64 systems. Platforms provide a 64-bit memory-mapped machine-mode timer compare register (`mtimecmp`). A machine timer interrupt becomes pending whenever `mtime` contains a value greater than or equal to `mtimecmp`, treating the values as unsigned integers. The interrupt remains posted until `mtimecmp` becomes greater than `mtime`(typically as a result of writing `mtimecmp`).The interrupt will only be taken if interrupts are enabled and the MTIE bit is set in the `mie` register. ![Machine time register (memory-mapped control register).](_images/diag-850b45304d1878447c2c983f3f29bec99a983c6a.svg) Figure 26\. Machine time register (memory-mapped control register). ![Machine time compare register (memory-mapped control register).](_images/diag-ed9eddb704a73478f37cbb41121d057703dd8b0d.svg) Figure 27\. Machine time compare register (memory-mapped control register). | | The timer facility is defined to use wall-clock time rather than a cycle counter to support modern processors that run with a highly variable clock frequency to save energy through dynamic voltage and frequency scaling. Accurate real-time clocks (RTCs) are relatively expensive to provide (requiring a crystal or MEMS oscillator) and have to run even when the rest of system is powered down, and so there is usually only one in a system located in a different frequency/voltage domain from the processors. Hence, the RTC must be shared by all the harts in a system and accesses to the RTC will potentially incur the penalty of a voltage-level-shifter and clock-domain crossing. It is thus more natural to expose mtime as a memory-mapped register than as a CSR. Lower privilege levels do not have their own timecmp registers. Instead, machine-mode software can implement any number of virtual timers on a hart by multiplexing the next timer interrupt into themtimecmp register. Simple fixed-frequency systems can use a single clock for both cycle counting and wall-clock time. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If the result of the comparison between `mtime` and `mtimecmp`changes, it is guaranteed to be reflected in MTIP eventually, but not necessarily immediately. | | A spurious timer interrupt might occur if an interrupt handler increments mtimecmp then immediately returns, because MTIP might not yet have fallen in the interim. All software should be written to assume this event is possible, but most software should assume this event is extremely unlikely. It is almost always more performant to incur an occasional spurious timer interrupt than to poll MTIP until it falls. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | In RV32, memory-mapped writes to `mtimecmp` modify only one 32-bit part of the register.The following code sequence sets a 64-bit `mtimecmp` value without spuriously generating a timer interrupt due to the intermediate value of the comparand: For RV64, naturally aligned 64-bit memory accesses to the `mtime` and`mtimecmp` registers are additionally supported and are atomic. Sample code for setting the 64-bit time comparand in RV32 assuming a little-endian memory system and that the registers live in a strongly ordered I/O region. Storing -1 to the low-order bits of `mtimecmp` prevents `mtimecmp` from temporarily becoming smaller than the lesser of the old and new values. # New comparand is in a1:a0. li t0, -1 la t1, mtimecmp sw t0, 0(t1) # No smaller than old value. sw a1, 4(t1) # No smaller than new value. sw a0, 0(t1) # New value. The `time` CSR is a read-only shadow of the memory-mapped `mtime` register. When XLEN=32, the `timeh` CSR is a read-only shadow of the upper 32 bits of the memory-mapped `mtime` register, while `time` shadows only the lower 32 bits of`mtime`.When `mtime` changes, it is guaranteed to be reflected in `time`and `timeh` eventually, but not necessarily immediately. ### [](#3-1-3-machine-mode-privileged-instructions)3.1.3\. Machine-Mode Privileged Instructions #### [](#3-1-3-1-environment-call-and-breakpoint)3.1.3.1\. Environment Call and Breakpoint ![svg](_images/svg-b85f17219e404c3a56e44aa02c3ecc9a8b03847c.svg) The ECALL instruction is used to make a request to the supporting execution environment. When executed in U-mode, S-mode, or M-mode, it generates an environment-call-from-U-mode exception, environment-call-from-S-mode exception, or environment-call-from-M-mode exception, respectively, and performs no other operation. | | ECALL generates a different exception for each originating privilege mode so that environment call exceptions can be selectively delegated. A typical use case for Unix-like operating systems is to delegate to S-mode the environment-call-from-U-mode exception but not the others. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The EBREAK instruction is used by debuggers to cause control to be transferred back to a debugging environment. Unless overridden by an external debug environment, EBREAK raises a breakpoint exception and performs no other operation. | | As described in the "C" Standard Extension for Compressed Instructions in Volume I of this manual, the C.EBREAK instruction performs the same operation as the EBREAK instruction. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ECALL and EBREAK cause the receiving privilege mode’s `epc` register to be set to the address of the ECALL or EBREAK instruction itself, _not_the address of the following instruction. As ECALL and EBREAK cause synchronous exceptions, they are not considered to retire, and should not increment the `minstret` CSR. #### [](#otherpriv)3.1.3.2\. Trap-Return Instructions Instructions to return from trap are encoded under the PRIV minor opcode. ![svg](_images/svg-ac662d7542e4caefcdd4ac0d98bdff55e9262bfa.svg) To return after handling a trap, there are separate trap return instructions per privilege level, MRET and SRET.MRET is always provided. SRET must be provided if supervisor mode is supported, and should raise an illegal-instruction exception otherwise. SRET should also raise an illegal-instruction exception when TSR=1 in `mstatus`, as described in [3.1.1.6.6\. Virtualization Support in mstatus Register](#virt-control). An _x_RET instruction can be executed in privilege mode _x_ or higher, where executing a lower-privilege _x_RET instruction will pop the relevant lower-privilege interrupt enable and privilege mode stack. Attempting to execute an _x_RET instruction in a mode less privileged than _x_ will raise an illegal-instruction exception. In addition to manipulating the privilege stack as described in [3.1.1.6.1\. Privilege and Global Interrupt-Enable Stack in mstatus register](#privstack),_x_RET sets the `pc` to the value stored in the `_x_epc` register. If the A extension is supported, the _x_RET instruction is allowed to clear any outstanding LR address reservation but is not required to.Trap handlers should explicitly clear the reservation if required (e.g., by using a dummy SC) before executing the _x_RET. | | If _x_RET instructions always cleared LR reservations, it would be impossible to single-step through LR/SC sequences using a debugger. | | ----------------------------------------------------------------------------------------------------------------------------------------- | #### [](#wfi)3.1.3.3\. Wait for Interrupt The Wait for Interrupt instruction (WFI) informs the implementation that the current hart can be stalled until an interrupt might need servicing. Execution of the WFI instruction can also be used to inform the hardware platform that suitable interrupts should preferentially be routed to this hart. WFI is available in all privileged modes, andoptionally available to U-mode. This instruction may raise an illegal-instruction exception when TW=1 in `mstatus`, as described in [3.1.1.6.6\. Virtualization Support in mstatus Register](#virt-control). ![svg](_images/svg-a54c2f49d09ef2fbb6004a88508793daf4d5d405.svg) If an enabled interrupt is present or later becomes present while the hart is stalled, the interrupt trap will be taken on the following instruction, i.e., execution resumes in the trap handler and `mepc` \= `pc` \+ 4. | | The following instruction takes the interrupt trap so that a simple return from the trap handler will execute code after the WFI instruction. | | ------------------------------------------------------------------------------------------------------------------------------------------------ | Implementations are permitted to resume execution for any reason, even if an enabled interrupt has not become pending. Hence, a legal implementation is to simply implement the WFI instruction as a NOP. | | If the implementation does not stall the hart on execution of the instruction, then the interrupt will be taken on some instruction in the idle loop containing the WFI, and on a simple return from the handler, the idle loop will resume execution. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The WFI instruction can also be executed when interrupts are disabled. The operation of WFI must be unaffected by the global interrupt bits in`mstatus` (MIE and SIE) and the delegation register `mideleg` (i.e., the hart must resume if a locally enabled interrupt becomes pending, even if it has been delegated to a less-privileged mode), but should honor the individual interrupt enables (e.g, MTIE) (i.e., implementations should avoid resuming the hart if the interrupt is pending but not individually enabled). WFI is also required to resume execution for locally enabled interrupts pending at any privilege level, regardless of the global interrupt enable at each privilege level. If the event that causes the hart to resume execution does not cause an interrupt to be taken, execution will resume at `pc` \+ 4, and software must determine what action to take, including looping back to repeat the WFI if there was no actionable event. | | By allowing wake-up when interrupts are disabled, an alternate entry point to an interrupt handler can be called that does not require saving the current context, as the current context can be saved or discarded before the WFI is executed. As implementations are free to implement WFI as a NOP, software must explicitly check for any relevant pending but disabled interrupts in the code following an WFI, and should loop back to the WFI if no suitable interrupt was detected. The mip or sip registers can be interrogated to determine the presence of any interrupt in machine or supervisor mode respectively. The operation of WFI is unaffected by the delegation register settings. WFI is defined so that an implementation can trap into a higher privilege mode, either immediately on encountering the WFI or after some interval to initiate a machine-mode transition to a lower power state, for example. The same "wait-for-event" template might be used for possible future extensions that wait on memory locations changing, or message arrival. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-3-4-custom-system-instructions)3.1.3.4\. Custom SYSTEM Instructions The subspace of the SYSTEM major opcode shown in [Figure 28](#customsys) is designated for custom use. It is recommended that these instructions use bits 29:28 to designate the minimum required privilege mode, as do other SYSTEM instructions. ![SYSTEM instruction encodings designated for custom use.](_images/diag-9a22e09f44a9318b24cb7215438cad5abea31086.svg) Figure 28\. SYSTEM instruction encodings designated for custom use. ### [](#reset)3.1.4\. Reset Upon reset, a hart’s privilege mode is set to M. The `mstatus` fields MIE and MPRV are reset to 0. If little-endian memory accesses are supported, the `mstatus`/`mstatush` field MBE is reset to 0. The `misa` register is reset to enable the maximal set of supported extensions, as described in [3.1.1.1\. Machine ISA (misa) Register](#misa). For implementations with the "A" standard extension, there is no valid load reservation. The `pc` is set to an implementation-defined reset vector. The `mcause` register is set to a value indicating the cause of the reset. Writable PMP registers’ A and L fields are set to 0, unless the platform mandates a different reset value for some PMP registers’ A and L fields. If the hypervisor extension is implemented, the`hgatp`.MODE and `vsatp`.MODE fields are reset to 0. If the Smrnmi extension is implemented, the `mnstatus`.NMIE field is reset to 0. No **WARL** field contains an illegal value. If the Zicfilp extension is implemented, the `mseccfg`.MLPE field is reset to 0.All other hart state is UNSPECIFIED. The `MML`, `MMWP`, and `RLB` fields of the `mseccfg` register are set to 0, unless the platform mandates a different reset value. The `mcause` values after reset have implementation-specific interpretation, but the value 0 should be returned on implementations that do not distinguish different reset conditions. Implementations that distinguish different reset conditions should only use 0 to indicate the most complete reset. The `USEED` and `SSEED` fields of the `mseccfg` CSR must have defined reset values. The system must not allow them to be in an undefined state after reset. | | Some designs may have multiple causes of reset (e.g., power-on reset, external hard reset, brownout detected, watchdog timer elapse, sleep-mode wake-up), which machine-mode software and debuggers may wish to distinguish. To avoid ambiguity, mcause reset values may alias mcause values following synchronous exceptions. There should be no ambiguity in this overlap, since on reset the pc is typically set to a different value than on other traps. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#nmi)3.1.5\. Non-Maskable Interrupts Non-maskable interrupts (NMIs) are only used for hardware error conditions, and cause an immediate jump to an implementation-defined NMI vector running in M-mode regardless of the state of a hart’s interrupt enable bits. The `mepc` register is written with the virtual address of the instruction that was interrupted, and `mcause` is set to a value indicating the source of the NMI. The NMI can thus overwrite state in an active machine-mode interrupt handler. The values written to `mcause` on an NMI are implementation-defined. The high Interrupt bit of `mcause` should be set to indicate that this was an interrupt. An Exception Code of 0 is reserved to mean "unknown cause" and implementations that do not distinguish sources of NMIs via the `mcause` register should return 0 in the Exception Code. Unlike resets, NMIs do not reset processor state, enabling diagnosis, reporting, and possible containment of the hardware error. ### [](#pma)3.1.6\. Physical Memory Attributes The physical memory map for a complete system includes various address ranges, some corresponding to memory regions and some to memory-mapped control registers, portions of which might not be accessible. Some memory regions might not support reads, writes, or execution; some might not support subword or subblock accesses; some might not support atomic operations; and some might not support cache coherence or might have different memory models. Similarly, memory-mapped control registers vary in their supported access widths, support for atomic operations, and whether read and write accesses have associated side effects. In RISC-V systems, these properties and capabilities of each region of the machine’s physical address space are termed _physical memory attributes_(PMAs). This section describes RISC-V PMA terminology and how RISC-V systems implement and check PMAs. PMAs are inherent properties of the underlying hardware and rarely change during system operation. Unlike physical memory protection values described in [3.1.7\. Physical Memory Protection](#pmp), PMAs do not vary by execution context. The PMAs of some memory regions are fixed at chip design time—for example, for an on-chip ROM. Others are fixed at board design time, depending, for example, on which other chips are connected to off-chip buses. Off-chip buses might also support devices that could be changed on every power cycle (cold pluggable) or dynamically while the system is running (hot pluggable). Some devices might be configurable at run time to support different uses that imply different PMAs—for example, an on-chip scratchpad RAM might be cached privately by one core in one end-application, or accessed as a shared non-cached memory in another end-application. Most systems will require that at least some PMAs are dynamically checked in hardware later in the execution pipeline after the physical address is known, as some operations will not be supported at all physical memory addresses, and some operations require knowing the current setting of a configurable PMA attribute. While many other architectures specify some PMAs in the virtual memory page tables and use the TLB to inform the pipeline of these properties, this approach injects platform-specific information into a virtualized layer and can cause system errors unless attributes are correctly initialized in each page-table entry for each physical memory region. In addition, the available page sizes might not be optimal for specifying attributes in the physical memory space, leading to address-space fragmentation and inefficient use of expensive TLB entries. For RISC-V, we separate out specification and checking of PMAs into a separate hardware structure, the _PMA checker_. In many cases, the attributes are known at system design time for each physical address region, and can be hardwired into the PMA checker.Where the attributes are run-time configurable, platform-specific memory-mapped control registers can be provided to specify these attributes at a granularity appropriate to each region on the platform (e.g., for an on-chip SRAM that can be flexibly divided between cacheable and uncacheable uses).PMAs are checked for any access to physical memory, including accesses that have undergone virtual to physical memory translation. To aid in system debugging, we strongly recommend that, where possible, RISC-V processors precisely trap physical memory accesses that fail PMA checks. Precisely trapped PMA violations manifest as instruction, load, or store access-fault exceptions, distinct from virtual-memory page-fault exceptions. Precise PMA traps might not always be possible, for example, when probing a legacy bus architecture that uses access failures as part of the discovery mechanism. In this case, error responses from peripheral devices will be reported as imprecise bus-error interrupts. PMAs must also be readable by software to correctly access certain devices or to correctly configure other hardware components that access memory, such as DMA engines. As PMAs are tightly tied to a given physical platform’s organization, many details are inherently platform-specific, as is the means by which software can learn the PMA values for a platform. Some devices, particularly legacy buses, do not support discovery of PMAs and so will give error responses or time out if an unsupported access is attempted. Typically, platform-specific machine-mode code will extract PMAs and ultimately present this information to higher-level less-privileged software using some standard representation. Where platforms support dynamic reconfiguration of PMAs, an interface will be provided to set the attributes by passing requests to a machine-mode driver that can correctly reconfigure the platform. For example, switching cacheability attributes on some memory regions might involve platform-specific operations, such as cache flushes, that are available only to machine-mode. #### [](#3-1-6-1-main-memory-versus-io-regions)3.1.6.1\. Main Memory versus I/O Regions The most important characterization of a given memory address range is whether it holds regular main memory or I/O devices. Regular main memory is required to have a number of properties, specified below, whereas I/O devices can have a much broader range of attributes. Memory regions that do not fit into regular main memory, for example, device scratchpad RAMs, are categorized as I/O regions. | | What previous versions of this specification termed _vacant_ regions are no longer a distinct category; they are now described as I/O regions that are not accessible (i.e. lacking read, write, and execute permissions). Main memory regions that are not accessible are also allowed. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-6-2-supported-access-type-pmas)3.1.6.2\. Supported Access Type PMAs Access types specify which access widths, from 8-bit byte to long multi-word burst, are supported, and also whether misaligned accesses are supported for each access width. | | Although software running on a RISC-V hart cannot directly generate bursts to memory, software might have to program DMA engines to access I/O devices and might therefore need to know which access sizes are supported. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Main memory regions always support read and write of all access widths required by the attached devices, and can specify whether instruction fetch is supported. | | Some platforms might mandate that all of main memory support instruction fetch. Other platforms might prohibit instruction fetch from some main memory regions. In some cases, the design of a processor or device accessing main memory might support other widths, but must be able to function with the types supported by the main memory. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | I/O regions can specify which combinations of read, write, or execute accesses to which data widths are supported. For systems with page-based virtual memory, I/O and memory regions can specify which combinations of hardware page-table reads and hardware page-table writes are supported. | | Unix-like operating systems generally require that all of cacheable main memory supports page-table walks. | | ------------------------------------------------------------------------------------------------------------- | #### [](#3-1-6-3-atomicity-pmas)3.1.6.3\. Atomicity PMAs Atomicity PMAs describes which atomic instructions are supported in this address region. Support for atomic instructions is divided into two categories: _LR/SC_ and _AMOs_. | | Some platforms might mandate that all of cacheable main memory support all atomic operations required by the attached processors. | | ------------------------------------------------------------------------------------------------------------------------------------ | ##### [](#3-1-6-3-1-amo-pma)3.1.6.3.1\. AMO PMA Within AMOs, there are four levels of support: _AMONone_, _AMOSwap_,_AMOLogical_, and _AMOArithmetic_. AMONone indicates that no AMO operations are supported. AMOSwap indicates that only `amoswap`instructions are supported in this address range. AMOLogical indicates that swap instructions plus all the logical AMOs (`amoand`, `amoor`,`amoxor`) are supported. AMOArithmetic indicates that all RISC-V AMOs defined by the A extension are supported. For each level of support, naturally aligned AMOs of a given width are supported if the underlying memory region supports reads and writes of that width. Main memory and I/O regions may only support a subset or none of the processor-supported atomic operations. __Table 12\. Classes of AMOs supported by I/O regions.__ | AMO Class | Supported Operations | | ------------------------------------- | ------------------------------------------------------------------------------------------ | | AMONoneAMOSwapAMOLogicalAMOArithmetic | _None_ amoswapabove + amoand, amoor, amoxorabove + amoadd, amomin, amomax, amominu,amomaxu | | | We recommend providing at least AMOLogical support for I/O regions where possible. | | ------------------------------------------------------------------------------------- | The Zacas extension defines three additional levels of support: AMOCASW, AMOCASD, and AMOCASQ. AMOCASW indicates that in addition to instructions indicated by AMOArithmetic level support, the `AMOCAS.W` instruction is supported. AMOCASD indicates that in addition to instructions indicated by AMOCASW level support, the `AMOCAS.D`instruction is supported. AMOCASQ indicates that in addition to instructions indicated by AMOCASD level support, the `AMOCAS.Q` instruction is supported. | | AMOCASW/D/Q require AMOArithmetic level support as the AMOCAS.W/D/Qinstructions require ability to perform an arithmetic comparison and a swap operation. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | The AMOs specified by the Zabha extension require the same level of support as the corresponding instructions in the A standard extension or the Zacas extension. ##### [](#3-1-6-3-2-reservability-pma)3.1.6.3.2\. Reservability PMA For _LR/SC_, there are three levels of support indicating combinations of the reservability and eventuality properties: _RsrvNone_,_RsrvNonEventual_, and _RsrvEventual_. RsrvNone indicates that no LR/SC operations are supported (the location is non-reservable). RsrvNonEventual indicates that the operations are supported (the location is reservable), but without the eventual success guarantee described in the unprivileged ISA specification. RsrvEventual indicates that the operations are supported and provide the eventual success guarantee. | | We recommend providing RsrvEventual support for main memory regions where possible. Most I/O regions will not support LR/SC accesses, as these are most conveniently built on top of a cache-coherence scheme, but some may support RsrvNonEventual or RsrvEventual. When LR/SC is used for memory locations marked RsrvNonEventual, software should provide alternative fall-back mechanisms used when lack of progress is detected. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-6-4-misaligned-atomicity-granule-pma)3.1.6.4\. Misaligned Atomicity Granule PMA The misaligned atomicity granule PMA provides constrained support for misaligned AMOs. This PMA, if present, specifies the size of a _misaligned atomicity granule_, a naturally aligned power-of-two number of bytes. Specific supported values for this PMA are represented by MAG_NN_, e.g., MAG16 indicates the misaligned atomicity granule is at least 16 bytes. The misaligned atomicity granule PMA applies only to AMOs, loads and stores defined in the base ISAs, and loads and stores of no more than XLEN bits defined in the F, D, and Q extensions, and compressed encodings thereof.For an instruction in that set, if all accessed bytes lie within the same misaligned atomicity granule, the instruction will not raise an exception for reasons of address alignment, and the instruction will give rise to only one memory operation for the purposes of RVWMO—​i.e., it will execute atomically. If a misaligned AMO accesses a region that does not specify a misaligned atomicity granule PMA, or if not all accessed bytes lie within the same misaligned atomicity granule, then an exception is raised. For regular loads and stores that access such a region or for which not all accessed bytes lie within the same atomicity granule, then either an exception is raised, or the access proceeds but is not guaranteed to be atomic. Implementations may raise access-fault exceptions instead of address-misaligned exceptions for some misaligned accesses, indicating the instruction should not be emulated by a trap handler. | | LR/SC instructions are unaffected by this PMA and so always raise an exception when misaligned. Vector memory accesses are also unaffected, so might execute non-atomically even when contained within a misaligned atomicity granule. Implicit accesses are similarly unaffected by this PMA. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-6-5-memory-ordering-pmas)3.1.6.5\. Memory-Ordering PMAs Regions of the address space are classified as either _main memory_ or_I/O_ for the purposes of ordering by the FENCE instruction and atomic-instruction ordering bits. Accesses by one hart to main memory regions are observable not only by other harts but also by other devices with the capability to initiate requests in the main memory system (e.g., DMA engines). Coherent main memory regions always have either the RVWMO or RVTSO memory model. Incoherent main memory regions have an implementation-defined memory model. Accesses by one hart to an I/O region are observable not only by other harts and bus mastering devices but also by the targeted I/O devices, and I/O regions may be accessed with either _relaxed_ or _strong_ordering. Accesses to an I/O region with relaxed ordering are generally observed by other harts and bus mastering devices in a manner similar to the ordering of accesses to an RVWMO memory region, as discussed in the I/O Ordering section in the RVWMO Explanatory Material appendix of Volume I of this specification. By contrast, accesses to an I/O region with strong ordering are generally observed by other harts and bus mastering devices in program order. Each strongly ordered I/O region specifies a numbered ordering channel, which is a mechanism by which ordering guarantees can be provided between different I/O regions. Channel 0 is used to indicate point-to-point strong ordering only, where only accesses by the hart to the single associated I/O region are strongly ordered. Channel 1 is used to provide global strong ordering across all I/O regions. Any accesses by a hart to any I/O region associated with channel 1 can only be observed to have occurred in program order by all other harts and I/O devices, including relative to accesses made by that hart to relaxed I/O regions or strongly ordered I/O regions with different channel numbers. In other words, any access to a region in channel 1 is equivalent to executing a `fence io,io` instruction before and after the instruction. Other larger channel numbers provide program ordering to accesses by that hart across any regions with the same channel number Systems might support dynamic configuration of ordering properties on each memory region. | | Strong ordering can be used to improve compatibility with legacy device driver code, or to enable increased performance compared to insertion of explicit ordering instructions when the implementation is known to not reorder accesses. Local strong ordering (channel 0) is the default form of strong ordering as it is often straightforward to provide if there is only a single in-order communication path between the hart and the I/O device. Generally, different strongly ordered I/O regions can share the same ordering channel without additional ordering hardware if they share the same interconnect path and the path does not reorder requests. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#3-1-6-6-coherence-and-cacheability-pmas)3.1.6.6\. Coherence and Cacheability PMAs Coherence is a property defined for a single physical address, and indicates that writes to that address by one agent will eventually be made visible to other coherent agents in the system. Coherence is not to be confused with the memory consistency model of a system, which defines what values a memory read can return given the previous history of reads and writes to the entire memory system. In RISC-V platforms, the use of hardware-incoherent regions is discouraged due to software complexity, performance, and energy impacts. The cacheability of a memory region should not affect the software view of the region except for differences reflected in other PMAs, such as main memory versus I/O classification, memory ordering, supported accesses and atomic operations, and coherence. For this reason, we treat cacheability as a platform-level setting managed by machine-mode software only. Where a platform supports configurable cacheability settings for a memory region, a platform-specific machine-mode routine will change the settings and flush caches if necessary, so the system is only incoherent during the transition between cacheability settings. This transitory state should not be visible to lower privilege levels. | | Coherence is straightforward to provide for a shared memory region that is not cached by any agent. The PMA for such a region would simply indicate it should not be cached in a private or shared cache. Coherence is also straightforward for read-only regions, which can be safely cached by multiple agents without requiring a cache-coherence scheme. The PMA for this region would indicate that it can be cached, but that writes are not supported. Some read-write regions might only be accessed by a single agent, in which case they can be cached privately by that agent without requiring a coherence scheme. The PMA for such regions would indicate they can be cached. The data can also be cached in a shared cache, as other agents should not access the region. If an agent can cache a read-write region that is accessible by other agents, whether caching or non-caching, a cache-coherence scheme is required to avoid use of stale values. In regions lacking hardware cache coherence (hardware-incoherent regions), cache coherence can be implemented entirely in software, but software coherence schemes are notoriously difficult to implement correctly and often have severe performance impacts due to the need for conservative software-directed cache-flushing. Hardware cache-coherence schemes require more complex hardware and can impact performance due to the cache-coherence probes, but are otherwise invisible to software. For each hardware cache-coherent region, the PMA would indicate that the region is coherent and which hardware coherence controller to use if the system has multiple coherence controllers. For some systems, the coherence controller might be an outer-level shared cache, which might itself access further outer-level cache-coherence controllers hierarchically. Most memory regions within a platform will be coherent to software, because they will be fixed as either uncached, read-only, hardware cache-coherent, or only accessed by one agent. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | If a PMA indicates non-cacheability, then accesses to that region must be satisfied by the memory itself, not by any caches. | | For implementations with a cacheability-control mechanism, the situation may arise that a program uncacheably accesses a memory location that is currently cache-resident. In this situation, the cached copy must be ignored. This constraint is necessary to prevent more-privileged modes’ speculative cache refills from affecting the behavior of less-privileged modes’ uncacheable accesses. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#3-1-6-7-idempotency-pmas)3.1.6.7\. Idempotency PMAs Idempotency PMAs describe whether reads and writes to an address region are idempotent. Main memory regions are assumed to be idempotent.For I/O regions, idempotency on reads and writes can be specified separately (e.g., reads are idempotent but writes are not). If accesses are non-idempotent, i.e., there is potentially a side effect on any read or write access, then speculative or redundant accesses must be avoided. For the purposes of defining the idempotency PMAs, changes in observed memory ordering created by redundant accesses are not considered a side effect. | | While hardware should always be designed to avoid speculative or redundant accesses to memory regions marked as non-idempotent, it is also necessary to ensure software or compiler optimizations do not generate spurious accesses to non-idempotent memory regions. Non-idempotent regions might not support misaligned accesses. Misaligned accesses to such regions should raise access-fault exceptions rather than address-misaligned exceptions, indicating that software should not emulate the misaligned access using multiple smaller accesses, which could cause unexpected side effects. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For non-idempotent regions, implicit reads and writes must not be performed early or speculatively, with the following exceptions. When a non-speculative implicit read is performed, an implementation is permitted to additionally read any of the bytes within a naturally aligned power-of-2 region containing the address of the non-speculative implicit read. Furthermore, when a non-speculative instruction fetch is performed, an implementation is permitted to additionally read any of the bytes within the _next_ naturally aligned power-of-2 region of the same size (with the address of the region taken modulo 2XLEN).The results of these additional reads may be used to satisfy subsequent early or speculative implicit reads.The size of these naturally aligned power-of-2 regions is implementation-defined, but, for systems with page-based virtual memory, must not exceed the smallest supported page size. ### [](#pmp)3.1.7\. Physical Memory Protection To support secure processing and contain faults, it is desirable to limit the physical addresses accessible by software running on a hart. An optional physical memory protection (PMP) unit provides per-hart machine-mode control registers to allow physical memory access privileges (read, write, execute) to be specified for each physical memory region. The PMP values are checked in parallel with the PMA checks described in [3.1.6\. Physical Memory Attributes](#pma). The granularity of PMP access control settings are platform-specific, but the standard PMP encoding supports regions as small as four bytes. Certain regions’ privileges can be hardwired—for example, some regions might only ever be visible in machine mode but in no lower-privilege layers. | | Platforms vary widely in demands for physical memory protection, and some platforms may provide other PMP structures in addition to or instead of the scheme described in this section. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | PMP checks are applied to all accesses whose effective privilege mode is S or U, including instruction fetches and data accesses in S and U mode, and data accesses in M-mode when the MPRV bit in `mstatus` is set and the MPP field in `mstatus` contains S or U. PMP checks are also applied to page-table accesses for virtual-address translation, for which the effective privilege mode is S. Optionally, PMP checks may additionally apply to M-mode accesses, in which case the PMP registers themselves are locked, so that even M-mode software cannot change them until the hart is reset. In effect, PMP can _grant_ permissions to S and U modes, which by default have none, and can _revoke_ permissions from M-mode, which by default has full permissions. PMP violations are always trapped precisely at the processor. #### [](#3-1-7-1-physical-memory-protection-csrs)3.1.7.1\. Physical Memory Protection CSRs PMP entries are described by an 8-bit configuration register and one MXLEN-bit address register. Some PMP settings additionally use the address register associated with the preceding PMP entry. Up to 64 PMP entries are supported. Implementations may implement zero, 16, or 64 PMP entries; the lowest-numbered PMP entries must be implemented first. All PMP CSR fields are **WARL** and may be read-only zero. PMP CSRs are only accessible to M-mode. The PMP configuration registers are densely packed into CSRs to minimize context-switch time. For RV32, sixteen CSRs, `pmpcfg0`–`pmpcfg15`, hold the configurations `pmp0cfg`–`pmp63cfg` for the 64 PMP entries, as shown in [Figure 29](#pmpcfg-rv32). For RV64, eight even-numbered CSRs, `pmpcfg0`, `pmpcfg2`, …, `pmpcfg14`, hold the configurations for the 64 PMP entries, as shown in[Figure 30](#pmpcfg-rv64). For RV64, the odd-numbered configuration registers, `pmpcfg1`, `pmpcfg3`, …, `pmpcfg15`, are illegal. | | RV64 harts use pmpcfg2, rather than pmpcfg1, to hold configurations for PMP entries 8-15\. This design reduces the cost of supporting multiple MXLEN values, since the configurations for PMP entries 8-11 appear in pmpcfg2\[31:0\] for both RV32 and RV64. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ![RV32 PMP configuration CSR layout.](_images/diag-efd252c46c30d454fa34f34e22da6366e5c313c2.svg) Figure 29\. RV32 PMP configuration CSR layout. ![RV64 PMP configuration CSR layout.](_images/diag-b47bcbc11126ed099881aa16d91f134279156f98.svg) Figure 30\. RV64 PMP configuration CSR layout. The PMP address registers are CSRs named `pmpaddr0`\-`pmpaddr63`. Each PMP address register encodes bits 33-2 of a 34-bit physical address for RV32, as shown in [Figure 31](#pmpaddr-rv32). For RV64, each PMP address register encodes bits 55-2 of a 56-bit physical address, as shown in [Figure 32](#pmpaddr-rv64). Not all physical address bits may be implemented, and so the `pmpaddr` registers are **WARL**. | | The Sv32 page-based virtual-memory scheme described in[Sv32: Page-Based 32-bit Virtual-Memory Systems](supervisor.html#sv32) supports 34-bit physical addresses for RV32, so the PMP scheme must support addresses wider than XLEN for RV32\. The Sv39 and Sv48 page-based virtual-memory schemes described in[Sv32: Page-Based 39-bit Virtual-Memory Systems](supervisor.html#sv39) and [Sv48: Page-Based 48-bit Virtual-Memory Systems](supervisor.html#sv48) support a 56-bit physical address space, so the RV64 PMP address registers impose the same limit. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ![PMP address register format, RV32.](_images/diag-ec67b4b34dae76306144656520ff9981f801f1cb.svg) Figure 31\. PMP address register format, RV32. ![PMP address register format, RV64.](_images/diag-9855fd151b9d924dfc27737cdee72a35b7d4f7cf.svg) Figure 32\. PMP address register format, RV64. [Figure 33](#pmpcfg) shows the layout of a PMP configuration register. The R, W, and X bits, when set, indicate that the PMP entry permits read, write, and instruction execution, respectively. When one of these bits is clear, the corresponding access type is denied. The R, W, and X fields form a collective **WARL** field for which the combinations with R=0 and W=1 are reserved. The remaining two fields, A and L, are described in the following sections. ![PMP configuration register format.](_images/diag-87e49044797677d35392804a47489c59fb69ce26.svg) Figure 33\. PMP configuration register format. Attempting to fetch an instruction from a PMP region that does not have execute permissions raises an instruction access-fault exception. Attempting to execute a load, load-reserved, or cache-block management instruction which accesses a physical address within a PMP region without read permissions raises a load access-fault exception. Attempting to execute a store, store-conditional, AMO, or cache-block zero instruction which accesses a physical address within a PMP region without write permissions raises a store access-fault exception. ##### [](#3-1-7-1-1-address-matching)3.1.7.1.1\. Address Matching The A field in a PMP entry’s configuration register encodes the address-matching mode of the associated PMP address register. The encoding of this field is shown in [Table 13](#pmpcfg-a). When A=0, this PMP entry is disabled and matches no addresses. Two other address-matching modes are supported: naturally aligned power-of-2 regions (NAPOT), including the special case of naturally aligned four-byte regions (NA4); and the top boundary of an arbitrary range (TOR). These modes support four-byte granularity. __Table 13\. Encoding of A field in PMP configuration registers.__ | A | Name | Description | | ---- | -------------- | ------------------------------------------------------------------------------------------------------------------- | | 0123 | OFFTORNA4NAPOT | Null region (disabled)Top of rangeNaturally aligned four-byte regionNaturally aligned power-of-two region, ≥8 bytes | NAPOT ranges make use of the low-order bits of the associated address register to encode the size of the range, as shown in[Table 14](#pmpcfg-napot). __Table 14\. NAPOT range encoding in PMP address and configuration registers.__ | pmpaddr | pmpcfg.A | Match type and size | | ----------------------------------------------------------------------------------------- | ---------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | yyyy…​yyyy yyyy…​yyy0 yyyy…​yy01 yyyy…​y011…​ yy01…​1111 y011…​1111 0111…​1111 1111…​1111 | NA4NAPOTNAPOTNAPOT…​NAPOTNAPOTNAPOTNAPOT | 4-byte NAPOT range8-byte NAPOT range16-byte NAPOT range32-byte NAPOT range…2XLEN\-byte NAPOT range2XLEN+1\-byte NAPOT range2XLEN+2\-byte NAPOT range2XLEN+3\-byte NAPOT range | If TOR is selected, the associated address register forms the top of the address range, and the preceding PMP address register forms the bottom of the address range. If PMP entry _i_'s A field is set to TOR, the entry matches any address _y_ such that `pmpaddri-1`≤_y_<`pmpaddri` (irrespective of the value of `pmpcfgi-1`). If PMP entry 0’s A field is set to TOR, zero is used for the lower bound, and so it matches any address `y XLEN. * Defined the format of the memory-mapped `msip` registers. Finally, the following clarifications and document improvements have been made since the last document release: * Transliterated the document from LaTeX into AsciiDoc. * Included all ratified extensions through March 2024. * Clarified that "platform- or custom-use" interrupts are actually "platform-use interrupts", where the platform can choose to make some custom. * Clarified semantics of explicit accesses to CSRs wider than XLEN bits. * Clarified that MXLEN≥SXLEN. * Clarified that WFI is not a HINT instruction. * Clarified that VS-stage page-table accesses set G-stage A/D bits. * Clarified ordering rules when PBMT=IO is used on main-memory regions. * Clarified ordering rules for hardware A/D bit updates. * Clarified that, for a given exception cause, `_x_tval` might sometimes be set to a nonzero value but sometimes not. * Clarified exception behavior of unimplemented or inaccessible CSRs. * Clarified that Svpbmt allows implementations to override additional PMAs. * Replaced the concept of vacant memory regions with inaccessible memory or I/O regions. * Clarified that timer and count-overflow interrupts' arrival in interrupt-pending registers is not immediate. * Clarified that MXR affects only explicit memory accesses. **_Preface to Version 20211203_** This document describes the RISC-V privileged architecture. This release, version 20211203, contains the following versions of the RISC-V ISA modules: | Module | Version | Status | | ---------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------- | ----------------------------------------------------------------------------- | | **Machine ISA** **Supervisor ISA** **Svnapot Extension** **Svpbmt Extension** **Svinval Extension** **Hypervisor ISA** | **1.12** **1.12** **1.0** **1.0** **1.0** **1.0** | **Ratified** **Ratified** **Ratified** **Ratified** **Ratified** **Ratified** | The following changes have been made since version 1.11, which, while not strictly backwards compatible, are not anticipated to cause software portability problems in practice: * Changed MRET and SRET to clear `mstatus`.MPRV when leaving M-mode. * Reserved additional `satp` patterns for future use. * Stated that the `scause` Exception Code field must implement bits 4–0 at minimum. * Relaxed I/O regions have been specified to follow RVWMO. The previous specification implied that PPO rules other than fences and acquire/release annotations did not apply. * Constrained the LR/SC reservation set size and shape when using page-based virtual memory. * PMP changes require an SFENCE.VMA on any hart that implements page-based virtual memory, even if VM is not currently enabled. * Allowed for speculative updates of page table entry A bits. * Clarify that if the address-translation algorithm non-speculatively reaches a PTE in which a bit reserved for future standard use is set, a page-fault exception must be raised. Additionally, the following compatible changes have been made since version 1.11: * Removed the N extension. * Defined the mandatory RV32-only CSR `mstatush`, which contains most of the same fields as the upper 32 bits of RV64’s `mstatus`. * Defined the mandatory CSR `mconfigptr`, which if nonzero contains the address of a configuration data structure. * Defined `mseccfg` and `mseccfgh` CSRs, which control the machine’s security configuration. * Defined `menvcfg`, `henvcfg`, and `senvcfg` CSRs (and RV32-only`menvcfgh` and `henvcfgh` CSRs), which control various characteristics of the execution environment. * Designated part of SYSTEM major opcode for custom use. * Permitted the unconditional delegation of less-privileged interrupts. * Added optional big-endian and bi-endian support. * Made priority of load/store/AMO address-misaligned exceptions implementation-defined relative to load/store/AMO page-fault and access-fault exceptions. * PMP reset values are now platform-defined. * An additional 48 optional PMP registers have been defined. * Slightly relaxed the atomicity requirement for A and D bit updates performed by the implementation. * Clarify the architectural behavior of address-translation caches * Added Sv57 and Sv57x4 address translation modes. * Software breakpoint exceptions are permitted to write either 0 or the`pc` to `_x_tval`. * Clarified that bare S-mode need not support the SFENCE.VMA instruction. * Specified relaxed constraints for implicit reads of non-idempotent regions. * Added the Svnapot Standard Extension, along with the N bit in Sv39, Sv48, and Sv57 PTEs. * Added the Svpbmt Standard Extension, along with the PBMT bits in Sv39, Sv48, and Sv57 PTEs. * Added the Svinval Standard Extension and associated instructions. Finally, the hypervisor architecture proposal has been extensively revised. **_Preface to Version 1.11_** This is version 1.11 of the RISC-V privileged architecture. The document contains the following versions of the RISC-V ISA modules: | Module | Version | Status | | --------------------------------------------------- | ----------------------- | --------------------------------- | | **Machine ISA** **Supervisor ISA** _Hypervisor ISA_ | **1.11** **1.11** _0.3_ | **Ratified** **Ratified** _Draft_ | Changes from version 1.10 include: * Moved Machine and Supervisor spec to **Ratified** status. * Improvements to the description and commentary. * Added a draft proposal for a hypervisor extension. * Specified which interrupt sources are reserved for standard use. * Allocated some synchronous exception causes for custom use. * Specified the priority ordering of synchronous exceptions. * Added specification that xRET instructions may, but are not required to, clear LR reservations if A extension present. * The virtual-memory system no longer permits supervisor mode to execute instructions from user pages, regardless of the SUM setting. * Clarified that ASIDs are private to a hart, and added commentary about the possibility of a future global-ASID extension. * SFENCE.VMA semantics have been clarified. * Made the `mstatus`.MPP field **WARL**, rather than **WLRL**. * Made the unused `_x_ip` fields **WPRI**, rather than **WIRI**. * Made the unused `misa` fields **WARL**, rather than **WIRI**. * Made the unused `pmpaddr` and `pmpcfg` fields **WARL**, rather than **WIRI**. * Required all harts in a system to employ the same PTE-update scheme as each other. * Rectified an editing error that misdescribed the mechanism by which`mstatus._x_IE` is written upon an exception. * Described scheme for emulating misaligned AMOs. * Specified the behavior of the `misa` and `_x_epc` registers in systems with variable IALIGN. * Specified the behavior of writing self-contradictory values to the`misa` register. * Defined the `mcountinhibit` CSR, which stops performance counters from incrementing to reduce energy consumption. * Specified semantics for PMP regions coarser than four bytes. * Specified contents of CSRs across XLEN modification. * Moved PLIC chapter into its own document. **_Preface to Version 1.10_** This is version 1.10 of the RISC-V privileged architecture proposal. Changes from version 1.9.1 include: * The previous version of this document was released under a Creative Commons Attribution 4.0 International License by the original authors, and this and future versions of this document will be released under the same license. * The explicit convention on shadow CSR addresses has been removed to reclaim CSR space. Shadow CSRs can still be added as needed. * The `mvendorid` register now contains the JEDEC code of the core provider as opposed to a code supplied by the Foundation. This avoids redundancy and offloads work from the Foundation. * The interrupt-enable stack discipline has been simplified. * An optional mechanism to change the base ISA used by supervisor and user modes has been added to the `mstatus` CSR, and the field previously called Base in `misa` has been renamed to `MXL` for consistency. * Clarified expected use of XS to summarize additional extension state status fields in `mstatus`. * Optional vectored interrupt support has been added to the `mtvec` and`stvec` CSRs. * The SEIP and UEIP bits in the `mip` CSR have been redefined to support software injection of external interrupts. * The `mbadaddr` register has been subsumed by a more general `mtval`register that can now capture bad instruction bits on an illegal-instruction fault to speed instruction emulation. * The machine-mode base-and-bounds translation and protection schemes have been removed from the specification as part of moving the virtual memory configuration to `sptbr` (now `satp`). Some of the motivation for the base and bound schemes are now covered by the PMP registers, but space remains available in `mstatus` to add these back at a later date if deemed useful. * In systems with only M-mode, or with both M-mode and U-mode but without U-mode trap support, the `medeleg` and `mideleg` registers now do not exist, whereas previously they returned zero. * Virtual-memory page faults now have `mcause` values distinct from physical-memory access faults. Page-fault exceptions can now be delegated to S-mode without delegating exceptions generated by PMA and PMP checks. * An optional physical-memory protection (PMP) scheme has been proposed. * The supervisor virtual memory configuration has been moved from the`mstatus` register to the `sptbr` register. Accordingly, the `sptbr`register has been renamed to `satp` (Supervisor Address Translation and Protection) to reflect its broadened role. * The SFENCE.VM instruction has been removed in favor of the improved SFENCE.VMA instruction. * The `mstatus` bit MXR has been exposed to S-mode via `sstatus`. * The polarity of the PUM bit in `sstatus` has been inverted to shorten code sequences involving MXR. The bit has been renamed to SUM. * Hardware management of page-table entry Accessed and Dirty bits has been made optional; simpler implementations may trap to software to set them. * The counter-enable scheme has changed, so that S-mode can control availability of counters to U-mode. * H-mode has been removed, as we are focusing on recursive virtualization support in S-mode. The encoding space has been reserved and may be repurposed at a later date. * A mechanism to improve virtualization performance by trapping S-mode virtual-memory management operations has been added. * The Supervisor Binary Interface (SBI) chapter has been removed, so that it can be maintained as a separate specification. **_Preface to Version 1.9.1_** This is version 1.9.1 of the RISC-V privileged architecture proposal. Changes from version 1.9 include: * Numerous additions and improvements to the commentary sections. * Change configuration string proposal to be use a search process that supports various formats including Device Tree String and flattened Device Tree. * Made `misa` optionally writable to support modifying base and supported ISA extensions. CSR address of `misa` changed. * Added description of debug mode and debug CSRs. * Added a hardware performance monitoring scheme. Simplified the handling of existing hardware counters, removing privileged versions of the counters and the corresponding delta registers. * Fixed description of SPIE in presence of user-level interrupts. Historical Rationale for Extensions ==================== ## [](#historical-rationale-for-extensions)Appendix A: Historical Rationale for Extensions This appendix contains the rationale for RISC-V ISA extensions at the time they were ratified. Unlike the ISA specification, this appendix is ordered chronologically, so as to convey the motivation and architectural reasoning underpinning each extension at the time of ratification. For extensions ratified prior to the conception of this appendix (ca. 2025), the rationale will be added over time. In cases where the rationale was not recorded, the authors and editors will synthesize it from the historical record. ### [](#smepmp%5Frationale)"Smepmp" Extension for PMP Enhancements for memory access and execution prevention in Machine mode 1. Since a CSR for security and / or global PMP behavior settings is not available with the current spec, we needed to define a new `mseccfg` CSR. This new CSR will allow us to add further security configuration options in the future and also allow developers to verify the existence of the new mechanisms defined on this extension. 2. There are use cases where developers want to enforce PMP rules in M-mode during the boot process, that are also able to modify, merge, and / or remove later on. Since a rule that is enforced in M-mode also needs to be locked (or else badly written or malicious M-mode software can remove it at any time), the only way for developers to approach this is to keep adding PMP rules to the chain and rely on rule priority. This is a waste of PMP rules and since it’s only needed during boot, `mseccfg.RLB` is a simple workaround that can be used temporarily and then disabled and locked down. Also when `mseccfg.MML` is set, according to 4b it’s not possible to add a _Shared-Region_ rule with executable privileges. So RLB can be set temporarily during the boot process to register such regions. Note that it’s still possible to register executable _Shared-Region_ rules using initial register settings (that may include `mseccfg.MML` being set and the rule being set on PMP registers) on **PMP reset**, without using RLB. | | **Be aware that RLB introduces a security vulnerability if left set after the boot process is over and in general it should be used with caution, even when used temporarily.** Having editable PMP rules in M-mode gives a false sense of security since it only takes a few malicious instructions to lift any PMP restrictions this way. It doesn’t make sense to have a security control in place and leave it unprotected. Rule Locking Bypass is only meant as a way to optimize the allocation of PMP rules, catch errors during debugging, and allow the bootrom/firmware to register executable _Shared-Region_ rules. If developers / vendors have no use for such functionality, they should never set mseccfg.RLB and if possible hard-wire it to 0\. In any case **RLB should be disabled and locked as soon as possible**. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | If mseccfg.RLB is not used and left unset, it will be locked as soon as a PMP rule/entry with the pmpcfg.L bit set is configured. | | ------------------------------------------------------------------------------------------------------------------------------------ | | | Since PMP rules with a higher priority override rules with a lower priority, locked rules must precede non-locked rules. | | --------------------------------------------------------------------------------------------------------------------------- | 3. With the current spec M-mode can access any memory region unless restricted by a PMP rule with the `pmpcfg.L` bit set. There are cases where this approach is overly permissive, and although it’s possible to restrict M-mode by adding PMP rules during the boot process, this can also be seen as a waste of PMP rules. Having the option to block anything by default, and use PMP as an allowlist for M-mode is considered a safer approach. This functionality may be used during the boot process or upon **PMP reset**, using initial register settings. 4. The current dual meaning of the `pmpcfg.L` bit that marks a rule as Locked and **enforced** on all modes is neither flexible nor clean. With the introduction of _Machine Mode Lock-down_ the `pmpcfg.L` bit distinguishes between rules that are **enforced** **only** in M-mode (_M-mode-only_) or **only** in S/U-modes (_S/U-mode-only_). The rule locking becomes part of the definition of an _M-mode-only_ rule, since when a rule is added in M mode, if not locked, can be modified or removed in a few instructions. On the other hand, S/U modes can’t modify PMP rules anyway so locking them doesn’t make sense. 1. This separation between _M-mode-only_ and _S/U-mode-only_ rules also allows us to distinguish which regions are to be used by processes in Machine mode (`pmpcfg.L == 1`) and which by Supervisor or User mode processes (`pmpcfg.L == 0`), in the same way the U bit on the Virtual Memory’s PTEs marks which Virtual Memory pages are to be used by User mode applications (U=1) and which by the Supervisor / OS (U=0). With this distinction in place we are able to implement memory access and execution prevention in M-mode for any physical memory region that is not _M-mode-only_. An attacker that manages to tamper with a memory region used by S/U mode, even after successfully tricking a process running in M-mode to use or execute that region, will fail to perform a successful attack since that region will be _S/U-mode-only_ hence any access when in M-mode will trigger an access exception. | | In order to support zero-copy transfers between M-mode and S/U-mode we need to either allow shared memory regions, or introduce a mechanism similar to the sstatus.SUM bit to temporary allow the high-privileged mode (in this case M-mode) to be able to perform loads and stores on the region of a less-privileged process (in this case S/U-mode). In our case after discussion within the group it seemed a better idea to follow the first approach and have this functionality encoded on a per-rule basis to avoid the risk of leaving a temporary, global bypass active when exiting M-mode, hence rendering memory access prevention useless. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Although it’s possible to use mstatus.MPRV in M-mode to read/write data on an _S/U-mode-only_ region using general purpose registers for copying, this will happen with S/U-mode permissions, honoring any MMU restrictions put in place by S-mode. Of course it’s still possible for M-mode to tamper with the page tables and / or add _S/U-mode-only_ rules and bypass the protections put in place by S-mode but if an attacker has managed to compromise M-mode to such extent, no security guarantees are possible in any way. **Also note that the threat model we present here assumes buggy software in M-mode, not compromised software**. We considered disabling mstatus.MPRV but it seemed too much and out of scope. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | _Shared-region_ rules can be used both for zero-copy data transfers and for sharing code segments. The latter may be used for example to allow S/U-mode to execute code by the vendor, that makes use of some vendor-specific ISA extension, without having to go through the firmware with an ecall. This is similar to the vDSO approach followed on Linux, that allows user space code to execute kernel code without having to perform a system call. To make sure that shared data regions can’t be executed and shared code regions can’t be modified, the encoding changes the meaning of the `pmpcfg.X bit`. In case of shared data regions, with the exception of the `pmpcfg.LRWX=1111` encoding, the `pmpcfg.X` bit marks the capability of S/U-mode to write to that region, so it’s not possible to encode an executable shared data region. In case of shared code regions, the `pmpcfg.X` bit marks the capability of M-mode to read from that region, and since `pmpcfg.RW=01` is used for encoding the shared region, it’s not possible to encode a shared writable code region. | | For adding _Shared-region_ rules with executable privileges to share code segments between M-mode and S/U-mode, mseccfg.RLB needs to be implemented, or else such rules can only be added together with mseccfg.MML being set on **PMP Reset**. That’s because the reserved encoding pmpcfg.RW=01 being used for _Shared-region_ rules is only defined when mseccfg.MML is set, and 4b prevents the addition of rules with executable privileges on M-mode after mseccfg.MML is set unless mseccfg.RLB is also set. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Using the pmpcfg.LRWX=1111 encoding for a locked shared read-only data region was decided later on, its initial meaning was an M-mode-only read/write/execute region. The reason for that change was that the already defined shared data regions were not locked, so r/w access to M-mode couldn’t be restricted. In the same way we have execute-only shared code regions for both modes, it was decided to also be able to allow a least-privileged shared data region for both modes. This approach allows for example to share the .text section of an ELF with a shared code region and the .rodata section with a locked shared data region, without allowing M-mode to modify .rodata. We also decided that having a locked read/write/execute region in M-mode doesn’t make much sense and could be dangerous, since M-mode won’t be able to add further restrictions there (as in the case of S/U-mode where S-mode can further limit access to an pmpcfg.LWRX=0111 region through the MMU), leaving the possibility of modifying an executable region in M-mode open. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | For encoding Shared-region rules initially we used one of the two reserved bits on pmpcfg (bit 5) but in order to avoid allocating an extra bit, since those bits are a very limited resource, it was decided to use the reserved R=0,W=1 combination. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 2. The idea with this restriction is that after the Firmware or the OS running in M-mode is initialized and `mseccfg.MML` is set, no new code regions are expected to be added since nothing else is expected to run in M-mode (everything else will run in S/U mode). Since we want to limit the attack surface of the system as much as possible, it makes sense to disallow any new code regions which may include malicious code, to be added/executed in M-mode. 3. In case `mseccfg.MMWP` is not set, M-mode can still access and execute any region not covered by a PMP rule. Since we try to prevent M-mode from executing malicious code and since an attacker may manage to place code on some region not covered by PMP (e.g. a directly-addressable flash memory), we need to ensure that M-mode can only execute the code segments initialized during firmware / OS initialization. 4. We are only using the encoding `pmpcfg.RW=01` together with `mseccfg.MML`, if `mseccfg.MML` is not set the encoding remains usable for future use. 8.1. "Smrnmi" Extension for Resumable Non-Maskable Interrupts, Version 1.0 ==================== ## [](#rnmi)8.1\. "Smrnmi" Extension for Resumable Non-Maskable Interrupts, Version 1.0 The base machine-level architecture supports only unresumable non-maskable interrupts (UNMIs), where the NMI jumps to a handler in machine mode, overwriting the current `mepc` and `mcause` register values. If the hart had been executing machine-mode code in a trap handler, the previous values in `mepc` and `mcause` would not be recoverable and so execution is not generally resumable. The Smrnmi extension adds support for resumable non-maskable interrupts (RNMIs) to RISC-V. The extension adds four new CSRs (`mnepc`, `mncause`,`mnstatus`, and `mnscratch`) to hold the interrupted state, and one new instruction, MNRET, to resume from the RNMI handler. ### [](#8-1-1-rnmi-interrupt-signals)8.1.1\. RNMI Interrupt Signals The `rnmi` interrupt signals are inputs to the hart. These interrupts have higher priority than any other interrupt or exception on the hart and cannot be disabled by software. Specifically, they are not disabled by clearing the `mstatus`.MIE register. ### [](#8-1-2-rnmi-handler-addresses)8.1.2\. RNMI Handler Addresses The RNMI interrupt trap handler address is implementation-defined. RNMI also has an associated exception trap handler address, which is implementation defined. | | For example, some implementations might use the address specified in mtvec as the RNMI exception trap handler. | | ----------------------------------------------------------------------------------------------------------------- | ### [](#8-1-3-rnmi-csrs)8.1.3\. RNMI CSRs This extension adds additional M-mode CSRs to enable a resumable non-maskable interrupt (RNMI). ![Resumable NMI scratch register `mnscratch`](_images/diag-0074ce2c284850b6e449ed273be6fcf4f41e6e44.svg) Figure 1\. Resumable NMI scratch register `mnscratch` The `mnscratch` CSR holds an MXLEN-bit read-write register which enables the RNMI trap handler to save and restore the context that was interrupted. ![Resumable NMI program counter `mnepc`.](_images/diag-43606b08b31719fe318ecd008e39787ca1cbab75.svg) Figure 2\. Resumable NMI program counter `mnepc`. The `mnepc` CSR is an MXLEN-bit read-write register which on entry to the RNMI trap handler holds the PC of the instruction that took the interrupt. The low bit of `mnepc` (`mnepc[0]`) is always zero. On implementations that support only IALIGN=32, the two low bits (`mnepc[1:0]`) are always zero. If an implementation allows IALIGN to be either 16 or 32 (by changing CSR `misa`, for example), then, whenever IALIGN=32, bit `mnepc[1]` is masked on reads so that it appears to be 0\. This masking occurs also for the implicit read by the MNRET instruction. Though masked, `mnepc[1]`remains writable when IALIGN=32. `mnepc` is a **WARL** register that must be able to hold all valid virtual addresses. It need not be capable of holding all possible invalid addresses. Prior to writing `mnepc`, implementations may convert an invalid address into some other invalid address that `mnepc` is capable of holding. ![Resumable NMI cause `mncause`.](_images/diag-c39c0e7ed4908fd54999df3cbb966f027c967eb6.svg) Figure 3\. Resumable NMI cause `mncause`. The `mncause` CSR holds the reason for the RNMI. If the reason is an interrupt, bit MXLEN-1 is set to 1, and the RNMI cause is encoded in the least-significant bits. If the reason is an interrupt and RNMI causes are not supported, bit MXLEN-1 is set to 1, and zero is written to the least-significant bits. If the reason is an exception within M-mode that results in a double trap as specified in the Smdbltrp extension, bit MXLEN-1 is set to 0 and the least-significant bits are set to the cause code corresponding to the exception that precipitated the double trap. ![Resumable NMI status register `mnstatus`.](_images/diag-e7c59bcf4e5dc16e7646527d049201b7253e39b0.svg) Figure 4\. Resumable NMI status register `mnstatus`. The `mnstatus` CSR holds a two-bit field, MNPP, which on entry to the RNMI trap handler holds the privilege mode of the interrupted context, encoded in the same manner as `mstatus`.MPP. It also holds a one-bit field, MNPV, which on entry to the RNMI trap handler holds the virtualization mode of the interrupted context, encoded in the same manner as`mstatus`.MPV. If the Zicfilp extension is implemented, `mnstatus` also holds the MNPELP field, which on entry to the RNMI trap handler holds the previous `ELP` state. When an RNMI trap is taken, MNPELP is set to `ELP` and `ELP` is set to 0. `mnstatus` also holds the NMIE bit. When NMIE=1, non-maskable interrupts are enabled. When NMIE=0, _all_ interrupts are disabled. When NMIE=0, the hart behaves as though `mstatus`.MPRV were clear, regardless of the current setting of `mstatus`.MPRV. Upon reset, NMIE contains the value 0. | | RNMIs are masked out of reset to give software the opportunity to initialize data structures and devices for subsequent RNMI handling. | | ----------------------------------------------------------------------------------------------------------------------------------------- | Software can set NMIE to 1, but attempts to clear NMIE have no effect. | | Normally, only reset sequences will explicitly set the NMIE bit. That the NMIE bit is settable does not suffice to support the nesting of RNMIs. To support this feature in a direct manner would have required allowing software to clear the NMIE bit—a design choice that would have contravened the concept of non-maskability. Software that wishes to minimize the latency until the next RNMI is taken can follow the top-half/bottom-half model, where the RNMI handler itself only enqueues a task to a task queue then returns. The bulk of the interrupt servicing is performed later, with RNMIs enabled. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | For the purposes of the WFI instruction, NMIE is a global interrupt enable, meaning that the setting of NMIE does not affect the operation of the WFI instruction. The other bits in `mnstatus` are _reserved_; software should write zeros and hardware implementations should return zeros. ### [](#8-1-4-mnret-instruction)8.1.4\. MNRET Instruction MNRET is an M-mode-only instruction that uses the values in `mnepc` and`mnstatus` to return to the program counter, privilege mode, and virtualization mode of the interrupted context. This instruction also sets `mnstatus`.NMIE. If MNRET changes the privilege mode to a mode less privileged than M, it also sets `mstatus`.MPRV to 0\. If the Zicfilp extension is implemented, then if the new privileged mode is _y_, MNRET sets `ELP` to the logical AND of _y_LPE (see [Landing-Pad-Enabled (LPE) State](priv-cfi.html#FCFIACT)) and `mnstatus`.MNPELP. ### [](#8-1-5-rnmi-operation)8.1.5\. RNMI Operation When an RNMI interrupt is detected, the interrupted PC is written to the`mnepc` CSR, the type of RNMI to the `mncause` CSR, and the privilege mode of the interrupted context to the `mnstatus` CSR. The`mnstatus`.NMIE bit is cleared, masking all interrupts. The hart then enters machine-mode and jumps to the RNMI trap handler address. The RNMI handler can resume original execution using the new MNRET instruction, which restores the PC from `mnepc`, the privilege mode from`mnstatus`, and also sets `mnstatus`.NMIE, which re-enables interrupts. If the hart encounters an exception while executing in M-mode with the `mnstatus`.NMIE bit clear, the actions taken are the same as if the exception had occurred while `mnstatus`.NMIE were set, except that the program counter is set to the RNMI exception trap handler address. | | The Smrnmi extension does not change the behavior of the MRET and SRET instructions. In particular, MRET and SRET are unaffected by themnstatus.NMIE bit, and their execution does not alter themnstatus.NMIE bit. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 9.1. "Smcdeleg/Ssccfg" Counter Delegation Extensions, Version 1.0 ==================== ## [](#smcdeleg)9.1\. "Smcdeleg/Ssccfg" Counter Delegation Extensions, Version 1.0 In modern “Rich OS” environments, hardware performance monitoring resources are managed by the kernel, kernel driver, and/or hypervisor. Counters may be configured with differing scopes, in some cases counting events system-wide, while in others counting events on behalf of a single virtual machine or application. In such environments, the latency of counter writes has a direct impact on overall profiling overhead as a result of frequent counter writes during: 1. Sample collection, to clear overflow indication, and reload overflowed counter(s) 2. Context switch, between processes, threads, containers, or virtual machines These extensions provide a means for M-mode to allow writing select counters and event selectors from S/HS-mode. The purpose is to avert transitions to and from M-mode that add latency to these performance critical supervisor/hypervisor code sections. These extensions also defines one new CSR, scountinhibit. For a Machine-level environment, extension **Smcdeleg** (‘Sm’ for Privileged architecture and Machine-level extension, ‘cdeleg’ for Counter Delegation) encompasses all added CSRs and all behavior modifications for a hart, over all privilege levels. For a Supervisor-level environment, extension **Ssccfg** (‘Ss’ for Privileged architecture and Supervisor-level extension, ‘ccfg’ for Counter Configuration) provides access to delegated counters, and to new supervisor-level state.For a RISC-V hardware platform, Smcdeleg and Ssccfg must always be implemented in tandem. ### [](#9-1-1-counter-delegation)9.1.1\. Counter Delegation The `mcounteren` register allows M-mode to provide the next-lower privilege mode with read access to select counters.When the Smcdeleg/Ssccfg extensions are enabled (`menvcfg`.CDE=1), it further allows M-mode to delegate select counters to S-mode. The `siselect` (and `vsiselect`) index range 0x40-0x5F is reserved for delegated counter access. When a counter _i_ is delegated (`mcounteren`\[_i_\]=1 and `menvcfg`.CDE=1), the register state associated with counter _i_ can be read or written via `sireg*`, while `siselect` holds 0x40+_i_. The counter state accessible via alias CSRs is shown in the table below. Table 1\. Indirect HPM State Mappings **`siselect` value** `sireg` **`sireg4`** **`sireg2`** **`sireg5`** 0x40 `cycle`1 `cycleh`1 `cyclecfg`14 `cyclecfgh`14 0x41 _See below_ 0x42 `instret`1 `instreth`1 `instretcfg`14 `instretcfgh`14 0x43 `hpmcounter3`2 `hpmcounter3h`2 `hpmevent3`2 `hpmevent3h`23 … … … … … 0x5F `hpmcounter31`2 `hpmcounter31h`2 `hpmevent31`2 `hpmevent31h`23 1 Depends on Zicntr support 2 Depends on Zihpm support 3 Depends on Sscofpmf support 4 Depends on Smcntrpmf support `hpmevent_i_` may represent a subset of the state accessed by the `mhpmevent_i_` register. Specifically, if Sscofpmf is implemented, event selector bit 62 (MINH) is read-only 0 when accessed through `sireg*`. Likewise, `cyclecfg` and `instretcfg` may represent a subset of the state accessed by the `mcyclecfg` and `minstretcfg` registers, respectively. If Smcntrpmf is implemented, counter configuration register bit 62 (MINH) is read-only 0 when accessed through `sireg*`. If extension Smstateen is implemented, refer to extensions Smcsrind/Sscsrind (["Smcsrind/Sscsrind" Indirect CSR Access](indirect-csr.html)) for how setting bit 60 of CSR`mstateen0` to zero prevents access to registers `siselect`, `sireg*`,`vsiselect`, and `vsireg*` from privileged modes less privileged than M-mode, and likewise how setting bit 60 of `hstateen0` to zero prevents access to `siselect` and `sireg*` (really `vsiselect` and `vsireg*`) from VS-mode. The remaining rules of this section apply only when access to a CSR is not blocked by `mstateen0`\[60\] = 0 or `hstateen0`\[60\] = 0. While the privilege mode is M or S and `siselect` holds a value in the range 0x40-0x5F, illegal-instruction exceptions are raised for the following cases: \* attempts to access any `sireg*` when `menvcfg`.CDE = 0; \* attempts to access `sireg3` or `sireg6`; \* attempts to access `sireg4` or `sireg5` when XLEN = 64; \* attempts to access `sireg*` when `siselect` \= 0x41, or when the counter selected by `siselect` is not delegated to S-mode (the corresponding bit in `mcounteren` \= 0). | | _The memory-mapped mtime register is not a performance monitoring counter to be managed by supervisor software, hence the special treatment of siselect value 0x41 described above._ | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For each `siselect` and `sireg*` combination defined in [Table 1](#indirect-hpm-state-mappings), the table further indicates the extensions upon which the underlying counter state depends.If any extension upon which the underlying state depends is not implemented, an attempt from M or S mode to access the given state through `sireg*` raises an illegal-instruction exception. If the hypervisor (H) extension is also implemented, then as specified by extensions Smcsrind/Sscsrind, a virtual-instruction exception is raised for attempts from VS-mode or VU-mode to directly access `vsiselect`or `vsireg*`, or attempts from VU-mode to access `siselect` or `sireg*`. Furthermore, while `vsiselect` holds a value in the range 0x40-0x5F: \* An attempt to access any `vsireg*` from M or S mode raises an illegal-instruction exception. \* An attempt from VS-mode to access any `sireg*` (really `vsireg*`) raises an illegal-instruction exception if `menvcfg`.CDE = 0, or a virtual-instruction exception if `menvcfg`.CDE = 1. ### [](#9-1-2-supervisor-counter-inhibit-scountinhibit-register)9.1.2\. Supervisor Counter Inhibit (`scountinhibit`) Register Smcdeleg/Ssccfg defines a new `scountinhibit` register, a masked alias of `mcountinhibit`. For counters delegated to S-mode, the associated `mcountinhibit` bits can be accessed via `scountinhibit`.For counters not delegated to S-mode, the associated bits in `scountinhibit` are read-only zero. When `menvcfg`.CDE=0, attempts to access `scountinhibit` raise an illegal-instruction exception. When Supervisor Counter Delegation is enabled, attempts to access `scountinhibit` from VS-mode or VU-mode raise a virtual-instruction exception. ### [](#9-1-3-virtualizing-scountovf)9.1.3\. Virtualizing `scountovf` For implementations that support Smcdeleg/Ssccfg, Sscofpmf, and the H extension, when `menvcfg`.CDE=1, attempts to read `scountovf` from VS-mode or VU-mode raise a virtual-instruction exception. ### [](#9-1-4-virtualizing-local-counter-overflow-interrupts)9.1.4\. Virtualizing Local-Counter-Overflow Interrupts For implementations that support Smcdeleg, Sscofpmf, and Smaia, the local-counter-overflow interrupt (LCOFI) bit (bit 13) in each of CSRs`mvip` and `mvien` is implemented and writable. For implementations that support Smcdeleg/Ssccfg, Sscofpmf, Smaia/Ssaia, and the H extension, the LCOFI bit (bit 13) in each of `hvip`and `hvien` is implemented and writable. | | _The hvip register is defined by the hypervisor (H) extension, while the mvip, mvien and hvien registers are defined by the Smaia/Ssaia extensions._ _By virtue of implementing hvip.LCOFI, it is implicit that the LCOFI bit (bit 13) in each of vsie and vsip is also implemented._ _Requiring support for the LCOFI bits listed above ensures that virtual LCOFIs can be delivered to an OS running in S-mode, and to a guest OS running in VS-mode. It is optional whether the LCOFI bit (bit 13) in each of mideleg and hideleg, which allows all LCOFIs to be delegated to S-mode and VS-mode, respectively, is implemented and writable._ | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 7.1. "Smcntrpmf" Cycle and Instret Privilege Mode Filtering, Version 1.0 ==================== ## [](#smcntrpmf)7.1\. "Smcntrpmf" Cycle and Instret Privilege Mode Filtering, Version 1.0 ### [](#7-1-1-introduction)7.1.1\. Introduction The cycle and instret counters serve to support user mode self-profiling usages, wherein a user can read the counter(s) twice and compute the delta(s) to evaluate user software performance and behavior. By default, these counters are not filtered by privilege mode, and thus they continue to increment while traps (e.g., page faults or interrupts) to more privileged code are handled. This causes two problems: * It introduces unpredictable noise to the counter values observed by the user. * It leaks information about privileged software execution to user mode. Smcntrpmf remedies these issues by introducing privilege mode filtering for the cycle and instret counters. ### [](#7-1-2-csrs)7.1.2\. CSRs #### [](#7-1-2-1-machine-counter-configuration-mcyclecfg-minstretcfg-registers)7.1.2.1\. Machine Counter Configuration (`mcyclecfg`, `minstretcfg`) Registers mcyclecfg and minstretcfg are 64-bit registers that configure privilege mode filtering for the cycle and instret counters, respectively. | 63 | 62 | 61 | 60 | 59 | 58 | 57:0 | | -- | ---- | ---- | ---- | ----- | ----- | ------ | | 0 | MINH | SINH | UINH | VSINH | VUINH | _WPRI_ | | Field | Description | | ----- | --------------------------------------------------------- | | MINH | If set, then counting of events in M-mode is inhibited | | SINH | If set, then counting of events in S/HS-mode is inhibited | | UINH | If set, then counting of events in U-mode is inhibited | | VSINH | If set, then counting of events in VS-mode is inhibited | | VUINH | If set, then counting of events in VU-mode is inhibited | When all _x_INH bits are zero, event counting is enabled in all modes. For each bit in 61:58, if the associated privilege mode is not implemented, the bit is read-only zero. For RV32, bits 63:32 of mcyclecfg can be accessed via the mcyclecfgh CSR, and bits 63:32 of minstretcfg can be accessed via the minstretcfgh CSR. The content of these registers may be accessible from Supervisor level if the Smcdeleg/Ssccfg extensions are implemented. | | The more natural CSR number for mcyclecfg would be 0x320, but that was allocated to mcountinhibit. This register format matches that specified for programmable counters by Sscofpmf. The bit position for the OF bit (bit 63) is read-only 0, since these counters do not generate local-counter-overflow interrupts on overflow. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#7-1-3-counter-behavior)7.1.3\. Counter Behavior The fundamental behavior of cycle and instret is modified in that counting does not occur while executing in an inhibited privilege mode. Further, the following defines how transitions between a non-inhibited privilege mode and an inhibited privilege mode are counted. The cycle counter will simply count CPU cycles while the CPU is in a non-inhibited privilege mode. Mode transition operations (traps and trap returns) may take multiple clock cycles, and the change of privilege mode may be reported as occurring in any one of those cycles (possibly different for each occurrence of a trap or trap return). | | The RISC-V ISA has no requirement that the number of cycles for a trap or trap return be the same for all occurrences. Implementations are free to determine the extent to which this number may be consistent and predictable (or not), and the same is true for the specific cycle in which privilege mode changes. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | For the instret counter, most instructions do not affect mode transitions, so for those the behavior is clear: instructions that retire in a non-inhibited mode increment instret, and instructions that retire in an inhibited mode do not. There are two types of instructions that can affect a privilege mode change: instructions that cause synchronous exceptions to a more privileged mode, and xRET instructions that return to a less privileged mode. The former are not considered to retire, and hence do not increment instret. The latter do retire, and should increment instret only if the originating privilege mode is not inhibited. | | The instret definition above is intended to ensure that the counter increments in a predictable fashion. For example, consider a scenario where minstretcfg is configured such that all modes other than U-mode are inhibited. A user mode load should increment only once, even if it takes a page fault or other exception. With this definition, the faulting execution of the load will not increment (it does not retire), the handler instructions will not increment (they execute in an inhibited mode), including the xRET (it arguably retires in a non-inhibited mode, but it originates in an inhibited mode). Only once the load is re-executed and retires will it increment instret. In cases where an instruction is emulated by software running in a privilege mode that is inhibited in minstretcfg, the emulation routine must emulate the instret increment. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | 11.1. "Smctr" Control Transfer Records Extension, Version 1.0 ==================== ## [](#smctr)11.1\. "Smctr" Control Transfer Records Extension, Version 1.0 A method for recording control flow transfer history is valuable not only for performance profiling but also for debugging. Control flow transfers refer to jump instructions (including function calls and returns), taken branch instructions, traps, and trap returns. Profiling tools, such as Linux perf, collect control transfer history when sampling software execution, thereby enabling tools, like AutoFDO, to identify hot paths for optimization. Control flow trace capabilities offer very deep transfer history, but the volume of data produced can result in significant performance overheads due to memory bandwidth consumption, buffer management, and decoder overhead. The Control Transfer Records (CTR) extension provides a method to record a limited history in register-accessible internal chip storage, with the intent of dramatically reducing the performance overhead and complexity of collecting transfer history. CTR defines a circular (FIFO) buffer. Each buffer entry holds a record for a single recorded control flow transfer. The number of records that can be held in the buffer depends upon both the implementation (the maximum supported depth) and the CTR configuration (the software selected depth). Only qualified transfers are recorded. Qualified transfers are those that meet the filtering criteria, which include the privilege mode and the transfer type. Recorded transfers are inserted at the write pointer, which is then incremented, while older recorded transfers may be overwritten once the buffer is full. Or the user can enable RAS (Return Address Stack) emulation mode, where only function calls are recorded, and function returns pop the last call record. The source PC, target PC, and some optional metadata (transfer type, elapsed cycles) are stored for each recorded transfer. The CTR buffer is accessible through an indirect CSR interface, such that software can specify which logical entry in the buffer it wishes to read or write. Logical entry 0 always corresponds to the youngest recorded transfer, followed by entry 1 as the next youngest, and so on. The machine-level extension, **Smctr**, encompasses all newly added Control Status Registers (CSRs), instructions, and behavior modifications for a hart across all privilege levels. The corresponding supervisor-level extension, **Ssctr**, is essentially identical to Smctr, except that it excludes machine-level CSRs and behaviors not intended to be directly accessible at the supervisor level. Smctr and Ssctr depend on both the implementation of S-mode and the Sscsrind extension. ### [](#CSRs)11.1.1\. CSRs #### [](#11-1-1-1-machine-control-transfer-records-control-register-mctrctl)11.1.1.1\. Machine Control Transfer Records Control Register (`mctrctl`) The `mctrctl` register is a 64-bit read/write register that enables and configures the CTR capability. ![Machine Control Transfer Records Control Register Format](_images/svg-060a4b8809aa69a9ebe0c92be8b4973e1ebf0e3f.svg) Figure 1\. Machine Control Transfer Records Control Register Format __Table 1\. Machine Control Transfer Records Control Register Field Definitions__ | Field | Description | | ------------- | ----------------------------------------------------------------------------------------------------------------------------------- | | M, S, U | Enable transfer recording in the selected privileged mode(s). | | RASEMU | Enables RAS (Return Address Stack) Emulation Mode. See [RAS (Return Address Stack) Emulation Mode](#ras-emulation-mode). | | MTE | Enables recording of traps to M-mode when M=0. See [External Traps](#external-traps). | | STE | Enables recording of traps to S-mode when S=0. See [External Traps](#external-traps). | | BPFRZ | Set sctrstatus.FROZEN on a breakpoint exception that traps to M-mode or S-mode. See [Freeze](#freeze). | | LCOFIFRZ | Set sctrstatus.FROZEN on local-counter-overflow interrupt (LCOFI) that traps to M-mode or S-mode. See [Freeze](#freeze). | | EXCINH | Inhibit recording of exceptions. See [Transfer Type Filtering](#transfer-type-filtering). | | INTRINH | Inhibit recording of interrupts. See [Transfer Type Filtering](#transfer-type-filtering). | | TRETINH | Inhibit recording of trap returns. See [Transfer Type Filtering](#transfer-type-filtering). | | NTBREN | Enable recording of not-taken branches. See [Transfer Type Filtering](#transfer-type-filtering). | | TKBRINH | Inhibit recording of taken branches. See [Transfer Type Filtering](#transfer-type-filtering). | | INDCALLINH | Inhibit recording of indirect calls. See [Transfer Type Filtering](#transfer-type-filtering). | | DIRCALLINH | Inhibit recording of direct calls. See [Transfer Type Filtering](#transfer-type-filtering). | | INDJMPINH | Inhibit recording of indirect jumps (without linkage). See [Transfer Type Filtering](#transfer-type-filtering). | | DIRJMPINH | Inhibit recording of direct jumps (without linkage). See [Transfer Type Filtering](#transfer-type-filtering). | | CORSWAPINH | Inhibit recording of co-routine swaps. See [Transfer Type Filtering](#transfer-type-filtering). | | RETINH | Inhibit recording of function returns. See [Transfer Type Filtering](#transfer-type-filtering). | | INDLJMPINH | Inhibit recording of other indirect jumps (with linkage). See [Transfer Type Filtering](#transfer-type-filtering). | | DIRLJMPINH | Inhibit recording of other direct jumps (with linkage). See [Transfer Type Filtering](#transfer-type-filtering). | | Custom\[3:0\] | WARL bits designated for custom use. The value 0 must correspond to standard behavior. See [Custom Extensions](#custom-extensions). | All fields are optional except for M, S, U, and BPFRZ. All unimplemented fields are read-only 0, while all implemented fields are writable. If the Sscofpmf extension is implemented, LCOFIFRZ must be writable. | | _Because the ROI of CTR is perceived to be low for RV32 implementations, CTR does not fully support RV32\. While control flow transfers in RV32 can be recorded, RV32 cannot access_ x_ctrctl_ _bits 63:32\. A future extension could add support for RV32 by adding 3 new CSRs (mctrctlh, sctrctlh, and vsctrctlh) to provide this access._ | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#11-1-1-2-supervisor-control-transfer-records-control-register-sctrctl)11.1.1.2\. Supervisor Control Transfer Records Control Register (`sctrctl`) The `sctrctl` register provides supervisor mode access to a subset of `mctrctl`. Bits 2 and 9 in `sctrctl` are read-only 0\. As a result, the M and MTE fields in `mctrctl` are not accessible through `sctrctl`. All other `mctrctl` fields are accessible through `sctrctl`. #### [](#11-1-1-3-virtual-supervisor-control-transfer-records-control-register-vsctrctl)11.1.1.3\. Virtual Supervisor Control Transfer Records Control Register (`vsctrctl`) If the H extension is implemented, the `vsctrctl` register is a 64-bit read/write register that is VS-mode’s version of supervisor register `sctrctl`. When V=1, `vsctrctl` substitutes for the usual `sctrctl`, so instructions that normally read or modify `sctrctl` actually access `vsctrctl` instead. ![Virtual Supervisor Control Transfer Records Control Register Format](_images/svg-b877f32b9db72dae57175fdbc5129fcca01f39a3.svg) Figure 2\. Virtual Supervisor Control Transfer Records Control Register Format __Table 2\. Virtual Supervisor Control Transfer Records Control Register Field Definitions__ | Field | Description | | -------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | | S | Enable transfer recording in VS-mode. | | U | Enable transfer recording in VU-mode. | | STE | Enables recording of traps to VS-mode when S=0. See [External Traps](#external-traps).. | | BPFRZ | Set sctrstatus.FROZEN on a breakpoint exception that traps to VS-mode. See [Freeze](#freeze). | | LCOFIFRZ | Set sctrstatus.FROZEN on local-counter-overflow interrupt (LCOFI) that traps to VS-mode. See [Freeze](#freeze). | | Other field definitions match those of sctrctl. The optional fields implemented in vsctrctl should match those implemented in sctrctl. | | | | _Unlike the CTR status register or the CTR entry registers, the CTR control register has a VS-mode version. This allows a guest to manage the CTR configuration directly, without requiring traps to HS-mode, while ensuring that the guest configuration (most notably the privilege mode enable bits) do not impact CTR behavior when V=0._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#11-1-1-4-supervisor-control-transfer-records-depth-register-sctrdepth)11.1.1.4\. Supervisor Control Transfer Records Depth Register (`sctrdepth`) The 32-bit `sctrdepth` register specifies the depth of the CTR buffer. ![Supervisor Control Transfer Records Depth Register Format](_images/svg-45bc5a8be341f505f29a383d293a450296a890fd.svg) Figure 3\. Supervisor Control Transfer Records Depth Register Format __Table 3\. Supervisor Control Transfer Records Depth Register Field Definitions__ | Field | Description | | ----- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | DEPTH | WARL field that selects the depth of the CTR buffer. Encodings: ‘000 - 16 ‘001 - 32 ‘010 - 64 ‘011 - 128 ‘100 - 256 '11x - reserved The depth of the CTR buffer dictates the number of entries to which the hardware records transfers. For a depth of N, the hardware records transfers to entries 0..N-1\. All [Entry Registers](#entry-registers) read as '0' and are read-only when the selected entry is in the range N to 255\. When the depth is increased, the newly accessible entries contain unspecified but legal values. It is implementation-specific which DEPTH value(s) are supported. | Attempts to access `sctrdepth` from VS-mode or VU-mode raise a virtual-instruction exception, unless CTR state enable access restrictions apply. See [State Enable Access Control](#state-enable-access-control). | | _It is expected that operating systems (OSs) will access sctrdepth only at boot, to select the maximum supported depth value. More frequent accesses may result in reduced performance in virtualization scenarios, as a result of traps from VS-mode incurred._ _There may be scenarios where software chooses to operate on only a subset of the entries, to reduce overhead. In such cases tools may choose to read only the lower entries, and OSs may choose to save/restore only on the lower entries while using SCTRCLR to clear the others._ _The value in configurable depth lies in supporting VM migration. It is expected that a platform spec may specify that one or more CTR depth values must be supported. A hypervisor may wish to restrict guests to using one of these required depths, in order to ensure that such guests can be migrated to any system that complies with the platform spec. The trapping behavior specified for VS-mode accesses to sctrdepth ensures that the hypervisor can impose such restrictions._ | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#11-1-1-5-supervisor-control-transfer-records-status-register-sctrstatus)11.1.1.5\. Supervisor Control Transfer Records Status Register (`sctrstatus`) The 32-bit `sctrstatus` register grants access to CTR status information and is updated by the hardware whenever CTR is active. CTR is active when the current privilege mode is enabled for recording and CTR is not frozen. ![Supervisor Control Transfer Records Status Register Format](_images/svg-ae2fbf1700f610a769629098077e7afa35bbd7b7.svg) Figure 4\. Supervisor Control Transfer Records Status Register Format __Table 4\. Supervisor Control Transfer Records Status Register Field Definitions__ | Field | Description | | ------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | WRPTR | WARL field that indicates the physical CTR buffer entry to be written next. It is incremented after new transfers are recorded (see [Behavior](#behavior)), though there are exceptions when _x_ctrctl.RASEMU=1, see [RAS (Return Address Stack) Emulation Mode](#ras-emulation-mode). For a given CTR depth (where depth = 2(DEPTH+4)), WRPTR wraps to 0 on an increment when the value matches depth-1, and to depth-1 on a decrement when the value is 0\. Bits above those needed to represent depth-1 (e.g., bits 7:4 for a depth of 16) are read-only 0\. On depth changes, WRPTR holds an unspecified but legal value. | | FROZEN | Inhibit transfer recording. See [Freeze](#freeze). | Undefined bits in `sctrstatus` are WPRI. Status fields may be added by future extensions, and software should ignore but preserve any fields that it does not recognize. Undefined bits must be implemented as read-only 0, unless a custom extension is implemented and enabled (see [Custom Extensions](#custom-extensions)). | | _Logical entry 0, accessed via sireg\* when siselect\=0x200, is always the physical buffer entry preceding the WRPTR entry. More generally, the physical buffer entry Y associated with logical entry X (X < depth) can be determined using the formula Y = (WRPTR - X - 1) % depth, where depth = 2(DEPTH+4). Logical entries >= depth are read-only 0._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | _Because the sctrstatus register is updated by hardware, writes should be performed with caution. If a multi-instruction read-modify-write to sctrstatus is performed while CTR is active, and between the read and write a qualified transfer or trap that causes CTR freeze completes, a hardware update could be lost. Software may wish to ensure that CTR is inactive before performing a read-modify-write, by ensuring that either sctrstatus.FROZEN=1, or that the current privilege mode is not enabled for recording._ _When restoring CTR state, sctrstatus should be written before CTR entry state is restored. This ensures that the software writes to logical CTR entries modify the proper physical entries._ | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | _Exposing the WRPTR provides a more efficient means for synthesizing CTR entries. If a qualified control transfer is emulated, the emulator can simply increment the WRPTR, then write the synthesized record to logical entry 0\. If a qualified function return is emulated while RASEMU=1, the emulator can clear ctrsource.V for logical entry 0, then decrement the WRPTR._ _Exposing the WRPTR may also allow support for Linux perf’s [stack stitching](https://lwn.net/Articles/802821) capability._ | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | _Smctr/Ssctr depends upon implementation of S-mode because much of CTR state is accessible only through S-mode CSRs. If, in the future, it becomes desirable to remove this dependency, an extension could add mctrdepth and mctrstatus CSRs that reflect the same state as sctrdepth and sctrstatus, respectively. Further, such an extension should make CTR entries accessible via miselect/mireg\*. See [Entry Registers](#entry-registers)._ | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#entry-registers)11.1.2\. Entry Registers Control transfer records are stored in a CTR buffer, such that each buffer entry stores information about a single transfer. The CTR buffer entries are logically accessed via the indirect register access mechanism defined by the Sscsrind extension. The `siselect` index range 0x200 through 0x2FF is reserved for CTR logical entries 0 through 255\. When `siselect` holds a value in this range, `sireg` provides access to `ctrsource`, `sireg2` provides access to `ctrtarget`, and `sireg3` provides access to `ctrdata`. `sireg4`, `sireg5`, and `sireg6` are read-only 0. When `vsiselect` holds a value in 0x200..0x2FF, the `vsireg*` registers provide access to the same CTR entry register state as the analogous `sireg*` registers. There is not a separate set of entry registers for V=1. See [State Enable Access Control](#state-enable-access-control) for cases where CTR accesses from S-mode and VS-mode may be restricted. #### [](#11-1-2-1-control-transfer-record-source-register-ctrsource)11.1.2.1\. Control Transfer Record Source Register (`ctrsource`) The `ctrsource` register contains the source program counter, which is the `pc` of the recorded control transfer instruction, or the epc of the recorded trap. The valid (V) bit is set by the hardware when a transfer is recorded in the selected CTR buffer entry, and implies that data in `ctrsource`, `ctrtarget`, and `ctrdata` is valid for this entry. `ctrsource` is an MXLEN-bit WARL register that must be able to hold all valid virtual or physical addresses that can serve as a `pc`. It need not be able to hold any invalid addresses; implementations may convert an invalid address into a valid address that the register is capable of holding. When XLEN < MXLEN, both explicit writes (by software) and implicit writes (for recorded transfers) will be zero-extended. ![Control Transfer Record Source Register Format for MXLEN=64](_images/svg-7e79cdd70a53ed065913e97e3c9824c81f9386f1.svg) Figure 5\. Control Transfer Record Source Register Format for MXLEN=64 | | _CTR entry registers are defined as MXLEN, despite the_ x_ireg\*_ _CSRs used to access them being XLEN, to ensure that entries recorded in RV64 are not truncated, as a result of CSR Width Modulation, on a transition to RV32._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#11-1-2-2-control-transfer-record-target-register-ctrtarget)11.1.2.2\. Control Transfer Record Target Register (`ctrtarget`) The `ctrtarget` register contains the target (destination) program counter of the recorded transfer. For a not-taken branch, `ctrtarget` holds the PC of the next sequential instruction following the branch. The optional MISP bit is set by the hardware when the recorded transfer is an instruction whose target or taken/not-taken direction was mispredicted by the branch predictor. MISP is read-only 0 when not implemented. `ctrtarget` is an MXLEN-bit WARL register that must be able to hold all valid virtual or physical addresses that can serve as a `pc`. It need not be able to hold any invalid addresses; implementations may convert an invalid address into a valid address that the register is capable of holding. When XLEN < MXLEN, both explicit writes (by software) and implicit writes (by recorded transfers) will be zero-extended. ![Control Transfer Record Target Register Format for MXLEN=64](_images/svg-a3816f4f0228d46e7ac5d7b90bd1acd6c127ab55.svg) Figure 6\. Control Transfer Record Target Register Format for MXLEN=64 #### [](#11-1-2-3-control-transfer-record-metadata-register-ctrdata)11.1.2.3\. Control Transfer Record Metadata Register (`ctrdata`) The `ctrdata` register contains metadata for the recorded transfer. This register must be implemented, though all fields within it are optional. Unimplemented fields are read-only 0\. `ctrdata` is a 64-bit register. ![Control Transfer Record Metadata Register Format](_images/svg-b74eca27e58ba638fe64a8c61169f17f4054507b.svg) Figure 7\. Control Transfer Record Metadata Register Format __Table 5\. Control Transfer Record Metadata Register Field Definitions__ | Field | Description | Access | | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | | TYPE\[3:0\] | Identifies the type of the control flow transfer recorded in the entry, using the encodings listed in [Table 8](#transfer-type-defs). Implementations that do not support this field will report 0. | WARL | | CCV | Cycle Count Valid. See [Cycle Counting](#cycle-counting). | WARL | | CC\[15:0\] | Cycle Count, composed of the Cycle Count Exponent (CCE, in CC\[15:12\]) and Cycle Count Mantissa (CCM, in CC\[11:0\]). See[Cycle Counting](#cycle-counting). | WARL | Undefined bits in `ctrdata` are WPRI. Undefined bits must be implemented as read-only 0, unless a [Custom extension](#custom-extension) is implemented and enabled. | | _Like the [Transfer Type Filter](#transfer-type-filtering) bits in mctrctl, the ctrdata.TYPE bits leverage the E-trace itype encodings._ | | ------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#11-1-3-instructions)11.1.3\. Instructions #### [](#supervisor-ctr-clear-instruction)11.1.3.1\. Supervisor CTR Clear Instruction ![svg](_images/svg-5e9f60b9b6bddba48714931a72c3484ceaedbe54.svg) The SCTRCLR instruction performs the following operations: * Zeroes all CTR [Entry Registers](#entry-registers), for all DEPTH values * Zeroes the CTR cycle counter and CCV (see [Cycle Counting](#cycle-counting)) Any read of `ctrsource`, `ctrtarget`, or `ctrdata` that follows SCTRCLR, such that it precedes the next qualified control transfer, will return the value 0\. Further, the first recorded transfer following SCTRCLR will have `ctrdata`.CCV=0. SCTRCLR raises an illegal-instruction exception in U-mode, and a virtual-instruction exception in VU-mode, unless CTR state enable access restrictions apply. See [State Enable Access Control](#state-enable-access-control). ### [](#state-enable-access-control)11.1.4\. State Enable Access Control When Smstateen is implemented, the `mstateen0`.CTR bit controls access to CTR register state from privilege modes less privileged than M-mode. When `mstateen0`.CTR=1, accesses to CTR register state behave as described in [CSRs](#CSRs) and [Entry Registers](#entry-registers) above, while SCTRCLR behaves as described in [Supervisor CTR Clear Instruction](#supervisor-ctr-clear-instruction). When `mstateen0`.CTR=0 and the privilege mode is less privileged than M-mode, the following operations raise an illegal-instruction exception: * Attempts to access `sctrctl`, `vsctrctl`, `sctrdepth`, or `sctrstatus` * Attempts to access `sireg*` when `siselect` is in 0x200..0x2FF, or `vsireg*` when `vsiselect` is in 0x200..0x2FF * Execution of the SCTRCLR instruction When `mstateen0`.CTR=0, qualified control transfers executed in privilege modes less privileged than M-mode will continue to implicitly update entry registers and `sctrstatus`. If the H extension is implemented and `mstateen0`.CTR=1, the `hstateen0`.CTR bit controls access to supervisor CTR state when V=1\. This state includes `sctrctl` (really `vsctrctl`), `sctrstatus`, and `sireg*` (really `vsireg*`) when `siselect` (really `vsiselect`) is in 0x200..0x2FF. `hstateen0`.CTR is read-only 0 when `mstateen0`.CTR=0. When `mstateen0`.CTR=1 and `hstateen0`.CTR=1, VS-mode accesses to supervisor CTR state behave as described in [CSRs](#CSRs) and [Entry Registers](#entry-registers) above, while SCTRCLR behaves as described in [Supervisor CTR Clear Instruction](#supervisor-ctr-clear-instruction). When `mstateen0`.CTR=1 and `hstateen0`.CTR=0, both VS-mode accesses to supervisor CTR state and VS-mode execution of SCTRCLR raise a virtual-instruction exception. | | _sctrdepth_ _is not included in the above list of supervisor CTR state controlled by hstateen0.CTR since accesses to sctrdepth from VS-mode raise a virtual-instruction exception regardless of the value of hstateen0.CTR._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When `hstateen0`.CTR=0, qualified control transfers executed while V=1 will continue to implicitly update entry registers and `sctrstatus`. | | _See [Indirect CSR](indirect-csr.html) for how bit 60 in mstateen0 and hstateen0 can also restrict access to sireg\*/siselect and vsireg\*/vsiselect from privilege modes less privileged than M-mode._ | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | _Implementations that support Smctr/Ssctr but not Smstateen/Ssstateen may observe reduced performance. Because Smctr/Ssctr introduces a significant number of new CSRs, it is desirable to avoid save/restore of CTR state when possible. A hypervisor is likely to leverage State Enable to trap on the initial guest access to CTR state, delegating CTR and enabling save/restore of guest CTR state only once the guest has begun to use it. Without Smstateen/Ssstateen, a hypervisor is required to save/restore guest CTR state on every context switch._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#behavior)11.1.5\. Behavior CTR records qualified control transfers. Control transfers are qualified if they meet the following criteria: * The current privilege mode is enabled * The transfer type is not inhibited * `sctrstatus`.FROZEN is not set * The transfer completes/retires Such qualified transfers update the [Entry Registers](#entry-registers) at logical entry 0\. As a result, older entries are pushed down the stack; the record previously in logical entry 0 moves to logical entry 1, the record in logical entry 1 moves to logical entry 2, and so on. If the CTR buffer is full, the oldest recorded entry (previously at entry depth-1) is lost. Recorded transfers will set the `ctrsource`.V bit to 1, and will update all implemented record fields. | | _In order to collect accurate and representative performance profiles while using CTR, it is recommended that hardware recording of control transfers incurs no added performance overhead, e.g., in the form of retirement or instruction execution restrictions that are not present when CTR is not active._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#11-1-5-1-privilege-mode-transitions)11.1.5.1\. Privilege Mode Transitions Transfers that change the privilege mode are a special case. What is recorded, if anything, depends on whether the source privilege mode and/or target privilege mode are enabled for recording, and on the transfer type (trap or trap return). Traps between enabled privilege modes are recorded as normal. Traps from a disabled privilege mode to an enabled privilege mode are partially recorded, such that the `ctrsource`.PC is 0. Traps from an enabled mode to a disabled mode, known as external traps, are not recorded by default. See [External Traps](#external-traps). for how they can be recorded. Trap returns have similar treatment. Trap returns between enabled privilege modes are recorded as normal. Trap returns from an enabled mode back to a disabled mode are partially recorded, such that `ctrtarget`.PC is 0. Trap returns from a disabled mode to an enabled mode are not recorded. | | _If privileged software is configuring CTR on behalf of less privileged software, it should ensure that its privilege mode enable bit (e.g., sctrctl.S for Supervisor software) is cleared before a trap return to the less privileged mode. Otherwise the trap return will be recorded, leaking the privileged source pc._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Recording in Debug Mode is always inhibited. Transfers into and out of Debug Mode are never recorded. The table below provides details on recording of privilege mode transitions. Standard dependencies on FROZEN and transfer type inhibits also apply, but are not covered by the table. __Table 6\. Trap and Trap Return Recording__ | **Transfer Type** | **Source Mode** | **Target Mode** | | | ----------------- | ---------------------------- | --------------- | ----------------------------------------------------------------------------------- | | **Enabled** | **Disabled** | | | | **Trap** | **Enabled** | Recorded. | External trap. Not recorded by default, but see [External Traps](#external-traps).. | | **Disabled** | Recorded, ctrsource.PC is 0. | Not recorded. | | | **Trap Return** | **Enabled** | Recorded. | Recorded, ctrtarget.PC is 0. | | **Disabled** | Not recorded. | Not recorded. | | ##### [](#11-1-5-1-1-virtualization-mode-transitions)11.1.5.1.1\. Virtualization Mode Transitions Transitions between VS/VU-mode and M/HS-mode are unique in that they effect a change in the active CTR control register, and hence the CTR configuration. What is recorded, if anything, on these virtualization mode transitions depends upon fields from both `[ms]ctrctl` and `vsctrctl`. * `mctrctl`.M, `sctrctl`.S, and `vsctrctl`.{S,U} are used to determine whether the source and target modes are enabled; * `mctrctl`.MTE, `sctrctl`.STE, and `vsctrctl`.STE are used to determine whether an external trap is recorded (see [External Traps](#external-traps).); * `sctrctl`.LCOFIFRZ and `sctrctl`.BPFRZ determine whether CTR becomes frozen (see [Freeze](#freeze)) * For all other `_x_ctrctl` fields, the value in `vsctrctl` is used. | | _Consider an exception that traps from VU-mode to HS-mode, with vsctrctl.U=1 and sctrctl.S=1\. Because both the source mode and target mode are enabled for recording, whether the trap is recorded then depends on the CTR configuration (e.g., the [Transfer Type Filter](#transfer-type-filtering) bits) in vsctrctl, not in sctrctl._ | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#external-traps)11.1.5.1.2\. External Traps External traps are traps from a privilege mode enabled for CTR recording to a privilege mode that is not enabled for CTR recording. By default external traps are not recorded, but privileged software running in the target mode of the trap can opt-in to allowing CTR to record external traps into that mode. The `_x_ctrctl`._x_TE bits allow M-mode, S-mode, and VS-mode to opt-in separately. External trap recording depends not only on the target mode, but on any intervening modes, which are modes that are more privileged than the source mode but less privileged than the target mode. Not only must the external trap enable bit for the target mode be set, but the external trap enable bit(s) for any intervening modes must also be set. See the table below for details. | | _Requiring intervening modes to be enabled for external traps simplifies software management of CTR. Consider a scenario where S-mode software is configuring CTR for U-mode contexts A and B, such that external traps (to any mode) are enabled for A but not for B. When switching between the two contexts, S-mode can simply toggle sctrctl.STE, rather than requiring a trap to M-mode to additionally toggle mctrctl.MTE._ _This method does not provide the flexibility to record external traps to a more privileged mode but not to all intervening mode(s). Because it is expected that profiling tools generally wish to observe all external traps or none, this is not considered a meaningful limitation._ | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 7\. External Trap Enable Requirements__ | Source Mode | Target Mode | External Trap Enable(s) Required | | ----------- | -------------------------------------- | -------------------------------- | | U-mode | S-mode | sctrctl.STE | | M-mode | mctrctl.MTE, sctrctl.STE | | | S-mode | M-mode | mctrctl.MTE | | VU-mode | VS-mode | vsctrctl.STE | | HS-mode | sctrctl.STE, vsctrctl.STE | | | M-mode | mctrctl.MTE, sctrctl.STE, vsctrctl.STE | | | VS-mode | HS-mode | sctrctl.STE | | M-mode | mctrctl.MTE, sctrctl.STE | | In records for external traps, the `ctrtarget`.PC is 0. | | _No mechanism exists for recording external trap returns, because the external trap record includes all relevant information, and gives the trap handler (e.g., an emulator) the opportunity to modify the record._ | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | _Note that external trap recording does not depend on EXCINH/INTRINH. Thus, when external traps are enabled, both external interrupts and external exceptions are recorded._ _STE allows recording of traps from U-mode to S-mode as well as from VS/VU-mode to HS-mode. The hypervisor can flip sctrctl.STE before entering a guest if it wants different behavior for U-to-S vs VS/VU-to-HS._ | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If external trap recording is implemented, `mctrctl`.MTE and `sctrctl`.STE must be implemented, while `vsctrctl`.STE must be implemented if the H extension is implemented. #### [](#transfer-type-filtering)11.1.5.2\. Transfer Type Filtering Default CTR behavior, when all transfer type filter bits (`_x_ctrctl`\[47:32\]) are unimplemented or 0, is to record all control transfers within enabled privileged modes. By setting transfer type filter bits, software can opt out of recording select transfer types, or opt into recording non-default operations. All transfer type filter bits are optional. | | _Because not-taken branches are not recorded by default, the polarity of the associated enable bit (NTBREN) is the opposite of other bits associated with transfer type filtering (TKBRINH, RETINH, etc). Non-default operations require opt-in rather than opt-out._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The transfer type filter bits leverage the type definitions specified in the[RISC-V Efficient Trace Spec v2.0](https://github.com/riscv-non-isa/riscv-trace-spec/releases/download/v2.0rc2/riscv-trace-spec.pdf) (Table 4.4 and Section 4.1.1). For completeness, the definitions are reproduced below. | | _Here "indirect" is used interchangeably with "uninferrable", which is used in the trace spec. Both imply that the target of the jump is not encoded in the opcode._ | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 8\. Control Transfer Type Definitions__ | Encoding | Transfer Type Name | | -------- | ---------------------------------- | | 0 | _Not used by CTR_ | | 1 | Exception | | 2 | Interrupt | | 3 | Trap return | | 4 | Not-taken branch | | 5 | Taken branch | | 6 | _reserved_ | | 7 | _reserved_ | | 8 | Indirect call | | 9 | Direct call | | 10 | Indirect jump (without linkage) | | 11 | Direct jump (without linkage) | | 12 | Co-routine swap | | 13 | Function return | | 14 | Other indirect jump (with linkage) | | 15 | Other direct jump (with linkage) | Encodings 8 through 15 refer to various encodings of jump instructions. The types are distinguished as described below. __Table 9\. Control Transfer Type Definitions__ | Transfer Type Name | Associated Opcodes | | ----------------------------------------- | ------------------------------------------------------------------------------ | | Indirect call | JALR _x1_, _rs_ where _rs_ != _x5_ | | JALR _x5_, _rs_ where _rs_ != _x1_ | | | C.JALR _rs1_ where _rs1_ != _x5_ | | | Direct call | JAL _x1_ | | JAL _x5_ | | | C.JAL | | | CM.JALT _index_ | | | Indirect jump (without linkage) | JALR _x0_, _rs_ where _rs_ != (_x1_ or _x5_) | | C.JR _rs1_ where _rs1_ != (_x1_ or _x5_) | | | Direct jump (without linkage) | JAL _x0_ | | C.J | | | CM.JT _index_ | | | Co-routine swap | JALR _x1_, _x5_ | | JALR _x5_, _x1_ | | | C.JALR _x5_ | | | Function return | JALR _rd_, _rs_ where _rs_ \== (_x1_ or _x5_) and _rd_ != (_x1_ or _x5_) | | C.JR _rs1_ where _rs1_ \== (_x1_ or _x5_) | | | CM.POPRET(Z) | | | Other indirect jump (with linkage) | JALR _rd_, _rs_ where _rs_ != (_x1_ or _x5_) and _rd_ != (_x0_, _x1_, or _x5_) | | Other direct jump (with linkage) | JAL _rd_ where _rd_ != (_x0_, _x1_, or _x5_) | | | _If implementation of any transfer type filter bit results in reduced software performance, perhaps due to additional retirement restrictions, it is strongly recommended that this reduced performance apply only when the bit is set. Alternatively, support for the bit may be omitted. Maintaining software performance for the default CTR configuration, when all transfer type bits are cleared, is recommended._ | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#cycle-counting)11.1.5.3\. Cycle Counting The `ctrdata` register may optionally include a count of CPU cycles elapsed since the prior CTR record. The elapsed cycle count value is represented by the CC field, which has a 12-bit mantissa component (Cycle Count Mantissa, or CCM) and a 4-bit exponent component (Cycle Count Exponent, or CCE). The elapsed cycle counter (CtrCycleCounter) increments at the same rate as the `mcycle` counter. Only cycles while CTR is active are counted, where active implies that the current privilege mode is enabled for recording and CTR is not frozen. The CC field is encoded such that CCE holds 0 if the CtrCycleCounter value is less than 4096, otherwise it holds the index of the most significant one bit in the CtrCycleCounter value, minus 11\. CCM holds CtrCycleCounter bits CCE+10:CCE-1. The elapsed cycle count can then be calculated by software using the following formula: if (CCE==0): return CCM else: return (212 + CCM) << CCE-1 endif The CtrCycleCounter is reset on writes to `_x_ctrctl`, and on execution of SCTRCLR, to ensure that any accumulated cycle counts do not persist across a context switch. An implementation that supports cycle counting must implement CCV and all CCM bits, but may implement 0..4 exponent bits in CCE. Unimplemented CCE bits are read-only 0\. For implementations that support transfer type filtering, it is recommended to implement at least 3 exponent bits. This allows capturing the full latency of most functions, when recording only calls and returns. The size of the CtrCycleCounter required to support each CCE width is given in the table below. __Table 10\. Cycle Counter Size Options__ | CCE bits | CtrCycleCounter bits | Max elapsed cycle value | | -------- | -------------------- | ----------------------- | | 0 | 12 | 4095 | | 1 | 13 | 8191 | | 2 | 15 | 32764 | | 3 | 19 | 524224 | | 4 | 27 | 134201344 | | | _When CCE>1, the granularity of the reported cycle count is reduced. For example, when CCE=3, the bottom 2 bits of the cycle counter are not reported, and thus the reported value increments only every 4 cycles. As a result, the reported value represents an undercount of elapsed cycles for most cases (when the unreported bits are non-zero). On average, the undercount will be (2CCE-1\-1)/2\. Software can reduce the average undercount to 0 by adding (2CCE-1\-1)/2 to each computed cycle count value when CCE>1._ _Though this compressed method of representation results in some imprecision for larger cycle count values, it produces meaningful area savings, reducing storage per entry from 27 bits to 16._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The CC value saturates when all implemented bits in CCM and CCE are 1. The CC value is valid only when the Cycle Count Valid (CCV) bit is set. If CCV=0, the CC value might not hold the correct count of elapsed active cycles since the last recorded transfer. The next record will have CCV=0 after a write to `_x_ctrctl`, or execution of SCTRCLR, since CtrCycleCounter is reset. CCV should additionally be cleared after any other implementation-specific scenarios where active cycles might not be counted in CtrCycleCounter. #### [](#ras-emulation-mode)11.1.5.4\. RAS (Return Address Stack) Emulation Mode When the optional `_x_ctrctl`.RASEMU bit is implemented and set to 1, transfer recording behavior is altered to emulate the behavior of a return-address stack (RAS). * Indirect and direct calls are recorded as normal * Function returns pop the most recent call, by decrementing the WRPTR then invalidating the WRPTR entry (by setting ctrsource.V=0). As a result, logical entry 0 is invalidated and moves to logical entry depth-1, while logical entries 1..depth-1 move to 0..depth-2. * Co-routine swaps affect both a return and a call. Logical entry 0 is overwritten, and WRPTR is not modified. * Other transfer types are inhibited * Transfer type filtering bits (`_x_ctrctl`\[47:32\]) and external trap enable bits (`_x_ctrctl`._x_TE) are ignored | | _Profiling tools often collect call stacks along with each sample. Stack walking, however, is a complex and often slow process that may require recompilation (e.g., -fno-omit-frame-pointer) to work reliably. With RAS emulation, tools can ask CTR hardware to save call stacks even for unmodified code._ _CTR RAS emulation has limitations. The CTR buffer will contain only partial stacks in cases where the call stack depth was greater than the CTR depth, CTR recording was enabled at a lower point in the call stack than main(), or where the CTR buffer was cleared since main()._ _The CTR stack may be corrupted in cases where calls and returns are not symmetric, such as with stack unwinding (e.g., setjmp/longjmp, C++ exceptions), where stale call entries may be left on the CTR stack, or user stack switching, where calls from multiple stacks may be intermixed._ | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | _As described in [Cycle Counting](#cycle-counting), when CCV=1, the CC field provides the elapsed cycles since the prior CTR entry was recorded. This introduces implementation challenges when RASEMU=1 because, for each recorded call, there may have been several recorded calls (and returns which “popped” them) since the prior remaining call entry was recorded (see [RAS (Return Address Stack) Emulation Mode](#ras-emulation-mode)). The implication is that returns that pop a call entry not only do not reset the cycle counter, but instead add the CC field from the popped entry to the counter. For simplicity, an implementation may opt to record CCV=0 for all calls, or those whose parent call was popped, when RASEMU=1._ | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#freeze)11.1.5.5\. Freeze When `sctrstatus`.FROZEN=1, transfer recording is inhibited. This bit can be set by hardware, as described below, or by software. When `sctrctl`.LCOFIFRZ=1 and a local-counter-overflow interrupt (LCOFI) traps (as a result of an HPM counter overflow) to M-mode or to S-mode, `sctrstatus`.FROZEN is set by hardware. This inhibits CTR recording until software clears FROZEN. The LCOFI trap itself is not recorded. | | _Freeze on LCOFI ensures that the execution path leading to the sampled instruction (xepc) is preserved, and that the local-counter-overflow interrupt (LCOFI) and associated Interrupt Service Routine (ISR) do not displace any recorded transfer history state. It is the responsibility of the ISR to clear FROZEN before xRET, if continued control transfer recording is desired._ _LCOFI refers only to architectural traps directly caused by a local counter overflow. If a local-counter-overflow interrupt is recognized without a trap, FROZEN is not automatically set. For instance, no freeze occurs if the LCOFI is pended while interrupts are masked, and software recognizes the LCOFI (perhaps by reading stopi or sip) and clears sip.LCOFIP before the trap is raised. As a result, some or all CTR history may be overwritten while handling the LCOFI. Such cases are expected to be very rare; for most usages (e.g., application profiling) privilege mode filtering is sufficient to ensure that CTR updates are inhibited while interrupts are handled in a more privileged mode._ | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Similarly, on a breakpoint exception that traps to M-mode or S-mode with `sctrctl`.BPFRZ=1, FROZEN is set by hardware. The breakpoint exception itself is not recorded. | | _Breakpoint exception refers to synchronous exceptions with a cause value of Breakpoint (3), regardless of source (ebreak, c.ebreak, Sdtrig); it does not include entry into Debug Mode, even in cores where this is implemented as an exception._ | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If the H extension is implemented, freeze behavior for LCOFIs and breakpoint exceptions that trap to VS-mode is determined by the LCOFIFRZ and BPFRZ values, respectively, in `vsctrctl`. This includes virtual LCOFIs pended by a hypervisor. | | _When a guest uses the SBI Supervisor Software Events (SSE) extension, the LCOFI will trap to HS-mode, which will then invoke a registered VS-mode LCOFI handler routine. If vsctrctl.LCOFIFRZ=1, the HS-mode handler will need to emulate the freeze by setting sctrstatus.FROZEN=1 before invoking the registered handler routine._ | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#custom-extensions)11.1.6\. Custom Extensions Any custom CTR extension must be associated with a non-zero value within the designated custom bits in `_x_ctrctl`. When the custom bits hold a non-zero value that enables a custom extension, the extension may alter standard CTR behavior, and may define new custom status fields within `sctrstatus` or the CTR [Entry Registers](#entry-registers). All custom status fields, and standard status fields whose behavior is altered by the custom extension, must revert to standard behavior when the custom bits hold zero. This includes read-only 0 behavior for any bits undefined by any implemented standard extensions. 10.1. "Smdbltrp" Double Trap Extension, Version 1.0 ==================== ## [](#smdbltrp)10.1\. "Smdbltrp" Double Trap Extension, Version 1.0 The Smdbltrp extension addresses a double trap (See [Double Trap Control in mstatus Register](machine.html#machine-double-trap)) in M-mode. When the Smrnmi extension (["Smrnmi" Extension for Resumable Non-Maskable Interrupts](rnmi.html#rnmi)) is implemented, it enables invocation of the RNMI handler on a double trap in M-mode to handle the critical error. If the Smrnmi extension is not implemented or if a double trap occurs during the RNMI handler’s execution, this extension helps transition the hart to a critical error state and enables signaling the critical error to the platform. To improve error diagnosis and resolution, this extension supports debugging harts in a critical error state. The extension introduces a mechanism to enter Debug Mode instead of asserting a critical-error signal to the platform when the hart is in a critical error state. See \[[91](../biblio/bibliography.html#bib-debug%5Fspec)\] for details. See [Double Trap Control in mstatus Register](machine.html#machine-double-trap) for the operational details. 6.1. "Smepmp" Extension for PMP Enhancements for memory access and execution prevention in Machine mode, Version 1.0 ==================== ## [](#smepmp)6.1\. "Smepmp" Extension for PMP Enhancements for memory access and execution prevention in Machine mode, Version 1.0 Being able to access the memory of a process running at a high privileged execution mode, such as the Supervisor or Machine mode, from a lower privileged mode such as the User mode, introduces an obvious attack vector since it allows for an attacker to perform privilege escalation, and tamper with the code and/or data of that process. A less obvious attack vector exists when the reverse happens, in which case an attacker instead of tampering with code and/or data that belong to a high-privileged process, can tamper with the memory of an unprivileged / less-privileged process and trick the high-privileged process to use or execute it. Two mechanisms combine to prevent this attack vector. The first one prevents the OS from accessing the memory of an unprivileged process unless a specific code path is followed, and the second one prevents the OS from executing the memory of an unprivileged process at all times. RISC-V already includes support for the former through the `sstatus.SUM` bit, and for the latter by always denying supervisor execution of virtual memory pages marked with the U bit. | | Terms: **PMP Entry**: A pair of pmpcfg\[i\] / pmpaddr\[i\] registers. **PMP Rule**: The contents of a pmpcfg register and its associated pmpaddr register(s), that encode a valid protected physical memory region, where pmpcfg\[i\].A != OFF, and if pmpcfg\[i\].A == TOR, pmpaddr\[i-1\] < pmpaddr\[i\]. **Ignored**: Any permissions set by a matching PMP rule are ignored, and _all_ accesses to the requested address range are allowed. **Enforced**: Only access types configured in the PMP rule matching the requested address range are allowed; failures will cause an access-fault exception. **Denied**: Any permissions set by a matching PMP rule are ignored, and _no_ accesses to the requested address range are allowed.; failures will cause an access-fault exception. **Locked**: A PMP rule/entry where the pmpcfg.L bit is set. **PMP reset**: A reset process where all PMP settings of the hart, including locked rules/settings, are re-initialized to a set of safe defaults, before releasing the hart (back) to the firmware / OS / application. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#smepmp%5Fthreat)6.1.1\. Threat model The rationale that guided development of this extension is included in Section ["Smepmp" Extension for PMP Enhancements for memory access and execution prevention in Machine mode](priv-rationale.html#smepmp%5Frationale). Without the Smepmp extension, it is not possible for a PMP rule to be **enforced** only on non-Machine modes and **denied** on Machine mode, in order to allow access to a memory region solely by less-privileged modes. It is only possible to have a **locked** rule that will be **enforced** on all modes, or a rule that will be **enforced** on non-Machine modes and be **ignored** by Machine mode. So for any physical memory region which is not protected with a Locked rule, Machine mode has unlimited access, including the ability to execute it. Without being able to protect less-privileged modes from Machine mode, it is not possible to prevent the mentioned attack vector. This becomes even more important for RISC-V than on other architectures, since implementations are allowed where a hart only has Machine and User modes available, so the whole OS will run on Machine mode instead of the non-existent Supervisor mode. In such implementations the attack surface is greatly increased, and the same kind of attacks performed on Supervisor mode and mitigated through the virtual-memory system, can be performed on Machine mode without any available mitigations. Even on implementations with Supervisor mode present attacks are still possible against the Firmware and/or the Secure Monitor running on Machine mode. ### [](#6-1-2-smepmp-physical-memory-protection-rules)6.1.2\. Smepmp Physical Memory Protection Rules To address the threat model outlined in Section [6.1.1\. Threat model](#smepmp%5Fthreat), this extension introduces the `RLB`, `MMWP`, and `MML` fields in the `mseccfg` CSR and their associated rules. See [Machine security configuration (mseccfg) register](machine.html#norm:mseccfg%5Fenc%5Fimg) for the detailed specification of these fields and the corresponding rules. The physical memory protection rules when `mseccfg.MML` is set to 1 are summarized in the truth table below. | Bits on _pmpcfg_ register | Result | | | | | | ------------------------- | ------ | - | - | ------------------------------------------------------------------------------ | ------------------------- | | L | R | W | X | M Mode | S/U Mode | | 0 | 0 | 0 | 0 | Inaccessible region (Access Exception) | | | 0 | 0 | 0 | 1 | Access Exception | Execute-only region | | 0 | 0 | 1 | 0 | Shared data region: Read/write on M mode, read-only on S/U mode | | | 0 | 0 | 1 | 1 | Shared data region: Read/write for both M and S/U mode | | | 0 | 1 | 0 | 0 | Access Exception | Read-only region | | 0 | 1 | 0 | 1 | Access Exception | Read/Execute region | | 0 | 1 | 1 | 0 | Access Exception | Read/Write region | | 0 | 1 | 1 | 1 | Access Exception | Read/Write/Execute region | | 1 | 0 | 0 | 0 | Locked inaccessible region\* (Access Exception) | | | 1 | 0 | 0 | 1 | Locked Execute-only region\* | Access Exception | | 1 | 0 | 1 | 0 | Locked Shared code region: Execute only on both M and S/U mode.\* | | | 1 | 0 | 1 | 1 | Locked Shared code region: Execute only on S/U mode, read/execute on M mode.\* | | | 1 | 1 | 0 | 0 | Locked Read-only region\* | Access Exception | | 1 | 1 | 0 | 1 | Locked Read/Execute region\* | Access Exception | | 1 | 1 | 1 | 0 | Locked Read/Write region\* | Access Exception | | 1 | 1 | 1 | 1 | Locked Shared data region: Read only on both M and S/U mode.\* | | **: \*Locked** rules cannot be removed or modified until a **PMP reset**, unless `mseccfg.RLB` is set. A visual representation of these rules is as follows: ![smepmp visual representation](_images/smepmp-visual-representation.png) ### [](#6-1-3-smepmp-software-discovery)6.1.3\. Smepmp software discovery Since all fields defined in `mseccfg` as part of this extension are locked when set (`MMWP`/`MML`) or locked when cleared (`RLB`), software can’t poll them for determining the presence of Smepmp. It is expected that BootROM will set `mseccfg.MMWP` and/or `mseccfg.MML` during early boot, before jumping to the firmware, so that the firmware will be able to determine the presence of Smepmp by reading `mseccfg` and checking the state of `mseccfg.MMWP` and `mseccfg.MML`. 4.1. "Smstateen/Ssstateen" Extensions, Version 1.0 ==================== ## [](#smstateen)4.1\. "Smstateen/Ssstateen" Extensions, Version 1.0 The implementation of optional RISC-V extensions has the potential to open covert channels between separate user threads, or between separate guest OSes running under a hypervisor. The problem occurs when an extension adds processor state — usually explicit registers, but possibly other forms of state — that the main OS or hypervisor is unaware of (and hence won’t context-switch) but that can be modified/written by one user thread or guest OS and perceived/examined/read by another. For example, the Advanced Interrupt Architecture (AIA) for RISC-V adds to a hart as many as ten supervisor-level CSRs (`siselect`, `sireg`, `stopi`,`sseteipnum`, `sclreipnum`, `sseteienum`, `sclreienum`, `sclaimei`, `sieh`, and `siph`) and provides also the option for hardware to be backward-compatible with older, pre-AIA software. Because an older hypervisor that is oblivious to the AIA will not know to swap any of the AIA’s new CSRs on context switches, the registers may then be used as a covert channel between multiple guest OSes that run atop this hypervisor. Although traditional practices might consider such a communication channel harmless, the intense focus on security today argues that a means be offered to plug such channels. The `f` registers of the RISC-V floating-point extensions and the `v` registers of the vector extension would similarly be potential covert channels between user threads, except for the existence of the FS and VS fields in the `sstatus`register. Even if an OS is unaware of, say, the vector extension and its `v`registers, access to those registers is blocked when the VS field is initialized to zero, either at machine level or by the OS itself initializing`sstatus`. Obviously, one way to prevent the use of new user-level CSRs as covert channels would be to add to `mstatus` or `sstatus` an "XS" field for each relevant extension, paralleling the V extension’s VS field. However, this is not considered a general solution to the problem due to the number of potential future extensions that may add small amounts of state. Even with a 64-bit`sstatus` (necessitating adding `sstatush` for RV32), it is not certain there are enough remaining bits in `sstatus` to accommodate all future user-level extensions. In any event, there is no need to strain `sstatus` (and add `sstatush`) for this purpose. The "enable" flags that are needed to plug covert channels are not generally expected to require swapping on context switches of user threads, making them a less-than-compelling candidate for inclusion in `sstatus`. Hence, a new place is provided for them instead. ### [](#4-1-1-state-enable-extensions)4.1.1\. State Enable Extensions The Smstateen and Ssstateen extensions collectively specify machine-mode and supervisor-mode features. The Smstateen extension specification comprises the mstateen\*, sstateen\*, and hstateen\* CSRs and their functionality. The Ssstateen extension specification comprises only the sstateen\* and hstateen\* CSRs and their functionality. For RV64 harts, this extension adds four new 64-bit CSRs at machine level:`mstateen0` (Machine State Enable 0), `mstateen1`, `mstateen2`, and `mstateen3`. If supervisor mode is implemented, another four CSRs are defined at supervisor level:`sstateen0`, `sstateen1`, `sstateen2`, and `sstateen3`. And if the hypervisor extension is implemented, another set of CSRs is added:`hstateen0`, `hstateen1`, `hstateen2`, and `hstateen3`. For RV32, there are CSR addresses for accessing the upper 32 bits of corresponding machine-level and hypervisor CSRs:`mstateen0h`, `mstateen1h`, `mstateen2h`, `mstateen3h`,`hstateen0h`, `hstateen1h`, `hstateen2h`, and `hstateen3h`. For the supervisor-level `sstateen` registers, high-half CSRs are not added at this time because it is expected the upper 32 bits of these registers will always be zeros, as explained later below. Each bit of a `stateen` CSR controls less-privileged access to an extension’s state, for an extension that was not deemed "worthy" of a full XS field in`sstatus` like the FS and VS fields for the F and V extensions. The number of registers provided at each level is four because it is believed that 4 \* 64 = 256 bits for machine and hypervisor levels, and 4 \* 32 = 128 bits for supervisor level, will be adequate for many years to come, perhaps for as long as the RISC-V ISA is in use. The exact number four is an attempted compromise between providing too few bits on the one hand and going overboard with CSRs that will never be used on the other. A possible future doubling of the number of `stateen` CSRs is covered later. The `stateen` registers at each level control access to state at all less-privileged levels, but not at its own level. This is analogous to how the existing `counteren` CSRs control access to performance counter registers. Just as with the `counteren` CSRs, when a `stateen` CSR prevents access to state by less-privileged levels, an attempt in one of those privilege modes to execute an instruction that would read or write the protected state raises an illegal-instruction exception, or, if executing in VS or VU mode and the circumstances for a virtual-instruction exception apply, raises a virtual-instruction exception instead of an illegal-instruction exception. When this extension is not implemented, all state added by an extension is accessible as defined by that extension. When a `stateen` CSR prevents access to state for a privilege mode, attempting to execute in that privilege mode an instruction that _implicitly_ updates the state without reading it may or may not raise an illegal-instruction or virtual-instruction exception. Such cases must be disambiguated by being explicitly specified one way or the other. In some cases, the bits of the `stateen` CSRs will have a dual purpose as enables for the ISA extensions that introduce the controlled state. Each bit of a supervisor-level `sstateen` CSR controls user-level access (from U-mode or VU-mode) to an extension’s state. The intention is to allocate the bits of `sstateen` CSRs starting at the least-significant end, bit 0, through to bit 31, and then on to the next-higher-numbered `sstateen` CSR. For every bit with a defined purpose in an `sstateen` CSR, the same bit is defined in the matching `mstateen` CSR to control access below machine level to the same state. The upper 32 bits of an `mstateen` CSR (or for RV32, the corresponding high-half CSR) control access to state that is inherently inaccessible to user level, so no corresponding enable bits in the supervisor-level `sstateen` CSR are applicable. The intention is to allocate bits for this purpose starting at the most-significant end, bit 63, through to bit 32, and then on to the next-higher `mstateen` CSR. If the rate that bits are being allocated from the least-significant end for `sstateen` CSRs is sufficiently low, allocation from the most-significant end of `mstateen` CSRs may be allowed to encroach on the lower 32 bits before jumping to the next-higher`mstateen` CSR. In that case, the bit positions of "encroaching" bits will remain forever read-only zeros in the matching `sstateen` CSRs. With the hypervisor extension, the `hstateen` CSRs have identical encodings to the `mstateen` CSRs, except controlling accesses for a virtual machine (from VS and VU modes). Each standard-defined bit of a `stateen` CSR is WARL and may be read-only zero or one, subject to the following conditions. Bits in any `stateen` CSR that are defined to control state that a hart doesn’t implement are read-only zeros for that hart. Likewise, all reserved bits not yet given a defined meaning are also read-only zeros. For every bit in an`mstateen` CSR that is zero (whether read-only zero or set to zero), the same bit appears as read-only zero in the matching `hstateen` and `sstateen` CSRs. For every bit in an `hstateen` CSR that is zero (whether read-only zero or set to zero), the same bit appears as read-only zero in `sstateen` when accessed in VS-mode. A bit in a supervisor-level `sstateen` CSR cannot be read-only one unless the same bit is read-only one in the matching `mstateen` CSR and, if it exists, in the matching `hstateen` CSR. A bit in an `hstateen` CSR cannot be read-only one unless the same bit is read-only one in the matching `mstateen` CSR. On reset, all writable `mstateen` bits are initialized by the hardware to zeros. If machine-level software changes these values, it is responsible for initializing the corresponding writable bits of the `hstateen` and `sstateen` CSRs to zeros too. Software at each privilege level should set its respective`stateen` CSRs to indicate the state it is prepared to allow less-privileged software to access. For OSes and hypervisors, this usually means the state that the OS or hypervisor is prepared to swap on a context switch, or to manage in some other way. For each `mstateen` CSR, bit 63 is defined to control access to the matching `sstateen` and `hstateen` CSRs. That is, bit 63 of `mstateen0` controls access to `sstateen0` and `hstateen0`; bit 63 of `mstateen1` controls access to`sstateen1` and `hstateen1`; etc. Likewise, bit 63 of each `hstateen`correspondingly controls access to the matching `sstateen` CSR. A hypervisor may need this control over accesses to the `sstateen` CSRs if it ever must emulate for a virtual machine an extension that is supposed to be affected by a bit in an `sstateen` CSR. Even if such emulation is uncommon, it should not be excluded. Machine-level software needs identical control to be able to emulate the hypervisor extension. That is, machine level needs control over accesses to the supervisor-level `sstateen` CSRs in order to emulate the `hstateen` CSRs, which have such control. Bit 63 of each `mstateen` CSR may be read-only zero only if the hypervisor extension is not implemented and the matching supervisor-level `sstateen` CSR is all read-only zeros. In that case, machine-level software should emulate attempts to access the affected `sstateen` CSR from S-mode, ignoring writes and returning zero for reads. Bit 63 of each `hstateen` CSR is always writable (not read-only). ### [](#4-1-2-state-enable-0-registers)4.1.2\. State Enable 0 Registers ![Machine State Enable 0 Register (`mstateen0`)](_images/svg-0b49af3729b0a7f2f020a30ba8c85c396cbb4e88.svg) Figure 1\. Machine State Enable 0 Register (`mstateen0`) ![Hypervisor State Enable 0 Register (`hstateen0`)](_images/svg-15c379a7e27825d306343aa3b47e2c16b7d1d5e7.svg) Figure 2\. Hypervisor State Enable 0 Register (`hstateen0`) ![Supervisor State Enable 0 Register (`sstateen0`)](_images/svg-1f7f310bac65284ddfa1d1cf0f0c6ad943006ec3.svg) Figure 3\. Supervisor State Enable 0 Register (`sstateen0`) The C bit controls access to any and all custom state.The C bit of these registers is not custom state itself; it is a standard field of a standard CSR, either `mstateen0`, `hstateen0`, or`sstateen0`. | | The requirements that non-standard extensions must meet to be conforming are not relaxed due solely to changes in the value of this bit. In particular, if software sets this bit but does not execute any custom instructions or access any custom state, the software must continue to execute as specified by all relevant RISC-V standards, or the hardware is not standard-conforming. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The FCSR bit controls access to `fcsr` for the case when floating-point instructions operate on `x` registers instead of `f` registers as specified by the Zfinx and related extensions (Zdinx, etc.). Whenever `misa.F` \= 1, FCSR bit of `mstateen0` is read-only zero (and hence read-only zero in `hstateen0` and`sstateen0` too). For convenience, when the `stateen` CSRs are implemented and`misa.F` \= 0, then if the FCSR bit of a controlling `stateen0` CSR is zero, all floating-point instructions cause an illegal-instruction exception (or virtual-instruction exception, if relevant), as though they all access `fcsr`, regardless of whether they really do. The JVT bit controls access to the `jvt` CSR provided by the Zcmt extension. The SE0 bit in `mstateen0` controls access to the `hstateen0`, `hstateen0h`, and the `sstateen0` CSRs. The SE0 bit in `hstateen0` controls access to the`sstateen0` CSR. The ENVCFG bit in `mstateen0` controls access to the `henvcfg`, `henvcfgh`, and the `senvcfg` CSRs. The ENVCFG bit in `hstateen0` controls access to the`senvcfg` CSRs. The CSRIND bit in `mstateen0` controls access to the `siselect`, `sireg*`,`vsiselect`, and the `vsireg*` CSRs provided by the Sscsrind extensions. The CSRIND bit in `hstateen0` controls access to the `siselect` and the`sireg*`, (really `vsiselect` and `vsireg*`) CSRs provided by the Sscsrind extensions. The IMSIC bit in `mstateen0` controls access to the IMSIC state, including CSRs `stopei` and `vstopei`, provided by the Ssaia extension. The IMSIC bit in`hstateen0` controls access to the guest IMSIC state, including CSRs `stopei`(really `vstopei`), provided by the Ssaia extension. | | Setting the IMSIC bit in hstateen0 to zero prevents a virtual machine from accessing the hart’s IMSIC the same as setting hstatus.VGEIN = 0. | | ----------------------------------------------------------------------------------------------------------------------------------------------- | The AIA bit in `mstateen0` controls access to all state introduced by the Ssaia extension and not controlled by either the CSRIND or the IMSIC bits. The AIA bit in `hstateen0` controls access to all state introduced by the Ssaia extension and not controlled by either the CSRIND or the IMSIC bits of `hstateen0`. The CONTEXT bit in `mstateen0` controls access to the `scontext` and`hcontext` CSRs provided by the Sdtrig extension. The CONTEXT bit in`hstateen0` controls access to the `scontext` CSR provided by the Sdtrig extension. The P1P13 bit in `mstateen0` controls access to the `hedelegh` introduced by Privileged Specification Version 1.13. The SRMCFG bit in `mstateen0` controls access to the `srmcfg` CSR introduced by the [Ssqosid](supervisor.html#ssqosid) extension. ### [](#4-1-3-usage)4.1.3\. Usage After the writable bits of the machine-level `mstateen` CSRs are initialized to zeros on reset, machine-level software can set bits in these registers to enable less-privileged access to the controlled state. This may be either because machine-level software knows how to swap the state or, more likely, because machine-level software isn’t swapping supervisor-level environments. (Recall that the main reason the `mstateen` CSRs must exist is so machine level can emulate the hypervisor extension. When machine level isn’t emulating the hypervisor extension, it is likely there will be no need to keep any implemented `mstateen` bits zero.) If machine level sets any writable `mstateen` bits to nonzero, it must initialize the matching `hstateen` CSRs, if they exist, by writing zeros to them. And if any`mstateen` bits that are set to one have matching bits in the `sstateen` CSRs, machine-level software must also initialize those `sstateen` CSRs by writing zeros to them. Ordinarily, machine-level software will want to set bit 63 of all `mstateen` CSRs, necessitating that it write zero to all `hstateen` CSRs. Software should ensure that all writable bits of `sstateen` CSRs are initialized to zeros when an OS at supervisor level is first entered. The OS can then set bits in these registers to enable user-level access to the controlled state, presumably because it knows how to context-swap the state. For the `sstateen` CSRs whose access by a guest OS is permitted by bit 63 of the corresponding `hstateen` CSRs, a hypervisor must include the `sstateen` CSRs in the context it swaps for a guest OS. When it starts a new guest OS, it must ensure the writable bits of those `sstateen` CSRs are initialized to zeros, and it must emulate accesses to any other `sstateen` CSRs. If software at any privilege level does not support multiple contexts for less-privilege levels, then it may choose to maximize less-privileged access to all state by writing a value of all ones to the `stateen` CSRs at its level (the`mstateen` CSRs for machine level, the `sstateen` CSRs for an OS, and the `hstateen`CSRs for a hypervisor), without knowing all the state to which it is granting access. This is justified because there is no risk of a covert channel between execution contexts at the less-privileged level when only one context exists at that level. This situation is expected to be common for machine level, and it might also arise, for example, for a type-1 hypervisor that hosts only a single guest virtual machine. | | If a need is anticipated, the set of stateen CSRs could in the future be doubled by adding these: 0x38C mstateen4, 0x39C mstateen4h 0x38D mstateen5, 0x39D mstateen5h 0x38E mstateen6, 0x39E mstateen6h 0x38F mstateen7, 0x39F mstateen7h 0x18C sstateen4 0x18D sstateen5 0x18E sstateen6 0x18F sstateen7 0x68C hstateen4, 0x69C hstateen4h 0x68D hstateen5, 0x69D hstateen5h 0x68E hstateen6, 0x69E hstateen6h 0x68F hstateen7, 0x69F hstateen7h These additional CSRs are not a definite part of the original proposal because it is unclear whether they will ever be needed, and it is believed the rate of consumption of bits in the first group, registers numbered 0-3, will be slow enough that any looming shortage will be perceptible many years in advance. At the moment, it is not known even how many years it may take to exhaust justmstateen0, sstateen0, and hstateen0. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 14.1. "Sscofpmf" Extension for Count Overflow and Mode-Based Filtering, Version 1.0 ==================== ## [](#Sscofpmf)14.1\. "Sscofpmf" Extension for Count Overflow and Mode-Based Filtering, Version 1.0 The current Privileged specification defines mhpmevent CSRs to select and control event counting by the associated hpmcounter CSRs, but provides no standardization of any fields within these CSRs. For at least Linux-class rich-OS systems it is desirable to standardize certain basic features that are broadly desired (and have come up over the past year plus on RISC-V lists, as well as have been the subject of past proposals). This enables there to be standard upstream software support that eliminates the need for implementations to provide their own custom software support. This extension serves to accomplish exactly this within the existing mhpmevent CSRs (and correspondingly avoids the unnecessary creation of whole new sets of CSRs - past just one new CSR). This extension sticks to addressing two basic well-understood needs that have been requested by various people. To make it easy to understand the deltas from the current Priv 1.11/1.12 specs, this is written as the actual exact changes to be made to existing paragraphs of Priv spec text (or additional paragraphs within the existing text). The extension name is "Sscofpmf" ('Ss' for Privileged arch and Supervisor-level extensions, and 'cofpmf' for Count OverFlow and Privilege Mode Filtering). Note that the new count overflow interrupt will be treated as a standard local interrupt that is assigned to bit 13 in the mip/mie/sip/sie registers. ### [](#14-1-1-count-overflow-control)14.1.1\. Count Overflow Control The following bits are added to `mhpmevent`: | 63 | 62 | 61 | 60 | 59 | 58 | 57 | 56 | | -- | ---- | ---- | ---- | ----- | ----- | ------ | ------ | | OF | MINH | SINH | UINH | VSINH | VUINH | _WPRI_ | _WPRI_ | | Field | Description | | ------ | ---------------------------------------------------------------------------- | | OF | Overflow status and interrupt disable bit that is set when counter overflows | | MINH | If set, then counting of events in M-mode is inhibited | | SINH | If set, then counting of events in S/HS-mode is inhibited | | UINH | If set, then counting of events in U-mode is inhibited | | VSINH | If set, then counting of events in VS-mode is inhibited | | VUINH | If set, then counting of events in VU-mode is inhibited | | _WPRI_ | Reserved | | _WPRI_ | Reserved | For each `x`INH bit, if the associated privilege mode is not implemented, the bit is read-only zero. Each of the five `x`INH bits, when set, inhibit counting of events while in privilege mode `x`. All-zeroes for these bits results in counting of events in all modes. The OF bit is set when the corresponding hpmcounter overflows, and remains set until written by software. Since hpmcounter values are unsigned values, overflow is defined as unsigned overflow of the implemented counter bits. Note that there is no loss of information after an overflow since the counter wraps around and keeps counting while the sticky OF bit remains set. If supervisor mode is implemented, the 32-bit scountovf register contains read-only shadow copies of the OF bits in all 29 mhpmevent registers. If an hpmcounter overflows while the associated OF bit is zero, then a "count overflow interrupt request" is generated. If the OF bit is one, then no interrupt request is generated. Consequently the OF bit also functions as a count overflow interrupt disable for the associated hpmcounter. Count overflow never results from writes to the mhpmcounter_n_ or mhpmevent_n_ registers, only from hardware increments of counter registers. This count-overflow-interrupt-request signal is treated as a standard local interrupt that corresponds to bit 13 in the `mip`/`mie`/`sip`/`sie` registers. The `mip`/`sip` LCOFIP and `mie`/`sie` LCOFIE bits are, respectively, the interrupt-pending and interrupt-enable bits for this interrupt. ('LCOFI' represents 'Local Count Overflow Interrupt'.) Generation of a count-overflow-interrupt request by an `hpmcounter` sets the associated OF bit. When an OF bit is set, it eventually, but not necessarily immediately, sets the LCOFIP bit in the `mip`/`sip` registers.The LCOFIP bit is cleared by software before servicing the count overflow interrupt resulting from one or more count overflows. The `mideleg` register controls the delegation of this interrupt to S-mode versus M-mode.# | | There are not separate overflow status and overflow interrupt enable bits. In practice, enabling overflow interrupt generation (by clearing the OF bit) is done in conjunction with initializing the counter to a starting value. Once a counter has overflowed, it and the OF bit must be reinitialized before another overflow interrupt can be generated. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Software can distinguish newly overflowed counters (yet to be serviced by an overflow interrupt handler) from overflowed counters that have already been serviced or that are configured to not generate an interrupt on overflow, by maintaining a bit mask reflecting which counters are active and due to eventually overflow. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#14-1-2-supervisor-count-overflow-scountovf-register)14.1.2\. Supervisor Count Overflow (`scountovf`) Register This extension adds the `scountovf` CSR, a 32-bit read-only register that contains shadow copies of the OF bits in the 29 mhpmevent CSRs (mhpmevent_3_ \- mhpmevent_31_) - where scountovf bit _X_ corresponds to mhpmevent_X_. This register enables supervisor-level overflow interrupt handler software to quickly and easily determine which counter(s) have overflowed (without needing to make an execution environment call or series of calls ultimately up to M-mode). Read access to bit _X_ is subject to the same mcounteren (or mcounteren and hcounteren) CSRs that mediate access to the hpmcounter CSRs by S-mode (or VS-mode). In M-mode, scountovf bit _X_ is always readable. In S/HS-mode, scountovf bit _X_ is readable when mcounteren bit_X_ is set, and otherwise reads as zero. Similarly, in VS mode, scountovf bit_X_ is readable when mcounteren bit _X_ and hcounteren bit _X_ are both set, and otherwise reads as zero. 17.1. "Ssdbltrp" Double Trap Extension, Version 1.0 ==================== ## [](#ssdbltrp)17.1\. "Ssdbltrp" Double Trap Extension, Version 1.0 The Ssdbltrp extension addresses a double trap (See [Double Trap Control in mstatus Register](machine.html#machine-double-trap)) privilege modes lower than M. It enables HS-mode to invoke a critical error handler in a virtual machine on a double trap in VS-mode. It also allows M-mode to invoke a critical error handler in the OS/Hypervisor on a double trap in S/HS-mode. The Ssdbltrp extension adds the `menvcfg`.DTE (See [Machine Environment Configuration (menvcfg) Register](machine.html#sec:menvcfg)) and the`sstatus`.SDT fields (See [Supervisor Status (sstatus) Register](supervisor.html#sstatus)). If the hypervisor extension is additionally implemented, then the extension adds the `henvcfg`.DTE (See[Hypervisor Environment Configuration Register (henvcfg)](hypervisor.html#sec:henvcfg)) and the `vsstatus`.SDT fields (See [Virtual Supervisor Status (vsstatus) Register](hypervisor.html#vsstatus)). See [Double Trap Control in sstatus Register](supervisor.html#supv-double-trap) for the operational details. 13.1. "Sstc" Extension for Supervisor-mode Timer Interrupts, Version 1.0 ==================== ## [](#Sstc)13.1\. "Sstc" Extension for Supervisor-mode Timer Interrupts, Version 1.0 The current Privileged arch specification only defines a hardware mechanism for generating machine-mode timer interrupts (based on the mtime and mtimecmp registers). With the resultant requirement that timer services for S-mode/HS-mode (and for VS-mode) have to all be provided by M-mode - via SBI calls from S/HS-mode up to M-mode (or VS-mode calls to HS-mode and then to M-mode). M-mode software then multiplexes these multiple logical timers onto its one physical M-mode timer facility, and the M-mode timer interrupt handler passes timer interrupts back down to the appropriate lower privilege mode. This extension serves to provide supervisor mode with its own CSR-based timer interrupt facility that it can directly manage to provide its own timer service (in the form of having its own `stimecmp` register) - thus eliminating the large overheads for emulating S/HS-mode timers and timer interrupt generation up in M-mode. Further, this extension adds a similar facility to the Hypervisor extension for VS-mode. The extension name is "Sstc" ('Ss' for Privileged arch and Supervisor-level extensions, and 'tc' for timecmp). This extension adds the S-level `stimecmp`CSR ([Supervisor Timer (stimecmp) Register](supervisor.html#stimecmp)) and the VS-level `vstimecmp` CSR ([Virtual Supervisor Timer (vstimecmp) Register](hypervisor.html#vstimecmp). This extension adds the `STCE` bit to the `menvcfg` ([Machine Environment Configuration (menvcfg) Register](machine.html#sec:menvcfg)) and `henvcfg`([Hypervisor Environment Configuration Register (henvcfg)](hypervisor.html#sec:henvcfg)) CSRs. 12.1. Supervisor-Level ISA, Version 1.13 ==================== ## [](#supervisor)12.1\. Supervisor-Level ISA, Version 1.13 This chapter describes the RISC-V supervisor-level architecture, which contains a common core that is used with various supervisor-level address translation and protection schemes. | | Supervisor mode is deliberately restricted in terms of interactions with underlying physical hardware, such as physical memory and device interrupts, to support clean virtualization. In this spirit, certain supervisor-level facilities, including requests for timer and interprocessor interrupts, are provided by implementation-specific mechanisms. In some systems, a supervisor execution environment (SEE) provides these facilities in a manner specified by a supervisor binary interface (SBI). Other systems supply these facilities directly, through some other implementation-defined mechanism. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#12-1-1-supervisor-csrs)12.1.1\. Supervisor CSRs A number of CSRs are provided for the supervisor. | | The supervisor should only view CSR state that should be visible to a supervisor-level operating system. In particular, there is no information about the existence (or non-existence) of higher privilege levels (machine level or other) visible in the CSRs accessible by the supervisor. Many supervisor CSRs are a subset of the equivalent machine-mode CSR, and the machine-mode chapter should be read first to help understand the supervisor-level CSR descriptions. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#sstatus)12.1.1.1\. Supervisor Status (`sstatus`) Register The `sstatus` register is an SXLEN-bit read/write register formatted as shown in [Figure 1](#sstatusreg-rv32) when SXLEN=32 and [Figure 2](#sstatusreg) when SXLEN=64\. The `sstatus`register keeps track of the processor’s current operating state. ![Supervisor-mode status (`sstatus`) register when SXLEN=32.](_images/svg-b6248fe153c498faef45742cc9949e46c7f7bd33.svg) Figure 1\. Supervisor-mode status (`sstatus`) register when SXLEN=32. ![Supervisor-mode status (`sstatus`) register when SXLEN=64.](_images/svg-f96a49ba3c356808d7a18da0f227529f7dde5569.svg) Figure 2\. Supervisor-mode status (`sstatus`) register when SXLEN=64. The SPP bit indicates the privilege level at which a hart was executing before entering supervisor mode. When a trap is taken, SPP is set to 0 if the trap originated from user mode, or 1 otherwise. When an SRET instruction (see [Other Privileged Instructions](machine.html#otherpriv)) is executed to return from the trap handler, the privilege level is set to user mode if the SPP bit is 0, or supervisor mode if the SPP bit is 1; SPP is then set to 0. The SIE bit enables or disables all interrupts in supervisor mode. When SIE is clear, interrupts are not taken while in supervisor mode. When the hart is running in user-mode, the value in SIE is ignored, and supervisor-level interrupts are enabled. The supervisor can disable individual interrupt sources using the `sie` CSR. The SPIE bit indicates whether supervisor interrupts were enabled prior to trapping into supervisor mode. When a trap is taken into supervisor mode, SPIE is set to SIE, and SIE is set to 0\. When an SRET instruction is executed, SIE is set to SPIE, then SPIE is set to 1. The `sstatus` register is a subset of the `mstatus` register. | | In a straightforward implementation, reading or writing any field insstatus is equivalent to reading or writing the homonymous field inmstatus. | | -------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#12-1-1-1-1-base-isa-control-in-sstatus-register)12.1.1.1.1\. Base ISA Control in `sstatus` Register The UXL field controls the value of XLEN for U-mode, termed _UXLEN_, which may differ from the value of XLEN for S-mode, termed _SXLEN_. The encoding of UXL is the same as that of the MXL field of `misa`, shown in[Encoding of UXL field in misa](machine.html#norm:misa%5Fmxl%5Fenc). When SXLEN=32, the UXL field does not exist, and UXLEN=32\. When SXLEN=64, it is a **WARL** field that encodes the current value of UXLEN. In particular, an implementation may make UXL be a read-only field whose value always ensures that UXLEN=SXLEN. If UXLEN≠SXLEN, instructions executed in the narrower mode must ignore source register operand bits above the configured XLEN, and must sign-extend results to fill the widest supported XLEN in the destination register. If UXLEN < SXLEN, user-mode instruction-fetch addresses and load and store effective addresses are taken modulo 2UXLEN. For example, when UXLEN=32 and SXLEN=64, user-mode memory accesses reference the lowest 4 GiB of the address space. Some HINT instructions are encoded as integer computational instructions that overwrite their destination register with its current value, e.g.,`c.addi x8, 0`. When such a HINT is executed with XLEN < SXLEN and bits SXLEN..XLEN of the destination register not all equal to bit XLEN-1, it is implementation-defined whether bits SXLEN..XLEN of the destination register are unchanged or are overwritten with copies of bit XLEN-1. | | This definition allows implementations to elide register write-back for some HINTs, while allowing them to execute other HINTs in the same manner as other integer computational instructions. The implementation choice is observable only by S-mode with SXLEN > UXLEN; it is invisible to U-mode. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#sum)12.1.1.1.2\. Memory Privilege in `sstatus` Register The MXR (Make eXecutable Readable) bit modifies the privilege with which loads access virtual memory. When MXR=0, only loads from pages marked readable (R=1 in [Figure 19](#sv32pte)) will succeed. When MXR=1, loads from pages marked either readable or executable (R=1 or X=1) will succeed. MXR has no effect when page-based virtual memory is not in effect. The SUM (permit Supervisor User Memory access) bit modifies the privilege with which S-mode loads and stores access virtual memory. When SUM=0, S-mode memory accesses to pages that are accessible by U-mode (U=1 in [Figure 19](#sv32pte)) will fault. When SUM=1, these accesses are permitted. SUM has no effect when page-based virtual memory is not in effect, nor when executing in U-mode. Note that S-mode can never execute instructions from user pages, regardless of the state of SUM. SUM is read-only 0 if `satp`.MODE is read-only 0. | | The SUM mechanism prevents supervisor software from inadvertently accessing user memory. Operating systems can execute the majority of code with SUM clear; the few code segments that should access user memory can temporarily set SUM. The SUM mechanism does not avail S-mode software of permission to execute instructions in user code pages. Legitimate use cases for execution from user memory in supervisor context are rare in general and nonexistent in POSIX environments. However, bugs in supervisors that lead to arbitrary code execution are much easier to exploit if the supervisor exploit code can be stored in a user buffer at a virtual address chosen by an attacker. Some non-POSIX single address space operating systems do allow certain privileged software to partially execute in supervisor mode, while most programs run in user mode, all in a shared address space. This use case can be realized by mapping the physical code pages at multiple virtual addresses with different permissions, possibly with the assistance of the instruction page-fault handler to direct supervisor software to use the alternate mapping. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#12-1-1-1-3-endianness-control-in-sstatus-register)12.1.1.1.3\. Endianness Control in `sstatus` Register The UBE bit is a **WARL** field that controls the endianness of explicit memory accesses made from U-mode, which may differ from the endianness of memory accesses in S-mode. An implementation may make UBE be a read-only field that always specifies the same endianness as for S-mode. UBE controls whether explicit load and store memory accesses made from U-mode are little-endian (UBE=0) or big-endian (UBE=1). UBE has no effect on instruction fetches, which are _implicit_ memory accesses that are always little-endian. For _implicit_ accesses to supervisor-level memory management data structures, such as page tables, S-mode endianness always applies and UBE is ignored. | | Standard RISC-V ABIs are expected to be purely little-endian-only or big-endian-only, with no accommodation for mixing endianness. Nevertheless, endianness control has been defined so as to permit an OS of one endianness to execute user-mode programs of the opposite endianness. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#12-1-1-1-4-previous-expected-landing-pad-elp-state-in-sstatus-register)12.1.1.1.4\. Previous Expected Landing Pad (ELP) State in `sstatus` Register Access to the `SPELP` field, added by Zicfilp, accesses the homonymous fields of `mstatus` when `V=0`, and the homonymous fields of `vsstatus`when `V=1`. ##### [](#supv-double-trap)12.1.1.1.5\. Double Trap Control in `sstatus` Register The S-mode-disable-trap (`SDT`) bit is a WARL field introduced by the Ssdbltrp extension to address double trap (See [Double Trap Control in mstatus Register](machine.html#machine-double-trap)) at privilege modes lower than M. When the `SDT` bit is set to 1 by an explicit CSR write, the `SIE` (Supervisor Interrupt Enable) bit is cleared to 0\. This clearing occurs regardless of the value written, if any, to the `SIE` bit by the same write. The `SIE` bit can only be set to 1 by an explicit CSR write if the `SDT` bit is being set to 0 by the same write or is already 0. When a trap is to be taken into S-mode, if the `SDT` bit is currently 0, it is then set to 1, and the trap is delivered as expected. However, if `SDT` is already set to 1, then this is an _unexpected trap_. In the event of an_unexpected trap_, a double-trap exception trap is delivered into M-mode. To deliver this trap, the hart writes registers, except `mcause` and `mtval2`, with the same information that the _unexpected trap_ would have written if it was taken into M-mode. The `mtval2` register is then set to what would be otherwise written into the `mcause` register by the _unexpected trap_. The `mcause`register is set to 16, the double-trap exception code. An `SRET` instruction sets the `SDT` bit to 0. | | After a trap handler has saved the state, such as scause, sepc, and stval, needed for resuming from the trap and is reentrant, it should clear the SDT bit. Resetting the SDT by an SRET enables the trap handler to detect a double trap that may occur during the tail phase, where it restores critical state to return from a trap. The consequence of this specification is that if a critical error condition was caused by a guest-page fault, then the GPA will not be available in mtval2when the double trap is delivered to M-mode. This condition arises if the HS-mode invokes a hypervisor virtual-machine load or store instruction whenSDT is 1 and the instruction raises a guest-page fault. The use of such an instruction in this phase of trap handling is not common. However, not recording the GPA is considered benign because, if required, it can still be obtained — albeit with added effort — through the process of walking the page tables. For a double trap that originates in VS-mode, M-mode should redirect the exception to HS-mode by copying the values of M-mode CSRs updated by the trap to HS-mode CSRs and should use an MRET to resume execution at the address in stvec. Supervisor Software Events (SSE), an extension to the SBI, provide a mechanism for supervisor software to register and service system events emanating from an SBI implementation, such as firmware or a hypervisor. In the event of a double trap, HS-mode and M-mode can utilize the SSE mechanism to invoke a critical-error handler in VS-mode or S/HS-mode, respectively. Additionally, the implementation of an SSE protocol can be considered as an optional measure to aid in the recovery from such critical errors. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#12-1-1-2-supervisor-trap-vector-base-address-stvec-register)12.1.1.2\. Supervisor Trap Vector Base Address (`stvec`) Register The `stvec` register is an SXLEN-bit read/write register that holds trap vector configuration, consisting of a vector base address (BASE) and a vector mode (MODE). ![Supervisor trap vector base address (`stvec`) register.](_images/diag-efe1113b8cf503c82d16fdf049c4d92c5c01cdb6.svg) Figure 3\. Supervisor trap vector base address (`stvec`) register. The BASE field in `stvec` is a field that can hold any valid virtual or physical address, subject to the following alignment constraints: the address must be 4-byte aligned, and MODE settings other than Direct might impose additional alignment constraints on the value in the BASE field. Note that the CSR contains only bits XLEN-1 through 2 of the address BASE. When used as an address, the lower two bits are filled with zeroes to obtain an XLEN-bit address that is always aligned on a 4-byte boundary. __Table 1\. Encoding of stvec MODE field.__ | Value | Name | Description | | ----- | -------------- | ---------------------------------------------------------------------------------------- | | 01≥2 | DirectVectored | All exceptions set pc to BASE.Asynchronous interrupts set pc to BASE+4×cause. _Reserved_ | The encoding of the MODE field is shown in[Table 1](#stvec-mode). When MODE=Direct, all traps into supervisor mode cause the `pc` to be set to the address in the BASE field. When MODE=Vectored, all synchronous exceptions into supervisor mode cause the `pc` to be set to the address in the BASE field, whereas interrupts cause the `pc` to be set to the address in the BASE field plus four times the interrupt cause number. For example, a supervisor-mode timer interrupt (see [Table 2](#scauses)) causes the `pc` to be set to BASE+`0x14`. Setting MODE=Vectored may impose a stricter alignment constraint on BASE. #### [](#12-1-1-3-supervisor-interrupt-sip-and-sie-registers)12.1.1.3\. Supervisor Interrupt (`sip` and `sie`) Registers The `sip` register is an SXLEN-bit read/write register containing information on pending interrupts, while `sie` is the corresponding SXLEN-bit read/write register containing interrupt enable bits. Interrupt cause number _i_ (as reported in CSR `scause`,[12.1.1.8\. Supervisor Cause (scause) Register](#scause)) corresponds with bit _i_ in both `sip` and`sie`. Bits 15:0 are allocated to standard interrupt causes only, while bits 16 and above are designated for platform use. ![Supervisor interrupt-pending register (`sip`).](_images/diag-b33254286a615a4c40783f32b57dae1d76549bb4.svg) Figure 4\. Supervisor interrupt-pending register (`sip`). ![Supervisor interrupt-enable register (`sie`).](_images/diag-b33254286a615a4c40783f32b57dae1d76549bb4.svg) Figure 5\. Supervisor interrupt-enable register (`sie`). An interrupt _i_ will trap to S-mode if both of the following are true: (a) either the current privilege mode is S and the SIE bit in the`sstatus` register is set, or the current privilege mode has less privilege than S-mode; and (b) bit _i_ is set in both `sip` and `sie`. These conditions for an interrupt trap to occur must be evaluated in a bounded amount of time from when an interrupt becomes, or ceases to be, pending in `sip`, and must also be evaluated immediately following the execution of an SRET instruction or an explicit write to a CSR on which these interrupt trap conditions expressly depend (including `sip`, `sie`and `sstatus`). Interrupts to S-mode take priority over any interrupts to lower privilege modes. Each individual bit in register `sip` may be writable or may be read-only. When bit _i_ in `sip` is writable, a pending interrupt _i_can be cleared by writing 0 to this bit. If interrupt _i_ can become pending but bit _i_ in `sip` is read-only, the implementation must provide some other mechanism for clearing the pending interrupt (which may involve a call to the execution environment). A bit in `sie` must be writable if the corresponding interrupt can ever become pending. Bits of `sie` that are not writable are read-only zero. The standard portions (bits 15:0) of registers `sip` and `sie` are formatted as shown in Figures [Figure 6](#sipreg-standard)and [Figure 7](#siereg-standard) respectively. ![Standard portion (bits 15:0) of `sip`.](_images/diag-d4ff8af3924ac6455f101b40509ed274c3bbf7cc.svg) Figure 6\. Standard portion (bits 15:0) of `sip`. ![Standard portion (bits 15:0) of `sie`.](_images/diag-bdd8f9144ff139ccde485f890014176740dbd614.svg) Figure 7\. Standard portion (bits 15:0) of `sie`. Bits `sip`.SEIP and `sie`.SEIE are the interrupt-pending and interrupt-enable bits for supervisor-level external interrupts. If implemented, SEIP is read-only in `sip`, and is set and cleared by the execution environment, typically through a platform-specific interrupt controller. Bits `sip`.STIP and `sie`.STIE are the interrupt-pending and interrupt-enable bits for supervisor-level timer interrupts. If implemented, STIP is read-only in`sip`. When the Sstc extension is not implemented, STIP is set and cleared by the execution environment. When the Sstc extension is implemented, STIP reflects the timer interrupt signal resulting from `stimecmp`. The `sip`.STIP bit, in response to timer interrupts generated by `stimecmp`, is set by writing`stimecmp` with a value that is less than or equal to `time`, and is cleared by writing `stimecmp` with a value greater than `time`. Bits `sip`.SSIP and `sie`.SSIE are the interrupt-pending and interrupt-enable bits for supervisor-level software interrupts. If implemented, SSIP is writable in `sip` and may also be set to 1 by a platform-specific interrupt controller. If the Sscofpmf extension is implemented, bits `sip`.LCOFIP and `sie`.LCOFIE are the interrupt-pending and interrupt-enable bits for local-counter-overflow interrupts. LCOFIP is read-write in `sip` and reflects the occurrence of a local counter-overflow overflow interrupt request resulting from any of the`mhpmevent_n_`.OF bits being set. If the Sscofpmf extension is not implemented, `sip`.LCOFIP and `sie`.LCOFIE are read-only zeros. | | Interprocessor interrupts are sent to other harts by implementation-specific means, which will ultimately cause the SSIP bit to be set in the recipient hart’s sip register. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Each standard interrupt type (SEI, STI, SSI, or LCOFI) may not be implemented, in which case the corresponding interrupt-pending and interrupt-enable bits are read-only zeros. All bits in `sip` and `sie` are **WARL** fields. The implemented interrupts may be found by writing one to every bit location in `sie`, then reading back to see which bit positions hold a one. | | The sip and sie registers are subsets of the mip and mieregisters. Reading any implemented field, or writing any writable field, of sip/sie effects a read or write of the homonymous field ofmip/mie. Bits 3, 7, and 11 of sip and sie correspond to the machine-mode software, timer, and external interrupts, respectively. Since most platforms will choose not to make these interrupts delegatable from M-mode to S-mode, they are shown as 0 in[Figure 6](#sipreg-standard) and [Figure 7](#siereg-standard). | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Multiple simultaneous interrupts destined for supervisor mode are handled in the following decreasing priority order: SEI, SSI, STI, LCOFI. #### [](#12-1-1-4-supervisor-timers-and-performance-counters)12.1.1.4\. Supervisor Timers and Performance Counters Supervisor software uses the same hardware performance monitoring facility as user-mode software, including the `time`, `cycle`, and`instret` CSRs. The implementation should provide a mechanism to modify the counter values. The implementation must provide a facility for scheduling timer interrupts in terms of the real-time counter, `time`. #### [](#12-1-1-5-counter-enable-scounteren-register)12.1.1.5\. Counter-Enable (`scounteren`) Register ![Counter-enable (`scounteren`) register](_images/diag-4f3e88c89fe616d4c5c6fead328f65267e7bbf37.svg) Figure 8\. Counter-enable (`scounteren`) register The counter-enable (`scounteren`) CSR is a 32-bit register that controls the availability of the hardware performance monitoring counters to U-mode. When the CY, TM, IR, or HPM_n_ bit in the `scounteren` register is clear, attempts to read the `cycle`, `time`, `instret`, or `hpmcountern`register while executing in U-mode will cause an illegal-instruction exception. When one of these bits is set, access to the corresponding register is permitted. `scounteren` must be implemented. However, any of the bits may be read-only zero, indicating reads to the corresponding counter will cause an exception when executing in U-mode. Hence, they are effectively**WARL** fields. | | The setting of a bit in mcounteren does not affect whether the corresponding bit in scounteren is writable. However, U-mode may only access a counter if the corresponding bits in scounteren andmcounteren are both set. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#12-1-1-6-supervisor-scratch-sscratch-register)12.1.1.6\. Supervisor Scratch (`sscratch`) Register The `sscratch` CSR is an SXLEN-bit read/write register, dedicated for use by the supervisor. Typically, `sscratch` is used to hold a pointer to the hart-local supervisor context while the hart is executing user code. At the beginning of a trap handler, software normally uses a CSRRW instruction to swap `sscratch` with an integer register to obtain an initial working register. ![Supervisor Scratch Register](_images/diag-34cfa3adb3a239400e2c15dbf150614b44e17cf5.svg) Figure 9\. Supervisor Scratch Register #### [](#12-1-1-7-supervisor-exception-program-counter-sepc-register)12.1.1.7\. Supervisor Exception Program Counter (`sepc`) Register `sepc` is an SXLEN-bit read/write CSR formatted as shown in[Figure 10](#epcreg). The low bit of `sepc` (`sepc[0]`) is always zero. On implementations that support only IALIGN=32, the two low bits (`sepc[1:0]`) are always zero. If an implementation allows IALIGN to be either 16 or 32 (by changing CSR `misa`, for example), then, whenever IALIGN=32, bit `sepc[1]` is masked on reads so that it appears to be 0\. This masking occurs also for the implicit read by the SRET instruction. Though masked, `sepc[1]`remains writable when IALIGN=32. `sepc` is a **WARL** register that must be able to hold all valid virtual addresses. It need not be capable of holding all possible invalid addresses. Prior to writing `sepc`, implementations may convert an invalid address into some other invalid address that `sepc` is capable of holding. When a trap is taken into S-mode, `sepc` is written with the virtual address of the instruction that was interrupted or that encountered the exception. Otherwise, `sepc` is never written by the implementation, though it may be explicitly written by software. ![Supervisor exception program counter register.](_images/diag-770d37cd4c46dfb59cb1d81b0b7b7cc8d673c7b7.svg) Figure 10\. Supervisor exception program counter register. #### [](#scause)12.1.1.8\. Supervisor Cause (`scause`) Register The `scause` CSR is an SXLEN-bit read-write register formatted as shown in [Figure 11](#scausereg). When a trap is taken into S-mode, `scause` is written with a code indicating the event that caused the trap. Otherwise, `scause` is never written by the implementation, though it may be explicitly written by software. The Interrupt bit in the `scause` register is set if the trap was caused by an interrupt. The Exception Code field contains a code identifying the last exception or interrupt. [Table 2](#scauses) lists the possible exception codes for the current supervisor ISAs. The Exception Code is a **WLRL** field. It is required to hold the values 0–31 (i.e., bits 4–0 must be implemented), but otherwise it is only guaranteed to hold supported exception codes. ![Supervisor Cause (`scause`) register.](_images/diag-4a4b41b5284beb2d109d680b16f80bb7dd7b4ab0.svg) Figure 11\. Supervisor Cause (`scause`) register. __Table 2\. Supervisor cause (scause) register values after trap. Synchronous exception priorities are given by [Table: Synchronous exception priority in decreasing priority order](machine.html#norm:exc%5Fpriority).__ | Interrupt | Exception Code | Description | | ----------------------- | ------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 1111111111 | 012-456-8910-121314-15≥16 | _Reserved_Supervisor software interrupt _Reserved_Supervisor timer interrupt _Reserved_Supervisor external interrupt _Reserved_Counter-overflow interrupt _Reserved_ _Designated for platform use_ | | 00000000000000000000000 | 012345678910-111213141516-17181920-2324-3132-4748-63≥64 | Instruction address misalignedInstruction access faultIllegal instructionBreakpointLoad address misalignedLoad access faultStore/AMO address misalignedStore/AMO access faultEnvironment call from U-modeEnvironment call from S-mode _Reserved_Instruction page faultLoad page fault _Reserved_Store/AMO page fault _Reserved_Software checkHardware error _Reserved_ _Designated for custom use_ _Reserved_ _Designated for custom use_ _Reserved_ | #### [](#12-1-1-9-supervisor-trap-value-stval-register)12.1.1.9\. Supervisor Trap Value (`stval`) Register The `stval` CSR is an SXLEN-bit read-write register formatted as shown in [Figure 12](#stvalreg). When a trap is taken into S-mode, `stval` is written with exception-specific information to assist software in handling the trap. Otherwise, `stval` is never written by the implementation, though it may be explicitly written by software. The hardware platform will specify which exceptions must set `stval`informatively, which may unconditionally set it to zero, and which may exhibit either behavior, depending on the underlying event that caused the exception. If `stval` is written with a nonzero value when a breakpoint, address-misaligned, access-fault, page-fault, or hardware-error exception occurs on an instruction fetch, load, or store, then `stval` will contain the faulting virtual address. On a breakpoint exception raised by an EBREAK or C.EBREAK instruction, `stval`is written with either zero or the virtual address of the instruction. ![Supervisor Trap Value register.](_images/diag-9c289a0641ce933ae2952c317b4ad77c3916d086.svg) Figure 12\. Supervisor Trap Value register. If `stval` is written with a nonzero value when a misaligned load or store causes an access-fault, page-fault, or hardware-error exception, then `stval` will contain the virtual address of the portion of the access that caused the fault. If `stval` is written with a nonzero value when an instruction access-fault, page-fault, or hardware-error exception occurs on a hart with variable-length instructions, then `stval` will contain the virtual address of the portion of the instruction that caused the fault, while`sepc` will point to the beginning of the instruction. The `stval` register can optionally also be used to return the faulting instruction bits on an illegal-instruction exception (`sepc` points to the faulting instruction in memory). If `stval` is written with a nonzero value when an illegal-instruction exception occurs, then `stval`will contain the shortest of: * the actual faulting instruction * the first ILEN bits of the faulting instruction * the first SXLEN bits of the faulting instruction The value loaded into `stval` on an illegal-instruction exception is right-justified and all unused upper bits are cleared to zero. On a trap caused by a software-check exception, the `stval` register holds the cause for the exception. The following encodings are defined: * 0 - No information provided. * 2 - Landing Pad Fault. Defined by the Zicfilp extension ([Landing Pad (Zicfilp)](priv-cfi.html#priv-forward)). * 3 - Shadow Stack Fault. Defined by the Zicfiss extension ([Shadow Stack (Zicfiss)](priv-cfi.html#priv-backward)). For other traps, `stval` is set to zero, but a future standard may redefine `stval`’s setting for other traps. `stval` is a **WARL** register that must be able to hold all valid virtual addresses and the value 0\. It need not be capable of holding all possible invalid addresses. Prior to writing `stval`, implementations may convert an invalid address into some other invalid address that`stval` is capable of holding. If the feature to return the faulting instruction bits is implemented, `stval` must also be able to hold all values less than 2_N_, where _N_ is the smaller of SXLEN and ILEN. #### [](#sec:senvcfg)12.1.1.10\. Supervisor Environment Configuration (`senvcfg`) Register The `senvcfg` CSR is an SXLEN-bit read/write register, formatted as shown in [Figure 13](#senvcfg), that controls certain characteristics of the U-mode execution environment. ![Supervisor environment configuration register (`senvcfg`) for RV64.](_images/svg-e86891e9384d8ce0f4bc2f23bc472535422cfa78.svg) Figure 13\. Supervisor environment configuration register (`senvcfg`) for RV64. ![Supervisor environment configuration register (`senvcfg`) for RV32.](_images/svg-bcc87728447213cd09f2797ec3b3d2fcfac80d5f.svg) Figure 14\. Supervisor environment configuration register (`senvcfg`) for RV32. If bit FIOM (Fence of I/O implies Memory) is set to one in `senvcfg`, FENCE instructions executed in U-mode are modified so the requirement to order accesses to device I/O implies also the requirement to order main memory accesses. [Table 3](#senvcfg-FIOM) details the modified interpretation of FENCE instruction bits PI, PO, SI, and SO in U-mode when FIOM=1. Similarly, for U-mode when FIOM=1, if an atomic instruction that accesses a region ordered as device I/O has its _aq_ and/or _rl_ bit set, then that instruction is ordered as though it accesses both device I/O and memory. If `satp`.MODE is read-only zero (always Bare), the implementation may make FIOM read-only zero. __Table 3\. Modified interpretation of FENCE predecessor and successor sets in U-mode when FIOM=1.__ | Instruction bit | Meaning when set | | --------------- | -------------------------------------------------------------------------------------------------------------- | | PIPO | Predecessor device input and memory reads (PR implied)Predecessor device output and memory writes (PW implied) | | SISO | Successor device input and memory reads (SR implied)Successor device output and memory writes (SW implied) | | | Bit FIOM exists for a specific circumstance when an I/O device is being emulated for U-mode and both of the following are true: (a) the emulated device has a memory buffer that should be I/O space but is actually mapped to main memory via address translation, and (b) multiple physical harts are involved in accessing this emulated device from U-mode. A hypervisor running in S-mode without the benefit of the hypervisor extension of ["H" Extension for Hypervisor Support](hypervisor.html#hypervisor) may need to emulate a device for U-mode if paravirtualization cannot be employed. If the same hypervisor provides a virtual machine (VM) with multiple virtual harts, mapped one-to-one to real harts, then multiple harts may concurrently access the emulated device, perhaps because: (a) the guest OS within the VM assigns device interrupt handling to one hart while the device is also accessed by a different hart outside of an interrupt handler, or (b) control of the device (or partial control) is being migrated from one hart to another, such as for interrupt load balancing within the VM. For such cases, guest software within the VM is expected to properly coordinate access to the (emulated) device across multiple harts using mutex locks and/or interprocessor interrupts as usual, which in part entails executing I/O fences. But those I/O fences may not be sufficient if some of the device \`\`I/O'' is actually main memory, unknown to the guest. Setting FIOM=1 modifies those fences (and all other I/O fences executed in U-mode) to include main memory, too. Software can always avoid the need to set FIOM by never using main memory to emulate a device memory buffer that should be I/O space. However, this choice usually requires trapping all U-mode accesses to the emulated buffer, which might have a noticeable impact on performance. The alternative offered by FIOM is sufficiently inexpensive to implement that we consider it worth supporting even if only rarely enabled. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The Zicboz extension adds the `CBZE` (Cache Block Zero instruction enable) field to `senvcfg`. The `CBZE` field controls execution of the cache block zero instruction (`CBO.ZERO`) in U-mode. Execution of `CBO.ZERO` in U-mode is enabled only if execution of the instruction is enabled for use in S-mode and `CBZE` is set to 1; otherwise, an illegal-instruction exception is raised. When the Zicboz extension is not implemented, `CBZE` is read-only zero. The Zicbom extension adds the `CBCFE` (Cache Block Clean and Flush instruction Enable) field to `senvcfg` to control execution of the `CBO.CLEAN` and`CBO.FLUSH` instructions in U-mode. Execution of these instructions in U-mode is enabled only if execution of these instructions is enabled for use in S-mode and`CBCFE` is set to 1; otherwise, an illegal-instruction exception is raised. When the Zicbom extension is not implemented, `CBCFE` is read-only zero. The Zicbom extension adds the `CBIE` (Cache Block Invalidate instruction Enable) WARL field to `senvcfg` to control execution of the `CBO.INVAL` instruction in U-mode. The encoding `10b` is reserved. When the Zicbom extension is not implemented, `CBIE` is read-only zero. Execution of `CBO.INVAL` in U-mode is enabled only if execution of the instruction is enabled for use in S-mode and`CBIE` is set to `01b` or `11b`; otherwise, an illegal-instruction exception is raised. If `CBO.INVAL` is enabled in S-mode to perform a flush operation, then when the instruction is enabled in U-mode it performs a flush operation, even if `CBIE`is set to `11b`. Otherwise, the instruction behaves as follows, depending on the`CBIE` encoding: * `01b` — The instruction is executed and performs a flush operation. * `11b` — The instruction is executed and performs an invalidate operation. If the Ssnpm extension is implemented, the `PMM` field enables or disables pointer masking (see [Pointer Masking Extensions](zpm.html)) for the next-lower privilege mode (U/VU), according to the values in [Table 4](#senvcfg-pmm-values). If Ssnpm is not implemented, `PMM` is read-only zero. The `PMM` field is read-only zero for RV32. __Table 4\. Legal values of PMM WARL field__ | Value | Description | | ----- | ---------------------------------------------------------------------- | | 00 | Pointer masking is disabled (PMLEN = 0) | | 01 | Reserved | | 10 | Pointer masking is enabled with PMLEN = XLEN - 57 (PMLEN = 7 on RV64) | | 11 | Pointer masking is enabled with PMLEN = XLEN - 48 (PMLEN = 16 on RV64) | The Zicfilp extension adds the `LPE` field in `senvcfg`. When the `LPE` field is set to 1, the Zicfilp extension is enabled in VU/U-mode. When the `LPE` field is 0, the Zicfilp extension is not enabled in VU/U-mode and the following rules apply to VU/U-mode: * The hart does not update the `ELP` state; it remains as `NO_LP_EXPECTED`. * The `LPAD` instruction operates as a no-op. The Zicfiss extension adds the `SSE` field in `senvcfg`. When the `SSE` field is set to 1, the Zicfiss extension is activated in VU/U-mode. When the `SSE` field is 0, the Zicfiss extension remains inactive in VU/U-mode, and the following rules apply: * 32-bit Zicfiss instructions will revert to their behavior as defined by Zimop. * 16-bit Zicfiss instructions will revert to their behavior as defined by Zcmop. * When `menvcfg.SSE` is one, `SSAMOSWAP.W/D` raises an illegal-instruction exception in U-mode and a virtual-instruction exception in VU-mode. #### [](#satp)12.1.1.11\. Supervisor Address Translation and Protection (`satp`) Register The `satp` CSR is an SXLEN-bit read/write register, formatted as shown in [Figure 15](#rv32satp) for SXLEN=32 and[Figure 16](#rv64satp) for SXLEN=64, which controls supervisor-mode address translation and protection. This register holds the physical page number (PPN) of the root page table, i.e., its supervisor physical address divided by 4 KiB; an address space identifier (ASID), which facilitates address-translation fences on a per-address-space basis; and the MODE field, which selects the current address-translation scheme. Further details on the access to this register are described in [Virtualization Support in mstatus Register](machine.html#virt-control). ![Supervisor address translation and protection (`satp`) register when SXLEN=32.](_images/diag-cf2342eda91ed518cf259a8b38fa84ecda4ac327.svg) Figure 15\. Supervisor address translation and protection (`satp`) register when SXLEN=32. | | Storing a PPN in satp, rather than a physical address, supports a physical address space larger than 4 GiB for RV32. The satp.PPN field might not be capable of holding all physical page numbers. Some platform standards might place constraints on the valuessatp.PPN may assume, e.g., by requiring that all physical page numbers corresponding to main memory be representable. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ![Supervisor address translation and protection (`satp`) register when SXLEN=64, for MODE values Bare, Sv39, Sv48, and Sv57.](_images/diag-074dc5fbd1ea1c221bf077c8cfa6f508fbbb791c.svg) Figure 16\. Supervisor address translation and protection (`satp`) register when SXLEN=64, for MODE values Bare, Sv39, Sv48, and Sv57. | | We store the ASID and the page table base address in the same CSR to allow the pair to be changed atomically on a context switch. Swapping them non-atomically could pollute the old virtual address space with new translations, or vice-versa. This approach also slightly reduces the cost of a context switch. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | [Table 5](#satp-mode) shows the encodings of the MODE field when SXLEN=32 and SXLEN=64\. When MODE=Bare, supervisor virtual addresses are equal to supervisor physical addresses, and there is no additional memory protection beyond the physical memory protection scheme described in [Physical Memory Protection](machine.html#pmp). To select MODE=Bare, software must write zero to the remaining fields of `satp` (bits 30–0 when SXLEN=32, or bits 59–0 when SXLEN=64). Attempting to select MODE=Bare with a nonzero pattern in the remaining fields has an UNSPECIFIED effect on the value that the remaining fields assume and an UNSPECIFIED effect on address translation and protection behavior. When SXLEN=32, the `satp` encodings corresponding to MODE=Bare and ASID\[8:7\]=3 are designated for custom use, whereas the encodings corresponding to MODE=Bare and ASID\[8:7\]≠3 are reserved for future standard use. When SXLEN=64, all `satp` encodings corresponding to MODE=Bare are reserved for future standard use. | | Version 1.11 of this standard stated that the remaining fields in satphad no effect when MODE=Bare. Making these fields reserved facilitates future definition of additional translation and protection modes, particularly in RV32, for which all patterns of the existing MODE field have already been allocated. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When SXLEN=32, the only other valid setting for MODE is Sv32, a paged virtual-memory scheme described in [12.1.3\. Sv32: Page-Based 32-bit Virtual-Memory Systems](#sv32). When SXLEN=64, three paged virtual-memory schemes are defined: Sv39, Sv48, and Sv57, described in [12.1.4\. Sv39: Page-Based 39-bit Virtual-Memory System](#sv39), [12.1.5\. Sv48: Page-Based 48-bit Virtual-Memory System](#sv48), and [12.1.6\. Sv57: Page-Based 57-bit Virtual-Memory System](#sv57), respectively. One additional scheme, Sv64, will be defined in a later version of this specification. The remaining MODE settings are reserved for future use and may define different interpretations of the other fields in `satp`. Implementations are not required to support all MODE settings, and if`satp` is written with an unsupported MODE, the entire write has no effect; no fields in `satp` are modified. The number of ASID bits is UNSPECIFIED and may be zero. The number of implemented ASID bits, termed _ASIDLEN_, may be determined by writing one to every bit position in the ASID field, then reading back the value in `satp` to see which bit positions in the ASID field hold a one. The least-significant bits of ASID are implemented first: that is, if ASIDLEN > 0, ASID\[ASIDLEN-1:0\] is writable. The maximal value of ASIDLEN, termed ASIDMAX, is 9 for Sv32 or 16 for Sv39, Sv48, and Sv57. __Table 5\. Encoding of satp MODE field.__ | SXLEN=32 | | | | -------------------- | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Value | Name | Description | | 01 | BareSv32 | No translation or protection.Page-based 32-bit virtual addressing (see [12.1.3\. Sv32: Page-Based 32-bit Virtual-Memory Systems](#sv32)). | | **SXLEN=64** | | | | Value | Name | Description | | 01-789101112-1314-15 | Bare\-Sv39Sv48Sv57Sv64\-\- | No translation or protection. _Reserved for standard use_Page-based 39-bit virtual addressing (see [12.1.4\. Sv39: Page-Based 39-bit Virtual-Memory System](#sv39)).Page-based 48-bit virtual addressing (see [12.1.5\. Sv48: Page-Based 48-bit Virtual-Memory System](#sv48)).Page-based 57-bit virtual addressing (see [12.1.6\. Sv57: Page-Based 57-bit Virtual-Memory System](#sv57)). _Reserved for page-based 64-bit virtual addressing._ _Reserved for standard use_ _Designated for custom use_ | | | For many applications, the choice of page size has a substantial performance impact. A large page size increases TLB reach and loosens the associativity constraints on virtually indexed, physically tagged caches. At the same time, large pages exacerbate internal fragmentation, wasting physical memory and possibly cache capacity. After much deliberation, we have settled on a conventional page size of 4 KiB for both RV32 and RV64\. We expect this decision to ease the porting of low-level runtime software and device drivers. The TLB reach problem is ameliorated by transparent superpage support in modern operating systems. \[[92](../biblio/bibliography.html#bib-transparent-superpages)\] Additionally, multi-level TLB hierarchies are quite inexpensive relative to the multi-level cache hierarchies whose address space they map. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The `satp` CSR is considered _active_ when the effective privilege mode is S-mode or U-mode. Executions of the address-translation algorithm may only begin using a given value of `satp` when `satp` is active. | | Translations that began while satp was active are not required to complete or terminate when satp is no longer active, unless an SFENCE.VMA instruction matching the address and ASID is executed. The SFENCE.VMA instruction must be used to ensure that updates to the address-translation data structures are observed by subsequent implicit reads to those structures by a hart. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Note that writing `satp` does not imply any ordering constraints between page-table updates and subsequent address translations, nor does it imply any invalidation of address-translation caches. If the new address space’s page tables have been modified, or if an ASID is reused, it may be necessary to execute an SFENCE.VMA instruction (see[12.1.2.1\. Supervisor Memory-Management Fence Instruction](#sfence.vma)) after, or in some cases before, writing`satp`. | | Not imposing upon implementations to flush address-translation caches upon satp writes reduces the cost of context switches, provided a sufficiently large ASID space. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#stimecmp)12.1.1.12\. Supervisor Timer (`stimecmp`) Register The `stimecmp` CSR is a 64-bit register and has 64-bit precision on all RV32 and RV64 systems. In RV32 only, accesses to the `stimecmp` CSR access the low 32 bits, while accesses to the `stimecmph` CSR access the high 32 bits of `stimecmp`. A supervisor timer interrupt becomes pending, as reflected in the STIP bit in the `mip` and `sip` registers whenever `time` contains a value greater than or equal to `stimecmp`, treating the values as unsigned integers. If the result of this comparison changes, it is guaranteed to be reflected in STIP eventually, but not necessarily immediately. The interrupt remains posted until `stimecmp` becomes greater than `time`, typically as a result of writing `stimecmp`. The interrupt will be taken based on the standard interrupt enable and delegation rules. | | A spurious timer interrupt might occur if an interrupt handler advancesstimecmp then immediately returns, because STIP might not yet have fallen in the interim. All software should be written to assume this event is possible, but most software should assume this event is extremely unlikely. It is almost always more performant to incur an occasional spurious timer interrupt than to poll STIP until it falls. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | In systems in which a supervisor execution environment (SEE) provides timer facilities via an SBI function call, this SBI call will continue to support requests to schedule a timer interrupt. The SEE will simply make use of stimecmp, changing its value as appropriate. This ensures compatibility with existing S-mode software that uses this SEE facility, while new S-mode software takes advantage of stimecmp directly.) | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#12-1-2-supervisor-instructions)12.1.2\. Supervisor Instructions In addition to the SRET instruction defined in [Trap-Return Instructions](machine.html#otherpriv), one new supervisor-level instruction is provided. #### [](#sfence.vma)12.1.2.1\. Supervisor Memory-Management Fence Instruction ![svg](_images/svg-923a12c8951e1924fa31f2b0e7f04f2135dee632.svg) The supervisor memory-management fence instruction SFENCE.VMA is used to synchronize updates to in-memory memory-management data structures with current execution. Instruction execution causes implicit reads and writes to these data structures; however, these implicit references are ordinarily not ordered with respect to explicit loads and stores.Executing an SFENCE.VMA instruction guarantees that any previous stores already visible to the current RISC-V hart are ordered before certain implicit references by subsequent instructions in that hart to the memory-management data structures. The specific set of operations ordered by SFENCE.VMA is determined by _rs1_ and _rs2_, as described below. SFENCE.VMA is also used to invalidate entries in the address-translation cache associated with a hart (see [12.1.3.2\. Virtual Address Translation Process](#sv32algorithm)). Further details on the behavior of this instruction are described in [Virtualization Support in mstatus Register](machine.html#virt-control) and [Physical Memory Protection and Paging](machine.html#pmp-vmem). | | The SFENCE.VMA is used to flush any local hardware caches related to address translation. It is specified as a fence rather than a TLB flush to provide cleaner semantics with respect to which instructions are affected by the flush operation and to support a wider variety of dynamic caching structures and memory-management schemes. SFENCE.VMA is also used by higher privilege levels to synchronize page table writes and the address translation hardware. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | SFENCE.VMA orders only the local hart’s implicit references to the memory-management data structures. | | Consequently, other harts must be notified separately when the memory-management data structures have been modified. One approach is to use 1) a local data fence to ensure local writes are visible globally, then 2) an interprocessor interrupt to the other thread, then 3) a local SFENCE.VMA in the interrupt handler of the remote thread, and finally 4) signal back to originating thread that operation is complete. This is, of course, the RISC-V analog to a TLB shootdown. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For the common case that the translation data structures have only been modified for a single address mapping (i.e., one page or superpage),_rs1_ can specify a virtual address within that mapping to effect a translation fence for that mapping only. Furthermore, for the common case that the translation data structures have only been modified for a single address-space identifier, _rs2_ can specify the address space. The behavior of SFENCE.VMA depends on _rs1_ and _rs2_ as follows: * If _rs1_\=`x0` and _rs2_\=`x0`, the fence orders all reads and writes made to any level of the page tables, for all address spaces. The fence also invalidates all address-translation cache entries, for all address spaces. * If _rs1_\=`x0` and _rs2_≠`x0`, the fence orders all reads and writes made to any level of the page tables, but only for the address space identified by integer register _rs2_. Accesses to _global_mappings (see [12.1.3.1\. Addressing and Memory Protection](#translation)) are not ordered. The fence also invalidates all address-translation cache entries matching the address space identified by integer register _rs2_, except for entries containing global mappings. * If _rs1_≠`x0` and _rs2_\=`x0`, the fence orders only reads and writes made to leaf page table entries corresponding to the virtual address in _rs1_, for all address spaces. The fence also invalidates all address-translation cache entries that contain leaf page table entries corresponding to the virtual address in _rs1_, for all address spaces. * If _rs1_≠`x0` and _rs2_≠`x0`, the fence orders only reads and writes made to leaf page table entries corresponding to the virtual address in _rs1_, for the address space identified by integer register _rs2_. Accesses to global mappings are not ordered. The fence also invalidates all address-translation cache entries that contain leaf page table entries corresponding to the virtual address in _rs1_ and that match the address space identified by integer register _rs2_, except for entries containing global mappings. If the value held in _rs1_ is not a valid virtual address, then the SFENCE.VMA instruction has no effect. No exception is raised in this case. | | It is always legal to over-fence, e.g., by fencing only based on a subset of the bits in _rs1_ and/or _rs2_, and/or by simply treating all SFENCE.VMA instructions as having _rs1_\=x0 and/or _rs2_\=x0. For example, simpler implementations can ignore the virtual address in _rs1_and the ASID value in _rs2_ and always perform a global fence. The choice not to raise an exception when an invalid virtual address is held in _rs1_ facilitates this type of simplification. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When _rs2_≠`x0`, bits SXLEN-1:ASIDMAX of the value held in _rs2_ are reserved for future standard use. Until their use is defined by a standard extension, they should be zeroed by software and ignored by current implementations. Furthermore, if ASIDLEN0_ and _pte_._ppn_\[_i_\-1:0\] ≠ 0, this is a misaligned superpage; stop and raise a page-fault exception corresponding to the original access type. 6. Determine if the requested memory access is allowed by the _pte_._u_ bit, given the current privilege mode and the value of the SUM and MXR fields of the **mstatus** register. If not, stop and raise a page-fault exception corresponding to the original access type. 7. Determine if the requested memory access is allowed by the _pte_._r_, _pte_._w_, and _pte_._x_ bits, given the Shadow Stack Memory Protection rules. If not, stop and raise an access-fault exception. 8. Determine if the requested memory access is allowed by the _pte_._r_, _pte_._w_, and _pte_._x_ bits. If not, stop and raise a page-fault exception corresponding to the original access type. 9. If _pte_._a_\=0, or if the original memory access is a store and _pte_._d_\=0: * If the Svade extension is implemented, stop and raise a page-fault exception corresponding to the original access type. * If a store to the PTE at address _a_+_va.vpn_\[_i_\]×PTESIZE would violate a PMA or PMP check, raise an access-fault exception corresponding to the original access type. * Perform the following steps atomically: * Compare _pte_ to the value of the PTE at address _a_+_va.vpn_\[_i_\]×PTESIZE. * If the values match, set _pte_._a_ to 1 and, if the original memory access is a store, also set _pte_._d_ to 1\. Then store _pte_ to the PTE at address _a_+_va.vpn_\[_i_\]×PTESIZE. * If the comparison fails, return to step 2. 10. The translation is successful. The translated physical address is given as follows: * _pa.pgoff_ \= _va.pgoff_. * If _i_\>0, then this is a superpage translation and _pa.ppn_\[_i_\-1:0\] = _va.vpn_\[_i_\-1:0\]. * _pa.ppn_\[LEVELS-1:_i_\] = _pte_._ppn_\[LEVELS-1:_i_\]. All implicit accesses to the address-translation data structures in this algorithm are performed using width PTESIZE. | | This implies, for example, that an Sv48 implementation may not use two separate 4 B reads to non-atomically access a single 8 B PTE, and that A/D bit updates performed by the implementation are treated as atomically updating the entire PTE, rather than just the A and/or D bit alone (even though the PTE value does not otherwise change). | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The results of implicit address-translation reads in step 2 may be held in a read-only, incoherent _address-translation cache_ but not shared with other harts. The address-translation cache may hold an arbitrary number of entries, including an arbitrary number of entries for the same address and ASID. Entries in the address-translation cache may then satisfy subsequent step 2 reads if the ASID associated with the entry matches the ASID loaded in step 0 or if the entry is associated with a_global_ mapping. To ensure that implicit reads observe writes to the same memory locations, an SFENCE.VMA instruction must be executed after the writes to flush the relevant cached translations. The address-translation cache cannot be used in step 9; accessed and dirty bits may only be updated in memory directly. | | It is permitted for multiple address-translation cache entries to co-exist for the same address. This represents the fact that in a conventional TLB hierarchy, it is possible for multiple entries to match a single address if, for example, a page is upgraded to a superpage without first clearing the original non-leaf PTE’s valid bit and executing an SFENCE.VMA with _rs1_\=x0, or if multiple TLBs exist in parallel at a given level of the hierarchy. In this case, just as if an SFENCE.VMA is not executed between a write to the memory-management tables and subsequent implicit read of the same address: it is unpredictable whether the old non-leaf PTE or the new leaf PTE is used, but the behavior is otherwise well defined. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Implementations may also execute the address-translation algorithm speculatively at any time, for any virtual address, as long as `satp` is active (as defined in [12.1.1.11\. Supervisor Address Translation and Protection (satp) Register](#satp)). Such speculative executions have the effect of pre-populating the address-translation cache. Speculative executions of the address-translation algorithm behave as non-speculative executions of the algorithm do, except that they must not set the dirty bit for a PTE, they must not trigger an exception, and they must not create address-translation cache entries if those entries would have been invalidated by any SFENCE.VMA instruction executed by the hart since the speculative execution of the algorithm began. | | For instance, it is illegal for both non-speculative and speculative executions of the translation algorithm to begin, read the level 2 page table, pause while the hart executes an SFENCE.VMA with_rs1_\=_rs2_\=x0, then resume using the now-stale level 2 PTE, as subsequent implicit reads could populate the address-translation cache with stale PTEs. In many implementations, an SFENCE.VMA instruction with _rs1_\=x0 will therefore either terminate all previously-launched speculative executions of the address-translation algorithm (for the specified ASID, if applicable), or simply wait for them to complete (in which case any address-translation cache entries created will be invalidated by the SFENCE.VMA as appropriate). Likewise, an SFENCE.VMA instruction with_rs1_≠x0 generally must either ensure that previously-launched speculative executions of the address-translation algorithm (for the specified ASID, if applicable) are prevented from creating new address-translation cache entries mapping leaf PTEs, or wait for them to complete. A consequence of implementations being permitted to read the translation data structures arbitrarily early and speculatively is that at any time, all page table entries reachable by executing the algorithm may be loaded into the address-translation cache. Although it would be uncommon to place page tables in non-idempotent memory, there is no explicit prohibition against doing so. Since the algorithm may only touch page tables reachable from the root page table indicated in satp, the range of addresses that an implementation’s page-table walker will touch is fully under supervisor control. The algorithm does not admit the possibility of ignoring high-order PPN bits for implementations with narrower physical addresses. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#sv39)12.1.4\. Sv39: Page-Based 39-bit Virtual-Memory System This section describes a simple paged virtual-memory system for SXLEN=64, which supports 39-bit virtual address spaces. The design of Sv39 follows the overall scheme of Sv32, and this section details only the differences between the schemes. | | We specified multiple virtual memory systems for RV64 to relieve the tension between providing a large address space and minimizing address-translation cost. For many systems, 39 bits of virtual-address space is ample, and so Sv39 suffices. Sv48 increases the virtual address space to 48 bits, but increases the physical memory capacity dedicated to page tables, the latency of page-table traversals, and the size of hardware structures that store virtual addresses. Sv57 increases the virtual address space, page table capacity requirement, and translation latency even further. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#addressing-and-memory-protection)12.1.4.1\. Addressing and Memory Protection Sv39 implementations support a 39-bit virtual address space, divided into pages. An Sv39 address is partitioned as shown in[Figure 20](#sv39va). Instruction fetch addresses and load and store effective addresses, which are 64 bits,must have bits 63–39 all equal to bit 38, or else a page-fault exception will occur. The 27-bit VPN is translated into a44-bit PPN via athree-level page table, while the12-bit page offset is untranslated. | | When mapping between narrower and wider addresses, RISC-V zero-extends a narrower physical address to a wider size. The mapping between 64-bit virtual addresses and the 39-bit usable address space of Sv39 is not based on zero extension but instead follows an entrenched convention that allows an OS to use one or a few of the most-significant bits of a full-size (64-bit) virtual address to quickly distinguish user and supervisor address regions. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ![Sv39 virtual address.](_images/diag-1abc9a2fed5d17636919fb7e458b6ce8b3ba4073.svg) Figure 20\. Sv39 virtual address. ![Sv39 physical address.](_images/diag-30f0b3822fd96cece4635639eca58633f7b7138e.svg) Figure 21\. Sv39 physical address. ![Sv39 page table entry.](_images/diag-beed6ade106087e4449846ecba47afe435ea5c8f.svg) Figure 22\. Sv39 page table entry. Sv39 page tables contain 29 page table entries (PTEs), eight bytes each. A page table is exactly the size of a page andmust always be aligned to a page boundary. The physical page number of the root page table is stored in the `satp`register’s PPN field. The PTE format for Sv39 is shown in [Figure 22](#sv39pte). Bits 9-0 have the same meaning as for Sv32\. Bit 63 is reserved for use by the Svnapot extension in [12.1.7\. "Svnapot" Extension for NAPOT Translation Contiguity, Version 1.0](#svnapot).If Svnapot is not implemented, bit 63 remains reserved and must be zeroed by software for forward compatibility, or else a page-fault exception is raised. Bits 62-61 are reserved for use by the Svpbmt extension in [12.1.8\. "Svpbmt" Extension for Page-Based Memory Types, Version 1.0](#svpbmt).If Svpbmt is not implemented, bits 62-61 remain reserved and must be zeroed by software for forward compatibility, or else a page-fault exception is raised. Bits 60-54 are reserved for future standard use and, until their use is defined by some standard extension,must be zeroed by software for forward compatibility. If any of these bits are set, a page-fault exception is raised. | | We reserved several PTE bits for a possible extension that improves support for sparse address spaces by allowing page-table levels to be skipped, reducing memory usage and TLB refill latency. These reserved bits may also be used to facilitate research experimentation. The cost is reducing the physical address space, but is presently ample. When it no longer suffices, the reserved bits that remain unallocated could be used to expand the physical address space. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Any level of PTE may be a leaf PTE, so in addition to 4 KiB pages, Sv39 supports2 MiB megapages and 1 GiB gigapages, each of which must be virtually and physically aligned to a boundary equal to its size.A page-fault exception is raised if the physical address is insufficiently aligned. The algorithm for virtual-to-physical address translation is the same as in [12.1.3.2\. Virtual Address Translation Process](#sv32algorithm), exceptLEVELS equals 3 andPTESIZE equals 8. ### [](#sv48)12.1.5\. Sv48: Page-Based 48-bit Virtual-Memory System This section describes a simple paged virtual-memory system forSXLEN=64, which supports48-bit virtual address spaces. Sv48 is intended for systems for which a 39-bit virtual address space is insufficient. It closely follows the design of Sv39, simply adding an additional level of page table, and so this chapter only details the differences between the two schemes. Implementations that support Sv48 must also support Sv39. | | Systems that support Sv48 can also support Sv39 at essentially no cost, and so should do so to maintain compatibility with supervisor software that assumes Sv39. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#addressing-and-memory-protection-1)12.1.5.1\. Addressing and Memory Protection Sv48 implementations support a 48-bit virtual address space, divided into pages. An Sv48 address is partitioned as shown in[Figure 23](#sv48va). Instruction fetch addresses and load and store effective addresses, which are 64 bits,must have bits 63–48 all equal to bit 47, or else a page-fault exception will occur. The 36-bit VPN is translated into a44-bit PPN via afour-level page table, while the12-bit page offset is untranslated. ![Sv48 virtual address.](_images/diag-78365ce07def203a46c5d41e86fe44b2b82fee68.svg) Figure 23\. Sv48 virtual address. ![Sv48 physical address.](_images/diag-90d0c94fee3eaf80b936cf9ee110021c4b79f4a0.svg) Figure 24\. Sv48 physical address. ![Sv48 page table entry.](_images/diag-cf758c77a8c18632ce647d2932269d694a5c030d.svg) Figure 25\. Sv48 page table entry. The PTE format for Sv48 is shown in [Figure 25](#sv48pte). Bits 63-54 and 9-0 have the same meaning as for Sv39.Any level of PTE may be a leaf PTE, so in addition to 4 KiB pages, Sv48 supports2 MiB megapages, 1 GiB gigapages, and 512 GiB terapages, each of whichmust be virtually and physically aligned to a boundary equal to its size.A page-fault exception is raised if the physical address is insufficiently aligned. The algorithm for virtual-to-physical address translation is the same as in [12.1.3.2\. Virtual Address Translation Process](#sv32algorithm), exceptLEVELS equals 4 andPTESIZE equals 8. ### [](#sv57)12.1.6\. Sv57: Page-Based 57-bit Virtual-Memory System This section describes a simple paged virtual-memory system designed forRV64 systems, which supports57-bit virtual address spaces. Sv57 is intended for systems for which a 48-bit virtual address space is insufficient. It closely follows the design of Sv48, simply adding an additional level of page table, and so this chapter only details the differences between the two schemes. Implementations that support Sv57 must also support Sv48. | | Systems that support Sv57 can also support Sv48 at essentially no cost, and so should do so to maintain compatibility with supervisor software that assumes Sv48. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#addressing-and-memory-protection-2)12.1.6.1\. Addressing and Memory Protection Sv57 implementations support a 57-bit virtual address space, divided into pages. An Sv57 address is partitioned as shown in[Figure 26](#sv57va). Instruction fetch addresses and load and store effective addresses, which are 64 bits,must have bits 63–57 all equal to bit 56, or else a page-fault exception will occur. The 45-bit VPN is translated into a44-bit PPN via afive-level page table, while the12-bit page offset is untranslated. ![Sv57 virtual address.](_images/diag-974886df6158fc60f28ef5296c6420d172094f4c.svg) Figure 26\. Sv57 virtual address. ![Sv57 physical address.](_images/diag-fd261849fbb4176194c52fd8848cae2f0a6e2e34.svg) Figure 27\. Sv57 physical address. ![Sv57 page table entry.](_images/diag-b4890c1ecb7d8e0dfdf4aa235ff3950218acb682.svg) Figure 28\. Sv57 page table entry. The PTE format for Sv57 is shown in [Figure 28](#sv57pte). Bits 63–54 and 9–0 have the same meaning as for Sv39.Any level of PTE may be a leaf PTE, so in addition to 4 KiB pages, Sv57 supports2 MiB megapages, 1 GiB gigapages, 512 GiB terapages, and 256 TiB petapages, each of whichmust be virtually and physically aligned to a boundary equal to its size.A page-fault exception is raised if the physical address is insufficiently aligned. The algorithm for virtual-to-physical address translation is the same as in [12.1.3.2\. Virtual Address Translation Process](#sv32algorithm), exceptLEVELS equals 5 andPTESIZE equals 8. ### [](#svnapot)12.1.7\. "Svnapot" Extension for NAPOT Translation Contiguity, Version 1.0 In Sv39, Sv48, and Sv57, when a PTE hasN=1, the PTE represents a translation that is part of a range of contiguous virtual-to-physical translations with the same values for PTE bits 5–0.Such ranges must be of a naturally aligned power-of-2 (NAPOT) granularity larger than the base page size. The Svnapot extension depends on the Sv39 extension. __Table 7\. Page table entry encodings when _pte_.N=1__ | i | _pte_._ppn_\[_i_\] | Description | _pte_._napot\_bits_ | | ------- | ----------------------------------------------------------------------- | ------------------------------------------------------------------------------ | ------------------- | | 00000≥1 | x xxxx xxx1 x xxxx xx1x x xxxx x1xx x xxxx 1000 x xxxx 0xxx x xxxx xxxx | _Reserved_ _Reserved_ _Reserved_64 KiB contiguous region _Reserved_ _Reserved_ | \-\-\-4\-\- | NAPOT PTEs behave identically to non-NAPOT PTEs within the address-translation algorithm in [12.1.3.2\. Virtual Address Translation Process](#sv32algorithm), except that: * If the encoding in _pte_ isvalid according to [Table 7](#ptenapot), then instead of returning the original value of _pte_,implicit reads of a NAPOT PTE return a copy of _pte_ in which _pte_._ppn_\[_i_\]\[_pte_._napot\_bits_\-1:0\] is replaced by _vpn_\[_i_\]\[_pte_._napot\_bits_\-1:0\]. If the encoding in _pte_ isreserved according to[Table 7](#ptenapot), then a page-fault exception must be raised. * Implicit reads of NAPOT page table entries may create address-translation cache entries mapping_a_ \+ _j_×PTESIZE to a copy of _pte_ in which_pte_._ppn_\[_i_\]\[_pte_._napot\_bits_\-1:0\] is replaced by_vpn\[i\]\[pte.napot\_bits_\-1:0\], for any or all _j_ such that_j_ \>> _napot\_bits_ \= _vpn_\[_i_\] >> _napot\_bits_, all for the address space identified in _satp_ as loaded by step 1. | | The motivation for a NAPOT PTE is that it can be cached in a TLB as one or more entries representing the contiguous region as if it were a single (large) page covered by a single translation. This compaction can help relieve TLB pressure in some scenarios. The encoding is designed to fit within the pre-existing Sv39, Sv48, and Sv57 PTE formats so as not to disrupt existing implementations or designs that choose not to implement the scheme. It is also designed so as not to complicate the definition of the address-translation algorithm. The address translation cache abstraction captures the behavior that would result from the creation of a single TLB entry covering the entire NAPOT region. It is also designed to be consistent with implementations that support NAPOT PTEs by splitting the NAPOT region into TLB entries covering any smaller power-of-two region sizes. For example, a 64 KiB NAPOT PTE might trigger the creation of 16 standard 4 KiB TLB entries, all with contents generated from the NAPOT PTE (even if the PTEs for the other 4 KiB regions have different contents). In typical usage scenarios, NAPOT PTEs in the same region will have the same attributes, same PPNs, and same values for bits 5-0\. RSW remains reserved for supervisor software control. It is the responsibility of the OS and/or hypervisor to configure the page tables in such a way that there are no inconsistencies between NAPOT PTEs and other NAPOT or non-NAPOT PTEs that overlap the same address range. If an update needs to be made, the OS generally should first mark all of the PTEs invalid, then issue SFENCE.VMA instruction(s) covering all 4 KiB regions within the range (either via a single SFENCE.VMA with _rs1_\=x0, or with multiple SFENCE.VMA instructions with _rs1_≠x0), then update the PTE(s), as described in [12.1.2.1\. Supervisor Memory-Management Fence Instruction](#sfence.vma), unless any inconsistencies are known to be benign. If any inconsistencies do exist, then the effect is the same as when SFENCE.VMA is used incorrectly: one of the translations will be chosen, but the choice is unpredictable. If an implementation chooses to use a NAPOT PTE (or cached version thereof), it might not consult the PTE directly specified by the algorithm in [12.1.3.2\. Virtual Address Translation Process](#sv32algorithm) at all. Therefore, the D and A bits may not be identical across all mappings of the same address range even in typical use cases The operating system must query all NAPOT aliases of a page to determine whether that page has been accessed and/or is dirty. If the OS manually sets the A and/or D bits for a page, it is recommended that the OS also set the A and/or D bits for other NAPOT aliases as appropriate in order to avoid unnecessary traps. Just as with normal PTEs, TLBs are permitted to cache NAPOT PTEs whose V (Valid) bit is clear. Depending on need, the NAPOT scheme may be extended to other intermediate page sizes and/or to other levels of the page table in the future. The encoding is designed to accommodate other NAPOT sizes should that need arise. For example: \_\_ i _pte_._ppn_\[_i_\] Description _pte_._napot\_bits_ 00000…​11…​ x xxxx xxx1 x xxxx xx10 x xxxx x100 x xxxx 1000 x xxx1 0000…​ x xxxx xxx1 x xxxx xx10…​ 8 KiB contiguous region16 KiB contiguous region32 KiB contiguous region64 KiB contiguous region128 KiB contiguous region…​4 MiB contiguous region8 MiB contiguous region…​ 12345…​12…​ In such a case, an implementation may or may not support all options. The discoverability mechanism for this extension would be extended to allow system software to determine which sizes are supported. Other sizes may remain deliberately excluded, so that PPN bits not being used to indicate a valid NAPOT region size (e.g., the least-significant bit of _pte_._ppn_\[_i_\]) may be repurposed for other uses in the future. However, in case finer-grained intermediate page size support proves not to be useful, we have chosen to standardize only 64 KiB support as a first step. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If the hypervisor extension is also implemented,Svnapot is also supported in G-stage translation. ### [](#svpbmt)12.1.8\. "Svpbmt" Extension for Page-Based Memory Types, Version 1.0 In Sv39, Sv48, and Sv57, bits 62-61 of a leaf page table entry indicate the use of page-based memory types that override the PMA(s) for the associated memory pages. The encoding for the PBMT bits is captured in[Table 8](#pbmt). The Svpbmt extension depends on the Sv39 extension. __Table 8\. Encodings for PBMT field in Sv39, Sv48, and Sv57 PTEs.__ | Mode | Value | Requested Memory Attributes | | --------- | ----- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | PMANCIO\- | 0123 | NoneNon-cacheable, idempotent, weakly-ordered (RVWMO), main memoryNon-cacheable, non-idempotent, strongly-ordered (I/O ordering), I/O _Reserved for future standard use_ | Implementations may override additional PMAs not explicitly listed in[Table 8](#pbmt).For example, to be consistent with the characteristics of a typical I/O region, a misaligned memory access to a page with PBMT=IO might raise an exception, even if the underlying region were main memory and the same access would have succeeded for PBMT=PMA. | | Future extensions may provide more and/or finer-grained control over which PMAs can be overridden. | | ----------------------------------------------------------------------------------------------------- | For non-leaf PTEs, bits 62-61 are reserved for future standard use.Until their use is defined by a standard extension, they must be cleared by software for forward compatibility, or else a page-fault exception is raised. For leaf PTEs, setting bits 62-61 to the value 3 is reserved for future standard use.Until this value is defined by a standard extension, using this reserved value in a leaf PTE raises a page-fault exception. When PBMT settings override a main memory page into I/O or vice versa,memory accesses to such pages obey the memory ordering rules of the final effective attribute, as follows. If the underlying physical memory attribute for a page is I/O, and the page has PBMT=NC, then accesses to that page obey RVWMO. However,accesses to such pages are considered to be _both_ I/O and main memory accesses for the purposes of FENCE, _.aq_, and _.rl_. If the underlying physical memory attribute for a page is main memory, and the page has PBMT=IO, thenaccesses to that page obey strong channel 0 I/O ordering rules. However,accesses to such pages are considered to be _both_ I/O and main memory accesses for the purposes of FENCE, _.aq_, and _.rl_. | | A device driver written to rely on I/O strong ordering rules will not operate correctly if the address range is mapped with PBMT=NC. As such, this configuration is discouraged. It will often still be useful to map physical I/O regions using PBMT=NC so that write combining and speculative accesses can be performed. Such optimizations will likely improve performance when applied with adequate care. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | When Svpbmt is used with non-zero PBMT encodings, it is possible for multiple virtual aliases of the same physical page to exist simultaneously with different memory attributes. It is also possible for a U-mode or S-mode mapping through a PTE with Svpbmt enabled to observe different memory attributes for a given region of physical memory than a concurrent access to the same page performed by M-mode or when MODE=Bare. In such cases, the behaviors dictated by the attributes (including coherence, which is otherwise unaffected) may be violated. Accessing the same location using different attributes that are both non-cacheable (e.g., NC and IO) does not cause loss of coherence, butmight result in weaker memory ordering than the stricter attribute ordinarily guarantees. Executing a`fence iorw, iorw` instruction between such accesses suffices to prevent loss of memory ordering. Accessing the same location using different cacheability attributesmay cause loss of coherence. Executing the following sequence between such accessesprevents both loss of coherence and loss of memory ordering: `fence iorw, iorw`, followed by `cbo.flush` to an address of that location, followed by a `fence iorw, iorw`. | | It follows that, if the same location might later be referenced using the original attributes, then this sequence must be repeated beforehand. In certain cases, a weaker sequence might suffice to prevent loss of coherence. These situations will be detailed following the forthcoming formalization of the interaction of the RVWMO memory model with the instructions in the Zicbom extension. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When two-stage address translation is enabled within the H extension, the page-based memory types are also applied in two stages. First,if `hgatp`.MODE is not equal to zero, non-zero G-stage PTE PBMT bits override the attributes in the PMA to produce an intermediate set of attributes. Otherwise, the PMAs serve as the intermediate attributes. Second,if `vsatp`.MODE is not equal to zero, non-zero VS-stage PTE PBMT bits override the intermediate attributes to produce the final set of attributes used by accesses to the page in question. Otherwise, the intermediate attributes are used as the final set of attributes. | | These final attributes apply to implicit and explicit accesses that are subject to both stages of address translation. For accesses that are not subject to the first stage of address translation, e.g. VS-stage page-table accesses, the intermediate attributes apply instead. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#svinval)12.1.9\. "Svinval" Extension for Fine-Grained Address-Translation Cache Invalidation, Version 1.0 The Svinval extension splits SFENCE.VMA, HFENCE.VVMA, and HFENCE.GVMA instructions into finer-grained invalidation and ordering operationsthat can be more efficiently batched or pipelined on certain classes of high-performance implementation. ![svg](_images/svg-03c79275ea03b9beabaf26303346b02d6ee9891e.svg) The SINVAL.VMA instruction invalidates any address-translation cache entries that an SFENCE.VMA instruction with the same values of _rs1_ and_rs2_ would invalidate.However, unlike SFENCE.VMA, SINVAL.VMA instructions are only ordered with respect to SFENCE.VMA, SFENCE.W.INVAL, and SFENCE.INVAL.IR instructions as defined below. ![svg](_images/svg-28f9eec40a4604917037e4435f14062f317554d2.svg) ![svg](_images/svg-0cdda65c316d98d1125b620c96b7f7fdc517d1a2.svg) The SFENCE.W.INVAL instruction guarantees that any previous stores already visible to the current RISC-V hart are ordered before subsequent SINVAL.VMA instructions executed by the same hart.The SFENCE.INVAL.IR instruction guarantees that any previous SINVAL.VMA instructions executed by the current hart are ordered before subsequent implicit references by that hart to the memory-management data structures. When executed in order (but not necessarily consecutively) by a single hart, the sequence SFENCE.W.INVAL, SINVAL.VMA, and SFENCE.INVAL.IR has the same effect as a hypothetical SFENCE.VMA instruction in which: * the values of _rs1_ and _rs2_ for the SFENCE.VMA are the same as those used in the SINVAL.VMA, * reads and writes prior to the SFENCE.W.INVAL are considered to be those prior to the SFENCE.VMA, and * reads and writes following the SFENCE.INVAL.IR are considered to be those subsequent to the SFENCE.VMA. ![svg](_images/svg-a70bc08a5f47d87340d758f102dcf52a71ebd8d6.svg) ![svg](_images/svg-792388fd5f74d87efe04f819cf8c98fb726d2a6f.svg) If the hypervisor extension is implemented, the Svinval extension also provides two additional instructions: HINVAL.VVMA and HINVAL.GVMA.These have the same semantics as SINVAL.VMA, except that they combine with SFENCE.W.INVAL and SFENCE.INVAL.IR to replace HFENCE.VVMA and HFENCE.GVMA, respectively, instead of SFENCE.VMA. In addition,HINVAL.GVMA uses VMIDs instead of ASIDs. SINVAL.VMA, HINVAL.VVMA, and HINVAL.GVMA require the same permissions and raise the same exceptions as SFENCE.VMA, HFENCE.VVMA, and HFENCE.GVMA, respectively.In particular, an attempt to execute any of these instructions in U-mode always raises an illegal-instruction exception. An attempt to execute SINVAL.VMA or HINVAL.GVMA in S-mode or HS-mode when`mstatus`.TVM=1 also raises an illegal-instruction exception. An attempt to execute HINVAL.VVMA or HINVAL.GVMA in VS-mode or VU-mode, or to execute SINVAL.VMA in VU-mode, raises a virtual-instruction exception. When `hstatus`.VTVM=1, an attempt to execute SINVAL.VMA in VS-mode also raises a virtual-instruction exception. Attempting to execute SFENCE.W.INVAL or SFENCE.INVAL.IR in U-moderaises an illegal-instruction exception. Doing so in VU-mode raises a virtual-instruction exception. SFENCE.W.INVAL and SFENCE.INVAL.IR are unaffected by the `mstatus`.TVM and`hstatus`.VTVM fields and hence are always permitted in S-mode and VS-mode. | | SFENCE.W.INVAL and SFENCE.INVAL.IR instructions do not need to be trapped when mstatus.TVM=1 or when hstatus.VTVM=1, as they only have ordering effects but no visible side effects. Trapping of the SINVAL.VMA instruction is sufficient to enable emulation of the intended overall TLB maintenance functionality. In typical usage, software will invalidate a range of virtual addresses in the address-translation caches by executing an SFENCE.W.INVAL instruction, executing a series of SINVAL.VMA, HINVAL.VVMA, or HINVAL.GVMA instructions to the addresses (and optionally ASIDs or VMIDs) in question, and then executing an SFENCE.INVAL.IR instruction. High-performance implementations will be able to pipeline the address-translation cache invalidation operations, and will defer any pipeline stalls or other memory ordering enforcement until an SFENCE.W.INVAL, SFENCE.INVAL.IR, SFENCE.VMA, HFENCE.GVMA, or HFENCE.VVMA instruction is executed. Simpler implementations may implement SINVAL.VMA, HINVAL.VVMA, and HINVAL.GVMA identically to SFENCE.VMA, HFENCE.VVMA, and HFENCE.GVMA, respectively, while implementing SFENCE.W.INVAL and SFENCE.INVAL.IR instructions as no-ops. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#sec:svadu)12.1.10\. "Svadu" Extension for Hardware Updating of A/D Bits, Version 1.0 The Svadu extension adds support and CSR controls for hardware updating of PTE A/D bits. If the Svadu extension is implemented, the `menvcfg`.ADUE field is writable. If the hypervisor extension is additionally implemented, the `henvcfg`.ADUE field is also writable.See [Machine Environment Configuration (menvcfg) Register](machine.html#sec:menvcfg) and [Hypervisor Environment Configuration Register (henvcfg)](hypervisor.html#sec:henvcfg) for the definitions of those fields. [12.1.3.1\. Addressing and Memory Protection](#translation) defines the semantics of hardware updating of A/D bits.When hardware updating of A/D bits is disabled, the Svade extension, which mandates exceptions when A/D bits need be set, instead takes effect.The Svade extension is also defined in [12.1.3.1\. Addressing and Memory Protection](#translation). ### [](#sec:svvptc)12.1.11\. "Svvptc" Extension for Obviating Memory-Management Instructions after Marking PTEs Valid, Version 1.0 When the Svvptc extension is implemented, explicit stores by a hart that update the Valid bit of leaf and/or non-leaf PTEs from 0 to 1 and are visible to a hart will eventually become visible within a bounded timeframe to subsequent implicit accesses by that hart to such PTEs. | | Svvptc relieves an operating system from executing certain memory-management instructions, such as SFENCE.VMA or SINVAL.VMA, which would normally be used to synchronize the hart’s address-translation caches when a memory-resident PTE is changed from Invalid to Valid. Synchronizing the hart’s address-translation caches with other forms of updates to a memory-resident PTE, including when a PTE is changed from Valid to Invalid, requires the use of suitable memory-management instructions. Svvptc guarantees that a change to a PTE from Invalid to Valid is made visible within a bounded time, thereby making the execution of these memory-management instructions redundant. The performance benefit of eliding these instructions outweighs the cost of an occasional gratuitous additional page fault that may occur. Depending on the microarchitecture, some possible ways to facilitate implementation of Svvptc include: not having any address-translation caches, not storing Invalid PTEs in the address-translation caches, automatically evicting Invalid PTEs using a bounded timer, or making address-translation caches coherent with store instructions that modify PTEs. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#sec:svrsw60t59b)12.1.12\. "Svrsw60t59b" Extension for PTE Reserved-for-Software Bits 60-59, Version 1.0 If the Svrsw60t59b extension is implemented, then bits 60-59 of the page table entries (PTEs) are reserved for use by supervisor software and are ignored by the implementation. If the Hypervisor (H) extension is also implemented, then bits 60-59 of the G-stage PTEs are reserved for use by supervisor software and are ignored by the implementation. The Svrsw60t59b extension depends on Sv39. | | Operating systems frequently use reserved bits within PTEs to store metadata for advanced memory management features. Embedding these metadata bits directly within the PTEs allows for fast access with minimal overhead, avoiding costly lookups in auxiliary data structures. By default, Sv39 and Sv39x4 require a page fault and a guest-page fault exception, respectively, to be raised if bits 60–59 are not zero. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#ssqosid)12.1.13\. "Ssqosid" Extension for Quality-of-Service (QoS) Identifiers, Version 1.0 Quality of Service (QoS) is defined as the minimal end-to-end performance guaranteed in advance by a service level agreement (SLA) to a workload. Performance metrics might include measures such as instructions per cycle (IPC), latency of service, etc. When multiple workloads execute concurrently on modern processors—equipped with large core counts, multiple cache hierarchies, and multiple memory controllers— the performance of any given workload becomes less deterministic, or even non-deterministic, due to shared resource contention. To manage performance variability, system software needs resource allocation and monitoring capabilities. These capabilities allow for the reservation of resources like cache and bandwidth, thus meeting individual performance targets while minimizing interference. For resource management, hardware should provide monitoring features that allow system software to profile workload resource consumption and allocate resources accordingly. To facilitate this, the QoS Identifiers extension (Ssqosid) introduces the`srmcfg` register, which configures a hart with two identifiers: a Resource Control ID (`RCID`) and a Monitoring Counter ID (`MCID`). These identifiers accompany each request issued by the hart to shared resource controllers. Additional metadata, like the nature of the memory access and the ID of the originating supervisor domain, can accompany `RCID` and `MCID`. Resource controllers may use this metadata for differentiated service such as a different capacity allocation for code storage vs. data storage. Resource controllers can use this data for security policies such as not exposing statistics of one security domain to another. These identifiers are crucial for the RISC-V Capacity and Bandwidth Controller QoS Register Interface (CBQRI) specification, which provides methods for setting resource usage limits and monitoring resource consumption. The `RCID`controls resource allocations, while the `MCID` is used for tracking resource usage. | | The Ssqosid extension does not require that S-mode mode be implemented. | | -------------------------------------------------------------------------- | #### [](#12-1-13-1-supervisor-resource-management-configuration-srmcfg-register)12.1.13.1\. Supervisor Resource Management Configuration (`srmcfg`) register The `srmcfg` register is an SXLEN-bit read/write register used to configure a Resource Control ID (`RCID`) and a Monitoring Counter ID (`MCID`). Both `RCID`and `MCID` are WARL fields. The register is formatted as shown in [Figure 29](#SRMCFG64)when SXLEN=64 and [Figure 30](#SRMCFG32) when SXLEN=32. The `RCID` and `MCID` accompany each request made by the hart to shared resource controllers. The `RCID` is used to determine the resource allocations (e.g., cache occupancy limits, memory bandwidth limits, etc.) to enforce. The `MCID`is used to identify a counter to monitor resource usage. ![Supervisor Resource Management Configuration (`srmcfg`) register for SXLEN=64](_images/diag-d5884a89e022bf008ca81070262a1f03c154befa.svg) Figure 29\. Supervisor Resource Management Configuration (`srmcfg`) register for SXLEN=64 ![Supervisor Resource Management Configuration (`srmcfg`) register for SXLEN=32](_images/diag-126fd69e59ba4cdba93a7ecfe8c76bcaf1aaacb8.svg) Figure 30\. Supervisor Resource Management Configuration (`srmcfg`) register for SXLEN=32 The `RCID` and `MCID` configured in the `srmcfg` CSR apply to all privilege modes of software execution on that hart by default, but this behavior may be overridden by future extensions. If extension Smstateen is implemented together with Ssqosid, then Ssqosid also requires the SRMCFG bit in `mstateen0` to be implemented. If `mstateen0`.SRMCFG is 0, attempts to access `srmcfg` in privilege modes less privileged than M-mode raise an illegal-instruction exception. If `mstateen0`.SRMCFG is 1 or if extension Smstateen is not implemented, attempts to access `srmcfg` when `V=1` raise a virtual-instruction exception. | | A reset value of 0 is suggested for the RCID field matching resource controllers' default behavior of associating all capacity with RCID=0. TheMCID reset value does not affect functionality and may be implementation-defined. Typically, fewer bits are allocated for RCID (e.g., to support tens of RCIDs) than for MCID (e.g., to support hundreds of MCIDs). A common RCID is usually used to group apps or VMs, pooling resource allocations to meet collective SLAs. If an SLA breach occurs, unique MCIDs enable granular monitoring, aiding decisions on resource adjustment, associating a different RCID with a subset of members, or migrating members to other machines. The larger pool of MCIDs speeds up this analysis. The RCID and MCID in srmcfg apply across all privilege levels on the hart. Typically, higher-privilege modes don’t modify srmcfg, as they often serve lower-privileged tasks. If differentiation is needed, higher privilege code can update srmcfg and restore it before returning to a lower privilege level. In VM environments, hypervisors usually manage resource allocations, keeping the Guest OS out of QoS flows. If needed, the hypervisor can virtualizesrmcfg CSR for a VM using the virtual-instruction exceptions triggered upon Guest access. If the direct selection of RCID and MCID by the VM becomes common and emulation overhead is an issue, future extensions may allow VS-mode to use a selector for a hypervisor-configured set of CSRs holding RCID andMCID values designated for that Guest OS use. During context switches, the supervisor may choose to execute with the srmcfgof the outgoing context to attribute the execution to it. Prior to restoring the new context, it switches to the new VM’s srmcfg. The supervisor can also use a separate configuration for execution not to be attributed to either contexts. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | 18.1. Pointer Masking Extensions, Version 1.0.0 ==================== ## [](#Zpm)18.1\. Pointer Masking Extensions, Version 1.0.0 ### [](#18-1-1-introduction)18.1.1\. Introduction RISC-V Pointer Masking (PM) is a feature that, when enabled, causes the CPU to ignore the upper bits of the effective address (these terms will be defined more precisely in the Background section). This allows these bits to be used in whichever way the application chooses. The version of the extension being described here specifically targets **tag checks**: When an address is accessed, the tag stored in the masked bits can be compared against a range-based tag. This is used for dynamic safety checkers such as HWASAN \[[93](../biblio/bibliography.html#bib-hwasan)\]. Such tools can be applied in all privilege modes (U, S, and M). HWASAN leverages tags in the upper bits of the address to identify memory errors such as use-after-free or buffer overflow errors. By storing a **pointer tag** in the upper bits of the address and checking it against a **memory tag** stored in a side table, it can identify whether a pointer is pointing to a valid location. Doing this without hardware support introduces significant overheads since the pointer tag needs to be manually removed for every conventional memory operation. Pointer masking support reduces these overheads. Pointer masking only adds the ability to ignore pointer tags during regular memory accesses. The tag checks themselves can be implemented in software or hardware. If implemented in software, pointer masking still provides performance benefits since non-checked accesses do not need to transform the address before every memory access. Hardware implementations are expected to provide even larger benefits due to performing tag checks out-of-band and hardening security guarantees derived from these checks. We anticipate that future extensions may build on pointer masking to support this functionality in hardware. It is worth mentioning that while HWASAN is the primary use-case for the current pointer masking extension, a number of other hardware/software features may be implemented leveraging Pointer Masking. Some of these use cases include sandboxing, object type checks and garbage collection bits in runtime systems. Note that the current version of the spec does not explicitly address these use cases, but future extensions may build on it to do so. While we describe the high-level concepts of pointer masking as if it was a single extension, it is, in reality, a family of extensions that implementations or profiles may choose to individually include or exclude (see [18.1.2.7\. Pointer Masking Extensions](#pointer-mask-extensions)). ### [](#18-1-2-background)18.1.2\. Background #### [](#definitions)18.1.2.1\. Definitions We now define basic terms. Note that these rely on the definition of an “ignore” transformation, which is defined in [18.1.2.2\. The “Ignore” Transformation](#sec-ignore-transform). * **Effective address (as defined in the RISC-V Base ISA):** A load/store effective address sent to the memory subsystem (e.g., as generated during the execution of load/store instructions). This does not include addresses corresponding to implicit accesses, such as page-table walks. * **Masked bits:** The upper PMLEN bits of an address, where PMLEN is a configurable parameter. We will use PMLEN consistently throughout this document to refer to this parameter. * **Transformed address:** An effective address after the ignore transformation has been applied. * **Address translation mode:** The MODE of the currently active address translation scheme as defined in the RISC-V privileged specification. This could, for example, refer to Bare, Sv39, Sv48, and Sv57\. In accordance with the privileged specification, non-Bare translation modes are referred to as virtual-memory schemes. For the purpose of this specification, M-mode translation is treated as equivalent to Bare. * **Address validity:** The RISC-V privileged spec defines validity of addresses based on the address translation mode that is currently in use (e.g., Sv57, Sv48, Sv39, etc.). For a virtual address to be valid, all bits in the unused portion of the address must be the same as the Most Significant Bit (MSB) of the used portion. For example, when page-based 48-bit virtual memory (Sv48) is used, load/store effective addresses, which are 64 bits, must have bits 63–48 all set to bit 47, or else a page-fault exception will occur. For physical addresses, validity means that bits XLEN-1 to PABITS are zero, where PABITS is the number of physical address bits supported by the processor. * **NVBITS:** The upper bits within a virtual address that have no effect on addressing memory and are only used for validity checks. These bits depend on the currently active address translation mode. For example, in Sv48, these are bits 63-48. * **VBITS:** The bits within a virtual address that affect which memory is addressed. These are the bits of an address which are used to index into page tables. #### [](#sec-ignore-transform)18.1.2.2\. The “Ignore” Transformation The ignore transformation differs depending on whether it applies to a virtual or physical address. For virtual addresses, it replaces the upper PMLEN bits with the sign extension of the PMLEN+1st bit. "Ignore" Transformation for virtual addresses, expressed in Verilog code. ```none transformed_effective_address = {{PMLEN{effective_address[XLEN-PMLEN-1]}}, effective_address[XLEN-PMLEN-1:0]} ``` | | If PMLEN is less than or equal to NVBITS for the largest supported address translation mode on a given architecture, this is equivalent to ignoring a subset of NVBITS. This enables cheap implementations that modify validity checks in the CPU instead of performing the sign extension. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When applied to a physical address, including guest-physical addresses (i.e., all cases except when the active satp register’s MODE field != Bare), the ignore transformation replaces the upper PMLEN bits with 0\. This includes both the case of running in M-mode and running in other privilege modes with Bare address translation mode. "Ignore" Transformation for physical addresses, expressed in Verilog code. ```none transformed_effective_address = {{PMLEN{0}}, effective_address[XLEN-PMLEN-1:0]} ``` | | This definition is consistent with the way that RISC-V already handles physical and virtual addresses differently. While the unused upper bits of virtual addresses are the sign-extension of the used bits (see the definition of "address validity" in [18.1.2.1\. Definitions](#definitions)), the equivalent bits in physical addresses are zero-extended. This is necessary due to their interactions with other mechanisms such as Physical Memory Protection (PMP). | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When pointer masking is enabled, the ignore transformation will be applied to every explicit memory access (e.g., loads/stores, atomics operations, and floating point loads/stores). The transformation **does not** apply to implicit accesses such as page-table walks or instruction fetches. The set of accesses that pointer masking applies to is described in [18.1.2.6\. Memory Accesses Subject to Pointer Masking](#memory-accesses-subject). | | Pointer masking does not change the underlying address generation logic or permission checks. Under a fixed address translation mode, it is semantically equivalent to replacing a subset of instructions (e.g., loads and stores) with an instruction sequence that applies the ignore operation to the target address of this instruction and then applies the instruction to the transformed address. References to address translation and other implementation details in the text are primarily to explain design decisions and common implementation patterns. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Note that pointer masking is purely an arithmetic operation on the address that makes no assumption about the meaning of the addresses it is applied to. Pointer masking with the same value of PMLEN always has the same effect for the same type of address (virtual or physical). This ensures that code that relies on pointer masking does not need to be aware of the environment it runs in once pointer masking has been enabled, as long as the value of PMLEN is known, and whether or not addresses are virtual or physical. For example, the same application or library code can run in user mode, supervisor mode or M-mode (with different address translation modes) without modification. | | A common scenario for such code is that addresses are generated by mmap system calls. This abstracts away the details of the underlying address translation mode from the application code. Software therefore needs to be aware of the value of PMLEN to ensure that its minimally required number of tag bits is supported. [18.1.2.4\. Determining the Value of PMLEN](#determining-the-value-of-pmlen) covers how this value is derived. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#18-1-2-3-example)18.1.2.3\. Example [Table 1](#pm-example) shows an example of the pointer masking transformation on a virtual address when PM is enabled for RV64 under Sv57 (PMLEN=7). __Table 1\. Example of PM address translation for RV64 under Sv57__ | Page-based profile | Sv57 on RV64 | | --------------------------------------------------- | ------------------------------------------------------------------------------- | | Effective Address | 0xABFFFFFF12345678NVBITS\[1010101\] VBITS\[11111111111111111111111110001…​000\] | | PMLEN | 7 | | Mask | 0x01FFFFFFFFFFFFFFNVBITS\[0000000\] VBITS\[11111111111111111111111111111…​111\] | | PMLEN+1st bit from the top (i.e., bit XLEN-PMLEN-1) | 1 | | Transformed effective address | 0xFFFFFFFF12345678NVBITS\[1111111\] VBITS\[11111111111111111111111110001…​000\] | If the address was a physical address rather than a virtual address with Sv57, the transformed address with PMLEN=7 would be 0x1FFFFFF12345678. #### [](#determining-the-value-of-pmlen)18.1.2.4\. Determining the Value of PMLEN From an implementation perspective, ignoring bits is deeply connected to the maximum virtual and physical address space supported by the processor (e.g., Bare, Sv48, Sv57). In particular, applying the above transformation is cheap if it covers only bits that are not used by **any** supported address translation mode (as it is equivalent to switching off validity checks). Masking NVBITS beyond those bits is more expensive as it requires ignoring them in the TLB tag, and even more expensive if the masked bits extend into the VBITS portion of the address (as it requires performing the actual sign extension). Similarly, when running in Bare or M mode, it is common for implementations to not use a particular number of bits at the top of the physical address range and fix them to zero. Applying the ignore transformation to those bits is cheap as well, since it will result in a valid physical address with all the upper bits fixed to 0. The current standard only supports PMLEN=XLEN-48 (i.e., PMLEN=16 in RV64) and PMLEN=XLEN-57 (i.e., PMLEN=7 in RV64). A setting has been reserved to potentially support other values of PMLEN in future standards. In such future standards, different supported values of PMLEN may be defined for each privilege mode (U/VU, S/HS, and M). | | Future versions of the pointer masking extension may introduce the ability to freely configure the value of PMLEN. The current extension does not define the behavior if PMLEN was different from the values defined above. In particular, there is no guarantee that a future pointer masking extension would define the ignore operation in the same way for those values of PMLEN. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#pointer-masking-privilege-modes)18.1.2.5\. Pointer Masking and Privilege Modes Pointer masking is controlled separately for different privilege modes. The subset of supported privilege modes is determined by the set of supported pointer masking extensions. Different privilege modes may have different pointer masking settings active simultaneously and the hardware will automatically apply the pointer masking settings of the currently active privilege mode. A privilege mode’s pointer masking setting is configured by bits in configuration registers of the next-higher privilege mode. Note that the pointer masking setting that is applied only depends on the active privilege mode, not on the address that is being masked. Some operating systems (e.g., Linux) may use certain bits in the address to disambiguate between different types of addresses (e.g., kernel and user-mode addresses). Pointer masking _does not_ take these semantics into account and is purely an arithmetic operation on the address it is given. | | Linux places kernel addresses in the upper half of the address space and user addresses in the lower half of the address space. As such, the MSB is often used to identify the type of a particular address. With pointer masking enabled, this role is now played by bit XLEN-PMLEN-1 and code that checks whether a pointer is a kernel or a user address needs to inspect this bit instead. For backward compatibility, it may be desirable that the MSB still indicates whether an address is a user or a kernel address. An operating system’s ABI may mandate this, but it does not affect the pointer masking mechanism itself. For example, the Linux ABI may choose to mandate that the MSB is not used for tagging and replicates bit XLEN-PMLEN-1 bit (note that for such a mechanism to be secure, the kernel needs to check the MSB of any user mode-supplied address and ensure that this invariant holds before using it; alternatively, it can apply the transformation from Listing 1 or 2 to ensure that the MSB is set to the correct value). | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#memory-accesses-subject)18.1.2.6\. Memory Accesses Subject to Pointer Masking Pointer masking applies to all explicit memory accesses. Currently, in the Base and Privileged ISAs, these are: * **Base Instruction Set**: LB, LH, LW, LBU, LHU, LWU, LD, SB, SH, SW, SD. * **Atomics**: All instructions in RV32A and RV64A. * **Floating Point**: FLW, FLD, FLQ, FSW, FSD, FSQ. * **Compressed**: All instructions mapping to any of the above, and C.LWSP, C.LDSP, C.FLWSP, C.FLDSP, C.SWSP, C.SDSP, C.FSWSP, C.FSDSP. * **Hypervisor Extension**: HLV.\*, HSV.\* (in some cases; see [Hypervisor Status (hstatus) Register](hypervisor.html#sec:hstatus)). * **Cache Management Operations**: All instructions in Zicbom, Zicbop and Zicboz. * **Vector Extension**: All vector load and store instructions in the ratified RVV 1.0 spec. * **Zicfiss Extension**: SSPUSH, C.SSPUSH, SSPOPCHK, C.SSPOPCHK, SSAMOSWAP.W/D. * **Assorted**: FENCE, FENCE.I (if the currently unused address fields become enabled in the future). | | This list will grow over time as new extensions introduce new instructions that perform explicit memory accesses. | | -------------------------------------------------------------------------------------------------------------------- | For other extensions, pointer masking applies to all explicit memory accesses by default. Future extensions may add specific language to indicate whether particular accesses are or are not included in pointer masking. | | It is worth noting that pointer masking is not applied to SFENCE.\*, HFENCE.\*, SINVAL.\*, or HINVAL.\*. When such an operation is invoked, it is the responsibility of the software to provide the correct address. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | MPRV and SPVP affect pointer masking as well, causing the pointer masking settings of the effective privilege mode to be applied. When MXR is in effect at the effective privilege mode where explicit memory access is performed, pointer masking does not apply. | | Note that this includes cases where page-based virtual memory is not in effect; i.e., although MXR has no effect on permissions checks when page-based virtual memory is not in effect, it is still used in determining whether or not pointer masking should be applied. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Cache Management Operations (CMOs) must respect and take into account pointer masking. Otherwise, a few serious security problems can appear, including: CBO.ZERO may work as a STORE operation. If pointer masking is not respected, it would be possible to write to memory bypassing the mask enforcement. If CMOs did not respect pointer masking, it would be possible to weaponize this in a side-channel attack. For example, U-mode would be able to flush a physical address (without masking) that it should not be permitted to. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Pointer masking only applies to accesses generated by instructions on the CPU (including CPU extensions such as an FPU). E.g., it does not apply to accesses generated by page-table walks, the IOMMU, or devices. | | Pointer Masking does not apply to DMA controllers and other devices. It is therefore the responsibility of the software to manually untag these addresses. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | Misaligned accesses are supported, subject to the same limitations as in the absence of pointer masking. The behavior is identical to applying the pointer masking transformation to every constituent aligned memory access. In other words, the accessed bytes should be identical to the bytes that would be accessed if the pointer masking transformation was individually applied to every byte of the access without pointer masking. This ensures that both hardware implementations and emulation of misaligned accesses in M-mode behave the same way, and that the M-mode implementation is identical whether or not pointer masking is enabled (e.g., such an implementation may leverage MPRV to apply the correct privilege mode’s pointer masking setting). No pointer masking operations are applied when software reads/writes to CSRs, including those meant to hold addresses. If software stores tagged addresses into such CSRs, data load or data store operations based on those addresses are subject to pointer masking only if they are explicit ([18.1.2.6\. Memory Accesses Subject to Pointer Masking](#memory-accesses-subject)) and pointer masking is enabled for the privilege mode that performs the access. The implemented WARL width of CSRs is unaffected by pointer masking (e.g., if a CSR supports 52 bits of valid addresses and pointer masking is supported with PMLEN=16, the necessary number of WARL bits remains 52 independently of whether pointer masking is enabled or disabled). In contrast to software writes, pointer masking, when applicable, **is applied** for hardware writes to a CSR (e.g., when the hardware writes the transformed address to `stval` when taking an exception). Pointer masking is also applied, when applicable, to the memory access address when matching address triggers in debug. For example, software is free to write a tagged or untagged address to `stvec`, but on trap delivery (e.g., due to an exception or interrupt), pointer masking **will not be applied** to the address of the trap handler. However, when delivering an exception, the hardware applies pointer masking to any address written into `stval` if pointer masking is applicable to that address. | | The rationale for this choice is that delivering the additional bits may add overheads in some hardware implementations. Further, pointer masking is configured per privilege mode, so all trap handlers in supervisor mode would need to be careful to configure pointer masking the same way as user mode or manually unmask (which is expensive). | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#pointer-mask-extensions)18.1.2.7\. Pointer Masking Extensions Pointer masking refers to a number of separate extensions, all of which are privileged. This approach is used to capture optionality of pointer masking features. Profiles and implementations may choose to support an arbitrary subset of these extensions and must define valid ranges for their corresponding values of PMLEN. **Extensions**: * **Ssnpm**: A supervisor-level extension that provides pointer masking for the next lower privilege mode (U-mode), and for VS- and VU-modes if the H extension is present. See [Supervisor Environment Configuration (senvcfg) Register](supervisor.html#sec:senvcfg), [Hypervisor Environment Configuration Register (henvcfg)](hypervisor.html#sec:henvcfg), [Hypervisor Status (hstatus) Register](hypervisor.html#sec:hstatus), and [Interaction with Pointer Masking](hypervisor.html#pm-two-stage). * **Smnpm**: A machine-level extension that provides pointer masking for the next lower privilege mode (S/HS if S-mode is implemented, or U-mode otherwise). See [Machine Environment Configuration (menvcfg) Register](machine.html#sec:menvcfg). * **Smmpm**: A machine-level extension that provides pointer masking for M-mode. See [Machine Security Configuration (mseccfg) Register](machine.html#sec:mseccfg). In addition, the pointer masking standard defines two extensions that describe an execution environment but have no bearing on hardware implementations. These extensions are intended to be used in profile specifications where a User profile or a Supervisor profile can only reference User level or Supervisor level pointer masking functionality, and not the associated CSR controls that exist at a higher privilege level (i.e., in the execution environment). * **Sspm**: An extension that indicates that there is pointer-masking support available in supervisor mode, with some facility provided in the supervisor execution environment to control pointer masking. * **Supm**: An extension that indicates that there is pointer-masking support available in user mode, with some facility provided in the application execution environment to control pointer masking. The precise nature of these facilities is left to the respective execution environment. Pointer masking only applies to RV64\. In RV32, trying to enable pointer masking will result in an illegal WARL write and not update the pointer masking configuration bits (see [Machine Security Configuration (mseccfg) Register](machine.html#sec:mseccfg), [Machine Environment Configuration (menvcfg) Register](machine.html#sec:menvcfg), [Hypervisor Environment Configuration Register (henvcfg)](hypervisor.html#sec:henvcfg), and [Supervisor Environment Configuration (senvcfg) Register](supervisor.html#sec:senvcfg) for details). The same is the case on RV64 or larger systems when UXL/SXL/MXL is set to 1 for the corresponding privilege mode. Note that in RV32, the CSR bits introduced by pointer masking are still present, for compatibility between RV32 and larger systems with UXL/SXL/MXL set to 1\. Setting UXL/SXL/MXL to 1 will clear the corresponding pointer masking configuration bits. | | Note that setting UXL/SXL/MXL to 1 and back to 0 does not preserve the previous values of the PMM bits. This includes the case of entering an RV32 virtual machine from an RV64 hypervisor and returning. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | Future extensions may introduce additional CSRs to allow different privilege modes to modify their own pointer masking settings. This may be required for future use cases in managed runtime systems that are not currently addressed as part of this extension. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#number-of-masked-bits)18.1.2.8\. Number of Masked Bits As described in [18.1.2.4\. Determining the Value of PMLEN](#determining-the-value-of-pmlen), the supported values of PMLEN may depend on the effective privilege mode. The current standard only defines PMLEN=XLEN-48 and PMLEN=XLEN-57, but this assumption may be relaxed in future extensions and profiles. Trying to enable pointer masking in an unsupported scenario represents an illegal write to the corresponding pointer masking enable bit and follows WARL semantics. Future profiles may choose to define certain combinations of privilege modes and supported values of PMLEN as mandatory. | | An option that was considered but discarded was to allow implementations to set PMLEN depending on the active addressing mode. For example, PMLEN could be set to 16 for Sv48 and to 25 for Sv39\. However, having a single value of PMLEN (e.g., setting PMLEN to 16 for both Sv39 and Sv48 rather than 25) facilitates TLB implementations in designs that support Sv39 and Sv48 but not Sv57\. 16 bits are sufficient for current pointer masking use cases but allow for a TLB implementation that matches against the same number of virtual tag bits independently of whether it is running with Sv39 or Sv48\. However, if Sv57 is supported, tag matching may need to be conditional on the current address translation mode. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 13.1. "A" Extension for Atomic Instructions, Version 2.1 ==================== ## [](#13-1-a-extension-for-atomic-instructions-version-2-1)13.1\. "A" Extension for Atomic Instructions, Version 2.1 The atomic-instruction extension, named "A", contains instructions that atomically read-modify-write memory to support synchronization between multiple RISC-V harts running in the same memory space. The two forms of atomic instruction provided are load-reserved/store-conditional instructions and atomic fetch-and-op memory instructions. Both types of atomic instruction support various memory consistency orderings including unordered, acquire, release, and sequentially consistent semantics. These instructions allow RISC-V to support the RCsc memory consistency model. \[[18](../biblio/bibliography.html#bib-gharachorloo90memoryconsistency)\] | | After much debate, the language community and architecture community appear to have finally settled on release consistency as the standard memory consistency model and so the RISC-V atomic support is built around this model. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The A extension comprises instructions provided by the Zaamo and Zalrsc extensions. ### [](#13-1-1-specifying-ordering-of-atomic-instructions)13.1.1\. Specifying Ordering of Atomic Instructions The base RISC-V ISA has a relaxed memory model, with the FENCE instruction used to impose additional ordering constraints. The address space is divided by the execution environment into memory and I/O domains, and the FENCE instruction provides options to order accesses to one or both of these two address domains. To provide more efficient support for release consistency \[[18](../biblio/bibliography.html#bib-gharachorloo90memoryconsistency)\], each atomic instruction has two bits, _aq_ and _rl_, used to specify additional memory ordering constraints as viewed by other RISC-V harts. The bits order accesses to one of the two address domains, memory or I/O, depending on which address domain the atomic instruction is accessing. No ordering constraint is implied to accesses to the other domain, and a FENCE instruction should be used to order across both domains. If both bits are clear, no additional ordering constraints are imposed on the atomic memory operation. If only the _aq_ bit is set, the atomic memory operation is treated as an _acquire_ access, i.e., no following memory operations on this RISC-V hart can be observed to take place before the acquire memory operation. If only the _rl_ bit is set, the atomic memory operation is treated as a _release_ access, i.e., the release memory operation cannot be observed to take place before any earlier memory operations on this RISC-V hart. If both the _aq_ and _rl_bits are set, the atomic memory operation is _sequentially consistent_and cannot be observed to happen before any earlier memory operations or after any later memory operations in the same RISC-V hart and to the same address domain. ### [](#sec:lrsc)13.1.2\. "Zalrsc" Extension for Load-Reserved/Store-Conditional Instructions ![svg](_images/svg-07c82b51f2efd59a5908edda3bb001962afc56ff.svg) Complex atomic memory operations on a single memory word or doubleword are performed with the load-reserved (LR) and store-conditional (SC) instructions. LR.W loads a word from the address in _rs1_, places the sign-extended value in _rd_, and registers a _reservation set_—a set of bytes that subsumes the bytes in the addressed word. SC.W conditionally writes a word in _rs2_ to the address in _rs1_: the SC.W succeeds only if the reservation is still valid and the reservation set contains the bytes being written. If the SC.W succeeds, the instruction writes the word in _rs2_ to memory, and it writes zero to _rd_. If the SC.W fails, the instruction does not write to memory, and it writes a nonzero value to _rd_. No SC.W instruction shall retire unless it passes memory permission checks, but it is UNSPECIFIED whether any side effects of implicit address translation and protection memory accesses (such as setting a page-table entry D bit) occur on a failed SC.W. For the purposes of memory protection, a failed SC.W may be treated like a store. Regardless of success or failure, executing an SC.W instruction invalidates any reservation held by this hart. LR.D and SC.D act analogously on doublewords and are only available on RV64\. For RV64, LR.W and SC.W sign-extend the value placed in _rd_. | | Both compare-and-swap (CAS) and LR/SC can be used to build lock-free data structures. After extensive discussion, we opted for LR/SC for several reasons: 1) CAS suffers from the ABA problem, which LR/SC avoids because it monitors all writes to the address rather than only checking for changes in the data value; 2) CAS would also require a new integer instruction format to support three source operands (address, compare value, swap value) as well as a different memory system message format, which would complicate microarchitectures; 3) Furthermore, to avoid the ABA problem, other systems provide a double-wide CAS (DW-CAS) to allow a counter to be tested and incremented along with a data word. This requires reading five registers and writing two in one instruction, and also a new larger memory system message type, further complicating implementations; 4) LR/SC provides a more efficient implementation of many primitives as it only requires one load as opposed to two with CAS (one load before the CAS instruction to obtain a value for speculative computation, then a second load as part of the CAS instruction to check if value is unchanged before updating). The main disadvantage of LR/SC over CAS is livelock, which we avoid, under certain circumstances, with an architected guarantee of eventual forward progress as described below. Another concern is whether the influence of the current x86 architecture, with its DW-CAS, will complicate porting of synchronization libraries and other software that assumes DW-CAS is the basic machine primitive. A possible mitigating factor is the recent addition of transactional memory instructions to x86, which might cause a move away from DW-CAS. More generally, a multi-word atomic primitive is desirable, but there is still considerable debate about what form this should take, and guaranteeing forward progress adds complexity to a system. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The failure code with value 1 encodes an unspecified failure. Other failure codes are reserved at this time. Portable software should only assume the failure code will be non-zero. | | We reserve a failure code of 1 to mean ''unspecified'' so that simple implementations may return this value using the existing multiplexer required for the SLT/SLTU instructions. More specific failure codes might be defined in future versions or extensions to the ISA. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For LR and SC, the Zalrsc extension requires that the address held in _rs1_be naturally aligned to the size of the operand (i.e., eight-byte aligned for _doublewords_ and four-byte aligned for _words_). If the address is not naturally aligned, an address-misaligned exception or an access-fault exception will be generated. The access-fault exception can be generated for a memory access that would otherwise be able to complete except for the misalignment, if the misaligned access should not be emulated. | | Emulating misaligned LR/SC sequences is impractical in most systems. Misaligned LR/SC sequences also raise the possibility of accessing multiple reservation sets at once, which present definitions do not provide for. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | An implementation can register an arbitrarily large reservation set on each LR, provided the reservation set includes all bytes of the addressed data word or doubleword. An SC can only pair with the most recent LR in program order. An SC may succeed only if no store from another hart to the reservation set can be observed to have occurred between the LR and the SC, and if there is no other SC between the LR and itself in program order. An SC may succeed only if no write from a device other than a hart to the bytes accessed by the LR instruction can be observed to have occurred between the LR and SC. Note this LR might have had a different effective address and data size, but reserved the SC’s address as part of the reservation set. | | Following this model, in systems with memory translation, an SC is allowed to succeed if the earlier LR reserved the same location using an alias with a different virtual address, but is also allowed to fail if the virtual address is different. To accommodate legacy devices and buses, writes from devices other than RISC-V harts are only required to invalidate reservations when they overlap the bytes accessed by the LR. These writes are not required to invalidate the reservation when they access other bytes in the reservation set. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The SC must fail if the address is not within the reservation set of the most recent LR in program order. The SC must fail if a store to the reservation set from another hart can be observed to occur between the LR and SC. The SC must fail if a write from some other device to the bytes accessed by the LR can be observed to occur between the LR and SC.(If such a device writes the reservation set but does not write the bytes accessed by the LR, the SC may or may not fail.) An SC must fail if there is another SC (to any address) between the LR and the SC in program order. The precise statement of the atomicity requirements for successful LR/SC sequences is defined by the Atomicity Axiom in[RVWMO Memory Consistency Model](rvwmo.html). | | The platform should provide a means to determine the size and shape of the reservation set. A platform specification may constrain the size and shape of the reservation set. A store-conditional instruction to a scratch word of memory should be used to forcibly invalidate any existing load reservation: during a preemptive context switch, and if necessary when changing virtual to physical address mappings, such as when migrating pages that might contain an active reservation. The invalidation of a hart’s reservation when it executes an LR or SC imply that a hart can only hold one reservation at a time, and that an SC can only pair with the most recent LR, and LR with the next following SC, in program order. This is a restriction to the Atomicity Axiom in[RVWMO Memory Consistency Model](rvwmo.html) that ensures software runs correctly on expected common implementations that operate in this manner. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | An SC instruction can never be observed by another RISC-V hart before the LR instruction that established the reservation. | | The LR/SC sequence can be given acquire semantics by setting the _aq_ bit on the LR instruction. The LR/SC sequence can be given release semantics by by setting the _rl_ bit on the SC instruction. Assuming suitable mappings for other atomic operations, setting the_aq_ bit on the LR instruction, and setting the_rl_ bit on the SC instruction makes the LR/SC sequence sequentially consistent in the C++ memory\_order\_seq\_cstsense. Such a sequence does not act as a fence for ordering ordinary load and store instructions before and after the sequence. Specific instruction mappings for other C++ atomic operations, or stronger notions of "sequential consistency", may require both bits to be set on either or both of the LR or SC instruction. If neither bit is set on either LR or SC, the LR/SC sequence can be observed to occur before or after surrounding memory operations from the same RISC-V hart. This can be appropriate when the LR/SC sequence is used to implement a parallel reduction operation. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Software should not set the _rl_ bit on an LR instruction unless the_aq_ bit is also set, nor should software set the _aq_ bit on an SC instruction unless the _rl_ bit is also set. LR._rl_ and SC._aq_instructions are not guaranteed to provide any stronger ordering than those with both bits clear, but may result in lower performance. | | Sample code for compare-and-swap function using LR/SC. \# a0 holds address of memory location \# a1 holds expected value \# a2 holds desired value \# a0 holds return value, 0 if successful, !0 otherwise cas: lr.w t0, (a0) # Load original value. bne t0, a1, fail # Doesn't match, so fail. sc.w t0, a2, (a0) # Try to update. bnez t0, cas # Retry if store-conditional failed. li a0, 0 # Set return to success. jr ra # Return. fail: li a0, 1 # Set return to failure. jr ra # Return. LR/SC can be used to construct lock-free data structures. An example using LR/SC to implement a compare-and-swap function is shown in[Sample code for compare-and-swap function using LR/SC.](#cas). If inlined, compare-and-swap functionality need only take four instructions. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#sec:lrscseq)13.1.3\. Eventual Success of Store-Conditional Instructions The Zalrsc extension defines _constrained LR/SC loops_, which have the following properties: * The loop comprises only an LR/SC sequence and code to retry the sequence in the case of failure, and must comprise at most 16 instructions placed sequentially in memory. * An LR/SC sequence begins with an LR instruction and ends with an SC instruction. The dynamic code executed between the LR and SC instructions can only contain instructions from the base ''I'' instruction set, excluding loads, stores, backward jumps, taken backward branches, JALR, FENCE, and SYSTEM instructions. Compressed forms of the aforementioned ''I'' instructions in the C (hence Zca) and Zcb extensions are also permitted. * The code to retry a failing LR/SC sequence can contain backwards jumps and/or branches to repeat the LR/SC sequence, but otherwise has the same constraint as the code between the LR and SC. * The LR and SC addresses must lie within a memory region with the_LR/SC eventuality_ property. The execution environment is responsible for communicating which regions have this property. * The SC must be to the same effective address and of the same data size as the latest LR executed by the same hart. LR/SC sequences that do not lie within constrained LR/SC loops are_unconstrained_. Unconstrained LR/SC sequences might succeed on some attempts on some implementations, but might never succeed on other implementations. | | We restricted the length of LR/SC loops to fit within 64 contiguous instruction bytes in the base ISA to avoid undue restrictions on instruction cache and TLB size and associativity. Similarly, we disallowed other loads and stores within the loops to avoid restrictions on data-cache associativity in simple implementations that track the reservation within a private cache. The restrictions on branches and jumps limit the time that can be spent in the sequence. Floating-point operations and integer multiply/divide were disallowed to simplify the operating system’s emulation of these instructions on implementations lacking appropriate hardware support. Software is not forbidden from using unconstrained LR/SC sequences, but portable software must detect the case that the sequence repeatedly fails, then fall back to an alternate code sequence that does not rely on an unconstrained LR/SC sequence. Implementations are permitted to unconditionally fail any unconstrained LR/SC sequence. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | \[#norm:constrained\_lrsc\_forward\_progress\]#If a hart _H_ enters a constrained LR/SC loop, the execution environment must guarantee that one of the following events eventually occurs: * _H_ or some other hart executes a successful SC to the reservation set of the LR instruction in _H_'s constrained LR/SC loops. * Some other hart executes an unconditional store or AMO instruction to the reservation set of the LR instruction in _H_'s constrained LR/SC loop, or some other device in the system writes to that reservation set. * _H_ executes a branch or jump that exits the constrained LR/SC loop. * _H_ traps.# | | Note that these definitions permit an implementation to fail an SC instruction occasionally for any reason, provided the aforementioned guarantee is not violated. As a consequence of the eventuality guarantee, if some harts in an execution environment are executing constrained LR/SC loops, and no other harts or devices in the execution environment execute an unconditional store or AMO to that reservation set, then at least one hart will eventually exit its constrained LR/SC loop. By contrast, if other harts or devices continue to write to that reservation set, it is not guaranteed that any hart will exit its LR/SC loop. Loads and load-reserved instructions do not by themselves impede the progress of other harts' LR/SC sequences. We note this constraint implies, among other things, that loads and load-reserved instructions executed by other harts (possibly within the same core) cannot impede LR/SC progress indefinitely. For example, cache evictions caused by another hart sharing the cache cannot impede LR/SC progress indefinitely. Typically, this implies reservations are tracked independently of evictions from any shared cache. Similarly, cache misses caused by speculative execution within a hart cannot impede LR/SC progress indefinitely. These definitions admit the possibility that SC instructions may spuriously fail for implementation reasons, provided progress is eventually made. One advantage of CAS is that it guarantees that some hart eventually makes progress, whereas an LR/SC atomic sequence could livelock indefinitely on some systems. To avoid this concern, we added an architectural guarantee of livelock freedom for certain LR/SC sequences. Earlier versions of this specification imposed a stronger starvation-freedom guarantee. However, the weaker livelock-freedom guarantee is sufficient to implement the C11 and C++11 languages, and is substantially easier to provide in some microarchitectural styles. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#sec:amo)13.1.4\. "Zaamo" Extension for Atomic Memory Operations ![svg](_images/svg-93d33ebfe58a072e40d275d64c8be10c9b98811e.svg) The atomic memory operation (AMO) instructions perform read-modify-write operations for multiprocessor synchronization and are encoded with an R-type instruction format. These AMO instructions atomically load a data value from the address in _rs1_, place the value into register _rd_, apply a binary operator to the loaded value and the original value in_rs2_, then store the result back to the original address in _rs1_. AMOs can either operate on _doublewords_ (RV64 only) or _words_ in memory. For RV64, 32-bit AMOs always sign-extend the value placed in _rd_, and ignore the upper 32 bits of the original value of _rs2_. For AMOs, the Zaamo extension requires that the address held in _rs1_ be naturally aligned to the size of the operand (i.e., eight-byte aligned for _doublewords_ and four-byte aligned for _words_). If the address is not naturally aligned, an address-misaligned exception or an access-fault exception will be generated. The access-fault exception can be generated for a memory access that would otherwise be able to complete except for the misalignment, if the misaligned access should not be emulated. The misaligned atomicity granule PMA, defined in Volume II of this manual, optionally relaxes this alignment requirement. If present, the misaligned atomicity granule PMA specifies the size of a misaligned atomicity granule, a power-of-two number of bytes. The misaligned atomicity granule PMA applies only to AMOs, loads and stores defined in the base ISAs, and loads and stores of no more than XLEN bits defined in the F, D, and Q extensions, and compressed encodings thereof. For an instruction in that set, if all accessed bytes lie within the same misaligned atomicity granule, the instruction will not raise an exception for reasons of address alignment, and the instruction will give rise to only one memory operation for the purposes of RVWMO—​i.e., it will execute atomically. The operations supported are swap, integer add, bitwise AND, bitwise OR, bitwise XOR, and signed and unsigned integer maximum and minimum. Without ordering constraints, these AMOs can be used to implement parallel reduction operations, where typically the return value would be discarded by writing to `x0`. | | We provided fetch-and-op style atomic primitives as they scale to highly parallel systems better than LR/SC or CAS. A simple microarchitecture can implement AMOs using the LR/SC primitives, provided the implementation can guarantee the AMO eventually completes. More complex implementations might also implement AMOs at memory controllers, and can optimize away fetching the original value when the destination is x0. The set of AMOs was chosen to support the C11/C++11 atomic memory operations efficiently, and also to support parallel reductions in memory. Another use of AMOs is to provide atomic updates to memory-mapped device registers (e.g., setting, clearing, or toggling bits) in the I/O space. The Zaamo extension enables microcontroller class implementations to utilize atomic primitives from the AMO subset of the A extension. Typically such implementations do not have caches and thus may not be able to naturally support the LR/SC instructions provided by the Zalrsc extension. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | To help implement multiprocessor synchronization, the AMOs optionally provide release consistency semantics. If the _aq_ bit is set, then no later memory operations in this RISC-V hart can be observed to take place before the AMO. Conversely, if the _rl_ bit is set, then other RISC-V harts will not observe the AMO before memory accesses preceding the AMO in this RISC-V hart. Setting both the _aq_ and the _rl_ bit on an AMO makes the sequence sequentially consistent, meaning that it cannot be reordered with earlier or later memory operations from the same hart. | | The AMOs were designed to implement the C11 and C++11 memory models efficiently. Although the FENCE R, RW instruction suffices to implement the _acquire_ operation and FENCE RW, W suffices to implement _release_, both imply additional unnecessary ordering as compared to AMOs with the corresponding _aq_ or _rl_ bit set. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | An example code sequence for a critical section guarded by a test-and-test-and-set spinlock is shown in Example [Sample code for mutual exclusion. a0 contains the address of the lock.](#critical). Note the first AMO is marked _aq_ to order the lock acquisition before the critical section, and the second AMO is marked _rl_ to order the critical section before the lock relinquishment. Sample code for mutual exclusion. a0 contains the address of the lock. li t0, 1 # Initialize swap value. again: lw t1, (a0) # Check if lock is held. bnez t1, again # Retry if held. amoswap.w.aq t1, t0, (a0) # Attempt to acquire lock. bnez t1, again # Retry if held. \# ... \# Critical section. \# ... amoswap.w.rl x0, x0, (a0) # Release lock by storing 0. We recommend the use of the AMO Swap idiom shown in [Sample code for mutual exclusion. a0 contains the address of the lock.](#critical) for both lock acquire and release to simplify the implementation of speculative lock elision. \[[19](../biblio/bibliography.html#bib-rajwar:2001:sle)\] | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The instructions in the "A" extension can be used to provide sequentially consistent loads and stores, but this constrains hardware reordering of memory accesses more than necessary. A C++ sequentially consistent load can be implemented as an LR with _aq_ set. However, the LR/SC eventual success guarantee may slow down concurrent loads from the same effective address. A sequentially consistent store can be implemented as an AMOSWAP that writes the old value to x0 and has _rl_ set. However the superfluous load may impose ordering constraints that are unnecessary for this use case. Specific compilation conventions may require both the _aq_ and _rl_bits to be set in either or both the LR and AMOSWAP instructions. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 30.1. Bit Manipulation Extensions ==================== ## [](#bits)30.1\. Bit Manipulation Extensions The bit-manipulation (bitmanip) extension collection is comprised of several component extensions to the base RISC-V architecture that are intended to provide some combination of code-size reduction, performance improvement, and energy reduction. While the instructions are intended for general use, some instructions are more useful in certain domains than in others. Hence, several smaller bitmanip extensions are provided. Each of these smaller extensions is grouped by common function and use case, and each has its own Zb\*-extension name. Each bitmanip extension includes a group of several bitmanip instructions that have similar purposes and can often share the same logic. Some instructions are available in only one extension, while others are available in several. The instructions have mnemonics and encodings that are independent of the extensions in which they appear. Thus, when implementing extensions with overlapping instructions, there is no redundancy in logic or encoding. The bitmanip extensions are defined for RV32 and RV64. The bitmanip extension follows the convention in RV64 that _w_\-suffixed instructions (without a dot before the _w_) ignore the upper 32 bits of their inputs, operate on the least-significant 32 bits as signed values, and produce a 32-bit signed result that is sign-extended to XLEN. Bitmanip instructions with the suffix _.uw_ have one operand that is an unsigned 32-bit value that is extracted from the least-significant 32 bits of the specified register. Other than that, these perform full-XLEN operations. Bitmanip instructions with the suffixes _.b_, _.h_, and _.w_ only look at the least-significant 8 bits, 16 bits, and 32 bits of the input (respectively) and produce an XLEN-wide result that is sign-extended or zero-extended, based on the specific instruction. The bit-manipulation instructions comprise the following extensions: * Zba: [Address generation instructions](#zba) * Zbb: [Basic bit-manipulation](#zbb) * Zbs: [Single-bit instructions](#zbs) * Zbc: [Carry-less multiplication](#zbc) * Zbkb: [Bit-manipulation for Cryptography](#zbkb) * Zbkc: [Carry-less multiplication for Cryptography](#zbkc) * Zbkx: [Crossbar permutations](#zbkx) Below is a list of all of the instructions that are included in these extensions, along with their specific mapping: | RV32 | RV64 | Mnemonic | Instruction | Zbb | Zbkb | Zbc | Zbkc | | ---- | -------------------------- | ---------------------------------------------------- | -------------------------------------------------- | --- | ---- | --- | ---- | | ✓ | ✓ | andn _rd_, _rs1_, _rs2_ | [AND with inverted operand](#insns-andn) | ✓ | ✓ | | | | ✓ | ✓ | brev8 _rd_, _rs_ | [Reverse bits in bytes](#insns-brev8) | ✓ | | | | | ✓ | ✓ | clmul _rd_, _rs1_, _rs2_ | [Carry-less multiply (low-part)](#insns-clmul) | ✓ | ✓ | | | | ✓ | ✓ | clmulh _rd_, _rs1_, _rs2_ | [Carry-less multiply (high-part)](#insns-clmulh) | ✓ | ✓ | | | | ✓ | ✓ | clmulr _rd_, _rs1_, _rs2_ | [Carry-less multiply (reversed)](#insns-clmulr) | ✓ | | | | | ✓ | ✓ | clz _rd_, _rs_ | [Count leading zero bits](#insns-clz) | ✓ | | | | | ✓ | clzw _rd_, _rs_ | [Count leading zero bits in word](#insns-clzw) | ✓ | | | | | | ✓ | ✓ | cpop _rd_, _rs_ | [Count set bits](#insns-cpop) | ✓ | | | | | ✓ | cpopw _rd_, _rs_ | [Count set bits in word](#insns-cpopw) | ✓ | | | | | | ✓ | ✓ | ctz _rd_, _rs_ | [Count trailing zero bits](#insns-ctz) | ✓ | | | | | ✓ | ctzw _rd_, _rs_ | [Count trailing zero bits in word](#insns-ctzw) | ✓ | | | | | | ✓ | ✓ | max _rd_, _rs1_, _rs2_ | [Maximum](#insns-max) | ✓ | | | | | ✓ | ✓ | maxu _rd_, _rs1_, _rs2_ | [Unsigned maximum](#insns-maxu) | ✓ | | | | | ✓ | ✓ | min _rd_, _rs1_, _rs2_ | [Minimum](#insns-min) | ✓ | | | | | ✓ | ✓ | minu _rd_, _rs1_, _rs2_ | [Unsigned minimum](#insns-minu) | ✓ | | | | | ✓ | ✓ | orc.b _rd_, _rs_ | [Bitwise OR-Combine, byte granule](#insns-orc%5Fb) | ✓ | | | | | ✓ | ✓ | orn _rd_, _rs1_, _rs2_ | [OR with inverted operand](#insns-orn) | ✓ | ✓ | | | | ✓ | ✓ | pack _rd_, _rs1_, _rs2_ | [Pack low halves of registers](#insns-pack) | ✓ | | | | | ✓ | ✓ | packh _rd_, _rs1_, _rs2_ | [Pack low bytes of registers](#insns-packh) | ✓ | | | | | ✓ | packw _rd_, _rs1_, _rs2_ | [Pack low 16-bits of registers (RV64)](#insns-packw) | ✓ | | | | | | ✓ | ✓ | rev8 _rd_, _rs_ | [Byte-reverse register](#insns-rev8) | ✓ | ✓ | | | | ✓ | ✓ | rol _rd_, _rs1_, _rs2_ | [Rotate left (Register)](#insns-rol) | ✓ | ✓ | | | | ✓ | rolw _rd_, _rs1_, _rs2_ | [Rotate Left Word (Register)](#insns-rolw) | ✓ | ✓ | | | | | ✓ | ✓ | ror _rd_, _rs1_, _rs2_ | [Rotate right (Register)](#insns-ror) | ✓ | ✓ | | | | ✓ | ✓ | rori _rd_, _rs1_, _shamt_ | [Rotate right (Immediate)](#insns-rori) | ✓ | ✓ | | | | ✓ | roriw _rd_, _rs1_, _shamt_ | [Rotate right Word (Immediate)](#insns-roriw) | ✓ | ✓ | | | | | ✓ | rorw _rd_, _rs1_, _rs2_ | [Rotate right Word (Register)](#insns-rorw) | ✓ | ✓ | | | | | ✓ | ✓ | sext.b _rd_, _rs_ | [Sign-extend byte](#insns-sext%5Fb) | ✓ | | | | | ✓ | ✓ | sext.h _rd_, _rs_ | [Sign-extend halfword](#insns-sext%5Fh) | ✓ | | | | | ✓ | unzip _rd_, _rs_ | [Bit deinterleave](#insns-unzip) | ✓ | | | | | | ✓ | ✓ | xnor _rd_, _rs1_, _rs2_ | [Exclusive NOR](#insns-xnor) | ✓ | ✓ | | | | ✓ | ✓ | zext.h _rd_, _rs_ | [Zero-extend halfword](#insns-zext%5Fh) | ✓ | | | | | ✓ | zip _rd_, _rs_ | [Bit interleave](#insns-zip) | ✓ | | | | | | RV32 | RV64 | Mnemonic | Instruction | Zba | Zbs | | ---- | ---------------------------- | ----------------------------------------------------------- | ---------------------------------------------- | --- | --- | | ✓ | add.uw _rd_, _rs1_, _rs2_ | [Add unsigned word](#insns-add%5Fuw) | ✓ | | | | ✓ | ✓ | bclr _rd_, _rs1_, _rs2_ | [Single-Bit Clear (Register)](#insns-bclr) | ✓ | | | ✓ | ✓ | bclri _rd_, _rs1_, _imm_ | [Single-Bit Clear (Immediate)](#insns-bclri) | ✓ | | | ✓ | ✓ | bext _rd_, _rs1_, _rs2_ | [Single-Bit Extract (Register)](#insns-bext) | ✓ | | | ✓ | ✓ | bexti _rd_, _rs1_, _imm_ | [Single-Bit Extract (Immediate)](#insns-bexti) | ✓ | | | ✓ | ✓ | binv _rd_, _rs1_, _rs2_ | [Single-Bit Invert (Register)](#insns-binv) | ✓ | | | ✓ | ✓ | binvi _rd_, _rs1_, _imm_ | [Single-Bit Invert (Immediate)](#insns-binvi) | ✓ | | | ✓ | ✓ | bset _rd_, _rs1_, _rs2_ | [Single-Bit Set (Register)](#insns-bset) | ✓ | | | ✓ | ✓ | bseti _rd_, _rs1_, _imm_ | [Single-Bit Set (Immediate)](#insns-bseti) | ✓ | | | ✓ | ✓ | sh1add _rd_, _rs1_, _rs2_ | [Shift left by 1 and add](#insns-sh1add) | ✓ | | | ✓ | sh1add.uw _rd_, _rs1_, _rs2_ | [Shift unsigned word left by 1 and add](#insns-sh1add%5Fuw) | ✓ | | | | ✓ | ✓ | sh2add _rd_, _rs1_, _rs2_ | [Shift left by 2 and add](#insns-sh2add) | ✓ | | | ✓ | sh2add.uw _rd_, _rs1_, _rs2_ | [Shift unsigned word left by 2 and add](#insns-sh2add%5Fuw) | ✓ | | | | ✓ | ✓ | sh3add _rd_, _rs1_, _rs2_ | [Shift left by 3 and add](#insns-sh3add) | ✓ | | | ✓ | sh3add.uw _rd_, _rs1_, _rs2_ | [Shift unsigned word left by 3 and add](#insns-sh3add%5Fuw) | ✓ | | | | ✓ | slli.uw _rd_, _rs1_, _imm_ | [Shift-left unsigned word (Immediate)](#insns-slli%5Fuw) | ✓ | | | ### [](#30-1-1-b-extension-for-bit-manipulation-version-1-0-0)30.1.1\. "B" Extension for Bit Manipulation, Version 1.0.0 The B standard extension comprises instructions provided by the Zba, Zbb, and Zbs extensions. ### [](#zba)30.1.2\. Zba: Extension for Address generation, Version 1.0.0 The Zba instructions can be used to accelerate the generation of addresses that index into arrays of basic types (halfword, word, doubleword) using both unsigned word-sized and XLEN-sized indices: a shifted index is added to a base address. The shift and add instructions do a left shift of 1, 2, or 3 because these are commonly found in real-world code and because they can be implemented with a minimal amount of additional hardware beyond that of the simple adder. This avoids lengthening the critical path in implementations. While the shift and add instructions are limited to a maximum left shift of 3, the slli instruction (from the base ISA) can be used to perform similar shifts for indexing into arrays of wider elements. The slli.uw — added in this extension — can be used when the index is to be interpreted as an unsigned word. The following instructions comprise the Zba extension: | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---------------------------- | ----------------------------------------------------------- | ---------------------------------------- | | ✓ | add.uw _rd_, _rs1_, _rs2_ | [Add unsigned word](#insns-add%5Fuw) | | | ✓ | ✓ | sh1add _rd_, _rs1_, _rs2_ | [Shift left by 1 and add](#insns-sh1add) | | ✓ | sh1add.uw _rd_, _rs1_, _rs2_ | [Shift unsigned word left by 1 and add](#insns-sh1add%5Fuw) | | | ✓ | ✓ | sh2add _rd_, _rs1_, _rs2_ | [Shift left by 2 and add](#insns-sh2add) | | ✓ | sh2add.uw _rd_, _rs1_, _rs2_ | [Shift unsigned word left by 2 and add](#insns-sh2add%5Fuw) | | | ✓ | ✓ | sh3add _rd_, _rs1_, _rs2_ | [Shift left by 3 and add](#insns-sh3add) | | ✓ | sh3add.uw _rd_, _rs1_, _rs2_ | [Shift unsigned word left by 3 and add](#insns-sh3add%5Fuw) | | | ✓ | slli.uw _rd_, _rs1_, _imm_ | [Shift-left unsigned word (Immediate)](#insns-slli%5Fuw) | | ### [](#zbb)30.1.3\. Zbb: Extension for Basic bit-manipulation, Version 1.0.0 #### [](#30-1-3-1-logical-with-negate)30.1.3.1\. Logical with negate | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | ----------------------- | ---------------------------------------- | | ✓ | ✓ | andn _rd_, _rs1_, _rs2_ | [AND with inverted operand](#insns-andn) | | ✓ | ✓ | orn _rd_, _rs1_, _rs2_ | [OR with inverted operand](#insns-orn) | | ✓ | ✓ | xnor _rd_, _rs1_, _rs2_ | [Exclusive NOR](#insns-xnor) | | | Implementation Hint The Logical with Negate instructions can be implemented by inverting the _rs2_ inputs to the base-required AND, OR, and XOR logic instructions. In some implementations, the inverter on rs2 used for subtraction can be reused for this purpose. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#30-1-3-2-count-leadingtrailing-zero-bits)30.1.3.2\. Count leading/trailing zero bits | RV32 | RV64 | Mnemonic | Instruction | | ---- | --------------- | ----------------------------------------------- | -------------------------------------- | | ✓ | ✓ | clz _rd_, _rs_ | [Count leading zero bits](#insns-clz) | | ✓ | clzw _rd_, _rs_ | [Count leading zero bits in word](#insns-clzw) | | | ✓ | ✓ | ctz _rd_, _rs_ | [Count trailing zero bits](#insns-ctz) | | ✓ | ctzw _rd_, _rs_ | [Count trailing zero bits in word](#insns-ctzw) | | #### [](#30-1-3-3-count-population)30.1.3.3\. Count population These instructions count the number of set bits (1-bits). This is also commonly referred to as population count. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---------------- | -------------------------------------- | ----------------------------- | | ✓ | ✓ | cpop _rd_, _rs_ | [Count set bits](#insns-cpop) | | ✓ | cpopw _rd_, _rs_ | [Count set bits in word](#insns-cpopw) | | #### [](#30-1-3-4-integer-minimummaximum)30.1.3.4\. Integer minimum/maximum The integer minimum/maximum instructions are arithmetic R-type instructions that return the smaller/larger of two operands. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | ----------------------- | ------------------------------- | | ✓ | ✓ | max _rd_, _rs1_, _rs2_ | [Maximum](#insns-max) | | ✓ | ✓ | maxu _rd_, _rs1_, _rs2_ | [Unsigned maximum](#insns-maxu) | | ✓ | ✓ | min _rd_, _rs1_, _rs2_ | [Minimum](#insns-min) | | ✓ | ✓ | minu _rd_, _rs1_, _rs2_ | [Unsigned minimum](#insns-minu) | #### [](#30-1-3-5-sign-extension-and-zero-extension)30.1.3.5\. Sign extension and zero extension These instructions perform the sign extension or zero extension of the least-significant 8 bits or 16 bits of the source register. These instructions replace the generalized idioms `slli rd,rs,(XLEN-) + srai` (for sign extension of 8-bit and 16-bit quantities) and `slli + srli` (for zero extension of 16-bit quantities). | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | ----------------- | --------------------------------------- | | ✓ | ✓ | sext.b _rd_, _rs_ | [Sign-extend byte](#insns-sext%5Fb) | | ✓ | ✓ | sext.h _rd_, _rs_ | [Sign-extend halfword](#insns-sext%5Fh) | | ✓ | ✓ | zext.h _rd_, _rs_ | [Zero-extend halfword](#insns-zext%5Fh) | #### [](#30-1-3-6-bitwise-rotation)30.1.3.6\. Bitwise rotation Bitwise rotation instructions are similar to the shift-logical operations from the base spec. However, where the shift-logical instructions shift in zeros, the rotate instructions shift in the bits that were shifted out of the other side of the value.Such operations are also referred to as ‘circular shifts’. | RV32 | RV64 | Mnemonic | Instruction | | ---- | -------------------------- | --------------------------------------------- | --------------------------------------- | | ✓ | ✓ | rol _rd_, _rs1_, _rs2_ | [Rotate left (Register)](#insns-rol) | | ✓ | rolw _rd_, _rs1_, _rs2_ | [Rotate Left Word (Register)](#insns-rolw) | | | ✓ | ✓ | ror _rd_, _rs1_, _rs2_ | [Rotate right (Register)](#insns-ror) | | ✓ | ✓ | rori _rd_, _rs1_, _shamt_ | [Rotate right (Immediate)](#insns-rori) | | ✓ | roriw _rd_, _rs1_, _shamt_ | [Rotate right Word (Immediate)](#insns-roriw) | | | ✓ | rorw _rd_, _rs1_, _rs2_ | [Rotate right Word (Register)](#insns-rorw) | | | | Architecture Explanation The rotate instructions were included to replace a common four-instruction sequence to achieve the same effect (neg; sll/srl; srl/sll; or) | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#30-1-3-7-or-combine)30.1.3.7\. OR Combine **orc.b** sets the bits of each byte in the result _rd_ to all zeros if no bit within the respective byte of _rs_ is set, or to all ones if any bit within the respective byte of _rs_ is set. One use-case is string-processing functions, such as **strlen** and **strcpy**, which can use **orc.b** to test for the terminating zero byte by counting the set bits in leading non-zero bytes in a word. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | ---------------- | -------------------------------------------------- | | ✓ | ✓ | orc.b _rd_, _rs_ | [Bitwise OR-Combine, byte granule](#insns-orc%5Fb) | #### [](#30-1-3-8-byte-reverse)30.1.3.8\. Byte-reverse **rev8** reverses the byte-ordering of _rs_. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | --------------- | ------------------------------------ | | ✓ | ✓ | rev8 _rd_, _rs_ | [Byte-reverse register](#insns-rev8) | ### [](#zbc)30.1.4\. Zbc: Extension for Carry-less multiplication, Version 1.0.0 Carry-less multiplication is the multiplication in the polynomial ring over GF(2). **clmul** produces the lower half of the carry-less product and **clmulh** produces the upper half of the 2×XLEN carry-less product. **clmulr** produces bits 2×XLEN−2:XLEN-1 of the 2×XLEN carry-less product. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | ------------------------- | ------------------------------------------------ | | ✓ | ✓ | clmul _rd_, _rs1_, _rs2_ | [Carry-less multiply (low-part)](#insns-clmul) | | ✓ | ✓ | clmulh _rd_, _rs1_, _rs2_ | [Carry-less multiply (high-part)](#insns-clmulh) | | ✓ | ✓ | clmulr _rd_, _rs1_, _rs2_ | [Carry-less multiply (reversed)](#insns-clmulr) | ### [](#zbs)30.1.5\. Zbs: Extension for Single-bit instructions, Version 1.0.0 The single-bit instructions provide a mechanism to set, clear, invert, or extract a single bit in a register. The bit is specified by its index. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | ------------------------ | ---------------------------------------------- | | ✓ | ✓ | bclr _rd_, _rs1_, _rs2_ | [Single-Bit Clear (Register)](#insns-bclr) | | ✓ | ✓ | bclri _rd_, _rs1_, _imm_ | [Single-Bit Clear (Immediate)](#insns-bclri) | | ✓ | ✓ | bext _rd_, _rs1_, _rs2_ | [Single-Bit Extract (Register)](#insns-bext) | | ✓ | ✓ | bexti _rd_, _rs1_, _imm_ | [Single-Bit Extract (Immediate)](#insns-bexti) | | ✓ | ✓ | binv _rd_, _rs1_, _rs2_ | [Single-Bit Invert (Register)](#insns-binv) | | ✓ | ✓ | binvi _rd_, _rs1_, _imm_ | [Single-Bit Invert (Immediate)](#insns-binvi) | | ✓ | ✓ | bset _rd_, _rs1_, _rs2_ | [Single-Bit Set (Register)](#insns-bset) | | ✓ | ✓ | bseti _rd_, _rs1_, _imm_ | [Single-Bit Set (Immediate)](#insns-bseti) | ### [](#zbkb)30.1.6\. Zbkb: Extension for Bit-manipulation for Cryptography, Version 1.0.0 This extension contains instructions essential for implementing common operations in cryptographic workloads. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ----- | ---------------------------------------------------- | ------------------------------------------- | | ✓ | ✓ | rol | [Rotate left (Register)](#insns-rol) | | ✓ | rolw | [Rotate Left Word (Register)](#insns-rolw) | | | ✓ | ✓ | ror | [Rotate right (Register)](#insns-ror) | | ✓ | ✓ | rori | [Rotate right (Immediate)](#insns-rori) | | ✓ | roriw | [Rotate right Word (Immediate)](#insns-roriw) | | | ✓ | rorw | [Rotate right Word (Register)](#insns-rorw) | | | ✓ | ✓ | andn | [AND with inverted operand](#insns-andn) | | ✓ | ✓ | orn | [OR with inverted operand](#insns-orn) | | ✓ | ✓ | xnor | [Exclusive NOR](#insns-xnor) | | ✓ | ✓ | pack | [Pack low halves of registers](#insns-pack) | | ✓ | ✓ | packh | [Pack low bytes of registers](#insns-packh) | | ✓ | packw | [Pack low 16-bits of registers (RV64)](#insns-packw) | | | ✓ | ✓ | brev8 | [Reverse bits in bytes](#insns-brev8) | | ✓ | ✓ | rev8 | [Byte-reverse register](#insns-rev8) | | ✓ | zip | [Bit interleave](#insns-zip) | | | ✓ | unzip | [Bit deinterleave](#insns-unzip) | | ### [](#zbkc)30.1.7\. Zbkc: Extension for Carry-less multiplication for Cryptography, Version 1.0.0 Carry-less multiplication is the multiplication in the polynomial ring over GF(2). This is a critical operation in some cryptographic workloads, particularly the AES-GCM authenticated encryption scheme. This extension provides only the instructions needed to efficiently implement the GHASH operation, which is part of this workload. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | ------------------------- | ------------------------------------------------ | | ✓ | ✓ | clmul _rd_, _rs1_, _rs2_ | [Carry-less multiply (low-part)](#insns-clmul) | | ✓ | ✓ | clmulh _rd_, _rs1_, _rs2_ | [Carry-less multiply (high-part)](#insns-clmulh) | ### [](#zbkx)30.1.8\. Zbkx: Extension for Crossbar permutations, Version 1.0.0 These instructions implement a "lookup table" for 4 and 8 bit elements inside the general purpose registers. _rs1_ is used as a vector of N-bit words, and _rs2_ as a vector of N-bit indices into _rs1_.Elements in _rs1_ are replaced by the indexed element in _rs2_, or zero if the index into _rs2_ is out of bounds. These instructions are useful for expressing N-bit to N-bit boolean operations, and implementing cryptographic code with secret dependent memory accesses (particularly SBoxes) such that the execution latency does not depend on the (secret) data being operated on. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | ------------------------- | ----------------------------------------------- | | ✓ | ✓ | xperm4 _rd_, _rs1_, _rs2_ | [Crossbar permutation (nibbles)](#insns-xperm4) | | ✓ | ✓ | xperm8 _rd_, _rs1_, _rs2_ | [Crossbar permutation (bytes)](#insns-xperm8) | ### [](#insns-b)30.1.9\. Instructions (in alphabetical order) The semantics of each instruction is expressed in a SAIL-like syntax. #### [](#insns-add%5Fuw)30.1.9.1\. add.uw Synopsis Add unsigned word Mnemonic add.uw _rd_, _rs1_, _rs2_ Pseudoinstructions zext.w _rd_, _rs1_ → add.uw _rd_, _rs1_, zero Encoding ![svg](_images/svg-d287b61d49587b4e48a65b5af64ced6314f31a10.svg) Description This instruction performs an XLEN-wide addition between _rs2_ and the zero-extended least-significant word of _rs1_. Operation ```sail let base = X(rs2); let index = EXTZ(X(rs1)[31..0]); X(rd) = base + index; ``` Included in | Extension | Minimum version | Lifecycle state | | --------------------------------------------- | --------------- | --------------- | | Zba ([Address generation instructions](#zba)) | 0.93 | Ratified | #### [](#insns-andn)30.1.9.2\. andn Synopsis AND with inverted operand Mnemonic andn _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-5037e8d0227804a70953db7915f2b4e3c251a9cc.svg) Description This instruction performs the bitwise logical AND operation between _rs1_ and the bitwise inversion of _rs2_. Operation ```sail X(rd) = X(rs1) & ~X(rs2); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | #### [](#insns-bclr)30.1.9.3\. bclr Synopsis Single-Bit Clear (Register) Mnemonic bclr _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-5197c3695ea7f562541ba30a2c5d87f312caac90.svg) Description This instruction returns _rs1_ with a single bit cleared at the index specified in _rs2_. The index is read from the lower log2(XLEN) bits of _rs2_. Operation ```sail let index = X(rs2) & (XLEN - 1); X(rd) = X(rs1) & ~(1 << index) ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------- | --------------- | --------------- | | Zbs ([Single-bit instructions](#zbs)) | v1.0 | Ratified | #### [](#insns-bclri)30.1.9.4\. bclri Synopsis Single-Bit Clear (Immediate) Mnemonic bclri _rd_, _rs1_, _shamt_ Encoding (RV32) ![svg](_images/svg-02e365d64002b0b09a8e4e574f7f4867993515a9.svg) Encoding (RV64) ![svg](_images/svg-60abd493e6ecf0fbdba20871f941562a0ef1cd82.svg) Description This instruction returns _rs1_ with a single bit cleared at the index specified in _shamt_. The index is read from the lower log2(XLEN) bits of _shamt_.For RV32, the encodings corresponding to shamt\[5\]=1 are reserved. Operation ```sail let index = shamt & (XLEN - 1); X(rd) = X(rs1) & ~(1 << index) ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------- | --------------- | --------------- | | Zbs ([Single-bit instructions](#zbs)) | v1.0 | Ratified | #### [](#insns-bext)30.1.9.5\. bext Synopsis Single-Bit Extract (Register) Mnemonic bext _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-eac455dfc9194ec1ab1ac700352b0a3380b6e1e5.svg) Description This instruction returns a single bit extracted from _rs1_ at the index specified in _rs2_. The index is read from the lower log2(XLEN) bits of _rs2_. Operation ```sail let index = X(rs2) & (XLEN - 1); X(rd) = (X(rs1) >> index) & 1; ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------- | --------------- | --------------- | | Zbs ([Single-bit instructions](#zbs)) | v1.0 | Ratified | #### [](#insns-bexti)30.1.9.6\. bexti Synopsis Single-Bit Extract (Immediate) Mnemonic bexti _rd_, _rs1_, _shamt_ Encoding (RV32) ![svg](_images/svg-aae56a19005d8ff5b81c52a0de91cb7fe8a54f16.svg) Encoding (RV64) ![svg](_images/svg-ff66988e340a0a8c2c3ffffd34bca0f276608680.svg) Description This instruction returns a single bit extracted from _rs1_ at the index specified in _shamt_. The index is read from the lower log2(XLEN) bits of _shamt_.For RV32, the encodings corresponding to shamt\[5\]=1 are reserved. Operation ```sail let index = shamt & (XLEN - 1); X(rd) = (X(rs1) >> index) & 1; ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------- | --------------- | --------------- | | Zbs ([Single-bit instructions](#zbs)) | v1.0 | Ratified | #### [](#insns-binv)30.1.9.7\. binv Synopsis Single-Bit Invert (Register) Mnemonic binv _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-4ee6a1489b3bffc306382d215843de5b07320baa.svg) Description This instruction returns _rs1_ with a single bit inverted at the index specified in _rs2_. The index is read from the lower log2(XLEN) bits of _rs2_. Operation ```sail let index = X(rs2) & (XLEN - 1); X(rd) = X(rs1) ^ (1 << index) ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------- | --------------- | --------------- | | Zbs ([Single-bit instructions](#zbs)) | v1.0 | Ratified | #### [](#insns-binvi)30.1.9.8\. binvi Synopsis Single-Bit Invert (Immediate) Mnemonic binvi _rd_, _rs1_, _shamt_ Encoding (RV32) ![svg](_images/svg-10d898da03c5334ffe8e910ac5b665e4f2dea43d.svg) Encoding (RV64) ![svg](_images/svg-4fdb41129dada3e8072fa7005628b28ec8ef38b3.svg) Description This instruction returns _rs1_ with a single bit inverted at the index specified in _shamt_. The index is read from the lower log2(XLEN) bits of _shamt_.For RV32, the encodings corresponding to shamt\[5\]=1 are reserved. Operation ```sail let index = shamt & (XLEN - 1); X(rd) = X(rs1) ^ (1 << index) ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------- | --------------- | --------------- | | Zbs ([Single-bit instructions](#zbs)) | v1.0 | Ratified | #### [](#insns-bset)30.1.9.9\. bset Synopsis Single-Bit Set (Register) Mnemonic bset _rd_, _rs1_,_rs2_ Encoding ![svg](_images/svg-7a799a829a0a205064eba0d27f32b3090b832c2d.svg) Description This instruction returns _rs1_ with a single bit set at the index specified in _rs2_. The index is read from the lower log2(XLEN) bits of _rs2_. Operation ```sail let index = X(rs2) & (XLEN - 1); X(rd) = X(rs1) | (1 << index) ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------- | --------------- | --------------- | | Zbs ([Single-bit instructions](#zbs)) | v1.0 | Ratified | #### [](#insns-bseti)30.1.9.10\. bseti Synopsis Single-Bit Set (Immediate) Mnemonic bseti _rd_, _rs1_,_shamt_ Encoding (RV32) ![svg](_images/svg-a3d3088a9efd3af4ca4a17cfa5f6ddd012677760.svg) Encoding (RV64) ![svg](_images/svg-bec3d17dc1959490b585b07d69a47629566c308e.svg) Description This instruction returns _rs1_ with a single bit set at the index specified in _shamt_. The index is read from the lower log2(XLEN) bits of _shamt_.For RV32, the encodings corresponding to shamt\[5\]=1 are reserved. Operation ```sail let index = shamt & (XLEN - 1); X(rd) = X(rs1) | (1 << index) ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------- | --------------- | --------------- | | Zbs ([Single-bit instructions](#zbs)) | v1.0 | Ratified | #### [](#insns-clmul)30.1.9.11\. clmul Synopsis Carry-less multiply (low-part) Mnemonic clmul _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-3ace2bb1d5afe94f7c9321cfdf1a432733accd80.svg) Description clmul produces the lower half of the 2·XLEN carry-less product. Operation ```sail let rs1_val = X(rs1); let rs2_val = X(rs2); let output : xlenbits = 0; foreach (i from 0 to (xlen - 1) by 1) { output = if ((rs2_val >> i) & 1) then output ^ (rs1_val << i); else output; } X[rd] = output ``` Included in | Extension | Minimum version | Lifecycle state | | ---------------------------------------------------------- | --------------- | --------------- | | Zbc ([Carry-less multiplication](#zbc)) | v1.0 | Ratified | | Zbkc ([Carry-less multiplication for Cryptography](#zbkc)) | v1.0 | Ratified | #### [](#insns-clmulh)30.1.9.12\. clmulh Synopsis Carry-less multiply (high-part) Mnemonic clmulh _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-6107d0ddbb95c0228642e006d50cb857f98e58dc.svg) Description clmulh produces the upper half of the 2·XLEN carry-less product. Operation ```sail let rs1_val = X(rs1); let rs2_val = X(rs2); let output : xlenbits = 0; foreach (i from 1 to xlen by 1) { output = if ((rs2_val >> i) & 1) then output ^ (rs1_val >> (xlen - i)); else output; } X[rd] = output ``` Included in | Extension | Minimum version | Lifecycle state | | ---------------------------------------------------------- | --------------- | --------------- | | Zbc ([Carry-less multiplication](#zbc)) | v1.0 | Ratified | | Zbkc ([Carry-less multiplication for Cryptography](#zbkc)) | v1.0 | Ratified | #### [](#insns-clmulr)30.1.9.13\. clmulr Synopsis Carry-less multiply (reversed) Mnemonic clmulr _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-144588cb61ea0fafddd576505d53579b34200c32.svg) Description **clmulr** produces bits 2·XLEN−2:XLEN-1 of the 2·XLEN carry-less product. Operation ```sail let rs1_val = X(rs1); let rs2_val = X(rs2); let output : xlenbits = 0; foreach (i from 0 to (xlen - 1) by 1) { output = if ((rs2_val >> i) & 1) then output ^ (rs1_val >> (xlen - i - 1)); else output; } X[rd] = output ``` | | Note The **clmulr** instruction is used to accelerate CRC calculations. The **r** in the instruction’s mnemonic stands for _reversed_, as the instruction is equivalent to bit-reversing the inputs, performing a **clmul**, then bit-reversing the output. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Included in | Extension | Minimum version | Lifecycle state | | --------------------------------------- | --------------- | --------------- | | Zbc ([Carry-less multiplication](#zbc)) | v1.0 | Ratified | #### [](#insns-clz)30.1.9.14\. clz Synopsis Count leading zero bits Mnemonic clz _rd_, _rs_ Encoding ![svg](_images/svg-74030ddfa6239a384ea769fb82f1bf4f091436d5.svg) Description This instruction counts the number of 0’s before the first 1, starting at the most-significant bit (i.e., XLEN-1) and progressing to bit 0\. Accordingly, if the input is 0, the output is XLEN, and if the most-significant bit of the input is a 1, the output is 0. Operation ```sail val HighestSetBit : forall ('N : Int), 'N >= 0. bits('N) -> int function HighestSetBit x = { foreach (i from (xlen - 1) to 0 by 1 in dec) if [x[i]] == 0b1 then return(i) else (); return -1; } let rs = X(rs); X[rd] = (xlen - 1) - HighestSetBit(rs); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | #### [](#insns-clzw)30.1.9.15\. clzw Synopsis Count leading zero bits in word Mnemonic clzw _rd_, _rs_ Encoding ![svg](_images/svg-0aea0581bd17e3cf1f173380258cfab612bf791d.svg) Description This instruction counts the number of 0’s before the first 1 starting at bit 31 and progressing to bit 0\. Accordingly, if the least-significant word is 0, the output is 32, and if the most-significant bit of the word (i.e., bit 31) is a 1, the output is 0. Operation ```sail val HighestSetBit32 : forall ('N : Int), 'N >= 0. bits('N) -> int function HighestSetBit32 x = { foreach (i from 31 to 0 by 1 in dec) if [x[i]] == 0b1 then return(i) else (); return -1; } let rs = X(rs); X[rd] = 31 - HighestSetBit(rs); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | #### [](#insns-cpop)30.1.9.16\. cpop Synopsis Count set bits Mnemonic cpop _rd_, _rs_ Encoding ![svg](_images/svg-cb428f8069dcf4274666071bc4b3b0e248ca26b1.svg) Description This instructions counts the number of 1’s (i.e., set bits) in the source register. Operation ```sail let bitcount = 0; let rs = X(rs); foreach (i from 0 to (xlen - 1) in inc) if rs[i] == 0b1 then bitcount = bitcount + 1 else (); X[rd] = bitcount ``` | | Software Hint This operation is known as population count, popcount, sideways sum, bit summation, or Hamming weight. The GCC builtin function \_\_builtin\_popcount (unsigned int x) is implemented by cpop on RV32 and by **cpopw** on RV64\. The GCC builtin function \_\_builtin\_popcountl (unsigned long x) for LP64 is implemented by **cpop** on RV64. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | #### [](#insns-cpopw)30.1.9.17\. cpopw Synopsis Count set bits in word Mnemonic cpopw _rd_, _rs_ Encoding ![svg](_images/svg-4c58e162882caadbe615b8e18eed769e33bdf658.svg) Description This instructions counts the number of 1’s (i.e., set bits) in the least-significant word of the source register. Operation ```sail let bitcount = 0; let val = X(rs); foreach (i from 0 to 31 in inc) if val[i] == 0b1 then bitcount = bitcount + 1 else (); X[rd] = bitcount ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | #### [](#insns-ctz)30.1.9.18\. ctz Synopsis Count trailing zeros Mnemonic ctz _rd_, _rs_ Encoding ![svg](_images/svg-793bdc506dbf49900ea89fdcac36f0204640a2eb.svg) Description This instruction counts the number of 0’s before the first 1, starting at the least-significant bit (i.e., 0) and progressing to the most-significant bit (i.e., XLEN-1). Accordingly, if the input is 0, the output is XLEN, and if the least-significant bit of the input is a 1, the output is 0. Operation ```sail val LowestSetBit : forall ('N : Int), 'N >= 0. bits('N) -> int function LowestSetBit x = { foreach (i from 0 to (xlen - 1) by 1 in dec) if [x[i]] == 0b1 then return(i) else (); return xlen; } let rs = X(rs); X[rd] = LowestSetBit(rs); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | #### [](#insns-ctzw)30.1.9.19\. ctzw Synopsis Count trailing zero bits in word Mnemonic ctzw _rd_, _rs_ Encoding ![svg](_images/svg-524a4a04bbad27bc4eb2fb1e5c9513fa66012e9d.svg) Description This instruction counts the number of 0’s before the first 1, starting at the least-significant bit (i.e., 0) and progressing to the most-significant bit of the least-significant word (i.e., 31). Accordingly, if the least-significant word is 0, the output is 32, and if the least-significant bit of the input is a 1, the output is 0. Operation ```sail val LowestSetBit32 : forall ('N : Int), 'N >= 0. bits('N) -> int function LowestSetBit32 x = { foreach (i from 0 to 31 by 1 in dec) if [x[i]] == 0b1 then return(i) else (); return 32; } let rs = X(rs); X[rd] = LowestSetBit32(rs); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | #### [](#insns-max)30.1.9.20\. max Synopsis Maximum Mnemonic max _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-f1d739999eabb70b93d031b92db63ec5d8b15539.svg) Description This instruction returns the larger of two signed integers. Operation ```sail let rs1_val = X(rs1); let rs2_val = X(rs2); let result = if rs1_val <_s rs2_val then rs2_val else rs1_val; X(rd) = result; ``` | | Software Hint Calculating the absolute value of a signed integer can be performed using the following sequence: **neg rD,rS** followed by **max rD,rS,rD**. When using this common sequence, it is suggested that they are scheduled with no intervening instructions so that implementations that are so optimized can fuse them together. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | #### [](#insns-maxu)30.1.9.21\. maxu Synopsis Unsigned maximum Mnemonic maxu _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-28dee91738ed95baa119cf39a15c943c321f50f0.svg) Description This instruction returns the larger of two unsigned integers. Operation ```sail let rs1_val = X(rs1); let rs2_val = X(rs2); let result = if rs1_val <_u rs2_val then rs2_val else rs1_val; X(rd) = result; ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | #### [](#insns-min)30.1.9.22\. min Synopsis Minimum Mnemonic min _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-41d337c1029a029a12826c7ef84acd8aad3ed2d0.svg) Description This instruction returns the smaller of two signed integers. Operation ```sail let rs1_val = X(rs1); let rs2_val = X(rs2); let result = if rs1_val <_s rs2_val then rs1_val else rs2_val; X(rd) = result; ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | #### [](#insns-minu)30.1.9.23\. minu Synopsis Unsigned minimum Mnemonic minu _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-6b03b98209913a0a84285f4eea463be01afba503.svg) Description This instruction returns the smaller of two unsigned integers. Operation ```sail let rs1_val = X(rs1); let rs2_val = X(rs2); let result = if rs1_val <_u rs2_val then rs1_val else rs2_val; X(rd) = result; ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | #### [](#insns-orc%5Fb)30.1.9.24\. orc.b Synopsis Bitwise OR-Combine, byte granule Mnemonic orc.b _rd_, _rs_ Encoding ![svg](_images/svg-8c20e009ae27789fbd745fb19b8c3a1861c24b26.svg) Description Combines the bits within each byte using bitwise logical OR. This sets the bits of each byte in the result _rd_ to all zeros if no bit within the respective byte of _rs_ is set, or to all ones if any bit within the respective byte of _rs_ is set. Operation ```sail let input = X(rs); let output : xlenbits = 0; foreach (i from 0 to (xlen - 8) by 8) { output[(i + 7)..i] = if input[(i + 7)..i] == 0 then 0b00000000 else 0b11111111; } X[rd] = output; ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | #### [](#insns-orn)30.1.9.25\. orn Synopsis OR with inverted operand Mnemonic orn _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-cd2eb5051fd881d76177c58af39944f79a35e27a.svg) Description This instruction performs the bitwise logical OR operation between _rs1_ and the bitwise inversion of _rs2_. Operation ```sail X(rd) = X(rs1) | ~X(rs2); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | #### [](#insns-pack)30.1.9.26\. pack Synopsis Pack the low halves of _rs1_ and _rs2_ into _rd_. Mnemonic pack _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-3b154237dc0d0a29f3ee4a786ba979924f77f0a0.svg) Description The pack instruction packs the XLEN/2-bit lower halves of _rs1_ and _rs2_ into_rd_, with _rs1_ in the lower half and _rs2_ in the upper half. Operation ```sail let lo_half : bits(xlen/2) = X(rs1)[xlen/2-1..0]; let hi_half : bits(xlen/2) = X(rs2)[xlen/2-1..0]; X(rd) = EXTZ(hi_half @ lo_half); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | | | For RV32, the pack instruction with _rs2_\=x0 is the zext.hinstruction. Hence, for RV32, any extension that contains the pack instruction also contains the zext.h instruction (but not necessarily the c.zext.hinstruction, which is only guaranteed to exist if both the Zcb and Zbb extensions are implemented). | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#insns-packh)30.1.9.27\. packh Synopsis Pack the low bytes of _rs1_ and _rs2_ into _rd_. Mnemonic packh _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-de742e02dd64298507cef284a266492885bc7eb3.svg) Description The packh instruction packs the least-significant bytes of_rs1_ and _rs2_ into the 16 least-significant bits of _rd_, zero extending the rest of _rd_. Operation ```sail let lo_half : bits(8) = X(rs1)[7..0]; let hi_half : bits(8) = X(rs2)[7..0]; X(rd) = EXTZ(hi_half @ lo_half); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | #### [](#insns-packw)30.1.9.28\. packw Synopsis Pack the low 16-bits of _rs1_ and _rs2_ into _rd_ on RV64. Mnemonic packw _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-c62d518868337ac528c94ddef4360d218992288c.svg) Description This instruction packs the low 16 bits of_rs1_ and _rs2_ into the 32 least-significant bits of _rd_, sign extending the 32-bit result to the rest of _rd_. This instruction only exists on RV64 based systems. Operation ```sail let lo_half : bits(16) = X(rs1)[15..0]; let hi_half : bits(16) = X(rs2)[15..0]; X(rd) = EXTS(hi_half @ lo_half); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | | | For RV64, the packw instruction with _rs2_\=x0 is the zext.hinstruction. Hence, for RV64, any extension that contains the packw instruction also contains the zext.h instruction (but not necessarily the c.zext.hinstruction, which is only guaranteed to exist if both the Zcb and Zbb extensions are implemented). | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#insns-rev8)30.1.9.29\. rev8 Synopsis Byte-reverse register Mnemonic rev8 _rd_, _rs_ Encoding (RV32) ![svg](_images/svg-693d173289db88cf4b8d29c9a4c08f5ecbe7f7e2.svg) Encoding (RV64) ![svg](_images/svg-6d52ea2a08d942672651db671ee4dca698d0b929.svg) Description This instruction reverses the order of the bytes in _rs_. Operation ```sail let input = X(rs); let output : xlenbits = 0; let j = xlen - 1; foreach (i from 0 to (xlen - 8) by 8) { output[i..(i + 7)] = input[(j - 7)..j]; j = j - 8; } X[rd] = output ``` | | Note The **rev8** mnemonic corresponds to different instruction encodings in RV32 and RV64. | | ---------------------------------------------------------------------------------------------- | | | Software Hint The byte-reverse operation is only available for the full register width. To emulate word-sized and halfword-sized byte-reversal, perform a rev8 rd,rs followed by a srai rd,rd,K, where K is XLEN-32 and XLEN-16, respectively. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | v1.0 | Ratified | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | #### [](#insns-brev8)30.1.9.30\. brev8 Synopsis Reverse the bits in each byte of a source register. Mnemonic brev8 _rd_, _rs_ Encoding ![svg](_images/svg-2fb21fe817b13f00b3265dee5e0d3eb405943c35.svg) Description This instruction reverses the order of the bits in every byte of a register. Operation ```sail result : xlenbits = EXTZ(0b0); foreach (i from 0 to sizeof(xlen) by 8) { result[i+7..i] = reverse_bits_in_byte(X(rs1)[i+7..i]); }; X(rd) = result; ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | #### [](#insns-rol)30.1.9.31\. rol Synopsis Rotate Left (Register) Mnemonic rol _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-0add38b10235680259e9ea3ba4b471b26a942d2d.svg) Description This instruction performs a rotate left of _rs1_ by the amount in least-significant log2(XLEN) bits of _rs2_. Operation ```sail let shamt = if xlen == 32 then X(rs2)[4..0] else X(rs2)[5..0]; let result = (X(rs1) << shamt) | (X(rs1) >> (xlen - shamt)); X(rd) = result; ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | 0.93 | Ratified | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | #### [](#insns-rolw)30.1.9.32\. rolw Synopsis Rotate Left Word (Register) Mnemonic rolw _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-f5b062418b2d8410926587fb72cd6fc0123e75e9.svg) Description This instruction performs a rotate left on the least-significant word of _rs1_ by the amount in least-significant 5 bits of _rs2_. The resulting word value is sign-extended by copying bit 31 to all of the more-significant bits. Operation ```sail let rs1 = EXTZ(X(rs1)[31..0]) let shamt = X(rs2)[4..0]; let result = (rs1 << shamt) | (rs1 >> (32 - shamt)); X(rd) = EXTS(result[31..0]); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | 0.93 | Ratified | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | #### [](#insns-ror)30.1.9.33\. ror Synopsis Rotate Right Mnemonic ror _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-30fb690178dfc1807bfbaadb3e95c705c92708a3.svg) Description This instruction performs a rotate right of _rs1_ by the amount in least-significant log2(XLEN) bits of _rs2_. Operation ```sail let shamt = if xlen == 32 then X(rs2)[4..0] else X(rs2)[5..0]; let result = (X(rs1) >> shamt) | (X(rs1) << (xlen - shamt)); X(rd) = result; ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | 0.93 | Ratified | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | #### [](#insns-rori)30.1.9.34\. rori Synopsis Rotate Right (Immediate) Mnemonic rori _rd_, _rs1_, _shamt_ Encoding (RV32) ![svg](_images/svg-938fcfa9f20d5623c60b26ec22337b55fe5692e6.svg) Encoding (RV64) ![svg](_images/svg-56357a7bc26ad075b2aa68de3ae2db30bc1d104d.svg) Description This instruction performs a rotate right of _rs1_ by the amount in the least-significant log2(XLEN) bits of _shamt_. For RV32, the encodings corresponding to shamt\[5\]=1 are reserved. Operation ```sail let shamt = if xlen == 32 then shamt[4..0] else shamt[5..0]; let result = (X(rs1) >> shamt) | (X(rs1) << (xlen - shamt)); X(rd) = result; ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | 0.93 | Ratified | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | #### [](#insns-roriw)30.1.9.35\. roriw Synopsis Rotate Right Word by Immediate Mnemonic roriw _rd_, _rs1_, _shamt_ Encoding ![svg](_images/svg-ade4a28eb2f571fc50a9a5567dbd7991a3b05a4c.svg) Description This instruction performs a rotate right on the least-significant word of _rs1_ by the amount in the least-significant log2(XLEN) bits of_shamt_. The resulting word value is sign-extended by copying bit 31 to all of the more-significant bits. Operation ```sail let rs1_data = EXTZ(X(rs1)[31..0]; let result = (rs1_data >> shamt) | (rs1_data << (32 - shamt)); X(rd) = EXTS(result[31..0]); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | 0.93 | Ratified | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | #### [](#insns-rorw)30.1.9.36\. rorw Synopsis Rotate Right Word (Register) Mnemonic rorw _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-565fec38b77826e629c17d1e46d0048fef634910.svg) Description This instruction performs a rotate right on the least-significant word of _rs1_ by the amount in least-significant 5 bits of _rs2_. The resultant word is sign-extended by copying bit 31 to all of the more-significant bits. Operation ```sail let rs1 = EXTZ(X(rs1)[31..0]) let shamt = X(rs2)[4..0]; let result = (rs1 >> shamt) | (rs1 << (32 - shamt)); X(rd) = EXTS(result); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | 0.93 | Ratified | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | #### [](#insns-sext%5Fb)30.1.9.37\. sext.b Synopsis Sign-extend byte Mnemonic sext.b _rd_, _rs_ Encoding ![svg](_images/svg-56c0f8cb67ef8494715fe3ea758045779c1200ee.svg) Description This instruction sign-extends the least-significant byte in the source to XLEN by copying the most-significant bit in the byte (i.e., bit 7) to all of the more-significant bits. Operation ```sail X(rd) = EXTS(X(rs)[7..0]); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | 0.93 | Ratified | #### [](#insns-sext%5Fh)30.1.9.38\. sext.h Synopsis Sign-extend halfword Mnemonic sext.h _rd_, _rs_ Encoding ![svg](_images/svg-884d6d756b0420aeed820f45382fcf146f477843.svg) Description This instruction sign-extends the least-significant halfword in _rs_ to XLEN by copying the most-significant bit in the halfword (i.e., bit 15) to all of the more-significant bits. Operation ```sail X(rd) = EXTS(X(rs)[15..0]); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | 0.93 | Ratified | #### [](#insns-sh1add)30.1.9.39\. sh1add Synopsis Shift left by 1 and add Mnemonic sh1add _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-3496adff925a90289b3082d858eab871a9416314.svg) Description This instruction shifts _rs1_ to the left by 1 bit and adds it to _rs2_. Operation ```sail X(rd) = X(rs2) + (X(rs1) << 1); ``` Included in | Extension | Minimum version | Lifecycle state | | --------------------------------------------- | --------------- | --------------- | | Zba ([Address generation instructions](#zba)) | 0.93 | Ratified | #### [](#insns-sh1add%5Fuw)30.1.9.40\. sh1add.uw Synopsis Shift unsigned word left by 1 and add Mnemonic sh1add.uw _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-f9219bafb82a7a4d365baa0fad2f815136e8f694.svg) Description This instruction performs an XLEN-wide addition of two addends. The first addend is _rs2_. The second addend is the unsigned value formed by extracting the least-significant word of _rs1_ and shifting it left by 1 place. Operation ```sail let base = X(rs2); let index = EXTZ(X(rs1)[31..0]); X(rd) = base + (index << 1); ``` Included in | Extension | Minimum version | Lifecycle state | | --------------------------------------------- | --------------- | --------------- | | Zba ([Address generation instructions](#zba)) | 0.93 | Ratified | #### [](#insns-sh2add)30.1.9.41\. sh2add Synopsis Shift left by 2 and add Mnemonic sh2add _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-f1b34ee72bee682b8ca55d8014eb2202ecd616de.svg) Description This instruction shifts _rs1_ to the left by 2 places and adds it to _rs2_. Operation ```sail X(rd) = X(rs2) + (X(rs1) << 2); ``` Included in | Extension | Minimum version | Lifecycle state | | --------------------------------------------- | --------------- | --------------- | | Zba ([Address generation instructions](#zba)) | 0.93 | Ratified | #### [](#insns-sh2add%5Fuw)30.1.9.42\. sh2add.uw Synopsis Shift unsigned word left by 2 and add Mnemonic sh2add.uw _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-5b7ae400d369bca899be827cd16283e0630e9a4f.svg) Description This instruction performs an XLEN-wide addition of two addends. The first addend is _rs2_. The second addend is the unsigned value formed by extracting the least-significant word of _rs1_ and shifting it left by 2 places. Operation ```sail let base = X(rs2); let index = EXTZ(X(rs1)[31..0]); X(rd) = base + (index << 2); ``` Included in | Extension | Minimum version | Lifecycle state | | --------------------------------------------- | --------------- | --------------- | | Zba ([Address generation instructions](#zba)) | 0.93 | Ratified | #### [](#insns-sh3add)30.1.9.43\. sh3add Synopsis Shift left by 3 and add Mnemonic sh3add _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-cae18020bb7c0c1b2b48a043f53e54400ea2254e.svg) Description This instruction shifts _rs1_ to the left by 3 places and adds it to _rs2_. Operation ```sail X(rd) = X(rs2) + (X(rs1) << 3); ``` Included in | Extension | Minimum version | Lifecycle state | | --------------------------------------------- | --------------- | --------------- | | Zba ([Address generation instructions](#zba)) | 0.93 | Ratified | #### [](#insns-sh3add%5Fuw)30.1.9.44\. sh3add.uw Synopsis Shift unsigned word left by 3 and add Mnemonic sh3add.uw _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-02a948fc567e1c58dad47667f66a51358b33154f.svg) Description This instruction performs an XLEN-wide addition of two addends. The first addend is _rs2_. The second addend is the unsigned value formed by extracting the least-significant word of _rs1_ and shifting it left by 3 places. Operation ```sail let base = X(rs2); let index = EXTZ(X(rs1)[31..0]); X(rd) = base + (index << 3); ``` Included in | Extension | Minimum version | Lifecycle state | | --------------------------------------------- | --------------- | --------------- | | Zba ([Address generation instructions](#zba)) | 0.93 | Ratified | #### [](#insns-slli%5Fuw)30.1.9.45\. slli.uw Synopsis Shift-left unsigned word (Immediate) Mnemonic slli.uw _rd_, _rs1_, _shamt_ Encoding ![svg](_images/svg-b962d32860639dddda8ba396d0f05c7274aac470.svg) Description This instruction takes the least-significant word of _rs1_, zero-extends it, and shifts it left by the immediate. Operation ```sail X(rd) = (EXTZ(X(rs)[31..0]) << shamt); ``` Included in | Extension | Minimum version | Lifecycle state | | --------------------------------------------- | --------------- | --------------- | | Zba ([Address generation instructions](#zba)) | 0.93 | Ratified | | | Architecture Explanation This instruction is the same as **slli** with **zext.w** performed on _rs1_ before shifting. | | ------------------------------------------------------------------------------------------------------------------------ | #### [](#insns-unzip)30.1.9.46\. unzip Synopsis Place odd and even bits of the source register into upper and lower halves of the destination register, respectively. Mnemonic unzip _rd_, _rs_ Encoding ![svg](_images/svg-021809b893e360d56ec1ac2e1a07b74fe181b9bb.svg) Description This instruction scatters all of the odd and even bits of a source word into the high and low halves of a destination word. It is the inverse of the [zip](scalar-crypto.html#insns-zip-sc) instruction. This instruction is available only on RV32. Operation ```sail foreach (i from 0 to xlen/2-1) { X(rd)[i] = X(rs1)[2*i] X(rd)[i+xlen/2] = X(rs1)[2*i+1] } ``` | | Software Hint This instruction is useful for implementing the SHA3 cryptographic hash function on a 32-bit architecture, as it implements the bit-interleaving operation used to speed up the 64-bit rotations directly. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Included in | Extension | Minimum version | Lifecycle state | | -------------------------------------------------------- | --------------- | --------------- | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) (RV32) | v1.0 | Ratified | #### [](#insns-xnor)30.1.9.47\. xnor Synopsis Exclusive NOR Mnemonic xnor _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-69198fb0cd20a10e882917a7e04916117cbc5945.svg) Description This instruction performs the bit-wise exclusive-NOR operation on _rs1_ and _rs2_. Operation ```sail X(rd) = ~(X(rs1) ^ X(rs2)); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------------------- | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | 0.93 | Ratified | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) | v1.0 | Ratified | #### [](#insns-xperm8)30.1.9.48\. xperm8 Synopsis Byte-wise lookup of indices into a vector in registers. Mnemonic xperm8 _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-6425b4a7e06c2628490e17ec51384fe9ff6ee36b.svg) Description The xperm8 instruction operates on bytes. The _rs1_ register contains a vector of XLEN/8 8-bit elements. The _rs2_ register contains a vector of XLEN/8 8-bit indexes. The result is each element in _rs2_ replaced by the indexed element in _rs1_, or zero if the index into _rs2_ is out of bounds. Operation ```sail val xperm8_lookup : (bits(8), xlenbits) -> bits(8) function xperm8_lookup (idx, lut) = { (lut >> (idx @ 0b000))[7..0] } function clause execute ( XPERM8 (rs2,rs1,rd)) = { result : xlenbits = EXTZ(0b0); foreach(i from 0 to xlen by 8) { result[i+7..i] = xperm8_lookup(X(rs2)[i+7..i], X(rs1)); }; X(rd) = result; RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------- | --------------- | --------------- | | Zbkx ([Crossbar permutations](#zbkx)) | v1.0 | Ratified | #### [](#insns-xperm4)30.1.9.49\. xperm4 Synopsis Nibble-wise lookup of indices into a vector. Mnemonic xperm4 _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-e633638e4cd9d0c4bfb0a8b0eb4925c256cf5c8f.svg) Description The xperm4 instruction operates on nibbles. The _rs1_ register contains a vector of XLEN/4 4-bit elements. The _rs2_ register contains a vector of XLEN/4 4-bit indexes. The result is each element in _rs2_ replaced by the indexed element in _rs1_, or zero if the index into _rs2_ is out of bounds. Operation ```sail val xperm4_lookup : (bits(4), xlenbits) -> bits(4) function xperm4_lookup (idx, lut) = { (lut >> (idx @ 0b00))[3..0] } function clause execute ( XPERM4 (rs2,rs1,rd)) = { result : xlenbits = EXTZ(0b0); foreach(i from 0 to xlen by 4) { result[i+3..i] = xperm4_lookup(X(rs2)[i+3..i], X(rs1)); }; X(rd) = result; RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------- | --------------- | --------------- | | Zbkx ([Crossbar permutations](#zbkx)) | v1.0 | Ratified | #### [](#insns-zext%5Fh)30.1.9.50\. zext.h Synopsis Zero-extend halfword Mnemonic zext.h _rd_, _rs_ Encoding (RV32) ![svg](_images/svg-e5193b2f300435576a16ceede38c54a2c68bb137.svg) Encoding (RV64) ![svg](_images/svg-a131c95e966566e7eb5c42ee69de1512cbeb3255.svg) Description This instruction zero-extends the least-significant halfword of the source to XLEN by inserting 0’s into all of the bits more significant than 15. Operation ```sail X(rd) = EXTZ(X(rs)[15..0]); ``` | | Note The **zext.h** mnemonic corresponds to different instruction encodings in RV32 and RV64. | | ------------------------------------------------------------------------------------------------ | Included in | Extension | Minimum version | Lifecycle state | | ------------------------------------ | --------------- | --------------- | | Zbb ([Basic bit-manipulation](#zbb)) | 0.93 | Ratified | #### [](#insns-zip)30.1.9.51\. zip Synopsis Interleave upper and lower halves of the source register into odd and even bits of the destination register, respectively. Mnemonic zip _rd_, _rs_ Encoding ![svg](_images/svg-a9c5f3131819c935d2e569aa509a89770202a0d0.svg) Description This instruction gathers bits from the high and low halves of the source word into odd/even bit positions in the destination word. It is the inverse of the [unzip](scalar-crypto.html#insns-unzip-sc) instruction. This instruction is available only on RV32. Operation ```sail foreach (i from 0 to xlen/2-1) { X(rd)[2*i] = X(rs1)[i] X(rd)[2*i+1] = X(rs1)[i+xlen/2] } ``` | | Software Hint This instruction is useful for implementing the SHA3 cryptographic hash function on a 32-bit architecture, as it implements the bit-interleaving operation used to speed up the 64-bit rotations directly. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Included in | Extension | Minimum version | Lifecycle state | | -------------------------------------------------------- | --------------- | --------------- | | Zbkb ([Bit-manipulation for Cryptography](#zbkb)) (RV32) | v1.0 | Ratified | 25.1. "BF16" Extensions for BFloat16-precision Floating-Point, Version 1.0 ==================== ## [](#bf16)25.1\. "BF16" Extensions for BFloat16-precision Floating-Point, Version 1.0 ### [](#BF16%5Fintroduction)25.1.1\. Introduction When FP16 (officially called binary16) was first introduced by the IEEE-754 standard, it was just an interchange format. It was intended as a space/bandwidth efficient encoding that would be used to transfer information. This is in line with the Zfhmin extension. However, there were some applications (notably graphics) that found that the smaller precision and dynamic range was sufficient for their space. So, FP16 started to see some widespread adoption as an arithmetic format. This is in line with the Zfh extension. While it was not the intention of '754 to have FP16 be an arithmetic format, it is supported by the standard. Even though the '754 committee recognized that FP16 was gaining popularity, the committee decided to hold off on making it a basic format in the 2019 release. This means that a '754 compliant implementation of binary floating point, which needs to support at least one basic format, cannot support only FP16 - it needs to support at least one of binary32, binary64, and binary128. Experts working in machine learning noticed that FP16 was a much more compact way of storing operands and often provided sufficient precision for them. However, they also found that intermediate values were much better when accumulated into a higher precision. The final computations were then typically converted back into the more compact FP16 encoding. This approach has become very common in machine learning (ML) inference where the weights and activations are stored in FP16 encodings. There was the added benefit that smaller multiplication blocks could be created for the FP16’s smaller number of significant bits. At this point, widening multiply-accumulate instructions became much more common. Also, more complicated dot product instructions started to show up including those that packed two FP16 numbers in a 32-bit register, multiplied these by another pair of FP16 numbers in another register, added these two products to an FP32 accumulate value in a 3rd register and returned an FP32 result. Experts working in machine learning at Google who continued to work with FP32 values noted that the least significant 16 bits of their mantissas were not always needed for good results, even in training. They proposed a truncated version of FP32, which was the 16 most significant bits of the FP32 encoding. This format was named BFloat16 (or BF16). The B in BF16, stands for Brain since it was initially introduced by the Google Brain team. Not only did they find that the number of significant bits in BF16 tended to be sufficient for their work (despite being fewer than in FP16), but it was very easy for them to reuse their existing data; FP32 numbers could be readily rounded to BF16 with a minimal amount of work. Furthermore, the even smaller number of the BF16 significant bits enabled even smaller multiplication blocks to be built. Similar to FP16, BF16 multiply-accumulate widening and dot-product instructions started to proliferate. ### [](#BF16%5Faudience)25.1.2\. Intended Audience Floating-point arithmetic is a specialized subject, requiring people with many different backgrounds to cooperate in its correct and efficient implementation. Where possible, we have written this specification to be understandable by all, though we recognize that the motivations and references to algorithms or other specifications and standards may be unfamiliar to those who are not domain experts. This specification anticipates being read and acted on by various people with different backgrounds. We have tried to capture these backgrounds here, with a brief explanation of what we expect them to know, and how it relates to the specification. We hope this aids people’s understanding of which aspects of the specification are particularly relevant to them, and which they may (safely!) ignore or pass to a colleague. Software developers These are the people we expect to write code using the instructions in this specification. They should understand the motivations for the instructions we include, and be familiar with most of the algorithms and outside standards to which we refer. Computer architects We expect architects to have some basic floating-point background. Furthermore, we expect architects to be able to examine our instructions for implementation issues, understand how the instructions will be used in context, and advise on how they best to fit the functionality. Digital design engineers & micro-architects These are the people who will implement the specification inside a core. Floating-point expertise is assumed as not all of the corner cases are pointed out in the specification. Verification engineers Responsible for ensuring the correct implementation of the extension in hardware. These people are expected to have some floating-point expertise so that they can identify and generate the interesting corner cases --- include exceptions --- that are common in floating-point architectures and implementations. These are by no means the only people concerned with the specification, but they are the ones we considered most while writing it. ### [](#BF16%5Fformat)25.1.3\. Number Format #### [](#25-1-3-1-bf16-operand-format)25.1.3.1\. BF16 Operand Format BF16 bits ![svg](_images/svg-04e4bf4bf16547c3cc427084877aa6d2a36322a7.svg) IEEE Compliance: While BF16 (also known as BFloat16) is not an IEEE-754 _standard_ format, it is a valid floating-point format as defined by IEEE-754\. There are three parameters that specify a format: radix (b), number of digits in the significand (p), and maximum exponent (emax). For BF16 these values are: __Table 1\. BF16 parameters__ | Parameter | Value | | --------------- | ----- | | radix (b) | 2 | | significand (p) | 8 | | emax | 127 | __Table 2\. Obligatory Floating Point Format Table__ | Format | Sign Bits | Expo Bits | fraction bits | padded 0s | encoding bits | expo max/bias | expo min | | ------ | --------- | --------- | ------------- | --------- | ------------- | ------------- | -------- | | FP16 | 1 | 5 | 10 | 0 | 16 | 15 | \-14 | | BF16 | 1 | 8 | 7 | 0 | 16 | 127 | \-126 | | TF32 | 1 | 8 | 10 | 13 | 32 | 127 | \-126 | | FP32 | 1 | 8 | 23 | 0 | 32 | 127 | \-126 | | FP64 | 1 | 11 | 52 | 0 | 64 | 1023 | \-1022 | | FP128 | 1 | 15 | 112 | 0 | 128 | 16,383 | \-16,382 | #### [](#25-1-3-2-bf16-behavior)25.1.3.2\. BF16 Behavior For these BF16 extensions, instruction behavior on BF16 operands is the same as for other floating-point instructions in the RISC-V ISA. For easy reference, some of this behavior is repeated here. ##### [](#25-1-3-2-1-subnormal-numbers)25.1.3.2.1\. Subnormal Numbers: Floating-point values that are too small to be represented as normal numbers, but can still be expressed by the format’s smallest exponent value with a "0" integer bit and at least one "1" bit in the trailing fractional bits are called subnormal numbers. Basically, the idea is there is a trade off of precision to support _gradual underflow_. All of the BF16 instructions in the extensions defined in this specification (i.e., Zfbfmin, Zvfbfmin, and Zvfbfwma) fully support subnormal numbers. That is, instructions are able to accept subnormal values as inputs and they can produce subnormal results. | | Future floating-point extensions, including those that operate on BF16 values, may chose not to support subnormal numbers. The comments about supporting subnormal BF16 values are limited to those instructions defined in this specification. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#25-1-3-2-2-infinities)25.1.3.2.2\. Infinities: Infinities are used to represent values that are too large to be represented by the target format. These are usually produced as a result of overflows (depending on the rounding mode), but can also be provided as inputs. Infinities have a sign associated with them: there are positive infinities and negative infinities. Infinities are important for keeping meaningless results from being operated upon. ##### [](#25-1-3-2-3-nans)25.1.3.2.3\. NaNs NaN stands for Not a Number. There are two types of NaNs: signalling (sNaN) and quiet (qNaN). No computational instruction will ever produce an sNaN; These are only provided as input data. Operating on an sNaN will cause an invalid operation exception. Operating on a Quiet NaN usually does not cause an exception. QNaNs are provided as the result of an operation when it cannot be represented as a number or infinity. For example, performing the square root of -1 will result in a qNaN because there is no real number that can represent the result. NaNs can also be used as inputs. NaNs include a sign bit, but the bit has no meaning. NaNs are important for keeping meaningless results from being operated upon. Except where otherwise explicitly stated, when the result of a floating-point operation is a qNaN, it is the RISC-V canonical NaN. For BF16, the RISC-V canonical NaN corresponds to the pattern of _0x7fc0_ which is the most significant 16 bits of the RISC-V single-precision canonical NaN. ##### [](#25-1-3-2-4-scalar-nan-boxing)25.1.3.2.4\. Scalar NaN Boxing RISC-V applies NaN boxing to scalar results and checks for NaN boxing when a floating-point operation --- even a vector-scalar operation --- consumes a value from a scalar floating-point register. If the value is properly NaN-boxed, its least significant bits are used as the operand, otherwise it is treated as if it were the canonical QNaN. NaN boxing is nothing more than putting the smaller encoding in the least significant bits of a register and setting all of the more significant bits to “1”. This matches the encoding of a qNaN (although not the canonical NaN) in the larger precision. Nan-boxing never affects the value of the operand itself, it just changes the bits of the register that are more significant than the operand’s most significant bit. ##### [](#25-1-3-2-5-rounding-modes)25.1.3.2.5\. Rounding Modes: As is the case with other floating-point instructions, the BF16 instructions support all 5 RISC-V Floating-point rounding modes. These modes can be specified in the `rm` field of scalar instructions as well as in the `frm` CSR __Table 3\. RISC-V Floating Point Rounding Modes__ | Rounding Mode | Mnemonic | Meaning | | ------------- | -------- | --------------------------------------- | | 000 | RNE | Round to Nearest, ties to Even | | 001 | RTZ | Round towards Zero | | 010 | RDN | Round Down (towards −∞) | | 011 | RUP | Round Up (towards +∞) | | 100 | RMM | Round to Nearest, ties to Max Magnitude | As with other scalar floating-point instructions, the rounding mode field`rm` can also take on the`DYN` encoding, which indicates that the instruction uses the rounding mode specified in the `frm` CSR. __Table 4\. Additional encoding for the rm field of scalar instructions__ | Rounding Mode | Mnemonic | Meaning | | ------------- | -------- | ---------------------------- | | 111 | DYN | select dynamic rounding mode | In practice, the default IEEE rounding mode (round to nearest, ties to even) is generally used for arithmetic. ##### [](#25-1-3-2-6-handling-exceptions)25.1.3.2.6\. Handling exceptions RISC-V supports IEEE-defined default exception handling. BF16 is no exception. Default exception handling, as defined by IEEE, is a simple and effective approach to producing results in exceptional cases. For the coder to be able to see what has happened, and take further action if needed, BF16 instructions set floating-point exception flags the same way as all other floating-point instructions in RISC-V. ###### [](#25-1-3-2-6-1-underflow)25.1.3.2.6.1\. Underflow The IEEE-defined underflow exception requires that a result be inexact and tiny, where tininess can be detected before or after rounding. In RISC-V, tininess is detected after rounding. It is important to note that the detection of tininess after rounding requires its own rounding that is different from the final result rounding. This tininess detection requires rounding as if the exponent were unbounded. This means that the input to the rounder is always a normal number. This is different from the final result rounding where the input to the rounder is a subnormal number when the value is too small to be represented as a normal number in the target format. The two different roundings can result in underflow being signalled for results that are rounded back to the normal range. As is defined in '754, under default exception handling, underflow is only signalled when the result is tiny and inexact. In such a case, both the underflow and inexact flags are raised. ### [](#BF16%5Fextensions)25.1.4\. Extensions The group of extensions introduced by the BF16 Instruction Set Extensions is listed here. Detection of individual BF16 extensions uses the unified software-based RISC-V discovery method. | | At the time of writing, these discovery mechanisms are still a work in progress. | | ----------------------------------------------------------------------------------- | The BF16 extensions defined in this specification (i.e., `Zfbfmin`,`Zvfbfmin`, and `Zvfbfwma`) depend on the single-precision floating-point extension`F`. Furthermore, the vector BF16 extensions (i.e.,`Zvfbfmin`, and`Zvfbfwma`) depend on the `"V"` Vector Extension for Application Processors or the `Zve32f` Vector Extension for Embedded Processors. As stated later in this specification, there exists a dependency between the newly defined extensions:`Zvfbfwma` depends on `Zfbfmin`and `Zvfbfmin`. This initial set of BF16 extensions provides very basic functionality including scalar and vector conversion between BF16 and single-precision values, and vector widening multiply-accumulate instructions. #### [](#zfbfmin)25.1.4.1\. `Zfbfmin` \- Scalar BF16 Converts This extension provides the minimal set of instructions needed to enable scalar support of the BF16 format. It enables BF16 as an interchange format as it provides conversion between BF16 values and FP32 values. This extension depends upon the single-precision floating-point extension`F`. This extension includes six instructions: the `FCVT.BF16.S` and `FCVT.S.BF16`instructions, defined below, and the `FLH`, `FSH`, `FMV.X.H`, and `FMV.H.X`instructions, defined in ["Zfh" Extension for Half-Precision Floating-Point](zfh.html#chap:zfh). | | While conversion instructions tend to include all supported formats, in these extensions we only support conversion between BF16 and FP32 as we are targeting a special use case. These extensions are intended to support the case where BF16 values are used as reduced precision versions of FP32 values, where use of BF16 provides a two-fold advantage for storage, bandwidth, and computation. In this use case, the BF16 values are typically multiplied by each other and accumulated into FP32 sums. These sums are typically converted to BF16 and then used as subsequent inputs. The operations on the BF16 values can be performed on the CPU or a loosely coupled coprocessor. Subsequent extensions might provide support for native BF16 arithmetic. Such extensions could add additional conversion instructions to allow all supported formats to be converted to and from BF16. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | BF16 addition, subtraction, multiplication, division, and square-root operations can be faithfully emulated by converting the BF16 operands to single-precision, performing the operation using single-precision arithmetic, and then converting back to BF16\. Performing BF16 fused multiply-addition using this method can produce results that differ by 1-ulp on some inputs for the RNE and RMM rounding modes. Conversions between BF16 and formats larger than FP32 can be emulated. Exact widening conversions from BF16 can be synthesized by first converting to FP32 and then converting from FP32 to the target precision. Conversions narrowing to BF16 can be synthesized by first converting to FP32 through a series of halving steps and then converting from FP32 to BF16\. As with the fused multiply-addition instruction described above, this method of converting values to BF16 can be off by 1-ulp on some inputs for the RNE and RMM rounding modes. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Mnemonic | Instruction | | ----------- | ------------------------------------------ | | FCVT.BF16.S | [Convert FP32 to BF16](#insns-fcvt.bf16.s) | | FCVT.S.BF16 | [Convert BF16 to FP32](#insns-fcvt.s.bf16) | | FLH | | | FSH | | | FMV.H.X | | | FMV.X.H | | #### [](#zvfbfmin)25.1.4.2\. `Zvfbfmin` \- Vector BF16 Converts This extension provides the minimal set of instructions needed to enable vector support of the BF16 format. It enables BF16 as an interchange format as it provides conversion between BF16 values and FP32 values. This extension depends upon `Zve32f` vector extension. | | While conversion instructions tend to include all supported formats, in these extensions we only support conversion between BF16 and FP32 as we are targeting a special use case. These extensions are intended to support the case where BF16 values are used as reduced precision versions of FP32 values, where use of BF16 provides a two-fold advantage for storage, bandwidth, and computation. In this use case, the BF16 values are typically multiplied by each other and accumulated into FP32 sums. These sums are typically converted to BF16 and then used as subsequent inputs. The operations on the BF16 values can be performed on the CPU or a loosely coupled coprocessor. Subsequent extensions might provide support for native BF16 arithmetic. Such extensions could add additional conversion instructions to allow all supported formats to be converted to and from BF16. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | BF16 addition, subtraction, multiplication, division, and square-root operations can be faithfully emulated by converting the BF16 operands to single-precision, performing the operation using single-precision arithmetic, and then converting back to BF16\. Performing BF16 fused multiply-addition using this method can produce results that differ by 1-ulp on some inputs for the RNE and RMM rounding modes. Conversions between BF16 and formats larger than FP32 can be faithfully emulated. Exact widening conversions from BF16 can be synthesized by first converting to FP32 and then converting from FP32 to the target precision. Conversions narrowing to BF16 can be synthesized by first converting to FP32 through a series of halving steps using vector round-towards-odd narrowing conversion instructions (_vfncvt.rod.f.f.w_). The final convert from FP32 to BF16 would use the desired rounding mode. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Mnemonic | Instruction | | ---------------- | ------------------------------------------------------ | | vfncvtbf16.f.f.w | [Vector convert FP32 to BF16](#insns-vfncvtbf16.f.f.w) | | vfwcvtbf16.f.f.v | [Vector convert BF16 to FP32](#insns-vfwcvtbf16.f.f.v) | #### [](#zvfbfwma)25.1.4.3\. `Zvfbfwma` \- Vector BF16 widening mul-add This extension provides a vector widening BF16 mul-add instruction that accumulates into FP32. This extension depends upon the `Zvfbfmin` extension and the `Zfbfmin` extension. | Mnemonic | Instruction | | ----------- | -------------------------------------------------------------- | | VFWMACCBF16 | [Vector BF16 widening multiply-accumulate](#insns-vfwmaccbf16) | ### [](#BF16%5Finsns)25.1.5\. Instructions #### [](#insns-fcvt.bf16.s)25.1.5.1\. fcvt.bf16.s Synopsis Convert FP32 value to a BF16 value Mnemonic fcvt.bf16.s rd, rs1 Encoding ![svg](_images/svg-c3344c3fe208c528d45a3eec99aeeacb903c2e36.svg) | | Encoding While the mnemonic of this instruction is consistent with that of the other RISC-V floating-point convert instructions, a new encoding is used in bits 24:20. BF16.S and H are used to signify that the source is FP32 and the destination is BF16. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Description Narrowing convert FP32 value to a BF16 value. Round according to the RM field. This instruction is similar to other narrowing floating-point-to-floating-point conversion instructions. Exceptions: Overflow, Underflow, Inexact, Invalid Included in: [Zfbfmin](#zfbfmin) #### [](#insns-fcvt.s.bf16)25.1.5.2\. fcvt.s.bf16 Synopsis Convert BF16 value to an FP32 value Mnemonic fcvt.s.bf16 rd, rs1 Encoding ![svg](_images/svg-691e04e8cad74b80b324a1ef856775bca9ae19b1.svg) | | Encoding While the mnemonic of this instruction is consistent with that of the other RISC-V floating-point convert instructions, a new encoding is used in bits 24:20 to indicate that the source is BF16. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Description Converts a BF16 value to an FP32 value. The conversion is exact. This instruction is similar to other widening floating-point-to-floating-point conversion instructions. | | If the input is normal or infinity, the BF16 encoded value is shifted to the left by 16 places and the least significant 16 bits are written with 0s. The result is NaN-boxed by writing the most significant FLEN\-32 bits with 1s. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Exceptions: Invalid Included in: [Zfbfmin](#zfbfmin) #### [](#insns-vfncvtbf16.f.f.w)25.1.5.3\. vfncvtbf16.f.f.w Synopsis Vector convert FP32 to BF16 Mnemonic vfncvtbf16.f.f.w vd, vs2, vm Encoding ![svg](_images/svg-41928a051987580541fccbf6b71413083ea860a3.svg) Reserved Encodings * `SEW` is any value other than 16 Arguments | Register | Direction | EEW | Definition | | -------- | --------- | --- | ----------- | | Vs2 | input | 32 | FP32 Source | | Vd | output | 16 | BF16 Result | Description Narrowing convert from FP32 to BF16\. Round according to the _frm_ register. This instruction is similar to `vfncvt.f.f.w` which converts a floating-point value in a 2\*SEW-width format into an SEW-width format. However, here the SEW-width format is limited to BF16. Exceptions: Overflow, Underflow, Inexact, Invalid Included in: [Zvfbfmin](#zvfbfmin) #### [](#insns-vfwcvtbf16.f.f.v)25.1.5.4\. vfwcvtbf16.f.f.v Synopsis Vector convert BF16 to FP32 Mnemonic vfwcvtbf16.f.f.v vd, vs2, vm Encoding ![svg](_images/svg-876e42a51e5e60b4b3cf9eb0f3a4fe8d544ec5db.svg) Reserved Encodings * `SEW` is any value other than 16 Arguments | Register | Direction | EEW | Definition | | -------- | --------- | --- | ----------- | | Vs2 | input | 16 | BF16 Source | | Vd | output | 32 | FP32 Result | Description Widening convert from BF16 to FP32\. The conversion is exact. This instruction is similar to `vfwcvt.f.f.v` which converts a floating-point value in an SEW-width format into a 2\*SEW-width format. However, here the SEW-width format is limited to BF16. | | If the input is normal or infinity, the BF16 encoded value is shifted to the left by 16 places and the least significant 16 bits are written with 0s. | | -------------------------------------------------------------------------------------------------------------------------------------------------------- | Exceptions: Invalid Included in: [Zvfbfmin](#zvfbfmin) #### [](#insns-vfwmaccbf16)25.1.5.5\. vfwmaccbf16 Synopsis Vector BF16 widening multiply-accumulate Mnemonic vfwmaccbf16.vv vd, vs1, vs2, vm vfwmaccbf16.vf vd, rs1, vs2, vm Encoding (Vector-Vector) ![svg](_images/svg-24a0efe51fc715d8764f3a9e99436dd6e65b1b06.svg) Encoding (Vector-Scalar) ![svg](_images/svg-3f9947ea21fbbc7a26407548674e230e2b11ab33.svg) Reserved Encodings * `SEW` is any value other than 16 Arguments | Register | Direction | EEW | Definition | | -------- | --------- | --- | --------------- | | Vd | input | 32 | FP32 Accumulate | | Vs1/rs1 | input | 16 | BF16 Source | | Vs2 | input | 16 | BF16 Source | | Vd | output | 32 | FP32 Result | Description This instruction performs a widening fused multiply-accumulate operation, where each pair of BF16 values are multiplied and their unrounded product is added to the corresponding FP32 accumulate value. The sum is rounded according to the _frm_ register. In the vector-vector version, the BF16 elements are read from `vs1`and `vs2` and FP32 accumulate value is read from `vd`. The FP32 result is written to the destination register `vd`. The vector-scalar version is similar, but instead of reading elements from `vs1`, a scalar BF16 value is read from the FPU register `rs1`. Exceptions: Overflow, Underflow, Inexact, Invalid Operation This `vfwmaccbf16.vv` instruction is equivalent to widening each of the BF16 inputs to FP32 and then performing an FMACC as shown in the following instruction sequence: ```asm vfwcvtbf16.f.f.v T1, vs1, vm vfwcvtbf16.f.f.v T2, vs2, vm vfmacc.vv vd, T1, T2, vm ``` Likewise, `vfwmaccbf16.vf` is equivalent to the following instruction sequence: ```asm fcvt.s.bf16 T1, rs1 vfwcvtbf16.f.f.v T2, vs2, vm vfmacc.vf vd, T1, T2, vm ``` Included in: [Zvfbfwma](#zvfbfwma) ### [](#25-1-6-bibliography)25.1.6\. Bibliography [754-2019 - IEEE Standard for Floating-Point Arithmetic](https://ieeexplore.ieee.org/document/8766229) [754-2008 - IEEE Standard for Floating-Point Arithmetic](https://ieeexplore.ieee.org/document/4610935) Bit Manipulation Extensions Assembly Code Examples ==================== ## [](#bit-manipulation-extensions-assembly-code-examples)Appendix A: Bit Manipulation Extensions Assembly Code Examples The following examples provide software optimization guidance. ### [](#strlen)strlen The **orc.b** instruction allows for the efficient detection of **NUL** bytes in an XLEN-sized chunk of data: * the result of **orc.b** on a chunk that does not contain any **NUL** bytes will be all-ones, and * after a bitwise-negation of the result of **orc.b**, the number of data bytes before the first **NUL** byte (if any) can be detected by **ctz**/**clz** (depending on the endianness of data). A full example of a **strlen** function, which uses these techniques and also demonstrates the use of it for unaligned/partial data, is the following: ```asm #include .text .globl strlen .type strlen, @function strlen: andi a3, a0, (SZREG-1) // offset andi a1, a0, -SZREG // align pointer .Lprologue: li a4, SZREG sub a4, a4, a3 // XLEN - offset slli a3, a3, 3 // offset * 8 REG_L a2, 0(a1) // chunk /* * Shift the partial/unaligned chunk we loaded to remove the bytes * from before the start of the string, adding NUL bytes at the end. */ #if __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__ srl a2, a2 ,a3 // chunk >> (offset * 8) #else sll a2, a2, a3 #endif orc.b a2, a2 not a2, a2 /* * Non-NUL bytes in the string have been expanded to 0x00, while * NUL bytes have become 0xff. Search for the first set bit * (corresponding to a NUL byte in the original chunk). */ #if __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__ ctz a2, a2 #else clz a2, a2 #endif /* * The first chunk is special: compare against the number of valid * bytes in this chunk. */ srli a0, a2, 3 bgtu a4, a0, .Ldone addi a3, a1, SZREG li a4, -1 .align 2 /* * Our critical loop is 4 instructions and processes data in 4 byte * or 8 byte chunks. */ .Lloop: REG_L a2, SZREG(a1) addi a1, a1, SZREG orc.b a2, a2 beq a2, a4, .Lloop .Lepilogue: not a2, a2 #if __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__ ctz a2, a2 #else clz a2, a2 #endif sub a1, a1, a3 add a0, a0, a1 srli a2, a2, 3 add a0, a0, a2 .Ldone: ret ``` ### [](#strcmp)strcmp ```asm #include .text .globl strcmp .type strcmp, @function strcmp: or a4, a0, a1 li t2, -1 and a4, a4, SZREG-1 bnez a4, .Lsimpleloop # Main loop for aligned strings .Lloop: REG_L a2, 0(a0) REG_L a3, 0(a1) orc.b t0, a2 bne t0, t2, .Lfoundnull addi a0, a0, SZREG addi a1, a1, SZREG beq a2, a3, .Lloop # Words don't match, and no null byte in first word. # Get bytes in big-endian order and compare. #if __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__ rev8 a2, a2 rev8 a3, a3 #endif # Synthesize (a2 >= a3) ? 1 : -1 in a branchless sequence. sltu a0, a2, a3 neg a0, a0 ori a0, a0, 1 ret .Lfoundnull: # Found a null byte. # If words don't match, fall back to simple loop. bne a2, a3, .Lsimpleloop # Otherwise, strings are equal. li a0, 0 ret # Simple loop for misaligned strings .Lsimpleloop: lbu a2, 0(a0) lbu a3, 0(a1) addi a0, a0, 1 addi a1, a1, 1 bne a2, a3, 1f bnez a2, .Lsimpleloop 1: sub a0, a2, a3 ret .size strcmp, .-strcmp ``` 28.1. "C" Extension for Compressed Instructions, Version 2.0 ==================== ## [](#compressed)28.1\. "C" Extension for Compressed Instructions, Version 2.0 This chapter describes the RISC-V standard compressed instruction-set extension, named "C", which reduces static and dynamic code size by adding short 16-bit instruction encodings for common operations. The C extension can be added to any of the base ISAs (RV32I, RV32E, RV64I, RV64E), and we use the generic term "RVC" to cover any of these. Typically, 50%-60% of the RISC-V instructions in a program can be replaced with RVC instructions, resulting in a 25%-30% code-size reduction. ### [](#28-1-1-overview)28.1.1\. Overview RVC uses a simple compression scheme that offers shorter 16-bit versions of common 32-bit RISC-V instructions when: * the immediate or address offset is small, or * one of the registers is the zero register (`x0`), the ABI link register (`x1`), or the ABI stack pointer (`x2`), or * the destination register and the first source register are identical, or * the registers used are the 8 most popular ones. The C extension is compatible with all other standard instruction extensions. The C extension allows 16-bit instructions to be freely intermixed with 32-bit instructions, with the latter now able to start on any 16-bit boundary, i.e., IALIGN=16. With the addition of the C extension, no instructions can raise instruction-address-misaligned exceptions. | | Removing the 32-bit alignment constraint on the original 32-bit instructions allows significantly greater code density. | | -------------------------------------------------------------------------------------------------------------------------- | The compressed instruction encodings are mostly common across RV32C and RV64C, but as shown in [Figure 1](#rvc-instr-table0), a few opcodes are used for different purposes depending on base ISA. For example, the wider address-space RV64C variant requires additional opcodes to compress loads and stores of 64-bit integer values, while RV32C uses the same opcodes to compress loads and stores of single-precision floating-point values. If the C extension is implemented, the appropriate compressed floating-point load and store instructions must be provided whenever the relevant standard floating-point extension (F and/or D) is also implemented. In addition, RV32C includes a compressed jump and link instruction to compress short-range subroutine calls, where the same opcode is used to compress ADDIW for RV64C. | | Double-precision loads and stores are a significant fraction of static and dynamic instructions, hence the motivation to include them in the RV32C and RV64C encoding. Although single-precision loads and stores are not a significant source of static or dynamic compression for benchmarks compiled for the currently supported ABIs, for microcontrollers that only provide hardware single-precision floating-point units and have an ABI that only supports single-precision floating-point numbers, the single-precision loads and stores will be used at least as frequently as double-precision loads and stores in the measured benchmarks. Hence, the motivation to provide compressed support for these in RV32C. Short-range subroutine calls are more likely in small binaries for microcontrollers, hence the motivation to include these in RV32C. Although reusing opcodes for different purposes for different base ISAs adds some complexity to documentation, the impact on implementation complexity is small even for designs that support multiple base ISAs. The compressed floating-point load and store variants use the same instruction format with the same register specifiers as the wider integer loads and stores. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | RVC was designed under the constraint that each RVC instruction expands into a single 32-bit instruction in either the base ISA (RV32I/E or RV64I/E) or the F and D standard extensions where present. Adopting this constraint has two main benefits: * Hardware designs can simply expand RVC instructions during decode, simplifying verification and minimizing modifications to existing microarchitectures. * Compilers can be unaware of the RVC extension and leave code compression to the assembler and linker, although a compression-aware compiler will generally be able to produce better results. | | We felt the multiple complexity reductions of a simple one-one mapping between C and base IFD instructions far outweighed the potential gains of a slightly denser encoding that added additional instructions only supported in the C extension, or that allowed encoding of multiple IFD instructions in one C instruction. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | It is important to note that the C extension is not designed to be a stand-alone ISA, and is meant to be used alongside a base ISA. | | Variable-length instruction sets have long been used to improve code density. For example, the IBM Stretch \[[22](../biblio/bibliography.html#bib-stretch)\], developed in the late 1950s, had an ISA with 32-bit and 64-bit instructions, where some of the 32-bit instructions were compressed versions of the full 64-bit instructions. Stretch also employed the concept of limiting the set of registers that were addressable in some of the shorter instruction formats, with short branch instructions that could only refer to one of the index registers. The later IBM 360 architecture \[[23](../biblio/bibliography.html#bib-ibm360)\] supported a simple variable-length instruction encoding with 16-bit, 32-bit, or 48-bit instruction formats. In 1963, CDC introduced the Cray-designed CDC 6600 \[[24](../biblio/bibliography.html#bib-cdc6600)\], a precursor to RISC architectures, that introduced a register-rich load-store architecture with instructions of two lengths, 15-bits and 30-bits. The later Cray-1 design used a very similar instruction format, with 16-bit and 32-bit instruction lengths. The initial RISC ISAs from the 1980s all picked performance over code size, which was reasonable for a workstation environment, but not for embedded systems. Hence, both ARM and MIPS subsequently made versions of the ISAs that offered smaller code size by offering an alternative 16-bit wide instruction set instead of the standard 32-bit wide instructions. The compressed RISC ISAs reduced code size relative to their starting points by about 25-30%, yielding code that was significantly smaller than 80x86\. This result surprised some, as their intuition was that the variable-length CISC ISA should be smaller than RISC ISAs that offered only 16-bit and 32-bit formats. Since the original RISC ISAs did not leave sufficient opcode space free to include these unplanned compressed instructions, they were instead developed as complete new ISAs. This meant compilers needed different code generators for the separate compressed ISAs. The first compressed RISC ISA extensions (e.g., ARM Thumb and MIPS16) used only a fixed 16-bit instruction size, which gave good reductions in static code size but caused an increase in dynamic instruction count, which led to lower performance compared to the original fixed-width 32-bit instruction size. This led to the development of a second generation of compressed RISC ISA designs with mixed 16-bit and 32-bit instruction lengths (e.g., ARM Thumb2, microMIPS, PowerPC VLE), so that performance was similar to pure 32-bit instructions but with significant code size savings. Unfortunately, these different generations of compressed ISAs are incompatible with each other and with the original uncompressed ISA, leading to significant complexity in documentation, implementations, and software tools support. Of the commonly used 64-bit ISAs, only PowerPC and microMIPS currently supports a compressed instruction format. It is surprising that the most popular 64-bit ISA for mobile platforms (ARM v8) does not include a compressed instruction format given that static code size and dynamic instruction fetch bandwidth are important metrics. Although static code size is not a major concern in larger systems, instruction fetch bandwidth can be a major bottleneck in servers running commercial workloads, which often have a large instruction working set. Benefiting from 25 years of hindsight, RISC-V was designed to support compressed instructions from the outset, leaving enough opcode space for RVC to be added as a simple extension on top of the base ISA (along with many other extensions). The philosophy of RVC is to reduce code size for embedded applications _and_ to improve performance and energy-efficiency for all applications due to fewer misses in the instruction cache. Waterman shows that RVC fetches 25%-30% fewer instruction bits, which reduces instruction cache misses by 20%-25%, or roughly the same performance impact as doubling the instruction cache size. \[[25](../biblio/bibliography.html#bib-waterman-ms)\] | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#28-1-2-compressed-instruction-formats)28.1.2\. Compressed Instruction Formats [Table 1](#rvc-form) shows the nine compressed instruction formats. CR, CI, and CSS can use any of the 32 RVI registers, but CIW, CL, CS, CA, and CB are limited to just 8 of them.[Table 2](#registers) lists these popular registers, which correspond to registers `x8` to `x15`. Note that there is a separate version of load and store instructions that use the stack pointer as the base address register, since saving to and restoring from the stack are so prevalent, and that they use the CI and CSS formats to allow access to all 32 data registers. CIW supplies an 8-bit immediate for the ADDI4SPN instruction. | | The RISC-V ABI was changed to make the frequently used registers map to registers 'x8-x15'. This simplifies the decompression decoder by having a contiguous naturally aligned set of register numbers, and is also compatible with the RV32E and RV64E base ISAs, which only have 16 integer registers. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Compressed register-based floating-point loads and stores also use the CL and CS formats respectively, with the eight registers mapping to `f8` to `f15`. | | _The standard RISC-V calling convention maps the most frequently used floating-point registers to registers f8 to f15, which allows the same register decompression decoding as for integer register numbers._ | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The formats were designed to keep bits for the two register source specifiers in the same place in all instructions, while the destination register field can move. When the full 5-bit destination register specifier is present, it is in the same place as in the 32-bit RISC-V encoding. Where immediates are sign-extended, the sign extension is always from bit 12\. Immediate fields have been scrambled, as in the base specification, to reduce the number of immediate multiplexers required. | | The immediate fields are scrambled in the instruction formats instead of in sequential order so that as many bits as possible are in the same position in every instruction, thereby simplifying implementations. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For many RVC instructions, zero-valued immediates are disallowed and`x0` is not a valid 5-bit register specifier. These restrictions free up encoding space for other instructions requiring fewer operand bits. __Table 1\. Compressed 16-bit RVC instruction formats__ | Format Meaning CR Register CI Immediate CSS Stack-relative Store CIW Wide Immediate CL Load CS Store CA Arithmetic CB Branch/Arithmetic CJ Jump | 15 14 13 12 11 10 9 8 7 6 5 4 3 2 1 0 funct4 rd/rs1 rs2 op funct3 imm rd/rs1 imm op funct3 imm rs2 op funct3 imm rd′ op funct3 imm rs1′ imm rd′ op funct3 imm rs1′ imm rs2′ op funct6 rd′/rs1′ funct2 rs2′ op funct3 offset rd′/rs1′ offset op funct3 jump target op | | ----------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 2\. Registers specified by the three-bit _rs1_′, _rs2_′, and _rd_′ fields of the CIW, CL, CS, CA, and CB formats.__ | RVC Register Number Integer Register Number Integer Register ABI Name Floating-Point Register Number Floating-Point Register ABI Name | 000 001 010 011 100 101 110 111 x8 x9 x10 x11 x12 x13 x14 x15 s0 s1 a0 a1 a2 a3 a4 a5 f8 f9 f10 f11 f12 f13 f14 f15 fs0 fs1 fa0 fa1 fa2 fa3 fa4 fa5 | | ------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#28-1-3-load-and-store-instructions)28.1.3\. Load and Store Instructions To increase the reach of 16-bit instructions, data-transfer instructions use zero-extended immediates that are scaled by the size of the data in bytes: ×4 for words, ×8 for double words, and ×16 for quad words. RVC provides two variants of loads and stores. One uses the ABI stack pointer, `x2`, as the base address and can target any data register. The other can reference one of 8 base address registers and one of 8 data registers. #### [](#28-1-3-1-stack-pointer-based-loads-and-stores)28.1.3.1\. Stack-Pointer-Based Loads and Stores ![svg](_images/svg-e18f1e1e5092d13a44025b07aad9b6da08b9b32f.svg) These instructions use the CI format. C.LWSP loads a 32-bit value from memory into register _rd_. It computes an effective address by adding the _zero_\-extended offset, scaled by 4, to the stack pointer, `x2`. It expands to `lw rd, offset(x2)`. C.LWSP is valid only when _rd_≠`x0`; the code points with _rd_\=`x0` are reserved. C.LDSP is an RV64C-only instruction that loads a 64-bit value from memory into register _rd_. It computes its effective address by adding the zero-extended offset, scaled by 8, to the stack pointer,`x2`. It expands to `ld rd, offset(x2)`. C.LDSP is valid only when_rd_≠`x0`; the code points with_rd_\=`x0` are reserved. C.FLWSP is an RV32FC-only instruction that loads a single-precision floating-point value from memory into floating-point register _rd_. It computes its effective address by adding the _zero_\-extended offset, scaled by 4, to the stack pointer, `x2`. It expands to`flw rd, offset(x2)`. C.FLDSP is an RV32DC/RV64DC-only instruction that loads a double-precision floating-point value from memory into floating-point register _rd_. It computes its effective address by adding the_zero_\-extended offset, scaled by 8, to the stack pointer, `x2`. It expands to `fld rd, offset(x2)`. ![svg](_images/svg-b48d94bf6f0d97e8165edd4f1a2bcfde7aafd511.svg) These instructions use the CSS format. C.SWSP stores a 32-bit value in register _rs2_ to memory. It computes an effective address by adding the _zero_\-extended offset, scaled by 4, to the stack pointer, `x2`. It expands to `sw rs2, offset(x2)`. C.SDSP is an RV64C-only instruction that stores a 64-bit value in register _rs2_ to memory. It computes an effective address by adding the_zero_\-extended offset, scaled by 8, to the stack pointer, `x2`. It expands to `sd rs2, offset(x2)`. C.FSWSP is an RV32FC-only instruction that stores a single-precision floating-point value in floating-point register _rs2_ to memory. It computes an effective address by adding the _zero_\-extended offset, scaled by 4, to the stack pointer, `x2`. It expands to`fsw rs2, offset(x2)`. C.FSDSP is an RV32DC/RV64DC-only instruction that stores a double-precision floating-point value in floating-point register _rs2_to memory. It computes an effective address by adding the_zero_\-extended offset, scaled by 8, to the stack pointer, `x2`. It expands to `fsd rs2, offset(x2)`. | | Register save/restore code at function entry/exit represents a significant portion of static code size. The stack-pointer-based compressed loads and stores in RVC are effective at reducing the save/restore static code size by a factor of 2 while improving performance by reducing dynamic instruction bandwidth. A common mechanism used in other ISAs to further reduce save/restore code size is load-multiple and store-multiple instructions. We considered adopting these for RISC-V but noted the following drawbacks to these instructions: These instructions complicate processor implementations. For virtual memory systems, some data accesses could be resident in physical memory and some could not, which requires a new restart mechanism for partially executed instructions. Unlike the rest of the RVC instructions, there is no IFD equivalent to Load Multiple and Store Multiple. Unlike the rest of the RVC instructions, the compiler would have to be aware of these load-multiple and store-multiple instructions to both allocate registers in the expected order and also to schedule the loads and stores contiguously and in the proper order, to maximize the chances of them being detected and replaced by an assembler or linker with the equivalent load-multiple or store-multiple compressed instruction. Simple microarchitectural implementations will constrain how other instructions can be scheduled around the load and store multiple instructions, leading to a potential performance loss. The desire for sequential register allocation might conflict with the featured registers selected for the CIW, CL, CS, CA, and CB formats. Furthermore, much of the gains can be realized in software by replacing prologue and epilogue code with subroutine calls to common prologue and epilogue code, a technique described in Section 5.6 of \[[26](../biblio/bibliography.html#bib-waterman-phd)\]. While reasonable architects might come to different conclusions, we decided to omit load and store multiple and instead use the software-only approach of calling save/restore millicode routines to attain the greatest code size reduction. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#28-1-3-2-register-based-loads-and-stores)28.1.3.2\. Register-Based Loads and Stores ![svg](_images/svg-a6c6023cddbfc673ff34f9d8b654c3fd6e2449a0.svg) These instructions use the CL format. C.LW loads a 32-bit value from memory into register`_rd′_`. It computes an effective address by adding the_zero_\-extended offset, scaled by 4, to the base address in register`_rs1′_`. It expands to `lw rd′, offset(rs1′)`. C.LD is an RV64C-only instruction that loads a 64-bit value from memory into register `_rd′_`. It computes an effective address by adding the _zero_\-extended offset, scaled by 8, to the base address in register `_rs1′_`. It expands to`ld rd′, offset(rs1′)`. C.FLW is an RV32FC-only instruction that loads a single-precision floating-point value from memory into floating-point register`_rd′_`. It computes an effective address by adding the_zero_\-extended offset, scaled by 4, to the base address in register`_rs1′_`. It expands to`flw rd′, offset(rs1′)`. C.FLD is an RV32DC/RV64DC-only instruction that loads a double-precision floating-point value from memory into floating-point register`_rd′_`. It computes an effective address by adding the_zero_\-extended offset, scaled by 8, to the base address in register`_rs1′_`. It expands to`fld rd′, offset(rs1′)`. ![svg](_images/svg-95a4adb6faa0b91163a7ca7355571d48945968c1.svg) These instructions use the CS format. C.SW stores a 32-bit value in register `_rs2′_` to memory. It computes an effective address by adding the _zero_\-extended offset, scaled by 4, to the base address in register `_rs1′_`. It expands to `sw rs2′, offset(rs1′)`. C.SD is an RV64C-only instruction that stores a 64-bit value in register `_rs2′_` to memory. It computes an effective address by adding the _zero_\-extended offset, scaled by 8, to the base address in register `_rs1′_`. It expands to`sd rs2′, offset(rs1′)`. C.FSW is an RV32FC-only instruction that stores a single-precision floating-point value in floating-point register `_rs2′_` to memory. It computes an effective address by adding the _zero_\-extended offset, scaled by 4, to the base address in register`_rs1′_`. It expands to`fsw rs2′, offset(rs1′)`. C.FSD is an RV32DC/RV64DC-only instruction that stores a double-precision floating-point value in floating-point register`_rs2′_` to memory. It computes an effective address by adding the _zero_\-extended offset, scaled by 8, to the base address in register `_rs1′_`. It expands to`fsd rs2′, offset(rs1′)`. ### [](#28-1-4-control-transfer-instructions)28.1.4\. Control Transfer Instructions RVC provides unconditional jump instructions and conditional branch instructions. As with base RVI instructions, the offsets of all RVC control transfer instructions are in multiples of 2 bytes. ![svg](_images/svg-1d6d90ff75a468d006ca9343b224bb7d7de17b44.svg) These instructions use the CJ format. C.J performs an unconditional control transfer. The offset is sign-extended and added to the `pc` to form the jump target address. C.J can therefore target a ±2 KiB range. C.J expands to`jal x0, offset`. C.JAL is an RV32C-only instruction that performs the same operation as C.J, but additionally writes the address of the instruction following the jump (`pc+2`) to the link register, `x1`. C.JAL expands to`jal x1, offset`. ![svg](_images/svg-ca378d779bca23f7bbbf5ada7e62f7c987cf1546.svg) These instructions use the CR format. C.JR (jump register) performs an unconditional control transfer to the address in register _rs1_. C.JR expands to `jalr x0, 0(rs1)`. C.JR is valid only when _rs1_≠`x0`; the code point with _rs1_\=`x0` is reserved. C.JALR (jump and link register) performs the same operation as C.JR, but additionally writes the address of the instruction following the jump (`pc`+2) to the link register, `x1`. C.JALR expands to`jalr x1, 0(rs1)`. C.JALR is valid only when_rs1_≠`x0`; the code point with_rs1_\=`x0` corresponds to the C.EBREAK instruction. | | Strictly speaking, C.JALR does not expand exactly to a base RVI instruction as the value added to the PC to form the link address is 2 rather than 4 as in the base ISA, but supporting both offsets of 2 and 4 bytes is only a very minor change to the base microarchitecture. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ![svg](_images/svg-386882b890f10d93eee7c836feee88331e5c9bb6.svg) These instructions use the CB format. C.BEQZ performs conditional control transfers. The offset is sign-extended and added to the `pc` to form the branch target address. It can therefore target a ±256 B range. C.BEQZ takes the branch if the value in register _rs1′_ is zero. It expands to `beq rs1′, x0, offset`. C.BNEZ is defined analogously, but it takes the branch if_rs1′_ contains a nonzero value. It expands to`bne rs1′, x0, offset`. ### [](#28-1-5-integer-computational-instructions)28.1.5\. Integer Computational Instructions RVC provides several instructions for integer arithmetic and constant generation. #### [](#28-1-5-1-integer-constant-generation-instructions)28.1.5.1\. Integer Constant-Generation Instructions The two constant-generation instructions both use the CI instruction format and can target any integer register. ![svg](_images/svg-db6d70581f491dfca4c6c2607b7bef9077753246.svg) C.LI loads the sign-extended 6-bit immediate, _imm_, into register _rd_. C.LI expands into `addi rd, x0, imm`. The C.LI code points with _rd_\=`x0` are HINTs. C.LUI loads the non-zero 6-bit immediate field into bits 17–12 of the destination register, clears the bottom 12 bits, and sign-extends bit 17 into all higher bits of the destination. C.LUI expands into`lui rd, imm`. C.LUI is valid only when_rd_≠`x2`, and when the immediate is not equal to zero. The code points with_imm_\=0 are reserved. The code points with _rd_\=`x2` and _imm_≠0 correspond to the C.ADDI16SP instruction. The code points with _rd_\=`x0` and _imm_≠0 are HINTs. #### [](#28-1-5-2-integer-register-immediate-operations)28.1.5.2\. Integer Register-Immediate Operations These integer register-immediate operations are encoded in the CI format and perform operations on an integer register and a 6-bit immediate. ![svg](_images/svg-bbb362623c0f88d5dbb743ebfc933be4e36eb4a2.svg) C.ADDI adds the non-zero sign-extended 6-bit immediate to the value in register _rd_ then writes the result to _rd_. C.ADDI expands into`addi rd, rd, imm`. The code points with _rd_≠0 and _imm_\=0 are HINTs. The code points with _rd_\=`x0` encode the C.NOP instruction, of which the code points with _imm_≠0 are HINTs. C.ADDIW is an RV64C-only instruction that performs the same computation but produces a 32-bit result, then sign-extends result to 64 bits. C.ADDIW expands into `addiw rd, rd, imm`. The immediate can be zero for C.ADDIW, where this corresponds to `sext.w rd`. C.ADDIW is valid only when _rd_≠`x0`; the code points with_rd_\=`x0` are reserved. C.ADDI16SP (add immediate to stack pointer) shares the opcode with C.LUI, but has a destination field of`x2`. C.ADDI16SP adds the non-zero sign-extended 6-bit immediate to the value in the stack pointer (`sp=x2`), where the immediate is scaled to represent multiples of 16 in the range \[-512, 496\]. C.ADDI16SP is used to adjust the stack pointer in procedure prologues and epilogues. It expands into `addi x2, x2, nzimm[9:4]`. C.ADDI16SP is valid only when_nzimm_≠0; the code point with _nzimm_\=0 is reserved. | | In the standard RISC-V calling convention, the stack pointer sp is always 16-byte aligned. | | --------------------------------------------------------------------------------------------- | ![svg](_images/svg-bf986cd7521109161e78530931e222b3c637ae50.svg) C.ADDI4SPN (add immediate to stack pointer, non-destructive) is a CIW-format instruction that adds a _zero_\-extended non-zero immediate, scaled by 4, to the stack pointer, `x2`, and writes the result to `rd′`. This instruction is used to generate pointers to stack-allocated variables, and expands to`addi rd′, x2, nzuimm[9:2]`. C.ADDI4SPN is valid only when_nzuimm_≠0; the code points with _nzuimm_\=0 are reserved. ![svg](_images/svg-4ceb865372085f49cea50da520cc223645c46d2d.svg) C.SLLI is a CI-format instruction that performs a logical left shift of the value in register _rd_ then writes the result to _rd_. The shift amount is encoded in the _shamt_ field. C.SLLI expands into `slli rd, rd, shamt[5:0]`. The C.SLLI code points with _shamt_\=0 or with _rd_\=`x0` are HINTs. For RV32C, _shamt\[5\]_ must be zero; the code points with _shamt\[5\]_\=1 are designated for custom extensions. ![svg](_images/svg-e63605912406813e686938b80b8f5e76981efe04.svg) C.SRLI is a CB-format instruction that performs a logical right shift of the value in register _rd′_ then writes the result to_rd′_. The shift amount is encoded in the _shamt_ field. C.SRLI expands into `srli rd′, rd′, shamt`. The C.SRLI code points with _shamt_\=0 are HINTs. For RV32C, _shamt\[5\]_ must be zero; the code points with _shamt\[5\]_\=1 are designated for custom extensions. C.SRAI is defined analogously to C.SRLI, but instead performs an arithmetic right shift. C.SRAI expands to`srai rd′, rd′, shamt`. | | Left shifts are usually more frequent than right shifts, as left shifts are frequently used to scale address values. Right shifts have therefore been granted less encoding space and are placed in an encoding quadrant where all other immediates are sign-extended. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ![svg](_images/svg-96adf533629add3edd308ef7105a8aecbfe20bf1.svg) C.ANDI is a CB-format instruction that computes the bitwise AND of the value in register _rd′_ and the sign-extended 6-bit immediate, then writes the result to _rd′_. C.ANDI expands to `andi rd′, rd′, imm`. #### [](#28-1-5-3-integer-register-register-operations)28.1.5.3\. Integer Register-Register Operations ![svg](_images/svg-0522b2220c99b36c1e028b418dc079fa0b55d25c.svg) These instructions use the CR format. C.MV copies the value in register _rs2_ into register _rd_. C.MV expands into `add rd, x0, rs2`. C.MV is valid only when_rs2_≠`x0`; the code points with _rs2_\=`x0` correspond to the C.JR instruction. The code points with _rs2_≠`x0` and _rd_\=`x0` are HINTs. | | _C.MV expands to a different instruction than the canonical MV pseudoinstruction, which instead uses ADDI. Implementations that handle MV specially, e.g. using register-renaming hardware, may find it more convenient to expand C.MV to MV instead of ADD, at slight additional hardware cost._ | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | C.ADD adds the values in registers _rd_ and _rs2_ and writes the result to register _rd_. C.ADD expands into `add rd, rd, rs2`. C.ADD is only valid when _rs2_≠`x0`; the code points with _rs2_\=`x0` correspond to the C.JALR and C.EBREAK instructions. The code points with _rs2_≠`x0` and _rd_\=`x0` are HINTs. ![svg](_images/svg-0ccc2817794606823407179191eb0fd1147858d9.svg) These instructions use the CA format. `C.AND` computes the bitwise `AND` of the values in registers_rd′_ and _rs2′_, then writes the result to register _rd′_. `C.AND` expands into`and rd′, rd′, rs2′`. `C.OR` computes the bitwise `OR` of the values in registers_rd′_ and _rs2′_, then writes the result to register _rd′_. `C.OR` expands into`or rd′, rd′, rs2′`. `C.XOR` computes the bitwise `XOR` of the values in registers_rd′_ and _rs2′_, then writes the result to register _rd′_. `C.XOR` expands into`xor rd′, rd′, rs2′`. `C.SUB` subtracts the value in register _rs2′_ from the value in register _rd′_, then writes the result to register _rd′_. `C.SUB` expands into`sub rd′, rd′, rs2′`. `C.ADDW` is an RV64C-only instruction that adds the values in registers _rd′_ and _rs2′_, then sign-extends the lower 32 bits of the sum before writing the result to register _rd′_. `C.ADDW` expands into`addw rd′, rd′, rs2′`. `C.SUBW` is an RV64C-only instruction that subtracts the value in register _rs2′_ from the value in register_rd′_, then sign-extends the lower 32 bits of the difference before writing the result to register _rd′_.`C.SUBW` expands into `subw rd′, rd′, rs2′`. | | This group of six instructions do not provide large savings individually, but do not occupy much encoding space and are straightforward to implement, and as a group provide a worthwhile improvement in static and dynamic compression. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#28-1-5-4-defined-illegal-instruction)28.1.5.4\. Defined Illegal Instruction ![svg](_images/svg-9cde54c08660055fa18c338b2e96302b49cb7a98.svg) A 16-bit instruction with all bits zero is permanently reserved as an illegal instruction. | | We reserve all-zero instructions to be illegal instructions to help trap attempts to execute zero-ed or non-existent portions of the memory space. The all-zero value should not be redefined in any non-standard extension. Similarly, we reserve instructions with all bits set to 1 (corresponding to very long instructions in the RISC-V variable-length encoding scheme) as illegal to capture another common value seen in non-existent memory regions. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#28-1-5-5-nop-instruction)28.1.5.5\. NOP Instruction ![svg](_images/svg-434d6eda5829ad04680d0622689ab6b207218d23.svg) `C.NOP` is a CI-format instruction that does not change any user-visible state, except for advancing the `pc` and incrementing any applicable performance counters. `C.NOP` expands to `nop`. The `C.NOP` code points with _imm_≠0 encode HINTs. #### [](#28-1-5-6-breakpoint-instruction)28.1.5.6\. Breakpoint Instruction ![svg](_images/svg-27922b9dac22e8a5762388dc50030c283177d2db.svg) Debuggers can use the `C.EBREAK` instruction, which expands to `ebreak`, to cause control to be transferred back to the debugging environment.`C.EBREAK` shares the opcode with the `C.ADD` instruction, but with _rd_ and_rs2_ both zero, thus can also use the `CR` format. ### [](#28-1-6-usage-of-c-instructions-in-lrsc-sequences)28.1.6\. Usage of C Instructions in LR/SC Sequences On implementations that support the C extension, compressed forms of the I instructions permitted inside constrained LR/SC sequences, as described in [HINT Instructions](rv32.html#rv32i-hints), are also permitted inside constrained LR/SC sequences. | | The implication is that any implementation that claims to support both the A and C extensions must ensure that LR/SC sequences containing valid C instructions will eventually complete. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#rvc-hints)28.1.7\. HINT Instructions A portion of the RVC encoding space is reserved for microarchitectural HINTs. Like the HINTs in the RV32I base ISA (see[HINT Instructions](rv32.html#rv32i-hints)), these instructions do not modify any architectural state, except for advancing the `pc` and any applicable performance counters. HINTs are executed as no-ops on implementations that ignore them. RVC HINTs are encoded as computational instructions that do not modify the architectural state, either because _rd_\=`x0` (e.g.`C.ADD _x0_, _t0_`), or because _rd_ is overwritten with a copy of itself (e.g. `C.ADDI _t0_, 0`). | | This HINT encoding has been chosen so that simple implementations can ignore HINTs altogether, and instead execute a HINT as a regular computational instruction that happens not to mutate the architectural state. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | RVC HINTs do not necessarily expand to their RVI HINT counterparts. For example, `C.ADD` _x0_, _a0_ might not encode the same HINT as`ADD` _x0_, _x0_, _a0_. | | The primary reason to not require an RVC HINT to expand to an RVI HINT is that HINTs are unlikely to be compressible in the same manner as the underlying computational instruction. Also, decoupling the RVC and RVI HINT mappings allows the scarce RVC HINT space to be allocated to the most popular HINTs, and in particular, to HINTs that are amenable to macro-op fusion. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | [Table 3](#rvc-t-hints) lists all RVC HINT code points. For RV32C, 78% of the HINT space is reserved for standard HINTs. The remainder of the HINT space is designated for custom HINTs; no standard HINTs will ever be defined in this subspace. __Table 3\. RVC HINT instructions.__ | Instruction | Constraints | Code Points | Purpose | | ----------- | ------------------------------- | -------------------- | -------------------------------------------------------------------------- | | C.NOP | _imm_≠0 | 63 | _Designated for future standard use_ | | C.ADDI | _rd_≠x0, _imm_\=0 | 31 | | | C.LI | _rd_\=x0 | 64 | | | C.LUI | _rd_\=x0, _imm_≠0 | 63 | | | C.MV | _rd_\=x0, _rs2_≠x0 | 31 | | | C.ADD | _rd_\=x0, _rs2_≠x0, _rs2_≠x2-x5 | 27 | | | C.ADD | _rd_\=x0, _rs2_\=x2-x5 | 4 | (rs2=x2) C.NTL.P1 (rs2=x3) C.NTL.PALL (rs2=x4) C.NTL.S1 (rs2=x5) C.NTL.ALL | | C.SLLI | _rd_\=x0 or _imm_\=0 | 63 (RV32), 95 (RV64) | _Designated for custom use_ | | C.SRLI | _imm_\=0 | 8 | | | C.SRAI | _imm_\=0 | 8 | | ### [](#28-1-8-rvc-instruction-set-listings)28.1.8\. RVC Instruction Set Listings [Table 4](#rvcopcodemap) shows a map of the major opcodes for RVC. Each row of the table corresponds to one quadrant of the encoding space. The last quadrant, which has the two least-significant bits set, corresponds to instructions wider than 16 bits, including those in the base ISAs. Several instructions are only valid for certain operands; when invalid, they are marked either _RES_to indicate that the opcode is reserved for future standard extensions;_Custom_ to indicate that the opcode is designated for custom extensions; or _HINT_ to indicate that the opcode is reserved for microarchitectural hints (see [28.1.7\. HINT Instructions](#rvc-hints)). __Table 4\. RVC opcode map instructions.__ | inst\[15:13\]inst\[1:0\] | **000** | **001** | **010** | **011** | **100** | **101** | **110** | **111** | | | ------------------------ | -------- | ---------- | ------- | ------------ | --------------- | ---------- | ------- | --------- | -------- | | 00 | ADDI4SPN | FLDFLD | LW | FLWLD | _Reserved_ | FSDFSD | SW | FSWSD | RV32RV64 | | 01 | ADDI | JALADDIW | LI | LUI/ADDI16SP | MISC-ALU | J | BEQZ | BNEZ | RV32RV64 | | 10 | SLLI | FLDSPFLDSP | LWSP | FLWSPLDSP | J\[AL\]R/MV/ADD | FSDSPFSDSP | SWSP | FSWSPSDSP | RV32RV64 | | 11 | \>16b | | | | | | | | | [Figure 1](#rvc-instr-table0), [Figure 2](#rvc-instr-table1), and [Figure 3](#rvc-instr-table2) list the RVC instructions. ![Instruction listing for RVC, Quadrant 0](_images/diag-1f91aa9c4226801a769871aec6df542e2f647a16.svg) Figure 1\. Instruction listing for RVC, Quadrant 0 ![Instruction listing for RVC, Quadrant 1](_images/diag-ac100c399208ef3eb92487b098c395e35d378dbf.svg) Figure 2\. Instruction listing for RVC, Quadrant 1 ![Instruction listing for RVC, Quadrant 2](_images/diag-8418563ff7fdb02d02c93d904f77a522a33c22e2.svg) Figure 3\. Instruction listing for RVC, Quadrant 2 Calling Convention for Vector State (Not authoritative - Placeholder Only) ==================== ## [](#calling-convention-for-vector-state-not-authoritative-placeholder-only)Appendix A: Calling Convention for Vector State (Not authoritative - Placeholder Only) | | This Appendix is only a placeholder to help explain the conventions used in the code examples, and is not considered frozen or part of the ratification process. The official RISC-V psABI document is being expanded to specify the vector calling conventions. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | In the RISC-V psABI, the vector registers `v0`\-`v31` are all caller-saved. The `vl` and `vtype` CSRs are also caller-saved. Procedures may assume that `vstart` is zero upon entry. Procedures may assume that `vstart` is zero upon return from a procedure call. | | Application software should normally not write vstart explicitly. Any procedure that does explicitly write vstart to a nonzero value must zero vstart before either returning or calling another procedure. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `vxrm` and `vxsat` fields of `vcsr` have thread storage duration. Executing a system call causes all caller-saved vector registers (`v0`\-`v31`, `vl`, `vtype`) and `vstart` to become unspecified. | | This scheme allows system calls that cause context switches to avoid saving and later restoring the vector registers. | | ------------------------------------------------------------------------------------------------------------------------ | | | Most OSes will choose to either leave these registers intact or reset them to their initial state to avoid leaking information across process boundaries. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | 20.1. "CMO" Extensions for Base Cache Management Operation ISA, Version 1.0.0 ==================== ## [](#cmo)20.1\. "CMO" Extensions for Base Cache Management Operation ISA, Version 1.0.0 ### [](#20-1-1-pseudocode-for-instruction-semantics)20.1.1\. Pseudocode for instruction semantics The semantics of each instruction in the [Instructions](#insns) chapter is expressed in a SAIL-like syntax. ### [](#intro-cmo)20.1.2\. Introduction _Cache-management operation_ (or _CMO_) instructions perform operations on copies of data in the memory hierarchy. In general, CMO instructions operate on cached copies of data, but in some cases, a CMO instruction may operate on memory locations directly. Furthermore, CMO instructions are grouped by operation into the following classes: * A _management_ instruction manipulates cached copies of data with respect to a set of agents that can access the data * A _zero_ instruction zeros out a range of memory locations, potentially allocating cached copies of data in one or more caches * A _prefetch_ instruction indicates to hardware that data at a given memory location may be accessed in the near future, potentially allocating cached copies of data in one or more caches This document introduces a base set of CMO ISA extensions that operate specifically on cache blocks or the memory locations corresponding to a cache block; these are known as _cache-block operation_ (or _CBO_) instructions. Each of the above classes of instructions represents an extension in this specification: * The _Zicbom_ extension defines a set of cache-block management instructions:`CBO.INVAL`, `CBO.CLEAN`, and `CBO.FLUSH` * The _Zicboz_ extension defines a cache-block zero instruction: `CBO.ZERO` * The _Zicbop_ extension defines a set of cache-block prefetch instructions:`PREFETCH.R`, `PREFETCH.W`, and `PREFETCH.I` The execution behavior of the above instructions is also modified by CSR state added by this specification. The remainder of this document provides general background information on CMO instructions and describes each of the above ISA extensions. | | _The term CMO encompasses all operations on caches or resources related to caches. The term CBO represents a subset of CMOs that operate only on cache blocks. The first CMO extensions only define CBOs._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#background)20.1.3\. Background This chapter provides information common to all CMO extensions. #### [](#memory-caches)20.1.3.1\. Memory and Caches A _memory location_ is a physical resource in a system uniquely identified by a_physical address_. An _agent_ is a logic block, such as a RISC-V hart, accelerator, I/O device, etc., that can access a given memory location. | | _A given agent may not be able to access all memory locations in a system, and two different agents may or may not be able to access the same set of memory locations._ | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | A _load operation_ (or _store operation_) is performed by an agent to consume (or modify) the data at a given memory location. Load and store operations are performed as a result of explicit memory accesses to that memory location. Additionally, a _read transfer_ from memory fetches the data at the memory location, while a _write transfer_ to memory updates the data at the memory location. A _cache_ is a structure that buffers copies of data to reduce average memory latency. Any number of caches may be interspersed between an agent and a memory location, and load and store operations from an agent may be satisfied by a cache instead of the memory location. | | _Load and store operations are decoupled from read and write transfers by caches. For example, a load operation may be satisfied by a cache without performing a read transfer from memory, or a store operation may be satisfied by a cache that first performs a read transfer from memory._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Caches organize copies of data into _cache blocks_, each of which represents a contiguous, naturally aligned power-of-two (or _NAPOT_) range of memory locations. A cache block is identified by any of the physical addresses corresponding to the underlying memory locations. The capacity and organization of a cache and the size of a cache block are both _implementation-specific_, and the execution environment provides software a means to discover information about the caches and cache blocks in a system. In the initial set of CMO extensions, the size of a cache block shall be uniform throughout the system. | | _In future CMO extensions, the requirement for a uniform cache block size may be relaxed._ | | --------------------------------------------------------------------------------------------- | Implementation techniques such as speculative execution or hardware prefetching may cause a given cache to allocate or deallocate a copy of a cache block at any time, provided the corresponding physical addresses are accessible according to the supported access type PMA and are cacheable according to the cacheability PMA. Allocating a copy of a cache block results in a read transfer from another cache or from memory, while deallocating a copy of a cache block may result in a write transfer to another cache or to memory depending on whether the data in the copy were modified by a store operation. Additional details are discussed in[Coherent Agents and Caches](#coherent-agents-caches). #### [](#20-1-3-2-cache-block-operations)20.1.3.2\. Cache-Block Operations A CBO instruction causes one or more operations to be performed on the cache blocks identified by the instruction. In general, a CBO instruction may identify one or more cache blocks; however, in the initial set of CMO extensions, CBO instructions identify a single cache block only. A cache-block management instruction performs one of the following operations, relative to the copy of a given cache block allocated in a given cache: * An _invalidate operation_ deallocates the copy of the cache block * A _clean operation_ performs a write transfer to another cache or to memory if the data in the copy of the cache block have been modified by a store operation * A _flush operation_ atomically performs a clean operation followed by an invalidate operation Additional details, including the actual operation performed by a given cache-block management instruction, are described in [Cache-Block Management Instructions](#Zicbom). A cache-block zero instruction performs a set of store operations that write zeros to the set of bytes corresponding to a cache block. Unless specified otherwise, the store operations generated by a cache-block zero instruction have the same general properties and behaviors that other store instructions in the architecture have. An implementation may or may not update the entire set of bytes atomically with a single store operation. Additional details are described in [Cache-Block Zero Instructions](#Zicboz). A cache-block prefetch instruction is a HINT to the hardware that software expects to perform a particular type of memory access in the near future. Additional details are described in [Cache-Block Prefetch Instructions](#Zicbop). ### [](#coherent-agents-caches)20.1.4\. Coherent Agents and Caches For a given memory location, a _set of coherent agents_ consists of the agents for which all of the following hold: * Store operations from all agents in the set appear to be serialized with respect to each other * Store operations from all agents in the set eventually appear to all other agents in the set * A load operation from an agent in the set returns data from a store operation from an agent in the set (or from the initial data in memory) The coherent agents within such a set shall access a given memory location with the same physical address and the same physical memory attributes; however, if the coherence PMA for a given agent indicates a given memory location is not coherent, that agent shall not be a member of a set of coherent agents with any other agent for that memory location and shall be the sole member of a set of coherent agents consisting of itself. An agent who is a member of a set of coherent agents is said to be _coherent_with respect to the other agents in the set. On the other hand, an agent who is_not_ a member is said to be _non-coherent_ with respect to the agents in the set. Caches introduce the possibility that multiple copies of a given cache block may be present in a system at the same time. An _implementation-specific_ mechanism keeps these copies coherent with respect to the load and store operations from the agents in the set of coherent agents. Additionally, if a coherent agent in the set executes a CBO instruction that specifies the cache block, the resulting operation shall apply to any and all of the copies in the caches that can be accessed by the load and store operations from the coherent agents. | | _An operation from a CBO instruction is defined to operate only on the copies of a cache block that are cached in the caches accessible by the explicit memory accesses performed by the set of coherent agents. This includes copies of a cache block in caches that are accessed only indirectly by load and store operations, e.g. coherent instruction caches._ | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The set of caches subject to the above mechanism form a _set of coherent caches_, and each coherent cache has the following behaviors, assuming all operations are performed by the agents in a set of coherent agents: * A coherent cache is permitted to allocate and deallocate copies of a cache block and perform read and write transfers as described in [Memory and Caches](#memory-caches) * A coherent cache is permitted to perform a write transfer to memory provided that a store operation has modified the data in the cache block since the most recent invalidate, clean, or flush operation on the cache block * At least one coherent cache is responsible for performing a write transfer to memory once a store operation has modified the data in the cache block until the next invalidate, clean, or flush operation on the cache block, after which no coherent cache is responsible (or permitted) to perform a write transfer to memory until the next store operation has modified the data in the cache block * A coherent cache is required to perform a write transfer to memory if a store operation has modified the data in the cache block since the most recent invalidate, clean, or flush operation on the cache block and if the next clean or flush operation requires a write transfer to memory | | _The above restrictions ensure that a "clean" copy of a cache block, fetched by a read transfer from memory and unmodified by a store operation, cannot later overwrite the copy of the cache block in memory updated by a write transfer to memory from a non-coherent agent._ | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | A non-coherent agent may initiate a cache-block operation that operates on the set of coherent caches accessed by a set of coherent agents. The mechanism to perform such an operation is _implementation-specific_. #### [](#20-1-4-1-memory-ordering)20.1.4.1\. Memory Ordering ##### [](#20-1-4-1-1-preserved-program-order)20.1.4.1.1\. Preserved Program Order The preserved program order (abbreviated _PPO_) rules are defined by the RVWMO memory ordering model. How the operations resulting from CMO instructions fit into these rules is described below. For cache-block management instructions, the resulting invalidate, clean, and flush operations behave as stores in the PPO rules subject to one additional overlapping address rule. Specifically, if _a_ precedes _b_ in program order, then _a_ will precede _b_ in the global memory order if: * _a_ is an invalidate, clean, or flush, _b_ is a load, and _a_ and _b_ access overlapping memory addresses | | _The above rule ensures that a subsequent load in program order never appears in the global memory order before a preceding invalidate, clean, or flush operation to an overlapping address._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Additionally, invalidate, clean, and flush operations are classified as W or O (depending on the physical memory attributes for the corresponding physical addresses) for the purposes of predecessor and successor sets in `FENCE`instructions. These operations are _not_ ordered by other instructions that order stores, e.g. `FENCE.I` and `SFENCE.VMA`. For cache-block zero instructions, the resulting store operations behave as stores in the PPO rules and are ordered by other instructions that order stores. Finally, for cache-block prefetch instructions, the resulting operations are_not_ ordered by the PPO rules nor are they ordered by any other ordering instructions. ##### [](#20-1-4-1-2-load-values)20.1.4.1.2\. Load Values An invalidate operation may change the set of values that can be returned by a load. In particular, an additional condition is added to the Load Value Axiom: * If an invalidate operation _i_ precedes a load _r_ and operates on a byte _x_returned by _r_, and no store to _x_ appears between _i_ and _r_ in program order or in the global memory order, then _r_ returns any of the following values for _x_: 1. If no clean or flush operations on _x_ precede _i_ in the global memory order, either the initial value of _x_ or the value of any store to _x_ that precedes_i_ 2. If no store to _x_ precedes a clean or flush operation on _x_ in the global memory order and if the clean or flush operation on _x_ precedes _i_ in the global memory order, either the initial value of _x_ or the value of any store to _x_ that precedes _i_ 3. If a store to _x_ precedes a clean or flush operation on _x_ in the global memory order and if the clean or flush operation on _x_ precedes _i_ in the global memory order, either the value of the latest store to _x_ that precedes the latest clean or flush operation on _x_ or the value of any store to _x_that both precedes _i_ and succeeds the latest clean or flush operation on _x_that precedes _i_ 4. The value of any store to _x_ by a non-coherent agent regardless of the above conditions | | _The first three bullets describe the possible load values at different points in the global memory order relative to clean or flush operations. The final bullet implies that the load value may be produced by a non-coherent agent at any time._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#20-1-4-2-traps)20.1.4.2\. Traps Execution of certain CMO instructions may result in traps due to CSR state, described in the [Control and Status Register State](#csr%5Fstate) section, or due to the address translation and protection mechanisms. The trapping behavior of CMO instructions is described in the following sections. ##### [](#20-1-4-2-1-illegal-instruction-and-virtual-instruction-exceptions)20.1.4.2.1\. Illegal-Instruction and Virtual-Instruction Exceptions Cache-block management instructions and cache-block zero instructions may raise illegal-instruction exceptions or virtual-instruction exceptions depending on the current privilege mode and the state of the CMO control registers described in the [Control and Status Register State](#csr%5Fstate) section. Cache-block prefetch instructions raise neither illegal-instruction exceptions nor virtual-instruction exceptions. ##### [](#20-1-4-2-2-page-fault-guest-page-fault-and-access-fault-exceptions)20.1.4.2.2\. Page-Fault, Guest-Page-Fault, and Access-Fault Exceptions Similar to load and store instructions, CMO instructions are explicit memory access instructions that compute an effective address. The effective address is ultimately translated into a physical address based on the privilege mode and the enabled translation mechanisms, and the CMO extensions impose the following constraints on the physical addresses in a given cache block: * The PMP access control bits shall be the same for _all_ physical addresses in the cache block, and if write permission is granted by the PMP access control bits, read permission shall also be granted * The PMAs shall be the same for _all_ physical addresses in the cache block, and if write permission is granted by the supported access type PMAs, read permission shall also be granted If the above constraints are not met, the behavior of a CBO instruction is UNSPECIFIED. | | _This specification assumes that the above constraints will typically be met for main memory regions and may be met for certain I/O regions._ | | ------------------------------------------------------------------------------------------------------------------------------------------------ | | | The access size for CMO instructions is equal to the size of the cache block, however in some cases that access can be decomposed into multiple memory operations. PMP checks are applied to each memory operation independently. For example a 64-byte **cbo.zero** that spans two 32-byte PMP regions would succeed if it was decomposed into two 32-byte memory operations (and the PMP access control bits are the same in both regions), but if performed as a single 64-byte memory operation it would cause an access fault. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The Zicboz extension introduces an additional supported access type PMA for cache-block zero instructions. Main memory regions are required to support accesses by cache-block zero instructions; however, I/O regions may specify whether accesses by cache-block zero instructions are supported. A cache-block management instruction is permitted to access the specified cache block whenever a load instruction or store instruction is permitted to access the corresponding physical addresses. If neither a load instruction nor store instruction is permitted to access the physical addresses, but an instruction fetch is permitted to access the physical addresses, whether a cache-block management instruction is permitted to access the cache block is UNSPECIFIED. If access to the cache block is not permitted, a cache-block management instruction raises a store page-fault or store guest-page-fault exception if address translation does not permit any access or raises a store access-fault exception otherwise. During address translation, the instruction also checks the accessed bit and may either raise an exception or set the bit as required. | | _The interaction between cache-block management instructions and instruction fetches will be specified in a future extension._ _As implied by omission, a cache-block management instruction does not check the dirty bit and neither raises an exception nor sets the bit._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | A cache-block zero instruction is permitted to access the specified cache block whenever a store instruction is permitted to access the corresponding physical addresses and when the PMAs indicate that cache-block zero instructions are a supported access type. If access to the cache block is not permitted, a cache-block zero instruction raises a store page-fault or store guest-page-fault exception if address translation does not permit write access or raises a store access-fault exception otherwise. During address translation, the instruction also checks the accessed and dirty bits and may either raise an exception or set the bits as required. A cache-block prefetch instruction is permitted to access the specified cache block whenever a load instruction, store instruction, or instruction fetch is permitted to access the corresponding physical addresses. If access to the cache block is not permitted, a cache-block prefetch instruction does not raise any exceptions and shall not access any caches or memory. During address translation, the instruction does _not_ check the accessed and dirty bits and neither raises an exception nor sets the bits. When a page-fault, guest-page-fault, or access-fault exception is taken, the relevant \*tval CSR is written with the faulting effective address (i.e. the value of _rs1_). | | _Like a load or store instruction, a CMO instruction may or may not be permitted to access a cache block based on the states of the MPRV, MPV, and MPP bits in mstatus and the SUM and MXR bits in mstatus, sstatus, andvsstatus._ _This specification expects that implementations will process cache-block management instructions like store/AMO instructions, so store/AMO exceptions are appropriate for these instructions, regardless of the permissions required._ | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#20-1-4-2-3-address-misaligned-exceptions)20.1.4.2.3\. Address-Misaligned Exceptions CMO instructions do _not_ generate address-misaligned exceptions. ##### [](#20-1-4-2-4-breakpoint-exceptions-and-debug-mode-entry)20.1.4.2.4\. Breakpoint Exceptions and Debug Mode Entry Unless otherwise defined by the debug architecture specification, the behavior of trigger modules with respect to CMO instructions is UNSPECIFIED. | | _For the Zicbom, Zicboz, and Zicbop extensions, this specification recommends the following common trigger module behaviors:_ Type 6 address match triggers, i.e. tdata1.type=6 and mcontrol6.select=0, should be supported Type 2 address/data match triggers, i.e. tdata1.type=2, should be unsupported The size of a memory access equals the size of the cache block accessed, and the compare values follow from the addresses of the NAPOT memory region corresponding to the cache block containing the effective address Unless an encoding for a cache block is added to the mcontrol6.size field, an address trigger should only match a memory access from a CBO instruction ifmcontrol6.size=0 _If the Zicbom extension is implemented, this specification recommends the following additional trigger module behaviors:_ Implementing address match triggers should be optional Type 6 data match triggers, i.e. tdata1.type=6 and mcontrol6.select=1, should be unsupported Memory accesses are considered to be stores, i.e. an address trigger matches only if mcontrol6.store=1 _If the Zicboz extension is implemented, this specification recommends the following additional trigger module behaviors:_ Implementing address match triggers should be mandatory Type 6 data match triggers, i.e. tdata1.type=6 and mcontrol6.select=1, should be supported, and implementing these triggers should be optional Memory accesses are considered to be stores, i.e. an address trigger matches only if mcontrol6.store=1 _If the Zicbop extension is implemented, this specification recommends the following additional trigger module behaviors:_ Implementing address match triggers should be optional Type 6 data match triggers, i.e. tdata1.type=6 and mcontrol6.select=1, should be unsupported Memory accesses may be considered to be loads or stores depending on the implementation, i.e. whether an address trigger matches on these instructions when mcontrol6.load=1 or mcontrol6.store=1 is _implementation-specific_ _This specification also recommends that the behavior of trigger modules with respect to the Zicboz extension should be defined in version 1.0 of the debug architecture specification. The behavior of trigger modules with respect to the Zicbom and Zicbop extensions is expected to be defined in future extensions._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ##### [](#20-1-4-2-5-hypervisor-extension)20.1.4.2.5\. Hypervisor Extension For the purposes of writing the `mtinst` or `htinst` register on a trap, the following standard transformation is defined for cache-block management instructions and cache-block zero instructions: ![svg](_images/svg-9a35c968285eb5e2a3fced1bf550dda737fb5e3b.svg) The `operation` field corresponds to the 12 most significant bits of the trapping instruction. | | _As described in the hypervisor extension, a zero may be written into mtinstor htinst instead of the standard transformation defined above._ | | ----------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#20-1-4-3-effects-on-constrained-lrsc-loops)20.1.4.3\. Effects on Constrained LR/SC Loops The following event is added to the list of events that satisfy the eventuality guarantee provided by constrained LR/SC loops, as defined in the A extension: * Some other hart executes a cache-block management instruction or a cache-block zero instruction to the reservation set of the LR instruction in _H_'s constrained LR/SC loop. | | _The above event has been added to accommodate cache coherence protocols that cannot distinguish between invalidations for stores and invalidations for cache-block management operations._ _Aside from the above event, CMO instructions neither change the properties of constrained LR/SC loops nor modify the eventuality guarantee provided by them. For example, executing a CMO instruction may cause a constrained LR/SC loop on any hart to fail periodically or may cause a unconstrained LR/SC sequence on the same hart to fail always. Additionally, executing a cache-block prefetch instruction does not impact the eventuality guarantee provided by constrained LR/SC loops executed on any hart._ | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#20-1-4-4-software-discovery)20.1.4.4\. Software Discovery The initial set of CMO extensions requires the following information to be discovered by software: * The size of the cache block for management and prefetch instructions * The size of the cache block for zero instructions * CBIE support at each privilege level Other general cache characteristics may also be specified in the discovery mechanism. ### [](#csr%5Fstate)20.1.5\. CSR controls for CMO instructions The xenvcfg registers control CBO instruction execution based on the current privilege mode and the state of the appropriate CSRs, as detailed below. A `CBO.INVAL` instruction executes or raises either an illegal-instruction exception or a virtual-instruction exception based on the state of the`xenvcfg.CBIE` fields: ```sail // illegal-instruction exceptions if (((priv_mode != M) && (menvcfg.CBIE == 00)) || ((priv_mode == U) && (senvcfg.CBIE == 00))) { } // virtual-instruction exceptions else if (((priv_mode == VS) && (henvcfg.CBIE == 00)) || ((priv_mode == VU) && ((henvcfg.CBIE == 00) || (senvcfg.CBIE == 00)))) { } // execute instruction else { if (((priv_mode != M) && (menvcfg.CBIE == 01)) || ((priv_mode == U) && (senvcfg.CBIE == 01)) || ((priv_mode == VS) && (henvcfg.CBIE == 01)) || ((priv_mode == VU) && ((henvcfg.CBIE == 01) || (senvcfg.CBIE == 01)))) { } else { } } ``` | | _Until a modified cache block has updated memory, a CBO.INVAL instruction may expose stale data values in memory if the CSRs are programmed to perform an invalidate operation. This behavior may result in a security hole if lower privileged level software performs an invalidate operation and accesses sensitive information in memory._ _To avoid such holes, higher privileged level software must perform either a clean or flush operation on the cache block before permitting lower privileged level software to perform an invalidate operation on the block. Alternatively, higher privileged level software may program the CSRs so that CBO.INVALeither traps or performs a flush operation in a lower privileged level._ | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | A `CBO.CLEAN` or `CBO.FLUSH` instruction executes or raises an illegal-instruction or virtual-instruction exception based on the state of the`xenvcfg.CBCFE` bits: ```sail // illegal-instruction exceptions if (((priv_mode != M) && !menvcfg.CBCFE) || ((priv_mode == U) && !senvcfg.CBCFE)) { } // virtual-instruction exceptions else if (((priv_mode == VS) && !henvcfg.CBCFE) || ((priv_mode == VU) && !(henvcfg.CBCFE && senvcfg.CBCFE))) { } // execute instruction else { } ``` Finally, a `CBO.ZERO` instruction executes or raises an illegal-instruction or virtual-instruction exception based on the state of the `xenvcfg.CBZE` bits: ```sail // illegal-instruction exceptions if (((priv_mode != M) && !menvcfg.CBZE) || ((priv_mode == U) && !senvcfg.CBZE)) { } // virtual-instruction exceptions else if (((priv_mode == VS) && !henvcfg.CBZE) || ((priv_mode == VU) && !(henvcfg.CBZE && senvcfg.CBZE))) { } // execute instruction else { } ``` The CBIE/CBCFE/CBZE fields in each `xenvcfg` register do not affect the read and write behavior of the same fields in the other `xenvcfg` registers. Each `xenvcfg` register is WARL; however, software should determine the legal values from the execution environment discovery mechanism. ### [](#extensions)20.1.6\. Extensions CMO instructions are defined in the following extensions: * [Cache-Block Management Instructions](#Zicbom) * [Cache-Block Zero Instructions](#Zicboz) * [Cache-Block Prefetch Instructions](#Zicbop) #### [](#Zicbom)20.1.6.1\. Cache-Block Management Instructions Cache-block management instructions enable software running on a set of coherent agents to communicate with a set of non-coherent agents by performing one of the following operations: * An invalidate operation makes data from store operations performed by a set of non-coherent agents visible to the set of coherent agents at a point common to both sets by deallocating all copies of a cache block from the set of coherent caches up to that point * A clean operation makes data from store operations performed by the set of coherent agents visible to a set of non-coherent agents at a point common to both sets by performing a write transfer of a copy of a cache block to that point provided a coherent agent performed a store operation that modified the data in the cache block since the previous invalidate, clean, or flush operation on the cache block * A flush operation atomically performs a clean operation followed by an invalidate operation In the Zicbom extension, the instructions operate to a point common to _all_agents in the system. In other words, an invalidate operation ensures that store operations from all non-coherent agents visible to agents in the set of coherent agents, and a clean operation ensures that store operations from coherent agents visible to all non-coherent agents. | | _The Zicbom extension does not prohibit agents that fall outside of the above architectural definition; however, software cannot rely on the defined cache operations to have the desired effects with respect to those agents._ _Future extensions may define different sets of agents for the purposes of performance optimization._ | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | These instructions operate on the cache block whose effective address is specified in _rs1_. The effective address is translated into a corresponding physical address by the appropriate translation mechanisms. The following instructions comprise the Zicbom extension: | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | ---------------- | -------------------------------------------- | | ✓ | ✓ | cbo.clean _base_ | [Cache Block Clean](#insns-cbo%5Fclean) | | ✓ | ✓ | cbo.flush _base_ | [Cache Block Flush](#insns-cbo%5Fflush) | | ✓ | ✓ | cbo.inval _base_ | [Cache Block Invalidate](#insns-cbo%5Finval) | | | _Cache-block management instructions ignore cacheability attributes and operate on the cache block irrespective of the PMA cacheable attribute and any Page-Based Memory Type (PBMT) downgrade from cacheable to non-cacheable._ | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#Zicboz)20.1.6.2\. Cache-Block Zero Instructions Cache-block zero instructions store zeros to the set of bytes corresponding to a cache block. An implementation may update the bytes in any order and with any granularity and atomicity, including individual bytes. | | _Cache-block zero instructions store zeros independently of whether data from the underlying memory locations are cacheable. In addition, this specification does not constrain how the bytes are written._ | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | These instructions operate on the cache block, or the memory locations corresponding to the cache block, whose effective address is specified in _rs1_.The effective address is translated into a corresponding physical address by the appropriate translation mechanisms. The following instructions comprise the Zicboz extension: | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | --------------- | ------------------------------------- | | ✓ | ✓ | cbo.zero _base_ | [Cache Block Zero](#insns-cbo%5Fzero) | #### [](#Zicbop)20.1.6.3\. Cache-Block Prefetch Instructions Cache-block prefetch instructions are HINTs to the hardware to indicate that software intends to perform a particular type of memory access in the near future. The types of memory accesses are instruction fetch, data read (i.e. load), and data write (i.e. store). These instructions operate on the cache block whose effective address is the sum of the base address specified in _rs1_ and the sign-extended offset encoded in_imm\[11:0\]_, where _imm\[4:0\]_ shall equal `0b00000`. The effective address is translated into a corresponding physical address by the appropriate translation mechanisms. | | _Cache-block prefetch instructions are encoded as ORI instructions with rd equal to 0b00000; however, for the purposes of effective address calculation, this field is also interpreted as imm\[4:0\] like a store instruction._ | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The following instructions comprise the Zicbop extension: | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | --------------------------- | ----------------------------------------------------------------- | | ✓ | ✓ | prefetch.i _offset_(_base_) | [Cache Block Prefetch for Instruction Fetch](#insns-prefetch%5Fi) | | ✓ | ✓ | prefetch.r _offset_(_base_) | [Cache Block Prefetch for Data Read](#insns-prefetch%5Fr) | | ✓ | ✓ | prefetch.w _offset_(_base_) | [Cache Block Prefetch for Data Write](#insns-prefetch%5Fw) | ### [](#insns)20.1.7\. Instructions #### [](#insns-cbo%5Fclean)20.1.7.1\. cbo.clean Synopsis Perform a clean operation on a cache block Mnemonic cbo.clean _offset_(_base_) Encoding ![svg](_images/svg-d4b47caa273c632ba3474227c4ddc1af5d65f705.svg) Description A **cbo.clean** instruction performs a clean operation on the cache block whose effective address is the base address specified in _rs1_. The offset operand may be omitted; otherwise, any expression that computes the offset shall evaluate to zero. The instruction operates on the set of coherent caches accessed by the agent executing the instruction. | | _When executing a **cbo.clean** instruction, an implementation may instead perform a flush operation, since the result of that operation is indistinguishable from the sequence of performing a clean operation just before deallocating all cached copies in the set of coherent caches._ | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#insns-cbo%5Fflush)20.1.7.2\. cbo.flush Synopsis Perform a flush operation on a cache block Mnemonic cbo.flush _offset_(_base_) Encoding ![svg](_images/svg-485e71f327ed07c1c420ce0688726954575611fb.svg) Description A **cbo.flush** instruction performs a flush operation on the cache block whose that contains the address specified in _rs1_. It is not required that _rs1_ is aligned to the size of a cache block. On faults, the faulting virtual address is considered to be the value in rs1, rather than the base address of the cache block. The instruction operates on the set of coherent caches accessed by the agent executing the instruction. The assembly _offset_ operand may be omitted. If it isn’t then any expression that computes the offset shall evaluate to zero. #### [](#insns-cbo%5Finval)20.1.7.3\. cbo.inval Synopsis Perform an invalidate operation on a cache block Mnemonic cbo.inval _offset_(_base_) Encoding ![svg](_images/svg-b8633422958e63603dab774850cf5ae792e62af4.svg) Description A **cbo.inval** instruction performs an invalidate operation on the cache block that contains the address specified in _rs1_. It is not required that _rs1_ is aligned to the size of a cache block. On faults, the faulting virtual address is considered to be the value in rs1, rather than the base address of the cache block. The instruction operates on the set of coherent caches accessed by the agent executing the instruction. Depending on CSR programming, the instruction may perform a flush operation instead of an invalidate operation. The assembly _offset_ operand may be omitted. If it isn’t then any expression that computes the offset shall evaluate to zero. | | _When executing a **cbo.inval** instruction, an implementation may instead perform a flush operation, since the result of that operation is indistinguishable from the sequence of performing a write transfer to memory just before performing an invalidate operation._ | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#insns-cbo%5Fzero)20.1.7.4\. cbo.zero Synopsis Store zeros to the full set of bytes corresponding to a cache block Mnemonic cbo.zero _offset_(_base_) Encoding ![svg](_images/svg-b53d4cf7b77afa80e944d62932e7021e7f1715dd.svg) Description A **cbo.zero** instruction performs stores of zeros to the full set of bytes corresponding to the cache block that contains the address specified in _rs1_. It is not required that _rs1_ is aligned to the size of a cache block. On faults, the faulting virtual address is considered to be the value in rs1, rather than the base address of the cache block. An implementation may or may not update the entire set of bytes atomically. The assembly _offset_ operand may be omitted. If it isn’t then any expression that computes the offset shall evaluate to zero. #### [](#insns-prefetch%5Fi)20.1.7.5\. prefetch.i Synopsis Provide a HINT to hardware that a cache block is likely to be accessed by an instruction fetch in the near future Mnemonic prefetch.i _offset_(_base_) Encoding ![svg](_images/svg-f9d95edf326e4e202e1dba2e64f1d72c10f0757d.svg) Description A **prefetch.i** instruction indicates to hardware that the cache block whose effective address is the sum of the base address specified in _rs1_ and the sign-extended offset encoded in _imm\[11:0\]_, where _imm\[4:0\]_ equals `0b00000`, is likely to be accessed by an instruction fetch in the near future. | | _An implementation may opt to cache a copy of the cache block in a cache accessed by an instruction fetch in order to improve memory access latency, but this behavior is not required._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#insns-prefetch%5Fr)20.1.7.6\. prefetch.r Synopsis Provide a HINT to hardware that a cache block is likely to be accessed by a data read in the near future Mnemonic prefetch.r _offset_(_base_) Encoding ![svg](_images/svg-04c35b2b1b271ab717acd55eaa5615b33982f1a7.svg) Description A **prefetch.r** instruction indicates to hardware that the cache block whose effective address is the sum of the base address specified in _rs1_ and the sign-extended offset encoded in _imm\[11:0\]_, where _imm\[4:0\]_ equals `0b00000`, is likely to be accessed by a data read (i.e. load) in the near future. | | _An implementation may opt to cache a copy of the cache block in a cache accessed by a data read in order to improve memory access latency, but this behavior is not required._ | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#insns-prefetch%5Fw)20.1.7.7\. prefetch.w Synopsis Provide a HINT to hardware that a cache block is likely to be accessed by a data write in the near future Mnemonic prefetch.w _offset_(_base_) Encoding ![svg](_images/svg-802d1461e50bc0c46e1d341f048b6aa81bcce089.svg) Description A **prefetch.w** instruction indicates to hardware that the cache block whose effective address is the sum of the base address specified in _rs1_ and the sign-extended offset encoded in _imm\[11:0\]_, where _imm\[4:0\]_ equals `0b00000`, is likely to be accessed by a data write (i.e. store) in the near future. | | _An implementation may opt to cache a copy of the cache block in a cache accessed by a data write in order to improve memory access latency, but this behavior is not required._ | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Preface ==================== ## [](#preface)Preface This document describes the RISC-V unprivileged architecture. It contains the following versions of the RISC-V ISA modules, all of which have been ratified: | Base | Version | Status | | --------------- | -------- | ------------ | | **RV32I** | **2.1** | **Ratified** | | **RV32E** | **2.0** | **Ratified** | | **RV64E** | **2.0** | **Ratified** | | **RV64I** | **2.1** | **Ratified** | | Extension | Version | Status | | **Zifencei** | **2.0** | **Ratified** | | **Zicsr** | **2.0** | **Ratified** | | **Zicntr** | **2.0** | **Ratified** | | **Zihintntl** | **1.0** | **Ratified** | | **Zihintpause** | **2.0** | **Ratified** | | **Zimop** | **1.0** | **Ratified** | | **Zicond** | **1.0** | **Ratified** | | **Zilsd** | **1.0** | **Ratified** | | **M** | **2.0** | **Ratified** | | **Zmmul** | **1.0** | **Ratified** | | **A** | **2.1** | **Ratified** | | **Zawrs** | **1.01** | **Ratified** | | **Zacas** | **1.0** | **Ratified** | | **Zabha** | **1.0** | **Ratified** | | **Zalasr** | **1.0** | **Ratified** | | **RVWMO** | **2.0** | **Ratified** | | **Ztso** | **1.0** | **Ratified** | | **CMO** | **1.0** | **Ratified** | | **F** | **2.2** | **Ratified** | | **D** | **2.2** | **Ratified** | | **Q** | **2.2** | **Ratified** | | **Zfh** | **1.0** | **Ratified** | | **Zfhmin** | **1.0** | **Ratified** | | **BF16** | **1.0** | **Ratified** | | **Zfa** | **1.0** | **Ratified** | | **Zfinx** | **1.0** | **Ratified** | | **Zdinx** | **1.0** | **Ratified** | | **Zhinx** | **1.0** | **Ratified** | | **Zhinxmin** | **1.0** | **Ratified** | | **C** | **2.0** | **Ratified** | | **Zce** | **1.0** | **Ratified** | | **Zclsd** | **1.0** | **Ratified** | | **B** | **1.0** | **Ratified** | | **V** | **1.0** | **Ratified** | | **Zbkb** | **1.0** | **Ratified** | | **Zbkc** | **1.0** | **Ratified** | | **Zbkx** | **1.0** | **Ratified** | | **Zk** | **1.0** | **Ratified** | | **Zks** | **1.0** | **Ratified** | | **Zvbb** | **1.0** | **Ratified** | | **Zvbc** | **1.0** | **Ratified** | | **Zvkg** | **1.0** | **Ratified** | | **Zvkned** | **1.0** | **Ratified** | | **Zvknhb** | **1.0** | **Ratified** | | **Zvksed** | **1.0** | **Ratified** | | **Zvksh** | **1.0** | **Ratified** | | **Zvkt** | **1.0** | **Ratified** | | **Zicfiss** | **1.0** | **Ratified** | | **Zicfilp** | **1.0** | **Ratified** | The changes in this version of the document include: * Addition of the Zalasr extension for Load-Acquire/Store-Release operations. **_Preface to Document Version 20250508_** This document describes the RISC-V unprivileged architecture. It contains the following versions of the RISC-V ISA modules, all of which have been ratified: | Base | Version | Status | | --------------- | -------- | ------------ | | **RV32I** | **2.1** | **Ratified** | | **RV32E** | **2.0** | **Ratified** | | **RV64E** | **2.0** | **Ratified** | | **RV64I** | **2.1** | **Ratified** | | Extension | Version | Status | | **Zifencei** | **2.0** | **Ratified** | | **Zicsr** | **2.0** | **Ratified** | | **Zicntr** | **2.0** | **Ratified** | | **Zihintntl** | **1.0** | **Ratified** | | **Zihintpause** | **2.0** | **Ratified** | | **Zimop** | **1.0** | **Ratified** | | **Zicond** | **1.0** | **Ratified** | | **Zilsd** | **1.0** | **Ratified** | | **M** | **2.0** | **Ratified** | | **Zmmul** | **1.0** | **Ratified** | | **A** | **2.1** | **Ratified** | | **Zawrs** | **1.01** | **Ratified** | | **Zacas** | **1.0** | **Ratified** | | **Zabha** | **1.0** | **Ratified** | | **RVWMO** | **2.0** | **Ratified** | | **Ztso** | **1.0** | **Ratified** | | **CMO** | **1.0** | **Ratified** | | **F** | **2.2** | **Ratified** | | **D** | **2.2** | **Ratified** | | **Q** | **2.2** | **Ratified** | | **Zfh** | **1.0** | **Ratified** | | **Zfhmin** | **1.0** | **Ratified** | | **BF16** | **1.0** | **Ratified** | | **Zfa** | **1.0** | **Ratified** | | **Zfinx** | **1.0** | **Ratified** | | **Zdinx** | **1.0** | **Ratified** | | **Zhinx** | **1.0** | **Ratified** | | **Zhinxmin** | **1.0** | **Ratified** | | **C** | **2.0** | **Ratified** | | **Zce** | **1.0** | **Ratified** | | **Zclsd** | **1.0** | **Ratified** | | **B** | **1.0** | **Ratified** | | **V** | **1.0** | **Ratified** | | **Zbkb** | **1.0** | **Ratified** | | **Zbkc** | **1.0** | **Ratified** | | **Zbkx** | **1.0** | **Ratified** | | **Zk** | **1.0** | **Ratified** | | **Zks** | **1.0** | **Ratified** | | **Zvbb** | **1.0** | **Ratified** | | **Zvbc** | **1.0** | **Ratified** | | **Zvkg** | **1.0** | **Ratified** | | **Zvkned** | **1.0** | **Ratified** | | **Zvknhb** | **1.0** | **Ratified** | | **Zvksed** | **1.0** | **Ratified** | | **Zvksh** | **1.0** | **Ratified** | | **Zvkt** | **1.0** | **Ratified** | | **Zicfiss** | **1.0** | **Ratified** | | **Zicfilp** | **1.0** | **Ratified** | The changes in this version of the document include: * The inclusion of all ratified extensions through May 2025. * Removal of all unratified material. * Addition of the BFloat16-precision Floating Point extension. * Addition of the Zabha extension for Byte and Halfword Atomic Memory Operations. **_Preface to Document Version 20240411_** This document describes the RISC-V unprivileged architecture. It contains the following versions of the RISC-V ISA modules: | Base | Version | Status | | --------------- | -------- | ------------ | | **RV32I** | **2.1** | **Ratified** | | **RV32E** | **2.0** | **Ratified** | | **RV64E** | **2.0** | **Ratified** | | **RV64I** | **2.1** | **Ratified** | | Extension | Version | Status | | **Zifencei** | **2.0** | **Ratified** | | **Zicsr** | **2.0** | **Ratified** | | **Zicntr** | **2.0** | **Ratified** | | **Zihintntl** | **1.0** | **Ratified** | | **Zihintpause** | **2.0** | **Ratified** | | **Zimop** | **1.0** | **Ratified** | | **Zicond** | **1.0** | **Ratified** | | **Zilsd** | **1.0** | **Ratified** | | **M** | **2.0** | **Ratified** | | **Zmmul** | **1.0** | **Ratified** | | **A** | **2.1** | **Ratified** | | **Zawrs** | **1.01** | **Ratified** | | **Zacas** | **1.0** | **Ratified** | | **Zabha** | **1.0** | **Ratified** | | **RVWMO** | **2.0** | **Ratified** | | **Ztso** | **1.0** | **Ratified** | | **CMO** | **1.0** | **Ratified** | | **F** | **2.2** | **Ratified** | | **D** | **2.2** | **Ratified** | | **Q** | **2.2** | **Ratified** | | **Zfh** | **1.0** | **Ratified** | | **Zfhmin** | **1.0** | **Ratified** | | **Zfa** | **1.0** | **Ratified** | | **Zfinx** | **1.0** | **Ratified** | | **Zdinx** | **1.0** | **Ratified** | | **Zhinx** | **1.0** | **Ratified** | | **Zhinxmin** | **1.0** | **Ratified** | | **C** | **2.0** | **Ratified** | | **Zce** | **1.0** | **Ratified** | | **Zclsd** | **1.0** | **Ratified** | | **B** | **1.0** | **Ratified** | | **V** | **1.0** | **Ratified** | | **Zbkb** | **1.0** | **Ratified** | | **Zbkc** | **1.0** | **Ratified** | | **Zbkx** | **1.0** | **Ratified** | | **Zk** | **1.0** | **Ratified** | | **Zks** | **1.0** | **Ratified** | | **Zvbb** | **1.0** | **Ratified** | | **Zvbc** | **1.0** | **Ratified** | | **Zvkg** | **1.0** | **Ratified** | | **Zvkned** | **1.0** | **Ratified** | | **Zvknhb** | **1.0** | **Ratified** | | **Zvksed** | **1.0** | **Ratified** | | **Zvksh** | **1.0** | **Ratified** | | **Zvkt** | **1.0** | **Ratified** | | **Zicfiss** | **1.0** | **Ratified** | | **Zicfilp** | **1.0** | **Ratified** | The changes in this version of the document include: * The inclusion of all ratified extensions through February 2025. * The draft Zam extension has been removed, in favor of the definition of a misaligned atomicity granule PMA. * The concept of vacant memory regions has been superseded by inaccessible memory or I/O regions. * The removal of unratified content, including the sketch of the RV128I base ISA. **_Preface to Document Version 20191213-Base-Ratified_** This document describes the RISC-V unprivileged architecture. The ISA modules marked **Ratified** have been ratified at this time. The modules marked _Frozen_ are not expected to change significantly before being put up for ratification. The modules marked _Draft_ are expected to change before ratification. The document contains the following versions of the RISC-V ISA modules: | Base | Version | Status | | ------------ | ------- | ------------ | | RVWMO | 2.0 | **Ratified** | | **RV32I** | **2.1** | **Ratified** | | **RV64I** | **2.1** | **Ratified** | | _RV32E_ | _1.9_ | _Draft_ | | _RV128I_ | _1.7_ | _Draft_ | | Extension | Version | Status | | **M** | **2.0** | **Ratified** | | **A** | **2.1** | **Ratified** | | **F** | **2.2** | **Ratified** | | **D** | **2.2** | **Ratified** | | **Q** | **2.2** | **Ratified** | | **C** | **2.0** | **Ratified** | | _Counters_ | _2.0_ | _Draft_ | | _L_ | _0.0_ | _Draft_ | | _B_ | _0.0_ | _Draft_ | | _J_ | _0.0_ | _Draft_ | | _T_ | _0.0_ | _Draft_ | | _P_ | _0.2_ | _Draft_ | | _V_ | _0.7_ | _Draft_ | | **Zicsr** | **2.0** | **Ratified** | | **Zifencei** | **2.0** | **Ratified** | | _Zam_ | _0.1_ | _Draft_ | | _Ztso_ | _0.1_ | _Frozen_ | The changes in this version of the document include: * The A extension, now version 2.1, was ratified by the board in December 2019. * Defined big-endian ISA variant. * Moved N extension for user-mode interrupts into Volume II. * Defined PAUSE hint instruction. **_Preface to Document Version 20190608-Base-Ratified_** This document describes the RISC-V unprivileged architecture. The RVWMO memory model has been ratified at this time. The ISA modules marked **Ratified**, have been ratified at this time. The modules marked_Frozen_ are not expected to change significantly before being put up for ratification. The modules marked _Draft_ are expected to change before ratification. The document contains the following versions of the RISC-V ISA modules: | Base | Version | Status | | ------------ | ------- | ------------ | | RVWMO | 2.0 | **Ratified** | | **RV32I** | **2.1** | **Ratified** | | **RV64I** | **2.1** | **Ratified** | | _RV32E_ | _1.9_ | _Draft_ | | _RV128I_ | _1.7_ | _Draft_ | | Extension | Version | Status | | **Zifencei** | **2.0** | **Ratified** | | **Zicsr** | **2.0** | **Ratified** | | **M** | **2.0** | **Ratified** | | _A_ | _2.0_ | Frozen | | **F** | **2.2** | **Ratified** | | **D** | **2.2** | **Ratified** | | **Q** | **2.2** | **Ratified** | | **C** | **2.0** | **Ratified** | | _Ztso_ | _0.1_ | _Frozen_ | | _Counters_ | _2.0_ | _Draft_ | | _L_ | _0.0_ | _Draft_ | | _B_ | _0.0_ | _Draft_ | | _J_ | _0.0_ | _Draft_ | | _T_ | _0.0_ | _Draft_ | | _P_ | _0.2_ | _Draft_ | | _V_ | _0.7_ | _Draft_ | | _Zam_ | _0.1_ | _Draft_ | The changes in this version of the document include: * Moved description to **Ratified** for the ISA modules ratified by the board in early 2019. * Removed the A extension from ratification. * Changed document version scheme to avoid confusion with versions of the ISA modules. * Incremented the version numbers of the base integer ISA to 2.1, reflecting the presence of the ratified RVWMO memory model and exclusion of FENCE.I, counters, and CSR instructions that were in previous base ISA. * Incremented the version numbers of the F and D extensions to 2.2, reflecting that version 2.1 changed the canonical NaN, and version 2.2 defined the NaN-boxing scheme and changed the definition of the FMIN and FMAX instructions. * Changed name of document to refer to "unprivileged" instructions as part of move to separate ISA specifications from platform profile mandates. * Added clearer and more precise definitions of execution environments, harts, traps, and memory accesses. * Defined instruction-set categories: _standard_, _reserved_, _custom_,_non-standard_, and _non-conforming_. * Removed text implying operation under alternate endianness, as alternate-endianness operation has not yet been defined for RISC-V. * Changed description of misaligned load and store behavior. The specification now allows visible misaligned address traps in execution environment interfaces, rather than just mandating invisible handling of misaligned loads and stores in user mode. Also, now allows access-fault exceptions to be reported for misaligned accesses (including atomics) that should not be emulated. * Moved FENCE.I out of the mandatory base and into a separate extension, with Zifencei ISA name. FENCE.I was removed from the Linux user ABI and is problematic in implementations with large incoherent instruction and data caches. However, it remains the only standard instruction-fetch coherence mechanism. * Removed prohibitions on using RV32E with other extensions. * Removed platform-specific mandates that certain encodings produce illegal-instruction exceptions in RV32E and RV64I chapters. * Counter/timer instructions are now not considered part of the mandatory base ISA, and so CSR instructions were moved into separate chapter and marked as version 2.0, with the unprivileged counters moved into another separate chapter. The counters are not ready for ratification as there are outstanding issues, including counter inaccuracies. * A CSR-access ordering model has been added. * Explicitly defined the 16-bit half-precision floating-point format for floating-point instructions in the 2-bit _fmt field._ * Defined the signed-zero behavior of FMIN._fmt_ and FMAX._fmt_, and changed their behavior on signaling-NaN inputs to conform to the`minimumNumber` and `maximumNumber` operations in the proposed IEEE 754-201x specification. * The memory consistency model, RVWMO, has been defined. * The "Zam" extension, which permits misaligned AMOs and specifies their semantics, has been defined. * The "Ztso" extension, which enforces a stricter memory consistency model than RVWMO, has been defined. * Improvements to the description and commentary. * Defined the term `IALIGN` as shorthand to describe the instruction-address alignment constraint. * Removed text of `P` extension chapter as now superseded by active task group documents. * Removed text of `V` extension chapter as now superseded by separate vector extension draft document. **_Preface to Document Version 2.2_** This is version 2.2 of the document describing the RISC-V user-level architecture. The document contains the following versions of the RISC-V ISA modules: | Base | _Version_ | _Draft Frozen?_ | | --------- | --------- | --------------- | | RV32I | 2.0 | Y | | RV32E | 1.9 | N | | RV64I | 2.0 | Y | | RV128I | 1.7 | N | | Extension | Version | Frozen? | | M | 2.0 | Y | | A | 2.0 | Y | | F | 2.0 | Y | | D | 2.0 | Y | | Q | 2.0 | Y | | L | 0.0 | N | | C | 2.0 | Y | | B | 0.0 | N | | J | 0.0 | N | | T | 0.0 | N | | P | 0.1 | N | | V | 0.7 | N | | N | 1.1 | N | To date, no parts of the standard have been officially ratified by the RISC-V Foundation, but the components labeled "frozen" above are not expected to change during the ratification process beyond resolving ambiguities and holes in the specification. The major changes in this version of the document include: * The previous version of this document was released under a Creative Commons Attribution 4.0 International License by the original authors, and this and future versions of this document will be released under the same license. * Rearranged chapters to put all extensions first in canonical order. * Improvements to the description and commentary. * Modified implicit hinting suggestion on `JALR` to support more efficient macro-op fusion of `LUI/JALR` and `AUIPC/JALR` pairs. * Clarification of constraints on load-reserved/store-conditional sequences. * A new table of control and status register (CSR) mappings. * Clarified purpose and behavior of high-order bits of `fcsr`. * Corrected the description of the `FNMADD`._fmt_ and `FNMSUB`._fmt_instructions, which had suggested the incorrect sign of a zero result. * Instructions `FMV.S.X` and `FMV.X.S` were renamed to `FMV.W.X` and `FMV.X.W`respectively to be more consistent with their semantics, which did not change. The old names will continue to be supported in the tools. * Specified behavior of narrower (64 bits to avoid moving the _rd_ specifier in very long instruction formats. * CSR instructions are now described in the base integer format where the counter registers are introduced, as opposed to only being introduced later in the floating-point section (and the companion privileged architecture manual). * The SCALL and SBREAK instructions have been renamed to `ECALL` and`EBREAK`, respectively. Their encoding and functionality are unchanged. * Clarification of floating-point NaN handling, and a new canonical NaN value. * Clarification of values returned by floating-point to integer conversions that overflow. * Clarification of `LR/SC` allowed successes and required failures, including use of compressed instructions in the sequence. * A new `RV32E` base ISA proposal for reduced integer register counts, supports `MAC` extensions. * A revised calling convention. * Relaxed stack alignment for soft-float calling convention, and description of the RV32E calling convention. * A revised proposal for the `C` compressed extension, version 1.9 . **_Preface to Version 2.0_** This is the second release of the user ISA specification, and we intend the specification of the base user ISA plus general extensions (i.e., IMAFD) to remain fixed for future development. The following changes have been made since Version 1.0 \[[3](../biblio/bibliography.html#bib-riscvtr)\] of this ISA specification. * The ISA has been divided into an integer base with several standard extensions. * The instruction formats have been rearranged to make immediate encoding more efficient. * The base ISA has been defined to have a little-endian memory system, with big-endian or bi-endian as non-standard variants. * Load-Reserved/Store-Conditional (`LR/SC`) instructions have been added in the atomic instruction extension. * `AMOs` and `LR/SC` can support the release consistency model. * The `FENCE` instruction provides finer-grain memory and I/O orderings. * An `AMO` for fetch-and-`XOR` (`AMOXOR`) has been added, and the encoding for`AMOSWAP` has been changed to make room. * The `AUIPC` instruction, which adds a 20-bit upper immediate to the `PC`, replaces the `RDNPC` instruction, which only read the current `PC` value. This results in significant savings for position-independent code. * The `JAL` instruction has now moved to the `U-Type` format with an explicit destination register, and the `J` instruction has been dropped being replaced by `JAL` with _rd_\=`x0`. This removes the only instruction with an implicit destination register and removes the `J-Type` instruction format from the base ISA. There is an accompanying reduction in `JAL`reach, but a significant reduction in base ISA complexity. * The static hints on the `JALR` instruction have been dropped. The hints are redundant with the _rd_ and _rs1_ register specifiers for code compliant with the standard calling convention. * The `JALR` instruction now clears the lowest bit of the calculated target address, to simplify hardware and to allow auxiliary information to be stored in function pointers. * The `MFTX.S` and `MFTX.D` instructions have been renamed to `FMV.X.S` and`FMV.X.D`, respectively. Similarly, `MXTF.S` and `MXTF.D` instructions have been renamed to `FMV.S.X` and `FMV.D.X`, respectively. * The `MFFSR` and `MTFSR` instructions have been renamed to `FRCSR` and `FSCSR`, respectively. `FRRM`, `FSRM`, `FRFLAGS`, and `FSFLAGS` instructions have been added to individually access the rounding mode and exception flags subfields of the `fcsr`. * The `FMV.X.S` and `FMV.X.D` instructions now source their operands from_rs1_, instead of _rs2_. This change simplifies datapath design. * `FCLASS.S` and `FCLASS.D` floating-point classify instructions have been added. * A simpler NaN generation and propagation scheme has been adopted. * For `RV32I`, the system performance counters have been extended to 64-bits wide, with separate read access to the upper and lower 32 bits. * Canonical `NOP` and `MV` encodings have been defined. * Standard instruction-length encodings have been defined for 48-bit, 64-bit, and >64-bit instructions. * Description of a 128-bit address space variant, `RV128`, has been added. * Major opcodes in the 32-bit base instruction format have been allocated for user-defined custom extensions. * A typographical error that suggested that stores source their data from _rd_ has been corrected to refer to _rs2_. 7.1. "Zicntr" and "Zihpm" Extensions for Counters, Version 2.0 ==================== ## [](#counters)7.1\. "Zicntr" and "Zihpm" Extensions for Counters, Version 2.0 RISC-V ISAs provide a set of up to thirty-two 64-bit performance counters and timers that are accessible via unprivileged XLEN-bit read-only CSR registers `0xC00`–`0xC1F` (when XLEN=32, the upper 32 bits are accessed via CSR registers `0xC80`–`0xC9F`). These counters are divided between the "Zicntr" and "Zihpm" extensions. ### [](#7-1-1-zicntr-extension-for-base-counters-and-timers)7.1.1\. "Zicntr" Extension for Base Counters and Timers The Zicntr standard extension comprises the first three of these counters (CYCLE, TIME, and INSTRET), which have dedicated functions (cycle count, real-time clock, and instructions retired, respectively). The Zicntr extension depends on the Zicsr extension. | | We recommend provision of these basic counters in implementations as they are essential for basic performance analysis, adaptive and dynamic optimization, and to allow an application to work with real-time streams. Additional counters in the separate Zihpm extension can help diagnose performance problems and these should be made accessible from user-level application code with low overhead. Some execution environments might prohibit access to counters, for example, to impede timing side-channel attacks. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ![svg](_images/svg-c5fa6f6235d53a40d7b7ba92c4b9c90f1285e6b2.svg) For base ISAs with XLEN≥64, CSR instructions can access the full 64-bit CSRs directly. In particular, the RDCYCLE, RDTIME, and RDINSTRET pseudoinstructions read the full 64 bits of the `cycle`,`time`, and `instret` counters. | | The counter pseudoinstructions are mapped to the read-onlycsrrs rd, counter, x0 canonical form, but the other read-only CSR instruction forms (based on CSRRC/CSRRSI/CSRRCI) are also legal ways to read these CSRs. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For base ISAs with XLEN=32, the Zicntr extension enables the three 64-bit read-only counters to be accessed in 32-bit pieces. The RDCYCLE, RDTIME, and RDINSTRET pseudoinstructions provide the lower 32 bits, and the RDCYCLEH, RDTIMEH, and RDINSTRETH pseudoinstructions provide the upper 32 bits of the respective counters. | | We required the counters be 64 bits wide, even when XLEN=32, as otherwise it is very difficult for software to determine if values have overflowed. The sample code given below shows how the full 64-bit width value can be safely read using the individual 32-bit width pseudoinstructions. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The RDCYCLE pseudoinstruction reads the low XLEN bits of the `cycle`CSR which holds a count of the number of clock cycles executed by the processor core on which the hart is running from an arbitrary start time in the past. RDCYCLEH is only present when XLEN=32 and reads bits 63-32 of the same cycle counter. The underlying 64-bit counter should never overflow in practice. The rate at which the cycle counter advances will depend on the implementation and operating environment. The execution environment should provide a means to determine the current rate (cycles/second) at which the cycle counter is incrementing. | | RDCYCLE is intended to return the number of cycles executed by the processor core, not the hart. Precisely defining what is a "core" is difficult given some implementation choices (e.g., AMD Bulldozer). Precisely defining what is a "clock cycle" is also difficult given the range of implementations (including software emulations), but the intent is that RDCYCLE is used for performance monitoring along with the other performance counters. In particular, where there is one hart/core, one would expect cycle-count/instructions-retired to measure CPI for a hart. Cores don’t have to be exposed to software at all, and an implementer might choose to pretend multiple harts on one physical core are running on separate cores with one hart/core, and provide separate cycle counters for each hart. This might make sense in a simple barrel processor (e.g., CDC 6600 peripheral processors) where inter-hart timing interactions are non-existent or minimal. Where there is more than one hart/core and dynamic multithreading, it is not generally possible to separate out cycles per hart (especially with SMT). It might be possible to define a separate performance counter that tried to capture the number of cycles a particular hart was running, but this definition would have to be very fuzzy to cover all the possible threading implementations. For example, should we only count cycles for which any instruction was issued to execution for this hart, and/or cycles any instruction retired, or include cycles this hart was occupying machine resources but couldn’t execute due to stalls while other harts went into execution? Likely, "all of the above" would be needed to have understandable performance stats. This complexity of defining a per-hart cycle count, and also the need in any case for a total per-core cycle count when tuning multithreaded code led to just standardizing the per-core cycle counter, which also happens to work well for the common single hart/core case. Standardizing what happens during "sleep" is not practical given that what "sleep" means is not standardized across execution environments, but if the entire core is paused (entirely clock-gated or powered-down in deep sleep), then it is not executing clock cycles, and the cycle count shouldn’t be increasing per the spec. There are many details, e.g., whether clock cycles required to reset a processor after waking up from a power-down event should be counted, and these are considered execution-environment-specific details. Even though there is no precise definition that works for all platforms, this is still a useful facility for most platforms, and an imprecise, common, "usually correct" standard here is better than no standard. The intent of RDCYCLE was primarily performance monitoring/tuning, and the specification was written with that goal in mind. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The RDTIME pseudoinstruction reads the low XLEN bits of the "time" CSR, which counts wall-clock real time that has passed from an arbitrary start time in the past. RDTIMEH is only present when XLEN=32 and reads bits 63-32 of the same real-time counter. The underlying 64-bit counter increments by one with each tick of the real-time clock, and, for realistic real-time clock frequencies, should never overflow in practice. The execution environment should provide a means of determining the period of a counter tick (seconds/tick). The period should be constant within a small error bound. The environment should provide a means to determine the accuracy of the clock (i.e., the maximum relative error between the nominal and actual real-time clock periods). | | On some simple platforms, cycle count might represent a valid implementation of RDTIME, in which case RDTIME and RDCYCLE may return the same result. It is difficult to provide a strict mandate on clock period given the wide variety of possible implementation platforms. The maximum error bound should be set based on the requirements of the platform. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The real-time clocks of all harts must be synchronized to within one tick of the real-time clock. | | As with other architectural mandates, it suffices to appear "as if" harts are synchronized to within one tick of the real-time clock, i.e., software is unable to observe that there is a greater delta between the real-time clock values observed on two harts. If, for example, the real-time clock increments at a frequency of 1 GHz, then all harts must appear to be synchronized to within 1 nsec. But it is also acceptable for this example implementation to only update the real-time clock at, say, a frequency of 100 MHz with increments of 10 ticks. As long as software cannot observe this seeming violation of the above synchronization requirement, and software always observes time across harts to be monotonically nondecreasing, then this implementation is compliant. A platform spec may then, for example, specify an apparent real-time clock tick frequency (e.g. 1 GHz) and also a minimum update frequency (e.g. 100 MHz) at which updated time values are guaranteed to be observable by software. Software may read time more frequently, but it should only observe monotonically nondecreasing values and it should observe a new value at least once every 10 ns (corresponding to the 100 MHz update frequency in this example). | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The RDINSTRET pseudoinstruction reads the low XLEN bits of the`instret` CSR, which counts the number of instructions retired by this hart from some arbitrary start point in the past. RDINSTRETH is only present when XLEN=32 and reads bits 63-32 of the same instruction counter. The underlying 64-bit counter should never overflow in practice. | | Instructions that cause synchronous exceptions, including ECALL and EBREAK, are not considered to retire and hence do not increment theinstret CSR. | | ------------------------------------------------------------------------------------------------------------------------------------------------------ | The following code sequence will read a valid 64-bit cycle counter value into `x3:x2`, even if the counter overflows its lower half between reading its upper and lower halves. Sample code for reading the 64-bit cycle counter when XLEN=32. ```asm. again: rdcycleh x3 rdcycle x2 rdcycleh x4 bne x3, x4, again ``` ### [](#7-1-2-zihpm-extension-for-hardware-performance-counters)7.1.2\. "Zihpm" Extension for Hardware Performance Counters The Zihpm extension comprises up to 29 additional unprivileged 64-bit hardware performance counters, `hpmcounter3-hpmcounter31`. When XLEN=32, the upper 32 bits of these performance counters are accessible via additional CSRs `hpmcounter3h- hpmcounter31h`. The Zihpm extension depends on the Zicsr extension. | | In some applications, it is important to be able to read multiple counters at the same instant in time. When run under a multitasking environment, a user thread can suffer a context switch while attempting to read the counters. One solution is for the user thread to read the real-time counter before and after reading the other counters to determine if a context switch occurred in the middle of the sequence, in which case the reads can be retried. We considered adding output latches to allow a user thread to snapshot the counter values atomically, but this would increase the size of the user context, especially for implementations with a richer set of counters. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The implemented number and width of these additional counters, and the set of events they count, are platform-specific. Accessing an unimplemented counter may cause an illegal-instruction exception or may return a constant value. If the configuration used to select the events counted by a counter is misconfigured, the counter may return a constant value. The execution environment should provide a means to determine the number and width of the implemented counters, and an interface to configure the events to be counted by each counter. | | For execution environments implemented on RISC-V privileged platforms, the privileged architecture manual describes privileged CSRs controlling access by lower privileged modes to these counters, and to set the events to be counted. Alternative execution environments (e.g., user-level-only software performance models) may provide alternative mechanisms to configure the events counted by the performance counters. It would be useful to eventually standardize event settings to count ISA-level metrics, such as the number of floating-point instructions executed for example, and possibly a few common microarchitectural metrics, such as "L1 instruction cache misses". | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 22.1. "D" Extension for Double-Precision Floating-Point, Version 2.2 ==================== ## [](#22-1-d-extension-for-double-precision-floating-point-version-2-2)22.1\. "D" Extension for Double-Precision Floating-Point, Version 2.2 This chapter describes the standard double-precision floating-point instruction-set extension, which is named "D" and adds double-precision floating-point computational instructions compliant with the IEEE 754-2008 arithmetic standard. The D extension depends on the base single-precision instruction subset F. ### [](#22-1-1-d-register-state)22.1.1\. D Register State The D extension widens the 32 floating-point registers, `f0-f31`, to 64 bits (FLEN=64 in [RISC-V standard F extension single-precision floating-point state](f-st-ext.html#fprs). The `f` registers can now hold either 32-bit or 64-bit floating-point values as described below in [22.1.2\. NaN Boxing of Narrower Values](#nanboxing). | | FLEN can be 32, 64, or 128 depending on which of the F, D, and Q extensions are supported. There can be up to four different floating-point precisions supported, including H, F, D, and Q. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#nanboxing)22.1.2\. NaN Boxing of Narrower Values When multiple floating-point precisions are supported, then valid values of narrower _n_\-bit types, _n_32, FCVT.W\[U\].S sign-extends the 32-bit result to the destination register width.FCVT.L\[U\].S and FCVT.S.L\[U\] are RV64-only instructions.If the rounded result is not representable in the destination format, it is clipped to the nearest value and the invalid flag is set. [Table 5](#int%5Fconv) gives the range of valid inputs for FCVT._int_.S and the behavior for invalid inputs. All floating-point to integer and integer to floating-point conversion instructions round according to the _rm_ field. A floating-point register can be initialized to floating-point positive zero using FCVT.S.W _rd_, `x0`, which will never set any exception flags. __Table 5\. Domains of float-to-integer conversions and behavior for invalid inputs__ | FCVT.W.S | FCVT.WU.S | FCVT.L.S | FCVT.LU.S | | | -------------------------------------- | --------- | -------- | --------- | ----- | | Minimum valid input (after rounding) | −231 | 0 | −263 | 0 | | Maximum valid input (after rounding) | 231−1 | 232−1 | 263−1 | 264−1 | | Output for out-of-range negative input | −231 | 0 | −263 | 0 | | Output for -∞ | −231 | 0 | −263 | 0 | | Output for out-of-range positive input | 231−1 | 232−1 | 263−1 | 264−1 | | Output for +∞ or NaN | 231−1 | 232−1 | 263−1 | 264−1 | All floating-point conversion instructions set the Inexact exception flag if the rounded result differs from the operand value and the Invalid exception flag is not set. ![svg](_images/svg-6e5e0d1572d9a0e05de079d28254257e4ae7ed36.svg) Floating-point to floating-point sign-injection instructions, FSGNJ.S, FSGNJN.S, and FSGNJX.S, produce a result that takes all bits except the sign bit from _rs1_. For FSGNJ, the result’s sign bit is _rs2_'s sign bit; for FSGNJN, the result’s sign bit is the opposite of _rs2_'s sign bit; and for FSGNJX, the sign bit is the XOR of the sign bits of _rs1_and _rs2_. Sign-injection instructions do not set floating-point exception flags, nor do they canonicalize NaNs. Note, FSGNJ.S _rx, ry, ry_ moves _ry_ to _rx_ (assembler pseudoinstruction FMV.S _rx, ry_); FSGNJN.S _rx, ry, ry_ moves the negation of _ry_ to _rx_ (assembler pseudoinstruction FNEG.S _rx, ry_); and FSGNJX.S _rx, ry, ry_ moves the absolute value of _ry_ to _rx_ (assembler pseudoinstruction FABS.S _rx, ry_). ![svg](_images/svg-2591776373d7209c5d709f100045f2f7c2465330.svg) | | The sign-injection instructions provide floating-point MV, ABS, and NEG, as well as supporting a few other operations, including the IEEEcopySign operation and sign manipulation in transcendental math function libraries. Although MV, ABS, and NEG only need a single register operand, whereas FSGNJ instructions need two, it is unlikely most microarchitectures would add optimizations to benefit from the reduced number of register reads for these relatively infrequent instructions. Even in this case, a microarchitecture can simply detect when both source registers are the same for FSGNJ instructions and only read a single copy. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Instructions are provided to move bit patterns between the floating-point and integer registers. FMV.X.W moves the single-precision value in floating-point register _rs1_ represented in IEEE 754-2008 encoding to the lower 32 bits of integer register _rd_. The bits are not modified in the transfer, and in particular, the payloads of non-canonical NaNs are preserved. For RV64, the higher 32 bits of the destination register are filled with copies of the floating-point number’s sign bit. FMV.W.X moves the single-precision value encoded in IEEE 754-2008 standard encoding from the lower 32 bits of integer register _rs1_ to the floating-point register _rd_. The bits are not modified in the transfer, and in particular, the payloads of non-canonical NaNs are preserved. | | The FMV.W.X and FMV.X.W instructions were previously called FMV.S.X and FMV.X.S. The use of W is more consistent with their semantics as an instruction that moves 32 bits without interpreting them. This became clearer after defining NaN-boxing. To avoid disturbing existing code, both the W and S versions will be supported by tools. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ![svg](_images/svg-f43dbe8c77818963355ae63ce6f5989bf5a77511.svg) | | The base floating-point ISA was defined so as to allow implementations to employ an internal recoding of the floating-point format in registers to simplify handling of subnormal values and possibly to reduce functional unit latency. To this end, the F extension avoids representing integer values in the floating-point registers by defining conversion and comparison operations that read and write the integer register file directly. This also removes many of the common cases where explicit moves between integer and floating-point registers are required, reducing instruction count and critical paths for common mixed-format code sequences. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#21-1-8-single-precision-floating-point-compare-instructions)21.1.8\. Single-Precision Floating-Point Compare Instructions Floating-point compare instructions (FEQ.S, FLT.S, FLE.S) perform the specified comparison between floating-point registers (_rs1_ \= _rs2_, _rs1_ < _rs2_,_rs1_ ≤ _rs2_) writing 1 to the integer register _rd_ if the condition holds, and 0 otherwise. FLT.S and FLE.S perform what the IEEE 754-2008 standard refers to as_signaling_ comparisons: that is, they set the invalid operation exception flag if either input is NaN. FEQ.S performs a _quiet_comparison: it only sets the invalid operation exception flag if either input is a signaling NaN. For all three instructions, the result is 0 if either operand is NaN. ![svg](_images/svg-f47c9efb18a201eaac373e9194bf1b6416895153.svg) | | The F extension provides a ≤ comparison, whereas the base ISAs provide a ≥ branch comparison. Because ≤ can be synthesized from ≥ and vice-versa, there is no performance implication to this inconsistency, but it is nevertheless an unfortunate incongruity in the ISA. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#21-1-9-single-precision-floating-point-classify-instruction)21.1.9\. Single-Precision Floating-Point Classify Instruction The FCLASS.S instruction examines the value in floating-point register_rs1_ and writes to integer register _rd_ a 10-bit mask that indicates the class of the floating-point number. The format of the mask is described in [Table 6](#fclass). The corresponding bit in _rd_ will be set if the property is true and clear otherwise. All other bits in_rd_ are cleared. Note that exactly one bit in _rd_ will be set. FCLASS.S does not set the floating-point exception flags. ![svg](_images/svg-d11bade1f026be22713c4cc93baf3dc133764e37.svg) __Table 6\. Format of result of FCLASS instruction.__ | _rd_ bit | Meaning | | -------- | ------------------------------------- | | 0 | _rs1_ is −∞. | | 1 | _rs1_ is a negative normal number. | | 2 | _rs1_ is a negative subnormal number. | | 3 | _rs1_ is −0. | | 4 | _rs1_ is +0. | | 5 | _rs1_ is a positive subnormal number. | | 6 | _rs1_ is a positive normal number. | | 7 | _rs1_ is +∞. | | 8 | _rs1_ is a signaling NaN. | | 9 | _rs1_ is a quiet NaN. | 1.1. Introduction ==================== ## [](#1-1-introduction)1.1\. Introduction RISC-V (pronounced "risk-five") is a new instruction-set architecture (ISA) that was originally designed to support computer architecture research and education, but which we now hope will also become a standard free and open architecture for industry implementations. Our goals in defining RISC-V include: * A completely _open_ ISA that is freely available to academia and industry. * A _real_ ISA suitable for direct native hardware implementation, not just simulation or binary translation. * An ISA that avoids "over-architecting" for a particular microarchitecture style (e.g., microcoded, in-order, decoupled, out-of-order) or implementation technology (e.g., full-custom, ASIC, FPGA), but which allows efficient implementation in any of these. * An ISA separated into a _small_ base integer ISA, usable by itself as a base for customized accelerators or for educational purposes, and optional standard extensions, to support general-purpose software development. * Support for the revised 2008 IEEE-754 floating-point standard. \[[4](../biblio/bibliography.html#bib-ieee754-2008)\] * An ISA supporting extensive ISA extensions and specialized variants. * Both 32-bit and 64-bit address space variants for applications, operating system kernels, and hardware implementations. * An ISA with support for highly parallel multicore or manycore implementations, including heterogeneous multiprocessors. * Optional _variable-length instructions_ to both expand available instruction encoding space and to support an optional _dense instruction encoding_ for improved performance, static code size, and energy efficiency. * A fully virtualizable ISA to ease hypervisor development. * An ISA that simplifies experiments with new privileged architecture designs. | | Commentary on our design decisions is formatted as in this paragraph. This non-normative text can be skipped if the reader is only interested in the specification itself. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The name RISC-V was chosen to represent the fifth major RISC ISA design from UC Berkeley (RISC-I \[[5](../biblio/bibliography.html#bib-risci-isca1981)\], RISC-II \[[6](../biblio/bibliography.html#bib-katevenis:1983)\], SOAR \[[7](../biblio/bibliography.html#bib-ungar:1984)\], and SPUR \[[8](../biblio/bibliography.html#bib-spur-jsscc1989)\] were the first four). We also pun on the use of the Roman numeral "V" to signify "variations" and "vectors", as support for a range of architecture research, including various data-parallel accelerators, is an explicit goal of the ISA design. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The RISC-V ISA is defined avoiding implementation details as much as possible (although commentary is included on implementation-driven decisions) and should be read as the software-visible interface to a wide variety of implementations rather than as the design of a particular hardware artifact. The RISC-V manual is structured in two volumes. This volume covers the design of the base _unprivileged_instructions, including optional unprivileged ISA extensions. Unprivileged instructions are those that are generally usable in all privilege modes in all privileged architectures, though behavior might vary depending on privilege mode and privilege architecture. The second volume provides the design of the first ("classic") privileged architecture. The manuals use IEC 80000-13:2008 conventions, with a byte of 8 bits. | | In the unprivileged ISA design, we tried to remove any dependence on particular microarchitectural features, such as cache line size, or on privileged architecture details, such as page translation. This is both for simplicity and to allow maximum flexibility for alternative microarchitectures or alternative privileged architectures. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#1-1-1-risc-v-hardware-platform-terminology)1.1.1\. RISC-V Hardware Platform Terminology A RISC-V hardware platform can contain one or more RISC-V-compatible processing cores together with other non-RISC-V-compatible cores, fixed-function accelerators, various physical memory structures, I/O devices, and an interconnect structure to allow the components to communicate. A component is termed a _core_ if it contains an independent instruction fetch unit. A RISC-V-compatible core might support multiple RISC-V-compatible hardware threads, or _harts_, through multithreading. A RISC-V core might have additional specialized instruction-set extensions or an added _coprocessor_. We use the term _coprocessor_ to refer to a unit that is attached to a RISC-V core and is mostly sequenced by a RISC-V instruction stream, but which contains additional architectural state and instruction-set extensions, and possibly some limited autonomy relative to the primary RISC-V instruction stream. We use the term _accelerator_ to refer to either a non-programmable fixed-function unit or a core that can operate autonomously but is specialized for certain tasks. In RISC-V systems, we expect many programmable accelerators will be RISC-V-based cores with specialized instruction-set extensions and/or customized coprocessors. An important class of RISC-V accelerators are I/O accelerators, which offload I/O processing tasks from the main application cores. The system-level organization of a RISC-V hardware platform can range from a single-core microcontroller to a many-thousand-node cluster of shared-memory manycore server nodes. Even small systems-on-a-chip might be structured as a hierarchy of multicomputers and/or multiprocessors to modularize development effort or to provide secure isolation between subsystems. ### [](#1-1-2-risc-v-software-execution-environments-and-harts)1.1.2\. RISC-V Software Execution Environments and Harts The behavior of a RISC-V program depends on the execution environment in which it runs. A RISC-V execution environment interface (EEI) defines the initial state of the program, the number and type of harts in the environment including the privilege modes supported by the harts, the accessibility and attributes of memory and I/O regions, the behavior of all legal instructions executed on each hart (i.e., the ISA is one component of the EEI), and the handling of any interrupts or exceptions raised during execution including environment calls. Examples of EEIs include the Linux application binary interface (ABI), or the RISC-V supervisor binary interface (SBI). The implementation of a RISC-V execution environment can be pure hardware, pure software, or a combination of hardware and software. For example, opcode traps and software emulation can be used to implement functionality not provided in hardware. Examples of execution environment implementations include: * "Bare metal" hardware platforms where harts are directly implemented by physical processor threads and instructions have full access to the physical address space. The hardware platform defines an execution environment that begins at power-on reset. * RISC-V operating systems that provide multiple user-level execution environments by multiplexing user-level harts onto available physical processor threads and by controlling access to memory via virtual memory. * RISC-V hypervisors that provide multiple supervisor-level execution environments for guest operating systems. * RISC-V emulators, such as Spike, QEMU or rv8, which emulate RISC-V harts on an underlying x86 system, and which can provide either a user-level or a supervisor-level execution environment. | | A bare hardware platform can be considered to define an EEI, where the accessible harts, memory, and other devices populate the environment, and the initial state is that at power-on reset. Generally, most software is designed to use a more abstract interface to the hardware, as more abstract EEIs provide greater portability across different hardware platforms. Often EEIs are layered on top of one another, where one higher-level EEI uses another lower-level EEI. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | From the perspective of software running in a given execution environment, a hart is a resource that autonomously fetches and executes RISC-V instructions within that execution environment. In this respect, a hart behaves like a hardware thread resource even if time-multiplexed onto real hardware by the execution environment. Some EEIs support the creation and destruction of additional harts, for example, via environment calls to fork new harts. The execution environment is responsible for ensuring the eventual forward progress of each of its harts. For a given hart, that responsibility is suspended while the hart is exercising a mechanism that explicitly waits for an event, such as the wait-for-interrupt instruction defined in Volume II of this specification; and that responsibility ends if the hart is terminated. The following events constitute forward progress: * The retirement of an instruction. * A trap, as defined in [1.1.6\. Exceptions, Traps, and Interrupts](#trap-defn). * Any other event defined by an extension to constitute forward progress. | | The term hart was introduced in the work on Lithe \[[9](../biblio/bibliography.html#bib-lithe-pan-hotpar09)\] and \[[10](../biblio/bibliography.html#bib-lithe-pan-pldi10)\] to provide a term to represent an abstract execution resource as opposed to a software thread programming abstraction. The important distinction between a hardware thread (hart) and a software thread context is that the software running inside an execution environment is not responsible for causing progress of each of its harts; that is the responsibility of the outer execution environment. So the environment’s harts operate like hardware threads from the perspective of the software inside the execution environment. An execution environment implementation might time-multiplex a set of guest harts onto fewer host harts provided by its own execution environment but must do so in a way that guest harts operate like independent hardware threads. In particular, if there are more guest harts than host harts then the execution environment must be able to preempt the guest harts and must not wait indefinitely for guest software on a guest hart to "yield" control of the guest hart. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#1-1-3-risc-v-isa-overview)1.1.3\. RISC-V ISA Overview A RISC-V ISA is defined as a base integer ISA, which must be present in any implementation, plus optional extensions to the base ISA. The base integer ISAs are very similar to that of the early RISC processors except with no branch delay slots and with support for optional variable-length instruction encodings. A base is carefully restricted to a minimal set of instructions sufficient to provide a reasonable target for compilers, assemblers, linkers, and operating systems (with additional privileged operations), and so provides a convenient ISA and software toolchain "skeleton" around which more customized processor ISAs can be built. Although it is convenient to speak of _the_ RISC-V ISA, RISC-V is actually a family of related ISAs, of which there are currently four base ISAs. Each base integer instruction set is characterized by the width of the integer registers and the corresponding size of the address space and by the number of integer registers. There are two primary base integer variants, RV32I and RV64I, described in[RV32E and RV64E Base Integer Instruction Sets](rv32e.html) and [RV64I Base Integer Instruction Set](rv64.html), which provide 32-bit or 64-bit address spaces respectively. We use the term XLEN to refer to the width of an integer register in bits (either 32 or 64). [RV32E and RV64E Base Integer Instruction Sets](rv32e.html) describes the RV32E and RV64E subset variants of the RV32I or RV64I base instruction sets respectively, which have been added to support small microcontrollers, and which have half the number of integer registers.The base integer instruction sets use a two’s-complement representation for signed integer values. | | Although 64-bit address spaces are a requirement for larger systems, we believe 32-bit address spaces will remain adequate for many embedded and client devices for decades to come and will be desirable to lower memory traffic and energy consumption. In addition, 32-bit address spaces are sufficient for educational purposes. A larger flat 128-bit address space might eventually be required and could be accommodated with a new RV128I base ISA within the existing RISC-V ISA framework. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The four base ISAs in RISC-V are treated as distinct base ISAs. A common question is why is there not a single ISA, and in particular, why is RV32I not a strict subset of RV64I? Some earlier ISA designs (SPARC, MIPS) adopted a strict superset policy when increasing address space size to support running existing 32-bit binaries on new 64-bit hardware. The main advantage of explicitly separating base ISAs is that each base ISA can be optimized for its needs without requiring to support all the operations needed for other base ISAs. For example, RV64I can omit instructions and CSRs that are only needed to cope with the narrower registers in RV32I. The RV32I variants can use encoding space otherwise reserved for instructions only required by wider address-space variants. The main disadvantage of not treating the design as a single ISA is that it complicates the hardware needed to emulate one base ISA on another (e.g., RV32I on RV64I). However, differences in addressing and illegal-instruction traps generally mean some mode switch would be required in hardware in any case even with full superset instruction encodings, and the different RISC-V base ISAs are similar enough that supporting multiple versions is relatively low cost. Although some have proposed that the strict superset design would allow legacy 32-bit libraries to be linked with 64-bit code, this is impractical in practice, even with compatible encodings, due to the differences in software calling conventions and system-call interfaces. The RISC-V privileged architecture provides fields in misa to control the unprivileged ISA at each level to support emulating different base ISAs on the same hardware. We note that newer SPARC and MIPS ISA revisions have deprecated support for running 32-bit code unchanged on 64-bit systems. A related question is why there is a different encoding for 32-bit adds in RV32I (ADD) and RV64I (ADDW)? The ADDW opcode could be used for 32-bit adds in RV32I and ADDD for 64-bit adds in RV64I, instead of the existing design which uses the same opcode ADD for 32-bit adds in RV32I and 64-bit adds in RV64I with a different opcode ADDW for 32-bit adds in RV64I. This would also be more consistent with the use of the same LW opcode for 32-bit load in both RV32I and RV64I. The very first versions of RISC-V ISA did have a variant of this alternate design, but the RISC-V design was changed to the current choice in January 2011\. Our focus was on supporting 32-bit integers in the 64-bit ISA, not on providing compatibility with the 32-bit ISA, and the motivation was to remove the asymmetry that arose from having not all opcodes in RV32I have a \*W suffix (e.g., ADDW, but AND not ANDW). In hindsight, this was perhaps not well-justified and a consequence of designing both ISAs at the same time as opposed to adding one later to sit on top of another, and also from a belief we had to fold platform requirements into the ISA spec which would imply that all the RV32I instructions would have been required in RV64I. It is too late to change the encoding now, but this is also of little practical consequence for the reasons stated above. It has been noted we could enable the \*W variants as an extension to RV32I systems to provide a common encoding across RV64I and a future RV32 variant. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | RISC-V has been designed to support extensive customization and specialization. Each base integer ISA can be extended with one or more optional instruction-set extensions. An extension may be categorized as either standard, custom, or non-conforming. For this purpose, we divide each RISC-V instruction-set encoding space (and related encoding spaces such as the CSRs) into three disjoint categories: _standard_,_reserved_, and _custom_. Standard extensions and encodings are defined by RISC-V International; any extensions not defined by RISC-V International are_non-standard_. Each base ISA and its standard extensions use only standard encodings, and shall not conflict with each other in their uses of these encodings. Reserved encodings are currently not defined but are saved for future standard extensions; once thus used, they become standard encodings. Custom encodings shall never be used for standard extensions and are made available for vendor-specific non-standard extensions. Non-standard extensions are either custom extensions, that use only custom encodings, or _non-conforming_ extensions, that use any standard or reserved encoding. Instruction-set extensions are generally shared but may provide slightly different functionality depending on the base ISA. We have also developed a naming convention for RISC-V base instructions and instruction-set extensions, described in detail in [ISA Extension Naming Conventions](naming.html). To support more general software development, a set of standard extensions are defined to provide integer multiply/divide, atomic operations, and single and double-precision floating-point arithmetic. The base integer ISA is named "I" (prefixed by RV32 or RV64 depending on integer register width), and contains integer computational instructions, integer loads, integer stores, and control-flow instructions. The standard integer multiplication and division extension is named "M", and adds instructions to multiply and divide values held in the integer registers. The standard atomic instruction extension, denoted by "A", adds instructions that atomically read, modify, and write memory for inter-processor synchronization. The standard single-precision floating-point extension, denoted by "F", adds floating-point registers, single-precision computational instructions, and single-precision loads and stores. The standard double-precision floating-point extension, denoted by "D", expands the floating-point registers, and adds double-precision computational instructions, loads, and stores. The standard "C" compressed instruction extension provides narrower 16-bit forms of common instructions. Beyond the base integer ISA and these standard extensions, we believe it is rare that a new instruction will provide a significant benefit for all applications, although it may be very beneficial for a certain domain. As energy efficiency concerns are forcing greater specialization, we believe it is important to simplify the required portion of an ISA specification. Whereas other architectures usually treat their ISA as a single entity, which changes to a new version as instructions are added over time, RISC-V will endeavor to keep the base and each standard extension constant over time, and instead layer new instructions as further optional extensions. For example, the base integer ISAs will continue as fully supported standalone ISAs, regardless of any subsequent extensions. ### [](#1-1-4-memory)1.1.4\. Memory A RISC-V hart has a single byte-addressable address space of 2XLEN bytes for all memory accesses. A _word_ of memory is defined as 32 bits (4 bytes). Correspondingly, a _halfword_ is 16 bits (2 bytes), a_doubleword_ is 64 bits (8 bytes), and a _quadword_ is 128 bits (16 bytes). The memory address space is circular, so that the byte at address 2XLEN−1 is adjacent to the byte at address zero. Accordingly, memory address computations done by the hardware ignore overflow and instead wrap around modulo 2XLEN. The execution environment determines the mapping of hardware resources into a hart’s address space. Different address ranges of a hart’s address space may (1) contain _main memory_, or (2) contain one or more _I/O devices_. Reads and writes of I/O devices may have visible side effects, but accesses to main memory cannot. Vacant address ranges are not a separate category but can be represented as either main memory or I/O regions that are not accessible. Although it is possible for the execution environment to call everything in a hart’s address space an I/O device, it is usually expected that some portion will be specified as main memory. When a RISC-V platform has multiple harts, the address spaces of any two harts may be entirely the same, or entirely different, or may be partly different but sharing some subset of resources, mapped into the same or different address ranges. | | For a purely "bare metal" environment, all harts may see an identical address space, accessed entirely by physical addresses. However, when the execution environment includes an operating system employing address translation, it is common for each hart to be given a virtual address space that is largely or entirely its own. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Executing each RISC-V machine instruction entails one or more memory accesses, subdivided into _implicit_ and _explicit_ accesses. For each instruction executed, an _implicit_ memory read (instruction fetch) is done to obtain the encoded instruction to execute. Many RISC-V instructions perform no further memory accesses beyond instruction fetch. Specific load and store instructions perform an _explicit_ read or write of memory at an address determined by the instruction. The execution environment may dictate that instruction execution performs other _implicit_ memory accesses (such as to implement address translation) beyond those documented for the unprivileged ISA. The execution environment determines what portions of the address space are accessible for each kind of memory access. For example, the set of locations that can be implicitly read for instruction fetch may or may not have any overlap with the set of locations that can be explicitly read by a load instruction; and the set of locations that can be explicitly written by a store instruction may be only a subset of locations that can be read. Ordinarily, if an instruction attempts to access memory at an inaccessible address, an exception is raised for the instruction. Except when specified otherwise, implicit reads that do not raise an exception and that have no side effects may occur arbitrarily early and speculatively, even before the machine could possibly prove that the read will be needed. For instance, a valid implementation could attempt to read all of main memory at the earliest opportunity, cache as many fetchable (executable) bytes as possible for later instruction fetches, and avoid reading main memory for instruction fetches ever again. To ensure that certain implicit reads are ordered only after writes to the same memory locations, software must execute specific fence or cache-control instructions defined for this purpose (such as the FENCE.I instruction defined in ["Zifencei" Extension for Instruction-Fetch Fence](zifencei.html)). The memory accesses (implicit or explicit) made by a hart may appear to occur in a different order as perceived by another hart or by any other agent that can access the same memory. This perceived reordering of memory accesses is always constrained, however, by the applicable memory consistency model. The default memory consistency model for RISC-V is the RISC-V Weak Memory Ordering (RVWMO), defined in[RVWMO Memory Consistency Model](rvwmo.html) and in appendices. Optionally, an implementation may adopt the stronger model of Total Store Ordering, as defined in ["Ztso" Extension for Total Store Ordering](ztso-st-ext.html). The execution environment may also add constraints that further limit the perceived reordering of memory accesses. Since the RVWMO model is the weakest model allowed for any RISC-V implementation, software written for this model is compatible with the actual memory consistency rules of all RISC-V implementations. As with implicit reads, software must execute fence or cache-control instructions to ensure specific ordering of memory accesses beyond the requirements of the assumed memory consistency model and execution environment. ### [](#1-1-5-base-instruction-length-encoding)1.1.5\. Base Instruction-Length Encoding The base RISC-V ISA has fixed-length 32-bit instructions that must be naturally aligned on 32-bit boundaries. However, the standard RISC-V encoding scheme is designed to support ISA extensions with variable-length instructions, where each instruction can be any number of 16-bit instruction _parcels_ in length and parcels are naturally aligned on 16-bit boundaries. The standard compressed ISA extension described in ["C" Extension for Compressed Instructions](c-st-ext.html) reduces code size by providing compressed 16-bit instructions and relaxes the alignment constraints to allow all instructions (16 bit and 32 bit) to be aligned on any 16-bit boundary to improve code density. We use the term IALIGN (measured in bits) to refer to the instruction-address alignment constraint the implementation enforces. IALIGN is 32 bits in the base ISA, but some ISA extensions, including the compressed ISA extension, relax IALIGN to 16 bits. IALIGN may not take on any value other than 16 or 32. We use the term ILEN (measured in bits) to refer to the maximum instruction length supported by an implementation, and which is always a multiple of IALIGN. For implementations supporting only a base instruction set, ILEN is 32 bits. Implementations supporting longer instructions have larger values of ILEN. All the 32-bit instructions in the base ISA have their lowest two bits set to `11`. The optional compressed 16-bit instruction-set extensions have their lowest two bits equal to `00`, `01`, or `10`. | | Given the code size and energy savings of a compressed format, we wanted to build in support for a compressed format to the ISA encoding scheme rather than adding this as an afterthought, but to allow simpler implementations we didn’t want to make the compressed format mandatory. We also wanted to optionally allow longer instructions to support experimentation and larger instruction-set extensions. Although our encoding convention required a tighter encoding of the core RISC-V ISA, this has several beneficial effects. An implementation of the standard IMAFD ISA need only hold the most-significant 30 bits in instruction caches (a 6.25% saving). On instruction cache refills, any instructions encountered with either low bit clear should be recoded into illegal 30-bit instructions before storing in the cache to preserve illegal-instruction exception behavior. Perhaps more importantly, by condensing our base ISA into a subset of the 32-bit instruction word, we leave more space available for non-standard and custom extensions. In particular, the base RV32I ISA uses less than 1/8 of the encoding space in the 32-bit instruction word. An implementation that does not require support for the standard compressed instruction extension can map 3 additional non-conforming 30-bit instruction spaces into the 32-bit fixed-width format, while preserving support for standard ≥32-bit instruction-set extensions. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Encodings with bits \[15:0\] all zeros are defined as illegal instructions. These instructions are considered to be of minimal length: 16 bits if any 16-bit instruction-set extension is present, otherwise 32 bits. The encoding with bits \[ILEN-1:0\] all ones is also illegal; this instruction is considered to be ILEN bits long. | | We consider it a feature that any length of instruction containing all zero bits is not legal, as this quickly traps erroneous jumps into zeroed memory regions. Similarly, we also reserve the instruction encoding containing all ones to be an illegal instruction, to catch the other common pattern observed with unprogrammed non-volatile memory devices, disconnected memory buses, or broken memory devices. Software can rely on a naturally aligned 32-bit word containing zero to act as an illegal instruction on all RISC-V implementations, to be used by software where an illegal instruction is explicitly desired. Defining a corresponding known illegal value for all ones is more difficult due to the variable-length encoding. Software cannot generally use the illegal value of ILEN bits of all 1s, as software might not know ILEN for the eventual target machine (e.g., if software is compiled into a standard binary library used by many different machines). Defining a 32-bit word of all ones as illegal was also considered, as all machines must support a 32-bit instruction size, but this requires the instruction-fetch unit on machines with ILEN >32 report an illegal-instruction exception rather than an access-fault exception when such an instruction borders a protection boundary, complicating variable-instruction-length fetch and decode. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | RISC-V base ISAs have either little-endian or big-endian memory systems, with the privileged architecture further defining bi-endian operation. Instructions are stored in memory as a sequence of 16-bit little-endian parcels, regardless of memory system endianness. Parcels forming one instruction are stored at increasing halfword addresses, with the lowest-addressed parcel holding the lowest-numbered bits in the instruction specification. | | We originally chose little-endian byte ordering for the RISC-V memory system because little-endian systems are currently dominant commercially (all x86 systems; iOS, Android, and Windows for ARM). A minor point is that we have also found little-endian memory systems to be more natural for hardware designers. However, certain application areas, such as IP networking, operate on big-endian data structures, and certain legacy code bases have been built assuming big-endian processors, so we have defined big-endian and bi-endian variants of RISC-V. We have to fix the order in which instruction parcels are stored in memory, independent of memory system endianness, to ensure that the length-encoding bits always appear first in halfword address order. This allows the length of a variable-length instruction to be quickly determined by an instruction-fetch unit by examining only the first few bits of the first 16-bit instruction parcel. We further make the instruction parcels themselves little-endian to decouple the instruction encoding from the memory system endianness altogether. This design benefits both software tooling and bi-endian hardware. Otherwise, for instance, a RISC-V assembler or disassembler would always need to know the intended active endianness, despite that in bi-endian systems, the endianness mode might change dynamically during execution. In contrast, by giving instructions a fixed endianness, it is sometimes possible for carefully written software to be endianness-agnostic even in binary form, much like position-independent code. The choice to have instructions be only little-endian does have consequences, however, for RISC-V software that encodes or decodes machine instructions. Big-endian JIT compilers, for example, must swap the byte order when storing to instruction memory. Once we had decided to fix on a little-endian instruction encoding, this naturally led to placing the length-encoding bits in the LSB positions of the instruction format to avoid breaking up opcode fields. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#trap-defn)1.1.6\. Exceptions, Traps, and Interrupts We use the term _exception_ to refer to an unusual condition occurring at run time associated with an instruction in the current RISC-V hart. We use the term _interrupt_ to refer to an external asynchronous event that may cause a RISC-V hart to experience an unexpected transfer of control. We use the term _trap_ to refer to the transfer of control to a trap handler caused by either an exception or an interrupt. The instruction descriptions in following chapters describe conditions that can raise an exception during execution. The general behavior of most RISC-V EEIs is that a trap to some handler occurs when an exception is signaled on an instruction (except for floating-point exceptions, which, in the standard floating-point extensions, do not cause traps). The manner in which interrupts are generated, routed to, and enabled by a hart depends on the EEI. | | Our use of "exception" and "trap" is compatible with that in the IEEE-754 floating-point standard. | | ----------------------------------------------------------------------------------------------------- | How traps are handled and made visible to software running on the hart depends on the enclosing execution environment. From the perspective of software running inside an execution environment, traps encountered by a hart at runtime can have four different effects: Contained Trap The trap is visible to, and handled by, software running inside the execution environment. For example, in an EEI providing both supervisor and user mode on harts, an ECALL by a user-mode hart will generally result in a transfer of control to a supervisor-mode handler running on the same hart. Similarly, in the same environment, when a hart is interrupted, an interrupt handler will be run in supervisor mode on the hart. Requested Trap The trap is a synchronous exception that is an explicit call to the execution environment requesting an action on behalf of software inside the execution environment. An example is a system call. In this case, execution may or may not resume on the hart after the requested action is taken by the execution environment. For example, a system call could remove the hart or cause an orderly termination of the entire execution environment. Invisible Trap The trap is handled transparently by the execution environment and execution resumes normally after the trap is handled. Examples include emulating missing instructions, handling non-resident page faults in a demand-paged virtual-memory system, or handling device interrupts for a different job in a multiprogrammed machine. In these cases, the software running inside the execution environment is not aware of the trap (we ignore timing effects in these definitions). Fatal Trap The trap represents a fatal failure and causes the execution environment to terminate execution. Examples include failing a virtual-memory page-protection check or allowing a watchdog timer to expire. Each EEI should define how execution is terminated and reported to an external environment. [Table 1](#trapcharacteristics) shows the characteristics of each kind of trap. __Table 1\. Characteristics of traps__ | Contained | Requested | Invisible | Fatal | | | ---------------------- | --------- | --------- | ----- | ---- | | Execution terminates | No | No1 | No | Yes | | Software is oblivious | No | No | Yes | Yes2 | | Handled by environment | No | Yes | Yes | Yes | 1 Termination may be requested 2 Imprecise fatal traps might be observable by software The EEI defines for each trap whether it is handled precisely, though the recommendation is to maintain preciseness where possible. Contained and requested traps can be observed to be imprecise by software inside the execution environment. Invisible traps, by definition, cannot be observed to be precise or imprecise by software running inside the execution environment. Fatal traps can be observed to be imprecise by software running inside the execution environment, if known-errorful instructions do not cause immediate termination. Because this document describes unprivileged instructions, traps are rarely mentioned. Architectural means to handle contained traps are defined in the privileged architecture manual, along with other features to support richer EEIs. Unprivileged instructions that are defined solely to cause requested traps are documented here. Invisible traps are, by their nature, out of scope for this document. Instruction encodings that are not defined here and not defined by some other means may cause a fatal trap. ### [](#1-1-7-unspecified-behaviors-and-values)1.1.7\. UNSPECIFIED Behaviors and Values The architecture fully describes what implementations must do and any constraints on what they may do. In cases where the architecture intentionally does not constrain implementations, the term UNSPECIFIED is explicitly used. The term UNSPECIFIED refers to a behavior or value that is intentionally unconstrained. The definition of these behaviors or values is open to extensions, platform standards, or implementations. Extensions, platform standards, or implementation documentation may provide normative content to further constrain cases that the base architecture defines as UNSPECIFIED. Like the base architecture, extensions should fully describe allowable behavior and values and use the term UNSPECIFIED for cases that are intentionally unconstrained. These cases may be constrained or defined by other extensions, platform standards, or implementations. 12.1. "M" Extension for Integer Multiplication and Division, Version 2.0 ==================== ## [](#mstandard)12.1\. "M" Extension for Integer Multiplication and Division, Version 2.0 This chapter describes the standard integer multiplication and division instruction extension, which is named `M` and contains instructions that multiply or divide values held in two integer registers. | | We separate integer multiply and divide out from the base to simplify low-end implementations, or for applications where integer multiply and divide operations are either infrequent or better handled in attached accelerators. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#mult-ops)12.1.1\. Multiplication Operations ![svg](_images/svg-8674a2181bfeefba02cea7452f88332419866f6b.svg) MUL performs an XLEN-bit×XLEN-bit multiplication of`rs1` by `rs2` and places the lower XLEN bits in the destination register. MULH, MULHU, and MULHSU perform the same multiplication but return the upper XLEN bits of the full 2×XLEN-bit product, for signed×signed, unsigned×unsigned, and `rs1`×unsigned `rs2` multiplication.If both the high and low bits of the same product are required, then the recommended code sequence is: `MULH[[S]U] rdh, rs1, rs2; MUL rdl, rs1, rs2` (source register specifiers must be in same order and `rdh` cannot be the same as `rs1` or `rs2`). Microarchitectures can then fuse these into a single multiply operation instead of performing two separate multiplies. | | MULHSU is used in multi-word signed multiplication to multiply the most-significant word of the multiplicand (which contains the sign bit) with the less-significant words of the multiplier (which are unsigned). | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | MULW is an RV64 instruction that multiplies the lower 32 bits of the source registers, placing the sign extension of the lower 32 bits of the result into the destination register. | | In RV64, MUL can be used to obtain the upper 32 bits of the 64-bit product, but signed arguments must be proper 32-bit signed values, whereas unsigned arguments must have their upper 32 bits clear. If the arguments are not known to be sign- or zero-extended, an alternative is to shift both arguments left by 32 bits, then use MULH\[\[S\]U\]. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#12-1-2-division-operations)12.1.2\. Division Operations ![svg](_images/svg-8f078043848d9bbe75c874eafffc75906ce2b7be.svg) DIV and DIVU perform an XLEN bits by XLEN bits signed and unsigned integer division of `rs1` by `rs2`, rounding towards zero. REM and REMU provide the remainder of the corresponding division operation. For REM, the sign of a nonzero result equals the sign of the dividend. | | For both signed and unsigned division, except in the case of overflow, it holds that dividend = divisor × quotient + remainder. | | ---------------------------------------------------------------------------------------------------------------------------------- | If both the quotient and remainder are required from the same division, the recommended code sequence is: `DIV[U] rdq, rs1, rs2; REM[U] rdr,` `rs1, rs2` (`rdq` cannot be the same as `rs1` or `rs2`). Microarchitectures can then fuse these into a single divide operation instead of performing two separate divides. DIVW and DIVUW are RV64 instructions that divide the lower 32 bits of`rs1` by the lower 32 bits of `rs2`, treating them as signed and unsigned integers, placing the 32-bit quotient in `rd`, sign-extended to 64 bits. REMW and REMUW are RV64 instructions that provide the corresponding signed and unsigned remainder operations. Both REMW and REMUW always sign-extend the 32-bit result to 64 bits, including on a divide by zero. The semantics for division by zero and division overflow are summarized in [Table 1](#divby0). The quotient of division by zero has all bits set, and the remainder of division by zero equals the dividend. Signed division overflow occurs only when the most-negative integer is divided by −1\. The quotient of a signed division with overflow is equal to the dividend, and the remainder is zero. Unsigned division overflow cannot occur. __Table 1\. Semantics for division by zero and division overflow. L is the width of the operation in bits: XLEN for DIV\[U\] and REM\[U\], or 32 for DIV\[U\]W and REM\[U\]W.__ | Condition | Dividend | Divisor | DIVU\[W\] | REMU\[W\] | DIV\[W\] | REM\[W\] | | -------------------------------------- | --------- | ------- | --------- | --------- | --------- | -------- | | Division by zeroOverflow (signed only) | _x_\-2L-1 | 0−1 | 2L\-1 \- | _x_ \- | −1 \-2L-1 | _x_ 0 | | | We considered raising exceptions on integer divide by zero, with these exceptions causing a trap in most execution environments. However, this would be the only arithmetic trap in the standard ISA (floating-point exceptions set flags and write default values, but do not cause traps) and would require language implementers to interact with the execution environment’s trap handlers for this case. Further, where language standards mandate that a divide-by-zero exception must cause an immediate control flow change, only a single branch instruction needs to be added to each divide operation, and this branch instruction can be inserted after the divide and should normally be very predictably not taken, adding little runtime overhead. The value of all bits set is returned for both unsigned and signed divide by zero to simplify the divider circuitry. The value of all 1s is both the natural value to return for unsigned divide, representing the largest unsigned number, and also the natural result for simple unsigned divider implementations. Signed division is often implemented using an unsigned division circuit and specifying the same overflow result simplifies the hardware. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#12-1-3-zmmul-extension-version-1-0)12.1.3\. `Zmmul` Extension, Version 1.0 The `Zmmul` extension implements the multiplication subset of the M extension. It adds all of the instructions defined in[Multiplication Operations](#mult-ops), namely: MUL, MULH, MULHU, MULHSU, and (for RV64 only) MULW. The encodings are identical to those of the corresponding M-extension instructions. `M` implies `Zmmul`. | | The Zmmul extension enables low-cost implementations that require multiplication operations but not division. For many microcontroller applications, division operations are too infrequent to justify the cost of divider hardware. By contrast, multiplication operations are more frequent, making the cost of multiplier hardware more justifiable. Simple FPGA soft cores particularly benefit from eliminating division but retaining multiplication, since many FPGAs provide hardwired multipliers but require dividers be implemented in soft logic. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | RVWMO Explanatory Material, Version 0.1 ==================== ## [](#rvwmo-explanatory-material-version-0-1)Appendix A: RVWMO Explanatory Material, Version 0.1 This section provides more explanation for RVWMO[RVWMO Memory Consistency Model](rvwmo.html), using more informal language and concrete examples. These are intended to clarify the meaning and intent of the axioms and preserved program order rules. This appendix should be treated as commentary; all normative material is provided in [RVWMO Memory Consistency Model](rvwmo.html) and in the rest of the main body of the ISA specification. All currently known discrepancies are listed in [Known Issues](#discrepancies). Any other discrepancies are unintentional. ### [](#whyrvwmo)Why RVWMO? Memory consistency models fall along a loose spectrum from weak to strong. Weak memory models allow more hardware implementation flexibility and deliver arguably better performance, performance per watt, power, scalability, and hardware verification overheads than strong models, at the expense of a more complex programming model. Strong models provide simpler programming models, but at the cost of imposing more restrictions on the kinds of (non-speculative) hardware optimizations that can be performed in the pipeline and in the memory system, and in turn imposing some cost in terms of power, area overhead, and verification burden. RISC-V has chosen the RVWMO memory model, a variant of release consistency. This places it in between the two extremes of the memory model spectrum. The RVWMO memory model enables architects to build simple implementations, aggressive implementations, implementations embedded deeply inside a much larger system and subject to complex memory system interactions, or any number of other possibilities, all while simultaneously being strong enough to support programming language memory models at high performance. To facilitate the porting of code from other architectures, some hardware implementations may choose to implement the Ztso extension, which provides stricter RVTSO ordering semantics by default. Code written for RVWMO is automatically and inherently compatible with RVTSO, but code written assuming RVTSO is not guaranteed to run correctly on RVWMO implementations. In fact, most RVWMO implementations will (and should) simply refuse to run RVTSO-only binaries. Each implementation must therefore choose whether to prioritize compatibility with RVTSO code (e.g., to facilitate porting from x86) or whether to instead prioritize compatibility with other RISC-V cores implementing RVWMO. Some fences and/or memory ordering annotations in code written for RVWMO may become redundant under RVTSO; the cost that the default of RVWMO imposes on Ztso implementations is the incremental overhead of fetching those fences (e.g., FENCE R,RW and FENCE RW,W) which become no-ops on that implementation. However, these fences must remain present in the code if compatibility with non-Ztso implementations is desired. ### [](#litmustests)Litmus Tests The explanations in this chapter make use of _litmus tests_, or small programs designed to test or highlight one particular aspect of a memory model. [Litmus sample](#litmus-sample) shows an example of a litmus test with two harts. As a convention for this figure and for all figures that follow in this chapter, we assume that `s0-s2` are pre-set to the same value in all harts and that `s0` holds the address labeled `x`, `s1` holds `y`, and `s2` holds `z`, where `x`, `y`, and `z`are disjoint memory locations aligned to 8 byte boundaries. All other registers and all referenced memory locations are presumed to be initialized to zero. Each figure shows the litmus test code on the left, and a visualization of one particular valid or invalid execution on the right. __Table 1\. A sample litmus test and one forbidden execution (a0=1).__ | Hart 0 Hart 1 ⋮ ⋮ li t1,1 li t4,4 (a) sw t1,0(s0) (e) sw t4,0(s0) ⋮ ⋮ li t2,2 (b) sw t2,0(s0) ⋮ ⋮ (c) lw a0,0(s0) ⋮ ⋮ li t3,3 li t5,5 (d) sw t3,0(s0) (f) sw t5,0(s0) ⋮ ⋮ | ![litmus sample](_images/graphviz/litmus_sample.png) | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------- | Litmus tests are used to understand the implications of the memory model in specific concrete situations. For example, in the litmus test of[Litmus sample](#litmus-sample), the final value of `a0`in the first hart can be either 2, 4, or 5, depending on the dynamic interleaving of the instruction stream from each hart at runtime. However, in this example, the final value of `a0` in Hart 0 will never be 1 or 3; intuitively, the value 1 will no longer be visible at the time the load executes, and the value 3 will not yet be visible by the time the load executes. We analyze this test and many others below. __Table 2\. A key for the litmus test diagrams drawn in this appendix__ | Edge | Full Name (and explanation) | | ----- | ---------------------------------------------------------------------------------------------- | | rf | Reads From (from each store to the loads that return a value written by that store) | | co | Coherence (a total order on the stores to each address) | | fr | From-Reads (from each load to co-successors of the store from which the load returned a value) | | ppo | Preserved Program Order | | fence | Orderings enforced by a FENCE instruction | | addr | Address Dependency | | ctrl | Control Dependency | | data | Data Dependency | The diagram shown to the right of each litmus test shows a visual representation of the particular execution candidate being considered. These diagrams use a notation that is common in the memory model literature for constraining the set of possible global memory orders that could produce the execution in question. It is also the basis for the _herd_ models presented in [Formal Axiomatic Specification in Herd](mm-formal.html#sec:herd). This notation is explained in[Table 2](#litmus-key). Of the listed relations, rf edges between harts, co edges, fr edges, and ppo edges directly constrain the global memory order (as do fence, addr, data, and some ctrl edges, via ppo). Other edges (such as intra-hart rf edges) are informative but do not constrain the global memory order. For example, in [Litmus sample](#litmus-sample), `a0=1`could occur only if one of the following were true: * (b) appears before (a) in global memory order (and in the coherence order co). However, this violates RVWMO PPO rule `ppo:→st`. The co edge from (b) to (a) highlights this contradiction. * (a) appears before (b) in global memory order (and in the coherence order co). However, in this case, the Load Value Axiom would be violated, because (a) is not the latest matching store prior to (c) in program order. The fr edge from (c) to (b) highlights this contradiction. Since neither of these scenarios satisfies the RVWMO axioms, the outcome`a0=1` is forbidden. Beyond what is described in this appendix, a suite of more than seven thousand litmus tests is available at. | | The litmus tests repository also provides instructions on how to run the litmus tests on RISC-V hardware and how to compare the results with the operational and axiomatic models. In the future, we expect to adapt these memory model litmus tests for use as part of the RISC-V compliance test suite as well. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#explaining-the-rvwmo-rules)Explaining the RVWMO Rules In this section, we provide explanation and examples for all of the RVWMO rules and axioms. #### [](#preserved-program-order-and-global-memory-order)Preserved Program Order and Global Memory Order Preserved program order represents the subset of program order that must be respected within the global memory order. Conceptually, events from the same hart that are ordered by preserved program order must appear in that order from the perspective of other harts and/or observers. Events from the same hart that are not ordered by preserved program order, on the other hand, may appear reordered from the perspective of other harts and/or observers. Informally, the global memory order represents the order in which loads and stores perform. The formal memory model literature has moved away from specifications built around the concept of performing, but the idea is still useful for building up informal intuition. A load is said to have performed when its return value is determined. A store is said to have performed not when it has executed inside the pipeline, but rather only when its value has been propagated to globally visible memory. In this sense, the global memory order also represents the contribution of the coherence protocol and/or the rest of the memory system to interleave the (possibly reordered) memory accesses being issued by each hart into a single total order agreed upon by all harts. The order in which loads perform does not always directly correspond to the relative age of the values those two loads return. In particular, a load _b_ may perform before another load _a_ to the same address (i.e., _b_ may execute before_a_, and _b_ may appear before _a_in the global memory order), but _a_ may nevertheless return an older value than _b_. This discrepancy captures (among other things) the reordering effects of buffering placed between the core and memory. For example, _b_ may have returned a value from a store in the store buffer, while _a_ may have ignored that younger store and read an older value from memory instead. To account for this, at the time each load performs, the value it returns is determined by the load value axiom, not just strictly by determining the most recent store to the same address in the global memory order, as described below. #### [](#loadvalueaxiom)Load value axiom | | [Load Value Axiom](rvwmo.html#ax-load): Each byte of each load _i_ returns the value written to that byte by the store that is the latest in global memory order among the following stores: Stores that write that byte and that precede i in the global memory order Stores that write that byte and that precede i in program order | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Preserved program order is _not_ required to respect the ordering of a store followed by a load to an overlapping address. This complexity arises due to the ubiquity of store buffers in nearly all implementations. Informally, the load may perform (return a value) by forwarding from the store while the store is still in the store buffer, and hence before the store itself performs (writes back to globally visible memory). Any other hart will therefore observe the load as performing before the store. Consider the [Table 3](#litms%5Fsb%5Fforward). When running this program on an implementation with store buffers, it is possible to arrive at the final outcome `a0=1, a1=0, a2=1, a3=0` as follows: __Table 3\. A store buffer forwarding litmus test (outcome permitted)__ | Hart 0 Hart 1 li t1, 1 li t1, 1 (a) sw t1,0(s0) (e) sw t1,0(s1) (b) lw a0,0(s0) (f) lw a2,0(s1) (c) fence r,r (g) fence r,r (d) lw a1,0(s1) (h) lw a3,0(s0) Outcome: a0=1, a1=0, a2=1, a3=0 | ![litmus sb fwd](_images/graphviz/litmus_sb_fwd.png) | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------- | * (a) executes and enters the first hart’s private store buffer * (b) executes and forwards its return value 1 from (a) in the store buffer * (c) executes since all previous loads (i.e., (b)) have completed * (d) executes and reads the value 0 from memory * (e) executes and enters the second hart’s private store buffer * (f) executes and forwards its return value 1 from (e) in the store buffer * (g) executes since all previous loads (i.e., (f)) have completed * (h) executes and reads the value 0 from memory * (a) drains from the first hart’s store buffer to memory * (e) drains from the second hart’s store buffer to memory Therefore, the memory model must be able to account for this behavior. To put it another way, suppose the definition of preserved program order did include the following hypothetical rule: memory access_a_ precedes memory access _b_ in preserved program order (and hence also in the global memory order) if_a_ precedes _b_ in program order and_a_ and _b_ are accesses to the same memory location, _a_ is a write, and _b_ is a read. Call this "Rule X". Then we get the following: * (a) precedes (b): by rule X * (b) precedes (d): by rule [4](rvwmo.html#overlapping-ordering) * (d) precedes (e): by the load value axiom. Otherwise, if (e) preceded (d), then (d) would be required to return the value 1\. (This is a perfectly legal execution; it’s just not the one in question) * (e) precedes (f): by rule X * (f) precedes (h): by rule [4](rvwmo.html#overlapping-ordering) * (h) precedes (a): by the load value axiom, as above. The global memory order must be a total order and cannot be cyclic, because a cycle would imply that every event in the cycle happens before itself, which is impossible. Therefore, the execution proposed above would be forbidden, and hence the addition of rule X would forbid implementations with store buffer forwarding, which would clearly be undesirable. Nevertheless, even if (b) precedes (a) and/or (f) precedes (e) in the global memory order, the only sensible possibility in this example is for (b) to return the value written by (a), and likewise for (f) and (e). This combination of circumstances is what leads to the second option in the definition of the load value axiom. Even though (b) precedes (a) in the global memory order, (a) will still be visible to (b) by virtue of sitting in the store buffer at the time (b) executes. Therefore, even if (b) precedes (a) in the global memory order, (b) should return the value written by (a) because (a) precedes (b) in program order. Likewise for (e) and (f). __Table 4\. The "PPOCA" store buffer forwarding litmus test (outcome permitted)__ | Hart 0 Hart 1 li t1, 1 li t1, 1 (a) sw t1,0(s0) LOOP: (b) fence w,w (d) lw a0,0(s1) (c) sw t1,0(s1) beqz a0, LOOP (e) sw t1,0(s2) (f) lw a1,0(s2) xor a2,a1,a1 add s0,s0,a2 (g) lw a2,0(s0) Outcome: a0=1, a1=1, a2=0 | ![litmus ppoca](_images/graphviz/litmus_ppoca.png) | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------- | Another test that highlights the behavior of store buffers is shown in[Table 4](#litmus%5Fppoca). In this example, (d) is ordered before (e) because of the control dependency, and (f) is ordered before (g) because of the address dependency. However, (e) is _not_necessarily ordered before (f), even though (f) returns the value written by (e). This could correspond to the following sequence of events: * (e) executes speculatively and enters the second hart’s private store buffer (but does not drain to memory) * (f) executes speculatively and forwards its return value 1 from (e) in the store buffer * (g) executes speculatively and reads the value 0 from memory * (a) executes, enters the first hart’s private store buffer, and drains to memory * (b) executes and retires * (c) executes, enters the first hart’s private store buffer, and drains to memory * (d) executes and reads the value 1 from memory * (e), (f), and (g) commit, since the speculation turned out to be correct * (e) drains from the store buffer to memory #### [](#atomicityaxiom)Atomicity axiom | | [Atomicity Axiom](rvwmo.html#ax-atom) (for Aligned Atomics): If r and w are paired load and store operations generated by aligned LR and SC instructions in a hart h, s is a store to byte x, and r returns a value written by s, then s must precede w in the global memory order, and there can be no store from a hart other than h to byte x following s and preceding w in the global memory order. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The RISC-V architecture decouples the notion of atomicity from the notion of ordering. Unlike architectures such as TSO, RISC-V atomics under RVWMO do not impose any ordering requirements by default. Ordering semantics are only guaranteed by the PPO rules that otherwise apply. RISC-V contains two types of atomics: AMOs and LR/SC pairs. These conceptually behave differently, in the following way. LR/SC behave as if the old value is brought up to the core, modified, and written back to memory, all while a reservation is held on that memory location. AMOs on the other hand conceptually behave as if they are performed directly in memory. AMOs are therefore inherently atomic, while LR/SC pairs are atomic in the slightly different sense that the memory location in question will not be modified by another hart during the time the original hart holds the reservation. __Table 5\. In all four (independent) instances, the final store-conditional instruction is permitted but not guaranteed to succeed.__ | (a) lr.d a0, 0(s0) | (a) lr.d a0, 0(s0) | (a) lr.w a0, 0(s0) | (a) lr.w a0, 0(s0) | | ---------------------- | ---------------------- | ---------------------- | ------------------ | | (b) sd t1, 0(s0) | (b) sw t1, 4(s0) | (b) sw t1, 4(s0) | (b) sw t1, 4(s0) | | (c) sc.d t3, t2, 0(s0) | (c) sc.d t3, t2, 0(s0) | (c) sc.w t3, t2, 0(s0) | (c) addi s0, s0, 8 | | (d) sc.w t3, t2, 0(s0) | | | | The atomicity axiom forbids stores from other harts from being interleaved in global memory order between an LR and the SC paired with that LR. The atomicity axiom does not forbid loads from being interleaved between the paired operations in program order or in the global memory order, nor does it forbid stores from the same hart or stores to non-overlapping locations from appearing between the paired operations in either program order or in the global memory order. For example, the SC instructions in [Table 5](#litmus%5Flrsdsc) may (but are not guaranteed to) succeed. None of those successes would violate the atomicity axiom, because the intervening non-conditional stores are from the same hart as the paired load-reserved and store-conditional instructions. This way, a memory system that tracks memory accesses at cache line granularity (and which therefore will see the four snippets of [Table 5](#litmus%5Flrsdsc) as identical) will not be forced to fail a store-conditional instruction that happens to (falsely) share another portion of the same cache line as the memory location being held by the reservation. The atomicity axiom also technically supports cases in which the LR and SC touch different addresses and/or use different access sizes; however, use cases for such behaviors are expected to be rare in practice. Likewise, scenarios in which stores from the same hart between an LR/SC pair actually overlap the memory location(s) referenced by the LR or SC are expected to be rare compared to scenarios where the intervening store may simply fall onto the same cache line. #### [](#mm-progress)Progress axiom | | [Progress Axiom](rvwmo.html#ax-prog): No memory operation may be preceded in the global memory order by an infinite sequence of other memory operations. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | The progress axiom ensures a minimal forward progress guarantee. It ensures that stores from one hart will eventually be made visible to other harts in the system in a finite amount of time, and that loads from other harts will eventually be able to read those values (or successors thereof). Without this rule, it would be legal, for example, for a spinlock to spin infinitely on a value, even with a store from another hart unlocking the spinlock. The progress axiom is intended not to impose any other notion of fairness, latency, or quality of service onto the harts in a RISC-V implementation. Any stronger notions of fairness are up to the rest of the ISA and/or up to the platform and/or device to define and implement. The forward progress axiom will in almost all cases be naturally satisfied by any standard cache coherence protocol. Implementations with non-coherent caches may have to provide some other mechanism to ensure the eventual visibility of all stores (or successors thereof) to all harts. #### [](#mm-overlap)Overlapping-Address Orderings ([Rules 1-3](rvwmo.html#overlapping-ordering)) | | [Rule 1](rvwmo.html#overlapping-ordering): b is a store, and a and b access overlapping memory addresses [Rule 2](rvwmo.html#overlapping-ordering): a and b are loads, x is a byte read by both a and b, there is no store to x between a and b in program order, and a and b return values for x written by different memory operations [Rule 3](rvwmo.html#overlapping-ordering): a is generated by an AMO or SC instruction, b is a load, and b returns a value written by a | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Same-address orderings where the latter is a store are straightforward: a load or store can never be reordered with a later store to an overlapping memory location. From a microarchitecture perspective, generally speaking, it is difficult or impossible to undo a speculatively reordered store if the speculation turns out to be invalid, so such behavior is simply disallowed by the model. Same-address orderings from a store to a later load, on the other hand, do not need to be enforced. As discussed in[Load value axiom](#loadvalueaxiom), this reflects the observable behavior of implementations that forward values from buffered stores to later loads. Same-address load-load ordering requirements are far more subtle. The basic requirement is that a younger load must not return a value that is older than a value returned by an older load in the same hart to the same address. This is often known as "CoRR" (Coherence for Read-Read pairs), or as part of a broader "coherence" or "sequential consistency per location" requirement. Some architectures in the past have relaxed same-address load-load ordering, but in hindsight this is generally considered to complicate the programming model too much, and so RVWMO requires CoRR ordering to be enforced. However, because the global memory order corresponds to the order in which loads perform rather than the ordering of the values being returned, capturing CoRR requirements in terms of the global memory order requires a bit of indirection. __Table 6\. Litmus test MP+fence.w.w+fre-rfi-addr (outcome permitted)__ | Hart 0 Hart 1 li t1, 1 li t2, 2 (a) sw t1,0(s0) (d) lw a0,0(s1) (b) fence w, w (e) sw t2,0(s1) (c) sw t1,0(s1) (f) lw a1,0(s1) (g) xor t3,a1,a1 (h) add s0,s0,t3 (i) lw a2,0(s0) Outcome: a0=1, a1=2, a2=0 | ![litmus mp fenceww fri rfi addr](_images/graphviz/litmus_mp_fenceww_fri_rfi_addr.png) | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------- | Consider the litmus test of [Table 6](#frirfi), which is one particular instance of the more general "fri-rfi" pattern. The term "fri-rfi" refers to the sequence (d), (e), (f): (d) "from-reads" (i.e., reads from an earlier write than) (e) which is the same hart, and (f) reads from (e) which is in the same hart. From a microarchitectural perspective, outcome `a0=1`, `a1=2`, `a2=0` is legal (as are various other less subtle outcomes). Intuitively, the following would produce the outcome in question: * (d) stalls (for whatever reason; perhaps it’s stalled waiting for some other preceding instruction) * (e) executes and enters the store buffer (but does not yet drain to memory) * (f) executes and forwards from (e) in the store buffer * (g), (h), and (i) execute * (a) executes and drains to memory, (b) executes, and (c) executes and drains to memory * (d) unstalls and executes * (e) drains from the store buffer to memory This corresponds to a global memory order of (f), (i), (a), (c), (d), (e). Note that even though (f) performs before (d), the value returned by (f) is newer than the value returned by (d). Therefore, this execution is legal and does not violate the CoRR requirements. Likewise, if two back-to-back loads return the values written by the same store, then they may also appear out-of-order in the global memory order without violating CoRR. Note that this is not the same as saying that the two loads return the same value, since two different stores may write the same value. __Table 7\. Litmus test RSW (outcome permitted)__ | Hart 0 Hart 1 li t1, 1 (d) lw a0,0(s1) (a) sw t1,0(s0) (e) xor t2,a0,a0 (b) fence w, w (f) add s4,s2,t2 (c) sw t1,0(s1) (g) lw a1,0(s4) (h) lw a2,0(s2) (i) xor t3,a2,a2 (j) add s0,s0,t3 (k) lw a3,0(s0) Outcome: a0=1, a1=v, a2=v, a3=0 | ![litmus rsw](_images/graphviz/litmus_rsw.png) | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------- | Consider the litmus test of [Table 7](#litmus-rsw). The outcome `a0=1`, `a1=v`, `a2=v`, `a3=0` (where _v_ is some value written by another hart) can be observed by allowing (g) and (h) to be reordered. This might be done speculatively, and the speculation can be justified by the microarchitecture (e.g., by snooping for cache invalidations and finding none) because replaying (h) after (g) would return the value written by the same store anyway. Hence assuming `a1` and `a2` would end up with the same value written by the same store anyway, (g) and (h) can be legally reordered. The global memory order corresponding to this execution would be (h),(k),(a),(c),(d),(g). Executions of the test in [Table 7](#litmus-rsw) in which `a1` does not equal `a2` do in fact require that (g) appears before (h) in the global memory order. Allowing (h) to appear before (g) in the global memory order would in that case result in a violation of CoRR, because then (h) would return an older value than that returned by (g). Therefore, [rule 2](rvwmo.html#overlapping-ordering) forbids this CoRR violation from occurring. As such, [rule 2](rvwmo.html#overlapping-ordering) strikes a careful balance between enforcing CoRR in all cases while simultaneously being weak enough to permit "RSW" and "fri-rfi" patterns that commonly appear in real microarchitectures. There is one more overlapping-address rule: [rule 3](rvwmo.html#overlapping-ordering) simply states that a value cannot be returned from an AMO or SC to a subsequent load until the AMO or SC has (in the case of the SC, successfully) performed globally. This follows somewhat naturally from the conceptual view that both AMOs and SC instructions are meant to be performed atomically in memory. However, notably, [rule 3](rvwmo.html#overlapping-ordering) states that hardware may not even non-speculatively forward the value being stored by an AMOSWAP to a subsequent load, even though for AMOSWAP that store value is not actually semantically dependent on the previous value in memory, as is the case for the other AMOs. The same holds true even when forwarding from SC store values that are not semantically dependent on the value returned by the paired LR. The three PPO rules above also apply when the memory accesses in question only overlap partially. This can occur, for example, when accesses of different sizes are used to access the same object. Note also that the base addresses of two overlapping memory operations need not necessarily be the same for two memory accesses to overlap. When misaligned memory accesses are being used, the overlapping-address PPO rules apply to each of the component memory accesses independently. #### [](#mm-fence)Fences ([Rule 4](rvwmo.html#overlapping-ordering)) | | Rule [4](rvwmo.html#overlapping-ordering): There is a FENCE instruction that orders a before b | | ------------------------------------------------------------------------------------------------- | By default, the FENCE instruction ensures that all memory accesses from instructions preceding the fence in program order (the "predecessor set") appear earlier in the global memory order than memory accesses from instructions appearing after the fence in program order (the "successor set"). However, fences can optionally further restrict the predecessor set and/or the successor set to a smaller set of memory accesses in order to provide some speedup. Specifically, fences have PR, PW, SR, and SW bits which restrict the predecessor and/or successor sets. The predecessor set includes loads (resp.stores) if and only if PR (resp.PW) is set. Similarly, the successor set includes loads (resp.stores) if and only if SR (resp.SW) is set. The FENCE encoding currently has nine non-trivial combinations of the four bits PR, PW, SR, and SW, plus one extra encoding FENCE.TSO which facilitates mapping of "acquire+release" or RVTSO semantics. The remaining seven combinations have empty predecessor and/or successor sets and hence are no-ops. Of the ten non-trivial options, six are commonly used in practice: * FENCE RW,RW * FENCE.TSO * FENCE RW,W * FENCE R,RW * FENCE R,R * FENCE W,W FENCE instructions using other combinations of PR, PW, SR, and SW are not normally used in the Linux or C++ memory models but are otherwise well defined. Finally, we note that since RISC-V uses a multi-copy atomic memory model, programmers can reason about fences bits in a thread-local manner. Fences in RISC-V are not cumulative, as they are in some non-multi-copy-atomic memory models. #### [](#sec:memory:acqrel)Explicit Synchronization ([Rules 5-8](rvwmo.html#overlapping-ordering)) | | [Rule 5](rvwmo.html#overlapping-ordering): a has an acquire annotation [Rule 6](rvwmo.html#overlapping-ordering): b has a release annotation [Rule 7](rvwmo.html#overlapping-ordering): a and b both have RCsc annotations [Rule 8](rvwmo.html#overlapping-ordering): a is paired with b | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | An _acquire_ operation, as would be used at the start of a critical section, requires all memory operations following the acquire in program order to also follow the acquire in the global memory order. This ensures, for example, that all loads and stores inside the critical section are up to date with respect to the synchronization variable being used to protect it. Acquire ordering can be enforced in one of two ways: with an acquire annotation, which enforces ordering with respect to just the synchronization variable itself, or with a FENCE R,RW, which enforces ordering with respect to all previous loads. A spinlock with atomics ```asm sd x1, (a1) # Arbitrary unrelated store ld x2, (a2) # Arbitrary unrelated load li t0, 1 # Initialize swap value. again: amoswap.w.aq t0, t0, (a0) # Attempt to acquire lock. bnez t0, again # Retry if held. # ... # Critical section. # ... amoswap.w.rl x0, x0, (a0) # Release lock by storing 0. sd x3, (a3) # Arbitrary unrelated store ld x4, (a4) # Arbitrary unrelated load ``` Consider [Example 1](#spinlock%5Fatomics). Because this example uses _aq_, the loads and stores in the critical section are guaranteed to appear in the global memory order after the AMOSWAP used to acquire the lock. However, assuming `a0`, `a1`, and `a2`point to different memory locations, the loads and stores in the critical section may or may not appear after the "Arbitrary unrelated load" at the beginning of the example in the global memory order. A spinlock with fences ```asm sd x1, (a1) # Arbitrary unrelated store ld x2, (a2) # Arbitrary unrelated load li t0, 1 # Initialize swap value. again: amoswap.w t0, t0, (a0) # Attempt to acquire lock. fence r, rw # Enforce "acquire" memory ordering bnez t0, again # Retry if held. # ... # Critical section. # ... fence rw, w # Enforce "release" memory ordering amoswap.w x0, x0, (a0) # Release lock by storing 0. sd x3, (a3) # Arbitrary unrelated store ld x4, (a4) # Arbitrary unrelated load ``` Now, consider the alternative in [Example 2](#spinlock%5Ffences). In this case, even though the AMOSWAP does not enforce ordering with an_aq_ bit, the fence nevertheless enforces that the acquire AMOSWAP appears earlier in the global memory order than all loads and stores in the critical section. Note, however, that in this case, the fence also enforces additional orderings: it also requires that the "Arbitrary unrelated load" at the start of the program appears earlier in the global memory order than the loads and stores of the critical section. (This particular fence does not, however, enforce any ordering with respect to the "Arbitrary unrelated store" at the start of the snippet.) In this way, fence-enforced orderings are slightly coarser than orderings enforced by _.aq_. Release orderings work exactly the same as acquire orderings, just in the opposite direction. Release semantics require all loads and stores preceding the release operation in program order to also precede the release operation in the global memory order. This ensures, for example, that memory accesses in a critical section appear before the lock-releasing store in the global memory order. Just as for acquire semantics, release semantics can be enforced using release annotations or with a FENCE RW,W operation. Using the same examples, the ordering between the loads and stores in the critical section and the "Arbitrary unrelated store" at the end of the code snippet is enforced only by the FENCE RW,W in [Example 2](#spinlock%5Ffences), not by the _rl_ in [Example 1](#spinlock%5Fatomics). With RCpc annotations alone, store-release-to-load-acquire ordering is not enforced. This facilitates the porting of code written under the TSO and/or RCpc memory models. To enforce store-release-to-load-acquire ordering, the code must use store-release-RCsc and load-acquire-RCsc operations so that PPO rule 7 applies. RCpc alone is sufficient for many use cases in C/C++ but is insufficient for many other use cases in C/C++, Java, and Linux, to name just a few examples; see [Memory Porting](#memory%5Fporting) for details. PPO rule 8 indicates that an SC must appear after its paired LR in the global memory order. This will follow naturally from the common use of LR/SC to perform an atomic read-modify-write operation due to the inherent data dependency. However, PPO rule 8 also applies even when the value being stored does not syntactically depend on the value returned by the paired LR. Lastly, we note that, as with fences, ordering annotations are not cumulative. #### [](#sec:memory:dependencies)Syntactic Dependencies ([Rules 9-11](rvwmo.html#overlapping-ordering)) | | [Rule 9](rvwmo.html#overlapping-ordering): b has a syntactic address dependency on a [Rule 10](rvwmo.html#overlapping-ordering): b has a syntactic data dependency on a [Rule 11](rvwmo.html#overlapping-ordering): b is a store, and b has a syntactic control dependency on a | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Dependencies from a load to a later memory operation in the same hart are respected by the RVWMO memory model. The Alpha memory model was notable for choosing _not_ to enforce the ordering of such dependencies, but most modern hardware and software memory models consider allowing dependent instructions to be reordered too confusing and counterintuitive. Furthermore, modern code sometimes intentionally uses such dependencies as a particularly lightweight ordering enforcement mechanism. The terms in [Syntactic Dependencies](rvwmo.html#mem-dependencies) work as follows. Instructions are said to carry dependencies from their source register(s) to their destination register(s) whenever the value written into each destination register is a function of the source register(s). For most instructions, this means that the destination register(s) carry a dependency from all source register(s). However, there are a few notable exceptions. In the case of memory instructions, the value written into the destination register ultimately comes from the memory system rather than from the source register(s) directly, and so this breaks the chain of dependencies carried from the source register(s). In the case of unconditional jumps, the value written into the destination register comes from the current `pc` (which is never considered a source register by the memory model), and so likewise, JALR (the only jump with a source register) does not carry a dependency from_rs1_ to _rd_. (c) has a syntactic dependency on both (a) and (b) via fflags, a destination register that both (a) and (b) implicitly accumulate into ```source%linenums (a) fadd f3,f1,f2 (b) fadd f6,f4,f5 (c) csrrs a0,fflags,x0 ``` The notion of accumulating into a destination register rather than writing into it reflects the behavior of CSRs such as `fflags`. In particular, an accumulation into a register does not clobber any previous writes or accumulations into the same register. For example, in[(c) has a syntactic dependency on both (a) and (b) via fflags, a destination register that both (a) and (b) implicitly accumulate into](#fflags), (c) has a syntactic dependency on both (a) and (b). Like other modern memory models, the RVWMO memory model uses syntactic rather than semantic dependencies. In other words, this definition depends on the identities of the registers being accessed by different instructions, not the actual contents of those registers. This means that an address, control, or data dependency must be enforced even if the calculation could seemingly be `optimized away`. This choice ensures that RVWMO remains compatible with code that uses these false syntactic dependencies as a lightweight ordering mechanism. A syntactic address dependency ```source%linenums ld a1,0(s0) xor a2,a1,a1 add s1,s1,a2 ld a5,0(s1) ``` For example, there is a syntactic address dependency from the memory operation generated by the first instruction to the memory operation generated by the last instruction in[A syntactic address dependency](#address), even though `a1` XOR`a1` is zero and hence has no effect on the address accessed by the second load. The benefit of using dependencies as a lightweight synchronization mechanism is that the ordering enforcement requirement is limited only to the specific two instructions in question. Other non-dependent instructions may be freely reordered by aggressive implementations. One alternative would be to use a load-acquire, but this would enforce ordering for the first load with respect to _all_ subsequent instructions. Another would be to use a FENCE R,R, but this would include all previous and all subsequent loads, making this option more expensive. A syntactic control dependency ```source%linenums lw x1,0(x2) bne x1,x0,next sw x3,0(x4) next: sw x5,0(x6) ``` Control dependencies behave differently from address and data dependencies in the sense that a control dependency always extends to all instructions following the original target in program order. Consider [A syntactic control dependency](#control1) the instruction at `next` will always execute, but the memory operation generated by that last instruction nevertheless still has a control dependency from the memory operation generated by the first instruction. Another syntactic control dependency ```source%linenums lw x1,0(x2) bne x1,x0,next next: sw x3,0(x4) ``` Likewise, consider [Another syntactic control dependency](#control2). Even though both branch outcomes have the same target, there is still a control dependency from the memory operation generated by the first instruction in this snippet to the memory operation generated by the last instruction. This definition of control dependency is subtly stronger than what might be seen in other contexts (e.g., C++), but it conforms with standard definitions of control dependencies in the literature. Notably, PPO rules [9-11](rvwmo.html#overlapping-ordering) are also intentionally designed to respect dependencies that originate from the output of a successful store-conditional instruction. Typically, an SC instruction will be followed by a conditional branch checking whether the outcome was successful; this implies that there will be a control dependency from the store operation generated by the SC instruction to any memory operations following the branch. PPO rule [11](rvwmo.html#ppo) in turn implies that any subsequent store operations will appear later in the global memory order than the store operation generated by the SC. However, since control, address, and data dependencies are defined over memory operations, and since an unsuccessful SC does not generate a memory operation, no order is enforced between unsuccessful SC and its dependent instructions. Moreover, since SC is defined to carry dependencies from its source registers to _rd_ only when the SC is successful, an unsuccessful SC has no effect on the global memory order. __Table 8\. A variant of the LB litmus test (outcome forbidden)__ | Initial values: 0(s0)=1; 0(s2)=1 Hart 0 Hart 1 (a) ld a0,0(s0) (e) ld a3,0(s2) (b) lr a1,0(s1) (f) sd a3,0(s0) (c) sc a2,a0,0(s1) (d) sd a2,0(s2) Outcome: a0=0, a3=0 | ![litmus lb lrsc](_images/graphviz/litmus_lb_lrsc.png) | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------ | In addition, the choice to respect dependencies originating at store-conditional instructions ensures that certain out-of-thin-air-like behaviors will be prevented. Consider[Table 8](#litmus%5Flb%5Flrsc). Suppose a hypothetical implementation could occasionally make some early guarantee that a store-conditional operation will succeed. In this case, (c) could return 0 to `a2` early (before actually executing), allowing the sequence (d), (e), (f), (a), and then (b) to execute, and then (c) might execute (successfully) only at that point. This would imply that (c) writes its own success value to `0(s1)`! Fortunately, this situation and others like it are prevented by the fact that RVWMO respects dependencies originating at the stores generated by successful SC instructions. We also note that syntactic dependencies between instructions only have any force when they take the form of a syntactic address, control, and/or data dependency. For example: a syntactic dependency between two`F` instructions via one of the `accumulating CSRs` in[Source and Destination Register Listings](rvwmo.html#source-dest-regs) does _not_ imply that the two `F` instructions must be executed in order. Such a dependency would only serve to ultimately set up later a dependency from both `F` instructions to a later CSR instruction accessing the CSR flag in question. #### [](#memory-ppopipeline)Pipeline Dependencies ([Rules 12-13](rvwmo.html#overlapping-ordering)) | | [Rule 12](rvwmo.html#overlapping-ordering): b is a load, and there exists some store m between a and b in program order such that m has an address or data dependency on a, and b returns a value written by m [Rule 13](rvwmo.html#overlapping-ordering): b is a store, and there exists some instruction m between a and b in program order such that m has an address dependency on a | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 9\. Because of PPO [rule 12](rvwmo.html#overlapping-ordering) and the data dependency from (d) to (e), (d) must also precede (f) in the global memory order (outcome forbidden)__ | Hart 0 Hart 1 li t1, 1 (d) lw a0, 0(s1) (a) sw t1,0(s0) (e) sw a0, 0(s2) (b) fence w, w (f) lw a1, 0(s2) (c) sw t1,0(s1) xor a2,a1,a1 add s0,s0,a2 (g) lw a3,0(s0) Outcome: a0=1, a3=0 | ![litmus datarfi](_images/graphviz/litmus_datarfi.png) | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------ | PPO rules [12](rvwmo.html#overlapping-ordering) and [13](rvwmo.html#overlapping-ordering) reflect behaviors of almost all real processor pipeline implementations. Rule [12](rvwmo.html#overlapping-ordering)states that a load cannot forward from a store until the address and data for that store are known. Consider [Table 9](#litmus%5Fdatarfi) (f) cannot be executed until the data for (e) has been resolved, because (f) must return the value written by (e) (or by something even later in the global memory order), and the old value must not be clobbered by the write-back of (e) before (d) has had a chance to perform. Therefore, (f) will never perform before (d) has performed. __Table 10\. Because of the extra store between (e) and (g), (d) no longer necessarily precedes (g) (outcome permitted)__ | Hart 0 Hart 1 li t1, 1 li t1, 1 (a) sw t1,0(s0) (d) lw a0, 0(s1) (b) fence w, w (e) sw a0, 0(s2) (c) sw t1,0(s1) (f) sw t1, 0(s2) (g) lw a1, 0(s2) xor a2,a1,a1 add s0,s0,a2 (h) lw a3,0(s0) Outcome: a0=1, a3=0 | ![litmus datacoirfi](_images/graphviz/litmus_datacoirfi.png) | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------ | If there were another store to the same address in between (e) and (f), as in [Table 11](#litmus:addrdatarfi%5Fno), then (f) would no longer be dependent on the data of (e) being resolved, and hence the dependency of (f) on (d), which produces the data for (e), would be broken. Rule [13](rvwmo.html#overlapping-ordering) makes a similar observation to the previous rule: a store cannot be performed at memory until all previous loads that might access the same address have themselves been performed. Such a load must appear to execute before the store, but it cannot do so if the store were to overwrite the value in memory before the load had a chance to read the old value. Likewise, a store generally cannot be performed until it is known that preceding instructions will not cause an exception due to failed address resolution, and in this sense, rule 13 can be seen as somewhat of a special case of rule [11](rvwmo.html#overlapping-ordering). __Table 11\. Because of the address dependency from (d) to (e), (d) also precedes (f) (outcome forbidden)__ | Hart 0 Hart 1 li t1, 1 (a) lw a0,0(s0) (d) lw a1, 0(s1) (b) fence rw,rw (e) lw a2, 0(a1) (c) sw s2,0(s1) (f) sw t1, 0(s0) Outcome: a0=1, a1=t | ![litmus addrpo](_images/graphviz/litmus_addrpo.png) | | --------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------- | Consider [Table 11](#litmus:addrdatarfi%5Fno) (f) cannot be executed until the address for (e) is resolved, because it may turn out that the addresses match; i.e., that `a1=s0`. Therefore, (f) cannot be sent to memory before (d) has executed and confirmed whether the addresses do indeed overlap. ### [](#beyond-main-memory)Beyond Main Memory RVWMO does not currently attempt to formally describe how FENCE.I, SFENCE.VMA, I/O fences, and PMAs behave. All of these behaviors will be described by future formalizations. In the meantime, the behavior of FENCE.I is described in ["Zifencei" Extension for Instruction-Fetch Fence](zifencei.html), the behavior of SFENCE.VMA is described in the RISC-V Instruction Set Privileged Architecture Manual, and the behavior of I/O fences and the effects of PMAs are described below. #### [](#coherence-and-cacheability)Coherence and Cacheability The RISC-V Privileged ISA defines Physical Memory Attributes (PMAs) which specify, among other things, whether portions of the address space are coherent and/or cacheable. See the RISC-V Privileged ISA Specification for the complete details. Here, we simply discuss how the various details in each PMA relate to the memory model: * Main memory vs.I/O, and I/O memory ordering PMAs: the memory model as defined applies to main memory regions. I/O ordering is discussed below. * Supported access types and atomicity PMAs: the memory model is simply applied on top of whatever primitives each region supports. * Cacheability PMAs: the cacheability PMAs in general do not affect the memory model. Non-cacheable regions may have more restrictive behavior than cacheable regions, but the set of allowed behaviors does not change regardless. However, some platform-specific and/or device-specific cacheability settings may differ. * Coherence PMAs: The memory consistency model for memory regions marked as non-coherent in PMAs is currently platform-specific and/or device-specific: the load-value axiom, the atomicity axiom, and the progress axiom all may be violated with non-coherent memory. Note however that coherent memory does not require a hardware cache coherence protocol. The RISC-V Privileged ISA Specification suggests that hardware-incoherent regions of main memory are discouraged, but the memory model is compatible with hardware coherence, software coherence, implicit coherence due to read-only memory, implicit coherence due to only one agent having access, or otherwise. * Idempotency PMAs: Idempotency PMAs are used to specify memory regions for which loads and/or stores may have side effects, and this in turn is used by the microarchitecture to determine, e.g., whether prefetches are legal. This distinction does not affect the memory model. #### [](#io-ordering)I/O Ordering For I/O, the load value axiom and atomicity axiom in general do not apply, as both reads and writes might have device-specific side effects and may return values other than the value "written" by the most recent store to the same address. Nevertheless, the following preserved program order rules still generally apply for accesses to I/O memory: memory access _a_ precedes memory access _b_ in global memory order if _a_ precedes _b_ in program order and one or more of the following holds: 1. _a_ precedes _b_ in preserved program order as defined in [RVWMO Memory Consistency Model](rvwmo.html), with the exception that acquire and release ordering annotations apply only from one memory operation to another memory operation and from one I/O operation to another I/O operation, but not from a memory operation to an I/O nor vice versa 2. _a_ and _b_ are accesses to overlapping addresses in an I/O region 3. _a_ and _b_ are accesses to the same strongly ordered I/O region 4. _a_ and _b_ are accesses to I/O regions, and the channel associated with the I/O region accessed by either_a_ or _b_ is channel 1 5. _a_ and _b_ are accesses to I/O regions associated with the same channel (except for channel 0) Note that the FENCE instruction distinguishes between main memory operations and I/O operations in its predecessor and successor sets. To enforce ordering between I/O operations and main memory operations, code must use a FENCE with PI, PO, SI, and/or SO, plus PR, PW, SR, and/or SW. For example, to enforce ordering between a write to main memory and an I/O write to a device register, a FENCE W,O or stronger is needed. Ordering memory and I/O accesses ```source%linenums sd t0, 0(a0) fence w,o sd a0, 0(a1) ``` When a fence is in fact used, implementations must assume that the device may attempt to access memory immediately after receiving the MMIO signal, and subsequent memory accesses from that device to memory must observe the effects of all accesses ordered prior to that MMIO operation. In other words, in [Ordering memory and I/O accesses](#wo), suppose `0(a0)` is in main memory and `0(a1)` is the address of a device register in I/O memory. If the device accesses `0(a0)` upon receiving the MMIO write, then that load must conceptually appear after the first store to `0(a0)` according to the rules of the RVWMO memory model. In some implementations, the only way to ensure this will be to require that the first store does in fact complete before the MMIO write is issued. Other implementations may find ways to be more aggressive, while others still may not need to do anything different at all for I/O and main memory accesses. Nevertheless, the RVWMO memory model does not distinguish between these options; it simply provides an implementation-agnostic mechanism to specify the orderings that must be enforced. Many architectures include separate notions of "ordering" and "completion" fences, especially as it relates to I/O (as opposed to regular main memory). Ordering fences simply ensure that memory operations stay in order, while completion fences ensure that predecessor accesses have all completed before any successors are made visible. RISC-V does not explicitly distinguish between ordering and completion fences. Instead, this distinction is simply inferred from different uses of the FENCE bits. For implementations that conform to the RISC-V Unix Platform Specification, I/O devices and DMA operations are required to access memory coherently and via strongly ordered I/O channels. Therefore, accesses to regular main memory regions that are concurrently accessed by external devices can also use the standard synchronization mechanisms. Implementations that do not conform to the Unix Platform Specification and/or in which devices do not access memory coherently will need to use mechanisms (which are currently platform-specific or device-specific) to enforce coherency. I/O regions in the address space should be considered non-cacheable regions in the PMAs for those regions. Such regions can be considered coherent by the PMA if they are not cached by any agent. The ordering guarantees in this section may not apply beyond a platform-specific boundary between the RISC-V cores and the device. In particular, I/O accesses sent across an external bus (e.g., PCIe) may be reordered before they reach their ultimate destination. Ordering must be enforced in such situations according to the platform-specific rules of those external devices and buses. ### [](#memory%5Fporting)Code Porting and Mapping Guidelines __Table 12\. Mappings from TSO operations to RISC-V operations__ | x86/TSO Operation | RVWMO Mapping | | ----------------- | ----------------------------------------------------------------------- | | Load | l{b\|h|w|d}; fence r,rw | | Store | fence rw,w; s{b\|h|w|d} | | Atomic RMW | amo.{w\|d}.aqrl OR loop:lr.{w|d}.aq; ; sc.{w|d}.aqrl; bnez loop | | Fence | fence rw,rw | [Table 12](#tsomappings) provides a mapping from TSO memory operations onto RISC-V memory instructions. Normal x86 loads and stores are all inherently acquire-RCpc and release-RCpc operations: TSO enforces all load-load, load-store, and store-store ordering by default. Therefore, under RVWMO, all TSO loads must be mapped onto a load followed by FENCE R,RW, and all TSO stores must be mapped onto FENCE RW,W followed by a store. TSO atomic read-modify-writes and x86 instructions using the LOCK prefix are fully ordered and can be implemented either via an AMO with both _aq_ and _rl_ set, or via an LR with _aq_ set, the arithmetic operation in question, an SC with both_aq_ and _rl_ set, and a conditional branch checking the success condition. In the latter case, the _rl_ annotation on the LR turns out (for non-obvious reasons) to be redundant and can be omitted. Alternatives to [Table 12](#tsomappings) are also possible. A TSO store can be mapped onto AMOSWAP with _rl_ set. However, since RVWMO PPO Rule [3](rvwmo.html#overlapping-ordering) forbids forwarding of values from AMOs to subsequent loads, the use of AMOSWAP for stores may negatively affect performance. A TSO load can be mapped using LR with _aq_ set: all such LR instructions will be unpaired, but that fact in and of itself does not preclude the use of LR for loads. However, again, this mapping may also negatively affect performance if it puts more pressure on the reservation mechanism than was originally intended. __Table 13\. Mappings from Power operations to RISC-V operations__ | Power Operation | RVWMO Mapping | | ----------------- | ------------------ | | Load | l{b\|h|w|d} | | Load-Reserve | lr.{w\|d} | | Store | s{b\|h|w|d} | | Store-Conditional | sc.{w\|d} | | lwsync | fence.tso | | sync | fence rw,rw | | isync | fence.i; fence r,r | [Table 13](#powermappings) provides a mapping from Power memory operations onto RISC-V memory instructions. Power ISYNC maps on RISC-V to a FENCE.I followed by a FENCE R,R; the latter fence is needed because ISYNC is used to define a "control+control fence" dependency that is not present in RVWMO. __Table 14\. Mappings from ARM operations to RISC-V operations__ | ARM Operation | RVWMO Mapping | | ----------------------- | ------------------------------------- | | Load | l{b\|h|w|d} | | Load-Acquire | fence rw, rw; l{b\|h|w|d}; fence r,rw | | Load-Exclusive | lr.{w\|d} | | Load-Acquire-Exclusive | lr.{w\|d}.aqrl | | Store | s{b\|h|w|d} | | Store-Release | fence rw,w; s{b\|h|w|d} | | Store-Exclusive | sc.{w\|d} | | Store-Release-Exclusive | sc.{w\|d}.rl | | dmb | fence rw,rw | | dmb.ld | fence r,rw | | dmb.st | fence w,w | | isb | fence.i; fence r,r | [Table 14](#armmappings) provides a mapping from ARM memory operations onto RISC-V memory instructions. Since RISC-V does not currently have plain load and store opcodes with _aq_ or _rl_annotations, ARM load-acquire and store-release operations should be mapped using fences instead. Furthermore, in order to enforce store-release-to-load-acquire ordering, there must be a FENCE RW,RW between the store-release and load-acquire; [Table 14](#armmappings)enforces this by always placing the fence in front of each acquire operation. ARM load-exclusive and store-exclusive instructions can likewise map onto their RISC-V LR and SC equivalents, but instead of placing a FENCE RW,RW in front of an LR with _aq_ set, we simply also set _rl_ instead. ARM ISB maps on RISC-V to FENCE.I followed by FENCE R,R similarly to how ISYNC maps for Power. __Table 15\. Mappings from Linux memory primitives to RISC-V primitives.__ | Linux Operation | RVWMO Mapping | | ------------------------------------------------------- | --------------------------------------------------- | | smp\_mb() | fence rw,rw | | smp\_rmb() | fence r,r | | smp\_wmb() | fence w,w | | dma\_rmb() | fence r,r | | dma\_wmb() | fence w,w | | mb() | fence iorw,iorw | | rmb() | fence ri,ri | | wmb() | fence wo,wo | | smp\_load\_acquire() | l{b\|h|w|d}; fence r,rw | | smp\_store\_release() | fence.tso; s{b\|h|w|d} | | Linux Construct | RVWMO AMO Mapping | | atomic relaxed | amo .{w\|d} | | atomic acquire | amo .{w\|d}.aq | | atomic release | amo .{w\|d}.rl | | atomic | amo .{w\|d}.aqrl | | Linux Construct | RVWMO LR/SC Mapping | | atomic relaxed | loop:lr.{w\|d}; ; sc.{w|d}; bnez loop | | atomic acquire | loop:lr.{w\|d}.aq; ; sc.{w|d}; bnez loop | | atomic release | loop:lr.{w\|d}; ; sc.{w|d}.aqrl\*; bnez loop OR | | fence.tso; loop:lr.{w\|d}; ; sc.{w|d}\*; bnez loop | | | atomic | loop:lr.{w\|d}.aq; ; sc.{w|d}.aqrl; bnez loop | With regards to [Table 15](#linuxmappings), other constructs (such as spinlocks) should follow accordingly. Platforms or devices with non-coherent DMA may need additional synchronization (such as cache flush or invalidate mechanisms); currently any such extra synchronization will be device-specific. [Table 15](#linuxmappings) provides a mapping of Linux memory ordering macros onto RISC-V memory instructions. The Linux fences`dma_rmb()` and `dma_wmb()` map onto FENCE R,R and FENCE W,W, respectively, since the RISC-V Unix Platform requires coherent DMA, but would be mapped onto FENCE RI,RI and FENCE WO,WO, respectively, on a platform with non-coherent DMA. Platforms with non-coherent DMA may also require a mechanism by which cache lines can be flushed and/or invalidated. Such mechanisms will be device-specific and/or standardized in a future extension to the ISA. The Linux mappings for release operations may seem stronger than necessary, but these mappings are needed to cover some cases in which Linux requires stronger orderings than the more intuitive mappings would provide. In particular, as of the time this text is being written, Linux is actively debating whether to require load-load, load-store, and store-store orderings between accesses in one critical section and accesses in a subsequent critical section in the same hart and protected by the same synchronization object. Not all combinations of FENCE RW,W/FENCE R,RW mappings with _aq_/_rl_ mappings combine to provide such orderings. There are a few ways around this problem, including: 1. Always use FENCE RW,W/FENCE R,RW, and never use _aq_/_rl_. This suffices but is undesirable, as it defeats the purpose of the _aq_/_rl_modifiers. 2. Always use _aq_/_rl_, and never use FENCE RW,W/FENCE R,RW. This does not currently work due to the lack of load and store opcodes with _aq_and _rl_ modifiers. 3. Strengthen the mappings of release operations such that they would enforce sufficient orderings in the presence of either type of acquire mapping. This is the currently recommended solution, and the one shown in [Table 15](#linuxmappings). RVWMO Mapping: (a) lw a0, 0(s0) (b) fence.tso // vs. fence rw,w (c) sd x0,0(s1) …​ loop: (d) amoswap.d.aq a1,t1,0(s1) bnez a1,loop (e) lw a2,0(s2) For example, the critical section ordering rule currently being debated by the Linux community would require (a) to be ordered before (e) in[Orderings between critical sections in Linux](#lkmm%5Fll). If that will indeed be required, then it would be insufficient for (b) to map as FENCE RW,W. That said, these mappings are subject to change as the Linux Kernel Memory Model evolves. Orderings between critical sections in Linux ```asm Linux Code: (a) int r0 = *x; (bc) spin_unlock(y, 0); .... .... (d) spin_lock(y); (e) int r1 = *z; RVWMO Mapping: (a) lw a0, 0(s0) (b) fence.tso // vs. fence rw,w (c) sd x0,0(s1) .... loop: (d) lr.d.aq a1,(s1) bnez a1,loop sc.d a1,t1,(s1) bnez a1,loop (e) lw a2,0(s2) ``` [Table 16](#c11mappings) provides a mapping of C11/C++11 atomic operations onto RISC-V memory instructions. If load and store opcodes with _aq_ and _rl_ modifiers are introduced, then the mappings in[Table 17](#c11mappings%5Fhypothetical) will suffice. Note however that the two mappings only interoperate correctly if`atomic_(memory_order_seq_cst)` is mapped using an LR that has both_aq_ and _rl_ set. Even more importantly, a [Table 16](#c11mappings) sequentially consistent store, followed by a [Table 17](#c11mappings%5Fhypothetical) sequentially consistent load can be reordered unless the [Table 16](#c11mappings) mapping of stores is strengthened by either adding a second fence or mapping the store to `amoswap.rl` instead. __Table 16\. Mappings from C/C++ primitives to RISC-V primitives.__ | C/C++ Construct | RVWMO Mapping | | ---------------------------------------------- | ------------------------------------- | | Non-atomic load | l{b\|h|w|d} | | atomic\_load(memory\_order\_relaxed) | l{b\|h|w|d} | | atomic\_load(memory\_order\_acquire) | l{b\|h|w|d}; fence r,rw | | atomic\_load(memory\_order\_seq\_cst) | fence rw,rw; l{b\|h|w|d}; fence r,rw | | Non-atomic store | s{b\|h|w|d} | | atomic\_store(memory\_order\_relaxed) | s{b\|h|w|d} | | atomic\_store(memory\_order\_release) | fence rw,w; s{b\|h|w|d} | | atomic\_store(memory\_order\_seq\_cst) | fence rw,w; s{b\|h|w|d} | | atomic\_thread\_fence(memory\_order\_acquire) | fence r,rw | | atomic\_thread\_fence(memory\_order\_release) | fence rw,w | | atomic\_thread\_fence(memory\_order\_acq\_rel) | fence.tso | | atomic\_thread\_fence(memory\_order\_seq\_cst) | fence rw,rw | | C/C++ Construct | RVWMO AMO Mapping | | atomic\_(memory\_order\_relaxed) | amo.{w\|d} | | atomic\_(memory\_order\_acquire) | amo.{w\|d}.aq | | atomic\_(memory\_order\_release) | amo.{w\|d}.rl | | atomic\_(memory\_order\_acq\_rel) | amo.{w\|d}.aqrl | | atomic\_(memory\_order\_seq\_cst) | amo.{w\|d}.aqrl | | C/C++ Construct | RVWMO LR/SC Mapping | | atomic\_(memory\_order\_relaxed) | loop:lr.{w\|d}; ; sc.{w|d}; | | bnez loop | | | atomic\_(memory\_order\_acquire) | loop:lr.{w\|d}.aq; ; sc.{w|d}; | | bnez loop | | | atomic\_(memory\_order\_release) | loop:lr.{w\|d}; ; sc.{w|d}.rl; | | bnez loop | | | atomic\_(memory\_order\_acq\_rel) | loop:lr.{w\|d}.aq; ; sc.{w|d}.rl; | | bnez loop | | | atomic\_(memory\_order\_seq\_cst) | loop:lr.{w\|d}.aqrl; ; | | sc.{w\|d}.rl; bnez loop | | __Table 17\. Hypothetical mappings from C/C++ primitives to RISC-V primitives, if native load-acquire and store-release opcodes are introduced.__ | C/C++ Construct | RVWMO Mapping | | ------------------------------------------------------------------------------------------------ | -------------------------------- | | Non-atomic load | l{b\|h|w|d} | | atomic\_load(memory\_order\_relaxed) | l{b\|h|w|d} | | atomic\_load(memory\_order\_acquire) | l{b\|h|w|d}.aq | | atomic\_load(memory\_order\_seq\_cst) | l{b\|h|w|d}.aq | | Non-atomic store | s{b\|h|w|d} | | atomic\_store(memory\_order\_relaxed) | s{b\|h|w|d} | | atomic\_store(memory\_order\_release) | s{b\|h|w|d}.rl | | atomic\_store(memory\_order\_seq\_cst) | s{b\|h|w|d}.rl | | atomic\_thread\_fence(memory\_order\_acquire) | fence r,rw | | atomic\_thread\_fence(memory\_order\_release) | fence rw,w | | atomic\_thread\_fence(memory\_order\_acq\_rel) | fence.tso | | atomic\_thread\_fence(memory\_order\_seq\_cst) | fence rw,rw | | C/C++ Construct | RVWMO AMO Mapping | | atomic\_(memory\_order\_relaxed) | amo.{w\|d} | | atomic\_(memory\_order\_acquire) | amo.{w\|d}.aq | | atomic\_(memory\_order\_release) | amo.{w\|d}.rl | | atomic\_(memory\_order\_acq\_rel) | amo.{w\|d}.aqrl | | atomic\_(memory\_order\_seq\_cst) | amo.{w\|d}.aqrl | | C/C++ Construct | RVWMO LR/SC Mapping | | atomic\_(memory\_order\_relaxed) | lr.{w\|d}; ; sc.{w|d} | | atomic\_(memory\_order\_acquire) | lr.{w\|d}.aq; ; sc.{w|d} | | atomic\_(memory\_order\_release) | lr.{w\|d}; ; sc.{w|d}.rl | | atomic\_(memory\_order\_acq\_rel) | lr.{w\|d}.aq; ; sc.{w|d}.rl | | atomic\_(memory\_order\_seq\_cst) | lr.{w\|d}.aq\* ; sc.{w|d}.rl | | \* must be lr.{w\|d}.aqrl in order to interoperate with code mapped per [Table 16](#c11mappings) | | Any AMO can be emulated by an LR/SC pair, but care must be taken to ensure that any PPO orderings that originate from the LR are also made to originate from the SC, and that any PPO orderings that terminate at the SC are also made to terminate at the LR. For example, the LR must also be made to respect any data dependencies that the AMO has, given that load operations do not otherwise have any notion of a data dependency. Likewise, the effect a FENCE R,R elsewhere in the same hart must also be made to apply to the SC, which would not otherwise respect that fence. The emulator may achieve this effect by simply mapping AMOs onto `lr.aq; ; sc.aqrl`, matching the mapping used elsewhere for fully ordered atomics. These C11/C++11 mappings require the platform to provide the following Physical Memory Attributes (as defined in the RISC-V Privileged ISA) for all memory: * main memory * coherent * AMOArithmetic * RsrvEventual Platforms with different attributes may require different mappings, or require platform-specific SW (e.g., memory-mapped I/O). ### [](#implementation-guidelines)Implementation Guidelines The RVWMO and RVTSO memory models by no means preclude microarchitectures from employing sophisticated speculation techniques or other forms of optimization in order to deliver higher performance. The models also do not impose any requirement to use any one particular cache hierarchy, nor even to use a cache coherence protocol at all. Instead, these models only specify the behaviors that can be exposed to software. Microarchitectures are free to use any pipeline design, any coherent or non-coherent cache hierarchy, any on-chip interconnect, etc., as long as the design only admits executions that satisfy the memory model rules. That said, to help people understand the actual implementations of the memory model, in this section we provide some guidelines on how architects and programmers should interpret the models' rules. Both RVWMO and RVTSO are multi-copy atomic (or_other-multi-copy-atomic_): any store value that is visible to a hart other than the one that originally issued it must also be conceptually visible to all other harts in the system. In other words, harts may forward from their own previous stores before those stores have become globally visible to all harts, but no early inter-hart forwarding is permitted. Multi-copy atomicity may be enforced in a number of ways. It might hold inherently due to the physical design of the caches and store buffers, it may be enforced via a single-writer/multiple-reader cache coherence protocol, or it might hold due to some other mechanism. Although multi-copy atomicity does impose some restrictions on the microarchitecture, it is one of the key properties keeping the memory model from becoming extremely complicated. For example, a hart may not legally forward a value from a neighbor hart’s private store buffer (unless of course it is done in such a way that no new illegal behaviors become architecturally visible). Nor may a cache coherence protocol forward a value from one hart to another until the coherence protocol has invalidated all older copies from other caches. Of course, microarchitectures may (and high-performance implementations likely will) violate these rules under the covers through speculation or other optimizations, as long as any non-compliant behaviors are not exposed to the programmer. As a rough guideline for interpreting the PPO rules in RVWMO, we expect the following from the software perspective: * programmers will use PPO rules [1](rvwmo.html#overlapping-ordering) and [4-8](rvwmo.html#overlapping-ordering) regularly and actively. * expert programmers will use PPO rules [9-11](rvwmo.html#overlapping-ordering) to speed up critical paths of important data structures. * even expert programmers will rarely if ever use PPO rules [2-3](rvwmo.html#overlapping-ordering) and[12-13](rvwmo.html#overlapping-ordering) directly. These are included to facilitate common microarchitectural optimizations (rule [2](rvwmo.html#overlapping-ordering)) and the operational formal modeling approach (rules [3](rvwmo.html#overlapping-ordering) and[12-13](rvwmo.html#overlapping-ordering)) described in [An Operational Memory Model](mm-formal.html#operational). They also facilitate the process of porting code from other architectures that have similar rules. We also expect the following from the hardware perspective: * PPO rules [1](rvwmo.html#overlapping-ordering) and [3-6](rvwmo.html#overlapping-ordering) reflect well-understood rules that should pose few surprises to architects. * PPO rule [2](rvwmo.html#overlapping-ordering) reflects a natural and common hardware optimization, but one that is very subtle and hence is worth double checking carefully. * PPO rule [7](rvwmo.html#overlapping-ordering) may not be immediately obvious to architects, but it is a standard memory model requirement * The load value axiom, the atomicity axiom, and PPO rules[8-13](rvwmo.html#overlapping-ordering) reflect rules that most hardware implementations will enforce naturally, unless they contain extreme optimizations. Of course, implementations should make sure to double check these rules nevertheless. Hardware must also ensure that syntactic dependencies are not `optimized away`. Architectures are free to implement any of the memory model rules as conservatively as they choose. For example, a hardware implementation may choose to do any or all of the following: * interpret all fences as if they were FENCE RW,RW (or FENCE IORW,IORW, if I/O is involved), regardless of the bits actually set * implement all fences with PW and SR as if they were FENCE RW,RW (or FENCE IORW,IORW, if I/O is involved), as PW with SR is the most expensive of the four possible main memory ordering components anyway * emulate _aq_ and _rl_ as described in [Code Porting and Mapping Guidelines](#memory%5Fporting) * enforcing all same-address load-load ordering, even in the presence of patterns such as `fri-rfi` and `RSW` * forbid any forwarding of a value from a store in the store buffer to a subsequent AMO or LR to the same address * forbid any forwarding of a value from an AMO or SC in the store buffer to a subsequent load to the same address * implement TSO on all memory accesses, and ignore any main memory fences that do not include PW and SR ordering (e.g., as Ztso implementations will do) * implement all atomics to be RCsc or even fully ordered, regardless of annotation Architectures that implement RVTSO can safely do the following: * Ignore all fences that do not have both PW and SR (unless the fence also orders I/O) * Ignore all PPO rules except for rules [4](rvwmo.html#overlapping-ordering) through [7](rvwmo.html#overlapping-ordering), since the rest are redundant with other PPO rules under RVTSO assumptions Other general notes: * Silent stores (i.e., stores that write the same value that already exists at a memory location) behave like any other store from a memory model point of view. Likewise, AMOs which do not actually change the value in memory (e.g., an AMOMAX for which the value in _rs2_ is smaller than the value currently in memory) are still semantically considered store operations. Microarchitectures that attempt to implement silent stores must take care to ensure that the memory model is still obeyed, particularly in cases such as RSW [Overlapping-Address Orderings (Rules 1-3)](#mm-overlap)which tend to be incompatible with silent stores. * Writes may be merged (i.e., two consecutive writes to the same address may be merged) or subsumed (i.e., the earlier of two back-to-back writes to the same address may be elided) as long as the resulting behavior does not otherwise violate the memory model semantics. The question of write subsumption can be understood from the following example: __Table 18\. Write subsumption litmus test, allowed execution__ | Hart 0 Hart 1 li t1, 3 li t3, 2 li t2, 1 (a) sw t1,0(s0) (d) lw a0,0(s1) (b) fence w, w (e) sw a0,0(s0) (c) sw t2,0(s1) (f) sw t3,0(s0) | ![litmus subsumption](_images/graphviz/litmus_subsumption.png) | | --------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------- | As written, if the load (d) reads value _1_, then (a) must precede (f) in the global memory order: * (a) precedes (c) in the global memory order because of rule 4 * (c) precedes (d) in the global memory order because of the Load Value axiom * (d) precedes (e) in the global memory order because of rule 10 * (e) precedes (f) in the global memory order because of rule 1 In other words the final value of the memory location whose address is in `s0` must be _2_ (the value written by the store (f)) and cannot be _3_ (the value written by the store (a)). A very aggressive microarchitecture might erroneously decide to discard (e), as (f) supersedes it, and this may in turn lead the microarchitecture to break the now-eliminated dependency between (d) and (f) (and hence also between (a) and (f)). This would violate the memory model rules, and hence it is forbidden. Write subsumption may in other cases be legal, if for example there were no data dependency between (d) and (e). #### [](#possible-future-extensions)Possible Future Extensions We expect that any or all of the following possible future extensions would be compatible with the RVWMO memory model: * "V" vector ISA extensions * "J" JIT extension * Native encodings for load and store opcodes with _aq_ and _rl_ set * Fences limited to certain addresses * Cache write-back/flush/invalidate/etc.instructions ### [](#discrepancies)Known Issues #### [](#mixedrsw)Mixed-size RSW __Table 19\. Mixed-size discrepancy (permitted by axiomatic models, forbidden by operational model)__ | Hart 0 | Hart 1 | | | | ------------------------------------- | ----------- | --- | ------------------------ | | li t1, 1 | li t1, 1 | | | | (a) | lw a0,0(s0) | (d) | lw a1,0(s1) | | (b) | fence rw,rw | (e) | amoswap.w.rl a2,t1,0(s2) | | (c) | sw t1,0(s1) | (f) | ld a3,0(s2) | | (g) | lw a4,4(s2) | | | | xor a5,a4,a4 | | | | | add s0,s0,a5 | | | | | (h) | sw t1,0(s0) | | | | Outcome: a0=1, a1=1, a2=0, a3=1, a4=0 | | | | __Table 20\. Mixed-size discrepancy (permitted by axiomatic models, forbidden by operational model)__ | Hart 0 | Hart 1 | | | | ------------------------- | ----------- | ------------ | ----------- | | li t1, 1 | li t1, 1 | | | | (a) | lw a0,0(s0) | (d) | ld a1,0(s1) | | (b) | fence rw,rw | (e) | lw a2,4(s1) | | (c) | sw t1,0(s1) | xor a3,a2,a2 | | | add s0,s0,a3 | | | | | (f) | sw t1,0(s0) | | | | Outcome: a0=1, a1=1, a2=0 | | | | __Table 21\. Mixed-size discrepancy (permitted by axiomatic models, forbidden by operational model)__ | Hart 0 | Hart 1 | | | | ----------------------------------- | ----------- | --- | ----------- | | li t1, 1 | li t1, 1 | | | | (a) | lw a0,0(s0) | (d) | sw t1,4(s1) | | (b) | fence rw,rw | (e) | ld a1,0(s1) | | (c) | sw t1,0(s1) | (f) | lw a2,4(s1) | | xor a3,a2,a2 | | | | | add s0,s0,a3 | | | | | (g) | sw t1,0(s0) | | | | Outcome: a0=1, a1=0x100000001, a2=1 | | | | There is a known discrepancy between the operational and axiomatic specifications within the family of mixed-size RSW variants shown in[Table 19](#rsw1)\-[Table 21](#rsw3). To address this, we may choose to add something like the following new PPO rule: Memory operation _a_ precedes memory operation_b_ in preserved program order (and hence also in the global memory order) if _a_ precedes _b_ in program order, _a_ and _b_ both access regular main memory (rather than I/O regions), _a_ is a load,_b_ is a store, there is a load _m_ between_a_ and _b_, there is a byte _x_that both _a_ and _m_ read, there is no store between _a_ and _m_ that writes to_x_, and _m_ precedes _b_ in PPO. In other words, in herd syntax, we may choose to add`(po-loc & rsw);ppo;[W]` to PPO. Many implementations will already enforce this ordering naturally. As such, even though this rule is not official, we recommend that implementers enforce it nevertheless in order to ensure forwards compatibility with the possible future addition of this rule to RVWMO. Formal Memory Model Specifications, Version 0.1 ==================== ## [](#formal-memory-model-specifications-version-0-1)Appendix A: Formal Memory Model Specifications, Version 0.1 To facilitate formal analysis of RVWMO, this chapter presents a set of formalizations using different tools and modeling approaches. Any discrepancies are unintended; the expectation is that the models describe exactly the same sets of legal behaviors. This appendix should be treated as commentary; all normative material is provided in [RVWMO Memory Model](rvwmo.html#memorymodel) and in the rest of the main body of the ISA specification. All currently known discrepancies are listed in [Known Issues](mm-eplan.html#discrepancies). Any other discrepancies are unintentional. ### [](#alloy)Formal Axiomatic Specification in Alloy We present a formal specification of the RVWMO memory model in Alloy (). This model is available online at. The online material also contains some litmus tests and some examples of how Alloy can be used to model check some of the mappings in [Code Porting and Mapping Guidelines](mm-eplan.html#memory%5Fporting). The RVWMO memory model formalized in Alloy (1/5: PPO) ```c // =RVWMO PPO= // Preserved Program Order fun ppo : Event->Event { // same-address ordering po_loc :> Store + rdw + (AMO + StoreConditional) <: rfi // explicit synchronization + ppo_fence + Acquire <: ^po :> MemoryEvent + MemoryEvent <: ^po :> Release + RCsc <: ^po :> RCsc + pair // syntactic dependencies + addrdep + datadep + ctrldep :> Store // pipeline dependencies + (addrdep+datadep).rfi + addrdep.^po :> Store } // the global memory order respects preserved program order fact { ppo in ^gmo } ``` The RVWMO memory model formalized in Alloy (2/5: Axioms) // =RVWMO axioms= // Load Value Axiom fun candidates[r: MemoryEvent] : set MemoryEvent { (r.~^gmo & Store & same_addr[r]) // writes preceding r in gmo + (r.^~po & Store & same_addr[r]) // writes preceding r in po } fun latest_among[s: set Event] : Event { s - s.~^gmo } pred LoadValue { all w: Store | all r: Load | w->r in rf <=> w = latest_among[candidates[r]] } // Atomicity Axiom pred Atomicity { all r: Store.~pair | // starting from the lr, no x: Store & same_addr[r] | // there is no store x to the same addr x not in same_hart[r] // such that x is from a different hart, and x in r.~rf.^gmo // x follows (the store r reads from) in gmo, and r.pair in x.^gmo // and r follows x in gmo } // Progress Axiom implicit: Alloy only considers finite executions pred RISCV_mm { LoadValue and Atomicity /* and Progress */ } The RVWMO memory model formalized in Alloy (3/5: model of memory) ```sml //Basic model of memory sig Hart { // hardware thread start : one Event } sig Address {} abstract sig Event { po: lone Event // program order } abstract sig MemoryEvent extends Event { address: one Address, acquireRCpc: lone MemoryEvent, acquireRCsc: lone MemoryEvent, releaseRCpc: lone MemoryEvent, releaseRCsc: lone MemoryEvent, addrdep: set MemoryEvent, ctrldep: set Event, datadep: set MemoryEvent, gmo: set MemoryEvent, // global memory order rf: set MemoryEvent } sig LoadNormal extends MemoryEvent {} // l{b|h|w|d} sig LoadReserve extends MemoryEvent { // lr pair: lone StoreConditional } sig StoreNormal extends MemoryEvent {} // s{b|h|w|d} // all StoreConditionals in the model are assumed to be successful sig StoreConditional extends MemoryEvent {} // sc sig AMO extends MemoryEvent {} // amo sig NOP extends Event {} fun Load : Event { LoadNormal + LoadReserve + AMO } fun Store : Event { StoreNormal + StoreConditional + AMO } sig Fence extends Event { pr: lone Fence, // opcode bit pw: lone Fence, // opcode bit sr: lone Fence, // opcode bit sw: lone Fence // opcode bit } sig FenceTSO extends Fence {} /* Alloy encoding detail: opcode bits are either set (encoded, e.g., * as f.pr in iden) or unset (f.pr not in iden). The bits cannot be used for * anything else */ fact { pr + pw + sr + sw in iden } // likewise for ordering annotations fact { acquireRCpc + acquireRCsc + releaseRCpc + releaseRCsc in iden } // don't try to encode FenceTSO via pr/pw/sr/sw; just use it as-is fact { no FenceTSO.(pr + pw + sr + sw) } ``` The RVWMO memory model formalized in Alloy (4/5: Basic model rules) ```scala // =Basic model rules= // Ordering annotation groups fun Acquire : MemoryEvent { MemoryEvent.acquireRCpc + MemoryEvent.acquireRCsc } fun Release : MemoryEvent { MemoryEvent.releaseRCpc + MemoryEvent.releaseRCsc } fun RCpc : MemoryEvent { MemoryEvent.acquireRCpc + MemoryEvent.releaseRCpc } fun RCsc : MemoryEvent { MemoryEvent.acquireRCsc + MemoryEvent.releaseRCsc } // There is no such thing as store-acquire or load-release, unless it's both fact { Load & Release in Acquire } fact { Store & Acquire in Release } // FENCE PPO fun FencePRSR : Fence { Fence.(pr & sr) } fun FencePRSW : Fence { Fence.(pr & sw) } fun FencePWSR : Fence { Fence.(pw & sr) } fun FencePWSW : Fence { Fence.(pw & sw) } fun ppo_fence : MemoryEvent->MemoryEvent { (Load <: ^po :> FencePRSR).(^po :> Load) + (Load <: ^po :> FencePRSW).(^po :> Store) + (Store <: ^po :> FencePWSR).(^po :> Load) + (Store <: ^po :> FencePWSW).(^po :> Store) + (Load <: ^po :> FenceTSO) .(^po :> MemoryEvent) + (Store <: ^po :> FenceTSO) .(^po :> Store) } // auxiliary definitions fun po_loc : Event->Event { ^po & address.~address } fun same_hart[e: Event] : set Event { e + e.^~po + e.^po } fun same_addr[e: Event] : set Event { e.address.~address } // initial stores fun NonInit : set Event { Hart.start.*po } fun Init : set Event { Event - NonInit } fact { Init in StoreNormal } fact { Init->(MemoryEvent & NonInit) in ^gmo } fact { all e: NonInit | one e.*~po.~start } // each event is in exactly one hart fact { all a: Address | one Init & a.~address } // one init store per address fact { no Init <: po and no po :> Init } ``` The RVWMO memory model formalized in Alloy (5/5: Auxiliaries) ```asm // po fact { acyclic[po] } // gmo fact { total[^gmo, MemoryEvent] } // gmo is a total order over all MemoryEvents //rf fact { rf.~rf in iden } // each read returns the value of only one write fact { rf in Store <: address.~address :> Load } fun rfi : MemoryEvent->MemoryEvent { rf & (*po + *~po) } //dep fact { no StoreNormal <: (addrdep + ctrldep + datadep) } fact { addrdep + ctrldep + datadep + pair in ^po } fact { datadep in datadep :> Store } fact { ctrldep.*po in ctrldep } fact { no pair & (^po :> (LoadReserve + StoreConditional)).^po } fact { StoreConditional in LoadReserve.pair } // assume all SCs succeed // rdw fun rdw : Event->Event { (Load <: po_loc :> Load) // start with all same_address load-load pairs, - (~rf.rf) // subtract pairs that read from the same store, - (po_loc.rfi) // and subtract out "fri-rfi" patterns } // filter out redundant instances and/or visualizations fact { no gmo & gmo.gmo } // keep the visualization uncluttered fact { all a: Address | some a.~address } // =Optional: opcode encoding restrictions= // the list of blessed fences fact { Fence in Fence.pr.sr + Fence.pw.sw + Fence.pr.pw.sw + Fence.pr.sr.sw + FenceTSO + Fence.pr.pw.sr.sw } pred restrict_to_current_encodings { no (LoadNormal + StoreNormal) & (Acquire + Release) } // =Alloy shortcuts= pred acyclic[rel: Event->Event] { no iden & ^rel } pred total[rel: Event->Event, bag: Event] { all disj e, f: bag | e->f in rel + ~rel acyclic[rel] } ``` ### [](#sec:herd)Formal Axiomatic Specification in Herd The tool herd takes a memory model and a litmus test as input and simulates the execution of the test on top of the memory model. Memory models are written in the domain specific language Cat. This section provides two Cat memory model of RVWMO. The first model,[riscv.cat, a herd version of the RVWMO memory model (2/3)](#herd2), follows the _global memory order_, Chapter [RVWMO Memory Consistency Model](rvwmo.html), definition of RVWMO, as much as is possible for a Cat model. The second model,[riscv.cat, an alternative herd presentation of the RVWMO memory model (3/3)](#herd3), is an equivalent, more efficient, partial order based RVWMO model. The simulator `herd` is part of the `diy` tool suite — see for software and documentation. The models and more are available online at . riscv-defs.cat, a herd definition of preserved program order (1/3) ```asm (*************) (* Utilities *) (*************) (* All fence relations *) let fence.r.r = [R];fencerel(Fence.r.r);[R] let fence.r.w = [R];fencerel(Fence.r.w);[W] let fence.r.rw = [R];fencerel(Fence.r.rw);[M] let fence.w.r = [W];fencerel(Fence.w.r);[R] let fence.w.w = [W];fencerel(Fence.w.w);[W] let fence.w.rw = [W];fencerel(Fence.w.rw);[M] let fence.rw.r = [M];fencerel(Fence.rw.r);[R] let fence.rw.w = [M];fencerel(Fence.rw.w);[W] let fence.rw.rw = [M];fencerel(Fence.rw.rw);[M] let fence.tso = let f = fencerel(Fence.tso) in ([W];f;[W]) | ([R];f;[M]) let fence = fence.r.r | fence.r.w | fence.r.rw | fence.w.r | fence.w.w | fence.w.rw | fence.rw.r | fence.rw.w | fence.rw.rw | fence.tso (* Same address, no W to the same address in-between *) let po-loc-no-w = po-loc \ (po-loc?;[W];po-loc) (* Read same write *) let rsw = rf^-1;rf (* Acquire, or stronger *) let AQ = Acq|AcqRel (* Release or stronger *) and RL = RelAcqRel (* All RCsc *) let RCsc = Acq|Rel|AcqRel (* Amo events are both R and W, relation rmw relates paired lr/sc *) let AMO = R & W let StCond = range(rmw) (*************) (* ppo rules *) (*************) (* Overlapping-Address Orderings *) let r1 = [M];po-loc;[W] and r2 = ([R];po-loc-no-w;[R]) \ rsw and r3 = [AMO|StCond];rfi;[R] (* Explicit Synchronization *) and r4 = fence and r5 = [AQ];po;[M] and r6 = [M];po;[RL] and r7 = [RCsc];po;[RCsc] and r8 = rmw (* Syntactic Dependencies *) and r9 = [M];addr;[M] and r10 = [M];data;[W] and r11 = [M];ctrl;[W] (* Pipeline Dependencies *) and r12 = [R];(addr|data);[W];rfi;[R] and r13 = [R];addr;[M];po;[W] let ppo = r1 | r2 | r3 | r4 | r5 | r6 | r7 | r8 | r9 | r10 | r11 | r12 | r13 ``` riscv.cat, a herd version of the RVWMO memory model (2/3) ```asm Total (* Notice that herd has defined its own rf relation *) (* Define ppo *) include "riscv-defs.cat" (********************************) (* Generate global memory order *) (********************************) let gmo0 = (* precursor: ie build gmo as an total order that include gmo0 *) loc & (W\FW) * FW | # Final write after any write to the same location ppo | # ppo compatible rfe # includes herd external rf (optimization) (* Walk over all linear extensions of gmo0 *) with gmo from linearizations(M\IW,gmo0) (* Add initial writes upfront -- convenient for computing rfGMO *) let gmo = gmo | loc & IW * (M\IW) (**********) (* Axioms *) (**********) (* Compute rf according to the load value axiom, aka rfGMO *) let WR = loc & ([W];(gmo|po);[R]) let rfGMO = WR \ (loc&([W];gmo);WR) (* Check equality of herd rf and of rfGMO *) empty (rf\rfGMO)|(rfGMO\rf) as RfCons (* Atomicity axiom *) let infloc = (gmo & loc)^-1 let inflocext = infloc & ext let winside = (infloc;rmw;inflocext) & (infloc;rf;rmw;inflocext) & [W] empty winside as Atomic ``` `riscv.cat`, an alternative herd presentation of the RVWMO memory model (3/3) ```asm Partial (***************) (* Definitions *) (***************) (* Define ppo *) include "riscv-defs.cat" (* Compute coherence relation *) include "cos-opt.cat" (**********) (* Axioms *) (**********) (* Sc per location *) acyclic co|rf|fr|po-loc as Coherence (* Main model axiom *) acyclic co|rfe|fr|ppo as Model (* Atomicity axiom *) empty rmw & (fre;coe) as Atomic ``` ### [](#operational)An Operational Memory Model This is an alternative presentation of the RVWMO memory model in operational style. It aims to admit exactly the same extensional behavior as the axiomatic presentation: for any given program, admitting an execution if and only if the axiomatic presentation allows it. The axiomatic presentation is defined as a predicate on complete candidate executions. In contrast, this operational presentation has an abstract microarchitectural flavor: it is expressed as a state machine, with states that are an abstract representation of hardware machine states, and with explicit out-of-order and speculative execution (but abstracting from more implementation-specific microarchitectural details such as register renaming, store buffers, cache hierarchies, cache protocols, etc.). As such, it can provide useful intuition. It can also construct executions incrementally, making it possible to interactively and randomly explore the behavior of larger examples, while the axiomatic model requires complete candidate executions over which the axioms can be checked. The operational presentation covers mixed-size execution, with potentially overlapping memory accesses of different power-of-two byte sizes. Misaligned accesses are broken up into single-byte accesses. The operational model, together with a fragment of the RISC-V ISA semantics (RV64I and A), are integrated into the `rmem` exploration tool (). `rmem` can explore litmus tests (see [Litmus Tests](mm-eplan.html#litmustests)) and small ELF binaries exhaustively, pseudorandomly and interactively. In `rmem`, the ISA semantics is expressed explicitly in Sail (see for the Sail language, and for the RISC-V ISA model), and the concurrency semantics is expressed in Lem (see for the Lem language). `rmem` has a command-line interface and a web-interface. The web-interface runs entirely on the client side, and is provided online together with a library of litmus tests:. The command-line interface is faster than the web-interface, specially in exhaustive mode. Below is an informal introduction of the model states and transitions. The description of the formal model starts in the next subsection. Terminology: In contrast to the axiomatic presentation, here every memory operation is either a load or a store. Hence, AMOs give rise to two distinct memory operations, a load and a store. When used in conjunction with `instruction`, the terms `load` and `store` refer to instructions that give rise to such memory operations. As such, both include AMO instructions. The term `acquire` refers to an instruction (or its memory operation) with the acquire-RCpc or acquire-RCsc annotation. The term `release` refers to an instruction (or its memory operation) with the release-RCpc or release-RCsc annotation. **Model states** Model states: A model state consists of a shared memory and a tuple of hart states. ![Diagram](_images/diag-65ec6465e14cb6224e7684808e67fc614df27844.svg) The shared memory state records all the memory store operations that have propagated so far, in the order they propagated (this can be made more efficient, but for simplicity of the presentation we keep it this way). Each hart state consists principally of a tree of instruction instances, some of which have been _finished_, and some of which have not. Non-finished instruction instances can be subject to _restart_, e.g. if they depend on an out-of-order or speculative load that turns out to be unsound. Conditional branch and indirect jump instructions may have multiple successors in the instruction tree. When such instruction is finished, any untaken alternative paths are discarded. Each instruction instance in the instruction tree has a state that includes an execution state of the intra-instruction semantics (the ISA pseudocode for this instruction). The model uses a formalization of the intra-instruction semantics in Sail. One can think of the execution state of an instruction as a representation of the pseudocode control state, pseudocode call stack, and local variable values. An instruction instance state also includes information about the instance’s memory and register footprints, its register reads and writes, its memory operations, whether it is finished, etc. **Model transitions** The model defines, for any model state, the set of allowed transitions, each of which is a single atomic step to a new abstract machine state. Execution of a single instruction will typically involve many transitions, and they may be interleaved in operational-model execution with transitions arising from other instructions. Each transition arises from a single instruction instance; it will change the state of that instance, and it may depend on or change the rest of its hart state and the shared memory state, but it does not depend on other hart states, and it will not change them. The transitions are introduced below and defined in [Transitions](#transitions), with a precondition and a construction of the post-transition model state for each. Transitions for all instructions: * [Fetch instruction](#fetch): This transition represents a fetch and decode of a new instruction instance, as a program order successor of a previously fetched instruction instance (or the initial fetch address). The model assumes the instruction memory is fixed; it does not describe the behavior of self-modifying code. In particular, the [Fetch instruction](#fetch) transition does not generate memory load operations, and the shared memory is not involved in the transition. Instead, the model depends on an external oracle that provides an opcode when given a memory location. * [Register write](#reg%5Fwrite): This is a write of a register value. * [Register read](#reg%5Fread): This is a read of a register value from the most recent program-order-predecessor instruction instance that writes to that register. * [Pseudocode internal step](#sail%5Finterp): This covers pseudocode internal computation: arithmetic, function calls, etc. * [Finish instruction](#finish): At this point the instruction pseudocode is done, the instruction cannot be restarted, memory accesses cannot be discarded, and all memory effects have taken place. For conditional branch and indirect jump instructions, any program order successors that were fetched from an address that is not the one that was written to the _pc_ register are discarded, together with the sub-tree of instruction instances below them. Transitions specific to load instructions: * [Initiate memory load operations](#initiate%5Fload): At this point the memory footprint of the load instruction is provisionally known (it could change if earlier instructions are restarted) and its individual memory load operations can start being satisfied. * [Satisfy memory load operation by forwarding from unpropogated stores](#sat%5Fby%5Fforwarding): This partially or entirely satisfies a single memory load operation by forwarding, from program-order-previous memory store operations. * [Satisfy memory load operation from memory](#sat%5Ffrom%5Fmem): This entirely satisfies the outstanding slices of a single memory load operation, from memory. * [Complete load operations](#complete%5Floads): At this point all the memory load operations of the instruction have been entirely satisfied and the instruction pseudocode can continue executing. A load instruction can be subject to being restarted until the transition. But, under some conditions, the model might treat a load instruction as non-restartable even before it is finished (e.g. see ). Transitions specific to store instructions: * [Initiate memory store operation footprints](#initiate%5Fstore%5Ffootprint): At this point the memory footprint of the store is provisionally known. * [Instantiate memory store operation values](#instantiate%5Fstore%5Fvalue): At this point the memory store operations have their values and program-order-successor memory load operations can be satisfied by forwarding from them. * [Commit store instruction](#commit%5Fstores): At this point the store operations are guaranteed to happen (the instruction can no longer be restarted or discarded), and they can start being propagated to memory. * [Propagate store operation](#prop%5Fstore): This propagates a single memory store operation to memory. * [Complete store operations](#complete%5Fstores): At this point all the memory store operations of the instruction have been propagated to memory, and the instruction pseudocode can continue executing. Transitions specific to `sc` instructions: * [Early sc fail](#early%5Fsc%5Ffail): This causes the `sc` to fail, either a spontaneous fail or because it is not paired with a program-order-previous `lr`. * [Paired sc](#paired%5Fsc): This transition indicates the `sc` is paired with an `lr` and might succeed. * [Commit and propagate store operation of an sc](#commit%5Fsc): This is an atomic execution of the transitions [Commit store instruction](#commit%5Fstores) and [Propagate store operation](#prop%5Fstore), it is enabled only if the stores from which the `lr` read from have not been overwritten. * [Late sc fail](#late%5Fsc%5Ffail): This causes the `sc` to fail, either a spontaneous fail or because the stores from which the `lr` read from have been overwritten. Transitions specific to AMO instructions: * [Satisfy, commit and propagate operations of an AMO](#do%5Famo): This is an atomic execution of all the transitions needed to satisfy the load operation, do the required arithmetic, and propagate the store operation. Transitions specific to fence instructions: * [Commit fence](#commit%5Ffence) The transitions labeled ○ can always be taken eagerly, as soon as their precondition is satisfied, without excluding other behavior; the ∙ cannot. Although [Fetch instruction](#fetch) is marked with a ∙, it can be taken eagerly as long as it is not taken infinitely many times. An instance of a non-AMO load instruction, after being fetched, will typically experience the following transitions in this order: 1. [Register read](#reg%5Fread) 2. [Initiate memory load operations](#initiate%5Fload) 3. [Satisfy memory load operation by forwarding from unpropagated stores](#sat%5Fby%5Fforwarding) and/or [Satisfy memory load operation from memory](#sat%5Ffrom%5Fmem) (as many as needed to satisfy all the load operations of the instance) 4. [Complete load operations](#complete%5Floads) 5. [Register write](#reg%5Fwrite) 6. [Finish instruction](#finish) Before, between, and after the transitions above, any number of[Pseudocode internal step](#sail%5Finterp) transitions may appear. In addition, a [Fetch instruction](#fetch) transition for fetching the instruction in the next program location will be available until it is taken. This concludes the informal description of the operational model. The following sections describe the formal operational model. #### [](#pseudocode%5Fexec)Intra-instruction Pseudocode Execution The intra-instruction semantics for each instruction instance is expressed as a state machine, essentially running the instruction pseudocode. Given a pseudocode execution state, it computes the next state. Most states identify a pending memory or register operation, requested by the pseudocode, which the memory model has to do. The states are (this is a tagged union; tags in small-caps): | Load\_mem(_kind_, _address_, _size_, _load\_continuation_) | \- memory load operation | | ---------------------------------------------------------- | --------------------------------- | | Early\_sc\_fail(_res\_continuation_) | \- allow sc to fail early | | Store\_ea(_kind_, _address_, _size_, _next\_state_) | \- memory store effective address | | Store\_memv(_mem\_value_, _store\_continuation_) | \- memory store value | | Fence(_kind_, _next\_state_) | \- fence | | Read\_reg(_reg\_name_, _read\_continuation_) | \- register read | | Write\_reg(_reg\_name_, _reg\_value_, _next\_state_) | \- register write | | Internal(_next\_state_) | \- pseudocode internal step | | Done | \- end of pseudocode | Here: * _mem\_value_ and _reg\_value_ are lists of bytes; * _address_ is an integer of XLEN bits; for load/store, _kind_ identifies whether it is `lr/sc`, acquire-RCpc/release-RCpc, acquire-RCsc/release-RCsc, acquire-release-RCsc; \* for fence, _kind_ identifies whether it is a normal or TSO, and (for normal fences) the predecessor and successor ordering bits; \* _reg\_name_ identifies a register and a slice thereof (start and end bit indices); and the continuations describe how the instruction instance will continue for each value that might be provided by the surrounding memory model (the _load\_continuation_ and _read\_continuation_ take the value loaded from memory and read from the previous register write, the_store\_continuation_ takes _false_ for an `sc` that failed and _true_ in all other cases, and _res\_continuation_ takes _false_ if the `sc` fails and _true_ otherwise). | | For example, given the load instruction lw x1,0(x2), an execution will typically go as follows. The initial execution state will be computed from the pseudocode for the given opcode. This can be expected to be Read\_reg(x2, _read\_continuation_). Feeding the most recently written value of register x2 (the instruction semantics will be blocked if necessary until the register value is available), say 0x4000, to_read\_continuation_ returns Load\_mem(plain\_load, 0x4000, 4,_load\_continuation_). Feeding the 4-byte value loaded from memory location 0x4000, say 0x42, to _load\_continuation_ returns Write\_reg(x1, 0x42, Done). Many Internal(_next\_state_) states may appear before and between the states above. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Notice that writing to memory is split into two steps, Store\_ea and Store\_memv: the first one makes the memory footprint of the store provisionally known, and the second one adds the value to be stored. We ensure these are paired in the pseudocode (Store\_ea followed by Store\_memv), but there may be other steps between them. | | It is observable that the Store\_ea can occur before the value to be stored is determined. For example, for the litmus test LB+fence.r.rw+data-po to be allowed by the operational model (as it is by RVWMO), the first store in Hart 1 has to take the Store\_ea step before its value is determined, so that the second store can see it is to a non-overlapping memory footprint, allowing the second store to be committed out of order without violating coherence. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The pseudocode of each instruction performs at most one store or one load, except for AMOs that perform exactly one load and one store. Those memory accesses are then split apart into the architecturally atomic units by the hart semantics (see [Initiate memory load operations](#initiate%5Fload) and [Initiate memory store operation footprints](#initiate%5Fstore%5Ffootprint) below). Informally, each bit of a register read should be satisfied from a register write by the most recent (in program order) instruction instance that can write that bit (or from the hart’s initial register state if there is no such write). Hence, it is essential to know the register write footprint of each instruction instance, which we calculate when the instruction instance is created (see the [Fetch instruction](#fetch) action of below). We ensure in the pseudocode that each instruction does at most one register write to each register bit, and also that it does not try to read a register value it just wrote. Data-flow dependencies (address and data) in the model emerge from the fact that each register read has to wait for the appropriate register write to be executed (as described above). #### [](#inst%5Fstate)Instruction Instance State Each instruction instance _\_i_ has a state comprising: * _program\_loc_, the memory address from which the instruction was fetched; * _instruction\_kind_, identifying whether this is a load, store, AMO, fence, branch/jump or a `simple` instruction (this also includes a_kind_ similar to the one described for the pseudocode execution states); * _src\_regs_, the set of source \_reg\_name\_s (including system registers), as statically determined from the pseudocode of the instruction; * _dst\_regs_, the destination \_reg\_name\_s (including system registers), as statically determined from the pseudocode of the instruction; * _pseudocode\_state_ (or sometimes just `state` for short), one of (this is a tagged union; tags in small-caps): | Plain(_isa\_state_) | \- ready to make a pseudocode transition | | ------------------------------------------- | ---------------------------------------- | | Pending\_mem\_loads(_load\_continuation_) | \- requesting memory load operation(s) | | Pending\_mem\_stores(_store\_continuation_) | \- requesting memory store operation(s) | * _reg\_reads_, the register reads the instance has performed, including, for each one, the register write slices it read from; * _reg\_writes_, the register writes the instance has performed; * _mem\_loads_, a set of memory load operations, and for each one the as-yet-unsatisfied slices (the byte indices that have not been satisfied yet), and, for the satisfied slices, the store slices (each consisting of a memory store operation and subset of its byte indices) that satisfied it. * _mem\_stores_, a set of memory store operations, and for each one a flag that indicates whether it has been propagated (passed to the shared memory) or not. * information recording whether the instance is committed, finished, etc. Each memory load operation includes a memory footprint (address and size). Each memory store operations includes a memory footprint, and, when available, a value. A load instruction instance with a non-empty _mem\_loads_, for which all the load operations are satisfied (i.e. there are no unsatisfied load slices) is said to be _entirely satisfied_. Informally, an instruction instance is said to have _fully determined data_ if the load (and `sc`) instructions feeding its source registers are finished. Similarly, it is said to have a _fully determined memory footprint_ if the load (and `sc`) instructions feeding its memory operation address register are finished. Formally, we first define the notion of _fully determined register write_: a register write_w_ from _reg\_writes_ of instruction instance_i_ is said to be _fully determined_ if one of the following conditions hold: 1. _i_ is finished; or 2. the value written by _w_ is not affected by a memory operation that _i_ has made (i.e. a value loaded from memory or the result of `sc`), and, for every register read that_i_ has made, that affects _w_, the register write from which _i_ read is fully determined (or_i_ read from the initial register state). Now, an instruction instance _i_ is said to have _fully determined data_ if for every register read _r_ from_reg\_reads_, the register writes that _r_ reads from are fully determined. An instruction instance _i_ is said to have a _fully determined memory footprint_ if for every register read_r_ from _reg\_reads_ that feeds into _i_’s memory operation address, the register writes that _r_ reads from are fully determined. | | The rmem tool records, for every register write, the set of register writes from other instructions that have been read by this instruction at the point of performing the write. By carefully arranging the pseudocode of the instructions covered by the tool we were able to make it so that this is exactly the set of register writes on which the write depends on. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#hart-state)Hart State The model state of a single hart comprises: * _hart\_id_, a unique identifier of the hart; * _initial\_register\_state_, the initial register value for each register; * _initial\_fetch\_address_, the initial instruction fetch address; * _instruction\_tree_, a tree of the instruction instances that have been fetched (and not discarded), in program order. #### [](#shared-memory-state)Shared Memory State The model state of the shared memory comprises a list of memory store operations, in the order they propagated to the shared memory. When a store operation is propagated to the shared memory it is simply added to the end of the list. When a load operation is satisfied from memory, for each byte of the load operation, the most recent corresponding store slice is returned. | | For most purposes, it is simpler to think of the shared memory as an array, i.e., a map from memory locations to memory store operation slices, where each memory location is mapped to a one-byte slice of the most recent memory store operation to that location. However, this abstraction is not detailed enough to properly handle the scinstruction. The RVWMO allows store operations from the same hart as thesc to intervene between the store operation of the sc and the store operations the paired lr read from. To allow such store operations to intervene, and forbid others, the array abstraction must be extended to record more information. Here, we use a list as it is very simple, but a more efficient and scalable implementations should probably use something better. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#transitions)Transitions Each of the paragraphs below describes a single kind of system transition. The description starts with a condition over the current system state. The transition can be taken in the current state only if the condition is satisfied. The condition is followed by an action that is applied to that state when the transition is taken, in order to generate the new system state. ##### [](#fetch)Fetch instruction A possible program-order-successor of instruction instance_i_ can be fetched from address _loc_ if: 1. it has not already been fetched, i.e., none of the immediate successors of _i_ in the hart’s _instruction\_tree_ are from_loc_; and 2. if _i_’s pseudocode has already written an address to_pc_, then _loc_ must be that address, otherwise _loc_ is: * for a conditional branch, the successor address or the branch target address; * for a (direct) jump and link instruction (`jal`), the target address; * for an indirect jump instruction (`jalr`), any address; and * for any other instruction, _i.program\_loc_+4. Action: construct a freshly initialized instruction instance_i'_ for the instruction in the program memory at _loc_, with state Plain(_isa\_state_), computed from the instruction pseudocode, including the static information available from the pseudocode such as its _instruction\_kind_, _src\_regs_, and _dst\_regs_, and add_i'_ to the hart’s _instruction\_tree_ as a successor of_i_. The possible next fetch addresses (_loc_) are available immediately after fetching _i_ and the model does not need to wait for the pseudocode to write to _pc_; this allows out-of-order execution, and speculation past conditional branches and jumps. For most instructions these addresses are easily obtained from the instruction pseudocode. The only exception to that is the indirect jump instruction (`jalr`), where the address depends on the value held in a register. In principle the mathematical model should allow speculation to arbitrary addresses here. The exhaustive search in the `rmem` tool handles this by running the exhaustive search multiple times with a growing set of possible next fetch addresses for each indirect jump. The initial search uses empty sets, hence there is no fetch after indirect jump instruction until the pseudocode of the instruction writes to _pc_, and then we use that value for fetching the next instruction. Before starting the next iteration of exhaustive search, we collect for each indirect jump (grouped by code location) the set of values it wrote to _pc_ in all the executions in the previous search iteration, and use that as possible next fetch addresses of the instruction. This process terminates when no new fetch addresses are detected. ##### [](#initiate%5Fload)Initiate memory load operations An instruction instance _i_ in state Plain(Load\_mem(_kind_,_address_, _size_, _load\_continuation_)) can always initiate the corresponding memory load operations. Action: 1. Construct the appropriate memory load operations _mlos_: * if _address_ is aligned to _size_ then _mlos_ is a single memory load operation of _size_ bytes from _address_; * otherwise, _mlos_ is a set of _size_ memory load operations, each of one byte, from the addresses_address_…​_address_+_size_−1. 2. set _mem\_loads_ of _i_ to _mlos_; and 3. update the state of _i_ to Pending\_mem\_loads(_load\_continuation_). In [Memory Model Primitives](rvwmo.html#rvwmo-primitives) it is said that misaligned memory accesses may be decomposed at any granularity. Here we decompose them to one-byte accesses as this granularity subsumes all others. ##### [](#sat%5Fby%5Fforwarding)Satisfy memory load operation by forwarding from unpropagated stores For a non-AMO load instruction instance _i_ in state Pending\_mem\_loads(_load\_continuation_), and a memory load operation_mlo_ in _i.mem\_loads_ that has unsatisfied slices, the memory load operation can be partially or entirely satisfied by forwarding from unpropagated memory store operations by store instruction instances that are program-order-before_i_ if: 1. all program-order-previous `fence` instructions with `.sr` and `.pw`set are finished; 2. for every program-order-previous `fence` instruction, _f_, with `.sr` and `.pr` set, and `.pw` not set, if _f_ is not finished then all load instructions that are program-order-before_f_ are entirely satisfied; 3. for every program-order-previous `fence.tso` instruction,_f_, that is not finished, all load instructions that are program-order-before _f_ are entirely satisfied; 4. if _i_ is a load-acquire-RCsc, all program-order-previous store-releases-RCsc are finished; 5. if _i_ is a load-acquire-release, all program-order-previous instructions are finished; 6. all non-finished program-order-previous load-acquire instructions are entirely satisfied; and 7. all program-order-previous store-acquire-release instructions are finished; Let _msoss_ be the set of all unpropagated memory store operation slices from non-`sc` store instruction instances that are program-order-before _i_ and have already calculated the value to be stored, that overlap with the unsatisfied slices of_mlo_, and which are not superseded by intervening store operations or store operations that are read from by an intervening load. The last condition requires, for each memory store operation slice_msos_ in _msoss_ from instruction_i'_: * that there is no store instruction program-order-between _i_and _i'_ with a memory store operation overlapping_msos_; and * that there is no load instruction program-order-between _i_and _i'_ that was satisfied from an overlapping memory store operation slice from a different hart. Action: 1. update _i.mem\_loads_ to indicate that_mlo_ was satisfied by _msoss_; and 2. restart any speculative instructions which have violated coherence as a result of this, i.e., for every non-finished instruction_i'_ that is a program-order-successor of _i_, and every memory load operation _mlo'_ of _i'_that was satisfied from _msoss'_, if there exists a memory store operation slice _msos'_ in _msoss'_, and an overlapping memory store operation slice from a different memory store operation in _msoss_, and _msos'_ is not from an instruction that is a program-order-successor of_i_, restart _i'_ and its _restart-dependents_. Where, the _restart-dependents_ of instruction _j_ are: * program-order-successors of _j_ that have data-flow dependency on a register write of _j_; * program-order-successors of _j_ that have a memory load operation that reads from a memory store operation of _j_(by forwarding); * if _j_ is a load-acquire, all the program-order-successors of _j_; * if _j_ is a load, for every `fence`, _f_, with`.sr` and `.pr` set, and `.pw` not set, that is a program-order-successor of _j_, all the load instructions that are program-order-successors of _f_; * if _j_ is a load, for every `fence.tso`, _f_, that is a program-order-successor of _j_, all the load instructions that are program-order-successors of _f_; and * (recursively) all the restart-dependents of all the instruction instances above. Forwarding memory store operations to a memory load might satisfy only some slices of the load, leaving other slices unsatisfied. A program-order-previous store operation that was not available when taking the transition above might make _msoss_ provisionally unsound (violating coherence) when it becomes available. That store will prevent the load from being finished (see [Finish instruction](#finish)), and will cause it to restart when that store operation is propagated (see [Propagate store operation](#prop%5Fstore)). A consequence of the transition condition above is that store-release-RCsc memory store operations cannot be forwarded to load-acquire-RCsc instructions: _msoss_ does not include memory store operations from finished stores (as those must be propagated memory store operations), and the condition above requires all program-order-previous store-releases-RCsc to be finished when the load is acquire-RCsc. ##### [](#sat%5Ffrom%5Fmem)Satisfy memory load operation from memory For an instruction instance _i_ of a non-AMO load instruction or an AMO instruction in the context of the [Satisfy, commit and propagate operations of an AMO](#do%5Famo) transition, any memory load operation _mlo_ in_i.mem\_loads_ that has unsatisfied slices, can be satisfied from memory if all the conditions of > are satisfied. Action: let _msoss_ be the memory store operation slices from memory covering the unsatisfied slices of _mlo_, and apply the action of [Satisfy memory operation by forwarding from unpropagates stores](#do%5Famo). | | Note that [Satisfy memory operation by forwarding from unpropagates stores](#do%5Famo) might leave some slices of the memory load operation unsatisfied, those will have to be satisfied by taking the transition again, or taking [Satisfy memory load operation from memory](#sat%5Ffrom%5Fmem). [Satisfy memory load operation from memory](#sat%5Ffrom%5Fmem), on the other hand, will always satisfy all the unsatisfied slices of the memory load operation. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#complete%5Floads)Complete load operations A load instruction instance _i_ in state Pending\_mem\_loads(_load\_continuation_) can be completed (not to be confused with finished) if all the memory load operations_i.mem\_loads_ are entirely satisfied (i.e. there are no unsatisfied slices). Action: update the state of _i_to Plain(_load\_continuation(mem\_value)_), where _mem\_value_ is assembled from all the memory store operation slices that satisfied_i.mem\_loads_. ##### [](#early%5Fsc%5Ffail)Early `sc` fail An `sc` instruction instance _i_ in state Plain(Early\_sc\_fail(_res\_continuation_)) can always be made to fail. Action: update the state of _i_ to Plain(_res\_continuation(false)_). ##### [](#paired%5Fsc)Paired `sc` An `sc` instruction instance _i_ in state Plain(Early\_sc\_fail(_res\_continuation_)) can continue its (potentially successful) execution if _i_ is paired with an `lr`. Action: update the state of _i_ to Plain(_res\_continuation(true)_). ##### [](#initiate%5Fstore%5Ffootprint)Initiate memory store operation footprints An instruction instance _i_ in state Plain(Store\_ea(_kind_,_address_, _size_, _next\_state_)) can always announce its pending memory store operation footprint. Action: 1. construct the appropriate memory store operations _msos_(without the store value): * if _address_ is aligned to _size_ then _msos_ is a single memory store operation of _size_ bytes to _address_; * otherwise, _msos_ is a set of _size_ memory store operations, each of one-byte size, to the addresses_address_…​_address_+_size_−1. 2. set _i.mem\_stores_ to _msos_; and 3. update the state of _i_ to Plain(_next\_state_). Note that after taking the transition above the memory store operations do not yet have their values. The importance of splitting this transition from the transition below is that it allows other program-order-successor store instructions to observe the memory footprint of this instruction, and if they don’t overlap, propagate out of order as early as possible (i.e. before the data register value becomes available). ##### [](#instantiate%5Fstore%5Fvalue)Instantiate memory store operation values An instruction instance _i_ in state Plain(Store\_memv(_mem\_value_, _store\_continuation_)) can always instantiate the values of the memory store operations_i.mem\_stores_. Action: 1. split _mem\_value_ between the memory store operations_i.mem\_stores_; and 2. update the state of _i_ to Pending\_mem\_stores(_store\_continuation_). ##### [](#commit%5Fstores)Commit store instruction An uncommitted instruction instance _i_ of a non-`sc` store instruction or an `sc` instruction in the context of the [Commit and propagate store operation of an sc](#commit%5Fsc)transition, in state Pending\_mem\_stores(_store\_continuation_), can be committed (not to be confused with propagated) if: 1. _i_ has fully determined data; 2. all program-order-previous conditional branch and indirect jump instructions are finished; 3. all program-order-previous `fence` instructions with `.sw` set are finished; 4. all program-order-previous `fence.tso` instructions are finished; 5. all program-order-previous load-acquire instructions are finished; 6. all program-order-previous store-acquire-release instructions are finished; 7. if _i_ is a store-release, all program-order-previous instructions are finished; 8. all program-order-previous memory access instructions have a fully determined memory footprint; 9. all program-order-previous store instructions, except for `sc` that failed, have initiated and so have non-empty _mem\_stores_; and 10. all program-order-previous load instructions have initiated and so have non-empty _mem\_loads_. Action: record that _i_ is committed. | | Notice that if condition[8](#commit%5Fstores) is satisfied the conditions[9](#commit%5Fstores) and[10](#commit%5Fstores) are also satisfied, or will be satisfied after taking some eager transitions. Hence, requiring them does not strengthen the model. By requiring them, we guarantee that previous memory access instructions have taken enough transitions to make their memory operations visible for the condition check of , which is the next transition the instruction will take, making that condition simpler. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#prop%5Fstore)Propagate store operation For a committed instruction instance _i_ in state Pending\_mem\_stores(_store\_continuation_), and an unpropagated memory store operation _mso_ in_i.mem\_stores_, _mso_ can be propagated if: 1. all memory store operations of program-order-previous store instructions that overlap with _mso_ have already propagated; 2. all memory load operations of program-order-previous load instructions that overlap with _mso_ have already been satisfied, and (the load instructions) are _non-restartable_ (see definition below); and 3. all memory load operations that were satisfied by forwarding_mso_ are entirely satisfied. Where a non-finished instruction instance _j_ is_non-restartable_ if: 1. there does not exist a store instruction _s_ and an unpropagated memory store operation _mso_ of _s_such that applying the action of the [Propagate store operation](#prop%5Fstore) transition to_mso_ will result in the restart of _j_; and 2. there does not exist a non-finished load instruction _l_and a memory load operation _mlo_ of _l_ such that applying the action of the [Satisfy memory load operation by forwarding from unpropagated stores](#sat%5Fby%5Fforwarding)/[Satisfy memory load operation from memory](#sat%5Ffrom%5Fmem) transition (even if_mlo_ is already satisfied) to _mlo_ will result in the restart of _j_. Action: 1. update the shared memory state with _mso_; 2. update _i.mem\_stores_ to indicate that_mso_ was propagated; and 3. restart any speculative instructions which have violated coherence as a result of this, i.e., for every non-finished instruction_i'_ program-order-after _i_ and every memory load operation _mlo'_ of _i'_ that was satisfied from _msoss'_, if there exists a memory store operation slice _msos'_ in _msoss'_ that overlaps with_mso_ and is not from _mso_, and_msos'_ is not from a program-order-successor of_i_, restart _i'_ and its _restart-dependents_(see [Satisfy memory load operation by forwarding from unpropagated stores](#sat%5Fby%5Fforwarding)). ##### [](#commit%5Fsc)Commit and propagate store operation of an `sc` An uncommitted `sc` instruction instance _i_, from hart_h_, in state Pending\_mem\_stores(_store\_continuation_), with a paired `lr` _i'_ that has been satisfied by some store slices _msoss_, can be committed and propagated at the same time if: 1. _i'_ is finished; 2. every memory store operation that has been forwarded to_i'_ is propagated; 3. the conditions of [Commit store instruction](#commit%5Fstores) is satisfied; 4. the conditions of [Propagate store instruction](#prop%5Fstore) is satisfied (notice that an `sc` instruction can only have one memory store operation); and 5. for every store slice _msos_ from _msoss_,_msos_ has not been overwritten, in the shared memory, by a store that is from a hart that is not _h_, at any point since _msos_ was propagated to memory. Action: 1. apply the actions of [Commit store instruction](#commit%5Fstores); and 2. apply the action of [Propagate store instruction](#prop%5Fstore). ##### [](#late%5Fsc%5Ffail)Late `sc` fail An `sc` instruction instance _i_ in state Pending\_mem\_stores(_store\_continuation_), that has not propagated its memory store operation, can always be made to fail. Action: 1. clear _i.mem\_stores_; and 2. update the state of _i_ to Plain(_store\_continuation(false)_). For efficiency, the `rmem` tool allows this transition only when it is not possible to take the [Commit and propagate store operation of an sc](#commit%5Fsc) transition. This does not affect the set of allowed final states, but when explored interactively, if the `sc`should fail one should use the [Early sc fail](#early%5Fsc%5Ffail) transition instead of waiting for this transition. ##### [](#complete%5Fstores)Complete store operations A store instruction instance _i_ in state Pending\_mem\_stores(_store\_continuation_), for which all the memory store operations in _i.mem\_stores_ have been propagated, can always be completed (not to be confused with finished). Action: update the state of _i_ to Plain(_store\_continuation(true)_). ##### [](#do%5Famo)Satisfy, commit and propagate operations of an AMO An AMO instruction instance _i_ in state Pending\_mem\_loads(_load\_continuation_) can perform its memory access if it is possible to perform the following sequence of transitions with no intervening transitions: 1. [Satisfy memory load operation from memory](#sat%5Ffrom%5Fmem) 2. [Complere load operations](#complete%5Floads) 3. [Pseudocode internal step](#sail%5Finterp) (zero or more times) 4. [Instantiate memory store operation values](#instantiate%5Fstore%5Fvalue) 5. [Commit store instruction](#commit%5Fstores) 6. [Propagate store operation](#prop%5Fstore) 7. [Complete store operations](#complete%5Fstores) and in addition, the condition of [Finish instruction](#finish), with the exception of not requiring_i_ to be in state Plain(Done), holds after those transitions. Action: perform the above sequence of transitions (this does not include [Finish instruction](#finish)), one after the other, with no intervening transitions. | | Notice that program-order-previous stores cannot be forwarded to the load of an AMO. This is simply because the sequence of transitions above does not include the forwarding transition. But even if it did include it, the sequence will fail when trying to do the [Propagate store operation](#prop%5Fstore) transition, as this transition requires all program-order-previous store operations to overlapping memory footprints to be propagated, and forwarding requires the store operation to be unpropagated. In addition, the store of an AMO cannot be forwarded to a program-order-successor load. Before taking the transition above, the store operation of the AMO does not have its value and therefore cannot be forwarded; after taking the transition above the store operation is propagated and therefore cannot be forwarded. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#commit%5Ffence)Commit fence A fence instruction instance _i_ in state Plain(Fence(_kind_, _next\_state_)) can be committed if: 1. if _i_ is a normal fence and it has `.pr` set, all program-order-previous load instructions are finished; 2. if _i_ is a normal fence and it has `.pw` set, all program-order-previous store instructions are finished; and 3. if _i_ is a `fence.tso`, all program-order-previous load and store instructions are finished. Action: 1. record that _i_ is committed; and 2. update the state of _i_ to Plain(_next\_state_). ##### [](#reg%5Fread)Register read An instruction instance _i_ in state Plain(Read\_reg(_reg\_name_, _read\_cont_)) can do a register read of_reg\_name_ if every instruction instance that it needs to read from has already performed the expected _reg\_name_ register write. Let _read\_sources_ include, for each bit of _reg\_name_, the write to that bit by the most recent (in program order) instruction instance that can write to that bit, if any. If there is no such instruction, the source is the initial register value from _initial\_register\_state_. Let_reg\_value_ be the value assembled from _read\_sources_. Action: 1. add _reg\_name_ to _i.reg\_reads_ with_read\_sources_ and _reg\_value_; and 2. update the state of _i_ to Plain(_read\_cont(reg\_value)_). ##### [](#reg%5Fwrite)Register write An instruction instance _i_ in state Plain(Write\_reg(_reg\_name_, _reg\_value_, _next\_state_)) can always do a_reg\_name_ register write. Action: 1. add _reg\_name_ to _i.reg\_writes_ with_deps_ and _reg\_value_; and 2. update the state of _i_ to Plain(_next\_state_). where _deps_ is a pair of the set of all _read\_sources_ from_i.reg\_reads_, and a flag that is true iff_i_ is a load instruction instance that has already been entirely satisfied. ##### [](#sail%5Finterp)Pseudocode internal step An instruction instance _i_ in state Plain(Internal(_next\_state_)) can always do that pseudocode-internal step. Action: update the state of _i_ to Plain(_next\_state_). ##### [](#finish)Finish instruction A non-finished instruction instance _i_ in state Plain(Done) can be finished if: 1. if _i_ is a load instruction: 1. all program-order-previous load-acquire instructions are finished; 2. all program-order-previous `fence` instructions with `.sr` set are finished; 3. for every program-order-previous `fence.tso` instruction,_f_, that is not finished, all load instructions that are program-order-before _f_ are finished; and 4. it is guaranteed that the values read by the memory load operations of _i_ will not cause coherence violations, i.e., for any program-order-previous instruction instance _i'_, let_cfp_ be the combined footprint of propagated memory store operations from store instructions program-order-between_i_ and _i'_, and _fixed memory store operations_ that were forwarded to _i_ from store instructions program-order-between _i_ and _i'_including _i'_, and let_/cfp_ be the complement of_cfp_ in the memory footprint of _i_. If _/cfp_ is not empty: 1. _i'_ has a fully determined memory footprint; 2. _i'_ has no unpropagated memory store operations that overlap with _/cfp_; and 3. if _i'_ is a load with a memory footprint that overlaps with _/cfp_, then all the memory load operations of _i'_ that overlap with_/cfp_ are satisfied and _i'_is _non-restartable_ (see the [Propagate store operation](#prop%5Fstore) transition for how to determined if an instruction is non-restartable). Here, a memory store operation is called fixed if the store instruction has fully determined data. 2. _i_ has a fully determined data; and 3. if _i_ is not a fence, all program-order-previous conditional branch and indirect jump instructions are finished. Action: 1. if _i_ is a conditional branch or indirect jump instruction, discard any untaken paths of execution, i.e., remove all instruction instances that are not reachable by the branch/jump taken in_instruction\_tree_; and 2. record the instruction as finished, i.e., set _finished_ to _true_. #### [](#limitations)Limitations * The model covers user-level RV64I and RV64A. In particular, it does not support the misaligned atomicity granule PMA or the total store ordering extension "Ztso". It should be trivial to adapt the model to RV32I/A and to the G, Q and C extensions, but we have never tried it. This will involve, mostly, writing Sail code for the instructions, with minimal, if any, changes to the concurrency model. * The model covers only normal memory accesses (it does not handle I/O accesses). * The model does not cover TLB-related effects. * The model assumes the instruction memory is fixed. In particular, the[Fetch instruction](#fetch) transition does not generate memory load operations, and the shared memory is not involved in the transition. Instead, the model depends on an external oracle that provides an opcode when given a memory location. * The model does not cover exceptions, traps and interrupts. 37.1. ISA Extension Naming Conventions ==================== ## [](#naming)37.1\. ISA Extension Naming Conventions This chapter describes the RISC-V ISA extension naming scheme that is used to concisely describe the set of instructions present in a hardware implementation, or the set of instructions used by an application binary interface (ABI). | | The RISC-V ISA is designed to support a wide variety of implementations with various experimental instruction-set extensions. We have found that an organized naming scheme simplifies software tools and documentation. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#37-1-1-case-sensitivity)37.1.1\. Case Sensitivity The ISA naming strings are case insensitive. ### [](#37-1-2-base-integer-isa)37.1.2\. Base Integer ISA RISC-V ISA strings begin with either RV32I, RV32E, RV64I, or RV64E, indicating the supported address space size in bits for the base integer ISA. ### [](#37-1-3-instruction-set-extension-names)37.1.3\. Instruction-Set Extension Names Standard ISA extensions are given a name consisting of a single letter. For example, the first four standard extensions to the integer bases are: "M" for integer multiplication and division, "A" for atomic memory instructions, "F" for single-precision floating-point instructions, and "D" for double-precision floating-point instructions. Any RISC-V instruction-set variant can be succinctly described by concatenating the base integer prefix with the names of the included extensions, e.g., "RV64IMAFD". We have also defined an abbreviation "G" to represent the "IMAFDZicsr\_Zifencei" base and extensions, as this is intended to represent our standard general-purpose ISA. Standard extensions to the RISC-V ISA are given other reserved letters, e.g., "Q" for quad-precision floating-point, or "C" for the 16-bit compressed instruction format. Some ISA extensions depend on the presence of other extensions, e.g., "D" depends on "F" and "F" depends on "Zicsr". These dependencies may be implicit in the ISA name: for example, RV32IF is equivalent to RV32IFZicsr, and RV32ID is equivalent to RV32IFD and RV32IFDZicsr. ### [](#37-1-4-underscores)37.1.4\. Underscores Underscores "\_" may be used to separate ISA extensions to improve readability and to provide disambiguation, e.g., "RV32I2\_M2\_A2". ### [](#37-1-5-additional-standard-unprivileged-extension-names)37.1.5\. Additional Standard Unprivileged Extension Names Standard unprivileged extensions can also be named by using a single "Z" followed by an alphanumeric name. The name must end with an alphabetical character. The second letter from the end cannot be numeric if the last letter is "p". For example, "Zifencei" names the instruction-fetch fence extension described in["Zifencei" Extension for Instruction-Fetch Fence](zifencei.html); "Zifencei2" and "Zifencei2p0" name version 2.0 of same. The first letter following the "Z" conventionally indicates the most closely related alphabetical extension category, IMAFDQLCBKJTPVH. For the "Zfa" extension for additional floating-point instructions, for example, the letter "f" indicates the extension is related to the "F" standard extension. If multiple "Z" extensions are named, they should be ordered first by category, then alphabetically within a category—for example, "Zicsr\_Zifencei\_Ztso". All multi-letter extensions, including those with the "Z" prefix, must be separated from other multi-letter extensions by an underscore, e.g., "RV32IMACZicsr\_Zifencei". ### [](#37-1-6-supervisor-level-instruction-set-extension-names)37.1.6\. Supervisor-level Instruction-Set Extension Names Standard extensions that extend the supervisor-level virtual-memory architecture are prefixed with the letters "Sv", followed by an alphanumeric name. Other standard extensions that extend the supervisor-level architecture are prefixed with the letters "Ss", followed by an alphanumeric name. The name must end with an alphabetical character. The second letter from the end cannot be numeric if the last letter is "p". These extensions are further defined in Volume II. The extensions "sv32", "sv39", "sv48", and "sv59" were defined before the rule against extension names ending in numbers was established. Standard supervisor-level extensions should be listed after standard unprivileged extensions, and like other multi-letter extensions, must be separated from other multi-letter extensions by an underscore. If multiple supervisor-level extensions are listed, they should be ordered alphabetically. ### [](#37-1-7-hypervisor-level-instruction-set-extension-names)37.1.7\. Hypervisor-level Instruction-Set Extension Names Standard extensions that extend the hypervisor-level architecture are prefixed with the letters "Sh". If multiple hypervisor-level extensions are listed, they should be ordered alphabetically. | | Many augmentations to the hypervisor-level architecture are more naturally defined as supervisor-level extensions, following the scheme described in the previous section. The "Sh" prefix is used by the few hypervisor-level extensions that have no supervisor-visible effects. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#37-1-8-machine-level-instruction-set-extension-names)37.1.8\. Machine-level Instruction-Set Extension Names Standard machine-level instruction-set extensions are prefixed with the letters "Sm". Standard machine-level extensions should be listed after standard lesser-privileged extensions, and like other multi-letter extensions, must be separated from other multi-letter extensions by an underscore. If multiple machine-level extensions are listed, they should be ordered alphabetically. ### [](#37-1-9-non-standard-extension-names)37.1.9\. Non-Standard Extension Names Non-standard extensions are named by using a single "X" followed by the alphanumeric name. The name must end with an alphabetic character. The second letter from the end cannot be numeric if the last letter is "p". For example, "Xhwacha" names the Hwacha vector-fetch ISA extension. Non-standard extensions must be listed after all standard extensions, and, like other multi-letter extensions, must be separated from other multi-letter extensions by an underscore. For example, an ISA with non-standard extensions Argle and Bargle may be named "RV64IZifencei\_Xargle\_Xbargle". If multiple non-standard extensions are listed, they should be ordered alphabetically. Like other multi-letter extensions, they should be separated from other multi-letter extensions by an underscore. ### [](#37-1-10-version-numbers)37.1.10\. Version Numbers Recognizing that instruction sets may expand or alter over time, we encode extension version numbers following the extension name. Version numbers are divided into major and minor version numbers, separated by a "p". If the minor version is "0", then "p0" can be omitted from the version string. To avoid ambiguity, no extension name may end with a number or a "p" preceded by a number. Because the "P" extension for Packed SIMD can be confused for the decimal point in a version number, it must be preceded by an underscore if it follows another extension with a version number. For example, "rv32i2p2" means version 2.2 of RV32I, whereas "rv32i2\_p2" means version 2.0 of RV32I with version 2.0 of the P extension. Changes in major version numbers imply a loss of backwards compatibility, whereas changes in only the minor version number must be backwards-compatible. For example, the original 64-bit standard ISA defined in release 1.0 of this manual can be written in full as "RV64I1p0M1p0A1p0F1p0D1p0", more concisely as "RV64I1M1A1F1D1". We introduced the version numbering scheme with the second release. Hence, we define the default version of a standard extension to be the version present at that time, e.g., "RV32I" is equivalent to "RV32I2". ### [](#37-1-11-subset-naming-convention)37.1.11\. Subset Naming Convention [Table 1](#isanametable) summarizes the standardized extension names. The table also defines the canonical order in which extension names must appear in the name string, with top-to-bottom in table indicating first-to-last in the name string, e.g., RV32IMACV is legal, whereas RV32IMAVC is not. __Table 1\. Standard ISA extension names.__ | Subset | Name | Implies | | ------------------------------------------------- | ----- | -------------------- | | Base ISA | | | | Integer | I | | | Reduced Integer | E | | | **Standard Unprivileged Extensions** | | | | Integer Multiplication and Division | M | Zmmul | | Atomics | A | | | Single-Precision Floating-Point | F | Zicsr | | Double-Precision Floating-Point | D | F | | General | G | IMAFDZicsr\_Zifencei | | Quad-Precision Floating-Point | Q | D | | 16-bit Compressed Instructions | C | | | B Extension | B | | | Packed-SIMD Extensions | P | | | Vector Extension | V | D | | Hypervisor Extension | H | | | **Additional Standard Unprivileged Extensions** | | | | Additional Standard unprivileged extensions "abc" | Zabc | | | **Standard Supervisor-Level Extensions** | | | | Supervisor-level extension "def" | Ssdef | | | **Standard Hypervisor-Level Extensions** | | | | Hypervisor-level extension "ghi" | Shghi | | | **Standard Machine-Level Extensions** | | | | Machine-level extension "jkl" | Smjkl | | | **Non-Standard Extensions** | | | | Non-standard extension "mno" | Xmno | | 23.1. "Q" Extension for Quad-Precision Floating-Point, Version 2.2 ==================== ## [](#23-1-q-extension-for-quad-precision-floating-point-version-2-2)23.1\. "Q" Extension for Quad-Precision Floating-Point, Version 2.2 This chapter describes the Q standard extension for 128-bit quad-precision binary floating-point instructions compliant with the IEEE 754-2008 arithmetic standard. The quad-precision binary floating-point instruction-set extension is named "Q"; it depends on the double-precision floating-point extension D. The floating-point registers are now extended to hold either a single, double, or quad-precision floating-point value (FLEN=128). The NaN-boxing scheme described in [NaN Boxing of Narrower Values](d-st-ext.html#nanboxing) is now extended recursively to allow a single-precision value to be NaN-boxed inside a double-precision value which is itself NaN-boxed inside a quad-precision value. ### [](#23-1-1-quad-precision-load-and-store-instructions)23.1.1\. Quad-Precision Load and Store Instructions New 128-bit variants of LOAD-FP and STORE-FP instructions are added, encoded with a new value for the funct3 width field. ![svg](_images/svg-2a16e859bb5c0fca32e52b3b4f44d98f3e6caa3d.svg) ![svg](_images/svg-2b334c20b68f921d79e5e277822a6e133384456d.svg) FLQ and FSQ are only guaranteed to execute atomically if the effective address is naturally aligned and XLEN=128. FLQ and FSQ do not modify the bits being transferred; in particular, the payloads of non-canonical NaNs are preserved. ### [](#23-1-2-quad-precision-computational-instructions)23.1.2\. Quad-Precision Computational Instructions A new supported format is added to the format field of most instructions, as shown in [Table 1](#fpextfmt) __Table 1\. Format field encoding.__ | _fmt_ field | Mnemonic | Meaning | | ----------- | -------- | ----------------------- | | 00 | S | 32-bit single-precision | | 01 | D | 64-bit double-precision | | 10 | H | 16-bit half-precision | | 11 | Q | 128-bit quad-precision | The quad-precision floating-point computational instructions are defined analogously to their double-precision counterparts, but operate on quad-precision operands and produce quad-precision results. ![svg](_images/svg-1a95a7294abe66b7b5523af0b0b2943db84a86bb.svg) ![svg](_images/svg-eaca638bb8b9f77d2f3d6ef3e1b39c19dc76c841.svg) ### [](#quad-compute)23.1.3\. Quad-Precision Convert and Move Instructions New floating-point-to-integer and integer-to-floating-point conversion instructions are added. These instructions are defined analogously to the double-precision-to-integer and integer-to-double-precision conversion instructions. FCVT.W.Q or FCVT.L.Q converts a quad-precision floating-point number to a signed 32-bit or 64-bit integer, respectively. FCVT.Q.W or FCVT.Q.L converts a 32-bit or 64-bit signed integer, respectively, into a quad-precision floating-point number.FCVT.WU.Q, FCVT.LU.Q, FCVT.Q.WU, and FCVT.Q.LU variants convert to or from unsigned integer values. FCVT.L\[U\].Q and FCVT.Q.L\[U\] are RV64-only instructions. Note FCVT.Q.L\[U\] always produces an exact result and is unaffected by rounding mode. ![svg](_images/svg-d88b4219a55f29d8ac2ed938270389bc162dd354.svg) New floating-point-to-floating-point conversion instructions are added. These instructions are defined analogously to the double-precision floating-point-to-floating-point conversion instructions. FCVT.S.Q or FCVT.Q.S converts a quad-precision floating-point number to a single-precision floating-point number, or vice-versa, respectively.FCVT.D.Q or FCVT.Q.D converts a quad-precision floating-point number to a double-precision floating-point number, or vice-versa, respectively. ![svg](_images/svg-4653d8c6b12c99eaa852fa93984e16156f5c3f31.svg) Floating-point to floating-point sign-injection instructions, FSGNJ.Q, FSGNJN.Q, and FSGNJX.Q are defined analogously to the double-precision sign-injection instruction. ![svg](_images/svg-14b1c2fe50e599c8139fe51f78dc16a9c60fdfca.svg) FMV.X.Q and FMV.Q.X instructions are not provided in RV32 or RV64, so quad-precision bit patterns must be moved to the integer registers via memory. ### [](#23-1-4-quad-precision-floating-point-compare-instructions)23.1.4\. Quad-Precision Floating-Point Compare Instructions The quad-precision floating-point compare instructions are defined analogously to their double-precision counterparts, but operate on quad-precision operands. ![svg](_images/svg-a4cc3c2f3b5bb151c91158a3df3a779fa63bdbdb.svg) ### [](#quad-float-compare)23.1.5\. Quad-Precision Floating-Point Classify Instruction The quad-precision floating-point classify instruction, FCLASS.Q, is defined analogously to its double-precision counterpart, but operates on quad-precision operands. ![svg](_images/svg-0fbdc9079ae6482318595fb8e617b3c7f19cbd69.svg) Historical Rationale for Extensions ==================== ## [](#historical-rationale-for-extensions)Appendix A: Historical Rationale for Extensions This appendix contains the rationale for RISC-V ISA extensions at the time they were ratified. Unlike the ISA specification, this appendix is ordered chronologically, so as to convey the motivation and architectural reasoning underpinning each extension at the time of ratification. For extensions ratified prior to the conception of this appendix (ca. 2025), the rationale will be added over time. In cases where the rationale was not recorded, the authors and editors will synthesize it from the historical record. ### [](#zihintpause-extension-for-pause-hint)"Zihintpause" Extension for Pause Hint The PAUSE instruction hints to a hart that it should temporarily reduce its rate of execution. It is normally used to save energy and execution resources while polling, e.g. while waiting for a spinlock to become free. Much of the debate surrounding this extension centered on whether a facility similar to x86’s MONITOR/MWAIT should instead be provided. We concluded that, even if such a facility were to be defined for RISC-V, it would not supplant PAUSE. PAUSE is more appropriate when polling for non-memory events, when polling for multiple events, or when software does not know precisely what events it is polling for. (Perhaps surprisingly, the latter case is ubiquitous, in part because it is the mechanism expected by the Linux kernel’s `cpu_relax` API.) ### [](#zicond-extension-for-integer-conditional-operations)"Zicond" Extension for Integer Conditional Operations Replacing unpredictable branches with conditional-select or conditional-move instructions can mitigate a class of costly branch mispredictions. Unfortunately, conditional-select instructions require three source operands. These instructions are a logical addition to ISAs that include three-source integer instructions for other reasons, but are too costly otherwise. Some ISAs have instead furnished conditional-move instructions, which consume less encoding space and avoid the extra register read in simple microarchitectures. Unfortunately, in register-renamed microarchitectures, these instructions incur costs simlar to conditional select, or require additional microarchitectural structures and micro-op-issue constraints. The Zicond extension was defined to solve the same problem as conditional select and conditional move, but with very little incremental cost for complex microarchitectures. It provides conditional-zero instructions, which read two source operands and, based upon the zeroness of the second operand, produce either the first operand or zero. These instructions can be used as part of a three-instruction sequence to synthesize conditional select. Several common conditional-execution idioms require only two instructions, as would be the case with conditional select or move, including conditional addition, subtraction, and bitwise AND, OR, and XOR. Two conditional-zero instructions are included: one that writes zero if the comparand is zero, and one that does so if the comparand is nonzero. Variants that perform magnitude comparisons with zero were considered but ultimately excluded for insufficient quantitative justification. ### [](#zacas-extension-for-atomic-compare-and-swap-cas-instructions)"Zacas" Extension for Atomic Compare-and-Swap (CAS) Instructions While compare-and-swap for XLEN wide data may be accomplished using LR/SC, the CAS atomic instructions scale better to highly parallel systems than LR/SC. Many lock-free algorithms, such as a lock-free queue, require manipulation of pointer variables. A simple CAS operation may not be sufficient to guard against what is commonly referred to as the ABA problem in such algorithms that manipulate pointer variables. To avoid the ABA problem, the algorithms associate a reference counter with the pointer variable and perform updates using a quadword compare and swap (of both the pointer and the counter). The double and quadword CAS instructions support implementation of algorithms for ABA problem avoidance. The CAS instruction supports the C++11 atomic compare and exchange operation. ### [](#zabha-extension-for-byte-and-halfword-atomic-memory-operations-version-1-0)"Zabha" Extension for Byte and Halfword Atomic Memory Operations, Version 1.0 The A-extension offers atomic memory operation (AMO) instructions for _words_,_doublewords_, and _quadwords_ (only for `AMOCAS`). The absence of atomic operations for subword data types necessitates emulation strategies. For bitwise operations, this emulation can be performed via word-sized bitwise AMO\* instructions. For non-bitwise operations, emulation is achievable using word-sized `LR`/`SC` instructions. Several limitations arise from this emulation approach: 1. In systems with large-scale or Non-Uniform Memory Access (NUMA) configurations, emulation based on `LR`/`SC` introduces issues related to scalability and fairness, particularly under conditions of high contention. 2. Emulation of narrower AMOs through wider AMO\* instructions on non-idempotent IO memory regions may result in unintended side effects. 3. Utilizing wider AMO\* instructions for emulating narrower AMOs risks activating extraneous breakpoints or watchpoints. 4. In the absence of native support for subword atomics, compilers often resort to inlining code sequences to provide the required emulation. This practice contributes to an increase in code size, with consequent impacts on system performance and memory utilization. The Zabha extension addresses these limitations by adding support for _byte_ and_halfword_ atomic memory operations to the RISC-V Unprivileged ISA. 36.1. RV32/64G Instruction Set Listings ==================== ## [](#rv32-64g)36.1\. RV32/64G Instruction Set Listings One goal of the RISC-V project is that it be used as a stable software development target. For this purpose, we define a combination of a base ISA (RV32I or RV64I) plus selected standard extensions (IMAFD, Zicsr, Zifencei) as a "general-purpose" ISA, and we use the abbreviation G for the IMAFDZicsr\_Zifencei combination of instruction-set extensions. This chapter presents opcode maps and instruction-set listings for RV32G and RV64G. __Table 1\. RISC-V base opcode map, inst\[1:0\]=11__ | inst\[4:2\] | 000 | 001 | 010 | 011 | 100 | 101 | 110 | 111 (>32b) | | ----------- | ------ | -------- | ---------- | -------- | ------ | ----- | ---------- | ---------- | | inst\[6:5\] | | | | | | | | | | 00 | LOAD | LOAD-FP | _custom-0_ | MISC-MEM | OP-IMM | AUIPC | OP-IMM-32 | _reserved_ | | 01 | STORE | STORE-FP | _custom-1_ | AMO | OP | LUI | OP-32 | _reserved_ | | 10 | MADD | MSUB | NMSUB | NMADD | OP-FP | OP-V | _custom-2_ | _reserved_ | | 11 | BRANCH | JALR | _reserved_ | JAL | SYSTEM | OP-VE | _custom-3_ | _reserved_ | [Table 1](#opcodemap) shows a map of the major opcodes for RVG. Opcodes marked as _reserved_should be avoided for custom instruction-set extensions as they might be used by future standard extensions. Major opcodes marked as _custom-0_through _custom-3_ will be avoided by future standard extensions and are recommended for use by custom instruction-set extensions within the base 32-bit instruction format. We believe RV32G and RV64G provide simple but complete instruction sets for a broad range of general-purpose computing. The optional compressed instruction set described in ["C" Extension for Compressed Instructions](c-st-ext.html) can be added (forming RV32GC and RV64GC) to improve performance, code size, and energy efficiency, though with some additional hardware complexity. As we move beyond IMAFDC into further instruction-set extensions, the added instructions tend to be more domain-specific and only provide benefits to a restricted class of applications, e.g., for multimedia or security. Unlike most commercial ISAs, the RISC-V ISA design clearly separates the base ISA and broadly applicable standard extensions from these more specialized additions. | 31 | 27 | 26 | 25 | 24 | 20 | 19 | 15 | 14 | 12 | 11 | 7 | 6 | 0 | | ------------------------ | --- | ------ | ------ | -------------- | ------ | ------ | -- | -- | -- | -- | - | - | - | | funct7 | rs2 | rs1 | funct3 | rd | opcode | R-type | | | | | | | | | imm\[11:0\] | rs1 | funct3 | rd | opcode | I-type | | | | | | | | | | imm\[11:5\] | rs2 | rs1 | funct3 | imm\[4:0\] | opcode | S-type | | | | | | | | | imm\[12\|10:5\] | rs2 | rs1 | funct3 | imm\[4:1\|11\] | opcode | B-type | | | | | | | | | imm\[31:12\] | rd | opcode | U-type | | | | | | | | | | | | imm\[20\|10:1|11|19:12\] | rd | opcode | J-type | | | | | | | | | | | | **RV32I Base Instruction Set** | | | | | | | | | ------------------------------ | ----- | ------- | ----- | -------------- | ------- | ------- | --------- | | imm\[31:12\] | rd | 0110111 | LUI | | | | | | imm\[31:12\] | rd | 0010111 | AUIPC | | | | | | imm\[20\|10:1|11|19:12\] | rd | 1101111 | JAL | | | | | | imm\[11:0\] | rs1 | 000 | rd | 1100111 | JALR | | | | imm\[12\|10:5\] | rs2 | rs1 | 000 | imm\[4:1\|11\] | 1100011 | BEQ | | | imm\[12\|10:5\] | rs2 | rs1 | 001 | imm\[4:1\|11\] | 1100011 | BNE | | | imm\[12\|10:5\] | rs2 | rs1 | 100 | imm\[4:1\|11\] | 1100011 | BLT | | | imm\[12\|10:5\] | rs2 | rs1 | 101 | imm\[4:1\|11\] | 1100011 | BGE | | | imm\[12\|10:5\] | rs2 | rs1 | 110 | imm\[4:1\|11\] | 1100011 | BLTU | | | imm\[12\|10:5\] | rs2 | rs1 | 111 | imm\[4:1\|11\] | 1100011 | BGEU | | | imm\[11:0\] | rs1 | 000 | rd | 0000011 | LB | | | | imm\[11:0\] | rs1 | 001 | rd | 0000011 | LH | | | | imm\[11:0\] | rs1 | 010 | rd | 0000011 | LW | | | | imm\[11:0\] | rs1 | 100 | rd | 0000011 | LBU | | | | imm\[11:0\] | rs1 | 101 | rd | 0000011 | LHU | | | | imm\[11:5\] | rs2 | rs1 | 000 | imm\[4:0\] | 0100011 | SB | | | imm\[11:5\] | rs2 | rs1 | 001 | imm\[4:0\] | 0100011 | SH | | | imm\[11:5\] | rs2 | rs1 | 010 | imm\[4:0\] | 0100011 | SW | | | imm\[11:0\] | rs1 | 000 | rd | 0010011 | ADDI | | | | imm\[11:0\] | rs1 | 010 | rd | 0010011 | SLTI | | | | imm\[11:0\] | rs1 | 011 | rd | 0010011 | SLTIU | | | | imm\[11:0\] | rs1 | 100 | rd | 0010011 | XORI | | | | imm\[11:0\] | rs1 | 110 | rd | 0010011 | ORI | | | | imm\[11:0\] | rs1 | 111 | rd | 0010011 | ANDI | | | | 0000000 | shamt | rs1 | 001 | rd | 0010011 | SLLI | | | 0000000 | shamt | rs1 | 101 | rd | 0010011 | SRLI | | | 0100000 | shamt | rs1 | 101 | rd | 0010011 | SRAI | | | 0000000 | rs2 | rs1 | 000 | rd | 0110011 | ADD | | | 0100000 | rs2 | rs1 | 000 | rd | 0110011 | SUB | | | 0000000 | rs2 | rs1 | 001 | rd | 0110011 | SLL | | | 0000000 | rs2 | rs1 | 010 | rd | 0110011 | SLT | | | 0000000 | rs2 | rs1 | 011 | rd | 0110011 | SLTU | | | 0000000 | rs2 | rs1 | 100 | rd | 0110011 | XOR | | | 0000000 | rs2 | rs1 | 101 | rd | 0110011 | SRL | | | 0100000 | rs2 | rs1 | 101 | rd | 0110011 | SRA | | | 0000000 | rs2 | rs1 | 110 | rd | 0110011 | OR | | | 0000000 | rs2 | rs1 | 111 | rd | 0110011 | AND | | | fm | pred | succ | rs1 | 000 | rd | 0001111 | FENCE | | 1000 | 0011 | 0011 | 00000 | 000 | 00000 | 0001111 | FENCE.TSO | | 0000 | 0001 | 0000 | 00000 | 000 | 00000 | 0001111 | PAUSE | | 000000000000 | 00000 | 000 | 00000 | 1110011 | ECALL | | | | 000000000001 | 00000 | 000 | 00000 | 1110011 | EBREAK | | | | 31 | 27 | 26 | 25 | 24 | 20 | 19 | 15 | 14 | 12 | 11 | 7 | 6 | 0 | | ----------- | --- | ------ | ------ | ---------- | ------ | ------ | -- | -- | -- | -- | - | - | - | | funct7 | rs2 | rs1 | funct3 | rd | opcode | R-type | | | | | | | | | imm\[11:0\] | rs1 | funct3 | rd | opcode | I-type | | | | | | | | | | imm\[11:5\] | rs2 | rs1 | funct3 | imm\[4:0\] | opcode | S-type | | | | | | | | | **RV64I Base Instruction Set (in addition to RV32I)** | | | | | | | | ----------------------------------------------------- | ----- | --- | --- | ---------- | ------- | ----- | | imm\[11:0\] | rs1 | 110 | rd | 0000011 | LWU | | | imm\[11:0\] | rs1 | 011 | rd | 0000011 | LD | | | imm\[11:5\] | rs2 | rs1 | 011 | imm\[4:0\] | 0100011 | SD | | 000000 | shamt | rs1 | 001 | rd | 0010011 | SLLI | | 000000 | shamt | rs1 | 101 | rd | 0010011 | SRLI | | 010000 | shamt | rs1 | 101 | rd | 0010011 | SRAI | | imm\[11:0\] | rs1 | 000 | rd | 0011011 | ADDIW | | | 0000000 | shamt | rs1 | 001 | rd | 0011011 | SLLIW | | 0000000 | shamt | rs1 | 101 | rd | 0011011 | SRLIW | | 0100000 | shamt | rs1 | 101 | rd | 0011011 | SRAIW | | 0000000 | rs2 | rs1 | 000 | rd | 0111011 | ADDW | | 0100000 | rs2 | rs1 | 000 | rd | 0111011 | SUBW | | 0000000 | rs2 | rs1 | 001 | rd | 0111011 | SLLW | | 0000000 | rs2 | rs1 | 101 | rd | 0111011 | SRLW | | 0100000 | rs2 | rs1 | 101 | rd | 0111011 | SRAW | | **RV32/RV64 _Zifencei_ Standard Extension** | | | | | | | ------------------------------------------- | --- | --- | -- | ------- | ------- | | imm\[11:0\] | rs1 | 001 | rd | 0001111 | FENCE.I | | **RV32/RV64 _Zicsr_ Standard Extension** | | | | | | | ---------------------------------------- | ---- | --- | -- | ------- | ------ | | csr | rs1 | 001 | rd | 1110011 | CSRRW | | csr | rs1 | 010 | rd | 1110011 | CSRRS | | csr | rs1 | 011 | rd | 1110011 | CSRRC | | csr | uimm | 101 | rd | 1110011 | CSRRWI | | csr | uimm | 110 | rd | 1110011 | CSRRSI | | csr | uimm | 111 | rd | 1110011 | CSRRCI | | **RV32M Standard Extension** | | | | | | | | ---------------------------- | --- | --- | --- | -- | ------- | ------ | | 0000001 | rs2 | rs1 | 000 | rd | 0110011 | MUL | | 0000001 | rs2 | rs1 | 001 | rd | 0110011 | MULH | | 0000001 | rs2 | rs1 | 010 | rd | 0110011 | MULHSU | | 0000001 | rs2 | rs1 | 011 | rd | 0110011 | MULHU | | 0000001 | rs2 | rs1 | 100 | rd | 0110011 | DIV | | 0000001 | rs2 | rs1 | 101 | rd | 0110011 | DIVU | | 0000001 | rs2 | rs1 | 110 | rd | 0110011 | REM | | 0000001 | rs2 | rs1 | 111 | rd | 0110011 | REMU | | **RV64M Standard Extension (in addition to RV32M)** | | | | | | | | --------------------------------------------------- | --- | --- | --- | -- | ------- | ----- | | 0000001 | rs2 | rs1 | 000 | rd | 0111011 | MULW | | 0000001 | rs2 | rs1 | 100 | rd | 0111011 | DIVW | | 0000001 | rs2 | rs1 | 101 | rd | 0111011 | DIVUW | | 0000001 | rs2 | rs1 | 110 | rd | 0111011 | REMW | | 0000001 | rs2 | rs1 | 111 | rd | 0111011 | REMUW | | 31 | 27 | 26 | 25 | 24 | 20 | 19 | 15 | 14 | 12 | 11 | 7 | 6 | 0 | | ------ | --- | --- | ------ | -- | ------ | ------ | -- | -- | -- | -- | - | - | - | | funct7 | rs2 | rs1 | funct3 | rd | opcode | R-type | | | | | | | | | **RV32A Standard Extension** | | | | | | | | | | ---------------------------- | -- | -- | ----- | --- | --- | -- | ------- | --------- | | 00010 | aq | rl | 00000 | rs1 | 010 | rd | 0101111 | LR.W | | 00011 | aq | rl | rs2 | rs1 | 010 | rd | 0101111 | SC.W | | 00001 | aq | rl | rs2 | rs1 | 010 | rd | 0101111 | AMOSWAP.W | | 00000 | aq | rl | rs2 | rs1 | 010 | rd | 0101111 | AMOADD.W | | 00100 | aq | rl | rs2 | rs1 | 010 | rd | 0101111 | AMOXOR.W | | 01100 | aq | rl | rs2 | rs1 | 010 | rd | 0101111 | AMOAND.W | | 01000 | aq | rl | rs2 | rs1 | 010 | rd | 0101111 | AMOOR.W | | 10000 | aq | rl | rs2 | rs1 | 010 | rd | 0101111 | AMOMIN.W | | 10100 | aq | rl | rs2 | rs1 | 010 | rd | 0101111 | AMOMAX.W | | 11000 | aq | rl | rs2 | rs1 | 010 | rd | 0101111 | AMOMINU.W | | 11100 | aq | rl | rs2 | rs1 | 010 | rd | 0101111 | AMOMAXU.W | | **RV64A Standard Extension (in addition to RV32A)** | | | | | | | | | | --------------------------------------------------- | -- | -- | ----- | --- | --- | -- | ------- | --------- | | 00010 | aq | rl | 00000 | rs1 | 011 | rd | 0101111 | LR.D | | 00011 | aq | rl | rs2 | rs1 | 011 | rd | 0101111 | SC.D | | 00001 | aq | rl | rs2 | rs1 | 011 | rd | 0101111 | AMOSWAP.D | | 00000 | aq | rl | rs2 | rs1 | 011 | rd | 0101111 | AMOADD.D | | 00100 | aq | rl | rs2 | rs1 | 011 | rd | 0101111 | AMOXOR.D | | 01100 | aq | rl | rs2 | rs1 | 011 | rd | 0101111 | AMOAND.D | | 01000 | aq | rl | rs2 | rs1 | 011 | rd | 0101111 | AMOOR.D | | 10000 | aq | rl | rs2 | rs1 | 011 | rd | 0101111 | AMOMIN.D | | 10100 | aq | rl | rs2 | rs1 | 011 | rd | 0101111 | AMOMAX.D | | 11000 | aq | rl | rs2 | rs1 | 011 | rd | 0101111 | AMOMINU.D | | 11100 | aq | rl | rs2 | rs1 | 011 | rd | 0101111 | AMOMAXU.D | | 31 | 27 | 26 | 25 | 24 | 20 | 19 | 15 | 14 | 12 | 11 | 7 | 6 | 0 | | ----------- | ------ | ------ | ------ | ---------- | ------ | ------ | ------- | -- | -- | -- | - | - | - | | funct7 | rs2 | rs1 | funct3 | rd | opcode | R-type | | | | | | | | | rs3 | funct2 | rs2 | rs1 | funct3 | rd | opcode | R4-type | | | | | | | | imm\[11:0\] | rs1 | funct3 | rd | opcode | I-type | | | | | | | | | | imm\[11:5\] | rs2 | rs1 | funct3 | imm\[4:0\] | opcode | S-type | | | | | | | | | **RV32F Standard Extension** | | | | | | | | | ---------------------------- | ----- | --- | --- | ---------- | ------- | --------- | -------- | | imm\[11:0\] | rs1 | 010 | rd | 0000111 | FLW | | | | imm\[11:5\] | rs2 | rs1 | 010 | imm\[4:0\] | 0100111 | FSW | | | rs3 | 00 | rs2 | rs1 | rm | rd | 1000011 | FMADD.S | | rs3 | 00 | rs2 | rs1 | rm | rd | 1000111 | FMSUB.S | | rs3 | 00 | rs2 | rs1 | rm | rd | 1001011 | FNMSUB.S | | rs3 | 00 | rs2 | rs1 | rm | rd | 1001111 | FNMADD.S | | 0000000 | rs2 | rs1 | rm | rd | 1010011 | FADD.S | | | 0000100 | rs2 | rs1 | rm | rd | 1010011 | FSUB.S | | | 0001000 | rs2 | rs1 | rm | rd | 1010011 | FMUL.S | | | 0001100 | rs2 | rs1 | rm | rd | 1010011 | FDIV.S | | | 0101100 | 00000 | rs1 | rm | rd | 1010011 | FSQRT.S | | | 0010000 | rs2 | rs1 | 000 | rd | 1010011 | FSGNJ.S | | | 0010000 | rs2 | rs1 | 001 | rd | 1010011 | FSGNJN.S | | | 0010000 | rs2 | rs1 | 010 | rd | 1010011 | FSGNJX.S | | | 0010100 | rs2 | rs1 | 000 | rd | 1010011 | FMIN.S | | | 0010100 | rs2 | rs1 | 001 | rd | 1010011 | FMAX.S | | | 1100000 | 00000 | rs1 | rm | rd | 1010011 | FCVT.W.S | | | 1100000 | 00001 | rs1 | rm | rd | 1010011 | FCVT.WU.S | | | 1110000 | 00000 | rs1 | 000 | rd | 1010011 | FMV.X.W | | | 1010000 | rs2 | rs1 | 010 | rd | 1010011 | FEQ.S | | | 1010000 | rs2 | rs1 | 001 | rd | 1010011 | FLT.S | | | 1010000 | rs2 | rs1 | 000 | rd | 1010011 | FLE.S | | | 1110000 | 00000 | rs1 | 001 | rd | 1010011 | FCLASS.S | | | 1101000 | 00000 | rs1 | rm | rd | 1010011 | FCVT.S.W | | | 1101000 | 00001 | rs1 | rm | rd | 1010011 | FCVT.S.WU | | | 1111000 | 00000 | rs1 | 000 | rd | 1010011 | FMV.W.X | | | **RV64F Standard Extension (in addition to RV32F)** | | | | | | | | --------------------------------------------------- | ----- | --- | -- | -- | ------- | --------- | | 1100000 | 00010 | rs1 | rm | rd | 1010011 | FCVT.L.S | | 1100000 | 00011 | rs1 | rm | rd | 1010011 | FCVT.LU.S | | 1101000 | 00010 | rs1 | rm | rd | 1010011 | FCVT.S.L | | 1101000 | 00011 | rs1 | rm | rd | 1010011 | FCVT.S.LU | | 31 | 27 | 26 | 25 | 24 | 20 | 19 | 15 | 14 | 12 | 11 | 7 | 6 | 0 | | ----------- | ------ | ------ | ------ | ---------- | ------ | ------ | ------- | -- | -- | -- | - | - | - | | funct7 | rs2 | rs1 | funct3 | rd | opcode | R-type | | | | | | | | | rs3 | funct2 | rs2 | rs1 | funct3 | rd | opcode | R4-type | | | | | | | | imm\[11:0\] | rs1 | funct3 | rd | opcode | I-type | | | | | | | | | | imm\[11:5\] | rs2 | rs1 | funct3 | imm\[4:0\] | opcode | S-type | | | | | | | | | **RV32D Standard Extension** | | | | | | | | | ---------------------------- | ----- | --- | --- | ---------- | ------- | --------- | -------- | | imm\[11:0\] | rs1 | 011 | rd | 0000111 | FLD | | | | imm\[11:5\] | rs2 | rs1 | 011 | imm\[4:0\] | 0100111 | FSD | | | rs3 | 01 | rs2 | rs1 | rm | rd | 1000011 | FMADD.D | | rs3 | 01 | rs2 | rs1 | rm | rd | 1000111 | FMSUB.D | | rs3 | 01 | rs2 | rs1 | rm | rd | 1001011 | FNMSUB.D | | rs3 | 01 | rs2 | rs1 | rm | rd | 1001111 | FNMADD.D | | 0000001 | rs2 | rs1 | rm | rd | 1010011 | FADD.D | | | 0000101 | rs2 | rs1 | rm | rd | 1010011 | FSUB.D | | | 0001001 | rs2 | rs1 | rm | rd | 1010011 | FMUL.D | | | 0001101 | rs2 | rs1 | rm | rd | 1010011 | FDIV.D | | | 0101101 | 00000 | rs1 | rm | rd | 1010011 | FSQRT.D | | | 0010001 | rs2 | rs1 | 000 | rd | 1010011 | FSGNJ.D | | | 0010001 | rs2 | rs1 | 001 | rd | 1010011 | FSGNJN.D | | | 0010001 | rs2 | rs1 | 010 | rd | 1010011 | FSGNJX.D | | | 0010101 | rs2 | rs1 | 000 | rd | 1010011 | FMIN.D | | | 0010101 | rs2 | rs1 | 001 | rd | 1010011 | FMAX.D | | | 0100000 | 00001 | rs1 | rm | rd | 1010011 | FCVT.S.D | | | 0100001 | 00000 | rs1 | rm | rd | 1010011 | FCVT.D.S | | | 1010001 | rs2 | rs1 | 010 | rd | 1010011 | FEQ.D | | | 1010001 | rs2 | rs1 | 001 | rd | 1010011 | FLT.D | | | 1010001 | rs2 | rs1 | 000 | rd | 1010011 | FLE.D | | | 1110001 | 00000 | rs1 | 001 | rd | 1010011 | FCLASS.D | | | 1100001 | 00000 | rs1 | rm | rd | 1010011 | FCVT.W.D | | | 1100001 | 00001 | rs1 | rm | rd | 1010011 | FCVT.WU.D | | | 1101001 | 00000 | rs1 | rm | rd | 1010011 | FCVT.D.W | | | 1101001 | 00001 | rs1 | rm | rd | 1010011 | FCVT.D.WU | | | **RV64D Standard Extension (in addition to RV32D)** | | | | | | | | --------------------------------------------------- | ----- | --- | --- | -- | ------- | --------- | | 1100001 | 00010 | rs1 | rm | rd | 1010011 | FCVT.L.D | | 1100001 | 00011 | rs1 | rm | rd | 1010011 | FCVT.LU.D | | 1110001 | 00000 | rs1 | 000 | rd | 1010011 | FMV.X.D | | 1101001 | 00010 | rs1 | rm | rd | 1010011 | FCVT.D.L | | 1101001 | 00011 | rs1 | rm | rd | 1010011 | FCVT.D.LU | | 1111001 | 00000 | rs1 | 000 | rd | 1010011 | FMV.D.X | | 31 | 27 | 26 | 25 | 24 | 20 | 19 | 15 | 14 | 12 | 11 | 7 | 6 | 0 | | ----------- | ------ | ------ | ------ | ---------- | ------ | ------ | ------- | -- | -- | -- | - | - | - | | funct7 | rs2 | rs1 | funct3 | rd | opcode | R-type | | | | | | | | | rs3 | funct2 | rs2 | rs1 | funct3 | rd | opcode | R4-type | | | | | | | | imm\[11:0\] | rs1 | funct3 | rd | opcode | I-type | | | | | | | | | | imm\[11:5\] | rs2 | rs1 | funct3 | imm\[4:0\] | opcode | S-type | | | | | | | | | **RV32Q Standard Extension** | | | | | | | | | ---------------------------- | ----- | --- | --- | ---------- | ------- | --------- | -------- | | imm\[11:0\] | rs1 | 100 | rd | 0000111 | FLQ | | | | imm\[11:5\] | rs2 | rs1 | 100 | imm\[4:0\] | 0100111 | FSQ | | | rs3 | 11 | rs2 | rs1 | rm | rd | 1000011 | FMADD.Q | | rs3 | 11 | rs2 | rs1 | rm | rd | 1000111 | FMSUB.Q | | rs3 | 11 | rs2 | rs1 | rm | rd | 1001011 | FNMSUB.Q | | rs3 | 11 | rs2 | rs1 | rm | rd | 1001111 | FNMADD.Q | | 0000011 | rs2 | rs1 | rm | rd | 1010011 | FADD.Q | | | 0000111 | rs2 | rs1 | rm | rd | 1010011 | FSUB.Q | | | 0001011 | rs2 | rs1 | rm | rd | 1010011 | FMUL.Q | | | 0001111 | rs2 | rs1 | rm | rd | 1010011 | FDIV.Q | | | 0101111 | 00000 | rs1 | rm | rd | 1010011 | FSQRT.Q | | | 0010011 | rs2 | rs1 | 000 | rd | 1010011 | FSGNJ.Q | | | 0010011 | rs2 | rs1 | 001 | rd | 1010011 | FSGNJN.Q | | | 0010011 | rs2 | rs1 | 010 | rd | 1010011 | FSGNJX.Q | | | 0010111 | rs2 | rs1 | 000 | rd | 1010011 | FMIN.Q | | | 0010111 | rs2 | rs1 | 001 | rd | 1010011 | FMAX.Q | | | 0100000 | 00011 | rs1 | rm | rd | 1010011 | FCVT.S.Q | | | 0100011 | 00000 | rs1 | rm | rd | 1010011 | FCVT.Q.S | | | 0100001 | 00011 | rs1 | rm | rd | 1010011 | FCVT.D.Q | | | 0100011 | 00001 | rs1 | rm | rd | 1010011 | FCVT.Q.D | | | 1010011 | rs2 | rs1 | 010 | rd | 1010011 | FEQ.Q | | | 1010011 | rs2 | rs1 | 001 | rd | 1010011 | FLT.Q | | | 1010011 | rs2 | rs1 | 000 | rd | 1010011 | FLE.Q | | | 1110011 | 00000 | rs1 | 001 | rd | 1010011 | FCLASS.Q | | | 1100011 | 00000 | rs1 | rm | rd | 1010011 | FCVT.W.Q | | | 1100011 | 00001 | rs1 | rm | rd | 1010011 | FCVT.WU.Q | | | 1101011 | 00000 | rs1 | rm | rd | 1010011 | FCVT.Q.W | | | 1101011 | 00001 | rs1 | rm | rd | 1010011 | FCVT.Q.WU | | | **RV64Q Standard Extension (in addition to RV32Q)** | | | | | | | | --------------------------------------------------- | ----- | --- | -- | -- | ------- | --------- | | 1100011 | 00010 | rs1 | rm | rd | 1010011 | FCVT.L.Q | | 1100011 | 00011 | rs1 | rm | rd | 1010011 | FCVT.LU.Q | | 1101011 | 00010 | rs1 | rm | rd | 1010011 | FCVT.Q.L | | 1101011 | 00011 | rs1 | rm | rd | 1010011 | FCVT.Q.LU | | 31 | 27 | 26 | 25 | 24 | 20 | 19 | 15 | 14 | 12 | 11 | 7 | 6 | 0 | | ----------- | ------ | ------ | ------ | ---------- | ------ | ------ | ------- | -- | -- | -- | - | - | - | | funct7 | rs2 | rs1 | funct3 | rd | opcode | R-type | | | | | | | | | rs3 | funct2 | rs2 | rs1 | funct3 | rd | opcode | R4-type | | | | | | | | imm\[11:0\] | rs1 | funct3 | rd | opcode | I-type | | | | | | | | | | imm\[11:5\] | rs2 | rs1 | funct3 | imm\[4:0\] | opcode | S-type | | | | | | | | | RV32Zfh Standard Extension | | | | | | | | | -------------------------- | ----- | --- | --- | ---------- | ------- | --------- | -------- | | imm\[11:0\] | rs1 | 001 | rd | 0000111 | FLH | | | | imm\[11:5\] | rs2 | rs1 | 001 | imm\[4:0\] | 0100111 | FSH | | | rs3 | 10 | rs2 | rs1 | rm | rd | 1000011 | FMADD.H | | rs3 | 10 | rs2 | rs1 | rm | rd | 1000111 | FMSUB.H | | rs3 | 10 | rs2 | rs1 | rm | rd | 1001011 | FNMSUB.H | | rs3 | 10 | rs2 | rs1 | rm | rd | 1001111 | FNMADD.H | | 0000010 | rs2 | rs1 | rm | rd | 1010011 | FADD.H | | | 0000110 | rs2 | rs1 | rm | rd | 1010011 | FSUB.H | | | 0001010 | rs2 | rs1 | rm | rd | 1010011 | FMUL.H | | | 0001110 | rs2 | rs1 | rm | rd | 1010011 | FDIV.H | | | 0101110 | 00000 | rs1 | rm | rd | 1010011 | FSQRT.H | | | 0010010 | rs2 | rs1 | 000 | rd | 1010011 | FSGNJ.H | | | 0010010 | rs2 | rs1 | 001 | rd | 1010011 | FSGNJN.H | | | 0010010 | rs2 | rs1 | 010 | rd | 1010011 | FSGNJX.H | | | 0010110 | rs2 | rs1 | 000 | rd | 1010011 | FMIN.H | | | 0010110 | rs2 | rs1 | 001 | rd | 1010011 | FMAX.H | | | 0100000 | 00010 | rs1 | rm | rd | 1010011 | FCVT.S.H | | | 0100010 | 00000 | rs1 | rm | rd | 1010011 | FCVT.H.S | | | 0100001 | 00010 | rs1 | rm | rd | 1010011 | FCVT.D.H | | | 0100010 | 00001 | rs1 | rm | rd | 1010011 | FCVT.H.D | | | 0100011 | 00010 | rs1 | rm | rd | 1010011 | FCVT.Q.H | | | 0100010 | 00011 | rs1 | rm | rd | 1010011 | FCVT.H.Q | | | 1010010 | rs2 | rs1 | 010 | rd | 1010011 | FEQ.H | | | 1010010 | rs2 | rs1 | 001 | rd | 1010011 | FLT.H | | | 1010010 | rs2 | rs1 | 000 | rd | 1010011 | FLE.H | | | 1110010 | 00000 | rs1 | 001 | rd | 1010011 | FCLASS.H | | | 1100010 | 00000 | rs1 | rm | rd | 1010011 | FCVT.W.H | | | 1100010 | 00001 | rs1 | rm | rd | 1010011 | FCVT.WU.H | | | 1110010 | 00000 | rs1 | 000 | rd | 1010011 | FMV.X.H | | | 1101010 | 00000 | rs1 | rm | rd | 1010011 | FCVT.H.W | | | 1101010 | 00001 | rs1 | rm | rd | 1010011 | FCVT.H.WU | | | 1111010 | 00000 | rs1 | 000 | rd | 1010011 | FMV.H.X | | | RV64Zfh Standard Extension (in addition to RV32Zfh) | | | | | | | | --------------------------------------------------- | ----- | --- | -- | -- | ------- | --------- | | 1100010 | 00010 | rs1 | rm | rd | 1010011 | FCVT.L.H | | 1100010 | 00011 | rs1 | rm | rd | 1010011 | FCVT.LU.H | | 1101010 | 00010 | rs1 | rm | rd | 1010011 | FCVT.H.L | | 1101010 | 00011 | rs1 | rm | rd | 1010011 | FCVT.H.LU | | Zawrs Standard Extension | | | | | | | ------------------------ | ----- | --- | ----- | ------- | ------- | | 000000001101 | 00000 | 000 | 00000 | 1110011 | WRS.NTO | | 000000011101 | 00000 | 000 | 00000 | 1110011 | WRS.STO | [Table 2](#rvgcsrnames) lists the CSRs that have currently been allocated CSR addresses. The timers, counters, and floating-point CSRs are the only CSRs defined in this specification. __Table 2\. RISC-V control and status register (CSR) address map.__ | Number | Privilege | Name | Description | | ------------------------------------------- | ---------- | -------- | ----------------------------------------------------------- | | Floating-Point Control and Status Registers | | | | | 0x001 | Read write | fflags | Floating-Point Accrued Exceptions. | | 0x002 | Read write | frm | Floating-Point Dynamic Rounding Mode. | | 0x003 | Read write | fcsr | Floating-Point Control and Status Register (frm \+ fflags). | | Counters and Timers | | | | | 0xC00 | Read-only | cycle | Cycle counter for RDCYCLE instruction. | | 0xC01 | Read-only | time | Timer for RDTIME instruction. | | 0xC02 | Read-only | instret | Instructions-retired counter for RDINSTRET instruction. | | 0xC80 | Read-only | cycleh | Upper 32 bits of cycle, RV32I only. | | 0xC81 | Read-only | timeh | Upper 32 bits of time, RV32I only. | | 0xC82 | Read-only | instreth | Upper 32 bits of instret, RV32I only. | 2.1. RV32I Base Integer Instruction Set, Version 2.1 ==================== ## [](#rv32)2.1\. RV32I Base Integer Instruction Set, Version 2.1 This chapter describes the RV32I base integer instruction set. | | RV32I was designed to be sufficient to form a compiler target and to support modern operating system environments. The ISA was also designed to reduce the hardware required in a minimal implementation. RV32I contains 40 unique instructions, though a simple implementation might cover the ECALL/EBREAK instructions with a single SYSTEM hardware instruction that always traps and might be able to implement the FENCE instruction as a NOP, reducing base instruction count to 38 total. RV32I can emulate almost any other ISA extension (except the A extension, which requires additional hardware support for atomicity). In practice, a hardware implementation including the machine-mode privileged architecture will also require the 6 CSR instructions. Subsets of the base integer ISA might be useful for pedagogical purposes, but the base has been defined such that there should be little incentive to subset a real hardware implementation beyond omitting support for misaligned memory accesses and treating all SYSTEM instructions as a single trap. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The standard RISC-V assembly language syntax is documented in the Assembly Programmer’s Manual \[[11](../biblio/bibliography.html#bib-riscv-asm-manual)\]. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Most of the commentary for RV32I also applies to the RV64I base. | | ------------------------------------------------------------------- | ### [](#2-1-1-programmers-model-for-base-integer-isa)2.1.1\. Programmers' Model for Base Integer ISA [Table 1](#gprs) shows the unprivileged state for the base integer ISA.For RV32I, the 32 `x` registers are each 32 bits wide, i.e., `XLEN=32`. Register `x0` is hardwired with all bits equal to 0. General purpose registers `x1-x31` hold values that various instructions interpret as a collection of Boolean values, or as two’s complement signed binary integers or unsigned binary integers. There is one additional unprivileged register: the program counter `pc`holds the address of the current instruction. __Table 1\. RISC-V base unprivileged integer register state.__ | XLEN-1 | 0 | | ------- | - | | x0/zero | | | x1 | | | x2 | | | x3 | | | x4 | | | x5 | | | x6 | | | x7 | | | x8 | | | x9 | | | x10 | | | x11 | | | x12 | | | x13 | | | x14 | | | x15 | | | x16 | | | x17 | | | x18 | | | x19 | | | x20 | | | x21 | | | x22 | | | x23 | | | x24 | | | x25 | | | x26 | | | x27 | | | x28 | | | x29 | | | x30 | | | x31 | | | pc | | | | There is no dedicated stack pointer or subroutine return address link register in the Base Integer ISA; the instruction encoding allows anyx register to be used for these purposes. However, the standard software calling convention uses register x1 to hold the return address for a call, with register x5 available as an alternate link register. The standard calling convention uses register x2 as the stack pointer. Hardware might choose to accelerate function calls and returns that usex1 or x5. See the descriptions of the JAL and JALR instructions. The optional compressed 16-bit instruction format is designed around the assumption that x1 is the return address register and x2 is the stack pointer. Software using other conventions will operate correctly but may have greater code size. The number of available architectural registers can have large impacts on code size, performance, and energy consumption. Although 16 registers would arguably be sufficient for an integer ISA running compiled code, it is impossible to encode a complete ISA with 16 registers in 16-bit instructions using a 3-address format. Although a 2-address format would be possible, it would increase instruction count and lower efficiency. We wanted to avoid intermediate instruction sizes (such as Xtensa’s 24-bit instructions) to simplify base hardware implementations, and once a 32-bit instruction size was adopted, it was straightforward to support 32 integer registers. A larger number of integer registers also helps performance on high-performance code, where there can be extensive use of loop unrolling, software pipelining, and cache tiling. For these reasons, we chose a conventional size of 32 integer registers for RV32I. Dynamic register usage tends to be dominated by a few frequently accessed registers, and register file implementations can be optimized to reduce access energy for the frequently accessed registers \[[12](../biblio/bibliography.html#bib-jtseng:sbbci)\]. The optional compressed 16-bit instruction format mostly only accesses 8 registers and hence can provide a dense instruction encoding, while additional instruction-set extensions could support a much larger register space (either flat or hierarchical) if desired. For resource-constrained embedded applications, we have defined the RV32E subset, which only has 16 registers ([RV32E and RV64E Base Integer Instruction Sets](rv32e.html)). | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#2-1-2-base-instruction-formats)2.1.2\. Base Instruction Formats In the base RV32I ISA, there are four core instruction formats (R/I/S/U), as shown in [Base instruction formats](#base%5Finstr). All are a fixed 32 bits in length. The base ISA has `IALIGN=32`, meaning that instructions must be aligned on a four-byte boundary in memory.An instruction-address-misaligned exception is generated on a taken branch or unconditional jump if the target address is not `IALIGN-bit` aligned. This exception is reported on the branch or jump instruction, not on the target instruction. No instruction-address-misaligned exception is generated for a conditional branch that is not taken. | | The alignment constraint for base ISA instructions is relaxed to a two-byte boundary when instruction extensions with 16-bit lengths or other odd multiples of 16-bit lengths are added (i.e., IALIGN=16). Instruction-address-misaligned exceptions are reported on the branch or jump that would cause instruction misalignment to help debugging, and to simplify hardware design for systems with IALIGN=32, where these are the only places where misalignment can occur. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The behavior upon decoding a reserved instruction is UNSPECIFIED. | | Some platforms may require that opcodes reserved for standard use raise an illegal-instruction exception. Other platforms may permit reserved opcode space be used for non-conforming extensions. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The RISC-V ISA keeps the source (_rs1_ and _rs2_) and destination (_rd_) registers at the same position in all formats to simplify decoding.Except for the 5-bit immediates used in CSR instructions ([CSR Instructions](zicsr.html#csrinsts)), immediates are always sign-extended, and are generally packed towards the leftmost available bits in the instruction and have been allocated to reduce hardware complexity. In particular, the sign bit for all immediates is always in bit 31 of the instruction to speed sign-extension circuitry. ![svg](_images/svg-5d875942584dea273a450c7a261a1ad5d1f3259f.svg) ![svg](_images/svg-6e79131cf8a4d97ce98f508a368da764a7b11e95.svg) ![svg](_images/svg-5d013bc40702f55441bc51b0f78d8e0364674976.svg) ![svg](_images/svg-1aa02bb0f4c4c7c2c08334976bd48bff3084fe01.svg) RISC-V base instruction formats. Each immediate subfield is labeled with the bit position (imm\[x\]) in the immediate value being produced, rather than the bit position within the instruction’s immediate field as is usually done. | | Decoding register specifiers is usually on the critical paths in implementations, and so the instruction format was chosen to keep all register specifiers at the same position in all formats at the expense of having to move immediate bits across formats (a property shared with RISC-IV aka. SPUR \[[8](../biblio/bibliography.html#bib-spur-jsscc1989)\]). In practice, most immediates are either small or require all XLEN bits. We chose an asymmetric immediate split (12 bits in regular instructions plus a special load-upper-immediate instruction with 20 bits) to increase the opcode space available for regular instructions. Immediates are sign-extended because we did not observe a benefit to using zero extension for some immediates as in the MIPS ISA and wanted to keep the ISA as simple as possible. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#2-1-3-immediate-encoding-variants)2.1.3\. Immediate Encoding Variants There are a further two variants of the instruction formats (B/J) based on the handling of immediates, as shown in [Base instruction formats immediate variants.](#baseinstformatsimm). ![svg](_images/svg-5d875942584dea273a450c7a261a1ad5d1f3259f.svg) ![svg](_images/svg-6e79131cf8a4d97ce98f508a368da764a7b11e95.svg) ![svg](_images/svg-5d013bc40702f55441bc51b0f78d8e0364674976.svg) ![svg](_images/svg-21498c6d6516a6ccebcacf081f799f7c40206aa8.svg) ![svg](_images/svg-fff5090c931f6fd03a317313ec25535245e8374f.svg) ![svg](_images/svg-14f2b6c489d6c020109952daa15bec93033adb93.svg) The only difference between the S and B formats is that the 12-bit immediate field is used to encode branch offsets in multiples of 2 in the B format. Instead of shifting all bits in the instruction-encoded immediate left by one in hardware as is conventionally done, the middle bits (imm\[10:1\]) and sign bit stay in fixed positions, while the lowest bit in S format (inst\[7\]) encodes a high-order bit in B format. Similarly, the only difference between the U and J formats is that the 20-bit immediate is shifted left by 12 bits to form U immediates and by 1 bit to form J immediates. The location of instruction bits in the U and J format immediates is chosen to maximize overlap with the other formats and with each other. [Immediate types](#immtypes) shows the immediates produced by each of the base instruction formats, and is labeled to show which instruction bit (inst\[_y_\]) produces each bit of the immediate value. ![svg](_images/svg-f8dfef33c3ef80332baa331a8e67b7a1475c2265.svg) ![svg](_images/svg-dcabb5476c0c0275c128bb81b9e6bd7a4304cd62.svg) ![svg](_images/svg-3a207a4fff82d1d4a3cb382aaef60aadc1552e38.svg) ![svg](_images/svg-0b42136e8bf9a754253a9ae2da62db368a8e0bb5.svg) ![Types of immediate produced by RISC-V instructions.](_images/svg-60de38432d3b44cb8fbd41a78322646b44e31e6f.svg) Figure 1\. Types of immediate produced by RISC-V instructions. The fields are labeled with the instruction bits used to construct their value. Sign extensions always uses inst\[31\]. | | Sign extension is one of the most critical operations on immediates (particularly for XLEN>32), and in RISC-V the sign bit for all immediates is always held in bit 31 of the instruction to allow sign extension to proceed in parallel with instruction decoding. Although more complex implementations might have separate adders for branch and jump calculations and so would not benefit from keeping the location of immediate bits constant across types of instruction, we wanted to reduce the hardware cost of the simplest implementations. By rotating bits in the instruction encoding of B and J immediates instead of using dynamic hardware multiplexers to multiply the immediate by 2, we reduce instruction signal fanout and immediate multiplexer costs by around a factor of 2\. The scrambled immediate encoding will add negligible time to static or ahead-of-time compilation. For dynamic generation of instructions, there is some small additional overhead, but the most common short forward branches have straightforward immediate encodings. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#2-1-4-integer-computational-instructions)2.1.4\. Integer Computational Instructions Most integer computational instructions operate on `XLEN` bits of values held in the integer register file. Integer computational instructions are either encoded as register-immediate operations using the I-type format or as register-register operations using the R-type format. The destination is register _rd_ for both register-immediate and register-register instructions. No integer computational instructions cause arithmetic exceptions. | | We did not include special instruction-set support for overflow checks on integer arithmetic operations in the base instruction set, as many overflow checks can be cheaply implemented using RISC-V branches. Overflow checking for unsigned addition requires only a single additional branch instruction after the addition:add t0, t1, t2; bltu t0, t1, overflow. For signed addition, if one operand’s sign is known, overflow checking requires only a single branch after the addition:addi t0, t1, +imm; blt t0, t1, overflow. This covers the common case of addition with an immediate operand. For general signed addition, three additional instructions after the addition are required, leveraging the observation that the sum should be less than one of the operands if and only if the other operand is negative. add t0, t1, t2 slti t3, t2, 0 slt t4, t0, t1 bne t3, t4, overflow In RV64I, checks of 32-bit signed additions can be optimized further by comparing the results of ADD and ADDW on the operands. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#2-1-4-1-integer-register-immediate-instructions)2.1.4.1\. Integer Register-Immediate Instructions ![svg](_images/svg-9f17986825ade937c6ac689b28d78b9d9de67e71.svg) ADDI adds the sign-extended 12-bit immediate to register _rs1_. Arithmetic overflow is ignored and the result is simply the low XLEN bits of the result.ADDI _rd, rs1, 0_ is used to implement the MV _rd, rs1_ assembler pseudoinstruction. SLTI (set less than immediate) places the value 1 in register _rd_ if register _rs1_ is less than the sign-extended immediate when both are treated as signed numbers, else 0 is written to _rd_. SLTIU is similar but compares the values as unsigned numbers (i.e., the immediate is first sign-extended to XLEN bits then treated as an unsigned number).Note, SLTIU _rd, rs1, 1_ sets _rd_ to 1 if _rs1_ equals zero, otherwise sets _rd_ to 0 (assembler pseudoinstruction SEQZ _rd, rs_). ANDI, ORI, XORI are logical operations that perform bitwise AND, OR, and XOR on register _rs1_ and the sign-extended 12-bit immediate and place the result in _rd_.Note, XORI _rd, rs1, -1_ performs a bitwise logical inversion of register _rs1_ (assembler pseudoinstruction NOT _rd, rs_). ![svg](_images/svg-bad579e5e2f2b9a4758b862b3ec451f7fa6319a8.svg) Shifts by a constant are encoded as a specialization of the I-type format. The operand to be shifted is in _rs1_, and the shift amount is encoded in the lower 5 bits of the I-immediate field. The right-shift type is encoded in bit 30.SLLI is a logical left shift (zeros are shifted into the lower bits); SRLI is a logical right shift (zeros are shifted into the upper bits); andSRAI is an arithmetic right shift (the original sign bit is copied into the vacated upper bits). ![svg](_images/svg-c0abc0bfc80eceb73d353e08a13e7a84ed863093.svg) LUI (load upper immediate) is used to build 32-bit constants and uses the U-type format.LUI places the 32-bit U-immediate value into the destination register _rd_, filling in the lowest 12 bits with zeros. AUIPC (add upper immediate to `pc`) is used to build `pc`\-relative addresses and uses the U-type format.AUIPC forms a 32-bit offset from the U-immediate, filling in the lowest 12 bits with zeros, adds this offset to the address of the AUIPC instruction, then places the result in register _rd_. | | The assembly syntax for lui and auipc does not represent the lower 12 bits of the U-immediate, which are always zero. The AUIPC instruction supports two-instruction sequences to access arbitrary offsets from the PC for both control-flow transfers and data accesses. The combination of an AUIPC and the 12-bit immediate in a JALR can transfer control to any 32-bit PC-relative address, while an AUIPC plus the 12-bit immediate offset in regular load or store instructions can access any 32-bit PC-relative data address. The current PC can be obtained by setting the U-immediate to 0\. Although a JAL +4 instruction could also be used to obtain the local PC (of the instruction following the JAL), it might cause pipeline breaks in simpler microarchitectures or pollute BTB structures in more complex microarchitectures. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#2-1-4-2-integer-register-register-instructions)2.1.4.2\. Integer Register-Register Instructions RV32I defines several arithmetic R-type operations. All operations read the _rs1_ and _rs2_ registers as source operands and write the result into register _rd_.The _funct7_ and _funct3_ fields select the type of operation. ![svg](_images/svg-0efc95832ccd70882f52ef5147ed5039f1ea4819.svg) ADD performs the addition of _rs1_ and _rs2_. SUB performs the subtraction of _rs2_ from _rs1_. Overflows are ignored and the low XLEN bits of results are written to the destination _rd_. SLT and SLTU perform signed and unsigned compares respectively, writing 1 to _rd_ if_rs1_ < _rs2_, 0 otherwise.Note, SLTU _rd_, _x0_, _rs2_ sets _rd_ to 1 if _rs2_ is not equal to zero, otherwise sets _rd_ to zero (assembler pseudoinstruction SNEZ _rd, rs_).AND, OR, and XOR perform bitwise logical operations. SLL, SRL, and SRA perform logical left, logical right, and arithmetic right shifts on the value in register _rs1_ by the shift amount held in the lower 5 bits of register _rs2_. #### [](#2-1-4-3-nop-instruction)2.1.4.3\. NOP Instruction ![svg](_images/svg-7fbc7aaa33d9f74764e1543657dc0fbdadf8665e.svg) The NOP instruction does not change any architecturally visible state, except for advancing the `pc` and incrementing any applicable performance counters. NOP is encoded as ADDI _x0, x0, 0_. | | NOPs can be used to align code segments to microarchitecturally significant address boundaries, or to leave space for inline code modifications. Although there are many possible ways to encode a NOP, we define a canonical NOP encoding to allow microarchitectural optimizations as well as for more readable disassembly output. The other NOP encodings are made available for ["Zfh" and "HINT Instructions](#rv32i-hints). ADDI was chosen for the NOP encoding as this is most likely to take fewest resources to execute across a range of systems (if not optimized away in decode). In particular, the instruction only reads one register. Also, an ADDI functional unit is more likely to be available in a superscalar design as adds are the most common operation. In particular, address-generation functional units can execute ADDI using the same hardware needed for base+offset address calculations, while register-register ADD or logical/shift operations require additional hardware. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#2-1-5-control-transfer-instructions)2.1.5\. Control Transfer Instructions RV32I provides two types of control transfer instructions: unconditional jumps and conditional branches.Control transfer instructions in RV32I do _not_ have architecturally visible delay slots. If an instruction access-fault or instruction page-fault exception occurs on the target of a jump or taken branch, the exception is reported on the target instruction, not on the jump or branch instruction. #### [](#2-1-5-1-unconditional-jumps)2.1.5.1\. Unconditional Jumps The jump and link (JAL) instruction uses the J-type format, where the J-immediate encodes a signed offset in multiples of 2 bytes.The offset is sign-extended and added to the address of the jump instruction to form the jump target address. Jumps can therefore target a ±1 MiB range.JAL stores the address of the instruction following the jump ('pc'+4) into register _rd_.The standard software calling convention uses 'x1' as the return address register and 'x5' as an alternate link register. | | The alternate link register supports calling millicode routines (e.g., those to save and restore registers in compressed code) while preserving the regular return address register. The register x5 was chosen as the alternate link register as it maps to a temporary in the standard calling convention, and has an encoding that is only one bit different than the regular link register. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Plain unconditional jumps (assembler pseudoinstruction J) are encoded as a JAL with _rd_\=`x0`. ![svg](_images/svg-cefe3ffbaf4d3366a0c285f44a38c33b9a4e8c4b.svg) The indirect jump instruction JALR (jump and link register) uses the I-type encoding.The target address is obtained by adding the sign-extended 12-bit I-immediate to the register _rs1_, then setting the least-significant bit of the result to zero. The address of the instruction following the jump (`pc`+4) is written to register _rd_.Register `x0` can be used as the destination if the result is not required. Plain unconditional indirect jumps (assembler pseudoinstruction JR) are encoded as a JALR with _rd_\=`x0`. Procedure returns in the standard calling convention (assembler pseudoinstruction RET) are encoded as a JALR with _rd_\=`x0`, _rs1_\=`x1`, and_imm_\=0. ![svg](_images/svg-de763b65945810102fc912b1b7817ef6b7b3b94d.svg) | | The unconditional jump instructions all use PC-relative addressing to help support position-independent code. The JALR instruction was defined to enable a two-instruction sequence to jump anywhere in a 32-bit absolute address range. A LUI instruction can first load _rs1_ with the upper 20 bits of a target address, then JALR can add in the lower bits. Similarly, AUIPC then JALR can jump anywhere in a 32-bit pc\-relative address range. Note that the JALR instruction does not treat the 12-bit immediate as multiples of 2 bytes, unlike the conditional branch instructions. This avoids one more immediate format in hardware. In practice, most uses of JALR will have either a zero immediate or be paired with a LUI or AUIPC, so the slight reduction in range is not significant. Clearing the least-significant bit when calculating the JALR target address both simplifies the hardware slightly and allows the low bit of function pointers to be used to store auxiliary information. Although there is potentially a slight loss of error checking in this case, in practice jumps to an incorrect instruction address will usually quickly raise an exception. When used with a base _rs1_\=x0, JALR can be used to implement a single instruction subroutine call to the lowest or highest address region from anywhere in the address space, which could be used to implement fast calls to a small runtime library. Alternatively, an ABI could dedicate a general-purpose register to point to a library elsewhere in the address space. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The JAL and JALR instructions will generate an instruction-address-misaligned exception if the target address is not aligned to a four-byte boundary. | | Instruction-address-misaligned exceptions are not possible on machines with IALIGN=16, e.g., those with the compressed instruction-set extension, C. | | ------------------------------------------------------------------------------------------------------------------------------------------------------- | Return-address prediction stacks are a common feature of high-performance instruction-fetch units, but require accurate detection of instructions used for procedure calls and returns to be effective. For RISC-V, hints as to the instructions' usage are encoded implicitly via the register numbers used. A JAL instruction should push the return address onto a return-address stack (RAS) only when _rd_ is 'x1' or`x5`. JALR instructions should push/pop a RAS as shown in [Table 2](#rashints). __Table 2\. Return-address stack prediction hints encoded in the register operands of a JALR instruction.__ | _rd_ is _x1/x5_ | _rs1_ is _x1/x5_ | _rd_\=_rs1_ | RAS action | | --------------- | ---------------- | ----------- | -------------- | | No | No | — | None | | No | Yes | — | Pop | | Yes | No | — | Push | | Yes | Yes | No | Pop, then push | | Yes | Yes | Yes | Push | | | Some other ISAs added explicit hint bits to their indirect-jump instructions to guide return-address stack manipulation. We use implicit hinting tied to register numbers and the calling convention to reduce the encoding space used for these hints. When two different link registers (x1 and x5) are given as _rs1_ and_rd_, then the RAS is both popped and pushed to support coroutines. If_rs1_ and _rd_ are the same link register (either x1 or x5), the RAS is only pushed to enable macro-op fusion of the sequences:lui ra, imm20; jalr ra, imm12(ra)\_ and \_auipc ra, imm20; jalr ra, imm12(ra) | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#2-1-5-2-conditional-branches)2.1.5.2\. Conditional Branches All branch instructions use the B-type instruction format.The 12-bit B-immediate encodes signed offsets in multiples of 2 bytes. The offset is sign-extended and added to the address of the branch instruction to give the target address.The conditional branch range is ±4 KiB. ![svg](_images/svg-a8224bf5157effe30b1680ab0e042981042b17c5.svg) Branch instructions compare two registers.BEQ and BNE take the branch if registers _rs1_ and _rs2_ are equal or unequal respectively. BLT and BLTU take the branch if _rs1_ is less than _rs2_, using signed and unsigned comparison respectively. BGE and BGEU take the branch if _rs1_ is greater than or equal to _rs2_, using signed and unsigned comparison respectively.Note, BGT, BGTU, BLE, and BLEU can be synthesized by reversing the operands to BLT, BLTU, BGE, and BGEU, respectively. | | Signed array bounds may be checked with a single BLTU instruction, since any negative index will compare greater than any nonnegative bound. | | ----------------------------------------------------------------------------------------------------------------------------------------------- | Software should be optimized such that the sequential code path is the most common path, with less-frequently taken code paths placed out of line. Software should also assume that backward branches will be predicted taken and forward branches as not taken, at least the first time they are encountered. Dynamic predictors should quickly learn any predictable branch behavior. Unlike some other architectures, the RISC-V jump (JAL with _rd_\=`x0`) instruction should always be used for unconditional branches instead of a conditional branch instruction with an always-true condition. RISC-V jumps are also PC-relative and support a much wider offset range than branches, and will not pollute conditional-branch prediction tables. | | The conditional branches were designed to include arithmetic comparison operations between two registers (as also done in PA-RISC, Xtensa, and MIPS R6), rather than use condition codes (x86, ARM, SPARC, PowerPC), or to only compare one register against zero (Alpha, MIPS), or two registers only for equality (MIPS). This design was motivated by the observation that a combined compare-and-branch instruction fits into a regular pipeline, avoids additional condition code state or use of a temporary register, and reduces static code size and dynamic instruction fetch traffic. Another point is that comparisons against zero require non-trivial circuit delay (especially after the move to static logic in advanced processes) and so are almost as expensive as arithmetic magnitude compares. Another advantage of a fused compare-and-branch instruction is that branches are observed earlier in the front-end instruction stream, and so can be predicted earlier. There is perhaps an advantage to a design with condition codes in the case where multiple branches can be taken based on the same condition codes, but we believe this case to be relatively rare. We considered but did not include static branch hints in the instruction encoding. These can reduce the pressure on dynamic predictors, but require more instruction encoding space and software profiling for best results, and can result in poor performance if production runs do not match profiling runs. We considered but did not include conditional moves or predicated instructions, which can effectively replace unpredictable short forward branches. Conditional moves are the simpler of the two, but are difficult to use with conditional code that might cause exceptions (memory accesses and floating-point operations). Predication adds additional flag state to a system, additional instructions to set and clear flags, and additional encoding overhead on every instruction. Both conditional move and predicated instructions add complexity to out-of-order microarchitectures, adding an implicit third source operand due to the need to copy the original value of the destination architectural register into the renamed destination physical register if the predicate is false. Also, static compile-time decisions to use predication instead of branches can result in lower performance on inputs not included in the compiler training set, especially given that unpredictable branches are rare, and becoming rarer as branch prediction techniques improve. We note that various microarchitectural techniques exist to dynamically convert unpredictable short forward branches into internally predicated code to avoid the cost of flushing pipelines on a branch mispredict \[[13](../biblio/bibliography.html#bib-heil-tr1996)\], \[[14](../biblio/bibliography.html#bib-klauser-1998)\], \[[15](../biblio/bibliography.html#bib-kim-micro2005)\] and have been implemented in commercial processors \[[16](../biblio/bibliography.html#bib-ibmpower7)\]. The simplest techniques just reduce the penalty of recovering from a mispredicted short forward branch by only flushing instructions in the branch shadow instead of the entire fetch pipeline, or by fetching instructions from both sides using wide instruction fetch or idle instruction fetch slots. More complex techniques for out-of-order cores add internal predicates on instructions in the branch shadow, with the internal predicate value written by the branch instruction, allowing the branch and following instructions to be executed speculatively and out-of-order with respect to other code. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The conditional branch instructions will generate an instruction-address-misaligned exception if the target address is not aligned to a four-byte boundary and the branch condition evaluates to true. If the branch condition evaluates to false, the instruction-address-misaligned exception will not be raised. | | Instruction-address-misaligned exceptions are not possible on machines with IALIGN=16, e.g., those with the compressed instruction-set extension, C. | | ------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#ldst)2.1.6\. Load and Store Instructions RV32I is a load-store architecture, where only load and store instructions access memory and arithmetic instructions only operate on CPU registers. RV32I provides a 32-bit address space that is byte-addressed. The EEI will define what portions of the address space are legal to access with which instructions (e.g., some addresses might be read only, or support word access only).Loads with a destination of`x0` must still raise any exceptions and cause any other side effects even though the load value is discarded. The EEI will define whether the memory system is little-endian or big-endian. In RISC-V, endianness is byte-address invariant. | | In a system for which endianness is byte-address invariant, the following property holds: if a byte is stored to memory at some address in some endianness, then a byte-sized load from that address in any endianness returns the stored value. In a little-endian configuration, multibyte stores write the least-significant register byte at the lowest memory byte address, followed by the other register bytes in ascending order of their significance. Loads similarly transfer the contents of the lesser memory byte addresses to the less-significant register bytes. In a big-endian configuration, multibyte stores write the most-significant register byte at the lowest memory byte address, followed by the other register bytes in descending order of their significance. Loads similarly transfer the contents of the greater memory byte addresses to the less-significant register bytes. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ![svg](_images/svg-55b16975ccd74c790db66243bd8e9253cb886cfd.svg) ![svg](_images/svg-75ac5b7e9cdf4f4060af8d22fce5258257a6bdf2.svg) Load and store instructions transfer a value between the registers and memory. Loads are encoded in the I-type format and stores are S-type.The effective address is obtained by adding register _rs1_ to the sign-extended 12-bit offset. Loads copy a value from memory to register _rd_. Stores copy the value in register _rs2_ to memory. The LW instruction loads a 32-bit value from memory into _rd_. LH loads a 16-bit value from memory, then sign-extends to 32-bits before storing in _rd_. LHU loads a 16-bit value from memory but then zero extends to 32-bits before storing in _rd_. LB and LBU are defined analogously for 8-bit values. The SW, SH, and SB instructions store 32-bit, 16-bit, and 8-bit values from the low bits of register _rs2_ to memory. Regardless of EEI, loads and stores whose effective addresses are naturally aligned shall not raise an address-misaligned exception. Loads and stores whose effective address is not naturally aligned to the referenced datatype (i.e., the effective address is not divisible by the size of the access in bytes) have behavior dependent on the EEI. An EEI may guarantee that misaligned loads and stores are fully supported, and so the software running inside the execution environment will never experience a contained or fatal address-misaligned trap. In this case, themisaligned loads and stores can be handled in hardware, orvia an invisible trap into the execution environment implementation, or possibly acombination of hardware and invisible trap depending on address. An EEI may not guarantee misaligned loads and stores are handled invisibly. In this case,loads and stores that are not naturally aligned may either complete execution successfully or raise an exception. The exception raised can be either an address-misaligned exception or an access-fault exception. For a memory access that would otherwise be able to complete except for the misalignment, an access-fault exception can be raised instead of an address-misaligned exception if the misaligned access should not be emulated, e.g., if accesses to the memory region have side effects. When an EEI does not guarantee misaligned loads and stores are handled invisibly, the EEI must define if exceptions caused by address misalignment result in a contained trap (allowing software running inside the execution environment to handle the trap) or a fatal trap (terminating execution). | | Misaligned accesses are occasionally required when porting legacy code, and help performance on applications when using any form of packed-SIMD extension or handling externally packed data structures. Our rationale for allowing EEIs to choose to support misaligned accesses via the regular load and store instructions is to simplify the addition of misaligned hardware support. One option would have been to disallow misaligned accesses in the base ISAs and then provide some separate ISA support for misaligned accesses, either special instructions to help software handle misaligned accesses or a new hardware addressing mode for misaligned accesses. Special instructions are difficult to use, complicate the ISA, and often add new processor state (e.g., SPARC VIS align address offset register) or complicate access to existing processor state (e.g., MIPS LWL/LWR partial register writes). In addition, for loop-oriented packed-SIMD code, the extra overhead when operands are misaligned motivates software to provide multiple forms of loop depending on operand alignment, which complicates code generation and adds to loop startup overhead. New misaligned hardware addressing modes take considerable space in the instruction encoding or require very simplified addressing modes (e.g., register indirect only). | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Even when misaligned loads and stores complete successfully, these accesses might run extremely slowly depending on the implementation (e.g., when implemented via an invisible trap). Furthermore, whereasnaturally aligned loads and stores are guaranteed to execute atomically, misaligned loads and stores might not, and hence require additional synchronization to ensure atomicity. | | We do not mandate atomicity for misaligned accesses so execution environment implementations can use an invisible machine trap and a software handler to handle some or all misaligned accesses. If hardware misaligned support is provided, software can exploit this by simply using regular load and store instructions. Hardware can then automatically optimize accesses depending on whether runtime addresses are aligned. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#fence)2.1.7\. Memory Ordering Instructions ![mem-order](_images/mem-order-5b7b23bbf33298adbc6def76ec51ceb04bec6dd9.svg) FENCE instructions are used to order device I/O and memory accesses as viewed by other RISC-V harts and external devices or coprocessors. Any combination of device input (I), device output (O), memory reads (R), and memory writes (W) may be ordered with respect to any combination of the same. Informally, no other RISC-V hart or external device can observe any operation in the _successor_ set following a FENCE before any operation in the _predecessor_ set preceding the FENCE.[RVWMO Memory Consistency Model](rvwmo.html) provides a precise description of the RISC-V memory consistency model. FENCE instructions also order memory reads and writes made by the hart as observed by memory reads and writes made by an external device. However, FENCE instructions do not order observations of events made by an external device using any other signaling mechanism. | | A device might observe an access to a memory location via some external communication mechanism, e.g., a memory-mapped control register that drives an interrupt signal to an interrupt controller. This communication is outside the scope of the FENCE ordering mechanism and hence FENCE instructions can provide no guarantee on when a change in the interrupt signal is visible to the interrupt controller. Specific devices might provide additional ordering guarantees to reduce software overhead but those are outside the scope of the RISC-V memory model. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The EEI will define what I/O operations are possible, and in particular, which memory addresses when accessed by load and store instructions will be treated and ordered as device input and device output operations respectively rather than memory reads and writes. For example, memory-mapped I/O devices will typically be accessed with uncached loads and stores that are ordered using the I and O bits rather than the R and W bits. Instruction-set extensions might also describe new I/O instructions that will also be ordered using the I and O bits in a FENCE instruction. __Table 3\. Fence mode encoding__ | _fm_ field | Mnemonic suffix | Meaning | | ---------- | --------------- | --------------------------------------------------------------------------------------- | | 0000 | _none_ | Normal Fence | | 1000 | .TSO | With FENCE RW,RW: exclude write-to-read ordering; otherwise: _Reserved for future use._ | | _other_ | _other_ | _Reserved for future use._ | The FENCE mode field _fm_ defines the semantics of the FENCE instruction.A `FENCE`(with _fm_\=`0000`) orders all memory operations in its predecessor set before all memory operations in its successor set. A `FENCE.TSO` instruction is encoded as a FENCE instruction with _fm_\=`1000`, _predecessor_\=`RW`, and _successor_\=`RW`.`FENCE.TSO` orders all load operations in its predecessor set before all memory operations in its successor set, and all store operations in its predecessor set before all store operations in its successor set. This leaves `non-AMO`store operations in the `FENCE.TSO’s` predecessor set unordered with`non-AMO` loads in its successor set. | | Because FENCE RW,RW imposes a superset of the orderings that FENCE.TSOimposes, it is correct to ignore the _fm_ field and implement FENCE.TSO as FENCE RW,RW. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | The unused fields in the FENCE instructions--_rs1_ and _rd_\--are reserved for finer-grain fences in future extensions. For forward compatibility, base implementations shall ignore these fields, and standard software shall zero these fields. Likewise, many _fm_ and predecessor/successor set settings are also reserved for future use. Base implementations shall treat all such reserved configurations as`FENCE` instructions (with _fm_\=`0000`), and standard software shall use only non-reserved configurations. | | We chose a relaxed memory model to allow high performance from simple machine implementations and from likely future coprocessor or accelerator extensions. We separate out I/O ordering from memory R/W ordering to avoid unnecessary serialization within a device-driver hart and also to support alternative non-memory paths to control added coprocessors or I/O devices. Simple implementations may additionally ignore the _predecessor_ and _successor_ fields and always execute a conservative FENCE on all operations. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#ecall-ebreak)2.1.8\. Environment Call and Breakpoints `SYSTEM` instructions are used to access system functionality that might require privileged access and are encoded using the I-type instruction format. These can be divided into two main classes: those that atomically read-modify-write control and status registers (CSRs), and all other potentially privileged instructions. CSR instructions are described in [CSR Instructions](zicsr.html#csrinsts), and the base unprivileged instructions are described in the following section. | | The SYSTEM instructions are defined to allow simpler implementations to always trap to a single software trap handler. More sophisticated implementations might execute more of each system instruction in hardware. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ![svg](_images/svg-b638a11ee7e95c4caedda73399822484c82e2451.svg) These two instructions cause a precise requested trap to the supporting execution environment. The `ECALL` instruction is used to make a service request to the execution environment.The `EEI` will define how parameters for the service request are passed, but usually these will be in defined locations in the integer register file. The `EBREAK` instruction is used to return control to a debugging environment. | | ECALL and EBREAK were previously named SCALL and SBREAK. The instructions have the same functionality and encoding, but were renamed to reflect that they can be used more generally than to call a supervisor-level operating system or debugger. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | EBREAK was primarily designed to be used by a debugger to cause execution to stop and fall back into the debugger. EBREAK is also used by the standard GCC compiler to mark code paths that should not be executed. Another use of EBREAK is to support "semihosting", where the execution environment includes a debugger that can provide services over an alternate system call interface built around the EBREAK instruction. Because the RISC-V base ISAs do not provide more than one EBREAK instruction, RISC-V semihosting uses a special sequence of instructions to distinguish a semihosting EBREAK from a debugger inserted EBREAK. slli x0, x0, 0x1f # Entry NOP ebreak # Break to debugger srai x0, x0, 7 # Exit NOP Note that these three instructions must be 32-bit-wide instructions, i.e., they mustn’t be among the compressed 16-bit instructions described in ["C" Extension for Compressed Instructions](c-st-ext.html). The shift NOP instructions are still considered available for use as HINTs. Semihosting is a form of service call and would be more naturally encoded as an ECALL using an existing ABI, but this would require the debugger to be able to intercept ECALLs, which is a newer addition to the debug standard. We intend to move over to using ECALLs with a standard ABI, in which case, semihosting can share a service ABI with an existing standard. We note that ARM processors have also moved to using SVC instead of BKPT for semihosting calls in newer designs. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#2-1-9-hint-instructions)2.1.9\. HINT Instructions RV32I reserves a large encoding space for HINT instructions, which are usually used to communicate performance hints to the microarchitecture. Like the NOP instruction, HINTs do not change any architecturally visible state, except for advancing the `pc` and any applicable performance counters. Implementations are always allowed to ignore the encoded hints. Most RV32I HINTs are encoded as integer computational instructions with_rd_\=`x0`. The other RV32I HINTs are encoded as FENCE instructions with a null predecessor or successor set and with _fm_\=0. | | These HINT encodings have been chosen so that simple implementations can ignore HINTs altogether, and instead execute a HINT as a regular instruction that happens not to mutate the architectural state. For example, ADD is a HINT if the destination register is x0; the five-bit_rs1_ and _rs2_ fields encode arguments to the HINT. However, a simple implementation can simply execute the HINT as an ADD of _rs1_ and _rs2_that writes x0, which has no architecturally visible effect. As another example, a FENCE instruction with a zero _pred_ field and a zero _fm_ field is a HINT; the _succ_, _rs1_, and _rd_ fields encode the arguments to the HINT. A simple implementation can simply execute the HINT as a FENCE that orders the null set of prior memory accesses before whichever subsequent memory accesses are encoded in the _succ_ field. Since the intersection of the predecessor and successor sets is null, the instruction imposes no memory orderings, and so it has no architecturally visible effect. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | [Table 4](#t-rv32i-hints) lists all RV32I HINT code points. 91% of the HINT space is reserved for standard HINTs. The remainder of the HINT space is designated for custom HINTs: no standard HINTs will ever be defined in this subspace. | | We anticipate standard hints to eventually include memory-system spatial and temporal locality hints, branch prediction hints, thread-scheduling hints, security tags, and instrumentation flags for simulation/emulation. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 4\. RV32I HINT instructions.__ | Instruction | Constraints | Code Points | Purpose | | ----------- | -------------------------------------------------------------- | ----------- | --------------------------------------------------------------------------- | | LUI | _rd_\=x0 | 220 | _Designated for future standard use_ | | AUIPC | _rd_\=x0 | 220 | | | ADDI | _rd_\=x0, and either _rs1_≠x0 or _imm_≠0 | 217−1 | | | ANDI | _rd_\=x0 | 217 | | | ORI | _rd_\=x0 | 217 | | | XORI | _rd_\=x0 | 217 | | | ADD | _rd_\=x0, _rs1_≠x0 | 210−32 | | | ADD | _rd_\=x0, _rs1_\=x0, _rs2_≠x2-x5 | 28 | | | ADD | _rd_\=x0, _rs1_\=x0, _rs2_\=x2-x5 | 4 | (_rs2_\=x2) NTL.P1(_rs2_\=x3) NTL.PALL(_rs2_\=x4) NTL.S1(_rs2_\=x5) NTL.ALL | | SLLI | _rd_\=x0, _rs1_\=x0, _shamt_\=31 | 1 | Semihosting entry marker | | SRAI | _rd_\=x0, _rs1_\=x0, _shamt_\=7 | 1 | Semihosting exit marker | | SUB | _rd_\=x0 | 210 | _Designated for future standard use_ | | AND | _rd_\=x0 | 210 | | | OR | _rd_\=x0 | 210 | | | XOR | _rd_\=x0 | 210 | | | SLL | _rd_\=x0 | 210 | | | SRL | _rd_\=x0 | 210 | | | SRA | _rd_\=x0 | 210 | | | FENCE | _rd_\=x0, _rs1_≠x0, _fm_\=0, and either _pred_\=0 or _succ_\=0 | 210−63 | | | FENCE | _rd_≠x0, _rs1_\=x0, _fm_\=0, and either _pred_\=0 or _succ_\=0 | 210−63 | | | FENCE | _rd_\=_rs1_\=x0, _fm_\=0, _pred_\=0, _succ_≠0 | 15 | | | FENCE | _rd_\=_rs1_\=x0, _fm_\=0, _pred_≠W, _succ_\=0 | 15 | | | FENCE | _rd_\=_rs1_\=x0, _fm_\=0, _pred_\=W, _succ_\=0 | 1 | PAUSE | | SLTI | _rd_\=x0 | 217 | _Designated for custom use_ | | SLTIU | _rd_\=x0 | 217 | | | SLLI | _rd_\=x0, and either _rs1_≠x0 or _shamt_≠31 | 210−1 | | | SRLI | _rd_\=x0 | 210 | | | SRAI | _rd_\=x0, and either _rs1_≠x0 or _shamt_≠7 | 210−1 | | | SLT | _rd_\=x0 | 210 | | | SLTU | _rd_\=x0 | 210 | | | | slli x0, x0, 0x1f and srai x0, x0, 7 were previously designated as custom HINTs, but they have been appropriated for use in semihosting calls, as described in [2.1.8\. Environment Call and Breakpoints](#ecall-ebreak). To reflect their usage in practice, the base ISA spec has been changed to designate them as standard HINTs. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 3.1. RV32E and RV64E Base Integer Instruction Sets, Version 2.0 ==================== ## [](#rv32e)3.1\. RV32E and RV64E Base Integer Instruction Sets, Version 2.0 This chapter describes the RV32E and RV64E base integer instruction sets, designed for microcontrollers in embedded systems. RV32E and RV64E are reduced versions of RV32I and RV64I, respectively: the only change is to reduce the number of integer registers to 16\. This chapter only outlines the differences between RV32E/RV64E and RV32I/RV64I, and so should be read after [RV32I Base Integer Instruction Set, Version 2.1](rv32.html) and [RV64I Base Integer Instruction Set, Version 2.1](rv64.html). | | RV32E was designed to provide an even smaller base core for embedded microcontrollers. There is also interest in RV64E for microcontrollers within large SoC designs, and to reduce context state for highly threaded 64-bit processors. Unless otherwise stated, standard extensions compatible with RV32I and RV64I are also compatible with RV32E and RV64E, respectively. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#3-1-1-rv32e-and-rv64e-programmers-model)3.1.1\. RV32E and RV64E Programmers’ Model RV32E and RV64E reduce the integer register count to 16 general-purpose registers, (`x0-x15`), where `x0` is a dedicated zero register. | | We have found that in the small RV32I core implementations, the upper 16 registers consume around one quarter of the total area of the core excluding memories, thus their removal saves around 25% core area with a corresponding core power reduction. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#3-1-2-rv32e-and-rv64e-instruction-set-encoding)3.1.2\. RV32E and RV64E Instruction Set Encoding RV32E and RV64E use the same instruction-set encoding as RV32I and RV64I respectively, except that only registers `x0-x15` are provided. All encodings specifying the other registers `x16-x31` are reserved. | | The previous draft of this chapter made all encodings using thex16-x31 registers available as custom. This version takes a more conservative approach, making these reserved so that they can be allocated between custom space or new standard encodings at a later date. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 4.1. RV64I Base Integer Instruction Set, Version 2.1 ==================== ## [](#rv64)4.1\. RV64I Base Integer Instruction Set, Version 2.1 This chapter describes the RV64I base integer instruction set, which builds upon the RV32I variant described in [RV32I Base Integer Instruction Set, Version 2.1](rv32.html). This chapter presents only the differences with RV32I, so should be read in conjunction with the earlier chapter. ### [](#4-1-1-register-state)4.1.1\. Register State RV64I widens the integer registers and supported user address space to 64 bits (XLEN=64 in [RISC-V base unprivileged integer register state](rv32.html#gprs)). ### [](#4-1-2-integer-computational-instructions)4.1.2\. Integer Computational Instructions Most integer computational instructions operate on XLEN-bit values. Additional instruction variants are provided to manipulate 32-bit values in RV64I, indicated by a 'W' suffix to the opcode. These "\*W" instructions ignore the upper 32 bits of their inputs and always produce 32-bit signed values, sign-extending them to 64 bits, i.e. bits XLEN-1 through 31 are equal. | | The compiler and calling convention maintain an invariant that all 32-bit values are held in a sign-extended format in 64-bit registers. Even 32-bit unsigned integers extend bit 31 into bits 63 through 32\. Consequently, conversion between unsigned and signed 32-bit integers is a no-op, as is conversion from a signed 32-bit integer to a signed 64-bit integer. Existing 64-bit wide SLTU and unsigned branch compares still operate correctly on unsigned 32-bit integers under this invariant. Similarly, existing 64-bit wide logical operations on 32-bit sign-extended integers preserve the sign-extension property. A few new instructions (ADD\[I\]W/SUBW/SxxW) are required for addition and shifts to ensure reasonable performance for 32-bit values. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#4-1-2-1-integer-register-immediate-instructions)4.1.2.1\. Integer Register-Immediate Instructions ![svg](_images/svg-a31d7ad37b38827d4e3cffab485145976b2c3d03.svg) ADDIW is an RV64I instruction that adds the sign-extended 12-bit immediate to register _rs1_ and produces the proper sign extension of a 32-bit result in _rd_. Overflows are ignored and the result is the low 32 bits of the result sign-extended to 64 bits. Note, ADDIW _rd, rs1, 0_writes the sign extension of the lower 32 bits of register _rs1_ into register _rd_ (assembler pseudoinstruction SEXT.W). ![svg](_images/svg-89922c26a50881c1e9eb5b0d038f197dc3fb92c7.svg) Shifts by a constant are encoded as a specialization of the I-type format using the same instruction opcode as RV32I. The operand to be shifted is in _rs1_, and the shift amount is encoded in the lower 6 bits of the I-immediate field for RV64I. The right-shift type is encoded in bit 30\. SLLI is a logical left shift (zeros are shifted into the lower bits); SRLI is a logical right shift (zeros are shifted into the upper bits); and SRAI is an arithmetic right shift (the original sign bit is copied into the vacated upper bits). ![svg](_images/svg-8f8300f7dbbc139755c77536430b10888ae71356.svg) SLLIW, SRLIW, and SRAIW are RV64I-only instructions that are analogously defined but operate on 32-bit values and sign-extend their 32-bit results to 64 bits. SLLIW, SRLIW, and SRAIW encodings with_imm\[5\] ≠ 0_ are reserved. | | Previously, SLLIW, SRLIW, and SRAIW with _imm\[5\] ≠ 0_were defined to cause illegal-instruction exceptions, whereas now they are marked as reserved. This is a backwards-compatible change. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ![svg](_images/svg-c0abc0bfc80eceb73d353e08a13e7a84ed863093.svg) LUI (load upper immediate) uses the same opcode as RV32I.LUI places the 32-bit U-immediate into register _rd_, filling in the lowest 12 bits with zeros. The 32-bit result is sign-extended to 64 bits. AUIPC (add upper immediate to `pc`) uses the same opcode as RV32I. AUIPC is used to build `pc`\-relative addresses and uses the U-type format.AUIPC forms a 32-bit offset from the U-immediate, filling in the lowest 12 bits with zeros, sign-extends the result to 64 bits, adds it to the address of the AUIPC instruction, then places the result in register _rd_. | | Note that the set of address offsets that can be formed by pairing LUI with LD, AUIPC with JALR, etc. in RV64I is \[−231−211, 231−211−1\]. | | --------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#4-1-2-2-integer-register-register-operations)4.1.2.2\. Integer Register-Register Operations ![svg](_images/svg-dd4eb281d1e740043dc217cedc320a457b30d75e.svg) ADDW and SUBW are RV64I-only instructions that are defined analogously to ADD and SUB but operate on 32-bit values and produce signed 32-bit results. Overflows are ignored, and the low 32-bits of the result is sign-extended to 64-bits and written to the destination register. SLL, SRL, and SRA perform logical left, logical right, and arithmetic right shifts on the value in register _rs1_ by the shift amount held in register _rs2_. In RV64I, only the low 6 bits of _rs2_ are considered for the shift amount. SLLW, SRLW, and SRAW are RV64I-only instructions that are analogously defined but operate on 32-bit values and sign-extend their 32-bit results to 64 bits. The shift amount is given by _rs2\[4:0\]_. ### [](#4-1-3-load-and-store-instructions)4.1.3\. Load and Store Instructions RV64I extends the address space to 64 bits. The execution environment will define what portions of the address space are legal to access. ![svg](_images/svg-55b16975ccd74c790db66243bd8e9253cb886cfd.svg) ![svg](_images/svg-75ac5b7e9cdf4f4060af8d22fce5258257a6bdf2.svg) The LD instruction loads a 64-bit value from memory into register _rd_for RV64I. The LW instruction loads a 32-bit value from memory and sign-extends this to 64 bits before storing it in register _rd_ for RV64I. The LWU instruction, on the other hand, zero-extends the 32-bit value from memory for RV64I. LH and LHU are defined analogously for 16-bit values, as are LB and LBU for 8-bit values. The SD, SW, SH, and SB instructions store 64-bit, 32-bit, 16-bit, and 8-bit values from the low bits of register _rs2_ to memory respectively. ### [](#rv64i-hints)4.1.4\. HINT Instructions All instructions that are microarchitectural HINTs in RV32I (see[RV32I Base Integer Instruction Set, Version 2.1](rv32.html)) are also HINTs in RV64I.# The additional computational instructions in RV64I expand both the standard and custom HINT encoding spaces. [Table 1](#rv64i-h) lists all RV64I HINT code points. 91% of the HINT space is reserved for standard HINTs. The remainder of the HINT space is designated for custom HINTs; no standard HINTs will ever be defined in this subspace. __Table 1\. RV64I HINT instructions.__ | Instruction | Constraints | Code Points | Purpose | | ----------- | -------------------------------------------------------------- | ----------- | -------------------------------------------------------------------------------------- | | LUI | _rd_\=x0 | 220 | _Designated for future standard use_ | | AUIPC | _rd_\=x0 | 220 | | | ADDI | _rd_\=x0, and either _rs1_≠\`x0\` or _imm_≠0 | 217−1 | | | ANDI | _rd_\=x0 | 217 | | | ORI | _rd_\=x0 | 217 | | | XORI | _rd_\=x0 | 217 | | | ADDIW | _rd_\=x0 | 217 | | | ADD | _rd_\=x0, _rs1_≠\`x0\` | 210−32 | | | ADD | _rd_\=x0, _rs1_\=x0, _rs2_≠\_x2\_-_x5_ | 28 | | | ADD | _rd_\=x0, _rs1_\=x0, _rs2_\=_x2_\-_x5_ | 4 | (_rs2_\=_x2_) NTL.P1 (_rs2_\=_x3_) NTL.PALL (_rs2_\=_x4_) NTL.S1 (_rs2_\=_x5_) NTL.ALL | | SLLI | _rd_\=x0, _rs1_\=x0, _shamt_\=31 | 1 | Semihosting entry marker | | SRAI | _rd_\=x0, _rs1_\=x0, _shamt_\=7 | 1 | Semihosting exit marker | | SUB | _rd_\=x0 | 210 | _Designated for future standard use_ | | AND | _rd_\=x0 | 210 | | | OR | _rd_\=x0 | 210 | | | XOR | _rd_\=x0 | 210 | | | SLL | _rd_\=x0 | 210 | | | SRL | _rd_\=x0 | 210 | | | SRA | _rd_\=x0 | 210 | | | ADDW | _rd_\=x0 | 210 | | | SUBW | _rd_\=x0 | 210 | | | SLLW | _rd_\=x0 | 210 | | | SRLW | _rd_\=x0 | 210 | | | SRAW | _rd_\=x0 | 210 | | | FENCE | _rd_\=x0, _rs1_≠x0, _fm_\=0, and either _pred_\=0 or _succ_\=0 | 210−63 | | | FENCE | _rd_≠x0, _rs1_\=x0, _fm_\=0, and either _pred_\=0 or _succ_\=0 | 210−63 | | | FENCE | _rd_\=_rs1_\=x0, _fm_\=0, _pred_\=0, _succ_≠0 | 15 | | | FENCE | _rd_\=_rs1_\=x0, _fm_\=0, _pred_≠W, _succ_\=0 | 15 | | | FENCE | _rd_\=_rs1_\=x0, _fm_\=0, _pred_\=W, _succ_\=0 | 1 | PAUSE | | SLTI | _rd_\=x0 | 217 | _Designated for custom use_ | | SLTIU | _rd_\=x0 | 217 | | | SLLI | _rd_\=x0, and either _rs1_≠x0 or _shamt_≠31 | 211−1 | | | SRLI | _rd_\=x0 | 211 | | | SRAI | _rd_\=x0, and either _rs1_≠x0 or _shamt_≠7 | 211−1 | | | SLLIW | _rd_\=x0 | 210 | | | SRLIW | _rd_\=x0 | 210 | | | SRAIW | _rd_\=x0 | 210 | | | SLT | _rd_\=x0 | 210 | | | SLTU | _rd_\=x0 | 210 | | | | slli x0, x0, 0x1f and srai x0, x0, 7 were previously designated as custom HINTs, but they have been appropriated for use in semihosting calls, as described in [Environment Call and Breakpoints](rv32.html#ecall-ebreak). To reflect their usage in practice, the base ISA spec has been changed to designate them as standard HINTs. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 18.1. RVWMO Memory Consistency Model, Version 2.0 ==================== ## [](#memorymodel)18.1\. RVWMO Memory Consistency Model, Version 2.0 This chapter defines the RISC-V memory consistency model. A memory consistency model is a set of rules specifying the values that can be returned by loads of memory. RISC-V uses a memory model called "RVWMO" (RISC-V Weak Memory Ordering) which is designed to provide flexibility for architects to build high-performance scalable designs while simultaneously supporting a tractable programming model. Under RVWMO, code running on a single hart appears to execute in order from the perspective of other memory instructions in the same hart, but memory instructions from another hart may observe the memory instructions from the first hart being executed in a different order. Therefore, multithreaded code may require explicit synchronization to guarantee ordering between memory instructions from different harts. The base RISC-V ISA provides a FENCE instruction for this purpose, described in [Memory Ordering Instructions](rv32.html#fence), while the atomics extension "A" additionally defines load-reserved/store-conditional and atomic read-modify-write instructions. The standard ISA extension for total store ordering "Ztso" (["Ztso" Extension for Total Store Ordering](ztso-st-ext.html)) augments RVWMO with additional rules specific to those extensions. The appendices to this specification provide both axiomatic and operational formalizations of the memory consistency model as well as additional explanatory material. | | This chapter defines the memory model for regular main memory operations. The interaction of the memory model with I/O memory, instruction fetches, FENCE.I, page-table walks, and SFENCE.VMA is not (yet) formalized. Some or all of the above may be formalized in a future revision of this specification. Future ISA extensions such as the V vector and J JIT extensions will need to be incorporated into a future revision as well. Memory consistency models supporting overlapping memory accesses of different widths simultaneously remain an active area of academic research and are not yet fully understood. The specifics of how memory accesses of different sizes interact under RVWMO are specified to the best of our current abilities, but they are subject to revision should new issues be uncovered. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#rvwmo)18.1.1\. Definition of the RVWMO Memory Model The RVWMO memory model is defined in terms of the _global memory order_, a total ordering of the memory operations produced by all harts. In general, a multithreaded program has many different possible executions, with each execution having its own corresponding global memory order. The global memory order is defined over the primitive load and store operations generated by memory instructions. It is then subject to the constraints defined in the rest of this chapter. Any execution satisfying all of the memory model constraints is a legal execution (as far as the memory model is concerned). #### [](#rvwmo-primitives)18.1.1.1\. Memory Model Primitives The _program order_ over memory operations reflects the order in which the instructions that generate each load and store are logically laid out in that hart’s dynamic instruction stream; i.e., the order in which a simple in-order processor would execute the instructions of that hart. Memory-accessing instructions give rise to _memory operations_. A memory operation can be either a _load operation_, a _store operation_, or both simultaneously. All memory operations are single-copy atomic: they can never be observed in a partially complete state. Each aligned memory instruction that accesses XLEN or fewer bits gives rise to exactly one memory operation, unless specified otherwise. An aligned AMO gives rise to a single memory operation that is both a load operation and a store operation simultaneously. | | Among instructions in RV32GC and RV64GC, the following are exceptions to the rule that an aligned memory instruction gives rise to exactly one memory operation: An unsuccessful SC instruction does not give rise to any memory operations. Floating-point load and store instructions that access more than XLEN bits (e.g., FLD/FSD in RV32) may each give rise to multiple memory operations. ISA extensions such as **V** (Vector) and the upcoming **P** (SIMD) may give rise to multiple memory operations. However, the memory model for these extensions has not yet been formalized. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | A misaligned load or store instruction may be decomposed into a set of component memory operations of any granularity. A floating-point load or store of more than XLEN bits may also be decomposed into a set of component memory operations of any granularity. The memory operations generated by such instructions are not ordered with respect to each other in program order, but they are ordered normally with respect to the memory operations generated by preceding and subsequent instructions in program order. The atomics extension "A" does not require execution environments to support misaligned atomic instructions at all. However, if misaligned atomics are supported via the misaligned atomicity granule PMA, then AMOs within an atomicity granule are not decomposed, nor are loads and stores defined in the base ISAs, nor are loads and stores of no more than XLEN bits defined in the F, D, and Q extensions. | | The decomposition of misaligned memory operations down to byte granularity facilitates emulation on implementations that do not natively support misaligned accesses. Such implementations might, for example, simply iterate over the bytes of a misaligned access one by one. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | An LR instruction and an SC instruction are said to be _paired_ if the LR precedes the SC in program order and if there are no other LR or SC instructions in between; the corresponding memory operations are said to be paired as well (except in case of a failed SC, where no store operation is generated). The complete list of conditions determining whether an SC must succeed, may succeed, or must fail is defined in[Load-Reserved/Store-Conditional Instructions](a-st-ext.html#sec:lrsc). Load and store operations may also carry one or more ordering annotations from the following set: "acquire-RCpc", "acquire-RCsc", "release-RCpc", and "release-RCsc". An AMO or LR instruction with_aq_ set has an "acquire-RCsc" annotation. An AMO or SC instruction with _rl_ set has a "release-RCsc" annotation. An AMO, LR, or SC instruction with both _aq_ and _rl_ set has both "acquire-RCsc" and "release-RCsc" annotations. For convenience, we use the term "acquire annotation" to refer to an acquire-RCpc annotation or an acquire-RCsc annotation. Likewise, a "release annotation" refers to a release-RCpc annotation or a release-RCsc annotation. An "RCpc annotation" refers to an acquire-RCpc annotation or a release-RCpc annotation. An _RCsc annotation_ refers to an acquire-RCsc annotation or a release-RCsc annotation. | | In the memory model literature, the term "RCpc" stands for release consistency with processor-consistent synchronization operations, and the term "RCsc" stands for release consistency with sequentially consistent synchronization operations. While there are many different definitions for acquire and release annotations in the literature, in the context of RVWMO these terms are concisely and completely defined by [Preserved Program Order](#ppo) rules 5-7. "RCpc" annotations are currently only used when implicitly assigned to every memory access per the standard extension "Ztso" (["Ztso" Extension for Total Store Ordering](ztso-st-ext.html)). Furthermore, although the ISA does not currently contain native load-acquire or store-release instructions, nor RCpc variants thereof, the RVWMO model itself is designed to be forwards-compatible with the potential addition of any or all of the above into the ISA in a future extension. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#mem-dependencies)18.1.1.2\. Syntactic Dependencies The definition of the RVWMO memory model depends in part on the notion of a syntactic dependency, defined as follows. In the context of defining dependencies, a _register_ refers either to an entire general-purpose register, some portion of a CSR, or an entire CSR. The granularity at which dependencies are tracked through CSRs is specific to each CSR and is defined in[18.1.2\. CSR Dependency Tracking Granularity](#csr-granularity). Syntactic dependencies are defined in terms of instructions' _source registers_, instructions' _destination registers_, and the way instructions _carry a dependency_ from their source registers to their destination registers. This section provides a general definition of all of these terms; however, [18.1.3\. Source and Destination Register Listings](#source-dest-regs) provides a complete listing of the specifics for each instruction. In general, a register _r_ other than `x0` is a _source register_ for an instruction _i_ if any of the following hold: * In the opcode of _i_, _rs1_, _rs2_, or _rs3_ is set to_r_ * _i_ is a CSR instruction, and in the opcode of_i_, _csr_ is set to _r_, unless _i_is CSRRW or CSRRWI and _rd_ is set to `x0` * _r_ is a CSR and an implicit source register for_i_, as defined in [18.1.3\. Source and Destination Register Listings](#source-dest-regs) * _r_ is a CSR that aliases with another source register for_i_ Memory instructions also further specify which source registers are_address source registers_ and which are _data source registers_. In general, a register _r_ other than `x0` is a _destination register_ for an instruction _i_ if any of the following hold: * In the opcode of _i_, _rd_ is set to _r_ * _i_ is a CSR instruction, and in the opcode of_i_, _csr_ is set to _r_, unless _i_is CSRRS or CSRRC and _rs1_ is set to `x0` or _i_ is CSRRSI or CSRRCI and uimm\[4:0\] is set to zero. * _r_ is a CSR and an implicit destination register for_i_, as defined in [18.1.3\. Source and Destination Register Listings](#source-dest-regs) * _r_ is a CSR that aliases with another destination register for _i_ Most non-memory instructions _carry a dependency_ from each of their source registers to each of their destination registers. However, there are exceptions to this rule; see [18.1.3\. Source and Destination Register Listings](#source-dest-regs). Instruction _j_ has a _syntactic dependency_ on instruction_i_ via destination register _s_ of_i_ and source register _r_ of _j_if either of the following hold: * _s_ is the same as _r_, and no instruction program-ordered between _i_ and _j_ has_r_ as a destination register * There is an instruction _m_ program-ordered between_i_ and _j_ such that all of the following hold: 1. _j_ has a syntactic dependency on _m_ via destination register _q_ and source register _r_ 2. _m_ has a syntactic dependency on _i_ via destination register _s_ and source register _p_ 3. _m_ carries a dependency from _p_ to_q_ Finally, in the definitions that follow, let _a_ and_b_ be two memory operations, and let _i_ and_j_ be the instructions that generate _a_ and_b_, respectively. _b_ has a _syntactic address dependency_ on _a_if _r_ is an address source register for _j_ and_j_ has a syntactic dependency on _i_ via source register _r_ _b_ has a _syntactic data dependency_ on _a_ if_b_ is a store operation, _r_ is a data source register for _j_, and _j_ has a syntactic dependency on _i_ via source register _r_ _b_ has a _syntactic control dependency_ on _a_if there is an instruction _m_ program-ordered between_i_ and _j_ such that _m_ is a branch or indirect jump and _m_ has a syntactic dependency on _i_. | | Generally speaking, non-AMO load instructions do not have data source registers, and unconditional non-AMO store instructions do not have destination registers. However, a successful SC instruction is considered to have the register specified in _rd_ as a destination register, and hence it is possible for an instruction to have a syntactic dependency on a successful SC instruction that precedes it in program order. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#18-1-1-3-preserved-program-order)18.1.1.3\. Preserved Program Order The global memory order for any given execution of a program respects some but not all of each hart’s program order. The subset of program order that must be respected by the global memory order is known as_preserved program order_. The complete definition of preserved program order is as follows (and note that AMOs are simultaneously both loads and stores): memory operation _a_ precedes memory operation _b_ in preserved program order (and hence also in the global memory order) if_a_ precedes _b_ in program order,_a_ and _b_ both access regular main memory (rather than I/O regions), and any of the following hold: * Overlapping-Address Orderings: 1. _b_ is a store, and_a_ and _b_ access overlapping memory addresses 2. _a_ and _b_ are loads,_x_ is a byte read by both _a_ and_b_, there is no store to _x_ between_a_ and _b_ in program order, and_a_ and _b_ return values for _x_written by different memory operations 3. _a_ is generated by an AMO or SC instruction, _b_ is a load, and_b_ returns a value written by _a_ * Explicit Synchronization 1. There is a FENCE instruction that orders _a_ before _b_ 2. _a_ has an acquire annotation 3. _b_ has a release annotation 4. _a_ and _b_ both have RCsc annotations 5. _a_ is paired with_b_ * Syntactic Dependencies 1. _b_ has a syntactic address dependency on _a_ 2. _b_ has a syntactic data dependency on _a_ 3. _b_ is a store, and_b_ has a syntactic control dependency on _a_ * Pipeline Dependencies 1. _b_ is a load, and there exists some store _m_ between_a_ and _b_ in program order such that_m_ has an address or data dependency on _a_, and _b_ returns a value written by _m_ 2. _b_ is a store, and there exists some instruction _m_ between _a_and _b_ in program order such that _m_ has an address dependency on _a_ #### [](#18-1-1-4-memory-model-axioms)18.1.1.4\. Memory Model Axioms An execution of a RISC-V program obeys the RVWMO memory consistency model only if there exists a global memory order conforming to preserved program order and satisfying the _load value axiom_, the _atomicity axiom_, and the _progress axiom_. ##### [](#ax-load)18.1.1.4.1\. Load Value Axiom Each byte of each load _i_ returns the value written to that byte by the store that is the latest in global memory order among the following stores: 1. Stores that write that byte and that precede _i_ in the global memory order 2. Stores that write that byte and that precede _i_ in program order ##### [](#ax-atom)18.1.1.4.2\. Atomicity Axiom If _r_ and _w_ are paired load and store operations generated by aligned LR and SC instructions in a hart_h_, _s_ is a store to byte _x_, and_r_ returns a value written by _s_, then_s_ must precede _w_ in the global memory order, and there can be no store from a hart other than _h_ to byte_x_ following _s_ and preceding _w_in the global memory order. | | The [Atomicity Axiom](#ax-atom) theoretically supports LR/SC pairs of different widths and to mismatched addresses, since implementations are permitted to allow SC operations to succeed in such cases. However, in practice, we expect such patterns to be rare, and their use is discouraged. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#ax-prog)18.1.1.4.3\. Progress Axiom No memory operation may be preceded in the global memory order by an infinite sequence of other memory operations. ### [](#csr-granularity)18.1.2\. CSR Dependency Tracking Granularity __Table 1\. Granularities at which syntactic dependencies are tracked through CSRs__ | Name | Portions Tracked as Independent Units | Aliases | | -------- | ------------------------------------- | --------------- | | _fflags_ | Bits 4, 3, 2, 1, 0 | _fcsr_ | | _frm_ | entire CSR | _fcsr_ | | _fcsr_ | Bits 7-5, 4, 3, 2, 1, 0 | _fflags_, _frm_ | Note: read-only CSRs are not listed, as they do not participate in the definition of syntactic dependencies. ### [](#source-dest-regs)18.1.3\. Source and Destination Register Listings This section provides a concrete listing of the source and destination registers for each instruction. These listings are used in the definition of syntactic dependencies in[18.1.1.2\. Syntactic Dependencies](#mem-dependencies). The term "accumulating CSR" is used to describe a CSR that is both a source and a destination register, but which carries a dependency only from itself to itself. Instructions carry a dependency from each source register in the "Source Registers" column to each destination register in the "Destination Registers" column, from each source register in the "Source Registers" column to each CSR in the "Accumulating CSRs" column, and from each CSR in the "Accumulating CSRs" column to itself, except where annotated otherwise. Key: * AAddress source register * DData source register * † The instruction does not carry a dependency from any source register to any destination register * ‡ The instruction carries dependencies from source register(s) to destination register(s) as specified __Table 2\. RV32I Base Integer Instruction Set__ | Source Registers | Destination Registers | Accumulating CSRs | | | --------------------------------------------------------------------------- | --------------------- | ----------------- | ---------------------- | | LUI | _rd_ | | | | AUIPC | _rd_ | | | | JAL | _rd_ | | | | JALR† | _rs1_ | _rd_ | | | BEQ | _rs1_, _rs2_ | | | | BNE | _rs1_, _rs2_ | | | | BLT | _rs1_, _rs2_ | | | | BGE | _rs1_, _rs2_ | | | | BLTU | _rs1_, _rs2_ | | | | BGEU | _rs1_, _rs2_ | | | | LB † | _rs1_ A | _rd_ | | | LH † | _rs1_ A | _rd_ | | | LW † | _rs1_ A | _rd_ | | | LBU † | _rs1_ A | _rd_ | | | LHU † | _rs1_ A | _rd_ | | | SB | _rs1_ A, _rs2_ D | | | | SH | _rs1_ A, _rs2_ D | | | | SW | _rs1_ A, _rs2_ D | | | | ADDI | _rs1_ | _rd_ | | | SLTI | _rs1_ | _rd_ | | | SLTIU | _rs1_ | _rd_ | | | XORI | _rs1_ | _rd_ | | | ORI | _rs1_ | _rd_ | | | ANDI | _rs1_ | _rd_ | | | SLLI | _rs1_ | _rd_ | | | SRLI | _rs1_ | _rd_ | | | SRAI | _rs1_ | _rd_ | | | ADD | _rs1_, _rs2_ | _rd_ | | | SUB | _rs1_, _rs2_ | _rd_ | | | SLL | _rs1_, _rs2_ | _rd_ | | | SLT | _rs1_, _rs2_ | _rd_ | | | SLTU | _rs1_, _rs2_ | _rd_ | | | XOR | _rs1_, _rs2_ | _rd_ | | | SRL | _rs1_, _rs2_ | _rd_ | | | SRA | _rs1_, _rs2_ | _rd_ | | | OR | _rs1_, _rs2_ | _rd_ | | | AND | _rs1_, _rs2_ | _rd_ | | | FENCE | | | | | FENCE.I | | | | | ECALL | | | | | EBREAK | | | | | CSRRW‡ | _rs1_, _csr_\* | _rd_, _csr_ | \*unless _rd_\=x0 | | ‡ carries a dependency from _rs1_ to _csr_ and from _csr_ to _rd_ | | | | | CSRRS‡ | _rs1_, _csr_ | _rd_, _csr_\* | \*unless _rs1_\=x0 | | CSRRC‡ | _rs1_, _csr_ | _rd_, _csr_\* | \*unless _rs1_\=x0 | | ‡ carries a dependency from _csr_ and _rs1_ to _csr_ and from _csr_ to _rd_ | | | | | CSRRWI ‡ | _csr_ \* | _rd_, _csr_ | \*unless _rd_\=_x0_ | | ‡ carries a dependency from _csr_ to _rd_ | | | | | CSRRSI ‡ | _csr_ | _rd_, _csr_\* | \*unless uimm\[4:0\]=0 | | CSRRCI ‡ | _csr_ | _rd_, _csr_\* | \*unless uimm\[4:0\]=0 | | ‡ carries a dependency from _csr_ to _rd_ and _csr_ | | | | __Table 3\. RV64I Base Integer Instruction Set__ | Source Registers | Destination Registers | Accumulating CSRs | | ---------------- | --------------------- | ----------------- | | _LWU_ † | _rs1_ A | _rd_ | | _LD_ † | _rs1_ A | _rd_ | | SD | _rs1_ A, _rs2_ D | | | SLLI | _rs1_ | _rd_ | | SRLI | _rs1_ | _rd_ | | SRAI | _rs1_ | _rd_ | | ADDIW | _rs1_ | _rd_ | | SLLIW | _rs1_ | _rd_ | | SRLIW | _rs1_ | _rd_ | | SRAIW | _rs1_ | _rd_ | | ADDW | _rs1_, _rs2_ | _rd_ | | SUBW | _rs1_, _rs2_ | _rd_ | | SLLW | _rs1_, _rs2_ | _rd_ | | SRLW | _rs1_, _rs2_ | _rd_ | | SRAW | _rs1_, _rs2_ | _rd_ | __Table 4\. RV32M Standard Extension__ | Source Registers | Destination Registers | Accumulating CSRs | | ---------------- | --------------------- | ----------------- | | MUL | _rs1_, _rs2_ | _rd_ | | MULH | _rs1_, _rs2_ | _rd_ | | MULHSU | _rs1_, _rs2_ | _rd_ | | MULHU | _rs1_, _rs2_ | _rd_ | | DIV | _rs1_, _rs2_ | _rd_ | | DIVU | _rs1_, _rs2_ | _rd_ | | REM | _rs1_, _rs2_ | _rd_ | | REMU | _rs1_, _rs2_ | _rd_ | __Table 5\. RV64M Standard Extension__ | Source Registers | Destination Registers | Accumulating CSRs | | ---------------- | --------------------- | ----------------- | | MULW | _rs1_, _rs2_ | _rd_ | | DIVW | _rs1_, _rs2_ | _rd_ | | DIVUW | _rs1_, _rs2_ | _rd_ | | REMW | _rs1_, _rs2_ | _rd_ | | REMUW | _rs1_, _rs2_ | _rd_ | __Table 6\. RV32A Standard Extension__ | Source Registers | Destination Registers | Accumulating CSRs | | | ---------------- | --------------------- | ----------------- | ---------------- | | LR.W† | _rs1_ A | _rd_ | | | SC.W† | _rs1_ A, _rs2_ D | _rd_ \* | \* if successful | | AMOSWAP.W† | _rs1_ A, _rs2_ D | _rd_ | | | AMOADD.W† | _rs1_ A, _rs2_ D | _rd_ | | | AMOXOR.W† | _rs1_ A, _rs2_ D | _rd_ | | | AMOAND.W† | _rs1_ A, _rs2_ D | _rd_ | | | AMOOR.W† | _rs1_ A, _rs2_D | _rd_ | | | AMOMIN.W† | _rs1_ A, _rs2_ D | _rd_ | | | AMOMAX.W† | _rs1_ A, _rs2_ D | _rd_ | | | AMOMINU.W† | _rs1_ A, _rs2_ D | _rd_ | | | AMOMAXU.W† | _rs1_ A, _rs2_ D | _rd_ | | __Table 7\. RV64A Standard Extension__ | Source Registers | Destination Registers | Accumulating CSRs | | | ---------------- | --------------------- | ----------------- | --------------- | | LR.D† | _rs1_ A | _rd_ | | | SC.D† | _rs1_ A, _rs2_ D | _rd_ \* | \*if successful | | AMOSWAP.D† | _rs1_ A, _rs2_ D | _rd_ | | | AMOADD.D† | _rs1_ A, _rs2_ D | _rd_ | | | AMOXOR.D† | _rs1_ A, _rs2_ D | _rd_ | | | AMOAND.D† | _rs1_ A, _rs2_D | _rd_ | | | AMOOR.D† | _rs1_ A, _rs2_D | _rd_ | | | AMOMIN.D† | _rs1_ A, _rs2_D | _rd_ | | | AMOMAX.D† | _rs1_ A, _rs2_D | _rd_ | | | AMOMINU.D† | _rs1_ A, _rs2_D | _rd_ | | | AMOMAXU.D† | _rs1_ A, _rs2_D | _rd_ | | __Table 8\. RV32F Standard Extension__ | Source Registers | Destination Registers | Accumulating CSRs | | | | ---------------- | -------------------------- | ----------------- | ------------------ | ----------- | | FLW† | _rs1_ A | _rd_ | | | | FSW | _rs1_ A, _rs2_D | | | | | FMADD.S | _rs1_, _rs2_, _rs3_, frm\* | _rd_ | NV, OF, UF, NX | \*if rm=111 | | FMSUB.S | _rs1_, _rs2_, _rs3_, frm\* | _rd_ | NV, OF, UF, NX | \*if rm=111 | | FNMSUB.S | _rs1_, _rs2_, _rs3_, frm\* | _rd_ | NV, OF, UF, NX | \*if rm=111 | | FNMADD.S | _rs1_, _rs2_, _rs3_, frm\* | _rd_ | NV, OF, UF, NX | \*if rm=111 | | FADD.S | _rs1_, _rs2_, frm\* | _rd_ | NV, OF, NX | \*if rm=111 | | FSUB.S | _rs1_, _rs2_, frm\* | _rd_ | NV, OF, NX | \*if rm=111 | | FMUL.S | _rs1_, _rs2_, frm\* | _rd_ | NV, OF, UF, NX | \*if rm=111 | | FDIV.S | _rs1_, _rs2_, frm\* | _rd_ | NV, DZ, OF, UF, NX | \*if rm=111 | | FSQRT.S | _rs1_, frm\* | _rd_ | NV, NX | \*if rm=111 | | FSGNJ.S | _rs1_, _rs2_ | _rd_ | | | | FSGNJN.S | _rs1_, _rs2_ | _rd_ | | | | FSGNJX.S | _rs1_, _rs2_ | _rd_ | | | | FMIN.S | _rs1_, _rs2_ | _rd_ | NV | | | FMAX.S | _rs1_, _rs2_ | _rd_ | NV | | | FCVT.W.S | _rs1_, frm\* | _rd_ | NV, NX | \*if rm=111 | | FCVT.WU.S | _rs1_, frm\* | _rd_ | NV, NX | \*if rm=111 | | FMV.X.W | _rs1_ | _rd_ | | | | FEQ.S | _rs1_, _rs2_ | _rd_ | NV | | | FLT.S | _rs1_, _rs2_ | _rd_ | NV | | | FLE.S | _rs1_, _rs2_ | _rd_ | NV | | | FCLASS.S | _rs1_ | _rd_ | | | | FCVT.S.W | _rs1_, frm\* | _rd_ | NX | \*if rm=111 | | FCVT.S.WU | _rs1_, frm\* | _rd_ | NX | \*if rm=111 | | FMV.W.X | _rs1_ | _rd_ | | | __Table 9\. RV64F Standard Extension__ | Source Registers | Destination Registers | Accumulating CSRs | | | | ---------------- | --------------------- | ----------------- | ------ | ----------- | | FCVT.L.S | _rs1_, frm\* | _rd_ | NV, NX | \*if rm=111 | | FCVT.LU.S | _rs1_, frm\* | _rd_ | NV, NX | \*if rm=111 | | FCVT.S.L | _rs1_, frm\* | _rd_ | NX | \*if rm=111 | | FCVT.S.LU | _rs1_, frm\* | _rd_ | NX | \*if rm=111 | __Table 10\. RV32D Standard Extension__ | Source Registers | Destination Registers | Accumulating CSRs | | | | ---------------- | -------------------------- | ----------------- | ------------------ | ----------- | | FLD† | _rs1_ A | _rd_ | | | | FSD | _rs1_ A, _rs2_D | | | | | FMADD.D | _rs1_, _rs2_, _rs3_, frm\* | _rd_ | NV, OF, UF, NX | \*if rm=111 | | FMSUB.D | _rs1_, _rs2_, _rs3_, frm\* | _rd_ | NV, OF, UF, NX | \*if rm=111 | | FNMSUB.D | _rs1_, _rs2_, _rs3_, frm\* | _rd_ | NV, OF, UF, NX | \*if rm=111 | | FNMADD.D | _rs1_, _rs2_, _rs3_, frm\* | _rd_ | NV, OF, UF, NX | \*if rm=111 | | FADD.D | _rs1_, _rs2_, frm\* | _rd_ | NV, OF, NX | \*if rm=111 | | FSUB.D | _rs1_, _rs2_, frm\* | _rd_ | NV, OF, NX | \*if rm=111 | | FMUL.D | _rs1_, _rs2_, frm\* | _rd_ | NV, OF, UF, NX | \*if rm=111 | | FDIV.D | _rs1_, _rs2_, frm\* | _rd_ | NV, DZ, OF, UF, NX | \*if rm=111 | | FSQRT.D | _rs1_, frm\* | _rd_ | NV, NX | \*if rm=111 | | FSGNJ.D | _rs1_, _rs2_ | _rd_ | | | | FSGNJN.D | _rs1_, _rs2_ | _rd_ | | | | FSGNJX.D | _rs1_, _rs2_ | _rd_ | | | | FMIN.D | _rs1_, _rs2_ | _rd_ | NV | | | FMAX.D | _rs1_, _rs2_ | _rd_ | NV | | | FCVT.S.D | _rs1_, frm\* | _rd_ | NV, OF, UF, NX | \*if rm=111 | | FCVT.D.S | _rs1_ | _rd_ | NV | | | FEQ.D | _rs1_, _rs2_ | _rd_ | NV | | | FLT.D | _rs1_, _rs2_ | _rd_ | NV | | | FLE.D | _rs1_, _rs2_ | _rd_ | NV | | | FCLASS.D | _rs1_ | _rd_ | | | | FCVT.W.D | _rs1_, frm\* | _rd_ | NV, NX | \*if rm=111 | | FCVT.WU.D | _rs1_, frm\* | _rd_ | NV, NX | \*if rm=111 | | FCVT.D.W | _rs1_ | _rd_ | | | | FCVT.D.WU | _rs1_ | _rd_ | | | __Table 11\. RV64D Standard Extension__ | Source Registers | Destination Registers | Accumulating CSRs | | | | ---------------- | --------------------- | ----------------- | ------ | ----------- | | FCVT.L.D | _rs1_, frm\* | _rd_ | NV, NX | \*if rm=111 | | FCVT.LU.D | _rs1_, frm\* | _rd_ | NV, NX | \*if rm=111 | | FMV.X.D | _rs1_ | _rd_ | | | | FCVT.D.L | _rs1_, frm\* | _rd_ | NX | \*if rm=111 | | FCVT.D.LU | _rs1_, frm\* | _rd_ | NX | \*if rm=111 | | FMV.D.X | _rs1_ | _rd_ | | | 32.1. Cryptography Extensions: Scalar & Entropy Source Instructions, Version 1.0.1 ==================== ## [](#crypto%5Fscalar%5Finstructions)32.1\. Cryptography Extensions: Scalar & Entropy Source Instructions, Version 1.0.1 ### [](#crypto%5Fscalar%5Fintroduction)32.1.1\. Introduction This document describes the _scalar_ cryptography extension for RISC-V. All instructions described herein use the general-purpose `X`registers, and obey the 2-read-1-write register access constraint. These instructions are designed to be lightweight and suitable for `32` and `64` bit base architectures; from embedded IoT class cores to large, application class cores which do not implement a vector unit. This document also describes the architectural interface to an Entropy Source, which can be used to generate cryptographic secrets. This is found in [32.1.4\. Entropy Source](#crypto%5Fscalar%5Fes). It also contains a mechanism allowing core implementers to provide_"Constant Time Execution"_ guarantees in [32.1.5\. Data Independent Execution Latency Subset: Zkt](#crypto%5Fscalar%5Fzkt). #### [](#crypto%5Fscalar%5Faudience)32.1.1.1\. Intended Audience Cryptography is a specialised subject, requiring people with many different backgrounds to cooperate in its secure and efficient implementation. Where possible, we have written this specification to be understandable by all, though we recognise that the motivations and references to algorithms or other specifications and standards may be unfamiliar to those who are not domain experts. This specification anticipates being read and acted on by various people with different backgrounds. We have tried to capture these backgrounds here, with a brief explanation of what we expect them to know, and how it relates to the specification. We hope this aids people’s understanding of which aspects of the specification are particularly relevant to them, and which they may (safely!) ignore or pass to a colleague. Cryptographers and cryptographic software developers These are the people we expect to write code using the instructions in this specification. They should understand fairly obviously the motivations for the instructions we include, and be familiar with most of the algorithms and outside standards to which we refer. We expect the sections on constant time execution ([32.1.5\. Data Independent Execution Latency Subset: Zkt](#crypto%5Fscalar%5Fzkt)) and the entropy source ([32.1.4\. Entropy Source](#crypto%5Fscalar%5Fes)) to be chiefly understood with their help. Computer architects We do not expect architects to have a cryptography background. We nonetheless expect architects to be able to examine our instructions for implementation issues, understand how the instructions will be used in context, and advise on how best to fit the functionality the cryptographers want to the ISA interface. Digital design engineers & micro-architects These are the people who will implement the specification inside a core. Again, no cryptography expertise is assumed, but we expect them to interpret the specification and anticipate any hardware implementation issues, e.g., where high-frequency design considerations apply, or where latency/area tradeoffs exist etc. In particular, they should be aware of the literature around efficiently implementing AES and SM4 SBoxes in hardware. Verification engineers Responsible for ensuring the correct implementation of the extension in hardware. No cryptography background is assumed. We expect them to identify interesting test cases from the specification. An understanding of their real-world usage will help with this. We do not expect verification engineers in this sense to be experts in entropy source design or certification, since this is a very specialised area. We do expect them however to identify all of the _architectural_test cases around the entropy source interface. These are by no means the only people concerned with the specification, but they are the ones we considered most while writing it. #### [](#crypto%5Fscalar%5Fsail%5Fspecifications)32.1.1.2\. Sail Specifications RISC-V maintains a[formal model](https://github.com/riscv/sail-riscv)of the ISA specification, implemented in the Sail ISA specification language \[[27](../biblio/bibliography.html#bib-sail)\]. Note that _Sail_ refers to the specification language itself, and that there is a _model of RISC-V_, written using Sail. It is not correct to refer to "the Sail model". This is ambiguous, given there are many models of different ISAs implemented using Sail. We refer to the Sail implementation of RISC-V as "the RISC-V Sail model". The Cryptography extension uses inline Sail code snippets from the actual model to give canonical descriptions of instruction functionality. Each instruction is accompanied by its expression in Sail, and includes calls to supporting functions which are too verbose to include directly in the specification. This supporting code is listed in[32.1.9\. Supporting Sail Code](#crypto%5Fscalar%5Fappx%5Fsail). The[Sail Manual](https://alasdair.github.io/manual.html)is recommended reading in order to best understand the code snippets. Note that this document contains only a subset of the formal model: refer to the formal model GitHub[repository](https://github.com/riscv/sail-riscv)for the complete model. #### [](#crypto%5Fscalar%5Fpolicies)32.1.1.3\. Policies In creating this proposal, we tried to adhere to the following policies: * Where there is a choice between: 1. supporting diverse implementation strategies for an algorithm or 2. supporting a single implementation style which is more performant / less expensive; the crypto extension will pick the more constrained but performant option. This fits a common pattern in other parts of the RISC-V specification, where recommended (but not required) instruction sequences for performing particular tasks are given as an example, such that both hardware and software implementers can optimise for only a single use-case. * The extension will be designed to support _existing_ standardised cryptographic constructs well. It will not try to support proposed standards, or cryptographic constructs which exist only in academia. Cryptographic standards which are settled upon concurrently with or after the RISC-V cryptographic extension standardisation will be dealt with by future additions to, or versions of, the RISC-V cryptographic standard extension. It is anticipated that the NIST Lightweight Cryptography contest and the NIST Post-Quantum Cryptography contest may be dealt with this way, depending on timescales. * Historically, there has been some discussion \[[28](../biblio/bibliography.html#bib-lsyrr:04)\] on how newly supported operations in general-purpose computing might enable new bases for cryptographic algorithms. The standard will not try to anticipate new useful low-level operations which _may_ be useful as building blocks for future cryptographic constructs. * Regarding side-channel countermeasures: Where relevant, proposed instructions must aim to remove the possibility of any timing side-channels. For side-channels based on power or electro-magnetic (EM) measurements, the extension will not aim to support countermeasures which are implemented above the ISA abstraction layer. Recommendations will be given where relevant on how micro-architectures can implement instructions in a power/EM side-channel resistant way. ### [](#crypto%5Fscalar%5Fextensions)32.1.2\. Extensions Overview The group of extensions introduced by the Scalar Cryptography Instruction Set Extension is listed here. Detection of individual cryptography extensions uses the unified software-based RISC-V discovery method. | | At the time of writing, these discovery mechanisms are still a work in progress. | | ----------------------------------------------------------------------------------- | | | A note on extension rationale Specialist encryption and decryption instructions are separated into different functional groups because some use cases (e.g., Galois/Counter Mode in TLS 1.3) do not require decryption functionality. The NIST and ShangMi algorithms suites are separated because their usefulness is heavily dependent on the countries a device is expected to operate in. NIST ciphers are a part of most standardised internet protocols, while ShangMi ciphers are required for use in China. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#zbkb-sc)32.1.2.1\. `Zbkb` \- Bitmanip instructions for Cryptography This extension contains bit-manipulation instructions that are particularly useful for cryptography, most of which are also in the `Zbb` extension. Please refer to [Bit-manipulation for Cryptography](b-st-ext.html#zbkb) for more information. #### [](#zbkc-sc)32.1.2.2\. `Zbkc` \- Carry-less multiply instructions Constant time carry-less multiply for Galois/Counter Mode. These are separated from the [Bit-manipulation for Cryptography](b-st-ext.html#zbkb) because they have a considerable implementation overhead which cannot be amortised across other instructions. Please refer to [Carry-less multiplication for Cryptography](b-st-ext.html#zbkc). #### [](#zbkx-sc)32.1.2.3\. `Zbkx` \- Crossbar permutation instructions These instructions are useful for implementing SBoxes in constant time, and potentially with DPA protections. These are separated from the [Bit-manipulation for Cryptography](b-st-ext.html#zbkb) because they have an implementation overhead which cannot be amortised across other instructions. Please refer to [Crossbar permutation instructions](b-st-ext.html#zbkx). #### [](#zknd)32.1.2.4\. `Zknd` \- NIST Suite: AES Decryption Instructions for accelerating the decryption and key-schedule functions of the AES block cipher. | RV32 | RV64 | Mnemonic | Instruction | | ---- | --------- | ----------------------------------------------------------- | ----------- | | ✓ | aes32dsi | [AES final round decrypt (RV32)](#insns-aes32dsi) | | | ✓ | aes32dsmi | [AES middle round decrypt (RV32)](#insns-aes32dsmi) | | | ✓ | aes64ds | [AES decrypt final round (RV64)](#insns-aes64ds) | | | ✓ | aes64dsm | [AES decrypt middle round (RV64)](#insns-aes64dsm) | | | ✓ | aes64im | [AES Decrypt KeySchedule MixColumns (RV64)](#insns-aes64im) | | | ✓ | aes64ks1i | [AES Key Schedule Instruction 1 (RV64)](#insns-aes64ks1i) | | | ✓ | aes64ks2 | [AES Key Schedule Instruction 2 (RV64)](#insns-aes64ks2) | | | | The [AES Key Schedule Instruction 1 (RV64)](#insns-aes64ks1i) and [AES Key Schedule Instruction 2 (RV64)](#insns-aes64ks2) instructions are present in both the [Zknd](#zknd) and [Zkne](#zkne) extensions. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#zkne)32.1.2.5\. `Zkne` \- NIST Suite: AES Encryption Instructions for accelerating the encryption and key-schedule functions of the AES block cipher. | RV32 | RV64 | Mnemonic | Instruction | | ---- | --------- | -------------------------------------------------------------- | ----------- | | ✓ | aes32esi | [AES final round encrypt (RV32)](#insns-aes32esi) | | | ✓ | aes32esmi | [AES middle round encrypt (RV32)](#insns-aes32esmi) | | | ✓ | aes64es | [AES encrypt final round instruction (RV64)](#insns-aes64es) | | | ✓ | aes64esm | [AES encrypt middle round instruction (RV64)](#insns-aes64esm) | | | ✓ | aes64ks1i | [AES Key Schedule Instruction 1 (RV64)](#insns-aes64ks1i) | | | ✓ | aes64ks2 | [AES Key Schedule Instruction 2 (RV64)](#insns-aes64ks2) | | | | The[aes64ks1i](#insns-aes64ks1i)and[aes64ks2](#insns-aes64ks2)instructions are present in both the [Zknd](#zknd) and [Zkne](#zkne) extensions. | | ------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#zknh)32.1.2.6\. `Zknh` \- NIST Suite: Hash Function Instructions Instructions for accelerating the SHA2 family of cryptographic hash functions, as specified in \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ----------- | ------------------------------------------------------- | ------------------------------------------------ | | ✓ | ✓ | sha256sig0 | [SHA2-256 Sigma0 instruction](#insns-sha256sig0) | | ✓ | ✓ | sha256sig1 | [SHA2-256 Sigma1 instruction](#insns-sha256sig1) | | ✓ | ✓ | sha256sum0 | [SHA2-256 Sum0 instruction](#insns-sha256sum0) | | ✓ | ✓ | sha256sum1 | [SHA2-256 Sum1 instruction](#insns-sha256sum1) | | ✓ | sha512sig0h | [SHA2-512 Sigma0 high (RV32)](#insns-sha512sig0h) | | | ✓ | sha512sig0l | [SHA2-512 Sigma0 low (RV32)](#insns-sha512sig0l) | | | ✓ | sha512sig1h | [SHA2-512 Sigma1 high (RV32)](#insns-sha512sig1h) | | | ✓ | sha512sig1l | [SHA2-512 Sigma1 low (RV32)](#insns-sha512sig1l) | | | ✓ | sha512sum0r | [SHA2-512 Sum0 (RV32)](#insns-sha512sum0r) | | | ✓ | sha512sum1r | [SHA2-512 Sum1 (RV32)](#insns-sha512sum1r) | | | ✓ | sha512sig0 | [SHA2-512 Sigma0 instruction (RV64)](#insns-sha512sig0) | | | ✓ | sha512sig1 | [SHA2-512 Sigma1 instruction (RV64)](#insns-sha512sig1) | | | ✓ | sha512sum0 | [SHA2-512 Sum0 instruction (RV64)](#insns-sha512sum0) | | | ✓ | sha512sum1 | [SHA2-512 Sum1 instruction (RV64)](#insns-sha512sum1) | | #### [](#zksed)32.1.2.7\. `Zksed` \- ShangMi Suite: SM4 Block Cipher Instructions Instructions for accelerating the SM4 Block Cipher. Note that unlike AES, this cipher uses the same core operation for encryption and decryption, hence there is only one extension for it. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | -------- | ----------------------------------------------- | | ✓ | ✓ | sm4ed | [SM4 Encrypt/Decrypt Instruction](#insns-sm4ed) | | ✓ | ✓ | sm4ks | [SM4 Key Schedule Instruction](#insns-sm4ks) | #### [](#zksh)32.1.2.8\. `Zksh` \- ShangMi Suite: SM3 Hash Function Instructions Instructions for accelerating the SM3 hash function. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | -------- | -------------------------------- | | ✓ | ✓ | sm3p0 | [SM3 P0 transform](#insns-sm3p0) | | ✓ | ✓ | sm3p1 | [SM3 P1 transform](#insns-sm3p1) | #### [](#zkr)32.1.2.9\. `Zkr` \- Entropy Source Extension The entropy source extension defines the `seed` CSR at address `0x015`. This CSR provides up to 16 physical `entropy` bits that can be used to seed cryptographic random bit generators. See [32.1.4\. Entropy Source](#crypto%5Fscalar%5Fes) for the normative specification and access control notes. [32.1.7\. Entropy Source Rationale and Recommendations](#crypto%5Fscalar%5Fappx%5Fes) contains design rationale and further recommendations to implementers. #### [](#zkn)32.1.2.10\. `Zkn` \- NIST Algorithm Suite This extension is shorthand for the following set of other extensions: | Included Extension | Description | | ------------------ | ---------------------------------------------- | | [Zbkb](#zbkb-sc) | Bitmanipulation instructions for cryptography. | | [Zbkc](#zbkc-sc) | Carry-less multiply instructions. | | [Zbkx](#zbkx-sc) | Cross-bar Permutation instructions. | | [Zkne](#zkne) | AES encryption instructions. | | [Zknd](#zknd) | AES decryption instructions. | | [Zknh](#zknh) | SHA2 hash function instructions. | A core which implements `Zkn` must implement all of the above extensions. #### [](#zks)32.1.2.11\. `Zks` \- ShangMi Algorithm Suite This extension is shorthand for the following set of other extensions: | Included Extension | Description | | ------------------ | ---------------------------------------------- | | [Zbkb](#zbkb-sc) | Bitmanipulation instructions for cryptography. | | [Zbkc](#zbkc-sc) | Carry-less multiply instructions. | | [Zbkx](#zbkx-sc) | Cross-bar Permutation instructions. | | [Zksed](#zksed) | SM4 block cipher instructions. | | [Zksh](#zksh) | SM3 hash function instructions. | A core which implements `Zks` must implement all of the above extensions. #### [](#zk)32.1.2.12\. `Zk` \- Standard scalar cryptography extension This extension is shorthand for the following set of other extensions: | Included Extension | Description | | ----------------------------- | --------------------------------------------- | | [Zkn](#zkn) | NIST Algorithm suite extension. | | [Zkr](#zkr) | Entropy Source extension. | | [Zkt](#crypto%5Fscalar%5Fzkt) | Data independent execution latency extension. | A core which implements `Zk` must implement all of the above extensions. #### [](#32-1-2-13-zkt-data-independent-execution-latency)32.1.2.13\. `Zkt` \- Data Independent Execution Latency This extension allows CPU implementers to indicate to cryptographic software developers that a subset of RISC-V instructions are guaranteed to be implemented such that their execution latency is independent of the data values they operate on. A complete description of this extension is found in[32.1.5\. Data Independent Execution Latency Subset: Zkt](#crypto%5Fscalar%5Fzkt). ### [](#crypto%5Fscalar%5Finsns)32.1.3\. Instructions #### [](#insns-aes32dsi)32.1.3.1\. aes32dsi Synopsis AES final round decryption instruction for RV32. Mnemonic aes32dsi rd, rs1, rs2, bs Encoding ![svg](_images/svg-8f8a38b97eef4940a4a1ccdf6b03aa1dcd60c677.svg) Description This instruction sources a single byte from `rs2` according to `bs`. To this it applies the inverse AES SBox operation, and XOR’s the result with`rs1`. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (AES32DSI (bs,rs2,rs1,rd)) = { let shamt : bits( 5) = bs @ 0b000; /* shamt = bs*8 */ let si : bits( 8) = (X(rs2)[31..0] >> shamt)[7..0]; /* SBox Input */ let so : bits(32) = 0x000000 @ aes_sbox_inv(si); let result : bits(32) = X(rs1)[31..0] ^ rol32(so, unsigned(shamt)); X(rd) = EXTS(result); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknd](#zknd) (RV32) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV32) | v1.0.0 | Ratified | | [Zk](#zk) (RV32) | v1.0.0 | Ratified | #### [](#insns-aes32dsmi)32.1.3.2\. aes32dsmi Synopsis AES middle round decryption instruction for RV32. Mnemonic aes32dsmi rd, rs1, rs2, bs Encoding ![svg](_images/svg-c907e44765d8c3761bdce516e5fed2cfe9d53e36.svg) Description This instruction sources a single byte from `rs2` according to `bs`. To this it applies the inverse AES SBox operation, and a partial inverse MixColumn, before XOR’ing the result with `rs1`. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (AES32DSMI (bs,rs2,rs1,rd)) = { let shamt : bits( 5) = bs @ 0b000; /* shamt = bs*8 */ let si : bits( 8) = (X(rs2)[31..0] >> shamt)[7..0]; /* SBox Input */ let so : bits( 8) = aes_sbox_inv(si); let mixed : bits(32) = aes_mixcolumn_byte_inv(so); let result : bits(32) = X(rs1)[31..0] ^ rol32(mixed, unsigned(shamt)); X(rd) = EXTS(result); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknd](#zknd) (RV32) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV32) | v1.0.0 | Ratified | | [Zk](#zk) (RV32) | v1.0.0 | Ratified | #### [](#insns-aes32esi)32.1.3.3\. aes32esi Synopsis AES final round encryption instruction for RV32. Mnemonic aes32esi rd, rs1, rs2, bs Encoding ![svg](_images/svg-2b5882c83346290b5521428246ddd2bf5195e3c2.svg) Description This instruction sources a single byte from `rs2` according to `bs`. To this it applies the forward AES SBox operation, before XOR’ing the result with `rs1`. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (AES32ESI (bs,rs2,rs1,rd)) = { let shamt : bits( 5) = bs @ 0b000; /* shamt = bs*8 */ let si : bits( 8) = (X(rs2)[31..0] >> shamt)[7..0]; /* SBox Input */ let so : bits(32) = 0x000000 @ aes_sbox_fwd(si); let result : bits(32) = X(rs1)[31..0] ^ rol32(so, unsigned(shamt)); X(rd) = EXTS(result); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zkne](#zkne) (RV32) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV32) | v1.0.0 | Ratified | | [Zk](#zk) (RV32) | v1.0.0 | Ratified | #### [](#insns-aes32esmi)32.1.3.4\. aes32esmi Synopsis AES middle round encryption instruction for RV32. Mnemonic aes32esmi rd, rs1, rs2, bs Encoding ![svg](_images/svg-9fd0d1f0e51106e25ce5dbcc624c88d66a785f14.svg) Description This instruction sources a single byte from `rs2` according to `bs`. To this it applies the forward AES SBox operation, and a partial forward MixColumn, before XOR’ing the result with `rs1`. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (AES32ESMI (bs,rs2,rs1,rd)) = { let shamt : bits( 5) = bs @ 0b000; /* shamt = bs*8 */ let si : bits( 8) = (X(rs2)[31..0] >> shamt)[7..0]; /* SBox Input */ let so : bits( 8) = aes_sbox_fwd(si); let mixed : bits(32) = aes_mixcolumn_byte_fwd(so); let result : bits(32) = X(rs1)[31..0] ^ rol32(mixed, unsigned(shamt)); X(rd) = EXTS(result); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zkne](#zkne) (RV32) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV32) | v1.0.0 | Ratified | | [Zk](#zk) (RV32) | v1.0.0 | Ratified | #### [](#insns-aes64ds)32.1.3.5\. aes64ds Synopsis AES final round decryption instruction for RV64. Mnemonic aes64ds rd, rs1, rs2 Encoding ![svg](_images/svg-6d5553dcd5b6432019332015c32c2b0009567b7a.svg) Description Uses the two 64-bit source registers to represent the entire AES state, and produces _half_ of the next round output, applying the Inverse ShiftRows and SubBytes steps. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. | | Note To Software Developers The following code snippet shows the final round of the AES block decryption.t0 and t1 hold the current round state.t2 and t3 hold the next round state. aes64ds t2, t0, t1 aes64ds t3, t1, t0 Note the reversed register order of the second instruction. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (AES64DS(rs2, rs1, rd)) = { let sr : bits(64) = aes_rv64_shiftrows_inv(X(rs2)[63..0], X(rs1)[63..0]); let wd : bits(64) = sr[63..0]; X(rd) = aes_apply_inv_sbox_to_each_byte(wd); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknd](#zknd) (RV64) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV64) | v1.0.0 | Ratified | | [Zk](#zk) (RV64) | v1.0.0 | Ratified | #### [](#insns-aes64dsm)32.1.3.6\. aes64dsm Synopsis AES middle round decryption instruction for RV64. Mnemonic aes64dsm rd, rs1, rs2 Encoding ![svg](_images/svg-2428d34aa6620a4bb8bf9c3a44f6346eeda86134.svg) Description Uses the two 64-bit source registers to represent the entire AES state, and produces _half_ of the next round output, applying the Inverse ShiftRows, SubBytes and MixColumns steps. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. | | Note To Software Developers The following code snippet shows one middle round of the AES block decryption.t0 and t1 hold the current round state.t2 and t3 hold the next round state. aes64dsm t2, t0, t1 aes64dsm t3, t1, t0 Note the reversed register order of the second instruction. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (AES64DSM(rs2, rs1, rd)) = { let sr : bits(64) = aes_rv64_shiftrows_inv(X(rs2)[63..0], X(rs1)[63..0]); let wd : bits(64) = sr[63..0]; let sb : bits(64) = aes_apply_inv_sbox_to_each_byte(wd); X(rd) = aes_mixcolumn_inv(sb[63..32]) @ aes_mixcolumn_inv(sb[31..0]); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknd](#zknd) (RV64) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV64) | v1.0.0 | Ratified | | [Zk](#zk) (RV64) | v1.0.0 | Ratified | #### [](#insns-aes64es)32.1.3.7\. aes64es Synopsis AES final round encryption instruction for RV64. Mnemonic aes64es rd, rs1, rs2 Encoding ![svg](_images/svg-ffb6a60569eaeb402426288354c5d72ad15c0cac.svg) Description Uses the two 64-bit source registers to represent the entire AES state, and produces _half_ of the next round output, applying the ShiftRows and SubBytes steps. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. | | Note To Software Developers The following code snippet shows the final round of the AES block encryption.t0 and t1 hold the current round state.t2 and t3 hold the next round state. aes64es t2, t0, t1 aes64es t3, t1, t0 Note the reversed register order of the second instruction. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (AES64ES(rs2, rs1, rd)) = { let sr : bits(64) = aes_rv64_shiftrows_fwd(X(rs2)[63..0], X(rs1)[63..0]); let wd : bits(64) = sr[63..0]; X(rd) = aes_apply_fwd_sbox_to_each_byte(wd); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zkne](#zkne) (RV64) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV64) | v1.0.0 | Ratified | | [Zk](#zk) (RV64) | v1.0.0 | Ratified | #### [](#insns-aes64esm)32.1.3.8\. aes64esm Synopsis AES middle round encryption instruction for RV64. Mnemonic aes64esm rd, rs1, rs2 Encoding ![svg](_images/svg-99651b71ad39e7b5a53a80581c90a77f5bc0093a.svg) Description Uses the two 64-bit source registers to represent the entire AES state, and produces _half_ of the next round output, applying the ShiftRows, SubBytes and MixColumns steps. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. | | Note To Software Developers The following code snippet shows one middle round of the AES block encryption.t0 and t1 hold the current round state.t2 and t3 hold the next round state. aes64esm t2, t0, t1 aes64esm t3, t1, t0 Note the reversed register order of the second instruction. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (AES64ESM(rs2, rs1, rd)) = { let sr : bits(64) = aes_rv64_shiftrows_fwd(X(rs2)[63..0], X(rs1)[63..0]); let wd : bits(64) = sr[63..0]; let sb : bits(64) = aes_apply_fwd_sbox_to_each_byte(wd); X(rd) = aes_mixcolumn_fwd(sb[63..32]) @ aes_mixcolumn_fwd(sb[31..0]); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zkne](#zkne) (RV64) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV64) | v1.0.0 | Ratified | | [Zk](#zk) (RV64) | v1.0.0 | Ratified | #### [](#insns-aes64im)32.1.3.9\. aes64im Synopsis This instruction accelerates the inverse MixColumns step of the AES Block Cipher, and is used to aid creation of the decryption KeySchedule. Mnemonic aes64im rd, rs1 Encoding ![svg](_images/svg-3e774f7f82f8a54a2c243bdc1d72d04e7b2c02c2.svg) Description The instruction applies the inverse MixColumns transformation to two columns of the state array, packed into a single 64-bit register. It is used to create the inverse cipher KeySchedule, according to the equivalent inverse cipher construction in \[[30](../biblio/bibliography.html#bib-nist:fips:197)\] (Page 23, Section 5.3.5). This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (AES64IM(rs1, rd)) = { let w0 : bits(32) = aes_mixcolumn_inv(X(rs1)[31.. 0]); let w1 : bits(32) = aes_mixcolumn_inv(X(rs1)[63..32]); X(rd) = w1 @ w0; RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknd](#zknd) (RV64) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV64) | v1.0.0 | Ratified | | [Zk](#zk) (RV64) | v1.0.0 | Ratified | #### [](#insns-aes64ks1i)32.1.3.10\. aes64ks1i Synopsis This instruction implements part of the KeySchedule operation for the AES Block cipher involving the SBox operation. Mnemonic aes64ks1i rd, rs1, rnum Encoding ![svg](_images/svg-9353ce092416a482831f92880d078681e9c7de33.svg) Description This instruction implements the rotation, SubBytes and Round Constant addition steps of the AES block cipher Key Schedule. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Note that `rnum` must be in the range `0x0..0xA`. The values `0xB..0xF` are reserved. Operation ```sail function clause execute (AES64KS1I(rnum, rs1, rd)) = { if(unsigned(rnum) > 10) then { handle_illegal(); RETIRE_SUCCESS } else { let tmp1 : bits(32) = X(rs1)[63..32]; let rc : bits(32) = aes_decode_rcon(rnum); /* round number -> round constant */ let tmp2 : bits(32) = if (rnum ==0xA) then tmp1 else ror32(tmp1, 8); let tmp3 : bits(32) = aes_subword_fwd(tmp2); let result : bits(64) = (tmp3 ^ rc) @ (tmp3 ^ rc); X(rd) = EXTZ(result); RETIRE_SUCCESS } } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zkne](#zkne) (RV64) | v1.0.0 | Ratified | | [Zknd](#zknd) (RV64) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV64) | v1.0.0 | Ratified | | [Zk](#zk) (RV64) | v1.0.0 | Ratified | #### [](#insns-aes64ks2)32.1.3.11\. aes64ks2 Synopsis This instruction implements part of the KeySchedule operation for the AES Block cipher. Mnemonic aes64ks2 rd, rs1, rs2 Encoding ![svg](_images/svg-baaa646b0f56fdddacd8c8cf92dc7a7f068a3038.svg) Description This instruction implements the additional XOR’ing of key words as part of the AES block cipher Key Schedule. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (AES64KS2(rs2, rs1, rd)) = { let w0 : bits(32) = X(rs1)[63..32] ^ X(rs2)[31..0]; let w1 : bits(32) = X(rs1)[63..32] ^ X(rs2)[31..0] ^ X(rs2)[63..32]; X(rd) = w1 @ w0; RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zkne](#zkne) (RV64) | v1.0.0 | Ratified | | [Zknd](#zknd) (RV64) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV64) | v1.0.0 | Ratified | | [Zk](#zk) (RV64) | v1.0.0 | Ratified | #### [](#insns-andn-sc)32.1.3.12\. andn Synopsis AND with inverted operand Mnemonic andn _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-5037e8d0227804a70953db7915f2b4e3c251a9cc.svg) Description This instruction performs the bitwise logical AND operation between _rs1_ and the bitwise inversion of _rs2_. Operation ```sail X(rd) = X(rs1) & ~X(rs2); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbb ([Zbb](b-st-ext.html#zbb)) | 1.0.0 | Ratified | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-brev8-sc)32.1.3.13\. brev8 Synopsis Reverse the bits in each byte of a source register. Mnemonic brev8 _rd_, _rs_ Encoding ![svg](_images/svg-2fb21fe817b13f00b3265dee5e0d3eb405943c35.svg) Description This instruction reverses the order of the bits in every byte of a register. Operation ```sail result : xlenbits = EXTZ(0b0); foreach (i from 0 to sizeof(xlen) by 8) { result[i+7..i] = reverse_bits_in_byte(X(rs1)[i+7..i]); }; X(rd) = result; ``` Included in | Extension | Minimum version | Lifecycle state | | ----------------------- | --------------- | --------------- | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-clmul-sc)32.1.3.14\. clmul Synopsis Carry-less multiply (low-part) Mnemonic clmul _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-3ace2bb1d5afe94f7c9321cfdf1a432733accd80.svg) Description clmul produces the lower half of the 2·XLEN carry-less product. Operation ```sail let rs1_val = X(rs1); let rs2_val = X(rs2); let output : xlenbits = 0; foreach (i from 0 to (xlen - 1) by 1) { output = if ((rs2_val >> i) & 1) then output ^ (rs1_val << i); else output; } X[rd] = output ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbc ([Zbc](b-st-ext.html#zbc)) | 1.0.0 | Ratified | | Zbkc ([Zbkc](#zbkc-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-clmulh-sc)32.1.3.15\. clmulh Synopsis Carry-less multiply (high-part) Mnemonic clmulh _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-6107d0ddbb95c0228642e006d50cb857f98e58dc.svg) Description clmulh produces the upper half of the 2·XLEN carry-less product. Operation ```sail let rs1_val = X(rs1); let rs2_val = X(rs2); let output : xlenbits = 0; foreach (i from 1 to xlen by 1) { output = if ((rs2_val >> i) & 1) then output ^ (rs1_val >> (xlen - i)); else output; } X[rd] = output ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbc ([Zbc](b-st-ext.html#zbc)) | 1.0.0 | Ratified | | Zbkc ([Zbkc](#zbkc-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-orn-sc)32.1.3.16\. orn Synopsis OR with inverted operand Mnemonic orn _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-cd2eb5051fd881d76177c58af39944f79a35e27a.svg) Description This instruction performs the bitwise logical OR operation between _rs1_ and the bitwise inversion of _rs2_. Operation ```sail X(rd) = X(rs1) | ~X(rs2); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbb ([Zbb](b-st-ext.html#zbb)) | v1.0.0 | Ratified | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-pack-sc)32.1.3.17\. pack Synopsis Pack the low halves of _rs1_ and _rs2_ into _rd_. Mnemonic pack _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-3b154237dc0d0a29f3ee4a786ba979924f77f0a0.svg) Description The pack instruction packs the XLEN/2-bit lower halves of _rs1_ and _rs2_ into_rd_, with _rs1_ in the lower half and _rs2_ in the upper half. Operation ```sail let lo_half : bits(xlen/2) = X(rs1)[xlen/2-1..0]; let hi_half : bits(xlen/2) = X(rs2)[xlen/2-1..0]; X(rd) = EXTZ(hi_half @ lo_half); ``` Included in | Extension | Minimum version | Lifecycle state | | ----------------------- | --------------- | --------------- | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-packh-sc)32.1.3.18\. packh Synopsis Pack the low bytes of _rs1_ and _rs2_ into _rd_. Mnemonic packh _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-de742e02dd64298507cef284a266492885bc7eb3.svg) Description And the packh instruction packs the least-significant bytes of_rs1_ and _rs2_ into the 16 least-significant bits of _rd_, zero extending the rest of _rd_. Operation ```sail let lo_half : bits(8) = X(rs1)[7..0]; let hi_half : bits(8) = X(rs2)[7..0]; X(rd) = EXTZ(hi_half @ lo_half); ``` Included in | Extension | Minimum version | Lifecycle state | | ----------------------- | --------------- | --------------- | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-packw-sc)32.1.3.19\. packw Synopsis Pack the low 16-bits of _rs1_ and _rs2_ into _rd_ on RV64. Mnemonic packw _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-94ceb467b69ed6f2eecccefc8b1976078d327a0b.svg) Description This instruction packs the low 16 bits of_rs1_ and _rs2_ into the 32 least-significant bits of _rd_, sign extending the 32-bit result to the rest of _rd_. This instruction only exists on RV64 based systems. Operation ```sail let lo_half : bits(16) = X(rs1)[15..0]; let hi_half : bits(16) = X(rs2)[15..0]; X(rd) = EXTS(hi_half @ lo_half); ``` Included in | Extension | Minimum version | Lifecycle state | | ----------------------- | --------------- | --------------- | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-rev8-sc)32.1.3.20\. rev8 Synopsis Byte-reverse register Mnemonic rev8 _rd_, _rs_ Encoding (RV32) ![svg](_images/svg-693d173289db88cf4b8d29c9a4c08f5ecbe7f7e2.svg) Encoding (RV64) ![svg](_images/svg-6d52ea2a08d942672651db671ee4dca698d0b929.svg) Description This instruction reverses the order of the bytes in _rs_. Operation ```sail let input = X(rs); let output : xlenbits = 0; let j = xlen - 1; foreach (i from 0 to (xlen - 8) by 8) { output[i..(i + 7)] = input[(j - 7)..j]; j = j - 8; } X[rd] = output ``` | | Note The **rev8** mnemonic corresponds to different instruction encodings in RV32 and RV64. | | ---------------------------------------------------------------------------------------------- | | | Software Hint The byte-reverse operation is only available for the full register width. To emulate word-sized and halfword-sized byte-reversal, perform a rev8 rd,rs followed by a srai rd,rd,K, where K is XLEN-32 and XLEN-16, respectively. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbb ([Zbb](b-st-ext.html#zbb)) | v1.0.0 | Ratified | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-rol-sc)32.1.3.21\. rol Synopsis Rotate Left (Register) Mnemonic rol _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-0add38b10235680259e9ea3ba4b471b26a942d2d.svg) Description This instruction performs a rotate left of _rs1_ by the amount in least-significant log2(XLEN) bits of _rs2_. Operation ```sail let shamt = if xlen == 32 then X(rs2)[4..0] else X(rs2)[5..0]; let result = (X(rs1) << shamt) | (X(rs1) >> (xlen - shamt)); X(rd) = result; ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbb ([Zbb](b-st-ext.html#zbb)) | v1.0.0 | Ratified | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-rolw-sc)32.1.3.22\. rolw Synopsis Rotate Left Word (Register) Mnemonic rolw _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-f5b062418b2d8410926587fb72cd6fc0123e75e9.svg) Description This instruction performs a rotate left on the least-significant word of _rs1_ by the amount in least-significant 5 bits of _rs2_. The resulting word value is sign-extended by copying bit 31 to all of the more-significant bits. Operation ```sail let rs1 = EXTZ(X(rs1)[31..0]) let shamt = X(rs2)[4..0]; let result = (rs1 << shamt) | (rs1 >> (32 - shamt)); X(rd) = EXTS(result[31..0]); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbb ([Zbb](b-st-ext.html#zbb)) | v1.0.0 | Ratified | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-ror-sc)32.1.3.23\. ror Synopsis Rotate Right Mnemonic ror _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-30fb690178dfc1807bfbaadb3e95c705c92708a3.svg) Description This instruction performs a rotate right of _rs1_ by the amount in least-significant log2(XLEN) bits of _rs2_. Operation ```sail let shamt = if xlen == 32 then X(rs2)[4..0] else X(rs2)[5..0]; let result = (X(rs1) >> shamt) | (X(rs1) << (xlen - shamt)); X(rd) = result; ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbb ([Zbb](b-st-ext.html#zbb)) | v1.0.0 | Ratified | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-rori-sc)32.1.3.24\. rori Synopsis Rotate Right (Immediate) Mnemonic rori _rd_, _rs1_, _shamt_ Encoding (RV32) ![svg](_images/svg-938fcfa9f20d5623c60b26ec22337b55fe5692e6.svg) Encoding (RV64) ![svg](_images/svg-56357a7bc26ad075b2aa68de3ae2db30bc1d104d.svg) Description This instruction performs a rotate right of _rs1_ by the amount in the least-significant log2(XLEN) bits of _shamt_. For RV32, the encodings corresponding to shamt\[5\]=1 are reserved. Operation ```sail let shamt = if xlen == 32 then shamt[4..0] else shamt[5..0]; let result = (X(rs1) >> shamt) | (X(rs1) << (xlen - shamt)); X(rd) = result; ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbb ([Zbb](b-st-ext.html#zbb)) | v1.0.0 | Ratified | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-roriw-sc)32.1.3.25\. roriw Synopsis Rotate Right Word by Immediate Mnemonic roriw _rd_, _rs1_, _shamt_ Encoding ![svg](_images/svg-ade4a28eb2f571fc50a9a5567dbd7991a3b05a4c.svg) Description This instruction performs a rotate right on the least-significant word of _rs1_ by the amount in the least-significant log2(XLEN) bits of_shamt_. The resulting word value is sign-extended by copying bit 31 to all of the more-significant bits. Operation ```sail let rs1_data = EXTZ(X(rs1)[31..0]; let result = (rs1_data >> shamt) | (rs1_data << (32 - shamt)); X(rd) = EXTS(result[31..0]); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbb ([Zbb](b-st-ext.html#zbb)) | v1.0.0 | Ratified | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-rorw-sc)32.1.3.26\. rorw Synopsis Rotate Right Word (Register) Mnemonic rorw _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-565fec38b77826e629c17d1e46d0048fef634910.svg) Description This instruction performs a rotate right on the least-significant word of _rs1_ by the amount in least-significant 5 bits of _rs2_. The resultant word is sign-extended by copying bit 31 to all of the more-significant bits. Operation ```sail let rs1 = EXTZ(X(rs1)[31..0]) let shamt = X(rs2)[4..0]; let result = (rs1 >> shamt) | (rs1 << (32 - shamt)); X(rd) = EXTS(result); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbb ([Zbb](b-st-ext.html#zbb)) | v1.0.0 | Ratified | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-sha256sig0)32.1.3.27\. sha256sig0 Synopsis Implements the Sigma0 transformation function as used in the SHA2-256 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha256sig0 rd, rs1 Encoding ![svg](_images/svg-8d7da1b452c5e2260c38a7f3c229f1a1d1233cbf.svg) Description This instruction is supported for both RV32 and RV64 base architectures. For RV32, the entire `XLEN` source register is operated on. For RV64, the low `32` bits of the source register are operated on, and the result sign extended to `XLEN` bits. Though named for SHA2-256, the instruction works for both the SHA2-224 and SHA2-256 parameterizations as described in \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (SHA256SIG0(rs1,rd)) = { let inb : bits(32) = X(rs1)[31..0]; let result : bits(32) = ror32(inb, 7) ^ ror32(inb, 18) ^ (inb >> 3); X(rd) = EXTS(result); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | ------------- | --------------- | --------------- | | [Zknh](#zknh) | v1.0.0 | Ratified | | [Zkn](#zkn) | v1.0.0 | Ratified | | [Zk](#zk) | v1.0.0 | Ratified | #### [](#insns-sha256sig1)32.1.3.28\. sha256sig1 Synopsis Implements the Sigma1 transformation function as used in the SHA2-256 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha256sig1 rd, rs1 Encoding ![svg](_images/svg-84f6861aa871ef896e2a95d14dca8d2c68952eaf.svg) Description This instruction is supported for both RV32 and RV64 base architectures. For RV32, the entire `XLEN` source register is operated on. For RV64, the low `32` bits of the source register are operated on, and the result sign extended to `XLEN` bits. Though named for SHA2-256, the instruction works for both the SHA2-224 and SHA2-256 parameterizations as described in \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (SHA256SIG1(rs1,rd)) = { let inb : bits(32) = X(rs1)[31..0]; let result : bits(32) = ror32(inb, 17) ^ ror32(inb, 19) ^ (inb >> 10); X(rd) = EXTS(result); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | ------------- | --------------- | --------------- | | [Zknh](#zknh) | v1.0.0 | Ratified | | [Zkn](#zkn) | v1.0.0 | Ratified | | [Zk](#zk) | v1.0.0 | Ratified | #### [](#insns-sha256sum0)32.1.3.29\. sha256sum0 Synopsis Implements the Sum0 transformation function as used in the SHA2-256 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha256sum0 rd, rs1 Encoding ![svg](_images/svg-7339adef06280c14fc2499b7f6a9dd09234ae863.svg) Description This instruction is supported for both RV32 and RV64 base architectures. For RV32, the entire `XLEN` source register is operated on. For RV64, the low `32` bits of the source register are operated on, and the result sign extended to `XLEN` bits. Though named for SHA2-256, the instruction works for both the SHA2-224 and SHA2-256 parameterizations as described in \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (SHA256SUM0(rs1,rd)) = { let inb : bits(32) = X(rs1)[31..0]; let result : bits(32) = ror32(inb, 2) ^ ror32(inb, 13) ^ ror32(inb, 22); X(rd) = EXTS(result); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | ------------- | --------------- | --------------- | | [Zknh](#zknh) | v1.0.0 | Ratified | | [Zkn](#zkn) | v1.0.0 | Ratified | | [Zk](#zk) | v1.0.0 | Ratified | #### [](#insns-sha256sum1)32.1.3.30\. sha256sum1 Synopsis Implements the Sum1 transformation function as used in the SHA2-256 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha256sum1 rd, rs1 Encoding ![svg](_images/svg-330da71a6aaa2377257558738382463bd96f7759.svg) Description This instruction is supported for both RV32 and RV64 base architectures. For RV32, the entire `XLEN` source register is operated on. For RV64, the low `32` bits of the source register are operated on, and the result sign extended to `XLEN` bits. Though named for SHA2-256, the instruction works for both the SHA2-224 and SHA2-256 parameterizations as described in \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (SHA256SUM1(rs1,rd)) = { let inb : bits(32) = X(rs1)[31..0]; let result : bits(32) = ror32(inb, 6) ^ ror32(inb, 11) ^ ror32(inb, 25); X(rd) = EXTS(result); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | ------------- | --------------- | --------------- | | [Zknh](#zknh) | v1.0.0 | Ratified | | [Zkn](#zkn) | v1.0.0 | Ratified | | [Zk](#zk) | v1.0.0 | Ratified | #### [](#insns-sha512sig0h)32.1.3.31\. sha512sig0h Synopsis Implements the _high half_ of the Sigma0 transformation, as used in the SHA2-512 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha512sig0h rd, rs1, rs2 Encoding ![svg](_images/svg-1d12c0c3e16bfb3aa85aa3978453da392ee6aed3.svg) Description This instruction is implemented on RV32 only. Used to compute the Sigma0 transform of the SHA2-512 hash function in conjunction with the [sha512sig0l](#insns-sha512sig0l) instruction. The transform is a 64-bit to 64-bit function, so the input and output are each represented by two 32-bit registers. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. | | Note to software developers The entire Sigma0 transform for SHA2-512 may be computed on RV32 using the following instruction sequence: sha512sig0l t0, a0, a1 sha512sig0h t1, a1, a0 | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (SHA512SIG0H(rs2, rs1, rd)) = { X(rd) = EXTS((X(rs1) >> 1) ^ (X(rs1) >> 7) ^ (X(rs1) >> 8) ^ (X(rs2) << 31) ^ (X(rs2) << 24) ); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknh](#zknh) (RV32) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV32) | v1.0.0 | Ratified | | [Zk](#zk) (RV32) | v1.0.0 | Ratified | #### [](#insns-sha512sig0l)32.1.3.32\. sha512sig0l Synopsis Implements the _low half_ of the Sigma0 transformation, as used in the SHA2-512 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha512sig0l rd, rs1, rs2 Encoding ![svg](_images/svg-437dcd0be70433926b01baf0b104e5d13a30f362.svg) Description This instruction is implemented on RV32 only. Used to compute the Sigma0 transform of the SHA2-512 hash function in conjunction with the [sha512sig0h](#insns-sha512sig0h) instruction. The transform is a 64-bit to 64-bit function, so the input and output are each represented by two 32-bit registers. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. | | Note to software developers The entire Sigma0 transform for SHA2-512 may be computed on RV32 using the following instruction sequence: sha512sig0l t0, a0, a1 sha512sig0h t1, a1, a0 | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (SHA512SIG0L(rs2, rs1, rd)) = { X(rd) = EXTS((X(rs1) >> 1) ^ (X(rs1) >> 7) ^ (X(rs1) >> 8) ^ (X(rs2) << 31) ^ (X(rs2) << 25) ^ (X(rs2) << 24) ); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknh](#zknh) (RV32) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV32) | v1.0.0 | Ratified | | [Zk](#zk) (RV32) | v1.0.0 | Ratified | #### [](#insns-sha512sig1h)32.1.3.33\. sha512sig1h Synopsis Implements the _high half_ of the Sigma1 transformation, as used in the SHA2-512 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha512sig1h rd, rs1, rs2 Encoding ![svg](_images/svg-a3f5abb80aead7a81c46ec08bbe147053a50895e.svg) Description This instruction is implemented on RV32 only. Used to compute the Sigma1 transform of the SHA2-512 hash function in conjunction with the [sha512sig1l](#insns-sha512sig1l) instruction. The transform is a 64-bit to 64-bit function, so the input and output are each represented by two 32-bit registers. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. | | Note to software developers The entire Sigma1 transform for SHA2-512 may be computed on RV32 using the following instruction sequence: sha512sig1l t0, a0, a1 sha512sig1h t1, a1, a0 | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (SHA512SIG1H(rs2, rs1, rd)) = { X(rd) = EXTS((X(rs1) << 3) ^ (X(rs1) >> 6) ^ (X(rs1) >> 19) ^ (X(rs2) >> 29) ^ (X(rs2) << 13) ); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknh](#zknh) (RV32) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV32) | v1.0.0 | Ratified | | [Zk](#zk) (RV32) | v1.0.0 | Ratified | #### [](#insns-sha512sig1l)32.1.3.34\. sha512sig1l Synopsis Implements the _low half_ of the Sigma1 transformation, as used in the SHA2-512 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha512sig1l rd, rs1, rs2 Encoding ![svg](_images/svg-7eb5a210f74347df986b549091e9d08ea16226f0.svg) Description This instruction is implemented on RV32 only. Used to compute the Sigma1 transform of the SHA2-512 hash function in conjunction with the [sha512sig1h](#insns-sha512sig1h) instruction. The transform is a 64-bit to 64-bit function, so the input and output are each represented by two 32-bit registers. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. | | Note to software developers The entire Sigma1 transform for SHA2-512 may be computed on RV32 using the following instruction sequence: sha512sig1l t0, a0, a1 sha512sig1h t1, a1, a0 | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (SHA512SIG1L(rs2, rs1, rd)) = { X(rd) = EXTS((X(rs1) << 3) ^ (X(rs1) >> 6) ^ (X(rs1) >> 19) ^ (X(rs2) >> 29) ^ (X(rs2) << 26) ^ (X(rs2) << 13) ); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknh](#zknh) (RV32) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV32) | v1.0.0 | Ratified | | [Zk](#zk) (RV32) | v1.0.0 | Ratified | #### [](#insns-sha512sum0r)32.1.3.35\. sha512sum0r Synopsis Implements the Sum0 transformation, as used in the SHA2-512 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha512sum0r rd, rs1, rs2 Encoding ![svg](_images/svg-ee3b35f1c516eafc3c400fe21334f6b63777707d.svg) Description This instruction is implemented on RV32 only. Used to compute the Sum0 transform of the SHA2-512 hash function. The transform is a 64-bit to 64-bit function, so the input and output is represented by two 32-bit registers. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. | | Note to software developers The entire Sum0 transform for SHA2-512 may be computed on RV32 using the following instruction sequence: sha512sum0r t0, a0, a1 sha512sum0r t1, a1, a0 Note the reversed source register ordering. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (SHA512SUM0R(rs2, rs1, rd)) = { X(rd) = EXTS((X(rs1) << 25) ^ (X(rs1) << 30) ^ (X(rs1) >> 28) ^ (X(rs2) >> 7) ^ (X(rs2) >> 2) ^ (X(rs2) << 4) ); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknh](#zknh) (RV32) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV32) | v1.0.0 | Ratified | | [Zk](#zk) (RV32) | v1.0.0 | Ratified | #### [](#insns-sha512sum1r)32.1.3.36\. sha512sum1r Synopsis Implements the Sum1 transformation, as used in the SHA2-512 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha512sum1r rd, rs1, rs2 Encoding ![svg](_images/svg-a58fd57f04fba5d5bd8fb4fbbb42e6f9dd049e6d.svg) Description This instruction is implemented on RV32 only. Used to compute the Sum1 transform of the SHA2-512 hash function. The transform is a 64-bit to 64-bit function, so the input and output is represented by two 32-bit registers. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. | | Note to software developers The entire Sum1 transform for SHA2-512 may be computed on RV32 using the following instruction sequence: sha512sum1r t0, a0, a1 sha512sum1r t1, a1, a0 Note the reversed source register ordering. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (SHA512SUM1R(rs2, rs1, rd)) = { X(rd) = EXTS((X(rs1) << 23) ^ (X(rs1) >> 14) ^ (X(rs1) >> 18) ^ (X(rs2) >> 9) ^ (X(rs2) << 18) ^ (X(rs2) << 14) ); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknh](#zknh) (RV32) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV32) | v1.0.0 | Ratified | | [Zk](#zk) (RV32) | v1.0.0 | Ratified | #### [](#insns-sha512sig0)32.1.3.37\. sha512sig0 Synopsis Implements the Sigma0 transformation function as used in the SHA2-512 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha512sig0 rd, rs1 Encoding ![svg](_images/svg-3c1b4953013c63878410fd923d6e1cdd89fbf379.svg) Description This instruction is supported for the RV64 base architecture. It implements the Sigma0 transform of the SHA2-512 hash function. \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (SHA512SIG0(rs1, rd)) = { X(rd) = ror64(X(rs1), 1) ^ ror64(X(rs1), 8) ^ (X(rs1) >> 7); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknh](#zknh) (RV64) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV64) | v1.0.0 | Ratified | | [Zk](#zk) (RV64) | v1.0.0 | Ratified | #### [](#insns-sha512sig1)32.1.3.38\. sha512sig1 Synopsis Implements the Sigma1 transformation function as used in the SHA2-512 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha512sig1 rd, rs1 Encoding ![svg](_images/svg-8e22b1ee36b0dd12cb5af8970eba42f98faccf68.svg) Description This instruction is supported for the RV64 base architecture. It implements the Sigma1 transform of the SHA2-512 hash function. \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (SHA512SIG1(rs1, rd)) = { X(rd) = ror64(X(rs1), 19) ^ ror64(X(rs1), 61) ^ (X(rs1) >> 6); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknh](#zknh) (RV64) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV64) | v1.0.0 | Ratified | | [Zk](#zk) (RV64) | v1.0.0 | Ratified | #### [](#insns-sha512sum0)32.1.3.39\. sha512sum0 Synopsis Implements the Sum0 transformation function as used in the SHA2-512 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha512sum0 rd, rs1 Encoding ![svg](_images/svg-7e02f8ce9c30951df9bc19b1f39c3ddaa9bd6399.svg) Description This instruction is supported for the RV64 base architecture. It implements the Sum0 transform of the SHA2-512 hash function. \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (SHA512SUM0(rs1, rd)) = { X(rd) = ror64(X(rs1), 28) ^ ror64(X(rs1), 34) ^ ror64(X(rs1) ,39); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknh](#zknh) (RV64) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV64) | v1.0.0 | Ratified | | [Zk](#zk) (RV64) | v1.0.0 | Ratified | #### [](#insns-sha512sum1)32.1.3.40\. sha512sum1 Synopsis Implements the Sum1 transformation function as used in the SHA2-512 hash function \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. Mnemonic sha512sum1 rd, rs1 Encoding ![svg](_images/svg-0d26c615c98ee35ce73b9f5046fcd5ed3115d178.svg) Description This instruction is supported for the RV64 base architecture. It implements the Sum1 transform of the SHA2-512 hash function. \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (SHA512SUM1(rs1, rd)) = { X(rd) = ror64(X(rs1), 14) ^ ror64(X(rs1), 18) ^ ror64(X(rs1) ,41); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | -------------------- | --------------- | --------------- | | [Zknh](#zknh) (RV64) | v1.0.0 | Ratified | | [Zkn](#zkn) (RV64) | v1.0.0 | Ratified | | [Zk](#zk) (RV64) | v1.0.0 | Ratified | #### [](#insns-sm3p0)32.1.3.41\. sm3p0 Synopsis Implements the _P0_ transformation function as used in the SM3 hash function \[[31](../biblio/bibliography.html#bib-gbt:sm3)\] \[[32](../biblio/bibliography.html#bib-iso:sm3)\]. Mnemonic sm3p0 rd, rs1 Encoding ![svg](_images/svg-5606d4d4145063185c8891a1e40b1c7816c55611.svg) Description This instruction is supported for the RV32 and RV64 base architectures. It implements the _P0_ transform of the SM3 hash function \[[31](../biblio/bibliography.html#bib-gbt:sm3)\] \[[32](../biblio/bibliography.html#bib-iso:sm3)\]. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. | | Supporting Material This instruction is based on work done in \[[33](../biblio/bibliography.html#bib-mjs:lwsha:20)\]. | | ------------------------------------------------------------------------------------------------------------------------ | Operation ```sail function clause execute (SM3P0(rs1, rd)) = { let r1 : bits(32) = X(rs1)[31..0]; let result : bits(32) = r1 ^ rol32(r1, 9) ^ rol32(r1, 17); X(rd) = EXTS(result); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | ------------- | --------------- | --------------- | | [Zksh](#zksh) | v1.0.0 | Ratified | | [Zks](#zks) | v1.0.0 | Ratified | #### [](#insns-sm3p1)32.1.3.42\. sm3p1 Synopsis Implements the _P1_ transformation function as used in the SM3 hash function \[[31](../biblio/bibliography.html#bib-gbt:sm3)\] \[[32](../biblio/bibliography.html#bib-iso:sm3)\]. Mnemonic sm3p1 rd, rs1 Encoding ![svg](_images/svg-7c3f19f607d38c1053edb60da8ae41261828c73a.svg) Description This instruction is supported for the RV32 and RV64 base architectures. It implements the _P1_ transform of the SM3 hash function \[[31](../biblio/bibliography.html#bib-gbt:sm3)\] \[[32](../biblio/bibliography.html#bib-iso:sm3)\]. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. | | Supporting Material This instruction is based on work done in \[[33](../biblio/bibliography.html#bib-mjs:lwsha:20)\]. | | ------------------------------------------------------------------------------------------------------------------------ | Operation ```sail function clause execute (SM3P1(rs1, rd)) = { let r1 : bits(32) = X(rs1)[31..0]; let result : bits(32) = r1 ^ rol32(r1, 15) ^ rol32(r1, 23); X(rd) = EXTS(result); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | ------------- | --------------- | --------------- | | [Zksh](#zksh) | v1.0.0 | Ratified | | [Zks](#zks) | v1.0.0 | Ratified | #### [](#insns-sm4ed)32.1.3.43\. sm4ed Synopsis Accelerates the block encrypt/decrypt operation of the SM4 block cipher \[[34](../biblio/bibliography.html#bib-gbt:sm4)\] \[[35](../biblio/bibliography.html#bib-iso:sm4)\]. Mnemonic sm4ed rd, rs1, rs2, bs Encoding ![svg](_images/svg-57bcd06fb87ed455cb656d1d4bd3eff6c83fd89a.svg) Description Implements a T-tables in hardware style approach to accelerating the SM4 round function. A byte is extracted from `rs2` based on `bs`, to which the SBox and linear layer transforms are applied, before the result is XOR’d with`rs1` and written back to `rd`. This instruction exists on RV32 and RV64 base architectures. On RV64, the 32-bit result is sign extended to XLEN bits. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (SM4ED (bs,rs2,rs1,rd)) = { let shamt : bits(5) = bs @ 0b000; /* shamt = bs*8 */ let sb_in : bits(8) = (X(rs2)[31..0] >> shamt)[7..0]; let x : bits(32) = 0x000000 @ sm4_sbox(sb_in); let y : bits(32) = x ^ (x << 8) ^ ( x << 2) ^ (x << 18) ^ ((x & 0x0000003F) << 26) ^ ((x & 0x000000C0) << 10); let z : bits(32) = rol32(y, unsigned(shamt)); let result: bits(32) = z ^ X(rs1)[31..0]; X(rd) = EXTS(result); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | --------------- | --------------- | --------------- | | [Zksed](#zksed) | v1.0.0 | Ratified | | [Zks](#zks) | v1.0.0 | Ratified | #### [](#insns-sm4ks)32.1.3.44\. sm4ks Synopsis Accelerates the Key Schedule operation of the SM4 block cipher \[[34](../biblio/bibliography.html#bib-gbt:sm4)\] \[[35](../biblio/bibliography.html#bib-iso:sm4)\]. Mnemonic sm4ks rd, rs1, rs2, bs Encoding ![svg](_images/svg-a3affdc02eb9ab7521d22c3fdc3e7a32b45aaddb.svg) Description Implements a T-tables in hardware style approach to accelerating the SM4 Key Schedule. A byte is extracted from `rs2` based on `bs`, to which the SBox and linear layer transforms are applied, before the result is XOR’d with`rs1` and written back to `rd`. This instruction exists on RV32 and RV64 base architectures. On RV64, the 32-bit result is sign extended to XLEN bits. This instruction must _always_ be implemented such that its execution latency does not depend on the data being operated on. Operation ```sail function clause execute (SM4KS (bs,rs2,rs1,rd)) = { let shamt : bits(5) = (bs @ 0b000); /* shamt = bs*8 */ let sb_in : bits(8) = (X(rs2)[31..0] >> shamt)[7..0]; let x : bits(32) = 0x000000 @ sm4_sbox(sb_in); let y : bits(32) = x ^ ((x & 0x00000007) << 29) ^ ((x & 0x000000FE) << 7) ^ ((x & 0x00000001) << 23) ^ ((x & 0x000000F8) << 13) ; let z : bits(32) = rol32(y, unsigned(shamt)); let result: bits(32) = z ^ X(rs1)[31..0]; X(rd) = EXTS(result); RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | --------------- | --------------- | --------------- | | [Zksed](#zksed) | v1.0.0 | Ratified | | [Zks](#zks) | v1.0.0 | Ratified | #### [](#insns-unzip-sc)32.1.3.45\. unzip Synopsis Place odd and even bits of the source register into upper and lower halves of the destination register, respectively. Mnemonic unzip _rd_, _rs_ Encoding ![svg](_images/svg-97edb370bfc8c5ffa9520d49467c4ff1b7f4368d.svg) Description This instruction scatters all of the odd and even bits of a source word into the high and low halves of a destination word. It is the inverse of the [zip](#insns-zip-sc) instruction. This instruction is available only on RV32. Operation ```sail foreach (i from 0 to xlen/2-1) { X(rd)[i] = X(rs1)[2*i] X(rd)[i+xlen/2] = X(rs1)[2*i+1] } ``` | | Software Hint This instruction is useful for implementing the SHA3 cryptographic hash function on a 32-bit architecture, as it implements the bit-interleaving operation used to speed up the 64-bit rotations directly. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbkb ([Zbkb](#zbkb-sc)) (RV32) | v1.0.0-rc4 | Ratified | #### [](#insns-xnor-sc)32.1.3.46\. xnor Synopsis Exclusive NOR Mnemonic xnor _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-69198fb0cd20a10e882917a7e04916117cbc5945.svg) Description This instruction performs the bit-wise exclusive-NOR operation on _rs1_ and _rs2_. Operation ```sail X(rd) = ~(X(rs1) ^ X(rs2)); ``` Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbb ([Zbb](b-st-ext.html#zbb)) | v1.0.0 | Ratified | | Zbkb ([Zbkb](#zbkb-sc)) | v1.0.0-rc4 | Ratified | #### [](#insns-xperm8-sc)32.1.3.47\. xperm8 Synopsis Byte-wise lookup of indices into a vector in registers. Mnemonic xperm8 _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-23bb204c0f89532d4d93dbeb2e87fb7cef67f698.svg) Description The xperm8 instruction operates on bytes. The _rs1_ register contains a vector of XLEN/8 8-bit elements. The _rs2_ register contains a vector of XLEN/8 8-bit indexes. The result is each element in _rs2_ replaced by the indexed element in _rs1_, or zero if the index into _rs2_ is out of bounds. Operation ```sail val xperm8_lookup : (bits(8), xlenbits) -> bits(8) function xperm8_lookup (idx, lut) = { (lut >> (idx @ 0b000))[7..0] } function clause execute ( XPERM8 (rs2,rs1,rd)) = { result : xlenbits = EXTZ(0b0); foreach(i from 0 to xlen by 8) { result[i+7..i] = xperm8_lookup(X(rs2)[i+7..i], X(rs1)); }; X(rd) = result; RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | --------------------------------- | --------------- | --------------- | | Zbkx ([Zbkx](b-st-ext.html#zbkx)) | v1.0 | Ratified | #### [](#insns-xperm4-sc)32.1.3.48\. xperm4 Synopsis Nibble-wise lookup of indices into a vector. Mnemonic xperm4 _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-e5856e4e32691a6664d923b27a4a8f85e24cda05.svg) Description The xperm4 instruction operates on nibbles. The _rs1_ register contains a vector of XLEN/4 4-bit elements. The _rs2_ register contains a vector of XLEN/4 4-bit indexes. The result is each element in _rs2_ replaced by the indexed element in _rs1_, or zero if the index into _rs2_ is out of bounds. Operation ```sail val xperm4_lookup : (bits(4), xlenbits) -> bits(4) function xperm4_lookup (idx, lut) = { (lut >> (idx @ 0b00))[3..0] } function clause execute ( XPERM4 (rs2,rs1,rd)) = { result : xlenbits = EXTZ(0b0); foreach(i from 0 to xlen by 4) { result[i+3..i] = xperm4_lookup(X(rs2)[i+3..i], X(rs1)); }; X(rd) = result; RETIRE_SUCCESS } ``` Included in | Extension | Minimum version | Lifecycle state | | --------------------------------- | --------------- | --------------- | | Zbkx ([Zbkx](b-st-ext.html#zbkx)) | v1.0 | Ratified | #### [](#insns-zip-sc)32.1.3.49\. zip Synopsis Interleave upper and lower halves of the source register into odd and even bits of the destination register, respectively. Mnemonic zip _rd_, _rs_ Encoding ![svg](_images/svg-7551362fcbac13f4e2c1ab5c2a42fe5bdcb6f1a5.svg) Description This instruction gathers bits from the high and low halves of the source word into odd/even bit positions in the destination word. It is the inverse of the [unzip](#insns-unzip-sc) instruction. This instruction is available only on RV32. Operation ```sail foreach (i from 0 to xlen/2-1) { X(rd)[2*i] = X(rs1)[i] X(rd)[2*i+1] = X(rs1)[i+xlen/2] } ``` | | Software Hint This instruction is useful for implementing the SHA3 cryptographic hash function on a 32-bit architecture, as it implements the bit-interleaving operation used to speed up the 64-bit rotations directly. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Included in | Extension | Minimum version | Lifecycle state | | ------------------------------ | --------------- | --------------- | | Zbkb ([Zbkb](#zbkb-sc)) (RV32) | v1.0.0-rc4 | Ratified | ### [](#crypto%5Fscalar%5Fes)32.1.4\. Entropy Source The `seed` CSR provides an interface to a NIST SP 800-90B \[[36](../biblio/bibliography.html#bib-tubake:18)\] or BSI AIS-31 \[[37](../biblio/bibliography.html#bib-kisc11)\] compliant physical Entropy Source (ES). An entropy source, by itself, is not a cryptographically secure Random Bit Generator (RBG), but can be used to build standard (and nonstandard) RBGs of many types with the help of symmetric cryptography. Expected usage is to condition (typically with SHA-2/3) the output from an entropy source and use it to seed a cryptographically secure Deterministic Random Bit Generator (DRBG) such as AES-based `CTR_DRBG` \[[38](../biblio/bibliography.html#bib-bake15)\]. The combination of an Entropy Source, Conditioning, and a DRBG can be used to create random bits securely \[[39](../biblio/bibliography.html#bib-bakero:21)\]. See [32.1.7\. Entropy Source Rationale and Recommendations](#crypto%5Fscalar%5Fappx%5Fes) for a non-normative description of a certification and self-certification procedures, design rationale, and more detailed suggestions on how the entropy source output can be used. #### [](#crypto%5Fscalar%5Fseed%5Fcsr)32.1.4.1\. The `seed` CSR `seed` is an unprivileged CSR located at address `0x015`. The 32-bit contents of `seed` are as follows: | Bits | Name | Description | | ----- | ---------- | -------------------------------------------------- | | 31:30 | OPST | Status: BIST (00), WAIT (01), ES16 (10), DEAD(11). | | 29:24 | _reserved_ | For future use by the RISC-V specification. | | 23:16 | _custom_ | Designated for custom and experimental use. | | 15: 0 | entropy | 16 bits of randomness, only when OPST=ES16. | Attempts to access the `seed` CSR using a read-only CSR-access instruction (`CSRRS`/`CSRRC` with _rs1_\=`x0` or `CSRRSI`/`CSRRCI` with _uimm_\=0) raise an illegal-instruction exception; any other CSR-access instruction may be used to access `seed`. The write value (in `rs1` or `uimm`) must be ignored by implementations. The purpose of the write is to signal polling and flushing. Software normally uses the instruction `csrrw rd, seed, x0` to read the `seed`CSR. Encoding ![svg](_images/svg-57188ebad0a4176f7cbe8d11290bf93710da536a.svg) The `seed` CSR is also access controlled by execution mode, and attempted read or write access will raise an illegal-instruction exception outside M mode unless access is explicitly granted. See [32.1.4.3\. Access Control to seed](#crypto%5Fscalar%5Fes%5Faccess) for more details. The status bits `seed[31:30]` \= `OPST` may be `ES16` (10), indicating successful polling, or one of three entropy polling failure statuses `BIST` (00), `WAIT` (01), or `DEAD` (11), discussed below. Each returned `seed[15:0]` \= `entropy` value represents unique randomness when `OPST`\=`ES16` (`seed[31:30]` \= `10`), even if its numerical value is the same as that of a previously polled `entropy` value. The implementation requirements of `entropy` bits are defined in [32.1.4.2\. Entropy Source Requirements](#crypto%5Fscalar%5Fes%5Freq). When `OPST` is not `ES16`, `entropy` must be set to 0\. An implementation may safely set reserved and custom bits to zeros. For security reasons, the interface guarantees that secret `entropy`words are not made available multiple times. Hence polling (reading) must also have the side effect of clearing (wipe-on-read) the `entropy` contents and changing the state to `WAIT` (unless there is `entropy`immediately available for `ES16`). Other states (`BIST`, `WAIT`, and `DEAD`) may be unaffected by polling. The Status Bits returned in `seed[31:30]`\=`OPST`: * `00` \- `BIST`indicates that Built-In Self-Test "on-demand" (BIST) testing is being performed. If `OPST` returns temporarily to `BIST` from any other state, this signals a non-fatal self-test alarm, which is non-actionable, apart from being logged. Such a `BIST` alarm must be latched until polled at least once to enable software to record its occurrence. * `01` \- `WAIT`means that a sufficient amount of entropy is not yet available. This is not an error condition and may (in fact) be more frequent than ES16 since physical entropy sources often have low bandwidth. * `10` \- `ES16`indicates success; the low bits `seed[15:0]` will have 16 bits of randomness (`entropy`), which is guaranteed to meet certain minimum entropy requirements, regardless of implementation. * `11` \- `DEAD`is an unrecoverable self-test error. This may indicate a hardware fault, a security issue, or (extremely rarely) a type-1 statistical false positive in the continuous testing procedures. In case of a fatal failure, an immediate lockdown may also be an appropriate response in dedicated security devices. **Example.** `0x8000ABCD` is a valid `ES16` status output, with `0xABCD`being the `entropy` value. `0xFFFFFFFF` is an invalid output (`DEAD`) with no `entropy` value. ![es state](_images/es_state.svg) Figure 1\. Entropy Source state transition diagram. Normally the operational state alternates between WAIT (no data) and ES16, which means that 16 bits of randomness (`entropy`) have been polled. BIST (Built-in Self-Test) only occurs after reset or to signal a non-fatal self-test alarm (if reached after WAIT or ES16). DEAD is an unrecoverable error state. #### [](#crypto%5Fscalar%5Fes%5Freq)32.1.4.2\. Entropy Source Requirements The output `entropy` (`seed[15:0]` in ES16 state) is not necessarily fully conditioned randomness due to hardware and energy limitations of smaller, low-powered implementations. However, minimum requirements are defined. The main requirement is that 2-to-1 cryptographic post-processing in 256-bit input blocks will yield 128-bit "full entropy" output blocks. Entropy source users may make this conservative assumption but are not prohibited from using more than twice the number of seed bits relative to the desired resulting entropy. An implementation of the entropy source should meet at least one of the following requirements sets in order to be considered a secure and safe design: * [32.1.4.2.1\. NIST SP 800-90B / FIPS 140-3 Requirements](#crypto%5Fscalar%5Fes%5Freq%5F90b): A physical entropy source meeting NIST SP 800-90B \[[36](../biblio/bibliography.html#bib-tubake:18)\] criteria with evaluated min-entropy of 192 bits for each 256 output bits (min-entropy rate 0.75). * [32.1.4.2.2\. BSI AIS-31 PTG.2 / Common Criteria Requirements](#crypto%5Fscalar%5Fes%5Freq%5Fptg2): A physical entropy source meeting the AIS-31 PTG.2 \[[37](../biblio/bibliography.html#bib-kisc11)\] criteria, implying average Shannon entropy rate 0.997\. The source must also meet the NIST 800-90B min-entropy rate 192/256 = 0.75. * [32.1.4.2.3\. Virtual Sources: Security Requirement](#crypto%5Fscalar%5Fes%5Freq%5Fvirt): A virtual entropy source is a DRBG seeded from a physical entropy source. It must have at least a 256-bit (Post-Quantum Category 5) internal security level. All implementations must signal initialization, test mode, and health alarms as required by respective standards. This may require the implementer to add non-standard (custom) test interfaces in a secure and safe manner, an example of which is described in [32.1.7.6\. Suggested GetNoise Test Interface](#crypto%5Fscalar%5Fes%5Fgetnoise) ##### [](#crypto%5Fscalar%5Fes%5Freq%5F90b)32.1.4.2.1\. NIST SP 800-90B / FIPS 140-3 Requirements All NIST SP 800-90B \[[36](../biblio/bibliography.html#bib-tubake:18)\] required components and health test mechanisms must be implemented. The entropy requirement is satisfied if 128 bits of _full entropy_ can be obtained from each 256-bit (16\*16 -bit) successful, but possibly non-consecutive `entropy` (ES16) output sequence using a vetted conditioning algorithm such as a cryptographic hash (See Section 3.1.5.1.1, SP 800-90B \[[36](../biblio/bibliography.html#bib-tubake:18)\]). In practice, a min-entropy rate of 0.75 or larger is required for this. Note that 128 bits of estimated input min-entropy does not yield 128 bits of conditioned, full entropy in SP 800-90B/C evaluation. Instead, the implication is that every 256-bit sequence should have min-entropy of at least 128+64 = 192 bits, as discussed in SP 800-90C \[[39](../biblio/bibliography.html#bib-bakero:21)\]; the likelihood of successfully "guessing" an individual 256-bit output sequence should not be higher than 2\-192 even with (almost) unconstrained amount of entropy source data and computational power. Rather than attempting to define all the mathematical and architectural properties that the entropy source must satisfy, we define that the physical entropy source be strong and robust enough to pass the equivalent of NIST SP 800-90 evaluation and certification for full entropy when conditioned cryptographically in ratio 2:1 with 128-bit output blocks. Even though the requirement is defined in terms of 128-bit full entropy blocks, we recommend 256-bit security. This can be accomplished by using at least 512 `entropy` bits to initialize a DRBG that has 256-bit security. ##### [](#crypto%5Fscalar%5Fes%5Freq%5Fptg2)32.1.4.2.2\. BSI AIS-31 PTG.2 / Common Criteria Requirements For alternative Common Criteria certification (or self-certification), AIS 31 PTG.2 class \[[37](../biblio/bibliography.html#bib-kisc11)\] (Sect. 4.3.) required hardware components and mechanisms must be implemented. In addition to AIS-31 PTG.2 randomness requirements (Shannon entropy rate of 0.997 as evaluated in that standard), the overall min-entropy requirement of remains, as discussed in [32.1.4.2.1\. NIST SP 800-90B / FIPS 140-3 Requirements](#crypto%5Fscalar%5Fes%5Freq%5F90b). Note that 800-90B min-entropy can be significantly lower than AIS-31 Shannon entropy. These two metrics should not be equated or confused with each other. ##### [](#crypto%5Fscalar%5Fes%5Freq%5Fvirt)32.1.4.2.3\. Virtual Sources: Security Requirement | | A virtual source is not an ISA compliance requirement. It is defined for the benefit of the RISC-V security ecosystem so that virtual systems may have a consistent level of security. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | A virtual source is not a physical entropy source but provides additional protection against covert channels, depletion attacks, and host identification in operating environments that can not be entirely trusted with direct access to a hardware resource. Despite limited trust, implementers should try to guarantee that even such environments have sufficient entropy available for secure cryptographic operations. A virtual source traps access to the `seed` CSR, emulates it, or otherwise implements it, possibly without direct access to a physical entropy source. The output can be cryptographically secure pseudorandomness instead of real entropy, but must have at least 256-bit security, as defined below. A virtual source is intended especially for guest operating systems, sandboxes, emulators, and similar use cases. As a technical definition, a random-distinguishing attack against the output should require computational resources comparable or greater than those required for exhaustive key search on a secure block cipher with a 256-bit key (e.g., AES 256). This applies to both classical and quantum computing models, but only classical information flows. The virtual source security requirement maps to Post-Quantum Security Category 5 \[[40](../biblio/bibliography.html#bib-ni16)\]. Any implementation of the `seed` CSR that limits the security strength shall not reduce it to less than 256 bits. If the security level is under 256 bits, then the interface must not be available. A virtual entropy source does not need to implement `WAIT` or `BIST` states. It should fail (`DEAD`) if the host DRBG or entropy source fails and there is insufficient seeding material for the host DRBG. #### [](#crypto%5Fscalar%5Fes%5Faccess)32.1.4.3\. Access Control to `seed` The Zkr extension adds the `SSEED` and `USEED` fields to the `mseccfg` CSR to control access to the `seed` CSR from U, S, or HS modes (see Privileged ISA specification). Systems should implement carefully considered access control policies from lower privilege modes to physical entropy sources. The system can trap attempted access to `seed` and feed a less privileged client_virtual entropy source_ data ([32.1.4.2.3\. Virtual Sources: Security Requirement](#crypto%5Fscalar%5Fes%5Freq%5Fvirt)) instead of invoking an SP 800-90B ([32.1.4.2.1\. NIST SP 800-90B / FIPS 140-3 Requirements](#crypto%5Fscalar%5Fes%5Freq%5F90b)) or PTG.2 ([32.1.4.2.2\. BSI AIS-31 PTG.2 / Common Criteria Requirements](#crypto%5Fscalar%5Fes%5Freq%5Fptg2)) _physical entropy source_. Emulated `seed`data generation is made with an appropriately seeded, secure software DRBG. See [32.1.7.3.5\. Security Considerations for Direct Hardware Access](#crypto%5Fscalar%5Fappx%5Fes%5Faccess) for security considerations related to direct access to entropy sources. Implementations may implement `mseccfg` such that `[s,u]seed` is a read-only constant value `0`. Software may discover if access to the `seed` CSR can be enabled in U and S mode by writing a `1` to `[s,u]seed` and reading back the result. ### [](#crypto%5Fscalar%5Fzkt)32.1.5\. Data Independent Execution Latency Subset: Zkt The Zkt extension attests that the machine has data-independent execution time for a safe subset of instructions. This property is commonly called_"constant-time"_ although should not be taken with that literal meaning. All currently defined cryptographic instructions (Zk\* and Zbk\* extensions) are on this list, together with a set of relevant supporting instructions from I, M, C, and B extensions. | | Note to software developers Failure to prevent leakage of sensitive parameters via the direct timing channel is considered a serious security vulnerability and will typically result in a CERT CVE security advisory. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#32-1-5-1-scope-and-goal)32.1.5.1\. Scope and Goal An "ISA contract" is made between a programmer and the RISC-V implementation that Zkt instructions do not leak information about processed secret data (plaintext, keying information, or other "sensitive security parameters" — FIPS 140-3 term) through differences in execution latency. Zkt does _not_define a set of instructions available in the core; it just restricts the behaviour of certain instructions if those are implemented. Currently, the scope of this document is within scalar RV32/RV64 processors. Vector cryptography instructions (and appropriate vector support instructions) will be added later, as will other security-related functions that wish to assert leakage-free execution latency properties. Loads, stores, conditional branches are excluded, along with a set of instructions that are rarely necessary to process secret data. Also excluded are instructions for which workarounds exist in standard cryptographic middleware due to the limitations of other ISA processors. The stated goal is that OpenSSL, BoringSSL (Android), the Linux Kernel, and similar trusted software will not have directly observable timing side channels when compiled and running on a Zkt-enabled RISC-V target. The Zkt extension explicitly states many of the common latency assumptions made by cryptography developers. Vendors do not have to implement all of the list’s instructions to be Zkt compliant; however, if they claim to have Zkt and implement any of the listed instructions, it must have data-independent latency. For example, many simple RV32I and RV64I cores (without Multiply, Compressed, Bitmanip, or Cryptographic extensions) are technically compliant with Zkt. A constant-time AES can be implemented on them using "bit-slice" techniques, but it will be excruciatingly slow when compared to implementation with AES instructions. There are no guarantees that even a bit-sliced cipher implementation (largely based on boolean logic instructions) is secure on a core without Zkt attestation. Out-of-order implementations adhering to Zkt are still free to fuse, crack, change or even ignore sequences of instructions, so long as the optimisations are applied deterministically, and not based on operand data. The guiding principle should be that no information about the data being operated on should be leaked based on the execution latency. | | It is left to future extensions or other techniques to tackle the problem of data-independent execution in implementations which advanced out-of-order capabilities which use value prediction, or which are otherwise data-dependent. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Note to software developers Programming techniques can only mitigate leakage directly caused by arithmetic, caches, and branches. Other ISAs have had micro-architectural issues such as Spectre, Meltdown, Speculative Store Bypass, Rogue System Register Read, Lazy FP State Restore, Bounds Check Bypass Store, TLBleed, and L1TF/Foreshadow, etc. See e.g.[NSA Hardware and Firmware Security Guidance](https://github.com/nsacyber/Hardware-and-Firmware-Security-Guidance) It is not within the remit of this proposal to mitigate these_micro-architectural_ leakages. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#32-1-5-2-background)32.1.5.2\. Background * Timing attacks are much more powerful than was realised before the 2010s, which has led to a significant mitigation effort in current cryptographic code-bases. * Cryptography developers use static and dynamic security testing tools to trace the handling of secret information and detect occasions where it influences a branch or is used for a table lookup. * Architectural testing for Zkt can be pragmatic and semi-formal;_security by design_ against basic timing attacks can usually be achieved via conscious implementation (of relevant iterative multi-cycle instructions or instructions composed of micro-ops) in way that avoids data-dependent latency. * Laboratory testing may utilize statistical timing attack leakage analysis techniques such as those described in ISO/IEC 17825 \[[41](../biblio/bibliography.html#bib-is16)\]. * Binary executables should not contain secrets in the instruction encodings (Kerckhoffs’s principle), so instruction timing may leak information about immediates, ordering of input registers, etc. There may be an exception to this in systems where a binary loader modifies the executable for purposes of relocation — and it is desirable to keep the execution location (PC) secret. This is why instructions such as LUI, AUIPC, and ADDI are on the list. * The rules used by audit tools are relatively simple to understand. Very briefly; we call the plaintext, secret keys, expanded keys, nonces, and other such variables "secrets". A secret variable (arithmetically) modifying any other variable/register turns that into a secret too. If a secret ends up in address calculation affecting a load or store, that is a violation. If a secret affects a branch’s condition, that is also a violation. A secret variable location or register becomes a non-secret via specific zeroization/sanitisation or by being declared ciphertext (or otherwise no-longer-secret information). In essence, secrets can only "touch" instructions on the Zkt list while they are secrets. #### [](#32-1-5-3-specific-instruction-rationale)32.1.5.3\. Specific Instruction Rationale * HINT instruction forms (typically encodings with _rd_\=`x0`) are excluded from the data-independent time requirement. * Floating point (F, D, Q, L extensions) are currently excluded from the constant-time requirement as they have very few applications in standardised cryptography. We may consider adding floating point add, sub, multiply as a constant time requirement for some floating point extension in case a specific algorithm (such as the PQC Signature algorithm Falcon) becomes critical. * Cryptographers typically assume division to be variable-time (while multiplication is constant time) and implement their Montgomery reduction routines with that assumption. * Zicsr, Zifencei are excluded. * Some instructions are on the list simply because we see no harm in including them in testing scope. #### [](#32-1-5-4-programming-information)32.1.5.4\. Programming Information For background information on secure programming "models", see: * Thomas Pornin: _"Why Constant-Time Crypto?"_ (A great introduction to timing assumptions.) * Jean-Philippe Aumasson: _"Guidelines for low-level cryptography software."_(A list of recommendations.) * Peter Schwabe: _"Timing Attacks and Countermeasures."_(Lecture slides — nice references.) * Adam Langley: _"ctgrind."_ (This is from 2010 but is still relevant.) * Kris Kwiatkowski: _"Constant-time code verification with Memory Sanitizer."_ * For early examples of timing attack vulnerabilities, see and related academic papers. #### [](#32-1-5-5-zkt-listings)32.1.5.5\. Zkt listings The following instructions are included in the `Zkt` subset They are listed here grouped by their original parent extension. | | Note to implementers You do not need to implement all of these instructions to implement Zkt. Rather, every one of these instructions that the core does implement must adhere to the requirements of Zkt. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#32-1-5-5-1-rvi-base-instruction-set)32.1.5.5.1\. RVI (Base Instruction Set) Only basic arithmetic and `slt*` (for carry computations) are included. The data-independent timing requirement does not apply to HINT instruction encoding forms of these instructions. | RV32 | RV64 | Mnemonic | | ---- | ------------------------ | ------------------------ | | ✓ | ✓ | lui _rd_, _imm_ | | ✓ | ✓ | auipc _rd_, _imm_ | | ✓ | ✓ | addi _rd_, _rs1_, _imm_ | | ✓ | ✓ | slti _rd_, _rs1_, _imm_ | | ✓ | ✓ | sltiu _rd_, _rs1_, _imm_ | | ✓ | ✓ | xori _rd_, _rs1_, _imm_ | | ✓ | ✓ | ori _rd_, _rs1_, _imm_ | | ✓ | ✓ | andi _rd_, _rs1_, _imm_ | | ✓ | ✓ | slli _rd_, _rs1_, _imm_ | | ✓ | ✓ | srli _rd_, _rs1_, _imm_ | | ✓ | ✓ | srai _rd_, _rs1_, _imm_ | | ✓ | ✓ | add _rd_, _rs1_, _rs2_ | | ✓ | ✓ | sub _rd_, _rs1_, _rs2_ | | ✓ | ✓ | sll _rd_, _rs1_, _rs2_ | | ✓ | ✓ | slt _rd_, _rs1_, _rs2_ | | ✓ | ✓ | sltu _rd_, _rs1_, _rs2_ | | ✓ | ✓ | xor _rd_, _rs1_, _rs2_ | | ✓ | ✓ | srl _rd_, _rs1_, _rs2_ | | ✓ | ✓ | sra _rd_, _rs1_, _rs2_ | | ✓ | ✓ | or _rd_, _rs1_, _rs2_ | | ✓ | ✓ | and _rd_, _rs1_, _rs2_ | | ✓ | addiw _rd_, _rs1_, _imm_ | | | ✓ | slliw _rd_, _rs1_, _imm_ | | | ✓ | srliw _rd_, _rs1_, _imm_ | | | ✓ | sraiw _rd_, _rs1_, _imm_ | | | ✓ | addw _rd_, _rs1_, _rs2_ | | | ✓ | subw _rd_, _rs1_, _rs2_ | | | ✓ | sllw _rd_, _rs1_, _rs2_ | | | ✓ | srlw _rd_, _rs1_, _rs2_ | | | ✓ | sraw _rd_, _rs1_, _rs2_ | | ##### [](#32-1-5-5-2-rvm-multiply)32.1.5.5.2\. RVM (Multiply) Multiplication is included; division and remaindering excluded. | RV32 | RV64 | Mnemonic | | ---- | ----------------------- | ------------------------- | | ✓ | ✓ | mul _rd_, _rs1_, _rs2_ | | ✓ | ✓ | mulh _rd_, _rs1_, _rs2_ | | ✓ | ✓ | mulhsu _rd_, _rs1_, _rs2_ | | ✓ | ✓ | mulhu _rd_, _rs1_, _rs2_ | | ✓ | mulw _rd_, _rs1_, _rs2_ | | ##### [](#32-1-5-5-3-rvc-compressed)32.1.5.5.3\. RVC (Compressed) Same criteria as in RVI. Organised by quadrants. | RV32 | RV64 | Mnemonic | | ---- | ------- | -------- | | ✓ | ✓ | c.nop | | ✓ | ✓ | c.addi | | ✓ | c.addiw | | | ✓ | ✓ | c.lui | | ✓ | ✓ | c.srli | | ✓ | ✓ | c.srai | | ✓ | ✓ | c.andi | | ✓ | ✓ | c.sub | | ✓ | ✓ | c.xor | | ✓ | ✓ | c.or | | ✓ | ✓ | c.and | | ✓ | c.subw | | | ✓ | c.addw | | | ✓ | ✓ | c.slli | | ✓ | ✓ | c.mv | | ✓ | ✓ | c.add | ##### [](#32-1-5-5-4-zcb-extension)32.1.5.5.4\. Zcb Extension These instructions are compressed versions of I and M instructions that are included in Zkt. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | -------- | -------------------------------------- | | ✓ | ✓ | c.mul | [c.mul](zc.html#insns-c%5Fmul) | | ✓ | ✓ | c.not | [c.not](zc.html#insns-c%5Fnot) | | ✓ | ✓ | c.zext.b | [c.zext.b](zc.html#insns-c%5Fzext%5Fb) | ##### [](#32-1-5-5-5-rvk-scalar-cryptography)32.1.5.5.5\. RVK (Scalar Cryptography) All K-specific instructions are included. Additionally, `seed` CSR latency should be independent of `ES16` state output`entropy` bits, as that is a sensitive security parameter. See [32.1.7.3.5\. Security Considerations for Direct Hardware Access](#crypto%5Fscalar%5Fappx%5Fes%5Faccess). | RV32 | RV64 | Mnemonic | Instruction | | ---- | ----------- | -------------------------------------------------------------- | ------------------------------------------------ | | ✓ | aes32dsi | [AES final round decrypt (RV32)](#insns-aes32dsi) | | | ✓ | aes32dsmi | [AES middle round decrypt (RV32)](#insns-aes32dsmi) | | | ✓ | aes32esi | [AES final round encrypt (RV32)](#insns-aes32esi) | | | ✓ | aes32esmi | [AES middle round encrypt (RV32)](#insns-aes32esmi) | | | ✓ | aes64ds | [AES decrypt final round (RV64)](#insns-aes64ds) | | | ✓ | aes64dsm | [AES decrypt middle round (RV64)](#insns-aes64dsm) | | | ✓ | aes64es | [AES encrypt final round instruction (RV64)](#insns-aes64es) | | | ✓ | aes64esm | [AES encrypt middle round instruction (RV64)](#insns-aes64esm) | | | ✓ | aes64im | [AES Decrypt KeySchedule MixColumns (RV64)](#insns-aes64im) | | | ✓ | aes64ks1i | [AES Key Schedule Instruction 1 (RV64)](#insns-aes64ks1i) | | | ✓ | aes64ks2 | [AES Key Schedule Instruction 2 (RV64)](#insns-aes64ks2) | | | ✓ | ✓ | sha256sig0 | [SHA2-256 Sigma0 instruction](#insns-sha256sig0) | | ✓ | ✓ | sha256sig1 | [SHA2-256 Sigma1 instruction](#insns-sha256sig1) | | ✓ | ✓ | sha256sum0 | [SHA2-256 Sum0 instruction](#insns-sha256sum0) | | ✓ | ✓ | sha256sum1 | [SHA2-256 Sum1 instruction](#insns-sha256sum1) | | ✓ | sha512sig0h | [SHA2-512 Sigma0 high (RV32)](#insns-sha512sig0h) | | | ✓ | sha512sig0l | [SHA2-512 Sigma0 low (RV32)](#insns-sha512sig0l) | | | ✓ | sha512sig1h | [SHA2-512 Sigma1 high (RV32)](#insns-sha512sig1h) | | | ✓ | sha512sig1l | [SHA2-512 Sigma1 low (RV32)](#insns-sha512sig1l) | | | ✓ | sha512sum0r | [SHA2-512 Sum0 (RV32)](#insns-sha512sum0r) | | | ✓ | sha512sum1r | [SHA2-512 Sum1 (RV32)](#insns-sha512sum1r) | | | ✓ | sha512sig0 | [SHA2-512 Sigma0 instruction (RV64)](#insns-sha512sig0) | | | ✓ | sha512sig1 | [SHA2-512 Sigma1 instruction (RV64)](#insns-sha512sig1) | | | ✓ | sha512sum0 | [SHA2-512 Sum0 instruction (RV64)](#insns-sha512sum0) | | | ✓ | sha512sum1 | [SHA2-512 Sum1 instruction (RV64)](#insns-sha512sum1) | | | ✓ | ✓ | sm3p0 | [SM3 P0 transform](#insns-sm3p0) | | ✓ | ✓ | sm3p1 | [SM3 P1 transform](#insns-sm3p1) | | ✓ | ✓ | sm4ed | [SM4 Encrypt/Decrypt Instruction](#insns-sm4ed) | | ✓ | ✓ | sm4ks | [SM4 Key Schedule Instruction](#insns-sm4ks) | ##### [](#32-1-5-5-6-rvb-bitmanip)32.1.5.5.6\. RVB (Bitmanip) The [Zbkb](#zbkb-sc), [Zbkc](#zbkc-sc) and [Zbkx](#zbkx-sc) extensions are included in their entirety. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ----- | ------------------------------------------------------- | --------------------------------------------------- | | ✓ | ✓ | clmul | [Carry-less multiply (low-part)](#insns-clmul-sc) | | ✓ | ✓ | clmulh | [Carry-less multiply (high-part)](#insns-clmulh-sc) | | ✓ | ✓ | xperm4 | [Crossbar permutation (nibbles)](#insns-xperm4-sc) | | ✓ | ✓ | xperm8 | [Crossbar permutation (bytes)](#insns-xperm8-sc) | | ✓ | ✓ | ror | [Rotate right (Register)](#insns-ror-sc) | | ✓ | ✓ | rol | [Rotate left (Register)](#insns-rol-sc) | | ✓ | ✓ | rori | [Rotate right (Immediate)](#insns-rori-sc) | | ✓ | rorw | [Rotate right Word (Register)](#insns-rorw-sc) | | | ✓ | rolw | [Rotate Left Word (Register)](#insns-rolw-sc) | | | ✓ | roriw | [Rotate right Word (Immediate)](#insns-roriw-sc) | | | ✓ | ✓ | andn | [AND with inverted operand](#insns-andn-sc) | | ✓ | ✓ | orn | [OR with inverted operand](#insns-orn-sc) | | ✓ | ✓ | xnor | [Exclusive NOR](#insns-xnor-sc) | | ✓ | ✓ | pack | [Pack low halves of registers](#insns-pack-sc) | | ✓ | ✓ | packh | [Pack low bytes of registers](#insns-packh-sc) | | ✓ | packw | [Pack low 16-bits of registers (RV64)](#insns-packw-sc) | | | ✓ | ✓ | brev8 | [Reverse bits in bytes](#insns-brev8-sc) | | ✓ | ✓ | rev8 | [Byte-reverse register](#insns-rev8-sc) | | ✓ | zip | [Bit interleave](#insns-zip-sc) | | | ✓ | unzip | [Bit deinterleave](#insns-unzip-sc) | | ### [](#crypto%5Fscalar%5Fappx%5Frationale)32.1.6\. Instruction Rationale This section contains various rationale, design notes and usage recommendations for the instructions in the scalar cryptography extension. It also tries to record how the designs of instructions were derived, or where they were contributed from. #### [](#32-1-6-1-aes-instructions)32.1.6.1\. AES Instructions The 32-bit instructions were derived from work in \[[42](../biblio/bibliography.html#bib-mjs:lwaes:20)\] and contributed to the RISC-V cryptography extension. The 64-bit instructions were developed collaboratively by task group members on our mailing list. Supporting material, including rationale and a design space exploration for all of the AES instructions in the specification can be found in the paper_"[The design of scalar AES Instruction Set Extensions for RISC-V](https://doi.org/10.46586/tches.v2021.i1.109-136)"_ \[[43](../biblio/bibliography.html#bib-mnpsw:20)\]. #### [](#32-1-6-2-sha2-instructions)32.1.6.2\. SHA2 Instructions These instructions were developed based on academic work at the University of Bristol as part of the XCrypto project \[[44](../biblio/bibliography.html#bib-mpp:19)\], and contributed to the RISC-V cryptography extension. The RV32 SHA2-512 instructions were based on this work, and developed in \[[33](../biblio/bibliography.html#bib-mjs:lwsha:20)\], before being contributed in the same way. #### [](#32-1-6-3-sm3-and-sm4-instructions)32.1.6.3\. SM3 and SM4 Instructions The SM4 instructions were derived from work in \[[42](../biblio/bibliography.html#bib-mjs:lwaes:20)\], and are hence very similar to the RV32 AES instructions. The SM3 instructions were inspired by the SHA2 instructions, and based on development work done in \[[33](../biblio/bibliography.html#bib-mjs:lwsha:20)\], before being contributed to the RISC-V cryptography extension. #### [](#crypto%5Fscalar%5Fzkb)32.1.6.4\. Bitmanip Instructions for Cryptography Many of the primitive operations used in symmetric key cryptography and cryptographic hash functions are well supported by the RISC-V Bitmanip extensions (see [Bitmanip](b-st-ext.html#bits)). | | This section repeats much of the information in[Zbkb](#zbkb-sc),[Zbkc](#zbkc-sc)and[Zbkx](#zbkx-sc), but includes more rationale. | | ------------------------------------------------------------------------------------------------------------------------------------ | We proposed that the scalar cryptographic extension _reuse_ a subset of the instructions from the Bitmanip extensions `Zb[abc]` directly. Specifically, this would mean that a core implementing_either_the scalar cryptographic extensions,_or_the `Zb[abc]`,_or_both, would be required to implement these instructions. ##### [](#32-1-6-4-1-rotations)32.1.6.4.1\. Rotations RV32, RV64: RV64 only: ror rd, rs1, rs2 rorw rd, rs1, rs2 rol rd, rs1, rs2 rolw rd, rs1, rs2 rori rd, rs1, imm roriw rd, rs1, imm See [zbkb](b-st-ext.html#zbkb) for details of these instructions. | | Notes to software developers Standard bitwise rotation is a primitive operation in many block ciphers and hash functions; it features particularly in the ARX (Add, Rotate, Xor) class of block ciphers and stream ciphers. Algorithms making use of 32-bit rotations: SHA256, AES (Shift Rows), ChaCha20, SM3. Algorithms making use of 64-bit rotations: SHA512, SHA3. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#32-1-6-4-2-bit-byte-permutations)32.1.6.4.2\. Bit & Byte Permutations RV32, RV64: brev8 rd, rs1 rev8 rd, rs1 See [zbkb](b-st-ext.html#zbkb) for details of these instructions. | | Notes to software developers Reversing bytes in words is very common in cryptography when setting a standard endianness for input and output data. Bit reversal within bytes is used for implementing the GHASH component of Galois/Counter Mode (GCM) \[[45](../biblio/bibliography.html#bib-nist:gcm)\]. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | RV32: zip rd, rs1 unzip rd, rs1 See [zbkb](b-st-ext.html#zbkb) for details of these instructions. | | Notes to software developers These instructions perform a bit-interleave (or de-interleave) operation, and are useful for implementing the 64-bit rotations in the SHA3 \[[46](../biblio/bibliography.html#bib-nist:fips:202)\] algorithm on a 32-bit architecture. On RV64, the relevant operations in SHA3 can be done natively using rotation instructions, so zip and unzip are not required. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#32-1-6-4-3-carry-less-multiply)32.1.6.4.3\. Carry-less Multiply RV32, RV64: clmul rd, rs1, rs2 clmulh rd, rs1, rs2 See [zbkc](b-st-ext.html#zbkc) for details of these instructions. See [32.1.5\. Data Independent Execution Latency Subset: Zkt](#crypto%5Fscalar%5Fzkt) for additional implementation requirements for these instructions, related to data independent execution latency. | | Notes to software developers As is mentioned there, obvious cryptographic use-cases for carry-less multiply are for Galois Counter Mode (GCM) block cipher operations. GCM is recommended by NIST as a block cipher mode of operation \[[45](../biblio/bibliography.html#bib-nist:gcm)\], and is the only _required_ mode for the TLS 1.3 protocol. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ##### [](#32-1-6-4-4-logic-with-negate)32.1.6.4.4\. Logic With Negate RV32, RV64: andn rd, rs1, rs2 orn rd, rs1, rs2 xnor rd, rs1, rs2 See [zbkb](b-st-ext.html#zbkb) for details of these instructions. These instructions are useful inside hash functions, block ciphers and for implementing software based side-channel countermeasures like masking. The `andn` instruction is also useful for constant time word-select in systems without the ternary Bitmanip `cmov` instruction. | | Notes to software developers In the context of Cryptography, these instructions are useful for: SHA3/Keccak Chi step, Bit-sliced function implementations, Software based power/EM side-channel countermeasures based on masking. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ##### [](#32-1-6-4-5-packing)32.1.6.4.5\. Packing RV32, RV64: RV64: pack rd, rs1, rs2 packw rd, rs1, rs2 packh rd, rs1, rs2 See [zbkb](b-st-ext.html#zbkb) for details of these instructions. | | Notes to software developers The pack\* instructions are useful for re-arranging halfwords within words, and generally getting data into the right shape prior to applying transforms. This is particularly useful for cryptographic algorithms which pass inputs around as (potentially unaligned) byte strings, but can operate on words made out of those byte strings. This occurs (for example) in AES when loading blocks and keys (which may not be word aligned) into registers to perform the round functions. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#32-1-6-4-6-crossbar-permutation-instructions)32.1.6.4.6\. Crossbar Permutation Instructions RV32, RV64: xperm4 rd, rs1, rs2 xperm8 rd, rs1, rs2 See [zbkx](b-st-ext.html#zbkx) for a complete description of these instructions. The `xperm4` instruction operates on nibbles.`GPR[rs1]` contains a vector of `XLEN/4` 4-bit elements.`GPR[rs2]` contains a vector of `XLEN/4` 4-bit indexes. The result is each element in `GPR[rs2]` replaced by the indexed element in `GPR[rs1]`, or zero if the index into `GPR[rs2]` is out of bounds. The `xperm8` instruction operates on bytes.`GPR[rs1]` contains a vector of `XLEN/8` 8-bit elements.`GPR[rs2]` contains a vector of `XLEN/8` 8-bit indexes. The result is each element in `GPR[rs2]` replaced by the indexed element in `GPR[rs1]`, or zero if the index into `GPR[rs2]` is out of bounds. | | Notes to software developers The instruction can be used to implement arbitrary bit permutations. For cryptography, they can accelerate bit-sliced implementations, permutation layers of block ciphers, masking based countermeasures and SBox operations. Lightweight block ciphers using 4-bit SBoxes include: PRESENT \[[47](../biblio/bibliography.html#bib-block:present)\], Rectangle \[[48](../biblio/bibliography.html#bib-block:rectangle)\], GIFT \[[49](../biblio/bibliography.html#bib-block:gift)\], Twine \[[50](../biblio/bibliography.html#bib-block:twine)\], Skinny, MANTIS \[[51](../biblio/bibliography.html#bib-block:skinny)\], Midori \[[52](../biblio/bibliography.html#bib-block:midori)\]. National ciphers using 8-bit SBoxes include: Camellia \[[53](../biblio/bibliography.html#bib-block:camellia)\] (Japan), Aria \[[54](../biblio/bibliography.html#bib-block:aria)\] (Korea), AES \[[30](../biblio/bibliography.html#bib-nist:fips:197)\] (USA, Belgium), SM4 \[[34](../biblio/bibliography.html#bib-gbt:sm4)\] (China) Kuznyechik (Russia). All of these SBoxes can be implemented efficiently, in constant time, using the xperm8 instruction\[[1](#%5Ffootnotedef%5F1 "View footnote.")\]. Note that this technique is also suitable for masking based side-channel countermeasures. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#crypto%5Fscalar%5Fappx%5Fes)32.1.7\. Entropy Source Rationale and Recommendations This **non-normative** appendix focuses on the rationale, security, self-certification, and implementation aspects of entropy sources. Hence we also discuss non-ISA system features that may be needed for cryptographic standards compliance and security testing. #### [](#32-1-7-1-checklists-for-design-and-self-certification)32.1.7.1\. Checklists for Design and Self-Certification The security of cryptographic systems is based on secret bits and keys. These bits need to be random and originate from cryptographically secure Random Bit Generators (RBGs). An Entropy Source (ES) is required to construct secure RBGs. While entropy source implementations do not have to be certified designs, RISC-V expects that they behave in a compatible manner and do not create unnecessary security risks to users. Self-evaluation and testing following appropriate security standards is usually needed to achieve this. * **ISA Architectural Tests.** Verify, to the extent possible, that RISC-V ISA requirements in this specification are correctly implemented. This includes the state transitions ([32.1.4\. Entropy Source](#crypto%5Fscalar%5Fes) and[32.1.7.6\. Suggested GetNoise Test Interface](#crypto%5Fscalar%5Fes%5Fgetnoise)), access control ([32.1.4.3\. Access Control to seed](#crypto%5Fscalar%5Fes%5Faccess)), and that `seed` ES16 `entropy` words can only be read destructively. The scope of RISC-V ISA architectural tests are those behaviors that are independent of the physical entropy source details. A smoke test ES module may be helpful in design phase. * **Technical justification for entropy.** This may take the form of a stochastic model or a heuristic argument that explains why the noise source output is from a random, rather than pseudorandom (deterministic) process, and is not easily predictable or externally observable. A complete physical model is not necessary; research literature can be cited. For example, one can show that a good ring oscillator noise derives an amount of physical entropy from local, spontaneously occurring Johnson-Nyquist thermal noise \[[55](../biblio/bibliography.html#bib-sa21)\], and is therefore not merely "random-looking". * **Entropy Source Design Review.** An entropy source is more than a noise source, and must have features such as health tests ([32.1.7.4\. Security Controls and Health Tests](#crypto%5Fscalar%5Fes%5Fsecurity%5Fcontrols)), a conditioner ([32.1.7.2.2\. Conditioning: Cryptographic and Non-Cryptographic](#crypto%5Fscalar%5Fappx%5Fes%5Fintro-cond)), and a security boundary with clearly defined interfaces. One may tabulate the SHALL statements of SP 800-90B \[[36](../biblio/bibliography.html#bib-tubake:18)\], FIPS 140-3 Implementation Guidance \[[56](../biblio/bibliography.html#bib-nicc21)\], AIS-31 \[[37](../biblio/bibliography.html#bib-kisc11)\] or other standards being used. Official and non-official checklist tables are available: * **Experimental Tests.** The raw noise source is subjected to entropy estimation as defined in NIST 800-90B, Section 3 \[[36](../biblio/bibliography.html#bib-tubake:18)\]. The interface described in [32.1.7.6\. Suggested GetNoise Test Interface](#crypto%5Fscalar%5Fes%5Fgetnoise) can used be to record datasets for this purpose. One also needs to show experimentally that the conditioner and health test components work appropriately to meet the ES16 output entropy requirements of [32.1.4.2\. Entropy Source Requirements](#crypto%5Fscalar%5Fes%5Freq). For SP 800-90B, NIST has made a min-entropy estimation package freely available: * **Resilience.** Above physical engineering steps should consider the operational environment of the device, which may be unexpected or hostile (actively attempting to exploit vulnerabilities in the design). See [32.1.7.5\. Implementation Strategies](#crypto%5Fscalar%5Fappx%5Fes%5Fimplementation) for a discussion of various implementation options. | | It is one of the goals of the RISC-V Entropy Source specification that a standard 90B Entropy Source Module or AIS-31 RNG IP may be licensed from a third party and integrated with a RISC-V processor design. Compared to older (FIPS 140-2) RNG and DRBG modules, an entropy source module may have a relatively small area (just a few thousand NAND2 gate equivalent). CMVP is introducing an "Entropy Source Validation Scope" which potentially allows 90B validations to be reused for different (FIPS 140-3) modules. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#32-1-7-2-standards-and-terminology)32.1.7.2\. Standards and Terminology As a fundamental security function, the generation of random numbers is governed by numerous standards and technical evaluation methods, the main ones being FIPS 140-3 \[[57](../biblio/bibliography.html#bib-ni19)\] \[[56](../biblio/bibliography.html#bib-nicc21)\] required for U.S. Federal use, and Common Criteria Methodology \[[58](../biblio/bibliography.html#bib-cr17)\] used in high-security evaluations internationally. Note that FIPS 140-3 is a significantly updated standard compared to its predecessor FIPS 140-2 and is only coming into use in the 2020s. These standards set many of the technical requirements for the RISC-V entropy source design, and we use their terminology if possible. ![es dataflow](_images/es_dataflow.svg) The `seed` CSR provides an Entropy Source (ES) interface, not a stateful random number generator. As a result, it can support arbitrary security levels. Cryptographic (AES, SHA-2/3) ISA Extensions can be used to construct high-speed DRBGs that are seeded from the entropy source. ##### [](#crypto%5Fscalar%5Fappx%5Fes%5Fintro-es)32.1.7.2.1\. Entropy Source (ES) Entropy sources are built by sampling and processing data from a noise source ([32.1.7.5.1\. Ring Oscillators](#crypto%5Fscalar%5Fappx%5Fes%5Fnoise%5Fsources)). We will only consider physical sources of true randomness in this work. Since these are directly based on natural phenomena and are subject to environmental conditions (which may be adversarial), they require features that monitor the "health" and quality of those sources. The requirements for physical entropy sources are specified in NIST SP 800-90B \[[36](../biblio/bibliography.html#bib-tubake:18)\] ([32.1.4.2.1\. NIST SP 800-90B / FIPS 140-3 Requirements](#crypto%5Fscalar%5Fes%5Freq%5F90b)) for U.S. Federal FIPS 140-3 \[[57](../biblio/bibliography.html#bib-ni19)\] evaluations and in BSI AIS-31 \[[59](../biblio/bibliography.html#bib-kisc01)\] \[[37](../biblio/bibliography.html#bib-kisc11)\] ([32.1.4.2.2\. BSI AIS-31 PTG.2 / Common Criteria Requirements](#crypto%5Fscalar%5Fes%5Freq%5Fptg2)) for high-security Common Criteria evaluations. There is some divergence in the types of health tests and entropy metrics mandated in these standards, and RISC-V enables support for both alternatives. ##### [](#crypto%5Fscalar%5Fappx%5Fes%5Fintro-cond)32.1.7.2.2\. Conditioning: Cryptographic and Non-Cryptographic Raw physical randomness (noise) sources are rarely statistically perfect, and some generate very large amounts of bits, which need to be "debiased" and reduced to a smaller number of bits. This process is called conditioning. A secure hash function is an example of a cryptographic conditioner. It is important to note that even though hashing may make any data look random, it does not increase its entropy content. Non-cryptographic conditioners and extractors such as von Neumann’s "debiased coin tossing" \[[60](../biblio/bibliography.html#bib-ne51)\] are easier to implement efficiently but may reduce entropy content (in individual bits removed) more than cryptographic hashes, which mix the input entropy very efficiently. However, they do not require cryptanalytic or computational hardness assumptions and are therefore inherently more future-proof. See [32.1.7.5.5\. Non-cryptographic Conditioners](#crypto%5Fscalar%5Fappx%5Fes%5Fnoncrypto) for a more detailed discussion. ##### [](#crypto%5Fscalar%5Fappx%5Fes%5Fintro-prng)32.1.7.2.3\. Pseudorandom Number Generator (PRNG) Pseudorandom Number Generators (PRNGs) use deterministic mathematical formulas to create abundant random numbers from a smaller amount of "seed" randomness. PRNGs are also divided into cryptographic and non-cryptographic ones. Non-cryptographic PRNGs, such as LFSRs and the linear-congruential generators found in many programming libraries, may generate statistically satisfactory random numbers but must never be used for cryptographic keying. This is because they are not designed to resist_cryptanalysis_; it is usually possible to take some output and mathematically derive the "seed" or the internal state of the PRNG from it. This is a security problem since knowledge of the state allows the attacker to compute future or past outputs. ##### [](#crypto%5Fscalar%5Fappx%5Fes%5Fintro-drbg)32.1.7.2.4\. Deterministic Random Bit Generator (DRBG) Cryptographic PRNGs are also known as Deterministic Random Bit Generators (DRBGs), a term used by SP 800-90A \[[38](../biblio/bibliography.html#bib-bake15)\]. A strong cryptographic algorithm such as AES \[[30](../biblio/bibliography.html#bib-nist:fips:197)\] or SHA-2/3 is used to produce random bits from a seed. The secret seed material is like a cryptographic key; determining the seed from the DRBG output is as hard as breaking AES or a strong hash function. This also illustrates that the seed/key needs to be long enough and come from a trusted Entropy Source. The DRBG should still be frequently refreshed (reseeded) for forward and backward security. #### [](#32-1-7-3-specific-rationale-and-considerations)32.1.7.3\. Specific Rationale and Considerations ##### [](#32-1-7-3-1-the-seed-csr)32.1.7.3.1\. The `seed` CSR See [32.1.4.1\. The seed CSR](#crypto%5Fscalar%5Fseed%5Fcsr). The interface was designed to be simple so that a vendor- and device-independent driver component (e.g., in Linux kernel, embedded firmware, or a cryptographic library) may use `seed` to generate truly random bits. An entropy source does not require a high-bandwidth interface; a single DRBG source initialization only requires 512 bits (256 bits of entropy), and DRBG output can be shared by any number of callers. Once initiated, a DRBG requires new entropy only to mitigate the risk of state compromise. From a security perspective, it is essential that the side effect of flushing the secret entropy bits occurs upon reading. Hence we mandate a write operation on this particular CSR. A blocking instruction may have been easier to use, but most users should be querying a (D)RBG instead of an entropy source. Without a polling-style mechanism, the entropy source could hang for thousands of cycles under some circumstances. A `wfi` or `pause`mechanism (at least potentially) allows energy-saving sleep on MCUs and context switching on higher-end CPUs. The reason for the particular `OPST = seed[31:0]` two-bit mechanism is to provide redundancy. The "fault" bit combinations `11` (`DEAD`) and `00`(`BIST`) are more likely for electrical reasons if feature discovery fails and the entropy source is actually not available. The 16-bit bandwidth was a compromise motivated by the desire to provide redundancy in the return value, some protection against potential Power/EM leakage (further alleviated by the 2:1 cryptographic conditioning discussed in [32.1.7.5.6\. Cryptographic Conditioners](#crypto%5Fscalar%5Fappx%5Fes%5Fcrypto-cond)), and the desire to have all of the bits "in the same place" on both RV32 and RV64 architectures for programming convenience. ##### [](#32-1-7-3-2-nist-sp-800-90b)32.1.7.3.2\. NIST SP 800-90B See [32.1.4.2.1\. NIST SP 800-90B / FIPS 140-3 Requirements](#crypto%5Fscalar%5Fes%5Freq%5F90b). SP 800-90C \[[39](../biblio/bibliography.html#bib-bakero:21)\] states that each conditioned block of n bits is required to have n+64 bits of input entropy to attain full entropy. Hence NIST SP 800-90B \[[36](../biblio/bibliography.html#bib-tubake:18)\] min-entropy assessment must guarantee at least 128 + 64 = 192 bits input entropy per 256-bit block (\[[39](../biblio/bibliography.html#bib-bakero:21)\], Sections 4.1\. and 4.3.2). Only then a hashing of 16 \* 16 = 256 bits from the entropy source will produce the desired 128 bits of full entropy. This follows from the specific requirements, threat model, and distinguishability proof contained in SP 800-90C \[[39](../biblio/bibliography.html#bib-bakero:21)\], Appendix A. The implied min-entropy rate is 192/256=12/16=0.75\. The expected Shannon entropy is much larger. In FIPS 140-3 / SP 800-90 classification, an RBG2(P) construction is a cryptographically secure RBG with continuous access to a physical entropy source (`seed`) and output generated by a fully seeded, secure DRBG. The entropy source can also be used to build RBG3 full entropy sources \[[39](../biblio/bibliography.html#bib-bakero:21)\]. The concatenation of output words corresponds to the `Get_ES_Bitstring` function. The 128-bit output block size was selected because that is the output size of the CBC-MAC conditioner specified in Appendix F of \[[36](../biblio/bibliography.html#bib-tubake:18)\] and also the smallest key size we expect to see in applications. If NIST SP 800-90B certification is chosen, the entropy source should implement at least the health tests defined in Section 4.4 of \[[36](../biblio/bibliography.html#bib-tubake:18)\]: the repetition count test and adaptive proportion test, or show that the same flaws will be detected by vendor-defined tests. ##### [](#32-1-7-3-3-bsi-ais-31)32.1.7.3.3\. BSI AIS-31 See [32.1.4.2.2\. BSI AIS-31 PTG.2 / Common Criteria Requirements](#crypto%5Fscalar%5Fes%5Freq%5Fptg2). PTG.2 is one of the security and functionality classes defined in BSI AIS 20/31 \[[37](../biblio/bibliography.html#bib-kisc11)\]. The PTG.2 source requirements work as a building block for other types of BSI generators (e.g., DRBGs, or PTG.3 TRNG with appropriate software post-processing). For validation purposes, the PTG.2 requirements may be mapped to security controls T1-3 ([32.1.7.4\. Security Controls and Health Tests](#crypto%5Fscalar%5Fes%5Fsecurity%5Fcontrols)) and the interface as follows: * P1 **\[PTG.2.1\]** Start-up tests map to T1 and reset-triggered (on-demand)`BIST` tests. * P2 **\[PTG.2.2\]** Continuous testing total failure maps to T2 and the`DEAD` state. * P3 **\[PTG.2.3\]** Online tests are continuous tests of T2 – entropy output is prevented in the `BIST` state. * P4 **\[PTG.2.4\]** Is related to the design of effective entropy source health tests, which we encourage. * P5 **\[PTG.2.5\]** Raw random sequence may be checked via the GetNoise interface ([32.1.7.6\. Suggested GetNoise Test Interface](#crypto%5Fscalar%5Fes%5Fgetnoise)). * P6 **\[PTG.2.6\]** Test Procedure A \[[37](../biblio/bibliography.html#bib-kisc11)\] (Sect 2.4.4.1) is a part of the evaluation process, and we suggest self-evaluation using these tests even if AIS-31 certification is not sought. * P7 **\[PTG.2.7\]** Average Shannon entropy of "internal random bits" exceeds 0.997. Note how P7 concerns Shannon entropy, not min-entropy as with NIST sources. Hence the min-entropy requirement needs to be also stated. PTG.2 modules built and certified to the AIS-31 standard can also meet the "full entropy" condition after 2:1 cryptographic conditioning, but not necessarily so. The technical validation process is somewhat different. ##### [](#32-1-7-3-4-virtual-sources)32.1.7.3.4\. Virtual Sources [32.1.4.2.3\. Virtual Sources: Security Requirement](#crypto%5Fscalar%5Fes%5Freq%5Fvirt). All sources that are not direct physical sources (meeting the SP 800-90B or the AIS-31 PTG.2 requirements) need to meet the security requirements of virtual entropy sources. It is assumed that a virtual entropy source is not a limiting, shared bandwidth resource (but a software DRBG). DRBGs can be used to feed other (virtual) DRBGs, but that does not increase the absolute amount of entropy in the system. The entropy source must be able to support current and future security standards and applications. The 256-bit requirement maps to "Category 5" of NIST Post-Quantum Cryptography (4.A.5 "Security Strength Categories" in \[[40](../biblio/bibliography.html#bib-ni16)\]) and TOP SECRET schemes in Suite B and the newer U.S. Government CNSA Suite \[[61](../biblio/bibliography.html#bib-ns15)\]. ##### [](#crypto%5Fscalar%5Fappx%5Fes%5Faccess)32.1.7.3.5\. Security Considerations for Direct Hardware Access [32.1.4.3\. Access Control to seed](#crypto%5Fscalar%5Fes%5Faccess). The ISA implementation and system design must try to ensure that the hardware-software interface minimizes avenues for adversarial information flow even if not explicitly forbidden in the specification. For security, virtualization requires both conditioning and DRBG processing of physical entropy output. It is recommended if a single physical entropy source is shared between multiple different virtual machines or if the guest OS is untrusted. A virtual entropy source is significantly more resistant to depletion attacks and also lessens the risk from covert channels. The direct `mseccfg.[s,u]seed` option allows one to draw a security boundary around a component in relation to Sensitive Security Parameter (SSP) flows, even if that component is not in M mode. This is helpful when implementing trusted enclaves. Such modules can enforce the entire key lifecycle from birth (in the entropy source) to death (zeroization) to occur without the key being passed across the boundary to external code. **Depletion.**Active polling may deny the entropy source to another simultaneously running consumer. This can (for example) delay the instantiation of that virtual machine if it requires entropy to initialize fully. **Covert Channels.**Direct access to a component such as the entropy source can be used to establish communication channels across security boundaries. Active polling from one consumer makes the resource unavailable WAIT instead of ES16 to another (which is polling infrequently). Such interactions can be used to establish low-bandwidth channels. **Hardware Fingerprinting.**An entropy source (and its noise source circuits) may have a uniquely identifiable hardware "signature." This can be harmless or even useful in some applications (as random sources may exhibit Physically Un-clonable Function (PUF) -like features) but highly undesirable in others (anonymized virtualized environments and enclaves). A DRBG masks such statistical features. **Side Channels.**Some of the most devastating practical attacks against real-life cryptosystems have used inconsequential-looking additional information, such as padding error messages \[[62](../biblio/bibliography.html#bib-bafoka:12)\] or timing information \[[63](../biblio/bibliography.html#bib-mosuei:20)\]. We urge implementers against creating unnecessary information flows via status or custom bits or to allow any other mechanism to disable or affect the entropy source output. All information flows and interaction mechanisms must be considered from an adversarial viewpoint: the fewer the better. As an example of side-channel analysis, we note that the entropy polling interface is typically not "constant time." One needs to analyze what kind of information is revealed via the timing oracle; one way of doing it is to model `seed` as a rejection sampler. Such a timing oracle can reveal information about the noise source type and entropy source usage, but not about the random output`entropy` bits themselves. If it does, additional countermeasures are necessary. #### [](#crypto%5Fscalar%5Fes%5Fsecurity%5Fcontrols)32.1.7.4\. Security Controls and Health Tests The primary purpose of a cryptographic entropy source is to produce secret keying material. In almost all cases, a hardware entropy source must implement appropriate _security controls_ to guarantee unpredictability, prevent leakage, detect attacks, and deny adversarial control over the entropy output or ts generation mechanism. Explicit security controls are required for security testing and certification. Many of the security controls built into the device are called "health checks." Health checks can take the form of integrity checks, start-up tests, and on-demand tests. These tests can be implemented in hardware or firmware, typically both. Several are mandated by standards such as NIST SP 800-90B \[[57](../biblio/bibliography.html#bib-ni19)\]. The choice of appropriate health tests depends on the certification target, system architecture, threat model, entropy source type, and other factors. Health checks are not intended for hardware diagnostics but for detecting security issues. Hence the default action in case of a failure should be aimed at damage control: Limiting further output and preventing weak crypto keys from being generated. We discuss three specific testing requirements T1-T3\. The testing requirement follows from the definition of an Entropy Source; without it, the module is simply a noise source and can’t be trusted to safely generate keying material. ##### [](#32-1-7-4-1-t1-on-demand-testing)32.1.7.4.1\. T1: On-demand testing A sequence of simple tests is invoked via resetting, rebooting, or powering up the hardware (not an ISA signal). The implementation will simply return `BIST` during the initial start-up self-test period; in any case, the driver must wait for them to finish before starting cryptographic operations. Upon failure, the entropy source will enter a no-output `DEAD` state. **Rationale.**Interaction with hardware self-test mechanisms from the software side should be minimal; the term "on-demand" does not mean that the end-user or application program should be able to invoke them in the field (the term is a throwback to an age of discrete, non-autonomous crypto devices with human operators). ##### [](#32-1-7-4-2-t2-continuous-checks)32.1.7.4.2\. T2: Continuous checks If an error is detected in continuous tests or environmental sensors, the entropy source will enter a no-output state. We define that a non-critical alarm is signaled if the entropy source returns to `BIST` state from live (`WAIT` or `ES16`) states. Critical failures will result in `DEAD` state immediately. A hardware-based continuous testing mechanism must not make statistical information externally available, and it must be zeroized periodically or upon demand via reset, power-up, or similar signal. **Rationale.**Physical attacks can occur while the device is running. The design should avoid guiding such active attacks by revealing detailed status information. Upon detection of an attack, the default action should be aimed at damage control — to prevent weak crypto keys from being generated. The statistical nature of some tests makes "type-1" false positives a possibility. There may also be requirements for signaling of non-fatal alarms; AIS 31 specifies "noise alarms" that can go off with non-negligible probability even if the device is functioning correctly; these can be signaled with `BIST`. There rarely is anything that can or should be done about a non-fatal alarm condition in an operator-free, autonomous system. The state of statistical runtime health checks (such as counters) is potentially correlated with some secret keying material, hence the zeroization requirement. ##### [](#32-1-7-4-3-t3-fatal-error-states)32.1.7.4.3\. T3: Fatal error states Since the security of most cryptographic operations depends on the entropy source, a system-wide "default deny" security policy approach is appropriate for most entropy source failures. A hardware test failure should at least result in the `DEAD` state and possibly reset/halt. It’s a show stopper: The entropy source (or its cryptographic client application) _must not_ be allowed to run if its secure operation can’t be guaranteed. **Rationale.**These tests can complement other integrity and tamper resistance mechanisms (See Chapter 18 of \[[64](../biblio/bibliography.html#bib-an20)\] for examples). Some hardware random generators are, by their physical construction, exposed to relatively non-adversarial environmental and manufacturing issues. However, even such "innocent" failure modes may indicate a _fault attack_ \[[65](../biblio/bibliography.html#bib-kascve13)\] and therefore should be addressed as a system integrity failure rather than as a diagnostic issue. Security architects will understand to use permanent or hard-to-recover "security-fuse" lockdowns only if the threshold of a test is such that the probability of false-positive is negligible over the entire device lifetime. ##### [](#32-1-7-4-4-information-flows)32.1.7.4.4\. Information Flows Some of the most devastating practical attacks against real-life cryptosystems have used inconsequential-looking additional information, such as padding error messages \[[62](../biblio/bibliography.html#bib-bafoka:12)\] or timing information \[[63](../biblio/bibliography.html#bib-mosuei:20)\]. In cryptography, such out-of-band information sources are called "oracles." To guarantee that no sensitive data is read twice and that different callers don’t get correlated output, it is required that hardware implements _wipe-on-read_ on the randomness pathway during each read (successful poll). For the same reasons, only complete and fully processed random words shall be made available via `entropy` (ES16 status of `seed`). This also applies to the raw noise source. The raw source interface has been delegated to an optional vendor-specific test interface. Importantly the test interface and the main interface should not be operational at the same time. > The noise source state shall be protected from adversarial knowledge or influence to the greatest extent possible. The methods used for this shall be documented, including a description of the (conceptual) security boundary’s role in protecting the noise source from adversarial observation or influence. — NIST SP 800-90B Noise Source Requirements An entropy source is a singular resource, subject to depletion and also covert channels \[[66](../biblio/bibliography.html#bib-evpo16)\]. Observation of the entropy can be the same as the observation of the noise source output, as cryptographic conditioning is mandatory only as a post-processing step. SP 800-90B and other security standards mandate protection of noise bits from observation and also influence. #### [](#crypto%5Fscalar%5Fappx%5Fes%5Fimplementation)32.1.7.5\. Implementation Strategies As a general rule, RISC-V specifies the ISA only. We provide some additional suggestions so that portable, vendor-independent middleware and kernel components can be created. The actual hardware implementation and certification are left to vendors and circuit designers; the discussion in this Section is purely informational. When considering implementation options and trade-offs, one must look at the entire information flow. 1. **A Noise Source** generates private, unpredictable signals from stable and well-understood physical random events. 2. **Sampling** digitizes the noise signal into a raw stream of bits. This raw data also needs to be protected by the design. 3. **Continuous health tests** ensure that the noise source and its environment meet their operational parameters. 4. **Non-cryptographic conditioners** remove much of the bias and correlation in input noise. 5. **Cryptographic conditioners** produce full entropy output, completely indistinguishable from ideal random. 6. **DRBG** takes in `>=256` bits of seed entropy as keying material and uses a "one way" cryptographic process to rapidly generate bits on demand (without revealing the seed/state). Steps 1-4 (possibly 5) are considered to be part of the Entropy Source (ES) and provided by the `seed` CSR. Adding the software-side cryptographic steps 5-6 and control logic complements it into a True Random Number Generator (TRNG). ##### [](#crypto%5Fscalar%5Fappx%5Fes%5Fnoise%5Fsources)32.1.7.5.1\. Ring Oscillators We will give some examples of common noise sources that can be implemented in the processor itself (using standard cells). The most common entropy source type in production use today is based on "free running" ring oscillators and their timing jitter. Here, an odd number of inverters is connected into a loop from which noise source bits are sampled in relation to a reference clock \[[67](../biblio/bibliography.html#bib-balumi:11)\]. The sampled bit sequence may be expected to be relatively uncorrelated (close to IID) if the sample rate is suitably low \[[37](../biblio/bibliography.html#bib-kisc11)\]. However, further processing is usually required. AMD \[[68](../biblio/bibliography.html#bib-am17)\], ARM \[[69](../biblio/bibliography.html#bib-ar17)\], and IBM \[[70](../biblio/bibliography.html#bib-libabo:13)\] are examples of ring oscillator TRNGs intended for high-security applications. There are related metastability-based generator designs such as Transition Effect Ring Oscillator (TERO) \[[71](../biblio/bibliography.html#bib-vadr10)\]. The differential/feedback Intel construction \[[72](../biblio/bibliography.html#bib-hakoma12)\] is slightly different but also falls into the same general metastable oscillator-based category. The main benefits of ring oscillators are: (1) They can be implemented with standard cell libraries without external components — and even on FPGAs \[[73](../biblio/bibliography.html#bib-vafiau:10)\], (2) there is an established theory for their behavior \[[74](../biblio/bibliography.html#bib-hale98)\] \[[75](../biblio/bibliography.html#bib-halile99)\] \[[67](../biblio/bibliography.html#bib-balumi:11)\], and (3) ample precedent exists for testing and certifying them at the highest security levels. Ring oscillators also have well-known implementation pitfalls. Their output is sometimes highly dependent on temperature, which must be taken into account in testing and modeling. If the ring oscillator construction is parallelized, it is important that the number of stages and/or inverters in each chain is suitable to avoid entropy reduction due to harmonic "Huyghens synchronization" \[[76](../biblio/bibliography.html#bib-ba86)\]. Such harmonics can also be inserted maliciously in a frequency injection attack, which can have devastating results \[[77](../biblio/bibliography.html#bib-mamo09)\]. Countermeasures are related to circuit design; environmental sensors, electrical filters, and usage of a differential oscillator may help. ##### [](#32-1-7-5-2-shot-noise)32.1.7.5.2\. Shot Noise A category of random sources consisting of discrete events and modeled as a Poisson process is called "shot noise." There’s a long-established precedent of certifying them; the AIS 31 document \[[37](../biblio/bibliography.html#bib-kisc11)\] itself offers reference designs based on noisy diodes. Shot noise sources are often more resistant to temperature changes than ring oscillators. Some of these generators can also be fully implemented with standard cells (The Rambus / Inside Secure generic TRNG IP \[[78](../biblio/bibliography.html#bib-ra20)\] is described as a Shot Noise generator). ##### [](#32-1-7-5-3-other-types-of-noise)32.1.7.5.3\. Other types of noise It may be possible to certify more exotic noise sources and designs, although their stochastic model needs to be equally well understood, and their CPU interfaces must be secure. See [32.1.7.5.8\. Quantum vs. Classical Random](#crypto%5Fscalar%5Fappx%5Fes%5Fquantum) for a discussion of Quantum entropy sources. ##### [](#crypto%5Fscalar%5Fappx%5Fes%5Fcont-tests)32.1.7.5.4\. Continuous Health Tests Health monitoring requires some state information related to the noise source to be maintained. The tests should be designed in a way that a specific number of samples guarantees a state flush (no hung states). We suggest flush size `W =< 1024` to match with the NIST SP 800-90B required tests (See Section 4.4 in \[[36](../biblio/bibliography.html#bib-tubake:18)\]). The state is also fully zeroized in a system reset. The two mandatory tests can be built with minimal circuitry. Full histograms are not required, only simple counter registers: repetition count, window count, and sample count. Repetition count is reset every time the output sample value changes; if the count reaches a certain cutoff limit, a noise alarm (`BIST`) or failure (`DEAD`) is signaled. The window counter is used to save every W’th output (typically `W` in { 512, 1024 }). The frequency of this reference sample in the following window is counted; cutoff values are defined in the standard. We see that the structure of the mandatory tests is such that, if well implemented, no information is carried beyond a limit of `W` samples. Section 4.5 of \[[36](../biblio/bibliography.html#bib-tubake:18)\] explicitly permits additional developer-defined tests, and several more were defined in early versions of FIPS 140-1 before being "crossed out." The choice of additional tests depends on the nature and implementation of the physical source. Especially if a non-cryptographic conditioner is used in hardware, it is possible that the AIS 31 \[[37](../biblio/bibliography.html#bib-kisc11)\] online tests are implemented by driver software. They can also be implemented in hardware. For some security profiles, AIS 31 mandates that their tolerances are set in a way that the probability of an alarm is at least 10\-6yearly under "normal usage." Such requirements are problematic in modern applications since their probability is too high for critical systems. There rarely is anything that can or should be done about a non-fatal alarm condition in an operator-free, autonomous system. However, AIS 31 allows the DRBG component to keep running despite a failure in its Entropy Source, so we suggest re-entering a temporary `BIST`state ([32.1.7.4\. Security Controls and Health Tests](#crypto%5Fscalar%5Fes%5Fsecurity%5Fcontrols)) to signal a non-fatal statistical error if such (non-actionable) signaling is necessary. Drivers and applications can react to this appropriately (or simply log it), but it will not directly affect the availability of the TRNG. A permanent error condition should result in `DEAD` state. ##### [](#crypto%5Fscalar%5Fappx%5Fes%5Fnoncrypto)32.1.7.5.5\. Non-cryptographic Conditioners As noted in [32.1.7.2.2\. Conditioning: Cryptographic and Non-Cryptographic](#crypto%5Fscalar%5Fappx%5Fes%5Fintro-cond), physical randomness sources generally require a post-processing step called _conditioning_ to meet the desired quality requirements, which are outlined in[32.1.4.2\. Entropy Source Requirements](#crypto%5Fscalar%5Fes%5Freq). The approach taken in this interface is to allow a combination of non-cryptographic and cryptographic filtering to take place. The first stage (hardware) merely needs to be able to distill the entropy comfortably above the necessary level. * One may take a set of bits from a noise source and XOR them together to produce a less biased (and more independent) bit. However, such an XOR may introduce "pseudorandomness" and make the output difficult to analyze. * The von Neumann extractor \[[60](../biblio/bibliography.html#bib-ne51)\] looks at consecutive pairs of bits, rejects 00 and 11, and outputs 0 or 1 for 01 and 10, respectively. It will reduce the number of bits to less than 25% of the original, but the output is provably unbiased (assuming independence). * Blum’s extractor \[[79](../biblio/bibliography.html#bib-bl86)\] can be used on sources whose behavior resembles N-state Markov chains. If its assumptions hold, it also removes dependencies, creating an independent and identically distributed (IID) source. * Other linear and non-linear correctors such as those discussed by Dichtl and Lacharme \[[80](../biblio/bibliography.html#bib-la08)\]. Note that the hardware may also implement a full cryptographic conditioner in the entropy source, even though the software driver still needs a cryptographic conditioner, too ([32.1.4.2\. Entropy Source Requirements](#crypto%5Fscalar%5Fes%5Freq)). **Rationale:**The main advantage of non-cryptographic extractors is in their energy efficiency, relative simplicity, and amenability to mathematical analysis. If well designed, they can be evaluated in conjunction with a stochastic model of the noise source itself. They do not require computational hardness assumptions. ##### [](#crypto%5Fscalar%5Fappx%5Fes%5Fcrypto-cond)32.1.7.5.6\. Cryptographic Conditioners For secure use, cryptographic conditioners are always required on the software side of the ISA boundary. They may also be implemented on the hardware side if necessary. In any case, the `entropy` ES16 output must always be compressed 2:1 (or more) before being used as keying material or considered "full entropy." Examples of cryptographic conditioners include the random pool of the Linux operating system, secure hash functions (SHA-2/3, SHAKE \[[46](../biblio/bibliography.html#bib-nist:fips:202)\] \[[29](../biblio/bibliography.html#bib-nist:fips:180:4)\]), and the AES / CBC-MAC construction in Appendix F, SP 800-90B \[[36](../biblio/bibliography.html#bib-tubake:18)\]. In some constructions, such as the Linux RNG and SHA-3/SHAKE \[[46](../biblio/bibliography.html#bib-nist:fips:202)\] based generators, the cryptographic conditioning and output (DRBG) generation are provided by the same component. **Rationale:**For many low-power targets constructions the type of hardware AES CBC-MAC conditioner used by Intel \[[81](../biblio/bibliography.html#bib-me18)\] and AMD \[[68](../biblio/bibliography.html#bib-am17)\] would be too complex and energy-hungry to implement solely to serve the `seed` CSR. On the other hand, simpler non-cryptographic conditioners may be too wasteful on input entropy if high-quality random output is required — (ARM TrustZone TRBG \[[69](../biblio/bibliography.html#bib-ar17)\] outputs only 10Kbit/sec at 200 MHz.) Hence a resource-saving compromise is made between hardware and software generation. ##### [](#crypto%5Fscalar%5Fappx%5Fes%5Fdrbgs)32.1.7.5.7\. The Final Random: DRBGs All random bits reaching end users and applications must come from a cryptographic DRBG. These are generally implemented by the driver component in software. The RISC-V AES and SHA instruction set extensions should be used if available since they offer additional security features such as timing attack resistance. Currently recommended DRBGs are defined in NIST SP 800-90A (Rev 1) \[[38](../biblio/bibliography.html#bib-bake15)\]: `CTR_DRBG`, `Hash_DRBG`, and `HMAC_DRBG`. Certification often requires known answer tests (KATs) for the symmetric components and the DRBG as a whole. These are significantly easier to implement in software than in hardware. In addition to the directly certifiable SP 800-90A DRBGs, a Linux-style random pool construction based on ChaCha20 \[[82](../biblio/bibliography.html#bib-mu20)\] can be used, or an appropriate construction based on SHAKE256 \[[46](../biblio/bibliography.html#bib-nist:fips:202)\]. These are just recommendations; programmers can adjust the usage of the CPU Entropy Source to meet future requirements. ##### [](#crypto%5Fscalar%5Fappx%5Fes%5Fquantum)32.1.7.5.8\. Quantum vs. Classical Random > The NCSC believes that classical RNGs will continue to meet our needs for government and military applications for the foreseeable future. — U.K. NCSC QRNG Guidance March 2020 A Quantum Random Number Generator (QRNG) is a TRNG whose source of randomness can be unambiguously identified to be a specific quantum phenomenon such as quantum state superposition, quantum state entanglement, Heisenberg uncertainty, quantum tunneling, spontaneous emission, or radioactive decay \[[83](../biblio/bibliography.html#bib-it19)\]. Direct quantum entropy is theoretically the best possible kind of entropy. A typical TRNG based on electronic noise is also largely based on quantum phenomena and is equally unpredictable - the difference is that the relative amount of quantum and classical physics involved is difficult to quantify for a classical TRNG. QRNGs are designed in a way that allows the amount of quantum-origin entropy to be modeled and estimated. This distinction is important in the security model used by QKD (Quantum Key Distribution) security mechanisms which can be used to protect the physical layer (such as fiber optic cables) against interception by using quantum mechanical effects directly. This security model means that many of the available QRNG devices do not use cryptographic conditioning and may fail cryptographic statistical requirements \[[84](../biblio/bibliography.html#bib-huhe20)\]. Many implementers may consider them to be entropy sources instead. Relatively little research has gone into QRNG implementation security, but many QRNG designs are arguably more susceptible to leakage than classical generators (such as ring oscillators) as they tend to employ external components and mixed materials. As an example, amplification of a photon detector signal may be observable in power analysis, which classical noise-based sources are designed to resist. ##### [](#32-1-7-5-9-post-quantum-cryptography)32.1.7.5.9\. Post-Quantum Cryptography PQC public-key cryptography standards \[[40](../biblio/bibliography.html#bib-ni16)\] do not require quantum-origin randomness, just sufficiently secure keying material. Recall that cryptography aims to protect the confidentiality and integrity of data itself and does not place any requirements on the physical communication channel (like QKD). Classical good-quality TRNGs are perfectly suitable for generating the secret keys for PQC protocols that are hard for quantum computers to break but implementable on classical computers. What matters in cryptography is that the secret keys have enough true randomness (entropy) and that they are generated and stored securely. Of course, one must avoid DRBGs that are based on problems that are easily solvable with quantum computers, such as factoring \[[85](../biblio/bibliography.html#bib-sh94)\] in the case of the Blum-Blum-Shub generator \[[86](../biblio/bibliography.html#bib-blblsh86)\]. Most symmetric algorithms are not affected as the best quantum attacks are still exponential to key size \[[87](../biblio/bibliography.html#bib-gr96)\]. As an example, the original Intel RNG \[[81](../biblio/bibliography.html#bib-me18)\], whose output generation is based on AES-128, can be attacked using Grover’s algorithm with approximately square-root effort \[[88](../biblio/bibliography.html#bib-janaro:20)\]. While even "64-bit" quantum security is extremely difficult to break, many applications specify a higher security requirement. NIST \[[40](../biblio/bibliography.html#bib-ni16)\] defines AES-128 to be "Category 1" equivalent post-quantum security, while AES-256 is "Category 5" (highest). We avoid this possible future issue by exposing direct access to the entropy source which can derive its security from information-theoretic assumptions only. #### [](#crypto%5Fscalar%5Fes%5Fgetnoise)32.1.7.6\. Suggested GetNoise Test Interface Compliance testing, characterization, and configuration of entropy sources require access to raw, unconditioned noise samples. This conceptual test interface is named GetNoise in Section 2.3.2 of NIST SP 800-90B \[[36](../biblio/bibliography.html#bib-tubake:18)\]. Since this type of interface is both necessary for security testing and also constitutes a potential backdoor to the cryptographic key generation process, we define a safety behavior that compliant implementations can have for temporarily disabling the entropy source `seed` CSR interface during test. In order for shared RISC-V self-certification scripts (and drivers) to accommodate the test interface in a secure fashion, we suggest that it is implemented as a custom, M-mode only CSR, denoted here as `mnoise`. This non-normative interface is not intended to be used as a source of randomness or for other production use. We define the semantics for single bit for this interface, `mnoise[31]`, which is named `NOISE_TEST`, which will affect the behavior of `seed`if implemented. When `NOISE_TEST = 1` in `mnoise`, the `seed` CSR must not return anything via `ES16`; it should be in `BIST` state unless the source is `DEAD`. When `NOISE_TEST` is again disabled, the entropy source shall return from `BIST` via an appropriate zeroization and self-test mechanism. The behavior of other input and output bits is largely left to the vendor (as they depend on the technical details of the physical entropy source), as is the address of the custom `mnoise` CSR. Other contents and behavior of the CSR only can be interpreted in the context of `mvendorid`, `marchid`, and`mimpid` CSR identifiers. When not implemented (e.g., in virtual machines), `mnoise` can permanently read zero (`0x00000000`) and ignore writes. When available, but `NOISE_TEST = 0`, `mnoise` can return a nonzero constant (e.g. `0x00000001`) but no noise samples. ![es noisetest](_images/es_noisetest.svg) Figure 2\. Entropy source can’t be read in test mode. In `NOISE_TEST` mode, the WAIT and ES16 states are unreachable, and no entropy is output. Implementation of test interfaces that directly affect ES16 entropy output from the `seed` CSR interface is discouraged. Such vendor test interfaces have been exploited in attacks. For example, an ECDSA \[[89](../biblio/bibliography.html#bib-nist:fips:186:4)\] signature process without sufficient entropy will not only create an insecure signature but can also reveal the secret signing key, that can be used for authentication forgeries by attackers. Hence even a temporary lapse in `entropy` security may have serious security implications. ### [](#crypto%5Fscalar%5Fappx%5Fmaterials)32.1.8\. Supplementary Materials While this document contains the specifications for the RISC-V cryptography extensions, numerous supplementary materials and example codes have also been developed. All of the materials related to the RISC-V Cryptography extension live in a GitHub Repository, located at * `doc/`Contains the source code for this document. * `doc/supp/`Contains supplementary information and recommendations for implementers of software and hardware. * `benchmarks/`Example software implementations. * `rtl/`Example Verilog implementations of each instruction. * `sail/`Formal model implementations in Sail. ### [](#crypto%5Fscalar%5Fappx%5Fsail)32.1.9\. Supporting Sail Code This section contains the supporting Sail code referenced by the instruction descriptions throughout the specification. The[Sail Manual](https://alasdair.github.io/manual.html)is recommended reading in order to best understand the supporting code. ```sail /* Auxiliary function for performing GF multiplication */ val xt2 : bits(8) -> bits(8) function xt2(x) = { (x << 1) ^ (if bit_to_bool(x[7]) then 0x1b else 0x00) } val xt3 : bits(8) -> bits(8) function xt3(x) = x ^ xt2(x) /* Multiply 8-bit field element by 4-bit value for AES MixCols step */ val gfmul : (bits(8), bits(4)) -> bits(8) function gfmul( x, y) = { (if bit_to_bool(y[0]) then x else 0x00) ^ (if bit_to_bool(y[1]) then xt2( x) else 0x00) ^ (if bit_to_bool(y[2]) then xt2(xt2( x)) else 0x00) ^ (if bit_to_bool(y[3]) then xt2(xt2(xt2(x))) else 0x00) } /* 8-bit to 32-bit partial AES Mix Column - forwards */ val aes_mixcolumn_byte_fwd : bits(8) -> bits(32) function aes_mixcolumn_byte_fwd(so) = { gfmul(so, 0x3) @ so @ so @ gfmul(so, 0x2) } /* 8-bit to 32-bit partial AES Mix Column - inverse*/ val aes_mixcolumn_byte_inv : bits(8) -> bits(32) function aes_mixcolumn_byte_inv(so) = { gfmul(so, 0xb) @ gfmul(so, 0xd) @ gfmul(so, 0x9) @ gfmul(so, 0xe) } /* 32-bit to 32-bit AES forward MixColumn */ val aes_mixcolumn_fwd : bits(32) -> bits(32) function aes_mixcolumn_fwd(x) = { let s0 : bits (8) = x[ 7.. 0]; let s1 : bits (8) = x[15.. 8]; let s2 : bits (8) = x[23..16]; let s3 : bits (8) = x[31..24]; let b0 : bits (8) = xt2(s0) ^ xt3(s1) ^ (s2) ^ (s3); let b1 : bits (8) = (s0) ^ xt2(s1) ^ xt3(s2) ^ (s3); let b2 : bits (8) = (s0) ^ (s1) ^ xt2(s2) ^ xt3(s3); let b3 : bits (8) = xt3(s0) ^ (s1) ^ (s2) ^ xt2(s3); b3 @ b2 @ b1 @ b0 /* Return value */ } /* 32-bit to 32-bit AES inverse MixColumn */ val aes_mixcolumn_inv : bits(32) -> bits(32) function aes_mixcolumn_inv(x) = { let s0 : bits (8) = x[ 7.. 0]; let s1 : bits (8) = x[15.. 8]; let s2 : bits (8) = x[23..16]; let s3 : bits (8) = x[31..24]; let b0 : bits (8) = gfmul(s0, 0xE) ^ gfmul(s1, 0xB) ^ gfmul(s2, 0xD) ^ gfmul(s3, 0x9); let b1 : bits (8) = gfmul(s0, 0x9) ^ gfmul(s1, 0xE) ^ gfmul(s2, 0xB) ^ gfmul(s3, 0xD); let b2 : bits (8) = gfmul(s0, 0xD) ^ gfmul(s1, 0x9) ^ gfmul(s2, 0xE) ^ gfmul(s3, 0xB); let b3 : bits (8) = gfmul(s0, 0xB) ^ gfmul(s1, 0xD) ^ gfmul(s2, 0x9) ^ gfmul(s3, 0xE); b3 @ b2 @ b1 @ b0 /* Return value */ } /* Turn a round number into a round constant for AES. Note that the AES64KS1I instruction is defined such that the r argument is always in the range 0x0..0xA. Values of rnum outside the range 0x0..0xA do not decode to the AES64KS1I instruction. The 0xA case is used specifically for the AES-256 KeySchedule, and this function is never called in that case. */ val aes_decode_rcon : bits(4) -> bits(32) function aes_decode_rcon(r) = { assert(r <_u 0xA); match r { 0x0 => 0x00000001, 0x1 => 0x00000002, 0x2 => 0x00000004, 0x3 => 0x00000008, 0x4 => 0x00000010, 0x5 => 0x00000020, 0x6 => 0x00000040, 0x7 => 0x00000080, 0x8 => 0x0000001b, 0x9 => 0x00000036, _ => internal_error(__FILE__, __LINE__, "Unexpected AES r") /* unreachable -- required to silence Sail warning */ } } /* SM4 SBox - only one sbox for forwards and inverse */ let sm4_sbox_table : vector(256, bits(8)) = [ 0xD6, 0x90, 0xE9, 0xFE, 0xCC, 0xE1, 0x3D, 0xB7, 0x16, 0xB6, 0x14, 0xC2, 0x28, 0xFB, 0x2C, 0x05, 0x2B, 0x67, 0x9A, 0x76, 0x2A, 0xBE, 0x04, 0xC3, 0xAA, 0x44, 0x13, 0x26, 0x49, 0x86, 0x06, 0x99, 0x9C, 0x42, 0x50, 0xF4, 0x91, 0xEF, 0x98, 0x7A, 0x33, 0x54, 0x0B, 0x43, 0xED, 0xCF, 0xAC, 0x62, 0xE4, 0xB3, 0x1C, 0xA9, 0xC9, 0x08, 0xE8, 0x95, 0x80, 0xDF, 0x94, 0xFA, 0x75, 0x8F, 0x3F, 0xA6, 0x47, 0x07, 0xA7, 0xFC, 0xF3, 0x73, 0x17, 0xBA, 0x83, 0x59, 0x3C, 0x19, 0xE6, 0x85, 0x4F, 0xA8, 0x68, 0x6B, 0x81, 0xB2, 0x71, 0x64, 0xDA, 0x8B, 0xF8, 0xEB, 0x0F, 0x4B, 0x70, 0x56, 0x9D, 0x35, 0x1E, 0x24, 0x0E, 0x5E, 0x63, 0x58, 0xD1, 0xA2, 0x25, 0x22, 0x7C, 0x3B, 0x01, 0x21, 0x78, 0x87, 0xD4, 0x00, 0x46, 0x57, 0x9F, 0xD3, 0x27, 0x52, 0x4C, 0x36, 0x02, 0xE7, 0xA0, 0xC4, 0xC8, 0x9E, 0xEA, 0xBF, 0x8A, 0xD2, 0x40, 0xC7, 0x38, 0xB5, 0xA3, 0xF7, 0xF2, 0xCE, 0xF9, 0x61, 0x15, 0xA1, 0xE0, 0xAE, 0x5D, 0xA4, 0x9B, 0x34, 0x1A, 0x55, 0xAD, 0x93, 0x32, 0x30, 0xF5, 0x8C, 0xB1, 0xE3, 0x1D, 0xF6, 0xE2, 0x2E, 0x82, 0x66, 0xCA, 0x60, 0xC0, 0x29, 0x23, 0xAB, 0x0D, 0x53, 0x4E, 0x6F, 0xD5, 0xDB, 0x37, 0x45, 0xDE, 0xFD, 0x8E, 0x2F, 0x03, 0xFF, 0x6A, 0x72, 0x6D, 0x6C, 0x5B, 0x51, 0x8D, 0x1B, 0xAF, 0x92, 0xBB, 0xDD, 0xBC, 0x7F, 0x11, 0xD9, 0x5C, 0x41, 0x1F, 0x10, 0x5A, 0xD8, 0x0A, 0xC1, 0x31, 0x88, 0xA5, 0xCD, 0x7B, 0xBD, 0x2D, 0x74, 0xD0, 0x12, 0xB8, 0xE5, 0xB4, 0xB0, 0x89, 0x69, 0x97, 0x4A, 0x0C, 0x96, 0x77, 0x7E, 0x65, 0xB9, 0xF1, 0x09, 0xC5, 0x6E, 0xC6, 0x84, 0x18, 0xF0, 0x7D, 0xEC, 0x3A, 0xDC, 0x4D, 0x20, 0x79, 0xEE, 0x5F, 0x3E, 0xD7, 0xCB, 0x39, 0x48 ] let aes_sbox_fwd_table : vector(256, bits(8)) = [ 0x63, 0x7c, 0x77, 0x7b, 0xf2, 0x6b, 0x6f, 0xc5, 0x30, 0x01, 0x67, 0x2b, 0xfe, 0xd7, 0xab, 0x76, 0xca, 0x82, 0xc9, 0x7d, 0xfa, 0x59, 0x47, 0xf0, 0xad, 0xd4, 0xa2, 0xaf, 0x9c, 0xa4, 0x72, 0xc0, 0xb7, 0xfd, 0x93, 0x26, 0x36, 0x3f, 0xf7, 0xcc, 0x34, 0xa5, 0xe5, 0xf1, 0x71, 0xd8, 0x31, 0x15, 0x04, 0xc7, 0x23, 0xc3, 0x18, 0x96, 0x05, 0x9a, 0x07, 0x12, 0x80, 0xe2, 0xeb, 0x27, 0xb2, 0x75, 0x09, 0x83, 0x2c, 0x1a, 0x1b, 0x6e, 0x5a, 0xa0, 0x52, 0x3b, 0xd6, 0xb3, 0x29, 0xe3, 0x2f, 0x84, 0x53, 0xd1, 0x00, 0xed, 0x20, 0xfc, 0xb1, 0x5b, 0x6a, 0xcb, 0xbe, 0x39, 0x4a, 0x4c, 0x58, 0xcf, 0xd0, 0xef, 0xaa, 0xfb, 0x43, 0x4d, 0x33, 0x85, 0x45, 0xf9, 0x02, 0x7f, 0x50, 0x3c, 0x9f, 0xa8, 0x51, 0xa3, 0x40, 0x8f, 0x92, 0x9d, 0x38, 0xf5, 0xbc, 0xb6, 0xda, 0x21, 0x10, 0xff, 0xf3, 0xd2, 0xcd, 0x0c, 0x13, 0xec, 0x5f, 0x97, 0x44, 0x17, 0xc4, 0xa7, 0x7e, 0x3d, 0x64, 0x5d, 0x19, 0x73, 0x60, 0x81, 0x4f, 0xdc, 0x22, 0x2a, 0x90, 0x88, 0x46, 0xee, 0xb8, 0x14, 0xde, 0x5e, 0x0b, 0xdb, 0xe0, 0x32, 0x3a, 0x0a, 0x49, 0x06, 0x24, 0x5c, 0xc2, 0xd3, 0xac, 0x62, 0x91, 0x95, 0xe4, 0x79, 0xe7, 0xc8, 0x37, 0x6d, 0x8d, 0xd5, 0x4e, 0xa9, 0x6c, 0x56, 0xf4, 0xea, 0x65, 0x7a, 0xae, 0x08, 0xba, 0x78, 0x25, 0x2e, 0x1c, 0xa6, 0xb4, 0xc6, 0xe8, 0xdd, 0x74, 0x1f, 0x4b, 0xbd, 0x8b, 0x8a, 0x70, 0x3e, 0xb5, 0x66, 0x48, 0x03, 0xf6, 0x0e, 0x61, 0x35, 0x57, 0xb9, 0x86, 0xc1, 0x1d, 0x9e, 0xe1, 0xf8, 0x98, 0x11, 0x69, 0xd9, 0x8e, 0x94, 0x9b, 0x1e, 0x87, 0xe9, 0xce, 0x55, 0x28, 0xdf, 0x8c, 0xa1, 0x89, 0x0d, 0xbf, 0xe6, 0x42, 0x68, 0x41, 0x99, 0x2d, 0x0f, 0xb0, 0x54, 0xbb, 0x16 ] let aes_sbox_inv_table : vector(256, bits(8)) = [ 0x52, 0x09, 0x6a, 0xd5, 0x30, 0x36, 0xa5, 0x38, 0xbf, 0x40, 0xa3, 0x9e, 0x81, 0xf3, 0xd7, 0xfb, 0x7c, 0xe3, 0x39, 0x82, 0x9b, 0x2f, 0xff, 0x87, 0x34, 0x8e, 0x43, 0x44, 0xc4, 0xde, 0xe9, 0xcb, 0x54, 0x7b, 0x94, 0x32, 0xa6, 0xc2, 0x23, 0x3d, 0xee, 0x4c, 0x95, 0x0b, 0x42, 0xfa, 0xc3, 0x4e, 0x08, 0x2e, 0xa1, 0x66, 0x28, 0xd9, 0x24, 0xb2, 0x76, 0x5b, 0xa2, 0x49, 0x6d, 0x8b, 0xd1, 0x25, 0x72, 0xf8, 0xf6, 0x64, 0x86, 0x68, 0x98, 0x16, 0xd4, 0xa4, 0x5c, 0xcc, 0x5d, 0x65, 0xb6, 0x92, 0x6c, 0x70, 0x48, 0x50, 0xfd, 0xed, 0xb9, 0xda, 0x5e, 0x15, 0x46, 0x57, 0xa7, 0x8d, 0x9d, 0x84, 0x90, 0xd8, 0xab, 0x00, 0x8c, 0xbc, 0xd3, 0x0a, 0xf7, 0xe4, 0x58, 0x05, 0xb8, 0xb3, 0x45, 0x06, 0xd0, 0x2c, 0x1e, 0x8f, 0xca, 0x3f, 0x0f, 0x02, 0xc1, 0xaf, 0xbd, 0x03, 0x01, 0x13, 0x8a, 0x6b, 0x3a, 0x91, 0x11, 0x41, 0x4f, 0x67, 0xdc, 0xea, 0x97, 0xf2, 0xcf, 0xce, 0xf0, 0xb4, 0xe6, 0x73, 0x96, 0xac, 0x74, 0x22, 0xe7, 0xad, 0x35, 0x85, 0xe2, 0xf9, 0x37, 0xe8, 0x1c, 0x75, 0xdf, 0x6e, 0x47, 0xf1, 0x1a, 0x71, 0x1d, 0x29, 0xc5, 0x89, 0x6f, 0xb7, 0x62, 0x0e, 0xaa, 0x18, 0xbe, 0x1b, 0xfc, 0x56, 0x3e, 0x4b, 0xc6, 0xd2, 0x79, 0x20, 0x9a, 0xdb, 0xc0, 0xfe, 0x78, 0xcd, 0x5a, 0xf4, 0x1f, 0xdd, 0xa8, 0x33, 0x88, 0x07, 0xc7, 0x31, 0xb1, 0x12, 0x10, 0x59, 0x27, 0x80, 0xec, 0x5f, 0x60, 0x51, 0x7f, 0xa9, 0x19, 0xb5, 0x4a, 0x0d, 0x2d, 0xe5, 0x7a, 0x9f, 0x93, 0xc9, 0x9c, 0xef, 0xa0, 0xe0, 0x3b, 0x4d, 0xae, 0x2a, 0xf5, 0xb0, 0xc8, 0xeb, 0xbb, 0x3c, 0x83, 0x53, 0x99, 0x61, 0x17, 0x2b, 0x04, 0x7e, 0xba, 0x77, 0xd6, 0x26, 0xe1, 0x69, 0x14, 0x63, 0x55, 0x21, 0x0c, 0x7d ] /* Lookup function - takes an index and a table, and retrieves the * x'th element of that table. Note that the Sail vector literals * start at index 255, and go down to 0. */ val sbox_lookup : (bits(8), vector(256, bits(8))) -> bits(8) function sbox_lookup(x, table) = { table[255 - unsigned(x)] } /* Easy function to perform a forward AES SBox operation on 1 byte. */ val aes_sbox_fwd : bits(8) -> bits(8) function aes_sbox_fwd(x) = sbox_lookup(x, aes_sbox_fwd_table) /* Easy function to perform an inverse AES SBox operation on 1 byte. */ val aes_sbox_inv : bits(8) -> bits(8) function aes_sbox_inv(x) = sbox_lookup(x, aes_sbox_inv_table) /* AES SubWord function used in the key expansion * - Applies the forward sbox to each byte in the input word. */ val aes_subword_fwd : bits(32) -> bits(32) function aes_subword_fwd(x) = { aes_sbox_fwd(x[31..24]) @ aes_sbox_fwd(x[23..16]) @ aes_sbox_fwd(x[15.. 8]) @ aes_sbox_fwd(x[ 7.. 0]) } /* AES Inverse SubWord function. * - Applies the inverse sbox to each byte in the input word. */ val aes_subword_inv : bits(32) -> bits(32) function aes_subword_inv(x) = { aes_sbox_inv(x[31..24]) @ aes_sbox_inv(x[23..16]) @ aes_sbox_inv(x[15.. 8]) @ aes_sbox_inv(x[ 7.. 0]) } /* Easy function to perform an SM4 SBox operation on 1 byte. */ val sm4_sbox : bits(8) -> bits(8) function sm4_sbox(x) = sbox_lookup(x, sm4_sbox_table) val aes_get_column : (bits(128), nat) -> bits(32) function aes_get_column(state,c) = (state >> (to_bits(7, 32 * c)))[31..0] /* 64-bit to 64-bit function which applies the AES forward sbox to each byte * in a 64-bit word. */ val aes_apply_fwd_sbox_to_each_byte : bits(64) -> bits(64) function aes_apply_fwd_sbox_to_each_byte(x) = { aes_sbox_fwd(x[63..56]) @ aes_sbox_fwd(x[55..48]) @ aes_sbox_fwd(x[47..40]) @ aes_sbox_fwd(x[39..32]) @ aes_sbox_fwd(x[31..24]) @ aes_sbox_fwd(x[23..16]) @ aes_sbox_fwd(x[15.. 8]) @ aes_sbox_fwd(x[ 7.. 0]) } /* 64-bit to 64-bit function which applies the AES inverse sbox to each byte * in a 64-bit word. */ val aes_apply_inv_sbox_to_each_byte : bits(64) -> bits(64) function aes_apply_inv_sbox_to_each_byte(x) = { aes_sbox_inv(x[63..56]) @ aes_sbox_inv(x[55..48]) @ aes_sbox_inv(x[47..40]) @ aes_sbox_inv(x[39..32]) @ aes_sbox_inv(x[31..24]) @ aes_sbox_inv(x[23..16]) @ aes_sbox_inv(x[15.. 8]) @ aes_sbox_inv(x[ 7.. 0]) } /* * AES full-round transformation functions. */ val getbyte : (bits(64), int) -> bits(8) function getbyte(x, i) = (x >> to_bits(6, i * 8))[7..0] val aes_rv64_shiftrows_fwd : (bits(64), bits(64)) -> bits(64) function aes_rv64_shiftrows_fwd(rs2, rs1) = { getbyte(rs1, 3) @ getbyte(rs2, 6) @ getbyte(rs2, 1) @ getbyte(rs1, 4) @ getbyte(rs2, 7) @ getbyte(rs2, 2) @ getbyte(rs1, 5) @ getbyte(rs1, 0) } val aes_rv64_shiftrows_inv : (bits(64), bits(64)) -> bits(64) function aes_rv64_shiftrows_inv(rs2, rs1) = { getbyte(rs2, 3) @ getbyte(rs2, 6) @ getbyte(rs1, 1) @ getbyte(rs1, 4) @ getbyte(rs1, 7) @ getbyte(rs2, 2) @ getbyte(rs2, 5) @ getbyte(rs1, 0) } /* 128-bit to 128-bit implementation of the forward AES ShiftRows transform. * Byte 0 of state is input column 0, bits 7..0. * Byte 5 of state is input column 1, bits 15..8. */ val aes_shift_rows_fwd : bits(128) -> bits(128) function aes_shift_rows_fwd(x) = { let ic3 : bits(32) = aes_get_column(x, 3); let ic2 : bits(32) = aes_get_column(x, 2); let ic1 : bits(32) = aes_get_column(x, 1); let ic0 : bits(32) = aes_get_column(x, 0); let oc0 : bits(32) = ic0[31..24] @ ic1[23..16] @ ic2[15.. 8] @ ic3[ 7.. 0]; let oc1 : bits(32) = ic1[31..24] @ ic2[23..16] @ ic3[15.. 8] @ ic0[ 7.. 0]; let oc2 : bits(32) = ic2[31..24] @ ic3[23..16] @ ic0[15.. 8] @ ic1[ 7.. 0]; let oc3 : bits(32) = ic3[31..24] @ ic0[23..16] @ ic1[15.. 8] @ ic2[ 7.. 0]; (oc3 @ oc2 @ oc1 @ oc0) /* Return value */ } /* 128-bit to 128-bit implementation of the inverse AES ShiftRows transform. * Byte 0 of state is input column 0, bits 7..0. * Byte 5 of state is input column 1, bits 15..8. */ val aes_shift_rows_inv : bits(128) -> bits(128) function aes_shift_rows_inv(x) = { let ic3 : bits(32) = aes_get_column(x, 3); /* In column 3 */ let ic2 : bits(32) = aes_get_column(x, 2); let ic1 : bits(32) = aes_get_column(x, 1); let ic0 : bits(32) = aes_get_column(x, 0); let oc0 : bits(32) = ic0[31..24] @ ic3[23..16] @ ic2[15.. 8] @ ic1[ 7.. 0]; let oc1 : bits(32) = ic1[31..24] @ ic0[23..16] @ ic3[15.. 8] @ ic2[ 7.. 0]; let oc2 : bits(32) = ic2[31..24] @ ic1[23..16] @ ic0[15.. 8] @ ic3[ 7.. 0]; let oc3 : bits(32) = ic3[31..24] @ ic2[23..16] @ ic1[15.. 8] @ ic0[ 7.. 0]; (oc3 @ oc2 @ oc1 @ oc0) /* Return value */ } /* Applies the forward sub-bytes step of AES to a 128-bit vector * representation of its state. */ val aes_subbytes_fwd : bits(128) -> bits(128) function aes_subbytes_fwd(x) = { let oc0 : bits(32) = aes_subword_fwd(aes_get_column(x, 0)); let oc1 : bits(32) = aes_subword_fwd(aes_get_column(x, 1)); let oc2 : bits(32) = aes_subword_fwd(aes_get_column(x, 2)); let oc3 : bits(32) = aes_subword_fwd(aes_get_column(x, 3)); (oc3 @ oc2 @ oc1 @ oc0) /* Return value */ } /* Applies the inverse sub-bytes step of AES to a 128-bit vector * representation of its state. */ val aes_subbytes_inv : bits(128) -> bits(128) function aes_subbytes_inv(x) = { let oc0 : bits(32) = aes_subword_inv(aes_get_column(x, 0)); let oc1 : bits(32) = aes_subword_inv(aes_get_column(x, 1)); let oc2 : bits(32) = aes_subword_inv(aes_get_column(x, 2)); let oc3 : bits(32) = aes_subword_inv(aes_get_column(x, 3)); (oc3 @ oc2 @ oc1 @ oc0) /* Return value */ } /* Applies the forward MixColumns step of AES to a 128-bit vector * representation of its state. */ val aes_mixcolumns_fwd : bits(128) -> bits(128) function aes_mixcolumns_fwd(x) = { let oc0 : bits(32) = aes_mixcolumn_fwd(aes_get_column(x, 0)); let oc1 : bits(32) = aes_mixcolumn_fwd(aes_get_column(x, 1)); let oc2 : bits(32) = aes_mixcolumn_fwd(aes_get_column(x, 2)); let oc3 : bits(32) = aes_mixcolumn_fwd(aes_get_column(x, 3)); (oc3 @ oc2 @ oc1 @ oc0) /* Return value */ } /* Applies the inverse MixColumns step of AES to a 128-bit vector * representation of its state. */ val aes_mixcolumns_inv : bits(128) -> bits(128) function aes_mixcolumns_inv(x) = { let oc0 : bits(32) = aes_mixcolumn_inv(aes_get_column(x, 0)); let oc1 : bits(32) = aes_mixcolumn_inv(aes_get_column(x, 1)); let oc2 : bits(32) = aes_mixcolumn_inv(aes_get_column(x, 2)); let oc3 : bits(32) = aes_mixcolumn_inv(aes_get_column(x, 3)); (oc3 @ oc2 @ oc1 @ oc0) /* Return value */ } ``` --- [1](#%5Ffootnoteref%5F1). 34.1. Control-flow Integrity (CFI) ==================== ## [](#34-1-control-flow-integrity-cfi)34.1\. Control-flow Integrity (CFI) Control-flow Integrity (CFI) capabilities help defend against Return-Oriented Programming (ROP) and Call/Jump-Oriented Programming (COP/JOP) style control-flow subversion attacks. These attack methodologies use code sequences in authorized modules, with at least one instruction in the sequence being a control transfer instruction that depends on attacker-controlled data either in the return stack or in memory used to obtain the target address for a call or jump. Attackers stitch these sequences together by diverting the control flow instructions (e.g., `JALR`, `C.JR`, `C.JALR`), from their original target address to a new target via modification in the return stack or in the memory used to obtain the jump/call target address. RV32/RV64 provides two types of control transfer instructions - unconditional jumps and conditional branches. Conditional branches encode an offset in the immediate field of the instruction and are thus direct branches that are not susceptible to control-flow subversion. Unconditional direct jumps using `JAL`transfer control to a target that is in a +/- 1 MiB range from the current `pc`. Unconditional indirect jumps using the `JALR` obtain their branch target by adding the sign extended 12-bit immediate encoded in the instruction to the`rs1` register. The RV32I/RV64I does not have a dedicated instruction for calling a procedure or returning from a procedure. A `JAL` or `JALR` may be used to perform a procedure call and `JALR` to return from a procedure. The RISC-V ABI however defines the convention that a `JAL`/`JALR` where `rd` (i.e. the link register) is `x1` or`x5` is a procedure call, and a `JALR` where `rs1` is the conventional link register (i.e. `x1` or `x5`) is a return from procedure. The architecture allows for using these hints and conventions to support return address prediction (See [Return-address stack prediction hints](rv32.html#rashints)). The RVC standard extension for compressed instructions provides unconditional jump and conditional branch instructions. The `C.J` and `C.JAL` instructions encode an offset in the immediate field of the instruction and thus are not susceptible to control-flow subversion. The `C.JR` and `C.JALR` RVC instructions perform an unconditional control transfer to the address in register `rs1`. The`C.JALR` additionally writes the address of the instruction following the jump (`pc+2`) to the link register `x1` and is a procedure call. The `C.JR` is a return from procedure if `rs1` is a conventional link register (i.e. `x1` or`x5`); else it is an indirect jump. The term _call_ is used to refer to a `JAL` or `JALR` instruction with a link register as destination, i.e., _rd_≠`x0`. Conventionally, the link register is`x1` or `x5`. A _call_ using `JAL` or `C.JAL` is termed a direct call. A`C.JALR` expands to `JALR x1, 0(rs1)` and is a _call_. A _call_ using `JALR` or`C.JALR` is termed an _indirect-call_. The term _return_ is used to refer to a `JALR` instruction with _rd_\=`x0` and with _rs1_\=`x1` or _rs1_\=`x5`. A `C.JR` instruction expands to`JALR x0, 0(rs1)` and is a _return_ if _rs1_\=`x1` or _rs1_\=`x5`. The term _indirect-jump_ is used to refer to a `JALR` instruction with _rd_\=`x0`and where the _rs1_ is not `x1` or `x5` (i.e., not a return). A `C.JR`instruction where _rs1_ is not `x1` or `x5` (i.e., not a return) is an_indirect-jump_. The Zicfiss and Zicfilp extensions build on these conventions and hints and provide backward-edge and forward-edge control flow integrity respectively. The Unprivileged ISA for Zicfilp extension is specified in [34.1.1\. Landing Pad (Zicfilp)](#unpriv-forward)and for the Unprivileged ISA for Zicfiss extension is specified in[34.1.2\. Shadow Stack (Zicfiss)](#unpriv-backward). The Privileged ISA for these extensions is specified in the Privileged ISA specification. ### [](#unpriv-forward)34.1.1\. Landing Pad (Zicfilp) To enforce forward-edge control-flow integrity, the Zicfilp extension introduces a landing pad (`LPAD`) instruction. The `LPAD` instruction must be placed at the program locations that are valid targets of indirect jumps or calls. The `LPAD`instruction (See [34.1.1.2\. Landing Pad Instruction](#LP%5FINST)) is encoded using the `AUIPC` major opcode with_rd_\=`x0`. Compilers emit a landing pad instruction as the first instruction of an address-taken function, as well as at any indirect jump targets. A landing pad instruction is not required in functions that are only reached using a direct call or direct jump. The landing pad is designed to provide integrity to control transfers performed using indirect calls and jumps, and this is referred to as forward-edge protection. When the Zicfilp is active, the hart tracks an expected landing pad (`ELP`) state that is updated by an _indirect\_call_ or _indirect\_jump_ to require a landing pad instruction at the target of the branch. If the instruction at the target is not a landing pad, then a software-check exception is raised. A landing pad may be optionally associated with a 20-bit label. With labeling enabled, the number of landing pads that can be reached from an indirect call or jump sites can be defined using programming language-based policies. Labeling of the landing pads enables software to achieve greater precision in pairing up indirect call/jump sites with valid targets. When labeling of landing pads is used, indirect call or indirect jump site can specify the expected label of the landing pad and thereby constrain the set of landing pads that may be reached from each indirect call or indirect jump site in the program. In the simplest form, a program can be built with a single label value to implement a coarse-grained version of forward-edge control-flow integrity. By constraining gadgets to be preceded by a landing pad instruction that marks the start of indirect callable functions, the program can significantly reduce the available gadget space. A second form of label generation may generate a signature, such as a MAC, using the prototype of the function. Programs that use this approach would further constrain the gadgets accessible from a call site to only indirectly callable functions that match the prototype of the called functions. Another approach to label generation involves analyzing the control-flow-graph (CFG) of the program, which can lead to even more stringent constraints on the set of reachable gadgets. Such programs may further use multiple labels per function, which means that if a function is called from two or more call sites, the functions can be labeled as being reachable from each of the call sites. For instance, consider two call sites A and B, where A calls the functions X and Y, and B calls the functions Y and Z. In a single label scheme, functions X, Y, and Z would need to be assigned the same label so that both call sites A and B can invoke the common function Y. This scheme would allow call site A to also call function Z and call site B to also call function X. However, if function Y was assigned two labels - one corresponding to call site A and the other to call site B, then Y can be invoked by both call sites, but X can only be invoked by call site A and Z can only be invoked by call site B. To support multiple labels, the compiler could generate a call-site-specific entry point for shared functions, with each entry point having its own landing pad instruction followed by a direct branch to the start of the function. This would allow the function to be labeled with multiple labels, each corresponding to a specific call site. A portion of the label space may be dedicated to labeled landing pads that are only valid targets of an indirect jump (and not an indirect call). The `LPAD` instruction uses the code points defined as HINTs for the `AUIPC`opcode. When Zicfilp is not active at a privilege level or when the extension is not implemented, the landing pad instruction executes as a no-op. A program that is built with `LPAD` instructions can thus continue to operate correctly, but without forward-edge control-flow integrity, on processors that do not support the Zicfilp extension or if the Zicfilp extension is not active. Compilers and linkers should provide an attribute flag to indicate if the program has been compiled with the Zicfilp extension and use that to determine if the Zicfilp extension should be activated. The dynamic loader should activate the use of Zicfilp extension for an application only if all executables (the application and the dependent dynamically linked libraries) used by that application use the Zicfilp extension. When Zicfilp extension is not active or not implemented, the hart does not require landing pad instructions at the targets of indirect calls/jumps, and the landing instructions revert to being no-ops. This allows a program compiled with landing pad instructions to operate correctly but without forward-edge control-flow integrity. The Zicfilp extensions may be activated for use individually and independently for each privilege mode. The Zicfilp extension depends on the Zicsr extension. #### [](#34-1-1-1-landing-pad-enforcement)34.1.1.1\. Landing Pad Enforcement To enforce that the target of an indirect call or indirect jump must be a valid landing pad instruction, the hart maintains an expected landing pad (`ELP`) state to determine if a landing pad instruction is required at the target of an indirect call or an indirect jump. The `ELP` state can be one of: * 0 - `NO_LP_EXPECTED` * 1 - `LP_EXPECTED` The `ELP` state is initialized to `NO_LP_EXPECTED` by the hart upon reset. The Zicfilp extension, when enabled, determines if an indirect call or an indirect jump must land on a landing pad, as specified in [Landing pad expected determination](#IND%5FCALL%5FJMP). If`is_lp_expected` is 1, then the hart updates the `ELP` to `LP_EXPECTED`. Landing pad expected determination is_lp_expected = ( (JALR || C.JR || C.JALR) && (rs1 != x1) && (rs1 != x5) && (rs1 != x7) ) ? 1 : 0; An indirect branch using `JALR`, `C.JALR`, or `C.JR` with `rs1` as `x7` is termed a software guarded branch. Such branches do not need to land on a`LPAD` instruction and thus do not set `ELP` to `LP_EXPECTED`. | | When the register source is a link register and the register destination isx0, then it’s a return from a procedure and does not require a landing pad at the target. When the register source and register destination are both link registers, then it is a semantically-direct-call. For example, the call offsetpseudoinstruction may expand to a two instruction sequence composed of alui ra, imm20 or a auipc ra, imm20 instruction followed by ajalr ra, imm12(ra) instruction where ra is the link register (either x1 orx5). Since the address of the procedure was not explicitly taken and the computed address is not obtained from mutable memory, such semantically-direct calls do not require a landing pad to be placed at the target. Compilers and JITers must use the semantically-direct calls only if the rs1 was computed as a PC-relative or an absolute offset to the symbol. The tail offset pseudoinstruction used to tail call a far-away procedure may also be expanded to a two instruction sequence composed of a lui x7, imm20 orauipc x7, imm20 followed by a jalr x0, x7. Since the address of the procedure was not explicitly taken and the computed address is not obtained from mutable memory, such semantically-direct tail-calls do not require a landing pad to be placed at the target. Software guarded branches may also be used by compilers to generate code for constructs like switch-cases. When using the software guarded branches, the compiler is required to ensure it has full control on the possible jump targets (e.g., by obtaining the targets from a read-only table in memory and performing bounds checking on the index into the table, etc.). | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The landing pad may be labeled. Zicfilp extension designates the register `x7`for use as the landing pad label register. To support labeled landing pads, the indirect call/jump sites establish an expected landing pad label (e.g., using the `LUI` instruction) in the bits 31:12 of the `x7` register. The `LPAD`instruction is encoded with a 20-bit immediate value called the landing-pad-label (`LPL`) that is matched to the expected landing pad label. When `LPL` is encoded as zero, the `LPAD` instruction does not perform the label check and in programs built with this single label mode of operation the indirect call/jump sites do not need to establish an expected landing pad label value in `x7`. When `ELP` is set to `LP_EXPECTED`, if the next instruction in the instruction stream is not 4-byte aligned, or is not `LPAD`, or if the landing pad label encoded in `LPAD` is not zero and does not match the expected landing pad label in bits 31:12 of the `x7` register, then a software-check exception (cause=18) with `_x_tval` set to "landing pad fault (code=2)" is raised else the `ELP` is updated to `NO_LP_EXPECTED`. | | The tracking of ELP and the requirement for a landing pad instruction at the target of indirect call and jump enables a processor implementation to significantly reduce or to prevent speculation to non-landing-pad instructions. Constraining speculation using this technique, greatly reduces the gadget space and increases the difficulty of using techniques such as branch-target-injection, also known as Spectre variant 2, which use speculative execution to leak data through side channels. The LPAD requires a 4-byte alignment to address the concatenation of two instructions A and B accidentally forming an unintended landing pad in the program. For example, consider a 32-bit instruction where the bytes 3 and 2 have a pattern of ?017h (for example, the immediate fields of a LUI, AUIPC, or a JAL instruction), followed by a 16-bit or a 32-bit instruction. When patterns that can accidentally form a valid landing pad are detected, the assembler or linker can force instruction A to be aligned to a 4-byte boundary to force the unintended LPAD pattern to become misaligned, and thus not a valid landing pad, or may use an alternate register allocation to prevent the accidental landing pad. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#LP%5FINST)34.1.1.2\. Landing Pad Instruction When Zicfilp is enabled, `LPAD` is the only instruction allowed to execute when the `ELP` state is `LP_EXPECTED`. If Zicfilp is not enabled then the instruction is a no-op. If Zicfilp is enabled, the `LPAD` instruction causes a software-check exception with `_x_tval` set to "landing pad fault (code=2)" if any of the following conditions are true: * The `pc` is not 4-byte aligned and `ELP` is `LP_EXPECTED`. * The `ELP` is `LP_EXPECTED` and the `LPL` is not zero and the `LPL` does not match the expected landing pad label in bits 31:12 of the `x7` register. If a software-check exception is not caused then the `ELP` is updated to`NO_LP_EXPECTED`. ![svg](_images/svg-ded55101008cc4964e182bfa901211874fffad41.svg) The operation of the `LPAD` instruction is as follows: `LPAD` operation if (xLPE == 1 && ELP == LP_EXPECTED) // If PC not 4-byte aligned then software-check exception if pc[1:0] != 0 raise software-check exception // If landing pad label not matched -> software-check exception else if (inst.LPL != x7[31:12] && inst.LPL != 0) raise software-check exception else ELP = NO_LP_EXPECTED else no-op endif \# <<< ### [](#unpriv-backward)34.1.2\. Shadow Stack (Zicfiss) The Zicfiss extension introduces a shadow stack to enforce backward-edge control-flow integrity. A shadow stack is a second stack used to store a shadow copy of the return address in the link register if it needs to be spilled. The shadow stack is designed to provide integrity to control transfers performed using a _return_, where the return may be from a procedure invoked using an indirect call or a direct call, and this is referred to as backward-edge protection. A program using backward-edge control-flow integrity has two stacks: a regular stack and a shadow stack. The shadow stack is used to spill the link register, if required, by non-leaf functions. An additional register, shadow-stack-pointer (`ssp`), is introduced in the architecture to hold the address of the top of the active shadow stack. The shadow stack, similar to the regular stack, grows downwards, from higher addresses to lower addresses. Each entry on the shadow stack is `XLEN`wide and holds the link register value. The `ssp` points to the top of the shadow stack, which is the address of the last element stored on the shadow stack. The shadow stack is architecturally protected from inadvertent corruptions and modifications, as detailed in the Privileged specification. The Zicfiss extension provides instructions to store and load the link register to/from the shadow stack and to check the integrity of the return address. The extension provides instructions to support common stack maintenance operations such as stack unwinding and stack switching. When Zicfiss is enabled, each function that needs to spill the link register, typically non-leaf functions, store the link register value to the regular stack and a shadow copy of the link register value to the shadow stack when the function is entered (the prologue). When such a function returns (the epilogue), the function loads the link register from the regular stack and the shadow copy of the link register from the shadow stack. Then, the link register value from the regular stack and the shadow link register value from the shadow stack are compared. A mismatch of the two values is indicative of a subversion of the return address control variable and causes a software-check exception. The Zicfiss instructions, except `SSAMOSWAP.W/D`, are encoded using a subset of May-Be-Operation instructions defined by the Zimop and Zcmop extensions.This subset of instructions revert to their Zimop/Zcmop defined behavior when the Zicfiss extension is not implemented or if the extension has not been activated. A program that is built with Zicfiss instructions can thus continue to operate correctly, but without backward-edge control-flow integrity, on processors that do not support the Zicfiss extension or if the Zicfiss extension is not active. The Zicfiss extension may be activated for use individually and independently for each privilege mode. Compilers should flag each object file (for example, using flags in the ELF attributes) to indicate if the object file has been compiled with the Zicfiss instructions. The linker should flag (for example, using flags in the ELF attributes) the binary/executable generated by linking objects as being compiled with the Zicfiss instructions only if all the object files that are linked have the same Zicfiss attributes. The dynamic loader should activate the use of Zicfiss extension for an application only if all executables (the application and the dependent dynamically-linked libraries) used by that application use the Zicfiss extension. An application that has the Zicfiss extension active may request the dynamic loader at runtime to load a new dynamic shared object (using dlopen() for example). If the requested object does not have the Zicfiss attribute then the dynamic loader, based on its policy (e.g., established by the operating system or the administrator) configuration, could either deny the request or deactivate the Zicfiss extension for the application. It is strongly recommended that the policy enforces a strict security posture and denies the request. The Zicfiss extension depends on the Zicsr, Zimop and Zaamo extensions. Furthermore, if the Zcmop extension is implemented, the Zicfiss extension also provides the`C.SSPUSH` and `C.SSPOPCHK` instructions. Moreover, use of Zicfiss in U-mode requires S-mode to be implemented. Use of Zicfiss in M-mode is not supported. #### [](#34-1-2-1-zicfiss-instructions-summary)34.1.2.1\. Zicfiss Instructions Summary The Zicfiss extension introduces the following instructions: * Push to the shadow stack (See [34.1.2.4\. Push to the Shadow Stack](#SS%5FPUSH)) * `SSPUSH x1` and `SSPUSH x5` \- encoded using `MOP.RR.7` * `C.SSPUSH x1` \- encoded using `C.MOP.1` * Pop from the shadow stack (See [34.1.2.5\. Pop from the Shadow Stack](#SS%5FPOP)) * `SSPOPCHK x1` and `SSPOPCHK x5` \- encoded using `MOP.R.28` * `C.SSPOPCHK x5` \- encoded using `C.MOP.5` * Read the value of `ssp` into a register (See [34.1.2.6\. Read ssp into a Register](#SSP%5FREAD)) * `SSRDP` \- encoded using `MOP.R.28` * Perform an atomic swap from a shadow stack location (See [34.1.2.7\. Atomic Swap from a Shadow Stack Location](#SSAMOSWAP)) * `SSAMOSWAP.W` and `SSAMOSWAP.D` Zicfiss does not use all encodings of `MOP.RR.7` or `MOP.R.28`. When a`MOP.RR.7` or `MOP.R.28` encoding is not used by the Zicfiss extension, the corresponding instruction adheres to its Zimop-defined behavior, unless redefined by another extension. #### [](#34-1-2-2-shadow-stack-pointer-ssp)34.1.2.2\. Shadow Stack Pointer (`ssp`) The `ssp` CSR is an unprivileged read-write (URW) CSR that reads and writes`XLEN` low order bits of the shadow stack pointer (`ssp`). The CSR address is 0x011. There is no high CSR defined as the `ssp` is always as wide as the `XLEN`of the current privilege mode. The bits 1:0 of `ssp` are read-only zero. If the UXLEN or SXLEN may never be 32, then the bit 2 is also read-only zero. #### [](#34-1-2-3-zicfiss-instructions)34.1.2.3\. Zicfiss Instructions #### [](#SS%5FPUSH)34.1.2.4\. Push to the Shadow Stack A shadow stack push operation is defined as decrement of the `ssp` by `XLEN/8`followed by a store of the value in the link register to memory at the new top of the shadow stack. ![svg](_images/svg-c2ca1bde05833cdb524b1cdf80ba167ca9b4626d.svg) ![svg](_images/svg-4e45669c7dfdc8a4516ca8ca83399595feb7d24e.svg) Only `x1` and `x5` registers are supported as `rs2` for `SSPUSH`. Zicfiss provides a 16-bit version of the `SSPUSH x1` instruction using the Zcmop defined `C.MOP.1` encoding. The `C.SSPUSH x1` expands to `SSPUSH x1`. The `SSPUSH` instruction and its compressed form `C.SSPUSH` can be used to push a link register on the shadow stack. The `SSPUSH` and `C.SSPUSH` instructions perform a store identically to the existing store instructions, with the difference that the base is implicitly `ssp` and the width is implicitly `XLEN`. The operation of the `SSPUSH` and `C.SSPUSH` instructions is as follows: `SSPUSH` and `C.SSPUSH` operation if (xSSE == 1) mem[ssp - (XLEN/8)] = X(src) # Store src value to ssp - XLEN/8 ssp = ssp - (XLEN/8) # decrement ssp by XLEN/8 endif The `ssp` is decremented by `SSPUSH` and `C.SSPUSH` only if the store to the shadow stack completes successfully. #### [](#SS%5FPOP)34.1.2.5\. Pop from the Shadow Stack A shadow stack pop operation is defined as an `XLEN` wide read from the current top of the shadow stack followed by an increment of the `ssp` by`XLEN/8`. ![svg](_images/svg-3f200ab6da34be7c8a974a335fb2584dd08f427b.svg) ![svg](_images/svg-054d91ebf0633808e90ae99e50d4f0775195513b.svg) Only `x1` and `x5` registers are supported as `rs1` for `SSPOPCHK`. Zicfiss provides a 16-bit version of the `SSPOPCHK x5` using the Zcmop defined `C.MOP.5`encoding. The `C.SSPOPCHK x5` expands to `SSPOPCHK x5`. Programs with a shadow stack push the return address onto the regular stack as well as the shadow stack in the prologue of non-leaf functions. When returning from these non-leaf functions, such programs pop the link register from the regular stack and pop a shadow copy of the link register from the shadow stack. The two values are then compared. If the values do not match, it is indicative of a corruption of the return address variable on the regular stack. The `SSPOPCHK` instruction, and its compressed form `C.SSPOPCHK`, can be used to pop the shadow return address value from the shadow stack and check that the value matches the contents of the link register, and if not cause a software-check exception with `_x_tval` set to "shadow stack fault (code=3)". While any register may be used as link register, conventionally the `x1` or `x5`registers are used. The shadow stack instructions are designed to be most efficient when the `x1` and `x5` registers are used as the link register. | | Return-address prediction stacks are a common feature of high-performance instruction-fetch units, but they require accurate detection of instructions used for procedure calls and returns to be effective. For RISC-V, hints as to the instructions' usage are encoded implicitly via the register numbers used. The return-address stack (RAS) actions to pop and/or push onto the RAS are specified in [Return-address stack prediction hints](rv32.html#rashints). Using x1 or x5 as the link register allows a program to benefit from the return-address prediction stacks. Additionally, since the shadow stack instructions are designed around the use of x1 or x5 as the link register, using any other register as a link register would incur the cost of additional register movements. Compilers, when generating code with backward-edge CFI, must protect the link register, e.g., x1 and/or x5, from arbitrary modification by not emitting unsafe code sequences. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Storing the return address on both stacks preserves the call stack layout and the ABI, while also allowing for the detection of corruption of the return address on the regular stack. The prologue and epilogue of a non-leaf function that uses shadow stacks is as follows: function\_entry: addi sp,sp,-8 # push link register x1 sd x1,(sp) # on regular stack sspush x1 # push link register x1 on shadow stack : ld x1,(sp) # pop link register x1 from regular stack addi sp,sp,8 sspopchk x1 # fault if x1 not equal to shadow \# return address ret This example illustrates the use of x1 register as the link register. Alternatively, the x5 register may also be used as the link register. A leaf function, a function that does not itself make function calls, does not need to spill the link register. Consequently, the return value may be held in the link register itself for the duration of the leaf function’s execution. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `C.SSPOPCHK`, and `SSPOPCHK` instructions perform a load identically to the existing load instructions, with the difference that the base is implicitly`ssp` and the width is implicitly `XLEN`. The operation of the `SSPOPCHK` and `C.SSPOPCHK` instructions is as follows: `SSPOPCHK` and `C.SSPOPCHK` operation if (xSSE == 1) temp = mem[ssp] # Load temp from address in ssp and if temp != X(src) # Compare temp to value in src and # cause a software-check exception # if they are not bitwise equal. # Only x1 and x5 may be used as src raise software-check exception else ssp = ssp + (XLEN/8) # increment ssp by XLEN/8. endif endif If the value loaded from the address in `ssp` does not match the value in `rs1`, a software-check exception (cause=18) is raised with `_x_tval` set to "shadow stack fault (code=3)". The software-check exception caused by `SSPOPCHK`/`C.SSPOPCHK` is lower in priority than a load/store/AMO access-fault exception. The `ssp` is incremented by `SSPOPCHK` and `C.SSPOPCHK` only if the load from the shadow stack completes successfully and no software-check exception is raised. | | The use of the compressed instruction C.SSPUSH x1 to push on the shadow stack is most efficient when the ABI uses x1 as the link register, as the link register may then be pushed without needing a register-to-register move in the function prologue. To use the compressed instruction C.SSPOPCHK x5, the function should pop the return address from regular stack into the alternate link register x5 and use the C.SSPOPCHK x5 to compare the return address to the shadow copy stored on the shadow stack. The function then uses C.JR x5 to jump to the return address. function\_entry: c.addi sp,sp,-8 # push link register x1 c.sd x1,(sp) # on regular stack c.sspush x1 # push link register x1 on shadow stack : c.ld x5,(sp) # pop link register x5 from regular stack c.addi sp,sp,8 c.sspopchk x5 # fault if x5 not equal to shadow return address c.jr x5 | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Store-to-load forwarding is a common technique employed by high-performance processor implementations. Zicfiss implementations may prevent forwarding from a non-shadow-stack store to the SSPOPCHK or the C.SSPOPCHK instructions. A non-shadow-stack store causes a fault if done to a page mapped as a shadow stack. However, such determination may be delayed till the PTE has been examined and thus may be used to transiently forward the data from such stores toSSPOPCHK or to C.SSPOPCHK. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#SSP%5FREAD)34.1.2.6\. Read `ssp` into a Register The `SSRDP` instruction is provided to move the contents of `ssp` to a destination register. ![svg](_images/svg-e8475c77571f1d17df824c173f7a8b4e63d20cbb.svg) Encoding _rd_ as `x0` is not supported for `SSRDP`. The operation of the `SSRDP` instructions is as follows: `SSRDP` operation if (xSSE == 1) X(dst) = ssp else X(dst) = 0 endif | | The property of Zimop writing 0 to the rd when the extension using Zimop is not implemented or not active may be used by to determine if Zicfiss extension is active. For example, functions that unwind shadow stacks may skip over the unwind actions by dynamically detecting if the Zicfiss extension is active. An example sequence such as the following may be used: ssrdp t0 # mv ssp to t0 beqz t0, zicfiss\_not\_active # zero is not a valid shadow stack \# pointer by convention \# Zicfiss is active : : zicfiss\_not\_active: To assist with the use of such code sequences, operating systems and runtimes must not locate shadow stacks at address 0. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | A common operation performed on stacks is to unwind them to support constructs like setjmp/longjmp, C++ exception handling, etc. A program that uses shadow stacks must unwind the shadow stack in addition to the stack used to store data. The unwind function must verify that it does not accidentally unwind past the bounds of the shadow stack. Shadow stacks are expected to be bounded on each end using guard pages. A guard page for a stack is a page that is not accessible by the process that owns the stack. To detect if the unwind occurs past the bounds of the shadow stack, the unwind may be done in maximal increments of 4 KiB, testing whether the ssp is still pointing to a shadow stack page or has unwound into the guard page. The following examples illustrate the use of shadow stack instructions to unwind a shadow stack. This example assumes that thesetjmp function itself does not push on to the shadow stack (being a leaf function, it is not required to). setjmp() { : : // read and save the shadow stack pointer to jmp\_buf asm("ssrdp %0" : "=r"(cur\_ssp):); jmp\_buf->saved\_ssp = cur\_ssp; : : } longjmp() { : // Read current shadow stack pointer and // compute number of call frames to unwind asm("ssrdp %0" : "=r"(cur\_ssp):); // Skip the unwind if backward-edge CFI not active asm("beqz %0, back\_cfi\_not\_active" : "=r"(cur\_ssp):); // Unwind the frames in a loop while ( jmp\_buf->saved\_ssp > cur\_ssp ) { // advance by a maximum of 4K at a time to avoid // unwinding past bounds of the shadow stack cur\_ssp = ( (jmp\_buf->saved\_ssp - cur\_ssp) >= 4096 ) ? (cur\_ssp + 4096) : jmp\_buf->saved\_ssp; asm("csrw ssp, %0" : : "r" (cur\_ssp)); // Test if unwound past the shadow stack bounds asm("sspush x5"); asm("sspopchk x5"); } back\_cfi\_not\_active: : } | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#SSAMOSWAP)34.1.2.7\. Atomic Swap from a Shadow Stack Location ![svg](_images/svg-2ae67d26f273286801a52e320bebad4559dbc469.svg) For RV32, `SSAMOSWAP.W` atomically loads a 32-bit data value from address of a shadow stack location in `rs1`, puts the loaded value into register `rd`, and stores the 32-bit value held in `rs2` to the original address in `rs1`.`SSAMOSWAP.D` (RV64 only) is similar to `SSAMOSWAP.W` but operates on 64-bit data values. `SSAMOSWAP.W` for RV32 and `SSAMOSWAP.D` (RV64 only) operation if privilege_mode != M && menvcfg.SSE == 0 raise illegal-instruction exception else if S-mode not implemented raise illegal-instruction exception else if privilege_mode == U && senvcfg.SSE == 0 raise illegal-instruction exception else if privilege_mode == VS && henvcfg.SSE == 0 raise virtual-instruction exception else if privilege_mode == VU && senvcfg.SSE == 0 raise virtual-instruction exception else X(rd) = mem[X(rs1)] mem[X(rs1)] = X(rs2) endif For RV64, `SSAMOSWAP.W` atomically loads a 32-bit data value from address of a shadow stack location in `rs1`, sign-extends the loaded value and puts it in`rd`, and stores the lower 32 bits of the value held in `rs2` to the original address in `rs1`. `SSAMOSWAP.W` for RV64 if privilege_mode != M && menvcfg.SSE == 0 raise illegal-instruction exception else if S-mode not implemented raise illegal-instruction exception else if privilege_mode == U && senvcfg.SSE == 0 raise illegal-instruction exception else if privilege_mode == VS && henvcfg.SSE == 0 raise virtual-instruction exception else if privilege_mode == VU && senvcfg.SSE == 0 raise virtual-instruction exception else temp[31:0] = mem[X(rs1)] X(rd) = SignExtend(temp[31:0]) mem[X(rs1)] = X(rs2)[31:0] endif Just as for AMOs in the A extension, `SSAMOSWAP.W/D` requires that the address held in `rs1` be naturally aligned to the size of the operand (i.e., eight-byte aligned for _doublewords_, and four-byte aligned for _words_). The same exception options apply if the address is not naturally aligned. Just as for AMOs in the A extension, `SSAMOSWAP.W/D` optionally provides release consistency semantics, using the `aq` and `rl` bits, to help implement multiprocessor synchronization. An `SSAMOSWAP.W/D` operation has acquire semantics if `aq=1` and release semantics if `rl=1`. | | Stack switching is a common operation in user programs as well as supervisor programs. When a stack switch is performed the stack pointer of the currently active stack is saved into a context data structure and the new stack is made active by loading a new stack pointer from a context data structure. When shadow stacks are active for a program, the program needs to additionally switch the shadow stack pointer. If the pointer to the top of the deactivated shadow stack is held in a context data structure, then it may be susceptible to memory corruption vulnerabilities. To protect the pointer value, the program may store it at the top of the deactivated shadow stack itself and thereby create a checkpoint. A legal checkpoint is defined as one that holds a value of X, where X is the address at which the checkpoint is positioned on the shadow stack. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | An example sequence to restore the shadow stack pointer from the new shadow stack and save the old shadow stack pointer on the old shadow stack is as follows: \# a0 hold pointer to top of new shadow stack to switch to stack\_switch: ssrdp ra beqz ra, 2f # skip if Zicfiss not active ssamoswap.d ra, x0, (a0) # ra=\*\[a0\] and \*\[a0\]=0 beq ra, a0, 1f # \[a0\] must be == \[ra\] unimp # else crash 1: addi ra, ra, XLEN/8 # pop the checkpoint csrrw ra, ssp, ra # swap ssp: ra=ssp, ssp=ra addi ra, ra, -(XLEN/8) # checkpoint = "old ssp - XLEN/8" ssamoswap.d x0, ra, (ra) # Save checkpoint at "old ssp - XLEN/8" 2: This sequence uses the ra register. If the privilege mode at which this sequence is executed can be interrupted, then the trap handler should save thera on the shadow stack itself. There it is guarded against tampering and can be restored prior to returning from the trap. When a new shadow stack is created by the supervisor, it needs to store a checkpoint at the highest address on that stack. This enables the shadow stack pointer to be switched using the process outlined in this note. TheSSAMOSWAP.W/D instruction can be used to store this checkpoint. When the old value at the memory location operated on by SSAMOSWAP.W/D is not required,rd can be set to x0. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 31.1. "V" Standard Extension for Vector Operations, Version 1.0 ==================== ## [](#vector)31.1\. "V" Standard Extension for Vector Operations, Version 1.0 | | _The base vector extension is intended to provide general support for data-parallel execution within the 32-bit instruction encoding space, with later vector extensions supporting richer functionality for certain domains._ | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#31-1-1-introduction)31.1.1\. Introduction [31.1.18\. Standard Vector Extensions](#sec-vector-extensions) lists the standard vector extensions and which instructions and element widths are supported by each extension. ### [](#31-1-2-implementation-defined-constant-parameters)31.1.2\. Implementation-defined Constant Parameters Each hart supporting a vector extension defines two parameters: 1. The maximum size in bits of a vector element that any operation can produce or consume, _ELEN_ ≥ 8, which must be a power of 2. 2. The number of bits in a single vector register, _VLEN_ ≥ ELEN, which must be a power of 2, and must be no greater than 216. Standard vector extensions ([31.1.18\. Standard Vector Extensions](#sec-vector-extensions)) and architecture profiles may set further constraints on _ELEN_ and _VLEN_. | | Future extensions may allow ELEN > VLEN by holding one element using bits from multiple vector registers, but this extension does not include this option. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The upper limit on VLEN allows software to know that indices will fit into 16 bits (largest VLMAX of 65,536 occurs for LMUL=8 and SEW=8 with VLEN=65,536). Any future extension beyond 64Kib per vector register will require new configuration instructions such that software using the old configuration instructions does not see greater vector lengths. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The vector extension supports writing binary code that under certain constraints will execute portably on harts with different values for the VLEN parameter, provided the harts support the required element types and instructions. | | Code can be written that will expose differences in implementation parameters. | | --------------------------------------------------------------------------------- | | | In general, thread contexts with active vector state cannot be migrated during execution between harts that have any difference in VLEN or ELEN parameters. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#31-1-3-vector-extension-programmers-model)31.1.3\. Vector Extension Programmer’s Model The vector extension adds 32 vector registers, and seven unprivileged CSRs (`vstart`, `vxsat`, `vxrm`, `vcsr`, `vtype`, `vl`, `vlenb`) to a base scalar RISC-V ISA. __Table 1\. New vector CSRs__ | Address | Privilege | Name | Description | | ------- | --------- | ------ | ---------------------------------------- | | 0x008 | URW | vstart | Vector start element index | | 0x009 | URW | vxsat | Fixed-Point Saturate Flag | | 0x00A | URW | vxrm | Fixed-Point Rounding Mode | | 0x00F | URW | vcsr | Vector control and status register | | 0xC20 | URO | vl | Vector length | | 0xC21 | URO | vtype | Vector data type register | | 0xC22 | URO | vlenb | VLEN/8 (vector register length in bytes) | | | The four CSR numbers 0x00B\-0x00E are tentatively reserved for future vector CSRs, some of which may be mirrored into vcsr. | | ------------------------------------------------------------------------------------------------------------------------------ | #### [](#31-1-3-1-vector-registers)31.1.3.1\. Vector Registers The vector extension adds 32 architectural vector registers,`v0`\-`v31` to the base scalar RISC-V ISA. Each vector register has a fixed VLEN bits of state. #### [](#31-1-3-2-vector-context-status-in-mstatus)31.1.3.2\. Vector Context Status in `mstatus` A vector context status field, `VS`, is added to `mstatus[10:9]` and shadowed in `sstatus[10:9]`. It is defined analogously to the floating-point context status field, `FS`. Attempts to execute any vector instruction, or to access the vector CSRs, raise an illegal-instruction exception when `mstatus.VS` is set to Off. When `mstatus.VS` is set to Initial or Clean, executing any instruction that changes vector state, including the vector CSRs, will change `mstatus.VS` to Dirty. Implementations may also change `mstatus.VS` from Initial or Clean to Dirty at any time, even when there is no change in vector state. | | Accurate setting of mstatus.VS is an optimization. Software will typically use VS to reduce context-swap overhead. | | --------------------------------------------------------------------------------------------------------------------- | If `mstatus.VS` is Dirty, `mstatus.SD` is 1; otherwise, `mstatus.SD` is set in accordance with existing specifications. Implementations may have a writable `misa.V` field. Analogous to the way in which the floating-point unit is handled, the `mstatus.VS`field may exist even if `misa.V` is clear. | | Allowing mstatus.VS to exist when misa.V is clear, enables vector emulation and simplifies handling of mstatus.VS in systems with writable misa.V. | | ----------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-3-3-vector-context-status-in-vsstatus)31.1.3.3\. Vector Context Status in `vsstatus` When the hypervisor extension is present, a vector context status field, `VS`, is added to `vsstatus[10:9]`. It is defined analogously to the floating-point context status field, `FS`. When V=1, both `vsstatus.VS` and `mstatus.VS` are in effect: attempts to execute any vector instruction, or to access the vector CSRs, raise an illegal-instruction exception when either field is set to Off. When V=1 and neither `vsstatus.VS` nor `mstatus.VS` is set to Off, executing any instruction that changes vector state, including the vector CSRs, will change both `mstatus.VS` and `vsstatus.VS` to Dirty. Implementations may also change `mstatus.VS` or `vsstatus.VS` from Initial or Clean to Dirty at any time, even when there is no change in vector state. If `vsstatus.VS` is Dirty, `vsstatus.SD` is 1; otherwise, `vsstatus.SD` is set in accordance with existing specifications. If `mstatus.VS` is Dirty, `mstatus.SD` is 1; otherwise, `mstatus.SD` is set in accordance with existing specifications. For implementations with a writable `misa.V` field, the `vsstatus.VS` field may exist even if `misa.V` is clear. #### [](#31-1-3-4-vector-type-vtype-register)31.1.3.4\. Vector Type (`vtype`) Register The read-only XLEN-wide _vector_ _type_ CSR, `vtype` provides the default type used to interpret the contents of the vector register file, and can only be updated by `vset_i_vl_i_` instructions. The vector type determines the organization of elements in each vector register, and how multiple vector registers are grouped. The`vtype` register also indicates how masked-off elements and elements past the current vector length in a vector result are handled. | | Allowing updates only via the vset\_i\_vl\_i\_ instructions simplifies maintenance of the vtype register state. | | ------------------------------------------------------------------------------------------------------------------ | The `vtype` register has five fields, `vill`, `vma`, `vta`,`vsew[2:0]`, and `vlmul[2:0]`. Bits `vtype[XLEN-2:8]` should be written with zero, and non-zero values in this field are reserved. ![svg](_images/svg-e8a84db94ab974b97caf28e1738cb4330d74acf3.svg) | | This diagram shows the layout for RV32 systems, whereas in general vill should be at bit XLEN-1. | | --------------------------------------------------------------------------------------------------- | __Table 2\. vtype register layout__ | Bits | Name | Description | | -------- | ------------ | ----------------------------------------------- | | XLEN-1 | vill | Illegal value if set | | XLEN-2:8 | 0 | Reserved if non-zero | | 7 | vma | Vector mask agnostic | | 6 | vta | Vector tail agnostic | | 5:3 | vsew\[2:0\] | Selected element width (SEW) setting | | 2:0 | vlmul\[2:0\] | Vector register group multiplier (LMUL) setting | | | A small implementation supporting ELEN=32 requires only seven bits of state in vtype: two bits for ma and ta, two bits for vsew\[1:0\] and three bits for vlmul\[2:0\]. The illegal value represented by vill can be internally encoded using the illegal 64-bit combination in vsew\[1:0\] without requiring an additional storage bit to hold vill. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Further standard and custom vector extensions may extend these fields to support a greater variety of data types. | | -------------------------------------------------------------------------------------------------------------------- | | | The primary motivation for the vtype CSR is to allow the vector instruction set to fit into a 32-bit instruction encoding space. A separate vset\_i\_vl\_i\_ instruction can be used to set vland/or vtype fields before execution of a vector instruction, and implementations may choose to fuse these two instructions into a single internal vector microop. In many cases, the vl and vtype values can be reused across multiple instructions, reducing the static and dynamic instruction overhead from the vset\_i\_vl\_i\_ instructions. It is anticipated that a future extended 64-bit instruction encoding would allow these fields to be specified statically in the instruction encoding. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#31-1-3-4-1-vector-selected-element-width-vsew20)31.1.3.4.1\. Vector Selected Element Width (`vsew[2:0]`) The value in `vsew` sets the dynamic _selected_ _element_ _width_(SEW). By default, a vector register is viewed as being divided into VLEN/SEW elements. __Table 3\. vsew\[2:0\] (selected element width) encoding__ | vsew\[2:0\] | SEW | | | | ----------- | --- | - | -------- | | 0 | 0 | 0 | 8 | | 0 | 0 | 1 | 16 | | 0 | 1 | 0 | 32 | | 0 | 1 | 1 | 64 | | 1 | X | X | Reserved | | | While it is anticipated the larger vsew\[2:0\] encodings (100\-111) will be used to encode larger SEW, the encodings are formally _reserved_ at this point. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 4\. Example VLEN = 128 bits__ | SEW | Elements per vector register | | --- | ---------------------------- | | 64 | 2 | | 32 | 4 | | 16 | 8 | | 8 | 16 | The supported element width may vary with LMUL. | | The current set of standard vector extensions do not vary supported element width with LMUL. Some future extensions may support larger SEWs only when bits from multiple vector registers are combined using LMUL. In this case, software that relies on large SEW should attempt to use the largest LMUL, and hence the fewest vector register groups, to increase the number of implementations on which the code will run. The vill bit in vtype should be checked after settingvtype to see if the configuration is supported, and an alternate code path should be provided if it is not. Alternatively, a profile can mandate the minimum SEW at each LMUL setting. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#vector-register-grouping)31.1.3.4.2\. Vector Register Grouping (`vlmul[2:0]`) Multiple vector registers can be grouped together, so that a single vector instruction can operate on multiple vector registers. The term_vector_ _register_ _group_ is used herein to refer to one or more vector registers used as a single operand to a vector instruction. Vector register groups can be used to provide greater execution efficiency for longer application vectors, but the main reason for their inclusion is to allow double-width or larger elements to be operated on with the same vector length as single-width elements. The vector length multiplier, _LMUL_, when greater than 1, represents the default number of vector registers that are combined to form a vector register group. Implementations must support LMUL integer values of 1, 2, 4, and 8. | | The vector architecture includes instructions that take multiple source and destination vector operands with different element widths, but the same number of elements. The effective LMUL (EMUL) of each vector operand is determined by the number of registers required to hold the elements. For example, for a widening add operation, such as add 32-bit values to produce 64-bit results, a double-width result requires twice the LMUL of the single-width inputs. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | LMUL can also be a fractional value, reducing the number of bits used in a single vector register. Fractional LMUL is used to increase the number of effective usable vector register groups when operating on mixed-width values. | | With only integer LMUL values, a loop operating on a range of sizes would have to allocate at least one whole vector register (LMUL=1) for the narrowest data type and then would consume multiple vector registers (LMUL>1) to form a vector register group for each wider vector operand. This can limit the number of vector register groups available. With fractional LMUL, the widest values need occupy only a single vector register while narrower values can occupy a fraction of a single vector register, allowing all 32 architectural vector register names to be used for different values in a vector loop even when handling mixed-width values. Fractional LMUL implies portions of vector registers are unused, but in some cases, having more shorter register-resident vectors improves efficiency relative to fewer longer register-resident vectors. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Implementations must provide fractional LMUL settings that allow the narrowest supported type to occupy a fraction of a vector register corresponding to the ratio of the narrowest supported type’s width to that of the largest supported type’s width. In general, the requirement is to support LMUL ≥ SEWMIN/ELEN, where SEWMIN is the narrowest supported SEW value and ELEN is the widest supported SEW value. In the standard extensions, SEWMIN\=8\. For standard vector extensions with ELEN=32, fractional LMULs of 1/2 and 1/4 must be supported. For standard vector extensions with ELEN=64, fractional LMULs of 1/2, 1/4, and 1/8 must be supported. | | When LMUL < SEWMIN/ELEN, there is no guarantee an implementation would have enough bits in the fractional vector register to store at least one element, as VLEN=ELEN is a valid implementation choice. For example, with VLEN=ELEN=32, and SEWMIN\=8, an LMUL of 1/8 would only provide four bits of storage in a vector register. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For a given supported fractional LMUL setting, implementations must support SEW settings between SEWMIN and LMUL \* ELEN, inclusive. The use of `vtype` encodings with LMUL < SEWMIN/ELEN is_reserved_, but implementations can set `vill` if they do not support these configurations. | | Requiring all implementations to set vill in this case would prohibit future use of this case in an extension, so to allow for a future definition of LMUL> d) + r`, where `r` depends on the rounding mode as specified in the following table. __Table 5\. vxrm encoding__ | vxrm\[1:0\] | Abbreviation | Rounding Mode | Rounding increment, r | | | ----------- | ------------ | ------------- | ------------------------------------------ | ----------------------------------- | | 0 | 0 | rnu | round-to-nearest-up (add +0.5 LSB) | v\[d-1\] | | 0 | 1 | rne | round-to-nearest-even | v\[d-1\] & (v\[d-2:0\]≠0 \| v\[d\]) | | 1 | 0 | rdn | round-down | 0 | | 1 | 1 | rod | round-to-odd (OR bits into LSB, aka "jam") | !v\[d\] & v\[d-1:0\]≠0 | The rounding functions: roundoff_unsigned(v, d) = (unsigned(v) >> d) + r roundoff_signed(v, d) = (signed(v) >> d) + r are used to represent this operation in the instruction descriptions below. #### [](#31-1-3-9-vector-fixed-point-saturation-flag-vxsat)31.1.3.9\. Vector Fixed-Point Saturation Flag (`vxsat`) The `vxsat` CSR has a single read-write least-significant bit (`vxsat[0]`) that indicates if a fixed-point instruction has had to saturate an output value to fit into a destination format. Bits `vxsat[XLEN-1:1]` should be written as zeros. The `vxsat` bit is mirrored in `vcsr`. #### [](#31-1-3-10-vector-control-and-status-vcsr-register)31.1.3.10\. Vector Control and Status (`vcsr`) Register The `vxrm` and `vxsat` separate CSRs can also be accessed via fields in the _XLEN_\-bit-wide vector control and status CSR, `vcsr`. __Table 6\. vcsr layout__ | Bits | Name | Description | | -------- | ----------- | ----------------------------------- | | XLEN-1:3 | Reserved | | | 2:1 | vxrm\[1:0\] | Fixed-point rounding mode | | 0 | vxsat | Fixed-point accrued saturation flag | #### [](#31-1-3-11-state-of-vector-extension-at-reset)31.1.3.11\. State of Vector Extension at Reset The vector extension must have a consistent state at reset. In particular, `vtype` and `vl` must have values that can be read and then restored with a single `vsetvl` instruction. | | It is recommended that at reset, vtype.vill is set, the remaining bits in vtype are zero, and vl is set to zero. | | ------------------------------------------------------------------------------------------------------------------- | The `vstart`, `vxrm`, `vxsat` CSRs can have arbitrary values at reset. | | Most uses of the vector unit will require an initial vset\_i\_vl\_i\_, which will reset vstart. The vxrm and vxsat fields should be reset explicitly in software before use. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The vector registers can have arbitrary values at reset. ### [](#31-1-4-mapping-of-vector-elements-to-vector-register-state)31.1.4\. Mapping of Vector Elements to Vector Register State The following diagrams illustrate how different width elements are packed into the bytes of a vector register depending on the current SEW and LMUL settings, as well as implementation VLEN. Elements are packed into each vector register with the least-significant byte in the lowest-numbered bits. The mapping was chosen to provide the simplest and most portable model for software, but might appear to incur large wiring cost for wider vector datapaths on certain operations. The vector instruction set was expressly designed to support implementations that internally rearrange vector data for different SEW to reduce datapath wiring costs, while externally preserving the simple software model. | | For example, microarchitectures can track the EEW with which a vector register was written, and then insert additional scrambling operations to rearrange data if the register is accessed with a different EEW. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-4-1-mapping-for-lmul-1)31.1.4.1\. Mapping for LMUL = 1 When LMUL=1, elements are simply packed in order from the least-significant to most-significant bits of the vector register. | | To increase readability, vector register layouts are drawn with bytes ordered from right to left with increasing byte address. Bits within an element are numbered in a little-endian format with increasing bit index from right to left corresponding to increasing magnitude. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | LMUL=1 examples. The element index is given in hexadecimal and is shown placed at the least-significant byte of the stored element. VLEN=32b Byte 3 2 1 0 SEW=8b 3 2 1 0 SEW=16b 1 0 SEW=32b 0 VLEN=64b Byte 7 6 5 4 3 2 1 0 SEW=8b 7 6 5 4 3 2 1 0 SEW=16b 3 2 1 0 SEW=32b 1 0 SEW=64b 0 VLEN=128b Byte F E D C B A 9 8 7 6 5 4 3 2 1 0 SEW=8b F E D C B A 9 8 7 6 5 4 3 2 1 0 SEW=16b 7 6 5 4 3 2 1 0 SEW=32b 3 2 1 0 SEW=64b 1 0 VLEN=256b Byte 1F1E1D1C1B1A19181716151413121110 F E D C B A 9 8 7 6 5 4 3 2 1 0 SEW=8b 1F1E1D1C1B1A19181716151413121110 F E D C B A 9 8 7 6 5 4 3 2 1 0 SEW=16b F E D C B A 9 8 7 6 5 4 3 2 1 0 SEW=32b 7 6 5 4 3 2 1 0 SEW=64b 3 2 1 0 #### [](#31-1-4-2-mapping-for-lmul-1)31.1.4.2\. Mapping for LMUL < 1 When LMUL < 1, only the first LMUL\*VLEN/SEW elements in the vector register are used. The remaining space in the vector register is treated as part of the tail, and hence must obey the vta setting. Example, VLEN=128b, LMUL=1/4 Byte F E D C B A 9 8 7 6 5 4 3 2 1 0 SEW=8b - - - - - - - - - - - - 3 2 1 0 SEW=16b - - - - - - 1 0 SEW=32b - - - 0 #### [](#31-1-4-3-mapping-for-lmul-1)31.1.4.3\. Mapping for LMUL > 1 When vector registers are grouped, the elements of the vector register group are packed contiguously in element order beginning with the lowest-numbered vector register and moving to the next-highest-numbered vector register in the group once each vector register is filled. LMUL > 1 examples VLEN=32b, SEW=8b, LMUL=2 Byte 3 2 1 0 v2*n 3 2 1 0 v2*n+1 7 6 5 4 VLEN=32b, SEW=16b, LMUL=2 Byte 3 2 1 0 v2*n 1 0 v2*n+1 3 2 VLEN=32b, SEW=16b, LMUL=4 Byte 3 2 1 0 v4*n 1 0 v4*n+1 3 2 v4*n+2 5 4 v4*n+3 7 6 VLEN=32b, SEW=32b, LMUL=4 Byte 3 2 1 0 v4*n 0 v4*n+1 1 v4*n+2 2 v4*n+3 3 VLEN=64b, SEW=32b, LMUL=2 Byte 7 6 5 4 3 2 1 0 v2*n 1 0 v2*n+1 3 2 VLEN=64b, SEW=32b, LMUL=4 Byte 7 6 5 4 3 2 1 0 v4*n 1 0 v4*n+1 3 2 v4*n+2 5 4 v4*n+3 7 6 VLEN=128b, SEW=32b, LMUL=2 Byte F E D C B A 9 8 7 6 5 4 3 2 1 0 v2*n 3 2 1 0 v2*n+1 7 6 5 4 VLEN=128b, SEW=32b, LMUL=4 Byte F E D C B A 9 8 7 6 5 4 3 2 1 0 v4*n 3 2 1 0 v4*n+1 7 6 5 4 v4*n+2 B A 9 8 v4*n+3 F E D C #### [](#sec-mapping-mixed)31.1.4.4\. Mapping across Mixed-Width Operations The vector ISA is designed to support mixed-width operations without requiring additional explicit rearrangement instructions. The recommended software strategy when operating on multiple vectors with different precision values is to modify `vtype` dynamically to keep SEW/LMUL constant (and hence VLMAX constant). The following example shows four different packed element widths (8b, 16b, 32b, 64b) in a VLEN=128b implementation. The vector register grouping factor (LMUL) is increased by the relative element size such that each group can hold the same number of vector elements (VLMAX=8 in this example) to simplify strip-mining code. Example VLEN=128b, with SEW/LMUL=16 Byte F E D C B A 9 8 7 6 5 4 3 2 1 0 vn - - - - - - - - 7 6 5 4 3 2 1 0 SEW=8b, LMUL=1/2 vn 7 6 5 4 3 2 1 0 SEW=16b, LMUL=1 v2*n 3 2 1 0 SEW=32b, LMUL=2 v2*n+1 7 6 5 4 v4*n 1 0 SEW=64b, LMUL=4 v4*n+1 3 2 v4*n+2 5 4 v4*n+3 7 6 The following table shows each possible constant SEW/LMUL operating point for loops with mixed-width operations. Each column represents a constant SEW/LMUL operating point. Entries in table are the LMUL values that yield that column’s SEW/LMUL value for the data width on that row. In each column, an LMUL setting for a data width indicates that it can be aligned with the other data widths in the same column that also have an LMUL setting, such that all have the same VLMAX. | SEW/LMUL | | | | | | | | | -------- | - | - | - | -- | --- | --- | --- | | 1 | 2 | 4 | 8 | 16 | 32 | 64 | | | SEW= 8 | 8 | 4 | 2 | 1 | 1/2 | 1/4 | 1/8 | | SEW= 16 | 8 | 4 | 2 | 1 | 1/2 | 1/4 | | | SEW= 32 | 8 | 4 | 2 | 1 | 1/2 | | | | SEW= 64 | 8 | 4 | 2 | 1 | | | | Larger LMUL settings can also used to simply increase vector length to reduce instruction fetch and dispatch overheads in cases where fewer vector register groups are needed. #### [](#sec-mask-register-layout)31.1.4.5\. Mask Register Layout A vector mask occupies only one vector register regardless of SEW and LMUL. Each element is allocated a single mask bit in a mask vector register. The mask bit for element _i_ is located in bit _i_ of the mask register, independent of SEW or LMUL. ### [](#31-1-5-vector-instruction-formats)31.1.5\. Vector Instruction Formats The instructions in the vector extension fit under two existing major opcodes (LOAD-FP and STORE-FP) and one new major opcode (OP-V). Vector loads and stores are encoded within the scalar floating-point load and store major opcodes (LOAD-FP/STORE-FP). The vector load and store encodings repurpose a portion of the standard scalar floating-point load/store 12-bit immediate field to provide further vector instruction encoding, with bit 25 holding the standard vector mask bit (see [31.1.5.3.1\. Mask Encoding](#sec-vector-mask-encoding)). Format for Vector Load Instructions under LOAD-FP major opcode ![svg](_images/svg-c33ada74287fb606746e0a642dc087939adf30d9.svg) ![svg](_images/svg-aa3cc9c26838d293d4e8cbe826b918bc85618969.svg) ![svg](_images/svg-cd3756d4a31d83b159759877e3fd24cb6dc068bb.svg) Format for Vector Store Instructions under STORE-FP major opcode ![svg](_images/svg-bf9683d62ac15a4dddefe469419a1bd5e429ba79.svg) ![svg](_images/svg-f168513db6e0a43315553951aeacf2adfa05bc8b.svg) ![svg](_images/svg-ec19e0ce6d098a6a9dd29792d4080fe0543be251.svg) Formats for Vector Arithmetic Instructions under OP-V major opcode ![svg](_images/svg-5396f7fd8303a709b8af78a1ffcc077dfb7d493d.svg) ![svg](_images/svg-04fc029599f10329d0c6a4d4693a4ec4cd207737.svg) ![svg](_images/svg-94a609226d17644b20ab3d9d5f851b2ed19df521.svg) ![svg](_images/svg-00be6421ad58b605c5400be73b016ee77bf1d240.svg) ![svg](_images/svg-14f2eb1960a5198ae2dbca451267c2c4ff2f4ef4.svg) ![svg](_images/svg-88068752b1e8e34a02b47c4916b891effb2c7cd4.svg) ![svg](_images/svg-473cfaddc7051d77875939d6948f66d6eb63f560.svg) Formats for Vector Configuration Instructions under OP-V major opcode ![svg](_images/svg-ee84885be4124111ac9561f141d550dd046cd24a.svg) ![svg](_images/svg-89c9546cb74625ef0d399ee14772d704b3aea9b7.svg) ![svg](_images/svg-c0d30e9cda23dcb00661ffa9333a2adf72634b06.svg) Vector instructions can have scalar or vector source operands and produce scalar or vector results, and most vector instructions can be performed either unconditionally or conditionally under a mask. Vector loads and stores move bit patterns between vector register elements and memory. Vector arithmetic instructions operate on values held in vector register elements. #### [](#31-1-5-1-scalar-operands)31.1.5.1\. Scalar Operands Scalar operands can be immediates, or taken from the `x` registers, the `f` registers, or element 0 of a vector register. Scalar results are written to an `x` or `f` register or to element 0 of a vector register. Any vector register can be used to hold a scalar regardless of the current LMUL setting. | | Zfinx ("F in X") is a new ISA extension where floating-point instructions take their arguments from the integer register file. The vector extension is also compatible with Zfinx, where the Zfinx vector extension has vector-scalar floating-point instructions taking their scalar argument from the x registers. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | We considered but did not pursue overlaying the f registers onv registers. The adopted approach reduces vector register pressure, avoids interactions with the standard calling convention, simplifies high-performance scalar floating-point design, and provides compatibility with the Zfinx ISA option. Overlaying f with vwould provide the advantage of lowering the number of state bits in some implementations, but complicates high-performance designs and would prevent compatibility with the Zfinx ISA option. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#sec-vec-operands)31.1.5.2\. Vector Operands Each vector operand has an _effective_ _element_ _width_ (EEW) and an_effective_ LMUL (EMUL) that is used to determine the size and location of all the elements within a vector register group. By default, for most operands of most instructions, EEW=SEW and EMUL=LMUL. Some vector instructions have source and destination vector operands with the same number of elements but different widths, so that EEW and EMUL differ from SEW and LMUL respectively but EEW/EMUL = SEW/LMUL.For example, most widening arithmetic instructions have a source group with EEW=SEW and EMUL=LMUL but have a destination group with EEW=2\*SEW and EMUL=2\*LMUL. Narrowing instructions have a source operand that has EEW=2\*SEW and EMUL=2\*LMUL but with a destination where EEW=SEW and EMUL=LMUL. Vector operands or results may occupy one or more vector registers depending on EMUL, but are always specified using the lowest-numbered vector register in the group. Using other than the lowest-numbered vector register to specify a vector register group is a reserved encoding. A vector register cannot be used to provide source operands with more than one EEW for a single instruction. A mask register source is considered to have EEW=1 for this constraint. An encoding that would result in the same vector register being read with two or more different EEWs, including when the vector register appears at different positions within two or more vector register groups, is reserved. | | In practice, there is no software benefit to reading the same register with different EEW in the same instruction, and this constraint reduces complexity for implementations that internally rearrange data dependent on EEW. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | A destination vector register group can overlap a source vector register group only if one of the following holds: * The destination EEW equals the source EEW. * The destination EEW is smaller than the source EEW, and the lowest-numbered register in the destination vector register group is the same as the lowest-numbered register in the source vector register group. (For example, when LMUL=1,`vnsrl.wi v0, v0, 3` is legal, but a destination of `v1` is not). * The destination EEW is greater than the source EEW, the source EMUL is at least 1, and the highest-numbered register in the destination vector register group is the same as the highest-numbered register in the source vector register group. (For example, when LMUL=8, `vzext.vf4 v0, v6` is legal, but a source of `v0`, `v2`, or `v4` is not). For the purpose of determining register group overlap constraints, mask elements have EEW=1. | | The overlap constraints are designed to support resumable exceptions in machines without register renaming. | | -------------------------------------------------------------------------------------------------------------- | Any instruction encoding that violates the overlap constraints is reserved. When source and destination registers overlap and have different EEW, the instruction is mask- and tail-agnostic, regardless of the setting of the`vta` and `vma` bits in `vtype`. The largest vector register group used by an instruction can not be greater than 8 vector registers (i.e., EMUL≤8), and if a vector instruction would require greater than 8 vector registers in a group, the instruction encoding is reserved. For example, a widening operation that produces a widened vector register group result when LMUL=8 is reserved as this would imply a result EMUL=16. Widened scalar values, e.g., input and output to a widening reduction operation, are held in the first element of a vector register and have EMUL=1. #### [](#31-1-5-3-vector-masking)31.1.5.3\. Vector Masking Masking is supported on many vector instructions. Element operations that are masked off (inactive) never generate exceptions. The destination vector register elements corresponding to masked-off elements are handled with either a mask-undisturbed or mask-agnostic policy depending on the setting of the `vma` bit in `vtype`([31.1.3.4.3\. Vector Tail Agnostic and Vector Mask Agnostic vta and vma](#sec-agnostic)). The mask value used to control execution of a masked vector instruction is always supplied by vector register `v0`. | | Masks are held in vector registers, rather than in a separate mask register file, to reduce total architectural state and to simplify the ISA. | | ------------------------------------------------------------------------------------------------------------------------------------------------- | | | Future vector extensions may provide longer instruction encodings with space for a full mask register specifier. | | ------------------------------------------------------------------------------------------------------------------- | The destination vector register group for a masked vector instruction cannot overlap the source mask register (`v0`), unless the destination vector register is being written with a mask value (e.g., compares) or the scalar result of a reduction. These instruction encodings are reserved. | | This constraint supports restart with a non-zero vstart value. | | ----------------------------------------------------------------- | Other vector registers can be used to hold working mask values, and mask vector logical operations are provided to perform predicate calculations. As specified in [31.1.3.4.3\. Vector Tail Agnostic and Vector Mask Agnostic vta and vma](#sec-agnostic), mask destination tail elements are always treated as tail-agnostic, regardless of the setting of `vta`. ##### [](#sec-vector-mask-encoding)31.1.5.3.1\. Mask Encoding Where available, masking is encoded in a single-bit `vm` field in the instruction (`inst[25]`). | vm | Description | | -- | ------------------------------------------ | | 0 | vector result, only where v0.mask\[i\] = 1 | | 1 | unmasked | Vector masking is represented in assembler code as another vector operand, with `.t` indicating that the operation occurs when`v0.mask[i]` is `1` (`t` for "true"). If no masking operand is specified, unmasked vector execution (`vm=1`) is assumed. vop.v* v1, v2, v3, v0.t # enabled where v0.mask[i]=1, vm=0 vop.v* v1, v2, v3 # unmasked vector operation, vm=1 | | Even though the current vector extensions only support one vector mask register v0 and only the true form of predication, the assembly syntax writes it out in full to be compatible with future extensions that might add a mask register specifier and support both true and complement mask values. The .t suffix on the masking operand also helps to visually encode the use of a mask. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The .mask suffix is not part of the assembly syntax. We only append it in contexts where a mask vector is subscripted, e.g., v0.mask\[i\]. | | --------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#sec-inactive-defs)31.1.5.4\. Prestart, Active, Inactive, Body, and Tail Element Definitions The destination element indices operated on during a vector instruction’s execution can be divided into three disjoint subsets. * The _prestart_ elements are those whose element index is less than the initial value in the `vstart` register. The prestart elements do not raise exceptions and do not update the destination vector register. * The _body_ elements are those whose element index is greater than or equal to the initial value in the `vstart` register, and less than the current vector length setting in `vl`. The body can be split into two disjoint subsets: * The _active_ elements during a vector instruction’s execution are the elements within the body and where the current mask is enabled at that element position. The active elements can raise exceptions and update the destination vector register group. * The _inactive_ elements are the elements within the body but where the current mask is disabled at that element position. The inactive elements do not raise exceptions and do not update any destination vector register group unless masked agnostic is specified (`vtype.vma`\=1), in which case inactive elements may be overwritten with 1s. * The _tail_ elements during a vector instruction’s execution are the elements past the current vector length setting specified in `vl`.The tail elements do not raise exceptions, and do not update any destination vector register group unless tail agnostic is specified (`vtype.vta`\=1), in which case tail elements may be overwritten with 1s, or with the result of the instruction in the case of mask-producing instructions except for mask loads. When LMUL < 1, the tail includes the elements past VLMAX that are held in the same vector register. for element index x prestart(x) = (0 <= x < vstart) body(x) = (vstart <= x < vl) tail(x) = (vl <= x < max(VLMAX,VLEN/SEW)) mask(x) = unmasked || v0.mask[x] == 1 active(x) = body(x) && mask(x) inactive(x) = body(x) && !mask(x) When `vstart` ≥ `vl`, there are no body elements, and no elements are updated in any destination vector register group, including that no tail elements are updated with agnostic values. | | As a consequence, when vl\=0, no elements, including agnostic elements, are updated in the destination vector register group regardless of vstart. | | ----------------------------------------------------------------------------------------------------------------------------------------------------- | Instructions that write an `x` register or `f` register do so even when `vstart` ≥ `vl`, including when `vl`\=0. | | Some instructions such as vslidedown and vrgather may read indices past vl or even VLMAX in source vector register groups. The general policy is to return the value 0 when the index is greater than VLMAX in the source vector register group. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#sec-vector-config)31.1.6\. Configuration-Setting Instructions (`vsetvli`/`vsetivli`/`vsetvl`) One of the common approaches to handling a large number of elements is "strip mining" where each iteration of a loop handles some number of elements, and the iterations continue until all elements have been processed. The RISC-V vector specification provides direct, portable support for this approach. The application specifies the total number of elements to be processed (the application vector length or AVL) as a candidate value for `vl`, and the hardware responds via a general-purpose register with the (frequently smaller) number of elements that the hardware will handle per iteration (stored in `vl`), based on the microarchitectural implementation and the `vtype` setting. A straightforward loop structure, shown in [31.1.6.4\. Example of strip mining and changes to SEW](#example-stripmine-sew), depicts the ease with which the code keeps track of the remaining number of elements and the amount per iteration handled by hardware. A set of instructions is provided to allow rapid configuration of the values in `vl` and `vtype` to match application needs. The`vset_i_vl_i_` instructions set the `vtype` and `vl` CSRs based on their arguments, and write the new value of `vl` into `rd`. vsetvli rd, rs1, vtypei # rd = new vl, rs1 = AVL, vtypei = new vtype setting vsetivli rd, uimm, vtypei # rd = new vl, uimm = AVL, vtypei = new vtype setting vsetvl rd, rs1, rs2 # rd = new vl, rs1 = AVL, rs2 = new vtype value Formats for Vector Configuration Instructions under OP-V major opcode ![svg](_images/svg-ee84885be4124111ac9561f141d550dd046cd24a.svg) ![svg](_images/svg-89c9546cb74625ef0d399ee14772d704b3aea9b7.svg) ![svg](_images/svg-c0d30e9cda23dcb00661ffa9333a2adf72634b06.svg) #### [](#31-1-6-1-vtype-encoding)31.1.6.1\. `vtype` encoding ![svg](_images/svg-e8a84db94ab974b97caf28e1738cb4330d74acf3.svg) | | This diagram shows the layout for RV32 systems, whereas in general vill should be at bit XLEN-1. | | --------------------------------------------------------------------------------------------------- | __Table 7\. vtype register layout__ | Bits | Name | Description | | -------- | ------------ | ----------------------------------------------- | | XLEN-1 | vill | Illegal value if set | | XLEN-2:8 | 0 | Reserved if non-zero | | 7 | vma | Vector mask agnostic | | 6 | vta | Vector tail agnostic | | 5:3 | vsew\[2:0\] | Selected element width (SEW) setting | | 2:0 | vlmul\[2:0\] | Vector register group multiplier (LMUL) setting | The new `vtype` value is encoded in the immediate fields of `vsetvli`and `vsetivli`, and in the `rs2` register for `vsetvl`. Suggested assembler names used for vset{i}vli vtypei immediate e8 # SEW=8b e16 # SEW=16b e32 # SEW=32b e64 # SEW=64b mf8 # LMUL=1/8 mf4 # LMUL=1/4 mf2 # LMUL=1/2 m1 # LMUL=1 m2 # LMUL=2 m4 # LMUL=4 m8 # LMUL=8 Examples: vsetvli t0, a0, e8, m1, ta, ma # SEW= 8, LMUL=1 vsetvli t0, a0, e8, m2, ta, ma # SEW= 8, LMUL=2 vsetvli t0, a0, e32, mf2, ta, ma # SEW=32, LMUL=1/2 The `vsetvl` variant operates similarly to `vsetvli` except that it takes a `vtype` value from `rs2` and can be used for context restore. ##### [](#31-1-6-1-1-unsupported-vtype-values)31.1.6.1.1\. Unsupported `vtype` Values If the `vtype` value is not supported by the implementation, then the `vill` bit is set in `vtype`, the remaining bits in `vtype` are set to zero, and the `vl` register is also set to zero. | | Earlier drafts required a trap when setting vtype to an illegal value. However, this would have added the first data-dependent trap on a CSR write to the ISA. Implementations could choose to trap when illegal values are written to vtype instead of setting vill, to allow emulation to support new configurations for forward-compatibility. The current scheme supports light-weight runtime interrogation of the supported vector unit configurations by checking if vill is clear for a given setting. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | A `vtype` value with `vill` set is treated as an unsupported configuration. Implementations must consider all bits of the `vtype` value to determine if the configuration is supported. An unsupported value in any location within the `vtype` value must result in `vill` being set. | | In particular, all XLEN bits of the register vtype argument to the vsetvl instruction must be checked. Implementations cannot ignore fields they do not implement. All bits must be checked to ensure that new code assuming unsupported vector features in vtypetraps instead of executing incorrectly on an older implementation. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-6-2-avl-encoding)31.1.6.2\. AVL encoding The new vector length setting is based on AVL, which for `vsetvli` and `vsetvl` is encoded in the `rs1` and `rd`fields as follows: __Table 8\. AVL used in vsetvli and vsetvl instructions__ | rd | rs1 | AVL value | Effect on vl | | --- | --- | -------------------- | ---------------------------------------------- | | \- | !x0 | Value in x\[rs1\] | Normal strip mining | | !x0 | x0 | \~0 | Set vl to VLMAX | | x0 | x0 | Value in vl register | Keep existing vl (of course, vtype may change) | When _rs1_ is not `x0`, the AVL is an unsigned integer held in the `x`register specified by _rs1_, and the new `vl` value is also written to the `x` register specified by _rd_. When _rs1_\=`x0` but _rd_≠`x0`, the maximum unsigned integer value (`~0`) is used as the AVL, and the resulting VLMAX is written to `vl` and also to the `x` register specified by `rd`. When _rs1_\=`x0` and _rd_\=`x0`, the instructions operate as if the current vector length in `vl` is used as the AVL, and the resulting value is written to `vl`, but not to a destination register. This form can only be used when VLMAX and hence `vl` is not actually changed by the new SEW/LMUL ratio. Use of the instructions with a new SEW/LMUL ratio that would result in a change of VLMAX is reserved. Use of the instructions is also reserved if `vill` was 1 beforehand. Implementations may set `vill` in either case. | | This last form of the instructions allows the vtype register to be changed while maintaining the current vl, provided VLMAX is not reduced. This design was chosen to ensure vl would always hold a legal value for current vtype setting. The current vl value can be read from the vl CSR. The vl value could be reduced by these instructions if the new SEW/LMUL ratio causes VLMAX to shrink, and so this case has been reserved as it is not clear this is a generally useful operation, and implementations can otherwise assume vl is not changed by these instructions to optimize their microarchitecture. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For the `vsetivli` instruction, the AVL is encoded as a 5-bit zero-extended immediate (0—​31) in the `rs1` field. | | The encoding of AVL for vsetivli is the same as for regular CSR immediate values. | | ------------------------------------------------------------------------------------ | | | The vsetivli instruction provides more compact code when the dimensions of vectors are small and known to fit inside the vector registers, in which case there is no strip-mining overhead. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#constraints-on-setting-vl)31.1.6.3\. Constraints on Setting `vl` The `vset_i_vl_i_` instructions first set VLMAX according to their `vtype`argument, then set `vl` obeying the following constraints: 1. `vl = AVL` if `AVL ≤ VLMAX` 2. `ceil(AVL / 2) ≤ vl ≤ VLMAX` if `AVL < (2 * VLMAX)` 3. `vl = VLMAX` if `AVL ≥ (2 * VLMAX)` 4. Deterministic on any given implementation for same input AVL and VLMAX values 5. These specific properties follow from the prior rules: 1. `vl = 0` if `AVL = 0` 2. `vl > 0` if `AVL > 0` 3. `vl ≤ VLMAX` 4. `vl ≤ AVL` 5. a value read from `vl` when used as the AVL argument to `vset_i_vl_i_` results in the same value in `vl`, provided the resultant VLMAX equals the value of VLMAX at the time that `vl` was read | | The vl setting rules are designed to be sufficiently strict to preserve vl behavior across register spills and context swaps forAVL ≤ VLMAX, yet flexible enough to enable implementations to improve vector lane utilization for AVL > VLMAX. For example, this permits an implementation to set vl = ceil(AVL / 2)for VLMAX < AVL < 2\*VLMAX in order to evenly distribute work over the last two iterations of a strip-mine loop. Requirement 2 ensures that the first strip-mine iteration of reduction loops uses the largest vector length of all iterations, even in the case of AVL < 2\*VLMAX. This allows software to avoid needing to explicitly calculate a running maximum of vector lengths observed during a strip-mined loop. Requirement 2 also allows an implementation to set vl to VLMAX for VLMAX < AVL < 2\*VLMAX | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#example-stripmine-sew)31.1.6.4\. Example of strip mining and changes to SEW The SEW and LMUL settings can be changed dynamically to provide high throughput on mixed-width operations in a single loop. # Example: Load 16-bit values, widen multiply to 32b, shift 32b result # right by 3, store 32b values. # On entry: # a0 holds the total number of elements to process # a1 holds the address of the source array # a2 holds the address of the destination array loop: vsetvli a3, a0, e16, m4, ta, ma # vtype = 16-bit integer vectors; # also update a3 with vl (# of elements this iteration) vle16.v v4, (a1) # Get 16b vector slli t1, a3, 1 # Multiply # elements this iteration by 2 bytes/source element add a1, a1, t1 # Bump pointer vwmul.vx v8, v4, x10 # Widening multiply into 32b in vsetvli x0, x0, e32, m8, ta, ma # Operate on 32b values vsrl.vi v8, v8, 3 vse32.v v8, (a2) # Store vector of 32b elements slli t1, a3, 2 # Multiply # elements this iteration by 4 bytes/destination element add a2, a2, t1 # Bump pointer sub a0, a0, a3 # Decrement count by vl bnez a0, loop # Any more? ### [](#sec-vector-memory)31.1.7\. Vector Loads and Stores Vector loads and stores move values between vector registers and memory. Vector loads and stores can be masked, and they only access memory or raise exceptions for active elements. Masked vector loads do not update inactive elements in the destination vector register group, unless masked agnostic is specified (`vtype.vma`\=1). All vector loads and stores may generate and accept a non-zero `vstart` value. #### [](#31-1-7-1-vector-loadstore-instruction-encoding)31.1.7.1\. Vector Load/Store Instruction Encoding Vector loads and stores are encoded within the scalar floating-point load and store major opcodes (LOAD-FP/STORE-FP). The vector load and store encodings repurpose a portion of the standard scalar floating-point load/store 12-bit immediate field to provide further vector instruction encoding, with bit 25 holding the standard vector mask bit (see [31.1.5.3.1\. Mask Encoding](#sec-vector-mask-encoding)). Format for Vector Load Instructions under LOAD-FP major opcode ![svg](_images/svg-c33ada74287fb606746e0a642dc087939adf30d9.svg) ![svg](_images/svg-aa3cc9c26838d293d4e8cbe826b918bc85618969.svg) ![svg](_images/svg-cd3756d4a31d83b159759877e3fd24cb6dc068bb.svg) Format for Vector Store Instructions under STORE-FP major opcode ![svg](_images/svg-bf9683d62ac15a4dddefe469419a1bd5e429ba79.svg) ![svg](_images/svg-f168513db6e0a43315553951aeacf2adfa05bc8b.svg) ![svg](_images/svg-ec19e0ce6d098a6a9dd29792d4080fe0543be251.svg) | Field | Description | | ------------------------- | ---------------------------------------------------------------------------------------------------------------------- | | rs1\[4:0\] | specifies x register holding base address | | rs2\[4:0\] | specifies x register holding stride | | vs2\[4:0\] | specifies v register holding address offsets | | vs3\[4:0\] | specifies v register holding store data | | vd\[4:0\] | specifies v register destination of load | | vm | specifies whether vector masking is enabled (0 = mask enabled, 1 = mask disabled) | | width\[2:0\] | specifies size of memory elements, and distinguishes from FP scalar | | mew | extended memory element width. See [31.1.7.3\. Vector Load/Store Width Encoding](#sec-vector-loadstore-width-encoding) | | mop\[1:0\] | specifies memory addressing mode | | nf\[2:0\] | specifies the number of fields in each segment, for segment load/stores | | lumop\[4:0\]/sumop\[4:0\] | are additional fields encoding variants of unit-stride instructions | Vector memory unit-stride and constant-stride operations directly encode EEW of the data to be transferred statically in the instruction to reduce the number of `vtype` changes when accessing memory in a mixed-width routine. Indexed operations use the explicit EEW encoding in the instruction to set the size of the indices used, and use SEW/LMUL to specify the data width. #### [](#31-1-7-2-vector-loadstore-addressing-modes)31.1.7.2\. Vector Load/Store Addressing Modes The vector extension supports unit-stride, constant-stride, and indexed (scatter/gather) addressing modes. Vector load/store base registers and strides are taken from the GPR `x` registers. The base effective address for all vector accesses is given by the contents of the `x` register named in `rs1`. Vector unit-stride operations access elements stored contiguously in memory starting from the base effective address. Vector constant-stride operations access the first memory element at the base effective address, and then access subsequent elements at address increments given by the byte offset contained in the `x` register specified by `rs2`. Vector indexed operations add the contents of each element of the vector offset operand specified by `vs2` to the base effective address to give the effective address of each element. The data vector register group has EEW=SEW, EMUL=LMUL, while the offset vector register group has EEW encoded in the instruction and EMUL=(EEW/SEW)\*LMUL. The vector offset operand is treated as a vector of byte-address offsets. | | The indexed operations can also be used to access fields within a vector of objects, where the vs2 vector holds pointers to the base of the objects and the scalar x register holds the offset of the member field in each object. Supporting this case is why the indexed operations were not defined to scale the element indices by the data EEW. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If the vector offset elements are narrower than XLEN, they are zero-extended to XLEN before adding to the base effective address. If the vector offset elements are wider than XLEN, the least-significant XLEN bits are used in the address calculation. If the implementation does not support the EEW of the offset elements, the instruction is reserved. | | A profile may place an upper limit on the maximum supported index EEW (e.g., only up to XLEN) smaller than ELEN. | | ------------------------------------------------------------------------------------------------------------------- | The vector addressing modes are encoded using the 2-bit `mop[1:0]`field. __Table 9\. encoding for loads__ | mop \[1:0\] | Description | Opcodes | | | ----------- | ----------- | ----------------- | ----------- | | 0 | 0 | unit-stride | VLE | | 0 | 1 | indexed-unordered | VLUXEI | | 1 | 0 | constant-stride | VLSE | | 1 | 1 | indexed-ordered | VLOXEI | __Table 10\. encoding for stores__ | mop \[1:0\] | Description | Opcodes | | | ----------- | ----------- | ----------------- | ----------- | | 0 | 0 | unit-stride | VSE | | 0 | 1 | indexed-unordered | VSUXEI | | 1 | 0 | constant-stride | VSSE | | 1 | 1 | indexed-ordered | VSOXEI | Vector unit-stride and constant-stride memory accesses do not guarantee ordering between individual element accesses. The vector indexed load and store memory operations have two forms, ordered and unordered. The indexed-ordered variants preserve element ordering on memory accesses. For unordered instructions (`mop[1:0]`!=11) there is no guarantee on element access order. If the accesses are to a strongly ordered IO region, the element accesses can be initiated in any order. | | To provide ordered vector accesses to a strongly ordered IO region, the ordered indexed instructions should be used. | | ----------------------------------------------------------------------------------------------------------------------- | For implementations with precise vector traps, exceptions on indexed-unordered stores must also be precise. Additional unit-stride vector addressing modes are encoded using the 5-bit `lumop` and `sumop` fields in the unit-stride load and store instruction encodings respectively. __Table 11\. lumop__ | lumop\[4:0\] | Description | | | | | | ------------ | ----------- | - | - | - | -------------------------------- | | 0 | 0 | 0 | 0 | 0 | unit-stride load | | 0 | 1 | 0 | 0 | 0 | unit-stride, whole register load | | 0 | 1 | 0 | 1 | 1 | unit-stride, mask load, EEW=8 | | 1 | 0 | 0 | 0 | 0 | unit-stride fault-only-first | | x | x | x | x | x | other encodings reserved | __Table 12\. sumop__ | sumop\[4:0\] | Description | | | | | | ------------ | ----------- | - | - | - | --------------------------------- | | 0 | 0 | 0 | 0 | 0 | unit-stride store | | 0 | 1 | 0 | 0 | 0 | unit-stride, whole register store | | 0 | 1 | 0 | 1 | 1 | unit-stride, mask store, EEW=8 | | x | x | x | x | x | other encodings reserved | The `nf[2:0]` field encodes the number of fields in each segment. For regular vector loads and stores, `nf`\=0, indicating that a single value is moved between a vector register group and memory at each element position. Larger values in the `nf` field are used to access multiple contiguous fields within a segment as described below in[31.1.7.8\. Vector Load/Store Segment Instructions](#sec-aos). The `nf[2:0]` field also encodes the number of whole vector registers to transfer for the whole vector register load/store instructions. #### [](#sec-vector-loadstore-width-encoding)31.1.7.3\. Vector Load/Store Width Encoding Vector loads and stores have an EEW encoded directly in the instruction. The corresponding EMUL is calculated as EMUL = (EEW/SEW)\*LMUL. If the EMUL would be out of range (EMUL>8 or EMUL<1/8), the instruction encoding is reserved. The vector register groups must have legal register specifiers for the selected EMUL, otherwise the instruction encoding is reserved. Vector unit-stride and constant-stride use the EEW/EMUL encoded in the instruction for the data values, while vector indexed loads and stores use the EEW/EMUL encoded in the instruction for the index values and the SEW/LMUL encoded in `vtype` for the data values. Vector loads and stores are encoded using width values that are not claimed by the standard scalar floating-point loads and stores. Implementations must provide vector loads and stores with EEWs corresponding to all supported SEW settings. Vector load/store encodings for unsupported EEW widths are reserved. __Table 13\. Width encoding for vector loads and stores.__ | mew | width \[2:0\] | Mem bits | Data Reg bits | Index bits | Opcodes | | | | | ------------------ | ------------- | -------- | ------------- | ---------- | ------- | ---- | -- | --------------- | | Standard scalar FP | x | 0 | 0 | 1 | 16 | FLEN | \- | FLH/FSH | | Standard scalar FP | x | 0 | 1 | 0 | 32 | FLEN | \- | FLW/FSW | | Standard scalar FP | x | 0 | 1 | 1 | 64 | FLEN | \- | FLD/FSD | | Standard scalar FP | x | 1 | 0 | 0 | 128 | FLEN | \- | FLQ/FSQ | | Vector 8b element | 0 | 0 | 0 | 0 | 8 | 8 | \- | VLxE8/VSxE8 | | Vector 16b element | 0 | 1 | 0 | 1 | 16 | 16 | \- | VLxE16/VSxE16 | | Vector 32b element | 0 | 1 | 1 | 0 | 32 | 32 | \- | VLxE32/VSxE32 | | Vector 64b element | 0 | 1 | 1 | 1 | 64 | 64 | \- | VLxE64/VSxE64 | | Vector 8b index | 0 | 0 | 0 | 0 | SEW | SEW | 8 | VLxEI8/VSxEI8 | | Vector 16b index | 0 | 1 | 0 | 1 | SEW | SEW | 16 | VLxEI16/VSxEI16 | | Vector 32b index | 0 | 1 | 1 | 0 | SEW | SEW | 32 | VLxEI32/VSxEI32 | | Vector 64b index | 0 | 1 | 1 | 1 | SEW | SEW | 64 | VLxEI64/VSxEI64 | | Reserved | 1 | X | X | X | \- | \- | \- | | Mem bits is the size of each element accessed in memory. Data reg bits is the size of each data element accessed in register. Index bits is the size of each index accessed in register. The `mew` bit (`inst[28]`) when set is expected to be used to encode expanded memory sizes of 128 bits and above, but these encodings are currently reserved. #### [](#31-1-7-4-vector-unit-stride-instructions)31.1.7.4\. Vector Unit-Stride Instructions # Vector unit-stride loads and stores # vd destination, rs1 base address, vm is mask encoding (v0.t or ) vle8.v vd, (rs1), vm # 8-bit unit-stride load vle16.v vd, (rs1), vm # 16-bit unit-stride load vle32.v vd, (rs1), vm # 32-bit unit-stride load vle64.v vd, (rs1), vm # 64-bit unit-stride load # vs3 store data, rs1 base address, vm is mask encoding (v0.t or ) vse8.v vs3, (rs1), vm # 8-bit unit-stride store vse16.v vs3, (rs1), vm # 16-bit unit-stride store vse32.v vs3, (rs1), vm # 32-bit unit-stride store vse64.v vs3, (rs1), vm # 64-bit unit-stride store Additional unit-stride mask load and store instructions are provided to transfer mask values to/from memory. These operate similarly to unmasked byte loads or stores (EEW=8), except that the effective vector length is `evl`\=ceil(`vl`/8) (i.e. EMUL=1), and the destination register is always written with a tail-agnostic policy. # Vector unit-stride mask load vlm.v vd, (rs1) # Load byte vector of length ceil(vl/8) # Vector unit-stride mask store vsm.v vs3, (rs1) # Store byte vector of length ceil(vl/8) `vlm.v` and `vsm.v` are encoded with the same `width[2:0]`\=0 encoding as`vle8.v` and `vse8.v`, but are distinguished by different`lumop` and `sumop` encodings. Since `vlm.v` and `vsm.v` operate as byte loads and stores,`vstart` is in units of bytes for these instructions. | | vlm.v and vsm.v respect the vill field in vtype, as they depend on vtype indirectly through its constraints on vl. | | --------------------------------------------------------------------------------------------------------------------- | | | The previous assembler mnemonics vle1.v and vse1.v were confusing as length was handled differently for these instructions versus other element load/store instructions. To avoid software churn, these older assembly mnemonics are being retained as aliases. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | The primary motivation to provide mask load and store is to support machines that internally rearrange data to reduce cross-datapath wiring. However, these instructions also provide a convenient mechanism to use packed bit vectors in memory as mask values, and also reduce the cost of mask spill/fill by reducing need to changevl. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-7-5-vector-constant-stride-instructions)31.1.7.5\. Vector Constant-Stride Instructions # Vector constant-stride loads and stores # vd destination, rs1 base address, rs2 byte constant-stride vlse8.v vd, (rs1), rs2, vm # 8-bit constant-stride load vlse16.v vd, (rs1), rs2, vm # 16-bit constant-stride load vlse32.v vd, (rs1), rs2, vm # 32-bit constant-stride load vlse64.v vd, (rs1), rs2, vm # 64-bit constant-stride load # vs3 store data, rs1 base address, rs2 byte constant-stride vsse8.v vs3, (rs1), rs2, vm # 8-bit constant-stride store vsse16.v vs3, (rs1), rs2, vm # 16-bit constant-stride store vsse32.v vs3, (rs1), rs2, vm # 32-bit constant-stride store vsse64.v vs3, (rs1), rs2, vm # 64-bit constant-stride store Negative and zero strides are supported. Element accesses within a constant-stride instruction are unordered with respect to each other. When `rs2`\=`x0`, then an implementation is allowed, but not required, to perform fewer memory operations than the number of active elements, and may perform different numbers of memory operations across different dynamic executions of the same static instruction. | | Compilers must be aware to not use the x0 form for rs2 when the immediate stride is 0 if the intent is to require all memory accesses are performed. | | ------------------------------------------------------------------------------------------------------------------------------------------------------- | When `rs2!=x0` and the value of `x[rs2]=0`, the implementation must perform one memory access for each active element (but these accesses will not be ordered). | | As with other architectural mandates, implementations must_appear_ to perform each memory access. Microarchitectures are free to optimize away accesses that would not be observed by another agent, for example, in idempotent memory regions obeying RVWMO. For non-idempotent memory regions, where by definition each access can be observed by a device, the optimization would not be possible. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | When repeating ordered vector accesses to the same memory address are required, then an ordered indexed operation can be used. | | --------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-7-6-vector-indexed-instructions)31.1.7.6\. Vector Indexed Instructions # Vector indexed loads and stores # Vector indexed-unordered load instructions # vd destination, rs1 base address, vs2 byte offsets vluxei8.v vd, (rs1), vs2, vm # unordered 8-bit indexed load of SEW data vluxei16.v vd, (rs1), vs2, vm # unordered 16-bit indexed load of SEW data vluxei32.v vd, (rs1), vs2, vm # unordered 32-bit indexed load of SEW data vluxei64.v vd, (rs1), vs2, vm # unordered 64-bit indexed load of SEW data # Vector indexed-ordered load instructions # vd destination, rs1 base address, vs2 byte offsets vloxei8.v vd, (rs1), vs2, vm # ordered 8-bit indexed load of SEW data vloxei16.v vd, (rs1), vs2, vm # ordered 16-bit indexed load of SEW data vloxei32.v vd, (rs1), vs2, vm # ordered 32-bit indexed load of SEW data vloxei64.v vd, (rs1), vs2, vm # ordered 64-bit indexed load of SEW data # Vector indexed-unordered store instructions # vs3 store data, rs1 base address, vs2 byte offsets vsuxei8.v vs3, (rs1), vs2, vm # unordered 8-bit indexed store of SEW data vsuxei16.v vs3, (rs1), vs2, vm # unordered 16-bit indexed store of SEW data vsuxei32.v vs3, (rs1), vs2, vm # unordered 32-bit indexed store of SEW data vsuxei64.v vs3, (rs1), vs2, vm # unordered 64-bit indexed store of SEW data # Vector indexed-ordered store instructions # vs3 store data, rs1 base address, vs2 byte offsets vsoxei8.v vs3, (rs1), vs2, vm # ordered 8-bit indexed store of SEW data vsoxei16.v vs3, (rs1), vs2, vm # ordered 16-bit indexed store of SEW data vsoxei32.v vs3, (rs1), vs2, vm # ordered 32-bit indexed store of SEW data vsoxei64.v vs3, (rs1), vs2, vm # ordered 64-bit indexed store of SEW data | | The assembler syntax for indexed loads and stores usesei_x_ instead of e_x_ to indicate the statically encoded EEW is of the index not the data. | | --------------------------------------------------------------------------------------------------------------------------------------------------- | | | The indexed operations mnemonics have a "U" or "O" to distinguish between unordered and ordered, while the other vector addressing modes have no character. While this is perhaps a little less consistent, this approach minimizes disruption to existing software, as VSXEI previously meant "ordered" - and the opcode can be retained as an alias during transition to help reduce software churn. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-7-7-unit-stride-fault-only-first-loads)31.1.7.7\. Unit-stride Fault-Only-First Loads The unit-stride fault-only-first load instructions are used to vectorize loops with data-dependent exit conditions ("while" loops). These instructions execute as a regular load except that they will only take a trap caused by a synchronous exception on element 0. If element 0 raises an exception, `vl` is not modified, and the trap is taken. If an element > 0 raises an exception, the corresponding trap is not taken, and the vector length `vl` is reduced to the index of the element that would have raised an exception. Load instructions may overwrite active destination vector register group elements past the element index at which the trap is reported. Similarly, fault-only-first load instructions may update active destination elements past the element that causes trimming of the vector length (but not past the original vector length). The values of these spurious updates do not have to correspond to the values in memory at the addressed memory locations. Non-idempotent memory locations can only be accessed when it is known the corresponding element load operation will not be restarted due to a trap or vector-length trimming. # Vector unit-stride fault-only-first loads # vd destination, rs1 base address, vm is mask encoding (v0.t or ) vle8ff.v vd, (rs1), vm # 8-bit unit-stride fault-only-first load vle16ff.v vd, (rs1), vm # 16-bit unit-stride fault-only-first load vle32ff.v vd, (rs1), vm # 32-bit unit-stride fault-only-first load vle64ff.v vd, (rs1), vm # 64-bit unit-stride fault-only-first load strlen example using unit-stride fault-only-first instruction # size_t strlen(const char *str) # a0 holds *str strlen: mv a3, a0 # Save start loop: vsetvli a1, x0, e8, m8, ta, ma # Vector of bytes of maximum length vle8ff.v v8, (a3) # Load bytes csrr a1, vl # Get bytes read vmseq.vi v0, v8, 0 # Set v0[i] where v8[i] = 0 vfirst.m a2, v0 # Find first set bit add a3, a3, a1 # Bump pointer bltz a2, loop # Not found? add a0, a0, a1 # Sum start + bump add a3, a3, a2 # Add index sub a0, a3, a0 # Subtract start address+bump ret | | There is a security concern with fault-on-first loads, as they can be used to probe for valid effective addresses. The unit-stride versions only allow probing a region immediately contiguous to a known region, and so reduce the security impact when used in unprivileged code. However, code running in S-mode can establish arbitrary page translations that allow probing of random guest physical addresses provided by a hypervisor. Constant-stride and scatter/gather fault-only-first instructions are not provided due to lack of encoding space, but they can also represent a larger security hole, allowing even unprivileged software to easily check multiple random pages for accessibility without experiencing a trap. This standard does not address possible security mitigations for fault-only-first instructions. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Even when an exception is not raised, implementations are permitted to process fewer than `vl` elements and reduce `vl` accordingly, but if `vstart`\=0 and`vl`\>0, then at least one element must be processed. When the fault-only-first instruction takes a trap due to an interrupt, implementations should not reduce `vl` and should instead set a `vstart` value. | | When the fault-only-first instruction would trigger a debug data-watchpoint trap on an element after the first, implementations should not reduce vl but instead should trigger the debug trap as otherwise the event might be lost. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#sec-aos)31.1.7.8\. Vector Load/Store Segment Instructions The vector load/store segment instructions move multiple contiguous fields in memory to and from consecutively numbered vector registers. | | The name "segment" reflects that the items moved are subarrays with homogeneous elements. These operations can be used to transpose arrays between memory and registers, and can support operations on "array-of-structures" datatypes by unpacking each field in a structure into a separate vector register. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The three-bit `nf` field in the vector instruction encoding is an unsigned integer that contains one less than the number of fields per segment, _NFIELDS_. __Table 14\. NFIELDS Encoding__ | nf\[2:0\] | NFIELDS | | | | --------- | ------- | - | - | | 0 | 0 | 0 | 1 | | 0 | 0 | 1 | 2 | | 0 | 1 | 0 | 3 | | 0 | 1 | 1 | 4 | | 1 | 0 | 0 | 5 | | 1 | 0 | 1 | 6 | | 1 | 1 | 0 | 7 | | 1 | 1 | 1 | 8 | The EMUL setting must be such that EMUL \* NFIELDS ≤ 8, otherwise the instruction encoding is reserved. | | The product ceil(EMUL) \* NFIELDS represents the number of underlying vector registers that will be touched by a segmented load or store instruction. This constraint makes this total no larger than 1/4 of the architectural register file, and the same as for regular operations with EMUL=8. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Each field will be held in successively numbered vector register groups. When EMUL>1, each field will occupy a vector register group held in multiple successively numbered vector registers, and the vector register group for each field must follow the usual vector register alignment constraints (e.g., when EMUL=2 and NFIELDS=4, each field’s vector register group must start at an even vector register, but does not have to start at a multiple of 8 vector register number). If the vector register numbers accessed by the segment load or store would increment past 31, then the instruction encoding is reserved. | | This constraint is to help allow for forward-compatibility with a possible future longer instruction encoding that has more addressable vector registers. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | The `vl` register gives the number of segments to move, which is equal to the number of elements transferred to each vector register group. Masking is also applied at the level of whole segments. For segment loads and stores, the individual memory accesses used to access fields within each segment are unordered with respect to each other even for ordered indexed segment loads and stores. The `vstart` value is in units of whole segments. If a trap occurs during access to a segment, it is implementation-defined whether a subset of the faulting segment’s accesses are performed before the trap is taken. ##### [](#31-1-7-8-1-vector-unit-stride-segment-loads-and-stores)31.1.7.8.1\. Vector Unit-Stride Segment Loads and Stores The vector unit-stride load and store segment instructions move packed contiguous segments into multiple destination vector register groups. | | Where the segments hold structures with heterogeneous-sized fields, software can later unpack individual structure fields using additional instructions after the segment load brings data into the vector registers. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The assembler prefixes `vlseg`/`vsseg` are used for unit-stride segment loads and stores respectively. # Format # In this syntax, equals NFIELDS and is an integer in the range [2, 8]. vlsege.v vd, (rs1), vm # Unit-stride segment load template vssege.v vs3, (rs1), vm # Unit-stride segment store template # Examples vlseg8e8.v vd, (rs1), vm # Load eight vector registers with eight byte fields. vsseg3e32.v vs3, (rs1), vm # Store packed vector of 3*4-byte segments from vs3,vs3+1,vs3+2 to memory For loads, the `vd` register will hold the first field loaded from the segment. For stores, the `vs3` register is read to provide the first field to be stored to each segment. # Example 1 # Memory structure holds packed RGB pixels (24-bit data structure, 8bpp) vsetvli a1, t0, e8, m1, ta, ma vlseg3e8.v v8, (a0), vm # v8 holds the red pixels # v9 holds the green pixels # v10 holds the blue pixels # Example 2 # Memory structure holds complex values, 32b for real and 32b for imaginary vsetvli a1, t0, e32, m1, ta, ma vlseg2e32.v v8, (a0), vm # v8 holds real # v9 holds imaginary There are also fault-only-first versions of the unit-stride instructions. # Template for vector fault-only-first unit-stride segment loads. vlsegeff.v vd, (rs1), vm # Unit-stride fault-only-first segment loads For fault-only-first segment loads, if an exception is detected partway through accessing the zeroth segment, the trap is taken. If an exception is detected partway through accessing a subsequent segment,`vl` is reduced to the index of that segment. In both cases, it is implementation-defined whether a subset of the segment is loaded. These instructions may overwrite destination vector register group elements past the point at which a trap is reported or past the point at which vector length is trimmed. ##### [](#31-1-7-8-2-vector-constant-stride-segment-loads-and-stores)31.1.7.8.2\. Vector Constant-Stride Segment Loads and Stores Vector constant-stride segment loads and stores move contiguous segments where each segment is separated by the byte-stride offset given in the `rs2`GPR argument. | | Negative and zero strides are supported. | | ------------------------------------------- | # Format vlssege.v vd, (rs1), rs2, vm # Constant-stride segment loads vsssege.v vs3, (rs1), rs2, vm # Constant-stride segment stores # Examples vsetvli a1, t0, e8, m1, ta, ma vlsseg3e8.v v4, (x5), x6 # Load bytes at addresses x5+i*x6 into v4[i], # and bytes at addresses x5+i*x6+1 into v5[i], # and bytes at addresses x5+i*x6+2 into v6[i]. # Examples vsetvli a1, t0, e32, m1, ta, ma vssseg2e32.v v2, (x5), x6 # Store words from v2[i] to address x5+i*x6 # and words from v3[i] to address x5+i*x6+4 Accesses to the fields within each segment can occur in any order, including the case where the byte stride is such that segments overlap in memory. ##### [](#31-1-7-8-3-vector-indexed-segment-loads-and-stores)31.1.7.8.3\. Vector Indexed Segment Loads and Stores Vector indexed segment loads and stores move contiguous segments where each segment is located at an address given by adding the scalar base address in the `rs1` field to byte offsets in vector register `vs2`. Both ordered and unordered forms are provided, where the ordered forms access segments in element order. However, even for the ordered form, accesses to the fields within an individual segment are not ordered with respect to each other. The data vector register group has EEW=SEW, EMUL=LMUL, while the index vector register group has EEW encoded in the instruction with EMUL=(EEW/SEW)\*LMUL. The EMUL \* NFIELDS ≤ 8 constraint applies to the data vector register group. # Format vluxsegei.v vd, (rs1), vs2, vm # Indexed-unordered segment loads vloxsegei.v vd, (rs1), vs2, vm # Indexed-ordered segment loads vsuxsegei.v vs3, (rs1), vs2, vm # Indexed-unordered segment stores vsoxsegei.v vs3, (rs1), vs2, vm # Indexed-ordered segment stores # Examples vsetvli a1, t0, e8, m1, ta, ma vluxseg3ei8.v v4, (x5), v3 # Load bytes at addresses x5+v3[i] into v4[i], # and bytes at addresses x5+v3[i]+1 into v5[i], # and bytes at addresses x5+v3[i]+2 into v6[i]. # Examples vsetvli a1, t0, e32, m1, ta, ma vsuxseg2ei32.v v2, (x5), v5 # Store words from v2[i] to address x5+v5[i] # and words from v3[i] to address x5+v5[i]+4 For vector indexed segment loads, the destination vector register groups cannot overlap the source vector register group (specified by`vs2`), else the instruction encoding is reserved. | | This constraint supports restart of indexed segment loads that raise exceptions partway through loading a structure. | | ----------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-7-9-vector-loadstore-whole-register-instructions)31.1.7.9\. Vector Load/Store Whole Register Instructions Format for Vector Load Whole Register Instructions under LOAD-FP major opcode ![svg](_images/svg-25d17e197cd864bb05b12d267495186c715e661d.svg) Format for Vector Store Whole Register Instructions under STORE-FP major opcode ![svg](_images/svg-99457d1c1e73f9018dda9697be7d866cb43ce218.svg) These instructions load and store whole vector register groups. | | These instructions are intended to be used to save and restore vector registers when the type or length of the current contents of the vector register is not known, or where modifying vl and vtypewould be costly. Examples include compiler register spills, vector function calls where values are passed in vector registers, interrupt handlers, and OS context switches. Software can determine the number of bytes transferred by reading the vlenb register. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The load instructions have an EEW encoded in the `mew` and `width`fields following the pattern of regular unit-stride loads. | | Because in-register byte layouts are identical to in-memory byte layouts, the same data is written to the destination register group regardless of EEW. Hence, it would have sufficed to provide only EEW=8 variants. The full set of EEW variants is provided so that the encoded EEW can be used as a hint to indicate the destination register group will next be accessed with this EEW, which aids implementations that rearrange data internally. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The vector whole register store instructions are encoded similar to unmasked unit-stride store of elements with EEW=8. The `nf` field encodes how many vector registers to load and store using the NFIELDS encoding (Figure [Table 14](#fig-nf)). The encoded number of registers must be a power of 2 and the vector register numbers must be aligned as with a vector register group, otherwise the instruction encoding is reserved. NFIELDS indicates the number of vector registers to transfer, numbered successively after the base. Only NFIELDS values of 1, 2, 4, 8 are supported, with other values reserved. When multiple registers are transferred, the lowest-numbered vector register is held in the lowest-numbered memory addresses and successive vector register numbers are placed contiguously in memory. The instructions operate with an effective vector length,`evl`\=NFIELDS\*VLEN/EEW, regardless of current settings in `vtype` and`vl`. The usual property that no elements are written if `vstart`≥ `vl` does not apply to these instructions. Similarly, the property that the instructions are reserved if `vstart`exceeds the largest element index for the current `vtype` setting does not apply. Instead, the instructions are reserved if `vstart` ≥ `evl`. The instructions operate similarly to unmasked unit-stride load and store instructions, with the base address passed in the scalar `x`register specified by `rs1`. Implementations are allowed to raise a misaligned address exception on whole register loads and stores if the base address is not naturally aligned to the larger of the size of the encoded EEW in bytes (EEW/8) or the implementation’s smallest supported SEW size in bytes (SEWMIN/8). | | Allowing misaligned exceptions to be raised based on non-alignment to the encoded EEW simplifies the implementation of these instructions. Some subset implementations might not support smaller SEW widths, so are allowed to report misaligned exceptions for the smallest supported SEW even if larger than encoded EEW. An extreme non-standard implementation might have SEWMIN\>XLEN for example. Software environments can mandate the minimum alignment requirements to support an ABI. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | # Format of whole register load and store instructions. vl1r.v v3, (a0) # Pseudoinstruction equal to vl1re8.v vl1re8.v v3, (a0) # Load v3 with VLEN/8 bytes held at address in a0 vl1re16.v v3, (a0) # Load v3 with VLEN/16 halfwords held at address in a0 vl1re32.v v3, (a0) # Load v3 with VLEN/32 words held at address in a0 vl1re64.v v3, (a0) # Load v3 with VLEN/64 doublewords held at address in a0 vl2r.v v2, (a0) # Pseudoinstruction equal to vl2re8.v vl2re8.v v2, (a0) # Load v2-v3 with 2*VLEN/8 bytes from address in a0 vl2re16.v v2, (a0) # Load v2-v3 with 2*VLEN/16 halfwords held at address in a0 vl2re32.v v2, (a0) # Load v2-v3 with 2*VLEN/32 words held at address in a0 vl2re64.v v2, (a0) # Load v2-v3 with 2*VLEN/64 doublewords held at address in a0 vl4r.v v4, (a0) # Pseudoinstruction equal to vl4re8.v vl4re8.v v4, (a0) # Load v4-v7 with 4*VLEN/8 bytes from address in a0 vl4re16.v v4, (a0) vl4re32.v v4, (a0) vl4re64.v v4, (a0) vl8r.v v8, (a0) # Pseudoinstruction equal to vl8re8.v vl8re8.v v8, (a0) # Load v8-v15 with 8*VLEN/8 bytes from address in a0 vl8re16.v v8, (a0) vl8re32.v v8, (a0) vl8re64.v v8, (a0) vs1r.v v3, (a1) # Store v3 to address in a1 vs2r.v v2, (a1) # Store v2-v3 to address in a1 vs4r.v v4, (a1) # Store v4-v7 to address in a1 vs8r.v v8, (a1) # Store v8-v15 to address in a1 | | We have considered adding a whole register mask load instruction (vl1rm.v) but have decided to omit from initial extension. The primary purpose would be to inform the microarchitecture that the data will be used as a mask. The same effect can be achieved with the following code sequence, whose cost is at most four instructions. Of these, the first could likely be removed as vl is often already in a scalar register, and the last might already be present if the following vector instruction needs a new SEW/LMUL. So, in best case only two instructions (of which only one performs vector operations) are needed to synthesize the effect of the dedicated instruction: | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | csrr t0, vl # Save current vl (potentially not needed) vsetvli t1, x0, e8, m8, ta, ma # Maximum VLMAX vlm.v v0, (a0) # Load mask register vsetvli x0, t0, # Restore vl (potentially already present) ### [](#31-1-8-vector-memory-alignment-constraints)31.1.8\. Vector Memory Alignment Constraints If an element accessed by a vector memory instruction is not naturally aligned to the size of the element, either the element is transferred successfully or an address-misaligned exception is raised on that element. Support for misaligned vector memory accesses is independent of an implementation’s support for misaligned scalar memory accesses. | | An implementation may have neither, one, or both scalar and vector memory accesses support some or all misaligned accesses in hardware. A separate PMA should be defined to determine if vector misaligned accesses are supported in the associated address range. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Vector misaligned memory accesses follow the same rules for atomicity as scalar misaligned memory accesses. ### [](#31-1-9-vector-memory-consistency-model)31.1.9\. Vector Memory Consistency Model Vector memory instructions appear to execute in program order on the local hart. Vector memory instructions follow RVWMO at the instruction level. If the Ztso extension is implemented, vector memory instructions additionally follow RVTSO at the instruction level. Except for vector indexed-ordered loads and stores, element operations are unordered within the instruction. Vector indexed-ordered loads and stores read and write elements from/to memory in element order respectively, obeying RVWMO at the element level. | | Ztso only imposes RVTSO at the instruction level; intra-instruction ordering follows RVWMO regardless of whether Ztso is implemented. | | ---------------------------------------------------------------------------------------------------------------------------------------- | | | More formal definitions required. | | ------------------------------------ | Instructions affected by the vector length register `vl` have a control dependency on `vl`, rather than a data dependency. Similarly, masked vector instructions have a control dependency on the source mask register, rather than a data dependency. | | Treating the vector length and mask as control rather than data typically matches the semantics of the corresponding scalar code, where branch instructions ordinarily would have been used. Treating the mask as control allows masked vector load instructions to access memory before the mask value is known, without the need for a misspeculation-recovery mechanism. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#31-1-10-vector-arithmetic-instruction-formats)31.1.10\. Vector Arithmetic Instruction Formats The vector arithmetic instructions use a new major opcode (OP-V = 10101112) which neighbors OP-FP. The three-bit `funct3` field is used to define sub-categories of vector instructions. Formats for Vector Arithmetic Instructions under OP-V major opcode ![svg](_images/svg-5396f7fd8303a709b8af78a1ffcc077dfb7d493d.svg) ![svg](_images/svg-04fc029599f10329d0c6a4d4693a4ec4cd207737.svg) ![svg](_images/svg-94a609226d17644b20ab3d9d5f851b2ed19df521.svg) ![svg](_images/svg-00be6421ad58b605c5400be73b016ee77bf1d240.svg) ![svg](_images/svg-14f2eb1960a5198ae2dbca451267c2c4ff2f4ef4.svg) ![svg](_images/svg-88068752b1e8e34a02b47c4916b891effb2c7cd4.svg) ![svg](_images/svg-473cfaddc7051d77875939d6948f66d6eb63f560.svg) #### [](#sec-arithmetic-encoding)31.1.10.1\. Vector Arithmetic Instruction encoding The `funct3` field encodes the operand type and source locations. __Table 15\. funct3__ | funct3\[2:0\] | Category | Operands | Type of scalar operand | | | | ------------- | -------- | -------- | ---------------------- | ---------------- | ---------------------------- | | 0 | 0 | 0 | OPIVV | vector-vector | N/A | | 0 | 0 | 1 | OPFVV | vector-vector | N/A | | 0 | 1 | 0 | OPMVV | vector-vector | N/A | | 0 | 1 | 1 | OPIVI | vector-immediate | imm\[4:0\] | | 1 | 0 | 0 | OPIVX | vector-scalar | GPR x register rs1 | | 1 | 0 | 1 | OPFVF | vector-scalar | FP f register rs1 | | 1 | 1 | 0 | OPMVX | vector-scalar | GPR x register rs1 | | 1 | 1 | 1 | OPCFG | scalars-imms | GPR x register rs1 & rs2/imm | Integer operations are performed using unsigned or two’s-complement signed integer arithmetic depending on the opcode. | | In this discussion, fixed-point operations are considered to be integer operations. | | -------------------------------------------------------------------------------------- | All standard vector floating-point arithmetic operations follow the IEEE-754/2008 standard. All vector floating-point operations use the dynamic rounding mode in the `frm` register. Use of the `frm` field when it contains an invalid rounding mode by any vector floating-point instruction—​even those that do not depend on the rounding mode, or when `vl`\=0, or when `vstart` ≥ `vl`\--is reserved. | | All vector floating-point code will rely on a valid value infrm. Implementations can make all vector FP instructions report exceptions when the rounding mode is invalid to simplify control logic. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Vector-vector operations take two vectors of operands from vector register groups specified by `vs2` and `vs1` respectively. Vector-scalar operations can have three possible forms. In all three forms,the vector register group operand is specified by `vs2`. The second scalar source operand comes from one of three alternative sources: 1. For integer operations, the scalar can be a 5-bit immediate, `imm[4:0]`, encoded in the `rs1` field. The value is sign-extended to SEW bits, unless otherwise specified. 2. For integer operations, the scalar can be taken from the scalar `x`register specified by `rs1`. If XLEN>SEW, the least-significant SEW bits of the `x` register are used, unless otherwise specified. If XLEN SEW, the value in the `f` registers is checked for a valid NaN-boxed value, in which case the least-significant SEW bits of the `f` register are used, else the canonical NaN value is used. Vector instructions where any floating-point vector operand’s EEW is not a supported floating-point type width (which includes when FLEN < SEW) are reserved. | | Some instructions _zero_\-extend the 5-bit immediate, and denote this by naming the immediate uimm in the assembly syntax. | | ----------------------------------------------------------------------------------------------------------------------------- | | | When adding a vector extension to the Zfinx/Zdinx/Zhinx extensions, floating-point scalar arguments are taken from the x registers. NaN-boxing is not supported in these extensions, and so operands narrower than XLEN bits are not checked for a NaN box; bits XLEN-1:EEW are ignored. For RV32\_Zdinx, EEW=64 scalar arguments are supplied by an x\-register pair. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Vector arithmetic instructions are masked under control of the `vm`field. # Assembly syntax pattern for vector binary arithmetic instructions # Operations returning vector results, masked by vm (v0.t, ) vop.vv vd, vs2, vs1, vm # integer vector-vector vd[i] = vs2[i] op vs1[i] vop.vx vd, vs2, rs1, vm # integer vector-scalar vd[i] = vs2[i] op x[rs1] vop.vi vd, vs2, imm, vm # integer vector-immediate vd[i] = vs2[i] op imm vfop.vv vd, vs2, vs1, vm # FP vector-vector operation vd[i] = vs2[i] fop vs1[i] vfop.vf vd, vs2, rs1, vm # FP vector-scalar operation vd[i] = vs2[i] fop f[rs1] | | In the encoding, vs2 is the first operand, while rs1/immis the second operand. This is the opposite to the standard scalar ordering. This arrangement retains the existing encoding conventions that instructions that read only one scalar register, read it fromrs1, and that 5-bit immediates are sourced from the rs1 field. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | # Assembly syntax pattern for vector ternary arithmetic instructions (multiply-add) # Integer operations overwriting sum input vop.vv vd, vs1, vs2, vm # vd[i] = vs1[i] * vs2[i] + vd[i] vop.vx vd, rs1, vs2, vm # vd[i] = x[rs1] * vs2[i] + vd[i] # Integer operations overwriting product input vop.vv vd, vs1, vs2, vm # vd[i] = vs1[i] * vd[i] + vs2[i] vop.vx vd, rs1, vs2, vm # vd[i] = x[rs1] * vd[i] + vs2[i] # Floating-point operations overwriting sum input vfop.vv vd, vs1, vs2, vm # vd[i] = vs1[i] * vs2[i] + vd[i] vfop.vf vd, rs1, vs2, vm # vd[i] = f[rs1] * vs2[i] + vd[i] # Floating-point operations overwriting product input vfop.vv vd, vs1, vs2, vm # vd[i] = vs1[i] * vd[i] + vs2[i] vfop.vf vd, rs1, vs2, vm # vd[i] = f[rs1] * vd[i] + vs2[i] | | For ternary multiply-add operations, the assembler syntax always places the destination vector register first, followed by either rs1or vs1, then vs2. This ordering provides a more natural reading of the assembler for these ternary operations, as the multiply operands are always next to each other. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#sec-widening)31.1.10.2\. Widening Vector Arithmetic Instructions A few vector arithmetic instructions are defined to be _widening_operations where the destination vector register group has EEW=2\*SEW and EMUL=2\*LMUL. These are generally given a `vw*` prefix on the opcode, or `vfw*` for vector floating-point instructions. The first vector register group operand can be either single or double-width. # Assembly syntax pattern for vector widening arithmetic instructions # Double-width result, two single-width sources: 2*SEW = SEW op SEW vwop.vv vd, vs2, vs1, vm # integer vector-vector vd[i] = vs2[i] op vs1[i] vwop.vx vd, vs2, rs1, vm # integer vector-scalar vd[i] = vs2[i] op x[rs1] # Double-width result, first source double-width, second source single-width: 2*SEW = 2*SEW op SEW vwop.wv vd, vs2, vs1, vm # integer vector-vector vd[i] = vs2[i] op vs1[i] vwop.wx vd, vs2, rs1, vm # integer vector-scalar vd[i] = vs2[i] op x[rs1] | | Originally, a w suffix was used on opcode, but this could be confused with the use of a w suffix to mean word-sized operations in doubleword integers, so the w was moved to prefix. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The floating-point widening operations were changed to vfw\*from vwf\* to be more consistent with any scalar widening floating-point operations that will be written as fw\*. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Widening instruction encodings must follow the constraints in[31.1.5.2\. Vector Operands](#sec-vec-operands). #### [](#sec-narrowing)31.1.10.3\. Narrowing Vector Arithmetic Instructions A few instructions are provided to convert double-width source vectors into single-width destination vectors. These instructions convert a vector register group specified by `vs2` with EEW/EMUL=2\*SEW/2\*LMUL to a vector register group with the current SEW/LMUL setting. Where there is a second source vector register group (specified by `vs1`), this has the same (narrower) width as the result (i.e., EEW=SEW). | | An alternative design decision would have been to treat SEW/LMUL as defining the size of the source vector register group. The choice here is motivated by the belief the chosen approach will require fewervtype changes. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Compare operations that set a mask register are also implicitly a narrowing operation. | | ----------------------------------------------------------------------------------------- | A `vn*` prefix on the opcode is used to distinguish these instructions in the assembler, or a `vfn*` prefix for narrowing floating-point opcodes. The double-width source vector register group is signified by a `w` in the source operand suffix (e.g., `vnsra.wv`) Assembly syntax pattern for vector narrowing arithmetic instructions # Single-width result vd, double-width source vs2, single-width source vs1/rs1 # SEW = 2*SEW op SEW vnop.wv vd, vs2, vs1, vm # integer vector-vector vd[i] = vs2[i] op vs1[i] vnop.wx vd, vs2, rs1, vm # integer vector-scalar vd[i] = vs2[i] op x[rs1] Narrowing instruction encodings must follow the constraints in[31.1.5.2\. Vector Operands](#sec-vec-operands). ### [](#sec-vector-integer)31.1.11\. Vector Integer Arithmetic Instructions A set of vector integer arithmetic instructions is provided. Unless otherwise stated, integer operations wrap around on overflow. #### [](#31-1-11-1-vector-single-width-integer-add-and-subtract)31.1.11.1\. Vector Single-Width Integer Add and Subtract Vector integer add and subtract are provided. Reverse-subtract instructions are also provided for the vector-scalar forms. # Integer adds. vadd.vv vd, vs2, vs1, vm # Vector-vector vadd.vx vd, vs2, rs1, vm # vector-scalar vadd.vi vd, vs2, imm, vm # vector-immediate # Integer subtract vsub.vv vd, vs2, vs1, vm # Vector-vector vsub.vx vd, vs2, rs1, vm # vector-scalar # Integer reverse subtract vrsub.vx vd, vs2, rs1, vm # vd[i] = x[rs1] - vs2[i] vrsub.vi vd, vs2, imm, vm # vd[i] = imm - vs2[i] | | A vector of integer values can be negated using a reverse-subtract instruction with a scalar operand of x0. An assembly pseudoinstruction vneg.v vd,vs \= vrsub.vx vd,vs,x0 is provided. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-11-2-vector-widening-integer-addsubtract)31.1.11.2\. Vector Widening Integer Add/Subtract The widening add/subtract instructions are provided in both signed and unsigned variants, depending on whether the narrower source operands are first sign- or zero-extended before forming the double-width sum. # Widening unsigned integer add/subtract, 2*SEW = SEW +/- SEW vwaddu.vv vd, vs2, vs1, vm # vector-vector vwaddu.vx vd, vs2, rs1, vm # vector-scalar vwsubu.vv vd, vs2, vs1, vm # vector-vector vwsubu.vx vd, vs2, rs1, vm # vector-scalar # Widening signed integer add/subtract, 2*SEW = SEW +/- SEW vwadd.vv vd, vs2, vs1, vm # vector-vector vwadd.vx vd, vs2, rs1, vm # vector-scalar vwsub.vv vd, vs2, vs1, vm # vector-vector vwsub.vx vd, vs2, rs1, vm # vector-scalar # Widening unsigned integer add/subtract, 2*SEW = 2*SEW +/- SEW vwaddu.wv vd, vs2, vs1, vm # vector-vector vwaddu.wx vd, vs2, rs1, vm # vector-scalar vwsubu.wv vd, vs2, vs1, vm # vector-vector vwsubu.wx vd, vs2, rs1, vm # vector-scalar # Widening signed integer add/subtract, 2*SEW = 2*SEW +/- SEW vwadd.wv vd, vs2, vs1, vm # vector-vector vwadd.wx vd, vs2, rs1, vm # vector-scalar vwsub.wv vd, vs2, vs1, vm # vector-vector vwsub.wx vd, vs2, rs1, vm # vector-scalar | | An integer value can be doubled in width using the widening add instructions with a scalar operand of x0. Assembly pseudoinstructions vwcvt.x.x.v vd,vs,vm \= vwadd.vx vd,vs,x0,vm andvwcvtu.x.x.v vd,vs,vm \= vwaddu.vx vd,vs,x0,vm are provided. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-11-3-vector-integer-extension)31.1.11.3\. Vector Integer Extension The vector integer extension instructions zero- or sign-extend a source vector integer operand with EEW less than SEW to fill SEW-sized elements in the destination. The EEW of the source is 1/2, 1/4, or 1/8 of SEW, while EMUL of the source is (EEW/SEW)\*LMUL. The destination has EEW equal to SEW and EMUL equal to LMUL. vzext.vf2 vd, vs2, vm # Zero-extend SEW/2 source to SEW destination vsext.vf2 vd, vs2, vm # Sign-extend SEW/2 source to SEW destination vzext.vf4 vd, vs2, vm # Zero-extend SEW/4 source to SEW destination vsext.vf4 vd, vs2, vm # Sign-extend SEW/4 source to SEW destination vzext.vf8 vd, vs2, vm # Zero-extend SEW/8 source to SEW destination vsext.vf8 vd, vs2, vm # Sign-extend SEW/8 source to SEW destination If the source EEW is not a supported width, or source EMUL would be below the minimum legal LMUL, the instruction encoding is reserved. | | Standard vector load instructions access memory values that are the same size as the destination register elements. Some application code needs to operate on a range of operand widths in a wider element, for example, loading a byte from memory and adding to an eight-byte element. To avoid having to provide the cross-product of the number of vector load instructions by the number of data types (byte, word, halfword, and also signed/unsigned variants), we instead add explicit extension instructions that can be used if an appropriate widening arithmetic instruction is not available. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-11-4-vector-integer-add-with-carry-subtract-with-borrow-instructions)31.1.11.4\. Vector Integer Add-with-Carry / Subtract-with-Borrow Instructions To support multi-word integer arithmetic, instructions that operate on a carry bit are provided. For each operation (add or subtract), two instructions are provided: one to provide the result (SEW width), and the second to generate the carry output (single bit encoded as a mask boolean). The carry inputs and outputs are represented using the mask register layout as described in [31.1.4.5\. Mask Register Layout](#sec-mask-register-layout). Due to encoding constraints, the carry input must come from the implicit `v0`register, but carry outputs can be written to any vector register that respects the source/destination overlap restrictions. `vadc` and `vsbc` add or subtract the source operands and the carry-in or borrow-in, and write the result to vector register `vd`. These instructions are encoded as masked instructions (`vm=0`), but they operate on and write back all body elements. Encodings corresponding to the unmasked versions (`vm=1`) are reserved. `vmadc` and `vmsbc` add or subtract the source operands, optionally add the carry-in or subtract the borrow-in if masked (`vm=0`), and write the resulting carry-out or borrow-out back to mask register `vd`. If unmasked (`vm=1`), there is no carry-in or borrow-in. These instructions operate on and write back all body elements, even if masked. Because these instructions produce a mask value, they always operate with a tail-agnostic policy. # Produce sum with carry. # vd[i] = vs2[i] + vs1[i] + v0.mask[i] vadc.vvm vd, vs2, vs1, v0 # Vector-vector # vd[i] = vs2[i] + x[rs1] + v0.mask[i] vadc.vxm vd, vs2, rs1, v0 # Vector-scalar # vd[i] = vs2[i] + imm + v0.mask[i] vadc.vim vd, vs2, imm, v0 # Vector-immediate # Produce carry out in mask register format # vd.mask[i] = carry_out(vs2[i] + vs1[i] + v0.mask[i]) vmadc.vvm vd, vs2, vs1, v0 # Vector-vector # vd.mask[i] = carry_out(vs2[i] + x[rs1] + v0.mask[i]) vmadc.vxm vd, vs2, rs1, v0 # Vector-scalar # vd.mask[i] = carry_out(vs2[i] + imm + v0.mask[i]) vmadc.vim vd, vs2, imm, v0 # Vector-immediate # vd.mask[i] = carry_out(vs2[i] + vs1[i]) vmadc.vv vd, vs2, vs1 # Vector-vector, no carry-in # vd.mask[i] = carry_out(vs2[i] + x[rs1]) vmadc.vx vd, vs2, rs1 # Vector-scalar, no carry-in # vd.mask[i] = carry_out(vs2[i] + imm) vmadc.vi vd, vs2, imm # Vector-immediate, no carry-in Because implementing a carry propagation requires executing two instructions with unchanged inputs, destructive accumulations will require an additional move to obtain correct results. # Example multi-word arithmetic sequence, accumulating into v4 vmadc.vvm v1, v4, v8, v0 # Get carry into temp register v1 vadc.vvm v4, v4, v8, v0 # Calc new sum vmmv.m v0, v1 # Move temp carry into v0 for next word The subtract with borrow instruction `vsbc` performs the equivalent function to support long word arithmetic for subtraction. There are no subtract with immediate instructions. # Produce difference with borrow. # vd[i] = vs2[i] - vs1[i] - v0.mask[i] vsbc.vvm vd, vs2, vs1, v0 # Vector-vector # vd[i] = vs2[i] - x[rs1] - v0.mask[i] vsbc.vxm vd, vs2, rs1, v0 # Vector-scalar # Produce borrow out in mask register format # vd.mask[i] = borrow_out(vs2[i] - vs1[i] - v0.mask[i]) vmsbc.vvm vd, vs2, vs1, v0 # Vector-vector # vd.mask[i] = borrow_out(vs2[i] - x[rs1] - v0.mask[i]) vmsbc.vxm vd, vs2, rs1, v0 # Vector-scalar # vd.mask[i] = borrow_out(vs2[i] - vs1[i]) vmsbc.vv vd, vs2, vs1 # Vector-vector, no borrow-in # vd.mask[i] = borrow_out(vs2[i] - x[rs1]) vmsbc.vx vd, vs2, rs1 # Vector-scalar, no borrow-in For `vmsbc`, the borrow is defined to be 1 iff the difference, prior to truncation, is negative. For `vadc` and `vsbc`, the instruction encoding is reserved if the destination vector register is `v0`. | | This constraint corresponds to the constraint on masked vector operations that overwrite the mask register. | | -------------------------------------------------------------------------------------------------------------- | #### [](#31-1-11-5-vector-bitwise-logical-instructions)31.1.11.5\. Vector Bitwise Logical Instructions # Bitwise logical operations. vand.vv vd, vs2, vs1, vm # Vector-vector vand.vx vd, vs2, rs1, vm # vector-scalar vand.vi vd, vs2, imm, vm # vector-immediate vor.vv vd, vs2, vs1, vm # Vector-vector vor.vx vd, vs2, rs1, vm # vector-scalar vor.vi vd, vs2, imm, vm # vector-immediate vxor.vv vd, vs2, vs1, vm # Vector-vector vxor.vx vd, vs2, rs1, vm # vector-scalar vxor.vi vd, vs2, imm, vm # vector-immediate | | With an immediate of -1, scalar-immediate forms of the vxorinstruction provide a bitwise NOT operation. This is provided as an assembler pseudoinstruction vnot.v vd,vs,vm \= vxor.vi vd,vs,-1,vm. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-11-6-vector-single-width-shift-instructions)31.1.11.6\. Vector Single-Width Shift Instructions A full set of vector shift instructions are provided, including logical shift left (`sll`), and logical (zero-extending `srl`) and arithmetic (sign-extending `sra`) shift right. The data to be shifted is in the vector register group specified by `vs2` and the shift amount value can come from a vector register group `vs1`, a scalar integer register `rs1`, or a zero-extended 5-bit immediate. Only the low lg2(SEW) bits of the shift-amount value are used to control the shift amount. # Bit shift operations vsll.vv vd, vs2, vs1, vm # Vector-vector vsll.vx vd, vs2, rs1, vm # vector-scalar vsll.vi vd, vs2, uimm, vm # vector-immediate vsrl.vv vd, vs2, vs1, vm # Vector-vector vsrl.vx vd, vs2, rs1, vm # vector-scalar vsrl.vi vd, vs2, uimm, vm # vector-immediate vsra.vv vd, vs2, vs1, vm # Vector-vector vsra.vx vd, vs2, rs1, vm # vector-scalar vsra.vi vd, vs2, uimm, vm # vector-immediate #### [](#31-1-11-7-vector-narrowing-integer-right-shift-instructions)31.1.11.7\. Vector Narrowing Integer Right Shift Instructions The narrowing right shifts extract a smaller field from a wider operand and have both zero-extending (`srl`) and sign-extending (`sra`) forms. The shift amount can come from a vector register group, or a scalar `x` register, or a zero-extended 5-bit immediate. The low lg2(2\*SEW) bits of the shift-amount value are used (e.g., the low 6 bits for a SEW=64-bit to SEW=32-bit narrowing operation). # Narrowing shift right logical, SEW = (2*SEW) >> SEW vnsrl.wv vd, vs2, vs1, vm # vector-vector vnsrl.wx vd, vs2, rs1, vm # vector-scalar vnsrl.wi vd, vs2, uimm, vm # vector-immediate # Narrowing shift right arithmetic, SEW = (2*SEW) >> SEW vnsra.wv vd, vs2, vs1, vm # vector-vector vnsra.wx vd, vs2, rs1, vm # vector-scalar vnsra.wi vd, vs2, uimm, vm # vector-immediate | | Future extensions might add support for versions that narrow to a destination that is 1/4 the width of the source. | | --------------------------------------------------------------------------------------------------------------------- | | | An integer value can be halved in width using the narrowing integer shift instructions with a scalar operand of x0. An assembly pseudoinstruction is provided vncvt.x.x.w vd,vs,vm \= vnsrl.wx vd,vs,x0,vm. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-11-8-vector-integer-compare-instructions)31.1.11.8\. Vector Integer Compare Instructions The following integer compare instructions write 1 to the destination mask register element if the comparison evaluates to true, and 0 otherwise. The destination mask vector is always held in a single vector register, with a layout of elements as described in[31.1.4.5\. Mask Register Layout](#sec-mask-register-layout). The destination mask vector register may be the same as the source vector mask register (`v0`). # Set if equal vmseq.vv vd, vs2, vs1, vm # Vector-vector vmseq.vx vd, vs2, rs1, vm # vector-scalar vmseq.vi vd, vs2, imm, vm # vector-immediate # Set if not equal vmsne.vv vd, vs2, vs1, vm # Vector-vector vmsne.vx vd, vs2, rs1, vm # vector-scalar vmsne.vi vd, vs2, imm, vm # vector-immediate # Set if less than, unsigned vmsltu.vv vd, vs2, vs1, vm # Vector-vector vmsltu.vx vd, vs2, rs1, vm # Vector-scalar # Set if less than, signed vmslt.vv vd, vs2, vs1, vm # Vector-vector vmslt.vx vd, vs2, rs1, vm # vector-scalar # Set if less than or equal, unsigned vmsleu.vv vd, vs2, vs1, vm # Vector-vector vmsleu.vx vd, vs2, rs1, vm # vector-scalar vmsleu.vi vd, vs2, imm, vm # Vector-immediate # Set if less than or equal, signed vmsle.vv vd, vs2, vs1, vm # Vector-vector vmsle.vx vd, vs2, rs1, vm # vector-scalar vmsle.vi vd, vs2, imm, vm # vector-immediate # Set if greater than, unsigned vmsgtu.vx vd, vs2, rs1, vm # Vector-scalar vmsgtu.vi vd, vs2, imm, vm # Vector-immediate # Set if greater than, signed vmsgt.vx vd, vs2, rs1, vm # Vector-scalar vmsgt.vi vd, vs2, imm, vm # Vector-immediate # Following two instructions are not provided directly # Set if greater than or equal, unsigned # vmsgeu.vx vd, vs2, rs1, vm # Vector-scalar # Set if greater than or equal, signed # vmsge.vx vd, vs2, rs1, vm # Vector-scalar The following table indicates how all comparisons are implemented in native machine code. Comparison Assembler Mapping Assembler Pseudoinstruction va < vb vmslt{u}.vv vd, va, vb, vm va <= vb vmsle{u}.vv vd, va, vb, vm va > vb vmslt{u}.vv vd, vb, va, vm vmsgt{u}.vv vd, va, vb, vm va >= vb vmsle{u}.vv vd, vb, va, vm vmsge{u}.vv vd, va, vb, vm va < x vmslt{u}.vx vd, va, x, vm va <= x vmsle{u}.vx vd, va, x, vm va > x vmsgt{u}.vx vd, va, x, vm va >= x see below va < i vmsle{u}.vi vd, va, i-1, vm vmslt{u}.vi vd, va, i, vm va <= i vmsle{u}.vi vd, va, i, vm va > i vmsgt{u}.vi vd, va, i, vm va >= i vmsgt{u}.vi vd, va, i-1, vm vmsge{u}.vi vd, va, i, vm va, vb vector register groups x scalar integer register i immediate | | The immediate forms of vmslt\_u\_.vi are not provided as the immediate value can be decreased by 1 and the vmsle\_u\_.vi variants used instead. The vmsle.vi range is -16 to 15, resulting in an effective vmslt.vi range of -15 to 16\. The vmsleu.vi range is 0 to 15 giving an effective vmsltu.vi range of 1 to 16 (Note,vmsltu.vi with immediate 0 is not useful as it is always false). | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | Because the 5-bit vector immediates are always sign-extended, when the high bit of the simm5 immediate is set, vmsleu.vi also supports unsigned immediate values in the range 2SEW\-16 to2SEW\-1, allowing corresponding vmsltu.vi compares against unsigned immediates in the range 2SEW\-15 to 2SEW. Note thatvmsltu.vi with immediate 2SEW is not useful as it is always true. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Similarly, `vmsge_u_.vi` is not provided and the compare is implemented using `vmsgt_u_.vi` with the immediate decremented by one. The resulting effective `vmsge.vi` range is -15 to 16, and the resulting effective `vmsgeu.vi` range is 1 to 16 (Note, `vmsgeu.vi` with immediate 0 is not useful as it is always true). | | The vmsgt forms for register scalar and immediates are provided to allow a single compare instruction to provide the correct polarity of mask value without using additional mask logical instructions. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | To reduce encoding space, the `vmsge_u_.vx` form is not directly provided, and so the `va ≥ x` case requires special treatment. | | The vmsge\_u\_.vx could potentially be encoded in a non-orthogonal way under the unused OPIVI variant of vmslt\_u\_. These would be the only instructions in OPIVI that use a scalar x register however. Alternatively, a further two funct6 encodings could be used, but these would have a different operand format (writes to mask register) than others in the same group of 8 funct6 encodings. The current PoR is to omit these instructions and to synthesize where needed as described below. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `vmsge_u_.vx` operation can be synthesized by reducing the value of `x` by 1 and using the `vmsgt_u_.vx` instruction, when it is known that this will not underflow the representation in `x`. Sequences to synthesize vmsge{u}.vx instruction va >= x, x > minimum addi t0, x, -1; vmsgt{u}.vx vd, va, t0, vm The above sequence will usually be the most efficient implementation, but assembler pseudoinstructions can be provided for cases where the range of `x` is unknown. unmasked va >= x pseudoinstruction: vmsge{u}.vx vd, va, x expansion: vmslt{u}.vx vd, va, x; vmnand.mm vd, vd, vd masked va >= x, vd != v0 pseudoinstruction: vmsge{u}.vx vd, va, x, v0.t expansion: vmslt{u}.vx vd, va, x, v0.t; vmxor.mm vd, vd, v0 masked va >= x, vd == v0 pseudoinstruction: vmsge{u}.vx vd, va, x, v0.t, vt expansion: vmslt{u}.vx vt, va, x; vmandn.mm vd, vd, vt masked va >= x, any vd pseudoinstruction: vmsge{u}.vx vd, va, x, v0.t, vt expansion: vmslt{u}.vx vt, va, x; vmandn.mm vt, v0, vt; vmandn.mm vd, vd, v0; vmor.mm vd, vt, vd The vt argument to the pseudoinstruction must name a temporary vector register that is not same as vd and which will be clobbered by the pseudoinstruction Compares effectively AND in the mask under a mask-undisturbed policy if the destination register is `v0`, e.g., # (a < b) && (b < c) in two instructions when mask-undisturbed vmslt.vv v0, va, vb # All body elements written vmslt.vv v0, vb, vc, v0.t # Only update at set mask Compares write mask registers, and so always operate under a tail-agnostic policy. #### [](#31-1-11-9-vector-integer-minmax-instructions)31.1.11.9\. Vector Integer Min/Max Instructions Signed and unsigned integer minimum and maximum instructions are supported. # Unsigned minimum vminu.vv vd, vs2, vs1, vm # Vector-vector vminu.vx vd, vs2, rs1, vm # vector-scalar # Signed minimum vmin.vv vd, vs2, vs1, vm # Vector-vector vmin.vx vd, vs2, rs1, vm # vector-scalar # Unsigned maximum vmaxu.vv vd, vs2, vs1, vm # Vector-vector vmaxu.vx vd, vs2, rs1, vm # vector-scalar # Signed maximum vmax.vv vd, vs2, vs1, vm # Vector-vector vmax.vx vd, vs2, rs1, vm # vector-scalar #### [](#31-1-11-10-vector-single-width-integer-multiply-instructions)31.1.11.10\. Vector Single-Width Integer Multiply Instructions The single-width multiply instructions perform a SEW-bit\*SEW-bit multiply to generate a 2\*SEW-bit product, then return one half of the product in the SEW-bit-wide destination. The `**mul**` versions write the low word of the product to the destination register, while the`**mulh**` versions write the high word of the product to the destination register. # Signed multiply, returning low bits of product vmul.vv vd, vs2, vs1, vm # Vector-vector vmul.vx vd, vs2, rs1, vm # vector-scalar # Signed multiply, returning high bits of product vmulh.vv vd, vs2, vs1, vm # Vector-vector vmulh.vx vd, vs2, rs1, vm # vector-scalar # Unsigned multiply, returning high bits of product vmulhu.vv vd, vs2, vs1, vm # Vector-vector vmulhu.vx vd, vs2, rs1, vm # vector-scalar # Signed(vs2)-Unsigned multiply, returning high bits of product vmulhsu.vv vd, vs2, vs1, vm # Vector-vector vmulhsu.vx vd, vs2, rs1, vm # vector-scalar | | There is no vmulhus.vx opcode to return high half of unsigned-vector \* signed-scalar product. The scalar can be splatted to a vector, then a vmulhsu.vv used. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The current vmulh\* opcodes perform simple fractional multiplies, but with no option to scale, round, and/or saturate the result. A possible future extension can consider variants of vmulh,vmulhu, vmulhsu that use the vxrm rounding mode when discarding low half of product. There is no possibility of overflow in these cases. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-11-11-vector-integer-divide-instructions)31.1.11.11\. Vector Integer Divide Instructions The divide and remainder instructions are equivalent to the RISC-V standard scalar integer multiply/divides, with the same results for extreme inputs. # Unsigned divide. vdivu.vv vd, vs2, vs1, vm # Vector-vector vdivu.vx vd, vs2, rs1, vm # vector-scalar # Signed divide vdiv.vv vd, vs2, vs1, vm # Vector-vector vdiv.vx vd, vs2, rs1, vm # vector-scalar # Unsigned remainder vremu.vv vd, vs2, vs1, vm # Vector-vector vremu.vx vd, vs2, rs1, vm # vector-scalar # Signed remainder vrem.vv vd, vs2, vs1, vm # Vector-vector vrem.vx vd, vs2, rs1, vm # vector-scalar | | The decision to include integer divide and remainder was contentious. The argument in favor is that without a standard instruction, software would have to pick some algorithm to perform the operation, which would likely perform poorly on some microarchitectures versus others. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | There is no instruction to perform a "scalar divide by vector" operation. | | ---------------------------------------------------------------------------- | #### [](#31-1-11-12-vector-widening-integer-multiply-instructions)31.1.11.12\. Vector Widening Integer Multiply Instructions The widening integer multiply instructions return the full 2\*SEW-bit product from an SEW-bit\*SEW-bit multiply. # Widening signed-integer multiply vwmul.vv vd, vs2, vs1, vm # vector-vector vwmul.vx vd, vs2, rs1, vm # vector-scalar # Widening unsigned-integer multiply vwmulu.vv vd, vs2, vs1, vm # vector-vector vwmulu.vx vd, vs2, rs1, vm # vector-scalar # Widening signed(vs2)-unsigned integer multiply vwmulsu.vv vd, vs2, vs1, vm # vector-vector vwmulsu.vx vd, vs2, rs1, vm # vector-scalar #### [](#31-1-11-13-vector-single-width-integer-multiply-add-instructions)31.1.11.13\. Vector Single-Width Integer Multiply-Add Instructions The integer multiply-add instructions are destructive and are provided in two forms, one that overwrites the addend or minuend (`vmacc`, `vnmsac`) and one that overwrites the first multiplicand (`vmadd`, `vnmsub`). The low half of the product is added or subtracted from the third operand. | | sac is intended to be read as "subtract from accumulator". The opcode is vnmsac to match the (unfortunately counterintuitive) floating-point fnmsub instruction definition. Similarly for thevnmsub opcode. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | # Integer multiply-add, overwrite addend vmacc.vv vd, vs1, vs2, vm # vd[i] = (vs1[i] * vs2[i]) + vd[i] vmacc.vx vd, rs1, vs2, vm # vd[i] = (x[rs1] * vs2[i]) + vd[i] # Integer multiply-sub, overwrite minuend vnmsac.vv vd, vs1, vs2, vm # vd[i] = -(vs1[i] * vs2[i]) + vd[i] vnmsac.vx vd, rs1, vs2, vm # vd[i] = -(x[rs1] * vs2[i]) + vd[i] # Integer multiply-add, overwrite multiplicand vmadd.vv vd, vs1, vs2, vm # vd[i] = (vs1[i] * vd[i]) + vs2[i] vmadd.vx vd, rs1, vs2, vm # vd[i] = (x[rs1] * vd[i]) + vs2[i] # Integer multiply-sub, overwrite multiplicand vnmsub.vv vd, vs1, vs2, vm # vd[i] = -(vs1[i] * vd[i]) + vs2[i] vnmsub.vx vd, rs1, vs2, vm # vd[i] = -(x[rs1] * vd[i]) + vs2[i] #### [](#31-1-11-14-vector-widening-integer-multiply-add-instructions)31.1.11.14\. Vector Widening Integer Multiply-Add Instructions The widening integer multiply-add instructions add the full 2\*SEW-bit product from a SEW-bit\*SEW-bit multiply to a 2\*SEW-bit value and produce a 2\*SEW-bit result. All combinations of signed and unsigned multiply operands are supported. # Widening unsigned-integer multiply-add, overwrite addend vwmaccu.vv vd, vs1, vs2, vm # vd[i] = (vs1[i] * vs2[i]) + vd[i] vwmaccu.vx vd, rs1, vs2, vm # vd[i] = (x[rs1] * vs2[i]) + vd[i] # Widening signed-integer multiply-add, overwrite addend vwmacc.vv vd, vs1, vs2, vm # vd[i] = (vs1[i] * vs2[i]) + vd[i] vwmacc.vx vd, rs1, vs2, vm # vd[i] = (x[rs1] * vs2[i]) + vd[i] # Widening signed-unsigned-integer multiply-add, overwrite addend vwmaccsu.vv vd, vs1, vs2, vm # vd[i] = (signed(vs1[i]) * unsigned(vs2[i])) + vd[i] vwmaccsu.vx vd, rs1, vs2, vm # vd[i] = (signed(x[rs1]) * unsigned(vs2[i])) + vd[i] # Widening unsigned-signed-integer multiply-add, overwrite addend vwmaccus.vx vd, rs1, vs2, vm # vd[i] = (unsigned(x[rs1]) * signed(vs2[i])) + vd[i] #### [](#31-1-11-15-vector-integer-merge-instructions)31.1.11.15\. Vector Integer Merge Instructions The vector integer merge instructions combine two source operands based on a mask. Unlike regular arithmetic instructions, the merge operates on all body elements (i.e., the set of elements from`vstart` up to the current vector length in `vl`). The `vmerge` instructions are encoded as masked instructions (`vm=0`). The instructions combine two sources as follows. At elements where the mask value is zero, the first operand is copied to the destination element, otherwise the second operand is copied to the destination element. The first operand is always a vector register group specified by `vs2`. The second operand is a vector register group specified by `vs1` or a scalar `x` register specified by `rs1` or a 5-bit sign-extended immediate. vmerge.vvm vd, vs2, vs1, v0 # vd[i] = v0.mask[i] ? vs1[i] : vs2[i] vmerge.vxm vd, vs2, rs1, v0 # vd[i] = v0.mask[i] ? x[rs1] : vs2[i] vmerge.vim vd, vs2, imm, v0 # vd[i] = v0.mask[i] ? imm : vs2[i] #### [](#31-1-11-16-vector-integer-move-instructions)31.1.11.16\. Vector Integer Move Instructions The vector integer move instructions copy a source operand to a vector register group. The `vmv.v.v` variant copies a vector register group, whereas the `vmv.v.x`and `vmv.v.i` variants _splat_ a scalar register or immediate to all active elements of the destination vector register group.These instructions are encoded as unmasked instructions (`vm=1`).The first operand specifier (`vs2`) must contain `v0`, and any other vector register number in `vs2` is _reserved_. vmv.v.v vd, vs1 # vd[i] = vs1[i] vmv.v.x vd, rs1 # vd[i] = x[rs1] vmv.v.i vd, imm # vd[i] = imm | | Mask values can be widened into SEW-width elements using a sequence vmv.v.i vd, 0; vmerge.vim vd, vd, 1, v0. | | --------------------------------------------------------------------------------------------------------------- | | | The vector integer move instructions share the encoding with the vector merge instructions, but with vm=1 and vs2=v0. | | ------------------------------------------------------------------------------------------------------------------------ | The form `vmv.v.v vd, vd`, which leaves body elements unchanged, can be used to indicate that the register will next be used with an EEW equal to SEW. | | Implementations that internally reorganize data according to EEW can shuffle the internal representation according to SEW. Implementations that do not internally reorganize data can dynamically elide this instruction (aside from resetting vstart to 0). | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The vmv.v.v vd, vd instruction is not a RISC-V HINT as a tail-agnostic setting may cause an architectural state change on some implementations. | | -------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#sec-vector-fixed-point)31.1.12\. Vector Fixed-Point Arithmetic Instructions The preceding set of integer arithmetic instructions is extended to support fixed-point arithmetic. A fixed-point number is a two’s-complement signed or unsigned integer interpreted as the numerator in a fraction with an implicit denominator. The fixed-point instructions are intended to be applied to the numerators; it is the responsibility of software to manage the denominators. An N-bit element can hold two’s-complement signed integers in the range -2N-1…​+2N-1\-1, and unsigned integers in the range 0 …​ +2N\-1\. The fixed-point instructions help preserve precision in narrow operands by supporting scaling and rounding, and can handle overflow by saturating results into the destination format range. | | The widening integer operations described above can also be used to avoid overflow. | | -------------------------------------------------------------------------------------- | #### [](#31-1-12-1-vector-single-width-saturating-add-and-subtract)31.1.12.1\. Vector Single-Width Saturating Add and Subtract Saturating forms of integer add and subtract are provided, for both signed and unsigned integers. If the result would overflow the destination, the result is replaced with the closest representable value, and the `vxsat` bit is set. # Saturating adds of unsigned integers. vsaddu.vv vd, vs2, vs1, vm # Vector-vector vsaddu.vx vd, vs2, rs1, vm # vector-scalar vsaddu.vi vd, vs2, imm, vm # vector-immediate # Saturating adds of signed integers. vsadd.vv vd, vs2, vs1, vm # Vector-vector vsadd.vx vd, vs2, rs1, vm # vector-scalar vsadd.vi vd, vs2, imm, vm # vector-immediate # Saturating subtract of unsigned integers. vssubu.vv vd, vs2, vs1, vm # Vector-vector vssubu.vx vd, vs2, rs1, vm # vector-scalar # Saturating subtract of signed integers. vssub.vv vd, vs2, vs1, vm # Vector-vector vssub.vx vd, vs2, rs1, vm # vector-scalar #### [](#31-1-12-2-vector-single-width-averaging-add-and-subtract)31.1.12.2\. Vector Single-Width Averaging Add and Subtract The averaging add and subtract instructions right shift the result by one bit and round off the result according to the setting in `vxrm`. Computation is performed in infinite precision before rounding and truncating.Both unsigned and signed versions are provided. For `vaaddu` and `vaadd` there can be no overflow in the result.For `vasub` and `vasubu`, overflow is ignored and the result wraps around. | | For vasub, overflow occurs only when subtracting the smallest number from the largest number under rnu or rne rounding. | | -------------------------------------------------------------------------------------------------------------------------- | # Averaging add # Averaging adds of unsigned integers. vaaddu.vv vd, vs2, vs1, vm # roundoff_unsigned(vs2[i] + vs1[i], 1) vaaddu.vx vd, vs2, rs1, vm # roundoff_unsigned(vs2[i] + x[rs1], 1) # Averaging adds of signed integers. vaadd.vv vd, vs2, vs1, vm # roundoff_signed(vs2[i] + vs1[i], 1) vaadd.vx vd, vs2, rs1, vm # roundoff_signed(vs2[i] + x[rs1], 1) # Averaging subtract # Averaging subtract of unsigned integers. vasubu.vv vd, vs2, vs1, vm # roundoff_unsigned(vs2[i] - vs1[i], 1) vasubu.vx vd, vs2, rs1, vm # roundoff_unsigned(vs2[i] - x[rs1], 1) # Averaging subtract of signed integers. vasub.vv vd, vs2, vs1, vm # roundoff_signed(vs2[i] - vs1[i], 1) vasub.vx vd, vs2, rs1, vm # roundoff_signed(vs2[i] - x[rs1], 1) #### [](#31-1-12-3-vector-single-width-fractional-multiply-with-rounding-and-saturation)31.1.12.3\. Vector Single-Width Fractional Multiply with Rounding and Saturation The signed fractional multiply instruction produces a 2\*SEW product of the two SEW inputs, then shifts the result right by SEW-1 bits, rounding these bits according to `vxrm`, then saturates the result to fit into SEW bits. If the result causes saturation, the `vxsat` bit is set. # Signed saturating and rounding fractional multiply # See vxrm description for rounding calculation vsmul.vv vd, vs2, vs1, vm # vd[i] = clip(roundoff_signed(vs2[i]*vs1[i], SEW-1)) vsmul.vx vd, vs2, rs1, vm # vd[i] = clip(roundoff_signed(vs2[i]*x[rs1], SEW-1)) | | When multiplying two N-bit signed numbers, the largest magnitude is obtained for -2N-1 \* -2N-1 producing a result +22N-2, which has a single (zero) sign bit when held in 2N bits. All other products have two sign bits in 2N bits. To retain greater precision in N result bits, the product is shifted right by one bit less than N, saturating the largest magnitude result but increasing result precision by one bit for all other products. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | We do not provide an equivalent fractional multiply where one input is unsigned, as these would retain all upper SEW bits and would not need to saturate. This operation is partly covered by thevmulhu and vmulhsu instructions, for the case where rounding is simply truncation (rdn). | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-12-4-vector-single-width-scaling-shift-instructions)31.1.12.4\. Vector Single-Width Scaling Shift Instructions These instructions shift the input value right, and round off the shifted out bits according to `vxrm`. The scaling right shifts have both zero-extending (`vssrl`) and sign-extending (`vssra`) forms. The data to be shifted is in the vector register group specified by `vs2`and the shift amount value can come from a vector register group`vs1`, a scalar integer register `rs1`, or a zero-extended 5-bit immediate. Only the low lg2(SEW) bits of the shift-amount value are used to control the shift amount. # Scaling shift right logical vssrl.vv vd, vs2, vs1, vm # vd[i] = roundoff_unsigned(vs2[i], vs1[i]) vssrl.vx vd, vs2, rs1, vm # vd[i] = roundoff_unsigned(vs2[i], x[rs1]) vssrl.vi vd, vs2, uimm, vm # vd[i] = roundoff_unsigned(vs2[i], uimm) # Scaling shift right arithmetic vssra.vv vd, vs2, vs1, vm # vd[i] = roundoff_signed(vs2[i],vs1[i]) vssra.vx vd, vs2, rs1, vm # vd[i] = roundoff_signed(vs2[i], x[rs1]) vssra.vi vd, vs2, uimm, vm # vd[i] = roundoff_signed(vs2[i], uimm) #### [](#31-1-12-5-vector-narrowing-fixed-point-clip-instructions)31.1.12.5\. Vector Narrowing Fixed-Point Clip Instructions The `vnclip` instructions are used to pack a fixed-point value into a narrower destination. The instructions support rounding, scaling, and saturation into the final destination format. The source data is in the vector register group specified by `vs2`. The scaling shift amount value can come from a vector register group `vs1`, a scalar integer register `rs1`, or a zero-extended 5-bit immediate. The low lg2(2\*SEW) bits of the vector or scalar shift-amount value (e.g., the low 6 bits for a SEW=64-bit to SEW=32-bit narrowing operation) are used to control the right shift amount, which provides the scaling. # Narrowing unsigned clip # SEW 2*SEW SEW vnclipu.wv vd, vs2, vs1, vm # vd[i] = clip(roundoff_unsigned(vs2[i], vs1[i])) vnclipu.wx vd, vs2, rs1, vm # vd[i] = clip(roundoff_unsigned(vs2[i], x[rs1])) vnclipu.wi vd, vs2, uimm, vm # vd[i] = clip(roundoff_unsigned(vs2[i], uimm)) # Narrowing signed clip vnclip.wv vd, vs2, vs1, vm # vd[i] = clip(roundoff_signed(vs2[i], vs1[i])) vnclip.wx vd, vs2, rs1, vm # vd[i] = clip(roundoff_signed(vs2[i], x[rs1])) vnclip.wi vd, vs2, uimm, vm # vd[i] = clip(roundoff_signed(vs2[i], uimm)) For `vnclipu`/`vnclip`, the rounding mode is specified in the `vxrm`CSR. Rounding occurs around the least-significant bit of the destination and before saturation. For `vnclipu`, the shifted rounded source value is treated as an unsigned integer and saturates if the result would overflow the destination viewed as an unsigned integer. | | There is no single instruction that can saturate a signed value into an unsigned destination. A sequence of two vector instructions that first removes negative numbers by performing a max against 0 using vmax then clips the resulting unsigned value into the destination using vnclipu can be used if setting vxsat value for negative numbers is not required. A vsetvli is required between these two instructions to change SEW. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For `vnclip`, the shifted rounded source value is treated as a signed integer and saturates if the result would overflow the destination viewed as a signed integer. If any destination element is saturated, the `vxsat` bit is set in the`vxsat` register. ### [](#sec-vector-float)31.1.13\. Vector Floating-Point Instructions The standard vector floating-point instructions treat elements as IEEE-754/2008-compatible values. If the EEW of a vector floating-point operand does not correspond to a supported IEEE floating-point type, the instruction encoding is reserved. | | Whether floating-point is supported, and for which element widths, is determined by the specific vector extension. The current set of extensions include support for 32-bit and 64-bit floating-point values. When 16-bit and 128-bit element widths are added, they will be also be treated as IEEE-754/2008-compatible values. Other floating-point formats may be supported in future extensions. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Vector floating-point instructions require the presence of base scalar floating-point extensions corresponding to the supported vector floating-point element widths. | | In particular, future vector extensions supporting 16-bit half-precision floating-point values will also require some scalar half-precision floating-point support. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If the floating-point unit status field `mstatus.FS` is `Off` then any attempt to execute a vector floating-point instruction will raise an illegal-instruction exception. Any vector floating-point instruction that modifies any floating-point extension state (i.e., floating-point CSRs or `f` registers) must set `mstatus.FS` to `Dirty`. If the hypervisor extension is implemented and V=1, the `vsstatus.FS` field is additionally in effect for vector floating-point instructions. If`vsstatus.FS` or `mstatus.FS` is `Off` then any attempt to execute a vector floating-point instruction will raise an illegal-instruction exception. Any vector floating-point instruction that modifies any floating-point extension state (i.e., floating-point CSRs or `f` registers) must set both `mstatus.FS` and `vsstatus.FS` to `Dirty`. The vector floating-point instructions have the same behavior as the scalar floating-point instructions with regard to NaNs. Scalar values for floating-point vector-scalar operations are sourced as described in [31.1.10.1\. Vector Arithmetic Instruction encoding](#sec-arithmetic-encoding). #### [](#31-1-13-1-vector-floating-point-exception-flags)31.1.13.1\. Vector Floating-Point Exception Flags A vector floating-point exception at any active floating-point element sets the standard FP exception flags in the `fflags` register. Inactive elements do not set FP exception flags. #### [](#31-1-13-2-vector-single-width-floating-point-addsubtract-instructions)31.1.13.2\. Vector Single-Width Floating-Point Add/Subtract Instructions # Floating-point add vfadd.vv vd, vs2, vs1, vm # Vector-vector vfadd.vf vd, vs2, rs1, vm # vector-scalar # Floating-point subtract vfsub.vv vd, vs2, vs1, vm # Vector-vector vfsub.vf vd, vs2, rs1, vm # Vector-scalar vd[i] = vs2[i] - f[rs1] vfrsub.vf vd, vs2, rs1, vm # Scalar-vector vd[i] = f[rs1] - vs2[i] #### [](#31-1-13-3-vector-widening-floating-point-addsubtract-instructions)31.1.13.3\. Vector Widening Floating-Point Add/Subtract Instructions # Widening FP add/subtract, 2*SEW = SEW +/- SEW vfwadd.vv vd, vs2, vs1, vm # vector-vector vfwadd.vf vd, vs2, rs1, vm # vector-scalar vfwsub.vv vd, vs2, vs1, vm # vector-vector vfwsub.vf vd, vs2, rs1, vm # vector-scalar # Widening FP add/subtract, 2*SEW = 2*SEW +/- SEW vfwadd.wv vd, vs2, vs1, vm # vector-vector vfwadd.wf vd, vs2, rs1, vm # vector-scalar vfwsub.wv vd, vs2, vs1, vm # vector-vector vfwsub.wf vd, vs2, rs1, vm # vector-scalar #### [](#31-1-13-4-vector-single-width-floating-point-multiplydivide-instructions)31.1.13.4\. Vector Single-Width Floating-Point Multiply/Divide Instructions # Floating-point multiply vfmul.vv vd, vs2, vs1, vm # Vector-vector vfmul.vf vd, vs2, rs1, vm # vector-scalar # Floating-point divide vfdiv.vv vd, vs2, vs1, vm # Vector-vector vfdiv.vf vd, vs2, rs1, vm # vector-scalar # Reverse floating-point divide vector = scalar / vector vfrdiv.vf vd, vs2, rs1, vm # scalar-vector, vd[i] = f[rs1]/vs2[i] #### [](#31-1-13-5-vector-widening-floating-point-multiply)31.1.13.5\. Vector Widening Floating-Point Multiply # Widening floating-point multiply vfwmul.vv vd, vs2, vs1, vm # vector-vector vfwmul.vf vd, vs2, rs1, vm # vector-scalar #### [](#31-1-13-6-vector-single-width-floating-point-fused-multiply-add-instructions)31.1.13.6\. Vector Single-Width Floating-Point Fused Multiply-Add Instructions All four varieties of fused multiply-add are provided, and in two destructive forms that overwrite one of the operands, either the addend or the first multiplicand. # FP multiply-accumulate, overwrites addend vfmacc.vv vd, vs1, vs2, vm # vd[i] = (vs1[i] * vs2[i]) + vd[i] vfmacc.vf vd, rs1, vs2, vm # vd[i] = (f[rs1] * vs2[i]) + vd[i] # FP negate-(multiply-accumulate), overwrites subtrahend vfnmacc.vv vd, vs1, vs2, vm # vd[i] = -(vs1[i] * vs2[i]) - vd[i] vfnmacc.vf vd, rs1, vs2, vm # vd[i] = -(f[rs1] * vs2[i]) - vd[i] # FP multiply-subtract-accumulator, overwrites subtrahend vfmsac.vv vd, vs1, vs2, vm # vd[i] = (vs1[i] * vs2[i]) - vd[i] vfmsac.vf vd, rs1, vs2, vm # vd[i] = (f[rs1] * vs2[i]) - vd[i] # FP negate-(multiply-subtract-accumulator), overwrites minuend vfnmsac.vv vd, vs1, vs2, vm # vd[i] = -(vs1[i] * vs2[i]) + vd[i] vfnmsac.vf vd, rs1, vs2, vm # vd[i] = -(f[rs1] * vs2[i]) + vd[i] # FP multiply-add, overwrites multiplicand vfmadd.vv vd, vs1, vs2, vm # vd[i] = (vs1[i] * vd[i]) + vs2[i] vfmadd.vf vd, rs1, vs2, vm # vd[i] = (f[rs1] * vd[i]) + vs2[i] # FP negate-(multiply-add), overwrites multiplicand vfnmadd.vv vd, vs1, vs2, vm # vd[i] = -(vs1[i] * vd[i]) - vs2[i] vfnmadd.vf vd, rs1, vs2, vm # vd[i] = -(f[rs1] * vd[i]) - vs2[i] # FP multiply-sub, overwrites multiplicand vfmsub.vv vd, vs1, vs2, vm # vd[i] = (vs1[i] * vd[i]) - vs2[i] vfmsub.vf vd, rs1, vs2, vm # vd[i] = (f[rs1] * vd[i]) - vs2[i] # FP negate-(multiply-sub), overwrites multiplicand vfnmsub.vv vd, vs1, vs2, vm # vd[i] = -(vs1[i] * vd[i]) + vs2[i] vfnmsub.vf vd, rs1, vs2, vm # vd[i] = -(f[rs1] * vd[i]) + vs2[i] | | While we considered using the two unused rounding modes in the scalar FP FMA encoding to provide a few non-destructive FMAs, these would complicate microarchitectures by being the only maskable operation with three inputs and separate output. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-13-7-vector-widening-floating-point-fused-multiply-add-instructions)31.1.13.7\. Vector Widening Floating-Point Fused Multiply-Add Instructions The widening floating-point fused multiply-add instructions all overwrite the wide addend with the result. The multiplier inputs are all SEW wide, while the addend and destination is 2\*SEW bits wide. # FP widening multiply-accumulate, overwrites addend vfwmacc.vv vd, vs1, vs2, vm # vd[i] = (vs1[i] * vs2[i]) + vd[i] vfwmacc.vf vd, rs1, vs2, vm # vd[i] = (f[rs1] * vs2[i]) + vd[i] # FP widening negate-(multiply-accumulate), overwrites addend vfwnmacc.vv vd, vs1, vs2, vm # vd[i] = -(vs1[i] * vs2[i]) - vd[i] vfwnmacc.vf vd, rs1, vs2, vm # vd[i] = -(f[rs1] * vs2[i]) - vd[i] # FP widening multiply-subtract-accumulator, overwrites addend vfwmsac.vv vd, vs1, vs2, vm # vd[i] = (vs1[i] * vs2[i]) - vd[i] vfwmsac.vf vd, rs1, vs2, vm # vd[i] = (f[rs1] * vs2[i]) - vd[i] # FP widening negate-(multiply-subtract-accumulator), overwrites addend vfwnmsac.vv vd, vs1, vs2, vm # vd[i] = -(vs1[i] * vs2[i]) + vd[i] vfwnmsac.vf vd, rs1, vs2, vm # vd[i] = -(f[rs1] * vs2[i]) + vd[i] #### [](#31-1-13-8-vector-floating-point-square-root-instruction)31.1.13.8\. Vector Floating-Point Square-Root Instruction This is a unary vector-vector instruction. # Floating-point square root vfsqrt.v vd, vs2, vm # Vector-vector square root #### [](#31-1-13-9-vector-floating-point-reciprocal-square-root-estimate-instruction)31.1.13.9\. Vector Floating-Point Reciprocal Square-Root Estimate Instruction # Floating-point reciprocal square-root estimate to 7 bits. vfrsqrt7.v vd, vs2, vm This is a unary vector-vector instruction that returns an estimate of 1/sqrt(x) accurate to 7 bits. | | An earlier draft version had used the assembler name vfrsqrte7but this was deemed to cause confusion with the e_x_ notation for element width. The earlier name can be retained as alias in tool chains for backward compatibility. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The following table describes the instruction’s behavior for all classes of floating-point inputs: | Input | Output | Exceptions raised | | ---------------- | ----------------------- | ----------------- | | \-∞ ≤ _x_ < -0.0 | canonical NaN | NV | | \-0.0 | \-∞ | DZ | | +0.0 | +∞ | DZ | | +0.0 < _x_ < +∞ | _estimate of 1/sqrt(x)_ | | | +∞ | +0.0 | | | qNaN | canonical NaN | | | sNaN | canonical NaN | NV | | | All positive normal and subnormal inputs produce normal outputs. | | ------------------------------------------------------------------- | | | The output value is independent of the dynamic rounding mode. | | ---------------------------------------------------------------- | For the non-exceptional cases, the low bit of the exponent and the six high bits of significand (after the leading one) are concatenated and used to address the following table. The output of the table becomes the seven high bits of the result significand (after the leading one); the remainder of the result significand is zero. Subnormal inputs are normalized and the exponent adjusted appropriately before the lookup. The output exponent is chosen to make the result approximate the reciprocal of the square root of the argument. More precisely, the result is computed as follows. Let the normalized input exponent be equal to the input exponent if the input is normal, or 0 minus the number of leading zeros in the significand otherwise. If the input is subnormal, the normalized input significand is given by shifting the input significand left by 1 minus the normalized input exponent, discarding the leading 1 bit. The output exponent equals floor((3\*B - 1 - the normalized input exponent) / 2), where B is the exponent bias. The output sign equals the input sign. The following table gives the seven MSBs of the output significand as a function of the LSB of the normalized input exponent and the six MSBs of the normalized input significand; the other bits of the output significand are zero. __Table 16\. vfrsqrt7.v common-case lookup table contents__ | exp\[0\] | sig\[MSB -: 6\] | sig\_out\[MSB -: 7\] | | -------- | --------------- | -------------------- | | 0 | 0 | 52 | | 0 | 1 | 51 | | 0 | 2 | 50 | | 0 | 3 | 48 | | 0 | 4 | 47 | | 0 | 5 | 46 | | 0 | 6 | 44 | | 0 | 7 | 43 | | 0 | 8 | 42 | | 0 | 9 | 41 | | 0 | 10 | 40 | | 0 | 11 | 39 | | 0 | 12 | 38 | | 0 | 13 | 36 | | 0 | 14 | 35 | | 0 | 15 | 34 | | 0 | 16 | 33 | | 0 | 17 | 32 | | 0 | 18 | 31 | | 0 | 19 | 30 | | 0 | 20 | 30 | | 0 | 21 | 29 | | 0 | 22 | 28 | | 0 | 23 | 27 | | 0 | 24 | 26 | | 0 | 25 | 25 | | 0 | 26 | 24 | | 0 | 27 | 23 | | 0 | 28 | 23 | | 0 | 29 | 22 | | 0 | 30 | 21 | | 0 | 31 | 20 | | 0 | 32 | 19 | | 0 | 33 | 19 | | 0 | 34 | 18 | | 0 | 35 | 17 | | 0 | 36 | 16 | | 0 | 37 | 16 | | 0 | 38 | 15 | | 0 | 39 | 14 | | 0 | 40 | 14 | | 0 | 41 | 13 | | 0 | 42 | 12 | | 0 | 43 | 12 | | 0 | 44 | 11 | | 0 | 45 | 10 | | 0 | 46 | 10 | | 0 | 47 | 9 | | 0 | 48 | 9 | | 0 | 49 | 8 | | 0 | 50 | 7 | | 0 | 51 | 7 | | 0 | 52 | 6 | | 0 | 53 | 6 | | 0 | 54 | 5 | | 0 | 55 | 4 | | 0 | 56 | 4 | | 0 | 57 | 3 | | 0 | 58 | 3 | | 0 | 59 | 2 | | 0 | 60 | 2 | | 0 | 61 | 1 | | 0 | 62 | 1 | | 0 | 63 | 0 | | 1 | 0 | 127 | | 1 | 1 | 125 | | 1 | 2 | 123 | | 1 | 3 | 121 | | 1 | 4 | 119 | | 1 | 5 | 118 | | 1 | 6 | 116 | | 1 | 7 | 114 | | 1 | 8 | 113 | | 1 | 9 | 111 | | 1 | 10 | 109 | | 1 | 11 | 108 | | 1 | 12 | 106 | | 1 | 13 | 105 | | 1 | 14 | 103 | | 1 | 15 | 102 | | 1 | 16 | 100 | | 1 | 17 | 99 | | 1 | 18 | 97 | | 1 | 19 | 96 | | 1 | 20 | 95 | | 1 | 21 | 93 | | 1 | 22 | 92 | | 1 | 23 | 91 | | 1 | 24 | 90 | | 1 | 25 | 88 | | 1 | 26 | 87 | | 1 | 27 | 86 | | 1 | 28 | 85 | | 1 | 29 | 84 | | 1 | 30 | 83 | | 1 | 31 | 82 | | 1 | 32 | 80 | | 1 | 33 | 79 | | 1 | 34 | 78 | | 1 | 35 | 77 | | 1 | 36 | 76 | | 1 | 37 | 75 | | 1 | 38 | 74 | | 1 | 39 | 73 | | 1 | 40 | 72 | | 1 | 41 | 71 | | 1 | 42 | 70 | | 1 | 43 | 70 | | 1 | 44 | 69 | | 1 | 45 | 68 | | 1 | 46 | 67 | | 1 | 47 | 66 | | 1 | 48 | 65 | | 1 | 49 | 64 | | 1 | 50 | 63 | | 1 | 51 | 63 | | 1 | 52 | 62 | | 1 | 53 | 61 | | 1 | 54 | 60 | | 1 | 55 | 59 | | 1 | 56 | 59 | | 1 | 57 | 58 | | 1 | 58 | 57 | | 1 | 59 | 56 | | 1 | 60 | 56 | | 1 | 61 | 55 | | 1 | 62 | 54 | | 1 | 63 | 53 | | | For example, when SEW=32, vfrsqrt7(0x00718abc (≈ 1.043e-38)) = 0x5f080000 (≈ 9.800e18), and vfrsqrt7(0x7f765432 (≈ 3.274e38)) = 0x1f820000 (≈ 5.506e-20). | | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | The 7 bit accuracy was chosen as it requires 0,1,2,3 Newton-Raphson iterations to converge to close to bfloat16, FP16, FP32, FP64 accuracy respectively. Future instructions can be defined with greater estimate accuracy. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#31-1-13-10-vector-floating-point-reciprocal-estimate-instruction)31.1.13.10\. Vector Floating-Point Reciprocal Estimate Instruction # Floating-point reciprocal estimate to 7 bits. vfrec7.v vd, vs2, vm | | An earlier draft version had used the assembler name vfrece7but this was deemed to cause confusion with e_x_ notation for element width. The earlier name can be retained as alias in tool chains for backward compatibility. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | This is a unary vector-vector instruction that returns an estimate of 1/x accurate to 7 bits. The following table describes the instruction’s behavior for all classes of floating-point inputs, where _B_ is the exponent bias: | Input (_x_) | Rounding Mode | Output (_y_ ≈ _1/x_) | Exceptions raised | | ---------------------------------------------- | ------------- | ----------------------------------------------- | ----------------- | | \-∞ | _any_ | \-0.0 | | | \-2B+1 < _x_ ≤ -2B (normal) | _any_ | \-2\-(B+1) ≥ _y_ \> -2\-B (subnormal, sig=01…​) | | | \-2B < _x_ ≤ -2B-1 (normal) | _any_ | \-2\-B ≥ _y_ \> -2\-B+1 (subnormal, sig=1…​) | | | \-2B-1 < _x_ ≤ -2\-B+1 (normal) | _any_ | \-2\-B+1 ≥ _y_ \> -2B-1 (normal) | | | \-2\-B+1 < _x_ ≤ -2\-B (subnormal, sig=1…​) | _any_ | \-2B-1 ≥ _y_ \> -2B (normal) | | | \-2\-B < _x_ ≤ -2\-(B+1) (subnormal, sig=01…​) | _any_ | \-2B ≥ _y_ \> -2B+1 (normal) | | | \-2\-(B+1) < _x_ < -0.0 (subnormal, sig=00…​) | RUP, RTZ | greatest-mag. negative finite value | NX, OF | | \-2\-(B+1) < _x_ < -0.0 (subnormal, sig=00…​) | RDN, RNE, RMM | \-∞ | NX, OF | | \-0.0 | _any_ | \-∞ | DZ | | +0.0 | _any_ | +∞ | DZ | | +0.0 < _x_ < 2\-(B+1) (subnormal, sig=00…​) | RUP, RNE, RMM | +∞ | NX, OF | | +0.0 < _x_ < 2\-(B+1) (subnormal, sig=00…​) | RDN, RTZ | greatest finite value | NX, OF | | 2\-(B+1) ≤ _x_ < 2\-B (subnormal, sig=01…​) | _any_ | 2B+1 \> _y_ ≥ 2B (normal) | | | 2\-B ≤ _x_ < 2\-B+1 (subnormal, sig=1…​) | _any_ | 2B \> _y_ ≥ 2B-1 (normal) | | | 2\-B+1 ≤ _x_ < 2B-1 (normal) | _any_ | 2B-1 \> _y_ ≥ 2\-B+1 (normal) | | | 2B-1 ≤ _x_ < 2B (normal) | _any_ | 2\-B+1 \> _y_ ≥ 2\-B (subnormal, sig=1…​) | | | 2B ≤ _x_ < 2B+1 (normal) | _any_ | 2\-B \> _y_ ≥ 2\-(B+1) (subnormal, sig=01…​) | | | +∞ | _any_ | +0.0 | | | qNaN | _any_ | canonical NaN | | | sNaN | _any_ | canonical NaN | NV | | | Subnormal inputs with magnitude at least 2\-(B+1) produce normal outputs; other subnormal inputs produce infinite outputs. Normal inputs with magnitude at least 2B-1 produce subnormal outputs; other normal inputs produce normal outputs. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The output value depends on the dynamic rounding mode when the overflow exception is raised. | | ----------------------------------------------------------------------------------------------- | For the non-exceptional cases, the seven high bits of significand (after the leading one) are used to address the following table. The output of the table becomes the seven high bits of the result significand (after the leading one); the remainder of the result significand is zero. Subnormal inputs are normalized and the exponent adjusted appropriately before the lookup. The output exponent is chosen to make the result approximate the reciprocal of the argument, and subnormal outputs are denormalized accordingly. More precisely, the result is computed as follows. Let the normalized input exponent be equal to the input exponent if the input is normal, or 0 minus the number of leading zeros in the significand otherwise. The normalized output exponent equals (2\*B - 1 - the normalized input exponent). If the normalized output exponent is outside the range \[-1, 2\*B\], the result corresponds to one of the exceptional cases in the table above. If the input is subnormal, the normalized input significand is given by shifting the input significand left by 1 minus the normalized input exponent, discarding the leading 1 bit. Otherwise, the normalized input significand equals the input significand. The following table gives the seven MSBs of the normalized output significand as a function of the seven MSBs of the normalized input significand; the other bits of the normalized output significand are zero. __Table 17\. vfrec7.v common-case lookup table contents__ | sig\[MSB -: 7\] | sig\_out\[MSB -: 7\] | | --------------- | -------------------- | | 0 | 127 | | 1 | 125 | | 2 | 123 | | 3 | 121 | | 4 | 119 | | 5 | 117 | | 6 | 116 | | 7 | 114 | | 8 | 112 | | 9 | 110 | | 10 | 109 | | 11 | 107 | | 12 | 105 | | 13 | 104 | | 14 | 102 | | 15 | 100 | | 16 | 99 | | 17 | 97 | | 18 | 96 | | 19 | 94 | | 20 | 93 | | 21 | 91 | | 22 | 90 | | 23 | 88 | | 24 | 87 | | 25 | 85 | | 26 | 84 | | 27 | 83 | | 28 | 81 | | 29 | 80 | | 30 | 79 | | 31 | 77 | | 32 | 76 | | 33 | 75 | | 34 | 74 | | 35 | 72 | | 36 | 71 | | 37 | 70 | | 38 | 69 | | 39 | 68 | | 40 | 66 | | 41 | 65 | | 42 | 64 | | 43 | 63 | | 44 | 62 | | 45 | 61 | | 46 | 60 | | 47 | 59 | | 48 | 58 | | 49 | 57 | | 50 | 56 | | 51 | 55 | | 52 | 54 | | 53 | 53 | | 54 | 52 | | 55 | 51 | | 56 | 50 | | 57 | 49 | | 58 | 48 | | 59 | 47 | | 60 | 46 | | 61 | 45 | | 62 | 44 | | 63 | 43 | | 64 | 42 | | 65 | 41 | | 66 | 40 | | 67 | 40 | | 68 | 39 | | 69 | 38 | | 70 | 37 | | 71 | 36 | | 72 | 35 | | 73 | 35 | | 74 | 34 | | 75 | 33 | | 76 | 32 | | 77 | 31 | | 78 | 31 | | 79 | 30 | | 80 | 29 | | 81 | 28 | | 82 | 28 | | 83 | 27 | | 84 | 26 | | 85 | 25 | | 86 | 25 | | 87 | 24 | | 88 | 23 | | 89 | 23 | | 90 | 22 | | 91 | 21 | | 92 | 21 | | 93 | 20 | | 94 | 19 | | 95 | 19 | | 96 | 18 | | 97 | 17 | | 98 | 17 | | 99 | 16 | | 100 | 15 | | 101 | 15 | | 102 | 14 | | 103 | 14 | | 104 | 13 | | 105 | 12 | | 106 | 12 | | 107 | 11 | | 108 | 11 | | 109 | 10 | | 110 | 9 | | 111 | 9 | | 112 | 8 | | 113 | 8 | | 114 | 7 | | 115 | 7 | | 116 | 6 | | 117 | 5 | | 118 | 5 | | 119 | 4 | | 120 | 4 | | 121 | 3 | | 122 | 3 | | 123 | 2 | | 124 | 2 | | 125 | 1 | | 126 | 1 | | 127 | 0 | If the normalized output exponent is 0 or -1, the result is subnormal: the output exponent is 0, and the output significand is given by concatenating a 1 bit to the left of the normalized output significand, then shifting that quantity right by 1 minus the normalized output exponent. Otherwise, the output exponent equals the normalized output exponent, and the output significand equals the normalized output significand. The output sign equals the input sign. | | For example, when SEW=32, vfrec7(0x00718abc (≈ 1.043e-38)) = 0x7e900000 (≈ 9.570e37), and vfrec7(0x7f765432 (≈ 3.274e38)) = 0x00214000 (≈ 3.053e-39). | | -------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The 7 bit accuracy was chosen as it requires 0,1,2,3 Newton-Raphson iterations to converge to close to bfloat16, FP16, FP32, FP64 accuracy respectively. Future instructions can be defined with greater estimate accuracy. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#31-1-13-11-vector-floating-point-minmax-instructions)31.1.13.11\. Vector Floating-Point MIN/MAX Instructions The vector floating-point `vfmin` and `vfmax` instructions have the same behavior as the corresponding scalar floating-point instructions in version 2.2 of the RISC-V F/D/Q extension: they perform the `minimumNumber`or `maximumNumber` operation on active elements. # Floating-point minimum vfmin.vv vd, vs2, vs1, vm # Vector-vector vfmin.vf vd, vs2, rs1, vm # vector-scalar # Floating-point maximum vfmax.vv vd, vs2, vs1, vm # Vector-vector vfmax.vf vd, vs2, rs1, vm # vector-scalar #### [](#31-1-13-12-vector-floating-point-sign-injection-instructions)31.1.13.12\. Vector Floating-Point Sign-Injection Instructions Vector versions of the scalar sign-injection instructions. The result takes all bits except the sign bit from the vector `vs2` operands. vfsgnj.vv vd, vs2, vs1, vm # Vector-vector vfsgnj.vf vd, vs2, rs1, vm # vector-scalar vfsgnjn.vv vd, vs2, vs1, vm # Vector-vector vfsgnjn.vf vd, vs2, rs1, vm # vector-scalar vfsgnjx.vv vd, vs2, vs1, vm # Vector-vector vfsgnjx.vf vd, vs2, rs1, vm # vector-scalar | | A vector of floating-point values can be negated using a sign-injection instruction with both source operands set to the same vector operand. An assembly pseudoinstruction is provided: vfneg.v vd,vs \= vfsgnjn.vv vd,vs,vs. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The absolute value of a vector of floating-point elements can be calculated using a sign-injection instruction with both source operands set to the same vector operand. An assembly pseudoinstruction is provided: vfabs.v vd,vs \= vfsgnjx.vv vd,vs,vs. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#31-1-13-13-vector-floating-point-compare-instructions)31.1.13.13\. Vector Floating-Point Compare Instructions These vector FP compare instructions compare two source operands and write the comparison result to a mask register. The destination mask vector is always held in a single vector register, with a layout of elements as described in [31.1.4.5\. Mask Register Layout](#sec-mask-register-layout). The destination mask vector register may be the same as the source vector mask register (`v0`). Compares write mask registers, and so always operate under a tail-agnostic policy. The compare instructions follow the semantics of the scalar floating-point compare instructions. `vmfeq` and `vmfne` raise the invalid operation exception only on signaling NaN inputs. `vmflt`, `vmfle`, `vmfgt`, and `vmfge` raise the invalid operation exception on both signaling and quiet NaN inputs. `vmfne` writes 1 to the destination element when either operand is NaN, whereas the other compares write 0 when either operand is NaN. # Compare equal vmfeq.vv vd, vs2, vs1, vm # Vector-vector vmfeq.vf vd, vs2, rs1, vm # vector-scalar # Compare not equal vmfne.vv vd, vs2, vs1, vm # Vector-vector vmfne.vf vd, vs2, rs1, vm # vector-scalar # Compare less than vmflt.vv vd, vs2, vs1, vm # Vector-vector vmflt.vf vd, vs2, rs1, vm # vector-scalar # Compare less than or equal vmfle.vv vd, vs2, vs1, vm # Vector-vector vmfle.vf vd, vs2, rs1, vm # vector-scalar # Compare greater than vmfgt.vf vd, vs2, rs1, vm # vector-scalar # Compare greater than or equal vmfge.vf vd, vs2, rs1, vm # vector-scalar Comparison Assembler Mapping Assembler pseudoinstruction va < vb vmflt.vv vd, va, vb, vm va <= vb vmfle.vv vd, va, vb, vm va > vb vmflt.vv vd, vb, va, vm vmfgt.vv vd, va, vb, vm va >= vb vmfle.vv vd, vb, va, vm vmfge.vv vd, va, vb, vm va < f vmflt.vf vd, va, f, vm va <= f vmfle.vf vd, va, f, vm va > f vmfgt.vf vd, va, f, vm va >= f vmfge.vf vd, va, f, vm va, vb vector register groups f scalar floating-point register | | Providing all forms is necessary to correctly handle unordered compares for NaNs. | | ------------------------------------------------------------------------------------ | | | C99 floating-point quiet compares can be implemented by masking the signaling compares when either input is NaN, as follows. When the comparand is a non-NaN constant, the middle two instructions can be omitted. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | # Example of implementing isgreater() vmfeq.vv v0, va, va # Only set where A is not NaN. vmfeq.vv v1, vb, vb # Only set where B is not NaN. vmand.mm v0, v0, v1 # Only set where A and B are ordered, vmfgt.vv v0, va, vb, v0.t # so only set flags on ordered values. | | In the above sequence, it is tempting to mask the second vmfeqinstruction and remove the vmand instruction, but this more efficient sequence incorrectly fails to raise the invalid exception when an element of va contains a quiet NaN and the corresponding element invb contains a signaling NaN. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-13-14-vector-floating-point-classify-instruction)31.1.13.14\. Vector Floating-Point Classify Instruction This is a unary vector-vector instruction that operates in the same way as the scalar classify instruction. vfclass.v vd, vs2, vm # Vector-vector The 10-bit mask produced by this instruction is placed in the least-significant bits of the result elements. The upper (SEW-10) bits of the result are filled with zeros. The instruction is only defined for SEW=16b and above, so the result will always fit in the destination elements. #### [](#31-1-13-15-vector-floating-point-merge-instruction)31.1.13.15\. Vector Floating-Point Merge Instruction A vector-scalar floating-point merge instruction is provided, whichoperates on all body elements from `vstart` up to the current vector length in `vl` regardless of mask value. The `vfmerge.vfm` instruction is encoded as a masked instruction (`vm=0`).At elements where the mask value is zero, the first vector operand is copied to the destination element, otherwise a scalar floating-point register value is copied to the destination element. vfmerge.vfm vd, vs2, rs1, v0 # vd[i] = v0.mask[i] ? f[rs1] : vs2[i] #### [](#31-1-13-16-vector-floating-point-move-instruction)31.1.13.16\. Vector Floating-Point Move Instruction The vector floating-point move instruction _splats_ a floating-point scalar operand to a vector register group. The instruction copies a scalar `f` register value to all active elements of a vector register group. This instruction is encoded as an unmasked instruction (`vm=1`).The instruction must have the `vs2` field set to `v0`, with all other values for `vs2` reserved. vfmv.v.f vd, rs1 # vd[i] = f[rs1] | | The vfmv.v.f instruction shares the encoding with the vfmerge.vfminstruction, but with vm=1 and vs2=v0. | | ---------------------------------------------------------------------------------------------------------- | #### [](#31-1-13-17-single-width-floating-pointinteger-type-convert-instructions)31.1.13.17\. Single-Width Floating-Point/Integer Type-Convert Instructions Conversion operations are provided to convert to and from floating-point values and unsigned and signed integers, where both source and destination are SEW wide. vfcvt.xu.f.v vd, vs2, vm # Convert float to unsigned integer. vfcvt.x.f.v vd, vs2, vm # Convert float to signed integer. vfcvt.rtz.xu.f.v vd, vs2, vm # Convert float to unsigned integer, truncating. vfcvt.rtz.x.f.v vd, vs2, vm # Convert float to signed integer, truncating. vfcvt.f.xu.v vd, vs2, vm # Convert unsigned integer to float. vfcvt.f.x.v vd, vs2, vm # Convert signed integer to float. The conversions follow the same rules on exceptional conditions as the scalar conversion instructions. The conversions use the dynamic rounding mode in `frm`, except for the `rtz`variants, which round towards zero. | | The rtz variants are provided to accelerate truncating conversions from floating-point to integer, as is common in languages like C and Java. | | ------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#31-1-13-18-widening-floating-pointinteger-type-convert-instructions)31.1.13.18\. Widening Floating-Point/Integer Type-Convert Instructions A set of conversion instructions is provided to convert between narrower integer and floating-point datatypes to a type of twice the width. vfwcvt.xu.f.v vd, vs2, vm # Convert float to double-width unsigned integer. vfwcvt.x.f.v vd, vs2, vm # Convert float to double-width signed integer. vfwcvt.rtz.xu.f.v vd, vs2, vm # Convert float to double-width unsigned integer, truncating. vfwcvt.rtz.x.f.v vd, vs2, vm # Convert float to double-width signed integer, truncating. vfwcvt.f.xu.v vd, vs2, vm # Convert unsigned integer to double-width float. vfwcvt.f.x.v vd, vs2, vm # Convert signed integer to double-width float. vfwcvt.f.f.v vd, vs2, vm # Convert single-width float to double-width float. These instructions have the same constraints on vector register overlap as other widening instructions (see [31.1.10.2\. Widening Vector Arithmetic Instructions](#sec-widening)). | | A double-width IEEE floating-point value can always represent a single-width integer exactly. | | ------------------------------------------------------------------------------------------------ | | | A double-width IEEE floating-point value can always represent a single-width IEEE floating-point value exactly. | | ------------------------------------------------------------------------------------------------------------------ | | | A full set of floating-point widening conversions is not supported as single instructions, but any widening conversion can be implemented as several doubling steps with equivalent results and no additional exception flags raised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-13-19-narrowing-floating-pointinteger-type-convert-instructions)31.1.13.19\. Narrowing Floating-Point/Integer Type-Convert Instructions A set of conversion instructions is provided to convert wider integer and floating-point datatypes to a type of half the width. vfncvt.xu.f.w vd, vs2, vm # Convert double-width float to unsigned integer. vfncvt.x.f.w vd, vs2, vm # Convert double-width float to signed integer. vfncvt.rtz.xu.f.w vd, vs2, vm # Convert double-width float to unsigned integer, truncating. vfncvt.rtz.x.f.w vd, vs2, vm # Convert double-width float to signed integer, truncating. vfncvt.f.xu.w vd, vs2, vm # Convert double-width unsigned integer to float. vfncvt.f.x.w vd, vs2, vm # Convert double-width signed integer to float. vfncvt.f.f.w vd, vs2, vm # Convert double-width float to single-width float. vfncvt.rod.f.f.w vd, vs2, vm # Convert double-width float to single-width float, # rounding towards odd. These instructions have the same constraints on vector register overlap as other narrowing instructions (see [31.1.10.3\. Narrowing Vector Arithmetic Instructions](#sec-narrowing)). | | A full set of floating-point narrowing conversions is not supported as single instructions. Conversions can be implemented in a sequence of halving steps. Results are equivalently rounded and the same exception flags are raised if all but the last halving step use round-towards-odd (vfncvt.rod.f.f.w). Only the final step should use the desired rounding mode. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | For vfncvt.rod.f.f.w, a finite value that exceeds the range of the destination format is converted to the destination format’s largest finite value with the same sign. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#31-1-14-vector-reduction-operations)31.1.14\. Vector Reduction Operations Vector reduction operations take a vector register group of elements and a scalar held in element 0 of a vector register, and perform a reduction using some binary operator, to produce a scalar result in element 0 of a vector register. The scalar input and output operands are held in element 0 of a single vector register, not a vector register group, so any vector register can be the scalar source or destination of a vector reduction regardless of LMUL setting. The destination vector register can overlap the source operands, including the mask register. | | Vector reductions read and write the scalar operand and result into element 0 of a vector register instead of a scalar register to avoid a loss of decoupling with the scalar processor, and to support future polymorphic use with future types not supported in the scalar unit. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Inactive elements from the source vector register group are excluded from the reduction, but the scalar operand is always included regardless of the mask values. The other elements in the destination vector register ( 0 < index < VLEN/SEW) are considered the tail and are managed with the current tail agnostic/undisturbed policy. If `vl`\=0, no operation is performed and the destination register is not updated. | | This choice of behavior for vl\=0 reduces implementation complexity as it is consistent with other operations on vector register state. For the common case that the source and destination scalar operand are the same vector register, this behavior also produces the expected result. For the uncommon case that the source and destination scalar operand are in different vector registers, this instruction will not copy the source into the destination when vl\=0\. However, it is expected that in most of these cases it will be statically known that vl is not zero. In other cases, a check forvl\=0 will have to be added to ensure that the source scalar is copied to the destination (e.g., by explicitly setting vl\=1 and performing a register-register copy). | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Traps on vector reduction instructions are always reported with a`vstart` of 0. Vector reduction operations raise an illegal-instruction exception if `vstart` is non-zero. The assembler syntax for a reduction operation is `vredop.vs`, where the `.vs` suffix denotes the first operand is a vector register group and the second operand is a scalar stored in element 0 of a vector register. #### [](#sec-vector-integer-reduce)31.1.14.1\. Vector Single-Width Integer Reduction Instructions All operands and results of single-width reduction instructions have the same SEW width. Overflows wrap around on arithmetic sums. # Simple reductions, where [*] denotes all active elements: vredsum.vs vd, vs2, vs1, vm # vd[0] = sum( vs1[0] , vs2[*] ) vredmaxu.vs vd, vs2, vs1, vm # vd[0] = maxu( vs1[0] , vs2[*] ) vredmax.vs vd, vs2, vs1, vm # vd[0] = max( vs1[0] , vs2[*] ) vredminu.vs vd, vs2, vs1, vm # vd[0] = minu( vs1[0] , vs2[*] ) vredmin.vs vd, vs2, vs1, vm # vd[0] = min( vs1[0] , vs2[*] ) vredand.vs vd, vs2, vs1, vm # vd[0] = and( vs1[0] , vs2[*] ) vredor.vs vd, vs2, vs1, vm # vd[0] = or( vs1[0] , vs2[*] ) vredxor.vs vd, vs2, vs1, vm # vd[0] = xor( vs1[0] , vs2[*] ) #### [](#sec-vector-integer-reduce-widen)31.1.14.2\. Vector Widening Integer Reduction Instructions The unsigned `vwredsumu.vs` instruction zero-extends the SEW-wide vector elements before summing them, then adds the 2\*SEW-width scalar element, and stores the result in a 2\*SEW-width scalar element. The `vwredsum.vs` instruction sign-extends the SEW-wide vector elements before summing them. For both `vwredsumu.vs` and `vwredsum.vs`, overflows wrap around. # Unsigned sum reduction into double-width accumulator vwredsumu.vs vd, vs2, vs1, vm # 2*SEW = 2*SEW + sum(zero-extend(SEW)) # Signed sum reduction into double-width accumulator vwredsum.vs vd, vs2, vs1, vm # 2*SEW = 2*SEW + sum(sign-extend(SEW)) #### [](#sec-vector-float-reduce)31.1.14.3\. Vector Single-Width Floating-Point Reduction Instructions # Simple reductions. vfredosum.vs vd, vs2, vs1, vm # Ordered sum vfredusum.vs vd, vs2, vs1, vm # Unordered sum vfredmax.vs vd, vs2, vs1, vm # Maximum value vfredmin.vs vd, vs2, vs1, vm # Minimum value | | Older assembler mnemonic vfredsum is retained as alias for vfredusum. | | ------------------------------------------------------------------------ | ##### [](#31-1-14-3-1-vector-ordered-single-width-floating-point-sum-reduction)31.1.14.3.1\. Vector Ordered Single-Width Floating-Point Sum Reduction The `vfredosum` instruction must sum the floating-point values in element order, starting with the scalar in `vs1[0]`\--that is, it performs the computation: vd[0] = (((vs1[0] + vs2[0]) + vs2[1]) + ...) + vs2[vl-1] where each addition operates identically to the scalar floating-point instructions in terms of raising exception flags and generating or propagating special values. | | The ordered reduction supports compiler auto-vectorization, while the unordered FP sum allows for faster implementations. | | ---------------------------------------------------------------------------------------------------------------------------- | When the operation is masked (`vm=0`), the masked-off elements do not affect the result or the exception flags. | | If no elements are active, no additions are performed, so the scalar invs1\[0\] is simply copied to the destination register, without canonicalizing NaN values and without setting any exception flags. This behavior preserves the handling of NaNs, exceptions, and rounding when auto-vectorizing a scalar summation loop. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#31-1-14-3-2-vector-unordered-single-width-floating-point-sum-reduction)31.1.14.3.2\. Vector Unordered Single-Width Floating-Point Sum Reduction The unordered sum reduction instruction, `vfredusum`, provides an implementation more freedom in performing the reduction. The implementation must produce a result equivalent to a reduction tree composed of binary operator nodes, with the inputs being elements from the source vector register group (`vs2`) and the source scalar value (`vs1[0]`). Each operator in the tree accepts two inputs and produces one result. Each operator first computes an exact sum as a RISC-V scalar floating-point addition with infinite exponent range and precision, then converts this exact sum to a floating-point format with range and precision each at least as great as the element floating-point format indicated by SEW, rounding using the currently active floating-point dynamic rounding mode and raising exception flags as necessary. A different floating-point range and precision may be chosen for the result of each operator. A node where one input is derived only from elements masked-off or beyond the active vector length may either treat that input as the additive identity of the appropriate EEW or simply copy the other input to its output. The rounded result from the root node in the tree is converted (rounded again, using the dynamic rounding mode) to the standard floating-point format indicated by SEW.An implementation is allowed to add an additional additive identity to the final result. The additive identity is +0.0 when rounding down (towards -∞) or -0.0 for all other rounding modes. The reduction tree structure must be deterministic for a given value in `vtype` and `vl`. | | As a consequence of this definition, implementations need not propagate NaN payloads through the reduction tree when no elements are active. In particular, if no elements are active and the scalar input is NaN, implementations are permitted to canonicalize the NaN and, if the NaN is signaling, set the invalid exception flag. Implementations are alternatively permitted to pass through the original NaN and set no exception flags, as withvfredosum. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The vfredosum instruction is a valid implementation of thevfredusum instruction. | | ----------------------------------------------------------------------------------- | ##### [](#31-1-14-3-3-vector-single-width-floating-point-max-and-min-reductions)31.1.14.3.3\. Vector Single-Width Floating-Point Max and Min Reductions The `vfredmin` and `vfredmax` instructions reduce the scalar argument in`vs1[0]` and active elements in `vs2` using the `minimumNumber` and`maximumNumber` operations, respectively. | | Floating-point max and min reductions should return the same final value and raise the same exception flags regardless of operation order. | | --------------------------------------------------------------------------------------------------------------------------------------------- | | | If no elements are active, the scalar in vs1\[0\] is simply copied to the destination register, without canonicalizing NaN values and without setting any exception flags. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#sec-vector-float-reduce-widen)31.1.14.4\. Vector Widening Floating-Point Reduction Instructions Widening forms of the sum reductions are provided that read and write a double-width reduction result. # Simple reductions. vfwredosum.vs vd, vs2, vs1, vm # Ordered sum vfwredusum.vs vd, vs2, vs1, vm # Unordered sum | | Older assembler mnemonic vfwredsum is retained as alias for vfwredusum. | | -------------------------------------------------------------------------- | The reduction of the SEW-width elements is performed as in the single-width reduction case, with the elements in `vs2` promoted to 2\*SEW bits before adding to the 2\*SEW-bit accumulator. | | vfwredosum.vs handles inactive elements and NaN payloads analogously to vfredosum.vs; vfwredusum.vs does so analogously to vfredusum.vs. | | ------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#sec-vector-mask)31.1.15\. Vector Mask Instructions Several instructions are provided to help operate on mask values held in a vector register. #### [](#sec-mask-register-logical)31.1.15.1\. Vector Mask-Register Logical Instructions Vector mask-register logical operations operate on mask registers. Each element in a mask register is a single bit, so these instructions all operate on single vector registers regardless of the setting of the `vlmul` field in `vtype`. They do not change the value of`vlmul`. The destination vector register may be the same as either source vector register. As with other vector instructions, the elements with indices less than`vstart` are unchanged, and `vstart` is reset to zero after execution. Vector mask logical instructions are always unmasked, so there are no inactive elements, and the encodings with `vm=0` are reserved. Mask elements past `vl`, the tail elements, are always updated with a tail-agnostic policy. vmand.mm vd, vs2, vs1 # vd.mask[i] = vs2.mask[i] && vs1.mask[i] vmnand.mm vd, vs2, vs1 # vd.mask[i] = !(vs2.mask[i] && vs1.mask[i]) vmandn.mm vd, vs2, vs1 # vd.mask[i] = vs2.mask[i] && !vs1.mask[i] vmxor.mm vd, vs2, vs1 # vd.mask[i] = vs2.mask[i] ^^ vs1.mask[i] vmor.mm vd, vs2, vs1 # vd.mask[i] = vs2.mask[i] || vs1.mask[i] vmnor.mm vd, vs2, vs1 # vd.mask[i] = !(vs2.mask[i] || vs1.mask[i]) vmorn.mm vd, vs2, vs1 # vd.mask[i] = vs2.mask[i] || !vs1.mask[i] vmxnor.mm vd, vs2, vs1 # vd.mask[i] = !(vs2.mask[i] ^^ vs1.mask[i]) | | The previous assembler mnemonics vmandnot and vmornot have been changed to vmandn and vmorn to be consistent with the equivalent scalar instructions. The old vmandnot and vmornotmnemonics can be retained as assembler aliases for compatibility. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Several assembler pseudoinstructions are defined as shorthand for common uses of mask logical operations: vmmv.m vd, vs => vmand.mm vd, vs, vs # Copy mask register vmclr.m vd => vmxor.mm vd, vd, vd # Clear mask register vmset.m vd => vmxnor.mm vd, vd, vd # Set mask register vmnot.m vd, vs => vmnand.mm vd, vs, vs # Invert bits | | The vmmv.m instruction was previously called vmcpy.m, but with new layout it is more consistent to name as a "mv" because bits are copied without interpretation. The vmcpy.m assembler pseudoinstruction can be retained for compatibility. For implementations that internally rearrange bits according to EEW, avmmv.m instruction with same source and destination can be used as idiom to force an internal reformat into a mask vector. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The set of eight mask logical instructions can generate any of the 16 possibly binary logical functions of the two input masks: | inputs | | | | | | ------ | - | - | - | ---- | | 0 | 0 | 1 | 1 | src1 | | 0 | 1 | 0 | 1 | src2 | | output | instruction | pseudoinstruction | | | | | ------ | ----------- | ----------------- | - | ------------------------ | ---------------- | | 0 | 0 | 0 | 0 | vmxor.mm vd, vd, vd | vmclr.m vd | | 1 | 0 | 0 | 0 | vmnor.mm vd, src1, src2 | | | 0 | 1 | 0 | 0 | vmandn.mm vd, src2, src1 | | | 1 | 1 | 0 | 0 | vmnand.mm vd, src1, src1 | vmnot.m vd, src1 | | 0 | 0 | 1 | 0 | vmandn.mm vd, src1, src2 | | | 1 | 0 | 1 | 0 | vmnand.mm vd, src2, src2 | vmnot.m vd, src2 | | 0 | 1 | 1 | 0 | vmxor.mm vd, src1, src2 | | | 1 | 1 | 1 | 0 | vmnand.mm vd, src1, src2 | | | 0 | 0 | 0 | 1 | vmand.mm vd, src1, src2 | | | 1 | 0 | 0 | 1 | vmxnor.mm vd, src1, src2 | | | 0 | 1 | 0 | 1 | vmand.mm vd, src2, src2 | vmmv.m vd, src2 | | 1 | 1 | 0 | 1 | vmorn.mm vd, src2, src1 | | | 0 | 0 | 1 | 1 | vmand.mm vd, src1, src1 | vmmv.m vd, src1 | | 1 | 0 | 1 | 1 | vmorn.mm vd, src1, src2 | | | 0 | 1 | 1 | 1 | vmor.mm vd, src1, src2 | | | 1 | 1 | 1 | 1 | vmxnor.mm vd, vd, vd | vmset.m vd | | | The vector mask logical instructions are designed to be easily fused with a following masked vector operation to effectively expand the number of predicate registers by moving values into v0 before use. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-15-2-vector-count-population-in-mask-vcpop-m)31.1.15.2\. Vector count population in mask `vcpop.m` vcpop.m rd, vs2, vm | | This instruction previously had the assembler mnemonic vpopc.mbut was renamed to be consistent with the scalar instruction. The assembler instruction alias vpopc.m is being retained for software compatibility. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The source operand is a single vector register holding mask register values as described in [31.1.4.5\. Mask Register Layout](#sec-mask-register-layout). The `vcpop.m` instruction counts the number of mask elements of the active elements of the vector source mask register that have the value 1 and writes the result to a scalar `x` register. The operation can be performed under a mask, in which case only the masked elements are counted. vcpop.m rd, vs2, v0.t # x[rd] = sum_i ( vs2.mask[i] && v0.mask[i] ) The `vcpop.m` instruction writes `x[rd]` even if `vl`\=0 (with the value 0, since no mask elements are active). Traps on `vcpop.m` are always reported with a `vstart` of 0. The`vcpop.m` instruction will raise an illegal-instruction exception if`vstart` is non-zero. #### [](#31-1-15-3-vfirst-find-first-set-mask-bit)31.1.15.3\. `vfirst` find-first-set mask bit vfirst.m rd, vs2, vm The `vfirst` instruction finds the lowest-numbered active element of the source mask vector that has the value 1 and writes that element’s index to a GPR. If no active element has the value 1, -1 is written to the GPR. | | Software can assume that any negative value (highest bit set) corresponds to no element found, as vector lengths will never reach 2(XLEN-1) on any implementation. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `vfirst.m` instruction writes `x[rd]` even if `vl`\=0 (with the value -1, since no mask elements are active). Traps on `vfirst` are always reported with a `vstart` of 0. The`vfirst` instruction will raise an illegal-instruction exception if`vstart` is non-zero. #### [](#31-1-15-4-vmsbf-m-set-before-first-mask-bit)31.1.15.4\. `vmsbf.m` set-before-first mask bit vmsbf.m vd, vs2, vm # Example 7 6 5 4 3 2 1 0 Element number 1 0 0 1 0 1 0 0 v3 contents vmsbf.m v2, v3 0 0 0 0 0 0 1 1 v2 contents 1 0 0 1 0 1 0 1 v3 contents vmsbf.m v2, v3 0 0 0 0 0 0 0 0 v2 0 0 0 0 0 0 0 0 v3 contents vmsbf.m v2, v3 1 1 1 1 1 1 1 1 v2 1 1 0 0 0 0 1 1 v0 vcontents 1 0 0 1 0 1 0 0 v3 contents vmsbf.m v2, v3, v0.t 0 1 x x x x 1 1 v2 contents The `vmsbf.m` instruction takes a mask register as input and writes results to a mask register. The instruction writes a 1 to all active mask elements before the first active source element that is a 1, then writes a 0 to that element and all following active elements. If there is no set bit in the active elements of the source vector, then all active elements in the destination are written with a 1. The tail elements in the destination mask register are updated under a tail-agnostic policy. Traps on `vmsbf.m` are always reported with a `vstart` of 0. The`vmsbf` instruction will raise an illegal-instruction exception if`vstart` is non-zero. The destination register cannot overlap the source register and, if masked, cannot overlap the mask register ('v0'). #### [](#31-1-15-5-vmsif-m-set-including-first-mask-bit)31.1.15.5\. `vmsif.m` set-including-first mask bit The vector mask set-including-first instruction is similar to set-before-first, except it also includes the element with a set bit. vmsif.m vd, vs2, vm # Example 7 6 5 4 3 2 1 0 Element number 1 0 0 1 0 1 0 0 v3 contents vmsif.m v2, v3 0 0 0 0 0 1 1 1 v2 contents 1 0 0 1 0 1 0 1 v3 contents vmsif.m v2, v3 0 0 0 0 0 0 0 1 v2 1 1 0 0 0 0 1 1 v0 vcontents 1 0 0 1 0 1 0 0 v3 contents vmsif.m v2, v3, v0.t 1 1 x x x x 1 1 v2 contents The tail elements in the destination mask register are updated under a tail-agnostic policy. Traps on `vmsif.m` are always reported with a `vstart` of 0. The`vmsif` instruction will raise an illegal-instruction exception if`vstart` is non-zero. The destination register cannot overlap the source register and, if masked, cannot overlap the mask register ('v0'). #### [](#31-1-15-6-vmsof-m-set-only-first-mask-bit)31.1.15.6\. `vmsof.m` set-only-first mask bit The vector mask set-only-first instruction is similar to set-before-first, except it only sets the first element with a bit set, if any. vmsof.m vd, vs2, vm # Example 7 6 5 4 3 2 1 0 Element number 1 0 0 1 0 1 0 0 v3 contents vmsof.m v2, v3 0 0 0 0 0 1 0 0 v2 contents 1 0 0 1 0 1 0 1 v3 contents vmsof.m v2, v3 0 0 0 0 0 0 0 1 v2 1 1 0 0 0 0 1 1 v0 vcontents 1 1 0 1 0 1 0 0 v3 contents vmsof.m v2, v3, v0.t 0 1 x x x x 0 0 v2 contents The tail elements in the destination mask register are updated under a tail-agnostic policy. Traps on `vmsof.m` are always reported with a `vstart` of 0. The`vmsof` instruction will raise an illegal-instruction exception if`vstart` is non-zero. The destination register cannot overlap the source register and, if masked, cannot overlap the mask register ('v0'). #### [](#31-1-15-7-example-using-vector-mask-instructions)31.1.15.7\. Example using vector mask instructions The following is an example of vectorizing a data-dependent exit loop. # char* strcpy(char *dst, const char* src) strcpy: mv a2, a0 # Copy dst li t0, -1 # Infinite AVL loop: vsetvli x0, t0, e8, m8, ta, ma # Max length vectors of bytes vle8ff.v v8, (a1) # Get src bytes csrr t1, vl # Get number of bytes fetched vmseq.vi v1, v8, 0 # Flag zero bytes vfirst.m a3, v1 # Zero found? add a1, a1, t1 # Bump pointer vmsif.m v0, v1 # Set mask up to and including zero byte. vse8.v v8, (a2), v0.t # Write out bytes add a2, a2, t1 # Bump pointer bltz a3, loop # Zero byte not found, so loop ret # char* strncpy(char *dst, const char* src, size_t n) strncpy: mv a3, a0 # Copy dst loop: vsetvli x0, a2, e8, m8, ta, ma # Vectors of bytes. vle8ff.v v8, (a1) # Get src bytes vmseq.vi v1, v8, 0 # Flag zero bytes csrr t1, vl # Get number of bytes fetched vfirst.m a4, v1 # Zero found? vmsbf.m v0, v1 # Set mask up to before zero byte. vse8.v v8, (a3), v0.t # Write out non-zero bytes bgez a4, zero_tail # Zero remaining bytes. sub a2, a2, t1 # Decrement count. add a3, a3, t1 # Bump dest pointer add a1, a1, t1 # Bump src pointer bnez a2, loop # Anymore? ret zero_tail: sub a2, a2, a4 # Subtract count on non-zero bytes. add a3, a3, a4 # Advance past non-zero bytes. vsetvli t1, a2, e8, m8, ta, ma # Vectors of bytes. vmv.v.i v0, 0 # Splat zero. zero_loop: vse8.v v0, (a3) # Store zero. sub a2, a2, t1 # Decrement count. add a3, a3, t1 # Bump pointer vsetvli t1, a2, e8, m8, ta, ma # Vectors of bytes. bnez a2, zero_loop # Anymore? ret #### [](#31-1-15-8-vector-iota-instruction)31.1.15.8\. Vector Iota Instruction The `viota.m` instruction reads a source vector mask register and writes to each element of the destination vector register group the sum of all the bits of elements in the mask register whose index is less than the element, e.g., a parallel prefix sum of the mask values. This instruction can be masked, in which case only the enabled elements contribute to the sum. viota.m vd, vs2, vm # Example 7 6 5 4 3 2 1 0 Element number 1 0 0 1 0 0 0 1 v2 contents viota.m v4, v2 # Unmasked 2 2 2 1 1 1 1 0 v4 result 1 1 1 0 1 0 1 1 v0 contents 1 0 0 1 0 0 0 1 v2 contents 2 3 4 5 6 7 8 9 v4 contents viota.m v4, v2, v0.t # Masked, vtype.vma=0 1 1 1 5 1 7 1 0 v4 results The result value is zero-extended to fill the destination element if SEW is wider than the result. If the result value would overflow the destination SEW, the least-significant SEW bits are retained. Traps on `viota.m` are always reported with a `vstart` of 0, andexecution is always restarted from the beginning when resuming after a trap handler. An illegal-instruction exception is raised if `vstart`is non-zero. The destination register group cannot overlap the source register and, if masked, cannot overlap the mask register (`v0`). The `viota.m` instruction can be combined with memory scatter instructions (indexed stores) to perform vector compress functions. # Compact non-zero elements from input memory array to output memory array # # size_t compact_non_zero(size_t n, const int* in, int* out) # { # size_t i; # int *p = out; # # for (i=0; i XLEN, the least-significant XLEN bits are transferred and the upper SEW-XLEN bits are ignored. If SEW < XLEN, the value is sign-extended to XLEN bits. | | vmv.x.s performs its operation even if vstart ≥ vl or vl\=0. | | --------------------------------------------------------------- | The `vmv.s.x` instruction copies the scalar integer register to element 0 of the destination vector register. If SEW < XLEN, the least-significant bits are copied and the upper XLEN-SEW bits are ignored. If SEW > XLEN, the value is sign-extended to SEW bits. The other elements in the destination vector register ( 0 < index < VLEN/SEW) are treated as tail elements using the current tail agnostic/undisturbed policy. If `vstart` ≥ `vl`, no operation is performed and the destination register is not updated. | | As a consequence, when vl\=0, no elements are updated in the destination vector register group, regardless of vstart. | | ------------------------------------------------------------------------------------------------------------------------ | The encodings corresponding to the masked versions (`vm=0`) of `vmv.x.s`and `vmv.s.x` are reserved. #### [](#sec-vector-float-move)31.1.16.2\. Floating-Point Scalar Move Instructions The floating-point scalar read/write instructions transfer a single value between a scalar `f` register and element 0 of a vector register. The instructions ignore LMUL and vector register groups. vfmv.f.s rd, vs2 # f[rd] = vs2[0] (rs1=0) vfmv.s.f vd, rs1 # vd[0] = f[rs1] (vs2=0) The `vfmv.f.s` instruction copies a single SEW-wide element from index 0 of the source vector register to a destination scalar floating-point register. | | vfmv.f.s performs its operation even if vstart ≥ vl or vl\=0. | | ---------------------------------------------------------------- | The `vfmv.s.f` instruction copies the scalar floating-point register to element 0 of the destination vector register. The other elements in the destination vector register ( 0 < index < VLEN/SEW) are treated as tail elements using the current tail agnostic/undisturbed policy. If `vstart` ≥ `vl`, no operation is performed and the destination register is not updated. | | As a consequence, when vl\=0, no elements are updated in the destination vector register group, regardless of vstart. | | ------------------------------------------------------------------------------------------------------------------------ | The encodings corresponding to the masked versions (`vm=0`) of `vfmv.f.s`and `vfmv.s.f` are reserved. #### [](#31-1-16-3-vector-slide-instructions)31.1.16.3\. Vector Slide Instructions The slide instructions move elements up and down a vector register group. | | The slide operations can be implemented much more efficiently than using the arbitrary register gather instruction. Implementations may optimize certain OFFSET values for vslideup and vslidedown. In particular, power-of-2 offsets may operate substantially faster than other offsets. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For all of the `vslideup`, `vslidedown`, `v[f]slide1up`, and`v[f]slide1down` instructions, if `vstart` ≥ `vl`, the instruction performs no operation and leaves the destination vector register unchanged. | | As a consequence, when vl\=0, no elements are updated in the destination vector register group, regardless of vstart. | | ------------------------------------------------------------------------------------------------------------------------ | The tail agnostic/undisturbed policy is followed for tail elements. The slide instructions may be masked, with mask element _i_controlling whether _destination_ element _i_ is written. The mask undisturbed/agnostic policy is followed for inactive elements. ##### [](#31-1-16-3-1-vector-slide-up-instructions)31.1.16.3.1\. Vector Slide-up Instructions vslideup.vx vd, vs2, rs1, vm # vd[i+x[rs1]] = vs2[i] vslideup.vi vd, vs2, uimm, vm # vd[i+uimm] = vs2[i] For `vslideup`, the value in `vl` specifies the maximum number of destination elements that are written. The start index (_OFFSET_) for the destination can be either specified using an unsigned integer in the`x` register specified by `rs1`, or a 5-bit immediate, zero-extended to XLEN bits. If XLEN > SEW, _OFFSET_ is _not_ truncated to SEW bits. Destination elements _OFFSET_ through `vl`\-1 are written if unmasked and if _OFFSET_ < `vl`. vslideup behavior for destination elements (vstart < vl) OFFSET is amount to slideup, either from x register or a 5-bit immediate 0 <= i < min(vl, max(vstart, OFFSET)) Unchanged max(vstart, OFFSET) <= i < vl vd[i] = vs2[i-OFFSET] if v0.mask[i] enabled vl <= i < VLMAX Follow tail policy The destination vector register group for `vslideup` cannot overlap the source vector register group, otherwise the instruction encoding is reserved. | | The non-overlap constraint avoids WAR hazards on the input vectors during execution, and enables restart with non-zerovstart. | | -------------------------------------------------------------------------------------------------------------------------------- | ##### [](#31-1-16-3-2-vector-slide-down-instructions)31.1.16.3.2\. Vector Slide-down Instructions vslidedown.vx vd, vs2, rs1, vm # vd[i] = vs2[i+x[rs1]] vslidedown.vi vd, vs2, uimm, vm # vd[i] = vs2[i+uimm] For `vslidedown`, the value in `vl` specifies the maximum number of destination elements that are written. The remaining elements past`vl` are handled according to the current tail policy ([31.1.3.4.3\. Vector Tail Agnostic and Vector Mask Agnostic vta and vma](#sec-agnostic)). The start index (_OFFSET_) for the source can be either specified using an unsigned integer in the `x` register specified by `rs1`, or a 5-bit immediate, zero-extended to XLEN bits. If XLEN > SEW, _OFFSET_ is _not_ truncated to SEW bits. vslidedown behavior for source elements for element i in slide (vstart < vl) 0 <= i+OFFSET < VLMAX src[i] = vs2[i+OFFSET] VLMAX <= i+OFFSET src[i] = 0 vslidedown behavior for destination element i in slide (vstart < vl) 0 <= i < vstart Unchanged vstart <= i < vl vd[i] = src[i] if v0.mask[i] enabled vl <= i < VLMAX Follow tail policy ##### [](#31-1-16-3-3-vector-slide-1-up)31.1.16.3.3\. Vector Slide-1-up Variants of slide are provided that only move by one element but which also allow a scalar integer value to be inserted at the vacated element position. vslide1up.vx vd, vs2, rs1, vm # vd[0]=x[rs1], vd[i+1] = vs2[i] The `vslide1up` instruction places the `x` register argument at location 0 of the destination vector register group, provided that element 0 is active, otherwise the destination element update follows the current mask agnostic/undisturbed policy. If XLEN < SEW, the value is sign-extended to SEW bits. If XLEN > SEW, the least-significant bits are copied over and the high XLEN-SEW bits are ignored. The remaining active `vl`\-1 elements are copied over from index _i_ in the source vector register group to index _i_+1 in the destination vector register group. The `vl` register specifies the maximum number of destination vector register elements updated with source values, and remaining elements past `vl` are handled according to the current tail policy ([31.1.3.4.3\. Vector Tail Agnostic and Vector Mask Agnostic vta and vma](#sec-agnostic)). vslide1up behavior when vl > 0 i < vstart unchanged 0 = i = vstart vd[i] = x[rs1] if v0.mask[i] enabled max(vstart, 1) <= i < vl vd[i] = vs2[i-1] if v0.mask[i] enabled vl <= i < VLMAX Follow tail policy The `vslide1up` instruction requires that the destination vector register group does not overlap the source vector register group. Otherwise, the instruction encoding is reserved. ##### [](#sec-vfslide1up)31.1.16.3.4\. Vector Floating-Point Slide-1-up Instruction vfslide1up.vf vd, vs2, rs1, vm # vd[0]=f[rs1], vd[i+1] = vs2[i] The `vfslide1up` instruction is defined analogously to `vslide1up`, but sources its scalar argument from an `f` register. ##### [](#31-1-16-3-5-vector-slide-1-down-instruction)31.1.16.3.5\. Vector Slide-1-down Instruction The `vslide1down` instruction copies the first `vl`\-1 active elements values from index _i_+1 in the source vector register group to index_i_ in the destination vector register group. The `vl` register specifies the maximum number of destination vector register elements written with source values, and remaining elements past `vl` are handled according to the current tail policy ([31.1.3.4.3\. Vector Tail Agnostic and Vector Mask Agnostic vta and vma](#sec-agnostic)). vslide1down.vx vd, vs2, rs1, vm # vd[i] = vs2[i+1], vd[vl-1]=x[rs1] The `vslide1down` instruction places the `x` register argument at location `vl`\-1 in the destination vector register, provided that element `vl-1` is active, otherwise the destination element update follows the current mask agnostic/undisturbed policy. If XLEN < SEW, the value is sign-extended to SEW bits. If XLEN > SEW, the least-significant bits are copied over and the high SEW-XLEN bits are ignored. vslide1down behavior i < vstart unchanged vstart <= i < vl-1 vd[i] = vs2[i+1] if v0.mask[i] enabled vstart <= i = vl-1 vd[vl-1] = x[rs1] if v0.mask[i] enabled vl <= i < VLMAX Follow tail policy | | The vslide1down instruction can be used to load values into a vector register without using memory and without disturbing other vector registers. This provides a path for debuggers to modify the contents of a vector register, albeit slowly, with multiple repeatedvslide1down invocations. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#sec-vfslide1down)31.1.16.3.6\. Vector Floating-Point Slide-1-down Instruction vfslide1down.vf vd, vs2, rs1, vm # vd[i] = vs2[i+1], vd[vl-1]=f[rs1] The `vfslide1down` instruction is defined analogously to `vslide1down`, but sources its scalar argument from an `f` register. #### [](#31-1-16-4-vector-register-gather-instructions)31.1.16.4\. Vector Register Gather Instructions The vector register gather instructions read elements from a first source vector register group at locations given by a second source vector register group. The index values in the second vector are treated as unsigned integers. The source vector can be read at any index < VLMAX regardless of `vl`. The maximum number of elements to write to the destination register is given by `vl`, and the remaining elements past `vl` are handled according to the current tail policy ([31.1.3.4.3\. Vector Tail Agnostic and Vector Mask Agnostic vta and vma](#sec-agnostic)). The operation can be masked, and the mask undisturbed/agnostic policy is followed for inactive elements. vrgather.vv vd, vs2, vs1, vm # vd[i] = (vs1[i] >= VLMAX) ? 0 : vs2[vs1[i]]; vrgatherei16.vv vd, vs2, vs1, vm # vd[i] = (vs1[i] >= VLMAX) ? 0 : vs2[vs1[i]]; The `vrgather.vv` form uses SEW/LMUL for both the data and indices. The `vrgatherei16.vv` form uses SEW/LMUL for the data in`vs2` but EEW=16 and EMUL = (16/SEW)\*LMUL for the indices in `vs1`. | | When SEW=8, vrgather.vv can only reference vector elements 0-255\. The vrgatherei16 form can index 64K elements, and can also be used to reduce the register capacity needed to hold indices when SEW > 16. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If an element index is out of range ( `vs1[i]` ≥ VLMAX ) then zero is returned for the element value. Vector-scalar and vector-immediate forms of the register gather are also provided. These read one element from the source vector at the given index, and write this value to the active elements of the destination vector register. The index value in the scalar register and the immediate, zero-extended to XLEN bits, are treated as unsigned integers. If XLEN > SEW, the index value is _not_ truncated to SEW bits. | | These forms allow any vector element to be "splatted" to an entire vector. | | ----------------------------------------------------------------------------- | vrgather.vx vd, vs2, rs1, vm # vd[i] = (x[rs1] >= VLMAX) ? 0 : vs2[x[rs1]] vrgather.vi vd, vs2, uimm, vm # vd[i] = (uimm >= VLMAX) ? 0 : vs2[uimm] For any `vrgather` instruction, the destination vector register group cannot overlap with the source vector register groups, otherwise the instruction encoding is reserved. #### [](#31-1-16-5-vector-compress-instruction)31.1.16.5\. Vector Compress Instruction The vector compress instruction allows elements selected by a vector mask register from a source vector register group to be packed into contiguous elements at the start of the destination vector register group. vcompress.vm vd, vs2, vs1 # Compress into vd elements of vs2 where vs1 is enabled The vector mask register specified by `vs1` indicates which of the first `vl` elements of vector register group `vs2` should be extracted and packed into contiguous elements at the beginning of vector register `vd`. The remaining elements of `vd` are treated as tail elements according to the current tail policy ([31.1.3.4.3\. Vector Tail Agnostic and Vector Mask Agnostic vta and vma](#sec-agnostic)). Example use of vcompress instruction 8 7 6 5 4 3 2 1 0 Element number 1 1 0 1 0 0 1 0 1 v0 8 7 6 5 4 3 2 1 0 v1 1 2 3 4 5 6 7 8 9 v2 vsetivli t0, 9, e8, m1, tu, ma vcompress.vm v2, v1, v0 1 2 3 4 8 7 5 2 0 v2 T\`vcompress\` is encoded as an unmasked instruction (`vm=1`). The equivalent masked instruction (`vm=0`) is reserved. The destination vector register group cannot overlap the source vector register group or the source mask register, otherwise the instruction encoding is reserved. A trap on a `vcompress` instruction is always reported with a`vstart` of 0. Executing a `vcompress` instruction with a non-zero`vstart` raises an illegal-instruction exception. | | Although possible, vcompress is one of the more difficult instructions to restart with a non-zero vstart, so assumption is implementations will choose not do that but will instead restart from element 0\. This does mean elements in destination register aftervstart will already have been updated. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#31-1-16-5-1-synthesizing-vdecompress)31.1.16.5.1\. Synthesizing `vdecompress` There is no inverse `vdecompress` provided, as this operation can be readily synthesized using iota and a masked vrgather: Desired functionality of 'vdecompress' 7 6 5 4 3 2 1 0 # vid e d c b a # packed vector of 5 elements 1 0 0 1 1 1 0 1 # mask vector of 8 elements p q r s t u v w # destination register before vdecompress e q r d c b v a # result of vdecompress # v0 holds mask # v1 holds packed data # v11 holds input expanded vector and result viota.m v10, v0 # Calc iota from mask in v0 vrgather.vv v11, v1, v10, v0.t # Expand into destination p q r s t u v w # v11 destination register e d c b a # v1 source vector 1 0 0 1 1 1 0 1 # v0 mask vector 4 4 4 3 2 1 1 0 # v10 result of viota.m e q r d c b v a # v11 destination after vrgather using viota.m under mask #### [](#31-1-16-6-whole-vector-register-move)31.1.16.6\. Whole Vector Register Move The `vmvr.v` instructions copy whole vector registers (i.e., all VLEN bits) and can copy whole vector register groups. The `nr` value in the opcode is the number of individual vector registers, NREG, to copy. The instructions operate as if EEW=SEW, EMUL = NREG, effective length `evl`\= EMUL \* VLEN/SEW. | | These instructions are intended to aid compilers to shuffle vector registers without needing to know or change vl. | | --------------------------------------------------------------------------------------------------------------------- | The usual property that no elements are written if `vstart` ≥ `vl`does not apply to these instructions. Similarly, the property that the instructions are reserved if `vstart`exceeds the largest element index for the current `vtype` setting does not apply. Instead, the instructions are reserved if `vstart` ≥ `evl`. | | If vd is equal to vs2, the instruction does not change any vector register state. Implementations that rearrange data internally can treat this instruction as a hint that the register group will next be accessed with an EEW equal to SEW. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The instruction is encoded as an OPIVI instruction. The number of vector registers to copy is encoded in the low three bits of the`simm` field (`simm[2:0]`) using the same encoding as the `nf[2:0]` field for memory instructions (Figure [Table 14](#fig-nf)), i.e., `simm[2:0]` \= NREG-1. The value of NREG must be 1, 2, 4, or 8, and values of `simm[4:0]`other than 0, 1, 3, and 7 are reserved. | | A future extension may support other numbers of registers to be moved. | | ------------------------------------------------------------------------- | | | The instruction uses the same funct6 encoding as the vsmulinstruction but with an immediate operand, and only the unmasked version (vm=1). This encoding is chosen as it is close to the related vmerge encoding, and it is unlikely the vsmul instruction would benefit from an immediate form. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | vmvr.v vd, vs2 # General form vmv1r.v v1, v2 # Copy v1=v2 vmv2r.v v10, v12 # Copy v10=v12; v11=v13 vmv4r.v v4, v8 # Copy v4=v8; v5=v9; v6=v10; v7=v11 vmv8r.v v0, v8 # Copy v0=v8; v1=v9; ...; v7=v15 The source and destination vector register numbers must be aligned appropriately for the vector register group size, and encodings with other vector register numbers are reserved. | | A future extension may relax the vector register alignment restrictions. | | --------------------------------------------------------------------------- | ### [](#31-1-17-exception-handling)31.1.17\. Exception Handling On a trap during a vector instruction (caused by either a synchronous exception or an asynchronous interrupt), the existing `*epc` CSR is written with a pointer to the trapping vector instruction, while the`vstart` CSR contains the element index on which the trap was taken. | | We chose to add a vstart CSR to allow resumption of a partially executed vector instruction to reduce interrupt latencies and to simplify forward-progress guarantees. This is similar to the scheme in the IBM 3090 vector facility. To ensure forward progress without the vstart CSR, implementations would have to guarantee an entire vector instruction can always complete atomically without generating a trap. This is particularly difficult to ensure in the presence of constant-stride or scatter/gather operations and demand-paged virtual memory. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-17-1-precise-vector-traps)31.1.17.1\. Precise vector traps | | We assume most supervisor-mode environments with demand-paging will require precise vector traps. | | ---------------------------------------------------------------------------------------------------- | Precise vector traps require that: 1. all instructions older than the trapping vector instruction have committed their results 2. no instructions newer than the trapping vector instruction have altered architectural state 3. any operations within the trapping vector instruction affecting result elements preceding the index in the `vstart` CSR have committed their results 4. no operations within the trapping vector instruction affecting elements at or following the `vstart` CSR have altered architectural state except if restarting and completing the affected vector instruction will nevertheless produce the correct final state. We relax the last requirement to allow elements following `vstart` to have been updated at the time the trap is reported, provided that re-executing the instruction from the given `vstart` will correctly overwrite those elements. In idempotent memory regions, vector store instructions may have updated elements in memory past the element causing a synchronous trap. Non-idempotent memory regions must not have been updated for indices equal to or greater than the element that caused a synchronous trap during a vector store instruction. Except where noted above, vector instructions are allowed to overwrite their inputs, and so in most cases, the vector instruction restart must be from the `vstart` element index. However, there are a number of cases where this overwrite is prohibited to enable execution of the vector instructions to be idempotent and hence restartable from an earlier index location. Implementations must ensure forward progress can be eventually guaranteed for the element or segment reported by `vstart`. #### [](#31-1-17-2-imprecise-vector-traps)31.1.17.2\. Imprecise vector traps Imprecise vector traps are traps that are not precise. In particular, instructions newer than `*epc` may have committed results, and instructions older than `*epc` may have not completed execution. Imprecise traps are primarily intended to be used in situations where reporting an error and terminating execution is the appropriate response. | | A profile might specify that interrupts are precise while other traps are imprecise. We assume many embedded implementations will generate only imprecise traps for vector instructions on fatal errors, as they will not require resumable traps. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Imprecise traps shall report the faulting element in `vstart` for traps caused by synchronous vector exceptions. There is no support for imprecise traps in the current standard extensions. #### [](#31-1-17-3-selectable-preciseimprecise-traps)31.1.17.3\. Selectable precise/imprecise traps Some profiles may choose to provide a privileged mode bit to select between precise and imprecise vector traps. Imprecise mode would run at high-performance but possibly make it difficult to discern error causes, while precise mode would run more slowly, but support debugging of errors albeit with a possibility of not experiencing the same errors as in imprecise mode. This mechanism is not defined in the current standard extensions. #### [](#31-1-17-4-swappable-traps)31.1.17.4\. Swappable traps Another trap mode can support swappable state in the vector unit, where on a trap, special instructions can save and restore the vector unit microarchitectural state, to allow execution to continue correctly around imprecise traps. This mechanism is not defined in the current standard extensions. | | A future extension might define a standard way of saving and restoring opaque microarchitectural state from a vector unit implementation to support context switching with imprecise traps. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#sec-vector-extensions)31.1.18\. Standard Vector Extensions This section describes the standard vector extensions. A set of smaller extensions intended for embedded use are named with a "Zve" prefix, while a larger vector extension designed for application processors is named as a single-letter V extension. A set of vector length extension names with prefix "Zvl" are also provided. The initial vector extensions are designed to act as a base for additional vector extensions in various domains, including cryptography and machine learning. #### [](#31-1-18-1-zvl-minimum-vector-length-standard-extensions)31.1.18.1\. Zvl\*: Minimum Vector Length Standard Extensions All standard vector extensions have a minimum required VLEN as described below. A set of vector length extensions are provided to increase the minimum vector length of a vector extension. | | The vector length extensions can be used to either specify additional software or architecture profile requirements, or to advertise hardware capabilities. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 18\. Vector length extensions__ | Extension | Minimum VLEN | | --------- | ------------ | | Zvl32b | 32 | | Zvl64b | 64 | | Zvl128b | 128 | | Zvl256b | 256 | | Zvl512b | 512 | | Zvl1024b | 1024 | | | Longer vector length extensions should follow the same pattern. | | ------------------------------------------------------------------ | | | Every vector length extension effectively includes all shorter vector length extensions. | | ------------------------------------------------------------------------------------------- | | | Explicit use of the Zvl32b extension string is not required for any standard vector extension as they all effectively mandate at least this minimum, but the string can be useful when stating hardware capabilities. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#31-1-18-2-zve-vector-extensions-for-embedded-processors)31.1.18.2\. Zve\*: Vector Extensions for Embedded Processors The following five standard extensions are defined to provide varying degrees of vector support and are intended for use with embedded processors. Any of these extensions can be added to base ISAs with XLEN=32 or XLEN=64. The table lists the minimum VLEN and supported EEWs for each extension as well as what floating-point types are supported. __Table 19\. Embedded vector extensions__ | Extension | Minimum VLEN | Supported EEW | FP32 | FP64 | | --------- | ------------ | ------------- | ---- | ---- | | Zve32x | 32 | 8, 16, 32 | N | N | | Zve32f | 32 | 8, 16, 32 | Y | N | | Zve64x | 64 | 8, 16, 32, 64 | N | N | | Zve64f | 64 | 8, 16, 32, 64 | Y | N | | Zve64d | 64 | 8, 16, 32, 64 | Y | Y | The Zve32f and Zve64x extensions depend on the Zve32x extension. The Zve64f extension depends on the Zve32f and Zve64x extensions. The Zve64d extension depends on the Zve64f extension. All Zve\* extensions have precise traps. | | There is currently no standard support for handling imprecise traps, so standard extensions have to provide precise traps. | | ----------------------------------------------------------------------------------------------------------------------------- | All Zve\* extensions provide support for EEW of 8, 16, and 32, and Zve64\* extensions also support EEW of 64. All Zve\* extensions support the vector configuration instructions ([31.1.6\. Configuration-Setting Instructions (vsetvli/vsetivli/vsetvl)](#sec-vector-config)). All Zve\* extensions support all vector load and store instructions ([31.1.7\. Vector Loads and Stores](#sec-vector-memory)), except Zve64\* extensions do not support EEW=64 for index values when XLEN=32. All Zve\* extensions support all vector integer instructions ([31.1.11\. Vector Integer Arithmetic Instructions](#sec-vector-integer)), except that the `vmulh` integer multiply variants that return the high word of the product (`vmulh.vv`,`vmulh.vx`, `vmulhu.vv`, `vmulhu.vx`, `vmulhsu.vv`, `vmulhsu.vx`) are not included for EEW=64 in Zve64\*. | | Producing the high-word of a product can take substantial additional gates for large EEW. | | -------------------------------------------------------------------------------------------- | All Zve\* extensions support all vector fixed-point arithmetic instructions ([31.1.12\. Vector Fixed-Point Arithmetic Instructions](#sec-vector-fixed-point)), except that `vsmul.vv` and`vsmul.vx` are not included in EEW=64 in Zve64\*. | | As with vmulh, vsmul requires a large amount of additional logic, and 64-bit fixed-point multiplies are relatively rare. | | --------------------------------------------------------------------------------------------------------------------------- | All Zve\* extensions support all vector integer single-width and widening reduction operations ([31.1.14.1\. Vector Single-Width Integer Reduction Instructions](#sec-vector-integer-reduce),[31.1.14.2\. Vector Widening Integer Reduction Instructions](#sec-vector-integer-reduce-widen)). All Zve\* extensions support all vector mask instructions ([31.1.15\. Vector Mask Instructions](#sec-vector-mask)). All Zve\* extensions support all vector permutation instructions ([31.1.16\. Vector Permutation Instructions](#sec-vector-permute)), except that Zve32x and Zve64x do not include those with floating-point operands, and Zve64f does not include those with EEW=64 floating-point operands. The Zve32x extension depends on the Zicsr extension. The Zve32f and Zve64f extensions depend upon the F extension, and implement all vector floating-point instructions ([31.1.13\. Vector Floating-Point Instructions](#sec-vector-float)) for floating-point operands with EEW=32\. Vector single-width floating-point reduction operations ([31.1.14.3\. Vector Single-Width Floating-Point Reduction Instructions](#sec-vector-float-reduce)) for EEW=32 are supported. The Zve64d extension depends upon the D extension, and implements all vector floating-point instructions ([31.1.13\. Vector Floating-Point Instructions](#sec-vector-float)) for floating-point operands with EEW=32 or EEW=64 (including widening instructions and conversions between FP32 and FP64). Vector single-width floating-point reductions ([31.1.14.3\. Vector Single-Width Floating-Point Reduction Instructions](#sec-vector-float-reduce)) for EEW=32 and EEW=64 are supported as well as widening reductions from FP32 to FP64. #### [](#31-1-18-3-v-vector-extension-for-application-processors)31.1.18.3\. V: Vector Extension for Application Processors The single-letter V extension is intended for use in application processor profiles. The `misa.v` bit is set for implementations providing `misa` and supporting V. The V vector extension has precise traps. The V vector extension depends upon the Zvl128b and Zve64d extensions. | | The value of 128 was chosen as a compromise for application processors. Providing a larger VLEN allows strip-mining code to be elided in some cases for short vectors, but also increases the size of the minimum implementation. Note that larger LMUL can be used to avoid strip mining for longer known-size application vectors at the cost of having fewer available vector register groups. For example, an LMUL of 8 allows vectors of up to sixteen 64-bit elements to be processed without strip mining using four vector register groups. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The V extension supports EEW of 8, 16, and 32, and 64. The V extension supports the vector configuration instructions ([31.1.6\. Configuration-Setting Instructions (vsetvli/vsetivli/vsetvl)](#sec-vector-config)). The V extension supports all vector load and store instructions ([31.1.7\. Vector Loads and Stores](#sec-vector-memory)), except the V extension does not support EEW=64 for index values when XLEN=32. The V extension supports all vector integer instructions ([31.1.11\. Vector Integer Arithmetic Instructions](#sec-vector-integer)). The V extension supports all vector fixed-point arithmetic instructions ([31.1.12\. Vector Fixed-Point Arithmetic Instructions](#sec-vector-fixed-point)). The V extension supports all vector integer single-width and widening reduction operations ([31.1.14.1\. Vector Single-Width Integer Reduction Instructions](#sec-vector-integer-reduce),[31.1.14.2\. Vector Widening Integer Reduction Instructions](#sec-vector-integer-reduce-widen)). The V extension supports all vector mask instructions ([31.1.15\. Vector Mask Instructions](#sec-vector-mask)). The V extension supports all vector permutation instructions ([31.1.16\. Vector Permutation Instructions](#sec-vector-permute)). The V extension depends upon the F and D extensions, and implements all vector floating-point instructions ([31.1.13\. Vector Floating-Point Instructions](#sec-vector-float)) for floating-point operands with EEW=32 or EEW=64 (including widening instructions and conversions between FP32 and FP64). Vector single-width floating-point reductions ([31.1.14.3\. Vector Single-Width Floating-Point Reduction Instructions](#sec-vector-float-reduce)) for EEW=32 and EEW=64 are supported as well as widening reductions from FP32 to FP64. | | As is the case with other RISC-V extensions, it is valid to include overlapping extensions in the same ISA string. For example, RV64GCV and RV64GCV\_Zve64f are both valid and equivalent ISA strings, as is RV64GCV\_Zve64f\_Zve32x\_Zvl128b. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-18-4-zvfhmin-vector-extension-for-minimal-half-precision-floating-point)31.1.18.4\. Zvfhmin: Vector Extension for Minimal Half-Precision Floating-Point The Zvfhmin extension provides minimal support for vectors of IEEE 754-2008 binary16 values, adding conversions to and from binary32\. When the Zvfhmin extension is implemented, the `vfwcvt.f.f.v` and`vfncvt.f.f.w` instructions become defined when SEW=16\. The EEW=16 floating-point operands of these instructions use the binary16 format. The Zvfhmin extension depends on the Zve32f extension. #### [](#31-1-18-5-zvfh-vector-extension-for-half-precision-floating-point)31.1.18.5\. Zvfh: Vector Extension for Half-Precision Floating-Point The Zvfh extension provides support for vectors of IEEE 754-2008 binary16 values.When the Zvfh extension is implemented, all instructions in[31.1.13\. Vector Floating-Point Instructions](#sec-vector-float), [31.1.14.3\. Vector Single-Width Floating-Point Reduction Instructions](#sec-vector-float-reduce),[31.1.14.4\. Vector Widening Floating-Point Reduction Instructions](#sec-vector-float-reduce-widen), [31.1.16.2\. Floating-Point Scalar Move Instructions](#sec-vector-float-move),[31.1.16.3.4\. Vector Floating-Point Slide-1-up Instruction](#sec-vfslide1up), and [31.1.16.3.6\. Vector Floating-Point Slide-1-down Instruction](#sec-vfslide1down)become defined when SEW=16. The EEW=16 floating-point operands of these instructions use the binary16 format. Additionally, conversions between 8-bit integers and binary16 values are provided. The floating-point-to-integer narrowing conversions (`vfncvt[.rtz].x[u].f.w`) and integer-to-floating-point widening conversions (`vfwcvt.f.x[u].v`) become defined when SEW=8. The Zvfh extension depends on the Zve32f and Zfhmin extensions. | | Requiring basic scalar half-precision support makes Zvfh’s vector-scalar instructions substantially more useful. We considered requiring more complete scalar half-precision support, but we reasoned that, for many half-precision vector workloads, performing the scalar computation in single-precision will suffice. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#vector-element-groups)31.1.19\. Vector Element Groups Some vector instructions treat operands as a vector of one or more_element_ _groups_, where each element group is a fixed number of elements. For example, complex numbers can be viewed as a two-element group (one real element and one imaginary element). As another example, the SHA-256 cryptographic instructions in the Zvknha extension operate on 128-bit values represented as a 4-element group of 32-bit elements. This section describes recommendations and terminology for generic instruction set design for vector instructions that operate on element groups. #### [](#31-1-19-1-element-group-size)31.1.19.1\. Element Group Size The _element_ _group_ _size_ (EGS) is the number of elements in one group, and must be a power-of-two (POT). | | Support for non-POT EGS was considered but causes many practical complications and so has been dropped. Error checking for vl is a little more difficult. For LMUL>1, non-POT EGSs will result in groups straddling the individual vector registers in a vector register group. Non-POT EGS can also cause large increases in the lowest-common-multiple of element group sizes, which adds constraints to vl setting in order to avoid splitting an element group across strip-mine iterations in vector-length-agnostic code. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The element group size is statically encoded in the instruction, often implicitly as part of the opcode. Vector instructions with EGS > VLMAX are reserved. | | The vector instructions in the base V vector ISA can be viewed as all having an element group size of 1 for all operands statically encoded in the instruction. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | Many operations only make sense with a certain number of elements per group (e.g., complex operations require a element group size of 2 and SHA-256 requires an element group size of 4). | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-19-2-setting-vl)31.1.19.2\. Setting `vl` Each source and destination operand to a vector instruction might be defined as either a single element group or a vector of element groups. When an operand is a vector of element groups, the `vl`setting must correspond to an integer multiple of the element group size, with other values of `vl` reserved. | | For example, a SHA-256 instruction would require that vl is a multiple of 4. | | ------------------------------------------------------------------------------- | When element group instructions are present, an additional constraint is placed on the setting of `vl` based on an AVL value (augmenting [31.1.6.3\. Constraints on Setting vl](#constraints-on-setting-vl)). EGSMAX is the largest EGS supported by the implementation. When AVL > VLMAX, the value of `vl` must be set to either VLMAX or a positive integer multiple of EGSMAX. | | As the base vector extension only has element group size of 1, this constraint is backwards-compatible. | | ---------------------------------------------------------------------------------------------------------- | | | This constraint prevents element groups being broken across strip-mining iterations in vector-length-agnostic code when a VLMAX-size vector would otherwise be able to accommodate a whole number of element groups. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | If EEW is encoded statically in the instruction, or if an instruction has multiple operands containing vectors of element groups with different EEW, an appropriate SEW must be chosen for vsetvlinstructions. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Additional constraints may be required for some element group instructions to ensure legal length values for all operands. | | ----------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-19-3-determining-eew)31.1.19.3\. Determining EEW The `vtype` SEW can be used to indicate or calculate the effective element size (EEW) of one or more operands of an element group instruction. Where the operand is an element group, SEW and EEW refer to the number of bits in each individual element within a group not the number of bits in the group as a whole. Alternatively, the opcode might encode EEW of all operands statically and ignore the value of SEW when the operation only makes sense for a single size on each operand. | | Many operations are only defined for one EEW, e.g., SHA-256 requires EEW=32\. Encoding EEWs statically in the instruction removes a dynamic dependency on the SEW value and the need to check for errors in SEW values. However, ignoring SEW also prevents reuse of the static opcode with a different dynamic SEW, and in many cases, the SEW setting will be needed for regular vector instructions used to process the individual elements in the vector. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-19-4-determining-emul)31.1.19.4\. Determining EMUL The `vtype` LMUL setting can be used to indicate or calculate the effective length multiplier (EMUL) for one or more operands. Element group instructions tend to exhibit a much wider range of relationships between various operand EEW/EMUL values. For example, an instruction might take a vector of length N of 4-element groups with EEW=8b and reduce each group to produce a vector length N of 1-element groups with EEW=32b. In this case, the input and output EMUL values are equal even though the EEW settings differ by a factor of 4. Each source and destination operand to a vector instruction may have a different element group size, different EMUL, and/or different EEW. #### [](#31-1-19-5-element-group-width)31.1.19.5\. Element Group Width The _element_ _group_ _width_ (EGW) is the number of bits in the element group as a whole. For example, the SHA-256 instructions in the Zvknha extension operate on an EGW of 128, with EGS=4 and EEW=32\. It is possible to use LMUL to concatenate multiple vector registers together to support larger EGW>VLEN. | | If software using large-EGW instructions need be portable across a range of implementations, some of which may have VLEN1, then software can only use a subset of the architectural registers. Profiles can set minimum VLEN requirements to inform authors of such software. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | Element group operations by their nature will gather data from across a wider portion of a vector datapath than regular vector instructions. Some element group instructions might allow temporal execution of individual element operations in a larger group, while others will require all EGW bits of a group to be presented to a functional unit at the same time. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#31-1-19-6-masking)31.1.19.6\. Masking No ratified extensions include masked element-group instructions. Future extensions might extend the element-group scheme to support element-level masking, or might define the concept of a _mask element group_(which might, e.g., update the destination element group if any mask bit in the mask element group is set). ### [](#31-1-20-vector-instruction-listing)31.1.20\. Vector Instruction Listing | Integer | Integer | FP | | | | | ------- | ------- | ------ | - | ----- | - | | funct3 | funct3 | funct3 | | | | | OPIVV | V | OPMVV | V | OPFVV | V | | OPIVX | X | OPMVX | X | OPFVF | F | | OPIVI | I | | | | | | funct6 | funct6 | funct6 | | | | | | | | | | | | ------ | ------ | ------------ | ---------- | -------- | ------ | ----------- | ------- | ------ | ------------ | ----- | ----- | ------- | | 000000 | V | X | I | vadd | 000000 | V | vredsum | 000000 | V | F | vfadd | | | 000001 | 000001 | V | vredand | 000001 | V | vfredusum | | | | | | | | 000010 | V | X | vsub | 000010 | V | vredor | 000010 | V | F | vfsub | | | | 000011 | X | I | vrsub | 000011 | V | vredxor | 000011 | V | vfredosum | | | | | 000100 | V | X | vminu | 000100 | V | vredminu | 000100 | V | F | vfmin | | | | 000101 | V | X | vmin | 000101 | V | vredmin | 000101 | V | vfredmin | | | | | 000110 | V | X | vmaxu | 000110 | V | vredmaxu | 000110 | V | F | vfmax | | | | 000111 | V | X | vmax | 000111 | V | vredmax | 000111 | V | vfredmax | | | | | 001000 | 001000 | V | X | vaaddu | 001000 | V | F | vfsgnj | | | | | | 001001 | V | X | I | vand | 001001 | V | X | vaadd | 001001 | V | F | vfsgnjn | | 001010 | V | X | I | vor | 001010 | V | X | vasubu | 001010 | V | F | vfsgnjx | | 001011 | V | X | I | vxor | 001011 | V | X | vasub | 001011 | | | | | 001100 | V | X | I | vrgather | 001100 | 001100 | | | | | | | | 001101 | 001101 | 001101 | | | | | | | | | | | | 001110 | X | I | vslideup | 001110 | X | vslide1up | 001110 | F | vfslide1up | | | | | 001110 | V | vrgatherei16 | | | | | | | | | | | | 001111 | X | I | vslidedown | 001111 | X | vslide1down | 001111 | F | vfslide1down | | | | | funct6 | funct6 | funct6 | | | | | | | | | | | ------ | ------ | --------- | -------- | ---------- | --------- | -------- | --------- | ------ | -------- | ------------ | ----- | | 010000 | V | X | I | vadc | 010000 | V | VWXUNARY0 | 010000 | V | VWFUNARY0 | | | 010000 | X | VRXUNARY0 | 010000 | F | VRFUNARY0 | | | | | | | | 010001 | V | X | I | vmadc | 010001 | 010001 | | | | | | | 010010 | V | X | vsbc | 010010 | V | VXUNARY0 | 010010 | V | VFUNARY0 | | | | 010011 | V | X | vmsbc | 010011 | 010011 | V | VFUNARY1 | | | | | | 010100 | 010100 | V | VMUNARY0 | 010100 | | | | | | | | | 010101 | 010101 | 010101 | | | | | | | | | | | 010110 | 010110 | 010110 | | | | | | | | | | | 010111 | V | X | I | vmerge/vmv | 010111 | V | vcompress | 010111 | F | vfmerge/vfmv | | | 011000 | V | X | I | vmseq | 011000 | V | vmandn | 011000 | V | F | vmfeq | | 011001 | V | X | I | vmsne | 011001 | V | vmand | 011001 | V | F | vmfle | | 011010 | V | X | vmsltu | 011010 | V | vmor | 011010 | | | | | | 011011 | V | X | vmslt | 011011 | V | vmxor | 011011 | V | F | vmflt | | | 011100 | V | X | I | vmsleu | 011100 | V | vmorn | 011100 | V | F | vmfne | | 011101 | V | X | I | vmsle | 011101 | V | vmnand | 011101 | F | vmfgt | | | 011110 | X | I | vmsgtu | 011110 | V | vmnor | 011110 | | | | | | 011111 | X | I | vmsgt | 011111 | V | vmxnor | 011111 | F | vmfge | | | | funct6 | funct6 | funct6 | | | | | | | | | | | | ------ | ------ | -------- | ------ | ------- | ------ | ------ | ----- | ------ | ------ | ------ | ------ | ------- | | 100000 | V | X | I | vsaddu | 100000 | V | X | vdivu | 100000 | V | F | vfdiv | | 100001 | V | X | I | vsadd | 100001 | V | X | vdiv | 100001 | F | vfrdiv | | | 100010 | V | X | vssubu | 100010 | V | X | vremu | 100010 | | | | | | 100011 | V | X | vssub | 100011 | V | X | vrem | 100011 | | | | | | 100100 | 100100 | V | X | vmulhu | 100100 | V | F | vfmul | | | | | | 100101 | V | X | I | vsll | 100101 | V | X | vmul | 100101 | | | | | 100110 | 100110 | V | X | vmulhsu | 100110 | | | | | | | | | 100111 | V | X | vsmul | 100111 | V | X | vmulh | 100111 | F | vfrsub | | | | 100111 | I | vmvr | | | | | | | | | | | | 101000 | V | X | I | vsrl | 101000 | 101000 | V | F | vfmadd | | | | | 101001 | V | X | I | vsra | 101001 | V | X | vmadd | 101001 | V | F | vfnmadd | | 101010 | V | X | I | vssrl | 101010 | 101010 | V | F | vfmsub | | | | | 101011 | V | X | I | vssra | 101011 | V | X | vnmsub | 101011 | V | F | vfnmsub | | 101100 | V | X | I | vnsrl | 101100 | 101100 | V | F | vfmacc | | | | | 101101 | V | X | I | vnsra | 101101 | V | X | vmacc | 101101 | V | F | vfnmacc | | 101110 | V | X | I | vnclipu | 101110 | 101110 | V | F | vfmsac | | | | | 101111 | V | X | I | vnclip | 101111 | V | X | vnmsac | 101111 | V | F | vfnmsac | | funct6 | funct6 | funct6 | | | | | | | | | | ------ | ------ | --------- | -------- | -------- | ------ | ------ | ---------- | -------- | ---------- | ------ | | 110000 | V | vwredsumu | 110000 | V | X | vwaddu | 110000 | V | F | vfwadd | | 110001 | V | vwredsum | 110001 | V | X | vwadd | 110001 | V | vfwredusum | | | 110010 | 110010 | V | X | vwsubu | 110010 | V | F | vfwsub | | | | 110011 | 110011 | V | X | vwsub | 110011 | V | vfwredosum | | | | | 110100 | 110100 | V | X | vwaddu.w | 110100 | V | F | vfwadd.w | | | | 110101 | 110101 | V | X | vwadd.w | 110101 | | | | | | | 110110 | 110110 | V | X | vwsubu.w | 110110 | V | F | vfwsub.w | | | | 110111 | 110111 | V | X | vwsub.w | 110111 | | | | | | | 111000 | 111000 | V | X | vwmulu | 111000 | V | F | vfwmul | | | | 111001 | 111001 | 111001 | | | | | | | | | | 111010 | 111010 | V | X | vwmulsu | 111010 | | | | | | | 111011 | 111011 | V | X | vwmul | 111011 | | | | | | | 111100 | 111100 | V | X | vwmaccu | 111100 | V | F | vfwmacc | | | | 111101 | 111101 | V | X | vwmacc | 111101 | V | F | vfwnmacc | | | | 111110 | 111110 | X | vwmaccus | 111110 | V | F | vfwmsac | | | | | 111111 | 111111 | V | X | vwmaccsu | 111111 | V | F | vfwnmsac | | | __Table 20\. VRXUNARY0 encoding space__ | vs2 | | | ----- | ------- | | 00000 | vmv.s.x | __Table 21\. VWXUNARY0 encoding space__ | vs1 | | | ----- | ------- | | 00000 | vmv.x.s | | 10000 | vcpop | | 10001 | vfirst | __Table 22\. VXUNARY0 encoding space__ | vs1 | | | ----- | --------- | | 00010 | vzext.vf8 | | 00011 | vsext.vf8 | | 00100 | vzext.vf4 | | 00101 | vsext.vf4 | | 00110 | vzext.vf2 | | 00111 | vsext.vf2 | __Table 23\. VRFUNARY0 encoding space__ | vs2 | | | ----- | -------- | | 00000 | vfmv.s.f | __Table 24\. VWFUNARY0 encoding space__ | vs1 | | | ----- | -------- | | 00000 | vfmv.f.s | __Table 25\. VFUNARY0 encoding space__ | vs1 | name | | --------------------- | ----------------- | | single-width converts | | | 00000 | vfcvt.xu.f.v | | 00001 | vfcvt.x.f.v | | 00010 | vfcvt.f.xu.v | | 00011 | vfcvt.f.x.v | | 00110 | vfcvt.rtz.xu.f.v | | 00111 | vfcvt.rtz.x.f.v | | widening converts | | | 01000 | vfwcvt.xu.f.v | | 01001 | vfwcvt.x.f.v | | 01010 | vfwcvt.f.xu.v | | 01011 | vfwcvt.f.x.v | | 01100 | vfwcvt.f.f.v | | 01110 | vfwcvt.rtz.xu.f.v | | 01111 | vfwcvt.rtz.x.f.v | | narrowing converts | | | 10000 | vfncvt.xu.f.w | | 10001 | vfncvt.x.f.w | | 10010 | vfncvt.f.xu.w | | 10011 | vfncvt.f.x.w | | 10100 | vfncvt.f.f.w | | 10101 | vfncvt.rod.f.f.w | | 10110 | vfncvt.rtz.xu.f.w | | 10111 | vfncvt.rtz.x.f.w | __Table 26\. VFUNARY1 encoding space__ | vs1 | name | | ----- | ---------- | | 00000 | vfsqrt.v | | 00100 | vfrsqrt7.v | | 00101 | vfrec7.v | | 10000 | vfclass.v | __Table 27\. VMUNARY0 encoding space__ | vs1 | | | ----- | ----- | | 00001 | vmsbf | | 00010 | vmsof | | 00011 | vmsif | | 10000 | viota | | 10001 | vid | 33.1. Cryptography Extensions: Vector Instructions, Version 1.0 ==================== ## [](#33-1-cryptography-extensions-vector-instructions-version-1-0)33.1\. Cryptography Extensions: Vector Instructions, Version 1.0 This document describes the Vector Cryptography extensions to the RISC-V Instruction Set Architecture. ### [](#crypto%5Fvector%5Fintroduction)33.1.1\. Introduction This document describes the RISC-V _vector_ cryptography extensions. All instructions described here are based on the Vector registers. The instructions are designed to be highly performant, with large application and server-class cores being the main target. A companion chapter _Volume I: Scalar & Entropy Source Instructions_, describes cryptographic instructions for smaller cores which do not implement the vector extension. #### [](#crypto%5Fvector%5Faudience)33.1.1.1\. Intended Audience Cryptography is a specialized subject, requiring people with many different backgrounds to cooperate in its secure and efficient implementation. Where possible, we have written this specification to be understandable by all, though we recognize that the motivations and references to algorithms or other specifications and standards may be unfamiliar to those who are not domain experts. This specification anticipates being read and acted on by various people with different backgrounds. We have tried to capture these backgrounds here, with a brief explanation of what we expect them to know, and how it relates to the specification. We hope this aids people’s understanding of which aspects of the specification are particularly relevant to them, and which they may (safely!) ignore or pass to a colleague. Cryptographers and cryptographic software developers These are the people we expect to write code using the instructions in this specification. They should understand the motivations for the instructions we include, and be familiar with most of the algorithms and outside standards to which we refer. Computer architects We do not expect architects to have a cryptography background. We nonetheless expect architects to be able to examine our instructions for implementation issues, understand how the instructions will be used in context, and advise on how best to fit the functionality the cryptographers want. Digital design engineers & micro-architects These are the people who will implement the specification inside a core. Again, no cryptography expertise is assumed, but we expect them to interpret the specification and anticipate any hardware implementation issues, e.g., where high-frequency design considerations apply, or where latency/area tradeoffs exist etc. In particular, they should be aware of the literature around efficiently implementing AES and SM4 SBoxes in hardware. Verification engineers These people are responsible for ensuring the correct implementation of the extensions in hardware. No cryptography background is assumed. We expect them to identify interesting test cases from the specification. An understanding of their real-world usage will help with this. These are by no means the only people concerned with the specification, but they are the ones we considered most while writing it. #### [](#crypto%5Fvector%5Fsail%5Fspecifications)33.1.1.2\. Sail Specifications RISC-V maintains a[formal model](https://github.com/riscv/sail-riscv)of the ISA specification, implemented in the Sail ISA specification language \[[27](../biblio/bibliography.html#bib-sail)\]. Note that _Sail_ refers to the specification language itself, and that there is a _model of RISC-V_, written using Sail. It was our intention to include actual Sail code in this specification. However, the Vector Crypto Sail model needs the Vector Sail model as a basis on which to build. This Vector Cryptography extensions specification was completed before there was an approved RISC-V Vector Sail Model. Therefore, we don’t have any Sail code to include in the instruction descriptions. Instead we have included Sail-like pseudocode. While we have endeavored to adhere to Sail syntax, we have taken some liberties for the sake of simplicity where we believe that that our intent is clear to the reader. | | Where variables are concatenated, the order shown is how they would appear in a vector register from left to right. For example, an element group specified as {a, b, e, f} would appear in a vector register with a having the highest element index of the group and f having the lowest index of the group. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For the sake of brevity, our pseudocode does not include the handling of masks or tail elements. We follow the _undisturbed_ and _agnostic_ policies for masks and tails as described in the **RISC-V "V" Vector Extension**specification. Furthermore, the code does not explicitly handle overlap and SEW constraints; these are, however, explicitly stated in the text. In many cases the pseudocode includes calls to supporting functions which are too verbose to include directly in the specification. This supporting code is listed in[33.1.6\. Supporting Sail Code](#crypto%5Fvector%5Fappx%5Fsail). The[Sail Manual](https://alasdair.github.io/manual.html)is recommended reading in order to best understand the code snippets. Also,[The Sail Programming Language: A Sail Cookbook](https://github.com/billmcspadden-riscv/sail/blob/cookbook%5Fbr/cookbook/doc/TheSailCookbook%5FComplete.pdf)is a good reference. For the latest RISC-V Sail model, refer to the formal model GitHub[repository](https://github.com/riscv/sail-riscv). #### [](#crypto%5Fvector%5Fpolicies)33.1.1.3\. Policies In creating this extension, we tried to adhere to the following policies: * Where there is a choice between: 1) supporting diverse implementation strategies for an algorithm or 2) supporting a single implementation style which is more performant / less expensive; the vector crypto extensions will pick the more constrained but performant option. This fits a common pattern in other parts of the RISC-V specifications, where recommended (but not required) instruction sequences for performing particular tasks are given as an example, such that both hardware and software implementers can optimize for only a single use-case. * The extensions will be designed to support _existing_ standardized cryptographic constructs well. It will not try to support proposed standards, or cryptographic constructs which exist only in academia. Cryptographic standards which are settled upon concurrently with or after the RISC-V vector cryptographic extensions standardization will be dealt with by future RISC-V vector cryptographic standard extensions. * Historically, there has been some discussion \[[28](../biblio/bibliography.html#bib-lsyrr:04)\] on how newly supported operations in general-purpose computing might enable new bases for cryptographic algorithms. The standard will not try to anticipate new useful low-level operations which _may_ be useful as building blocks for future cryptographic constructs. * Regarding side-channel countermeasures: Where relevant, proposed instructions must aim to remove the possibility of any timing side-channels. All instructions shall be implemented with data-independent timing. That is, the latency of the execution of these instructions shall not vary with different input values. #### [](#crypto-vector-element-groups)33.1.1.4\. Element Groups Many vector crypto instructions operate on operands that are wider than elements (which are currently limited to 64 bits wide). Typically, these operands are 128- and 256-bits wide. In many cases, these operands are comprised of smaller operands that are combined (for example, each SHA-2 operand is comprised of 4 words). However, in other cases these operands are a single value (for example, in the AES round instructions, each operand is 128-bit block or round key). We treat these operands as a vector of one or more _element groups_ as defined in [Vector Element Groups](v-st-ext.html#vector-element-groups). Each vector crypto instruction that operates on element groups explicitly specifies their three defining parameters: EGW, EGS, and EEW. | Instruction Group | Extension | EGW | EEW | EGS | | ----------------- | --------------------- | --- | --- | --- | | AES | [Zvkned](#zvkned) | 128 | 32 | 4 | | SHA256 | [zvknh\[ab\]](#zvknh) | 128 | 32 | 4 | | SHA512 | [zvknhb](#zvknh) | 256 | 64 | 4 | | GCM | [Zvkg](#zvkg) | 128 | 32 | 4 | | SM4 | [Zvksed](#zvksed) | 128 | 32 | 4 | | SM3 | [Zvksh](#zvksh) | 256 | 32 | 8 | | | Element Group Width (EGW) - total number of bits in an element group Effective Element Width (EEW) - number of bits in each element Element Group Size (EGS) - number of elements in an element group | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For all of the vector crypto instructions in this specification, `EEW`\=`SEW`. | | The required SEW for each cryptographic instruction was chosen to match what is typically needed for other instructions when implementing the targeted algorithm. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * A **Vector Element Group** is a vector of one or more element groups. * A **Scalar Element Group** is a single element group. Element groups can be formed across registers in implementations where`VLEN`< `EGW` by using an `LMUL`\>1. | | Since the **vector extension for application processors** requires a minimum of VLEN of 128, at most such implementations would require LMUL=2 to form the largest element groups in this specification. However, implementations with a smaller VLEN, such as embedded designs, will requires a larger LMULto form the necessary element groups. It is important to keep in mind that this reduces the number of register groups available such that it may be difficult or impossible to write efficient code for the intended cryptographic algorithms. For example, an implementation with VLEN\=32 would need to set LMUL\=8 to create a 256-bit element group for SM3. This would mean that there would only be 4 register groups, 3 of which would be consumed by a single SM3 message-expansion instruction. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | As with all vector instructions, the number of elements processed is specified by the vector length `vl`. The number of element groups operated upon is then `vl`/`EGS`. Likewise the starting element group is `vstart`/`EGS`. See [33.1.1.5\. Instruction Constraints](#crypto-vector-instruction-constraints) for limitations on `vl` and `vstart`for vector crypto instructions. #### [](#crypto-vector-instruction-constraints)33.1.1.5\. Instruction Constraints All standard vector instruction constraints specified by RVV 1.0 apply to Vector Crypto instructions. In addition to those constraints a few additional specific constraints are introduced. The following is a quick reference for the various constraints of specific Vector Crypto instructions. vl and vstart constraints Since `vl` and `vstart` refer to elements, Vector Crypto instructions that use elements groups (See [33.1.1.4\. Element Groups](#crypto-vector-element-groups)) require that these values are an integer multiple of the Element Group Size (`EGS`). * Instructions that violate the `vl` or `vstart` requirements are _reserved_. | Instructions | EGS | | ------------ | --- | | vaes\* | 4 | | vsha2\* | 4 | | vg\* | 4 | | vsm3\* | 8 | | vsm4\* | 4 | LMUL constraints For element-group instructions, `LMUL`\*`VLEN` must always be at least as large as `EGW`, otherwise an_illegal-instruction exception_ is raised, even if `vl`\=0. | Instructions | SEW | EGW | | ------------ | --- | --- | | vaes\* | 32 | 128 | | vsha2\* | 32 | 128 | | vsha2\* | 64 | 256 | | vg\* | 32 | 128 | | vsm3\* | 32 | 256 | | vsm4\* | 32 | 128 | SEW constraints Some Vector Crypto instructions are only defined for a specific `SEW`. In such a case all other `SEW` values are _reserved_. | Instructions | Required SEW | | --------------- | ------------ | | vaes\* | 32 | | Zvknha: vsha2\* | 32 | | Zvknhb: vsha2\* | 32 or 64 | | vclmul\[h\] | 64 | | vg\* | 32 | | vsm3\* | 32 | | vsm4\* | 32 | Vector/Scalar constraints This specification defines new vector/scalar (.vs) instructions that uses **Scalar Element Groups**. The **Scalar Element Group** operand has `EMUL = ceil(EGW / VLEN)`. | | Scalar element group operands do not need to be aligned to LMUL for any implementation with VLEN >= EGW. | | ----------------------------------------------------------------------------------------------------------- | In the case of the `.vs` instructions defined in this specification, `vs2` holds a 128-bit scalar element group. For implementations with `VLEN` ≥ 128, `vs2` refers to a single register. Thus, the `vd` register group must not overlap the `vs2` register. However, in implementations where `VLEN` < 128, `vs2` refers to a register group comprised of the number of registers needed to hold the 128-bit scalar element group. In this case, the `vd` register group must not overlap this `vs2` register group. | Instruction | Register | Cannot Overlap | | ------------ | -------- | -------------- | | vaes\*.vs | vs2 | vd | | vsm4r.vs | vs2 | vd | | vsha2c\[hl\] | vs1, vs2 | vd | | vsha2ms | vs1, vs2 | vd | | vsm3me | vs2 | vd | | vsm3c | vs2 | vd | #### [](#crypto-vector-scalar-instructions)33.1.1.6\. Vector-Scalar Instructions The RISC-V Vector Extension defines three encodings for Vector-Scalar operations which get their scalar operand from a GPR or FP register: * OPIVX: Scalar GPR _x_ register * OPFVF: Scalar FP _f_ register * OPMVX: Scalar GPR _x_ register However, the Vector Extensions include Vector Reduction Operations which can also be considered Vector-Scalar operations because a scalar operand is provided from element 0 of vector register `vs1`. The vector operand is provided in vector register group `vs2`. These reduction operations all use the `.vs` suffix in their mnemonics. Additionally, the reduction operations all produce a scalar result in element 0 of the destination register, `vd`. The Vector Crypto Extensions define Vector-Scalar instructions that are similar to these Vector Reduction Operations in that they get a scalar operand from a vector register. However, they differ in that they get a scalar element group (see [33.1.1.4\. Element Groups](#crypto-vector-element-groups)) from `vs2` and they return _vector_ results to `vd`, which is also a source vector operand.These Vector-Scalar crypto instructions also use the `.vs` suffix in their mnemonics. | | We chose to use vs2 as the scalar operand, and vd as the vector operand, so that we could use the vs1specifier as additional encoding bits for these instructions. This allows these instructions to have a much smaller encoding footprint, leaving more rooms for other instructions in the future. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | These instructions enable a single key, specified as a scalar element group in `vs2`, to be applied to each element group of register group `vd`. | | Scalar element groups will occupy at most a single register in application processors. However, in implementations where VLEN<128, they will occupy 2 (VLEN=64) or 4 (VLEN=32) registers. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | It is common for multiple AES encryption rounds (for example) to be performed in parallel with the same round key (e.g. in counter modes). Rather than having to first splat the common key across the whole vector group, these vector-scalar crypto instructions allow the round key to be specified as a scalar element group. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#crypto-vector-software-portability)33.1.1.7\. Software Portability The following contains some guidelines that enable the portability of vector-crypto-based code to implementations with different values for `VLEN` Application Processors Application processors are expected to follow the V-extension and will therefore have `VLEN` ≥ 128. Since most of the _cryptography-specific_ instructions have an `EGW`\=128, nothing special needs to be done for these instructions to support implementations with `VLEN`\=128. However, the SHA-512 and SM3 instructions have an `EGW`\=256\. Implementations with `VLEN` \= 128, require that`LMUL` is doubled for these instructions in order to create 256-bit elements across a pair of registers. Code written with this doubling of `LMUL` will not affect the results returned by implementations with `VLEN` ≥ 256 because `vl` controls how many element groups are processed. Therefore, we recommend that libraries that implement SHA-512 and SM3 employ this doubling of `LMUL` to ensure that the software can run on all implementation with `VLEN` ≥ 128. While the doubling of `LMUL` for these instructions is _safe_ for implementations with `VLEN` ≥ 256, it may be less optimal as it will result in unnecessary register pressure and might exact a performance penalty in some microarchitectures. Therefore, we suggest that in addition to providing portable code for SHA-512 and SM3, libraries should also include more optimal code for these instructions when `VLEN` ≥ 256. | Algorithm | Instructions | VLEN | LMUL | | --------- | ------------ | ---- | ---- | | SHA-512 | vsha2\* | 64 | vl/2 | | SM3 | vsm3\* | 32 | vl/4 | Embedded Processors Embedded processors will typically have implementations with `VLEN` < 128\. This will require code to be written with larger `LMUL` values to enable the element groups to be formed. The `.vs` instructions require scalar element groups of `EGW`\=128\. On implementations with `VLEN` < 128, these scalar element groups will necessarily be formed across registers. This is different from most scalars in vector instructions that typically consume part of a single register. We recommend that different code be available for `VLEN`\=32 and `VLEN`\=64, as code written for `VLEN`\=32 will likely be too burdensome for `VLEN`\=64 implementations. ### [](#crypto%5Fvector%5Fextensions)33.1.2\. Extensions Overview The section introduces all of the extensions in the Vector Cryptography Instruction Set Extension Specification. The [Zvknhb](#zvknh) and [Zvbc](#zvbc) Vector Crypto Extensions --and accordingly the composite extensions [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkng](#zvkng), and [Zvksc](#zvksc)\-- depend on Zve64x. All of the other Vector Crypto Extensions depend on `Zve32x`. Note: If `Zve32x` is supported then `Zvkb` or `Zvbb` provide support for EEW of 8, 16, and 32\. If `Zve64x` is supported then `Zvkb` or `Zvbb` also add support for EEW 64. All _cryptography-specific_ instructions defined in this Vector Crypto specification (i.e., those in [Zvkned](#zvkned), [Zvknh\[ab\]](#zvknh), [Zvkg](#zvkg), [Zvksed](#zvksed) and [Zvksh](#zvksh) but _not_ [Zvbb](#zvbb),[Zvkb](#zvkb), or [Zvbc](#zvbc)) shall be executed with data-independent execution latency as defined in the[RISC-V Scalar Cryptography Extensions specification](scalar-crypto.html#crypto%5Fscalar%5Finstructions).It is important to note that the Vector Crypto instructions are independent of the implementation of the `Zkt` extension and do not require that `Zkt` is implemented. This specification includes a [Zvkt](#zvkt) extension that, when implemented, requires certain vector instructions (including [Zvbb](#zvbb), [Zvkb](#zvkb), and [Zvbc](#zvbc)) to be executed with data-independent execution latency. Detection of individual cryptography extensions uses the unified software-based RISC-V discovery method. | | At the time of writing, these discovery mechanisms are still a work in progress. | | ----------------------------------------------------------------------------------- | #### [](#zvbb)33.1.2.1\. `Zvbb` \- Vector Basic Bit-manipulation Vector basic bit-manipulation instructions. | | This extension is a superset of the [Zvkb](#zvkb) extension. | | --------------------------------------------------------------- | | Mnemonic | Instruction | | ------------------ | -------------------------------------------------- | | vandn.\[vv,vx\] | [Vector And-Not](#insns-vandn) | | vbrev.v | [Vector Reverse Bits in Elements](#insns-vbrev) | | vbrev8.v | [Vector Reverse Bits in Bytes](#insns-vbrev8) | | vrev8.v | [Vector Reverse Bytes](#insns-vrev8) | | vclz.v | [Vector Count Leading Zeros](#insns-vclz) | | vctz.v | [Vector Count Trailing Zeros](#insns-vctz) | | vcpop.v | [Vector Population Count](#insns-vcpop) | | vrol.\[vv,vx\] | [Vector Rotate Left](#insns-vrol) | | vror.\[vv,vx,vi\] | [Vector Rotate Right](#insns-vror) | | vwsll.\[vv,vx,vi\] | [Vector Widening Shift Left Logical](#insns-vwsll) | #### [](#zvbc)33.1.2.2\. `Zvbc` \- Vector Carry-less Multiplication General purpose carry-less multiplication instructions which are commonly used in cryptography and hashing (e.g., Elliptic curve cryptography, GHASH, CRC). These instructions are only defined for `SEW`\=64. | Mnemonic | Instruction | | ----------------- | ------------------------------------------------------------- | | vclmul.\[vv,vx\] | [Vector Carry-less Multiply](#insns-vclmul) | | vclmulh.\[vv,vx\] | [Vector Carry-less Multiply Return High Half](#insns-vclmulh) | #### [](#zvkb)33.1.2.3\. `Zvkb` \- Vector Cryptography Bit-manipulation Vector bit-manipulation instructions that are essential for implementing common cryptographic workloads securely & efficiently. | | This Zvkb extension is a proper subset of the Zvbb extension. Zvkb allows for vector crypto implementations without incurring the cost of implementing the additional bitmanip instructions in the Zvbb extension: vbrev.v, vclz.v, vctz.v, vcpop.v, and vwsll.\[vv,vx,vi\]. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Mnemonic | Instruction | | ----------------- | --------------------------------------------- | | vandn.\[vv,vx\] | [Vector And-Not](#insns-vandn) | | vbrev8.v | [Vector Reverse Bits in Bytes](#insns-vbrev8) | | vrev8.v | [Vector Reverse Bytes](#insns-vrev8) | | vrol.\[vv,vx\] | [Vector Rotate Left](#insns-vrol) | | vror.\[vv,vx,vi\] | [Vector Rotate Right](#insns-vror) | #### [](#zvkg)33.1.2.4\. `Zvkg` \- Vector GCM/GMAC Instructions to enable the efficient implementation of GHASHH which is used in Galois/Counter Mode (GCM) and Galois Message Authentication Code (GMAC). All of these instructions work on 128-bit element groups comprised of four 32-bit elements. GHASHH is defined in the "Recommendation for Block Cipher Modes of Operation: Galois/Counter Mode (GCM) and GMAC" \[[45](../biblio/bibliography.html#bib-nist:gcm)\] (NIST Specification). | | GCM is used in conjunction with block ciphers (e.g., AES and SM4) to encrypt a message and provide authentication. GMAC is used to provide authentication of a message without encryption. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | To help avoid side-channel timing attacks, these instructions shall be implemented with data-independent timing. The number of element groups to be processed is `vl`/`EGS`.`vl` must be set to the number of `SEW=32` elements to be processed and therefore must be a multiple of `EGS=4`. Likewise, `vstart` must be a multiple of `EGS=4`. | SEW | EGW | Mnemonic | Instruction | | --- | --- | -------- | ----------------------------------------- | | 32 | 128 | vghsh.vv | [Vector GHASH Add-Multiply](#insns-vghsh) | | 32 | 128 | vgmul.vv | [Vector GHASH Multiply](#insns-vgmul) | #### [](#zvkned)33.1.2.5\. `Zvkned` \- NIST Suite: Vector AES Block Cipher Instructions for accelerating encryption, decryption and key-schedule functions of the AES block cipher as defined in Federal Information Processing Standards Publication 197 All of these instructions work on 128-bit element groups comprised of four 32-bit elements. For the best performance, it is suggested that these instruction be implemented on systems with `VLEN`\>=128\. On systems with `VLEN`<128, element groups may be formed by concatenating 32-bit elements from two or four registers by using an LMUL =2 and LMUL=4 respectively. To help avoid side-channel timing attacks, these instructions shall be implemented with data-independent timing. The number of element groups to be processed is `vl`/`EGS`.`vl` must be set to the number of `SEW=32` elements to be processed and therefore must be a multiple of `EGS=4`. Likewise, `vstart` must be a multiple of `EGS=4`. | SEW | EGW | Mnemonic | Instruction | | --- | --- | ---------------- | ---------------------------------------------------- | | 32 | 128 | vaesef.\[vv,vs\] | [Vector AES encrypt final round](#insns-vaesef) | | 32 | 128 | vaesem.\[vv,vs\] | [Vector AES encrypt middle round](#insns-vaesem) | | 32 | 128 | vaesdf.\[vv,vs\] | [Vector AES decrypt final round](#insns-vaesdf) | | 32 | 128 | vaesdm.\[vv,vs\] | [Vector AES decrypt middle round](#insns-vaesdm) | | 32 | 128 | vaeskf1.vi | [Vector AES-128 Forward KeySchedule](#insns-vaeskf1) | | 32 | 128 | vaeskf2.vi | [Vector AES-256 Forward KeySchedule](#insns-vaeskf2) | | 32 | 128 | vaesz.vs | [Vector AES round zero](#insns-vaesz) | #### [](#zvknh)33.1.2.6\. `Zvknh[ab]` \- NIST Suite: Vector SHA-2 Secure Hash Instructions for accelerating SHA-2 as defined in FIPS PUB 180-4 Secure Hash Standard (SHS) `SEW` differentiates between SHA-256 (`SEW`\=32) and SHA-512 (`SEW`\=64). * SHA-256: these instructions work on 128-bit element groups comprised of four 32-bit elements. * SHA-512: these instructions work on 256-bit element groups comprised of four 64-bit elements. | SEW | EGW | SHA-2 | Extension | | --- | --- | ------- | -------------- | | 32 | 128 | SHA-256 | Zvknha, Zvknhb | | 64 | 256 | SHA-512 | Zvknhb | * Zvknhb supports SHA-256 and SHA-512. * Zvknha supports only SHA-256. SHA-256 implementations with VLEN < 128 require LMUL>1 to combine 32-bit elements from register groups to provide all four elements of the element group. SHA-512 implementations with VLEN < 256 require LMUL>1 to combine 64-bit elements from register groups to provide all four elements of the element group. To help avoid side-channel timing attacks, these instructions shall be implemented with data-independent timing. The number of element groups to be processed is `vl`/`EGS`.`vl` must be set to the number of `SEW` elements to be processed and therefore must be a multiple of `EGS=4`. Likewise, `vstart` must be a multiple of `EGS=4`. | Mnemonic | Instruction | | --------------- | ----------------------------------------------- | | vsha2ms.vv | [Vector SHA-2 Message Schedule](#insns-vsha2ms) | | vsha2c\[hl\].vv | [Vector SHA-2 Compression](#insns-vsha2c) | #### [](#zvksed)33.1.2.7\. `Zvksed` \- ShangMi Suite: SM4 Block Cipher Instructions for accelerating encryption, decryption, and key-schedule functions of the SM4 block cipher. The SM4 block cipher is specified in _32907-2016: {SM4} Block Cipher Algorithm_ There are other various sources available that describe the SM4 block cipher. While not the final version of the standard,[RFC 8998 ShangMi (SM) Cipher Suites for TLS 1.3](https://www.rfc-editor.org/rfc/rfc8998.html)is useful and easy to access. All of these instructions work on 128-bit element groups comprised of four 32-bit elements. To help avoid side-channel timing attacks, these instructions shall be implemented with data-independent timing. The number of element groups to be processed is `vl`/`EGS`.`vl` must be set to the number of `SEW=32` elements to be processed and therefore must be a multiple of `EGS=4`. Likewise, `vstart` must be a multiple of `EGS=4`. | SEW | EGW | Mnemonic | Instruction | | --- | --- | --------------- | ---------------------------------------- | | 32 | 128 | vsm4k.vi | [Vector SM4 Key Expansion](#insns-vsm4k) | | 32 | 128 | vsm4r.\[vv,vs\] | [SM4 Block Cipher Rounds](#insns-vsm4r) | #### [](#zvksh)33.1.2.8\. `Zvksh` \- ShangMi Suite: SM3 Secure Hash Instructions for accelerating functions of the SM3 Hash Function. The SM3 secure hash algorithm is specified in _32905-2016: SM3 Cryptographic Hash Algorithm_ There are other various sources available that describe the SM3 secure hash. While not the final version of the standard,[RFC 8998 ShangMi (SM) Cipher Suites for TLS 1.3](https://www.rfc-editor.org/rfc/rfc8998.html)is useful and easy to access. All of these instructions work on 256-bit element groups comprised of eight 32-bit elements. Implementations with VLEN < 256 require LMUL>1 to combine 32-bit elements from register groups to provide all eight elements of the element group. To help avoid side-channel timing attacks, these instructions shall be implemented with data-independent timing. The number of element groups to be processed is `vl`/`EGS`.`vl` must be set to the number of `SEW=32` elements to be processed and therefore must be a multiple of `EGS=8`. Likewise, `vstart` must be a multiple of `EGS=8`. | SEW | EGW | Mnemonic | Instruction | | --- | --- | --------- | -------------------------------------- | | 32 | 256 | vsm3me.vv | [SM3 Message Expansion](#insns-vsm3me) | | 32 | 256 | vsm3c.vi | [SM3 Compression](#insns-vsm3c) | #### [](#zvkn)33.1.2.9\. `Zvkn` \- NIST Algorithm Suite This extension is shorthand for the following set of other extensions: | Included Extension | Description | | ------------------ | ----------------- | | Zvkned | [Zvkned](#zvkned) | | Zvknhb | [Zvknhb](#zvknh) | | Zvkb | [Zvkb](#zvkb) | | Zvkt | [Zvkt](#zvkt) | | | While Zvkg and Zvbc are not part of this extension, it is recommended that at least one of them is implemented with this extension to enable efficient AES-GCM. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#zvknc)33.1.2.10\. `Zvknc` \- NIST Algorithm Suite with carry-less multiply This extension is shorthand for the following set of other extensions: | Included Extension | Description | | ------------------ | ------------- | | Zvkn | [Zvkn](#zvkn) | | Zvbc | [Zvbc](#zvbc) | | | This extension combines the NIST Algorithm Suite with the vector carry-less multiply extension to enable AES-GCM. | | -------------------------------------------------------------------------------------------------------------------- | #### [](#zvkng)33.1.2.11\. `Zvkng` \- NIST Algorithm Suite with GCM This extension is shorthand for the following set of other extensions: | Included Extension | Description | | ------------------ | ------------- | | Zvkn | [Zvkn](#zvkn) | | Zvkg | [Zvkg](#zvkg) | | | This extension combines the NIST Algorithm Suite with the GCM/GMAC extension to enable high-performance AES-GCM. | | ------------------------------------------------------------------------------------------------------------------- | #### [](#zvks)33.1.2.12\. `Zvks` \- ShangMi Algorithm Suite This extension is shorthand for the following set of other extensions: | Included Extension | Description | | ------------------ | ----------------- | | Zvksed | [Zvksed](#zvksed) | | Zvksh | [Zvksh](#zvksh) | | Zvkb | [Zvkb](#zvkb) | | Zvkt | [Zvkt](#zvkt) | | | While Zvkg and Zvbc are not part of this extension, it is recommended that at least one of them is implemented with this extension to enable efficient SM4-GCM. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#zvksc)33.1.2.13\. `Zvksc` \- ShangMi Algorithm Suite with carry-less multiplication This extension is shorthand for the following set of other extensions: | Included Extension | Description | | ------------------ | ------------- | | Zvks | [Zvks](#zvks) | | Zvbc | [Zvbc](#zvbc) | | | This extension combines the ShangMi Algorithm Suite with the vector carry-less multiply extension to enable SM4-GCM. | | ----------------------------------------------------------------------------------------------------------------------- | #### [](#zvksg)33.1.2.14\. `Zvksg` \- ShangMi Algorithm Suite with GCM This extension is shorthand for the following set of other extensions: | Included Extension | Description | | ------------------ | ------------- | | Zvks | [Zvks](#zvks) | | Zvkg | [Zvkg](#zvkg) | | | This extension combines the ShangMi Algorithm Suite with the GCM/GMAC extension to enable high-performance SM4-GCM. | | ---------------------------------------------------------------------------------------------------------------------- | #### [](#zvkt)33.1.2.15\. `Zvkt` \- Vector Data-Independent Execution Latency The Zvkt extension requires all implemented instructions from the following list to be executed with data-independent execution latency as defined in the[RISC-V Scalar Cryptography Extensions specification](scalar-crypto.html#crypto%5Fscalar%5Finstructions). Data-independent execution latency (DIEL) applies to all _data operands_ of an instruction, even those that are not a part of the body or that are inactive. However, DIEL does not apply to other values such as vl, vtype, and the mask (when used to control execution of a masked vector instruction). Also, DIEL does not apply to constant values specified in the instruction encoding such as the use of the zero register (`x0`), and, in the case of immediate forms of an instruction, the values in the immediate fields (i.e., imm, and uimm). In some cases --- which are explicitly specified in the lists below --- operands that are used as control rather than data are exempt from DIEL. | | DIEL helps protect against side-channel timing attacks that are used to determine data values that are intended to be kept secret. Such values include cryptographic keys, plain text, and partially encrypted text. DIEL is not intended to keep software (and cryptographic algorithms contained therein) secret as it is assumed that an adversary would already know these. This is why DIEL doesn’t apply to constants embedded in instruction encodings. It is important that the _values_ of elements that are not in the body or that are masked off do not affect the execution latency of the instruction. Sometimes such elements contain data that also needs to be kept secret. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#33-1-2-15-1-all-zvbb-instructions)33.1.2.15.1\. All [Zvbb](#zvbb) instructions * vandn.v\[vx\] * vclz.v * vcpop.v * vctz.v * vbrev.v * vbrev8.v * vrev8.v * vrol.v\[vx\] * vror.v\[vxi\] * vwsll.\[vv,vx,vi\] | | All [Zvkb](#zvkb) instructions are also covered by DIEL as they are a proper subset of [Zvbb](#zvbb) | | ------------------------------------------------------------------------------------------------------- | ##### [](#33-1-2-15-2-all-zvbc-instructions)33.1.2.15.2\. All [Zvbc](#zvbc) instructions * vclmul\[h\].v\[vx\] ##### [](#33-1-2-15-3-addsub)33.1.2.15.3\. add/sub * v\[r\]sub.v\[vx\] * vadd.v\[ivx\] * vsub.v\[vx\] * vwadd\[u\].\[vw\]\[vx\] * vwsub\[u\].\[vw\]\[vx\] ##### [](#33-1-2-15-4-addsub-with-carry)33.1.2.15.4\. add/sub with carry * vadc.v\[ivx\]m * vmadc.v\[ivx\]\[m\] * vmsbc.v\[vx\]m * vsbc.v\[vx\]m ##### [](#33-1-2-15-5-compare-and-set)33.1.2.15.5\. compare and set * vmseq.v\[vxi\] * vmsgt\[u\].v\[xi\] * vmsle\[u\].v\[xi\] * vmslt\[u\].v\[xi\] * vmsne.v\[ivx\] ##### [](#33-1-2-15-6-copy)33.1.2.15.6\. copy * vmv.s.x * vmv.v.\[ivxs\] * vmv\[1248\]r.v ##### [](#33-1-2-15-7-extend)33.1.2.15.7\. extend * vsext.vf\[248\] * vzext.vf\[248\] ##### [](#33-1-2-15-8-logical)33.1.2.15.8\. logical * vand.v\[ivx\] * vm\[n\]or.mm * vmand\[n\].mm * vmnand.mm * vmorn.mm * vmx\[n\]or.mm * vor.v\[ivx\] * vxor.v\[ivx\] ##### [](#33-1-2-15-9-multiply)33.1.2.15.9\. multiply * vmul\[h\].v\[vx\] * vmulh\[s\]u.v\[vx\] * vwmul.v\[vx\] * vwmul\[s\]u.v\[vx\] ##### [](#33-1-2-15-10-multiply-add)33.1.2.15.10\. multiply-add * vmacc.v\[vx\] * vmadd.v\[vx\] * vnmsac.v\[vx\] * vnmsub.v\[vx\] * vwmacc.v\[vx\] * vwmacc\[s\]u.v\[vx\] * vwmaccus.vx ##### [](#33-1-2-15-11-integer-merge)33.1.2.15.11\. Integer Merge * vmerge.v\[ivx\]m ##### [](#33-1-2-15-12-permute)33.1.2.15.12\. permute In the `.vv` and `.xv` forms of the `vrgather[ei16]` instructions, the values in `vs1` and `rs1` are used for control and therefore are exempt from DIEL. * vrgather.v\[ivx\] * vrgatherei16.vv ##### [](#33-1-2-15-13-shift)33.1.2.15.13\. shift * vnsr\[al\].w\[ivx\] * vsll.v\[ivx\] * vsr\[al\].v\[ivx\] ##### [](#33-1-2-15-14-slide)33.1.2.15.14\. slide * vslide1\[up|down\].vx * vfslide1\[up|down\].vf In the vslide\[up|down\].vx instructions, the value in `rs1`is used for control (i.e., slide amount) and therefore is exempt from DIEL. * vslide\[up|down\].v\[ix\] | | The following instructions are not affected by Zvkt: **All storage operations** **All floating-point operations** add/sub saturate vsadd\[u\].v\[ivx\] vssub\[u\].v\[vx\] clip vnclip\[u\].w\[ivx\] compress vcompress.vm divide vdiv\[u\].v\[vx\] vrem\[u\].v\[vx\] average vaadd\[u\].v\[vx\] vasub\[u\].v\[vx\] mask Op vcpop.m vfirst.m vid.v viota.m vms\[bio\]f.m min/max vmax\[u\].v\[vx\] vmin\[u\].v\[vx\] Multiply-saturate vsmul.v\[vx\] reduce vredsum.vs vwredsum\[u\].vs vred\[and\|or|xor\].vs vred\[min|max\]\[u\].vs shift round vssra.v\[ivx\] vssrl.v\[ivx\] vset vsetivli vsetvl\[i\] | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#crypto%5Fvector%5Finsns)33.1.3\. Instructions #### [](#insns-vaesdf)33.1.3.1\. vaesdf.\[vv,vs\] Synopsis Vector AES final-round decryption Mnemonic vaesdf.vv vd, vs2 vaesdf.vs vd, vs2 Encoding (Vector-Vector) ![svg](_images/svg-0436f29b326c9ce26cc63b0de9cfe34a50adf4cf.svg) Encoding (Vector-Scalar) ![svg](_images/svg-d3556e94803f25207153ff4891b787e4d63440ed.svg) Reserved Encodings * `SEW` is any value other than 32 * Only for the `.vs` form: the `vd` register group overlaps the `vs2` scalar element group Arguments | Register | Direction | EGW | EGS | EEW | Definition | | -------- | --------- | --- | --- | --- | --------------- | | Vd | input | 128 | 4 | 32 | round state | | Vs2 | input | 128 | 4 | 32 | round key | | Vd | output | 128 | 4 | 32 | new round state | Description A final-round AES block cipher decryption is performed. The InvShiftRows and InvSubBytes steps are applied to each round state element group from `vd`. This is then XORed with the round key in either the corresponding element group in `vs2` (vector-vector form) or scalar element group in `vs2` (vector-scalar form). This instruction must always be implemented such that its execution latency does not depend on the data being operated upon. Operation ```sail function clause execute (VAESDF(vs2, vd, suffix)) = { if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { eg_len = (vl/EGS) eg_start = (vstart/EGS) foreach (i from eg_start to eg_len-1) { let keyelem = if suffix == "vv" then i else 0; let state : bits(128) = get_velem(vd, EGW=128, i); let rkey : bits(128) = get_velem(vs2, EGW=128, keyelem); let sr : bits(128) = aes_shift_rows_inv(state); let sb : bits(128) = aes_subbytes_inv(sr); let ark : bits(128) = sb ^ rkey; set_velem(vd, EGW=128, i, ark); } RETIRE_SUCCESS } } ``` Included in [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkned](#zvkned), [Zvkng](#zvkng) #### [](#insns-vaesdm)33.1.3.2\. vaesdm.\[vv,vs\] Synopsis Vector AES middle-round decryption Mnemonic vaesdm.vv vd, vs2 vaesdm.vs vd, vs2 Encoding (Vector-Vector) ![svg](_images/svg-749e00d875e60dc4a0c39a25ee7ee81b67daa98f.svg) Encoding (Vector-Scalar) ![svg](_images/svg-2c576bc989f66fd2cc3425102010d491e59fea01.svg) Reserved Encodings * `SEW` is any value other than 32 * Only for the `.vs` form: the `vd` register group overlaps the `vs2` scalar element group Arguments | Register | Direction | EGW | EGS | EEW | Definition | | -------- | --------- | --- | --- | --- | --------------- | | Vd | input | 128 | 4 | 32 | round state | | Vs2 | input | 128 | 4 | 32 | round key | | Vd | output | 128 | 4 | 32 | new round state | Description A middle-round AES block cipher decryption is performed. The InvShiftRows and InvSubBytes steps are applied to each round state element group from `vd`. This is then XORed with the round key in either the corresponding element group in `vs2` (vector-vector form) or the scalar element group in `vs2` (vector-scalar form). The result is then applied to theInvMixColumns step. This instruction must always be implemented such that its execution latency does not depend on the data being operated upon. Operation ```sail function clause execute (VAESDM(vs2, vd, suffix)) = { if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { eg_len = (vl/EGS) eg_start = (vstart/EGS) foreach (i from eg_start to eg_len-1) { let keyelem = if suffix == "vv" then i else 0; let state : bits(128) = get_velem(vd, EGW=128, i); let rkey : bits(128) = get_velem(vs2, EGW=128, keyelem); let sr : bits(128) = aes_shift_rows_inv(state); let sb : bits(128) = aes_subbytes_inv(sr); let ark : bits(128) = sb ^ rkey; let mix : bits(128) = aes_mixcolumns_inv(ark); set_velem(vd, EGW=128, i, mix); } RETIRE_SUCCESS } } ``` Included in [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkned](#zvkned), [Zvkng](#zvkng) #### [](#insns-vaesef)33.1.3.3\. vaesef.\[vv,vs\] Synopsis Vector AES final-round encryption Mnemonic vaesef.vv vd, vs2 vaesef.vs vd, vs2 Encoding (Vector-Vector) ![svg](_images/svg-8d9d4f71862a13a026ec0403bcc4f72daeee1ba2.svg) Encoding (Vector-Scalar) ![svg](_images/svg-c8d84c7886d494468202ae9b4acc72a814f03509.svg) Reserved Encodings * `SEW` is any value other than 32 * Only for the `.vs` form: the `vd` register group overlaps the `vs2` scalar element group Arguments | Register | Direction | EGW | EGS | EEW | Definition | | -------- | --------- | --- | --- | --- | --------------- | | vd | input | 128 | 4 | 32 | round state | | vs2 | input | 128 | 4 | 32 | round key | | vd | output | 128 | 4 | 32 | new round state | Description A final-round encryption function of the AES block cipher is performed. The SubBytes and ShiftRows steps are applied to each round state element group from `vd`. This is then XORed with the round key in either the corresponding element group in `vs2` (vector-vector form) or the scalar element group in `vs2` (vector-scalar form). This instruction must always be implemented such that its execution latency does not depend on the data being operated upon. Operation ```sail function clause execute (VAESEF(vs2, vd, suffix) = { if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { eg_len = (vl/EGS) eg_start = (vstart/EGS) foreach (i from eg_start to eg_len-1) { let keyelem = if suffix == "vv" then i else 0; let state : bits(128) = get_velem(vd, EGW=128, i); let rkey : bits(128) = get_velem(vs2, EGW=128, keyelem); let sb : bits(128) = aes_subbytes_fwd(state); let sr : bits(128) = aes_shift_rows_fwd(sb); let ark : bits(128) = sr ^ rkey; set_velem(vd, EGW=128, i, ark); } RETIRE_SUCCESS } } ``` Included in [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkned](#zvkned), [Zvkng](#zvkng) #### [](#insns-vaesem)33.1.3.4\. vaesem.\[vv,vs\] Synopsis Vector AES middle-round encryption Mnemonic vaesem.vv vd, vs2 vaesem.vs vd, vs2 Encoding (Vector-Vector) ![svg](_images/svg-a5d4ab5d27058748594a674d09d39b75d53ab7a7.svg) Encoding (Vector-Scalar) ![svg](_images/svg-3d2a97fe21c1c911b72a5d62e501ff415387f83c.svg) Reserved Encodings * `SEW` is any value other than 32 * Only for the `.vs` form: the `vd` register group overlaps the `vs2` scalar element group Arguments | Register | Direction | EGW | EGS | EEW | Definition | | -------- | --------- | --- | --- | --- | --------------- | | Vd | input | 128 | 4 | 32 | round state | | Vs2 | input | 128 | 4 | 32 | Round key | | Vd | output | 128 | 4 | 32 | new round state | Description A middle-round encryption function of the AES block cipher is performed. The SubBytes, ShiftRows, and MixColumns steps are applied to each round state element group from `vd`. This is then XORed with the round key in either the corresponding element group in `vs2` (vector-vector form) or the scalar element group in `vs2` (vector-scalar form). This instruction must always be implemented such that its execution latency does not depend on the data being operated upon. Operation ```sail function clause execute (VAESEM(vs2, vd, suffix)) = { if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { eg_len = (vl/EGS) eg_start = (vstart/EGS) foreach (i from eg_start to eg_len-1) { let keyelem = if suffix == "vv" then i else 0; let state : bits(128) = get_velem(vd, EGW=128, i); let rkey : bits(128) = get_velem(vs2, EGW=128, keyelem); let sb : bits(128) = aes_subbytes_fwd(state); let sr : bits(128) = aes_shift_rows_fwd(sb); let mix : bits(128) = aes_mixcolumns_fwd(sr); let ark : bits(128) = mix ^ rkey; set_velem(vd, EGW=128, i, ark); } RETIRE_SUCCESS } } ``` Included in [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkned](#zvkned), [Zvkng](#zvkng) #### [](#insns-vaeskf1)33.1.3.5\. vaeskf1.vi Synopsis Vector AES-128 Forward KeySchedule generation Mnemonic vaeskf1.vi vd, vs2, uimm Encoding ![svg](_images/svg-d9aef1b16793a1208a181657fdaa9a0100104fa8.svg) Reserved Encodings * `SEW` is any value other than 32 Arguments | Register | Direction | EGW | EGS | EEW | Definition | | -------- | --------- | --- | --- | --- | ------------------ | | uimm | input | \- | \- | \- | Round Number (rnd) | | Vs2 | input | 128 | 4 | 32 | Current round key | | Vd | output | 128 | 4 | 32 | Next round key | Description A single round of the forward AES-128 KeySchedule is performed. The next round key is generated word by word from the current round key element group in `vs2` and the immediately previous word of the round key. The least significant word is generated using the most significant word of the current round key as well as a round constant which is selected by the round number. The round number, which ranges from 1 to 10, comes from `uimm[3:0]`;`uimm[4]` is ignored. The out-of-range `uimm[3:0]` values of 0 and 11-15 are mapped to in-range values by inverting `uimm[3]`. Thus, 0 maps to 8, and 11-15 maps to 3-7\. The round number is used to specify a round constant which is used in generating the first round key word. This instruction must always be implemented such that its execution latency does not depend on the data being operated upon. | | We chose to map out-of-range round numbers to in-range values as this allows the instruction’s behavior to be fully defined for all values of uimm\[4:0\] with minimal extra logic. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```Sail function clause execute (VAESKF1(rnd, vd, vs2)) = { if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { // project out-of-range immediates onto in-range values if( (unsigned(rnd[3:0]) > 10) | (rnd[3:0] = 0)) then rnd[3] = ~rnd[3] eg_len = (vl/EGS) eg_start = (vstart/EGS) let r : bits(4) = rnd-1; foreach (i from eg_start to eg_len-1) { let CurrentRoundKey[3:0] : bits(128) = get_velem(vs2, EGW=128, i); let w[0] : bits(32) = aes_subword_fwd(aes_rotword(CurrentRoundKey[3])) XOR aes_decode_rcon(r) XOR CurrentRoundKey[0] let w[1] : bits(32) = w[0] XOR CurrentRoundKey[1] let w[2] : bits(32) = w[1] XOR CurrentRoundKey[2] let w[3] : bits(32) = w[2] XOR CurrentRoundKey[3] set_velem(vd, EGW=128, i, w[3:0]); } RETIRE_SUCCESS } } ``` Included in [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkned](#zvkned), [Zvkng](#zvkng) #### [](#insns-vaeskf2)33.1.3.6\. vaeskf2.vi Synopsis Vector AES-256 Forward KeySchedule generation Mnemonic vaeskf2.vi vd, vs2, uimm Encoding ![svg](_images/svg-0c072004bd58e6c75f21724d7bac8eaa4135c6fd.svg) Reserved Encodings * `SEW` is any value other than 32 Arguments | Register | Direction | EGW | EGS | EEW | Definition | | -------- | --------- | --- | --- | --- | ------------------ | | Vd | input | 128 | 4 | 32 | Previous Round key | | uimm | input | \- | \- | \- | Round Number (rnd) | | Vs2 | input | 128 | 4 | 32 | Current Round key | | Vd | output | 128 | 4 | 32 | Next round key | Description A single round of the forward AES-256 KeySchedule is performed. The next round key is generated word by word from the previous round key element group in `vd` and the immediately previous word of the round key. The least significant word of the next round key is generated by applying a function to the most significant word of the current round key and then XORing the result with the round constant. The round number is used to select the round constant as well as the function. The round number, which ranges from 2 to 14, comes from `uimm[3:0]`;`uimm[4]` is ignored. The out-of-range `uimm[3:0]` values of 0-1 and 15 are mapped to in-range values by inverting `uimm[3]`. Thus, 0-1 maps to 8-9, and 15 maps to 7. This instruction must always be implemented such that its execution latency does not depend on the data being operated upon. | | We chose to map out-of-range round numbers to in-range values as this allows the instruction’s behavior to be fully defined for all values of uimm\[4:0\] with minimal extra logic. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```Sail function clause execute (VAESKF2(rnd, vd, vs2)) = { if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { // project out-of-range immediates into in-range values if((unsigned(rnd[3:0]) < 2) | (unsigned(rnd[3:0]) > 14)) then rnd[3] = ~rnd[3] eg_len = (vl/EGS) eg_start = (vstart/EGS) foreach (i from eg_start to eg_len-1) { let CurrentRoundKey[3:0] : bits(128) = get_velem(vs2, EGW=128, i); let RoundKeyB[3:0] : bits(128) = get_velem(vd, EGW=128, i); // Previous round key let w[0] : bits(32) = if (rnd[0]==1) then aes_subword_fwd(CurrentRoundKey[3]) XOR RoundKeyB[0]; else aes_subword_fwd(aes_rotword(CurrentRoundKey[3])) XOR aes_decode_rcon((rnd>>1) - 1) XOR RoundKeyB[0]; w[1] : bits(32) = w[0] XOR RoundKeyB[1] w[2] : bits(32) = w[1] XOR RoundKeyB[2] w[3] : bits(32) = w[2] XOR RoundKeyB[3] set_velem(vd, EGW=128, i, w[3:0]); } RETIRE_SUCCESS } } ``` Included in [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkned](#zvkned), [Zvkng](#zvkng) #### [](#insns-vaesz)33.1.3.7\. vaesz.vs Synopsis Vector AES round zero encryption/decryption Mnemonic vaesz.vs vd, vs2 Encoding (Vector-Scalar) ![svg](_images/svg-dea1b30b5938622d80627b6f21acf94c37efa8f4.svg) Reserved Encodings * `SEW` is any value other than 32 * The `vd` register group overlaps the `vs2` register Arguments | Register | Direction | EGW | EGS | EEW | Definition | | -------- | --------- | --- | --- | --- | --------------- | | vd | input | 128 | 4 | 32 | round state | | vs2 | input | 128 | 4 | 32 | round key | | vd | output | 128 | 4 | 32 | new round state | Description A round-0 AES block cipher operation is performed. This operation is used for both encryption and decryption. There is only a `.vs` form of the instruction. `Vs2` holds a scalar element group that is used as the round key for all of the round state element groups.The new round state output of each element group is produced by XORing the round key with each element group of `vd`. This instruction must always be implemented such that its execution latency does not depend on the data being operated upon. | | This instruction is needed to avoid the need to "splat" a 128-bit vector register group when the round key is the same for all 128-bit "lanes". Such a splat would typically be implemented with a vrgather instruction which would hurt performance in many implementations. This instruction only exists in the .vs form because the .vv form would be identical to the vxor.vv vd, vs2, vd instruction. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (VAESZ(vs2, vd) = { if(((vstart%EGS)<>0) | (LMUL*VLEN < EGW)) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { eg_len = (vl/EGS) eg_start = (vstart/EGS) foreach (i from eg_start to eg_len-1) { let state : bits(128) = get_velem(vd, EGW=128, i); let rkey : bits(128) = get_velem(vs2, EGW=128, 0); let ark : bits(128) = state ^ rkey; set_velem(vd, EGW=128, i, ark); } RETIRE_SUCCESS } } ``` Included in [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkned](#zvkned), [Zvkng](#zvkng) #### [](#insns-vandn)33.1.3.8\. vandn.\[vv,vx\] Synopsis Bitwise And-Not Mnemonic vandn.vv vd, vs2, vs1, vm vandn.vx vd, vs2, rs1, vm Encoding (Vector-Vector) ![svg](_images/svg-9fc7ef19d2258965792b1f3cd783e30f43811d09.svg) Encoding (Vector-Scalar) ![svg](_images/svg-c3b69b17898da72e1b692fd59c94e3d1baa27b45.svg) Vector-Vector Arguments | Register | Direction | Definition | | -------- | --------- | -------------------- | | Vs1 | input | Op1 (to be inverted) | | Vs2 | input | Op2 | | Vd | output | Result | Vector-Scalar Arguments | Register | Direction | Definition | | -------- | --------- | -------------------- | | Rs1 | input | Op1 (to be inverted) | | Vs2 | input | Op2 | | Vd | output | Result | Description A bitwise _and-not_ operation is performed. Each bit of `Op1` is inverted and logically ANDed with the corresponding bits in `vs2`. In the vector-scalar version, `Op1` is the sign-extended or truncated value in scalar register `rs1`. In the vector-vector version, `Op1` is `vs1`. | | Note on necessity of instruction This instruction is performance-critical to SHA3\. Specifically, the Chi step of the FIPS 202 Keccak Permutation. Emulating it via 2 instructions is expected to have significant performance impact. The .vv form of the instruction is what is needed for SHA3; the .vx form was added for completeness. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | There is no .vi version of this instruction because the same functionality can be achieved by using an inversion of the immediate value with the vand.vi instruction. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Operation ```sail function clause execute (VANDN(vs2, vs1, vd, suffix)) = { foreach (i from vstart to vl-1) { let op1 = match suffix { "vv" => get_velem(vs1, SEW, i), "vx" => sext_or_truncate_to_sew(X(vs1)) }; let op2 = get_velem(vs2, SEW, i); set_velem(vd, EEW=SEW, i, ~op1 & op2); } RETIRE_SUCCESS } ``` Included in [Zvbb](#zvbb), [Zvkb](#zvkb), [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkng](#zvkng), [Zvks](#zvks) [Zvksc](#zvksc), [Zvksg](#zvksg) #### [](#insns-vbrev)33.1.3.9\. vbrev.v Synopsis Vector Reverse Bits in Elements Mnemonic vbrev.v vd, vs2, vm Encoding (Vector) ![svg](_images/svg-ae6788e0281df2944e04378cde723594e59db59f.svg) Arguments | Register | Direction | Definition | | -------- | --------- | --------------------------- | | Vs2 | input | Input elements | | Vd | output | Elements with bits reversed | Description A bit reversal is performed on the bits of each element. Operation ```sail function clause execute (VBREV(vs2)) = { foreach (i from vstart to vl-1) { let input = get_velem(vs2, SEW, i); let output : bits(SEW) = 0; foreach (i from 0 to SEW-1) let output[SEW-1-i] = input[i]; set_velem(vd, SEW, i, output) } RETIRE_SUCCESS } ``` Included in [Zvbb](#zvbb) #### [](#insns-vbrev8)33.1.3.10\. vbrev8.v Synopsis Vector Reverse Bits in Bytes Mnemonic vbrev8.v vd, vs2, vm Encoding (Vector) ![svg](_images/svg-617fa36e04cb89f63cdacc5815f8ae89a5c6f4f7.svg) Arguments | Register | Direction | Definition | | -------- | --------- | -------------------------------- | | Vs2 | input | Input elements | | Vd | output | Elements with bit-reversed bytes | Description A bit reversal is performed on the bits of each byte. | | This instruction is commonly used for GCM when the zvkg extension is not implemented. This byte-wise instruction is defined for all SEWs to eliminate the need to change SEW when operating on wider elements. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (VBREV8(vs2)) = { foreach (i from vstart to vl-1) { let input = get_velem(vs2, SEW, i); let output : bits(SEW) = 0; foreach (i from 0 to SEW-8 by 8) let output[i+7..i] = reverse_bits_in_byte(input[i+7..i]); set_velem(vd, SEW, i, output) } RETIRE_SUCCESS } ``` Included in [Zvbb](#zvbb), [Zvkb](#zvkb), [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkng](#zvkng), [Zvks](#zvks) [Zvksc](#zvksc), [Zvksg](#zvksg) #### [](#insns-vclmul)33.1.3.11\. vclmul.\[vv,vx\] Synopsis Vector Carry-less Multiply by vector or scalar - returning low half of product. Mnemonic vclmul.vv vd, vs2, vs1, vm vclmul.vx vd, vs2, rs1, vm Encoding (Vector-Vector) ![svg](_images/svg-0af662c5b6e9ca5c407896f11b38338a1e685944.svg) Encoding (Vector-Scalar) ![svg](_images/svg-d533c639b139c5b5c54985dd00a369523bac34ff.svg) Reserved Encodings * `SEW` is any value other than 64 Arguments | Register | Direction | Definition | | -------- | --------- | ---------------------- | | Vs1/Rs1 | input | multiplier | | Vs2 | input | multiplicand | | Vd | output | carry-less product low | Description Produces the low half of 128-bit carry-less product. Each 64-bit element in the `vs2` vector register is carry-less multiplied by either each 64-bit element in `vs1` (vector-vector), or the 64-bit value from integer register `rs1` (vector-scalar). The result is the least significant 64 bits of the carry-less product. | | The 64-bit carry-less multiply instructions can be used for implementing GCM in the absence of the zvkg extension. We do not make these instructions exclusive as the 64-bit carry-less multiply is readily derived from the instructions in the zvkg extension and can have utility in other areas. Likewise, we treat other SEW values as reserved so as not to preclude future extensions from using this opcode with different element widths. For example, a future extension might define an SEW\=32 version of this instruction to enable Zve32\* implementations to have vector carry-less multiplication instructions. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (VCLMUL(vs2, vs1, vd, suffix)) = { foreach (i from vstart to vl-1) { let op1 : bits (64) = if suffix =="vv" then get_velem(vs1,i) else zext_or_truncate_to_sew(X(vs1)); let op2 : bits (64) = get_velem(vs2,i); let product : bits (64) = clmul(op1,op2,SEW); set_velem(vd, i, product); } RETIRE_SUCCESS } function clmul(x, y, width) = { let result : bits(width) = zeros(); foreach (i from 0 to (width - 1)) { if y[i] == 1 then result = result ^ (x << i); } result } ``` Included in [Zvbc](#zvbc), [Zvknc](#zvknc), [Zvksc](#zvksc) #### [](#insns-vclmulh)33.1.3.12\. vclmulh.\[vv,vx\] Synopsis Vector Carry-less Multiply by vector or scalar - returning high half of product. Mnemonic vclmulh.vv vd, vs2, vs1, vm vclmulh.vx vd, vs2, rs1, vm Encoding (Vector-Vector) ![svg](_images/svg-fea30f387ccb854ba59801e5154ba5a784f04b48.svg) Encoding (Vector-Scalar) ![svg](_images/svg-3e461e13c74a687a2a8ce663609ae55ce8523bc9.svg) Reserved Encodings * `SEW` is any value other than 64 Arguments | Register | Direction | Definition | | -------- | --------- | ----------------------- | | Vs1 | input | multiplier | | Vs2 | input | multiplicand | | Vd | output | carry-less product high | Description Produces the high half of 128-bit carry-less product. Each 64-bit element in the `vs2` vector register is carry-less multiplied by either each 64-bit element in `vs1` (vector-vector), or the 64-bit value from integer register `rs1` (vector-scalar). The result is the most significant 64 bits of the carry-less product. Operation ```sail function clause execute (VCLMULH(vs2, vs1, vd, suffix)) = { foreach (i from vstart to vl-1) { let op1 : bits (64) = if suffix =="vv" then get_velem(vs1,i) else zext_or_truncate_to_sew(X(vs1)); let op2 : bits (64) = get_velem(vs2, i); let product : bits (64) = clmulh(op1, op2, SEW); set_velem(vd, i, product); } RETIRE_SUCCESS } function clmulh(x, y, width) = { let result : bits(width) = 0; foreach (i from 1 to (width - 1)) { if y[i] == 1 then result = result ^ (x >> (width - i)); } result } ``` Included in [Zvbc](#zvbc), [Zvknc](#zvknc), [Zvksc](#zvksc) #### [](#insns-vclz)33.1.3.13\. vclz.v Synopsis Vector Count Leading Zeros Mnemonic vclz.v vd, vs2, vm Encoding (Vector) ![svg](_images/svg-6d28a25fd429b4d3e9678b20514bef0a3a4eeaf9.svg) Arguments | Register | Direction | Definition | | -------- | --------- | -------------------------- | | Vs2 | input | Input elements | | Vd | output | Count of leading zero bits | Description A leading zero count is performed on each element. The result for zero-valued inputs is the value SEW. Operation ```sail function clause execute (VCLZ(vs2)) = { foreach (i from vstart to vl-1) { let input = get_velem(vs2, SEW, i); for (j = (SEW - 1); j >= 0; j--) if [input[j]] == 0b1 then break; set_velem(vd, SEW, i, SEW - 1 - j) } RETIRE_SUCCESS } ``` Included in [Zvbb](#zvbb) #### [](#insns-vcpop)33.1.3.14\. vcpop.v Synopsis Count the number of bits set in each element Mnemonic vcpop.v vd, vs2, vm Encoding (Vector) ![svg](_images/svg-1a487b445aea78b0eafcbf948f4b53c8f8f8c04d.svg) Arguments | Register | Direction | Definition | | -------- | --------- | ----------------- | | Vs2 | input | Input elements | | Vd | output | Count of bits set | Description A population count is performed on each element. Operation ```sail function clause execute (VCPOP(vs2)) = { foreach (i from vstart to vl-1) { let input = get_velem(vs2, SEW, i); let output : bits(SEW) = 0; for (j = 0; j < SEW; j++) output = output + input[j]; set_velem(vd, SEW, i, output) } RETIRE_SUCCESS } ``` Included in [Zvbb](#zvbb) #### [](#insns-vctz)33.1.3.15\. vctz.v Synopsis Vector Count Trailing Zeros Mnemonic vctz.v vd, vs2, vm Encoding (Vector) ![svg](_images/svg-3d4d69ca52b755bdcb3577205930274e12983e39.svg) Arguments | Register | Direction | Definition | | -------- | --------- | --------------------------- | | Vs2 | input | Input elements | | Vd | output | Count of trailing zero bits | Description A trailing zero count is performed on each element. Operation ```sail function clause execute (VCTZ(vs2)) = { foreach (i from vstart to vl-1) { let input = get_velem(vs2, SEW, i); for (j = 0; j < SEW; j++) if [input[j]] == 0b1 then break; set_velem(vd, SEW, i, j) } RETIRE_SUCCESS } ``` Included in [Zvbb](#zvbb) #### [](#insns-vghsh)33.1.3.16\. vghsh.vv Synopsis Vector Add-Multiply over GHASH Galois-Field Mnemonic vghsh.vv vd, vs2, vs1 Encoding ![svg](_images/svg-5cb822857c693491e34377f59e0fd737417de939.svg) Reserved Encodings * `SEW` is any value other than 32 Arguments | Register | Direction | EGW | EGS | SEW | Definition | | -------- | --------- | --- | --- | --- | ------------------- | | Vd | input | 128 | 4 | 32 | Partial hash (Yi) | | Vs1 | input | 128 | 4 | 32 | Cipher text (Xi) | | Vs2 | input | 128 | 4 | 32 | Hash Subkey (H) | | Vd | output | 128 | 4 | 32 | Partial-hash (Yi+1) | Description A single "iteration" of the GHASHH algorithm is performed. This instruction treats all of the inputs and outputs as 128-bit polynomials and performs operations over GF\[2\]. It produces the next partial hash (Yi+1) by adding the current partial hash (Yi) to the cipher text block (Xi) and then multiplying (over GF(2128)) this sum by the Hash Subkey (H). The multiplication over GF(2128) is a carry-less multiply of two 128-bit polynomials modulo GHASH’s irreducible polynomial (x128 \+ x7 \+ x2 \+ x + 1). The operation can be compactly defined as Yi+1 \= ((Yi ^ Xi) · H) The NIST specification (see [Zvkg](#zvkg)) orders the coefficients from left to right x0x1x2…​x127for a polynomial x0 \+ x1u +x2 u2 \+ …​ + x127u127. This can be viewed as a collection of byte elements in memory with the byte containing the lowest coefficients (i.e., 0,1,2,3,4,5,6,7) residing at the lowest memory address. Since the bits in the bytes are reversed, this instruction internally performs bit swaps within bytes to put the bits in the standard ordering (e.g., 7,6,5,4,3,2,1,0). This instruction must always be implemented such that its execution latency does not depend on the data being operated upon. | | We are bit-reversing the bytes of inputs and outputs so that the intermediate values are consistent with the NIST specification. These reversals are inexpensive to implement as they unconditionally swap bit positions and therefore do not require any logic. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Since the same hash subkey H will typically be used repeatedly on a given message, a future extension might define a vector-scalar version of this instruction wherevs2 is the scalar element group. This would help reduce register pressure when LMUL \> 1. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```pseudocode function clause execute (VGHSH(vs2, vs1, vd)) = { // operands are input with bits reversed in each byte if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { eg_len = (vl/EGS) eg_start = (vstart/EGS) foreach (i from eg_start to eg_len-1) { let Y = get_velem(vd,EGW=128,i); // current partial-hash let X = get_velem(vs1,EGW=128,i); // block cipher output let H = brev8(get_velem(vs2,EGW=128,i)); // Hash subkey let Z : bits(128) = 0; let S = brev8(Y ^ X); for (int bit = 0; bit < 128; bit++) { if bit_to_bool(S[bit]) Z ^= H bool reduce = bit_to_bool(H[127]); H = H << 1; // left shift H by 1 if (reduce) H ^= 0x87; // Reduce using x^7 + x^2 + x^1 + 1 polynomial } let result = brev8(Z); // bit reverse bytes to get back to GCM standard ordering set_velem(vd, EGW=128, i, result); } RETIRE_SUCCESS } } ``` Included in [Zvkg](#zvkg), [Zvkng](#zvkng), [Zvksg](#zvksg) #### [](#insns-vgmul)33.1.3.17\. vgmul.vv Synopsis Vector Multiply over GHASH Galois-Field Mnemonic vgmul.vv vd, vs2 Encoding ![svg](_images/svg-eb4deead0ddbe25943eccb5e88c2c63a5277156b.svg) Reserved Encodings * `SEW` is any value other than 32 Arguments | Register | Direction | EGW | EGS | SEW | Definition | | -------- | --------- | --- | --- | --- | ------------ | | Vd | input | 128 | 4 | 32 | Multiplier | | Vs2 | input | 128 | 4 | 32 | Multiplicand | | Vd | output | 128 | 4 | 32 | Product | Description A GHASHH multiply is performed. This instruction treats all of the inputs and outputs as 128-bit polynomials and performs operations over GF\[2\]. It produces the product over GF(2128) of the two 128-bit inputs. The multiplication over GF(2128) is a carry-less multiply of two 128-bit polynomials modulo GHASH’s irreducible polynomial (x128 \+ x7 \+ x2 \+ x + 1). The NIST specification (see [Zvkg](#zvkg)) orders the coefficients from left to right x0x1x2…​x127for a polynomial x0 \+ x1u +x2 u2 \+ …​ + x127u127. This can be viewed as a collection of byte elements in memory with the byte containing the lowest coefficients (i.e., 0,1,2,3,4,5,6,7) residing at the lowest memory address. Since the bits in the bytes are reversed, This instruction internally performs bit swaps within bytes to put the bits in the standard ordering (e.g., 7,6,5,4,3,2,1,0). This instruction must always be implemented such that its execution latency does not depend on the data being operated upon. | | We are bit-reversing the bytes of inputs and outputs so that the intermediate values are consistent with the NIST specification. These reversals are inexpensive to implement as they unconditionally swap bit positions and therefore do not require any logic. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Since the same multiplicand will typically be used repeatedly on a given message, a future extension might define a vector-scalar version of this instruction wherevs2 is the scalar element group. This would help reduce register pressure when LMUL \> 1. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | This instruction is identical to vghsh.vv with vs1=0\. This instruction is often used in GHASH code. In some cases it is followed by an XOR to perform a multiply-add. Implementations may choose to fuse these two instructions to improve performance on GHASH code that doesn’t use the add-multiply form of the vghsh.vv instruction. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```pseudocode function clause execute (VGMUL(vs2, vs1, vd)) = { // operands are input with bits reversed in each byte if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { eg_len = (vl/EGS) eg_start = (vstart/EGS) foreach (i from eg_start to eg_len-1) { let Y = brev8(get_velem(vd,EGW=128,i)); // Multiplier let H = brev8(get_velem(vs2,EGW=128,i)); // Multiplicand let Z : bits(128) = 0; for (int bit = 0; bit < 128; bit++) { if bit_to_bool(Y[bit]) Z ^= H bool reduce = bit_to_bool(H[127]); H = H << 1; // left shift H by 1 if (reduce) H ^= 0x87; // Reduce using x^7 + x^2 + x^1 + 1 polynomial } let result = brev8(Z); set_velem(vd, EGW=128, i, result); } RETIRE_SUCCESS } } ``` Included in [Zvkg](#zvkg), [Zvkng](#zvkng), [Zvksg](#zvksg) #### [](#insns-vrev8)33.1.3.18\. vrev8.v Synopsis Vector Reverse Bytes Mnemonic vrev8.v vd, vs2, vm Encoding (Vector) ![svg](_images/svg-51d89973e1c95fe78acebcdfe68c2230fd245e7d.svg) Arguments | Register | Direction | Definition | | -------- | --------- | ---------------------- | | Vs2 | input | Input elements | | Vd | output | Byte-reversed elements | Description A byte reversal is performed on each element of `vs2`, effectively performing an endian swap. | | This element-wise endian swapping is needed for several cryptographic algorithms including SHA2 and SM3. | | ----------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (VREV8(vs2)) = { foreach (i from vstart to vl-1) { input = get_velem(vs2, SEW, i); let output : SEW = 0; let j = SEW - 1; foreach (k from 0 to (SEW - 8) by 8) { output[k..(k + 7)] = input[(j - 7)..j]; j = j - 8; set_velem(vd, SEW, i, output) } RETIRE_SUCCESS } ``` Included in [Zvbb](#zvbb), [Zvkb](#zvkb), [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkng](#zvkng), [Zvks](#zvks) [Zvksc](#zvksc), [Zvksg](#zvksg) #### [](#insns-vrol)33.1.3.19\. vrol.\[vv,vx\] Synopsis Vector rotate left by vector/scalar. Mnemonic vrol.vv vd, vs2, vs1, vm vrol.vx vd, vs2, rs1, vm Encoding (Vector-Vector) ![svg](_images/svg-19c0a8a3417042fd8f62bfc2ac0df76dfe2941c0.svg) Encoding (Vector-Scalar) ![svg](_images/svg-a072ab7c86ae05b0eea6cdfd841d844ee60eddcd.svg) Vector-Vector Arguments | Register | Direction | Definition | | -------- | --------- | ------------- | | Vs1 | input | Rotate amount | | Vs2 | input | Data | | Vd | output | Rotated data | Vector-Scalar Arguments | Register | Direction | Definition | | -------- | --------- | ------------- | | Rs1 | input | Rotate amount | | Vs2 | input | Data | | Vd | output | Rotated data | Description A bitwise left rotation is performed on each element of `vs2` The elements in `vs2` are rotated left by the rotate amount specified by either the corresponding elements of `vs1` (vector-vector), or integer register `rs1`(vector-scalar). Only the low log2(`SEW`) bits of the rotate-amount value are used, all other bits are ignored. | | There is no immediate form of this instruction (i.e., vrol.vi) as the same result can be achieved by negating the rotate amount and using the immediate form of rotate right instruction (i.e., vror.vi). | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Operation ```sail function clause execute (VROL_VV(vs2, vs1, vd)) = { foreach (i from vstart to vl - 1) { set_velem(vd, EEW=SEW, i, get_velem(vs2, i) <<< (get_velem(vs1, i) & (SEW-1)) ) } RETIRE_SUCCESS } function clause execute (VROL_VX(vs2, rs1, vd)) = { foreach (i from vstart to vl - 1) { set_velem(vd, EEW=SEW, i, get_velem(vs2, i) <<< (X(rs1) & (SEW-1)) ) } RETIRE_SUCCESS } ``` Included in [Zvbb](#zvbb), [Zvkb](#zvkb), [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkng](#zvkng), [Zvks](#zvks) [Zvksc](#zvksc), [Zvksg](#zvksg) #### [](#insns-vror)33.1.3.20\. vror.\[vv,vx,vi\] Synopsis Vector rotate right by vector/scalar/immediate. Mnemonic vror.vv vd, vs2, vs1, vm vror.vx vd, vs2, rs1, vm vror.vi vd, vs2, uimm, vm Encoding (Vector-Vector) ![svg](_images/svg-e15ba3ecbb8f69e0606ecc46015034c20e7260c8.svg) Encoding (Vector-Scalar) ![svg](_images/svg-c434188a88a73537c8cbb7ed0279a85f6b303ac0.svg) Encoding (Vector-Immediate) ![svg](_images/svg-a067c944244844a5fa089155a2d798c4c9b6f7c4.svg) Vector-Vector Arguments | Register | Direction | Definition | | -------- | --------- | ------------- | | Vs1 | input | Rotate amount | | Vs2 | input | Data | | Vd | output | Rotated data | Vector-Scalar/Immediate Arguments | Register | Direction | Definition | | -------- | --------- | ------------- | | Rs1/imm | input | Rotate amount | | Vs2 | input | Data | | Vd | output | Rotated data | Description A bitwise right rotation is performed on each element of `vs2`. The elements in `vs2` are rotated right by the rotate amount specified by either the corresponding elements of `vs1` (vector-vector), integer register `rs1`(vector-scalar), or an immediate value (vector-immediate). Only the low log2(`SEW`) bits of the rotate-amount value are used, all other bits are ignored. Operation ```sail function clause execute (VROR_VV(vs2, vs1, vd)) = { foreach (i from vstart to vl - 1) { set_velem(vd, EEW=SEW, i, get_velem(vs2, i) >>> (get_velem(vs1, i) & (SEW-1)) ) } RETIRE_SUCCESS } function clause execute (VROR_VX(vs2, rs1, vd)) = { foreach (i from vstart to vl - 1) { set_velem(vd, EEW=SEW, i, get_velem(vs2, i) >>> (X(rs1) & (SEW-1)) ) } RETIRE_SUCCESS } function clause execute (VROR_VI(vs2, uimm[5:0], vd)) = { foreach (i from vstart to vl - 1) { set_velem(vd, EEW=SEW, i, get_velem(vs2, i) >>> (uimm[5:0] & (SEW-1)) ) } RETIRE_SUCCESS } ``` Included in [Zvbb](#zvbb), [Zvkb](#zvkb), [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkng](#zvkng), [Zvks](#zvks) [Zvksc](#zvksc), [Zvksg](#zvksg) #### [](#insns-vsha2c)33.1.3.21\. vsha2c\[hl\].vv Synopsis Vector SHA-2 two rounds of compression. Mnemonic vsha2ch.vv vd, vs2, vs1 vsha2cl.vv vd, vs2, vs1 Encoding (Vector-Vector) High part ![svg](_images/svg-aef17ff995653d1cecd330709d1da1a86823d3fc.svg) Encoding (Vector-Vector) Low part ![svg](_images/svg-081b53055c613deda6a7664e0eefa2c590afdbdc.svg) Reserved Encodings * `zvknha`: `SEW` is any value other than 32 * `zvknhb`: `SEW` is any value other than 32 or 64 * The `vd` register group overlaps with either `vs1` or `vs2` Arguments | Register | Direction | EGW | EGS | EEW | Definition | | -------- | --------- | ------ | --- | --- | --------------------------------- | | Vd | input | 4\*SEW | 4 | SEW | current state {c, d, g, h} | | Vs1 | input | 4\*SEW | 4 | SEW | MessageSched plus constant\[3:0\] | | Vs2 | input | 4\*SEW | 4 | SEW | current state {a, b, e, f} | | Vd | output | 4\*SEW | 4 | SEW | next state {a, b, e, f} | Description * `SEW`\=32: 2 rounds of SHA-256 compression are performed (`zvknha` and `zvknhb`) * `SEW`\=64: 2 rounds of SHA-512 compression are performed (`zvknhb`) Two words of `vs1` are processed with the 8 words of current state held in `vd` and `vs2` to perform two rounds of hash computation producing four words of the next state. | | Note to software developers The NIST standard (see [zvknh\[ab\]](#zvknh)) requires the final hash to be in big-endian byte ordering within SEW-sized words. Since this instruction treats all words as little-endian, software needs to perform an endian swap on the final output of this instruction after all of the message blocks have been processed. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The vsha2ch version of this instruction uses the two most significant message schedule words from the element group in vs1while the vsha2cl version uses the two least significant message schedule words. Otherwise, these versions of the instruction are identical. Having a high and low version of this instruction typically improves performance when interleaving independent hashing operations (i.e., when hashing several files at once). | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Preventing overlap between vd and vs1 or vs2 simplifies implementation with VLEN < EGW. This restriction does not have any coding impact since proper implementation of the algorithm requires that vd, vs1 and vs2 each are different registers. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (VSHA2c(vs2, vs1, vd)) = { if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { eg_len = (vl/EGS) eg_start = (vstart/EGS) foreach (i from eg_start to eg_len-1) { let {a @ b @ e @ f} : bits(4*SEW) = get_velem(vs2, 4*SEW, i); let {c @ d @ g @ h} : bits(4*SEW) = get_velem(vd, 4*SEW, i); let MessageSchedPlusC[3:0] : bits(4*SEW) = get_velem(vs1, 4*SEW, i); let {W1, W0} == VSHA2cl ? MessageSchedPlusC[1:0] : MessageSchedPlusC[3:2]; // l vs h difference is the words selected let T1 : bits(SEW) = h + sum1(e) + ch(e,f,g) + W0; let T2 : bits(SEW) = sum0(a) + maj(a,b,c); h = g; g = f; f = e; e = d + T1; d = c; c = b; b = a; a = T1 + T2; T1 = h + sum1(e) + ch(e,f,g) + W1; T2 = sum0(a) + maj(a,b,c); h = g; g = f; f = e; e = d + T1; d = c; c = b; b = a; a = T1 + T2; set_velem(vd, 4*SEW, i, {a @ b @ e @ f}); } RETIRE_SUCCESS } } function sum0(x) = { match SEW { 32 => rotr(x,2) XOR rotr(x,13) XOR rotr(x,22), 64 => rotr(x,28) XOR rotr(x,34) XOR rotr(x,39) } } function sum1(x) = { match SEW { 32 => rotr(x,6) XOR rotr(x,11) XOR rotr(x,25), 64 => rotr(x,14) XOR rotr(x,18) XOR rotr(x,41) } } function ch(x, y, z) = ((x & y) ^ ((~x) & z)) function maj(x, y, z) = ((x & y) ^ (x & z) ^ (y & z)) function ROTR(x,n) = (x >> n) | (x << SEW - n) ``` Included in [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkng](#zvkng), [zvknh\[ab\]](#zvknh) #### [](#insns-vsha2ms)33.1.3.22\. vsha2ms.vv Synopsis Vector SHA-2 message schedule. Mnemonic vsha2ms.vv vd, vs2, vs1 Encoding (Vector-Vector) ![svg](_images/svg-b2d06559aaa52915dce8503073077abcea54704b.svg) Reserved Encodings * `zvknha`: `SEW` is any value other than 32 * `zvknhb`: `SEW` is any value other than 32 or 64 * The `vd` register group overlaps with either `vs1` or `vs2` Arguments | Register | Direction | EGW | EGS | EEW | Definition | | -------- | --------- | ------ | --- | --- | -------------------------------------------------- | | Vd | input | 4\*SEW | 4 | SEW | Message words {W\[3\], W\[2\], W\[1\], W\[0\]} | | Vs2 | input | 4\*SEW | 4 | SEW | Message words {W\[11\], W\[10\], W\[9\], W\[4\]} | | Vs1 | input | 4\*SEW | 4 | SEW | Message words {W\[15\], W\[14\], -, W\[12\]} | | Vd | output | 4\*SEW | 4 | SEW | Message words {W\[19\], W\[18\], W\[17\], W\[16\]} | Description * `SEW`\=32: Four rounds of SHA-256 message schedule expansion are performed (`zvknha` and `zvknhb`) * `SEW`\=64: Four rounds of SHA-512 message schedule expansion are performed (`zvknhb`) Eleven of the last 16 `SEW`\-sized message-schedule words from `vd` (oldest), `vs2`, and `vs1` (most recent) are processed to produce the next 4 message-schedule words. | | Note to software developers The first 16 SEW-sized words of the message schedule come from the _message block_in big-endian byte order. Since this instruction treats all words as little endian, software is required to endian swap these words. All of the subsequent message schedule words are produced by this instruction and therefore do not require an endian swap. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Note to software developers Software is required to pack the words into element groups as shown above in the arguments table. The indices indicate the relate age with lower indices indicating older words. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Note to software developers The {W11, W10, W9, W4} element group can easily be formed by using a vector vmerge instruction with the appropriate mask (for example with vl=4 and 4b0001as the 4 mask bits) vmerge.vvm {W11, W10, W9, W4}, {W11, W10, W9, W8}, {W7, W6, W5, W4}, V0 | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | Preventing overlap between vd and vs1 or vs2 simplifies implementation with VLEN < EGW. This restriction does not have any coding impact since proper implementation of the algorithm requires that vd, vs1 and vs2 each contain different portions of the message schedule. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (VSHA2ms(vs2, vs1, vd)) = { // SEW32 = SHA-256 // SEW64 = SHA-512 if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { eg_len = (vl/EGS) eg_start = (vstart/EGS) foreach (i from eg_start to eg_len-1) { {W[3] @ W[2] @ W[1] @ W[0]} : bits(EGW) = get_velem(vd, EGW, i); {W[11] @ W[10] @ W[9] @ W[4]} : bits(EGW) = get_velem(vs2, EGW, i); {W[15] @ W[14] @ W[13] @ W[12]} : bits(EGW) = get_velem(vs1, EGW, i); W[16] = sig1(W[14]) + W[9] + sig0(W[1]) + W[0]; W[17] = sig1(W[15]) + W[10] + sig0(W[2]) + W[1]; W[18] = sig1(W[16]) + W[11] + sig0(W[3]) + W[2]; W[19] = sig1(W[17]) + W[12] + sig0(W[4]) + W[3]; set_velem(vd, EGW, i, {W[19] @ W[18] @ W[17] @ W[16]}); } RETIRE_SUCCESS } } function sig0(x) = { match SEW { 32 => (ROTR(x,7) XOR ROTR(x,18) XOR SHR(x,3)), 64 => (ROTR(x,1) XOR ROTR(x,8) XOR SHR(x,7))); } } function sig1(x) = { match SEW { 32 => (ROTR(x,17) XOR ROTR(x,19) XOR SHR(x,10), 64 => ROTR(x,19) XOR ROTR(x,61) XOR SHR(x,6)); } } function ROTR(x,n) = (x >> n) | (x << SEW - n) function SHR (x,n) = x >> n ``` Included in [Zvkn](#zvkn), [Zvknc](#zvknc), [Zvkng](#zvkng), [zvknh\[ab\]](#zvknh) #### [](#insns-vsm3c)33.1.3.23\. vsm3c.vi Synopsis Vector SM3 Compression Mnemonic vsm3c.vi vd, vs2, uimm Encoding ![svg](_images/svg-1420dacd869e5e6d020b0211712a3ed89315b40f.svg) Reserved Encodings * `SEW` is any value other than 32 * The `vd` register group overlaps with the `vs2` register group Arguments | Register | Direction | EGW | EGS | EEW | Definition | | -------- | --------- | --- | --- | --- | --------------------------------------------------- | | Vd | input | 256 | 8 | 32 | Current state {H,G.F,E,D,C,B,A} | | uimm | input | \- | \- | \- | round number (rnds) | | Vs2 | input | 256 | 8 | 32 | Message words {-,-,w\[5\],w\[4\],-,-,w\[1\],w\[0\]} | | Vd | output | 256 | 8 | 32 | Next state {H,G.F,E,D,C,B,A} | Description Two rounds of SM3 compression are performed. The current state of eight 32-bit words is read in as an element group from `vd`. Eight 32-bit message words are read in as an element group from `vs2`, although only four of them are used. All of the 32-bit input words are byte-swapped from big endian to little endian. These inputs are processed somewhat differently based on the round group (as specified in rnds), and the next state is generated as an element group of eight 32-bit words. The next state of eight 32-bit words are generated, swapped from little endian to big endian, and are returned in an eight-element group. The round number is provided by the 5-bit `rnds` unsigned immediate. Legal values are 0 - 31 and indicate which group of two rounds are being performed. For example, if rnds=1, then rounds 2 and 3 are being performed. | | The round number is used in the rotation of the constant as well to inform the behavior which differs between rounds 0-15 and rounds 16-63. | | ---------------------------------------------------------------------------------------------------------------------------------------------- | | | The endian byte swapping of the input and output words enables us to align with the SM3 specification without requiring that software perform these swaps. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Preventing overlap between vd and vs2 simplifies implementation with VLEN < EGW. This restriction does not have any coding impact since proper implementation of the algorithm requires that vd and vs2 each are different registers. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (VSM3C(rnds, vs2, vd)) = { if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { eg_len = (vl/EGS) eg_start = (vstart/EGS) foreach (i from eg_start to eg_len-1) { // load state let {Hi @ Gi @ Fi @ Ei @ Di @ Ci @ Bi @ Ai} : bits(256) : bits(256) = (get_velem(vd, 256, i)); //load message schedule let {u_w7 @ u_w6 @ w5i @ w4i @ u_w3 @ u_w2 @ w1i @ w0i} : bits(256) = (get_velem(vs2, 256, i)); // u_w inputs are unused // perform endian swap let H : bits(32) = rev8(Hi); let G : bits(32) = rev8(Gi); let F : bits(32) = rev8(Fi); let E : bits(32) = rev8(Ei); let D : bits(32) = rev8(Di); let C : bits(32) = rev8(Ci); let B : bits(32) = rev8(Bi); let A : bits(32) = rev8(Ai); let w5 = : bits(32) rev8(w5i); let w4 = : bits(32) rev8(w4i); let w1 = : bits(32) rev8(w1i); let w0 = : bits(32) rev8(w0i); let x0 :bits(32) = w0 ^ w4; // W'[0] let x1 :bits(32) = w1 ^ w5; // W'[1] let j = 2 * rnds; let ss1 : bits(32) = ROL32(ROL32(A, 12) + E + ROL32(T_j(j), j % 32), 7); let ss2 : bits(32) = ss1 ^ ROL32(A, 12); let tt1 : bits(32) = FF_j(A, B, C, j) + D + ss2 + x0; let tt2 : bits(32) = GG_j(E, F, G, j) + H + ss1 + w0; D = C; let : bits(32) C1 = ROL32(B, 9); B = A; let A1 : bits(32) = tt1; H = G; let G1 : bits(32) = ROL32(F, 19); F = E; let E1 : bits(32) = P_0(tt2); j = 2 * rnds + 1; ss1 = ROL32(ROL32(A1, 12) + E1 + ROL32(T_j(j), j % 32), 7); ss2 = ss1 ^ ROL32(A1, 12); tt1 = FF_j(A1, B, C1, j) + D + ss2 + x1; tt2 = GG_j(E1, F, G1, j) + H + ss1 + w1; D = C1; let C2 : bits(32) = ROL32(B, 9); B = A1; let A2 : bits(32) = tt1; H = G1; let G2 = : bits(32) ROL32(F, 19); F = E1; let E2 = : bits(32) P_0(tt2); // Update the destination register - swap back to big endian let result : bits(256) = {rev8(G1) @ rev8(G2) @ rev8(E1) @ rev8(E2) @ rev8(C1) @ rev8(C2) @ rev8(A1) @ rev8(A2)}; set_velem(vd, 256, i, result); } RETIRE_SUCCESS } } function FF1(X, Y, Z) = ((X) ^ (Y) ^ (Z)) function FF2(X, Y, Z) = (((X) & (Y)) | ((X) & (Z)) | ((Y) & (Z))) function FF_j(X, Y, Z, J) = (((J) <= 15) ? FF1(X, Y, Z) : FF2(X, Y, Z)) function GG1(X, Y, Z) = ((X) ^ (Y) ^ (Z)) function GG2(X, Y, Z) = (((X) & (Y)) | ((~(X)) & (Z))) . function GG_j(X, Y, Z, J) = (((J) <= 15) ? GG1(X, Y, Z) : GG2(X, Y, Z)) function T_j(J) = (((J) <= 15) ? (0x79CC4519) : (0x7A879D8A)) function P_0(X) = ((X) ^ ROL32((X), 9) ^ ROL32((X), 17)) ``` Included in [Zvks](#zvks), [Zvksc](#zvksc), [Zvksg](#zvksg), [Zvksh](#zvksh) #### [](#insns-vsm3me)33.1.3.24\. vsm3me.vv Synopsis Vector SM3 Message Expansion Mnemonic vsm3me.vv vd, vs2, vs1 Encoding ![svg](_images/svg-c378acaaf1493883cd05b0eb6bcf2ba040be9ef1.svg) Reserved Encodings * `SEW` is any value other than 32 * The `vd` register group overlaps with the `vs2` register group. Arguments | Register | Direction | EGW | EGS | EEW | Definition | | -------- | --------- | --- | --- | --- | ------------------------ | | Vs1 | input | 256 | 8 | 32 | Message words W\[7:0\] | | Vs2 | input | 256 | 8 | 32 | Message words W\[15:8\] | | Vd | output | 256 | 8 | 32 | Message words W\[23:16\] | Description Eight rounds of SM3 message expansion are performed. The sixteen most recent 32-bit message words are read in as two eight-element groups from `vs1` and `vs2`. Each of these words is swapped from big endian to little endian. The next eight 32-bit message words are generated, swapped from little endian to big endian, and are returned in an eight-element group. | | The endian byte swapping of the input and output words enables us to align with the SM3 specification without requiring that software perform these swaps. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Preventing overlap between vd and vs2 simplifies implementations with VLEN < EGW. This restriction should not have any coding impact since the algorithm requires these values to be preserved for generating the next 8 words. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (VSM3ME(vs2, vs1)) = { if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { eg_len = (vl/EGS) eg_start = (vstart/EGS) foreach (i from eg_start to eg_len-1) { let w[7:0] : bits(256) = get_velem(vs1, 256, i); let w[15:8] : bits(256) = get_velem(vs2, 256, i); // Byte Swap inputs from big-endian to little-endian let w15 = rev8(w[15]); let w14 = rev8(w[14]); let w13 = rev8(w[13]); let w12 = rev8(w[12]); let w11 = rev8(w[11]); let w10 = rev8(w[10]); let w9 = rev8(w[9]); let w8 = rev8(w[8]); let w7 = rev8(w[7]); let w6 = rev8(w[6]); let w5 = rev8(w[5]); let w4 = rev8(w[4]); let w3 = rev8(w[3]); let w2 = rev8(w[2]); let w1 = rev8(w[1]); let w0 = rev8(w[0]); // Note that some of the newly computed words are used in later invocations. let w[16] = ZVKSH_W(w0 @ w7 @ w13 @ w3 @ w10); let w[17] = ZVKSH_W(w1 @ w8 @ w14 @ w4 @ w11); let w[18] = ZVKSH_W(w2 @ w9 @ w15 @ w5 @ w12); let w[19] = ZVKSH_W(w3 @ w10 @ w16 @ w6 @ w13); let w[20] = ZVKSH_W(w4 @ w11 @ w17 @ w7 @ w14); let w[21] = ZVKSH_W(w5 @ w12 @ w18 @ w8 @ w15); let w[22] = ZVKSH_W(w6 @ w13 @ w19 @ w9 @ w16); let w[23] = ZVKSH_W(w7 @ w14 @ w20 @ w10 @ w17); // Byte swap outputs from little-endian back to big-endian let w16 : Bits(32) = rev8(W[16]); let w17 : Bits(32) = rev8(W[17]); let w18 : Bits(32) = rev8(W[18]); let w19 : Bits(32) = rev8(W[19]); let w20 : Bits(32) = rev8(W[20]); let w21 : Bits(32) = rev8(W[21]); let w22 : Bits(32) = rev8(W[22]); let w23 : Bits(32) = rev8(W[23]); // Update the destination register. set_velem(vd, 256, i, {w23 @ w22 @ w21 @ w20 @ w19 @ w18 @ w17 @ w16}); } RETIRE_SUCCESS } } function P_1(X) ((X) ^ ROL32((X), 15) ^ ROL32((X), 23)) function ZVKSH_W(M16, M9, M3, M13, M6) = \ (P1( (M16) ^ (M9) ^ ROL32((M3), 15) ) ^ ROL32((M13), 7) ^ (M6)) ``` Included in [Zvks](#zvks), [Zvksc](#zvksc), [Zvksg](#zvksg), [Zvksh](#zvksh) #### [](#insns-vsm4k)33.1.3.25\. vsm4k.vi Synopsis Vector SM4 KeyExpansion Mnemonic vsm4k.vi vd, vs2, uimm Encoding ![svg](_images/svg-f6e3390f9eadcc2217beb6365662ea4b99efaa1c.svg) Reserved Encodings * `SEW` is any value other than 32 Arguments | Register | Direction | EGW | EGS | EEW | Definition | | -------- | --------- | --- | --- | --- | ------------------------------ | | uimm | input | \- | \- | \- | Round group (rnd) | | Vs2 | input | 128 | 4 | 32 | Current 4 round keys rK\[0:3\] | | Vd | output | 128 | 4 | 32 | Next 4 round keys rK'\[0:3\] | Description Four rounds of the SM4 Key Expansion are performed. Four round keys are read in as a 4-element group from `vs2`. Each of the next four round keys are generated by iteratively XORing the last three round keys with a constant that is indexed by the Round Group Number, performing a byte-wise substitution, and then performing XORs between rotated versions of this value and the corresponding current round key. The Round group number (`rnd`) comes from `uimm[2:0]`; the bits in `uimm[4:3]` are ignored. Round group numbers range from 0 to 7 and indicate which group of four round keys are being generated. Round Keys range from 0-31\. For example, if `rnd`\=1, then round keys 4, 5, 6, and 7 are being generated. | | Software needs to generate the initial round keys. This is done by XORing the 128-bit encryption key with the system parameters: FK\[0:3\] | | --------------------------------------------------------------------------------------------------------------------------------------------- | __Table 1\. System Parameters__ | FK | constant | | -- | -------- | | 0 | A3B1BAC6 | | 1 | 56AA3350 | | 2 | 677D9197 | | 3 | B27022DC | | | Implementation Hint The round constants (CK) can be generated on the fly fairly cheaply. If the bytes of the constants are assigned an incrementing index from 0 to 127, the value of each byte is equal to its index multiplied by 7 modulo 256\. Since the results are all limited to 8 bits, the modulo operation occurs for free: B\[n\] = n + 2n + 4n; \= 8n + \~n + 1; | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```sail function clause execute (vsm4k(uimm, vs2)) = { if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { eg_len = (vl/EGS) eg_start = (vstart/EGS) let B : bits(32) = 0; let S : bits(32) = 0; let rk4 : bits(32) = 0; let rk5 : bits(32) = 0; let rk6 : bits(32) = 0; let rk7 : bits(32) = 0; let rnd : bits(3) = uimm[2:0]; // Lower 3 bits foreach (i from eg_start to eg_len-1) { let (rk3 @ rk2 @ rk1 @ rk0) : bits(128) = get_velem(vs2, 128, i); B = rk1 ^ rk2 ^ rk3 ^ ck(4 * rnd); S = sm4_subword(B); rk4 = ROUND_KEY(rk0, S); B = rk2 ^ rk3 ^ rk4 ^ ck(4 * rnd + 1); S = sm4_subword(B); rk5 = ROUND_KEY(rk1, S); B = rk3 ^ rk4 ^ rk5 ^ ck(4 * rnd + 2); S = sm4_subword(B); rk6 = ROUND_KEY(rk2, S); B = rk4 ^ rk5 ^ rk6 ^ ck(4 * rnd + 3); S = sm4_subword(B); rk7 = ROUND_KEY(rk3, S); // Update the destination register. set_velem(vd, EGW=128, i, (rk7 @ rk6 @ rk5 @ rk4)); } RETIRE_SUCCESS } } val round_key : bits(32) -> bits(32) function ROUND_KEY(X, S) = ((X) ^ ((S) ^ ROL32((S), 13) ^ ROL32((S), 23))) // SM4 Constant Key (CK) let ck : list(bits(32)) = [| 0x00070E15, 0x1C232A31, 0x383F464D, 0x545B6269, 0x70777E85, 0x8C939AA1, 0xA8AFB6BD, 0xC4CBD2D9, 0xE0E7EEF5, 0xFC030A11, 0x181F262D, 0x343B4249, 0x50575E65, 0x6C737A81, 0x888F969D, 0xA4ABB2B9, 0xC0C7CED5, 0xDCE3EAF1, 0xF8FF060D, 0x141B2229, 0x30373E45, 0x4C535A61, 0x686F767D, 0x848B9299, 0xA0A7AEB5, 0xBCC3CAD1, 0xD8DFE6ED, 0xF4FB0209, 0x10171E25, 0x2C333A41, 0x484F565D, 0x646B7279 |] }; ``` Included in [Zvks](#zvks), [Zvksc](#zvksc), [Zvksed](#zvksed), [Zvksg](#zvksg) #### [](#insns-vsm4r)33.1.3.26\. vsm4r.\[vv,vs\] Synopsis Vector SM4 Rounds Mnemonic vsm4r.vv vd, vs2 vsm4r.vs vd, vs2 Encoding (Vector-Vector) ![svg](_images/svg-ca751c7e470bc683eaaeff3bcab016a89c3dbfaf.svg) Encoding (Vector-Scalar) ![svg](_images/svg-c377a78ab9ebd53bf89d9c6224405430add5e6e5.svg) Reserved Encodings * `SEW` is any value other than 32 * Only for the `.vs` form: the `vd` register group overlaps the `vs2` register Arguments | Register | Direction | EGW | EGS | EEW | Definition | | -------- | --------- | --- | --- | --- | ---------------------- | | Vd | input | 128 | 4 | 32 | Current state X\[0:3\] | | Vs2 | input | 128 | 4 | 32 | Round keys rk\[0:3\] | | Vd | output | 128 | 4 | 32 | Next state X'\[0:3\] | Description Four rounds of SM4 Encryption/Decryption are performed. The four words of current state are read as a 4-element group from 'vd' and the round keys are read from either the corresponding 4-element group in `vs2` (vector-vector form) or the scalar element group in `vs2`(vector-scalar form). The next four words of state are generated by iteratively XORing the last three words of the state with the corresponding round key, performing a byte-wise substitution, and then performing XORs between rotated versions of this value and the corresponding current state. | | In SM4, encryption and decryption are identical except that decryption consumes the round keys in the reverse order. | | ----------------------------------------------------------------------------------------------------------------------- | | | For the first four rounds of encryption, the _current state_ is the plain text. For the first four rounds of decryption, the _current state_ is the cipher text. For all subsequent rounds, the _current state_ is the _next state_ from the previous four rounds. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Operation ```pseudocode function clause execute (VSM4R(vd, vs2)) = { if(LMUL*VLEN < EGW) then { handle_illegal(); // illegal-instruction exception RETIRE_FAIL } else { eg_len = (vl/EGS) eg_start = (vstart/EGS) let B : bits(32) = 0; let S : bits(32) = 0; let rk0 : bits(32) = 0; let rk1 : bits(32) = 0; let rk2 : bits(32) = 0; let rk3 : bits(32) = 0; let x0 : bits(32) = 0; let x1 : bits(32) = 0; let x2 : bits(32) = 0; let x3 : bits(32) = 0; let x4 : bits(32) = 0; let x5 : bits(32) = 0; let x6 : bits(32) = 0; let x7 : bits(32) = 0; let keyelem : bits(32) = 0; foreach (i from eg_start to eg_len-1) { keyelem = if suffix == "vv" then i else 0; {rk3 @ rk2 @ rk1 @ rk0} : bits(128) = get_velem(vs2, EGW=128, keyelem); {x3 @ x2 @ x1 @ x0} : bits(128) = get_velem(vd, EGW=128, i); B = x1 ^ x2 ^ x3 ^ rk0; S = sm4_subword(B); x4 = sm4_round(x0, S); B = x2 ^ x3 ^ x4 ^ rk1; S = sm4_subword(B); x5= sm4_round(x1, S); B = x3 ^ x4 ^ x5 ^ rk2; S = sm4_subword(B); x6 = sm4_round(x2, S); B = x4 ^ x5 ^ x6 ^ rk3; S = sm4_subword(B); x7 = sm4_round(x3, S); set_velem(vd, EGW=128, i, (x7 @ x6 @ x5 @ x4)); } RETIRE_SUCCESS } } val sm4_round : bits(32) -> bits(32) function sm4_round(X, S) = \ ((X) ^ ((S) ^ ROL32((S), 2) ^ ROL32((S), 10) ^ ROL32((S), 18) ^ ROL32((S), 24))) ``` Included in [Zvks](#zvks), [Zvksc](#zvksc), [Zvksed](#zvksed), [Zvksg](#zvksg) #### [](#insns-vwsll)33.1.3.27\. vwsll.\[vv,vx,vi\] Synopsis Vector widening shift left logical by vector/scalar/immediate. Mnemonic vwsll.vv vd, vs2, vs1, vm vwsll.vx vd, vs2, rs1, vm vwsll.vi vd, vs2, uimm, vm Encoding (Vector-Vector) ![svg](_images/svg-0a7394233471b3a363fdae1fc6cdac4ed01cc8ac.svg) Encoding (Vector-Scalar) ![svg](_images/svg-717098c95cd2151bf91aa529db7596490080c84c.svg) Encoding (Vector-Immediate) ![svg](_images/svg-9b819e3fec0029975dccc35a34ea05402957b95d.svg) Vector-Vector Arguments | Register | Direction | Definition | | -------- | --------- | ------------ | | Vs1 | input | Shift amount | | Vs2 | input | Data | | Vd | output | Shifted data | Vector-Scalar/Immediate Arguments | Register | Direction | EEW | Definition | | -------- | --------- | ------ | ------------ | | Rs1/imm | input | SEW | Shift amount | | Vs2 | input | SEW | Data | | Vd | output | 2\*SEW | Shifted data | Description A widening logical shift left is performed on each element of `vs2`. The elements in `vs2` are zero-extended to 2\*`SEW` bits, then shifted left by the shift amount specified by either the corresponding elements of `vs1` (vector-vector), integer register `rs1`(vector-scalar), or an immediate value (vector-immediate). Only the low log2(2\*`SEW`) bits of the shift-amount value are used, all other bits are ignored. Operation ```sail function clause execute (VWSLL_VV(vs2, vs1, vd)) = { foreach (i from vstart to vl - 1) { set_velem(vd, EEW=2*SEW, i, get_velem(vs2, i) << (get_velem(vs1, i) & ((2*SEW)-1)) ) } RETIRE_SUCCESS } function clause execute (VWSLL_VX(vs2, rs1, vd)) = { foreach (i from vstart to vl - 1) { set_velem(vd, EEW=2*SEW, i, get_velem(vs2, i) << (X(rs1) & ((2*SEW)-1)) ) } RETIRE_SUCCESS } function clause execute (VWSLL_VI(vs2, uimm[4:0], vd)) = { foreach (i from vstart to vl - 1) { set_velem(vd, EEW=2*SEW, i, get_velem(vs2, i) << (uimm[4:0] & ((2*SEW)-1)) ) } RETIRE_SUCCESS } ``` Included in [Zvbb](#zvbb) ### [](#crypto%5Fvector%5Finstructions)33.1.4\. Crypto Vector Cryptographic Instructions OP-VE (0x77) Crypto Vector instructions except Zvbb and Zvbc | Integer | Integer | FP | | | | | ------- | ------- | ------ | - | ----- | - | | funct3 | funct3 | funct3 | | | | | OPIVV | V | OPMVV | V | OPFVV | V | | OPIVX | X | OPMVX | X | OPFVF | F | | OPIVI | I | | | | | | funct6 | funct6 | funct6 | | | | ------ | ------ | ------ | ----------- | ------ | | 100000 | 100000 | V | vsm3me | 100000 | | 100001 | 100001 | V | vsm4k.vi | 100001 | | 100010 | 100010 | V | vaeskf1.vi | 100010 | | 100011 | 100011 | 100011 | | | | 100100 | 100100 | 100100 | | | | 100101 | 100101 | 100101 | | | | 100110 | 100110 | 100110 | | | | 100111 | 100111 | 100111 | | | | 101000 | 101000 | V | **VAES.vv** | 101000 | | 101001 | 101001 | V | **VAES.vs** | 101001 | | 101010 | 101010 | V | vaeskf2.vi | 101010 | | 101011 | 101011 | V | vsm3c.vi | 101011 | | 101100 | 101100 | V | vghsh | 101100 | | 101101 | 101101 | V | vsha2ms | 101101 | | 101110 | 101110 | V | vsha2ch | 101110 | | 101111 | 101111 | V | vsha2cl | 101111 | __Table 2\. VAES.vv and VAES.vs encoding space__ | vs1 | | | ----- | ------ | | 00000 | vaesdm | | 00001 | vaesdf | | 00010 | vaesem | | 00011 | vaesef | | 00111 | vaesz | | 10000 | vsm4r | | 10001 | vgmul | ### [](#crypto%5Fvector%5Finstructions%5FZvbb%5FZvbc)33.1.5\. Vector Bitmanip and Carry-less Multiply Instructions OP-V (0x57)**Zvbb**, **Zvkb**, and **Zvbc** Vector instructions **in bold** | Integer | Integer | FP | | | | | ------- | ------- | ------ | - | ----- | - | | funct3 | funct3 | funct3 | | | | | OPIVV | V | OPMVV | V | OPFVV | V | | OPIVX | X | OPMVX | X | OPFVF | F | | OPIVI | I | | | | | | funct6 | funct6 | funct6 | | | | | | | | | | | | ------ | ------ | ------------ | ---------- | ----------- | ------ | ----------- | ------- | ---------- | ------------ | ----- | ----- | ------- | | 000000 | V | X | I | vadd | 000000 | V | vredsum | 000000 | V | F | vfadd | | | 000001 | V | X | **vandn** | 000001 | V | vredand | 000001 | V | vfredusum | | | | | 000010 | V | X | vsub | 000010 | V | vredor | 000010 | V | F | vfsub | | | | 000011 | X | I | vrsub | 000011 | V | vredxor | 000011 | V | vfredosum | | | | | 000100 | V | X | vminu | 000100 | V | vredminu | 000100 | V | F | vfmin | | | | 000101 | V | X | vmin | 000101 | V | vredmin | 000101 | V | vfredmin | | | | | 000110 | V | X | vmaxu | 000110 | V | vredmaxu | 000110 | V | F | vfmax | | | | 000111 | V | X | vmax | 000111 | V | vredmax | 000111 | V | vfredmax | | | | | 001000 | 001000 | V | X | vaaddu | 001000 | V | F | vfsgnj | | | | | | 001001 | V | X | I | vand | 001001 | V | X | vaadd | 001001 | V | F | vfsgnjn | | 001010 | V | X | I | vor | 001010 | V | X | vasubu | 001010 | V | F | vfsgnjx | | 001011 | V | X | I | vxor | 001011 | V | X | vasub | 001011 | | | | | 001100 | V | X | I | vrgather | 001100 | V | X | **vclmul** | 001100 | | | | | 001101 | 001101 | V | X | **vclmulh** | 001101 | | | | | | | | | 001110 | X | I | vslideup | 001110 | X | vslide1up | 001110 | F | vfslide1up | | | | | 001110 | V | vrgatherei16 | | | | | | | | | | | | 001111 | X | I | vslidedown | 001111 | X | vslide1down | 001111 | F | vfslide1down | | | | | funct6 | funct6 | funct6 | | | | | | | | | | | ------ | ------ | --------- | -------- | ---------- | --------- | -------- | --------- | ------ | -------- | ------------ | ----- | | 010000 | V | X | I | vadc | 010000 | V | VWXUNARY0 | 010000 | V | VWFUNARY0 | | | 010000 | X | VRXUNARY0 | 010000 | F | VRFUNARY0 | | | | | | | | 010001 | V | X | I | vmadc | 010001 | 010001 | | | | | | | 010010 | V | X | vsbc | 010010 | V | VXUNARY0 | 010010 | V | VFUNARY0 | | | | 010011 | V | X | vmsbc | 010011 | 010011 | V | VFUNARY1 | | | | | | 010100 | V | X | **vror** | 010100 | V | VMUNARY0 | 010100 | | | | | | 010101 | V | X | **vrol** | 010101 | 010101 | | | | | | | | 01010x | I | **vror** | | | | | | | | | | | 010110 | 010110 | 010110 | | | | | | | | | | | 010111 | V | X | I | vmerge/vmv | 010111 | V | vcompress | 010111 | F | vfmerge/vfmv | | | 011000 | V | X | I | vmseq | 011000 | V | vmandn | 011000 | V | F | vmfeq | | 011001 | V | X | I | vmsne | 011001 | V | vmand | 011001 | V | F | vmfle | | 011010 | V | X | vmsltu | 011010 | V | vmor | 011010 | | | | | | 011011 | V | X | vmslt | 011011 | V | vmxor | 011011 | V | F | vmflt | | | 011100 | V | X | I | vmsleu | 011100 | V | vmorn | 011100 | V | F | vmfne | | 011101 | V | X | I | vmsle | 011101 | V | vmnand | 011101 | F | vmfgt | | | 011110 | X | I | vmsgtu | 011110 | V | vmnor | 011110 | | | | | | 011111 | X | I | vmsgt | 011111 | V | vmxnor | 011111 | F | vmfge | | | | funct6 | funct6 | funct6 | | | | | | | | | | | | ------ | -------- | ------ | ------ | ------- | ------ | ------ | ----- | ------ | ------ | ------ | ------ | ------- | | 100000 | V | X | I | vsaddu | 100000 | V | X | vdivu | 100000 | V | F | vfdiv | | 100001 | V | X | I | vsadd | 100001 | V | X | vdiv | 100001 | F | vfrdiv | | | 100010 | V | X | vssubu | 100010 | V | X | vremu | 100010 | | | | | | 100011 | V | X | vssub | 100011 | V | X | vrem | 100011 | | | | | | 100100 | 100100 | V | X | vmulhu | 100100 | V | F | vfmul | | | | | | 100101 | V | X | I | vsll | 100101 | V | X | vmul | 100101 | | | | | 100110 | 100110 | V | X | vmulhsu | 100110 | | | | | | | | | 100111 | V | X | vsmul | 100111 | V | X | vmulh | 100111 | F | vfrsub | | | | I | vmvr | | | | | | | | | | | | | 101000 | V | X | I | vsrl | 101000 | 101000 | V | F | vfmadd | | | | | 101001 | V | X | I | vsra | 101001 | V | X | vmadd | 101001 | V | F | vfnmadd | | 101010 | V | X | I | vssrl | 101010 | 101010 | V | F | vfmsub | | | | | 101011 | V | X | I | vssra | 101011 | V | X | vnmsub | 101011 | V | F | vfnmsub | | 101100 | V | X | I | vnsrl | 101100 | 101100 | V | F | vfmacc | | | | | 101101 | V | X | I | vnsra | 101101 | V | X | vmacc | 101101 | V | F | vfnmacc | | 101110 | V | X | I | vnclipu | 101110 | 101110 | V | F | vfmsac | | | | | 101111 | V | X | I | vnclip | 101111 | V | X | vnmsac | 101111 | V | F | vfnmsac | | funct6 | funct6 | funct6 | | | | | | | | | | ------ | ------ | --------- | -------- | --------- | ------ | ------ | ---------- | -------- | ---------- | ------ | | 110000 | V | vwredsumu | 110000 | V | X | vwaddu | 110000 | V | F | vfwadd | | 110001 | V | vwredsum | 110001 | V | X | vwadd | 110001 | V | vfwredusum | | | 110010 | 110010 | V | X | vwsubu | 110010 | V | F | vfwsub | | | | 110011 | 110011 | V | X | vwsub | 110011 | V | vfwredosum | | | | | 110100 | 110100 | V | X | vwaddu.w | 110100 | V | F | vfwadd.w | | | | 110101 | V | X | I | **vwsll** | 110101 | V | X | vwadd.w | 110101 | | | 110110 | 110110 | V | X | vwsubu.w | 110110 | V | F | vfwsub.w | | | | 110111 | 110111 | V | X | vwsub.w | 110111 | | | | | | | 111000 | 111000 | V | X | vwmulu | 111000 | V | F | vfwmul | | | | 111001 | 111001 | 111001 | | | | | | | | | | 111010 | 111010 | V | X | vwmulsu | 111010 | | | | | | | 111011 | 111011 | V | X | vwmul | 111011 | | | | | | | 111100 | 111100 | V | X | vwmaccu | 111100 | V | F | vfwmacc | | | | 111101 | 111101 | V | X | vwmacc | 111101 | V | F | vfwnmacc | | | | 111110 | 111110 | X | vwmaccus | 111110 | V | F | vfwmsac | | | | | 111111 | 111111 | V | X | vwmaccsu | 111111 | V | F | vfwnmsac | | | __Table 3\. VXUNARY0 encoding space__ | vs1 | | | ----- | ---------- | | 00010 | vzext.vf8 | | 00011 | vsext.vf8 | | 00100 | vzext.vf4 | | 00101 | vsext.vf4 | | 00110 | vzext.vf2 | | 00111 | vsext.vf2 | | 01000 | **vbrev8** | | 01001 | **vrev8** | | 01010 | **vbrev** | | 01100 | **vclz** | | 01101 | **vctz** | | 01110 | **vcpop** | ### [](#crypto%5Fvector%5Fappx%5Fsail)33.1.6\. Supporting Sail Code This section contains the supporting Sail code referenced by the instruction descriptions throughout the specification. The[Sail Manual](https://alasdair.github.io/manual.html)is recommended reading in order to best understand the supporting code. ```sail /* Auxiliary function for performing GF multiplication */ val xt2 : bits(8) -> bits(8) function xt2(x) = { (x << 1) ^ (if bit_to_bool(x[7]) then 0x1b else 0x00) } val xt3 : bits(8) -> bits(8) function xt3(x) = x ^ xt2(x) /* Multiply 8-bit field element by 4-bit value for AES MixCols step */ val gfmul : (bits(8), bits(4)) -> bits(8) function gfmul( x, y) = { (if bit_to_bool(y[0]) then x else 0x00) ^ (if bit_to_bool(y[1]) then xt2( x) else 0x00) ^ (if bit_to_bool(y[2]) then xt2(xt2( x)) else 0x00) ^ (if bit_to_bool(y[3]) then xt2(xt2(xt2(x))) else 0x00) } /* 8-bit to 32-bit partial AES Mix Column - forwards */ val aes_mixcolumn_byte_fwd : bits(8) -> bits(32) function aes_mixcolumn_byte_fwd(so) = { gfmul(so, 0x3) @ so @ so @ gfmul(so, 0x2) } /* 8-bit to 32-bit partial AES Mix Column - inverse*/ val aes_mixcolumn_byte_inv : bits(8) -> bits(32) function aes_mixcolumn_byte_inv(so) = { gfmul(so, 0xb) @ gfmul(so, 0xd) @ gfmul(so, 0x9) @ gfmul(so, 0xe) } /* 32-bit to 32-bit AES forward MixColumn */ val aes_mixcolumn_fwd : bits(32) -> bits(32) function aes_mixcolumn_fwd(x) = { let s0 : bits (8) = x[ 7.. 0]; let s1 : bits (8) = x[15.. 8]; let s2 : bits (8) = x[23..16]; let s3 : bits (8) = x[31..24]; let b0 : bits (8) = xt2(s0) ^ xt3(s1) ^ (s2) ^ (s3); let b1 : bits (8) = (s0) ^ xt2(s1) ^ xt3(s2) ^ (s3); let b2 : bits (8) = (s0) ^ (s1) ^ xt2(s2) ^ xt3(s3); let b3 : bits (8) = xt3(s0) ^ (s1) ^ (s2) ^ xt2(s3); b3 @ b2 @ b1 @ b0 /* Return value */ } /* 32-bit to 32-bit AES inverse MixColumn */ val aes_mixcolumn_inv : bits(32) -> bits(32) function aes_mixcolumn_inv(x) = { let s0 : bits (8) = x[ 7.. 0]; let s1 : bits (8) = x[15.. 8]; let s2 : bits (8) = x[23..16]; let s3 : bits (8) = x[31..24]; let b0 : bits (8) = gfmul(s0, 0xE) ^ gfmul(s1, 0xB) ^ gfmul(s2, 0xD) ^ gfmul(s3, 0x9); let b1 : bits (8) = gfmul(s0, 0x9) ^ gfmul(s1, 0xE) ^ gfmul(s2, 0xB) ^ gfmul(s3, 0xD); let b2 : bits (8) = gfmul(s0, 0xD) ^ gfmul(s1, 0x9) ^ gfmul(s2, 0xE) ^ gfmul(s3, 0xB); let b3 : bits (8) = gfmul(s0, 0xB) ^ gfmul(s1, 0xD) ^ gfmul(s2, 0x9) ^ gfmul(s3, 0xE); b3 @ b2 @ b1 @ b0 /* Return value */ } val aes_decode_rcon : bits(4) -> bits(32) function aes_decode_rcon(r) = { match r { 0x0 => 0x00000001, 0x1 => 0x00000002, 0x2 => 0x00000004, 0x3 => 0x00000008, 0x4 => 0x00000010, 0x5 => 0x00000020, 0x6 => 0x00000040, 0x7 => 0x00000080, 0x8 => 0x0000001b, 0x9 => 0x00000036, 0xA => 0x00000000, 0xB => 0x00000000, 0xC => 0x00000000, 0xD => 0x00000000, 0xE => 0x00000000, 0xF => 0x00000000 } } /* SM4 SBox - only one sbox for forwards and inverse */ let sm4_sbox_table : list(bits(8)) = [| 0xD6, 0x90, 0xE9, 0xFE, 0xCC, 0xE1, 0x3D, 0xB7, 0x16, 0xB6, 0x14, 0xC2, 0x28, 0xFB, 0x2C, 0x05, 0x2B, 0x67, 0x9A, 0x76, 0x2A, 0xBE, 0x04, 0xC3, 0xAA, 0x44, 0x13, 0x26, 0x49, 0x86, 0x06, 0x99, 0x9C, 0x42, 0x50, 0xF4, 0x91, 0xEF, 0x98, 0x7A, 0x33, 0x54, 0x0B, 0x43, 0xED, 0xCF, 0xAC, 0x62, 0xE4, 0xB3, 0x1C, 0xA9, 0xC9, 0x08, 0xE8, 0x95, 0x80, 0xDF, 0x94, 0xFA, 0x75, 0x8F, 0x3F, 0xA6, 0x47, 0x07, 0xA7, 0xFC, 0xF3, 0x73, 0x17, 0xBA, 0x83, 0x59, 0x3C, 0x19, 0xE6, 0x85, 0x4F, 0xA8, 0x68, 0x6B, 0x81, 0xB2, 0x71, 0x64, 0xDA, 0x8B, 0xF8, 0xEB, 0x0F, 0x4B, 0x70, 0x56, 0x9D, 0x35, 0x1E, 0x24, 0x0E, 0x5E, 0x63, 0x58, 0xD1, 0xA2, 0x25, 0x22, 0x7C, 0x3B, 0x01, 0x21, 0x78, 0x87, 0xD4, 0x00, 0x46, 0x57, 0x9F, 0xD3, 0x27, 0x52, 0x4C, 0x36, 0x02, 0xE7, 0xA0, 0xC4, 0xC8, 0x9E, 0xEA, 0xBF, 0x8A, 0xD2, 0x40, 0xC7, 0x38, 0xB5, 0xA3, 0xF7, 0xF2, 0xCE, 0xF9, 0x61, 0x15, 0xA1, 0xE0, 0xAE, 0x5D, 0xA4, 0x9B, 0x34, 0x1A, 0x55, 0xAD, 0x93, 0x32, 0x30, 0xF5, 0x8C, 0xB1, 0xE3, 0x1D, 0xF6, 0xE2, 0x2E, 0x82, 0x66, 0xCA, 0x60, 0xC0, 0x29, 0x23, 0xAB, 0x0D, 0x53, 0x4E, 0x6F, 0xD5, 0xDB, 0x37, 0x45, 0xDE, 0xFD, 0x8E, 0x2F, 0x03, 0xFF, 0x6A, 0x72, 0x6D, 0x6C, 0x5B, 0x51, 0x8D, 0x1B, 0xAF, 0x92, 0xBB, 0xDD, 0xBC, 0x7F, 0x11, 0xD9, 0x5C, 0x41, 0x1F, 0x10, 0x5A, 0xD8, 0x0A, 0xC1, 0x31, 0x88, 0xA5, 0xCD, 0x7B, 0xBD, 0x2D, 0x74, 0xD0, 0x12, 0xB8, 0xE5, 0xB4, 0xB0, 0x89, 0x69, 0x97, 0x4A, 0x0C, 0x96, 0x77, 0x7E, 0x65, 0xB9, 0xF1, 0x09, 0xC5, 0x6E, 0xC6, 0x84, 0x18, 0xF0, 0x7D, 0xEC, 0x3A, 0xDC, 0x4D, 0x20, 0x79, 0xEE, 0x5F, 0x3E, 0xD7, 0xCB, 0x39, 0x48 |] let aes_sbox_fwd_table : list(bits(8)) = [| 0x63, 0x7c, 0x77, 0x7b, 0xf2, 0x6b, 0x6f, 0xc5, 0x30, 0x01, 0x67, 0x2b, 0xfe, 0xd7, 0xab, 0x76, 0xca, 0x82, 0xc9, 0x7d, 0xfa, 0x59, 0x47, 0xf0, 0xad, 0xd4, 0xa2, 0xaf, 0x9c, 0xa4, 0x72, 0xc0, 0xb7, 0xfd, 0x93, 0x26, 0x36, 0x3f, 0xf7, 0xcc, 0x34, 0xa5, 0xe5, 0xf1, 0x71, 0xd8, 0x31, 0x15, 0x04, 0xc7, 0x23, 0xc3, 0x18, 0x96, 0x05, 0x9a, 0x07, 0x12, 0x80, 0xe2, 0xeb, 0x27, 0xb2, 0x75, 0x09, 0x83, 0x2c, 0x1a, 0x1b, 0x6e, 0x5a, 0xa0, 0x52, 0x3b, 0xd6, 0xb3, 0x29, 0xe3, 0x2f, 0x84, 0x53, 0xd1, 0x00, 0xed, 0x20, 0xfc, 0xb1, 0x5b, 0x6a, 0xcb, 0xbe, 0x39, 0x4a, 0x4c, 0x58, 0xcf, 0xd0, 0xef, 0xaa, 0xfb, 0x43, 0x4d, 0x33, 0x85, 0x45, 0xf9, 0x02, 0x7f, 0x50, 0x3c, 0x9f, 0xa8, 0x51, 0xa3, 0x40, 0x8f, 0x92, 0x9d, 0x38, 0xf5, 0xbc, 0xb6, 0xda, 0x21, 0x10, 0xff, 0xf3, 0xd2, 0xcd, 0x0c, 0x13, 0xec, 0x5f, 0x97, 0x44, 0x17, 0xc4, 0xa7, 0x7e, 0x3d, 0x64, 0x5d, 0x19, 0x73, 0x60, 0x81, 0x4f, 0xdc, 0x22, 0x2a, 0x90, 0x88, 0x46, 0xee, 0xb8, 0x14, 0xde, 0x5e, 0x0b, 0xdb, 0xe0, 0x32, 0x3a, 0x0a, 0x49, 0x06, 0x24, 0x5c, 0xc2, 0xd3, 0xac, 0x62, 0x91, 0x95, 0xe4, 0x79, 0xe7, 0xc8, 0x37, 0x6d, 0x8d, 0xd5, 0x4e, 0xa9, 0x6c, 0x56, 0xf4, 0xea, 0x65, 0x7a, 0xae, 0x08, 0xba, 0x78, 0x25, 0x2e, 0x1c, 0xa6, 0xb4, 0xc6, 0xe8, 0xdd, 0x74, 0x1f, 0x4b, 0xbd, 0x8b, 0x8a, 0x70, 0x3e, 0xb5, 0x66, 0x48, 0x03, 0xf6, 0x0e, 0x61, 0x35, 0x57, 0xb9, 0x86, 0xc1, 0x1d, 0x9e, 0xe1, 0xf8, 0x98, 0x11, 0x69, 0xd9, 0x8e, 0x94, 0x9b, 0x1e, 0x87, 0xe9, 0xce, 0x55, 0x28, 0xdf, 0x8c, 0xa1, 0x89, 0x0d, 0xbf, 0xe6, 0x42, 0x68, 0x41, 0x99, 0x2d, 0x0f, 0xb0, 0x54, 0xbb, 0x16 |] let aes_sbox_inv_table : list(bits(8)) = [| 0x52, 0x09, 0x6a, 0xd5, 0x30, 0x36, 0xa5, 0x38, 0xbf, 0x40, 0xa3, 0x9e, 0x81, 0xf3, 0xd7, 0xfb, 0x7c, 0xe3, 0x39, 0x82, 0x9b, 0x2f, 0xff, 0x87, 0x34, 0x8e, 0x43, 0x44, 0xc4, 0xde, 0xe9, 0xcb, 0x54, 0x7b, 0x94, 0x32, 0xa6, 0xc2, 0x23, 0x3d, 0xee, 0x4c, 0x95, 0x0b, 0x42, 0xfa, 0xc3, 0x4e, 0x08, 0x2e, 0xa1, 0x66, 0x28, 0xd9, 0x24, 0xb2, 0x76, 0x5b, 0xa2, 0x49, 0x6d, 0x8b, 0xd1, 0x25, 0x72, 0xf8, 0xf6, 0x64, 0x86, 0x68, 0x98, 0x16, 0xd4, 0xa4, 0x5c, 0xcc, 0x5d, 0x65, 0xb6, 0x92, 0x6c, 0x70, 0x48, 0x50, 0xfd, 0xed, 0xb9, 0xda, 0x5e, 0x15, 0x46, 0x57, 0xa7, 0x8d, 0x9d, 0x84, 0x90, 0xd8, 0xab, 0x00, 0x8c, 0xbc, 0xd3, 0x0a, 0xf7, 0xe4, 0x58, 0x05, 0xb8, 0xb3, 0x45, 0x06, 0xd0, 0x2c, 0x1e, 0x8f, 0xca, 0x3f, 0x0f, 0x02, 0xc1, 0xaf, 0xbd, 0x03, 0x01, 0x13, 0x8a, 0x6b, 0x3a, 0x91, 0x11, 0x41, 0x4f, 0x67, 0xdc, 0xea, 0x97, 0xf2, 0xcf, 0xce, 0xf0, 0xb4, 0xe6, 0x73, 0x96, 0xac, 0x74, 0x22, 0xe7, 0xad, 0x35, 0x85, 0xe2, 0xf9, 0x37, 0xe8, 0x1c, 0x75, 0xdf, 0x6e, 0x47, 0xf1, 0x1a, 0x71, 0x1d, 0x29, 0xc5, 0x89, 0x6f, 0xb7, 0x62, 0x0e, 0xaa, 0x18, 0xbe, 0x1b, 0xfc, 0x56, 0x3e, 0x4b, 0xc6, 0xd2, 0x79, 0x20, 0x9a, 0xdb, 0xc0, 0xfe, 0x78, 0xcd, 0x5a, 0xf4, 0x1f, 0xdd, 0xa8, 0x33, 0x88, 0x07, 0xc7, 0x31, 0xb1, 0x12, 0x10, 0x59, 0x27, 0x80, 0xec, 0x5f, 0x60, 0x51, 0x7f, 0xa9, 0x19, 0xb5, 0x4a, 0x0d, 0x2d, 0xe5, 0x7a, 0x9f, 0x93, 0xc9, 0x9c, 0xef, 0xa0, 0xe0, 0x3b, 0x4d, 0xae, 0x2a, 0xf5, 0xb0, 0xc8, 0xeb, 0xbb, 0x3c, 0x83, 0x53, 0x99, 0x61, 0x17, 0x2b, 0x04, 0x7e, 0xba, 0x77, 0xd6, 0x26, 0xe1, 0x69, 0x14, 0x63, 0x55, 0x21, 0x0c, 0x7d |] /* Lookup function - takes an index and a list, and retrieves the * x'th element of that list. */ val sbox_lookup : (bits(8), list(bits(8))) -> bits(8) function sbox_lookup(x, table) = { match (x, table) { (0x00, t0::tn) => t0, ( y, t0::tn) => sbox_lookup(x - 0x01, tn) } } /* Easy function to perform a forward AES SBox operation on 1 byte. */ val aes_sbox_fwd : bits(8) -> bits(8) function aes_sbox_fwd(x) = sbox_lookup(x, aes_sbox_fwd_table) /* Easy function to perform an inverse AES SBox operation on 1 byte. */ val aes_sbox_inv : bits(8) -> bits(8) function aes_sbox_inv(x) = sbox_lookup(x, aes_sbox_inv_table) /* AES SubWord function used in the key expansion * - Applies the forward sbox to each byte in the input word. */ val aes_subword_fwd : bits(32) -> bits(32) function aes_subword_fwd(x) = { aes_sbox_fwd(x[31..24]) @ aes_sbox_fwd(x[23..16]) @ aes_sbox_fwd(x[15.. 8]) @ aes_sbox_fwd(x[ 7.. 0]) } /* AES Inverse SubWord function. * - Applies the inverse sbox to each byte in the input word. */ val aes_subword_inv : bits(32) -> bits(32) function aes_subword_inv(x) = { aes_sbox_inv(x[31..24]) @ aes_sbox_inv(x[23..16]) @ aes_sbox_inv(x[15.. 8]) @ aes_sbox_inv(x[ 7.. 0]) } /* Easy function to perform an SM4 SBox operation on 1 byte. */ val sm4_sbox : bits(8) -> bits(8) function sm4_sbox(x) = sbox_lookup(x, sm4_sbox_table) val aes_get_column : (bits(128), nat) -> bits(32) function aes_get_column(state,c) = (state >> (to_bits(7, 32 * c)))[31..0] /* 64-bit to 64-bit function which applies the AES forward sbox to each byte * in a 64-bit word. */ val aes_apply_fwd_sbox_to_each_byte : bits(64) -> bits(64) function aes_apply_fwd_sbox_to_each_byte(x) = { aes_sbox_fwd(x[63..56]) @ aes_sbox_fwd(x[55..48]) @ aes_sbox_fwd(x[47..40]) @ aes_sbox_fwd(x[39..32]) @ aes_sbox_fwd(x[31..24]) @ aes_sbox_fwd(x[23..16]) @ aes_sbox_fwd(x[15.. 8]) @ aes_sbox_fwd(x[ 7.. 0]) } /* 64-bit to 64-bit function which applies the AES inverse sbox to each byte * in a 64-bit word. */ val aes_apply_inv_sbox_to_each_byte : bits(64) -> bits(64) function aes_apply_inv_sbox_to_each_byte(x) = { aes_sbox_inv(x[63..56]) @ aes_sbox_inv(x[55..48]) @ aes_sbox_inv(x[47..40]) @ aes_sbox_inv(x[39..32]) @ aes_sbox_inv(x[31..24]) @ aes_sbox_inv(x[23..16]) @ aes_sbox_inv(x[15.. 8]) @ aes_sbox_inv(x[ 7.. 0]) } /* * AES full-round transformation functions. */ val getbyte : (bits(64), int) -> bits(8) function getbyte(x, i) = (x >> to_bits(6, i * 8))[7..0] val aes_rv64_shiftrows_fwd : (bits(64), bits(64)) -> bits(64) function aes_rv64_shiftrows_fwd(rs2, rs1) = { getbyte(rs1, 3) @ getbyte(rs2, 6) @ getbyte(rs2, 1) @ getbyte(rs1, 4) @ getbyte(rs2, 7) @ getbyte(rs2, 2) @ getbyte(rs1, 5) @ getbyte(rs1, 0) } val aes_rv64_shiftrows_inv : (bits(64), bits(64)) -> bits(64) function aes_rv64_shiftrows_inv(rs2, rs1) = { getbyte(rs2, 3) @ getbyte(rs2, 6) @ getbyte(rs1, 1) @ getbyte(rs1, 4) @ getbyte(rs1, 7) @ getbyte(rs2, 2) @ getbyte(rs2, 5) @ getbyte(rs1, 0) } /* 128-bit to 128-bit implementation of the forward AES ShiftRows transform. * Byte 0 of state is input column 0, bits 7..0. * Byte 5 of state is input column 1, bits 15..8. */ val aes_shift_rows_fwd : bits(128) -> bits(128) function aes_shift_rows_fwd(x) = { let ic3 : bits(32) = aes_get_column(x, 3); let ic2 : bits(32) = aes_get_column(x, 2); let ic1 : bits(32) = aes_get_column(x, 1); let ic0 : bits(32) = aes_get_column(x, 0); let oc0 : bits(32) = ic3[31..24] @ ic2[23..16] @ ic1[15.. 8] @ ic0[ 7.. 0]; let oc1 : bits(32) = ic0[31..24] @ ic3[23..16] @ ic2[15.. 8] @ ic1[ 7.. 0]; let oc2 : bits(32) = ic1[31..24] @ ic0[23..16] @ ic3[15.. 8] @ ic2[ 7.. 0]; let oc3 : bits(32) = ic2[31..24] @ ic1[23..16] @ ic0[15.. 8] @ ic3[ 7.. 0]; (oc3 @ oc2 @ oc1 @ oc0) /* Return value */ } /* 128-bit to 128-bit implementation of the inverse AES ShiftRows transform. * Byte 0 of state is input column 0, bits 7..0. * Byte 5 of state is input column 1, bits 15..8. */ val aes_shift_rows_inv : bits(128) -> bits(128) function aes_shift_rows_inv(x) = { let ic3 : bits(32) = aes_get_column(x, 3); /* In column 3 */ let ic2 : bits(32) = aes_get_column(x, 2); let ic1 : bits(32) = aes_get_column(x, 1); let ic0 : bits(32) = aes_get_column(x, 0); let oc0 : bits(32) = ic1[31..24] @ ic2[23..16] @ ic3[15.. 8] @ ic0[ 7.. 0]; let oc1 : bits(32) = ic2[31..24] @ ic3[23..16] @ ic0[15.. 8] @ ic1[ 7.. 0]; let oc2 : bits(32) = ic3[31..24] @ ic0[23..16] @ ic1[15.. 8] @ ic2[ 7.. 0]; let oc3 : bits(32) = ic0[31..24] @ ic1[23..16] @ ic2[15.. 8] @ ic3[ 7.. 0]; (oc3 @ oc2 @ oc1 @ oc0) /* Return value */ } /* Applies the forward sub-bytes step of AES to a 128-bit vector * representation of its state. */ val aes_subbytes_fwd : bits(128) -> bits(128) function aes_subbytes_fwd(x) = { let oc0 : bits(32) = aes_subword_fwd(aes_get_column(x, 0)); let oc1 : bits(32) = aes_subword_fwd(aes_get_column(x, 1)); let oc2 : bits(32) = aes_subword_fwd(aes_get_column(x, 2)); let oc3 : bits(32) = aes_subword_fwd(aes_get_column(x, 3)); (oc3 @ oc2 @ oc1 @ oc0) /* Return value */ } /* Applies the inverse sub-bytes step of AES to a 128-bit vector * representation of its state. */ val aes_subbytes_inv : bits(128) -> bits(128) function aes_subbytes_inv(x) = { let oc0 : bits(32) = aes_subword_inv(aes_get_column(x, 0)); let oc1 : bits(32) = aes_subword_inv(aes_get_column(x, 1)); let oc2 : bits(32) = aes_subword_inv(aes_get_column(x, 2)); let oc3 : bits(32) = aes_subword_inv(aes_get_column(x, 3)); (oc3 @ oc2 @ oc1 @ oc0) /* Return value */ } /* Applies the forward MixColumns step of AES to a 128-bit vector * representation of its state. */ val aes_mixcolumns_fwd : bits(128) -> bits(128) function aes_mixcolumns_fwd(x) = { let oc0 : bits(32) = aes_mixcolumn_fwd(aes_get_column(x, 0)); let oc1 : bits(32) = aes_mixcolumn_fwd(aes_get_column(x, 1)); let oc2 : bits(32) = aes_mixcolumn_fwd(aes_get_column(x, 2)); let oc3 : bits(32) = aes_mixcolumn_fwd(aes_get_column(x, 3)); (oc3 @ oc2 @ oc1 @ oc0) /* Return value */ } /* Applies the inverse MixColumns step of AES to a 128-bit vector * representation of its state. */ val aes_mixcolumns_inv : bits(128) -> bits(128) function aes_mixcolumns_inv(x) = { let oc0 : bits(32) = aes_mixcolumn_inv(aes_get_column(x, 0)); let oc1 : bits(32) = aes_mixcolumn_inv(aes_get_column(x, 1)); let oc2 : bits(32) = aes_mixcolumn_inv(aes_get_column(x, 2)); let oc3 : bits(32) = aes_mixcolumn_inv(aes_get_column(x, 3)); (oc3 @ oc2 @ oc1 @ oc0) /* Return value */ } /* Performs the word rotation for AES key schedule */ val aes_rotword : bits(32) -> bits(32) function aes_rotword(x) = { let a0 : bits (8) = x[ 7.. 0]; let a1 : bits (8) = x[15.. 8]; let a2 : bits (8) = x[23..16]; let a3 : bits (8) = x[31..24]; (a0 @ a3 @ a2 @ a1) /* Return Value */ } val brev : bits(SEW) -> bits(SEW) function brev(x) = { let output : bits(SEW) = 0; foreach (i from 0 to SEW-8 by 8) output[i+7..i] = reverse_bits_in_byte(input[i+7..i]); output /* Return Value */ } val reverse_bits_in_byte : bits(8) -> bits(8) function reverse_bits_in_byte(x) = { let output : bits(8) = 0; foreach (i from 0 to 7) output[i] = x[7-i]); output /* Return Value */ } val rev8 : bits(SEW) -> bits(SEW) function rev8(x) = { // endian swap let output : bits(SEW) = 0; let j = SEW - 1; foreach (k from 0 to (SEW - 8) by 8) { output[k..(k + 7)] = x[(j - 7)..j]; j = j - 8; output /* Return Value */ } RETIRE_SUCCESS val rol32 : bits(32) -> bits(32) function ROL32(x,n) = (X << N) | (X >> (32 - N)) val sm4_subword : bits(32) -> bits(32) function sm4_subword(x) = { sm4_sbox(x[31..24]) @ sm4_sbox(x[23..16]) @ sm4_sbox(x[15.. 8]) @ sm4_sbox(x[ 7.. 0]) } ``` Vector Assembly Code Examples ==================== ## [](#vector-assembly-code-examples)Appendix A: Vector Assembly Code Examples The following are provided as non-normative text to help explain the vector ISA. ### [](#vector-vector-add-example)Vector-vector add example # vector-vector add routine of 32-bit integers # void vvaddint32(size_t n, const int*x, const int*y, int*z) # { for (size_t i=0; iThis Inner Loop Header: Depth=1 add s9, a2, s6 vsetvli s1, zero, e8,m1,ta,mu vle8.v v25, (s9) add s1, a3, s6 vle8.v v26, (s1) vadd.vv v25, v26, v25 add s1, a1, s6 vse8.v v25, (s1) add s9, a5, s10 vsetvli s1, zero, e64,m8,ta,mu vle64.v v8, (s9) add s1, a6, s10 vle64.v v16, (s1) add s1, a7, s10 vle64.v v24, (s1) add s1, s3, s10 vle64.v v0, (s1) sd a0, -112(s0) ld a0, -128(s0) vs8r.v v0, (a0) # Spill LMUL=8 add s9, t6, s10 add s11, t5, s10 add ra, t2, s10 add s1, t3, s10 vle64.v v0, (s9) ld s9, -136(s0) vs8r.v v0, (s9) # Spill LMUL=8 vle64.v v0, (s11) ld s9, -144(s0) vs8r.v v0, (s9) # Spill LMUL=8 vle64.v v0, (ra) ld s9, -160(s0) vs8r.v v0, (s9) # Spill LMUL=8 vle64.v v0, (s1) ld s1, -152(s0) vs8r.v v0, (s1) # Spill LMUL=8 vadd.vv v16, v16, v8 ld s1, -128(s0) vl8r.v v8, (s1) # Reload LMUL=8 vadd.vv v8, v8, v24 ld s1, -136(s0) vl8r.v v24, (s1) # Reload LMUL=8 ld s1, -144(s0) vl8r.v v0, (s1) # Reload LMUL=8 vadd.vv v24, v0, v24 ld s1, -128(s0) vs8r.v v24, (s1) # Spill LMUL=8 ld s1, -152(s0) vl8r.v v0, (s1) # Reload LMUL=8 ld s1, -160(s0) vl8r.v v24, (s1) # Reload LMUL=8 vadd.vv v0, v0, v24 add s1, a4, s10 vse64.v v16, (s1) add s1, s2, s10 vse64.v v8, (s1) vadd.vv v8, v8, v16 add s1, t4, s10 ld s9, -128(s0) vl8r.v v16, (s9) # Reload LMUL=8 vse64.v v16, (s1) add s9, t0, s10 vadd.vv v8, v8, v16 vle64.v v16, (s9) add s1, t1, s10 vse64.v v0, (s1) vadd.vv v8, v8, v0 vsll.vi v16, v16, 1 vadd.vv v8, v8, v16 vse64.v v8, (s9) add s6, s6, s7 add s10, s10, s8 bne s6, s4, .LBB0_4 If instead of using LMUL=1 for the 8-bit computation, the compiler is allowed to use a fractional LMUL=1/2, then the 64-bit computations can be performed using LMUL=4 (note that the same ratio of 64-bit elements and 8-bit elements is preserved as in the previous example). Now the compiler has 8 available registers to perform register allocation, resulting in no spill code, as shown in the loop below: .LBB0_4: # %vector.body # =>This Inner Loop Header: Depth=1 add s9, a2, s6 vsetvli s1, zero, e8,mf2,ta,mu // LMUL=1/2 ! vle8.v v25, (s9) add s1, a3, s6 vle8.v v26, (s1) vadd.vv v25, v26, v25 add s1, a1, s6 vse8.v v25, (s1) add s9, a5, s10 vsetvli s1, zero, e64,m4,ta,mu // LMUL=4 vle64.v v28, (s9) add s1, a6, s10 vle64.v v8, (s1) vadd.vv v28, v8, v28 add s1, a7, s10 vle64.v v8, (s1) add s1, s3, s10 vle64.v v12, (s1) add s1, t6, s10 vle64.v v16, (s1) add s1, t5, s10 vle64.v v20, (s1) add s1, a4, s10 vse64.v v28, (s1) vadd.vv v8, v12, v8 vadd.vv v12, v20, v16 add s1, t2, s10 vle64.v v16, (s1) add s1, t3, s10 vle64.v v20, (s1) add s1, s2, s10 vse64.v v8, (s1) add s9, t4, s10 vadd.vv v16, v20, v16 add s11, t0, s10 vle64.v v20, (s11) vse64.v v12, (s9) add s1, t1, s10 vse64.v v16, (s1) vsll.vi v20, v20, 1 vadd.vv v28, v8, v28 vadd.vv v28, v28, v12 vadd.vv v28, v28, v16 vadd.vv v28, v28, v20 vse64.v v28, (s11) add s6, s6, s7 add s10, s10, s8 bne s6, s4, .LBB0_4 16.1. "Zabha" Extension for Byte and Halfword Atomic Memory Operations, Version 1.0 ==================== ## [](#16-1-zabha-extension-for-byte-and-halfword-atomic-memory-operations-version-1-0)16.1\. "Zabha" Extension for Byte and Halfword Atomic Memory Operations, Version 1.0 The A-extension offers atomic memory operation (AMO) instructions for _words_,_doublewords_, and _quadwords_ (only for `AMOCAS`). The absence of atomic operations for subword data types necessitates emulation strategies. For bitwise operations, this emulation can be performed via word-sized bitwise AMO\* instructions. For non-bitwise operations, emulation is achievable using word-sized `LR`/`SC` instructions. Several limitations arise from this emulation approach: 1. In systems with large-scale or Non-Uniform Memory Access (NUMA) configurations, emulation based on `LR`/`SC` introduces issues related to scalability and fairness, particularly under conditions of high contention. 2. Emulation of narrower AMOs through wider AMO\* instructions on non-idempotent IO memory regions may result in unintended side effects. 3. Utilizing wider AMO\* instructions for emulating narrower AMOs risks activating extraneous breakpoints or watchpoints. 4. In the absence of native support for subword atomics, compilers often resort to inlining code sequences to provide the required emulation. This practice contributes to an increase in code size, with consequent impacts on system performance and memory utilization. The Zabha extension addresses these limitations by adding support for _byte_ and_halfword_ atomic memory operations to the RISC-V Unprivileged ISA. The Zabha extension depends upon the Zaamo standard extension. ### [](#16-1-1-byte-and-halfword-atomic-memory-operation-instructions)16.1.1\. Byte and Halfword Atomic Memory Operation Instructions Zabha extension provides the `AMO[ADD|AND|OR|XOR|SWAP|MIN[U]|MAX[U]].[B|H]`instructions. If Zacas extension is also implemented, Zabha further provides the`AMOCAS.[B|H]` instructions. ![zabha-ext-wavedrom-reg](_images/zabha-ext-wavedrom-reg-03dfe23087d49235276711642d51aa7df7792a99.svg) Byte and halfword AMOs always sign-extend the value placed in `rd`, and ignore the bits of the original value in `rs2`. The`AMOCAS.[B|H]` instructions similarly ignore the bits of the original value in `rd`. Similar to the AMOs specified in the A extension, the Zabha extension mandates that the address contained in the `rs1` register must be naturally aligned to the size of the operand. The same exception options as specified in the A extension are applicable in cases where the address is not naturally aligned. Similar to the AMOs specified in the A and Zacas extensions, the AMOs in the Zabha extension optionally provide release consistency semantics, using the `aq`and `rl` bits, to help implement multiprocessor synchronization. | | Zabha omits _byte_ and _halfword_ support for LR and SC due to low utility. | | ------------------------------------------------------------------------------ | 15.1. "Zacas" Extension for Atomic Compare-and-Swap (CAS) Instructions, Version 1.0.0 ==================== ## [](#15-1-zacas-extension-for-atomic-compare-and-swap-cas-instructions-version-1-0-0)15.1\. "Zacas" Extension for Atomic Compare-and-Swap (CAS) Instructions, Version 1.0.0 Compare-and-Swap (CAS) provides an easy and typically faster way to perform thread synchronization operations when supported as a hardware instruction. CAS is typically used by lock-free and wait-free algorithms. This extension defines CAS instructions to operate on 32-bit, 64-bit, and 128-bit (RV64 only) data values. The Zacas extension depends upon the Zaamo extension. ### [](#15-1-1-worddoublewordquadword-cas-amocas-wdq-instructions)15.1.1\. Word/Doubleword/Quadword CAS (AMOCAS.W/D/Q) Instructions ![svg](_images/svg-01fe6c8647c106825955b54f1d9f8f7883e13011.svg) For RV32, `AMOCAS.W` atomically loads a 32-bit data value from address in `rs1`, compares the loaded value to the 32-bit value held in `rd`, and if the comparison is bitwise equal, then stores the 32-bit value held in `rs2` to the original address in `rs1`. The value loaded from memory is placed into register `rd`. The operation performed by `AMOCAS.W` for RV32 is as follows: temp = mem[X(rs1)] if ( temp == X(rd) ) mem[X(rs1)] = X(rs2) X(rd) = temp `AMOCAS.D` is similar to `AMOCAS.W` but operates on 64-bit data values. For RV32, `AMOCAS.D` atomically loads 64-bits of a data value from address in`rs1`, compares the loaded value to a 64-bit value held in a register pair consisting of `rd` and `rd+1`, and if the comparison is bitwise equal, then stores the 64-bit value held in the register pair `rs2` and `rs2+1` to the original address in `rs1`. The value loaded from memory is placed into the register pair `rd` and `rd+1`. The instruction requires the first register in the pair to be even numbered; encodings with odd numbered registers specified in `rs2` and `rd` are reserved. When the first register of a source register pair is `x0`, then both halves of the pair read as zero. When the first register of a destination register pair is `x0`, then the entire register result is discarded and neither destination register is written.The operation performed by `AMOCAS.D` for RV32 is as follows: temp0 = mem[X(rs1)+0] temp1 = mem[X(rs1)+4] comp0 = (rd == x0) ? 0 : X(rd) comp1 = (rd == x0) ? 0 : X(rd+1) swap0 = (rs2 == x0) ? 0 : X(rs2) swap1 = (rs2 == x0) ? 0 : X(rs2+1) if ( temp0 == comp0 ) && ( temp1 == comp1 ) mem[X(rs1)+0] = swap0 mem[X(rs1)+4] = swap1 endif if ( rd != x0 ) X(rd) = temp0 X(rd+1) = temp1 endif For RV64, `AMOCAS.W` atomically loads a 32-bit data value from address in`rs1`, compares the loaded value to the lower 32 bits of the value held in `rd`, and if the comparison is bitwise equal, then stores the lower 32 bits of the value held in `rs2` to the original address in `rs1`. The 32-bit value loaded from memory is sign-extended and is placed into register `rd`. The operation performed by `AMOCAS.W` for RV64 is as follows: temp[31:0] = mem[X(rs1)] if ( temp[31:0] == X(rd)[31:0] ) mem[X(rs1)] = X(rs2)[31:0] X(rd) = SignExtend(temp[31:0]) For RV64, `AMOCAS.D` atomically loads 64-bits of a data value from address in`rs1`, compares the loaded value to a 64-bit value held in `rd`, and if the comparison is bitwise equal, then stores the 64-bit value held in `rs2` to the original address in `rs1`. The value loaded from memory is placed into register`rd`. The operation performed by `AMOCAS.D` for RV64 is as follows: temp = mem[X(rs1)] if ( temp == X(rd) ) mem[X(rs1)] = X(rs2) X(rd) = temp `AMOCAS.Q` (RV64 only) atomically loads 128-bits of a data value from address in`rs1`, compares the loaded value to a 128-bit value held in a register pair consisting of `rd` and `rd+1`, and if the comparison is bitwise equal, then stores the 128-bit value held in the register pair `rs2` and `rs2+1` to the original address in `rs1`. The value loaded from memory is placed into the register pair `rd` and `rd+1`. The instruction requires the first register in the pair to be even numbered; encodings with odd numbered registers specified in`rs2` and `rd` are reserved. When the first register of a source register pair is `x0`, then both halves of the pair read as zero. When the first register of a destination register pair is `x0`, then the entire register result is discarded and neither destination register is written. The operation performed by`AMOCAS.Q` is as follows: temp0 = mem[X(rs1)+0] temp1 = mem[X(rs1)+8] comp0 = (rd == x0) ? 0 : X(rd) comp1 = (rd == x0) ? 0 : X(rd+1) swap0 = (rs2 == x0) ? 0 : X(rs2) swap1 = (rs2 == x0) ? 0 : X(rs2+1) if ( temp0 == comp0 ) && ( temp1 == comp1 ) mem[X(rs1)+0] = swap0 mem[X(rs1)+8] = swap1 endif if ( rd != x0 ) X(rd) = temp0 X(rd+1) = temp1 endif | | Some algorithms may load the previous data value of a memory location into the register used as the compare data value source by a Zacas instruction. When using a Zacas instruction that uses a register pair to source the compare value, the two registers may be loaded using two individual loads. The two individual loads may read an inconsistent pair of values but that is not an issue since theAMOCAS operation itself uses an atomic load-pair from memory to obtain the data value for its comparison. The following example code sequence illustrates the use of AMOCAS.D in a RV32 implementation to atomically increment a 64-bit counter. \# a0 - address of the counter. increment: lw a2, (a0) # Load current counter value using lw a3, 4(a0) # two individual loads. retry: mv a6, a2 # Save the low 32 bits of the current value. mv a7, a3 # Save the high 32 bits of the current value. addi a4, a2, 1 # Increment the low 32 bits. sltu a1, a4, a2 # Determine if there is a carry out. add a5, a3, a1 # Add the carry if any to high 32 bits. amocas.d.aqrl a2, a4, (a0) bne a2, a6, retry # If amocas.d failed then retry bne a3, a7, retry # using current values loaded by amocas.d. ret | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Just as for AMOs in the A extension, `AMOCAS.W/D/Q` requires that the address held in `rs1` be naturally aligned to the size of the operand (i.e., 16-byte aligned for _quadwords_, eight-byte aligned for _doublewords_, and four-byte aligned for _words_). And the same exception options apply if the address is not naturally aligned. Just as for AMOs in the A extension, the `AMOCAS.W/D/Q` optionally provide release consistency semantics, using the `aq` and `rl` bits, to help implement multiprocessor synchronization. The memory operation performed by an`AMOCAS.W/D/Q`, when successful, has acquire semantics if `aq` bit is 1 and has release semantics if `rl` bit is 1. The memory operation performed by an`AMOCAS.W/D/Q`, when not successful, has acquire semantics if `aq` bit is 1 but does not have release semantics, regardless of `rl`. A FENCE instruction may be used to order the memory read access and, if produced, the memory write access by an `AMOCAS.W/D/Q` instruction. | | An unsuccessful AMOCAS.W/D/Q may either not perform a memory write or may write back the old value loaded from memory. The memory write, if produced, does not have release semantics, regardless of rl. Irrespective of whether a write is actually performed, the instruction is treated as an AMO for the purposes of the RVWMO PPO rules. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | An `AMOCAS.W/D/Q` instruction always requires write permissions. | | The following example code sequence illustrates the use of AMOCAS.Q to implement the _enqueue_ operation for a non-blocking concurrent queue using the algorithm outlined in \[[20](../biblio/bibliography.html#bib-queue)\]. The algorithm atomically operates on a pointer and its associated modification counter using the AMOCAS.Q instruction to avoid the ABA problem. \# Enqueue operation of a non-blocking concurrent queue. \# Data structures used by the queue: \# structure pointer\_t {ptr: node\_t \*, count: uint64\_t} \# structure node\_t {next: pointer\_t, value: data type} \# structure queue\_t {Head: pointer\_t, Tail: pointer\_t} \# Inputs to the procedure: \# a0 - address of Tail variable \# a4 - address of a new node to insert at tail enqueue: ld a6, (a0) # a6 = Tail.ptr ld a7, 8(a0) # a7 = Tail.count ld a2, (a6) # a2 = Tail.ptr->next.ptr ld a3, 8(a6) # a3 = Tail.ptr->next.count ld t1, (a0) ld t2, 8(a0) bne a6, t1, enqueue # Retry if Tail & next are not consistent bne a7, t2, enqueue # Retry if Tail & next are not consistent bne a2, x0, move\_tail # Was tail pointing to the last node? mv t1, a2 # Save Tail.ptr->next.ptr mv t2, a3 # Save Tail.ptr->next.count addi a5, a3, 1 # Link the node at the end of the list amocas.q.aqrl a2, a4, (a6) bne a2, t1, enqueue # Retry if CAS failed bne a3, t2, enqueue # Retry if CAS failed addi a5, a7, 1 # Update Tail to the inserted node amocas.q.aqrl a6, a4, (a0) ret # Enqueue done move\_tail: # Tail was not pointing to the last node addi a3, a7, 1 # Try to swing Tail to the next node amocas.q.aqrl a6, a2, (a0) j enqueue # Retry | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | 17.1. "Zalasr" Atomic Load-Acquire and Store-Release Instructions, Version 1.0 ==================== ## [](#17-1-zalasr-atomic-load-acquire-and-store-release-instructions-version-1-0)17.1\. "Zalasr" Atomic Load-Acquire and Store-Release Instructions, Version 1.0 The Zalasr (Load-Acquire and Store-Release) extension provides load-acquire and store-release instructions in RISC-V. These can be important for high performance designs by enabling finer-grained synchronisation than is possible with fences alone, by providing a unidirectional fence. Load-acquire and store-release are widely used in language-level memory models: both the Java and C++ memory models make use of acquire-release semantics, and C++'s `atomic` provides primitives that are meant to map directly to load-acquire and store-release instructions. The Zalasr extension builds on the atomic support provided by the Zaamo (Atomic Memory Operations), Zalrsc (Load-Reserved and Store-Conditional), and Zabha (Byte and Halfword Atomic Memory Operations) extensions by providing additional atomic operations (although it can be implemented independently of them). All of the AMO operations in Zaamo (and Zabha) are read-modify-write operations that both load and store. The Zalrsc extension provides operations that are only loads or stores. However, since it is designed to perform an atomic operation on a single memory word or doubleword, the loads and stores are designed to be paired. The load-reserved implies that a future store-conditional will follow while store-conditional requires that there was a previous load-reserved without other intervening loads or stores. Therefore, the Zalrsc extension does not provide a general atomic and ordered load or store. Zalasr fills this gap by offering truly standalone atomic and ordered loads and stores. The Zalasr instructions are atomic loads and stores that support ordering annotations. With the combination of Zaamo, Zabha, and Zalasr all C++ atomic operations can be supported with single instructions. ### [](#17-1-1-load-acquire-and-store-release-instructions)17.1.1\. Load-Acquire and Store-Release Instructions The Zalasr instructions always sign-extend the value placed in _rd_ and ignore the upper bits of the value of _rs2_. The instructions in the Zalasr extension require that the address held in _rs1_ be naturally aligned to the size in bytes (2width) of the operand. If the address is not naturally aligned, an address-misaligned exception or an access-fault exception will be generated. The access-fault exception can be generated for a memory access that would otherwise be able to complete except for the misalignment, if the misaligned access should not be emulated. The misaligned atomicity granule PMA, defined in Volume II of this manual, optionally relaxes this alignment requirement. If all accessed bytes lie within the same misaligned atomicity granule, the instruction will not raise an exception for reasons of address alignment, and the instruction will give rise to only one memory operation for the purposes of RVWMO—i.e., it will execute atomically. ### [](#insns-ldatomic)17.1.2\. Load Acquire Synopsis The load-acquire instruction atomically loads a 2width\-byte value from the address in _rs1_ and places the sign-extended value into the register _rd_, subject to the ordering annotations specified in the instruction. Mnemonic lb.{aq,aqrl} _rd_, (_rs1_) lh.{aq,aqrl} _rd_, (_rs1_) lw.{aq,aqrl} _rd_, (_rs1_) ld.{aq,aqrl} _rd_, (_rs1_) Encoding ![svg](_images/svg-5b9a9fe017ba4b0995a69ff0644b5639fadde3da.svg) Description This instruction loads 2width bytes of memory from rs1 atomically and writes the result into rd. If the size (2width+3) is less than XLEN, it is sign-extended to fill the destination register. This load must have the ordering annotation _aq_ and may have ordering annotation _rl_ encoded in the instruction. The instruction always has an "acquire-RCsc" annotation, and if the bit _rl_ is set the instruction has a "release-RCsc" annotation. The versions without the _aq_ bit set are RESERVED. LD.{AQ, AQRL} is RV64-only. | | The _aq_ bit is mandatory because the two encodings that would be produced are not seen as useful at this time. The version with neither the _aq_ nor the _rl_ bit set would correspond to a load with no ordering annotations that was guaranteed to be performed atomically. This can be achieved with ordinary load instructions by suitably aligning pointers. The version with only the _rl_ bit would correspond to load-release. Load-release has theoretical applications in seqlocks, but is not supported in language-level memory models and so is not included. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#insns-sdatomic)17.1.3\. Store Release Synopsis The store-release instruction atomically stores the 2width\-byte value from the low bits of register _rs2_ to the address in _rs1_, subject to the ordering annotations specified in the instruction. Mnemonic sb.{rl,aqrl} _rs2_, (_rs1_) sh.{rl,aqrl} _rs2_, (_rs1_) sw.{rl,aqrl} _rs2_, (_rs1_) sd.{rl,aqrl} _rs2_, (_rs1_) Encoding ![svg](_images/svg-c906b6a9d84bf2fa00d265b08050050ebd597eca.svg) Description This instruction stores 2width bytes of memory from rs1 atomically. This store must have ordering annotation _rl_ and may have ordering annotation _aq_ encoded in the instruction. The instruction always has an "release-RCsc" annotation, and if the bit _aq_ is set the instruction has a "acquire-RCsc" annotation. The versions without the _rl_ bit set are RESERVED. SD.{RL, AQRL} is RV64-only. | | The _rl_ bit is mandatory because the two encodings that would be produced are not seen as useful at this time. The version with neither the _aq_ nor the _rl_ bit set would correspond to a store with no ordering annotations that was guaranteed to be performed atomically. This can be achieved with ordinary store instructions by suitably aligned pointers. The version with only the _aq_ bit would correspond to store-acquire. Store-acquire has theoretical applications in seqlocks, but is not supported in language-level memory models and so is not included. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 14.1. "Zawrs" Extension for Wait-on-Reservation-Set instructions, Version 1.01 ==================== ## [](#14-1-zawrs-extension-for-wait-on-reservation-set-instructions-version-1-01)14.1\. "Zawrs" Extension for Wait-on-Reservation-Set instructions, Version 1.01 The Zawrs extension defines a pair of instructions to be used in polling loops that allows a core to enter a low-power state and wait on a store to a memory location. Waiting for a memory location to be updated is a common pattern in many use cases such as: 1. Contenders for a lock waiting for the lock variable to be updated. 2. Consumers waiting on the tail of an empty queue for the producer to queue work/data. The producer may be code executing on a RISC-V hart, an accelerator device, an external I/O agent. 3. Code waiting on a flag to be set in memory indicative of an event occurring. For example, software on a RISC-V hart may wait on a "done" flag to be set in memory by an accelerator device indicating completion of a job previously submitted to the device. Such use cases involve polling on memory locations, and such busy loops can be a wasteful expenditure of energy. To mitigate the wasteful looping in such usages, a `WRS.NTO` (WRS-with-no-timeout) instruction is provided. Instead of polling for a store to a specific memory location, software registers a reservation set that includes all the bytes of the memory location using the `LR` instruction.Then a subsequent `WRS.NTO` instruction would cause the hart to temporarily stall execution in a low-power state until a store occurs to the reservation set or an interrupt is observed. Sometimes the program waiting on a memory update may also need to carry out a task at a future time or otherwise place an upper bound on the wait. To support such use cases a second instruction `WRS.STO` (WRS-with-short-timeout) is provided that works like `WRS.NTO` but bounds the stall duration to an implementation-define short timeout such that the stall is terminated on the timeout if no other conditions have occurred to terminate the stall. The program using this instruction may then determine if its deadline has been reached. | | The instructions in the Zawrs extension are only useful in conjunction with the LR instruction, which is provided by the Zalrsc component of the A extension. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#Zawrs)14.1.1\. Wait-on-Reservation-Set Instructions The `WRS.NTO` and `WRS.STO` instructions cause the hart to temporarily stall execution in a low-power state as long as the reservation set is valid and no pending interrupts, even if disabled, are observed. For `WRS.STO` the stall duration is bounded by an implementation defined short timeout. These instructions are available in all privilege modes. ![svg](_images/svg-92bd2960c14bdb39610458fd6f079cbada57aeee.svg) Hart execution may be stalled while the following conditions are all satisfied: 1. The reservation set is valid 2. If `WRS.STO`, a "short" duration since start of stall has not elapsed 3. No pending interrupt is observed (see the rules below) While stalled, an implementation is permitted to occasionally terminate the stall and complete execution for any reason. `WRS.NTO` and `WRS.STO` instructions follow the rules of the `WFI` instruction for resuming execution on a pending interrupt. When the `TW` (Timeout Wait) bit in `mstatus` is set and `WRS.NTO` is executed in any privilege mode other than M mode, and it does not complete within an implementation-specific bounded time limit, the `WRS.NTO` instruction will cause an illegal-instruction exception. When executing in VS or VU mode, if the `VTW` bit is set in `hstatus`, the`TW` bit in `mstatus` is clear, and the `WRS.NTO` does not complete within an implementation-specific bounded time limit, the `WRS.NTO` instruction will cause a virtual-instruction exception. | | Since the WRS.STO and WRS.NTO instructions can complete execution for reasons other than stores to the reservation set, software will likely need a means of looping until the required stores have occurred. The duration of a WRS.STO instruction’s timeout may vary significantly within and among implementations. In typical implementations this duration should be roughly in the range of 10 to 100 times an on-chip cache miss latency or a cacheless access to main memory. WRS.NTO, unlike WFI, is not specified to cause an illegal-instruction exception if executed in U-mode when the governing TW bit is 0\. WFI is typically not expected to be used in U-mode and on many systems may promptly cause an illegal-instruction exception if used at U-mode. Unlike WFI,WRS.NTO is expected to be used by software in U-mode when waiting on memory but without a deadline for that wait. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 29.1. "Zc*" Extension for Code Size Reduction, Version 1.0.0 ==================== ## [](#Zc)29.1\. "Zc\*" Extension for Code Size Reduction, Version 1.0.0 ### [](#29-1-1-zc-overview)29.1.1\. Zc\* Overview Zc\* is a group of extensions that define subsets of the existing C extension (Zca, Zcd, Zcf) and new extensions which only contain 16-bit encodings. Zcm\* all reuse the encodings for _c.fld_, _c.fsd_, _c.fldsp_, _c.fsdsp_. __Table 1\. Zc\* extension overview__ | Instruction | Zca | Zcf | Zcd | Zcb | Zcmp | Zcmt | | ----------------------------------------------------------------------------------------------------------------------------------------- | ---- | --- | --- | --- | ---- | ---- | | **The Zca extension is added as way to refer to instructions in the C extension that do not include the floating-point loads and stores** | | | | | | | | C excl. c.f\* | yes | | | | | | | **The Zcf extension is added as a way to refer to compressed single-precision floating-point load/stores** | | | | | | | | c.flw | rv32 | | | | | | | c.flwsp | rv32 | | | | | | | c.fsw | rv32 | | | | | | | c.fswsp | rv32 | | | | | | | **The Zcd extension is added as a way to refer to compressed double-precision floating-point load/stores** | | | | | | | | c.fld | yes | | | | | | | c.fldsp | yes | | | | | | | c.fsd | yes | | | | | | | c.fsdsp | yes | | | | | | | **Simple operations for use on all architectures** | | | | | | | | c.lbu | yes | | | | | | | c.lh | yes | | | | | | | c.lhu | yes | | | | | | | c.sb | yes | | | | | | | c.sh | yes | | | | | | | c.zext.b | yes | | | | | | | c.sext.b | yes | | | | | | | c.zext.h | yes | | | | | | | c.sext.h | yes | | | | | | | c.zext.w | yes | | | | | | | c.mul | yes | | | | | | | c.not | yes | | | | | | | **PUSH/POP and double move which overlap with _c.fsdsp_. Complex operations intended for embedded CPUs** | | | | | | | | cm.push | yes | | | | | | | cm.pop | yes | | | | | | | cm.popret | yes | | | | | | | cm.popretz | yes | | | | | | | cm.mva01s | yes | | | | | | | cm.mvsa01 | yes | | | | | | | **Table jump which overlaps with _c.fsdsp_. Complex operations intended for embedded CPUs** | | | | | | | | cm.jt | yes | | | | | | | cm.jalt | yes | | | | | | ### [](#C)29.1.2\. C The C extension is the superset of the following extensions: * Zca * Zcf if F is specified (RV32 only) * Zcd if D is specified As C defines the same instructions as Zca, Zcf, and Zcd, the rule is that: * C always implies Zca * C+F implies Zcf (RV32 only) * C+D implies Zcd ### [](#29-1-3-zce)29.1.3\. Zce The Zce extension is intended to be used for microcontrollers, and includes all relevant Zc extensions. * Specifying Zce on RV32 without F includes Zca, Zcb, Zcmp, Zcmt * Specifying Zce on RV32 with F includes Zca, Zcb, Zcmp, Zcmt _and_ Zcf * Specifying Zce on RV64 always includes Zca, Zcb, Zcmp, Zcmt * Zcf doesn’t exist for RV64 Therefore common ISA strings can be updated as follows to include the relevant Zc extensions, for example: * RV32IMC becomes RV32IM\_Zce * RV32IMCF becomes RV32IMF\_Zce ### [](#misaC)29.1.4\. MISA.C MISA.C is set if the following extensions are selected: * Zca and not F * Zca, Zcf and F (but not D) is specified (RV32 only) * Zca, Zcf and Zcd if D is specified (RV32 only) * this configuration excludes Zcmp, Zcmt * Zca, Zcd if D is specified (RV64 only) * this configuration excludes Zcmp, Zcmt ### [](#29-1-5-zca)29.1.5\. Zca The Zca extension is added as way to refer to instructions in the C extension that do not include the floating-point loads and stores. Therefore it _excluded_ all 16-bit floating point loads and stores: _c.flw_, _c.flwsp_, _c.fsw_, _c.fswsp_, _c.fld_, _c.fldsp_, _c.fsd_, _c.fsdsp_. | | the C extension only includes F/D instructions when D and F are also specified | | --------------------------------------------------------------------------------- | ### [](#29-1-6-zcf-rv32-only)29.1.6\. Zcf (RV32 only) Zcf is the existing set of compressed single precision floating point loads and stores: _c.flw_, _c.flwsp_, _c.fsw_, _c.fswsp_. Zcf is only relevant to RV32, it cannot be specified for RV64. The Zcf extension depends on the [Zca](#29-1-5-zca) and F extensions. ### [](#29-1-7-zcd)29.1.7\. Zcd Zcd is the existing set of compressed double precision floating point loads and stores: _c.fld_, _c.fldsp_, _c.fsd_, _c.fsdsp_. The Zcd extension depends on the [Zca](#29-1-5-zca) and D extensions. ### [](#29-1-8-zcb)29.1.8\. Zcb Zcb has simple code-size saving instructions which are easy to implement on all CPUs. All encodings are currently reserved for all architectures, and have no conflicts with any existing extensions. | | Zcb can be implemented on _any_ CPU as the instructions are 16-bit versions of existing 32-bit instructions from the application class profile. | | -------------------------------------------------------------------------------------------------------------------------------------------------- | The Zcb extension depends on the [Zca](#29-1-5-zca) extension. As shown on the individual instruction pages, many of the instructions in Zcb depend upon another extension being implemented. For example, _c.mul_ is only implemented if M or Zmmul is implemented, and _c.sext.b_ is only implemented if Zbb is implemented. The _c.mul_ encoding uses the CA register format along with other instructions such as _c.sub_, _c.xor_ etc. | | _c.sext.w_ is a pseudoinstruction for _c.addiw rd, 0_ (RV64) | | --------------------------------------------------------------- | | RV32 | RV64 | Mnemonic | Instruction | | ---- | --------------- | -------------------------------------------------------- | ------------------------------------------------------------ | | yes | yes | c.lbu _rd'_, uimm(_rs1'_) | [Load unsigned byte, 16-bit encoding](#insns-c%5Flbu) | | yes | yes | c.lhu _rd'_, uimm(_rs1'_) | [Load unsigned halfword, 16-bit encoding](#insns-c%5Flhu) | | yes | yes | c.lh _rd'_, uimm(_rs1'_) | [Load signed halfword, 16-bit encoding](#insns-c%5Flh) | | yes | yes | c.sb _rs2'_, uimm(_rs1'_) | [Store byte, 16-bit encoding](#insns-c%5Fsb) | | yes | yes | c.sh _rs2'_, uimm(_rs1'_) | [Store halfword, 16-bit encoding](#insns-c%5Fsh) | | yes | yes | c.zext.b _rsd'_ | [Zero extend byte, 16-bit encoding](#insns-c%5Fzext%5Fb) | | yes | yes | c.sext.b _rsd'_ | [Sign extend byte, 16-bit encoding](#insns-c%5Fsext%5Fb) | | yes | yes | c.zext.h _rsd'_ | [Zero extend halfword, 16-bit encoding](#insns-c%5Fzext%5Fh) | | yes | yes | c.sext.h _rsd'_ | [Sign extend halfword, 16-bit encoding](#insns-c%5Fsext%5Fh) | | yes | c.zext.w _rsd'_ | [Zero extend word, 16-bit encoding](#insns-c%5Fzext%5Fw) | | | yes | yes | c.not _rsd'_ | [Bitwise not, 16-bit encoding](#insns-c%5Fnot) | | yes | yes | c.mul _rsd'_, _rs2'_ | [Multiply, 16-bit encoding](#insns-c%5Fmul) | ### [](#Zcmp)29.1.9\. Zcmp The Zcmp extension is a set of instructions which may be executed as a series of existing 32-bit RISC-V instructions. This extension reuses some encodings from _c.fsdsp_. Therefore it is _incompatible_ with [Zcd](#29-1-7-zcd), which is included when C and D extensions are both present. | | Zcmp is primarily targeted at embedded class CPUs due to implementation complexity. Additionally, it is not compatible with application class profiles. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | The Zcmp extension depends on the [Zca](#29-1-5-zca) extension. The PUSH/POP assembly syntax uses several variables, the meaning of which are: * _reg\_list_ is a list containing 1 to 13 registers (ra and 0 to 12 s registers) * valid values: {ra}, \\{ra, s0}, \\{ra, s0-s1}, \\{ra, s0-s2}, …​, \\{ra, s0-s8}, \\{ra, s0-s9}, \\{ra, s0-s11} * note that \\{ra, s0-s10} is _not_ valid, giving 12 lists not 13 for better encoding * _stack\_adj_ is the total size of the stack frame. * valid values vary with register list length and the specific encoding, see the instruction pages for details. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | ------------------------------------ | ------------------------------------------------------------------- | | yes | yes | cm.push _{reg\_list}, -stack\_adj_ | [cm.push](#insns-cm%5Fpush) | | yes | yes | cm.pop _{reg\_list}, stack\_adj_ | [cm.pop](#insns-cm%5Fpop) | | yes | yes | cm.popret _{reg\_list}, stack\_adj_ | [cm.popret](#insns-cm%5Fpopret) | | yes | yes | cm.popretz _{reg\_list}, stack\_adj_ | [cm.popretz](#insns-cm%5Fpopretz) | | yes | yes | cm.mva01s _rs1', rs2'_ | [Move two s0-s7 registers into a0-a1](#insns-cm%5Fmva01s) | | yes | yes | cm.mvsa01 _r1s', r2s'_ | [Move a0-a1 into two different s0-s7 registers](#insns-cm%5Fmvsa01) | ### [](#Zcmt)29.1.10\. Zcmt Zcmt adds the table jump instructions and also adds the jvt CSR. The jvt CSR requires a state enable if Smstateen is implemented. See [jvt CSR, table jump base vector and control register](#csrs-jvt) for details. This extension reuses some encodings from _c.fsdsp_. Therefore it is _incompatible_ with [Zcd](#29-1-7-zcd), which is included when C and D extensions are both present. | | Zcmt is primarily targeted at embedded class CPUs due to implementation complexity. Additionally, it is not compatible with RVA profiles. | | -------------------------------------------------------------------------------------------------------------------------------------------- | The Zcmt extension depends on the [Zca](#29-1-5-zca) and Zicsr extensions. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | --------------- | ------------------------------------------- | | yes | yes | cm.jt _index_ | [Jump via table](#insns-cm%5Fjt) | | yes | yes | cm.jalt _index_ | [Jump and link via table](#insns-cm%5Fjalt) | ### [](#Zc%5Fformats)29.1.11\. Zc instruction formats Several instructions in this specification use the following new instruction formats. | Format | instructions | 15:10 | 9 | 8 | 7 | 6 | 5 | 4 | 3 | 2 | 1 | 0 | | ------ | --------------------- | ------ | -------- | ------ | ----- | ---- | -- | - | - | - | - | - | | CLB | c.lbu | funct6 | rs1' | uimm | rd' | op | | | | | | | | CSB | c.sb | funct6 | rs1' | uimm | rs2' | op | | | | | | | | CLH | c.lhu, c.lh | funct6 | rs1' | funct1 | uimm | rd' | op | | | | | | | CSH | c.sh | funct6 | rs1' | funct1 | uimm | rs2' | op | | | | | | | CU | c.\[sz\]ext.\*, c.not | funct6 | rd'/rs1' | funct5 | op | | | | | | | | | CMMV | cm.mvsa01 cm.mva01s | funct6 | r1s' | funct2 | r2s' | op | | | | | | | | CMJT | cm.jt cm.jalt | funct6 | index | op | | | | | | | | | | CMPP | cm.push\*, cm.pop\* | funct6 | funct2 | urlist | spimm | op | | | | | | | | | c.mul uses the existing CA format. | | ------------------------------------- | ### [](#Zcb%5Finstructions)29.1.12\. Zcb instructions #### [](#insns-c%5Flbu)29.1.12.1\. c.lbu Synopsis Load unsigned byte, 16-bit encoding Mnemonic c.lbu _rd'_, _uimm_(_rs1'_) Encoding (RV32, RV64): ![svg](_images/svg-1945561bbf8fdcbce6e5b815ad424f2cca0a4ee2.svg) The immediate offset is formed as follows: ```sail uimm[31:2] = 0; uimm[1] = encoding[5]; uimm[0] = encoding[6]; ``` Description This instruction loads a byte from the memory address formed by adding _rs1'_ to the zero extended immediate _uimm_. The resulting byte is zero extended to XLEN bits and is written to _rd'_. | | _rd'_ and _rs1'_ are from the standard 8-register set x8-x15. | | ---------------------------------------------------------------- | Prerequisites None Operation ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. X(rdc) = EXTZ(mem[X(rs1c)+EXTZ(uimm)][7..0]); ``` #### [](#insns-c%5Flhu)29.1.12.2\. c.lhu Synopsis Load unsigned halfword, 16-bit encoding Mnemonic c.lhu _rd'_, _uimm_(_rs1'_) Encoding (RV32, RV64): ![svg](_images/svg-8e462ffff47c5aa82fa19f1f397053723b490172.svg) The immediate offset is formed as follows: ```sail uimm[31:2] = 0; uimm[1] = encoding[5]; uimm[0] = 0; ``` Description This instruction loads a halfword from the memory address formed by adding _rs1'_ to the zero extended immediate _uimm_. The resulting halfword is zero extended to XLEN bits and is written to _rd'_. | | _rd'_ and _rs1'_ are from the standard 8-register set x8-x15. | | ---------------------------------------------------------------- | Prerequisites None Operation ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. X(rdc) = EXTZ(load_mem[X(rs1c)+EXTZ(uimm)][15..0]); ``` #### [](#insns-c%5Flh)29.1.12.3\. c.lh Synopsis Load signed halfword, 16-bit encoding Mnemonic c.lh _rd'_, _uimm_(_rs1'_) Encoding (RV32, RV64): ![svg](_images/svg-c6a953d9d2f7d4ec2703dcfa1c5db61495bb75b8.svg) The immediate offset is formed as follows: ```sail uimm[31:2] = 0; uimm[1] = encoding[5]; uimm[0] = 0; ``` Description This instruction loads a halfword from the memory address formed by adding _rs1'_ to the zero extended immediate _uimm_. The resulting halfword is sign extended to XLEN bits and is written to _rd'_. | | _rd'_ and _rs1'_ are from the standard 8-register set x8-x15. | | ---------------------------------------------------------------- | Prerequisites None Operation ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. X(rdc) = EXTS(load_mem[X(rs1c)+EXTZ(uimm)][15..0]); ``` #### [](#insns-c%5Fsb)29.1.12.4\. c.sb Synopsis Store byte, 16-bit encoding Mnemonic c.sb _rs2'_, _uimm_(_rs1'_) Encoding (RV32, RV64): ![svg](_images/svg-35484c3bc40f2c9b43e4d742c5fb67421231517c.svg) The immediate offset is formed as follows: ```sail uimm[31:2] = 0; uimm[1] = encoding[5]; uimm[0] = encoding[6]; ``` Description This instruction stores the least significant byte of _rs2'_ to the memory address formed by adding _rs1'_ to the zero extended immediate _uimm_. | | _rs1'_ and _rs2'_ are from the standard 8-register set x8-x15. | | ----------------------------------------------------------------- | Prerequisites None Operation ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. mem[X(rs1c)+EXTZ(uimm)][7..0] = X(rs2c) ``` #### [](#insns-c%5Fsh)29.1.12.5\. c.sh Synopsis Store halfword, 16-bit encoding Mnemonic c.sh _rs2'_, _uimm_(_rs1'_) Encoding (RV32, RV64): ![svg](_images/svg-38fe08f2617944263157fae83dc1e8ee853420c7.svg) The immediate offset is formed as follows: ```sail uimm[31:2] = 0; uimm[1] = encoding[5]; uimm[0] = 0; ``` Description This instruction stores the least significant halfword of _rs2'_ to the memory address formed by adding _rs1'_ to the zero extended immediate _uimm_. | | _rs1'_ and _rs2'_ are from the standard 8-register set x8-x15. | | ----------------------------------------------------------------- | Prerequisites None Operation ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. mem[X(rs1c)+EXTZ(uimm)][15..0] = X(rs2c) ``` #### [](#insns-c%5Fzext%5Fb)29.1.12.6\. c.zext.b Synopsis Zero extend byte, 16-bit encoding Mnemonic c.zext.b _rd'/rs1'_ Encoding (RV32, RV64): ![svg](_images/svg-b57fb9f9a30874c8ef6d75c48a2bfcd91a0f2007.svg) Description This instruction takes a single source/destination operand. It zero-extends the least-significant byte of the operand to XLEN bits by inserting zeros into all of the bits more significant than 7. | | _rd'/rs1'_ is from the standard 8-register set x8-x15. | | --------------------------------------------------------- | Prerequisites None 32-bit equivalent: ```asm andi rd'/rs1', rd'/rs1', 0xff ``` | | The SAIL module variable for _rd'/rs1'_ is called _rsdc_. | | ------------------------------------------------------------ | Operation ```sail X(rsdc) = EXTZ(X(rsdc)[7..0]); ``` #### [](#insns-c%5Fsext%5Fb)29.1.12.7\. c.sext.b Synopsis Sign extend byte, 16-bit encoding Mnemonic c.sext.b _rd'/rs1'_ Encoding (RV32, RV64): ![svg](_images/svg-96ae5e4b8055be3253586a7f7a62bc755050d552.svg) Description This instruction takes a single source/destination operand. It sign-extends the least-significant byte in the operand to XLEN bits by copying the most-significant bit in the byte (i.e., bit 7) to all of the more-significant bits. | | _rd'/rs1'_ is from the standard 8-register set x8-x15. | | --------------------------------------------------------- | Prerequisites Zbb is also required. | | The SAIL module variable for _rd'/rs1'_ is called _rsdc_. | | ------------------------------------------------------------ | Operation ```sail X(rsdc) = EXTS(X(rsdc)[7..0]); ``` #### [](#insns-c%5Fzext%5Fh)29.1.12.8\. c.zext.h Synopsis Zero extend halfword, 16-bit encoding Mnemonic c.zext.h _rd'/rs1'_ Encoding (RV32, RV64): ![svg](_images/svg-0254a498086bad21900c5c3c260e1f8acd534fa7.svg) Description This instruction takes a single source/destination operand. It zero-extends the least-significant halfword of the operand to XLEN bits by inserting zeros into all of the bits more significant than 15. | | _rd'/rs1'_ is from the standard 8-register set x8-x15. | | --------------------------------------------------------- | Prerequisites Zbb is also required. | | The SAIL module variable for _rd'/rs1'_ is called _rsdc_. | | ------------------------------------------------------------ | Operation ```sail X(rsdc) = EXTZ(X(rsdc)[15..0]); ``` #### [](#insns-c%5Fsext%5Fh)29.1.12.9\. c.sext.h Synopsis Sign extend halfword, 16-bit encoding Mnemonic c.sext.h _rd'/rs1'_ Encoding (RV32, RV64): ![svg](_images/svg-8aad6996bcbe2ed79da5157ab38508348bb34780.svg) Description This instruction takes a single source/destination operand. It sign-extends the least-significant halfword in the operand to XLEN bits by copying the most-significant bit in the halfword (i.e., bit 15) to all of the more-significant bits. | | _rd'/rs1'_ is from the standard 8-register set x8-x15. | | --------------------------------------------------------- | Prerequisites Zbb is also required. | | The SAIL module variable for _rd'/rs1'_ is called _rsdc_. | | ------------------------------------------------------------ | Operation ```sail X(rsdc) = EXTS(X(rsdc)[15..0]); ``` #### [](#insns-c%5Fzext%5Fw)29.1.12.10\. c.zext.w Synopsis Zero extend word, 16-bit encoding Mnemonic c.zext.w _rd'/rs1'_ Encoding (RV64): ![svg](_images/svg-b31a255d898b564936ece4f75f0f59e0d87a5a93.svg) Description This instruction takes a single source/destination operand. It zero-extends the least-significant word of the operand to XLEN bits by inserting zeros into all of the bits more significant than 31. | | _rd'/rs1'_ is from the standard 8-register set x8-x15. | | --------------------------------------------------------- | Prerequisites Zba is also required. 32-bit equivalent: ```asm add.uw rd'/rs1', rd'/rs1', zero ``` | | The SAIL module variable for _rd'/rs1'_ is called _rsdc_. | | ------------------------------------------------------------ | Operation ```sail X(rsdc) = EXTZ(X(rsdc)[31..0]); ``` #### [](#insns-c%5Fnot)29.1.12.11\. c.not Synopsis Bitwise not, 16-bit encoding Mnemonic c.not _rd'/rs1'_ Encoding (RV32, RV64): ![svg](_images/svg-1a774fc4390f2953c585722873d64db545ce71f4.svg) Description This instruction takes the one’s complement of _rd'/rs1'_ and writes the result to the same register. | | rd'/rs1' is from the standard 8-register set x8-x15. | | ------------------------------------------------------- | Prerequisites None 32-bit equivalent: ```asm xori rd'/rs1', rd'/rs1', -1 ``` | | The SAIL module variable for _rd'/rs1'_ is called _rsdc_. | | ------------------------------------------------------------ | Operation ```sail X(rsdc) = X(rsdc) XOR -1; ``` #### [](#insns-c%5Fmul)29.1.12.12\. c.mul Synopsis Multiply, 16-bit encoding Mnemonic c.mul _rsd'_, _rs2'_ Encoding (RV32, RV64): ![svg](_images/svg-32f838dcac9e0e4e1a9c58f4f1188cd9eed05c19.svg) Description This instruction multiplies XLEN bits of the source operands from _rsd'_ and _rs2'_ and writes the lowest XLEN bits of the result to _rsd'_. | | _rd'/rs1'_ and _rs2'_ are from the standard 8-register set x8-x15. | | --------------------------------------------------------------------- | Prerequisites M or Zmmul must be configured. | | The SAIL module variable for _rd'/rs1'_ is called _rsdc_, and for _rs2'_ is called _rs2c_. | | --------------------------------------------------------------------------------------------- | Operation ```sail let result_wide = to_bits(2 * sizeof(xlen), signed(X(rsdc)) * signed(X(rs2c))); X(rsdc) = result_wide[(sizeof(xlen) - 1) .. 0]; ``` ### [](#insns-pushpop)29.1.13\. PUSH/POP register instructions These instructions are collectively referred to as PUSH/POP: * [cm.push](#insns-cm%5Fpush) * [cm.pop](#insns-cm%5Fpop) * [cm.popret](#insns-cm%5Fpopret) * [cm.popretz](#insns-cm%5Fpopretz) The term PUSH refers to _cm.push_. The term POP refers to _cm.pop_. The term POPRET refers to _cm.popret and cm.popretz_. Common details for these instructions are in this section. #### [](#29-1-13-1-pushpop-functional-overview)29.1.13.1\. PUSH/POP functional overview PUSH, POP, POPRET are used to reduce the size of function prologues and epilogues. 1. The PUSH instruction * adjusts the stack pointer to create the stack frame * pushes (stores) the registers specified in the register list to the stack frame 2. The POP instruction * pops (loads) the registers in the register list from the stack frame * adjusts the stack pointer to destroy the stack frame 3. The POPRET instructions * pop (load) the registers in the register list from the stack frame * _cm.popretz_ also moves zero into _a0_ as the return value * adjust the stack pointer to destroy the stack frame * execute a _ret_ instruction to return from the function #### [](#29-1-13-2-example-usage)29.1.13.2\. Example usage This example gives an illustration of the use of PUSH and POPRET. The function _processMarkers_ in the EMBench benchmark picojpeg in the following file on github: [libpicojpeg.c](https://github.com/embench/embench-iot/blob/master/src/picojpeg/libpicojpeg.c) The prologue and epilogue compile with GCC10 to: ```asm 0001098a : 1098a: 711d addi sp,sp,-96 ;#cm.push(1) 1098c: c8ca sw s2,80(sp) ;#cm.push(2) 1098e: c6ce sw s3,76(sp) ;#cm.push(3) 10990: c4d2 sw s4,72(sp) ;#cm.push(4) 10992: ce86 sw ra,92(sp) ;#cm.push(5) 10994: cca2 sw s0,88(sp) ;#cm.push(6) 10996: caa6 sw s1,84(sp) ;#cm.push(7) 10998: c2d6 sw s5,68(sp) ;#cm.push(8) 1099a: c0da sw s6,64(sp) ;#cm.push(9) 1099c: de5e sw s7,60(sp) ;#cm.push(10) 1099e: dc62 sw s8,56(sp) ;#cm.push(11) 109a0: da66 sw s9,52(sp) ;#cm.push(12) 109a2: d86a sw s10,48(sp);#cm.push(13) 109a4: d66e sw s11,44(sp);#cm.push(14) ... 109f4: 4501 li a0,0 ;#cm.popretz(1) 109f6: 40f6 lw ra,92(sp) ;#cm.popretz(2) 109f8: 4466 lw s0,88(sp) ;#cm.popretz(3) 109fa: 44d6 lw s1,84(sp) ;#cm.popretz(4) 109fc: 4946 lw s2,80(sp) ;#cm.popretz(5) 109fe: 49b6 lw s3,76(sp) ;#cm.popretz(6) 10a00: 4a26 lw s4,72(sp) ;#cm.popretz(7) 10a02: 4a96 lw s5,68(sp) ;#cm.popretz(8) 10a04: 4b06 lw s6,64(sp) ;#cm.popretz(9) 10a06: 5bf2 lw s7,60(sp) ;#cm.popretz(10) 10a08: 5c62 lw s8,56(sp) ;#cm.popretz(11) 10a0a: 5cd2 lw s9,52(sp) ;#cm.popretz(12) 10a0c: 5d42 lw s10,48(sp);#cm.popretz(13) 10a0e: 5db2 lw s11,44(sp);#cm.popretz(14) 10a10: 6125 addi sp,sp,96 ;#cm.popretz(15) 10a12: 8082 ret ;#cm.popretz(16) ``` with the GCC option _\-msave-restore_ the output is the following: ```asm 0001080e : 1080e: 73a012ef jal t0,11f48 <__riscv_save_12> 10812: 1101 addi sp,sp,-32 ... 10862: 4501 li a0,0 10864: 6105 addi sp,sp,32 10866: 71e0106f j 11f84 <__riscv_restore_12> ``` with PUSH/POPRET this reduces to ```asm 0001080e : 1080e: b8fa cm.push \{ra,s0-s11},-96 ... 10866: bcfa cm.popretz \{ra,s0-s11}, 96 ``` The prologue / epilogue reduce from 60-bytes in the original code, to 14-bytes with _\-msave-restore_, and to 4-bytes with PUSH and POPRET. As well as reducing the code-size PUSH and POPRET eliminate the branches from calling the millicode _save/restore_ routines and so may also perform better. | | The calls to _/_ become 64-bit when the target functions are out of the ±1 MB range, increasing the prologue/epilogue size to 22-bytes. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | POP is typically used in tail-calling sequences where _ret_ is not used to return to _ra_ after destroying the stack frame. | | ------------------------------------------------------------------------------------------------------------------------------ | ##### [](#pushpop-areg-list)29.1.13.2.1\. Stack pointer adjustment handling The instructions all automatically adjust the stack pointer by enough to cover the memory required for the registers being saved or restored. Additionally the _spimm_ field in the encoding allows the stack pointer to be adjusted in additional increments of 16-bytes. There is only a small restricted range available in the encoding; if the range is insufficient then a separate _c.addi16sp_ can be used to increase the range. ##### [](#29-1-13-2-2-register-list-handling)29.1.13.2.2\. Register list handling There is no support for the _\\{ra, s0-s10}_ register list without also adding _s11_. Therefore the _\\{ra, s0-s11}_ register list must be used in this case. #### [](#pushpop-idempotent-memory)29.1.13.3\. PUSH/POP Fault handling Correct execution requires that _sp_ refers to idempotent memory (also see [Non-idempotent memory handling](#pushpop%5Fnon-idem-mem)), because the core must be able to handle traps detected during the sequence. The entire PUSH/POP sequence is re-executed after returning from the trap handler, and multiple traps are possible during the sequence. If a trap occurs during the sequence then _xEPC_ is updated with the PC of the instruction, _xTVAL_ (if not read-only-zero) updated with the bad address if it was an access fault and _xCAUSE_ updated with the type of trap. | | It is implementation defined whether interrupts can also be taken during the sequence execution. | | --------------------------------------------------------------------------------------------------- | #### [](#pushpop-software-view)29.1.13.4\. Software view of execution ##### [](#29-1-13-4-1-software-view-of-the-push-sequence)29.1.13.4.1\. Software view of the PUSH sequence From a software perspective the PUSH sequence appears as: * A sequence of stores writing the bytes required by the pseudocode * The bytes may be written in any order. * The bytes may be grouped into larger accesses. * Any of the bytes may be written multiple times. * A stack pointer adjustment | | If an implementation allows interrupts during the sequence, and the interrupt handler uses _sp_ to allocate stack memory, then any stores which were executed before the interrupt may be overwritten by the handler. This is safe because the memory is idempotent and the stores will be re-executed when execution resumes. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The stack pointer adjustment must only be committed only when it is certain that the entire PUSH instruction will commit. Stores may also return imprecise faults from the bus. It is platform defined whether the core implementation waits for the bus responses before continuing to the final stage of the sequence, or handles errors responses after completing the PUSH instruction. For example: ```asm cm.push \{ra, s0-s5}, -64 ``` Appears to software as: ```asm # any bytes from sp-1 to sp-28 may be written multiple times before # the instruction completes therefore these updates may be visible in # the interrupt/exception handler below the stack pointer sw s5, -4(sp) sw s4, -8(sp) sw s3,-12(sp) sw s2,-16(sp) sw s1,-20(sp) sw s0,-24(sp) sw ra,-28(sp) # this must only execute once, and will only execute after all stores # completed without any precise faults, therefore this update is only # visible in the interrupt/exception handler if cm.push has completed addi sp, sp, -64 ``` ##### [](#29-1-13-4-2-software-view-of-the-poppopret-sequence)29.1.13.4.2\. Software view of the POP/POPRET sequence From a software perspective the POP/POPRET sequence appears as: * A sequence of loads reading the bytes required by the pseudocode. * The bytes may be loaded in any order. * The bytes may be grouped into larger accesses. * Any of the bytes may be loaded multiple times. * A stack pointer adjustment * An optional `li a0, 0` * An optional `ret` If a trap occurs during the sequence, then any loads which were executed before the trap may update architectural state. The loads will be re-executed once the trap handler completes, so the values will be overwritten. Therefore it is permitted for an implementation to update some of the destination registers before taking a fault. The optional `li a0, 0`, stack pointer adjustment and optional `ret` must only be committed only when it is certain that the entire POP/POPRET instruction will commit. For POPRET once the stack pointer adjustment has been committed the `ret` must execute. For example: ```asm cm.popretz \{ra, s0-s3}, 32; ``` Appears to software as: ```asm # any or all of these load instructions may execute multiple times # therefore these updates may be visible in the interrupt/exception handler lw s3, 28(sp) lw s2, 24(sp) lw s1, 20(sp) lw s0, 16(sp) lw ra, 12(sp) # these must only execute once, will only execute after all loads # complete successfully all instructions must execute atomically # therefore these updates are not visible in the interrupt/exception handler li a0, 0 addi sp, sp, 32 ret ``` #### [](#pushpop%5Fnon-idem-mem)29.1.13.5\. Non-idempotent memory handling An implementation may have a requirement to issue a PUSH/POP instruction to non-idempotent memory. If the core implementation does not support PUSH/POP to non-idempotent memories, the core may use an idempotency PMA to detect it and take a load (POP/POPRET) or store (PUSH) access-fault exception in order to avoid unpredictable results. Software should only use these instructions on non-idempotent memory regions when software can tolerate the required memory accesses being issued repeatedly in the case that they cause exceptions. #### [](#29-1-13-6-example-rv32i-pushpop-sequences)29.1.13.6\. Example RV32I PUSH/POP sequences The examples are included show the load/store series expansion and the stack adjustment. Examples of _cm.popret_ and _cm.popretz_ are not included, as the difference in the expanded sequence from _cm.pop_ is trivial in all cases. ##### [](#29-1-13-6-1-cm-push-ra-s0-s2-64)29.1.13.6.1\. cm.push \\{ra, s0-s2}, -64 Encoding: _rlist_\=7, _spimm_\=3 expands to: ```asm sw s2, -4(sp); sw s1, -8(sp); sw s0, -12(sp); sw ra, -16(sp); addi sp, sp, -64; ``` ##### [](#29-1-13-6-2-cm-push-ra-s0-s11-112)29.1.13.6.2\. cm.push \\{ra, s0-s11}, -112 Encoding: _rlist_\=15, _spimm_\=3 expands to: ```asm sw s11, -4(sp); sw s10, -8(sp); sw s9, -12(sp); sw s8, -16(sp); sw s7, -20(sp); sw s6, -24(sp); sw s5, -28(sp); sw s4, -32(sp); sw s3, -36(sp); sw s2, -40(sp); sw s1, -44(sp); sw s0, -48(sp); sw ra, -52(sp); addi sp, sp, -112; ``` ##### [](#29-1-13-6-3-cm-pop-ra-16)29.1.13.6.3\. cm.pop {ra}, 16 Encoding: _rlist_\=4, _spimm_\=0 expands to: ```asm lw ra, 12(sp); addi sp, sp, 16; ``` ##### [](#29-1-13-6-4-cm-pop-ra-s0-s3-48)29.1.13.6.4\. cm.pop \\{ra, s0-s3}, 48 Encoding: _rlist_\=8, _spimm_\=1 expands to: ```asm lw s3, 44(sp); lw s2, 40(sp); lw s1, 36(sp); lw s0, 32(sp); lw ra, 28(sp); addi sp, sp, 48; ``` ##### [](#29-1-13-6-5-cm-pop-ra-s0-s4-64)29.1.13.6.5\. cm.pop \\{ra, s0-s4}, 64 Encoding: _rlist_\=9, _spimm_\=2 expands to: ```asm lw s4, 60(sp); lw s3, 56(sp); lw s2, 52(sp); lw s1, 48(sp); lw s0, 44(sp); lw ra, 40(sp); addi sp, sp, 64; ``` #### [](#insns-cm%5Fpush)29.1.13.7\. cm.push Synopsis Create stack frame: store ra and 0 to 12 saved registers to the stack frame, optionally allocate additional stack space. Mnemonic cm.push _{reg\_list}, -stack\_adj_ Encoding (RV32, RV64): ![svg](_images/svg-533cc80019c6a1cb3c65758f651a2fe237302279.svg) | | _rlist_ values 0 to 3 are reserved for a future EABI variant called _cm.push.e_ | | ---------------------------------------------------------------------------------- | Assembly Syntax: ```asm cm.push \{reg_list}, -stack_adj cm.push {xreg_list}, -stack_adj ``` The variables used in the assembly syntax are defined below. ```sail RV32E: switch (rlist){ case 4: \{reg_list="ra"; xreg_list="x1";} case 5: \{reg_list="ra, s0"; xreg_list="x1, x8";} case 6: \{reg_list="ra, s0-s1"; xreg_list="x1, x8-x9";} default: reserved(); } stack_adj = stack_adj_base + spimm * 16; ``` ```sail RV32I, RV64: switch (rlist){ case 4: \{reg_list="ra"; xreg_list="x1";} case 5: \{reg_list="ra, s0"; xreg_list="x1, x8";} case 6: \{reg_list="ra, s0-s1"; xreg_list="x1, x8-x9";} case 7: \{reg_list="ra, s0-s2"; xreg_list="x1, x8-x9, x18";} case 8: \{reg_list="ra, s0-s3"; xreg_list="x1, x8-x9, x18-x19";} case 9: \{reg_list="ra, s0-s4"; xreg_list="x1, x8-x9, x18-x20";} case 10: \{reg_list="ra, s0-s5"; xreg_list="x1, x8-x9, x18-x21";} case 11: \{reg_list="ra, s0-s6"; xreg_list="x1, x8-x9, x18-x22";} case 12: \{reg_list="ra, s0-s7"; xreg_list="x1, x8-x9, x18-x23";} case 13: \{reg_list="ra, s0-s8"; xreg_list="x1, x8-x9, x18-x24";} case 14: \{reg_list="ra, s0-s9"; xreg_list="x1, x8-x9, x18-x25";} //note - to include s10, s11 must also be included case 15: \{reg_list="ra, s0-s11"; xreg_list="x1, x8-x9, x18-x27";} default: reserved(); } stack_adj = stack_adj_base + spimm * 16; ``` ```sail RV32E: stack_adj_base = 16; Valid values: stack_adj = [16|32|48|64]; ``` ```sail RV32I: switch (rlist) { case 4.. 7: stack_adj_base = 16; case 8..11: stack_adj_base = 32; case 12..14: stack_adj_base = 48; case 15: stack_adj_base = 64; } Valid values: switch (rlist) { case 4.. 7: stack_adj = [16|32|48| 64]; case 8..11: stack_adj = [32|48|64| 80]; case 12..14: stack_adj = [48|64|80| 96]; case 15: stack_adj = [64|80|96|112]; } ``` ```sail RV64: switch (rlist) { case 4.. 5: stack_adj_base = 16; case 6.. 7: stack_adj_base = 32; case 8.. 9: stack_adj_base = 48; case 10..11: stack_adj_base = 64; case 12..13: stack_adj_base = 80; case 14: stack_adj_base = 96; case 15: stack_adj_base = 112; } Valid values: switch (rlist) { case 4.. 5: stack_adj = [ 16| 32| 48| 64]; case 6.. 7: stack_adj = [ 32| 48| 64| 80]; case 8.. 9: stack_adj = [ 48| 64| 80| 96]; case 10..11: stack_adj = [ 64| 80| 96|112]; case 12..13: stack_adj = [ 80| 96|112|128]; case 14: stack_adj = [ 96|112|128|144]; case 15: stack_adj = [112|128|144|160]; } ``` Description This instruction pushes (stores) the registers in _reg\_list_ to the memory below the stack pointer, and then creates the stack frame by decrementing the stack pointer by _stack\_adj_, including any additional stack space requested by the value of _spimm_. | | All ABI register mappings are for the UABI. An EABI version is planned once the EABI is frozen. | | -------------------------------------------------------------------------------------------------- | For further information see [PUSH/POP Register Instructions](#insns-pushpop). Stack Adjustment Calculation: _stack\_adj\_base_ is the minimum number of bytes, in multiples of 16-byte address increments, required to cover the registers in the list. _spimm_ is the number of additional 16-byte address increments allocated for the stack frame. The total stack adjustment represents the total size of the stack frame, which is _stack\_adj\_base_ added to _spimm_ scaled by 16, as defined above. Prerequisites None 32-bit equivalent: No direct equivalent encoding exists Operation The first section of pseudocode may be executed multiple times before the instruction successfully completes. ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. if (XLEN==32) bytes=4; else bytes=8; addr=sp-bytes; for(i in 27,26,25,24,23,22,21,20,19,18,9,8,1) { //if register i is in xreg_list if (xreg_list[i]) { switch(bytes) { 4: asm("sw x[i], 0(addr)"); 8: asm("sd x[i], 0(addr)"); } addr-=bytes; } } ``` The final section of pseudocode executes atomically, and only executes if the section above completes without any exceptions or interrupts. ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. sp-=stack_adj; ``` #### [](#insns-cm%5Fpop)29.1.13.8\. cm.pop Synopsis Destroy stack frame: load ra and 0 to 12 saved registers from the stack frame, deallocate the stack frame. Mnemonic cm.pop _{reg\_list}, stack\_adj_ Encoding (RV32, RV64): ![svg](_images/svg-e341cebc191f49bdedf0edde18b625cef6add4ff.svg) | | _rlist_ values 0 to 3 are reserved for a future EABI variant called _cm.pop.e_ | | --------------------------------------------------------------------------------- | Assembly Syntax: ```asm cm.pop \{reg_list}, stack_adj cm.pop {xreg_list}, stack_adj ``` The variables used in the assembly syntax are defined below. ```sail RV32E: switch (rlist){ case 4: \{reg_list="ra"; xreg_list="x1";} case 5: \{reg_list="ra, s0"; xreg_list="x1, x8";} case 6: \{reg_list="ra, s0-s1"; xreg_list="x1, x8-x9";} default: reserved(); } stack_adj = stack_adj_base + spimm * 16; ``` ```sail RV32I, RV64: switch (rlist){ case 4: \{reg_list="ra"; xreg_list="x1";} case 5: \{reg_list="ra, s0"; xreg_list="x1, x8";} case 6: \{reg_list="ra, s0-s1"; xreg_list="x1, x8-x9";} case 7: \{reg_list="ra, s0-s2"; xreg_list="x1, x8-x9, x18";} case 8: \{reg_list="ra, s0-s3"; xreg_list="x1, x8-x9, x18-x19";} case 9: \{reg_list="ra, s0-s4"; xreg_list="x1, x8-x9, x18-x20";} case 10: \{reg_list="ra, s0-s5"; xreg_list="x1, x8-x9, x18-x21";} case 11: \{reg_list="ra, s0-s6"; xreg_list="x1, x8-x9, x18-x22";} case 12: \{reg_list="ra, s0-s7"; xreg_list="x1, x8-x9, x18-x23";} case 13: \{reg_list="ra, s0-s8"; xreg_list="x1, x8-x9, x18-x24";} case 14: \{reg_list="ra, s0-s9"; xreg_list="x1, x8-x9, x18-x25";} //note - to include s10, s11 must also be included case 15: \{reg_list="ra, s0-s11"; xreg_list="x1, x8-x9, x18-x27";} default: reserved(); } stack_adj = stack_adj_base + spimm * 16; ``` ```sail RV32E: stack_adj_base = 16; Valid values: stack_adj = [16|32|48|64]; ``` ```sail RV32I: switch (rlist) { case 4.. 7: stack_adj_base = 16; case 8..11: stack_adj_base = 32; case 12..14: stack_adj_base = 48; case 15: stack_adj_base = 64; } Valid values: switch (rlist) { case 4.. 7: stack_adj = [16|32|48| 64]; case 8..11: stack_adj = [32|48|64| 80]; case 12..14: stack_adj = [48|64|80| 96]; case 15: stack_adj = [64|80|96|112]; } ``` ```sail RV64: switch (rlist) { case 4.. 5: stack_adj_base = 16; case 6.. 7: stack_adj_base = 32; case 8.. 9: stack_adj_base = 48; case 10..11: stack_adj_base = 64; case 12..13: stack_adj_base = 80; case 14: stack_adj_base = 96; case 15: stack_adj_base = 112; } Valid values: switch (rlist) { case 4.. 5: stack_adj = [ 16| 32| 48| 64]; case 6.. 7: stack_adj = [ 32| 48| 64| 80]; case 8.. 9: stack_adj = [ 48| 64| 80| 96]; case 10..11: stack_adj = [ 64| 80| 96|112]; case 12..13: stack_adj = [ 80| 96|112|128]; case 14: stack_adj = [ 96|112|128|144]; case 15: stack_adj = [112|128|144|160]; } ``` Description This instruction pops (loads) the registers in _reg\_list_ from stack memory, and then adjusts the stack pointer by _stack\_adj_. | | All ABI register mappings are for the UABI. An EABI version is planned once the EABI is frozen. | | -------------------------------------------------------------------------------------------------- | For further information see [PUSH/POP Register Instructions](#insns-pushpop). Stack Adjustment Calculation: _stack\_adj\_base_ is the minimum number of bytes, in multiples of 16-byte address increments, required to cover the registers in the list. _spimm_ is the number of additional 16-byte address increments allocated for the stack frame. The total stack adjustment represents the total size of the stack frame, which is _stack\_adj\_base_ added to _spimm_ scaled by 16, as defined above. Prerequisites None 32-bit equivalent: No direct equivalent encoding exists Operation The first section of pseudocode may be executed multiple times before the instruction successfully completes. ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. if (XLEN==32) bytes=4; else bytes=8; addr=sp+stack_adj-bytes; for(i in 27,26,25,24,23,22,21,20,19,18,9,8,1) { //if register i is in xreg_list if (xreg_list[i]) { switch(bytes) { 4: asm("lw x[i], 0(addr)"); 8: asm("ld x[i], 0(addr)"); } addr-=bytes; } } ``` The final section of pseudocode executes atomically, and only executes if the section above completes without any exceptions or interrupts. ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. sp+=stack_adj; ``` #### [](#insns-cm%5Fpopretz)29.1.13.9\. cm.popretz Synopsis Destroy stack frame: load ra and 0 to 12 saved registers from the stack frame, deallocate the stack frame, move zero into a0, return to ra. Mnemonic cm.popretz _{reg\_list}, stack\_adj_ Encoding (RV32, RV64): ![svg](_images/svg-b3689d303951881df175b71d8b7ec23cbe4dbb3f.svg) | | _rlist_ values 0 to 3 are reserved for a future EABI variant called _cm.popretz.e_ | | ------------------------------------------------------------------------------------- | Assembly Syntax: ```sail cm.popretz \{reg_list}, stack_adj cm.popretz {xreg_list}, stack_adj ``` ```sail RV32E: switch (rlist){ case 4: \{reg_list="ra"; xreg_list="x1";} case 5: \{reg_list="ra, s0"; xreg_list="x1, x8";} case 6: \{reg_list="ra, s0-s1"; xreg_list="x1, x8-x9";} default: reserved(); } stack_adj = stack_adj_base + spimm * 16; ``` ```sail RV32I, RV64: switch (rlist){ case 4: \{reg_list="ra"; xreg_list="x1";} case 5: \{reg_list="ra, s0"; xreg_list="x1, x8";} case 6: \{reg_list="ra, s0-s1"; xreg_list="x1, x8-x9";} case 7: \{reg_list="ra, s0-s2"; xreg_list="x1, x8-x9, x18";} case 8: \{reg_list="ra, s0-s3"; xreg_list="x1, x8-x9, x18-x19";} case 9: \{reg_list="ra, s0-s4"; xreg_list="x1, x8-x9, x18-x20";} case 10: \{reg_list="ra, s0-s5"; xreg_list="x1, x8-x9, x18-x21";} case 11: \{reg_list="ra, s0-s6"; xreg_list="x1, x8-x9, x18-x22";} case 12: \{reg_list="ra, s0-s7"; xreg_list="x1, x8-x9, x18-x23";} case 13: \{reg_list="ra, s0-s8"; xreg_list="x1, x8-x9, x18-x24";} case 14: \{reg_list="ra, s0-s9"; xreg_list="x1, x8-x9, x18-x25";} //note - to include s10, s11 must also be included case 15: \{reg_list="ra, s0-s11"; xreg_list="x1, x8-x9, x18-x27";} default: reserved(); } stack_adj = stack_adj_base + spimm * 16; ``` ```sail RV32E: stack_adj_base = 16; Valid values: stack_adj = [16|32|48|64]; ``` ```sail RV32I: switch (rlist) { case 4.. 7: stack_adj_base = 16; case 8..11: stack_adj_base = 32; case 12..14: stack_adj_base = 48; case 15: stack_adj_base = 64; } Valid values: switch (rlist) { case 4.. 7: stack_adj = [16|32|48| 64]; case 8..11: stack_adj = [32|48|64| 80]; case 12..14: stack_adj = [48|64|80| 96]; case 15: stack_adj = [64|80|96|112]; } ``` ```sail RV64: switch (rlist) { case 4.. 5: stack_adj_base = 16; case 6.. 7: stack_adj_base = 32; case 8.. 9: stack_adj_base = 48; case 10..11: stack_adj_base = 64; case 12..13: stack_adj_base = 80; case 14: stack_adj_base = 96; case 15: stack_adj_base = 112; } Valid values: switch (rlist) { case 4.. 5: stack_adj = [ 16| 32| 48| 64]; case 6.. 7: stack_adj = [ 32| 48| 64| 80]; case 8.. 9: stack_adj = [ 48| 64| 80| 96]; case 10..11: stack_adj = [ 64| 80| 96|112]; case 12..13: stack_adj = [ 80| 96|112|128]; case 14: stack_adj = [ 96|112|128|144]; case 15: stack_adj = [112|128|144|160]; } ``` Description This instruction pops (loads) the registers in _reg\_list_ from stack memory, adjusts the stack pointer by _stack\_adj_, moves zero into a0 and then returns to _ra_. | | All ABI register mappings are for the UABI. An EABI version is planned once the EABI is frozen. | | -------------------------------------------------------------------------------------------------- | For further information see [PUSH/POP Register Instructions](#insns-pushpop). Stack Adjustment Calculation: _stack\_adj\_base_ is the minimum number of bytes, in multiples of 16-byte address increments, required to cover the registers in the list. _spimm_ is the number of additional 16-byte address increments allocated for the stack frame. The total stack adjustment represents the total size of the stack frame, which is _stack\_adj\_base_ added to _spimm_ scaled by 16, as defined above. Prerequisites None 32-bit equivalent: No direct equivalent encoding exists Operation The first section of pseudocode may be executed multiple times before the instruction successfully completes. ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. if (XLEN==32) bytes=4; else bytes=8; addr=sp+stack_adj-bytes; for(i in 27,26,25,24,23,22,21,20,19,18,9,8,1) { //if register i is in xreg_list if (xreg_list[i]) { switch(bytes) { 4: asm("lw x[i], 0(addr)"); 8: asm("ld x[i], 0(addr)"); } addr-=bytes; } } ``` The final section of pseudocode executes atomically, and only executes if the section above completes without any exceptions or interrupts. | | The _li a0, 0_ **could** be executed more than once, but is included in the atomic section for convenience. | | -------------------------------------------------------------------------------------------------------------- | ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. asm("li a0, 0"); sp+=stack_adj; asm("ret"); ``` #### [](#insns-cm%5Fpopret)29.1.13.10\. cm.popret Synopsis Destroy stack frame: load ra and 0 to 12 saved registers from the stack frame, deallocate the stack frame, return to ra. Mnemonic cm.popret _{reg\_list}, stack\_adj_ Encoding (RV32, RV64): ![svg](_images/svg-6486253fa6b237777186c1b1f2db258728379c24.svg) | | _rlist_ values 0 to 3 are reserved for a future EABI variant called _cm.popret.e_ | | ------------------------------------------------------------------------------------ | Assembly Syntax: ```sail cm.popret \{reg_list}, stack_adj cm.popret {xreg_list}, stack_adj ``` The variables used in the assembly syntax are defined below. ```sail RV32E: switch (rlist){ case 4: \{reg_list="ra"; xreg_list="x1";} case 5: \{reg_list="ra, s0"; xreg_list="x1, x8";} case 6: \{reg_list="ra, s0-s1"; xreg_list="x1, x8-x9";} default: reserved(); } stack_adj = stack_adj_base + spimm * 16; ``` ```sail RV32I, RV64: switch (rlist){ case 4: \{reg_list="ra"; xreg_list="x1";} case 5: \{reg_list="ra, s0"; xreg_list="x1, x8";} case 6: \{reg_list="ra, s0-s1"; xreg_list="x1, x8-x9";} case 7: \{reg_list="ra, s0-s2"; xreg_list="x1, x8-x9, x18";} case 8: \{reg_list="ra, s0-s3"; xreg_list="x1, x8-x9, x18-x19";} case 9: \{reg_list="ra, s0-s4"; xreg_list="x1, x8-x9, x18-x20";} case 10: \{reg_list="ra, s0-s5"; xreg_list="x1, x8-x9, x18-x21";} case 11: \{reg_list="ra, s0-s6"; xreg_list="x1, x8-x9, x18-x22";} case 12: \{reg_list="ra, s0-s7"; xreg_list="x1, x8-x9, x18-x23";} case 13: \{reg_list="ra, s0-s8"; xreg_list="x1, x8-x9, x18-x24";} case 14: \{reg_list="ra, s0-s9"; xreg_list="x1, x8-x9, x18-x25";} //note - to include s10, s11 must also be included case 15: \{reg_list="ra, s0-s11"; xreg_list="x1, x8-x9, x18-x27";} default: reserved(); } stack_adj = stack_adj_base + spimm * 16; ``` ```sail RV32E: stack_adj_base = 16; Valid values: stack_adj = [16|32|48|64]; ``` ```sail RV32I: switch (rlist) { case 4.. 7: stack_adj_base = 16; case 8..11: stack_adj_base = 32; case 12..14: stack_adj_base = 48; case 15: stack_adj_base = 64; } Valid values: switch (rlist) { case 4.. 7: stack_adj = [16|32|48| 64]; case 8..11: stack_adj = [32|48|64| 80]; case 12..14: stack_adj = [48|64|80| 96]; case 15: stack_adj = [64|80|96|112]; } ``` ```sail RV64: switch (rlist) { case 4.. 5: stack_adj_base = 16; case 6.. 7: stack_adj_base = 32; case 8.. 9: stack_adj_base = 48; case 10..11: stack_adj_base = 64; case 12..13: stack_adj_base = 80; case 14: stack_adj_base = 96; case 15: stack_adj_base = 112; } Valid values: switch (rlist) { case 4.. 5: stack_adj = [ 16| 32| 48| 64]; case 6.. 7: stack_adj = [ 32| 48| 64| 80]; case 8.. 9: stack_adj = [ 48| 64| 80| 96]; case 10..11: stack_adj = [ 64| 80| 96|112]; case 12..13: stack_adj = [ 80| 96|112|128]; case 14: stack_adj = [ 96|112|128|144]; case 15: stack_adj = [112|128|144|160]; } ``` Description This instruction pops (loads) the registers in _reg\_list_ from stack memory, adjusts the stack pointer by _stack\_adj_ and then returns to _ra_. | | All ABI register mappings are for the UABI. An EABI version is planned once the EABI is frozen. | | -------------------------------------------------------------------------------------------------- | For further information see [PUSH/POP Register Instructions](#insns-pushpop). Stack Adjustment Calculation: _stack\_adj\_base_ is the minimum number of bytes, in multiples of 16-byte address increments, required to cover the registers in the list. _spimm_ is the number of additional 16-byte address increments allocated for the stack frame. The total stack adjustment represents the total size of the stack frame, which is _stack\_adj\_base_ added to _spimm_ scaled by 16, as defined above. Prerequisites None 32-bit equivalent: No direct equivalent encoding exists Operation The first section of pseudocode may be executed multiple times before the instruction successfully completes. ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. if (XLEN==32) bytes=4; else bytes=8; addr=sp+stack_adj-bytes; for(i in 27,26,25,24,23,22,21,20,19,18,9,8,1) { //if register i is in xreg_list if (xreg_list[i]) { switch(bytes) { 4: asm("lw x[i], 0(addr)"); 8: asm("ld x[i], 0(addr)"); } addr-=bytes; } } ``` The final section of pseudocode executes atomically, and only executes if the section above completes without any exceptions or interrupts. ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. sp+=stack_adj; asm("ret"); ``` #### [](#insns-cm%5Fmvsa01)29.1.13.11\. cm.mvsa01 Synopsis Move a0-a1 into two registers of s0-s7 Mnemonic cm.mvsa01 _r1s'_, _r2s'_ Encoding (RV32, RV64): ![svg](_images/svg-574e9410060980afe1a40696abb82242bf391af8.svg) | | For the encoding to be legal _r1s'_ != _r2s'_. | | ------------------------------------------------- | Assembly Syntax: ```asm cm.mvsa01 r1s', r2s' ``` Description This instruction moves _a0_ into _r1s'_ and _a1_ into _r2s'_. _r1s'_ and _r2s'_ must be different.The execution is atomic, so it is not possible to observe state where only one of _r1s'_ or _r2s'_ has been updated. The encoding uses _sreg_ number specifiers instead of _xreg_ number specifiers to save encoding space.The mapping between them is specified in the pseudocode below. | | The _s_ register mapping is taken from the UABI, and may not match the currently unratified EABI. _cm.mvsa01.e_ may be included in the future. | | ------------------------------------------------------------------------------------------------------------------------------------------------- | Prerequisites None 32-bit equivalent: No direct equivalent encoding exists. Operation ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. if (RV32E && (r1sc>1 || r2sc>1)) { reserved(); } xreg1 = {r1sc[2:1]>0,r1sc[2:1]==0,r1sc[2:0]}; xreg2 = {r2sc[2:1]>0,r2sc[2:1]==0,r2sc[2:0]}; X[xreg1] = X[10]; X[xreg2] = X[11]; ``` #### [](#insns-cm%5Fmva01s)29.1.13.12\. cm.mva01s Synopsis Move two s0-s7 registers into a0-a1 Mnemonic cm.mva01s _r1s'_, _r2s'_ Encoding (RV32, RV64): ![svg](_images/svg-9d9662a432d9c46953c886ebe0772a518e778737.svg) Assembly Syntax: ```asm cm.mva01s r1s', r2s' ``` Description This instruction moves _r1s'_ into _a0_ and _r2s'_ into _a1_.The execution is atomic, so it is not possible to observe state where only one of _a0_ or _a1_ have been updated. The encoding uses _sreg_ number specifiers instead of _xreg_ number specifiers to save encoding space.The mapping between them is specified in the pseudocode below. | | The _s_ register mapping is taken from the UABI, and may not match the currently unratified EABI. _cm.mva01s.e_ may be included in the future. | | ------------------------------------------------------------------------------------------------------------------------------------------------- | Prerequisites None 32-bit equivalent: No direct equivalent encoding exists. Operation ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. if (RV32E && (r1sc>1 || r2sc>1)) { reserved(); } xreg1 = {r1sc[2:1]>0,r1sc[2:1]==0,r1sc[2:0]}; xreg2 = {r2sc[2:1]>0,r2sc[2:1]==0,r2sc[2:0]}; X[10] = X[xreg1]; X[11] = X[xreg2]; ``` ### [](#insns-tablejump)29.1.14\. Table Jump Overview _cm.jt_ ([Jump via table](#insns-cm%5Fjt)) and _cm.jalt_ ([Jump and link via table](#insns-cm%5Fjalt)) are referred to as table jump. Table jump uses a 256-entry XLEN wide table in instruction memory to contain function addresses. The table must be a minimum of 64-byte aligned. Table entries follow the current data endianness. This is different from normal instruction fetch which is always little-endian. _cm.jt_ and _cm.jalt_ encodings index the table, giving access to functions within the full XLEN wide address space. This is used as a form of dictionary compression to reduce the code size of _jal_ / _auipc+jalr_ / _jr_ / _auipc+jr_ instructions. Table jump allows the linker to replace the following instruction sequences with a _cm.jt_ or _cm.jalt_ encoding, and an entry in the table: * 32-bit _j_ calls * 32-bit _jal_ ra calls * 64-bit _auipc+jr_ calls to fixed locations * 64-bit _auipc+jalr ra_ calls to fixed locations * The _auipc+jr/jalr_ sequence is used because the offset from the PC is out of the ±1 MB range. If a return address stack is implemented, then as _cm.jalt_ is equivalent to _jal ra_, it pushes to the stack. #### [](#29-1-14-1-jvt)29.1.14.1\. jvt The base of the table is in the jvt CSR (see [jvt CSR, table jump base vector and control register](#csrs-jvt)), each table entry is XLEN bits. If the same function is called with and without linking then it must have two entries in the table. This is typically caused by the same function being called with and without tail calling. #### [](#tablejump-fault-handling)29.1.14.2\. Table Jump Fault handling For a table jump instruction, the table entry that the instruction selects is considered an extension of the instruction itself. Hence, the execution of a table jump instruction involves two instruction fetches, the first to read the instruction (_cm.jt_/_cm.jalt_) and the second to read from the jump vector table (JVT). Both instruction fetches are _implicit_ reads, and both require execute permission; read permission is irrelevant. It is recommended that the second fetch be ignored for hardware triggers and breakpoints. Memory writes to the jump vector table require an instruction barrier (_fence.i_) to guarantee that they are visible to the instruction fetch. Multiple contexts may have different jump vector tables. JVT may be switched between them without an instruction barrier if the tables have not been updated in memory since the last _fence.i_. If an exception occurs on either instruction fetch, xEPC is set to the PC of the table jump instruction, xCAUSE is set as expected for the type of fault and xTVAL (if not set to zero) contains the fetch address which caused the fault. #### [](#csrs-jvt)29.1.14.3\. jvt CSR Synopsis Table jump base vector and control register Address: 0x017 Permissions: URW Format (RV32): ![svg](_images/svg-f52d7f22ac9cc7a7a4b770c35ab6318f1a1e76aa.svg) Format (RV64): ![svg](_images/svg-37d77d1b494eeccc692acc07d7fecc744d801b35.svg) Description The _jvt_ register is an XLEN-bit **WARL** read/write register that holds the jump table configuration, consisting of the jump table base address (BASE) and the jump table mode (MODE). If [29.1.10\. Zcmt](#Zcmt) is implemented then _jvt_ must also be implemented, but can contain a read-only value. If _jvt_ is writable, the set of values the register may hold can vary by implementation. The value in the BASE field must always be aligned on a 64-byte boundary. Note that the CSR contains only bits XLEN-1 through 6 of the address _base_. When computing jump-table accesses, the lower six bits of _base_ are filled with zeroes to obtain an XLEN-bit jump-table base address _jvt.base_ that is always aligned on a 64-byte boundary. _jvt.base_ is a virtual address, whenever virtual memory is enabled. The memory pointed to by _jvt.base_ is treated as instruction memory for the purpose of executing table jump instructions, implying execute access permission. __Table 2\. _jvt.mode_ definition__ | jvt.mode | Comment | | -------- | ------------------------------------ | | 000000 | Jump table mode | | others | **reserved for future standard use** | _jvt.mode_ is a **WARL** field, so can only be programmed to modes which are implemented. Therefore the discovery mechanism is to attempt to program different modes and read back the values to see which are available. Jump table mode _must_ be implemented. | | in future the RISC-V Unified Discovery method will report the available modes. | | --------------------------------------------------------------------------------- | Architectural State: _jvt_ CSR adds architectural state to the system software context (such as an OS process), therefore must be saved/restored on context switches. <<< #### [](#insns-cm%5Fjt)29.1.14.4\. cm.jt Synopsis jump via table Mnemonic cm.jt _index_ Encoding (RV32, RV64): ![svg](_images/svg-1071503e2d1f9a15172c6b55a3bf17c42dd2d62d.svg) | | For this encoding to decode as _cm.jt_, _index<32_, otherwise it decodes as _cm.jalt_, see [Jump and link via table](#insns-cm%5Fjalt). | | ------------------------------------------------------------------------------------------------------------------------------------------ | | | If jvt.mode = 0 (Jump Table Mode) then _cm.jt_ behaves as specified here. If jvt.mode is a reserved value, then _cm.jt_ is also reserved. In the future other defined values of jvt.mode may change the behaviour of _cm.jt_. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Assembly Syntax: ```asm cm.jt index ``` Description _cm.jt_ reads an entry from the jump vector table in memory and jumps to the address that was read. For further information see [Table Jump Overview](#insns-tablejump). Prerequisites None 32-bit equivalent: No direct equivalent encoding exists. Operation ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. # table_address is temporary internal state, it doesn't represent a real register # InstMemory is byte indexed switch(XLEN) { 32: table_address[XLEN-1:0] = jvt.base + (index<<2); 64: table_address[XLEN-1:0] = jvt.base + (index<<3); } //fetch from the jump table pc = InstMemory[table_address][XLEN-1:0]&~0x1; // Clear bit 0. ``` #### [](#insns-cm%5Fjalt)29.1.14.5\. cm.jalt Synopsis jump via table with optional link Mnemonic cm.jalt _index_ Encoding (RV32, RV64): ![svg](_images/svg-1071503e2d1f9a15172c6b55a3bf17c42dd2d62d.svg) | | For this encoding to decode as _cm.jalt_, _index>=32_, otherwise it decodes as _cm.jt_, see [Jump via table](#insns-cm%5Fjt). | | -------------------------------------------------------------------------------------------------------------------------------- | | | If jvt.mode = 0 (Jump Table Mode) then _cm.jalt_ behaves as specified here. If jvt.mode is a reserved value, then _cm.jalt_ is also reserved. In the future other defined values of jvt.mode may change the behaviour of _cm.jalt_. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Assembly Syntax: ```asm cm.jalt index ``` Description _cm.jalt_ reads an entry from the jump vector table in memory and jumps to the address that was read, linking to _ra_. For further information see [Table Jump Overview](#insns-tablejump). Prerequisites None 32-bit equivalent: No direct equivalent encoding exists. Operation ```sail //This is not SAIL, it's pseudocode. The SAIL hasn't been written yet. # table_address is temporary internal state, it doesn't represent a real register # InstMemory is byte indexed switch(XLEN) { 32: table_address[XLEN-1:0] = jvt.base + (index<<2); 64: table_address[XLEN-1:0] = jvt.base + (index<<3); } //fetch from the jump table ra = pc+2; pc = InstMemory[table_address][XLEN-1:0]&~0x1; // Clear bit 0. ``` 26.1. "Zfa" Extension for Additional Floating-Point Instructions, Version 1.0 ==================== ## [](#zfa)26.1\. "Zfa" Extension for Additional Floating-Point Instructions, Version 1.0 This chapter describes the Zfa standard extension, which adds instructions for immediate loads, IEEE 754-2019 minimum and maximum operations, round-to-integer operations, and quiet floating-point comparisons. For RV32D, the Zfa extension also adds instructions to transfer double-precision floating-point values to and from integer registers, and for RV64Q, it adds analogous instructions for quad-precision floating-point values. The Zfa extension depends on the F extension. ### [](#26-1-1-load-immediate-instructions)26.1.1\. Load-Immediate Instructions The FLI.S instruction loads one of 32 single-precision floating-point constants, encoded in the _rs1_ field, into floating-point register_rd_. The correspondence of _rs1_ field values and single-precision floating-point values is shown in [Table 1](#tab:flis). FLI.S is encoded like FMV.W.X, but with _rs2_\=1. __Table 1\. Immediate values loaded by the FLI.S instruction.__ | _rs1_ | Value | Sign | Exponent | Significand | | ----- | ------------------------- | ---- | -------- | ----------- | | 0 | −1.0 | 1 | 01111111 | 000…​000 | | 1 | _Minimum positive normal_ | 0 | 00000001 | 000…​000 | | 2 | 1.0 × 2−16 | 0 | 01101111 | 000…​000 | | 3 | 1.0 × 2−15 | 0 | 01110000 | 000…​000 | | 4 | 1.0 × 2−8 | 0 | 01110111 | 000…​000 | | 5 | 1.0 × 2−7 | 0 | 01111000 | 000…​000 | | 6 | 0.0625 (2−4) | 0 | 01111011 | 000…​000 | | 7 | 0.125 (2−3) | 0 | 01111100 | 000…​000 | | 8 | 0.25 | 0 | 01111101 | 000…​000 | | 9 | 0.3125 | 0 | 01111101 | 010…​000 | | 10 | 0.375 | 0 | 01111101 | 100…​000 | | 11 | 0.4375 | 0 | 01111101 | 110…​000 | | 12 | 0.5 | 0 | 01111110 | 000…​000 | | 13 | 0.625 | 0 | 01111110 | 010…​000 | | 14 | 0.75 | 0 | 01111110 | 100…​000 | | 15 | 0.875 | 0 | 01111110 | 110…​000 | | 16 | 1.0 | 0 | 01111111 | 000…​000 | | 17 | 1.25 | 0 | 01111111 | 010…​000 | | 18 | 1.5 | 0 | 01111111 | 100…​000 | | 19 | 1.75 | 0 | 01111111 | 110…​000 | | 20 | 2.0 | 0 | 10000000 | 000…​000 | | 21 | 2.5 | 0 | 10000000 | 010…​000 | | 22 | 3 | 0 | 10000000 | 100…​000 | | 23 | 4 | 0 | 10000001 | 000…​000 | | 24 | 8 | 0 | 10000010 | 000…​000 | | 25 | 16 | 0 | 10000011 | 000…​000 | | 26 | 128 (27) | 0 | 10000110 | 000…​000 | | 27 | 256 (28) | 0 | 10000111 | 000…​000 | | 28 | 215 | 0 | 10001110 | 000…​000 | | 29 | 216 | 0 | 10001111 | 000…​000 | | 30 | +∞ | 0 | 11111111 | 000…​000 | | 31 | _Canonical NaN_ | 0 | 11111111 | 100…​000 | | | The preferred assembly syntax for entries 1, 30, and 31 is min, inf, and nan, respectively. For entries 0 through 29 (including entry 1), the assembler will accept decimal constants in C-like syntax. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The set of 32 constants was chosen by examining floating-point libraries, including the C standard math library, and to optimize fixed-point to floating-point conversion. Entries 8-22 follow a regular encoding pattern. No entry sets mantissa bits other than the two most significant ones. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If the D extension is implemented, FLI.D performs the analogous operation, but loads a double-precision value into floating-point register _rd_. Note that entry 1 (corresponding to the minimum positive normal value) has a numerically different value for double-precision than for single-precision. FLI.D is encoded like FLI.S, but with_fmt_\=D. If the Q extension is implemented, FLI.Q performs the analogous operation, but loads a quad-precision value into floating-point register_rd_. Note that entry 1 (corresponding to the minimum positive normal value) has a numerically different value for quad-precision. FLI.Q is encoded like FLI.S, but with _fmt_\=Q. If the Zfh or Zvfh extension is implemented, FLI.H performs the analogous operation, but loads a half-precision floating-point value into register _rd_. Note that entry 1 (corresponding to the minimum positive normal value) has a numerically different value for half-precision. Furthermore, since 216 is not representable in half-precision floating-point, entry 29 in the table instead loads positive infinity—i.e., it is redundant with entry 30\. FLI.H is encoded like FLI.S, but with _fmt_\=H. | | Additionally, since 2−16 and 2−15 are subnormal in half-precision, entry 1 is numerically greater than entries 2 and 3 for FLI.H. | | ------------------------------------------------------------------------------------------------------------------------------------ | The FLI._fmt_ instructions never set any floating-point exception flags. ### [](#26-1-2-minimum-and-maximum-instructions)26.1.2\. Minimum and Maximum Instructions The FMINM.S and FMAXM.S instructions are defined like the FMIN.S and FMAX.S instructions, except that if either input is NaN, the result is the canonical NaN. If the D extension is implemented, FMINM.D and FMAXM.D instructions are analogously defined to operate on double-precision numbers. If the Zfh extension is implemented, FMINM.H and FMAXM.H instructions are analogously defined to operate on half-precision numbers. If the Q extension is implemented, FMINM.Q and FMAXM.Q instructions are analogously defined to operate on quad-precision numbers. These instructions are encoded like their FMIN and FMAX counterparts, but with instruction bit 13 set to 1. | | These instructions implement the IEEE 754-2019 minimum and maximum operations. | | --------------------------------------------------------------------------------- | ### [](#26-1-3-round-to-integer-instructions)26.1.3\. Round-to-Integer Instructions The FROUND.S instruction rounds the single-precision floating-point number in floating-point register _rs1_ to an integer, according to the rounding mode specified in the instruction’s _rm_ field. It then writes that integer, represented as a single-precision floating-point number, to floating-point register _rd_. Zero and infinite inputs are copied to_rd_ unmodified. Signaling NaN inputs cause the invalid operation exception flag to be set; no other exception flags are set. FROUND.S is encoded like FCVT.S.D, but with _rs2_\=4. The FROUNDNX.S instruction is defined similarly, but it also sets the inexact exception flag if the input differs from the rounded result and is not NaN. FROUNDNX.S is encoded like FCVT.S.D, but with _rs2_\=5. If the D extension is implemented, FROUND.D and FROUNDNX.D instructions are analogously defined to operate on double-precision numbers. They are encoded like FCVT.D.S, but with _rs2_\=4 and 5, respectively, If the Zfh extension is implemented, FROUND.H and FROUNDNX.H instructions are analogously defined to operate on half-precision numbers. They are encoded like FCVT.H.S, but with _rs2_\=4 and 5, respectively, If the Q extension is implemented, FROUND.Q and FROUNDNX.Q instructions are analogously defined to operate on quad-precision numbers. They are encoded like FCVT.Q.S, but with _rs2_\=4 and 5, respectively, | | The FROUNDNX._fmt_ instructions implement the IEEE 754-2019 roundToIntegralExact operation, and the FROUND._fmt_ instructions implement the other operations in the roundToIntegral family. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#26-1-4-modular-convert-to-integer-instruction)26.1.4\. Modular Convert-to-Integer Instruction The FCVTMOD.W.D instruction is defined similarly to the FCVT.W.D instruction, with the following differences. FCVTMOD.W.D always rounds towards zero. Bits 31:0 are taken from the rounded, unbounded two’s complement result, then sign-extended to XLEN bits and written to integer register _rd_. ±∞ and NaN are converted to zero. Floating-point exception flags are raised the same as they would be for FCVT.W.D with the same input operand. This instruction is only provided if the D extension is implemented. It is encoded like FCVT.W.D, but with the rs2 field set to 8 and the _rm_field set to 1 (RTZ). Other _rm_ values are _reserved_. | | The assembly syntax requires the RTZ rounding mode to be explicitly specified, i.e., fcvtmod.w.d rd, rs1, rtz. The FCVTMOD.W.D instruction was added principally to accelerate the processing of JavaScript Numbers. Numbers are double-precision values, but some operators implicitly truncate them to signed integers mod 232. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#26-1-5-move-instructions)26.1.5\. Move Instructions For RV32 only, if the D extension is implemented, the FMVH.X.D instruction moves bits 63:32 of floating-point register _rs1_ into integer register _rd_. It is encoded in the OP-FP major opcode with_funct3_\=0, _rs2_\=1, and _funct7_\=1110001. | | FMVH.X.D is used in conjunction with the existing FMV.X.W instruction to move a double-precision floating-point number to a pair of x-registers. | | --------------------------------------------------------------------------------------------------------------------------------------------------- | For RV32 only, if the D extension is implemented, the FMVP.D.X instruction moves a double-precision number from a pair of integer registers into a floating-point register. Integer registers _rs1_ and_rs2_ supply bits 31:0 and 63:32, respectively; the result is written to floating-point register _rd_. FMVP.D.X is encoded in the OP-FP major opcode with _funct3_\=0 and _funct7_\=1011001. For RV64 only, if the Q extension is implemented, the FMVH.X.Q instruction moves bits 127:64 of floating-point register _rs1_ into integer register _rd_. It is encoded in the OP-FP major opcode with_funct3_\=0, _rs2_\=1, and _funct7_\=1110011. | | FMVH.X.Q is used in conjunction with the existing FMV.X.D instruction to move a quad-precision floating-point number to a pair of x-registers. | | ------------------------------------------------------------------------------------------------------------------------------------------------- | For RV64 only, if the Q extension is implemented, the FMVP.Q.X instruction moves a double-precision number from a pair of integer registers into a floating-point register. Integer registers _rs1_ and_rs2_ supply bits 63:0 and 127:64, respectively; the result is written to floating-point register _rd_. FMVP.Q.X is encoded in the OP-FP major opcode with _funct3_\=0 and _funct7_\=1011011. ### [](#26-1-6-comparison-instructions)26.1.6\. Comparison Instructions The FLEQ.S and FLTQ.S instructions are defined like the FLE.S and FLT.S instructions, except that quiet NaN inputs do not cause the invalid operation exception flag to be set. If the D extension is implemented, FLEQ.D and FLTQ.D instructions are analogously defined to operate on double-precision numbers. If the Zfh extension is implemented, FLEQ.H and FLTQ.H instructions are analogously defined to operate on half-precision numbers. If the Q extension is implemented, FLEQ.Q and FLTQ.Q instructions are analogously defined to operate on quad-precision numbers. These instructions are encoded like their FLE and FLT counterparts, but with instruction bit 14 set to 1. | | We do not expect analogous comparison instructions will be added to the vector ISA, since they can be reasonably efficiently emulated using masking. | | ------------------------------------------------------------------------------------------------------------------------------------------------------- | 24.1. "Zfh" and "Zfhmin" Extensions for Half-Precision Floating-Point, Version 1.0 ==================== ## [](#chap:zfh)24.1\. "Zfh" and "Zfhmin" Extensions for Half-Precision Floating-Point, Version 1.0 This chapter describes the Zfh standard extension for 16-bit half-precision binary floating-point instructions compliant with the IEEE 754-2008 arithmetic standard. The Zfh extension depends on the single-precision floating-point extension, F. The NaN-boxing scheme described in [NaN Boxing of Narrower Values](d-st-ext.html#nanboxing) is extended to allow a half-precision value to be NaN-boxed inside a single-precision value (which may be recursively NaN-boxed inside a double- or quad-precision value when the D or Q extension is present). | | This extension primarily provides instructions that consume half-precision operands and produce half-precision results. However, it is also common to compute on half-precision data using higher intermediate precision. Although this extension provides explicit conversion instructions that suffice to implement that pattern, future extensions might further accelerate such computation with additional instructions that implicitly widen their operands—e.g., half×half+single→single—or implicitly narrow their results—e.g., half+single→half. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#24-1-1-half-precision-load-and-store-instructions)24.1.1\. Half-Precision Load and Store Instructions New 16-bit variants of LOAD-FP and STORE-FP instructions are added, encoded with a new value for the funct3 width field. ![svg](_images/svg-43b5c2b0a97281843159ef82964b366feac332bd.svg) ![svg](_images/svg-cbf3aef7c9e9452cdb40ee9820c3cca33bb798c9.svg) FLH and FSH are only guaranteed to execute atomically if the effective address is naturally aligned. FLH and FSH do not modify the bits being transferred; in particular, the payloads of non-canonical NaNs are preserved. FLH NaN-boxes the result written to _rd_, whereas FSH ignores all but the lower 16 bits in _rs2_. ### [](#24-1-2-half-precision-computational-instructions)24.1.2\. Half-Precision Computational Instructions A new supported format is added to the format field of most instructions, as shown in [Table 1](#tab:fpextfmth). __Table 1\. Format field encoding.__ | _fmt_ field | Mnemonic | Meaning | | ----------- | -------- | ----------------------- | | 00 | S | 32-bit single-precision | | 01 | D | 64-bit double-precision | | 10 | H | 16-bit half-precision | | 11 | Q | 128-bit quad-precision | The half-precision floating-point computational instructions are defined analogously to their single-precision counterparts, but operate on half-precision operands and produce half-precision results. ![svg](_images/svg-46361031bb67f0d9c15ff2d4a33cc02bd86b95e5.svg) ![svg](_images/svg-99b93bf80c94c94f0ec817b44bdb5a7165378a42.svg) ### [](#24-1-3-half-precision-conversion-and-move-instructions)24.1.3\. Half-Precision Conversion and Move Instructions New floating-point-to-integer and integer-to-floating-point conversion instructions are added. These instructions are defined analogously to the single-precision-to-integer and integer-to-single-precision conversion instructions. FCVT.W.H or FCVT.L.H converts a half-precision floating-point number to a signed 32-bit or 64-bit integer, respectively. FCVT.H.W or FCVT.H.L converts a 32-bit or 64-bit signed integer, respectively, into a half-precision floating-point number.FCVT.WU.H, FCVT.LU.H, FCVT.H.WU, and FCVT.H.LU variants convert to or from unsigned integer values. FCVT.L\[U\].H and FCVT.H.L\[U\] are RV64-only instructions. ![svg](_images/svg-48aa1e6d42b94d88e20f7fcd27e6f716202f453c.svg) New floating-point-to-floating-point conversion instructions are added. These instructions are defined analogously to the double-precision floating-point-to-floating-point conversion instructions. FCVT.S.H or FCVT.H.S converts a half-precision floating-point number to a single-precision floating-point number, or vice-versa, respectively. If the D extension is present, FCVT.D.H or FCVT.H.D converts a half-precision floating-point number to a double-precision floating-point number, or vice-versa, respectively. If the Q extension is present, FCVT.Q.H or FCVT.H.Q converts a half-precision floating-point number to a quad-precision floating-point number, or vice-versa, respectively. ![svg](_images/svg-0ecff240910136246b1ea4cf5796e2cf903635d6.svg) Floating-point to floating-point sign-injection instructions, FSGNJ.H, FSGNJN.H, and FSGNJX.H are defined analogously to the single-precision sign-injection instruction. ![svg](_images/svg-9062450444720ccfbffb4f64412ffad9b70c2027.svg) Instructions are provided to move bit patterns between the floating-point and integer registers. FMV.X.H moves the half-precision value in floating-point register _rs1_ to a representation in IEEE 754-2008 standard encoding in integer register _rd_, filling the upper XLEN-16 bits with copies of the floating-point number’s sign bit. FMV.H.X moves the half-precision value encoded in IEEE 754-2008 standard encoding from the lower 16 bits of integer register _rs1_ to the floating-point register _rd_, NaN-boxing the result. FMV.X.H and FMV.H.X do not modify the bits being transferred; in particular, the payloads of non-canonical NaNs are preserved. ![svg](_images/svg-86412cb7a6df56627d86059a8dbc406ee8faf887.svg) ### [](#flt-pt-to-int-move)24.1.4\. Half-Precision Floating-Point Compare Instructions The half-precision floating-point compare instructions are defined analogously to their single-precision counterparts, but operate on half-precision operands. ![svg](_images/svg-97b3530202d91cc76da0d43426f7bd7ff5cc0883.svg) ### [](#half-pr-flt-pt-compare)24.1.5\. Half-Precision Floating-Point Classify Instruction The half-precision floating-point classify instruction, FCLASS.H, is defined analogously to its single-precision counterpart, but operates on half-precision operands. ![svg](_images/svg-ce65c8430d703053d0c5f1b1bccb74e9706e2300.svg) ### [](#half-pr-flt-class)24.1.6\. "Zfhmin" Standard Extension for Minimal Half-Precision Floating-Point This section describes the Zfhmin standard extension, which provides minimal support for 16-bit half-precision binary floating-point instructions. The Zfhmin extension is a subset of the Zfh extension, consisting only of data transfer and conversion instructions. Like Zfh, the Zfhmin extension depends on the single-precision floating-point extension, F. The expectation is that Zfhmin software primarily uses the half-precision format for storage, performing most computation in higher precision. The Zfhmin extension includes the following instructions from the Zfh extension: FLH, FSH, FMV.X.H, FMV.H.X, FCVT.S.H, and FCVT.H.S. If the D extension is present, the FCVT.D.H and FCVT.H.D instructions are also included. If the Q extension is present, the FCVT.Q.H and FCVT.H.Q instructions are additionally included. | | Zfhmin does not include the FSGNJ.H instruction, because it suffices to instead use the FSGNJ.S instruction to move half-precision values between floating-point registers. Half-precision addition, subtraction, multiplication, division, and square-root operations can be faithfully emulated by converting the half-precision operands to single-precision, performing the operation using single-precision arithmetic, then converting back to half-precision. \[[21](../biblio/bibliography.html#bib-roux:hal-01091186)\] Performing half-precision fused multiply-addition using this method incurs a 1-ulp error on some inputs for the RNE and RMM rounding modes. Conversion from 8- or 16-bit integers to half-precision can be emulated by first converting to single-precision, then converting to half-precision. Conversion from 32-bit integer can be emulated by first converting to double-precision. If the D extension is not present and a 1-ulp error under RNE or RMM is tolerable, 32-bit integers can be first converted to single-precision instead. The same remark applies to conversions from 64-bit integers without the Q extension. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 27.1. "Zfinx", "Zdinx", "Zhinx", "Zhinxmin" Extensions for Floating-Point in Integer Registers, Version 1.0 ==================== ## [](#sec:zfinx)27.1\. "Zfinx", "Zdinx", "Zhinx", "Zhinxmin" Extensions for Floating-Point in Integer Registers, Version 1.0 This chapter defines the "Zfinx" extension (pronounced "z-f-in-x") that provides instructions similar to those in the standard floating-point F extension for single-precision floating-point instructions but which operate on the `x` registers instead of the `f`registers. This chapter also defines the "Zdinx", "Zhinx", and "Zhinxmin" extensions that provide similar instructions for other floating-point precisions. | | The F extension uses separate f registers for floating-point computation, to reduce register pressure and simplify the provision of register-file ports for wide superscalars. However, the additional architectural state increases the minimal implementation cost. By eliminating the f registers, the Zfinx extension substantially reduces the cost of simple RISC-V implementations with floating-point instruction-set support. Zfinx also reduces context-switch cost. In general, software that assumes the presence of the F extension is incompatible with software that assumes the presence of the Zfinx extension, and vice versa. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The Zfinx extension adds all of the instructions that the F extension adds, _except_ for the transfer instructions FLW, FSW, FMV.W.X, FMV.X.W, C.FLW\[SP\], and C.FSW\[SP\]. | | Zfinx software uses integer loads and stores to transfer floating-point values from and to memory. Transfers between registers use either integer arithmetic or floating-point sign-injection instructions. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The Zfinx variants of these F-extension instructions have the same semantics, except that whenever such an instruction would have accessed an `f` register, it instead accesses the `x` register with the same number. The Zfinx extension depends on the "Zicsr" extension for control and status register access. ### [](#27-1-1-processing-of-narrower-values)27.1.1\. Processing of Narrower Values Floating-point operands of width _w_ < XLEN bits occupy bits _w_\-1:0 of an `x` register. Floating-point operations on _w_\-bit operands ignore operand bits XLEN-1: _w_. Floating-point operations that produce _w_ < XLEN-bit results fill bits XLEN-1: _w_ with copies of bit _w_\-1 (the sign bit). | | The NaN-boxing scheme employed in the f registers was designed to efficiently support recoded floating-point formats. Recoding is less practical for Zfinx, though, since the same registers hold both floating-point and integer operands. Hence, the need for NaN boxing is diminished. Sign-extending 32-bit floating-point numbers when held in RV64 xregisters is compatible with the existing RV64 calling conventions, which leave bits 63-32 undefined when passing a 32-bit floating point value in x registers. To keep the architecture more regular, we extend this pattern to 16-bit floating-point numbers in both RV32 and RV64. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#27-1-2-zdinx)27.1.2\. Zdinx The Zdinx extension provides analogous double-precision floating-point instructions. The Zdinx extension depends upon the Zfinx extension. The Zdinx extension adds all of the instructions that the D extension adds, _except_ for the transfer instructions FLD, FSD, FMV.D.X, FMV.X.D, C.FLD\[SP\], and C.FSD\[SP\]. The Zdinx variants of these D-extension instructions have the same semantics, except that whenever such an instruction would have accessed an `f` register, it instead accesses the `x` register with the same number. ### [](#27-1-3-processing-of-wider-values)27.1.3\. Processing of Wider Values Double-precision operands in RV32Zdinx are held in aligned `x`\-register pairs, i.e., register numbers must be even. Use of misaligned (odd-numbered) registers for double-width floating-point operands is_reserved_. Regardless of endianness, the lower-numbered register holds the low-order bits, and the higher-numbered register holds the high-order bits: e.g., bits 31:0 of a double-precision operand in RV32Zdinx might be held in register `x14`, with bits 63:32 of that operand held in`x15`. When a double-width floating-point result is written to `x0`, the entire write takes no effect: e.g., for RV32Zdinx, writing a double-precision result to `x0` does not cause `x1` to be written. When `x0` is used as a double-width floating-point operand, the entire operand is zero—i.e., `x1` is not accessed. | | Load-pair and store-pair instructions are contained in a separate extension (see [Extensions for Load/Store pair for RV32](zilsd.html#sec:zilsd)). In case this is not available, transferring double-precision operands in RV32Zdinx from or to memory requires two loads or stores. Register moves need only a single FSGNJ.D instruction, however. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#27-1-4-zhinx)27.1.4\. Zhinx The Zhinx extension provides analogous half-precision floating-point instructions. The Zhinx extension depends upon the Zfinx extension. The Zhinx extension adds all of the instructions that the Zfh extension adds, _except_ for the transfer instructions FLH, FSH, FMV.H.X, and FMV.X.H. The Zhinx variants of these Zfh-extension instructions have the same semantics, except that whenever such an instruction would have accessed an `f` register, it instead accesses the `x` register with the same number. ### [](#27-1-5-zhinxmin)27.1.5\. Zhinxmin The Zhinxmin extension provides minimal support for 16-bit half-precision floating-point instructions that operate on the `x`registers. The Zhinxmin extension depends upon the Zfinx extension. The Zhinxmin extension includes the following instructions from the Zhinx extension: FCVT.S.H and FCVT.H.S. If the Zdinx extension is present, the FCVT.D.H and FCVT.H.D instructions are also included. | | In the future, an RV64Zqinx quad-precision extension could be defined analogously to RV32Zdinx. An RV32Zqinx extension could also be defined but would require quad-register groups. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#27-1-6-privileged-architecture-implications)27.1.6\. Privileged Architecture Implications In the standard privileged architecture defined in Volume II, the`mstatus` field FS is hardwired to 0 if the Zfinx extension is implemented, and FS no longer affects the trapping behavior of floating-point instructions or `fcsr` accesses. The `misa` bits F, D, and Q are hardwired to 0 when the Zfinx extension is implemented. | | A future discoverability mechanism might be used to probe the existence of the Zfinx, Zhinx, and Zdinx extensions. | | --------------------------------------------------------------------------------------------------------------------- | 11.1. "Zicond" Extension for Integer Conditional Operations, Version 1.0.0 ==================== ## [](#Zicond)11.1\. "Zicond" Extension for Integer Conditional Operations, Version 1.0.0 The Zicond extension defines two R-type instructions that support branchless conditional operations. | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | ---------------------------- | ------------------------------------------------------------------- | | ✓ | ✓ | czero.eqz _rd_, _rs1_, _rs2_ | [Conditional zero, if condition is equal to zero](#insns-czero-eqz) | | ✓ | ✓ | czero.nez _rd_, _rs1_, _rs2_ | [Conditional zero, if condition is nonzero](#insns-czero-nez) | ### [](#11-1-1-instructions-in-alphabetical-order)11.1.1\. Instructions (in alphabetical order) #### [](#insns-czero-eqz)11.1.1.1\. czero.eqz Synopsis Moves zero to a register _rd_, if the condition _rs2_ is equal to zero, otherwise moves _rs1_ to _rd_. Mnemonic czero.eqz _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-34389227572362e4a64211391b05200befc2cb7b.svg) Description If _rs2_ contains the value zero, this instruction writes the value zero to _rd_. Otherwise, this instruction copies the contents of _rs1_ to _rd_. This instruction carries a syntactic dependency from both _rs1_ and _rs2_ to _rd_. Furthermore, if the Zkt extension is implemented, this instruction’s timing is independent of the data values in _rs1_ and _rs2_. SAIL code ```sail let condition = X(rs2); result : xlenbits = if (condition == zeros()) then zeros() else X(rs1); X(rd) = result; ``` #### [](#insns-czero-nez)11.1.1.2\. czero.nez Synopsis Moves zero to a register _rd_, if the condition _rs2_ is nonzero, otherwise moves _rs1_ to _rd_. Mnemonic czero.nez _rd_, _rs1_, _rs2_ Encoding ![svg](_images/svg-7d786258fac0cfe4283212b4dde9b0da28236053.svg) Description If _rs2_ contains a nonzero value, this instruction writes the value zero to _rd_. Otherwise, this instruction copies the contents of _rs1_ to _rd_. This instruction carries a syntactic dependency from both _rs1_ and _rs2_ to _rd_. Furthermore, if the Zkt extension is implemented, this instruction’s timing is independent of the data values in _rs1_ and _rs2_. SAIL code ```sail let condition = X(rs2); result : xlenbits = if (condition != zeros()) then zeros() else X(rs1); X(rd) = result; ``` ### [](#11-1-2-usage-examples)11.1.2\. Usage examples The instructions from this extension can be used to construct sequences that perform conditional-arithmetic, conditional-bitwise-logical, and conditional-select operations. #### [](#11-1-2-1-instruction-sequences)11.1.2.1\. Instruction sequences | Operation | Instruction sequence | Length | | --------------------------------------------------------------------------- | ------------------------------------------------------------------------ | ----------------------------- | | **Conditional add, if zero** rd = (rc == 0) ? (rs1 + rs2) : rs1 | czero.nez rd, rs2, rc add rd, rs1, rd | 2 insns | | **Conditional add, if non-zero** rd = (rc != 0) ? (rs1 + rs2) : rs1 | czero.eqz rd, rs2, rc add rd, rs1, rd | | | **Conditional subtract, if zero** rd = (rc == 0) ? (rs1 - rs2) : rs1 | czero.nez rd, rs2, rc sub rd, rs1, rd | | | **Conditional subtract, if non-zero** rd = (rc != 0) ? (rs1 - rs2) : rs1 | czero.eqz rd, rs2, rc sub rd, rs1, rd | | | **Conditional bitwise-or, if zero** rd = (rc == 0) ? (rs1 \| rs2) : rs1 | czero.nez rd, rs2, rc or rd, rs1, rd | | | **Conditional bitwise-or, if non-zero** rd = (rc != 0) ? (rs1 \| rs2) : rs1 | czero.eqz rd, rs2, rc or rd, rs1, rd | | | **Conditional bitwise-xor, if zero** rd = (rc == 0) ? (rs1 ^ rs2) : rs1 | czero.nez rd, rs2, rc xor rd, rs1, rd | | | **Conditional bitwise-xor, if non-zero** rd = (rc != 0) ? (rs1 ^ rs2) : rs1 | czero.eqz rd, rs2, rc xor rd, rs1, rd | | | **Conditional bitwise-and, if zero** rd = (rc == 0) ? (rs1 & rs2) : rs1 | and rd, rs1, rs2 czero.eqz rtmp, rs1, rc or rd, rd, rtmp | 3 insns(requires 1 temporary) | | **Conditional bitwise-and, if non-zero** rd = (rc != 0) ? (rs1 & rs2) : rs1 | and rd, rs1, rs2 czero.nez rtmp, rs1, rc or rd, rd, rtmp | | | **Conditional select, if zero** rd = (rc == 0) ? rs1 : rs2 | czero.nez rd, rs1, rc czero.eqz rtmp, rs2, rc add rd, rd, rtmp | | | **Conditional select, if non-zero** rd = (rc != 0) ? rs1 : rs2 | czero.eqz rd, rs1, rc czero.nez rtmp, rs2, rc add rd, rd, rtmp | | 6.1. "Zicsr" Extension for Control and Status Register (CSR) Instructions, Version 2.0 ==================== ## [](#csrinsts)6.1\. "Zicsr" Extension for Control and Status Register (CSR) Instructions, Version 2.0 RISC-V defines a separate address space of 4096 Control and Status registers associated with each hart. This chapter defines the full set of CSR instructions that operate on these CSRs. | | While CSRs are primarily used by the privileged architecture, there are several uses in unprivileged code including for counters and timers, and for floating-point status. The counters and timers are no longer considered mandatory parts of the standard base ISAs, and so the CSR instructions required to access them have been moved out of [RV32I Base Integer Instruction Set, Version 2.1](rv32.html) into this separate chapter. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#6-1-1-csr-instructions)6.1.1\. CSR Instructions All CSR instructions atomically read-modify-write a single CSR, whose CSR specifier is encoded in the 12-bit _csr_ field of the instruction held in bits 31-20\. The immediate forms use a 5-bit zero-extended immediate encoded in the _rs1_ field. ![svg](_images/svg-e61af5754aca1bf1e34eb29951401f81b92c5b83.svg) The CSRRW (Atomic Read/Write CSR) instruction atomically swaps values in the CSRs and integer registers. CSRRW reads the old value of the CSR, zero-extends the value to XLEN bits, then writes it to integer register_rd_. The initial value in _rs1_ is written to the CSR. If _rd_\=`x0`, then the instruction shall not read the CSR and shall not cause any of the side effects that might occur on a CSR read. The CSRRS (Atomic Read and Set Bits in CSR) instruction reads the value of the CSR, zero-extends the value to XLEN bits, and writes it to integer register _rd_. The initial value in integer register _rs1_ is treated as a bit mask that specifies bit positions to be set in the CSR. Any bit that is high in _rs1_ will cause the corresponding bit to be set in the CSR, if that CSR bit is writable. The CSRRC (Atomic Read and Clear Bits in CSR) instruction reads the value of the CSR, zero-extends the value to XLEN bits, and writes it to integer register _rd_. The initial value in integer register _rs1_ is treated as a bit mask that specifies bit positions to be cleared in the CSR. Any bit that is high in _rs1_ will cause the corresponding bit to be cleared in the CSR, if that CSR bit is writable. | | Since CSRRS and CSRRC perform a read-modify-write operation, any bits that read as a different value to their underlying value may be modified by these instructions even if the corresponding bit is not set in _rs1_. For example, pmpaddr_n_\[G-1\] may have an underlying value of 1 but read as 0\. Executing CSRRC or CSRRS to modify a different bit will cause 0 to be read from pmpaddr_n_\[G-1\] and then written back, updating the underlying value to 0. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | For both CSRRS and CSRRC, if _rs1_\=`x0`, then the instruction will not write to the CSR at all, and so shall not cause any of the side effects that might otherwise occur on a CSR write, nor raise illegal-instruction exceptions on accesses to read-only CSRs. Both CSRRS and CSRRC always read the addressed CSR and cause any read side effects regardless of_rs1_ and _rd_ fields. Note that if _rs1_ specifies a register other than `x0`, and that register holds a zero value, the instruction will not action any attendant per-field side effects, but will action any side effects caused by writing to the entire CSR. A CSRRW with _rs1_\=`x0` will attempt to write zero to the destination CSR. The CSRRWI, CSRRSI, and CSRRCI variants are similar to CSRRW, CSRRS, and CSRRC respectively, except they update the CSR using an XLEN-bit value obtained by zero-extending a 5-bit unsigned immediate (uimm\[4:0\]) field encoded in the _rs1_ field instead of a value from an integer register.For CSRRSI and CSRRCI, if the uimm\[4:0\] field is zero, then these instructions will not write to the CSR, and shall not cause any of the side effects that might otherwise occur on a CSR write, nor raise illegal-instruction exceptions on accesses to read-only CSRs. For CSRRWI, if _rd_\=`x0`, then the instruction shall not read the CSR and shall not cause any of the side effects that might occur on a CSR read.Both CSRRSI and CSRRCI will always read the CSR and cause any read side effects regardless of _rd_ and _rs1_ fields. __Table 1\. Conditions determining whether a CSR instruction reads or writes the specified CSR.__ | **Register operand** | | | | | | --------------------- | ---------- | ----------- | --------- | ---------- | | Instruction | _rd_ is x0 | _rs1_ is x0 | Reads CSR | Writes CSR | | CSRRW | Yes | \- | No | Yes | | CSRRW | No | \- | Yes | Yes | | CSRRS/CSRRC | \- | Yes | Yes | No | | CSRRS/CSRRC | \- | No | Yes | Yes | | **Immediate operand** | | | | | | Instruction | _rd_ is x0 | _uimm_\=0 | Reads CSR | Writes CSR | | CSRRWI | Yes | \- | No | Yes | | CSRRWI | No | \- | Yes | Yes | | CSRRSI/CSRRCI | \- | Yes | Yes | No | | CSRRSI/CSRRCI | \- | No | Yes | Yes | [Table 1](#csrsideeffects) summarizes the behavior of the CSR instructions with respect to whether they read and/or write the CSR. In addition to side effects that occur as a consequence of reading or writing a CSR, individual fields within a CSR might have side effects when written. The CSRRW\[I\] instructions action side effects for all such fields within the written CSR. The CSRRS\[I\] and CSRRC\[I\] instructions only action side effects for fields for which the _rs1_ or _uimm_ argument has at least one bit set corresponding to that field. | | As of this writing, no standard CSRs have side effects on field writes. Hence, whether a standard CSR access has any side effects can be determined solely from the opcode. Defining CSRs with side effects on field writes is not recommended. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For any event or consequence that occurs due to a CSR having a particular value, if a write to the CSR gives it that value, the resulting event or consequence is said to be an _indirect effect_ of the write. Indirect effects of a CSR write are not considered by the RISC-V ISA to be side effects of that write. | | An example of side effects for CSR accesses would be if reading from a specific CSR causes a light bulb to turn on, while writing an odd value to the same CSR causes the light to turn off. Assume writing an even value has no effect. In this case, both the read and write have side effects controlling whether the bulb is lit, as this condition is not determined solely from the CSR value. (Note that after writing an odd value to the CSR to turn off the light, then reading to turn the light on, writing again the same odd value causes the light to turn off again. Hence, on the last write, it is not a change in the CSR value that turns off the light.) On the other hand, if a bulb is rigged to light whenever the value of a particular CSR is odd, then turning the light on and off is not considered a side effect of writing to the CSR but merely an indirect effect of such writes. More concretely, the RISC-V privileged architecture defined in Volume II specifies that certain combinations of CSR values cause a trap to occur. When an explicit write to a CSR creates the conditions that trigger the trap, the trap is not considered a side effect of the write but merely an indirect effect. Standard CSRs do not have any side effects on reads. Standard CSRs may have side effects on writes. Custom extensions might add CSRs for which accesses have side effects on either reads or writes. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Some CSRs, such as the instructions-retired counter, `instret`, may be modified as side effects of instruction execution. In these cases, if a CSR access instruction reads a CSR, it reads the value prior to the execution of the instruction. If a CSR access instruction writes such a CSR, the explicit write is done instead of the update from the side effect. In particular, a value written to `instret` by one instruction will be the value read by the following instruction. The assembler pseudoinstruction to read a CSR, CSRR _rd, csr_, is encoded as CSRRS _rd, csr, x0_. The assembler pseudoinstruction to write a CSR, CSRW _csr, rs1_, is encoded as CSRRW _x0, csr, rs1_, while CSRWI_csr, uimm_, is encoded as CSRRWI _x0, csr, uimm_. Further assembler pseudoinstructions are defined to set and clear bits in the CSR when the old value is not required: CSRS/CSRC _csr, rs1_; CSRSI/CSRCI _csr, uimm_. #### [](#6-1-1-1-csr-access-ordering)6.1.1.1\. CSR Access Ordering Each RISC-V hart normally observes its own CSR accesses, including its implicit CSR accesses, as performed in program order. In particular, unless specified otherwise, a CSR access is performed after the execution of any prior instructions in program order whose behavior modifies or is modified by the CSR state and before the execution of any subsequent instructions in program order whose behavior modifies or is modified by the CSR state. Furthermore, an explicit CSR read returns the CSR state before the execution of the instruction, while an explicit CSR write suppresses and overrides any implicit writes or modifications to the same CSR by the same instruction. Likewise, any side effects from an explicit CSR access are normally observed to occur synchronously in program order. Unless specified otherwise, the full consequences of any such side effects are observable by the very next instruction, and no consequences may be observed out-of-order by preceding instructions. (Note the distinction made earlier between side effects and indirect effects of CSR writes.) For the RVWMO memory consistency model ([RVWMO Memory Consistency Model](rvwmo.html)), CSR accesses are weakly ordered by default, so other harts or devices may observe CSR accesses in an order different from program order. In addition, CSR accesses are not ordered with respect to explicit memory accesses, unless a CSR access modifies the execution behavior of the instruction that performs the explicit memory access or unless a CSR access and an explicit memory access are ordered by either the syntactic dependencies defined by the memory model or the ordering requirements defined by the Memory-Ordering PMAs section in Volume II of this manual. To enforce ordering in all other cases, software should execute a FENCE instruction between the relevant accesses. For the purposes of the FENCE instruction, CSR read accesses are classified as device input (I), and CSR write accesses are classified as device output (O). | | Informally, the CSR space acts as a weakly ordered memory-mapped I/O region, as defined by the Memory-Ordering PMAs section in Volume II of this manual. As a result, the order of CSR accesses with respect to all other accesses is constrained by the same mechanisms that constrain the order of memory-mapped I/O accesses to such a region. These CSR-ordering constraints are imposed to support ordering main memory and memory-mapped I/O accesses with respect to CSR accesses that are visible to, or affected by, devices or other harts. Examples include the time, cycle, and mcycle CSRs, in addition to CSRs that reflect pending interrupts, like mip and sip. Note that implicit reads of such CSRs (e.g., taking an interrupt because of a change in mip) are also ordered as device input. Most CSRs (including, e.g., the fcsr) are not visible to other harts; their accesses can be freely reordered in the global memory order with respect to FENCE instructions without violating this specification. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The hardware platform may define that accesses to certain CSRs are strongly ordered, as defined by the Memory-Ordering PMAs section in Volume II of this manual. Accesses to strongly ordered CSRs have stronger ordering constraints with respect to accesses to both weakly ordered CSRs and accesses to memory-mapped I/O regions. | | The rules for the reordering of CSR accesses in the global memory order should probably be moved to [RVWMO Memory Consistency Model](rvwmo.html) concerning the RVWMO memory consistency model. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 5.1. "Zifencei" Extension for Instruction-Fetch Fence, Version 2.0 ==================== ## [](#zifencei)5.1\. "Zifencei" Extension for Instruction-Fetch Fence, Version 2.0 This chapter defines the "Zifencei" extension, which includes the FENCE.I instruction that provides explicit synchronization between writes to instruction memory and instruction fetches on the same hart. Currently, this instruction is the only standard mechanism to ensure that stores visible to a hart will also be visible to its instruction fetches. | | We considered but did not include a "store instruction word" instruction as in \[[17](../biblio/bibliography.html#bib-majc)\]. JIT compilers may generate a large trace of instructions before a single FENCE.I, and amortize any instruction cache snooping/invalidation overhead by writing translated instructions to memory regions that are known not to reside in the I-cache. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --- | | The FENCE.I instruction was designed to support a wide variety of implementations. A simple implementation can flush the local instruction cache and the instruction pipeline when the FENCE.I is executed. A more complex implementation might snoop the instruction (data) cache on every data (instruction) cache miss, or use an inclusive unified private L2 cache to invalidate lines from the primary instruction cache when they are being written by a local store instruction. If instruction and data caches are kept coherent in this way, or if the memory system consists of only uncached RAMs, then just the fetch pipeline needs to be flushed at a FENCE.I. The FENCE.I instruction was previously part of the base I instruction set. Two main issues are driving moving this out of the mandatory base, although at time of writing it is still the only standard method for maintaining instruction-fetch coherence. First, it has been recognized that on some systems, FENCE.I will be expensive to implement and alternate mechanisms are being discussed in the memory model task group. In particular, for designs that have an incoherent instruction cache and an incoherent data cache, or where the instruction cache refill does not snoop a coherent data cache, both caches must be completely flushed when a FENCE.I instruction is encountered. This problem is exacerbated when there are multiple levels of I and D cache in front of a unified cache or outer memory system. Second, the instruction is not powerful enough to make available at user level in a Unix-like operating system environment. The FENCE.I only synchronizes the local hart, and the OS can reschedule the user hart to a different physical hart after the FENCE.I. This would require the OS to execute an additional FENCE.I as part of every context migration. For this reason, the standard Linux ABI has removed FENCE.I from user-level and now requires a system call to maintain instruction-fetch coherence, which allows the OS to minimize the number of FENCE.I executions required on current systems and provides forward-compatibility with future improved instruction-fetch coherence mechanisms. Future approaches to instruction-fetch coherence under discussion include providing more restricted versions of FENCE.I that only target a given address specified in _rs1_, and/or allowing software to use an ABI that relies on machine-mode cache-maintenance operations. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ![svg](_images/svg-9bea6e98625e14ac6492d8dc82c67940810c8daa.svg) The FENCE.I instruction is used to synchronize the instruction and data streams. RISC-V does not guarantee that stores to instruction memory will be made visible to instruction fetches on a RISC-V hart until that hart executes a FENCE.I instruction. A FENCE.I instruction ensures that a subsequent instruction fetch on a RISC-V hart will see any previous data stores already visible to the same RISC-V hart. FENCE.I does _not_ensure that other RISC-V harts' instruction fetches will observe the local hart’s stores in a multiprocessor system. To make a store to instruction memory visible to all RISC-V harts, the writing hart also has to execute a data FENCE before requesting that all remote RISC-V harts execute a FENCE.I. A FENCE.I instruction orders all explicit memory accesses that precede the FENCE.I in program order before all instruction fetches that follow the FENCE.I in program order. | | In the following litmus test, for example, the outcome a0\=1, a1\=0 on the consumer hart is forbidden, assuming little-endian RV32IC harts: Initially, flag = 0\. Producer hart: Consumer hart: la t0, patch\_me la t2, flag li t1, 0x4585 lw a0, (t2) sh t1, (t0) # patch\_me := c.li a1, 1 fence.i fence w, w # order flag write patch\_me: la t0, flag c.li a1, 0 li t1, 1 sw t1, (t0) # flag := 1 Note that this example is only meant to illustrate the aforementioned ordering property. In a realistic producer-consumer code-generation scheme, the consumer would loop until flag becomes 1 before executing the FENCE.I instruction. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | An instruction fetch is always ordered before any explicit memory accesses that instruction gives rise to. The unused fields in the FENCE.I instruction, _funct12_, _rs1_, and_rd_, are reserved for finer-grain fences in future extensions. For forward compatibility, base implementations shall ignore these fields, and standard software shall zero these fields. | | Because FENCE.I only orders stores with a hart’s own instruction fetches, application code should only rely upon FENCE.I if the application thread will not be migrated to a different hart. The EEI can provide mechanisms for efficient multiprocessor instruction-stream synchronization. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 8.1. "Zihintntl" Extension for Non-Temporal Locality Hints, Version 1.0 ==================== ## [](#chap:zihintntl)8.1\. "Zihintntl" Extension for Non-Temporal Locality Hints, Version 1.0 The NTL instructions are HINTs that indicate that the explicit memory accesses of the immediately subsequent instruction (henceforth "target instruction") exhibit poor temporal locality of reference.The NTL instructions do not change architectural state, nor do they alter the architecturally visible effects of the target instruction. Four variants are provided: The NTL.P1 instruction indicates that the target instruction does not exhibit temporal locality within the capacity of the innermost level of private cache in the memory hierarchy. NTL.P1 is encoded as ADD _x0, x0, x2_. The NTL.PALL instruction indicates that the target instruction does not exhibit temporal locality within the capacity of any level of private cache in the memory hierarchy. NTL.PALL is encoded as ADD _x0, x0, x3_. The NTL.S1 instruction indicates that the target instruction does not exhibit temporal locality within the capacity of the innermost level of shared cache in the memory hierarchy. NTL.S1 is encoded as ADD _x0, x0, x4_. The NTL.ALL instruction indicates that the target instruction does not exhibit temporal locality within the capacity of any level of cache in the memory hierarchy. NTL.ALL is encoded as ADD _x0, x0, x5_. | | The NTL instructions can be used to avoid cache pollution when streaming data or traversing large data structures, or to reduce latency in producer-consumer interactions. A microarchitecture might use the NTL instructions to inform the cache replacement policy, or to decide which cache to allocate into, or to avoid cache allocation altogether. For example, NTL.P1 might indicate that an implementation should not allocate a line in a private L1 cache, but should allocate in L2 (whether private or shared). In another implementation, NTL.P1 might allocate the line in L1, but in the least-recently used state. NTL.ALL will typically inform implementations not to allocate anywhere in the cache hierarchy. Programmers should use NTL.ALL for accesses that have no exploitable temporal locality. Like any HINTs, these instructions may be freely ignored. Hence, although they are described in terms of cache-based memory hierarchies, they do not mandate the provision of caches. Some implementations might respect these HINTs for some memory accesses but not others: e.g., implementations that implement LR/SC by acquiring a cache line in the exclusive state in L1 might ignore NTL instructions on LR and SC, but might respect NTL instructions for AMOs and regular loads and stores. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | [Table 1](#ntl-portable) lists several software use cases and the recommended NTL variant that _portable_ software—i.e., software not tuned for any specific implementation’s memory hierarchy—should use in each case. __Table 1\. Recommended NTL variant for portable software to employ in various scenarios.__ | Scenario | Recommended NTL variant | | -------------------------------------------------------------- | ----------------------- | | Access to a working set between and in size | NTL.P1 | | Access to a working set between and in size | NTL.PALL | | Access to a working set greater than in size | NTL.S1 | | Access with no exploitable temporal locality (e.g., streaming) | NTL.ALL | | Access to a contended synchronization variable | NTL.PALL | | | The working-set sizes listed in [Table 1](#ntl-portable) are not meant to constrain implementers' cache-sizing decisions. Cache sizes will obviously vary between implementations, and so software writers should only take these working-set sizes as rough guidelines. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | [Table 2](#ntl) lists several sample memory hierarchies and recommends how each NTL variant maps onto each cache level. The table also recommends which NTL variant that implementation-tuned software should use to avoid allocating in a particular cache level. For example, for a system with a private L1 and a shared L2, it is recommended that NTL.P1 and NTL.PALL indicate that temporal locality cannot be exploited by the L1, and that NTL.S1 and NTL.ALL indicate that temporal locality cannot be exploited by the L2\. Furthermore, software tuned for such a system should use NTL.P1 to indicate a lack of temporal locality exploitable by the L1, or should use NTL.ALL indicate a lack of temporal locality exploitable by the L2. If the C or Zca extension is provided, compressed variants of these HINTs are also provided: C.NTL.P1 is encoded as C.ADD _x0, x2_; C.NTL.PALL is encoded as C.ADD _x0, x3_; C.NTL.S1 is encoded as C.ADD _x0, x4_; and C.NTL.ALL is encoded as C.ADD _x0, x5_. The NTL instructions affect all memory-access instructions except the cache-management instructions in the Zicbom extension. | | As of this writing, there are no other exceptions to this rule, and so the NTL instructions affect all memory-access instructions defined in the base ISAs and the A, F, D, Q, C, and V standard extensions, as well as those defined within the hypervisor extension in Volume II. The NTL instructions can affect cache-management operations other than those in the Zicbom extension. For example, NTL.PALL followed by CBO.ZERO might indicate that the line should be allocated in L3 and zeroed, but not allocated in L1 or L2. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 2\. Mapping of NTL variants to various memory hierarchies.__ | Memory hierarchy | Recommended mapping of NTLvariant to actual cache level | Recommended NTL variant forexplicit cache management | | | | | | | | ------------------------------ | ------------------------------------------------------- | ---------------------------------------------------- | --- | -- | --- | ---- | ----- | ---- | | P1 | PALL | S1 | ALL | L1 | L2 | L3 | L4/L5 | | | Common Scenarios | | | | | | | | | | No caches | \--- | none | | | | | | | | Private L1 only | L1 | L1 | L1 | L1 | ALL | \--- | \--- | \--- | | Private L1; shared L2 | L1 | L1 | L2 | L2 | P1 | ALL | \--- | \--- | | Private L1; shared L2/L3 | L1 | L1 | L2 | L3 | P1 | S1 | ALL | \--- | | Private L1/L2 | L1 | L2 | L2 | L2 | P1 | ALL | \--- | \--- | | Private L1/L2; shared L3 | L1 | L2 | L3 | L3 | P1 | PALL | ALL | \--- | | Private L1/L2; shared L3/L4 | L1 | L2 | L3 | L4 | P1 | PALL | S1 | ALL | | Uncommon Scenarios | | | | | | | | | | Private L1/L2/L3; shared L4 | L1 | L3 | L4 | L4 | P1 | P1 | PALL | ALL | | Private L1; shared L2/L3/L4 | L1 | L1 | L2 | L4 | P1 | S1 | ALL | ALL | | Private L1/L2; shared L3/L4/L5 | L1 | L2 | L3 | L5 | P1 | PALL | S1 | ALL | | Private L1/L2/L3; shared L4/L5 | L1 | L3 | L4 | L5 | P1 | P1 | PALL | ALL | When an NTL instruction is applied to a prefetch hint in the Zicbop extension, it indicates that a cache line should be prefetched into a cache that is _outer_ from the level specified by the NTL. | | For example, in a system with a private L1 and shared L2, NTL.P1 followed by PREFETCH.R might prefetch into L2 with read intent. To prefetch into the innermost level of cache, do not prefix the prefetch instruction with an NTL instruction. In some systems, NTL.ALL followed by a prefetch instruction might prefetch into a cache or prefetch buffer internal to a memory controller. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Software is discouraged from following an NTL instruction with an instruction that does not explicitly access memory. Nonadherence to this recommendation might reduce performance but otherwise has no architecturally visible effect. In the event that a trap is taken on the target instruction, implementations are discouraged from applying the NTL to the first instruction in the trap handler. Instead, implementations are recommended to ignore the HINT in this case. | | If an interrupt occurs between the execution of an NTL instruction and its target instruction, execution will normally resume at the target instruction. That the NTL instruction is not re-executed does not change the semantics of the program. Some implementations might prefer not to process the NTL instruction until the target instruction is seen (e.g., so that the NTL can be fused with the memory access it modifies). Such implementations might preferentially take the interrupt before the NTL, rather than between the NTL and the memory access. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | Since the NTL instructions are encoded as ADDs, they can be used within LR/SC loops without voiding the forward-progress guarantee. But, since using other loads and stores within an LR/SC loop _does_ void the forward-progress guarantee, the only reason to use an NTL within such a loop is to modify the LR or the SC. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 9.1. "Zihintpause" Extension for Pause Hint, Version 2.0 ==================== ## [](#zihintpause)9.1\. "Zihintpause" Extension for Pause Hint, Version 2.0 The PAUSE instruction is a HINT that indicates the current hart’s rate of instruction retirement should be temporarily reduced or paused. The duration of its effect must be bounded and may be zero. | | Software can use the PAUSE instruction to reduce energy consumption while executing spin-wait code sequences. Multithreaded cores might temporarily relinquish execution resources to other harts when PAUSE is executed. It is recommended that a PAUSE instruction generally be included in the code sequence for a spin-wait loop. The duration of a PAUSE instruction’s effect may vary significantly within and among implementations. In typical implementations this duration should be much less than the time to perform a context switch, probably more on the rough order of an on-chip cache miss latency or a cacheless access to main memory. A series of PAUSE instructions can be used to create a cumulative delay loosely proportional to the number of PAUSE instructions. In spin-wait loops in portable code, however, only one PAUSE instruction should be used before re-evaluating loop conditions, else the hart might stall longer than optimal on some implementations, degrading system performance. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | PAUSE is encoded as a FENCE instruction with _pred_\=`W`, _succ_\=`0`, _fm_\=`0`,_rd_\=`x0`, and _rs1_\=`x0`. | | PAUSE is encoded as a hint within the FENCE opcode because some implementations are expected to deliberately stall the PAUSE instruction until outstanding memory transactions have completed. Because the successor set is null, however, PAUSE does not _mandate_ any particular memory ordering—hence, it truly is a HINT. Like other FENCE instructions, PAUSE cannot be used within LR/SC sequences without voiding the forward-progress guarantee. The choice of a predecessor set of W is arbitrary, since the successor set is null. Other HINTs similar to PAUSE might be encoded with other predecessor sets. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 35.1. "Zilsd", "Zclsd" Extensions for Load/Store pair for RV32, Version 1.0 ==================== ## [](#sec:zilsd)35.1\. "Zilsd", "Zclsd" Extensions for Load/Store pair for RV32, Version 1.0 The Zilsd & Zclsd extensions provide load/store pair instructions for RV32, reusing the existing RV64 doubleword load/store instruction encodings. Operands containing `src` for store instructions and `dest` for load instructions are held in aligned `x`\-register pairs, i.e., register numbers must be even. Use of misaligned (odd-numbered) registers for these operands is _reserved_. Regardless of endianness, the lower-numbered register holds the low-order bits, and the higher-numbered register holds the high-order bits: e.g., bits 31:0 of an operand in Zilsd might be held in register `x14`, with bits 63:32 of that operand held in `x15`. ### [](#zilsd)35.1.1\. Load/Store pair instructions (Zilsd) The Zilsd extension adds the following RV32-only instructions: | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | ------------------- | ----------------------------------------------------------------- | | yes | no | ld rd, offset(rs1) | [Load doubleword to register pair, 32-bit encoding](#insns-ld) | | yes | no | sd rs2, offset(rs1) | [Store doubleword from register pair, 32-bit encoding](#insns-sd) | As the access size is 64-bit, accesses are only considered naturally aligned for effective addresses that are a multiple of 8\. In this case, these instructions are guaranteed to not raise an address-misaligned exception. Even if naturally aligned, the memory access might not be performed atomically. If the effective address is a multiple of 4, then each word access is required to be performed atomically. The following table summarizes the required behavior: | Alignment | Word accesses guaranteed atomic? | Can cause misaligned trap? | | ----------- | -------------------------------- | -------------------------- | | 8 B | yes | no | | 4 B not 8 B | yes | yes | | else | no | yes | To ensure resumable trap handling is possible for the load instructions, the base register must have its original value if a trap is taken. The other register in the pair can have been updated. This affects x2 for the stack pointer relative instruction and rs1 otherwise. | | If an implementation performs a doubleword load access atomically and the register file implements write-back for even/odd register pairs, the mentioned atomicity requirements are inherently fulfilled. Otherwise, an implementation either needs to delay the write-back until the write can be performed atomically, or order sequential writes to the registers to ensure the requirement above is satisfied. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#zclsd)35.1.2\. Compressed Load/Store pair instructions (Zclsd) Zclsd depends on Zilsd and Zca. It has overlapping encodings with Zcf and is thus incompatible with Zcf. Zclsd adds the following RV32-only instructions: | RV32 | RV64 | Mnemonic | Instruction | | ---- | ---- | ----------------------- | ---------------------------------------------------------------------------------------- | | yes | no | c.ldsp rd, offset(sp) | [Stack-pointer based load doubleword to register pair, 16-bit encoding](#insns-cldsp) | | yes | no | c.sdsp rs2, offset(sp) | [Stack-pointer based store doubleword from register pair, 16-bit encoding](#insns-csdsp) | | yes | no | c.ld rd', offset(rs1') | [Load doubleword to register pair, 16-bit encoding](#insns-cld) | | yes | no | c.sd rs2', offset(rs1') | [Store doubleword from register pair, 16-bit encoding](#insns-csd) | ### [](#35-1-3-use-of-x0-as-operand)35.1.3\. Use of x0 as operand LD instructions with destination `x0` are processed as any other load, but the result is discarded entirely and x1 is not written. For C.LDSP, usage of `x0` as the destination is reserved. If using `x0` as `src` of SD or C.SDSP, the entire 64-bit operand is zero — i.e., register `x1` is not accessed. C.LD and C.SD instructions can only use `x8-15`. ### [](#35-1-4-exception-handling)35.1.4\. Exception Handling For the purposes of RVWMO and exception handling, LD and SD instructions are considered to be misaligned loads and stores, with one additional constraint:an LD or SD instruction whose effective address is a multiple of 4 gives rise to two 4-byte memory operations. | | This definition permits LD and SD instructions giving rise to exactly one memory access, regardless of alignment. If instructions with 4-byte-aligned effective address are decomposed into two 32b operations, there is no constraint on the order in which the operations are performed and each operation is guaranteed to be atomic. These decomposed sequences are interruptible. Exceptions might occur on subsequent operations, making the effects of previous operations within the same instruction visible. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Software should make no assumptions about the number or order of accesses these instructions might give rise to, beyond the 4-byte constraint mentioned above. For example, an interrupted store might overwrite the same bytes upon return from the interrupt handler. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#35-1-5-instructions)35.1.5\. Instructions #### [](#insns-ld)35.1.5.1\. ld Synopsis Load doubleword to even/odd register pair, 32-bit encoding Mnemonic ld rd, offset(rs1) Encoding (RV32) ![svg](_images/svg-2884039455a6d8a979405d136c99f1d8072f4a97.svg) Description Loads a 64-bit value into registers `rd` and `rd+1`. The effective address is obtained by adding register rs1 to the sign-extended 12-bit offset. Included in: [Zilsd](#zilsd) #### [](#insns-sd)35.1.5.2\. sd Synopsis Store doubleword from even/odd register pair, 32-bit encoding Mnemonic sd rs2, offset(rs1) Encoding (RV32) ![svg](_images/svg-5528737016f07cba11ee03d56db941f8d23ba997.svg) Description Stores a 64-bit value from registers `rs2` and `rs2+1`. The effective address is obtained by adding register rs1 to the sign-extended 12-bit offset. Included in: [Zilsd](#zilsd) #### [](#insns-cldsp)35.1.5.3\. c.ldsp Synopsis Stack-pointer based load doubleword to even/odd register pair, 16-bit encoding Mnemonic c.ldsp rd, offset(sp) Encoding (RV32) ![svg](_images/svg-bfdc112f3f8ca136c2f431834a6147b5f48db604.svg) Description Loads stack-pointer relative 64-bit value into registers `rd'` and `rd'+1`. It computes its effective address by adding the zero-extended offset, scaled by 8, to the stack pointer, `x2`. It expands to `ld rd, offset(x2)`. C.LDSP is only valid when _rd_≠x0; the code points with _rd_\=x0 are reserved. Included in: [Zclsd](#zclsd) #### [](#insns-csdsp)35.1.5.4\. c.sdsp Synopsis Stack-pointer based store doubleword from even/odd register pair, 16-bit encoding Mnemonic c.sdsp rs2, offset(sp) Encoding (RV32) ![svg](_images/svg-b521e6a813277bc26c3d467c98e3836828c16d07.svg) Description Stores a stack-pointer relative 64-bit value from registers `rs2'` and `rs2'+1`. It computes an effective address by adding the _zero_\-extended offset, scaled by 8, to the stack pointer, `x2`. It expands to `sd rs2, offset(x2)`. Included in: [Zclsd](#zclsd) #### [](#insns-cld)35.1.5.5\. c.ld Synopsis Load doubleword to even/odd register pair, 16-bit encoding Mnemonic c.ld rd', offset(rs1') Encoding (RV32) ![svg](_images/svg-0c754ade6278c1efefcf099b81e0e39b621cacc0.svg) Description Loads a 64-bit value into registers `rd'` and `rd'+1`. It computes an effective address by adding the zero-extended offset, scaled by 8, to the base address in register rs1'. Included in: [Zclsd](#zclsd) #### [](#insns-csd)35.1.5.6\. c.sd Synopsis Store doubleword from even/odd register pair, 16-bit encoding Mnemonic c.sd rs2', offset(rs1') Encoding (RV32) ![svg](_images/svg-127d4912863eaaf23fe0704b90a8c6c41208cf8d.svg) Description Stores a 64-bit value from registers `rs2'` and `rs2'+1`. It computes an effective address by adding the zero-extended offset, scaled by 8, to the base address in register rs1'. It expands to `sd rs2', offset(rs1')`. Included in: [Zclsd](#zclsd) 10.1. "Zimop" Extension for May-Be-Operations, Version 1.0 ==================== ## [](#zimop)10.1\. "Zimop" Extension for May-Be-Operations, Version 1.0 This chapter defines the "Zimop" extension, which introduces the concept of instructions that _may be operations_ (MOPs). MOPs are initially defined to simply write zero to `x[rd]`, but are designed to be redefined by later extensions to perform some other action. The Zimop extension defines an encoding space for 40 MOPs. | | It is sometimes desirable to define instruction-set extensions whose instructions, rather than raising illegal-instruction exceptions when the extension is not implemented, take no useful action (beyond writing x\[rd\]). For example, programs with control-flow integrity checks can execute correctly on implementations without the corresponding extension, provided the checks are simply ignored. Implementing these checks as MOPs allows the same programs to run on implementations with or without the corresponding extension. Although similar in some respects to HINTs, MOPs cannot be encoded as HINTs, because unlike HINTs, MOPs are allowed to alter architectural state. Because MOPs may be redefined by later extensions, standard software should not execute a MOP unless it is deliberately targeting an extension that has redefined that MOP. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The Zimop extension defines 32 MOP instructions named MOP.R._n_, where_n_ is an integer between 0 and 31, inclusive. Unless redefined by another extension, these instructions simply write 0 to`x[rd]`. Their encoding allows future extensions to define them to read `x[rs1]`, as well as write `x[rd]`. ![svg](_images/svg-7f1803709ce827f61f63daf6a127000e08423c3a.svg) The Zimop extension additionally defines 8 MOP instructions named MOP.RR._n_, where _n_ is an integer between 0 and 7, inclusive. Unless redefined by another extension, these instructions simply write 0 to `x[rd]`. Their encoding allows future extensions to define them to read `x[rs1]` and `x[rs2]`, as well as write `x[rd]`. ![svg](_images/svg-6e98e6a222e99907046cd3cf8aa19c7a7f1fced9.svg) | | The recommended assembly syntax for MOP.R._n_ is MOP.R._n_ rd, rs1, with any x\-register specifier being valid for either argument. Similarly for MOP.RR._n_, the recommended syntax is MOP.RR._n_ rd, rs1, rs2\. The extension that redefines a MOP may define an alternate assembly mnemonic. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | These MOPs are encoded in the SYSTEM major opcode in part because it is expected their behavior will be modulated by privileged CSR state. | | --------------------------------------------------------------------------------------------------------------------------------------------- | | | These MOPs are defined to write zero to x\[rd\], rather than performing no operation, to simplify instruction decoding and to allow testing the presence of features by branching on the zeroness of the result. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The MOPs defined in the Zimop extension do not carry a syntactic dependency from `x[rs1]` or `x[rs2]` to `x[rd]`, though an extension that redefines the MOP may impose such a requirement. | | Not carrying a syntactic dependency relieves straightforward implementations of reading x\[rs1\] and x\[rs2\]. | | ----------------------------------------------------------------------------------------------------------------- | ### [](#10-1-1-zcmop-compressed-may-be-operations-extension-version-1-0)10.1.1\. "Zcmop" Compressed May-Be-Operations Extension, Version 1.0 This section defines the "Zcmop" extension, which defines eight 16-bit MOP instructions named C.MOP._n_, where _n_ is an odd integer between 1 and 15, inclusive. C.MOP._n_ is encoded in the reserved encoding space corresponding to C.LUI x_n_, 0, as shown in [Table 1](#norm:c-mop%5Fenc). Unlike the MOPs defined in the Zimop extension, the C.MOP._n_ instructions are defined to _not_ write any register.Their encoding allows future extensions to define them to read register`x[_n_]`. The Zcmop extension depends upon the Zca extension. ![svg](_images/svg-5b86065f01ee794eb2fc368880d5d94298921b48.svg) | | Very few suitable 16-bit encoding spaces exist. This space was chosen because it already has unusual behavior with respect to the rd/rs1field—​it encodes c.addi16sp when the field contains x2\--and is therefore of lower value for most purposes. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 1\. C.MOP._n_ instruction encoding.__ | Mnemonic | Encoding | Redefinable to read register | | -------- | ---------------- | ---------------------------- | | C.MOP.1 | 0110000010000001 | x1 | | C.MOP.3 | 0110000110000001 | x3 | | C.MOP.5 | 0110001010000001 | x5 | | C.MOP.7 | 0110001110000001 | x7 | | C.MOP.9 | 0110010010000001 | x9 | | C.MOP.11 | 0110010110000001 | x11 | | C.MOP.13 | 0110011010000001 | x13 | | C.MOP.15 | 0110011110000001 | x15 | | | The recommended assembly syntax for C.MOP._n_ is simply the nullary C.MOP._n_. The possibly accessed register is implicitly x_n_. | | ------------------------------------------------------------------------------------------------------------------------------------ | | | The expectation is that each Zcmop instruction is equivalent to some Zimop instruction, but the choice of expansion (if any) is left to the extension that redefines the MOP. Note, a Zcmop instruction that does not write a value can expand into a write to x0. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 19.1. "Ztso" Extension for Total Store Ordering, Version 1.0 ==================== ## [](#ztso)19.1\. "Ztso" Extension for Total Store Ordering, Version 1.0 This chapter defines the "Ztso" extension for the RISC-V Total Store Ordering (RVTSO) memory consistency model. RVTSO is defined as a delta from RVWMO, which is defined in [RVWMO Memory Consistency Model](rvwmo.html). | | _The Ztso extension is meant to facilitate the porting of code originally written for the x86 or SPARC architectures, both of which use TSO by default. It also supports implementations which inherently provide RVTSO behavior and want to expose that fact to software._ | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | RVTSO makes the following adjustments to RVWMO: * All load operations behave as if they have an acquire-RCpc annotation * All store operations behave as if they have a release-RCpc annotation. * All AMOs behave as if they have both acquire-RCsc and release-RCsc annotations. | | _These rules render all PPO rules except[4-7](rvwmo.html#overlapping-ordering) redundant. They also make redundant any non-I/O fences that do not have both PW and SR set. Finally, they also imply that no memory operation will be reordered past an AMO in either direction._ _In the context of RVTSO, as is the case for RVWMO, the storage ordering annotations are concisely and completely defined by PPO rules[5-7](rvwmo.html#overlapping-ordering). In both of these memory models, it is the [Load Value Axiom](rvwmo.html#ax-load) that allows a hart to forward a value from its store buffer to a subsequent (in program order) load—that is to say that stores can be forwarded locally before they are visible to other harts._ | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Additionally, if the Ztso extension is implemented, then vector memory instructions in the V extension and Zve family of extensions follow RVTSO at the instruction level.The Ztso extension does not strengthen the ordering of intra-instruction element accesses. In spite of the fact that Ztso adds no new instructions to the ISA, code written assuming RVTSO will not run correctly on implementations not supporting Ztso. Binaries compiled to run only under Ztso should indicate as such via a flag in the binary, so that platforms which do not implement Ztso can simply refuse to run them. 1.1. Changes since Public Review version 0.8 ==================== ## [](#1-1-changes-since-public-review-version-0-8)1.1\. Changes since Public Review version 0.8 * Clarified that profile name can be used as ISA base string * Renamed Ssptead to Svade * Fixed Ssu64xl to make supporting UXL=64 mandatory * Added section listing new extension names in profiles document * Added new extension name Sscounterenw * Removed outdated text on Zicntr/Zihpm ratification plan 4.1. Components of a Profile ==================== ## [](#4-1-components-of-a-profile)4.1\. Components of a Profile ### [](#4-1-1-profile-family)4.1.1\. Profile Family Every profile is a member of a _profile_ _family_. A profile family is a set of profiles that share the same base ISA but which vary in highest-supported privilege mode. The initial two types of family are: * generic unprivileged instructions (I) * application processors running rich operating systems (A) | | More profile families may be added over time. | | ------------------------------------------------ | A profile family may be updated no more than annually, and the release calendar year is treated as part of the profile family name. Each profile family is described in more detail below. ### [](#4-1-2-profile-privilege-mode)4.1.2\. Profile Privilege Mode RISC-V has a layered architecture supporting multiple privilege modes, and most RISC-V platforms support more than one privilege mode. Software is usually written assuming a particular privilege mode during execution. For example, application code is written assuming it will be run in user mode, and kernel code is written assuming it will be run in supervisor mode. | | Software can be run in a mode different than the one for which it was written. For example, privileged code using privileged ISA features can be run in a user-mode execution environment, but will then cause traps into the enclosing execution environment when privileged instructions are executed. This behavior might be exploited, for example, to emulate a privileged execution environment using a user-mode execution environment. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The profile for a privilege mode describes the ISA features for an execution environment that has the eponymous privilege mode as the most-privileged mode available, but also includes all supported lower-privilege modes. In general, available instructions vary by privilege mode, and the behavior of RISC-V instructions can depend on the current privilege mode. For example, an S-mode profile includes U-mode as well as S-mode and describes the behavior of instructions when running in different modes in an S-mode execution environment, such as how an `ecall` instruction in U-mode causes a contained trap into an S-mode handler whereas an `ecall` in S-mode causes a requested trap out to the execution environment. A profile may specify that certain conditions will cause a requested trap (such as an `ecall` made in the highest-supported privilege mode) or fatal trap to the enclosing execution environment. The profile does not specify the behavior of the enclosing execution environment in handling requested or fatal traps. | | In particular, a profile does not specify the set of ECALLs available in the outer execution environment. This should be documented in the appropriate binary interface to the outer execution environment (e.g., Linux user ABI, or RISC-V SEE). | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | In general, a profile can be implemented by an execution environment using any hardware or software technique that provides compatible functionality, including pure software emulation. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | A profile does not specify any invisible traps. | | In particular, a profile does not constrain how invisible traps to a more-privileged mode can be used to emulate profile features. | | ------------------------------------------------------------------------------------------------------------------------------------- | A more-privileged profile can always support running software to implement a less-privileged profile from the same profile family. For example, a platform supporting the S-mode profile can run a supervisor-mode operating system that provides user-mode execution environments supporting the U-mode profile. | | Instructions in a U-mode profile, which are all executed in user mode, have potentially different behaviors than instructions executed in user mode in an S-mode profile. For this reason, a U-mode profile cannot be considered a subset of an S-mode profile. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#4-1-3-profile-isa-features)4.1.3\. Profile ISA Features An architecture profile has a mandatory ratified base instruction set (RV32I or RV64I for the current profiles). The profile also includes ratified ISA extensions placed into two categories: 1. Mandatory 2. Optional As the name implies, _Mandatory_ _ISA_ _extensions_ are a required part of the profile. Implementations of the profile must provide these. The combination of the profile base ISA plus the mandatory ISA extensions are termed the profile _mandates_, and software using the profile can assume these always exist. The _Optional_ category (also known as _options_) contains extensions that may be added as options, and which are expected to be generally supported as options by the software ecosystem for this profile. | | The level of "support" for an Optional extension will likely vary greatly among different software components supporting a profile. Users would expect that software claiming compatibility with a profile would make use of any available supported options, but as a bare minimum software should not report errors or warnings when supported options are present in a system. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | An optional extension may comprise many individually named and ratified extensions but a profile option requires all constituent extensions are present. In particular, unless explicitly listed as a profile option, individual extensions are not by themselves a profile option even when required as part of a profile option. For example, the Zbkb extension is not by itself a profile option even though it is a required component of the Zkn option. | | Profile optional extensions are intended to capture the granularity at which the broad software ecosystem is expected to cope with combinations of extensions. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | All components of a ratified profile must themselves have been ratified. Platforms may provide a discovery mechanism to determine what optional extensions are present. Extensions that are not explicitly listed in the mandatory or optional categories are termed _non-profile_ extensions, and are not considered parts of the profile. Some non-profile extensions can be added to an implementation without conflicting with the mandatory or optional components of a profile. In this case, the implementation is still compatible with the profile even though additional non-profile extensions are present. Other non-profile extensions added to an implementation might alter or conflict with the behavior of the mandatory or optional extensions in a profile, in which case the implementation would not be compatible with the profile. | | Extensions that are released after a given profile is released are by definition non-profile extensions. For example, mandatory or optional profile extensions for a new profile might be prototyped as non-profile extensions on an earlier profile. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#4-1-4-profile-naming-convention)4.1.4\. Profile Naming Convention A profile name is a string comprised of, in order: 1. Prefix **RV** for RISC-V. 2. A specific profile family name string. Initially a single letter (**I**, **M**, or **A**), but later profiles may have longer family name strings. 3. A numeric string giving the first complete calendar year for which the profile is ratified, represented as number of years after year 2000, i.e., **20** for profiles built on specifications ratified during 2019\. The year string will be longer than two digits in the next century. 4. A privilege mode (**U**, **S**, **M**). Hypervisor support is treated as an option. 5. A base ISA XLEN specifier (**32**, **64**). The initial profiles based on specifications ratified in 2019 are: * RVI20U32 basic unprivileged instructions for RV32I * RVI20U64 basic unprivileged instructions for RV64I * RVA20U64, RVA20S64 64-bit application-processor profiles | | Profile names are embeddable into RISC-V ISA naming strings. This implies that there will be no standard ISA extension with a name that matches the profile naming convention. This allows tools that process the RISC-V ISA naming string to parse and/or process a combined string. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at. Copyright 2023 by RISC-V International. New ISA Extensions ==================== ## [](#new-isa-extensions)New ISA Extensions This profile specification introduces the following new extension names for existing features, but none require new features: * **Ziccif**: Main memory supports instruction fetch with atomicity requirement * **Ziccrse**: Main memory supports forward progress on LR/SC sequences * **Ziccamoa**: Main memory supports all atomics in A * **Zicclsm**: Main memory supports misaligned loads/stores * **Za64rs**: Reservation set size of 64 bytes * **Za128rs**: Reservation set size of 128 bytes * **Zic64b**: Cache block size isf 64 bytes * **Svbare**: Bare mode virtual-memory translation supported * **Svade**: Raise exceptions on improper A/D bits * **Ssccptr**: Main memory supports page table reads * **Sscounterenw**: Support writeable enables for any supported counter * **Sstvecd**: `stvec` supports Direct mode * **Sstvala**: `stval` provides all needed values * **Ssu64xl**: UXLEN=64 must be supported * **Ssstateen**: Supervisor-mode view of the state-enable extension * **Shcounterenw**: Support writeable enables for any supported counter * **Shvstvala**: `vstval` provides all needed values * **Shtvala**: `htval` provides all needed values * **Shvstvecd**: `vstvec` supports Direct mode * **Shvsatpa**: `vsatp` supports all modes supported by `satp` * **Shgatpa**: SvNNx4 mode supported for all modes supported by `satp`, as well as Bare Glossary of ISA Extensions ==================== ## [](#glossary-of-isa-extensions)Glossary of ISA Extensions The following unprivileged ISA extensions are defined in Volume I of the [RISC-V Instruction Set Manual](../../isa/v20260120/unpriv/unpriv-index.html). * M Extension for Integer Multiplication and Division * A Extension for Atomic Memory Operations * F Extension for Single-Precision Floating-Point * D Extension for Double-Precision Floating-Point * Q Extension for Quad-Precision Floating-Point * C Extension for Compressed Instructions * Zifencei Instruction-Fetch Synchronization Extension * Zicsr Extension for Control and Status Register Access * Zicntr Extension for Basic Performance Counters * Zihpm Extension for Hardware Performance Counters * Zihintpause Pause Hint Extension * Zfh Extension for Half-Precision Floating-Point * Zfhmin Minimal Extension for Half-Precision Floating-Point * Zfinx Extension for Single-Precision Floating-Point in x-registers * Zdinx Extension for Double-Precision Floating-Point in x-registers * Zhinx Extension for Half-Precision Floating-Point in x-registers * Zhinxmin Minimal Extension for Half-Precision Floating-Point in x-registers The following privileged ISA extensions are defined in Volume II of the [RISC-V Instruction Set Manual](../../isa/v20260120/priv/priv-index.html). * Sv32 Page-based Virtual Memory Extension, 32-bit * Sv39 Page-based Virtual Memory Extension, 39-bit * Sv48 Page-based Virtual Memory Extension, 48-bit * Sv57 Page-based Virtual Memory Extension, 57-bit * Svpbmt, Page-Based Memory Types * Svnapot, NAPOT Translation Contiguity * Svinval, Fine-Grained Address-Translation Cache Invalidation * Hypervisor Extension * Sm1p11, Machine Architecture v1.11 * Sm1p12, Machine Architecture v1.12 * Ss1p11, Supervisor Architecture v1.11 * Ss1p12, Supervisor Architecture v1.12 The following extensions have not yet been incorporated into the RISC-V Instruction Set Manual; the hyperlinks lead to their separate specifications. * [Zba Address Computation Extension](https://github.com/riscv/riscv-bitmanip) * [Zbb Bit Manipulation Extension](https://github.com/riscv/riscv-bitmanip) * [Zbc Carryless Multiplication Extension](https://github.com/riscv/riscv-bitmanip) * [Zbs Single-Bit Manipulation Extension](https://github.com/riscv/riscv-bitmanip) * [Zbkb Extension for Bit Manipulation for Cryptography](https://github.com/riscv/riscv-crypto) * [Zbkc Extension for Carryless Multiplication for Cryptography](https://github.com/riscv/riscv-crypto) * [Zbkx Crossbar Permutation Extension](https://github.com/riscv/riscv-crypto) * [Zk Standard Scalar Cryptography Extension](https://github.com/riscv/riscv-crypto) * [Zkn NIST Cryptography Extension](https://github.com/riscv/riscv-crypto) * [Zknd AES Decryption Extension](https://github.com/riscv/riscv-crypto) * [Zkne AES Encryption Extension](https://github.com/riscv/riscv-crypto) * [Zknh SHA2 Hashing Extension](https://github.com/riscv/riscv-crypto) * [Zkr Entropy Source Extension](https://github.com/riscv/riscv-crypto) * [Zks ShangMi Cryptography Extension](https://github.com/riscv/riscv-crypto) * [Zksed SM4 Block Cypher Extension](https://github.com/riscv/riscv-crypto) * [Zksh SM3 Hashing Extension](https://github.com/riscv/riscv-crypto) * [Zkt Extension for Data-Independent Execution Latency](https://github.com/riscv/riscv-crypto) * [V Extension for Vector Computation](https://github.com/riscv/riscv-v-spec) * [Zve32x Extension for Embedded Vector Computation (32-bit integer)](https://github.com/riscv/riscv-v-spec) * [Zve32f Extension for Embedded Vector Computation (32-bit integer, 32-bit FP)](https://github.com/riscv/riscv-v-spec) * [Zve32d Extension for Embedded Vector Computation (32-bit integer, 64-bit FP)](https://github.com/riscv/riscv-v-spec) * [Zve64x Extension for Embedded Vector Computation (64-bit integer)](https://github.com/riscv/riscv-v-spec) * [Zve64f Extension for Embedded Vector Computation (64-bit integer, 32-bit FP)](https://github.com/riscv/riscv-v-spec) * [Zve64d Extension for Embedded Vector Computation (64-bit integer, 64-bit FP)](https://github.com/riscv/riscv-v-spec) * [Zicbom Extension for Cache-Block Management](https://github.com/riscv/riscv-CMOs) * [Zicbop Extension for Cache-Block Prefetching](https://github.com/riscv/riscv-CMOs) * [Zicboz Extension for Cache-Block Zeroing](https://github.com/riscv/riscv-CMOs) * [Sstc Extension for Supervisor-mode Timer Interrupts](https://github.com/riscv/riscv-time-compare) * [Sscofpmf Extension for Count Overflow and Mode-Based Filtering](https://github.com/riscv/riscv-count-overflow) * [Smstateen Extension for State-enable](https://github.com/riscv/riscv-state-enable) RISC-V Profiles ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-profiles)RISC-V Profiles Version 1.0, April 2, 2023: This document is in Ratified state. | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 2.1. Introduction. ==================== ## [](#2-1-introduction)2.1\. Introduction. RISC-V was designed to provide a highly modular and extensible instruction set, and includes a large and growing set of standard extensions. In addition, users may add their own custom extensions. This flexibility can be used to highly optimize a specialized design by including only the exact set of ISA features required for an application, but the same flexibility also leads to a combinatorial explosion in possible ISA choices. Profiles specify a much smaller common set of ISA choices that capture the most value for most users, and which thereby enable the software community to focus resources on building a rich software ecosystem with application and operating system portability across different implementations. | | Another pragmatic concern is the long and unwieldy ISA strings required to encode common sets of extensions, which will continue to grow as new extensions are defined. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Each profile is built on a standard base ISA plus a set of mandatory ISA extensions, and provides a small set of standard ISA options to extend the mandatory components. Profiles provide a convenient shorthand for describing the ISA portions of hardware and software platforms, and also guide the development of common software toolchains shared by different platforms that use the same profile. The intent is that the software ecosystem focus on supporting the profiles' mandatory base and standard options, instead of attempting to support every possible combination of individual extensions. Similarly, hardware vendors should aim to structure their offerings around standard profiles to increase the likelihood their designs will have mainstream software support. | | Profiles are not intended to prohibit the use of combinations of individual ISA extensions or the addition of custom extensions, which can continue to be used for more specialized applications albeit without the expectation of widespread software support or portability between hardware platforms. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | As RISC-V evolves over time, the set of ISA features will grow, and new platforms will be added that may need different profiles. To manage this evolution, RISC-V is adopting a model of regular annual releases of new ISA profiles, following an ISA roadmap managed by the RISC-V Technical Steering Committee. The architecture profiles will also be used for branding and to advertise compatibility with the RISC-V standard. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | This document describes the general structure of RISC-V architecture profiles and also the specifics of the first few profiles: RVI20 is a generic RISC-V unprivileged software profile, and RVA20 and RVA22 are architecture profiles for application processors. 3.1. Profiles versus Platforms ==================== ## [](#3-1-profiles-versus-platforms)3.1\. Profiles versus Platforms Profiles only describe ISA features, not a complete execution environment. A _software_ _platform_ is a specification for an execution environment, in which software targeted for that software platform can run. A _hardware_ _platform_ is a specification for a hardware system (which can be viewed as a physical realization of an execution environment). Both software and hardware platforms include specifications for many features beyond details of the ISA used by RISC-V harts in the platform (e.g., boot process, calling convention, behavior of environment calls, discovery mechanism, presence of certain memory-mapped hardware devices, etc.). Architecture profiles factor out ISA-specific definitions from platform definitions to allow ISA profiles to be reused across different platforms, and to be used by tools (e.g., compilers) that are common across many different platforms. A platform can add additional constraints on top of those in a profile. For example, mandating an extension that is a standard option in the underlying profile, or constraining some implementation-specific parameter in the profile to lie within a certain range. A platform cannot remove mandates or reduce other requirements in a profile. | | A new profile should be proposed if existing profiles do not match the needs of a new platform. | | -------------------------------------------------------------------------------------------------- | 6.1. RVA20 Profiles ==================== ## [](#6-1-rva20-profiles)6.1\. RVA20 Profiles The RVA20 profiles are intended to be used for 64-bit application processors running rich OS stacks. Only user-mode (RVA20U64) and supervisor-mode (RVA20S64) profiles are specified in this family. | | There is no machine-mode profile currently defined for application processor families. A machine-mode profile for application processors would only be used in specifying platforms for portable machine-mode software. Given the relatively low volume of portable M-mode software in this domain, the wide variety of potential M-mode code, and the very specific needs of each type of M-mode software, we are not specifying individual M-mode ISA requirements in the A-family profiles. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Only XLEN=64 application processor profiles are currently defined. It would be possible to also define very similar XLEN=32 variants. | | ---------------------------------------------------------------------------------------------------------------------------------------- | ### [](#6-1-1-rva20u64-profile)6.1.1\. RVA20U64 Profile The RVA20U64 profile specifies the ISA features available to user-mode execution environments in 64-bit applications processors. This is the most important profile within the application processor family in terms of the amount of software that targets this profile. RVA20U64 has one optional extension (Zihpm). #### [](#6-1-1-1-rva20u64-mandatory-base)6.1.1.1\. RVA20U64 Mandatory Base RV64I is the mandatory base ISA for RVA20U64, and is little-endian. As per the unprivileged architecture specification, the `ecall`instruction causes a requested trap to the execution environment. The `fence.tso` instruction is mandatory. | | The fence.tso instruction was incorrectly described as optional in the 2019 ratified specifications. However, fence.tso is encoded within the standard fence encoding such that implementations must treat it as a simple global fence if they do not natively support TSO-ordering optimizations. As software can always assume without any penalty that fence.tso is being exploited by a hardware implementation, there is no advantage to making the instruction a profile option. Later versions of the unprivileged ISA specifications correctly indicate that fence.tso is mandatory. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#6-1-1-2-rva20u64-mandatory-extensions)6.1.1.2\. RVA20U64 Mandatory Extensions * **M** Integer multiplication and division. * **A** Atomic instructions. * **F** Single-precision floating-point instructions. * **D** Double-precision floating-point instructions. * **C** Compressed Instructions. * **Zicsr** CSR instructions. These are implied by presence of Zicntr or F. * **Zicntr** Basic counters. * **Ziccif** Main memory regions with both the cacheability and coherence PMAs must support instruction fetch, and any instruction fetches of naturally aligned power-of-2 sizes up to min(ILEN,XLEN) (i.e., 32 bits for RVA20) are atomic. | | Ziccif is a new extension name capturing this feature. The fetch atomicity requirement facilitates runtime patching of aligned instructions. | | ----------------------------------------------------------------------------------------------------------------------------------------------- | * **Ziccrse** Main memory regions with both the cacheability and coherence PMAs must support RsrvEventual. | | Ziccrse is a new extension name capturing this feature. | | ---------------------------------------------------------- | * **Ziccamoa** Main memory regions with both the cacheability and coherence PMAs must support AMOArithmetic. | | Ziccamoa is a new extension name capturing this feature. | | ----------------------------------------------------------- | * **Za128rs** Reservation sets must be contiguous, naturally aligned, and at most 128 bytes in size. | | Za128rs is a new extension name capturing this feature. The minimum reservation set size is effectively determined by the size of atomic accesses in the A extension. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | * **Zicclsm** Misaligned loads and stores to main memory regions with both the cacheability and coherence PMAs must be supported. | | This introduces a new extension name for this feature. This requires misaligned support for all regular load and store instructions (including scalar and vector) but not AMOs or other specialized forms of memory access. Even though mandated, misaligned loads and stores might execute extremely slowly. Standard software distributions should assume their existence only for correctness, not for performance. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#6-1-1-3-rva20u64-optional-extensions)6.1.1.3\. RVA20U64 Optional Extensions * **Zihpm** Hardware performance counters. | | Hardware performance counters are a supported option in RVA20\. The number of counters is platform-specific. | | --------------------------------------------------------------------------------------------------------------- | | | The rationale to not make Q an optional extension is that quad-precision floating-point is unlikely to be implemented in hardware, and so we do not require or expect A-profile software to expend effort optimizing use of Q instructions in case they are present. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Zifencei is not classed as a supported option in the user-mode profile because it is not sufficient by itself to produce the desired effect in a multiprogrammed multiprocessor environment without OS support, and so the instruction cache flush should always be performed using an OS call rather than using the fence.i instruction.fence.i semantics can be expensive to implement for some hardware memory hierarchy designs, and so alternative non-standard instruction-cache coherence mechanisms can be used behind the OS abstraction. A separate extension is being developed for more general and efficient instruction cache coherence. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The execution environment must provide a means to synchronize writes to instruction memory with instruction fetches, the implementation of which likely relies on the Zifencei extension. For example, RISC-V Linux supplies the \_\_riscv\_flush\_icache system call and a corresponding vDSO call. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#6-1-1-4-rva20u64-recommendations)6.1.1.4\. RVA20U64 Recommendations Recommendations are not strictly mandated but are included to guide implementers making design choices. Implementations are strongly recommended to raise illegal-instruction exceptions on attempts to execute unimplemented opcodes. ### [](#6-1-2-rva20s64-profile)6.1.2\. RVA20S64 Profile The RVA20S64 profile specifies the ISA features available to a supervisor-mode execution environment in 64-bit applications processors. RVA20S64 is based on privileged architecture version 1.11. RVA20S64 has one unprivileged option (Zihpm) and one privileged option (Sv48). #### [](#6-1-2-1-rva20s64-mandatory-base)6.1.2.1\. RVA20S64 Mandatory Base RV64I is the mandatory base ISA for RVA20S64, and is little-endian. The `ecall` instruction operates as per the unprivileged architecture specification. An `ecall` in user mode causes a contained trap to supervisor mode. An `ecall` in supervisor mode causes a requested trap to the execution environment. #### [](#6-1-2-2-rva20s64-mandatory-extensions)6.1.2.2\. RVA20S64 Mandatory Extensions The following unprivileged extensions are mandatory: * The RVA20S64 mandatory unprivileged extensions include all the mandatory unprivileged extensions in RVA20U64. * **Zifencei** Instruction-Fetch Fence. | | Zifencei is mandated as it is the only standard way to support instruction-cache coherence in RVA20 application processors. A new instruction-cache coherence mechanism is under development which might be added as an option in the future. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The following privileged extensions are mandatory: * **Ss1p11** Privileged Architecture version 1.11. * **Svbare** The `satp` mode Bare must be supported. | | This is a new extension name for this feature. | | ------------------------------------------------- | * **Sv39** Page-Based 39-bit Virtual-Memory System. * **Svade** Page-fault exceptions are raised when a page is accessed when A bit is clear, or written when D bit is clear. | | This is a new extension name for this feature. | | ------------------------------------------------- | * **Ssccptr** Main memory regions with both the cacheability and coherence PMAs must support hardware page-table reads. | | This is a new extension name for this feature. | | ------------------------------------------------- | * **Sstvecd** `stvec.MODE` must be capable of holding the value 0 (Direct). When`stvec.MODE=Direct`, `stvec.BASE` must be capable of holding any valid four-byte-aligned address. | | This is a new extension name for this feature. | | ------------------------------------------------- | * **Sstvala** `stval` must be written with the faulting virtual address for load, store, and instruction page-fault, access-fault, and misaligned exceptions, and for breakpoint exceptions other than those caused by execution of the `ebreak` or `c.ebreak` instructions. For illegal-instruction exceptions, `stval` must be written with the faulting instruction. | | This is a new extension name for this feature. | | ------------------------------------------------- | #### [](#6-1-2-3-rva20s64-optional-extensions)6.1.2.3\. RVA20S64 Optional Extensions RVA20S64 has one unprivileged option. * **Zihpm** Hardware performance counters. | | The number of counters is platform-specific. | | ----------------------------------------------- | RVA20S64 has the following privileged options: * **Sv48** Page-Based 48-bit Virtual-Memory System. * **Ssu64xl** `sstatus.UXL` must be capable of holding the value 2 (i.e., UXLEN=64 must be supported). | | This is a new extension name for this feature. | | ------------------------------------------------- | 7.1. RVA22 Profiles ==================== ## [](#7-1-rva22-profiles)7.1\. RVA22 Profiles The RVA22 profiles are intended to be used for 64-bit application processors running rich OS stacks. Only user-mode (RVA22U64) and supervisor-mode (RVA22S64) profiles are specified in this family. ### [](#7-1-1-rva22u64-profile)7.1.1\. RVA22U64 Profile The RVA22U64 profile specifies the ISA features available to user-mode execution environments in 64-bit applications processors. This is the most important profile within the application processor family in terms of the amount of software that targets this profile. #### [](#7-1-1-1-rva22u64-mandatory-base)7.1.1.1\. RVA22U64 Mandatory Base RV64I is the mandatory base ISA for RVA22U64, including mandatory `fence.tso`, and is little-endian. | | Later versions of the RV64I unprivileged ISA specification ratified in 2021 made clear that fence.tso is mandatory. | | ---------------------------------------------------------------------------------------------------------------------- | As per the unprivileged architecture specification, the `ecall`instruction causes a requested trap to the execution environment. #### [](#7-1-1-2-rva22u64-mandatory-extensions)7.1.1.2\. RVA22U64 Mandatory Extensions The following mandatory extensions were present in RVA20U64. * **M** Integer multiplication and division. * **A** Atomic instructions. * **F** Single-precision floating-point instructions. * **D** Double-precision floating-point instructions. * **C** Compressed Instructions. * **Zicsr** CSR instructions. These are implied by presence of F. * **Zicntr** Base counters and timers. * **Zihpm** Hardware performance counters. * **Ziccif** Main memory regions with both the cacheability and coherence PMAs must support instruction fetch, and any instruction fetches of naturally aligned power-of-2 sizes up to min(ILEN,XLEN) (i.e., 32 bits for RVA22) are atomic. * **Ziccrse** Main memory regions with both the cacheability and coherence PMAs must support RsrvEventual. | | Ziccrse is a new extension name capturing this feature. | | ---------------------------------------------------------- | * **Ziccamoa** Main memory regions with both the cacheability and coherence PMAs must support AMOArithmetic. | | Ziccamoa is a new extension name capturing this feature. | | ----------------------------------------------------------- | * **Zicclsm** Misaligned loads and stores to main memory regions with both the cacheability and coherence PMAs must be supported. | | This is a new extension name for this feature. Even though mandated, misaligned loads and stores might execute extremely slowly. Standard software distributions should assume their existence only for correctness, not for performance. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The following mandatory feature was further restricted in RVA22U64: * **Za64rs** Reservation sets are contiguous, naturally aligned, and a maximum of 64 bytes. | | This is a new extension name capturing this feature. The maximum reservation size has been reduced to match the required cache block size. The minimum reservation size is effectively set by the instructions in the mandatory A extension. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The following mandatory extensions are new for RVA22U64. * **Zihintpause** Pause instruction. | | While the pause instruction is a HINT can be implemented as a NOP and hence trivially supported by hardware implementers, its inclusion in the mandatory extension list signifies that software should use the instruction whenever it would make sense and that implementors are expected to exploit this information to optimize hardware execution. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * **Zba** Address computation. * **Zbb** Basic bit manipulation. * **Zbs** Single-bit instructions. * **Zic64b** Cache blocks must be 64 bytes in size, naturally aligned in the address space. | | This is a new extension name for this feature. While the general RISC-V specifications are agnostic to cache block size, selecting a common cache block size simplifies the specification and use of the following cache-block extensions within the application processor profile. Software does not have to query a discovery mechanism and/or provide dynamic dispatch to the appropriate code. We choose 64 bytes at it is effectively an industry standard. Implementations may use longer cache blocks to reduce tag cost provided they use 64-byte sub-blocks to remain compatible. Implementations may use shorter cache blocks provided they sequence cache operations across the multiple cache blocks comprising a 64-byte block to remain compatible. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * **Zicbom** Cache-Block Management Operations. * **Zicbop** Cache-Block Prefetch Operations. | | As with other HINTS, the inclusion of prefetches in the mandatory set of extensions indicates that software should generate these instructions where they are expected to be useful, and hardware is expected to exploit that information. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * **Zicboz** Cache-Block Zero Operations. * **Zfhmin** Half-Precision Floating-point transfer and convert. | | Zfhmin is a small extension that adds support to load/store and convert IEEE FP16 numbers to and from IEEE FP32 format. The hardware cost for this extension is low, and mandating the extension avoids adding an option to the profile. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * **Zkt** Data-independent execution time. | | Zkt requires a certain subset of integer instructions execute with data-independent latency. Mandating this feature enables portable libraries for safe basic cryptographic operations. It is expected that application processors will naturally have this property and so implementation cost is low, if not zero, in most systems that would support RVA22. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#7-1-1-3-rva22u64-optional-extensions)7.1.1.3\. RVA22U64 Optional Extensions RVA22U64 has four profile options (Zfh, V, Zkn, Zks): * **Zfh** Half-Precision Floating-Point. | | A future profile might mandate Zfh. | | -------------------------------------- | * **V** Vector Extension. | | The smaller vector extensions (Zve32f, Zve32x, Zve64d, Zve64f, Zve64x) are not provided as separately supported profile options. The full V extension is specified as the only supported profile option. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | A future profile might mandate V. | | ------------------------------------ | * **Zkn** Scalar Crypto NIST Algorithms. * **Zks** Scalar Crypto ShangMi Algorithms. | | The scalar crypto extensions are expected to be superseded by vector crypto standards in future profiles, and the scalar extensions may be removed as supported options once vector crypto is present. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The smaller component scalar crypto extensions (Zbc, Zbkb, Zbkc, Zbkx, Zknd, Zkne, Zknh, Zksed, Zksh) are not provided as separate options in the profile. Profile implementers should provide all of the instructions in a given algorithm suite as part of the Zkn or Zks supported options. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Access to the entropy source (Zkr) in a system is usually carefully controlled. While the design supports unprivileged access to the entropy source, this is unlikely to be commonly used in an application processor, and so Zkr was not added as a profile option. This also means the roll-up Zk was not added as a profile option. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The Zfinx, Zdinx, Zhinx, Zhinxmin extensions are incompatible with the profile mandates to support the F and D extensions. | | ----------------------------------------------------------------------------------------------------------------------------- | #### [](#7-1-1-4-rva22u64-recommendations)7.1.1.4\. RVA22U64 Recommendations Recommendations are not strictly mandated but are included to guide implementers making design choices. Implementations are strongly recommended to raise illegal-instruction exceptions on attempts to execute unimplemented opcodes. ### [](#7-1-2-rva22s64-profile)7.1.2\. RVA22S64 Profile The RVA22S64 profile specifies the ISA features available to a supervisor-mode execution environment in 64-bit applications processors. RVA22S64 is based on privileged architecture version 1.12. #### [](#7-1-2-1-rva22s64-mandatory-base)7.1.2.1\. RVA22S64 Mandatory Base RV64I is the mandatory base ISA for RVA22S64, including mandatory`fence.tso`, and is little-endian. | | Later versions of the RV64I unprivileged ISA specification ratified in 2021 made clear that fence.tso is mandatory. | | ---------------------------------------------------------------------------------------------------------------------- | The `ecall` instruction operates as per the unprivileged architecture specification. An `ecall` in user mode causes a contained trap to supervisor mode. An `ecall` in supervisor mode causes a requested trap to the execution environment. #### [](#7-1-2-2-rva22s64-mandatory-extensions)7.1.2.2\. RVA22S64 Mandatory Extensions The following unprivileged extensions are mandatory: * The RVA22S64 mandatory unprivileged extensions include all the mandatory unprivileged extensions in RVA22U64. * **Zifencei** Instruction-Fetch Fence. | | Zifencei is mandated as it is the only standard way to support instruction-cache coherence in RVA22 application processors. A new instruction-cache coherence mechanism is under development which might be added as an option in the future. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The following privileged extensions are mandatory: * **Ss1p12** Privileged Architecture version 1.12. | | Ss1p12 supersedes Ss1p11. | | ---------------------------- | * **Svbare** The `satp` mode Bare must be supported. | | This is a new extension name for this feature. | | ------------------------------------------------- | * **Sv39** Page-Based 39-bit Virtual-Memory System. * **Svade** Page-fault exceptions are raised when a page is accessed when A bit is clear, or written when D bit is clear. * **Ssccptr** Main memory regions with both the cacheability and coherence PMAs must support hardware page-table reads. * **Sstvecd** `stvec.MODE` must be capable of holding the value 0 (Direct). When `stvec.MODE=Direct`, `stvec.BASE` must be capable of holding any valid four-byte-aligned address. * **Sstvala** stval must be written with the faulting virtual address for load, store, and instruction page-fault, access-fault, and misaligned exceptions, and for breakpoint exceptions other than those caused by execution of the EBREAK or C.EBREAK instructions. For illegal-instruction exceptions, stval must be written with the faulting instruction. * **Sscounterenw** For any hpmcounter that is not read-only zero, the corresponding bit in scounteren must be writable. | | This is new extension name capturing this feature. | | ----------------------------------------------------- | * **Svpbmt** Page-Based Memory Types * **Svinval** Fine-Grained Address-Translation Cache Invalidation #### [](#7-1-2-3-rva22s64-optional-extensions)7.1.2.3\. RVA22S64 Optional Extensions RVA22S64 has four unprivileged options (Zfh, V, Zkn, Zks) from RVA22U64, and eight privileged options (Sv48, Sv57, Svnapot, Ssu64xl, Sstc, Sscofpmf, Zkr, H). The privileged optional extensions are: * **Sv48** Page-Based 48-bit Virtual-Memory System. * **Sv57** Page-Based 57-bit Virtual-Memory System. * **Svnapot** NAPOT Translation Contiguity | | It is expected that Svnapot will be mandatory in the next profile release. | | ----------------------------------------------------------------------------- | * **Ssu64xl** `sstatus.UXL` must be capable of holding the value 2 (i.e., UXLEN=64 must be supported). | | This is a new extension name for this feature. | | ------------------------------------------------- | * **Sstc** supervisor-mode timer interrupts. | | Sstc was not made mandatory in RVA22S64 as it is a more disruptive change affecting system-level architecture, and will take longer for implementations to adopt. It is expected to be made mandatory in the next profile release. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * **Sscofpmf** Count Overflow and Mode-Based Filtering. | | Platforms may choose to mandate the presence of Sscofpmf. | | ------------------------------------------------------------ | * **Zkr** Entropy CSR. | | Technically, Zk is also a privileged-mode option capturing that Zkr, Zkn, and Zkt are all implemented. However, the Zk rollup is less descriptive than specifying the individual extensions explicitly. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * **H** The hypervisor extension. When the hypervisor extension is implemented, the following are also mandatory: * **Ssstateen** Supervisor-mode view of the state-enable extension. The supervisor-mode (`sstateen0-3`) and hypervisor-mode (`hstateen0-3`) state-enable registers must be provided. | | The Smstateen extension specification is an M-mode extension as it includes M-mode features, but the supervisor-mode visible components of the extension are named as the Ssstateen extension. Only Ssstateen is mandated in the RVA22S64 profile when the hypervisor extension is implemented. These registers are not mandated or supported options without the hypervisor extension, as there are no RVA22S64 supported options with relevant state to control in the absence of the hypervisor extension. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * **Shcounterenw** For any `hpmcounter` that is not read-only zero, the corresponding bit in `hcounteren` must be writable. | | This is a new extension name for this feature. | | ------------------------------------------------- | * **Shvstvala** `vstval` must be written in all cases described above for `stval`. | | This is a new extension name for this feature. | | ------------------------------------------------- | * **Shtvala** `htval` must be written with the faulting guest physical address in all circumstances permitted by the ISA. | | This is a new extension name for this feature. | | ------------------------------------------------- | * **Shvstvecd** `vstvec.MODE` must be capable of holding the value 0 (Direct). When `vstvec.MODE`\=Direct, `vstvec.BASE` must be capable of holding any valid four-byte-aligned address. | | This is a new extension name for this feature. | | ------------------------------------------------- | * **Shvsatpa** All translation modes supported in `satp` must be supported in `vsatp`. | | This is a new extension name for this feature. | | ------------------------------------------------- | * **Shgatpa** For each supported virtual memory scheme SvNN supported in`satp`, the corresponding hgatp SvNNx4 mode must be supported. The`hgatp` mode Bare must also be supported. | | This is a new extension name for this feature. | | ------------------------------------------------- | #### [](#7-1-2-4-rva22s64-recommendations)7.1.2.4\. RVA22S64 Recommendations * Implementations are strongly recommended to raise illegal-instruction exceptions when attempting to execute unimplemented opcodes. 5.1. RVI20 Profiles ==================== ## [](#5-1-rvi20-profiles)5.1\. RVI20 Profiles The RVI20 profiles document the initial set of unprivileged instructions. These provide a generic target for software toolchains and represent the minimum level of compatibility with RISC-V ratified standards. The two profiles RVI20U32 and RVI20U64 correspond to the RV32I and RV64I base ISAs respectively. | | These are designed as _unprivileged_ profiles as opposed to_user_\-_mode_ profiles. Code using this profile can run in any privilege mode, and so requested and fatal traps may be horizontal traps into an execution environment running in the same privilege mode. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#5-1-1-rvi20u32)5.1.1\. RVI20U32 RVI20U32 specifies the ISA features available to generic unprivileged execution environments. #### [](#5-1-1-1-rvi20u32-mandatory-base)5.1.1.1\. RVI20U32 Mandatory Base RV32I is the mandatory base ISA for RVI20U32, and is little-endian. As per the unprivileged architecture specification, the `ecall`instruction causes a requested trap to the execution environment. Misaligned loads and stores might not be supported. The `fence.tso` instruction is mandatory. | | The fence.tso instruction was incorrectly described as optional in the 2019 ratified specifications. However, fence.tso is encoded within the standard fence encoding such that implementations must treat it as a simple global fence if they do not natively support TSO-ordering optimizations. As software can always assume without any penalty that fence.tso is being exploited by a hardware implementation, there is no advantage to making the instruction an option. Later versions of the unprivileged ISA specifications correctly indicate that fence.tso is mandatory. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#5-1-1-2-rvi20u32-mandatory-extensions)5.1.1.2\. RVI20U32 Mandatory Extensions There are no mandatory extensions for RVI20U32. #### [](#5-1-1-3-rvi20u32-optional-extensions)5.1.1.3\. RVI20U32 Optional Extensions * **M** Integer multiplication and division. * **A** Atomic instructions. * **F** Single-precision floating-point instructions. * **D** Double-precision floating-point instructions. | | The rationale to not include Q as an optional extension is that quad-precision floating-point is unlikely to be implemented in hardware, and so we do not require or expect software to expend effort optimizing use of Q instructions in case they are present. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * **C** Compressed Instructions. * **Zifencei** Instruction-fetch fence instruction. * Misaligned loads and stores may be supported. * **Zicntr** Basic counters. | | The Zicsr extension is not supported independent of the Zicntr or F extensions. | | ---------------------------------------------------------------------------------- | * **Zihpm** Hardware performance counters. ### [](#5-1-2-rvi20u64)5.1.2\. RVI20U64 RVI20U64 specifies the ISA features available to generic unprivileged execution environments. #### [](#5-1-2-1-rvi20u64-mandatory-base)5.1.2.1\. RVI20U64 Mandatory Base RV64I is the mandatory base ISA for RVI20U64, and is little-endian. As per the unprivileged architecture specification, the `ecall`instruction causes a requested trap to the execution environment. Misaligned loads and stores might not be supported. The `fence.tso` instruction is mandatory. | | The fence.tso instruction was incorrectly described as optional in the 2019 ratified specifications. However, fence.tso is encoded within the standard fence encoding such that implementations must treat it as a simple global fence if they do not natively support TSO-ordering optimizations. As software can always assume without any penalty that fence.tso is being exploited by a hardware implementation, there is no advantage to making the instruction a profile option. Later versions of the unprivileged ISA specifications correctly indicate that fence.tso is mandatory. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#5-1-2-2-rvi20u64-mandatory-extensions)5.1.2.2\. RVI20U64 Mandatory Extensions There are no mandatory extensions for RVI20U64. #### [](#5-1-2-3-rvi20u64-optional-extensions)5.1.2.3\. RVI20U64 Optional Extensions * **M** Integer multiplication and division. * **A** Atomic instructions. * **F** Single-precision floating-point instructions. * **D** Double-precision floating-point instructions. | | The rationale to not include Q as a profile option is that quad-precision floating-point is unlikely to be implemented in hardware, and so we do not require or expect software to expend effort optimizing use of Q instructions in case they are present. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * **C** Compressed Instructions. * **Zifencei** Instruction-fetch fence instruction. * Misaligned loads and stores may be supported. * **Zicntr** Basic counters. | | The Zicsr extension is not supported independent of the Zicntr or F extensions. | | ---------------------------------------------------------------------------------- | * **Zihpm** Hardware performance counters. 1.1. Copyright and license information ==================== ## [](#1-1-copyright-and-license-information)1.1\. Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at. Copyright 2024 by RISC-V International. Glossary of ISA Extensions ==================== ## [](#glossary-of-isa-extensions)Glossary of ISA Extensions The following unprivileged ISA extensions are defined in Volume I of the h[RISC-V Instruction Set Manual](../../isa/v20260120/unpriv/unpriv-index.html). * M Extension for Integer Multiplication and Division * A Extension for Atomic Instructions * F Extension for Single-Precision Floating-Point * D Extension for Double-Precision Floating-Point * H Hypervisor Extension * Q Extension for Quad-Precision Floating-Point * C Extension for Compressed Instructions * B Extension for Bit Manipulation * V Extension for Vector Computation * Zifencei Instruction-Fetch Fence Extension * Zicsr Extension for Control and Status Register Access * Zicntr Extension for Basic Performance Counters * Zihpm Extension for Hardware Performance Counters * Zihintpause Pause Hint Extension * Zfh Extension for Half-Precision Floating-Point * Zfhmin Minimal Extension for Half-Precision Floating-Point * Zfinx Extension for Single-Precision Floating-Point in x-registers * Zdinx Extension for Double-Precision Floating-Point in x-registers * Zhinx Extension for Half-Precision Floating-Point in x-registers * Zhinxmin Minimal Extension for Half-Precision Floating-Point in x-registers * Zba Address Computation Extension * Zbb Bit Manipulation Extension * Zbc Carryless Multiplication Extension * Zbs Single-Bit Manipulation Extension * Zk Standard Scalar Cryptography Extension * Zkn NIST Cryptography Extension * Zknd AES Decryption Extension * Zkne AES Encryption Extension * Zknh SHA2 Hashing Extension * Zkr Entropy Source Extension * Zks ShangMi Cryptography Extension * Zksed SM4 Block Cypher Extension * Zksh SM3 Hashing Extension * Zkt Extension for Data-Independent Execution Latency * Zicbom Extension for Cache-Block Management * Zicbop Extension for Cache-Block Prefetching * Zicboz Extension for Cache-Block Zeroing * Zawrs Wait-on-reservation-set instructions * Zacas Extension for Atomic Compare-and-Swap (CAS) instructions * Zabha Extension for Byte and Halfword Atomic Memory Operations * Zbkb Extension for Bit Manipulation for Cryptography * Zbkc Extension for Carryless Multiplication for Cryptography * Zbkx Crossbar Permutation Extension * Zvbb - Vector Basic Bit-manipulation * Zvbc - Vector Carryless Multiplication * Zvkng - NIST Algorithm Suite with GCM * Zvksg - ShangMi Algorithm Suite with GCM * Zvkt - Vector Data-Independent Execution Latency The following privileged ISA extensions are defined in Volume II of the [RISC-V Instruction Set Manual](../../isa/v20260120/priv/priv-index.html). * Sv32 Page-based Virtual Memory Extension, 32-bit * Sv39 Page-based Virtual Memory Extension, 39-bit * Sv48 Page-based Virtual Memory Extension, 48-bit * Sv57 Page-based Virtual Memory Extension, 57-bit * Svpbmt, Page-Based Memory Types * Svnapot, NAPOT Translation Contiguity * Svinval, Fine-Grained Address-Translation Cache Invalidation * Hypervisor Extension * Sm1p11, Machine Architecture v1.11 * Sm1p12, Machine Architecture v1.12 * Ss1p11, Supervisor Architecture v1.11 * Ss1p12, Supervisor Architecture v1.12 * Ss1p13, Supervisor Architecture v1.13 * Sstc Extension for Supervisor-mode Timer Interrupts * Sscofpmf Extension for Count Overflow and Mode-Based Filtering * Smstateen/Ssstateen Extension for State-enable * Svvptc Obviating Memory-management Instructions after Marking PTEs valid * Svadu Hardware Updating of A/D Bits The following extensions have not yet been incorporated into the RISC-V Instruction Set Manual; the hyperlinks lead to their separate specifications. * [Zve32x Extension for Embedded Vector Computation (32-bit integer)](https://github.com/riscv/riscv-v-spec) * [Zve32f Extension for Embedded Vector Computation (32-bit integer, 32-bit FP)](https://github.com/riscv/riscv-v-spec) * [Zve32d Extension for Embedded Vector Computation (32-bit integer, 64-bit FP)](https://github.com/riscv/riscv-v-spec) * [Zve64x Extension for Embedded Vector Computation (64-bit integer)](https://github.com/riscv/riscv-v-spec) * [Zve64f Extension for Embedded Vector Computation (64-bit integer, 32-bit FP)](https://github.com/riscv/riscv-v-spec) * [Zve64d Extension for Embedded Vector Computation (64-bit integer, 64-bit FP)](https://github.com/riscv/riscv-v-spec) * **Ziccif**: Main memory supports instruction fetch with atomicity requirement * **Ziccrse**: Main memory supports forward progress on LR/SC sequences * **Ziccamoa**: Main memory supports all atomics in A * **Ziccamoc** Main memory supports atomics in Zacas * **Zicclsm**: Main memory supports misaligned loads/stores * **Zama16b**: Misaligned loads, stores, and AMOs to main memory regions that do not cross a naturally aligned 16-byte boundary are atomic. * **Za64rs**: Reservation set size of at most 64 bytes * **Za128rs**: Reservation set size of at most 128 bytes * **Zic64b**: Cache block size is 64 bytes * **Svbare**: Bare mode virtual-memory translation supported * **Svade**: Raise exceptions on improper A/D bits * **Ssccptr**: Main memory supports page table reads * **Sscounterenw**: Support writeable enables for any supported counter * **Sstvecd**: `stvec` supports Direct mode * **Sstvala**: `stval` provides all needed values * **Ssu64xl**: UXLEN=64 must be supported * **Sha**: Augmented hypervisor extension * **Shcounterenw**: Support writeable enables for any supported counter * **Shvstvala**: `vstval` provides all needed values * **Shtvala**: `htval` provides all needed values * **Shvstvecd**: `vstvec` supports Direct mode * **Shvsatpa**: `vsatp` supports all modes supported by `satp` * **Shgatpa**: SvNNx4 mode supported for all modes supported by `satp`, as well as Bare * **Ssstrict**: Unimplemented reserved encodings raise illegal instruction exceptions and no non-conforming extension are present RVA23 Profile ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#rva23-profile)RVA23 Profile Version 1.0, 2024-10-17: This document is in Ratified state. | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 3.1. Profile-Defined Extensions ==================== ## [](#3-1-profile-defined-extensions)3.1\. Profile-Defined Extensions This profile, as with earlier profiles, includes several new extensions defined directly in the profile text. These profile-defined extensions name optional features or combinations of features that are already present in ratified specifications, but that were not previously explicitly named. Once the profile is ratified, these extension definitions will move into the appropriate sections of the combined ISA manual. The combined ISA manual was not available at the start of this profile definition. Future profile proposals will be presented as an update to the combined ISA manual, with new profile-defined extensions provided as edits to the appropriate sections of the combined ISA manual. 2.1. RVA Profiles Rationale ==================== ## [](#2-1-rva-profiles-rationale)2.1\. RVA Profiles Rationale RISC-V was designed to provide a highly modular and extensible instruction set and includes a large and growing set of standard extensions, where each standard extension is a bundle of instruction-set features. This is no different than other industry ISAs that continue to add new ISA features. Unlike other ISAs, however, RISC-V has a broad set of contributors and implementers, and also allows users to add their own custom extensions. For some deep embedded markets, highly customized processor configurations are desirable for efficiency, and all software is compiled, ported, and/or developed in-house by the same organization for that specific processor configuration. However, for other markets that expect a substantial fraction of software to be delivered to end-customers in binary form, compatibility across multiple implementations from different RISC-V vendors is required. The RVIA ISA extension ratification process ensures that all processor vendors have agreed to the specification of a standard extension if present. However, by themselves, the ISA extension specifications do not guarantee that a certain set of standard extensions will be present in all implementations. **The primary goal of the RVA profiles is to align processor vendors targeting binary software markets, so software can rely on the existence of a certain set of ISA features in a particular generation of RISC-V implementations.** Alignment is not only for compatibility, but also to ensure RISC-V is competitive in these markets. The binary app markets are also generally those with the most competitive performance requirements (e.g., mobile, client, server). RVIA cannot mandate the ISA features that a RISC-V binary software ecosystem should use, as each ecosystem will typically select the lowest-common denominator they empirically observe in the deployed devices in their target markets. But RVIA can align hardware vendors to support a common set of features in each generation through the RVA profiles. Without proactive alignment through RVA profiles, RISC-V will be uncompetitive, as even if a particular vendor implements a certain feature, if other vendors do not, then binary distributions will not generally use that feature and all implementations will suffer. While certain features may be discoverable, and alternate code provided in case of presence/absence of a feature, the added cost to support such options is only justified for certain limited cases, and binary app markets will not support a wide range of optional features, particularly for the nascent RISC-V binary app ecosystems. To maintain alignment and increase RISC-V competitiveness over time, the mandatory set of extensions must increase over time in successive generations of RVA profile. (RVA profiles may eventually have to deprecate previously mandatory instructions, but that is unlikely in the near future.) Note that the RISC-V ISA will continue to evolve, regardless of whether a given software ecosystem settles on a certain generation of profile as the baseline for their ecosystem for many years or even decades. There are many existing binary software ecosystems, which will migrate to RISC-V and evolve at different rates, and more new ones will doubtless be created over the hopefully long lifetime of RISC-V. High-performance application processors require considerable investment, and no single binary app ecosystem can justify the development costs of these processors, especially for RISC-V in its early stage of adoption. While the heart of the profile is the set of mandatory extensions, there are several kinds of optional extension that serve important roles in the profile. The first kind are _localized_ _options_, whose presence or use necessarily differs along geo-political and/or jurisdictional boundaries, with crypto being the obvious example. These will always be optional. At least for crypto, discovery has been found to be perfectly acceptable to handle this optionality on other architectures, as the use of the extensions is well contained in certain libraries. The second kind of optional extension is a _development_ _option_, which represents a new ISA extension in an early part of its lifecycle but which is intended to become mandatory in a later generation of the RVA profile. Processor vendors and software toolchain providers will have varying development schedules, and providing an optional phase in a new extension’s lifecycle provides some flexibility while maintaining overall alignment, and is particularly appropriate when hardware or software development for the extension is complex. Denoting an extension as a _development_ _option_ signals to the community that development should be prioritized for such extensions as they will become mandatory. The third kind of optional extension are _expansion_ _options_, which are those that may have a large implementation cost but are not always needed in a particular platform, and which can be readily handled by discovery. These are also intended to remain available as expansion options in future versions of the profile. Several supervisor-mode extensions fall into this category, e.g., Sv57, which has a notable PPA impact over Sv48 and is not needed on smaller platforms. Some unprivileged extensions that may fall into this category are possible future matrix extensions. These have large implementation costs, and use of matrix instructions can be readily supported with discovery and alternate math libraries. The fourth kind of optional extensions are _transitory_ _options_, where it is not clear if the extension will change to a mandatory, localized, or expansion option, or be possibly dropped over time. Cryptography provides some examples where earlier cyphers have been broken and are now deprecated. RVIA used this mechanism to enable scalar crypto until vector crypto was ready. Software security features may also be in this category, with examples of deprecated security features occuring in other architectures. As another example, the recent avalanche of new numeric datatypes for AI/ML may eventually subside with a few survivors actually being used longer term. Denoting an option as transitory signals to the community that this extension may be removed in a future profile, though the time scale may span many years. Except for the localized options, it could be argued that other three kinds of option could be left out of profiles. Binary distributions of applications willing to invest in discovery can use an optional extension, and customers compiling their own applications can take advantage of the feature on a particular implementation, even when that system is mostly running binary distributions that ignore the new extension. However, there is value in providing guidance to align hardware vendors and software developers around what extensions are worth implementing and worth discovering, by designating only a few important features as profile options and limiting their granularity. 4.1. RVA23 Profiles ==================== ## [](#4-1-rva23-profiles)4.1\. RVA23 Profiles The RVA23 profiles are intended to align implementations of RISC-V 64-bit application processors to allow binary software ecosystems to rely on a large set of guaranteed extensions and a small number of discoverable coarse-grain options. It is explicitly a non-goal of RVA23 to allow more hardware implementation flexibility by supporting only a minimal set of features and a large number of fine-grain extensions. Only user-mode (RVA23U64) and supervisor-mode (RVA23S64) profiles are specified in this family. ### [](#4-1-1-rva23u64-profile)4.1.1\. RVA23U64 Profile The RVA23U64 profile specifies the ISA features available to user-mode execution environments in 64-bit applications processors. This is the most important profile within the application processor family in terms of the amount of software that targets this profile. #### [](#4-1-1-1-rva23u64-mandatory-base)4.1.1.1\. RVA23U64 Mandatory Base RV64I is the mandatory base ISA for RVA23U64 and is little-endian. As per the unprivileged architecture specification, the `ECALL`instruction causes a requested trap to the execution environment. #### [](#4-1-1-2-rva23u64-mandatory-extensions)4.1.1.2\. RVA23U64 Mandatory Extensions The following mandatory extensions were present in RVA22U64. * **M** Integer multiplication and division. * **A** Atomic instructions. * **F** Single-precision floating-point instructions. * **D** Double-precision floating-point instructions. * **C** Compressed instructions. * **B** Bit-manipulation instructions. * **Zicsr** CSR instructions. These are implied by presence of F. * **Zicntr** Base counters and timers. * **Zihpm** Hardware performance counters. * **Ziccif** Main memory regions with both the cacheability and coherence PMAs must support instruction fetch, and any instruction fetches of naturally aligned power-of-2 sizes up to min(ILEN,XLEN) (i.e., 32 bits for RVA23) are atomic. * **Ziccrse** Main memory regions with both the cacheability and coherence PMAs must support RsrvEventual. * **Ziccamoa** Main memory regions with both the cacheability and coherence PMAs must support all atomics in A. * **Zicclsm** Misaligned loads and stores to main memory regions with both the cacheability and coherence PMAs must be supported. * **Za64rs** Reservation sets are contiguous, naturally aligned, and a maximum of 64 bytes. * **Zihintpause** Pause hint. * **Zic64b** Cache blocks must be 64 bytes in size, naturally aligned in the address space. * **Zicbom** Cache-block management instructions. * **Zicbop** Cache-block prefetch instructions. * **Zicboz** Cache-Block Zero Instructions. * **Zfhmin** Half-precision floating-point. * **Zkt** Data-independent execution latency. The following mandatory extensions are new in RVA23U64: * **V** Vector extension. | | V was optional in RVA22U64. | | ------------------------------ | * **Zvfhmin** Vector minimal half-precision floating-point. * **Zvbb** Vector basic bit-manipulation instructions. * **Zvkt** Vector data-independent execution latency. * **Zihintntl** Non-temporal locality hints. * **Zicond** Integer conditional operations. * **Zimop** may-be-operations. * **Zcmop** Compressed may-be-operations. * **Zcb** Additional compressed instructions. * **Zfa** Additional floating-Point instructions. * **Zawrs** Wait-on-reservation-set instructions. * **Supm** Pointer masking, with the execution environment providing a means to select PMLEN=0 and PMLEN=7 at minimum. #### [](#4-1-1-3-rva23u64-optional-extensions)4.1.1.3\. RVA23U64 Optional Extensions ##### [](#4-1-1-3-1-localized-options)4.1.1.3.1\. Localized Options The following localized options are new in RVA23U64: * **Zvkng** Vector crypto NIST algorithms with GCM. * **Zvksg** Vector crypto ShangMi algorithms with GCM. | | The scalar crypto extensions Zkn and Zks that were options in RVA22 are not options in RVA23\. The goal is for both hardware and software vendors to move to use vector crypto, as vectors are now mandatory and vector crypto is substantially faster than scalar crypto. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | We have included only the Zvkng/Zvksg options with GCM to standardize on a higher performance crypto alternative. Zvbc is listed as a development option for use in other algorithms, and will become mandatory. Scalar Zbc is now listed as an expansion option, i.e., it will probably not become mandatory. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#4-1-1-3-2-development-options)4.1.1.3.2\. Development Options The following are new development options intended to become mandatory in a future RVA profile. * **Zabha** Byte and halfword atomic memory operations. * **Zacas** Compare-and-Swap instructions. * **Ziccamoc** Main memory regions with both the cacheability and coherence PMAs must provide `AMOCASQ` level PMA support. | | Ziccamoc is a new profile-defined extension that ensures Compare and Swap instructions are properly supported in main memory regions. The extension will be added to the PMA section of the privileged architecture manual. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | * **Zvbc** Vector carryless multiplication. * **Zama16b** Misaligned loads, stores, and AMOs to main memory regions that do not cross a naturally aligned 16-byte boundary are atomic. | | Zama16b is a new profile-defined extension that represents the presence of the new Misaligned Atomicity Granule feature added in Sm1p13\. The extension will be added to the PMA section of the privileged architecture manual. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#4-1-1-3-3-expansion-options)4.1.1.3.3\. Expansion Options The following expansion options were also present in RVA22U64: * **Zfh** Scalar half-precision floating-point. The following are new expansion options in RVA23U64: * **Zbc** Scalar carryless multiply. * **Zicfilp** Landing Pads. * **Zicfiss** Shadow Stack. * **Zvfh** Vector half-precision floating-point. * **Zfbfmin** Scalar BF16 converts. * **Zvfbfmin** Vector BF16 converts. * **Zvfbfwma** Vector BF16 widening mul-add. ##### [](#4-1-1-3-4-transitory-options)4.1.1.3.4\. Transitory Options There are no transitory options in RVA23U64. | | Scalar crypto is no longer an option in RVA23U64, though the Zbc extension has now been exposed as an expansion option. | | -------------------------------------------------------------------------------------------------------------------------- | #### [](#4-1-1-4-rva23u64-recommendations)4.1.1.4\. RVA23U64 Recommendations Implementations are strongly recommended to raise illegal-instruction exceptions on attempts to execute unimplemented opcodes. ### [](#4-1-2-rva23s64-profile)4.1.2\. RVA23S64 Profile The RVA23S64 profile specifies the ISA features available to a supervisor-mode execution environment in 64-bit applications processors. RVA23S64 is based on privileged architecture version 1.13. #### [](#4-1-2-1-rva23s64-mandatory-base)4.1.2.1\. RVA23S64 Mandatory Base RV64I is the mandatory base ISA for RVA23S64 and is little-endian. The `ECALL` instruction operates as per the unprivileged architecture specification. An `ECALL` in user mode causes a contained trap to supervisor mode. An `ECALL` in supervisor mode causes a requested trap to the execution environment. #### [](#4-1-2-2-rva23s64-mandatory-extensions)4.1.2.2\. RVA23S64 Mandatory Extensions The following unprivileged extensions are mandatory: * The RVA23S64 mandatory unprivileged extensions include all the mandatory unprivileged extensions in RVA23U64. * **Zifencei** Instruction-Fetch Fence. | | Zifencei is mandated as it is the only standard way to support instruction-cache coherence in RVA23 application processors. A new instruction-cache coherence mechanism is under development (tentatively named Zjid) which might be added as an option in the future. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The following privileged extensions are mandatory: * **Ss1p13** Supervisor architecture version 1.13. | | Ss1p13 supersedes Ss1p12. | | ---------------------------- | The following privileged extensions were also mandatory in RVA22S64: * **Svbare** The `satp` mode Bare must be supported. * **Sv39** Page-based 39-bit virtual-Memory system. * **Svade** Page-fault exceptions are raised when a page is accessed when A bit is clear, or written when D bit is clear. * **Ssccptr** Main memory regions with both the cacheability and coherence PMAs must support hardware page-table reads. * **Sstvecd** `stvec.MODE` must be capable of holding the value 0 (Direct). When `stvec.MODE=Direct`, `stvec.BASE` must be capable of holding any valid four-byte-aligned address. * **Sstvala** `stval` must be written with the faulting virtual address for load, store, and instruction page-fault, access-fault, and misaligned exceptions, and for breakpoint exceptions other than those caused by execution of the `EBREAK` or `C.EBREAK` instructions. For virtual-instruction and illegal-instruction exceptions, `stval` must be written with the faulting instruction. * **Sscounterenw** For any `hpmcounter` that is not read-only zero, the corresponding bit in `scounteren` must be writable. * **Svpbmt** Page-based memory types * **Svinval** Fine-grained address-translation cache invalidation. The following are new mandatory extensions: * **Svnapot** NAPOT translation contiguity. | | Svnapot was optional in RVA22. | | --------------------------------- | * **Sstc** supervisor-mode timer interrupts. | | Sstc was optional in RVA22. | | ------------------------------ | * **Sscofpmf** count overflow and mode-based filtering. * **Ssnpm** Pointer masking, with `senvcfg.PME` and `henvcfg.PME` supporting, at minimum, settings PMLEN=0 and PMLEN=7. * **Ssu64xl** `sstatus.UXL` must be capable of holding the value 2 (i.e., UXLEN=64 must be supported). | | Ssu64xl was optional in RVA22. | | --------------------------------- | * **Sha** The augmented hypervisor extension. | | Sha is a new profile-defined extension that captures the full set of features that are mandated to be supported along with the H extension. There is no change to the features added by including the hypervisor extension in a profile—​the new name is solely to simplify the text of the profiles. The definition has been added to the RVA22 profile text, where the hypervisor extension was first added, but will be added to the hypervisor section of the combined ISA manual. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | **Sha** comprises the following extensions: * **H** The hypervisor extension. * **Ssstateen** Supervisor-mode view of the state-enable extension. The supervisor-mode (`sstateen0-3`) and hypervisor-mode (`hstateen0-3`) state-enable registers must be provided. * **Shcounterenw** For any `hpmcounter` that is not read-only zero, the corresponding bit in `hcounteren` must be writable. * **Shvstvala** `vstval` must be written in all cases described above for `stval`. * **Shtvala** `htval` must be written with the faulting guest physical address in all circumstances permitted by the ISA. * **Shvstvecd** `vstvec.MODE` must be capable of holding the value 0 (Direct). When `vstvec.MODE`\=Direct, `vstvec.BASE` must be capable of holding any valid four-byte-aligned address. * **Shvsatpa** All translation modes supported in `satp` must be supported in `vsatp`. * **Shgatpa** For each supported virtual memory scheme SvNN supported in`satp`, the corresponding hgatp SvNNx4 mode must be supported. The`hgatp` mode Bare must also be supported. | | The augmented hypervisor extension (exactly equivalet to Sha) was optional in RVA22. | | --------------------------------------------------------------------------------------- | #### [](#4-1-2-3-rva23s64-optional-extensions)4.1.2.3\. RVA23S64 Optional Extensions ##### [](#4-1-2-3-1-localized-options)4.1.2.3.1\. Localized Options There are no privileged localized options in RVA23S64. ##### [](#4-1-2-3-2-development-options)4.1.2.3.2\. Development Options There are no privileged development options in RVA23S64. ##### [](#4-1-2-3-3-expansion-options)4.1.2.3.3\. Expansion Options The following privileged expansion options were present in RVA22S64: * **Sv48** Page-based 48-bit virtual-memory system. * **Sv57** Page-based 57-bit virtual-memory system. * **Zkr** Entropy CSR. The following are new privileged expansion options in RVA23S64 * **Svadu** Hardware A/D bit updates. * **Sdtrig** Debug triggers. * **Ssstrict** No non-conforming extensions are present. Attempts to execute unimplemented opcodes or access unimplemented CSRs in the standard or reserved encoding spaces raises an illegal instruction exception that results in a contained trap to the supervisor-mode trap handler. | | Ssstrict is a new profile-defined extension that restricts the behavior of reserved encoding spaces. The extension will be added to the supervisor chapter of the privileged architecture. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Ssstrict does not prescribe behavior for the custom encoding spaces or CSRs. | | ------------------------------------------------------------------------------- | | | Ssstrict definition applies to the execution environment claiming to be RVA23-compatible, which must have the hypervisor extension. That execution environment will take a contained trap to supervisor-mode (however that trap is implemented, including, but not limited to, emulation/delegation in the outer execution environment). Ssstrict (and all the other RVA23 mandates and options) do not apply to any guest VMs run by a hypervisor. An RVA23 hypervisor can provide guest VMs that are also RVA23-compatible but with an expanded set of emulated standard instructions. An RVA23 hypervisor can also choose to implement guest VMs that are not RVA23 compatible (e.g., lacking H, or only RVA20). | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * **Svvptc** Transitions from invalid to valid PTEs will be visible in bounded time without an explicit memory-management fence. * **Sspm** Supervisor-mode pointer masking, with the supervisor execution environment providing a means to select PMLEN=0 and PMLEN=7 at minimum. ##### [](#4-1-2-3-4-transitory-options)4.1.2.3.4\. Transitory Options There are no privileged transitory options in RVA23S64. #### [](#4-1-2-4-rva23s64-recommendations)4.1.2.4\. RVA23S64 Recommendations * Implementations are strongly recommended to raise illegal-instruction exceptions when attempting to execute unimplemented opcodes or access unimplemented CSRs. 1.1. Copyright and license information ==================== ## [](#1-1-copyright-and-license-information)1.1\. Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at. Copyright 2024 by RISC-V International. Glossary of ISA Extensions ==================== ## [](#glossary-of-isa-extensions)Glossary of ISA Extensions The following unprivileged ISA extensions are defined in Volume I of the [RISC-V Instruction Set Manual](../../isa/v20260120/unpriv/unpriv-index.html). * M Extension for Integer Multiplication and Division * A Extension for Atomic Memory Instructions * F Extension for Single-Precision Floating-Point * D Extension for Double-Precision Floating-Point * H Hypervisor Extension * Q Extension for Quad-Precision Floating-Point * C Extension for Compressed Instructions * B Extension for Bit Manipulation * V Extension for Vector Computation * Zifencei Instruction-Fetch Fence Extension * Zicsr Extension for Control and Status Register Access * Zicntr Extension for Basic Performance Counters * Zihpm Extension for Hardware Performance Counters * Zihintpause Pause Hint Extension * Zfh Extension for Half-Precision Floating-Point * Zfhmin Minimal Extension for Half-Precision Floating-Point * Zfinx Extension for Single-Precision Floating-Point in x-registers * Zdinx Extension for Double-Precision Floating-Point in x-registers * Zhinx Extension for Half-Precision Floating-Point in x-registers * Zhinxmin Minimal Extension for Half-Precision Floating-Point in x-registers * Zba Address Computation Extension * Zbb Bit Manipulation Extension * Zbc Carryless Multiplication Extension * Zbs Single-Bit Manipulation Extension * Zk Standard Scalar Cryptography Extension * Zkn NIST Cryptography Extension * Zknd AES Decryption Extension * Zkne AES Encryption Extension * Zknh SHA2 Hashing Extension * Zkr Entropy Source Extension * Zks ShangMi Cryptography Extension * Zksed SM4 Block Cypher Extension * Zksh SM3 Hashing Extension * Zkt Extension for Data-Independent Execution Latency * Zicbom Extension for Cache-Block Management * Zicbop Extension for Cache-Block Prefetching * Zicboz Extension for Cache-Block Zeroing * Zawrs Wait-on-reservation-set instructions * Zacas Extension for Atomic Compare-and-Swap (CAS) instructions * Zabha Extension for Byte and Halfword Atomic Memory Operations * Zbkb Extension for Bit Manipulation for Cryptography * Zbkc Extension for Carryless Multiplication for Cryptography * Zbkx Crossbar Permutation Extension * Zvbb - Vector Basic Bit-manipulation * Zvbc - Vector Carryless Multiplication * Zvkng - NIST Algorithm Suite with GCM * Zvksg - ShangMi Algorithm Suite with GCM * Zvkt - Vector Data-Independent Execution Latency The following privileged ISA extensions are defined in Volume II of the [RISC-V Instruction Set Manual](../../isa/v20260120/priv/priv-index.html) * Sv32 Page-based Virtual Memory Extension, 32-bit * Sv39 Page-based Virtual Memory Extension, 39-bit * Sv48 Page-based Virtual Memory Extension, 48-bit * Sv57 Page-based Virtual Memory Extension, 57-bit * Svpbmt, Page-Based Memory Types * Svnapot, NAPOT Translation Contiguity * Svinval, Fine-Grained Address-Translation Cache Invalidation * Hypervisor Extension * Sm1p11, Machine Architecture v1.11 * Sm1p12, Machine Architecture v1.12 * Ss1p11, Supervisor Architecture v1.11 * Ss1p12, Supervisor Architecture v1.12 * Ss1p13, Supervisor Architecture v1.13 * Sstc Extension for Supervisor-mode Timer Interrupts * Sscofpmf Extension for Count Overflow and Mode-Based Filtering * Smstateen/Ssstateen Extension for State-enable * Svvptc Obviating Memory-management Instructions after Marking PTEs valid * Svadu Hardware Updating of A/D Bits The following extensions have not yet been incorporated into the RISC-V Instruction Set Manual; the hyperlinks lead to their separate specifications. * [Zve32x Extension for Embedded Vector Computation (32-bit integer)](https://github.com/riscv/riscv-v-spec) * [Zve32f Extension for Embedded Vector Computation (32-bit integer, 32-bit FP)](https://github.com/riscv/riscv-v-spec) * [Zve32d Extension for Embedded Vector Computation (32-bit integer, 64-bit FP)](https://github.com/riscv/riscv-v-spec) * [Zve64x Extension for Embedded Vector Computation (64-bit integer)](https://github.com/riscv/riscv-v-spec) * [Zve64f Extension for Embedded Vector Computation (64-bit integer, 32-bit FP)](https://github.com/riscv/riscv-v-spec) * [Zve64d Extension for Embedded Vector Computation (64-bit integer, 64-bit FP)](https://github.com/riscv/riscv-v-spec) * **Ziccif**: Main memory supports instruction fetch with atomicity requirement * **Ziccrse**: Main memory supports forward progress on LR/SC sequences * **Ziccamoa**: Main memory supports all atomics in A * **Ziccamoc** Main memory supports atomics in Zacas * **Zicclsm**: Main memory supports misaligned loads/stores * **Zama16b**: Misaligned loads, stores, and AMOs to main memory regions that do not cross a naturally aligned 16-byte boundary are atomic. * **Za64rs**: Reservation set size of at most 64 bytes * **Za128rs**: Reservation set size of at most 128 bytes * **Zic64b**: Cache block size is 64 bytes * **Svbare**: Bare mode virtual-memory translation supported * **Svade**: Raise exceptions on improper A/D bits * **Ssccptr**: Main memory supports page table reads * **Sscounterenw**: Support writeable enables for any supported counter * **Sstvecd**: `stvec` supports Direct mode * **Sstvala**: `stval` provides all needed values * **Ssu64xl**: UXLEN=64 must be supported * **Sha**: Augmented hypervisor extension * **Shcounterenw**: Support writeable enables for any supported counter * **Shvstvala**: `vstval` provides all needed values * **Shtvala**: `htval` provides all needed values * **Shvstvecd**: `vstvec` supports Direct mode * **Shvsatpa**: `vsatp` supports all modes supported by `satp` * **Shgatpa**: SvNNx4 mode supported for all modes supported by `satp`, as well as Bare * **Ssstrict**: Unimplemented reserved encodings raise illegal instruction exceptions and no non-conforming extension are present RVB23 Profiles ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#rvb23-profiles)RVB23 Profiles Version 1.0, 2024-10-17 This document is in Ratified state. | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 2.1. Introduction ==================== ## [](#2-1-introduction)2.1\. Introduction This document specifies the RVB23 profile family. RVB23 is the first major release of the RVB series of RISC-V Application Processor Profile. RVB profiles are intended to be used for customized 64-bit application processors that will run rich OS stacks, but usually as a custom build of standard OS source-code distributions. The approach is to provide a large guaranteed set of relatively inexpensive and/or widely beneficial features but allow optionality for more expensive and/or more targeted extensions. Unlike the RVA profiles, it is explicitly a non-goal of RVB profiles to provide a single standard ISA interface supporting a wide variety of binary kernel and binary application software distributions. However, individual software ecosystems may build upon RVB profiles to produce a more targeted standard interface for a certain market. 3.1. Profile-Defined Extensions ==================== ## [](#3-1-profile-defined-extensions)3.1\. Profile-Defined Extensions RVB23 has been ratified alongside RVA23, and the same set of new profile-defined extensions defined in RVA23 are present in RVB23\. These profile-defined extensions will soon move to the combined ISA manual. Future releases of RVA and RVB profiles might not proceed through ratification at the same time, and future profile-defined extensions will be presented as an edit to the combined ISA manual. 4.1. RVB23 Profiles ==================== ## [](#4-1-rvb23-profiles)4.1\. RVB23 Profiles Only user-mode (RVB23U64) and supervisor-mode (RVB23S64) profiles are specified in this family. ### [](#4-1-1-rvb23u64-profile)4.1.1\. RVB23U64 Profile The RVB23U64 profile specifies the ISA features available to user-mode execution environments in 64-bit RVB applications processors. #### [](#4-1-1-1-rvb23u64-mandatory-base)4.1.1.1\. RVB23U64 Mandatory Base RV64I is the mandatory base ISA for RVB23U64 and is little-endian. As per the unprivileged architecture specification, the `ECALL`instruction causes a requested trap to the execution environment. #### [](#4-1-1-2-rvb23u64-mandatory-extensions)4.1.1.2\. RVB23U64 Mandatory Extensions The following mandatory extensions in RVB23U64 were also mandatory in RVA22U64. * **M** Integer multiplication and division. * **A** Atomic instructions. * **F** Single-precision floating-point instructions. * **D** Double-precision floating-point instructions. * **C** Compressed instructions. * **B** Bit-manipulation instructions. * **Zicsr** CSR instructions. These are implied by presence of F. * **Zicntr** Base counters and timers. * **Zihpm** Hardware performance counters. * **Ziccif** Main memory regions with both the cacheability and coherence PMAs must support instruction fetch, and any instruction fetches of naturally aligned power-of-2 sizes up to min(ILEN,XLEN) (i.e., 32 bits for RVB23) are atomic. * **Ziccrse** Main memory regions with both the cacheability and coherence PMAs must support RsrvEventual. * **Ziccamoa** Main memory regions with both the cacheability and coherence PMAs must support all atomics in A. * **Zicclsm** Misaligned loads and stores to main memory regions with both the cacheability and coherence PMAs must be supported. * **Za64rs** Reservation sets are contiguous, naturally aligned, and a maximum of 64 bytes. * **Zihintpause** Pause hint. * **Zic64b** Cache blocks must be 64 bytes in size, naturally aligned in the address space. * **Zicbom** Cache-block management instructions. * **Zicbop** Cache-block prefetch instructions. * **Zicboz** Cache-block zero instructions. * **Zkt** Data-independent execution latency. The following mandatory extensions are also present in RVA23U64: * **Zihintntl** Non-temporal locality hints. * **Zicond** Integer conditional operations. * **Zimop** May-be-operations. * **Zcmop** Compressed may-be-operations. * **Zcb** Additional compressed instructions. * **Zfa** Additional floating-point instructions. * **Zawrs** Wait-on-reservation-set instructions. #### [](#4-1-1-3-rvb23u64-optional-extensions)4.1.1.3\. RVB23U64 Optional Extensions RVB23U64 has 18 profile options listed below. ##### [](#4-1-1-3-1-localized-options)4.1.1.3.1\. Localized Options The following extensions are localized options in both RVA23U64 and RVB23U64: * **Zvkng** Vector crypto NIST Algorithms with GCM. * **Zvksg** Vector crypto ShangMi Algorithms with GCM. The following extensions options are localized options in RVB23U64 but are not present in RVA23U64: * **Zvkg** Vector GCM/GMAC instructions. * **Zvknc** Vector crypto NIST algorithms with carryless multiply. * **Zvksc** Vector crypto ShangMi algorithms with carryless multiply. | | RVA profiles mandate the higher-performing but more expensive GHASH options when adding vector crypto. To reduce implementation cost, RVB profiles also allow these carryless multiply options (Zvknc and Zvksc) to implement GCM efficiently, with GHASH available as a separate option. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * **Zkn** Scalar crypto NIST algorithms. * **Zks** Scalar crypto ShangMi algorithms. | | RVA23 profiles drop support for scalar crypto as an option, as the vector extension is now mandatory in RVA23\. RVB23 profiles support scalar crypto, as the vector extension is optional in RVB23. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ##### [](#4-1-1-3-2-development-options)4.1.1.3.2\. Development Options The following are new development options intended to become mandatory in a later RVB profile: * **Zabha** Byte and halfword atomic memory operations. * **Zacas** Compare-and-Swap instructions. * **Ziccamoc** Main memory regions with both the cacheability and coherence PMAs must provide `AMOCASQ` level PMA support. * **Zama16b** Misaligned loads, stores, and AMOs to main memory regions that do not cross a naturally aligned 16-byte boundary are atomic. ##### [](#4-1-1-3-3-expansion-options)4.1.1.3.3\. Expansion Options The following are expansion options in RVB23U64, but are mandatory in RVA23U64. * **Zfhmin** Half-precision floating-point. * **V** Vector extension. | | Unclear if other Zve\* extensions should also be supported in RVB. | | --------------------------------------------------------------------- | * **Zvfhmin** Vector minimal half-precision floating-point. * **Zvbb** Vector basic bit-manipulation instructions. * **Zvkt** Vector data-independent execution latency. * **Supm** Pointer masking, with the execution environment providing a means to select PMLEN=0 and PMLEN=7 at minimum. The following extensions are expansion options in both RVA23U64 and RVB23U64: * **Zfh** Scalar half-precision floating-point. * **Zbc** Scalar carryless multiplication. * **Zicfilp** Landing Pads. * **Zicfiss** Shadow Stack. * **Zvfh** Vector half-precision floating-point. * **Zfbfmin** Scalar BF16 converts. * **Zvfbfmin** Vector BF16 converts. * **Zvfbfwma** Vector BF16 widening mul-add. The following are expansion options for RVB23U64 as they are not intended to be made mandatory in future RVB profiles, but are listed as RVA23U64 development options as they are intended to become mandatory in future RVA profiles. * **Zvbc** Vector carryless multiplication. ##### [](#4-1-1-3-4-transitory-options)4.1.1.3.4\. Transitory Options There are no transitory options in RVB23U64. #### [](#4-1-1-4-rvb23u64-recommendations)4.1.1.4\. RVB23U64 Recommendations Implementations are strongly recommended to raise illegal-instruction exceptions on attempts to execute unimplemented opcodes. ### [](#4-1-2-rvb23s64-profile)4.1.2\. RVB23S64 Profile The RVB23S64 profile specifies the ISA features available to a supervisor-mode execution environment in 64-bit applications processors. RVB23S64 is based on privileged architecture version 1.13. | | Priv 1.13 is still being defined. | | ------------------------------------ | #### [](#4-1-2-1-rvb23s64-mandatory-base)4.1.2.1\. RVB23S64 Mandatory Base RV64I is the mandatory base ISA for RVB23S64 and is little-endian. The `ECALL` instruction operates as per the unprivileged architecture specification. An `ECALL` in user mode causes a contained trap to supervisor mode. An `ECALL` in supervisor mode causes a requested trap to the execution environment. #### [](#4-1-2-2-rvb23s64-mandatory-extensions)4.1.2.2\. RVB23S64 Mandatory Extensions The following unprivileged extensions are mandatory: * The RVB23S64 mandatory unprivileged extensions include all the mandatory unprivileged extensions in RVB23U64. * **Zifencei** Instruction-Fetch Fence. | | Zifencei is mandated as it is the only standard way to support instruction-cache coherence in RVB23 application processors. A new instruction-cache coherence mechanism is under development (tentatively named Zjid) which might be added as an option in the future. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The following privileged extensions are mandatory, and are also mandatory in RVA23S64. * **Ss1p13** Supervisor architecture version 1.13. | | Ss1p13 supersedes Ss1p12 but is not yet ratified. | | ---------------------------------------------------- | * **Svnapot** NAPOT translation contiguity. | | Svnapot is very low cost to provide, so is made mandatory even in RVB. | | ------------------------------------------------------------------------- | * **Svbare** The `satp` mode Bare must be supported. * **Sv39** Page-Based 39-bit Virtual-Memory System. * **Svade** Page-fault exceptions are raised when a page is accessed when A bit is clear, or written when D bit is clear. * **Ssccptr** Main memory regions with both the cacheability and coherence PMAs must support hardware page-table reads. * **Sstvecd** `stvec.MODE` must be capable of holding the value 0 (Direct). When `stvec.MODE=Direct`, `stvec.BASE` must be capable of holding any valid four-byte-aligned address. * **Sstvala** `stval` must be written with the faulting virtual address for load, store, and instruction page-fault, access-fault, and misaligned exceptions, and for breakpoint exceptions other than those caused by execution of the `EBREAK` or `C.EBREAK` instructions. For virtual-instruction and illegal-instruction exceptions, `stval` must be written with the faulting instruction. * **Sscounterenw** For any `hpmcounter` that is not read-only zero, the corresponding bit in `scounteren` must be writable. * **Svpbmt** Page-based memory types. * **Svinval** Fine-grained address-translation cache invalidation. * **Sstc** supervisor-mode timer interrupts. * **Sscofpmf** Count overflow and mode-based filtering. * **Ssu64xl** `sstatus.UXL` must be capable of holding the value 2 (i.e., UXLEN=64 must be supported). #### [](#4-1-2-3-rvb23s64-optional-extensions)4.1.2.3\. RVB23S64 Optional Extensions RVB23S64 has the same unprivileged options as RVB23U64, The privileged options in RVB23S64 are listed in the following sections. ##### [](#4-1-2-3-1-localized-options)4.1.2.3.1\. Localized Options There are no privileged localized options in RVB23S64. ##### [](#4-1-2-3-2-development-options)4.1.2.3.2\. Development Options There are no privileged development options in RVB23S64. ##### [](#4-1-2-3-3-expansion-options)4.1.2.3.3\. Expansion Options The following are privileged expansion options in RVB23S64, but are mandatory in RVA23S64: * **Ssnpm** Pointer masking, with `senvcfg.PME` supporting at minimum, settings PMLEN=0 and PMLEN=7. * **Sha** The augmented hypervisor extension. When the hypervisor extension is implemented, the following are also mandatory: * If the hypervisor extension is implemented and pointer masking (Ssnpm) is supported then `henvcfg.PME` must support at minimum, settings PMLEN=0 and PMLEN=7. The following are privileged expansion options in RVB23S64 that are also privileged expansion options in RVA23S64: * **Sv48** Page-based 48-bit virtual-memory system. * **Sv57** Page-based 57-bit virtual-memory system. * **Svadu** Hardware A/D bit updates. * **Zkr** Entropy CSR. * **Sdtrig** Debug triggers. * **Ssstrict** No non-conforming extensions are present. Attempts to execute unimplemented opcodes or access unimplemented CSRs in the standard or reserved encoding spaces raises an illegal instruction exception that results in a contained trap to the supervisor-mode trap handler. | | Ssstrict does not prescribe behavior for the custom encoding spaces or CSRs. | | ------------------------------------------------------------------------------- | | | Ssstrict definition applies to the execution environment claiming to be RVA23-compatible, which must have the hypervisor extension. That execution environment will take a contained trap to supervisor-mode (however that trap is implemented, including, but not limited to, emulation/delegation in the outer execution environment). Ssstrict (and all the other RVA23 mandates and options) do not apply to any guest VMs run by a hypervisor. An RVA23 hypervisor can provide guest VMs that are also RVA23-compatible but with an expanded set of emulated standard instructions. An RVA23 hypervisor can also choose to implement guest VMs that are not RVA23 compatible (e.g., lacking H, or only RVA20). | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | * **Svvptc** Transitions from invalid to valid PTEs will be visible in bounded time without an explicit memory-management fence. * **Sspm** Supervisor-mode pointer masking, with the supervisor execution environment providing a means to select PMLEN=0 and PMLEN=7 at minimum. #### [](#4-1-2-4-rvb23s64-recommendations)4.1.2.4\. RVB23S64 Recommendations * Implementations are strongly recommended to raise illegal-instruction exceptions when attempting to execute unimplemented opcodes. Contributors ==================== ## [](#contributors)Contributors Contributors to all versions of this specification in alphabetical order (please contact the editor to suggest corrections): Krste Asanović, Paul Donahue, Greg Favor, John Hauser, James Kenney, David Kruckemyer, Shubu Mukherjee, Stefan O’Rear, Vernon Pang, Anup Patel, Josh Scheid, Ved Shanbhogue, and Andrew Waterman. Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This document is released under a Creative Commons Attribution 4.0 International License. . The RISC-V Advanced Interrupt Architecture ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#the-risc-v-advanced-interrupt-architecture)The RISC-V Advanced Interrupt Architecture Editors: John Hauser Version 1.0, Revised 20250312 | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 4.1. Advanced Platform-Level Interrupt Controller (APLIC) ==================== ## [](#AdvPLIC)4.1\. Advanced Platform-Level Interrupt Controller (APLIC) In a RISC-V system, a Platform-Level Interrupt Controller (PLIC) handles external interrupts that are signaled through wires rather than by MSIs. When the RISC-V harts in a system do not have IMSICs, the harts themselves do not support MSIs, and all external interrupts to such harts must pass through a PLIC. But even in machines where harts have IMSICs and most interrupts are communicated via MSIs, it is not unusual for some device interrupts still to be signaled by dedicated wires. In particular, for devices (or device controllers) that do not otherwise need to initiate bus transactions in the system, the cost of supporting MSIs is especially high, so wired interrupts are a frugal alternative. Wired interrupts also continue to be universally supported by all current computer platforms, unlike MSIs, making another reason for many commodity devices or controllers to choose wired interrupts over MSIs, unless conforming to a standard like PCI Express that dictates MSIs. This chapter specifies an _Advanced PLIC_ (APLIC) that is not backward compatible with the earlier RISC-V PLIC. Full conformance to the Advanced Interrupt Architecture requires the APLIC. However, a workable system can be built substituting the older PLIC instead, assuming only wired interrupts to harts, not MSIs. | | We intend eventually to provide a free example parameterized implementation of an APLIC, written in portable SystemVerilog, that we expect will be suitable for many RISC-V systems without modification. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | A draft specification exists for a _Duo-PLIC_ that is software-configurable to act as either an original RISC-V PLIC or an APLIC. However, at this time, it appears unlikely that RISC-V International will ever ratify the Duo-PLIC specification as a standard. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | In a machine without IMSICs, every RISC-V hart accepts interrupts from exactly one PLIC or APLIC that is the _external interrupt controller_ for that hart. A hart’s external interrupt controller (the PLIC or APLIC) signals interrupts to the hart through a dedicated connection, usually a wire, for each privilege level that the hart may receive interrupts. (Recall[Figure 1\. Traditional delivery of wired interrupts to harts without support for MSIs.](intro.html#intrsWithoutIMSICs)). A system without IMSICs will typically have only one PLIC or APLIC, serving as the external interrupt controller for all RISC-V harts. | | Because every RISC-V hart without an IMSIC has exactly one PLIC or APLIC as its external interrupt controller, a system with multiple APLICs must partition the harts into disjoint subsets, making each APLIC the external interrupt controller for a separate subset of the harts. While not prohibited, this arrangement is likely to be less efficient than having all harts share a single APLIC. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | RISC-V harts that employ IMSICs as their external interrupt controllers can receive external interrupts only in the form of MSIs. In that case, the role of an APLIC is to convert wired interrupts into MSIs for harts. (Recall [Figure 2\. Interrupt delivery by MSIs when harts have IMSICs for receiving them](intro.html#intrsWithIMSICs).) The APLIC is said to _forward_ incoming wire-signaled interrupts to harts by sending MSIs to the harts. When harts have IMSICs to support MSIs, a system may easily contain multiple APLICs for converting wired interrupts into MSIs, with each APLIC forwarding interrupts from a different subset of devices. Multiple APLICs are presumably more likely to arise when groups of devices are physically distant from one another, perhaps even on separate chips (including chiplets in a multi-chip module). ### [](#4-1-1-interrupt-sources-and-identities)4.1.1\. Interrupt sources and identities An individual APLIC supports a fixed number of _interrupt sources_, corresponding exactly with the set of physical incoming interrupt wires at the APLIC. Most often, each source’s incoming wire is connected to the output interrupt wire from a single device or device controller. (For level-sensitive interrupts, the interrupt outputs of multiple devices or controllers may be combined to drive the incoming wire of a single interrupt source at an APLIC. An interrupt source’s incoming wire might also be simply tied high or low, if, for example, the source will always be configured as Detached. See[4.1.5.2\. Source configurations (sourcecfg\[1\]–sourcecfg\[1023\])](#AdvPLIC-reg-sourcecfg) for a description of _source modes_.) Each of an APLIC’s interrupt sources has a fixed unique _identity number_ in the range 1 to , where is the total number of sources at the APLIC. The number zero is not a valid interrupt identity number at an APLIC. The maximum number of interrupt sources an APLIC may support is 1023. When an APLIC delivers interrupts directly to harts at a given privilege level (rather than forwarding interrupts as MSIs), the APLIC is the external interrupt controller for the harts at that privilege level, and the interrupt identities at the APLIC become directly the _minor identities_ for external interrupts at the harts. On the other hand, when an APLIC forwards interrupts by MSIs, software configures a new interrupt identity number for the outgoing MSIs of each source. Consequently, in this case, the source identity numbers at a given APLIC only distinguish the incoming interrupts at the APLIC and have no relevance outside the APLIC. ### [](#4-1-2-interrupt-domains)4.1.2\. Interrupt domains An APLIC supports one or more _interrupt domains_, each associated with a subset of RISC-V harts at one privilege level (machine or supervisor level). The harts within an interrupt domain are those that the domain can interrupt at the corresponding privilege level. Each domain has its own memory-mapped control region in the machine’s address space that appears to control a complete, separate APLIC, though in fact all domain interfaces together access a single combined interrupt controller. [Figure 1](#AdvPLIC-ex-1Domain) through[Figure 3](#AdvPLIC-ex-3Domains) depict some possible hierarchies of interrupt domains implemented by an APLIC in a RISC-V system. The first figure represents a minimal system that has a single hart not supporting supervisor mode, with a single interrupt domain for machine level on that hart. The next figure, [Figure 2](#AdvPLIC-ex-2Domains), shows a basic arrangement for a larger system designed for symmetric multiprocessing (SMP), with multiple harts that all implement supervisor mode. In such cases, the APLIC will usually provide a separate interrupt domain for supervisor level, as the figure portrays. This supervisor-level interrupt domain allows an operating system, running in S-mode on the multiple harts, to have direct control over the interrupts it receives, avoiding the need to call upon M-mode to exercise that control. ![AdvPLIC ex 1Domain](_images/AdvPLIC-ex-1Domain.png) Figure 1\. Example of a RISC-V system that has a single hart implementing only M-mode, with a single machine-level interrupt domain for that hart. An APLIC’s interrupt domains are arranged in a tree hierarchy, with the root domain always being at machine level. Incoming interrupt wires arrive first at the root domain. Each domain may then selectively delegate all or a subset of interrupt sources to its child domains in the hierarchy. Within a given APLIC, interrupt source numbers are invariant across all domains, so source identity number always refers to the same source in every domain, corresponding to incoming wire number . For an interrupt domain below the root, interrupt sources not delegated down to that domain appear to the domain as being not implemented. [Figure 3](#AdvPLIC-ex-3Domains) shows a hierarchy of three interrupt domains, two at machine level and one at supervisor level. The arrangement in the figure, when combined with PMP (physical memory protection), allows machine-level software to isolate a selection of interrupts exclusively for hart 0, beyond the reach of the four application harts, even at machine level. ![AdvPLIC ex 2Domains](_images/AdvPLIC-ex-2Domains.png) Figure 2\. An example system with four harts that implement M-mode and S-mode, with two APLIC interrupt domains, one each for machine and supervisor levels. ![AdvPLIC ex 3Domains](_images/AdvPLIC-ex-3Domains.png) Figure 3\. A RISC-V system that extends the example of [Figure 2](#AdvPLIC-ex-2Domains) with a fifth M-mode-only "manager" hart, with a separate machine-level interrupt domain above the other domains. | | In order for the harts within an interrupt domain to have direct control over the interrupts from the domain, the harts must be cooperatively controlled by software at the same privilege level. In particular, a single operating system should control all of the harts associated with a supervisor-level interrupt domain. In the examples of [Figure 2](#AdvPLIC-ex-2Domains) and [Figure 3](#AdvPLIC-ex-3Domains), control of the APLIC’s supervisor-level interrupt domain could not be safely split among multiple independent OSes. Given the domain hierarchies depicted in the figures, if it were necessary to partition the application harts for multiple OSes, machine-level software would need to prevent direct OS access to the supervisor-level interrupt domain and instead provide SBI services for controlling APLIC interrupts or, alternatively, emulate the control interfaces of separate supervisor-level interrupt domains, one for each OS. Note that such emulation might still make use of the APLIC’s physical supervisor-level interrupt domain, but under the control of machine-level software. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | An APLIC’s interrupt domain hierarchy satisfies these rules: * The root domain is at machine level. * The parent of any supervisor-level interrupt domain is a machine-level domain that includes at least the same harts (but at machine level, obviously). The parent domain may have a larger set of harts at machine level. * For each interrupt domain, interrupts from the domain are signaled to harts all by the same method, either by wire or by MSIs, not by a mixture of methods among the harts. When a RISC-V hart’s external interrupt controller is an APLIC, not an IMSIC, the hart can be within only one interrupt domain of this APLIC at each privilege level. On the other hand, a hart that has an IMSIC for its external interrupt controller may, at each privilege level, be in multiple APLIC interrupt domains, even those of the same APLIC, and may potentially receive MSIs from multiple different APLICs in the machine. A platform might give software a way to choose between multiple interrupt domain hierarchies for any given APLIC. Any such configurability is outside the scope of this specification, but should be available to machine level only. ### [](#4-1-3-hart-index-numbers)4.1.3\. Hart index numbers Within a given interrupt domain, each of the domain’s harts has a unique_index number_ in the range 0 to (= 16,383). The index number a domain associates with a hart may or may not have any relationship to the unique hart identifier ("hart ID") that the Privileged Architecture assigns to the hart. Two different interrupt domains may employ a different mapping of index numbers to the same set of harts. However, if any of an APLIC’s interrupt domains can forward interrupts by MSI, then all machine-level domains of the APLIC share a common mapping of index numbers to harts. | | For efficiency, implementations should prefer small integers for hart index numbers. | | --------------------------------------------------------------------------------------- | ### [](#4-1-4-overview-of-interrupt-control-for-a-single-domain)4.1.4\. Overview of interrupt control for a single domain Each interrupt domain implemented by an APLIC has its own separate physical control interface that is memory-mapped in the machine’s address space, allowing access to each domain to be easily regulated by both PMP (physical memory protection) and page-based address translation. The control interfaces of all interrupt domains have a common structure. In most respects, every domain appears to software as though it were a root domain, without visibility of the domains above it in the hierarchy. An individual interrupt domain has the following components for each interrupt source at the APLIC: * Source configuration. This determines whether the specific source is active in the domain and, if so, how the incoming wire is to be interpreted, such as level-sensitive or edge-sensitive. For a source that is inactive in the domain, source configuration controls any delegation to a child domain. * Interrupt-pending and interrupt-enable bits. For an inactive source, these two bits are read-only zeros. Otherwise, the pending bit records an interrupt that arrived and has not yet been signaled or forwarded, while the enable bit determines whether interrupts from this source should currently be delivered, or should remain pending. * Target selection. For an active source, target selection determines the hart to receive the interrupt and either the interrupt’s priority or the new interrupt identity when forwarding as an MSI. For interrupt domains that deliver interrupts directly to harts rather than forwarding by MSIs, the domain has a final set of components for controlling interrupt delivery to harts, one instance per hart in the domain. | | Although an APLIC with multiple interrupt domains may appear to duplicate the per-source state listed above (source configuration, etc.) by a factor equal to the number of domains, in fact, APLIC implementations can exploit the fact that each source is ultimately active in only one domain. In all domains to which a specific interrupt source has not been delegated, the state associated with the source appears as read-only zeros, requiring no physical register bits. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#AdvPLIC-domainControlRegion)4.1.5\. Memory-mapped control region for an interrupt domain For each interrupt domain that an APLIC supports, there is a dedicated memory-mapped control region for managing interrupts in that domain. This control region is a multiple of 4 KiB in size and aligned to a 4-KiB address boundary. The smallest valid control region is 16 KiB. An interrupt domain’s control region is populated by a set of 32-bit registers. The first 16 KiB contains the registers listed in[Table 1](#TableAdvPLIC-domainControlRegion). __Table 1\. The registers of the first 16 KiB of an interrupt domain’s memory-mapped control region.__ | offset | size | register name | | | ------ | ------- | ----------------- | -------------------------------------- | | 0x0000 | 4 bytes | domaincfg | | | 0x0004 | 4 bytes | sourcecfg\[1\] | | | 0x0008 | 4 bytes | sourcecfg\[2\] | | | … | … | | | | 0x0FFC | 4 bytes | sourcecfg\[1023\] | | | 0x1BC0 | 4 bytes | mmsiaddrcfg | (machine-level interrupt domains only) | | 0x1BC4 | 4 bytes | mmsiaddrcfgh | ” | | 0x1BC8 | 4 bytes | smsiaddrcfg | ” | | 0x1BCC | 4 bytes | smsiaddrcfgh | ” | | 0x1C00 | 4 bytes | setip\[0\] | | | 0x1C04 | 4 bytes | setip\[1\] | | | … | … | | | | 0x1C7C | 4 bytes | setip\[31\] | | | 0x1CDC | 4 bytes | setipnum | | | 0x1D00 | 4 bytes | in\_clrip\[0\] | | | 0x1D04 | 4 bytes | in\_clrip\[1\] | | | … | … | | | | 0x1D7C | 4 bytes | in\_clrip\[31\] | | | 0x1DDC | 4 bytes | clripnum | | | 0x1E00 | 4 bytes | setie\[0\] | | | 0x1E04 | 4 bytes | setie\[1\] | | | … | … | | | | 0x1E7C | 4 bytes | setie\[31\] | | | 0x1EDC | 4 bytes | setienum | | | 0x1F00 | 4 bytes | clrie\[0\] | | | 0x1F04 | 4 bytes | clrie\[1\] | | | … | … | | | | 0x1F7C | 4 bytes | clrie\[31\] | | | 0x1FDC | 4 bytes | clrienum | | | 0x2000 | 4 bytes | setipnum\_le | | | 0x2004 | 4 bytes | setipnum\_be | | | 0x3000 | 4 bytes | genmsi | | | 0x3004 | 4 bytes | target\[1\] | | | 0x3008 | 4 bytes | target\[2\] | | | … | … | | | | 0x3FFC | 4 bytes | target\[1023\] | | Starting at offset `0x4000`, an interrupt domain’s control region may optionally have an array of _interrupt delivery control_ (IDC) structures, one for each potential hart index number in the range 0 to some maximum that is at least as large as the maximum hart index number for the interrupt domain. IDC structures are used only when the domain is configured to deliver interrupts directly to harts instead of being forwarded by MSIs. An interrupt domain that supports only interrupt forwarding by MSIs and not the direct delivery of interrupts by the APLIC does not need IDC structures in its control region. The first IDC structure, if any, is for the hart with index number 0; the second is for the hart with index number 1; and so forth. Each IDC structure is 32 bytes and has these defined registers: | offset | size | register name | | ------ | ------- | ------------- | | 0x00 | 4 bytes | idelivery | | 0x04 | 4 bytes | iforce | | 0x08 | 4 bytes | ithreshold | | 0x18 | 4 bytes | topi | | ox1C | 4 bytes | claimi | IDC structures are packed contiguously, 32 bytes per structure, so the offset from the beginning of an interrupt domain’s control region to its second IDC structure (hart index 1), if it exists, is `0x4020`; the offset to the third IDC structure (hart index 2), if it exists, is `0x4040`; etc. The array of IDC structures may include some for _potential_ hart index numbers that are not _actual_ hart index numbers in the domain. For example, the first IDC structure is always for hart index 0, but 0 is not necessarily a valid index number for any hart in the domain. For each IDC structure in the array that does not correspond to a valid hart index number in the domain, the IDC structure’s registers may (or may not) be all read-only zeros. Aside from the registers in[Table 1](#TableAdvPLIC-domainControlRegion)and those listed above for IDC structures, all other bytes in an interrupt domain’s control region are reserved and are implemented as read-only zeros. Only naturally aligned 32-bit simple reads and writes are supported within an interrupt domain’s control region. Writes to read-only bytes are ignored. For other forms of accesses (other sizes, misaligned accesses, or AMOs), implementations should preferably report an access fault or bus error but must otherwise ignore the access. The registers of the first 16 KiB of an interrupt domain’s control region (all but the IDC structures) are documented individually below. IDC structures are documented later, in[4.1.8\. Interrupt delivery directly by the APLIC](#AdvPLIC-directMode), "Interrupt delivery directly by the APLIC." #### [](#AdvPLIC-reg-domaincfg)4.1.5.1\. Domain configuration (`domaincfg`) The `domaincfg` register has this format: | bits 31:24 | read-only 0x80 | | ---------- | -------------- | | bit 8 | IE | | bit 7 | read-only 0 | | bit 2 | DM (**WARL**) | | bit 0 | BE (**WARL**) | All other register bits are reserved and read as zeros. Bit IE (Interrupt Enable) is a global enable for all active interrupt sources at this interrupt domain. Only when IE = 1 are pending-and-enabled interrupts actually signaled or forwarded to harts. The value of bit IE affects only whether interrupts are delivered to harts. It has no effect on any other APLIC state, including the interrupt-enable and interrupt-pending bits of interrupt sources and IDC registers `idelivery`, `topi`, and `claimi`. Field DM (Delivery Mode) is **WARL** and determines how this interrupt domain delivers interrupts to harts. The two possible values for DM are: | 0 = | direct delivery mode | | --- | -------------------- | | 1 = | MSI delivery mode | In _direct delivery mode_, interrupts are prioritized and signaled directly to harts by the APLIC itself. In _MSI delivery mode_, interrupts are forwarded by the APLIC as MSIs to harts, presumably for further handling by IMSICs at those harts. A given APLIC implementation may support either or both of these delivery modes for each interrupt domain. If the interrupt domain’s harts have IMSICs, then unless the relevant interrupt files of those IMSICs support value `0x40000000` for register `eidelivery`, setting DM to zero (direct delivery mode) will have the same effect as setting IE to zero. See [External interrupt delivery enable register (eidelivery)](IMSIC.html#IMSIC-reg-eidelivery)and [4.1.8.2\. Interrupt delivery and handling](#AdvPLIC-directMode-intrDelivery). BE (Big-Endian) is a **WARL** field that determines the byte order for most registers in the interrupt domain’s memory-mapped control region. If BE = 0, byte order is little-endian, and if BE = 1, it is big-endian. For RISC-V systems that support only little-endian, BE may be read-only zero, and for those that support only big-endian, BE may be read-only one. For bi-endian systems, BE is writable. Field BE affects the byte order of accesses to the `domaincfg` register itself, just as for other registers in the interrupt domain’s control region. To deal with this fact, the read-only value in `domaincfg’s` most-significant byte, bits 31:24, serves two purposes. First, for any read of `domaincfg`, the register’s correct byte order is easily determined from the four-byte value obtained: When interpreted in the correct byte order, bit 31 is one, and in the wrong order, bit 31 is zero. Second, if the value of BE is uncertain (prior to software initializing the interrupt domain, presumably), an 8-bit value can be safely written to `domaincfg` by writing ( <<24)| , where <<24 represents shifting left by 24 bits, and the vertical bar (|) represents bitwise logical OR. After `domaincfg` is written once, the value of BE should then be known, so subsequent writes should not need to repeat the same trick. At system reset, all writable bits in `domaincfg` are initialized to zero, including IE. If an implementation supports additional forms of reset for the APLIC, it is implementation-defined (or possibly platform-defined) how these other resets may affect `domaincfg`. #### [](#AdvPLIC-reg-sourcecfg)4.1.5.2\. Source configurations (`sourcecfg[1]–sourcecfg[1023]`) For each possible interrupt source , register `sourcecfg[ ]` controls the _source mode_ for source in this interrupt domain as well as any delegation of the source to a child domain. When source is not implemented, or appears in this domain not to be implemented, `sourcecfg[ ]` is read-only zero. If source was not delegated to this domain and is then changed (at the parent domain) to become delegated to this domain, `sourcecfg[ ]` remains zero until successfully written with a nonzero value. Bit 10 of `sourcecfg[ ]` is a 1-bit field called D (Delegate). If D = 1, source is delegated to a child domain, and if D = 0, it is not delegated to a child domain. Interpretation of the rest of `sourcecfg[ ]` depends on field D. When interrupt source is delegated to a child domain, `sourcecfg[ ]` has this format: | bit 10 | D, =1 | | -------- | ---------------------- | | bits 9:0 | Child Index (**WLRL**) | All other register bits are reserved and read as zeros. Child Index is a **WLRL** field that specifies the interrupt domain to which this source is delegated. For an interrupt domain with child domains, this field must be able to hold integer values in the range 0 to . Each interrupt domain has a fixed mapping from these index numbers to child domains. If an interrupt domain has no children in the domain hierarchy, bit D cannot be set to one in any `sourcecfg` register for that domain. For such a leaf domain, attempting to write a `sourcecfg` register with a value that has bit 10 = 1 causes the entire register to be set to zero instead. When interrupt source is not delegated to a child domain, `sourcecfg[ ]` has this format: | bit 10 | D, =0 | | -------- | ------------- | | bits 2:0 | SM (**WARL**) | All other register bits are reserved and read as zeros. The SM (Source Mode) field is **WARL** and controls whether the interrupt source is active in this domain, and if so, what values or transitions on the incoming wire are interpreted as interrupts. The values allowed for SM and their meanings are listed in[Table 2](#TableAdvPLIC-sourcecfg-SM). Inactive (zero) is always supported for field SM. Implementations are free to choose, independently for each interrupt source, what other values are supported for SM. __Table 2\. Encoding of the SM (Source Mode) field of a sourcecfg register when bit D = 0__ | Value | Name | Description | | ----- | -------- | ---------------------------------------------------------- | | 0 | Inactive | Inactive in this domain (and not delegated) | | 1 | Detached | Active, detached from the source wire | | 2–3 | — | _Reserved_ | | 4 | Edge1 | Active, edge-sensitive; interrupt asserted on rising edge | | 5 | Edge0 | Active, edge-sensitive; interrupt asserted on falling edge | | 6 | Level1 | Active, level-sensitive; interrupt asserted when high | | 7 | Level0 | Active, level-sensitive; interrupt asserted when low | An interrupt source is inactive in the interrupt domain if either the source is delegated to a child domain (D = 1) or it is not delegated (D = 0) and SM is Inactive. Whenever interrupt source is inactive in an interrupt domain, the corresponding interrupt-pending and interrupt-enable bits within the domain are read-only zeros, and register `target[ ]` is also read-only zero. If source is changed from inactive to an active mode, the interrupt source’s pending and enable bits remain zeros, unless set automatically for a reason specified later in this section or in[4.1.7\. Precise effects on interrupt-pending bits](#AdvPLIC-pendingBits), and the defined subfields of `target[ ]` obtain UNSPECIFIED values. When a source is configured as Detached, its wire input is ignored; however, the interrupt-pending bit may still be set by a write to a `setip` or `setipnum` register. (This mode can be useful for receiving MSIs, for example.) An edge-sensitive source can be configured to recognize an incoming interrupt on either a rising edge (low-to-high transition) or a falling edge (high-to-low transition). When configured for a falling edge (mode Edge0), the source is said to be _inverted_. A level-sensitive source can be configured to interpret either a high level (1) or a low level (0) on the wire as the assertion of an interrupt. When configured for a low level (mode Level0), the source is said to be _inverted_. For an interrupt source that is configured as edge-sensitive or level-sensitive, define | _rectified input value_ \= (incoming wire value) XOR (source is inverted). | | -------------------------------------------------------------------------- | For a source that is inactive or Detached, the _rectified input value_is zero. Any write to a `sourcecfg` register might (or might not) cause the corresponding interrupt-pending bit to be set to one if the rectified input value is high (= 1) under the new source mode. A write to a `sourcecfg` register will not by itself cause a pending bit to be cleared except when the source is made inactive. (But see [4.1.7\. Precise effects on interrupt-pending bits](#AdvPLIC-pendingBits).) #### [](#AdvPLIC-reg-mmsiaddrcfg)4.1.5.3\. Machine MSI address configuration (`mmsiaddrcfg` and `mmsiaddrcfgh`) For machine-level interrupt domains, registers `mmsiaddrcfg` and `mmsiaddrcfgh` may optionally provide parameters used to determine the addresses to write outgoing MSIs. If no interrupt domain of the APLIC supports MSI delivery mode (`domaincfg`.DM is read-only zero for all domains), these two registers are not implemented for any domain. Otherwise, they are implemented for the root domain, and may or may not be implemented for other machine-level domains. For domains not at machine level, they are never implemented. When a domain does not implement `mmsiaddrcfg` and `mmsiaddrcfgh`, the eight bytes at their locations are simply read-only zeros like other reserved bytes. Registers `mmsiaddrcfg` and `mmsiaddrcfgh` are potentially writable only for the root domain. For all other machine-level domains that implement them, they are read-only. When implemented, `mmsiaddrcfg` has this format: | bits 31:0 | Low Base PPN (**WARL**) | | --------- | ----------------------- | and `mmsiaddrcfgh` has this format: | bit 31 | L | | ---------- | ------------------------ | | bits 28:24 | HHXS (**WARL**) | | bits 22:20 | LHXS (**WARL**) | | bits 18:16 | HHXW (**WARL**) | | bits 15:12 | LHXW (**WARL**) | | bits 11:0 | High Base PPN (**WARL**) | All other bits of `mmsiaddrcfgh` are reserved and read as zeros. Fields High Base PPN from `mmsiaddrcfgh` and Low Base PPN from `mmsiaddrcfg` concatenate to form a 44-bit Base PPN (Physical Page Number). The use of this value and fields HHXS (High Hart Index Shift), LHXS (Low Hart Index Shift), HHXW (High Hart Index Width), and LHXW (Low Hart Index Width) for determining target addresses for MSIs is described later, in[4.1.9.1\. Addresses and data for outgoing MSIs](#AdvPLIC-MSIAddrs). When `mmsiaddrcfg` and `mmsiaddrcfgh` are writable (root domain only), all fields other than L are **WARL**. An implementation is free to choose what values are supported. Typically, some bits are writable while others are read-only constants. In the extreme, the values of all fields may be entirely constant, fixed by the implementation. If bit L in `mmsiaddrcfgh` is set to one, `mmsiaddrcfg` and `mmsiaddrcfgh` are _locked_, and writes to the registers are ignored, making the registers effectively read-only. When L = 1, the other fields in `mmsiaddrcfg` and `mmsiaddrcfgh` may optionally all read as zeros. In that case, if these other fields were given nonzero values when L was first set in the root domain, their values are retained internally by the APLIC but become no longer visible by reading `mmsiaddrcfg` and `mmsiaddrcfgh`. Setting `mmsiaddrcfgh`.L to one also locks registers `smsiaddrcfg` and `smsiaddrcfgh` described in the next subsection, if those registers are implemented as well. For the root domain, L is initialized at system reset to either zero or one, whichever is deemed appropriate for the specific APLIC implementation. If reset initializes L to one, either the other fields are hardwired by the APLIC to constants, or the APLIC has a different means, outside of this standard, for determining the addresses of outgoing MSI writes. In the latter case, the other fields in `mmsiaddrcfg` and `mmsiaddrcfgh` may all read as zeros, so registers `mmsiaddrcfg` and `mmsiaddrcfgh` have only read-only values zero and `0x80000000`respectively. Any time `mmsiaddrcfg` or `mmsiaddrcfgh` has a different value (not zero or `0x80000000`respectively), the addresses for outgoing MSI writes directed to machine level must be derivable from the visible values of these registers, as specified in [4.1.9.1\. Addresses and data for outgoing MSIs](#AdvPLIC-MSIAddrs). For machine-level domains that are not the root domain, if these registers are implemented, bit L is always one, and the other fields either are read-only copies of `mmsiaddrcfg` and `mmsiaddrcfgh` from the root domain, or are all zeros. | | Giving software the ability to arbitrarily determine the addresses to which MSIs are sent, even if allowed only for machine level, permits bypassing physical memory protection (PMP). For APLICs that support MSI delivery mode, it is recommended, if feasible, that the APLIC internally hardwire the physical addresses for all target IMSICs, putting those addresses beyond the reach of software to change. However, not all APLIC implementations will be able to follow that recommendation. It is expected that most systems will arrange the physical addresses of target IMSICs in a simple linear correspondence with hart index numbers. (See [Arrangement of the memory regions of multiple interrupt files](IMSIC.html#IMSIC-systemMemRegions).) Registers mmsiaddrcfg and mmsiaddrcfgh (along with smsiaddrcfg and smsiaddrcfgh from the next subsection) allow sufficiently trusted machine-level software, early after system reset, to configure the pattern of physical addresses for target IMSICs and then lock this configuration against subsequent tampering. APLICs that actually hardwire the IMSIC addresses internally can implement these registers simply as read-only with values zero and 0x80000000. Or, if the IMSIC addresses must be configured by software but the formula is too complex for registers mmsiaddrcfg and mmsiaddrcfgh to handle, again the registers can be implemented simply as read-only with values zero and 0x80000000, and a separate, custom mechanism supplied for configuring the IMSIC addresses. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If an APLIC supports additional forms of reset besides system reset, it is implementation-defined (or possibly platform-defined) how these other resets may affect `mmsiaddrcfg` and `mmsiaddrcfgh` (as well as `smsiaddrcfg` and `smsiaddrcfgh`) in the root domain. However, it must not be possible for insufficiently privileged software to use a localized reset to unlock these registers by changing bit L back to zero. For this reason, it is likely that only a complete system reset affects these registers, and any other resets do not. #### [](#AdvPLIC-reg-smsiaddrcfg)4.1.5.4\. Supervisor MSI address configuration (`smsiaddrcfg` and `smsiaddrcfgh`) For machine-level interrupt domains, registers `smsiaddrcfg` and `smsiaddrcfgh` may optionally provide parameters used by supervisor-level domains to determine the addresses to write outgoing MSIs. Registers `smsiaddrcfg` and `smsiaddrcfgh` are implemented by a domain if the domain implements `mmsiaddrcfg` and `mmsiaddrcfgh`and the APLIC has at least one supervisor-level interrupt domain. If the registers are not implemented, the eight bytes at their locations are simply read-only zeros like other reserved bytes. Like `mmsiaddrcfg` and `mmsiaddrcfgh`, registers `smsiaddrcfg` and `smsiaddrcfgh` are potentially writable only for the root domain. For all other machine-level domains that implement them, they are read-only. When implemented, `smsiaddrcfg` has this format: | bits 31:0 | Low Base PPN (**WARL**) | | --------- | ----------------------- | and `smsiaddrcfgh` has this format: | bits 22:20 | LHXS (**WARL**) | | ---------- | ------------------------ | | bits 11:0 | High Base PPN (**WARL**) | All other bits of `smsiaddrcfgh` are reserved and read as zeros. Fields High Base PPN from `smsiaddrcfgh` and Low Base PPN from `smsiaddrcfg` concatenate to form a 44-bit Base PPN (Physical Page Number). The use of this value and field LHXS (Low Hart Index Shift) for determining target addresses for MSIs is described later, in [4.1.9.1\. Addresses and data for outgoing MSIs](#AdvPLIC-MSIAddrs). When `smsiaddrcfg` and `smsiaddrcfgh` are writable (root domain only), all fields are **WARL**. An implementation is free to choose what values are supported, just as for `mmsiaddrcfg` and `mmsiaddrcfgh`. If register `mmsiaddrcfgh` of the domain has bit L set to one, then `smsiaddrcfg` and `smsiaddrcfgh` are _locked_ as read-only alongside `mmsiaddrcfg` and `mmsiaddrcfgh`. When `mmsiaddrcfgh.L` \= 1, if the readable values of `mmsiaddrcfg` and `mmsiaddrcfgh` are zero and `0x80000000` respectively—because their other fields are hidden—then `smsiaddrcfg` and `smsiaddrcfgh` are hidden also and read as zeros. For the root domain only, if `mmsiaddrcfgh.L` \= 1 and the MSI-address-configuration fields are hidden (so `mmsiaddrcfgh` reads as `0x80000000` and registers `mmsiaddrcfg`, `smsiaddrcfg`, and `smsiaddrcfgh` all read as zeros), then whatever values `smsiaddrcfg` and `smsiaddrcfgh` had when `mmsiaddrcfgh`.L was first set are retained internally by the APLIC, though those values are no longer visible by reading the registers. Alternatively, if system reset initializes `mmsiaddrcfgh.L` \= 1 in the root domain, and if all MSI-address-configuration fields never appear as anything other than zeros, then the APLIC implementation has some other, possibly nonstandard, means for determining the addresses of outgoing MSIs, as discussed in the previous subsection,[4.1.5.3\. Machine MSI address configuration (mmsiaddrcfg and mmsiaddrcfgh)](#AdvPLIC-reg-mmsiaddrcfg). Any time `mmsiaddrcfg` and `mmsiaddrcfgh` are not read-only zero and `0x80000000` respectively, the addresses for outgoing MSI writes directed to supervisor level must be derivable from the visible values of registers `mmsiaddrcfgh`, `smsiaddrcfg`, and `smsiaddrcfgh`, as specified in[4.1.9.1\. Addresses and data for outgoing MSIs](#AdvPLIC-MSIAddrs). For machine-level domains that are not the root domain, if `smsiaddrcfg` and `smsiaddrcfgh` are implemented and are not read-only zeros, then they are read-only copies of the same registers from the root domain. #### [](#4-1-5-5-set-interrupt-pending-bits-setip0-setip31)4.1.5.5\. Set interrupt-pending bits (`setip[0]`\-`setip[31]`) Reading or writing `setip[ ]` register reads or potentially modifies the pending bits for interrupt sources through . For an implemented interrupt source within that range, the pending bit for source corresponds with register bit ( ). A read of a `setip` register returns the pending bits of the corresponding interrupt sources. Bit positions in the result value that do not correspond to an implemented interrupt source (such as bit 0 of `setip[0]`) are zeros. On a write to a `setip` register, for each bit that is one in the 32-bit value written, if that bit position corresponds to an active interrupt source, the interrupt-pending bit for that source is set to one if possible. See[4.1.7\. Precise effects on interrupt-pending bits](#AdvPLIC-pendingBits) for exactly when a pending bit may be set by writing to a `setip` register. #### [](#4-1-5-6-set-interrupt-pending-bit-by-number-setipnum)4.1.5.6\. Set interrupt-pending bit by number (`setipnum`) If is an active interrupt source number in the domain, writing 32-bit value to register `setipnum` causes the pending bit for source to be set to one if possible. See[4.1.7\. Precise effects on interrupt-pending bits](#AdvPLIC-pendingBits) for exactly when a pending bit may be set by writing to `setipnum`. A write to `setipnum` is ignored if the value written is not an active interrupt source number in the domain. A read of `setipnum` always returns zero. #### [](#4-1-5-7-rectified-inputs-clear-interrupt-pending-bits-in%5Fclrip0-in%5Fclrip31)4.1.5.7\. Rectified inputs, clear interrupt-pending bits (`in_clrip[0]`\-`in_clrip[31]`) Reading register `in_clrip[ ]` returns the rectified input ([4.1.5.2\. Source configurations (sourcecfg\[1\]–sourcecfg\[1023\])](#AdvPLIC-reg-sourcecfg)) for interrupt sources through , while writing `in_clrip[ ]` potentially modifies the pending bits for the same sources. For an implemented interrupt source within the specified range, source corresponds with register bit ( ). A read of an `in_clrip` register returns the rectified input values of the corresponding interrupt sources. Bit positions in the result value that do not correspond to an implemented interrupt source (such as bit 0 of `in_clrip[0]`) are zeros. On a write to an `in_clrip` register, for each bit that is one in the 32-bit value written, if that bit position corresponds to an active interrupt source, the interrupt-pending bit for that source is cleared if possible. See[4.1.7\. Precise effects on interrupt-pending bits](#AdvPLIC-pendingBits) for exactly when a pending bit may be cleared by writing to an `in_clrip` register. #### [](#4-1-5-8-clear-interrupt-pending-bit-by-number-clripnum)4.1.5.8\. Clear interrupt-pending bit by number (`clripnum`) If is an active interrupt source number in the domain, writing 32-bit value to register `clripnum` causes the pending bit for source to be cleared if possible. See[4.1.7\. Precise effects on interrupt-pending bits](#AdvPLIC-pendingBits) for exactly when a pending bit may be cleared by writing to `clripnum`. A write to `clripnum` is ignored if the value written is not an active interrupt source number in the domain. A read of `clripnum` always returns zero. #### [](#4-1-5-9-set-interrupt-enable-bits-setie0-setie31)4.1.5.9\. Set interrupt-enable bits (`setie[0]`\-`setie[31]`) Reading or writing register `setie[ ]` reads or potentially modifies the enable bits for interrupt sources through . For an implemented interrupt source within that range, the enable bit for source corresponds with register bit . A read of a `setie` register returns the enable bits of the corresponding interrupt sources. Bit positions in the result value that do not correspond to an implemented interrupt source (such as bit 0 of `setie[0]`) are zeros. On a write to a `setie` register, for each bit that is one in the 32-bit value written, if that bit position corresponds to an active interrupt source, the interrupt-enable bit for that source is set to one. #### [](#4-1-5-10-set-interrupt-enable-bit-by-number-setienum)4.1.5.10\. Set interrupt-enable bit by number (`setienum`) If is an active interrupt source number in the domain, writing 32-bit value to register `setienum` causes the enable bit for source to be set to one. A write to `setienum` is ignored if the value written is not an active interrupt source number in the domain. A read of `setienum` always returns zero. #### [](#4-1-5-11-clear-interrupt-enable-bits-clrie0-clrie31)4.1.5.11\. Clear interrupt-enable bits (`clrie[0]`\-`clrie[31]`) Writing register `clrie[ ]` potentially modifies the enable bits for interrupt sources through . For an implemented interrupt source within that range, the enable bit for source corresponds with register bit . On a write to a `clrie` register, for each bit that is one in the 32-bit value written, the interrupt-enable bit for that source is cleared. A read of a `clrie` register always returns zero. #### [](#4-1-5-12-clear-interrupt-enable-bit-by-number-clrienum)4.1.5.12\. Clear interrupt-enable bit by number (`clrienum`) If is an active interrupt source number in the domain, writing 32-bit value to register `clrienum` causes the enable bit for source to be cleared. A write to `clrienum` is ignored if the value written is not an active interrupt source number in the domain. A read of `clrienum` always returns zero. #### [](#4-1-5-13-set-interrupt-pending-bit-by-number-little-endian-setipnum%5Fle)4.1.5.13\. Set interrupt-pending bit by number, little-endian (`setipnum_le`) Register `setipnum_le` acts identically to `setipnum` except that byte order is always little-endian, as though field BE (Big-Endian) of register `domaincfg` is zero. For systems that are big-endian-only, with `domaincfg`.BE hardwired to one, `setipnum_le` need not be implemented, in which case the four bytes at this offset are simply read-only zeros like other reserved bytes. `setipnum_le` may be used as a write port for MSIs. #### [](#4-1-5-14-set-interrupt-pending-bit-by-number-big-endian-setipnum%5Fbe)4.1.5.14\. Set interrupt-pending bit by number, big-endian (`setipnum_be`) Register `setipnum_be` acts identically to `setipnum` except that byte order is always big-endian, as though field BE (Big-Endian) of register `domaincfg` is one. For systems that are little-endian-only, with `domaincfg`.BE hardwired to zero, `setipnum_be` need not be implemented, in which case the four bytes at this offset are simply read-only zeros like other reserved bytes. For systems built mainly for big-endian byte order, `setipnum_be` may be useful as a write port for MSIs from some devices. #### [](#AdvPLIC-reg-genmsi)4.1.5.15\. Generate MSI (`genmsi`) When the interrupt domain is configured in MSI delivery mode (`domaincfg`.DM = 1), register `genmsi` can be used to cause an _extempore_ MSI to be sent from the APLIC to a hart. The main purpose for this function is to assist in establishing a temporary known ordering between a hart’s writes to the APLIC’s registers and the transmission of MSIs from the APLIC to the hart, as explained later in [4.1.9.3\. Synchronizing interactions between a hart and the APLIC](#AdvPLIC-MSISync). | | For other purposes, sending an MSI to a hart is usually better done by writing directly to the hart’s IMSIC, rather than employing an APLIC as an intermediary. Use of the genmsi register should be minimized to avoid it becoming a bottleneck. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Register `genmsi` has this format: | bits 31:18 | Hart Index (**WLRL**) | | ---------- | --------------------- | | bits 12 | Busy (read-only) | | bits 10:0 | EIID (**WARL**) | All other register bits are reserved and read as zeros. The Busy bit is ordinarily zero (false), but a write to `genmsi` causes Busy to become one (true), indicating an extempore MSI is pending. The Hart Index field specifies the destination hart, and EIID (External Interrupt Identity) specifies the data value for the MSI. Fields Hart Index and EIID have the same formats and behavior as in a `target` register, documented in the next subsection, [4.1.5.16\. Interrupt targets (target\[1\]-target\[1023\])](#AdvPLIC-reg-target). For a machine-level interrupt domain, an extempore MSI is sent to the destination hart at machine level, and for a supervisor-level interrupt domain, an extempore MSI is sent to the destination hart at supervisor level. A pending extempore MSI should be sent by the APLIC with minimal delay. Once it has left the APLIC and the APLIC is able to accept a new write to `genmsi` for another extempore MSI, Busy reverts to false. All MSIs previously sent from this APLIC to the same hart must be visible at the hart’s IMSIC before the extempore MSI becomes visible at the hart’s IMSIC. While Busy is true, writes to `genmsi` are ignored. Extempore MSIs are not affected by the IE bit of the domain’s `domaincfg` register. An extempore MSI is sent even if `domaincfg`.IE = 0. When the interrupt domain is configured in direct delivery mode (`domaincfg`.DM = 0), register `genmsi` is read-only zero. #### [](#AdvPLIC-reg-target)4.1.5.16\. Interrupt targets (`target[1]-target[1023]`) If interrupt source is inactive in this domain, register `target[ ]` is read-only zero. If source is active, `target[ ]` determines the hart to which interrupts from the source are signaled or forwarded. The exact interpretation of `target[ ]` depends on the delivery mode configured by field DM of register `domaincfg`. If `domaincfg`.DM is changed, the `target` registers for all active interrupt sources within the domain obtain UNSPECIFIED values in all fields defined for the new delivery mode. ##### [](#4-1-5-16-1-active-source-direct-delivery-mode)4.1.5.16.1\. Active source, direct delivery mode For an active interrupt source , if the domain is configured in direct delivery mode (`domaincfg`.DM = 0), then register `target[ ]` has this format: | bits 31:18 | Hart Index (**WLRL**) | | ---------- | --------------------- | | bits 7:0 | IPRIO (**WARL**) | All other register bits are reserved and read as zeros. Hart Index is a **WLRL** field that specifies the hart to which interrupts from this source will be delivered. Field IPRIO (Interrupt Priority) specifies the _priority number_ for the interrupt source. This field is a **WARL** unsigned integer of _IPRIOLEN_ bits, where IPRIOLEN is a constant parameter for the given APLIC, in the range of 1 to 8\. Only values 1 through are allowed for IPRIO, not zero. A write to a `target` register sets IPRIO equal to bits :0 of the 32-bit value written, unless those bits are all zeros, in which case the priority number is set to 1 instead. (If IPRIOLEN = 1, these rules cause IPRIO to be effectively read-only with value 1.) Smaller priority numbers convey higher priority. When interrupt sources have equal priority number, the source with the lowest identity number has the highest priority. | | Interrupt priorities are encoded as integers, with smaller numbers denoting higher priority, to match the encoding of priorities by IMSICs. | | ---------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#4-1-5-16-2-active-source-msi-delivery-mode)4.1.5.16.2\. Active source, MSI delivery mode For an active interrupt source , if the domain is configured in MSI delivery mode (`domaincfg`.DM = 1), then register `target[ ]` has this format: | bits 31:18 | Hart Index (**WLRL**) | | ---------- | ---------------------- | | bits 17:12 | Guest Index (**WLRL**) | | bits 10:0 | EIID (**WARL**) | Bit 11 is reserved and reads as zero. The Hart Index field specifies the hart to which interrupts from this source will be forwarded. If the interrupt domain is at supervisor level and the domain’s harts implement the H extension, then Guest Index is a **WLRL** field that must be able to hold all integer values in the range 0 through GEILEN. (Parameter _GEILEN_ is defined by the H extension.) Otherwise, field Guest Index is read-only zero. For a supervisor-level interrupt domain, a nonzero Guest Index is the number of the target hart’s guest interrupt file to which MSIs will be sent. When Guest Index is zero, MSIs from a supervisor-level domain are forwarded to the target hart at supervisor level. For a machine-level domain, Guest Index is read-only zero, and MSIs are forwarded to a target hart always at machine level. Together, fields Hart Index and Guest Index of register `target[ ]` determine the address for MSIs forwarded for interrupt source . The remaining field EIID (External Interrupt Identity) specifies the data value for those MSIs, eventually becoming the minor identity for an external interrupt at the target hart. If the interrupt domain’s harts have IMSIC interrupt files that implement distinct interrupt identities ([Interrupt files and interrupt identities](IMSIC.html#IMSIC-intrFilesAndIdents)), then EIID is a \-bit unsigned integer field, where . EIID is thus able to hold at least values 0 through . A write to a `target`register sets the implemented bits of EIID equal to the least-significant bits of the 32-bit value written. ### [](#4-1-6-reset)4.1.6\. Reset Upon reset of an APLIC, all its state becomes valid and consistent but otherwise UNSPECIFIED, except for: * the domaincfg register of each interrupt domain ([4.1.5.1\. Domain configuration (domaincfg)](#AdvPLIC-reg-domaincfg)); * possibly the MSI address configuration registers of machine-level interrupt domains ([4.1.5.3\. Machine MSI address configuration (mmsiaddrcfg and mmsiaddrcfgh)](#AdvPLIC-reg-mmsiaddrcfg) and [4.1.5.4\. Supervisor MSI address configuration (smsiaddrcfg and smsiaddrcfgh)](#AdvPLIC-reg-smsiaddrcfg)); and * the Busy bit of each interrupt domain’s `genmsi` register, if it exists ([4.1.5.15\. Generate MSI (genmsi)](#AdvPLIC-reg-genmsi)). ### [](#AdvPLIC-pendingBits)4.1.7\. Precise effects on interrupt-pending bits An attempt to set or clear an interrupt source’s pending bit by writing to a register in the interrupt domain’s control region may or may not be successful, depending on the corresponding source mode, the interrupt domain’s delivery mode, and the state of the source’s rectified input value (defined in [4.1.5.2\. Source configurations (sourcecfg\[1\]–sourcecfg\[1023\])](#AdvPLIC-reg-sourcecfg)). The following enumerates all the circumstances when a pending bit is set or cleared for a given source mode. If the source mode is Detached: * The pending bit is set to one only by a relevant write to a `setip` or `setipnum` register. * The pending bit is cleared when the interrupt is claimed at the APLIC or forwarded by MSI, or by a relevant write to an `in_clrip` register or to `clripnum`. If the source mode is Edge1 or Edge0: * The pending bit is set to one by a low-to-high transition in the rectified input value, or by a relevant write to a `setip` or `setipnum` register. * The pending bit is cleared when the interrupt is claimed at the APLIC or forwarded by MSI, or by a relevant write to an `in_clrip` register or to `clripnum`. If the source mode is Level1 or Level0 and the interrupt domain is configured in direct delivery mode (`domaincfg`.DM = 0): * The pending bit is set to one whenever the rectified input value is high. The pending bit cannot be set by a write to a `setip` or `setipnum` register. * The pending bit is cleared whenever the rectified input value is low. The pending bit is not cleared by a claim of the interrupt at the APLIC, nor can it be cleared by a write to an `in_clrip` register or to `clripnum`. If the source mode is Level1 or Level0 and the interrupt domain is configured in MSI delivery mode (`domaincfg`.DM = 1): * The pending bit is set to one by a low-to-high transition in the rectified input value. The pending bit may also be set by a relevant write to a `setip` or `setipnum` register when the rectified input value is high, but not when the rectified input value is low. * The pending bit is cleared whenever the rectified input value is low, when the interrupt is forwarded by MSI, or by a relevant write to an `in_clrip` register or to `clripnum`. | | When an interrupt domain is in direct delivery mode, the pending bit for a level-sensitive source is always just a copy of the rectified input value. Even in MSI delivery mode, the pending bit for a level-sensitive source is never set (= 1) when the rectified input value is low. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | In addition to the rules above, a write to a `sourcecfg` register can cause the source’s interrupt-pending bit to be set to one, as specified in[4.1.5.2\. Source configurations (sourcecfg\[1\]–sourcecfg\[1023\])](#AdvPLIC-reg-sourcecfg). ### [](#AdvPLIC-directMode)4.1.8\. Interrupt delivery directly by the APLIC When an interrupt domain is in direct delivery mode (`domaincfg`.DM = 0), interrupts are delivered from the APLIC to harts by a unique signal to each hart, usually a dedicated wire. In this case, the domain’s memory-mapped control region contains at the end an array of interrupt delivery control (IDC) structures, one IDC structure per potential hart index. The first IDC structure is for the domain’s hart with index 0; the second is for the hart with index 1; etc. #### [](#AdvPLIC-IDC)4.1.8.1\. Interrupt delivery control (IDC) structure Each IDC structure is 32 bytes (naturally aligned to a 32-byte address boundary) and has these defined registers: | offset | size | register name | | ------ | ------- | ------------- | | 0x00 | 4 bytes | idelivery | | ox04 | 4 bytes | iforce | | 0x08 | 4 bytes | ithreshold | | 0x18 | 4 bytes | topi | | 0x1C | 4 bytes | claimi | If the IDC structure is for a hart index number that is not valid for any actual hart in the interrupt domain, then these registers may optionally be all read-only zeros. Otherwise, the registers are documented individually below. | | A particular APLIC might be built to support up to some maximum number of harts without complete knowledge of the set of hart index numbers the system will employ in each interrupt domain. In that case, for the hart index numbers that are unused, the APLIC may have IDC structures that are functional within the APLIC (not read-only zeros) but simply left unconnected to any physical harts. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#4-1-8-1-1-interrupt-delivery-enable-idelivery)4.1.8.1.1\. Interrupt delivery enable (`idelivery`) `idelivery` is a **WARL** register that controls whether interrupts that are targeted to the corresponding hart are delivered to the hart so they appear as a pending interrupt in the hart’s `mip` CSR. Only two values are currently defined for `idelivery`: | 0 = | interrupt delivery is disabled | | --- | ------------------------------ | | 1 = | interrupt delivery is enabled | The `idelivery` register affects only whether interrupts are delivered to the relevant hart. It has no effect on any other APLIC state, including IDC registers `topi` and `claimi`. If an IDC structure is for a nonexistent hart (i.e., corresponding to a hart index number that is not valid for any actual hart in the interrupt domain), setting `idelivery` to 1 does not deliver interrupts to any hart. ##### [](#4-1-8-1-2-interrupt-force-iforce)4.1.8.1.2\. Interrupt force (`iforce`) `iforce` is a **WARL** register useful for testing. Only values 0 and 1 are allowed. Setting `iforce` \= 1 forces an interrupt to be asserted to the corresponding hart whenever both the IE field of `domaincfg` is one and interrupt delivery is enabled to the hart by the `idelivery` register. When `topi` is zero, this creates a _spurious external interrupt_ for the hart. When a read of register `claimi` returns an interrupt identity of zero (indicating a spurious interrupt), `iforce` is automatically cleared to zero. ##### [](#4-1-8-1-3-interrupt-enable-threshold-ithreshold)4.1.8.1.3\. Interrupt enable threshold (`ithreshold`) `ithreshold` is a **WLRL** register that determines the minimum interrupt priority (maximum priority number) for an interrupt to be signaled to the corresponding hart. Register `ithreshold` implements exactly IPRIOLEN bits, and thus is capable of holding all priority numbers from 0 to . When `ithreshold` is a nonzero value , interrupt sources with priority numbers and higher do not contribute to signaling interrupts to the hart, as though those sources were not enabled, regardless of the settings of their interrupt-enable bits. When `ithreshold` is zero, all enabled interrupt sources can contribute to signaling interrupts to the hart. ##### [](#4-1-8-1-4-top-interrupt-topi)4.1.8.1.4\. Top interrupt (`topi`) `topi` is a read-only register whose value indicates the current highest-priority pending-and-enabled interrupt targeted to this hart that also exceeds the priority threshold specified by `ithreshold`, if not zero. A read of `topi` returns zero either if no interrupt that is targeted to this hart is both pending and enabled, or if `ithreshold` is not zero and no pending-and-enabled interrupt targeted to this hart has a priority number less than the value of `ithreshold`. Otherwise, the value returned from a read of `topi` has this format: | bits 25:16 | Interrupt identity (source number) | | ---------- | ---------------------------------- | | bits 7:0 | Interrupt priority | All other bit positions are zeros. The interrupt identity reported in `topi` is the minor identity for an external interrupt at the target hart. The value of `topi` is not affected by `domaincfg`.IE or by `idelivery`. Writes to `topi` are ignored. ##### [](#4-1-8-1-5-claim-top-interrupt-claimi)4.1.8.1.5\. Claim top interrupt (`claimi`) Register `claimi` has the same value as `topi`. When this value is not zero, reading `claimi` has the simultaneous side effect of clearing the pending bit for the reported interrupt identity, if possible. See[4.1.7\. Precise effects on interrupt-pending bits](#AdvPLIC-pendingBits) for exactly when the pending bit is cleared by a read of `claimi`. A read from `claimi` that returns a value of zero has the simultaneous side effect of setting the `iforce` register to zero. Writes to `claimi` are ignored. #### [](#AdvPLIC-directMode-intrDelivery)4.1.8.2\. Interrupt delivery and handling When an interrupt domain is configured so the APLIC delivers interrupts directly to harts (field DM of `domaincfg` is zero), the APLIC supplies the_external interrupt_ signals, at the domain’s privilege level, for all harts of the domain, so long as one of the following is true: (a) the harts do not have IMSICs, or (b) the `eidelivery` registers of the relevant IMSIC interrupt files are set to `0x40000000` ([External interrupt delivery enable register (eidelivery)](IMSIC.html#IMSIC-reg-eidelivery)). For a machine-level domain, the interrupt signals from the APLIC appear as bit MEIP (Machine External Interrupt-Pending) in each hart’s `mip` CSR. For a supervisor-level domain, the interrupt signals appear as bit SEIP (Supervisor External Interrupt-Pending) in each hart’s `mip` and `sip` CSRs. Each interrupt signal may be arbitrarily delayed traveling from the APLIC to the proper hart. At the APLIC, each interrupt signal to a hart is derived from the IE field of register `domaincfg` and the current state of the hart’s IDC structure in the memory-mapped control region for the domain. If either `domaincfg`.IE = 0 or interrupt delivery to the hart is disabled by the `idelivery` register (`idelivery` \= 0), the interrupt signal is held de-asserted. When `domaincfg`.IE = 1 and interrupt delivery is enabled (`idelivery` \= 1), the interrupt signal is asserted whenever either register `iforce` or `topi` is not zero. Due to likely delay in the communication between an APLIC and a hart, it may happen that an external interrupt trap is taken, yet no interrupt is pending and enabled for the hart when a read of the hart’s `claimi` register actually occurs. In such a circumstance, the interrupt identity reported by the claim will be zero, resulting in an apparent _spurious interrupt_from the APLIC. Portable software must be prepared for the possibility of spurious interrupts at the APLIC, which can safely be ignored and should be rare. For testing purposes, a spurious interrupt can be triggered for a hart by setting an IDC structure’s `iforce` register to 1. A trap handler solely for external interrupts via an APLIC could be written roughly as follows: | save processor registers | | ------------------------------------------------------------------- | | i = read register claimi from the hart’s IDC structure at the APLIC | | i = i>>16 | | call the interrupt handler for external interrupt (minor identity) | | restore processor registers | | return from trap | To account for spurious interrupts, this pseudocode assumes there is an interrupt handler for "external interrupt 0" which does nothing. ### [](#4-1-9-interrupt-forwarding-by-msis)4.1.9\. Interrupt forwarding by MSIs In MSI delivery mode (`domaincfg`.DM = 1), an interrupt domain forwards interrupts to target harts by MSIs. An MSI is sent for a specific source only when the source’s corresponding pending and enable bits are both one and the IE field of register `domaincfg` is also one. If and when an MSI is sent, the source’s interrupt pending bit is cleared. #### [](#AdvPLIC-MSIAddrs)4.1.9.1\. Addresses and data for outgoing MSIs To forward interrupts by MSIs, an APLIC must know the MSI target address for each hart. For any given system, these addresses are fixed and should be hardwired into the APLIC if possible. However, some APLIC implementations may require that software supply the MSI target addresses. In that case, the root domain’s registers `mmsiaddrcfg`, `mmsiaddrcfgh`, `smsiaddrcfg`, and `smsiaddrcfgh` ([4.1.5.3\. Machine MSI address configuration (mmsiaddrcfg and mmsiaddrcfgh)](#AdvPLIC-reg-mmsiaddrcfg)and [4.1.5.4\. Supervisor MSI address configuration (smsiaddrcfg and smsiaddrcfgh)](#AdvPLIC-reg-smsiaddrcfg)) may be used to configure the MSI addresses for all interrupt domains. Alternatively MSI addresses may be configured by some custom means outside this standard. If MSI target addresses must be configured by software, this should be done only from a suitably privileged execution mode, typically just once, early after system reset. For a machine-level interrupt domain, if MSI target addresses are determined by `mmsiaddrcfg` and `mmsiaddrcfgh`, then the address for an outgoing MSI for interrupt source is computed from those registers and from the Hart Index field of register `target[ ]` as follows: | g = (Hart Index>>LHXW) & (2HHXW \- 1) | | | ------------------------------------------ | --------------- | | h = Hart Index & (2LHXW \-1) | | | MSI address = ( Base PPN \| (g<<(HHXS+12)) | (h<>12. For a supervisor-level domain, if MSI target addresses are determined by the root domain’s configuration registers (`smsiaddrcfg` and others), then to construct the address for an outgoing MSI for interrupt source , the Hart Index from register `target[ ]` must first be converted into the index number that machine-level domains use for the same hart. (These numbers are often the same, but they may not be.) The address for the MSI is then computed using this machine-level hart index together with the Base PPN and LHXS values from `smsiaddrcfg` and `smsiaddrcfgh`, the other fields (HHXW, LHXW, and HHXS) from `mmsiaddrcfgh`, and the Guest Index from `target[ ]`, as follows: | g = (machine-level hart index>>LHXW) & (2HHXW \- 1) | | | | --------------------------------------------------- | --------- | ---------------- | | h = machine-level hart index & (2LHXW \- 1) | | | | MSI address = (Base PPN \| (g<<(HHXS + 12)) | (h<>12. The data for an outgoing MSI write is taken from the EIID field of `target[ ]`, zero-extended to 32 bits. An MSI’s 32-bit data is always written in little-endian byte order, regardless of the BE field of the domain’s `domaincfg`register. #### [](#4-1-9-2-special-consideration-for-level-sensitive-interrupt-sources)4.1.9.2\. Special consideration for level-sensitive interrupt sources As soon as a level-sensitive interrupt is forwarded by MSI, the APLIC clears the pending bit for the interrupt source and then ignores the source until its incoming signal has been de-asserted. Clearing the pending bit when an MSI is sent is obviously necessary to avoid a constant stream of repeated MSIs from the APLIC to the target hart for the same interrupt. However, after an interrupt service routine has addressed a cause found for the interrupt, the incoming interrupt wire might remain asserted at the APLIC for another reason, despite that the interrupt’s pending bit at the APLIC was cleared and will remain so without intervention from software. If the interrupt service routine then exits without further action, a continued interrupt from this source might never receive attention. To avoid dropping interrupts in this way, the interrupt service routine for a level-sensitive interrupt may do one of the following before exiting: The first option is to test whether the interrupt wire into the APLIC is still asserted, by reading the appropriate `in_clrip` register at the APLIC. If the incoming interrupt is still asserted, the body of the interrupt service routine may be repeated to find and address an additional interrupt cause before the source wire is tested again. Once the incoming wire is observed not asserted, the interrupt service routine may safely exit, as any new interrupt assertion will cause the pending bit to become set and a new MSI sent to the hart. A second option is for the interrupt service routine to write the APLIC’s source identity number for the interrupt to the domain’s `setipnum`register just before exiting. This will cause the interrupt’s pending bit to be set to one again if the source is still asserting an interrupt, but not if the source is not asserting an interrupt. #### [](#AdvPLIC-MSISync)4.1.9.3\. Synchronizing interactions between a hart and the APLIC When an APLIC sends an MSI to a hart, there is an unspecified travel delay before the MSI is observed at the hart’s IMSIC. Consequently, after an APLIC’s configuration is changed by writing to an APLIC register, harts may continue to see MSIs arrive from the APLIC from the time before the write, for an unspecified amount of time. It is sometimes necessary to know when no more of these late MSIs can arrive. For example, if a hart will be turned off ("powered down"), all interrupts directed to it must be redirected to other harts, which may involve reconfiguring one or more APLICs. Even after the APLICs are reconfigured, the hart still cannot be safely turned off until it is known no more MSIs are destined for it. The `genmsi` register ([4.1.5.15\. Generate MSI (genmsi)](#AdvPLIC-reg-genmsi)) exists to allow software to determine when all earlier MSIs have arrived at a hart. To use `genmsi` for this purpose, software can dedicate one external interrupt identity at each hart’s IMSIC interrupt file solely for APLIC synchronization. Assuming there are multiple harts, an APLIC’s `genmsi` register should also be protected by a standard mutual-exclusion lock. The following sequence can then be used to synchronize between an APLIC and a specific hart: 1. At the hart’s IMSIC, clear the pending bit for the specific minor interrupt identity used exclusively for APLIC synchronization. 2. Acquire the shared lock for the APLIC’s `genmsi` register. 3. Write `genmsi` to generate an MSI to the hart with interrupt identity . 4. Repeatedly read `genmsi` until bit Busy is false. 5. Release the lock for `genmsi`. 6. Repeatedly read the pending bit for minor interrupt identity at the hart’s IMSIC until it is found set. The loops of steps 4 and 6 are expected normally to succeed very quickly, often on the first or second attempt. When this sequence is complete, all earlier MSIs from the APLIC must also have arrived at the hart’s IMSIC. 2.1. Control and Status Registers (CSRs) Added to Harts ==================== ## [](#CSRs)2.1\. Control and Status Registers (CSRs) Added to Harts For each privilege level at which a RISC-V hart can take interrupt traps, the Advanced Interrupt Architecture adds CSRs for interrupt control and handling. ### [](#2-1-1-machine-level-csrs)2.1.1\. Machine-level CSRs [Table 1](#CSRs-M) lists both the CSRs added for machine level and existing machine-level CSRs whose size is changed by the Advanced Interrupt Architecture. Existing CSRs `mie`, `mip`, and `mideleg` are widened to 64 bits to support a total of 64 interrupt causes. For RV32, the _high-half_ CSRs listed in the table allow access to the upper 32 bits of registers `mideleg`, `mie`, `mvien`, `mvip`, and `mip`. The Advanced Interrupt Architecture requires that these high-half CSRs exist for RV32, but the bits they access may all be merely read-only zeros. CSRs `miselect` and `mireg` provide a window for accessing multiple registers beyond the CSRs in [Table 1](#CSRs-M). The value of `miselect` determines which register is currently accessible through alias CSR `mireg`. `miselect` is a **WARL** register, and it must support a minimum range of values depending on the implemented features. When an IMSIC is not implemented, `miselect` must be able to hold at least any 6-bit value in the range 0 to 0`x3F`. When an IMSIC is implemented, `miselect` must be able to hold any 8-bit value in the range 0 to 0`xFF`. The Advanced Interrupt Architecture makes use of these subranges of values for `miselect`: | 0x30-0x3F | major interrupt priorities | | --------- | ---------------------------------------- | | 0x70-0xFF | external interrupts (only with an IMSIC) | __Table 1\. Machine-level CSRs added or widened by the Advanced Interrupt Architecture.__ | Number | Privilege | Width | Name | Description | | ----------------------------------------------------- | --------------- | -------------- | ------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Machine-Level Window to Indirectly Accessed Registers | | | | | | 0x3500x351 | MRWMRW | XLENXLEN | miselect mireg | Machine indirect register selectMachine indirect register alias | | Machine-Level Interrupts | | | | | | 0x3040x3440x35C0xFB0 | MRWMRWMRWMRO | 6464MXLENMXLEN | mie mip mtopei mtopi | Machine interrupt-enable bitsMachine interrupt-pending bitsMachine top external interrupt (only with an IMSIC)Machine top interrupt | | Delegated and Virtual Interrupts for Supervisor Level | | | | | | 0x3030x3080x309 | MRWMRWMRW | 646464 | mideleg mvien mvip | Machine interrupt delegationMachine virtual interrupt enablesMachine virtual interrupt-pending bits | | Machine-Level High-Half CSRs (RV32 only) | | | | | | 0x3130x3140x3180x3190x354 | MRWMRWMRWMRWMRW | 3232323232 | midelegh mieh mvienh mviph miph | Upper 32 bits of mideleg (only with S-mode)Upper 32 bits of mieUpper 32 bits of mvien (only with S-mode)Upper 32 bits of mvip (only with S-mode)Upper 32 bits of mip | Values of `miselect` with the most-significant bit set (bit`XLEN - 1` \= `1`) are designated for custom use, presumably for accessing custom registers through `mireg`. If `XLEN` changes, the most-significant bit of `miselect` moves to the new position, retaining its value from before. An implementation is not required to support any custom values for `miselect`. Other `miselect` values are reserved for other RISC-V extensions. | | RISC-V extension Smcsrind generalizes the mechanism of indirect register access provided by miselect and mireg. | | ------------------------------------------------------------------------------------------------------------------ | Normally, the range for external interrupts, 0x70-0xFF, is populated only when an IMSIC is implemented; else, attempts to access `mireg` when `miselect` is in this range cause an illegal instruction exception. The contents of the external-interrupts region are documented in[Incoming MSI Controller (IMSIC)](IMSIC.html#IMSIC) on the IMSIC. CSR `mtopei` also exists only when an IMSIC is implemented, so is documented in[Incoming MSI Controller (IMSIC)](IMSIC.html#IMSIC) along with the indirectly accessed IMSIC registers. CSR `mtopi` reports the highest-priority interrupt that is pending and enabled for machine level, as specified in [Machine top interrupt CSR (mtopi)](MSLevel.html#mtopi). When S-mode is implemented, CSRs `mvien` and `mvip` support interrupt filtering and virtual interrupts for supervisor level. These facilities are explained in [Interrupt filtering and virtual interrupts for supervisor level](MSLevel.html#virtIntrs-S). If extension Smcsrind is also implemented, then when `miselect` has a value in the range 0x30-0x3F or 0x70-0xFF, attempts to access alias CSRs `mireg2` through `mireg6` raise an illegal instruction exception. ### [](#2-1-2-supervisor-level-csrs)2.1.2\. Supervisor-level CSRs [Table 2](#CSRs-S) lists the supervisor-level CSRs that are added and existing CSRs that are widened to 64 bits, if the hart implements S-mode. The functions of these registers all match their machine-level counterparts. __Table 2\. Supervisor-level CSRs added or widened by the Advanced Interrupt Architecture.__ | Number | Privilege | Width | Name | Description | | -------------------------------------------------------- | ------------ | -------------- | -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | Supervisor-Level Window to Indirectly Accessed Registers | | | | | | 0x1500x151 | SRWSRW | XLENXLEN | siselect sireg | Supervisor indirect register selectSupervisor indirect register alias | | Supervisor-Level Interrupts | | | | | | 0x1040x1440x15C0xDB0 | SRWSRWSRWSRO | 6464SXLENSXLEN | sie sip stopei stopi | Supervisor interrupt-enable bitsSupervisor interrupt-pending bitsSupervisor top external interrupt (only with an IMSIC)Supervisor top interrupt | | Supervisor-Level High-Half CSRs (RV32 only) | | | | | | 0x1140x154 | SRWSRW | 3232 | sieh siph | Upper 32 bits of sieUpper 32 bits of sip | The space of registers accessible through the `siselect`/`sireg` window is separate from but parallels that of machine level, being for supervisor-level interrupts instead of machine-level interrupts. The subranges of values used for `siselect` are once again these: | 0x30-0x3F | major interrupt priorities | | --------- | ---------------------------------------- | | 0x70-0xFF | external interrupts (only with an IMSIC) | For maximum compatibility, it is recommended that `siselect` support at least a 9-bit range, `0` to `0x1FF`, regardless of whether an IMSIC exists. | | Because the VS CSR vsiselect ([2.1.3\. Hypervisor and VS CSRs](#hypervisor-vs-csrs)) always has at least 9 bits, and like other VS CSRs, vsiselect substitutes for siselect when executing in a virtual machine (VS-mode or VU-mode), implementing a smaller range forsiselect allows software to discover it is not running in a virtual machine. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Like `miselect`, values of `siselect` with the most-significant bit set (bit XLEN - 1 = 1) are designated for custom use. If XLEN changes, the most-significant bit of `siselect` moves to the new position, retaining its value from before. An implementation is not required to support any custom values for `siselect`. Other `siselect` values are reserved for other RISC-V extensions. | | At supervisor level, extension Sscsrind generalizes the mechanism of indirect register access provided by siselect and sireg, as well as the parallel at VS-level provided by vsiselect and vsireg, described in the next subsection. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Note that the widths of 'siselect' and 'sireg' are always the current XLEN rather than SXLEN. Hence, for example, if MXLEN = 64 and SXLEN = 32, then these registers are 64 bits when the current privilege mode is M (running RV64 code) but 32 bits when the privilege mode is S (RV32 code). CSR `stopei` is described with the IMSIC in [Incoming MSI Controller (IMSIC)](IMSIC.html#IMSIC). Register `stopi` reports the highest-priority interrupt that is pending and enabled for supervisor level, as specified in[Supervisor top interrupt CSR (stopi)](MSLevel.html#stopi). If extension Sscsrind is also implemented, then when `siselect` has a value in the range `0x30-0x3F` or `0x70-0xFF`, attempts to access alias CSRs `sireg2` through `sireg6` raise an illegal instruction exception (unless executing in a virtual machine, covered in the next section). ### [](#hypervisor-vs-csrs)2.1.3\. Hypervisor and VS CSRs If a hart implements the H extension, then the hypervisor and VS CSRs listed in [Table 3](#CSRs-hypervisor) are also either added or widened to 64 bits. The new hypervisor CSRs in the table (`hvien`, `hvictl` , `hviprio1`, and `hviprio2`) augment `hvip` for injecting interrupts into VS level. The use of these registers is covered in [Interrupts for Virtual Machines (VS Level)](VSLevel.html#VSLevel) on interrupts for virtual machines. The new VS CSRs (`vsiselect`, `vsireg`, `vstopei`, and `vstopi`) all match supervisor CSRs, and substitute for those supervisor CSRs when executing in a virtual machine (in VS-mode or VU-mode). CSR `vsiselect` is required to support at least a 9-bit range of `0` to `0x1FF`, whether or not an IMSIC is implemented. As with `siselect`, values of `vsiselect` with the most-significant bit set (bit XLEN - 1 = 1) are designated for custom use. If XLEN changes, the most-significant bit of `vsiselect` moves to the new position, retaining its value from before. Like `siselect` and `sireg`, the widths of `vsiselect` and `vsireg` are always the current XLEN rather than VSXLEN. Hence, for example, if HSXLEN = 64 and VSXLEN = 32, then these registers are 64 bits when accessed by a hypervisor in HS-mode (running RV64 code) but 32 bits for a guest OS in VS-mode (RV32 code). __Table 3\. Hypervisor and VS CSRs added or widened by the Advanced Interrupt Architecture. (Parameter HSXLEN is just another name for SXLEN for hypervisor-extended S-mode).__ | Number | Privilege | Width | Name | Description | | -------------------------------------------------------------------- | --------------------- | ----------------- | ----------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Delegated and Virtual Interrupts, Interrupt Priorities, for VS Level | | | | | | 0x6030x6080x6090x6450x6460x647 | HRWHRWHRWHRWHRWHRW | 6464HSXLEN646464 | hideleg hvien hvictl hvip hviprio1 hviprio2 | Hypervisor interrupt delegationHypervisor virtual interrupt enablesHypervisor virtual interrupt controlHypervisor virtual interrupt-pending bitsHypervisor VS-level interrupt prioritiesHypervisor VS-level interrupt priorities | | VS-Level Window to Indirectly Accessed Registers | | | | | | 0x2500x251 | HRWHRW | XLENXLEN | vsiselect vsireg | Virtual supervisor indirect register selectVirtual supervisor indirect register alias | | VS-Level Interrupts | | | | | | 0x2040x2440x25C 0xEB0 | HRWHRWHRW HRO | 6464VSXLEN VSXLEN | vsie vsip vstopei vstopi | Virtual supervisor interrupt-enable bitsVirtual supervisor interrupt-pending bitsVirtual supervisor top external interrupt (only with an IMSIC)Virtual supervisor top interrupt | | Hypervisor and VS-Level High-Half CSRs (RV32 only) | | | | | | 0x6130x6180x6550x6560x6570x2140x254 | HRWHRWHRWHRWHRWHRWHRW | 32323232323232 | hidelegh hvienh hviph hviprio1h hviprio2h vsieh vsiph | Upper 32 bits of hidelegUpper 32 bits of hvienUpper 32 bits of hvipUpper 32 bits of hviprio1Upper 32 bits of hviprio2Upper 32 bits of vsieUpper 32 bits of vsip | The space of registers selectable by `vsiselect` is more limited than for machine and supervisor levels: | 0x030-0x03F | inaccessible | | ----------- | ------------------------------------------------- | | 0x070-0x0FF | external interrupts (IMSIC only), or inaccessible | Other `vsiselect` values are reserved for other RISC-V extensions. For alias CSRs `sireg` and `vsireg`, the H extension’s usual rules for when to raise a virtual instruction exception (based on whether an instruction is _HS-qualified_) are not applicable. The rules given in this section for `sireg` and `vsireg` apply instead, unless overridden by the requirements of [2.1.5\. Access control by the state-enable CSRs](#CSRs-stateen), which take precedence over this section when extension Smstateen is also implemented. A virtual instruction exception is raised for attempts from VS-mode or VU-mode to directly access `vsireg`, or attempts from VU-mode to access `sireg`. When `vsiselect` has the number of an _inaccessible_ register, attempts from M-mode or HS-mode to access `vsireg` raise an illegal instruction exception, and attempts from VS-mode to access `sireg` (really `vsireg`) raise a virtual instruction exception. | | Requiring a range of 0-0x1FF for vsiselect, even though most or all of the space is reserved or inaccessible, permits a hypervisor to emulate indirectly accessed registers in the implemented range, including registers that are not currently defined but may be standardized in the future. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The indirectly accessed registers for external interrupts (numbers 0x70-0xFF) are accessible only when field VGEIN of `hstatus` is the number of an implemented guest external interrupt, not zero. If VGEIN is not the number of an implemented guest external interrupt (including the case when no IMSIC is implemented), then all indirect register numbers in the ranges 0x030-0x03F and 0x070-0x0FF designate an inaccessible register at VS level. Along the same lines, when `hstatus.VGEIN` is not the number of an implemented guest external interrupt, attempts from M-mode or HS-mode to access CSR `vstopei` raise an illegal instruction exception, and attempts from VS-mode to access `stopei` raise a virtual instruction exception. If extension Sscsrind is also implemented, then when `vsiselect` has a value in the range 0x30-0x3F or 0x70-0xFF, attempts from M-mode or HS-mode to access alias CSRs `vsireg2` through `vsireg6` raise an illegal instruction exception, and attempts from VS-mode to access `sireg2` through `sireg6` raise a virtual instruction exception. ### [](#2-1-4-virtual-instruction-exceptions)2.1.4\. Virtual instruction exceptions Following the default rules for the H extension, attempts from VS-mode to directly access a hypervisor or VS CSR other than `vsireg`, or from VU-mode to access any supervisor-level CSR (including hypervisor and VS CSRs) other than `sireg` or `vsireg`, usually raise not an illegal instruction exception but instead a virtual instruction exception. For details, see the H extension documentation. Instructions that read/write CSR `stopei` or `vstopei` are considered to be _HS-qualified_ unless all of following are true: the hart has an IMSIC, extension Smstateen is implemented, and bit 58 of `mstateen0` is zero. (See the next section, [2.1.5\. Access control by the state-enable CSRs](#CSRs-stateen), about `mstateen0`.) For `sireg` and `vsireg`, see both the previous section, [2.1.3\. Hypervisor and VS CSRs](#hypervisor-vs-csrs), and the next, [2.1.5\. Access control by the state-enable CSRs](#CSRs-stateen), for when a virtual instruction exception is required instead of an illegal instruction exception. ### [](#CSRs-stateen)2.1.5\. Access control by the state-enable CSRs If extension Smstateen is implemented together with the Advanced Interrupt Architecture (AIA), three bits of state-enable register `mstateen0` control access to AIA-added state from privilege modes less privileged than M-mode: | bit 60 CSRIND: CSRs siselect, sireg, vsiselect, and vsireg | | ---------------------------------------------------------------------------------------- | | bit 59 AIA: all other state added by the AIA and not controlled by bits CSRIND and IMSIC | | bit 58 IMSIC: all IMSIC state, including CSRs stopei and vstopei | If one of these bits is zero in `mstateen0`, an attempt to access the corresponding state from a privilege mode less privileged than M-mode results in an illegal instruction trap. As always, the state-enable CSRs do not affect the accessibility of any state when in M-mode, only in less privileged modes. For more explanation, see the documentation for extension Smstateen. The AIA bit controls access to AIA CSRs `siph`, `sieh`, `stopi`, `hidelegh`, `hvien`/`hvienh`, `hviph`, `hvictl`, `hviprio1`/`hviprio1h`, `hviprio2`/`hviprio2h`, `vsiph`, `vsieh`, and `vstopi`, as well as to the supervisor-level interrupt priorities accessed through `siselect` \+ `sireg` (the `iprio` array of [Configuring priorities of major interrupts at supervisor level](MSLevel.html#intrPrios-S)). The IMSIC bit is implemented in `mstateen0` only if the hart has an IMSIC. If the H extension is also implemented, this bit does not affect the behavior or accessibility of hypervisor CSRs `hgeip` and `hgeie`, or field VGEIN of `hstatus`. In particular, guest external interrupts from an IMSIC continue to be visible to HS-mode in `hgeip` even when `mstateen0`.IMSIC is zero. | | An earlier, pre-ratification draft of Smstateen said that when mstateen0.IMSIC is zero, registers hgeip and hgeie and field VGEIN of hstatus are all read-only zeros. That effect is no longer correct. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If the hart does not have an IMSIC, the IMSIC bit of `mstateen0` is read-only zero, but Smstateen has no effect on attempts to access the nonexistent IMSIC state. | | This means in particular that, when the hart does not have an IMSIC, the following raise a virtual instruction exception as described in [Table 3](#CSRs-hypervisor), not an illegal instruction exception, despite that mstateen0.IMSIC is zero: attempts from VS-mode to access sireg (really vsireg) while vsiselect has a value in the range 0x70–0xFF; and attempts from VS-mode to access stopei (really vstopei). | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If the CSRIND bit of `mstateen0` is one, then regardless of any other `mstateen` bits (including the AIA and IMSIC bits of `mstateen0`), a virtual instruction exception is raised as described in [2.1.3\. Hypervisor and VS CSRs](#hypervisor-vs-csrs) for all attempts from VS-mode or VU-mode to directly access `vsireg`, and for all attempts from VU-mode to access `sireg`. This behavior is overridden only when `mstateen0`.CSRIND is zero. If the H extension is implemented, the same three bits are defined also in hypervisor CSR `hstateen0` but concern only the state potentially accessible to a virtual machine executing in privilege modes VS and VU: | bit 60 CSRIND: CSRs siselect and sireg (really vsiselect and vsireg) | | -------------------------------------------------------------------------------------------- | | bit 59 AIA: CSRs siph and sieh (RV32 only) and stopi (really vsiph, vsieh, and vstopi) | | bit 58 IMSIC: all state of IMSIC guest interrupt files, including CSR stopei(really vstopei) | If one of these bits is zero in `hstateen0`, and the same bit is one in `mstateen0`, then an attempt to access the corresponding state from VS or VU-mode raises a virtual instruction exception. (But note that, for high-half CSRs `siph` and `sieh`, this applies only when XLEN = 32\. When XLEN > 32, an attempt to access `siph` or `sieh` raises an illegal instruction exception as usual, not a virtual instruction exception.) If the CSRIND bit is one in `mstateen0` but is zero in `hstateen0`, then all attempts from VS or VU-mode to access `siselect` or `sireg` raise a virtual instruction exception, not an illegal instruction exception, regardless of the value of `vsiselect` or any other `mstateen` bits. The IMSIC bit is implemented in `hstateen0` only if the hart has an IMSIC. Furthermore, even with an IMSIC, `hstateen0`.IMSIC may (or may not) be read-only zero if the IMSIC has no _guest interrupt files_ for guest external interrupts ([Incoming MSI Controller (IMSIC)](IMSIC.html#IMSIC)). When this bit is zero (whether read-only zero or set to zero), a virtual machine is prevented from accessing the hart’s IMSIC the same as when `hstatus.VGEIN` \= 0. Extension Ssstateen is defined as the supervisor-level view of Smstateen. Therefore, the combination of Ssaia and Ssstateen incorporates the bits defined above for `hstateen0` but not those for `mstateen0`, since machine-level CSRs are not visible to supervisor level. 3.1. Incoming MSI Controller (IMSIC) ==================== ## [](#IMSIC)3.1\. Incoming MSI Controller (IMSIC) An Incoming MSI Controller (IMSIC) is an optional RISC-V hardware component that is closely coupled with a hart, one IMSIC per hart. An IMSIC receives and records incoming message-signaled interrupts (MSIs) for a hart, and signals to the hart when there are pending and enabled interrupts to be serviced. An IMSIC has one or more memory-mapped registers in the machine’s address space for receiving MSIs. Aside from those memory-mapped registers, software interacts with an IMSIC primarily through several RISC-V CSRs at the attached hart. ### [](#IMSIC-intrFilesAndIdents)3.1.1\. Interrupt files and interrupt identities In a RISC-V system, MSIs are directed not just to a specific hart but to a specific privilege level of a specific hart, such as machine or supervisor level. Furthermore, when a hart implements the H extension, an IMSIC may optionally allow MSIs to be directed to a specific virtual hart at virtual supervisor level (VS level). For each privilege level and each virtual hart to which MSIs may be directed at a hart, the hart’s IMSIC contains a separate _interrupt file_. Assuming a hart implements supervisor mode, its IMSIC has at least two interrupt files, one for machine level and the other for supervisor level. When a hart also implements the H extension, its IMSIC may have additional interrupt files for virtual harts, called_guest interrupt files_. The number of guest interrupt files an IMSIC has for virtual harts is exactly _GEILEN_, the number of supported guest external interrupts, as defined by the H extension. Each individual interrupt file consists mainly of two arrays of bits of the same size, one array for recording MSIs that have arrived but are not yet serviced (interrupt-pending bits), and the other array for specifying which interrupts the hart will currently accept (interrupt-enable bits). Each bit position in the two arrays corresponds with a different interrupt _identity number_ by which MSIs from different sources are distinguished at an interrupt file. Because an IMSIC is the external interrupt controller for a hart, an interrupt file’s interrupt identities become the _minor identities_ for external interrupts at the attached hart. The number of interrupt identities supported by an interrupt file (and hence the number of active bits in each array) is one less than a multiple of 64, and may be a minimum of 63 and a maximum of 2047. | | Platform standards may increase the minimum number of interrupt identities that must be implemented by each interrupt file. | | ------------------------------------------------------------------------------------------------------------------------------ | When an interrupt file supports distinct interrupt identities, valid identity numbers are between 1 and inclusive. The identity numbers within this range are said to be implemented by the interrupt file; numbers outside this range are not implemented. The number zero is never a valid interrupt identity. IMSIC hardware does not assume any connection between the interrupt identity numbers at one interrupt file and those at another interrupt file. Software is commonly expected to assign the same interrupt identity number to different MSI sources at different interrupt files, without coordination across interrupt files. Thus the total number of MSI sources that can be separately distinguished within a system is potentially the product of the number of interrupt identities at a single interrupt file times the total number of interrupt files in the system, over all harts. It is not necessarily the case that all interrupt files in a system are the same size (implement the same number of interrupt identities). For a given hart, the interrupt files for guest external interrupts must all be the same size, but the interrupt files at machine level and at supervisor level may differ in size from those of guest external interrupts, and from each other. Likewise, the interrupt files of different harts may be different sizes. A platform might provide a means for software to configure the number of interrupt files in an IMSIC and/or their sizes, such as by allowing a smaller interrupt file at machine level to be traded for a larger one at supervisor level, or vice versa, for example. Any such configurability is outside the scope of this specification. It is recommended, however, that only machine level be given the power to change the number and sizes of interrupt files in an IMSIC. ### [](#MSIEncoding)3.1.2\. MSI encoding Established standards (in particular, for PCI and PCI Express) dictate that an individual message-signaled interrupt (MSI) from a device takes the form of a naturally aligned 32-bit write by the device, with the address and value both configured at the device (or device controller) by software. Depending on the versions of the standards to which a device or controller conforms, the address might be restricted to the lower 4-GiB (32-bit) range, and the value written might be limited to a 16-bit range, with the upper 16 bits always being zeros. When RISC-V harts have IMSICs, an MSI from a device is normally sent directly to an individual hart that was selected by software to handle the interrupt (presumably based on some interrupt affinity policy). An MSI is directed to a specific privilege level, or to a specific virtual hart, via the corresponding interrupt file that exists in the receiving hart’s IMSIC. The MSI write address is the physical address of a particular word-size register that is physically connected to the target interrupt file. The MSI write data is simply the identity number of the interrupt to be made pending in that interrupt file (becoming eventually the minor identity for an external interrupt to the attached hart). By configuring an MSI’s address and data at a device, system software fully controls: (a) which hart receives a particular device interrupt, (b) the target privilege level or virtual hart, and (c) the identity number that represents the MSI in the target interrupt file. Elements a and b are determined by which interrupt file is targeted by the MSI address, while element c is communicated by the MSI data. | | As the maximum interrupt identity number an IMSIC can support is 2047, a 16-bit limit on MSI data values presents no problem. | | -------------------------------------------------------------------------------------------------------------------------------- | When the H extension is implemented and a device is being managed directly by a guest operating system, MSI addresses from the device are initially guest physical addresses, as they are configured at the device by the guest OS. These guest addresses must be translated by an IOMMU, which gets configured by the hypervisor to redirect those MSIs to the interrupt files for the correct guest external interrupts. For more on this topic, see [IOMMU Support for MSIs to Virtual Machines](IOMMU.html#IOMMU). ### [](#3-1-3-interrupt-priorities)3.1.3\. Interrupt priorities Within a single interrupt file, interrupt priorities are determined directly from interrupt identity numbers. Lower identity numbers have higher priority. | | Because MSIs give software complete control over the assignment of identity numbers in an interrupt file, software is free to select identity numbers that reflect the relative priorities desired for interrupts. It is true that software could adjust interrupt priorities more dynamically if interrupt files included an array of priority numbers to assign to each interrupt identity. However, we believe that such additional flexibility would not be utilized often enough to justify the extra hardware expense. In fact, for many systems currently employing MSIs, it is common practice for software to ignore interrupt priorities entirely and act as though all interrupts had equal priority. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | An interrupt file’s lowest identity numbers have been given the highest priorities, not the reverse order, because it is only for the highest-priority interrupts that priority order may need to be carefully managed, yet it is the low-numbered identities, 1 through 63 (or perhaps 1 through 127), that are guaranteed to exist across all systems. Consider, for example, that an interrupt file’s highest-priority interrupt—presumably the most time-critical—is always identity number 1\. If priority order were reversed, the highest-priority interrupt would have different identity numbers on different machines, depending on how many identities are implemented by interrupt files. The ability for software to assign fixed identity numbers to the highest-priority interrupts is considered worth any discomfort that may be felt from interrupt priorities being the reverse of the natural number order. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#3-1-4-reset-and-revealed-state)3.1.4\. Reset and revealed state Upon reset of an IMSIC, all the state of its interrupt files becomes valid and consistent but otherwise UNSPECIFIED, except possibly for the `eidelivery` register of machine-level and supervisor-level interrupt files, as specified in[3.1.8.1\. External interrupt delivery enable register (eidelivery)](#IMSIC-reg-eidelivery). If an IMSIC contains a supervisor-level interrupt file and software at the attached hart enables S-mode that was previously disabled (e.g. by changing bit S of CSR `misa` from zero to one), all state of the supervisor-level interrupt file is valid and consistent but otherwise UNSPECIFIED. Likewise, if an IMSIC contains guest interrupt files and software at the attached hart enables the H extension that was previously disabled (e.g. by changing bit H of `misa` from zero to one), all state of the IMSIC’s guest interrupt files is valid and consistent but otherwise UNSPECIFIED. ### [](#IMSIC-memRegion)3.1.5\. Memory region for an interrupt file Each interrupt file in an IMSIC has one or two memory-mapped 32-bit registers for receiving MSI writes. These memory-mapped registers are located within a naturally aligned 4-KiB region (a page) of physical address space that exists for the interrupt file, i.e., one page per interrupt file. The layout of an interrupt-file’s memory region is: | offset | size | register name | | ------ | ------- | ------------- | | 0x000 | 4 bytes | seteipnum\_le | | 0x004 | 4 bytes | seteipnum\_be | All other bytes in an interrupt file’s 4-KiB memory region are reserved and must be implemented as read-only zeros. Only naturally aligned 32-bit simple reads and writes are supported within an interrupt file’s memory region. Writes to read-only bytes are ignored. For other forms of accesses (other sizes, misaligned accesses, or AMOs), an IMSIC implementation should preferably report an access fault or bus error but must otherwise ignore the access. If is an implemented interrupt identity number, writing value in little-endian byte order to `seteipnum_le` (Set External Interrupt-Pending bit by Number, Little-Endian) causes the pending bit for interrupt to be set to one. A write to `seteipnum_le` is ignored if the value written is not an implemented interrupt identity number in little-endian byte order. For systems that support big-endian byte order, if is an implemented interrupt identity number, writing value in big-endian byte order to `seteipnum_be` (Set External Interrupt-Pending bit by Number, Big-Endian) causes the pending bit for interrupt to be set to one. A write to `seteipnum_be` is ignored if the value written is not an implemented interrupt identity number in big-endian byte order. Systems that support only little-endian byte order may choose to ignore all writes to `seteipnum_be`. In most systems, `seteipnum_le` is the write port for MSIs directed to this interrupt file. For systems built mainly for big-endian byte order, `seteipnum_be` may serve as the write port for MSIs directed to this interrupt file from some devices. A read of `seteipnum_le` or `seteipnum_be` returns zero in all cases. When not ignored, writes to an interrupt file’s memory region are guaranteed to be reflected in the interrupt file eventually, but not necessarily immediately. For a single interrupt file, the effects of multiple writes (stores) to its memory region, though arbitrarily delayed, always occur in the same order as the _global memory order_ of the stores as defined by the RISC-V Unprivileged ISA. | | In most circumstances, any delay between the completion of a write to an interrupt file’s memory region and the effect of the write on the interrupt file is indistinguishable from other delays in the memory system. However, if a hart writes to a seteipnum\_le or seteipnum\_be register of its own IMSIC, then a delay between the completion of the store instruction and the consequent setting of an interrupt-pending bit in the interrupt file may be visible to the hart. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#IMSIC-systemMemRegions)3.1.6\. Arrangement of the memory regions of multiple interrupt files Each interrupt file that an IMSIC implements has its own memory region as described in the previous section, occupying exactly one 4-KiB page of machine address space. When practical, the memory pages of the machine-level interrupt files of all IMSICs should be located together in one part of the address space, and the memory pages of all supervisor-level and guest interrupt files should similarly be located together in another part of the address space, according to the rules below. | | The main reason for separating the machine-level interrupt files from the other interrupt files in the address space is so harts that implement physical memory protection (PMP) can grant supervisor-level access to all supervisor-level and guest interrupt files using only a single PMP table entry. If the memory pages for machine-level interrupt files are instead interleaved with those of lower-privilege interrupt files, the number of PMP table entries needed for granting supervisor-level access to all non-machine-level interrupt files could equal the number of harts in the system. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If a machine’s construction dictates that harts be subdivided into groups, with each group relegated to its own portion of the address space, then the best that can be achieved is to locate together the machine-level interrupt files of each group of harts separately, and likewise locate together the supervisor-level and guest interrupt files of each group of harts separately. This situation is further addressed later below. | | A system may divide harts into groups in the address space because each group exists on a separate chip (or chiplet in a multi-chip module), and weaving together the address spaces of the multiple chips is impractical. In that case, granting supervisor-level access to all non-machine-level interrupt files takes one PMP table entry per group. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | For the purpose of locating the memory pages of interrupt files in the address space, assume each hart (or each hart within a group) has a unique hart number that may or may not be related to the unique hart identifiers ("hart IDs") that the Privileged Architecture assigns to harts. For convenient addressing, the memory pages of all machine-level interrupt files (or all those of a single group of harts) should be arranged so that the address of the machine-level interrupt file for hart number is given by the formula for some integer constants and . If the largest hart number is , let , the number of bits needed to represent any hart number. Then the base address should be aligned to a address boundary, so always equals | , where the vertical bar (|) represents bitwise logical OR. The smallest that can be is 12, with being the size of one 4-KiB page. If , the start of the memory page for each machine-level interrupt file is aligned not just to a 4-KiB page but to a stricter address boundary. Within the \-size address range through , every 4-KiB page that is not occupied by a machine-level interrupt file should be filled with 32-bit words of read-only zeros, such that any read of an aligned word returns zero and any write to an aligned word is ignored. The memory pages of all supervisor-level interrupt files (or all those of a single group of harts) should similarly be arranged so that the address of the supervisor-level interrupt file for hart number is for some integer constants and , with the base address being aligned to a address boundary. If an IMSIC implements guest interrupt files, the memory pages for the IMSIC’s supervisor-level interrupt file and for its guest interrupt files should be contiguous, starting with the supervisor-level interrupt file at the lowest address and followed by the guest interrupt files, ordered by guest interrupt number. Schematically, the memory pages should be ordered contiguously as S, , , , … where S is the page for the supervisor-level interrupt file and each is the page for the interrupt file of guest interrupt number . Consequently, the smallest that constant can be is , recalling that GEILEN for each IMSIC is the number of guest interrupt files the IMSIC implements. Within the \-size address range through , every 4-KiB page that is not occupied by an interrupt file (supervisor-level or guest) should be filled with 32-bit words of read-only zeros. When a system divides harts into groups, each in its own separate portion of the address space, the memory page addresses of interrupt files should follow the formulas for machine-level interrupt files, and for supervisor-level interrupt files, with being a _group number_, being a hart number relative to the group, and being another integer constant but usually much larger. If the largest group number is , let , the number of bits needed to represent any group number. Besides being multiples of and respectively, and should be chosen so & and & where an ampersand (&) represents bitwise logical AND. This ensures that always equals ( ) | | ( ), and always equals ( ) | | ( ). Infilling with read-only-zero pages is expected only within each group, not between separate groups. Specifically, if is any integer between 0 and inclusive, then within the address ranges, through , and through , pages not occupied by an interrupt file should be read-only zeros. See also [Addresses and data for outgoing MSIs](AdvPLIC.html#AdvPLIC-MSIAddrs) for the default algorithms an Advanced PLIC may use to determine the destination addresses of outgoing MSIs, which should be the addresses of IMSIC interrupt files. ### [](#3-1-7-csrs-for-external-interrupts-via-an-imsic)3.1.7\. CSRs for external interrupts via an IMSIC Software accesses a hart’s IMSIC primarily through the CSRs introduced in [Control and Status Registers (CSRs) Added to Harts](CSRs.html#CSRs). There is a separate set of CSRs for each implemented privilege level that can receive interrupts. The machine-level CSRs interact with the IMSIC’s machine-level interrupt file, while, if supervisor mode is implemented, the supervisor-level CSRs interact with the IMSIC’s supervisor-level interrupt file. When an IMSIC has guest interrupt files, the VS CSRs interact with a single guest interrupt file, selected by the VGEIN field of CSR `hstatus`. For machine level, the relevant CSRs are `miselect`, `mireg`, and `mtopei`. When supervisor mode is implemented, the set of supervisor-level CSRs matches those of machine level: `siselect`, `sireg`, and `stopei`. And when the H extension is implemented, there are three corresponding VS CSRs: `vsiselect`, `vsireg`, and `vstopei`. As explained in [Control and Status Registers (CSRs) Added to Harts](CSRs.html#CSRs), registers `miselect` and `mireg` provide indirect access to additional machine-level registers. Likewise for supervisor-level `siselect` and `sireg`, and VS-level `vsiselect` and `vsireg` . In each case, a value of the **_\*iselect_** _CSR_ (`miselect`, `siselect` , or `vsiselect)`) in the range 0x70-0xFF selects a register of the corresponding IMSIC interrupt file, either the machine-level interrupt file (`miselect`), the supervisor-level interrupt file (`siselect`), or a guest interrupt file (`vsiselect`). Interrupt files at each level act identically. For a given privilege level, values of the `*iselect` CSR in the range `0x70-0xFF` select these registers of the corresponding interrupt file: | 0x70 | eidelivery | | ---- | ----------- | | 0x72 | eithreshold | | 0x80 | eip0 | | 0x81 | eip1 | | …​ | …​ | | 0xBF | eip63 | | 0xC0 | eie0 | | 0xC1 | eie1 | | …​ | …​ | | 0xFF | eie63 | Register numbers 0x71 and 0x73-0x7F are reserved. When an `_*iselect_` _CSR_ has one of these values, reads from the matching `_*ireg_` _CSR_ (`mireg`, `sireg`, or `vsireg`) return zero, and writes to the `_*ireg_` _CSR_ are ignored. (For `vsiselect` and `vsireg`, all accesses depend on `hstatus`.VGEIN being the valid number of a guest interrupt file.) Registers `eip0` through `eip63` contain the pending bits for all implemented interrupt identities, and are collectively called the `_eip_` _array_. Registers `eie0` through `eie63` contain the enable bits for the same interrupt identities, and are collectively called the `_eie_` _array_. The indirectly accessed interrupt-file registers and CSRs `mtopei`, `stopei`, and `vstopei` are all documented in more detail in the next two sections. ### [](#3-1-8-indirectly-accessed-interrupt-file-registers)3.1.8\. Indirectly accessed interrupt-file registers This section describes the registers of an interrupt file that are accessed indirectly through a `_*iselect_` _CSR_ (`miselect`, `siselect`, or `vsiselect`) and its partner `_*ireg_` _CSR_ (`mireg`, `sireg`, or `vsireg`). The width of these indirect accesses is always the current XLEN, 32 bits for RV32 code, or 64 bits for RV64 code. #### [](#IMSIC-reg-eidelivery)3.1.8.1\. External interrupt delivery enable register (`eidelivery`) `eidelivery` is a **WARL** register that controls whether interrupts from this interrupt file are delivered from the IMSIC to the attached hart so they appear as a pending external interrupt in the hart’s `mip` or `hgeip` CSR. Register `eidelivery` may optionally also support the direct delivery of interrupts from a PLIC (Platform-Level Interrupt Controller) or APLIC (Advanced PLIC) to the attached hart. Three possible values are currently defined for `eidelivery`: | 0 = | Interrupt delivery is disabled | | ------------ | ------------------------------------------------------------- | | 1 = | Interrupt delivery from the interrupt file is enabled | | 0x40000000 = | Interrupt delivery from a PLIC or APLIC is enabled (optional) | If `eidelivery` supports value 0x40000000, then a specific PLIC or APLIC in the system may act as an alternate external interrupt controller for the attached hart at the same privilege level as this interrupt file. When `eidelivery` is 0x40000000, the interrupt file functions the same as though `eidelivery` is 0, and the PLIC or APLIC replaces the interrupt file in supplying pending external interrupts at this privilege level at the hart. Guest interrupt files do not support value 0x40000000 for `eidelivery`. Reset initializes `eidelivery` to 0x40000000 if that value is supported; otherwise, `eidelivery` has an UNSPECIFIED valid value (0 or 1) after reset. | | eidelivery value 0x40000000 supports system software that is oblivious to IMSICs and assumes instead that the external interrupt controller is a PLIC or APLIC. Such software may exist either because it predates the existence of IMSICs or because bypassing IMSICs is believed to reduce programming effort. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `eidelivery` register affects only whether an external interrupt appears in a hart’s `*ip` register (MEI or SEI in `mip` or `sip`, or a bit in `hgeip`) and what the source of such an interrupt may be (either the interrupt file or a separate external interrupt controller such as an APLIC). It has no effect on other state within the interrupt file, or on any `*topei` CSR (`mtopei`, `stopei`, or `vstopei`). #### [](#3-1-8-2-external-interrupt-enable-threshold-register-eithreshold)3.1.8.2\. External interrupt enable threshold register (`eithreshold`) `eithreshold` is a **WLRL** register that determines the minimum interrupt priority (maximum interrupt identity number) allowing an interrupt to be signaled from this interrupt file to the attached hart. If is the maximum implemented interrupt identity number for this interrupt file,`eithreshold` must be capable of holding all values between 0 and , inclusive. When `eithreshold` is a nonzero value , interrupt identities and higher do not contribute to signaling interrupts, as though those identities were not enabled, regardless of the settings of their corresponding interrupt-enable bits in the `eie` array. When `eithreshold` is zero, all enabled interrupt identities contribute to signaling interrupts from the interrupt file. #### [](#3-1-8-3-external-interrupt-pending-registers-eip0-eip63)3.1.8.3\. External interrupt-pending registers (`eip0`\-`eip63`) When the current XLEN = 32, register `eip` contains the pending bits for interrupts with identity numbers through . For an implemented interrupt identity within that range, the pending bit for interrupt is bit of `eip` . When the current XLEN = 64, the odd-numbered registers `eip1`, `eip3`, … `eip63` do not exist. In that case, if the `*iselect` CSR is an odd value in the range 0x81–0xBF, an attempt to access the matching `*ireg` CSR raises an illegal instruction exception, unless done in VS-mode, in which case it raises a virtual instruction exception. For even , register `eip` contains the pending bits for interrupts with identity numbers through . For an implemented interrupt identity within that range, the pending bit for interrupt is bit of `eip` . Bit positions in a valid `eip` register that don’t correspond to a supported interrupt identity (such as bit 0 of `eip0`) are read-only zeros. #### [](#3-1-8-4-external-interrupt-enable-registers-eie0-eie63)3.1.8.4\. External interrupt-enable registers (`eie0`\-`eie63`) When the current XLEN = 32, register `eie` contains the enable bits for interrupts with identity numbers through . For an implemented interrupt identity within that range, the enable bit for interrupt is bit ( ) of `eie` . When the current XLEN = 64, the odd-numbered registers `eie1`, `eie3`, … `eie63` do not exist. In that case, if the `*iselect` CSR is an odd value in the range `0xC1`–`0xFF`, an attempt to access the matching `*ireg` CSR raises an illegal instruction exception, unless done in VS-mode, in which case it raises a virtual instruction exception. For even , register`eie` contains the enable bits for interrupts with identity numbers through . For an implemented interrupt identity within that range, the enable bit for interrupt is bit ( ) of `eie` . Bit positions in a valid `eie` register that don’t correspond to a supported interrupt identity (such as bit 0 of `eie0`) are read-only zeros. ### [](#3-1-9-top-external-interrupt-csrs-mtopei-stopei-vstopei)3.1.9\. Top external interrupt CSRs (`mtopei`, `stopei`, `vstopei`) CSR `mtopei` interacts directly with an IMSIC’s machine-level interrupt file. If supervisor mode is implemented, CSR `stopei` interacts directly with the supervisor-level interrupt file. And if the H extension is implemented and field VGEIN of `hstatus` is the number of an implemented guest interrupt file, `vstopei` interacts with the chosen guest interrupt file. The value of a `_*topei_` _CSR_ (`mtopei`, `stopei`, or `vstopei`) indicates the interrupt file’s current highest-priority pending-and-enabled interrupt that also exceeds the priority threshold specified by its `eithreshold` register if `eithreshold` is not zero. Interrupts with lower identity numbers have higher priorities. A read of a `*topei` CSR returns zero either if no interrupt is both pending in the interrupt file’s `eip` array and enabled in its `eie` array, or if `eithreshold` is not zero and no pending-and-enabled interrupt has an identity number less than the value of `eithreshold`. Otherwise, the value returned from a read of `*topei` has this format: | bits 26:16 | Interrupt identity | | ---------- | ------------------------------------- | | bits 10:0 | Interrupt priority (same as identity) | All other bit positions are zeros. The interrupt identity reported in a `*topei` CSR is the minor identity for an external interrupt at the hart. | | The redundancy in the value read from a \*topei CSR is consistent with the Advanced PLIC, which returns both an interrupt identity number and its priority in the same format as above, but with the two components being independent of one another. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The value of a `*topei` CSR is not affected by an interrupt file’s`eidelivery` register or by any of `mie`, `sie`, `hie`, `hgeie`, or `vsie`. A write to a `*topei` CSR _claims_ the reported interrupt identity by clearing its pending bit in the interrupt file. The value written is ignored; rather, the current readable value of the register determines which interrupt-pending bit is cleared. Specifically, when a `*topei` CSR is written, if the register value has interrupt identity in bits 26:16, then the interrupt file’s pending bit for interrupt is cleared. When a `*topei` CSR’s value is zero, a write to the register has no effect. If a read and write of a `*topei` CSR are done together by a single CSR instruction (CSRRW, CSRRS, or CSRRC), the value returned by the read indicates the pending bit that is cleared. | | It is almost always a mistake to write to a \*topei CSR without a simultaneous read to learn which interrupt was claimed. Note especially, if a read of a \*topei register and a subsequent write to the register are done by two separate CSR instructions, then a higher-priority interrupt may become newly pending-and-enabled in the interrupt file between the two instructions, causing the write to clear the pending bit of the new interrupt and not the one reported by the read. Once the pending bit of the new interrupt is cleared, the interrupt is lost. If it is necessary first to read a \*topei CSR and then subsequently claim the interrupt as a separate step, the claim can be safely done by clearing the pending bit in the eip array via \*siselect and \*sireg, instead of writing to \*topei. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#3-1-10-interrupt-delivery-and-handling)3.1.10\. Interrupt delivery and handling An IMSIC’s interrupt files supply _external interrupt_ signals to the attached hart, one interrupt signal per interrupt file. The interrupt signal from a machine-level interrupt file appears as bit MEIP in CSR `mip`, and the interrupt signal from a supervisor-level interrupt file appears as bit SEIP in `mip` and `sip`. Interrupt signals from any guest interrupt files appear as the active bits in hypervisor CSR `hgeip`. When interrupt delivery is disabled by an interrupt file’s `eidelivery` register (`eidelivery` \= 0), the interrupt signal from the interrupt file is held de-asserted (false). When interrupt delivery from an interrupt file is enabled (`eidelivery` \= 1), its interrupt signal is asserted if and only if the interrupt file has a pending-and-enabled interrupt that also exceeds the priority threshold specified by `eithreshold`, if not zero. Changes to the state of an interrupt file are guaranteed to be reflected in the relevant interrupt-pending bit in CSR`mip` or `hgeip` eventually, but not necessarily immediately. A trap handler solely for external interrupts via an IMSIC could be written roughly as follows: | save processor registers | | ----------------------------------------------------------------------------- | | i\=read CSR mtopei or stopei, and write simultaneously to claim the interrupt | | i\= i>>16 | | call the interrupt handler for external interrupt i (minor identity) | | restore processor registers | | return from trap | The combined read and write of `mtopei` or `stopei` in the second step can be done by a single CSRRW machine instruction, `csrrw` _rd_, `mtopei/stopei`, `x0` where _rd_ is the destination register for value _i_. 8.1. IOMMU Support for MSIs to Virtual Machines ==================== ## [](#IOMMU)8.1\. IOMMU Support for MSIs to Virtual Machines The existence of an IOMMU in a system makes it possible for a guest operating system, running in a virtual machine, to be given direct control of an I/O device with only minimal hypervisor intervention. A guest OS with direct control of a device will program the device with guest physical addresses, because that is all the OS knows. When the device then performs memory accesses using those addresses, an IOMMU is responsible for translating those guest physical addresses into machine physical addresses, referencing address-translation data structures supplied by the hypervisor. To handle MSIs from a device controlled by a guest OS, an IOMMU must be able to redirect those MSIs to a guest interrupt file in an IMSIC. Systems that do not have IMSICs with guest interrupt files do not need to implement the facilities described in this chapter. Because MSIs from devices are simply memory writes, they would naturally be subject to the same address translation that an IOMMU applies to other memory writes. However, the Advanced Interrupt Architecture requires that IOMMUs treat MSIs directed to virtual machines specially, in part to simplify software, and in part to allow optional support for_memory-resident interrupt files_. This chapter uses the term _IOMMU_ in a generic sense that encompasses all translation and transaction processing services required to virtualize device accesses and is concerned only with how an IOMMU recognizes and processes MSIs directed to virtual machines. Most other functions and details of an IOMMU are beyond the scope of this standard, and must be specified elsewhere. | | The RISC-V IOMMU Architecture Specification provides a detailed description of the IOMMU architecture, dividing translation and transaction processing functionality into blocks such as IOMMU, IO Bridge, etc. and describing how those blocks are integrated into a system. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If a single physical I/O device can be subdivided for control by multiple separate device drivers, each sub-device is referred to here as one device. ### [](#IOMMU-deviceContexts)8.1.1\. Device contexts at an IOMMU The following assumptions are made about the IOMMUs in a system: * For each I/O device connected to the system through an IOMMU, software can configure at the IOMMU a _device context_, which associates with the device a specific virtual address space and any other per-device parameters the IOMMU may support. By giving devices each their own separate device context at an IOMMU, each device can be individually configured for a separate operating system, which may be a guest OS or the main (host) OS. On every memory access initiated by a device, hardware indicates to the IOMMU the originating device by some form of unique device identifier, which the IOMMU uses to locate the appropriate device context within data structures supplied by software. For PCI, for example, the originating device may be identified by the unique triple of PCI bus number, device number, and function number. * An IOMMU optionally translates the addresses of a device’s memory accesses using address-translation data structures—typically page tables—specified by software via the corresponding device context. The smallest granularity of address translation implemented by all IOMMUs is not larger than a 4-KiB page, matching that of standard RISC-V address-translation page tables. (An IOMMU may in fact employ page tables in the same format as the page-based address translation defined by the RISC-V Privileged Architecture, but this is not required.) The Advanced Interrupt Architecture adds to device contexts these fields, as needed: * an _MSI address mask_ and _address pattern_, used together to identify pages in the guest physical address space that are the destinations of MSIs; and * the real physical address of an _MSI page table_ for controlling the translation and/or conversion of MSIs from the device. The MSI address mask and address pattern are each unsigned integers with the same width as guest physical page numbers, i.e., 12 bits narrower than the maximum supported width of a guest physical address. Their use is explained in [8.1.4\. Identification of page addresses of a VM’s interrupt files](#IOMMU-identIncomingMSIs). A device context’s MSI page table is separate from the usual address-translation data structures used to translate other memory accesses from the same device. The form and function of MSI page tables are the subject of most of the rest of this chapter. | | A device context is given an independent page table for MSIs for two reasons: First, hypervisors running under Linux or a similar OS can benefit from separate control of MSI translations to help simplify the case when virtual harts are migrated from one physical hart to another. As noted in [\[virtHartMigration\]](#virtHartMigration), when a virtual hart’s interrupt files are mapped to guest interrupt files in the real machine, migration of the virtual hart causes the physical guest interrupt files underlying those virtual interrupt files to change. However, because on other systems (not RISC-V) the migration of a virtual hart does not affect the mapping from guest physical addresses to real physical addresses, the internal functions of Linux that perform this migration are not set up to modify an IOMMU’s address-translation tables to adjust for the changing physical locations of RISC-V virtual interrupt files. Giving a hypervisor control of a separate MSI translation table at an IOMMU bypasses this limitation. The MSI page table can be modified at will by the hypervisor and/or by the subsystem that manages interrupts without coordinating with the many other OS components concerned with regular address translation. Second, specifying a separate MSI page table facilitates the use of_memory-resident interrupt files_ (MRIFs), which are introduced in[8.1.3\. Memory-resident interrupt files](#IOMMU-MRIFs). A dedicated MSI page table can easily support a special table entry format for MRIFs ([8.1.5.2\. MSI PTE, MRIF mode](#IOMMU-MSIPTE-MRIF)) that would be entirely foreign and difficult to retrofit to any other address-translation data structures. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#8-1-2-translation-of-addresses-for-msis-from-devices)8.1.2\. Translation of addresses for MSIs from devices To support the delivery of MSIs from I/O devices directly to RISC-V virtual machines without hypervisor intervention, an IOMMU must be able to translate the guest physical address of an MSI to the real physical address of an IMSIC’s guest interrupt file in the machine, as illustrated in [Figure 1](#IOMMU-guestIntrFiles). This address translation is controlled by the MSI page table configured in the appropriate device context at the IOMMU. Because every interrupt file, real or virtual, occupies a naturally aligned 4-KiB page of address space, the required address translation is from a virtual (guest) page address to a physical page address, the same as supported by regular RISC-V page-based address translation. ![IOMMU guestIntrFiles](_images/IOMMU-guestIntrFiles.png) Figure 1\. Translation of a device-sourced MSI that a guest OS intended to go to a (virtual) IMSIC interrupt file in the OS’s virtual machine. Referencing an MSI page table supplied by the controlling hypervisor, the IOMMU redirects the MSI to a guest interrupt file of the real machine. Memory writes from a device are recognized as MSIs by the destination address of the write. If an IOMMU determines that a 32-bit write is to the location of a (virtual) interrupt file in the relevant virtual machine, the write is considered an MSI within the VM, else not. The exact formula for recognizing MSIs is documented in[8.1.4\. Identification of page addresses of a VM’s interrupt files](#IOMMU-identIncomingMSIs). | | Although the translation of MSIs is controlled by its own separate page table, the fact that MSI translations are at the same page granularity as regular RISC-V address translations implies that an address translation cache within an IOMMU requires little modification to also cache MSI translations. Only on a translation cache miss does the IOMMU need to treat MSIs significantly differently than other memory accesses from the same device, to choose the correct translation table and to access and interpret the table properly. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#IOMMU-MRIFs)8.1.3\. Memory-resident interrupt files An IOMMU may optionally support memory-resident interrupt files (MRIFs). If implemented, the use of memory-resident interrupt files can greatly increase the number of virtual harts that can be given direct control of one or more physical devices in a system, assuming the rest of the system can still handle the added load. Without memory-resident interrupt files, the number of virtual RISC-V harts that can directly receive MSIs from devices is limited by the total number of guest interrupt files implemented by all IMSICs in the system, because all MSIs to RISC-V harts must go through IMSICs. For a single RISC-V hart, the number of guest interrupt files is the _GEILEN_ parameter defined by the H extension, which can be at most 31 for RV32 and 63 for RV64. With the use of memory-resident interrupt files, on the other hand, the total number of virtual RISC-V harts able to receive device MSIs is almost unbounded, constrained only by the amount of real physical memory and the additional processing time needed to handle them. As its name implies, a memory-resident interrupt file is located in memory instead of within an IMSIC. [Figure 2](#IOMMU-MRIF) depicts how an IOMMU can record an incoming MSI in an MRIF. When properly configured by a hypervisor, an IOMMU recognizes certain incoming MSIs as intended for a specific virtual interrupt file, and records each such MSI by setting an interrupt-pending bit stored within the MRIF data structure in ordinary memory. After each MSI is recorded in an MRIF, the IOMMU also sends a_notice MSI_ to the hypervisor to inform it that the MRIF contents may have changed. ![IOMMU MRIF](_images/IOMMU-MRIF.png) Figure 2\. Recording an incoming MSI into a memory-resident interrupt file (MRIF) instead of sending it to a guest interrupt file as in [Figure 1](#IOMMU-guestIntrFiles). While a memory-resident interrupt file provides a place to record MSIs, it cannot interrupt a hart directly the way an IMSIC’s guest interrupt files can. The notice MSIs that hypervisors receive only indicate that a virtual hart _might_ need interrupting; a hypervisor is responsible for examining the MRIF contents each time to determine whether actually to interrupt the virtual hart. Furthermore, whereas an IMSIC’s guest interrupt file can directly act as a supervisor-level interrupt file for a virtual hart, keeping a virtual hart’s interrupt file in an MRIF while the virtual hart executes requires that the hypervisor emulate a supervisor-level interrupt file for the virtual hart, hiding the underlying MRIF. Depending on how often the virtual hart touches its interrupt file and the implementation’s level of support for MRIFs, the cost of this emulation may be significant. Consequently, MRIFs are expected most often to be used for virtual harts that are more-or-less "swapped out" of a physical hart due to being idle, or nearly so. When a hypervisor determines that an MSI that landed in an MRIF should wake up a particular virtual hart that was idle, the virtual hart can be assigned a guest interrupt file in an IMSIC and its interrupt file moved from the MRIF into this guest interrupt file before the virtual hart is resumed. The process of allocating a guest interrupt file for the newly wakened virtual hart may of course force the interrupt file of another virtual hart to be evicted to its own MRIF. | | Not all systems need to accommodate large numbers of idle virtual harts. Many batch-processing servers, for example, strive to keep all virtual worker threads as busy as possible from start to finish, throttled only by I/O delays and limits on processing resources. In such environments, support for MRIFs may not be useful, so long as parameter GEILEN is not too small. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | An IOMMU can have one of these three levels of support for memory-resident interrupt files: * no memory-resident interrupt files; * memory-resident interrupt files without atomic update; or * memory-resident interrupt files with atomic update. Memory-resident interrupt files are most efficient when the memory system supports logical atomic memory operations (AMOs) corresponding to RISC-V instructions AMOAND and AMOOR, for memory accesses made both from harts and from the IOMMU. The AMOAND and AMOOR operations are required for_atomic update_ of a memory-resident interrupt file. A reduced level of support is possible without AMOs, relying solely on basic memory reads and writes. #### [](#IOMMU-MRIFFormat)8.1.3.1\. Format of a memory-resident interrupt file A memory-resident interrupt file occupies 512 bytes of memory, naturally aligned to a 512-byte address boundary. The 512 bytes are organized as an array of 32 pairs of 64-bit doublewords, 64 doublewords in all. Each doubleword is in little-endian byte order (even for systems where all harts are big-endian-only). | | Big-endian-configured harts that make use of MRIFs are expected to implement the REV8 byte-reversal instruction defined by standard RISC-V extension Zbb, or pay the cost of endianness conversion using a sequence of instructions. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The pairs of doublewords contain the interrupt-pending and interrupt-enable bits for external interrupt identities 1-2047, in this arrangement: | offset | size | contents | | ------ | ------- | -------------------------------------------------- | | 0x000 | 8 bytes | interrupt-pending bits for (minor) identities 1-63 | | 0x008 | 8 bytes | interrupt-enable bits for identities 1-63 | | 0x010 | 8 bytes | interrupt-pending bits for identities 64-127 | | 0x018 | 8 bytes | interrupt-enable bits for identities 64-127 | | … | … | | | 0x1F0 | 8 bytes | interrupt-pending bits for identities 1984-2047 | | 0x1F8 | 8 bytes | interrupt-enable bits for identities 1984-2047 | In general, the pair of doublewords at address offsets and for integer contain the interrupt-pending and interrupt-enable bits for external interrupt minor identities in the range to . For identity in this range, bit of the first (even) doubleword is the interrupt-pending bit, and the same bit of the second (odd) doubleword is the interrupt-enable bit. | | The interrupt-pending and interrupt-enable bits are stored interleaved by doublewords within an MRIF to facilitate the possibility of an IOMMU examining the relevant enable bit to determine whether to send a notice MSI after updating a pending bit, rather than the default behavior of always sending a notice MSI after an update without regard for the interrupt-enable bits. The memory arrangement matters only when MRIFs are supported without atomic update. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Bit 0 of the first doubleword of an MRIF stores a faux interrupt-pending bit for nonexistent interrupt 0\. If a write from an I/O device appears to be an MSI that should be stored in an MRIF, yet the data to write (the interrupt identity) is zero, the IOMMU acts as though zero were a valid interrupt identity, setting bit 0 of the target MRIF’s first doubleword and sending a notice MSI as usual. All MRIFs are the size to accommodate 2047 valid interrupt identities, the maximum allowed for an IMSIC interrupt file. If a system’s actual IMSICs have interrupt files that implement only interrupt identities, , then the contents of MRIFs for identities greater than may be ignored by software. IOMMUs, however, treat every MRIF as though all interrupt identities in the range 0-2047 are valid, even as software ignores invalid identity 0 and all identities greater than . | | There is no need to specify to an IOMMU a desired size for an MRIF smaller than 2047 valid interrupt identities. The only use an IOMMU would make of this information would be to discard any MSIs indicating an interrupt identity greater than . If devices are properly configured by software, such errant MSIs should not occur; but even if they do, it is just as effective for software to ignore spurious interrupt identities _after_ they have been recorded in an MRIF as for an IOMMU to discard them before recording them in the MRIF. It is likewise unnecessary for IOMMUs to check for and discard MSIs indicating an invalid interrupt identity of zero. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | #### [](#8-1-3-2-recording-of-incoming-msis-to-memory-resident-interrupt-files)8.1.3.2\. Recording of incoming MSIs to memory-resident interrupt files The data component of an MSI write specifies the interrupt identity to raise in the destination interrupt file. (Recall[MSI encoding](IMSIC.html#MSIEncoding).) This data may be in little-endian or big-endian byte order. If an IOMMU supports memory-resident interrupt files, it can store to an MRIF MSIs of the same endianness that the machine’s IMSICs accept. All IMSIC interrupt files are required to accept MSIs in little-endian byte order written to memory-mapped register `seteipnum_le` ([Memory region for an interrupt file](IMSIC.html#IMSIC-memRegion)). IMSIC interrupt files may also accept MSIs in big-endian byte order if register `seteipnum_be` is implemented alongside `seteipnum_le`. If the interrupt identity indicated by an MSI’s data (when interpreted in the correct byte order) is in the range 0-2047, an IOMMU stores the MSI to an MRIF by setting to one the interrupt-pending bit in the MRIF for that identity. If atomic update is supported for MRIFs, the pending bit is set using an AMOOR operation, else it is set using a non-atomic read-modify-write sequence. After the interrupt-pending bit is set in the MRIF, the IOMMU sends the notice MSI that software has configured for the MRIF. The exact process of storing an MSI to an MRIF is specified more precisely in [8.1.5.2\. MSI PTE, MRIF mode](#IOMMU-MSIPTE-MRIF), which covers MSI page table entries configured in _MRIF mode_. | | It is an open question whether an IOMMU might optionally examine the matching interrupt-enable bit within a destination MRIF to decide whether to send a notice MSI after setting an interrupt-pending bit. Currently, an IOMMU is required always to send a notice MSI after storing an MSI to an MRIF, even when the corresponding enable bit for the interrupt identity is zero. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#8-1-3-3-use-of-memory-resident-interrupt-files-with-atomic-update)8.1.3.3\. Use of memory-resident interrupt files with atomic update To make use of a memory-resident interrupt file with support for atomic update, software must have memory locations to save an IMSIC interrupt file’s `eidelivery` and `eithreshold` registers, in addition to the MRIF structure itself from [8.1.3.1\. Format of a memory-resident interrupt file](#IOMMU-MRIFFormat). Moving a virtual hart’s interrupt file from an IMSIC into an MRIF involves these steps: 1. Prepare the MRIF by zeroing all of its interrupt-pending bits (the even doublewords) and by copying the IMSIC interrupt file’s `eie` array to the MRIF’s interrupt-enable bits (the odd doublewords). 2. Save to memory the existing values of the IMSIC interrupt file’s registers `eidelivery` and `eithreshold`, and set `eidelivery` \= 0. 3. Modify all relevant translation tables at IOMMUs so that MSIs for this virtual interrupt file are now stored in the MRIF. If necessary, synchronize with all IOMMUs to ensure that no straggler MSIs will arrive at the IMSIC interrupt file after this step. 4. Logically OR the contents of the IMSIC interrupt file’s `eip` array into the interrupt-pending bits of the MRIF, using AMOOR operations. Once this sequence is complete, the IMSIC interrupt file is no longer in use. Each time a notice MSI arrives indicating that an MSI has been stored in the MRIF, the controlling hypervisor should scan the MRIF’s interrupt-pending and interrupt-enable bits to determine if any enabled interrupt is now both pending and enabled and thus should interrupt the virtual hart. With atomic update of MRIFs, a virtual hart may continue executing with its interrupt file contained in an MRIF, so long as the hypervisor emulates for the virtual hart a proper IMSIC interrupt file to hide the underlying MRIF. Hypervisor software can safely set and clear the interrupt-pending and interrupt-enable bits of the MRIF using AMOOR and AMOAND operations, even as an IOMMU may be storing incoming MSIs into the same MRIF. | | If an IOMMU is ever configured to examine an MRIF’s interrupt-enable bits to decide whether to send notice MSIs, then modifying those enable bits will generally require coordination with the IOMMU. But so long as IOMMUs ignore the interrupt-enable bits as is currently assumed, the bits can be changed by software without risk. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | To move the same interrupt file from the MRIF back to an IMSIC: 1. At the new IMSIC interrupt file, set `eidelivery` \= 0, and zero the `eip` array. 2. Modify all relevant translation tables at IOMMUs so that MSIs for this virtual interrupt file are now sent to the IMSIC interrupt file. If necessary, synchronize with all IOMMUs to ensure that no straggler MSIs will be stored in the MRIF after this step. 3. Logically OR the interrupt-pending bits from the MRIF into the IMSIC interrupt file, using instruction CSRS to write to the `eip` array. Also, copy the interrupt-enable bits from the MRIF to the IMSIC interrupt file’s `eie` array. 4. Load the IMSIC interrupt file’s registers `eithreshold` and `eidelivery` with the values that were earlier saved. #### [](#8-1-3-4-use-of-memory-resident-interrupt-files-without-atomic-update)8.1.3.4\. Use of memory-resident interrupt files without atomic update Without support for atomic update, the use of memory-resident interrupt files is similar to the atomic-update case of the previous subsection, but with some added complexities. First, if the I/O devices that a virtual hart controls are behind multiple IOMMUs, then multiple MRIF structures are needed, one per IOMMU, not just a single MRIF structure. Furthermore, in addition to locations for storing `eidelivery` and `eithreshold`, software needs a place for a complete copy of the interrupt file’s implemented `eip` array, apart from the MRIFs. While a virtual interrupt file is in memory, its interrupt-pending bits will be split across all the MRIFs and the saved `eip` array. The interrupt-enable bits may exist only in the MRIFs. To move a virtual hart’s interrupt file from an IMSIC into memory, with one MRIF per IOMMU: 1. Prepare all MRIFs by zeroing their interrupt-pending bits (the even doublewords) and by copying the IMSIC interrupt file’s `eie` array to the MRIFs' interrupt-enable bits (the odd doublewords). 2. Save to memory the existing values of the IMSIC interrupt file’s registers `eidelivery` and `eithreshold`, and set `eidelivery` \= 0. 3. At each IOMMU, modify all relevant translation tables so that MSIs for this virtual interrupt file are now stored in the individual MRIF matched to the IOMMU. If necessary, synchronize with all IOMMUs to ensure that no straggler MSIs will arrive at the IMSIC interrupt file after this step. 4. Dump the IMSIC interrupt file’s `eip` array to its separate location outside the MRIFs. Once this sequence is complete, the IMSIC interrupt file is no longer in use. While a virtual hart’s interrupt file remains in memory, an interrupt identity’s true pending bit is the logical OR of its bit in all MRIFs and its bit in the saved `eip` array. All pending bits in the MRIFs start as zeros, but interrupts may become pending there as MSIs for this virtual hart arrive at IOMMUs and are stored in the corresponding MRIFs. Without atomic update of MRIFs, an interrupt-pending bit is not easily cleared in an MRIF. (Clearing a single pending bit in one MRIF requires that a new MRIF be allocated and initialized and the corresponding IOMMU reconfigured to store MSIs into the new MRIF.) For this reason, it may or may not be practical to have a virtual hart execute while keeping one of its interrupt files in memory. When an MRIF records an interrupt that should wake a virtual hart, the simplest strategy is to always move the interrupt file back into an IMSIC’s guest interrupt file before resuming execution of the virtual hart. To transfer an interrupt file from memory back to an IMSIC: 1. At the new IMSIC interrupt file, set `eidelivery` \= 0, and zero the `eip` array. 2. Modify all relevant translation tables at IOMMUs so that MSIs for this virtual interrupt file are now sent to the IMSIC interrupt file. If necessary, synchronize with all IOMMUs to ensure that no straggler MSIs will be stored in MRIFs after this step. 3. Merge by bitwise logical OR the interrupt-pending bits of all MRIFs and the saved `eip` array, and logically OR these merged bits into the IMSIC interrupt file, using instruction CSRS to write to the `eip` array. Also, copy the interrupt-enable bits from one of the MRIFs to the IMSIC interrupt file’s `eie` array. 4. Load the IMSIC interrupt file’s registers `eithreshold` and `eidelivery` with the values that were earlier saved. #### [](#8-1-3-5-allocation-of-guest-interrupt-files-for-receiving-notice-msis)8.1.3.5\. Allocation of guest interrupt files for receiving notice MSIs The processing a hypervisor does in response to notice MSIs can be minimized by assigning a separate interrupt identity for each MRIF, so the identity encoded in a notice MSI always indicates which one MRIF may have changed. However, if there are very many MRIFs (potentially in the thousands), a hypervisor may run short of interrupt identities within the supervisor-level interrupt files available in IMSICs. In that case, the hypervisor can increase its supply of interrupt identities by allocating one or more of the IMSICs’ guest interrupt files to itself for the purpose of receiving notice MSIs. | | Although guest interrupt files exist primarily to act as supervisor-level interrupt files for virtual harts, the IMSIC hardware does not police exactly how they are used by software. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#IOMMU-identIncomingMSIs)8.1.4\. Identification of page addresses of a VM’s interrupt files When an I/O device is configured directly by a guest operating system, MSIs from the device are expected to be targeted to virtual IMSICs within the guest OS’s virtual machine, using guest physical addresses that are inappropriate and unsafe for the real machine. An IOMMU must recognize certain incoming writes from such devices as MSIs and convert them as needed for the real machine. (Recall[Figure 1](#IOMMU-guestIntrFiles).) MSIs originating from a single device that require conversion are expected to have been configured at the device by a single guest OS running within one RISC-V virtual machine. Assuming the VM itself conforms to the Advanced Interrupt Architecture, MSIs are sent to virtual harts within the VM by writing to the memory-mapped registers of the interrupt files of virtual IMSICs. Each of these virtual interrupt files occupies a separate 4-KiB page in the VM’s guest physical address space, the same as real interrupt files do in a real machine’s physical address space. A write to a guest physical address can thus be recognized as an MSI to a virtual hart if the write is to a page occupied by an interrupt file of a virtual IMSIC within the VM. The MSI address mask and address pattern specified in a device context ([8.1.1\. Device contexts at an IOMMU](#IOMMU-deviceContexts)) are used to identify the 4-KiB pages of virtual interrupt files in the guest physical address space of the relevant VM. An incoming 32-bit write made by a device is recognized as an MSI write to a virtual interrupt file if the destination guest physical page matches the supplied address pattern in all bit positions that are zeros in the supplied address mask. In detail, a memory access to guest physical address is an access to a virtual interrupt file’s memory-mapped page if ((A >> 12) & \~address mask) = (address pattern & \~address mask) where >> 12 represents shifting right by 12 bits, an ampersand (&) represents bitwise logical AND, and "\~address mask" is the bitwise logical complement of the address mask. When a memory access is found to be to a virtual interrupt file, an_interrupt file number_ is extracted from the original guest physical address as interrupt file number = extract(A >> 12, address mask) Here, extract( , ) is a "bit extract" that discards all bits from whose matching bits in the same positions in the mask are zeros, and packs the remaining bits from contiguously at the least-significant end of the result, keeping the same bit order as and filling any other bits at the most-significant end of the result with zeros. For example, if the bits of and are \= a b c d e f g h \= 1 0 1 0 0 1 1 0 then the value of extract( , ) has bits 0 0 0 0 a c f g. ### [](#8-1-5-msi-page-tables)8.1.5\. MSI page tables When an IOMMU determines that a memory access is to a virtual interrupt file as specified in the previous section, the access is translated or converted by consulting the MSI page table configured for the device, instead of using the regular translation data structures that apply to all other memory accesses from the same device. An MSI page table is a flat array of MSI page table entries (MSI PTEs), each 16 bytes. MSI page tables have no multi-level hierarchy like regular RISC-V page tables do. Rather, every MSI PTE is a leaf entry specifying the translation or conversion of accesses made to a particular 4-KiB guest physical page that a virtual interrupt file occupies (or may occupy) in the relevant virtual machine. To select an individual MSI PTE from an MSI page table, the PTE array is indexed by the interrupt file number extracted from the destination guest physical address of the incoming memory access by the formula of the previous section. Each MSI PTE may specify either the address of a real guest interrupt file that substitutes for the targeted virtual interrupt file (as in[Figure 1](#IOMMU-guestIntrFiles)), or a memory-resident interrupt file in which to store incoming MSIs for the virtual interrupt file (as in [Figure 2](#IOMMU-MRIF)). The number of entries in an MSI page table is where is the number of bits that are ones in the MSI address mask used to extract the interrupt file number from the destination guest physical address. If an MSI page table has 256 or fewer entries, the start of the table is aligned to a 4-KiB page address in real physical memory. If an MSI page table has entries, the table must be naturally aligned to a address boundary. If an MSI page table is not aligned as required, all entries in the table appear to an IOMMU as UNSPECIFED, and any address an IOMMU may compute and use for reading an individual MSI PTE from the table is also UNSPECIFIED. Every 16-byte MSI PTE is interpreted as two 64-bit doublewords. If an IOMMU also references standard RISC-V page tables, defined by the RISC-V Privileged Architecture, for regular address translation, then the byte order for each of the two doublewords in memory, little-endian or big-endian, should be the same as the endianness of the regular RISC-V page tables configured for the same device context. Otherwise, the endianness of the doublewords of an MSI PTE is implementation-defined. Bit 0 of the first doubleword of an MSI PTE is field V (Valid). When V = 0, the PTE is invalid, and all other bits of both doublewords are ignored by an IOMMU, making them free for software to use. If V = 1, bit 63 of the first doubleword is field C (Custom), designated for custom use. If an MSI PTE has V = 1 and C = 1, interpretation of the rest of the PTE is implementation-defined. If V = 1 and the custom-use bit C = 0, then bits 2:1 of the first doubleword contain field M (Mode). If M = 3, the MSI PTE specifies_basic translate mode_ for accesses to the page, and if M = 1, it specifies _MRIF mode_. Values of 0 and 2 for M are reserved. The interpretation of an MSI PTE for each of the two defined modes is detailed further in the next two subsections. #### [](#8-1-5-1-msi-pte-basic-translate-mode)8.1.5.1\. MSI PTE, basic translate mode When an MSI PTE has fields V = 1, C = 0, and M = 3 (basic translate mode), the PTE’s complete format is: | First doubleword: | bit 63 | C, = 0 | | ------------------ | ------- | ------ | | bits 53:10 | PPN | | | bits 2:1 | M, = 3 | | | bit 0 | V, = 1 | | | Second doubleword: | ignored | | All other bits of the first doubleword are reserved and must be set to zeros by software. The second doubleword is ignored by an IOMMU so is free for software to use. A memory access within the page covered by the MSI PTE is translated by replacing the access’s original address bits 12 and above (the guest physical page number) with field PPN (Physical Page Number) from the PTE, while retaining the original address bits 11:0 (the page offset). This translated address is either zero-extended or clipped at the upper end as needed to make it the width of a real physical address for the machine. The original memory access from the device is then passed onward to the memory system with the new address. An MSI PTE in basic translate mode allows a hypervisor to route an MSI write intended for a virtual interrupt file to go instead to a guest interrupt file of a real IMSIC in the machine. | | An IOMMU that also employs standard RISC-V page tables for regular address translation can maximize the overlap between the handling of MSI PTEs and regular RISC-V leaf PTEs as follows: For RV64, the first doubleword of an MSI PTE in basic translate mode has the same encoding as a regular RISC-V leaf PTE for Sv39, Sv48, Sv57, Sv39x4, Sv48x4, or Sv57x4 page-based address translation, with PTE fields D, A, G, U, and X all zeros and W = R = 1\. Hence, the MSI PTE’s first doubleword appears the same as a regular PTE that grants read and write permission (R = W = 1) but not execute permissions (X = 0). This same-encoded regular PTE would translate an MSI write the same as the actual MSI PTE, except that what would be the PTE’s accessed (A), dirty (D), and user (U) bits are all zeros. An IOMMU needs to treat only these three bits differently for an MSI PTE versus a regular RV64 leaf PTE. The address computation used to select a PTE from a regular RISC-V page table must be modified to select an MSI PTE’s first doubleword from an MSI page table. However, the extraction of an interrupt file number from a guest physical address to obtain the index for accessing the MSI page table already creates an unavoidable difference in PTE addressing. For RV32, the lower 32-bit word of an MSI PTE’s first doubleword has the same format as a leaf PTE for Sv32 or Sv32x4 page-based address translation, except again for what would be PTE bits A, D, and U, which must be treated differently. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#IOMMU-MSIPTE-MRIF)8.1.5.2\. MSI PTE, MRIF mode If memory-resident interrupt files are supported and an MSI PTE has fields V = 1, C = 0, and M = 1 (MRIF mode), the PTE’s complete format is: | First doubleword: | bit 63 | C, = 0 | | ------------------ | -------------------- | --------- | | bits 53:7 | MRIF Address\[55:9\] | | | bits 2:1 | M, = 1 | | | bit 0 | V, = 1 | | | Second doubleword: | bit 60 | NID\[10\] | | bits 53:10 | NPPN | | | bits 9:0 | NID\[9:0\] | | All other PTE bits are reserved and must be set to zeros by software. The PTE’s MRIF Address field provides bits 55:9 of the physical address of a memory-resident interrupt file in which to store incoming MSIs, referred to as the _destination MRIF_. As every memory-resident interrupt file is naturally aligned to a 512-byte address boundary, bits 8:0 of the destination MRIF’s address must be zero and are not specified in the PTE. Field NPPN (Notice Physical Page Number) and the two NID (Notice Identifier) fields together specify a destination and value for a_notice MSI_ that is sent after each time the destination MRIF is updated as a result of consulting this PTE to store an incoming MSI. | | Typically, NPPN will be the page address of an IMSIC’s interrupt file in the real machine, and NID will be the interrupt identity to make pending in that interrupt file to indicate that the destination MRIF may have changed. However, NPPN is not required to be a valid interrupt file address, and an IOMMU must not attempt to restrict it to only such addresses. Any page address must be accepted for NPPN. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Memory accesses by I/O devices to addresses within a page covered by an MRIF-mode PTE are handled by the IOMMU instead of being passed through to the memory system. If a memory access, read or write, is not for 32 bits of data, or if the access address is not aligned to a 4-byte boundary (including accesses that straddle the page boundary), the access should be aborted as unsupported. For a naturally aligned 32-bit read, the IOMMU should preferably return zero as the read value but may alternatively abort the access. A naturally aligned 32-bit write is either interpreted as an MSI, resulting in an update of the destination MRIF, or is discarded. When the IMSIC interrupt files in the system implement memory-mapped register `seteipnum_be` for receiving MSIs in big-endian byte order ([Memory region for an interrupt file](IMSIC.html#IMSIC-memRegion)), then an IOMMU must be able to store MSIs in both little-endian and big-endian byte orders to the destination MRIF. If the IMSIC interrupt files in the system do not implement register `seteipnum_be`, an IOMMU should ordinarily store only little-endian MSIs to the destination MRIF. The data of an incoming MSI is assumed to be in little-endian byte order if bit 2 of the destination address is zero, and in big-endian byte order if bit 2 of the destination address is one. If a naturally aligned 32-bit write is to guest physical address within a page covered by an MRIF-mode PTE, and if the write data is when interpreted in the byte order indicated by bit 2 of , then the write is processed as follows: If either \[11:3\] or \[31:11\] is not zero, or if bit 2 of is one and big-endian MSIs are not supported, then the incoming write is accepted but discarded. Else, the original incoming write is recognized as an MSI and is replaced by one of the following memory accesses, setting the interrupt-pending bit that corresponds to the interrupt identity in the destination MRIF to one: * an atomic AMOOR operation, if atomic updates are supported; or * a non-atomic read-modify-write sequence, if atomic updates are not supported. Once the MRIF update operation is visible to all agents in the system, the 11-bit NID value is zero-extended to 32 bits, and this value is written to the address NPPN<<12 (i.e., physical page number NPPN, page offset zero) in little-endian byte order. | | While IOMMUs are expected typically to cache MSI PTEs that are configured in basic translate mode (M = 3), they might not cache PTEs configured in MRIF mode (M = 1). Two reasons together justify not caching MSI PTEs in MRIF mode: First, the information and actions required to store an MSI to an MRIF are far different than normal address translation; and second, by their nature, MSIs to MRIFs should occur less frequently. Hence, an IOMMU might perform MRIF-mode processing solely as an extension of cache-miss page table walks, leaving its address translation cache oblivious to MRIF-mode MSI PTEs. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 7.1. Interprocessor Interrupts (IPIs) ==================== ## [](#IPIs)7.1\. Interprocessor Interrupts (IPIs) By default, unless a platform has a different mechanism for interprocessor interrupts (IPIs), the base RISC-V Privileged Architecture specifies that a machine with multiple harts must provide for each hart an implementation-defined memory address that can be written to signal a machine-level _software interrupt_ (major code 3) at that hart. IPIs at machine level can thus be sent to any hart as machine-level software interrupts. | | A RISC-V software interrupt acts only as a minimal "doorbell" signal. Software at the receiving hart is responsible for recognizing an incoming software interrupt as an IPI and decoding its purpose further, usually making use of additional data stored by the sender in ordinary memory. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The same kind of mechanism (but with a different set of memory addresses) may or may not exist for signaling supervisor-level software interrupts (major code 1) at remote harts as well. If not directly supported in this way, a supervisor-level software interrupt is typically sent to another hart instead through an environment call from supervisor mode to machine mode. An operating system running in S-mode thus invokes a specific SBI function for delivering a software interrupt to another hart, causing machine-level software at the originating hart to send a machine-level IPI to the destination hart, where software then sets the supervisor-level software interrupt-pending bit (SSIP) in CSR `mip`. When harts have IMSICs, instead of using the base Privileged Architecture’s mechanism for signaling software interrupts at remote harts, an IPI can be sent to a hart by writing to the destination hart’s IMSIC, the same as a regular message-signaled interrupt (MSI). In that case, an incoming IPI appears at the destination hart as an _external interrupt_ routed through the IMSIC, rather than as a software interrupt as before. However, so long as the same software (e.g. an operating system or machine monitor) is in control at both endpoints of an IPI, source and destination, there should be no reason for a destination hart to misinterpret the purpose of an incoming external interrupt that represents an IPI. If harts do not have IMSICs, then the method specified by the base Privileged Architecture is assumed to be used for IPIs, signaling software interrupts at destination harts. On the other hand, when harts have IMSICs, the machinery for triggering software interrupts at remote harts is redundant with the capabilities of the IMSICs, so it is downgraded from a requirement to an option, useful perhaps only to provide software compatibility across a range of RISC-V systems, with and without IMSICs. If a machine implements IMSICs and not the earlier software-interrupt mechanism, then the bits of CSRs `mip` and `mie` for machine-level software interrupts, MSIP and MSIE, are hardwired to zero in harts. | | If a machine implements IMSICs but not the software-interrupt mechanism, the latter can still be fully emulated at supervisor level for S-mode or VS-mode, by trapping on writes to the special memory addresses that should signal supervisor-level software interrupts at remote harts. On such a trap, software can send a higher-level IPI via IMSIC to the destination hart, where the higher-level software then can set the SSIP bit in sip at the intended privilege level, S or VS. Similarly, SBI environment calls for sending IPIs can easily continue to be supported without clients being at all aware of a change in the underlying hardware for delivering IPIs between harts. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | When software sends IPIs by writing MSIs to the IMSICs of other harts, programmers should consider also the need to execute a FENCE instruction before each store instruction that writes such an MSI. In the absence of FENCEs, many systems guarantee to preserve the order of a hart’s loads and stores only to/from individual devices, not among multiple devices, and not at all for accesses to main memory. With such a system, it must be remembered that each IMSIC is likely to be considered a separate device among the many. For example, if hart A wants to notify hart B that it has completed a task involving accesses to some I/O device, hart A may need to execute a FENCE before sending an MSI to B’s IMSIC, to ensure that all of A’s accesses to the device have actually completed before the MSI could arrive at B. Similarly, if hart A stores anything to memory that should be visible at hart B, a FENCE is likely needed before a subsequent store sending an MSI to B’s IMSIC. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | 5.1. Interrupts for Machine and Supervisor Levels ==================== ## [](#MSLevel)5.1\. Interrupts for Machine and Supervisor Levels The RISC-V Privileged Architecture defines several major identities in the range 0-15 for interrupts at a hart, including machine-level and supervisor-level external interrupts (numbers 11 and 9), machine- and supervisor-level timer interrupts (7 and 5), and machine- and supervisor-level software interrupts (3 and 1). Beyond these major labels, the _external_ interrupts at each privilege level are given secondary, minor identities by an external interrupt controller such as an APLIC or IMSIC, distinguishing interrupts from different devices or causes. These minor identities for external interrupts were covered in[Incoming MSI Controller (IMSIC)](IMSIC.html#IMSIC) and [Advanced Platform-Level Interrupt Controller (APLIC)](AdvPLIC.html#AdvPLIC) specifying the IMSIC and APLIC components. The Advanced Interrupt Architecture reserves another 24 major interrupt identities for additional _local interrupts_ that arise within or in close proximity to the hart, often for reporting errors. A mechanism is also defined that allows software to selectively delegate both local and custom interrupts to the next less-privileged level, or in some cases to inject entirely virtual interrupts into a less-privileged level. Lastly, an optional facility lets software assign priorities to major interrupts (such as the timer and software interrupts, and any local interrupts) such that they may mix with the priorities set for external interrupts by a PLIC, APLIC, or IMSIC. ### [](#majorIntrs)5.1.1\. Defined major interrupts and default priorities [Table 1](#TablemajorIntrs) lists all the major interrupts currently defined for RISC-V harts that conform to this Advanced Interrupt Architecture (AIA). Besides the major interrupts specified by the RISC-V Privileged Architecture, the AIA adds interrupt numbers 35 and 43 as local interrupts for low- and high-priority _RAS events_. __Table 1\. The standard major interrupt codes, listed in default priority order__ | Default priority order | Major interrupt numbers | Description | | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------- | | Highest Lowest | 43 | Local interrupt: high-priority RAS event | | 11, 3, 79, 1, 51210, 2, 613 | Machine interrupts: external, software, timerSupervisor interrupts: external, software, timerSupervisor guest external interruptVS interrupts: external, software, timerLocal interrupt: counter overflow | | | 35 | Local interrupt: low-priority RAS event | | The default priority order in [Table 1](#TablemajorIntrs) is applicable only when multiple major interrupts would trap to the same privilege mode. Interrupt traps to a more-privileged mode always have priority over traps to a less-privileged mode. __Table 2\. Categorization of current and future major interrupts.__ | Major interrupt numbers | Category | | | ----------------------- | ---------------------------------------------------------------------------------------- | -------------------------------------- | | 0-1213-15 | Not Local interruptsLocal interrupts | Assigned by thePrivileged Architecture | | 16-2324-3132-47≥ 48 | Local interrupts _Designated for custom use_Local interrupts _Designated for custom use_ | | Of the major interrupts controlled by the base Privileged Architecture (numbers 0-15), the AIA categorizes the counter overflow interrupt (code 13) as a _local interrupt_. It is assumed furthermore that any future definitions for reserved interrupt numbers 14 and 15 will also be local interrupts. Besides the two RAS interrupts, the AIA additionally reserves major interrupt numbers in the ranges 16-23 and 32-47 for standard local interrupts that other RISC-V extensions may define. The remaining major interrupts allocated to the Privileged Architecture, numbers 0-12, are categorized as not local interrupts. Taken altogether,[Table 2](#TablemajorIntrCategories) summarizes the AIA’s categorization of all major interrupt identities. _RAS_ is an abbreviation for _Reliability, Availability, and Serviceability_. Typically a RAS event corresponds to the detection of corrupted data (e.g. as a result of a soft or hard error) and/or the use of such data. The high-priority RAS event local interrupt may, for example, signal an occurrence of an urgent uncorrected error that needs action from a RAS error handler to contain the error and, if possible, to recover from it. The low-priority RAS event local interrupt may, for example, be triggered by non-urgent deferred or corrected errors. The AIA does not itself require that detected RAS events trigger one of the two local interrupts defined for this purpose. Systems are free to report any or all RAS events another way, such as by external interrupts routed through an APLIC or IMSIC, or by custom interrupts. | | In all likelihood, the method for reporting a particular RAS event will depend on where in the system the event is detected. The AIA defines local interrupt numbers for RAS events so systems have a standard way to report such events when detected locally at a hart, without depending solely on external or custom interrupts. As always, platform standards may further constrain how a system reports events, whether RAS events or other. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | For the standard local interrupts not defined by the base RISC-V Privileged Architecture (numbers 16-23 and 32-47), the current plan is to assign default priorities in the order shown in this table: Highest Lowest 47, 23, 46, 45, 22, 44,43, 21, 42, 41, 20, 40 11, 3, 79, 1, 51210, 2, 613 Machine interrupts: external, software, timerSupervisor interrupts: external, software, timerSupervisor guest external interruptVS interrupts: external, software, timerCounter overflow interrupt 39, 19, 38, 37, 18, 36,35, 17, 34, 33, 16, 32 Among interrupts 16-23, a higher interrupt number conveys higher default priority, and likewise for interrupts 32-47\. These two groups are interleaved together in the complete order, and the Privileged Architecture’s standard interrupts, 0-15, are inserted into the middle of the sequence. This proposed default priority order is arranged so that interrupts 0-31 can potentially be an adequate subset on their own for 32-bit RISC-V systems. In actuality, future RISC-V extensions may or may not stick to this plan for the default priority order of interrupts they define. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | In addition to the existing major interrupts of[Table 1](#TablemajorIntrs), the following local interrupts are tentatively proposed, listed in order of decreasing default priority: 23 Bus or system error 45 Per-core high-power or over-temperature event 17 Debug/trace interrupt These local interrupts are expected to be specified by other RISC-V extensions. Be aware, this list is not final and may change as the relevant extensions are developed and ratified. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | If a future version of the RISC-V Privileged Architecture defines interrupt 0, the Advanced Interrupt Architecture needs it to have a default priority lower than certain external interrupts. See [5.1.2.2\. Machine top interrupt CSR (mtopi)](#mtopi)and [5.1.4.2\. Supervisor top interrupt CSR (stopi)](#stopi) on CSRs mtopi and stopi. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Interrupt numbers 24-31 and 48 and higher are all designated for custom use. If a hart implements any custom interrupts, their positions in the default priority order must be documented for the hart. | | While many of the standard registers such as mip and mie have space for major interrupts only in the range 0-63, custom interrupts with numbers 64 and above are conceivable with added custom support. CSRs mtopi([5.1.2.2\. Machine top interrupt CSR (mtopi)](#mtopi)) and stopi ([5.1.4.2\. Supervisor top interrupt CSR (stopi)](#stopi)) allow for major interrupt numbers potentially as large as 4095. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When a hart supports the arbitrary configuration of interrupt priorities by software (described in later sections), the default priority order still remains relevant for breaking ties when two interrupt sources are assigned the same priority number. ### [](#5-1-2-interrupts-at-machine-level)5.1.2\. Interrupts at machine level For whichever standard local interrupts are implemented, the corresponding bits in CSRs `mip` and `mie` must be writable, and the corresponding bits in `mideleg` (if that CSR exists because supervisor mode is implemented) must each either be writable or be hardwired to zero. An occurrence of a local interrupt event causes the interrupt-pending bit in `mip` to be set to one. This bit then remains set until cleared by software. As established by the base RISC-V Privileged Architecture, an interrupt traps to M-mode whenever all of the following are true: (a) either the current privilege mode is M-mode and machine-level interrupts are enabled by the MIE bit of `mstatus`, or the current privilege mode has less privilege than M-mode; (b) matching bits in `mip` and `mie` are both one; and (c) if `mideleg` exists, the corresponding bit in `mideleg` is zero. When multiple interrupt causes are ready to trigger simultaneously, the interrupt taken first is determined by priority order, which may be the default order specified in the previous section [5.1.1\. Defined major interrupts and default priorities](#majorIntrs), or may be a modified order configured by software. #### [](#intrPrios-M)5.1.2.1\. Configuring priorities of major interrupts at machine level The machine-level priorities for major interrupts 0-63 may be configured by a set of registers accessed through the `miselect` and `mireg` CSRs introduced in[Control and Status Registers (CSRs) Added to Harts](CSRs.html#CSRs). When XLEN = 32, sixteen of these registers are defined, listed below with their `miselect` addresses: | 0x30 | iprio0 | | ---- | ------- | | 0x31 | iprio1 | | …​ | …​ | | 0x3F | iprio15 | Each register controls the priorities of four interrupts, with one 8-bit byte per interrupt. For a number in the range 0-15, register `iprio` controls the priorities of interrupts through , formatted as follows: | bits 7:0 | Priority number for interrupt | | ---------- | ----------------------------- | | bits 15:8 | Priority number for interrupt | | bits 23:16 | Priority number for interrupt | | bits 31:24 | Priority number for interrupt | When XLEN = 64, only the even-numbered registers exist: | 0x30 | iprio0 | | ---- | ------- | | 0x32 | iprio2 | | …​ | …​ | | 0x3E | iprio14 | Each register controls the priorities of eight interrupts. For even in the range 0-14, register `iprio` controls the priorities of interrupts through , formatted as follows: | bits 7:0 | Priority number for interrupt | | ---------- | ----------------------------- | | bits 15:8 | Priority number for interrupt | | bits 23:16 | Priority number for interrupt | | bits 31:24 | Priority number for interrupt | | bits 39:32 | Priority number for interrupt | | bits 47:40 | Priority number for interrupt | | bits 55:48 | Priority number for interrupt | | bits 63:56 | Priority number for interrupt | When XLEN = 64 and `miselect` is an odd value in the range `0x31`\-`0x3F`, attempting to access `mireg` raises an illegal instruction exception. The valid registers `iprio0`\-`iprio15` are known collectively as the `_iprio_` _array_ for machine level. The width of priority numbers for external interrupts is _IPRIOLEN_. This parameter is affected by the main external interrupt controller for the hart, whether a PLIC, APLIC, or IMSIC. For an APLIC, IPRIOLEN is in the range 1-8 as specified in [Advanced Platform-Level Interrupt Controller (APLIC)](AdvPLIC.html#AdvPLIC) on the APLIC. For an IMSIC, IPRIOLEN is 6, 7, or 8\. IPRIOLEN may be 6 only if the number of external interrupt identities implemented by the IMSIC is 63\. IPRIOLEN may be 7 only if the number of external interrupt identities implemented by the IMSIC is no more than 127\. IPRIOLEN may be 8 for any IMSIC, regardless of the number of external interrupt identities implemented. Each byte of a valid `iprio` register is either a read-only zero or a **WARL** unsigned integer field implementing exactly IPRIOLEN bits. For a given interrupt number, if the corresponding bit in `mie` is read-only zero, then the interrupt’s priority number in the `iprio` array must be read-only zero as well. The priority number for a machine-level external interrupt (bits 31:24 of register `iprio2`) must also be read-only zero. Aside from these two restrictions, implementations may freely choose which priority number fields are settable and which are read-only zeros. If all bytes in the `iprio` array are read-only zeros, priorities can be configured only for external interrupts, not for any other interrupts. | | Platform standards may require that priorities be configurable for certain interrupt causes. | | ----------------------------------------------------------------------------------------------- | The `iprio` array accessed via `miselect` and `mireg` affects the prioritization of interrupts only when they trap to M-mode. When an interrupt’s priority number in the array is zero (either read-only zero or set to zero), its priority is the default order from [5.1.1\. Defined major interrupts and default priorities](#majorIntrs). Setting an interrupt’s priority number instead to a nonzero value gives that interrupt nominally the same priority as a machine-level external interrupt with priority number . For a major interrupt that defaults to a higher priority than machine external interrupts, setting its priority number to a nonzero value _lowers_ its priority. For a major interrupt that defaults to a lower priority than machine external interrupts, setting its priority number to a nonzero value _raises_ its priority. When two interrupt causes have been assigned the same nominal priority, ties are broken by the default priority order. [Table 3](#TableintrPrios-M) summarizes the effect of priority numbers on interrupt priority. | | When a hart has an IMSIC supporting more than 255 minor identities for external interrupts, the only non-default priorities that can be configured for other interrupts are those corresponding to external interrupt identities 1-255, not those of identities 256 or higher. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 3\. Effect of the machine-level iprio array on the priorities of interrupts taken in M-mode. For interrupts with the same priority number, the default order of [5.1.1\. Defined major interrupts and default priorities](#majorIntrs) prevails.__ | Interrupts with default priority above machine external interrupts | Machine external interrupts | Interrupts with default priority below machine external interrupts | | | ------------------------------------------------------------------ | -------------------------------------------- | ------------------------------------------------------------------ | -------------------------------------------- | | Priorityorder | Priority number in machine-level iprio array | Priority number from interrupt controller (APLIC or IMSIC) | Priority number in machine-level iprio array | | Highest | 0 | | | | 12…​254255 | 12…​254255 | 12…​254255 | | | 256 and above (IMSIC only) | | | | | Lowest | 0 | | | | | Implementing the priority configurability of this section requires that a RISC-V hart’s external interrupt controller communicate to the hart not only the existence of a pending-and-enabled external interrupt but also the interrupt’s priority number. Typically this implies that the width of the connection for signaling an external interrupt to the hart is not just a single wire as usual but now wires. It is expected that many systems will forego priority configurability of major interrupts and simply have the array be all read-only zeros. Systems that need this priority configurability can try to arrange for each hart’s external interrupt controller to be relatively close to the hart, by, for example, limiting the system to at most a few small cores connected to an APLIC, or alternatively by giving every hart its own IMSIC. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If supported, setting the priority number for supervisor-level external interrupts (bits 15:8 of `iprio2`) to a nonzero value has the effect of giving the entire category of supervisor external interrupts nominally the same priority as a machine external interrupt with priority number . But note that this applies only to the case when supervisor external interrupts trap to M-mode. (Because supervisor guest external interrupts and VS-level external interrupts are required to be delegated to supervisor level when the H extension is implemented, the machine-level priority numbers for these interrupts are always ignored and should be read-only zeros.) If the system has an original PLIC for backward compatibility with older software, reset should initialize the machine-level `iprio` array to all zeros. #### [](#mtopi)5.1.2.2\. Machine top interrupt CSR (`mtopi`) Machine-level CSR `mtopi` is read-only with width MXLEN. A read of `mtopi` returns information about the highest-priority pending-and-enabled interrupt for machine level, in this format: | bits 27:16 | IID | | ---------- | ----- | | bits 7:0 | IPRIO | All other bits of `mtopi` are reserved and read as zeros. The value of `mtopi` is zero unless there is an interrupt pending in `mip` and enabled in `mie` that is not delegated to a less-privileged level. When there is a pending-and-enabled major interrupt for machine level, field IID (Interrupt Identity) is the major identity number of the highest-priority interrupt, and field IPRIO indicates its priority. If all bytes of the machine-level `iprio` array are read-only zeros, a simplified implementation of field IPRIO is allowed in which its value is always 1 whenever `mtopi` is not zero. Otherwise, when `mtopi` is not zero, if the priority number for the reported interrupt is in the range 1 to 255, IPRIO is simply that number. If the interrupt’s priority number is zero or greater than 255, IPRIO is set to either 0 or 255 as follows: * If the interrupt’s priority number is greater than 255, then IPRIO is 255 (lowest representable priority). * If the interrupt’s priority number is zero and interrupt number IID has a default priority higher than a machine external interrupt, then IPRIO is 0 (highest priority). * If the interrupt’s priority number is zero and interrupt number IID has a default priority lower than a machine external interrupt, then IPRIO is 255 (lowest representable priority). | | To ensure that mtopi is never zero when an interrupt is pending and enabled for machine level, if major interrupt 0 can trap to M-mode, it must have a default priority lower than a machine external interrupt. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The value of `mtopi` is not affected by the global interrupt enable MIE in CSR `mstatus`. The RISC-V Privileged Architecture ensures that, when the value of `mtopi` is not zero, a trap is taken to M-mode for the interrupt indicated by field IID if either the current privilege mode is M and `mstatus`.MIE is one, or the current privilege mode has less privilege than M-mode. The trap itself does not cause the value of `mtopi` to change. The following pseudocode shows how a machine-level trap handler might read `mtopi` to avoid redundant restoring and saving of processor registers when an interrupt arrives during the handling of another trap (either a synchronous exception or an earlier interrupt): ```c save processor registers i = read CSR mcause if (i >= 0) { handle synchronous exception i restore mstatus if necessary } if (mstatus.MPIE == 1) { loop { i = read CSR mtopi if (i == 0) exit loop i = i>>16 call the interrupt handler for major interrupt i } } restore processor registers return from trap ``` (This example can be further optimized, but with an increase in complexity.) In order for this algorithm to function correctly, `mstatus`.MPIE must be set to 1 before executing an MRET that changes the privilege mode. | | Assuming mstatus is saved and restored by trap handlers at entry and exit as is common, it is sufficient to set mstatus.MPIE = 1 only once, before the first use of MRET that changes privilege mode. After an MRET, a trap back to M-mode will restore mstatus.MPIE = 1; and if the trap handler preserves mstatus, it will still be true before the next MRET that ends the handler. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#virtIntrs-S)5.1.3\. Interrupt filtering and virtual interrupts for supervisor level When supervisor mode is implemented, the Advanced Interrupt Architecture adds a facility for software filtering of interrupts and for virtual interrupts, making use of new CSRs `mvien` (Machine Virtual Interrupt Enables) and `mvip` (Machine Virtual Interrupt-Pending bits). _Interrupt filtering_permits a supervisor-level interrupt (SEI or SSI) or local or custom interrupt to trap to M-mode and then be selectively delegated by software to supervisor level, even while the corresponding bit in `mideleg`remains zero. The same hardware may also, under the right circumstances, allow machine level to assert _virtual interrupts_ to supervisor level that have no connection to any real interrupt events. Just as with CSRs `mip`, `mie`, and `mideleg`, each bit of registers `mvien` and `mvip` corresponds with an interrupt number in the range 0-63\. When a bit in `mideleg` is zero and the matching bit in `mvien` is one, then the same bit position in `sip` is an alias for the corresponding bit in `mvip`. A bit in `sip` is read-only zero when the corresponding bits in `mideleg` and `mvien` are both zero. The combined effects of `mideleg` and `mvien` on `sip` and `sie` are summarized in[Table 4](#TableintrFilteringForS). __Table 4\. The effects of mideleg and mvien on sip and sie (except for the H extension’s VS-level interrupts, which appear in hip and hie instead of sip and sie). A bit in mvien can be set to 1 only for major interrupts 1, 9, and 13-63\. For interrupts 0-12, some aliases of mip bits in sip may be read-only copies, as specified by the base Privileged Architecture.__ | mideleg\[ \] | mvien\[ \] | sip\[ \] | sie\[ \] | | ------------ | ---------- | ------------------ | ----------------- | | 0 | 0 | Read-only 0 | Read-only 0 | | 0 | 1 | Alias of mvip\[ \] | Writable | | 1 | \- | Alias of mip\[ \] | Alias of mie\[ \] | | | The name of CSR mvien is not "mvie" because the function of this register is more analogous to mcounteren than to mie. The bits of mvien control whether the virtual interrupt-pending bits in register mvip are active and visible at supervisor level. This is different than how the usual interrupt-enable bits (such as in mie) mask pending interrupts. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | A bit in `sie` is writable if and only if the corresponding bit is set in either `mideleg` or `mvien`. When an interrupt is delegated by `mideleg`, the writable bit in `sie` is an alias of the corresponding bit in `mie`; else it is an independent writable bit. As usual, bits that are not writable in `sie` must be read-only zeros. If a bit of `mideleg` is zero and the corresponding bit in `mvien` is changed from zero to one, then the value of the matching bit in `sie` becomes UNSPECIFIED. Likewise, if a bit of `mvien` is one and the corresponding bit in `mideleg` is changed from one to zero, the value of the matching bit in `sie` again becomes UNSPECIFIED. For interrupt numbers 13-63, implementations may freely choose which bits of `mvien` are writable and which bits are read-only zero or one. If such a bit in `mvien` is read-only zero (preventing the virtual interrupt from being enabled), the same bit should be read-only zero in `mvip`. All other bits for interrupts 13-63 must be writable in `mvip`. | | Platform standards or other extensions may require that bits of mvien for certain interrupt causes be writable, or be read-only zero or one. | | ----------------------------------------------------------------------------------------------------------------------------------------------- | The bits of `mvien` for supervisor software interrupts (code 1) and supervisor external interrupts (code 9) are each either writable or read-only zero; they cannot be read-only ones. The other bits of `mvien` for interrupts 0-12 are reserved and must be read-only zeros. It is strongly recommended that bit 9 of `mvien` be writable. Furthermore, if bit 1 (SSIP) of `mip` can be set automatically by an interrupt controller and not just by explicit writes to `mip` or `sip`, it is strongly recommended that bit 1 of `mvien` also be writable. When bit 1 of `mvien` is zero, bit 1 of `mvip` is an alias of the same bit (SSIP) of `mip`. But when bit 1 of `mvien` is one, bit 1 of `mvip` is a separate writable bit independent of `mip`.SSIP. When the value of bit 1 of `mvien` is changed from zero to one, the value of bit 1 of `mvip` becomes UNSPECIFIED. Bit 5 of `mvip` is an alias of the same bit (STIP) in `mip` when that bit is writable in `mip`. When STIP is not writable in `mip` (such as when `menvcfg`.STCE = 1), bit 5 of `mvip` is read-only zero. When bit 9 of `mvien` is zero, bit 9 of `mvip` is an alias of the software-writable bit 9 of `mip` (SEIP). But when bit 9 of `mvien` is one, bit 9 of `mvip` is a writable bit independent of `mip`.SEIP. Unlike for bit 1, changing the value of bit 9 of `mvien`does not affect the value of bit 9 of `mvip`. | | The base Privileged Architecture specifies unusual read/write behavior for what it calls the software-writable SEIP bit of register mip. When bit 9 of mvien is zero, bit 9 of mvip makes the software-writable SEIP bit of mip directly accessible by itself. Furthermore, as explained below, setting bit 9 of mvien to one separates the software-writable SEIP bit from mip entirely, so it is then just a writable bit in mvip. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Except for bits 1, 5, and 9 as specified above, the bits of `mvip` in the range 12:0 are reserved and must be read-only zeros. The value of bit 9 of `mvien` has some additional consequences for supervisor external interrupts: * When bit 9 of `mvien` is zero, the software-writable SEIP bit (bit 9 of `mvip`) interacts with reads and writes of `mip` in the way specified by the base RISC-V Privileged Architecture. In particular, for most purposes, the value of bit 9 of `mvip` is logically ORed into the readable value of `mip`.SEIP. But when bit 9 of `mvien` is one, bit SEIP in `mip` is read-only and does not include the value of bit 9 of `mvip`. Rather, the value of `mip`.SEIP is simply the supervisor external interrupt signal from the hart’s external interrupt controller (APLIC or IMSIC). * If the hart has an IMSIC, then when bit 9 of `mvien` is one, attempts from S-mode to explicitly access the supervisor-level interrupt file raise an illegal instruction exception. The exception is raised for attempts to access CSR `stopei`, or to access `sireg` when `siselect` has a value in the range `0x70`\-`0xFF`. Accesses to guest interrupt files (through `vstopei` or `viselect`+`vsireg`) are not affected. When the H extension is implemented, if a bit is zero in the same position in both `mideleg` and `mvien`, then that bit is read-only zero in `hideleg` (in addition to being read-only zero in `sip`, `sie`, `hip`, and `hie`). But if a bit for one of interrupts 13-63 is a one in either `mideleg` or `mvien`, then the same bit in `hideleg` may be writable or may be read-only zero, depending on the implementation. No bits in `hideleg` are ever read-only ones. The H extension further constrains bits 12:0 of `hideleg`. When supervisor mode is implemented, the minimal required implementation of `mvien` and `mvip` has all bits being read-only zeros except for `mvip` bits 1 and 9, and sometimes bit 5, each of which is an alias of an existing writable bit in `mip`. (Although, as noted, it is strongly recommended that bit 9 of `mvien` also be writable.) When supervisor mode is not implemented, registers `mvien` and `mvip` do not exist. ### [](#intrs-S)5.1.4\. Interrupts at supervisor level If a standard local interrupt becomes pending (= 1) in `sip`, the bit in `sip` is writable and will remain set until cleared by software. Just as for machine level, the taking of interrupt traps at supervisor level remains essentially the same as specified by the base RISC-V Privileged Architecture. An interrupt traps into S-mode (or HS-mode) whenever all of the following are true: (a) either the current privilege mode is S-mode and supervisor-level interrupts are enabled by the SIE bit of `sstatus`, or the current privilege mode has less privilege than S-mode; (b) matching bits in `sip` and `sie` are both one, or, if the H extension is implemented, matching bits in `hip` and `hie` are both one; and (c) if the H extension is implemented, the corresponding bit in `hideleg` is zero. #### [](#intrPrios-S)5.1.4.1\. Configuring priorities of major interrupts at supervisor level Supervisor-level priorities for major interrupts 0-63 are optionally configurable in an array of supervisor-level `iprio` registers accessed through `siselect` and `sireg`. This array has the same structure when XLEN = 32 or 64 as does the machine-level `iprio` array. To summarize, when XLEN = 32, there are sixteen 32-bit registers with these `siselect` addresses: | 0x30 | iprio0 | | ---- | ------- | | 0x31 | iprio1 | | …​ | …​ | | 0x3F | iprio15 | Each register controls the priorities of four interrupts, one 8-bit byte per interrupt. When XLEN = 64, only the even-numbered registers exist: | 0x30 | iprio0 | | ---- | ------- | | 0x32 | iprio2 | | …​ | …​ | | 0x3E | iprio14 | Each register controls the priorities of eight interrupts. If XLEN = 64 and `siselect` is an odd value in the range `0x31`\-`0x3F`, attempting to access `sireg` raises an illegal instruction exception. The valid registers `iprio0`\-`iprio15` are known collectively as the `_iprio_` _array_ for supervisor level. Each byte of a valid `iprio` register is either a read-only zero or a **WARL** unsigned integer field implementing exactly IPRIOLEN bits. For a given interrupt number, if the corresponding bit is not writable either in `sie` or, if the H extension is implemented, in `hie`, then the interrupt’s priority number in the supervisor-level `iprio` array must be read-only zero as well. The priority number for a supervisor-level external interrupt (bits 15:8 of `iprio2`) must also be read-only zero. Aside from these two restrictions, implementations may freely choose which priority number fields are settable and which are read-only zeros. | | As always, platform standards may require that priorities be configurable for certain interrupt causes. | | ---------------------------------------------------------------------------------------------------------- | | | It is expected that many higher-end systems will not support the ability to configure the priorities of major interrupts at supervisor level as described in this section. Linux in particular is not designed to take advantage of such facilities if provided. The iprio array must be accessible but may simply be all read-only zeros. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The supervisor-level `iprio` array accessed via `siselect` and `sireg` affects the prioritization of interrupts only when they trap to S-mode. When an interrupt’s priority number in the array is zero (either read-only zero or set to zero), its priority is the default order from [5.1.1\. Defined major interrupts and default priorities](#majorIntrs). Setting an interrupt’s priority number instead to a nonzero value gives that interrupt nominally the same priority as a supervisor-level external interrupt with priority number . For an interrupt that defaults to a higher priority than supervisor external interrupts, setting its priority number to a nonzero value lowers its priority. For an interrupt that defaults to a lower priority than supervisor external interrupts, setting its priority number to a nonzero value raises its priority. When two interrupt causes have been assigned the same nominal priority, ties are broken by the default priority order. [Table 5](#TableintrPrios-S) summarizes the effect of priority numbers on interrupt priority. If supported, setting the priority number for VS-level external interrupts (bits 23:16 of `iprio2`) to a nonzero value has the effect of giving the entire category of VS external interrupts nominally the same priority as a supervisor external interrupt with priority number , when VS external interrupts trap to S-mode. __Table 5\. Effect of the supervisor-level iprio array on the priorities of interrupts taken in S-mode. For interrupts with the same priority number, the default order of [5.1.1\. Defined major interrupts and default priorities](#majorIntrs) prevails.__ | Interrupts with default priority above supervisor external interrupts | Supervisor external interrupts | Interrupts with default priority below supervisor external interrupts | | | --------------------------------------------------------------------- | ----------------------------------------------- | --------------------------------------------------------------------- | ----------------------------------------------- | | Priorityorder | Priority number in supervisor-level iprio array | Priority number from interrupt controller (APLIC or IMSIC) | Priority number in supervisor-level iprio array | | Highest | 0 | | | | 12…​254255 | 12…​254255 | 12…​254255 | | | 256 and above (IMSIC only) | | | | | Lowest | 0 | | | If bit 9 for a supervisor external interrupt (SEI) is one in `mideleg` or `mvien` and in `mvip`, causing `sip`.SEIP to be one, but there is no supervisor-level interrupt from the hart’s external interrupt controller (APLIC or IMSIC), then a priority number for the SEI is not supplied by the external interrupt controller as usual. In that case, the SEI is assigned a priority number of 256. If the system has an original PLIC for backward compatibility with older software, reset should initialize the supervisor-level `iprio` array to all zeros. #### [](#stopi)5.1.4.2\. Supervisor top interrupt CSR (`stopi`) Supervisor-level CSR `stopi` is read-only with width SXLEN. A read of `stopi` returns information about the highest-priority pending-and-enabled interrupt for supervisor level, in this format: | bits 27:16 | IID | | ---------- | ----- | | bits 7:0 | IPRIO | All other bits of `stopi` are reserved and read as zeros. The value of `stopi` is zero unless: (a) there is an interrupt that is both pending in `sip` and enabled in `sie`, or, if the H extension is implemented, both pending in `hip` and enabled in `hie`; and (b) the interrupt is not delegated to a less-privileged level (by `hideleg`, if the H extension is implemented). When there is a pending-and-enabled major interrupt for supervisor level, field IID is the major identity number of the highest-priority interrupt, and field IPRIO indicates its priority. If all bytes of the supervisor-level `iprio` array are read-only zeros, a simplified implementation of field IPRIO is allowed in which its value is always 1 whenever `stopi` is not zero. Otherwise, when `stopi` is not zero, if the priority number for the reported interrupt is in the range 1 to 255, IPRIO is simply that number. If the interrupt’s priority number is zero or greater than 255, IPRIO is set to either 0 or 255 as follows: * If the interrupt’s priority number is greater than 255, then IPRIO is 255 (lowest representable priority). * If the interrupt’s priority number is zero and interrupt number IID has a default priority higher than a supervisor external interrupt, then IPRIO is 0 (highest priority). * If the interrupt’s priority number is zero and interrupt number IID has a default priority lower than a supervisor external interrupt, then IPRIO is 255 (lowest representable priority). | | To ensure that stopi is never zero when an interrupt is pending and enabled for supervisor level, if major interrupt 0 can trap to S-mode, it must have a default priority lower than a supervisor external interrupt. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The value of `stopi` is not affected by the global interrupt enable SIE in CSR `sstatus`. The RISC-V Privileged Architecture ensures that, when the value of `stopi` is not zero, a trap is taken to S-mode for the interrupt indicated by field IID if either the current privilege mode is S and `sstatus`.SIE is one, or the current privilege mode has less privilege than S-mode. The trap itself does not cause the value of `stopi` to change. The following pseudocode shows how a supervisor-level trap handler might read `stopi` to avoid redundant restoring and saving of processor registers when an interrupt arrives during the handling of another trap (either a synchronous exception or an earlier interrupt): ```c save processor registers i = read CSR scause if (i >= 0) { handle synchronous exception i restore sstatus if necessary } if (sstatus.SPIE == 1) { loop { i = read CSR stopi if (i == 0) exit loop i = i>>16 call the interrupt handler for major interrupt i } } restore processor registers return from trap ``` (This example can be further optimized, but with an increase in complexity.) In order for this algorithm to function correctly, `sstatus`.SPIE must be set to 1 before executing an SRET that changes the privilege mode. | | Assuming sstatus is saved and restored by trap handlers at entry and exit as is common, it is sufficient to set sstatus.SPIE = 1 only once, before the first use of SRET that changes privilege mode. After an SRET, a trap back to S-mode will restore sstatus.SPIE = 1; and if the trap handler preserves sstatus, it will still be true before the next SRET that ends the handler. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#5-1-5-wfi-wait-for-interrupt-instruction)5.1.5\. WFI (Wait for Interrupt) instruction The RISC-V Privileged Architecture specifies that instruction WFI (Wait for Interrupt) may suspend execution at a hart until an interrupt is pending for the hart. The Advanced Interrupt Architecture (AIA) redefines when execution must resume following a WFI. According to the base RISC-V Privileged Architecture, instruction execution must resume from a WFI whenever any interrupt is both pending and enabled in CSRs `mip` and `mie`, ignoring any delegation indicated by `mideleg`. With the AIA, this succinct rule is no longer appropriate, due to the mechanisms the AIA adds for virtual interrupts. Instead, execution must resume from a WFI whenever an interrupt is pending at any privilege level (regardless of whether the interrupt privilege level is higher or lower than the hart’s current privilege mode). An interrupt is pending at machine level if register `mtopi` is not zero. If S-mode is implemented, an interrupt is pending at supervisor level if `stopi` is not zero. And if the H extension is implemented, an interrupt is pending at VS level if `vstopi` ([Virtual supervisor top interrupt CSR (vstopi)](VSLevel.html#vstopi)) is not zero. | | The AIA’s rule for WFI gives the same behavior as the base Privileged Architecture’s rule when mvien \= 0 and, if the H extension is implemented, also hvien \= 0 and hvictl.VTI = 0, thus disabling all virtual interrupts not visible in mip. (The AIA’s hypervisor registers are covered in the next chapter, "Interrupts for Virtual Machines (VS Level)".) | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | 6.1. Interrupts for Virtual Machines (VS Level) ==================== ## [](#VSLevel)6.1\. Interrupts for Virtual Machines (VS Level) When the H extension is implemented, a hart’s set of possible privilege modes includes the _virtual supervisor_ (VS) and _virtual user_ (VU) modes for hosting virtual harts. The Advanced Interrupt Architecture adds to the H extension new interrupt facilities aligned with those described earlier for supervisor-level interrupts. As introduced in [Control and Status Registers (CSRs) Added to Harts](CSRs.html#CSRs), several hypervisor and VS CSRs are added: `hvien`, `hvictl`, `hviprio1`, `hviprio2`, `vsiselect`, `vsireg`, `vstopei`, and `vstopi`. (And for RV32, the following high-half CSRs are also added: `hidelegh`, `hvienh`, `hviph`, `hviprio1h`, `hviprio2h`, `vsiph` and `vsieh`.) As always, when executing in VS-mode or VU-mode, the VS CSRs substitute for the corresponding supervisor CSRs. To give software that runs in a virtual machine the appearance of executing on a real machine that implements the Advanced Interrupt Architecture at supervisor level, responsibility is shared between hypervisor software and the hardware facilities described in this chapter. While some behaviors can be handled directly by hardware, others require significant emulation by the hypervisor, sometimes with hardware assistance. ### [](#6-1-1-vs-level-external-interrupts-with-a-guest-interrupt-file)6.1.1\. VS-level external interrupts with a guest interrupt file When a hart implements the H extension, it is recommended that the hart also have an IMSIC with guest interrupt files. Assuming guest interrupt files are available, each can be assigned to a virtual hart at the physical hart to act as the supervisor-level interrupt file for that virtual hart. If there are guest interrupt files, then virtual harts at that physical hart may each have a physical guest interrupt file to serve as its (virtual) supervisor-level interrupt file. The guest interrupt file for the current virtual hart is always indicated by the VGEIN field of CSR `hstatus`. When VGEIN is not the valid number of a guest interrupt file, the current virtual hart has no guest interrupt file to act as its supervisor-level interrupt file. When `hstatus`.VGEIN is the valid number of a guest interrupt file, values of `vsiselect` in the range `0x70`\-`0xFF` select registers of this guest interrupt file, just as values of `siselect` in the same range select registers of the IMSIC’s true supervisor-level interrupt file. The registers of an interrupt file that are accessed indirectly through `vsiselect` and `vsireg` are documented in[Incoming MSI Controller (IMSIC)](IMSIC.html#IMSIC) on the IMSIC, along with IMSIC-only CSR `vstopei`. Because all IMSIC interrupt files act identically, the guest interrupt file that a virtual hart accesses through CSRs `siselect`, `sireg`, and `stopei` is indistinguishable from a true supervisor-level interrupt file as seen from S-mode (or HS-mode). In addition to an IMSIC at each hart, a virtual machine may also need to see a PLIC or APLIC. However, unlike an IMSIC’s ability to provide physical guest interrupt files for virtual harts, a PLIC or APLIC must be emulated for a virtual machine by the hypervisor. | | The Advanced Interrupt Architecture does not currently include hardware assistance for virtualizing an APLIC. For small numbers of harts, such hardware would be substantially larger than that required to implement guest interrupt files for an IMSIC. It is assumed that most high-performance I/O can be done through devices that can send MSIs directly to guest interrupt files (such as devices attached through a PCI Express interconnect). For the types of devices whose interrupts must go through a (virtual) APLIC, the overhead cost of emulating the APLIC is expected to be less significant. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When a virtual hart appears to have an IMSIC because a guest interrupt file is assigned to it, all external interrupts, real or emulated, destined for the virtual hart must go through this perceived IMSIC. A hypervisor can easily inject an emulated external interrupt into the guest interrupt file selected by `hstatus`.VGEIN by setting a bit in the interrupt-pending array indirectly accessed through `vsiselect` and `vsireg`. When a virtual hart has a guest interrupt file, a hypervisor is not normally expected to set bit VSEIP in CSR `hvip`. In the special case that an emulated APLIC for a virtual machine has a wired interrupt source that equates to an actual interrupt source of a real APLIC, if software running in this virtual machine configures its virtual APLIC to forward interrupts from that source as MSIs to a specific virtual hart, the hypervisor can configure the real APLIC to forward the actual interrupts directly as MSIs to the virtual hart’s guest interrupt file. In this way, although the hypervisor must trap and emulate the virtual machine’s memory accesses that configure the forwarding of interrupts at the virtual APLIC, the interrupts themselves can be converted automatically into real MSIs for the guest interrupt file, without the hypervisor being invoked for each arriving interrupt. #### [](#6-1-1-1-direct-control-of-a-device-by-a-guest-os)6.1.1.1\. Direct control of a device by a guest OS To ensure proper support for interrupts, two conditions must be met before a hypervisor may allow a guest OS running in a virtual machine to directly control a physical device that sends MSIs: First, each virtual hart must have a guest interrupt file assigned to it, giving each its own apparent IMSIC within the virtual machine. Second, interrupts from the device must be signaled by wire through an APLIC that can translate these interrupts into MSIs, or the system must have an IOMMU that can translate the addresses of MSI memory writes made by the device itself. If a guest OS directly controls a device capable of sending MSIs, it will naturally configure MSIs at the device with the guest physical addresses the OS sees for the IMSICs of its virtual harts, not knowing that these addresses are only virtual. When the device performs a memory write for an MSI, the destination address of this write must be translated by the IOMMU from the guest physical address assigned by the guest OS into the true physical address of the target guest interrupt file, using a translation table supplied by the hypervisor. By design, the translation an IOMMU must do for device MSIs is fundamentally no different than the address translation the IOMMU already must perform for other memory accesses from the same device, converting guest physical addresses into true physical addresses. Because each virtual hart is assigned a dedicated, physical guest interrupt file that is indistinguishable from a true supervisor-level interrupt file, no translation is needed for the data of an MSI write, which specifies the interrupt’s identity number in the target interrupt file. #### [](#virtHartMigration)6.1.1.2\. Migrating a virtual hart to a different guest interrupt file When it is necessary to move a virtual hart from one physical hart to another, if the virtual hart uses a guest interrupt file, the specific guest interrupt file assigned to it must change from the one in use at the old physical hart to a different one at the new physical hart. Because each guest interrupt file is physically tied to a single physical hart, a virtual hart cannot bring its guest interrupt file with it when it moves. The process of migrating a virtual hart from one guest interrupt file to another is more complex than moving most other state held by the virtual hart. After the destination guest interrupt file has been chosen at the new physical hart, the following steps are recommended: 1. At the old interrupt file, save to memory the values of registers `eidelivery` and`eithreshold`, and set `eidelivery` \= 0. 2. At the new interrupt file, set `eidelivery` \= 0, and zero all implemented interrupt-pending bits (the `eip` array). 3. Modify the relevant translation tables at all IOMMUs so that MSIs for this virtual interrupt file are now sent to the new physical interrupt file. Likewise, if any interrupts at an APLIC are forwarded by MSIs to the old interrupt file, reconfigure the APLIC to send them to the new interrupt file. As needed, synchronize with all IOMMUs and APLICs to ensure that no straggler MSIs will arrive at the old interrupt file after this step. Synchronizing with an APLIC can be accomplished using the algorithm of [Synchronizing interactions between a hart and the APLIC](AdvPLIC.html#AdvPLIC-MSISync). 4. At the old interrupt file, dump to memory all implemented interrupt-pending and interrupt-enable bits (the `eip` and `eie` arrays). After this step is done, the old interrupt file is no longer in use. 5. At the new interrupt file, logically OR the interrupt-pending bits that were saved in step 4 into the new interrupt file, using instruction CSRS to write to the `eip` array. Also, load the interrupt-enable bits that were saved in step 4 into the `eie` array. 6. At the new interrupt file, load registers `eithreshold` and `eidelivery` with the values that were saved in step 1. Resuming execution of the virtual hart at the new physical hart is not recommended until the entire interrupt file has been fully migrated. | | Resuming execution of the virtual hart before the interrupt file is fully migrated could allow software running in the virtual machine to see multiple MSIs arriving from a single device in an order that should not happen. While this would rarely matter in practice, it runs the risk of wedging a device driver that depends (perhaps inadvertently) on a valid ordering of events. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#6-1-2-vs-level-external-interrupts-without-a-guest-interrupt-file)6.1.2\. VS-level external interrupts without a guest interrupt file Although it is recommended that harts implementing the hypervisor extension also have IMSICs with guest interrupt files, this is not a requirement. Even assuming guest interrupt files exist, it may happen that there are more virtual harts at a physical hart than guest interrupt files, leaving some virtual harts without one. In either case, a hypervisor must emulate an external interrupt controller for a virtual hart without the benefit of a guest interrupt file allocated to the virtual hart. When emulating an external interrupt controller for a virtual hart, if configurable interrupt priority is not supported for the virtual hart other than for external interrupts, then external interrupts may be asserted to VS level simply by setting bit VSEIP in `hvip`, as defined by the H extension. However, to emulate both an external interrupt controller and priority configurability for non-external interrupts, a hypervisor must make use of CSR `hvictl` (Hypervisor Virtual Interrupt Control), described later in the next section. ### [](#6-1-3-interrupts-at-vs-level)6.1.3\. Interrupts at VS level #### [](#6-1-3-1-configuring-priorities-of-major-interrupts-at-vs-level)6.1.3.1\. Configuring priorities of major interrupts at VS level Like for supervisor level, the Advanced Interrupt Architecture optionally allows major VS-level interrupts to be configured by software to intermix in priority with VS-level external interrupts. As documented in [Interrupts at supervisor level](MSLevel.html#intrs-S), interrupt priorities for supervisor level are configured by the `iprio` array accessed indirectly through CSRs `siselect` and `sireg`. The `siselect` addresses for the `iprio` array registers are `0x30`\-`0x3F`. VS level has its own `vsiselect` and `vsireg`, but unlike supervisor level, there are no registers at `vsiselect` addresses `0x30`\-`0x3F`. When `vsiselect` has a value in the range `0x30`\-`0x3F`, an attempt from VS-mode to access `sireg` (really `vsireg`) causes a virtual instruction exception. To give a virtual hart the illusion of an array of `iprio` registers accessed through `siselect` and `sireg`, a hypervisor must emulate the VS-level `iprio` array when accesses to `sireg` from VS-mode cause virtual instruction traps. Instead of a physical VS-level `iprio` array, a separate hardware mechanism is provided for configuring the priorities of a subset of interrupts for VS level, using hypervisor CSRs `hviprio1` and `hviprio2`. The subset of major interrupt numbers whose priority may be configured in hardware are these: | 1 | Supervisor software interrupt | | ----- | ---------------------------------------- | | 5 | Supervisor timer interrupt | | 13 | Counter overflow interrupt | | 14-23 | _Reserved for standard local interrupts_ | For interrupts directed to VS level, software-configurable priorities are not supported in hardware for standard local interrupts in the range 32-48. | | For custom interrupts, priority configurability may be supported in hardware by custom CSRs, expanding upon hviprio1 and hviprio2 for standard interrupts. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | Registers `hviprio1` and `hviprio2` have these formats: `hviprio1`: | bits 7:0 | _Reserved for priority number for interrupt 0; reads as zero_ | | ---------- | -------------------------------------------------------------- | | bits 15:8 | Priority number for interrupt 1, supervisor software interrupt | | bits 23:16 | _Reserved for priority number for interrupt 4; reads as zero_ | | bits 31:24 | Priority number for interrupt 5, supervisor timer interrupt | | bits 39:32 | _Reserved for priority number for interrupt 8; reads as zero_ | | bits 47:40 | Priority number for interrupt 13, counter overflow interrupt | | bits 55:48 | Priority number for interrupt 14 | | bits 63:56 | Priority number for interrupt 15 | `hviprio2`: | bits 7:0 | Priority number for interrupt 16 | | ---------- | -------------------------------- | | bits 15:8 | Priority number for interrupt 17 | | bits 23:16 | Priority number for interrupt 18 | | bits 31:24 | Priority number for interrupt 19 | | bits 39:32 | Priority number for interrupt 20 | | bits 47:40 | Priority number for interrupt 21 | | bits 55:48 | Priority number for interrupt 22 | | bits 63:56 | Priority number for interrupt 23 | Each priority number in `hviprio1` and `hviprio2` is a **WARL** unsigned integer field that is either read-only zero or implements a minimum of IPRIOLEN bits or 6 bits, whichever is larger, and preferably all 8 bits. Implementations may freely choose which priority number fields are read-only zeros, but all other fields must implement the same number of integer bits. A minimal implementation of these CSRs has them both be read-only zeros. A hypervisor can choose to employ registers `hviprio1` and `hviprio2` when emulating the (virtual) supervisor-level `iprio` array accessed indirectly through `siselect` and `sireg` (really`vsiselect` and `vsireg`) for a virtual hart. For interrupts not in the subset supported by`hviprio1` and `hviprio2`, the priority number bytes in the emulated `iprio` array can be read-only zeros. | | Providing hardware support for configurable priority for only a subset of major interrupts at VS level is a compromise. The utility of being able to control interrupt priorities at VS level is arguably illusory when all traps to M-mode and HS-mode—both interrupts and synchronous exceptions—have absolute priority, and when each virtual hart may also be competing for resources against other virtual harts well beyond its control. Nevertheless, priority configurability has been made possible for the most likely subset of interrupts, while minimizing the number of added CSRs that must be swapped on a virtual hart switch. Major interrupts outside the priority-configurable subset can still be directed to VS level, but their priority will simply be the default order defined in [Defined major interrupts and default priorities](MSLevel.html#majorIntrs). | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | If a hypervisor really must emulate configurability of priority for interrupts beyond the subset supported by `hviprio1` and `hviprio2`, it can do so with extra effort by setting bit VTI of CSR `hvictl`, described in the next subsection. #### [](#6-1-3-2-virtual-interrupts-for-vs-level)6.1.3.2\. Virtual interrupts for VS level Assuming a virtual hart does not need configurable priority for major interrupts beyond the subset supported in hardware by `hviprio1` and `hviprio2`, a hypervisor can assert interrupts to the virtual hart using CSRs `hvien` (Hypervisor Virtual-Interrupt-Enable) and `hvip` (Hypervisor Virtual-Interrupt-Pending bits). These CSRs affect interrupts for VS level much the same way that `mvien`and `mvip` do for supervisor level, as explained in[Interrupt filtering and virtual interrupts for supervisor level](MSLevel.html#virtIntrs-S). Each bit of registers `hvien` and `hvip` corresponds with an interrupt number in the range 0-63\. Bits 12:0 of `hvien` are reserved and must be read-only zeros, while bits 12:0 of `hvip` are defined by the H extension. Specifically, bits 10, 6, and 2 of `hvip` are writable bits that correspond to VS-level external interrupts (VSEIP), VS-level timer interrupts (VSTIP), and VS-level software interrupts (VSSIP), respectively. The following applies only to the CSR bits for interrupt numbers 13-63: When a bit in `hideleg` is one, then the same bit position in `vsip` is an alias for the corresponding bit in `sip`. Else, when a bit in `hideleg` is zero and the matching bit in `hvien` is one, the same bit position in `vsip` is an alias for the corresponding bit in `hvip`. A bit in `vsip` is read-only zero when the corresponding bits in `hideleg` and `hvien`are both zero. The combined effects of `hideleg` and `hvien` on `vsip` and `vsie` are summarized in [Table 1](#intrFilteringForVS). __Table 1\. The effects of hideleg and hvien on vsip and vsie for major interrupts 13-63.__ | hideleg\[ \] | hvien\[ \] | vsip\[ \] | vsie\[ \] | | ------------ | ---------- | ------------------ | ----------------- | | 0 | 0 | Read-only 0 | Read-only 0 | | 0 | 1 | Alias of hvip\[ \] | Writable | | 1 | \- | Alias of sip\[ \] | Alias of sie\[ \] | For interrupt numbers 13-63, a bit in `vsie` is writable if and only if the corresponding bit is set in either `hideleg` or `hvien`. When an interrupt is delegated by `hideleg`, the writable bit in `vsie` is an alias of the corresponding bit in `sie`; else it is an independent writable bit. The H extension specifies when bits 12:0 of `vsie` are aliases of bits in `hie`. As usual, bits that are not writable in `vsie` must be read-only zeros. If a bit of `hideleg` is zero and the corresponding bit in `hvien` is changed from zero to one, then the value of the matching bit in `vsie` becomes UNSPECIFIED. Likewise, if a bit of `hvien` is one and the corresponding bit in `hideleg` is changed from one to zero, the value of the matching bit in `vsie` again becomes UNSPECIFIED. For interrupt numbers 13-63, implementations may freely choose which bits of `hvien` are writable and which bits are read-only zero or one. If such a bit in `hvien` is read-only zero (preventing the virtual interrupt from being enabled), the same bit should be read-only zero in `hvip`. All other bits for interrupts 13-63 must be writable in `hvip`. CSR `hvictl` (Hypervisor Virtual Interrupt Control) provides further flexibility for injecting interrupts into VS level in situations not fully supported by the facilities described thus far, but only with more active involvement of the hypervisor. A hypervisor must use `hvictl` for any of the following: * asserting for VS level a major interrupt not supported by `hvien` and `hvip`; * implementing configurability of priorities at VS level for major interrupts beyond those supported by `hviprio1` and `hviprio2`; or * emulating an external interrupt controller for a virtual hart without the use of an IMSIC’s guest interrupt file, while also supporting configurable priorities both for external interrupts and for major interrupts to the virtual hart. | | Among the possible uses, hvictl is needed for a hypervisor to fully emulate HS-mode in VS-mode, which is a requirement for the hosting of nested hypervisors without paravirtualization. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The format of `hvictl` is: | bit 30 | VTI | | ---------- | -------------- | | bits 27:16 | IID (**WARL**) | | bit 9 | DPR | | bit 8 | IPRIOM | | bits 7:0 | IPRIO | All other bits of `hvictl` are reserved and read as zeros. When bit VTI (Virtual Trap Interrupt control) = 1, attempts from VS-mode to explicitly access CSRs `sip` and `sie` (or, for RV32 only, `siph` and `sieh`) cause a virtual instruction exception. Furthermore, for any given CSR, if there is some circumstance in which a write to the register may cause a bit of `vsip` to change from one to zero, excluding bit 9 for external interrupts (SEIP), then when VTI = 1, a virtual instruction exception is raised also for any attempt by the guest to write this register. Both the value being written to the CSR and the value of `vsip` (before or after) are ignored for determining whether to raise the exception. (Hence a write would not actually need to change a bit of `vsip` from one to zero for the exception to be raised.) In particular, if register `vstimecmp` is implemented (from extension Sstc), then attempts from VS-mode to write to `stimecmp` (or, for RV32 only, `stimecmph`) cause a virtual instruction exception when VTI = 1. | | For the standard local interrupts (major identities 13-23 and 32-47), and for software interrupts (SSI), the corresponding interrupt-pending bits in vsip are defined as "sticky," meaning a guest can clear them only by writing directly to sip (really vsip). Among the standard-defined interrupts, that leaves only timer interrupts (STI), which can potentially be cleared in vsip by writing a new value to vstimecmp. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | All `hvictl` fields together can affect the value of CSR `vstopi` (Virtual Supervisor Top Interrupt) and therefore the interrupt identity reported in `vscause` when an interrupt traps to VS-mode. IID is a **WARL** unsigned integer field with at least 6 implemented bits, while IPRIO is always the full 8 bits. If bits are implemented for IID, then all values 0 through are supported, and a write to `hvictl` sets IID equal to bits ( ):16 of the value written. For a virtual interrupt specified for VS level by `hvictl`, if VTI = 1 and , field DPR (Default Priority Rank) determines the interrupt’s presumed default priority order relative to a (virtual) supervisor external interrupt (SEI), major identity 9, as follows: | 0 = interrupt has higher default priority than an SEI | | ----------------------------------------------------- | | 1 = interrupt has lower default priority than an SEI | When `hvictl`.IID = 9, DPR is ignored. | | Register hvictl has no effect on any of mip, sip, hip, or vsip; it affects only vstopi and the trapping of some instructions. | | -------------------------------------------------------------------------------------------------------------------------------- | #### [](#vstopi)6.1.3.3\. Virtual supervisor top interrupt CSR (`vstopi`) Read-only CSR `vstopi` is VSXLEN bits wide and has the same format as `stopi`: | bits 27:16 IID | | -------------- | | bits 7:0 IPRIO | `vstopi` returns information about the highest-priority interrupt for VS level, found from among these candidates (prefixed by + signs): * if bit 9 is one in both `vsip` and `vsie`, `hstatus`.VGEIN is the valid number of a guest interrupt file, and `vstopei` is not zero: * \+ a supervisor external interrupt (code 9) with the priority number indicated by `vstopei`; * if bit 9 is one in both `vsip` and `vsie`, `hstatus`.VGEIN = 0, and `hvictl` fields IID = 9 and : * \+ a supervisor external interrupt (code 9) with priority number `hvictl`.IPRIO; * if bit 9 is one in both `vsip` and `vsie`, and neither of the first two cases applies: * \+ a supervisor external interrupt (code 9) with priority number 256; * if `hvictl`.VTI = 0: * \+ the highest-priority pending-and-enabled major interrupt indicated by `vsip` and `vsie`other than a supervisor external interrupt (code 9), using the priority numbers assigned by `hviprio1` and `hviprio2`; * if `hvictl` fields VTI = 1 and : * \+ the major interrupt specified by `hvictl` fields IID, DPR, and IPRIO. In the list above, all "supervisor" external interrupts are virtual, directed to VS level, having major code 9 at VS level. | | The list of candidate interrupts can be reduced to two finalists relatively easily by observing that the first three list items are mutually exclusive of one another, and the remaining two items are also mutually exclusive of one another. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | When hvictl.VTI = 1, the absence of an interrupt for VS level can be indicated only by setting hvictl.IID = 9\. Software might want to use the pair IID = 9, IPRIO = 0 generally to represent _no interrupt_ in hvictl. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When no interrupt candidates satisfy the conditions of the list above,`vstopi` is zero. Else, `vstopi` fields IID and IPRIO are determined by the highest-priority interrupt from among the candidates. The usual priority order for supervisor level applies, as specified by[Effect of the supervisor-level iprio array on the priorities of interrupts taken in S-mode](MSLevel.html#TableintrPrios-S), except that priority numbers are taken from the candidate list above, not from the supervisor-level `iprio` array. Ties in nominal priority are broken as usual by the default priority order from[The standard major interrupt codes, listed in default priority order](MSLevel.html#TablemajorIntrs), unless `hvictl` fields VTI = 1 and (last item in the candidate list above), in which case default priority order is determined solely by`hvictl`.DPR. If bit IPRIOM (IPRIO Mode) of `hvictl` is zero, IPRIO in `vstopi` is 1; else, if the priority number for the highest-priority candidate is within the range 1 to 255, IPRIO is that value; else, IPRIO is set to either 0 or 255 in the manner documented for `stopi` in [Supervisor top interrupt CSR (stopi)](MSLevel.html#stopi). #### [](#6-1-3-4-interrupt-traps-to-vs-mode)6.1.3.4\. Interrupt traps to VS-mode The Advanced Interrupt Architecture modifies the H extension such that an interrupt is pending at VS level if and only if `vstopi` is not zero. CSRs `vsip` and `vsie` do not by themselves determine whether a VS-level interrupt is pending, though they may do so indirectly through their effect on `vstopi`. Whenever `vstopi` is not zero, if either the current privilege mode is VS-mode and the SIE bit in CSR `vsstatus` is one, or the current privilege mode is VU-mode, a trap is taken to VS-mode for the interrupt indicated by field IID of `vstopi`. The Exception Code field of `vscause` must implement at least as many bits as needed to represent the largest value that field IID of `vstopi` can have for the given hart. Preface ==================== ## [](#preface)Preface This document describes an Advanced Interrupt Architecture (AIA) for RISC-V systems. This specification was ratified by the RISC-V International Association in June of 2023. The table below indicates which chapters of this document specify extensions to the RISC-V ISA (instruction set architecture) and which are non-ISA. | Chapter | ISA? | | -------------------------------------------------------- | ---- | | 1\. Introduction | — | | 2\. Control and Status Registers (CSRs) Added to Harts | Yes | | 3\. Incoming MSI Controller (IMSIC) | Yes | | 4\. Advanced Platform-Level Interrupt Controller (APLIC) | No | | 5\. Interrupts for Machine and Supervisor Levels | Yes | | 6\. Interrupts for Virtual Machines (VS Level) | Yes | | 7\. Interprocessor Interrupts (IPIs) | No | | 8\. IOMMU Support for MSIs to Virtual Machines | No | **Changes for version 20250312** Made the following clarifications to AIA 1.0: * Where there are irreconcilable conflicts between the AIA and other implemented RISC-V extensions, the AIA usually has priority by default. * Deference is given to extension Smcsrind/Sscsrind (indirectly accessed CSRs). * Names are given to the bits defined in `mstateen0` and `hstateen0`when extension Smstateen/Ssstateen is also implemented. * An IMSIC interrupt file’s `eidelivery` register affects only whether an interrupt appears in a hart’s `mip` or `hgeip` register. * IMSIC CSRs `mtopei`, `stopei`, and `vstopei` are not affected by the values of `mie`, `sie`, `hie`, `hgeie`, or `vsie`. * There may be a visible delay between a change of state of an IMSIC interrupt file and its effect on a bit in `mip`, `sip`, or `hgeip`. * An APLIC’s `idelivery` registers and the IE bits of its `domaincfg` registers affect only whether pending-and-enabled interrupts are delivered to harts. * The default priority order for major interrupts is applicable only when multiple interrupts would trap to the same privilege mode. * The example pseudocode given for handling major interrupts at M-level and S-level has additional requirements not mentioned previously. * An interrupt priority number in the S-level `iprio` array may be writable (not read-only zero) if the correponding bit is writable in either `sie` or `hie`. * If a supervisor external interrupt (SEI) is injected from M-level when there is no actual interrupt from an external interrupt controller, the injected SEI is assigned an S-level priority number of 256. * CSR `hvictl` affects only `vstopi` and the trapping of some instructions, not `mip`, `sip`, `hip`, or `vsip`. **Changes for the ratified version 1.0** Resolved some inconsistencies in [Control and Status Registers (CSRs) Added to Harts](CSRs.html#CSRs) about when to raise a virtual instruction exception versus an illegal instruction exception. **Changes for RC5 (Release Candidate 5)** Better aligned the rules for indirectly accessed registers with the hypervisor extension and with forthcoming extension Smcsrind/Sscsrind. In particular, when `vsiselect` has a reserved value, attempts to access `sireg` from a virtual machine (VS or VU-mode) should preferably raise an illegal instruction exception instead of a virtual instruction exception. Added clarification about the term _IOMMU_ used in [IOMMU Support for MSIs to Virtual Machines](IOMMU.html#IOMMU). Added clarification about MSI write replaced by MRIF update and notice MSI sent after the update. **Changes for RC4** For alignment with other forthcoming RISC-V ISA extensions, the widths of the indirect-access CSRs, `miselect`, `mireg`, `siselect`,`sireg`, `vsiselect`, and `vsireg`, were changed to all be the current XLEN rather than being tied to their respective privilege levels (previously MXLEN for `miselect` and `mireg`, SXLEN for `siselect`and `sireg`, and VSXLEN for `vsiselect` and `vsireg`). Changed the description (but not the actual function) of _high-half_CSRs and their partner CSRs to match the latest RISC-V Privileged ISA specification. (An example of a high-half CSR is `miph`, and its partner here is `mip`.) **Changes for RC3** Removed the still-draft Duo-PLIC chapter to a separate document. Allocated major interrupts 35 and 43 for signaling RAS events ([Defined major interrupts and default priorities](MSLevel.html#majorIntrs)). In [Interrupt filtering and virtual interrupts for supervisor level](MSLevel.html#virtIntrs-S) added the options for bits 1 and 9 to be writable in CSR `mvien`, and specified the effects of setting each of these bits. Upgraded [IOMMU Support for MSIs to Virtual Machines](IOMMU.html#IOMMU) ("IOMMU Support") to the _frozen_ state. **Changes for RC2** Clarified that field IID of CSR `hvictl` must support all unsigned integer values of the number of bits implemented for that field, and that writes to `hvictl` always set IID in the most straightforward way. A comment was added to [Interprocessor Interrupts (IPIs)](IPIs.html#IPIs) warning about the possible need for FENCE instructions when IPIs are sent to other harts by writing MSIs to those harts' IMSICs. 1.1. Introduction ==================== ## [](#ch:intro)1.1\. Introduction This document specifies the Advanced Interrupt Architecture for RISC-V, consisting of: (a) an extension to the standard RISC-V Privileged Architecture; (b) two standard interrupt controllers for RISC-V systems, an Advanced Platform-Level Interrupt Controller (APLIC) and an Incoming Message-Signaled Interrupt Controller (IMSIC); and (c) requirements on other system components concerning interrupts. | | Commentary on our design decisions, implementation options, and application is formatted as in this paragraph, and can be skipped if the reader is only interested in the specification itself. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#1-1-1-goals)1.1.1\. Goals The RISC-V Advanced Interrupt Architecture has these goals: * Build upon the interrupt-handling functionality of the RISC-V Privileged Architecture, minimizing the replacement of existing functionality. * Provide facilities for RISC-V systems to work directly with message-signaled interrupts (MSIs) as employed by PCI Express and other device standards, in addition to basic wired interrupts. * For wired interrupts, define a new Platform-Level Interrupt Controller (the Advanced PLIC, or APLIC) that has an independent control interface for each level of privilege (such as RISC-V machine and supervisor levels), and that can convert wired interrupts into MSIs for systems supporting MSIs. * Expand the framework for local interrupts at a RISC-V hart. * Optionally allow software to configure the relative priorities of all sources of interrupts to a RISC-V hart (including the standard timer and software interrupts, among others), instead of being limited just to the ability of a separate interrupt controller to prioritize external interrupts only. * When harts implement the Privileged Architecture’s H extension, provide sufficient assistance for virtualizing these same interrupt facilities for virtual machines. * With the help of an IOMMU (I/O memory management unit) for redirecting MSIs, maximize the opportunities and ability for a guest operating system running in a virtual machine to have direct control of devices with minimal involvement of a hypervisor. * Avoid having the interrupt hardware be a limiter on the number of virtual machines. * Achieve all of the above with the best possible compromises between speed, efficiency, and flexibility of implementation. This initial version of the Advanced Interrupt Architecture is focused primarily on the needs of larger, high-performance RISC-V systems. Support is not currently defined for the following interrupt-handling features that are useful for minimizing interrupt response times in so-called "real-time" systems but are less appropriate for high-speed processor cores: * the option to give each interrupt source at a hart a separate trap entry address; * automatic stacking of register values on interrupt trap entry, and restoration on exit; and * automatic preemption (nesting) of interrupts at a hart, based on priority. It is intended that such features optimizing for smaller and/or real-time systems can be developed as a follow-on extension, either separately or as part of a future version of the interrupt architecture of this document. ### [](#1-1-2-limits)1.1.2\. Limits In its current version, the RISC-V Advanced Interrupt Architecture can support RISC-V symmetric multiprocessing (SMP) systems with up to 16,384 harts. If the harts are 64-bit (RV64) and implement the H extension, and if all features of the Advanced Interrupt Architecture are fully implemented as well, then for each physical hart there may be up to 63 active virtual harts and potentially thousands of additional idle (swapped-out) virtual harts, where each virtual hart has direct control of one or more physical devices. [Table 1](#overalllimits) summarizes the main limits on the numbers of harts, both physical and virtual, and the numbers of distinct interrupt identities that may be supported with the Advanced Interrupt Architecture. | | We assume that any single RISC-V computer (or any single node in a cluster or distributed system) with many thousands of physical harts will probably need an interrupt infrastructure adapted to the machine’s specific organization, which we do not attempt to predict. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 1\. Absolute limits on the numbers of harts and interrupt identities in a system. Individual implementations are likely to have smaller limits.__ | Maximum | Requirements | | | ------------------------------------------------------------------------------------- | ------------------------ | ------------------------------------------------------------------- | | Physical harts | 16,384 | | | Active virtual harts having direct control of a device, per physical hart | 31 for RV32, 63 for RV64 | RISC-V H extension; IMSICs with guest interrupt files; and an IOMMU | | Idle (swapped-out) virtual harts having direct control of a device, per physical hart | potentially thousands | An IOMMU with support for memory-resident interrupt files | | Wired interrupts at a single APLIC | 1023 | | | Distinct identities usable for MSIs at each hart (physical or virtual) | 2047 | IMSICs | ### [](#1-1-3-overview-of-main-components)1.1.3\. Overview of main components A RISC-V system’s overall architecture for signaling interrupts depends on whether it is built mainly for message-signaled interrupts (MSIs) or for more traditional wired interrupts. In systems with full support for MSIs, every hart has an _Incoming MSI Controller_ (IMSIC) that serves as the hart’s own private interrupt controller for external interrupts. Conversely, in systems based primarily on traditional wired interrupts, harts do not have IMSICs. Larger systems, and especially those with PCI devices, are expected to fully support MSIs by giving harts IMSICs, whereas many smaller systems may continue to be best served with wired interrupts and simpler harts without IMSICs. #### [](#1-1-3-1-external-interrupts-without-imsics)1.1.3.1\. External interrupts without IMSICs When RISC-V harts do not have Incoming MSI Controllers, external interrupts are signaled to harts through dedicated wires. In that case, an _Advanced Platform-Level Interrupt Controller_ (APLIC) acts as a traditional central hub for interrupts, routing and prioritizing external interrupts for each hart as illustrated in [Figure 1](#intrsWithoutIMSICs). Interrupts may be selectively routed either to machine level or to supervisor level at each hart. The APLIC is specified in[Advanced Platform-Level Interrupt Controller (APLIC)](AdvPLIC.html#AdvPLIC). Without IMSICs, the current Advanced Interrupt Architecture does not support the direct signaling of external interrupts to virtual machines, even when RISC-V harts implement the H extension. Instead, an interrupt must be sent to the relevant hypervisor, which can then choose to inject a virtual interrupt into the virtual machine. ![intrsWithoutIMSICs](_images/intrsWithoutIMSICs.png) Figure 1\. Traditional delivery of wired interrupts to harts without support for MSIs. ![intrsWithIMSICs](_images/intrsWithIMSICs.png) Figure 2\. Interrupt delivery by MSIs when harts have IMSICs for receiving them. #### [](#1-1-3-2-external-interrupts-with-imsics)1.1.3.2\. External interrupts with IMSICs To be able to receive message-signaled interrupts (MSIs), each RISC-V hart must have an Incoming MSI Controller (IMSIC) as shown in [Figure 2](#intrsWithIMSICs). Fundamentally, a message-signaled interrupt is simply a memory write to a specific address that hardware accepts as indicating an interrupt. To that end, every IMSIC is assigned one or more distinct addresses in the machine’s address space, and when a write is made to one of those addresses in the expected format, the receiving IMSIC interprets the write as an external interrupt for the respective hart. Because all IMSICs have unique addresses in the machine’s physical address space, every IMSIC can receive MSI writes from any agent (hart or device) with permission to write to it. IMSICs have separate addresses for MSIs directed to machine and supervisor levels, in part so the ability to signal interrupts at each privilege level can be separately granted or denied by controlling write permissions at the different addresses, and in part to better support virtualizability (pretending that one privilege level is a higher level). MSIs intended for a hart at a specific privilege level are recorded within the IMSIC in an _interrupt file_, which consists mainly of an array of interrupt-pending bits and a matching array of interrupt-enable bits, the latter indicating which individual interrupts the hart is currently prepared to receive. IMSIC units are fully defined in [Incoming MSI Controller (IMSIC)](IMSIC.html#IMSIC). The format of MSIs used by the RISC-V Advanced Interrupt Architecture is described in that chapter, [MSI encoding](IMSIC.html#MSIEncoding). When the harts in a RISC-V system have IMSICs, the system will normally still contain an APLIC, but its role is changed. Instead of signaling interrupts to harts directly by wires as in [Figure 1](#intrsWithoutIMSICs), an APLIC converts incoming wired interrupts into MSI writes that are sent to harts via their IMSIC units. Each MSI is sent to a single target hart according to the APLIC’s configuration set by software. If RISC-V harts implement the H extension, IMSICs may have additional _guest interrupt files_ for delivering interrupts to virtual machines. Besides [Incoming MSI Controller (IMSIC)](IMSIC.html#IMSIC) on the IMSIC, see [Interrupts for Virtual Machines (VS Level)](VSLevel.html#VSLevel) which specifically covers interrupts to virtual machines. If the system also contains an IOMMU to perform address translation of memory accesses made by I/O devices, then MSIs from those same devices may require special handling. This topic is addressed in [IOMMU Support for MSIs to Virtual Machines](IOMMU.html#IOMMU). #### [](#1-1-3-3-other-interrupts)1.1.3.3\. Other interrupts In addition to external interrupts from I/O devices, the RISC-V Privileged Architecture specifies a few other major classes of interrupts for harts. The Privileged Architecture’s timer interrupts remain supported in full, and software interrupts remain at least partly supported, although neither appears in [Figure 1](#intrsWithoutIMSICs)and [Figure 2](#intrsWithIMSICs). For the specifics on software interrupts, refer to [Interprocessor Interrupts (IPIs)](IPIs.html#IPIs). The Advanced Interrupt Architecture adds considerable support for _local interrupts_ at a hart, whereby a hart essentially interrupts itself in response to asynchronous events, usually errors. Local interrupts remain contained within a hart (or close to it), so like standard RISC-V timer and software interrupts, they do not pass through an APLIC or IMSIC. ### [](#1-1-4-interrupt-identities-at-a-hart)1.1.4\. Interrupt identities at a hart The RISC-V Privileged Architecture gives every interrupt cause at a hart a distinct _major identity number_, which is the Exception Code automatically written to CSR `mcause` or `scause` on an interrupt trap. Interrupt causes that are standardized by the base Privileged Architecture have major identities in the range 0-15, while numbers 16 and higher are officially available for platform standards or for custom use. The Advanced Interrupt Architecture claims further authority over identity numbers in the ranges 16-23 and 32-47, leaving numbers in the range 24-31 and all major identities 48 and higher still free for custom use.[Table 2](#interruptIdents) characterizes all major interrupt identities with this extension. __Table 2\. Major and minor identities for all interrupt causes at a hart. Major identities 0-15 are the purview of the base Privileged Architecture.__ | Major identity | Minor identity | | | ------------------- | ----------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | \- | _Reserved by base Privileged Architecture_ | | 123 | \-\-\- | Supervisor software interruptVirtual supervisor software interruptMachine software interrupt | | 4 | \- | _Reserved by base Privileged Architecture_ | | 567 | \-\-\- | Supervisor timer interruptVirtual supervisor timer interruptMachine timer interrupt | | 8 | \- | _Reserved by base Privileged Architecture_ | | 91011 | Determined byexternal interruptcontroller | Supervisor external interruptVirtual supervisor external interruptMachine external interrupt | | 121314-15 | \-\-\- | Supervisor guest external interruptCounter overflow interrupt _Reserved by base Privileged Architecture_ | | 16-23 | \- | _Reserved for standard local interrupts_ | | 24-31 | \- | _Designated for custom use_ | | 32-343536-424344-47 | \-\-\-\-\- | _Reserved for standard local interrupts_Low-priority RAS event interrupt _Reserved for standard local interrupts_High-priority RAS event interrupt _Reserved for standard local interrupts_ | | ≥48 | \- | _Designated for custom use_ | Interrupts from most I/O devices are conveyed to a hart by the _external interrupt controller_ for the hart, which is either the hart’s IMSIC ([Figure 2](#intrsWithIMSICs)) or an APLIC ([Figure 1](#intrsWithoutIMSICs)). As[Table 2](#interruptIdents) shows, external interrupts at a given privilege level all share a single major identity number: 11 for machine level, 9 for supervisor level, and 10 for VS-level. External interrupts from different causes are distinguished from one another at a hart by their _minor identity numbers_ supplied by the external interrupt controller. Other interrupt causes besides external interrupts might also have their own minor identities. However, this document has need to discuss minor identities only with regard to external interrupts. The local interrupts defined by the Advanced Interrupt Architecture and their handling are covered mainly in [Interrupts for Machine and Supervisor Levels](MSLevel.html#MSLevel), "Interrupts for Machine and Supervisor Levels." ### [](#1-1-5-selection-of-harts-to-receive-an-interrupt)1.1.5\. Selection of harts to receive an interrupt Each signaled interrupt is delivered to only one hart at one privilege level, usually determined by software in one way or another. Unlike some other architectures, the RISC-V Advanced Interrupt Architecture provides no standard hardware mechanism for the broadcast or multicast of interrupts to multiple harts. For local interrupts, and for any "virtual" interrupts that software injects into less-privileged levels at a hart, the interrupts are entirely a local affair at the hart and are never visible to other harts. The RISC-V Privileged Architecture’s timer interrupts are also uniquely tied to individual harts. For other interrupts, received by a hart from sources outside the hart, each interrupt signal (whether delivered by wire or by an MSI) is configured by software to go to only a single hart. To send an interprocessor interrupt (IPI) to multiple harts, the originating hart need only execute a loop, sending an individual IPI to each destination hart. For IPIs to a single destination hart, see[Interprocessor Interrupts (IPIs)](IPIs.html#IPIs). | | The effort that a source hart expends in sending individual IPIs to multiple destinations will invariably be dwarfed by the combined effort at the receiving harts to handle those interrupts. Hence, providing an automated mechanism for IPI multicast could be expected to reduce a system’s total overall work only modestly at best. With a very large number of harts, a hardware mechanism for IPI multicast must contend with the question of how exactly software specifies the intended destination set with each use, and furthermore, the actual physical delivery of IPIs may not differ that much from the software version. We do not exclude the future possibility of an optional hardware mechanism for multicast IPI, but only if a significant advantage can be demonstrated in real use. As of 2020, Linux has been observed not to make use of multicast IPI hardware even on systems that have it. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | In the rare event that a single interrupt from an I/O device needs to be communicated to multiple harts, the interrupt must be sent to a single hart which can then signal the other harts by IPIs. | | We contend that the need to communicate an I/O interrupt to multiple harts is sufficiently rare that standardizing hardware support for multicast cannot be justified in this case. Along with multicast delivery, other architectures support an option for "1-of- " delivery of interrupts, whereby the hardware chooses a single destination hart from among a configured set of harts, with the goal of automatic load balancing of interrupt handling among the harts. Experiments in the 2010s called into question the utility of 1-of- modes in practice, showing that software could often do a better job of load balancing than the hardware algorithms implemented in actual chips. Linux was consequently modified to discontinue using 1-of- interrupt delivery even on systems that have it. We remain open to the argument that hardware load balancing of interrupt handling may be beneficial for certain specialized markets, such as networking. However, the claims made so far in this regard do not justify requiring support for 1-of- delivery in all RISC-V servers. With more evidence, some mechanism for 1-of- delivery might become a future option. The original Platform-Level Interrupt Controller (PLIC) for RISC-V is configurable so each interrupt source signals external interrupts to any subset of the harts, potentially all harts. When multiple harts receive an external interrupt from a single cause at the PLIC, the first hart to_claim_ the interrupt at the PLIC is the one responsible for servicing it. Usually this sets up a race, where the subset of harts configured to receive the multicast interrupt all take an external interrupt trap simultaneously and compete to be the first to claim the interrupt at the PLIC. The intention is to provide a form of 1-of- interrupt delivery. However, for all the harts that fail to win the claim, the interrupt trap becomes wasted effort. For the reasons already given, the Advanced PLIC supports sending each signaled interrupt to only a single hart chosen by software, not to multiple harts. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#1-1-6-isa-extensions-smaia-and-ssaia)1.1.6\. ISA extensions Smaia and Ssaia The Advanced Interrupt Architecture (AIA) defines two names for extensions to the RISC-V instruction set architecture (ISA), one for machine-level execution environments, and another for supervisor-level environments. For a machine-level environment, extension **Smaia**encompasses all added CSRs and all modifications to interrupt response behavior that the AIA specifies for a hart, over all privilege levels. For a supervisor-level environment, extension **Ssaia** is essentially the same as Smaia except excluding the machine-level CSRs and behavior not directly visible to supervisor level. Extensions Smaia and Ssaia cover only those AIA features that impact the ISA at a hart. Although the following are described or discussed in this document as part of the AIA, they are not implied by Smaia or Ssaia because the components are categorized as non-ISA: APLICs, IOMMUs, and any mechanisms for initiating interprocessor interrupts apart from writing to IMSICs. As revealed in subsequent chapters, the exact set of CSRs and behavior added by the AIA, and hence implied by Smaia or Ssaia, depends on the base ISA’s XLEN (RV32 or RV64), on whether S-mode and the H extension are implemented, and on whether the hart has an IMSIC. But individual AIA extension names are not provided for each possible valid subset. Rather, the different combinations are inferable from the intersection of features indicated (such as RV64I + S-mode + Smaia, but without the H extension). Software development tools like compilers and assemblers need not be concerned about whether an IMSIC exists but should just allow attempts to access the IMSIC CSRs (described in [Control and Status Registers (CSRs) Added to Harts](CSRs.html#CSRs)and [Incoming MSI Controller (IMSIC)](IMSIC.html#IMSIC)) if Smaia or Ssaia is indicated. Without an actual IMSIC, such attempts may trap, but that is not a problem for the development tools. If extension Smaia/Ssaia is implemented, then anywhere that the AIA specification has an irreconcilable conflict with the requirements of another implemented RISC-V extension, the AIA is intended to have priority, unless the other extension explicitly extends or overrides the AIA. | | Extension Smcsrind/Sscsrind explicitly extends the AIA’s facility for indirect CSR access provided by the \*iselect and \*ireg CSRs described in the next chapter. Hence, if Smcsrind/Sscsrind is also implemented, any perceived conflicts between it and the AIA should be resolved in favor of Smcsrind/Sscsrind. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Bibliography ==================== ## [](#bibliography)Bibliography \[1\] _Clarification updates to IOMMU v20250828_. \[Online\]. Available: \[2\] _Clarification updates to IOMMU v20250625_. \[Online\]. Available: \[3\] _Clarification updates to IOMMU v1.0.1_. \[Online\]. Available: \[4\] _Clarification updates to IOMMU v1.0.0_. \[Online\]. Available: \[5\] _PCI Express® Base Specification Revision 6.0_, . \[Online\]. Available: \[6\] _RISC-V Advanced Interrupt Architecture_. \[Online\]. Available: \[7\] _RISC-V Instruction Set Manual, Volume II: Privileged Architecture_, . \[Online\]. Available: \[8\] _RISC-V Shadow Stacks and Landing Pads_. \[Online\]. Available: \[9\] _PCI Code and ID Assignment Specification Revision 1.1_, . \[Online\]. Available: \[10\] D. B. Kristof and E. Stijn and E. Lieven, "Per-Thread Cycle Accounting in Multicore Processors", _ACM Trans. Archit. Code Optim._, vol. 9, no. 4, jan 2013\. \[Online\]. Available: . \[11\] L. David and C. Liqun and G. Rama and R. Parthasarathy and K. Christos, "Heracles: Improving Resource Efficiency at Scale" in _Proceedings of the 42nd Annual International Symposium on Computer Architecture_, ISCA '15\. New York, NY, USA:, Association for Computing Machinery, 2015, pp. 450–462, Available: . \[12\] _RISC-V Quality-of-Service (QoS) Identifiers_. \[Online\]. Available: \[13\] _RISC-V Capacity and Bandwidth QoS Register Interface_. \[Online\]. Available: Contributors ==================== ## [](#contributors)Contributors This RISC-V specification has been contributed to directly or indirectly by (in alphabetical order): Aaron Durbin, Allen Baum, Anup Patel, Daniel Gracia Pérez, David Kruckemyer, Greg Favor, Ahmad Fawal, Guerney D Hunt, John Hauser, Josh Scheid, Matt Evans, Manuel Rodriguez, Nick Kossifidis, Paul Donahue, Paul Walmsley, Perrine Peresse, Philipp Tomsich, Rieul Ducousso, Scott Nelson, Siqi Zhao, Sunil V.L, Tomasz Jeznach, Vassilis Papaefstathiou, Vedvyas Shanbhogue Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at. Copyright 2023 by RISC-V International. RISC-V IOMMU Architecture Specification ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-iommu-architecture-specification)RISC-V IOMMU Architecture Specification Version v1.0.1, Revised 2026-02-22 | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Data Structures ==================== ## [](#DATA%5FSTRUCTURES)Data Structures A data structure called device-context (`DC`) is used by the IOMMU to associate a device with an address space and to hold other per-device parameters used by the IOMMU to perform address translations. A radix-tree data structure called device directory table (DDT) that is traversed using the `device_id` is used to locate the `DC`. The address space used by a device may require second-stage address translation and protection when the control of the device is passed through to a Guest OS. A Guest OS may optionally provide a first-stage page table for translating IOVA used by a device controlled by the Guest OS to a GPA. When the use of a first-stage is not required, then it may be effectively disabled by selecting the first-stage address translation scheme to be `Bare`. The second-stage is used to translate the GPA to a SPA. When the control of the device is retained by the hypervisor or Host OS itself then only the first-stage suffices to perform necessary address translations and protections; the second-stage scheme may be effectively disabled for the device by programming the second-stage address translation scheme to be `Bare`. When second-stage address translation is not Bare, the `DC` holds the PPN of the root second-stage page table; a guest-soft-context-ID (`GSCID`), which facilitates invalidation of cached address translations on a per-virtual-machine basis; and the second-stage address translation scheme. Some devices support multiple process contexts where each context may be associated with a different process and thus a different virtual address space. The context in such devices may be configured with a `process_id` that identifies the address space. When making a memory access, such devices signal the `process_id` along with the `device_id` to identify the accessed address space. An example of such a device may be a GPU that supports multiple process contexts, where each context is associated with a different user process, such that the GPU may access memory using the virtual address provided by the user process itself. To support selecting an address space associated with the`process_id`, the `DC` holds the PPN of the root Process Directory Table (PDT), a radix-tree data structure, indexed using fields of the `process_id` to locate a data structure called the Process Context (`PC`). When a PDT is active, the controls for first-stage address translation are held in the (`PC`). When a PDT is not active, the controls for first-stage address translation are held in the `DC` itself. The first-stage address translation controls include the PPN of the root first-stage page table; a process-soft-context-ID (`PSCID`), which facilitates invalidation of cached address translations on a per-address-space basis; and the first-stage address translation scheme. To handle MSIs from a device controlled by a guest OS, an IOMMU must be able to redirect those MSIs to a guest interrupt file in an IMSIC. Because MSIs from devices are simply memory writes, they would naturally be subject to the same address translation that an IOMMU applies to other memory writes. However, the IOMMU architecture may treat MSIs directed to virtual machines specially, in part to simplify software, and in part to allow optional support for memory-resident interrupt files. To support this capability, the architecture adds to the device contexts an MSI address mask and address pattern, used together to identify pages in the guest physical address space that are the destinations of MSIs; and the real physical address of an MSI page table for controlling the translation and/or conversion of MSIs from the device. The IOMMU support for MSIs to virtual machines is specified by the Advanced Interrupt Architecture specification. The `DC` further holds controls for the type of transactions that a device is allowed to generate. One example of such a control is whether the device is allowed to use the PCIe defined Address Translation Service (ATS) \[[5](bibliography.html#bib-pci)\]. Two formats of the device-context structure are supported: * **Base Format** \- is 32-bytes in size used when the special treatment of MSI as specified in [Process to translate addresses of MSIs](#MSI%5FTRANS) is not supported by the IOMMU. * **Extended Format** \- is 64-bytes in size and extends the base format `DC` with additional fields to translate MSIs as specified in [Process to translate addresses of MSIs](#MSI%5FTRANS). If `capabilities.MSI_FLAT` is 1 then the Extended Format is used else the Base Format is used. The DDT used to locate the `DC` may be configured to be a 1, 2, or 3 level radix-tree depending on the maximum width of the `device_id` supported. The partitioning of the `device_id` to obtain the device directory indexes (DDI) to traverse the DDT radix-tree are as follows: ![Base format `device_id` partitioning](_images/diag-9485f1554155a09a7c9de691afc887761c5dd07a.svg) Figure 1\. Base format `device_id` partitioning ![Extended format `device_id` partitioning](_images/diag-4263d256154f91f234913e4be7f9491bb2f7c778.svg) Figure 2\. Extended format `device_id` partitioning The PDT may be configured to be a 1, 2, or 3 level radix-tree depending on the maximum width of the `process_id` supported by that device. The partitioning of the `process_id` to obtain the process directory indices (PDI) to traverse the PDT radix-tree are as follows: ![`process_id` partitioning for PDT radix-tree traversal](_images/diag-4c696a04599cd901ef47be97385abc5206217a06.svg) Figure 3\. `process_id` partitioning for PDT radix-tree traversal | | The process\_id partitioning is designed to require a maximum of 4 KiB, a page, of memory for each process directory table. The root of the table when using a 20-bit wide process\_id is not fully populated. The option of making the root table occupy 32 KiB was considered but not adopted as these tables are allocated at run time and contiguous memory allocation larger than a page may stress the Guest and hypervisor memory allocators. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | All RISC-V IOMMU implementations are required to support DDT and PDT located in main memory. Supporting data structures in I/O memory is not required but is not prohibited by this specification. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#device-directory-table-ddt)Device-Directory-Table (DDT) The DDT is a 1, 2, or 3-level radix-tree indexed using the device directory index (DDI) bits of the `device_id` to locate a `DC`. The following diagrams illustrate the DDT radix-tree. The PPN of the root device-directory-table is held in a memory-mapped register called the device-directory-table pointer (`ddtp`). Each valid non-leaf (`NL`) entry is 8-bytes in size and holds the PPN of the next device-directory-table. A valid leaf device-directory-table entry holds the device-context (`DC`). ![ddt ext](_images/ddt-ext.svg) Figure 4\. Three, two and single-level device directory with extended format `DC` ![ddt base](_images/ddt-base.svg) Figure 5\. Three, two and single-level device directory with base format `DC` #### [](#non-leaf-ddt-entry)Non-leaf DDT entry A valid (`V==1`) non-leaf DDT entry provides the PPN of the next level DDT. ![Non-leaf device-directory-table entry](_images/diag-6be7974d0458dc0d2a716548507460afd44de2ba.svg) Figure 6\. Non-leaf device-directory-table entry #### [](#leaf-ddt-entry)Leaf DDT entry The leaf DDT page is indexed by `DDI[0]` and holds the device-context (`DC`). In base-format the `DC` is 32-bytes. In extended-format the `DC` is 64-bytes. ![Base-format device-context](_images/diag-f9692f54d72643c7306777389a79d8ad8e66d695.svg) Figure 7\. Base-format device-context ![Extended-format device-context](_images/diag-633af555e15b468ee38888f00fb39c1f2be5ad18.svg) Figure 8\. Extended-format device-context The `DC` is interpreted as four 64-bit doublewords in base-format and as eight 64-bit doublewords in extended-format. The byte order of each of the doublewords in memory, little-endian or big-endian, is the endianness as determined by `fctl.BE` ([iommu\_registers.adoc#FCTRL](iommu%5Fregisters.html#FCTRL)). The IOMMU may read the `DC` fields in any order. #### [](#device-context-fields)Device-context fields ##### [](#translation-control-tc)Translation control (`tc`) ![Translation control (`tc`) field](_images/diag-d786cb86384281b3fe97fb7ca7672d7fe38d8828.svg) Figure 9\. Translation control (`tc`) field `DC` is valid if the `V` bit is 1; If it is 0, all other bits in `DC` are don’t-care and may be freely used by software. If the IOMMU supports PCIe ATS specification \[[5](bibliography.html#bib-pci)\] (see `capabilities`register), the `EN_ATS` bit is used to enable ATS transaction processing. If`EN_ATS` is set to 1, IOMMU supports the following inbound transactions; otherwise they are treated as unsupported requests. * Translated read for execute transaction * Translated read transaction * Translated write/AMO transaction * PCIe ATS Translation Request * PCIe ATS Invalidation Completion Message If the `EN_ATS` bit is 1 and the `T2GPA` bit is set to 1 the IOMMU performs the two-stage address translation to determine the permissions and the size of the translation to be provided in the completion of a PCIe ATS Translation Request from the device. However, the IOMMU returns a GPA, instead of a SPA, as the translation of an IOVA in the response. In this mode of operation, the ATC in the device caches a GPA as a translation for an IOVA and uses the GPA as the address in subsequent translated memory access transactions. Usually, translated requests use a SPA and need no further translation to be performed by the IOMMU. However when `T2GPA` is 1, translated requests from a device use a GPA and are translated by the IOMMU using the second-stage page table to a SPA. The `T2GPA`control enables a hypervisor to contain DMA from a device, even if the device misuses the ATS capability and attempts to access memory that is not associated with the VM. | | When T2GPA is enabled, the addresses provided to the device in response to a PCIe ATS Translation Request cannot be directly routed by the I/O fabric (e.g. PCI switches) that connect the device to other peer devices and to host. Such addresses also cannot be routed within the device when peer-to-peer transactions within the device (e.g. between functions of a device) are supported. Use of T2GPA set to 1 may not be compatible with devices that implement caches tagged by the translated address returned in response to a PCIe ATS Translation Request. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Hypervisors that configure T2GPA to 1 must ensure through protocol-specific means that translated accesses are routed through the host such that the IOMMU may translate the GPA and then route the transaction based on PA to memory or to a peer device. For PCIe, for example, the Access Control Service (ACS) must be configured to always redirect peer-to-peer (P2P) requests upstream to the host. As an alternative to setting T2GPA to 1, the hypervisor may establish a trust relationship with the device if authentication protocols are supported by the device. For PCIe, for example, the PCIe component measurement and authentication (CMA) capability provides a mechanism to verify the device’s configuration and firmware/executable (Measurement) and hardware identities (Authentication) to establish such a trust relationship. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | If `EN_PRI` bit is 0, then PCIe "Page Request" messages from the device are invalid requests. A "Page Request" message received from a device is responded to with a "Page Request Group Response" message. Normally, a software handler generates this response message. However, under some conditions the IOMMU itself may generate a response. For IOMMU-generated "Page Request Group Response" messages the PRG-response-PASID-required (`PRPR`) bit when set to 1 indicates that the IOMMU response message should include a PASID if the associated "Page Request" had a PASID. | | Functions that support PASID and have the "PRG Response PASID Required" capability bit set to 1, expect that "Page Request Group Response" messages will contain a PASID if the associated "Page Request" message had a PASID. If the capability bit is 0, the function does not expect PASID on any "Page Request Group Response" message and the behavior of the function if it receives the response with a PASID is undefined. The PRPR bit should be configured with the value held in the "PRG Response PASID Required" capability bit. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | Setting the disable-translation-fault (`DTF`) bit to 1 disables reporting of faults encountered in the address translation process. Setting `DTF` to 1 does not disable error responses from being generated to the device in response to faulting transactions. Setting `DTF` to 1 does not disable reporting of faults from the IOMMU that are not related to the address translation process. The faults that are not reported when `DTF` is 1 are listed in [iommu\_in\_memory\_queues.adoc#FAULT\_CAUSE](iommu%5Fin%5Fmemory%5Fqueues.html#FAULT%5FCAUSE). | | A hypervisor may set DTF to 1 to disable fault reporting when it has identified conditions that may lead to a flurry of errors such as due to an abnormal termination of a virtual machine. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `DC.fsc` field holds the context for first-stage translation. If the`PDTV` bit is 1, the field holds the process-directory table pointer (`pdtp`). If the `PDTV` bit is 0, the `DC.fsc` field holds (`iosatp`). The `PDTV` bit is expected to be set to 1 when `DC` is associated with a device that supports multiple process contexts and thus generates a valid `process_id`with its memory accesses. For PCIe, for example, if the request has a PASID then the PASID is used as the `process_id`. When `PDTV` is 1, the `DPE` bit may set to 1 to enable the use of 0 as the default value of `process_id` for translating requests without a valid`process_id`. When `PDTV` is 0, the `DPE` bit is reserved for future standard extension. The IOMMU supports the 1 setting of `GADE` and `SADE` bits if`capabilities.AMO_HWAD` is 1\. When `capabilities.AMO_HWAD` is 0, these bits are reserved. If `GADE` is 1, the IOMMU updates A and D bits in second-stage PTEs atomically. If `GADE` is 0, the IOMMU causes a guest-page-fault corresponding to the original access type if the A bit is 0 or if the memory access is a store and the D bit is 0. If `SADE` is 1, the IOMMU updates A and D bits in first-stage PTEs atomically. If`SADE` is 0, the IOMMU causes a page-fault corresponding to the original access type if the A bit is 0 or if the memory access is a store and the D bit is 0. If `SBE` is 0, implicit memory accesses to PDT entries and first-stage PTEs are little-endian else they are big-endian. The supported values of `SBE` are the same as that of the `fctl.BE` field. The `SXL` field controls the supported paged virtual-memory schemes as defined in [Table 3](#IOSATP%5FMODE%5FENC-0) and [Table 4](#IOSATP%5FMODE%5FENC-1). If `fctl.GXL` is 1 then the`SXL` field must be 1; otherwise the legal values for the `SXL` field are the same as those for the `fctl.GXL` field. When `SXL` is 1, the following rules apply: * If the first-stage is not Bare, then a page fault corresponding to the original access type occurs if the `IOVA` has bits beyond bit 31 set to 1. * If the second-stage is not Bare, then a guest page fault corresponding to the original access type occurs if the incoming GPA has bits beyond bit 33 set to 1. ##### [](#io-hypervisor-guest-address-translation-and-protection-iohgatp)IO hypervisor guest address translation and protection (`iohgatp`) ![IO hypervisor guest address translation and protection (`iohgatp`) field](_images/diag-6d7c372d1fb41a60f328cfe61fa707206b05d14d.svg) Figure 10\. IO hypervisor guest address translation and protection (`iohgatp`) field The `iohgatp` field holds the PPN of the root second-stage page table and a virtual machine identified by a guest soft-context ID (`GSCID`), to facilitate address-translation fences on a per-virtual-machine basis. If multiple devices are associated to a VM with a common second-stage page table, the hypervisor is expected to program the same `GSCID` in each `iohgatp`. The `MODE` field is used to select the second-stage address translation scheme. The second-stage page table formats are as defined by the Privileged specification. The `fctl.GXL` field controls the supported address-translation schemes for guest physical addresses as defined in [Table 1](#IOHGATP%5FMODE%5FENC-0) and[Table 2](#IOHGATP%5FMODE%5FENC-1). The `iohgatp` `MODE` field identifies the paged virtual-memory schemes and its encodings are as follows: __Table 1\. Encodings of iohgatp.MODE field when fctl.GXL=0__ | Value | Name | Description | | ----- | ------ | --------------------------------------------------------------- | | 0 | Bare | No translation or protection. | | 1-7 | — | Reserved for standard use. | | 8 | Sv39x4 | Page-based 41-bit virtual addressing (2-bit extension of Sv39). | | 9 | Sv48x4 | Page-based 50-bit virtual addressing (2-bit extension of Sv48). | | 10 | Sv57x4 | Page-based 59-bit virtual addressing (2-bit extension of Sv57). | | 11-15 | — | Reserved for standard use. | __Table 2\. Encodings of iohgatp.MODE field when fctl.GXL=1__ | Value | Name | Description | | ----- | ------ | --------------------------------------------------------------- | | 0 | Bare | No translation or protection. | | 1-7 | — | Reserved for standard use. | | 8 | Sv32x4 | Page-based 34-bit virtual addressing (2-bit extension of Sv32). | | 9-15 | — | Reserved for standard use. | Implementations are not required to support all defined mode settings for`iohgatp`. The IOMMU only needs to support the modes also supported by the MMU in the harts integrated into the system or a subset thereof. The root page table as determined by `iohgatp.PPN` is 16 KiB and must be aligned to a 16-KiB boundary. | | The GSCID field of iohgatp identifies an address space. If an identicalGSCID is configured in two DC when the second-stage page-table referenced by the two DC are not identical then it is unpredictable whether the IOMMU uses the PTEs from the first page table or the second page table. These are the only expected behaviors. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#DC%5FTA)Translation attributes (`ta`) ![Translation attributes (`ta`) field](_images/diag-880135da2d406ecb2acf0b90f093c2a29c25cf1f.svg) Figure 11\. Translation attributes (`ta`) field The `PSCID` field of `ta` provides the process soft-context ID that identifies the address-space of the process. `PSCID` facilitates address-translation fences on a per-address-space basis. The `PSCID` field in `ta` is used as the address-space ID if `DC.tc.PDTV` is 0 and the `iosatp.MODE` field is not `Bare`. When `DC.tc.PDTV` is 1, the `PSCID` field in `ta` is ignored. The `RCID` and `MCID` fields are added by the QoS ID extension. If`capabilities.QOSID` is 0, these bits are reserved and must be set to 0\. IOMMU-initiated requests for accessing the following data structures use the value configured in the `RCID` and `MCID` fields of `DC.ta`. * Process directory table (`PDT`) * Second-stage page table * First-stage page table * MSI page table * Memory-resident interrupt file (`MRIF`) The `RCID` and `MCID` configured in `DC.ta` are provided to the IO bridge on successful address translations. The IO bridge should associate these QoS IDs with device-initiated requests. ##### [](#first-stage-context-fsc)First-Stage context (`fsc`) If `DC.tc.PDTV` is 0, the `DC.fsc` field holds the `iosatp` that provides the controls for first-stage address translation and protection. ![IO Supervisor address translation and prot. (`iosatp`) field](_images/diag-4839abeede695faecc082029624b887ff199b79a.svg) Figure 12\. IO Supervisor address translation and prot. (`iosatp`) field The first-stage page table formats are as defined by the Privileged specification. The `DC.tc.SXL` field controls the supported paged virtual-memory schemes. The `iosatp.MODE` identifies the paged virtual-memory schemes and is encoded as defined in [Table 3](#IOSATP%5FMODE%5FENC-0) and [Table 4](#IOSATP%5FMODE%5FENC-1). The `iosatp.PPN`field holds the PPN of the root page of a first-stage page table. When second-stage address translation is not `Bare`, the `iosatp.PPN` is a guest PPN. The GPA of the root page is then converted by guest physical address translation process, as controlled by the `iohgatp`, into a supervisor physical address. __Table 3\. Encodings of iosatp.MODE field when DC.tc.SXL=0__ | Value | Name | Description | | ----- | ---- | ------------------------------------- | | 0 | Bare | No translation or protection. | | 1-7 | — | Reserved for standard use. | | 8 | Sv39 | Page-based 39-bit virtual addressing. | | 9 | Sv48 | Page-based 48-bit virtual addressing. | | 10 | Sv57 | Page-based 57-bit virtual addressing. | | 11-13 | — | Reserved for standard use. | | 14-15 | — | Designated for custom use. | __Table 4\. Encodings of iosatp.MODE field when DC.tc.SXL=1__ | Value | Name | Description | | ----- | ---- | ------------------------------------- | | 0 | Bare | No translation or protection. | | 1-7 | — | Reserved for standard use. | | 8 | Sv32 | Page-based 32-bit virtual addressing. | | 9-15 | — | Reserved for standard use. | When `DC.tc.PDTV` is 1, the `DC.fsc` field holds the process-directory table pointer (`pdtp`). When the device supports multiple process contexts, selected by the `process_id`, the PDT is used to determine the first-stage page table and associated `PSCID` for virtual address translation and protection. The `pdtp` field holds the PPN of the root PDT and the `MODE` field that determines the number of levels of the PDT. ![Process-directory table pointer (`pdtp`) field](_images/diag-a25509c120a23fbda2a8aefe812e0f6ffd0215db.svg) Figure 13\. Process-directory table pointer (`pdtp`) field When second-stage address translation is not Bare, the `pdtp.PPN` field holds a guest PPN. The GPA of the root PDT is then converted by guest physical address translation process, as controlled by the `iohgatp`, into a supervisor physical address. Translating addresses of PDT using a second-stage page table, allows the PDT to be held in memory allocated by the guest OS and allows the guest OS to directly edit the PDT to associate a virtual-address space identified by a first-stage page table with a `process_id`. __Table 5\. Encodings of pdtp.MODE field__ | Value | Name | Description | | ----- | ---- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | Bare | No first-stage address translation or protection. | | 1 | PD8 | 8-bit process ID enabled. The directory has 1 levels with 256 entries.The bits 19:8 of process\_id must be 0. | | 2 | PD17 | 17-bit process ID enabled. The directory has 2 levels. The root PDT page has 512 entries and leaf level has 256 entries. The bits 19:17 of process\_id must be 0. | | 3 | PD20 | 20-bit process ID enabled. The directory has 3 levels. The root PDT has 8 entries and the next non-leaf level has 512 entries. The leaf level has 256 entries. | | 4-13 | — | Reserved for standard use. | | 14-15 | — | Designated for custom use. | ##### [](#msi-page-table-pointer-msiptp)MSI page table pointer (`msiptp`) ![MSI page table pointer (`msiptp`) field](_images/diag-a25509c120a23fbda2a8aefe812e0f6ffd0215db.svg) Figure 14\. MSI page table pointer (`msiptp`) field The `msiptp.PPN` field holds the PPN of the root MSI page table used to direct an MSI to a guest interrupt file in an IMSIC. The MSI page table formats are defined by the Advanced Interrupt Architecture specification. The `msiptp.MODE` field is used to select the MSI address translation scheme. __Table 6\. Encodings of msiptp.MODE field__ | Value | Name | Description | | ----- | ---- | -------------------------------------------------------------------------------------------------------- | | 0 | Off | Recognition of accesses to a virtual interrupt file using MSI address mask and pattern is not performed. | | 1 | Flat | Flat MSI page table | | 2-13 | — | Reserved for standard use. | | 14-15 | — | Designated for custom use. | When `DC.iohgatp.MODE` is `Bare`, the `msiptp.MODE` must be set to `Off`. ##### [](#MSI%5FID)MSI address mask (`msi_addr_mask`) and pattern (`msi_addr_pattern`) ![MSI address mask (`msi_addr_mask`) field](_images/diag-97cb2ee3c09f4011236452190129741366bd2bbb.svg) Figure 15\. MSI address mask (`msi_addr_mask`) field ![MSI address pattern (`msi_addr_pattern`) field](_images/diag-709203d36b2e0165319f6b371306252f1183958b.svg) Figure 16\. MSI address pattern (`msi_addr_pattern`) field The MSI address mask (`msi_addr_mask`) and pattern (`msi_addr_pattern`) fields are used to identify the 4-KiB pages of virtual interrupt files in the guest physical address space of the relevant VM. An incoming memory access made by a device is recognized as an access to a virtual interrupt file if the destination guest physical page matches the supplied address pattern in all bit positions that are zeros in the supplied address mask. In detail, a memory access to guest physical address `A` is recognized as an access to a virtual interrupt file’s memory-mapped page if: `(A >> 12) & ~msi_addr_mask = (msi_addr_pattern & ~msi_addr_mask)` where >> 12 represents shifting right by 12 bits, an ampersand (&) represents bitwise logical AND, and `~msi_addr_mask` is the bitwise logical complement of the address mask. While the MSI address mask and pattern fields are 52 bits wide, if , then bits are reserved for future standard use and must be set to zero by software. MGPAW is determined as follows: * If `capabilities.Sv57x4` is 1, then MGPAW = 59 * Else if `capabilities.Sv48x4` is 1, then MGPAW = 50 * Else if `capabilities.Sv39x4` is 1, then MGPAW = 41 * Else if `capabilities.Sv32x4` is 1, then MGPAW = 34 * Otherwise, MGPAW = `capabilities.PAS` #### [](#DC%5FMISCONFIG)Device-context configuration checks A `DC` with `DC.tc.V=1` is considered as misconfigured if any of the following conditions are true. If misconfigured then, stop and report "DDT entry misconfigured" (cause = 259). 1. If any bits or encodings that are reserved for future standard use are set. 2. `capabilities.ATS` is 0 and `DC.tc.EN_ATS`, or `DC.tc.EN_PRI`, or `DC.tc.PRPR` is 1 3. `DC.tc.EN_ATS` is 0 and `DC.tc.T2GPA` is 1 4. `DC.tc.EN_ATS` is 0 and `DC.tc.EN_PRI` is 1 5. `DC.tc.EN_PRI` is 0 and `DC.tc.PRPR` is 1 6. `capabilities.T2GPA` is 0 and `DC.tc.T2GPA` is 1 7. `DC.tc.T2GPA` is 1 and `DC.iohgatp.MODE` is `Bare` 8. `DC.tc.PDTV` is 1 and `DC.fsc.pdtp.MODE` is not a supported mode 1. `capabilities.PD20` is 0 and `DC.fsc.pdtp.MODE` is `PD20` 2. `capabilities.PD17` is 0 and `DC.fsc.pdtp.MODE` is `PD17` 3. `capabilities.PD8` is 0 and `DC.fsc.pdtp.MODE` is `PD8` 9. `DC.tc.PDTV` is 0 and `DC.fsc.iosatp.MODE` encoding is not a valid encoding as determined by [Table 3](#IOSATP%5FMODE%5FENC-0) and[Table 4](#IOSATP%5FMODE%5FENC-1). 10. `DC.tc.PDTV` is 0 and `DC.tc.SXL` is 0 `DC.fsc.iosatp.MODE`is not one of the supported modes 1. `capabilities.Sv39` is 0 and `DC.fsc.iosatp.MODE` is `Sv39` 2. `capabilities.Sv48` is 0 and `DC.fsc.iosatp.MODE` is `Sv48` 3. `capabilities.Sv57` is 0 and `DC.fsc.iosatp.MODE` is `Sv57` 11. `DC.tc.PDTV` is 0 and `DC.tc.SXL` is 1 `DC.fsc.iosatp.MODE`is not one of the supported modes 1. `capabilities.Sv32` is 0 and `DC.fsc.iosatp.MODE` is `Sv32` 12. `DC.tc.PDTV` is 0 and `DC.tc.DPE` is 1 13. `DC.iohgatp.MODE` encoding is not a valid encoding as determined by [Table 1](#IOHGATP%5FMODE%5FENC-0) and [Table 2](#IOHGATP%5FMODE%5FENC-1). 14. `fctl.GXL` is 0 and `DC.iohgatp.MODE` is not a supported mode 1. `capabilities.Sv39x4` is 0 and `DC.iohgatp.MODE` is `Sv39x4` 2. `capabilities.Sv48x4` is 0 and `DC.iohgatp.MODE` is `Sv48x4` 3. `capabilities.Sv57x4` is 0 and `DC.iohgatp.MODE` is `Sv57x4` 15. `fctl.GXL` is 1 and `DC.iohgatp.MODE` is not a supported mode 1. `capabilities.Sv32x4` is 0 and `DC.iohgatp.MODE` is `Sv32x4` 16. `capabilities.MSI_FLAT` is 1 and `DC.msiptp.MODE` is not `Off`and not `Flat` 17. `DC.iohgatp.MODE` is not `Bare` and the root page table determined by`DC.iohgatp.PPN` is not aligned to a 16-KiB boundary. 18. `capabilities.AMO_HWAD` is 0 and `DC.tc.SADE` or `DC.tc.GADE` is 1 19. `capabilities.END` is 0 and `fctl.BE != DC.tc.SBE` 20. `DC.tc.SXL` value is not a legal value. If `fctl.GXL` is 1, then`DC.tc.SXL` must be 1\. If `fctl.GXL` is 0 and is writable, then`DC.tc.SXL` may be 0 or 1\. If `fctl.GXL` is 0 and is not writable then `DC.tc.SXL` must be 0. 21. `DC.tc.SBE` value is not a legal value. If `fctl.BE` is writable then `DC.tc.SBE` may be 0 or 1\. If `fctl.BE` is not writable then`DC.tc.SBE` must be the same as `fctl.BE`. 22. `capabilities.QOSID` is 1 and `DC.ta.RCID` or `DC.ta.MCID` values are wider than that supported by the IOMMU. When `DC.iohgatp.MODE` is `Bare`, `DC.msiptp.MODE` must be set to `Off` by software. All other settings are reserved. Implementations are recommended to stop and report "DDT entry misconfigured" (cause = 259) if a reserved setting is detected. | | Some DC fields hold supervisor physical addresses or guest physical addresses. Some implementations may verify the validity of the addresses - e.g. the supervisor physical address is not wider than that supported as determined by capabilities.PAS, etc. at the time of locating theDC. Such implementations may cause a "DDT entry misconfigured" (cause = 259) fault. Other implementations only detect such addresses to be invalid when the data structure referenced by these fields needs to be accessed. Such implementations may detect access-violation faults in the process of making the access. An earlier version of the specification did not recommend implementations to check that msiptp.MODE was set to Off when iohgatp.MODE was Bare. When iohgatp.MODE is Bare, second-stage address translation is effectively disabled and no valid GSCID exists to associate translations from an MSI page table with a VM address space. In such cases, software must set msiptp.MODE to Off. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#process-directory-table-pdt)Process-Directory-Table (PDT) The PDT is a 1, 2, or 3-level radix-tree indexed using the process directory index (`PDI`) bits of the `process_id`. The following diagrams illustrate the PDT radix-tree. The root process-directory page number is located using the process-directory-table pointer (`pdtp`) field of the device-context. Each non-leaf (`NL`) entry provides the PPN of the next level process-directory-table. The leaf process-directory-table entry holds the process-context (`PC`). ![pdt](_images/pdt.svg) Figure 17\. Three, two and single-level process directory #### [](#non-leaf-pdt-entry)Non-leaf PDT entry A valid (`V==1`) non-leaf PDT entry holds the PPN of the next-level PDT. ![Non-leaf process-directory-table entry](_images/diag-7575fedc945c1870ba234a59ef61e17c2d7ba9ba.svg) Figure 18\. Non-leaf process-directory-table entry #### [](#leaf-pdt-entry)Leaf PDT entry The leaf PDT page is indexed by `PDI[0]` and holds the 16-byte process-context (`PC`). ![Process-context](_images/diag-0451fb55cd8546aed2c59899e95bf7aa63c439e8.svg) Figure 19\. Process-context The `PC` is interpreted as two 64-bit doublewords. The byte order of each of the doublewords in memory, little-endian or big-endian, is the endianness as determined by `DC.tc.SBE`. The IOMMU may read the `PC` fields in any order. #### [](#process-context-fields)Process-context fields ##### [](#translation-attributes-ta)Translation attributes (`ta`) ![Translation attributes (`ta`) field](_images/diag-0311d31280abd3cfc335cc8a56a5809e999274dc.svg) Figure 20\. Translation attributes (`ta`) field `PC` is valid if the `V` bit is 1; If it is 0, all other bits in `PC` are don’t care and may be freely used by software. When Enable-Supervisory-access (`ENS`) is 1, transactions requesting supervisor privilege are allowed with this `process_id` else the transaction is treated as an unsupported request. When `ENS` is 1, the `SUM` (permit Supervisor User Memory access) bit modifies the privilege with which supervisor privilege transactions access virtual memory. When `SUM` is 0, supervisor privilege transactions to pages mapped with`U` bit in PTE set to 1 are disallowed. When `ENS` is 1, supervisor privilege transactions that read with execute intent to pages mapped with `U` bit in PTE set to 1 are disallowed, regardless of the value of `SUM`. The software assigned process soft-context ID (`PSCID`) is used as the address space ID for the process identified by the first-stage page table when first-stage address translation is not Bare. ##### [](#first-stage-context-fsc-2)First-Stage context (`fsc`) ![Process First-Stage context](_images/diag-a25509c120a23fbda2a8aefe812e0f6ffd0215db.svg) Figure 21\. Process First-Stage context The `PC.fsc` field provides the controls for first-stage address translation and protection. The `PC.fsc.MODE` is used to determine the first-stage paged virtual-memory scheme and its encodings are as defined in [Table 3](#IOSATP%5FMODE%5FENC-0) and[Table 4](#IOSATP%5FMODE%5FENC-1). The `DC.tc.SXL` field controls the supported paged virtual-memory schemes. When `PC.fsc.MODE` is not `Bare`, the `PC.fsc.PPN` field holds the PPN of the root page of a first-stage page table. When second-stage address translation is not Bare, the `PC.fsc.PPN` field holds a guest PPN of the root of a first-stage page table. Addresses of the first-stage page table entries are then converted by guest physical address translation process, as controlled by the `DC.iohgatp`, into a supervisor physical address. A guest OS may thus directly edit the first-stage page table to limit access by the device to a subset of its memory and specify permissions for the device accesses. | | The PC.ta.PSCID identifies an address space. If an identicalPSCID is configured in two PC when the page-table referenced by the two PCare not identical then it is unpredictable whether the IOMMU uses the PTEs from the first page table or the second page table. These are the only expected behaviors. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#PC%5FMISCONFIG)Process-context configuration checks A `PC` with `PC.ta.V=1` is considered as misconfigured if any of the following conditions are true. If misconfigured then stop and report "PDT entry misconfigured" (cause = 267). 1. If any bits or encoding that are reserved for future standard use are set 2. `PC.fsc.MODE` encoding is not valid as determined by [Table 3](#IOSATP%5FMODE%5FENC-0) and[Table 4](#IOSATP%5FMODE%5FENC-1). 3. `DC.tc.SXL` is 0 and `PC.fsc.MODE` is not one of the supported modes 1. `capabilities.Sv39` is 0 and `PC.fsc.MODE` is `Sv39` 2. `capabilities.Sv48` is 0 and `PC.fsc.MODE` is `Sv48` 3. `capabilities.Sv57` is 0 and `PC.fsc.MODE` is `Sv57` 4. `DC.tc.SXL` is 1 and `PC.fsc.MODE` is not one of the supported modes 1. `capabilities.Sv32` is 0 and `PC.fsc.MODE` is `Sv32` | | Some PC fields hold supervisor physical addresses or guest physical addresses. Some implementations may verify the validity of the addresses - e.g. the supervisor physical address is not wider than that supported as determined by capabilities.PAS, etc. at the time of locating the PC. Such implementations may cause a "PDT entry misconfigured" (cause = 267) fault. Other implementations only detect such addresses to be invalid when the data structure referenced by these fields needs to be accessed. Such implementations may detect access-violation faults in the process of making the access. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#P2IOVA)Process to translate an IOVA The process to translate an IOVA uses the hardware IDs (`device_id` and`process_id`) to locate the Device-Context and the Process-Context. The Device-context and Process-context provide the root PPN of the page tables,`PSCID`, `GSCID`, and other control parameters that affect the address translation and protection process. When address translation caches ([Caching in-memory data structures](#CACHING)) are implemented, the translation process may use the `GSCID` and`PSCID` to associate the cached translations with their address spaces. The process to translate an `IOVA` is as follows: 1. If `ddtp.iommu_mode == Off` then stop and report "All inbound transactions disallowed" (cause = 256). 2. If `ddtp.iommu_mode == Bare` and any of the following conditions hold then stop and report "Transaction type disallowed" (cause = 260); else go to step 20 with translated address same as the `IOVA`. 1. Transaction type is a Translated request (read, write/AMO, read-for-execute) or is a PCIe ATS Translation request. 3. If `capabilities.MSI_FLAT` is 0 then the IOMMU uses base-format device context. Let `DDI[0]` be `device_id[6:0]`, `DDI[1]` be `device_id[15:7]`, and`DDI[2]` be `device_id[23:16]`. 4. If `capabilities.MSI_FLAT` is 1 then the IOMMU uses extended-format device context. Let `DDI[0]` be `device_id[5:0]`, `DDI[1]` be `device_id[14:6]`, and`DDI[2]` be `device_id[23:15]`. 5. If the `device_id` is wider than that supported by the IOMMU mode, as determined by the following checks then stop and report "Transaction type disallowed" (cause = 260). 1. `ddtp.iommu_mode` is `2LVL` and `DDI[2]` is not 0 2. `ddtp.iommu_mode` is `1LVL` and either `DDI[2]` is not 0 or `DDI[1]` is not 0 6. Use `device_id` to then locate the device-context (`DC`) as specified in[Process to locate the Device-context](#GET%5FDC). 7. If any of the following conditions hold then stop and report "Transaction type disallowed" (cause = 260). 1. Transaction type is a Translated request (read, write/AMO, read-for-execute) or is a PCIe ATS Translation request and `DC.tc.EN_ATS` is 0. 2. Transaction has a valid `process_id` and `DC.tc.PDTV` is 0. 3. Transaction has a valid `process_id` and `DC.tc.PDTV` is 1 and the`process_id` is wider than that supported by `pdtp.MODE`. 4. Transaction type is not supported by the IOMMU. 8. If request is a Translated request and `DC.tc.T2GPA` is 0 then the translation process is complete. Go to step 20. 9. If request is a Translated request and `DC.tc.T2GPA` is 1 then the IOVA is a GPA. Go to step 17 with following page table information: 1. Let `A` be the `IOVA` (the `IOVA` is a GPA). 2. Let `iosatp.MODE` be `Bare` 1. The `PSCID` value is not used when first-stage is Bare. 3. Let `iohgatp` be the value in the `DC.iohgatp` field 10. If `DC.tc.PDTV` is set to 0 then go to step 17 with the following page table information: 1. Let `iosatp.MODE` be the value in the `DC.fsc.MODE` field 2. Let `iosatp.PPN` be the value in the `DC.fsc.PPN` field 3. Let `PSCID` be the value in the `DC.ta.PSCID` field 4. Let `iohgatp` be the value in the `DC.iohgatp` field 11. If `DPE` is 1 and there is no `process_id` associated with the transaction then let `process_id` be the default value of 0. 12. If `DPE` is 0 and there is no `process_id` associated with the transaction then then go to step 17 with the following page table information: 1. Let `iosatp.MODE` be `Bare` 1. The `PSCID` value is not used when first-stage is Bare. 2. Let `iohgatp` be the value in the `DC.iohgatp` field 13. If `DC.fsc.pdtp.MODE = Bare` then go to step 17 with the following page table information: 1. Let `iosatp.MODE` be `Bare` 1. The `PSCID` value is not used when first-stage is Bare. 2. Let `iohgatp` be value in `DC.iohgatp` field 14. Locate the process-context (`PC`) as specified in [Process to locate the Process-context](#GET%5FPC). 15. if any of the following conditions hold then stop and report "Transaction type disallowed" (cause = 260). 1. The transaction requests supervisor privilege but `PC.ta.ENS` is not set. 16. Go to step 17 with the following page table information: 1. Let `iosatp.MODE` be the value in the `PC.fsc.MODE` field 2. Let `iosatp.PPN` be the value in the `PC.fsc.PPN` field 3. Let `PSCID` be the value in the `PC.ta.PSCID` field 4. Let `iohgatp` be the value in the `DC.iohgatp` field 17. Use the process specified in Section "Two-Stage Address Translation" of the RISC-V Privileged specification \[[7](bibliography.html#bib-priv)\] to determine the GPA accessed by the transaction. If a fault is detected by the first stage address translation process then stop and report the fault. If the translation process is completed successfully then let `A` be the translated GPA. 18. If MSI address translations using MSI page tables is enabled (i.e., `DC.msiptp.MODE != Off`) then the MSI address translation process specified in [Process to translate addresses of MSIs](#MSI%5FTRANS) is invoked. If the GPA `A` is not determined to be the address of a virtual interrupt file then the process continues at step 19\. If a fault is detected by the MSI address translation process then stop and report the fault else the process continues at step 20. 19. Use the second-stage address translation process specified in Section "Two-Stage Address Translation" of the RISC-V Privileged specification \[[7](bibliography.html#bib-priv)\] to translate the GPA `A` to determine the SPA accessed by the transaction. If a fault is detected by the address translation process then stop and report the fault. 20. Translation process is complete When checking the `U` bit in a second-stage PTE, the transaction is treated as not requesting supervisor privilege. The `pte.xwr=010` encoding, as specified by the Zicfiss \[[8](bibliography.html#bib-cfi)\] extension for the Shadow Stack page type in single-stage and VS-stage page tables, remains a reserved encoding for IO transactions. When the translation process reports a fault, and the request is an Untranslated request or a Translated request, the IOMMU requests the IO bridge to abort the transaction. Guidelines for handling faulting transactions in the IO bridge are provided in [iommu\_hw\_guidelines.adoc#IOBR\_FAULT\_RESP](iommu%5Fhw%5Fguidelines.html#IOBR%5FFAULT%5FRESP). The fault may be reported using the fault/event reporting mechanism and fault record formats specified in [iommu\_in\_memory\_queues.adoc#FAULT\_QUEUE](iommu%5Fin%5Fmemory%5Fqueues.html#FAULT%5FQUEUE). If the fault was detected by a PCIe ATS Translation Request then the IOMMU may provide a PCIe protocol defined response instead of reporting fault to software or causing an abort. The handling of faulting PCIe ATS Translation Requests is specified in [PCIe ATS translation request handling](#ATS%5FFAULTS). #### [](#GET%5FDC)Process to locate the Device-context The process to locate the Device-context for transaction using its `device_id`is as follows: 1. Let `a` be `ddtp.PPN x 212` and let `i = LEVELS - 1`. When`ddtp.iommu_mode` is `3LVL`, `LEVELS` is three. When `ddtp.iommu_mode` is`2LVL`, `LEVELS` is two. When `ddtp.iommu_mode` is `1LVL`, `LEVELS` is one. 2. If `i == 0` go to step 8. 3. Let `ddte` be the value of the eight bytes at address `a + DDI[i] x 8`. If accessing`ddte` violates a PMA or PMP check, then stop and report "DDT entry load access fault" (cause = 257). 4. If `ddte` access detects a data corruption (a.k.a. poisoned data), then stop and report "DDT data corruption" (cause = 268). 5. If `ddte.V == 0`, stop and report "DDT entry not valid" (cause = 258). 6. If any bits or encoding that are reserved for future standard use are set within `ddte`, stop and report "DDT entry misconfigured" (cause = 259). 7. Let `i = i - 1` and let `a = ddte.PPN x 212`. Go to step 2. 8. Let `DC` be the value of `DC_SIZE` bytes at address `a + DDI[0] * DC_SIZE`. If`capabilities.MSI_FLAT` is 1 then `DC_SIZE` is 64-bytes else it is 32-bytes. If accessing `DC` violates a PMA or PMP check, then stop and report "DDT entry load access fault" (cause = 257). If `DC` access detects a data corruption (a.k.a. poisoned data), then stop and report "DDT data corruption" (cause = 268). 9. If `DC.tc.V == 0`, stop and report "DDT entry not valid" (cause = 258). 10. If the `DC` is misconfigured as determined by rules outlined in[Device-context configuration checks](#DC%5FMISCONFIG) then stop and report "DDT entry misconfigured" (cause = 259). 11. The device-context has been successfully located. #### [](#GET%5FPC)Process to locate the Process-context The device-context provides the PDT root page PPN (`pdtp.ppn`). When`DC.iohgatp.mode` is not `Bare`, `pdtp.PPN` as well as `pdte.PPN` are Guest Physical Addresses (GPA) which must be translated into Supervisor Physical Addresses (SPA) using the second-stage page table pointed to by `DC.iohgatp`. The memory accesses to the PDT are treated as implicit read memory accesses by the second-stage. However, any guest-page fault exception raised by the second stage is always reported using the original access type (instruction, load, or store/AMO). An access fault in the second stage is reported as "PDT entry load access fault" (`cause = 265`). If the second-stage accesses detect data corruption (i.e., poisoned data), it is reported as "PDT data corruption" (`cause = 269`). The process to locate the Process-context for a transaction using its`process_id` is as follows: 1. Let `a` be `pdtp.PPN x 212` and let `i = LEVELS - 1`. When`pdtp.MODE` is `PD20`, `LEVELS` is three. When `pdtp.MODE` is`PD17`, `LEVELS` is two. When `pdtp.MODE` is `PD8`, `LEVELS` is one. 2. If `i != 0`, then let `a = a + PDI[2] × 8`; otherwise, let`a = a + PDI[0] × 16`. 3. If `DC.iohgatp.mode != Bare`, then `a` is a GPA. Invoke the process to translate `a` to a SPA as an implicit memory access. If faults occur during second-stage address translation of `a` then stop and report the fault detected by the second-stage address translation process. The translated `a` is used in subsequent steps. 4. If `i == 0` go to step 10. 5. Let `pdte` be the value of the eight bytes at address `a`. If accessing `pdte` violates a PMA or PMP check, then stop and report "PDT entry load access fault" (cause = 265). 6. If `pdte` access detects a data corruption (a.k.a. poisoned data), then stop and report "PDT data corruption" (cause = 269). 7. If `pdte.V == 0`, stop and report "PDT entry not valid" (cause = 266). 8. If any bits or encoding that are reserved for future standard use are set within `pdte`, stop and report "PDT entry misconfigured" (cause = 267). 9. Let `i = i - 1` and let `a = pdte.PPN x 212`. Go to step 2. 10. Let `PC` be the value of the 16-bytes at address `a`. If accessing `PC`violates a PMA or PMP check, then stop and report "PDT entry load access fault" (cause = 265). If `PC` access detects a data corruption (a.k.a. poisoned data), then stop and report "PDT data corruption" (cause = 269). 11. If `PC.ta.V == 0`, stop and report "PDT entry not valid" (cause = 266). 12. If the `PC` is misconfigured as determined by rules outlined in[Process-context configuration checks](#PC%5FMISCONFIG) then stop and report "PDT entry misconfigured" (cause = 267). 13. The Process-context has been successfully located. #### [](#MSI%5FTRANS)Process to translate addresses of MSIs When an I/O device is configured directly by a guest operating system, MSIs from the device are expected to be targeted to virtual IMSICs within the guest OS’s virtual machine, using guest physical addresses that are inappropriate and unsafe for the real machine. An IOMMU must recognize certain incoming writes from such devices as MSIs and convert them as needed for the real machine. MSIs originating from a single device that require conversion are expected to have been configured at the device by a single guest OS running within one RISC-V virtual machine. Assuming the VM itself conforms to the RISC-V Advanced Interrupt Architecture \[[6](bibliography.html#bib-aia)\], MSIs are sent to virtual harts within the VM by writing to the memory-mapped registers of the interrupt files of virtual IMSICs. Each of these virtual interrupt files occupies a separate 4-KiB page in the VM’s guest physical address space, the same as real interrupt files do in a real machine’s physical address space. A write to a guest physical address can thus be recognized as an MSI to a virtual hart if the write is to a page occupied by an interrupt file of a virtual IMSIC within the VM. When MSI address translation is supported (`capabilities.MSI_FLAT`, [iommu\_registers.adoc#CAP](iommu%5Fregisters.html#CAP)), the process to identify an incoming `IOVA` as the address of a virtual interrupt file and translating the address using the MSI page table is as follows: 1. Let `A` be the `GPA` 2. Let `DC` be the device-context located using the `device_id` of the device using the process outlined in [Process to locate the Device-context](#GET%5FDC). 3. Determine if the address `A` is an access to a virtual interrupt file as specified in [MSI address mask (msi\_addr\_mask) and pattern (msi\_addr\_pattern)](#MSI%5FID). 4. If the address is not determined to be that of a virtual interrupt file then stop this process and instead use the regular translation data structures to do the address translation. 5. Extract an interrupt file number `I` from `A` as`I = extract(A >> 12, DC.msi_addr_mask)`. The bit extract function`extract(x, y)` discards all bits from `x` whose matching bits in the same positions in the mask `y` are zeros, and packs the remaining bits from `x`contiguously at the least-significant end of the result, keeping the same bit order as `x` and filling any other bits at the most-significant end of the result with zeros. For example, if the bits of `x` and `y` are: * `x = a b c d e f g h` * `y = 1 0 1 0 0 1 1 0` * then the value of `extract(x, y)` has bits `0 0 0 0 a c f g`. 6. Let `m` be `(DC.msiptp.PPN x 212)`. 7. Let `msipte` be the value of sixteen bytes at address `(m | (I x 16))`. If accessing `msipte` violates a PMA or PMP check, then stop and report "MSI PTE load access fault" (cause = 261). 8. If `msipte` access detects a data corruption (a.k.a. poisoned data), then stop and report "MSI PT data corruption" (cause = 270). 9. If `msipte.V == 0`, then stop and report "MSI PTE not valid" (cause = 262). 10. If `msipte.C == 1`, then further processing to interpret the PTE is implementation defined. 11. If `msipte.C == 0` then the process is outlined in subsequent steps. 12. If `msipte.M == 0` or `msipte.M == 2`, then stop and report "MSI PTE misconfigured" (cause = 263). 13. If `msipte.M == 3` the PTE is in basic translate mode and the translation process is as follows: 1. If any bits or encoding that are reserved for future standard use are set within `msipte`, stop and report "MSI PTE misconfigured" (cause = 263). 2. Compute the translated address as `msipte.PPN << 12 | A[11:0]`. 14. If `msipte.M == 1` the PTE is in MRIF mode and the translation process is as follows: 1. If `capabilities.MSI_MRIF == 0`, stop and report "MSI PTE misconfigured" (cause = 263). 2. If any bits or encoding that are reserved for future standard use are set within `msipte`, stop and report "MSI PTE misconfigured" (cause = 263). 3. The address of the destination MRIF is `msipte.MRIF_Address[55:9] * 512`. 4. The destination address of the notice MSI is `msipte.NPPN << 12`. 5. Let `NID` be `(msipte.N10 << 10) | msipte.N[9:0]`. The data value for notice MSI is the 11-bit `NID` value zero-extended to 32-bits. 15. The access permissions associated with the translation determined through this process are equivalent to that of a regular RISC-V second-stage PTE with`R`\=`W`\=`U`\=1 and `X`\=0\. Similar to a second-stage PTE, when checking the `U`bit, the transaction is treated as not requesting supervisor privilege. 1. If the transaction is an Untranslated or Translated read-for-execute then stop and report "Instruction access fault" (cause = 1). 16. MSI address translation process is complete. | | Unlike regular RISC-V leaf PTEs, MSI PTEs do not have an accessed (A) or dirty (D) bit. An IOMMU may treat an MSI PTE **as if** the A and D bits are always set to 1. In MRIF mode, the Advanced Interrupt Architecture Specification defines the operation to store the incoming MSIs into the destination MRIF and to generate the notice MSI. These operations may be performed by the IOMMU itself or the IOMMU may provide the destination MRIF address, the notice MSI address, and the notice MSI data value to the I/O bridge in response to the translation request and the operations may be performed by the I/O bridge. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#iommu-updating-of-pte-accessed-a-and-dirty-d-updates)IOMMU updating of PTE accessed (A) and dirty (D) updates When `capabilities.AMO_HWAD` is 1, the IOMMU supports updating the A and D bits in PTEs atomically. When updating of A and D bits in second-stage PTEs is enabled (`DC.tc.GADE=1`) and/or updating of A and D bits in first-stage PTEs is enabled (`DC.tc.SADE=1`) the following rules apply: 1. The A and/or D bit updates by the IOMMU must follow the rules specified by the Privileged specification for validity, permission checking, and atomicity. 2. The PTE update must be globally visible before a memory access using the translated address provided by the IOMMU becomes globally visible. Specifically, when a translated address is provided to a device in an ATS Translation completion, the PTE update must be globally visible before a memory access from the device using the translated address becomes globally visible. | | The A and D bits are never cleared by the IOMMU. If the supervisor software does not rely on accessed and/or dirty bits, e.g. if it does not swap memory pages to secondary storage or if the pages are being used to map I/O space, it should set them to 1 in the PTE to improve performance. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#faults-from-virtual-address-translation-process)Faults from virtual address translation process Faults detected during the two-stage address translation specified in the RISC-V Privileged specification \[[7](bibliography.html#bib-priv)\] cause the IOVA translation process to stop and report the detected fault. ### [](#ATS%5FFAULTS)PCIe ATS translation request handling ATS \[[5](bibliography.html#bib-pci)\] translation requests that encounter a configuration error results in a Completer Abort (CA) response to the requester. The following cause codes belong to this category: * Instruction access fault (cause = 1) * Read access fault (cause = 5) * Write/AMO access fault (cause = 7) * MSI PTE load access fault (cause = 261) * MSI PTE misconfigured (cause = 263) * PDT entry load access fault (cause = 265) * PDT entry misconfigured (cause = 267) If there is a permanent error or if ATS transactions are disabled then an Unsupported Request (UR) response is generated. The following cause codes belong to this category: * All inbound transactions disallowed (cause = 256) * DDT entry load access fault (cause = 257) * DDT entry not valid (cause = 258) * DDT entry misconfigured (cause = 259) * Transaction type disallowed (cause = 260) When translation could not be completed due to the following causes a Success Response with R and W bits set to 0 is generated. No faults are logged in the fault queue on these errors. The translated address returned with such completions is `UNSPECIFIED`. * Instruction page fault (cause = 12) * Read page fault (cause = 13) * Write/AMO page fault (cause = 15) * Instruction guest page fault (cause = 20) * Read guest-page fault (cause = 21) * Write/AMO guest-page fault (cause = 23) * PDT entry not valid (cause = 266) * MSI PTE not valid (cause = 262) If the translation request has a PASID with "Privilege Mode Requested" field set to 0, or the request does not have a PASID then the request does not target privileged memory. If the U-bit that indicates if the memory is accessible to user mode is 0 then a Success response with R and W bits set to 0 is generated. If the translation request has a PASID with "Privilege Mode Requested" field set to 1, then the request targets privileged memory. If the U-bit that indicates if the page is accessible to user mode is 1 and the `SUM` bit in the `ta` field of the process-context is 0 then a Success response with R and W bits set to 0 is generated. If the translation could be successfully completed but the requested permissions are not present in either stage (Execute requested but no execute permission; no-write not requested and no write permission; no read permission) then a Success response is returned with the denied permission (R, W or X) set to 0 and the other permission bits set to the value determined from the page tables. The X permission is granted only if the R permission is also granted and the execute permission was requested. Execute-only translations are not compatible with PCIe ATS as PCIe requires read permission to be granted if the execute permission is granted. When a Success response is generated for an ATS translation request, no fault records are reported to software through the fault/event reporting mechanism, even when the response indicates no access was granted or some permissions were denied. Conversely, when a UR or CA response is generated for an ATS translation request, the corresponding fault is reported to software through the fault/event reporting mechanism. If the translation request is successfully completed and the address is determined to be an MSI address using the rules defined by the [MSI address mask (msi\_addr\_mask) and pattern (msi\_addr\_pattern)](#MSI%5FID), but the MSI PTE is configured in MRIF mode, a Success response is generated with the U bit (Untranslated access only) set to 1\. The U bit being set to 1 in the response instructs the device that it must use only Untranslated requests to access the implied 4 KiB memory range. The R, W, and Exe bits in the response indicate the granted permissions. | | When a MSI PTE is configured in MRIF mode, a MSI write with data value Drequires the IOMMU to set the interrupt-pending bit for interrupt identity Din the MRIF. A translation request from a device to a GPA that is mapped through a MRIF mode MSI PTE is not eligible to receive a translated address. This is accomplished by setting "Untranslated Access Only" (U) field of the returned response to 1. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The translation range size returned in a Success response to an ATS translation request, when either stages of address translation are Bare, is implementation-defined. However, it is recommended that the translation range size be large, such as 2 MiB or 1 GiB. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When a Success response is generated for an ATS translation request, the setting of the Priv, N, CXL.io, Global, and AMA fields is as follows: * Priv field of the ATS translation completion is always set to 0 if the request does not have a PASID. When a PASID is present then the Priv field is set to the value in "Privilege Mode Requested" field as the permissions provided correspond to those the privilege mode indicate in the request. * N field of the ATS translation completion is always set to 0\. The device may use other means to determine if the No-snoop flag should be set in the translated requests. * Global field is set to the value determined from the first-stage page tables if translation could be successfully completed and the request had a PASID present. In all other cases, including MSI address translations, this field is set to 0. * If requesting device is not a CXL device then CXL.io is set to 0. * If requesting device is a CXL type 1 or type 2 device * If the address is determined to be a MSI then the CXL.io bit is set to 1. * Else if `T2GPA` is 1 in the device context then the CXL.io bit is set to 1. * Else if the memory type, as determined by the Svpbmt extension, is NC or IO then the CXL.io bit is set to 1\. If the memory type is PMA then the determination of the setting of this bit is `UNSPECIFIED`. If the Svpbmt extension is not supported then the setting of this bit is `UNSPECIFIED`. * In all other cases the setting of this bit is `UNSPECIFIED`. * The AMA field is by default set to 000b. The IOMMU may support an implementation-specific method to provide other encodings. | | The IO bridge may override the CXL.io bit in the ATS translation completion based on the PMA of the translated address. Other implementations may provide an implementation-defined method for determining PMA for the translated address to set the CXL.io bit. Use of T2GPA set to 1 may not be compatible with CXL type 1 or type 2 devices as they use the CXL.cache protocol to implement caches tagged by the translated address returned in response to a PCIe ATS Translation Request. The IOMMU may not be invoked for translating addresses in CXL.cache transactions. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#ATS%5FPRI)PCIe ATS Page Request handling To process a "Page Request" or "Stop Marker" message \[[5](bibliography.html#bib-pci)\], the IOMMU first locates the device-context—​using the procedure outlined in steps 1 through 5 of [Process to translate an IOVA](#P2IOVA)\--to determine if ATS and PRI are enabled for the requester. If ATS and PRI are enabled, i.e. `EN_ATS` and `EN_PRI` are both set to 1, the IOMMU queues the message into an in-memory queue called the page-request-queue (`PQ`) (See [iommu\_in\_memory\_queues.adoc#PRQ](iommu%5Fin%5Fmemory%5Fqueues.html#PRQ)). Following suitable processing of the "Page Request", a software handler may generate a "Page Request Group Response" message to the device. When PRI is enabled for a device, the IOMMU may still be unable to report "Page Request" or "Stop Marker" messages through the `PQ` due to error conditions such as the queue being disabled, queue being full, or the IOMMU encountering access faults when attempting to access queue memory. These error conditions are specified in [iommu\_in\_memory\_queues.adoc#PRQ](iommu%5Fin%5Fmemory%5Fqueues.html#PRQ). If the `ddtp.iommu_mode` is `Bare` or is `Off`, then the IOMMU cannot locate a device-context for the requester. If `EN_PRI` is set to 0, or `EN_ATS` is set to 0, or if the IOMMU is unable to locate the `DC` to determine the `EN_PRI` configuration, or the request could not be queued into `PQ` then the IOMMU behavior depends on the type of "Page Request". * If the "Page Request" does not require a response, i.e. the "Last Request in PRG" field of the message is set to 0, then such messages are silently discarded. "Stop Marker" messages do not require a response and are always silently discarded on such errors. * If the "Page Request" needs a response, then the IOMMU itself may generate a "Page Request Group Response" message to the device. When the IOMMU generates the response, the status field of the response depends on the cause of the error. If a fault condition prevents locating a valid device context then the `PRPR` value assumed is 0. The status is set to Response Failure if the following faults are encountered: * `ddtp.iommu_mode` is `Off` (cause = 256) * DDT entry load access fault (cause = 257) * DDT entry misconfigured (cause = 259) * DDT entry not valid (cause = 258) * Page-request queue is not enabled (`pqcsr.pqen == 0` or `pqcsr.pqon == 0`) * Page-request queue encountered a memory access fault (`pqcsr.pqmf == 1`) The status is set to Invalid Request if the following faults are encountered: * `ddtp.iommu_mode` is `Bare` (cause = 260) * `EN_PRI` is set to 0 (cause = 260) * `device_id` is wider than that supported by the IOMMU mode (cause = 260) The status is set to Success if no other faults were encountered but the "Page Request" could not be queued due to the page-request queue being full (`pqt == pqh - 1`) or had a overflow (`pqcsr.pqof == 1`). | | When SR-IOV VF is used as a unit of allocation, a hypervisor may disable page requests from one of the virtual functions by setting EN\_PRI to 0\. However the page-request interface is shared by the PF and all VFs. The IOMMU protocol specific logic classifies this condition (cause = 260) as a non-catastrophic failure, an Invalid Request, in its response to avoid the shared PRI in the device being disabled for all PFs/VFs. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | A "Stop Marker" is encoded as a "Page Request" with a PASID but with the L, W, and R fields set to 1, 0, and 0 respectively. | | ------------------------------------------------------------------------------------------------------------------------------- | For IOMMU-generated "Page Request Group Response" messages that have status Invalid Request or Success, the PRG-response-PASID-required (`PRPR`) bit when set to 1 indicates that the IOMMU response message should include a PASID if the associated "Page Request" had a PASID. For IOMMU-generated "Page Request Group Response" with response code set to Response Failure, if the "Page Request" had a PASID then response is generated with a PASID. No faults are logged in the fault queue for PCIe ATS "Page Request" messages for the following conditions: * Page-request queue is not enabled (`pqcsr.pqen == 0` or `pqcsr.pqon == 0`) * Page-request queue encountered a memory access fault (`pqcsr.pqmf == 1`) * "Page Request" could not be queued due to the page-request queue being full (`pqt == pqh - 1`) or had a overflow (`pqcsr.pqof == 1`). ### [](#CACHING)Caching in-memory data structures To speed up Direct Memory Access (DMA) translations, the IOMMU may make use of translation caches to hold entries from device-directory-table, process-directory-table, first-stage and second-stage translation tables, and MSI page tables. These caches are collectively referred to as the IOMMU Address Translation Caches (IOATC). This specification does not allow the caching of first/second-stage PTEs whose`V` (valid) bit is clear, non-leaf DDT entries whose `V` (valid) bit is clear, Device-context whose `V` (valid) bit is clear, non-leaf PDT entries whose `V`(valid) bit is clear, Process-context whose `V` (valid) bit is clear, or MSI PTEs whose `V` bit is clear. These IOATC do not observe modifications to the in-memory data structures using explicit loads and stores by RISC-V harts or by device DMA. Software must use the IOMMU commands to invalidate the cached data structure entries using IOMMU commands to synchronize the IOMMU operations to observe updates to in-memory data structures. A simpler implementation may not implement IOATC for some or any of the in-memory data structures. The IOMMU commands may use one or more IDs to tag the cached entries to identify a specific entry or a group of entries. __Table 7\. Identifiers used to tag IOATC entries__ | Data Structure cached | IDs used to tag entries | Invalidation command | | ------------------------------------------------------ | ----------------------- | ---------------------------------------------------------- | | Device Directory Table | device\_id | [IODIR.INVAL\_DDT](iommu%5Fin%5Fmemory%5Fqueues.html#IDDT) | | Process Directory Table | device\_id, process\_id | [IODIR.INVAL\_PDT](iommu%5Fin%5Fmemory%5Fqueues.html#IPDT) | | First-stage page table (when second-stage is not Bare) | GSCID, PSCID, and IOVA | [IOTINVAL.VMA](iommu%5Fin%5Fmemory%5Fqueues.html#IVMA) | | First-stage page table (when second-stage is Bare) | PSCID, and IOVA | [IOTINVAL.VMA](iommu%5Fin%5Fmemory%5Fqueues.html#IVMA) | | Second-stage page table | GSCID, GPA | [IOTINVAL.GVMA](iommu%5Fin%5Fmemory%5Fqueues.html#IGVMA) | | MSI page table | GSCID, GPA | [IOTINVAL.GVMA](iommu%5Fin%5Fmemory%5Fqueues.html#IGVMA) | ### [](#updating-in-memory-data-structure-entries)Updating in-memory data structure entries The RISC-V memory model requires memory access from a hart to be single-copy atomic. When RV32 is implemented the size of a single-copy atomic memory access is up to 32-bits. When RV64 is implemented the size of a single-copy atomic memory access is up to 64-bits. The size of a single-copy atomic memory access implemented by the IOMMU is `UNSPECIFIED` but is required to be at least 32-bits if all of the harts in the system implement RV32 and is required to be at least 64-bits if any of the harts in the system implement RV64. The IOMMU data structure entries have a `V` bit that when set to 1 indicates that the entry is valid. Software is allowed to make updates to a data structure entry that has the `V`bit set to 1\. However, some rules as outlined below must be followed. * It may be unsafe for software to partially update the fields of a valid data structure entry, as it is legal for an IOMMU to read the entry at any time, including when only some of the partial updates have taken effect. * For an update to an IOMMU data structure entry to be atomically observed by the IOMMU, software must use a store that results in a single memory operation. * If the update to a field will make the field inconsistent with another field of the entry then software must first set the `V` field to 0 and use the commands outlined in [Caching in-memory data structures](#CACHING) to invalidate any previous copies of that entry that may be in IOMMU caches before updating other fields of that entry. * The IOMMU is not required to immediately observe the software update to an entry. Software must use the commands outlined in [Caching in-memory data structures](#CACHING) to invalidate any previous copies of that entry that may be in IOMMU caches to synchronize the updates to the entry with the operation of the IOMMU. | | If a data structure entry is changed, the IOMMU may use the old value of the entry or the new value of the entry and the choice is unpredictable until software uses the commands outlined in [Caching in-memory data structures](#CACHING) to invalidate any previous copies of that entry that may be in IOMMU caches to synchronize updates to the entry with the operation of the IOMMU. These are the only behaviors expected. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#endianness-of-in-memory-data-structures)Endianness of in-memory data structures The RISC-V memory model specifies byte-invariance for the entire address space. When mixed-endian mode of operation is supported, the IO bridge and the IOMMU must implement byte-invariant addressing such that a byte access to a given address accesses the same memory location in both little-endian and big-endian mode of operation. The endianness of implicit memory access to in-memory data structures is determined by `fctl.BE` or by `DC.tc.SBE` as follows: __Table 8\. Endianness of memory access to data structures__ | Data Structure | Controlled by | | ----------------------- | ------------- | | Device directory table | fctl.BE | | Second-stage page table | fctl.BE | | MSI page table | fctl.BE | | Process directory Table | DC.tc.SBE | | First-stage page table | DC.tc.SBE | | | The PSCID field of first-stage context, along with the GSCID (when two-stage address translation is active), identifies an address space. Configuring an identical GSCID and PSCID in two DC but with different SBE is not expected and if done may lead to the IOMMU interpreting a first-stage PTE as big-endian or little-endian. These are the only behaviors expected. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | Software must use an appropriate software sequence to swap bytes as necessary to create a mutually agreed to data representation when sharing data with an IO agent that does not share its endianness. Software must use an LR/SC sequence to perform atomic operations in non-native endian format when the data shared with such IO agents must be accessed atomically. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Debug support ==================== ## [](#debug)Debug support To support software debug, the IOMMU may provide an optional register interface that may be used by software to request IOMMU to perform an address translation. The IOMMU supports this capability when `capabilities.DBG` is 1\. The interface consists of two set of registers; translation-request registers that are used by software to program an IOVA and other inputs needed by the process to translate an IOVA ([iommu\_data\_structures.adoc#P2IOVA](iommu%5Fdata%5Fstructures.html#P2IOVA)) as an Untranslated Request. The result of the translation, if the process completes successfully, is reported through the translation-response registers. If the process stops due to faults then the faults are reported normally in the fault-queue and the translation-response registers updated with a failure indicator. If the IOVA is determined to be that of a virtual interrupt file ([iommu\_data\_structures.adoc#MSI\_ID](iommu%5Fdata%5Fstructures.html#MSI%5FID)) and the corresponding MSI PTE is in MRIF mode, then the process stops and reports a "Transaction type disallowed" (cause = 260) fault. When the process to translate an IOVA is invoked for this purpose, the IOMMU may or may not cache first-stage PTEs, second-stage PTEs, DDT entries, PDT entries, or MSI PTEs accessed for the translation process in the IOATC. The IOMMU is allowed to use any PTEs or directory structure entries that may already be cached in the IOATC. The IOMMU may update the Accessed (A) and/or Dirty (D) bits in the PTEs used for the translation process if supported by the IOMMU. When the IOMMU implements a HPM, the HPM counters may be updated normally by the IOMMU. For the purpose of counting in the HPM, these requests are treated as Untranslated Requests. The translation-request interface consists of the following 64-bit WARL registers: * `tr_req_iova` ([iommu\_registers.adoc#TRR\_IOVA](iommu%5Fregisters.html#TRR%5FIOVA)) * `tr_req_ctl` ([iommu\_registers.adoc#TRR\_CTRL](iommu%5Fregisters.html#TRR%5FCTRL)) The translation-response interface consists of a single 64-bit RO register`tr_response` ([iommu\_registers.adoc#TRR\_RSP](iommu%5Fregisters.html#TRR%5FRSP)) To request a translation, the `tr_req_iova` register is written first with the desired IOVA and the `tr_req_ctl` register is written next. The 'Go/Busy\` bit is set in `tr_req_ctl` to indicate a valid request in the registers. The`Go/Busy` bit is a read-write-1-to-set (RW1S) bit that once set cannot be cleared by writing the register. The `Go/Busy` bit will be cleared to 0 by the IOMMU when the process completes (successfully or due to encountering a fault). When the `Go/Busy` bit goes from 1 to 0, a response is valid in the `tr_response`register. When the `Go/Busy` bit is 1, the IOMMU behavior is `UNSPECIFIED` if: * The `tr_req_iova` or `tr_req_ctl` are modified. * IOMMU configurations, such as `ddtp.iommu_mode`, are modified. The time to complete a translation request through this debug interface is`UNSPECIFIED` but is required to be finite. If the IOMMU is serving translation requests from the IO bridge when a request is made through this register interface then the time to complete the request may be longer than when the IOMMU is otherwise idle. | | The debug interface is optional but recommended to be implemented to aid software debug and to implement architectural compliance tests. | | ------------------------------------------------------------------------------------------------------------------------------------------- | IOMMU Extensions ==================== ## [](#extensions)IOMMU Extensions This chapter specifies the following standard extensions to the IOMMU Base Architecture: | Specification | Version | Status | | ------------------------------------------------------------ | ------- | ------------ | | [**Quality-of-Service (QoS) Identifiers Extension**](#QOSID) | **1.0** | **Ratified** | | [**Non-leaf PTE Invalidation Extension**](#NLINV) | **1.0** | **Ratified** | | [**Address Range Invalidation Extension**](#ARINV) | **1.0** | **Ratified** | | [**PTE Reserved-for-Software Bits 60-59**](#SVRSW) | **1.0** | **Ratified** | ### [](#QOSID)Quality-of-Service (QoS) Identifiers Extension, Version 1.0 Quality of Service (QoS) is defined as the minimal end-to-end performance guaranteed in advance by a service level agreement (SLA) to a workload. Performance metrics might include measures such as instructions per cycle (IPC), latency of service, etc. When multiple workloads execute concurrently on modern processors — equipped with large core counts, multiple cache hierarchies, and multiple memory controllers — the performance of any given workload becomes less deterministic, or even non-deterministic, due to shared resource contention \[[10](bibliography.html#bib-ptcamp)\]. To manage performance variability, system software needs resource allocation and monitoring capabilities. These capabilities allow for the reservation of resources like cache and bandwidth, thus meeting individual performance targets while minimizing interference \[[11](bibliography.html#bib-heracles)\]. For resource management, hardware should provide monitoring features that allow system software to profile workload resource consumption and allocate resources accordingly. To facilitate this, the QoS Identifiers ISA extension (Ssqosid) \[[12](bibliography.html#bib-ssqosid)\] introduces the `srmcfg` register, which configures a hart with two identifiers: a Resource Control ID (`RCID`) and a Monitoring Counter ID (`MCID`). These identifiers accompany each request issued by the hart to shared resource controllers. These identifiers are crucial for the RISC-V Capacity and Bandwidth Controller QoS Register Interface \[[13](bibliography.html#bib-cbqri)\], which provides methods for setting resource usage limits and monitoring resource consumption. The `RCID` controls resource allocations, while the `MCID` is used for tracking resource usage. The IOMMU QoS ID extension provides a method to associate QoS IDs with requests to access resources by the IOMMU, as well as with devices governed by it. This complements the Ssqosid extension that provides a method to associate QoS IDs with requests originated by the RISC-V harts. Assocating QoS IDs with device and IOMMU originated requests is required for effective monitoring and allocation of shared resources. The IOMMU `capabilities` register ([iommu\_registers.adoc#CAP](iommu%5Fregisters.html#CAP)) is extended with a `QOSID` field which enumerates support for associating QoS IDs with requests made through the IOMMU. When `capabilities.QOSID` is 1, the memory-mapped register layout is extended to add a register named `iommu_qosid` ([iommu\_registers.adoc#IOQOSID](iommu%5Fregisters.html#IOQOSID)). This register is used to configure the Quality of Service (QoS) IDs associated with IOMMU-originated requests. The `ta` field of the device context ([iommu\_data\_structures.adoc#DC\_TA](iommu%5Fdata%5Fstructures.html#DC%5FTA)) is extended with two fields, `RCID` and `MCID`, to configure the QoS IDs to associate with requests originated by the devices. #### [](#reset-behavior)Reset Behavior If the reset value for `ddtp.iommu_mode` field is `Bare`, then the`iommu_qosid.RCID` field must have a reset value of 0. | | At reset, it is required that the RCID field of iommu\_qosid is set to 0 if the IOMMU is in Bare mode, as typically the resource controllers in the SoC default to a reset behavior of associating all capacity or bandwidth to theRCID value of 0\. When the reset value of the ddtp.iommu\_mode is not Bare, the iommu\_qosid register should be initialized by software before changing the mode to allow DMA. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#sizing-qos-identifiers)Sizing QoS Identifiers The size (or width) of `RCID` and `MCID`, as fields in registers or in data structures, supported by the IOMMU must be at least as large as that supported by any RISC-V application processor hart in the system. #### [](#iommu-atc-capacity-allocation-and-monitoring)IOMMU ATC Capacity Allocation and Monitoring Some IOMMUs might support capacity allocation and usage monitoring in the IOMMU address translation cache (IOATC) by implementing the capacity controller register interface. Additionally, some IOMMUs might support multiple IOATCs, each potentially having different capacities. In scenarios where multiple IOATCs are implemented, such as an IOATC for each supported page size, the IOMMU can implement a capacity controller register interface for each IOATC to facilitate individual capacity allocation. ### [](#NLINV)Non-leaf PTE Invalidation Extension, Version 1.0 The RISC-V IOMMU Version 1.0 specification provides commands to invalidate leaf page table entries from address translation caches when performing an address-specific invalidation operation. The non-leaf PTE invalidation extension provides commands to optionally also invalidate non-leaf PTE entries from the address translation caches when performing an address-specific invalidation operation. The non-leaf PTE invalidation extension is implemented if the `capabilities.NL`(bit 42) is 1\. When the `capabilities.NL` bit is 1, a non-leaf (`NL`) field is defined at bit 34 in the `IOTINVAL.VMA` and `IOTINVAL.GVMA` commands by this extension. When the `capabilities.NL` bit is 0, bit 34 remains reserved. | | The non-leaf PTE invalidation extension enables optimizations in shared virtual addressing use cases by providing the ability to invalidate non-leaf PTEs corresponding to the IOVA being invalidated from the IOMMU address translation caches. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If the address range invalidation extension is also implemented, the `NL`operand applies to the address range determined by the `ADDR` and `S` operands. #### [](#non-leaf-pte-invalidation-by-iotinval-vma)Non-leaf PTE Invalidation by `IOTINVAL.VMA` * When the `AV` operand is 0, the `NL` operand is ignored and the `IOTINVAL.VMA`command operations are as specified in RISC-V IOMMU Version 1.0 specification. * When the `AV` operand is 1 and the `NL` operand is 0, the `IOTINVAL.VMA`command operations are as specified in RISC-V IOMMU Version 1.0 specification. * When both the `AV` and `NL` operands are 1, the `IOTINVAL.VMA` command performs the following operations: * When `GV=0` and `PSCV=0`: Invalidates information cached from all levels of first-stage page table entries corresponding to the IOVA in the `ADDR`operand for all host address spaces, including entries containing global mappings. * When `GV=0` and `PSCV=1`: Invalidates information cached from all levels of first-stage page table entries corresponding to the IOVA in the `ADDR`operand and the host address space identified by the `PSCID` operand, except for entries containing global mappings. * When `GV=1` and `PSCV=0`: Invalidates information cached from all levels of first-stage page table entries corresponding to the IOVA in the `ADDR`operand for all VM address spaces associated with the `GSCID` operand, including entries that contain global mappings. * When `GV=1` and `PSCV=1`: Invalidates information cached from all levels of first-stage page table entries corresponding to the IOVA in the `ADDR`operand and the VM address space identified by the `PSCID` and `GSCID`operands, except for entries containing global mappings. #### [](#non-leaf-pte-invalidation-by-iotinval-gvma)Non-leaf PTE Invalidation by `IOTINVAL.GVMA` * When the `GV` operand is 0, both the `AV` and `NL` operands are ignored and the `IOTINVAL.GVMA` command operations are as specified in RISC-V IOMMU Version 1.0 specification. * When the `GV` operand is 1 and the `AV` operand is 0, the `NL` operand is ignored and the `IOTINVAL.GVMA` command operations are as specified in RISC-V IOMMU Version 1.0 specification. * When the `GV` and `AV` operands are 1 and the `NL` operand is 0, the`IOTINVAL.GVMA` command operations are as specified in RISC-V IOMMU Version 1.0 specification. * When `GV`, `AV`, and `NL` are all 1, the `IOTINVAL.GVMA` command performs the following operations: * Invalidates information cached from all levels of second-stage page table entries corresponding to the guest-physical address in the `ADDR` operand and the VM address spaces identified by the `GSCID` operand. ### [](#ARINV)Address Range Invalidation Extension, Version 1.0 The address range invalidation extension enables specifying a range of addresses in an IOMMU ATC invalidation command, reducing the number of commands queued to the IOMMU. This facility is especially useful when superpages are employed in page tables. The address range invalidation extension is implemented if `capabilities.S` (bit 43) is 1\. When `capabilities.S` is 1, a range-size (`S`) operand is defined at bit 73 in the `IOTINVAL.VMA` and `IOTINVAL.GVMA` commands by this extension. When the `capabilities.S` bit is 0, bit 73 remains reserved. When the `GV` operand is 0, both the `AV` and `S` operands are ignored by the`IOTINVAL.GVMA` command. When the `AV` operand is 0, the `S` operand is ignored in both the `IOTINVAL.VMA` and `IOTINVAL.GVMA` commands. When the `S` operand is ignored or set to 0, the operations of the `IOTINVAL.VMA` and `IOTINVAL.GVMA`commands are as specified in the RISC-V IOMMU Version 1.0 specification. When the `S` operand is not ignored and is 1, the `ADDR` operand represents a NAPOT range encoded in the operand itself. Starting from bit position 0 of the `ADDR` operand, if the first 0 bit is at position `X`, the range size is`2(X+1) * 4` KiB. When `X` is 0, the size of the range is 8 KiB. If the `S` operand is not ignored and is 1 and all bits of the `ADDR` operand are 1, the behavior is UNSPECIFIED. If the `S` operand is not ignored and is 1 and the most significant bit of the`ADDR` operand is 0 while all other bits are 1, the specified address range covers the entire address space. | | The NAPOT range encoding used by this extension follows the convention used by PCIe ATS Invalidation Requests to denote address ranges. This convention is also used to encode the translation range size in tr\_response ([iommu\_registers.adoc#TRR\_RSP](iommu%5Fregisters.html#TRR%5FRSP)) register. Simpler implementations may invalidate all address-translation cache entries when the S bit is set to 1. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#SVRSW)PTE Reserved-for-Software Bits 60-59, Version 1.0 The Svrsw60t59b extension is implemented if `capabilities.Svrsw60t59b` (bit 14) is set to 1. Hardware guidelines ==================== ## [](#hw%5Fguidelines)Hardware guidelines This section provides guidelines to the system/hardware integrator of the IOMMU in the platform. ### [](#integrating-an-iommu-as-a-pcie-device)Integrating an IOMMU as a PCIe device The IOMMU may be constructed as a PCIe device itself and be discoverable as a dedicated PCIe function with PCIe defined Base Class 08h, Sub-Class 06h, and Programming Interface 00h \[[9](bibliography.html#bib-pci-cls)\]. Such IOMMU must map the IOMMU registers defined in this specification as PCIe BAR mapped registers. The IOMMU may support MSI or MSI-X or both. When MSI-X is supported, the MSI-X capability block must point to the `msi_cfg_tbl` in BAR mapped registers such that system software can configure MSI address and data pairs for each message supported by the IOMMU. The MSI-X PBA may be located in the same BAR or another BAR of the IOMMU. The IOMMU is recommended to support MSI-X capability. ### [](#faults-from-pma-and-pmp)Faults from PMA and PMP The IO bridge may invoke a PMA and/or a PMP checker on memory accesses from IO devices or those generated by the IOMMU implicitly to access the in-memory data structures. When a memory access violates a PMA check or violates a PMP check, the IO bridge may abort the memory access as specified in[Aborting transactions](#IOBR%5FFAULT%5FRESP). ### [](#IOBR%5FFAULT%5FRESP)Aborting transactions If the aborted transaction is an IOMMU-initiated implicit memory access then the IO bridge signals such access faults to the IOMMU itself. The details of such signaling is implementation defined. If the aborted transaction is a write then the IO bridge may discard the write; the details of how the write is discarded are implementation defined. If the IO protocol requires a response for write transactions (e.g., AXI) then a response as defined by the IO protocol may be generated by the IO bridge (e.g., SLVERR on BRESP - Write Response channel). For PCIe, for example, write transactions are posted and no response is returned when a write transaction is discarded. If the faulting transaction is a read then the device expects a completion. The IO bridge may provide a completion to the device. The data, if returned, in such completion is implementation defined; usually it is a fixed value such as all 0 or all 1\. A status code may be returned to the device in the completion to indicate this condition. For AXI, for example, the completion status is provided by SLVERR on RRESP (Read Data channel). For PCIe, for example, the completion status field may be set to "Unsupported Request" (UR) or "Completer Abort" (CA). ### [](#RAS)Reliability, Availability, and Serviceability (RAS) The IOMMU may support a RAS architecture that specifies the methods for enabling error detection, logging the detected errors (including their severity, nature, and location), and configuring means to report the error to an error handler. Some errors, such as those in the IOATC, may be correctable by reloading the cached in-memory data structures when the error is detected. Such errors are not expected to affect the functioning of the IOMMU. Some errors may corrupt critical internal state of the IOMMU and such errors may lead the IOMMU to a failed state. Examples of such state may include registers such as the `ddtp`, `cqb`, etc. On entering such a failed state, the IOMMU may request the IO bridge to abort all incoming transactions. Some errors, such as corruptions that occur within the internal data paths of the IOMMU, may not be correctable but the effects of such errors may be contained to the transaction being processed by the IOMMU. As part of processing a transaction, the IOMMU may need to read data from in-memory data structures such as the DDT, PDT, or first/second-stage page tables. The provider (a memory controller or a cache) of the data may detect that the data requested has an uncorrectable error and signal that the data is corrupted and defer the error to the IOMMU. Such technique to defer the handling of the corrupted data to the consumer of the data is also commonly known as data poisoning. The effects of such errors may be contained to the transaction that caused the corrupted data to be accessed. In the cases where the error affects the transaction being processed but otherwise allows the IOMMU to continue providing service, the IOMMU may abort (see [Aborting transactions](#IOBR%5FFAULT%5FRESP)) the transaction and report the the fault by queuing a fault record in the `FQ`. For PCIe, for example, a "Completer Abort (CA)" response is appropriate to abort the transaction. The following cause codes are used to report such faulting transactions: * DDT data corruption (cause = 268) * PDT data corruption (cause = 269) * MSI PT data corruption (cause = 270) * MSI MRIF data corruption (cause = 271) * Internal data-path error (cause = 272) * First/second-stage PT data corruption (cause = 274) If the IO bridge is not capable of signaling such deferred errors uniquely from other errors that prevent the IOMMU from accessing in-memory data structures then the IOMMU may report such errors as access faults instead of using the differentiated data corruption cause codes. In-memory queue interface ==================== ## [](#in-memory-queue-interface)In-memory queue interface Software and IOMMU interact using 3 in-memory queue data structures. * A command-queue (`CQ`) used by software to queue commands to the IOMMU. * A fault/event queue (`FQ`) used by IOMMU to bring faults and events to software’s attention. * A page-request queue (`PQ`) used by IOMMU to report “Page Request” messages received from PCIe devices. This queue is supported if the IOMMU supports PCIe \[[5](bibliography.html#bib-pci)\] defined Page Request Interface. ![IOMMU in-memory queues](_images/diag-31d49ae9e61f4553e40e1b0324306429d60e8a21.svg) Figure 1\. IOMMU in-memory queues Each queue is a circular buffer with a head controlled by the consumer of data from the queue and a tail controlled by the producer of data into the queue. IOMMU is the producer of records into `PQ` and `FQ` and controls the tail register. IOMMU is the consumer of commands produced by software into the CQ and controls the head register. The tail register holds the index into the queue where the next entry will be written by the producer. The head register holds the index into the queue where the consumer will read the next entry to process. A queue is empty if the head is equal to the tail. A queue is full if the tail is the head minus one. The head and tail wrap around when they reach the end of the circular buffer. The producer of data must ensure that the data written to a queue and the tail update are ordered such that the consumer that observes an update to the tail register must also observe all data produced into the queue between the offsets determined by the head and the tail. | | All RISC-V IOMMU implementations are required to support in-memory queues located in main memory. Supporting in-memory queues in I/O memory is not required but is not prohibited by this specification. The implication of the queue being considered full when tail is head minus one is that the effective size of the queue is one less than the number of entries in the queue. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#command-queue-cq)Command-Queue (CQ) Command queue is used by software to queue commands to be processed by the IOMMU. Each command is 16 bytes. The PPN of the base of this in-memory queue and the size of the queue is configured into a memory-mapped register called command-queue base (`cqb`). The tail of the command-queue resides in a software-controlled read/write memory-mapped register called command-queue tail (`cqt`). The `cqt` is an index into the next command queue entry that software will write. Subsequent to writing the command(s), software advances the `cqt` by the count of the number of commands written. The head of the command-queue resides in a read-only memory-mapped IOMMU controlled register called command-queue head (`cqh`). The `cqh` is an index into the command queue that IOMMU should process next. Subsequent to reading each command the IOMMU may advance the `cqh` by 1\. If `cqh` \== `cqt`, the command-queue is empty. If `cqt` \== (`cqh` \- 1) the command-queue is full. When an error bit or the `fence_w_ip` bit in `cqcsr` is 1, the command-queue interrupt pending (`cip`) bit is set in the `ipsr` if interrupts from command-queue are enabled (i.e. `cqcsr.cie` is 1). IOMMU commands are grouped into a major command group determined by the `opcode`and within each group the `func3` field specifies the function invoked by that command. The `opcode` defines the format of the operand fields. One or more of those fields may be used by the specific function invoked. The `opcode`encodings 64 to 127 are designated for custom use. ![Format of an IOMMU command](_images/diag-a734136ada374562d1c914cb1eed26f5dd980a3d.svg) Figure 2\. Format of an IOMMU command The commands are interpreted as two 64-bit doublewords. The byte order of each of the doublewords in memory, little-endian or big-endian, is the endianness as determined by `fctl.BE` ([iommu\_registers.adoc#FCTRL](iommu%5Fregisters.html#FCTRL)). The following command opcodes are defined: __Table 1\. IOMMU command opcodes__ | opcode | Encoding | Description | | -------- | -------- | ----------------------------------------------------------- | | IOTINVAL | 1 | IOMMU page-table cache invalidation commands. | | IOFENCE | 2 | IOMMU command-queue fence commands. | | IODIR | 3 | IOMMU directory cache invalidation commands. | | ATS | 4 | IOMMU PCIe \[[5](bibliography.html#bib-pci)\] ATS commands. | | Reserved | 5-63 | Reserved for future standard use. | | Custom | 64-127 | Designated for custom use. | All undefined functions of command opcodes 0 through 63 are reserved for future standard use. A command is determined to be illegal if it uses a reserved encoding or if a reserved bit is set to 1\. A command is unsupported if it is defined but not implemented as determined by the IOMMU `capabilities` register. If an illegal or unsupported command is fetched and decoded by the command-queue then the command-queue sets the `cqcsr.cmd_ill` bit and stops processing commands from the command-queue. To re-enable command processing software should clear the`cmd_ill` bit by writing 1 to it. #### [](#iommu-page-table-cache-invalidation-commands)IOMMU Page-Table cache invalidation commands ![Diagram](_images/diag-ae32292a3a7975698aecbf21cacf6467bcd0bfea.svg) IOMMU operations cause implicit reads to PDT, first-stage and second-stage page tables. To reduce latency of such reads, the IOMMU may cache entries from the first-stage and/or second-stage page tables in the IOMMU-address-translation-cache (IOATC). These caches might not observe modifications performed by software to these data structures in memory. The IOMMU translation-table cache invalidation commands, `IOTINVAL.VMA` and`IOTINVAL.GVMA` synchronize updates to in-memory first-stage and second-stage page table data structures respectively with the operation of the IOMMU and invalidate the matching IOATC entries. The `GV` operand indicates if the Guest-Soft-Context ID (`GSCID`) operand is valid. The `PSCV` operand indicates if the Process Soft-Context ID (`PSCID`) operand is valid. Setting `PSCV` to 1 is allowed only for `IOTINVAL.VMA`. The`AV` operand indicates if the address (`ADDR`) operand is valid. When `GV` is 0, the translations associated with the host (i.e. those where the second-stage is Bare) are operated on. When `GV` is 0, the `GSCID` operand is ignored. When `AV` is 0, the `ADDR` operand is ignored. When `PSCV` operand is 0, the`PSCID` operand is ignored. When the `AV` operand is set to 1, if the `ADDR`operand specifies an invalid address, the command may or may not perform any invalidations. The definition of the `NL` bit is provided by the non-leaf PTE invalidation extension [iommu\_extensions.adoc#NLINV](iommu%5Fextensions.html#NLINV). The definition of the `S` bit is provided by the address range invalidation extension [iommu\_extensions.adoc#ARINV](iommu%5Fextensions.html#ARINV). | | When an invalid address is specified, an implementation may either complete the command with no effect or may complete the command using an alternate, yetUNSPECIFIED, legal value for the address. Note that entries may generally be invalidated from the address translation cache at any time. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | `IOTINVAL.VMA` ensures that previous stores made to the first-stage page tables by the harts are observed by the IOMMU before all subsequent implicit reads from IOMMU to the corresponding first-stage page tables. __Table 2\. IOTINVAL.VMA operands and operations__ | GV | AV | PSCV | Operation | | -- | -- | ---- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | 0 | 0 | Invalidates all address-translation cache entries, including those that contain global mappings, for all host address spaces. | | 0 | 0 | 1 | Invalidates all address-translation cache entries for the host address space identified by PSCID operand, except for entries containing global mappings. | | 0 | 1 | 0 | Invalidates all address-translation cache entries that contain first-stage leaf page table entries, including those that contain global mappings, corresponding to the IOVA inADDR operand, for all host address spaces. | | 0 | 1 | 1 | Invalidates all address-translation cache entries that contain first-stage leaf page table entries corresponding to the IOVA in ADDR operand and that match the host address space identified by PSCID operand, except for entries containing global mappings. | | 1 | 0 | 0 | Invalidates all address-translation cache entries, including those that contain global mappings, for all VM address spaces associated with GSCID operand. | | 1 | 0 | 1 | Invalidates all address-translation cache entries for the VM address space identified by PSCID and GSCID operands, except for entries containing global mappings. | | 1 | 1 | 0 | Invalidates all address-translation cache entries that contain first-stage leaf page table entries, including those that contain global mappings, corresponding to the IOVA inADDR operand, for all VM address spaces associated with theGSCID operand. | | 1 | 1 | 1 | Invalidates all address-translation cache entries that contain first-stage leaf page table entries corresponding to the IOVA in ADDR operand, for the VM address space identified by PSCID and GSCID operands, except for entries containing global mappings. | `IOTINVAL.GVMA` ensures that previous stores made to the second-stage page tables are observed before all subsequent implicit reads from IOMMU to the corresponding second-stage page tables. Setting `PSCV` to 1 with `IOTINVAL.GVMA`is illegal. __Table 3\. IOTINVAL.GVMA operands and operations__ | GV | AV | Operation | | -- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | ignored | Invalidates information cached from any level of the second-stage page table, for all VM address spaces. | | 1 | 0 | Invalidates information cached from any level of the second-stage page tables, but only for VM address spaces identified by the GSCID operand. | | 1 | 1 | Invalidates information cached from leaf second-stage page table entries corresponding to the guest-physical-address inADDR operand, but only for VM address spaces identified by the GSCID operand. | | | Conceptually, an implementation might contain two address-translation caches: one that maps guest virtual addresses to guest physical addresses, and another that maps guest physical addresses to supervisor physical addresses.IOTINVAL.GVMA need not invalidate the former cache, but it must invalidate entries from the latter cache that match the IOTINVAL.GVMA address andGSCID operands. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | More commonly, implementations contain address-translation caches that map guest virtual addresses directly to supervisor physical addresses, removing a level of indirection. For such implementations, any entry whose guest virtual address maps to a guest physical address that matches the IOTINVAL.GVMAaddress and GSCID arguments must be invalidated. Selectively invalidating entries in this fashion requires tagging them with the guest physical address, which is costly, and so a common technique is to invalidate all entries that match the GSCID argument, regardless of the address argument. Simpler implementations may ignore the operand of IOTINVAL.VMA and/orIOTINVAL.GVMA and perform a global invalidation of all address-translation entries. Some implementations may cache an identity-mapped translation for the stage of address translation operating in Bare mode. Since these identity mappings are invariably correct, an explicit invalidation is unnecessary. A consequence of this specification is that an implementation may use any translation for an address that was valid at any time since the most recentIOTINVAL that subsumes that address. In particular, if a leaf PTE is modified but a subsuming IOTINVAL is not executed, either the old translation or the new translation will be used, but the choice is unpredictable. The behavior is otherwise well-defined. In a conventional TLB design, it is possible for multiple entries to match a single address if, for example, a page is upgraded to a larger page without first clearing the original non-leaf PTE’s valid bit and executing anIOTINVAL.VMA or IOTINVAL.GVMA as applicable with AV=0. In this case, a similar remark applies: it is unpredictable whether the old non-leaf PTE or the new leaf PTE is used, but the behavior is otherwise well defined. Another consequence of this specification is that it is generally unsafe to update a PTE using a set of stores of a width less than the width of the PTE, as it is legal for the implementation to read the PTE at any time, including when only some of the partial stores have taken effect. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#iommu-command-queue-fence-commands)IOMMU Command-queue Fence commands ![Diagram](_images/diag-311ba9da75bf724864e5f1da4e83b5e8d1c10e7f.svg) The IOMMU fetches commands from the CQ in order but the IOMMU may execute the fetched commands out of order. The IOMMU advancing `cqh` is not a guarantee that the commands fetched by the IOMMU have been executed or committed. A `IOFENCE.C` command completion, as determined by `cqh` advancing past the index of the `IOFENCE.C` command in the CQ, guarantees that all previous commands fetched from the CQ have been completed and committed. If the `IOFENCE.C` times out waiting on completion of previous commands that are specified to have a timeout, then the `cmd_to` bit in `cqcsr` [iommu\_registers.adoc#CSR](iommu%5Fregisters.html#CSR) is set to signal this condition. The `cqh` holds the index of the `IOFENCE.C` that timed out and all previous commands that are not specified to have a timeout have been completed and committed. | | In this version of the specification, only the ATS.INVAL command is specified to have a timeout. | | --------------------------------------------------------------------------------------------------- | The commands may be used to order memory accesses from I/O devices connected to the IOMMU as viewed by the IOMMU, other RISC-V harts, and external devices or co-processors. The `PR` bit, when set to 1, can be used to request that the IOMMU ensure that all previous read requests from devices that have already been processed by the IOMMU be committed to a global ordering point such that they can be observed by all RISC-V harts and IOMMUs in the system. The `PW` bit, when set to 1, can be used to request that the IOMMU ensure that all previous write requests from devices that have already been processed by the IOMMU be committed to a global ordering point such that they can be observed by all RISC-V harts and IOMMUs in the system. The wire-signaled-interrupts (`WSI`) bit when set to 1 causes a wired-interrupt from the command queue to be generated (by setting `cqcsr.fence_w_ip` \- [iommu\_registers.adoc#CSR](iommu%5Fregisters.html#CSR)) on completion of `IOFENCE.C`. This bit is reserved if the IOMMU does not support wired-interrupts or wired-interrupts have not been enabled (i.e., `fctl.WSI == 0`). | | Software should ensure that all previous read and writes processed by the IOMMU have been committed to a global ordering point before reclaiming memory that was previously made accessible to a device. A safe sequence for such memory reclamation is to first update the page tables to disallow access to the memory from the device and then use the IOTINVAL.VMA or IOTINVAL.GVMA appropriately to synchronize the IOMMU with the update to the page table. As part of the synchronization if the memory reclaimed was previously made read accessible to the device then request ordering of all previous reads; else if the memory reclaimed was previously made write accessible to the device then request ordering of all previous reads and writes. Ordering previous reads may be required if the reclaimed memory will be used to hold data that must not be made visible to the device. The IOFENCE.C with PR and/or PW set to 1 only ensures that requests that have been already processed by the IOMMU are committed to the global ordering point. Software must perform an interconnect-specific fence action if there is a need to ensure that all in-flight requests from a device that have not yet been processed by the IOMMU are observed. For PCIe, for example, a completion from device in response to a read from the device memory has the property of ensuring that previous posted writes are observed by the IOMMU as completions may not pass previous posted writes. The ordering guarantees are made for accesses to main-memory. For accesses to I/O memory, the ordering guarantees are implementation and I/O protocol defined. Simpler implementations may unconditionally order all previous memory accesses globally. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The `AV` command operand indicates if `ADDR[63:2]` and `DATA` operands are valid. If `AV`\=1, the IOMMU writes `DATA` to memory at a 4-byte aligned address`ADDR[63:2] * 4` as a 4-byte store when the command completes. When `AV` is 0, the `ADDR[63:2]` and `DATA` operands are ignored. If the attempt to perform this write encounters a memory fault, the `cmd_mf` bit in `cqcsr` [iommu\_registers.adoc#CSR](iommu%5Fregisters.html#CSR) is set to signal this condition, and the `cqh` holds the index of the `IOFENCE.C` that encountered such a memory fault and did not complete. | | Software may configure the ADDR\[63:2\] command operand to specify the address of the seteipnum\_le/seteipnum\_be register in an IMSIC to cause an external interrupt notification on IOFENCE.C completion. Alternatively, software may program ADDR\[63:2\] to a memory location and use IOFENCE.C to set a flag in memory indicating command completion. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#iommu-directory-cache-invalidation-commands)IOMMU directory cache invalidation commands ![Diagram](_images/diag-2907d0b1d5506609160eb00d0a58c0a8273ad9f6.svg) IOMMU operations cause implicit reads to DDT and/or PDT. To reduce latency of such reads, the IOMMU may cache entries from the DDT and/or PDT in IOMMU directory caches. These caches might not observe modifications performed by software to these data structures in memory. The IOMMU DDT cache invalidation command, `IODIR.INVAL_DDT`, synchronizes updates to DDT with the operation of the IOMMU and flushes the matching cached entries. The IOMMU PDT cache invalidation command, `IODIR.INVAL_PDT`, synchronizes updates to PDT with the operation of the IOMMU and flushes the matching cached entries. The `DV` operand indicates if the device ID (`DID`) operand is valid. The `DV`operand must be 1 for `IODIR.INVAL_PDT` else the command is illegal. When `DV`operand is 1, the value of the `DID` operand must not be wider than that supported by the `ddtp.iommu_mode`. `IODIR.INVAL_DDT` guarantees that any previous stores made by a RISC-V hart to the DDT are observed before all subsequent implicit reads from IOMMU to DDT. If `DV` is 0, then the command invalidates all DDT and PDT entries cached for all devices; the `DID` operand is ignored. If `DV` is 1, then the command invalidates cached leaf-level DDT entry for the device identified by `DID`operand and all associated PDT entries. The `PID` operand is reserved for the`IODIR.INVAL_DDT` command. `IODIR.INVAL_PDT` guarantees that any previous stores made by a RISC-V hart to the PDT are observed before all subsequent implicit reads from IOMMU to PDT. The command invalidates cached leaf PDT entry for the specified `PID` and `DID`. The `PID` operand of `IODIR.INVAL_PDT` must not be wider than the width supported by the IOMMU (see [iommu\_registers.adoc#CAP](iommu%5Fregisters.html#CAP)). | | Some fields in the Device-context or Process-context may be guest-physical addresses. An implementation when caching the device-context or process-context may cache these fields after translating them to a supervisor physical address. Other implementations may cache them as guest-physical addresses and translate them to supervisor physical addresses using a second-stage page table just prior to accessing memory referenced by these addresses. If second-stage page tables used for these translations are modified, software must issue the appropriate IODIR command as some implementations may choose to cache the translated supervisor physical address pointer in the IOMMU directory caches. The IOTINVAL command has no effect on the IOMMU directory caches. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#iommu-pcie-ats-commands)IOMMU PCIe ATS commands This command is supported if `capabilities.ATS` is set to 1. ![Diagram](_images/diag-e302886e4607b659e6977a58bb1981e4de087c04.svg) The `ATS.INVAL` command instructs the IOMMU to send an “Invalidation Request” message to the PCIe device function identified by `RID`. An “Invalidation Request” message is used to clear a specific subset of the address range from the address translation cache in a device function. The`ATS.INVAL` command completes when an “Invalidation Completion” response message is received from the device or a protocol-defined timeout occurs while waiting for a response. The IOMMU may advance the `cqh` and fetch more commands from CQ while a response is awaited. If a timeout occurs, it is reported when a subsequent `IOFENCE.C` command is executed. | | Software that needs to know if the invalidation operation completed on the device may use the IOMMU command-queue fence command (IOFENCE.C) to wait for the responses to all prior “Invalidation Request” messages. The IOFENCE.C is guaranteed to not complete before all previously fetched commands were executed and completed. A previously fetched ATS command to invalidate device ATC does not complete until either the request times out or a valid response is received from the device. If one or more ATS invalidation commands preceding the IOFENCE.C have timed out, then software may make the CQ operational again and resubmit the invalidation commands that may have timed out. If the ATS.INVAL commands queued before the IOFENCE.C were directed at multiple devices then software may resubmit these commands as ATS.INVAL and IOFENCE.C pairs to identify the device that caused the timeout. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `ATS.PRGR` command instructs the IOMMU to send a “Page Request Group Response” message to the PCIe device function identified by the `RID`. The “Page Request Group Response” message is used by system hardware and/or software to communicate with the device functions page-request interface to signal completion of a “Page Request”, or the catastrophic failure of the interface. If the `PV` operand is set to 1, the message is generated with a PASID with the PASID field set to the `PID` operand. if `PV` operand is set to 0, then the`PID` operand is ignored and the message is generated without a PASID. The `PAYLOAD` operand of the command is used to form the message body and its fields are as specified by the PCIe specification \[[5](bibliography.html#bib-pci)\]. The `PAYLOAD` field is formatted as follows: ![`PAYLOAD` of an `ATS.INVAL` command](_images/diag-393efad3a28f0a1ce114d1a6d3d004540c471442.svg) Figure 3\. `PAYLOAD` of an `ATS.INVAL` command ![`PAYLOAD` of an `ATS.PRGR` command](_images/diag-1b672a5420e8c405611e5b7688995f3f3d73656f.svg) Figure 4\. `PAYLOAD` of an `ATS.PRGR` command If the `DSV` operand is 1, then a valid destination segment number is specified by the `DSEG` operand. If the `DSV` operand is 0, then the `DSEG` operand is ignored. | | A Hierarchy is a PCI Express I/O interconnect topology, wherein the Configuration Space addresses, referred to as the tuple of Bus/Device/Function Numbers, are unique. In some contexts, a Hierarchy is also called a Segment, and in Flit Mode, the Segment number is sometimes included in the ID of a Function. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#FAULT%5FQUEUE)Fault/Event-Queue (`FQ`) Fault/Event queue is an in-memory queue data structure used to report events and faults raised when processing transactions. Each fault record is 32 bytes. The PPN of the base of this in-memory queue and the size of the queue is configured into a memory-mapped register called fault-queue base (`fqb`). The tail of the fault-queue resides in an IOMMU controlled read-only memory-mapped register called `fqt`. The `fqt` is an index into the next fault record that IOMMU will write in the fault-queue. Subsequent to writing the record, the IOMMU advances the `fqt` by 1\. The head of the fault-queue resides in a read/write memory-mapped software controlled register called `fqh`. The `fqh`is an index into the fault record that SW should process next. Subsequent to processing fault record(s) software advances the `fqh` by the count of the number of fault records processed. If `fqh` \== `fqt`, the fault-queue is empty. If`fqt` \== (`fqh` \- 1) the fault-queue is full. The fault records are interpreted as four 64-bit doublewords. The byte order of each of the doublewords in memory, little-endian or big-endian, is the endianness as determined by `fctl.BE` ([iommu\_registers.adoc#FCTRL](iommu%5Fregisters.html#FCTRL)). ![Fault-queue record](_images/diag-fdc51ddd32406d6b5baf1f14083b2d48a13cbd86.svg) Figure 5\. Fault-queue record The `CAUSE` is a code indicating the cause of the fault/event. __Table 4\. Fault record CAUSE field encodings__ | CAUSE | Description | Reported if DTF is 1? | | ----- | ------------------------------------- | --------------------- | | 1 | Instruction access fault | No | | 4 | Read address misaligned | No | | 5 | Read access fault | No | | 6 | Write/AMO address misaligned | No | | 7 | Write/AMO access fault | No | | 12 | Instruction page fault | No | | 13 | Read page fault | No | | 15 | Write/AMO page fault | No | | 20 | Instruction guest page fault | No | | 21 | Read guest-page fault | No | | 23 | Write/AMO guest-page fault | No | | 256 | All inbound transactions disallowed | Yes | | 257 | DDT entry load access fault | Yes | | 258 | DDT entry not valid | Yes | | 259 | DDT entry misconfigured | Yes | | 260 | Transaction type disallowed | No | | 261 | MSI PTE load access fault | No | | 262 | MSI PTE not valid | No | | 263 | MSI PTE misconfigured | No | | 264 | MRIF access fault | No | | 265 | PDT entry load access fault | No | | 266 | PDT entry not valid | No | | 267 | PDT entry misconfigured | No | | 268 | DDT data corruption | Yes | | 269 | PDT data corruption | No | | 270 | MSI PT data corruption | No | | 271 | MSI MRIF data corruption | No | | 272 | Internal data path error | Yes | | 273 | IOMMU MSI write access fault | Yes | | 274 | First/second-stage PT data corruption | No | The `CAUSE` encodings 275 through 2047 are reserved for future standard use and the encodings 2048 through 4095 are designated for custom use. Encodings between 0 and 275 that are not specified in [Table 4](#FAULT%5FCAUSE) are reserved for future standard use. If a fault condition prevents locating a valid device context then the `DTF`value assumed for reporting such faults is 0. The `TTYP` field reports inbound transaction type. __Table 5\. Fault record TTYP field encodings__ | TTYP | Description | | ------- | ------------------------------------------------- | | 0 | None. Fault not caused by an inbound transaction. | | 1 | Untranslated read for execute transaction | | 2 | Untranslated read transaction | | 3 | Untranslated write/AMO transaction | | 4 | Reserved | | 5 | Translated read for execute transaction | | 6 | Translated read transaction | | 7 | Translated write/AMO transaction | | 8 | PCIe ATS Translation Request | | 9 | PCIe Message Request | | 10 - 31 | Reserved | | 31 - 63 | Designated for custom use | If the `TTYP` is a transaction with an IOVA, the IOVA is reported in `iotval`. If the `TTYP` is a PCIe message request, the message code of the PCIe message is reported in `iotval`. If `TTYP` is 0, the values reported in `iotval` and`iotval2` fields are as defined by the `CAUSE`. | | The IOVA is partitioned into a virtual page number (VPN) and page offset. Whereas the VPN is translated into a physical page number (PPN) by the address translation process, the page offset is not required for this process. The IO bridge in some implementations may not provide the page offset part of theIOVA to the IOMMU and the IOMMU may report the page offset in iotval as 0\. Likewise, an IOMMU may report the page offset of a GPA in iotval2 as 0. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | `DID` holds the `device_id` of the transaction. If `PV` is 0, then `PID` and`PRIV` are 0\. If `PV` is 1, the `PID` holds a `process_id` of the transaction and if the privilege of the transaction was Supervisor then the `PRIV` bit is 1 else it’s 0\. The `DID`, `PV`, `PID`, and `PRIV` fields are 0 if `TTYP` is 0. If the `CAUSE` is a guest-page fault then bits 63:2 of the zero-extended guest-physical-address are reported in `iotval2[63:2]`. If bit 0 of `iotval2` is 1, then the guest-page-fault was caused by an implicit memory access for first-stage address translation. If bit 0 of `iotval2` is 1, and the implicit access was a write then bit 1 of `iotval2` is set to 1 else it is set to 0. | | The bit 1 of iotval2 is set for the case where the implementation supports hardware updating of A/D bits and the implicit memory access was attempted to automatically update A and/or D in first-stage page tables. All other implicit memory accesses for first-stage address translation will be reads. If the hardware updating of A/D bits is not implemented, the _write_ case will never arise. When the second-stage is not Bare, the memory accesses for reading PDT entries to locate the Process-context are implicit memory accesses for first-stage address translation. If a guest-page fault was caused by implicit memory access to read PDT entries, then bit 0 of iotval2 is reported as 1 and bit 1 as 0. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The IOMMU may be unable to report faults through the fault-queue due to error conditions such as the fault-queue being full or the IOMMU encountering access faults when attempting to access the queue memory. A memory-mapped fault control and status register (`fqcsr`) holds information about such faults. If the fault-queue full condition is detected, the IOMMU sets the fault-queue overflow (`fqof`) bit in fqcsr. If the IOMMU encounters a fault in accessing the fault-queue memory, the IOMMU sets the fault-queue memory access fault (`fqmf`) bit in `fqcsr`. While either error bit is set in `fqcsr`, the IOMMU discards the record that led to the fault and all further fault records. When an error bit in `fqcsr` is 1 or when a new fault record is produced in the fault-queue, the fault interrupt pending (`fip`) bit is set in `ipsr` if interrupts from the fault-queue are enabled i.e. `fqcsr.fie` is 1. The IOMMU may identify multiple requests as having detected an identical fault. In such cases the IOMMU may report each of those faults individually, or report the fault for a subset, including one, of requests. ### [](#PRQ)Page-Request-Queue (`PQ`) Page-request queue is an in-memory queue data structure used to report PCIe ATS “Page Request” and "Stop Marker" messages \[[5](bibliography.html#bib-pci)\] to software. The base PPN of this in-memory queue and the size of the queue is configured into a memory-mapped register called page-request queue base (`pqb`). Each Page-Request record is 16 bytes. The tail of the queue resides in an IOMMU controlled read-only memory-mapped register called `pqt`. The `pqt` holds an index into the queue where the next page-request message will be written by the IOMMU. Subsequent to writing the message, the IOMMU advances the `pqt` by 1. The head of the queue resides in a software controlled read/write memory-mapped register called `pqh`. The `pqh` holds an index into the queue where the next page-request message will be received by software. Subsequent to processing the message(s) software advances the `pqh` by the count of the number of messages processed. If `pqh` \== `pqt`, the page-request queue is empty. If `pqt` \== (`pqh` \- 1) the page-request queue is full. The IOMMU may be unable to report "Page Request" messages through the queue due to error conditions such as the queue being disabled, queue being full, or the IOMMU encountering access faults when attempting to access queue memory. A memory-mapped page-request queue control and status register (`pqcsr`) is used to hold information about such faults. On a page queue full condition the page-request-queue overflow (`pqof`) bit is set in `pqcsr`. If the IOMMU encountered a fault in accessing the queue memory, the page-request-queue memory access fault (`pqmf`) bit is set in `pqcsr`. While either error bit is set in`pqcsr`, the IOMMU discards all subsequent "Page Request" messages, including the message that caused the error bits to be set. "Page request" messages that do not require a response, i.e. those with the "Last Request in PRG" field is 0, are silently discarded. "Page request" messages that require a response, i.e. those with "Last Request in PRG" field set to 1 and are not "Stop Marker" messages, may be auto-completed by an IOMMU generated “Page Request Group Response” message as specified in [iommu\_data\_structures.adoc#ATS\_PRI](iommu%5Fdata%5Fstructures.html#ATS%5FPRI). When an error bit in `pqcsr` is 1 or when a new message is produced in the queue, the page-request-queue interrupt pending (`pip`) bit is set in the `ipsr` if interrupts from page-request-queue are enabled i.e. `pqcsr.pie` is 1. ![Page-request-queue record](_images/diag-0f8bf9395990798aba8332ee1c726a064c8c3c98.svg) Figure 6\. Page-request-queue record The `DID` field holds the requester ID from the message. The `PID` field is valid if `PV` is 1 and reports the PASID from message. `PRIV` is set to 0 if the message did not have a PASID, otherwise it holds the “Privilege Mode Requested” bit from the TLP. The `EXEC` bit is set to 0 if the message did not have a PASID, otherwise it reports the “Execute Requested” bit from the TLP. All other fields are set to 0\. The payload of the “Page Request” message (bytes 0x08 through 0x0F of the message) is held in the `PAYLOAD` field. If `R` and `W` are both 0 and`L` is 1, the message is "Stop Marker". The page-request-queue records are interpreted as two 64-bit doublewords. The byte order of each of the doublewords in memory, little-endian or big-endian, is the endianness as determined by `fctl.BE` ([iommu\_registers.adoc#FCTRL](iommu%5Fregisters.html#FCTRL)). endianness as determined by `fctl.BE` ([iommu\_registers.adoc#FCTRL](iommu%5Fregisters.html#FCTRL)). The `PAYLOAD` holds the message body and its fields are as specified by the PCIe specification \[[5](bibliography.html#bib-pci)\]. The `PAYLOAD` field is formatted as follows: ![`PAYLOAD` of a "Page request" message](_images/diag-80c3ce8c9e35e08c414865a7621dd8c9a302f4c8.svg) Figure 7\. `PAYLOAD` of a "Page request" message Introduction ==================== ## [](#intro)Introduction The Input-Output Memory Management Unit (IOMMU), sometimes referred to as a System MMU (SMMU), is a system-level Memory Management Unit (MMU) that connects direct-memory-access-capable Input/Output (I/O) devices to system memory. For each I/O device connected to the system through an IOMMU, software can configure at the IOMMU a device context, which associates with the device a specific virtual address space and other per-device parameters. By giving each device its own separate device context at an IOMMU, each device can be individually configured for a separate operating system, which may be a guest OS or the main (host) OS. On every memory access initiated by a device, the IOMMU identifies the originating device by some form of unique device identifier, which the IOMMU then uses to locate the appropriate device context within data structures supplied by software. For PCIe \[[5](bibliography.html#bib-pci)\], for example, the originating device may be identified by the unique 16-bit triplet of PCI bus number (8-bit), device number (5-bit), and function number (3-bit) (collectively known as routing identifier or RID) and optionally up to 8-bit segment number when the IOMMU supports multiple Hierarchies. This specification refers to such unique device identifier as `device_id` and supports up to 24-bit wide identifiers. | | A Hierarchy is a PCI Express I/O interconnect topology, wherein the Configuration Space addresses, referred to as the tuple of Bus/Device/Function Numbers, are unique. In some contexts, a Hierarchy is also called a Segment, and in Flit Mode, the Segment number is sometimes included in the ID of a Function. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Some devices may support shared virtual addressing which is the ability to share process address spaces with devices. Sharing process address spaces with devices allows to rely on core kernel memory management for DMA, removing some complexity from application and device drivers. After binding to a device, applications can instruct it to perform DMA on statically or dynamically allocated buffers. To support such addressing, software can configure one or more process contexts into the device context. Every memory access initiated by such a device is accompanied by a unique process identifier, which the IOMMU uses in conjunction with the unique device identifier to locate the appropriate process context configured by software in the device context. For PCIe, for example, the process context may be identified by the unique 20-bit process address space identifier (PASID). This specification refers to such unique process identifiers as `process_id` and supports up to 20-bit wide identifiers. The IOMMU employs a two-stage address translation process to translate the IOVA to an SPA and to enforce memory protections for the DMA. To perform address translation and memory protection the IOMMU uses same page table formats as used by the CPU’s MMU for the first-stage and second-stage address translation. Using the same page table formats as the CPU’s MMU removes some of the memory management complexity for DMA. Use of an identical format also allows the same page tables to be used simultaneously by both the CPU MMU and the IOMMU. Although there is no option to disable two-stage address translation, either stage may be effectively disabled by configuring the virtual memory scheme for that stage to be `Bare` i.e. perform no address translation or memory protection. The virtual memory scheme employed by the IOMMU may be configured individually per device in the IOMMU. Devices perform DMA using an I/O virtual address (IOVA). Depending on the virtual memory scheme selected for a device, the IOVA used by the device may be a supervisor physical address (SPA), guest physical address (GPA), or a virtual address (VA). If the virtual memory scheme selected for both stages is `Bare` then the IOVA is a SPA. There is no address translation or protection performed by the IOMMU. If the virtual memory scheme selected for first-stage is `Bare` but the scheme for the second-stage is not `Bare` then the IOVA is a GPA. The first-stage is effectively disabled. The second-stage translates the GPA to SPA and enforces the configured memory protections. Such a configuration would be typically employed when the device control is passed through to a virtual machine but the Guest OS in the VM does not use first-stage address translation to further constrain memory accesses from such devices. Comparing to a RISC-V hart, this configuration is analogous to two-stage address translation being in effect on a RISC-V hart with the G-stage active and the VS-stage set to Bare. If the virtual memory scheme selected for first-stage is not `Bare` but the scheme for the second-stage is `Bare` then IOVA is a VA. The second-stage is effectively disabled. The first-stage translates the VA to a SPA and enforces the configured memory protections. This configuration would be typically employed when the IOMMU is used by a native OS or when the control of the device is retained by the hypervisor itself. Comparing to a RISC-V hart, this configuration is analogous to single-stage address translation being in effect on a RISC-V hart. If the virtual memory scheme selected for neither stage is `Bare` then the IOVA is a VA. Two-stage address translation is in effect. The first-stage translates the VA to a GPA and the second-stage translates the GPA to a SPA. Each stage enforces the configured memory protections. Such a configuration would be typically be employed when the device control is passed-through to a virtual machine and the Guest OS in the VM uses the first-stage address translation to further constrain the memory accessed by such devices and associated privileges and memory protections. Comparing to a RISC-V hart, this configuration is analogous to two-stage address translation being in effect on a RISC-V hart with both G-stage and VS-stage active (not Bare). DMA address translation in the IOMMU has certain performance implications for DMA accesses as the access time may be lengthened by the time required to determine the SPA using the software provided data structures. Similar overheads in the CPU MMU are mitigated typically through the use of a translation look-aside buffer (TLB) to cache these address translations such that they may be re-used to reduce the translation overhead on subsequent accesses. The IOMMU may employ similar address translation caches, referred as IOMMU Address Translation Cache (IOATC). The IOMMU provides mechanisms for software to synchronize the IOATC with the memory resident data structures used for address translation when they are modified. Software may configure the device context with a software defined context identifier called guest soft-context identifier (`GSCID`) to indicate that a collection of devices are assigned to the same VM and thus access a common virtual address space. Software may configure the process context with a software defined context identifier called process soft-context identifier (`PSCID`) to identify a collection of processes that share a common virtual address space. The IOMMU may use the `GSCID` and `PSCID` to tag entries in the IOATC to avoid duplication and simplify invalidation operations. Some devices may participate in the translation process and provide a device side ATC (DevATC) for its own memory accesses. By providing a DevATC, the device shares the translation caching responsibility and thereby reduce probability of "thrashing" in the IOATC. The DevATC may be sized by the device to suit its unique performance requirements and may also be used by the device to optimize DMA latency by prefetching translations. Such mechanisms require close cooperation of the device and the IOMMU using a protocol. For PCIe, for example, the Address Translation Services (ATS) protocol may be used by the device to request translations to cache in the DevATC and to synchronize it with updates made by software address translation data structures. The device participating in the address translation process also enables the use of I/O page faults to avoid the core kernel memory manager from having to make all physical memory that may be accessed by the device resident at all times. For PCIe, for example, the device may implement the Page Request Interface (PRI) to dynamically request the memory manager to make a page resident if it discovers the page for which it requested a translation was not available. An IOMMU may support specialized software interfaces and protocols with the device to enable services such as PCIe ATS and PCIe PRI \[[5](bibliography.html#bib-pci)\]. In systems built with an Incoming Message-Signaled Interrupt Controller (IMSIC), the IOMMU may be programmed by the hypervisor to direct message-signaled interrupts (MSI) from devices controlled by the guest OS to a guest interrupt file in an IMSIC. Because MSIs from devices are simply memory writes, they would naturally be subject to the same address translation that an IOMMU applies to other memory writes. However, the RISC-V Advanced Interrupt Architecture \[[6](bibliography.html#bib-aia)\] requires that IOMMUs treat MSIs directed to virtual machines specially, in part to simplify software, and in part to allow optional support for memory-resident interrupt files. The device context is configured by software with parameters to identify memory accesses to a virtual interrupt file and to be translated using a MSI address translation table configured by software in the device context. ### [](#glossary)Glossary __Table 1\. Terms and definitions__ | Term | Definition | | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | AIA | RISC-V Advanced Interrupt Architecture \[[6](bibliography.html#bib-aia)\]. | | ATS / PCIe ATS | Address Translation Services: A PCIe protocol to support DevATC \[[5](bibliography.html#bib-pci)\]. | | CXL | Compute Express Link bus standard. | | DC / Device Context | A hardware representation of state that identifies a device and the VM to which the device is assigned. | | DDT | Device-directory-table: A radix-tree structure traversed using the unique device identifier to locate the Device Context structure. | | DDI | Device-directory-index: A sub-field of the unique device identifier used as a index into a leaf or non-leaf DDT structure. | | Device ID | An identification number that is up to 24-bits to identify the source of a DMA or interrupt request. For PCIe devices this is the routing identifier (RID) \[[5](bibliography.html#bib-pci)\]. | | DevATC | An address translation cache at the device. | | DMA | Direct Memory Access. | | GPA | Guest Physical Address: An address in the virtualized physical memory space of a virtual machine. | | GSCID | Guest soft-context identifier: An identification number used by software to uniquely identify a collection of devices assigned to a virtual machine. An IOMMU may tag IOATC entries with the GSCID. Device contexts programmed with the same GSCID must also be programmed with identical second-stage page tables. | | Guest | Software in a virtual machine. | | HPM | Hardware Performance Monitor. | | Hypervisor | Software entity that controls virtualization. | | ID | Identifier. | | IMSIC | Incoming Message-signaled Interrupt Controller. | | IOATC | IOMMU Address Translation Cache: cache in IOMMU that caches data structures used for address translations. | | IOVA | I/O Virtual Address: Virtual address for DMA by devices. | | MSI | Message Signaled Interrupts. | | OS | Operating System. | | PASID | Process Address Space Identifier: It identifies the address space of a process. The PASID value is provided in the PASID TLP prefix of the request. | | PBMT | Page-Based Memory Types. | | PC | Process Context. | | PCIe | Peripheral Component Interconnect Express bus standard \[[5](bibliography.html#bib-pci)\]. | | PDI | Process-directory-index: a sub field of the unique process identifier used to index into a leaf or non-leaf PDT structure. | | PDT | Process-directory-table: A radix tree data structure traversed using the unique Process identifier to locate the process context structure. | | PMA | Physical Memory Attributes. | | PMP | Physical Memory Protection. | | PPN | Physical Page Number. | | PRI | Page Request Interface - a PCIe protocol \[[5](bibliography.html#bib-pci)\] that enables devices to request OS memory manager services to make pages resident. | | Process ID | An identification number that is up to 20-bits to identify a process context. For PCIe devices this is the PASID \[[5](bibliography.html#bib-pci)\]. | | PSCID | Process soft-context identifier: An identification number used by software to identify a unique address space. The IOMMU may tag IOATC entries with PSCID. | | PT | Page Table. | | PTE | Page Table Entry. A leaf or non-leaf entry in a page table. | | Reserved | A register or data structure field reserved for future use. Reserved fields in data structures must be set to 0 by software. Software must ignore reserved fields in registers and preserve the value held in these fields when writing values to other fields in the same register. | | RID / PCIe RID | PCIe routing identifier \[[5](bibliography.html#bib-pci)\]. | | RO | Read-only - Register bits are read-only and cannot be altered by software. Where explicitly defined, these bits are used to reflect changing hardware state, and as a result bit values can be observed to change at run time. If the optional feature that would Set the bits is not implemented, the bits must be hardwired to Zero | | RW | Read-Write - Register bits are read-write and are permitted to be either Set or Cleared by software to the desired state. If the optional feature that is associated with the bits is not implemented, the bits are permitted to be hardwired to Zero. | | RW1C | Write-1-to-clear status - Register bits indicate status when read. A Set bit indicates a status event which is Cleared by writing a 1b. Writing a 0b to RW1C bits has no effect. If the optional feature that would Set the bit is not implemented, the bit must be read-only and hardwired to Zero | | RW1S | Read-Write-1-to-set - register bits indicate status when read. The bit may be Set by writing 1b. Writing a 0b to RW1S bits has no effect. If the optional feature that introduces the bit is not implemented, the bit must be read-only and hardwired to Zero | | SOC | System on a chip, also referred as system-on-a-chip and system-on-chip. | | SPA | Supervisor Physical Address: Physical address used to to access memory and memory-mapped resources. | | TLB | Translation Lookaside Buffer. A cache that stores virtual-to-physical address translations to reduce translation latency. On a TLB hit, the translation is completed without accessing the first-stage and/or second-stage page tables. On a TLB miss, a page table walk is performed. Some implementations cache non-leaf levels of the page tables, reducing the number of walks required. | | TLP | Transaction Layer Packet. | | VA | Virtual Address. | | VM | Virtual Machine: An efficient, isolated duplicate of a real computer system. In this specification it refers to the collection of resources and state that is accessible when a RISC-V hart supporting the hypervisor extension executes with the virtualization mode set to 1. | | VMM | Virtual Machine Monitor. Also referred to as hypervisor. | | VS | Virtual Supervisor: Supervisor privilege in virtualization mode. | | Walk | A single memory access by the IOMMU to load a table entry. Each entry load—​leaf or non-leaf—​is one walk. The number of walks required to load the leaf entry depends on the number of table levels and may be fewer if non-leaf levels are already cached. | | WARL | Write Any values, Reads Legal values: Attribute of a register field that is only defined for a subset of bit encodings, but allow any value to be written while guaranteeing to return a legal value whenever read. | | WPRI | Writes Preserve values, Reads Ignore values: Attribute of a register field that is reserved for future use. | ### [](#usage-models)Usage models #### [](#non-virtualized-os)Non-virtualized OS A non-virtualized OS may use the IOMMU for the following significant system-level functionalities: 1. Protect the operating system from bad memory accesses from errant devices 2. Support 32-bit devices in 64-bit environment (avoidance of bounce buffers) 3. Support mapping of contiguous virtual addresses to an underlying fragmented physical addresses (avoidance of scatter/gather lists) 4. Support shared virtual addressing In the absence of an IOMMU a device could access any memory, such as privileged memory, and cause malicious or unintended corruptions. This may be due to hardware bugs, device driver bugs, or due to malicious software/hardware. The IOMMU offers a mechanism for the OS to defend against such unintended corruptions by limiting the memory that can be accessed by devices. As depicted in [Figure 1](#fig:device-isolation) the OS may configure the IOMMU with a page table to translate the IOVA and thereby limit the addresses that may be accessed to those allowed by the page table. ![non virt OS](_images/non-virt-OS.svg) Figure 1\. Device isolation in non-virtualized OS Legacy 32-bit devices cannot access the memory above 4 GiB. The IOMMU, through its address remapping capability, offers a simple mechanism for the device to directly access any address in the system (with appropriate access permission). Without an IOMMU, the OS must resort to copying data through buffers (also known as bounce buffers) allocated in memory below 4 GiB. In this scenario the IOMMU improves the system performance. The IOMMU can be useful to perform scatter/gather DMA as it permits to allocate large regions of memory for I/O without the need for all of the memory to be contiguous. A contiguous virtual address range can map to such fragmented physical addresses and the device programmed with the virtual address range. The IOMMU can be used to support shared virtual addressing which is the ability to share a process address space with devices. The virtual addresses used for DMA are then translated by the IOMMU to an SPA. When the IOMMU is used by a non-virtualized OS, the first-stage suffices to provide the required address translation and protection function and the second-stage may be set to Bare. #### [](#hypervisor)Hypervisor IOMMU makes it possible for a guest operating system, running in a virtual machine, to be given direct control of an I/O device with only minimal hypervisor intervention. A guest OS with direct control of a device will program the device with guest physical addresses, because that is all the OS knows. When the device then performs memory accesses using those addresses, an IOMMU is responsible for translating those guest physical addresses into supervisor physical addresses, referencing address-translation data structures supplied by the hypervisor. [Figure 2](#fig:dma-translation-direct-device-assignment) illustrates the concept. The device D1 is directly assigned to VM-1 and device D2 is directly assigned to VM-2\. The VMM configures a second-stage page table to be used for each device and restricts the memory that can be accessed by D1 to VM-1 associated memory and from D2 to VM-2 associated memory. ![hypervisor](_images/hypervisor.svg) Figure 2\. DMA translation to enable direct device assignment To handle MSIs from a device controlled by a guest OS, the hypervisor configures an IOMMU to redirect those MSIs to a guest interrupt file in an IMSIC (see [Figure 3](#MSI%5FREDIR)) or to a memory-resident interrupt file. The IOMMU is responsible to use the MSI address-translation data structures supplied by the hypervisor to perform the MSI redirection. Because every interrupt file, real or virtual, occupies a naturally aligned 4-KiB page of address space, the required address translation is from a virtual (guest) page address to a physical page address, the same as supported by regular RISC-V page-based address translation. ![msi imsic](_images/msi-imsic.svg) Figure 3\. MSI address translation to direct guest programmed MSI to IMSIC guest interrupt files #### [](#guest-os)Guest OS The hypervisor may provide a virtual IOMMU facility, through hardware emulation or by enlightening the guest OS to use a software interface with the Hypervisor (also known as para-virtualization). The guest OS may then use the facilities provided by the virtual IOMMU to avail the same benefits as those discussed for a non-virtualized OS through the use of a first-stage page table that it controls. The hypervisor establishes a second-stage page table that it controls to virtualize the address space for the virtual machine and to contain memory accesses from the devices passed through to the VM to the memory associated with the VM. With two-stage address translations active, the IOVA is first translated to a GPA using the first-stage page tables managed by the guest OS and the GPA translated to a SPA using the second-stage page tables managed by the hypervisor. [Figure 4](#fig:iommu-for-guest-os) illustrates the concept. ![guest OS](_images/guest-OS.svg) Figure 4\. Address translation in IOMMU for Guest OS The IOMMU is configured to perform address translation using a first-stage and second-stage page table for device D1\. The second-stage is typically used by the hypervisor to translate GPA to SPA and limit the device D1 to memory associated with VM-1\. The first-stage is typically configured by the Guest OS to translate a VA to a GPA and contain device D1 access to a subset of VM-1 memory. For device D2 only the second-stage is active and the first-stage is set to Bare. The host OS or hypervisor may also retain a device, such as D3, for its own use. The first-stage suffices to provide the required address translation and protection function for device D3 and the second-stage is set to Bare. ### [](#placement-and-data-flow)Placement and data flow [Figure 5](#fig:example-soc-with-iommu) shows an example of a typical system on a chip (SOC) with RISC-V hart(s). The SOC incorporates memory controllers and several IO devices. This SOC also incorporates two instances of the IOMMU. A device may be directly connected to the IO Bridge and the system interconnect or may be connected through a Root Port when a IO protocol transaction to system interconnect transaction translation is required. In case of PCIe \[[5](bibliography.html#bib-pci)\], for example, the Root Port is a PCIe port that maps a portion of a hierarchy through an associated virtual PCI-PCI bridge and maps the PCIe IO protocol transactions to the system interconnect transactions. The first IOMMU instance, IOMMU 0 (associated with the IO Bridge 0), interfaces a Root Port to the system fabric/interconnect. One or more endpoint devices are interfaced to the SoC through this Root Port. In the case of PCIe, the Root Port incorporates an ATS interface to the IOMMU that is used to support the PCIe ATS protocol by the IOMMU. The example shows an endpoint device with a device side ATC (DevATC) that holds translations obtained by the device from IOMMU 0 using the PCIe ATS protocol \[[5](bibliography.html#bib-pci)\]. When such IO-protocol-to-system-fabric-protocol translation using a Root Port is not required, the devices may interface directly with the system fabric. The second IOMMU instance, IOMMU 1 (associated with the IO Bridge 1), illustrates interfacing devices (IO Devices A and B) to the system fabric without the use of a Root Port. The IO Bridge is placed between the device(s) and the system interconnect to process DMA transactions. IO Devices may perform DMA transactions using IO Virtual Addresses (VA, GVA or GPA). The IO Bridge invokes the associated IOMMU to translate the IOVA to a Supervisor Physical Addresses (SPA). The IOMMU is not invoked for outbound transactions. ![placement](_images/placement.svg) Figure 5\. Example of IOMMUs integration in SoC. The IOMMU is invoked by the IO Bridge for address translation and protection for inbound transactions. The data associated with the inbound transactions is not processed by the IOMMU. The IOMMU behaves like a look-aside IP to the IO Bridge and has several interfaces (see [Figure 6](#fig:iommu-interfaces)): * Host interface: it is an interface to the IOMMU for the harts to access its memory-mapped registers and perform global configuration and/or maintenance operations. * Device Translation Request interface: it is an interface, which receives the translation requests from the IO Bridge. On this interface the IO Bridge provides information about the request such as: 1. The hardware identities associated with transaction - the `device_id` and if applicable the `process_id` and its validity. The IOMMU uses the hardware identities to retrieve the context information to perform the requested address translations. 2. The IOVA and the type of the transaction (Translated or Untranslated). 3. Whether the request is for a read, write, execute, or an atomic operation. 1. Execute requested must be explicitly associated with the request (e.g., using a PCIe PASID). When not explicitly requested, the default must be 0. 4. The privilege mode associated with the request. When a privilege mode is not explicitly associated with the request (e.g., using a PCIe PASID), the default privilege mode must be User. For requests without a `process_id` the privilege mode must be User. 5. The number of bytes accessed by the request. 6. The IO Bridge may also provide some additional opaque information (e.g. tags) that are not interpreted by the IOMMU but returned along with the response from the IOMMU to the IO Bridge. As the IOMMU is allowed to complete translation requests out of order, such information may be used by the IO Bridge to correlate completions to previous requests. * Data Structure interface: it is used by the IOMMU for implicit access to memory. It is a requester interface to the IO Bridge and is used to fetch the required data structure from main memory. This interface is used to access: 1. The device and process directories to get the context information and translation rules. 2. The first-stage and/or second-stage page table entries to translate the IOVA. 3. The in-memory queues (command-queue, fault-queue, and page-request-queue) used to interface with software. * Device Translation Completion interface: it is an interface which provides the completion response from the IOMMU for previously requested address translations. The completion interface may provide information such as: 1. The status of the request, indicating if the request completed successfully or a fault occurred. 2. If the request was completed successfully; the Supervisor Physical Address (SPA). 3. Opaque information (e.g. tags), if applicable, associated with the request. 4. The page-based memory types (PBMT), if Svpbmt is supported, obtained from the IOMMU address translation page tables. The IOMMU provides the page-based memory type as resolved between the first-stage and second-stage page table entries. * ATS interface: The ATS interface, if the optional PCIe ATS capability is supported by the IOMMU, is used to communicate with ATS capable endpoints through the PCIe Root Port. This interface is used: 1. To receive ATS translation requests from the endpoints and to return the completions to the endpoints. The Root Port may provide an indication if the endpoint originating the request is a CXL type 1 or type 2 device. 2. To send ATS "Invalidation Request" messages to the endpoints and to receive the "Invalidation Completion" messages from the endpoints. 3. To receive "Page Request" and "Stop Marker" messages from the endpoints and to send "Page Request Group Response" messages to the endpoints. The interfaces related to recording an incoming MSI in a memory-resident interrupt file (MRIF) (See RISC-V Advanced Interrupt Architecture \[[6](bibliography.html#bib-aia)\]) are implementation-specific. The partitioning of responsibility between the IOMMU and the IO bridge for recording the incoming MSI in an MRIF and generating the associated _notice_ MSI are implementation-specific. ![interfaces](_images/interfaces.svg) Figure 6\. IOMMU interfaces. Similar to the RISC-V harts, physical memory attributes (PMA) and physical memory protection (PMP) checks must be completed on all inbound IO transactions even when the IOMMU is in bypass (`Bare` mode). The placement and integration of the PMA and PMP checkers is a platform choice. PMA and PMP checkers reside outside the IOMMU. The example above is showing them in the IO Bridge. Implicit accesses by the IOMMU itself through the Data Structure interface are checked by the PMA checker. PMAs are tightly tied to a given physical platform’s organization, and many details are inherently platform-specific. The memory accesses performed by the IOMMU using the Data Structure interface need not be ordered in general with the device-initiated memory accesses. | | The IOMMU may generate implicit memory accesses on the Data Structure interface to access data structures needed to perform the address translations. Such accesses must not be blocked by the original device-initiated memory access. The IO bridge may perform ordering of memory accesses on the Data Structure interface to satisfy the necessary hazard checks and other rules as defined by the IO bridge and the system interconnect. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The IOMMU provides the resolved PBMT (PMA, IO, NC) along with the translated address on the device translation completion interface to the IO Bridge. The PMA checker in the IO Bridge may use the provided PBMT to override the PMA(s) for the associated memory pages. The PMP checker may use the hardware ID of the bus access initiator to determine physical memory access privileges. As the IOMMU itself is a bus access initiator for its implicit accesses, the IOMMU hardware ID may be used by the PMP checker to select the appropriate access control rules. | | The IOMMU does not validate the authenticity of the hardware IDs provided by the IO bridge. The IO bridge and/or the root ports must include suitable mechanisms to authenticate the hardware IDs. In some SOCs this may be trivially achieved as a property of the devices being integrated into the SOC and their IDs being immutable. For PCIe, for example, the PCIe defined Access Control Services (ACS) Source Validation capabilities may be used to authenticate the hardware IDs. Other implementation-specific methods in the IO bridge may be provided to perform such authentication. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#iommu-features)IOMMU features Version 1.0 of the RISC-V IOMMU specification supports the following features: * Memory-based device context to locate parameters and address translation structures. The device context is located using the hardware-provided unique `device_id`. The supported `device_id` width may be up to 24 bits. * Memory-based process context to locate parameters and address translation structures using hardware-provided unique `process_id`. The supported`process_id` may be up to 20 bits. * 16-bit GSCIDs and 20-bit PSCIDs. * Two-stage address translation. * Page based virtual-memory system as specified by the RISC-V Privileged specification \[[7](bibliography.html#bib-priv)\] to allow software flexibility to either use a common page table for the CPU MMU as well as the IOMMU or to use a separate page table for the IOMMU. * Up to 57-bit virtual-address width, 56-bit system-physical-address, and 59-bit guest-physical-address width. * Hardware updating of PTE Accessed and Dirty bits. * Identifying memory accesses to a virtual interrupt file and MSI address translation using MSI page tables specified by the RISC-V Advanced Interrupt Architecture \[[6](bibliography.html#bib-aia)\]. * Svnapot and Svpbmt extensions. * PCIe ATS and PRI services \[[5](bibliography.html#bib-pci)\]. Support for translating an IOVA to a GPA instead of a SPA in response to a translation request. * A hardware performance monitor (HPM). * MSI and wire-signaled interrupts to request service from software. * A register interface for software to request an address translation to support debug. Features supported by the IOMMU may be discovered using the `capabilities`register [iommu\_registers.adoc#CAP](iommu%5Fregisters.html#CAP). Preface ==================== ## [](#preface)Preface **_Preface to Version 20260222_** This document describes the RISC-V IOMMU architecture. This release, version 20260222, includes the following versions of the RISC-V IOMMU Base Architecture specification and standard extensions: | Specification | Version | Status | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------- | ---------------------------------------------------------------- | | **RISC-V IOMMU Base Architecture Specification** **Quality-of-Service (QoS) Identifiers Extension** **Non-leaf PTE Invalidation Extension** **Address Range Invalidation Extension** **PTE Reserved-for-Software Bits 60-59** | **1.0** **1.0** **1.0** **1.0** **1.0** | **Ratified** **Ratified** **Ratified** **Ratified** **Ratified** | The following backward-compatible changes—​comprising a set of clarifications and corrections—​have been made since version 20250828: * Corrected typographic errors and made editorial updates. * Reworded the guidance on consistently updating valid data structures. * Clarified procedure to use for determining if ATS and PRI are enabled and included missing fault for device ID wider than that supported by the IOMMU. * Clarified that pmip is set when OF transitions from 0 to 1 due to counter overflow. * Clarified that Priv is ignored when PV is 0 in tr\_req\_ctl. * Clarified that the behavior is unspecified if Exe=1/NW=0 in tr\_req\_ctl. * Clarified memory access by IOFENCE and for IOMMU-generated MSI follow fctl.BE endianess. The following change has been made which, while not strictly backwards compatible, is not expected to cause software portability issues in practice: * When `iohgatp.MODE` is `Bare`, the `msiptp.MODE` must be set to `Off`. These changes were made through PR#617, \[[1](bibliography.html#bib-pr617)\]. **_Preface to Version 20250828_** This document describes the RISC-V IOMMU architecture. This release, version 20250828, includes the following versions of the RISC-V IOMMU Base Architecture specification and standard extensions: | Specification | Version | Status | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------- | ---------------------------------------------------------------- | | **RISC-V IOMMU Base Architecture Specification** **Quality-of-Service (QoS) Identifiers Extension** **Non-leaf PTE Invalidation Extension** **Address Range Invalidation Extension** **PTE Reserved-for-Software Bits 60-59** | **1.0** **1.0** **1.0** **1.0** **1.0** | **Ratified** **Ratified** **Ratified** **Ratified** **Ratified** | The following backward-compatible changes—​comprising a set of clarifications and corrections—​have been made since version 20250620: * Corrected typographic errors and made editorial updates. * Clarified the types of faults that may be caused by G-stage due to implicit PDT accesses. * Updated the software guideline indicating that wired-signaled interrupts are supported when `IGS` is either `WSI` or `BOTH`. * Clarified that ATS Translation responses with `U=1` include the granted permissions. * Clarified that MSI PTEs do not include A/D bits, but these bits may be assumed to be `1`. * Included definitions for TLB and Walk in the Glossary. The following change has been made which, while not strictly backwards compatible, is not expected to cause software portability issues in practice: * While the MSI address mask and pattern fields are 52 bits wide, any bits beyond the maximum GPA width supported by the IOMMU are reserved for future standard use. These changes were made through PR#569, \[[2](bibliography.html#bib-pr569)\]. **_Preface to Version 20250620_** This document describes the RISC-V IOMMU architecture. This release, version 20250620, includes the following versions of the RISC-V IOMMU Base Architecture specification and standard extensions: | Specification | Version | Status | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------- | --------------------------------------------------- | | **RISC-V IOMMU Base Architecture Specification** **Quality-of-Service (QoS) Identifiers Extension** **Non-leaf PTE Invalidation Extension** **Address Range Invalidation Extension** | **1.0** **1.0** **1.0** **1.0** | **Ratified** **Ratified** **Ratified** **Ratified** | The following backward-compatible changes—​comprising a set of clarifications and corrections—​have been made since version 20240901: * Typographic errors have been corrected, and editorial updates have been made. * Clarified that the translation size is implementation-defined when both stages are bare. * Clarified that the size of a queue is one less than the number of its entries. These changes were made through PR#441, \[[3](bibliography.html#bib-pr441)\]. **_Preface to Version 20240901_** Chapters 2 through 8 of this document form the RISC-V IOMMU Base Architecture Specification. Chapter 9 includes the standard extensions to the base architecture. This release, version 20240901, contains the following versions of the RISC-V IOMMU Base Architecture specification and standard extensions: | Specification | Version | Status | | --------------------------------------------------------------------------------------------------- | --------------- | ------------------------- | | **RISC-V IOMMU Base Architecture specification** **Quality-of-Service (QoS) Identifiers Extension** | **1.0** **1.0** | **Ratified** **Ratified** | The following backward-compatible changes, comprising a set of clarifications and corrections, have been made since version 1.0.0: * A set of typographic errors and editorial updates were made. * Translations cached, if any, in `Bare` mode do not require invalidation. * Clarified that memory faults encountered by commands also set the `cqmf` flag. * Values tested by algorithms in SW Guidelines are before modifications made by the algorithms. * Included SW guidelines for modifying non-leaf PDT entries. * Clarified the behavior for in-flight transactions observed at the time of `ddtp`write operations. * Clarified the behavior when `IOTINVAL` is invoked with an invalid address. * Stated that faults leading to UR/CA ATS responses are reported in the Fault Queue. * Added a detailed description of the `capabilities.PAS` field. * SW guidelines for changing IOMMU modes and programming `tr_req_ctl` and HPM counters. * PCIe ATS Translation Resp. grants execute permission only if requested. * Clarified the handling of hardware implementations that internally split 8-byte transactions. * Shadow stack encodings introduced by Zicfiss are reserved for IOMMU use. * Listed the fault codes reported for faults detected by Page Request. * Updated Fig 31 to remove the unused Destination ID field for ATS.PRGR * Included a software guideline for IOMMU emulation. These changes were made through PR#243 \[[4](bibliography.html#bib-pr243)\]. **_Preface to Version 1.0.0_** * Ratified version of the RISC-V IOMMU Architecture Specification. Memory-mapped register interface ==================== ## [](#memory-mapped-register-interface)Memory-mapped register interface The IOMMU provides a memory-mapped programming interface. The memory-mapped registers of each IOMMU are located within a naturally aligned 4-KiB region (a page) of physical address space. The IOMMU behavior for register accesses where the address is not aligned to the size of the access, or if the access spans multiple registers, or if the size of the access is not 4 bytes or 8 bytes, is `UNSPECIFIED`. A 4 byte access to an IOMMU register must be single-copy atomic. Whether an 8 byte access to an IOMMU register is single-copy atomic is `UNSPECIFIED`, and such an access may appear, internally to the IOMMU, as if two separate 4 byte accesses — first to the high half and second to the low half — were performed. | | The 8-byte IOMMU registers are defined in such a way that software can perform two individual 4-byte accesses, or hardware can perform two independent 4-byte transactions resulting from an 8-byte access, to the high and low halves of the register, in that order, as long as the register semantics, with regard to side-effects, are respected between the two software accesses, or two hardware transactions, respectively. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The IOMMU registers have little-endian byte order, even for systems where all harts are big-endian-only. | | Big-endian-configured harts that make use of an IOMMU are expected to implement the REV8 byte-reversal instruction defined by the Zbb extension. If REV8 is not implemented, then endianness conversion may be implemented using a sequence of instructions. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If a register is optional, as determined by the corresponding `capabilities`register bit being 0, then a read from the memory-mapped register offset of the register returns 0 and writes to that offset are ignored. Registers and register fields designated for future standard use must be read-only zero. For forward compatibility, registers and register fields designated for custom use should be implemented as read-only zero if the implementation does not define a custom use for them. ### [](#register-layout)Register layout __Table 1\. IOMMU Memory-mapped register layout__ | Offset | Name | Size | Description | Is Optional? | | ------ | ------------- | ---- | -------------------------------------------- | ------------------------ | | 0 | capabilities | 8 | [Capabilities of the IOMMU](#CAP) | No | | 8 | fctl | 4 | [Features control](#FCTRL) | No | | 12 | custom | 4 | Designated For custom use | | | 16 | ddtp | 8 | [Device directory table pointer](#DDTP) | No | | 24 | cqb | 8 | [Command-queue base](#CQB) | No | | 32 | cqh | 4 | [Command-queue head](#CQH) | No | | 36 | cqt | 4 | [Command-queue tail](#CQT) | No | | 40 | fqb | 8 | [Fault-queue base](#FQB) | No | | 48 | fqh | 4 | [Fault-queue head](#FQH) | No | | 52 | fqt | 4 | [Fault-queue tail](#FQT) | No | | 56 | pqb | 8 | [Page-request-queue base](#PQB) | if capabilities.ATS==0 | | 64 | pqh | 4 | [Page-request-queue head](#PQH) | if capabilities.ATS==0 | | 68 | pqt | 4 | [Page-request-queue tail](#PQT) | if capabilities.ATS==0 | | 72 | cqcsr | 4 | [Command-queue CSR](#CSR) | No | | 76 | fqcsr | 4 | [Fault-queue CSR](#FQCSR) | No | | 80 | pqcsr | 4 | [Page-request-queue CSR ](#PQCSR) | if capabilities.ATS==0 | | 84 | ipsr | 4 | [Interrupt pending status register](#IPSR) | No | | 88 | iocountovf | 4 | [HPM counter overflows](#OVF) | if capabilities.HPM==0 | | 92 | iocountinh | 4 | [HPM counter inhibits](#INH) | if capabilities.HPM==0 | | 96 | iohpmcycles | 8 | [HPM cycles counter](#CYC) | if capabilities.HPM==0 | | 104 | iohpmctr1-31 | 248 | [HPM event counters](#CTR) | if capabilities.HPM==0 | | 352 | iohpmevt1-31 | 248 | [HPM event selector](#EVT) | if capabilities.HPM==0 | | 600 | tr\_req\_iova | 8 | [Translation-request IOVA](#TRR%5FIOVA) | if capabilities.DBG==0 | | 608 | tr\_req\_ctl | 8 | [Translation-request control](#TRR%5FCTRL) | if capabilities.DBG==0 | | 616 | tr\_response | 8 | [Translation-request response](#TRR%5FRSP) | if capabilities.DBG==0 | | 624 | iommu\_qosid | 4 | [IOMMU QoS ID](#IOQOSID) | if capabilities.QOSID==0 | | 628 | Reserved | 60 | Reserved for future use (WPRI) | | | 688 | custom | 72 | Designated for custom use (WARL) | | | 760 | icvec | 8 | [Interrupt cause to vector register](#ICVEC) | No | | 768 | msi\_cfg\_tbl | 256 | [MSI Configuration Table](#MSI) | if capabilities.IGS==WSI | | 1024 | Reserved | 3072 | Reserved for standard use | | ### [](#reset-behavior)Reset behavior The reset value is 0 for the following registers fields. * `cqcsr` \- `cqen`, `cqie`, `cqon`, and `busy` * `fqcsr` \- `fqen`, `fqie`, `fqon`, and `busy` * `pqcsr` \- `pqen`, `pqie`, `pqon`, and `busy` * `tr_req_ctl.Go/Busy` * `ddtp.busy` The reset value is 0 for the following registers. * `ipsr` Reset value for `ddtp.iommu_mode` field must be either `Off` or `Bare`. After a reset the caches ([iommu\_data\_structures.adoc#CACHING](iommu%5Fdata%5Fstructures.html#CACHING)) must have no valid entries. | | The reset value for the iommu\_mode is recommended to be Off. | | ---------------------------------------------------------------- | The reset value is `UNSPECIFIED` for all other registers and/or fields. ### [](#CAP)IOMMU capabilities (`capabilities`) The `capabilities` register is a read-only register reporting features supported by the IOMMU. Each field if not clear indicates the presence of that feature in the IOMMU. At reset, the register shall contain the IOMMU supported features. ![IOMMU capabilities register fields](_images/diag-bf50f2dff17907df3442c4c2fd13cfc631d5263a.svg) Figure 1\. IOMMU capabilities register fields | Bits | Field | Attribute | Description | | ----- | ----------- | --------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 7:0 | version | RO | The version field holds the version of the specification implemented by the IOMMU. The low nibble is used to hold the minor version of the specification and the upper nibble is used to hold the major version of the specification. For example, an implementation that supports version 1.0 of the specification reports 0x10. | | 8 | Sv32 | RO | Page-based 32-bit virtual addressing is supported. | | 9 | Sv39 | RO | Page-based 39-bit virtual addressing is supported. | | 10 | Sv48 | RO | Page-based 48-bit virtual addressing is supported. When Sv48 is set, Sv39 must be set. | | 11 | Sv57 | RO | Page-based 57-bit virtual addressing is supported When Sv57 is set, Sv48 must be set. | | 13:12 | reserved | RO | Reserved for standard use. | | 14 | Svrsw60t59b | RO | PTE Reserved-for-Software Bits 60-59. | | 15 | Svpbmt | RO | Page-based memory types. | | 16 | Sv32x4 | RO | Page-based 34-bit virtual addressing for second-stage address translation is supported. | | 17 | Sv39x4 | RO | Page-based 41-bit virtual addressing for second-stage address translation is supported. | | 18 | Sv48x4 | RO | Page-based 50-bit virtual addressing for second-stage address translation is supported. | | 19 | Sv57x4 | RO | Page-based 59-bit virtual addressing for second-stage address translation is supported. | | 20 | reserved | RO | Reserved for standard use. | | 21 | AMO\_MRIF | RO | Atomic updates to MRIF is supported. | | 22 | MSI\_FLAT | RO | MSI address translation using Pass-through mode MSI PTE is supported. | | 23 | MSI\_MRIF | RO | MSI address translation using MRIF mode MSI PTE is supported. | | 24 | AMO\_HWAD | RO | Atomic updates to PTE accessed (A) and dirty (D) bit is supported. | | 25 | ATS | RO | PCIe Address Translation Services (ATS) and page-request interface (PRI) \[[5](bibliography.html#bib-pci)\] is supported. | | 26 | T2GPA | RO | Returning guest-physical-address in ATS translation completions is supported. | | 27 | END | RO | When 0, IOMMU supports one endianness (either little or big). When 1, IOMMU supports both endianness. The endianness is defined in the fctl register. | | 29:28 | IGS | RO | IOMMU interrupt generation support. Value Name Description 0 MSI IOMMU supports only message- signaled-interrupt generation. 1 WSI IOMMU supports only wire- signaled-interrupt generation. 2 BOTH IOMMU supports both MSI and WSI generation. The interrupt generation method must be defined in the fctl register. 3 0 Reserved for standard use | | 30 | HPM | RO | IOMMU implements a hardware performance monitor. | | 31 | DBG | RO | IOMMU supports the translation-request interface | | 37:32 | PAS | RO | Physical Address Size supported by the IOMMU. | | 38 | PD8 | RO | One level PDT with 8-bit process\_id supported. | | 39 | PD17 | RO | Two level PDT with 17-bit process\_id supported. | | 40 | PD20 | RO | Three level PDT with 20-bit process\_id supported. | | 41 | QOSID | RO | Associating QoS IDs with requests is supported. | | 42 | NL | RO | Non-leaf PTE invalidation extension is supported. | | 43 | S | RO | Address range invalidation extension is supported. | | 55:44 | reserved | RO | Reserved for standard use. | | 63:56 | custom | RO | Designated for custom use. | When `HPM` is 1, the `iohpmcycles` and the `iohpmctr1` registers must be present and be at least 32-bits wide. At least one method, `MSI` or `WSI`, of generating interrupts from the IOMMU must be supported. IOMMU implementations must support the Svnapot standard extension for NAPOT Translation Contiguity. The physical address space addressable by the IOMMU ranges from 0 to . | | Hypervisor may provide an SW emulated IOMMU to allow the guest to manage the first-stage page tables for fine grained control on memory accessed by guest controlled devices. A hypervisor that provides such an emulated IOMMU to the guest may retain control of the second-stage address translation and clear the SvNx4 fields of the emulated capabilities register. A hypervisor that provides such an emulated IOMMU to the guest may retain control of the MSI page tables used to direct MSIs to guest interrupt files in an IMSIC or to a memory-resident-interrupt-file and clear the MSI\_FLAT andMSI\_MRIF fields of the emulated capabilities register. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | The AMO\_HWAD/AMO\_MRIF bits do not indicate support for device-initiated atomic memory operations. Support for device-initiated atomic memory operations must be discovered through other means. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The IOMMU is designed to provide a highly modular and extensible set of capabilities allowing implementations to include only the exact set of capabilities required for an application. In addition, implementations may add their own custom extensions to the IOMMU. The IOMMU must support all the virtual memory extensions that are supported by any of the harts in the system. RISC-V platform specifications may mandate a set of IOMMU capabilities that must be provided by an implementation to be compliant to those specifications. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#FCTRL)Features-control register (`fctl`) This register must be readable in any implementation. An implementation may allow one or more fields in the register to be writable to support enabling or disabling the feature controlled by that field. If software enables or disables a feature when the IOMMU is not OFF (i.e. when `ddtp.iommu_mode != Off`) then the IOMMU behavior is `UNSPECIFIED`. If software enables or disables a feature when the IOMMU in-memory queues are enabled (i.e. `cqcsr.cqon/cqen == 1`, `fqcsr.fqon/fqen == 1`, or`pqcsr.pqon/pqen == 1`) then the IOMMU behavior is `UNSPECIFIED`. ![Feature-control register fields](_images/diag-a3d5682765e6d99358e062afc34bf87e121aa2cc.svg) Figure 2\. Feature-control register fields | Bits | Field | Attribute | Description | | ----- | -------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | BE | WARL | When 0, IOMMU accesses to memory resident data structures, as specified in [iommu\_data\_structures.adoc#ENDIAN\_CONFIG](iommu%5Fdata%5Fstructures.html#ENDIAN%5FCONFIG), accesses made by the IOMMU during command processing or for MSI generation, and accesses to in-memory queues are performed as little-endian accesses and when 1 as big-endian accesses. | | 1 | WSI | WARL | When 1, IOMMU interrupts are signaled as wire-signaled-interrupts else they are signaled as message-signaled-interrupts. | | 2 | GXL | WARL | Controls the address-translation schemes that may be used for guest physical addresses as defined in [iommu\_data\_structures.adoc#IOHGATP\_MODE\_ENC-0](iommu%5Fdata%5Fstructures.html#IOHGATP%5FMODE%5FENC-0) and[iommu\_data\_structures.adoc#IOHGATP\_MODE\_ENC-1](iommu%5Fdata%5Fstructures.html#IOHGATP%5FMODE%5FENC-1). | | 15:3 | reserved | WPRI | Reserved for standard use. | | 31:16 | custom | WPRI | Designated for custom use. | ### [](#DDTP)Device-directory-table pointer (`ddtp`) ![Device-directory-table pointer register fields](_images/diag-3eaf67e68c91e42b695cf84917d75a0b2e5e7cc0.svg) Figure 3\. Device-directory-table pointer register fields | Bits | Field | Attribute | Description | | ----- | ----------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 3:0 | iommu\_mode | WARL | The IOMMU may be configured to be in the following modes: Value Name Description 0 Off No inbound memory transactions are allowed by the IOMMU. 1 Bare No translation or protection. All inbound memory accesses are passed through. 2 1LVL One-level device-directory-table 3 2LVL Two-level device-directory-table 4 3LVL Three-level device-directory-table 5-13 reserved Reserved for standard use. 14-15 custom Designated for custom use. | | 4 | busy | RO | A write to ddtp.iommu\_mode may require the IOMMU to perform many operations that may not occur synchronously to the write. When a write is observed by the ddtp.iommu\_mode, the busy bit is set to 1\. When the busy bit is 1, behavior of additional writes to the ddtp isUNSPECIFIED. Some implementations may ignore the second write and others may perform the actions determined by the second write. Software must verify that the busy bit is 0 before writing to the ddtp. If the busy bit reads 0 then the IOMMU has completed the operations associated with the previous write to ddtp.iommu\_mode. An IOMMU that can complete these operations synchronously may hard-wire this bit to 0. | | 9:5 | reserved | WPRI | Reserved for standard use | | 53:10 | PPN | WARL | Holds the PPN of the root page of the device-directory-table. | | 63:54 | reserved | WPRI | Reserved for standard use | The device-context is 64-bytes in size if `capabilities.MSI_FLAT` is 1 else it is 32-bytes. When the `iommu_mode` is `Bare` or `Off`, the `PPN` field is don’t-care. When in `Bare` mode only Untranslated requests are allowed. Translated requests, Translation request, and PCIe message transactions are unsupported. All IOMMUs must support `Off` and `Bare` mode. An IOMMU is allowed to support a subset of directory-table levels and device-context widths. At a minimum one of the modes must be supported. When the `iommu_mode` field value is changed to `Off` the IOMMU guarantees that in-flight transactions, observed at the time of the write to this field, from devices connected to the IOMMU will either be processed with the configurations applicable to the old value of the `iommu_mode` field or be aborted ([iommu\_hw\_guidelines.adoc#IOBR\_FAULT\_RESP](iommu%5Fhw%5Fguidelines.html#IOBR%5FFAULT%5FRESP)). It also ensures that all transactions and previous requests from devices that have already been processed by the IOMMU are committed to a global ordering point such that they can be observed by all RISC-V harts, devices, and IOMMUs in the platform. Software must not change the `PPN` field value when transitioning the `iommu_mode` to `Off`. The IOMMU behavior of writing `iommu_mode` to `1LVL`, `2LVL`, or `3LVL`, when the previous value of the `iommu_mode` is not `Off` or `Bare` is `UNSPECIFIED`. To change DDT levels, the IOMMU must first be transitioned to `Bare` or `Off`state. The behavior resulting from changing the `iommu_mode` to `Bare` when the previous value of the `iommu_mode` was not `Off` is `UNSPECIFIED`. When an IOMMU is transitioned to `Bare` or `Off` state, the IOMMU may retain information cached from in-memory data structures such as page tables, DDT, PDT, etc. Software must use suitable invalidation commands to invalidate cached entries. | | In RV32, only the low order 32-bits of the register (22-bit PPN and 4-bit iommu\_mode) need to be written. | | ------------------------------------------------------------------------------------------------------------- | ### [](#CQB)Command-queue base (`cqb`) This 64-bit register (RW) holds the PPN of the root page of the command-queue and number of entries in the queue. Each command is 16 bytes. The IOMMU behavior on writing `cqb` when `cqcsr.busy` or `cqon` bits are 1 is`UNSPECIFIED`. The software recommended sequence to change `cqb` is to first disable the command-queue by clearing `cqen` and wait for both `cqcsr.busy` and`cqon` to be 0 before changing the `cqb`. The status of bits `31:cqb.LOG2SZ` in`cqt` following a write to `cqb` is 0 and the bits `cqb.LOG2SZ-1:0` in `cqt`assume a valid but otherwise `UNSPECIFIED` value. ![Command-queue base register fields](_images/diag-e42548f50fb7ac84cb476e64eed446e62804ec2f.svg) Figure 4\. Command-queue base register fields | Bits | Field | Attribute | Description | | ----- | -------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 4:0 | LOG2SZ-1 | WARL | The LOG2SZ-1 field holds the number of entries in command-queue as a log to base 2 minus 1\. A value of 0 indicates a queue of 2 entries. Each IOMMU command is 16-bytes. If the command-queue has 256 or fewer entries then the base address of the queue is always aligned to 4-KiB. If the command-queue has more than 256 entries then the command-queue base address must be naturally aligned to2LOG2SZ x 16. | | 9:5 | reserved | WPRI | Reserved for standard use | | 53:10 | PPN | WARL | Holds the PPN of the root page of the in-memory command-queue used by software to queue commands to the IOMMU. If the base address as determined by PPN is not aligned as required, all entries in the queue appear to an IOMMU as UNSPECIFIED and any address an IOMMU may compute and use for accessing an entry in the queue is also UNSPECIFIED. | | 63:54 | reserved | WPRI | Reserved for standard use | | | In RV32, only the low order 32-bits of the register (22-bit PPN and 5-bit LOG2SZ-1) need to be written. | | ---------------------------------------------------------------------------------------------------------- | ### [](#CQH)Command-queue head (`cqh`) This 32-bit register (RO) holds the index into the command-queue where the IOMMU will fetch the next command. ![Command-queue head register fields](_images/diag-a3cf5cac70c011cf10cdac580a4f5867d04f6227.svg) Figure 5\. Command-queue head register fields | Bits | Field | Attribute | Description | | ---- | ----- | --------- | ------------------------------------------------------------------------------------------------ | | 31:0 | index | RO | Holds the index into the command-queue from where the next command will be fetched by the IOMMU. | ### [](#CQT)Command-queue tail (`cqt`) This 32-bit register (RW) holds the index into the command-queue where the software queues the next command for the IOMMU. ![Command-queue tail register fields](_images/diag-a3cf5cac70c011cf10cdac580a4f5867d04f6227.svg) Figure 6\. Command-queue tail register fields | Bits | Field | Attribute | Description | | ---- | ----- | --------- | -------------------------------------------------------------------------------------------------------------------------- | | 31:0 | index | WARL | Holds the index into the command-queue where software queues the next command for IOMMU. OnlyLOG2SZ-1:0 bits are writable. | ### [](#FQB)Fault queue base (`fqb`) This 64-bit register (RW) holds the PPN of the root page of the fault-queue and number of entries in the queue. Each fault record is 32 bytes. The IOMMU behavior on writing `fqb` when `fqcsr.busy` or `fqon` bits are 1 is`UNSPECIFIED`. The software recommended sequence to change `fqb` is to first disable the fault-queue by clearing `fqen` and wait for both `fqcsr.busy` and`fqon` to be 0 before changing the `fqb`. The status of bits `31:fqb.LOG2SZ`in `fqh` following a write to `fqb` is 0 and the bits `fqb.LOG2SZ-1:0` in `fqh`assume a valid but otherwise `UNSPECIFIED` value. ![Fault queue base register fields](_images/diag-541e503d61771a86af1a3b67dd44280dd0e2db47.svg) Figure 7\. Fault queue base register fields | Bits | Field | Attribute | Description | | ----- | -------- | --------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 4:0 | LOG2SZ-1 | WARL | The LOG2SZ-1 field holds the number of entries in the fault-queue as a log-to-base-2 minus 1\. A value of 0 indicates a queue of 2 entries. Each fault record is 32-bytes. If the fault-queue has 128 or fewer entries then the base address of the queue is always aligned to 4-KiB. If the fault-queue has more than 128 entries then the fault-queue base address must be naturally aligned to 2LOG2SZ x 32. | | 9:5 | reserved | WPRI | Reserved for standard use | | 53:10 | PPN | WARL | Holds the PPN of the root page of the in-memory fault-queue used by IOMMU to queue fault record. If the base address as determined by PPN is not aligned as required, all entries in the queue appear to an IOMMU as UNSPECIFIED and any address an IOMMU may compute and use for accessing an entry in the queue is alsoUNSPECIFIED. | | 63:54 | reserved | WPRI | Reserved for standard use | | | In RV32, only the low order 32-bits of the register (22-bit PPN and 5-bit LOG2SZ-1) need to be written. | | ---------------------------------------------------------------------------------------------------------- | ### [](#FQH)Fault queue head (`fqh`) This 32-bit register (RW) holds the index into the fault-queue where the software will fetch the next fault record. ![Fault queue head register fields](_images/diag-a3cf5cac70c011cf10cdac580a4f5867d04f6227.svg) Figure 8\. Fault queue head register fields | Bits | Field | Attribute | Description | | ---- | ----- | --------- | ----------------------------------------------------------------------------------------------------------------------- | | 31:0 | index | WARL | Holds the index into the fault-queue from which software reads the next fault record. OnlyLOG2SZ-1:0 bits are writable. | ### [](#FQT)Fault queue tail (`fqt`) This 32-bit register (RO) holds the index into the fault-queue where the IOMMU queues the next fault record. ![Fault queue tail register fields](_images/diag-a3cf5cac70c011cf10cdac580a4f5867d04f6227.svg) Figure 9\. Fault queue tail register fields | Bits | Field | Attribute | Description | | ---- | ----- | --------- | ------------------------------------------------------------------------------ | | 31:0 | index | RO | Holds the index into the fault-queue where IOMMU writes the next fault record. | ### [](#PQB)Page-request-queue base (`pqb`) This 64-bit register (WARL) holds the PPN of the root page of the page-request-queue and number of entries in the queue. Each "Page Request" message is 16 bytes. The IOMMU behavior on writing `pqb` when `pqcsr.busy` or `pqon` bits are 1 is`UNSPECIFIED`. The software recommended sequence to change `pqb` is to first disable the page-request-queue by clearing `pqen` and wait for both `pqcsr.busy`and `pqon` to be 0 before changing the `pqb`. The status of bits `31:pqb.LOG2SZ`in `pqh` following a write to `pqb` is 0 and the bits `pqb.LOG2SZ-1:0` in `pqh`assume a valid but otherwise `UNSPECIFIED` value. ![Page-Request-queue base register fields](_images/diag-da57316fb2ed0d55a89c405be9db5ecb32a474ca.svg) Figure 10\. Page-Request-queue base register fields | Bits | Field | Attribute | Description | | ----- | -------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 4:0 | LOG2SZ-1 | WARL | The LOG2SZ-1 field holds the number of entries in the page-request-queue as a log-to-base-2 minus 1\. A value of 0 indicates a queue of 2 entries. Each page-request is 16-bytes. If the page-request-queue has 256 or fewer entries then the base address of the queue is always aligned to 4-KiB. If the page-request-queue has more than 256 entries then the page-request-queue base address must be naturally aligned to 2LOG2SZ x 16. | | 9:5 | reserved | WPRI | Reserved for standard use | | 53:10 | PPN | WARL | Holds the PPN of the root page of the in-memory page-request-queue used by IOMMU to queue "Page Request" messages. If the base address as determined by PPN is not aligned as required, all entries in the queue appear to an IOMMU as UNSPECIFIED and any address an IOMMU may compute and use for accessing an entry in the queue is also UNSPECIFIED. | | 63:54 | reserved | WPRI | Reserved for standard use | | | In RV32, only the low order 32-bits of the register (22-bit PPN and 5-bit LOG2SZ-1) need to be written. | | ---------------------------------------------------------------------------------------------------------- | ### [](#PQH)Page-request-queue head (`pqh`) This 32-bit register (RW) holds the index into the page-request-queue where software will fetch the next page-request. ![Page-request-queue head register fields](_images/diag-a3cf5cac70c011cf10cdac580a4f5867d04f6227.svg) Figure 11\. Page-request-queue head register fields | Bits | Field | Attribute | Description | | ---- | ----- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------- | | 31:0 | index | WARL | Holds the index into the page-request-queue from which software reads the next "Page Request" message. Only LOG2SZ-1:0 bits are writable. | ### [](#PQT)Page-request-queue tail (`pqt`) This 32-bit register (RO) holds the index into the page-request-queue where the IOMMU writes the next page-request. ![Page-request-queue tail register fields](_images/diag-a3cf5cac70c011cf10cdac580a4f5867d04f6227.svg) Figure 12\. Page-request-queue tail register fields | Bits | Field | Attribute | Description | | ---- | ----- | --------- | ----------------------------------------------------------------------------------------------- | | 31:0 | index | RO | Holds the index into the page-request-queue where IOMMU writes the next "Page Request" message. | ### [](#CSR)Command-queue CSR (`cqcsr`) This 32-bit register (RW) is used to control the operations and report the status of the command-queue. ![Command-queue CSR register fields](_images/diag-4e65a49236fe164344a985c83479bf26ecb69ea8.svg) Figure 13\. Command-queue CSR register fields | Bits | Field | Attribute | Description | | ----- | ------------ | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 0 | cqen | RW | The command-queue-enable bit enables the command- queue when set to 1. Changing cqen from 0 to 1 sets the cqh register and the cqcsr bits cmd\_ill,cmd\_to,cqmf, fence\_w\_ip to 0\. The command-queue may take some time to be active following setting thecqen to 1\. During this delay the busy bit is 1\. When the command queue is active, the cqon bit reads 1\. When cqen is changed from 1 to 0, the command queue may stay active (with busy asserted) until the commands already fetched from the command-queue are being processed and/or there are outstanding implicit loads from the command-queue. When the command-queue turns off the cqon bit reads 0. When the cqon bit reads 0, the IOMMU guarantees that no implicit memory accesses to the command queue are in-flight and the command-queue will not generate new implicit loads to the queue memory. | | 1 | cie | RW | Command-queue-interrupt-enable bit enables generation of interrupts from command-queue when set to 1. | | 7:2 | reserved | WPRI | Reserved for standard use | | 8 | cqmf | RW1C | If command-queue access to fetch a command or a memory access made by a command leads to a memory fault, then the command-queue-memory-fault bit is set to 1, and the command-queue stalls until this bit is cleared. To re-enable command processing, software should clear this bit by writing 1. | | 9 | cmd\_to | RW1C | If the execution of a command leads to a timeout (e.g. a command to invalidate device ATC may timeout waiting for a completion), then the command-queue sets the cmd\_to bit and stops processing from the command-queue. To re-enable command processing, software should clear this bit by writing 1. | | 10 | cmd\_ill | RW1C | If an illegal or unsupported command is fetched and decoded by the command-queue then the command-queue sets the cmd\_ill bit and stops processing from the command-queue. To re-enable command processing software should clear this bit by writing 1. | | 11 | fence\_w\_ip | RW1C | An IOMMU that supports wire-signaled-interrupts sets the fence\_w\_ip bit to indicate completion of an IOFENCE.C command. To re-enable interrupts on IOFENCE.C completion, software should clear this bit by writing 1\. This bit is reserved if the IOMMU does not support wire-signaled-interrupts or wire-signaled-interrupts are not enabled (i.e., fctl.WSI == 0). | | 15:12 | reserved | WPRI | Reserved for standard use | | 16 | cqon | RO | The command-queue is active if cqon is 1. | | 17 | busy | RO | A write to cqcsr may require the IOMMU to perform many operations that may not occur synchronously to the write. When a write is observed by thecqcsr, the busy bit is set to 1. When the busy bit is 1, behavior of additional writes to the cqcsr is UNSPECIFIED. Some implementations may ignore the second write and others may perform the actions determined by the second write. Software must verify that the busy bit is 0 before writing to the cqcsr. An IOMMU that can complete these operations synchronously may hard-wire this bit to 0. | | 27:18 | reserved | WPRI | Reserved for standard use. | | 31:28 | custom | WPRI | Designated for custom use. | When `cmd_ill` or `cqmf` is 1 in `cqcsr`, the `cqh` references the command in the CQ that caused the error. Previous commands may have completed, timed out, or their execution aborted by the IOMMU. | | If software makes the CQ operational again after a cmd\_ill or cqmf error, then software should resubmit the commands submitted since the last IOFENCE.Cthat successfully completed. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `cmd_to` bit is set when a `IOFENCE.C` command detects that one or more previous commands that are specified to have timeouts have timed out but all other commands previous to the `IOFENCE.C` have completed. When `cmd_to` is 1,`cqh` references the `IOFENCE.C` command that detected the timeout. | | Command-queue being empty does not imply that all commands fetched from the command-queue have been completed. When the command-queue is requested to be disabled, an implementation may either complete the already fetched commands or abort execution of those commands. Software must use an IOFENCE.C command to wait for all previous commands to be committed, if so desired, before turning off the command-queue. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#FQCSR)Fault queue CSR (`fqcsr`) This 32-bit register (RW) is used to control the operations and report the status of the fault-queue. ![Fault queue CSR register fields](_images/diag-3017af1a6b625cdbff938fae5d741547a99f6b1a.svg) Figure 14\. Fault queue CSR register fields | Bits | Field | Attribute | Description | | ----- | -------- | --------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | fqen | RW | The fault-queue enable bit enables the fault-queue when set to 1. Changing fqen from 0 to 1 sets the fqt register and the fqcsr bits fqof and fqmf to 0\. The fault-queue may take some time to be active following setting the fqen to 1\. During this delay the busy bit is 1\. When the fault queue is active, the fqon bit reads 1. When fqen is changed from 1 to 0, the fault-queue may stay active (with busy asserted) until in-flight fault-recording is completed. When the fault-queue is off the fqon bit reads 0. When fqon reads 0, the IOMMU guarantees that there are no in-flight implicit writes to the fault-queue in progress and that no new fault records will be written to the fault-queue. | | 1 | fie | RW | Fault queue interrupt enable bit enables generation of interrupts from fault-queue when set to 1. | | 7:2 | reserved | WPRI | Reserved for standard use | | 8 | fqmf | RW1C | The fqmf bit is set to 1 if the IOMMU encounters an access fault when storing a fault record to the fault queue. The fault-record that was attempted to be written is discarded and no more fault records are generated until software clears the fqmf bit by writing 1 to the bit. | | 9 | fqof | RW1C | The fault-queue-overflow bit is set to 1 if the IOMMU needs to queue a fault record but the fault-queue is full (i.e., fqt == fqh - 1). The fault-record is discarded and no more fault records are generated until software clears fqof by writing 1 to the bit. | | 15:10 | reserved | WPRI | Reserved for standard use | | 16 | fqon | RO | The fault-queue is active if fqon reads 1. | | 17 | busy | RO | Write to fqcsr may require the IOMMU to perform many operations that may not occur synchronously to the write. When a write is observed by the fqcsr, the busy bit is set to 1\. When the busy bit is 1, behavior of additional writes to the fqcsr areUNSPECIFIED. Some implementations may ignore the second write and others may perform the actions determined by the second write. Software should ensure that the busy bit is 0 before writing to the fqcsr. An IOMMU that can complete controls synchronously may hard-wire this bit to 0. | | 27:18 | reserved | WPRI | Reserved for standard use. | | 31:28 | custom | WPRI | Designated for custom use. | ### [](#PQCSR)Page-request-queue CSR (`pqcsr`) This 32-bit register (RW) is used to control the operations and report the status of the page-request-queue. ![Page-request-queue CSR register fields](_images/diag-e96413e8c774b700567b41fb80ccd2aded6781a2.svg) Figure 15\. Page-request-queue CSR register fields | Bits | Field | Attribute | Description | | ----- | -------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 0 | pqen | RW | The page-request-enable bit enables the page-request-queue when set to 1. Changing pqen from 0 to 1, sets the pqt register and the pqcsr bits pqmf and pqof to 0\. The page-request-queue may take some time to be active following setting the pqen to 1\. During this delay the busy bit is 1\. When the page-request-queue is active, the pqon bit reads 1. When pqen is changed from 1 to 0, the page-request-queue may stay active (with busy asserted) until in-flight page-request writes are completed. When the page-request-queue turns off, the pqon bit reads 0. When pqon reads 0, the IOMMU guarantees that there are no older in-flight implicit writes to the queue memory and no further implicit writes will be generated to the queue memory. The IOMMU may respond to “Page Request” messages received when page-request-queue is off or in the process of being turned off, as specified in[iommu\_data\_structures.adoc#ATS\_PRI](iommu%5Fdata%5Fstructures.html#ATS%5FPRI). | | 1 | pie | RW | The page-request-queue-interrupt-enable bit when set to 1, enables generation of interrupts from page-request-queue. | | 7:2 | reserved | WPRI | Reserved for standard use | | 8 | pqmf | RW1C | The pqmf bit is set to 1 if the IOMMU encounters an access fault when storing a "Page Request" message to the page-request-queue. The "Page Request" message that caused the pqmf or pqof error and all subsequent "Page Request" messages are discarded until software clears thepqof and/or pqmf bits by writing 1 to it. The IOMMU may respond to “Page Request” messages that caused the pqof or pqmf bit to be set and all subsequent “Page Request” messages received while these bits are 1 as specified in[iommu\_data\_structures.adoc#ATS\_PRI](iommu%5Fdata%5Fstructures.html#ATS%5FPRI). | | 9 | pqof | RW1C | The page-request-queue-overflow bit is set to 1 if the page-request queue overflows i.e. IOMMU needs to queue a "Page Request" message but the page-request queue is full (i.e., pqt == pqh - 1). The "Page Request" message that caused the pqmf or pqof error and all subsequent "Page Request" messages are discarded until software clears thepqof and/or pqmf bits by writing 1 to it. The IOMMU may respond to “Page Request” messages that caused the pqof or pqmf bit to be set and all subsequent “Page Request” messages received while these bits are 1 as specified in[iommu\_data\_structures.adoc#ATS\_PRI](iommu%5Fdata%5Fstructures.html#ATS%5FPRI). | | 15:10 | reserved | WPRI | Reserved for standard use | | 16 | pqon | RO | The page-request is active when pqon reads 1. | | 17 | busy | RO | A write to pqcsr may require the IOMMU to perform many operations that may not occur synchronously to the write. When a write is observed by the pqcsr, the busy bit is set to 1. When the busy bit is 1, behavior of additional writes to the pqcsr are UNSPECIFIED. Some implementations may ignore the second write and others may perform the actions determined by the second write. Software should ensure that thebusy bit is 0 before writing to the pqcsr. An IOMMU that can complete controls synchronously may hard-wire this bit to 0 | | 27:18 | reserved | WPRI | Reserved for standard use | | 31:28 | custom | WPRI | Designated for custom use. | ### [](#IPSR)Interrupt pending status register (`ipsr`) This 32-bit register (RW1C) reports the pending interrupts which require software service. Each interrupt-pending bit in the register corresponds to a interrupt source in the IOMMU. The interrupt-pending bit in the register once set to 1 stays 1 till software clears that interrupt-pending bit by writing 1 to clear it. When `fctl.WSI` is 1, the interrupt-pending bit drives the wire selected by the corresponding `icvec` field to signal an interrupt. When `fctl.WSI` is 0, the IOMMU signals interrupts using messages. MSI have edge semantics and an interrupt message is generated when an interrupt-pending bit transitions from 0 to 1\. The address and data for the message are obtained from the `msi_cfg_tbl` entry selected by the `icvec` field corresponding to the interrupt-pending bit. ![Interrupt pending status register fields](_images/diag-f1540e2e25d03945fea846eea76758016b58cc20.svg) Figure 16\. Interrupt pending status register fields __Table 2\. Interrupt pending status register fields__ | Bits | Field | Attribute | Description | | ----- | -------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | cip | RW1C | The command-queue-interrupt-pending bit is set to 1 if cqcsr.cie is 1 and any of the following are true: cqcsr.fence\_w\_ip is 1. cqcsr.cmd\_ill is 1. cqcsr.cmd\_to is 1. cqcsr.cqmf is 1. | | 1 | fip | RW1C | The fault-queue-interrupt-pending bit is set to 1 if fqcsr.fie is 1 and any of the following are true: fqcsr.fqof is 1. fqcsr.fqmf is 1. A new record is produced in the FQ. | | 2 | pmip | RW1C | The performance-monitoring-interrupt-pending bit is set to 1 when the OF bit in iohpmcycles or in any of theiohpmevt1-31 registers transitions from 0 to 1 due to a counter overflow. | | 3 | pip | RW1C | The page-request-queue-interrupt-pending is set to 1 if pqcsr.pie is 1 and any of the following are true: pqcsr.pqof is 1. pqcsr.pqmf is 1. A new message is produced in the PQ. | | 7:4 | reserved | WPRI | Reserved for standard use. | | 15:8 | custom | WPRI | Designated for custom use. | | 31:16 | reserved | WPRI | Reserved for standard use | If a bit in `ipsr` is 1 then a write of 1 to the bit transitions the bit from 1→0\. If the conditions to set that bit are still present (See [Table 2](#IPSR%5FFIELD)) or if they occur after the bit is cleared then that bit transitions again from 0→1. ### [](#OVF)Performance-monitoring counter overflow status (`iocountovf`) The performance-monitoring counter overflow status is a 32-bit read-only register that contains shadow copies of the OF bits in the `iohpmevt1-31`registers - where `iocountovf` bit X corresponds to `iohpmevtX` and bit 0 corresponds to the `OF` bit of `iohpmcycles`. This register enables overflow interrupt handler software to quickly and easily determine which counter(s) have overflowed. ![Performance-monitoring counter overflow status register fields](_images/diag-a0317eff5b83cdefd66cf4defd1c9c15bdb4355d.svg) Figure 17\. Performance-monitoring counter overflow status register fields | Bits | Field | Attribute | Description | | ---- | ----- | --------- | ----------------------------- | | 0 | CY | RO | Shadow of iohpmcycles.OF | | 31:1 | HPM | RO | Shadow of iohpmevt\[1-31\].OF | ### [](#INH)Performance-monitoring counter inhibits (`iocountinh`) The performance-monitoring counter inhibits is a 32-bit WARL register that contains bits to inhibit the corresponding counters from counting. Bit X when set inhibits counting in `iohpmctrX` and bit 0 inhibits counting in`iohpmcycles`. ![Performance-monitoring counter inhibits register fields](_images/diag-a0317eff5b83cdefd66cf4defd1c9c15bdb4355d.svg) Figure 18\. Performance-monitoring counter inhibits register fields | Bits | Field | Attribute | Description | | ---- | ----- | --------- | -------------------------------------------------------------------- | | 0 | CY | RW | When set, iohpmcycles counter is inhibited from counting. | | 31:1 | HPM | WARL | When bit X is set, then counting of events iniohpmctrX is inhibited. | | | When the iohpmcycles counter is not needed, it is desirable to conditionally inhibit it to reduce energy consumption. Providing a single register to inhibit all counters allows a) one or more counters to be atomically programmed with events to count b) one or more counters to be sampled atomically. To initialize an event counter or the cycles counter to a desired value, it should be first inhibited if it is enabled to count. This measure ensures that it does not count during the update process. The inhibition should be removed after the register has been programmed with the desired value. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#CYC)Performance-monitoring cycles counter (`iohpmcycles`) This 64-bit register is a free running clock cycle counter. There is no associated `iohpmevt0`. ![Performance-monitoring cycles counter register fields](_images/diag-fd76fa68ea9a2c0855eb3f73dfe3c46c642a9a27.svg) Figure 19\. Performance-monitoring cycles counter register fields | Bits | Field | Attribute | Description | | ---- | ------- | --------- | --------------------- | | 62:0 | counter | WARL | Cycles counter value. | | 63 | OF | RW | Overflow | The `OF` bit is set when the `iohpmcycles` counter overflows, and remains set until cleared by software. Since `iohpmcycles` value is an unsigned value, overflow is defined as unsigned overflow. Note that there is no loss of information after an overflow since the counter wraps around and keeps counting while the sticky `OF` bit remains set. If the `iohpmcycles` counter overflows when the `OF` bit is zero, then a HPM Counter Overflow interrupt is generated by setting `ipsr.pmip` bit to 1\. If the `OF` bit is already one, then no interrupt request is generated. Consequently the `OF` bit also functions as a count overflow interrupt disable for the`iohpmcycles`. ### [](#CTR)Performance-monitoring event counters (`iohpmctr1-31`) These registers are 64-bit WARL counter registers. ![Performance-monitoring event counters register fields](_images/diag-90e3ebd61a0bf019ef0df71a6441d645f9a38495.svg) Figure 20\. Performance-monitoring event counters register fields | Bits | Field | Attribute | Description | | ---- | ------- | --------- | -------------------- | | 63:0 | counter | WARL | Event counter value. | ### [](#EVT)Performance-monitoring event selectors (`iohpmevt1-31`) These performance-monitoring event registers are 64-bit RW registers. When a transaction processed by the IOMMU causes an event that is programmed to count in a counter then the counter is incremented. In addition to matching events, the event selector may be programmed with additional filters based on`device_id`, `process_id`, `GSCID`, and `PSCID` such that the counter is incremented conditionally based on the transaction matching these additional filters. When such `device_id` based filtering is used, the match may be configured to be a precise match or a partial match. A partial match allows transactions with a range of IDs to be counted by the counter. ![Performance-monitoring event selector register fields](_images/diag-6e2ef357a1d1983e4c8d8fb4579835a18f5128d2.svg) Figure 21\. Performance-monitoring event selector register fields | Bits | Field | Attribute | Description | | ----- | ---------- | --------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 14:0 | eventID | WARL | Indicates the event to count. A value of 0 indicates no events are counted. Encodings 1 to 16383 are reserved for standard events defined in the [Table 5](#Event%5Flist). Encodings 16384 to 32767 are for designated for custom use. When eventID is changed, including to 0, the counter retains its value. | | 15 | DMASK | RW | When set to 1, partial matching of theDID\_GSCID is performed for the transaction. The lower bits of the DID\_GSCID all the way to the first low order 0 bit (including the 0 bit position itself) are masked. | | 35:16 | PID\_PSCID | RW | process\_id if IDT is 0,PSCID if IDT is 1 | | 59:36 | DID\_GSCID | RW | device\_id if IDT is 0,GSCID if IDT is 1. | | 60 | PV\_PSCV | RW | If set, only transactions with matchingprocess\_id or PSCID (based on the Filter ID Type) are counted. | | 61 | DV\_GSCV | RW | If set, only transactions with matchingdevice\_id or GSCID (based on the Filter ID Type) are counted. | | 62 | IDT | RW | Filter ID Type: This field indicates the type of ID to filter on. When 0, the DID\_GSCID field holds a device\_id and the PID\_PSCID field holds a process\_id. When 1, theDID\_GSCID field holds a GSCID andPID\_PSCID field holds a PSCID. | | 63 | OF | RW | Overflow status or Interrupt disable | The table below summarizes the filtering option for events that support filtering by IDs. __Table 3\. filtering options__ | **IDT** | **DV\_GSCV** | **PV\_PSCV** | **Operation** | | ------- | ------------ | ------------ | -------------------------------------------------------------------------------------------------------------------------------- | | 0/1 | 0 | 0 | Counter increments. No ID based filtering. | | 0 | 0 | 1 | If the transaction has a validprocess\_id, counter increments if process\_id matches PID\_PSCID. | | 0 | 1 | 0 | Counter increments if device\_id matches DID\_GSCID. | | 0 | 1 | 1 | If the transaction has a validprocess\_id, counter increments ifdevice\_id matches DID\_GSCID andprocess\_id matches PID\_PSCID. | | 1 | 0 | 1 | If the transaction has a validPSCID, counter increments if the PSCID of that process matchesPID\_PSCID. | | 1 | 1 | 0 | Counter increments if GSCID is valid and matches DID\_GSCID. | | 1 | 1 | 1 | Counter increments if GSCID is valid and matches DID\_GSCID and if PSCID is valid and matches PID\_PSCID. | When filtering by `device_id` or `GSCID` is selected and the event supports ID based filtering, the DMASK field can be used to configure a partial match. When DMASK is set to 1, partial matching of the `DID_GSCID` is performed for the transaction. The lower bits of the `DID_GSCID` all the way to the first low order 0 bit (including the 0 bit position itself) are masked. The following example illustrates the use of DMASK and filtering by `device_id`. __Table 4\. DMASK with IDT set to device\_id based filtering__ | DMASK | DID\_GSCID | **Comment** | | ----- | -------------------------- | ----------------------------- | | 0 | yyyyyyyy yyyyyyyy yyyyyyyy | One specific seg:bus:dev:func | | 1 | yyyyyyyy yyyyyyyy yyyyy011 | seg:bus:dev - any func | | 1 | yyyyyyyy yyyyyyyy 01111111 | seg:bus - any dev:func | | 1 | yyyyyyyy 01111111 11111111 | seg - any bus:dev:func | The following table lists the standard events that can be counted: __Table 5\. Standard Events list__ | **eventID** | **Event counted** | **IDT settings supported** | | ----------- | ----------------------------- | -------------------------- | | 0 | Do not count | | | 1 | Untranslated requests | 0 | | 2 | Translated requests | 0 | | 3 | ATS Translation requests | 0 | | 4 | TLB miss | 0/1 | | 5 | Device Directory Walks | 0 | | 6 | Process Directory Walks | 0 | | 7 | First-stage Page Table Walks | 0/1 | | 8 | Second-stage Page Table Walks | 0/1 | | 9 - 16383 | reserved for future standard | \- | When the programmed `IDT` setting is not supported for an event then the associated counter does not increment. The `OF` bit is set when the corresponding `iohpmctr1-31` counter overflows, and remains set until cleared by software. Since `iohpmctr1-31` values are unsigned values, overflow is defined as unsigned overflow. Note that there is no loss of information after an overflow since the counter wraps around and keeps counting while the sticky `OF` bit remains set. If a `iohpmctr1-31` counter overflows when the associated `OF` bit is zero, then a HPM Counter Overflow interrupt is generated by setting `ipsr.pmip` bit to 1\. If the `OF` bit is already one, then no interrupt request is generated. Consequently the `OF` bit also functions as a count overflow interrupt disable for the associated `iohpmctr1-31`. | | There are not separate overflow status and overflow interrupt enable bits. In practice, enabling overflow interrupt generation (by clearing the OF bit) is done in conjunction with initializing the counter to a starting value. Once a counter has overflowed, it and the OF bit must be reinitialized before another overflow interrupt can be generated. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | In RV32, memory-mapped writes to iohpmevt1-31 modify only one 32-bit part of the register. The following sequence may be used to update the register without counting events spuriously due to the intermediate value of the register: Write the low order 32-bits to set eventID to 0. Write the high order 32-bits with the new desired values. Write the low order 32-bits the new desired values, including that of theeventID field. Alternatively, the counter may first be inhibited such that no events count during the update and the inhibit removed after the register has been programmed with the desired value. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | If capabilities.HPM is 1 then a minimum of one programmable event counter besides the cycles counter is required to comply with this specification. One counter may be used in a time multiplexed manner to sample events but such analysis may take longer to complete. The IOMMU, unlike the CPU MMU, services multiple streams of IO and the HPM may be used by a performance analyst to analyze one or more of those streams concurrently. Typically, a performance analyst may require four programmable counters to count events for an IO stream. To support concurrent analysis of at least two streams of IO it is recommended to support seven programmable counters. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#TRR%5FIOVA)Translation-request IOVA (`tr_req_iova`) The `tr_req_iova` is a 64-bit register used to implement a translation-request interface for debug. This register is present when`capabilities.DBG == 1`. ![Translation-request IOVA register fields](_images/diag-32a657882dc44165640c3ce607467ceefd86f3ba.svg) Figure 22\. Translation-request IOVA register fields | Bits | Field | Attribute | Description | | ----- | -------- | --------- | ---------------------------- | | 11:0 | reserved | WPRI | Reserved for standard use | | 63:12 | vpn | WARL | The IOVA virtual page number | ### [](#TRR%5FCTRL)Translation-request control (`tr_req_ctl`) The `tr_req_ctl` is a 64-bit WARL register used to implement a translation-request interface for debug. This register is present when`capabilities.DBG == 1`. ![Translation-request control register fields](_images/diag-6171071f230b8d5d3ef8288aa3c96f2186f3225d.svg) Figure 23\. Translation-request control register fields | Bits | Field | Attribute | Description | | ----- | -------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | Go/Busy | RW1S | This bit is set to indicate a valid request has been setup in thetr\_req\_iova/tr\_req\_ctl registers for the IOMMU to translate. The IOMMU indicates completion of the requested translation by clearing this bit to 0\. On completion, the results of the translation are in the tr\_response register. | | 1 | Priv | WARL | If set to 1, Privileged Mode access is requested; otherwise, Privileged Mode access is not requested. When PV is 0, this field is ignored and no Privileged Mode access is requested. | | 2 | Exe | WARL | If set to 1, execute permission is requested else execute permission is not requested. If Exe is set to 1, then NW must be also be set to 1; otherwise the behavior is UNSPECIFIED. | | 3 | NW | WARL | If set to 1, read permission is requested. If set to 0, both read and write permissions are requested. | | 11:4 | reserved | WPRI | Reserved for standard use | | 31:12 | PID | WARL | If PV is 1, this field provides theprocess\_id input for this translation request. If PV is 0 then this field is not used. | | 32 | PV | WARL | If set to 1, the PID field of the register is valid and provides the process\_id for this translation request. If set to 0 then the PID field is not used and a process\_id is not valid for this translation request. | | 35:33 | reserved | WPRI | Reserved for standard use. | | 39:36 | custom | WPRI | Designated for custom use. | | 63:40 | DID | WARL | This field provides the device\_id for this translation request. | | | In RV32, the high half of the register should be written first, followed by the low half, which includes the Go/Busy bit, to initiate a translation. | | ------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#TRR%5FRSP)Translation-response (`tr_response`) The `tr_response` is a 64-bit RO register used to hold the results of a translation requested using the translation-request interface. This register is present when `capabilities.DBG == 1`. ![Translation-response register fields](_images/diag-0dbd72804d0ef1d6175880cf672efa5b0cb0833e.svg) Figure 24\. Translation-response register fields | Bits | Field | Attribute | Description | | ----- | -------- | --------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | fault | RO | If the process to translate the IOVA detects a fault then the fault field is set to 1\. The detected fault may be reported through the fault-queue. | | 6:1 | reserved | RO | Reserved for standard use | | 8:7 | PBMT | RO | Memory type determined for the translation using the PBMT fields in the first-stage and/or the second-stage page tables used for the translation. This value of this field isUNSPECIFIED if the fault field is 1. | | 9 | S | RO | Translation range size field, when set to 1 indicates that the translation applies to a range that is larger than 4 KiB and the size of the translation range is encoded in thePPN field. The value of this field isUNSPECIFIED if the fault field is 1. | | 53:10 | PPN | RO | If the fault bit is 0, then this field provides the PPN determined as a result of translating the vpn in tr\_req\_iova. If the fault bit is 1, then the value of this field is UNSPECIFIED. If the S bit is 0, then the size of the translation is 4 KiB - a page. If the S bit is 1, then the translation resulted in a superpage, and the size of the superpage is encoded in the PPN itself. If scanning from bit position 0 to bit position 43, the first bit with a value of 0 at position X, then the superpage size is2X+1 \* 4 KiB. If X is not 0, then all bits at position 0 through X-1 are each encoded with a value of 1. Table 6\. Example of encoding of super page size in PPN PPN S Size yyyy…​.yyyy yyyy yyyy 0 4 KiB yyyy…​.yyyy yyyy 0111 1 64 KiB yyyy…​.yyy0 1111 1111 1 2 MiB yyyy…​.yy01 1111 1111 1 4 MiB | | 59:54 | reserved | RO | Reserved for standard use. | | 63:60 | custom | RO | Designated for custom use. | | | An IOMMU implementation is not required to report a superpage translation or support reporting all possible superpage sizes. An implementation is allowed to report a 4 KiB translation corresponding to the requestedvpn or report a translation size that is smaller than the superpage size configured in the page tables. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#IOQOSID)IOMMU QoS ID (`iommu_qosid`) The `iommu_qosid` register fields are defined as follows: ![`iommu_qosid` register fields](_images/diag-c4837722535117106755568dc19ec27db080323f.svg) Figure 25\. `iommu_qosid` register fields | Bits | Field | Attribute | Description | | ----- | -------- | --------- | ---------------------------------- | | 11:0 | RCID | WARL | RCID for IOMMU-initiated requests. | | 15:12 | reserved | WPRI | Reserved for standard use. | | 27:16 | MCID | WARL | MCID for IOMMU-initiated requests. | | 31:28 | reserved | WPRI | Reserved for standard use. | IOMMU-initiated requests for accessing the following data structures use the value programmed in the `RCID` and `MCID` fields of the `iommu_qosid` register. * Device directory table (`DDT`) * Fault queue (`FQ`) * Command queue (`CQ`) * Page-request queue (`PQ`) * IOMMU-initiated MSI (Message-signaled interrupts) When `ddtp.iommu_mode == Bare`, all device-originated requests are associated with the QoS IDs configured in the `iommu_qosid` register. ### [](#ICVEC)Interrupt-cause-to-vector register (`icvec`) Interrupt-cause-to-vector register maps a cause to a vector. All causes can be mapped to the same vector or a cause can be given a unique vector. The vector is used: 1. By an IOMMU that generates interrupts as MSIs, to index into MSI configuration table (`msi_cfg_tbl`) to determine the MSI to generate. An IOMMU is capable of generating interrupts as a MSI if `capabilities.IGS==MSI`or if `capabilities.IGS==BOTH`. When `capabilities.IGS==BOTH` the IOMMU may be configured to generate interrupts as MSI by setting `fctl.WSI` to 0. 2. By an IOMMU that generates WSI, to determine the wire to signal the interrupt. An IOMMU is capable of generating wire-signaled- interrupts if `capabilities.IGS==WSI` or if `capabilities.IGS==BOTH`. When`capabilities.IGS==BOTH` the IOMMU may be configured to generate wire-signaled- interrupts by setting `fctl.WSI` to 1. If an implementation only supports a single vector then all bits of this register may be hardwired to 0 (WARL). Likewise if only two vectors are supported then only bit 0 for each cause could be writable. ![Interrupt-cause-to-vector register fields](_images/diag-ad1453e8b46e7077878d0dd577d9be137be77777.svg) Figure 26\. Interrupt-cause-to-vector register fields | Bits | Field | Attribute | Description | | ----- | -------- | --------- | ------------------------------------------------------------------------------------------------------------------------- | | 3:0 | civ | WARL | The command-queue-interrupt-vector (civ) is the vector number assigned to the command-queue-interrupt. | | 7:4 | fiv | WARL | The fault-queue-interrupt-vector (fiv) is the vector number assigned to the fault-queue-interrupt. | | 11:8 | pmiv | WARL | The performance-monitoring-interrupt-vector (pmiv) is the vector number assigned to the performance-monitoring-interrupt. | | 15:12 | piv | WARL | The page-request-queue-interrupt-vector (piv) is the vector number assigned to the page-request-queue-interrupt. | | 31:16 | reserved | WPRI | Reserved for standard use. | | 63:32 | custom | WPRI | Designated for custom use. | ### [](#MSI)MSI configuration table (`msi_cfg_tbl`) An IOMMU that supports generating IOMMU-originated interrupts (i.e., `capabilities.IGS == MSI` or `capabilities.IGS == BOTH`) as MSIs implements a MSI configuration table that is indexed by the vector from `icvec`to determine a MSI table entry. Each MSI table entry for interrupt vector `x`has three registers `msi_addr_x`, `msi_data_x`, and `msi_vec_ctl_x`. These registers are hardwired to 0 if `capabilities.IGS == WSI`. If an access fault is detected on a MSI write using `msi_addr_x`, then the IOMMU reports a "IOMMU MSI write access fault" (cause 273) fault, with `TTYP` set to 0 and `iotval` set to the value of `msi_addr_x`. __Table 7\. MSI configuration table structure__ | bit 63 | bit 0 | Byte Offset | | ------------------------ | --------------------- | ----------- | | Entry 0: Message address | +000h | | | Entry 0: Vector Control | Entry 0: Message Data | +008h | | Entry 1: Message address | +010h | | | Entry 1: Vector Control | Entry 1: Message Data | +018h | | …​ | +020h | | ![Message address register fields](_images/diag-3aec10aa6a37dd6474da8948960523e8fe9f7635.svg) Figure 27\. Message address register fields | Bits | Field | Attribute | Description | | ----- | -------- | --------- | ------------------------------------- | | 1:0 | 0 | RO | Fixed to 0 | | 55:2 | ADDR | WARL | Holds the 4-byte aligned MSI address. | | 63:56 | reserved | WPRI | Reserved for standard use. | ![Message data register fields](_images/diag-93a3d2ad3a7d4fff70115b85ecc44d3083698ba3.svg) Figure 28\. Message data register fields | Bits | Field | Attribute | Description | | ---- | ----- | --------- | ------------------ | | 31:0 | data | WARL | Holds the MSI data | ![Vector control register fields](_images/diag-4bdeba9e5b83657977b13d489cf88575411c0371.svg) Figure 29\. Vector control register fields | Bits | Field | Attribute | Description | | ---- | -------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | M | RW | When the mask bit M is 1, the corresponding interrupt vector is masked and the IOMMU is prohibited from sending the associated message. Pending messages for that vector are later generated if the corresponding mask bit is cleared to 0. | | 31:1 | reserved | WPRI | Reserved for standard use. | Software guidelines ==================== ## [](#sw%5Fguidelines)Software guidelines This section provides guidelines to software developers on the correct and expected sequence of using the IOMMU interfaces. The behavior of the IOMMU if these guidelines are not followed is implementation defined. ### [](#reading-and-writing-iommu-registers)Reading and writing IOMMU registers Read or write access to IOMMU registers must follow the following rules: * Address of the access must be aligned to the size of the access. * The access must not span multiple registers. * Registers that are 64-bit wide may be accessed using either a 32-bit or a 64-bit access. * Registers that are 32-bit wide must only be accessed using a 32-bit access. ### [](#guidelines-for-initialization)Guidelines for initialization The guidelines for initializing the IOMMU are as follows: 1. Read the `capabilities` register to discover the capabilities of the IOMMU. 2. Stop and report failure if `capabilities.version` is not supported. 3. Read the feature control register (`fctl`). 4. Stop and report failure if big-endian memory access is needed and the`capabilities.END` field is 0 (i.e. only one endianness) and `fctl.BE` is 0 (i.e. little endian). 5. If big-endian memory access is needed and the `capabilities.END` field is 1 (i.e. both endiannesses supported), set `fctl.BE` to 1 (i.e. big endian) if the field is not already 1. 6. Stop and report failure if wire-signaled-interrupts are needed for IOMMU initiated interrupts and `capabilities.IGS` is neither `WSI` nor `BOTH`. 7. If wire-signaled-interrupts are needed for IOMMU initiated interrupts and`capabilities.IGS` is `BOTH`, set `fctl.WSI` to 1 if the field is not already 1. 8. Stop and report failure if other required capabilities (e.g. virtual-addressing modes, MSI translation, etc.) are not supported. 9. The `icvec` register is used to program an interrupt vector for each interrupt cause. Determine the number of vectors supported by the IOMMU by writing 0xF to each field and reading back the number of writable bits. If the number of writable bits is `N` then the number of supported vectors is`2N`. For each cause `C` associate a vector `V` with the cause. `V` is a number between 0 and `(2N - 1)`. 10. If the IOMMU is configured to use wired interrupts, then each vector `V`corresponds to an interrupt wire connected to a platform level interrupt controller (e.g. APLIC). Determine the interrupt controller configuration register to be programmed for each such wire using configuration information provided by configuration mechanisms such as device tree and program the interrupt controller. 11. If the IOMMU is configured to use MSI, then each vector `V` is an index into the `msi_cfg_tbl`. For each vector `V`, allocate a MSI address `A` and an interrupt identity `D`. Configure the `msi_addr_V` register with value `A`,`msi_data_V` register with value `D`. Configure the interrupt mask `M` in`msi_vec_ctl_V` register appropriately. 12. To program the command queue, first determine the number of entries `N` needed in the command queue. The number of entries in the command queue must be a power of two. Allocate a `N` x 16-bytes sized memory buffer that is naturally aligned to the greater of 4-KiB or `N` x 16-bytes. Let `k=log2(N)` and `B`be the physical page number (PPN) of the allocated memory buffer. Program the command queue registers as follows: * `temp_cqb_var.PPN = B` * `temp_cqb_var.LOG2SZ-1 = (k - 1)` * `cqb = temp_cqb_var` * `cqt = 0` * `cqcsr.cqen = 1` * Poll on `cqcsr.cqon` until it reads 1 13. To program the fault queue, first determine the number of entries N needed in the fault queue. The number of entries in the fault queue is always a power of two. Allocate a `N` x 32-bytes sized memory buffer that is naturally aligned to the greater of 4-KiB or `N` x 32-bytes. Let `k=log2(N)` and `B`be the PPN of the allocated memory buffer. Program the fault queue registers as follows: * `temp_fqb_var.PPN = B` * `temp_fqb_var.LOG2SZ-1 = (k - 1)` * `fqb = temp_fqb_var` * `fqh = 0` * `fqcsr.fqen = 1` * Poll on `fqcsr.fqon` until it reads 1 14. To program the page-request queue, first determine the number of entries `N`needed in the page-request queue. The number of entries in the page-request queue is always a power of two. Allocate a `N` x 16-bytes sized buffer that is naturally aligned to the greater of 4-KiB or `N` x 16-bytes. Let `k=log2(N)`and `B` be the PPN of the allocated memory buffer. Program the page-request queue registers as follows: * `temp_pqb_var.PPN = B` * `temp_pqb_var.LOG2SZ-1 = (k - 1)` * `pqb = temp_pqb_var` * `pqh = 0` * `pqcsr.pqen = 1` * Poll on `pqcsr.pqon` until it reads 1 15. To program the DDT pointer, first determine the supported `device_id` width `Dw`and the format of the device-context data structure. If `capabilities.MSI` is 0, then the IOMMU uses base-format device-contexts else extended-format device-contexts are used. Allocate a page (4 KiB) of memory to use as the root table of the DDT. Initialize the allocated memory to all 0\. Let `B` be the PPN of the allocated memory. Determine the mode `M` of the DDT based on `Dw`and the IOMMU device-contexts format as follows: * Determine the values supported by `ddtp.iommu_mode` by writing legal values and reading it to see if the value was retained. Stop and report a failure if the supported modes do not support the required `Dw`. * If extended-format device-contexts are used then 1. If `Dw` is less than or equal to 6-bits and `1LVL` is supported then `M = 1LVL` 2. If `Dw` is less than or equal to 15-bits and `2LVL` is supported then `M = 2LVL` 3. If `Dw` is less than or equal to 24-bits and `3LVL` is supported then `M = 3LVL` * If base-format device-contexts are used then 1. If `Dw` is less than or equal to 7-bits and `1LVL` is supported then `M = 1LVL` 2. If `Dw` is less than or equal to 16-bits and `2LVL` is supported then `M = 2LVL` 3. If `Dw` is less than or equal to 24-bits and `3LVL` is supported then `M = 3LVL` * Program the `ddtp` register as follows: 1. `temp_ddtp_var.iommu_mode = M` 2. `temp_ddtp_var.PPN = B` 3. `ddtp = temp_ddtp_var` The IOMMU is initialized and may be now be configured with device-contexts for devices in scope of the IOMMU. ### [](#guidelines-for-invalidations)Guidelines for invalidations This section provides guidelines to software on the invalidation commands to send to the IOMMU through the `CQ` when modifying the IOMMU in-memory data structures. Software must perform the invalidation after the update is globally visible. The ordering on stores provided by FENCE instructions and the acquire/ release bits on atomic instructions also orders the data structure updates associated with those stores as observed by IOMMU. A `IOFENCE.C` command may be used by software to ensure that all previous commands fetched from the `CQ` have been completed and committed. The `PR`and/or `PW` bits may be set to 1 in the `IOFENCE.C` command to request that all previous read and/or write requests, that have already been processed by the IOMMU, be committed to a global ordering point as part of the `IOFENCE.C`command. In subsequent sections, when an algorithm step tests values in the in-memory data structures to determine the type of invalidation operation to perform, the data values tested are the old values i.e. values before a change is made. #### [](#DC%5FCHANGE)Changing device directory table entry If software changes a leaf-level DDT entry (i.e, a device context (`DC`), of device with `device_id = D`) then the following invalidations must be performed: * `IODIR.INVAL_DDT` with `DV=1` and `DID=D` * If `DC.iohgatp.MODE != Bare` * `IOTINVAL.VMA` with `GV=1`, `AV=PSCV=0`, and `GSCID=DC.iohgatp.GSCID` * `IOTINVAL.GVMA` with `GV=1`, `AV=0`, and `GSCID=DC.iohgatp.GSCID` * else * If `DC.tc.PDTV==1` * `IOTINVAL.VMA` with `GV=AV=PSCV=0` * else if `DC.fsc.MODE != Bare` * `IOTINVAL.VMA` with `GV=AV=0` and `PSCV=1`, and `PSCID=DC.ta.PSCID` If software changes a non-leaf-level DDT entry the following invalidations must be performed: * `IODIR.INVAL_DDT` with `DV=0` Between a change to the DDT entry and when an invalidation command to invalidate the cached entry is processed by the IOMMU, the IOMMU may use the old value or the new value of the entry. #### [](#PC%5FCHANGE)Changing process directory table entry If software changes a leaf-level PDT entry (i.e, a process context (`PC`), for`device_id=D` and `process_id=P`) then the following invalidations must be performed: * `IODIR.INVAL_PDT` with `DV=1`, `DID=D` and `PID=P` * If `DC.iohgatp.MODE != Bare` * `IOTINVAL.VMA` with `GV=1`, `AV=0`, `PSCV=1`, `GSCID=DC.iohgatp.GSCID`, and `PSCID=PC.PSCID` * else * `IOTINVAL.VMA` with `GV=0`, `AV=0`, `PSCV=1`, and `PSCID=PC.PSCID` If software changes a non-leaf-level PDT entry the following invalidations must be performed: * `IODIR.INVAL_DDT` with `DV=1` and `DID=D` Between a change to the PDT entry and when an invalidation command to invalidate the cached entry is processed by the IOMMU, the IOMMU may use the old value or the new value of the entry. #### [](#MSI%5FPT%5FCHANGE)Changing MSI page table entry If software changes a MSI page-table entry identified by interrupt file number `I` that corresponds to an untranslated MSI address `A` then the following invalidations must be performed: * `IOTINVAL.GVMA` with `GV=AV=1`, `ADDR[63:12]=A[63:12]` and`GSCID=DC.iohgatp.GSCID` To invalidate all cache entries from a MSI page table the following invalidations must be performed: * `IOTINVAL.GVMA` with `GV=1`, `AV=0`, and `GSCID=DC.iohgatp.GSCID` Between a change to the MSI PTE and when an invalidation command to invalidate the cached PTE is processed by the IOMMU, the IOMMU may use the old PTE value or the new PTE value. An `IOFENCE.C` command with `PW=1` may be used to to ensure that all previous writes, including MSI writes, that have been previously processed by the IOMMU are committed to a global ordering point such that they can be observed by all RISC-V harts and IOMMUs in the system. #### [](#changing-second-stage-page-table-entry)Changing second-stage page table entry If software changes a leaf second-stage page-table entry of a VM where the change affects translation for a guest-PPN `G` then the following invalidations must be performed: * `IOTINVAL.GVMA` with `GV=AV=1`, `GSCID=DC.iohgatp.GSCID`, and `ADDR[63:12]=G` If software changes a non-leaf second-stage page-table entry of a VM then the following invalidations must be performed: * `IOTINVAL.GVMA` with `GV=1`, `AV=0`, `GSCID=DC.iohgatp.GSCID` The `DC` has fields that hold a guest-PPN. An implementation may translate such fields to a supervisor-PPN as part of caching the `DC`. If the second-stage page table update affects translation of guest-PPN held in the `DC` then software must invalidate all such cached `DC` using `IODIR.INVAL_DDT` with `DV=1` and`DID` set to the corresponding `device_id`. Alternatively, an`IODIR.INVAL_DDT` with `DV=0` may be used to invalidate all cached `DC`. Between a change to the second-stage PTE and when an invalidation command to invalidate the cached PTE is processed by the IOMMU, the IOMMU may use the old PTE value or the new PTE value. #### [](#changing-first-stage-page-table-entry)Changing first-stage page table entry A `DC` may be configured with a first-stage page table (when `DC.tc.PDTV=0`) or a directory of first-stage page tables selected using `process_id` from a process-directory-table (when `DC.tc.PDTV=1`). When a change is made to a first-stage page table, and the second-stage is Bare, then software must perform invalidations using `IOTINVAL.VMA` with`GV=0` and `AV` and `PSCV` operands appropriate for the modification as specified in [iommu\_in\_memory\_queues.adoc#IVMA](iommu%5Fin%5Fmemory%5Fqueues.html#IVMA). When a change is made to a first-stage page table, and the second-stage is not Bare, then software must perform invalidations using `IOTINVAL.VMA` with`GV=1`, `GSCID=DC.iohgatp.GSCID` and `AV` and `PSCV` operands appropriate for the modification as specified in [iommu\_in\_memory\_queues.adoc#IVMA](iommu%5Fin%5Fmemory%5Fqueues.html#IVMA). Between a change to the first-stage PTE and when an invalidation command to invalidate the cached PTE is processed by the IOMMU, the IOMMU may use the old PTE value or the new PTE value. #### [](#accessed-adirty-d-bit-updates-and-page-promotions)Accessed (A)/Dirty (D) bit updates and page promotions When IOMMU supports hardware-managed A and D bit updates, if software clears the A and/or D bit in the first-stage and/or second-stage PTEs then software must invalidate corresponding PTE entries that may be cached by the IOMMU. If such invalidations are not performed, then the IOMMU may not set these bits when processing subsequent transactions that use such entries. When software upgrades a page in a first-stage PT and/or a second-stage PT to a superpage without first clearing the original non-leaf PTE’s valid bit and invalidating cached translations in the IOMMU then it is possible for the IOMMU to cache multiple entries that match a single address. The IOMMU may use either the old non-leaf PTE or the new non-leaf PTE but the behavior is otherwise well defined. When promoting and/or demoting page sizes, software must ensure that the original and new PTEs have identical permission and memory type attributes and the physical address that is determined as a result of translation using either the original or the new PTE is otherwise identical for any given input. The only PTE update supported by the IOMMU without first clearing the V bit in the original PTE and executing a appropriate `IOTINVAL` command is to do a page size promotion or demotion. The behavior of the IOMMU if other attributes are changed in this fashion is implementation defined. #### [](#device-address-translation-cache-invalidations)Device Address Translation Cache invalidations When first-stage and/or second-stage page tables are modified, invalidations may be needed to the DevATC in the devices that may have cached translations from the modified page tables. Invalidation of such page tables requires generating ATS invalidations using `ATS.INVAL` command. Software must specify the `PAYLOAD`following the rules defined in PCIe ATS specifications \[[5](bibliography.html#bib-pci)\]. If software generates ATS invalidate requests at a rate that exceeds the average DevATC service rate then flow control mechanisms may be triggered by the device to throttle the rate. A side effect of this is congestion spreading to other channels and links which could lead to performance degradation. An ATS capable device publishes the maximum number of invalidations it can buffer before causing back-pressure through the Queue Depth field of the ATS capability structure. When the device is virtualized using PCIe SR-IOV, this queue depth is shared among all the VFs of the device. Software must limit the number of outstanding ATS invalidations queued to the device advertised limit. The `RID` field is used to specify the routing ID of the ATS invalidation request message destination. A PASID specific invalidation may be performed by setting `PV=1` and specifying the PASID in `PID`. When the IOMMU supports multiple segments then the `RID` must be qualified by the destination segment number by setting `DSV=1` with the segment number provided in `DSEG`. When ATS protocol is enabled for a device, the IOMMU may still cache translations in its IOATC in addition to providing translations to the DevATC. Software must not skip IOMMU translation cache invalidations even when ATS is enabled in the device context of the device. Since a translation request from the DevATC may be satisfied by the IOMMU from the IOATC, to ensure correct operation software must first invalidate the IOATC before sending invalidations to the DevATC. #### [](#caching-invalid-entries)Caching invalid entries This specification does not allow the caching of first/second-stage PTEs whose`V` (valid) bit is clear, non-leaf DDT entries whose `V` (valid) bit is clear, Device-context whose `V` (valid) bit is clear, non-leaf PDT entries whose `V`(valid) bit is clear, Process-context whose `V` (valid) bit is clear, or MSI PTEs whose `V` bit is clear. Software need not perform invalidations when changing the `V` bit in these entries from 0 to 1. #### [](#guidelines-for-emulating-an-iommu)Guidelines for emulating an IOMMU Certain uses may involve emulating a RISC-V IOMMU. In such cases, the emulator may require the IOMMU driver to notify the emulator for efficient operation when updates are made to in-memory data structure entries, including when making such entries valid. Queueing an appropriate invalidation command when making such updates is a common way to provide notifications to the emulator. While usually an invalidation is not required when marking an invalid entry as valid, the emulator may indicate the need to invoke such invalidation commands for emulation efficiency purposes through a suitable flag in the device tree or ACPI table describing such emulated IOMMU instances. ### [](#reconfiguring-pmas)Reconfiguring PMAs Where platforms support dynamic reconfiguration of PMAs, a machine-mode driver is usually provided that can correctly configure the platform. In some platforms that might involve platform-specific operations and if the IOMMU must participate in these operations then platform-specific operations in the IOMMU are used by the machine-mode driver to perform such reconfiguration. ### [](#guidelines-for-handling-interrupts-from-iommu)Guidelines for handling interrupts from IOMMU IOMMU may generate an interrupt from the `CQ`, the `FQ`, the `PQ`, or the HPM. Each interrupt source may be configured with a unique vector or a vector may be shared among one or more interrupt sources. The interrupt may be delivered as a MSI or a wire-signaled-interrupt. The interrupt handler may perform the following actions: 1. Read the `ipsr` register to determine the source of the pending interrupts 2. If the `ipsr.cip` bit is set then an interrupt is pending from the `CQ`. 1. Read the `cqcsr` register. 2. Determine if an error caused the interrupt and if so, the cause of the error by examining the state of the `cmd_to`, `cmd_ill`, and `cqmf` bits. If any of these bits are set then the `CQ` encountered an error and command processing is temporarily disabled. 3. If errors have occurred, correct the cause of the error and clear the bits corresponding to the corrected errors in `cqcsr` by writing 1 to the bits. 1. Clearing all error indication bits in `cqcsr` re-enables command processing. 4. An IOMMU that supports wired-interrupts may be requested to generate an interrupt from the command queue on completion of a `IOFENCE.C` command. This cause is indicated by the `fence_w_ip` bit. Note that command processing does not stop when `fence_w_ip` is set to 1\. Software handler may re-enable interrupts from `CQ` on `IOFENCE.C` completions by clearing this bit by writing 1 to it. 5. Clear `ipsr.cip` by writing 1 to the bit. 3. If the `ipsr.fip` bit is set then an interrupt is pending from the `FQ`. 1. Read the `fqcsr` register. 2. Determine if an error caused the interrupt and if so, the cause of the error by examining the state of the `fqmf` and `fqof` bits. If either of these bits are set then the `FQ` encountered an error and fault/event reporting is temporarily disabled. 3. If errors have occurred, correct the cause of the error and clear the bits corresponding to the corrected errors in `fqcsr` by writing 1 to the bits. 1. Clearing all error indication bits in `fqcsr` re-enables fault/event reporting. 4. Clear `ipsr.fip` by writing 1 to the bit. 5. Read the `fqt` and `fqh` registers. 6. If value of `fqt` is not equal to value of `fqh` then the `FQ` is not empty and contains fault/event reports that need processing. 7. Process pending fault/event reports that need processing and remove them from the `FQ` by advancing the `fqh` by the number of records processed. 4. If the `ipsr.pip` bit is set then an interrupt is pending from the `PQ`. 1. Read the `pqcsr` register. 2. Determine if an error caused the interrupt and if so, the cause of the error by examining the state of the `pqmf` and `pqof` bits. If either of these bits are set then the `PQ` encountered an error and "Page Request" reporting is temporarily disabled. 3. If errors have occurred, correct the cause of the error and clear the bits corresponding to the corrected errors in `pqcsr` by writing 1 to the bits. 1. Clearing all error indication bits in `pqcsr` re-enables "Page Request" reporting. 4. Clear `ipsr.pip` by writing 1 to the bit. 5. Read the `pqt` and `pqh` registers. 6. If value of `pqt` is not equal to the value of `pqh` then the `PQ` is not empty and contains "Page Request" messages that need processing. 7. Process pending "Page Request" messages that need processing and remove them from the `PQ` by advancing the `pqh` by the number of records processed. 1. When a `PQ` overflow condition occurs, software may observe incomplete page-request groups due to the "Page Request" messages being dropped. The IOMMU might have automatically responded (see [iommu\_data\_structures.adoc#ATS\_PRI](iommu%5Fdata%5Fstructures.html#ATS%5FPRI)) to a dropped "Page Request" in such groups if the "Last Request in PRG" flag was set to 1 in the message. Software should ignore and not service the such incomplete groups. 2. The automatic response to the "Page Request" with "Last request in PRG" set to 1 on a `PQ` overflow is expected to cause the device to retry the ATS translation request. However, since the IOMMU generated response was without actually resolving the condition that caused the "Page Request" to be originally sent by the device, this will likely lead to the device sending the "Page Request" messages again. These retried messages may now be stored in the `PQ` if the overflow condition has been corrected by creating space in the `PQ`. 5. If `ipsr.pmip` bit is set then an interrupt is pending from the HPM. 1. Clear `ipsr.pmip` by writing 1 to the bit. 2. Process the performance monitoring counter overflows. ### [](#guidelines-for-enabling-and-disabling-ats-andor-pri)Guidelines for enabling and disabling ATS and/or PRI To enable ATS and/or PRI: 1. Place the device in an idle state such that no transactions are generated by the device. 2. If the device-context for the device is already valid then first mark the device-context as invalid and queue commands to the IOMMU to invalidate all cached first/second-stage page table entries, DDT entries, MSI PT entries (if required), and PDT entries (if required). 3. Program the device-context with `EN_ATS` set to 1 and if required the `T2GPA`field set to 1\. Set `EN_PRI` to 1 if required. If `EN_PRI` is set to 1 then set `PRPR` to 1 if required. 4. Mark the device-context as valid. 5. Enable device to use ATS and if required enable the PRI. To disable ATS and/or PRI: 1. Place the device in an idle state such that no transactions are generated by the device. 2. Disable ATS and/or PRI at the device 3. Set `EN_ATS` and/or `EN_PRI` to 0 in the device-context. If `EN_ATS` is set to 0 then set `EN_PRI` and `T2GPA` to 0\. If `EN_PRI` is set to 0 then set `PRPR`to 0. 4. Queue commands to the IOMMU to invalidate all cached first/second-stage page table entries, DDT entries, MSI PT entries (if required), and PDT entries (if required). 5. Queue commands to the IOMMU to invalidate DevATC by generating Invalidation Request messages. 6. Enable DMA operations in the device. RISC-V Platform-Level Interrupt Controller Specification ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-platform-level-interrupt-controller-specification)RISC-V Platform-Level Interrupt Controller Specification RISC-V Task Group Version 1.0.0, 3/2023 | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Change Log ==================== ## [](#change-log)Change Log ### [](#version-1-0-0)version 1.0.0 * 2023-3-11 * Ratified version ### [](#version-1-0-0%5Frc6)version 1.0.0\_rc6 * 2023-1-5 * Remove H/U mode from this spec. ### [](#version-1-0-0%5Frc5)version 1.0.0\_rc5 * 2022-10-29 * Clarify the "register width" when access to memory map register. * Combine Register Width chapter and Memory Map chapter. ### [](#version-1-0-0%5Frc4)version 1.0.0\_rc4 * 2022-8-27 * Update specification to Frozen state. ### [](#version-1-0-0%5Frc3)version 1.0.0\_rc3 * 2022-6-11 * Revised Copyright and license information * Added Andrew Waterman and Krste Asanovic as contributors who did the original design and wrote the spec. ### [](#version-1-0-0%5Frc2)version 1.0.0\_rc2 * 2022-5-30 * Revised the copyright statements to follow section 3.3 in Appendix A "INTELLECTUAL PROPERTY RIGHTS POLICY " of the RISC-V Regulations. * Correct offsets of PLIC Interrupt Enable Bits memory map. ### [](#version-1-0-0%5Frc1)version 1.0.0\_rc1 * 2022-4-16 * The pre-public review version Contributors ==================== ## [](#contributors)Contributors The contributor to RISC-V PLIC specification in alphabetical order: Abner Chang <[abner.chang@hpe.com](mailto:abner.chang@hpe.com)\> Andrew Waterman <[andrew@sifive.com](mailto:andrew@sifive.com)\> Bin Meng <[bmeng.cn@gmail.com](mailto:bmeng.cn@gmail.com)\> Drew Barbier <[drew@sifive.com](mailto:drew@sifive.com)\> Jeff Scheel <[jeff@riscv.org](mailto:jeff@riscv.org)\> Jessica Clarke <[jrtc27@jrtc27.com](mailto:jrtc27@jrtc27.com)\> Jinyan Xu <[phantom@zju.edu.cn](mailto:phantom@zju.edu.cn)\> Krste Asanovic <[krste@sifive.com](mailto:krste@sifive.com)\> Palmer Dabbelt <[palmer@dabbelt.com](mailto:palmer@dabbelt.com)\> Robert Balas <[balasr@iis.ee.ethz.ch](mailto:balasr@iis.ee.ethz.ch)\> 1.1. Introduction ==================== ## [](#1-1-introduction)1.1\. Introduction This specification delineates the operation parameters according the general PLIC architecture defined in the RISC-V platform-level interrupt controller (PLIC) specification (was removed from [RISC-V Privileged Spec v1.11-draft](https://github.com/riscv/riscv-isa-manual/releases/download/draft-20181201-5449851/riscv-privileged.pdf)) to work in the context of RISC-V systems. The PLIC multiplexes various device interrupts onto the external interrupt lines of Hart contexts, with hardware support for interrupt priorities. PLIC supports up-to 1023 interrupts (0 is reserved) and 15872 contexts, but the actual number of interrupts and context depends on the PLIC implementation. However, the implementation must adhere to the offset of each register within the PLIC operation parameters. The PLIC which claimed as PLIC-Compliant standard PLIC should follow the implementations mentioned in sections below. Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This RISC-V PLIC specification is © 2017-2023 RISC-V international This document is released under a Creative Commons Attribution 4.0 International License. . Please cite as: “RISC-V Platform-Level Interrupt Controller Specification", RISC-V International This document is a derivative of the "The RISC-V Instruction Set Manual, Volume II: Privileged Architecture, Document version 1.9.1" released under following license: © 2010–2017 Andrew Waterman, Yunsup Lee, Rimas Aviˇzienis, David Patterson, Krste Asanovi ́c. Creative Commons Attribution 4.0 International License. 8.1. Interrupt Claim Process ==================== ## [](#8-1-interrupt-claim-process)8.1\. Interrupt Claim Process Sometime after a target receives an interrupt notification, it might decide to service the interrupt. The target sends an interrupt claim message to the PLIC core, which will usually be implemented as a non-idempotent memory-mapped I/O control register read. On receiving a claim message, the PLIC core will atomically determine the ID of the highest-priority pending interrupt for the target and then clear down the corresponding source’s IP bit. The PLIC core will then return the ID to the target. The PLIC core will return an ID of zero, if there were no pending interrupts for the target when the claim was serviced. After the highest-priority pending interrupt is claimed by a target and the corresponding IP bit is cleared, other lower-priority pending interrupts might then become visible to the target, and so the PLIC EIP bit might not be cleared after a claim. The interrupt handler can check the local meip/seip/ueip bits before exiting the handler, to allow more efficient service of other interrupts without first restoring the interrupted context and taking another interrupt trap. It is always legal for a hart to perform a claim even if the EIP is not set. In particular, a hart could set the threshold value to maximum to disable interrupt notifications and instead poll for active interrupts using periodic claim requests, though a simpler approach to implement polling would be to clear the external interrupt enable in the corresponding xie register for privilege mode x. The PLIC can perform an interrupt claim by reading the `claim/complete`register, which returns the ID of the highest priority pending interrupt or zero if there is no pending interrupt. A successful claim will also atomically clear the corresponding pending bit on the interrupt source. The PLIC can perform a claim at any time and the claim operation is not affected by the setting of the priority threshold register. The Interrupt Claim Process register is context based and is located at (4K alignment + 4) starts from offset 0x200000. | **PLIC Register Block Name** | **Function** | **Register Block Size in Byte** | **Description** | | ---------------------------- | ------------------------------------------ | ----------------------------------------- | ------------------------------------------------------------------ | | Interrupt Claim Register | Interrupt Claim Process for 15872 contexts | 4096 \* 15872 = 65011712(0x3e00000) bytes | This is the register used to acquire interrupt ID for each context | **PLIC Interrupt Claim Process Memory Map** 0x200004: Interrupt Claim Process for context 0 0x201004: Interrupt Claim Process for context 1 0x202004: Interrupt Claim Process for context 2 0x203004: Interrupt Claim Process for context 3 ... ... ... 0x3FFF004: Interrupt Claim Process for context 15871 Interrupt Completion ==================== ## [](#interrupt-completion)Interrupt Completion The PLIC signals it has completed executing an interrupt handler by writing the interrupt ID it received from the claim to the `claim/complete` register. The PLIC does not check whether the completion ID is the same as the last claim ID for that target. If the completion ID does not match an interrupt source that is currently enabled for the target, the completion is silently ignored. After a handler has completed service of an interrupt, the associated gateway must be sent an interrupt completion message, usually as a write to a non-idempotent memory-mapped I/O control register. The gateway will only forward additional interrupts to the PLIC core after receiving the completion message. The Interrupt Completion registers are context based and located at the same address with Interrupt Claim Process register, which is at (4K alignment + 4) starts from offset 0x200000. | **PLIC Register Block Name** | **Registers** | **Register Block Size in Byte** | **Description** | | ----------------------------- | --------------------------------------- | ----------------------------------------- | ------------------------------------------------------- | | Interrupt Completion Register | Interrupt Completion for 15872 contexts | 4096 \* 15872 = 65011712(0x3e00000) bytes | This is register to write to complete Interrupt process | **PLIC Interrupt Completion Memory Map** 0x200004: Interrupt Completion for context 0 0x201004: Interrupt Completion for context 1 0x202004: Interrupt Completion for context 2 0x203004: Interrupt Completion for context 3 ... ... ... 0x3FFF004: Interrupt Completion for context 15871 6.1. Interrupt Enables ==================== ## [](#6-1-interrupt-enables)6.1\. Interrupt Enables Each global interrupt can be enabled by setting the corresponding bit in the`enables` register. The `enables` registers are accessed as a contiguous array of 32-bit registers, packed the same way as the `pending` bits. Bit 0 of enable register 0 represents the non-existent interrupt ID 0 and is hardwired to 0\. PLIC has 15872 Interrupt Enable blocks for the contexts. How PLIC organizes interrupts for the contexts (Hart and privilege mode) is out of RISC-V PLIC specification scope, however it must be spec-out in vendor’s PLIC specification. (_A large number of potential IE bits might be hardwired to zero in cases where some interrupt sources can only be routed to a subset of targets. A larger number of bits might be wired to 1 for an embedded device with fixed interrupt routing. Interrupt priorities, thresholds, and hart-internal interrupt masking provide considerable flexibility in ignoring external interrupts even if a global interrupt source is always enabled._) The base address of Interrupt Enable Bits block within PLIC Memory Map region is fixed at 0x002000. | **PLIC Register Block Name** | **Function** | **Register Block Size in Byte** | **Description** | | ---------------------------- | ----------------------------------------------------------------------- | --------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Interrupt Enable Bits | Interrupt Enable Bit of Interrupt Source #0 to #1023 for 15872 contexts | (1024 / 8) \* 15872 = 2031616(0x1f0000) bytes | This is a continuously memory block contains PLIC Interrupt Enable Bits of 15872 contexts. Each Interrupt Enable Bit occupies 1-bit from this register block and total 15872 Interrupt Enable Bit blocks | **PLIC Interrupt Enable Bits Memory Map** 0x002000: Interrupt Source #0 to #31 Enable Bits on context 0 ... 0x00207C: Interrupt Source #992 to #1023 Enable Bits on context 0 0x002080: Interrupt Source #0 to #31 Enable Bits on context 1 ... 0x0020FC: Interrupt Source #992 to #1023 Enable Bits on context 1 0x002100: Interrupt Source #0 to #31 Enable Bits on context 2 ... 0x00217C: Interrupt Source #992 to #1023 Enable Bits on context 2 0x002180: Interrupt Source #0 to #31 Enable Bits on context 3 ... 0x0021FC: Interrupt Source #992 to #1023 Enable Bits on context 3 ... ... ... 0x1F1F80: Interrupt Source #0 to #31 on context 15871 ... 0x1F1FFC: Interrupt Source #992 to #1023 on context 15871 5.1. Interrupt Pending Bits ==================== ## [](#5-1-interrupt-pending-bits)5.1\. Interrupt Pending Bits The current status of the interrupt source pending bits in the PLIC core can be read from the pending array, organized as 32-bit register. The pending bit for interrupt ID N is stored in bit (N mod 32) of word (N/32). Bit 0 of word 0, which represents the non-existent interrupt source 0, is hardwired to zero. A pending bit in the PLIC core can be cleared by setting the associated enable bit then performing a claim. The base address of Interrupt Pending Bits block within PLIC Memory Map region is fixed at 0x001000. | **PLIC Register Block Name** | **Function** | **Register Block Size in Byte** | **Description** | | ---------------------------- | -------------------------------------------------- | ------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | Interrupt Pending Bits | Interrupt Pending Bit of Interrupt Source #0 to #N | 1024 / 8 = 128(0x80) bytes | This is a continuously memory block contains PLIC Interrupt Pending Bits. Each Interrupt Pending Bit occupies 1-bit from this register block. | **PLIC Interrupt Pending Bits Memory Map** 0x001000: Interrupt Source #0 to #31 Pending Bits ... 0x00107C: Interrupt Source #992 to #1023 Pending Bits 3.1. Memory Map ==================== ## [](#3-1-memory-map)3.1\. Memory Map The `base address of PLIC Memory Map` is platform implementation-specific. The memory-mapped registers specified in this chapter have a width of 32-bits. The bits are accessed atomically with LW and SW instructions. **PLIC Memory Map** base + 0x000000: Reserved (interrupt source 0 does not exist) base + 0x000004: Interrupt source 1 priority base + 0x000008: Interrupt source 2 priority ... base + 0x000FFC: Interrupt source 1023 priority base + 0x001000: Interrupt Pending bit 0-31 base + 0x00107C: Interrupt Pending bit 992-1023 ... base + 0x002000: Enable bits for sources 0-31 on context 0 base + 0x002004: Enable bits for sources 32-63 on context 0 ... base + 0x00207C: Enable bits for sources 992-1023 on context 0 base + 0x002080: Enable bits for sources 0-31 on context 1 base + 0x002084: Enable bits for sources 32-63 on context 1 ... base + 0x0020FC: Enable bits for sources 992-1023 on context 1 base + 0x002100: Enable bits for sources 0-31 on context 2 base + 0x002104: Enable bits for sources 32-63 on context 2 ... base + 0x00217C: Enable bits for sources 992-1023 on context 2 ... base + 0x1F1F80: Enable bits for sources 0-31 on context 15871 base + 0x1F1F84: Enable bits for sources 32-63 on context 15871 base + 0x1F1FFC: Enable bits for sources 992-1023 on context 15871 ... base + 0x1FFFFC: Reserved base + 0x200000: Priority threshold for context 0 base + 0x200004: Claim/complete for context 0 base + 0x200008: Reserved ... base + 0x200FFC: Reserved base + 0x201000: Priority threshold for context 1 base + 0x201004: Claim/complete for context 1 ... base + 0x3FFF000: Priority threshold for context 15871 base + 0x3FFF004: Claim/complete for context 15871 base + 0x3FFF008: Reserved ... base + 0x3FFFFFC: Reserved Sections below describe the control register blocks of PLIC operation parameters. 2.1. RISC-V PLIC Operation Parameters ==================== ## [](#2-1-risc-v-plic-operation-parameters)2.1\. RISC-V PLIC Operation Parameters General PLIC operation parameter register blocks are defined in this spec, those are: * **Interrupt Priorities registers:** The interrupt priority for each interrupt source. * **Interrupt Pending Bits registers:** The interrupt pending status of each interrupt source. * **Interrupt Enables registers:** The enablement of interrupt source of each context. * **Priority Thresholds registers:** The interrupt priority threshold of each context. * **Interrupt Claim registers:** The register to acquire interrupt source ID of each context. * **Interrupt Completion registers:** The register to send interrupt completion message to the associated gateway. Below is the figure of PLIC Operation Parameter Block Diagram, PLIC Operation Parameter Block Diagram ![PLICArch](_images/PLICArch.jpg) 4.1. Interrupt Priorities ==================== ## [](#4-1-interrupt-priorities)4.1\. Interrupt Priorities Interrupt priorities are small unsigned integers, with a platform-specific maximum number of supported levels. The priority value 0 is reserved to mean "never interrupt", and interrupt priority increases with increasing integer values. Each global interrupt source has an associated interrupt priority held in a memory-mapped register. Different interrupt sources need not support the same set of priority values. A valid implementation can hardwire all input priority levels. Interrupt source priority registers should be WARL fields to allow software to determine the number and position of read-write bits in each priority specification, if any. To simplify discovery of supported priority values, each priority register must support any combination of values in the bits that are variable within the register, i.e., if there are two variable bits in the register, all four combinations of values in those bits must operate as valid priority levels. If PLIC supports Interrupt Priorities, then each PLIC interrupt source can be assigned a priority by writing to its 32-bit memory-mapped `priority` register. A priority value of 0 is reserved to mean "never interrupt" and effectively disables the interrupt. Priority 1 is the lowest active priority while the maximum level of priority depends on PLIC implementation. Ties between global interrupts of the same priority are broken by the Interrupt ID; interrupts with the lowest ID have the highest effective priority. The base address of Interrupt Source Priority block within PLIC Memory Map region is fixed at 0x000000. | **PLIC Register Block Name** | **Function** | **Register Block Size in Byte** | **Description** | | ---------------------------- | ------------------------------------- | ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Interrupt Source Priority | Interrupt Source Priority #0 to #1023 | 1024 \* 4 = 4096(0x1000) bytes | This is a continuously memory block which contains PLIC Interrupt Source Priority. Total 1024 Interrupt Source Priority in this memory block. Interrupt Source Priority #0 is reserved which indicates it does not exist. | **PLIC Interrupt Source Priority Memory Map** 0x000000: Reserved (interrupt source 0 does not exist) 0x000004: Interrupt source 1 priority 0x000008: Interrupt source 2 priority ... 0x000FFC: Interrupt source 1023 priority 7.1. Priority Thresholds ==================== ## [](#7-1-priority-thresholds)7.1\. Priority Thresholds PLIC provides context based `threshold register` for the settings of a interrupt priority threshold of each context. The `threshold register` is a WARL field. The PLIC will mask all PLIC interrupts of a priority less than or equal to `threshold`. For example, a `threshold` value of zero permits all interrupts with non-zero priority. The base address of Priority Thresholds register block is located at 4K alignment starts from offset 0x200000. | **PLIC Register Block Name** | **Function** | **Register Block Size in Byte** | **Description** | | ---------------------------- | ------------------------------------- | ----------------------------------------- | -------------------------------------------------------------------- | | Priority Threshold | Priority Threshold for 15872 contexts | 4096 \* 15872 = 65011712(0x3e00000) bytes | This is the register of Priority Thresholds setting for each context | **PLIC Interrupt Priority Thresholds Memory Map** 0x200000: Priority threshold for context 0 0x201000: Priority threshold for context 1 0x202000: Priority threshold for context 2 0x203000: Priority threshold for context 3 ... ... ... 0x3FFF000: Priority threshold for context 15871 RISC-V Server SoC Specification ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-server-soc-specification)RISC-V Server SoC Specification Server SoC Task Group Version v1.0, 2025-02-21: Ratified | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at. Copyright 2022-2025 by RISC-V International. Bibliography ==================== ## [](#bibliography)Bibliography \[1\] _RISC-V Boot and Runtime Services Specification (BRS)_, . \[Online\]. Available: \[2\] _RISC-V Security Model_, . \[Online\]. Available: \[3\] _Key words for use in RFCs to Indicate Requirement Levels_. \[Online\]. Available: \[4\] _PCI Express® Base Specification Revision 6.0_, . \[Online\]. Available: \[5\] _Advanced Configuration and Power Interface (ACPI) Specification_. \[Online\]. Available: \[6\] _Unified Extensible Firmware Interface_. \[Online\]. Available: \[7\] _RISC-V Instruction Set Manual, Volume I: Unprivileged Architecture_, . \[Online\]. Available: \[8\] _RISC-V Advanced Interrupt Architecture_. \[Online\]. Available: \[9\] _RISC-V IOMMU Architecture Specification_. \[Online\]. Available: \[10\] _PCI Code and ID Assignment Specification Revision 1.1_, . \[Online\]. Available: \[11\] _RISC-V RAS error record register interface_. \[Online\]. Available: \[12\] _RISC-V Capacity and Bandwidth QoS Register Interface_. \[Online\]. Available: \[13\] _RISC-V Instruction Set Manual, Volume II: Privileged Architecture_, . \[Online\]. Available: \[14\] _Redfish specification 1.18.0_. \[Online\]. Available: \[15\] _PLDM base specification 1.1.0_. \[Online\]. Available: \[16\] _MCTP base specification 1.3.1_. \[Online\]. Available: \[17\] _Security protocol and data model (SPDM) specification 1.2.1_. \[Online\]. Available: \[18\] _Secured messages using SPDM specification 1.1.0_. \[Online\]. Available: \[19\] _Intelligent Platform Management Interface (IPMI) 2.0_. \[Online\]. Available: \[20\] _Datacenter Secure Control Module Specification_. \[Online\]. Available: \[21\] _TPM 2.0 Library_. \[Online\]. Available: Contributors ==================== ## [](#contributors)Contributors This RISC-V specification has been contributed to directly or indirectly by (in alphabetical order): Aaron Durbin, Andrea Bolognani, Andrei Warkentin, Andrew Jones, Beeman Strong, Cameron McNairy, Greg Favor, Heinrich Schuchardt, Isaac Chute, Jon Masters, Ken Dockser, Krste Asanovic, Manu Gulati, Mark Hayter, Michael Klinglesmith, Paul Walmsley, Ravi Sahita, Shaolin Xie, Shubu Mukherjee, Sibaranjan Pattnayak, Ved Shanbhogue 1.1. Introduction ==================== ## [](#intro)1.1\. Introduction The RISC-V server SoC specification defines a standardized set of capabilities that portable system software such as operating systems and hypervisors, can rely on being present in a RISC-V server SoC. A server is a computing system designed to manage and distribute resources, services, and data to other computers or devices on a network. It is often referred to as a 'server' because it serves or provides information and resources upon request. Such computing systems are designed to operate continually and have higher requirements for capabilities such as RAS, security, performance, and quality of service. Examples of servers include web servers, file servers, database servers, mail servers, game servers, and more. This specification focuses on defining requirements for general-purpose server computing systems that may be used for one or more of these purposes. ![riscv server platform](_images/riscv-server-platform.png) Figure 1\. Components of a RISC-V Server Platform The RISC-V server platform is defined as the collection of SoC hardware, platform firmware, boot/runtime services, and security services. The platform provides hardware interfaces (e.g., harts, timers, interrupt controllers, PCIe root ports, etc.) to portable system software. It also offers a set of standardized RISC-V boot and runtime services \[[1](server%5Fsoc%5Fbibliography.html#bib-brs)\] based on the UEFI and ACPI standards. To support provisioning and platform management, it interfaces with a baseboard management controller (BMC) through both in-band and out-of-band (OOB) management interfaces. The in-band management interfaces support the use of standard manageability specifications like MCTP, PLDM, IPMI, and Redfish for provisioning and management of the operating system executing on the platform. The OOB interface supports the use of standard manageability specifications like MCTP, PLDM, Redfish, and IPMI for functions such as power management, telemetry, debug, and provisioning. The RISC-V security model \[[2](server%5Fsoc%5Fbibliography.html#bib-sec)\] includes guidelines and requirements for aspects such as debug authorization, secure/measured boot, firmware updates, firmware resilience, and confidential computing, among others. The platform firmware, typically operating at privilege level M, is considered part of the platform and is usually expected to be customized and tailored to meet the requirements of the SoC hardware (e.g., initialization of address decoders, memory controllers, RAS, etc.). This specification standardizes the requirements for the hardware interfaces and capabilities (e.g., timers, interrupt controllers, PCIe root complexes, RAS, QoS, in-band management, etc.) provided by the SoC to software executing on the application processor harts at privilege levels below M. It enables OS and hypervisor vendors to support such SoCs with a single binary OS image distribution model. The requirements posed by this specification represent a standard set of infrastructural capabilities, encompassing areas where divergence is typically unnecessary and where novelty is absent across implementations. To be compliant with this specification, the SoC MUST support all mandatory rules and MUST support the listed versions of the specifications. This standard set of capabilities MAY be extended by a specific implementation with additional standard or custom capabilities, including compatible later versions of listed standard specifications. Portable system software MUST support the specified mandatory capabilities to be compliant with this specification. The rules in this specification use the following format: | ID# | Rule | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | CAT\_NNN | The CAT is a category prefix that logically groups the rules and is followed by 3 digits - NNN \- assigning a numeric ID to the rule. The rules use the key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" that are to be interpreted as described in RFC 2119 \[[3](server%5Fsoc%5Fbibliography.html#bib-rfc%5F2119)\] when, and only when, they appear in all capitals, as shown here. When these words are not capitalized, they have their normal English meanings. | | _A rule or a group of rules may be followed by non-normative text providing context or justification for the rule. The non-normative text may also be used to reference sources that are the origin of the rule._ | | This specification groups the rules in the following broad categories: * Clocks and Timers * Interrupt Controllers * IOMMU * PCIe subsystem * Reliability, Availability, and Serviceability * Quality of Service * Performance monitoring * Security ### [](#1-1-1-glossary)1.1.1\. Glossary Most terminology has the standard RISC-V meaning. This table captures other terms used in the document. Terms in the document prefixed by 'PCIe' have the meaning defined in the PCI Express (PCIe) Base Specification \[[4](server%5Fsoc%5Fbibliography.html#bib-pci)\] (even if they are not in this table). __Table 1\. Terms and definitions__ | Term | Definition | | ---------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ACPI | Advanced Configuration and Power Interface \[[5](server%5Fsoc%5Fbibliography.html#bib-acpi)\]. | | ACS | Follows PCI Express. Access Control Services. A set of capabilities used to provide controls over routing of PCIe transactions. | | AER | Advanced Error Reporting. Follows PCI Express. A PCIe defined error reporting paradigm. | | AIA | RISC-V Advanced Interrupt Architecture. | | ATS | Follows PCI Express. Address Translation Services. | | BAR or Base Address Register | Follows PCI Express. A register that is used by hardware to show the amount of system memory needed by a PCIe function and used by system software to set the base address of the allocated space. | | BMC | Baseboard Management Controller. A motherboard resident management controller that provides functions for platform management. | | CXL | Compute Express Link bus standard. | | DMA | Direct Memory Access. | | DMTF | Distributed Management Task Force. Industry association for promoting systems management and interoperability. | | ECAM | Follows PCI Express. Enhanced Configuration Access Method. A mechanism to allow addressing of Configuration Registers for PCIe functions. In addition to the PCI Express Base Specification, see the detailed rules in this specification. | | EP, EP=1 | Follows PCI Express. Also called Data Poisoning. EP is an error flag that accompanies data in some PCIe transactions to indicate the data is known to contain an error. Defined in PCI Express Base Specification 6.0 section 2.7.2\. Unless otherwise blocked, the poison associated with the data must continue to propagate in the SoC internal interconnect. | | GPA | Guest Physical Address: An address in the virtualized physical memory space of a virtual machine. | | Guest | Software in a virtual machine. | | Hierarchy ID or Segment ID | Follows PCI Express. An identifier of a PCIe Hierarchy within which the Requester IDs are unique. | | Host Bridge | Part of a SoC that connects host CPUs and memory to PCIe root ports, RCiEP, and non-PCIe devices integrated in the SoC. The host bridge is placed between the device(s) and the platform interconnect to process DMA transactions. IO Devices may perform DMA transactions using IO Virtual Addresses (VA, GVA or GPA). The host bridge invokes the associated IOMMU to translate the IOVA to Supervisor Physical Addresses (SPA). | | HPM | Hardware Performance Monitor. | | Hypervisor | Software entity that controls virtualization. | | ID | Identifier. | | IMSIC | Incoming Message-signaled Interrupt Controller. | | IO Bridge | See host bridge. | | IOVA | I/O Virtual Address: Virtual address for DMA by devices. | | MCTP | Follows DMTF Standard. Management Component Transport Protocol used for communication between components of a platform management system. | | MSI | Message Signaled Interrupts. | | NUMA | Non-uniform memory access. | | OS | Operating System. | | PASID | Follows PCI Express. Process Address Space Identifier: It identifies the address space of a process. The PASID value is provided in the PASID TLP prefix of the request. | | PBMT | Page-Based Memory Types. | | PRI | Page Request Interface. Follows PCI Express. A PCIe protocol that enables devices to request OS memory manager services to make pages resident. | | RCiEP | Root Complex Integrated Endpoint. Follows PCI Express. An internal peripheral that enumerates and behaves as specified in the PCIe standard. | | RCEC | Follows PCI Express. Root Complex Event Collector. A block for collecting errors and PME messages in a standard way from various internal peripherals. | | RID or Requester ID | Follows PCI Express. An identifier that uniquely identifies the requester within a PCIe Hierarchy. Needs to be extended with a Hierarchy ID to ensure it is unique across the platform. | | Root Complex, RC | Follows PCI Express. Part of the SoC that includes the Host Bridge, Root Port, and RCiEP. | | Root Port, RP | Follows PCI Express. A PCIe port in a Root Complex used to map a Hierarchy Domain using a PCI-PCI bridge. | | P2P or peer-to-peer | Follows PCI Express. Transfer of data directly from one device to another. If the devices are under different PCIe Root Ports or are internal to the SoC this may involve data movement across the SoC internal interconnect. | | PLDM | Follows DMTF standard. Platform Level Data Model. | | PMA | Physical Memory Attributes. | | PMP | Physical Memory Protection. | | Significant Cache | A large cache that might have significant impact on performance. This specification recommendeds that a cache with a capacity larger than 32 KiB be considered a significant cache if it has a significant impact on performance. | | SMBIOS | System Management BIOS. | | SoC | System on a chip, also referred as system-on-a-chip and system-on-chip. | | SPA | Supervisor Physical Address: Physical address used to to access memory and memory-mapped resources. | | SPDM | Follows DMTF Standard. Security Protocols and Data Models. A standard for authentication, attestation and key exchange to assist in providing infrastructure security enablement. | | SR-IOV | Follows PCI Express. Single-Root I/O Virtualization. | | TLP | Follows PCI Express. Transaction Layer Packet. Defined by Chapter 2 of the PCI Express Base Specification. | | QoS | Quality of Service. Quality of Service (QoS) is defined as the minimal end-to-end performance that is guaranteed in advance by a service level agreement (SLA) to a workload. | | UEFI | Unified Extensible Firmware Interface. \[[6](server%5Fsoc%5Fbibliography.html#bib-uefi)\] | | UR, CA | Follows PCI Express. Error returns to an access made to a PCIe hierarchy. | | VM | Virtual Machine. | 2.1. Server SoC Requirements ==================== ## [](#2-1-server-soc-requirements)2.1\. Server SoC Requirements ### [](#2-1-1-clocks-and-timers)2.1.1\. Clocks and Timers | ID# | Rule | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | CTI\_010 | The time CSR MUST increment at a constant frequency and the count MUST be in units of 1 ns. The frequency at which the CSR provides an updated time value MUST be at least 100 MHz. | | _The Zicntr extension \[[7](server%5Fsoc%5Fbibliography.html#bib-unpriv)\] requires the real-time clocks of all harts to be synchronized to within one tick of the real-time clock._ | | | CTI\_020 | The time counter MUST appear to be always on and MUST appear to not lose its count across hart low power idle states, including when the hart is powered off. | | _This rule does not apply to system power states such as G3 (power off), S3 (Suspend to RAM), or S4 (Hibernate)._ _Losing time count across hart low power idle states may lead to the hart losing time synchronization with other application processor harts, potentially causing unexpected behaviors and/or system instability._ _Information about whether a hart low power idle state retains timer context may be determined by the OS/hypervisors using information provided by the ACPI \_LPI object or equivalent mechanisms._ | | ### [](#2-1-2-interrupt-controllers)2.1.2\. Interrupt Controllers This section specifies the requirements on the interrupt controllers used to deliver external interrupts to the RISC-V application processor harts. | ID# | Rule | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | IIC\_010 | The RISC-V Advanced Interrupt Architecture \[[8](server%5Fsoc%5Fbibliography.html#bib-aia)\] MUST be supported. | | IIC\_020 | External interrupts MUST be signaled to a hart as message-signaled interrupts (MSI). | | _Since Message Signaled Interrupts (MSI) are implemented as memory writes, they facilitate a simplified enforcement of producer-consumer ordering rules. Specifically, interrupts issued by a device following a write operation must be processed only after the previous write operations have been completed and observed. Similarly, interrupts issued by a device must be observed before any subsequent read completions generated by the device._ _MSI is the preferred mechanism for interrupt signaling in PCIe due to its efficiency and support for low-latency communication between devices and harts. By adopting MSI, systems can achieve faster and more reliable interrupt handling, essential for high-performance computing environments._ | | | IIC\_030 | The Incoming Message-signaled Interrupt Controller (IMSIC) MUST implement an interrupt file for S-mode. | | IIC\_040 | The IMSIC MUST support at least 5 VS-mode interrupt files. | | _Supporting 5 VS-mode interrupt files for a hart allows context switching between up to 5 virtual CPUs (vCPU) on a hart without needing to swap the contents of the interrupt file out to memory. This is particularly beneficial when devices are directly assigned to virtual machines (VMs), as swapping out the context of an IMSIC interrupt file may result in longer latencies due to the need to redirect device interrupts to a memory-resident interrupt file._ | | | IIC\_050 | The S-mode interrupt file MUST support at least 255 interrupt identities. | | IIC\_060 | The VS-mode interrupt files MUST support at least 63 interrupt identities. | | IIC\_070 | The memory regions designated for IMSIC interrupt files MUST have the following PMAs: Not cacheable, non-idempotent, coherent, strongly-ordered (I/O ordering) channel 0 I/O region Support for 4-byte aligned reads and writes. | | IIC\_080 | If the SoC implements devices that use wire-signaled interrupts then the SoC MUST implement an APLIC as specified by the RISC-V AIA specification and MUST use the APLIC to convert the wire-signaled interrupts into MSIs. If implemented, the APLIC MUST support: Supervisor interrupt domain. GEILEN values matching those implemented by the harts. MSI delivery mode. Extempore MSI generation using the genmsi register. | | _SoC devices using wire-signaled interrupts must implement the rules related to ordering of interrupts vs. older read/writes from devices as specified by the device and/or bus interface specifications that such devices conform to. See also SID\_010._ | | ### [](#IOMMU)2.1.3\. Input-Output Memory Management Unit (IOMMU) | ID# | Rule | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | IOM\_010 | All IOMMUs in the SoC MUST support the RISC-V IOMMU specification \[[9](server%5Fsoc%5Fbibliography.html#bib-iommu)\]. | | IOM\_020 | All DMA capable peripherals (RCiEP and non-PCIe devices) and all PCIe root ports accessible by software on the RISC-V application processor harts MUST be governed by an IOMMU. Initiators, such as the following, are exempt from this rule: Interrupt controllers, such as the APLIC. IOMMUs. System Bus Access blocks of Debug Modules. Controllers, including the root of trust (RoT) controllers, power management controllers, or other SoC management controllers, when they access resources reserved for their use. | | _DMA capable peripherals being governed by an IOMMU allows OS/hypervisors to restrict DMA originating from such devices to a subset of memory to enhance security and software fault tolerance. The address translation capability provided by the IOMMU enables usages such as passthrough of such devices to virtual machines, shared virtual addressing, etc._ _The number of IOMMUs implemented in the SoC to satisfy rule IOM\_020 is UNSPECIFIED._ | | | IOM\_030 | The IOMMU governing a PCIe root port MUST support at least 16-bit wide device IDs. | | IOM\_040 | An IOMMU that does not govern a PCIe root port MUST support a device ID width required to support all requester IDs originated by the devices governed by that IOMMU. | | IOM\_050 | The IOMMU MUST implement all the page based virtual memory system modes and extensions that are implemented by the RISC-V application processor harts in the SoC. | | _The page based virtual memory system modes supported by the IOMMU are enumerated in the IOMMU capabilities register._ | | | IOM\_070 | The IOMMU SHOULD support pass-through mode and MRIF mode MSI address translation. | | IOM\_080 | When MRIF mode MSI address translation is supported, the IOMMU MUST support atomic updates to the MRIF (enumerated by 1 setting ofcapabilities.AMO\_MRIF). | | IOM\_090 | IOMMU governing PCIe root ports SHOULD support PCIe address translation services (ATS). | | _High performance devices such as DPU/SmartNICs, GPUs, and FPGAs, utilized in server platforms rely on ATS and Page Request services to achieve high throughput and low-latency I/O. Supporting ATS is also required for efficiently accommodating usage models such as Shared Virtual Addressing and direct work submission from user mode._ | | | IOM\_100 | IOMMU governing PCIe root ports SHOULD support the T2GPA mode of operation with ATS if ATS is supported. | | _The T2GPA control enables a hypervisor to prevent DMA from a device, even if the device misuses the ATS capability and attempts to access memory that is not explicitly authorized by the page tables governing that device’s memory accesses. The threat model could also include a man-in-the-middle on the PCIe link inserting ATS-translated requests to access memory that was not previously authorized. As an alternative to setting T2GPA to 1, the hypervisor might establish a trust relationship with the device if authentication protocols such as SPDM are supported by the device. For PCIe, for example, the PCIe Component Measurement and Authentication (CMA) capability provides a mechanism to verify the device’s configuration and firmware/executable (Measurement) and hardware identities (Authentication). This mechanism establishes such a trust relationship, and the PCIe link may be integrity-protected using PCIe integrity and data encryption (IDE) to defend against a man-in-the-middle adversary._ | | | IOM\_110 | IOMMU governing RCiEP MUST support PCIe address translation services (ATS) if any of the RCiEPs governed by the IOMMU support the ATS capability. | | IOM\_120 | IOMMU governing RCiEP MAY support the T2GPA mode of operation with ATS if ATS is supported. | | _The threats associated with misuse of ATS or malicious insertion of ATS translated requests by a man-in-the-middle may not be present with RCiEP being integrated in the SoC._ | | | IOM\_130 | IOMMU MUST support MSI and MAY support wire-signaled interrupts for external interrupts originated by the IOMMU itself. | | IOM\_140 | IOMMU MUST support little-endian memory access to its in-memory data structures. | | IOM\_150 | IOMMU MAY support big-endian mode memory access to its in-memory data structures. | | _The IOMMU memory-mapped registers always have a little-endian byte order._ | | | IOM\_160 | IOMMU MAY support the PCIe PASID capability. | | IOM\_170 | IOMMU that supports PASID capability MUST support 20-bit PASID width and MAY support 8-bit and 17-bit PASID widths. | | _PCIe specification strongly recommends that hardware implement the maximum width of 20 bits to ensure interoperability with system software. See also the implementation note on PASID width homogeneity in the PCIe specification 6.0 section 6.20.2.2._ | | | IOM\_180 | IOMMU SHOULD support a hardware performance monitor (HPM). | | _The HPM is a valuable tool for system integrators for performance monitoring and optimizations. An IOMMU is highly recommended to provide an HPM._ | | | IOM\_190 | An IOMMU that supports an HPM MUST support the cycles counter. | | IOM\_200 | An IOMMU that supports an HPM MUST incorporate at least 4 event counters. | | _A typical performance analysis operation may involve simultaneously counting the number of translation requests, IOATC misses, and page table walks. An HPM with sufficient number of event counters ensures accurate and comprehensive data collection, enabling detailed performance analysis and optimization._ | | | IOM\_210 | The cycles counter and the event counters MUST be at least 40 bits wide. | | IOM\_220 | The IOMMU SHOULD support the software debug capabilities enumerated by DBG field in the capabilities register. | | IOM\_230 | The physical address width supported by the IOMMU MUST be greater than or equal to the physical address width supported by the RISC-V application processor harts in the SoC. | | _Having the physical address width greater than or equal to the width supported the harts in the SoC enables use of all addressable memory for I/O and facilitates the sharing of page tables between the hart MMU and the IOMMU._ | | | IOM\_240 | The reset default of the iommu\_mode MUST be Off. | | _The IOMMU disallowing DMA unconditionally following reset due to the mode being Off allows the SoC firmware and software to enable DMA when suitable security protections as required have been established. The IOMMU mode being Off at reset does not pose a significant issue to SoC firmware that needs to employ DMA (e.g., for firmware loading) as that firmware may program the mode in the appropriate IOMMU prior to programming the peripheral governed by that IOMMU to perform a DMA._ | | | IOM\_250 | An IOMMU that is implemented as an RCiEP MUST use base class 08H and subclass 06H \[[10](server%5Fsoc%5Fbibliography.html#bib-pci-cls)\]. | | _The base class 08H and sub-class 06H are designated by PCIe for use by an IOMMU. Implementing the IOMMU as a PCIe device allows an operating system to determine a driver for the IOMMU and to assign resources such as interrupt vectors to the IOMMU in a PCIe compatible manner._ | | | IOM\_260 | The host bridge MUST enforce the physical memory attribute checks and physical memory protection checks on memory accesses originated by the IOMMU and signal detected access violations to the IOMMU. | | _These checks are analogous to the PMA and PMP checks performed by the RISC-V hart. The host bridge (also known as IO bridge) invokes the IOMMU for address translations. To perform the operations requested by the host bridge the IOMMU may need to access in-memory data structures such as the device directory table and page tables. The physical memory protection limit access from IOMMUs to phusical addresses to support secure processing and contain faults. These checks allow restricting the IOMMU to only have access to the same memory that the hart software that programs the IOMMU has access to. The IOMMU specification requires an IOMMU to support locating IOMMUs in-memory data structures, in-memory queues, and page tables in memory address ranges that hold main memory. Support for locating these in I/O memory is not required._ | | | IOM\_270 | An IOMMU MUST support 24-bit device IDs if the IOMMU governs multiple PCIe root ports that may be part of different PCIe hierarchies. | | _An IOMMU governing PCIe root ports uses requester ID (RID) - the tuple of bus/device/function numbers (or just bus/function numbers, if the PCIe ARI option is used) - to locate a device context to use for address translation and protection. The 16-bit RID uniquely identifies a requester within a hierarchy. This RID needs to be augmented with the Hierarchy ID (also known as segment ID) - an 8-bit number - to uniquely identify a requester across PCIe hierarchies._ | | | IOM\_280 | The host bridge MUST provide the PCIe RID as the bits 15:0 of thedevice\_id input to the IOMMU for requests from PCIe EPs and RCiEP. | | IOM\_290 | When the IOMMU supports 24-bit device IDs, the host bridge MUST specify the segment number associated with the PCIe hierarchy from which requests were received as the bits 23:16 of the device\_id to the IOMMU. | | IOM\_300 | The determination of device\_id input to an IOMMU for requests originating from non-PCIe devices is UNSPECIFIED. If PCIe and non-PCIe endpoints/RCiEP are governed by the same IOMMU, the SoC MUST ensure that there is no overlap between any device\_id associated with non-PCIe devices with any device\_id formed using the PCIe RID (and if applicable the segment ID). | | IOM\_310 | The host bridge MUST provide the 20-bit PASID from the PCIe PASID TLP Prefix as the process\_id input to the IOMMU along with an indication about the validity of the process\_id input. When theprocess\_id is indicated as valid, the host bridge MUST additionally provide the "Execute Requested" and the "Privilege Mode Requested" bits from the PASID TLP prefix as input to the IOMMU. When process\_id input is indicated as not valid, the host bridge MUST set the "Execute Requested" and "Privilege Mode Requested" inputs to 0. | | _The host bridge providing the full 20-bit value without truncation from the PASID TLP prefix to the IOMMU enables the IOMMU to determine if the PASID value is wider than supported by the current configuration of the process directory table for that device and generate a fault notification if so._ | | | IOM\_320 | The determination of process\_id, "Execute Requested", and "Privilege Mode Requested" inputs to an IOMMU for requests originating from non-PCIe devices is UNSPECIFIED. | ### [](#2-1-4-pcie-subsystem)2.1.4\. PCIe Subsystem A PCIe subsystem consists of a root complex with a collection of root ports, root complex event collectors (RCECs), root complex register blocks (RCRBs), and root complex integrated end points (RCiEPs). The root complex implements a host bridge to connect the PCIe root ports, RCECs, RCRBs, and RCiEP, to the CPU and system memory in the SoC through an interconnect. ![riscv server rc](_images/riscv-server-rc.svg) Figure 1\. PCIe root complex One or more root ports in a root complex may be part of a hierarchy where a hierarchy is a PCI Express I/O interconnect topology, wherein the Configuration Space addresses, referred to as the tuple of Bus/Device/Function Numbers (or just Bus/Function Numbers, for PCIe ARI cases), are unique. These addresses are used for Configuration Request routing, Completion routing, some Message routing, and for other purposes. In some contexts a Hierarchy is also called a Segment, and in Flit Mode, the Segment number is sometimes also included in the ID of a Function. Each root port in a hierarchy originates a hierarchy domain i.e. a part of a Hierarchy originating from a single Root Port. The root ports are PCI-PCI bridges that bridge a primary PCIe bus to a range of secondary and subordinate buses. In some SoCs, PCIe devices may be integrated in the same package/die as the root complex. Examples of such devices are network controllers, USB host controllers, NVMe controllers, AHCI controllers, etc. Such SoC integrated devices may be presented to software using one of the following options: 1. Presented to software as a PCIe endpoint (EP; See section 1.3.2.2 of the PCIe 6.0 specification) connected to a PCIe root port (See example of such an endpoint connected to root port 3 in [Figure 1](#fig:RISC-V-Server-RC)). Such PCIe endpoints must comply with the PCIe specified rules for endpoints. 2. Presented to software as a root complex integrated endpoint (RCiEP; See section 1.3.2.3 of the PCIe 6.0 specification). Such PCIe endpoints must comply with the PCIe specified rules for RCiEP. Implementing integrated devices that perform as RCiEP or EP allows the use of standardized PCIe frameworks for memory and interrupt resource allocation, virtualization (SR-IOV), ATS/PRI for shared virtual addressing, trusted IO using SPDM/TDISP, RAS frameworks like data poisoning and AER, power management, etc. The host bridge is placed between the device(s) and the system interconnect to process DMA transactions. Devices perform DMA transactions using IO Virtual Addresses (VA, GVA or GPA). The host bridge invokes the associated IOMMU to translate the IOVA to Supervisor Physical Addresses (SPA). | RCI\_010 | The PCIe root ports, host bridges, RCRBs, and RCECs in the root compplex MUST implement all software visible rules defined by the PCIe specification 6.0 for the root complex as applicable. | | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#2-1-4-1-enhanced-configuration-access-method-ecam)2.1.4.1\. Enhanced Configuration Access Method (ECAM) Each PCIe endpoint and the PCIe root port itself implement a set of memory mapped configuration registers that are accessed using the PCIe enhanced configuration access method (ECAM). The memory mapped ECAM address range for a hierarchy is up to 256 MiB in size and the base address of the range is naturally aligned to the size. Each PCIe function is associated with a 4 KiB page in this range such that the address bits (20+b):20 where b=0 to 7 identify the bus number of that function (see also recommendations in the PCIe specification 6.0 section 7.2.2), the address bits 19:15 identify the device number, and the address bits 14:12 identify the function number. The host bridge in conjunction with the SoC boot firmware maps the ECAM address range to the hierarchy domain originating at each PCIe root port. | ID# | Rule | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ECM\_010 | The ECAM address ranges MUST have the following physical memory attributes (PMAs): Not cacheable, non-idempotent, coherent, strongly-ordered (I/O ordering) channel 0 I/O region One, two, and four byte naturally aligned read and write MUST be supported and MUST result in a single PCIe Configuration Request. | | _See also the implementation note on root complex requirements for generating configuration requests in section 7.2.2 of PCIe specification 6.0._ | | | ECM\_020 | Writes to the ECAM address range from a RISC-V hart MUST be non-posted and the write MUST complete at the hart only after a completion is received from the function hosting the accessed configuration register. | | _Besides performing a write, software executing on a hart must not require any additional actions to achieve this property._ _This rule satisfies the processor and host bridge implementation requirement mentioned in the “Ordering Considerations for the Enhanced Configuration Access Mechanism” implementation note of the PCIe 6.0 specification._ | | | ECM\_030 | The ECAM address range for a hierarchy MUST be contiguous and the base address of the range MUST be naturally aligned to the size of the ECAM address range associated with the hierarchy. | | ECM\_040 | A SoC MAY support multiple hierarchies. When multiple hierarchies are supported, the ECAM address range of the hierarchies MUST NOT overlap, but they are not required to be contiguous. | | ECM\_050 | The configuration space of the PCIe root ports MUST be associated with the primary bus number of the hierarchy associated with the root port. | | _PCIe root ports are PCI-PCI bridges that bridge the primary bus to the secondary/subordinate buses. The root port itself enumerates as a PCI-PCI bridge device on the primary bus. The collection of primary, secondary, and subordinate buses are part of a single hierarchy domain that originates at that PCIe root port._ | | | ECM\_060 | The configuration space of functions on the primary bus MUST be accessible irrespective of the state of the corresponding PCIe link. | | _Discovery and activation of the PCIe link requires accessing the configuration space registers of the PCIe root port itself and the PCIe root port is a PCI-PCI bridge device on the primary bus._ | | | ECM\_070 | The PCIe root port MUST support the PCIe Configuration RRS software (CRS) visibility enable control. | | _The number of times a configuration request is retried on an RRS response is UNSPECIFIED._ | | | ECM\_080 | Read and/or write to the ECAM range of the hierarchy domain originating at a root port MUST generate PCIe configuration transactions as type 0 or type 1 configuration transactions following the rules specified for ECAM in PCIe specification 6.0. | | _Determination of the type of configuration transaction based on whether the access is to the primary, secondary or subordinate buses may involve logic in the host bridge to work in conjunction with the root port PCIe controller. See also Alternative Routing-ID Interpretation in PCIe specification 6.0 section 6.13 for rules related to converting type 1 configuration requests into type 0 configuration request based on the traditional Device Number field being 0\. Specifically, when ARI forwarding is disabled, write accesses to configuration space of Device Number greater than 0 must be silently dropped, and read accesses must be responded to with all 1s data._ | | | ECM\_090 | Read access to ECAM address range from a RISC-V hart MUST be responded with all 1s data if any of the following conditions are TRUE: Access is to a non-existent function on the primary bus of a hierarchy domain. Accessed bus is not part of any of the hierarchy domains. An Unsupported Request or Completer Abort response was received. A completion timeout occurs. Access targets a function downstream of a root port whose link is not in DL\_Active state. A PCIe RRS response was received on each retry of the configuration read and CRS software visibility is not enabled. PCIe CRS software visibility is enabled, but the access does not target the vendor ID register, and a RRS response was received on each retry of the configuration read. | | _The data response to the Vendor ID register on receipt of an RRS response MUST follow the PCIe defined rules. See also the recommendations in PCIe specification 6.0 section 2.3.2._ | | | ECM\_100 | Write access from a RISC-V hart to configuration registers of a non-existent function on the primary bus MUST be dropped (silently ignored or discarded) and the write completed. Such accesses MUST NOT lead to any other behavior (e.g., hangs, deadlocks, etc.). | | ECM\_110 | Poisoned data received from completers (EP=1) MUST be forwarded to the requesting RISC-V hart as poisoned data unless such forwarding is disallowed (e.g., SoC does not support data poisoning or forwarding of poisoned data is disabled though implementation defined means). If forwarding of poisoned data is disallowed then the poisoned data MUST be replaced with all 1s data. | #### [](#2-1-4-2-pcie-memory-space)2.1.4.2\. PCIe Memory Space | ID# | Rule | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | MMS\_010 | The SoC MUST support designating, for each hierarchy domain, one or more ranges of system physical addresses that may be used for mapping memory space of endpoints in that hierarchy domain using the 64-bit wide base address registers (BARs) of the endpoints. | | MMS\_020 | SoC MUST support designating, for each hierarchy domain, at least one system physical address range for mapping memory space of endpoints in that hierarchy domain using 32-bit wide BARs of the endpoint. | | _The ranges suitable for mapping using 32-bit BARs are also sometimes termed as the low MMIO ranges and those suitable for use with 64-bit BARs termed as high MMIO ranges. \_The bit 3 of the Base Address Register used to called the “Prefetchable” bit and required PCIe functions to support 64-bit addressing for any BAR that requested "Prefetchable" memory space. The "Removing Prefetchable Terminology" ECN [Removing Prefetchable Terminology ECN](server%5Fsoc%5Fbibliography.html#PCI%5FPREF) reworks the PCIe Base Specification to remove Prefetchable terminology to more accurately reflect modern device and system requirements._ | | | MMS\_030 | The system physical address ranges designated for mapping endpoint memory spaces have the following physical memory attribute (PMAs): MUST be not cacheable, non-idempotent, coherent, strongly-ordered (I/O ordering) channel 0 I/O region. MUST support all aligned and unaligned access sizes that can be generated by data requests from any of the RISC-V application processor harts in the SoC or by peer endpoints, including those of type RCiEP. MAY support atomics, instruction fetch, and page walks. Naturally aligned data requests of size up to 8 bytes from the RISC-V application processor harts in the SoC or by peer endpoints, including those of type RCiEP, MUST result in a single PCIe Memory Request to the target device. | | _Software may use the Svpbmt extension to override the PMA to NC if such an override is compatible with the restricted programming model of the device._ _See also the implementation note on optimizations based on restricted programming mode in section 2.3.1 of PCIe specification 6.0._ _See also first/last DW byte enable rules in section 2.2.5 of PCIe specification 6.0._ | | | MMS\_040 | A load from a RISC-V application processor hart to memory ranges designated for mapping memory spaces of endpoints or RCiEP MUST complete with an all 1s response and MUST NOT lead to any abnormal behavior (e.g., hangs, deadlocks, etc.) if any of the following are TRUE: Address is not within any of the following address ranges: Address range defined by memory base/limit or 64-bit memory base/limit registers of any root port. BAR (including when EA capability is used) mapped range of any RCiEP. BAR (including when EA capability is used) mapped range of any root port. The PCIe link of the root port to which the access is routed is not active. Including due to the root port entering downstream port containment state. A UR or a CA response is received from the completer. A completion timeout occurs. | | _The 64-bit memory base/limit register was previously called Prefetchable Memory Base/Limit. The concept of “Prefetchable” MMIO was originally needed to control PCI-PCI Bridges, which were allowed/encouraged to prefetch Memory Read data in prefetchable regions. The original intent of the Prefetchable/Non-Prefetchable distinction was focused on PCI behaviors, and was not intended for software use in determining memory attributes and/or coding techniques. The "Removing Prefetchable Terminology" ECN[Removing Prefetchable Terminology ECN](server%5Fsoc%5Fbibliography.html#PCI%5FPREF) reworks the PCIe Base Specification to remove Prefetchable terminology._ _See also the implementation note on optimizations based on restricted programming mode in section 2.3.1 of PCIe specification 6.0._ | | | MMS\_050 | A store from a RISC-V application processor hart to memory ranges designated for mapping memory space of endpoints or RCiEP MUST be dropped (silently ignored or discarded) and MUST NOT lead to any abnormal behavior (e.g., hangs, deadlocks, etc.) if any of the following are TRUE: Address is not within any of the following address ranges: Address range defined by memory base/limit or 64-bit memory base/limit registers of any root port. BAR (including when EA capability is used) mapped range of any RCiEP. BAR (including when EA capability is used) mapped range of any root port. The PCIe link of the root port to which the access is routed is not active. Including due to the root port entering downstream port containment state. | | MMS\_060 | Poisoned data received from completers (EP=1) MUST be forwarded to the requester PCIe device (a RCiEP or an endpoint) as poisoned data unless such forwarding is disallowed (e.g., poisoned TLP egress blocking). | | MMS\_070 | Poisoned data received from completers (EP=1) MUST be forwarded to a requester RISC-V hart as poisoned data unless such forwarding is disallowed through implementation defined means. When such forwarding is disallowed, then the poisoned data MUST be replaced with all 1s data. | | MMS\_080 | SoC MUST NOT use EA capability to indicate memory resources for allocation to endpoints downstream of a PCIe root port. | #### [](#2-1-4-3-access-control-services-acs)2.1.4.3\. Access Control Services (ACS) The PCIe ACS provides controls on routing of PCIe TLPs. ACS controls may be used to determine whether the TLP should be routed normally, blocked, or redirected. These controls may be applicable to the root complex, switches, multi-function devices, and SR-IOV capable devices. | ID# | Rule | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ACS\_010 | PCIe root ports and SoC integrated downstream switch ports MUST support the following PCIe access control services (ACS) controls: ACS source validation. ACS translation blocking. ACS I/O request blocking. | | ACS\_020 | If a PCIe root port or a SoC-integrated downstream switch port implements a memory BAR, then it SHOULD support the PCIe ACS DSP memory target access control. | | _The ACS DSP memory target access control can be used to prevent unauthorized accesses to protected memory spaces such as the PCIe root port’s BAR mapped registers._ | | | ACS\_030 | Root ports and SoC-integrated downstream switch ports that support direct routing between root ports or direct routing from ingress to egress port of a root port MUST support the following PCIe ACS controls: ACS P2P request redirect. ACS P2P completion redirect. ACS upstream forwarding. ACS direct translated P2P. | | ACS\_040 | Root ports and SoC-integrated downstream switch ports that support direct routing between root ports or direct routing from ingress to egress port of a root port SHOULD also support ACS P2P egress control. | | _More commonly, P2P routing is accomplished by forwarding the TLP to the host bridge for routing. For further information, refer to the application note accompanying Fig 2-14 and Section 1.3.1 of the PCIe specification 6.0._ | | #### [](#2-1-4-4-address-routed-transactions)2.1.4.4\. Address Routed Transactions The rules in this section apply to treatment in the root complex of TLPs that are routed by address. An address carried in such transactions may be the address of a host memory location or the address of a location in the memory space of an endpoint or RCiEP. | ID# | Rule | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ADR\_010 | The host bridge MUST request IOMMU translations for addresses (Translated, Untranslated, or a PCIe ATS address translation request) used in the request by endpoints and RCiEPs. | | _The IOMMU must be invoked even for Translated requests to allow determination of whether the requester is configured by software to use Translated requests._ _When the IOMMU operates in the T2GPA mode, it provides a GPA as the translated address in response to a PCIe ATS address translation requests. In this mode of operation, the IOMMU must be invoked by the host bridge for Translated requests to translate the GPA to an SPA._ _When ACS direct translated P2P controls are enabled, the Translated requests may not be routed through the host bridge. In such cases, if direct P2P routing of these requests is not desired, due to security and/or functional reasons (e.g., when operating in T2GPA mode), software should utilize the ACS controls to direct these requests to the root complex._ | | | ADR\_020 | The host bridge MUST enforce physical memory attribute checks and physical memory protection checks on the translated address provided by the IOMMU and MUST treat violating requests as Unsupported Requests. | | _These checks are analogous to the PMA and PMP checks performed by the RISC-V hart._ | | | ADR\_030 | For Translated and Untranslated requests, the host bridge MUST use the translated addresses provided by the IOMMU to determine whether the transaction is targeting host memory or peer device memory. | | ADR\_040 | The host bridge MAY support devices accessing peer devices' memory. If peer device memory access is not enabled (either by design or configuration), then such accesses MUST be responded to with a UR/CA response. The host bridge MUST NOT cause any other errors (e.g., hang, deadlock, etc.) when rejecting access by a device to a peer device’s memory. | | _A virtual machine may violate the peer-to-peer access policies and/or configurations enforced by the hypervisor and/or SoC firmware, which prohibit peer device memory accesses. In instances where a VM configures devices passed through to it to perform peer memory accesses, such attempts must not result in system instabilities (e.g., hangs, deadlocks, etc.) or errors. Compliance with this directive ensures system resilience against unauthorized access attempts, maintaining operational integrity._ | | | ADR\_050 | When a posted or non-posted-with-data request from a device is allowed to access peer device memory, then any poisoned data (EP=1) MUST be forwarded as poisoned data, unless such forwarding is disallowed (e.g., due to poisoned TLP egress blocking or lack of support for data poisoning in the SoC). | | ADR\_060 | Host memory writes resulting from posted or non-posted-with-data requests with poisoned data (EP=1) MUST mark such data as poisoned in the host memory. | | ADR\_070 | Host memory reads that encounter uncorrectable data errors detected within the SoC MUST result in a response with poisoned data (EP=1) if transmission of poisoned TLPs is not blocked (see also section 2.7.2.1 of PCIe specification 6.0). | #### [](#2-1-4-5-id-routed-transactions)2.1.4.5\. ID Routed Transactions The rules in this section apply to treatment in the root complex of TLPs that are routed by ID. Such requests may be Configuration requests, ID routed messages or completions. | ID# | Rule | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------- | | IDR\_010 | Configuration requests from endpoints and RCiEP MUST be treated as Unsupported Requests. | | IDR\_020 | P2P routing of PCIe VDM between root ports within or across hierarchies SHOULD be supported. | | _MCTP transport protocols using PCIe VDM are used by the BMC to manage PCIe/CXL devices. These messages are used to support manageability protocols such as PLDM, NVMe-MI, Redfish, etc. Supporting P2P routing of VDMs such as those carrying MCTP protocol messages enables greater system design flexibility in supporting these management protocols._ | | | IDR\_030 | P2P routing of PCIe VDM to/from RCIeP MAY be supported. | #### [](#2-1-4-6-cacheability-and-coherence)2.1.4.6\. Cacheability and Coherence | ID# | Rule | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | CCS\_010 | The host bridge MUST enforce PCIe memory ordering rules and SHOULD support the relaxed ordering (RO) and ID-based ordering (IDO). | | _An implementation may occasionally or never permit the relaxations allowed by RO and/or IDO attributes. Such implementations will result in a more conservative interpretation of the ordering rules, but they will not result in a violation of the ordering rules._ | | | CCS\_020 | Writes to host or device memory using the RO attribute set to 0 MUST be observed by other harts and bus mastering devices in the order in which the write was received by the PCIe root port or the host bridge, ensuring that all previous writes are globally observed before the RO=0 write is globally observed. | | CCS\_030 | The host bridge MUST enforce the idempotency, coherence, cacheability, and access type physical memory attributes of the accessed memory and perform any reordering or combining of PCIe transactions only if the combination of physical memory attributes and TLP-specified memory ordering attributes allow it. | | CCS\_040 | The host bridge SHOULD implement hardware enforced cache coherency, irrespective of the “No Snoop” attribute in the TLP, unless it has been configured through UNSPECIFIED means to not enforce coherency for TLPs with “No Snoop” attribute set to 1. | | _A PCIe requester is permitted to set the “No Snoop” in transactions it initiates that do not require hardware enforced cache coherency. Host bridges that do not support isochronous VCs or can meet deadlines with hardware enforced coherency may always enforce coherency. Enforcing cache coherency is always conservative and will not lead to data corruption._ _Modern systems with integrated memory controllers and snoop directories may not require the use of “No Snoop” to meet the latency targets as memory regions accessed for isochronous operations would usually be device exclusive. PCIe requires a function to guarantee that addresses accessed using “No Snoop” set to 1 are not cached in any of the caches and software that instructs a device to perform “No Snoop” transactions must only do so when it can provide this guarantee._ _Some caches in a SoC may perform clean evictions to memory. In such SoCs, if the addresses used by the non-snooped transactions may be cached (e.g., due to speculative accesses from a hart), then such clean evictions may cause data corruption, even if the caches were explicitly cleaned by software using the cache management operations. To ensure data integrity, software should use memory that has such non-cacheable PMA or use the Svpbmt extension to override the PMA to NC/IO, thereby implementing the guarantee required by the PCIe specification when using the “No Snoop” attribute set to 1\. If the Svpbmt extension was used to override the PMA, then use of cache management operations defined by Zicbom extension may be necessary to flush data that might already be cached._ _See also section 7.5.3.4 of PCIe specification 6.0._ | | | CCS\_050 | The host bridge MUST NOT violate the coherence physical memory attribute if the “No Snoop” attribute in the TLP is 0. | | CCS\_060 | The interpretation of the TLP processing hints (TPH) by the SoC isUNSPECIFIED. | | _A future extension of the RISC-V IOMMU specification may define a standard interpretation of the TPH including the use of ATS memory attributes (AMA) for performing cache management._ | | #### [](#2-1-4-7-message-signaled-interrupts)2.1.4.7\. Message signaled interrupts A message signaled interrupt (MSI or MSI-X) is the preferred interrupt signaling mechanism in PCIe. | ID# | Rule | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------- | | MSI\_010 | Message Signaled Interrupts MUST be supported. | | MSI\_020 | SoC MUST NOT support INTx virtual wire based interrupt signaling. | | _PCIe supports INTx emulation to support legacy PCI interrupt mechanisms. Modern SoC and devices are not expected be limited by the lack of this emulation mode._ | | #### [](#2-1-4-8-precision-time-measurement-ptm)2.1.4.8\. Precision Time Measurement (PTM) | ID# | Rule | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | PTM\_010 | PCIe root ports MAY support PCIe PTM capability. | | _Several applications such as instrumentation, media servers, telecom servers, etc. require high precision monitoring and tracking of time. The PCIe PTM protocol supports synchronization of multiple devices/functions to a common shared PTM master time provided by the PTM root._ | | | PTM\_020 | When PCIe PTM capability is supported, the SoC MUST make the PTM master time available to the operating system. | | _The mechanism to make the master time available to the operating system is implementation specific._ _Making PTM master time available to software enables software to translate timing information between local time and PTM master time and thereby enable coordination of events across multiple PCIe devices._ | | | PTM\_030 | When PCIe PTM capability is supported, the PTM master time MUST be 64-bit wide. | | PTM\_040 | When PCIe PTM capability is supported, the PTM master time MUST use the same or higher resolution clock than the clock used to incrementtime CSR of the RISC-V application processor harts. | #### [](#2-1-4-9-errorevent-reporting)2.1.4.9\. Error/Event Reporting | ID# | Rule | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | AER\_010 | PCIe root ports MUST support advanced error reporting (AER) capability for reporting errors from connected devices or the errors detected by the root port itself. | | _AER capability defines more robust error reporting as compared to the baseline error reporting capability._ | | | AER\_020 | PCIe root ports MUST support the downstream port containment (DPC) capability. | | AER\_030 | PCIe root ports MUST support the RP PIO controls. | | _The root port programmed I/O (PIO) controls enable fine-grained control over handling of non-posted requests that encounter errors and allows handling of such errors as either uncorrectable or advisory based on policies established by the operating system._ | | | AER\_040 | A RCiEP in the SoC SHOULD support the AER capability if it detects any of the errors defined by PCIe specification 6.0 (See section 6.2.7). | | AER\_050 | A RCiEP in the SoC MUST support the AER capability if it supports the ACS capability. | | AER\_060 | SoC MUST implement one or more PCIe RCEC in the root complex if any of the RCiEP implement the AER capability or implement PME signaling. | | AER\_070 | The PCIe RCEC implemented in a SoC MUST implement the RCEC endpoint association extended capability. | | AER\_080 | PCIe root port configuration registers MUST NOT be affected, except as required to update status associated with the transition to DL\_Down (see also section 2.9.1 of PCIe specification 6.0). | | _Retaining port configurations on transition to DL\_Down state is important to support hot-plug._ | | #### [](#2-1-4-10-vendor-specific-registers)2.1.4.10\. Vendor Specific Registers | ID# | Rule | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | VSR\_010 | Vendor specific registers in the root ports, host bridge, RCiEP, and RCRB MUST be implemented using one or more of the following capabilities: Vendor specific capability. Vendor specific extended capability. Designated Vendor Specific extended capability. | | VSR\_020 | SoC MUST NOT require hypervisor and/or operating system interaction with PCIe configuration space registers that are not defined by an industry standard. Non-standard vendor specific registers, if implemented in the PCIe configuration space, must only be used by the SoC firmware. | | _Some industry standards such a CXL may define standard DVSEC structures in the PCIe configuration space._ _The preferred way to implement device/SoC vendor specific registers that need to be used by drivers in the run-time environment is to implement them in the memory space of the device. Certain operating systems and hypervisors may disallow and/or require mediating access to the PCIe configuration space of devices. See also the implementation note in the PCIe specification 6.0 section 7.2.2.2._ | | #### [](#2-1-4-11-soc-integrated-pcie-devices)2.1.4.11\. SoC-Integrated PCIe Devices | ID# | Rule | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | SID\_010 | SoC-integrated PCIe devices MUST implement all software visible rules defined by the PCIe specification 6.0 for an EP or RCiEP as applicable. | | _Implementing integrated devices as RCiEP or EP allows the use of standardized frameworks for memory and interrupt resource allocation, virtualization (SR-IOV), ATS/PRI, shared virtual addressing, trusted IO using SPDM/TDISP, participate in RAS frameworks like data poisoning and AER, power management, etc._ | | | SID\_020 | SoC-integrated PCIe devices MUST NOT require the use of I/O space or I/O transactions. | | SID\_030 | SoC integrated PCIe devices that cache address translations MUST implement the PCIe ATS capability if the address translation cache needs management by the operating system or hypervisors. | | SID\_040 | SoC-integrated PCIe devices that support PCIe SR-IOV capability SHOULD support the MSI-X capability. | | _MSI-X capability enables virtual machines to assign interrupt resources to virtual functions without needing access to the configuration space of the function. Access to the configuration space of the virtual function is usually mediated by the hypervisor._ | | | SID\_050 | SoC-integrated PCIe devices MAY support the PASID capability. When PASID capability is supported, the devices SHOULD support a 20-bit wide PASID. | | _Endpoints are recommended to support a 20-bit wide PASID to ensure interoperability with system software. See also the implementation note on PASID width homogeneity in the PCIe specification 6.0 section 6.20.2.2._ | | | SID\_060 | SoC-integrated PCIe devices (a multi-function device or an SR-IOV capable device) that support P2P traffic among functions (including among SR-IOV virtual functions) of the device MUST support the following PCIe ACS controls: ACS P2P request redirect. ACS P2P completion redirect. ACS direct translated P2P. | | SID\_070 | If the BAR registers are implemented by SoC-integrated PCIe devices then they MUST be programmable. The Memory Space Indicator (bit 0) of such BAR registers MUST be 1, and they SHOULD support being mapped anywhere in the 64-bit memory space. | | SID\_080 | RCiEP MAY support the PCIe enhanced allocation (EA) capability for fixed allocation of memory resources. If EA capability is used then the BEI of the entries MUST be one of 0 through 5 or 9 through 14 and their primary/secondary properties must be one of 0 through 4 or 0xFF. | | SID\_090 | SoC-integrated PCIe devices MUST support the PCIe defined baseline error reporting capability and MAY support PCIe Advanced Error Reporting capability. If PCIe ACS controls are supported then the PCIe Advanced Error Reporting capability MUST be supported. | | _See PCIe specification 6.0 section 7.5.1.1.14._ | | | SID\_100 | A RCiEP that supports PCIe Advanced Error Reporting MUST be associated with a Root Complex Event Collector. | ### [](#2-1-5-reliability-availability-and-serviceability-ras)2.1.5\. Reliability, Availability, and Serviceability (RAS) | ID# | Rule | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | RAS\_010 | The level of RAS implemented by the SoC is UNSPECIFIED. | | _The level of RAS implemented by an SoC depends on the reliability goals established for the SoC, which are commonly measured using metrics such as failure-in-time (FIT) and defects-per-million (DPM). Achieving these goals requires a combination of fault prevention, error detection, and error correction techniques._ _This specification strongly recommends the implementation of error detection and correction codes for storage elements like significant caches and memories. Furthermore, it suggests utilizing mechanisms such as single-symbol (SSC) ECC in DRAM controllers to address failure scenarios, such as when all bits in a single DRAM device experience a failure._ _Additionally, this specification encourages the adoption of mechanisms like periodic scrubbing, also known as patrol scrubbing. These mechanisms proactively identify and rectify errors before they accumulate to a critical point, surpassing the capability of the implemented error correction codes. For instance, this could involve addressing situations where single bit errors escalate into double bit errors, surpassing the correction code’s capacity._ | | | RAS\_020 | SoC SHOULD support the generation, storage, and forwarding of poisoned data. The granularity at which data is poisoned isUNSPECIFIED. | | _When an uncorrected data error is detected by a component, it might allow potentially corrupted data to reach the data requester, but with an associated poison indicator. These errors are referred to as uncorrected deferred errors (UDE), as they enable the detecting component to continue functioning and postpone addressing the error until a later time, assuming the poisoned data gets consumed. If a component (such as a hart, an IOMMU, a device, etc.) consumes the poisoned data, it triggers an uncorrected urgent error (UUE), leading to the invocation of a recovery handler for immediate remedial actions, as further deferral of the error is not feasible._ _The technique of data poisoning facilitates delaying the handling of uncorrected errors until the moment the corrupted data is actually consumed. Data poisoning offers a more precise identification of the software and/or hardware component affected by the data corruption. This specificity allows for targeted recovery actions that impact only the affected components._ _To ensure the integrity of the poisoned data indicator when stored, error detection and correction codes should be applied. This practice prevents subsequent errors from leading to the silent consumption of the corrupted data._ _Data poisoning also empowers the implementation of error containment features supported by industry standards like PCIe and CXL._ _For more detailed discussions on the treatment of faults and errors, refer to the RISC-V RERI specification._ | | | RAS\_030 | If poisoned data needs to be transmitted from a first component to a second component that lacks the ability to manage poison, the first component MUST trigger an critical uncorrected error report instead of silently transmitting the corrupted data. | | _Some components serve as intermediaries through which data passes. For instance, a PCIe/CXL port acts as an intermediary that receives data from memory but doesn’t consume it; rather, it forwards the data to an endpoint. In such cases, the intermediary component might encounter poisoned data. While this component can propagate the error and avoid logging an error, a different scenario arises when the destination component (such as a PCIe endpoint) cannot handle poison. In such situations, the originating component must trigger an urgent error signal instead of transmitting the poisoned data without the associated poison indicator. Failing to do so would breach the containment of the corrupted data during propagation._ | | | RAS\_040 | The SoC SHOULD support the RISC-V RAS error record register interface (RERI) \[[11](server%5Fsoc%5Fbibliography.html#bib-reri)\] for error logging and signaling. | | RAS\_050 | When RERI is supported, the RAS error records MUST include the capability to individually enable error signaling for each severity - Uncorrected Error Critical (UEC), Uncorrected Error Deferred (UED), and Corrected Error (CE) - of error that could be logged in that specific error record. | | _Configurable enables provide software with the flexibility of using an event-based or polling-based error logging for both corrected errors and deferred errors. Typically, software operates in an event-based mode for critical errors, as these errors necessiate immediate remedial action when they arise._ | | | RAS\_060 | If RERI is supported, RAS error records MUST preserve the state of logged error information (including status, address, information, supplemental information, and timestamp) across a RAS-initiated reset. The state of RAS error records MAY persist across other types of implementation-defined resets. After a reset, including a RAS-initiated reset, the state of the control register in the RAS error record is considered UNSPECIFIED. | | _Some errors may lead a hardware component to enter a failure mode in which it becomes incapable of servicing additional requests- colloquially termed 'jammed' or 'wedged'. In these situations, the SoC may require a reset to restore it to an operational state (a RAS-initiated reset). Preserving the RAS error records through such resets enables the SoC firmware and system software to retrieve these error records during boot following such a reset, facilitating logging and analysis._ | | | RAS\_070 | If RERI is supported, the RAS error records MAY support error record injection, which is intended to facilitate RAS handler verification. | | _Verifying the correct implementation of RAS handlers presents a formidable challenge, given the impracticality of deterministically inducing all potential errors within the SoC to validate the RAS handler’s adherence to desired recovery protocols. An unverified RAS handler can lead to undesired behavior during error occurrences, potentially reducing SoC availability or affecting its serviceability._ _To address this, error record injection offers a convenient method for conducting such verification. It allows the introduction of a range of error signatures, which can then be signaled and observed. While hardware error injection techniques also offer a means of verification (e.g., methods to intentionally corrupt a data location protected by an error detection code), providing open access to these capabilities for software use might not align with security and stability concerns._ | | | RAS\_080 | If RERI is supported, then the hardware components in the SoC that support error correction MUST incorporate a corrected error counter within their respective error records. Additionally, these components MUST support the signaling of counter overflows. | | _Counting corrected errors offers a more precise assessment of system reliability. Enabling signaling upon counter overflow empowers software to define a suitable threshold for logging and analysis of these corrected errors._ _Certain hardware units might maintain a history of corrected errors and increment the corrected error counter only if the error differs from a previously reported one. Additionally, some hardware units could incorporate low-pass filters like leaky buckets, which regulate the rate at which corrected errors are reported and counted. This rule pertains to corrected errors tracked by the error record once the hardware component determines reporting and counting based on its specific filtering rules._ | | ### [](#2-1-6-quality-of-service)2.1.6\. Quality of Service Quality of Service (QoS) refers to the minimum end-to-end performance that a service level agreement (SLA) guarantees to an application in advance. QoS capabilities within the SoC offer mechanisms that system software can leverage to manage interference to an application, effectively diminishing performance variability caused by other applications' utilization of shared resources such as cache capacity, memory bandwidth, interconnect bandwidth, power consumption, and more. | ID# | Rule | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | QOS\_010 | The SoC SHOULD incorporate QoS mechanisms to mitigate unwarranted performance interference that arises when multiple workloads access shared resources like caches and system memory. | | QOS\_020 | The SoC SHOULD integrate support for the RISC-V capacity and bandwidth controller register interface (CBQRI) \[[12](server%5Fsoc%5Fbibliography.html#bib-cbqri)\] in significant shared caches and the memory controllers. | | QOS\_030 | If CBQRI is supported, RISC-V harts within the application processors of the SoC MUST include support for the srmcfg CSR. Furthermore, this CSR MUST support a minimum of 16 RCIDs and at least 32 MCIDs. The count of RCID and MCID that can be used in the SoC SHOULD scale with the number of RISC-V harts in the SoC. | | _The srmcfg CSR is provided by the Ssqosid extension \[[13](server%5Fsoc%5Fbibliography.html#bib-priv)\]._ | | | QOS\_040 | If CBQRI is supported, the IOMMUs in the SoC SHOULD incorporate support for the CBQRI-defined extension, enabling the association of RCID and MCID with requests initiated by devices and the IOMMU. | | QOS\_050 | If CBQRI is supported, significant caches such as the last-level cache in the SoC SHOULD support cache capacity allocation. | | QOS\_060 | If CBQRI is supported, significant caches such as the last-level cache in the SoC SHOULD incorporate support for monitoring cache capacity usage. | | QOS\_070 | If CBQRI is supported, the memory controllers within the SoC SHOULD include support for bandwidth allocation. | | QOS\_080 | If CBQRI is supported, the memory controllers in the SoC SHOULD include support for monitoring bandwidth usage. | | _The method employed by the SoC for bandwidth throttling and control is specific to its implementation. It is advisable for the implementation to utilize a scheme that results in a deviation of no more than +/- 10 % from the target set by system software through the CBQRI interface._ | | | QOS\_090 | If CBQRI is supported, the count of RCID and MCID supported by capacity controllers, bandwidth controllers, and all RISC-V application processor harts in the SoC MUST be consistent. | | _Portable system software could opt to limit itself to accommodating the minimum count of RCID and MCID across the controllers. This approach avoids the complexity of dealing with unequal numbers of RCID and MCID across controllers, which would otherwise necessitate intricate allocations and constraints on workload placement._ | | | QOS\_100 | If CBQRI is supported, the monitoring counters in the capacity and bandwidth controllers MUST be sufficiently wide to not overflow when sampled at a rate of 1 Hz. | | _As an illustration, consider an HBM3 memory interface that can facilitate data transfers at a rate of up to 1 TB/s. This scenario would necessitate a 34-bit counter to prevent overflow when sampled at a frequency of 1 Hz._ | | ### [](#2-1-7-manageability)2.1.7\. Manageability This section outlines the guidelines for RISC-V server SoCs to incorporate a standardized set of protocols and standards for server management. The SoC interfaces with a baseboard management controller (BMC) through in-band and out-of-band (OOB) management agents. The in-band management agents execute on the RISC-V application processor harts and the out-of-band management agents execute on a management controller in the SoC. The out-of-band management interface facilitates the monitoring of sensors (e.g., temperature, power, etc.), parameter control (e.g., power limits, etc.), and logging (e.g., RAS error records, etc.) by the BMC without participation of software on the application processor harts. The in-band management interface facilitates system configuration (e.g., boot order, memory domains, secure boot, network, etc.), and event log collection through management agents in the OS and/or firmware that executes on the application processor harts. This specification strongly recommends the use of the DMTF Redfish \[[14](server%5Fsoc%5Fbibliography.html#bib-dsp0266)\], DMTF Platform Level Data Model (PLDM) \[[15](server%5Fsoc%5Fbibliography.html#bib-dsp0240)\], and DMTF Management Component Transport Protocol (MCTP) \[[16](server%5Fsoc%5Fbibliography.html#bib-dsp0236)\]) protocols for in-band and out-of-band server management. This specification strongly recommends the use of DMTF specified Security Protocol and Data Model (SPDM) \[[17](server%5Fsoc%5Fbibliography.html#bib-dsp0274)\] for device attestation and using SPDM encrypted messages \[[18](server%5Fsoc%5Fbibliography.html#bib-dsp0277)\] for secure in-band and out-of-band communication with the BMC. SPDM authentication protocols support establishing a trust relationship between the manageability agents in the SoC and the BMC. Use of SPDM secured messages enables preserving the confidentiality and integrity of data exchanged between the BMC and the manageability agents in the SoC. The specification recommends supporting Intelligent Platform Management Interface (IPMI) \[[19](server%5Fsoc%5Fbibliography.html#bib-ipmi20)\] due to the widespread use of this protocol for server management functions such as credentials provisioning and remote power control. This specification recommends the RISC-V server SoC to support open standards for server management through supporting integration with technologies such as the datacenter-ready secure control module (DC-SCM) \[[20](server%5Fsoc%5Fbibliography.html#bib-dc-scm)\] specified by the Open Compute Project for server management, security, and control features. Adhering to the industry standard management protocols such as those specified by DMTF and OCP allows server platforms built with RISC-V server SoCs to seamlessly integrate into the server management frameworks and tools employed by data centers and enterprises. | ID# | Rule | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | | MNG\_010 | The SoC SHOULD incorporate support for an x1 PCIe lane, preferably Gen 5, but at least Gen 3, to establish a connection with the BMC. | | _This interface is commonly linked to a BMC as a PCIe endpoint, serving various purposes. These include facilitating host-to-BMC communication for tasks like video output (e.g., remote KVM support), MCTP transport over PCIe VDM, and hosting a USB controller. The BMC might also support remote presence capabilities, like remote media redirection and support for keyboard and mouse functions through virtual USB._ _The in-band network interfaces serve as communication channels for system software to interact with the BMC. This interaction employs protocols like the Redfish host interface._ _Furthermore, the PCIe interface to the BMC empowers the BMC, using SoC-routed PCIe VDMs, to utilize these VDMs for transmitting MCTP messages. These messages manage platform devices, including network controllers, NVMe controllers, FPGAs, GPUs, and more._ | | | MNG\_020 | The SoC SHOULD support the use of I2C based IPMI SSIF for in-band management agents in the SoC to communicate with the BMC. | | MNG\_030 | The SoC SHOULD incorporate support for utilizing a UART connection to the BMC, enabling the provision of a host debug console. | ### [](#2-1-8-performance-monitoring)2.1.8\. Performance Monitoring | ID# | Rule | | -------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | SPM\_010 | Significant caches within the SoC SHOULD incorporate an HPM capable of counting: Cache lookups for reads Cache misses on reads Cache lookups for writes Cache misses on writes | | _It is recommended that a cache with a capacity that is approximately 16 KiB or larger be considered a significant cache._ | | | SPM\_020 | The memory controllers within the SoC SHOULD incorporate an HPM capable of counting: Read bandwidth Write bandwidth | | SPM\_030 | The PCIe ports within the SoC SHOULD incorporate an HPM capable of counting: Read bandwidth (from system memory) Write bandwidth (to system memory) | | SPM\_040 | The SoC SHOULD incorporate an HPM capable of counting the average latency of a read request from a memory requester (e.g., a hart, a PCIe host bridge, etc.) in the SoC. | | _Bandwidth and latency are the most commonly used performance metrics to guide workload placement and tuning._ | | | SPM\_050 | If the SoC supports NUMA configurations, then the HPM for SPM\_010, SPM\_020, SPM\_030, and SPM\_040 SHOULD support filtering the counting based on whether the request is to local memory or to remote memory. | | SPM\_060 | All PCIe Gen6 ports within the SoC SHOULD incorporate support for the Flit performance measurement extended capability defined by PCIe specification 6.0. | Please refer to [Input-Output Memory Management Unit (IOMMU)](#IOMMU) for details on the IOMMU performance monitoring rules. ### [](#2-1-9-security-requirements)2.1.9\. Security Requirements | ID# | Rule | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | SEC\_010 | The Server SoC MUST implement a hardware RoT as the _primary_ root of trust. | | _A root of trust (RoT) is the foundation on which all secure operations of a system depend. A hardware RoT is a dedicated and possibly isolated trusted subsystem that can provide stronger protections against physical and logical attacks._ | | | SEC\_020 | The PCIe root ports within the SoC SHOULD support PCIe Integrity and Data Encryption (IDE) capability. | | _The IDE extension adds optional capabilities to perform hardware encryption and integrity checks on packets transferred across PCIe links. This addition provides confidentiality, integrity, and replay protection against hardware-level attacks._ | | | SEC\_030 | The SoC SHOULD support encryption of off-chip DRAM using a transient memory encryption key that has at least 256-bit key lengths. | | _Off-chip memory encryption provides protection to critical assets in memory such as credentials, data encryption keys, and other secrets._ | | | SEC\_040 | The cryptographic modules used to implement PCIe and off-chip DRAM encryption SHOULD comply with security requirements specified by relevant security standards from national standards laboratories. | | _FIPS 140-3 is an example of such a standard_ | | | SEC\_050 | The SoC SHOULD have the capability of interfacing with a Trusted Platform Module (TPM) that adheres to the TPM 2.0 Library specification \[[21](server%5Fsoc%5Fbibliography.html#bib-tpm20)\]. | | _A TPM enhances security by providing secure storage for sensitive information such as credentials and passwords, cryptographic operations and protection against tampering or unauthorized access._ | | Efficient Trace for RISC-V ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#efficient-trace-for-risc-v)Efficient Trace for RISC-V Gajinder Panesar, Iain Robertson Version 2.0, 6/2024: Ratified state | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at. Copyright 2025 by RISC-V International. Preface ==================== ## [](#preface)Preface **_Preface to Version 20250616_** Clarifications only - no changes to normative behaviour. * Updated [\[packets\]](#packets) and [\[fragments\]](#fragments) to reference the [Unformatted Trace & Diagnostic Data Packet Encapsulation for RISC-V Specification](https://github.com/riscv-non-isa/e-trace-encap/releases/latest/); * Removed ambiguity between 'last' and 'final'. Last was previously used to mean both the instruction before the current one, and the final instruction traced; * Clarified behaviour when trace-on and trace-off triggers occur both occur in the same cycle (see [\[sec:trigger\]](#sec:trigger)); * Clarified that full synchronization also takes place on a privilege change (see [\[sec:synchronization\]](#sec:synchronization) and [\[sec:thaddr\]](#sec:thaddr)); * Reworded jump classifications in jump classifications in [\[sec:InstructionInterfaceRequirements\]](#sec:InstructionInterfaceRequirements) to align with terminology used in other specifications; * Clarified when to issue sync-support packet when trace is enabled (see [\[sec:format33\]](#sec:format33)); * Clarified that branch\_map can be set to 31 any time an address needs to be reported and the branch map is full (see [\[sec:format1\]](#sec:format1)); * Updated reference encoding algorithm to generalize resync behaviour and add missing trap conditions (see [\[Algorithm\]](#Algorithm)); * Clarified that implicit returns are not stored in the jump target cache (see [\[sec:jump-cache\]](#sec:jump-cache)); * Updated decoder to remove ECALL, EBREAK and C.EBREAK from is\_uninferrable\_jump function. Including them is harmless but unnecessary, as these instrucitons don’t retire (see [\[Decoder\]](#Decoder)); * Corrected description of why the EPC is not always reported (see [\[sec:thaddr\]](#sec:thaddr)). * Minor clarifications, formatting and typo fixes. **_Preface to Version 20240419_** Formatting and typo fixes. **_Preface to Version 20240305_** First version in AsciiDoc format. **_Preface to Version 20231215_** Clarifications only - no changes to normative behaviour. * Control field definitions removed from section 2, which now references the [RISC-V Trace Control Interface Specification](#https://github.com/riscv-non-isa/e-trace-encap/releases/latest/) * Added detail on handling of multi-load/store instructions for data trace to [\[sec:DataInterfaceRequirements\]](#sec:DataInterfaceRequirements). * Removed references to tail-calls in jump classifications in [\[sec:InstructionInterfaceRequirements\]](#sec:InstructionInterfaceRequirements). * Corrected typos where `lrid` was inadvertently refered to by an earlier name (`index`) in [\[sec:data-loadstore\]](#sec:data-loadstore). * Corrected reference decoder in [\[Decoder\]](#Decoder) to cover a corner-case related to trap returns. **_Preface to Version 2.0_** Ratified version of the Efficient Trace for RISC-V specification. 3.1. Branch Trace ==================== ## [](#BranchTrace)3.1\. Branch Trace Instruction delta tracing, also known as branch tracing, works by tracking execution from a known start address by sending information about the deltas taken by the program. Deltas are typically introduced by jump, call, return and branch type instructions, although interrupts and exceptions are also types of deltas. Instruction delta tracing provides an efficient encoding of an instruction sequence by exploiting the deterministic way the processor behaves based on the program it is executing. The approach relies on an offline copy of the program binary being available to the decoder, so it is generally unsuitable for either dynamic (self-modifying) programs or those where access to the program binary is prohibited. While the program binary is sufficient, access to the assembly or higher-level source code will improve the ability of the decoder to present the decoded trace in the debugger by annotating the traced instructions with source code line numbers and labels, variable names etc. This approach can be extended to cope with small sections of deterministically dynamic code by arranging for the decoder to request instruction memory from the target. Memory lookups generally lead to a prohibitive reduction in performance, although they are suitable for examining modest jump tables, such as the exception/interrupt vector pointers of an operating system which may be adjusted at boot up and when services are registered. Both static and dynamically linked programs can be traced using this approach. Statically linked programs are straightforward as they generally operate in a known address space, often mapping directly to physical memory. Dynamically linked programs require the debugger to keep track of memory allocation operations using either trace or stop-mode debugging. ### [](#TraceConcepts)3.1.1\. Instruction delta trace concepts #### [](#SequentialInstructions)3.1.1.1\. Sequential instructions For instruction set architectures such as RISC-V where all instructions are executed unconditionally or at least their execution can be determined based on the program binary, the instructions between the deltas are assumed to be executed sequentially. Consequently, there is no need to report them in the trace. The trace only needs to contain whether branches were taken or not, the addresses of taken indirect jumps, or other program counter discontinuities. #### [](#uninfpc)3.1.1.2\. Uninferable PC discontinuities An uninferable program counter discontinuity is a program counter change that can not be inferred from the program binary alone. For these cases, the instruction delta trace must include a destination address: the address of the next valid instruction. Indirect jumps are an example of this, where the next instruction address is determined by the contents of a register rather than a constant embedded in the program binary. In this case, the address of the instruction following the jump (also known as the jump target) must be traced. Interrupts and exceptions are another form of uninferable PC discontinuity; these are discussed in detail below. #### [](#3-1-1-3-branches)3.1.1.3\. Branches A branch is an instruction where a jump is conditional on the value of a register or a flag. For a decoder to able to follow program flow, the trace must include whether a branch was taken or not. For a direct branch, where the destination address is encoded in the program binary (either as a constant, or as a constant offset from the program counter), no further information is required. Direct branches are the only type of branch that is supported by the RISC-V ISA. #### [](#interruptsexceptions)3.1.1.4\. Interrupts and exceptions Interrupts are a different type of delta that generally occur asynchronously to the program’s execution rather than intentionally as a result of a specific instruction or event. Exceptions can be thought of in the same way, even though they can be typically linked back to a specific instruction address. The decoder generally does not know where an interrupt occurred in the instruction sequence, so the trace must report the address where normal program flow ceased, as well as give an indication of the asynchronous destination which may be as simple as reporting the exception type. When an interrupt or exception occurs, the final instruction retired beforehand must be traced. Following this the next valid instruction address (the first of the trap handler) must be traced. Note: not all exceptions and interrupts cause traps (see[\[sec:terminology\]](#sec:terminology) for definitions). Most notably, floating point exceptions and disabled interrupts do not trap. If an exception or interrupt doesn’t trap, the program counter does not change. So, there is no need to trace all exceptions/interrupts, just traps. In this document, interrupts and exceptions are only traced when they cause traps to be taken. #### [](#sec:synchronization)3.1.1.5\. Synchronization In order to make the trace robust there must be regular synchronization points within the trace. Synchronization is accomplished by sending a full valued instruction address (and potentially a context identifier). The decoder and debugger may also benefit from sending the reason for synchronizing. The frequency of synchronization is a trade-off between robustness and trace bandwidth. The instruction trace encoder needs to synchronise fully: * For the first instruction traced after reset or resume from halt; * Any time that an instruction is traced and the previous instruction was not traced; * If the instruction is the first of an interrupt service routine or exception handler; * If the privilege level changes; * After a prolonged period of time. #### [](#sec:endoftrace)3.1.1.6\. End of trace If tracing stops for any reason, the address of the final traced instruction must be output. Some examples of why tracing may stop are: * The hart may be halted (entered debug mode); * The hart may be reset; * Encoding may be stopped (for example via a _Trace-off_ trigger - see[\[sec:trigger\]](#sec:trigger)); * The matching criteria for any filtering capabilities implemented by the encoder may no longer be met; * The encoder may be disabled. ### [](#optional)3.1.2\. Optional and run-time configurable modes An instruction trace encoder may support multiple tracing modes. To ensure that the decoder treats the incoming packets correctly, it needs to be informed of the current active configuration. The configuration is reported by a packet that is issued by the encoder whenever the encoder configuration is changed. Here are common examples of such modes: * delta address mode: program counter discontinuities are encoded as differences instead of absolute address values. * full address mode: program counter discontinuities are encoded as absolute address values. * implicit exception mode: the destination address of an exception (i.e. the address of the exception trap) is assumed to be known by the decoder, and thus not encoded in the trace. * Sequentially inferable jump mode: The target of an indirect jump can be inferred by considering the combined effect of two instructions. * implicit return mode: the destination address of function call returns is derived from a call stack, and thus not encoded in the trace. * branch prediction mode: branches that are predicted correctly by an encoder branch predictor (and an identical copy in the decoder) are not encoded as taken/non-taken, but as a more efficient branch count number. * Jump target cache mode: Rather than reporting the address of an uninferable jump target, efficiency can be improved by caching recent jump targets, and reporting the cache entry index instead. Modes may have associated parameters; see [\[tab:iparameters\]](#tab:iparameters) for further details. All modes are optional apart from delta address mode, which must be supported. #### [](#sec:delta-address)3.1.2.1\. Delta address mode Related parameters: None In delta address mode, addresses are encoded as the difference between the actual address of the current instruction and the actual address of the instruction reported in the previous packet that contained an address. This differential encoding requires fewer bits than the full address, and thus results in more efficient trace compression. #### [](#sec:full-address)3.1.2.2\. Full address mode Related parameters: None In full address mode, all addresses in the trace are encoded as absolute addresses instead of in differential form. This kind of encoding is always less efficient, but it can be a useful debugging aid for software decoder developers. #### [](#sec:implicit-exception)3.1.2.3\. Implicit exception mode Related parameters: None The RISC-V Privileged ISA specification stores exception handler base addresses in the **_stvec/vstvec/mtvec_** CSR registers. In some RISC-V implementations, the lower address bits are stored in the**_scause/vscause/mcause_** CSR registers. By default, both the **_\*tvec_** and **_\*cause_** values are reported when an exception or interrupt occurs. The implicit exception mode omits **_\*tvec_** (the trap handler address), from the trace and thus improves efficiency. This mode can only be used if the decoder can infer the address of the trap handler from just the exception cause. #### [](#sec:si-jump)3.1.2.4\. Sequentially inferable jump mode Related parameters: _sijump\_p_. By default, the target of an indirect jump is always considered an uninferable PC discontinuity. However, if the register that specifies the jump target was loaded with a constant then it can be considered inferable under some circumstances. The hart must identify jumps with sequentially inferable targets and provide this information separately to the encoder. The final decision as to whether to treat the jump as inferable or not must be made by the encoder. Both the constant load and the jump must be traced in order for the decoder to be able to infer the jump target. See [\[JumpClasses\]](#JumpClasses) for details of what constitutes a sequentially inferable jump. #### [](#sec:implicit-return)3.1.2.5\. Implicit return mode Related parameters: _call\_counter\_size\_p_, _return\_stack\_size\_p_,_itype\_width\_p_. Although a function return is usually an indirect jump, well behaved programs return to the point in the program from which the function was called using a standard calling convention. For those programs, it is possible to determine the execution path without being explicitly notified of the destination address of the return. The implicit return mode can result in very significant improvements in trace encoder efficiency. Returns can only be treated as inferable if the associated call has already been reported in an earlier packet. The encoder must ensure that this is the case. This can be accomplished by utilizing a counter to keep track of the number of nested calls being traced. The counter increments on calls (but not tail calls), and decrements on returns (see[\[JumpClasses\]](#JumpClasses) for definitions). The counter will not over or underflow, and is reset to 0 whenever a synchronization packet is sent. Returns will be treated as inferable and will not generate a trace packet if the count is non-zero (i.e. the associated call was already reported in an earlier packet). Such a scheme is low cost, and will work as long as programs are "well behaved". The encoder does not check that the return address is actually that of the instruction following the associated call. As such, any program that modifies return addresses cannot be traced using this mode with this minimal implementation. Alternatively, the encoder can maintain a stack of expected return addresses, and only treat a return as inferable if the actual return address matches the prediction. This is fully robust for all programs, but is more expensive to implement. In this case, if a return address does not match the prediction, it must be reported explicitly via a packet, along with the number of return addresses currently on the stack. This ensures that the decoder can determine which return is being reported. #### [](#sec:branch-prediction)3.1.2.6\. Branch prediction mode Related parameters: _bpred\_size\_p_. Without branch prediction, the outcome of each executed branch is stored in a branch map: a bit vector in which the taken/non-taken status of each branch is stored in chronological order. While this encoding is efficient, at 1 bit per branch, there are some cases where this can still result in a relatively large volume of trace packets. For example: * Executing tight loops of code containing no uninferable jumps. Each iteration of the loop will add a bit to the branch map; * Sitting in an idle loop waiting for an interrupt. This produces large amounts of trace when nothing of any interest is actually happening! * Breakpoints, which in some implementations also spin in an idle loop. A significant coding efficiency can be obtained by the addition of a branch predictor in the encoder. To keep the encoder and decoder synchronized, a predictor with identical behavior will need to be implemented in the decoder software. The predictor shall comprise a lookup table of 2_bpred\_size\_p_entries. Each entry is indexed by bits _bpred\_size\_p_:1 of the instruction address (or _bpred\_size\_p_+1:2 if compressed instructions aren’t supported), and each contains a 2-bit prediction state: * 00: predict not taken, transition to 01 if prediction fails; * 01: predict not taken, transition to 00 if prediction succeeds, else 11; * 11: predict taken, transition to 10 if prediction fails; * 10: predict taken, transition to 11 if prediction succeeds, else 00. The MSB represents the predicted outcome, the LSB the most recent actual outcome. The prediction must fail twice for the predicted value to change. The lookup table entries are initialized to 01 when a synchronization packet is sent. #### [](#sec:jump-cache)3.1.2.7\. Jump target cache mode Related parameters: _cache\_size\_p_. By default, the target address of an uninferable jump is output in the trace, usually in differential form. If the same function is called repeatedly, (for example, in a loop), the same address will be output repeatedly. An efficiency gain can be obtained by the addition of a jump target cache to the encoder. To keep the encoder and decoder synchronized, a cache with identical behavior will need to be implemented in the decoder software. Even a small cache can provide significant improvement. The cache shall comprise 2_cache\_size\_p_ entries, each of which can contain an instruction address. The addresses stored in the cache are the targets of uninferable jumps. It will be direct mapped, with each entry indexed by bits _cache\_size\_p_:1 of the instruction address (or_cache\_size\_p_+1:2 if compressed instructions aren’t supported). Each uninferable jump target is first compared with the entry in the cache at the index derived from the jump target address. If it is found in the cache, the index number is traced rather than the target address. If it is not found in the cache, the entry at that index is replaced with the current instruction address. Note that if implicit return mode is enabled, most function return targets are inferable and are not output in the trace. Any such inferrable return target must not be stored in the cache. This effectively puts the implicit return evaluation in series before the jump target cache, and this may present timing challenges. This can be mitigated by invalidating the cache entry that corresponds to a mis-predicted return rather than updating the entry with the actual return address. Mis-predicted returns are rare so the performance impact will be negligible. The cache entries are all invalidated when a synchronization packet is sent. 2.1. Encoder Control ==================== ## [](#encoderControl)2.1\. Encoder Control The fields required to control a Trace Encoder are defined in the[RISC-V Trace Control Interface Specification](https://github.com/riscv-non-isa/e-trace-encap/releases/latest/), which is intended to apply to any and all RISC-V trace encoders, regardless of encoding protocol. This chapter details which of those fields apply to E-Trace. To avoid replication, descriptions are not provided here; additional E-Trace specific context or clarification is provided only where required. How fields are organized and accessed (e.g packet based or memory mapped) is outside the scope of this document. If a memory mapped approach is adopted, the 'Trace Component Register Map' from the[RISC-V Trace Control Interface Specification](https://github.com/riscv-non-isa/e-trace-encap/releases/latest/) should be used. Note: Upto and including the E-Trace v2.0.0 specification, which predated the creation of the RISC-V Trace Control Interface Specification, the full field definitions were included in this chapter. For versions later than this, the field definitions have simply moved from this specification to the RISC-V Trace Control Interface Specification, without any change to their meaning. However, in order to create a more widely applicable protocol agnostic specification it has been necessary to change the field names in the process. The applicability of fields for E-trace is categorized as follows: * N: Not applicable * M: Mandatory * O: Optional * MD: Mandatory if data trace is supported * OD: Optional for data trace ### [](#sec:ctl-basic)2.1.1\. Basic Control The following fields control basic encoding behavior. __Table 1\. Basic Control__ | **Field** | **Applicability** | **E-Trace Specific Details** | | --------------------------- | ----------------- | ---------------------------------------------------------------------------------- | | **trTeActive** | M | | | **trTeEnable** | M | | | **trTeInstTracing** | M | | | **trTeDataTracing** | MD | | | **trTeInstTrigEnable** | O | | | **trTeDataTrigEnable** | OD | | | **trTeInstStallOrOverflow** | O | | | **trTeDataStallOrOverflow** | OD | | | **trTeInstStallEn** | O | | | **trTeDataStallEn** | OD | | | **trTeEmpty** | O | Recommended if the trace datapath requires manual flushing when trace is disabled. | | **trTeDataDrop** | OD | | | **trTeDataDropEn** | OD | | | **trTeInhibitSrc** | O | | | **trTeInstSyncMode** | M | If hardcoded, must be to a non-zero value. | | **trTeInstSyncMax** | M | May be hardcoded. | | **trTeFormat** | M | Must be set to 0 (denoting E-Trace format). | | **trTeVerMajor** | M | | | **trTeVerMinor** | M | | | **trTeCompType** | M | | | **trTeProtocolMajor** | M | Must be 0 to indicate this version (2.0.x) of the E-Trace protocol. | | **trTeProtocolMinor** | M | Must be 0. | | **trTeSrcID** | O | | | **trTeSrcBits** | O | | ### [](#sec:ctl-modes)2.1.2\. Optional Modes See [\[optional\]](#optional) for details of the modes covered in this section. __Table 2\. Optional and run-time configurable modes.__ | **Field** | **Applicability** | **E-Trace Specific Details** | | ------------------------------ | ----------------- | ---------------------------- | | **trTeInstNoAddrDiff** | O | | | **trTeInstNoTrapAddr** | O | | | **trTeInstEnSequentialJump** | O | | | **trTeInstEnImplicitReturn** | O | | | **trTeInstEnBranchPrediction** | O | | | **trTeInstJumpTargetCache** | O | | | **trTeDataNoValue** | OD | | | **trTeDataNoAddr** | OD | | | **trTeDataAddrCompress** | OD | | | **trTeContext** | N | Hardcode to 0. | | **trTeInstMode** | N | Hardcode to 7. | | **trTeInstImplicitReturnMode** | N | Hardcode to 0. | | **trTeInstEnRepeatedHistory** | N | Hardcode to 0. | | **trTeInstEnAllJumps** | N | Hardcode to 0. | | **trTeInstExtendAddrMSB** | N | Hardcode to 0. | ### [](#sec:ctl-filter)2.1.3\. Filtering See [\[ch:filtering\]](#ch:filtering) for details of the filtering capabilities covered in this section. __Table 3\. Trace filtering selection__ | **Field** | **Applicability** | **E-Trace Specific Details** | | ------------------- | ----------------- | ---------------------------- | | **trTeInstFilters** | O | | | **trTeDataFilters** | OD | | | **trTeFilter…​** | O | | | **trTeComp…​** | O | | | **trTeTrig…​** | N | Hardcode to 0. | 8.1. Data Trace Encoder Output Packets ==================== ## [](#dataTracePackets)8.1\. Data Trace Encoder Output Packets Data trace packets must be differentiated from instruction trace packets, and the means by which this is accomplished is dependent on the trace transport infrastructure. Several possibilities exist: One option is for instruction and data trace to be issued using different IDs (for example, if using ATB transport, different **ATID** values). Alternatively, an additional field as part of the packet encapsulation can be used (Siemens uses a 2-bit **msg\_type** field to differentiate different trace types from the same source). By default, all data trace packets include both address and data. However, provision is made for run-time configuration options to exclude either the address or the data, in order to minimize trace bandwidth. For example, if filtering has been configured to only trace from a specific data access address there is no need to report the address in the trace. Alternatively, the user may want to know which locations are accessed but not care about the data value. Information about whether address or data are omitted is not encoded in the packets themselves as it does not change dynamically, and to do so would reduce encoding efficiency. The run-time configuration should be reported in the Format 3, subformat 3 support packet (see [\[sec:format33\]](#sec:format33)). The following sections include examples for all three cases. As outlined in [\[sec:DataInterfaceRequirements\]](#sec:DataInterfaceRequirements), two different signaling protocols between the RISC-V hart and the encoder are supported: _unified_ and _split_. Accordingly, both unified and split trace packets are defined. | | In the following tables, "clog2" is an abbreviation for "ceiling of log2". | | ----------------------------------------------------------------------------- | ### [](#sec:data-loadstore)8.1.1\. Load and Store #### [](#sec:loadstore-format)8.1.1.1\. format field Types of data trace packets are differentiated by the **format** field. This field is 2 bits wide if only unified loads and stores are supported, or 3 bits otherwise. Unified loads and split load request phase share the same code because the encoder will support one or the other, indicated by a discoverable parameter. Data accesses aligned to their size (e.g. 32-bit loads aligned to 32-bit word boundaries) are expected to be commonplace, and in such cases, encoding efficiency can be improved by not reporting the redundant LSBs of the address. __Table 1\. Packet format for Unified load or store, with address and data__ | **Field name** | **Bits** | **Description** | | -------------- | --------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 2 or 3 | Transaction type:000: Unified load or split load address, aligned001: Unified load or split load address, unaligned010: Store, aligned address011: Store, unaligned address(other codes select other packet formats) | | **size** | max(1, clog2(clog2( _data\_width\_p_/8 + 1))) | Transfer size is 2**size** bytes | | **diff** | 2 | 00: Full address and data (sync)01: Differential address, XOR-compressed data10: Differential address, full data11: Differentail address, differential data | | **data\_len** | **size** | Number of bytes of data is **data\_len** \+ 1 | | **data** | 8 \* (**data\_len** \+ 1) | Data | | **address** | _daddress\_width\_p_ | Byte address if format is unaligned, otherwise shift left by **size** to recover byte address | __Table 2\. Packet format for Unified load or store, with address only__ | **Field name** | **Bits** | **Description** | | -------------- | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 2 or 3 | Transaction type000: Unified load or split load address, aligned001: Unified load or split load address, unaligned010: Store, aligned address011: Store, unaligned address(other codes select other packet formats) | | **size** | max(1, clog2(clog2( _data\_width\_p_/8 + 1))) | Transfer size is 2**size** bytes | | **diff** | 1 | 0: Full address (sync)1: Differential address | | **address** | _daddress\_width\_p_ | Byte address if format is unaligned, otherwise shift left by **size** to recover byte address | __Table 3\. Packet format for Unified load or store, with data only__ | **Field name** | **Bits** | **Description** | | -------------- | --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 2 or 3 | Transaction type000: Unified load or split load address, aligned001: Unified load or split load address, unaligned010: Store, aligned address011: Store, unaligned address(other codes select other packet formats) | | **size** | max(1, clog2(clog2( _data\_width\_p_/8 + 1))) | Transfer size is 2**size** bytes | | **diff** | 1 or 2 | 00: Full data (sync)01: Compressed data (XOR if 2 bits)10: reserved11 : Differential data | | **data** | _data\_width\_p_ | Data | __Table 4\. Packet format for Split load - Address only__ | **Field name** | **Bits** | **Description** | | -------------- | --------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 3 | Transaction type000: Unified load or split load address, aligned001: Unified load or split load address, unaligned(other codes select other packet formats) | | **size** | max(1, clog2(clog2( _data\_width\_p_/8 + 1))) | Transfer size is 2**size** bytes | | **lrid** | _lrid\_width\_p_ | Load request ID | | **diff** | 1 | 0: Full address (sync)1: Differential address | | **address** | _daddress\_width\_p_ | Byte address if format is unaligned, otherwise shift left by **size** to recover byte address | __Table 5\. Packet format for Split load - Data only__ | **Field name** | **Bits** | **Description** | | -------------- | --------------------------------------------- | ----------------------------------------------------------------------------- | | **format** | 3 | Transaction type100: split load data(other codes select other packet formats) | | **size** | max(1, clog2(clog2( _data\_width\_p_/8 + 1))) | Transfer size is 2**size** bytes | | **lrid** | _lrid\_width\_p_ | Load request ID | | **resp** | 2 | 00: Error (no data)01: XOR-compressed data10: Full data11: Differential data | | **data** | _data\_width\_p_ | Data | #### [](#sec:loadstore-size)8.1.1.2\. size field The width of this field is 2 bits if max size is 64-bits (_data\_width\_p_< 128), 3 bits if wider. #### [](#sec:loadstore-diff)8.1.1.3\. diff field Unlike instruction trace, compression options for data trace are somewhat limited. Following a synchronization instruction trace packet, the first data trace packet for a given access size must include the full (unencoded) data access address. Thereafter, the address may be reported differentially (i.e. address of this data access, minus the address of the previous data access of the same size). Similarly, following a synchronization instruction trace packet, the first data trace packet for a given access size must include the full (unencoded) data value. Beyond this, data may be encoded or unencoded depending on whichever results in the most efficient represenation. Implementors may chose to offer one of XOR or differential compression, or both. XOR compression will be simpler to implement, and avoids the need for performing subtraction of large values. If only one data compression type is offered, the **diff** field can be 1 bit wide rather than 2 for [Table 3](#tab:te%5Fdatadx0y2). #### [](#sec:loadstore-datalen)8.1.1.4\. data\_len field However the data is compressed, upper bytes that are all the same value do not need to be included in the packet; the decoder can recreate the full-width value by sign extending from the most significant received bit. In cases where **data** is not the final field in the packet, the width of **data** is indicated by this field. ### [](#sec:data-atomic)8.1.2\. Atomic #### [](#sec:atomic-size)8.1.2.1\. size field Strictly, **size** could be just one bit as atomics are currently either 32 or 64 bits. Defining as per regular loads and stores provisions for future extensions (proprietary or otherwise) that support smaller atomics. __Table 6\. Packet format for Unified atomic with address and data__ | **Field name** | **Bits** | **Description** | | -------------- | --------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 3 | Transaction type110: Unified atomic or split atomic address(other codes other packet formats) | | **subtype** | 3 | Atomic sub-type000: Swap001: ADD010: AND011: OR100: XOR101: MAX110: MIN111: reserved | | **size** | max(1, clog2(clog2( _data\_width\_p_/8 + 1))) | Transfer size is 2**size** bytes | | **diff** | 2 | 00: Full address and data (sync)01: Differential address, XOR-compressed data10: Differential address, full data11: Differential address, differential data | | **op\_len** | **size** | Number of bytes of operand is **op\_len** \+ 1 | | **operand** | 8 \* (**op\_len** \+ 1) | Operand. Value from rs2 before operator applied | | **data\_len** | **size** | Number of bytes of data is **data\_len** \+ 1 | | **data** | 8 \* (**data\_len** \+ 1) | Data | | **address** | _daddress\_width\_p_ | Address, aligned and encoded as per size | __Table 7\. Packet format for Unified atomic with address only__ | **Field name** | **Bits** | **Description** | | -------------- | --------------------------------------------- | ----------------------------------------------------------------------------------------------------- | | **format** | 3 | Transaction type110: Unified atomic or split atomic address(other codes other packet formats) | | **subtype** | 3 | Atomic sub-type000: Swap001: ADD010: AND011: OR100: XOR101: MAX110: MIN111: conditional store failure | | **size** | max(1, clog2(clog2( _data\_width\_p_/8 + 1))) | Transfer size is 2**size** bytes | | **diff** | 1 | 0: Full address1: Differential address | | **address** | _daddress\_width\_p_ | Address, aligned and encoded as per size | __Table 8\. Packet format for Unified atomic with data only__ | **Field name** | **Bits** | **Description** | | -------------- | --------------------------------------------- | --------------------------------------------------------------------------------------------- | | **format** | 3 | Transaction type110: Unified atomic or split atomic address(other codes other packet formats) | | **subtype** | 3 | Atomic sub-type000: Swap001: ADD010: AND011: OR100: XOR101: MAX110: MIN111: reserved | | **size** | max(1, clog2(clog2( _data\_width\_p_/8 + 1))) | Transfer size is 2**size** bytes | | **diff** | 1 or 2 | 00: Full data (sync)01: Compressed data (XOR if 2 bits)10: reserved11: Differential data | | **op\_len** | **size** | Number of bytes of operand is **op\_len** \+ 1 | | **operand** | 8 \* (**op\_len** \+ 1) | Operand. Value from rs2 before operator applied | | **data** | _data\_width\_p_ | Data | #### [](#sec:atomic-diff)8.1.2.2\. diff field See [8.1.1.3\. diff field](#sec:loadstore-diff). #### [](#sec:atomic-operand)8.1.2.3\. operand field The operand value for the atomic operation. Uncompressed, although upper bytes that are all the same value do not need to be included in the packet; the decoder can recreate the full-width value by sign extending from the most significant received bit; see [8.1.2.4\. data\_len and op\_len fields](#sec:atomic-datalen). #### [](#sec:atomic-datalen)8.1.2.4\. data\_len and op\_len fields Width of **data and \*operand** fields respectively. See **[8.1.1.4\. data\_len field](#sec:loadstore-datalen).** __Table 9\. Packet format for Split atomic with operand only__ | **Field name** | **Bits** | **Description** | | -------------- | --------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 3 | Transaction type110: Unified atomic or split atomic address(other codes other packet formats) | | **subtype** | 3 | Atomic sub-type000: Swap001: ADD010: AND011: OR100: XOR101: MAX110: MIN111: reserved | | **size** | max(1, clog2(clog2( _data\_width\_p_/8 + 1))) | Transfer size is 2**size** bytes | | **lrid** | _lrid\_width\_p_ | Load request ID | | **diff** | 1 or 2 | 00: Full address and data (sync)01: Differential address, XOR-compressed data10: Differential address, full data11: Differential address, differential data | | **op\_len** | **size** | Number of bytes of operand is **op\_len** \+ 1 | | **operand** | 8 \* (**op\_len** \+ 1) | Operand. Value from rs2 before operator applied | | **address** | _daddress\_width\_p_ | Address, aligned and encoded as per size | __Table 10\. Packet format for Split atomic load data only__ | **Field name** | **Bits** | **Description** | | -------------- | ------------------------- | ----------------------------------------------------------------------------------------------------------------- | | **format** | 3 | Transaction type110: Split atomic data other codes other packet formats | | **lrid** | _lrid\_width\_p_ | Load request ID | | **resp** | 2 | 00: Error (no data)01: XOR-compressed data10: full data11: differential data | | **data\_len** | **size** | Number of bytes of operand is _data\_len + 1_. Not included if resp indicates an error (sign-extend **resp** MSB) | | **data** | 8 \* (**data\_len** \+ 1) | Data. Not included if resp indicates an error (sign-extend **resp** MSB) | ### [](#sec:data-csr)8.1.3\. CSR __Table 11\. Packet format for Unified CSR, with address, data and operand__ | **Field name** | **Bits** | **Description** | | -------------- | ----------------------- | ----------------------------------------------------------------------------------------- | | **format** | 3 | Transaction type101: CSR(other codes other packet formats) | | **subtype** | 2 | CSR sub-type00: RW01: RS10: RC11: reserved | | **diff** | 1 or 2 | 00: Full data (sync)01: Compressed data (XOR if 2 bits)10: reserved11 : Differential data | | **data\_len** | 2 or 3 | Number of bytes of data is **data\_len** \+ 1 | | **data** | 8 \* (**data\_len** 1) | Data | | **addr\_msbs** | 6 | Address\[11:6\] | | **op\_len** | 2 or 3 | Number of bytes of operand is **op\_len** \+ 1 | | **operand** | 8 \* (**op\_len** \+ 1) | Operand. Value from rs1 before operator applied | | **addr\_lsbs** | 6 | Address\[5:0\] | #### [](#sec:csr-diff)8.1.3.1\. diff field See [8.1.1.3\. diff field](#sec:loadstore-diff). #### [](#sec:csr-operand)8.1.3.2\. operand field See [8.1.2.3\. operand field](#sec:atomic-operand). #### [](#sec:csr-datalen)8.1.3.3\. data\_len and op\_len fields 2 bits wide if hart has 32-bit CSRs, 3 bits if 64-bit. Width of **data**and **operand** fields respectively. See [8.1.1.4\. data\_len field](#sec:loadstore-datalen). #### [](#sec:csr-addr)8.1.3.4\. addr fields The address is split into two parts, with the 6 LSBs output last as these are more likely to compress away. __Table 12\. Packet format for Unified CSR, with address and read-only data (as determined by addr\[11:10\] = 11)__ | **Field name** | **Bits** | **Description** | | -------------- | ------------------------- | ----------------------------------------------------------------------------------------- | | **format** | 3 | Transaction type101: CSR other codes other packet formats | | **subtype** | 2 | CSR sub-type00: RW01: RS10: RC11: reserved | | **diff** | 1 or 2 | 00: Full data (sync)01: Compressed data (XOR if 2 bits)10: reserved11 : Differential data | | **data\_len** | 2 or 3 | Number of bytes of data is **data\_len** \+ 1 | | **data** | 8 \* (**data\_len** \+ 1) | Data | | **addr\_msbs** | 6 | Address\[11:6\] | | **addr\_lsbs** | 6 | Address\[5:0\] | __Table 13\. Packet format for Unified CSR, with address only__ | **Field name** | **Bits** | **Description** | | -------------- | -------- | -------------------------------------------------------- | | **format** | 3 | Transaction type101: CSRother codes other packet formats | | **subtype** | 3 | CSR sub-type00: RW01: RS10: RC11: reserved | | **diff** | 0 or 1 | 0: Full address1: Differential address | | **addr\_msbs** | 6 | Address\[11:6\] | | **addr\_lsbs** | 6 | Address\[5:0\] | 11.1. Decoder ==================== ## [](#Decoder)11.1\. Decoder This decoder implementation assumes there is no branch predictor or return address stack (_return\_stack\_size\_p_ and _bpred\_size\_p_ both zero). Reference Python implementations of both the encoder and decoder can be found at . ### [](#11-1-1-decoder-pseudo-code)11.1.1\. Decoder pseudo code # global variables global pc # Reconstructed program counter global last_pc # PC of previous instruction global branches = 0 # Number of branches to process global branch_map = 0 # Bit vector of not taken/taken (1/0) status # for branches global bool stop_at_last_branch = FALSE # Flag to indicate reconstruction is to end at # the final branch global bool inferred_address = FALSE # Flag to indicate that reported address from # format 0/1/2 was not following an uninferable # jump (and is therefore inferred) global bool start_of_trace = TRUE # Flag indicating 1st trace packet still # to be processed global address # Reconstructed address from te_inst messages global privilege # Privilege from te_inst messages global options # Operating mode flags global array return_stack # Array holding return address stack global irstack_depth = 0 # Depth of the return address stack # Process te_inst packet. Call each time a te_inst packet is received # function process_te_inst (te_inst) if (te_inst.format == 3) if (te_inst.subformat == 3) # Support packet process_support(te_inst) return if (te_inst.subformat == 2) # Context packet return if (te_inst.subformat == 1) # Trap packet report_trap(te_inst) if (!te_inst.interrupt) # Exception report_epc(exception_address(te_inst)) if (!te_inst.thaddr) # Trap only - nothing retired return inferred_address = FALSE address = (te_inst.address << discovery_response.iaddress_lsb) if (te_inst.subformat == 1 or start_of_trace) branches = 0 branch_map = 0 if (is_branch(get_instr(address))) # 1 unprocessed branch if this instruction is a branch branch_map = branch_map | (te_inst.branch << branches) branches++ if (te_inst.subformat == 0 and !start_of_trace) follow_execution_path(address, te_inst) else pc = address report_pc(pc) last_pc = pc # previous pc not known but ensures correct # operation for is_sequential_jump() privilege = te_inst.privilege start_of_trace = FALSE irstack_depth = 0 else # Duplicated at top of next page to show continuity else # Duplicate of last line from previous page to show continuity if (start_of_trace) # This should not be possible! ERROR: Expecting trace to start with format 3 return if (te_inst.format == 2 or te_inst.branches != 0) stop_at_last_branch = FALSE if (options.full_address) address = (te_inst.address << discovery_response.iaddress_lsb) else address += (te_inst.address << discovery_response.iaddress_lsb) if (te_inst.format == 1) stop_at_last_branch = (te_inst.branches == 0) # Branch map will contain <= 1 branch (1 if last reported instruction was a branch) branch_map = branch_map | (te_inst.branch_map << branches) if (te_inst.branches == 0) branches += 31 else branches += te_inst.branches follow_execution_path(address, te_inst) # Follow execution path to reported address # function follow_execution_path(address, te_inst) local previous_address = pc local stop_here = FALSE while (TRUE) if (inferred_address) # iterate again from previously reported address to # find second occurrence stop_here = next_pc(previous_address) report_pc(pc) if (stop_here) inferred_address = FALSE else stop_here = next_pc(address) report_pc(pc) if (branches == 1 and is_branch(get_instr(pc)) and stop_at_last_branch) # Reached final branch - stop here (do not follow to next instruction as # we do not yet know whether it retires) stop_at_last_branch = FALSE return if (stop_here) # Reached reported address following an uninferable discontinuity - stop here if (unprocessed_branches(pc)) ERROR: unprocessed branches return if (te_inst.format != 3 and pc == address and !stop_at_last_branch and (te_inst.notify != get_preceding_bit(te_inst, "notify")) and !unprocessed_branches(pc)) # All branches processed, and reached reported address due to notification, # not as an uninferable jump target return if (te_inst.format != 3 and pc == address and !stop_at_last_branch and !is_uninferable_discon(get_instr(last_pc)) and (te_inst.updiscon == get_preceding_bit(te_inst, "updiscon")) and !unprocessed_branches()) and ((te_inst.irreport == get_previous_bit(te_inst, "irreport")) or te_inst.irdepth == irstack_depth)) # All branches processed, and reached reported address, but not as an # uninferable jump target # Stop here for now, though flag indicates this may not be # final retired instruction inferred_address = TRUE return if (te_inst.format == 3 and pc == address and !unprocessed_branches(pc) and (te_inst.privilege == privilege or is_return_from_trap(get_instr(last_pc)))) # All branches processed, and reached reported address return # Compute next PC # function next_pc (address) local instr = get_instr(pc) local this_pc = pc local stop_here = FALSE if (is_inferable_jump(instr)) pc += instr.imm else if (is_sequential_jump(instr, last_pc)) # lui/auipc followed by # jump using same register pc = sequential_jump_target(pc, last_pc) else if (is_implicit_return(instr)) pc = pop_return_stack() else if (is_uninferable_discon(instr)) if (stop_at_last_branch) ERROR: unexpected uninferable discontinuity else pc = address stop_here = TRUE else if (is_taken_branch(instr)) pc += instr.imm else pc += instruction_size(instr) if (is_call(instr)) push_return_stack(this_pc) last_pc = this_pc return stop_here # Process support packet # function process_support (te_inst) local stop_here = FALSE options = te_inst.options if (te_inst.qual_status != no_change) start_of_trace = TRUE # Trace ended, so get ready to start again if (te_inst.qual_status == ended_ntr and inferred_address) local previous_address = pc inferred_address = FALSE while (TRUE) stop_here = next_pc(previous_address) report_pc(pc) if (stop_here) return return # Determine if instruction is a branch, adjust branch count/map, # and return taken status # function is_taken_branch (instr) local bool taken = FALSE if (!is_branch(instr)) return FALSE if (branches == 0) ERROR: cannot resolve branch else taken = !branch_map[0] branches-- branch_map >> 1 return taken # Determine if instruction is a branch # function is_branch (instr) if ((instr.opcode == BEQ) or (instr.opcode == BNE) or (instr.opcode == BLT) or (instr.opcode == BGE) or (instr.opcode == BLTU) or (instr.opcode == BGEU) or (instr.opcode == C.BEQZ) or (instr.opcode == C.BNEZ)) return TRUE return FALSE # Determine if instruction is an inferable jump # function is_inferable_jump (instr) if ((instr.opcode == JAL) or (instr.opcode == C.JAL) or (instr.opcode == C.J)) return TRUE return FALSE # Determine if instruction is an uninferable jump # function is_uninferable_jump (instr) if ((instr.opcode == JALR and instr.rs1 != 0) or (instr.opcode == C.JALR) or (instr.opcode == C.JR)) return TRUE return FALSE # Determine if instruction is a return from trap # function is_return_from_trap (instr) if ((instr.opcode == URET) or (instr.opcode == SRET) or (instr.opcode == MRET) or (instr.opcode == DRET)) return TRUE return false # Determine if instruction is an uninferrable discontinuity # function is_uninferrable_discon (instr) if (is_uninferrable_jump(instr) or is_return_from_trap (instr)) return TRUE return FALSE # Determine if instruction is a sequentially inferable jump # function is_sequential_jump (instr, prev_addr) if (not (is_uninferable_jump(instr) and options.sijump)) return FALSE local prev_instr = get_instr(prev_addr) if((prev_instr.opcode == AUIPC) or (prev_instr.opcode == LUI) or (prev_instr.opcode == C.LUI)) return (instr.rs1 == prev_instr.rd) return FALSE # Find the target of a sequentially inferable jump # function sequential_jump_target (addr, prev_addr) local instr = get_instr(addr) local prev_instr = get_instr(prev_addr) local target = 0 if (prev_instr.opcode == AUIPC) target = prev_addr target += prev_instr.imm if (instr.opcode == JALR) target += instr.imm return target # Determine if instruction is a call # # - excludes tail calls as they do not push an address onto the return stack function is_call (instr) if ((instr.opcode == JALR and instr.rd == 1) or (instr.opcode == C.JALR) or (instr.opcode == JAL and instr.rd == 1) or (instr.opcode == C.JAL)) return TRUE return FALSE # Determine if instruction return address can be implicitly inferred # function is_implicit_return (instr) if (options.implicit_return == 0) # Implicit return mode disabled return FALSE if ((instr.opcode == JALR and instr.rs1 == 1 and instr.rd == 0) or (instr.opcode == C.JR and instr.rs1 == 1)) if ((te_inst.irreport != get_preceding_bit(te_inst, "irreport")) and te_inst.irdepth == irstack_depth) return FALSE return (irstack_depth > 0) return FALSE #Check for unprocessed branches # function unprocessed_branches (address) # Check all branches processed (except 1 if this instruction is a branch) return (branches != (is_branch(get_instr(address)) ? 1 : 0)) # Push address onto return stack # function push_return_stack (address) if (options.implicit_return == 0) # Implicit return mode disabled return local irstack_depth_max = discovery_response.return_stack_size ? 2**discovery_response.return_stack_size : 2**discovery_response.call_counter_size local instr = get_instr(address) local link = address if (irstack_depth == irstack_depth_max) # Delete oldest entry from stack to make room for new entry added below irstack_depth-- for (i = 0; i < irstack_depth; i++) return_stack[i] = return_stack[i+1] link += instruction_size(instr) return_stack[irstack_depth] = link irstack_depth++ return # Pop address from return stack # function pop_return_stack () irstack_depth-- # function not called if irstack_depth is 0, so no need # to check for underflow local link = return_stack[irstack_depth] return link # Return the address of an exception # function exception_address(te_inst) local instr = get_instr(pc) if (is_uninferable_discon(instr) and !te_inst.thaddr) return te_inst.address if (instr.opcode == ECALL) or (instr.opcode == EBREAK) or (instr.opcode == C.EBREAK)) return pc return next_pc(pc) # Report ecause and tval (user to populate if desired) # function report_trap(te_inst) return # Report program counter value (user to populate if desired) # function report_pc(address) return # Report exception program counter value (user to populate if desired) # function report_epc(address) return 10.1. Parameters and Discovery ==================== ## [](#10-1-parameters-and-discovery)10.1\. Parameters and Discovery This document defines a number of parameters for describing aspects of the encoder such as the widths of buses, the presence or absence of optional features and the size of resources, as listed in[Table 1](#tab:iparameters) and [Table 2](#tab:dparameters). Depending on the implementation, some parameters may be inherently fixed whilst others may be passed in to the design by some means. __Table 1\. Parameters to the encoder - instruction trace__ | **Parameter name** | **Range** | **Description** | | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | | _arch\_p_ | The architecture specification version with which the encoder is compliant (0 for initial version). | | | _blocks\_p_ | Number of times **iretire**, **itype** etc. are replicated | | | _bpred\_size\_p_ | Number of entries in the branch predictor is 2**bpred\_size\_p**. Minimum number of entries is 2, so a value of 0 indicates that there is no branch predictor implemented. | | | _cache\_size\_p_ | Number of entries in the jump target cache is 2**cache\_size\_p**. Minimum number of entries is 2, so a value of 0 indicates that there is no jump target cache implemented. | | | _call\_counter\_size\_p_ | Number of bits in the nested call counter is 2**call\_counter\_size\_p**. Minimum number of entries is 2, so a value of 0 indicates that there is no implicit return call counter implemented. | | | _ctype\_width\_p_ | Width of the **ctype** bus | | | _context\_width\_p_ | Width of **context** bus | | | _time\_width\_p_ | Width of **time** bus | | | _ecause\_width\_p_ | Width of **exception cause** bus | | | _ecause\_choice\_p_ | Number of bits of exception cause to match using multiple choice | | | _f0s\_width\_p_ | Width of the **subformat** field in format 0 _te\_inst_packets (see [\[sec:f0s\]](#sec:f0s)). | | | _filter\_context\_p_ | 0 or 1 | Filtering on context supported when 1 | | _filter\_time\_p_ | 0 or 1 | Filtering on time supported when 1 | | _filter\_excint\_p_ | Filtering on exception cause or interrupt supported when non\_zero. Number of nested exceptions supported is 2filter\_excint\_p | | | _filter\_privilege\_p_ | 0 or 1 | Filtering on privilege supported when 1 | | _filter\_tval\_p_ | 0 or 1 | Filtering on trap value supported when 1 (provided _filter\_excint\_p_ is non-zero) | | _iaddress\_lsb\_p_ | LSB of instruction **address** bus to trace. 1 if compressed instructions are supported, 2 otherwise | | | _iaddress\_width\_p_ | Width of instruction **address** bus. This is the same as _DXLEN_ | | | _iretire\_width\_p_ | Width of the **iretire** bus | | | _ilastsize\_width\_p_ | Width of the **ilastsize** bus | | | _itype\_width\_p_ | Width of the **itype** bus | | | _nocontext\_p_ | 0 or 1 | Exclude context from _te\_inst_ packets if 1 | | _notime\_p_ | 0 or 1 | Exclude time from _te\_inst_ packets if 1 | | _privilege\_width\_p_ | Width of **privilege** bus | | | _retires\_p_ | Maximum number of instructions that can be retired per block | | | _return\_stack\_size\_p_ | Number of entries in the return address stack is 2**return\_stack\_size\_p**. Minimum number of entries is 2, so a value of 0 indicates that there is no implicit return stack implemented. | | | _sijump\_p_ | 0 or 1 | **sijump** is used to identify sequentially inferable jumps | | _impdef\_width\_p_ | Width of **implementation-defined input** bus | | __Table 2\. Parameters to the encoder - data trace__ | **Parameter name** | **Range** | **Description** | | ----------------------- | -------------------------------- | --------------- | | _daddress\_width\_p_ | Width of the **daddress** bus | | | _dblock\_width\_p_ | Width of the **dblock** bus | | | _data\_width\_p_ | Width of the **data** bus | | | _dsize\_width\_p_ | Width of the **dsize** bus | | | _dtype\_width\_p_ | Width of the **dtype** bus | | | _iaddr\_lsbs\_width\_p_ | Width of the **iaddr\_lsbs** bus | | | _lrid\_width\_p_ | Width of the **lrid** bus | | | _lresp\_width\_p_ | Width of the **lresp** bus | | | _ldata\_width\_p_ | Width of the **ldata** bus | | | _sdata\_width\_p_ | Width of the **sdata** bus | | ### [](#sec:disco)10.1.1\. Discovery of encoder parameters To operate correctly, the decoder must be able to determine some of the encoder’s parameters at runtime, in the form of discoverable attributes. These parameters must be discoverable by the decoder, or else be fixed at the default value (in other words, if an encoder does not make a particular parameter discoverable, it must implement only the default value of that parameter, which the decoder will also use). [Table 3](#tab:requiredAttributes) lists the required discoverable attributes for instruction trace. To access the discoverable attributes, some external entity, for example a debugger or a supervisory hart, must request it from the encoder. The encoder will provide the discovery information in one or more different formats. The preferred format is a packet which is sent over the trace infrastructure. Another format would be allowing the external entity to read the values from some register or memory mapped space maintained by the encoder. [10.1.2\. Example ipxact description](#sec:ipxact) gives an example of how this may be accomplished. __Table 3\. Required instruction trace attributes__ | **Name** | **Default** | **Parameter mapping** | | --------------------- | ----------- | -------------------------- | | _arch_ | 0 | _arch\_p_ | | _bpred\_size_ | 0 | _bpred\_size\_p_ | | _cache\_size_ | 0 | _cache\_size\_p_ | | _call\_counter\_size_ | 0 | _call\_counter\_size\_p_ | | _context\_width_ | 0 | _context\_width\_p_ \- 1 | | _time\_width_ | 0 | _time\_width\_p_ \- 1 | | _ecause\_width_ | 3 | _ecause\_width\_p_ \- 1 | | _f0s\_width_ | 0 | _f0s\_width\_p_ | | _iaddress\_lsb_ | 0 | _iaddress\_lsb\_p_ \- 1 | | _iaddress\_width_ | 31 | _iaddress\_width\_p_ \- 1 | | _nocontext_ | 1 | _nocontext_ | | _notime_ | 1 | _notime_ | | _privilege\_width_ | 1 | _privilege\_width\_p_ \- 1 | | _return\_stack\_size_ | 0 | _return\_stack\_size\_p_ | | _sijump_ | 0 | _sijump\_p_ | For ease of use it is further recommended that all of the encoder’s parameters be mapped to discoverable attributes, even if not directly required by the decoder. In particular, attributes related to filtering capabilities. [Table 4](#tab:optionalAttributes)lists the attributes associated with the filtering recommendations discussed in [\[ch:filtering\]](#ch:filtering), [Table 5](#tab:otherAttributes) lists attributes related to other instruction trace parameters mentioned in this document, and [Table 6](#tab:dataAttributes) lists attributes related to data trace. __Table 4\. Optional filtering attributes__ | **Name** | **Default** | **Parameter mapping** | | ------------------- | ----------- | --------------------- | | _comparators_ | 0 | _comparators\_p_ \- 1 | | _filters_ | 0 | _filters\_p_ \- 1 | | _ecause\_choice_ | 5 | _ecause\_choice\_p_ | | _filter\_context_ | 1 | _filter\_context\_p_ | | _filter\_time_ | 1 | _filter\_time\_p_ | | _filter\_excint_ | 1 | _filter\_excint\_p_ | | _filter\_privilege_ | 1 | _filter\_privilegep_ | | _filter\_tval_ | 1 | _filter\_tval\_p_ | __Table 5\. Other recommended attributes__ | **Name** | **Default** | **Description** | | ------------------ | ----------- | -------------------------- | | _ctype\_width_ | 0 | _ctype\_width\_p_ \- 1 | | _ilastsize\_width_ | 0 | _ilastsize\_width\_p_ \- 1 | | _itype\_width_ | 3 | _itype\_width\_p_ \- 1 | | _iretire\_width_ | 1 | _iretire\_width\_p_ \- 1 | | _retires_ | 0 | _retires\_p_ \- 1 | | _impdef\_width_ | 0 | _impdef\_width\_p_ \- 1 | __Table 6\. Data trace attributes__ | **Name** | **Default** | **Description** | | -------------------- | ----------- | ---------------------------- | | _daddress\_width_ | 31 | _daddress\_width\_p_ \- 1 | | _dblock\_width_ | 0 | _dblock\_width\_p_ \- 1 | | _data\_width_ | 31 | _data\_width\_p_ \- 1 | | _dsize\_width_ | 2 | _dsize\_width\_p_ \- 1 | | _dtype\_width_ | 0 | _dtype\_width\_p_ \- 1 | | _iaddr\_lsbs\_width_ | 0 | _iaddr\_lsbs\_width\_p_ \- 1 | | _lrid\_width_ | 0 | _lrid\_width\_p_ \- 1 | | _lresp\_width_ | 0 | _lresp\_width\_p_ \- 1 | | _ldata\_width_ | 31 | _ldata\_width\_p_ \- 1 | | _sdata\_width_ | 31 | _sdata\_width\_p_ \- 1 | ### [](#sec:ipxact)10.1.2\. Example ipxact description This section provides an example of discovery information represented in the ipxact form. ```xml Siemens TraceEncoder TraceEncoder 0.8 TraceEncoderRegisterMap >TraceEncoderRegisterAddressBlock 0 128 64 discovery_info_0 'h0 64 read-only version text 0 4 minor_revision text 4 4 arch text 8 4 bpred_size text 12 4 cache_size text 16 4 call_counter_size text 20 3 comparators text 23 3 context_type_width text 26 5 context_width text 31 5 ecause_choice text 36 3 ecause_width text 39 4 filters text 43 4 filter_context text 47 1 filter_excint text 48 4 filter_privilege text 52 1 filter_tval text 53 1 filter_impdef text 54 1 f0s_width text 55 2 iaddress_lsb text 57 2 discovery_info_1 'h4 64 read-only iaddress_width text 0 7 ilastsize_width text 7 7 itype_width text 14 7 iretire_width text 21 7 nocontext text 28 1 privilege_width text 29 2 retires text 31 3 return_stack_size text 34 4 sijump text 38 1 taken_branches text 39 4 impdef_width text 43 5 8 ``` 9.1. Reference Compressed Branch Trace Algorithm ==================== ## [](#Algorithm)9.1\. Reference Compressed Branch Trace Algorithm The contents of this chapter are informative only. A reference algorithm for compressed branch trace is given in[Figure 1](#fig:algo). In the diagram, the following terms are used: * _te\_inst._ The name of the packet type emitted by the encoder (see[\[packets\]](#packets)); * _inst._ Abbreviation for 'instruction'; * _trap. Exception or interrupt signalled;_ * _updiscon._ Uninferable PC discontinuity. This identifies an instruction that causes the program counter to be changed by an amount that cannot be predicted from the source code alone (**itype** values 8, 10, 12 or 14); * _Qualified?_ An instruction that meets the filtering criteria is qualified, and will be traced; * _Branch?_ Is the instruction a branch or not (**itype** values 4 or 5); * _branch map._ A vector where each bit represents the outcome of a branch. A 0 indicates the branch was taken, a 1 indicates that it was not; * _ppccd._ Privilege has changed, or context has changed and needs to be reported precisely or treated as an uninferable PC discontinuity (see[\[tab:context-type\]](#tab:context-type)); * _ppccd\_br._ As above, but branch map not empty; * _ntf._ Trace notify trigger (see [\[tab:debugModuleTriggerSupport\]](#tab:debugModuleTriggerSupport)); * _cci._ context change that can be reported imprecisely (see[\[tab:context-type\]](#tab:context-type)); * _rpt\_br._ Report branches due to full branch map or misprediction; * _branches._ The number of branches encountered but not yet reported to the decoder; * _pbc._ Correctly predicted branches count (always zero if branch predictor disabled or not present); * _trep_ Previous trap already reported with **thaddr** \= 0 because it was preceded by an updiscon or immediately followed by another exception; * _resync count._ A counter used to keep track of when it is necessary to send a synchronization packet (see [9.1.2\. Resynchronisation](#sec:resync)); * _resync2\_br._ The resync FSM is in state 2 and there are entries in the branch map that have not yet been output (see[9.1.2\. Resynchronisation](#sec:resync)). * _resync3._ The resync FSM is in state 3 (see [9.1.2\. Resynchronisation](#sec:resync)); [Figure 1](#fig:algo) shows instruction by instruction behavior, as would be seen in a single-retirement system only. Whilst the core to encoder interface allows the RISC-V hart to provide information on multiple retiring instructions simultaneously, the resultant packet sequence generated by the encoder must be the same as if retiring one instruction at a time. Note that even with a single-retirement system it is possible to retire an instruction and report a trap simultaneously (**itype** \= 1 or 2 and **iretire** \= 1). In this case the flow diagram must be traversed twice, first for the retired instruction, and then for the trap. A 3-stage pipeline within the encoder is assumed, such that the encoder has visibility of the current, previous and next instructions. All packets are generated using information relating to the current instruction. The orange diamonds indicate decisions based on the previous instruction, the green diamond indicates a decision based on the next instruction, and all other diamonds are based on the current instruction. Additionally, the encoder can generate one further packet type, not shown on the diagram for clarity. The _support_ packet (format 3, subformat 3 - see [\[sec:format33\]](#sec:format33)) is sent when: * The encoder is enabled or disabled, or its configuration is changed, to inform the decoder of the operating mode of the encoder; * After the final qualified instruction has been traced, to inform the decoder that tracing has stopped; * If trace packets are lost (for example if the buffer into which packets are being written fills up), in this situation, the 1st packet loaded into the buffer when space next becomes available must be a_support_ packet. Following this, tracing will resume with a sync packet. Note: if the **halted** or **reset** sideband signals are asserted (see[\[tab:ingress-side-band\]](#tab:ingress-side-band)) the encoder will behave as if it has received an unqualified instruction (output _te\_inst_ reporting the address of the previous instruction, followed by_te\_support_); ![algo](_images/algo.png) Figure 1\. Instruction delta trace algorithm ### [](#9-1-1-format-selection)9.1.1\. Format selection In all cases but two, the packet format is determined only by a 'yes' outcome from the associated decision. When reporting branch information on its own (without an address), the choice between format 1 and format 0, subformat 0 depends on the number of correctly predicted branches (this will be 0 if the predictor is not supported, or is disabled). No packets are generated until there are at least 31 branches to report. Format 1 is used if the outcome of at least one of those 31 branches was not predicted correctly. If all were predicted correctly, nothing is output at this time, and the encoder continues to count correctly predicted branch outcomes. As soon as one of the branch outcomes is not correctly predicted, the encoder will output a format 0, subformat 0 packet. See also[\[sec:format0\]](#sec:format0). The choice between formats for the "format 0/1/2" case in the middle of the diagram also needs further explanation. * If the number of correctly predicted branches is 31 or more, then format 0, subformat 0 is always used; * Else, if the jump target cache is supported and enabled, and the address being reported is in the cache, then normally format 0, subformat 1 will be used, reporting the cache index associated with the address. This will include branch information if there are any branches to report. However, the encoder may chose to output the equivalent format 1 or 2 packet (containing the differential address, with or without branch information) if that will result in a shorter packet (see[\[sec:format0\]](#sec:format0)); * Else, if there are branches to report, format 1 is used, otherwise format 2. Packet formats 0, 1 and 2 are organized so that the address is usually the final field. Minimizing the number of bits required to represent the address reduces the total packet size and significantly improves efficiency. See [\[packets\]](#packets). ### [](#sec:resync)9.1.2\. Resynchronisation Per [\[sec:synchronization\]](#sec:synchronization), a format 3 synchronisation packet must be output after "a prolonged period of time". The exact mechanism for determining this is not specified, but options might be to count the number of _te\_inst_ packets emitted, or the number of clock cycles elapsed, since the previous synchronization message was sent. When the resync is required, the primary objective is to output a format 3 packet, so that the decoder can start tracing from that point without needing any of the history. However, if the decoder is already synced, then it is also required that it can continue to follow the execution path up to and through the format 3 packet seamlessly. As such, before outputting a format 3 packet, it is necessary to output a format 0/1 packet for the preceding instruction if there are any unreported branches (because format 3 does not contain a branch map). There are several supported options for incrementing the resync timer (packets, cycle or instruction half-words), and as such updates to the timer do not necessarily coincide with instruction retirements. A small, independent FSM can be used to ensure the required packets are output in the required order, as follows: * **State 1**: resync\_count < max\_resync. Transition to **State 2** when resync\_count >= max\_resync * **State 2**: output packet of unreported branches if required, transition to **state 3** * **State 3**: output sync packet, reset resync\_count and transition to **State 1**. The format 3 will be sent if the resync timer has been exceeded. On the cycle before this (when the resync timer value has been exactly reached), a format 1 will be generated if the branch map is not empty. ### [](#rec:multiretcon)9.1.3\. Multiple retirement considerations As noted earlier in this section, for a single-retirement system the reference algorithm is applied to each retired instruction. When instructions are retired in blocks, only the first and last instruction in a block need be considered, as all those in between are "uninteresting", and will have no effect on the encoder’s state (their route through [Figure 1](#fig:algo) does not pass through any of the rectangular boxes). In most cases, either the first or last instruction of a block (but not both) is interesting, meaning that the encoder does not need to generate more than one packet from a block. However, there are a few cases where this is not true, and it is possible that the encoder will need to generate two packets from the same block. For example, the first instruction in a block must generate a packet if it is the first traced instruction. However, if the block also indicates an exception or interrupt (**itype**\= 1 or 2), then the last instruction in the block must also generate a packet. As generating multiple packets per cycle would significatly complicate the encoder, and as situations such as this will only occur infrequently, some elastic buffering in the encoder is the preferred approach. This will allow subsequent blocks to be queued whilst the encoder generates two successive packets from a block. The encoder can drain the elastic buffer any time there is a cycle when the hart doesn’t report anything, or if there is a block with **itype** \= 0 (which is uninteresting to the encoder). There are pathological cases where consecutive blocks could require packets to be generated from both first and last instructions, but elastic buffering is only required if the blocks are also input on consecutive cycles. In practice there are very few cases where this can occur. The worst so far identified case is a variation on the example above, where the exception is an ecall, and that in turn encounters some other form of exception or interrupt in the first few instructions of the trap handler: * Block 1: **itype** \= 1 (ecall), **iretires** \> 1\. Generate packet from first instruction (first traced), and last instruction (last before ecall); * Block 2: **itype** \= 1 or 2 (some other exception or interrupt),**iretires** \> 0\. Generate packet from first instruction (ecall trap handler), and last instruction (last before other exception or interrupt); * Block 3: Generate packet from first instruction (other exception or interrupt trap handler) Because the ecall is known to the hart’s fetch unit and can be predicted, it may be possible for block 2 to occur the cycle after block 1\. However, it is reasonable to assume that the other exception or interrupt will not be predictable, and as a result there will be several cycles between blocks 2 and 3, which will allow the encoder to 'catch up'. It is recommended that encoders implement sufficient elastic buffering to handle this case, and if for some reason the elastic buffer overflows, it should issue a support packet indicating trace lost. 12.1. Example code and packets ==================== ## [](#12-1-example-code-and-packets)12.1\. Example code and packets In the following examples **_ret_** is referred to as uninferable, this is only true if implicit-return mode is off 1. **Call to debug\_printf(), from 80001a84, in main():** 00000000800019e8
: ........: ... 80001a80: f6d42423 {sw a3,-152(s0)} 80001a84: ef4ff0ef {jal x1,80001178} PC: 80001a84 →80001178 The target of the **_jal_** is inferable, thus NO te\_inst packet is sent. 0000000080001178 : 80001178: 7139 {addi sp,sp,-64} 8000117a: ... 2. Return from debug\_printf(): 80001186: ... 80001188: 6121 {addi sp,sp,64} 8000118a: 8082 {ret} PC: 8000118a →80001a88 The target of the **_ret_** is uninferable, thus a **_te\_inst_** packet IS sent:**_te\_inst_**\[format=2 (ADDR\_ONLY): address=0x80001a88, updiscon=0\] 80001a88: 00000597 {auipc a1,0x0}} 80001a8c: 65058593 {addi a1,a1,1616}} # 800020d8 3. **exiting from Func\_2(), with a final taken branch, followed by a _ret_** 00000000800010b6 : ........: .... 800010da: 4781 {li a5,0} 800010dc: 00a05863 {blez a0,800010ec} PC: 800010dc →800010ec, add branch TAKEN to branch\_map, but no packet sent yet. branches = 0; branch\_map = 0; branch\_map = 0 <: ........: .... 80001b8a: f4442603 {lw a2,-188(s0)} 80001b8e: .... 4. **3 branches, then a function return back to Proc\_1()** 0000000080001100 : ........: .... 80001112: c080 {sw s0,0(s1)} 80001114: 4785 {li a5,1} 80001116: 02f40463 {beq s0,a5,8000113e } PC: 80001116 →8000111a, add branch NOT taken to branch\_map, but no packet sent yet. branches = 0; branch\_map = 0; branch\_map = 1 <} PC: 8000111a →8000111c, add branch NOT taken to branch\_map, but no packet sent yet. branch\_map = 1 <} PC: 8000111e →8000115e, add branch TAKEN to branch\_map, but no packet sent yet. branch\_map = 0 <: ........: .... 80001258: 00093783 {ld a5,0(s2)} 8000125c: .... PC: 80001168 →80001258 The target of the **_ret_** is uninferable, thus a **_te\_inst_** packet is sent, with THREE branches in the branch\_map **_te\_inst_**\[ format=1 (DIFF\_DELTA): branches=3, branch\_map=0x3, address=0x80001258 ( \=0x148), updiscon=0 \] 5. **A complex example with 2 branches, 2 jal, and a ret** 00000000800011d6 : ........: .... 8000121c: 441c {lw a5,8(s0)} 8000121e: c795 {beqz a5,8000124a} PC: 8000121e →8000124a, add branch TAKEN to branch\_map, but no packet sent yet. branches = 0; branch\_map = 0; branch\_map = 0 < PC: 80001254 →80001100 The target of the **_jal_** is inferable, thus no **_te\_inst_** packet needs be sent. 0000000080001100 : 80001100: 1101 {addi sp,sp,-32} 80001102: e822 {sd s0,16(sp)} 80001104: e426 {sd s1,8(sp)} 80001106: ec06 {sd ra,24(sp)} 80001108: 842a {mv s0,a0} 8000110a: 84ae {mv s1,a1} 8000110c: fedff0ef {jal x1,800010f8} PC: 8000110c →800010f8 The target of the **_jal_** is inferable, thus no **_te\_inst_** packet needs to be sent. 00000000800010f8 : 800010f8: 1579 {addi a0,a0,-2} 800010fa: 00153513 {seqz a0,a0} 800010fe: 8082 {ret} PC: 800010fe →80001110 The target of the **_ret_** is uninferable, thus a **_te\_inst_** packet will be sent shortly. 0000000080001100 : ........: .... 80001110: c115 {beqz a0,80001134} 80001112: .... PC: 80001110 →80001112, add branch NOT TAKEN to branch\_map. branch\_map = 1 <, =, !=, etc) independently selectable for each comparator; * Secondary match value may be used as a mask for the primary comparator; * The two comparators can be combined in several ways: P, P&&S, !(P&&S), latch (set on P clear on S); * Each comparator can also be used to explcitly report a particular instruction address (i.e. generate a watchpoint). Each filter can specify filtering against instruction and optionally data trace inputs from the HART, and offers: * Require up to 3 run-time selectable comparator units to match; * Multiple choice selection for **priv** and **cause** inputs (and **dtype**if data trace is supported); * Masked matching for **interrupt** and **impdef** inputs. Allowing for up to 3 comparators allows for simultaneous matching on Address, Trap value and context (unlikely, but should not be architecturally precluded). The filtering configuration fields are detailed in [\[encoderControl\]](#encoderControl). These support the architecture described above, though will also support simpler implementations, for example where the comparator function is more tightly coupled with each filter, or where filtering is provided on only some inputs (such as just instruction address). 13.1. Code fragment and transport ==================== ## [](#fragments)13.1\. Code fragment and transport This section shows fragments of code, and associated data from one of the architectural tests in the repository. For the individual fragments the ingress signals are shown and the corresponding packets generated. It further shows how the packets are transported via on-chip transport fabric. The fragments shown below are extracted from the test whilst it is being executed. In order to give some context to the fragment of interest, code prior to and after the fragment is also given. ### [](#13-1-1-illegal-opcode-test)13.1.1\. Illegal Opcode test In this example the test executes an illegal opcode (at line labelled 14) and traps. We show the output from the patched spike execution in line 30\. The input signals to the encoder are shown in lines labelled 38-46\. The HART will have set the signals shown in line 42 when the illegal instruction is executed and as can be seen it is not retired. Lines labelled 53, 56 and 59 show the packets output from the encoder for this fragment. #### [](#13-1-1-1-code-fragment)13.1.1.1\. Code fragment 1: ************************************************************************************* 2: ****************** Fragment 0x80000222 - 0x80000226:illegal_opcode ****************** 3: ************************************************************************************* 4: KEY: ">" means pre-fragment execution, "<" means post-fragment execution 5: ^^^^^^^^^^^^^^^^^^^^^^^^^^ Part 1 of 1 ^^^^^^^^^^^^^^^^^^^^^^^^^^ 6: 7: elf: 8: > 0000000080000104 : 9: > 80000104: 00000297 auipc t0,0x0 10: > 80000108: 11e28293 addi t0,t0,286 # 80000222 11: > 8000010c: 8282 jr t0 12: > 80000154: 9282 jalr t0 13: 0000000080000222 : 14: 80000222: 0000 unimp 15: 80000224: 0000 unimp 16: 80000226: b709 j 80000128 17: < 00000000800001b0 : 18: < 800001b0: a805 j 800001e0 19: < 00000000800001e0 : 20: < 800001e0: 342023f3 csrr t2,mcause 21: < 800001e4: fff0031b addiw t1,zero,-1 22: < 800001e8: 137e slli t1,t1,0x3f 23: 24: trace_spike: 25: ******** Data from br_j_asm.spike_pc_trace line 5029 ******** 26: > ADDRESS=80000154, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 27: > ADDRESS=80000104, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 28: > ADDRESS=80000108, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 29: > ADDRESS=8000010c, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 30: ADDRESS=80000222, PRIVILEGE=3, EXCEPTION=1, ECAUSE=2, TVAL=0, INTERRUPT=0 31: < ADDRESS=800001b0, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 32: < ADDRESS=800001e0, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 33: < ADDRESS=800001e4, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 34: < ADDRESS=800001e8, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 35: 36: encoder_input: 37: ******** Data from br_j_asm.encoder_input line 5029 ******** 38: > UNINFERABLE_JUMP, cause=0, tval=0, priv=3, iaddr_0=80000154, context=0, ctype=0, ilastsize_0=2 39: > ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=80000104, context=0, ctype=0, ilastsize_0=4 40: > ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=80000108, context=0, ctype=0, ilastsize_0=4 41: > UNINFERABLE_JUMP, cause=0, tval=0, priv=3, iaddr_0=8000010c, context=0, ctype=0, ilastsize_0=2 42: EXCEPTION, cause=2, tval=0, priv=3, iaddr_0=80000222, context=0, ctype=0, ilastsize_0=2, ----------> NOT RETIRED 43: < ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=800001b0, context=0, ctype=0, ilastsize_0=2 44: < ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=800001e0, context=0, ctype=0, ilastsize_0=4 45: < ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=800001e4, context=0, ctype=0, ilastsize_0=4 46: < ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=800001e8, context=0, ctype=0, ilastsize_0=2 47: 48: te_inst: 49: ******** Data from br_j_asm.te_inst_annotated line 5071 ******** 50: > next=80000154 curr=80000150 prev=8000014c 51: > next=80000104 curr=80000154 prev=80000150 52: > next=80000108 curr=80000104 prev=80000154 53: > format=1, address=80000104, branches=1, branch_map=0, irreport=0, notify=0, updiscon=0, Reason[prev_updiscon] Payload[05 04 01 00 80 00] 54: > next=8000010c curr=80000108 prev=80000104 55: next=80000222 curr=8000010c prev=80000108 56: format=2, address=8000010c, irreport=0, notify=0, updiscon=0, Reason[exc_only] Payload[32 04 00 00 02] 57: < next=800001b0 curr=80000222 prev=8000010c 58: < format=3, subformat=TRAP, address=80000222, branch=1, context=0, ecause=2, interrupt=0, privilege=3, thaddr=0, tval=0, Reason[prev_updiscon, curr_exc_only] Payload[77 00 00 00 00 81 88 00 00 20] 59: < format=3, subformat=START, address=800001b0, branch=1, context=0, privilege=3, Reason[exception_prev, reported] Payload[73 00 00 00 00 6c 00 00 10] 60: < next=800001e4 curr=800001e0 prev=800001b0 61: < next=800001e8 curr=800001e4 prev=800001e0 #### [](#13-1-1-2-packet-data)13.1.1.2\. Packet data The output from the encoder for the fragment of interest is given in line 56\. The least significant byte is output first, this means 32 is byte 0, 04 is byte 1 and and the final value 02 is byte 4. #### [](#13-1-1-3-siemens-transport)13.1.1.3\. Siemens transport The packet is encapsulated according to the [Unformatted Trace & Diagnostic Data Packet Encapsulation for RISC-V Specification](https://github.com/riscv-non-isa/e-trace-encap/releases/latest/), with the following attributes: * Header - 1 byte * SrcID - N bits. As an example use 6 bits and the value of 1. * This example has no timestamp * A 2-bit type field with ’10’ meaning instruction trace * trace\_payload - \[0x32 0x04 0x00 0x00 0x02\] Since the Siemens transport is byte stream based the data seen will be: `[0x06][0x81][0x32 0x04 0x00 0x00 0x02]` #### [](#13-1-1-4-atb-transport)13.1.1.4\. ATB transport The packet is encapsulated according to the [Unformatted Trace & Diagnostic Data Packet Encapsulation for RISC-V Specification](https://github.com/riscv-non-isa/e-trace-encap/releases/latest/), with the following attributes: * Header - 1 byte carried via ATDATA bus. * SrcID - 0 bits as the SrcID is carried via the ATID bus. * This example has no timestamp * No type field (encoder has no data trace support) * trace\_payload - \[0x32 0x04 0x00 0x00 0x02\] carried via ATDATA bus. Assuming a 32 bit ATB transport results in the following ATB transfers `[ATID=1] [ATBYTES = 3] [ATDATA = 0x00043205]` `[ATID=1] [ATBYTES = 1] [ATDATA = 0x00000200]` ### [](#13-1-2-timer-long-loop)13.1.2\. Timer Long Loop #### [](#13-1-2-1-code-fragment)13.1.2.1\. Code fragment 1: ************************************************************************************** 2: ****************** Fragment 0x800001a2 - 0x800001b0:timer_long_loop ****************** 3: ************************************************************************************** 4: KEY: ">" means pre-fragment execution, "<" means post-fragment execution 5: ^^^^^^^^^^^^^^^^^^^^^^^^^^ Part 443 of 445 ^^^^^^^^^^^^^^^^^^^^^^^^^^ 6: 7: elf: 8: > 80000194: fab50ce3 beq a0,a1,8000014c 9: > 80000198: 40430333 sub t1,t1,tp 10: > 8000019c: 34402473 csrr s0,mip 11: > 800001a0: 8c21 xor s0,s0,s0 12: 800001a2: 300024f3 csrr s1,mstatus 13: 800001a6: 8ca5 xor s1,s1,s1 14: 800001a8: fe0310e3 bnez t1,80000188 15: 800001ac: bfb5 j 80000128 16: 800001ae: 0001 nop 17: 00000000800001b0 : 18: 800001b0: a805 j 800001e0 19: < 00000000800001e0 : 20: < 800001e0: 342023f3 csrr t2,mcause 21: < 800001e4: fff0031b addiw t1,zero,-1 22: < 800001e8: 137e slli t1,t1,0x3f 23: < 800001ea: 031d addi t1,t1,7 24: 25: trace_spike: 26: ******** Data from br_j_asm.spike_pc_trace line 5000 ******** 27: > ADDRESS=80000194, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 28: > ADDRESS=80000198, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 29: > ADDRESS=8000019c, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 30: > ADDRESS=800001a0, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 31: ADDRESS=800001a2, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 32: ADDRESS=800001a6, PRIVILEGE=3, EXCEPTION=1, ECAUSE=8000000000000007, TVAL=0, INTERRUPT=1 33: ADDRESS=800001b0, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 34: < ADDRESS=800001e0, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 35: < ADDRESS=800001e4, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 36: < ADDRESS=800001e8, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 37: < ADDRESS=800001ea, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 38: 39: encoder_input: 40: ******** Data from br_j_asm.encoder_input line 5000 ******** 41: > NONTAKEN_BRANCH, cause=0, tval=0, priv=3, iaddr_0=80000194, context=0, ctype=0, ilastsize_0=4 42: > ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=80000198, context=0, ctype=0, ilastsize_0=4 43: > ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=8000019c, context=0, ctype=0, ilastsize_0=4 44: > ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=800001a0, context=0, ctype=0, ilastsize_0=2 45: ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=800001a2, context=0, ctype=0, ilastsize_0=4 46: INTERRUPT, cause=7, tval=0, priv=3, iaddr_0=800001a6, context=0, ctype=0, ilastsize_0=2, ----------> NOT RETIRED 47: ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=800001b0, context=0, ctype=0, ilastsize_0=2 48: < ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=800001e0, context=0, ctype=0, ilastsize_0=4 49: < ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=800001e4, context=0, ctype=0, ilastsize_0=4 50: < ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=800001e8, context=0, ctype=0, ilastsize_0=2 51: < ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=800001ea, context=0, ctype=0, ilastsize_0=2 52: 53: te_inst: 54: ******** Data from br_j_asm.te_inst_annotated line 5038 ******** 55: > next=80000194 curr=80000192 prev=80000190 56: > next=80000198 curr=80000194 prev=80000192 57: > next=8000019c curr=80000198 prev=80000194 58: > next=800001a0 curr=8000019c prev=80000198 59: next=800001a2 curr=800001a0 prev=8000019c 60: next=800001a6 curr=800001a2 prev=800001a0 61: format=1, address=800001a2, branches=15, branch_map=21845, irreport=0, notify=0, updiscon=0, Reason[exc_only] Payload[bd aa aa 68 00 00 20] 62: next=800001b0 curr=800001a6 prev=800001a2 63: < next=800001e0 curr=800001b0 prev=800001a6 64: < format=3, subformat=TRAP, address=800001b0, branch=1, context=0, ecause=7, interrupt=1, privilege=3, thaddr=1, Reason[prev_exception] Payload[77 00 00 00 80 33 6c 00 00 20] 65: < next=800001e4 curr=800001e0 prev=800001b0 66: < next=800001e8 curr=800001e4 prev=800001e0 67: < next=800001ea curr=800001e8 prev=800001e4 #### [](#13-1-2-2-packet-data)13.1.2.2\. Packet data The output from the encoder for the fragment of interest is given in line 61\. The least significant byte is output first, this means 77 is byte 0, 00 is byte 1 and and the final value 20 is byte 9. #### [](#13-1-2-3-siemens-transport)13.1.2.3\. Siemens transport The packet is encapsulated according to the Trace & Diagnostic Data Packet Encapsulation for RISC-V Specification\], with the following attributes: * Header - 1 byte * SrcID - N bits. As an example use 6 bits and the value of A. * This example has no timestamp * A 2-bit type field with ’10’ meaning instruction trace * trace\_payload - \[0xBD 0xAA 0xAA 0x68 0x00 0x00 0x20\] `[0x8][0x8A][0xBD 0xAA 0xAA 0x68 0x00 0x00 0x20]` #### [](#13-1-2-4-atb-transport)13.1.2.4\. ATB transport The packet is encapsulated according to the [Unformatted Trace & Diagnostic Data Packet Encapsulation for RISC-V Specification](https://github.com/riscv-non-isa/e-trace-encap/releases/latest/), with the following attributes: * Header - 1 byte carried via the ATDATA bus. * SrcID - 0 bits as the SrcID is carried via the ATID bus. * This example has no timestamp * No type field (encoder has no data trace support) * trace\_payload - \[0xBD 0xAA 0xAA 0x68 0x00 0x00 0x20\] carried via the ATDATA bus Assuming at 32 bit ATB transport results in the following ATB transfers `[ATID=0xA] [ATBYTES = 3] [ATDATA = 0xAAAABD07]` `[ATID=0xA] [ATBYTES = 3] [ATDATA = 0x20000068]` ### [](#13-1-3-startup-xrle)13.1.3\. Startup xrle #### [](#13-1-3-1-code-fragment)13.1.3.1\. Code fragment 1: *********************************************************************************** 2: ****************** Fragment 0x20010522 - 0x20010528:startup_xrle ****************** 3: *********************************************************************************** 4: KEY: ">" means pre-fragment execution, "<" means post-fragment execution 5: ^^^^^^^^^^^^^^^^^^^^^^^^^^ Part 1 of 1 ^^^^^^^^^^^^^^^^^^^^^^^^^^ 6: 7: elf: 8: 20010522
: 9: 20010522: 1141 addi sp,sp,-16 10: 20010524: c606 sw ra,12(sp) 11: 20010526: c422 sw s0,8(sp) 12: 20010528: 0800 addi s0,sp,16 13: < 2001052a: 800107b7 lui a5,0x80010 14: < 2001052e: 6721 lui a4,0x8 15: < 20010530: e8670713 addi a4,a4,-378 # 7e86 <__heap_size+0x7686> 16: < 20010534: 1ae7aa23 sw a4,436(a5) # 800101b4 <_sp+0xfffffbfc> 17: 18: trace_spike: 19: ******** Data from xrle.spike_pc_trace line 2 ******** 20: ADDRESS=20010522, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 21: ADDRESS=20010524, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 22: ADDRESS=20010526, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 23: ADDRESS=20010528, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 24: < ADDRESS=2001052a, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 25: < ADDRESS=2001052e, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 26: < ADDRESS=20010530, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 27: < ADDRESS=20010534, PRIVILEGE=3, EXCEPTION=0, ECAUSE=0, TVAL=0, INTERRUPT=0 28: 29: encoder_input: 30: ******** Data from xrle.encoder_input line 2 ******** 31: ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=20010522, context=0, ctype=0, ilastsize_0=2 32: ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=20010524, context=0, ctype=0, ilastsize_0=2 33: ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=20010526, context=0, ctype=0, ilastsize_0=2 34: ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=20010528, context=0, ctype=0, ilastsize_0=2 35: < ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=2001052a, context=0, ctype=0, ilastsize_0=4 36: < ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=2001052e, context=0, ctype=0, ilastsize_0=2 37: < ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=20010530, context=0, ctype=0, ilastsize_0=4 38: < ITYPE_NONE, cause=0, tval=0, priv=3, iaddr_0=20010534, context=0, ctype=0, ilastsize_0=4 39: 40: te_inst: 41: ******** Data from xrle.te_inst_annotated line 2 ******** 42: > format=3, subformat=SUPPORT, enable=1, encoder_mode=0, options=4, qual_status=0 Payload[1f 04] 43: next=20010522 44: next=20010524 curr=20010522 45: format=3, subformat=START, address=20010522, branch=1, context=0, privilege=3, Reason[ppccd] Payload[73 00 00 00 00 91 82 00 10] 46: next=20010526 curr=20010524 prev=20010522 47: next=20010528 curr=20010526 prev=20010524 48: < next=2001052a curr=20010528 prev=20010526 49: < next=2001052e curr=2001052a prev=20010528 50: < next=20010530 curr=2001052e prev=2001052a 51: < next=20010534 curr=20010530 prev=2001052e #### [](#13-1-3-2-packet-data)13.1.3.2\. Packet data The output from the encoder for the fragment of interest is given in line 45\. The least significant byte is output first, this means 73 is byte 0, 00 is byte 1 and and the final value 10 is byte 8. #### [](#13-1-3-3-siemens-transport)13.1.3.3\. Siemens transport The packet is encapsulated according to the [Unformatted Trace & Diagnostic Data Packet Encapsulation for RISC-V Specification](https://github.com/riscv-non-isa/e-trace-encap/releases/latest/), with the following attributes: * Header - 1 byte * SrcID - N bits. As an example use 6 bits and the value of 5. * This example has no timestamp * A 2-bit type field with ’10’ meaning instruction trace * trace\_payload - \[0x73 0x00 0x00 0x00 0x00 0x91 0x82 0x00 0x10\] `[0xA][0x85][0x73 0x00 0x00 0x00 0x00 0x91 0x82 0x00 0x10]` #### [](#13-1-3-4-atb-transport)13.1.3.4\. ATB transport The packet is encapsulated according to the [Unformatted Trace & Diagnostic Data Packet Encapsulation for RISC-V Specification](https://github.com/riscv-non-isa/e-trace-encap/releases/latest/), with the following attributes: * Header - 1 byte carried via the ATDATA bus. * SrcID - 0 bits as the SrcID is carried via the ATID bus. * This example has no timestamp * No type field (encoder has no data trace support) * trace\_payload - \[0x73 0x00 0x00 0x00 0x00 0x91 0x82 0x00 0x10\] carried via the ATDATA bus. Assuming at 32 bit ATB transport results in the following ATB transfers `[ATID=0x5] [ATBYTES = 3] [ATDATA = 0x00007309]` `[ATID=0x5] [ATBYTES = 3] [ATDATA=0x82910000]` `[ATID=0x5] [ATBYTES = 1] [ATDATA = 0x00001000]` 14.1. Future Directions ==================== ## [](#futureIdeas)14.1\. Future Directions This chapter captures ideas and enhancements that may be useful for to consider in future versions of the E-Trace specification. ### [](#14-1-1-vector)14.1.1\. Vector Now that the vector extension has been ratified it would be interesting to look at extending E-Trace to support instruction and data trace for vector operations. ### [](#14-1-2-inter-instruction-cycle-counts)14.1.2\. Inter-instruction cycle counts In this mode the encoder will trace where the hart is stalling by reporting the number of cycles between successive instruction retirements. 4.1. Hart to encoder interface ==================== ## [](#Interface)4.1\. Hart to encoder interface ### [](#sec:InstructionInterfaceRequirements)4.1.1\. Instruction Trace Interface requirements This section describes in general terms the information which must be passed from the RISC-V hart to the trace encoder for the purposes of Instruction Trace, and distinguishes between what is mandatory, and what is optional. The following information is mandatory: * The number of instructions that are being retired; * Whether there has been an exception or interrupt, and if so the cause (from the **_scause/vscause/mcause_** etc. CSR) and trap value (from the**_stval/vstval/mtval_** etc. CSR). The register set to output should be the set that is updated as a result of the exception (i.e. the set associated with the privilege level immediately following the exception); * The current privilege level of the RISC-V hart; * The _instruction\_type_ of retired instructions for: * Jumps with a target that cannot be inferred from the source code; * Taken and nontaken branches; * Return from exception or interrupt (**_\*ret_** instructions). * The _instruction\_address_ for: * Jumps with a target that _cannot_ be inferred from the source code; * The instruction retired immediately after a jump with a target that_cannot_ be inferred from the source code (also referred to as the target or destination of the jump); * Taken and nontaken branches; * The last instruction retired before an exception or interrupt; * The first instruction retired following an exception or interrupt; * The last instruction retired before a privilege change; * The first instruction retired following a privilege change. The following information is optional: * Context or Time information: * The context and/or hart ID and/or time; * The type of action to take when context or time data changes. * The _instruction\_type_ of instructions for: * Calls with a target that _cannot_ be inferred from the source code; * Calls with a target that _can_ be inferred from the source code; * Other jumps without linkage with a target that _cannot_ be inferred from the source code; * Other jumps without linkage with a target that _can_ be inferred from the source code; * Returns with a target that _cannot_ be inferred from the source code; * Returns with a target that _can_ be inferred from the source code; * Co-routine swap; * Other jumps with linkage which don’t fit any of the above classifications with a target that _cannot_ be inferred from the source code; * Other jumps with linkage which don’t fit any of the above classifications with a target that _can_ be inferred from the source code. * If context or time is supported then the _instruction\_address_ for: * The last instruction retired before a context or a time change; * The first instruction retired following a context or time change. * Whether jump targets are sequentially inferable or not. The mandatory information is the bare-minimum required to implement the branch trace algorithm outlined in [\[Algorithm\]](#Algorithm). The optional information facilitates alternative or improved trace algorithms: * Implicit return mode (see[\[sec:implicit-return\]](#sec:implicit-return)) requires the encoder to keep track of the number of nested function calls, and to do this it must be aware of all calls and returns regardless of whether the target can be inferred or not; * A simpler algorithm useful for basic code profiling would only report function calls and returns, again regardless of whether the target can be inferred or not; * Branch prediction techniques can be used to further improve the encoder efficiency, particularly for loops (see[\[sec:branch-prediction\]](#sec:branch-prediction)). This requires the encoder to be aware of the address of all branches, whether they are taken or not. * Uninferable jumps can be treated as inferable (which don’t need to be reported in the trace output) if both the jump and the preceding instruction which loads the target into a register have been traced. #### [](#JumpClasses)4.1.1.1\. Jump classification and target inference Jumps are classified as _inferable_, or _uninferable_. An _inferable_jump has a target which can be deduced from the binary executable or representation thereof (e.g. ELF). For the purposes of this specification, the following strict definition applies: If the target of a jump is supplied via a constant embedded within the jump opcode, it is classified as _inferable_. Jumps which are not_inferable_ are by definition _uninferable_. However, there are some jump targets which can still be deduced from the binary executable by considering pairs of instructions even though by the above definition they are classified as uninferable. Specifically, when the source register for the jump instruction is supplied via * an **_lui_** or **_c.lui_** (a register which contains a constant), or * an **_auipc_** (a register which contains a constant offset from the PC). Such jump targets are classified as _sequentially inferable_ if the pair of instructions are retired consecutively (i.e. the **_auipc_**, **_lui_**or **_c.lui_** immediately precedes the jump). Note: the restriction that the instructions are retired consecutively is necessary in order to minimize the additional signalling needed between the hart and the encoder, and should have a minimal impact on trace efficiency as it is anticipated that consecutive execution will be the norm. Support for sequentially inferable jumps is optional. Jumps may optionally be further classified according to the recommended calling convention: * _Calls_: * **_jal_** x1; * **_jal_** x5; * **_jalr_** x1, rs where rs != x5; * **_jalr_** x5, rs where rs != x1; * **_c.jalr_** rs1 where rs1 != x5; * **_c.jal_**. * _Other jumps without linkage_: * **_jal_** x0; * **_c.j_**; * **_jalr_** x0, rs where rs != x1 and rs != x5; * **_c.jr_** rs1 where rs1 != x1 and rs1 != x5. * _Returns_: * **_jalr_** rd, rs where (rs == x1 or rs == x5) and rd != x1 and rd != x5; * **_c.jr_** rs1 where rs1 == x1 or rs1 == x5. * _Co-routine swap_: * **_jalr_** x1, x5; * **_jalr_** x5, x1; * **_c.jalr_** x5. * _Other jumps with linkage_: * **_jal_** rd where rd != x0 and rd != x1 and rd != x5; * **_jalr_** rd, rs where rs != x1 and rs != x5 and rd != x0 and rd != x1 and rd != x5. #### [](#sec:relationship)4.1.1.2\. Relationship between RISC-V core and the encoder The encoder is intended to encode the instructions executed on a single hart. It is however commonplace for a RISC-V core to contain multiple harts. This can be supported by the core in several different ways: * Implement a separate instance of the interface per hart. Each instance can be connected to a separate encoder instance, allowing all harts to be traced concurrently. Alternatively, external muxing may be used in conjunction with a single encoder in order to trace one particular hart at a time; * Implement a single interface for the core, with muxing inside the core to select which hart to connect to the interface. (Whilst it is technically feasible to use a single encoder with multiple harts operating in a fine-grained multi-threaded configuration, the frequent context changes that would occur as a result of thread-switching would result in extremely poor encoding efficiency, and so this configuration is not recommended.) ### [](#sec:InstructionTraceInterface)4.1.2\. Instruction Trace Interface This section describes the interface between a RISC-V hart and the trace encoder that conveys the information described in the section[4.1.1\. Instruction Trace Interface requirements](#sec:InstructionInterfaceRequirements). Signals are assigned to one of the following groups: * M: Mandatory. The interface must include an instance of this signal. * O: Optional. The interface may include an instance of this signal. * MR: Mandatory, may be replicated. For harts that can retire a maximum of N "special" instructions per clock cycle, the interface must include N instances of this signal. * OR: Optional, may be replicated. For harts that can retire a maximum of N "special" instructions per clock cycle, the interface must include zero or N instances of this signal. * BR: Block, may be replicated. Mandatory for harts that can retire multiple instructions in a block. Replication as per OR. If omitted, the interface must include SR group signals instead. * SR: Single, may be replicated. Mandatory for harts that can only retire one instruction in a block. Replication as per OR (see[4.1.2.2\. Alternative multiple-retirement interface configurations](#sec:alt-multi)). If omitted, the interface must include BR group signals instead. "Special" instructions are those that require **itype** to be non-zero. __Table 1\. Instruction interface signals__ | **Signal** | **Group** | **Function** | | --------------------------------------- | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **itype**\[_itype\_width\_p_\-1:0\] | MR | Termination type of the instruction block. Encoding given in [Table 4](#tab:itype) (see [4.1.1.1\. Jump classification and target inference](#JumpClasses) for definitions of codes 6 - 15). | | **cause**\[_ecause\_width\_p_\-1:0\] | M | Exception or interrupt cause (**_scause/ vscause/mcause_**). Ignored unless **itype**\=1 or 2. | | **tval**\[_iaddress\_width\_p_\-1:0\] | M | The associated trap value, e.g. the faulting virtual address for address exceptions, as would be written to the **stval/vstval/mtval** CSR. Future optional extensions may define **tval** to provide ancillary information in cases where it currently supplies zero. Ignored unless **itype**\=1. | | **priv**\[_privilege\_width\_p_\-1:0\] | M | Privilege level for all instructions retired on this cycle. Encoding given in[Table 5](#tab:priv). Codes 4-7 optional. | | **iaddr**\[_iaddress\_width\_p_\-1:0\] | MR | The address of the 1st instruction retired in this block. Invalid if **iretire**\=0 unless **itype**\=1, in which case it indicates the address of the instruction which incurred the exception. | | **context**\[_context\_width\_p_\-1:0\] | O | Context for all instructions retired on this cycle. | | **time**\[_time\_width\_p_\-1:0\] | O | Time generated by the core. | | **ctype**\[_ctype\_width\_p_\-1:0\] | O | Reporting behavior for **context**. Encoding given in Table [Table 6](#tab:context-type). Codes 2-3 optional. | | **sijump** | OR | If **itype** indicates that this block ends with an uninferable discontinuity, setting this signal to 1 indicates that it is sequentially inferable and may be treated as inferable by the encoder if the preceding **_auipc_**, **_lui_** or **_c.lui_** has been traced. Ignored for **itype** codes other than 6, 8, 10, 12 or 14. | [Table 1](#tab:common-ingress) and [Table 2](#tab:multi-ingress) list the signals in the interface designed to efficiently support retirement of multiple instructions per cycle. The following discussion describes the multiple-retirement behavior. However, for harts that can only retire one instruction at a time, the signalling can be simplified, and this is discussed subsequently in [4.1.2.1\. Simplifications for single-retirement](#sec:single-retire). __Table 2\. Instruction interface signals - multiple retirement per block__ | **Signal** | **Group** | **Function** | | ------------------------------------------- | --------- | ---------------------------------------------------------------------- | | **iretire**\[_iretire\_width\_p_\-1:0\] | BR | Number of halfwords represented by instructions retired in this block. | | **ilastsize**\[_ilastsize\_width\_p_\-1:0\] | BR | The size of the last retired instruction is 2**ilastsize** half-words. | __Table 3\. Instruction interface signals - single retirement per block__ | **Signal** | **Group** | **Function** | | ------------------------------------------ | --------- | ------------------------------------------------------------------------------------ | | **iretire**\[0:0\] | SR | Number of instructions retired in this block (0 or 1). | | **ilastsize**\[_ilastsize\_width\_p-1_:0\] | SR | The size of the last retired instruction in this block is 2**ilastsize** half-words. | __Table 4\. Instruction Type (**itype**) encoding__ | **Value** | **Description** | | --------- | ----------------------------------------------------------------------------------------------- | | 0 | Last instruction in the block is none of the other named **itype**codes | | 1 | Exception. An exception that traps occurred following the last retired instruction in the block | | 2 | Interrupt. An interrupt that traps occurred following the last retired instruction in the block | | 3 | Exception or interrupt return | | 4 | Nontaken branch | | 5 | Taken branch | | 6 | Uninferable jump if _itype\_width\_p_ is 3, reserved otherwise | | 7 | reserved | | 8 | Uninferable call | | 9 | Inferrable call | | 10 | Uninferable jump | | 11 | Inferrable jump | | 12 | Co-routine swap | | 13 | Return | | 14 | Other uninferable jump | | 15 | Other inferable jump | __Table 5\. Privilege level (**priv**) encoding__ | **Value** | **Description** | | --------- | --------------- | | 0 | U | | 1 | S/HS | | 2 | reserved | | 3 | M | | 4 | D (debug mode) | | 5 | VU | | 6 | VS | | 7 | reserved | The information presented in a block represents a contiguous sequence of instructions starting at **iaddr**, all of which retired in the same cycle. Note if **itype** is 1 or 2 (indicating an exception or an interrupt), the number of instructions retired may be zero. **cause** and**tval** are only defined if **itype** is 1 or 2\. If **iretire**\=0 and**itype**\=0, the values of all other signals are undefined. **iretire** contains the number of (16-bit) half-words represented by instructions retired in this block, and **ilastsize** the size of the last instruction. Half-words rather than instruction count enables the encoder to easily compute the address of the last instruction in the block without having access to the size of every instruction in the block. **itype** can be 3 or 4 bits wide. If _itype\_width\_p_ is 3, a single code (6) is used to indicate all uninferable jumps. This is simpler to implement, but precludes use of the implicit return mode (see[\[sec:implicit-return\]](#sec:implicit-return)), which requires jump types to be fully classified. Note that when _itype\_width\_p_ is 3, **itype** \= 0 is used for inferrable calls. However, inferrable calls must still be the final instruction retired in a block, otherwise the block would not be comprised of contiguous instructions. Whilst **iaddr** is typically a virtual address, it does not affect the encoder’s behavior if it is a physical address. For harts that can retire a maximum of N non-zero **itype** values per clock cycle, the signal groups MR, OR and either BR or SR must be replicated N times. Typically N is determined by the maximum number of branches that can be retired per clock cycle. Signal group 0 represents information about the oldest instruction block, and group N-1 represents the newest instruction block. The interface supports no more than one privilege change, context change, exception or interrupt per cycle and so signals in groups M and O are not replicated. Furthermore, **itype**can only take the value 1 or 2 in one of the signal groups, and this must be the newest valid group (i.e. **iretire** and **itype** must be zero for higher numbered groups). If fewer than N groups are required in a cycle, then lower numbered groups must be used first. For example, if there is one branch, use only group 0, if there are two branches, instructions up to the 1st branch must be reported in group 0 and instructions up to the 2nd branch must be reported in group 1 and so on. **sijump** is optional and may be omitted if the hart does not implement the logic to detect sequentially inferable jumps. If the encoder offers an **sijump** input it must also provide a parameter to indicate whether the input is connected to a hart that implements this capability, or tied off. This is to ensure the decoder can be made aware of the hart’s capability. Enabling sequentially inferable jump mode in the encoder and decoder when the hart does not support it will prevent correct reconstruction by the decoder. The **context** and/or the **time** field can be used to convey any additional information to the decoder. For example: * The address space and virtual machine IDs (**ASID** and **VMID**respectively). Where present it is recommended these values be wired to bits \[15:0\] and \[29:16\]; * The software thread ID; * The process ID from an operating system; * It could be used to convey the values of CSRs to the decoder by setting **context** to the CSR number and value when a CSR is written; * In cases where a single encoder is being shared amongst multiple harts (see [4.1.1.2\. Relationship between RISC-V core and the encoder](#sec:relationship)), it could also be used to indicate the hart ID, in cases where the hart ID can be changed dynamically. * Time from within the hart [Table 6](#tab:context-type) specifies the actions for the various **ctype** values. A typical behavior would be for this signal to remain zero except on the 1st retirement after a context change or when a time value should be reported. _ctype\_width\_p_ may be 1 or 2\. The reduced width option only provides support for reporting context changes imprecisely. __Table 6\. Context type **ctype** values and corresponding actions__ | **Type** | **Value** | **Actions** | | ----------------------------------------------- | --------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Unreported | 0 | No action (don’t report context). | | Report context imprecisely | 1 | An example would be a SW thread or operating system process change. Report the new context value at the earliest convenient opportunity. It is reported without any address information, and the assumption is that the precise point of context change can be deduced from the source code (e.g. a CSR write). | | Report context precisely | 2 | Report the address of the 1st instruction retired in this block, and the new context. If there were unreported branches beforehand, these need to be reported first. Treated the same as a privilege change. | | Report context as an asynchronous discontinuity | 3 | An example would be a change of hart. Need to report the last instruction retired on the previous context, as well as the 1st on the new context. Treated the same as an exception. | #### [](#sec:single-retire)4.1.2.1\. Simplifications for single-retirement For harts that can only retire one instruction at a time, the interface can be simplified to the signals listed in[Table 1](#tab:common-ingress) and [Table 3](#tab:single-ingress). The simplifications can be summarized as follows: * **iretire** simply indicates whether an instruction retired or not; **Note:** **ilastsize** is still needed in order to determine the address of the next instruction, as this is the predicted return address for implicit return mode (see [\[sec:implicit-return\]](#sec:implicit-return)). The parameter _retires\_p_ which indicates to the encoder the maximum number of instructions that can be retired per cycle can be used by an encoder capable of supporting single or multiple retirement to select the appropriate interpretation of **iretire**. #### [](#sec:alt-multi)4.1.2.2\. Alternative multiple-retirement interface configurations For a hart that can retire multiple instructions per cycle, but no more than one branch, the preferred solution is to use one instance of signals from groups BR, MR and OR. However, if the hart can retire N branches in a cycle, N instances of signals from groups MR, OR and either SR or BR must be used (each instance can be either a single instruction or a block). If the hart can retire N instructions per cycle, but only one branch, it is allowed (though not recommended) to provide explicit details of every instruction retired by using N instances of signals from groups SR, MR and OR. #### [](#4-1-2-3-optional-sideband-signals)4.1.2.3\. Optional sideband signals Optional sideband signals may be included to provide additional functionality, as described in[Table 7](#tab:ingress-side-band) and[Table 8](#tab:egress-side-band). Note, any user defined information that needs to be output by the encoder will need to be applied via the **context** input. __Table 7\. Optional sideband encoder input signals__ | **Signal** | **Group** | **Function** | | ------------------------------------- | -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **impdef**\[_impdef\_width\_p_\-1:0\] | O | Implementation defined sideband signals. A typical use for these would be for filtering (see[\[ch:filtering\]](#ch:filtering). | | **trigger**\[2+:0\] | \[1:0\]: O\[2+\]: OR | A pulse on bit 0 will cause the encoder to start tracing, and continue until further notice, subject to other filtering criteria also being met. A pulse on bit 1 will cause the encoder to stop tracing until further notice. See [4.1.2.4\. Using trigger outputs from the Debug Module](#sec:trigger)). | | **halted** | O | Hart is halted. Upon assertion, the encoder will output a packet to report the address of the last instruction retired before halting, followed by a support packet to indicate that tracing has stopped. Upon deassertion, the encoder will start tracing again, commencing with a synchronization packet. **Note:** If this signal is not provided, it is strongly recommended that Debug mode can be signalled via a 3-bit **privilege** signal. This will allow tracing in Debug mode to be controlled via the optional filtering capabilities. | | **reset** | O | Hart is in reset. Provided the encoder is in a different reset domain to the hart, this allows the encoder to indicate that tracing has ended on entry to reset, and restarted on exit. Behavior is as described above for halt. | __Table 8\. Optional sideband encoder output signals__ | **Signal** | **Group** | **Function** | | ---------- | --------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **stall** | O | Stall request to hart. Some applications may require lossless trace, which can be achieved by using this signal to stall the hart if the trace encoder is unable to output a trace packet (for example due to back-pressure from the packet transport infrastructure). | #### [](#sec:trigger)4.1.2.4\. Using trigger outputs from the Debug Module The debug module of the RISC-V hart may have a trigger unit. This defines a match control register (**_mcontrol_**) containing a 4-bit**action** field, and reserves codes 2 - 5 of this field for trace use. These action codes are hereby defined as shown in table[Table 9](#tab:debugModuleTriggerSupport). If implemented, each action must generate a pulse on an output from the hart, on the same cycle as the instruction which caused the trigger is retired. __Table 9\. Debug Module trigger support (**_mcontrol_ action**)__ | **Value** | **Description** | | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 2 | _Trace-on_. This should be connected to **trigger\[0\]** if the encoder provides it. | | 3 | _Trace-off_. This should be connected to **trigger\[1\]** if the encoder provides it. | | 4 | _Trace-notify_. This should be connected to **trigger\[1 + _blocks_:**2\] if the encoder provides it. This will cause the encoder to output a packet containing the address of the last instruction in the block if it is enabled. One bit per block. | Trace-on and Trace-off actions provide a means for the hart to control when tracing starts and stops. It is recommended that tracing starts from the oldest instruction retired in the cycle that Trace-on is asserted, and stops following the newest instruction retired in the cycle that Trace-off is asserted (subject to any optional filtering). It follows from this that: * if tracing is enabled and trace-off occurs on the cycle before trace-on, then tracing will continue unimpeded (i.e. it stays on); * if tracing is disabled and trace-on and trace-off triggers occur simultaneously then only the instructions retired in that cycle will be traced. Trace-notify provides means to ensure that a specified instruction is explicitly reported (subject to any optional filtering). This capability is sometimes known as a watchpoint. #### [](#4-1-2-5-example-retirement-sequences)4.1.2.5\. Example retirement sequences __Table 10\. Example 1 : 9 Instructions retired over four cycles, 2 branches__ | **Retired** | **Instruction Trace Block** | | ---------------------------------------------------------------- | ----------------------------------------------- | | 1000: **_divuw_**1004: **_add_**1008: **_or_**100C: **_c.jalr_** | **iretire**\=7, **iaddr**\=0x1000, **itype**\=8 | | 0940: **_addi_**0944: **_c.beq_** | **iretire**\=3, **iaddr**\=0x0940, **itype**\=4 | | 0946: **_c.bnez_** | **iretire**\=1, **iaddr**\=0x0946, **itype**\=5 | | 0988: **_lbu_**098C: **_csrrw_** | **iretire**\=4, **iaddr**\=0x0988, **itype**\=0 | ### [](#sec:DataInterfaceRequirements)4.1.3\. Data Trace Interface requirements This section describes in general terms the information which must be passed from the RISC-V hart to the trace encoder for the purposes of Data Trace, and distinguishes between what is mandatory, and what is optional. If Data Trace is not needed in a system then there is no requirement for the RISC-V hart to supply any of the signals in [4.1.4\. Data Trace Interface](#sec:DataTraceInterface). Data trace supports up to four data access types: load, store, atomic and CSR. Support for both atomic and CSR accesses are independently optional. The signalling protocol can take one of two forms, depending on the needs of the RISC-V hart: _unified_ or _split_. Unified is the simplest form, suitable for simpler, in-order harts. In this form, all information about a data access is signalled by the RISC-V hart in the same cycle that the associated data access instruction is reported on the instruction trace interface. For harts with out of order or speculative execution capabilities, many loads may be in progress simultaneously, and this approach is not practical as it would require the hart to maintain a large amount of state relating to all the in-progress loads. For this reason, the interface also supports splitting loads into two parts: * The _request_ phase provides all the information about the load that originates from the hart (address, size, etc.) when the instruction retires; * The _response_ phase provides the data and response status when it has been returned to the hart from the memory system. The two parts of a split load are associated by use of a transaction ID. The Zc (code-size reduction) extension introduced push and pop instructions (_cm.push_, _cm.pop_, _cm.popret_ and _cm.popretz_) that each result in multiple loads or stores. To allow the resulting loads or stores to be associated with the correct instruction, these multi-memory-access instructions (and any other future instructions with similar characteristics) must be reported on the instruction trace interface multiple times (once for each individual load or store) using**itype** 0 except for the final load or store, which must retire using the natural **itype** for the instruction (for example, a _cm.popret_instruction must use **itype** 13 for the final load to signal the return). The instruction address reported will be the same for each occurrence. The following illustrations show the retirement sequences when a single_cm.push_ or _cm.popret_ is used to push or pop 4 registers from the stack. They assume a RISC-V to encoder interface that can report a block of 1 or more retired instructions and one load or store per cycle. Each comprises 4 elements, and shows the instruction information reported for each load and store. As detailed in section[4.1.2\. Instruction Trace Interface](#sec:InstructionTraceInterface), this takes the form of the address of an instruction, the length of the block (1 for a single instruction) and the type of the last instruction in the block. In each element, ’Block’ indicates a block of 1 or more instructions (i.e. could also be a single instruction), whereas ’Single’ indicates a single instruction (i.e. a block with a length of 1). A _cm.push_ is equivalent to 4 store instructions: 1. Block - last instruction is _cm.push_, **itype** 0 (data trace interface reports 1st store); 2. Single - _cm.push_, **itype** 0 (data trace interface reports 2nd store); 3. Single - _cm.push_, **itype** 0 (data trace interface reports 3rd store); 4. Block - 1st instruction is _cm.push_, **itype** dependent on last instruction in block (data trace interface reports 4th store); A _cm.popret_ is equivalent to 4 loads and a return: 1. Block - last instruction is _cm.popret_, **itype** 0 (data trace interface reports 1st load); 2. Single - _cm.popret_, **itype** 0 (data trace interface reports 2nd load); 3. Single - _cm.popret_, **itype** 0 (data trace interface reports 3rd load); 4. Single - _cm.popret_, **itype** 13 (data trace interface reports 4th load); If an exception occurs part way through the sequence of loads or stores initiated by such an instruction, and the instruction is re-executed after the exception handler has been serviced, the load or store sequence must recommence from the beginning. | | This is required for data trace only. If data trace is not implemented, the push or pop may instead be reported just once in the normal way when all associated loads or stores complete successfully. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#sec:DataTraceInterface)4.1.4\. Data Trace Interface This section describes the interface between a RISC-V hart and the trace encoder that conveys the information described in the[4.1.3\. Data Trace Interface requirements](#sec:DataInterfaceRequirements). Signals are assigned to one of the following groups: * M: Mandatory. The interface must include an instance of this signal; * U: Unified. Mandatory for unified signalling; * S: Split. Mandatory for split load signalling; * O: Optional. The interface may include an instance of this signal. All signals in M, U and O groups are only valid when **dretire** is high. Signals in the S group are valid as indicated in table[Table 11](#tab:data-ingress). For harts that can retire a maximum of M data accesses per cycle, the implemented signal groups must be replicated M times. If fewer than M groups are required in a cycle, then lower numbered groups must be used first. For example, if there is one data access, use only group 0. __Table 11\. Data interface signals__ | **Signal** | **Group** | **Function** | | ----------------------------------------------- | --------- | ----------------------------------------------------------------------------------------------------------------------------- | | **dretire** | M | Data access retired (when high) | | **dtype**\[_dtype\_width\_p_\-1:0\] | M | Data access type. Encoding given in[Table 12](#tab:dtype) | | **daddr**\[_daddress\_width\_p_\-1:0\] | M | The data access address | | **dsize**\[_dsize\_width\_p_\-1:0\] | M | The data access size is 2**dsize**bytes | | **data**\[_data\_width\_p_\-1:0\] | U | The data | | **iaddr\_lsbs**\[_iaddr\_lsbs\_width\_p_\-1:0\] | O | LSBs of the data access instruction address. Required if _retires\_p_ \> 1 | | **dblock**\[_dblock\_width\_p_\-1:0\] | O | Instruction block in which the data access instruction is retired. Required if there are replicated instruction block signals | | **lrid**\[_lrid\_width\_p_\-1:0\] | S | Load request ID. Valid when **dretire**is high | | **lresp**\[_lresp\_width\_p_\-1:0\] | S | Load response:0: None1: reserved2: Okay. Load successful; **ldata** valid3: Error. Load failed; **ldata** not valid | | **lid**\[_lrid\_width\_p_\-1:0\] | S | Split Load ID. Valid when **lresp** is non-zero | | **sdata**\[_sdata\_width\_p_\-1:0\] | S | Store data. Valid when **dretire** is high and access is a store (**dtype** is 1) or atomic (**dtype** is 8 - 14). | | **ldata**\[_ldata\_width\_p_\-1:0\] | S | Load data. Valid when **lresp** is non-zero | __Table 12\. Data access type (**dtype**) encoding__ | **Value** | **Description** | | --------- | ------------------------- | | 0 | Load | | 1 | Store | | 2 | reserved | | 3 | reserved | | 4 | CSR read-write | | 5 | CSR read-set | | 6 | CSR read-clear | | 7 | reserved | | 8 | Atomic swap | | 9 | Atomic add | | 10 | Atomic AND | | 11 | Atomic OR | | 12 | Atomic XOR | | 13 | Atomic max | | 14 | Atomic min | | 15 | Conditional store failure | The maximum value of _dtype\_width\_p_ is 4\. However, if only loads and stores are supported, _dtype\_width\_p_ can be 1\. If CSRs are supported but atomics are not, _dtype\_width\_p_ can be 3. Atomic and CSR accesses have either both load and store data, or store data and an operand. For CSRs and unified atomics, both values are reported via **data**, with the store data in the LSBs and the load data or operand in the MSBs. _lrid\_width\_p_ is determined by the maximum number of loads that can be in progress simultaneously, such that at any one time there can be no more than one load in progress with a given ID. **iaddr\_lsbs** and **dblock** are provided to support filtering of which data accesses to trace based on their instruction address. This is best illustrated by considering the following instruction sequence: 1. load 2. 3. load 4. 5. Suppose the hart is capable of retiring up to 4 instructions in a cycle, via a single block. Instruction trace is enabled throughout, but the requirement is to collect data trace for the 1st load (instruction 1), and filtering is configured to match the address of this instruction only. However, information about instruction addresses is passed to the encoder at the block level, and the block boundaries are invisible to the decoder. For instruction trace, all instructions in a block are traced if any of the instructions in that block match the filtering criteria. That is fine for instruction trace - the address of the 1st and final traced instruction are output explicitly. There will be some fuzziness about precisely what those addresses will be depending on where the block boundaries fall, but this is not a concern as everything is always self-consistent. However, that is not the case for data trace. Consider two scenarios: * Case 1: 1st block contains instructions 1, 2, 3; second block contains 4, 5 * Case 2: 1st block contains instructions 1, 2; second block contains 3, 4, 5 Given that **iretire** is non-zero in the same cycle that the data access retires, the encoder knows the address of the 1st and last instructions in a block, but does not know precisely where in the block the data access is. In both cases, the first block matches the filtering criteria (it contains the address of instruction 1), and the second block does not. But if the encoder traced all the data accesses in the matching block, then in case 1 it would trace both instructions 1 and 3, whereas in the second case it would trace only instruction 1\. The decoder has no visibility of the block boundaries so cannot account for this. It is expecting only instruction 1 to be traced, and so may misinterpret instruction 3\. If this code is in a loop for example, it will assume that the 2nd traced load is in fact instruction 1 from the next loop iteration, rather than instruction 3 from this iteration. Providing the LSBs of the data access instruction address allows the decoder to determine precisely whether the data access should be traced or not, and removes the dependency on the block sizes and boundaries. The number of bits required is one more bit than the number required to index within the block because blocks can start on any half-word boundary. For harts that replicate the block signals to allow multiple blocks to retire per cycle it is also necessary to indicate which block each data access is associated with, so the encoder knows which block address to combine with the LSBs in order to construct the actual data access instruction address. 1 bit for 2 blocks per cycle, 2 bits for 4, and so on. 1.1. Introduction ==================== ## [](#sec:intro)1.1\. Introduction In complex systems understanding program behavior is not easy. Unsurprisingly in such systems, software sometimes does not behave as expected. This may be due to a number of factors, for example, interactions with other cores, software, peripherals, realtime events, poor implementations or some combination of all of the above. It is not always possible to use a debugger to observe behavior of a running system as this is intrusive. Providing visibility of program execution is important. This needs to be done without swamping the system with vast amounts of data. One method of achieving this is via a Processor Branch Trace. This works by tracking execution from a known start address and sending messages about the address deltas taken by the program. These deltas are typically introduced by jump, call, return and branch type instructions, although interrupts and exceptions are also types of deltas. Conceptually, the system has one or more of the following fundamental components: * A core with an instruction trace interface that outputs all relevant information to allow the successful creation of a processor branch trace and more. This is a high bandwidth interface: in most implementations, it will supply a large amount of data (instruction address, instruction type, context information, …​) for each core execution clock cycle; * A hardware encoder that connects to this instruction trace interface and compresses the information into lower bandwidth trace packets; * A transmission channel to transmit or a memory to store these trace packets; * A decoder, usually software on an external PC, that takes in the trace packets and, with knowledge of the program binary that’s running on the originating hart, reconstructs the program flow. This decoding step can be done off-line or in real-time while the hart is executing. In RISC-V, all instructions are executed unconditionally or at least their execution can be determined based on the program binary. The instructions between the deltas can all be assumed to be executed sequentially. Because of this, there is no need to report sequential instructions in the trace, only whether the branches were taken or not and the address of taken indirect branches or jumps. If the program counter is changed by an amount that cannot be determined from the execution binary, the trace decoder needs to be given the destination address (i.e. the address of the next valid instruction). Examples of this are indirect branches or jumps, where the next instruction address is determined by the contents of a register rather than a constant embedded in the program binary. Interrupts generally occur asynchronously to the program’s execution rather than intentionally as a result of a specific instruction or event. Exceptions can be thought of in the same way, even though they can be typically linked back to a specific instruction address. The decoder generally does not know where an interrupt occurs in the instruction sequence, so the trace encoder must report the address where normal program flow ceased, as well as give an indication of the asynchronous destination which may be as simple as reporting the exception type. When an interrupt or exception occurs, or the processor is halted, the final instruction retired beforehand must be included in the trace. This document serves to specify the ingress port (the signals between the RISC-V core and the encoder), compressed branch trace algorithm and the packet format used to encapsulate the compressed branch trace information. ### [](#sec:terminology)1.1.1\. Terminology The following terms have a specific meaning in this specification. * **ATB**: Arm trace bus * **branch**: an instruction which conditionally changes the execution flow * **CSR**: control/status register * **decoder**: a piece of software that takes the trace packets emitted by the encoder and reconstructs the execution flow of the code executed by the RISC-V hart * **delta**: a change in the program counter that is other than the difference between two instructions placed consecutively in memory * **discontinuity**: another name for ’delta’ (see above) * **E-Trace**: Abbreviation of _Efficient Trace for RISC-V_. * **ELF**: executable and linkable format * **encoder**: a piece of hardware that takes in instruction execution information from a RISC-V hart and transforms it into trace packets * **EPC**: exception program counter * **exception**: an unusual condition occurring at run time associated with an instruction in a RISC-V hart * **hart**: a RISC-V hardware thread * **final instruction**: the instruction retired immediately before tracing stops. * **last instruction**: in a block of sequentially retired instructions, the instruction at the end of the sequence * **interrupt**: an external asynchronous event that may cause a RISC-V hart to experience an unexpected transfer of control * **ISA**: instruction set architecture * **jump**: an instruction which unconditionally changes the execution flow * **direct jump**: an instruction which unconditionally changes the execution flow by changing the PC by a constant value * **indirect jump**: an instruction which unconditionally changes the execution flow by changing the PC to a computed value * **inferable jump**: a jump where the target address is supplied via a constant embedded within the jump opcode * **uninferable jump**: a jump which is not inferable (see above) * **LSB**: least significant bit * **MSB**: most significant bit * **packet**: the atomic unit of encoded trace information emitted by the encoder * **PC**: program counter * **program counter**: a register containing the address of the instruction being executed * **retire**: the final stage of executing an instruction, when the machine state is updated (sometimes referred to as ’commit’ or ’graduate’) * **trap**: the transfer of control to a trap handler caused by either an exception or an interrupt * **updiscon**: contraction of ’uninferable PC discontinuity’ ### [](#1-1-2-nomenclature)1.1.2\. Nomenclature In the following sections items in **bold** are signals or fields within a packet. Items in **_bold italics_** are mnemonics for instructions or CSRs defined in the RISC-V ISA Items in _italics_ with names ending _’\_p’_ refer to parameters either built into the hardware or configurable hardware values. 7.1. Instruction Trace Encoder Output Packets ==================== ## [](#packets)7.1\. Instruction Trace Encoder Output Packets The bulk of this section describes the payload of packets output from the Instruction Trace Encoder. The infrastructure used to transport these packets is outside the scope of this document, and as such the manner in which packets are encapsulated for transport is not mandated, but the [Unformatted Trace & Diagnostic Data Packet Encapsulation for RISC-V Specification](https://github.com/riscv-non-isa/e-trace-encap/releases/latest/) is recommended. Chapter 3 of that document specifically addresses the encapsulation requirements for E-Trace. Any encapsulation must include the following information: * The packet type; * The packet length, in bytes; * The packet payload. Two example transport schemes are the Siemens Messaging Infrastructure, and the Arm AMBA Trace Bus. The Unformatted Trace & Diagnostic Data Packet Encapsulation for RISC-V Specification is well suited to either of these. The remainder of this section describes the contents of the payload portion which should be independent of the infrastructure. In each table, the fields are listed in transmission order: first field in the table is transmitted first, and multi-bit fields are transmitted LSB first. This packet payload format is used to output encoded instruction trace. Three different formats are used according to the needs of the encoding algorithm. The following tables show the format of the payload - i.e. excluding any encapsulation. In order to achieve best performance, actual packet lengths may be adjusted using 'sign based compression'. At the very minimum this should be applied to the address field of format 1 and 2 packets, but ideally will be applied to the whole packet, regardless of format. This technique eliminates identical bits from the most significant end of the packet, and adjusts the length of the packet accordingly. A decoder receiving this shortened packet can reconstruct the original full-length packet by sign-extending from the most significant received bit. Where the payload length given in the following tables, or after applying sign-based compression, is not a multiple of whole bytes in length, the payload must be sign-extended to the nearest byte boundary. Whilst offering maximum encoding efficiency, variable length packets can present some challenges, specifically in terms of identifying where the boundaries between packets occur either when packed packets are written to memory, or when packets are streamed offchip via a communications channel. Two potential solutions to this are as follows: * If the maximum packet payload length is 2N\-1 (for example, if N is 5, then the maximum length is 31 bytes), and the minimum packet payload length is 1, then a sequence of at least 2N zero bytes cannot occur within a packet payload, and therefore the first non-zero byte seen after a sequence of at least 2N zero bytes must be the first byte of a packet. This approach can be used for alignment in either memory or a data stream; * An alternative approach suitable for packets written to memory is to divide memory into blocks of M bytes (e.g. 1kbyte blocks), and write packets to memory such that the first byte in every block is always the first byte of a packet. This means packets cannot span block boundaries, and so zero bytes must be used to pad between the end of the last message in a block and the block boundary. ### [](#sec:format3)7.1.1\. Format 3 packets Format 3 packets are used for synchronization, traps, reporting context and supporting information. There are 4 sub-formats. Throughout this document, the term "synchronization packet" is used. This refers specifically to format 3, subformat 0 and subformat 1 packets. ### [](#sec:format30)7.1.2\. Format 3 subformat 0 - Synchronisation This packet contains all the information the decoder needs to fully identify an instruction. It is sent for the first traced instruction (unless that instruction also happens to be the first in a trap handler), and when resynchronization has been scheduled by expiry of the resynchronisation timer. __Table 1\. Packet format 3, subformat 0__ | **Field name** | **Bits** | **Description** | | -------------- | ------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 2 | 11 (sync): synchronisation | | **subformat** | 2 | 00 (start): Start of tracing, or resync | | **branch** | 1 | Set to 0 if the address points to a branch instruction, and the branch was taken. Set to 1 if the instruction is not a branch or if the branch is not taken. | | **privilege** | _privilege\_width\_p_ | The privilege level of the reported instruction | | **time** | _time\_width\_p_ or 0 if _notime\_p_ is 1 | The time value. | | **context** | _context\_width\_p_, or 0 if _nocontext\_p_ is 1 | The instruction context. | | **address** | _iaddress\_width\_p - iaddress\_lsb\_p_ | Full instruction address. Address alignment is determined by _iaddress\_lsb\_p_: **address**must be left shifted by _iaddress\_lsb\_p_ in order to recreate original byte address. | #### [](#7-1-2-1-format-3-branch-field)7.1.2.1\. Format 3 **branch** field This bit indicates the taken/not taken status in the case where the reported address points to a branch instruction. Overall efficiency would be slightly improved if this bit was removed, and the branch status was instead "carried over" and reported in the next _te\_inst_packet. This was considered, but there are several pathological cases where this approach fails. Consider for example the situation where the first traced instruction is a branch, and this is then followed immediately by an exception. This results in format 3 packets being generated on two consecutive instructions. The second packet does not contain a branch map, so there is no way to report the branch status of the 1st branch, apart from by inserting a format 1 packet in between. There are two issues with this: * It would require the generation of 2 packets on the same cycle, which adds significant additional complexity to the encoder; * It would complicate the algorithm shown in[\[fig:algo\]](#fig:algo). ### [](#sec:format31)7.1.3\. Format 3 subformat 1 - Trap This packet also contains all the information the decoder needs to fully identify an instruction. It is sent following an exception or interrupt, and includes the cause, the 'trap value' (for exceptions), and the address of the trap handler, or of the exception itself - [7.1.3.1\. Format 3 **thaddr**, **address** and **privilege** fields](#sec:thaddr). If the implicit exception mode is enabled (see[\[sec:implicit-exception\]](#sec:implicit-exception)), the trap handler address is omitted if **thaddr** is 1. __Table 2\. Packet format 3, subformat 1__ | **Field name** | **Bits** | **Description** | | -------------- | ------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 2 | 11 (sync): synchronisation | | **subformat** | 2 | 01 (trap): Exception or interrupt cause and trap handler address. | | **branch** | 1 | Set to 0 if the address points to a branch instruction, and the branch was taken. Set to 1 if the instruction is not a branch or if the branch is not taken. | | **privilege** | _privilege\_width\_p_ | The privilege level of the reported instruction. | | **time** | _time\_width\_p_ or 0 if _notime\_p_ is 1 | The time value. | | **context** | _context\_width\_p_, or 0 if _nocontext\_p_ is 1 | The instruction context | | **ecause** | _ecause\_width\_p_ | Exception or interrupt cause. | | **interrupt** | 1 | Interrupt. | | **thaddr** | 1 | When set to 1, **address** points to the trap handler address. When set to 0, **address** points to the EPC for an exception at the target of an updiscon, and is undefined for other exceptions and interrupts. | | **address** | _iaddress\_width\_p - iaddress\_lsb\_p_ | Full instruction address. Address alignment is determined by _iaddress\_lsb\_p_: **address**must be left shifted by\_iaddress-lsb\_p\_ in order to recreate original byte address. | | **tval** | _iaddress\_width\_p_ | Value from appropriate**utval/stval/vstval/mtval** CSR. Field omitted for interrupts | #### [](#sec:thaddr)7.1.3.1\. Format 3 **thaddr**, **address** and **privilege** fields If an exception occurs at the target of an uninferable PC discontinuity, the value of the EPC cannot be infered from the program binary, and so**address** contains the EPC and **thaddr** is set to 0\. In this case, the trap handler address will be reported via a subsequent format 3, subformat 0 packet. An exception occuring on the 1st traced instruction is treated in the same way. Usually when an exception or interrupt occurs, the cause is reported along with the 1st address of the trap handler, when that instruction retires. In this case, **thaddr** is 1\. However, if a second interrupt or exception occurs immediately, details of this must still be reported, even though the 1st instruction of the handler hasn’t retired. In this situation, **thaddr** is 0, and **address** is undefined (unless it contains the EPC as outlined in the previous paragraph). Note: The reason for not reporting the EPC for all exceptions when **thaddr**is 0 is that it can be inferred by the decoder using the exception cause. Where an interrupt or exception causes a privilege change, this change comes into force from the start of the trap handler. As such, the privilege value reported when **thaddr** is 0 will be the privilege level prior to any change. #### [](#7-1-3-2-format-3-tval-field)7.1.3.2\. Format 3 **tval** field This field reports the "trap value" from the appropriate**utval/stval/vstval/mtval** CSR, the meaning of which is dependent on the nature of the exception. It is omitted from the packet for interrupts. ### [](#sec:format32)7.1.4\. Format 3 subformat 2 - Context This packet contains only the context and/or the timestamp, and is output when the context value changes and can be reported imprecisely (see [\[tab:context-type\]](#tab:context-type)). __Table 3\. Packet format 3, subformat 2__ | **Field name** | **Bits** | **Description** | | -------------- | ------------------------------------------------ | --------------------------------------- | | **format** | 2 | 11 (sync): synchronisation | | **subformat** | 2 | 10 (context): Context change | | **privilege** | _privilege\_width\_p_ | The privilege level of the new context. | | **time** | _time\_width\_p_ or 0 if _notime\_p_ is 1 | The time value | | **context** | _context\_width\_p_, or 0 if _nocontext\_p_ is 1 | The instruction context. | ### [](#sec:format33)7.1.5\. Format 3 subformat 3 - Support This packet provides supporting information to aid the decoder. It must be issued: * When trace is enabled, and before the first sync packet, in order to ensure the decoder is aware of how the encoder is configured. This could be as late as when **trTeInstTracing** becomes 1, but it is recommended this be sent as soon as **trTeEnable** changes from 0 to 1\. This reduces the likelihood of having to generate two packets (support and sync-start) at the point tracing actually starts; * When tracing ceases for any reason (**trTeEnable** or **trTeInstTracing** set to 0, trace-off trigger, halt, reset, loss of filter qualification, etc.), in order to inform the decoder that the preceding packet reported the address of the final traced instruction; * If one or more trace packets cannot be sent (for example, due back-pressure from the packet transport infrastructure); * If the operating mode of the encoder changes such that the the information output in a previous support packet no longer applies (i.e. if **encoder\_mode**, **ioptions**, **doptions** or **denable** change). The **ioptions** and **doptions** fields are placeholders that must be replaced by an implementation specific set of individual bits - one for each of the optional modes supported by the encoder. __Table 4\. Packet format 3, subformat 3__ | **Field name** | **Bits** | **Description** | | ----------------- | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 2 | 11 (sync): synchronisation | | **subformat** | 2 | 11 (support): Supporting information for the decoder | | **ienable** | 1 | Indicates if the instruction trace encoder is enabled | | **encoder\_mode** | _E_ | Identifies trace algorithm Details and number of bits implementation dependent. Currently Branch trace is the only mode defined, indicated by the value 0. | | **qual\_status** | 2 | 00 (no\_change): No change to filter qualification01 (ended\_rep): Qualification ended, preceding **te\_inst** sent explicitly to indicate final qualification instruction10 (trace\_lost): One or more instruction trace packets lost.11 (ended\_ntr): Qualification ended, preceding **te\_inst** would have been sent anyway due to an updiscon, even if it wasn’t the final qualified instruction | | **ioptions** | _N_ | Values of all instruction trace run-time configuration bitsNumber of bits and definitions implementation dependent. Examples might be\- 'sequentially inferred jumps' Don’t report the targets of sequentially inferable jumps\- 'implicit return' Don’t report function return addresses\- 'implicit exception' Exclude address from format 3, sub-format 1 _te\_inst_ packets if trap vector can be determined from_ecause_\- 'branch prediction' Branch predictor enabled\- 'jump target cache' Jump target cache enabled\- 'full address' Always output full addresses (SW debug option) | | **denable** | 1 | Indicates if the data trace is enabled (if supported) | | **dloss** | 1 | One of more data trace packets lost (if supported) | | **doptions** | M | Values of all data trace run-time configuration bits Number of bits and definitions implementation dependent. Examples might be\- 'no data' Exclude data (just report addresses)\- 'no addr' Exclude address (just report data) | #### [](#sec:qual-status)7.1.5.1\. Format 3 subformat 3 **qual\_status** field When tracing ends, the encoder reports the address of the final traced instruction, and follows this with a format 3, subformat 3 (supporting information) packet. Two codes are provided for indicating that tracing has ended: **ended\_rep** and **ended\_ntr**. This relates to exactly the same ambiguous case described in detail in [7.1.6.2\. Format 2 **notify** and **updiscon** fields](#sec:updiscon), and in principle, the mechanism described in that section can be used to disambiguate when the final traced instruction is at looplabel. However, that mechanism relies on knowing when creating the format 1/2 packet, that a format 3 packet will be generated from the next instruction. This is possible because the encoding algorithm uses a 3-stage pipe with access to the previous, current and next instructions. However, decoding that the next instruction is a privilege change or exception is straightforward, but determining whether the next instruction meets the filtering criteria is much more involved, and this information won’t typically be available, at least not without adding an additional pipeline stage, which is expensive. This means a different mechanism is required, and that is provided by having two codes to indicate that tracing has ended: * **ended\_rep** indicates that the preceding packet would not have been issued if tracing hadn’t ended, which means that tracing stopped after executing looplabel in the 1st loop iteration; * **ended\_ntr** indicates that the preceding packet would have been issued anyway because of an uninferable PC discontinuity, which means that tracing stopped after executing looplabel in the 2nd loop iteration; If the encoder implementation does have early access to the filtering results, and the designer chooses to use the **updiscon** bit when the last qualified instruction is also the instruction following an uninferable PC discontinuity, loss of qualification should always be indicated using **ended\_rep**. ### [](#sec:format2)7.1.6\. Format 2 packets This packet contains only an instruction address, and is used when the address of an instruction must be reported, and there is no unreported branch information. The address is in differential format unless full address mode is enabled (see [\[sec:full-address\]](#sec:full-address)). __Table 5\. Packet format 2__ | **Field name** | **Bits** | **Description** | | -------------- | ---------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 2 | 10 (addr-only): differential address and no branch information | | **address** | _iaddress\_width\_p - iaddress\_lsb\_p_ | Differential instruction address. | | **notify** | 1 | If the value of this bit is different from the MSB of**address**, it indicates that this packet is reporting an instruction that is not the target of an uninferable discontinuity because a notification was requested via **trigger\[2\]** (see [\[sec:trigger\]](#sec:trigger)). | | **updiscon** | 1 | If the value of this bit is different from **notify**, it indicates that this packet is reporting the instruction following an uninferable discontinuity and is also the instruction before an exception, privilege change or resync (i.e. it will be followed immediately by a format 3 _te\_inst_). | | **irreport** | 1 | If the value of this bit is different from **updiscon**, it indicates that this packet is reporting an instruction that is either: following a return because its address differs from the predicted return address at the top of the implicit\_return return address stack, or the last retired before an exception, interrupt, privilege change or resync because it is necessary to report the current address stack depth or nested call count. | | **irdepth** | _return\_stack\_size\_p + (return\_stack\_size\_p > 0 ? 1 : 0) + call\_counter\_size\_p_ | If the value of **irreport** is different from**updiscon**, this field indicates the number of entries on the return address stack (i.e. the entry number of the return that failed) or nested call count. If **irreport** is the same value as **updiscon**, all bits in this field will also be the same value as **updiscon**. | #### [](#sec:notify)7.1.6.1\. Format 2 **notify** field This bit is encoded so that most of the time it will take the same value as the MSB of the **address** field, and will therefore compress away, having no impact on the encoding efficiency. It is required in order to cover the case where an address is reported as a result of a notification request, signalled by setting the **trigger\[2\]** input to 1. #### [](#sec:updiscon)7.1.6.2\. Format 2 **notify** and **updiscon** fields These bits are encoded so that most of the time they will compress away, having no impact on efficiency, by taking on the same value as the preceding bit in the packet (**notify** is normally the same value as the MSB of the **address** field, and **updiscon** is normally the same value as**notify**). They are required in order to cover a pathological case where otherwise the decoding software would not be able to reconstruct the program execution unambiguously. Consider the following code fragment: looplabel -4: *_opcode A_* looplabel : *_opcode B_* looplabel +4: *_opcode C_* : looplabel +N *_JALR_* # Jump to looplabel This is a loop with an indirect jump back to the next iteration. This is an uninferable discontinuity, and will be reported via a format 1 or 2 packet. Note however that the initial entry into the loop is fall-through from the instruction at looplabel - 4, and will not be reported explicitly. This means that when reconstructing the execution path of the program, the looplabel address is encountered twice. On first glance, it appears that the decoder can determine when it reaches the loop label for the 1st time that this is not the end of execution, because the preceding instruction was not one that can cause an uninferable discontinuity. It can therefore continue reconstructing the execution path until it reaches the **_JALR_**, from where it can deduce that **_opcode B_** at looplabel is the final retired instruction. However, there are circumstances where this approach does not work. For example, consider the case where there is an exception at looplabel + 4\. In this case, the decoder cannot tell whether this occurred during the 1st or 2nd loop iterations, without additional information from the encoder. This is the purpose of the **updiscon** field. In more detail: There are four scenarios to consider: 1. Code executes through to the end of the 1st loop iteration, and the encoder reports looplabel using format 1/2 following the **_JALR_**, then carries on executing the 2nd pass of the loop. In this case **updiscon**\== 7.2\. **notify**. The next packet will be a format 1/2; 2. Code executes through to the end of the 1st loop iteration and jumps back to looplabel, but there is then an exception, privilege change or resync in the second iteration at looplabel + 4\. In this case, the encoder reports looplabel using format 1/2 following the **_JALR_**, with**updiscon** \== !**notify**, and the next packet is a format 3; 3. An exception occurs immediately after the 1st execution of looplabel. In this case, the encoder reports looplabel using format 0/1/2 with**updiscon** \== **notify**, and the next packet is a format 3; 4. The hart requests the encoder to notify retirement of the instruction at looplabel. In this case, the encoder reports the 1st execution of looplabel with **notify** \== !**address\[MSB\]**, and subsequent executions with **notify** \== **address\[MSB\]** (because they would have been reported anyway as a result of the **_JALR_**). Looking at this from the perspective of the decoder, the decoder receives a format 1/2 reporting the address of the 1st instruction in the loop (looplabel). It follows the execution path from the previous reported address, until it reaches looplabel. Because looplabel is not preceded by an uninferable discontinuity, it must take the value of**notify** and **updiscon** into consideration, and may need to wait for the next packet in order to determine whether it has reached the most recently retired instruction: * If **updiscon** \== !**notify**, this indicates case 2\. The decoder must continue until it encounters looplabel a 2nd time; * If **updiscon** \== **notify**, the decoder cannot yet distinguish cases 1 and 3, and must wait for the next packet. * If the next packet is a format 3, this is case 3\. The decoder has already reached the correct instruction; * If the next packet is a format 1/2, this is case 1\. The decoder must continue until it encounters looplabel a 2nd time. * If **notify** \== !**address\[MSB\]**, this indicates case 4, 1st iteration. The decoder has reached the correct instruction. This example uses an exception at looplabel + 4, but anything that could cause a format 3 for looplabel + 4 would result in the same behavior: a privilege change, or the expiry of the resync timer. It could also occur if looplabel was the final traced instruction (because tracing was disabled for some reason). See [7.1.5.1\. Format 3 subformat 3 **qual\_status** field](#sec:qual-status) for further discussion of this point. | | Correct decoder behavior could have been achieved by implementing the **notify** bit only, setting it to the inverse of**address\[MSB\]** whenever an address is reported and it is not the instruction following an uninferable discontinuity. However, this would have been much less efficient, as this would have required **notify** to be different from **address\[MSB\]** the majority of the time when outputting a format 1/2 before an exception, interrupt or resync (as the probability of this instruction being the target of an uninferable jump is low). Using 2 separate bits results in superior compression. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#sec:irxx)7.2.1\. Format 2 **irreport** and **irdepth** These bits are encoded so that most of the time they will take the same value as the **updiscon** field, and will therefore compress away, having no impact on the encoding efficiency. If implicit\_return mode is enabled, the encoder keeps track of the number of traced nested calls, either as a simple count (_call\_counter\_size\_p_ non-zero) or a stack of predicted return addresses (_return\_stack\_size\_p_ non-zero). Where a stack of predicted return addresses is implemented, the predicted return addresses are compared with the actual return addresses, and a _te\_inst_ packet will be generated with **irreport** set to the opposite value to **updiscon** if a misprediction occurs. In some cases it is also necessary to report the current stack depth or call count if the packet is reporting the instruction immediately before an exception, interrupt, privilege change or resync. There are two cases of concern: * If the reported address is the instruction following a return, and it is not mis-predicted, the encoder must report the current stack depth or call count if it is non-zero. Without this, the decoder would attempt to follow the execution path until it encountered the reported address from the outermost nested call; * If the reported address is not the instruction following a return, the encoder must report the current stack depth or call count unless: * There have been no returns since the previous call (in which case the decoder will correctly stop in the innermost call), or * There has been at least one branch since the previous return (in which case the decoder will correctly stop in the call where there are no unprocessed branches). Without this, the decoder would follow the execution path until it encountered the reported address, and in most cases this would be the correct point. However, this cannot be guaranteed for recursive functions, as the reported address will occur multiple times in the execution path. ### [](#sec:format1)7.2.1\. Format 1 packets This packet includes branch information, and is used when either the branch information must be reported (for example because the branch map is full), or when the address of an instruction must be reported, and there has been at least one branch since the previous packet. If included, the address is in differential format unless full address mode is enabled (see [\[sec:full-address\]](#sec:full-address)). __Table 6\. Packet format 1 - address, branch map__ | **Field name** | **Bits** | **Description** | | --------------- | ---------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 2 | 01 (diff-delta): includes branch information and may include differential address | | **branches** | 5 | Number of valid bits **branch\_map**. The number of bits of **branch\_map** is determined as follows:0: (cannot occur for this format)1: 1 bit2-3: 3 bits4-7: 7 bits8-15: 15 bits16-31: 31 bitsFor example if branches = 12, **branch\_map** is 15 bits long, and the 12 LSBs are valid. | | **branch\_map** | Determined by **branches** field. | An array of bits indicating whether branches are taken or not. Bit 0 represents the oldest branch instruction executed. For each bit: : branch taken : branch not taken | | **address** | _iaddress\_width\_p - iaddress\_lsb\_p_ | Differential instruction address. | | **notify** | 1 | If the value of this bit is different from the MSB of**address**, it indicates that this packet is reporting an instruction that is not the target of an uninferable discontinuity because a notification was requested via **trigger\[2\]** (see [\[sec:trigger\]](#sec:trigger)). | | **updiscon** | 1 | If the value of this bit is different from the MSB of**notify**, it indicates that this packet is reporting the instruction following an uninferable discontinuity and is also the instruction before an exception, privilege change or resync (i.e. it will be followed immediately by a format 3 _te\_inst_). | | **irreport** | 1 | If the value of this bit is different from **updiscon**, it indicates that this packet is reporting an instruction that is either: following a return because its address differs from the predicted return address at the top of the implicit\_return return address stack, or the last retired before an exception, interrupt, privilege change or resync because it is necessary to report the current address stack depth or nested call count. | | **irdepth** | _return\_stack\_size\_p + (return\_stack\_size\_p > 0 ? 1 : 0) + call\_counter\_size\_p_ | If the value of **irreport** is different from**updiscon**, this field indicates the number of entries on the return address stack (i.e. the entry number of the return that failed) or nested call count. If **irreport** is the same value as **updiscon**, all bits in this field will also be the same value as **updiscon**. | __Table 7\. Packet format 1 - no address, branch map__ | **Field name** | **Bits** | **Description** | | --------------- | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 2 | 01 (diff-delta): includes branch information and may include differential address | | **branches** | 5 | Number of valid bits in **branch\_map**. The length of**branch\_map** is determined as follows:0: 31 bits, no **address** in packet1-31: (cannot occur for this format) | | **branch\_map** | 31 | An array of bits indicating whether branches are taken or not. Bit 0 represents the oldest branch instruction executed. For each bit: : branch taken : branch not taken | #### [](#7-2-1-1-format-1-updiscon-field)7.2.1.1\. Format 1 **updiscon** field See [7.1.6.2\. Format 2 **notify** and **updiscon** fields](#sec:updiscon). #### [](#7-2-1-2-format-1-branch%5Fmap-field)7.2.1.2\. Format 1 **branch\_map** field When the branch map becomes full it must be reported, but in most cases there is no need to report an address. This is indicated by setting**branches** to 0\. If the address does need to be reported for any reason (for example because the instruction immediately prior to the final branch causes an uninferable discontinuity) **branches** is set to 31. The choice of sizes (1, 3, 7, 15, 31) is designed to minimize efficiency loss. On average there will be some 'wasted' bits because the number of branches to report is less than the selected size of the **branch\_map**field. Using a tapered set of sizes means that the number of wasted bits will on average be less for shorter packets. If the number of branches between updiscons is randomly distributed then the probability of generating packets with large branch counts will be lower, in which case increased waste for longer packets will have less overall impact. Furthermore, the rate at which packets are generated can be higher for lower branch counts, and so reducing waste for this case will improve overall bandwidth at times where it is most important. #### [](#7-2-1-3-format-1-irreport-and-irdepth-fields)7.2.1.3\. Format 1 **irreport** and **irdepth** fields See [7.2.1\. Format 2 **irreport** and **irdepth**](#sec:irxx). ### [](#sec:format0)7.2.2\. Format 0 packets This format is intended for optional efficiency extensions. Currently two extensions are defined, for reporting counts of correctly predicted branches, and for reporting the jump target cache index. If branch prediction is supported and is enabled, then there is a choice of whether to output a full branch map (via format 1), or a count of correctly predicted branches. The count format is used if the number of correctly predicted branches is at least 31\. If there are 31 unreported branches (i.e. the branch map is full), but not all of them were predicted correctly, then the branch map will be output. A branch count will be output under the following conditions: * A branch is mis-predicted. The count value will be the number of correctly predicted branches, minus 31\. No address information is provided - it is implicitly that of the branch which failed prediction; * An updiscon, interrupt or exception requires the encoder to output an address. In this case the encoder will output the branch count (number of correctly predicted branches, minus 31); * The branch count reaches its maximum value (0xffff ffff). Strictly speaking an address isn’t required for this case, but is included to avoid having to distinguish the packet format from the case above. It will occur so rarely that the bandwidth impact can be ignored. Note: the decoder is reliant on the maximum value in order to identify this special case, so the 32-bit counter must be implemented in full. If a jump target cache is supported and enabled, and the address to report following an updiscon is in the cache then the encoder can output the cache index using format 0, subformat 1\. However, the encoder may still choose to output the differential address using format 1 or 2 if the resulting packet is shorter. This may occur if the differential address is zero, or very small. __Table 8\. Packet format 0, subformat 0 - no address, branch count__ | **Field name** | **Bits** | **Description** | | ----------------- | -------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 2 | 00 (opt-ext): formats for optional efficiency extensions | | **subformat** | See [7.2.2.1\. Format 0 subformat field](#sec:f0s) | 0 (correctly predicted branches) | | **branch\_count** | 32 | Count of the number of correctly predicted branches, minus 31. | | **branch\_fmt** | 2 | 00 (no-addr): Packet does not contain an **address**, and the branch following the previous correct prediction failed.01 - 11: (cannot occur for this format) | __Table 9\. Packet format 0, subformat 0 - address, branch count__ | **Field name** | **Bits** | **Description** | | ----------------- | ---------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 2 | 00 (opt-ext): formats for optional efficiency extensions | | **subformat** | See [7.2.2.1\. Format 0 subformat field](#sec:f0s) | 0 (correctly predicted branches) | | **branch\_count** | 32 | Count of the number of correctly predicted branches, minus 31. | | **branch\_fmt** | 2 | 10 (addr): Packet contains an **address**. If this points to a branch instruction, then the branch was predicted correctly.11 (addr-fail): Packet contains an **address** that points to a branch which failed the prediction.00, 01: (cannot occur for this format) | | **address** | _iaddress\_width\_p - iaddress\_lsb\_p_ | Differential instruction address. | | **notify** | 1 | If the value of this bit is different from the MSB of**address**, it indicates that this packet is reporting an instruction that is not the target of an uninferable discontinuity because a notification was requested via **trigger\[2\]** (see[\[sec:trigger\]](#sec:trigger)). | | **updiscon** | 1 | If the value of this bit is different from **notify**, it indicates that this packet is reporting the instruction following an uninferable discontinuity and is also the instruction before an exception, privilege change or resync (i.e. it will be followed immediately by a format 3 _te\_inst_). | | **irreport** | 1 | If the value of this bit is different from **updiscon**, it indicates that this packet is reporting an instruction that is either: following a return because its address differs from the predicted return address at the top of the implicit\_return return address stack, or the last retired before an exception, interrupt, privilege change or resync because it is necessary to report the current address stack depth or nested call count. | | **irdepth** | _return\_stack\_size\_p + (return\_stack\_size\_p > 0 ? 1 : 0) + call\_counter\_size\_p_ | If the value of **irreport** is different from**updiscon**, this field indicates the number of entries on the return address stack (i.e. the entry number of the return that failed) or nested call count. If **irreport** is the same value as **updiscon**, all bits in this field will also be the same value as **updiscon**. | __Table 10\. Packet format 0, subformat 1 - jump target index, branch map__ | **Field name** | **Bits** | **Description** | | --------------- | ---------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | **format** | 2 | 00 (opt-ext): formats for optional efficiency extensions | | **subformat** | See [7.2.2.1\. Format 0 subformat field](#sec:f0s) | 1 (jump target cache) | | **index** | _cache\_size\_p_ | Jump target cache index of entry containing target address. | | **branches** | 5 | Number of valid bits in **branch\_map**. The length of**branch\_map** is determined as follows:+ 0: (cannot occur for this format)1: 1 bit2-3: 3 bits4-7: 7 bits8-15: 15 bits16-31: 31 bitsFor example if branches = 12, **branch\_map** is 15 bits long, and the 12 LSBs are valid. | | **branch\_map** | Determined by **branches** field. | An array of bits indicating whether branches are taken or not. Bit 0 represents the oldest branch instruction executed. For each bit: : branch taken : branch not taken | | **irreport** | 1 | If the value of this bit is different from**branch\_map\[MSB\]**, it indicates that this packet is reporting an instruction that is either: following a return because its address differs from the predicted return address at the top of the implicit\_return return address stack, or the last retired before an exception, interrupt, privilege change or resync because it is necessary to report the current address stack depth or nested call count. | | **irdepth** | _return\_stack\_size\_p + (return\_stack\_size\_p > 0 ? 1 : 0) + call\_counter\_size\_p_ | If the value of **irreport** is different from**branch\_map\[MSB\]**, this field indicates the number of entries on the return address stack (i.e. the entry number of the return that failed) or nested call count. If **irreport** is the same value as**branch\_map\[MSB\]**, all bits in this field will also be the same value as**branch\_map\[MSB\]**. | __Table 11\. Packet format 0, subformat 1 - jump target index, no branch map__ | **Field name** | **Bits** | **Description** | | -------------- | ---------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **format** | 2 | 00 (opt-ext): formats for optional efficiency extensions | | **subformat** | See [7.2.2.1\. Format 0 subformat field](#sec:f0s) | 1 (jump target cache) | | **index** | _cache\_size\_p_ | Jump target cache index of entry containing target address. | | **branches** | 5 | Number of valid bits in **branch\_map**. The length of**branch\_map** is determined as follows: : no **branch\_map** in packet -31: (cannot occur for this format) | | **irreport** | 1 | If the value of this bit is different from**branches\[MSB\]**, it indicates that this packet is reporting an instruction that is either: following a return because its address differs from the predicted return address at the top of the implicit\_return return address stack, or the last retired before an exception, interrupt, privilege change or resync because it is necessary to report the current address stack depth or nested call count. | | **irdepth** | _return\_stack\_size\_p + (return\_stack\_size\_p > 0 ? 1 : 0) + call\_counter\_size\_p_ | If the value of **irreport** is different from**branches\[MSB\]**, this field indicates the number of entries on the return address stack (i.e. the entry number of the return that failed) or nested call count. If **irreport** is the same value as**branches\[MSB\]**, all bits in this field will also be the same value as**branches\[MSB\]**. | #### [](#sec:f0s)7.2.2.1\. Format 0 subformat field The width of this field depends on the number of optional formats supported. Currently, two optional formats are defined (correctly predicted branches and jump target cache). The width is specified by the_f0s\_width_ discovery field (see [\[sec:disco\]](#sec:disco)). If multiple optional formats are supported, the field width must be non-zero. However, if only one optional format is supported, the field can be omitted, and the value of the field inferred from the **options**field in the support packet (see [7.1.5\. Format 3 subformat 3 - Support](#sec:format33). This provision allows additional formats to be added in future without reducing the efficiency of the existing formats. #### [](#7-2-2-2-format-0-branch%5Ffmt-field)7.2.2.2\. Format 0 **branch\_fmt** field This is encoded so that when no address is required it will be zero, allowing the upper bits of the **branch\_count** field to be compressed away. When a branch count is reported without an address it is because a branch has failed the prediction. However, when an address is reported along with a branch count, it will be because the packet was initiated by an uninferable discontinuity, an exception, or because a branch has been encountered that increments **branch\_count** to 0xffff\_ffff. For the latter case, the reported address will always be for a branch, and in the former cases it may be. If it is a branch, it is necessary to be explicit about whether or not the prediction was met or not. If it is met, then the reported address is that of the last correctly predicted branch. #### [](#7-2-2-3-format-0-irreport-and-irdepth-fields)7.2.2.3\. Format 0 **irreport** and **irdepth** fields These bits are encoded so that most of the time they will take the same value as the immediately preceding bit (**updiscon**, **branch\_map\[MSB\]** or**branches\[MSB\]** depending on the specific packet format). Purpose and behavior is as described in [7.2.1\. Format 2 **irreport** and **irdepth**](#sec:irxx). For the jump target cache (subformat 1), they are included to allow return addresses that fail the implicit return prediction but which reside in the jump target cache to be reported using this format. An implementation could omit these if all implicit return failures are reported using format 1. 6.1. Timestamping ==================== ## [](#ch:timestamping)6.1\. Timestamping The support for Timestamps is optional and so the contents of this chapter are informative only. In many systems it is desirable to periodically insert a timestamp packet into the trace stream, effectively marking that point in the stream with a time value. This can be used to judge "time" between various points in the trace stream and, more notably, to be able to correlate trace streams from different harts (i.e. this point in hart A’s stream occurred at roughly the same time as that point in hart B’s trace stream). The former helps one to judge performance of sections of code execution (to the granularity of timestamp insertion). The latter helps debugging multi-hart MP problems. An implementation may have the following: * A timestamp is (up to) a 64-bit time value. * Configurable options for generating timestamp values such as a hart’s 'time' values or 'cycle' values. * Options could may also include things like taking 'time' values with the low 4 or 8 bits dropped off which would create a coarser granularity time values * Timestamp generation may be enabled or disabled. If enabled, a timestamp packet would be generated periodically which may be based on configurable interval or rate, e.g. once every 2n items where 'n' and 'items' are configurable among some limited set of choices. The choices could be: * Time * Time scaled down. An implementation specific scaled or divided down derivative of time. This may be useful in providing a smaller coarser graularity values * Time Interpolated up. An implementation specific interpolated up derivative of time. This may be useful in providing higher resolution time values * Cycle * Implementation specific * A timestamp packet may also be generated in conjunction with a sync packet * Timestamp packets are highly compressible and variable in size depending on the number of low bits of the current value that have changed wrt the last emitted timestamp value. If timestamp packets are emitted rarely (but not as rare as sync packets), then they will tend to be, say, 2-4 bytes in size (still much less than the full up to 64-bit size). If timestamp packets are emitted somewhat frequently, then they will tend to be 1-2 bytes in size. If timestamp packets are emitted very frequently, then they will tend to be <1 byte in size. Timestamp values associated with sync packets would always be the full implemented size. RISC-V Capacity and Bandwidth QoS Register Interface ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-capacity-and-bandwidth-qos-register-interface)RISC-V Capacity and Bandwidth QoS Register Interface RISC-V CBQRI Task Group Version v1.0, 2024-07-02: Ratified | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at. Copyright 2024 by RISC-V International. 4.1. Bandwidth-controller QoS Register Interface ==================== ## [](#BC%5FQOS)4.1\. Bandwidth-controller QoS Register Interface Controllers, such as memory controllers, that support bandwidth allocation and bandwidth usage monitoring provide a memory-mapped bandwidth-controller QoS register interface. __Table 1\. Bandwidth-controller QoS Register Layout (size and offset are in bytes)__ | Offset | Name | Size | Description | Optional? | | ------ | ----------------- | ---- | ------------------------------------------- | --------- | | 0 | bc\_capabilities | 8 | [Capabilities](#BC%5FCAP) | No | | 8 | bc\_mon\_ctl | 8 | [Usage monitoring control](#BC%5FMCTL) | Yes | | 16 | bc\_mon\_ctr\_val | 8 | [Monitoring counter value](#BC%5FMCTR) | Yes | | 24 | bc\_alloc\_ctl | 8 | [Bandwidth allocation control](#BC%5FALLOC) | Yes | | 32 | bc\_bw\_alloc | 8 | [Bandwidth allocation](#BC%5FBMASK) | Yes | The reset value is 0 for the following register fields. * `bc_mon_ctl.BUSY` field * `bc_alloc_ctl.BUSY` field The reset value is `UNSPECIFIED` for all other register fields. The bandwidth controllers at reset must allocate all available bandwidth to`RCID` value of 0\. When the bandwidth controller supports bandwidth allocation per access-type, the access-type value of 0 of `RCID=0` is allocated all available bandwidth, while all other access-types associated with that `RCID`share the bandwidth allocation with `AT=0`. The bandwidth allocation for all other `RCID` values is `UNSPECIFIED`. The bandwidth controller behavior in handling a request with a non-zero `RCID` value before configuring the bandwidth controller with bandwidth allocation for that `RCID` is also `UNSPECIFIED`. ### [](#BC%5FCAP)4.1.1\. Bandwidth-controller Capabilities The `bc_capabilities` register is a read-only register that holds the bandwidth-controller capabilities. ![Bandwidth-controller Capabilities Register](_images/diag-a4a69ee6f09305554dc691b79852eb20e7cc61ee.svg) Figure 1\. Bandwidth-controller Capabilities Register The `VER` field holds the version of the specification implemented by the bandwidth controller. The low nibble is used to hold the minor version of the specification and the upper nibble is used to hold the major version of the specification. For example, an implementation that supports version 1.0 of the specification reports `0x10`. The `NBWBLKS` field holds the total number of available bandwidth blocks in the controller. The bandwidth represented by each bandwidth block is`UNSPECIFIED`. The bandwidth controller supports reserving bandwidth in multiples of a bandwidth block, which enables proportional allocation of bandwidth. Bandwidth controllers can limit the maximum bandwidth that can be reserved to a value smaller than the `NBWBLKS` value. The `MRBWB` field reports the maximum number of bandwidth blocks that can be reserved. | | The bandwidth controller meters the bandwidth usage by a workload to determine if it is exceeding its allocations and, if necessary, take measures to throttle the workload’s bandwidth usage. Therefore, the instantaneous bandwidth used by a workload either exceeds or falls short of the configured allocation. QoS capabilities are statistical in nature and are typically designed to enforce the configured bandwidth over larger time windows. By not allowing all available bandwidth blocks to be reserved for allocation, the bandwidth controller can handle such transient inaccuracies. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When the `RCID`\-prefixed mode (`RPFX`) is 1, the controller operates in `RPFX`mode. The parameter `P` (prefixed bits) indicates the number of least significant bits of the `MCID` carried in the request that should be prefixed by the RCID. The controller uses `RCID` in the requests along with `P` number of least significant bits of the `MCID` to compute an effective `MCID` ([\[EMCID\]](#EMCID)). This `MCID` is used to identify the monitoring counter. Legal values of `P` range from 0 to 12. If `RPFX` is 0, `P` is set to 0, and the effective `MCID` is the same as the `MCID`in the request. ### [](#BC%5FMCTL)4.1.2\. Bandwidth Usage Monitoring Control The `bc_mon_ctl` register controls the monitoring of bandwidth usage by a`MCID`. When the controller does not support bandwidth usage monitoring, the`bc_mon_ctl` register is read-only zero. ![Bandwidth Usage Monitoring Control Register (`bc_mon_ctl`)](_images/diag-162d092e66bc48cb21fbd7a767a4dfd60241b10e.svg) Figure 2\. Bandwidth Usage Monitoring Control Register (`bc_mon_ctl`) Bandwidth controllers that support bandwidth usage monitoring implement a usage monitoring counter for each supported `MCID`. The usage monitoring counter might be configured to count a monitoring event. When an event matching the event configured for the `MCID` occurs, then the monitoring counter is updated. The event matching can optionally be filtered by the access-type. The monitoring counter for bandwidth usage counts the number of bytes transferred by requests matching the monitoring event as the requests go past the monitoring point. The `OP`, `AT`, `MCID`, and `EVT_ID` fields of the register are WARL fields. The `OP` field is used to instruct the controller to perform an operation listed in [Table 2](#BC%5FMON%5FOP). __Table 2\. Bandwidth Usage Monitoring Operations (OP)__ | Operation | Encoding | Description | | ------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | | — | 0 | Reserved for future standard use. | | CONFIG\_EVENT | 1 | Configure the counter selected by MCID to count the event selected by EVT\_ID, AT, and ATV. The EVT\_ID encodings are listed in [Table 3](#BC%5FEVT%5FID). | | READ\_COUNTER | 2 | Snapshot the value of the counter selected byMCID into bc\_mon\_ctr\_val register. TheEVT\_ID, AT, and ATV fields are not used by this operation. | | — | 3-23 | Reserved for future standard use. | | — | 24-31 | Designated for custom use. | The `CONFIG_EVENT` operation uses the `EVT_ID` operand to program the identifier of the event to count in the monitoring counter, selected by `MCID`. The `AT` field is used to program the access-type to count, and its validity is indicated by the `ATV` field. When `ATV` is 0, the counter counts requests with all access-types, and the `AT` value is ignored. __Table 3\. Bandwidth Monitoring Event ID (EVT\_ID)__ | Event ID | Encoding | Description | | ------------------------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------- | | None | 0 | Counter does not count. | | Total Read and Write byte count | 1 | Counter is incremented by the number of bytes transferred by a matching read or write request as the requests go past the monitor. | | Total Read byte count | 2 | Counter is incremented by the number of bytes transferred by a matching read request as the requests go past the monitor. | | Total Write byte count | 3 | Counter is incremented by the number of bytes transferred by a matching write request as the requests go past the monitor. | | — | 4-127 | Reserved for future standard use. | | — | 128-256 | Designated for custom use. | When the `EVT_ID` for a `MCID` is programmed with a non-zero and legal value, the counter is reset to 0 and starts counting matching events for requests with the matching `MCID` and `AT` (if `ATV` is 1). However, if the `EVT_ID` is configured as 0, the counter stops counting. A controller that does not support monitoring by access-type can hardwire the`ATV` and the `AT` fields to 0, indicating that the counter counts requests with all access-types. When the `bc_mon_ctl` register is written, the controller may need to perform several actions that may not complete synchronously with the write. A write to the `bc_mon_ctl` register sets the read-only `BUSY` bit to 1, indicating that the controller is performing the requested operation. When the `BUSY` bit reads 0, the operation is complete, and the read-only `STATUS` field provides a status value (see [Table 4](#BC%5FMON%5FSTS) for details). Written values to the `BUSY` and the `STATUS`fields are ignored. An implementation that can complete the operation synchronously with the write may hardwire the `BUSY` bit to 0\. The state of the`BUSY` bit, when not hardwired to 0, shall only change in response to a write to the register. The `STATUS` field remains valid until a subsequent write to the`bc_mon_ctl` register. __Table 4\. bc\_mon\_ctl.STATUS Field Encodings__ | STATUS | Description | | ------ | -------------------------------------------------- | | 0 | Reserved | | 1 | The operation was successfully completed. | | 2 | An invalid operation (OP) was requested. | | 3 | An operation was requested for an invalid MCID. | | 4 | An operation was requested for an invalid EVT\_ID. | | 5 | An operation was requested for an invalid AT. | | 6-63 | Reserved for future standard use. | | 64-127 | Designated for custom use. | When the `BUSY` bit is set to 1, the behavior of writes to the `bc_mon_ctl` is`UNSPECIFIED`. Some implementations may ignore the second write, while others may perform the operation determined by the second write. To ensure proper operation, software must first verify that the `BUSY` bit is 0 before writing the `bc_mon_ctl` register. ### [](#BC%5FMCTR)4.1.3\. Bandwidth Monitoring Counter Value The `bc_mon_ctr_val` is a read-only register that holds a snapshot of the counter selected by a `READ_COUNTER` operation. When the controller does not support bandwidth usage monitoring, the `bc_mon_ctr_val` register always reads as zero. ![Bandwidth Monitoring Counter Value Register (`bc_mon_ctr_val`)](_images/diag-ec48b38e971c869e86b7a87095690d44bef9849c.svg) Figure 3\. Bandwidth Monitoring Counter Value Register (`bc_mon_ctr_val`) The counter is valid if the `INV` field is 0\. The counter may be marked`INV` if, for `UNSPECIFIED` reasons, the controller determines the count to be not valid. Such counters may become valid in the future. Additionally, if an unsigned integer overflow of the counter occurs, then the `OVF` bit is set to 1. | | A counter may be marked as INV if the controller has not been able to establish an accurate counter value for the monitored event. | | ------------------------------------------------------------------------------------------------------------------------------------- | The counter provides the number of bytes transferred by requests matching the`EVT_ID` as they go past the monitoring point. A bandwidth value may be determined by reading the byte count value at two instances of time `T1` and`T2`. If the value of the counter at time `T1` was `B1`, and at time `T2` is`B2`, then the bandwidth can be calculated using [equation (1)](#eq-3). The frequency of the time source is represented by . The width of the counter is `UNSPECIFIED` but the unimplemented bits must be read-only zero. | | While the width of the counter is UNSPECIFIED, it is recommended to be wide enough to prevent more than one overflow per sample when the sampling frequency is 1 Hz. If an overflow was detected then software may discard that sample and reset the counter and overflow indication by reprogramming the event using CONFIG\_EVENToperation. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#BC%5FALLOC)4.1.4\. Bandwidth Allocation Control The `bc_alloc_ctl` register is used to control the allocation of bandwidth to an`RCID` per `AT`. If a controller does not support bandwidth allocation, then the register is read-only zero. If the controller does not support bandwidth allocation per access-type, then the `AT` field is read-only zero. ![Bandwidth Allocation Control Register (`bc_alloc_ctl`)](_images/diag-1c4adea840207d816fcbe59d0dc7abfe92f9c651.svg) Figure 4\. Bandwidth Allocation Control Register (`bc_alloc_ctl`) The `OP` field instructs the bandwidth controller to perform an operation listed in [Table 5](#BC%5FALLOC%5FOP). The `bc_alloc_ctl` register is used in conjunction with the`bc_bw_alloc` register to perform bandwidth allocation operations. If the requested operation uses the operands configured in `bc_bw_alloc`, software must first program the `bc_bw_alloc` register with the operands for the operation before requesting the operation. __Table 5\. Bandwidth Allocation Operations (OP)__ | Operation | Encoding | Description | | ------------- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | — | 0 | Reserved for future standard use. | | CONFIG\_LIMIT | 1 | Establishes reserved bandwidth allocation for requests by RCID and of access-type AT. The bandwidth allocation is specified in bc\_bw\_alloc. | | READ\_LIMIT | 2 | Reads back the previously configured bandwidth allocation for requests by RCID and of access-type AT. The current configured allocation is written to bc\_bw\_alloc on completion of the operation. | | — | 3-23 | Reserved for future standard use. | | — | 24-31 | Designated for custom use. | A bandwidth allocation must be configured for each access-type supported by the controller. When differentiated bandwidth allocation based on access-type is not required, one of the access-types may be designated to hold a default bandwidth allocation, and the other access-types can be configured to share the allocation with the default access-type. If bandwidth is not allocated for each access-type supported by the controller, the behavior is `UNSPECIFIED`. When the `bc_alloc_ctl` register is written, the controller may need to perform several actions that may not complete synchronously with the write. A write to the `bc_alloc_ctl` sets the read-only `BUSY` bit to 1 indicating the controller is performing the requested operation. When the `BUSY` bit reads 0, the operation is complete, and the read-only `STATUS` field provides a status value (see[Table 6](#BC%5FALLOC%5FSTS) for details). Written values to the `BUSY` and the `STATUS`fields are ignored. An implementation that can complete the operation synchronously with the write may hardwire the `BUSY` bit to 0\. The state of the`BUSY` bit, when not hardwired to 0, shall only change in response to a write to the register. The `STATUS` field remains valid until a subsequent write to the`bc_alloc_ctl` register. __Table 6\. bc\_alloc\_ctl.STATUS Field Encodings__ | STATUS | Description | | ------ | ----------------------------------------------------------------- | | 0 | Reserved | | 1 | The operation was successfully completed. | | 2 | An invalid operation (OP) was requested. | | 3 | An operation was requested for an invalid RCID. | | 4 | An operation was requested for an invalid AT. | | 5 | An invalid or unsupported reserved bandwidth block was requested. | | 6-63 | Reserved for future standard use. | | 64-127 | Designated for custom use. | ### [](#BC%5FBMASK)4.1.5\. Bandwidth Allocation Configuration The `bc_bw_alloc` is used to program reserved bandwidth blocks (`Rbwb`) for an`RCID` for requests of access-type `AT` using the `CONFIG_LIMIT` operation. If a controller does not support bandwidth allocation, then the `bc_bw_alloc` register is read-only zero. The `bc_bw_alloc` holds the previously configured reserved bandwidth blocks for an `RCID` and `AT` on successful completion of the `READ_LIMIT` operation. Bandwidth is allocated in multiples of bandwidth blocks, and the value in `Rbwb`must be at least 1 and must not exceed `MRBWB` value. Otherwise, the `CONFIG_LIMIT`operation fails with `STATUS=5`. Additionally, the sum of `Rbwb` allocated across all `RCIDs` must not exceed `MRBWB` value, or the `CONFIG_LIMIT` operation fails with `STATUS=5`. ![Bandwidth Allocation Configuration Register (`bc_bw_alloc`)](_images/diag-61cb4dd05825edaa1f75b64ac51838378fed6b8a.svg) Figure 5\. Bandwidth Allocation Configuration Register (`bc_bw_alloc`) The `Rbwb`, `Mweight`, `sharedAT`, and `useShared` are all WARL fields. Bandwidth allocation is typically enforced by the bandwidth controller over finite accounting windows. The process involves measuring the bandwidth consumption over an accounting window and determining if an `RCID` is exceeding its bandwidth allocations for each access-types. The specifics of how the accounting window is implemented are `UNSPECIFIED`, but is expected to provide a statistically accurate control of the bandwidth usage over a few accounting intervals. The `Rbwb` field represents the bandwidth that is made available to an `RCID` for requests that match `AT`, even when all other `RCID` are using their full allocation of bandwidth. The bandwidth allocation scales linearly with the number of bandwidth blocks programmed into `Rbwb`. If there is non-reserved or unused bandwidth available in an accounting interval, `RCIDs` may compete for additional bandwidth. The non-reserved or unused bandwidth is proportionately shared among the competing `RCIDs` by using the configured `Mweight` parameter, which is a number between 0 and 255\. A larger weight implies a greater fraction of the bandwidth. A weight of 0 implies that the configured limit is a hard limit, and the use of unused or non-reserved bandwidth is not allowed. Sharing of non-reserved bandwidth is not differentiated by access-type identifier. Therefore, the `Mweight` parameter must be programmed identically for all access-type identifiers. If this parameter is programmed differently for each access-type identifier, then the controller can use the parameter configured for any of the identifiers, but the behavior is otherwise well defined. When the `Mweight` parameter is not set to 0, the amount of unused bandwidth allocated to `RCID=x` during contention with another `RCID` that is also permitted to use unused bandwidth is determined by dividing the `Mweight` of`RCID=x` by the sum of the `Mweight` of all other contending `RCIDs`. This ratio `P` is determined by [equation (2)](#eq-4). | | The bandwidth enforcement is typically work-conserving, allowing unused bandwidth to be used by requestors that are enabled to use it, even if they have consumed their Rbwb allotment. When contending for unused bandwidth, the weighted share is typically computed among the RCIDs that are actively generating requests in that accounting interval and have a non-zero weight programmed. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If unique bandwidth allocation is not required for an access-type identifier, then the`useShared` parameter can be set to 1 for a `CONFIG_LIMIT` operation. When`useShared` is set to 1, the `sharedAT` field specifies the access-type identifer with which the bandwidth allocation is shared by the access-type identifier in`bc_alloc_ctl.AT`. In this case, the `Rbwb` and `Mweight` fields are ignored, and the configurations of the access-type identifier in `sharedAT` are applied. If the access-type identifier specified by `sharedAT` does not have unique bandwidth allocation, meaning that it has not been configured with `useShared=0`, then the behavior is `UNSPECIFIED`. The `useShared` and `sharedAT` fields are read-only zero if the bandwidth controller does not support bandwidth allocation per access-type identifier. | | When unique bandwidth allocation for an access-type identifier is not required, then one or more identifiers might be configured with a shared bandwidth allocation. For example, consider a bandwidth controller that supports 3 access-type identifers. The access-type identifier 0 and 1 of RCID 3 are configured with unique bandwidth allocations and the access-type identifier 2 is configured to share bandwidth allocation with identifier 1\. The example configuration is illustrated in the following table: Rbwb Mweight useShared sharedAT RCID=3, AT=0 100 16 0 N/A RCID=3, AT=1 50 16 0 N/A RCID=3, AT=2 N/A N/A 1 1 | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Bibliography ==================== ## [](#bibliography)Bibliography \[1\] D. B. Kristof and E. Stijn and E. Lieven, "Per-Thread Cycle Accounting in Multicore Processors", _ACM Trans. Archit. Code Optim._, vol. 9, no. 4, jan 2013\. \[Online\]. Available: . \[2\] L. David and C. Liqun and G. Rama and R. Parthasarathy and K. Christos, "Heracles: Improving Resource Efficiency at Scale" in _Proceedings of the 42nd Annual International Symposium on Computer Architecture_, ISCA '15\. New York, NY, USA:, Association for Computing Machinery, 2015, pp. 450–462, Available: . \[3\] _RISC-V Quality-of-Service (QoS) Identifiers_. \[Online\]. Available: \[4\] _RISC-V State Enable Extension_. \[Online\]. Available: \[5\] _RISC-V IOMMU Architecture Specification_. \[Online\]. Available: \[6\] T. N.C. and F. J.K., "Facilitating level three cache studies using set sampling" in _2000 Winter Simulation Conference Proceedings (Cat. No.00CH37165)_, vol. 1, no. . 2000, pp. 471-479 vol.1. 3.1. Capacity-controller QoS Register Interface ==================== ## [](#CC%5FQOS)3.1\. Capacity-controller QoS Register Interface Controllers, such as cache controllers, that support capacity allocation and usage monitoring provide a memory-mapped capacity-controller QoS register interface. The capacity controller allocates capacity in fixed multiples of _capacity units_. A group of these _capacity units_ is referred to as a _capacity block_. One or more _capacity blocks_ can be allocated to a workload. When a workload requests capacity allocation, the capacity is allocated by using _capacity units_situated within the _capacity blocks_ assigned to the workload. Capacity blocks can also be shared among one or more workloads. Optionally, the capacity controller might allow configuration of a limit on the maximum number of _capacity units_ that can be occupied in the _capacity blocks_ allocated to a specific workload. | | For example, a cache controller allocates capacity in multiples of cache blocks. In this context, a cache block serves as a _capacity unit_, and a group of cache blocks forms a _capacity block_. A cache controller supporting capacity allocation _by ways_ might define a _capacity block_ to be the cache blocks in one way of the cache. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The capacity allocation affects the decision regarding which _capacity blocks_to use when a new _capacity unit_ is requested by a workload, but usually does not affect other operations of the controller. | | For example, when a request is made to a cache controller, the request involves scanning the entire cache to determine if the requested data is present. If the data is located, then the request is fulfilled with this data, even if the cache block containing the data was initially allocated for a different workload. The data continues to reside in the same cache block. Consequently, the cache lookup function remains unaffected by the capacity allocation constraints set for the workload that initiated the request. Conversely, if the data is not found, a new cache block must be allocated. This allocation is executed by using the _capacity blocks_ assigned to the workload that made the request. Hence, a workload might only trigger evictions within _capacity blocks_ designated to it, but can access shared data in _capacity blocks_ allocated to other workloads. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | __Table 1\. Capacity-controller QoS Register Layout (size and offset are in bytes)__ | Offset | Name | Size | Description | Optional? | | ------ | ----------------- | ----- | ------------------------------------------- | --------- | | 0 | cc\_capabilities | 8 | [Capabilities ](#CC%5FCAP) | No | | 8 | cc\_mon\_ctl | 8 | [Usage monitoring control](#CC%5FMCTL) | Yes | | 16 | cc\_mon\_ctr\_val | 8 | [Monitoring counter value](#CC%5FMCTR) | Yes | | 24 | cc\_alloc\_ctl | 8 | [Capacity allocation control ](#CC%5FALLOC) | Yes | | 32 | cc\_block\_mask | BMW/8 | [Capacity block mask ](#CC%5FBMASK) | Yes | | N | cc\_cunits | 8 | [Capacity units count](#CC%5FCUNITS) | Yes | The size of the `cc_block_mask` register is determined by the `NCBLKS` field of the `cc_capabilities` register but is always a multiple of 8 bytes. The formula for determination of `BMW` is defined in [3.1.5\. Capacity Block Mask (cc\_block\_mask)](#CC%5FBMASK). The offset `N` is determined as `32 + BMW/8`. The reset value is 0 for the following register fields. * `cc_mon_ctl.BUSY` field * `cc_alloc_ctl.BUSY` field The reset value is `UNSPECIFIED` for all other registers fields. The capacity controllers at reset must allocate all available capacity to `RCID`value of 0\. When the capacity controller supports capacity allocation per access-type, then all available capacity is shared by all the access-type for`RCID=0`. The capacity allocation for all other `RCID` values is `UNSPECIFIED`. The capacity controller behavior for handling a request with a non-zero `RCID`value before configuring the capacity controller with capacity allocation for that `RCID` is `UNSPECIFIED`. ### [](#CC%5FCAP)3.1.1\. Capacity-controller Capabilities The `cc_capabilities` register is a read-only register that holds the capacity-controller capabilities. ![Capacity-controller Capabilities Register](_images/diag-ab2459cfbc6d51417953349d85dff96813659f68.svg) Figure 1\. Capacity-controller Capabilities Register The `VER` field holds the version of the specification implemented by the capacity controller. The low nibble is used to hold the minor version of the specification and the upper nibble is used to hold the major version of the specification. For example, an implementation that supports version 1.0 of the specification reports 0x10. The `NCBLKS` field holds the total number of allocatable _capacity blocks_ in the controller. The capacity represented by an allocatable _capacity block_ is`UNSPECIFIED`. The capacity controllers support allocating capacity in fixed multiples of an allocatable _capacity block_. | | For example, a cache controller that defines a way of the cache as a _capacity block_ may report the number of ways as the number of allocatable _capacity blocks_. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If `CUNITS` is 1, the controller supports specifying a limit on the _capacity units_ that can be occupied by an `RCID` in _capacity blocks_ allocated to it. If `FRCID` is 1, the controller supports an operation to flush and deallocate the _capacity blocks_ occupied by an `RCID`. When the `RCID`\-prefixed mode (`RPFX`) is 1, the controller operates in `RPFX`mode. The parameter `P` (prefixed bits) indicates the number of least significant bits of the `MCID` carried in the request that should be prefixed by the RCID. The controller uses `RCID` in the requests along with `P` number of least significant bits of the `MCID` to compute an effective `MCID` ([\[EMCID\]](#EMCID)). This `MCID` is used to identify the monitoring counter. Legal values of `P` range from 0 to 12. If `RPFX` is 0, `P` is set to 0, and the effective `MCID` is the same as the `MCID`in the request. ### [](#CC%5FMCTL)3.1.2\. Capacity Usage Monitoring Control The `cc_mon_ctl` register is used to control monitoring of capacity usage by a`MCID`. When the controller does not support capacity usage monitoring the`cc_mon_ctl` register is read-only zero. ![Capacity Usage Monitoring Control Register (`cc_mon_ctl`)](_images/diag-162d092e66bc48cb21fbd7a767a4dfd60241b10e.svg) Figure 2\. Capacity Usage Monitoring Control Register (`cc_mon_ctl`) Capacity controllers that support capacity usage monitoring implement a usage monitoring counter for each supported `MCID`. The usage monitoring counter can be configured to count a monitoring event. When an event matching the event configured for the `MCID` occurs, then the monitoring counter is updated. The event matching might optionally be filtered by the access-type identifier. The `OP`, `AT`, `ATV`, `MCID`, and `EVT_ID` fields of the register are WARL fields. The `OP` field is used to instruct the controller to perform an operation listed in [Table 2](#CC%5FMON%5FOP). __Table 2\. Capacity Usage Monitoring Operations (OP)__ | Operation | Encoding | Description | | ------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | | — | 0 | Reserved for future standard use. | | CONFIG\_EVENT | 1 | Configure the counter selected by MCID to count the event selected by EVT\_ID, AT, and ATV. The EVT\_ID encodings are listed in [Table 3](#CC%5FEVT%5FID). | | READ\_COUNTER | 2 | Snapshot the value of the counter selected byMCID into cc\_mon\_ctr\_val register. TheEVT\_ID, AT, and ATV fields are not used by this operation. | | — | 3-23 | Reserved for future standard use. | | — | 24-31 | Designated for custom use. | The `EVT_ID` field is used to program the identifier of the event to count in the monitoring counter selected by `MCID`. The `AT` field (See [\[AT\_ENC\]](#AT%5FENC)) is used to program the access-type identifier to count, and its validity is indicated by the`ATV` field. When `ATV` is 0, the counter counts requests with all access-type identifiers, and the `AT` value is ignored. __Table 3\. Capacity Usage Monitoring Event ID (EVT\_ID)__ | Event ID | Encoding | Description | | --------- | -------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | None | 0 | Counter does not count. | | Occupancy | 1 | Counter is incremented by 1 when a request with a matching MCID and AT allocates a unit of capacity. The counter is decremented by 1 when a unit of capacity is de-allocated. | | — | 2-127 | Reserved for future standard use. | | — | 128-256 | Designated for custom use. | When the `EVT_ID` for a `MCID` is programmed with a non-zero and legal value by using the `CONFIG_EVENT` operation, the counter is reset to 0 and starts counting matching events for requests with the matching `MCID` and `AT` (if `ATV` is 1). However, if the `EVT_ID` is programmed to 0, the counter stops counting. A controller that does not support monitoring by access-type identifier can hardwire the`ATV` and the `AT` fields to 0, indicating that the counter counts requests with all access-types identifiers. When the `cc_mon_ctl` register is written, the controller can perform several actions that might not complete synchronously with the write. A write to the `cc_mon_ctl` sets the read-only `BUSY` bit to 1, indicating the controller is performing the requested operation. When the `BUSY` bit reads 0, the operation is complete, and the read-only `STATUS` field provides a status value (see[Table 4](#CC%5FMON%5FSTS) for details). Written values to the `BUSY` and the `STATUS`fields are ignored. An implementation that can complete the operation synchronously with the write may hardwire the `BUSY` bit to 0\. The state of the`BUSY` bit, when not hardwired to 0, shall only change in response to a write to the register. The `STATUS` field remains valid until a subsequent write to the`cc_mon_ctl` register. __Table 4\. cc\_mon\_ctl.STATUS Field Encodings__ | STATUS | Description | | ------ | -------------------------------------------------- | | 0 | Reserved | | 1 | The operation was successfully completed. | | 2 | An invalid operation (OP) was requested. | | 3 | An operation was requested for an invalid MCID. | | 4 | An operation was requested for an invalid EVT\_ID. | | 5 | An operation was requested for an invalid AT. | | 6-63 | Reserved for future standard use. | | 64-127 | Designated for custom use. | When the `BUSY` bit is set to 1, the behavior of writes to the `cc_mon_ctl` is`UNSPECIFIED`. Some implementations ignore the second write, while others might perform the operation determined by the second write. To ensure proper operation, software must first verify that the `BUSY` bit is 0 before writing the `cc_mon_ctl` register. ### [](#CC%5FMCTR)3.1.3\. Capacity Usage Monitoring Counter Value The `cc_mon_ctr_val` is a read-only register that holds a snapshot of the counter that is selected by the `READ_COUNTER` operation. When the controller does not support capacity usage monitoring, the `cc_mon_ctr_val` register always reads as zero. ![Capacity Usage Monitoring Counter Value Register (`cc_mon_ctr_val`)](_images/diag-0f774051a3dcdd794ba4449388a8cd44a92bcb87.svg) Figure 3\. Capacity Usage Monitoring Counter Value Register (`cc_mon_ctr_val`) The counter is valid if the `INV` field is 0\. The counter is marked `INV` if the controller determines the count to be not valid for `UNSPECIFIED` reasons. The counters marked `INV` can become valid in future. The counter shall not decrement below zero. If an event occur that would otherwise result in a negative value, the counter continues to hold a value of 0. | | Following a reset of the counter to zero, a capacity de-allocation attempts to drive its value below zero. This scenario occurs when the MCID is reassigned to a new workload, yet the capacity controller continues to hold capacity that was initially allocated by the previous workload. In such cases, the counter shall not decrement below zero and shall remain at zero. After a brief period of execution for the new workload post-counter reset, the counter value is expected to stabilize to reflect the capacity usage of this new workload. Some implementations might not store the MCID of the request that caused the capacity to be allocated with every unit of capacity in the controller to optimize for the storage overheads. Such controllers, in turn, rely on statistical sampling to report the capacity usage by tagging only a subset of the capacity units. Set-sampling is a technique commonly used in caches to estimate the cache occupancy with a relatively small sample size. The basic idea behind set-sampling is to select a subset of the cache sets and monitor only those sets. By keeping track of the hits and misses in the monitored sets, it is possible to estimate the overall cache occupancy with a high degree of accuracy. The size of the subset needed to obtain accurate estimates depends on various factors, such as the size of the cache, the cache access patterns, and the desired accuracy level. Research \[[6](qos%5Fbiblio.html#bib-ssample)\] shows that set-sampling can provide statistically accurate estimates with a relatively small sample size, such as 10% or less, depending on the cache properties and sampling technique used. When the controller has not observed enough samples to provide an accurate value in the monitoring counter, it might report the counter as being INVuntil more accurate measurements are available. This state helps to prevent inaccurate or misleading data from being used in capacity planning or other decision-making processes. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#CC%5FALLOC)3.1.4\. Capacity Allocation Control The `cc_alloc_ctl` register is used to configure allocation of capacity to an`RCID` per access type (`AT`). The `OP`, `RCID` and `AT` fields in this register are WARL. If a controller does not support capacity allocation, then this register is read-only zero. If the controller does not support capacity allocation per access type, then the `AT` field is read-only zero. ![Capacity Allocation Control Register (`cc_alloc_ctl`)](_images/diag-1c4adea840207d816fcbe59d0dc7abfe92f9c651.svg) Figure 4\. Capacity Allocation Control Register (`cc_alloc_ctl`) The `OP` field is used to instruct the capacity controller to perform an operation listed in [Table 5](#CC%5FALLOC%5FOP). Some operations necessitate the specification of the _capacity blocks_ to act upon. For such operations, the targeted _capacity blocks_ are designated in the form of a bitmask in the`cc_block_mask` register. Additionally, certain operations require the _capacity unit_ limit to be defined in the `cc_cunits` register. To execute operations that require a capacity block mask and/or a capacity unit limit, software must first program the `cc_block_mask` and/or the `cc_cunits` register, followed by initiating the operation with the `cc_alloc_ctl` register. __Table 5\. Capacity Allocation Operations (OP)__ | Operation | Encoding | Description | | ------------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | — | 0 | Reserved for future standard use. | | CONFIG\_LIMIT | 1 | Configure a capacity allocation for requests byRCID and of access type AT. The _capacity blocks_ allocation is specified in thecc\_block\_mask register, and a limit on capacity units is specified in the cc\_cunits register. | | READ\_LIMIT | 2 | Read back the previously configured capacity allocation for requests by RCID and of access-type AT. The configured _capacity block_ allocation is returned as a bit-mask in thecc\_block\_mask register, and the configured limit on _capacity units_ is available in the cc\_cunits register on successful completion of the operation. | | FLUSH\_RCID | 3 | Flushes the _capacity units_ used by the specifiedRCID and access-type AT. This operation is supported if the capabilities.FRCID bit is 1\. The cc\_block\_mask and cc\_cunits registers are not used for this operation. The configured _capacity block_ allocation or the_capacity unit_ limit is not changed by this operation. | | — | 4-23 | Reserved for future standard use. | | — | 24-31 | Designated for custom use. | Capacity controllers enumerate the allocatable _capacity blocks_ in the `NCBLKS`field of the `cc_capabilities` register. The `cc_block_mask` register is programmed with a bit-mask value, where each bit represents a _capacity block_ for the operation. If configuring _capacity unit_ limits is supported (for example,`cc_capabilities.CUNIT=1`), then a limit on the _capacity unit_ that can be occupied in the allocated capacity blocks can be programmed in the `cc_cunits`register. If configuring limits is not supported, then the controller allows the use of all _capacity units_ in the allocated _capacity blocks_. A value of zero programmed into `cc_cunits` indicates that no limits shall be enforced on_capacity unit_ allocation. A capacity allocation must be configured for each supported access type by the controller. An implementation that does not support capacity allocation per access type can hardwire the `AT` field to 0 and associate the same capacity allocation configuration for requests with all access types. When capacity allocation per access type is supported, identical limits can be configured for two or more access types, if different capacity allocation per access type is not required. If capacity is not allocated for each access type supported by the controller, the behavior is `UNSPECIFIED`. | | A cache controller that supports capacity allocation indicates the number of allocatable _capacity blocks_ in cc\_capabilities.NCBLKS field. For example, consider a cache with NCBLKS=8. In this example, the RCID=5 is allocated _capacity blocks_ numbered 0 and 1 for requests with access type AT=0, and _capacity blocks_ numbered 2 for requests with access typeAT=1. The RCID=3 in this example is allocated _capacity blocks_numbered 3 and 4 for both AT=0 and AT=1 access types as separate capacity allocation by access type is not required for this workload. Further in this example, the RCID=6 has been configured with the same _capacity block_allocations as RCID=3. This configuration implies that they share a common capacity allocation in this cache, but might be associated with different RCID to allow differentiated treatment in another capacity and/or bandwidth controller. 7 6 5 4 3 2 1 0 RCID=3, AT=0 0 0 0 1 1 0 0 0 RCID=3, AT=1 0 0 0 1 1 0 0 0 RCID=5, AT=0 0 0 0 0 0 0 1 1 RCID=5, AT=1 0 0 0 0 0 1 0 0 RCID=6, AT=0 0 0 0 1 1 0 0 0 RCID=6, AT=1 0 0 0 1 1 0 0 0 Some controllers allow setting a limit on _capacity units_ in allocated capacity blocks. In exclusive allocations, like for RCID=5, the limit can be the capacity block’s maximum capacity. For shared allocations, such as betweenRCID=3 and RCID=6, individual limits can be set. For example, if two capacity blocks represent 100 units and RCID=3 has a 30-unit limit whileRCID=6 has a 70-unit limit, they can use 30% and 70% of the shared capacity blocks, respectively. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `FLUSH_RCID` operation can incur a long latency to complete. However, the`RCID` can submit new requests to the controller while it is being flushed. Additionally, the controller is allowed to deallocate capacity that was allocated after the operation was initiated. | | For cache controllers, the FLUSH\_RCID operation perfoms an operation similar to that performed by the RISC-V CBO.FLUSH instruction on each cache block that is part of the allocation configured for the RCID. The FLUSH\_RCID operation can be used as part of reclaiming a previously allocated RCID and associating it with a new workload. When such a reallocation is performed, the capacity controllers might have capacity allocated by the old workload and thus for a short warm-up duration, the capacity controller might be enforcing capacity allocation limits that reflect the usage by the old workload. Such warm-up durations are typically not statistically significant, but if that is not desired, then the FLUSH\_RCID operation can be used to flush and evict capacity allocated by the old workload. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When the `cc_alloc_ctl` register is written, the controller might perform several actions that might not complete synchronously with the write. A write to the `cc_alloc_ctl` sets the read-only `BUSY` bit to 1, indicating the controller is performing the requested operation. When the `BUSY` bit reads 0, the operation is complete, and the read-only `STATUS` field provides a status value ([Table 6](#CC%5FALLOC%5FSTS)) of the requested operation. Values that are written to the `BUSY` and the `STATUS` fields are always ignored. An implementation that can complete the operation synchronously with the write might hardwire the `BUSY` bit to 0\. The state of the `BUSY` bit, when not hardwired to 0, shall change only in response to a write to the register. The `STATUS` field remains valid until a subsequent write to the `cc_alloc_ctl` register. __Table 6\. cc\_alloc\_ctl.STATUS Field Encodings__ | STATUS | Description | | ------ | --------------------------------------------------- | | 0 | Reserved | | 1 | The operation was successfully completed. | | 2 | An invalid or unsupported operation (OP) requested. | | 3 | An operation was requested for an invalid RCID. | | 4 | An operation was requested for an invalid AT. | | 5 | An invalid _capacity block_ mask was specified. | | 6-63 | Reserved for future standard use. | | 64-127 | Designated for custom use. | When the `BUSY` bit is set to 1, the behavior of writes to the `cc_alloc_ctl`register, `cc_cunits` register, or to the `cc_block_mask` register is`UNSPECIFIED`. Some implementations might ignore the second write and others might perform the operation determined by the second write. To ensure proper operation, software must verify that `BUSY` bit is 0 before writing any of these registers. ### [](#CC%5FBMASK)3.1.5\. Capacity Block Mask (`cc_block_mask`) The `cc_block_mask` is a WARL register. If the controller does not support capacity allocation, for example, `NCBLKS` is 0, then this register is read-only 0. The register has `NCBLKS` bits, each corresponding to one allocatable_capacity block_ in the controller. The width of this register is variable, but always a multiple of 64 bits. The bitmap width in bits (`BMW`) is determined by the following equation. The division operation in this equation is an integer division. Bits `NCBLKS-1:0` are read-write, and the bits `BMW-1:NCBLKS` are read-only zero. The process of configuring capacity allocation for an `RCID` and `AT` begins by programming the `cc_block_mask` register with a bit-mask value that identifies the_capacity blocks_ to be allocated and, if supported, by programming the`cc_cunits` register with a limit on the capacity units that might be occupied in those capacity blocks. Next, the `cc_alloc_ctl register` is written to request a`CONFIG_LIMIT` operation for the `RCID` and `AT`. After a capacity allocation limit is established, a request can be allocated capacity in the _capacity blocks_ allocated to the `RCID` and `AT` associated with the request. It is important to note that some implementations might require at least one _capacity block_ to be allocated by using `cc_block_mask` when allocating capacity; otherwise, the operation fails with `STATUS=5`. Overlapping _capacity block_masks among `RCID` and/or `AT` are allowed to be configured. | | A multiway set-associative cache controller that supports capacity allocation _by ways_ can advertise NCBLKS as the number of ways per set in the cache. To allocate capacity in such a cache for an RCID and AT, a subset of ways must be selected and a mask of the selected ways must be programmed in cc\_block\_mask field when the CONFIG\_LIMIT operation is requested. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | To read the _capacity block_ allocation for an `RCID` and `AT`, the controller provides the `READ_LIMIT` operation, which can be requested by writing to the`cc_alloc_ctl` register. When the operation completes successfully, the`cc_block_mask` register holds the configured _capacity block_ allocation. ### [](#CC%5FCUNITS)3.1.6\. Capacity Units The `cc_cunits` register is a read-write WARL register. If the controller does not support capacity allocation (for example, `NCBLKS` is set to 0), this register shall be read-only zero. If the controller does not support configuring limits on _capacity units_ that may be occupied in the allocated _capacity blocks_ (for example,`cc_capabilities.CUNITS=0`), then this register shall be read-only zero. In such cases, the controller allows the utilization of all available _capacity units_ by an `RCID` within the _capacity blocks_ allocated to it. If the controller supports configuring limits on _capacity units_ that might be occupied in the allocated _capacity blocks_ (for example, `cc_capabilities.CUNITS=1`) then this register sets an upper limit on the number of _capacity units_ that can be occupied by an `RCID` in the _capacity blocks_ allocated for an `AT`. A value of zero specified in the `cc_cunits` register indicates that no limit is configured. The sum of the `cc_cunits` configured for the `RCID` sharing a _capacity block_allocation may exceed the _capacity units_ represented by that _capacity block_allocation. | | When multiple RCID instances share a _capacity block_ allocation, thecc\_cunits register can be employed to set an upper limit on the number of_capacity units_ each RCID can occupy. For instance, consider a group of four RCID instances configured to share a set of _capacity blocks_, representing a total of 100 capacity units. EachRCID can be configured with a limit of 30 capacity units, ensuring that no individual RCID exceeds 30% of the total shared _capacity units_. The capacity controller might enforce these limits through various techniques. Examples include: Refraining from allocating new capacity units to an RCID that reached its limit. Evicting previously allocated capacity units when a new allocation is required. These methods are not exhaustive and can be applied either individually or in combination to maintain _capacity unit_ limits. When the limit on the _capacity units_ is reached or is about to be reached, the capacity controller can initiate additional operations. These could include throttling certain activities (for example, prefetches) of the corresponding workload requests. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | To read the _capacity unit_ limit for an `RCID` and `AT`, the controller provides the `READ_LIMIT` operation that can be requested by writing to the`cc_alloc_ctl` register. When the operation completes successfully, the`cc_cunits` register holds the configured _capacity unit_ allocation limit. Contributors ==================== ## [](#contributors)Contributors This RISC-V specification has been contributed to directly or indirectly by: Aaron Durbin, Adrien Ricciardi, Allen Baum, Allison Randal, Ambika Krishnamoorthy, Beeman Strong, Daniel Gracia Pérez, Roy Franz, David Kruckemeyer, David Weaver, Derek Hower, Drew Fustini, Eric Shiu, Greg Favor, Perrine Peresse, Ravi Sahita, Ved Shanbhogue 6.1. Hardware Guidelines ==================== ## [](#QOS%5FHW%5FGUIDE)6.1\. Hardware Guidelines ### [](#QOS%5FSIZING)6.1.1\. Sizing QoS Identifiers In a typical implementation, the number of `RCID` bits implemented (for example, to support 10s of `RCIDs`) might be smaller than the number of `MCID` bits implemented (for example, to support 100s of `MCIDs`). It is a typical usage to associate a group of applications/VMs with a common`RCID` and thus sharing a common pool of resource allocations. The resource allocations for the `RCID` is established to meet the SLA objectives of all members of the group. If SLA objectives of one or more members of the group stop being met, the resource usage of one or more members of the group might be monitored by associating them with a unique `MCID` and this iterative analysis process used to determine the optimal strategy - increasing resources allocated to the `RCID`, moving some members to a different `RCID`, migrating some members away to another machine, and so on - for restoring the SLA. Having a sufficiently large pool of `MCID` speeds up this analysis. | | To maximize flexibility in the allocation of QoS IDs to workloads, it is recommended that all resource controllers in the system support an identical number of RCID and MCID, as well as a uniform mode of operation — either direct or RCID-prefixed — for determining the effective MCID. Uniformity ensures that software is not constrained by the lowest common denominator of ID support when requests are processed by multiple controllers, such as caches, fabrics, and memory controllers. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#6-1-2-sizing-monitoring-counters)6.1.2\. Sizing Monitoring Counters Typically software samples the monitoring counters periodically to monitor capacity and bandwidth usage. The width of the monitoring counters is recommended to be wide enough to not cause more than one overflow per sample when sampled at a frequency of 1 Hz. 2.1. QoS Identifiers ==================== ## [](#QOS%5FID)2.1\. QoS Identifiers Monitoring or allocation of resources requires a way to identify the originator of the request to access the resource. CBQRI and the Ssqosid extension provides a mechanism by which a workload can be associated with a resource control ID (`RCID`) and a monitoring counter ID (`MCID`) that accompany each request made by the workload to shared resources. To provide differentiated services to workloads, CBQRI defines a mechanism to configure resource usage limits, in the form of capacity or bandwidth, per supported access type, for an `RCID` in the resource controllers that control accesses to such shared resources. To monitor the resource utilization by a workload, CBQRI defines a mechanism to configure counters identified by the `MCID` to count events in the resource controllers that control accesses to such shared resources. [\[QOS\_SIZING\]](#QOS%5FSIZING) discusses guidelines for sizing the QoS IDs and the need for differentiated IDs for monitoring. All supported `RCID` and `MCID` can be actively used in the system at any instance. ### [](#EMCID)2.1.1\. Associating `RCID` and `MCID` with requests The `RCID` in the request is used by the resource controllers to determine the resource allocations (for example, cache occupancy limits, memory bandwidth limits, and so on) to enforce. The `MCID` in the request is used by the resource controllers to identify the ID of a counter to monitor resource usage (for example, cache occupancy, memory bandwidth, and so on). Two modes of operation are supported by CBQRI. In the direct mode, the`MCID` carried with the request is directly used by the controller to identify the counter and is the effective `MCID`. In the RCID-prefixed mode, the controller identifies the counter for monitoring using an effective `MCID`computed as: . Legal values of `P` range from 0 to 12 and are enumerated in the capability register of the controller. Software should use the effective `MCID` as the`MCID` operand to the controller for operations on the monitoring counters. #### [](#2-1-1-1-risc-v-hart-initiated-requests-ssqosid)2.1.1.1\. RISC-V hart initiated requests (Ssqosid) The Ssqosid extension \[[3](qos%5Fbiblio.html#bib-ssqosid)\] introduces a read/write S/HS-mode register (`srmcfg`) to configure QoS Identifiers to be used with requests made by the hart to shared resources. If Smstateen \[[4](qos%5Fbiblio.html#bib-stateen)\] is implemented then bit 55 of `mstateen0` controls access to `srmcfg` from privilege modes less than M. #### [](#2-1-1-2-device-initiated-requests)2.1.1.2\. Device initiated requests A RISC-V IOMMU \[[5](qos%5Fbiblio.html#bib-iommu)\] extension to support configuring QoS identifiers is specified in [\[QOS\_IOMMU\]](#QOS%5FIOMMU). If the system supports an IOMMU with this extension, the IOMMU can be configured with the `RCID` and `MCID` to associate with requests from devices and from the IOMMU itself. If the system does not support an IOMMU with this extension, then the association of `RCID` and `MCID` with requests from devices becomes implementation-defined. Such methods can include, but are not limited to, one of the following examples: * Devices can be configured with an `RCID` and `MCID` for requests originating from the device, provided the device implementation and the bus protocol used by the device support such capabilities. The method to configure the QoS identifiers into devices remains `UNSPECIFIED`. * Where the device does not natively support being configured with an `RCID`and `MCID`, the implementation might provide a shim at the device interface. This shim can be configured with the `RCID` and `MCID` to associate with requests originating from the device. The method to configure such QoS identifiers into a shim is `UNSPECIFIED`. ### [](#2-1-2-access-type-at)2.1.2\. Access type (`AT`) In some usages, in addition to providing differentiated service among workloads, the ability to differentiate between resource usage for accesses made by the same workload might be required. For example, the capacity allocated in a shared cache for code storage might be differentiated from the capacity allocated for data storage and thereby avoid code from being evicted from such shared cache due to a data access. When differentiation based on access type (for example, code vs. data) is supported the requests also carry an access-type (`AT`) indicator. The resource controllers can be configured with separate capacity and/or bandwidth allocations for each supported access type. CBQRI defines a 3-bit `AT` field, encoded as specified in[Table 1](#AT%5FENC), in the register interface to configure differentiated resource allocation and monitoring for each `AT`. __Table 1\. Encodings of AT field__ | Value | Name | Description | | ----- | -------- | --------------------------------- | | 0 | Data | Requests to access data. | | 1 | Code | Requests for code execution. | | 2-5 | Reserved | Reserved for future standard use. | | 6-7 | Custom | Designated for custom use. | For unsupported `AT` values the resource controller behaves as if `AT` was 0. 1.1. Introduction ==================== ## [](#intro)1.1\. Introduction Quality of Service (QoS) is defined as the minimal end-to-end performance that is guaranteed in advance by a service level agreement (SLA) to a workload. A workload can be a single application, a group of applications, a virtual machine, a group of virtual machines, or a combination of those. The performance is measured in the form of metrics such as instructions per cycle (IPC), latency of servicing work, etc. Various factors such as the available cache capacity, memory bandwidth, interconnect bandwidth, CPU cycles, system memory, and so on affect the performance of a computing system that runs multiple workloads concurrently. Furthermore, during arbitration for shared resources, the prioritization of the workloads' requests against other competing requests might also affect the performance of the workload. Such interference due to resource sharing can lead to unpredictable workload performance \[[1](qos%5Fbiblio.html#bib-ptcamp)\]. When multiple workloads are running concurrently on modern processors with large core counts, multiple cache hierarchies, and multiple memory controllers, the performance of a workload can become less deterministic or even non-deterministic. This situation occurs because the worload performance depends on the behavior of the other workloads in the machine that contend for the shared resources, which can lead to interference. In many deployment scenarios, such as public cloud servers, the workload owner might not be in control of the type and placement of other workloads in the platform. System software can control some of these resources available to the workload, such as the number of hardware threads made available for execution, the amount of system memory allocated to the workload, and the number of CPU cycles provided for execution. System software needs additional tools to control interference to a workload and thereby reduce the variability in the performance that is experienced by one workload due to other workloads' cache capacity usage, memory bandwidth usage, interconnect bandwidth usage, and so on through a resource allocation capability. The resource allocation capability enables system software to reserve capacity and/or bandwidth to meet the performance goals of the workload. Such controls can improve the utilization of the system by co-locating workloads while minimizing the interference caused by one workload to another \[[2](qos%5Fbiblio.html#bib-heracles)\]. Effective use of the resource allocation capability requires the hardware to provide a resource monitoring capability by which the resource requirements of a workload needed to meet a certain performance goal can be characterized. A typical use model involves profiling the resource usage of the workload by using the resource monitoring capability and to establish resource allocations for the workload by using the resource allocation capability. The allocation might be in the form of capacity or bandwidth, depending on the type of resource. For caches, TLBs, and directories, the resource allocation is in the form of storage capacity. For interconnects and memory controllers, the resource allocation is in the form of bandwidth. Workloads generate different types of accesses to shared resources. For example, some of the requests might be for accessing instructions and other requests might be to access data that is operated on by the workload. Certain data accesses might not have temporal locality, whereas others can have a high probability of reuse. In some cases, you can provide differentiated treatment to each type of access by providing unique resource allocation to each access type. RISC-V Capacity and Bandwidth Controller QoS Register Interface (CBQRI) specification specifies the following identifiers and interfaces: 1. QoS identifiers to identify workloads that originate requests to the shared resources. These QoS identifiers include an identifier for resource allocation configurations and an identifier for the monitoring counters used to monitor resource usage. These identifiers accompany each request that is made by the workload to the shared resource. [\[QOS\_ID\]](#QOS%5FID) specifies the mechanism to associate the identifiers with workloads. 2. Access-type identifiers to accompany request to access a shared resource to allow differentiated treatment of each access-type (for example, code vs. data,). The access-types are defined in [\[QOS\_ID\]](#QOS%5FID). 3. Register interface for capacity allocation in controllers, such as shared caches, directories, and so on. The capacity allocation register interface is specified in [\[CC\_QOS\]](#CC%5FQOS). 4. Register interface for capacity usage monitoring. The capacity usage monitoring register interface is specified in [\[CC\_QOS\]](#CC%5FQOS). 5. Register interface for bandwidth allocation in controllers, such as interconnect and memory controllers. The bandwidth allocation register interface is specified in [\[BC\_QOS\]](#BC%5FQOS). 6. Register interface for bandwidth usage monitoring. The bandwidth usage monitoring register interface is specified in [\[BC\_QOS\]](#BC%5FQOS). The capacity and bandwidth controller register interfaces for resource allocation and usage monitoring are defined as memory-mapped registers. Each controller that supports CBQRI provides a set of registers that are memory-mapped, starting at an 8-byte aligned physical address. The memory-mapped registers can be accessed by using naturally aligned 4-byte or 8-byte memory accesses. The controller behavior for register accesses where the address is not aligned to the size of the access, or if the access spans multiple registers, or if the size of the access is not 4 bytes or 8 bytes, is `UNSPECIFIED`. A 4-byte access to a register must be single-copy atomic. Whether an 8-byte access to a CBQRI register is single-copy atomic is considered `UNSPECIFIED`. This type of access can appear, internally to the CBQRI implementation, as if two separate 4-byte accesses are performed. | | The CBQRI registers are defined so that software can perform two individual 4 byte accesses, or hardware can perform two independent 4 byte transactions resulting from an 8 byte access, to the high and low halves of the register as long as the register semantics, with regards to side-effects, are respected between the two software accesses, or two hardware transactions, respectively. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The controller registers use little-endian byte order (even if all harts are big-endian-only). | | Big-endian-configured harts that make use of the register interface might implement the REV8 byte-reversal instruction defined by the Zbb extension. IfREV8 is not implemented, then endianness conversion might be implemented by using a sequence of instructions. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | A controller can support a subset of capabilities that are defined by CBQRI. When a capability is not supported, the registers and/or fields used to configure and/or control such capabilities are hardwired to `0`. Each controller supports a capabilities register to enumerate the supported capabilities. 5.1. IOMMU Extension for QoS ID ==================== ## [](#QOS%5FIOMMU)5.1\. IOMMU Extension for QoS ID A method to associate QoS IDs with requests to access resources by the Input-Output Memory Management Unit (IOMMU), as well as with devices governed by it, is required for effective monitoring and allocation. This section specifies a RISC-V I OMMU \[[5](qos%5Fbiblio.html#bib-iommu)\] extension for the following goals: * Configure and associate QoS IDs for device-originated requests. * Configure and associate QoS IDs for IOMMU-originated requests. The size (or width) of `RCID` and `MCID`, as fields in registers or in data structures, supported by the IOMMU must be at least as large as that supported by any RISC-V application processor hart in the system. ### [](#5-1-1-iommu-registers)5.1.1\. IOMMU Registers The specified memory-mapped register layout defines a new IOMMU register named`iommu_qosid`. This register is used to configure the Quality of Service (QoS) IDs associated with IOMMU-originated requests. The register is 4 bytes in size and is located at an offset of 624 from the beginning of the memory-mapped region. __Table 1\. IOMMU Memory-mapped Register Layout__ | Offset | Name | Size | Description | Is Optional? | | ------ | ------------ | ---- | ------------------------------ | ------------ | | 624 | iommu\_qosid | 4 | QoS IDs for IOMMU requests. | Yes | | 628 | Reserved | 60 | Reserved for future use (WPRI) | | #### [](#5-1-1-1-reset-behavior)5.1.1.1\. Reset Behavior If the reset value for `ddtp.iommu_mode` field is `Bare`, then the`iommu_qosid.RCID` field must have a reset value of 0. | | At reset, it is required that the RCID field of iommu\_qosid is set to 0 if the IOMMU is in Bare mode, as typically the resource controllers in the SoC default to a reset behavior of associating all capacity or bandwidth to theRCID value of 0\. When the reset value of the ddtp.iommu\_mode is not Bare, the iommu\_qosid register should be initialized by software before changing the mode to allow DMA. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#5-1-1-2-iommu-capabilities)5.1.1.2\. IOMMU Capabilities The IOMMU `capabilities` register is extended with a new field, `QOSID`, which enumerates support for associating QoS IDs with requests made through the IOMMU. ![IOMMU Capabilities Register](_images/diag-8fedf47ed65b4b85827d8c6395d82ac72cf74f30.svg) Figure 1\. IOMMU Capabilities Register | Bits | Field | Attribute | Description | | ---- | ----- | --------- | ----------------------------------------------- | | 41 | QOSID | RO | Associating QoS IDs with requests is supported. | #### [](#5-1-1-3-iommu-qos-id)5.1.1.3\. IOMMU QoS ID The `iommu_qosid` register fields are defined as follows: ![`iommu_qosid` register fields](_images/diag-c4837722535117106755568dc19ec27db080323f.svg) Figure 2\. `iommu_qosid` register fields | Bits | Field | Attribute | Description | | ----- | -------- | --------- | ---------------------------------- | | 11:0 | RCID | WARL | RCID for IOMMU-initiated requests. | | 15:12 | reserved | WPRI | Reserved for standard use. | | 27:16 | MCID | WARL | MCID for IOMMU-initiated requests. | | 31:28 | reserved | WPRI | Reserved for standard use. | IOMMU-initiated requests for accessing the following data structures use the value programmed in the `RCID` and `MCID` fields of the `iommu_qosid` register. * Device directory table (`DDT`) * Fault queue (`FQ`) * Command queue (`CQ`) * Page-request queue (`PQ`) * IOMMU-initiated MSI (Message-signaled interrupts) When `ddtp.iommu_mode == Bare`, all device-originated requests are associated with the QoS IDs configured in the `iommu_qosid` register. ### [](#5-1-2-device-context-fields)5.1.2\. Device-context Fields The `ta` field of the device context is extended with two new fields, `RCID`and `MCID`, to configure the QoS IDs to associate with requests originated by the devices. ![Translation Attributes (`ta`) Field](_images/diag-665ffb960a23a4bd93e5f5584ae8dc7af50b3dfc.svg) Figure 3\. Translation Attributes (`ta`) Field IOMMU-initiated requests for accessing the following data structures use the value configured in the `RCID` and `MCID` fields of `DC.ta`. * Process directory table (`PDT`) * Second-stage page table * First-stage page table * MSI page table * Memory-resident interrupt file (`MRIF`) The `RCID` and `MCID` configured in `DC.ta` are provided to the IO bridge on successful address translations. The IO bridge should associate these QoS IDs with device-initiated requests. If `capabilities.QOSID` is 1 and `DC.ta.RCID` or `DC.ta.MCID` is wider than that supported by the IOMMU, a `DC` with `DC.tc.V=1` is considered misconfigured. In this case, the IOMMU should stop and report "DDT entry misconfigured" (cause = 259). ### [](#5-1-3-iommu-atc-capacity-allocation-and-monitoring)5.1.3\. IOMMU ATC Capacity Allocation and Monitoring Some IOMMUs might support capacity allocation and usage monitoring in the IOMMU address translation cache (IOATC) by implementing the capacity controller register interface. Additionally, some IOMMUs might support multiple IOATCs, each potentially having different capacities. In scenarios where multiple IOATCs are implemented, such as an IOATC for each supported page size, the IOMMU can implement a capacity controller register interface for each IOATC to facilitate individual capacity allocation. 7.1. Software Guidelines ==================== ## [](#QOS%5FSW%5FGUIDE)7.1\. Software Guidelines ### [](#7-1-1-reporting-capacity-and-bandwidth-controllers)7.1.1\. Reporting Capacity and Bandwidth Controllers The capability and bandwidth controllers that are present in the system should be reported to operating systems using methods such as ACPI and/or device tree. For each capacity and bandwidth controller, the following information should be reported using these methods: * Type of controller (for example, cache, interconnect, memory, and so on) * Location of the register programming interface for the controller * Placement and topology describing the hart and IO bridges that share the resources controlled by the controller * The number of QoS identifiers supported by the controller * For memory bandwidth controllers, the controlled memory regions. These regions might be described in the form of NUMA domains or proximity domains. * If a controller is part of a set of controllers that collectively control a shared resource such as memory bandwidth of a memory region, then information to identify all members of the set should be reported. * Constraints imposed by the controllers, such as the minimum number of capacity or bandwidth blocks per RCID. ### [](#7-1-2-context-switching-qos-identifiers)7.1.2\. Context Switching QoS Identifiers Typically, the contents of the `srmcfg` CSR are updated with a new `RCID`and/or `MCID` by the HS/S-mode scheduler if the `RCID` and/or `MCID` of the new workload (a process or a VM) is not same as that of the previous workload. A context switch usually involves saving the context associated with the workload being switched away from and restoring the context of the workload being switched to. Such a context switch might be invoked in response to an explicit call from the workload (for example, as a function of an `ECALL` invocation) or can be done asynchronously (for example, in response to a timer interrupt). In such cases the scheduler might want to execute with the `srmcfg` configuration of the workload being switched away from such that this execution is attributed to the workload being switched away from and then prior to restoring the new workloads context, first switch to the `srmcfg` configuration appropriate for the workload being switched to such that all of that execution is attributed to the new workload. Further in this context switch process, if the scheduler intends some of its execution to be attributed to neither the outgoing workload nor the incoming workload, then the scheduler might switch to a new`srmcfg` configuration that is different from that of either of the workloads for the duration of such execution. ### [](#7-1-3-qos-configurations-for-virtual-machines)7.1.3\. QoS Configurations for Virtual Machines Usually for virtual machines the resource allocations are configured by the hypervisor. Usually the Guest OS in a virtual machine does not participate in the QoS flows as the Guest OS does not know the physical capabilities of the platform or the resource allocations for other virtual machines in the system. If a use case requires it, a hypervisor might virtualize the QoS capability to a VM by virtualizing the memory-mapped CBQRI register interface and virtualizing the virtual-instruction exception on access to `srmcfg` CSR by the Guest OS. | | If the use of directly selecting among a set of RCID and/or MCID by a VM becomes more prevalent and the overhead of virtualizing the srmcfg CSR using the virtual instruction exception is not acceptable then a future extension can be introduced where the RCID/MCID attempted to be written by VS mode are used as a selector for a set of RCID/MCID that the hypervisor configures in a set of HS mode CSRs. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#7-1-4-qos-identifiers-for-supervisor-and-machine-mode)7.1.4\. QoS Identifiers for Supervisor and Machine Mode The `RCID` and `MCID` configured in `srmcfg` also apply to execution in S/HS-mode, but this is typically not an issue. Usually, S/HS-mode execution occurs to provide services, such as through an ABI, to software executing at lower privilege. Because the S/HS-mode invocation provides a service for the lower privilege mode, the S/HS-mode software might not opt to modify the`srmcfg` CSR. Similarly, The `RCID` and `MCID` configured in `srmcfg` also apply to execution in M-mode, but this is typically not an issue either. Usually, M-mode execution occurs to provide services, such as through the SBI interface, to software executing at lower privilege. Because the M-mode invocation provides a service for the lower privilege mode, the M-mode software might not opt to modify the `srmcfg` CSR. If separate `RCID` and/or `MCID` are needed during software execution in M/S/HS-mode, then the M/S/HS-mode software might update the `srmcfg` CSR and restore it before returning to lower privilege mode execution. The statistical nature of QoS capabilities means that the brief duration, such as the few instructions in the M/S/HS-mode trap handler entry point, during which the trap handler might execute with the `RCID` and/or `MCID` established for lower privilege mode operation might not have a significant statistical impact. RISC-V RERI Architecture Specification ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-reri-architecture-specification)RISC-V RERI Architecture Specification RERI Task Group Version v1.0, 2024-05-24: Ratified | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at. Copyright 2022 - 2024 by RISC-V International. 3.1. Bibliography ==================== ## [](#3-1-bibliography)3.1\. Bibliography \[1\] _PCI Express® Base Specification Revision 6.0_, . \[Online\]. Available: \[2\] _Compute Express® Link (CXL) Specification Revision 3.0_, . \[Online\]. Available: \[3\] A. Algirdas and L. Jean-Claude and R. B. a. . . . . . . . . . . . L. Carl, "Basic Concepts and Taxonomy of Dependable and Secure Computing", _IEEE Trans. Dependable Secur. Comput._, vol. 1, no. 1, jan 2004\. pp. 11–33, \[Online\]. Available: . \[4\] S. Marc et al., "Addressing Failures in Exascale Computing", _Int. J. High Perform. Comput. Appl._, vol. 28, no. 2, may 2014\. pp. 129–173, \[Online\]. Available: . \[5\] K. Yoongu et al., "Flipping Bits in Memory without Accessing Them: An Experimental Study of DRAM Disturbance Errors" in _Proceeding of the 41st Annual International Symposium on Computer Architecuture_. IEEE Press, 2014, pp. 361–372. \[6\] Radojkovic and Petar, "Towards Resilient EU HPC Systems: A Blueprint" in _Proceedings of the 16th ACM International Conference on Computing Frontiers_. New York, NY, USA:, Association for Computing Machinery, 2019, pp. 339, Available: . \[7\] C. Franck and A. Geist and G. William and K. Sanjay and K. Bill and S. Marc, "Toward Exascale Resilience: 2014 Update", _Supercomput. Front. Innov.: Int. J._, vol. 1, no. 1, apr 2014\. pp. 5–28, \[Online\]. Available: . \[8\] S. Bianca and P. Eduardo and W. Wolf-Dietrich, "DRAM Errors in the Wild: A Large-Scale Field Study", _Commun. ACM_, vol. 54, no. 2, feb 2011\. pp. 100–107, \[Online\]. Available: . \[9\] S. Vilas and L. Dean, "A Study of DRAM Failures in the Field" in _International Conference on High Performance Computing, Networking, Storage and Analysis (SC)_. 2012, pp. 76:1—​76:11. \[10\] Z. Darko et al., "DRAM Errors in the Field: A Statistical Approach" in _International Symposium on Memory Systems (MEMSYS)_. 2019. \[11\] H. A. A. and S. I. A. and S. Bianca, "Cosmic Rays Don’t Strike Twice: Understanding the Nature of DRAM Errors and the Implications for System Design" in _International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS)_. 2012. \[12\] M. Justin and W. Qiang and K. Sanjeev and M. Onur, "Revisiting Memory Errors in Large-Scale Production Data Centers: Analysis and Modeling of New Trends from the Field" in _IEEE/IFIP International Conference on Dependable Systems and Networks (DSN)_. 2015, pp. 415—​426. \[13\] T. Dong and C. Peter and T. Zuheir and S. M. W., "Assessment of the Effect of Memory Page Retirement on System RAS Against Hardware Faults" in _International Conference on Dependable Systems and Networks (DSN)_. 2006. \[14\] D. Xiaoming et al., "Fault-Aware Prediction-Guided Page Offlining for Uncorrectable Memory Error Prevention" in _International Conference on Computer Design (ICCD)_. 2021. \[15\] M. C. Di and K. Zbigniew and I. R. K. and B. Fabio and F. Joseph and K. William, "Lessons Learned from the Analysis of System Failures at Petascale: The Case of Blue Waters" in _International Conference on Dependable Systems and Networks (DSN)_. 2014, pp. 610—​621. \[16\] D. Xiaoming and L. Cong and Z. Shen and Y. Mao and L. Jing, "Predicting Uncorrectable Memory Errors for Proactive Replacement: An Empirical Study on Large-Scale Field Data" in _European Dependable Computing Conference (EDCC)_. 2020. \[17\] _RISC-V Instruction Set Manual, Volume II: Privileged Architecture_, . \[Online\]. Available: \[18\] _RISC-V Instruction Set Manual, Volume I: Unprivileged ISA_, . \[Online\]. Available: Contributors ==================== ## [](#contributors)Contributors This RISC-V specification has been contributed to directly or indirectly by (in alphabetical order): Aaron Durbin, Allen Baum, Andrew Walter, Anup Patel, Cameron McNairy, Dimitris Gizopoulos, Daniele Rossi, David Kruckemeyer, Dhaval Sharma, Greg Favor, Himanshu Chauhan, Holger Blasum, Mark Hill, Nicasio Canino, Paul Donahue, Petar Radojkovic, Shubu Mukherjee, Vedvyas Shanbhogue, Xiaohan Ma 2.1. Error Reporting ==================== ## [](#2-1-error-reporting)2.1\. Error Reporting Components, such as a RISC-V hart or a memory controller, in a system that support error detection may implement one or more banks of error records. Each error bank may implement one or more error records. Each error record corresponds to one or more hardware units of the component and reports errors detected by those hardware units. A hardware unit may implement multiple error records. One or more error records may be valid at any given time due to one or more hardware units in the component detecting an error or due to a hardware unit having detected one or more errors. Each error bank is memory-mapped starting at an 8-byte aligned physical address and may include up to 63 error records. Each error record is a set of registers used to control that error record and to report status, address, and other information relevant to the error recorded in that error record. | | Implementations may use a coarser alignment for the start address of an error bank. For example, some implementations may locate the error bank within a naturally aligned 4-KiB region (a page) of physical address space for each error bank, i.e., one page per bank. Coarser alignments may enable register decoding to be implemented without a hardware adder circuit. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The behavior for register accesses where the address is not aligned to the size of the access, or if the access spans multiple registers, or if the size of the access is not 4 bytes or 8 bytes, is `UNSPECIFIED`. An aligned 4-byte access to a RERI register must be single-copy atomic. Whether an 8-byte access to an RERI register is single-copy atomic is `UNSPECIFIED`, and such an access may appear, internally to the RERI implementation, as if two separate 4-byte accesses were performed. | | The RERI registers are defined in such a way that software can perform two individual 4 byte accesses, or hardware can perform two independent 4 byte transactions resulting from an 8 byte access, to the high and low halves of the register as long as the register’s semantics, with regards to side-effects, are respected between the two software accesses, or two hardware transactions, respectively. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The RERI registers have little-endian byte order (even for systems where all harts are big-endian-only). | | Big-endian-configured harts using RERI may implement the REV8 byte-reversal instruction defined by the Zbb extension. If REV8 is not implemented, then endianness conversion may be implemented using a sequence of instructions. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | An implementation-specific response occurs if the error bank and/or record is unavailable (e.g., powered down) to memory-mapped accesses. For example, an error bank and/or record may respond with all zero data on reads and may ignore writes. Other implementations may, for example, signal an error response on the attempted transaction. An error bank that is otherwise available for memory-mapped accesses must respond with all zero data on reads and must ignore writes to unimplemented registers in the page. ### [](#2-1-1-register-layout)2.1.1\. Register Layout The error bank registers are organized as a 64-byte header providing information about the error bank followed by an array of 64-byte error records. The offset of the error record numbered `i` in the bank is (64 + `i` \* 64) where `i` may range from 0 to 62. __Table 1\. Error bank Memory-mapped register layout__ | Offset | Name | Size | Description | | ------------- | ------------------ | ---- | ---------------------------------------------------- | | 0 | vendor\_n\_imp\_id | 8 | Vendor and implementation ID. | | 8 | bank\_info | 8 | Error bank information. | | 16 | valid\_summary | 8 | Summary of valid error records. | | 24 | Reserved | 32 | Reserved for future standard use. | | 56 | Custom | 8 | Designated for custom use. | | 64 + 64 \* i | control\_i | 8 | Control register of error record i. | | 72 + 64 \* i | status\_i | 8 | Status register of error record i. | | 80 + 64 \* i | addr\_info\_i | 8 | Address-or-info. register of error record i. | | 88 + 64 \* i | info\_i | 8 | Information register of error record i. | | 96 + 64 \* i | suppl\_info\_i | 8 | Supplemental information register of error record i. | | 104 + 64 \* i | timestamp\_i | 8 | Timestamp register of error record i. | | 112 + 64 \* i | Reserved | 16 | Reserved for future standard use. | All registers and register fields defined by this specification are WARL unless noted otherwise. While all registers and register fields of an error bank and the error records in an error bank must exist, is legal to implement a register and/or register field of as read-only zero or a read-only legal value if they are not required to report errors information in an implementation. | | The number of error banks, the number of error records in an error bank and the amount of information reported in an error record may be implemented to meet the needs of the implementation. The error records are only required to implement the registers and register fields needed to report error information that is legally produced by the implementation. A minimal implementation with one error bank, which contains one error record only consumes 128 bytes of address space. In terms of storage, the minimal implementation requires only two bits of storage, for the v (valid) bit and the rdip (read-in-progress) bit, in the status\_i register in the single error record. All other register fields of the bank header and error record are WARL and may be hardwired to read-only zero or read-only one as appropriate. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#2-1-2-reset-behavior)2.1.2\. Reset Behavior The reset value is `UNSPECIFIED` for RERI registers. The registers of an error bank may preserve their value across certain types of reset. For example, a warm reset or a RAS initiated reset may preserve the register values whereas a cold reset may reset the values back to their initial state. | | Under normal circumstances, when an error is signaled, the RAS handler retrieves the logged errors to process the error condition. In some cases, the RAS handler may not be able to do such processing. For example, the system may be unable to support execution of the RAS handler and cause a RAS initiated reset. Preserving the information logged in error records across such resets allows reporting of unhandled errors that occurred in a previous boot of the system. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | All registers in an error bank must have the same reset behavior. ### [](#2-1-3-error-bank-header-registers)2.1.3\. Error Bank Header Registers #### [](#2-1-3-1-vendor-and-implementation-id-vendor%5Fn%5Fimp%5Fid)2.1.3.1\. Vendor and Implementation ID (`vendor_n_imp_id`) The `vendor_n_imp_id` register is a read-only register and its layout is: ![Vendor and implementation ID](_images/svg-c1ba819ec801750bb2aad29441ee2b310d066635.svg) Figure 1\. Vendor and implementation ID The `vendor_id` field follows the encoding as defined by `mvendorid` CSR and provides the JEDEC manufacturer ID of the provider of the component hosting the error bank. A value of 0 may be returned to indicate the field is not implemented or that this is a non-commercial implementation. The `imp_id` provides a unique identity, defined by the vendor, to identify the component and revisions of the component implementation hosting the error bank. A value of 0 may be returned to indicate that the field is not implemented. The value returned should reflect the design of the component itself and not of the surrounding system. | | The vendor\_id and the imp\_id are expected to be used as a identifier to determine the format of fields and encodings that are UNSPECIFIED by this specification. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#2-1-3-2-error-bank-information-bank%5Finfo)2.1.3.2\. Error Bank Information (`bank_info`) The `bank_info` is a read-only register and its layout is as follows: ![Error bank information](_images/svg-5e170184bf335d4f69de8011fec6d9f0db30a911.svg) Figure 2\. Error bank information The `version` field returns the version of the architectural register layout specification implemented by the error bank. The version defined by this specification is 0x01\. The encodings 0xF0 through 0xFF of this field are designated for custom use. The `layout` field along with the `version` field indicates the layout of the registers in the error bank and the error records. The `layout` encoding 0 indicates the registers are arranged and have meaning as defined by this specification. | | The offset of the version and the layout fields in the error bank shall not change across versions of the specification or the layouts defined by a version. Software should first read the version and layout fields and use the values to determine the register layout. The layout field may be used for future standard extensions to define segment specific extensions to the error bank and/or the error records. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `inst_id` field identifies a unique instance of an error bank, within a package or at least a silicon die, of the component; ideally unique in the whole system. The `inst_id` is defined by the vendor of the system as a unique identifier for the component. A value of 0 may be returned to indicate the field is not implemented. | | The inst\_id is expected to be collected and logged as part of the RAS error logs. These may allow the vendor of the silicon to make inferences about the instances of the components that may be vulnerable. As these values differ between vendors of the system and even among systems provided by the same vendor, these are not expected to be useful to the majority of software besides software intimately familiar with that system implementation. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `n_err_recs` field indicates the number of error records implemented by the error bank. The field is allowed to have an unsigned value between 1 and 63\. The error records of an error bank are located in the memory mapped region reserved for the error bank such that the first error record is at offset 64 and the last error record at offset (64 + 63 \* `n_err_recs`). #### [](#2-1-3-3-summary-of-valid-error-records-valid%5Fsummary)2.1.3.3\. Summary of Valid Error Records (`valid_summary`) The `valid_summary` is a read-only register and its layout is as follows: ![Summary of valid error records](_images/svg-d005408d0bb00c26010570367d9881776b7c699d.svg) Figure 3\. Summary of valid error records The `sv` bit when 1 indicates that the `valid_bitmap` provides a summary of the`valid` bits from the status registers of this error bank. If this bit is 0 then the error bank does not provide a summary of valid bits and the`valid_bitmap` is 0. | | If SV is 1, then software may use the valid\_bitmap to determine which error records in the bank are valid. If this bit is 0 then software must read thestatus\_register\_i of each implemented error record in this bank to determine if there is a valid error logged in that error record. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#2-1-4-error-record-registers)2.1.4\. Error Record Registers #### [](#2-1-4-1-control-register-control%5Fi)2.1.4.1\. Control Register (`control_i`) The `control_i` is a read/write WARL register used to control error reporting by the corresponding error record in the error bank. The layout of this register is as follows: ![Control register](_images/svg-b9a157fcfe4dd17e43d310ae63551e9ab585f008.svg) Figure 4\. Control register Error reporting functionality in the error record is enabled if the error-logging-and-signaling-enable (`else`) field is set to 1\. The `else` field is WARL and may default to 1 or 0 at reset. When `else` is 1, the hardware unit logs and signals errors in the error record. When `else` is 0, any signaling associated with prior logged errors remains unaffected, the hardware unit does not log and signal new errors in the error record, and it is `UNSPECIFIED`whether the hardware unit continues detecting and correcting errors. | | When error reporting is disabled, the hardware unit may continue to silently correct detected errors and when correction is not possible provide corrupt data to the consumers of the data. Alternatively an implementation may disable error detection altogether when error reporting is disabled. It is recommended that implementations continue performing error correction even when error reporting is disabled. It is recommended that a hardware component continue to produce error detection and correction codes on data generated by or stored in the hardware component even when error reporting is disabled. It is recommended hardware components continue to use containment techniques like data poisoning even when error reporting is disabled. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The `ces`, `ueds`, and `uecs` are WARL fields used to enable signaling of CE, UED, and UEC respectively when they are logged (i.e. when `else` is 1). Enables for unsupported classes of errors may be hardwired to 0\. The encodings of these fields are specified in [Table 2](#ERR%5FSIG%5FENABLES). __Table 2\. Error signaling enable field encodings__ | **Encoding** | **Error signal** | | ------------ | -------------------------------------------- | | 0 | Signaling is disabled. | | 1 | Signal using a Low-priority RAS signal. | | 2 | Signal using a High-priority RAS signal. | | 3 | Signal using a platform specific RAS signal. | The RAS signals are usually used to notify a RAS handler. The physical manifestation of the signal is `UNSPECIFIED` by this specification. The information carried by the signal is `UNSPECIFIED` by this specification. | | The error signaling enables typically default to 0 - disabled - at reset to allow a RAS handler an opportunity to initialize itself for handling RAS signals and to initialize the hardware units that generate the RAS signals before error reporting is enabled. The signal generated by the error record may in addition to causing an interrupt/event notification be also used to carry additional information to aid the RAS handler in the platform. The RAS handler may be implemented by a RISC-V application processor hart in the system, a dedicated RAS handling micro-controller, a Finite-State Machine (FSM), etc. The error signals may be configured, through platform specific means, to notify a RAS handler in the platform. For example, the High-priority RAS signal may be configured to cause a High-priority RAS local interrupt, an external interrupt, or an Non-Maskable Interrupt (NMI) and the Low-priority RAS signal may be configured to cause a Low-priority RAS local interrupt or an external interrupt. When error class and/or priority-specific RAS handlers are implemented, these handlers must take into consideration the possibility that an error record intended for a handler could be overwritten by an error of higher severity or priority — which also triggers a signal to another RAS handler for the new error — in the period between the first signal’s generation and its examination of the error record by the first RAS handler. In such instances, the first RAS handler may find an error record that is not intended for it. This handler may choose to disregard this error record as spurious from its perspective, and leave it to be handled by the other RAS handler. It may also note that an error occurred that concerns it, but information for the error is no longer available. Similarly, spurious signals may arise if the fields controlling the type of signal generated by an error record are modified while either the v field or the ceco field in the status\_i register is set to 1. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If the error record supports corrected-error counting then the corrected-error-counting-enable (`cece`) field, when set to 1, enables counting corrected errors in the corrected-error-counter (`cec`) field of the status register `status_i` of the error record. The `cec` is a counter that holds an unsigned integer count. When `cece` is 0, the `cec` does not count and retains its value. If corrected error counting is not supported in the error record then`cece` and `cec` may be hardwired to 0\. An overflow of `cec` is signaled using the signal configured in the `ces` field. When `cece` is 1, the logging of a CE in the error record does not cause an error signal and an error signal configured in `ces` occurs only on a `cec` overflow that sets the `ceco` bit. The set-read-in-progress (`srdp`) field, when written with a value of 1, causes the `rdip` (read-in-progress) bit of the associated `status_i` register to be set. The `srdp` field always returns 0 on read. The `rdip` field in the`status_i` register is set to 1 by hardware when an error is recorded in an invalid error record causing the `v` field to change from 0 to 1\. The `rdip`field is cleared to 0 by hardware when a new error updates any field of a valid (`v=1`) error record. The status-register-invalidate (`sinv`) bit, when written with a value of 1, causes the `v` (valid) field of the associated `status_i` register to be cleared if the `rdip` field in the `status_i` register is also 1\. The `sinv`field always returns 0 on read. The `sinv` field enables software to read out and invalidate an error record without needing to explicitly write the`status_i` register. Qualifying the clearing of the `v` field with `rdip` field being 1 prevents losing information about an overwrite that might have occurred while reading of the error record is in progress. If the `sinv` and `srdp` are both written to 1 together then the `rdip` bit is set and the `v` bit is cleared to 0. | | Software may determine if the error record was read atomically by first reading the registers of the error record, then clearing the valid in status\_i by writing 1 to control\_i.sinv and then reading the status\_i register again to determine if the v field was cleared to 0\. If the v field is still 1 but the rdip field is 0 then it is indicative of an overwrite that may have occurred during the process of reading the error record. If the v field is 1 and the rdip is also 1 then it indicates a new error was recorded after thev field was cleared; but the read of the error record to collect the previous error was atomic. If an overwrite occurred during the process of reading the error record then the process may be repeated, after setting the rdip field, to read the latest reported error. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The error-injection-delay (`eid`) is a WARL field used to control error record injection. When `eid` is written with a value greater than 0, the `eid` starts counting down, at an implementation defined rate, till the value reaches a count of 0\. Writing a value of 0 disables the counter. If error injection is not supported by the error record then the `eid` field may be hardwired to 0\. When`eid` reaches a count of 0, the status register is made valid by setting the`status_i.v` bit to 1\. The `status_i.v` transition from 0 to 1 generates a RAS signal corresponding to the class of error (CE, UED, or UEC) setup in the`status_i` register. The counter continues to count even if the `status_i`register was overwritten by a hardware detected error before the `eid` counts down to 0. | | Software may setup the error record registers with desired values of the error record to be injected and then program eid to cause the status\_i register to be marked valid when eid count reaches 0. The error record injection capability only injects an error record and not an error into the hardware itself. The error record injection capability is expected to be used to test the RAS handlers and is not intended to be used for verification of the hardware implementation itself. Other implementation specific mechanisms may be provided to generate and/or emulate hardware error conditions. When hardware error injection capabilities are implemented, the implementation should ensure that these capabilities cannot be misused to maliciously inject hardware errors that may lead to security issues. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#2-1-4-2-status-register-status%5Fi)2.1.4.2\. Status Register (`status_i`) The `status_i` is a read-write WARL register that reports errors detected by the hardware unit. ![Status register](_images/svg-99fdd4c73edb9e1e00af3064eb39236f9aede299.svg) Figure 5\. Status register The error record holds a valid error log if the valid (`v`) field is 1\. The`status_i` register does not accept a software write when the `v` field is 1. If the detected error was corrected then `ce` is set to 1\. If the detected error could not be corrected but was deferred then `ued` is set to 1\. If the detected error could not be corrected or deferred and thus needs immediate handling by an RAS handler, then the `uec` bit is set to 1\. If the error record does not log a class of errors (e.g., does not support UED), then the corresponding bit may be hardwired to 0\. If the bits corresponding to more than one error class are set to 1 then the error record holds information about the highest severity error class among the bits set. The error record may be used to provide an informational update by setting the `v` bit to 1 and setting `ce`, `ued`, and`uec` bits to 0\. Such informational updates are lower severity than a CE but are signaled using the signal configured in `control_i.ces`. When `v` is 1, if more errors of the same class as the error currently logged in the error record occur then the multiple-occurrence (`mo`) bit is set to indicate the multiple occurrence of errors of the same severity. See [2.1.5\. Error Record Overwrite Rules](#OVERWRITE%5FRULES)for rules on overwriting the error record in such cases. Each error of an error class (CE, UED, or UEC) that may be logged in an error record may be associated with a priority which is a number between 0 and 3; priority value of 3 being the highest priority and priority value of 0 being the lowest priority. The priority values indicate relative priority among errors of the same error class and therefore represent sub-classes of errors. Among errors of different error classes the priority values are unrelated. | | Some implementations may report errors from more than one sources into a single error records. Such implementations may prioritize reporting of error from one source over the other using the pri associated with the error when both sources simultaneously detect an error of the same class (e.g., CE). The priority is also used to determine if a new error may overwrite a previously reported error of the same error class in the error record. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The priority (`pri`) field in the error record indicates the priority of the currently logged error in the error record. The `pri` is a WARL field and an implementation may support only a subset of legal values for this field and an implementation that does not support reporting of a priority per error may hardwire this field to 0. The error record overwrite rules use the error class (CE, UED, or UEC) and the error priority (`pri`) as specified in [2.1.5\. Error Record Overwrite Rules](#OVERWRITE%5FRULES). When an UEC occurs the containable (`c`) bit may be set to 1 to indicate that the error has not propagated beyond the boundaries of the hardware unit that detected the error and thus may be **containable** through recovery actions (e.g., terminating the computation, etc.) carried out by the RAS handler. The `c` bit is WARL. For error classes other than UEC, the interpretation of the `c` bit may be specified in a future standard extension. For a RISC-V hart, some UEC may cause a Hardware Error exception \[[17](reri%5Fbibliography.html#bib-priv)\]. A Hardware Error is a synchronous exception, triggered when corrupted or uncorrectable data is accessed, either explicitly or implicitly, by an instruction. In this context, "data" encompasses all types of information used within a RISC-V hart. | | For example, a RISC-V hart by causing the precise hardware error exception on attempts to consume corrupted/poisoned data may contain the error to the program currently executing on the hart. Such errors may be reported with the c bit set to 1 indicating that the interrupted context may be restarted if the RAS handler is able to perform a suitable recovery operation. The _x_epc CSR on delivery of the hardware error exception holds the address of the instruction that attempted to access corrupted data, while the _x_tval CSR is either set to 0 or holds the virtual address of an instruction fetch, load, or store that attempted to access corrupted data. While the c bit indicates that the error may be containable the RAS handler may or may not be able to recover the system from such errors. The RAS handler must make the recovery determination based on additional information provided in the error record such as the address of the memory where corruption was detected. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The address-or-info-type (`ait`) is a WARL field that indicates the type of information reported in the `addr_info_i` register. An error record that does not report information in this field may hardwire this field to 0\. The encodings of the `ait` field are listed in [Table 3](#AIT%5FENCODINGS). __Table 3\. Address-or-information type encodings__ | **Encoding** | **Description** | | ------------ | ------------------------------------------------------------------------------ | | 0 | None. The contents of the addr\_info\_i register areUNSPECIFIED when ait is 0. | | 1 | Supervisor Physical Address (SPA). | | 2 | Guest Physical Address (GPA). | | 3 | Virtual Address (VA). | | 4-15 | Component-specific address or information. | | | Component-specific information types, as defined in the range 4-15 of the aitfield, may be used to report component-specific addresses or other component-specific information in the register. The component-specific addresses may include information such as a local bus address or a Dynamic Random-Access Memory (DRAM) address. The interpretation of such information is component-specific. When a standard address type (a VA, SPA, or GPA) is reported in theaddr\_info\_i register, additional non-redundant information about the location accessed using the address (e.g., cache set and way, etc.) may be reported in the info\_i and/or the suppl\_info\_i registers. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The transaction-type (`tt`) is a WARL field to report the type of transaction that detected the error and its encodings are listed in [Table 4](#TT%5FENCODINGS). An error record that does not report transaction types may hardwire this field to 0. __Table 4\. Transaction type encodings__ | **Encoding** | **Description** | | ------------ | --------------------------------- | | 0 | Unspecified or not applicable. | | 1 | Designated for custom use. | | 2-3 | Reserved for future standard use. | | 4 | Explicit read. | | 5 | Explicit write. | | 6 | Implicit read. | | 7 | Implicit write. | For a RISC-V hart, the Unprivileged specification \[[18](reri%5Fbibliography.html#bib-upriv)\] defines memory accesses by instructions as either explicit or implicit. An Implicit read or write is an access that may be implicitly performed by hardware to perform an explicit operation. For example, a load or store instruction executed by the hart may perform implicit memory accesses to page table data structures. Instruction memory accesses by a hart are termed as implicit accesses by the Unprivileged specification. However, for the purposes of error reporting, only the implicit accesses to data structures, such as the (guest) page tables that are used to determine the address of the instructions to be fetched, are termed as implicit accesses. The read to fetch the instruction bytes themselves is classified as an explicit read. | | Implementations may report additional information about the transaction (e.g., whether speculative, on-demand vs. prefetch, etc.) in the info\_i and/orsuppl\_info\_i registers. A non-hart component may also perform implicit accesses in order to process an explicit transaction. For example, processing a memory transaction may require a fabric component to implicitly access a routing table data structure. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If the detected error reports additional information in the `info_i` register then the information-valid (`iv`) field is set to 1\. If the detected error reports additional supplemental information in the `suppl_info_i` register then supplemental-information-valid (`siv`) field is set to 1\. The `iv` and/or `siv`fields may be hardwired to 0 if the error record does not provide information in`info_i` and/or `suppl_info_i` registers. When `iv` is 0, the value in `info_i`register is `UNSPECIFIED`. When `siv` is 0, the value in `suppl_info_i` register is `UNSPECIFIED`. If the error record holds a timestamp of when the last error was logged in the`timestamp_i` register then the timestamp-valid (`tsv`) field is set to 1\. This field may be hardwired to 0 if the error record does not report a timestamp with the error. When `tsv` field is 0, the value in `timestamp_i` register is`UNSPECIFIED`. The `scrub` bit is valid when a CE is logged and when set to 1 indicates that the storage location that held the data value has been updated with the corrected value (i.e., the data has been scrubbed). In an implementation that cannot make this distinction then it may conservatively report this field as 0\. When the error record is not associated with storage elements (e.g., correcting errors detected on bus transactions) this field may be hardwired to 0\. If this property is unconditionally true for a hardware unit then this field may be hardwired to 1\. For error classes other than CE, the interpretation of the `c`bit may be specified in a future standard extension. The error-code (`ec`) is a WARL field that holds an error code that provides a description of the detected error. Standard `ec` encodings are defined in[Table 5](#EC%5FENCODINGS). If an error record detects an error that does not correspond to a standard `ec` encoding then such errors may be reported using a custom encoding. The custom encodings have the most significant bit set to 1 to differentiate them from the standard encodings. The read-in-progress (`rdip`) field is set to 1 by hardware when a new error is recorded in an invalid status register and is cleared to 0 by hardware when a valid status register is overwritten. When the `control_i.sinv` field is written to 1, the `v` field is cleared to 0 only if the `rdip` field is 1\. Gating the clearing of the `v` field by the `rdip` field being 1 allows software to detect an overwrite that may occur while it is in process of reading an error record. An error record that supports the 1 setting of the `cece` field in `control_i`, implements a corrected-error-counter in the `cec` field. The `cec` is a WARL field. When `cece` is 1, the `cec` is incremented on each CE. If an unsigned integer overflow occurs on an `cec` increment then the corrected-error-counter-overflow (`ceco`) field is set to 1\. The `cec`continues to count following an overflow. The `cec` and `ceco` fields hold valid data and continue to count even when the `v` field is 0. | | Some hardware units may maintain a history of CE and may report a CE and may increment the cec only if the error is not identical to a previously reported CE. Some hardware units may implement low pass filters (e.g., leaky buckets) that throttle the rate at which CE are reported and counted. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | To invalidate a valid error record (presumably after having first read the error record), software should write 1 to the control\_i.sinv control bit to clear the v bit in the status\_i register of the error record. Using the sinvcontrol to clear the v bit, as compared to an explicit write to the register, avoids overwriting the cec and ceco fields (which typically want to be maintained across logged errors). If software needs to initialize the cec and/or ceco, then a software write to the status\_i register is appropriate. Before performing the write, software should first check for and read any valid error record, invalidate the error record, and then write the register with the new cec and/or ceco value and with v=0. If status\_i register write was not accepted due to hardware writing a new error into the record and setting the v field to 1, then software should repeat this process. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | When an UEC or UED error is logged in an error record, the `cec` and `ceco`fields of the error record are not modified and retain their values. #### [](#2-1-4-3-address-or-information-register-addr%5Finfo%5Fi)2.1.4.3\. Address-or-Information Register (`addr_info_i`) The `addr_info_i` WARL register reports the address or other information associated with the detected error when `status_i.ait` is not 0\. If`status_i.ait` is 0, the value in this register is `UNSPECIFIED`. An implementation that does not report information in this register may hardwire this register to 0\. Some fields of this register may be hardwired to zero if the field is unused to report any type of address or information. When an address (a VA, GPA, or an SPA) is reported in this register, to the extent possible, the error record should capture all significant parts of the address. However, as a function of the type of error being logged some address fields may be zeroes. Some of the highest address bits may be fixed or may be sign-extensions or may be zero-extensions of the next lowest address bit depending on the type of address reported. When component specific information is reported in this register, the interpretation of the information is component specific. #### [](#2-1-4-4-information-register-info%5Fi)2.1.4.4\. Information Register (`info_i`) The `info_i` WARL register provides additional information about the error when`status_i.iv` is 1\. If `status_i.iv` is 0, the value in this register is`UNSPECIFIED`. An implementation that does not report any additional information may hardwire this register to 0. The format of the register is `UNSPECIFIED` by this specification. This field may be interpreted using the error code in `status_i.ec` along with implementation defined format and rules. | | This register may be used to report information for guiding recovery, error nature (transient/permanent), error location (set/way, parity group, ECC syndrome), and other details (protocol FSM state, assertion failures). Components that are or monitor field replaceable units may log information in this register to identify the failing component. For example, a memory controller may log the DIMM channel, bank, column, row, rank, subRank, device ID, etc. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#2-1-4-5-supplemental-information-register-suppl%5Finfo%5Fi)2.1.4.5\. Supplemental Information Register (`suppl_info_i`) The `suppl_info_i` WARL register provides additional information about the error when `status_i.siv` is 1\. This information may supplement the information provided in `info_i` register. If `status_i.siv` is 0, the value in this register is `UNSPECIFIED`. An implementation that does not report any supplemental information may hardwire this register to 0. The format of the register is `UNSPECIFIED` by this specification. This field may be interpreted using the error code in `status_i.ec` along with implementation specific and implementation defined format and rules. #### [](#2-1-4-6-timestamp-register-timestamp%5Fi)2.1.4.6\. Timestamp Register (`timestamp_i`) The `timestamp_i` WARL register provides a timestamp for the last error recorded in the error record if `status_i.tsv` is 1\. When `status.tsv` is 0, the value in this register is `UNSPECIFIED`. An implementation that does not report a timestamp may hardwire this register to 0\. Some fields of the register may be hardwired to zero if the field is unused to report the timestamp. The nature, frequency, and resolution of the timestamp are `UNSPECIFIED`. | | The timestamp may be constructed by a hardware unit using mechanism such as sampling a local cycles counter (e.g., the cycles counter of a RISC-V hart, a global counter (e.g, mtime, etc.), or other implementation specific means. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#OVERWRITE%5FRULES)2.1.5\. Error Record Overwrite Rules When a hardware unit detects an error and its error record is not valid, it writes the error record with the error information and marks the record as valid. However, if the error record is already valid, owing to an earlier detected but unprocessed error, the decision to overwrite the error record with new error information is determined by the new error’s severity and/or priority. The overwrite rules allow a higher severity error to overwrite a lower severity error. UEC has the highest severity, followed by UED, then CE, and finally, informational. When the two errors have the same severity the priority of the errors (as determined by `status_i.pri`) is used to determine if the error record is overwritten. Higher priority errors overwrite the lower priority errors. When an error record is overwritten by a higher severity error (UED/CE by UEC, UED by UEC, or CE by UEC/UED), the status bits indicating the severity of the older errors are retained (i.e., are sticky). When an error writes or overwrites an error record, the `status_i.cec` and`status_i.ceco` fields update from CEs and retain value for errors of other severity. When implemented, `cec` counts CE occurrences; unsigned integer overflow on `cec` increment sets `ceco` to 1. Whenever a new error writes to or overwrites an error record, the signal configured in the `control_i` register for its severity level is asserted. When`status_i.ceco` changes from 0 to 1, the signal configured in `control_i.ces` is asserted. Error record writing rules Let new_status be the value to be recorded in status_i register for the new error overwrite = FALSE if status_i.v == 1 // There is a valid first error recorded if ( severity(new_error) > severity(status_i) ) // Higher severity errors overwrite less severe errors and clear mo status_i.mo = 0 overwrite = TRUE endif if ( severity(new_status) == severity(status_i) ) // Second errors of the same severity set MO status_i.mo = 1 // Second error of same severity overwrites previous error if it // has higher priority (status_i.pri). if ( new_status.pri > status_i.pri ) overwrite = TRUE; endif endif // previous error status bits are retained (sticky) but rdip bit is cleared. status_i.rdip = 0 status_i.uec |= new_status.uec status_i.ued |= new_status.ued status_i.ce |= new_status.ce else // No valid error recorded; new error logged, clearing sticky history // and MO bit, and rdip is set. status_i.rdip = 1 status_i.uec = new_status.uec status_i.ued = new_status.ued & ~new_status.uec status_i.ce = new_status.ce & ~new_status.uec & ~new_status.ued status_i.mo = 0 overwrite = TRUE; endif if ( overwrite = TRUE ) status_i.pri = new_status.pri status_i.c = new_status.c status_i.tt = new_status.tt status_i.ait = new_status.ait status_i.iv = new_status.iv status_i.siv = new_status.siv status_i.tsv = new_status.tsv status_i.scrub = new_status.scrub status_i.ec = new_status.ec // Update addr_info_i, info_i, suppl_info_i, and timestamp_i with new // error information, if valid. status_i.v = 1 endif If the `status_i.v`, `status_i.mo`, and `status_i.uec` are all 1 then the RAS handler should preferably restart the system to bring it to a correct state as an UEC record has been lost. If the `status_i.v` and `status_i.mo` are 1 but`status_i.uec` is 0 (i.e., the logged error is a UED or a CE) then the RAS handler may keep the system operational. If multiple errors occur simultaneously then they may be recorded individually in any order and the rules outlined in [Error record writing rules](#REC%5FWRITE%5FRULE) lead to the highest severity error among them being retained in the error record. When the error record registers are written by an error, all registers that are written must be written with information related to that error. | | When multiple errors occur simultaneously, some implementations may choose to record each error individually following the rules outlined in[Error record writing rules](#REC%5FWRITE%5FRULE). Other implementations may however choose to only record the highest severity error or when they have the same severity the highest priority error. And yet another implementation may choose to record one of the errors as determined by implementation specific rules. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#2-1-6-error-reporting-defined-by-other-standards)2.1.6\. Error Reporting Defined by Other Standards Standards such as PCIe \[[1](reri%5Fbibliography.html#bib-pci)\] and CXL \[[2](reri%5Fbibliography.html#bib-cxl)\] define standardized error reporting architectures such as the PCIe Advanced Error Reporting (AER). Specifications such as CXL define a standardized set of RAS requirements for hosts and devices. The RISC-V RERI specification complements the error reporting architecture defined by these standards with a RISC-V standard for reporting errors for components that are not PCIe/CXL components. There may also be other error reporting mechanisms, possibly custom, that are employed alongside the RERI specification. | | The RISC-V system components such as PCIe root ports or PCIe Root Complex Event Collectors may themselves implement error reporting compliant with the RISC-V RERI specification and thus provide a unified error reporting mechanism in such systems. For example, a root complex event collector may support an error record to report errors logged in the Advanced Error Reporting (AER) log registers. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#2-1-7-error-code-encodings)2.1.7\. Error Code Encodings __Table 5\. Error code encodings__ | **Encoding** | **Description** | | ------------ | ------------------------------------------------------------------------------------ | | 0 | None | | 1 | Other unspecified error occurred | | 2 | Corrupted data access (e.g., attempt to consume poisoned data) error | | 3 | Cache block data (e.g., ECC error on cache data) error | | 4 | Cache scrubbing detected (e.g., ECC error on cache data) error | | 5 | Cache address/control state (e.g., parity error tag or state) error | | 6 | Cache unspecified error | | 7 | Snoop-filter/directory address/control state (e.g., ECC error on tag or state) error | | 8 | Snoop-filter/directory unspecified error | | 9 | TLB/Page-walk cache data (e.g., ECC error on TLB data) error | | 10 | TLB/Page-walk cache address/control state (e.g., ECC error on TLB tag) error | | 11 | TLB/Page-walk cache unspecified error | | 12 | Hart state error (e.g., ECC error on CSRs or x/f/v registers) | | 13 | Interrupt controller state (e.g., ECC error on interrupt pending/enable state) error | | 14 | Interconnect data (e.g., ECC error on data bus) error | | 15 | Interconnect other (e.g., parity error on address bus) error | | 16 | Internal watchdog error | | 17 | Internal datapath, memory, or execution units error (e.g, ALU datapath parity) | | 18 | System memory command/address bus error | | 19 | System memory unspecified error | | 20 | System memory data (e.g., ECC error in SDRAM or HBM) error | | 21 | System Memory scrubbing detected error | | 22 | Protocol Error - illegal input/output error | | 23 | Protocol Error - illegal/unexpected state error | | 24 | Protocol Error - timeout error | | 25 | System internal controller (power management, security, etc.) error | | 26 | Deferred error pass-through (e.g., forwarding poisoned data) not supported | | 27 | PCIe/CXL detected (e.g., logged into PCIe AER, CXL.mem error log, etc.) errors | | 28 - 63 | Reserved for future standard use | | 64 - 255 | Designated for custom use | 1.1. Introduction ==================== ## [](#intro)1.1\. Introduction The RAS Error Record Register Interface (RERI) specification augments Reliability, Availability, and Serviceability (RAS) features in the SoC with a standard mechanism for reporting errors by means of a memory-mapped register interface to enable error reporting, provide the facility to log the detected errors (including their severity, nature, and location), and configuring means to signal the error to a RAS handler component. The RAS handler may use this information to determine suitable recovery actions that may include terminating the computation (e.g., terminating a process), restarting parts or all of the system, etc. to recover from the errors. Additionally, this specification shall support software-initiated error logging, reporting, and testing of RAS handlers. Lastly, this specification shall provide maximal flexibility to implement error handling and coexists with RAS frameworks defined by other standards such as PCIe \[[1](reri%5Fbibliography.html#bib-pci)\] and CXL \[[2](reri%5Fbibliography.html#bib-cxl)\]. A system is an entity that interacts with other entities such as other systems, software, operators, etc. to deliver one or more services in its role as a service provider. A system may itself be a consumer of one or more services provided by one or more other systems. A system thus is a collection of interacting components that implement one or more functions to provide a service. A service is the behavior as perceived by the consumers of the service. A system may implement the service as one or more functions in the system. The functions used to compose the service may be implemented by one or more components in the system. A service is described as a set of states that can be observed by the consumer of the service. The set of states observed by the consumer of the service may be further dependent on a set of internal states of the functions that implement the service. A service is said to be correct if the set of states observed by the consumer of the service match the specification of that service. The specifications of a service may include its functional behavior, performance goals, security objectives, and RAS requirements. Reliability of a system as a function of time is the probability it continues to provide correct service and may be characterized by metrics such as mean time between failures (MTBF). The services provided by a reliable system fail on faults instead of silently producing incorrect results. Reliable systems incorporate methods to detect occurrence of errors and to signal the errors to the consumers of the service. Availability of a system as a function of time is the probability that the system provides the expected service and is a measure of tolerance to errors. Systems may increase their availability by minimizing the impact of the errors in one part of the system to the rest of the system. This may be achieved by means such as error correction, redundancy, state checkpoints and rollbacks, error prediction, and error containment. Serviceability is a measure of time to restore the service to correct operation with minimal disruption to the consumers of the service. These may be achieved by means such as identifying and reporting failures and supporting mechanisms to repair and bring the system back online. ### [](#1-1-1-faults-and-errors)1.1.1\. Faults and Errors A fault is an incorrect state resulting from failures of components or due to interference from the environment in which the system operates. A fault is permanent if it reflects an irreversible change to the observable system state and is transient otherwise. A permanent fault may occur due to a physical defect or due to a flaw in the design of the functions implementing the service itself. A transient fault may occur due to temporary environmental conditions (cosmic rays, voltage glitches, etc.) or due to instability (e.g. marginal hardware). Some faults that occur in a component may be dormant and only affect the internal state of the component. Such dormant faults however may turn into active faults when that internal state is used by the computation process in that component and produce an error. An error is detected when its presence is indicated by an error message or signal. Software faults may similarly cause errors that cause the service provided by the system to deviate from its specification. Well known software engineering and reliability techniques may be employed to prevent, detect and recover from software errors. Software errors are not in the scope of this specification. Software should not have the ability to induce hardware errors. A service failure occurs when the service deviates from its specification due to errors. A reliable system deals with errors through one or more of the following techniques \[[3](reri%5Fbibliography.html#bib-ras%5Ftax)\] \[[4](reri%5Fbibliography.html#bib-exa%5Fras)\]: * Fault prevention * Error detection and correction * Error prediction ### [](#1-1-2-fault-prevention)1.1.2\. Fault Prevention Fault prevention involves use of techniques that reduce or prevent errors that may occur after the product has been shipped. These may be accomplished through the use of high quality in product design, technology selection, materials selection, and manufacturing time screening for defects. Through the use of systematic design, technology selection, and manufacturing tests many errors such as those induced by electric fields, temperature stress, switching/coupling noise (e.g. DRAM RowHammer \[[5](reri%5Fbibliography.html#bib-rham)\] effect), incorrect V/F operating points, insufficient guard bands, meta-stability, etc. can be prevented. Faults that are not prevented may manifest as errors during operation of the system. Errors that are not detected may still lead to a service failure. For example, an undetected error in an adder used to produce the address of a load may produce a bad address which causes the load to incur an exception and lead to a service failure. Some undetected errors however may not manifest as exceptions and cause a service failure due to silent data corruption. For example, a circuit performing encryption of a database may silently cause an error in the ciphertext produced leading to the entire database being left in a state where it cannot be decrypted. Such undetected errors that do not lead to a service failure are called Silent Data Errors (SDE). The impact of SDE is generally much higher than errors that lead to a service failure. A resilient system attempts to minimize the probability of SDE to the largest extent possible by implementing error detection capabilities. ### [](#1-1-3-error-detection-and-correction)1.1.3\. Error Detection and Correction Error detection involves the use of coding and protocols to detect errors \[[6](reri%5Fbibliography.html#bib-eu%5Fhpc)\] \[[7](reri%5Fbibliography.html#bib-exa%5F2014)\]. For example, caches with error correcting codes, TLB entries with parity protection, buses with parity protection on transaction fields, circuitry to detect unexpected and/or illegal encodings, gray codes, voltage sensors, clock/PLL monitors, timing margin sensors, etc. Some components such as memory controllers may actively attempt to detect errors using techniques such as periodic background scrubbing or on-demand scrubbing. Error correction involves the use of techniques to correct the detected errors. Error correction may be performed by employing error correcting codes and protocols. For example, a processor cache may employ error correcting codes (ECC) to detect and correct errors. Some components may recover from errors by using protocols that involve a retry. For example, a TLB that detects an error may invalidate the entry and attempt to refill it from the page tables, a receiver on a bus that detects an error may request the transmitter to retransmit the transaction, etc. Error correction is thus complete when the error is either corrected or it does not recur on retry. Such errors that were corrected by the hardware are called **Corrected Errors (CE)**. Errors that could not be corrected are called uncorrected errors. A component that detects an uncorrected error may allow possibly corrupted data to propagate to the requester of the data but associate an indicator (e.g., poison) with the data. Such errors are said to be **Uncorrected Errors Deferred (UED)** as they allow the component to continue operation and defer dealing with the error to a later point in time if the data corrupted by the error is consumed. Deferring errors allows deferring the error handling to an ultimate consumer of the corrupted data that may be able to provide more precise information to a RAS handler about the contexts affected by the corruption and thus enable more precise error recovery actions by the RAS handler. The component that detected and deferred the error may notify a RAS handler by reporting the UED but such a UED does not need an immediate remedial action to be performed by the RAS handler. For example, a memory controller may detect an uncorrectable ECC error on data in memory but since there is no immediate consumer of the data the memory controller may just mark the data as poisoned and defer the error handling to a component that requests the data. If the poisoned data is never consumed then deferred errors are benign. If the poisoned data is completely overwritten with new data then the associated poison is cleared. If the poisoned data is only partially written then the data continues to be marked as poisoned. A component that detects an uncorrected error may be unable to defer the handling of the error by techniques such as poisoning. Such errors are said to be **Uncorrected Errors Critical (UEC)** and a RAS handler is invoked as immediate remedial actions are required. For example, a cache controller may detect an uncorrectable ECC error on the memory used to hold cache tags and since such errors cannot be attributed to any particular data element these errors may be classified as UEC. If poisoned data is attempted to be consumed by a component (e.g. a hart, an IOMMU, a device, etc.) then an UEC occurs as immediate remedial actions are required and further deferral of the error is not possible. A component that signals a request for execution of an RAS handler for an UEC may indicate that the error has not propagated beyond the boundaries of the component that detected the error and thus may be **containable** through recovery actions (e.g., terminating the computation, etc.) carried out by the RAS handler. Some components act as an intermediary through which the data passes through. For example, a PCIe/CXL port is an intermediary component that by itself does not consume the data it receives from memory but forwards the data to the endpoint. In such cases the component may receive the data with a deferred error. Such a component may propagate the error and not log an error by itself. However, if the component to which the data is being propagated (e.g. a PCIe endpoint) is not capable of handling poison then the former component must signal a UEC instead of propagating the corrupted data, as the act of propagation breaks containment of the error. An error detected by a component may lead to a failure mode where the component may not be able to service requests anymore (e.g. colloquially called jammed, wedged, etc.). For example, an error in the hart pipeline may cause the hart to stop committing instructions, a fabric may be in a state where it cannot process any further requests, the link connecting the memory module to the host may have failed, etc. In such cases invoking a RAS handler may not be useful as the RAS handler itself may need to generate requests to the failed component to perform the recovery actions. Components in such failed states may use an implementation-defined signal to a system recovery controller (e.g., a Baseboard Management Controller (BMC), an on-chip service controller, etc.) to initiate a RAS-handling reset to restart the component, sub-system, or the system itself to restore correct service operations. ### [](#1-1-4-error-prediction)1.1.4\. Error Prediction Error prediction involves the use of corrected errors as a predictor of future uncorrectable permanent failures or other systemic issues, such as marginality due to aging. Monitoring corrected errors may facilitate the avoidance of future service failures. Studies indicate that the probability of an uncorrected DRAM error is elevated if the DIMM previously experienced corrected errors \[[8](reri%5Fbibliography.html#bib-dram%5Fwild)\] \[[9](reri%5Fbibliography.html#bib-sri%5F2012)\] \[[10](reri%5Fbibliography.html#bib-ziv%5F2019)\]. Such reasoning is used by system protection mechanisms, which utilize simple heuristics for offlining potentially failing memory pages \[[11](reri%5Fbibliography.html#bib-hwa%5F2012)\] \[[12](reri%5Fbibliography.html#bib-mez%5F2015)\] \[[13](reri%5Fbibliography.html#bib-tang%5F2006)\] \[[14](reri%5Fbibliography.html#bib-du%5F2021)\] or for replacing compromised DIMMs \[[15](reri%5Fbibliography.html#bib-mar%5F2014)\] \[[8](reri%5Fbibliography.html#bib-dram%5Fwild)\] \[[16](reri%5Fbibliography.html#bib-du%5F2020)\]. Reporting of detected and corrected hardware errors is requisite for any quantitative analysis of system resilience and for the prediction of future uncorrected errors \[[6](reri%5Fbibliography.html#bib-eu%5Fhpc)\]. This prediction capability facilitates the deployment of preventive mechanisms, such as pre-failure alerts in High-Performance Computing (HPC) cluster management software, thus mitigating the costs associated with unscheduled outages and system repairs. Components of a resilient system may also include corrected error counters to count the corrections performed. Such components may additionally include a fixed or programmable threshold to notify a RAS handler when the number of corrected errors surpasses the threshold. ### [](#1-1-5-reri-features)1.1.5\. RERI Features Version 1.0 of the RISC-V RERI specification supports the following features: * Error severity classes and standard error codes. * Standard register format and addressing for memory-mapped error-record registers and error-record banks. * Rules for prioritized overwriting of valid error records with new error records. * Corrected error counting. * Error record injection for RAS handler testing. This specification is intended to accommodate a wide variety of systems designs and needs - from high-end server-class systems to low-end embedded systems. This is accomplished through providing implementation flexibility and options - both within the registers of an error record and the number of error records in an error bank, and with respect to the association between hardware components and error errors/banks. ### [](#1-1-6-glossary)1.1.6\. Glossary __Table 1\. Terms and definitions__ | Term | Definition | | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | AER | Advanced Error Reporting. A PCIe capability to support advanced error control and reporting. | | BMC | Baseboard Management Controller. | | CE | Corrected Error. | | Custom | A register or data structure field designated for custom use. Software that is not aware of the custom use must ignore custom fields and preserve value held in these fields when writing values to other fields in the same register. | | CXL | Compute Express Link bus standard. | | Data | In this specification data refers broadly to all forms of information being stored or transferred in a computing system. In the case of a CPU, for example, this encompasses information that may be treated as instructions that are fetched and executed, as well as data that is loaded and stored. | | DIMM | Dual-In-line Memory Module. A packaging arrangement of memory devices on a socketable substrate. | | DRAM | Dynamic random-access memory. Devices made using Dynamic RAM circuit configurations that have data storage that must be refreshed periodically. | | ECC | Error Correcting Code. | | Error Reporting | Error reporting is the process of logging information (including their severity, nature, and location) about a detected error in an error record and signaling, if required, the occurrence of the error to an appropriate RAS handler. | | FSM | Finite-State Machine. An abstract machine that can be in exactly one of a finite number of states at any time. | | GPA | Guest Physical Address. See Priv. specification. | | HPC | High-performance Computing. High-Performance Computing (HPC) refers to the use of parallel processing techniques to solve complex computational problems. It enables faster data processing and simulation by leveraging multiple processors or servers. | | ID | Identifier. | | IOMMU | Input-Output Memory Management Unit. A system-level Memory Management Unit (MMU) that connects direct-memory-access capable Input/Output (I/O) devices to system memory. | | NMI | Non-Maskable interrupt. See Priv. specification. | | OS | Operating System. | | PLL | Phase-Locked Loop. A control system that generates an output signal whose phase is related to the phase of an input signal. PLLs are commonly used to perform clock synthesis. | | PCIe | Peripheral Component Interconnect Express bus standard. | | RAS | Reliability, Availability, and Serviceability. | | RERI | RAS error record register interface. | | Reserved | A register or data structure field reserved for future use. Reserved fields in data structures must be set to 0 by software. Software must ignore reserved fields in registers and preserve the value held in these fields when writing values to other fields in the same register. | | RO | Read-Only - Register bits are read-only and cannot be altered by software. Where explicitly defined, these bits are used to reflect changing hardware state, and as a result bit values can be observed to change at run time. If the optional feature that would Set the bits is not implemented, the bits must be hardwired to Zero | | RW | Read-Write - Register bits are read-write and are permitted to be either Set or Cleared by software to the desired state. If the optional feature that is associated with the bits is not implemented, the bits are permitted to be hardwired to Zero. | | RW1C | Write-1-to-Clear status - Register bits indicate status when read. A Set bit indicates a status event which is Cleared by writing a 1b. Writing a 0b to RW1C bits has no effect. If the optional feature that would Set the bit is not implemented, the bit must be read-only and hardwired to Zero | | RW1S | Read-Write-1-to-Set - register bits indicate status when read. The bit may be Set by writing 1b. Writing a 0b to RW1S bits has no effect. If the optional feature that introduces the bit is not implemented, the bit must be read-only and hardwired to Zero | | SDE | Silent Data Error. | | SOC | System On a Chip, also referred as System-On-a-Chip and System-On-Chip. | | SPA | Supervisor Physical Address. See Priv. specification. | | TLB | Translation Lookaside Buffer. | | VA | Virtual Address. See Priv. specification. | | UED | Uncorrected Error Deferred. | | UEC | Uncorrected Error Critical. | | WARL | Write Any values, Reads Legal values: Attribute of a register field that is only defined for a subset of bit encodings, but allow any value to be written while guaranteeing to return a legal value whenever read. | | WPRI | Writes Preserve values, Reads Ignore values: Attribute of a register field that is reserved for future standard use. | RISC-V Debug Specification ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-debug-specification)RISC-V Debug Specification RISC-V Debug Task Group Version v1.0, 2025-02-21: Ratified | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. | | ------------------------------------------------------------------------------------------------ | 4.1. Sdext (ISA Extension) ==================== ## [](#core%5Fdebug)4.1\. Sdext (ISA Extension) This chapter describes the Sdext ISA extension. It must be implemented to make external debug work, and is only useful in conjunction with external debug. Modifications to the RISC-V core to support debug are kept to a minimum. There is a special execution mode (Debug Mode) and a few extra CSRs. The DM takes care of the rest. In order to be compatible with this specification an implementation must implement everything described in this chapter that is not explicitly listed as optional. If Sdext is implemented and Sdtrig is not implemented, then accessing any of the Sdtrig CSRs must raise an illegal instruction exception. ### [](#debugmode)4.1.1\. Debug Mode Debug Mode is a special processor mode used only when a hart is halted for external debugging. Because the hart is halted, there is no forward progress in the normal instruction stream. How Debug Mode is implemented is not specified here. When executing code due to an abstract command, the hart stays in Debug Mode and the following apply: 1. All implemented instructions operate just as they do in M-mode, unless an exception is mentioned in this list. 2. All operations are executed with machine mode privilege, except that additional Debug Mode CSRs are accessible and `mprv` in `mstatus` may be ignored according to [mprven](#dcsr-mprven). Full permission checks, or a relaxed set of permission checks, will apply according to [relaxedpriv](debug%5Fmodule.html#abstractcs-relaxedpriv). 3. All interrupts (including NMI) are masked. 4. Traps don’t take place. Instead, they end execution of the program buffer and the hart remains in Debug Mode. Because they do not trap to M-mode, they do not update registers such as , `mepc`, `mcause`,`mtval`, `mtval2`, and `mtinst`. The same is true for the equivalent privileged registers that are updated when trapping to other modes. Registers that may be updated as part of execution before the exception are allowed to be updated. For example, vector load/store instructions which raise exceptions may partially update the destination register and set `vstart` appropriately. 5. Triggers don’t match or fire. 6. If [stopcount](#dcsr-stopcount) is 0 then counters continue. If it is 1 then counters are stopped. 7. If [stoptime](#dcsr-stoptime) is 0 then `time` continues to update. If it is 1 then `time` will not update. It will resynchronize with `time` after leaving Debug Mode. 8. Instructions that place the hart into a stalled state act as a `nop`. This includes `wfi`, `wrs.sto`, and `wrs.nto`. 9. Almost all instructions that change the privilege mode have UNSPECIFIED behavior. This includes `ecall`, `mret`, `sret`, and `uret`. (To change the privilege mode, the debugger can write [prv](#dcsr-prv) and [v](#dcsr-v) in [dcsr](#csr-dcsr)). The only exception is`ebreak`, which ends execution of the Program Buffer when executed. 10. All control transfer instructions may act as illegal instructions if their destination is in the Program Buffer. If one such instruction acts as an illegal instruction, all such instructions must act as illegal instructions. 11. All control transfer instructions may act as illegal instructions if their destination is outside the Program Buffer. If one such instruction acts as an illegal instruction, all such instructions must act as illegal instructions. 12. Instructions that depend on the value of the PC (e.g. `auipc`) may act as illegal instructions. 13. When the Zicfilp extension is implemented, the `ELP` state is`NO_LP_EXPECTED` and is not updated by any instructions. LPAD instruction executes as a no-op. 14. Effective XLEN is DXLEN. 15. Forward progress is guaranteed. | | When [mprven](#dcsr-mprven), the external debugger can set MPRV and MPP appropriately to have hardware perform memory accesses with the appropriate endianness, address translation, permission checks, and PMP/PMA checks (subject to [relaxedpriv](debug%5Fmodule.html#abstractcs-relaxedpriv)). This is also the only way to access all of physical memory when 34-bit physical addresses are supported on a Sv32 hart. If hardware ties [mprven](#dcsr-mprven) to 0 then the external debugger is expected to simulate all the effects of MPRV, including any extensions that affect memory accesses. For these reasons it is recommended to tie [mprven](#dcsr-mprven) to 1. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#4-1-2-load-reservedstore-conditional-instructions)4.1.2\. Load-Reserved/Store-Conditional Instructions The reservation registered by an `lr` instruction on a memory address may be lost when entering Debug Mode or while in Debug Mode. This means that there may be no forward progress if Debug Mode is entered between`lr` and `sc` pairs. | | This is a behavior that debug users must be aware of. If they have a breakpoint set between a lr and sc pair, or are stepping through such code, the sc may never succeed. Fortunately in general use there will be very few instructions in such a sequence, and anybody debugging it will quickly notice that the reservation is not occurring. The solution in that case is to set a breakpoint on the first instruction after the sc and run to it. A higher level debugger may choose to automate this. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#4-1-3-wait-for-interrupt-instruction)4.1.3\. Wait for Interrupt Instruction If halt is requested while `wfi` is executing, then the hart must leave the stalled state, completing this instruction’s execution, and then enter Debug Mode. ### [](#4-1-4-wait-on-reservation-set-instructions)4.1.4\. Wait-on-Reservation-Set Instructions If halt is requested while `wrs.sto` or `wrs.nto` is executing, then the hart must leave the stalled state, completing this instruction’s execution, and then enter Debug Mode. ### [](#4-1-5-single-step)4.1.5\. Single Step #### [](#stepbit)4.1.5.1\. Step Bit In Dcsr This method is only available to external debuggers, and is the preferred way to single step. An external debugger can cause a halted hart to execute a single instruction or trap and then re-enter Debug Mode by setting [step](#dcsr-step) before resuming. If [step](#dcsr-step) is set when a hart resumes then it will single step, regardless of the reason for resuming. If control is transferred to a trap handler while executing the instruction, then Debug Mode is re-entered immediately after the PC is changed to the trap handler, and the appropriate `tval` and `cause`registers are updated. In this case none of the trap handler is executed, and if the cause was a pending interrupt no instructions might be executed at all. If executing or fetching the instruction causes a trigger to fire with action=1, Debug Mode is re-entered immediately after that trigger has fired. In that case [cause](#dcsr-cause) is set to 2 (trigger) instead of 4 (single step). Whether the instruction is executed or not depends on the specific configuration of the trigger. If the instruction that is executed causes the PC to change to an address where an instruction fetch causes an exception, that exception does not occur until the next time the hart is resumed. Similarly, a trigger at the new address does not fire until the hart actually attempts to execute that instruction. If the instruction being stepped over would normally stall the hart, then instead the instruction is treated as a `nop`. This includes `wfi`,`wrs.sto`, and `wrs.nto`. #### [](#stepicount)4.1.5.2\. Icount Trigger Native debuggers won’t have access to [dcsr](#csr-dcsr), but can use the [icount](Sdtrig.html#csr-icount) trigger by setting [count](Sdtrig.html#icount-count) to 1. This approach does have some limitations: 1. Interrupts will fire as usual. Debuggers that want to disable interrupts while stepping must disable them by changing `mstatus`, and specially handle instructions that read `mstatus`. 2. `wfi` instructions are not treated specially and might take a very long time to complete. This mechanism cleanly supports a system which supports multiple privilege levels, where the OS or a debug stub runs in M-Mode while the program being debugged runs in a less privileged mode. Systems that only support M-Mode can use [icount](Sdtrig.html#csr-icount) as well, but [count](Sdtrig.html#icount-count) must be able to count several instructions (depending on the software implementation). See[\[nativestep\]](#nativestep). ### [](#4-1-6-reset)4.1.6\. Reset If the halt signal (driven by the hart’s halt request bit in the Debug Module) or [hasresethaltreq](debug%5Fmodule.html#dmstatus-hasresethaltreq) are asserted when a hart comes out of reset, the hart must enter Debug Mode before executing any instructions, but after performing any initialization that would usually happen before the first instruction is executed. ### [](#4-1-7-halt)4.1.7\. Halt When a hart halts: 1. [cause](#dcsr-cause) is updated. 2. [prv](#dcsr-prv) and [v](#dcsr-v) are set to reflect current privilege mode and virtualization mode. 3. If the Zicfilp extension is implemented, [pelp](#dcsr-pelp) is set to the current`ELP` state and `ELP` is set to `NO_LP_EXPECTED` 4. [dpc](#csr-dpc) is set to the next instruction that should be executed. 5. If the current instruction can be partially executed and should be restarted to complete, then the relevant state for that is updated. E.g. if a halt occurs during a partially executed vector instruction, then`vstart` is updated, and [dpc](#csr-dpc) is updated to the address of the partially executed instruction. This is analogous to how vector instructions behave for exceptions. 6. The hart enters Debug Mode. ### [](#4-1-8-resume)4.1.8\. Resume When a hart resumes: 1. `pc` changes to the value stored in [dpc](#csr-dpc). 2. The current privilege mode and virtualization mode are changed to that specified by [prv](#dcsr-prv) and [v](#dcsr-v). 3. If the Zicfilp extension is enabled at the new privilege mode, the current`ELP` state is changed to that specified by [pelp](#dcsr-pelp) else it is set to`NO_LP_EXPECTED`. [pelp](#dcsr-pelp) is set to `NO_LP_EXPECTED`. 4. If the new privilege mode is less privileged than M-mode, `MPRV` in `mstatus` is cleared. 5. If the Smdbltrp extension is implemented and the new privilege mode is not M, then the `MDT` bit is set to 0. 6. If the Ssdbltrp extension is implemented and the new privilege mode is U, VS, or VU, then `sstatus.SDT` is set to 0\. Additionally, if it is VU, then`vsstatus.SDT` is also set to 0. 7. The hart is no longer in debug mode. ### [](#debreg)4.1.9\. Core Debug Registers The supported Core Debug Registers must be implemented for each hart that can be debugged. They are CSRs, accessible using the RISC-V `csr`opcodes and optionally also using abstract debug commands. Attempts to access an unimplemented Core Debug Register raise an illegal instruction exception. These registers are only accessible from Debug Mode. __Table 1\. Core Debug Registers__ | Address | Name | Section | | ------- | ------------------------------------------------------ | ---------------------------------------------------------------- | | 0x7b0 | Debug Control and Status ([dcsr](#csr-dcsr)) | [Debug Control and Status (dcsr, at 0x7b0)](#csr-dcsr) | | 0x7b1 | Debug PC ([dpc](#csr-dpc)) | [Debug PC (dpc, at 0x7b1)](#csr-dpc) | | 0x7b2 | Debug Scratch Register 0 ([dscratch0](#csr-dscratch0)) | [Debug Scratch Register 0 (dscratch0, at 0x7b2)](#csr-dscratch0) | | 0x7b3 | Debug Scratch Register 1 ([dscratch1](#csr-dscratch1)) | [Debug Scratch Register 1 (dscratch1, at 0x7b3)](#csr-dscratch1) | #### [](#csr-dcsr)Debug Control and Status (dcsr, at 0x7b0) Upon entry into Debug Mode, [v](#dcsr-v) and [prv](#dcsr-prv) are updated with the privilege level the hart was previously in, and [cause](#dcsr-cause)is updated with the reason for Debug Mode entry. Other than these fields and [nmip](#dcsr-nmip), the other fields of [dcsr](#csr-dcsr) are only writable by the external debugger. [Table 2](#tab:dcsrcausepriority) shows the priorities of reasons for entering Debug Mode. Implementations should implement priorities as shown in the table. For compatibility with old versions of this spec, resethaltreq and haltreq are allowed to be at different positions than shown as long as: 1. resethaltreq is higher priority than haltreq 2. the relative order of the other four causes is maintained __Table 2\. Priority of reasons for entering Debug Mode from highest to lowest.__ | [cause](#dcsr-cause) encoding | Cause | | ----------------------------- | --------------------------------------------------------------------- | | 5 | resethaltreq | | 6 | halt group | | 3 | haltreq | | 2 | trigger (See [\[tab:priority\]](#tab:priority) for detailed priority) | | 1 | ebreak | | 4 | step | | | Note that mcontrol/mcontrol6 triggers which fire after the instruction which hit the trigger are considered to be high priority causes on the subsequent instruction. Therefore, an execute trigger with timing=after on an ebreak instruction is lower priority than the ebreak itself because the trigger will fire after the ebreak instruction. For the same reason, if a single instruction is stepped with both icount and [step](#dcsr-step) then the [step](#dcsr-step) has priority. See [\[tab:priority\]](#tab:priority) for the relative priorities of triggers with respect to the ebreak instruction. Most multi-hart implementations will probably hardwire [stoptime](#dcsr-stoptime)to 0, as the implementation can get complicated and the benefit is small. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | This CSR is read/write. ![Diagram](_images/diag-e8e0e041c424fb631a3d64f836cea68d767b7c5a.svg) ![Diagram](_images/diag-fc38fb790946fa50a681d5e00797b6a7a30198f3.svg) | Field | Description | Access | Reset | | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------- | ------ | | debugver | 0 (none): There is no debug support. 4 (1.0): Debug support exists as it is described in this document. 15 (custom): There is debug support, but it does not conform to any available version of this spec. | **R** | Preset | | extcause | When [cause](#dcsr-cause) is 7, this optional field contains the value of a more specific halt reason than "other." Otherwise it contains 0. 0 (critical error): The hart entered a critical error state, as defined in the Smdbltrp extension. All other values are reserved for future versions of this spec, or for use by other RISC-V extensions. | **R** | 0 | | cetrig | This bit is part of Smdbltrp and only exists when that extension is implemented. 0 (disabled): A hart in a critical error state does not enter Debug Mode but instead asserts the critical-error signal to the platform. 1 (enabled): A hart in a critical error state enters Debug Mode instead of asserting the critical-error signal to the platform. Upon such entry into Debug Mode, the cause field is set to 7, and the extcause field is set to 0, indicating a critical error triggered the Debug Mode entry. This cause has the highest priority among all reasons for entering Debug Mode. Resuming from Debug Mode following an entry from the critical error state returns the hart to the critical error state. When [cetrig](#dcsr-cetrig) is 1, resuming from Debug Mode following an entry due to a critical error will result in an immediate re-entry into Debug Mode due to the critical error. The debugger may resume with [cetrig](#dcsr-cetrig) set to 0 to allow the platform defined actions on critical-error signal to occur. Other possible actions include initiating a hart or platform reset using the Debug Module reset control. | **WARL** | 0 | | pelp | This bit is part of Zicfilp and only exists when that extension is implemented. 0 (NO\_LP\_EXPECTED): No landing pad instruction expected. 1 (LP\_EXPECTED): A landing pad instruction is expected. | **WARL** | 0 | | ebreakvs | 0 (exception): ebreak instructions in VS-mode behave as described in the Privileged Spec. 1 (debug mode): ebreak instructions in VS-mode enter Debug Mode. This bit is hardwired to 0 if the hart does not support virtualization mode. | **WARL** | 0 | | ebreakvu | 0 (exception): ebreak instructions in VU-mode behave as described in the Privileged Spec. 1 (debug mode): ebreak instructions in VU-mode enter Debug Mode. This bit is hardwired to 0 if the hart does not support virtualization mode. | **WARL** | 0 | | ebreakm | 0 (exception): ebreak instructions in M-mode behave as described in the Privileged Spec. 1 (debug mode): ebreak instructions in M-mode enter Debug Mode. | **R/W** | 0 | | ebreaks | 0 (exception): ebreak instructions in S-mode behave as described in the Privileged Spec. 1 (debug mode): ebreak instructions in S-mode enter Debug Mode. This bit is hardwired to 0 if the hart does not support S-mode. | **WARL** | 0 | | ebreaku | 0 (exception): ebreak instructions in U-mode behave as described in the Privileged Spec. 1 (debug mode): ebreak instructions in U-mode enter Debug Mode. This bit is hardwired to 0 if the hart does not support U-mode. | **WARL** | 0 | | stepie | 0 (interrupts disabled): Interrupts (including NMI) are disabled during single stepping with [step](#dcsr-step) set. This value should be supported. 1 (interrupts enabled): Interrupts (including NMI) are enabled during single stepping with [step](#dcsr-step) set. Implementations may hard wire this bit to 0\. In that case interrupt behavior can be emulated by the debugger. The debugger must not change the value of this bit while the hart is running. | **WARL** | 0 | | stopcount | 0 (normal): Increment counters as usual. 1 (freeze): Don’t increment any hart-local counters while in Debug Mode or on ebreak instructions that cause entry into Debug Mode. These counters include the instret CSR. On single-hart corescycle should be stopped, but on multi-hart cores it must keep incrementing. An implementation may hardwire this bit to 0 or 1. | **WARL** | Preset | | stoptime | 0 (normal): time continues to reflect mtime. 1 (freeze): time is frozen at the time that Debug Mode was entered. When leaving Debug Mode, time will reflect the latest value of mtime again. While all harts have [stoptime](#dcsr-stoptime)\=1 and are in Debug Mode,mtime is allowed to stop incrementing. An implementation may hardwire this bit to 0 or 1. | **WARL** | Preset | | cause | Explains why Debug Mode was entered. When there are multiple reasons to enter Debug Mode in a single cycle, hardware should set [cause](#dcsr-cause) to the cause with the highest priority. See [Table 2](#tab:dcsrcausepriority) for priorities. 1 (ebreak): An ebreak instruction was executed. 2 (trigger): A Trigger Module trigger fired with action=1. 3 (haltreq): The debugger requested entry to Debug Mode using [haltreq](debug%5Fmodule.html#dmcontrol-haltreq). 4 (step): The hart single stepped because [step](#dcsr-step) was set. 5 (resethaltreq): The hart halted directly out of reset due to \`resethaltreq\` It is also acceptable to report 3 when this happens. 6 (group): The hart halted because it’s part of a halt group. Harts may report 3 for this cause instead. 7 (other): The hart halted for a reason other than the ones mentioned above.[extcause](#dcsr-extcause) may contain a more specific reason. | **R** | 0 | | v | Extends the prv field with the virtualization mode the hart was operating in when Debug Mode was entered. The encoding is described in [Table 5](#tab:privmode). A debugger can change this value to change the hart’s virtualization mode when exiting Debug Mode. This bit is hardwired to 0 on harts that do not support virtualization mode. | **WARL** | 0 | | mprven | 0 (disabled): mprv in mstatus is ignored in Debug Mode. 1 (enabled): mprv in mstatus takes effect in Debug Mode. Implementing this bit is optional. It may be tied to either 0 or 1. | **WARL** | Preset | | nmip | When set, there is a Non-Maskable-Interrupt (NMI) pending for the hart. Since an NMI can indicate a hardware error condition, reliable debugging may no longer be possible once this bit becomes set. This is implementation-dependent. | **R** | 0 | | step | When set and not in Debug Mode, the hart will only execute a single instruction and then enter Debug Mode. See [4.1.5.1\. Step Bit In Dcsr](#stepbit)for details. The debugger must not change the value of this bit while the hart is running. | **R/W** | 0 | | prv | Contains the privilege mode the hart was operating in when Debug Mode was entered. The encoding is described in [Table 5](#tab:privmode). A debugger can change this value to change the hart’s privilege mode when exiting Debug Mode. Not all privilege modes are supported on all harts. If the encoding written is not supported or the debugger is not allowed to change to it, the hart may change to any supported privilege mode. | **WARL** | 3 | #### [](#csr-dpc)Debug PC (dpc, at 0x7b1) Upon entry to debug mode, [dpc](#csr-dpc) is updated with the virtual address of the next instruction to be executed. The behavior is described in more detail in [Table 3](#tab:dpc). __Table 3\. Virtual address in DPC.__ | Cause | Virtual Address in DPC | | -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ebreak | Address of the ebreak instruction | | single step | Address of the instruction that would be executed next if no debugging was going on. Ie. pc \+ 4 for 32-bit instructions that don’t change program flow, the destination PC on taken jumps/branches, etc. | | trigger module | The address of the next instruction to be executed at the time that debug mode was entered. If the trigger is [mcontrol](Sdtrig.html#csr-mcontrol) and [timing](Sdtrig.html#mcontrol-timing) is 0 or if the trigger is[mcontrol6](Sdtrig.html#csr-mcontrol6) and hit1 is 0, this corresponds to the address of the instruction which caused the trigger to fire. | | halt request | Address of the next instruction to be executed at the time that debug mode was entered. | Executing the Program Buffer may cause the value of [dpc](#csr-dpc) to become UNSPECIFIED. If that is the case, it must be possible to read/write[dpc](#csr-dpc) using an abstract command with [postexec](debug%5Fmodule.html#accessregister-postexec) not set. The debugger must attempt to save [dpc](#csr-dpc) between halting and executing a Program Buffer, and then restore [dpc](#csr-dpc) before leaving Debug Mode. | | Allowing [dpc](#csr-dpc) to become UNSPECIFIED upon Program Buffer execution allows for direct implementations that don’t have a separate PC register, and do need to use the PC when executing the Program Buffer. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If the Access Register abstract command supports reading [dpc](#csr-dpc) while the hart is running, then the value read should be the address of a recently executed instruction. If the Access Register abstract command supports writing [dpc](#csr-dpc) while the hart is running, then the executing program should jump to the written address shortly after the write occurs. The writability of [dpc](#csr-dpc) follows the same rules as `mepc` as defined in the Privileged Spec. In particular, [dpc](#csr-dpc) must be able to hold all valid virtual addresses and the writability of the low bits depends on IALIGN. When resuming, the hart’s PC is updated to the virtual address stored in[dpc](#csr-dpc). A debugger may write [dpc](#csr-dpc) to change where the hart resumes. This CSR is read/write. ![Diagram](_images/diag-6de3af4b6e5c401fae75e901c91b6bbb84e8e6b6.svg) #### [](#csr-dscratch0)Debug Scratch Register 0 (dscratch0, at 0x7b2) Optional scratch register that can be used by implementations that need it. A debugger must not write to this register unless [hartinfo](debug%5Fmodule.html#dm-hartinfo)explicitly mentions it (the Debug Module may use this register internally). This CSR is read/write. ![Diagram](_images/diag-9303aeb8c1e4714d323a0a68125effd4bb96b35c.svg) #### [](#csr-dscratch1)Debug Scratch Register 1 (dscratch1, at 0x7b3) Optional scratch register that can be used by implementations that need it. A debugger must not write to this register unless [hartinfo](debug%5Fmodule.html#dm-hartinfo)explicitly mentions it (the Debug Module may use this register internally). This CSR is read/write. ![Diagram](_images/diag-dce304802bfb65ef5cfccd27ab0431dad6cde921.svg) ### [](#virtreg)4.1.10\. Virtual Debug Registers A virtual register is one that doesn’t exist directly in the hardware, but that the debugger exposes as if it does. Debug software should implement them, but hardware can skip this section. Virtual registers exist to give users access to functionality that’s not part of standard debuggers without requiring them to carefully modify debug registers while the debugger is also accessing those same registers. __Table 4\. Virtual Core Debug Registers__ | Address | Name | Section | | ------- | ----------------------------------- | ----------------------------------------------- | | virtual | Privilege Mode ([priv](#virt-priv)) | [Privilege Mode (priv, at virtual)](#virt-priv) | #### [](#virt-priv)Privilege Mode (priv, at virtual) Users can read this register to inspect the privilege mode that the hart was running in when the hart halted. Users can write this register to change the privilege mode that the hart will run in when it resumes. This register contains [prv](#dcsr-prv) and [v](#dcsr-v) from [dcsr](#csr-dcsr), but in a place that the user is expected to access. The user should not access [dcsr](#csr-dcsr) directly, because doing so might interfere with the debugger. __Table 5\. Privilege Mode and Virtualization Mode Encoding__ | H extension supported | v | prv | Abbreviation | Name | | --------------------- | - | --- | ------------ | ---------------------------------- | | No | 0 | 0 | U-mode | User mode | | No | 0 | 1 | S-mode | Supervisor mode | | No | 0 | 3 | M-mode | Machine mode | | Yes | 0 | 0 | U-mode | User mode | | Yes | 0 | 1 | HS-mode | Hypervisor-enabled supervisor mode | | Yes | 0 | 3 | M-mode | Machine mode | | Yes | 1 | 0 | VU-mode | Virtual user mode | | Yes | 1 | 1 | VS-mode | Virtual supervisor mode | ![Diagram](_images/diag-783580259a36aeb92c9e6129f0218b3043c685f5.svg) | Field | Description | Access | Reset | | ----- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ----- | | v | Contains the virtualization mode the hart was operating in when Debug Mode was entered. The encoding is described in [Table 5](#tab:privmode), and matches the virtualization mode encoding from the Privileged Spec. A user can write this value to change the hart’s virtualization mode when exiting Debug Mode. | **WARL** | 0 | | prv | Contains the privilege mode the hart was operating in when Debug Mode was entered. The encoding is described in [Table 5](#tab:privmode), and matches the privilege mode encoding from the Privileged Spec. A user can write this value to change the hart’s privilege mode when exiting Debug Mode. | **R/W** | 0 | 5.1. Sdtrig (ISA Extension) ==================== ## [](#trigger)5.1\. Sdtrig (ISA Extension) This chapter describes the Sdtrig ISA extension, which can be implemented independently of functionality described in the other chapters. It consists exclusively of the Trigger Module (TM). Triggers can cause a breakpoint exception, entry into Debug Mode, or a trace action without having to execute a special instruction. This makes them invaluable when debugging code from ROM. They can trigger on execution of instructions at a given memory address, or on the address/data in loads/stores. If Sdtrig is implemented, the Trigger Module must support at least one trigger. Accessing trigger CSRs that are not used by any of the implemented triggers must result in an illegal instruction exception. M-Mode and Debug Mode accesses to trigger CSRs that are used by any of the implemented triggers must succeed, regardless of the current type of the currently selected trigger. A trigger matches when the conditions that it specifies (e.g. a load from a specific address) are met. A trigger fires when a trigger that matches performs the action configured for that trigger. Triggers do not fire while in Debug Mode. ### [](#5-1-1-enumeration)5.1.1\. Enumeration Each trigger may support a variety of features. A debugger can build a list of all triggers and their features as follows: 1. Write 0 to [tselect](#csr-tselect). If this results in an illegal instruction exception, then there are no triggers implemented. 2. Read back [tselect](#csr-tselect) and check that it contains the written value. If not, exit the loop. 3. Read [tinfo](#csr-tinfo). 4. If that caused an exception, the debugger must read [tdata1](#csr-tdata1) to discover the type. (If [type](#tdata1-type) is 0, this trigger doesn’t exist. Exit the loop.) 5. If [info](#tinfo-info) is 1, this trigger doesn’t exist. Exit the loop. 6. Otherwise, the selected trigger supports the types discovered in [info](#tinfo-info). 7. Repeat, incrementing the value in [tselect](#csr-tselect). | | The above algorithm reads back [tselect](#csr-tselect) so that implementations which have triggers only need to implement bits of [tselect](#csr-tselect). The algorithm checks [tinfo](#csr-tinfo) and [type](#tdata1-type) in case the implementation has bits of [tselect](#csr-tselect) but fewer than triggers. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#5-1-2-actions)5.1.2\. Actions Triggers can be configured to take one of several actions when they fire. [Table 1](#tab:action) lists all options. __Table 1\. [action](#mcontrol-action) encoding__ | Value | Description | | ----- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | Raise a breakpoint exception. (Used when software wants to use the trigger module without an external debugger attached.) xepc must contain the virtual address of the next instruction that must be executed to preserve the program flow. | | 1 | Enter Debug Mode. [dpc](Sdext.html#csr-dpc) must contain the virtual address of the next instruction that must be executed to preserve the program flow.This action is only legal when the trigger’s [dmode](#mcontrol-dmode) is 1\. Since[tdata1](#csr-tdata1) is WARL, hardware must prevent it from containing [dmode](#tdata1-dmode)\=0 and action=1.This action can only be supported if Sdext is implemented on the hart. | | 2 | Trace on, described in the trace specification. | | 3 | Trace off, described in the trace specification. | | 4 | Trace notify, described in the trace specification. | | 5 | Reserved for use by the trace specification. | | 8 - 9 | Send a signal to TM external trigger output 0 or 1 (respectively). | | other | Reserved for future use. | | | Actions 8 and 9 are intended to increment custom event counters, but these signals could also be brought to outputs for use by external logic. | | ------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#5-1-3-priority)5.1.3\. Priority [Table 2](#tab:priority) lists the synchronous exceptions from the Privileged Spec, and where the various types of triggers fit in. The first 3 columns come from the Privileged Spec, and the final column shows where triggers fit in. Priorities in the table are separated by horizontal lines, so e.g. etrigger and itrigger have the same priority. If this table contradicts the table in the Privileged Spec, then the latter takes precedence. This table only applies if triggers are precise. Otherwise triggers will fire some indeterminate time after the event, and the priority is irrelevant. When triggers are chained, the priority is the lowest priority of the triggers in the chain. __Table 2\. Synchronous exception priority in decreasing priority order.__ | Priority | Exception Code | Description | Trigger | | -------------------- | -------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ | ------- | | _Highest_ | 3333 | etriggericountitriggermcontrol/mcontrol6 after (on previous instruction) | | | 3 | Instruction address breakpoint | mcontrol/mcontrol6 execute address before | | | 12, 20, 1 | During instruction address translation: First encountered page fault, guest-page fault, or access fault | | | | 1 | With physical address for instruction: Instruction access fault | | | | 3 | mcontrol/mcontrol6 execute data before | | | | 22208, 9, 10, 1133 | Illegal instructionVirtual instructionInstruction address misalignedEnvironment callEnvironment breakLoad/Store/AMO address breakpoint | mcontrol/mcontrol6 load/store address before, store data before | | | 4, 6 | Optionally: Load/Store/AMO address misaligned | | | | 13, 15, 21, 23, 5, 7 | During address translation for an explicit memory access: First encountered page fault, guest-page fault, or access fault | | | | 5, 7 | With physical address for an explicit memory access: Load/store/AMO access fault | | | | 4, 6 | If not higher priority: Load/store/AMO address misaligned | | | | _Lowest_ | 3 | mcontrol/mcontrol6 load data before | | When multiple triggers in the same priority fire at once, [hit](#mcontrol-hit) (if implemented) is set for all of them. If more than one of these triggers has [action](#mcontrol-action)\=0 then `tval` is updated in accordance with one of them, but which one is UNSPECIFIED . If one of these triggers has the "enter Debug Mode" action (1) and another trigger has the "raise a breakpoint exception" action (0), the preferred behavior is to have both actions take place. It is implementation-dependent which of the two happens first. This ensures both that the presence of an external debugger doesn’t affect execution and that a trigger set by user code doesn’t affect the external debugger. If this is not implemented, then the hart must enter Debug Mode and ignore the breakpoint exception. In the latter case, [hit](#mcontrol-hit) of the trigger whose action is 0 must still be set, giving a debugger an opportunity to handle this case. Since triggers that have an action other than 0 or 1 don’t affect the execution of the hart, they are not mentioned in the priority table. Such triggers fire independently from those that have an action of 0 or 1. ### [](#nativetrigger)5.1.4\. Native Triggers Triggers can be used for native debugging when [action](#mcontrol-action)\=0\. If supported by the hart and desired by the debugger, triggers will often be programmed to have [m](#mcontrol-m)\=0 so that when they fire they cause a breakpoint exception to trap to a more privileged mode. That breakpoint exception can either be taken in M-mode or it can be delegated to a less privileged mode. However, it is possible for triggers to fire in the same mode that the resulting exception will be handled in. In these cases such a trigger may cause a breakpoint exception while already in a trap handler. This might leave the hart unable to resume normal execution because state such as `mcause` and `mepc` would be overwritten. | | In particular, when [action](#mcontrol-action)\=0: mcontrol and mcontrol6 triggers with [m](#mcontrol-m)\=1 can cause a breakpoint exception that is taken from M-mode to M-mode (regardless of delegation). mcontrol and mcontrol6 triggers with [s](#mcontrol-s)\=1 can cause a breakpoint exception that is taken from S-mode to S-mode if medeleg \[3\]=1. mcontrol6 triggers with [vs](#mcontrol6-vs)\=1 can cause a breakpoint exception that is taken from VS-mode to VS-mode if medeleg \[3\]=1 and hedeleg \[3\]=1 . icount triggers with [m](#mcontrol-m)\=1can cause a breakpoint exception that is taken from M-mode to M-mode (regardless of delegation). icount triggers with [s](#mcontrol-s)\=1 can cause a breakpoint exception that is taken from S-mode to S-mode if medeleg \[3\]=1 . icount triggers with [vs](#mcontrol6-vs)\=1 can cause a breakpoint exception that is taken from VS-mode to VS-mode if medeleg \[3\]=1 and hedeleg \[3\]=1. etrigger and itrigger triggers will always be taken from a trap handler before the first instruction of the handler. If etrigger/itrigger is set to trigger on exception/interrupt X and if X is delegated to mode Y then the trigger will cause a breakpoint exception that is taken from mode Y to mode Y unless breakpoint exceptions are delegated to a more privileged mode than Y. tmexttrigger triggers are asynchronous and may occur in any mode and at any time. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Harts that support triggers with [action](#mcontrol-action)\=0 should implement one of the following two solutions to solve the problem of reentrancy: 1. The hardware prevents triggers with [action](#mcontrol-action)\=0 from matching or firing while in M-mode and while `MIE` in `mstatus` is 0\. If `medeleg` \[3\]=1 then it prevents triggers with [action](#mcontrol-action)\=0 from matching or firing while in S-mode and while `SIE` in `sstatus` is 0\. If `medeleg` \[3\]=1 and `hedeleg` \[3\]=1 then it prevents triggers with [action](#mcontrol-action)\=0 from matching or firing while in VS-mode and while `SIE` in `vstatus` is 0. 2. [mte](#tcontrol-mte) and [mpte](#tcontrol-mpte) in [tcontrol](#csr-tcontrol) is implemented. `medeleg` \[3\] is hard-wired to 0. | | The first option has the limitation that interrupts might be disabled at times when a user still might want triggers to fire. It has the benefit that breakpoints are not required to be handled in M-mode. The second option has the benefit that it only disables triggers during the trap handler, though it requires specific software support for this debug feature in the M-mode trap handlers. It can only work if breakpoints are not delegated to less privileged modes and therefore targets primarily implementations without S-mode. Because [tcontrol](#csr-tcontrol) is not accessible to S-mode, the second option can not be extended to accommodate delegation without adding additional S-mode and VS-mode CSRs. Both options prevent etrigger and itrigger from having any effect on exceptions and interrupts that are handled in M-mode. They also prevent triggering during some initial portion of each handler. Debuggers should use other mechanisms to debug these cases, such as patching the handler or setting a breakpoint on the instruction after MIE is cleared. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#5-1-5-memory-access-triggers)5.1.5\. Memory Access Triggers [mcontrol](#csr-mcontrol) and [mcontrol6](#csr-mcontrol6) both enable triggers on memory accesses. This section describes for both of them how certain corner cases are treated. #### [](#5-1-5-1-a-extension)5.1.5.1\. A Extension If the A extension is supported, then triggers on loads/stores treat them as follows: 1. `lr` instructions are loads. 2. Successful `sc` instructions are stores. 3. It is UNSPECIFIED whether failing `sc` instructions are stores or not. 4. Each AMO instruction is a load for the read portion of the operation. The address is always available to trigger on, although the value loaded might not be, depending on the hardware implementation. 5. Each AMO instruction is a store for the write portion of the operation. The address is always available to trigger on. Whether data store triggers match on AMOs is UNSPECIFIED. 6. If the destination register of any load or AMO is `zero` then it is UNSPECIFIED whether a data load trigger will match. #### [](#5-1-5-2-combined-accesses)5.1.5.2\. Combined Accesses Some instructions lead a hart to perform multiple memory accesses. This includes vector loads and stores, as well as `cm.push` and `cm.pop`instructions. The Trigger Module should match such accesses as if they all happened individually. E.g. a vector load should be treated as if it performed multiple loads of size SEW (selected element width), and`cm.push` should be treated as if it performed multiple stores of size XLEN. #### [](#5-1-5-3-cache-operations)5.1.5.3\. Cache Operations Cache operations are infrequently performed, and code that uses them can have hard-to-find bugs. For the purposes of debug triggers, two classes of cache operations must match as stores: 1. Cache operations that enable software to maintain coherence between otherwise non-coherent implicit and explicit memory accesses. 2. Cache operations that perform block writes of constant data. Only triggers with [size](#mcontrol6-size)\=0 and [select](#mcontrol6-select)\=0 will match. Since cache operations affect multiple addresses, there are multiple possible values to compare against. Implementations must implement one of the following options. From most desirable to least desirable, they are: 1. Every address from the effective address rounded down to the nearest cache block boundary (inclusive) to the effective address rounded up to the nearest cache block boundary (exclusive) is a compare value. 2. The effective address rounded down to the nearest cache block boundary is a compare value. 3. The effective address of the instruction is a compare value. Cache operations encoded as HINTs do not match debug triggers. | | The above language intends to capture the trigger behavior with respect to the cache operations to be introduced in a forthcoming I/D consistency extension. For RISC-V Base Cache Management Operation ISA Extensions 1.0.1, this means the following: cbo.clean, cbo.flush, and cbo.inval match as if they are stores because they affect consistency. cbo.zero matches as if it is a store because it performs a block write of constant data. The prefetch instructions don’t match at all. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#5-1-5-4-address-matches)5.1.5.4\. Address Matches For address matches without a mask, [tdata2](#csr-tdata2) must be able to hold all valid addresses in all supported translation modes. That means that after writing any of these valid addresses, the exact same value XLEN-wide value is read back, including any high bits. An implementation may be able to optimize the storage required, depending on the widest addresses it supports. | | If physical addresses are less than XLEN bits wide, they are zero-extended. If virtual addresses are less than XLEN bits wide, they are sign-extended.[tdata2](#csr-tdata2) must be implemented with enough bits of storage to represent the full range of supported physical and virtual address values when read by software and used by hardware. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ##### [](#5-1-5-4-1-invalid-addresses)5.1.5.4.1\. Invalid Addresses If [tdata2](#csr-tdata2) can hold any invalid addresses, then writes of an invalid address that can not be represented as-is should be converted to a different invalid address that can be represented. For invalid instruction fetch addresses and load and store effective addresses, the compare value may be changed to a different invalid address. In addition, an implementation may choose to inhibit all trigger matching against invalid addresses, especially if there is no support for storage of any invalid address values in tdata2. ### [](#multistate)5.1.6\. Multiple State Change Instructions An instruction that performs multiple architectural state changes (e.g., register updates and/or memory accesses) might cause a trigger to fire at an intermediate point in its execution. As a result, architectural state changes up to that point might have been performed, while subsequent state changes, starting from the event that activated the trigger, might not have been. The definition of such an instruction will specify the order in which architectural state changes take place. Alternatively, it may state that partial execution is not allowed, implying that a mid-execution trigger must prevent any architectural state changes from occurring. Debuggers won’t be aware if an instruction has been partially executed. When they resume execution, they will execute the same instruction once more. Therefore, it’s crucial that partially executing the instruction and then executing it again leaves the hart in a state closely resembling the state it would have been in if the instruction had only been executed once. ### [](#5-1-7-trigger-module-registers)5.1.7\. Trigger Module Registers These registers are CSRs, accessible using the RISC-V `csr` opcodes and optionally also using abstract debug commands. They are the only mechanism to access the triggers. Almost all trigger functionality is optional. All `tdata` registers follow write-any-read-legal semantics. If a debugger writes an unsupported configuration, the register will read back a value that is supported (which may simply be a disabled trigger). This means that a debugger must always read back values it writes to `tdata` registers, unless it already knows what is supported. Writes to one `tdata`register must not modify the contents of other `tdata` registers, nor the configuration of any trigger besides the one that is currently selected. The combination of these rules means that a debugger cannot simply set a trigger by writing [tdata1](#csr-tdata1), then [tdata2](#csr-tdata2), etc. The current value of [tdata2](#csr-tdata2) might not be legal with the new value of [tdata1](#csr-tdata1). To help with this situation, it is guaranteed that writing 0 to [tdata1](#csr-tdata1) disables the trigger, and leaves it in a state where [tdata2](#csr-tdata2) and [tdata3](#csr-tdata3) can be written with any value that makes sense for any trigger type supported by this trigger. As a result, a debugger can write any supported trigger as follows: 1. Write 0 to [tdata1](#csr-tdata1). (This will result in [tdata1](#csr-tdata1) containing a non-zero value, since the register is **WARL**.) 2. Write desired values to [tdata2](#csr-tdata2) and [tdata3](#csr-tdata3). 3. Write desired value to [tdata1](#csr-tdata1). Code that restores CSR context of triggers that might be configured to fire in the current privilege mode must use this same sequence to restore the triggers. This avoids the problem of a partially written trigger firing at a different time than is expected. Attempts to access an unimplemented Trigger Module Register raise an illegal instruction exception. The Trigger Module registers, except [mscontext](#csr-mscontext), [scontext](#csr-scontext), and [hcontext](#csr-hcontext), are only accessible in machine and Debug Mode to prevent untrusted user code from causing entry into Debug Mode without the OS’s permission. In this section XLEN refers to the effective XLEN in the current execution mode. On systems where XLEN values can differ between modes, this is handled as follows. Fields retain their values regardless of XLEN, which only affects where in the register these fields appear (e.g. [type](#tdata1-type)). Some fields are wider when XLEN is 64 than when it is 32 (e.g.[svalue](#textra32-svalue)). The high bits in such fields retain their value but are not readable when XLEN is 32\. A modification of a register when XLEN is 32 clears any inaccessible bits in that register. __Table 3\. Trigger Module Registers__ | Address | Name | Section | | ------- | -------------------------------------------------------- | ------------------------------------------------------------------ | | 0x5a8 | Supervisor Context ([scontext](#csr-scontext)) | [Supervisor Context (scontext, at 0x5a8)](#csr-scontext) | | 0x6a8 | Hypervisor Context ([hcontext](#csr-hcontext)) | [Hypervisor Context (hcontext, at 0x6a8)](#csr-hcontext) | | 0x7a0 | Trigger Select ([tselect](#csr-tselect)) | [Trigger Select (tselect, at 0x7a0)](#csr-tselect) | | 0x7a1 | Trigger Data 1 ([tdata1](#csr-tdata1)) | [Trigger Data 1 (tdata1, at 0x7a1)](#csr-tdata1) | | 0x7a1 | Match Control ([mcontrol](#csr-mcontrol)) | [Match Control (mcontrol, at 0x7a1)](#csr-mcontrol) | | 0x7a1 | Match Control Type 6 ([mcontrol6](#csr-mcontrol6)) | [Match Control Type 6 (mcontrol6, at 0x7a1)](#csr-mcontrol6) | | 0x7a1 | Instruction Count ([icount](#csr-icount)) | [Instruction Count (icount, at 0x7a1)](#csr-icount) | | 0x7a1 | Interrupt Trigger ([itrigger](#csr-itrigger)) | [Interrupt Trigger (itrigger, at 0x7a1)](#csr-itrigger) | | 0x7a1 | Exception Trigger ([etrigger](#csr-etrigger)) | [Exception Trigger (etrigger, at 0x7a1)](#csr-etrigger) | | 0x7a1 | External Trigger ([tmexttrigger](#csr-tmexttrigger)) | [External Trigger (tmexttrigger, at 0x7a1)](#csr-tmexttrigger) | | 0x7a2 | Trigger Data 2 ([tdata2](#csr-tdata2)) | [Trigger Data 2 (tdata2, at 0x7a2)](#csr-tdata2) | | 0x7a3 | Trigger Data 3 ([tdata3](#csr-tdata3)) | [Trigger Data 3 (tdata3, at 0x7a3)](#csr-tdata3) | | 0x7a3 | Trigger Extra (RV32) ([textra32](#csr-textra32)) | [Trigger Extra (RV32) (textra32, at 0x7a3)](#csr-textra32) | | 0x7a3 | Trigger Extra (RV64) ([textra64](#csr-textra64)) | [Trigger Extra (RV64) (textra64, at 0x7a3)](#csr-textra64) | | 0x7a4 | Trigger Info ([tinfo](#csr-tinfo)) | [Trigger Info (tinfo, at 0x7a4)](#csr-tinfo) | | 0x7a5 | Trigger Control ([tcontrol](#csr-tcontrol)) | [Trigger Control (tcontrol, at 0x7a5)](#csr-tcontrol) | | 0x7a8 | Machine Context ([mcontext](#csr-mcontext)) | [Machine Context (mcontext, at 0x7a8)](#csr-mcontext) | | 0x7aa | Machine Supervisor Context ([mscontext](#csr-mscontext)) | [Machine Supervisor Context (mscontext, at 0x7aa)](#csr-mscontext) | #### [](#csr-tselect)Trigger Select (tselect, at 0x7a0) This register determines which trigger is accessible through the other Trigger Module registers. It is optional if no triggers are implemented. The set of accessible triggers must start at 0, and be contiguous. This register is **WARL**. Writes of values greater than or equal to the number of supported triggers may result in a different value in this register than what was written or may point to a trigger where [type](#tdata1-type)\=0\. To verify that what they wrote is a valid index, debuggers can read back the value and check that [tselect](#csr-tselect) holds what they wrote and read [tdata1](#csr-tdata1) to see that [type](#tdata1-type) is non-zero. Since triggers can be used both by Debug Mode and M-mode, the external debugger must restore this register if it modifies it. This CSR is read/write. ![Diagram](_images/diag-5711d1b3f16bdc389e98f7a9197c94c60d7f5ad0.svg) #### [](#csr-tdata1)Trigger Data 1 (tdata1, at 0x7a1) This register provides access to the trigger selected by [tselect](#csr-tselect). The reset values listed here apply to every underlying trigger. This register is optional if no triggers are implemented. Writing 0 to this register must result in a trigger that is disabled. If this trigger supports multiple types, then the hardware should disable it by changing [type](#tdata1-type) to 15. This CSR is read/write. ![Diagram](_images/diag-4e40dfebc9895aaab10769d9c4fddcded81ba0f1.svg) | Field | Description | Access | Reset | | ----- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ------ | | type | 0 (none): There is no trigger at this [tselect](#csr-tselect). 1 (legacy): The trigger is a legacy SiFive address match trigger. These should not be implemented and aren’t further documented here. 2 (mcontrol): The trigger is an address/data match trigger. The remaining bits in this register act as described in [mcontrol](#csr-mcontrol). 3 (icount): The trigger is an instruction count trigger. The remaining bits in this register act as described in [icount](#csr-icount). 4 (itrigger): The trigger is an interrupt trigger. The remaining bits in this register act as described in [itrigger](#csr-itrigger). 5 (etrigger): The trigger is an exception trigger. The remaining bits in this register act as described in [etrigger](#csr-etrigger). 6 (mcontrol6): The trigger is an address/data match trigger. The remaining bits in this register act as described in [mcontrol6](#csr-mcontrol6). This is similar to a type 2 trigger, but provides additional functionality and should be used instead of type 2 in newer implementations. 7 (tmexttrigger): The trigger is a trigger source external to the TM. The remaining bits in this register act as described in [tmexttrigger](#csr-tmexttrigger). 12—​14 (custom): These trigger types are available for non-standard use. 15 (disabled): This trigger is disabled. In this state, [tdata2](#csr-tdata2) and[tdata3](#csr-tdata3) can be written with any value that is supported for any of the types this trigger implements. The remaining bits in this register, except for [dmode](#tdata1-dmode), are ignored. Other values are reserved for future use. | **WARL** | Preset | | dmode | If [type](#tdata1-type) is 0, then this bit is hard-wired to 0. 0 (both): Both Debug and M-mode can write the tdata registers at the selected [tselect](#csr-tselect). 1 (dmode): Only Debug Mode can write the tdata registers at the selected [tselect](#csr-tselect). Writes from other modes are ignored. This bit is only writable from Debug Mode. In ordinary use, external debuggers will always set this bit when configuring a trigger. When clearing this bit, debuggers should also set the action field (whose location depends on [type](#tdata1-type)) to something other than 1. | **WARL** | 0 | | data | If [type](#tdata1-type) is 0, then this field is hard-wired to 0. Trigger-specific data. | **WARL** | Preset | #### [](#csr-tdata2)Trigger Data 2 (tdata2, at 0x7a2) This register provides access to the trigger selected by [tselect](#csr-tselect). The reset values listed here apply to every underlying trigger. Trigger-specific data. It is optional if no implemented triggers use it. If the trigger is disabled, then this register can be written with any value supported by any of the trigger types supported by this trigger. If XLEN is less than DXLEN, writes to this register are sign-extended. This CSR is read/write. ![Diagram](_images/diag-12a12435c776e0719f909da32dc7a4bbac9cf4e0.svg) #### [](#csr-tdata3)Trigger Data 3 (tdata3, at 0x7a3) This register provides access to the trigger selected by [tselect](#csr-tselect). The reset values listed here apply to every underlying trigger. Trigger-specific data. It is optional if no implemented triggers use it. If the trigger is disabled, then this register can be written with any value supported by any of the trigger types supported by this trigger. If XLEN is less than DXLEN, writes to this register are sign-extended. This CSR is read/write. ![Diagram](_images/diag-12a12435c776e0719f909da32dc7a4bbac9cf4e0.svg) #### [](#csr-tinfo)Trigger Info (tinfo, at 0x7a4) This register provides access to the trigger selected by [tselect](#csr-tselect). The reset values listed here apply to every underlying trigger. This register is optional if no triggers are implemented, or if[type](#tdata1-type) is not writable and [version](#tinfo-version) would be 0\. In this case the debugger can read the only supported type from[tdata1](#csr-tdata1). Writing this read/write CSR has no effect. ![Diagram](_images/diag-0dca6420ba09198caf7e4c73fcce5bee42a68e52.svg) | Field | Description | Access | Reset | | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | ------ | | version | Contains the version of the Sdtrig extension implemented. 0 (0): Supports triggers as described in this spec at commit 5a5c078, made on February 2, 2023. In these older versions: [mcontrol6](#csr-mcontrol6) has a timing bit identical to [timing](#mcontrol-timing) [hit0](#mcontrol6-hit0) behaves just as [hit](#mcontrol-hit). [hit1](#mcontrol6-hit1) is read-only 0. Encodings for [size](#mcontrol6-size) for access sizes larger than 64 bits are different. 1 (1): Supports triggers as described in the ratified version 1.0 of this document. | **R** | Preset | | info | One bit for each possible [type](#tdata1-type) enumerated in [tdata1](#csr-tdata1). Bit N corresponds to type N. If the bit is set, then that type is supported by the currently selected trigger. If the currently selected trigger doesn’t exist, this field contains 1. | **R** | Preset | #### [](#csr-tcontrol)Trigger Control (tcontrol, at 0x7a5) This optional register is only accessible in M-mode and Debug Mode and provides various control bits related to triggers. This CSR is read/write. ![Diagram](_images/diag-946427ef3de288b8f4c172a2f65c0edcca867f7f.svg) | Field | Description | Access | Reset | | ----- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ----- | | mpte | M-mode previous trigger enable field. [mpte](#tcontrol-mpte) and [mte](#tcontrol-mte) provide one solution to a problem regarding triggers with action=0 firing in M-mode trap handlers. See[5.1.4\. Native Triggers](#nativetrigger) for more details. When any trap into M-mode is taken, [mpte](#tcontrol-mpte) is set to the value of[mte](#tcontrol-mte). | **WARL** | 0 | | mte | M-mode trigger enable field. 0 (disabled): Triggers with action=0 do not match/fire while the hart is in M-mode. 1 (enabled): Triggers do match/fire while the hart is in M-mode. When any trap into M-mode is taken, [mte](#tcontrol-mte) is set to 0\. When mret is executed, [mte](#tcontrol-mte) is set to the value of [mpte](#tcontrol-mpte). | **WARL** | 0 | #### [](#csr-hcontext)Hypervisor Context (hcontext, at 0x6a8) This optional register may be implemented only if the H extension is implemented. If it is implemented, [mcontext](#csr-mcontext) must also be implemented. This register is only accessible in HS-Mode, M-mode and Debug Mode. If Smstateen is implemented, then accessibility of in HS-Mode is controlled by `mstateenzero[57]`. This register is an alias of the [mcontext](#csr-mcontext) register, providing access to the [hcontext](#mcontext-hcontext) field from HS-Mode. #### [](#csr-scontext)Supervisor Context (scontext, at 0x5a8) This optional register is only accessible in S/HS-mode, VS-mode, M-mode and Debug Mode. Accessibility of this CSR is controlled by `mstateenzero[57]` and`hstateenzero[57]` in the Smstateen extension. Enabling [scontext](#csr-scontext)can be a security risk in a virtualized system with a hypervisor that does not swap [scontext](#csr-scontext). This CSR is read/write. ![Diagram](_images/diag-b0d568a138b6b99a648784f47d9db769327e63f1.svg) | Field | Description | Access | Reset | | ----- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------- | ----- | | data | Supervisor mode software can write a context number to this register, which can be used to set triggers that only fire in that specific context. An implementation may tie any number of high bits in this field to 0\. It’s recommended to implement 16 bits on RV32 and 32 bits on RV64. | **WARL** | 0 | #### [](#csr-mcontext)Machine Context (mcontext, at 0x7a8) This register must be implemented if [hcontext](#csr-hcontext) is implemented, and is optional otherwise. It is only accessible in M-mode and Debug mode. | | [hcontext](#mcontext-hcontext) is primarily useful to set triggers on hypervisor systems that only fire when a given VM is executing. It is also useful in systems where M-Mode implements something like a hypervisor directly. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | This CSR is read/write. ![Diagram](_images/diag-3977696203a65299ef1ed0c31f501fd4d5585511.svg) | Field | Description | Access | Reset | | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ----- | | hcontext | M-Mode or HS-Mode (using [hcontext](#csr-hcontext)) software can write a context number to this register, which can be used to set triggers that only fire in that specific context. An implementation may tie any number of upper bits in this field to 0\. If the H extension is not implemented, it’s recommended to implement 6 bits on RV32 and 13 bits on RV64 (as visible through the[mcontext](#csr-mcontext) register). If the H extension is implemented, it’s recommended to implement 7 bits on RV32 and 14 bits on RV64. | **WARL** | 0 | #### [](#csr-mscontext)Machine Supervisor Context (mscontext, at 0x7aa) This optional register is an alias for [scontext](#csr-scontext). It is only accessible in S/HS-mode, M-mode and Debug Mode. It is included for backward compatibility with version 0.13. | | The encoding of this CSR does not conform to the CSR Address Mapping Convention in the Privileged Spec. It is expected that new implementations will not support this encoding and that new debuggers will not use this CSR if [scontext](#csr-scontext) is available. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#csr-mcontrol)Match Control (mcontrol, at 0x7a1) This register provides access to the trigger selected by [tselect](#csr-tselect). The reset values listed here apply to every underlying trigger. This register is accessible as [tdata1](#csr-tdata1) when [type](#tdata1-type) is 2\. This trigger type is deprecated. It is included for backward compatibility with version 0.13. | | This trigger type only supports a subset of features of the newer[mcontrol6](#csr-mcontrol6). It is expected that new implementations will not support this trigger type and that new debuggers will not use it if[mcontrol6](#csr-mcontrol6) is available. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Address and data trigger implementation are heavily dependent on how the processor core is implemented. To accommodate various implementations, execute, load, and store address/data triggers may fire at whatever point in time is most convenient for the implementation. The debugger may request specific timings as described in [timing](#mcontrol-timing).[Table 4](#tab:hwbp%5Ftiming) suggests timings for the best user experience. A chain of triggers that don’t all have the same [timing](#mcontrol-timing)value will never fire. That means to implement the suggestions in[Table 4](#tab:hwbp%5Ftiming), both timings should be supported on load address triggers that can be chained with a load data trigger. The Privileged Spec says that breakpoint exceptions that occur on instruction fetches, loads, or stores update the `tval` CSR with either zero or the faulting virtual address. The faulting virtual address for an mcontrol trigger with [action](#mcontrol-action)\=0 is the address being accessed and which caused that trigger to fire. If multiple mcontrol triggers are chained then the faulting virtual address is the address which caused any of the chained triggers to fire. If [textra32](#csr-textra32) or [textra64](#csr-textra64) are implemented for this trigger, it only matches when the conditions set there are satisfied. This CSR is read/write. ![Diagram](_images/diag-89390d0910a4ce43c5fcc2860caf5b5f2b848193.svg) ![Diagram](_images/diag-3ac51e9dee8e58363b9704a99185d74df58c2f91.svg) | Field | Description | Access | Reset | | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ------ | | maskmax | Specifies the largest naturally aligned powers-of-two (NAPOT) range supported by the hardware when [match](#mcontrol-match) is 1\. The value is the logarithm base 2 of the number of bytes in that range. A value of 0 indicates [match](#mcontrol-match) 1 is not supported. A value of 63 corresponds to the maximum NAPOT range, which is 263 bytes in size. | **R** | Preset | | sizehi | This field only exists when XLEN is at least 64\. It contains the 2 high bits of the access size. The low bits come from [sizelo](#mcontrol-sizelo). See [sizelo](#mcontrol-sizelo) for how this is used. | **WARL** | 0 | | hit | If this bit is implemented then it must become set when this trigger fires and may become set when this trigger matches. The trigger’s user can set or clear it at any time. It is used to determine which trigger(s) matched. If the bit is not implemented, it is always 0 and writing it has no effect. | **WARL** | 0 | | select | This bit determines the contents of the XLEN-bit compare values. 0 (address): There is at least one compare value and it contains the lowest virtual address of the access. It is recommended that there are additional compare values for the other accessed virtual addresses. (E.g. on a 32-bit read from 0x4000, the lowest address is 0x4000 and the other addresses are 0x4001, 0x4002, and 0x4003.) 1 (data): There is exactly one compare value and it contains the data value loaded or stored, or the instruction executed. Any bits beyond the size of the data access will contain 0. | **WARL** | 0 | | timing | 0 (before): The action for this trigger will be taken just before the instruction that triggered it is retired, but after all preceding instructions are retired. xepc or [dpc](Sdext.html#csr-dpc) (depending on [action](#mcontrol-action)) must be set to the virtual address of the instruction that matched. If this is combined with [load](#mcontrol-load) and[select](#mcontrol-select)\=1 then a memory access will be performed (including any side effects of performing such an access) even though the load will not update its destination register. Debuggers should consider this when setting such breakpoints on, for example, memory-mapped I/O addresses. If an instruction matches this trigger and the instruction performs multiple memory accesses, it is UNSPECIFIED which memory accesses have completed before the trigger fires. 1 (after): The action for this trigger will be taken after the instruction that triggered it is retired. It should be taken before the next instruction is retired, but it is better to implement triggers imprecisely than to not implement them at all. xepc or[dpc](Sdext.html#csr-dpc) (depending on [action](#mcontrol-action)) must be set to the virtual address of the next instruction that must be executed to preserve the program flow. Most hardware will only implement one timing or the other, possibly dependent on [select](#mcontrol-select), [execute](#mcontrol-execute),[load](#mcontrol-load), and [store](#mcontrol-store). This bit primarily exists for the hardware to communicate to the debugger what will happen. Hardware may implement the bit fully writable, in which case the debugger has a little more control. Data load triggers with [timing](#mcontrol-timing) of 0 will result in the same load happening again when the debugger lets the hart run. For data load triggers, debuggers must first attempt to set the breakpoint with[timing](#mcontrol-timing) of 1. If a trigger with [timing](#mcontrol-timing) of 0 matches, it is implementation-dependent whether that prevents a trigger with[timing](#mcontrol-timing) of 1 matching as well. | **WARL** | 0 | | sizelo | This field contains the 2 low bits of the access size. The high bits come from [sizehi](#mcontrol-sizehi). The combined value is interpreted as follows: 0 (any): The trigger will attempt to match against an access of any size. The behavior is only well-defined if [select](#mcontrol-select)\=0, or if the access size is XLEN. 1 (8bit): The trigger will only match against 8-bit memory accesses. 2 (16bit): The trigger will only match against 16-bit memory accesses or execution of 16-bit instructions. 3 (32bit): The trigger will only match against 32-bit memory accesses or execution of 32-bit instructions. 4 (48bit): The trigger will only match against execution of 48-bit instructions. 5 (64bit): The trigger will only match against 64-bit memory accesses or execution of 64-bit instructions. 6 (80bit): The trigger will only match against execution of 80-bit instructions. 7 (96bit): The trigger will only match against execution of 96-bit instructions. 8 (112bit): The trigger will only match against execution of 112-bit instructions. 9 (128bit): The trigger will only match against 128-bit memory accesses or execution of 128-bit instructions. An implementation must support the value of 0, but all other values are optional. When an implementation supports address triggers ([select](#mcontrol-select)\=0), it is recommended that those triggers support every access size that the hart supports, as well as for every instruction size that the hart supports. Implementations such as RV32D or RV64V are able to perform loads and stores that are wider than XLEN. Custom extensions may also support instructions that are wider than XLEN. Because[tdata2](#csr-tdata2) is of size XLEN, there is a known limitation that data value triggers ([select](#mcontrol-select)\=1) can only be supported for access sizes up to XLEN bits. When an implementation supports data value triggers ([select](#mcontrol-select)\=1), it is recommended that those triggers support every access size up to XLEN that the hart supports, as well as for every instruction length up to XLEN that the hart supports. | **WARL** | 0 | | action | The action to take when the trigger fires. The values are explained in [Table 1](#tab:action). | **WARL** | 0 | | chain | 0 (disabled): When this trigger matches, the configured action is taken. 1 (enabled): While this trigger does not match, it prevents the trigger with the next index from matching. A trigger chain starts on the first trigger with chain\=1 after a trigger with chain\=0, or simply on the first trigger if that has chain\=1\. It ends on the first trigger after that which haschain\=0\. This final trigger is part of the chain. The action on all but the final trigger is ignored. The action on that final trigger will be taken if and only if all the triggers in the chain match at the same time. Debuggers should not terminate a chain with a trigger with a different type. It is undefined when exactly such a chain fires. Because [chain](#mcontrol-chain) affects the next trigger, hardware must zero it in writes to [mcontrol](#csr-mcontrol) that set [dmode](#tdata1-dmode) to 0 if the next trigger has[dmode](#tdata1-dmode) of 1\. In addition hardware should ignore writes to [mcontrol](#csr-mcontrol) that set[dmode](#tdata1-dmode) to 1 if the previous trigger has both [dmode](#tdata1-dmode) of 0 and[chain](#mcontrol-chain) of 1\. Debuggers must avoid the latter case by checking[chain](#mcontrol-chain) on the previous trigger if they’re writing [mcontrol](#csr-mcontrol). Implementations that wish to limit the maximum length of a trigger chain (eg. to meet timing requirements) may do so by zeroing[chain](#mcontrol-chain) in writes to [mcontrol](#csr-mcontrol) that would make the chain too long. | **WARL** | 0 | | match | 0 (equal): Matches when any compare value equals [tdata2](#csr-tdata2). 1 (napot): Matches when the top M bits of any compare value match the topM bits of [tdata2](#csr-tdata2).M is XLEN-1 minus the index of the least-significant bit containing 0 in [tdata2](#csr-tdata2). Debuggers should only write values to [tdata2](#csr-tdata2) such that M \+ [maskmax](#mcontrol-maskmax) ≥ XLENand M \> 0, otherwise it’s undefined on what conditions the trigger will match. 2 (ge): Matches when any compare value is greater than (unsigned) or equal to [tdata2](#csr-tdata2). 3 (lt): Matches when any compare value is less than (unsigned)[tdata2](#csr-tdata2). 4 (mask low): Matches when of any compare value equals of [tdata2](#csr-tdata2) after of the compare value is ANDed withXLEN-1: of [tdata2](#csr-tdata2). 5 (mask high): Matches when XLEN-1: of any compare value equals of [tdata2](#csr-tdata2) afterXLEN-1: of the compare value is ANDed withXLEN-1: of [tdata2](#csr-tdata2). 8 (not equal): Matches when [match](#mcontrol-match)\=0 would not match. 9 (not napot): Matches when [match](#mcontrol-match)\=1 would not match. 12 (not mask low): Matches when [match](#mcontrol-match)\=4 would not match. 13 (not mask high): Matches when [match](#mcontrol-match)\=5 would not match. Other values are reserved for future use. All comparisons only look at the lower XLEN (in the current mode) bits of the compare values and of [tdata2](#csr-tdata2). When [select](#mcontrol-select)\=1 and access size is N, this is further reduced, and comparisons only look at the lower N bits of the compare values and of [tdata2](#csr-tdata2). | **WARL** | 0 | | m | When set, enable this trigger in M-mode. | **WARL** | 0 | | s | When set, enable this trigger in S/HS-mode. This bit is hard-wired to 0 if the hart does not support S-mode. | **WARL** | 0 | | u | When set, enable this trigger in U-mode. This bit is hard-wired to 0 if the hart does not support U-mode. | **WARL** | 0 | | execute | When set, the trigger fires on the virtual address or opcode of an instruction that is executed. | **WARL** | 0 | | store | When set, the trigger fires on the virtual address or data of any store. | **WARL** | 0 | | load | When set, the trigger fires on the virtual address or data of any load. | **WARL** | 0 | #### [](#csr-mcontrol6)Match Control Type 6 (mcontrol6, at 0x7a1) This register provides access to the trigger selected by [tselect](#csr-tselect). The reset values listed here apply to every underlying trigger. This register is accessible as [tdata1](#csr-tdata1) when [type](#tdata1-type) is 6. Implementing this trigger as described here requires that[version](#tinfo-version) is 1 or higher, which in turn means [tinfo](#csr-tinfo) must be implemented. This replaces mcontrol in newer implementations and serves to provide additional functionality. Address and data trigger implementation are heavily dependent on how the processor core is implemented. To accommodate various implementations, execute, load, and store address/data triggers may fire at whatever point in time is most convenient for the implementation. [Table 4](#tab:hwbp%5Ftiming) suggests timings for the best user experience. The underlying principle is that firing just before the instruction gives a user more insight, so is preferable. However, depending on the instruction and conditions, it might not be possible to evaluate the trigger until the instruction has partially executed. In that case it is better to let the instruction retire before the trigger fires, to avoid extra memory accesses which might affect the state of the system. __Table 4\. Suggested Trigger Timings__ | Match Type | Suggested Trigger Timing | | --------------------------- | ------------------------ | | Execute Address | Before | | Execute Instruction | Before | | Execute Address+Instruction | Before | | Load Address | Before | | Load Data | After | | Load Address+Data | After | | Store Address | Before | | Store Data | Before | | Store Address+Data | Before | A chain of triggers must only fire if every trigger in the chain was matched by the same instruction. The Privileged Spec says that breakpoint exceptions that occur on instruction fetches, loads, or stores update the `tval` CSR with either zero or the faulting virtual address. The faulting virtual address for an mcontrol6 trigger with [action](#mcontrol6-action)\=0 is the address being accessed and which caused that trigger to fire. If multiple mcontrol6 triggers are chained then the faulting virtual address is the address which caused any of the chained triggers to fire. In implementations that support [match](#mcontrol6-match) mode 1 (NAPOT), not all NAPOT ranges may be supported. All NAPOT ranges between and are supported where . The value of maskmax6 can be determined by the debugger via the following sequence: 1. Write [tdata1](#csr-tdata1)\=0, in case the current [tdata2](#csr-tdata2) value is not supported with mcontrol6 triggers. 2. Write [tdata2](#csr-tdata2)\=0, which is always supported with mcontrol6 triggers. 3. Write [tdata1](#csr-tdata1) with [type](#tdata1-type)\=mcontrol6 and [match](#mcontrol6-match)\=1. 4. Read [match](#mcontrol6-match). If it is not 1 then NAPOT matching is not supported. 5. Write all ones to [tdata2](#csr-tdata2). 6. Read [tdata2](#csr-tdata2). The value of maskmax6 is the index of the most significant 0 bit plus 1. If [textra32](#csr-textra32) or [textra64](#csr-textra64) are implemented for this trigger, it only matches when the conditions set there are satisfied. | | [uncertain](#mcontrol6-uncertain) and [uncertainen](#mcontrol6-uncertainen) exist to accommodate systems where not every memory access is fully observed by the Trigger Module. Possible examples include data values in far AMOs, and the address/data/size of accesses by instructions that perform multiple memory accesses, such as vector, push, and pop instructions. While the uncertain mechanism exists to deal with these situations, it can lead to an unusable number of false positives. Users will get a much better debug experience if the TM does have perfect visibility into the details of every memory access. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | This CSR is read/write. ![Diagram](_images/diag-2acff6b78beef3a2f9aa1703ab31545906dd0b86.svg) ![Diagram](_images/diag-d4900b81f53b377970866286b5ae2b3588c784e1.svg) | Field | Description | Access | Reset | | ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------- | ----- | | uncertain | If implemented, the TM updates this field every time the trigger fires. 0 (certain): The trigger that fired satisfied the configured conditions, or this bit is not implemented. 1 (uncertain): The trigger that fired might not have perfectly satisfied the configured conditions. Due to the implementation the hardware cannot be certain. | **WARL** | 0 | | vs | When set, enable this trigger in VS-mode. This bit is hard-wired to 0 if the hart does not support virtualization mode. | **WARL** | 0 | | vu | When set, enable this trigger in VU-mode. This bit is hard-wired to 0 if the hart does not support virtualization mode. | **WARL** | 0 | | hit0 | If they are implemented, [hit1](#mcontrol6-hit1) (MSB) and[hit0](#mcontrol6-hit0) (LSB) combine into a single 2-bit field. The TM updates this field when the trigger fires. After the debugger has seen the update, it will normally write 0 to this field to so it can see future changes. If either of the bits is not implemented, the unimplemented bits will be read-only 0. 0 (false): The trigger did not fire. 1 (before): The trigger fired before the instruction that matched it was retired, but after all preceding instructions are retired. This explicitly allows for instructions to be partially executed, as described in [5.1.6\. Multiple State Change Instructions](#multistate). xepc or [dpc](Sdext.html#csr-dpc) (depending on [action](#mcontrol6-action)) must be set to the virtual address of the instruction that matched. 2 (after): The trigger fired after the instruction that triggered and at least one additional instruction were retired.xepc or [dpc](Sdext.html#csr-dpc) (depending on [action](#mcontrol6-action)) must be set to the virtual address of the next instruction that must be executed to preserve the program flow. 3 (immediately after): The trigger fired just after the instruction that triggered it was retired, but before any subsequent instructions were executed.xepc or [dpc](Sdext.html#csr-dpc) (depending on [action](#mcontrol6-action)) must be set to the virtual address of the next instruction that must be executed to preserve the program flow. If the instruction performed multiple memory accesses, all of them have been completed. | **WARL** | 0 | | select | This bit determines the contents of the XLEN-bit compare values. 0 (address): There is at least one compare value and it contains the lowest virtual address of the access. In addition, it is recommended that there are additional compare values for the other accessed virtual addresses match. (E.g. on a 32-bit read from 0x4000, the lowest address is 0x4000 and the other addresses are 0x4001, 0x4002, and 0x4003.) 1 (data): There is exactly one compare value and it contains the data value loaded or stored, or the instruction executed. Any bits beyond the size of the data access will contain 0. | **WARL** | 0 | | size | 0 (any): The trigger will attempt to match against an access of any size. The behavior is only well-defined if [select](#mcontrol6-select)\=0, or if the access size is XLEN. 1 (8bit): The trigger will only match against 8-bit memory accesses. 2 (16bit): The trigger will only match against 16-bit memory accesses or execution of 16-bit instructions. 3 (32bit): The trigger will only match against 32-bit memory accesses or execution of 32-bit instructions. 4 (48bit): The trigger will only match against execution of 48-bit instructions. 5 (64bit): The trigger will only match against 64-bit memory accesses or execution of 64-bit instructions. 6 (128bit): The trigger will only match against 128-bit memory accesses or execution of 128-bit instructions. An implementation must support the value of 0, but all other values are optional. When an implementation supports address triggers ([select](#mcontrol6-select)\=0), it is recommended that those triggers support every access size that the hart supports, as well as for every instruction size that the hart supports. Implementations such as RV32D or RV64V are able to perform loads and stores that are wider than XLEN. Custom extensions may also support instructions that are wider than XLEN. Because[tdata2](#csr-tdata2) is of size XLEN, there is a known limitation that data value triggers ([select](#mcontrol6-select)\=1) can only be supported for access sizes up to XLEN bits. When an implementation supports data value triggers ([select](#mcontrol6-select)\=1), it is recommended that those triggers support every access size up to XLEN that the hart supports, as well as for every instruction length up to XLEN that the hart supports. | **WARL** | 0 | | action | The action to take when the trigger fires. The values are explained in [Table 1](#tab:action). | **WARL** | 0 | | chain | 0 (disabled): When this trigger matches, the configured action is taken. 1 (enabled): While this trigger does not match, it prevents the trigger with the next index from matching. A trigger chain starts on the first trigger with chain\=1 after a trigger with chain\=0, or simply on the first trigger if that has chain\=1\. It ends on the first trigger after that which haschain\=0\. This final trigger is part of the chain. The action on all but the final trigger is ignored. The action on that final trigger will be taken if and only if all the triggers in the chain match at the same time. Debuggers should not terminate a chain with a trigger with a different type. It is undefined when exactly such a chain fires. Because [chain](#mcontrol6-chain) affects the next trigger, hardware must zero it in writes to [mcontrol6](#csr-mcontrol6) that set [dmode](#tdata1-dmode) to 0 if the next trigger has[dmode](#tdata1-dmode) of 1\. In addition hardware should ignore writes to [mcontrol6](#csr-mcontrol6) that set[dmode](#tdata1-dmode) to 1 if the previous trigger has both [dmode](#tdata1-dmode) of 0 and[chain](#mcontrol6-chain) of 1\. Debuggers must avoid the latter case by checking[chain](#mcontrol6-chain) on the previous trigger if they’re writing [mcontrol6](#csr-mcontrol6). Implementations that wish to limit the maximum length of a trigger chain (eg. to meet timing requirements) may do so by zeroing[chain](#mcontrol6-chain) in writes to [mcontrol6](#csr-mcontrol6) that would make the chain too long. | **WARL** | 0 | | match | 0 (equal): Matches when any compare value equals [tdata2](#csr-tdata2). 1 (napot): Matches when the top M bits of any compare value match the topM bits of [tdata2](#csr-tdata2).M is XLEN-1 minus the index of the least-significant bit containing 0 in [tdata2](#csr-tdata2).[tdata2](#csr-tdata2) is **WARL** and if bits maskmax6-1:0 are written with all ones then bit maskmax6-1 will be set to 0 while the values of bits maskmax6-2:0are UNSPECIFIED. Legal values for [tdata2](#csr-tdata2) require M + maskmax6 ≥ XLEN and M \> 0\. See above for how to determine maskmax6. 2 (ge): Matches when any compare value is greater than (unsigned) or equal to [tdata2](#csr-tdata2). 3 (lt): Matches when any compare value is less than (unsigned)[tdata2](#csr-tdata2). 4 (mask low): Matches when of any compare value equals of [tdata2](#csr-tdata2) after of the compare value is ANDed withXLEN-1: of [tdata2](#csr-tdata2). 5 (mask high): Matches when XLEN-1: of any compare value equals of [tdata2](#csr-tdata2) afterXLEN-1: of the compare value is ANDed withXLEN-1: of [tdata2](#csr-tdata2). 8 (not equal): Matches when [match](#mcontrol6-match) \=0 would not match. 9 (not napot): Matches when [match](#mcontrol6-match) \=1 would not match. 12 (not mask low): Matches when [match](#mcontrol6-match) \=4 would not match. 13 (not mask high): Matches when [match](#mcontrol6-match) \=5 would not match. Other values are reserved for future use. All comparisons only look at the lower XLEN (in the current mode) bits of the compare values and of [tdata2](#csr-tdata2). When [select](#mcontrol-select)\=1 and access size is N, this is further reduced, and comparisons only look at the lower N bits of the compare values and of [tdata2](#csr-tdata2). | **WARL** | 0 | | m | When set, enable this trigger in M-mode. | **WARL** | 0 | | uncertainen | 0 (disabled): This trigger will only match if the hardware can perfectly evaluate it. 1 (enabled): This trigger will match if it’s possible that it would match if the Trigger Module had perfect information about the operations being performed. | **WARL** | 0 | | s | When set, enable this trigger in S/HS-mode. This bit is hard-wired to 0 if the hart does not support S-mode. | **WARL** | 0 | | u | When set, enable this trigger in U-mode. This bit is hard-wired to 0 if the hart does not support U-mode. | **WARL** | 0 | | execute | When set, the trigger fires on the virtual address or opcode of an instruction that is executed. | **WARL** | 0 | | store | When set, the trigger fires on the virtual address or data of any store. | **WARL** | 0 | | load | When set, the trigger fires on the virtual address or data of any load. | **WARL** | 0 | #### [](#csr-icount)Instruction Count (icount, at 0x7a1) This register provides access to the trigger selected by [tselect](#csr-tselect). The reset values listed here apply to every underlying trigger. This register is accessible as [tdata1](#csr-tdata1) when [type](#tdata1-type) is 3. This trigger matches when: 1. An instruction retires after having been fetched in a privilege mode where the trigger is enabled. This explicitly includes all RET instructions from various modes. 2. A trap is taken from a privilege mode where the trigger is enabled. This explicitly includes traps taken due to interrupts. If more than one of the above events occur during a single instruction execution, the trigger still only matches once for that instruction. | | For use in single step, icount must match for traps where the instruction will not be reexecuted after the handler, such as illegal instructions that are emulated by privileged software and the instruction being emulated never retires. Ideally, icount would not match for traps where the instruction will later be retried by the handler, such as page faults where privileged software modifies the page tables and returns to the faulting instruction which ultimately retires. Trying to distinguish the two cases leads to complex rules, so instead the rule is simply that all traps match. See also [\[stepicount\]](#stepicount). | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | When [count](#icount-count) is greater than 1 and the trigger matches, then[count](#icount-count) is decremented by 1. When [count](#icount-count) is 1 and the trigger matches, then [pending](#icount-pending)becomes set. In addition [count](#icount-count) will become 0 unless it is hard-wired to 1. The only exception to the above is when the instruction the trigger matched on is a write to the icount trigger. In that case [pending](#icount-pending) might or might not become set if [count](#icount-count) was 1\. Afterwards [count](#icount-count)contains the newly written value. When [count](#icount-count) is 0 it stays at 0 until explicitly written. When [pending](#icount-pending) is set, the trigger fires just before any further instructions are executed in a mode where the trigger is enabled. As the trigger fires, [pending](#icount-pending) is cleared. In addition, if [count](#icount-count) is hard-wired to 1 then [m](#icount-m),[s](#icount-s), [u](#icount-u), [vs](#icount-vs), and [vu](#icount-vu) are all cleared. If the trigger fires with [action](#icount-action)\=0 then zero is written to the`tval` CSR on the breakpoint trap. | | The intent of [pending](#icount-pending) is to cleanly handle the case where[action](#icount-action) is 0, [m](#icount-m) is 0, [u](#icount-u) is 1,[count](#icount-count) is 1, and the U-mode instruction being executed causes a trap into M-mode. In that case we want the entire M-mode handler to be executed, and the debug trap to be taken before the next U-mode instruction. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | This trigger type is intended to be used as a single step for software monitor programs or native debug. Systems that support multiple privilege modes that want to debug software running in lower privilege modes don’t need to support [count](#icount-count) greater than 1. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If [textra32](#csr-textra32) or [textra64](#csr-textra64) are implemented for this trigger, it only matches when the conditions set there are satisfied. This CSR is read/write. ![Diagram](_images/diag-bc85a065e1f7061082517c1075afac8ac83963a0.svg) | Field | Description | Access | Reset | | ------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ----- | | vs | When set, enable this trigger in VS-mode. This bit is hard-wired to 0 if the hart does not support virtualization mode. | **WARL** | 0 | | vu | When set, enable this trigger in VU-mode. This bit is hard-wired to 0 if the hart does not support virtualization mode. | **WARL** | 0 | | hit | If this bit is implemented, the hardware sets it when this trigger fires. The trigger’s user can set or clear it at any time. It is used to determine which trigger(s) fires. If the bit is not implemented, it is always 0 and writing it has no effect. | **WARL** | 0 | | count | The trigger will generally fire after [count](#icount-count) instructions in enabled modes have been executed. See above for the precise behavior. | **WARL** | 1 | | m | When set, enable this trigger in M-mode. | **WARL** | 0 | | pending | This bit becomes set when [count](#icount-count) is decremented from 1 to 0\. It is cleared when the trigger fires, which will happen just before executing the next instruction in one of the enabled modes. | **R/W** | 0 | | s | When set, enable this trigger in S/HS-mode. This bit is hard-wired to 0 if the hart does not support S-mode. | **WARL** | 0 | | u | When set, enable this trigger in U-mode. This bit is hard-wired to 0 if the hart does not support U-mode. | **WARL** | 0 | | action | The action to take when the trigger fires. The values are explained in [Table 1](#tab:action). | **WARL** | 0 | #### [](#csr-itrigger)Interrupt Trigger (itrigger, at 0x7a1) This register provides access to the trigger selected by [tselect](#csr-tselect). The reset values listed here apply to every underlying trigger. This register is accessible as [tdata1](#csr-tdata1) when [type](#tdata1-type) is 4. This trigger can fire when an interrupt trap is taken. It can be enabled for individual interrupt numbers by setting the bit corresponding to the interrupt number in [tdata2](#csr-tdata2). The interrupt number is interpreted in the mode that the trap handler executes in. (E.g. virtualized interrupt numbers are not the same in every mode.) In addition the trigger can be enabled for non-maskable interrupts using[nmi](#itrigger-nmi). | | If XLEN is 32, then it is not possible to set a trigger for interrupts with Exception Code larger than 31\. A future version of the RISC-V Privileged Spec will likely define interrupt Exception Codes 32 through 47\. Some of those numbers are already being used by the RISC-V Advanced Interrupt Architecture. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Hardware may only support a subset of interrupts for this trigger. A debugger must read back [tdata2](#csr-tdata2) after writing it to confirm the requested functionality is actually supported. When the trigger matches, it fires after the trap occurs, just before the first instruction of the trap handler is executed. If[action](#itrigger-action)\=0, the standard CSRs are updated for taking the breakpoint trap, and zero is written to the relevant `tval` CSR. If the breakpoint trap does not go to a higher privilege mode, this will lose CSR information for the original trap. See[5.1.4\. Native Triggers](#nativetrigger) for more information about this case. If [textra32](#csr-textra32) or [textra64](#csr-textra64) are implemented for this trigger, it only matches when the conditions set there are satisfied. This CSR is read/write. ![Diagram](_images/diag-96268af8978ffb7c65595819bcdf02d5f47cdf65.svg) | Field | Description | Access | Reset | | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ----- | | hit | If this bit is implemented, the hardware sets it when this trigger matches. The trigger’s user can set or clear it at any time. It is used to determine which trigger(s) matched. If the bit is not implemented, it is always 0 and writing it has no effect. | **WARL** | 0 | | vs | When set, enable this trigger for interrupts that are taken from VS mode. This bit is hard-wired to 0 if the hart does not support virtualization mode. | **WARL** | 0 | | vu | When set, enable this trigger for interrupts that are taken from VU mode. This bit is hard-wired to 0 if the hart does not support virtualization mode. | **WARL** | 0 | | nmi | When set, non-maskable interrupts cause this trigger to fire if the trigger is enabled for the current mode. | **WARL** | 0 | | m | When set, enable this trigger for interrupts that are taken from M mode. | **WARL** | 0 | | s | When set, enable this trigger for interrupts that are taken from S/HS mode. This bit is hard-wired to 0 if the hart does not support S-mode. | **WARL** | 0 | | u | When set, enable this trigger for interrupts that are taken from U mode. This bit is hard-wired to 0 if the hart does not support U-mode. | **WARL** | 0 | | action | The action to take when the trigger fires. The values are explained in [Table 1](#tab:action). | **WARL** | 0 | #### [](#csr-etrigger)Exception Trigger (etrigger, at 0x7a1) This register provides access to the trigger selected by [tselect](#csr-tselect). The reset values listed here apply to every underlying trigger. This register is accessible as [tdata1](#csr-tdata1) when [type](#tdata1-type) is 5. This trigger may fire on up to XLEN of the Exception Codes defined in`mcause` (described in the Privileged Spec, with Interrupt=0). Those causes are configured by writing the corresponding bit in [tdata2](#csr-tdata2). (E.g. to trap on an illegal instruction, the debugger sets bit 2 in[tdata2](#csr-tdata2).) | | If XLEN is 32, then it is not possible to set a trigger on Exception Codes higher than 31\. A future version of the RISC-V Privileged Spec will likely define Exception Codes 32 through 47. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Hardware may support only a subset of exceptions. A debugger must read back [tdata2](#csr-tdata2) after writing it to confirm the requested functionality is actually supported. When the trigger matches, it fires after the trap occurs, just before the first instruction of the trap handler is executed. If[action](#etrigger-action)\=0, the standard CSRs are updated for taking the breakpoint trap, and zero is written to the relevant `tval` CSR. If the breakpoint trap does not go to a higher privilege mode, this will lose CSR information for the original trap. See[5.1.4\. Native Triggers](#nativetrigger) for more information about this case. If [textra32](#csr-textra32) or [textra64](#csr-textra64) are implemented for this trigger, it only matches when the conditions set there are satisfied. This CSR is read/write. ![Diagram](_images/diag-5aa6130fbdf4b31484a1047ec960d40e8bdeb2f8.svg) | Field | Description | Access | Reset | | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ----- | | hit | If this bit is implemented, the hardware sets it when this trigger matches. The trigger’s user can set or clear it at any time. It is used to determine which trigger(s) matched. If the bit is not implemented, it is always 0 and writing it has no effect. | **WARL** | 0 | | vs | When set, enable this trigger for exceptions that are taken from VS mode. This bit is hard-wired to 0 if the hart does not support virtualization mode. | **WARL** | 0 | | vu | When set, enable this trigger for exceptions that are taken from VU mode. This bit is hard-wired to 0 if the hart does not support virtualization mode. | **WARL** | 0 | | m | When set, enable this trigger for exceptions that are taken from M mode. | **WARL** | 0 | | s | When set, enable this trigger for exceptions that are taken from S/HS mode. This bit is hard-wired to 0 if the hart does not support S-mode. | **WARL** | 0 | | u | When set, enable this trigger for exceptions that are taken from U mode. This bit is hard-wired to 0 if the hart does not support U-mode. | **WARL** | 0 | | action | The action to take when the trigger fires. The values are explained in [Table 1](#tab:action). | **WARL** | 0 | #### [](#csr-tmexttrigger)External Trigger (tmexttrigger, at 0x7a1) This register provides access to the trigger selected by [tselect](#csr-tselect). The reset values listed here apply to every underlying trigger. This register is accessible as [tdata1](#csr-tdata1) when [type](#tdata1-type) is 7. This trigger fires when any selected TM external trigger input signals. Up to 16 TM external trigger inputs coming from other blocks outside the TM, (e.g. signaling an hpmcounter overflow) can be selected. Hardware may support none or just a few TM external trigger inputs (starting with TM external trigger input 0 and continuing sequentially). Unsupported inputs are hardwired to be inactive. If the trigger fires with [action](#tmexttrigger-action)\=0 then zero is written to the`tval` CSR on the breakpoint trap. This trigger fires asynchronously but it is subject to delegation by medeleg\[3\] like the other triggers. The TM external trigger input can signal when the trigger is prevented from firing due to one of the mechanisms in [5.1.4\. Native Triggers](#nativetrigger). An implementation may either ignore the signal altogether when it cannot fire (dropping the trigger event) or it may hold the action as pending and fire the trigger once it is legal to do so. | | [intctl](#tmexttrigger-intctl) is intended to be used by the clicinttrigmechanism from the Core-Local Interrupt Controller (CLIC) RISC-V Privileged Architecture Extensions. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | This CSR is read/write. ![Diagram](_images/diag-f81ba36b09ccc44e2f4ccb5fd9e13b8926706ae4.svg) | Field | Description | Access | Reset | | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ----- | | hit | If this bit is implemented, the hardware sets it when this trigger matches. The trigger’s user can set or clear it at any time. It is used to determine which trigger(s) matched. If the bit is not implemented, it is always 0 and writing it has no effect. | **WARL** | 0 | | intctl | This optional bit, when set, causes this trigger to fire whenever an attached interrupt controller signals a trigger. | **WARL** | 0 | | select | Selects any combination of up to 16 TM external trigger inputs that cause this trigger to fire. | **WARL** | 0 | | action | The action to take when the trigger fires. The values are explained in [Table 1](#tab:action). | **WARL** | 0 | #### [](#csr-textra32)Trigger Extra (RV32) (textra32, at 0x7a3) This register provides access to the trigger selected by [tselect](#csr-tselect). The reset values listed here apply to every underlying trigger. This register is accessible as [tdata3](#csr-tdata3) when [type](#tdata1-type) is 2, 3, 4, 5, or 6 and XLEN=32. If DXLEN >= 64, then this register provides access to the low bits of each field defined in [textra64](#csr-textra64). Writes to this register will clear the high bits of the corresponding fields in [textra64](#csr-textra64). All functionality in this register is optional. Any number of upper bits of [mhvalue](#textra32-mhvalue) and [svalue](#textra32-svalue) may be tied to 0.[mhselect](#textra32-mhselect) and [sselect](#textra32-sselect) may only support 0 (ignore). Byte-granular comparison of [scontext](#csr-scontext) to [svalue](#textra32-svalue)allows [scontext](#csr-scontext) to be defined to include more than one element of comparison. For example, software instrumentation can program the [scontext](#csr-scontext) value to be the concatenation of different ID contexts such as process ID and thread ID. The user can then program byte compares based on [sbytemask](#textra32-sbytemask)to include one or more of the contexts in the compare. Byte masking only applies to [scontext](#csr-scontext) comparison; i.e when [sselect](#textra32-sselect) is 1. | | Note that sselect and mhselect filtering apply in all modes, including M-mode and S-mode. If desired, debuggers can use a trigger’s mode filtering bits to restrict the matching to modes where it considers ASID/VMID/scontext/hcontext to be active. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | This CSR is read/write. ![Diagram](_images/diag-b6381fa173fad76fc4a9b007285002edb778fd1f.svg) | Field | Description | Access | Reset | | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ----- | | mhvalue | Data used together with [mhselect](#textra32-mhselect). | **WARL** | 0 | | mhselect | 0 (ignore): Ignore [mhvalue](#textra32-mhvalue). 4 (mcontext): This trigger will only match or fire if the low bits of[mcontext](#csr-mcontext)/[hcontext](#csr-hcontext) equal [mhvalue](#textra32-mhvalue). 1, 5 (mcontext\_select): This trigger will only match or fire if the low bits of[mcontext](#csr-mcontext)/[hcontext](#csr-hcontext) equal {[mhvalue](#textra32-mhvalue), mhselect\[2\]}. 2, 6 (vmid\_select): This trigger will only match or fire if VMID in hgatp equals the lower VMIDMAX (defined in the Privileged Spec) bits of {[mhvalue](#textra32-mhvalue), mhselect\[2\]}. 3, 7 (reserved): Reserved. If the H extension is not supported, the only legal values are 0 and 4. | **WARL** | 0 | | sbytemask | When the least significant bit of this field is 1, it causes bits 7:0 in the comparison to be ignored, when [sselect](#textra32-sselect)\=1\. When the next most significant bit of this field is 1, it causes bits 15:8 to be ignored in the comparison, when [sselect](#textra32-sselect)\=1. | **WARL** | 0 | | svalue | Data used together with [sselect](#textra32-sselect). This field should be tied to 0 when S-mode is not supported. | **WARL** | 0 | | sselect | 0 (ignore): Ignore [svalue](#textra32-svalue). 1 (scontext): This trigger will only match or fire if the low bits of[scontext](#csr-scontext) equal [svalue](#textra32-svalue). 2 (asid): This trigger will only match or fire if: the mode is VS-mode or VU-mode and ASID in vsatpequals the lower ASIDMAX (defined in the Privileged Spec) bits of [svalue](#textra32-svalue). in all other modes, ASID in satp equals the lower ASIDMAX (defined in the Privileged Spec) bits of[svalue](#textra32-svalue). This field should be tied to 0 when S-mode is not supported. | **WARL** | 0 | #### [](#csr-textra64)Trigger Extra (RV64) (textra64, at 0x7a3) This register provides access to the trigger selected by [tselect](#csr-tselect). The reset values listed here apply to every underlying trigger. This register is accessible as [tdata3](#csr-tdata3) when [type](#tdata1-type) is 2, 3, 4, 5, or 6 and XLEN=64\. The function of the fields are defined above, in [textra32](#csr-textra32). This register retains its value when XLEN changes. When XLEN=32 some of the bits can be accessed through [textra32](#csr-textra32). Byte-granular comparison of [scontext](#csr-scontext) to [svalue](#textra64-svalue) in[textra64](#csr-textra64) allows [scontext](#csr-scontext) to be defined to include more than one element of comparison. For example, software instrumentation can program the [scontext](#csr-scontext) value to be the concatenation of different ID contexts such as process ID and thread ID. The user can then program byte compares based on[sbytemask](#textra64-sbytemask) to include one or more of the contexts in the compare. Byte masking only applies to [scontext](#csr-scontext) comparison; i.e when [sselect](#textra64-sselect) is 1. This CSR is read/write. ![Diagram](_images/diag-c0b0fd6ad2fabb2bdb06b907570087671dd9f67b.svg) | Field | Description | Access | Reset | | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ----- | | sbytemask | When the least significant bit of this field is 1, it causes bits 7:0 in the comparison to be ignored, when [sselect](#textra64-sselect)\=1\. Likewise, the second bit controls the comparison of bits 15:8, third bit controls the comparison of bits 23:16, and fourth bit controls the comparison of bits 31:24. | **WARL** | 0 | 3.1. Debug Module (DM) (non-ISA extension) ==================== ## [](#dm)3.1\. Debug Module (DM) (non-ISA extension) The Debug Module implements a translation interface between abstract debug operations and their specific implementation. It might support the following operations: 1. Give the debugger necessary information about the implementation. (Required) 2. Allow any individual hart to be halted and resumed. (Required) 3. Provide status on which harts are halted. (Required) 4. Provide abstract read and write access to a halted hart’s GPRs. (Required) 5. Provide access to a reset signal that allows debugging from the very first instruction after reset. (Required) 6. Provide a mechanism to allow debugging harts immediately out of reset (regardless of the reset cause). (Optional) 7. Provide abstract access to non-GPR hart registers. (Optional) 8. Provide a Program Buffer to force the hart to execute arbitrary instructions. (Optional) 9. Allow multiple harts to be halted, resumed, and/or reset at the same time. (Optional) 10. Allow memory access from a hart’s point of view. (Optional) 11. Allow direct System Bus Access. (Optional) 12. Group harts. When any hart in the group halts, they all halt. (Optional) 13. Respond to external triggers by halting each hart in a configured group. (Optional) 14. Signal an external trigger when a hart in a group halts. (Optional) In order to be compatible with this specification an implementation must: 1. Implement all the required features listed above. 2. Implement at least one of Program Buffer, System Bus Access, or Abstract Access Memory command mechanisms. 3. Do at least one of: 1. Implement the Program Buffer. 2. Implement abstract access to all registers that are visible to software running on the hart including all the registers that are present on the hart and listed in [Table 3](#tab:regno). 3. Implement abstract access to at least all GPRs, [dcsr](Sdext.html#csr-dcsr), and [dpc](Sdext.html#csr-dpc), and advertise the implementation as conforming to the "Minimal RISC-V Debug Specification", instead of the "RISC-V Debug Specification". A single DM can debug up to harts. ### [](#dmi)3.1.1\. Debug Module Interface (DMI) Debug Modules are subordinates on a bus called the Debug Module Interface (DMI). The bus manager is the Debug Transport Module(s). The Debug Module Interface can be a trivial bus with one manager and one subordinate (see [implementations.adoc#tab:dmi\_signals](implementations.html#tab:dmi%5Fsignals)), or use a more full-featured bus like TileLink or the AMBA Advanced Peripheral Bus. The details are left to the system designer. The DMI uses between 7 and 32 address bits. Each address points at a single 32-bit register that can be read or written. The bottom of the address space is used for the first (and usually only) DM. Extra space can be used for custom debug devices, other cores, additional DMs, etc. If there are additional DMs on this DMI, the base address of the next DM in the DMI address space is given in [nextdm](#dm-nextdm). The Debug Module is controlled via register accesses to its DMI address space. ### [](#reset)3.1.2\. Reset Control There are two methods that allow a debugger to reset harts. [ndmreset](#dmcontrol-ndmreset) resets all the harts in the hardware platform, as well as all other parts of the hardware platform except for the Debug Modules, Debug Transport Modules, and Debug Module Interface. Exactly what is affected by this reset is implementation dependent, but it must be possible to debug programs from the first instruction executed. [hartreset](#dmcontrol-hartreset) resets all the currently selected harts. In this case an implementation may reset more harts than just the ones that are selected. The debugger can discover which other harts are reset (if any) by selecting them and checking [anyhavereset](#dmstatus-anyhavereset) and [allhavereset](#dmstatus-allhavereset). To perform either of these resets, the debugger first asserts the bit, and then clears it. The actual reset may start as soon as the bit is asserted, but may start an arbitrarily long time after the bit is deasserted. The reset itself may also take an arbitrarily long time. While the reset is on-going, harts are either in the running state, indicating it’s possible to perform some abstract commands during this time, or in the unavailable state, indicating it’s not possible to perform any abstract commands during this time. Once a hart’s reset is complete, `havereset` becomes set. When a hart comes out of reset and [haltreq](#dmcontrol-haltreq) or `resethaltreq`are set, the hart will immediately enter Debug Mode (halted state). Otherwise, if the hart was initially running it will execute normally (running state) and if the hart was initially halted it should now be running but may be halted. | | There is no general, reliable way for the debugger to know when reset has actually begun. | | -------------------------------------------------------------------------------------------- | The Debug Module’s own state and registers should only be reset at power-up and while [dmactive](#dmcontrol-dmactive) in [dmcontrol](#dm-dmcontrol) is 0\. If there is another mechanism to reset the DM, this mechanism must also reset all the harts accessible to the DM. Due to clock and power domain crossing issues, it might not be possible to perform arbitrary DMI accesses across hardware platform reset. While[ndmreset](#dmcontrol-ndmreset) or any external reset is asserted, the only supported DM operations are reading/writing [dmcontrol](#dm-dmcontrol) and reading[ndmresetpending](#dmstatus-ndmresetpending). The behavior of other accesses is undefined. When harts have been reset, they must set a sticky `havereset` state bit. The conceptual `havereset` state bits can be read for selected harts in [anyhavereset](#dmstatus-anyhavereset) and [allhavereset](#dmstatus-allhavereset) in [dmstatus](#dm-dmstatus). These bits must be set regardless of the cause of the reset. The `havereset` bits for the selected harts can be cleared by writing 1 to [ackhavereset](#dmcontrol-ackhavereset) in [dmcontrol](#dm-dmcontrol). The `havereset` bits might or might not be cleared when [dmactive](#dmcontrol-dmactive) is low. ### [](#selectingharts)3.1.3\. Selecting Harts Up to harts can be connected to a single DM. Commands issued to the DM only apply to the currently selected harts. To enumerate all the harts, a debugger must first determine `HARTSELLEN`by writing all ones to \`hartsel\` (assuming the maximum size) and reading back the value to see which bits were actually set. Then it selects each hart starting from 0 until either [anynonexistent](#dmstatus-anynonexistent) in [dmstatus](#dm-dmstatus) is 1, or the highest index (depending on `HARTSELLEN`) is reached. The debugger can discover the mapping between hart indices and `mhartid` by using the interface to read `mhartid`, or by reading the hardware platform’s configuration structure. #### [](#3-1-3-1-selecting-a-single-hart)3.1.3.1\. Selecting a Single Hart All debug modules must support selecting a single hart. The debugger can select a hart by writing its index to \`hartsel\`. Hart indexes start at 0 and are contiguous until the final index. #### [](#hartarraymask)3.1.3.2\. Selecting Multiple Harts Debug Modules may implement a Hart Array Mask register to allow selecting multiple harts at once. The th bit in the Hart Array Mask register applies to the hart with index . If the bit is 1 then the hart is selected. Usually a DM will have a Hart Array Mask register exactly wide enough to select all the harts it supports, but it’s allowed to tie any of these bits to 0. The debugger can set bits in the hart array mask register using [hawindowsel](#dm-hawindowsel) and [hawindow](#dm-hawindow), then apply actions to all selected harts by setting [hasel](#dmcontrol-hasel). If this feature is supported, multiple harts can be halted, resumed, and reset simultaneously. The state of the hart array mask register is not affected by setting or clearing [hasel](#dmcontrol-hasel). Execution of Abstract Commands ignores this mechanism and only applies to the hart selected by \`hartsel\`. ### [](#3-1-4-hart-dm-states)3.1.4\. Hart DM States Every hart that can be selected is in exactly one of the following four DM states: non-existent, unavailable, running, or halted. Which state the selected harts are in is reflected by [allnonexistent](#dmstatus-allnonexistent), [anynonexistent](#dmstatus-anynonexistent), [allunavail](#dmstatus-allunavail), [anyunavail](#dmstatus-anyunavail), [allrunning](#dmstatus-allrunning), [anyrunning](#dmstatus-anyrunning), [allhalted](#dmstatus-allhalted), and [anyhalted](#dmstatus-anyhalted). Harts are nonexistent if they will never be part of this hardware platform, no matter how long a user waits. E.g. in a simple single-hart hardware platform only one hart exists, and all others are nonexistent. Debuggers may assume that a hardware platform has no harts with indexes higher than the first nonexistent one. Harts are unavailable if they might exist/become available at a later time, or if there are other harts with higher indexes than this one. Harts may be unavailable for a variety of reasons including being reset, temporarily powered down, and not being plugged into the hardware platform. That means harts might become available or unavailable at any time, although these events should be rare in hardware platforms built to be easily debugged. There are no guarantees about the state of the hart when it becomes available. Hardware platforms with very large number of harts may permanently disable some during manufacturing, leaving holes in the otherwise continuous hart index space. In order to let the debugger discover all harts, they must show up as unavailable even if there is no chance of them ever becoming available. Harts are running when they are executing normally, as if no debugger was attached. This includes being in a low power mode or waiting for an interrupt, as long as a halt request will result in the hart being halted. Harts are halted when they are in Debug Mode, only performing tasks on behalf of the debugger. Which states a hart that is reset goes through is implementation dependent. Harts may be unavailable while reset is asserted, and some time after reset is deasserted. They might transition to running for some time after reset is deasserted. Finally they end up either running or halted, depending on [haltreq](#dmcontrol-haltreq) and `resethaltreq`. ### [](#runcontrol)3.1.5\. Run Control For every hart, the Debug Module tracks 4 conceptual bits of state: halt request, resume ack, halt-on-reset request, and hart reset. (The hart reset and halt-on-reset request bits are optional.) These 4 bits reset to 0, except for resume ack, which may reset to either 0 or 1\. The DM receives halted, running, and havereset signals from each hart. The debugger can observe the state of resume ack in [allresumeack](#dmstatus-allresumeack) and [anyresumeack](#dmstatus-anyresumeack), and the state of halted, running, and havereset signals in [allhalted](#dmstatus-allhalted), [anyhalted](#dmstatus-anyhalted), [allrunning](#dmstatus-allrunning), [anyrunning](#dmstatus-anyrunning), [allhavereset](#dmstatus-allhavereset), and [anyhavereset](#dmstatus-anyhavereset). The state of the other bits cannot be observed directly. When a debugger writes 1 to [haltreq](#dmcontrol-haltreq), each selected hart’s halt request bit is set. When a running hart, or a hart just coming out of reset, sees its halt request bit high, it responds by halting, deasserting its running signal, and asserting its halted signal. Halted harts ignore their halt request bit. When a debugger writes 1 to [resumereq](#dmcontrol-resumereq), each selected hart’s resume ack bit is cleared and each selected, halted hart is sent a resume request. Harts respond by resuming, clearing their halted signal, and asserting their running signal. At the end of this process the resume ack bit is set. These status signals of all selected harts are reflected in [allresumeack](#dmstatus-allresumeack), [anyresumeack](#dmstatus-anyresumeack), [allrunning](#dmstatus-allrunning), and [anyrunning](#dmstatus-anyrunning). Resume requests are ignored by running harts. When halt or resume is requested, a hart must respond in less than one second, unless it is unavailable. (How this is implemented is not further specified. A few clock cycles will be a more typical latency). The DM can implement optional halt-on-reset bits for each hart, which it indicates by setting [hasresethaltreq](#dmstatus-hasresethaltreq) to 1\. This means the DM implements the [setresethaltreq](#dmcontrol-setresethaltreq) and [clrresethaltreq](#dmcontrol-clrresethaltreq) bits. Writing 1 to [setresethaltreq](#dmcontrol-setresethaltreq) sets the halt-on-reset request bit for each selected hart. When a hart’s halt-on-reset request bit is set, the hart will immediately enter debug mode on the next deassertion of its reset. This is true regardless of the reset’s cause. The hart’s halt-on-reset request bit remains set until cleared by the debugger writing 1 to [clrresethaltreq](#dmcontrol-clrresethaltreq) while the hart is selected, or by DM reset. If the DM is reset while a hart is halted, it is UNSPECIFIED whether that hart resumes. Debuggers should use [resumereq](#dmcontrol-resumereq) to explicitly resume harts before clearing [dmactive](#dmcontrol-dmactive) and disconnecting. ### [](#hrgroups)3.1.6\. Halt Groups, Resume Groups, and External Triggers An optional feature allows a debugger to place harts into two kinds of groups: halt groups and resume groups. It is also possible to add external triggers to a halt and resume groups. At any given time, each hart and each trigger is a member of exactly one halt group and exactly one resume group. In both halt and resume groups, group 0 is special. Harts in group 0 halt/resume as if groups aren’t implemented at all. When any hart in a halt group halts: 1. That hart halts normally, with [cause](Sdext.html#dcsr-cause) reflecting the original cause of the halt. 2. All the other harts in the halt group that are running will quickly halt. [cause](Sdext.html#dcsr-cause) for those harts should be set to 6, but may be set to 3\. Other harts in the halt group that are halted but have started the process of resuming must also quickly become halted, even if they do resume briefly. 3. Any external triggers in that group are notified. Adding a hart to a halt group does not automatically halt that hart, even if other harts in the group are already halted. When an external trigger that’s a member of the halt group fires: 1. All the harts in the halt group that are running will quickly halt. [cause](Sdext.html#dcsr-cause) for those harts should be set to 6, but may be set to 3\. Other harts in the halt group that are halted but have started the process of resuming must also quickly become halted, even if they do resume briefly. When any hart in a resume group resumes: 1. All the other harts in that group that are halted will quickly resume as soon as any currently executing abstract commands have completed. Each hart in the group sets its resume ack bit as soon as it has resumed. Harts that are in the process of halting should complete that process and stay halted. 2. Any external triggers in that group are notified. Adding a hart to a resume group does not automatically resume that hart, even if other harts in the group are currently running. When an external trigger that’s a member of the resume group fires: 1. All the harts in that group that are halted will quickly resume as soon as any currently executing abstract commands have completed. Each hart in the group sets its resume ack bit as soon as it has resumed. Harts that are in the process of halting should complete that process and stay halted. External triggers are abstract concepts that can signal the DM and/or receive signals from the DM. This configuration is done through [dmcs2](#dm-dmcs2), where external triggers are referred to by a number. Commonly, external triggers are capable of sending a signal from the hardware platform into the DM, as well as receiving a signal from the DM to take their own action on. It is also allowable for an external trigger to be input-only or output-only. By convention external triggers 0-7 are bidirectional, triggers 8-11 are input-only, and triggers 12-15 are output-only but this is not required. | | External triggers could be used to implement near simultaneous halting/resuming of all cores in a hardware platform, when not all cores are RISC-V cores. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | When the DM is reset, all harts must be placed in the lowest-numbered halt and resume groups that they can be in. (This will usually be group 0.) Some designs may choose to hardcode hart groups to a group other than group 0, meaning it is never possible to halt or resume just a single hart. This is explicitly allowed. In that case it must be possible to discover the groups by using [dmcs2](#dm-dmcs2) even if it’s not possible to change the configuration. ### [](#abstractcommands)3.1.7\. Abstract Commands The DM supports a set of abstract commands, most of which are optional. Depending on the implementation, the debugger may be able to perform some abstract commands even when the selected hart is not halted. Debuggers can only determine which abstract commands are supported by a given hart in a given state (running, halted, or held in reset) by attempting them and then looking at [cmderr](#abstractcs-cmderr) in [abstractcs](#dm-abstractcs) to see if they were successful. Commands may be supported with some options set, but not with other options set. If a command has unsupported options set or if bits that are defined as 0 aren’t 0, then the DM must set [cmderr](#abstractcs-cmderr) to 2 (not supported). | | Example: Every DM must support the Access Register command, but might not support accessing CSRs. If the debugger requests to read a CSR in that case, the command will return "not supported". | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Debuggers execute abstract commands by writing them to [command](#dm-command). They can determine whether an abstract command is complete by reading [busy](#abstractcs-busy) in [abstractcs](#dm-abstractcs). If the debugger starts a new command while [busy](#abstractcs-busy) is set, [cmderr](#abstractcs-cmderr) becomes 1 (busy), the currently executing command still gets to run to completion, but any error generated by the currently executing command is lost. After completion, [cmderr](#abstractcs-cmderr) indicates whether the command was successful or not. Commands may fail because a hart is not halted, not running, unavailable, or because they encounter an error during execution. If the command takes arguments, the debugger must write them to the`data` registers before writing to [command](#dm-command). If a command returns results, the Debug Module must ensure they are placed in the `data` registers before [busy](#abstractcs-busy)is cleared. Which `data` registers are used for the arguments is described in [Table 1](#tab:datareg). In all cases the least-significant word is placed in the lowest-numbered `data` register. The argument width depends on the command being executed, and is DXLEN where not explicitly specified. __Table 1\. Use of Data Registers__ | Argument Width | arg0/return value | arg1 | arg2 | | -------------- | ------------------------- | ------------ | ------------- | | 32 | [data0](#dm-data0) | data1 | data2 | | 64 | [data0](#dm-data0), data1 | data2, data3 | data4, data5 | | 128 | [data0](#dm-data0)\-data3 | data4\-data7 | data8\-data11 | | | The Abstract Command interface is designed to allow a debugger to write commands as fast as possible, and then later check whether they completed without error. In the common case the debugger will be much slower than the target and commands succeed, which allows for maximum throughput. If there is a failure, the interface ensures that no commands execute after the failing one. To discover which command failed, the debugger has to look at the state of the DM (e.g. contents of [data0](#dm-data0)) or hart (e.g. contents of a register modified by a Program Buffer program) to determine which one failed. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | While an abstract command is executing ([busy](#abstractcs-busy) in [abstractcs](#dm-abstractcs) is high), a debugger must not change \`hartsel\`, and must not write 1 to [haltreq](#dmcontrol-haltreq), [resumereq](#dmcontrol-resumereq), [ackhavereset](#dmcontrol-ackhavereset), [setresethaltreq](#dmcontrol-setresethaltreq), or [clrresethaltreq](#dmcontrol-clrresethaltreq). The hardware should not rely on this debugger behavior, but should enforce it by ignoring writes to these bits while [busy](#abstractcs-busy) is high. If an abstract command does not complete in the expected time and appears to be hung, the debugger can try to reset the hart (using [hartreset](#dmcontrol-hartreset) or [ndmreset](#dmcontrol-ndmreset)). If that doesn’t clear [busy](#abstractcs-busy), then it can try resetting the Debug Module (using [dmactive](#dmcontrol-dmactive)). If an abstract command is started while the selected hart is unavailable or if a hart becomes unavailable while executing an abstract command, then the Debug Module may terminate the abstract command, setting [busy](#abstractcs-busy) low, and [cmderr](#abstractcs-cmderr) to 4 (halt/resume). Alternatively, the command could just appear to be hung ([busy](#abstractcs-busy) never goes low). #### [](#3-1-7-1-abstract-command-listing)3.1.7.1\. Abstract Command Listing This section describes each of the different abstract commands and how their fields should be interpreted when they are written to [command](#dm-command). Each abstract command is a 32-bit value. The top 8 bits contain [cmdtype](#command-cmdtype) which determines the kind of command. [Table 2](#tab:cmdtype) lists all commands. __Table 2\. Meaning of [cmdtype](#command-cmdtype)__ | [cmdtype](#command-cmdtype) | Command | | --------------------------- | --------------------------------------------- | | 0 | [Access Register Command](#ac-accessregister) | | 1 | [Quick Access](#ac-quickaccess) | | 2 | [Access Memory Command](#ac-accessmemory) | ##### [](#ac-accessregister)`Access Register` This command gives the debugger access to CPU registers and allows it to execute the Program Buffer. It performs the following sequence of operations: 1. If [write](#accessregister-write) is clear and [transfer](#accessregister-transfer) is set, then copy data from the register specified by [regno](#accessregister-regno) into the `arg0` region of`data`, and perform any side effects that occur when this register is read from M-mode. 2. If [write](#accessregister-write) is set and [transfer](#accessregister-transfer) is set, then copy data from the`arg0` region of `data` into the register specified by[regno](#accessregister-regno), and perform any side effects that occur when this register is written from M-mode. 3. If [aarpostincrement](#accessregister-aarpostincrement) and[transfer](#accessregister-transfer) are set, increment[regno](#accessregister-regno). [regno](#accessregister-regno) may also be incremented if [aarpostincrement](#accessregister-aarpostincrement) is set and[transfer](#accessregister-transfer) is clear. 4. Execute the Program Buffer, if [postexec](#accessregister-postexec) is set. If any of these operations fail, [cmderr](#abstractcs-cmderr) is set and none of the remaining steps are executed. An implementation may detect an upcoming failure early, and fail the overall command before it reaches the step that would cause failure. If the failure is that the requested register does not exist in the hart, [cmderr](#abstractcs-cmderr) must be set to 3 (exception). Debug Modules must implement this command and must support read and write access to all GPRs when the selected hart is halted. Debug Modules may optionally support accessing other registers, or accessing registers when the hart is running. It is recommended that if one register in a group is accessible, then all registers in that group are accessible, but each individual register (aside from GPRs) may be supported differently across read, write, and halt status. Registers might not be accessible if they wouldn’t be accessible by M mode code currently running. (E.g. `fflags` might not be accessible when `mstatus.FS` is 0.) If this is the case, the debugger is responsible for changing state to make the registers accessible. The Core Debug Registers ([\[debreg\]](#debreg)) should be accessible if abstract CSR access is implemented. __Table 3\. Abstract Register Numbers__ | Numbers | Group Description | | --------------- | -------------------------------------------------------------------------- | | 0x0000 — 0x0fff | CSRs. The \`\`PC'' can be accessed here through [dpc](Sdext.html#csr-dpc). | | 0x1000 — 0x101f | GPRs | | 0x1020 — 0x103f | Floating point registers | | 0xc000 — 0xffff | Reserved for non-standard extensions and internal use. | | | The encoding of [aarsize](#accessregister-aarsize) was chosen to match [sbaccess](#sbcs-sbaccess) in [sbcs](#dm-sbcs). | | ------------------------------------------------------------------------------------------------------------------------- | This command modifies `arg0` only when a register is read. The other `data` registers are not changed. ![Diagram](_images/diag-c27e6f81b8dbe2930d33fbaeed4b2dfc0537022e.svg) | Field | Description | | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | cmdtype | This is 0 to indicate Access Register Command. | | aarsize | 2 (32bit): Access the lowest 32 bits of the register. 3 (64bit): Access the lowest 64 bits of the register. 4 (128bit): Access the lowest 128 bits of the register. If [aarsize](#accessregister-aarsize) specifies a size larger than the register’s actual size, then the access must fail. If a register is accessible, then reads of [aarsize](#accessregister-aarsize)less than or equal to the register’s actual size must be supported. Writing less than the full register may be supported, but what happens to the high bits in that case is UNSPECIFIED. This field controls the Argument Width as referenced in[Table 1](#tab:datareg). | | aarpostincrement | 0 (disabled): No effect. This variant must be supported. 1 (enabled): After a successful register access, [regno](#accessregister-regno) is incremented. Incrementing past the highest supported value causes [regno](#accessregister-regno) to become UNSPECIFIED. Supporting this variant is optional. It is undefined whether the increment happens when [transfer](#accessregister-transfer) is 0. | | postexec | 0 (disabled): No effect. This variant must be supported, and is the only supported one if [progbufsize](#abstractcs-progbufsize) is 0. 1 (enabled): Execute the program in the Program Buffer exactly once after performing the transfer, if any. Supporting this variant is optional. | | transfer | 0 (disabled): Don’t do the operation specified by [write](#accessregister-write). 1 (enabled): Do the operation specified by [write](#accessregister-write). This bit can be used to just execute the Program Buffer without having to worry about placing valid values into [aarsize](#accessregister-aarsize) or [regno](#accessregister-regno). | | write | When [transfer](#accessregister-transfer) is set: 0 (arg0): Copy data from the specified register into arg0 portion of data. 1 (register): Copy data from arg0 portion of data into the specified register. | | regno | Number of the register to access, as described in[Table 3](#tab:regno).[dpc](Sdext.html#csr-dpc) may be used as an alias for PC if this command is supported on a non-halted hart. | ##### [](#ac-quickaccess)`Quick Access` Perform the following sequence of operations: 1. If the hart is halted, the command sets [cmderr](#abstractcs-cmderr) to \`\`halt/resume'' and does not continue. 2. Halt the hart. If the hart halts for some other reason (e.g. breakpoint), the command sets [cmderr](#abstractcs-cmderr) to \`\`halt/resume'' and does not continue. 3. Execute the Program Buffer. If an exception occurs, [cmderr](#abstractcs-cmderr) is set to \`\`exception,'' the Program Buffer execution ends, and the hart is halted with [cause](Sdext.html#dcsr-cause) set to 3. 4. If the Program Buffer executed without an exception, then resume the hart. Implementing this command is optional. This command does not touch the `data` registers. ![Diagram](_images/diag-f825a9b60ffbf2af4685acb541ebb563bff2d59a.svg) | Field | Description | | ------- | ------------------------------------------- | | cmdtype | This is 1 to indicate Quick Access command. | ##### [](#ac-accessmemory)`Access Memory` This command lets the debugger perform memory accesses, with the exact same memory view and permissions as performing loads/stores on the selected hart. This includes access to hart-local memory-mapped registers, etc. The command performs the following sequence of operations: 1. Copy data from the memory location specified in `arg1` into the`arg0` portion of `data`, if [write](#accessregister-write) is clear. 2. Copy data from the `arg0` portion of `data` into the memory location specified in `arg1`, if [write](#accessregister-write) is set. 3. If [aampostincrement](#accessmemory-aampostincrement) is set, increment `arg1`. If any of these operations fail, [cmderr](#abstractcs-cmderr) is set and none of the remaining steps are executed. An access may only fail if the hart, running M-mode code, might encounter that same failure when it attempts the same access. An implementation may detect an upcoming failure early, and fail the overall command before it reaches the step that would cause failure. Debug Modules may optionally implement this command and may support read and write access to memory locations when the selected hart is running or halted. If this command supports memory accesses while the hart is running, it must also support memory accesses while the hart is halted. | | The encoding of [aamsize](#accessmemory-aamsize) was chosen to match [sbaccess](#sbcs-sbaccess) in [sbcs](#dm-sbcs). | | ----------------------------------------------------------------------------------------------------------------------- | This command modifies `arg0` only when memory is read. It modifies`arg1` only if [aampostincrement](#accessmemory-aampostincrement) is set. The other `data`registers are not changed. ![Diagram](_images/diag-d21e8d22644c19ef37a98a7e5733767af3ab5d9e.svg) | Field | Description | | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | cmdtype | This is 2 to indicate Access Memory Command. | | aamvirtual | An implementation does not have to implement both virtual and physical accesses, but it must fail accesses that it doesn’t support. 0 (physical): Addresses are physical (to the hart they are performed on). 1 (virtual): Addresses are virtual, and translated the way they would be from M-mode, with MPRV set. Debug Modules on systems without address translation (i.e. virtual addresses equal physical) may optionally allow [aamvirtual](#accessmemory-aamvirtual) set to 1, which would produce the same result as that same abstract command with [aamvirtual](#accessmemory-aamvirtual) cleared. | | aamsize | 0 (8bit): Access the lowest 8 bits of the memory location. 1 (16bit): Access the lowest 16 bits of the memory location. 2 (32bit): Access the lowest 32 bits of the memory location. 3 (64bit): Access the lowest 64 bits of the memory location. 4 (128bit): Access the lowest 128 bits of the memory location. | | aampostincrement | After a memory access has completed, if this bit is 1, incrementarg1 (which contains the address used) by the number of bytes encoded in [aamsize](#accessmemory-aamsize). Supporting this variant is optional, but highly recommended for performance reasons. | | write | 0 (arg0): Copy data from the memory location specified in arg1 into the low bits of arg0. The value of the remaining bits ofarg0 are UNSPECIFIED. 1 (memory): Copy data from the low bits of arg0 into the memory location specified in arg1. | | target-specific | These bits are reserved for target-specific uses. | ### [](#programbuffer)3.1.8\. Program Buffer To support executing arbitrary instructions on a halted hart, a Debug Module can include a Program Buffer that a debugger can write small programs to. DMs that support all necessary functionality using abstract commands only may choose to omit the Program Buffer. A debugger can write a small program to the Program Buffer, and then execute it exactly once with the Access Register Abstract Command, setting the [postexec](#accessregister-postexec) bit in [command](#dm-command). The debugger can write whatever program it likes (including jumps out of the Program Buffer), but the program must end with `ebreak` or `c.ebreak`. An implementation may support an implicit`ebreak` that is executed when a hart runs off the end of the Program Buffer. This is indicated by [impebreak](#dmstatus-impebreak). With this feature, a Program Buffer of just 2 32-bit words can offer efficient debugging. While these programs are executed, the hart does not leave Debug Mode (see [Sdext.adoc#debugmode](Sdext.html#debugmode)). If an exception is encountered during execution of the Program Buffer, no more instructions are executed, the hart remains in Debug Mode, and [cmderr](#abstractcs-cmderr) is set to 3 (`exception error`). If the debugger executes a program that doesn’t terminate with an `ebreak` instruction, the hart will remain in Debug Mode and the debugger will lose control of the hart. If [progbufsize](#abstractcs-progbufsize) is 1 then the following apply: 1. [impebreak](#dmstatus-impebreak) must be 1. 2. If the debugger writes a compressed instruction into the Program Buffer, it must be placed into the lower 16 bits and accompanied by a compressed`nop` in the upper 16 bits. | | This requirement on the debugger for the case of [progbufsize](#abstractcs-progbufsize) equal to 1 is to accommodate hardware designs that prefer to stuff instructions directly into the pipeline when halted, instead of having the Program Buffer exist in the address space somewhere. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The Program Buffer may be implemented as RAM which is accessible to the hart. A debugger can determine if this is the case by executing small programs that attempt to write and read back relative to `pc` while executing from the Program Buffer. If so, the debugger has more flexibility in what it can do with the program buffer. ### [](#3-1-9-overview-of-hart-debug-states)3.1.9\. Overview of Hart Debug States [Figure 1](#abstract%5Fsm) shows a conceptual view of the states passed through by a hart during run/halt debugging as influenced by the different fields of [dmcontrol](#dm-dmcontrol), [abstractcs](#dm-abstractcs), [abstractauto](#dm-abstractauto), and [command](#dm-command). ![abstract commands](_images/abstract_commands.png) Figure 1\. Run/Halt Debug State Machine for single-hart hardware platforms. As only a small amount of state is visible to the debugger, the states and transitions are conceptual. ### [](#systembusaccess)3.1.10\. System Bus Access A debugger can access memory from a hart’s point of view using a Program Buffer or the Abstract Access Memory command. (Both these features are optional.) A Debug Module may also include a System Bus Access block to provide memory access without involving a hart, regardless of whether Program Buffer is implemented. The System Bus Access block uses physical addresses. The System Bus Access block may support 8-, 16-, 32-, 64-, and 128-bit accesses. [Table 4](#sbdatabits) shows which bits in `sbdata` are used for each access size. __Table 4\. System Bus Data Bits__ | Access Size | Data Bits | | ----------- | ---------------------------------------------------------------------------------------------- | | 8 | [sbdata0](#dm-sbdata0) bits 7:0 | | 16 | [sbdata0](#dm-sbdata0) bits 15:0 | | 32 | [sbdata0](#dm-sbdata0) | | 64 | [sbdata1](#dm-sbdata1), [sbdata0](#dm-sbdata0) | | 128 | [sbdata3](#dm-sbdata3), [sbdata2](#dm-sbdata2), [sbdata1](#dm-sbdata1), [sbdata0](#dm-sbdata0) | Depending on the microarchitecture, data accessed through System Bus Access might not always be coherent with that observed by each hart. It is up to the debugger to enforce coherency if the implementation does not. This specification does not define a standard way to do this. Possibilities may include writing to special memory-mapped locations, or executing special instructions via the Program Buffer. | | Implementing a System Bus Access block has several benefits even when a Debug Module also implements a Program Buffer. First, it is possible to access memory in a running system with minimal impact. Second, it may improve performance when accessing memory. Third, it may provide access to devices that a hart does not have access to. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#3-1-11-minimally-intrusive-debugging)3.1.11\. Minimally Intrusive Debugging Depending on the task it is performing, some harts can only be halted very briefly. There are several mechanisms that allow accessing resources in such a running system with a minimal impact on the running hart. First, an implementation may allow some abstract commands to execute without halting the hart. Second, the Quick Access abstract command can be used to halt a hart, quickly execute the contents of the Program Buffer, and let the hart run again. Combined with instructions that allow Program Buffer code to access the `data` registers, as described in [hartinfo](#dm-hartinfo), this can be used to quickly perform a memory or register access. For some hardware platforms this will be too intrusive, but many hardware platforms that can’t be halted can bear an occasional hiccup of a hundred or less cycles. Third, if the System Bus Access block is implemented, it can be used while a hart is running to access system memory. ### [](#3-1-12-security)3.1.12\. Security To protect intellectual property it may be desirable to lock access to the Debug Module. To allow access during a manufacturing process and not afterwards, a reasonable solution could be to add a fuse bit to the Debug Module that can be used to permanently disable it. Since this is technology specific, it is not further addressed in this spec. Another option is to allow the DM to be unlocked only by users who have an access key. Between [authenticated](#dmstatus-authenticated), [authbusy](#dmstatus-authbusy), and [authdata](#dm-authdata) arbitrarily complex authentication mechanism can be supported. When [authenticated](#dmstatus-authenticated) is clear, the DM must not interact with the rest of the hardware platform, nor expose details about the harts connected to the DM. All DM registers should read 0, while writes should be ignored, with the following mandatory exceptions: 1. [authenticated](#dmstatus-authenticated) in [dmstatus](#dm-dmstatus) is readable. 2. [authbusy](#dmstatus-authbusy) in [dmstatus](#dm-dmstatus) is readable. 3. [version](Sdtrig.html#tinfo-version) in [dmstatus](#dm-dmstatus) is readable. 4. [dmactive](#dmcontrol-dmactive) in [dmcontrol](#dm-dmcontrol) is readable and writable. 5. [authdata](#dm-authdata) is readable and writable. Implementations where it’s not possible to unlock the DM by using [authdata](#dm-authdata) should not implement that register. ### [](#3-1-13-version-detection)3.1.13\. Version Detection To detect the version of the Debug Module with a minimum of side effects, use the following procedure: 1. Read [dmcontrol](#dm-dmcontrol). 2. If [dmactive](#dmcontrol-dmactive) is 0 or [ndmreset](#dmcontrol-ndmreset) is 1: 1. Write [dmcontrol](#dm-dmcontrol), preserving [hartreset](#dmcontrol-hartreset), [hasel](#dmcontrol-hasel), [hartsello](#dmcontrol-hartsello), and [hartselhi](#dmcontrol-hartselhi) from the value that was read, setting [dmactive](#dmcontrol-dmactive), and clearing all the other bits. 2. Read [dmcontrol](#dm-dmcontrol) until [dmactive](#dmcontrol-dmactive) is high. 3. Read [dmstatus](#dm-dmstatus), which contains [version](#dmstatus-version). If it was necessary to clear [ndmreset](#dmcontrol-ndmreset), this might have the following side effects: 1. [haltreq](#dmcontrol-haltreq) is cleared, potentially preventing a halt request made by a previous debugger from taking effect. 2. [resumereq](#dmcontrol-resumereq) is cleared, potentially preventing a resume request made by a previous debugger from taking effect. 3. [ndmreset](#dmcontrol-ndmreset) is deasserted, releasing the hardware platform from reset if a previous debugger had set it. 4. [dmactive](#dmcontrol-dmactive) is asserted, releasing the DM from reset. This in itself is not observable by any harts. This procedure is guaranteed to work in future versions of this spec. The meaning of the [dmcontrol](#dm-dmcontrol) bits where [hartreset](#dmcontrol-hartreset), [hasel](#dmcontrol-hasel), [hartsello](#dmcontrol-hartsello), and [hartselhi](#dmcontrol-hartselhi) currently reside might change, but preserving them will have no side effects. Clearing the bits of [dmcontrol](#dm-dmcontrol) not explicitly mentioned here will have no side effects beyond the ones mentioned above. ### [](#debbus)3.1.14\. Debug Module Registers The registers described in this section are accessed over the DMI bus. Each DM has a base address (which is 0 for the first DM). The register addresses below are offsets from this base address. Debug Module DMI Registers that are unimplemented or not mentioned in the table below return 0 when read. Writing them has no effect. __Table 5\. Debug Module Debug Bus Registers__ | Address | Name | Section | | ------- | ------------------------------------------------------------------ | --------------------------------------------------------------------------- | | 0x04 | Abstract Data 0 ([data0](#dm-data0)) | [Abstract Data 0 (data0, at 0x04)](#dm-data0) | | 0x05 | Abstract Data 1 (data1) | | | 0x06 | Abstract Data 2 (data2) | | | 0x07 | Abstract Data 3 (data3) | | | 0x08 | Abstract Data 4 (data4) | | | 0x09 | Abstract Data 5 (data5) | | | 0x0a | Abstract Data 6 (data6) | | | 0x0b | Abstract Data 7 (data7) | | | 0x0c | Abstract Data 8 (data8) | | | 0x0d | Abstract Data 9 (data9) | | | 0x0e | Abstract Data 10 (data10) | | | 0x0f | Abstract Data 11 (data11) | | | 0x10 | Debug Module Control ([dmcontrol](#dm-dmcontrol)) | [Debug Module Control (dmcontrol, at 0x10)](#dm-dmcontrol) | | 0x11 | Debug Module Status ([dmstatus](#dm-dmstatus)) | [Debug Module Status (dmstatus, at 0x11)](#dm-dmstatus) | | 0x12 | Hart Info ([hartinfo](#dm-hartinfo)) | [Hart Info (hartinfo, at 0x12)](#dm-hartinfo) | | 0x13 | Halt Summary 1 ([haltsum1](#dm-haltsum1)) | [Halt Summary 1 (haltsum1, at 0x13)](#dm-haltsum1) | | 0x14 | Hart Array Window Select ([hawindowsel](#dm-hawindowsel)) | [Hart Array Window Select (hawindowsel, at 0x14)](#dm-hawindowsel) | | 0x15 | Hart Array Window ([hawindow](#dm-hawindow)) | [Hart Array Window (hawindow, at 0x15)](#dm-hawindow) | | 0x16 | Abstract Control and Status ([abstractcs](#dm-abstractcs)) | [Abstract Control and Status (abstractcs, at 0x16)](#dm-abstractcs) | | 0x17 | Abstract Command ([command](#dm-command)) | [Abstract Command (command, at 0x17)](#dm-command) | | 0x18 | Abstract Command Autoexec ([abstractauto](#dm-abstractauto)) | [Abstract Command Autoexec (abstractauto, at 0x18)](#dm-abstractauto) | | 0x19 | Configuration Structure Pointer 0 ([confstrptr0](#dm-confstrptr0)) | [Configuration Structure Pointer 0 (confstrptr0, at 0x19)](#dm-confstrptr0) | | 0x1a | Configuration Structure Pointer 1 ([confstrptr1](#dm-confstrptr1)) | [Configuration Structure Pointer 1 (confstrptr1, at 0x1a)](#dm-confstrptr1) | | 0x1b | Configuration Structure Pointer 2 ([confstrptr2](#dm-confstrptr2)) | [Configuration Structure Pointer 2 (confstrptr2, at 0x1b)](#dm-confstrptr2) | | 0x1c | Configuration Structure Pointer 3 ([confstrptr3](#dm-confstrptr3)) | [Configuration Structure Pointer 3 (confstrptr3, at 0x1c)](#dm-confstrptr3) | | 0x1d | Next Debug Module ([nextdm](#dm-nextdm)) | [Next Debug Module (nextdm, at 0x1d)](#dm-nextdm) | | 0x1f | Custom Features ([custom](#dm-custom)) | [Custom Features (custom, at 0x1f)](#dm-custom) | | 0x20 | Program Buffer 0 ([progbuf0](#dm-progbuf0)) | [Program Buffer 0 (progbuf0, at 0x20)](#dm-progbuf0) | | 0x21 | Program Buffer 1 (progbuf1) | | | 0x22 | Program Buffer 2 (progbuf2) | | | 0x23 | Program Buffer 3 (progbuf3) | | | 0x24 | Program Buffer 4 (progbuf4) | | | 0x25 | Program Buffer 5 (progbuf5) | | | 0x26 | Program Buffer 6 (progbuf6) | | | 0x27 | Program Buffer 7 (progbuf7) | | | 0x28 | Program Buffer 8 (progbuf8) | | | 0x29 | Program Buffer 9 (progbuf9) | | | 0x2a | Program Buffer 10 (progbuf10) | | | 0x2b | Program Buffer 11 (progbuf11) | | | 0x2c | Program Buffer 12 (progbuf12) | | | 0x2d | Program Buffer 13 (progbuf13) | | | 0x2e | Program Buffer 14 (progbuf14) | | | 0x2f | Program Buffer 15 (progbuf15) | | | 0x30 | Authentication Data ([authdata](#dm-authdata)) | [Authentication Data (authdata, at 0x30)](#dm-authdata) | | 0x32 | Debug Module Control and Status 2 ([dmcs2](#dm-dmcs2)) | [Debug Module Control and Status 2 (dmcs2, at 0x32)](#dm-dmcs2) | | 0x34 | Halt Summary 2 ([haltsum2](#dm-haltsum2)) | [Halt Summary 2 (haltsum2, at 0x34)](#dm-haltsum2) | | 0x35 | Halt Summary 3 ([haltsum3](#dm-haltsum3)) | [Halt Summary 3 (haltsum3, at 0x35)](#dm-haltsum3) | | 0x37 | System Bus Address 127:96 ([sbaddress3](#dm-sbaddress3)) | [System Bus Address 127:96 (sbaddress3, at 0x37)](#dm-sbaddress3) | | 0x38 | System Bus Access Control and Status ([sbcs](#dm-sbcs)) | [System Bus Access Control and Status (sbcs, at 0x38)](#dm-sbcs) | | 0x39 | System Bus Address 31:0 ([sbaddress0](#dm-sbaddress0)) | [System Bus Address 31:0 (sbaddress0, at 0x39)](#dm-sbaddress0) | | 0x3a | System Bus Address 63:32 ([sbaddress1](#dm-sbaddress1)) | [System Bus Address 63:32 (sbaddress1, at 0x3a)](#dm-sbaddress1) | | 0x3b | System Bus Address 95:64 ([sbaddress2](#dm-sbaddress2)) | [System Bus Address 95:64 (sbaddress2, at 0x3b)](#dm-sbaddress2) | | 0x3c | System Bus Data 31:0 ([sbdata0](#dm-sbdata0)) | [System Bus Data 31:0 (sbdata0, at 0x3c)](#dm-sbdata0) | | 0x3d | System Bus Data 63:32 ([sbdata1](#dm-sbdata1)) | [System Bus Data 63:32 (sbdata1, at 0x3d)](#dm-sbdata1) | | 0x3e | System Bus Data 95:64 ([sbdata2](#dm-sbdata2)) | [System Bus Data 95:64 (sbdata2, at 0x3e)](#dm-sbdata2) | | 0x3f | System Bus Data 127:96 ([sbdata3](#dm-sbdata3)) | [System Bus Data 127:96 (sbdata3, at 0x3f)](#dm-sbdata3) | | 0x40 | Halt Summary 0 ([haltsum0](#dm-haltsum0)) | [Halt Summary 0 (haltsum0, at 0x40)](#dm-haltsum0) | | 0x70 | Custom Features 0 ([custom0](#dm-custom0)) | [Custom Features 0 (custom0, at 0x70)](#dm-custom0) | | 0x71 | Custom Features 1 (custom1) | | | 0x72 | Custom Features 2 (custom2) | | | 0x73 | Custom Features 3 (custom3) | | | 0x74 | Custom Features 4 (custom4) | | | 0x75 | Custom Features 5 (custom5) | | | 0x76 | Custom Features 6 (custom6) | | | 0x77 | Custom Features 7 (custom7) | | | 0x78 | Custom Features 8 (custom8) | | | 0x79 | Custom Features 9 (custom9) | | | 0x7a | Custom Features 10 (custom10) | | | 0x7b | Custom Features 11 (custom11) | | | 0x7c | Custom Features 12 (custom12) | | | 0x7d | Custom Features 13 (custom13) | | | 0x7e | Custom Features 14 (custom14) | | | 0x7f | Custom Features 15 (custom15) | | #### [](#dm-dmstatus)Debug Module Status (dmstatus, at 0x11) This register reports status for the overall Debug Module as well as the currently selected harts, as defined in [hasel](#dmcontrol-hasel). Its address will not change in the future, because it contains [version](#dmstatus-version). This entire register is read-only. ![Diagram](_images/diag-6ac93ab32d3adcdea0b861b6801c4cc1b27f949d.svg) ![Diagram](_images/diag-c25030db883a32a4af25a8d8b5a84e896a87ee12.svg) ![Diagram](_images/diag-6e35566a0f66d5e796a7c7b1308cd42bb59638d2.svg) | Field | Description | Access | Reset | | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------ | ------ | | ndmresetpending | 0 (false): Unimplemented, or [ndmreset](#dmcontrol-ndmreset) is zero and no ndmreset is currently in progress. 1 (true): [ndmreset](#dmcontrol-ndmreset) is currently nonzero, or there is an ndmreset in progress. | **R** | \- | | stickyunavail | 0 (current): The per-hart unavail bits reflect the current state of the hart. 1 (sticky): The per-hart unavail bits are sticky. Once they are set, they will not clear until the debugger acknowledges them using [ackunavail](#dmcontrol-ackunavail). | **R** | Preset | | impebreak | If 1, then there is an implicit ebreak instruction at the non-existent word immediately after the Program Buffer. This saves the debugger from having to write the ebreak itself, and allows the Program Buffer to be one word smaller. This must be 1 when [progbufsize](#abstractcs-progbufsize) is 1. | **R** | Preset | | allhavereset | This field is 1 when all currently selected harts have been reset and reset has not been acknowledged for any of them. | **R** | \- | | anyhavereset | This field is 1 when at least one currently selected hart has been reset and reset has not been acknowledged for that hart. | **R** | \- | | allresumeack | This field is 1 when all currently selected harts have their resume ack bit set. | **R** | \- | | anyresumeack | This field is 1 when any currently selected hart has its resume ack bit set. | **R** | \- | | allnonexistent | This field is 1 when all currently selected harts do not exist in this hardware platform. | **R** | \- | | anynonexistent | This field is 1 when any currently selected hart does not exist in this hardware platform. | **R** | \- | | allunavail | This field is 1 when all currently selected harts are unavailable, or (if [stickyunavail](#dmstatus-stickyunavail) is 1) were unavailable without that being acknowledged. | **R** | \- | | anyunavail | This field is 1 when any currently selected hart is unavailable, or (if [stickyunavail](#dmstatus-stickyunavail) is 1) was unavailable without that being acknowledged. | **R** | \- | | allrunning | This field is 1 when all currently selected harts are running. | **R** | \- | | anyrunning | This field is 1 when any currently selected hart is running. | **R** | \- | | allhalted | This field is 1 when all currently selected harts are halted. | **R** | \- | | anyhalted | This field is 1 when any currently selected hart is halted. | **R** | \- | | authenticated | 0 (false): Authentication is required before using the DM. 1 (true): The authentication check has passed. On components that don’t implement authentication, this bit must be preset as 1. | **R** | Preset | | authbusy | 0 (ready): The authentication module is ready to process the next read/write to [authdata](#dm-authdata). 1 (busy): The authentication module is busy. Accessing [authdata](#dm-authdata) results in unspecified behavior. [authbusy](#dmstatus-authbusy) only becomes set in immediate response to an access to[authdata](#dm-authdata). | **R** | 0 | | hasresethaltreq | 1 if this Debug Module supports halt-on-reset functionality controllable by the [setresethaltreq](#dmcontrol-setresethaltreq) and [clrresethaltreq](#dmcontrol-clrresethaltreq) bits. 0 otherwise. | **R** | Preset | | confstrptrvalid | 0 (invalid): [confstrptr0](#dm-confstrptr0)\--[confstrptr3](#dm-confstrptr3) hold information which is not relevant to the configuration structure. 1 (valid): [confstrptr0](#dm-confstrptr0)\--[confstrptr3](#dm-confstrptr3) hold the address of the configuration structure. | **R** | Preset | | version | 0 (none): There is no Debug Module present. 1 (0.11): There is a Debug Module and it conforms to version 0.11 of this specification. 2 (0.13): There is a Debug Module and it conforms to version 0.13 of this specification. 3 (1.0): There is a Debug Module and it conforms to version 1.0 of this specification. 15 (custom): There is a Debug Module but it does not conform to any available version of this spec. | **R** | 3 | #### [](#dm-dmcontrol)Debug Module Control (dmcontrol, at 0x10) This register controls the overall Debug Module as well as the currently selected harts, as defined in [hasel](#dmcontrol-hasel). Throughout this document we refer to \`hartsel\`, which is [hartselhi](#dmcontrol-hartselhi)combined with [hartsello](#dmcontrol-hartsello). While the spec allows for 20 \`hartsel\` bits, an implementation may choose to implement fewer than that. The actual width of \`hartsel\` is called `HARTSELLEN`. It must be at least 0 and at most 20\. A debugger should discover `HARTSELLEN` by writing all ones to \`hartsel\` (assuming the maximum size) and reading back the value to see which bits were actually set. Debuggers must not change \`hartsel\` while an abstract command is executing. Hardware should enforce this by ignoring changes to \`hartsel\` while [busy](#abstractcs-busy) is set. | | There are separate [setresethaltreq](#dmcontrol-setresethaltreq) and [clrresethaltreq](#dmcontrol-clrresethaltreq) bits so that it is possible to write [dmcontrol](#dm-dmcontrol) without changing the halt-on-reset request bit for each selected hart, when not all selected harts have the same configuration. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | On any given write, a debugger may only write 1 to at most one of the following bits: [resumereq](#dmcontrol-resumereq), [hartreset](#dmcontrol-hartreset), [ackhavereset](#dmcontrol-ackhavereset),[setresethaltreq](#dmcontrol-setresethaltreq), and [clrresethaltreq](#dmcontrol-clrresethaltreq). The others must be written 0. \`resethaltreq\` is an optional internal bit of per-hart state that cannot be read, but can be written with [setresethaltreq](#dmcontrol-setresethaltreq) and [clrresethaltreq](#dmcontrol-clrresethaltreq). \`keepalive\` is an optional internal bit of per-hart state. When it is set, it suggests that the hardware should attempt to keep the hart available for the debugger, e.g. by keeping it from entering a low-power state once powered on. Even if the bit is implemented, hardware might not be able to keep a hart available. The bit is written through [setkeepalive](#dmcontrol-setkeepalive) and[clrkeepalive](#dmcontrol-clrkeepalive). For forward compatibility, [version](#dmstatus-version) will always be readable when bit 1 ([ndmreset](#dmcontrol-ndmreset)) is 0 and bit 0 ([dmactive](#dmcontrol-dmactive)) is 1. ![Diagram](_images/diag-a91720fc8423a02e71783603b4a658e2d208e1c6.svg) ![Diagram](_images/diag-fd5c324d7507e627bf8b3c8a4e61fc4a117b1cff.svg) | Field | Description | Access | Reset | | --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ----- | | haltreq | Writing 0 clears the halt request bit for all currently selected harts. This may cancel outstanding halt requests for those harts. Writing 1 sets the halt request bit for all currently selected harts. Running harts will halt whenever their halt request bit is set. Writes apply to the new value of \`hartsel\` and [hasel](#dmcontrol-hasel). Writes to this bit should be ignored while an abstract command is executing. | **WARZ** | \- | | resumereq | Writing 1 causes the currently selected harts to resume once, if they are halted when the write occurs. It also clears the resume ack bit for those harts. [resumereq](#dmcontrol-resumereq) is ignored if [haltreq](#dmcontrol-haltreq) is set. Writes apply to the new value of \`hartsel\` and [hasel](#dmcontrol-hasel). Writes to this bit should be ignored while an abstract command is executing. | **W1** | \- | | hartreset | This optional field writes the reset bit for all the currently selected harts. To perform a reset the debugger writes 1, and then writes 0 to deassert the reset signal. While this bit is 1, the debugger must not change which harts are selected. If this feature is not implemented, the bit always stays 0, so after writing 1 the debugger can read the register back to see if the feature is supported. Writes apply to the new value of \`hartsel\` and [hasel](#dmcontrol-hasel). | **WARL** | 0 | | ackhavereset | 0 (nop): No effect. 1 (ack): Clears havereset for any selected harts. Writes apply to the new value of \`hartsel\` and [hasel](#dmcontrol-hasel). Writes to this bit should be ignored while an abstract command is executing. | **W1** | \- | | ackunavail | 0 (nop): No effect. 1 (ack): Clears unavail for any selected harts that are currently available. Writes apply to the new value of \`hartsel\` and [hasel](#dmcontrol-hasel). | **W1** | \- | | hasel | Selects the definition of currently selected harts. 0 (single): There is a single currently selected hart, that is selected by \`hartsel\`. 1 (multiple): There may be multiple currently selected harts — the hart selected by \`hartsel\`, plus those selected by the hart array mask register. An implementation which does not implement the hart array mask register must tie this field to 0\. A debugger which wishes to use the hart array mask register feature should set this bit and read back to see if the functionality is supported. | **WARL** | 0 | | hartsello | The low 10 bits of \`hartsel\`: the DM-specific index of the hart to select. This hart is always part of the currently selected harts. | **WARL** | 0 | | hartselhi | The high 10 bits of \`hartsel\`: the DM-specific index of the hart to select. This hart is always part of the currently selected harts. | **WARL** | 0 | | setkeepalive | This optional field sets \`keepalive\` for all currently selected harts, unless [clrkeepalive](#dmcontrol-clrkeepalive) is simultaneously set to 1. Writes apply to the new value of \`hartsel\` and [hasel](#dmcontrol-hasel). | **W1** | \- | | clrkeepalive | This optional field clears \`keepalive\` for all currently selected harts. Writes apply to the new value of \`hartsel\` and [hasel](#dmcontrol-hasel). | **W1** | \- | | setresethaltreq | This optional field writes the halt-on-reset request bit for all currently selected harts, unless [clrresethaltreq](#dmcontrol-clrresethaltreq) is simultaneously set to 1\. When set to 1, each selected hart will halt upon the next deassertion of its reset. The halt-on-reset request bit is not automatically cleared. The debugger must write to [clrresethaltreq](#dmcontrol-clrresethaltreq) to clear it. Writes apply to the new value of \`hartsel\` and [hasel](#dmcontrol-hasel). If [hasresethaltreq](#dmstatus-hasresethaltreq) is 0, this field is not implemented. Writes to this bit should be ignored while an abstract command is executing. | **W1** | \- | | clrresethaltreq | This optional field clears the halt-on-reset request bit for all currently selected harts. Writes apply to the new value of \`hartsel\` and [hasel](#dmcontrol-hasel). Writes to this bit should be ignored while an abstract command is executing. | **W1** | \- | | ndmreset | This bit controls the reset signal from the DM to the rest of the hardware platform. The signal should reset every part of the hardware platform, including every hart, except for the DM and any logic required to access the DM. To perform a hardware platform reset the debugger writes 1, and then writes 0 to deassert the reset. | **R/W** | 0 | | dmactive | This bit serves as a reset signal for the Debug Module itself. After changing the value of this bit, the debugger must poll[dmcontrol](#dm-dmcontrol) until [dmactive](#dmcontrol-dmactive) has taken the requested value before performing any action that assumes the requested [dmactive](#dmcontrol-dmactive)state change has completed. Hardware may take an arbitrarily long time to complete activation or deactivation and will indicate completion by setting [dmactive](#dmcontrol-dmactive) to the requested value. During this time, the DM may ignore any register writes. 0 (inactive): The module’s state, including authentication mechanism, takes its reset values (the [dmactive](#dmcontrol-dmactive) bit is the only bit which can be written to something other than its reset value). Any accesses to the module may fail. Specifically, [version](#dmstatus-version) might not return correct data. When this value is written, the DM may ignore any other bits written to \`dmcontrol\` in the same write. 1 (active): The module functions normally. No other mechanism should exist that may result in resetting the Debug Module after power up. To place the Debug Module into a known state, a debugger should write 0 to [dmactive](#dmcontrol-dmactive), poll until [dmactive](#dmcontrol-dmactive) is observed 0, write 1 to [dmactive](#dmcontrol-dmactive), and poll until [dmactive](#dmcontrol-dmactive) is observed 1. Implementations may pay attention to this bit to further aid debugging, for example by preventing the Debug Module from being power gated while debugging is active. | **R/W** | 0 | #### [](#dm-hartinfo)Hart Info (hartinfo, at 0x12) This register gives information about the hart currently selected by \`hartsel\`. This register is optional. If it is not present it should read all-zero. If this register is included, the debugger can do more with the Program Buffer by writing programs which explicitly access the `data` and/or `dscratch`registers. This entire register is read-only. ![Diagram](_images/diag-c6b6ad070c9c018d4f57a66909426b30d95932af.svg) | Field | Description | Access | Reset | | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | ------ | | nscratch | Number of dscratch registers available for the debugger to use during program buffer execution, starting from [dscratch0](Sdext.html#csr-dscratch0). The debugger can make no assumptions about the contents of these registers between commands. | **R** | Preset | | dataaccess | 0 (csr): The data registers are shadowed in the hart by CSRs. Each CSR is DXLEN bits in size, and corresponds to a single argument, per [Table 1](#tab:datareg). 1 (memory): The data registers are shadowed in the hart’s memory map. Each register takes up 4 bytes in the memory map. | **R** | Preset | | datasize | If [dataaccess](#hartinfo-dataaccess) is 0: Number of CSRs dedicated to shadowing the data registers. If [dataaccess](#hartinfo-dataaccess) is 1: Number of 32-bit words in the memory map dedicated to shadowing the data registers. Since there are at most 12 data registers, the value in this register must be 12 or smaller. | **R** | Preset | | dataaddr | If [dataaccess](#hartinfo-dataaccess) is 0: The number of the first CSR dedicated to shadowing the data registers. If [dataaccess](#hartinfo-dataaccess) is 1: Address of RAM where the data registers are shadowed. This address is sign extended giving a range of -2048 to 2047, easily addressed with a load or store usingx0 as the address register. | **R** | Preset | #### [](#dm-hawindowsel)Hart Array Window Select (hawindowsel, at 0x14) This register selects which of the 32-bit portion of the hart array mask register (see [3.1.3.2\. Selecting Multiple Harts](#hartarraymask)) is accessible in [hawindow](#dm-hawindow). ![Diagram](_images/diag-a3fc8c64266edd9dd5fd2a77c79e7e104edb29da.svg) | Field | Description | Access | Reset | | ----------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ----- | | hawindowsel | The high bits of this field may be tied to 0, depending on how large the array mask register is. E.g. on a hardware platform with 48 harts only bit 0 of this field may actually be writable. | **WARL** | 0 | #### [](#dm-hawindow)Hart Array Window (hawindow, at 0x15) This register provides R/W access to a 32-bit portion of the hart array mask register (see [3.1.3.2\. Selecting Multiple Harts](#hartarraymask)). The position of the window is determined by [hawindowsel](#dm-hawindowsel). I.e. bit 0 refers to hart [hawindowsel](#dm-hawindowsel) , while bit 31 refers to hart[hawindowsel](#dm-hawindowsel) . Since some bits in the hart array mask register may be constant 0, some bits in this register may be constant 0, depending on the current value of [hawindowsel](#dm-hawindowsel). ![Diagram](_images/diag-d4030493b7889094f5006d045064eb1b0170840e.svg) #### [](#dm-abstractcs)Abstract Control and Status (abstractcs, at 0x16) Writing this register while an abstract command is executing causes[cmderr](#abstractcs-cmderr) to become 1 (busy) once the command completes ([busy](#abstractcs-busy) becomes 0). | | [datacount](#abstractcs-datacount) must be at least 1 to support RV32 harts, 2 to support RV64 harts, or 4 to support RV128 harts. | | ------------------------------------------------------------------------------------------------------------------------------------- | ![Diagram](_images/diag-6f6cc8d0de2886075ef7045e4ef9d8248bb08eac.svg) | Field | Description | Access | Reset | | ----------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------- | ------ | | progbufsize | Size of the Program Buffer, in 32-bit words. Valid sizes are 0 - 16. | **R** | Preset | | busy | 0 (ready): There is no abstract command currently being executed. 1 (busy): An abstract command is currently being executed. This bit is set as soon as [command](#dm-command) is written, and is not cleared until that command has completed. | **R** | 0 | | relaxedpriv | This optional bit controls whether program buffer and abstract memory accesses are performed with the exact and full set of permission checks that apply based on the current architectural state of the hart performing the access, or with a relaxed set of permission checks (e.g. PMP restrictions are ignored). The details of the latter are implementation-specific. 0 (full checks): Full permission checks apply. 1 (relaxed checks): Relaxed permission checks apply. | **WARL** | Preset | | cmderr | Gets set if an abstract command fails. The bits in this field remain set until they are cleared by writing 1 to them. No abstract command is started until the value is reset to 0. This field only contains a valid value if [busy](#abstractcs-busy) is 0. 0 (none): No error. 1 (busy): An abstract command was executing while [command](#dm-command),[abstractcs](#dm-abstractcs), or [abstractauto](#dm-abstractauto) was written, or when one of the data or progbuf registers was read or written. This status is only written if [cmderr](#abstractcs-cmderr) contains 0. 2 (not supported): The command in [command](#dm-command) is not supported. It may be supported with different options set, but it will not be supported at a later time when the hart or system state are different. 3 (exception): An exception occurred while executing the command (e.g. while executing the Program Buffer). 4 (halt/resume): The abstract command couldn’t execute because the hart wasn’t in the required state (running/halted), or unavailable. 5 (bus): The abstract command failed due to a bus error (e.g. alignment, access size, or timeout). 6 (reserved): Reserved for future use. 7 (other): The command failed for another reason. | **R/W1C** | 0 | | datacount | Number of data registers that are implemented as part of the abstract command interface. Valid sizes are 1 — 12. | **R** | Preset | #### [](#dm-command)Abstract Command (command, at 0x17) Writes to this register cause the corresponding abstract command to be executed. Writing this register while an abstract command is executing causes[cmderr](#abstractcs-cmderr) to become 1 (busy) once the command completes (busy becomes 0). If [cmderr](#abstractcs-cmderr) is non-zero, writes to this register are ignored. | | [cmderr](#abstractcs-cmderr) inhibits starting a new command to accommodate debuggers that, for performance reasons, send several commands to be executed in a row without checking [cmderr](#abstractcs-cmderr) in between. They can safely do so and check [cmderr](#abstractcs-cmderr) at the end without worrying that one command failed but then a later command (which might have depended on the previous one succeeding) passed. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ![Diagram](_images/diag-359cefb5985654832642f91ccf2304cda9d75337.svg) | Field | Description | Access | Reset | | ------- | -------------------------------------------------------------------------------------------- | -------- | ----- | | cmdtype | The type determines the overall functionality of this abstract command. | **WARZ** | 0 | | control | This field is interpreted in a command-specific manner, described for each abstract command. | **WARZ** | 0 | #### [](#dm-abstractauto)Abstract Command Autoexec (abstractauto, at 0x18) This register is optional. Including it allows more efficient burst accesses. A debugger can detect whether it is supported by setting bits and reading them back. If this register is implemented then bits corresponding to implemented progbuf and data registers must be writable. Other bits must be hard-wired to 0. If this register is written while an abstract command is executing then the write is ignored and[cmderr](#abstractcs-cmderr) becomes 1 (busy) once the command completes (busy becomes 0). ![Diagram](_images/diag-2cb67ab161a34575800d6a8e8d8c511d6d0169d3.svg) | Field | Description | Access | Reset | | --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ----- | | autoexecprogbuf | When a bit in this field is 1, read or write accesses to the corresponding progbuf word cause the DM to act as if the current value in [command](#dm-command) was written there again after the access to progbuf completes. | **WARL** | 0 | | autoexecdata | When a bit in this field is 1, read or write accesses to the corresponding data word cause the DM to act as if the current value in [command](#dm-command) was written there again after the access to data completes. | **WARL** | 0 | #### [](#dm-confstrptr0)Configuration Structure Pointer 0 (confstrptr0, at 0x19) When [confstrptrvalid](#dmstatus-confstrptrvalid) is set, reading this register returns bits 31:0 of the configuration structure pointer. Reading the other `confstrptr`registers returns the upper bits of the address. When system bus access is implemented, this must be an address that can be used with the System Bus Access module. Otherwise, this must be an address that can be used to access the configuration structure from the hart with ID 0. If [confstrptrvalid](#dmstatus-confstrptrvalid) is 0, then the `confstrptr` registers hold identifier information which is not further specified in this document. The configuration structure itself is a data structure of the same format as the data structure pointed to by `mconfigptr` as described in the Privileged Spec. This entire register is read-only. ![Diagram](_images/diag-f66de0aacb34d1ce06c6e967770416dd40d365d7.svg) #### [](#dm-confstrptr1)Configuration Structure Pointer 1 (confstrptr1, at 0x1a) When [confstrptrvalid](#dmstatus-confstrptrvalid) is set, reading this register returns bits 63:32 of the configuration structure pointer. See [confstrptr0](#dm-confstrptr0) for more details. This entire register is read-only. ![Diagram](_images/diag-f66de0aacb34d1ce06c6e967770416dd40d365d7.svg) #### [](#dm-confstrptr2)Configuration Structure Pointer 2 (confstrptr2, at 0x1b) When [confstrptrvalid](#dmstatus-confstrptrvalid) is set, reading this register returns bits 95:64 of the configuration structure pointer. See [confstrptr0](#dm-confstrptr0) for more details. This entire register is read-only. ![Diagram](_images/diag-f66de0aacb34d1ce06c6e967770416dd40d365d7.svg) #### [](#dm-confstrptr3)Configuration Structure Pointer 3 (confstrptr3, at 0x1c) When [confstrptrvalid](#dmstatus-confstrptrvalid) is set, reading this register returns bits 127:96 of the configuration structure pointer. See [confstrptr0](#dm-confstrptr0) for more details. This entire register is read-only. ![Diagram](_images/diag-f66de0aacb34d1ce06c6e967770416dd40d365d7.svg) #### [](#dm-nextdm)Next Debug Module (nextdm, at 0x1d) If there is more than one DM accessible on this DMI, this register contains the base address of the next one in the chain, or 0 if this is the last one in the chain. This entire register is read-only. ![Diagram](_images/diag-f66de0aacb34d1ce06c6e967770416dd40d365d7.svg) #### [](#dm-data0)Abstract Data 0 (data0, at 0x04) [data0](#dm-data0) through data11 are registers that may be read or changed by abstract commands. [datacount](#abstractcs-datacount) indicates how many of them are implemented, starting at [data0](#dm-data0), counting up.[Table 1](#tab:datareg) shows how abstract commands use these registers. Accessing these registers while an abstract command is executing causes[cmderr](#abstractcs-cmderr) to be set to 1 (busy) if it is 0. Attempts to write them while [busy](#abstractcs-busy) is set does not change their value. The values in these registers might not be preserved after an abstract command is executed. The only guarantees on their contents are the ones offered by the command in question. If the command fails, no assumptions can be made about the contents of these registers. ![Diagram](_images/diag-f2d717e90e466d6b6d8c7ac7e3d3b05b946496b3.svg) #### [](#dm-progbuf0)Program Buffer 0 (progbuf0, at 0x20) [progbuf0](#dm-progbuf0) through progbuf15 must provide write access to the optional program buffer. It may also be possible for the debugger to read from the program buffer through these registers. If reading is not supported, then all reads return 0. [progbufsize](#abstractcs-progbufsize) indicates how many `progbuf` registers are implemented starting at [progbuf0](#dm-progbuf0), counting up. Accessing these registers while an abstract command is executing causes[cmderr](#abstractcs-cmderr) to be set to 1 (busy) if it is 0. Attempts to write them while [busy](#abstractcs-busy) is set does not change their value. ![Diagram](_images/diag-f2d717e90e466d6b6d8c7ac7e3d3b05b946496b3.svg) #### [](#dm-authdata)Authentication Data (authdata, at 0x30) This register serves as a 32-bit serial port to/from the authentication module. When [authbusy](#dmstatus-authbusy) is clear, the debugger can communicate with the authentication module by reading or writing this register. There is no separate mechanism to signal overflow/underflow. ![Diagram](_images/diag-f2d717e90e466d6b6d8c7ac7e3d3b05b946496b3.svg) #### [](#dm-dmcs2)Debug Module Control and Status 2 (dmcs2, at 0x32) This register contains DM control and status bits that didn’t easily fit in [dmcontrol](#dm-dmcontrol) and [dmstatus](#dm-dmstatus). All are optional. If halt groups are not implemented, then [group](#dmcs2-group) will always be 0 when [grouptype](#dmcs2-grouptype) is 0. If resume groups are not implemented, then [grouptype](#dmcs2-grouptype) will remain 0 even after 1 is written there. The DM external triggers available to add to halt groups may be the same as or distinct from the DM external triggers available to add to resume groups. ![Diagram](_images/diag-3d07395c0f4b12d1c1e1c83f4b178474ed45d46a.svg) | Field | Description | Access | Reset | | ------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------- | ------ | | grouptype | 0 (halt): The remaining fields in this register configure halt groups. 1 (resume): The remaining fields in this register configure resume groups. | **WARL** | 0 | | dmexttrigger | This field contains the currently selected DM external trigger. If a non-existent trigger value is written here, the hardware will change it to a valid one or 0 if no DM external triggers exist. | **WARL** | 0 | | group | When [hgselect](#dmcs2-hgselect) is 0, contains the group of the hart specified by \`hartsel\`. When [hgselect](#dmcs2-hgselect) is 1, contains the group of the DM external trigger selected by [dmexttrigger](#dmcs2-dmexttrigger). The value written to this field is ignored unless [hgwrite](#dmcs2-hgwrite)is also written 1. Group numbers are contiguous starting at 0, with the highest number being implementation-dependent, and possibly different between different group types. Debuggers should read back this field after writing to confirm they are using a hart group that is supported. If groups aren’t implemented, then this entire field is 0. | **WARL** | preset | | hgwrite | When 1 is written and [hgselect](#dmcs2-hgselect) is 0, for every selected hart the DM will change its group to the value written to [group](#dmcs2-group), if the hardware supports that group for that hart. Implementations may also change the group of a minimal set of unselected harts in the same way, if that is necessary due to a hardware limitation. When 1 is written and [hgselect](#dmcs2-hgselect) is 1, the DM will change the group of the DM external trigger selected by [dmexttrigger](#dmcs2-dmexttrigger)to the value written to [group](#dmcs2-group), if the hardware supports that group for that trigger. Writing 0 has no effect. | **W1** | \- | | hgselect | 0 (harts): Operate on harts. 1 (triggers): Operate on DM external triggers. If there are no DM external triggers, this field must be tied to 0. | **WARL** | 0 | #### [](#dm-haltsum0)Halt Summary 0 (haltsum0, at 0x40) Each bit in this read-only register indicates whether one specific hart is halted or not. Unavailable/nonexistent harts are not considered to be halted. This register might not be present if fewer than 2 harts are connected to this DM. The LSB reflects the halt status of hart {hartsel\[19:5\],5’h0}, and the MSB reflects halt status of hart {hartsel\[19:5\],5’h1f}. This entire register is read-only. ![Diagram](_images/diag-6c8abd34675729d546cf598369a0d344d0f0c53f.svg) #### [](#dm-haltsum1)Halt Summary 1 (haltsum1, at 0x13) Each bit in this read-only register indicates whether any of a group of harts is halted or not. Unavailable/nonexistent harts are not considered to be halted. This register might not be present if fewer than 33 harts are connected to this DM. The LSB reflects the halt status of harts {hartsel\[19:10\],10’h0} through {hartsel\[19:10\],10’h1f}. The MSB reflects the halt status of harts {hartsel\[19:10\],10’h3e0} through {hartsel\[19:10\],10’h3ff}. This entire register is read-only. ![Diagram](_images/diag-5ebf680e3e773a9616e6a8a25c843d42f5aedc7c.svg) #### [](#dm-haltsum2)Halt Summary 2 (haltsum2, at 0x34) Each bit in this read-only register indicates whether any of a group of harts is halted or not. Unavailable/nonexistent harts are not considered to be halted. This register might not be present if fewer than 1025 harts are connected to this DM. The LSB reflects the halt status of harts {hartsel\[19:15\],15’h0} through {hartsel\[19:15\],15’h3ff}. The MSB reflects the halt status of harts {hartsel\[19:15\],15’h7c00} through {hartsel\[19:15\],15’h7fff}. This entire register is read-only. ![Diagram](_images/diag-0bf0db8429c95dff6e195a228a2d69766b438744.svg) #### [](#dm-haltsum3)Halt Summary 3 (haltsum3, at 0x35) Each bit in this read-only register indicates whether any of a group of harts is halted or not. Unavailable/nonexistent harts are not considered to be halted. This register might not be present if fewer than 32769 harts are connected to this DM. The LSB reflects the halt status of harts 20’h0 through 20’h7fff. The MSB reflects the halt status of harts 20’hf8000 through 20’hfffff. This entire register is read-only. ![Diagram](_images/diag-f966beca56b0c0e6b4e83724ddba1be474d08a45.svg) #### [](#dm-sbcs)System Bus Access Control and Status (sbcs, at 0x38) ![Diagram](_images/diag-77712693974dbfecfdc44b6adc5799d040affdab.svg) ![Diagram](_images/diag-ead554efa269a45940d813ac35ea535f325c6629.svg) | Field | Description | Access | Reset | | --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------- | ------ | | sbversion | 0 (legacy): The System Bus interface conforms to mainline drafts of this spec older than 1 January, 2018. 1 (1.0): The System Bus interface conforms to this version of the spec. Other values are reserved for future versions. | **R** | 1 | | sbbusyerror | Set when the debugger attempts to read data while a read is in progress, or when the debugger initiates a new access while one is already in progress (while [sbbusy](#sbcs-sbbusy) is set). It remains set until it’s explicitly cleared by the debugger. While this field is set, no more system bus accesses can be initiated by the Debug Module. | **R/W1C** | 0 | | sbbusy | When 1, indicates the system bus manager is busy. (Whether the system bus itself is busy is related, but not the same thing.) This bit goes high immediately when a read or write is requested for any reason, and does not go low until the access is fully completed. Writes to [sbcs](#dm-sbcs) while [sbbusy](#sbcs-sbbusy) is high result in undefined behavior. A debugger must not write to [sbcs](#dm-sbcs) until it reads[sbbusy](#sbcs-sbbusy) as 0. | **R** | 0 | | sbreadonaddr | When 1, every write to [sbaddress0](#dm-sbaddress0) automatically triggers a system bus read at the new address. | **R/W** | 0 | | sbaccess | Select the access size to use for system bus accesses. 0 (8bit): 8-bit 1 (16bit): 16-bit 2 (32bit): 32-bit 3 (64bit): 64-bit 4 (128bit): 128-bit If [sbaccess](#sbcs-sbaccess) has an unsupported value when the DM starts a bus access, the access is not performed and [sberror](#sbcs-sberror) is set to 4. | **R/W** | 2 | | sbautoincrement | When 1, sbaddress is incremented by the access size (in bytes) selected in [sbaccess](#sbcs-sbaccess) after every system bus access. | **R/W** | 0 | | sbreadondata | When 1, every read from [sbdata0](#dm-sbdata0) automatically triggers a system bus read at the (possibly auto-incremented) address. | **R/W** | 0 | | sberror | When the Debug Module’s system bus manager encounters an error, this field gets set. The bits in this field remain set until they are cleared by writing 1 to them. While this field is non-zero, no more system bus accesses can be initiated by the Debug Module. An implementation may report \`\`Other'' (7) for any error condition. 0 (none): There was no bus error. 1 (timeout): There was a timeout. 2 (address): A bad address was accessed. 3 (alignment): There was an alignment error. 4 (size): An access of unsupported size was requested. 7 (other): Other. | **R/W1C** | 0 | | sbasize | Width of system bus addresses in bits. (0 indicates there is no bus access support.) | **R** | Preset | | sbaccess128 | 1 when 128-bit system bus accesses are supported. | **R** | Preset | | sbaccess64 | 1 when 64-bit system bus accesses are supported. | **R** | Preset | | sbaccess32 | 1 when 32-bit system bus accesses are supported. | **R** | Preset | | sbaccess16 | 1 when 16-bit system bus accesses are supported. | **R** | Preset | | sbaccess8 | 1 when 8-bit system bus accesses are supported. | **R** | Preset | #### [](#dm-sbaddress0)System Bus Address 31:0 (sbaddress0, at 0x39) If [sbasize](#sbcs-sbasize) is 0, then this register is not present. When the system bus manager is busy, writes to this register will set[sbbusyerror](#sbcs-sbbusyerror) and don’t do anything else. If [sberror](#sbcs-sberror) is 0, [sbbusyerror](#sbcs-sbbusyerror) is 0, and [sbreadonaddr](#sbcs-sbreadonaddr) is set then writes to this register start the following: 1. Set [sbbusy](#sbcs-sbbusy). 2. Perform a bus read from the new value of `sbaddress`. 3. If the read succeeded and [sbautoincrement](#sbcs-sbautoincrement) is set, increment `sbaddress`. 4. Clear [sbbusy](#sbcs-sbbusy). ![Diagram](_images/diag-a4b266f5557d7c97a14557595d84333d931c144a.svg) | Field | Description | Access | Reset | | ------- | -------------------------------------------------------- | ------- | ----- | | address | Accesses bits 31:0 of the physical address in sbaddress. | **R/W** | 0 | #### [](#dm-sbaddress1)System Bus Address 63:32 (sbaddress1, at 0x3a) If [sbasize](#sbcs-sbasize) is less than 33, then this register is not present. When the system bus manager is busy, writes to this register will set[sbbusyerror](#sbcs-sbbusyerror) and don’t do anything else. ![Diagram](_images/diag-a4b266f5557d7c97a14557595d84333d931c144a.svg) | Field | Description | Access | Reset | | ------- | -------------------------------------------------------------------------------------------------- | ------- | ----- | | address | Accesses bits 63:32 of the physical address in sbaddress (if the system address bus is that wide). | **R/W** | 0 | #### [](#dm-sbaddress2)System Bus Address 95:64 (sbaddress2, at 0x3b) If [sbasize](#sbcs-sbasize) is less than 65, then this register is not present. When the system bus manager is busy, writes to this register will set[sbbusyerror](#sbcs-sbbusyerror) and don’t do anything else. ![Diagram](_images/diag-a4b266f5557d7c97a14557595d84333d931c144a.svg) | Field | Description | Access | Reset | | ------- | -------------------------------------------------------------------------------------------------- | ------- | ----- | | address | Accesses bits 95:64 of the physical address in sbaddress (if the system address bus is that wide). | **R/W** | 0 | #### [](#dm-sbaddress3)System Bus Address 127:96 (sbaddress3, at 0x37) If [sbasize](#sbcs-sbasize) is less than 97, then this register is not present. When the system bus manager is busy, writes to this register will set[sbbusyerror](#sbcs-sbbusyerror) and don’t do anything else. ![Diagram](_images/diag-a4b266f5557d7c97a14557595d84333d931c144a.svg) | Field | Description | Access | Reset | | ------- | --------------------------------------------------------------------------------------------------- | ------- | ----- | | address | Accesses bits 127:96 of the physical address in sbaddress (if the system address bus is that wide). | **R/W** | 0 | #### [](#dm-sbdata0)System Bus Data 31:0 (sbdata0, at 0x3c) If all of the `sbaccess` bits in [sbcs](#dm-sbcs) are 0, then this register is not present. Any successful system bus read updates `sbdata`. If the width of the read access is less than the width of `sbdata`, the contents of the remaining high bits may take on any value. If either [sberror](#sbcs-sberror) or [sbbusyerror](#sbcs-sbbusyerror) isn’t 0 then accesses do nothing. If the bus manager is busy then accesses set [sbbusyerror](#sbcs-sbbusyerror), and don’t do anything else. Writes to this register start the following: 1. Set [sbbusy](#sbcs-sbbusy). 2. Perform a bus write of the new value of `sbdata` to `sbaddress`. 3. If the write succeeded and [sbautoincrement](#sbcs-sbautoincrement) is set, increment `sbaddress`. 4. Clear [sbbusy](#sbcs-sbbusy). Reads from this register start the following: 1. "Return" the data. 2. Set [sbbusy](#sbcs-sbbusy). 3. If [sbreadondata](#sbcs-sbreadondata) is set: 1. Perform a system bus read from the address contained in`sbaddress`, placing the result in `sbdata`. 2. If [sbautoincrement](#sbcs-sbautoincrement) is set and the read was successful, increment `sbaddress`. 4. Clear [sbbusy](#sbcs-sbbusy). Only [sbdata0](#dm-sbdata0) has this behavior. The other `sbdata` registers have no side effects. On systems that have buses wider than 32 bits, a debugger should access [sbdata0](#dm-sbdata0) after accessing the other `sbdata`registers. ![Diagram](_images/diag-f2d717e90e466d6b6d8c7ac7e3d3b05b946496b3.svg) | Field | Description | Access | Reset | | ----- | ----------------------------- | ------- | ----- | | data | Accesses bits 31:0 of sbdata. | **R/W** | 0 | #### [](#dm-sbdata1)System Bus Data 63:32 (sbdata1, at 0x3d) If [sbaccess64](#sbcs-sbaccess64) and [sbaccess128](#sbcs-sbaccess128) are 0, then this register is not present. If the bus manager is busy then accesses set [sbbusyerror](#sbcs-sbbusyerror), and don’t do anything else. ![Diagram](_images/diag-f2d717e90e466d6b6d8c7ac7e3d3b05b946496b3.svg) | Field | Description | Access | Reset | | ----- | --------------------------------------------------------------- | ------- | ----- | | data | Accesses bits 63:32 of sbdata (if the system bus is that wide). | **R/W** | 0 | #### [](#dm-sbdata2)System Bus Data 95:64 (sbdata2, at 0x3e) This register only exists if [sbaccess128](#sbcs-sbaccess128) is 1. If the bus manager is busy then accesses set [sbbusyerror](#sbcs-sbbusyerror), and don’t do anything else. ![Diagram](_images/diag-f2d717e90e466d6b6d8c7ac7e3d3b05b946496b3.svg) | Field | Description | Access | Reset | | ----- | --------------------------------------------------------------- | ------- | ----- | | data | Accesses bits 95:64 of sbdata (if the system bus is that wide). | **R/W** | 0 | #### [](#dm-sbdata3)System Bus Data 127:96 (sbdata3, at 0x3f) This register only exists if [sbaccess128](#sbcs-sbaccess128) is 1. If the bus manager is busy then accesses set [sbbusyerror](#sbcs-sbbusyerror), and don’t do anything else. ![Diagram](_images/diag-f2d717e90e466d6b6d8c7ac7e3d3b05b946496b3.svg) | Field | Description | Access | Reset | | ----- | ---------------------------------------------------------------- | ------- | ----- | | data | Accesses bits 127:96 of sbdata (if the system bus is that wide). | **R/W** | 0 | #### [](#dm-custom)Custom Features (custom, at 0x1f) This optional register may be used for non-standard features. Future version of the debug spec will not use this address. #### [](#dm-custom0)Custom Features 0 (custom0, at 0x70) The optional [custom0](#dm-custom0) through custom15 registers may be used for non-standard features. Future versions of the debug spec will not use these addresses. 6.1. Debug Transport Module (DTM) (non-ISA extension) ==================== ## [](#dtm)6.1\. Debug Transport Module (DTM) (non-ISA extension) Debug Transport Modules provide access to the DM over one or more transports (e.g. JTAG or USB). There may be multiple DTMs in a single hardware platform. Ideally every component that communicates with the outside world includes a DTM, allowing a hardware platform to be debugged through every transport it supports. For instance a USB component could include a DTM. This would trivially allow any hardware platform to be debugged over USB. All that is required is that the USB module already in use also has access to the Debug Module Interface. Using multiple DTMs at the same time is not supported. It is left to the user to ensure this does not happen. This specification defines a JTAG DTM in [6.1.1\. JTAG Debug Transport Module](#sec:jtagdtm). Additional DTMs may be added in future versions of this specification. An implementation can be compatible with this specification without implementing any of this section. In that case it must be advertised as conforming to "RISC-V Debug Specification, with custom DTM." If the JTAG DTM described here is implemented, it must be advertised as conforming to the "RISC-V Debug Specification, with JTAG DTM."" ### [](#sec:jtagdtm)6.1.1\. JTAG Debug Transport Module This Debug Transport Module is based around a normal JTAG Test Access Port (TAP). The JTAG TAP allows access to arbitrary JTAG registers by first selecting one using the JTAG instruction register (IR), and then accessing it through the JTAG data register (DR). #### [](#6-1-1-1-jtag-background)6.1.1.1\. JTAG Background JTAG refers to IEEE Std 1149.1-2013\. It is a standard that defines test logic that can be included in an integrated circuit to test the interconnections between integrated circuits, test the integrated circuit itself, and observe or modify circuit activity during the component’s normal operation. This specification uses the latter functionality. The JTAG standard defines a Test Access Port (TAP) that can be used to read and write a few custom registers, which can be used to communicate with debug hardware in a component. #### [](#6-1-1-2-jtag-dtm-registers)6.1.1.2\. JTAG DTM Registers JTAG TAPs used as a DTM must have an IR of at least 5 bits. When the TAP is reset, IR must default to 00001, selecting the IDCODE instruction. A full list of JTAG registers along with their encoding is in[Table 1](#tab:jtag%5Fregisters). If the IR actually has more than 5 bits, then the encodings in[Table 1](#tab:jtag%5Fregisters) should be extended with 0’s in their most significant bits, except for the 0x1f encoding of BYPASS, which must be extended with 1’s in the most significant bits. The only regular JTAG registers a debugger might use are BYPASS and IDCODE, but this specification leaves IR space for many other standard JTAG instructions. Unimplemented instructions must select the BYPASS register. __Table 1\. JTAG DTM TAP Registers__ | Address | Name | Description | Section | | ------- | ----------------------------------------------- | -------------------------------------- | -------------------------------------------------------- | | 0x00 | [bypass](#dtm-bypass) | JTAG recommends this encoding | | | 0x01 | [idcode](#dtm-idcode) | To identify a specific silicon version | [IDCODE (at 0x01)](#dtm-idcode) | | 0x10 | DTM Control and Status ([dtmcs](#dtm-dtmcs)) | For Debugging | [DTM Control and Status (dtmcs, at 0x10)](#dtm-dtmcs) | | 0x11 | Debug Module Interface Access ([dmi](#dtm-dmi)) | For Debugging | [Debug Module Interface Access (dmi, at 0x11)](#dtm-dmi) | | 0x12 | reserved (bypass) | Reserved for future RISC-V debugging | | | 0x13 | reserved (bypass) | Reserved for future RISC-V debugging | | | 0x14 | reserved (bypass) | Reserved for future RISC-V debugging | | | 0x15 | reserved (bypass) | Reserved for future RISC-V standards | | | 0x16 | reserved (bypass) | Reserved for future RISC-V standards | | | 0x17 | reserved (bypass) | Reserved for future RISC-V standards | | | 0x1f | [bypass](#dtm-bypass) | JTAG requires this encoding | [BYPASS (at 0x1f)](#dtm-bypass) | #### [](#dtm-idcode)`IDCODE` (at 0x01) This register is selected (in IR) when the TAP state machine is reset. Its definition is exactly as defined in IEEE Std 1149.1-2013. This entire register is read-only. ![Diagram](_images/diag-5cc0157aa6d822d0267b32118e279ddfb8b6783c.svg) | Field | Description | Access | Reset | | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | ------ | | Version | Identifies the release version of this part. | **R** | Preset | | PartNumber | Identifies the designer’s part number of this part. | **R** | Preset | | ManufId | Identifies the designer/manufacturer of this part. Bits 6:0 must be bits 6:0 of the designer/manufacturer’s Identification Code as assigned by JEDEC Standard JEP106\. Bits 10:7 contain the modulo-16 count of the number of continuation characters (0x7f) in that same Identification Code. | **R** | Preset | #### [](#dtm-dtmcs)DTM Control and Status (dtmcs, at 0x10) The size of this register will remain constant in future versions so that a debugger can always determine the version of the DTM. ![Diagram](_images/diag-832347df00b446c97ba1a81ca23e405280cbf748.svg) | Field | Description | Access | Reset | | ------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | ------ | | errinfo | This optional field may provide additional detail about an error that occurred when communicating with a DM. It is updated whenever[op](#dmi-op) is updated by the hardware or when 1 is written to[dmireset](#dtmcs-dmireset). 0 (not implemented): This field is not implemented. 1 (dmi error): There was an error between the DTM and DMI. 2 (communication error): There was an error between the DMI and a DMI subordinate. 3 (device error): The DMI subordinate reported an error. 4 (unknown): There is no error to report, or no further information available about the error. This is the reset value if the field is implemented. Other values are reserved for future use by this specification. | **R** | 4 | | dtmhardreset | Writing 1 to this bit does a hard reset of the DTM, causing the DTM to forget about any outstanding DMI transactions, and returning all registers and internal state to their reset value. In general this should only be used when the Debugger has reason to expect that the outstanding DMI transaction will never complete (e.g. a reset condition caused an inflight DMI transaction to be cancelled). | **W1** | \- | | dmireset | Writing 1 to this bit clears the sticky error state and resets[errinfo](#dtmcs-errinfo), but does not affect outstanding DMI transactions. | **W1** | \- | | idle | This is a hint to the debugger of the minimum number of cycles a debugger should spend in Run-Test/Idle after every DMI scan to avoid a \`busy' return code ([dmistat](#dtmcs-dmistat) of 3). A debugger must still check [dmistat](#dtmcs-dmistat) when necessary. 0: It is not necessary to enter Run-Test/Idle at all. 1: Enter Run-Test/Idle and leave it immediately. 2: Enter Run-Test/Idle and stay there for 1 cycle before leaving. And so on. | **R** | Preset | | dmistat | Read-only alias of [op](#dmi-op). | **R** | 0 | | abits | The size of [address](#dmi-address) in [dmi](#dtm-dmi). | **R** | Preset | | version | 0 (0.11): Version described in spec version 0.11. 1 (1.0): Version described in spec versions 0.13 and 1.0. 15 (custom): Version not described in any available version of this spec. | **R** | 1 | #### [](#dtm-dmi)Debug Module Interface Access (dmi, at 0x11) This register allows access to the Debug Module Interface (DMI). In Update-DR, the DTM starts the operation specified in [op](#dmi-op) unless the current status reported in [op](#dmi-op) is sticky. In Capture-DR, the DTM updates [data](debug%5Fmodule.html#sbdata0-data) with the result from that operation, updating [op](#dmi-op) if the current [op](#dmi-op) isn’t sticky. See [\[dmiaccess\]](#dmiaccess) for examples of how this is used. | | The still-in-progress status is sticky to accommodate debuggers that batch together a number of scans, which must all be executed or stop as soon as there’s a problem. For instance a series of scans may write a Debug Program and execute it. If one of the writes fails but the execution continues, then the Debug Program may hang or have other unexpected side effects. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ![Diagram](_images/diag-9cdd36123e1627b68fb34eb4bfbe63ac53ee2627.svg) | Field | Description | Access | Reset | | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------- | ----- | | address | Address used for DMI access. In Update-DR this value is used to access the DM over the DMI.[op](#dmi-op) defines what this register contains after every possible operation. | **R/W** | 0 | | data | The data to send to the DM over the DMI during Update-DR, and the data returned from the DM as a result of the previous operation. | **R/W** | 0 | | op | When the debugger writes this field, it has the following meaning: 0 (nop): Ignore [data](debug%5Fmodule.html#sbdata0-data) and [address](debug%5Fmodule.html#sbaddress0-address). Don’t send anything over the DMI during Update-DR. This operation should never affect DMI busy or error status. The address and data reported in the following Capture-DR are undefined. This operation leaves the values in [address](#dmi-address) and [data](#dmi-data)UNSPECIFIED. 1 (read): Read from [address](#dmi-address). When this operation succeeds, [address](#dmi-address) contains the address that was read from, and [data](#dmi-data) contains the value that was read. 2 (write): Write [data](#dmi-data) to [address](#dmi-address). This operation leaves the values in [address](#dmi-address) and [data](#dmi-data)UNSPECIFIED. 3 (reserved): Reserved. When the debugger reads this field, it means the following: 0 (success): The previous operation completed successfully. 1 (reserved): Reserved. 2 (failed): A previous operation failed. The data scanned into [dmi](#dtm-dmi) in this access will be ignored. This status is sticky and can be cleared by writing [dmireset](#dtmcs-dmireset) in [dtmcs](#dtm-dtmcs). This indicates that the DM itself or the DMI responded with an error. There are no specified cases in which the DM would respond with an error, and DMI is not required to support returning errors. If a debugger sees this status, there might be additional information in [errinfo](#dtmcs-errinfo). 3 (busy): A DMI operation was attempted while a prior DMI operation was still in progress. The data scanned into [dmi](#dtm-dmi) in this access will be ignored. This status is sticky and can be cleared by writing[dmireset](#dtmcs-dmireset) in [dtmcs](#dtm-dtmcs). If a debugger sees this status, it needs to give the target more TCK edges between Update-DR and Capture-DR. The simplest way to do that is to add extra transitions in Run-Test/Idle. | **R/W** | 0 | #### [](#dtm-bypass)`BYPASS` (at 0x1f) 1-bit register that has no effect. It is used when a debugger does not want to communicate with this TAP. This entire register is read-only. ![Diagram](_images/diag-384bd371721d4235c02b312b20cdd445d7a86797.svg) #### [](#6-1-1-3-jtag-connector)6.1.1.3\. JTAG Connector ##### [](#6-1-1-3-1-recommended-jtag-connector)6.1.1.3.1\. Recommended JTAG Connector To make it easy to acquire debug hardware, this spec recommends a connector that is compatible with the MIPI-10 .05 inch connector specification, as described in MIPI Debug & Trace Connector Recommendations, Version 1.20, 2 July 2021. The connector has .05 inch spacing, gold-plated male header with .016 inch thick hardened copper or beryllium bronze square posts (SAMTEC FTSH or equivalent). Female connectors are compatible gold connectors. Viewing the male header from above (the pins pointing at your eye), a target’s connector looks as it does in [Table 2](#tab:mipiten). The function of each pin is described in [Table 3](#tab:pinout). __Table 2\. MIPI 10-pin JTAG + nRESET Connector Diagram__ | VREF DEBUG | 1 | 2 | TMS | | ---------- | - | -- | ------ | | GND | 3 | 4 | TCK | | GND | 5 | 6 | TDO | | GND or KEY | 7 | 8 | TDI | | GND | 9 | 10 | nRESET | If a hardware platform requires nTRST then it is permissible to reuse the nRESET pin as the nTRST signal, resulting in a MIPI 10-pin JTAG nTRST connector. ##### [](#6-1-1-3-2-alternate-jtag-connector)6.1.1.3.2\. Alternate JTAG Connector The MIPI-10 connector should provide plenty of signals for all modern hardware. If a design does need legacy JTAG signals, then the MIPI-20 connector should be used. Pins whose functionality isn’t needed may be left unconnected. Its physical connector is virtually identical to MIPI-10, except that it’s twice as long, supporting twice as many pins. Its pinout is shown in [Table 4](#tab:mipitwenty). The function of each pin is described in[Table 3](#tab:pinout). __Table 3\. JTAG Connector Pin Functions__ | Essential | GND | Connected to ground. | | ----------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | TCK | JTAG TCK signal, driven by the debug adapter. | | | TDI | JTAG TDI signal, driven by the debug adapter. | | | TDO | JTAG TDO signal, driven by the target. | | | TMS | JTAG TMS signal, driven by the debug adapter. | | | VREF DEBUG | Reference voltage for logic high. | | | Recommended | nRESET | Open drain active low reset signal, usually driven by the debug adapter. The signal may be used bi-directional to drive or sense the target reset signal.Asserting reset should reset any RISC-V cores as well as any other peripherals on the PCB. It should not reset the debug logic. This pin is optional but strongly encouraged.nRESET should never be connected to the TAP reset, otherwise the debugger might not be able to debug through a reset to discover the cause of a crash or to maintain execution control after the reset. | | KEY | This pin may be cut on the male and plugged on the female header to ensure the header is always plugged in correctly. It is, however, recommended to use this pin as an additional ground, to allow for fastest TCK speeds. A shrouded connector should be used to prevent the cable from being plugged in incorrectly. | | | Advanced | EXT | Reserved for custom use. Could be an input or an output. | | TRIGIN | Not used by this specification, to be driven by debug adapter. (Can be used for extended functions like UART or boot mode selection by some debug adapters). | | | TRIGOUT | Not used by this specification, driven by the target. | | | Specialized | nTRST | Test reset, driven by the debug adapter. Asserting nTRST initializes the JTAG DTM asynchronously. It is used in systems where the JTAG DTM is not ready to be used after a normal power up. This signal is sometimes called TRST\*. | | Legacy | RTCK | Return test clock, driven by the target. A target may relay the TCK signal here once it has processed it, allowing a debugger to adjust its TCK frequency in response.This signal should only be used to support legacy components that rely on this functionality. | | nTRST\_PD | Test reset pull-down, driven by the debug adapter. Same function as nTRST, but with pull-down resistor on target.This signal should only be used to support legacy components that rely on this functionality. | | __Table 4\. MIPI 20-pin JTAG Connector Diagram__ | VREF DEBUG | 1 | 2 | TMS | | ---------- | -- | -- | --------------- | | GND | 3 | 4 | TCK | | GND | 5 | 6 | TDO | | GND or KEY | 7 | 8 | TDI | | GND | 9 | 10 | nRESET | | GND | 11 | 12 | GND or RTCK | | GND | 13 | 14 | NC or nTRST\_PD | | GND | 15 | 16 | nTRST or NC | | GND | 17 | 18 | TRIGIN or NC | | GND | 19 | 20 | TRIGOUT or GND | #### [](#6-1-1-4-cjtag)6.1.1.4\. cJTAG This spec does not have specific recommendations on how to use the cJTAG protocol. When implementing cJTAG access to a JTAG DTM, the MIPI 10-pin Narrow JTAG connector should be used. Pins whose functionality isn’t needed may be left unconnected. Viewing the male header from above (the pins pointing at your eye), a target’s connector looks as it does in [Table 5](#tab:mipicjtag). __Table 5\. MIPI 10-pin Narrow JTAG Connector Diagram__ | VREF DEBUG | 1 | 2 | TMSC | | ---------- | - | -- | --------------- | | GND | 3 | 4 | TCKC | | GND | 5 | 6 | EXT or NC | | GND or KEY | 7 | 8 | NC or nTRST\_PD | | GND | 9 | 10 | nRESET | Hardware Implementations ==================== ## [](#sec:implementations)Appendix A: Hardware Implementations Below are two possible implementations. A designer could choose one, mix and match, or come up with their own design. ### [](#abstract-command-based)Abstract Command Based Halting happens by stalling the hart execution pipeline. Muxes on the register file(s) allow for accessing GPRs and CSRs using the Access Register abstract command. Memory is accessed using the Abstract Access Memory command or through System Bus Access. This implementation could allow a debugger to collect information from the hart even when that hart is unable to execute instructions. ### [](#execution%5Fbased)Execution Based This implementation only implements the Access Register abstract command for GPRs on a halted hart, and relies on the Program Buffer for all other operations. It uses the hart’s existing pipeline and ability to execute from arbitrary memory locations to avoid modifications to a hart’s datapath. When the halt request bit is set, the Debug Module raises a special interrupt to the selected harts. This interrupt causes each hart to enter Debug Mode and jump to a defined memory region that is serviced by the DM and is only accessible to the harts in Debug Mode. Accesses to this memory should be uncached to avoid side effects from debugging operations. When taking this jump, `pc` is saved to [dpc](Sdext.html#csr-dpc) and [cause](Sdext.html#dcsr-cause) is updated in [dcsr](Sdext.html#csr-dcsr). This jump is similar to a trap but it is not architecturally considered a trap, so for instance doesn’t count as a trap for trigger behavior. The code in the Debug Module causes the hart to execute a "park loop." In the park loop the hart writes its `mhartid` to a memory location within the Debug Module to indicate that it is halted. To allow the DM to individually control one out of several halted harts, each hart polls for flags in a DM-controlled memory location to determine whether the debugger wants it to execute the Program Buffer or perform a resume. To execute an abstract command, the DM first populates some internal words of program buffer according to [command](debug%5Fmodule.html#dm-command). When [transfer](debug%5Fmodule.html#accessregister-transfer) is set, the DM populates these words with `lw , 0x400(zero)` or `sw , 0x400(zero)`. 64- and 128-bit accesses use `ld`/`sd` and `lq`/`sq` respectively. If [transfer](debug%5Fmodule.html#accessregister-transfer) is not set, the DM populates these instructions as `nop’s`. If [postexec](debug%5Fmodule.html#accessregister-postexec) is set, execution continues to the debugger-controlled Program Buffer, otherwise the DM causes an `ebreak` to execute immediately. When `ebreak` is executed (indicating the end of the Program Buffer code) the hart returns to its park loop. If an exception is encountered, the hart jumps to an address within the Debug Module. The code there causes the hart to write to the Debug Module indicating an exception. Then the hart jumps back to the park loop. The DM infers from the write that there was an exception, and sets [cmderr](debug%5Fmodule.html#abstractcs-cmderr) appropriately. Typically the hart will execute a `fence` instruction before entering the park loop, to ensure that any effects from the abstract command, such as a write to [data0](debug%5Fmodule.html#dm-data0), take effect before the DM returns [busy](debug%5Fmodule.html#abstractcs-busy) to 0. To resume execution, the debug module sets a flag which causes the hart to execute a `dret`. `dret` is an instruction that only has meaning while in Debug Mode and not executing from the Program Buffer. Its recommended encoding is 0x7b200073\. When `dret` is executed, `pc` is restored from [dpc](Sdext.html#csr-dpc) and normal execution resumes at the privilege set by [prv](Sdext.html#dcsr-prv) and [v](Sdext.html#dcsr-v), and the ELP state set by [pelp](Sdext.html#dcsr-pelp). [data0](debug%5Fmodule.html#dm-data0) etc. are mapped into regular memory at an address relative to with only a 12-bit `imm`. The exact address is an implementation detail that a debugger must not rely on. For example, the `data` registers might be mapped to `0x400`. For additional flexibility, [progbuf0](debug%5Fmodule.html#dm-progbuf0), etc. are mapped into regular memory immediately preceding [data0](debug%5Fmodule.html#dm-data0), in order to form a contiguous region of memory which can be used for either program execution or data transfer. The PMP must not disallow fetches, loads, or stores in the address range associated with the Debug Module when the hart is in Debug Mode, regardless of how the PMP is configured. The same is true of PMA. Without this guarantee, the park loop would enter an infinite loop of traps and debug would not be possible. ### [](#dmi%5Fsignals)Debug Module Interface Signals As stated in section [debug\_module.adoc#dmi](debug%5Fmodule.html#dmi) the details of the DMI are left to the system designer. It is quite often the case that only one DTM and one DM is implemented. In this case it might be useful to comply with the signals suggested in [Table 1](#tab:dmi%5Fsignals), which is the implementation used in the open-source[rocket-chip](https://github.com/chipsalliance/rocket-chip/blob/375045a7db1bdc7b4f7851f1a59b3f10a2b922ff/src/main/scala/devices/debug/Debug.scala#L170)RISC-V core. The DTM can start a request when the DM sets REQ\_READY to 1\. When this is the case REQ\_OP can be set to 1 for a read or 2 for a write request. The desired address is driven with the REQ\_ADDRESS signal. Finally REQ\_VALID is set high, indicating to the DM that a valid request is pending. The DM must respond to a request from the DTM when RSP\_READY is high. The status of the response is indicated by the RSP\_OP signal (see [op](dtm.html#dmi-op)). The data of the response is driven to RSP\_DATA. A pending response is signalled by setting RSP\_VALID. __Table 1\. Signals for the suggested DMI between one DTM and one DM__ | Signal | Width | Source | Description | | ------------ | ----------------------------- | ------ | --------------------------------------------------- | | REQ\_VALID | 1 | DTM | Indicates that a valid request is pending | | REQ\_READY | 1 | DM | Indicates that the DM is able to process a request | | REQ\_ADDRESS | [abits](dtm.html#dtmcs-abits) | DTM | Requested address | | REQ\_DATA | 32 | DTM | Requested data | | REQ\_OP | 2 | DTM | Same meaning as the [op](dtm.html#dmi-op) field | | RSP\_VALID | 1 | DM | Indicates that a valid respond is pending | | RSP\_READY | 1 | DTM | Indicates that the DTM is able to process a respond | | RSP\_DATA | 32 | DM | Response data | | RSP\_OP | 2 | DM | Same meaning as the [op](dtm.html#dmi-op) field | 1.1. Introduction ==================== ## [](#intro)1.1\. Introduction When a design progresses from simulation to hardware implementation, a user’s control and understanding of the system’s current state drops dramatically. To help bring up and debug low level software and hardware, it is critical to have good debugging support built into the hardware. When a robust OS is running on a core, software can handle many debugging tasks. However, in many scenarios, hardware support is essential. This document outlines a standard architecture for debug support on RISC-V hardware platforms. This architecture allows a variety of implementations and tradeoffs, which is complementary to the wide range of RISC-V implementations. At the same time, this specification defines common interfaces to allow debugging tools and components to target a variety of hardware platforms based on the RISC-V ISA. System designers may choose to add additional hardware debug support, but this specification defines a standard interface for common functionality. ### [](#1-1-1-terminology)1.1.1\. Terminology **advanced feature** An advanced feature for advanced users. Most users will not be able to take advantage of it. **AMO** Atomic Memory Operation. **BYPASS** JTAG instruction that selects a single bit data register, also called BYPASS. **component** A RISC-V core, or other part of a hardware platform. Typically all components will be connected to a single system bus. **CSR** Control and Status Register. **DM** Debug Module (see [debug\_module.adoc#dm](debug%5Fmodule.html#dm)). **DMI** Debug Module Interface (see [debug\_module.adoc#dmi](debug%5Fmodule.html#dmi)). **DR** JTAG Data Register. **DTM** Debug Transport Module (see [dtm.adoc#dtm](dtm.html#dtm)). **DXLEN** Debug XLEN, which is the widest XLEN a hart supports, ignoring the current value of `mxl` in `misa`. **ELP** Expected landing pad state, define by the Zicfilp extension. **essential feature** An essential feature must be present in order for debug to work correctly. **GPR** General Purpose Register. **hardware platform** A single system consisting of one or more _components_. **hart** A hardware thread in a RISC-V core. **IDCODE** 32-bit Identification CODE, and a JTAG instruction that returns the IDCODE value. **IR** JTAG Instruction Register. **JTAG** Refers to work done by IEEE’s Joint Test Action Group, described in IEEE 1149.1. **legacy feature** A legacy feature should only be implemented to support legacy hardware that is present in a system. **Minimal RISC-V Debug Specification** A subset of the full Debug Specification that allows for very small implementations. See [debug\_module.adoc#dm](debug%5Fmodule.html#dm). **NAPOT** Naturally Aligned Power-Of-Two. **NMI** Non-Maskable Interrupt. **physical address** address that is directly usable on the system bus. **recommended feature** A recommended feature is not required for debug to work correctly, but it is so useful that it should not be omitted without good reason. **SBA** System Bus Access (see [debug\_module.adoc#systembusaccess](debug%5Fmodule.html#systembusaccess)). **specialized feature** A specialized feature, that only makes sense in the context of some specific hardware. **TAP** Test Access Port, defined in IEEE 1149.1. **TM** Trigger Module (see [Sdtrig.adoc#trigger](Sdtrig.html#trigger)). **virtual address** An address as a hart sees it. If the hart is using address translation this may be different from the physical address. If there is no translation then it will be the same. **xepc** The exception program counter CSR (e.g. `mepc`) that is appropriate for the mode being trapped to. ### [](#1-1-2-context)1.1.2\. Context This specification attempts to support all RISC-V ISA extensions that have, roughly, been ratified through the first half of 2023\. In particular, though, this specification specifically addresses features in the following extensions: 1. A 2. C 3. D 4. F 5. H 6. Sm1p13 7. Smstateen 8. Ss1p13 9. V 10. Zawrs 11. Zcmp 12. Zicbom 13. Zicbop 14. Zicboz 15. Zicsr #### [](#1-1-2-1-versions)1.1.2.1\. Versions Version 0.13 of this document was ratified by the RISC-V Foundation’s board. Versions 0.13. are bug fix releases to that ratified specification. Version 0.14 was a working version that was never officially ratified. Version 1.0 is almost entirely forwards and backwards compatible with Version 0.13. ##### [](#1-1-2-1-1-bugfixes-from-0-13-to-1-0)1.1.2.1.1\. Bugfixes from 0.13 to 1.0 Changes that fix a bug in the spec: 1. Fix order of operations described in [sbdata0](debug%5Fmodule.html#dm-sbdata0).[#392](https://github.com/riscv/riscv-debug-spec/pull/392) 2. Resume ack is set after resume, in [debug\_module.adoc#runcontrol](debug%5Fmodule.html#runcontrol).[#400](https://github.com/riscv/riscv-debug-spec/pull/400) 3. [sselect](Sdtrig.html#textra32-sselect) applies to [svalue](Sdtrig.html#textra32-svalue) . [#402](https://github.com/riscv/riscv-debug-spec/pull/402) 4. [mte](Sdtrig.html#tcontrol-mte) only applies when action=0.[#411](https://github.com/riscv/riscv-debug-spec/pull/411) 5. [aamsize](debug%5Fmodule.html#accessmemory-aamsize) does not affect Argument Width.[#420](https://github.com/riscv/riscv-debug-spec/pull/420) 6. Clarify that harts halt out of reset if [haltreq](debug%5Fmodule.html#dmcontrol-haltreq) \=1.[#419](https://github.com/riscv/riscv-debug-spec/pull/419) ##### [](#1-1-2-1-2-incompatible-changes-from-0-13-to-1-0)1.1.2.1.2\. Incompatible Changes from 0.13 to 1.0 Changes that are not backwards-compatible. Debuggers or hardware implementations that implement 0.13 will have to change something in order to implement 1.0: 1. Make haltsum0 optional if there is only one hart.[#505](https://github.com/riscv/riscv-debug-spec/pull/505) 2. System bus autoincrement only happens if an access actually takes place. ([sbdata0](debug%5Fmodule.html#dm-sbdata0)) [#507](https://github.com/riscv/riscv-debug-spec/pull/507) 3. Bump [version](Sdtrig.html#tinfo-version) to 3\. [#512](https://github.com/riscv/riscv-debug-spec/pull/512) 4. Require debugger to poll [dmactive](debug%5Fmodule.html#dmcontrol-dmactive) after lowering it.[#566](https://github.com/riscv/riscv-debug-spec/pull/566) 5. Add [pending](Sdtrig.html#icount-pending) to [icount](Sdtrig.html#csr-icount) . [#574](https://github.com/riscv/riscv-debug-spec/pull/574) 6. When a selected trigger is disabled, [tdata2](Sdtrig.html#csr-tdata2) and [tdata3](Sdtrig.html#csr-tdata3) can be written with any value supported by any of the types this trigger supports.[#721](https://github.com/riscv/riscv-debug-spec/pull/721) 7. [tcontrol](Sdtrig.html#csr-tcontrol) fields only apply to breakpoint traps, not any trap.[#723](https://github.com/riscv/riscv-debug-spec/pull/723) 8. If [version](Sdtrig.html#tinfo-version) is greater than 0, then [hit0](Sdtrig.html#mcontrol6-hit0) (previously called [mcontrol6](Sdtrig.html#csr-mcontrol6).`hit`) now contains 0 when a trigger fires more than one instruction after the instruction that matched. (This information is now reflected in [hit1](Sdtrig.html#mcontrol6-hit1).)[#795](https://github.com/riscv/riscv-debug-spec/pull/795) 9. If [version](Sdtrig.html#tinfo-version) is greater than 0, then bit 20 of [mcontrol6](Sdtrig.html#csr-mcontrol6) is no longer used for timing information. (Previously the bit was called [mcontrol6](Sdtrig.html#csr-mcontrol6).`timing`.)[#807](https://github.com/riscv/riscv-debug-spec/pull/807) 10. If [version](Sdtrig.html#tinfo-version) is greater than 0, then the encodings of [size](Sdtrig.html#mcontrol6-size) for sizes greater than 64 bit have changed.[#807](https://github.com/riscv/riscv-debug-spec/pull/807) ##### [](#1-1-2-1-3-minor-changes-from-0-13-to-1-0)1.1.2.1.3\. Minor Changes from 0.13 to 1.0 Changes that slightly modify defined behavior. Technically backwards incompatible, but unlikely to be noticeable: 1. [stopcount](Sdext.html#dcsr-stopcount) only applies to hart-local counters.[#405](https://github.com/riscv/riscv-debug-spec/pull/405) 2. [version](debug%5Fmodule.html#dmstatus-version) may be invalid when [dmactive](debug%5Fmodule.html#dmcontrol-dmactive)\=0.[#414](https://github.com/riscv/riscv-debug-spec/pull/414) 3. Address triggers ([mcontrol](Sdtrig.html#csr-mcontrol)) may fire on any accessed address.[#421](https://github.com/riscv/riscv-debug-spec/pull/421) 4. All Trigger Module registers ([\[tab:trigger\]](#tab:trigger)) are optional. [#431](https://github.com/riscv/riscv-debug-spec/pull/431) 5. When extending IR, [bypass](dtm.html#dtm-bypass) still is all ones.[#437](https://github.com/riscv/riscv-debug-spec/pull/437) 6. [ebreaks](Sdext.html#dcsr-ebreaks) and [ebreaku](Sdext.html#dcsr-ebreaku) are WARL. [#458](https://github.com/riscv/riscv-debug-spec/pull/458) 7. NMIs are disabled by [stepie](Sdext.html#dcsr-stepie).[#465](https://github.com/riscv/riscv-debug-spec/pull/465) 8. R/W1C fields should be cleared by writing every bit high.[#472](https://github.com/riscv/riscv-debug-spec/pull/472) 9. Specify trigger priorities in [Sdtrig.adoc#tab:priority](Sdtrig.html#tab:priority) relative to exceptions.[#478](https://github.com/riscv/riscv-debug-spec/pull/478) 10. Time may pass before [dmactive](debug%5Fmodule.html#dmcontrol-dmactive) becomes high.[#500](https://github.com/riscv/riscv-debug-spec/pull/500) 11. Clear MPRV when resuming into lower privilege mode.[#503](https://github.com/riscv/riscv-debug-spec/pull/503) 12. Halt state may not be preserved across reset.[#504](https://github.com/riscv/riscv-debug-spec/pull/504) 13. Hardware should clear trigger action when [dmode](Sdtrig.html#tdata1-dmode) is cleared and action is 1.[#501](https://github.com/riscv/riscv-debug-spec/pull/501) 14. Change quick access exceptions to halt the target in [\[ac-quickaccess\]](#ac-quickaccess).[#585](https://github.com/riscv/riscv-debug-spec/pull/585) 15. Writing 0 to [tdata1](Sdtrig.html#csr-tdata1) forces a state where [tdata2](Sdtrig.html#csr-tdata2) and [tdata3](Sdtrig.html#csr-tdata3) are writable.[#598](https://github.com/riscv/riscv-debug-spec/pull/598) 16. Solutions to deal with reentrancy in [Sdtrig.adoc#nativetrigger](Sdtrig.html#nativetrigger) prevent triggers from_matching_, not merely _firing_. This primarily affects [icount](Sdtrig.html#csr-icount) behavior.[#722](https://github.com/riscv/riscv-debug-spec/pull/722) 17. Attempts to access an unimplemented CSR raise an illegal instruction exception. [#791](https://github.com/riscv/riscv-debug-spec/pull/791) ##### [](#1-1-2-1-4-new-features-from-0-13-to-1-0)1.1.2.1.4\. New Features from 0.13 to 1.0 New backwards-compatible feature that did not exist before: 1. Add halt groups and external triggers in [debug\_module.adoc#hrgroups](debug%5Fmodule.html#hrgroups).[#404](https://github.com/riscv/riscv-debug-spec/pull/404) 2. Reserve some DMI space for non-standard use. See [custom](debug%5Fmodule.html#dm-custom), and [custom0](debug%5Fmodule.html#dm-custom0) through `custom15`.[#406](https://github.com/riscv/riscv-debug-spec/pull/406) 3. Reserve trigger [type](Sdtrig.html#tdata1-type) values for non-standard use.[#417](https://github.com/riscv/riscv-debug-spec/pull/417) 4. Add [nmi](Sdtrig.html#itrigger-nmi) bit to [itrigger](Sdtrig.html#csr-itrigger). [#408](https://github.com/riscv/riscv-debug-spec/pull/408)and [#709](https://github.com/riscv/riscv-debug-spec/pull/709) 5. Recommend matching on every accessed address.[#449](https://github.com/riscv/riscv-debug-spec/pull/449) 6. Add resume groups in [debug\_module.adoc#hrgroups](debug%5Fmodule.html#hrgroups).[#506](https://github.com/riscv/riscv-debug-spec/pull/506) 7. Add [relaxedpriv](debug%5Fmodule.html#abstractcs-relaxedpriv) . [#536](https://github.com/riscv/riscv-debug-spec/pull/536) 8. Move [scontext](Sdtrig.html#csr-scontext), renaming original to [mscontext](Sdtrig.html#csr-mscontext), and create [hcontext](Sdtrig.html#csr-hcontext).[#535](https://github.com/riscv/riscv-debug-spec/pull/535) 9. Add [mcontrol6](Sdtrig.html#csr-mcontrol6), deprecating [mcontrol](Sdtrig.html#csr-mcontrol).[#538](https://github.com/riscv/riscv-debug-spec/pull/538) 10. Add hypervisor support: [ebreakvs](Sdext.html#dcsr-ebreakvs), [ebreakvu](Sdext.html#dcsr-ebreakvu), [v](Sdext.html#dcsr-v), [hcontext](Sdtrig.html#csr-hcontext), [mcontrol](Sdtrig.html#csr-mcontrol), [mcontrol6](Sdtrig.html#csr-mcontrol6), and [priv](Sdext.html#virt-priv).[#549](https://github.com/riscv/riscv-debug-spec/pull/549) 11. Optionally make [anyunavail](debug%5Fmodule.html#dmstatus-anyunavail) and [allunavail](debug%5Fmodule.html#dmstatus-allunavail) sticky, controlled by [stickyunavail](debug%5Fmodule.html#dmstatus-stickyunavail).[#520](https://github.com/riscv/riscv-debug-spec/pull/520) 12. Add [tmexttrigger](Sdtrig.html#csr-tmexttrigger) to support trigger module external trigger inputs.[#543](https://github.com/riscv/riscv-debug-spec/pull/543) 13. Describe [mcontrol](Sdtrig.html#csr-mcontrol) and [mcontrol6](Sdtrig.html#csr-mcontrol6) behavior with atomic instructions.[#561](https://github.com/riscv/riscv-debug-spec/pull/561) 14. Trigger hit bits must be set on fire, may be set on match.[#593](https://github.com/riscv/riscv-debug-spec/pull/593) 15. Add [sbytemask](Sdtrig.html#textra32-sbytemask) and [sbytemask](Sdtrig.html#textra32-sbytemask) to [textra32](Sdtrig.html#csr-textra32) and [textra64](Sdtrig.html#csr-textra64).[#588](https://github.com/riscv/riscv-debug-spec/pull/588) 16. Allow debugger to request harts stay alive with keepalive bit in[setkeepalive](debug%5Fmodule.html#dmcontrol-setkeepalive).[#592](https://github.com/riscv/riscv-debug-spec/pull/592) 17. Add [ndmresetpending](debug%5Fmodule.html#dmstatus-ndmresetpending) to allow a debugger to determine when ndmreset is complete.[#594](https://github.com/riscv/riscv-debug-spec/pull/594) 18. Add [intctl](Sdtrig.html#tmexttrigger-intctl) to support triggers from an interrupt controller.[#599](https://github.com/riscv/riscv-debug-spec/pull/599) ##### [](#1-1-2-1-5-incompatible-changes-during-1-0-stable)1.1.2.1.5\. Incompatible Changes During 1.0 Stable Backwards-incompatible changes between two versions that are both called 1.0 stable. 1. [nmi](Sdtrig.html#itrigger-nmi) was moved from [etrigger](Sdtrig.html#csr-etrigger) to [itrigger](Sdtrig.html#csr-itrigger), and is now subject to the mode bits in that trigger. 2. [#728](https://github.com/riscv/riscv-debug-spec/pull/728) introduced Message Registers, which were later removed in[#878](https://github.com/riscv/riscv-debug-spec/pull/878). 3. It may not be possible to read the contents of the Program Buffer using the `progbuf` registers.[#731](https://github.com/riscv/riscv-debug-spec/pull/731) 4. [tcontrol](Sdtrig.html#csr-tcontrol) fields apply to all traps, not just breakpoint traps. This reverts[#723](https://github.com/riscv/riscv-debug-spec/pull/723).[#880](https://github.com/riscv/riscv-debug-spec/pull/880) ##### [](#1-1-2-1-6-incompatible-changes-between-1-0-0-rc1-and-1-0-0-rc2)1.1.2.1.6\. Incompatible Changes Between 1.0.0-rc1 and 1.0.0-rc2 Backwards-incompatible changes between 1.0.0-rc1 and 1.0.0-rc2. 1. [#981](https://github.com/riscv/riscv-debug-spec/pull/981) made[scontext](Sdtrig.html#csr-scontext).[data](Sdtrig.html#scontext-data), [mcontext](Sdtrig.html#csr-mcontext).[hcontext](Sdtrig.html#mcontext-hcontext),[sbytemask](Sdtrig.html#textra64-sbytemask), and [textra64](Sdtrig.html#csr-textra64).`svalue` narrower. This avoids confusion about the contents of [scontext](Sdtrig.html#csr-scontext) and [mcontext](Sdtrig.html#csr-mcontext) when XLEN is reduced and increased again. ### [](#1-1-3-about-this-document)1.1.3\. About This Document #### [](#1-1-3-1-structure)1.1.3.1\. Structure This document contains two parts. The main part of the document is the specification, which is given in the numbered chapters. The second part of the document is a set of appendices. The information in the appendices is intended to clarify and provide examples, but is not part of the actual specification. #### [](#1-1-3-2-isa-vs-non-isa)1.1.3.2\. ISA vs. non-ISA This specification contains both ISA and non-ISA parts. The ISA parts define self-contained ISA extensions. The other parts of the document describe the non-ISA external debug extension. Chapters whose contents are solely one or the other are labeled as such in their title. Chapters without such a label apply to both ISA and non-ISA. #### [](#1-1-3-3-register-definition-format)1.1.3.3\. Register Definition Format All register definitions in this document follow the format shown below. A simple graphic shows which fields are in the register. The upper and lower bit indices are shown to the top left and top right of each field. The total number of bits in the field are shown below it. After the graphic follows a table which for each field lists its name, description, allowed accesses, and reset value. The allowed accesses are listed in [Table 1](#tab:access). The reset value is either a constant or "Preset." The latter means it is an implementation-specific legal value. Parts of the register which are currently unused are labeled with the number 0\. Software must only write 0 to those fields, and ignore their value while reading. Hardware must return 0 when those fields are read, and ignore the value written to them. | | This behavior enables us to use those fields later without having to increase the values in the version fields. | | ------------------------------------------------------------------------------------------------------------------ | Names of registers and their fields are hyperlinks to their definition, and are also listed in the [\[index\]](#index). ##### [](#shortname)Long Name (shortname, at 0x123) ![Diagram](_images/diag-308495061acf103c26ff9e0b21b6ae8d4bcabb4f.svg) | Field | Description | Access | Reset | | ----- | ------------------------------------------- | ------- | ----- | | field | Description of what this field is used for. | **R/W** | 15 | __Table 1\. Register Access Abbreviations__ | R | Read-only. | | ----- | ------------------------------------------------------------------------------------------------------------------------------------------------- | | R/W | Read/Write. | | R/W1C | Read/Write Ones to Clear. Writing 0 to every bit has no effect. Writing 1 to every bit clears the field. The result of other writes is undefined. | | WARZ | Write any, read zero. A debugger may write any value. When read this field returns 0. | | W1 | Write-only. Only writing 1 has an effect. When read the returned value should be 0. | | WARL | Write any, read legal. A debugger may write any value. If a value is unsupported, the implementation converts the value to one that is supported. | ### [](#1-1-4-background)1.1.4\. Background There are several use cases for dedicated debugging hardware, both in native debug and external debug. Native debug (sometimes called self-hosted debug) refers to debug software running on a RISC-V platform which debugs the same platform. The optional Trigger Module provides features that are useful for native debug. External debug refers to debug software running somewhere else, debugging the RISC-V platform via a debug transport like JTAG. The entire document provides features that are useful for external debug. This specification addresses the use cases listed below. Implementations can choose not to implement every feature, which means some use cases might not be supported. * Accessing hardware on a hardware platform without a working CPU. (External debug.) * Bootstrapping a hardware platform to test, configure, and program components before there is any executable code path in the hardware platform. (External debug.) * Debugging low-level software in the absence of an OS or other software. (External debug.) * Debugging issues in the OS itself. (External or native debug.) * Debugging processes running on an OS. (Native or external debug.) ### [](#1-1-5-supported-features)1.1.5\. Supported Features The debug interface described in this specification supports the following features: 1. All hart registers (including CSRs) can be read/written. 2. Memory can be accessed either from the hart’s point of view, through the system bus directly, or both. 3. RV32, RV64, and future RV128 are all supported. 4. Any hart in the hardware platform can be independently debugged. 5. A debugger can discover almost \[[1](#%5Ffootnotedef%5F1 "View footnote.")\] everything it needs to know itself, without user configuration. 6. Each hart can be debugged from the very first instruction executed. 7. A RISC-V hart can be halted when a software breakpoint instruction is executed. 8. Hardware single-step can execute one instruction at a time. 9. Debug functionality is independent of the debug transport used. 10. The debugger does not need to know anything about the microarchitecture of the harts it is debugging. 11. Arbitrary subsets of harts can be halted and resumed simultaneously. (Optional) 12. Arbitrary instructions can be executed on a halted hart. That means no new debug functionality is needed when a core has additional or custom instructions or state, as long as there exist programs that can move that state into GPRs. (Optional) 13. Registers can be accessed without halting. (Optional) 14. A running hart can be directed to execute a short sequence of instructions, with little overhead. (Optional) 15. A system bus manager allows memory access without involving any hart. (Optional) 16. A RISC-V hart can be halted when a trigger matches the PC, read/write address/data, or an instruction opcode. (Optional) 17. Harts can be grouped, and harts in the same group will all halt when any of them halts. These groups can also react to or notify external triggers. (Optional) This document does not suggest a strategy or implementation for hardware test, debugging or error detection techniques. Scan, built-in self test (BIST), etc. are out of scope of this specification, but this specification does not intend to limit their use in RISC-V systems. It is possible to debug code that uses software threads, but there is no special debug support for it. --- [1](#%5Ffootnoteref%5F1). Notable exceptions include information about the memory map and peripherals. 2.1. System Overview ==================== ## [](#overview)2.1\. System Overview [Figure 1](#systemoverview) shows the main components of Debug Support. Blocks shown in dotted lines are optional. The user interacts with the Debug Host (e.g. laptop), which is running a debugger (e.g. gdb). The debugger communicates with a Debug Translator (e.g. OpenOCD, which may include a hardware driver) to communicate with Debug Transport Hardware (e.g. Olimex USB-JTAG adapter). The Debug Transport Hardware connects the Debug Host to the hardware platform’s Debug Transport Module (DTM). The DTM provides access to one or more Debug Modules (DMs) using the Debug Module Interface (DMI). Each hart in the hardware platform is controlled by exactly one DM. Harts may be heterogeneous. There is no further limit on the hart-DM mapping, but usually all harts in a single core are controlled by the same DM. In most hardware platforms there will only be one DM that controls all the harts in the hardware platform. DMs provide run control of their harts in the hardware platform. Abstract commands provide access to GPRs. Additional registers are accessible through abstract commands or by writing programs to the optional Program Buffer. The Program Buffer allows the debugger to execute arbitrary instructions on a hart. This mechanism can also be used to access memory. An optional system bus access block allows memory accesses without using a RISC-V hart to perform the access. Each RISC-V hart may implement a Trigger Module. When trigger conditions are met, harts will halt and inform the debug module that they have halted. ![overview](_images/overview.png) Figure 1\. RISC-V Debug System Overview Change Log ==================== ## [](#change-log)Change Log PDF generated on: 2024-07-05 15:21:36 UTC Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at. Copyright 2024 by RISC-V International. 1.1. Contributors ==================== ## [](#1-1-contributors)1.1\. Contributors This RISC-V specification has been contributed to directly or indirectly by: * Iain Robertson (Siemens) <[iain.robertson@siemens.com](mailto:iain.robertson@siemens.com)\> - author * Paul Donahue (Ventana) - reviews * Michael Schleinkofer (Lauterbach) - reviews * Beeman Strong (Rivos) - reviews * Robert Chyla (MIPS) - reviews * Ved Shanbhogue (Rivos) - reviews Packet Encapsulation for E-Trace ==================== ## [](#chapter3)Packet Encapsulation for E-Trace [\[chapter2\]](#chapter2) describes the packet encapsulation in general terms. This chapter describes how that applies in the context of E-Trace. The [RISC-V Trace Control Interface Specification](https://github.com/riscv-non-isa/tg-nexus-trace/blob/master/pdfs/RISC-V-Trace-Control-Interface.pdf) includes several fields relevant for the construction and synchronization of packets, as described in the following sections. ### [](#srcid)**srcID** * **trTeInhibitSrc** in the _trTeControl_ register of a trace encoder indicates whether a **srcID** is present or absent in the encapsulated packets; * **trTeSrcBits** in the _trTeInstFeatures_ register indicates the length of the **srcID**; * **trTeSrcID** in the _trTeInstFeatures_ register indicates the value of the **srcID** associated with that particular source. ### [](#timestamp)**timestamp** * When **trTsEnable** in the _trTsControl_ register is 1, a timestamp may be optionally included, via the **timestamp** field in the encapsulated packets. When present, **header.extend** will be 1 as described in [\[\_section\_Timestamp\]](#%5Fsection%5FTimestamp); * **trTsWidth** in the _trTsControl_ register indicates the number of bits in **timestamp** fields. Note that some E-Trace packets may optionally include a **time** field, as an alternative method of providing time information. Presence or absence of this is fixed for a given system based on a discoverable parameter, but is not run-time configurable. ### [](#type)**type** The capabilities of the trace source will determine the minimum width of the **type** field in the **payload** field-group. For E-Trace: * If only instruction trace is supported, the minimum width is 0 (i.e. the field is omitted) * If both instruction and data trace are supported, the minimum width is 1, encoded as * 0: Instruction trace packet * 1: Data trace packet * A width of 2 or more is permited for applications where packet types other than instruction and data trace are required. Encoding for this case is application specific and not mandated by this standard. ### [](#synchronization)Synchronization * **trPibAsyncFreq**, **trRamAsyncFreq** and **trAtbBridgeAsyncFreq** in the _trPibControl_, _trRamControl_ and _trAtbBridgeControl_ registers respectively are used to determine the interval between insertion of synchronization sequences. 3.1. Packet Encapsulation ==================== ## [](#chapter2)3.1\. Packet Encapsulation Two types of encapsulation are defined: _normal_ and _null_. A transmitted stream of encapsulated packets comprises a mixture of _normal_ and _null_ packets. Each packet is atomic, and must be transmitted in its entirety before another packet can be sent. ### [](#3-1-1-normal-encapsulation-structure)3.1.1\. Normal Encapsulation Structure The normal encapsulation structure is comprised of four field-groups as shown in [Table 1](#%5Ftable%5FGroups). In this and the following tables, the field-groups and fields are listed in transmission order: the uppermost field-group or field in a table is transmitted first, and multi-bit fields are transmitted least significant bit first. __Table 1\. Encapsulation Field Groups__ | Group Name | \# Bits | Description | | ------------- | ------- | ------------------------------------------------------------------- | | **header** | 8 | Encapsulation header. See [3.1.1.1\. Header](#%5Fsection%5FHeader). | | **srcID** | 0 - 16 | Source ID. See [3.1.1.2\. srcID](#%5Fsection%5FsrcID). | | **timestamp** | T\*8 | Time stamp. See [3.1.1.3\. Timestamp](#%5Fsection%5FTimestamp). | | **payload** | 1-248 | Packet payload. See [3.1.1.4\. Payload](#%5Fsection%5FPayload). | The groups are defined in the following sections: #### [](#%5Fsection%5FHeader)3.1.1.1\. Header The header is a single byte comprising the fields defined in [Table 2](#%5Ftable%5FHeader). __Table 2\. Header Fields__ | Field Name | \# Bits | Description | | ---------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **length** | 5 | Encapsulated payload length. A value of _L_ indicates an _L_ byte payload. Must be > 0 - see [3.1.2\. Null Encapsulation](#%5Fsection%5Fnull%5Fencapsulation). | | **flow** | 2 | Flow indicator. This can be used to direct packets to a particular sink in systems where multiple sinks exist, and those sinks include the ability to accept or discard packets based on the flow value. | | **extend** | 1 | Indicates presence of timestamp when 1\. Must be 0 if timestamp width is 0. | #### [](#%5Fsection%5FsrcID)3.1.1.2\. srcID The **srcID** field identifies the source of the packet. It can be between 0 and 16 bits in length. This length must be fixed and discoverable for a given system. The **srcID** may be omitted (i.e. zero bits in length) if there is only one source in the system, or if the transport scheme includes a sideband bus for the source ID (for example, ATB). When present, an 8-bit **srcID** will be sufficient for most use cases, and is simplest in terms of determining the packet length, keeping all field groups aligned to byte boundaries. However, the length of the field can be reduced to improve efficiency for small systems, or increased if required for larger systems. This is explained in more detail in [3.1.1.5\. Packet Length](#%5Fsection%5Fpacket%5Flength). #### [](#%5Fsection%5FTimestamp)3.1.1.3\. Timestamp The **timestamp** field provides a means to include time information with every packet. It is included in the encapsulation if **header.extend** is 1\. When included, the timestamp must be T bytes in length. The length must be discoverable, and fixed for a given system. Timestamps may be omitted either because time is not of interest to the user, or if time information is already included within the encapsulated payload. #### [](#%5Fsection%5FPayload)3.1.1.4\. Payload The encapsulation payload can be up to 248 bits (31 bytes) in length, and comprises the fields shown in table [Table 3](#%5Ftable%5Fpayload). __Table 3\. Payload Fields__ | Field Name | \# Bits | Description | | ------------------ | ------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **type** | ≥ 0 | Packet type. May be eliminated (i.e. width set to 0) for sources with only one packet type. Length must be fixed for a given **srcID**, and discoverable if > 0. | | **trace\_payload** | ≤ R | Packet payloads such as those defined for E-Trace.Maximum value of R is defined as 248 - Y - **srcID**%8, where Y is the length of the **type** field. See [3.1.1.5\. Packet Length](#%5Fsection%5Fpacket%5Flength) for details of the relationship between **srcID** and payload length. | #### [](#%5Fsection%5Fpacket%5Flength)3.1.1.5\. Packet Length Encapsulated packets are a number of whole bytes in length, the exact number depending on the sizes of the **srcID**, **timestamp** (if present) and **header.length**: Packet length = 1 + S + (T \* **header.extend**) + **header.length** S and T are discoverable constants; S is the number of whole bytes of **srcID**: int(#bits(**srcID**)/8). For the case where the size of **srcID** is a multiple of 8 bits, **header.length** is simply the number of bits of payload rounded up to the nearest multiple of 8 and expressed in bytes: **header.length** \= ceiling(#bits(**payload**)/8) However, if the **srcID** is not a multiple of 8 bits the remaining **srcID** bits not accounted for by 'S' are instead included when determining the value of **header.length**. Thus the more general definition for any **srcID** size is: **header.length** \= ceiling((#bits(**payload**) + #bits(**srcID**)%8)/8) In this way, the maximum payload length is reduced by up to 7 when the **srcID** is not a multiple of 8 bits. In cases where the number of bits of **payload** \+ **srcID** is not a multiple of 8, some padding bits are required. These must be placed in the most significant bits of the final byte of the packet. Their value is "don’t care", but they must not "leak" information (for example, the previous contents of an intermediate buffer that may relate to a different trace session which the current recipient of trace is not authorised to receive). ### [](#%5Fsection%5Fnull%5Fencapsulation)3.1.2\. Null Encapsulation For _normal_ encapsulation, the **header.length** field is at least 1, and the overall length of the encapsulated packet will be at least 2 bytes (**header** plus 1 byte of payload). A _null_ packet is specified as consisting exclusively of one **header** byte with its **length** field set to zero, explicitly indicating the packet’s total size as one byte. The **extend** field is used to distinguish 2 different types of null packet, which are defined as follows: * **extend** \= 0: _null.idle_ * **extend** \= 1: _null.alignment_ Usually, _null.idle_ will be used. _null.alignment_ is used for synchronization, as described in the following section. Insertion of _null_ packets typically occurs at a trace sink where there are no sideband signals accompanying the data stream to identify valid data. Packets emitted from a trace source are generally transported over some form of on-chip transport (e.g. ATB) that includes sideband signalling to indicate when data is valid. In this situation when there is no data to send, valid is simply deasserted. That said, the **flow** field definition in _null_ packets is unchanged, so _null_ packets can be routed from a trace source to a specific sink if required. When generated at a sink, the **flow** field value is unimportant and is typically 0\. If the sink is generating a bit stream (i.e. the byte boundaries are not inherently known to the recipient) then the **flow** field must be zero in all _null_ packets within the generated bitstream. ### [](#3-1-3-synchronization)3.1.3\. Synchronization In a data stream comprised of packets, it’s a requirement to be able to determine where packets start and end, when starting from an arbitrary point, without knowledge of the full packet history. This can be achieved by inserting a synchronization sequence into the packet stream. This sequence is comprised of a sufficiently long sequence of _null_ packets. A 'null' byte is defined as a byte with the 5LSBs all zero, which may be a _null_ packet, or may be part of a _normal_ packet. The longest run N of ‘null’ bytes possible within a _normal_ packet is: N = 31 + T + S (see [3.1.1.3\. Timestamp](#%5Fsection%5FTimestamp) and [3.1.1.5\. Packet Length](#%5Fsection%5Fpacket%5Flength) for definitions of T and S respectively) Therefore, in a sequence of N or more ‘null’ bytes, the first N 'null' bytes may actually be part of a packet. However, any 'null' bytes after this must be _null_ packets, and the 1st non-null byte seen after this must therefore be the 1st byte of a _normal_ packet. For unframed data streams such as PIB, a _null.alignment_ packet must be transmitted as the final _null_ before a _normal_ packet. Strictly speaking this is necessary only if the data stream is sent via an interface less than 8 bits wide, but for simplicity this is mandatory for any width. The single 1 at the end of this sequence uniquely identifies the byte boundary, and what follows as the start of a packet. For example, for two _normal_ packets with M _nulls_ between them, this would comprise M-1 _null.idles_ and 1 _null.alignment_ (M > 0). For framed data streams which incorporate synchronization information in their own framing such as MIPI TWP (aka ARM Trace Formatter Protocol) or USB there is no requirement to include _null.alignment_ packets. The synchronization requirements are summarized in the following rules: * A synchronization sequence must have a length of N+1 bytes (N defined above), comprising: * For unframed data streams, N consecutive _null.idle_ packets, directly followed by one _null.alignment_ packet; * For framed data streams, N consecutive _null.idle_ packets, directly followed by one _null.idle_ or _null.alignment_ packet. Synchronization sequences are typically inserted periodically. In addition, a sufficiently long run of _null_ packets (due to a lack of _normal_ packets to send) may also serve as an 'opportunistic' synchronization sequence. For unframed data streams, this requires _null.alignment_ packets to be included, either as every (N+1)th _null_, or as the final _null_. For writing unframed data to memory, alternative synchronisation mechanisms may also be employed. For example, by dividing memory into blocks of known size, and requiring that packets do not straddle block boundaries. The first byte of every block will therefore be the start of a packet. Details of such schemes are out of scope of this specification. Unformatted Trace & Diagnostic Data Packet Encapsulation for RISC-V ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#unformatted-trace-diagnostic-data-packet-encapsulation-for-risc-v)Unformatted Trace & Diagnostic Data Packet Encapsulation for RISC-V Authors: Iain Robertson Version v1.0, 2024-07-05: Ratified | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 2.1. Introduction ==================== ## [](#intro)2.1\. Introduction The [Efficient Trace for RISC-V](https://github.com/riscv-non-isa/riscv-trace-spec/releases/download/v2.0rc2/riscv-trace-spec.pdf) (E-trace) standard defines packet payloads for instruction and data trace but does not fully define how this should be encapsulated into fully formed packets for transport, nor how instruction and data trace should be differentiated. Chapter 7 gives some illustrative examples but this is insufficiently detailed and informative only. Although the primary motivation for developing this standard was to define an encapsulation format for E-Trace packets that allows tools to parse and decode them in a standard manner, the encapsulation format defined in this document is agnostic to the packet payload structure and meaning and so can be used for any kind of unformatted data. In addition to E-trace, it could also be used for a wide variety of other uses, for example: performance counter metrics, trace or other diagnostic data from a bus fabric monitor or on-chip logic analyser. It is not suitable for data that has already been formatted into packets, such as N-Trace, which inserts a 2-bit MSEO formatting code into each byte. This specification defines an encapsulation format suitable for use with a variety of transport mechanisms, including but not limited to AMBA Advanced Trace Bus (ATB) and Siemens' Messaging Infrastructure. Examples of how trace packets can be routed for transport is given in the 'Trace Components' subsection of the [RISC-V Trace Control Interface Specification](https://github.com/riscv-non-isa/tg-nexus-trace/blob/master/docs/RISC-V-Trace-Control-Interface.adoc). ### [](#2-1-1-glossary)2.1.1\. Glossary * **ATB** \- Advanced Trace Bus, a protocol described in ARM document IHI0032B; * **E-Trace** \- Abbreviation for [Efficient Trace for RISC-V](https://github.com/riscv-non-isa/riscv-trace-spec/releases/download/v2.0rc2/riscv-trace-spec.pdf); * **PIB** \- Pin Interface Block, a parallel or serial off-chip trace port feeding into a trace probe, as defined in the [RISC-V Trace Control Interface Specification](https://github.com/riscv-non-isa/tg-nexus-trace/blob/master/docs/RISC-V-Trace-Control-Interface.adoc); * **N-Trace** \- Abbreviation for [RISC-V N-Trace (Nexus-based trace) Specification](https://github.com/riscv-non-isa/tg-nexus-trace/blob/master/docs/RISC-V-N-Trace.adoc) * **Trace Encoder** \- Hardware module that accepts execution information from a hart and generates a stream of trace packets; * **TFP** \- Trace Formatter protocol, a trace framing protocol described in ARM document IHI0029E. Also adopted by MIPI as Trace Wrapper Protocol (TWP); * **TWP** \- See **TFP**. RISC-V N-Trace (Nexus-based Trace) Specification ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-n-trace-nexus-based-trace-specification)RISC-V N-Trace (Nexus-based Trace) Specification RISC-V N-Trace Task Group Version 1.0, Nov 21, 2024: Ratified state | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 13.1. Additional Material ==================== ## [](#13-1-additional-material)13.1\. Additional Material **Trace Bandwidth Considerations** SRC field (if enabled) may change the otherwise optimal layout of [Fields in Messages](#Fields in Messages). **Validation Considerations** ResourceFull message with I-CNT full is rare and may not be experienced in normal code. Simplest way to generate is to have an infinite loop and (rare) interrupt handler. This loop should increment a register or memory location - this value should correspond to total accumulated I-CNT. **Potential Future Enhancements** Table below is proposing some future enhancements for N-Trace messages. These were discussed during the development of the N-Trace specification. __Table 1\. Future Enhancements__ | Enhancement | Conformance | Notes | | -------------------------- | ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Instrumentation Data Trace | Nexus Compatible | Very likely (Nexus defines appropriate messages). It will require software to be instrumented by code sending data using trace infrastructure (Arm CoreSight ITM enabled many use-cases). | | Selective Data Trace | Nexus Compatible | Very likely (Nexus defines appropriate messages). It will allow sending some data in response to triggers (from debug module or external). | | Full Data Trace | Nexus Compatible | Likely (E-Trace supports it), but necessary bandwidth may be a problem. | | Smaller field sizes | Nexus Extension | Unlikely (too much of a change). Some of the fields may be made shorter (as not all cases are needed), but it may not be justified. | | System Bus Trace | Nexus Compatible | Likely (Nexus defines appropriate messages and there is a need for more than trace of harts). | | Additional TCODE | Nexus Extension | Possible, but more real-life examples are needed to justify it. | | Single MSEO bit | Nexus Compatible | Unlikely to be considered. It may provide (12.5% instead of 25% MSEO overhead), but it is more complex to handle by both encoder and decoders. | | More MDO bits | Nexus Compatible | Very unlikely to be considered. To keep byte alignment, 14 or 22 or 30-bit MDO may be considered. Even 14-bit will cause a lot of 'wasted' bits. | | | Each of the above enhancements should be first prototyped and validated using reference C encoder/decoder. | | ------------------------------------------------------------------------------------------------------------- | Change Log ==================== ## [](#change-log)Change Log PDF generated on: 2026-07-17 19:42:40 UTC ### [](#version-1-0-ratified)Version 1.0 (Ratified) * 2024-11-21 * Ratified state (concent identical as in 1.0\_rc51 PDF) Contributors ==================== ## [](#contributors)Contributors Key contributors to RISC-V N-Trace (Nexus-based Trace) specification in alphabetical order: Bruce Ableidinger (SiFive) ⇒ Initial SiFive donation, reviews Robert Chyla (IAR, SiFive, MIPS) ⇒ Most topics, editing, publishing Ernie Edgar (SiFive) ⇒ Initial SiFive donation, reviews Jay Gamoneda (NXP) ⇒ Reviews, contributing, editing Markus Goehrle (Lauterbach) ⇒ Reviews, updates Ved Shanbhogue (Rivos) ⇒ Detailed Architecture Review Committee notes Nino Vidovic (Segger) ⇒ Reviews 4.1. N-Trace Specific Trace Controls ==================== ## [](#4-1-n-trace-specific-trace-controls)4.1\. N-Trace Specific Trace Controls This chapter describes how fields and bits from Trace Encoder control registers (named using **trTe…​** pattern) are influencing N-Trace encoder and N-Trace protocol messages. N-Trace specific clarifications, in addition to description in [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) specification are provided. | | The table below does not provide names of Trace Encoder control registers as names of bits/fields used in Trace Control Interface are unique. | | ------------------------------------------------------------------------------------------------------------------------------------------------ | __Table 1\. Trace Encoder Parameters and Controls__ | Trace Control Field | Applicability | Description | | -------------------------- | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | trTeActive | **Required** | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | trTeEnable | **Required** | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | trTeInstTracing | **Required** | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | trTeEmpty | **Required** | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | trTeInstMode | **Required** | **3:** Generate instruction trace using [BTM](#mode%5FBTM) (Branch Trace Messaging) mode. **6:** Generate instruction trace using [HTM](#mode%5FHTM) (History Trace Messaging) mode. **0, 7:** See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. **1-2, 4-5:** Reserved for future N-Trace use.At least a value of **3** or **6** must be settable. | | trTeContext | Optional | Controls generation of [Ownership](#msg2%5FOwnership) messages. | | trTeInstTrigEnable | Optional | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | trTeInstStallOrOverflow | **Required** | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | trTeInstStallEna | Optional | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | trTeInhibitSrc | Optional | Controls generation of [SRC](#field%5FSRC) field. | | trTeInstSyncMode | **Required** | Controls generation of [Synchronizing Messages](#Synchronizing Messages) with [SYNC](#field%5FSYNC) field=2. | | trTeInstSyncMax | **Required** | Controls generation of [Synchronizing Messages](#Synchronizing Messages) with [SYNC](#field%5FSYNC) field=2. | | trTeFormat | **Required** | Must be set to **1** (which denotes N-Trace format). | | trTeVerMajor | **Required** | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | trTeVerMinor | **Required** | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | trTeCompType | **Required** | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | trTeProtocolMajor | **Required** | **Must be 1** to encode this version (1.0) of N-Trace protocol. Value different than 1 is considered a non-compatible version and must be rejected by the trace tool if it is only compliant with version 1.0 of the N-trace protocol. | | trTeProtocolMinor | **Required** | **Must be 0** to encode this version (1.0) of N-Trace protocol. When trTeProtocolMajor is 1, values other than 0 are considered down compatible extension and should be accepted by the trace tool. Any future non-compatible feature should be specifically enabled (by new control bits), so older tools (which never set these new bits) should work with it. | | trTeInstNoAddrDiff | Not applicable | Must be hard coded as **0**. | | trTeInstNoTrapAddr | Not applicable | Must be hard coded as **0**. | | trTeInstEnSequentialJump | Optional | See [Sequential Jump Optimization](#Sequential Jump Optimization) chapter. | | trTeInstEnImplicitReturn | Optional | See [Implicit Return Optimization](#Implicit Return Optimization) chapter. | | trTeInstEnBranchPrediction | Not applicable | Must be hard coded as **0**. | | trTeInstEnJumpTargetCache | Not applicable | Must be hard coded as **0**. | | trTeInstImplicitReturnMode | Optional | See [Implicit Return Optimization](#Implicit Return Optimization) chapter. | | trTeInstEnRepeatedHistory | Optional | See [Repeated History Optimization](#Repeated History Optimization) chapter. | | trTeInstEnAllJumps | Optional | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | trTeInstExtendAddrMSB | Optional | See [Virtual Addresses Optimization](#Virtual Addresses Optimization) chapter. | | trTeSrcID | Optional | Controls generation of [SRC](#field%5FSRC) field. | | trTeSrcBits | Optional | Controls generation of [SRC](#field%5FSRC) field. | | trTeInstFilters | Optional | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | trTeDataImplemented | Not applicable | **Must be 0** as IEEE-5001 Nexus Standard data trace messages are not part of version 1.0 of N-Trace specification. | | **Other** trTeData…​ | Not applicable | **Must be 0** as IEEE-5001 Nexus Standard defines data trace messages, future versions of N-Trace may allow these (as an optional extension). | | **All** trTeTrig…​ | Optional | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | **All** trTeFilter…​ | Optional | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | **All** trTeComp…​ | Optional | See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. | | trTsEnable | Optional | Part of potentially shared Timestamp Unit controls generation of [TSTAMP](#field%5FTSTAMP) field. See [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification for details of the Timestamp Unit. | Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at Copyright 2019-2024 by RISC-V International. 11.1. N-Trace Decoding Guidelines ==================== ## [](#11-1-n-trace-decoding-guidelines)11.1\. N-Trace Decoding Guidelines To reconstruct the program control flow using the N-Trace encoded stream of messages (as any other compressed trace) access to opcodes of instructions which were executed is necessary. This is usually done by providing an ELF file of a program being executed, but it can also be read-out from the target. Three types of information are needed: 1. Size of each instruction (16-bit or 32-bit). 2. Types of all instructions (corresponding to 'itype' signal on trace ingress port - based on analysis of opcodes). 3. For direct unconditional jumps and direct conditional branches an offset (to jump/branch destination) encoded in an opcode. Decoding must start from a [synchronizing message](#Synchronizing Messages). The synchronizing message provides the complete PC in the [F-ADDR](#field%5FF-ADDR) field. Transfers relative to this PC may then be inferred using subsequent messages till a new PC is transmitted in a subsequent synchronizing message. | | To provide partial decoding of big trace, messages with [F-ADDR](#field%5FF-ADDR) are transmitted periodically. Periodic [F-ADDR](#field%5FF-ADDR) transmission is also needed to decode trace from small, circular buffers. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#11-1-1-decoding-algorithm-principles)11.1.1\. Decoding Algorithm Principles To reconstruct the control flow of the program from N-Trace messages do the following: * Copy **HIST** and **I-CNT** fields (if available) to corresponding registers. * Handle [HIST](#field%5FHIST) register (while not empty): * Analyze code from the current PC through direct (inferable) unconditional jumps (all types) and direct conditional branches (each direct conditional branch will 'consume' a single bit from the **HIST** register). * Each encountered instruction should subtract 1 or 2 (INST\_LEN/16) from I-CNT (depending on the size of that instruction). * At the end (after the least significant bit from **HIST** is processed), the PC will be of the instruction executed after the last conditional branch (either taken or not-taken). * Handle [I-CNT](#field%5FI-CNT) register (while greater than 0x0): * Analyze code from current PC through direct (inferable) unconditional jumps (all types) - each encountered direct conditional branch must be treated as not-taken. * Each encountered instruction should subtract 1 or 2 (INST\_LEN/16) from I-CNT (depending on the size of that instruction). * It will reach either indirect, unconditional jump or I-CNT will become 0 to denote that some other 'event' (like exception, interrupt, trace off, trigger etc.) happened. * In BTM mode, direct conditional branch may be reached as last instruction. Next PC should be the destination address of that taken branch. * At the last step the [F-ADDR](#field%5FF-ADDR) or [U-ADDR](#field%5FU-ADDR) field (if available) should be applied. * This is either a destination address of indirect unconditional jump or an address of an exception/interrupt handler. * This will be the next PC where analysis of the next trace message should start. * Handle the [B-CNT](#field%5FB-CNT) or [RCODE](#field%5FRCODE) fields by repeating processing of the fields from previous trace message. | | Phrase **inferable unconditional jumps (all types)** include indirect unconditional jumps, which may be inferable. Extra fields like [SYNC](#field%5FSYNC)/[B-TYPE](#field%5FB-TYPE) only provide extra details, but are NOT essential for a decoder to reconstruct the PC flow. See [N-Trace Reference Code](https://github.com/riscv-non-isa/tg-nexus-trace/tree/main/refcode/c) for simple but fully functional implementation. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#11-1-2-decoding-trace-from-multiple-harts)11.1.2\. Decoding trace from multiple harts Decoder assigned to a specific hart should process only those messages tagged with a [SRC](#field%5FSRC) value corresponding to that hart. To facilitate this, all encoders operating within the same trace stream must configure the `trTeSrcBits` field identically to ensure a consistent source identifier bit width, and must each be assigned a unique `trTeSrcID` field value. This arrangement ensures that messages can be accurately attributed to their originating hart, allowing for precise and isolated trace analysis per hart. ### [](#11-1-3-decoding-trace-of-operating-systems)11.1.3\. Decoding trace of operating systems In case of complex operating systems (Linux etc.), where code consists of several independently built programs and libraries, decoders must be aware of different program images (e.g., ELF files) at different locations. [Ownership](#msg%5FOwnership) messages should provide enough context. Decoders must be also aware of assignment of **scontext/hcontext** values for programs and processes being traced. Operating systems may decide to migrate single process to different cores/harts. It may also be the case, when different threads from the same process (sharing code …​) will run in the same time on more than one core/hart. ### [](#11-1-4-decoding-self-modifying-or-jit-just-in-time-compiled-code)11.1.4\. Decoding self-modifying or JIT (Just In Time compiled) code Trace encoder is just encoding a stream of instructions passed by ingress port from the hart running it, but decoder must be aware of types of all instructions being executed. In case of self-modifying code (or JIT code), binary image (at moment of execution) must be available to decoder. How this can be done is not in the scope of this specification. 8.1. Field Encoding and Calculation Techniques ==================== ## [](#8-1-field-encoding-and-calculation-techniques)8.1\. Field Encoding and Calculation Techniques This chapter describes in detail how key fields (I-CNT, HIST, U-ADDR/F-ADDR and TSTAMP) are calculated and encoded. ### [](#8-1-1-address-compression)8.1.1\. Address Compression Address transmissions is compliant with the IEEE-5001 Nexus Standard (most significant bit 0-s skipped) with optional extension allowing to skip identical most significant bits. See [Virtual Addresses Optimization](#Virtual Addresses Optimization) chapter below for clarifications. Rules when generating addresses: * Only execution addresses (as seen by the hart) are reported. When virtual memory system is enabled these are virtual addresses. * The [F-ADDR](#field%5FF-ADDR) field is the full address associated with the trace event, provides a starting point for reconstructing relative addresses. * The [U-ADDR](#field%5FU-ADDR) field is a compressed address that is relative to the previous trace message with an address field. It is generated by XORing the address with the previous message. * To decode the full address from the relative address (U-ADDR) can be XORed with the previously decoded full address. * Address fields are sent beginning with bit 1 since all execution addresses are on a 2-byte boundaries (the least significant bit is always 0 and never sent). **Address XOR Calculation Examples** ============================================================================================== | Address | U-ADDR XOR calculations | F-ADDR/U-ADDR field sent | New REF | | | | | Address | ============================================================================================== |0x3FC04 | | F-ADDR=1_1111_1110_0000_0010=0x1FE02 | 0x3FC04 | ---------------------------------------------------------------------------------------------- |0x3F368 | REF =0011_1111_1100_0000_0100 | | | | | addr=0011_1111_0011_0110_1000 | | | | | XOR =0000_0000_1111_0110_1100 | U-ADDR=111_1011_0110=0x7B6 | 0x3F368 | ---------------------------------------------------------------------------------------------- |0x3E100 | REF =0011_1111_0011_0110_1000 | | | | | addr=0011_1110_0001_0000_0000 | | | | | XOR =0000_0001_0010_0110_1000 | U-ADDR=1001_0011_0100=0x934 | 0x3E100 | ============================================================================================== ### [](#8-1-2-hist-field-generation)8.1.2\. HIST Field Generation When operating in HTM mode, the encoder does not generate messages for conditional branches. Instead, it maintains a HIST register or accumulator to record the outcomes of these branches, whether taken or not-taken. Each conditional branch contributes a single bit to the HIST register, as follows: * A bit with a value of 1 is appended at the least significant position for a taken conditional branch. * A bit with a value of 0 is appended at the least significant position for a not-taken conditional branch. The HIST register may be implemented as a left-shift register. Initially, when the HIST register is empty, bit 0 of the register is set to 1, with all other bits set to 0\. Subsequent conditional branches cause the register to shift left, recording each taken or not-taken outcome in bit 0. Examples: Binary(MSB-LSB): 101=0x5 (two direct conditional branches, not-taken and taken) Binary(MSB-LSB): 1111=0xF (three direct conditional branches, all three taken) Binary(MSB-LSB): 10000=0x10 (four direct conditional branches, all four not-taken) Binary(MSB-LSB): 1=0x1 (no direct conditional branches at all) After transmission of the HIST field, the register is reset to its initial, empty state. Decoders must initiate the interpretation of the HIST field starting from the second most significant bit. The most significant bit, designated as the stop-bit, is invariably set to 1\. This second most significant bit—immediately following the stop-bit—encodes the outcome of the first conditional branch captured in the HIST register. Conversely, the least significant bit represents the outcome of the last conditional branch prior to the transmission of the HIST register. #### [](#8-1-2-1-hist-field-full)8.1.2.1\. HIST Field Full The transition of the most significant bit in the HIST register from 0 to 1 indicates the register is full. At this point, the entire register, including the most significant bit — which serves as the stop-bit — is transmitted using a [ResourceFull](#msg2%5FResourceFull) message with the [RCODE](#field%5FRCODE) field set to either 1 or 2. When a HIST register is full and its value is the same as that of the HIST field transmitted in previous [ResourceFull](#msg2%5FResourceFull) message, then the encoder may increment an internal **HREPEAT** counter (history repeat counter) instead of generating a ResourceFull message if the Repeated History Optimization is enabled. See [Repeated History Optimization](#Repeated History Optimization) chapter for further details. | | Trace decoders do not have to be aware about the actual size of the HIST field implemented by the encoder, however, to allow efficient implementation of trace encoders (and allowing HIST pattern detection) this N-Trace specification limits HIST field size to max 32-bits. Longer HIST fields would not provide much of a gain and would make repeated HIST field detection more costly (in terms of hardware resources). | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#8-1-3-i-cnt-details)8.1.3\. I-CNT Details The I-CNT field, present in most messages, transmits the value of the I-CNT counter, which counts the number of halfwords used to encode retired instructions. The I-CNT counter in the trace encoder is reset to 0, in accordance with the IEEE-5001 Nexus Standard, under one of the following two conditions: * When tracing starts or is restarted for any reason. * After the I-CNT counter value has been transmitted in a message. Every retired instruction MUST increment I-CNT counter by 1 (for 16-bit instruction) or by 2 (for 32-bit instruction). Specifically: * If an instruction is explicitly changing the PC (as jump or return), that instruction itself MUST update the I-CNT. * Instructions that either raise exceptions or are interrupted prior to retirement do not increment the I-CNT counter. | | In case of longer instructions (48-bit, 64-bit, …​) (future ISA standards or custom) I-CNT may increment by 3 or more. | | ------------------------------------------------------------------------------------------------------------------------- | When I-CNT counter is full (reaches its maximum value or overflow bit is set) it can be reported in one of two ways: * By using a [ResourceFull](#msg%5FResourceFull) message with [RCODE](#field%5FRCODE)\=0\. This method is applicable to both BTM and HTM. * Optionally, by using a [synchronizing message](#Synchronizing Messages) with **SYNC=4 (Sequential Instruction Counter)**. It may be only used in [BTM](#mode%5FBTM) mode. | | Overflow bit allows efficient handling of cases, when single ingress port cycle reports bigger I-CNT (several instructions retired). Reporting maximum value (exactly) is not required and smaller or bigger value may be reported instead. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#8-1-3-1-example-of-i-cnt-handling-in-btm-mode)8.1.3.1\. Example of I-CNT Handling in BTM mode As an illustration, let’s consider the following piece of pseudo-code (specific operations are abstracted as "…​" as they do not matter for this example): 0x100: c.add ... ; 16-bit instruction 0x102: b... 0x200 ; 32-bit instruction (direct conditional branch) 0x106: add ... ; 32-bit instruction 0x10A: b... 0x300 ; 32-bit instruction (direct conditional branch) 0x10E: c.add ... ; 16-bit instruction 0x110: add ... ; 32-bit instruction 0x114: c.ebreak ; 16-bit breakpoint (to stop the code) ... 0x200: c.add ... ; 16-bit instruction 0x202: c.ebreak ; 16-bit breakpoint (to stop the code) ... 0x300: add ... ; 32-bit instruction 0x304: c.ebreak ; 16-bit breakpoint (to stop the code) | | In the description below a range specified as <0x100..0x105> means that addresses 0x100 and 0x105 are both included in the address range. | | -------------------------------------------------------------------------------------------------------------------------------------------- | Let’s assume we start a trace from address 0x100\. The [ProgTraceSync](#msg%5FProgTraceSync) message with **I-CNT=0** and F-ADDR=0x80 (encoding an address 0x100) should be generated. Let’s analyze a collected trace of above program (in [BTM](#mode%5FBTM) mode) executed three times (each time with different flow). 1. First direct conditional branch at address 0x102 is taken. * A [DirectBranch](#msg%5FDirectBranch) message with **I-CNT=3** should be generated. It means, that a code block from <0x100..0x105> (as 6=2\*3) was executed and a direct conditional branch at the end of this block was taken. Decoder will know PC=0x200 from an opcode of the direct conditional branch at an address 0x102. * Next message should be [ProgTraceCorrelation](#msg%5FProgTraceCorrelation) with **I-CNT=1** describing range <0x200..0x201> till **C.EBREAK** instruction. 2. First direct conditional branch at address 0x102 is not taken and second direct conditional branch at address 0x10A is taken. * A [DirectBranch](#msg%5FDirectBranch) message with **I-CNT=7** should be generated. It means, that a code block from <0x100..0x10D> (as 0xE=2\*7) was executed and a direct conditional branch at the end of this block was taken. Decoder will know PC=0x300 from an opcode of the direct conditional branch at an address 0x10A. * Next message should be [ProgTraceCorrelation](#msg%5FProgTraceCorrelation) with **I-CNT=2** describing a range <0x300..0x303> till **C.EBREAK** instruction. 3. Both direct conditional branches (at 0x102 and 0x10A) are not taken. * In this case only [ProgTraceCorrelation](#msg%5FProgTraceCorrelation) with **I-CNT=10** should be generated. It is describing a range <0x100..0x113> (as 0x14=10\*2) till **C.EBREAK** instructions. | | Decoder must analyze every instruction in each code block being processed to know its size. It cannot skip to the end of the block by calculating **PC+I-CNT\*2** as it is UNKNOWN what is the size of the last instruction retired in that block. It may be (compressed) 16-bit or 32-bit (not-compressed) direct conditional branch. Without knowing an instruction size, the offset encoded in that direct conditional branch cannot be determined and the next PC (after a branch) cannot be calculated. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Above we analyzed some I-CNT values. Let’s consider other I-CNT values. * **I-CNT=1** is a correct value. * The only valid reason to generate a message with I-CNT=1 would be an exception (or interrupt) at an instruction at address 0x102. * In this case an encoder should generate an [IndirectBranch](#msg%5FIndirectBranch) or [IndirectBranchSync](#msg%5FIndirectBranchSync) message with I-CNT=1, B-TYPE=1 (exception) and U-ADDR/F-ADDR field encoding an address of an exception/interrupt handler. * **I-CNT=5** is also correct. * It means that exception/interrupt happened before an instruction at an address 0x10A (after instruction at 0x106). * **I-CNT=0** is also possible. * It should be generated when an interrupt was pending before we started the code (and trace) and instruction at address 0x100 was not executed/retired. * Another reason for I-CNT=0 may be a case, where instruction at address 0x100 will generate page fault or is illegal. | | Values of **I-CNT=4 or 6 or 9** are **INCORRECT** as it would mean that only half of corresponding 32-bit instruction was executed/retired. Decoders must report such incorrect I-CNT values and immediately abandon the decoding as it means that either an encoder is not conforming to this specification or a trace was captured incorrectly. Decoding may resume at the next [synchronizing message](#Synchronizing Messages), but it is not mandatory for all decoders to do so. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#8-1-3-2-example-of-i-cnt-handling-in-htm-mode)8.1.3.2\. Example of I-CNT Handling in HTM mode When the encoder is operating in [HTM](#mode%5FHTM) mode, I-CNT should be incremented at every retired instruction the same way as for BTM mode. However direct conditional branches (from code piece above …​) will NOT generate any trace messages, but each of them will add a bit to the HIST field. Example [code](#ICNT%5Fcode) (used to illustrate BTM trace) may generate messages with the following fields (for all three runs): 1. First direct conditional branch at address 0x102 is taken. * **I-CNT=4, HIST=0x3** (0b1\_1). Most significant bit=1 is stop bit, bit pattern '1' means that first direct conditional branch was taken. Encoder should continue till an address 0x200 (as the first direct conditional branch encountered was reported as taken) as I-CNT=3 describes a <0x100..0x105> range. Remaining I-CNT=1 describes a <0x200..0x201> range. 2. First direct conditional branch at address 0x102 is not taken and second direct conditional branch at address 0x10A is taken. * **I-CNT=9, HIST=0x5** (0b1\_01). Most significant bit=1 is stop bit, bit pattern '01' means that first direct conditional branch was not taken and second direct conditional branch was taken. Encoder should continue till an address 0x300 (as the second direct conditional branch encountered was reported as taken) as I-CNT=7 describes a <0x100..0x10D> range. Remaining I-CNT=2 describes a <0x300..0x303> range. 3. Both direct conditional branches (at 0x102 and 0x10A) are not taken. * **I-CNT=10, HIST-0x4** (0b1\_00). Most significant bit=1 is stop bit, bit pattern '00' means that two direct conditional branches were not taken. Encoder should continue till an address 0x114 as I-CNT=10 describes a code in a <0x100..0x113> range. #### [](#8-1-3-3-examples-of-i-cnt-field-full-generation)8.1.3.3\. Examples of I-CNT Field Full Generation Let’s consider the following example code: 0x100: c.add ... ; 16-bit instruction 0x102: b... 0x200 ; 32-bit instruction (direct conditional branch) 0x106: add ... ; 32-bit instruction 0x10A: add ... ; 32-bit instruction 0x10E: add ... ; 32-bit instruction 0x112: add ... ; 32-bit instruction 0x116: add ... ; 32-bit instruction 0x11A: c.add ... ; 16-bit instruction 0x11C: c.ebreak ; 16-bit breakpoint (to stop the code) and let’s assume (for simplicity) that the I-CNT counter is 4-bit wide (most significant bit being an overflow flag) and that direct conditional branch at an address 0x102 is not taken (so code will run from address 0x100 till breakpoint at address 0x11C). Trace with **Resource Full** message (HTM mode shown): * [ProgTraceSync](#msg2%5FProgTraceSync) (start of trace) * SYNC=3 (Exit from Debug Mode), I-CNT=0 (nothing executed as we are starting) * F-ADDR=0x80 (encoding starting address 0x100) * [ResourceFull](#msg2%5FResourceFull) (I-CNT overflown to 9 at an address 0x112) * RCODE=0 (I-CNT counter is full), **RDATA\[0\]=9** (I-CNT value overflown value) * [ProgTraceCorrelation](#msg2%5FProgTraceCorrelation) (describes entire <0x100..0x11C> range) * EVCODE=0 (Entry into Debug Mode), CDF=1 (I-CNT and HIST fields follow) * **I-CNT=5** (see note below), HIST=0x2 (one not-taken direct conditional branch) Trace with **SYNC=Sequential Instruction Counter** (BTM mode only): * [ProgTraceSync](#msg2%5FProgTraceSync) (start of trace) * SYNC=3 (Exit from Debug Mode), I-CNT=0 (nothing executed as we are starting) * F-ADDR=0x80 (encoding starting address 0x100) * [ProgTraceSync](#msg2%5FProgTraceSync) (I-CNT overflown to 9 at an address 0x112) * SYNC=4 (Sequential Instruction Counter), **I-CNT=9** (see note below) * F-ADDR=0x89 (encoding address 0x112) * [ProgTraceCorrelation](#msg2%5FProgTraceCorrelation) (describes <0x112..0x11C> range) * EVCODE=0 (Entry into Debug Mode), CDF=0 (only I-CNT field follows) * **I-CNT=5** (see note below) **Notes (for both trace options)** * Overflown **I-CNT=9** (or **RDATA\[0\]=9**) field describes <0x100..0x112> range (18 bytes long). * The **I-CNT=5** field describes <0x112..0x11C> range (12 bytes long). * In both cases total I-CNT is 9+5=14, what describes the entire <0x100..0x11C> range. * Debug Mode is entered before C.EBREAK instruction (as it never retires), so C.EBREAK is NOT included in I-CNT. * Using **ResourceFull** generates smaller, more compressed trace. * In real life examples it will allow generation of repeated history patterns and even better trace compression. * Using **SYNC=Sequential Instruction Counter** generates bigger trace (as potentially long F-ADDR field is reported). ### [](#8-1-4-synchronizing-messages)8.1.4\. Synchronizing Messages Synchronizing messages are messages with a [SYNC](#field%5FSYNC) field. That field identifies the reason for synchronization and such messages include the [F-ADDR](#field%5FF-ADDR) (full address) field to synchronize the PC with the PC observed by the encoder. All synchronizing messages MUST fully reset the encoder state, so decoding can be started from any of synchronizing messages. | | Trace requires different types of synchronization on different abstraction levels. Two major categories of synchronization are: **Instruction trace synchronization**: allows the trace decoder to synchronize onto an ongoing instruction trace stream. This is done via synchronizing messages, which are described in this chapter in more detail. **Message alignment synchronization**: allows the trace decoder to detect the trace message boundaries (i.e. start and end of a trace message) within a trace stream. This kind of synchronization is not described in this chapter. It can be done via idle cycles, and is described in the [PIB Idle Cycles Explained](#PIB Idle Cycles Explained) chapter in more detail. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 1\. SYNC Field Values__ | Value | Name | Required | Description | | ------ | ------------------------------ | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | External Trace Trigger | No | This message serves as a marker of external trigger. If trace is enabled by an external trigger SYNC=5 should be used. | | 1 | Exit from Reset | No | Core was reset without stopping (by watchdog for example). Address should be a reset vector. The HIST and I-CNT may be used to determine the PC of the last instruction retired before reset. | | 2 | Periodic Synchronization | Yes | Just periodic instruction trace synchronization (to allow decoding the trace from the middle or when circular RAM buffer was wrapped around overwriting part of earlier trace). The interval for periodic instruction trace synchronization gets configured via [trTeInstSyncMode](#trTeInstSyncMode) and [trTeInstSyncMax](#trTeInstSyncMax). | | 3 | Exit from Debug Mode | Yes | Very first synchronizing message after exit from debug mode. If trace is disabled (at exit from debug more) no messages should be generated. | | 4 | Sequential Instruction Counter | No | Generated when I-CNT counter is full. See [I-CNT Details](#I-CNT Details) chapter. | | 5 | Trace Enable | No | Generated when trace is re-enabled after a gap caused by trace being disabled (e.g. due to trace filters). This must not be used for exit from debug mode (in which case SYNC=3 must be used). | | 6 | Trace Event | No | Serves as a marker when debug watchpoint with action=4 triggered. See **RISC-V Debug Specification** for watchpoint setting details. | | 7 | Restart from FIFO overrun | Yes | First synchronization after a gap caused by an internal FIFO overun. Some trace messages before this synchronization message were lost. | | 8 | Reserved | \- | For future standard use. | | 9 | Exit from Power-down | No | When the hart is restarted after powered down. Similar to SYNC=1 (Exit from Reset) described above. | | 10..13 | Reserved | \- | For future standard use. | | 14..15 | Reserved | \- | For vendor defined codes. | Decoders should report synchronization SYNC field values from messages (including reserved codes) as it provides a reason for the program flow change. * All synchronizing messages fully reset the encoder state, so decoding can be started from this message. * Before resetting the encoder state, the trace up to the current location must be emitted (it includes HIST, I-CNT, HREPEAT and B-CNT counters). * All synchronizing messages emit an absolute [TSTAMP](#field%5FTSTAMP) field (if enabled), so decoder may calculate full/absolute timestamps from this message forward. * An [Ownership](#msg%5FOwnership) messages (if enabled) must be emitted immediately after all synchronizing messages. * Some synchronizing messages not related to code being executed (periodic, notifications etc.) may be emitted between indirect jumps. In such a case field [B-TYPE](#field%5FB-TYPE)\=0 will be emitted, but it will not mean indirect flow change. Periodic Synchronization are generated to allow easier decoding (not necessarily from the start of collected trace) and may only be reported when desired by the user (for debugging). | | Periodic Synchronization (SYNC=2) messages may not be precise and may be delayed if any other SYNC message (for example Sequential Instruction Counter, SYNC=4) is sent. In such a case, Periodic Synchronization may be even skipped as decoding may start from any Synchronizing Message. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#8-1-4-1-examples-of-synchronizing-messages)8.1.4.1\. Examples of Synchronizing Messages The following cases are created to help illustrate the type of N-trace [synchronizing message](#Synchronizing Messages) generated for different scenarios. Events which may occur while a hart is running or halted: ![example key](_images/example_key.PNG) **Case1: Enable/disable debug while tracing:** ![case1 enable disable debug while tracing](_images/case1_enable_disable_debug_while_tracing.PNG) **Case2: Enable trace while in debug:** ![case2 enable trace while in debug](_images/case2_enable_trace_while_in_debug.PNG) **Case3: Disable trace while in debug:** ![case3 disable trace while in debug](_images/case3_disable_trace_while_in_debug.PNG) **Case4: Sync trigger event (internal or external):** ![case4 sync trigger event](_images/case4_sync_trigger_event.PNG) **Case5: Enable and disable while in debug:** ![case5 enable disable while in debug](_images/case5_enable_disable_while_in_debug.PNG) **Case6: Periodic synchronization:** * First possibility provides choice of messages generated at exact periodic synchronization event `P`. * Second provides a choice of messages which may be generated delayed after the periodic event `P`. ![case6 periodic](_images/case6_periodic.PNG) **Superscript notes:** 1. ProgramTraceSync message may be replaced with DirectBranchSync, IndirectBranchHistSync, IndirectBranchHistSync. 2. ProgramTraceSync message may be generated for a SYNC event, however, HIST information will not be reported. For HTM mode, the IndirectBranchHistSync or IndirectBranchSync message with SYNC=6 (Trace Event) should be used to ensure no trace data is lost. 3. Next available **…​Branch…​** message upgraded to **…​Branch…​Sync** counterpart, so SYNC code is reported. ### [](#8-1-5-timestamp-reporting)8.1.5\. Timestamp Reporting Timestamp reporting must be enabled by [trTsEnable](#trTsEnable) trace control bit. If timestamp is enabled, all [Synchronizing Messages](#Synchronizing Messages) include an absolute timestamp value with upper zeroes suppressed. Other message types with timestamp emit the timestamp as relative offset from last reported (absolute or relative) timestamp. | | The TSTAMP field is a variable-length field, and most significant bits set to 0 will not be transmitted. This approach provides good compression for both relative and absolute timestamps. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | To reconstruct the full timestamp, software begins at a [synchronizing message](#Synchronizing Messages) and stores the TSTAMP value found there, zero-extended to the full timestamp width. Shortly after starting a trace session, even a 64-bit timestamp will typically require far less than 64 bits to transmit. Software extracts the compressed TSTAMP from each message thereafter and adds it with the previous decompressed timestamp to obtain the full timestamp value associated with this message. The following rules must be observed: * If timestamps are enabled, ALL [Synchronizing Messages](#Synchronizing Messages) must include an absolute TSTAMP value. * It is not required for all non-synchronizing messages to always report a timestamp. Doing so may be opted for saving trace bandwidth or in the case of sending back-to-back messages. * The absolute timestamp cannot exceed 64 bits (even with 1ps resolution, 64-bit counters will overflow in about 584 years). * Implementations may choose a smaller counter. Trace tools may assume timestamp will not overflow in a single session, although adding support for overflow is not significantly challenging. * It is suggested that in multi-hart systems, all Trace Encoders use a shared timestamp (for better trace correlation), but it is not mandatory. * In all cases, when an address is provided, the timestamp should reflect the time when an event leading to that address occurred. | | If the above is not feasible, timestamps should be at least reported consistently, ensuring that the time distance between distant events (for example, a periodic timer interrupt) can be reliably calculated. It is necessary to assure that the time reported at exceptions/interrupt handlers reflects the moment when exception or interrupt was observed. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#8-1-6-corner-cases-and-sequences)8.1.6\. Corner Cases and Sequences Normal program flow generates a sequence of messages with I-CNT>0 (reporting at least 1 instruction retired), some HIST fields (to report direct conditional branches) and F-ADDR/U-ADDR fields (to report uninferable unconditional flow changes). However, sometimes normal flow is interrupted (by exception or interrupt) or some other extra event (trigger/enable/disable) happens and sequence of messages or values of some fields may be a bit unusual. Table below is trying to explain some corner cases. __Table 2\. Corner Cases__ | Sequence of events | Messages Generated | | ----------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Back to back return | Second message should have [I-CNT](#field%5FI-CNT)\=1 or 2 (depending on the size of the second return instruction). | | Other back to back jumps or branches | Same as above (depending on the size of a second instruction). | | Back to back exceptions | Second message with [B-TYPE](#field%5FB-TYPE)\=2 or 1 (Exception) and [I-CNT](#field%5FI-CNT)\=0 (nothing executed in between). | | Exception at interrupt destination | Same as above. | | Pending interrupt at debug mode exit | [ProgTraceSync](#msg2%5FProgTraceSync) with [SYNC](#field%5FSYNC)\=3 followed by message with [B-TYPE](#field%5FB-TYPE)\=3 or 1 (Interrupt). | | Exception at first instruction traced | [ProgTraceSync](#msg2%5FProgTraceSync) with [SYNC](#field%5FSYNC)\=3 followed by a message with [B-TYPE](#field%5FB-TYPE)\=2 or 1 (Exception). | | Trace starts disabled | [ProgTraceCorrelation](#msg2%5FProgTraceCorrelation) with [EVCODE](#field%5FEVCODE)\=4 (Trace Disabled). Once trace is enabled message with [SYNC](#field%5FSYNC)\=5 (Trace Enable). | | Hart halted with trace disabled | [ProgTraceCorrelation](#msg2%5FProgTraceCorrelation) with [EVCODE](#field%5FEVCODE)\=0 (Enter Debug mode) and [I-CNT](#field%5FI-CNT)\=0 (nothing executed). | | Exception/Interrupt immediately following trap return | Usual messages describing instructions up to return from trap (MRET/SRET) instruction.Synchronizing message with an address of trap return with [I-CNT](#field%5FI-CNT)\=0 (as nothing executed after a trap return).Optionally, an [Ownership](#msg%5FOwnership) messages describing privilege level after return from a trap.Synchronizing message with an address of interrupt/exception handler and appropriate SYNC code.Optionally, an [Ownership](#msg%5FOwnership) messages describing privilege level of new exception/interrupt handler. | 10.1. Rules of Generating Messages ==================== ## [](#10-1-rules-of-generating-messages)10.1\. Rules of Generating Messages This chapter explicitly addresses 16-bit and 32-bit instructions as defined in the currently ratified RISC-V instruction set. Nonetheless, the guidelines provided herein are applicable to any instruction size that is a multiple of 16-bit, should such instructions be defined in the future. **Main Rules** 1. **Inferable Instructions**: This category includes instructions that do not perform control transfers or are direct jumps. The subsequent program counter (PC) for these instructions can be determined through static analysis of the binary code. Because these instructions exhibit a predictable execution flow, they are termed "inferable," and no trace is generated for them. 2. **Uninferable Instructions**: This category comprises conditional branches and indirect jumps, including return and indirect calls. Due to the unpredictability of the next PC as determined through static analysis alone, uninferable instructions require trace. 3. **Interrupts and Exceptions**: Control flow changes caused by interrupts and exceptions necessitate trace generation. These events alter the flow in an unpredictable manner, like uninferable instructions, thereby requiring their occurrences to be traced. **Detailed Rules** 1. If tracing is started (or restarted after it was disabled), a [ProgTraceSync](#msg%5FProgTraceSync) message is generated. * This message specifies the reason for the start in the [SYNC](#field%5FSYNC) field and includes full address in the [F-ADDR](#field%5FF-ADDR) field. 2. A retired 16-bit instruction increments the [I-CNT](#field%5FI-CNT) counter by 1, while a retired 32-bit instruction increments it by 2. 3. The following types of instructions allow trace decoders to determine the next PC and encoder should not generate any trace for them. * Instruction which is not control transfer instructions should advance PC to the next instruction (increment by 2 or 4). * Direct (inferable) unconditional jump should set next PC to jump destination (PC plus an offset obtained from opcode). * Not-taken direct conditional branch (in BTM mode) should advance PC to the next instruction (increment by 2 or 4). 4. Indirect, unconditional jump instruction is handled as: * In BTM mode, an [IndirectBranch](#msg2%5FIndirectBranch) message is generated. * In HTM mode, an [IndirectBranchHist](#msg2%5FIndirectBranchHist) message is generated. Should the [HIST](#field%5FHIST) field be empty, an [IndirectBranch](#msg2%5FIndirectBranch) message may optionally be generated instead. 5. Direct, conditional branch instruction is handled as: * In BTM mode, a [DirectBranch](#msg%5FDirectBranch) message is generated, but only if the branch is taken. * In HTM mode, the outcome of the branch (1 for taken or 0 for not taken) is appended as a single bit into the branch history buffer ([HIST](#field%5FHIST) register). 6. When tracing is stopped or disabled, a [ProgTraceCorrelation](#msg%5FProgTraceCorrelation) message is generated. * This message included a reason for stopping or disabling (specified in the [EVCODE](#field%5FEVCODE) field), the [I-CNT](#field%5FI-CNT) and an optional [HIST](#field%5FHIST) field. These details allow for the calculation of the last PC. 7. When a generated message includes [I-CNT](#field%5FI-CNT) counter value or [HIST](#field%5FHIST) register value, the corresponding counter and/or register are reset. * If the I-CNT counter is full, a [ResourceFull](#msg%5FResourceFull) message, indicating that I-CNT counter is full, is generated. Subsequently, the I-CNT is reset. * Similarily, if the HIST register reaches it capacity, a [ResourceFull](#msg%5FResourceFull) message, specifying that the HIST register is full, is generated. The HIST register is then reset. **Extended Rules** These rules are augmenting the above rules if the corresponding configuration setting is set. 1. Call and return instructions may optionally be handled as described in the [Implicit Return Optimization](#Implicit Return Optimization) chapter and may generate no trace. 2. By default, the target of an indirect unconditional jump is always considered an uninferable PC discontinuity. However, if the register that specifies the jump target was loaded with a constant then it can be considered inferable under some circumstances. * Such instruction sequences may be detected and in such a case no trace is generated. * This optional feature is described in detail in the [Sequential Jump Optimization](#Sequential Jump Optimization) chapter. ### [](#10-1-1-custom-instructions)10.1.1\. Custom Instructions Custom instructions (or any future ratified instructions) which are not changing PC flow do not require any special treatment. Trace decoders should only look at instructions which may change PC flow and for all other instructions only advance PC (+2 or +4). Custom instruction which may change a PC (other than simple advance to next instruction) should be traced in one of the following ways: * If the PC just advances to the next instruction, it should only increment I-CNT. Decoder will just advance the PC. * If the program flow changes as result of a custom instruction, the custom instruction should be traced as an indirect unconditional jump (even if it is not an indirect unconditional jump). That way, the destination address will be reported (as F-ADDR or U-ADDR fields). Decoder will change PC to an address specified in this message. Such an approach will NOT require changes/adaptation in trace decoders. To illustrate this let’s consider the following piece of code with custom instruction **XYZ**: 0x100: add ... ; 32-bit instruction 0x104: XYZ ; 32-bit instruction (custom conditional branch to 0x200 - it does not matter if direct or indirect ...) 0x108: c.add ... ; 16-bit instruction 0x10A: c.ebreak ; 16-bit breakpoint (to stop the code) ... 0x200: c.add ... ; 16-bit instruction 0x202: c.ebreak ; 16-bit breakpoint (to stop the code) It can be traced as follows (exact type of messages do not matter): * Single message (if branch was not taken) * **I-CNT=5** ⇒ Instruction XYZ did not change the flow and code in range <0x100..0x10A) got executed * Two messages (if branch was taken) * **I-CNT=4**, **F-ADDR=0x100** (denote address 0x200) ⇒ Code in range <0x100..0x108) got executed and next PC after instruction XYZ is 0x200 * **I-CNT=1** ⇒ Code in range <0x200..0x202) got executed next | | If custom instruction will generate some other trace (for example some new type of direct conditional branch which may add HIST bit), decoders must be extended to be aware about the type of this custom instruction. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | If a custom instruction cannot be mapped into one of existing **itype** encodings, it may use custom encoding. In such a case encoder (and decoder …​) must be enhanced. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#10-1-2-pseudo-code-of-simple-n-trace-encoder)10.1.2\. Pseudo-code of Simple N-Trace Encoder Code below is a simplified part of actual C-code used by the reference encoder (in C). It defines two functions: * NTraceEncoderInit(void) - initialize state of encoder * NTraceEncoderHandleRetired(uint64\_t `addr`, uint32\_t `flags`) - handle single retired instruction * `addr` \- address of retired instruction * `info` \- information about instruction (type, size, taken/not-taken) ```c // Use N-Trace TCODE messages #define NEXUS_TCODE_Ownership 2 #define NEXUS_TCODE_DirectBranch 3 #define NEXUS_TCODE_IndirectBranch 4 #define NEXUS_TCODE_Error 8 #define NEXUS_TCODE_ProgTraceSync 9 #define NEXUS_TCODE_DirectBranchSync 11 #define NEXUS_TCODE_IndirectBranchSync 12 #define NEXUS_TCODE_ResourceFull 27 #define NEXUS_TCODE_IndirectBranchHist 28 #define NEXUS_TCODE_IndirectBranchHistSync 29 #define NEXUS_TCODE_RepeatBranch 30 #define NEXUS_TCODE_ProgTraceCorrelation 33 // Functions/macros which encode bits in 'info' (example...) #define INFO_LINEAR 0x1 // Linear (plain instruction or not-taken BRANCH) #define INFO_4 0x2 // If not 4, it must be 2 on RISC-V #define INFO_INDIRECT 0x8 // Possible for most types above #define INFO_BRANCH 0x10 // Always direct on RISC-V (may have LINEAR too) #define InfoIsBranchTaken(info) (!((info) & INFO_LINEAR)) #define InfoIsSize32(info) ((info) & INFO_4) #define InfoIsBranch(info) ((info) & INFO_BRANCH) #define InfoIsIndirect(info) ((info) & INFO_INDIRECT) // Function which emit N-Trace messages (all are empty here) void EmitFix(int nbits, uint32_t value); // Emit fixed-size field void EmitVar(uint64_t value); // Emit variable size field void EmitEnd(); // Terminate message // Encoder configuration options const bool enco_opt_branch_history = true; // Configuration option const uint32_t enco_opt_limICNT = 0x10000; // Limit of ICNT (max is 6+6+4 bits) const uint32_t enco_opt_limHIST = 0x40000000; // Limit of HIST (max is 5*6 bits) // Encoder state variables static uint32_t encoNextEmit = 0; // TCODE to be emitted next time static uint32_t encoICNT = 0; // ICNT accumulated static uint32_t encoHIST = 1; // HIST accumulated (most significant bit is guardian bit) static uint64_t encoADDR = 0; // Last emitted address void NTraceEncoderInit() { encoADDR = 0; encoICNT = 0; // Empty ICNT and HIST encoHIST = 1; encoNextEmit = NEXUS_TCODE_ProgTraceSync; } void NTraceEncoderHandleRetired(uint64_t addr, uint32_t info) { // Optionally emit what was determined previously if (encoNextEmit != 0) { EmitFix(6, encoNextEmit); // Emit TCODE (as determined) // Emit message fields (accordingly ...) if (encoNextEmit == NEXUS_TCODE_ProgTraceSync) { EmitFix(4, 1); // Emit SYNC=1 (4-bit) EmitVar(encoICNT); // Emit ICNT (variable) EmitVar(addr >> 1); // Emit FADDR (variable) } else if (encoNextEmit == NEXUS_TCODE_IndirectBranchHist || encoNextEmit == NEXUS_TCODE_IndirectBranch) { EmitFix(2, 0); // Emit BTYPE=0 (2-bit) EmitVar(encoICNT); // Emit ICNT (variable) EmitVar((encoADDR ^ addr) >> 1); // Emit UADDR (variable) if (encoNextEmit == NEXUS_TCODE_IndirectBranchHist) { EmitVar(encoHIST); // Emit HIST (variable) } } else if (encoNextEmit == NEXUS_TCODE_DirectBranch) { EmitVar(encoICNT); // Emit ICNT (variable) } EmitEnd(); // It will mark last entry with MSEO=11 and flush it if (encoNextEmit != NEXUS_TCODE_DirectBranch) { encoADDR = addr; // This is new address } encoNextEmit = 0; // Only one time encoICNT = 0; // Start from 'empty' ICNT and HIST encoHIST = 1; } // Update ICNT uint32_t prevICNT = encoICNT; // In case ICNT will overflow now, we need to emit previous value ... if (InfoIsSize32(info)) encoICNT += 2; else encoICNT += 1; // Determine type of message (only if this is branch or indirect ...) if (InfoIsBranch(info)) { if (enco_opt_branch_history) { // Update branch history buffer (add least significant bit) if (InfoIsBranchTaken(info)) encoHIST = (encoHIST << 1) | 1; // Mark branch as taken else encoHIST = (encoHIST << 1) | 0; // Mark branch as not-taken } else { if (InfoIsBranchTaken(info)) encoNextEmit = NEXUS_TCODE_DirectBranch; // Emit destination address (next retired) else ; // Not-taken branch is considered as linear instruction } } else if (InfoIsIndirect(info)) { if (enco_opt_branch_history) encoNextEmit = NEXUS_TCODE_IndirectBranchHist; // Emit destination address (next retired) else encoNextEmit = NEXUS_TCODE_IndirectBranch; // Emit destination address (next retired) } // Optionally emit ICNT full if (encoICNT > enco_opt_limICNT) // Instruction count overflown? { // Emit ResourceFull with ICNT before this instruction EmitFix(6, NEXUS_TCODE_ResourceFull); EmitFix(4, 0); // RCODE=0 (ICNT full) EmitVar(prevICNT); // RDATA=ICNT (before overflown) EmitEnd(); // It will mark last entry with MSEO=11 and flush it // Set ICNT for this instruction if (InfoIsSize32(info)) encoICNT = 2; else encoICNT = 1; } // Optionally emit HIST full if (encoHIST & enco_opt_limHIST) // Is HIST buffer overflown? { // Emit history BEFORE this instruction (remove least significant bit) EmitFix(6, NEXUS_TCODE_ResourceFull); EmitFix(4, 1); // RCODE=1 (HIST full) EmitVar(encoHIST >> 1); // RDATA=HIST (before overflown) EmitEnd(); // It will mark last entry with MSEO=11 and flush it // Keep single HIST for this branch (guardian | single least significant bit from encoHIST) encoHIST = (0x1 << 1) | (encoHIST & 0x1); } } ``` 2.1. Trace Ingress Port ==================== ## [](#2-1-trace-ingress-port)2.1\. Trace Ingress Port N-Trace uses the same ingress port as specified in [E-Trace Specification](#E-Trace%5FSpecification) (chapter **4 Instruction Trace Interface**). * As this specification does not define the data trace yet, sub-chapters **4.3 Data Trace Interface requirements** and **4.4 Data Trace Interface** are not applicable. * It is an ambition to extract single, shared **RISC-V Trace Ingress Port** specifications (combining this chapter with relevant E-Trace chapter). * Names of 'itype' values used in this specification are a bit different than names in E-Trace specification. These names were unconditionally enforced by ARC (during review phase) as compulsory in all relevant specifications from now on. The table below provides a detailed mapping of causes for terminating an instruction block to the corresponding **itype** encoding. It could be used during development of ingress port logic inside of a hart. For some instructions operands matter - for example **JALR rd,rs1** instruction may generate 5 different, distinct **itype** values. __Table 1\. Generating itype for different instructions__ | Instruction | Condition/Notes | itype Value/Name | | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------- | | Exception in instruction | An exception trap that occurred following the final retired instruction in the block. | 1 = Exception | | EBREAK, ECALL, C.EBREAK | An exception trap that occurred following the final retired instruction in the block due to these instructions. These instructions do not retire. | 1 = Exception | | Interrupted instruction | An interrupt trap occurred following the final retired instruction in the block. | 2 = Interrupt | | MRET, SRET | Return from an exception or interrupt handler. | 3 = Trap return | | [Conditional branch](#itype%5Fbranch) | Not-taken direct, conditional branch. | 4 = Not-taken branch | | [Conditional branch](#itype%5Fbranch) | Taken direct, conditional branch. | 5 = Taken branch | | Any other instruction | All other instructions that are not directly listed in this table. | 0 = No special type | | **Values of itype ([3-bit](#itype%5F3%5F4)) (without [Implicit Return Optimization](#Implicit Return Optimization)**) | | | | JAL rd | Any direct jump/call. | 0 = No special type | | JALR rd, rs | Any indirect jump/call. | 6 = Indirect jump (with or without linkage) | | C.J or C.JAL | C extension has direct jump/calls only. | 0 = No special type | | CM.JT | Defined by [Zcmt](#zcmt) extension. | 0 = No special type | | CM.JALT | Defined by [Zcmt](#zcmt) extension. | 0 = No special type | | CM.POPRET\* | Defined by **Zcmp** extension. | 6 = Indirect jump (with or without linkage) | | **Values of itype ([4-bit](#itype%5F3%5F4)) (needed for [Implicit Return Optimization](#Implicit Return Optimization)**). [link](#link) means **x1** or **x5**. | | | | JAL rd | rd = link | 9 = Direct call | | rd = **x0** | 11 = Direct jump (without linkage) | | | rd != link and rd != **x0** | 15 = Other direct jump (with linkage) | | | JALR rd, rs | rd = link and rs != link | 8 = Indirect call | | rd = link and rs = link and rd = rs | 8 = Indirect call | | | rd = link and rs = link and rd != rs | 12 = Co-routine swap | | | rd != link and rs = link | 13 = Function return | | | rd = **x0** and rs != link | 10 = Indirect jump (without linkage) | | | rd != link and rd != **x0** and rs != link | 14 = Other indirect jump (with linkage) | | | C.JAL | Expands to JAL x1, offset | 9 = Direct call | | C.JALR rs | rs = **x5** | 12 = Co-routine swap | | rs != **x5** | 8 = Indirect call | | | C.JR rs | rs = link | 13 = Function return | | rs != link | 10 = Indirect jump (without linkage) | | | C.J | Expands to JAL x0, offset | 11 = Direct jump (without linkage) | | CM.JT | Defined by [Zcmt](#zcmt) extension. | 11 = Direct jump (without linkage) | | CM.JALT | Defined by [Zcmt](#zcmt) extension. | 9 = Direct call | | CM.POPRET\* | Defined by **Zcmp** extension. | 13 = Function return | | | Branches (**itype**\=4, 5) are always conditional, direct branches. In RISC-V ISA all jumps, calls, returns are always unconditional. | | ---------------------------------------------------------------------------------------------------------------------------------------- | | | Extended 4-bit **itype** (codes 8..15) are only necessary when [Implicit Return Optimization](#Implicit Return Optimization) is implemented. | | ----------------------------------------------------------------------------------------------------------------------------------------------- | | | Symbol link means register **x1** or **x5** as specified in **The RISC-V Instruction Set Manual, Volume I: Unprivileged ISA** document. | | ------------------------------------------------------------------------------------------------------------------------------------------ | | | Jump instructions (CM.JT and CM.JALT) defined by ratified **Zcmt** extension are handled as direct (inferable) jumps as jump tables are assumed to be static and known to the decoder. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Table below defines how N-Trace encoder should handle different 3-bit **itype** values on trace ingress port. __Table 2\. Handling of 3-bit itype values__ | # | itype | Encoder Action | | - | --------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | No special type | Only update [I-CNT](#field%5FI-CNT) field. | | 1 | Exception | Update [I-CNT](#field%5FI-CNT) field.Emit Indirect Branch message with [B-TYPE](#field%5FB-TYPE)\=2 or 1. **IMPORTANT:** An address emitted is known at the next valid ingress port cycle. | | 2 | Interrupt | Update [I-CNT](#field%5FI-CNT) field.Emit Indirect Branch message with [B-TYPE](#field%5FB-TYPE)\=3 or 1. **IMPORTANT:** An address emitted is known at the next valid ingress port cycle. | | 3 | Trap return | Update [I-CNT](#field%5FI-CNT) field.Emit Indirect Branch message with [B-TYPE](#field%5FB-TYPE)\=0. **IMPORTANT:** An address emitted is known at the next valid ingress port cycle. | | 4 | Not-taken branch | **For [BTM](#mode%5FBTM) mode:** Only update [I-CNT](#field%5FI-CNT) field. **For [HTM](#mode%5FHTM) mode:** Update [I-CNT](#field%5FI-CNT) field. Add 0 as least significant bit to [HIST](#field%5FHIST) field. | | 5 | Taken branch | **For [BTM](#mode%5FBTM) mode:** Update [I-CNT](#field%5FI-CNT) field. Generate [DirectBranch](#msg%5FDirectBranch) message. **For [HTM](#mode%5FHTM) mode:**Update [I-CNT](#field%5FI-CNT) field.Add 1 as least significant bit to [HIST](#field%5FHIST) field. | | 6 | Indirect jump (with or without linkage) | Update [I-CNT](#field%5FI-CNT) field.Emit Indirect Branch message with [B-TYPE](#field%5FB-TYPE)\=0. **IMPORTANT:** An address emitted is known at the next valid ingress port cycle. | | 7 | Reserved | \- | When the **itype** input of ingress port is 4-bit wide, the Indirect jump (with or without linkage) **itype=6** should not be generated and one of the following values should be generated instead. Encoder must handle call stack action as described in the [Implicit Return Optimization](#Implicit Return Optimization) chapter (if enabled). __Table 3\. Handling of 4-bit itype values__ | # | itype | Encoder Action | Stack Action | | -- | ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------ | | 8 | Indirect call | Update [I-CNT](#field%5FI-CNT) field. Emit Indirect Branch message with [B-TYPE](#field%5FB-TYPE)\=0. | Push | | 9 | Direct call | Only update [I-CNT](#field%5FI-CNT) field. | Push | | 10 | Indirect jump (without linkage) | Update [I-CNT](#field%5FI-CNT) field.Emit Indirect Branch message with [B-TYPE](#field%5FB-TYPE)\=0. [Same handing](#same%5Fhandling) as **itype=14**. | \- | | 11 | Direct jump (without linkage) | Only update [I-CNT](#field%5FI-CNT) field. [Same handing](#same%5Fhandling) as **itype=15**. | \- | | 12 | Co-routine swap | Update [I-CNT](#field%5FI-CNT) field.If Pop does not returns the same address as PC at next valid ingress port cycle, emit Indirect Branch message with [B-TYPE](#field%5FB-TYPE)\=0. | Pop,Push | | 13 | Return | Update [I-CNT](#field%5FI-CNT) field.If Pop does not returns the same address as PC at next valid ingress port cycle, emit Indirect Branch message with [B-TYPE](#field%5FB-TYPE)\=0. | Pop | | 14 | Other indirect jump (with linkage) | Update [I-CNT](#field%5FI-CNT) field.Emit Indirect Branch message with [B-TYPE](#field%5FB-TYPE)\=0. [Same handing](#same%5Fhandling) as **itype=10**. | \- | | 15 | Other direct jump (with linkage) | Only update [I-CNT](#field%5FI-CNT) field. [Same handing](#same%5Fhandling) as **itype=11**. | \- | | | N-Trace messages do not differentiate instructions classified as **…​ jump (with linkage)** and **…​ jump (without linkage)**, so both N-Trace ingress ports and N-Trace encoders implementations may ignore differences between **with/without linkage** values. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | If optional [trTeInstEnAllJumps](#trTeInstEnAllJumps) bit is set, trace ingress port is required to report **itype**\=5 (Taken branch) for all direct unconditional jumps, which are normally reported as **itype** \= 0 or 15. | | The N-Trace encoder does not require **cause** and **tval** ingress port signals, which are valid only for exceptions and interrupts, as these details are not reported in N-Trace messages. Instead, N-Trace solely provides the address of the exception or interrupt handler. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Since almost every ingress port cycle updates I-CNT, there is a possibility of overflow. For more information, see [I-CNT Details](#I-CNT Details) chapter regarding I-CNT management and overflow handling. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 1.1. Introduction to N-Trace ==================== ## [](#1-1-introduction-to-n-trace)1.1\. Introduction to N-Trace This **RISC-V N-Trace (Nexus based trace) Specification** is based on the well-established **IEEE-5001 Nexus Standard** tailored to support the trace of RISC-V ISA cores, harts and SoC/MCU designs. It serves multiple audiences: * N-Trace encoder logic/IP developers. * Validation teams testing of N-Trace implementation. * Debug and trace tools developers. * Software programmers utilizing the trace for debugging and performance tuning of RISC-V-based systems. This specification, together with the **RISC-V Trace Control Interface Specification** and **RISC-V Trace Connectors Specification** provide a complete, end-to-end, trace system for RISC-V based SoC. A trace ingress port, which serves as the connection between the RISC-V hart and the trace system, is defined in the ratified **Efficient Trace for RISC-V Specification**. This port enables the RISC-V hart to communicate execution information to the trace system. The N-Trace encoder is responsible for encoding an execution flow into a stream of trace messages. This document describes an appropriate selection of N-Trace messages compatible with the original IEEE-5001 Nexus Standard. The primary objective was to define the program flow trace messages. Extensions have been introduced to enable better trace compression. Future versions may include IEEE-5001 Nexus-compatible data and bus trace. The registers controlling the N-trace decoder are defined by the **RISC-V Trace Control Interface Specification**. This specification is shared with E-trace, so not all registers and register fields are supported by N-trace. Trace connectors defined by IEEE-5001 Nexus Standard were debug oriented, so could not be directly applied. Instead, industry standard MIPI-compliant connectors are defined in **RISC-V Trace Connectors Specification**. These connectors are pure extensions of debug-only, MIPI-compliant connectors defined by ratified **RISC-V Debug Specification**. ### [](#1-1-1-related-specifications)1.1.1\. Related Specifications This document provides reference to separated documents developed together with this **RISC-V N-Trace Specification**: * **RISC-V Trace Control Interface Specification** \- Defines RISC-V trace control interface. * This document is intended to be shared with ratified **Efficient Trace for RISC-V Specification**. * **RISC-V Trace Connectors Specification** \- Defines RISC-V trace connectors (for external trace probes). Ratified **Efficient Trace for RISC-V Specification** defines RISC-V Trace Ingress Port signals (chapter **4 Instruction Trace Interface**). At the moment of this writing this is version 2.0 (ratified May 5-th 2022). | | In the future trace ingress port may be defined in separated document - in such a a case reference to E-Trace specification will not be necessary. | | ----------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#1-1-2-trace-encoder-interfaces)1.1.2\. Trace Encoder Interfaces The diagram below shows one possible implementation with only a single RISC-V hart. In a system with multiple cores/harts the **Trace Ingress Port**, **Trace Encoder Control** and **Trace Encoder** blocks should be replicated for each hart. The main **Trace Control Layer** controlling other (shared) components in the trace system is not replicated. ![Trace Encoder Interfaces](_images/diag-45734442c9d464e2412a4c8a2bff9d9e3a95e730.svg) Figure 1\. Trace Encoder Interfaces | | Placement of the Trace Encoder and Trace Control Layer are implementation dependent. | | --------------------------------------------------------------------------------------- | ### [](#1-1-3-definitions-and-terminology)1.1.3\. Definitions and Terminology __Table 1\. Terms Used In This Specification__ | Term | Definition | | ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Message | N-Trace messages are sequences of bytes. First byte of every message includes the TCODE field, which defines the type of information carried in the message and its format. When messages are transmitted or stored, a protocol, described in [N-Trace Transmission Protocol](#N-Trace Transmission Protocol) chapter, defines the start and the end of each message. | | Field | A field is a distinct piece of the information contained within a message, and messages may contain one or more fields (in addition to the first TCODE field). Fields can be either of fixed-length or variable-length. Several fields may be packed into single byte and single field may span multiple bytes. Definitions of all fields can be found in [Fields in Messages](#Fields in Messages) chapter. | | Variable-length Field | Specifying that a field is variable-length (**Var** used as field size definition) means that the message must contain the field, but the field’s size may vary from a minimum of 1 bit. When messages are transmitted or stored, variable-length fields must end on a byte boundary. If necessary, they must zero-fill bit positions beyond the highest order bit of the variable-length data. Because variable-length fields may be of different lengths in messages of the same type, when messages are transmitted or stored, a protocol, described in [N-Trace Transmission Protocol](#N-Trace Transmission Protocol) chapter, defines the end of each variable-length field. | | Configurable Field | Configurable field (**Cfg** used as field size) means that existence and size of this field depends on some configuration setting. See [N-Trace Specific Trace Controls](#N-Trace Specific Trace Controls) chapter for details. | | N-Trace | IEEE-5001 Nexus Standard Based Trace for RISC-V (as defined by this specification). | | E-Trace | Efficient Trace for RISC-V (as defined by [E-Trace Specification](#E-Trace%5FSpecification)). | | Unconditional Jump | On RISC-V ISA all jump instructions are always unconditional, but these two words are always used together to avoid any confusions with the term 'branch' used by the IEEE-5001 Nexus Standard. The two main sub-categories of unconditional jumps that are relevant for tracing are: direct unconditional jump and indirect unconditional jump. | | Direct Conditional Branch | On RISC-V ISA all branch instructions are always direct and conditional (and also relative), but these three words are always used together to avoid confusion with the term 'branch' used by the IEEE-5001 Nexus Standard. | 7.1. N-Trace Messages (Details) ==================== ## [](#7-1-n-trace-messages-details)7.1\. N-Trace Messages (Details) This chapter provides a detailed description of all N-Trace messages. Overview of all fields in all messages is provided in the [Fields in Messages](#Fields in Messages) table. Common fields are described in the [Common Fields](#Common Fields) chapter, but fields specific to message **TCODE** values are explained here. Size of field in **Bits** column may be one or more of the following values: * **n (1..6)** \- This is an **n**\-bits wide, fixed-length field. * **Var** \- This is a variable-length, at least 1-bit wide field. * **Cfg** \- Size of this field depends on configuration setting (**Cfg** fields are always optional). Each message has its own table showing all fields in that message. | | The IEEE-5001 Nexus Standard presents tables with **TCODE** (which is sent first) in the last row. In contrast, this specification shows [Fields in Messages](#Fields in Messages) in the order they are sent (the first field sent is described first), aligning with the order of storage, processing, and text dumps. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#msg2%5FOwnership)7.1.1\. Ownership Message This message furnishes the requisite context (privileged mode and Context ID, as assigned by the operating system or hypervisor), enabling the decoder to correlate program flow with distinct code segments associated with various programs. Activation of this feature requires explicit enabling of the [trTeContext](#trTeContext) control bit. Reporting of this information occurs under one of the following three conditions: * Upon the retirement of an instruction that writes to the **scontext/hcontext** CSR (as reported via `priv` and `context` field on an ingress port). * In the event of a trap or trap return that results in a change in privilege mode (including **ECALL** and **EBREAK** instructions). * Following any trace [synchronizing message](#Synchronizing Messages). | | Should **hcontext** be implemented, the protocol requires two consecutive messages: the first presenting **hcontext** information and the second **scontext** information. This sequence is important for enabling the decoder to identify the code associated with a specific process. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | If tracing multiple OS-es, main decoder may route messages to an OS-specific decoder after seeing **hcontext** and the **scontext** (which follows) will be decoded by decoder determined by **hcontext**. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 1\. Ownership Message Fields__ | Bits | Name | Description | | ------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 6 | TCODE | Value=2(0x2). Standard Transfer Code ([TCODE](#field%5FTCODE)) field. | | Cfg | SRC | Standard Message Source ([SRC](#field%5FSRC)) field. | | Var | PROCESS | This is a variable-length field, which encodes **V** and **PRV** privilege mode bits as well as **scontext/hcontext** CSR values. Details are provided below. | | Var,Cfg | TSTAMP | Standard Timestamp ([TSTAMP](#field%5FTSTAMP)) field. | **Explanations and Notes** Field PROCESS is encoded as 4 sub-fields (FORMAT, PRV, V, CONTEXT). Bit layout is defined in RTL-like syntax as follows: PROCESS[x+5:0] = {CONTEXT[x:0], V[0], PRV[1:0], FORMAT[1:0]} __Table 2\. Encoding of PROCESS field (in LSB to MSB order)__ | Reason | FORMAT\[1:0\] | PRV\[1:0\] | V\[0\] | CONTEXT\[x:0\] | | --------------------------- | ------------- | ---------- | ------ | ------------------ | | V and/or PRV change | 00 | Yes | Yes | — | | Reserved | 01 | — | — | — | | Sync or **scontext** change | 10 | Yes | Yes | **scontext** value | | Sync or **hcontext** change | 11 | Yes | Yes | **hcontext** value | Encodings of **V/PRV** follow ISA privilege mode encodings and are encoded as follows: U-mode: V=0, PRV[1:0]=00 S-mode: V=0, PRV[1:0]=01 M-mode: V=0, PRV[1:0]=11 VU-mode: V=1, PRV[1:0]=00 VS-mode: V=1, PRV[1:0]=01 All unused encodings are reserved. Examples: PROCESS=0x3B2 = 0b11101_1_00_10 => scontext=0x1D,V=1,PRV[1:0]=00 (VU-mode) PROCESS=0xC 0b0_11_00 => V=0,PRV[1:0]=11 (M-mode) ### [](#msg2%5FDirectBranch)7.1.2\. DirectBranch Message It is applicable to [BTM](#mode%5FBTM) mode only. This message is generated when the taken direct conditional branch has retired. __Table 3\. Direct Branch Message Fields__ | Bits | Name | Description | | ------- | ------ | --------------------------------------------------------------------- | | 6 | TCODE | Value=3(0x3). Standard Transfer Code ([TCODE](#field%5FTCODE)) field. | | Cfg | SRC | Standard Message Source ([SRC](#field%5FSRC)) field. | | Var | I-CNT | Standard Instruction Count ([I-CNT](#field%5FI-CNT)) field. | | Var,Cfg | TSTAMP | Standard Timestamp ([TSTAMP](#field%5FTSTAMP)) field. | **Explanations and Notes** Last instruction in the code block (or blocks) with all inferable instructions (described by I-CNT) is a taken, direct conditional branch instruction. Next PC is determined by decoding the conditional branch instruction opcode to determine the encoded signed offset and adding it to the address of the conditional branch instruction. | | Not-taken direct conditional branches and direct unconditional jumps increment I-CNT but do not generate any trace. Direct unconditional jumps change the PC to the destination address of such a jump. The I-CNT enables determination of the PC of the last instruction in the code block(s). | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#msg2%5FIndirectBranch)7.1.3\. IndirectBranch Message It is applicable to [BTM](#mode%5FBTM) mode only. This message is generated under two conditions: * An instruction that causes an indirect unconditional control flow change has retired. * A trap due to an interrupt or exception is delivered. __Table 4\. Indirect Branch Message Fields__ | Bits | Name | Description | | ------- | ------ | --------------------------------------------------------------------- | | 6 | TCODE | Value=4(0x4). Standard Transfer Code ([TCODE](#field%5FTCODE)) field. | | Cfg | SRC | Standard Message Source ([SRC](#field%5FSRC)) field. | | 2 | B-TYPE | Standard Branch Type ([B-TYPE](#field%5FB-TYPE)) field. | | Var | I-CNT | Standard Instruction Count ([I-CNT](#field%5FI-CNT)) field. | | Var | U-ADDR | Standard Unique Address ([U-ADDR](#field%5FU-ADDR)) field. | | Var,Cfg | TSTAMP | Standard Timestamp ([TSTAMP](#field%5FTSTAMP)) field. | **Explanations and Notes** The last instruction within the code block(s), as specified by the I-CNT field, either represents an indirect unconditional control flow change (i.e., jump, call, or return) or this packet is generated in response to an exception or interrupt reported on the ingress port. The next PC is determined by applying the [Address Compression](#Address Compression) rules to the U-ADDR field present in this message. | | Not-taken conditional branches and direct unconditional jumps do not generate any trace. However, they do increase the I-CNT. Additionally, direct unconditional jumps modify the PC to the destination address specified in the instruction. Consequently, the PC of the last instruction in a code block(s) can be determined. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#msg2%5FError)7.1.4\. Error Message An error message must be generated in the event of an internal messages FIFO overflow, resulting in the loss of a trace message. __Table 5\. Error Message Fields__ | Bits | Name | Description | | ------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 6 | TCODE | Value=8(0x8). Standard Transfer Code ([TCODE](#field%5FTCODE)) field. | | Cfg | SRC | Standard Message Source ([SRC](#field%5FSRC)) field. | | 4 | ETYPE | Standard Error Type (a subset of IEEE-5001 Nexus Standard encoding): **0:** A FIFO overrun has resulted in the loss of one or more messages. **1..7:** Reserved. **8..15:** Designated for Vendor Defined Error(s). | | Var | ECODE | Standard Error Code (a subset of IEEE-5001 Nexus Standard encoding). A bit mask that when not equal to 0 may have one or more bits set as follows to indicate errors: **0:** Exact reason unknown/not provided. **xxxxxxx1:** Reserved. **xxxxxx1x:** Reserved (for data trace in future). **xxxxx1xx:** Program Trace Message(s) lost. **xxxx1xxx:** Ownership Trace Message(s) lost. **xxx1xxxx:** Reserved. **xx1xxxxx:** Reserved (for data trace in future). **x1xxxxxx:** Reserved. **1xxxxxxx:** Vendor Defined Message(s) lost. **IMPORTANT:** The field must be generated even if the reported value is always 0, to guarantee that the TSTAMP field aligns at the byte boundary. | | Var,Cfg | TSTAMP | Standard Timestamp ([TSTAMP](#field%5FTSTAMP)) field. | **Explanations and Notes** Error Message must be sent immediately prior to a [synchronizing message](#Synchronizing Messages) as soon as space is available in the Trace Encoder output queue. It is recommended that the timestamp reported in the message corresponds to the moment when the first trace message was dropped; however, this is not a requirement. | | This message **is required** as otherwise decoder (even though restart after FIFO overflow is signaled) would not be aware that trace was lost in case of the following sequence of events: Trace is turned off by trigger (or from any other reason). Message reporting 'trace off' event is lost (due to lack of space for it). Here Error Message should be generated (as soon as there is a room) Trace is never restarted. Trace is stopped (this will not generate any trace as trace is turned off). In the above case, Error Message will be the last message in trace stream. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#msg2%5FProgTraceSync)7.1.5\. ProgTraceSync Message __Table 6\. Program Trace Synchronization Message Fields__ | Bits | Name | Description | | ------- | ------ | --------------------------------------------------------------------- | | 6 | TCODE | Value=9(0x9). Standard Transfer Code ([TCODE](#field%5FTCODE)) field. | | Cfg | SRC | Standard Message Source ([SRC](#field%5FSRC)) field. | | 4 | SYNC | Standard Synchronization Reason ([SYNC](#field%5FSYNC)) field. | | Var | I-CNT | Standard Instruction Count ([I-CNT](#field%5FI-CNT)) field. | | Var | F-ADDR | Standard Full Address ([F-ADDR](#field%5FF-ADDR)) field. | | Var,Cfg | TSTAMP | Standard Timestamp ([TSTAMP](#field%5FTSTAMP)) field. | **Explanations and Notes** This message is produced at the start or restart of trace. In such instances, the I-CNT field is required to be set to 0\. However, under certain conditions associated with the SYNC parameter (e.g., `External Trace Trigger`), the I-CNT field may not be zero. Instead, it serves to pinpoint the precise Program Counter (PC) location at which the specified trigger or event occurred. Additionally, the F-ADDR field provides the complete PC address at the moment the trigger was activated. This message may be also generated on linear code for certain synchronization events as described in [Synchronizing Message](#Synchronizing Messages) chapter. ### [](#msg2%5FDirectBranchSync)7.1.6\. DirectBranchSync Message __Table 7\. Direct Branch with Sync Message Fields__ | Bits | Name | Description | | ------- | ------ | ---------------------------------------------------------------------- | | 6 | TCODE | Value=11(0xB). Standard Transfer Code ([TCODE](#field%5FTCODE)) field. | | Cfg | SRC | Standard Message Source ([SRC](#field%5FSRC)) field. | | 4 | SYNC | Standard Synchronization Reason ([SYNC](#field%5FSYNC)) field. | | Var | I-CNT | Standard Instruction Count ([I-CNT](#field%5FI-CNT)) field. | | Var | F-ADDR | Standard Full Address ([F-ADDR](#field%5FF-ADDR)) field. | | Var,Cfg | TSTAMP | Standard Timestamp ([TSTAMP](#field%5FTSTAMP)) field. | **Explanations and Notes** This message is produced under the same conditions as the [DirectBranch](#msg2%5FDirectBranch) message. However, it further includes details on the reason for synchronization via the SYNC field, as well as the full Program Counter (PC) address through the F-ADDR field. This message may be also generated on linear code for certain synchronization events as described in [Synchronizing Message](#Synchronizing Messages) chapter. ### [](#msg2%5FIndirectBranchSync)7.1.7\. IndirectBranchSync Message __Table 8\. Indirect Branch with Sync Message Fields__ | Bits | Name | Description | | ------- | ------ | ---------------------------------------------------------------------- | | 6 | TCODE | Value=12(0xC). Standard Transfer Code ([TCODE](#field%5FTCODE)) field. | | Cfg | SRC | Standard Message Source ([SRC](#field%5FSRC)) field. | | 4 | SYNC | Standard Synchronization Reason ([SYNC](#field%5FSYNC)) field. | | 2 | B-TYPE | Standard Branch Type ([B-TYPE](#field%5FB-TYPE)) field. | | Var | I-CNT | Standard Instruction Count ([I-CNT](#field%5FI-CNT)) field. | | Var | F-ADDR | Standard Full Address ([F-ADDR](#field%5FF-ADDR)) field. | | Var,Cfg | TSTAMP | Standard Timestamp ([TSTAMP](#field%5FTSTAMP)) field. | **Explanations and Notes** This message is generated in the same conditions as [IndirectBranch](#msg2%5FIndirectBranch) message, but additionally provides a reason for synchronization (SYNC field) and full PC (F-ADDR field). This message may be also generated (with [B-TYPE](#field%5FB-TYPE)\=0 field) on linear code for certain synchronization events as described in [Synchronizing Message](#Synchronizing Messages) chapter. ### [](#msg2%5FResourceFull)7.1.8\. ResourceFull Message This message is emitted when either the HIST register is full, or the I-CNT counter became full for a given encoder implementation. This mechanism ensures that no information is lost, as it enables the decoder to reconstruct larger I-CNT and HIST fields by concatenating or adding the emitted values. __Table 9\. Resource Full Message Fields__ | Bits | Name | Description | | ------- | ----------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 6 | TCODE | Value=27(0x1B). Standard Transfer Code ([TCODE](#field%5FTCODE)) field. | | Cfg | SRC | Standard Message Source ([SRC](#field%5FSRC)) field. | | 4 | RCODE | Standard Resource Code field (defines a meaning of RDATA fields). **0:** I-CNT counter has reached max value and is reported in the RDATA\[0\] field. See [I-CNT Details](#I-CNT Details) chapter. **1:** HIST field is full and is reported in the RDATA\[0\] field. See [HIST Field Full](#HIST Field Full) chapter for more details. **2**: **Extension:** HIST field is full and is repeated. RDATA\[0\] field holds HIST value and RDATA\[1\] field holds HREPEAT (History Repeat) value. This optional extension can be enabled via the [trTeInstEnRepeatedHistory](#trTeInstEnRepeatedHistory) control bit. **3..7:** Reserved for future encodings. **8..15:** Designated for vendor specific encodings. | | Var | RDATA \[0\] | Standard For RCODE=0, this is the I-CNT field. For RCODE=1 this is the HIST field (with most significant bit=1 being stop-bit). **Extension:** For RCODE=2 this is the HIST field (with most significant bit=1 being stop-bit). | | Var,Cfg | RDATA \[1\] | **Extension:** When RCODE=2 is reported this field includes HREPEAT (History Repeat) count. | | Var,Cfg | TSTAMP | Standard Timestamp ([TSTAMP](#field%5FTSTAMP)) field. | **Explanations and Notes** When RCODE is set to 1, this signifies that the HIST register is full and will not be repeated. Under these circumstances, the HIST field generally encapsulates the maximum number of history bits implemented within the HIST register. Nonetheless, implementations may opt to include any quantity of history bits in this field, with the range extending from a minimum of 2 bits up to the maximum defined by [NTRACE\_MAX\_HIST](#NTRACE%5FMAX%5FHIST) bits. Should the I-CNT counter and the HIST register simultaneously reach their respective capacity limits, it is mandatory to emit two successive ResourceFull messages. ### [](#msg2%5FIndirectBranchHist)7.1.9\. IndirectBranchHist Message __Table 10\. Indirect Branch History Message Fields__ | Bits | Name | Description | | ------- | ------ | ----------------------------------------------------------------------- | | 6 | TCODE | Value=28(0x1C). Standard Transfer Code ([TCODE](#field%5FTCODE)) field. | | Cfg | SRC | Standard Message Source ([SRC](#field%5FSRC)) field. | | 2 | B-TYPE | Standard Branch Type ([B-TYPE](#field%5FB-TYPE)) field. | | Var | I-CNT | Standard Instruction Count ([I-CNT](#field%5FI-CNT)) field. | | Var | U-ADDR | Standard Unique Address ([U-ADDR](#field%5FU-ADDR)) field. | | Var | HIST | Standard Branch History ([HIST](#field%5FHIST)) field. | | Var,Cfg | TSTAMP | Standard Timestamp ([TSTAMP](#field%5FTSTAMP)) field. | **Explanations and Notes** Last instruction in the code block (or blocks) (described by HIST and I-CNT fields) is indirect unconditional control flow change (jump, call, return) instruction or this message is generated when exception or interrupt is reported in the ingress port. See [HIST Field Generation](#HIST Field Generation) and [I-CNT Details](#I-CNT Details) chapters for clarifications. Next PC is determined by applying the [Address Compression](#Address Compression) rules using the U-ADDR field in this message. ### [](#msg2%5FIndirectBranchHistSync)7.1.10\. IndirectBranchHistSync Message __Table 11\. Indirect Branch History with Sync Message Fields__ | Bits | Name | Description | | ------- | ------ | ----------------------------------------------------------------------- | | 6 | TCODE | Value=29(0x1D). Standard Transfer Code ([TCODE](#field%5FTCODE)) field. | | Cfg | SRC | Standard Message Source ([SRC](#field%5FSRC)) field. | | 4 | SYNC | Standard Synchronization Reason ([SYNC](#field%5FSYNC)) field. | | 2 | B-TYPE | Standard Branch Type ([B-TYPE](#field%5FB-TYPE)) field. | | Var | I-CNT | Standard Instruction Count ([I-CNT](#field%5FI-CNT)) field. | | Var | F-ADDR | Standard Full Address ([F-ADDR](#field%5FF-ADDR)) field. | | Var | HIST | Standard Branch History ([HIST](#field%5FHIST)) field. | | Var,Cfg | TSTAMP | Standard Timestamp ([TSTAMP](#field%5FTSTAMP)) field. | **Explanations and Notes** This message is generated in the same conditions as [IndirectBranchHist](#msg2%5FIndirectBranchHist) message. However, it further includes details on the reason for synchronization via the SYNC field, as well as the full Program Counter (PC) address through the F-ADDR field. This message may be also generated (with [B-TYPE](#field%5FB-TYPE)\=0 field) on linear code for certain synchronization events as described in [Synchronizing Message](#Synchronizing Messages) chapter. ### [](#msg2%5FRepeatBranch)7.1.11\. RepeatBranch Message __Table 12\. Repeat Branch Message Fields__ | Bits | Name | Description | | ------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 6 | TCODE | Value=30(0x1E). Standard Transfer Code ([TCODE](#field%5FTCODE)) field. | | Cfg | SRC | Standard Message Source ([SRC](#field%5FSRC)) field. | | Var | B-CNT | Standard Branch Count field. Number of times the previous branch message (without a [SYNC](#field%5FSYNC) field) is repeated. Generated if I-CNT, HIST and target address is the same as in the previous branch message. | | Var,Cfg | TSTAMP | Standard Timestamp ([TSTAMP](#field%5FTSTAMP)) field. | **Explanations and Notes** This message is reported when an identical (direct or indirect) branch message is encountered (just to save trace bandwidth). Trace decoder should just repeat handling of previous branch message B-CNT times. ### [](#msg2%5FProgTraceCorrelation)7.1.12\. ProgTraceCorrelation Message This message is emitted when the trace is disabled or stopped. __Table 13\. Program Trace Correlation Message Fields__ | Bits | Name | Description | | ------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 6 | TCODE | Value=33(0x21). Standard Transfer Code ([TCODE](#field%5FTCODE)) field. | | Cfg | SRC | Standard Message Source ([SRC](#field%5FSRC)) field. | | 4 | EVCODE | Standard Reason to generate Program Correlation: **0:** Entry into Debug Mode. Required (do not send 4 instead!). **1:** Entry into Low-power Mode. Optional. **2..3:** Reserved for data trace. **4:** Program Trace Disabled (hart may be still running). Optional. **5..7:** Reserved for future extensions of N-Trace specification. **8..15:** Designated for vendor specific encodings. | | 2 | CDF | Standard number of CDATA fields following it: **0:** Only I-CNT field follows and there is no HIST field. **1:** I-CNT field and single CDATA (HIST) field (for HTM trace). **2..3:** Reserved for future extensions of N-Trace specification.In BTM trace mode CDF must be 0\. In HTM trace mode CDF must be 1 (even if HIST field is empty, encoded as 0x1). | | Var | I-CNT | Standard Instruction Count ([I-CNT](#field%5FI-CNT)) field. | | Var,Cfg | HIST | Standard Branch History ([HIST](#field%5FHIST)) field. **This field must be present in HTM mode**, so decoder does not need to read CDF to determine its existence. | | Var,Cfg | TSTAMP | Standard Timestamp ([TSTAMP](#field%5FTSTAMP)) field. | **Explanations and Notes** It provides a reason (in EVCODE field) plus I-CNT and HIST fields, which allows the decoder to determine the PC where an execution or the trace stopped. This message includes the EVCODE field, which specifies the reason for generating this message, alongside the I-CNT and HIST fields. These fields collectively enable the decoder to accurately identify the PC location where execution or tracing was halted. 6.1. N-Trace Messages (Overview) ==================== ## [](#6-1-n-trace-messages-overview)6.1\. N-Trace Messages (Overview) | | The terminology Indirect Branch as used by the IEEE-5001 Nexus Standard may lead to confusion, given that the RISC-V ISA exclusively permits direct conditional branches, which are always relative. Furthermore, the RISC-V ISA makes a distinction between 'jump' (unconditional flow change) and 'branch' (conditional flow change), a differentiation not observed in Nexus terminology, where any flow change, including exceptions and interrupts, is uniformly referred to as a 'branch'. This specification employs the terms 'branch' and 'jump' as defined by RISC-V ISA. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#6-1-1-fields-in-messages)6.1.1\. Fields in Messages The table presented below enumerates all message types that can be generated, with each row comprehensively defining the fields associated with a particular message type. Fields that are present in different messages are consistently ordered. Message field attributes are described using the following terminology: * **\[n\]**: A fixed-length field that is **n** bits wide. * **\[Var\]**: A variable-length, non-empty (at least 1-bit wide), field. * **\[Cfg\]**: A configurable field, where the existence and size depend on the encoder configuration options. __Table 1\. Fields in Messages__ | Message ID/Field \[size\] | [TCODE](#field%5FTCODE) \[6\] | [SRC](#field%5FSRC) \[Cfg\] | [SYNC](#field%5FSYNC) \[4\] | [B-TYPE](#field%5FB-TYPE) \[2\] | Other fields | [I-CNT](#field%5FI-CNT) \[Var\] | [x-ADDR](#Address Compression) \[Var\] | [HIST](#field%5FHIST) \[Var\] | [TSTAMP](#field%5FTSTAMP) \[Var,Cfg\] | | -------------------------------------------------------- | ----------------------------- | --------------------------- | ------------------------------------------------------------------------ | ------------------------------- | ------------------------- | ------------------------------- | -------------------------------------- | ----------------------------- | ------------------------------------- | | [Ownership](#msg2%5FOwnership) | 2 | Cfg | [PROCESS](#field%5FPROCESS) **\[Var\]** | Cfg | | | | | | | [DirectBranch](#msg2%5FDirectBranch) | 3 | Cfg | Yes | Cfg | | | | | | | [IndirectBranch](#msg2%5FIndirectBranch) | 4 | Cfg | Yes | Yes | [U-ADDR](#field%5FU-ADDR) | Cfg | | | | | [Error](#msg2%5FError) | 8 | Cfg | [ETYPE](#field%5FETYPE) **\[4\]** \+ [ECODE](#field%5FECODE) **\[Var\]** | Cfg | | | | | | | [ProgTraceSync](#msg2%5FProgTraceSync) | 9 | Cfg | Yes | Yes | [F-ADDR](#field%5FF-ADDR) | Cfg | | | | | [DirectBranchSync](#msg2%5FDirectBranchSync) | 11 | Cfg | Yes | Yes | [F-ADDR](#field%5FF-ADDR) | Cfg | | | | | [IndirectBranchSync](#msg2%5FIndirectBranchSync) | 12 | Cfg | Yes | Yes | Yes | [F-ADDR](#field%5FF-ADDR) | Cfg | | | | [ResourceFull](#msg2%5FResourceFull) | 27 | Cfg | [RCODE](#field%5FRCODE) **\[4\]** \+ [RDATA](#field%5FRDATA) **\[Var\]** | Cfg | | | | | | | [IndirectBranchHist](#msg2%5FIndirectBranchHist) | 28 | Cfg | Yes | Yes | [U-ADDR](#field%5FU-ADDR) | Yes | Cfg | | | | [IndirectBranchHistSync](#msg2%5FIndirectBranchHistSync) | 29 | Cfg | Yes | Yes | Yes | [F-ADDR](#field%5FF-ADDR) | Yes | Cfg | | | [RepeatBranch](#msg2%5FRepeatBranch) | 30 | Cfg | [B-CNT](#field%5FB-CNT) **\[Var\]** | Cfg | | | | | | | [ProgTraceCorrelation](#msg2%5FProgTraceCorrelation) | 33 | Cfg | [EVCODE](#field%5FEVCODE) **\[4\]** \+ [CDF](#field%5FCDF) **\[2\]** | Yes | **Cfg** | Cfg | | | | | [Vendor Defined](#msg%5Fother) | 56..62 | Cfg | Designated for use by Vendor Defined messages | | | | | | | | [Reserved](#msg%5Fother) | other | Cfg | Reserved for future extensions of N-Trace specification | | | | | | | | | Any message may include the optional [TSTAMP](#field%5FTSTAMP) **\[Var,Cfg\]** field as the very last field of a message. It must be enabled by [trTsEnable](#trTsEnable) control bit. Timestamp field always starts at byte-boundary (as it is always preceded by variable-length field). See [Timestamp Reporting](#Timestamp Reporting) chapter for more details. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Messages marked as **Reserved** or **Vendor Defined** should be ignored by decoders interested in program flow only. However, decoders should provide an option to display/dump them and/or generate a warning as such a message may be seen when trace capture is corrupted.**Vendor Defined** messages can be used for prototyping, debugging, validation and maintenance purposes. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Reference code header [NexRvMsg.h](https://github.com/riscv-non-isa/tg-nexus-trace/blob/main/refcode/c/NexRvMsg.h) defines all N-Trace messages in machine-readable format. Here is small snipped from this file as an example: ```c // Naming: // NEXM=Nexus Message, BEG/END=Beginning/End of definition. // SRC=Message source (system-field). Name of an option given. // FLD/VAR=Fixed/variable size field. // ADR=Special case of variable field (without least significant bit). // CFG=Configurable, Name of an option given. NEXM_BEG(IndirectBranchSync, 12) NEXM_SRC(SrcBits) // Configurable NEXM_FLD(SYNC, 4) NEXM_FLD(BTYPE, 2) NEXM_VAR(ICNT) NEXM_ADR(FADDR) NEXM_VAR(TSTAMP) NEXM_END() NEXM_BEG(ResourceFull, 27) NEXM_SRC(SrcBits) // Configurable NEXM_FLD(RCODE, 4) NEXM_VAR(RDATA) NEXM_VAR_CFG(HREPEAT, EnaRepeatedHistory) // Configurable NEXM_VAR(TSTAMP) NEXM_END() NEXM_BEG(IndirectBranchHist, 28) NEXM_SRC(SrcBits) // Configurable NEXM_FLD(BTYPE, 2) NEXM_VAR(ICNT) NEXM_ADR(UADDR) NEXM_VAR(HIST) NEXM_VAR(TSTAMP) NEXM_END() ``` | | Reference code is using plain C-style identifiers, so the field name as **B-TYPE** will become **BTYPE**. | | ------------------------------------------------------------------------------------------------------------ | ### [](#6-1-2-common-fields)6.1.2\. Common Fields The table below provides details for fields which are used in more than one message type. Fields which are present in only one message are described with each message. __Table 2\. Details of Common Fields__ | Name | Bits | Description | Values/Notes | | -------------------------------- | ------- | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Fields used in many messages** | | | | | TCODE | 6 | Transfer Code | Message header that identifies the number and/or size of fields to be transferred, and how to interpret each of the fields following it. | | SRC | **Cfg** | Source of Message Transmission | Width of SRC field is defined by [trTeSrcBits](#trTeSrcBits) control field and it may be enabled/disabled by [trTeInhibitSrc](#trTeInhibitSrc) control bit. This optional field is used to identify the source of the message transmission. In configurations that comprise only a single hart, this field need not be transmitted. For devices that comprise multiple harts, this field must be transmitted (if enabled) as part of the message to identify the source of the message transmission. The transmitted SRC field size should be the same for all enabled trace encoders sharing a trace stream. | | SYNC | 4 | Reason for Synchronization | Encodings and details are provided in the [Synchronizing Messages](#Synchronizing Messages) chapter. **NOTE:** The SYNC field is always sent together with the [F-ADDR](#field%5FF-ADDR) field, so decoding may start from a message containing the SYNC field. | | B-TYPE | 2 | Branch Type | Reason for indirect flow change: **0:** Indirect control flow change (jump, call or return) or in a [synchronizing message](#Synchronizing Messages) not related to code execution. **1:** Exception or interrupt (if the encoder is not capable of reporting 2 and 3). **2:** **Extension:** Exception **3:** **Extension:** Interrupt **NOTE:** Either 1-only or both 2 and 3 should be implemented and consistently reported. Extended values 2 and 3 allow trace tools to distinguish exceptions and interrupts easily. | | I-CNT | **Var** | Instruction Count | As RISC-V allows variable-length instructions, this is the number of 16-bit (INST\_LEN/2) instruction units executed/retired since the I-CNT counter was transmitted or reset. See [I-CNT Details](#I-CNT Details) chapter for more details. | | F-ADDR | **Var** | Full Target Address | Full PC without the least significant bit. The least significant bit is not reported as it is always 0\. See [Address Compression](#Address Compression) chapter for more details. **NOTE:** The F-ADDR field is always sent together with the [SYNC](#field%5FSYNC) field. | | U-ADDR | **Var** | Unique part of Target Address | Unique part of PC address (XOR with recently reported address). See [Address Compression](#Address Compression) chapter for more details.The U-ADDR field is always sent together with the [B-TYPE](#field%5FB-TYPE) field. | | HIST | **Var** | Direct Branch History map | Most significant bit (always 1) serves as a 'stop-bit', the least significant bit denotes the last direct conditional branch. See [HIST Field Generation](#HIST Field Generation) chapter for more details. | | TSTAMP | **Var** | Timestamp (optional) | It must be enabled by [trTsEnable](#trTsEnable) control bit. See [Timestamp Reporting](#Timestamp Reporting) chapter for more details. | IEEE-5001 Nexus Standard does not define limits for variable-length fields, but N-Trace provides some limits. It will help to write efficient decoding software but is not limiting hardware in any way. __Table 3\. Maximum Field Sizes__ | Field | Symbol | Bits | Description | | -------------- | -------------------- | ---- | ------------------------------------------------------------------------------------------------------------------------------------ | | SRC | NTRACE\_MAX\_SRC | 12 | Determined by size of Trace Control register field. Enough for 4096 (4K) trace sources. | | I-CNT | NTRACE\_MAX\_ICNT | 22 | Usually a smaller value will be sufficient. An overflow bit may be used for efficient I-CNT full detection. | | F-ADDR, U-ADDR | NTRACE\_MAX\_ADDR | 63 | Only 63 bits suffice as the least significant bit of an instruction address is always 0 and does not need to be reported. | | HIST | NTRACE\_MAX\_HIST | 32 | It includes stop-bit. This size is optimal for not wasting any bits in very often used [ResourceFull](#msg%5FResourceFull) messages. | | TSTAMP | NTRACE\_MAX\_TSTAMP | 64 | It is certainly big enough. It corresponds to architecture defined timer and cycle count registers. | | HREPEAT | NTRACE\_MAX\_HREPEAT | 18 | Assure some trace is periodically generated for very long loops. | | B-CNT | NTRACE\_MAX\_BCNT | 18 | Assure some trace is periodically generated for very long loops. | 5.1. Main N-Trace Trace Modes ==================== ## [](#5-1-main-n-trace-trace-modes)5.1\. Main N-Trace Trace Modes RISC-V N-Trace defines two instruction trace modes: * **Branch Trace Messaging (BTM)** \- each taken direct conditional branch generates a minimum two-byte message. However, repeated branches can be aggregated and reported as a single message with a count, rather than numerous identical messages. * **History Trace Messaging (HTM)** \- every direct conditional branch, whether taken or not-taken, contributes a single bit to the history buffer, significantly enhancing the trace efficiency. The encoder is required to implement at least one of these modes. Both may be supported, but is not required. | | Above modes correspond to the following IEEE-5001 Nexus Standard instruction trace modes: **Branch Trace Messaging using Traditional Messages** **Branch Trace Messaging using Branch History Messages** | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The IEEE-5001 Nexus Standard defines different conformance levels. These levels are not directly applicable to N-Trace as Nexus levels always include debug levels. Different N-Trace options are provided in [N-Trace Specific Trace Controls](#N-Trace Specific Trace Controls) chapter. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 12.1. IEEE-5001 Nexus Standard Compliance ==================== ## [](#12-1-ieee-5001-nexus-standard-compliance)12.1\. IEEE-5001 Nexus Standard Compliance The IEEE-5001 Nexus Standard provides a lot of flexibility and in general N-Trace can be considered fully compatible. There is one incompatible, small change: * Field [ECODE](#field%5FECODE) is variable-length field (to assure TSTAMP field is on byte boundary). Several compatible extensions are described in preceding chapters and are marked with **Extension:** marker. Each of them is disabled by default and must be directly enabled. 9.1. Optimization Extensions ==================== ## [](#9-1-optimization-extensions)9.1\. Optimization Extensions N-Trace messages are defined as a strict subset of IEEE-5001 Nexus Standard messages. However, to provide better compression some optional extensions are defined. Each of them should be by default disabled and specifically enabled to allow simpler decoder to decode non fully optimized trace. Table [Details\_Control\_Parameters](#Details%5FControl%5FParameters) describes all control bits to enable these optimizations. ### [](#9-1-1-sequential-jump-optimization)9.1.1\. Sequential Jump Optimization This optimization must be enabled by [trTeInstEnSequentialJump](#trTeInstEnSequentialJump) control bit. By default, the target of an indirect unconditional jump is always considered an uninferable PC discontinuity. However, if the register that specifies the jump target was loaded with a constant then it can be considered inferable under some circumstances. The hart must identify indirect unconditional jumps with sequentially inferable targets and provide this information separately to the encoder. The final decision as to whether to treat the indirect unconditional jump as inferable or not must be made by the encoder. Both the constant load and the indirect unconditional jump must be traced as consecutive instructions in the same message for the decoder to be able to infer the indirect unconditional jump target. Some jump targets that are supplied via: * an **LUI** or **C.LUI** (a register which contains a constant), or * an **AUIPC** (a register which contains a constant offset from the PC). Such indirect unconditional jump targets are classified as sequentially inferable if the pair of instructions are retired consecutively (i.e. the **AUIPC**, **LUI** or **C.LUI** immediately precedes the indirect unconditional jump). When decoder is processing instructions (always forward) it must encounter the **AUIPC**, **LUI** or **C.LUI** immediately directly before **JR** and then calculate target address of a jump. I-CNT in that message must span over both (consecutive) instruction. | | The restriction that the instructions must be retired consecutively is necessary to minimize the additional signals needed between the hart and the encoder, and should have a minimal impact on trace efficiency as it is anticipated that consecutive execution will be the norm. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#9-1-2-implicit-return-optimization)9.1.2\. Implicit Return Optimization This optimization must be enabled by the [trTeInstImplicitReturnMode](#trTeInstImplicitReturnMode) control field different than 0. Although a function return is usually an indirect unconditional jump, most programs return to the point in the program from which the function was called using a standard calling convention. For those programs, it is possible to determine the execution path without being explicitly notified of the destination addresses of the returns. The implicit return mode can result in very significant improvements in trace encoder efficiency. Returns can only be treated as inferable if the associated call has already been reported in an earlier message. The encoder must ensure that this is the case. There are 3 possible ways of handling return address stack (values of [trTeInstImplicitReturnMode](#trTeInstImplicitReturnMode) control field): **Simple counting ([trTeInstImplicitReturnMode](#trTeInstImplicitReturnMode)\=1)** This can be accomplished by utilizing a counter to keep track of the number of nested calls being traced. The counter increments on calls and decrements on returns. The counter will not over or underflow, and is reset to 0 whenever a synchronizing message is sent. Returns will be treated as inferable and will not generate a trace message if the count is non-zero (i.e. the associated call was already reported in an earlier message). Such a scheme is low cost, and will work as long as programs are "well behaved". The encoder will not be able to check that the return address is that of the instruction following the associated call. As such, any program that modifies return addresses cannot be traced using this mode with this minimal implementation. Due to these limitations **this is NOT recommended implementation**. **Stack with Full Addresses ([trTeInstImplicitReturnMode](#trTeInstImplicitReturnMode)\=3)** The encoder maintains a stack of expected return addresses (created when call is encountered), and only treat a return as inferable if the actual return address matches the value on the stack. This is fully robust for all programs but is more expensive to implement. In this case, if a return address does not match the prediction, it must be reported explicitly via a message. This ensures that the decoder can determine which return is being reported. This method may use shadow stack if implemented by the core. **Stack with Partial Addresses ([trTeInstImplicitReturnMode](#trTeInstImplicitReturnMode)\=2)** Call stack maintained by encoder may not include all addresses, but only keep some least significant part of it and use them to compare if return is matching the call or not. Changes that program making incorrect return will return to address with the same least significant portion are very slim. | | Decoder does not need to know what actual depth of the call stack is implemented by encoder but for efficiency reasons it should assume max depth. N-Trace implementation should never implement call stack deeper than 32 levels. Such deep calls will be most likely interrupted by other events/messages (like periodic SYNC). | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#9-1-3-repeated-history-optimization)9.1.3\. Repeated History Optimization This optimization must be enabled by the [trTeInstEnRepeatedHistory](#trTeInstEnRepeatedHistory) control bit. A typical loop either has a direct conditional branch at the start of a loop (which must be typically 'taken' to terminate the loop) or has a direct conditional branch at the end of the loop (which must be typically taken to repeat the loop). In the first case, the direct conditional branch is not taken most of the time and taken once at the end. In the second case, the direct conditional branch is taken most of the time, but not taken at the end of the loop. Loops with many iterations such as those in functions like memcpy/strcpy have identical flow in each iteraction. Instead of sending the same history bits many times, repeated patterns can be detected and counted. This is a big saving! As an example, a memcpy of 4MB buffer using 32-bit transfers will execute at least 1M of direct conditional branches and 1M of history bits must be included in trace (it is a lot of trace). The IEEE-5001 Nexus Standard defines a [Repeat Branch](#msg%5FRepeatBranch) message. This message will provide a single [B-CNT](#field%5FB-CNT) (Branch Count) field instead of generating many identical [Direct Branch](#msg%5FDirectBranch) messages. But this message cannot be used in [HTM](#mode%5FHTM) mode as repeated messages (Direct Branch) do not include the HIST field. To allow generation of repeated history of direct conditional branches in HTM mode an extra encoding for [RCODE](#field%5FRCODE)\=2 in [Resource Full](#msg%5FResourceFull) message is added. | | It is allowed to generate any sequence of [Resource Full](#msg%5FResourceFull) messages as long as the logically concatenated sequence of (repeated or not …​) HIST bits (excluding most significant stop-bit\[s\]) is the same. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Tracing of such simple, long loops would benefit from generating special messages/fields which provide counters of taken/not-taken direct conditional branches (in a way like [Repeat Branch](#msg%5FRepeatBranch) message). But this approach will not work with more complex code with a conditional statement (or several of them) inside of a loop. In such a case, it is desired to detect repeated sequences of taken/not-taken direct conditional branches and instead generate many messages with HIST fields, generate a message consisting of a HIST pattern and repeat count. Let’s assume that we have a loop, which generates a long sequence of repeated taken/not-taken direct conditional branches. Trace may generate [Resource Full](#msg%5FResourceFull) messages with the following HIST records: Msg#1: TCODE=27 (ResourceFull) RCODE=1 (full HIST record is provided as RDATA) RDATA=0b1_01_0101_0101_0101_0101_0101_0101_0101 = 0x55555555 (stop-bit + pattern 01 repeated 15 times) Msg#2: TCODE=27 (ResourceFull) RCODE=1 (full HIST record is provided as RDATA) RDATA=0b1_01_0101_0101_0101_0101_0101_0101_0101 = 0x55555555 (stop-bit + pattern 01 repeated 15 times) ... Msg#10: TCODE=27 (ResourceFull) RCODE=1 (full HIST record is provided as RDATA) RDATA=0b1_01_0101_0101_0101_0101_0101_0101_0101 = 0x55555555 (stop-bit + pattern 01 repeated 15 times) Instead of generating many messages with identical HIST record, encoder can detect repeated pattern and generate the following single message: Msg#1: TCODE=27 (ResourceFull) RCODE=2 (full HIST record is provided as RDATA and repeat count is provided as HREPEAT field) RDATA=0b1_01_0101_0101_0101_0101_0101_0101_0101 = 0x55555555 (stop-bit + pattern 01 repeated 15 times) HREPEAT=10 (Repeat Count=10 instead 10 messages) Above example shows a 2-bit pattern, but using the same technique it can be expanded to any size of pattern. The exact way to detect these patterns is not specified as it does not change encoding of messages. So, it is possible to generate the following, a bit smaller, message: Msg#1: TCODE=27 (ResourceFull) RCODE=2 (full HIST record is provided as RDATA and repeat count is provided as HREPEAT field) RDATA=0b1_01 = 0x5 (stop-bit + single pattern 01) HREPEAT=150 (Repeat Count is bigger, but pattern is smaller) | | This type of compression (reporting shorter patterns and larger counts) may not be practical as it may save only a little. Trace is compressed a lot already and it really should not matter if we report 150 iterations of a loop in 6 or 7 bytes. Example above is provided to assure that trace encoders must handle this type of trace compression. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | When number of repeated branches is bigger than max HREPEAT counter value then several consecutive messages with max HREPEAT value should be generated. Total count represented by all these messages (sum of all HREPEAT fields) will be a number of repeated branch history message. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | HREPEAT counter should not have too many bits as it is not desired to not generate any trace messages for longer periods of time. Bigger HREPEAT will not make compression better but will produce timestamp rarely. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#9-1-4-virtual-addresses-optimization)9.1.4\. Virtual Addresses Optimization This optimization must be enabled by [trTeInstExtendAddrMSB](#trTeInstExtendAddrMSB) control bit. | | Normally (without the above bit enabled or implemented), addresses with many most significant bits set to 1 will be sent as long messages (as variable size fields skip only the most significant 0-s). An address,**0xFFFF\_FFFF\_8000\_31F4**, a real address from the Linux kernel, will be encoded as F-ADDR = **0x7FFF\_FFFF\_C000\_18FA** (with the least significant 0-bit skipped). Such a 63-bit variable field value will require 11 bytes to be sent (as we have 6 MDO bits in each byte). | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The following additional rules are used when [trTeInstExtendAddrMSB](#trTeInstExtendAddrMSB) control bit is implemented and set: * The encoder may skip any number of most significant identical bits in the U-ADDR/F-ADDR fields. However, it must ensure that if any bits are skipped, then the number of transmitted bits is a multiple of the MDO size. Additionally, the most significant transmitted bit must have the same value as the skipped bits. * If F-ADDR/U-ADDR field is received by decoder, then the last (most significant) bit of the very last MDO record must be extended up to bit#63 or bit#31 (depending on XLEN of the core). It is like sign-extension, but it is NOT a sign bit. * This method does NOT require a trace decoder to know what a virtual memory system mode is or if an address is physical or virtual. The decoder must look at the most significant bit of the last MDO in F-ADDR/U-ADDR field and either extend or not. * Simple implementations may not implement an enable bit and always send full address. Benefits of using it on 32-bit cores is small, so it may not be implemented. This way of encoding allows an encoder to efficiently send: * Any physical address. * Any virtual address (in any mode). * Any illegal address. Trace encoder must implement a most significant bit detection (skipping identical 1-s or 0-s in addition to skipping identical 0-s as for any other variable size field) while sending F-ADDR/U-ADDR field. Trace decoders must do it in reverse order, which means that a sign extension (if needed) must be done after collecting the last MDO bit in an F-ADDR/U-ADDR field. Calculation of full address (as defined in [Address Compression](#Address Compression) chapter above) must be done after sign extension of U-ADDR field. **Example Encodings** **Non-extended address (most significant MDO bit = 0)** MDO_MSEO #byte: 543210 <- MDO bit index (bit#5 is most significant bit) ------------------- #0: 111111_00 #1: 111111_00 #2: 111111_00 #3: 111111_00 #4: 111111_00 #5: 011111_01 <- Last MDO+MSO byte. Most significant bit #5 is 0, so NO extension. F-ADDR field=0x7_FFFF_FFFF, Encoded address=0xF_FFFF_FFFE **Extended address (most significant MDO bit = 1)** MDO_MSEO #byte: 543210 <- MDO bit index (bit#5 is most significant bit) ------------------- #0: 111111_00 #1: 111111_00 #2: 111111_00 #3: 111111_00 #4: 011111_00 #5: 111100_01 <- Last MDO+MSEO byte. Most significant bit #5 is 1, so WITH extension. F-ADDR field=0xF_1FFF_FFFF, Encoded address=0xFFFF_FFFE_3FFF_FFFE **Non-extended address (extra MDO with all 0-s prevents extension)** MDO_MSEO #byte: 543210 <- MDO bit index (bit#5 is most significant bit) ------------------- #0: 111111_00 #1: 111111_00 #2: 111111_00 #3: 111111_00 #4: 111111_00 #5: 111111_00 #6: 000000_01 <- Last MDO+MSEO byte. Most significant bit #5 is 0, so NO extension. F-ADDR field=0xF_FFFF_FFFF, Encoded address=0x1F_FFFF_FFFE **Non-extended full 64-bit address (invalid address)** MDO_MSEO #byte: 543210 <- MDO bit index (bit#5 is most significant bit) ------------------- #0: 111111_00 #1: 111111_00 #2: 111111_00 #3: 111111_00 #4: 111111_00 #5: 111111_00 #6: 111111_00 #7: 111111_00 #8: 111111_00 #9: 111111_00 #10: 000101_01 <- Last MDO+MSEO byte. Most significant bit #5 is 0, so NO extension. F-ADDR field=0x5FFF_FFFF_FFFF_FFFF, Encoded address=0xBFFF_FFFF_FFFF_FFFE | | Address **0xBFFF\_FFFF\_FFFF\_FFFF** is NOT a legal address in any RISC-V virtual memory modes as it does not have all most significant bits identical. But such an address may be encountered as result of a bug and as such should be reported. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 3.1. N-Trace Transmission Protocol ==================== ## [](#3-1-n-trace-transmission-protocol)3.1\. N-Trace Transmission Protocol The IEEE-5001 Nexus Standard defines a trace messaging protocol using several **MDO** (Message Data Out) signals and one or two flag signals known as **MSEO** (Message Start/End Out). A Nexus message is sent or stored in a record composed of **MDO** and **MSEO**. N-Trace specification defines 6-bit **MDO** and 2-bit **MSEO** so both fit in a single byte. * It allows easy storage in memory as well as sending using 1-bit/ 2-bit/ 4-bit/ 8-bit/ 16-bit parallel transport (which is supported by many existing trace probes and connectors). * Decoding software may work on bytes and 32-bit/64-bit words and expect MSEO bits at two least significant bits of each byte. N-Trace message transmission protocol is a strict subset of IEEE-5001 Nexus Standard trace messaging protocol. __Table 1\. N-Trace subset__ | Protocol Feature | Nexus Standard | N-Trace (strict subset of Nexus) | | ----------------------------- | -------------- | ------------------------------------------------------------------- | | Number of **MSEO** bits | 1 or 2 | 2 | | Number of **MDO** bits | At least 1 | 6 | | Total (**MDO**+**MSEO**) bits | At least 2 | 8 (one byte) | | Order (transmitted or stored) | Vendor defined | **MSEO** before **MDO**, least significant bit for each field first | | Max field size | Not specified | 64 bits (some 32 bits or less) | | Max standard message size | Not specified | 38 bytes (maximum sum of all fields) | The maximum standard message size of 38 bytes in this version of the specification is to transmit [IndirectBranchHistSync](#msg%5FIndirectBranchHistSync) message which includes TCODE/ SRC/ SYNC/ B-TYPE(5 bytes total), I-CNT(30 bits, 5 bytes), F-ADDR(63 bits, 11 bytes), HIST(32 bits, 6 bytes) TSTAMP(64 bits, 11 bytes). While implementations may have a shorter maximum message size (e.g. due to I-CNT being smaller), they must assure that the internal FIFOs are designed to hold at least two maximum sized messages that the implementation can produce. While decoding software may be designed to avoid dynamic memory allocation, it must nonetheless be robust enough to handle messages of any size. This is to account for scenarios when a trace memory could be corrupted, such as a trace consisting entirely of zeros, which could be interpreted as an unusually long variable-length field. Custom messages and fields may carry different payloads and may be larger than 64 bits and 38 bytes. ### [](#3-1-1-mseo-sequences)3.1.1\. MSEO Sequences **MSEO\[1:0\]** bits (located in the least significant bits of each byte) are defined by the follow rules: * The first byte of a message sends the least significant bits of the message and is indicated by **MSEO\[1:0\]=00**. * Bytes occupied by fixed-length fields are sent using **MSEO\[1:0\]=00**. * The last byte of a variable-length field, that is not last byte of a message, is indicated by **MSEO\[1:0\]=01**. * A variable-length field in a message always ends on a byte boundary (zero extended as needed). * The non-last bytes of a variable-length fields are indicated by **MSEO\[1:0\]=00**. * The last byte of a message is indicated by **MSEO\[1:0\]=11**. * It also implies an end of the last (fixed-length or variable-length) field of a message. * Idle bytes (between messages or used as padding) are indicated by **MSEO\[1:0\]=11** and **MDO\[5:0\]=111111** (entire byte is **0xFF**). * Value of **MSEO\[1:0\]=10** is reserved for future extensions. The table below provides possible sequences of **MSEO\[1:0\]** bits (to expand above rules - **highlighted** MSEO represent the actual function): __Table 2\. Transitions of MSEO Bits__ | MSEO Function | Previous-**Current** MSEO\[1:0\] Sequence | | ---------------------------- | ----------------------------------------- | | Start of message | 11-**00** | | Middle of field | 00 (or 01)-**00** | | End of variable-length field | 00 (or 01)-**01** | | End of message | 00 (or 01)-**11** | | Idle (no message) | 11-**11** | | Reserved | 11-**01** | | Reserved | any-**10** | | | Original IEEE-5001 Nexus Standard defines the MSEO protocol as follows: Two 1\-s followed by one 0 indicates the start of a message. 0 followed by two or more 1\-s indicates the end of a message. 0 followed by 1 followed by 0 indicates the end of a variable-length field. 0\-s at all other clocks during transmission of a message. 1\-s at all clocks during no message transmission (idle). Dual MSEO protocol (utilized by this N-Trace specification) is a two-pin mode of this general (single and dual) MSEO protocol definition. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#3-1-2-unified-n-trace-message-structure)3.1.2\. Unified N-Trace Message Structure Each N-Trace message has identical structure (100% compatible with IEEE-5001 Nexus Standard): * Very first field is always fixed-length **TCODE** (Transport Code) which defines the meaning and format of subsequent fields. * In case of simultaneous tracing from more than one hart, the second field is always fixed-length **SRC** (Message Source) field, which provides a unique ID of message source. * This field allows trace decoders to separate messages from different trace sources (Trace Encoders, harts) without knowing any details of each of the messages. * This method can be used to handle different (opaque) trace or debug or performance data using N-Trace transport/storage/export infrastructure. * One or more (fixed-length or variable-length) payload fields. Sequence and selection of these fields depend on the value of **TCODE** field. * In some rare cases one of preceding fields may define number of following fields. * Very last field is (optional) variable-length **TSTAMP** (Timestamp) field. * It may be possible to generate and analyze timestamps in a unified (simpler) way. ### [](#3-1-3-order-of-bits-and-bytes)3.1.3\. Order of bits and bytes Order of bits and bytes: * Trace messages/packets are considered as sequences of bytes and are always transmitted with least significant bits/bytes first. * IEEE-5001 Nexus Standard MSEO bits are transmitted on the least significant part and bit#0 first. * Idle state must be transmitted as all 1s MSEO and MDO bits. * For transmission on a 16bit interface (e.g. PIB 16-bit mode), the first byte of message/packet is transmitted on the least significant part and the MSEO of the second/odd byte is transmitted on bits #8-#9 and MDO on bits #10-#15. | | Above rules allow receiving trace probes to skip idle messages. | | ------------------------------------------------------------------ | ### [](#3-1-4-pib-idle-cycles-explained)3.1.4\. PIB Idle Cycles Explained This chapter describes N-Trace specific details about the transmission via a Pin Interface Block (PIB), as it is described in the [RISC-V Trace Control Interface](#RISC-V%5FTrace%5FControl%5FInterface) Specification. Trace messages may start on any (positive or negative) edge of trace clock. | | Once a message is started all bits of that message must be transmitted on consecutive trace clock edges (both positive and negative). | | ---------------------------------------------------------------------------------------------------------------------------------------- | Said so, an idle sequence may be sent using any number of trace clock edges (positive or negative). To explain this let’s assume the following serially transmitted (in 1-bit PIB mode) sequences of bits (MSEO\[0\] bit being first on the left): * < `11` DDDDDD> - 8 bits in a last byte of a message (`11` \= MSEO, DDDDDD = DATA bits) * < `1*n` \> - sequence of `n`\-bits long idle bits (each must be `1`) * < `00` TTTTTT> - 8 bits in a first byte of a message (`00` \= MSEO, TTTTTTT = TCODE bits) The following 4 example sequences are all valid: * …​ < `11` DDDDDD> < `00` TTTTTT> …​ ⇒ No idle bits/cycles between consecutive messages. * …​ < `11` DDDDDD> < `1*2` \> < `00` TTTTTT> …​ ⇒ Two (even) idle bits. * …​ < `11` DDDDDD> < `1*3` \> < `00` TTTTTT> …​ ⇒ Three (odd) idle bits (second message starts at another trace clock edge). * …​ < `11` DDDDDD> < `1*8` \> < `00` TTTTTT> …​ ⇒ 8 idle bits (idle sequence can be considered as byte 0xFF). Some implementations may always send idle sequences using even (or even multiple of 8) number of trace clocks - in such a case all messages will always start on a positive or negative trace clock. But conformant trace probes must handle any number of idle clocks. | | The trace probe needs to be able to synchronize with the trace stream and to detect trace message boundaries. This procedure is sometimes referred to as "message alignment synchronization" or "alignment-sync". For 8-bit or 16-bit trace idle cycles are not required (to detect an alignment) as MSEO bits are in well-defined positions and trace probes will know where is a start of a message. For 1-bit, 2-bit and 4-bit trace modes PIB must generate at least one idle byte to allow trace probes to detect which bit is the first MSEO bit of a message. How it is done is not defined in this specification. Here are two possible implementations: Generate at least one idle byte periodically in a trace stream anywhere between messages (PIB is aware about message boundaries as end of message has MSEO=11 bits). Always add an extra idle byte before sending synchronizing messages. It will guarantee that boundaries of every synchronizing message are always detectable and decoding may start from it. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#3-1-5-n-trace-message-example)3.1.5\. N-Trace Message Example Table below shows one N-Trace message with several fields. It is an output from N-Trace dump tool (part of N-Trace reference C code) with an added **Explanation** column. __Table 3\. MDO and MSEO Encoding Example__ | Byte | MDO \[5:0\] | MSEO \[1:0\] | Decoded (by reference tool) | Explanation | | -------------------------------------------------------------------------------------------------------------------------------- | ----------- | ------------ | -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0xFF | 111111 | 11 | Idle | Most likely idle but can also be the last byte of the previous message. | | 0x70 | 011100 | 00 | TCODE\[6\] = 28 - IndirectBranchHist | First byte, all 6 MDO bits have TCODE. | | Here we could have an SRC field (it would shift the start of B-TYPE). | | | | | | 0xD0 | 110100 | 00 | B-TYPE\[2\] = 0x0 | This is a 2-bit (fixed-length) field. As B-TYPE is a fixed-length field, four most significant bits are part of the next field (I-CNT). | | 0x1D | 000111 | 01 | I-CNT\[10\] = 0x7D | This is a second byte of the 10-bit (value 0x7D) variable-length I-CNT field. Four least significant bits (0b1101=0xD) are defined in previous MDO. Three most significant bits are all 0-s as variable-length field uses all 6 MDO bits. | | 0x1D | 000111 | 01 | U-ADDR\[6\] = 0x7 | This is a single byte variable-length U-ADDR field (with three most significant 0-s). | | 0xF8 | 111110 | 00 | Normal transfer of new field (6 least significant bits). | | | 0xFF | 111111 | 11 | HIST\[12\] = 0xFFE | Last byte of message. It implies the end of the 12-bit HIST field. In this field we do not have any extra most significant 0-s. | | Here optional TSTAMP field could be sent(Previous MSEO should became 01 encoding end of HIST field, but not end of the message). | | | | | | 0xFF | 111111 | 11 | Idle | This is idle as this is the second byte with MSEO=11 (NOTE: Last byte of message is also 0xFF). | RISC-V Trace Connectors Specification ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-trace-connectors-specification)RISC-V Trace Connectors Specification RISC-V N-Trace Task Group Version 1.0, Nov 21, 2024: Ratified state | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 4.1. Adapters, multiple connectors and on-board debug considerations ==================== ## [](#4-1-adapters-multiple-connectors-and-on-board-debug-considerations)4.1\. Adapters, multiple connectors and on-board debug considerations It is often seen that some evaluation boards provide more than one standard connector. This is not only costly, but also not necessary as most trace and debug probe vendors provide passive adapters or cables to adapt different pinouts as part of standard offering. In case several connectors must be used, the highest performance connector should be placed as the closest one to trace MCU pins. For example, if you want to have Mictor for high-speed trace and MIPI10 for casual-debug (and/or slow serial trace), Mictor should have all JTAG and trace signals connected. All JTAG signals should go 'through' that Mictor connector and go to the MIPI10 connector. All high-speed trace signals should not go any further than to Mictor connector pins. In rare cases, when more than one trace connector is desired, it is suggested to place 0R/DNP resistors to reduce fanout on trace lines. Be aware that every PCB trace disruption (via, test-point, resistor) will cause reflections and signal degradation. It is also very important to provide good GND on all GND pins for high quality high-quality trace. Assure all trace lines on PCB are of similar length and have identical impedance. In case trace pins are shared as functional IO, make sure that it is possible to cut-out devices connected to trace data lines (via 0R resistors or solder bridges - jumpers are not recommended as these provide additional signal degradation). In case scoping of trace signals is necessary, it is suggested to have a good GND test point (where wire can be soldered) close to where scope can be connected. MIPI Alliance White Paper (referenced at the beginning) provides extra details as far as routing signal trace on target PCB. In case when on-board circuitry is used for debugging, that circuitry should monitor the GNDDetect pin (MIPI20/MIPI10 #9). In case GND is detected there, it means that external debug probe is connected to that connector and in such a case on-board debug chip should tri-state all its outputs and disable all pull-up/pull-down on all pins, so external debug probe operation will not be disturbed by on-board debug circuitry. Change Log ==================== ## [](#change-log)Change Log PDF generated on: 2026-07-17 19:42:42 UTC ### [](#version-1-0-ratified)Version 1.0 (Ratified) * 2024-11-21 * Ratified state (concent identical as in 1.0\_rc51 PDF) Contributors ==================== ## [](#contributors)Contributors Key contributors to RISC-V Trace Connectors specification in alphabetical order: Bruce Ableidinger (SiFive) ⇒ Working with MIPI Alliance, reviews Robert Chyla (IAR, SiFive) ⇒ Most topics, editing, publishing Markus Goehrle (Lauterbach) ⇒ Dual voltage, reviews Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at Copyright 2019-2024 by RISC-V International. 1.1. Debug and Trace Connectors ==================== ## [](#1-1-debug-and-trace-connectors)1.1\. Debug and Trace Connectors This specification provides a small, optional extension to connectors described in [MIPI Debug & Trace Connectors Recommendations White Paper, Version 1.20, 2 July 2021](https://resources.mipi.org/download-mipi-whitepaper-debug-trace-connector). These optional extensions are as follows: * Clarifying dual voltage debug and trace via Mictor 38 connector (re-defining obsolete pin #14). * Allowing MIPI20 `TRC_DATA[2]` and `TRC_DATA[3]` to be optionally used as TRIGIN/TRIGOUT pins. * Defining some signal as (optional) serial trace and/or application UART. MIPI Alliance positively reviewed proposals for above optional extensions and hopefully will adopt them in the future. Additionally, the following MIPI20 signals are clarified as follows: * MIPI20 pin#12 RTCK signal is not applicable to RISC-V. If the target is not providing a parallel trace, the target connector must provide GND on this pin. If parallel trace is used, pin#12 must be used as a `TRC_CLK` signal. * MIPI20 pin#14 nTRST\_PD signal is not applicable for RISC-V. If nTRST is really needed, pin#16 nTRST should be used. Above two signals were present in older RISC-V Debug Specification but were never implemented/used. This specification also adds the following option (described in dedicated chapter): * Defining MIPI20 pins #11 and #13 as optional TgtPwr+Cap pins (to supply 5V to power-up a small, evaluation target board). This option is already supported by several debug and trace probe vendors. 3.1. Mictor 38-bit Debug and Trace Connector ==================== ## [](#3-1-mictor-38-bit-debug-and-trace-connector)3.1\. Mictor 38-bit Debug and Trace Connector Mictor-38 connector as defined by MIPI Alliance has all signals from MIPI20 connector and adds up to 16 bits of parallel trace and defines more trigger pins. Mictor-38 connector is also designed for high-speed trace (it is rated for 400MHz double edge captures). Mictor-38 connector provides also an option to have different reference voltages for debug and trace. __Table 1\. Mictor-38 Connector Layout__ | Signal | Ref Voltage | Odd Pin# | Even Pin# | Ref Voltage | Signal | | ------------------- | ----------- | -------- | --------- | ------------ | ---------------------- | | NC | 1 | 2 | NC | | | | NC | 3 | 4 | NC | | | | GND | 5 | 6 | Trace | **TRC\_CLK** | | | TRIGIN | Debug | 7 | 8 | Debug | TRIGOUT | | nRESET | Debug | 9 | 10 | Trace | **EXTTRIG** | | TDO | Debug | 11 | 12 | Trace | **VREF\_TRACE** | | GND | 13 | 14 | Debug | VREF\_DEBUG | | | TCK / TCKC | Debug | 15 | 16 | Trace | **TRC\_DATA\[7\]** | | TMS / TMSC | Debug | 17 | 18 | Trace | **TRC\_DATA\[6\]** | | TDI | Debug | 19 | 20 | Trace | **TRC\_DATA\[5\]** | | nTRST | Debug | 21 | 22 | Trace | **TRC\_DATA\[4\]** | | **TRC\_DATA\[15\]** | Trace | 23 | 24 | Trace | **TRC\_DATA\[3\]** | | **TRC\_DATA\[14\]** | Trace | 25 | 26 | Trace | **TRC\_DATA\[2\]** | | **TRC\_DATA\[13\]** | Trace | 27 | 28 | Trace | **TRC\_DATA\[1\]** | | **TRC\_DATA\[12\]** | Trace | 29 | 30 | Trace | Logic '0' (GND) | | **TRC\_DATA\[11\]** | Trace | 31 | 32 | Trace | Logic '0' (GND) | | **TRC\_DATA\[10\]** | Trace | 33 | 34 | Trace | **Logic '1'** | | **TRC\_DATA\[9\]** | Trace | 35 | 36 | Trace | **EXT** / **TRC\_CTL** | | **TRC\_DATA\[8\]** | Trace | 37 | 38 | Trace | **TRC\_DATA\[0\]** | | | Above table is using names compatible with MIPI specification (however MIPI specification shows rows of pins starting from 38 down to 1). | | -------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#3-1-1-explanation-for-additional-pins-comparing-to-mipi20)3.1.1\. Explanation for additional pins (comparing to MIPI20) All debug signals share alternate functions as defined for the MIPI20 connector. __Table 2\. Micror-38 additional pins (comparing to MIPI20 defined above)__ | Pin# | Pin Name | Explanation (comparing to MIPI20) | | ---- | ---------------------- | --------------------------------------------------------------------------------------------------------------------------- | | 7 | TRIGIN | Same as MIPI20 #18 alternative pin function but not shared with trace. | | 8 | TRIGOUT | Same as MIPI20 #20 alternative pin function but not shared with trace. | | 10 | **EXTTRIG** | External trace trigger from target (some trace probes may use it). | | 21 | nTRST | Same as MIPI20 #16 alternative pin function but not shared with trace. | | 36 | **EXT** / **TRC\_CTL** | Not applicable (should be 0). May be also used to denote valid/idle state, but it may not be supported by all trace probes. | ### [](#3-1-2-dual-voltage-different-for-debug-and-different-for-trace-configurations)3.1.2\. Dual voltage (different for debug and different for trace) configurations Sometimes (due to speed reasons) it may be beneficial to drive SoC trace pins with different (usually lower) voltage then the debug signals. Such a configuration may be supported using a single Mictor connector or two connectors (Mictor for trace only and MIPI for debug only). Be aware that two different voltages may not be supported by simpler trace probes. **Single voltage - single Mictor (Recommended)** * Mictor #12: VREF\_TRACE=VREF\_DEBUG (Required) * Mictor #14: VREF\_DEBUG (Recommended, see NOTE \*1 below) or NC **Single voltage - trace via Mictor, debug via extra JTAG connector (NOT Recommended)** * Mictor #12: VREF\_TRACE=VREF\_DEBUG (Required) * Mictor #14: NC (Recommended, see NOTE #1 below) or VREF\_DEBUG * Mictor JTAG pins: Connected or NC (Recommended, see NOTE #2 below) * JTAG connector VTREF (#1): VREF\_DEBUG (Required) * JTAG connector JTAG pins: Connected (Required) **Dual voltage - single Mictor (NOT Recommended)** * Mictor #12: VREF\_TRACE (Required) * Mictor #14: VREF\_DEBUG via jumper on PCB (Required, see NOTE #3 below) **Dual voltage - trace via Mictor, debug via extra connector (Recommended)** * Mictor #12: VREF\_TRACE (Required) * Mictor #14: NC (Required, see NOTE #3 below) * Mictor JTAG pins: NC (Required, see NOTE #4 below) * JTAG connector VTREF (#1): VREF\_DEBUG (Required) * JTAG connector JTAG pins: Connected (Required) | | **#1** Jumper (on PCB) between Mictor pin#14 and VREF\_DEBUG rail on PCB can be used to select NC or VREF\_DEBUG. Some trace probes (such as TRACE32 from Lauterbach) require VTREF\_DEBUG to be present on pin #14. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | **#2** If JTAG pins are NC, JTAG quality/speed may be better as there will be no stubs introduced by extra routing on PCB. | | ----------------------------------------------------------------------------------------------------------------------------- | | | **#3** Jumper provides extra safety in case a trace probe/adapter which does not support dual voltage is used. Before fitting this jumper, make sure the probe/adapter you are using is NOT shorting Mictor pin#12/#14 internally. If this is the case, two voltage rails may be shorted and the target may be permanently damaged. Some trace probes (such as TRACE32 from Lauterbach) require VTREF\_DEBUG to be present on pin #14. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | **#4** All JTAG pins should be NC for a reason mentioned in NOTE 2\. But mainly to make sure that there will be only a single voltage present on this connector. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | **EXTRA NOTES (related to debug and trace voltages)** 1. Lower voltage allows faster trace, but it is then more critical to have correct PCB design. 2. Allowed reference voltage ranges (for JTAG and trace) are different for different probes. 3. Lower voltage used for trace may be a good choice with FPGA-based development boards. * Trace pins may be available on an FPGA bank, which is setup for lower IO voltage. 4. When high-speed trace is important Mictor-38 should be the only debug and trace connector on the PCB. * In case two connectors are used, trace signals should have routing priority. * Many probe vendors provide adapters from Mictor to standard JTAG-only connectors, so non-trace probes can be used with target/PCB with Mictor-only connector. 5. Not all trace probes which support the Mictor-38 connector are capable of handling dual voltage tracing. * At the moment of this writing at least I-jet-Trace-A/R/M (by IAR Systems) and Trace32 (by Lauterbach) probes support such a mode (in both single Mictor and two Mictor + JTAG connectors). 6. It is not recommended to add buffers on PCB to adjust JTAG (usually higher) voltage to trace voltage. * It not only affects signal quality but also introduces extra delays, which may create problems for simple probes. * It is very hard to properly handle fast switching of bidirectional signals, so cJTAG and SWD debug protocols may never reliably work. * It makes PCB more complicated without good reason. ### [](#3-1-3-explanation-for-mictor-38-pins-30323436)3.1.3\. Explanation for Mictor-38 pins #30/32/34/36 It may be hard to understand why `**TRC_DATA[0]**` is not together with other `**TRC_DATA[1..15]**` signals and why pins #30/32/34 have specific fixed values (Logic '0' or Logic '1'). This is caused by the desire to provide compatibility with initial versions of Arm trace. These older versions used these 4 pins to denote idle state. Modern trace probes ignore these signals, but just in case they do not, it is better and safer to provide logic level as above. As `**TRC_CTL**` is not used, it should be tied to 0 on PCB but may be optionally used as an extra external trace trigger (from target to probe). 2.1. MIPI20 Debug and Trace Connector ==================== ## [](#2-1-mipi20-debug-and-trace-connector)2.1\. MIPI20 Debug and Trace Connector This connector is an extension of a MIPI10 and MIPI20 connectors as defined by ratified**RISC-V External Debug Support, Version 0.13.2, Mar 22 2019** or newer. This connector adds 1-bit/2-bit/4-bit parallel trace and serial trace options on the same physical MIPI20 connector. * Trace related pins (added by this specification) are `**highlighted**`. * All JTAG/cJTAG pins have the same meaning as defined in the Debug Specification. * Notation "A / B" as signal name is used to separate two alternative functions of the same pin. __Table 1\. MIPI20 Connector Layout__ | Signal | Odd Pin# | Even Pin# | Signal | | ---------------- | -------- | --------- | ------------------------------------------------ | | VREF | 1 | 2 | TMS / TMSC | | GND | 3 | 4 | TCK / TCKC | | GND | 5 | 6 | TDO / **SerialTrace** (primary) | | GND or KEY | 7 | 8 | TDI | | GNDDetect | 9 | 10 | nRESET | | GND / TgtPwr+Cap | 11 | 12 | **TRC\_CLK** | | GND / TgtPwr+Cap | 13 | 14 | **TRC\_DATA\[0\]** / **SerialTrace** (secondary) | | GND | 15 | 16 | **TRC\_DATA\[1\]** / nTRST | | GND | 17 | 18 | **TRC\_DATA\[2\]** / TRIGIN | | GND | 19 | 20 | **TRC\_DATA\[3\]** / TRIGOUT | | | **SerialTrace** via pin #6 (TDO signal) is considered primary but it requires cJTAG interface to be used for debugging. Pin #14 is secondary and can be used when JTAG is used for debugging. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | | Smaller MIPI10 version of this connector (pins #1 .. #10 only) can provide **SerialTrace** via pin #6 (TDO signal) when cJTAG is used for debugging. | | ------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 2\. Details of MIPI20 Signals__ | Pin# | Pin Name | Explanation | | ---- | ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 1 | VREF | Reference voltage for all other pins and signals (single voltage for debug and trace). | | 2 | TMS / TMSC | JTAG TMS (from probe to target) or cJTAG TMSC (bi-directional) signal. | | 4 | TCK / TCKC | JTAG TCK (from probe to target) or cJTAG TCKC (from probe to target) signal. | | 6 | TDO / **SerialTrace** | Either JTAG TDO (from target to probe) or serial trace (from target to probe) available in case cJTAG is used for debugging. | | 7 | GND or KEY | May be removed pin (to prevent wrong insertion for non-shrouded connectors and cable with plug in pin#7). In case the pin is not removed, it must be GND on the target side. | | 8 | TDI | JTAG TDI (from probe to target) signal | | 9 | GNDDetect | Must be GND on the probe. On-board debug circuitry can use this pin to disable itself when the external debug probe is connected. If not used for that purpose it must be GND on the target side. | | 10 | nRESET | Active-low, open-drain SoC reset signal driven and monitored by the debug probe. Some debug probes may monitor this signal to handle and report resets from the target. | | 11 | GND / TgtPwr+Cap | In standard, most common configuration, these must be connected to GND. See below for explanation of optional TgtPwr+Cap function. | | 12 | **TRC\_CLK** | Parallel trace clock (from target to probe). | | 13 | GND / TgtPwr+Cap | In standard, most common configuration, these must be connected to GND. See below for explanation of optional TgtPwr+Cap function. | | 14 | **TRC\_DATA\[0\]** / **SerialTrace** | Either parallel trace signal (from target to probe) or serial trace (from target to probe). | | 16 | **TRC\_DATA\[1\]** / nTRST | Either parallel trace signal (from target to probe) or in case nTRST signal is needed this pin can be used as nTRST. NOTE: Still 1-bit parallel or serial trace is possible. | | 18 | **TRC\_DATA\[2\]** / TRIGIN | Either parallel trace signal (from target to probe) or input debug trigger (from probe to target) or application UART (from probe to target). | | 20 | **TRC\_DATA\[3\]** / TRIGOUT | Either parallel trace signal (from target to probe) or output debug trigger (from target to probe) or application UART (from target to probe). | ### [](#2-1-1-possible-use-of-trigintrigout-or-tditdo-for-an-application-uart)2.1.1\. Possible use of TRIGIN/TRIGOUT or TDI/TDO for an application UART Some debug probes may allow definition of pin functions and provide a virtual UART port/terminal for the target. UART is often needed for testing and production and having both debug and UART on a single connector is desired. Supporting UART over TRIGIN/TRIGOUT pins will limit parallel trace to 1-bit or 2-bit options. Supporting UART over TDI/TDO pins will require 2-pin cJTAG to be used as a debug interface. ### [](#2-1-2-explanation-of-tgtpwrcap-option-for-pins1113)2.1.2\. Explanation of TgtPwr+Cap option for pins#11/#13 | | This chapter explains optional use of MIPI20 pins #11/#13 to power-up small evaluations boards. This optional functionality is already provided by several debug and trace probe vendors. If you are not interested in such a functionality, you may skip reading this chapter and simply connect these pins to GND on the target PCB. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Meaning of optional TgtPwr+Cap function of pins #11/#13 is often misunderstood, so it deserves a more elaborated explanation. When the target cannot be powered from MIPI20 both these pins must be GND (as most of the pins on the odd side of MIPI20 connector). Another function of these pins (TgtPwr+Cap) is to provide target power supply voltage into the evaluation target. This way to power-up evaluation target is equivalent to power from the USB connector VBUS, so the expected voltage is around 5V. Target should not assume this voltage is regulated - the same way as voltage provided by USB cable is. Max current taken from these pins should not be larger than 100mA. | | Some debug probes may provide regulated voltage and dynamically measure total power consumption by the target via TgtPwr pins. | | --------------------------------------------------------------------------------------------------------------------------------- | Target boards should use jumper/switch to select board power-source (either from MIPI20 or USB connector). It is recommended to use a jumper/switch layout preventing both sources to be enabled at the same time. | | It is specifically **FORBIDDEN** to short together 5V power from USB (VBUS) and MIPI20 (pins#11/13) on target PCB. It will allow handling a case when a trace/debug probe or adapter has both pin#11/#13 connected to GND. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | It is possible to use two diodes (instead of jumpers) to auto-select the 5V power source and prevent back-feeding voltage from one source to the other, but it is not recommended as diodes will provide additional voltage drop. Term **TgtPwr+Cap** means that if these pins are used to provide power to the target, it must have a capacitor (as close to the pin as possible) to improve the quality of adjacent TRC\_CLK and TRC\_DATA pins. Another term for using a capacitor on the supply pin is an "AC ground" or "high frequency ground". We recommend 10pf capacitors placed extremely close to pins#11/#13. | | Leaving these pins not connected (NC) as can be seen on some schematics, is not a very good option when trace is used. There is simply not enough GND around TRC\_CLK and TRC\_DATA\[0\] signals. Some leave it as NC as they perhaps worry that debug probes may provide voltage there and it will create problems. Debug probes which support TgtPwr function provide GND detection and/or current protection and will disable TgtPwr voltage once detecting that target has these pins shorted to GND. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | No matter what pins #11 and #13 must be **always** connected - it is NOT possible that one of them will function as GND and second as TgtPwr. If you are in doubt, your board may have a jumper to either isolate these pins (NC) or connect them to GND or use them as target power. A jumper with 3 pins **A-B-C** should work. Middle pin **B** should go to MIPI20 pins#11/#13, the left pin **A** should be GND and the right pin **C** should be the 5V rail on the target (via another 3-way jumper allowing to select 5V from MIPI20 or USB VBUS). This allows to select one of three configuration options: * Jumper between **A-B** ⇒ MIPI20 pins #11/#13 are connected to GND. * Jumper between **B-C** ⇒ MIPI20 pins #11/#13 will be able to supply 5V power to the target. * No jumper ⇒ MIPI20 pins #11/#13 are left NC (**this is not a recommended option**). | | It is not possible to have both GND and 5V connections enabled at the same time as two jumpers cannot physically fit into 3 pins. | | ------------------------------------------------------------------------------------------------------------------------------------ | 5.1. Rationale (looking at Nexus standard) ==================== ## [](#5-1-rationale-looking-at-nexus-standard)5.1\. Rationale (looking at Nexus standard) Nexus standard does NOT define any small connectors with focus on trace as Nexus defines message-based debug interface and it requires more pins than JTAG. Namely: * S26x 1-104068-2, Low performance trace (1 MDO signal). * S40x 1-104549-6, Low performance trace (6 MDO signals - labeled as "not recommended"). * S50x 104549-7, Low performance trace (8 MDO signals). As the smallest Nexus-recommended connector with reasonable trace bandwidth has 50 pins these are not practical as trace connectors. So, it was decided that connectors defined by MIPI and Arm will be used for the RISC-V trace. * There are a lot of hardware trace probes, which are being used for debugging and tracing of Arm cores. Arm defines two standard connectors for trace: * Based on MIPI 20-pin connector (defined by MIPI) - this is for medium-performance tracing (4-bit, 100+ MHz double edge captures, max trace bandwidth 800Mbps or even higher for some high-performance trace probes). * Based on Mictor 38-pin connector (defined by MIPI) - this is for high-performance tracing (16-bit, up to 400MHz double edge, max trace bandwidth 12.8Gbps). * In July 2021 MIPI Alliance (following recommendations by Nexus TG group) released White Paper updating recommendations for debug and trace connectors and allowing 1/2/4-bit trace via MIPI20 connector. RISC-V Trace Control Interface Specification ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-trace-control-interface-specification)RISC-V Trace Control Interface Specification RISC-V N-Trace Task Group Version 1.0, Nov 21, 2024: Ratified state | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Additional Material ==================== ## [](#additional-material)Additional Material ### [](#minimal-implementation)Minimal Implementation This (non-normative) chapter gives an of what needs to be done to put together complete RISC-V trace implementation without getting familiar with every detail of every register. **Minimal General Registers/Fields** These requirements are applicable to the entire trace sub-system. * One `Trace Encoder` per hart being traced is required. * At least one of Trace RAM or Trace PIB sinks or Trace ATB Bridge is required as the final destination of an encoded RISC-V trace. * Implementations providing custom transport only are NOT considered fully compliant with this specification as custom trace tools will be needed. * Each trace component in a system is required to implement `tr??Control` and `tr??Impl` registers. * `tr??Active` bit must be settable to 0 or 1, although reset itself is NOT required. * `tr??Enable` bit must be settable to 0 or 1 and must support flushing (if applicable) when changed from 1 to 0. * `tr??Empty` bit must read as 0 when the trace component has some trace data internally buffered (if trace component it not buffering any trace data, this bit may be hard coded to 1). * `tr??VerMajor`, `tr??VerMinor` and `tr??CompType` must be implemented. **Minimal Trace Encoder Register/Fields** * Bit `trTeInstTracing` must be implemented (to start/stop instruction trace output from Trace Encoder). * One of `trTeInstMode` \= 3 (Baseline instruction trace) or 6 (Optimized instruction trace) must be implemented (can be a hard coded value). * At least one of the non-0 values of `trTeInstSyncMode` must be settable (or hard coded). * Field `trTeFormat` must correspond to implemented trace protocol (0 for E-Trace or 1 for N-Trace). * Fields `trTeProtocolMajor` and `trTeProtocolMinor` must return versions of implemented protocol. * All other registers/fields/bits may be tied to 0. **Minimal Trace RAM Sink Register/Fields** SRAM mode only: * Bit `trRamHasSRAM` must be tied to 1 and `trRamMode` must be tied to 0. * Bit `trRamWrap` must be implemented. * Register `trRamLimitLow` must be implemented but can be hard coded to value '2^M-4' (address 0x..FC). * Register `trRamWPLow` must at least accept a write of 0. * Register `trRamRPLow` must accept any 32-bit aligned value in inclusive range < 0 .. `trRamLimitLow` \>. * If width of access to SRAM is wider than 32-bits any 32-bit aligned value of `trRamRP` must be allowed and reads must be buffered. * Register `trRamData` may be implemented for reading only. * All other registers/fields/bits may be tied to 0. SMEM mode only: * Bit `trRamHasSMEM` must be tied to 1 and `trRamMode` must be tied to 1. * Bit `trRamWrap` must be implemented. * Register `trRamStart` must be implemented but can be hard coded to value '2^N' (address 0x..00). * Register `trRamLimit` must be implemented but can be hard coded to value '2^N + 2^M-4' (address 0x..FC). * Registers `trRamWP` must accept any 32-bit aligned value in inclusive range < `trRamStart` .. `trRamLimit` \>. * All other registers/fields/bits may be tied to 0. **Minimal Trace PIB Sink Register/Fields** It is hard to define required mode as it depends on SoC bandwidth requirements and capabilities, but some general guidance may be provided. * 4-bit mode is supported by most (if not all) trace probes and less expensive MIPI20 connectors can be used. * 1-bit and 2-bit modes should be only used when there are critical constraints on the number of MCU pins. Not all trace probes may support these modes. * Serial mode should be only considered when either limited trace is required, or cores run slowly. Not all trace probes may support this mode and max allowed speeds may vary. * Manchester encoding is self-synchronizing and may provide a more reliable trace. However UART mode may provide better bandwidth. It is suggested to support both. * 8-bit and 16-bit modes will provide better bandwidth, but require a more expensive Mictor connector and only more advanced trace probe models may support it. * It is suggested to provide as fast as possible trace logic clock, and allow a trace tool to set the divider in the `trPibDivider` field. * For TRC\_CLK frequencies higher than 50MHz, it is suggested to provide a calibration mode. * If possible, implement `trPibClkCenter` for better flexibility. **Minimal ATB Bridge Register/Fields** * Field `trAtbBridgeID` must be settable by trace tool (hard coded ID may not be handled by all trace tools). ### [](#reset-and-discovery)Reset and Discovery This chapter describes what trace tools should do to reset and discover trace features. | | Trace tools must be provided with base addresses of all trace components. | | ---------------------------------------------------------------------------- | Only the `tr??Active` and the `tr??Enable` bits are reset to 0 on power-up. These `tr??Active` bits act as independent resets for the respected trace components: * `trTeActive` \- reset for Trace Encoder component (this will disable encoder from single hart) * `trFunnelActive` \- reset for Trace Funnel component * `trPibActive` \- reset for PIB component (resets Pin Interface Block only) * `trRamActive` \- reset for RAM component (resets RAM Sink only) * `trAtbBridgeActive` \- resets ATB Bridge component (resets ATB Bridge interface) * `trTsActive` \- resets the Timestamp Unit sub-component (resets timestamp generation logic) When component is held in reset (`tr??Active` is 0), then `tr??Enable` bit must be reset to 0 as well (what makes component disabled). Releasing components from reset (by setting `tr??Active` to 1) may take time - debug tools should monitor (with reasonable timeout) if the appropriate bit changed from 0 to 1. As not all fields/registers are affected by reset (defined as [Undef](#Undef)), the trace tools must initialize (usually, a write of a value 0) several registers to assure that trace component is in a predictable state. | | Some of the reset values are defined as [SD](#SD) (system dependent) and these values should reset as well and each time to the same value as would be after power-up. Most of fields/registers have [Undef](#Undef) specified in reset behavior of the field. It should not prevent some implementations from resetting these. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Reset and Discovery should be performed as follows: * Reset the component by setting `tr??Active` \= 0. * This should be done by writing a value 0x0 to `tr??Control` register. * Read-back and wait until `tr??Active` \= 0 is read, which means that a component reached a reset state. * Release a component from reset by setting `tr??Active` \= 1. * This should be done by writing a value 0x1 to `tr??Control` register. This write will reset most of other enable/mode bits in this register and all WARL and read-only fields will be set to defaults. * Read-back and wait until `tr??Active` \= 1 is read, which means that a component was released from reset. * In this moment `tr??Enable` is set to 0 and the component is not yet enabled. * Component clock should be enabled to allow programming of other registers. * Optionally save `tr??Control` register as it holds all reset values of all fields. It may be cached/shadowed, and trace tool may execute faster write-only (instead a read-modify-write) operations. * Handle `tr??VerMinor/Major` as described in 'Versioning of Components' chapter. * If `tr??VerMajor` is 0 (for Trace Encoder component) either handle it as pre-ratified/initial version 0 or generate fatal error with an appropriate error message. * Read `tr??Impl` and compare `tr??ComType` field with expected value. * Set some WARL fields and read back to discover supported component configuration - make sure the component is NOT enabled (by setting `tr??Enable` to 1) by mistake. * Configure some initial values in all needed registers/fields. Optionally Read-back to assure these are set properly. The table below is showing what registers needs to be written to have each trace component in identical, predictable state. __Table 1\. **Trace Components Reset**__ | **Component** | **Register** | **Write** | **Notes** | | --------------------- | ------------------ | -------------------------------------------- | ---------------------------------------- | | Trace Encoder | trTeControl | 0x1 | Release from reset and set all defaults. | | trTeInstFeatures | 0x0 | Force minimal sub-set of features. | | | trTeInstFilters | 0x0 | Disable all filters. | | | trTeDataControl | 0x0 | Disable data trace. | | | trTeDataFilters | 0x0 | Disable filters for data trace. | | | trTeTrigDbgControl | 0x0 | Disable all triggers from Debug module. | | | trTeTrigExtInControl | 0x0 | Disable all external trigger inputs. | | | trTeTrigExtOutControl | 0x0 | Do not general any external trigger outputs. | | | trTsControl | 0x0 | Keep timestamp unit in reset. | | | Trace Funnel | trFunnelControl | 0x1 | Release from reset and set all defaults. | | trFunnelDisInput | 0x0 | Make sure all inputs are enabled. | | | trTsControl | 0x0 | Keep timestamp unit in reset. | | | Trace RAM Sink | trRamControl | 0x1 | Release from reset and set all defaults. | | Trace PIB Sink | trPibControl | 0x1 | Release from reset and set all defaults. | | Trace ATB Bridge | trAtbBridgeControl | 0x1 | Release from reset and set all defaults. | As we are dealing with several independent components, it is important to assure that the component which is in reset (or powered down) is keeping its outputs on safe values, so garbage trace data is not emitted. In general, it is safer to power-up and enable components starting from sinks/bridges, followed by Trace Funnels and Trace Encoders as last. Each implementation should test this sequence to assure trace tools are working seamlessly. | | Pre-release version of this specification specified that most of fields and registers were reset. It was suggested (by Architecture Review Committee) that reset logic is made minimal to follow RISC-V ISA style. Older implementations (which reset more bits) are still compatible with a ratified version. These implementations do not have to change as it is perfectly OK to reset more fields and registers. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#enabling-and-disabling)Enabling and Disabling Enabling should work as follows: * Release all needed components from reset by setting `tr??Active` \= 1 as described above. * Set desired mode and verify if that mode is set (regardless of discovery results). * For RAM Sink: * Setup needed addresses (if possible and desired) as these may not reset. * For PIB Sink: * Calibrate PIB (if possible and desired). * Start physical trace capture (trace probe dependent). * Configure RAM Sink/PIB Sink/ATB Bridge in appropriate mode. * Verify if a particular mode is set. * Set main enable for RAM Sink/PIB Sink/ATB Bridge component by setting `tr??Enable` \= 1. * Read back and wait for confirmation (`tr??Enable` \= 1). * Enable Trace Funnel\[s\] in the same way. * Configure and Enable Trace Encoder\[s\] in the same way (last should be writing `trTeEnable` \= 1 followed by reading to verify that it is set). * The `trTeInstMode` WARL field should be set to 6 - it may revert to different mode. * Either manually set `trTeInstTracing=1` and/or `trTeDataTracing=1` bits or set triggers to start the trace. * Start hart\[s\] to be traced (hart could be already running as well - in this case trace will be generated in the moment when `trTeInstTracing` or `trTeDataTracing` bit is set). * Periodically read `trTeControl` for status of trace (as it may stop by itself due to triggers). * If RAM Sink was configured with `trRamStopOnWrap` \= 1, read `trRamEnable` to see if RAM capture was stopped. | | Discovery may not be necessary to enable and test the trace during development of SoC. However, a discovery must be possible and should be tested by SoC designer - this is necessary for trace tools to work with that SoC without any customization. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Trace tools may verify a particular setting once per session, so subsequent starts of trace may be faster. | | ------------------------------------------------------------------------------------------------------------- | | | Trace tools should provide configuration settings allowing more verbose logging mode during discovery and initialization, so potential compatibility issues may be solved. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Disabling the trace should work as follows: * It is essential to disable the trace from encoders associated with stopped harts as entering debug mode is NOT flushing any trace pipelines. * Disable and flush trace starting from Trace Encoders, then Trace Funnels and finally Trace Sinks or Trace Bridges. * Set `tr??Enable` \= 0 and wait for `tr??Enable` \= 0 and `tr??Empty` \= 1 for each trace component. * It is important to do it in that order as otherwise data may be lost. * Stop physical capture if PIB sink was enabled (trace probe dependent). * Read the trace. * For RAM Trace Sink read `trRamWP` \- depending on `trRamWrap` bit, you may read trace from two ranges. * For RAM Trace Sink in SRAM mode, set `trRamRP` and read `trRamData` multiple times. * For RAM Trace Sink in SMEM mode, read trace from system memory using memory read. * For PIB Trace Sink read trace from trace probe. * For ATB Bridge, read trace using Coresight components (ETB/TMC/TPIU). Decoding trace * Decoder (in most cases) must have access to code which is running on device either by reading it from device or from a file containing the code (binary/hex/srec/ELF). * The trace collected by trace probes can be read and decoded while a trace is being captured (this is called trace streaming mode). * There is no guarantee that the last trace packet is completed until the trace is properly flushed and disabled. * Decoding of the trace should never affect code being traced. ### [](#pre-ratifiedinitial-interface-version)Pre-ratified/Initial Interface Version The value of `trTeVerMajor` as 0 means this is the pre-ratified/initial version of this trace control interface. Initially this specification was kept highly compatible, but after the decision to split all components into 4K regions it was very hard to track and list all changes and appropriate chapter was removed. The migration path from 'ver 0' (for both IP providers and tool vendors) should not be hard as the main concepts remain unchanged. Original donation from SiFive (which describes implementation of pre-ratified/initial version 0) can be found here:[RISC-V-Trace-Control-Interface-Proposed-20200612.pdf](https://lists.riscv.org/g/tech-nexus/files/RISC-V-Trace-Control-Interface-Proposed-20200612.pdf) | | Not all trace tools may support pre-ratified/initial version 0\. But all such tools should reject a version 0 with a very clear message. | | ------------------------------------------------------------------------------------------------------------------------------------------- | Trace ATB Bridge ==================== ## [](#trace-atb-bridge)Trace ATB Bridge Some SoCs may have an Advanced Trace Bus (ATB) infrastructure to manage trace produced by other components. In such systems, it may be desired to route entire RISC-V trace stream to the ATB through an ATB Bridge. This module manages the interface to ATB, generating ATB trace records that encapsulate RISC-V trace produced by the Trace Encoder\[s\] and/or Trace Funnel\[s\]. There is a control register that includes trace on/off control and a field allowing software to set the ID to be used on the ATB bus. This ID allows software to extract entire RISC-V trace from the combined trace. This interface is compatible with AMBA 4 ATB v1.1. __Table 1\. **Register: trAtbBridgeControl: ATB Bridge Control Register (trAtbBridgeBase+0x000)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ----------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 0 | trAtbBridgeActive | Primary activate/reset for the ATB Bridge. When 0, the ATB Bridge may have clocks gated off or be powered down, and other register locations may be inaccessible. Hardware may take an arbitrarily long time to process power-up and power-down and will indicate completion when the read value of this bit matches what was written. See [Reset and Discovery](#Reset and Discovery) chapter for more details. | RW | 0 | | 1 | trAtbBridgeEnable | **1:** ATB Bridge enabled. Setting trAtbBridgeEnable to 0 flushes any queued trace data to ATB. See [Enabling and Disabling](#Enabling and Disabling) chapter for more details. | RW | 0 | | 2 | — | Reserved | — | 0 | | 3 | trAtbBridgeEmpty | Reads 1 when ATB Bridge internal buffers are empty | RO | 1 | | 7:4 | — | Reserved | — | 0 | | 14:8 | trAtbBridgeID | ID of this node on ATB. Values of 0x00 and 0x70-0x7F are reserved by the ATB specification and should not be used. | RW | [Undef](#Undef) | | 31:15 | — | Reserved | — | 0 | __Table 2\. **Register: trAtbBridgeImpl: ATB Bridge Implementation Register (trAtbBridgeBase+0x004)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 3:0 | trAtbBridgeVerMajor | ATB Bridge Component Major Version. Value 1 means the component is compliant with this document. | RO | 1 | | 7:4 | trAtbBridgeVerMinor | ATB Bridge Component Minor Version. Value 0 means the component is compliant with this document. | RO | 0 | | 11:8 | trAtbBridgeCompType | ATB Bridge Component Type (ATB Bridge) | RO | 0xE | | 14:12 | trAtbBridgeAsyncFreq | **0:** Alignment synchronization (Async) packets disabled (may be the only choice for some protocols) **1-7:** Different levels of alignment synchronization (bigger number, bigger distance).Details should be defined the specification of each trace protocol. | WARL | [Undef](#Undef) | | 23:15 | — | Reserved for future versions of this standard | — | 0 | | 31:24 | — | Reserved for vendor specific implementation details | — | [SD](#SD) | An implementation determines the data widths of the connection from the Trace Encoder or Trace Funnel and of the ATB port. ATB Bridge may optionally insert ATB alignment synchronization packets (controlled by `trAtbBridgeAsyncFreq` field) allowing trace decoding software to detect ATB packet boundaries. Not all protocols may require it. Change Log ==================== ## [](#change-log)Change Log PDF generated on: 2026-07-17 19:42:42 UTC ### [](#version-1-0-ratified)Version 1.0 (Ratified) * 2024-11-21 * Ratified state (concent identical as in 1.0\_rc51 PDF) Contributors ==================== ## [](#contributors)Contributors Key contributors to RISC-V Trace Control Interface specification in alphabetical order: Bruce Ableidinger (SiFive) ⇒ Initial SiFive donation, reviews Robert Chyla (IAR, SiFive, MIPS) ⇒ Most topics, editing, publishing Ernie Edgar (SiFive) ⇒ Initial SiFive donation, reviews Jay Gamoneda (NXP) ⇒ Reviews, editing and updating after ARC review Markus Goehrle (Lauterbach) ⇒ Reviews, updates Iain Robertson (UltraSoC, Siemens) ⇒ E-Trace compatibility, filtering chapter, reviews Ved Shanbhogue (Rivos) ⇒ Detailed Architecture Review Committee notes Nino Vidovic (Segger) ⇒ Reviews Trace Control Interface Overview ==================== ## [](#trace-control-interface-overview)Trace Control Interface Overview The Trace Control interface consists of a set of 32-bit registers. The control interface can be used to set up and control a trace session, retrieve collected trace, and control any trace system components. ### [](#trace-components)Trace Components Each Trace Component is controlled by a set of 32-bit registers occupying up to 4KB of an address space. Base address of each trace component must be aligned on the 4KB boundary. Each hart being traced must have its own separate Trace Encoder control component. A system with multiple harts must allow generating messages with a field indicating which hart is responsible for that message. This specification defines the following trace components (**_N_** in at the end of symbol name denotes 0-based index of the component) __Table 1\. **Trace Components**__ | **Component Name** | **Component Type** (value=symbol) | **Base Address** (symbol #**_N_**) | | ------------------ | --------------------------------- | ---------------------------------- | | Trace Encoder | 0x1=TRCOMP\_ENCODER | trBaseEncoder**_N_** | | Trace Funnel | 0x8=TRCOMP\_FUNNEL | trBaseFunnel**_N_** | | Trace RAM Sink | 0x9=TRCOMP\_RAMSINK | trBaseRamSink**_N_** | | Trace PIB Sink | 0xA=TRCOMP\_PIBSINK | trBasePibSink**_N_** | | Trace ATB Bridge | 0xE=TRCOMP\_ATBBRIDGE | trBaseAtbBridge**_N_** | | | This specification does NOT address the discovery of base addresses of trace components. These base addresses (symbols in above table) must be specified as part of trace tool configuration. Connections between different trace components must be also defined. Future versions of this specification may allow a single base address to be sufficient to access and discover all trace components in the system. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#connections-between-components)Connections Between Components Different components must be connected via internal busses and/or FIFO buffers. This specification does not define this interconnect logic, but the following rules must be followed: * Each component sending a trace message/packet must assure the entire packet can be accepted by the destination component (or pushed into the FIFO buffer). * Sending a partial packet is NEVER allowed as it will not be possible to process and decode such a trace. * If a component cannot send an entire message/packet it must wait until it is possible to do so. * Tracing is typically required to be non-intrusive, and if the Trace Encoder cannot keep up with the hart it should drop the packet and wait for the receiver to be ready. * Once trace is allowed to resume it must issue an instruction trace synchronization message/packet so the decoder will be aware that some (unknown) amount of trace has been lost. * It is advisable to drain the trace pipeline to some hysteresis level before resuming - otherwise a lot of short chunks of trace may be produced. * Optionally the Trace Encoder may be configured to stall the hart to avoid trace data loss. * To prevent trace overflows the following techniques can be used: * Add a FIFO capable of holding several trace messages/packets to mitigate bursts of trace data. * Use wider internal busses to provide more bandwidth. * Make sure funnels and sinks provide the same or more bandwidth than encoders. * Use triggers to create trace windows/ranges to limit amount of trace data - especially in multi-core configurations. __Table 2\. **Allowed Connections Between Components**__ | **Input** | **Output** | **Description** | | ---------------- | ---------------- | ---------------------------------------------------------------------------------------------------------------- | | Ingress Port | Trace Encoder | Ingress Port (from hart) providing raw trace to be encoded | | Trace Encoder | Trace RAM Sink | Single hart tracing to RAM buffer | | Trace Encoder | Trace PIB Sink | Single hart tracing via pins | | Trace Encoder | Trace ATB Bridge | Single hart tracing to Arm ATB infrastructure | | Trace Encoder | Trace Funnel | Sending trace from single hart to Trace Funnel (to be combined from other RISC-V trace) | | Trace Funnel | Trace Funnel | Sending combined trace from multiple harts to higher level Trace Funnel (to be combined from other RISC-V trace) | | Trace Funnel | Trace RAM Sink | Sending combined trace from multiple harts to RAM buffer | | Trace Funnel | Trace PIB Sink | Sending combined trace from multiple harts via pins | | Trace Funnel | Trace ATB Bridge | Sending combined trace from multiple harts to Arm ATB infrastructure | | Trace ATB Bridge | Arm ATB bus | Sending trace to ATB (to combine RISC-V trace with other Arm components on the system) | | | Sending RISC-V trace to Arm CoreSight infrastructure is allowed (via ATB Bridge), but this specification does not specify how to transport trace data from other Arm CoreSight components in the system using RISC-V Trace sub-system. One of possible ways of doing so would be to create a custom trace component, configure it to encapsulate it as custom N-Trace trace messages and connect it as input to one of trace funnels. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#example-component-connection-diagrams)Example Component Connection Diagrams ![Simplest trace: Single Hart, Trace Encoder and Trace Sink/Bridge](_images/diag-3c71f65e64b6f17a911c4e9d05f5605a4cf9a403.svg) Figure 1\. Simplest trace: Single Hart, Trace Encoder and Trace Sink/Bridge ![Multi-hart trace: Three harts, three Encoders, single Funnel and single Sink/Bridge](_images/diag-e6960bbbac87caa9fcd6e53d37ca9f802e4cf07b.svg) Figure 2\. Multi-hart trace: Three harts, three Encoders, single Funnel and single Sink/Bridge ![Multi-cluster trace: two three-hart clusters with top-level Funnel and Sink/Bridge](_images/diag-b7ca382e406fcfb06a67a481a1f86ef9a262db71.svg) Figure 3\. Multi-cluster trace: two three-hart clusters with top-level Funnel and Sink/Bridge ![Local RAM Sink: Three-hart cluster plus extra hart with own RAM Sink (in SRAM mode)](_images/diag-434e2ddcf03728641ba00b25982c42d0bdf30bc3.svg) Figure 4\. Local RAM Sink: Three-hart cluster plus extra hart with own RAM Sink (in SRAM mode) | | Trace data from **Trace Encoder #4** may be combined with trace from other 3 Trace Encoders. But it may be also sent to dedicated **Trace RAM Sink** \- in such a case corresponding input to **Trace Funnel (top)** should be disabled. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#accessing-trace-control-registers)Accessing Trace Control Registers For the access method to the trace control registers, it makes a difference whether these registers shall be accessed by an external debug/trace tool, or by an internal debugger running on the chip. Trace control register access by an external debugger (this is the most common use case): * External debuggers must be able to access all trace control registers independent of whether the traced harts are running or halted. That is why for external debuggers, the recommended access method for memory-mapped control registers is memory accesses through the RISC-V debug module using SBA (System Bus Access) as defined in the RISC-V Debug Specification. Trace control register access by an internal debugger: * Through loads and stores performed by one or more harts in the system. Mapping the control interface into physical memory accessible from a hart allows that hart to manage a trace session independently from an external debugger. A hart may act as an internal debugger or may act in cooperation with an external debugger. Two possible use models are collecting crash information in the field and modifying trace collection parameters during execution. If a system has physical memory protection (PMP), a range can be configured to restrict access to the trace system from hart(s). | | Additional control path(s) may also be implemented, such as extra JTAG registers or devices, a dedicated DMI debug bus or message-passing network. Such an access (which is NOT based on System Bus) may require custom implementation by trace probe vendors as this specification only mandates probe vendors to provide access via SBA commands. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ### [](#register-map)Trace Component Register Map Each block of 32-bit registers (for each component) has the following layout: __Table 3\. **Register Layout for Component**__ | **Address Offset** | **Register Name** | **Compliance** | **Description** | | ------------------ | ----------------- | -------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | | 0x000 | tr??Control | Required | Main control register for this trace component | | 0x004 | tr??Impl | Required | Trace Implementation information for this trace component | | 0x008 - 0x00F | extra controls | Optional | Extra controls for this trace component (named differently) | | 0x010 - 0xDFF | — | Optional | Additional registers (specific for the type of a component). All not used registers are reserved and should read as 0 and ignore writes. | | 0xE00 - 0xFFF | — | Optional | Registers reserved for implementation/vendor specific details. May allow identification of components on a system bus. | | | Each component has a tr??Active bit in the tr??Control register. Accesses to other registers are unspecified when the tr??Active bit is 0. | | --------------------------------------------------------------------------------------------------------------------------------------------- | Each trace component has a `tr??Impl` register (at address offset 0x4) allowing trace component version and trace component type to be identified. This register allows debug tools to confirm the component type and potentially adjust tool behavior by looking at component versions. | | Each component may have a different version. Initial version of this specification defines all components to specify component version as 1.0 (major=1, minor=0). | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Registers in the 4KB range that are not implemented are reserved and read as 0 and ignore writes. Most trace control registers are optional. Some WARL fields may be hard coded to any value (including 0). It allows different implementations to provide different functionality. Both N-Trace and E-Trace encoders are controlled by the same set of bits/fields in the same `trTe???` registers - as almost every register, field, bit is optional this provides good flexibility in implementation. All other trace components are shared between different trace encoders (N-Trace and E-Trace). #### [](#summary-of-trace-encoder-registers)Summary of Trace Encoder Registers __Table 4\. **Trace Encoder Registers (trTe??, trTs??)**__ | **Address Offset** | **Register Name** | **Compliance** | **Description** | | ------------------------------------------------------ | --------------------- | -------------- | ------------------------------------------------------------- | | 0x000 | trTeControl | Required | Trace Encoder control register | | 0x004 | trTeImpl | Required | Trace Encoder implementation information | | 0x008 | trTeInstFeatures | Optional | Extra instruction trace encoder features and trace source IDs | | 0x00C | trTeInstFilters | Optional | Mask of filters to qualify an instruction trace | | **_Data trace control (trTeData??)_** | | | | | 0x010 | trTeDataControl | Optional | Data trace control and features | | 0x014 - 0x018 | — | Reserved | Reserved for data trace related future standard extension | | 0x01C | trTeDataFilters | Optional | Mask of filters to qualify data trace | | **_Reserved_** | | | | | 0x020 - 0x03F | — | Reserved | Reserved for future standard extension | | **_Timestamp control (trTs??)_** | | | | | 0x040 | trTsControl | Optional | Timestamp control register | | 0x044 | — | Reserved | Reserved for future timestamp related standard extension | | 0x048 | trTsCounterLow | Optional | Lower 32 bits of timestamp counter | | 0x04C | trTsCounterHigh | Optional | Upper bits of timestamp counter | | **_Trigger control (trTeTrig??)_** | | | | | 0x050 | trTeTrigDbgControl | Optional | Debug Triggers control register | | 0x054 | trTeTrigExtInControl | Optional | External Triggers Input control register | | 0x058 | trTeTrigExtOutControl | Optional | External Triggers Output control register | | **_Reserved_** | | | | | 0x060 - 0x3FF | — | Reserved | Reserved for future standard extension | | **_Filters & comparators (trTeFilter??, trTeComp??)_** | | | | | 0x400 - 0x5FF | trTeFilter?? | Optional | Trace Encoder Filter Registers | | 0x600 - 0x7FF | trTeComp?? | Optional | Trace Encoder Comparator Registers | #### [](#summary-of-trace-ram-sink-registers)Summary of Trace RAM Sink Registers __Table 5\. **Trace RAM Sink Registers (trRam??)**__ | **Address Offset** | **Register Name** | **Compliance** | **Description** | | ------------------ | ----------------- | -------------- | ------------------------------------------------------------------------- | | 0x000 | trRamControl | Required | RAM Sink control register | | 0x004 | trRamImpl | Required | RAM Sink Implementation information | | 0x008 - 0x00F | — | Reserved | Reserved for more control registers | | 0x010 | trRamStartLow | Required | Lower 32 bits of start address of circular trace buffer | | 0x014 | trRamStartHigh | Optional | Upper bits of start address of circular trace buffer | | 0x018 | trRamLimitLow | Required | Lower 32 bits of end address of circular trace buffer | | 0x01C | trRamLimitHigh | Optional | Upper bits of end address of circular trace buffer | | 0x020 | trRamWPLow | Required | Lower 32 bits of current write location for trace data in circular buffer | | 0x024 | trRamWPHigh | Optional | Upper bits of current write location for trace data in circular buffer | | 0x028 | trRamRPLow | Optional | Lower 32 bits of access pointer for trace readback | | 0x02C | trRamRPHigh | Optional | Upper bits of access pointer for trace readback | | 0x030 - 0x03F | — | Reserved | Reserved for more control registers | | 0x040 | trRamData | Optional | Read/write access to SRAM trace memory (32-bit data) | #### [](#summary-of-trace-pib-sink-registers)Summary of Trace PIB Sink Registers __Table 6\. **Trace PIB Sink Registers (trPib??)**__ | **Address Offset** | **Register Name** | **Compliance** | **Description** | | ------------------ | ----------------- | -------------- | ----------------------------------------- | | 0x000 | trPibControl | Required | Trace PIB Sink control register | | 0x004 | trPibImpl | Required | Trace PIB Sink Implementation information | #### [](#summary-of-trace-funnel-registers)Summary of Trace Funnel Registers __Table 7\. **Trace Funnel Registers (trFunnel??, trTs??)**__ | **Address Offset** | **Register Name** | **Compliance** | **Description** | | -------------------------------- | ----------------- | -------------- | --------------------------------------- | | 0x000 | trFunnelControl | Required | Trace Funnel control register | | 0x004 | trFunnelImpl | Required | Trace Funnel Implementation information | | 0x008 | trFunnelDisInput | Optional | Disable individual funnel inputs | | 0x00C - 0x03F | — | Reserved | Reserved for more control registers | | **_Timestamp control (trTs??)_** | | | | | 0x040 | trTsControl | Optional | Timestamp control register | | 0x044 | — | Reserved | Reserved for extra timestamp control | | 0x048 | trTsCounterLow | Optional | Lower 32 bits of timestamp counter | | 0x04C | trTsCounterHigh | Optional | Upper bits of timestamp counter | | | Funnels may optionally be a source of timestamp and/or forward timestamp to Trace Encoders in the system. This way several Trace Encoders may share timestamp and trace from several harts may be time-correlated. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#summary-of-trace-atb-bridge-registers)Summary of Trace ATB Bridge Registers __Table 8\. **Trace ATB Bridge Registers (trAtbBridge??)**__ | **Address Offset** | **Register Name** | **Compliance** | **Description** | | ------------------ | ------------------ | -------------- | ------------------------------------------- | | 0x000 | trAtbBridgeControl | Required | Trace ATB Bridge control register | | 0x004 | trAtbBridgeImpl | Required | Trace ATB Bridge Implementation information | Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at Copyright 2019-2024 by RISC-V International. Trace Encoder Control Interface ==================== ## [](#trace-encoder-control-interface)Trace Encoder Control Interface Many features of the Trace Encoder (TE for short) are optional. In most cases, optional features are enabled using a WARL (write any, read legal) register field. A debugger can determine if optional feature is present by writing to the register field and reading back the result. __Table 1\. **Register: trTeControl: Trace Encoder Control Register (trBaseEncoder+0x000)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------ | --------------- | | 0 | trTeActive | Primary activate/reset bit for the TE. When 0, the TE may have clocks gated off or be powered down, and other register locations may be inaccessible. Hardware may take an arbitrarily long time to process power-up and power-down and will indicate completion when the read value of this bit matches what was written. See [Reset and Discovery](#Reset and Discovery) chapter for more details. | RW | 0 | | 1 | trTeEnable | **1:** Trace Encoder is enabled. Allows trTeInstTracing and trTeDataTracing to turn tracing on and off. Setting trTeEnable to 0 flushes any queued trace data to the sink or funnel attached to this encoder. This bit can be set to 1 only by direct writing to it. This write of 1 should be done after all other settings are done. See [Enabling and Disabling](#Enabling and Disabling) chapter for more details. | RW | 0 | | 2 | trTeInstTracing | **1:** Instruction trace is being generated. Written from a trace tool (after a write to trTeEnable) or controlled by triggers. When trTeInstTracing=1, instruction trace data may be subject to additional filtering in some implementations (additional trTeInstMode settings). | RW | [Undef](#Undef) | | 3 | trTeEmpty | Reads as 1 when all generated trace have been emitted. | RO | 1 | | 6:4 | trTeInstMode | Instruction trace generation mode **0:** Full Instruction trace is disabled, but other trace (data trace) may be emitted. **1-2:** Protocol defined trace mode. **3:** Baseline instruction trace (for example [Branch Trace](#Branch Trace Messaging)). **4-5:** Protocol defined trace mode. **6:** Optimized instruction trace (for example [Branch History Trace](#Branch History Messaging)). **7:** Reserved for vendor-defined instruction trace mode.NOTE: When non-supported mode (different than 0) is set, it cannot revert to 0 but MUST revert to supported non-0 mode. | WARL | [Undef](#Undef) | | 8:7 | — | Reserved | — | 0 | | 9 | trTeContext | Enable sending trace messages/fields with scontext/mcontext values and/or privilege levels. | WARL | [Undef](#Undef) | | 10 | — | Reserved | — | 0 | | 11 | trTeInstTrigEnable | **1:** Allows trTeInstTracing to be set or cleared by Trace-on and Trace-off signals generated by the corresponding trigger module. | WARL | [Undef](#Undef) | | 12 | trTeInstStallOrOverflow | Set to 1 by hardware when trace buffer overflow (also known as trace lost) occurs, or when the TE requests a hart stall. Clears to 0 at TE reset or when the trace is enabled (trTeEnable set to 1). Write 1 to clear. | RW1C | [Undef](#Undef) | | 13 | trTeInstStallEna | **0:** If TE cannot send a message, the message is dropped. The protocol dependent overflow instruction trace synchronization message/packet is generated when the trace is restarted, so the decoder will know that trace is lost and must reset any internal decoder state. **1:** If TE cannot send a message, the hart is stalled until it can. With this option execution of instructions by the hart may be intrusively affected, but in many cases it is acceptable. | WARL | [Undef](#Undef) | | 14 | — | Reserved | — | 0 | | 15 | trTeInhibitSrc | **0:** Messages/packets generated by the trace encoder include a message source field if the source width held in trTeSrcBits is not 0. **1:** Disable inclusion of source field in trace messages/packets. | WARL | [Undef](#Undef) | | 17:16 | trTeInstSyncMode | Select the periodic instruction trace synchronization message/packet generation mechanism. At least one non-zero mechanism must be implemented. **0:** Off **1:** Count trace messages/packets **2:** Count hart clock cycles **3:** Count instruction 16-bit half-wordsOnce the max value of periodic counter is reached, an instruction trace synchronization message/packet should be generated. | WARL | [Undef](#Undef) | | 19:18 | — | Reserved | — | 0 | | 23:20 | trTeInstSyncMax | The maximum interval (in units determined by trTeInstSyncMode) between instruction trace synchronization messages/packets. Generate synchronization when count reaches 2^(trTeInstSyncMax+4). If an instruction trace synchronization message/packet is generated for another reason, the internal counter should be reset. | WARL | [Undef](#Undef) | | 26:24 | trTeFormat | Trace recording/protocol format: **0:** Format defined by Efficient Trace for RISC-V (E-Trace) Specification **1:** Format defined by RISC-V N-Trace (Nexus-based Trace) Specification **2-6:** Reserved for future formats **7:** Vendor-specific format | WARL | [Undef](#Undef) | | 31:27 | — | Reserved | — | 0 | | | Writing to this register while trace is enabled may unintentionally change a value of trTeInstTracing bit because that bit may dynamically change by triggers. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 2\. **Register: trTeImpl: Trace Encoder Implementation Register (trBaseEncoder+0x004)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ----------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------- | | 3:0 | trTeVerMajor | Trace Encoder Component Major Version. Value 1 means the component is compliant with this document. Value 0 means pre-ratified/initial version - see 'Pre-ratified/Initial Interface Version' chapter at the end. | RO | 1 | | 7:4 | trTeVerMinor | Trace Encoder Component Minor Version. Value 0 means the component is compliant with this document. | RO | 0 | | 11:8 | trTeCompType | Trace Encoder Component Type (Trace Encoder) | RO | 0x1 | | 15:12 | — | Reserved for future versions of this standard | — | 0 | | 19:16 | trTeProtocolMajor | Trace Protocol Major Version. As specified by specification governing trTeFormat. | RO | [SD](#SD) | | 23:20 | trTeProtocolMinor | Trace Protocol Minor Version. As specified by specification governing trTeFormat. | RO | [SD](#SD) | | 31:24 | — | Reserved for vendor specific implementation details | — | [SD](#SD) | | | trTeProtocol?? fields are separated from trTeVer?? as we may have the same control interface, but protocol itself may be extended with new packets/ messages/ fields. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | __Table 3\. **Register: trTeInstFeatures: Trace Instruction Features Register (trBaseEncoder+0x008)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 0 | trTeInstNoAddrDiff | When set, trace messages/packets always carry a full address. | WARL | [Undef](#Undef) | | 1 | trTeInstNoTrapAddr | When set, do not include trap handler address in trap messages/packets. | WARL | [Undef](#Undef) | | 2 | trTeInstEnSequentialJump | When set, treat sequentially inferrable jumps as inferable PC discontinuities. | WARL | [Undef](#Undef) | | 3 | trTeInstEnImplicitReturn | When set, treat returns as inferable PC discontinuities when returning from a recent call on a stack. Field trTeInstImplicitReturnMode below provides more details. | WARL | [Undef](#Undef) | | 4 | trTeInstEnBranchPrediction | When set, Branch Predictor based compression is enabled. | WARL | [Undef](#Undef) | | 5 | trTeInstEnJumpTargetCache | When set, Jump Target Cache based compression is enabled. | WARL | [Undef](#Undef) | | 7:6 | trTeInstImplicitReturnMode | Defines how the decoder is handling stack of return addresses (if enabled by trTeInstEnImplicitReturn bit): **0:** Implicit Return mode is not supported, or implementation is not reporting how it is implemented. **1:** Simple level counting without the return address comparing. **2:** Partial (LSB portion of return address) compare (smaller logic cost than 3 below, but in most cases adequate as chances to have an incorrect return address with same LSB bits is very slim). **3:** Full address comparing (always assures skipped return addresses are the same as addresses deducted from call instruction). Implementation may take advantage of RAS (Return Address Stack) if implemented by the hart. | WARL | [Undef](#Undef) | | 8 | trTeInstEnRepeatedHistory | Enable repeated branch history/map detection when set. | WARL | [Undef](#Undef) | | 9 | trTeInstEnAllJumps | Enable emitting of trace message or add history/map bit for direct unconditional/inferable control flow changes (jumps or calls). Normally these instructions do not generate any trace as the decoder can determine the next instruction. Trace will not compress well but timestamp accuracy will be better - may be used when profiling loops. | WARL | [Undef](#Undef) | | 10 | trTeInstExtendAddrMSB | When set, allow extended handing of MSB address bits. Encoding details are trace protocol dependent. | WARL | [Undef](#Undef) | | 15:11 | — | Reserved for additional instruction trace control/status bits | — | 0 | | 27:16 | trTeSrcID | Trace source ID assigned to this trace encoder. If trTeSrcBits is not 0 and trace source is not disabled by trTeInhibitSrc, then trace messages from this TE will all include a trace source field of trTeSrcBits bits and all messages from this TE will use this value as trace source field. | WARL | [Undef](#Undef) | | 31:28 | trTeSrcBits | The number of bits in the trace source field (0..12), unless disabled by trTeInhibitSrc. Some trace protocols may require that this field is identical for all enabled trace encoders within the same trace stream. | WARL | [Undef](#Undef) | | | Applicability of different trTeInst?? fields for each trace encoding protocol is described in a document which defines the protocol (and not all fields are applicable to all protocols). | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 4\. **Register: trTeInstFilters: Trace Instruction Filters Register (trBaseEncoder+0x00C)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------ | --------------- | | 15:0 | trTeInstFilters | Determine which filters defined in [Trace Encoder Filter Registers](#trace-encoder-filter-registers) chapter qualify an instruction trace. If bit **_n_** is a 1 then instructions will be traced when filter **_n_** matches. If all bits are 0, all instructions are traced. | WARL | [Undef](#Undef) | | 31:16 | — | Reserved | — | 0 | __Table 5\. **Register: trTeDataControl: Data Trace Control Register (trBaseEncoder+0x010)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------ | --------------- | | 0 | trTeDataImplemented | Read as 1 if data trace is implemented. | RO | [SD](#SD) | | 1 | trTeDataTracing | **1:** Data trace is being generated. Written from a trace tool or controlled by triggers. When trTeDataTracing\=1, data trace may be subject to additional filtering in some implementations. | WARL | [Undef](#Undef) | | 2 | trTeDataTrigEnable | Global enable/disable for data trace triggers | WARL | [Undef](#Undef) | | 3 | trTeDataStallOrOverflow | Set to 1 by hardware when data trace causes trace buffer overflow, or when the TE requests a hart stall due to data trace. Clears to 0 at TE reset or when the trace is enabled (trTeEnable set to 1). Write 1 to clear. | RW1C | [Undef](#Undef) | | 4 | trTeDataStallEna | **0:** If TE cannot send data trace messages, an overflow message is generated when the trace is restarted. **1:** If TE cannot send data trace messages, the hart is stalled until it can. | WARL | [Undef](#Undef) | | 5 | trTeDataDrop | Written to 1 by hardware when the data trace packet was dropped (if enabled). Clears to 0 at TE reset or when the trace is enabled (trTeEnable set to 1). Write 1 to clear. | RW1C | [Undef](#Undef) | | 6 | trTeDataDropEna | **1:** Allow temporary suppression of data trace (at some watermark level) to prevent trace overflow or stall. This way instruction trace will have higher priority. | WARL | [Undef](#Undef) | | 15:7 | — | Reserved for additional data trace control/status bits. | — | 0 | | 16 | trTeDataNoValue | When set, omit data values from data trace packets. | WARL | [Undef](#Undef) | | 17 | trTeDataNoAddr | When set, omit data address from data trace packets. | WARL | [Undef](#Undef) | | 19:18 | trTeDataAddrCompress | Data trace address compression selection: **0:** Only send full (unmodified) addresses **1:** Use XOR compression **2:** Use differential compression **3:** Protocol defined address compression | WARL | [Undef](#Undef) | | 31:20 | — | Reserved | — | 0 | | | Writing to this register while trace is enabled may unintentionally change a value of trTeDataTracing bit because that bit may dynamically change by triggers. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Applicability of different trTeData?? fields for each trace encoding protocol is described in a document which defines the protocol (and not all fields are applicable to all protocols). | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 6\. **Register: trTeDataFilters: Trace Data Filters Register (trBaseEncoder+0x01C)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 15:0 | trTeDataFilters | Determine which filters defined in [Trace Encoder Filter Registers](#trace-encoder-filter-registers) chapter qualify data trace. If bit **_n_** is a 1 then data accessed will be traced when filter **_n_** matches. If all bits are 0, all data accesses are traced. | WARL | [Undef](#Undef) | | 31:16 | — | Reserved | — | 0 | ### [](#timestamp-unit)Timestamp Unit Timestamp Unit is an optional sub-component present in either Trace Encoder or Trace Funnel. An implementation may choose from several modes of timestamps: * **Internal System** \- fixed clock in a system (such as bus clock) is used to increment the timestamp counter (for both Trace Encoders and Trace Funnels) * **Internal Core** \- core clock is used to increment the timestamp counter (only for Trace Encoders) * **Shared** \- shares timestamp with another Trace Encoder or Trace Funnel * **External** \- accepts a binary timestamp value from an outside source such as ARM CoreSight™ trace (for both Trace Encoders and Trace Funnels) Implementations may have no timestamp, one timestamp mode, or more than one mode. The WARL field `trTsMode` is used to determine the system capability and to set the desired timestamp mode. The width of the timestamp is implementation dependent, typically 40 or 48 bits (40-bit timestamp will overflow every 4.7 minutes assuming 1GHz timestamp clock). In a system with Funnels, typically all the Funnels are built with a Timestamp Unit. The top-level Funnel is the source of the timestamp (Internal System or External) and all the Encoders and other Funnels have a Shared timestamp. This assures that all timestamps in the system are the same and trace from different harts may be time-correlated. To perform the forwarding function, the mid-level Funnels must be programmed with `trFunnelActive` \= 1 (which is natural as all trace messages must pass through that funnel). An Internal System or Core timestamp unit may include a timestamp clock pre-scaler divider, which can extend the range of a narrower timestamp and uses less power but has less resolution. In a system with an Internal Core timestamp counter (implemented in Trace Encoder associated with a hart) an optional control bit is provided to stop the counter when the hart is halted by a debugger. __Table 7\. **Register: trBaseEncoder/Funnel+0x040 trTsControl: Timestamp Control Register**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ----------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------- | --------------- | | 0 | trTsActive | Primary activate/reset bit for timestamp unit. This must either be RW or, if separated reset for timestamp component is not implemented, a read-only copy of the corresponding trTeActive or trFunnelActive bit. See [Reset and Discovery](#Reset and Discovery) chapter for more details. | WARL | SD | | 1 | trTsCount | **Internal System or Core** timestamp only. **1:** counter runs, **0:** counter stopped. | WARL | [Undef](#Undef) | | 2 | trTsReset | **Internal System or Core** timestamp only.Write 1 to reset the timestamp counter. | W1 | — | | 3 | trTsRunInDebug | **Internal Core** timestamp only. **1:** counter runs when hart is halted (in debug mode), **0:** stopped | WARL | [Undef](#Undef) | | 6:4 | trTsMode | Mode used by Timestamp unit: **0:** None **1:** External **2:** Internal System **3:** Internal Core **4:** Shared **5-7:** Vendor-specific mode | WARL | [Undef](#Undef) | | 7 | — | Reserved | — | 0 | | 9:8 | trTsPrescale | **Internal System or Core** timestamp only.Prescale timestamp input clock by 2^(2\*trTsPrescale). It will be divided by 1, 4, 16, 64 respectively. | WARL | [Undef](#Undef) | | 14:10 | — | Reserved | — | 0 | | 15 | trTsEnable | Enable for timestamp field in trace messages/packets (for Trace Encoder only). | WARL | [Undef](#Undef) | | 23:16 | Vendor-specific bits to control what message/packet types include timestamp fields. | WARL | [Undef](#Undef) | | | 29:24 | trTsWidth | Width of timestamp in bits (0..63) | RO | [SD](#SD) | | 31:30 | — | Reserved | — | 0 | __Table 8\. **Register: trTsCounterLow: Timestamp Counter Lower Bits (trBaseEncoder/Funnel+0x048)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | -------------- | ----------------------------------- | ------ | --------- | | 31:0 | trTsCounterLow | Lower 32 bits of timestamp counter. | RO | 0 | __Table 9\. **Register: trTsCounterHigh: Timestamp Counter Upper Bits (trBaseEncoder/Funnel+0x04C)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | --------------- | ----------------------------------------------- | ------ | --------- | | 31:0 | trTsCounterHigh | Upper bits of timestamp counter, zero-extended. | RO | 0 | ### [](#trace-encoder-triggers)Trace Encoder Triggers #### [](#debug-trigger-module)Debug Trigger Module Debug module triggers are signals from the hart that a trigger was hit, but the action associated with that trigger is a trace-related action. Action identifiers 2-5 are reserved for trace actions in the RISC-V Debug Specification, where triggers are defined. Actions 2-4 are defined by the Efficient Trace for RISC-V (E-Trace) Specification. The desired action is written to the `action` field of the Match Control `mcontrol` CSR (0x7a1). As not all harts may support all trace actions, the debugger should read back the `mcontrol` CSR after setting the desired trace action to verify that the option exists. __Table 10\. **Debug Trigger Actions**__ | **Trigger Action (from debug spec)** | **Effect** | | ------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | Breakpoint exception (as defined in RISC-V Debug Specification) | | 1 | Debug exception (as defined in RISC-V Debug Specification) | | 2 | **Trace-on action**When trTeInstTrigEnable \= 1 it will start instruction tracing (trTeInstTracing → 1).When trTeDataTrigEnable \= 1 it will start data tracing (trTeDataTracing → 1). | | 3 | **Trace-off action**When trTeInstTrigEnable \= 1 it will stop instruction tracing (trTeInstTracing → 0).When trTeDataTrigEnable \= 1 it will stop data tracing (trTeDataTracing → 0). | | 4 | **Trace-notify action**If tracing is active (trTeInstTracing \= 1), then the encoder generates a packet with the current PC and, if enabled, a timestamp. | | 5 | **Vendor-specific trace action** (optional) | If there are vendor-specific features that require control, the `trTeTrigDbgControl` register is used. __Table 11\. **Register: trTeTrigDbgControl: Debug Trigger Control Register (trBaseEncoder+0x050)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------------ | ----------------------------- | ------ | --------------- | | 31:0 | trTeTrigDbgControl | Vendor-specific trigger setup | WARL | [Undef](#Undef) | #### [](#external-trace-triggers)External Trace Triggers The TE may be configured with up to 8 external trigger inputs for controlling trace. These are in addition to the external triggers present in the Debug Module when Halt Groups are implemented. The specific hardware signals comprising an external trigger are implementation dependent. External Trigger Outputs may also be present. A trigger out may be generated by trace starting, trace stopping, a watchpoint, or by other system-specific events. __Table 12\. **Register: trTeTrigExtInControl: External Trigger Input Control Register (trBaseEncoder+0x054)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 3:0 | trTeTrigExtInAction0 | Select action to perform when external trigger input #0 fires. If external trigger input #0 does not exist, then its action is fixed at 0. **0:** No action **1:** Reserved **2:** **Trace-on action**When trTeInstTrigEnable \= 1 it will start instruction tracing (trTeInstTracing → 1).When trTeDataTrigEnable \= 1 it will start data tracing (trTeDataTracing → 1). **3:** **Trace-off action**When trTeInstTrigEnable \= 1 it will stop instruction tracing (trTeInstTracing → 0).When trTeDataTrigEnable \= 1 it will stop data tracing (trTeDataTracing → 0). **4:** **Trace-notify action**If tracing is active (trTeInstTracing \= 1), then the encoder generates a packet with the current PC and, if enabled, a timestamp. **5-15:** Reserved | WARL | [Undef](#Undef) | | 31:4 | trTeTrigExtInAction**N** | Select actions (as defined for bits 3-0) for external trigger input #**N** (1..7). If an external trigger input does not exist, then its action is fixed at 0. | WARL | [Undef](#Undef) | __Table 13\. **Register: trTeTrigExtOutControl: External Trigger Output Control Register (trBaseEncoder+0x058)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 3:0 | trTeTrigExtOutEvent0 | Bitmap to select which event(s) cause external trigger #0 output to fire. If external trigger output #0 does not exist, then all bits are fixed at 0\. Bits 2 and 3 may be fixed at 0 if the corresponding feature is not implemented. **Bit 0:**Start trace transition (trTeInstTracing 0 → 1) will fire the trigger. **Bit 1:**Stop trace transition (trTeInstTracing 1 → 0) will fire the trigger. **Bit 2-3:**Vendor-specific event (optional) | WARL | [Undef](#Undef) | | 31:4 | trTeTrigExtOutEvent**N** | Select events for external trigger output #**N** (1..7). If an external trigger output does not exist, then its event bits are fixed at 0 | WARL | [Undef](#Undef) | #### [](#triggers-precedence)Triggers Precedence It is implementation dependent what happens when triggers (from debug module or external) with conflicting actions occur simultaneously (signaled at the same ingress port cycle) or if triggers occur too frequently. It is recommended that tracing starts from the oldest instruction retired in the cycle that Trace-on is asserted, and stops following the newest instruction retired in the cycle that Trace-off is asserted. ### [](#trace-encoder-filter-registers)Trace Encoder Filter Registers All registers with offsets 0x400 .. 0x7FC are designated for additional trace encoder filter options (context, addresses, modes, etc.). Trace encoder filters are an optional feature that can be used to control the generated trace in various ways. The registers below divide the filter logic into filters and comparators to provide maximum flexibility at low cost. The number of filters and comparators depends on the system. Each filter unit can specify filtering against instruction and optionally against data trace inputs from the hart. When filter _i_ is implemented, the registers `trTeFilter_i_Control` and `trTeInstFilters` must be implemented to enable it. And to apply filter _i_ to the data trace, the `trTeDataFilters` register must also be present. And if a match bit in the `trTeFilter_i_Control` register can be set to 1 (= enabling a filter option), the corresponding register from the bit’s description must have a correct value already set as otherwise the trigger may fire unintentionally. Each of the mentioned comparator units is a pair of comparators (primary and secondary, or P and S), so a limited range can be matched with a single comparator unit if needed. Each enabled filter define independent condition where trace is enabled - if several filters are enabled they act as `logical OR`. Several conditions for single filter act as `logical AND`. | | Filter and comparator registers refer to values of some signals (as **priv**, **itype**, **ecause**, **dtype**, **dsize**, …​) available on Trace Ingress Port. See E-Trace specification for details of encoding of these values. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 14\. **Register: trTeFilter??: Trace Encoder Filter Registers (trBaseEncoder+0x400..0x5FF)**__ | **Address Offset** | **Register Name** | **Compliance** | **Description** | | ------------------ | ----------------------------- | -------------- | -------------------------------------------- | | 0x400 + 0x20\*_i_ | trTeFilter_i_Control | Optional | Filter _i_ control | | 0x404 + 0x20\*_i_ | trTeFilter_i_MatchInst | Optional | Filter _i_ instruction match control | | 0x408 + 0x20\*_i_ | trTeFilter_i_MatchEcauseLow | Optional | Filter _i_ Ecause match control (bits 31:0) | | 0x40C + 0x20\*_i_ | trTeFilter_i_MatchEcauseHigh | Optional | Filter _i_ Ecause match control (bits 63:32) | | 0x410 + 0x20\*_i_ | trTeFilter_i_MatchValueImpdef | Optional | Filter _i_ impdef value | | 0x414 + 0x20\*_i_ | trTeFilter_i_MatchMaskImpdef | Optional | Filter _i_ impdef mask | | 0x418 + 0x20\*_i_ | trTeFilter_i_MatchData | Optional | Filter _i_ Data trace match control | | 0x41C + 0x20\*_i_ | — | Optional | Reserved | __Table 15\. **Register: trTeComp??: Trace Encoder Comparator Registers (trBaseEncoder+0x600..0x6FF)**__ | **Address Offset** | **Register Name** | **Compliance** | **Description** | | ------------------ | --------------------- | -------------- | ------------------------------------------- | | 0x600 + 0x20\*_j_ | trTeComp_j_Control | Optional | Comparator _j_ control | | 0x604 + 0x20\*_j_ | — | Optional | Reserved | | 0x608 + 0x20\*_j_ | — | Optional | Reserved | | 0x60c + 0x20\*_j_ | — | Optional | Reserved | | 0x610 + 0x20\*_j_ | trTeComp_j_PmatchLow | Optional | Comparator _j_ primary match (bits 31:0) | | 0x614 + 0x20\*_j_ | trTeComp_j_PmatchHigh | Optional | Comparator _j_ primary match (bits 63:32) | | 0x618 + 0x20\*_j_ | trTeComp_j_SmatchLow | Optional | Comparator _j_ secondary match (bits 31:0) | | 0x61C + 0x20\*_j_ | trTeComp_j_SmatchHigh | Optional | Comparator _j_ secondary match (bits 63:32) | __Table 16\. **Register: trTeFilter_i_Control : Filter _i_ Control Register (trBaseEncoder+0x400 + 0x20_i_)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 0 | trTeFilterEnable | Overall filter enable for filter #_i_ | WARL | [Undef](#Undef) | | 1 | trTeFilterMatchPrivilege | When set, match privilege levels specified by trTeFilterMatchChoicePrivilege field for filter #_i_. | WARL | [Undef](#Undef) | | 2 | trTeFilterMatchEcause | When set, start matching from exception cause codes specified by trTeFilterMatchChoiceEcause field for filter #_i_, and stop matching upon return from the 1st matching exception. | WARL | [Undef](#Undef) | | 3 | trTeFilterMatchInterrupt | When set, start matching from either an interrupt or exception as specified bytrTeFilterMatchValueInterrupt field for filter #_i_, and stop matching upon return from the 1st matching trap. | WARL | [Undef](#Undef) | | 4 | trTeFilterMatchComp1 | When set, the output of the comparator selected by trTeFilterComp1 must be true for the filter to match. | WARL | [Undef](#Undef) | | 7:5 | trTeFilterComp1 | Specifies the comparator unit to use for the 1st comparison. | WARL | [Undef](#Undef) | | 8 | trTeFilterMatchComp2 | When set, the output of the comparator selected by trTeFilterComp2 must be true for the filter to match. | WARL | [Undef](#Undef) | | 11:9 | trTeFilterComp2 | Specifies the comparator unit to use for the 2nd comparison. | WARL | [Undef](#Undef) | | 12 | trTeFilterMatchComp3 | When set, the output of the comparator selected by trTeFilterComp3 must be true for the filter to match. | WARL | [Undef](#Undef) | | 15:13 | trTeFilterComp3 | Specifies the comparator unit to use for the 3rd comparison. | WARL | [Undef](#Undef) | | 16 | trTeFilterMatchImpdef | When set, match **impdef** values as specified by trTeFilterMatchValueImpdef andtrTeFilterMatchMaskImpdef fields for filter #_i_. | WARL | [Undef](#Undef) | | 23:17 | — | Reserved | — | 0 | | 24 | trTeFilterMatchDtype | When set, match **dtype** values as specified by trTeFilterMatchChoiceDtype field for filter #_i_. | WARL | [Undef](#Undef) | | 25 | trTeFilterMatchDsize | When set, match **dsize** values as specified by trTeFilterMatchChoiceDsize field for filter #_i_. | WARL | [Undef](#Undef) | | 31:26 | — | Reserved | — | 0 | | | Handling of trTeFilterMatchEcause and trTeFilterMatchInterrupt should include a count of nested traps. The size of the counter is implementation dependent. If the number of nested traps exceeds the number that can be counted, the counter will saturate, meaning that the filtering will turn off prematurely. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 17\. **Register: trTeFilter_i_MatchInst : Filter _i_ Instruction Match Control Register (trBaseEncoder+0x404 + 0x20_i_)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 7:0 | trTeFilterMatchChoicePrivilege | When trTeFilterMatchPrivilege field for filter #_i_ is set, match all privilege levels for which the corresponding bit is set. For example, if bit N is 1, then match if the **priv** value at ingress port is N. Setting several bits allow matching several privileges. | WARL | [Undef](#Undef) | | 8 | trTeFilterMatchValueInterrupt | When trTeFilterMatchInterrupt field for filter #_i_ is set, match **itype** of 2 or 1 depending on whether this bit is 1 or 0 respectively. | WARL | [Undef](#Undef) | | 31:9 | — | Reserved | — | 0 | __Table 18\. **Register: trTeFilter_i_MatchEcauseLow : Filter _i_ Ecause Match Control (low) Register (trBaseEncoder+0x408 + 0x20_i_)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 31:0 | trTeFilterMatchChoiceEcauseLow | When trTeFilterMatchEcause field for filter #_i_ is set, match all excepion causes for which the corresponding bit is set. If bit N is 1, then match if the **ecause** is N. | WARL | [Undef](#Undef) | __Table 19\. **Register: trTeFilter_i_MatchEcauseHigh : Filter _i_ Ecause Match Control (high) Register (trBaseEncoder+0x40C + 0x20_i_)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------------------------- | -------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 31:0 | trTeFilterMatchChoiceEcauseHigh | Stores bits 63:32 to allow matching of higher **ecause** codes. If bit N is 1, then match if the **ecause** is N+32. | WARL | [Undef](#Undef) | __Table 20\. **Register: trTeFilter_i_MatchValueImpdef : Filter _i_ Impdef Match Value Register (trBaseEncoder+0x410 + 0x20_i_)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 31:0 | trTeFilterMatchValueImpdef | When trTeFilterMatchimpdef field for filter #_i_ is set, match if (**impdef** & trTeFilterMatchMaskImpdef) == (trTeFilterMatchValueImpdef & trTeFilterMatchMaskImpdef). | WARL | [Undef](#Undef) | __Table 21\. **Register: trTeFilter_i_MatchMaskImpdef : Filter _i_ Impdef Match Mask Register (trBaseEncoder+0x414 + 0x20_i_)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 31:0 | trTeFilterMatchMaskImpdef | When trTeFilterMatchimpdef field for filter #_i_ is set, match if (**impdef** & trTeFilterMatchMaskImpdef) == (trTeFilterMatchValueImpdef & trTeFilterMatchMaskImpdef). | WARL | [Undef](#Undef) | __Table 22\. **Register: trTeFilter_i_MatchData : Filter _i_ Data Match Control Register (trBaseEncoder+0x418 + 0x20_i_)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 15:0 | trTeFilterMatchChoiceDtype | When trTeFilterMatchDtype field for filter #_i_ is set, match all data access types for which the corresponding bit is set. For example, if bit N is 1, then match if the **dtype** value is N. | WARL | [Undef](#Undef) | | 23:16 | trTeFilterMatchChoiceDsize | When trTeFilterMatchDsize field for filter #_i_ is set, match all data access sizes for which the corresponding bit is set. For example, if bit N is 1, then match if the **dsize** value is N. | WARL | [Undef](#Undef) | | 31:24 | — | Reserved | — | 0 | __Table 23\. **Register: trTeComp_j_Control : Comparator _j_ Control Register (trBaseEncoder+0x600 + 0x20_j_)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 1:0 | trTeCompPInput | Determines which input to compare against the primary comparator. **0:** **iaddr** **1:** **context** **2:** **tval** **3:** **daddr** | WARL | [Undef](#Undef) | | 3:2 | trTeCompSInput | Determines which input to compare against the secondary comparator. Same encoding as trTeCompPInput. | WARL | [Undef](#Undef) | | 6:4 | trTeCompPFunction | Selects the primary comparator function. Primary result is true if input selected via trTeCompPInput is: **0:** equal to trTeCompPMatch **1:** not equal to trTeCompPMatch **2:** less than trTeCompPMatch **3:** less than or equal to trTeCompPMatch **4:** greater than trTeCompPMatch **5:** greater than or equal to trTeCompPMatch **6:** Result always false (input ignored). Prime latch to 1 if trTeCompMatchMode is 3 **7:** Result always true (input ignored) | WARL | [Undef](#Undef) | | 7 | — | Reserved | — | 0 | | 10:8 | trTeCompSFunction | Selects the secondary comparator function. Secondary result is true if input selected via trTeCompSInput is: **0:** equal to trTeCompSMatch **1:** not equal to trTeCompSMatch **2:** less than trTeCompSMatch **3:** less than or equal to trTeCompSMatch **4:** greater than trTeCompSMatch **5:** greater than or equal to trTeCompSMatch **6:** Result always true (input ignored). Use trTeCompSMatch as a mask for trTeCompPMatch **7:** Result always true (input ignored) | WARL | [Undef](#Undef) | | 11 | — | Reserved | — | 0 | | 13:12 | trTeCompMatchMode | Selects the match condition used to assert the overall comparator output **0:** primary result true **1:** primary and secondary result both true: (P && S) **2:** Either primary or secondary result does not match: !(P && S) **3:** Set when primary result is true and continue to assert until instruction after secondary result is true | WARL | [Undef](#Undef) | | 14 | trTeCompPNotify | Generate a trace packet explicitly reporting the address of the final instruction in a block that causes a primary match. This is also known as a watchpoint. Requires trTeCompPInput to be 0, and has no effect otherwise. | WARL | [Undef](#Undef) | | 15 | trTeCompSNotify | Generate a trace packet explicitly reporting the address of the final instruction in a block that causes a secondary match. This is also known as a watchpoint. Requires trTeCompSInput to be 0, and has no effect otherwise. | WARL | [Undef](#Undef) | | 31:16 | — | Reserved | — | 0 | | | Comparisions are performed as unsigned numbers. Only bits from an input signal (as defined by trTeCompPInput and/or trTeCompSInput fields), should be compared. Additional most significant bits from the trTeComp_j_PMatchLow/High registers must be ignored. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 24\. **Register: trTeComp_j_PMatchLow : Comparator _j_ Primary match (low) Register (trBaseEncoder+0x610 + 0x20_j_)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ----------------- | ------------------------------------------------------- | ------ | --------------- | | 31:0 | trTeCompPMatchLow | The match value for the primary comparator (bits 31:0). | WARL | [Undef](#Undef) | __Table 25\. **Register: trTeComp_j_PMatchHigh : Comparator _j_ Primary match (high) Register (trBaseEncoder+0x614 + 0x20_j_)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------------ | -------------------------------------------------------- | ------ | --------------- | | 31:0 | trTeCompPMatchHigh | The match value for the primary comparator (bits 63:32). | WARL | [Undef](#Undef) | __Table 26\. **Register: trTeComp_j_SMatchLow : Comparator _j_ Secondary match (low) Register (trBaseEncoder+0x618 + 0x20_j_)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ----------------- | --------------------------------------------------------- | ------ | --------------- | | 31:0 | trTeCompSMatchLow | The match value for the secondary comparator (bits 31:0). | WARL | [Undef](#Undef) | __Table 27\. **Register: trTeComp_j_SMatchHigh : Comparator _j_ Secondary match (high) Register (trBaseEncoder+0x61C + 0x20_j_)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------------ | ---------------------------------------------------------- | ------ | --------------- | | 31:0 | trTeCompSMatchHigh | The match value for the secondary comparator (bits 63:32). | WARL | [Undef](#Undef) | Trace Funnel ==================== ## [](#trace-funnel)Trace Funnel The Trace Funnel combines messages/packets from multiple sources into a single trace stream. It is implementation dependent how many incoming messages/packets are accepted before it is switching to another input source and in what order. But a continuous stream of messages/packets at one input cannot cause other inputs to not be handled. Suggested implementation would be to process just a single message/packet from each input in a round-robin fashion. __Table 1\. **Register: trFunnelControl: Trace Funnel Control Register (trBaseFunnel+0x000)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------- | | 0 | trFunnelActive | Primary activate/reset bit for trace funnel. When 0, the Trace Funnel may have clocks gated off or be powered down, and other register locations may be inaccessible. Hardware may take an arbitrarily long time to process power-up and power-down and will indicate completion when the read value of this bit matches what was written. See [Reset and Discovery](#Reset and Discovery) chapter for more details. | RW | 0 | | 1 | trFunnelEnable | **1:** Trace Funnel enabled. Setting trFunnelEnable to 0 flushes any queued trace data to output. See [Enabling and Disabling](#Enabling and Disabling) chapter for more details. | RW | 0 | | 2 | — | Reserved | — | 0 | | 3 | trFunnelEmpty | Reads 1 when Trace Funnel internal buffers are empty | RO | 1 | | 31:4 | — | Reserved | — | 0 | __Table 2\. **Register: trFunnelImpl: Trace Funnel Implementation Register (trBaseFunnel+0x004)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ---------------- | -------------------------------------------------------------------------------------------------- | ------ | --------- | | 3:0 | trFunnelVerMajor | Trace Funnel Component Major Version. Value 1 means the component is compliant with this document. | RO | 1 | | 7:4 | trFunnelVerMinor | Trace Funnel Component Minor Version. Value 0 means the component is compliant with this document. | RO | 0 | | 11:8 | trFunnelCompType | Trace Funnel Component Type (Trace Funnel) | RO | 0x8 | | 23:12 | — | Reserved for future versions of this standard | — | 0 | | 31:24 | — | Reserved for vendor specific implementation details | — | [SD](#SD) | __Table 3\. **Register: trFunnelDisInput: Disable Individual Funnel Inputs (trBaseFunnel+0x008)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ---------------- | ------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 15:0 | trFunnelDisInput | **1:** Funnel input **#n** (bit position in register) is disabled. Incoming messages are read from diabled input but discarded. | WARL | [Undef](#Undef) | | 31:16 | — | Reserved | — | 0 | | | trFunnelDisInput register is optional. When not implemented (or never set) it will read as 0, which means that all inputs are always enabled. When implemented, it can be set to 0xFFFF to detect which inputs may be disabled in that trace funnel. Disabling inputs is needed when a single trace encoder may provide output to more than one possible active destination/sink. This register can be also used by trace tools to easily configure a trace in complex systems. Without the ability to disable individual funnel inputs, the trace tool must assure all trace sources which should not be traced are disabled. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#timestamp-unit)Timestamp Unit Trace Funnel may optionally include Timestamp Unit. It is described inside of the Trace Encoder chapter above. Introduction ==================== ## [](#introduction)Introduction This document presents a standardized control interface for RISC-V trace infrastructure (such as trace encoders, trace funnels, trace sinks, …​) for the _Efficient Trace for RISC-V Version 2.0 Specification_ and for the _RISC-V N-Trace (Nexus-based Trace) Specification Version 1.0.0_. Standardized control interface allows trace control software development tools to be used interchangeably with any RISC-V device implementing processor and/or data trace. Instruction Trace is a system that collects a history of processor execution, along with other events. The trace system may be set up and controlled using a register-based interface. Hart execution activity appears on the Ingress Port and feeds into a Trace Encoder where it is compressed and formatted into trace messages. The Trace Encoder transmits trace messages to a Trace Sink. In multi-core systems, each hart has its own Trace Encoder, and typically all will connect to a Trace Funnel that aggregates the trace data from multiple sources and sends the data to a single destination. This specification does not define the hardware interconnection between the hart and Trace Encoder, as this is defined in the _Efficient Trace for RISC-V Specification Version 2.0_. This document also does not define the hardware interconnection between the Trace Encoder and Trace Funnel, or between the Trace Encoder/Funnel and Trace Sink. This specification allows a wide range of implementations including low-gate-count minimal instruction trace and systems with only instrumentation trace. Implementation choices include whether to support instruction trace, data trace, instrumentation trace, timestamps, external triggers, various trace sink types, and various optimization tradeoffs between gate count, features, and bandwidth requirements. ### [](#glossary)Glossary **Trace Encoder (TE for short)** \- Hardware module that accepts execution information from a hart and generates a stream of trace messages/packets. **Trace Message/Packet** \- Depending on protocol different names can be used, but it means the same. It is considered as a continuous sequence of (usually bytes) describing program and/or data flow and other events. **Trace Funnel** \- Hardware module that combines trace streams from multiple trace sources (Trace Encoders and/or other Trace Funnels) into a single output stream of trace messages/packets. **Trace Sink** \- Hardware module that accepts a stream of trace messages/packets and records them into the memory or forwards them onward in some format. **Trace Decoder** \- Software program that takes a recorded trace (from a Trace Sink) and produces a readable execution history. **RO** \- Denotes read-only bit/field - it does not mean it will return the same value each time when read. **RW** \- Denotes read-write bit/field - value being read may not be the same as what was written as some fields may change their values because of other reasons. **RW1C** \- Denotes bit/field, which can be read but you must write 1 to clear it (writing 0 will be ignored). It is used for sticky status bits to assure that these are cleared by deliberate action (write 1). **WARL** \- Denotes Write any, read legal bit/field/register. If a non-legal value is written, the written value is converted to a value that is supported. That value should deterministically depend on the illegal written value and the architectural state of the trace sub-system. **W1** \- Denotes write-only bit, which performs an action when 1 is written to it. **SD** \- Reset value of a field/register is system dependent - these fields should always have the same values at trace component reset. In many cases this may be the only value supported. **Undef** \- This field/register may not reset. Trace tool must write correct value before enabling the trace component. **ATB** \- Advanced Trace Bus, a protocol described in ARM document _AMBA ATB Protocol Specification_. This is one of alternative methods to send the trace (in addition to native Trace Sinks defined in this specification). **PIB** \- Pin Interface Block, a parallel or serial off-chip trace port feeding into a trace probe. **??** \- Used in names refer to identical fields/registers in different components. For example `tr??Active` may mean `trTeActive` or `trTsActive`. Trace PIB Sink ==================== ## [](#trace-pib-sink)Trace PIB Sink Trace data may be sent to chip pins through an interface called the Pin Interface Block (PIB). This interface typically operates at a few hundred MHz and can sometimes be higher with careful constraints and board layout or by using LVDS or other high-speed signal protocol. PIB may consist of just one signal and in this configuration may be called SWT (Serial-Wire Trace). Alternative configurations include a trace clock (TRC\_CLK) and 1/2/4/8/16 parallel trace data signals (TRC\_DATA) timed to that trace clock. WARL register fields are used to determine specific PIB capabilities. The modes and behavior described here are intended to be compatible with trace probes available in the market. **PIB Register Interface** __Table 1\. **Register: trPibControl: PIB Sink Control Register (trBasePib+0x000)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 0 | trPibActive | Primary activate/reset bit for PIB Sink component. When 0, the PIB Sink may have clocks gated off or be powered down, and other register locations may be inaccessible. Hardware may take an arbitrarily long time to process power-up and power-down and will indicate completion when the read value of this bit matches what was written. See [Reset and Discovery](#Reset and Discovery) chapter for more details. | RW | 0 | | 1 | trPibEnable | **0:** PIB does not accept input but holds output(s) at idle state defined by pibMode. **1:** Enable PIB to generate output. See [Enabling and Disabling](#Enabling and Disabling) chapter for more details. | RW | 0 | | 2 | — | Reserved | — | 0 | | 3 | trPibEmpty | Reads 1 when PIB internal buffers are empty. | RO | 1 | | 7:4 | trPibMode | Select mode for output pins. Allowed values are described in the Allowed PIB Configurations table below. | WARL | [Undef](#Undef) | | 8 | trPibClkCenter | In parallel modes, adjust TRC\_CLK timing to the center of the bit period. This can be set only if trPibMode selects one of the parallel protocols. | WARL | [Undef](#Undef) | | 9 | trPibCalibrate | Set this to 1 to generate a repeating calibration pattern to help tune a probe’s signal delays, bit rate, etc. In this mode input to the sink is not consumed. The calibration pattern is described below. | WARL | [Undef](#Undef) | | 11:10 | — | Reserved | — | 0 | | 14:12 | trPibAsyncFreq | **0:** Alignment synchronization (Async) packets disabled (may be the only choice for some protocols) **1-7:** Different levels of alignment synchronization (bigger number, bigger distance).Details should be defined in the specification of each trace protocol. | WARL | [Undef](#Undef) | | 15 | — | Reserved | — | 0 | | 31:16 | trPibDivider | Timebase selection for the PIB module. The input clock is divided by trPibDivider \+ 1\. PIB data is sent at either this divided rate or 1/2 of this rate, depending on trPibMode. Width is implementation dependent. After the PIB reset value of this field should be set to safe (not too fast clock) setting for a particular SoC. Trace tools may set smaller values to utilize higher bandwidth. | WARL | [Undef](#Undef) | __Table 2\. **Register: trPibImpl: Trace PIB Implementation Register (trBasePib+0x004)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------- | ---------------------------------------------------------------------------------------------------- | ------ | --------- | | 3:0 | trPibVerMajor | Trace PIB Sink Component Major Version. Value 1 means the component is compliant with this document. | RO | 1 | | 7:4 | trPibVerMinor | Trace PIB Sink Component Minor Version. Value 0 means the component is compliant with this document. | RO | 0 | | 11:8 | trPibCompType | Trace PIB Sink Component Type (PIB Sink) | RO | 0xA | | 23:12 | — | Reserved for future versions of this standard | — | 0 | | 31:24 | — | Reserved for vendor specific implementation details | — | [SD](#SD) | Software can determine what modes are available by attempting to write each mode setting to the WARL field `trPibMode` and reading back to see if the value was accepted. __Table 3\. **Allowed PIB Configurations**__ | **Mode** | **trPibMode** | **trPibClkCenter** | **Bit rate** | | ----------------------- | ------------- | ------------------ | ------------ | | Off | 0 | X | — | | SWT Manchester | 4 | X | 1/2 | | SWT UART | 5 | X | 1 | | TRC\_CLK + 1 TRC\_DATA | 8 | 0 | 1 | | TRC\_CLK + 2 TRC\_DATA | 9 | 0 | 1 | | TRC\_CLK + 4 TRC\_DATA | 10 | 0 | 1 | | TRC\_CLK + 8 TRC\_DATA | 11 | 0 | 1 | | TRC\_CLK + 16 TRC\_DATA | 12 | 0 | 1 | | TRC\_CLK + 1 TRC\_DATA | 8 | 1 | 1/2 | | TRC\_CLK + 2 TRC\_DATA | 9 | 1 | 1/2 | | TRC\_CLK + 4 TRC\_DATA | 10 | 1 | 1/2 | | TRC\_CLK + 8 TRC\_DATA | 11 | 1 | 1/2 | | TRC\_CLK + 16 TRC\_DATA | 12 | 1 | 1/2 | Since the PIB supports many different modes, it is necessary to follow a particular programming sequence: * Activate the PIB by setting `trPibActive`. * Set the `trPibMode`, `trPibDivider`, `trPibClkCenter`, and `trPibCalibrate` fields. This will set the TRC\_DATA outputs to the quiescent state (whether that is high or low depends on `trPibMode`) and start TRC\_CLK running. * Activate the receiving device, such as a trace probe. Allow time for PLL to sync up, if using a PLL with a parallel PIB mode. * Set `trPibEnable`. This enables the PIB to generate output either immediately (calibration mode) or when the Trace Encoder or Trace Funnel begins sending trace messages/packets. ### [](#order-of-bits-and-bytes)Order of bits and bytes * Trace messages/packets are considered as sequences of bytes and are always transmitted with least significant bits/bytes first. * In 16-bit mode (`trPibMode` \== 12) the byte transmitted on bits #0-#7 is considered first and most significant bits#8-#15 are transmitting second byte. * Idle sequences (no message/packet to be sent) are transmitted between messages. * Idle sequence depends on trace protocol and must allow detection of the start of first byte of message/packet following the idle sequence. * Idle sequences may be different and should be defined by trace protocols. ### [](#pib-parallel-protocol)PIB Parallel Protocol Traditionally, off-chip trace has used this protocol. There are several parallel data signals (TRC\_DATA0..15) and one continuously-running trace clock (TRC\_CLK). The data rate of parallel signals can be much higher than either of the serial-wire protocols. This protocol is oriented to send full, variable length trace messages/packets rather than fixed-width trace words. When a message start is detected, this sample and possibly the next few (depending on the width of TRC\_DATA) are collected until a complete byte has been received. Bytes are transmitted least significant bit first, with TRC\_DATA\[0\] representing the least significant bit in each beat of data. The receiver continues collecting bytes until a complete message has been received. The criteria for this depends on the trace format. After the last byte of a message, the data signals may then go to their idle state or a new message may begin in the next trace clock edge. #### [](#pib-clock-center)PIB Clock Center The trace clock, TRC\_CLK, normally has edges coincident with the TRC\_DATA edges. Typically, a trace probe will delay trace data or use a PLL to recover a sampling clock that is twice the frequency of TRC\_CLK and shifted 90 degrees so that its rising edges occur near the center of each bit period. If the PIB implementation supports it, the debugger can set `trPibClkCenter` to change the timing of TRC\_CLK so that there is a TRC\_CLK edge at the center of each bit period on TRC\_DATA. Note that this option cuts the data rate in half relative to normal parallel mode and still requires the probe to sample TRC\_DATA on both edges of TRC\_CLK. This example shows 8-bit parallel mode with `trPibClkCenter` \= 0 transmitting a 5-byte message/packet followed by a 2-byte message/packet. ![image](_images/RISC-V-Trace-Control-Interface-images/pib-ref0.png) And an example showing 8-bit parallel mode transmitting a 4-byte packet with `trPibClkCenter` \= 1 ![image](_images/RISC-V-Trace-Control-Interface-images/pib-ref1.png) ### [](#swt-manchester-protocol)SWT Manchester Protocol In this mode, the PIB outputs complete trace messages encapsulated between a start bit and a stop bit. Each bit period is divided into 2 phases and the sequential values of the TRC\_DATA\[0\] pin during those 2 phases denote the bit value. Bits of the message are transmitted LSB first. The idle state of TRC\_DATA\[0\] is low in this mode. __Table 4\. **Manchester Encoding Patterns**__ | **Bit** | **Phase 1** | **Phase 2** | | --------- | ----------- | ----------- | | start | 1 | 0 | | logic 0 | 0 | 1 | | logic 1 | 1 | 0 | | stop/idle | 0 | 0 | ![image](_images/RISC-V-Trace-Control-Interface-images/swt-manchester.jpg) ### [](#swt-uart-protocol)SWT UART Protocol In UART protocol, the PIB outputs bytes of a trace message encapsulated in a 10-bit packet consisting of a low start bit, 8 data bits, LSB first, and a high stop bit. Another packet may begin immediately following the stop bit or there may be an idle period between packets. When no data is being sent, TRC\_DATA\[0\] is high in this mode. ![image](_images/RISC-V-Trace-Control-Interface-images/swt-uart.jpg) ### [](#calibration-mode)Calibration Mode In optional calibration mode, the PIB transmits a repeating pattern. Probes can use this to automatically tune input delays due to skew on different PIB signal lines and to adjust to the transmitter’s data rate (`trPibDivider` and `trPibClkCenter`). Calibration patterns for each mode are listed below. __Table 5\. **PIB Calibration Patterns**__ | **Mode** | **Calibration Bytes** | **Wire Sequence** | | ---------------- | ----------------------- | ---------------------------------------------- | | UART, Manchester | AA 55 00 FF | alternating 1/0, then all 0, then all 1 | | 1-bit parallel | AA 55 00 FF | alternating 1/0, then all 0, then all 1 | | 2-bit parallel | 66 66 CC 33 | 2, 1, 2, 1, 2, 1, 2, 1, 0, 3, 0, 3, 3, 0, 3, 0 | | 4-bit parallel | 5A 5A F0 0F | A, 5, A, 5, 0, F, F, 0 | | 8-bit parallel | AA 55 00 FF | AA, 55, 00, FF | | 16-bit parallel | AA AA 55 55 00 00 FF FF | AAAA, 5555, 0000, FFFF | | | Calibration mode may be used even by probes which do not support calibration of trace just to assure trace routing on PCB is correct and PIB is correctly enabled. It may be also possible to use calibration mode to check trace signal routing from SoC using scope or logic analyzer. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Trace Protocols and Trace Control ==================== ## [](#trace-protocols-and-trace-control)Trace Protocols and Trace Control There are two standard RISC-V trace protocols which will utilize this **RISC-V Trace Control Interface**: * **RISC-V N-Trace (Nexus-based Trace) Specification** * Version 1.0 to be ratified together with this specification. * **Efficient Trace for RISC-V Specification** * Version 2.0 (ratified May 5-th 2022). This specification together with details provided in any of above documents should be considered as a complete guideline for any standard RISC-V trace implementation. Trace is controlled by set of 32-bit memory-mapped registers. Not all trace protocols and components must support all registers, bits, fields and options. This document includes a chapter [Minimal Implementation](#Minimal Implementation) which describes the smallest possible set of registers and fields, but each message protocol supported by this standard must clarify the exact meaning of supported registers/fields and bits as some of them define. Trace RAM Sink ==================== ## [](#trace-ram-sink)Trace RAM Sink Trace RAM Sink may be instantiated or configured to support storing trace into dedicated SRAM or system memory. SRAM mode is using dedicated local memory inside of RAM sink, while system memory mode (SMEM mode) is accessing memory via system bus (care should be taken to not overwrite application code or data - it is usually done by reserving part of system memory for trace). Dedicated SRAM memory must be read via dedicated `trRamData` register, while memory in SMEM mode should be read as any other memory on system bus - for example using SBA (System Bus Access) access mode as defined in the RISC-V Debug Specification. Trace data is placed in memory in LSB order (first byte of trace packet/data is placed on LSB). Be aware that in case trace memory wraps around some protocols may require additional synchronization data - it is usually done by periodically generating a sequence of alignment synchronization bytes which cannot be part of any valid packet. Specification of each trace protocol must define it. __Table 1\. **Register: trRamControl: Trace RAM Sink Control Register (trBaseRam+0x000)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------ | --------------- | | 0 | trRamActive | Primary activate/reset bit for Trace RAM Sink. When 0, the Trace RAM Sink may have clocks gated off or be powered down, and other register locations may be inaccessible. Hardware may take an arbitrarily long time to process power-up and power-down and will indicate completion when the read value of this bit matches what was written. See [Reset and Discovery](#Reset and Discovery) chapter for more details. | RW | 0 | | 1 | trRamEnable | **1:** Trace RAM Sink enabled. Setting trRamEnable to 0 flushes any queued trace data to memory (idle bytes/packet may be appended after the last message/packet to assure memory access alignment). See [Enabling and Disabling](#Enabling and Disabling) chapter for more details. Enabling trace CANNOT change any of trRamStart/Limit/WP/RP?? registers. Disabling trace may update trRamWP?? because of flushing. | RW | 0 | | 2 | — | Reserved | — | 0 | | 3 | trRamEmpty | Reads 1 when Trace RAM Sink internal buffers are empty, which means that all trace data is flushed. | RO | 1 | | 4 | trRamMode | **0:** This RAM Sink will operate in SRAM mode **1:** This RAM Sink will operate in SMEM mode | WARL | [Undef](#Undef) | | 7:5 | — | Reserved | — | 0 | | 8 | trRamStopOnWrap | **1:** Disable storing trace to RAM (trRamEnable → 0) when the circular buffer gets full. Sink should stop accepting new messages which may result in an overflow or stall condition at an encoder. | WARL | [Undef](#Undef) | | 10:9 | trRamMemFormat | **0:** Memory is formatted as plain bytes **1-2:** Reserved for future formats **3:** Reserved for custom memory format | WARL | [Undef](#Undef) | | 11 | — | Reserved | — | 0 | | 14:12 | trRamAsyncFreq | **0:** Alignment synchronization (Async) packets disabled (may be the only choice for some protocols) **1-7:** Different levels of alignment synchronization (bigger number, bigger distance).Details should be defined in the specification of each trace protocol. | WARL | [Undef](#Undef) | | 31:15 | — | Reserved | — | 0 | __Table 2\. **Register: trRamImpl: Trace RAM Sink Implementation Register (trBaseRamSink+0x004)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------- | ---------------------------------------------------------------------------------------------------- | ------ | --------- | | 3:0 | trRamVerMajor | Trace RAM Sink Component Major Version. Value 1 means the component is compliant with this document. | RO | 1 | | 7:4 | trRamVerMinor | Trace RAM Sink Component Minor Version. Value 0 means the component is compliant with this document. | RO | 0 | | 11:8 | trRamCompType | Trace RAM Sink Component Type (RAM Sink) | RO | 0x9 | | 12 | trRamHasSRAM | This RAM Sink supports SRAM mode | RO | [SD](#SD) | | 13 | trRamHasSMEM | This RAM Sink supports SMEM (System Memory) mode | RO | [SD](#SD) | | 23:14 | — | Reserved for future versions of this standard | — | 0 | | 31:24 | — | Reserved for vendor specific implementation details | — | [SD](#SD) | | | Single RAM Sink may support both SRAM and SMEM modes, but not both may be enabled at the same time. It is also possible to have more than one RAM Sink in a system. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 3\. **Register: trRamStartLow: Trace RAM Sink Start Register (trBaseRamSink+0x010)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------ | ----------------------------- | | 1:0 | — | Always 0 (two LSB of 32-bit address) | RO | 0 | | 31:2 | trRamStartLow | Byte address of start of trace sink circular buffer. It is always aligned on at least a 32-bit/4-byte boundary. An SRAM sink will usually have trRamStartLow fixed at 0. | WARL | [Undef](#Undef) or fixed to 0 | For a bus with an address larger than 32-bit, corresponding `High` registers define the MSB part of such a larger address. __Table 4\. **Register: trRamStartHigh: Trace RAM Sink Start High Bits Register (trBaseRamSink+0x014)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | -------------- | ----------------------------------------------- | ------ | --------------- | | 31:0 | trRamStartHigh | High order bits (63:32) of trRamStart register. | WARL | [Undef](#Undef) | __Table 5\. **Register: trRamLimitLow: Trace RAM Sink Limit Register (trBaseRamSink+0x018)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------ | --------------- | | 1:0 | — | Always 0 (two LSB of 32-bit address) | RO | 0 | | 31:2 | trRamLimitLow | Highest absolute 32-bit part of address of trace circular buffer. The trRamWP register is reset to trRamStart after a trace word has been written to this address. | WARL | [Undef](#Undef) | __Table 6\. **Register: trRamLimitHigh: Trace RAM Sink Limit High Bits Register (trBaseRamSink+0x01C)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | -------------- | ----------------------------------------------- | ------ | --------------- | | 31:0 | trRamLimitHigh | High order bits (63:32) of trRamLimit register. | WARL | [Undef](#Undef) | __Table 7\. **Register: trRamWPLow: Trace RAM Sink Write Pointer Register (trBaseRamSink+0x020)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 0 | trRamWrap | Set to 1 by hardware when trRamWP wraps. It is only set to 0 if trRamWPLow is written | WARL | [Undef](#Undef) | | 1 | — | Always 0 (bit B1 of 32-bit address) | RO | 0 | | 31:2 | trRamWPLow | Absolute 32-bit part of address in trace sink memory where next trace message will be written. Fixed to a natural boundary. After a trace word write occurs while trRamWP \= trRamLimit, trRamWP is set to trRamStart. | WARL | [Undef](#Undef) | __Table 8\. **Register: trRamWPHigh: Trace RAM Sink Write Pointer High Bits Register (trBaseRamSink+0x024)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ----------- | -------------------------------------------- | ------ | --------------- | | 31:0 | trRamWPHigh | High order bits (63:32) of trRamWP register. | WARL | [Undef](#Undef) | __Table 9\. **Register: trRamRPLow: Trace RAM Sink Read Pointer Register (trBaseRamSink+0x028)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------ | --------------- | | 1:0 | — | Always 0 (two LSB of 32-bit address) | RO | 0 | | 31:2 | trRamRPLow | Absolute 32-bit part of address in trace circular memory buffer visible through trRamData. trRamRP auto-increments following an access to trRamData. After a trace word read occurs while trRamRP \= trRamLimit, trRamRP is set to trRamStart. Required for SRAM mode and optional for SMEM mode. | WARL | [Undef](#Undef) | __Table 10\. **Register: trRamRPHigh: Trace RAM Sink Read Pointer High Bits Register (trBaseRamSink+0x02C)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | ----------- | -------------------------------------------- | ------ | --------------- | | 31:0 | trRamRPHigh | High order bits (63:32) of trRamRP register. | WARL | [Undef](#Undef) | __Table 11\. **Register: trRamData: Trace RAM Sink Data Register (trBaseRamSink+0x040)**__ | **Bit** | **Field** | **Description** | **RW** | **Reset** | | ------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------- | --------------- | | 31:0 | trRamData | Read (and optional write) value for trace sink memory access. SRAM is always accessed by 32-bit words through this path regardless of the actual width of the sink memory. Required for SRAM mode and optional for SMEM mode. | R or RW | [Undef](#Undef) | | | When trace capture was wrapped around (trRamWrap \= 1) beginning of trace is not available and oldest packets/messages in the trace buffer (starting at address in trRamWP) will most likely not be complete. Trace decoders must look for the start of a message. Also when trace is stopped on wrap around, the very last message recorded in trace memory may not be complete. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | The table below shows typical Trace RAM Sink configurations. Implementing other configurations is not suggested as trace tools may not support it without adjustments. __Table 12\. **Typical Trace RAM Sink Configurations**__ | **Mode** | **trRamStart** | **trRamLimit** | **trRamWP** | **trRamRP** | **trRamData** | | ------------ | ------------------- | ------------------------------------------------------------------------------- | ----------- | --------------- | --------------- | | SRAM | 0 | Hard coded to max size (2^M - A) at reset, but can be possibly trimmed | Required | Required | Required | | SMEM Generic | Any (2^N aligned) | Any (trRamStart \+ 2^M - A) - must be set by trace tool | Required | Not implemented | Not implemented | | SMEM Fixed | Fixed (2^N aligned) | Fixed to max size at reset (trRamStart \+ 2^M - A), but can be possibly trimmed | Required | Not implemented | Not implemented | | | Value A means alignment which depends on memory access width. If we have memory access width of 32-bits, A=4 and value of trRamLimit register should be 0x…​FC. Some implementations may impose bigger alignment of trace data (to allow more efficient transfer rates) for SMEM mode. For SRAM mode A must be 4 as access to trace via trRamData is always 32-bits wide. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#accessing-and-detecting-ram-sink-registers)Accessing and Detecting RAM Sink Registers Trace tool should start interacting with Trace RAM Sink by releasing RAM Sink from reset by setting `trRamActive` \= 1 and waiting for this bit to be set. After that it should verify `trRamEmpty` \= 1, read `trRamImpl` and verify `trRamCompType` and `trRamVer??` fields. Values of `trRamHasSRAM/SMEM` fields will provide main types of RAM Sink being implemented. Later `trRamMode` should be set (depending on desired RAM Sink mode). It is important to set this field first as other registers may behave differently for SRAM and SMEM modes. In SRAM mode, the trace memory is dedicated for trace storage and `trRamStart??` registers should not be settable (usually both not implemented and return 0). `trRamLimitLow` register may be either hardcoded (to reflect physical SRAM size) or writable (allowing trimming RAM size allowing faster wrap-around or sharing the same memory with some other components in the system). The `trRamLimitHigh` register should not be implemented as it is not practical to have more than 4GB of dedicated on-chip RAM storage. Detection of valid ranges of each `trRamStart??` and `trRamLimit??` registers should be performed by writing 0 and 0xFFFFFFFF. After setting 0, the lowest possible value must be set. After setting 0xFFFFFFFF the highest possible value must be set. If the highest value for `trRamStartHigh` or `trRamLimitHigh` is 0, it means the register is NOT implemented. Some implementations may provide different limits for different start addresses, so the trace tool should always set `trRamStart??` registers first - this option can be used when a particular implementation has two different RAM regions (each with different physical memory size). Not every value may be settable in `trRamStart/Limit` registers. Value written may be trimmed (for example aligned on a particular 2^N boundary) and a trace tool should verify values being written. In case accepted values are different from what was provided by the user, a message should be printed which may allow the user to adjust (possbly suboptimal) settings. Registers `trRamStart??` and `trRamLimit??` are usually set at the beginning of a debug/trace session and never changed. | | In SMEM mode (trRamMode \= 1) trace tool should never set trRamStart?? and trRamLimit?? registers outside of range provided by the user as otherwise raw trace being written to memory may corrupt running code and/or data or stack. This type of errors may be very difficult to diagnose as in complex system code (or data) being overwritten by trace may be used way, way later after actual corruption was made. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Having both `trRamStart/Limit??` registers set, the tool should try to set `trRamRP??` to the same value as `trRamLimit??`. If it is settable, it means that the `trRamData` register should be used to read the trace. Otherwise collected trace must be read using normal, physical memory accesses (in range defined by `trRamStart/Limit??` registers). Before enabling RAM Trace Sink (by setting `trRamEnable` \= 1) the trace tool should set `trRamWP??` registers (usually to the same values as in `trRamStart??` register). Enabling trace must NOT change any of `trRamStart/Size/WP/RP??` registers. Just after the trace is enabled `trRamWP??` may change because of trace being added to trace memory. After trace is enabled and active (`trRamEnable` \= 1 or `trRamEmpty` \= 0), the trace tool should NOT write any of `trRamStart/Limit/WP??` registers. Setting `trRamRP` and reading `trRamData` may be attempted while trace is active, but support for reading SRAM trace while trace is active may not always be implemented. In such a case write to `trRamRP` must be ignored and `trRamData` read must not advance `trRamRP`. Reading the trace in the SMEM mode via normal memory reads is always allowed. | | Even if reading trace (while trace is active) is implemented, circular trace buffer may be overwritten even several times, so values being read by trRamData will be of no use. However, when trace is started/stopped by infrequent triggers, reading SRAM trace may be useful. Also, the very last packet in memory may be incomplete as the last trace word may be buffered inside (and trRamEmpty \= 0 will be observed). | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | Trace RAM Sink may implement writing trace by writing to trRamData, but this mode is usable only for testing, so will most likely not be implemented. Trace tool is not required to support writing to the trRamData register. | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Trace System Overview ==================== ## [](#trace-system-overview)Trace System Overview This section briefly describes features of the Trace Encoder and other trace components as background for understanding some of the control interface register fields. ### [](#trace-encoder)Trace Encoder By monitoring the Ingress Port, the Trace Encoder determines when a program flow discontinuity has occurred and whether the discontinuity is inferable or non-inferable. An inferable discontinuity is one for which the Trace Decoder can statically determine the destination, such as a direct branch instruction in which the destination or offset is included in the opcode. Non-inferable discontinuities include all other types such as interrupts, exceptions, and indirect jump instructions. #### [](#branch-trace-messaging)Branch Trace Messaging Branch Trace Messaging is the simplest, baseline form of instruction trace. Each program counter discontinuity results in one trace message, either a Direct or Indirect Branch Message. Linear instructions (or sequences of linear instructions) do not directly generate any trace messages/packets but overflow of counters (or exceptions) may generate corresponding packets/messages - these messages are infrequent and will not affect trace compression. Indirect Branch Messages normally contain a compressed address to reduce bandwidth. The Trace Encoder emits a Branch With Sync Message containing the complete instruction address under certain conditions. This message type is a variant of the Direct or Indirect Branch Message and includes a full address and a field indicating the reason for the Sync. #### [](#branch-history-messaging)Branch History Messaging Both the Efficient Trace for RISC-V (E-Trace) Specification and the RISC-V N-Trace (Nexus-based Trace) specification define systems of messages intended to improve compression by reporting only whether conditional branches are taken by encoding each branch outcome in a single taken/not-taken bit. The destinations of non-inferable jumps and calls are reported as compressed addresses. Much better optimized compression can be achieved, but an encoder implementation will typically require more hardware. #### [](#other-optimizations)Other Optimizations Several other optimizations are possible to improve trace compression. These are optional for any Trace Encoder and there should be a way to disable optimizations in case the trace system is used with code that does not follow recommended API rules. Examples of optimizations are a Return-address stack, Branch repetition, Statically inferable jump, and Branch prediction. ### [](#trace-sinks)Trace Sinks The Trace Encoder transmits completed messages to a Trace Sink. This specification defines a number of different sink types, all optional, and allows an implementation to define other sink types. A Trace Encoder must have at least one sink or funnel attached to it. | | Trace messages/packets are sequences of bytes. In case of wider sink width, some padding/idle bytes (or additional formatting) may be added by the sink. N-Trace format allows any number of idle bytes between messages. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | #### [](#sram-sink)SRAM Sink The Trace Encoder packs trace messages into fixed-width trace words (usually bytes). These are then stored in a dedicated RAM, typically located on-chip, in a circular-buffer fashion. When the RAM has filled, it may optionally be stopped, or it may wrap and overwrite earlier trace data. #### [](#system-memory-sink)System Memory Sink The Trace Encoder packs trace messages into fixed-width trace words (usually bytes). These are then stored in a range of system memory reserved for trace using a DMA-type bus controller in a circular-buffer fashion. When the memory range has been filled, it may optionally be stopped, or it may wrap and overwrite earlier trace data. This type of sink may also be used to transmit trace off-chip through, for example, a PCIe or USB port. #### [](#pib-sink)PIB Sink The Trace Encoder sends trace messages to the PIB Sink. Each message is transmitted off-chip (as sequence of bytes) using a specific protocol described later. ### [](#atb-bridge)ATB Bridge The ATB Bridge allows sending RISC-V trace to Arm CoreSight infrastructure (instead of RISC-V compliant sink defined in this document) as an ATB initiator. ATB Bridge is not needed for RISC-V only systems. ATB width is byte aligned (8, 16, 32, 64, 128) which allows transport of trace messages/packets defined as sequence of bytes. ### [](#trace-funnel)Trace Funnel The Trace Encoder may send trace messages to a Trace Funnel. The Funnel aggregates the trace from each of its inputs (either RISC-V Trace Encoder or another Trace Funnel) and sends the combined trace stream to its designated Trace Sink or ATB Bridge, which is one or more of the sink types above. | | It is assumed that each input to the funnel (Trace Encoder or another Trace Funnel) has a unique message source ID defined (trTeSrcID field in the trTeControl register). | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Versioning of Components ==================== ## [](#versioning-of-components)Versioning of Components Each component has a `tr??Impl` register, which includes two 4-bit `tr??VerMinor` and `tr??VerMajor` fields. These fields are guaranteed to be present in all future revisions of a standard, so trace tools will be able to discover a component version and act accordingly. * Value 0 as `tr??VerMajor` is NOT allowed (due to compatibility reasons). * Different components may report different versions (as some components may be updated more often than others). * The major version `tr??VerMajor` field is incremented when the modification breaks backward compatibility. * The minor version `tr??VerMinor` field is incremented when the modification maintains backward compatibility (for example adding a new field) - for that reason software should always write 0 to reserved bits in registers. * Version 15.x is reserved for non-compatible version encoding. * Version n.15 should be used as experimental (in development) implementation. Software tools must report the version number as two decimal numbers _major.minor_ \- initial version of this specification is defined as **_1.0_**. | | Trace software should handle versions as follows (let’s assume hypothetical version 2.3 was defined as current version in moment of release of trace software) 0.x ⇒ Reject as not supported or generate a warning and handle as pre-ratified/initial version 0. 2.3 ⇒ Accept silently. 2.2 ⇒ Accept silently (and trim features or not allow users to set newer features). 2.4 ⇒ Generate a warning but continue using 2.3 features. 2.15 ⇒ Generate an "experimental version" warning but continue using 2.3 features. 1.x ⇒ Generate a warning and continue or reject as an obsolete (referring to last debugger supporting this version). 3.x ⇒ Generate a fatal error that this future version is not compatible with existing software and possibly redirect to the tool update page. Displayed messages should report component name, component base address and current and supported version numbers. It is suggested to display the full hexadecimal value of tr??Impl register as it may aid in debugging of possibly incorrect/incompatible component configuration. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC-BY-SA-4.0). The full license text is available at. Copyright 2024 by RISC-V International. RISC-V Semihosting ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-semihosting)RISC-V Semihosting Version 1.0, 21st February 2025: This document is Ratified. | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Bibliography ==================== ## [](#bibliography)Bibliography \[1\] _Semihosting for AArch32 and AArch64 2023Q3_, 2023\. \[Online\]. Available: \[2\] _The RISC-V Instruction Set Manual Volume I: Unprivileged Architecture_, 2024\. \[Online\]. Available: 2.1. RISC-V Semihosting Binary Interface ==================== ## [](#2-1-risc-v-semihosting-binary-interface)2.1\. RISC-V Semihosting Binary Interface The RISC-V semihosting binary interface consist of a breakpoint instruction sequence and a mechanism to pass parameters which are defined by the following sub-sections. ### [](#2-1-1-semihosting-breakpoint-instruction-sequence)2.1.1\. Semihosting Breakpoint Instruction Sequence Semihosting operations are requested using a sequence of instructions including `EBREAK`. Because the RISC-V base ISA does not provide more than one `EBREAK` instruction, RISC-V semihosting uses a special sequence of instructions to distinguish a semihosting `EBREAK` from a debugger inserted`EBREAK`. The [RISC-V Semihosting Breakpoint Sequence](#breakpoint%5Finsns) shows the instruction sequence used to invoke a semihosting operation. RISC-V Semihosting Breakpoint Sequence slli x0, x0, 0x1f # 0x01f01013 Entry NOP ebreak # 0x00100073 Break to debugger srai x0, x0, 7 # 0x40705013 Exit NOP These three instructions must be 32-bit wide instructions. This sequence is applicable to all RISC-V base ISAs. If address translation and protection is enabled for the semihosting caller then the semihosting instruction sequence and data passed via memory must be paged in else the behavior of the semihosting call is UNSPECIFIED. | | The SLLI, EBREAK, and SRAI instructions are part of the ratified RV32E, RV32I, RV64E and RV64I (aka Base Integer Instruction Set) specifications \[[2](bibliography.html#bib-riscvunprivref)\] hence these instructions are present on almost all RISC-V platforms. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The SLLI and SRAI instruction based NOPs which serve as semihosting marker have been randomly selected from the Base Integer Instruction Set since these are designated for custom use and unlikely to appear in real life code. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ### [](#2-1-2-semihosting-parameters)2.1.2\. Semihosting Parameters The type of semihosting operation and its parameters are specified using general purpose registers. The OPERATION NUMBER is specified in the `a0`register, and the PARAMETER is specified in the `a1` register, whereas the RETURN VALUE is available in the `a0` register. All registers and data block fields are XLEN wide. Contributors ==================== ## [](#contributors)Contributors This RISC-V specification has been contributed to directly or indirectly by: * Krste Asanovic <[krste@sifive.com](mailto:krste@sifive.com)\> * Palmer Dabbelt <[palmer@dabbelt.com](mailto:palmer@dabbelt.com)\> * Liviu Ionescu <[ilg@livius.net](mailto:ilg@livius.net)\> * Keith Packard <[keith.packard@sifive.com](mailto:keith.packard@sifive.com)\> * Megan Wachs <[megan@sifive.com](mailto:megan@sifive.com)\> * Anup Patel <[apatel@ventanamicro.com](mailto:apatel@ventanamicro.com)\> * Ved Shanbhogue <[ved@rivosinc.com](mailto:ved@rivosinc.com)\> 1.1. Introduction ==================== ## [](#intro)1.1\. Introduction Semihosting is a technique where an application running in a debug or simulation environment can access elements of the system hosting the debugger or simulator including console, file system, time and other functions. This allows for diagnostics, interaction and measurement of a target system without requiring significant infrastructure to exist in that target environment. The RISC-V semihosting specification adopts the design of the ARM semihosting specification \[[1](bibliography.html#bib-armsemihostingref)\] to minimize the development effort. The services defined by the ARM semihosting specification \[[1](bibliography.html#bib-armsemihostingref)\] are portable across different architectures, and only the mechanism of invoking a semihosting service (aka semihosting binary interface) is archicture specific. The [Figure 1](#fig%5Fintro1) below shows an architecture independent high-level view of semihosting usage. The RISC-V semihosting specification only defines the semihosting binary interface for RISC-V platforms and all other aspects of semihosting are defined by the ARM semihosting specification \[[1](bibliography.html#bib-armsemihostingref)\]. ![intro image1](_images/intro-image1.png) Figure 1\. Generic Semihosting Usage Flow Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at. Copyright 2024 by RISC-V International. RISC-V Boot and Runtime Services Specification (BRS) ==================== ![risc v logo](_images/risc-v_logo.svg) ## [](#risc-v-boot-and-runtime-services-specification-brs)RISC-V Boot and Runtime Services Specification (BRS) BRS Task Group Version 1.0, 29th August 2025: This document is Ratified. | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | 6.1. BRS-I ACPI Requirements ==================== ## [](#acpi)6.1\. BRS-I ACPI Requirements The _Advanced Configuration and Power Interface Specification_ provides the OS-centric view of system configuration, various hardware resources, events and power management. This section defines the BRS-I mandatory and optional ACPI requirements on top of existing ACPI \[[3](bibliography.html#bib-acpi)\] and UEFI \[[10](bibliography.html#bib-uefi)\] specification requirements. Additional non-normative guidance may be found in the [firmware implementation guidance](#acpi-guidance)section. | | All content in this section is optional and recommended for BRS-B. | | --------------------------------------------------------------------- | | ID# | Rule | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ACPI\_010 | Be 64-bits clean. RSDT MUST NOT be implemented, with RsdtAddress in RSDP set to 0. 32-bit address fields MUST be 0. | | _[See additional guidance](#acpi-guidance-64bit-clean)._ | | | ACPI\_020 | MUST implement the hardware-reduced ACPI mode (no FACS table). | | _[See additional guidance](#acpi-guidance-hw-reduced)._ | | | ACPI\_030 | The Processor Properties Table (PPTT) MUST be implemented, even on systems with a simple hart topology. | | ACPI\_040 | The PCI Memory-mapped Configuration Space (MCFG) table MUST NOT be present if it violates \[[17](bibliography.html#bib-pcifw)\]. | | _Only compatible PCIe segments, exposed via ECAM (Enhanced Configuration Access Mechanism), may be described in the MCFG. The MCFG MUST NOT require vendor-specific OS support. See PCI Services (\[[3](bibliography.html#bib-acpi)\], Section 4) for more ACPI requirements relating to PCIe support. [See additional guidance](#acpi-guidance-pcie)._ | | | ACPI\_050 | A Serial Port Console Redirection Table \[[19](bibliography.html#bib-spcr)\] MUST be present on systems, where the graphics hardware is not present or not made available to an OS loader via the standard UEFI EFI\_GRAPHICS\_OUTPUT\_PROTOCOL interface. | | _In these cases, the table provides essential configuration for an early OS boot console._ | | | ACPI\_060 | An SPCR table, if present, MUST meet the following requirements: Revision 4 or later of SPCR. For NS16550-compatible UARTs: Use Interface Type 0x12 (16550-compatible with parameters defined in Generic Address Structure). There MUST be a matching AML device object with \_HID (Hardware ID) or \_CID (Compatible ID) RSCV0003. | | _See [additional guidance](#acpi-guidance-spcr)_. | | ### [](#acpi-aml)6.1.1\. BRS-I ACPI Methods and Objects This section lists additional requirements for ACPI methods and objects. [See additional guidance](#acpi-guidance-aml). | ID# | Rule | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | AML\_010 | The Current Resource Setting (\_CRS) device method for a PCIe Root Complex SHOULD NOT return any descriptors for I/O ranges (such as created by ASL macros WordIO, DWordIO, QWordIO, IO, FixedIO, or ExtendedIO). | | _Legacy PCI I/O BARs are uncommon in modern PCIe devices and support for PCI I/O space may complicate configuration of PCIe Root Complex hardware in a compliant manner._ | | | AML\_020 | The Possible Resource Settings (\_PRS) and Set Resource Settings (\_SRS) device method SHOULD NOT be implemented. | | _ACPI resource descriptors are typically used to describe devices with fixed I/O regions that do not change. Flexible resource assignment is not supported by most modern ACPI OSes._ | | | AML\_030 | Per-hart device objects MUST be defined under \\\_SB (System Bus) namespace and not in the deprecated \\\_PR (Processors) namespace. | | AML\_040 | Systems supporting OS-directed hart performance control and power management MUST expose these via Collaborative Processor Performance Control (CPPC, \[[3](bibliography.html#bib-acpi)\] Section 8.4.6). | | AML\_050 | Processor idle states MUST be described using Low Power Idle (\_LPI, \[[3](bibliography.html#bib-acpi)\] Section 8.4.3). | | AML\_060 | Systems with a Real-Time Clock on an OS-managed bus (e.g. I2C, subject to arbitration issues due to access to the bus by the OS) MUST implement the Time and Alarm Device (TAD) with functioning \_GRT and \_SRT methods, and the \_GCP method returning bit 2 set (i.e. get/set real time features implemented). | | _Also see [URT\_020](#uefi-rtc)_. | | | AML\_070 | Systems implementing a TAD MUST be functional without additional system-specific OS drivers. | | _In situations where the Time and Alarm Device (TAD) depends on a vendor-specific OS driver for correct function (SPI, I2C, etc), the TAD MUST be functional if the OS driver is not loaded. That is, when a dependent driver is loaded, an AML method switches further accesses to go through the driver-backed OperationRegion._ | | | AML\_080 | PLIC and APLIC device objects MUST support the Global System Interrupt Base (\_GSB, \[[3](bibliography.html#bib-acpi)\] Section 6.2.7) object.[See additional guidance](#acpi-guidance-gsi-namespace). | | AML\_090 | UART device objects with ID RSCV0003 MUST implement [Properties for UART Devices](#acpi-props-uart). | | AML\_100 | PLIC/APLIC namespace devices MUST be present in the ACPI namespace whenever corresponding MADT entries are present. [See RVI ACPI IDs](#acpi-ids). | | _Also see [AML\_080](#acpi-irq-gsb) and [additional guidance](#acpi-guidance-gsi-namespace)_. | | ### [](#acpi-ids)RVI-specific ACPI IDs ACPI ID is used in the `_HID` (Hardware ID), `_CID` (Compatible ID) or`_SUB` (Subsystem ID) objects as described in the ACPI Specification for devices, that do not have a standard enumeration mechanism. The ACPI ID consists of two parts: a vendor identifier followed by a product identifier. Vendor IDs consist of 4 characters, each character being either an uppercase letter (A-Z) or a numeral (0-9). The vendor ID SHOULD be unique across the Industry and registered by the UEFI forum. For RVI standard devices, `RSCV` is the vendor ID registered. Vendor-specific devices can use an appropriate vendor ID registered for the manufacturer. Product IDs are always four-character hexadecimal numbers (0-9 and A-F). The device manufacturer is responsible for assigning this identifier to each product model. This document contains the canonical list of ACPI IDs for the namespace devices that adhere to the RVI specifications. The RVI task groups may make pull requests against this repository to request the allocation of ACPI ID for any new device. | ACPI ID | Device | | ----------------------------------------------------------------------- | ------------------------------------------------------------------------- | | RSCV0001 | RISC-V Platform-Level Interrupt Controller (PLIC) | | RSCV0002 | RISC-V Advanced Platform-Level Interrupt Controller (APLIC) | | RSCV0003 | NS16550 UART compatible with an SPCR definition using Interface Type 0x12 | | RSCV0004 | RISC-V IOMMU implemented as a platform device | | RSCV0005 | RISC-V SBI Message Proxy (MPXY) Mailbox Controller | | RSCV0006 | RISC-V RPMI System MSI Interrupt Controller | | _Also see [ACPI Device Properties for UART Devices](#acpi-props-uart)._ | | ### [](#acpi-props)RVI-specific ACPI Device Properties This section is used to define the `_DSD` device properties \[[20](bibliography.html#bib-dsd)\] in the `rscv-` namespace. Where explicit values are provided in a property definition, only these values must be used. System behavior with any other values is undefined. | Property | Type | Description | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---- | ----------- | | _Currently, there are no properties defined in the rscv- namespace. Request for new property names in the rscv- namespace should be made as a git pull request to this table._ | | | ### [](#acpi-props-uart)ACPI Device Properties for UART Devices Generic 16550-compatible UART devices can have device properties in the global name space since Operating Systems are already using them. | Property | Type | Description | | ---------------------------------------------------------------------------------------------------- | ------- | ------------------------------------------------------------------------------------ | | clock-frequency | Integer | Clock feeding the IP block in Hz. | | _A value of zero will preclude the ability to set the baud rate, or to configure a disabled device._ | | | | reg-offset | Integer | Offset to apply to the register map base address from the start of the registers. | | reg-shift | Integer | Quantity to shift the register offsets by. | | reg-io-width | Integer | The size (in bytes) of the register accesses that should be performed on the device. | | _1, 2, 4 or 8._ | | | | fifo-size | Integer | The FIFO size (in bytes). | Bibliography ==================== ## [](#bibliography)Bibliography \[1\] _Key words for use in RFCs to Indicate Requirement Levels_. \[Online\]. Available: \[2\] _PCI Express® Base Specification Revision 6.0_, . \[Online\]. Available: \[3\] _Advanced Configuration and Power Interface Specification 6.6_. \[Online\]. Available: \[4\] _DeviceTree_. \[Online\]. Available: \[5\] _Embedded Base Boot Requirements Specification 2.1.0_. \[Online\]. Available: \[6\] _RISC-V Profile_. \[Online\]. Available: \[7\] _RISC-V Platform Management Interface Specification_. \[Online\]. Available: \[8\] _RISC-V Supervisor Binary Interface Specification_. \[Online\]. Available: \[9\] _System Management BIOS (SMBIOS) Reference Specification 3.7.0_. \[Online\]. Available: \[10\] _Unified Extensible Firmware Interface Specification 2.11_. \[Online\]. Available: \[11\] _RISC-V "stimecmp / vstimecmp" Extension_, 2021\. \[Online\]. Available: \[12\] _The RISC-V Advanced Interrupt Architecture_, 2023\. \[Online\]. Available: \[13\] _RISC-V Indirect CSR Access (Smcsrind/Sscsrind)_, 2023\. \[Online\]. Available: \[14\] _RISC-V Supervisor Counter Delegation Specification (Smcdeleg/Ssccfg)_, 2024\. \[Online\]. Available: \[15\] _UEFI memory mitigations_. \[Online\]. Available: \[16\] _UEFI Platform Initialization Specification 1.9_. \[Online\]. Available: \[17\] _PCI Firmware Specification Revision 3.3_. \[Online\]. Available: \[18\] _TCG EFI Platform Specification_. \[Online\]. Available: \[19\] _Serial Port Console Redirection Table (SPCR)_. \[Online\]. Available: \[20\] _\_DSD (Device Specific Data) Implementation Guide_. \[Online\]. Available: Contributors ==================== ## [](#contributors)Contributors This RISC-V specification has been contributed to directly or indirectly by (in alphabetical order): Aaron Durbin, Andrei Warkentin, Andrew Jones, Anup Patel, Atish Patra, Beeman Strong, Darius Rad, Heinrich Schuchardt, Haibo Xu, Jamie Iles, John Hauser, Radim Krčmář, Rahul Pathak, Paul Walmsley, Samuel Holland, Sia Jee Heng, Sunil V L, Vedvyas Shanbhogue 3.1. Hart Requirements ==================== ## [](#hart)3.1\. Hart Requirements A compliant system includes a RISC-V application processor and the requirements in this section apply solely to harts in the application processors of a system. The BRS specification is minimally prescriptive on the RISC-V hart requirements. It is anticipated that detailed requirements will be driven by target market segment and product/solution requirements. | ID# | Rule | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- | | HR\_010 | The RISC-V application processor harts MUST be compliant to RVA20S64 profile \[[6](bibliography.html#bib-profile)\]. | | _The BRS governs the interactions between 64-bit OS supervisor-mode software and 64-bit firmware. These are minimum requirements allowing for the wide variety of existing and future hart implementations to be supported. It is expected that operating systems and hypervisors may impose additional profile/ISA requirements, depending on the use-case and application._ | | 1.1. Introduction ==================== ## [](#intro)1.1\. Introduction The _RISC-V Boot and Runtime Services Specification_ (BRS) defines a standardized set of software capabilities, that portable system software, such as operating systems and hypervisors, can rely on being present in an implementation to utilize in acts of device discovery, OS boot and hand-off, system management, and other operations. The BRS specification is targeting systems that implement S/U privilege modes, and optionally the HS privilege mode. This is the expected deployment for OSVs and system vendors in a typical ecosystem covering client systems up through server systems where software is provided by different vendors than the system vendor. This specification standardizes the requirements for software interfaces and capabilities by building on top of relevant industry and ratified RISC-V standards. ### [](#1-1-1-releases)1.1.1\. Releases It is expected that the BRS will periodically release a new specification. The determination of a new release will be based on the evaluation of significant changes to its underlying dependencies. ### [](#1-1-2-approach-to-solutions)1.1.2\. Approach to Solutions The BRS focuses on two solutions in the form of what is deemed a recipe. Each recipe contains the requirements needed to fulfill each solution. The requirements of each recipe will be marked accordingly with a unique identifier. The recipes are BRS-I (Interoperable) and BRS-B (Bespoke). ### [](#1-1-3-testing-and-conformance)1.1.3\. Testing and Conformance To be compliant with this specification, an implementation MUST support all mandatory rules and MUST support the listed versions of the specifications. This standard set of capabilities MAY be extended by a specific implementation with additional standard or custom capabilities, including compatible later versions of listed standard specifications. Portable system software MUST support the specified mandatory capabilities to be compliant with this specification. The rules in this specification use the following format: | ID# | Rule | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | CAT\_NNN | The CAT is a category prefix that logically groups the rules and is followed by 3 digits - NNN \- assigning a numeric ID to the rule. The rules use the key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" that are to be interpreted as described in RFC 2119 \[[1](bibliography.html#bib-rfc%5F2119)\] when, and only when, they appear in all capitals, as shown here. When these words are not capitalized, they have their normal English meanings. | | _A rule or a group of rules may be followed by non-normative text providing context or justification for the rule. The non-normative text may also be used to reference sources that are the origin of the rule._ | | ### [](#1-1-4-glossary)1.1.4\. Glossary Most terminology has the standard RISC-V meaning. This table captures other terms used in the document. Terms in the document prefixed by **PCIe** have the meaning defined in the _PCI Express Base Specification_ \[[2](bibliography.html#bib-pci)\] (even if they are not in this table). __Table 1\. Terms and definitions__ | Term | Definition | | ------- | ------------------------------------------------------------------------------------------------ | | ACPI | _Advanced Configuration and Power Interface Specification_ \[[3](bibliography.html#bib-acpi)\]. | | BRS | _RISC-V Boot and Runtime Services Specification_. This document. | | BRS-I | Boot and Runtime Services recipe targeting interoperation across different software suppliers. | | BRS-B | Boot and Runtime Services recipe using a bespoke solution. | | DT | _Device Tree_ \[[4](bibliography.html#bib-dt)\]. | | EBBR | _Embedded Base Boot Requirements Specification_ \[[5](bibliography.html#bib-ebbr)\]. | | OSV | Operating System Vendor. | | OS | Operating System or Hypervisor. | | Profile | _RISC-V Profile_ \[[6](bibliography.html#bib-profile)\]. | | RPMI | _RISC-V Platform Management Interface_ \[[7](bibliography.html#bib-rpmi)\]. | | RVI | RISC-V International. | | SBI | _RISC-V Supervisor Binary Interface Specification_ \[[8](bibliography.html#bib-sbi)\]. | | SMBIOS | _System Management BIOS (SMBIOS) Reference Specification_ \[[9](bibliography.html#bib-smbios)\]. | | SoC | System on a chip, a combination of processor and supporting chipset logic in single package. | | UEFI | _Unified Extensible Firmware Interface Specification_ \[[10](bibliography.html#bib-uefi)\]. | 2.1. Recipes ==================== ## [](#recipes)2.1\. Recipes In this context, a recipe is a collection of firmware specification requirements that hardware, firmware, and software providers can implement to increase the likelihood that software written to the recipe will run predictably on all conforming devices. The BRS specification defines two recipes: BRS-I (for "Interoperable") and BRS-B (for "Bespoke"). ### [](#2-1-1-brs-i-recipe)2.1.1\. BRS-I Recipe The BRS-I recipe aims to simplify end-user experiences, software compatibility and OS distribution, by defining a common specification for boot and runtime interfaces. BRS-I is expected to be used by general-purpose compute devices such as servers, desktops, laptops and other devices with industry expectations on silicon vendor, OS and software ecosystem interoperability. BRS-I enables operating system providers to build a single **generic** operating system image that can be**successfully booted** on compliant systems. **Generic** means not requiring system-specific customizations - only an implementation of BRS-I requirements. **Successfully boot** means basic system configuration, sufficient for detecting the need for system-specific drivers and loading such drivers. It is understood that systems will deliver features beyond those covered by BRS-I. However, software written against a specific version of BRS-I must run, unaltered, without **anomalous and unexpected behavior** on systems that include such features and that are compliant to that specific version of BRS-I. Such behavior, caused by factors entirely unknown to a generic OS, is hard to diagnose and always results in a terrible user experience that negatively affects the value of the whole RISC-V standards-based ecosystem. **Anomalous and unexpected behavior** is taken to mean system instability and worst-case behavior for non-specialized workloads, but does not include suboptimal/unoptimized behavior or missing I/O or accelerator drivers. Any additional firmware features that cause anomalous and unexpected behavior must be disabled by default, and only enabled by opt-in. [See additional guidance](#recipe-brs-i-guidance). __Table 1\. BRS-I Recipe Overview__ | Profile | UEFI | ACPI | DT | SBI | SMBIOS | | ------------ | -------- | ------- | ----------------- | ------- | --------- | | \>= RVA20S64 | \>= 2.10 | \>= 6.6 | optional, >= v0.3 | \>= 2.0 | \>= 3.7.0 | ### [](#2-1-2-brs-b-recipe)2.1.2\. BRS-B Recipe BRS-B is intended for cases where only a minimal level of firmware interaction is mandated, focusing primarily on the boot process. The BRS-B recipe is the simpler of the two BRS recipes. It is expected to be used by software that is tailored to specific devices. Examples include many types of mobile devices, devices with real time response requirements, or embedded devices running rich operating systems with custom distributions. __Table 2\. BRS-B Recipe Overview__ | Profile | UEFI | ACPI | DT | SBI | SMBIOS | | ------------ | -------------------------------------------------- | ---------------- | ----------------- | ------- | ------------------ | | \>= RVA20S64 | EBBR, >= 2.1.0 \[[5](bibliography.html#bib-ebbr)\] | optional, >= 6.6 | optional, >= v0.3 | \>= 2.0 | optional, >= 3.7.0 | Either ACPI or DT may be used to describe hardware to the OS, but never both at the same time. 4.1. SBI Requirements ==================== ## [](#sbi)4.1\. SBI Requirements The _RISC-V Supervisor Binary Interface Specification_ (SBI) \[[8](bibliography.html#bib-sbi)\] defines an interface between the supervisor mode and the next higher privilege mode. This section defines the mandatory SBI version and extensions implemented by the higher privilege mode in order to be compatible with this specification. | ID# | Rule | | -------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- | | SBI\_010 | The SBI implementation MUST conform to SBI v2.0 or later. | | SBI\_020 | The SBI implementation MUST implement the Hart State Management (HSM) extension. | | _HSM is used by an OS for starting up, stopping, suspending and querying the status of secondary harts._ | | Certain requirements are conditional on the presence of RISC-V ISA extensions or system features. | ID# | Rule | | -------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | SBI\_030 | The Timer Extension (TIME) MUST be implemented, if the RISC-V "stimecmp / vstimecmp" Extension (Sstc, \[[11](bibliography.html#bib-sstc)\]) is not available. | | SBI\_040 | The S-Mode IPI Extension (sPI) MUST be implemented, if the Incoming MSI Controller (IMSIC, \[[12](bibliography.html#bib-aia)\]) is not available. | | SBI\_050 | The RFENCE Extension (RFNC) extension MUST be implemented, if the Incoming MSI Controller (IMSIC, \[[12](bibliography.html#bib-aia)\]) is not available. | | SBI\_060 | The Performance Monitoring Extension (PMU) MUST be implemented, if the counter delegation-related S-Mode ISA extensions (Sscsrind \[[13](bibliography.html#bib-sscsrind)\] and Ssccfg \[[14](bibliography.html#bib-smcdeleg)\]) are not present. | | SBI\_070 | The Debug Console Extension (DBCN) MUST be implemented if the ACPI SPCR table references Interface Type 0x15. | 7.1. BRS-I SMBIOS Requirements ==================== ## [](#smbios)7.1\. BRS-I SMBIOS Requirements The _System Management BIOS (SMBIOS) Reference Specification_ defines a standard format for presenting management information about an implementation, mostly focusing on hardware components. This section defines the BRS-I mandatory and optional SMBIOS requirements on top of existing \[[9](bibliography.html#bib-smbios)\] specification requirements. Additional non-normative guidance may be found in the [firmware implementation guidance](#smbios-guidance) section. | | All content in this section is optional and recommended for BRS-B. | | --------------------------------------------------------------------- | | | The structures and fields in this section are defined in a manner consistent with the DMTF specification language (\[[9](bibliography.html#bib-smbios)\]). | | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | ID# | Rule | | ------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | SMBIOS\_010 | A Baseboard/Module Information (Type 02) structure SHOULD be implemented. | | _This relaxes the SMBIOS specification requirement._ | | | SMBIOS\_020 | Processor Information (Type 04) structures, meeting the additional [7.1.1\. Type 04 Processor Information](#smbios-type04) clarifications, MUST be implemented. | | _This supersedes the RISC-V specific language in the SMBIOS specification (\[[9](bibliography.html#bib-smbios)\], Section 7.5.3.5)._ | | | SMBIOS\_030 | Port Connector Information (Type 08) structures SHOULD be implemented. | | SMBIOS\_040 | BIOS Language Information (Type 13) structures SHOULD be implemented. | | SMBIOS\_050 | An IPMI Device Information (Type 38) structure MUST be implemented, when an IPMIv1.0 host interface is present. | | SMBIOS\_060 | System Power Supply (Type 39) structures SHOULD be implemented. | | SMBIOS\_070 | Onboard Devices Extended Information (Type 41) structures SHOULD be implemented. | | SMBIOS\_080 | A Redfish Host Interface (Type 42) structure MUST be implemented, when a Redfish host interface is present. | | SMBIOS\_090 | A TPM Device (Type 43) structure MUST be implemented, when a TPM is present. | | SMBIOS\_100 | Processor Additional Information (Type 44) structures MUST be implemented. | | _See the [structure definitions below](#smbios-type44)_. | | | SMBIOS\_110 | Firmware Inventory Information (Type 45) structures SHOULD be implemented. | ### [](#smbios-type04)7.1.1\. Type 04 Processor Information | | The information in this section supersedes the definitions in (\[[9](bibliography.html#bib-smbios)\], Section 7.5.3.4). | | -------------------------------------------------------------------------------------------------------------------------- | A processor is a grouping of harts in a physical package. In modern designs this MAY mean an SoC. For RISC-V class CPUs, the `Processor ID` field contains two `DWORD`\-formatted values describing the overall physical processor package vendor and version. For some implementations this may also be known as the SoC ID. The first `DWORD` (offsets 08h-0Bh) is the JEP-106 code for the vendor, where bits 6:0 is the ID without the parity and bits 31:7 represent the number of continuation codes. The second `DWORD` (offsets 0Ch-0Fh) reflects vendor-specific part versioning. For hart-specific vendor and revision information, please see [7.1.2\. Type 44 Processor-Specific Data](#smbios-type44). ### [](#smbios-type44)7.1.2\. Type 44 Processor-Specific Data The processor-specific data structure fields are defined to follow the standard Processor-Specific Block fields (\[[9](bibliography.html#bib-smbios)\], Section 7.45.1). The structure is valid for processors declared with `Processor Type` 07h (64-bit RISC-V) only. A Type 44 structure needs to be provided for every hart meeting [\[hart\]](#hart) requirements. | Offset | Version | Name | Length | Value | Description | | ------ | ------- | ------------------------- | ------ | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 00h | 0100h | Revision | WORD | Varies | See [7.1.3\. Processor-Specific Data Structure Versioning](#smbios-psd-ver). | | 02h | 0100h | Hart ID | QWORD | Varies | The ID of this RISC-V hart. | | 0Ah | 0100h | Machine Vendor ID | QWORD | Varies | The vendor ID of this RISC-V hart. | | 12h | 0100h | Machine Architecture ID | QWORD | Varies | Base microarchitecture of the hart. Value of 0 is possible to indicate the field is not implemented. The combination of Machine Architecture ID and Machine Vendor ID should uniquely identify the type of hart microarchitecture that is implemented. | | 1Ah | 0100h | Machine Implementation ID | QWORD | Varies | Unique encoding of the version of the processor implementation. | ### [](#smbios-psd-ver)7.1.3\. Processor-Specific Data Structure Versioning The processor-specific data structure begins with a revision field to allow for future extensibility in a backwards-compatible manner. The minor revision is to be incremented anytime new fields are added in a backwards-compatible manner. The major revision is to be incremented on backwards-incompatible changes. | Version | Bits 15:8+ Major revision | Bits 7:0+ Minor revision | Combined | Description | | ------- | ------------------------- | ------------------------ | -------- | ---------------------------- | | v1.0 | 01h | 00h | 0100h | First BRS-defined definition | 5.1. BRS-I UEFI Requirements ==================== ## [](#uefi)5.1\. BRS-I UEFI Requirements The _Unified Extensible Firmware Interface Specification_ (UEFI) describes the interface between the OS and the supervisor-mode firmware. This section defines the BRS-I mandatory and optional UEFI rules on top of existing \[[10](bibliography.html#bib-uefi)\] specification requirements. Additional non-normative guidance may be found in the[firmware implementation guidance](#uefi-guidance) section. | | All content in this section is optional and recommended for BRS-B. | | --------------------------------------------------------------------- | | ID# | Rule | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | UEFI\_010 | MUST implement a 64-bit UEFI firmware. | | UEFI\_020 | MUST meet the 3rd Party UEFI Certificate Authority (CA) requirements on UEFI memory mitigations \[[15](bibliography.html#bib-msueficarequirements)\]. | | UEFI\_030 | MUST meet the following memory map rules: The default memory space attribute is EFI\_MEMORY\_WB. Paged virtual-memory scheme MUST be configured by the firmware with identity mapping and MUST support EFI\_MEMORY\_ATTRIBUTE\_PROTOCOL protocol. Only use EfiRuntimeServicesData memory type for describing any SMBIOS data structures. | | _Paged virtual memory scheme is required for platform protection use cases before handing off to the OS._ | | | UEFI\_040 | An implementation MAY comply with the _UEFI Platform Initialization Specification_ \[[16](bibliography.html#bib-uefi-pi)\]. | | UEFI\_050 | All hart manipulation internal to a firmware implementation SHOULD be done before completion of the EFI\_EVENT\_GROUP\_READY\_TO\_BOOT event. Firmware MUST place all secondary harts in an offline state before completion of the EFI\_EVENT\_GROUP\_READY\_TO\_BOOT event. | | _This ensures an OS loader is entered with an OS-compatible state for all harts.The OS loader and/or the OS may resume the secondary harts, if required, as part of their boot and join sequence._ | | | UEFI\_060 | The implementation MUST declare the EFI\_CONFORMANCE\_PROFILES\_UEFI\_SPEC\_GUID conformance profile. | | _The EFI\_CONFORMANCE\_PROFILES\_UEFI\_SPEC\_GUID conformance profile MUST be declared, as the BRS requirements are a superset of UEFI \[[10](bibliography.html#bib-uefi)\] (Section 2.6)._ | | | UEFI\_070 | The implementation MUST declare the EFI\_CONFORMANCE\_PROFILE\_BRS\_1\_0\_SPEC\_GUID conformance profile ({ 0x05453310, 0x0545, 0x0545, { 0x05, 0x45, 0x33, 0x05, 0x45, 0x33, 0x05, 0x45 }}). | | _Only a system fully compliant to the requirements in this section MUST declare the EFI\_CONFORMANCE\_PROFILE\_BRS\_1\_0\_SPEC\_GUID conformance profile._ | | | UEFI\_080 | A Device Tree MUST only be exposed to the OS if no actual hardware description is included in the DT. | | _Such a "dummy" DT could be installed by firmware, as a UEFI configuration table entry of type EFI\_DTB\_TABLE\_GUID, to provide necessary hand-off info to an OS, for example, to provide RAM disk information (e.g. via /chosen/linux,initrd-start)._ | | ### [](#5-1-1-brs-i-io-specific-requirements)5.1.1\. BRS-I I/O-specific Requirements | ID# | Rule | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | UIO\_010 | Systems implementing PCIe MUST always initialize all root complex hardware and perform resource assignment for all endpoints and usable hotplug-capable switches in the system, even in a boot scenario from a non-PCIe boot device. | | _This is a stronger requirement than the PCI Firmware Specification firmware/OS device hand-off state (\[[17](bibliography.html#bib-pcifw)\] Section 3.5). [See additional guidance](#uefi-guidance-pcie)._ | | | UIO\_020 | Systems implementing EFI\_GRAPHICS\_OUTPUT\_PROTOCOL SHOULD configure the frame buffer to be directly accessible. | | _That is, EFI\_GRAPHICS\_PIXEL\_FORMAT is not PixelBltOnly and FrameBufferBase is reported as a valid hart memory-mapped I/O address._ | | ### [](#uefi-rt)5.1.2\. BRS-I UEFI Runtime Services | ID# | Rule | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | URT\_010 | Systems without a Real-Time Clock (RTC), but with an equivalent alternate source for the current time, MUST meet the following requirements: GetTime() MUST be implemented. SetTime() MUST return EFI\_UNSUPPORTED, and be appropriately described in the EFI\_RT\_PROPERTIES\_TABLE. | | _[See additional guidance](#uefi-guidance-rt)_. | | | URT\_020 | Systems with a Real-Time Clock on an OS-managed bus (e.g. I2C, subject to arbitration issues due to access to the bus by the OS) MUST meet the following requirements: GetTime() and SetTime() MUST return EFI\_UNSUPPORTED, when called after the UEFI boot services have been exited, and must operate on the same hardware as the [ACPI TAD](#acpi-tad) before UEFI boot services are exited. GetTime() and SetTime() MUST be appropriately described in the EFI\_RT\_PROPERTIES\_TABLE. | | URT\_030 | The UEFI ResetSystem() runtime service MUST be implemented. | | _The OS MUST call the ResetSystem() runtime service call to reset or shutdown the system, preferring this to SBI, ACPI or other system-specific mechanisms. This allows for systems to perform any required system tasks on the way out (e.g. servicing UpdateCapsule() or persisting non-volatile variables in some systems)._ | | | URT\_040 | The non-volatile UEFI variables MUST persist across calls to the ResetSystem() runtime service call. | | _This rule is included in this specification to address a common mistake in implementing the UEFI requirements for non-volatile variables, even though it may appear redundant with the existing UEFI specification._ | | | URT\_050 | UEFI runtime services MUST be able to update the UEFI variables directly without the aid of an OS. | | _UEFI variables are normally saved in a dedicated storage which is not directly accessible by the operating system._ | | ### [](#5-1-3-brs-i-security-requirements)5.1.3\. BRS-I Security Requirements | ID# | Rule | | --------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | USEC\_010 | Systems implementing a TPM MUST implement the _TCG EFI Protocol Specification_ \[[18](bibliography.html#bib-tcgefiplat)\]. | | USEC\_020 | Systems with UEFI secure boot MUST support a minimum of 128 KiB of non-volatile storage for UEFI variables. | | USEC\_030 | For systems with UEFI secure boot, the maximum supported variable size MUST be at least 64 KiB. | | USEC\_040 | For systems with UEFI secure boot, the db signature database variable (EFI\_IMAGE\_SECURITY\_DATABASE) MUST be created with EFI\_VARIABLE\_TIME\_BASED\_AUTHENTICATED\_WRITE\_ACCESS, to prevent rollback attacks. | | USEC\_050 | For systems with UEFI secure boot, the dbx signature database variable (EFI\_IMAGE\_SECURITY\_DATABASE1) MUST be created with EFI\_VARIABLE\_TIME\_BASED\_AUTHENTICATED\_WRITE\_ACCESS, to prevent rollback attacks. | See additional [requirements for UEFI runtime services](#uefi-rt). ### [](#5-1-4-brs-i-firmware-update)5.1.4\. BRS-I Firmware Update | ID# | Rule | | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | UFU\_010 | Systems with in-band firmware updates MUST do so either via UpdateCapsule() UEFI runtime service (\[[10](bibliography.html#bib-uefi)\] Section 8.5.3) or via _Delivery of Capsules via file on Mass Storage Device_ (\[[10](bibliography.html#bib-uefi)\] Section 8.5.5). | | _In-band means the firmware running on a hart updates itself._ | | | UFU\_020 | Systems implementing in-band firmware updates via UpdateCapsule() MUST accept updates in the _Firmware Management Protocol Data Capsule Structure_ format as described in _Delivering Capsules Containing Updates to Firmware Management Protocol_ \[[10](bibliography.html#bib-uefi)\] (Section 23.3). | | UFU\_030 | Systems implementing in-band firmware updates via UpdateCapsule() MUST provide an ESRT \[[10](bibliography.html#bib-uefi)\] (Section 23.4) describing every firmware image that is updated in-band. | | UFU\_040 | Systems implementing in-band firmware updates via UpdateCapsule() MAY return EFI\_UNSUPPORTED, when called after the UEFI boot services have been exited. | | _[See additional guidance](#uefi-guidance-firmware-update)_. | | Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at. Copyright 2023 by RISC-V International. RISC-V Functional Fixed Hardware Specification ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-functional-fixed-hardware-specification)RISC-V Functional Fixed Hardware Specification RISC-V Platform Specification Task Group Version v1.0.1, 2024-10-10: Ratified | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Contributors ==================== ## [](#contributors)Contributors This RISC-V specification has been contributed to directly or indirectly by: * Andrew Jones [ajones@ventanamicro.com](mailto:ajones@ventanamicro.com) * Atish Patra [atishp@rivosinc.com](mailto:atishp@rivosinc.com) * Rafael Sene [rafael@riscv.org](mailto:rafael@riscv.org) * Vasudevan Srinivasan [vasu@rivosinc.com](mailto:vasu@rivosinc.com) Change Log ==================== ## [](#change-log)Change Log ### [](#version-1-0-0)Version 1.0.0 * Initial version #### [](#version-1-0-1)Version 1.0.1 * Only print the specification state on the title page, not the entire text. _CPC Object Examples ==================== ## [](#%5Fcpc-object-examples)\_CPC Object Examples ```C Device (C000) { // HART0 Name (_HID, “ACPI0007”) Name (_CPC, Package () { 23, // NumEntries 3, // Revision 120, // Highest Performance 100, // Nominal Performance 40, // Lowest Nonlinear Performance 20, // Lowest Performance ResourceTemplate () { // Guaranteed Performance Register Register(SystemMemory, 0, 0, 0, 0) // NULL }, ResourceTemplate () { // Desired Performance Register Register(FFixedHW, 64, 0, 0x1000_0000_0000_0005, QWord) }, ResourceTemplate () { // Minimum Performance Register Register(SystemMemory, 0, 0, 0, 0) // NULL }, ResourceTemplate () { // Maximum Performance Register Register(SystemMemory, 0, 0, 0, 0) // NULL }, ResourceTemplate () { // Performance Reduction Tolerance Register Register(SystemMemory, 0, 0, 0, 0) // NULL }, ResourceTemplate () { // Time Window Register Register(FFixedHW, 64, 0, 0x1000_0000_0000_0009, QWord) }, ResourceTemplate () { // Counter Wraparound Time Register(SystemMemory, 0, 0, 0, 0) // NULL }, ResourceTemplate () { // Reference Performance Counter Register Register(FFixedHW, 64, 0, 0x2000_0000_0000_0C01, QWord) }, ResourceTemplate () { // Delivered Performance Counter Register Register(FFixedHW, 64, 0, 0x1000_0000_0000_000C, QWord) }, ResourceTemplate () { // Performance Limited Register Register(FFixedHW, 64, 0, 0x1000_0000_0000_000D, QWord) }, ResourceTemplate () { // CPPC EnableRegister Register(SystemMemory, 0, 0, 0, 0) // NULL }, ResourceTemplate () { // Autonomous Selection Enable Register(SystemMemory, 0, 0, 0, 0) // NULL }, ResourceTemplate () { // AutonomousActivityWindowRegister Register(SystemMemory, 0, 0, 0, 0) // NULL }, ResourceTemplate () { // EnergyPerformancePreferenceRegister Register(SystemMemory, 0, 0, 0, 0) // NULL }, 1, // Reference Performance 20, // Lowest Frequency 100, // Nominal Frequency } ) } ``` 1.1. Introduction ==================== ## [](#1-1-introduction)1.1\. Introduction RISC-V systems which use Advanced Configuration and Power Interface (ACPI) require additional specifications for some ACPI object fields, typically those of type “Resource Descriptor”. A Functional Fixed Hardware (FFH) specification provides those additional specifications. The use cases addressed by this FFH are as follows: * Lower Power Idle States (LPI), \[ACPI 6.5\] section 8.4.3 * Collaborative Processor Performance Control (CPPC), \[ACPI 6.5\] section 8.4.6 ### [](#1-1-1-terms-and-abbreviations)1.1.1\. Terms and Abbreviations This specification uses the following terms and abbreviations: | Term | Meaning | | ---- | ------------------------------------------------------------ | | ACPI | Advanced Configuration and Power Interface Specification | | ASL | ACPI Source Language | | CPC | Continuous Performance Control | | CPPC | Collaborative Processor Performance Control | | FFH | Functional Fixed Hardware | | HSM | Hart State Management | | LPI | Low Power Idle | | OSPM | Operating System-directed configuration and Power Management | | SBI | Supervisor Binary Interface | ### [](#1-1-2-references)1.1.2\. References | Reference | Description | | ------------ | -------------------------------------------------------------------------------------------------------------- | | \[ACPI 6.5\] | Advanced Configuration and Power Interface Specification version 6.5 | | \[SBI\] | RISC-V Supervisor Binary Interface Specification version 2.0 | _LPI Object Examples ==================== ## [](#%5Flpi-object-examples)\_LPI Object Examples ```C Device (C000) { // HART0 Name (_HID, “ACPI0007”) Name (_LPI, Package () { 0, // Revision 0, // LevelID 3, // Count // LPI1 Package () { 1, // Min Residency (us) 1, // Worst case wakeup latency (us) 1, // Flags 0, // Arch. Context Lost Flags 100, // Residency Counter Frequency 0, // Enabled Parent State ResourceTemplate () { // Entry Method Register(FFixedHW, 64, 0, 0x0000_0000_0000_0000, QWord) }, ResourceTemplate () { // Residency Counter Register Register(SystemMemory, 0, 0, 0, 0) // NULL }, ResourceTemplate () { // Usage Counter Register Register(SystemMemory, 0, 0, 0, 0) // NULL }, // State Name “RISC-V WFI” }, // LPI2 Package () { 10, // Min Residency (us) 10, // Worst case wakeup latency (us) 1, // Flags 0, // Arch. Context Lost Flags 100, // Residency Counter Frequency 1, // Enabled Parent State ResourceTemplate () { // Entry Method Register(FFixedHW, 64, 0, 0x1000_0000_0000_0000, QWord) }, ResourceTemplate () { // Residency Counter Register Register(SystemMemory, 0, 0, 0, 0) // NULL }, ResourceTemplate () { // Usage Counter Register Register(SystemMemory, 0, 0, 0, 0) // NULL }, // State Name “RISC-V RET_DEFAULT” }, // LPI3 Package () { 3500, // Min Residency (us) 100, // Worst case wakeup latency (us) 1, // Flags 0, // Arch. Context Lost Flags 100, // Residency Counter Frequency 1, // Enabled Parent State ResourceTemplate () { // Entry Method Register(FFixedHW, 64, 0, 0x1000_0000_8000_0000, QWord) }, ResourceTemplate () { // Residency Counter Register Register(SystemMemory, 0, 0, 0, 0) // NULL }, ResourceTemplate () { // Usage Counter Register Register(SystemMemory, 0, 0, 0, 0) // NULL }, // State Name “RISC-V NONRET_DEFAULT” } } ) } ``` 2.1. RISC-V FFH Resource Descriptor Encoding ==================== ## [](#2-1-risc-v-ffh-resource-descriptor-encoding)2.1\. RISC-V FFH Resource Descriptor Encoding Resource descriptors, which are formatted per the “Generic Register Descriptor” definition (\[ACPI 6.5\] section 6.4.3.7, also see ASL Register in section 19.6.114), are used by ACPI tables to provide the OSPM read and/or write access to platform-specific “registers”. A resource descriptor may be a fixed value, hardware address, function or function parameter identifier, or any other encoding of its bits. While encodings are specific to their applications, RISC-V uses a common top-level encoding whenever possible. That encoding is: * Register bit width is 64 * The uppermost 4 bits are used to specify the type of identifier represented in the remaining bits. The possible types are described in table[Table 1](#table%5Fffh%5Fresource%5Fdescriptor%5Fidentifier%5Ftype) __Table 1\. RISC-V FFH Resource Descriptor Identifier Type__ | Type | Description | | ---- | ---------------------------------------- | | 0x0 | None / Other | | 0x1 | Bits\[31:0\] represent an SBI identifier | | 0x2 | Bits\[11:0\] represent a CSR identifier | **NOTE:** While it is possible for identical RISC-V FFH Resource Descriptor addresses to appear across ACPI tables, the addresses may not share the same semantics. Each address must be interpreted per its respective ACPI table type. 3.1. Use Cases ==================== ## [](#3-1-use-cases)3.1\. Use Cases ### [](#3-1-1-lower-power-idle-states)3.1.1\. Lower Power Idle States It is desirable for RISC-V system harts to transition to lower power states when they go idle. ACPI provides Lower Power Idle State (\_LPI) objects which support an operating system’s implementation of transitioning to lower power states. The following subsections specify some fields of the \_LPI object for RISC-V systems. Refer to the ACPI specification for all remaining fields. #### [](#3-1-1-1-entry-method)3.1.1.1\. Entry Method The \_LPI object uses a “Resource Descriptor”, which is formatted per[Chapter 2](ffh%5Fintroduction.html#resource%5Fdescriptor%5Fencoding), to specify the entry method for a lower power state. RISC-V systems set the resource descriptor as specified in[Table 1](#table%5Flpi%5Fentry%5Fmethod): __Table 1\. LPI Entry Method__ | Field | Value | | --------------------- | ------------------------------------------------------------------------------------------------------------------------------- | | _Address Space_ | 0x7F (FFixedHW) | | _Register Bit Width_ | 64 | | _Register Bit Offset_ | 0 | | _Access Size_ | 4 (QWord) | | _Register Address_ | As specified in the Bits\[63:60\], Bits\[59:32\], and Bits\[31:0\] columns of[Table 2](#table%5Flpi%5Fentry%5Fmethod%5Faddress) | __Table 2\. LPI Entry Method Address__ | Bits\[63:60\] (Type) | Bits\[59:32\] | Bits\[31:0\] | Description | | -------------------- | ------------- | ------------------------- | -------------------------------------------- | | 0x0 | 0x000\_0000 | 0x0000\_0000 | WFI | | 0x1 | 0x000\_0000 | SBI HSM hart suspend type | Suspend the hart using the SBI HSM extension | All other encodings for LPI entry methods with _Address Space_ set to_FFixedHW_ (0x7f) are reserved for future use. #### [](#3-1-1-2-arch-context-lost-flags)3.1.1.2\. Arch. Context Lost Flags \_LPI objects also provide an _Arch. Context Lost Flags_ field, which is a 32-bit integer, and may be used to indicate what processor context is lost. RISC-V \_LPI objects must have appropriate flags set in this field as specified in [Table 3](#table%5Flpi%5Farch%5Fcontext%5Flost%5Fflags): __Table 3\. LPI Arch. Context Lost Flags__ | Bit offset | Bit width | Description | | ---------- | --------- | ---------------------------------------- | | 0 | 1 | Set when the hart timer context is lost. | | 1 | 31 | Reserved. Must be zero. | [Appemdix A](#lpi%5Fexamples) provides examples for both a WFI entry method and SBI HSM hart suspend entry methods. ### [](#3-1-2-collaborative-processor-performance-control)3.1.2\. Collaborative Processor Performance Control ACPI describes the Collaborative Processor Performance Control (CPPC) mechanism, which is an abstract and flexible mechanism for the operating system to collaborate with an entity in the platform to manage the performance of the harts. The platform entity may be the hart itself, the platform chipset, or a separate controller. The ACPI \_CPC object provides a way for the operating system to transition the hart into a performance state selected from an abstract, continuous range of values. Fields in the \_CPC object may be static integers or “Resource Descriptors”. The following subsection specifies a RISC-V system “Resource Descriptor” for the \_CPC object. #### [](#3-1-2-1-%5Fcpc-object-resource-descriptor)3.1.2.1\. \_CPC Object Resource Descriptor The \_CPC object may use a “Resource Descriptor”, which is formatted per[Chapter 2](ffh%5Fintroduction.html#resource%5Fdescriptor%5Fencoding), for many of its fields. When using an FFH Resource Descriptor for a \_CPC field, it must be formatted as specified in[Table 4](#table%5Fcpc%5Fresource%5Fdescriptor): __Table 4\. \_CPC Resource Descriptor__ | Field | Value | | --------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | _Address Space_ | 0x7F (FFixedHW) | | _Register Bit Width_ | 64 | | _Register Bit Offset_ | 0 | | _Access Size_ | 4 (QWord) | | _Register Address_ | As specified in the Bits\[63:60\], Bits\[59:32\], Bits\[31:12\] and Bits\[11:0\] columns of[Table 5](#table%5Fcpc%5Fregister%5Faddress) | __Table 5\. \_CPC Register Address__ | Bits\[63:60\] (Type) | Bits\[59:32\] | Bits\[31:12\] | Bits\[11:0\] | Description | | -------------------- | ------------- | -------------------- | --------------- | ----------- | | 0x1 | 0x000\_0000 | SBI CPPC Register ID | SBI CPPC access | | | 0x2 | 0x000\_0000 | 0x00000 | CSR number | CSR access | All other encodings for \_CPC Resource Descriptors with _Address Space_set to _FFixedHW_ (0x7f) are reserved for future use. [Appendix B](ffh%5Flpi%5Fexamples.html#cppc%5Fexamples) provides examples for both a CSR access and an SBI CPPC access. Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at. Copyright 2024-2025 by RISC-V International. RISC-V IO Mapping Table (RIMT) ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-io-mapping-table-rimt)RISC-V IO Mapping Table (RIMT) Version v1.0, 2025-03-31: Ratified | | This document is in the [Ratified state](https://lf-riscv.atlassian.net/wiki/display/HOME/Specification+States) No changes are allowed. Any desired or needed changes can be the subject of a follow-on new extension. Ratified extensions are never revised. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Bibliography ==================== ## [](#bibliography)Bibliography \[1\] _RISC-V IOMMU Specification 1.0.0_. \[Online\]. Available: \[2\] _Advanced Configuration and Power Interface Specification 6.5_. \[Online\]. Available: \[3\] _PCI Firmware Specification 3.3_. \[Online\]. Available: Changelog ==================== ## [](#changelog)Changelog * **Version v1.0** * Version update to indicate "Ratified" state. * **Version v0.99** * Version update to indicate "Ratification Ready" state. * **Version v0.92** * Update the copyright year to include 2025. * **Version v0.91** * Addressed public review feedback. * Updated version and contributor list. * Ready for TSC sign-off. * **Version 1.0.0-rc5** * Updated document state to Frozen. * **Version 1.0.0-rc4** * Added source ID overlap restriction details. * Addressed other ARC feedback for RC3 version. * **Version 1.0.0-rc3** * Addressed feedback from ARC review. * **Version 1.0.0-rc2** * Draft for ARC review. * Addressed internal review feedback. * Used IEEE style bibliography. * Allowed HW ID to be valid for PCIe IOMMU as well. * **Version 1.0.0-rc1** * Draft for internal review. * Added ID mapping examples. * Documentation template changes. * Addressed PRS TG feedback. * **Version 0.0.1** * Initial draft for PRS TG review Contributors ==================== ## [](#contributors)Contributors This RISC-V specification has been contributed to directly or indirectly by: * Aaron Durbin <[adurbin@rivosinc.com](mailto:adurbin@rivosinc.com)\> * Andrew Jones <[ajones@ventanamicro.com](mailto:ajones@ventanamicro.com)\> * Anup Patel <[apatel@ventanamicro.com](mailto:apatel@ventanamicro.com)\> * Atish Kumar Patra <[atishp@rivosinc.com](mailto:atishp@rivosinc.com)\> * Heinrich Schuchardt <[heinrich.schuchardt@canonical.com](mailto:heinrich.schuchardt@canonical.com)\> * Samuel Holland <[samuel.holland@sifive.com](mailto:samuel.holland@sifive.com)\> * Sebastien Boeuf <[seb@rivosinc.com](mailto:seb@rivosinc.com)\> * Sunil V L <[sunilvl@ventanamicro.com](mailto:sunilvl@ventanamicro.com)\> * Tomasz Jeznach <[tjeznach@rivosinc.com](mailto:tjeznach@rivosinc.com)\> * Ved Shanbhogue <[ved@rivosinc.com](mailto:ved@rivosinc.com)\> 1.1. Introduction ==================== ## [](#1-1-introduction)1.1\. Introduction The RISC-V IO Mapping Table (RIMT) provides information about the RISC-V IOMMU \[[1](bibliography.html#bib-iommu-spec)\] and the relationship between the IO topology and the IOMMU in ACPI \[[2](bibliography.html#bib-acpi-spec)\] based RISC-V platforms. The RIMT identifies which components are behind IOMMU and how they are connected together. RISC-V IOMMU can be implemented as either a PCIe device or a platform device. 3.1. ID Mapping Examples ==================== ## [](#Mapping-Examples)3.1\. ID Mapping Examples __Table 1\. PCIe device ID mapping example__ | **Source ID Base** | **Number of IDs** | **Destination Device ID Base** | **Destination IOMMU Offset** | **Flags** | | ------------------ | ----------------- | ------------------------------ | ---------------------------- | --------- | | 0x0000 | 0x10 | 0x0 | IOMMU0\_OFFSET\_IN\_RIMT | 0 | | 0x0100 | 0x10 | 0x10 | IOMMU0\_OFFSET\_IN\_RIMT | 0 | __Table 2\. Platform device ID mapping example__ | **Source ID Base** | **Number of IDs** | **Destination Device ID Base** | **Destination IOMMU Offset** | **Flags** | | ------------------ | ----------------- | ------------------------------ | ---------------------------- | --------- | | 0x0000 | 0x1 | 0x20 | IOMMU0\_OFFSET\_IN\_RIMT | 0 | 2.1. RISC-V IO Mapping Table (RIMT) ==================== ## [](#2-1-risc-v-io-mapping-table-rimt)2.1\. RISC-V IO Mapping Table (RIMT) The [Table 1](#rimt) shows the structure of RIMT. Apart from the basic header, RIMT can contain several nodes. Each node represents a component, which can be an IOMMU, a PCIe root complex, or a platform device. __Table 1\. RISC-V IO Mapping Table__ | **Field** | **Byte Length** | **Byte Offset** | **Description** | | ------------------------- | --------------- | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Signature | 4 | 0 | 'RIMT' signature for the RISC-V IO Mapping Table. | | Length | 4 | 4 | The length of the table, in bytes, of the entire RIMT. | | Revision | 1 | 8 | 1 | | Checksum | 1 | 9 | The entire table must sum to zero. | | OEMID | 6 | 10 | OEM ID. | | OEM Table ID | 8 | 16 | For the RIMT, the table ID is the manufacturer model ID. | | OEM Revision | 4 | 24 | OEM revision of the RIMT for the supplied OEM Table ID. | | Creator ID | 4 | 28 | The vendor ID of the utility that created the table. | | Creator Revision | 4 | 32 | The revision of the utility that created the table. | | Number of RIMT Nodes | 4 | 36 | Number of nodes in the RIMT nodes array. | | Offset to RIMT Node Array | 4 | 40 | The offset from the start of this table to the first node in RIMT node array. | | Reserved | 4 | 44 | Must be zero. | | RIMT Node Array | \- | 48 | List of RIMT nodes in the platform. Nodes listed may be one of the types listed in[Table 2](#rimt%5Fnode%5Fstructure). This structure for node types is defined in the following sections. | ### [](#2-1-1-rimt-node-structure-types)2.1.1\. RIMT node structure types RIMT node structures can be broadly classified as two types: one is the actual IOMMU node structure and the other is the device node structure for devices bound to an IOMMU. The device node structure can be further classified as PCIe root complex and platform device structures bound to an IOMMU. For example, in a system with a single IOMMU, RIMT should have at least two nodes. One for the IOMMU itself and another for the devices behind this particular IOMMU. [Table 2](#rimt%5Fnode%5Fstructure) lists possible types for those structures. __Table 2\. RIMT Node Types__ | **Value** | **Description** | | --------- | ----------------------------------------------------------------- | | 0 | RISC-V IOMMU Node. See [Table 3](#iommu%5Fnode%5Fstructure) | | 1 | PCIe Root Complex Node. See [Table 5](#rc%5Fnode%5Fstructure) | | 2 | Platform Device Node. See [Table 7](#platform%5Fnode%5Fstructure) | | 3-255 | Reserved | #### [](#2-1-1-1-iommu-node)2.1.1.1\. IOMMU Node The IOMMU can be implemented as a platform device or as a PCIe device. The IOMMU node is the structure in RIMT used to report the configuration and capabilities of each IOMMU in the system. __Table 3\. IOMMU Node__ | **Field** | **Byte Length** | **Byte Offset** | **Description** | | --------------------------- | --------------- | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Type | 1 | 0 | 0 - IOMMU Node. | | Revision | 1 | 1 | 1 | | Length | 2 | 2 | The length of this structure. | | Reserved | 2 | 4 | Must be zero. | | ID | 2 | 6 | Unique ID of this node in the RIMT that can be used to locate it in the RIMT node array. It can be simply the array index in the RIMT node array. | | Hardware ID | 8 | 8 | ACPI ID of the IOMMU when it is a platform device or PCIe ID (Vendor ID + Device ID) for the PCIe IOMMU device. This field adheres to the**\_HID** format described by the ACPI \[[2](bibliography.html#bib-acpi-spec)\] specification. | | Base Address | 8 | 16 | Base address of the IOMMU registers. This field is valid for only an IOMMU that is a platform device. If IOMMU is a PCIe device, the base address of the IOMMU registers may be discovered from or programmed into the PCIe BAR of the IOMMU. | | Flags | 4 | 24 | Bit 0: IOMMU is a PCIe device 1: The IOMMU is implemented as a PCIe device. 0: The IOMMU is implemented as a platform device. Bit 1: Proximity Domain valid 1: The Proximity Domain field has a valid value. 0: The Proximity Domain field does not have a valid value. Bit \[31-2\]: Reserved, must be zero | | Proximity Domain | 4 | 28 | The Proximity Domain to which this IOMMU belongs. This is valid only when the "Proximity Domain Valid" flag is set. For optimal IOMMU performance, the in-memory data structures used by the IOMMU may be located in memory from this proximity domain. | | PCIe Segment number | 2 | 32 | If the IOMMU is implemented as a PCIe device (Bit 0 of Flags is 1), then this field holds the PCIe segment where this IOMMU is located. | | PCIe B/D/F | 2 | 34 | If the IOMMU is implemented as a PCIe device (Bit 0 of Flags is 1), then this field provides the Bus/Device/Function of the IOMMU. | | Number of interrupt wires | 2 | 36 | An IOMMU may signal IOMMU initiated interrupts by using wires or as message signaled interrupts (MSI). When the IOMMU supports signaling interrupts by using wires, this field provides the number of interrupt wires. This field must be 0 if the IOMMU does not support wire-based interrupt generation. | | Interrupt wire array offset | 2 | 38 | The offset from the start of this node entry to the first entry of the Interrupt Wire Array. This field is valid only if "Number of interrupt wires" is not 0. | | List of interrupt wires. | | | | | Interrupt wire array | 8 \* N | 40 | Array of Interrupt Wire Structures where N is the number of elements in the array. See [Table 4](#interrupt%5Fwire%5Fstructure). | __Table 4\. Interrupt Wire Structure__ | **Field** | **Byte Length** | **Byte Offset** | **Description** | | ---------------- | --------------- | --------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Interrupt Number | 4 | 0 | Interrupt number. This should be a Global System Interrupt (GSI) number. These are wired interrupts with GSI numbers mapping to a particular PLIC or APLIC. The OSPM determines the mapping of the Global System Interrupts by determining how many interrupt inputs each PLIC or APLIC supports and by determining the global system interrupt base for each PLIC / APLIC. | | Flags | 4 | 4 | Bit 0: Interrupt Mode 0: Edge Triggered. 1: Level Triggered. Bit 1: Interrupt Polarity 0: Active Low. 1: Active High. Bit \[31-2\]: Reserved, must be zero | #### [](#2-1-1-2-pcie-root-complex-node)2.1.1.2\. PCIe Root Complex Node The PCIe root complex node is the logical PCIe root complex that can be used to represent an entire physical root complex, an RCiEP/set of RCiEPs, a standalone PCIe device, or the hierarchy following a PCIe host bridge. __Table 5\. PCIe Root Complex Node__ | **Field** | **Byte Length** | **Byte Offset** | **Description** | | ----------------------- | --------------- | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Type | 1 | 0 | 1 - PCIe Root Complex Node. | | Revision | 1 | 1 | 1 | | Length | 2 | 2 | The length of this structure. | | Reserved | 2 | 4 | Must be zero. | | ID | 2 | 6 | Unique ID of this node in the RIMT that can be used to locate it in the RIMT node array. It can be simply the array index in the RIMT node array. | | Flags | 4 | 8 | Bit 0: ATS support 0: ATS is not supported in this root complex. 1: ATS supported in this root complex. Bit 1: PRI support 0: PRI is not supported in this root complex. 1: PRI is supported in this root complex. Bit \[31-2\]: Reserved, must be zero | | Reserved | 2 | 12 | Must be zero. | | PCIe Segment number | 2 | 14 | The PCIe segment number, as in MCFG and as returned by \_SEG method in the ACPI namespace. | | ID mapping array offset | 2 | 16 | The offset from the start of this node to the start of the ID mapping array. | | Number of ID mappings | 2 | 18 | Number of elements in the ID mapping array. | | List of ID mappings | | | | | ID mapping array | 20 \* N | 20 | Array of ID mapping structures where N is the number of ID mapping structures. See [Table 6](#id%5Fmapping%5Fstructure). | The ID mapping structure provides information about how devices are connected to an IOMMU. The devices can be natively identified by a source ID, but the platform can use a remapped ID to identify transactions from the device to the IOMMU. For PCIe devices, source ID is the 16-bit triplet of PCIe bus number (8-bit), device number (5-bit), and function number (3-bit) (collectively known as routing identifier or RID). A range of source IDs must map to a single IOMMU only. If there are multiple root complexes with the same PCIe segment number, then their source ID ranges must not overlap. For each ACPI device object of the root complex that belongs to the same PCIe segment, the firmware must include the Device Specific Method (\_DSM), Function Index 5, for preserving boot configurations as defined by the PCI Firmware Specification \[[3](bibliography.html#bib-pci-fw-spec)\]. The \_DSM method must return zero to indicate that the Operating System must preserve PCIe resource assignments made by the firmware at boot time. For platform devices, source ID is the implementation specific ID and managed by the device driver. Each ID mapping array entry provides a mapping from a range of source IDs to the corresponding device IDs that will be used at the input to the IOMMU. See [Chapter 3](mapping-example.html) for an example of ID mapping structures. __Table 6\. ID Mapping Structure__ | **Field** | **Byte Length** | **Byte Offset** | **Description** | | -------------------------- | --------------- | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Source ID Base | 4 | 0 | The base of a range of source IDs mapped by this entry to a range of device IDs that will be used at input to the IOMMU. | | Number of IDs | 4 | 4 | Number of IDs in the range. The range must include the IDs of devices that may be enumerated later during OS boot (For example, SR-IOV Virtual Functions). | | Destination Device ID Base | 4 | 8 | The base of the destination ID range as mapped by this entry. This is the**device\_id** as defined by the RISC-V IOMMU specification \[[1](bibliography.html#bib-iommu-spec)\] | | Destination IOMMU Offset | 4 | 12 | The destination IOMMU that is associated with these IDs. This field is the offset of the RISC-V IOMMU node from the start of the RIMT table. | | Flags | 4 | 16 | Bit 0: ATS Required 0: ATS does not need to be enabled for the device to function. 1: ATS needs to be enabled for the device to function. Bit 1: PRI Required 0: PRI does not need to be enabled for the device to function. 1: PRI needs to be enabled for the device to function. Bit \[31-2\]: Reserved, must be zero | #### [](#2-1-1-3-platform-device-node)2.1.1.3\. Platform Device Node There may be non-PCIe platform devices that are enumerated by using Differentiated System Description Table(DSDT). These devices can have one or more source IDs in the mapping table, but they can have their own scheme to define the source IDs. Hence, those source IDs can be unique to only the ACPI platform device. The interpretation of those source IDs is expected to be managed by the platform device’s device driver. __Table 7\. Platform Device Node__ | **Field** | **Byte Length** | **Byte Offset** | **Description** | | ----------------------- | --------------- | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | | Type | 1 | 0 | 2 - Platform Device Node. | | Revision | 1 | 1 | 1 | | Length | 2 | 2 | The length of this structure. | | Reserved | 2 | 4 | Must be zero. | | ID | 2 | 6 | Unique ID of this node in the RIMT that can be used to locate it in the RIMT node array. It can be simply the array index in the RIMT node array. | | ID mapping array offset | 2 | 8 | The offset from the start of this node to the start of the ID mapping array. | | Number of ID mappings | 2 | 10 | Number of elements in the ID mapping array. | | Device Object Name | M | 12 | Null terminated ASCII string. Full path to the device object in the ACPI namespace. | | Padding | P | 12 + M | Pad with zeros to align the ID mapping array at 4-byte offset. | | List of ID mappings. | | | | | ID Mapping Array | 20 \* N | 12 + M + P | Array of ID mapping structures where N is the number of ID mapping structures. See [Table 6](#id%5Fmapping%5Fstructure). | Terms and Abbreviations ==================== ## [](#terms-and-abbreviations)Terms and Abbreviations This specification uses the following terms and abbreviations: | Term | Meaning | | ----- | -------------------------------------------------------- | | ACPI | Advanced Configuration and Power Interface Specification | | APLIC | Advanced Platform-Level Interrupt Controller | | IOMMU | Input-Output Memory Management Unit | | PLIC | Platform-Level Interrupt Controller | | RCiEP | Root Complex Integrated End Point | Copyright and license information ==================== ## [](#copyright-and-license-information)Copyright and license information This specification is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). The full license text is available at. Copyright 2023-2025 by RISC-V International. RISC-V Platform Management Interface Specification (RPMI) ==================== ![RISCV](../../common/_images/risc-v_logo.svg) ## [](#risc-v-platform-management-interface-specification-rpmi)RISC-V Platform Management Interface Specification (RPMI) Version v1.0, 2025-07-16: Ratified | | This document is in the [Ratified state](http://riscv.org/spec-state) No changes are allowed. Any necessary or desired modifications must be addressed through a follow-on extension. Ratified extensions are never revised. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | Bibliography ==================== ## [](#bibliography)Bibliography \[1\] _RISC-V Supervisor Binary Interface Specification v3.0_. \[Online\]. Available: \[2\] _libRPMI_. \[Online\]. Available: \[3\] _Advanced Configuration and Power Interface Specification v6.6_. \[Online\]. Available: \[4\] _UEFI Platform Initialization Specification_. \[Online\]. Available: Changelog ==================== ## [](#changelog)Changelog ### [](#version-1-0)Version 1.0 * The RPMI specification version 1.0 with the foundations for the RPMI Message Protocol, RPMI Transport (shared memory based) and RPMI Service Groups for system control and management. Contributors ==================== ## [](#contributors)Contributors This RISC-V specification has been contributed to directly or indirectly by: Anup Patel <[apatel@ventanamicro.com](mailto:apatel@ventanamicro.com)\> Himanshu Chauhan <[hchauhan@ventanamicro.com](mailto:hchauhan@ventanamicro.com)\> Joshua Yeong <[joshua.yeong@starfivetech.com](mailto:joshua.yeong@starfivetech.com)\> Ley Foon Tan <[leyfoon.tan@starfivetech.com](mailto:leyfoon.tan@starfivetech.com)\> Rahul Pathak <[rpathak@ventanamicro.com](mailto:rpathak@ventanamicro.com)\> Samuel Holland <[samuel.holland@sifive.com](mailto:samuel.holland@sifive.com)\> Subrahmanya Lingappa <[slingappa@ventanamicro.com](mailto:slingappa@ventanamicro.com)\> Yong Li <[yong.li@intel.com](mailto:yong.li@intel.com)\> 1.1. Introduction ==================== ## [](#intro)1.1\. Introduction Today’s platforms pose challenges in terms of manageability and controllability, where the operating system (OS) needs to support a variety of hardware with a variety of connected devices. The extra complexity and demand to manage and control the platform along with executing sophisticated workloads is a challenge for the application processors (APs) running a general purpose OS. To address this challenge, platforms today contain one or more microcontrollers which can offload various platform management and control tasks. This document describes the RISC-V Platform Management Interface (RPMI), which is an OS-agnostic, firmware-agnostic, scalable and extensible interface for platform management and control from dedicated microcontrollers (also referred to as platform microcontroller or PuC). The RPMI defines a message based communication between multiple application processors (APs) and platform microcontrollers (PuCs) for system management and control. This message based communication can also be virtualized by the machine-mode firmware or hypervisors using the SBI MPXY extension \[[1](bibliography.html#bib-sbi)\]. All RPMI capabilities and services provided by platform microcontroller are discoverable at runtime which allows adding new capabilities and services in the future. ### [](#1-1-1-rpmi-abstractions)1.1.1\. RPMI Abstractions **RPMI Transport**: An RPMI transport represents the mechanism by which the messages are exchanged between the application processor and the platform microcontroller. An RPMI transport instance is associated to a particular RISC-V privilege level of the application processors and it must be accessed only by that RISC-V privilege level. **RPMI Messaging Protocol**: The RPMI messaging protocol defines different types of RPMI messages and the format for each RPMI message type. **RPMI Service Groups**: The services provided by the platform microcontrollers to the application processors are grouped into RPMI service groups based on functionality. Each RPMI service group specifies the RISC-V privilege levels from which the application processor can access it. Platform vendors can implement custom RPMI service groups. **RPMI Client**: An RPMI client is a software or a driver running on the application processor which is capable of sending and receiving RPMI messages. **RPMI Context**: An RPMI context consists of an RPMI transport instance, RPMI message protocol layer, a mandatory RPMI BASE service group and other optional RPMI service groups. An RPMI context is associated with a RISC-V privilege level which matches the included RPMI transport instance. The RPMI service groups included in an RPMI context must be accessible from the RISC-V privilege level associated with that RPMI context. The RPMI is designed to work with a single or multi-tenant topology as shown in the [Figure 1](#fig%5Fintro%5Ftrans%5Ftopology) below whereas the high-level architecture is shown in the [Figure 2](#fig%5Fintro%5Fhigh%5Flevel%5Farch) below. ![transport topologies](_images/transport-topologies.png) Figure 1\. Transport for M-Mode and S-Mode ![highlevel arch](_images/highlevel-arch.png) Figure 2\. High Level Architecture 3.1. Messaging Protocol ==================== ## [](#3-1-messaging-protocol)3.1\. Messaging Protocol The RPMI messaging protocol includes all the RPMI messages exchanged over a RPMI transport channel. ### [](#3-1-1-message-types)3.1.1\. Message Types The RPMI messaging protocol defines three types of RPMI messages namely:**REQUEST**, **ACKNOWLEDGEMENT** and **NOTIFICATION**. The [Table 1](#messaging%5Fmessage%5Ftypes%5Ftable)below summarize all RPMI message types. An **RPMI request message** represents a specific control and management task which needs to be performed and it is also referred to as an **RPMI service**. Multiple related RPMI services are grouped logically into an **RPMI service group** such as Clock, Voltage, Performance, etc. Depending on the RPMI service, a RPMI request message may carry data required to perform the control and management task. An RPMI request message may have an associated response which is sent back as an **RPMI acknowledgement message** on the same RPMI transport channel. The RPMI acknowledgement message carries the status and optional response data from an RPMI request after it has been processed. An RPMI request message which has an associated RPMI acknowledgement message is referred to as a **NORMAL REQUEST** otherwise it is referred to as a **POSTED REQUEST**. An **RPMI notification message** is an asynchronous message from the platform microcontroller to the application processors which is used to inform the later about certain events that have occurred in the system. There is no response required for an RPMI notification message from the application processors. __Table 1\. RPMI Message Types__ | Message Type | Message Subtypes | Description | | --------------- | --------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | | REQUEST | NORMAL REQUEST Request with Acknowledgement. POSTED REQUEST Request without Acknowledgement. | Messages for requesting a service from the platform microcontroller. | | ACKNOWLEDGEMENT | _Not applicable_ | Response message corresponding to a NORMAL REQUEST message. | | NOTIFICATION | _Not applicable_ | Asynchronous messages from the platform microcontroller representing system events. | ### [](#3-1-2-message-format)3.1.2\. Message Format An RPMI message consists of a fixed `8-byte` message header followed by a variable sized optional message data as show in the [Figure 1](#messaging%5Fformat)below. The byte ordering of an RPMI message is defined by the underlying RPMI transport. ![700](_images/message-format.png) Figure 1\. RPMI Message Format #### [](#3-1-2-1-message-layout-tables)3.1.2.1\. Message Layout Tables The RPMI message header and message data are split into multiple **words**, where each word is `4-byte` wide and indexed starting from `0`. The RPMI message layout is presented throughout the RPMI specification in the form of tables as shown in the [Table 2](#table%5Fmessage%5Flayout%5Ftable%5Fexample) below. Some of the columns listed below may be omitted in the layout tables if not required. __Table 2\. Message Layout Table Example__ | Word | Name | Type | Description | | ----------------------------------------------------------------------------------------- | ------------------------------------------------------- | ---------------------------------------- | -------------------------------------------- | | Index of the 4-byte word at which the field starts in the message header or message data. | Name of the field. Name may be omitted if not required. | Type of field, eg: int32 or uint32, etc. | Description and interpretation of the field. | #### [](#3-1-2-2-message-header)3.1.2.2\. Message Header The layout of the `8-byte` wide RPMI message header is shown in the[Table 3](#table%5Fmessage%5Fheader) below. The RPMI message header provide a unique identity to the corresponding RPMI message withing an RPMI context. __Table 3\. RPMI Message Header__ | Word | Description | | ---- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | Bits Name Description \[31:24\] FLAGS Message flags. FLAGS\[7:4\]: Reserved and must be 0. FLAGS\[3\]: Reserved for RPMI transport. Refer the corresponding RPMI transport chapter for more details. FLAGS\[2:0\]: Message Type. 0b000: NORMAL\_REQUEST. 0b001: POSTED\_REQUEST. 0b010: ACKNOWLEDGEMENT. 0b011: NOTIFICATION. 0b100 - 0b111: Reserved for future use. \[23:16\] SERVICE\_ID Service ID.8-bit wide identifier representing a RPMI service. This identifier is unique within a given RPMI service group. \[15:0\] SERVICEGROUP\_ID Service group ID.16-bit wide unique identifier representing a RPMI service group. | | 1 | Bits Name Description \[31:16\] TOKEN Message token.16-bit number for a RPMI message. \[15:0\] DATALEN Message data length.Stores the size of the message data in bytes. The value stored in this field must be a multiple of 4-byte or 0 if no message data is present. | For an RPMI normal request message, the `TOKEN`, `SERVICEGROUP_ID`, and`SERVICE_ID` fields of the RPMI acknowledgement message must have the same values as corresponding fields in the RPMI request message. The `DATALEN`field of the RPMI acknowledgement message must be set according to the data carried by this acknowledgement. | | The message token will help the application processors to keep track of the origin of the request when it receives a response. This is useful when the multiple application processors are sharing the same queues. For example, two different application processors may send the same type of request message with the same SERVICEGROUP\_ID and SERVICE\_ID. When the response messages for both requests are received from the platform microcontroller, the token helps distinguish which response belongs to which request. For other message types such as RPMI posted request and RPMI notification messages, the implementations may use the token for debugging or logging purposes. | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The RPMI specification recommends monotonically increasing token numbers and the token number can be initialized from any value without any constraints. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | For an RPMI notification message, the platform microcontroller will set appropriate values for the `TOKEN`, `SERVICEGROUP_ID`, and `DATALEN` fields whereas the `SERVICE_ID` field must be always set to `0x0`. #### [](#3-1-2-3-message-data)3.1.2.3\. Message Data The message data of an RPMI message is optional and variable sized. The maximum message data size of an RPMI message depends on the underlying RPMI transport implementation. The message data carries different information based on the RPMI message type: * An RPMI request message carries data required to perform the control and management task. * An RPMI acknowledgement message carries the status and optional response data. * An RPMI notification message carries an array of RPMI events. The message data format for RPMI request message and RPMI acknowledgement message is defined separately for each RPMI service. The message data format for RPMI notification message is defined in the [\[Notifications\]](#Notifications). An RPMI acknowledgement message must have a signed `STATUS` field as the first 4-byte word of the message data containing an error code defined in the [\[Possible Error Codes\]](#Possible Error Codes). An RPMI service where the response data exceeds the maximum message data size can use multipart RPMI acknowledgement messages. If a physical address is passed in the message data of any message type, then it refers to the physical address space of the application processor. ### [](#3-1-3-notifications)3.1.3\. Notifications The platform microcontroller can use RPMI notification message to notify application processors about system events which are also referred to as**RPMI events**. An RPMI notification message has no associated response or acknowledgement from application processors. Multiple RPMI events can be packed into a single RPMI notification message depending on the space available in the message data. Each RPMI event may have additional data associated with it based on the type of RPMI event. Any action required for handling an RPMI event depends on the application processors. The format of an RPMI notification message in shown in the [Figure 2](#messaging%5Fnotif%5Fformat)below. The RPMI events are defined separately for each RPMI service group. An RPMI service group must have a `ENABLE_NOTIFICATION` service with a fixed`SERVICE_ID=0x01` which can be used by the application processors to enable or disable notification messages for a particular RPMI event defined by the RPMI service groups. By default, notifications are disabled for all RPMI events of an RPMI service group. The platform microcontroller only sends RPMI notification messages for RPMI events which are enabled by the application processors. If multiple RPMI events are supported by an RPMI service group then the application processors must enable to each RPMI event individually. ![500](_images/notification-format.png) Figure 2\. RPMI Notification Message Format #### [](#3-1-3-1-events)3.1.3.1\. Events An RPMI event consists of a header containing two fields: `EVENT_ID (8-bit)`and `EVENT_DATALEN (16-bit)`. An RPMI event may have associated data whose size is specified in the `EVENT_DATALEN` field of the header and this data size must be a multiple of `4-byte`. The number of RPMI events that can be stored in a single RPMI notification message depends on the maximum RPMI message data size. The `DATALEN` field in the RPMI message header represents the aggregate size of all RPMI events included in RPMI message data. The [Table 4](#table%5Fnotification%5Fmessage%5Fformat) below defines the format of an RPMI event whereas the [Figure 3](#messaging%5Fevent%5Fformat) below shows a pictorial view of an RPMI event. The format of the event data for each RPMI event is defined separately by an RPMI service group. If multiple RPMI events are packed into a single RPMI notification message then the ordering of RPMI events within the RPMI notification message is implementation defined. The platform microcontroller is not required to include every occurrence of an event of the same type in a notification message. Instead, the platform microcontroller must only include the most recently occurred event of the same type. __Table 4\. Event Format__ | Word | Name | Description | | ---- | ----------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | EVENT\_HDR | 32-bit field represents a single event. Bits Name Description \[31:24\] _Reserved_ _Reserved_ and must be 0. \[23:16\] EVENT\_ID Unique identifier for an event in a service group. \[15:0\] EVENT\_DATALEN 16-bit field to store event data size in bytes. | | 1 | EVENT\_DATA | Event data whose size is specified by EVENT\_DATALEN. | ![800](_images/event-header.png) Figure 3\. Event Header ### [](#3-1-4-possible-error-codes)3.1.4\. Possible Error Codes The [Table 5](#table%5Ferror%5Fcodes) below lists the error codes which can be returned by an RPMI service in the `STATUS` field of the RPMI acknowledgement message. __Table 5\. RPMI Error Codes__ | Name | Error Code | Description | | ------------------------- | ------------------------------------ | -------------------------------------------------------------------------------------------------------------- | | RPMI\_SUCCESS | 0 | Service has been completed successfully. | | RPMI\_ERR\_FAILED | \-1 | Failed due to general error. | | RPMI\_ERR\_NOT\_SUPPORTED | \-2 | Service or feature is not supported. | | RPMI\_ERR\_INVALID\_PARAM | \-3 | One or more parameters passed are invalid. | | RPMI\_ERR\_DENIED | \-4 | Requested operation denied due to insufficient permissions or failed dependency check. | | RPMI\_ERR\_INVALID\_ADDR | \-5 | One or more addresses are invalid. | | RPMI\_ERR\_ALREADY | \-6 | Operation already in progress or state changed already for which the operation was performed. | | RPMI\_ERR\_EXTENSION | \-7 | Error in extension implementation that violates the extension specification or the extension version mismatch. | | RPMI\_ERR\_HW\_FAULT | \-8 | Failed due to hardware fault. | | RPMI\_ERR\_BUSY | \-9 | Service cannot be completed due to system or device is busy. | | RPMI\_ERR\_INVALID\_STATE | \-10 | Invalid state. | | RPMI\_ERR\_BAD\_RANGE | \-11 | Bad or invalid range. | | RPMI\_ERR\_TIMEOUT | \-12 | Failed due to timeout. | | RPMI\_ERR\_IO | \-13 | Input/Output error. | | RPMI\_ERR\_NO\_DATA | \-14 | Data not available. | | \-15 to -127 | _Reserved_. | | | < -127 | _Vendor or Implementation specific_. | | 5.1. Integration with SBI MPXY Extension ==================== ## [](#5-1-integration-with-sbi-mpxy-extension)5.1\. Integration with SBI MPXY Extension A platform with a limited number of RPMI transport instances can share an M-mode RPMI transport instance with the supervisor software using the SBI MPXY extension \[[1](bibliography.html#bib-sbi)\]. An M-mode firmware or hypervisor can also virtualize RPMI message communication for the supervisor software using the SBI MPXY extension. As shown in the [Figure 1](#mpxy%5Frpmi%5Fintegration) below, the SBI implementation acts as a **RPMI proxy**for the supervisor software when sending RPMI messages through an SBI MPXY channel. ![350](_images/mpxy-rpmi.png) Figure 1\. RPMI and SBI MPXY Integration The RPMI communication via the SBI MPXY extension must satisfy the following requirements: 1. The SBI MPXY channel must correspond to a single RPMI service group which is allowed in S-mode except BASE and CPPC service groups. The [\[table\_service\_groups\]](#table%5Fservice%5Fgroups)list the RPMI service groups allowed in S-mode. 2. The SBI MPXY channel corresponding to the RPMI SYSTEM\_MSI service group must not support the P2A doorbell system MSI. 3. The `message_id` parameter passed to the `sbi_mpxy_send_message_with_response()`and `sbi_mpxy_send_message_without_response()` must represent the `SERVICE_ID` of an RPMI service belonging to the RPMI service group bound to the SBI MPXY channel. 4. The format of the message data passed via the SBI MPXY shared memory to the`sbi_mpxy_send_message_with_response()` and `sbi_mpxy_send_message_without_response()`must match the request data format of the RPMI service represented by the`message_id` parameter. 5. The format of the response message data returned in the SBI MPXY shared memory by the `sbi_mpxy_send_message_with_response()` must match the response data format of the RPMI service represented by the `message_id` parameter. 6. The format of the protocol-specific data returned in the SBI MPXY shared memory by the `sbi_mpxy_get_notification_events()` must match the RPMI notifications message data format. 7. The SBI MPXY channel must support message protocol attributes listed in the[Table 1](#table%5Frpmi%5Fmpxy%5Fattributes) below. __Table 1\. RPMI Message Protocol Attributes of an SBI MPXY Channel__ | Attribute Name | Attribute ID | Access | Description | | ----------------------- | ------------ | ------ | ---------------------------- | | SERVICEGROUP\_ID | 0x80000000 | RO | RPMI service group ID. | | SERVICEGROUP\_VERSION | 0x80000001 | RO | RPMI service group version. | | IMPLEMENTATION\_ID | 0x80000002 | RO | RPMI implementation ID. | | IMPLEMENTATION\_VERSION | 0x80000003 | RO | RPMI implementation version. | 4.1. Service Groups ==================== ## [](#4-1-service-groups)4.1\. Service Groups An RPMI service group is a collection of RPMI services that are logically grouped based on functionality. For example, all the voltage related services are grouped into a voltage service group. The functionality implemented by certain RPMI service groups may impact the architectural state of application processors due to this each RPMI service group specifies the RISC-V privilege levels of the application processor which can be access it. For example, the clock service groups can be accessed from M-mode and S-mode but the HSM service group can be only accessed from M-mode. All RPMI service groups except the BASE service group are optional. If the `BASE_PROBE_SERVICE_GROUP` service indicates that a service group is implemented then the RPMI service group version must conform to the RPMI specification version returned by the `BASE_GET_SPEC_VERSION` service. All implemented RPMI service groups must satisfy the following requirements: 1. The RPMI service group must be accessible from the RISC-V privilege level associated with the RPMI context which includes it. 2. All RPMI services of the RPMI service groups must be supported except the dedicated notification service (`SERVICE_ID = 0x00`) which is reserved for RPMI notification messages. A RPMI service group may implement its RPMI services partially only if it also defines a mechanism to discover supported RPMI services. 3. The RPMI service group must implement a dedicated RPMI service with`SERVICE_ID = 0x01` to subscribe for event notifications. | | The RPMI services listed within each RPMI service group do not have a specific order. Additionally, the sequence in which services are defined in the specification does not necessarily reflect the order in which they should be invoked. | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | This specification defines standard RPMI service groups and RPMI services with the provision to add more service groups as required in the future. The RPMI specification also provides experimental service group IDs space for development of service group until a standard service group ID is allocated. The platform vendors can provide implementation specific RPMI service groups. The [Table 1](#table%5Fservice%5Fgroups) table below lists all standard RPMI service groups defined by this specification. __Table 1\. RPMI Service Groups__ | Service Group ID | RPMI Version (Major:Minor) | Service Group Name | Allowed RISC-V Privilege Levels on Application Processors | | ---------------- | ---------------------------------------- | ----------------------- | --------------------------------------------------------- | | 0x0001 | 1.0 | BASE | M-mode, S-mode | | 0x0002 | 1.0 | SYSTEM\_MSI | M-mode, S-mode | | 0x0003 | 1.0 | SYSTEM\_RESET | M-mode | | 0x0004 | 1.0 | SYSTEM\_SUSPEND | M-mode | | 0x0005 | 1.0 | HART\_STATE\_MANAGEMENT | M-mode | | 0x0006 | 1.0 | CPPC | M-mode, S-mode | | 0x0007 | 1.0 | VOLTAGE | M-mode, S-mode | | 0x0008 | 1.0 | CLOCK | M-mode, S-mode | | 0x0009 | 1.0 | DEVICE\_POWER | M-mode, S-mode | | 0x000A | 1.0 | PERFORMANCE | M-mode, S-mode | | 0x000B | 1.0 | MANAGEMENT\_MODE | M-mode, S-mode | | 0x000C | 1.0 | RAS\_AGENT | M-mode, S-mode | | 0x000D | 1.0 | REQUEST\_FORWARD | M-mode, S-mode | | 0x000E - 0x7BFF | _Reserved for Future Use_ | | | | 0x7C00 - 0x7FFF | _Experimental Service Groups_ | | | | 0x8000 - 0xFFFF | _Implementation Specific Service Groups_ | | | ### [](#service-group-base-servicegroup%5Fid-0x0001)Service Group: BASE (SERVICEGROUP\_ID: 0x0001) The BASE service group is mandatory and provides the following services: * Initial handshaking between the application processor and the platform microcontroller. * Discovering the RPMI implementation version information. * Discovering the implementation of a service group. * Discovering platform specific information. The following table lists the services in the BASE service group: __Table 2\. BASE Services__ | Service ID | Service Name | Request Type | | ---------- | ---------------------------------- | --------------- | | 0x01 | BASE\_ENABLE\_NOTIFICATION | NORMAL\_REQUEST | | 0x02 | BASE\_GET\_IMPLEMENTATION\_VERSION | NORMAL\_REQUEST | | 0x03 | BASE\_GET\_IMPLEMENTATION\_ID | NORMAL\_REQUEST | | 0x04 | BASE\_GET\_SPEC\_VERSION | NORMAL\_REQUEST | | 0x05 | BASE\_GET\_PLATFORM\_INFO | NORMAL\_REQUEST | | 0x06 | BASE\_PROBE\_SERVICE\_GROUP | NORMAL\_REQUEST | | 0x07 | BASE\_GET\_ATTRIBUTES | NORMAL\_REQUEST | #### [](#rpmi-implementation-ids)RPMI Implementation IDs The RPMI specification defines space for standard implementation IDs and for experimental implementation IDs. The experimental implementation IDs can be used by the implementations until a standard implementation ID is assigned to it. The RPMI implementations that have been assigned a standard implementation ID are listed in the table below. __Table 3\. RPMI Implementation IDs__ | Implementation ID | Name | | ----------------------- | ---------------------------------------------- | | 0x00000000 | libRPMI \[[2](bibliography.html#bib-librpmi)\] | | 0x00000001 - 0x7FFFFFFF | _Reserved for Future Use_ | | 0x80000000 - 0xFFFFFFFF | _Experimental Implementation IDs_ | #### [](#base-notifications)Notifications This service is used by the platform microcontroller to send the asynchronous message of type notification to the application processor. The message transfers the events defined by this service group. The events defined are listed in the below table. __Table 4\. BASE Service Group Events__ | Event ID | Event Name | Event Data | Description | | -------- | ---------------------- | ---------- | --------------------------------------------------------------------------------------------------------------------------------------- | | 0x01 | REQUEST\_HANDLE\_ERROR | NA | This event indicates that the platform microcontroller is unable to serve the message requests and acknowledgements are not guaranteed. | #### [](#service-base%5Fenable%5Fnotification-service%5Fid-0x01)Service: BASE\_ENABLE\_NOTIFICATION (SERVICE\_ID: 0x01) This service allows the application processor to subscribe to `BASE`service group notifications. The platform may optionally support notifications for events that may occur. The platform microcontroller can send these notification messages to the application processor if they are implemented and the application processor has subscribed to them. The supported events are described in [Notifications](#base-notifications). __Table 5\. Request Data__ | Word | Name | Type | Description | | ---- | ---------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | EVENT\_ID | uint32 | The event to be subscribed for notification. | | 1 | REQ\_STATE | uint32 | Requested event notification state.Change or query the current state of EVENT\_ID notification. 0: Disable. 1: Enable. 2: Return current state. Any other values of REQ\_STATE field other than the defined ones are reserved for future use. | __Table 6\. Response Data__ | Word | Name | Type | Description | | ---- | -------------- | ------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Event is subscribed successfully. RPMI\_ERR\_INVALID\_PARAM EVENT\_ID or REQ\_STATE is invalid. RPMI\_ERR\_NOT\_SUPPORTED Notification for the EVENT\_ID is not supported. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | CURRENT\_STATE | uint32 | Current EVENT\_ID notification state. 0: Notification is disabled. 1: Notification is enabled. In case of REQ\_STATE = 0 or 1, the CURRENT\_STATE will return the requested state.In case of an error, the value of CURRENT\_STATE is unspecified. | #### [](#service-base%5Fget%5Fimplementation%5Fversion-service%5Fid-0x02)Service: BASE\_GET\_IMPLEMENTATION\_VERSION (SERVICE\_ID: 0x02) This service is used to get the RPMI implementation version of the platform microcontroller. The version returned is a 32-bit composite number containing the `MAJOR` and `MINOR` version numbers. __Table 7\. Request Data__ | NA | | -- | __Table 8\. Response Data__ | Word | Name | Type | Description | | ---- | ------------- | ------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS RPMI implementation version returned successfully. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | IMPL\_VERSION | uint32 | Implementation version. Bits Description \[31:16\] MAJOR number. \[15:0\] MINOR number. | #### [](#service-base%5Fget%5Fimplementation%5Fid-service%5Fid-0x03)Service: BASE\_GET\_IMPLEMENTATION\_ID (SERVICE\_ID: 0x03) This service is used to get a 32-bit RPMI implementation ID assigned to the software that implements the RPMI specification. Every implementation ID is unique and listed in the [Table 3](#table%5Fbase%5Frpmi%5Fimpl%5Fid). __Table 9\. Request Data__ | NA | | -- | __Table 10\. Response Data__ | Word | Name | Type | Description | | ---- | -------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS RPMI implementation ID returned successfully. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | IMPL\_ID | uint32 | Implementation ID. | #### [](#service-base%5Fget%5Fspec%5Fversion-service%5Fid-0x04)Service: BASE\_GET\_SPEC\_VERSION (SERVICE\_ID: 0x04) This service is used to get the implemented RPMI specification version. The version returned is a 32-bit composite number containing the `MAJOR` and`MINOR` version numbers. __Table 11\. Request Data__ | NA | | -- | __Table 12\. Response Data__ | Word | Name | Type | Description | | ---- | ------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS RPMI specification version returned successfully. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes) | | 1 | SPEC\_VERSION | uint32 | RPMI specification version. Bits Description \[31:16\] MAJOR number. \[15:0\] MINOR number. | #### [](#service-base%5Fget%5Fplatform%5Finfo-service%5Fid-0x05)Service: BASE\_GET\_PLATFORM\_INFO (SERVICE\_ID: 0x05) This service is used to get additional platform information if available. __Table 13\. Request Data__ | NA | | -- | __Table 14\. Response Data__ | Word | Name | Type | Description | | ---- | ----------------- | -------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Platform information returned successfully. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | PLATFORM\_ID\_LEN | uint32 | Platform Identifier field length in bytes. | | 2 | PLATFORM\_ID | uint8\[PLATFORM\_ID\_LEN\] | Platform Identifier.Up to PLATFORM\_ID\_LEN bytes NULL terminated ASCII string. The use and interpretation of this field is implementation-defined. It can be used to convey details such as the vendor ID, vendor name, specific product model, revision, or configuration of the hardware. | #### [](#service-base%5Fprobe%5Fservice%5Fgroup-service%5Fid-0x06)Service: BASE\_PROBE\_SERVICE\_GROUP (SERVICE\_ID: 0x06) This service is used to probe the implementation of a service group and to obtain the implemented service group version. The service group version is a 32-bit composite number containing the `MAJOR` and `MINOR` numbers. If the service group is successfully probed then the implemented service group version is returned in the `SERVICE_GROUP_VERSION` field. Otherwise it returns`0`. __Table 15\. Request Data__ | Word | Name | Type | Description | | ---- | ---------------- | ------ | ----------------- | | 0 | SERVICEGROUP\_ID | uint32 | Service group ID. | __Table 16\. Response Data__ | Word | Name | Type | Description | | ---- | ----------------------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service probed successfully. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes) | | 1 | SERVICE\_GROUP\_VERSION | uint32 | Service group version. Bits Description \[31:16\] MAJOR number. \[15:0\] MINOR number. | #### [](#service-base%5Fget%5Fattributes-service%5Fid-0x07)Service: BASE\_GET\_ATTRIBUTES (SERVICE\_ID: 0x07) This service is used to discover additional features supported by the BASE service group. __Table 17\. Request Data__ | NA | | -- | __Table 18\. Response Data__ | Word | Name | Type | Description | | ---- | ------ | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Attributes returned successfully. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | FLAGS0 | uint32 | Bits Description \[31:2\] _Reserved_ and must be 0. \[1\] RPMI context privilege level. 0b1: M-mode. 0b0: S-mode. \[0\] Event notification support in platform. 0b1: Supported. 0b0: Not supported. | | 2 | FLAGS1 | uint32 | _Reserved_ and must be 0. | | 3 | FLAGS2 | uint32 | _Reserved_ and must be 0. | | 4 | FLAGS3 | uint32 | _Reserved_ and must be 0. | ### [](#service-group-system%5Fmsi-servicegroup%5Fid-0x0002)Service Group - SYSTEM\_MSI (SERVICEGROUP\_ID: 0x0002) The SYSTEM\_MSI service group defines services to allow application processors to receive MSIs upon system events such as P2A doorbell, graceful shutdown/reboot request, CPU hotplug event, memory hotplug event, etc. The number of system MSIs supported by this service group is fixed and referred to as `SYS_NUM_MSI`. Each system MSI is associated with a unique index which is also referred to as `SYS_MSI_INDEX` where `0 <​= SYS_MSI_INDEX < SYS_NUM_MSI`. The association between system events and system MSI index (aka `SYS_MSI_INDEX`) is platform specific and must be discovered using hardware description mechanisms such as device tree or ACPI. The system MSI state is 32-bit word also referred to as `SYS_MSI_STATE` which includes whether the system MSI is enabled/disabled and whether system MSI is currently pending at the platform microcontroller. The [Table 19](#table%5Fsysmsi%5Fstate)below shows the encoding of `SYS_MSI_STATE`. | | A system MSI can be pending for several reasons. For example, if the MSI target address and data are not configured, or if the configured MSI target address is not valid. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 19\. System MSI State__ | Bits | Permission | Description | | -------- | ---------- | ---------------------------------------------------------------- | | \[31:2\] | NA | _Reserved_ and must be 0. | | \[1\] | Read-Only | MSI pending state. 0b1: MSI is pending. 0b0: MSI is not pending. | | \[0\] | Read-Write | MSI enable state. 0b1: MSI enabled. 0b0: MSI disabled. | The platform microcontroller can only send a pending system MSI if it is enabled and the configured with a valid MSI target address. The system MSI can be enabled/disabled using the `SYSMSI_SET_MSI_STATE` service whereas the system MSI target configuration can be done using the `SYSMSI_SET_MSI_TARGET`service. | | If the system MSI target address is IMSIC, then the application processors will directly receive the system MSI whereas if the system MSI target address is setipnum register of a APLIC domain then the application processors receive it as wired interrupt. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | The [Table 20](#table%5Fsysmsi%5Fservices) below lists the services in the SYSTEM\_MSI service group: __Table 20\. SYSTEM\_MSI Services__ | Service ID | Service Name | Request Type | | ---------- | ---------------------------- | --------------- | | 0x01 | SYSMSI\_ENABLE\_NOTIFICATION | NORMAL\_REQUEST | | 0x02 | SYSMSI\_GET\_ATTRIBUTES | NORMAL\_REQUEST | | 0x03 | SYSMSI\_GET\_MSI\_ATTRIBUTES | NORMAL\_REQUEST | | 0x04 | SYSMSI\_SET\_MSI\_STATE | NORMAL\_REQUEST | | 0x05 | SYSMSI\_GET\_MSI\_STATE | NORMAL\_REQUEST | | 0x06 | SYSMSI\_SET\_MSI\_TARGET | NORMAL\_REQUEST | | 0x07 | SYSMSI\_GET\_MSI\_TARGET | NORMAL\_REQUEST | #### [](#system-msi-notifications)Notifications This service group does not support any events for notification. #### [](#service-sysmsi%5Fenable%5Fnotification-service%5Fid-0x01)Service: SYSMSI\_ENABLE\_NOTIFICATION (SERVICE\_ID: 0x01) This service allows the application processor to subscribe to `SYSTEM_MSI`service group notifications. The platform may optionally support notifications for events that may occur. The platform microcontroller can send these notification messages to the application processor if they are implemented and the application processor has subscribed to them. The supported events are described in [Notifications](#system-msi-notifications). __Table 21\. Request Data__ | Word | Name | Type | Description | | ---- | ---------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | EVENT\_ID | uint32 | The event to be subscribed for notification. | | 1 | REQ\_STATE | uint32 | Requested event notification state.Change or query the current state of EVENT\_ID notification. 0: Disable. 1: Enable. 2: Return current state. Any other values of REQ\_STATE field other than the defined ones are reserved for future use. | __Table 22\. Response Data__ | Word | Name | Type | Description | | ---- | -------------- | ------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Event is subscribed successfully. RPMI\_ERR\_INVALID\_PARAM EVENT\_ID or REQ\_STATE is invalid. RPMI\_ERR\_NOT\_SUPPORTED Notification for the EVENT\_ID is not supported. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | CURRENT\_STATE | uint32 | Current EVENT\_ID notification state. 0: Notification is disabled. 1: Notification is enabled. In case of REQ\_STATE = 0 or 1, the CURRENT\_STATE will return the requested state.In case of an error, the value of CURRENT\_STATE is unspecified. | #### [](#service-sysmsi%5Fget%5Fattributes-service%5Fid-0x02)Service: SYSMSI\_GET\_ATTRIBUTES (SERVICE\_ID: 0x02) This service is used to discover attributes of the system MSI service group. __Table 23\. Request Data__ | NA | | -- | __Table 24\. Response Data__ | Word | Name | Type | Description | | ---- | ------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | SYS\_NUM\_MSI | uint32 | Number of system MSIs. | | 2 | FLAGS0 | uint32 | _Reserved_ and must be 0. | | 3 | FLAGS1 | uint32 | _Reserved_ and must be 0. | #### [](#service-sysmsi%5Fget%5Fmsi%5Fattributes-service%5Fid-0x03)Service: SYSMSI\_GET\_MSI\_ATTRIBUTES (SERVICE\_ID: 0x03) This service is used to discover attributes of a particular system MSI. __Table 25\. Request Data__ | Word | Name | Type | Description | | ---- | --------------- | ------ | ------------------------ | | 0 | SYS\_MSI\_INDEX | uint32 | Index of the system MSI. | __Table 26\. Response Data__ | Word | Name | Type | Description | | ---- | -------------- | ----------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. RPMI\_ERR\_INVALID\_PARAM SYS\_MSI\_INDEX value is greater than SYS\_NUM\_MSI. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | FLAGS0 | uint32 | Bits Description \[31:1\] _Reserved_ and must be 0. \[0\] Preferred privilege level for MSI handling. 0b1: M-mode. 0b0: M-mode or S-mode. | | 2 | FLAGS1 | uint32 | _Reserved_ and must be 0. | | 3:6 | SYS\_MSI\_NAME | uint8\[16\] | System MSI name, a NULL-terminated ASCII string up to 16-bytes. | #### [](#srvgrp%5Fsysmsi%5Fset%5Fmsi%5Fstate)Service: SYSMSI\_SET\_MSI\_STATE (SERVICE\_ID: 0x04) This service is used to update the state of a system MSI. Specifically, it allows application processors to enable or disable a system MSI. The read-only bits of the system MSI state are not updated by this service. __Table 27\. Request Data__ | Word | Name | Type | Description | | ---- | --------------- | ------ | ------------------------------------------------------------------- | | 0 | SYS\_MSI\_INDEX | uint32 | Index of the system MSI. | | 1 | SYS\_MSI\_STATE | uint32 | System MSI state as defined in [Table 19](#table%5Fsysmsi%5Fstate). | __Table 28\. Response Data__ | Word | Name | Type | Description | | ---- | ------ | ----- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS MSI is enabled or disabled successfully. RPMI\_ERR\_INVALID\_PARAM SYS\_MSI\_INDEX value is greater than SYS\_NUM\_MSI orSYS\_MSI\_STATE value is reserved or invalid. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | #### [](#srvgrp%5Fsysmsi%5Fget%5Fmsi%5Fstate)Service: SYSMSI\_GET\_MSI\_STATE (SERVICE\_ID: 0x05) This service is used to get the current state of a system MSI. __Table 29\. Request Data__ | Word | Name | Type | Description | | ---- | --------------- | ------ | ------------------------ | | 0 | SYS\_MSI\_INDEX | uint32 | Index of the system MSI. | __Table 30\. Response Data__ | Word | Name | Type | Description | | ---- | --------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS MSI state is returned successfully. RPMI\_ERR\_INVALID\_PARAM SYS\_MSI\_INDEX value is greater than SYS\_NUM\_MSI. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | SYS\_MSI\_STATE | uint32 | System MSI state as defined in [Table 19](#table%5Fsysmsi%5Fstate). | #### [](#srvgrp%5Fsysmsi%5Fset%5Fmsi%5Ftarget)Service: SYSMSI\_SET\_MSI\_TARGET (SERVICE\_ID: 0x06) This service is used to configure the target address and data of a system MSI. __Table 31\. Request Data__ | Word | Name | Type | Description | | ---- | ----------------------- | ------ | -------------------------------- | | 0 | SYS\_MSI\_INDEX | uint32 | Index of the system MSI. | | 1 | SYS\_MSI\_ADDRESS\_LOW | uint32 | Lower 32-bit of the MSI address. | | 2 | SYS\_MSI\_ADDRESS\_HIGH | uint32 | Upper 32-bit of the MSI address. | | 3 | SYS\_MSI\_DATA | uint32 | 32-bit MSI data. | __Table 32\. Response Data__ | Word | Name | Type | Description | | ---- | ------ | ----- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS MSI address and data are configured successfully. RPMI\_ERR\_INVALID\_PARAM SYS\_MSI\_INDEX value is greater than SYS\_NUM\_MSI. RPMI\_ERR\_INVALID\_ADDR MSI target address is invalid or it is not 4-byte aligned. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | #### [](#srvgrp%5Fsysmsi%5Fget%5Fmsi%5Ftarget)Service: SYSMSI\_GET\_MSI\_TARGET (SERVICE\_ID: 0x07) This service is used to get the current target address and data of a system MSI. __Table 33\. Request Data__ | Word | Name | Type | Description | | ---- | --------------- | ------ | ------------------------ | | 0 | SYS\_MSI\_INDEX | uint32 | Index of the system MSI. | __Table 34\. Response Data__ | Word | Name | Type | Description | | ---- | ----------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS MSI target details returned successfully. RPMI\_ERR\_INVALID\_PARAM SYS\_MSI\_INDEX value is greater than SYS\_NUM\_MSI. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | SYS\_MSI\_ADDRESS\_LOW | uint32 | Lower 32-bit of the MSI address. | | 2 | SYS\_MSI\_ADDRESS\_HIGH | uint32 | Upper 32-bit of the MSI address. | | 3 | SYS\_MSI\_DATA | uint32 | 32-bit MSI data. | ### [](#service-group-system%5Freset-servicegroup%5Fid-0x0003)Service Group - SYSTEM\_RESET (SERVICEGROUP\_ID: 0x0003) This service group defines services for system-level reset or shutdown. The architectural reset types that are supported by default are System Cold Reset and System Shutdown. System Cold Reset, also known as power-on-reset, involves power cycling the entire system. Upon a successful system cold reset, all devices undergo power cycling in an implementation-defined sequence similar to the initial power-on sequence of the system. System Shutdown results in all components/devices in the system losing power. Currently, the application processor is the only entity that can request the system shutdown, which means that for the platform microcontroller, it is not necessary to categorize it as a graceful or forceful shutdown. In the case of a shutdown request, it is implicit for the platform microcontroller that the application processor has prepared itself for a successful shutdown. The following table lists the services in the SYSTEM\_RESET service group: __Table 35\. SYSTEM\_RESET Services__ | Service ID | Service Name | Request Type | | ---------- | ---------------------------- | --------------- | | 0x01 | SYSRST\_ENABLE\_NOTIFICATION | NORMAL\_REQUEST | | 0x02 | SYSRST\_GET\_ATTRIBUTES | NORMAL\_REQUEST | | 0x03 | SYSRST\_RESET | POSTED\_REQUEST | #### [](#section-reset-types)Reset Types RPMI supports reset types and their values as defined by SBI specification. Refer to [**SRST System Reset Types**](https://github.com/riscv-non-isa/riscv-sbi-doc/blob/master/src/ext-sys-reset.adoc#table%5Fsrst%5Fsystem%5Freset%5Ftypes)in the RISC-V SBI Specification \[[1](bibliography.html#bib-sbi)\] for the `RESET_TYPE`. #### [](#system-reset-notifications)Notifications This service group does not support any events for notification. #### [](#service-sysrst%5Fenable%5Fnotification-service%5Fid-0x01)Service: SYSRST\_ENABLE\_NOTIFICATION (SERVICE\_ID: 0x01) This service allows the application processor to subscribe to `SYSTEM_RESET`service group notifications. The platform may optionally support notifications for events that may occur. The platform microcontroller can send these notification messages to the application processor if they are implemented and the application processor has subscribed to them. The supported events are described in [Notifications](#system-reset-notifications). __Table 36\. Request Data__ | Word | Name | Type | Description | | ---- | ---------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | EVENT\_ID | uint32 | The event to be subscribed for notification. | | 1 | REQ\_STATE | uint32 | Requested event notification state.Change or query the current state of EVENT\_ID notification. 0: Disable. 1: Enable. 2: Return current state. Any other values of REQ\_STATE field other than the defined ones are reserved for future use. | __Table 37\. Response Data__ | Word | Name | Type | Description | | ---- | -------------- | ------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Event is subscribed successfully. RPMI\_ERR\_INVALID\_PARAM EVENT\_ID or REQ\_STATE is invalid. RPMI\_ERR\_NOT\_SUPPORTED Notification for the EVENT\_ID is not supported. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | CURRENT\_STATE | uint32 | Current EVENT\_ID notification state. 0: Notification is disabled. 1: Notification is enabled. In case of REQ\_STATE = 0 or 1, the CURRENT\_STATE will return the requested state.In case of an error, the value of CURRENT\_STATE is unspecified. | #### [](#service-sysrst%5Fget%5Fattributes-service%5Fid-0x02)Service: SYSRST\_GET\_ATTRIBUTES (SERVICE\_ID: 0x02) This service is used to discover the attributes of a reset type. The attribute flags indicates if a `RESET_TYPE` is supported or not apart from the System Shutdown and System Cold Reset which are mandatory and supported by default. System Warm Reset support can be discovered with this service. __Table 38\. Request Data__ | Word | Name | Type | Description | | ---- | ----------- | ------ | ---------------------------------------------------------------------- | | 0 | RESET\_TYPE | uint32 | Reset type.Refer [Reset Types](#section-reset-types) for more details. | __Table 39\. Response Data__ | Word | Name | Type | Description | | ---- | ------ | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Attributes returned successfully. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | FLAGS | uint32 | Reset type attributes. Bits Description \[31:1\] _Reserved_ and must be 0. \[0\] Reset type support. 0b1: Supported. 0b0: Not supported. | #### [](#service-sysrst%5Freset-service%5Fid-0x03)Service: SYSRST\_RESET (SERVICE\_ID: 0x03) This service is used to initiate the system reset or system shutdown. The application processor must only request supported reset types, discovered using the `SYSRST_GET_ATTRIBUTES` service except for System Shutdown and System Cold Reset which are supported by default. This service does not return a response. If an invalid reset type is provided, the reset request is ignored. If the system reset fails, the resulting system state is unspecified. __Table 40\. Request Data__ | Word | Name | Type | Description | | ---- | ----------- | ------ | ---------------------------------------------------------------------- | | 0 | RESET\_TYPE | uint32 | Reset type.Refer [Reset Types](#section-reset-types) for more details. | __Table 41\. Response Data__ | NA | | -- | ### [](#service-group-system%5Fsuspend-servicegroup%5Fid-0x0004)Service Group - SYSTEM\_SUSPEND (SERVICEGROUP\_ID: 0x0004) This service group defines services used to request platform microcontroller to transition the system into a suspend state, also called a sleep state. The suspend state `SUSPEND_TO_RAM` is supported by default by the platform and if the application processor requests for `SUSPEND_TO_RAM`, it’s implicit for the platform microcontroller that all the application processors except the one requesting are in `STOPPED` state and necessary state saving in the RAM has been complete. The following table lists the services in the SYSTEM\_SUSPEND service group: __Table 42\. SYSTEM\_SUSPEND Services__ | Service ID | Service Name | Request Type | | ---------- | ----------------------------- | --------------- | | 0x01 | SYSSUSP\_ENABLE\_NOTIFICATION | NORMAL\_REQUEST | | 0x02 | SYSSUSP\_GET\_ATTRIBUTES | NORMAL\_REQUEST | | 0x03 | SYSSUSP\_SUSPEND | NORMAL\_REQUEST | #### [](#section-suspend-types)Suspend Types RPMI supports suspend types and their values as defined by SBI specification. Refer to [**SBI System Sleep Types**](https://github.com/riscv-non-isa/riscv-sbi-doc/blob/master/src/ext-sys-suspend.adoc#table%5Fsusp%5Fsleep%5Ftypes)in the RISC-V SBI Specification \[[1](bibliography.html#bib-sbi)\] for the `SUSPEND_TYPE` definition. #### [](#system-suspend-notifications)Notifications This service group does not support any events for notification. #### [](#service-syssusp%5Fenable%5Fnotification-service%5Fid-0x01)Service: SYSSUSP\_ENABLE\_NOTIFICATION (SERVICE\_ID: 0x01) This service allows the application processor to subscribe to `SYSTEM_SUSPEND`service group notifications. The platform may optionally support notifications for events that may occur. The platform microcontroller can send these notification messages to the application processor if they are implemented and the application processor has subscribed to them. The supported events are described in [Notifications](#system-suspend-notifications). __Table 43\. Request Data__ | Word | Name | Type | Description | | ---- | ---------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | EVENT\_ID | uint32 | The event to be subscribed for notification. | | 1 | REQ\_STATE | uint32 | Requested event notification state.Change or query the current state of EVENT\_ID notification. 0: Disable. 1: Enable. 2: Return current state. Any other values of REQ\_STATE field other than the defined ones are reserved for future use. | __Table 44\. Response Data__ | Word | Name | Type | Description | | ---- | -------------- | ------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Event is subscribed successfully. RPMI\_ERR\_INVALID\_PARAM EVENT\_ID or REQ\_STATE is invalid. RPMI\_ERR\_NOT\_SUPPORTED Notification for the EVENT\_ID is not supported. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | CURRENT\_STATE | uint32 | Current EVENT\_ID notification state. 0: Notification is disabled. 1: Notification is enabled. In case of REQ\_STATE = 0 or 1, the CURRENT\_STATE will return the requested state.In case of an error, the value of CURRENT\_STATE is unspecified. | #### [](#service-syssusp%5Fget%5Fattributes-service%5Fid-0x02)Service: SYSSUSP\_GET\_ATTRIBUTES (SERVICE\_ID: 0x02) This service is used to discover the attributes of a suspend type. The attribute flags for a suspend type indicate whether a `SUSPEND_TYPE` is supported. Additionally, the flags specify whether a `SUSPEND_TYPE` supports a resume address. __Table 45\. Request Data__ | Word | Name | Type | Description | | ---- | ------------- | ------ | ---------------------------------------------------------------------------- | | 0 | SUSPEND\_TYPE | uint32 | Suspend type.Refer [Suspend Types](#section-suspend-types) for more details. | __Table 46\. Response Data__ | Word | Name | Type | Description | | ---- | ------ | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Attributes returned successfully. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | FLAGS | uint32 | Suspend type attributes. Bits Description \[31:2\] _Reserved_ and must be 0. \[1\] Resume Address Support.If a SUSPEND\_TYPE supports custom resume address which platform must configure for the resuming application processor. 0b1: Supported. 0b0: Not supported. \[0\] Suspend type Support. 0b1: Supported. 0b0: Not supported. | #### [](#service-syssusp%5Fsuspend-service%5Fid-0x03)Service: SYSSUSP\_SUSPEND (SERVICE\_ID: 0x03) This service is used to request the platform microcontroller to transition the system in a suspend state. This service returns successfully when the platform microcontroller accepts the system suspend request. The application processor which called this service must then enter into a quiesced state such as WFI. The platform microcontroller will transition the system to the requested`SUSPEND_TYPE` upon the successful transition of the application processor into the supported quiesced state. The mechanism for detecting the quiesced state of the application processor is platform specific. The application processor must only request supported suspend types, discovered using the `SYSSUSP_GET_ATTRIBUTES` service. If a suspend type does not support the custom resume address that the application processor can discover through the `SYSSUSP_GET_ATTRIBUTES` service then the `RESUME_ADDR_LOW` and `RESUME_ADDR_HIGH` will be ignored and the application processor will resume from the `pc` (program counter) after the instruction that put the application processor in the quiesced state, such as the `WFI` instruction. __Table 47\. Request Data__ | Word | Name | Type | Description | | ---- | ------------------ | ------ | ---------------------------------------------------------------------------- | | 0 | HART\_ID | uint32 | Hart ID of the calling hart. | | 1 | SUSPEND\_TYPE | uint32 | Suspend type.Refer [Suspend Types](#section-suspend-types) for more details. | | 2 | RESUME\_ADDR\_LOW | uint32 | Lower 32-bit address. | | 3 | RESUME\_ADDR\_HIGH | uint32 | Upper 32-bit address. | __Table 48\. Response Data__ | Word | Name | Type | Description | | ---- | ------ | ----- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. Suspend request has been accepted. RPMI\_ERR\_INVALID\_PARAM HART\_ID or SUSPEND\_TYPE is invalid. RPMI\_ERR\_INVALID\_ADDR Resume address is invalid. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | ### [](#service-group-hart%5Fstate%5Fmanagement-servicegroup%5Fid-0x0005)Service Group - HART\_STATE\_MANAGEMENT (SERVICEGROUP\_ID: 0x0005) This service group defines services to control and manage the application processor (hart) power states. Hart power states include power on, power off, suspend modes, etc. A hart is identified by a 32-bit identifier called `HART_ID`. In a platform, depending on the sharing of power controls and common resources, the harts can be grouped in a hierarchical topology to form cores, clusters, nodes, etc. In such cases the power state change for a hart can affect the entire hierarchical group in which the hart is located, requiring coordination for the power state change. RPMI supports the coordination mechanisms and hart power states defined by the RISC-V SBI Specification \[[1](bibliography.html#bib-sbi)\]. The following table lists the services in the HART\_STATE\_MANAGEMENT service group: __Table 49\. HART\_STATE\_MANAGEMENT Services__ | Service ID | Service Name | Request Type | | ---------- | ------------------------- | --------------- | | 0x01 | HSM\_ENABLE\_NOTIFICATION | NORMAL\_REQUEST | | 0x02 | HSM\_GET\_HART\_STATUS | NORMAL\_REQUEST | | 0x03 | HSM\_GET\_HART\_LIST | NORMAL\_REQUEST | | 0x04 | HSM\_GET\_SUSPEND\_TYPES | NORMAL\_REQUEST | | 0x05 | HSM\_GET\_SUSPEND\_INFO | NORMAL\_REQUEST | | 0x06 | HSM\_HART\_START | NORMAL\_REQUEST | | 0x07 | HSM\_HART\_STOP | NORMAL\_REQUEST | | 0x08 | HSM\_HART\_SUSPEND | NORMAL\_REQUEST | #### [](#section-hart-states)Hart States Hart HSM states and the HSM state machine supported by the RPMI are defined in the RISC-V SBI Specification \[[1](bibliography.html#bib-sbi)\]. Refer to[**HSM States**](https://github.com/riscv-non-isa/riscv-sbi-doc/blob/master/src/ext-hsm.adoc#table%5Fhsm%5Fstates). From a hart perspective a start state means hart has started execution of instructions and stop state means that hart is not executing the instructions. The platform can implement the stop state either by powering down the hart or just putting the hart in a platform supported low-power state. #### [](#section-hart-suspend-types)Hart Suspend Types The RPMI supports the hart suspend types encoding as defined in RISC-V SBI Specification \[[1](bibliography.html#bib-sbi)\]. Refer to [**HSM Suspend Types**](https://github.com/riscv-non-isa/riscv-sbi-doc/blob/master/src/ext-hsm.adoc#table%5Fhsm%5Fhart%5Fsuspend%5Ftypes). The values for the platform supported suspend types are discovered through a service defined in this service group. #### [](#hsm-notifications)Notifications This service group does not support any events for notification. #### [](#service-hsm%5Fenable%5Fnotification-service%5Fid-0x01)Service: HSM\_ENABLE\_NOTIFICATION (SERVICE\_ID: 0x01) This service allows the application processor to subscribe to `HART_STATE_MANAGEMENT`service group notifications. The platform may optionally support notifications for events that may occur. The platform microcontroller can send these notification messages to the application processor if they are implemented and the application processor has subscribed to them. The supported events are described in [Notifications](#hsm-notifications). __Table 50\. Request Data__ | Word | Name | Type | Description | | ---- | ---------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | EVENT\_ID | uint32 | The event to be subscribed for notification. | | 1 | REQ\_STATE | uint32 | Requested event notification state.Change or query the current state of EVENT\_ID notification. 0: Disable. 1: Enable. 2: Return current state. Any other values of REQ\_STATE field other than the defined ones are reserved for future use. | __Table 51\. Response Data__ | Word | Name | Type | Description | | ---- | -------------- | ------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Event is subscribed successfully. RPMI\_ERR\_INVALID\_PARAM EVENT\_ID or REQ\_STATE is invalid. RPMI\_ERR\_NOT\_SUPPORTED Notification for the EVENT\_ID is not supported. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | CURRENT\_STATE | uint32 | Current EVENT\_ID notification state. 0: Notification is disabled. 1: Notification is enabled. In case of REQ\_STATE = 0 or 1, the CURRENT\_STATE will return the requested state.In case of an error, the value of CURRENT\_STATE is unspecified. | #### [](#service-hsm%5Fget%5Fhart%5Fstatus-service%5Fid-0x02)Service: HSM\_GET\_HART\_STATUS (SERVICE\_ID: 0x02) This service returns the current HSM state of a hart. If a hart is in an invalid state that is not a defined HSM state, an error code will be set in the `STATUS` field. __Table 52\. Request Data__ | Word | Name | Type | Description | | ---- | -------- | ------ | ----------- | | 0 | HART\_ID | uint32 | Hart ID. | __Table 53\. Response Data__ | Word | Name | Type | Description | | ---- | ----------- | ------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. RPMI\_ERR\_INVALID\_PARAM HART\_ID is invalid. RPMI\_ERR\_INVALID\_STATE Hart is in invalid state. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | HART\_STATE | uint32 | Hart state.Refer [Hart States](#section-hart-states) for more details. | #### [](#service-hsm%5Fget%5Fhart%5Flist-service%5Fid-0x03)Service: HSM\_GET\_HART\_LIST (SERVICE\_ID: 0x03) This service retrieves the list of Hart IDs managed by this service group. If the number of words required for all available Hart IDs exceeds the number of words that can be returned in one acknowledgement message then the platform microcontroller will set the `REMAINING` and `RETURNED` fields accordingly and only return the Hart IDs which can be accommodated. The application processor may need to call this service again with the appropriate `START_INDEX` until the`REMAINING` field returns `0`. __Table 54\. Request Data__ | Word | Name | Type | Description | | ---- | ------------ | ------ | --------------------------- | | 0 | START\_INDEX | uint32 | Start index of the Hart ID. | __Table 55\. Response Data__ | Word | Name | Type | Description | | ---- | ------------- | ------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. RPMI\_ERR\_INVALID\_PARAM START\_INDEX is invalid. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | REMAINING | uint32 | Remaining number of Hart IDs to be returned. | | 2 | RETURNED | uint32 | Number of Hart IDs returned in this request. | | 3 | HART\_ID\[N\] | uint32 | Hart IDs list. | #### [](#service-hsm%5Fget%5Fsuspend%5Ftypes-service%5Fid-0x04)Service: HSM\_GET\_SUSPEND\_TYPES (SERVICE\_ID: 0x04) This service gets the list of all supported suspend types for a hart. The suspend types in the list must be ordered based on increasing power savings. If the number of words required for all available suspend types exceeds the number of words that can be returned in one acknowledgement message then the platform microcontroller will set the `REMAINING` and `RETURNED` fields accordingly and only return the suspend types which can be accommodated. The application processor may need to call this service again with the appropriate `START_INDEX` until the `REMAINING` field returns `0`. The attributes and details of each suspend type can be discovered using the`HSM_GET_SUSPEND_INFO` service. __Table 56\. Request Data__ | Word | Name | Type | Description | | ---- | ------------ | ------ | ------------------------------------------------------------------------------------------------------------------ | | 0 | START\_INDEX | uint32 | Start index of the Hart ID. 0 for the first call, subsequent calls will use the next index of the remaining items. | __Table 57\. Response Data__ | Word | Name | Type | Description | | ---- | ------------------ | ------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. RPMI\_ERR\_INVALID\_PARAM START\_INDEX is invalid. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | REMAINING | uint32 | Remaining number of suspend types to be returned. | | 2 | RETURNED | uint32 | Number of suspend types returned in this request. | | 3 | SUSPEND\_TYPE\[N\] | uint32 | Suspend types.Refer [Hart Suspend Types](#section-hart-suspend-types) for more details. | #### [](#service-hsm%5Fget%5Fsuspend%5Finfo-service%5Fid-0x05)Service: HSM\_GET\_SUSPEND\_INFO (SERVICE\_ID: 0x05) This service is used to get the attributes of a suspend type. The attributes of a suspend type include various associated latencies. The entry latency for a suspend type is the maximum amount of time that the hart requires to transition from the execution state to the low-power suspend state when the hart invokes an quiesced state entry mechanism such as WFI. The exit latency is the time required by the hart to transition from the suspend state to execution state after the wakeup event. For each suspend state there is a point of no return after which the suspend state transition cannot be reversed. The wakeup latency is the maximum time required by the hart to transition from the point of no return to the execution state. If the platform returns `0` in `WAKEUP_LATENCY` then the application processor can use the `(ENTRY_LATENCY + EXIT_LATENCY)` as the wakeup latency. The minimum residency of a suspend state is the minimum time the application processor must remain in that suspend state to become energy efficient compared to the shallower suspend state. | | The energy is also consumed while entering and exiting a suspend state. The application processor must spend time equal to or more than minimum residency to justify the energy cost of entering and exiting that suspend state. | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | | The application processor entering into a deeper suspend state with a high minimum residency will incur a longer wakeup latency due to more time required to exit that suspend state. If the predicted idle time by the application processor is less than the minimum residency of a suspend state, it should select the shallower suspend state to minimize the wakeup latency to achieve energy savings. | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | __Table 58\. Request Data__ | Word | Name | Type | Description | | ---- | ------------- | ------ | -------------------------------------------------------------------------------------- | | 0 | SUSPEND\_TYPE | uint32 | Suspend type.Refer [Hart Suspend Types](#section-hart-suspend-types) for more details. | __Table 59\. Response Data__ | Word | Name | Type | Description | | ---- | --------------- | ------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. RPMI\_ERR\_INVALID\_PARAM SUSPEND\_TYPE is invalid. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | FLAGS | uint32 | Bits Description \[31: 1\] _Reserved_ and must be 0. \[0\] Local timer running status. 0b1: Local timer stops when the hart is suspended. 0b0: Local timer does not stop when hart is suspended. | | 2 | ENTRY\_LATENCY | uint32 | Entry latency in microseconds. | | 3 | EXIT\_LATENCY | uint32 | Exit latency in microseconds. | | 4 | WAKEUP\_LATENCY | uint32 | Wakeup latency in microseconds. | | 5 | MIN\_RESIDENCY | uint32 | Minimum residency time in microseconds. | #### [](#service-hsm%5Fhart%5Fstart-service%5Fid-0x06)Service: HSM\_HART\_START (SERVICE\_ID: 0x06) This service is used to start the execution on a hart identified by `HART_ID`. This service requires a start address which is the physical address from which the target hart will start execution. Successful completion of this service means that the hart has started execution from the specified start address. __Table 60\. Request Data__ | Word | Name | Type | Description | | ---- | ----------------- | ------ | ----------------------------------------- | | 0 | HART\_ID | uint32 | Hart ID of the target hart to be started. | | 1 | START\_ADDR\_LOW | uint32 | Lower 32-bit of the start address. | | 2 | START\_ADDR\_HIGH | uint32 | Upper 32-bit of the start address. | __Table 61\. Response Data__ | Word | Name | Type | Description | | ---- | ------ | ----- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully and hart has started. RPMI\_ERR\_INVALID\_PARAM HART\_ID or start address is invalid. RPMI\_ERR\_ALREADY Hart is already in transition to start state or has already started. RPMI\_ERR\_DENIED Hart is not in stopped state. RPMI\_ERR\_HW\_FAULT Failed due to hardware fault. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | #### [](#service-hsm%5Fhart%5Fstop-service%5Fid-0x07)Service: HSM\_HART\_STOP (SERVICE\_ID: 0x07) This service stops the execution on the calling hart. The mechanism for stopping the hart is platform specific. The hart can be powered down, if supported, or put into the deepest available sleep state. This service returns successful if the platform microcontroller has successfully acknowledged that the target hart can be stopped. The hart upon successful acknowledgement can perform the final context saving if required and must enter into a quiesced state such as WFI which can be detected and allow the platform microcontroller to proceed to stop the hart. The mechanism to detect the hart quiesced state by the platform microcontroller is platform specific. Once the hart is stopped, it can only be restarted by explicitly invoking the`HSM_HART_START` service call explicitly by any other hart. __Table 62\. Request Data__ | Word | Name | Type | Description | | ---- | -------- | ------ | ---------------------------- | | 0 | HART\_ID | uint32 | Hart ID of the calling hart. | __Table 63\. Response Data__ | Word | Name | Type | Description | | ---- | ------ | ----- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully and hart is stopped. RPMI\_ERR\_ALREADY Hart is already in transition to stop state or has already stopped. RPMI\_ERR\_DENIED Hart is not in start state. RPMI\_ERR\_HW\_FAULT Failed due to hardware failure. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | #### [](#service-hsm%5Fhart%5Fsuspend-service%5Fid-0x08)Service: HSM\_HART\_SUSPEND (SERVICE\_ID: 0x08) This service is used to put a hart in a low-power suspend state supported by the platform. Each suspend type is a 32-bit value which is discovered through the`HSM_GET_SUSPEND_TYPES` service. This service returns successful if the platform microcontroller has successfully acknowledged that the target hart can be put into the requested `SUSPEND_TYPE`state. The target hart after the successful acknowledgement must enter into a quiesced state such as WFI which can be detected and allow the platform microcontroller complete the suspend state transition. The mechanism to detect the hart quiesced state by the platform microcontroller is platform specific. For non-retentive suspend state the hart will resume its execution from the provided resume address. __Table 64\. Request Data__ | Word | Name | Type | Description | | ---- | ------------------ | ------ | -------------------------------------------------------------------------------------- | | 0 | HART\_ID | uint32 | Hart ID of the calling hart. | | 1 | SUSPEND\_TYPE | uint32 | Suspend type.Refer [Hart Suspend Types](#section-hart-suspend-types) for more details. | | 2 | RESUME\_ADDR\_LOW | uint32 | Lower 32-bit of the resume address.Only used for non-retentive suspend types. | | 3 | RESUME\_ADDR\_HIGH | uint32 | Upper 32-bit of the resume address.Only used for non-retentive suspend types. | __Table 65\. Response Data__ | Word | Name | Type | Description | | ---- | ------ | ----- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. RPMI\_ERR\_INVALID\_PARAM HART\_ID or SUSPEND\_TYPE is invalid. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | ### [](#service-group-cppc-servicegroup%5Fid-0x0006)Service Group - CPPC (SERVICEGROUP\_ID: 0x0006) This service group defines the services to control application processor performance by managing a set of registers per application processor that are used for performance management and control. The ACPI CPPC (Collaborative Processor Performance Control) is an abstract and flexible mechanism that allows application processor to collaborate with the platform microcontroller to control the performance. The CPPC extension defined in the RISC-V SBI specification \[[1](bibliography.html#bib-sbi)\] defines the register IDs for the standard CPPC registers, along with additional registers also required by the application processor. The ACPI CPPC specification \[[3](bibliography.html#bib-acpi)\] provides the details of the CPPC registers and also provides details on the performance control mechanism through CPPC. This service group works with the abstract performance scale defined by the ACPI CPPC and is managed by the platform which is responsible for the conversion between the abstract performance level and the internal performance operating point. The platform may have multiple application processors that share the actual performance controls like clock, voltage regulator and others depending on the platform. In such cases a performance level change for one application processor will affect the entire the group sharing the controls. Its the responsibility of the power and performance management software running on the application processor and the platform to coordinate and manage the group level performance changes. The following table lists the services in the CPPC service group: __Table 66\. CPPC Services__ | Service ID | Service Name | Request Type | | ---------- | -------------------------------- | --------------- | | 0x01 | CPPC\_ENABLE\_NOTIFICATION | NORMAL\_REQUEST | | 0x02 | CPPC\_PROBE\_REG | NORMAL\_REQUEST | | 0x03 | CPPC\_READ\_REG | NORMAL\_REQUEST | | 0x04 | CPPC\_WRITE\_REG | NORMAL\_REQUEST | | 0x05 | CPPC\_GET\_FAST\_CHANNEL\_REGION | NORMAL\_REQUEST | | 0x06 | CPPC\_GET\_FAST\_CHANNEL\_OFFSET | NORMAL\_REQUEST | | 0x07 | CPPC\_GET\_HART\_LIST | NORMAL\_REQUEST | #### [](#cppc-fast-channel)CPPC Fast-channel The CPPC service group defines the fast-channels to be used by the application processor to request performance changes and to obtain performance change feedback for an application processor from the platform microcontroller. A fast-channel shared memory layout is specific to the CPPC service group. The data written in a fast-channel do not follow the conventional RPMI message format. The simple data format supported by the fast-channel allows faster processing of the performance change requests made through a fast-channel and faster read of the performance feedback values supported over the fast-channel. The CPPC service group defines two types of fast-channel for each application processor. If fast-channels are supported then each application processor must be assigned both types fast-channel. ##### [](#performance-request-fast-channel)Performance Request Fast-channel In this fast-channel the application processor will either write the desired performance level in case of normal mode or the minimum and maximum performance level in case of Autonomous (CPPC2) mode in the fast-channel. Otherwise the application processor can call the service`CPPC_WRITE_REG` for the `DesiredPerformanceRegister` or`MinimumPerformanceRegister` and `MaximumPerformanceRegister`. The supported values in this fast-channel which depends on the CPPC mode, either normal or autonomous mode is discoverable through `CPPC_GET_FAST_CHANNEL_REGION`service. The size of this fast-channel type is `8 bytes`. __Table 67\. CPPC Performance Request Fast-channel Layout__ | CPPC Mode | Layout | | ----------------------- | ----------------------------------------------------------------------------------- | | Normal Mode | Offset Value (32-bit) 0x0 Desired performance level. 0x4 _Reserved_ and must be 0. | | Autonomous (CPPC2) mode | Offset Value (32-bit) 0x0 Minimum performance level. 0x4 Maximum performance level. | ##### [](#performance-feedback-fast-channel)Performance Feedback Fast-channel In this fast-channel the application processor will read the supported value for estimating the delivered performance as performance feedback for an application processor. The application processor current frequency (Hz) is used for performance feedback in this fast-channel. The platform microcontroller must write the frequency of an application processor in the fast-channel whenever it changes. The size of this fast-channel type is `8 bytes`. __Table 68\. CPPC Performance Feedback Fast-channel Layout__ | Offset | Value (32-bit) | | ------ | ----------------------------------- | | 0x0 | Current frequency low 32-bit (Hz). | | 0x4 | Current frequency high 32-bit (Hz). | ##### [](#cppc-fast-channel-shared-memory-region)CPPC Fast-channel Shared Memory Region The size of the shared memory region containing all the fast-channels for all the managed application processors must be a `power-of-2`. The `base-address` and `size`(bytes) of this shared memory region can be discovered using the service `CPPC_GET_FAST_CHANNEL_REGION`. The `base-address` of the shared memory region must be aligned to `8 bytes` which is maximum size of a fast-channel in both the types. The offsets of fast-channels of both types for an application processor are aligned to `8 bytes`. The offset of both fast-channel types in the shared memory region can be discovered through service `CPPC_GET_FAST_CHANNEL_OFFSET`. The offsets discovered can be added to the `base-address` of the shared memory region to form the address of Performance Request fast-channel and Performance Feedback fast-channel for an application processor. ##### [](#performance-request-fast-channel-doorbell)Performance Request Fast-channel Doorbell A doorbell can also be supported for this fast-channel type which is shared between all the application processors. The doorbell, if supported must be a memory mapped register with write access. The doorbell details and attributes such as doorbell register address, doorbell write value can be discovered by the application processor through the`CPPC_GET_FAST_CHANNEL_REGION` service. The doorbell register address is the physical address of the register. The doorbell write value is the value which must be written in the doorbell register to trigger the doorbell interrupt. The width of the doorbell write value must be equal to the doorbell register width. | | The write value may also contains other set bits which must persist on every write to the doorbell register. | | --------------------------------------------------------------------------------------------------------------- | #### [](#cppc-notifications)Notifications This service group does not support any events for notification. #### [](#service-cppc%5Fenable%5Fnotification-service%5Fid-0x01)Service: CPPC\_ENABLE\_NOTIFICATION (SERVICE\_ID: 0x01) This service allows the application processor to subscribe to `CPPC`service group notifications. The platform may optionally support notifications for events that may occur. The platform microcontroller can send these notification messages to the application processor if they are implemented and the application processor has subscribed to them. The supported events are described in [Notifications](#cppc-notifications). __Table 69\. Request Data__ | Word | Name | Type | Description | | ---- | ---------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | EVENT\_ID | uint32 | The event to be subscribed for notification. | | 1 | REQ\_STATE | uint32 | Requested event notification state.Change or query the current state of EVENT\_ID notification. 0: Disable. 1: Enable. 2: Return current state. Any other values of REQ\_STATE field other than the defined ones are reserved for future use. | __Table 70\. Response Data__ | Word | Name | Type | Description | | ---- | -------------- | ------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Event is subscribed successfully. RPMI\_ERR\_INVALID\_PARAM EVENT\_ID or REQ\_STATE is invalid. RPMI\_ERR\_NOT\_SUPPORTED Notification for the EVENT\_ID is not supported. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | CURRENT\_STATE | uint32 | Current EVENT\_ID notification state. 0: Notification is disabled. 1: Notification is enabled. In case of REQ\_STATE = 0 or 1, the CURRENT\_STATE will return the requested state.In case of an error, the value of CURRENT\_STATE is unspecified. | #### [](#service-cppc%5Fprobe%5Freg-service%5Fid-0x02)Service: CPPC\_PROBE\_REG (SERVICE\_ID: 0x02) This service is used to probe a CPPC register implementation status for a application processor. If the CPPC register `reg_id` is implemented then the length in bits is returned in `REG_LENGTH` field. If the register is not supported or invalid then the `REG_LENGTH` will be `0`. __Table 71\. Request Data__ | Word | Name | Type | Description | | ---- | -------- | ------ | ----------------- | | 0 | REG\_ID | uint32 | CPPC register ID. | | 1 | HART\_ID | uint32 | Hart ID. | __Table 72\. Response Data__ | Word | Name | Type | Description | | ---- | ----------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS CPPC register probed successfully. RPMI\_ERR\_INVALID\_PARAM HART\_ID or REG\_ID is invalid. RPMI\_ERR\_NOT\_SUPPORTED REG\_ID is not supported. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | REG\_LENGTH | uint32 | Register length (bits). | #### [](#service-cppc%5Fread%5Freg-service%5Fid-0x03)Service: CPPC\_READ\_REG (SERVICE\_ID: 0x03) This service is used to read a CPPC register. If the fast-channels are supported, a read of the `DesiredPerformanceRegister` or`MinimumPerformanceRegister` and `MaximumPerformanceRegister` through this service will return the current desired performance level or minimum and maximum performance level limit depending on the CPPC mode from the fast-channel of a application processor. __Table 73\. Request Data__ | Word | Name | Type | Description | | ---- | -------- | ------ | ----------------- | | 0 | REG\_ID | uint32 | CPPC register ID. | | 1 | HART\_ID | uint32 | Hart ID. | __Table 74\. Response Data__ | Word | Name | Type | Description | | ---- | ---------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. RPMI\_ERR\_INVALID\_PARAM HART\_ID or REG\_ID is invalid. RPMI\_ERR\_NOT\_SUPPORTED REG\_ID is not supported. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | DATA\_LOW | uint32 | Lower 32-bit of the data. | | 2 | DATA\_HIGH | uint32 | Upper 32-bit of data. This will be 0 if the register is of 32-bit length. | #### [](#service-cppc%5Fwrite%5Freg-service%5Fid-0x04)Service: CPPC\_WRITE\_REG (SERVICE\_ID: 0x04) This service is used to write a CPPC register. If the fast-channels are supported the application processor must only write desired performance level in the fast-channel instead of writing into the`DesiredPerformanceRegister` through this service. Similarly, in case of the autonomous mode the application processor must write minimum and maximum limit levels into the fast-channel instead of calling this service for`MinimumPerformanceRegister` and `MaximumPerformanceRegister`. Otherwise the writes to these registers may be ignored. __Table 75\. Request Data__ | Word | Name | Type | Description | | ---- | ---------- | ------ | -------------------------------------------------------------------------- | | 0 | REG\_ID | uint32 | CPPC register ID. | | 1 | HART\_ID | uint32 | Hart ID. | | 2 | DATA\_LOW | uint32 | Lower 32-bit of data. | | 3 | DATA\_HIGH | uint32 | Upper 32-bit of data. This is ignored if the register is of 32-bit length. | __Table 76\. Response Data__ | Word | Name | Type | Description | | ---- | ------ | ----- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. RPMI\_ERR\_INVALID\_PARAM HART\_ID or REG\_ID is invalid. RPMI\_ERR\_NOT\_SUPPORTED REG\_ID is not supported. RPMI\_ERR\_DENIED REG\_ID is read only. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | #### [](#service-cppc%5Fget%5Ffast%5Fchannel%5Fregion-service%5Fid-0x05)Service: CPPC\_GET\_FAST\_CHANNEL\_REGION (SERVICE\_ID: 0x05) This service is used to get the details of the shared memory region containing all the fast-channels, attributes of the fast-channel and the details of the doorbell if supported. The doorbell details are unspecified and considered invalid if the Performance Request fast-channel doorbell (`FLAGS[0] = 0`) is not supported and must not be used. __Table 77\. Request Data__ | NA | | -- | __Table 78\. Response Data__ | Word | Name | Type | Description | | ---- | ------------------ | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. RPMI\_ERR\_NOT\_SUPPORTED Fast-channels not supported. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | FLAGS | uint32 | Bits Description \[31:5\] _Reserved_ and must be 0. \[4:3\] CPPC mode. 0b00: Normal mode. Desired performance level for performance change. 0b01: Autonomous mode. Performance limit change. Platform can choose the level in the requested limit. 0b10 - 0b11: Reserved. \[2:1\] Performance Request fast-channel doorbell register width. 0b00: 8-bit. 0b01: 16-bit. 0b10: 32-bit. 0b11: Reserved. \[0\] Performance Request fast-channel doorbell support. 0b1: Supported. 0b0: Not supported. | | 2 | REGION\_ADDR\_LOW | uint32 | Lower 32-bit of the fast-channels shared memory region physical address. | | 3 | REGION\_ADDR\_HIGH | uint32 | Upper 32-bit of the fast-channels shared memory region physical address. | | 4 | REGION\_SIZE\_LOW | uint32 | Lower 32-bit of the fast-channels shared memory region size. | | 5 | REGION\_SIZE\_HIGH | uint32 | Upper 32-bit of the fast-channels shared memory region size. | | 6 | DB\_ADDR\_LOW | uint32 | Lower 32-bit of doorbell register address for Performance Request fast-channel. | | 7 | DB\_ADDR\_HIGH | uint32 | Upper 32-bit of doorbell register address for Performance Request fast-channel. | | 8 | DB\_WRITE\_VALUE | uint32 | 32-bit doorbell write value for Performance Request fast-channel.If the doorbell register width is less than 32-bit, the lower bits in this field equal to the doorbell register width must be used as write value. | #### [](#service-cppc%5Fget%5Ffast%5Fchannel%5Foffset-service%5Fid-0x06)Service: CPPC\_GET\_FAST\_CHANNEL\_OFFSET (SERVICE\_ID: 0x06) This service is used to get the offsets of Performance Request fast-channel and Performance Feedback fast-channel for an application processor in the shared memory region containing all the fast-channels. __Table 79\. Request Data__ | Word | Name | Type | Description | | ---- | -------- | ------ | ----------- | | 0 | HART\_ID | uint32 | Hart ID. | __Table 80\. Response Data__ | Word | Name | Type | Description | | ---- | ---------------------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. RPMI\_ERR\_INVALID\_PARAM HART\_ID is invalid. RPMI\_ERR\_NOT\_SUPPORTED Fast-channels not supported. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | PERF\_REQUEST\_OFFSET\_LOW | uint32 | Lower 32-bit of a Performance Request fast-channel offset. | | 2 | PERF\_REQUEST\_OFFSET\_HIGH | uint32 | Upper 32-bit of a Performance Request fast-channel offset. | | 3 | PERF\_FEEDBACK\_OFFSET\_LOW | uint32 | Lower 32-bit of a Performance Feedback fast-channel offset. | | 4 | PERF\_FEEDBACK\_OFFSET\_HIGH | uint32 | Upper 32-bit of a Performance Feedback fast-channel offset. | #### [](#service-cppc%5Fget%5Fhart%5Flist-service%5Fid-0x07)Service: CPPC\_GET\_HART\_LIST (SERVICE\_ID: 0x07) This service retrieves the list of Hart IDs managed by this service group for performance control. If the number of words required for all available Hart IDs exceeds the number of words that can be returned in one acknowledgement message then the platform microcontroller will set the `REMAINING` and `RETURNED` fields accordingly and only return the Hart IDs which can be accommodated. The application processor may need to call this service again with the appropriate `START_INDEX` until the`REMAINING` field returns `0`. __Table 81\. Request Data__ | Word | Name | Type | Description | | ---- | ------------ | ------ | -------------------------- | | 0 | START\_INDEX | uint32 | Starting index of Hart ID. | __Table 82\. Response Data__ | Word | Name | Type | Description | | ---- | ------------- | ------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. RPMI\_ERR\_INVALID\_PARAM START\_INDEX is invalid. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | REMAINING | uint32 | Remaining number of Hart IDs to be returned. | | 2 | RETURNED | uint32 | Number of Hart IDs returned in this request. | | 3 | HART\_ID\[N\] | uint32 | Hart IDs. | ### [](#service-group-voltage-servicegroup%5Fid-0x0007)Service Group - VOLTAGE (SERVICEGROUP\_ID: 0x0007) This service group is used to control the voltage level of the voltage domains. A voltage domain is the logical grouping of one or more devices powered by a single controllable voltage source. A system may have multiple voltage domains and the services defined in this service group are used to manage and control the voltage levels of these voltage domains. Each voltage domain is identified by`DOMAIN_ID` which is a 32-bit integer starting from `0`. The following table lists the services in the VOLTAGE service group: __Table 83\. VOLTAGE Services__ | Service ID | Service Name | Request Type | | ---------- | ---------------------------- | --------------- | | 0x01 | VOLT\_ENABLE\_NOTIFICATION | NORMAL\_REQUEST | | 0x02 | VOLT\_GET\_NUM\_DOMAINS | NORMAL\_REQUEST | | 0x03 | VOLT\_GET\_ATTRIBUTES | NORMAL\_REQUEST | | 0x04 | VOLT\_GET\_SUPPORTED\_LEVELS | NORMAL\_REQUEST | | 0x05 | VOLT\_SET\_CONFIG | NORMAL\_REQUEST | | 0x06 | VOLT\_GET\_CONFIG | NORMAL\_REQUEST | | 0x07 | VOLT\_SET\_LEVEL | NORMAL\_REQUEST | | 0x08 | VOLT\_GET\_LEVEL | NORMAL\_REQUEST | #### [](#voltage-level-format-section)Voltage Level Format There are two types of voltage level formats supported in the VOLTAGE service group. The voltage levels are represented as a group. ##### [](#discrete-format)Discrete Format A set of discrete voltage levels arranged in a sequence, starting from the lowest value at the lowest index and increasing sequentially to higher levels. The following table shows the structure of the discrete format. ```c [voltage0, voltage1, voltage2, ... , voltage(N-1)] where: voltage0 < voltage1 < voltage2 < ... < voltage(N-1) ``` | Word | Name | Description | | ---- | ------- | ------------------------------------------ | | 0 | VOLTAGE | Discrete voltage level in microvolts (uV). | ##### [](#linear-range-format)Linear Range Format A linear range of voltage levels with a constant step size. The following table shows the structure of the linear range voltage format. ```c [voltage_min, voltage_max, voltage_step] ``` Multi-linear range format can be supported by having multiple `linear range` tuple arranged in an continuous array format. ```c [voltage_min0, voltage_max0, voltage_step0], [voltage_min1, voltage_max1, voltage_step1], ... [voltage_min(N-1), voltage_max(N-1), voltage_step(N-1)], ``` The format must be packed sequentially such that `voltage_max0 < voltage_min1, voltage_max1 < voltage_min2` and so on. Each linear range is considered as a single voltage level. | Word | Name | Description | | ---- | ------------ | --------------------------------------------------- | | 0 | VOLTAGE\_MIN | Lower boundary of voltage level in microvolts (uV). | | 1 | VOLTAGE\_MAX | Upper boundary of voltage level in microvolts (uV). | | 2 | STEP | Step size in microvolts (uV). | #### [](#voltage-notifications)Notifications This service group does not support any events for notification. #### [](#service-volt%5Fenable%5Fnotification-service%5Fid-0x01)Service: VOLT\_ENABLE\_NOTIFICATION (SERVICE\_ID: 0x01) This service allows the application processor to subscribe to `VOLTAGE`service group notifications. The platform may optionally support notifications for events that may occur. The platform microcontroller can send these notification messages to the application processor if they are implemented and the application processor has subscribed to them. The supported events are described in [Notifications](#voltage-notifications). __Table 84\. Request Data__ | Word | Name | Type | Description | | ---- | ---------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | EVENT\_ID | uint32 | The event to be subscribed for notification. | | 1 | REQ\_STATE | uint32 | Requested event notification state.Change or query the current state of EVENT\_ID notification. 0: Disable. 1: Enable. 2: Return current state. Any other values of REQ\_STATE field other than the defined ones are reserved for future use. | __Table 85\. Response Data__ | Word | Name | Type | Description | | ---- | -------------- | ------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Event is subscribed successfully. RPMI\_ERR\_INVALID\_PARAM EVENT\_ID or REQ\_STATE is invalid. RPMI\_ERR\_NOT\_SUPPORTED Notification for the EVENT\_ID is not supported. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | CURRENT\_STATE | uint32 | Current EVENT\_ID notification state. 0: Notification is disabled. 1: Notification is enabled. In case of REQ\_STATE = 0 or 1, the CURRENT\_STATE will return the requested state.In case of an error, the value of CURRENT\_STATE is unspecified. | #### [](#service-volt%5Fget%5Fnum%5Fdomains-service%5Fid-0x02)Service: VOLT\_GET\_NUM\_DOMAINS (SERVICE\_ID: 0x02) This service is used to query the number of voltage domains available in the system. __Table 86\. Request Data__ | NA | | -- | __Table 87\. Response Data__ | Word | Name | Type | Description | | ---- | ------------ | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | NUM\_DOMAINS | uint32 | Number of voltage domains. | #### [](#service-volt%5Fget%5Fattributes-service%5Fid-0x03)Service: VOLT\_GET\_ATTRIBUTES (SERVICE\_ID: 0x03) Each domain may support multiple voltage levels, which are permitted by the domain for operation. The number of levels indicates the total count of voltage levels supported within a voltage domain. Transition latency denotes the maximum time required for the voltage to stabilize upon a change in the regulator. The `FLAGS`field encodes the voltage format supported by the hardware, including discrete and linear range formats." The `NUM_LEVELS` field returns the number of discrete voltage in case discrete format and number of linear range tuple in linear range voltage format. Each domain can support only one voltage level format. Additional voltage formats can be accommodated in the future if required. __Table 88\. Request Data__ | Word | Name | Type | Description | | ---- | ---------- | ------ | ------------------ | | 0 | DOMAIN\_ID | uint32 | Voltage domain ID. | __Table 89\. Response Data__ | Word | Name | Type | Description | | ---- | -------------- | ----------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 0 | STATUS | int32 | Return error code. Error Code Description RPMI\_SUCCESS Service completed successfully. RPMI\_ERR\_INVALID\_PARAM DOMAIN\_ID is invalid. Other errors [\[table\_error\_codes\]](#table%5Ferror%5Fcodes). | | 1 | FLAGS | uint32 | Bits Description \[31:4\] _Reserved_ and must be 0. \[3:1\] Voltage format.Refer to [Voltage Level Format](#voltage-level-format-section) for more details. 0b000: Discrete format. 0b001: Linear range format. 0b010 - 0b111: Reserved. \[0\] Voltage domain control support. 0b0: Voltage domain can be enabled/disabled. 0b1: Voltage domain is always-on, voltage value can be changed in the supported voltage range. | | 2 | NUM\_LEVELS | uint32 | The number of voltage levels (number of arrays) supported by the domain. If the voltage level format is a linear range, then each linear range is considered a single voltage level.