Performance metrics for AMD EPYC™ “Zen 2” core architecture processors
| Metric Group | Metric | Description |
|---|---|---|
| ipc | Utilization (%) | Percentage of time the core was running, that is non-idle time. |
| Eff Freq | Core Effective Frequency (CEF) without halted cycles over the sampling period, reported in GHz. The metric is based on CEF = (APERF / TSC) * P0Freq. APERF is incremented in proportion to the actual number of core cycles while the core is in C6 state. | |
| IPC | Instructions Per Cycle (IPC) is the average number of instructions retired per CPU cycle. This is measured using Core PMC events PMCx0C0 [Retired Instructions] and PMCx076 [CPU Clocks not Halted]. These PMC events are counted in both OS and User mode. | |
| CPI | Cycles Per Instruction (CPI) is the multiplicative inverse of IPC metric. This is one of the basic performance metrics indicating how cache misses, branch mis-predictions, memory latencies, and other bottlenecks are affecting the execution of an application. A lower CPI value is better. | |
| Branch Mis-prediction Ratio | The ratio between mis-predicted branches and retired branch instructions. | |
| fp | Retired SSE/AVX Flops (GFLOPs) | The number of retired SSE/AVX FLOPs. |
| Mixed SSE/AVX Stalls | Mixed SSE/AVX stalls. This metric is in per thousand instructions (PTI). | |
| l1 | IC(32B) Fetch Miss Ratio | Instruction cache fetch miss ratio. |
| DC Access | All data cache (DC) accesses. This metric is in PTI. | |
| l2 | L2 Access | All the L2 cache accesses. This metric is in PTI. |
| L2 Access from IC Miss | The L2 cache accesses from IC miss. This metric is in PTI. | |
| L2 Access from DC Miss | The L2 cache accesses from DC miss. This metric is in PTI. | |
| L2 Access from HWPF | The L2 cache accesses from L2 hardware pre-fetching. This metric is in PTI. | |
| L2 Miss | All the L2 cache misses. This metric is in PTI. | |
| L2 Miss from IC Miss | The L2 cache misses from IC miss. This metric is in PTI. | |
| L2 Miss from DC Miss | The L2 cache misses from DC miss. This metric is in PTI. | |
| L2 Miss from HWPF | The L2 cache misses from L2 hardware pre-fetching. This metric is in PTI. | |
| L2 Hit | All the L2 cache hits. This metric is in PTI. | |
| L2 Hit from IC Miss | The L2 cache hits from IC miss. This metric is in PTI. | |
| L2 Hit from DC Miss | The L2 cache hits from DC miss. This metric is in PTI. | |
| L2 Hit from HWPF | The L2 cache hits from L2 hardware pre-fetching. This metric is in PTI. | |
| tlb | L1 ITLB Miss | The instruction fetches the misses in the L1 Instruction Translation Lookaside Buffer (ITLB), but hit in the L2- ITLB plus the ITLB reloads originating from page table walker. The table walk requests are made for L1-ITLB miss and L2-ITLB misses. This metric is in PTI. |
| L2 ITLB Miss | The number of ITLB reloads from page table walker due to L1-ITLB and L2-ITLB misses. This metric is in PTI. | |
| L1 DTLB Miss | The number of L1 Data Translation Lookaside Buffer (DTLB) misses from load store micro-ops. This event counts both L2-DTLB hit and L2-DTLB miss. This metric is in PTI. | |
| L2 DTLB Miss | The number of L2 Data Translation Lookaside Buffer (DTLB)missed from load store micro-ops. This metric is in PTI. | |
| l3 | L3 Access | The count of L3 cache accesses. |
| L3 Miss | The L3 cache miss. This metric is in PTI. | |
| L3 Miss (%) | The L3 cache miss percentage. This metric is in PTI. | |
| Ave L3 Miss Latency | Average L3 miss latency in core cycles. | |
| Memory |
Mem Ch-A RdBw (GB/s) Mem Ch-A WrBw (GB/s) ... |
Memory Read and Write bandwidth in GB/s for all the channels. |
| xgmi |
xGMI0 BW (GB/s) xGMI1 BW (GB/s) xGMI2 BW (GB/s) xGMI3 BW (GB/s) |
Approximate xGMI outbound data bytes in GB/s for all the remote links. |
| pcie |
PCIe0 (GB/s) PCIe1 (GB/s) PCIe2 (GB/s) PCIe3 (GB/s) |
Approximate PCIe bandwidth in GB/s. |
Performance Metrics for AMD EPYCTM “Zen 3” Core Architecture Processors
| Metric Group | Metric | Description |
|---|---|---|
| ipc | Utilization(%) | Percentage of time the core was running, that is non-idle time. |
| Eff Freq | Core Effective Frequency (CEF) without halted cycles over the sampling period, reported in GHz. The metric is based on CEF = (APERF / TSC) * P0Freq. APERF is incremented in proportion to the actual number of core cycles while the core is in C6 state. | |
| IPC | Instructions Per Cycle (IPC)is the average number of instructions retired per CPU cycle. This is measured using Core PMC events PMCx0C0 [Retired Instructions] and PMCx076 [CPU Clocks not Halted]. These PMC events are counted in both OS and User mode. | |
| CPI | Cycles Per Instruction (CPI) is the multiplicative inverse of IPC metric. This is one of the basic performance metrics indicating how cache misses, branch mis-predictions, memory latencies, and other bottlenecks are affecting the execution of an application. A lower CPI value is better. | |
| Branch Mis-prediction Ratio | The ratio between mis-predicted branches and retired branch instructions. | |
| fp | Retired SSE/AVX Flops (GFLOPs) | The number of retired SSE/AVX FLOPs. |
| Mixed SSE/AVX Stalls | Mixed SSE/AVX stalls. This metric is in per thousand instructions (PTI). | |
| l1 | IC (32B) Fetch Miss Ratio | Instruction cache fetch miss ratio. |
| Op Cache (64B) Fetch Miss Ratio | Operation cache fetch miss ratio. | |
| IC Access | All instruction cache accesses. This metric is in PTI. | |
| IC Miss | The instruction cache miss. This metric is in PTI. | |
| DC Access | All the DC accesses. This metric is in PTI. | |
| l2 | L2 Access | All the L2 cache accesses. This metric is in PTI. |
| L2 Access from IC Miss | The L2 cache accesses from IC miss. This metric is in PTI. | |
| L2 Access from DC Miss | The L2 cache accesses from DC miss. This metric is in PTI. | |
| L2 Access from HWPF | The L2 cache accesses from L2 hardware pre-fetching. This metric is in PTI. | |
| L2 Miss | All the L2 cache misses. This metric is in PTI. | |
| L2 Miss from IC Miss | The L2 cache misses from IC miss. This metric is in PTI. | |
| L2 Miss from DC Miss | The L2 cache misses from DC miss. This metric is in PTI. | |
| L2 Miss from HWPF | The L2 cache misses from L2 hardware pre-fetching. This metric is in PTI. | |
| L2 Hit | All the L2 cache hits. This metric is in PTI. | |
| L2 Hit from IC Miss | The L2 cache hits from IC miss. This metric is in PTI. | |
| L2 Hit from DC Miss | The L2 cache hits from DC miss. This metric is in PTI. | |
| L2 Hit from HWPF | The L2 cache hits from L2 hardware pre-fetching. This metric is in PTI. | |
| tlb | L1 ITLB Miss | The instruction fetches the misses in the L1 Instruction Translation Lookaside Buffer (ITLB), but hit in the L2- ITLB plus the ITLB reloads originating from page table walker. The table walk requests are made for L1-ITLB miss and L2-ITLB misses. This metric is in PTI. |
| L2 ITLB Miss | The number of ITLB reloads from page table walker due to L1-ITLB and L2-ITLB misses. This metric is in PTI. | |
| L1 DTLB Miss | The number of L1 Data Translation Lookaside Buffer (DTLB) misses from load store micro-ops. This event counts both L2-DTLB hit and L2-DTLB miss. This metric is in PTI. | |
| L2 DTLB Miss | The number of L2 Data Translation Lookaside Buffer (DTLB)missed from load store micro-ops. This metric is in PTI. | |
| All TLBs Flushed | All the TLBs flushed. This metric is in PTI. | |
| dc | DC Fills from Same CCX | The number of DC fills from local L2 cache to the core or different L2 cache in the same CCX or L3 cache that belongs to the CCX. This metric is in PTI. |
| DC Fills from different CCX in same node | The number of DC fills from cache of different CCX in the same package (node). This metric is in PTI. | |
| DC Fills from Local Memory | The number of DC fills from DRAM or IO connected in the same package (node). This metric is in PTI. | |
| DC Fills from Remote CCX Cache | The number of DC fills from cache of CCX in the different package (node). This metric is in PTI. | |
| DC Fills from Remote Memory | The number of DC fills from DRAM or IO connected in the different package (node). This metric is in PTI. | |
| All DC Fills | The total number of DC fills from all the data sources. This metric is in PTI. | |
| l3 | L3 Access | The L3 cache accesses. This metric is in PTI. |
| L3 Miss | The L3 cache miss. This metric is in PTI. | |
| L3 Miss (%) | The L3 cache miss percentage. This metric is in PTI. | |
| Ave L3 Miss Latency | The average L3 miss latency in core cycles. | |
| Memory |
Mem Ch-A RdBw (GB/s) Mem Ch-A WrBw (GB/s) ... |
Memory Read and Write bandwidth in GB/s for all the channels. |
| xgmi |
xGMI0 BW (GB/s) xGMI1 BW (GB/s) xGMI2 BW (GB/s) xGMI3 BW (GB/s) |
Approximate xGMI outbound data bytes in GB/s for all the remote links. |
| swpfdc |
SwPfDC Fills from DRAM or IO connected in remote node (pti) SwPfDC Fills from CCX Cache in remote node (pti) SwPfDC Fills from DRAM or IO connected in local node (pti) SwPfDC Fills from Cache of another CCX in local node (pti) SwPfDC Fills from L3 or different L2 in same CCX (pti) SwPfDC Fills from L2 (pti) |
Software prefetch data cache from various nodes and CCX. |
| hwpfdc |
HwPfDC Fills from DRAM or IO connected in remote node (pti) HwPfDC Fills from CCX Cache in remote node (pti) HwPfDC Fills from DRAM or IO connected in local node (pti) HwPfDC Fills from Cache of another CCX in local node (pti) HwPfDC Fills from L3 or different L2 in same CCX (pti) HwPfDC Fills From L2 (pti) |
Hardware prefetch data cache from various nodes and CCX. |
Performance Metrics for AMD EPYC™ “Zen 4” and later Core Architecture Processors
| Metric Group | Metric | Description |
|---|---|---|
| ipc | Utilization (%) | Percentage of time the core was running, that is non-idle time. |
| System Time (%) | Percentage of time in kernel mode. | |
| User Time (%) | Percentage of time in user mode. | |
| System instruction (%) | Percentage of retired instruction in kernel mode. | |
| User instructions (%) | Percentage of retired instructions in user mode. | |
| Eff Freq (MHz) | Core Effective Frequency (CEF) without halted cycles over the sampling period, reported in MHz. The metric is based on CEF = (APERF / TSC) * P0Freq. APERF is incremented in proportion to the actual number of core cycles while the core is in C6 state. | |
| IPC (Sys + User) | Instructions Per Cycle (IPC) is the average number of instructions retired per CPU cycle. This is measured using Core PMC events PMCx0C0 [Retired Instructions] and PMCx076 [CPU Clocks not Halted]. These PMC events are counted in both OS and User mode. | |
| IPC (Sys) | Instructions in kernel mode per cpu cycles in kernel mode. IPC of kernel mode. | |
| IPC (User) | Instructions in user mode per cpu cycles in user mode. IPC of user mode. | |
| CPI (Sys + User) | Cycles Per Instruction (CPI) is the multiplicative inverse of IPC metric. This is one of the basic performance metrics indicating how cache misses, branch mis-predictions, memory latencies, and other bottlenecks are affecting the execution of an application. A lower CPI value is better. | |
| CPI (Sys) | Cycles per Instructions for kernel mode. | |
| CPI (User) | Cycles per Instructions for user mode. | |
| Giga Instructions Per Sec | The number of retired giga instructions per second | |
| Locked Instructions (pti) | The number of retired lock instructions in PTI. | |
| Retired Branches (pti) | The number of retired branch instructions in PTI. | |
| Retired Branches Mispredicted (pti) | The number of retired mis-predicted branch instructions in PTI. | |
| fp | Retired SSE/AVX Flops (GFLOPs) | The number of retired SSE/AVX FLOPs. |
| FP Dispatch Faults (PTI) | The number of floating point instruction dispatch faults. This metric is in per thousand instructions (PTI). | |
| avx_imix | Packed 512-bit FP Ops Retired (%) | Percentage of retired packed 512-bit floating point operations out of retired floating-point operations. |
| Packed 256-bit FP Ops Retired (%) | Percentage of retired packed 256-bit floating point operations out of retired floating-point operations. | |
| Packed 128-bit FP Ops Retired (%) | Percentage of retired packed 128-bit floating point operations out of retired floating-point operations. | |
| Scalar/MMX/x87 FP Ops Retired (%) | Percentage of retired scalar, mmx and x87 floating point operations out of retired floating-point operations. | |
| SSE/AVX Instructions Retired (pti) | The number of retired SSE/AVX per thousand instructions. | |
| MMX Instructions Retired (pti) | The number of retired MMX per thousand instructions. | |
| x87 Instructions Retired (pti) | The number of retired x87 per thousand instructions. | |
| l1 | IC (32B) Fetch Miss Ratio | Instruction cache fetch miss ratio. |
| Op Cache Fetch Miss Ratio | Operation cache (64B) fetch miss ratio. | |
| IC Access (PTI) | Instruction cache access in PTI. | |
| IC Miss (PTI) | Instruction cache Miss in PTI. | |
| DC Access (PTI) | All the data cache (DC) accesses. This metric is in PTI. | |
| dc | All DC Fills (pti) | Total number of data cache fills in PTI. |
| DC Fills From Same CCX (pti) | Total number of data cache fills from the same CCX in PTI. | |
| DC Fills From different CCX in same node (pti) | Total number of data cache fills from different CCX in the same Numa node in PTI. | |
| DC Fills From Local Memory (pti) | Total number of data cache fills from local memory in PTI. | |
| DC Fills From Remote CCX Cache (pti) | Total number of data cache fills from CCX in the remote Numa node in PTI. | |
| DC Fills From Remote Memory (pti) | Total number of data cache fills from remote memory in PTI. | |
| Remote DRAM Reads % | Percentage of data cache fills from remote memory out of total number of number of data cache fills. | |
| l2 | L2 Access | All the L2 cache accesses. This metric is in PTI. |
| L2 Access from IC Miss | The L2 cache accesses from IC miss. This metric is in PTI. | |
| L2 Access from DC Miss | The L2 cache accesses from DC miss. This metric is in PTI. | |
| L2 Access from L2 HWPF | The L2 cache accesses from L2 hardware pre-fetching. This metric is in PTI. | |
| L2 Miss | All the L2 cache misses. This metric is in PTI. | |
| L2 Miss from IC Miss | The L2 cache misses from IC miss. This metric is in PTI. | |
| L2 Miss from DC Miss | The L2 cache misses from DC miss. This metric is in PTI. | |
| L2 Miss from HWPF | The L2 cache misses from L2 hardware pre-fetching. This metric is in PTI. | |
| L2 Hit | All the L2 cache hits. This metric is in PTI. | |
| L2 Hit from IC Miss | The L2 cache hits from IC miss. This metric is in PTI. | |
| L2 Hit from DC Miss | The L2 cache hits from DC miss. This metric is in PTI. | |
| L2 Hit from HWPF | The L2 cache hits from L2 hardware pre-fetching. This metric is in PTI. | |
| tlb | L1 ITLB Miss | The instruction fetches the misses in the L1 Instruction Translation Lookaside Buffer (ITLB), but hit in the L2- ITLB plus the ITLB reloads originating from page table walker. The table walk requests are made for L1-ITLB miss and L2-ITLB misses. This metric is in PTI. |
| L2 ITLB Miss | The number of ITLB reloads from page table walker due to L1-ITLB and L2-ITLB misses. This metric is in PTI. | |
| L1 DTLB Miss | The number of L1 Data Translation Lookaside Buffer (DTLB) misses from load store micro-ops. This event counts both L2-DTLB hit and L2-DTLB miss. This metric is in PTI. | |
| L2 DTLB Miss | The number of L2 Data Translation Lookaside Buffer (DTLB) missed from load store micro-ops. This metric is in PTI. | |
| All TLBs Flushed | All the flushed TLBs. | |
| l3 | L3 Access | The number of L3 cache accesses. |
| L3 Hit % | Percentage of L3 Hit. | |
| L3 Miss | The number of L3 cache misses. | |
| L3 Miss (%) | Percentage of L3 cache miss. | |
| AveL3 Miss Latency | Average L3 miss latency in core cycles. | |
| L3 Access (pti) | The number of L3 cache accesses per thousand instructions. | |
| L3 Miss (pti) | The number of L3 cache misses per thousand instructions. | |
| Memory | Total Memory Bw (GB/s) | Total read and write memory bandwidth. |
|
Local DRAM Read Data Bytes (GB/s) Local DRAM Write Data Bytes (GB/s) |
DRAM read and write data bytes for a local processor. | |
| Remote DRAM Read Data Bytes (GB/s) | DRAM read and write data bytes for a remote processor. | |
| Remote DRAM Write Data Bytes (GB/s) | ||
| Mem Ch-A RdBw (GB/s) | Memory read and write bandwidth in GB/s for all the channels. | |
| Mem Ch-A WrBw (GB/s) | ||
| ccm_bw | Local Socket Inbound Data to CPU Moderator (CCM) 0 at Interface 0 (GB/s) | Reports data traffic to CCM at interfaces 0 and 1. |
| Local Socket Outbound Data from CPU Moderator (CCM) 0 at Interface 0 (GB/s) | ||
| Remote Socket Inbound Data to CPU Moderator (CCM) 0 at Interface 0 (GB/s) | ||
| Remote Socket Outbound Data from CPU Moderator (CCM) 0 at Interface 0 (GB/s) | ||
| Local Inbound Read Data Bytes(GB/s) | Local inbound data bytes to the CPU, for example, read data. | |
| Local Outbound Write Data Bytes (GB/s) | Local outbound data bytes from the CPU, for example, write data. | |
| Remote Inbound Read Data Bytes(GB/s) | Remote socket inbound data bytes to the CPU, for example, read data. | |
| Remote Outbound Write Data Bytes (GB/s) | Remote socket outbound data bytes from the CPU for example, write data. | |
| xgmi | xGMI Outbound Data Bytes (GB/s) | Total outbound data bytes in Gigabytes per second. |
|
dma (not available in AMD “Zen 1”, AMD “Zen 2”, and AMD “Zen 3” processors) |
Total Upstream DMA Read Write Data Bytes (GB/s) | Total upstream DMA including read and write. |
| Local Upstream DMA Read Data Bytes (GB/s) | Local upstream DMA read data bytes. | |
| Local Upstream DMA Write Data Bytes (GB/s) | Local upstream DMA write data bytes. | |
| Remote Upstream DMA Read Data Bytes (GB/s) | Remote socket upstream DMA read data bytes | |
| Remote Upstream DMA Write Data Bytes (GB/s) | Remote socket upstream DMA write data bytes. | |
| pcie |
PCIe0 (GB/s) PCIe1 (GB/s) PCIe2 (GB/s) PCIe3 (GB/s) |
Approximate PCIe bandwidth in GB/s. |
|
Total PCIE Bandwidth (GB/s) Total PCIE Rd Bandwidth (GB/s) Total PCIE Wr Bandwidth (GB/s) Total PCIE Bandwidth Local (GB/s) Total PCIE Bandwidth Remote (GB/s) Total PCIE Rd Bandwidth Local (GB/s) Total PCIE Wr Bandwidth Local (GB/s) Total PCIE Rd Bandwidth Remote (GB/s) Total PCIE Wr Bandwidth Remote (GB/s) Quad 0 PCIE Rd Bandwidth Local (GB/s) Quad 0 PCIE Wr Bandwidth Local (GB/s) Quad 0 PCIE Rd Bandwidth Remote (GB/s) Quad 0 PCIE Wr Bandwidth Remote (GB/s) Quad 1 PCIE Rd Bandwidth Local (GB/s) Quad 1 PCIE Wr Bandwidth Local (GB/s) Quad 1 PCIE Rd Bandwidth Remote (GB/s) Quad 1 PCIE Wr Bandwidth Remote (GB/s) Quad 2 PCIE Rd Bandwidth Local (GB/s) Quad 2 PCIE Wr Bandwidth Local (GB/s) Quad 2 PCIE Rd Bandwidth Remote (GB/s) Quad 2 PCIE Wr Bandwidth Remote (GB/s) Quad 3 PCIE Rd Bandwidth Local (GB/s) Quad 3 PCIE Wr Bandwidth Local (GB/s) Quad 3 PCIE Rd Bandwidth Remote (GB/s) Quad 3 PCIE Wr Bandwidth Remote (GB/s) |
PCIe bandwidth for read and write transactions, local, and remote node bandwidth. Per quad PCIe bandwidth. | |
| swpfdc |
SwPf DC Fills from DRAM or IO connected in remote node (pti) SwPfDC Fills from CCX Cache in remote node (pti) SwPf DC Fills from DRAM or IO connected in local node (pti) SwPfDC Fills from Cache of another CCX in local node (pti) SwPfDC Fills from L3 or different L2 in same CCX (pti) SwPfDC Fills from L2 (pti) |
Software prefetch data cache from various nodes and CCX. |
| hwpfdc |
HwPfDC Fills from DRAM or IO connected in remote node (pti) HwPfDC Fills from CCX Cache in remote node (pti) HwPf DC Fills from DRAM or IO connected in local node (pti) HwPfDC Fills from Cache of another CCX in local node (pti) HwPfDC Fills from L3 or different L2 in same CCX (pti) HwPfDC Fills From L2 (pti) |
Hardware prefetch data cache from various nodes and CCX. |
| pipeline_util | Total_Dispatch_Slots | Up to 6 instructions can be dispatched in one cycle. |
| SMT_Disp_contention | Percentage of unused dispatch slots as other thread was selected. | |
| Frontend_Bound | Percentage of dispatch slots that remained unused as the front end did not supply enough instructions/operations. | |
| Bad_Speculation | Percentage of unused dispatch slots as other thread was selected. | |
| Backend_Bound | Percentage of dispatch slots that remained unused because of the back end stalls. | |
| Retiring | Percentage of dispatch slots used by the retired operations. | |
| IPC | Instructions per cycle. | |
| Frontend_Bound.Latency | Percentage of dispatch slots that remained unused because of a latency bottleneck in the front end, such as Instruction Cache or ITLB misses. | |
| Frontend_Bound.BW | Percentage of dispatch slots that remained unused because of a bandwidth bottleneck in the front end, such as decode bandwidth or Op Cache fetch bandwidth. | |
| Bad_Speculation.Mispredicts | Percentage of dispatched ops that were flushed due to branch mis-predicts. | |
| Bad_Speculation.Pipeline_R estarts | Percentage of dispatched ops that were flushed due to the pipeline restarts (resyncs). | |
| Backend_Bound.Memory | Percentage of dispatched slots that remained unused because of stalls due to the memory subsystem. | |
| Backend_Bound.CPU | Percentage of dispatched slots that remained unused because of stalls not related to the memory subsystem. | |
| Retiring.Fastpath | Percentage of dispatch slots used by the retired fastpath operations. | |
| Retiring.Microcode | Percentage of dispatch slots used by the retired microcode operations. |