Metrics - Metrics - 57368

uProf User Guide

Document_ID
57368
Release_Date
2025-06-09
Revision
5.1 English
The performance metrics for AMD EPYC™ “Zen 2”, “Zen 3”, “Zen 4”, and "Zen 5" core architecture processors are listed here.

Performance metrics for AMD EPYC™ “Zen 2” core architecture processors

Table 1. Performance Metrics for AMD EPYC™ “Zen 2”
Metric Group Metric Description
ipc Utilization (%) Percentage of time the core was running, that is non-idle time.
Eff Freq Core Effective Frequency (CEF) without halted cycles over the sampling period, reported in GHz. The metric is based on CEF = (APERF / TSC) * P0Freq. APERF is incremented in proportion to the actual number of core cycles while the core is in C6 state.
IPC Instructions Per Cycle (IPC) is the average number of instructions retired per CPU cycle. This is measured using Core PMC events PMCx0C0 [Retired Instructions] and PMCx076 [CPU Clocks not Halted]. These PMC events are counted in both OS and User mode.
CPI Cycles Per Instruction (CPI) is the multiplicative inverse of IPC metric. This is one of the basic performance metrics indicating how cache misses, branch mis-predictions, memory latencies, and other bottlenecks are affecting the execution of an application. A lower CPI value is better.
Branch Mis-prediction Ratio The ratio between mis-predicted branches and retired branch instructions.
fp Retired SSE/AVX Flops (GFLOPs) The number of retired SSE/AVX FLOPs.
Mixed SSE/AVX Stalls Mixed SSE/AVX stalls. This metric is in per thousand instructions (PTI).
l1 IC(32B) Fetch Miss Ratio Instruction cache fetch miss ratio.
DC Access All data cache (DC) accesses. This metric is in PTI.
l2 L2 Access All the L2 cache accesses. This metric is in PTI.
L2 Access from IC Miss The L2 cache accesses from IC miss. This metric is in PTI.
L2 Access from DC Miss The L2 cache accesses from DC miss. This metric is in PTI.
L2 Access from HWPF The L2 cache accesses from L2 hardware pre-fetching. This metric is in PTI.
L2 Miss All the L2 cache misses. This metric is in PTI.
L2 Miss from IC Miss The L2 cache misses from IC miss. This metric is in PTI.
L2 Miss from DC Miss The L2 cache misses from DC miss. This metric is in PTI.
L2 Miss from HWPF The L2 cache misses from L2 hardware pre-fetching. This metric is in PTI.
L2 Hit All the L2 cache hits. This metric is in PTI.
L2 Hit from IC Miss The L2 cache hits from IC miss. This metric is in PTI.
L2 Hit from DC Miss The L2 cache hits from DC miss. This metric is in PTI.
L2 Hit from HWPF The L2 cache hits from L2 hardware pre-fetching. This metric is in PTI.
tlb L1 ITLB Miss The instruction fetches the misses in the L1 Instruction Translation Lookaside Buffer (ITLB), but hit in the L2- ITLB plus the ITLB reloads originating from page table walker. The table walk requests are made for L1-ITLB miss and L2-ITLB misses. This metric is in PTI.
L2 ITLB Miss The number of ITLB reloads from page table walker due to L1-ITLB and L2-ITLB misses. This metric is in PTI.
L1 DTLB Miss The number of L1 Data Translation Lookaside Buffer (DTLB) misses from load store micro-ops. This event counts both L2-DTLB hit and L2-DTLB miss. This metric is in PTI.
L2 DTLB Miss The number of L2 Data Translation Lookaside Buffer (DTLB)missed from load store micro-ops. This metric is in PTI.
l3 L3 Access The count of L3 cache accesses.
L3 Miss The L3 cache miss. This metric is in PTI.
L3 Miss (%) The L3 cache miss percentage. This metric is in PTI.
Ave L3 Miss Latency Average L3 miss latency in core cycles.
Memory

Mem Ch-A RdBw (GB/s)

Mem Ch-A WrBw (GB/s)

...

Memory Read and Write bandwidth in GB/s for all the channels.
xgmi

xGMI0 BW (GB/s)

xGMI1 BW (GB/s)

xGMI2 BW (GB/s)

xGMI3 BW (GB/s)

Approximate xGMI outbound data bytes in GB/s for all the remote links.
pcie

PCIe0 (GB/s)

PCIe1 (GB/s)

PCIe2 (GB/s)

PCIe3 (GB/s)

Approximate PCIe bandwidth in GB/s.

Performance Metrics for AMD EPYCTM “Zen 3” Core Architecture Processors

Table 2. Performance Metrics for AMD EPYC™ “Zen 3”
Metric Group Metric Description
ipc Utilization(%) Percentage of time the core was running, that is non-idle time.
Eff Freq Core Effective Frequency (CEF) without halted cycles over the sampling period, reported in GHz. The metric is based on CEF = (APERF / TSC) * P0Freq. APERF is incremented in proportion to the actual number of core cycles while the core is in C6 state.
IPC Instructions Per Cycle (IPC)is the average number of instructions retired per CPU cycle. This is measured using Core PMC events PMCx0C0 [Retired Instructions] and PMCx076 [CPU Clocks not Halted]. These PMC events are counted in both OS and User mode.
CPI Cycles Per Instruction (CPI) is the multiplicative inverse of IPC metric. This is one of the basic performance metrics indicating how cache misses, branch mis-predictions, memory latencies, and other bottlenecks are affecting the execution of an application. A lower CPI value is better.
Branch Mis-prediction Ratio The ratio between mis-predicted branches and retired branch instructions.
fp Retired SSE/AVX Flops (GFLOPs) The number of retired SSE/AVX FLOPs.
Mixed SSE/AVX Stalls Mixed SSE/AVX stalls. This metric is in per thousand instructions (PTI).
l1 IC (32B) Fetch Miss Ratio Instruction cache fetch miss ratio.
Op Cache (64B) Fetch Miss Ratio Operation cache fetch miss ratio.
IC Access All instruction cache accesses. This metric is in PTI.
IC Miss The instruction cache miss. This metric is in PTI.
DC Access All the DC accesses. This metric is in PTI.
l2 L2 Access All the L2 cache accesses. This metric is in PTI.
L2 Access from IC Miss The L2 cache accesses from IC miss. This metric is in PTI.
L2 Access from DC Miss The L2 cache accesses from DC miss. This metric is in PTI.
L2 Access from HWPF The L2 cache accesses from L2 hardware pre-fetching. This metric is in PTI.
L2 Miss All the L2 cache misses. This metric is in PTI.
L2 Miss from IC Miss The L2 cache misses from IC miss. This metric is in PTI.
L2 Miss from DC Miss The L2 cache misses from DC miss. This metric is in PTI.
L2 Miss from HWPF The L2 cache misses from L2 hardware pre-fetching. This metric is in PTI.
L2 Hit All the L2 cache hits. This metric is in PTI.
L2 Hit from IC Miss The L2 cache hits from IC miss. This metric is in PTI.
L2 Hit from DC Miss The L2 cache hits from DC miss. This metric is in PTI.
L2 Hit from HWPF The L2 cache hits from L2 hardware pre-fetching. This metric is in PTI.
tlb L1 ITLB Miss The instruction fetches the misses in the L1 Instruction Translation Lookaside Buffer (ITLB), but hit in the L2- ITLB plus the ITLB reloads originating from page table walker. The table walk requests are made for L1-ITLB miss and L2-ITLB misses. This metric is in PTI.
L2 ITLB Miss The number of ITLB reloads from page table walker due to L1-ITLB and L2-ITLB misses. This metric is in PTI.
L1 DTLB Miss The number of L1 Data Translation Lookaside Buffer (DTLB) misses from load store micro-ops. This event counts both L2-DTLB hit and L2-DTLB miss. This metric is in PTI.
L2 DTLB Miss The number of L2 Data Translation Lookaside Buffer (DTLB)missed from load store micro-ops. This metric is in PTI.
All TLBs Flushed All the TLBs flushed. This metric is in PTI.
dc DC Fills from Same CCX The number of DC fills from local L2 cache to the core or different L2 cache in the same CCX or L3 cache that belongs to the CCX. This metric is in PTI.
DC Fills from different CCX in same node The number of DC fills from cache of different CCX in the same package (node). This metric is in PTI.
DC Fills from Local Memory The number of DC fills from DRAM or IO connected in the same package (node). This metric is in PTI.
DC Fills from Remote CCX Cache The number of DC fills from cache of CCX in the different package (node). This metric is in PTI.
DC Fills from Remote Memory The number of DC fills from DRAM or IO connected in the different package (node). This metric is in PTI.
All DC Fills The total number of DC fills from all the data sources. This metric is in PTI.
l3 L3 Access The L3 cache accesses. This metric is in PTI.
L3 Miss The L3 cache miss. This metric is in PTI.
L3 Miss (%) The L3 cache miss percentage. This metric is in PTI.
Ave L3 Miss Latency The average L3 miss latency in core cycles.
Memory

Mem Ch-A RdBw (GB/s)

Mem Ch-A WrBw (GB/s)

...

Memory Read and Write bandwidth in GB/s for all the channels.
xgmi

xGMI0 BW (GB/s)

xGMI1 BW (GB/s)

xGMI2 BW (GB/s)

xGMI3 BW (GB/s)

Approximate xGMI outbound data bytes in GB/s for all the remote links.
swpfdc

SwPfDC Fills from DRAM or IO connected in remote node (pti)

SwPfDC Fills from CCX Cache in remote node (pti)

SwPfDC Fills from DRAM or IO connected in local node (pti)

SwPfDC Fills from Cache of another CCX in local node (pti)

SwPfDC Fills from L3 or different L2 in same CCX (pti)

SwPfDC Fills from L2 (pti)

Software prefetch data cache from various nodes and CCX.
hwpfdc

HwPfDC Fills from DRAM or IO connected in remote node (pti)

HwPfDC Fills from CCX Cache in remote node (pti)

HwPfDC Fills from DRAM or IO connected in local node (pti)

HwPfDC Fills from Cache of another CCX in local node (pti)

HwPfDC Fills from L3 or different L2 in same CCX (pti)

HwPfDC Fills From L2 (pti)

Hardware prefetch data cache from various nodes and CCX.

Performance Metrics for AMD EPYC™ “Zen 4” and later Core Architecture Processors

Table 3. Performance Metrics for AMD EPYC™ “Zen 4” and later versions
Metric Group Metric Description
ipc Utilization (%) Percentage of time the core was running, that is non-idle time.
System Time (%) Percentage of time in kernel mode.
User Time (%) Percentage of time in user mode.
System instruction (%) Percentage of retired instruction in kernel mode.
User instructions (%) Percentage of retired instructions in user mode.
Eff Freq (MHz) Core Effective Frequency (CEF) without halted cycles over the sampling period, reported in MHz. The metric is based on CEF = (APERF / TSC) * P0Freq. APERF is incremented in proportion to the actual number of core cycles while the core is in C6 state.
IPC (Sys + User) Instructions Per Cycle (IPC) is the average number of instructions retired per CPU cycle. This is measured using Core PMC events PMCx0C0 [Retired Instructions] and PMCx076 [CPU Clocks not Halted]. These PMC events are counted in both OS and User mode.
IPC (Sys) Instructions in kernel mode per cpu cycles in kernel mode. IPC of kernel mode.
IPC (User) Instructions in user mode per cpu cycles in user mode. IPC of user mode.
CPI (Sys + User) Cycles Per Instruction (CPI) is the multiplicative inverse of IPC metric. This is one of the basic performance metrics indicating how cache misses, branch mis-predictions, memory latencies, and other bottlenecks are affecting the execution of an application. A lower CPI value is better.
CPI (Sys) Cycles per Instructions for kernel mode.
CPI (User) Cycles per Instructions for user mode.
Giga Instructions Per Sec The number of retired giga instructions per second
Locked Instructions (pti) The number of retired lock instructions in PTI.
Retired Branches (pti) The number of retired branch instructions in PTI.
Retired Branches Mispredicted (pti) The number of retired mis-predicted branch instructions in PTI.
fp Retired SSE/AVX Flops (GFLOPs) The number of retired SSE/AVX FLOPs.
FP Dispatch Faults (PTI) The number of floating point instruction dispatch faults. This metric is in per thousand instructions (PTI).
avx_imix Packed 512-bit FP Ops Retired (%) Percentage of retired packed 512-bit floating point operations out of retired floating-point operations.
Packed 256-bit FP Ops Retired (%) Percentage of retired packed 256-bit floating point operations out of retired floating-point operations.
Packed 128-bit FP Ops Retired (%) Percentage of retired packed 128-bit floating point operations out of retired floating-point operations.
Scalar/MMX/x87 FP Ops Retired (%) Percentage of retired scalar, mmx and x87 floating point operations out of retired floating-point operations.
SSE/AVX Instructions Retired (pti) The number of retired SSE/AVX per thousand instructions.
MMX Instructions Retired (pti) The number of retired MMX per thousand instructions.
x87 Instructions Retired (pti) The number of retired x87 per thousand instructions.
l1 IC (32B) Fetch Miss Ratio Instruction cache fetch miss ratio.
Op Cache Fetch Miss Ratio Operation cache (64B) fetch miss ratio.
IC Access (PTI) Instruction cache access in PTI.
IC Miss (PTI) Instruction cache Miss in PTI.
DC Access (PTI) All the data cache (DC) accesses. This metric is in PTI.
dc All DC Fills (pti) Total number of data cache fills in PTI.
DC Fills From Same CCX (pti) Total number of data cache fills from the same CCX in PTI.
DC Fills From different CCX in same node (pti) Total number of data cache fills from different CCX in the same Numa node in PTI.
DC Fills From Local Memory (pti) Total number of data cache fills from local memory in PTI.
DC Fills From Remote CCX Cache (pti) Total number of data cache fills from CCX in the remote Numa node in PTI.
DC Fills From Remote Memory (pti) Total number of data cache fills from remote memory in PTI.
Remote DRAM Reads % Percentage of data cache fills from remote memory out of total number of number of data cache fills.
l2 L2 Access All the L2 cache accesses. This metric is in PTI.
L2 Access from IC Miss The L2 cache accesses from IC miss. This metric is in PTI.
L2 Access from DC Miss The L2 cache accesses from DC miss. This metric is in PTI.
L2 Access from L2 HWPF The L2 cache accesses from L2 hardware pre-fetching. This metric is in PTI.
L2 Miss All the L2 cache misses. This metric is in PTI.
L2 Miss from IC Miss The L2 cache misses from IC miss. This metric is in PTI.
L2 Miss from DC Miss The L2 cache misses from DC miss. This metric is in PTI.
L2 Miss from HWPF The L2 cache misses from L2 hardware pre-fetching. This metric is in PTI.
L2 Hit All the L2 cache hits. This metric is in PTI.
L2 Hit from IC Miss The L2 cache hits from IC miss. This metric is in PTI.
L2 Hit from DC Miss The L2 cache hits from DC miss. This metric is in PTI.
L2 Hit from HWPF The L2 cache hits from L2 hardware pre-fetching. This metric is in PTI.
tlb L1 ITLB Miss The instruction fetches the misses in the L1 Instruction Translation Lookaside Buffer (ITLB), but hit in the L2- ITLB plus the ITLB reloads originating from page table walker. The table walk requests are made for L1-ITLB miss and L2-ITLB misses. This metric is in PTI.
L2 ITLB Miss The number of ITLB reloads from page table walker due to L1-ITLB and L2-ITLB misses. This metric is in PTI.
L1 DTLB Miss The number of L1 Data Translation Lookaside Buffer (DTLB) misses from load store micro-ops. This event counts both L2-DTLB hit and L2-DTLB miss. This metric is in PTI.
L2 DTLB Miss The number of L2 Data Translation Lookaside Buffer (DTLB) missed from load store micro-ops. This metric is in PTI.
All TLBs Flushed All the flushed TLBs.
l3 L3 Access The number of L3 cache accesses.
L3 Hit % Percentage of L3 Hit.
L3 Miss The number of L3 cache misses.
L3 Miss (%) Percentage of L3 cache miss.
AveL3 Miss Latency Average L3 miss latency in core cycles.
L3 Access (pti) The number of L3 cache accesses per thousand instructions.
L3 Miss (pti) The number of L3 cache misses per thousand instructions.
Memory Total Memory Bw (GB/s) Total read and write memory bandwidth.

Local DRAM Read Data Bytes (GB/s)

Local DRAM Write Data Bytes (GB/s)

DRAM read and write data bytes for a local processor.
Remote DRAM Read Data Bytes (GB/s) DRAM read and write data bytes for a remote processor.
Remote DRAM Write Data Bytes (GB/s)
Mem Ch-A RdBw (GB/s) Memory read and write bandwidth in GB/s for all the channels.
Mem Ch-A WrBw (GB/s)
ccm_bw Local Socket Inbound Data to CPU Moderator (CCM) 0 at Interface 0 (GB/s) Reports data traffic to CCM at interfaces 0 and 1.
Local Socket Outbound Data from CPU Moderator (CCM) 0 at Interface 0 (GB/s)
Remote Socket Inbound Data to CPU Moderator (CCM) 0 at Interface 0 (GB/s)
Remote Socket Outbound Data from CPU Moderator (CCM) 0 at Interface 0 (GB/s)
Local Inbound Read Data Bytes(GB/s) Local inbound data bytes to the CPU, for example, read data.
Local Outbound Write Data Bytes (GB/s) Local outbound data bytes from the CPU, for example, write data.
Remote Inbound Read Data Bytes(GB/s) Remote socket inbound data bytes to the CPU, for example, read data.
Remote Outbound Write Data Bytes (GB/s) Remote socket outbound data bytes from the CPU for example, write data.
xgmi xGMI Outbound Data Bytes (GB/s) Total outbound data bytes in Gigabytes per second.

dma

(not available in AMD “Zen 1”, AMD “Zen 2”, and AMD “Zen 3” processors)

Total Upstream DMA Read Write Data Bytes (GB/s) Total upstream DMA including read and write.
Local Upstream DMA Read Data Bytes (GB/s) Local upstream DMA read data bytes.
Local Upstream DMA Write Data Bytes (GB/s) Local upstream DMA write data bytes.
Remote Upstream DMA Read Data Bytes (GB/s) Remote socket upstream DMA read data bytes
Remote Upstream DMA Write Data Bytes (GB/s) Remote socket upstream DMA write data bytes.
pcie

PCIe0 (GB/s)

PCIe1 (GB/s)

PCIe2 (GB/s)

PCIe3 (GB/s)

Approximate PCIe bandwidth in GB/s.

Total PCIE Bandwidth (GB/s)

Total PCIE Rd Bandwidth (GB/s)

Total PCIE Wr Bandwidth (GB/s)

Total PCIE Bandwidth Local (GB/s)

Total PCIE Bandwidth Remote (GB/s)

Total PCIE Rd Bandwidth Local (GB/s)

Total PCIE Wr Bandwidth Local (GB/s)

Total PCIE Rd Bandwidth Remote (GB/s)

Total PCIE Wr Bandwidth Remote (GB/s)

Quad 0 PCIE Rd Bandwidth Local (GB/s)

Quad 0 PCIE Wr Bandwidth Local (GB/s)

Quad 0 PCIE Rd Bandwidth Remote (GB/s)

Quad 0 PCIE Wr Bandwidth Remote (GB/s)

Quad 1 PCIE Rd Bandwidth Local (GB/s)

Quad 1 PCIE Wr Bandwidth Local (GB/s)

Quad 1 PCIE Rd Bandwidth Remote (GB/s)

Quad 1 PCIE Wr Bandwidth Remote (GB/s)

Quad 2 PCIE Rd Bandwidth Local (GB/s)

Quad 2 PCIE Wr Bandwidth Local (GB/s)

Quad 2 PCIE Rd Bandwidth Remote (GB/s)

Quad 2 PCIE Wr Bandwidth Remote (GB/s)

Quad 3 PCIE Rd Bandwidth Local (GB/s)

Quad 3 PCIE Wr Bandwidth Local (GB/s)

Quad 3 PCIE Rd Bandwidth Remote (GB/s)

Quad 3 PCIE Wr Bandwidth Remote (GB/s)

PCIe bandwidth for read and write transactions, local, and remote node bandwidth. Per quad PCIe bandwidth.
swpfdc

SwPf DC Fills from DRAM or IO connected in remote node (pti)

SwPfDC Fills from CCX Cache in remote node (pti)

SwPf DC Fills from DRAM or IO connected in local node (pti)

SwPfDC Fills from Cache of another CCX in local node (pti)

SwPfDC Fills from L3 or different L2 in same CCX (pti)

SwPfDC Fills from L2 (pti)

Software prefetch data cache from various nodes and CCX.
hwpfdc

HwPfDC Fills from DRAM or IO connected in remote node (pti)

HwPfDC Fills from CCX Cache in remote node (pti) HwPf DC Fills from DRAM or IO connected in local node (pti)

HwPfDC Fills from Cache of another CCX in local node (pti)

HwPfDC Fills from L3 or different L2 in same CCX (pti)

HwPfDC Fills From L2 (pti)

Hardware prefetch data cache from various nodes and CCX.
pipeline_util Total_Dispatch_Slots Up to 6 instructions can be dispatched in one cycle.
SMT_Disp_contention Percentage of unused dispatch slots as other thread was selected.
Frontend_Bound Percentage of dispatch slots that remained unused as the front end did not supply enough instructions/operations.
Bad_Speculation Percentage of unused dispatch slots as other thread was selected.
Backend_Bound Percentage of dispatch slots that remained unused because of the back end stalls.
Retiring Percentage of dispatch slots used by the retired operations.
IPC Instructions per cycle.
Frontend_Bound.Latency Percentage of dispatch slots that remained unused because of a latency bottleneck in the front end, such as Instruction Cache or ITLB misses.
Frontend_Bound.BW Percentage of dispatch slots that remained unused because of a bandwidth bottleneck in the front end, such as decode bandwidth or Op Cache fetch bandwidth.
Bad_Speculation.Mispredicts Percentage of dispatched ops that were flushed due to branch mis-predicts.
Bad_Speculation.Pipeline_R estarts Percentage of dispatched ops that were flushed due to the pipeline restarts (resyncs).
Backend_Bound.Memory Percentage of dispatched slots that remained unused because of stalls due to the memory subsystem.
Backend_Bound.CPU Percentage of dispatched slots that remained unused because of stalls not related to the memory subsystem.
Retiring.Fastpath Percentage of dispatch slots used by the retired fastpath operations.
Retiring.Microcode Percentage of dispatch slots used by the retired microcode operations.
Note: Memory channels are available with package level.