Versal devices provide many new dedicated IP blocks, such as the NoC, DDR memory controller, CPM, and AI Engines. These dedicated IP blocks deliver the next level of system-level performance per Watt with high bandwidth data movement and interfaces. The integration of these new dedicated IP blocks upgrades the Versal device programmable logic (PL). This upgrade makes the PL more efficient with the silicon area while maintaining similar PL performance. As a result, following are key differences when working with Versal devices:
- Dedicated IP blocks now support many common hardware functions. These functions previously mapped to the PL in older architectures. This saves significant PL resources.
- Delay distribution of the PL routing interconnect and CLB as well as clock skew and jitter characteristics differ from previous architectures. This difference leads to some logic paths becoming faster and some logic paths becoming slower. The subsequent sections cover the key CLB and clocking differences.
- Next generation applications need more PL RAM resources (including the silicon efficient UltraRAM). This includes special block columns. These increases introduce additional routing delay variations.
When migrating PL functions to Versal devices, legacy RTL designs can require tuning to reduce logic levels around carry operators. It needs to rebalance logic levels between pipeline registers to achieve the same average programmable logic fabric performance as previous generations on equivalent device speedgrades. For hardware design recommendations, refer to the Versal Adaptive SoC Hardware, IP, and Platform Development Methodology Guide (UG1387). For timing closure recommendations, refer to the Versal Adaptive SoC System Integration and Validation Methodology Guide (UG1388).
Traditional FMAX benchmarking compares maximum PL clock speeds of RTL designs. These comparisons are made across different target technologies. This method is not appropriate for evaluating Versal adaptive SoCs against previous generation FPGAs and SoCs for the following reasons:
- Versal architecture optimizes for adaptive acceleration. Therefore, focusing on PL clock speeds does not account for the advantage of the Versal device dedicated IP blocks. Instead, AMD recommends focusing the comparison on system-level compute and throughput metrics.
- The AMD Vitis™ environment or the AMD Vivado™ IP integrator designs the new Versal adaptive SoC high-level building blocks rather than inferring them from RTL. Comparing RTL designs overestimates the Versal device PL utilization, failing to account for utilization and power savings from using the Versal device dedicated IP blocks.