Making the correct decisions for partitioning your application, using the information in the previous sections, enables the most efficient use of the Versal adaptive SoC. Therefore, it delivers the best performance. Efficient use of all engines ensures best performance. The AI Engine runs with high utilization, delivering high performance at low power. This results in high performance per watt for high compute applications.
Where AI Engine core utilization is low, consider running additional low utilization kernels on the same AI Engine tile. These kernels are not able to run simultaneously, but if the application allows it, this method can improve overall AI Engine array efficiency.
Clock gating of unused tiles is also possible within the AI Engine array and is turned on by default. When no components are enabled, the tile is considered unused. This means no memory banks, interconnect (including route-thrus), or AI Engine core usage. Where necessary, you can use the bounding box constructor to guide the placement of your kernels to ensure maximum clock gating can be achieved. For more information, refer to the AI Engine Tools and Flows User Guide (UG1076).