The embedded AI Engine system comprises the embedded processor in Versal adaptive SoC and the acceleration logic built into two key categories of acceleration components. The two categories are the traditional PL (LUTs, BRAMs, URAMs, DSPs) and the AI Engines. For Versal adaptive SoC, the embedded compute system comprises the processing system with application processors and real-time processors. For this design type, the use models vary. They range from a sophisticated embedded software stack to a simple bare-metal stack only required to support programming of the acceleration units.
An embedded AI Engine system design runs a software stack. This stack is executed on the built-in embedded processor. The processor serves as an overall control plane for the kernels running on the acceleration units. The flexible runtime (XRT) application programming interfaces (APIs) manage the data transfer between the embedded processors and Versal adaptive SoC. These APIs also have function calls for managing the acceleration units.
The embedded AI Engine interfaces with the PL within a Versal device entirely through hardware streaming interfaces. It can also target systems that leverage the embedded Arm processing subsystem and use the AI Engine and PL, as shown in the following figure.