Next-generation wireless, data center, automotive, and industrial applications demand significant increases in compute acceleration without increasing board area and while remaining power efficient. With Moore's Law and Dennard Scaling no longer following their traditional trajectories, simply porting existing hardware architectures to next-generation process nodes is not sufficient to meet the system-level performance requirements of these applications – architectural innovation is critical to overcome the diminishing gains from process scaling. To address this need, AMD has developed an innovative processing technology, the AI Engine, as part of the Versal architecture.
AI Engines are very long instruction word (VLIW), single instruction multiple data (SIMD) vector processors optimized for both machine learning and advanced signal processing workloads. AI Engines are arranged in 2D arrays within a given device, with the number of individual processors varying by device, allowing for scalability to meet the compute needs of a breadth of applications. AMD offers two primary types of AI Engines within the Versal portfolio: AI Engines (AIE), which are balanced between signal processing and machine learning workloads, and AI Engines-Machine Learning (AIE-ML), which are optimized for machine learning workloads. Although both types of AI Engines can be used for either machine learning or signal processing functions, differences in native data type support, memory, and other capabilities impact which workloads each AI Engine type is best suited for.
AI Engines-Machine Learning enable high performance per watt with low latency for inference workloads. AIE-ML v2 is the second generation of the AI Engines-Machine Learning architecture. AIE-ML v2 offers several architectural enhancements over the first-generation AIE-ML architecture, including up to two times compute/tile, new native data type support, and improved energy efficiency.
This architecture manual details the specifics of the AIE-ML v2 architecture.