The AI Engine API provides a aie::mmul class template for a vector-based matrix
multiplication. Multiple intermediate matrix multiplication results are accumulated
to give the final result. For more details on the supported matrix multiplication
shapes (M*K*N) and data types, see Matrix Multiplication in the
AI
Engine API User Guide (UG1529).
The aie::mmul operations mul and mac accept
row-major format data for the vector-based matrix multiplication. Then for the
mac operation of aie::mmul, arrange the data by M*K or
K*N. This data shuffling can be done either in
the PL or AI Engine.
This section gives an example of A(64 * 64) x B(64 *
64) matrix multiplication. The data type is int8 x int8. The matrix multiplication shape 4*16*8 is chosen for aie::mmul
operations.
The input data is assumed to be in row-major format. The data is input to the
matrix multiplication kernel as 4*16 matrix and
16*8 matrix. Prior to the matrix multiplication
kernel, the input data is shuffled.
For example, before shuffling, matrix A(64 *
64) is stored in memory with a0,
a1, …, a63,
a64,…, a4096
in order. The aie::mmul operations will use shapes
4*16. The matrix A is partitioned into smaller matrix sized 4*16. For the smaller matrix A00,
a0 to a15,
a64 to a79,
a128 to a143,
and a192 to a207
should be fetched sequentially for aie::mmul. So,
the purpose of the data shuffle is to put a0 to
a15, a64 to
a79, a128 to
a143, and a192
to a207 into continuous storage for the matrix
multiplication kernel. This data shuffling is shown in the following figure.
Similarly, the output data is shuffled. The following figure shows the graph of the design.