AOCL-DLP leverages AMD CPU features through runtime detection:
ISA |
Available On |
Enables |
|---|---|---|
AVX2 / FMA3 |
AMD Zen1+ |
f32 GEMM, bf16 fallback |
AVX512 |
AMD Zen4+ |
Wider vectors, bf16 x s4/u4 |
AVX512_VNNI |
AMD Zen4+ |
Accelerated integer GEMM |
AVX512_BF16 |
AMD Zen4+ |
Native bfloat16 operations |
AVX512_FP16 |
AMD Zen5+ |
Native half-precision GEMM |
The library automatically selects the best available kernel at runtime. You can override
this with the AOCL_ENABLE_INSTRUCTIONS environment variable (see
Dynamic Dispatch).