AOCL-DLP supports multiple precision formats to balance accuracy and performance:
Terminology |
C Type |
Bits |
Typical Use |
|---|---|---|---|
f32 |
|
32 |
Training, high-accuracy inference |
f16 |
|
16 |
Memory-efficient inference (native on Zen5+) |
bf16 |
|
16 |
Training and inference (good range, lower precision) |
s8 / u8 |
|
8 |
Quantized weights and activations |
s4 / u4 |
packed in |
4 |
Extreme weight quantization |
s32 / u32 |
|
32 |
Accumulation and intermediate results |
Supported GEMM Data Type Combinations:
Input A |
Input B |
Output C |
Accumulator |
Function Suffix |
|---|---|---|---|---|
u8/s8 |
s8 |
s32/s8/u8/f32/bf16 |
s32 |
<u8|s8>s8s32o<s32|s8|u8|f32|bf16> |
bf16/f32 |
s8 |
s32/s8/u8/f32/bf16 |
s32 |
<bf16|f32>s8s32o<s32|s8|u8|f32|bf16> |
bf16 |
s4/u4 |
f32/bf16 |
f32 |
bf16<s4|u4>f32o<f32|bf16> |
bf16 |
bf16 |
f32/bf16 |
f32 |
bf16bf16f32o<f32|bf16> |
f32 |
f32 |
f32 |
f32 |
f32f32f32of32 |
f16 |
f16 |
f16 |
f16 |
f16f16f16of16 |
u8 |
s4 |
s32 |
s32 |
u8s4s32os32 |
Note
u8s4s32os32only hasreorderandget_reorder_buf_sizeAPIs (no GEMM).s8s8s32o<f32|bf16>also has_sym_quantvariants for symmetric quantization.Mixed-precision reorder:
f32obf16(converts f32 input to bf16 reordered output).
For detailed information about data types, see the Types API Reference and the Library Overview Wiki.