You need to check more than matrix multiplication size support. Verify that the two loads, the store, and the compute remain equally optimized.
You can view a complete efficiency table, including matrix load and vector compute details, here: Performance Table