Vulkan 1.4.352 has been released with the new VK_NV_cooperative_matrix_decode_vector extension developed by NVIDIA. This extension builds on VK_NV_cooperative_matrix2 and allows decoding multiple matrix elements per invocation rather than one at a time, improving efficiency when unpacking quantized weight formats in groups. NVIDIA has already released beta drivers (596.54 for Windows, 595.44.08 for Linux) supporting this extension, which benefits machine learning workloads in Vulkan.

1m read timeFrom phoronix.com
Post cover image
166 Impressions