Generic scalar DTRMM was the remaining Level-3 gap (~15 GFLOPS vs ~23 DGEMM). Keep SIMD for double only; single-precision TRMM already auto-vectorized and a shared S+D kernel slowed SGEMM. Signed-off-by: Julien Jerphanion <git@jjerphan.xyz>