Benchmarking nested do loops, MATMUL and Blas DGEMM

222 · Fortran Lang Discourse · Aug. 8, 2026, 8:45 a.m.
Summary
This blog post benchmarks matrix multiplication using Fortran, comparing cache-friendly nested loops, the built-in MATMUL function, and OpenBLAS DGEMM with varying thread numbers. The results show that two threads for DGEMM yield the best performance, while MATMUL remains faster than the explicit loops. The post highlights the importance of thread management and provides code for performing these benchmarks.