Commit Graph
10761 Commits
Author SHA1 Message Date
Martin Kroeker b08df585f6 Merge pull request #5972 from Orcina-Ltd/asum-alignment-determinism
Title: kernel/x86_64: make AVX-512 asum/sum kernels independent of buffer alignment
2026-08-16 16:42:53 +02:00
Martin Kroeker 377753094d Merge pull request #5978 from martin-frbg/issue5976
Fix CMake cross-compilation to ARMV9SME (or DYNAMIC_ARCH containing same)
2026-08-15 20:13:54 +02:00
Martin Kroeker 7779b52f99 Merge pull request #5942 from HecaiYuan/develop
loongarch64: fix segfaults in copy kernels and adjust dsyrk block size
2026-08-15 18:37:08 +02:00
Martin Kroeker d7d350317b Merge pull request #5977 from martin-frbg/issue5975
Expressly restore the ARM64 generic OMATCOPY CT/RT kernels to plain C
2026-08-15 17:49:28 +02:00
Martin Kroeker 98425fe1bc Fix misspelling of ARMV9SME target 2026-08-15 14:14:57 +02:00
Martin Kroeker e7083596c9 Expressly restore OMATCOPY CT/RT kernels to plain C 2026-08-15 14:09:44 +02:00
Martin Kroeker e0cabe9b59 Merge pull request #5971 from martin-frbg/issue5841
[WIP] Add ARMv9.2 SME GEMM kernels ported from vlovero's project
2026-08-15 09:14:17 +02:00
Martin Kroeker b83ffe61f3 Increase timeout for OSX DYNAMIC_ARCH job 2026-08-14 23:33:52 +02:00
Martin Kroeker 2d4212947e Increase KC 2026-08-14 22:10:20 +02:00
Martin Kroeker 7313f3ff6e Increase KC 2026-08-14 22:09:04 +02:00
Martin Kroeker 17d99746fe Merge branch 'OpenMathLib:develop' into issue5841 2026-08-14 15:00:23 +02:00
Martin Kroeker 40fea772be Merge pull request #5974 from martin-frbg/ci-macos15
Azure CI: Move mac jobs from deprecated macOS-14 image to macOS-15
2026-08-14 14:59:57 +02:00
Martin Kroeker a79dce7976 Keep the ios-armv7 job at xcode16.2/sdk 18.2 as 16.4 appears to have dropped armv7 2026-08-14 10:22:22 +02:00
Martin Kroeker d43c87d317 Update macOS SDK versions as well 2026-08-14 01:05:25 +02:00
Martin Kroeker 0e79f73488 Move mac jobs from deprecated macOS-14 image to macOS-15 2026-08-14 00:12:42 +02:00
Martin Kroeker 29de61484d Merge pull request #5973 from moluopro/fix-loongarch64-dsdot-accumulator
LoongArch: Fix DSDOT accumulator initialization
2026-08-13 23:13:29 +02:00
Martin Kroeker 9884c480ea Add casts to pacify homebrew-llvm 2026-08-13 22:20:23 +02:00
Martin Kroeker f285f47ab1 Merge branch 'develop' into issue5841 2026-08-13 20:04:43 +02:00
Martin Kroeker 305bd67178 Add +sme-f64f64 to build flags of VortexM4 and ARMV9SME 2026-08-13 18:51:35 +02:00
Martin Kroeker f8830b66e3 Add sme-f64f64 capability to VortexM4 and ARMV9SME build flags 2026-08-13 18:48:02 +02:00
Martin Kroeker 81dd859785 Move declarations of the ARM64 SME kernels to the appropriate headers 2026-08-13 18:43:48 +02:00
Martin Kroeker 3ead57fd2b Improve clobber lists and interfaces 2026-08-13 18:41:33 +02:00
Martin Kroeker c073f087b4 Add ARM64 SME GEMM kernels 2026-08-13 18:39:56 +02:00
Martin Kroeker 9724481b59 Add declarations for ARM64 SME GEMM kernels 2026-08-13 18:38:55 +02:00
moluopro fc6f4a3b46 CI: Re-enable LoongArch DSDOT test with Clang 2026-08-13 23:01:10 +08:00
moluopro 8f2a8fe318 CI: Re-enable LoongArch DSDOT test with GCC 2026-08-13 23:00:47 +08:00
moluopro 404f288a9d LoongArch: Fix DSDOT accumulator initialization 2026-08-13 23:00:41 +08:00
David Heffernan d793be85b4 kernel/x86_64: make AVX-512 asum/sum kernels independent of buffer alignment
The skylakex/cooperlake d/s/c/z asum and c/z sum microkernels peel leading
elements until the input pointer reaches a 64-byte boundary (a scalar loop in
dasum/sasum, a masked header load in the complex variants) before entering an
aligned-load accumulator loop. The peel count depends on the buffer address
mod 64, so the grouping of the sum into accumulators - and therefore the
rounding of the result - depends on where the caller's buffer happens to sit
in memory. The same data at a different address can give a bitwise-different
sum.

That address dependence surfaced as non-reproducibility in OrcaFlex: LAPACK's
dstein scales each inverse-iteration eigenvector by 1/dasum(...) over a heap
array whose alignment varies with allocation history, so eigenvectors from
identical inputs differed run to run in the last bits, which zero-tolerance
regression comparison flags.

Fix by dropping the alignment peel and using unaligned loads throughout, so
the summation order is a function of the length alone. On AVX-512 hardware
unaligned load instructions on addresses that happen to be aligned cost the
same as aligned loads; only genuinely split cache lines pay a small penalty,
negligible for these level-1 reductions.
2026-08-13 13:57:49 +01:00
Martin Kroeker 8d73a856fb Merge pull request #5970 from hugomeiland/cortexa72-dgemm-6x8
ARM64: Cortex-A72 DGEMM 6×8 microkernel and blocking
2026-08-12 18:12:23 +02:00
Martin Kroeker 4ae369ac0a Make the compute kernel static 2026-08-12 10:54:27 +02:00
Martin Kroeker 1ed99815fa Add SME GEMM kernels 2026-08-12 10:53:07 +02:00
Hugo MeilandandCursor 4e923d1ba6 Keep CORTEXA72 out of default DYNAMIC_ARCH
Per review: A72 was dropped from the default DYNAMIC_CORE list in
#4389 to limit arm64 binary size. Restore the A57 alias for default
DYNAMIC_ARCH; TARGET=CORTEXA72 and DYNAMIC_LIST=CORTEXA72 still get
the dedicated 6x8 kernels.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-11 18:52:45 +02:00
Martin Kroeker e24e0da779 Add sme-f64f64 to ARMV9SME archflags too 2026-08-11 17:58:17 +02:00
Martin Kroeker 39b56e9a05 Add ARM64 SME GEMM kernels 2026-08-11 12:16:04 +02:00
Martin Kroeker cf1b7c1ad9 Clean up non-OpenMP build and add CBLAS GEMM benchmark 2026-08-11 12:14:58 +02:00
Martin Kroeker 78f03216de Add SME GEMM kernels ported from vlovero's ARMv9.2-GEMM project 2026-08-11 12:12:29 +02:00
Martin Kroeker f2dc74796f Integrate SME GEMM kernels 2026-08-11 12:10:47 +02:00
Martin Kroeker 2f915c5e23 Add f64f64 extension to VortexM4 options 2026-08-11 12:10:03 +02:00
Martin Kroeker 5aa2c41e8d Credit Vincent Lovero for his ARM SME kernel work 2026-08-11 12:07:46 +02:00
Hugo MeilandandCursor df6032c375 Alias gotoblas_CORTEXA72 to ARMV8 on Darwin DYNAMIC_ARCH
Darwin's DYNAMIC_CORE only builds ARMV8/NEOVERSEN1/ARMV9SME/VORTEXM4,
so an unconditional extern gotoblas_CORTEXA72 left Apple M builds with
an undefined symbol. Mirror the CORTEXA57 Darwin alias.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-09 12:47:18 +02:00
Hugo MeilandandCursor 2b69faff24 Add generic neg_tcopy_6 for DGEMM_UNROLL_M=6
DYNAMIC_ARCH builds CORTEXA72 as a separate kernel and pull
dneg_tcopy from generic/neg_tcopy_$(DGEMM_UNROLL_M).c. Width 6 was
missing (only 1/2/4/8/16 existed), which broke the arm64 Graviton
Cirun and Azure DYNAMIC_ARM64 jobs.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-09 12:16:01 +02:00
Hugo MeilandandCursor 23f832c464 Wire CORTEXA72 6x8 blocking and DYNAMIC_ARCH dispatch
Give CORTEXA72 its own param.h block (UNROLL 6x8, P=120 Q=240, R=4096
shared-L2 / R=768 single-core), add it to DYNAMIC_CORE, and stop
aliasing gotoblas_CORTEXA72 to A57 so DYNAMIC_ARCH can select the new
kernels on MIDR 0xd08.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-09 10:31:35 +02:00
Hugo MeilandandCursor 560caa56e6 Add Cortex-A72 DGEMM 6x8 microkernel with MR=6 packers and TRSM
TARGET=CORTEXA72 previously reused the A57 8x4 path. Add a dedicated
6x8 NEON ukernel, contiguous MR=6 panel packers (stock gemm_*copy_6 is
4+2), and UNROLL_M=6-aware TRSM kernels so HPL/dtrsm does not corrupt
the heap. DTRMM falls back to generic 2x2 until a matching kernel exists.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-09 10:31:35 +02:00
Martin Kroeker 05518e0f11 Merge pull request #5969 from martin-frbg/lapack1346
Raise SE2,SEP test thresholds to account for xSTEINs orthogonality guarantee (Reference-LAPACK PR 1346)
2026-08-09 00:28:18 +02:00
Martin Kroeker 7fdb527565 Merge pull request #5968 from martin-frbg/lapack1342
Fix invalid reads in the single&double precision DMD tests (Reference-LAPACK PR 1342)
2026-08-09 00:27:58 +02:00
Martin Kroeker b8d6a7bef8 Merge pull request #5967 from martin-frbg/lapack1338
Fix spurious SGEBAL/DGEBAL test failure caused by wrong metric (Reference-LAPACK PR 1338)
2026-08-08 23:05:16 +02:00
Martin Kroeker 9ff46235ea Merge pull request #5966 from Ka-zam/relapack-sytrf-workspace
Fix workspace size in ReLAPACK sytrf/hetrf
2026-08-08 23:04:44 +02:00
Martin Kroeker 330063dec8 Raise threshold to account for orthogonality guarantee (Reference-LAPACK PR 1346) 2026-08-08 18:25:41 +02:00
Martin Kroeker 2f24d51c5d Raise threshold to accound for orthogonality guarantee (Reference-LAPACK PR 1346) 2026-08-08 18:24:25 +02:00
Martin Kroeker b22e3244e4 Fix read beyond the array bounds (Reference-LAPACK PR 1342) 2026-08-08 18:17:04 +02:00