Releases · ggml-org/llama.cpp

06 Dec 15:08

f334b79

b7306 Latest

Latest

Warning

Release Format Update: Linux releases will soon use .tar.gz archives instead of .zip. Please make the necessary changes to your deployment scripts.

HIP: fix RDNA3 FP16/BF16 matrix multiplication (#17817)

macOS/iOS:

Linux:

Windows:

Assets 22

cudart-llama-bin-win-cuda-12.4-x64.zip

sha256:8c79a9b226de4b3cacfd1f83d24f962d0773be79f1e7b75c6af4ded7e32ae1d6

373 MB 2025-12-06T15:08:08Z
llama-b7306-bin-macos-arm64.tar.gz

sha256:e0a8ad44a312c77bd96310182909ab30b0f1ac0126c523efe3598b437adb51bc

13.2 MB 2025-12-06T15:08:20Z
llama-b7306-bin-macos-arm64.zip

sha256:64291139e6b4a138a057bed73db4be1d51ad5c0cf7fcd385d0536101b4d7dcb7

13.2 MB 2025-12-06T15:08:21Z
llama-b7306-bin-macos-x64.tar.gz

sha256:a8d0aa2457587acc999ec3bb7ceb7ea4a969b844e4988dd0c628b9235167b27a

36.2 MB 2025-12-06T15:08:22Z
llama-b7306-bin-macos-x64.zip

sha256:07fbddd6908cef1cbe2443734b5a3dbf02c48905eb82fa3621953d42669c1d22

36.1 MB 2025-12-06T15:08:24Z
llama-b7306-bin-ubuntu-s390x.tar.gz

sha256:2c61944ba0b1a6959efee7875628e997841a778c5203a7687d86f687c66a16ec

17.4 MB 2025-12-06T15:08:26Z
llama-b7306-bin-ubuntu-s390x.zip

sha256:adb0f31f8b004a02b161a6612ce6e89d8f84a044112851d3ee69569b5a612dd3

15.1 MB 2025-12-06T15:08:27Z
llama-b7306-bin-ubuntu-vulkan-x64.tar.gz

sha256:aefdd2db3798595f6dca3f1232b21280c330d60d8d533733dcf4b55d0d5635a2

30 MB 2025-12-06T15:08:28Z
llama-b7306-bin-ubuntu-vulkan-x64.zip

sha256:4b9259788e093c5189b88c007c4c52949b6d366d34ee0896532c714d3745a6bf

30 MB 2025-12-06T15:08:30Z
llama-b7306-bin-ubuntu-x64.tar.gz

sha256:588d31cfd9eb3bb8d349517324228c6b4e70f43d051b9c326f498e36b975939b

15.3 MB 2025-12-06T15:08:32Z
Source code (zip)

2025-12-06T12:45:36Z
Source code (tar.gz)

2025-12-06T12:45:36Z

06 Dec 13:50

github-actions

b7302

7b43f55

b7302

Warning

Release Format Update: Linux releases will soon use .tar.gz archives instead of .zip. Please make the necessary changes to your deployment scripts.

ggml : improve error handling for search path existence checks (#17653)

Improve error handling for search path existence checks

Refactor existence checks for search paths using std::error_code to handle potential errors.

Improve cache file existence check with error code

Update fs::exists to use std::error_code for error handling.

Simplify existence check for search paths

Simplify existence check for search paths

Fix logging path in error message for posix_stat
Update ggml/src/ggml-backend-reg.cpp

Co-authored-by: Aman Gupta [email protected]

Adapt to the coding standard

Co-authored-by: Aman Gupta [email protected]

macOS/iOS:

Linux:

Windows:

Assets 22

06 Dec 13:16

github-actions

b7301

444f00b

b7301

Warning

Release Format Update: Linux releases will soon use .tar.gz archives instead of .zip. Please make the necessary changes to your deployment scripts.

llama : remove quantization sanity check (#17788)

llama : remove quantization sanity check

This commit removes the quantization sanity check for attention layers.

The motivation for this is that there are model that are hybrid models
that have recurrent layers, experts layers, and attention layers. For
these models the current check fails as the experts layers are not
taking into account. After consideration, it was decided that this check
is not strictly necessary, and can be removed to allow for more flexible
model architectures.

llama : remove unused pruned_attention_w and is_clip_model vars

macOS/iOS:

Linux:

Windows:

Assets 22

06 Dec 11:16

github-actions

b7300

2960eb2

b7300

Warning

Release Format Update: Linux releases will soon use .tar.gz archives instead of .zip. Please make the necessary changes to your deployment scripts.

vulkan: Use one row per workgroup for f32 mmv (#17711)

The MoE models have a mul_mat_vec with very small m (32, 64, 128) right before
the topk_moe selection. Running multiple rows per wg doesn't utilize the SMs
well. I think even for larger m, f32 is so bandwidth-limited that running
multiple rows doesn't help.

macOS/iOS:

Linux:

Windows:

Assets 22

06 Dec 11:08

github-actions

b7298

c6c5e85

b7298

Warning

Release Format Update: Linux releases will soon use .tar.gz archives instead of .zip. Please make the necessary changes to your deployment scripts.

vulkan: support solve_tri with larger N/K values (#17781)

Split N into chunks to fit into shared memory.
If K > 128, use a larger workgroup with enough invocations.
Add perf tests matching qwen3next.

macOS/iOS:

Linux:

Windows:

Assets 22

06 Dec 10:24

github-actions

b7296

8ce774a

b7296

Warning

Release Format Update: Linux releases will soon use .tar.gz archives instead of .zip. Please make the necessary changes to your deployment scripts.

metal : fix build(#17799)

metal : fix build
tests : fix context destruction

macOS/iOS:

Linux:

Windows:

Assets 22

05 Dec 16:00

github-actions

b7285

6016d0b

b7285

Warning

Release Format Update: Linux releases will soon use .tar.gz archives instead of .zip. Please make the necessary changes to your deployment scripts.

HIP : fix RDNA4 build (#17792)

macOS/iOS:

Linux:

Windows:

Assets 22

05 Dec 04:27

github-actions

b7278

03d9a77

b7278

Warning

Release Format Update: Linux releases will soon use .tar.gz archives instead of .zip. Please make the necessary changes to your deployment scripts.

ci : transform release binary root dir in tar to llama-bXXXX (#17773)

transform release binary root dir in tar to llama-bXXXX
bsdtar supports -s instead of --transform

macOS/iOS:

Linux:

Windows:

Assets 22

05 Dec 01:21

github-actions

b7276

96fe9ba

b7276

Warning

Release Format Update: Linux releases will soon use .tar.gz archives instead of .zip. Please make the necessary changes to your deployment scripts.

Add support for CUMSUM and TRI for CUDA. (#17584)

Add support for CUMSUM and TRI for CUDA.
Minor optimizations.
Correct warp_prefix_inclusive_sum in float2 variant to return float2
Optimize TRI
Whitespace
Fix strides.
Implement double loop
Whitespace
Fix HIP compilation bugs
Optimizations + big case performance tests
Implement using CUB with fallback to custom kernel
Remove error message.
Fixes from code review
Comment out CPU-unsupported F16/BF16 cases to fix CI
Fine, you win :P
Fix last cast, use NO_DEVICE_CODE and GGML_UNUSED_VARS
Vary warp-size based on physical warp size
Add GGML_UNUSED_VARS in tri as well
Use constexpr and call prefix_inclusive with warp_size template param
Update ggml/src/ggml-cuda/cumsum.cu

Co-authored-by: Johannes Gäßler [email protected]

Apply suggestions from code review

Co-authored-by: Johannes Gäßler [email protected]

Change to tid % warp_size
Fix strides; hardcode mask; add ggml_lane_mask_t
Missing renames, remove unused get_warp_mask(), explicit calls to ggml_cuda_info()
Too hasty...

Co-authored-by: Johannes Gäßler [email protected]

macOS/iOS:

Linux:

Windows:

Assets 22

04 Dec 23:04

github-actions

b7275

bde188d

b7275

Warning

Release Format Update: Linux releases will soon use .tar.gz archives instead of .zip. Please make the necessary changes to your deployment scripts.

metal: TRI, FILL, EXPM1, SOFTPLUS (#16623)

feat(wip): Port initial TRI impl from pervious work

The kernel does not work and is not optimized, but the
code compiles and runs, so this will be the starting point
now that the core op has been merged.

Branch: ggml-cumsum-tri

Signed-off-by: Gabe Goodhart [email protected]

fix: Remove argument for constant val override

This was added in the original draft, but later removed. With this, the
kernel now passes tests.

Branch: ggml-cumsum-tri

Signed-off-by: Gabe Goodhart [email protected]

feat: Move the ttype conditional to templating to avoid conditional in kernel

Branch: ggml-cumsum-tri

Signed-off-by: Gabe Goodhart [email protected]

fix: Type fixes

Signed-off-by: Gabe Goodhart [email protected]
Co-authored-by: Georgi Gerganov [email protected]

Co-authored-by: Georgi Gerganov [email protected]

feat: Add softplus for metal

Branch: ggml-cumsum-tri

Signed-off-by: Gabe Goodhart [email protected]

feat: Add EXPM1 for metal

Branch: ggml-cumsum-tri

Signed-off-by: Gabe Goodhart [email protected]

feat: Add FILL for metal

Branch: ggml-cumsum-tri

Signed-off-by: Gabe Goodhart [email protected]

refactor: Branchless version of tri using _ggml_vec_tri_cmp as a mask

Branch: ggml-cumsum-tri

Signed-off-by: Gabe Goodhart [email protected]

fix: Remove unused arguments

Branch: ggml-cumsum-tri

Signed-off-by: Gabe Goodhart [email protected]

refactor: Use select instead of branch for softplus non-vec

Branch: ggml-cumsum-tri

Signed-off-by: Gabe Goodhart [email protected]

Signed-off-by: Gabe Goodhart [email protected]
Co-authored-by: Georgi Gerganov [email protected]

macOS/iOS:

Linux:

Windows:

Assets 22

Releases: ggml-org/llama.cpp

b7306

Uh oh!

b7302

Uh oh!

b7301

Uh oh!

b7300

Uh oh!

b7298

Uh oh!

b7296

Uh oh!

b7285

Uh oh!

b7278

Uh oh!

b7276

Uh oh!

b7275

Uh oh!