Try our Chrome extension
Easily add the current web-page from your browser directly into your changedetection.io tool, more great features coming soon!Changedetection.io needs your support!
You can help us by supporting changedetection.io on these platforms;
- Rate us at AlternativeTo.net
- Star us on GitHub
- Follow us at Twitter/X
- G2 Software reviews
- Check us out on LinkedIn
- And tell your friends and colleagues :)
The more popular changedetection.io is, the more time we can dedicate to adding amazing features!
Many thanks :)
changedetection.io team
Aún no hace unos segundos.
False
Aún no hace unos segundos
hace 4 horasIr a instantánea única
Skip to content
Navigation Menu
Sign in Appearance settings
* Platform
+ AI CODE CREATION
o GitHub Copilot Write better code with AI
o GitHub Copilot app Direct agents from issue to merge
o MCP Registry Integrate external tools
+ DEVELOPER WORKFLOWS
o Actions Automate any workflow
o Codespaces Instant dev environments
o Issues Plan and track work
o Code Review Manage code changes
o Code Quality Enforce quality at merge
+ APPLICATION SECURITY
o GitHub Advanced Security Find and fix vulnerabilities
o Code security Secure your code as you build
o Secret protection Stop leaks before they start
+ EXPLORE
o Why GitHub
o Documentation
o Blog
o Changelog
o Marketplace
View all features
* Solutions
+ BY COMPANY SIZE
o Enterprises
o Small and medium teams
o Startups
o Nonprofits
+ BY USE CASE
o App Modernization
o DevSecOps
o DevOps
o CI/CD
o View all use cases
+ BY INDUSTRY
o Healthcare
o Financial services
o Manufacturing
o Government
o View all industries
View all solutions
* Resources
+ EXPLORE BY TOPIC
o AI
o Software Development
o DevOps
o Security
o View all topics
+ EXPLORE BY TYPE
o Customer stories
o Events & webinars
o Ebooks & reports
o Business insights
o GitHub Skills
+ SUPPORT & SERVICES
o Documentation
o Customer support
o Community forum
o Trust center
o Partners
View all resources
* Open Source
+ COMMUNITY
o GitHub Sponsors Fund open source developers
+ PROGRAMS
o Security Lab
o Maintainer Community
o Accelerator
o GitHub Stars
o Archive Program
+ REPOSITORIES
o Topics
o Trending
o Collections
* Enterprise
+ ENTERPRISE SOLUTIONS
o Enterprise platform AI-powered developer platform
+ AVAILABLE ADD-ONS
o GitHub Advanced Security Enterprise-grade security features
o Copilot for Business Enterprise-grade AI features
o Premium Support Enterprise-grade 24/7 support
* Pricing
Type / to search
Sign in
Sign up Appearance settings
You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert
Uh oh!
There was an error while loading. Please reload this page.
ggml-org / llama.cpp Public
* Notifications You must be signed in to change notification settings
* Fork 21.5k
* Star 123k
* Code
* Issues 687
* Issues 688
* Pull requests 1.3k
* Discussions
* Actions
* Projects
* Wiki
* Security and quality 13
* Insights
Additional navigation options
* Code
* Issues
* Pull requests
* Discussions
* Actions
* Projects
* Wiki
* Security and quality
* Insights
Releases: ggml-org/llama.cpp
Releases Tags
Releases · ggml-org/llama.cpp
Release list
* b10333
* b10332
* b10331
* b10330
* b10329
* b10328
* b10327
* b10326
* b10322
* b10321
Previous Next
Jump to release
* b10333
* b10332
* b10331
* b10330
* b10329
* b10328
* b10327
* b10326
* b10322
* b10321
Previous Next
b10333
b10333 Latest
Latest
Compare
Choose a tag to compare
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
No results found
View all tags
github-actions released this 09 Aug 11:21
b10333
0865990
This commit was created on GitHub.com and signed with GitHub’s verified signature.
GPG key ID: B5690EEEBB952194
Verified
Learn about vigilant mode.
ggml-cpu : fix missing Q5_0 dispatch in SpaceMiT backend (#26792)
Website:
* https://llama.app
macOS/iOS:
* macOS Apple Silicon (arm64)
* macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
* macOS Intel (x64)
* iOS XCFramework
Linux:
* Ubuntu x64 (CPU)
* Ubuntu arm64 (CPU)
* Ubuntu s390x (CPU)
* Ubuntu x64 (Vulkan)
* Ubuntu arm64 (Vulkan)
* Ubuntu x64 (ROCm 7.2)
* Ubuntu x64 (OpenVINO)
* Ubuntu x64 (SYCL FP32)
* Ubuntu x64 (SYCL FP16)
Android:
* Android arm64 (CPU)
Windows:
* Windows x64 (CPU)
* Windows arm64 (CPU)
* Windows arm64 (OpenCL Adreno)
* Windows x64 (CUDA 12) - CUDA 12.4 DLLs
* Windows x64 (CUDA 13) - CUDA 13.3 DLLs
* Windows x64 (Vulkan)
* Windows x64 (OpenVINO)
* Windows x64 (SYCL)
* Windows x64 (HIP)
openEuler:
* DISABLED
* openEuler x86 (310p)
* openEuler x86 (910b, ACL Graph)
* openEuler aarch64 (310p)
* openEuler aarch64 (910b, ACL Graph)
UI:
* UI
Assets 27
* cudart-llama-bin-win-cuda-12.4-x64.zip
sha256:8c79a9b226de4b3cacfd1f83d24f962d0773be79f1e7b75c6af4ded7e32ae1d6
373 MB 2026-08-09T11:21:51Z
* cudart-llama-bin-win-cuda-13.3-x64.zip
sha256:1462a050eb4c684921ba51dcc4cc488a036674c3e73e9945ee705b854808d03e
373 MB 2026-08-09T11:22:05Z
* llama-b10333-bin-android-arm64.tar.gz
sha256:bcb40c8c432e6947bd9bf1094a070edd40c5546ee4c2d85d4b1d289553442f93
73.4 MB 2026-08-09T11:22:16Z
* llama-b10333-bin-macos-arm64.tar.gz
sha256:e5d67c5264107e3c14d3bf2aee349365bb2b85ae99bb077a0cb974a1c4c2741a
10.5 MB 2026-08-09T11:22:19Z
* llama-b10333-bin-macos-x64.tar.gz
sha256:6ffd9e0b9b2e3ab6ccfba332a74b968b0fef891f01bd4747d3d75bc7393877ea
10.8 MB 2026-08-09T11:22:20Z
* llama-b10333-bin-ubuntu-arm64.tar.gz
sha256:95da1a0f7538340f625b0301593ee63c046ec0a155c74f21444ea4d43bca79a1
12.8 MB 2026-08-09T11:22:21Z
* llama-b10333-bin-ubuntu-openvino-2026.2.1-x64.tar.gz
sha256:6601b30a9f1fa69f1ccbefff9e9a4ad4bcb82853fa489a5d62cf141091b8be21
97.2 MB 2026-08-09T11:22:22Z
* llama-b10333-bin-ubuntu-rocm-7.2-x64.tar.gz
sha256:bae09231fe6f965764c227b6969ea2edbf2505814a27522a1b7051d32bb1903c
124 MB 2026-08-09T11:22:25Z
* llama-b10333-bin-ubuntu-s390x.tar.gz
sha256:2fba779ca4b60c8dbdc2bb06af3dbe8dfbb944826ab2b3faf53bb42ba184154a
14.8 MB 2026-08-09T11:22:29Z
* llama-b10333-bin-ubuntu-sycl-fp16-x64.tar.gz
sha256:89f64bdeced31a2f80e9909dee4da11c26ed39f59a3664f4a3c4221dbe082124
50.7 MB 2026-08-09T11:22:31Z
* Source code (zip)
2026-08-09T10:16:53Z
* Source code (tar.gz)
2026-08-09T10:16:53Z
* Show all 27 assets Loading
Uh oh!
There was an error while loading. Please reload this page.
👍 1 carlosjln reacted with thumbs up emoji 🚀 6 nalinduash, luoluo74751-ux, Chroma01, little-sparkleeesss, FabioLeitao, and carlosjln reacted with rocket emoji
All reactions
* 👍 1 reaction
* 🚀 6 reactions
6 people reacted
b10332
b10332
Compare
Choose a tag to compare
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
No results found
View all tags
github-actions released this 09 Aug 10:48
b10332
61141f1
This commit was created on GitHub.com and signed with GitHub’s verified signature.
GPG key ID: B5690EEEBB952194
Verified
Learn about vigilant mode.
ci: rm GGML_HIP_ROCWMMA_FATTN (#26760)
Signed-off-by: Aaron Teo aaron.teo1@ibm.com
Website:
* https://llama.app
macOS/iOS:
* macOS Apple Silicon (arm64)
* macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
* macOS Intel (x64)
* iOS XCFramework
Linux:
* Ubuntu x64 (CPU)
* Ubuntu arm64 (CPU)
* Ubuntu s390x (CPU)
* Ubuntu x64 (Vulkan)
* Ubuntu arm64 (Vulkan)
* Ubuntu x64 (ROCm 7.2)
* Ubuntu x64 (OpenVINO)
* Ubuntu x64 (SYCL FP32)
* Ubuntu x64 (SYCL FP16)
Android:
* Android arm64 (CPU)
Windows:
* Windows x64 (CPU)
* Windows arm64 (CPU)
* Windows arm64 (OpenCL Adreno)
* Windows x64 (CUDA 12) - CUDA 12.4 DLLs
* Windows x64 (CUDA 13) - CUDA 13.3 DLLs
* Windows x64 (Vulkan)
* Windows x64 (OpenVINO)
* Windows x64 (SYCL)
* Windows x64 (HIP)
openEuler:
* DISABLED
* openEuler x86 (310p)
* openEuler x86 (910b, ACL Graph)
* openEuler aarch64 (310p)
* openEuler aarch64 (910b, ACL Graph)
UI:
* UI
Assets 27
Loading
Uh oh!
There was an error while loading. Please reload this page.
All reactions
b10331
b10331
Compare
Choose a tag to compare
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
No results found
View all tags
github-actions released this 08 Aug 23:26
b10331
7ba604f
This commit was created on GitHub.com and signed with GitHub’s verified signature.
GPG key ID: B5690EEEBB952194
Verified
Learn about vigilant mode.
server: report the isolate working directory from get_info (#26773)
* server: report the isolate working directory from get_info
Without an explicit cwd, get_info fell back to the server process
working directory even when a tools runtime was configured. That named a
host path no tool would ever run in, since an isolate starts in a
directory of its own.
It now asks the isolate for its working directory in that case, and
keeps the process one only when the tools run on the host.
* remove redundant comment
Co-authored-by: Xuan-Son Nguyen thichthat@gmail.com
Website:
* https://llama.app
macOS/iOS:
* macOS Apple Silicon (arm64)
* macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
* macOS Intel (x64)
* iOS XCFramework
Linux:
* Ubuntu x64 (CPU)
* Ubuntu arm64 (CPU)
* Ubuntu s390x (CPU)
* Ubuntu x64 (Vulkan)
* Ubuntu arm64 (Vulkan)
* Ubuntu x64 (ROCm 7.2)
* Ubuntu x64 (OpenVINO)
* Ubuntu x64 (SYCL FP32)
* Ubuntu x64 (SYCL FP16)
Android:
* Android arm64 (CPU)
Windows:
* Windows x64 (CPU)
* Windows arm64 (CPU)
* Windows arm64 (OpenCL Adreno)
* Windows x64 (CUDA 12) - CUDA 12.4 DLLs
* Windows x64 (CUDA 13) - CUDA 13.3 DLLs
* Windows x64 (Vulkan)
* Windows x64 (OpenVINO)
* Windows x64 (SYCL)
* Windows x64 (HIP)
openEuler:
* DISABLED
* openEuler x86 (310p)
* openEuler x86 (910b, ACL Graph)
* openEuler aarch64 (310p)
* openEuler aarch64 (910b, ACL Graph)
UI:
* UI
Assets 27
Loading
Uh oh!
There was an error while loading. Please reload this page.
❤️ 2 nalinduash and sjpoo reacted with heart emoji
All reactions
* ❤️ 2 reactions
2 people reacted
b10330
b10330
Compare
Choose a tag to compare
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
No results found
View all tags
github-actions released this 08 Aug 17:20
b10330
687e778
This commit was created on GitHub.com and signed with GitHub’s verified signature.
GPG key ID: B5690EEEBB952194
Verified
Learn about vigilant mode.
CUDA: fuse rms_norm + mul + rope (+ view + set_rows) (#26767)
* CUDA: fuse rms_norm + mul + rope (+ view + set_rows)
* tests: add broadcast weight case to rms_norm_mul_rope
* CUDA: check memory ranges before rms_norm rope fusion
* CUDA: check memory ranges in rope set_rows fusion
Website:
* https://llama.app
macOS/iOS:
* macOS Apple Silicon (arm64)
* macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
* macOS Intel (x64)
* iOS XCFramework
Linux:
* Ubuntu x64 (CPU)
* Ubuntu arm64 (CPU)
* Ubuntu s390x (CPU)
* Ubuntu x64 (Vulkan)
* Ubuntu arm64 (Vulkan)
* Ubuntu x64 (ROCm 7.2)
* Ubuntu x64 (OpenVINO)
* Ubuntu x64 (SYCL FP32)
* Ubuntu x64 (SYCL FP16)
Android:
* Android arm64 (CPU)
Windows:
* Windows x64 (CPU)
* Windows arm64 (CPU)
* Windows arm64 (OpenCL Adreno)
* Windows x64 (CUDA 12) - CUDA 12.4 DLLs
* Windows x64 (CUDA 13) - CUDA 13.3 DLLs
* Windows x64 (Vulkan)
* Windows x64 (OpenVINO)
* Windows x64 (SYCL)
* Windows x64 (HIP)
openEuler:
* DISABLED
* openEuler x86 (310p)
* openEuler x86 (910b, ACL Graph)
* openEuler aarch64 (310p)
* openEuler aarch64 (910b, ACL Graph)
UI:
* UI
Assets 27
Loading
Uh oh!
There was an error while loading. Please reload this page.
❤️ 1 sc2proton2012 reacted with heart emoji
All reactions
* ❤️ 1 reaction
1 person reacted
b10329
b10329
Compare
Choose a tag to compare
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
No results found
View all tags
github-actions released this 08 Aug 16:00
b10329
18f7ad7
This commit was created on GitHub.com and signed with GitHub’s verified signature.
GPG key ID: B5690EEEBB952194
Verified
Learn about vigilant mode.
server, ui: only offer a working directory when a tool reads it (#26762)
The working directory chip showed up as soon as the server exposed any
builtin tool, so a server started with just get_datetime, or a user who
turned every filesystem tool off in the settings, still got a control
that nothing would read.
Tools now declare whether they resolve their paths and run against the
working directory, next to the write permission they already publish in
the /tools listing. The WebUI shows the chip and enables the /cwd
command only when at least one such tool is both served and left
enabled.
Website:
* https://llama.app
macOS/iOS:
* macOS Apple Silicon (arm64)
* macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
* macOS Intel (x64)
* iOS XCFramework
Linux:
* Ubuntu x64 (CPU)
* Ubuntu arm64 (CPU)
* Ubuntu s390x (CPU)
* Ubuntu x64 (Vulkan)
* Ubuntu arm64 (Vulkan)
* Ubuntu x64 (ROCm 7.2)
* Ubuntu x64 (OpenVINO)
* Ubuntu x64 (SYCL FP32)
* Ubuntu x64 (SYCL FP16)
Android:
* Android arm64 (CPU)
Windows:
* Windows x64 (CPU)
* Windows arm64 (CPU)
* Windows arm64 (OpenCL Adreno)
* Windows x64 (CUDA 12) - CUDA 12.4 DLLs
* Windows x64 (CUDA 13) - CUDA 13.3 DLLs
* Windows x64 (Vulkan)
* Windows x64 (OpenVINO)
* Windows x64 (SYCL)
* Windows x64 (HIP)
openEuler:
* DISABLED
* openEuler x86 (310p)
* openEuler x86 (910b, ACL Graph)
* openEuler aarch64 (310p)
* openEuler aarch64 (910b, ACL Graph)
UI:
* UI
Assets 27
Loading
Uh oh!
There was an error while loading. Please reload this page.
❤️ 1 sc2proton2012 reacted with heart emoji
All reactions
* ❤️ 1 reaction
1 person reacted
b10328
b10328
Compare
Choose a tag to compare
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
No results found
View all tags
github-actions released this 08 Aug 15:22
b10328
dd2c7c4
This commit was created on GitHub.com and signed with GitHub’s verified signature.
GPG key ID: B5690EEEBB952194
Verified
Learn about vigilant mode.
server: add initial tool isolation support (via docker) (#26507)
* server: add initial tool isolation support (via docker)
* add docs
* adapt get_info
* py: fix type check
* cont
* separate tools_io_sandbox / tools_io_docker
* rename sandbox --> isolate
* x-tool-docker --> x-tool-runtime
Co-authored-by: Pascal admin@serveurperso.com
Website:
* https://llama.app
macOS/iOS:
* macOS Apple Silicon (arm64)
* macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
* macOS Intel (x64)
* iOS XCFramework
Linux:
* Ubuntu x64 (CPU)
* Ubuntu arm64 (CPU)
* Ubuntu s390x (CPU)
* Ubuntu x64 (Vulkan)
* Ubuntu arm64 (Vulkan)
* Ubuntu x64 (ROCm 7.2)
* Ubuntu x64 (OpenVINO)
* Ubuntu x64 (SYCL FP32)
* Ubuntu x64 (SYCL FP16)
Android:
* Android arm64 (CPU)
Windows:
* Windows x64 (CPU)
* Windows arm64 (CPU)
* Windows arm64 (OpenCL Adreno)
* Windows x64 (CUDA 12) - CUDA 12.4 DLLs
* Windows x64 (CUDA 13) - CUDA 13.3 DLLs
* Windows x64 (Vulkan)
* Windows x64 (OpenVINO)
* Windows x64 (SYCL)
* Windows x64 (HIP)
openEuler:
* DISABLED
* openEuler x86 (310p)
* openEuler x86 (910b, ACL Graph)
* openEuler aarch64 (310p)
* openEuler aarch64 (910b, ACL Graph)
UI:
* UI
Assets 27
Loading
Uh oh!
There was an error while loading. Please reload this page.
👍 1 WANDGAMES-cmd reacted with thumbs up emoji ❤️ 1 sc2proton2012 reacted with heart emoji
All reactions
* 👍 1 reaction
* ❤️ 1 reaction
2 people reacted
b10327
b10327
Compare
Choose a tag to compare
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
No results found
View all tags
github-actions released this 08 Aug 06:04
b10327
69bf643
This commit was created on GitHub.com and signed with GitHub’s verified signature.
GPG key ID: B5690EEEBB952194
Verified
Learn about vigilant mode.
CUDA: fix thread/block count in quantized cpy kernel launches (#26731)
* CUDA: fix thread/block count in quantized cpy kernel launches
* tests: add uneven block count cpy case
Website:
* https://llama.app
macOS/iOS:
* macOS Apple Silicon (arm64)
* macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
* macOS Intel (x64)
* iOS XCFramework
Linux:
* Ubuntu x64 (CPU)
* Ubuntu arm64 (CPU)
* Ubuntu s390x (CPU)
* Ubuntu x64 (Vulkan)
* Ubuntu arm64 (Vulkan)
* Ubuntu x64 (ROCm 7.2)
* Ubuntu x64 (OpenVINO)
* Ubuntu x64 (SYCL FP32)
* Ubuntu x64 (SYCL FP16)
Android:
* Android arm64 (CPU)
Windows:
* Windows x64 (CPU)
* Windows arm64 (CPU)
* Windows arm64 (OpenCL Adreno)
* Windows x64 (CUDA 12) - CUDA 12.4 DLLs
* Windows x64 (CUDA 13) - CUDA 13.3 DLLs
* Windows x64 (Vulkan)
* Windows x64 (OpenVINO)
* Windows x64 (SYCL)
* Windows x64 (HIP)
openEuler:
* DISABLED
* openEuler x86 (310p)
* openEuler x86 (910b, ACL Graph)
* openEuler aarch64 (310p)
* openEuler aarch64 (910b, ACL Graph)
UI:
* UI
Assets 27
Loading
Uh oh!
There was an error while loading. Please reload this page.
👍 2 little-sparkleeesss and nalinduash reacted with thumbs up emoji
All reactions
* 👍 2 reactions
2 people reacted
b10326
b10326
Compare
Choose a tag to compare
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
No results found
View all tags
github-actions released this 07 Aug 21:23
b10326
3653e6d
This commit was created on GitHub.com and signed with GitHub’s verified signature.
GPG key ID: B5690EEEBB952194
Verified
Learn about vigilant mode.
tts: account for the vocoder pass in the timings line (#26733)
get_output runs the waveform work the pipeline defers to it, from a
single trailing window to a full pass depending on the model. Measuring
it keeps the reported total and the audio to process ratio honest.
Website:
* https://llama.app
macOS/iOS:
* macOS Apple Silicon (arm64)
* macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
* macOS Intel (x64)
* iOS XCFramework
Linux:
* Ubuntu x64 (CPU)
* Ubuntu arm64 (CPU)
* Ubuntu s390x (CPU)
* Ubuntu x64 (Vulkan)
* Ubuntu arm64 (Vulkan)
* Ubuntu x64 (ROCm 7.2)
* Ubuntu x64 (OpenVINO)
* Ubuntu x64 (SYCL FP32)
* Ubuntu x64 (SYCL FP16)
Android:
* Android arm64 (CPU)
Windows:
* Windows x64 (CPU)
* Windows arm64 (CPU)
* Windows arm64 (OpenCL Adreno)
* Windows x64 (CUDA 12) - CUDA 12.4 DLLs
* Windows x64 (CUDA 13) - CUDA 13.3 DLLs
* Windows x64 (Vulkan)
* Windows x64 (OpenVINO)
* Windows x64 (SYCL)
* Windows x64 (HIP)
openEuler:
* DISABLED
* openEuler x86 (310p)
* openEuler x86 (910b, ACL Graph)
* openEuler aarch64 (310p)
* openEuler aarch64 (910b, ACL Graph)
UI:
* UI
Assets 27
Loading
Uh oh!
There was an error while loading. Please reload this page.
👍 1 little-sparkleeesss reacted with thumbs up emoji
All reactions
* 👍 1 reaction
1 person reacted
b10322
b10322
Compare
Choose a tag to compare
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
No results found
View all tags
github-actions released this 07 Aug 19:51
b10322
f8e3026
This commit was created on GitHub.com and signed with GitHub’s verified signature.
GPG key ID: B5690EEEBB952194
Verified
Learn about vigilant mode.
sycl: coalesce the ssm_conv window loads (#26612)
test-backend-ops perf -o SSM_CONV on an Arc Pro B70, interleaved A/B against
master, 6 reps, us/run:
ne_a=[515,3328,1,1] ne_b=[4,3328,1,1] n_t=512 97.68 -> 52.95 1.85x
ne_a=[937,8192,1,1] ne_b=[4,8192,1,1] n_t=934 516.16 -> 276.13 1.87x
ne_a=[4,3328,1,1] ne_b=[4,3328,1,1] n_t=1 2.73 -> 2.71 flat
llama-bench on qwen35 27B Q4_K - Medium (48 of its 64 blocks run ssm_conv),
-ngl 99 -fa 1 -ctk f16 -ctv f16, interleaved passes of r=3:
-b 2048 -ub 2048 pp2048 1045.1 / 1043.5 / 1043.7 -> 1069.5 / 1066.3 / 1065.9 +2.2%
-b 2048 -ub 512 pp2048 771.8 / 772.7 -> 785.5 / 786.6 +1.8%
-b 2048 -ub 512 tg128 23.81 / 23.88 -> 23.87 / 23.86 flat
Website:
* https://llama.app
macOS/iOS:
* macOS Apple Silicon (arm64)
* macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
* macOS Intel (x64)
* iOS XCFramework
Linux:
* Ubuntu x64 (CPU)
* Ubuntu arm64 (CPU)
* Ubuntu s390x (CPU)
* Ubuntu x64 (Vulkan)
* Ubuntu arm64 (Vulkan)
* Ubuntu x64 (ROCm 7.2)
* Ubuntu x64 (OpenVINO)
* Ubuntu x64 (SYCL FP32)
* Ubuntu x64 (SYCL FP16)
Android:
* Android arm64 (CPU)
Windows:
* Windows x64 (CPU)
* Windows arm64 (CPU)
* Windows arm64 (OpenCL Adreno)
* Windows x64 (CUDA 12) - CUDA 12.4 DLLs
* Windows x64 (CUDA 13) - CUDA 13.3 DLLs
* Windows x64 (Vulkan)
* Windows x64 (OpenVINO)
* Windows x64 (SYCL)
* Windows x64 (HIP)
openEuler:
* DISABLED
* openEuler x86 (310p)
* openEuler x86 (910b, ACL Graph)
* openEuler aarch64 (310p)
* openEuler aarch64 (910b, ACL Graph)
UI:
* UI
Assets 27
Loading
Uh oh!
There was an error while loading. Please reload this page.
👍 2 jreng02 and PrimeDirective8 reacted with thumbs up emoji 👀 1 cantosun99 reacted with eyes emoji
All reactions
* 👍 2 reactions
* 👀 1 reaction
3 people reacted
b10321
b10321
Compare
Choose a tag to compare
Sorry, something went wrong.
Filter
Loading
Sorry, something went wrong.
Uh oh!
There was an error while loading. Please reload this page.
No results found
View all tags
github-actions released this 07 Aug 19:07
b10321
a194a75
This commit was created on GitHub.com and signed with GitHub’s verified signature.
GPG key ID: B5690EEEBB952194
Verified
Learn about vigilant mode.
metal : fix NORM/RMS_NORM for row lengths that leave a partial simdgroup (#26708)
ggml_metal_op_norm sized the threadgroup with
nth = std::min(nth, args.ne00_t), which can leave nth not a multiple of
the simdgroup size. The kernels finish their row reduction with a
cross-simdgroup step where each lane of the last simdgroup reads one
per-simdgroup partial sum out of shmem_f32:
if (tiisg == 0) { shmem_f32[sgitg] = sumf; } threadgroup_barrier(mem_flags::mem_threadgroup); sumf = shmem_f32[tiisg]; sumf = simd_sum(sumf);
When the last simdgroup is partial it has fewer lanes than the
threadgroup has simdgroups, so the tail of the partial sums is never
read and the row sum is too small. For ne00_t = 33 nth becomes 33: two
simdgroups, but only one lane in the second, so one of the two partial
sums is dropped. The mean and variance are then wrong for the whole row.
Round ne00_t up to a whole number of simdgroups instead. Rounding up
rather than dropping the clamp keeps the threadgroup as small as
possible: deleting the line would raise nth to the next power of two
(ne00_t = 544 -> 1024 instead of 544), which costs idle lanes on 26 row
lengths below 8192 that were already correct, including 1536 and 3584.
GGML_OP_NORM is affected as well as GGML_OP_RMS_NORM - both dispatch
through ggml_metal_op_norm.
No mainstream LLM hidden size hits this: ne00_t is ne00/4 on the
vectorized path, so 4096, 8192, 2048 and friends all give a multiple of
32. It is reachable from other norm shapes, e.g. 320-channel norms.
Add NORM and RMS_NORM cases for ne0 = 33, 132 and 260 across the
existing eps values. 33 exercises the scalar path and 132/260 the
vectorized one, since only those divide by 4.
Before, on M3 Pro:
test-backend-ops test -b MTL0 -o NORM 25/50 test-backend-ops test -b MTL0 -o RMS_NORM 26/51
After:
test-backend-ops test -b MTL0 -o NORM 50/50 test-backend-ops test -b MTL0 -o RMS_NORM 51/51 test-backend-ops test -b MTL0 13943/13943
Website:
* https://llama.app
macOS/iOS:
* macOS Apple Silicon (arm64)
* macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
* macOS Intel (x64)
* iOS XCFramework
Linux:
* Ubuntu x64 (CPU)
* Ubuntu arm64 (CPU)
* Ubuntu s390x (CPU)
* Ubuntu x64 (Vulkan)
* Ubuntu arm64 (Vulkan)
* Ubuntu x64 (ROCm 7.2)
* Ubuntu x64 (OpenVINO)
* Ubuntu x64 (SYCL FP32)
* Ubuntu x64 (SYCL FP16)
Android:
* Android arm64 (CPU)
Windows:
* Windows x64 (CPU)
* Windows arm64 (CPU)
* Windows arm64 (OpenCL Adreno)
* Windows x64 (CUDA 12) - CUDA 12.4 DLLs
* Windows x64 (CUDA 13) - CUDA 13.3 DLLs
* Windows x64 (Vulkan)
* Windows x64 (OpenVINO)
* Windows x64 (SYCL)
* Windows x64 (HIP)
openEuler:
* DISABLED
* openEuler x86 (310p)
* openEuler x86 (910b, ACL Graph)
* openEuler aarch64 (310p)
* openEuler aarch64 (910b, ACL Graph)
UI:
* UI
Assets 27
Loading
Uh oh!
There was an error while loading. Please reload this page.
All reactions
Previous 1 2 3 4 5 … 99 100 Next
Previous Next
Footer
© 2026 GitHub, Inc.
Footer navigation
* Terms
* Privacy
* Security
* Status
* Community
* Docs
* Contact
* Manage cookies
* Do not share my personal information
You can’t perform that action at this time.
Consejo: Resalte el texto para compartir o agregar a las listas de ignorados.
— Download difference patch
Por ahora, las diferencias se realizan en texto, no gráficamente, solo está disponible la última captura de pantalla.
La captura de pantalla requiere que Playwright/WebDriver esté habilitado