llama.cpp builds b1886 through b7445 contain a race condition use-after-free vulnerability in the LLaMA-Android JNI wrapper where bench_1model() and free_1context() lack synchronization, allowing Thread A to operate on freed memory while Thread B concurrently frees the llama_context. Attackers can exploit this by performing heap spray with attacker-controlled data containing a fake vtable to hijack the vtable pointer at offset +0x30, causing llama_batch_allocr::clear() to dereference arbitrary memory and achieve remote code execution.
https://github.com/ggml-org/llama.cpp/releases/tag/b7446
https://github.com/ggml-org/llama.cpp/commit/5c0d18881e0e9794c96b2602736b758bac9d9388
https://github.com/Vladimir-tokarev-cyera/llama-cpp-security-patches
Published: 2026-08-06
Updated: 2026-08-07
Base Score: 6.2
Vector: CVSS2#AV:L/AC:H/Au:N/C:C/I:C/A:C
Severity: Medium
Base Score: 7
Vector: CVSS:3.1/AV:L/AC:H/PR:N/UI:R/S:U/C:H/I:H/A:H
Severity: High
Base Score: 7.3
Vector: CVSS:4.0/AV:L/AC:H/AT:N/PR:N/UI:P/VC:H/VI:H/VA:H/SC:N/SI:N/SA:N
Severity: High
EPSS: 0.00158