Note
No IDE or local setup required. This repository is optimized for fully AI-assisted development using Claude Code. No local toolchain, no IDE, nothing to install — everything works completely through Claude.
Security:
Java Bindings for llama.cpp
Forked from kherud/java-llama.cpp: many thanks to @kherud for the great work!
Inference of Meta's LLaMA model (and others) in pure C/C++.
You are welcome to contribute
- Features
- Quick Start
2.1 No Setup required
2.2 Setup required - Documentation
3.1 Example
3.2 Inference
3.3 Chat Completion
3.4 Infilling
3.5 Embeddings & Reranking
3.6 Raw JSON Endpoints
3.7 Local agent: terminal, browser, IDE via ACP - Android
- Feature Ideas
- Text completion (blocking and streaming) with full control over sampling parameters.
- OpenAI-compatible chat completion with automatic chat-template application, including streaming and tool/function calling support via the upstream server.
- Embeddings (single and native-batched via
embed(Collection<String>)) and reranking for retrieval pipelines. - Runtime LoRA adapter control — list the loaded adapters and change their scales at runtime without reloading the model (
getLoraAdapters()/setLoraAdapters(Map)), the typed counterpart of the upstreamGET/POST /lora-adaptersendpoints. - Text-to-speech (
TextToSpeech) over llama.cpp's Qwen3-TTS pipeline (mtmd_helper::gen_audio), returning WAV audio. - In-JVM GGUF quantization (
LlamaQuantizer) over llama.cpp'sllama_model_quantize— convert a GGUF to another quantization scheme without shelling out tollama-quantize. - Kolibri-1 (Aleph Alpha's German/English reasoning MoE, architecture
kolibri1) before upstream llama.cpp supports it, through a carried patch that loads the GGUFs of both community converters. See Kolibri-1. - Decision models (
handleSystemOne) — llama.cpp's TypeSafe-compatible/v1/systemoneAPI: typedchoice/score/noulquestions about a state, answered with probabilities in one forward pass, no token generated (laya, julia-1, lev, openjev, kev and the later decision models). See Decision models. - Infilling (fill-in-the-middle) for code models.
- Tokenize / detokenize and JSON-schema → grammar conversion.
- Raw JSON endpoint handlers mirroring the upstream llama.cpp HTTP server (
/completions,/v1/completions,/embeddings,/infill,/tokenize,/detokenize). - Two runnable HTTP server modes, one entry point. The classes jar's
Main-ClassisServerLauncher, which dispatches on the--jllama-openai-compatflag. Without it,jbang --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64 net.ladenthin:llama:5.2.0 -m model.gguf --port 8080runs the full upstream llama.cpp server (embedded WebUI, every llama-server flag forwarded) hosted insidelibjllamaover JNI — no separatellama-server.exe. With it, the same line plus--jllama-openai-compat --model model.gguf --port 8080runs the Java-transport, zero-extra-dependency OpenAI-compatible server (OpenAiCompatServer, streaming SSE) instead. Both are also runnable directly by class name viajava -cp … net.ladenthin.llama.server.{NativeServer,OpenAiCompatServer}; see Running the server from Maven Central. - Model metadata access (
getModelMeta()) and server management (metrics, slot save/restore, runtime thread reconfiguration). - Conversation checkpoints —
Session.checkpoint(...)/rewind(...)/fork(...)branch and roll back a chat (KV-cache slot save/restore + transcript snapshot) without re-prefilling. - GGUF metadata inspection without loading the model (
GgufInspector— pure Java, reads header + key/value table only, big-endian aware). - Distributed inference over RPC — offload a model's layers to llama.cpp RPC servers on other machines (
ModelParameters.setRpcServers(...)/--rpc host:port), and serve this machine's devices to them withRpcServer(the in-JVMrpc-server). See Distributed inference over RPC. - Local agent (
llama-atmosphere-agent, release asset, JDK 21+) — a fully offline agent on top of this library that reads and edits files and, if allowed, runs commands. One agent session, three ways to use it: a terminal (full console or line-oriented--plain, fine over SSH/PuTTY), a browser (--web, token-protected, loopback by default — reach it from elsewhere through an SSH tunnel), and IDEs over the Agent Client Protocol (--acp: JetBrains IDEs and Zed natively, VS Code through an ACP extension). See Local agent. - Multi-model router mode (
--models-dir+ per-request model selection, managed via the typedRouterClient) and attach mode (NativeServer(LlamaModel, ...)serves an already-loaded model over the full upstream HTTP frontend — one copy of the weights). - HTTPS built in — model downloads from
https://URLs (--model-url,-hf) and a TLS-capable embedded server (--ssl-key-file/--ssl-cert-file) on every desktop platform, with BoringSSL linked statically: no OpenSSL to install, and no dependency on one (Android and s390x ship without SSL). - Pre-built native binaries for Linux (x86-64, aarch64, s390x), macOS (arm64, Metal included), Windows (x86-64, x86, arm64) and Android (arm64, x86-64), plus GPU backends (CUDA, Vulkan, OpenCL, ROCm/HIP, SYCL, OpenVINO) — one natives jar each, all loadable side by side with automatic CPU fallback; see Choosing the natives jars. Android additionally ships as the
llama-androidAAR with the optionalllama-kotlincoroutines façade.
Access this library via Maven (released versions on Maven Central). llama-platform brings
the classes plus the CPU natives of every desktop platform:
<dependency>
<groupId>net.ladenthin</groupId>
<artifactId>llama-platform</artifactId>
<version>5.2.0</version>
<type>pom</type>
</dependency>(<type>pom</type> is required in Maven: llama-platform is a dependency list, not a jar.
In Gradle it is just implementation("net.ladenthin:llama-platform:5.2.0").)
Note
This layout starts with 5.2.0. Up to 5.1.0, net.ladenthin:llama was one jar that
carried the CPU natives of every platform, and each GPU classifier was a complete
replacement for it.
There are multiple examples.
examples/jbang/Chat.java is a one-file console chat that
JBang runs straight from the repository, with any JDK 8+ and a GGUF of an
instruction-tuned model:
jbang https://github.com/bernardladenthin/java-llama.cpp/blob/main/examples/jbang/Chat.java model.ggufIts //DEPS lines name the classes jar and the CPU natives jar of every desktop platform (the jars
llama-platform names; JBang treats a pom dependency as a BOM and puts nothing of it on the
classpath), and the loader picks this machine's. Copy the file as a starting point; a GPU backend is
one more //DEPS line (see Choosing the natives jars).
Every push to main publishes a snapshot to the Sonatype Central snapshot repository.
To use the latest snapshot, add the repository and dependency to your pom.xml:
<repositories>
<repository>
<id>sonatype-snapshots</id>
<url>https://central.sonatype.com/repository/maven-snapshots/</url>
<snapshots><enabled>true</enabled></snapshots>
<releases><enabled>false</enabled></releases>
</repository>
</repositories>
<dependency>
<groupId>net.ladenthin</groupId>
<artifactId>llama-platform</artifactId>
<version>5.2.0-SNAPSHOT</version>
<type>pom</type>
</dependency>No credentials are required — the repository is publicly readable.
We support CPU inference for the following platforms out of the box:
- Linux x86-64, aarch64, s390x
- macOS aarch64 (Apple silicon, with Metal)
- Windows x86-64, aarch64
- Android aarch64, x86-64 (see Importing in Android)
If any of these match your platform, you can include the Maven dependency and get started.
net.ladenthin:llama is the Java classes only. The native code ships as separate jars of the
same artifact, selected by a Maven <classifier> of the form <backend>-<os>-<arch>, of two kinds:
- One library jar per platform (
cpu-<os>-<arch>,metal-macos-aarch64):libjllamawith llama.cpp, ggml's shared libraries and the CPU backend modules (one per instruction-set level, the best one for the running CPU is loaded at start). Exactly one is loaded, the one of the running platform;llama-platformcollects them for every desktop platform. - GPU module jars (
cuda13-…,rocm-…,sycl-…,vulkan-…,opencl-…,openvino-…): each holds only ggml's backend module for that GPU (libggml-cuda.so,ggml-vulkan.dll, …). They are additive: the loader puts every GPU module jar of the platform next to the library, ggml loads each module whose vendor runtime is installed and registers its devices, and a module whose runtime is missing is skipped with a log line — the model then runs on the CPU, so a GPU jar on a machine without that GPU just works. Several GPU jars together are fine (llama.cpp lists every device they bring, the same GPU seen through two backends once). A module jar alone is not enough: it needs the library jar of its platform, from the same release — the loader compares the build stamp (jllama-build.txt, the llama.cpp tag) of every module jar with the library's and refuses a mismatch before loading anything, which the module's own export check at load time cannot see.
Each jar holds exactly one directory, net/ladenthin/llama/<OS>/<ARCH>/<backend>/, so any
combination can share one classpath. The start-up log names the library, the modules put in place
and the backends and devices ggml registered.
llama-platform is the classes jar plus the library jars of every desktop platform. For GPU
acceleration, add the module jar for your GPU next to it:
<dependency>
<groupId>net.ladenthin</groupId>
<artifactId>llama-platform</artifactId>
<version>5.2.0</version>
<type>pom</type>
</dependency>
<!-- Add any natives jar from the table below. Example shown: CUDA 13 on Linux x86-64. -->
<dependency>
<groupId>net.ladenthin</groupId>
<artifactId>llama</artifactId>
<version>5.2.0</version>
<classifier>cuda13-linux-x86-64</classifier>
</dependency>To ship less, depend on net.ladenthin:llama (the classes) plus only the natives jars of the
platforms you target, e.g. cpu-linux-x86-64.
| Classifier | Backend | Target platform | Runtime requirement |
|---|---|---|---|
cpu-linux-x86-64 |
CPU | Linux x86-64 | A JDK 8+ JVM; glibc ≥ 2.28 (manylinux_2_28: RHEL 8, Ubuntu 20.04, Debian 10 and later). Ships one CPU backend module per instruction-set level (x86-64 baseline, SSE4.2, AVX, AVX2, AVX-512, AVX-VNNI, AMX), of which ggml loads the best for the running CPU at start-up — a CPU without AVX2 works, and an AVX-512/AMX machine uses its kernels. |
cpu-linux-aarch64 |
CPU | Linux aarch64 | glibc ≥ 2.28 (RHEL 8, Ubuntu 20.04, Debian 10, Amazon Linux 2023 and later) — built in the manylinux_2_28 image on an arm64 runner. Ships one CPU backend module per ARM feature level (armv8.0 up to armv9.2: dotprod, fp16, SVE, i8mm, SVE2, SME), chosen at start-up like the x86-64 ones. |
cpu-linux-s390x |
CPU | Linux s390x (IBM Z, big-endian) | A JDK 8+ JVM. |
cpu-windows-x86-64 |
CPU | Windows x86-64 | A JDK 8+ JVM; Windows 10 or newer. Ships one CPU backend module per instruction-set level (as the Linux x86-64 jar: x86-64 baseline, SSE4.2, AVX, AVX2, AVX-512, AVX-VNNI, AMX), of which ggml loads the best for the running CPU at start-up. Built with clang, hybrid CRT: the STL and vcruntime are static (no msvcp140.dll / vcruntime140.dll, no VC++ redistributable), the Universal CRT is the OS one (ucrtbase.dll), so Microsoft's UCRT security updates apply. |
cpu-windows-aarch64 |
CPU | Windows on ARM (Snapdragon X / Surface) | A JDK 8+ JVM; Windows 10 or newer. Built natively on windows-11-arm with clang-cl, same hybrid CRT as the x86-64 jar; one CPU backend module (ggml builds no ARM variant set for Windows), so the Vulkan and OpenCL module jars can join it. |
metal-macos-aarch64 |
Metal + CPU | macOS aarch64 (Apple silicon) | A JDK 8+ JVM. |
cpu-android-aarch64 / cpu-android-x86-64 |
CPU | Android | For Android use the llama-android AAR; these jars are the same libraries (with the CPU modules, aarch64 one per ARM feature level) for other Android JVM setups. |
cuda13-linux-x86-64 |
CUDA 13 | Linux x86-64 with NVIDIA GPU | NVIDIA driver + CUDA 13 runtime libraries (libcudart.so.13, libcublas.so.13). |
cuda13-windows-x86-64 |
CUDA 13 | Windows x86-64 with NVIDIA GPU | NVIDIA driver + CUDA 13 Toolkit (cudart64_13.dll, cublas64_13.dll, cublasLt64_13.dll on PATH). |
vulkan-linux-x86-64 |
Vulkan | Linux x86-64 with a Vulkan 1.2+ GPU (NVIDIA / AMD / Intel) | A Vulkan runtime (libvulkan.so.1), which current GPU drivers install. The most portable Linux GPU option. glibc ≈ 2.39 (built on ubuntu-latest). |
vulkan-linux-aarch64 |
Vulkan | Linux aarch64 with a Vulkan 1.2+ GPU | A Vulkan runtime (libvulkan.so.1). glibc ≥ 2.39. |
vulkan-windows-x86-64 |
Vulkan | Windows x86-64 with a Vulkan 1.2+ GPU | A Vulkan runtime (vulkan-1.dll), which current GPU drivers install. The most portable Windows GPU option. |
vulkan-windows-aarch64 |
Vulkan | Windows on ARM (Snapdragon X) with a Vulkan 1.2+ GPU | A Vulkan runtime (vulkan-1.dll), which current GPU drivers install. Built natively on windows-11-arm with clang-cl, like upstream's Windows arm64 Vulkan release. |
opencl-windows-x86-64 |
OpenCL | Windows x86-64 with an OpenCL 2.0+ GPU | A vendor OpenCL ICD (OpenCL.dll). The GGML OpenCL backend is Adreno-tuned; on desktop GPUs CUDA or Vulkan are better supported. |
opencl-windows-aarch64 |
OpenCL (Adreno) | Windows on ARM (Snapdragon X) | The Adreno driver's OpenCL ICD (OpenCL.dll). |
opencl-android-aarch64 |
OpenCL (Adreno) | Android aarch64 with Adreno GPU | A device OpenCL ICD (libOpenCL.so); see also the llama-android-opencl AAR. |
rocm-linux-x86-64 |
ROCm / HIP | Linux x86-64 with AMD GPU | An AMD ROCm 10 runtime (libamdhip64.so, librocblas.so, libhipblas.so) — built against ROCm 10.0 (TheRock), like upstream llama.cpp; every GPU TheRock builds for Linux, Instinct included (gfx900/gfx906/gfx90c/gfx1153 best effort — built, but not release-ready in ROCm 10). |
rocm-windows-x86-64 |
ROCm / HIP | Windows x86-64 with AMD GPU | The AMD ROCm 10 runtime DLLs (amdhip64.dll, rocblas.dll, hipblas.dll) on PATH; every Radeon target TheRock builds for Windows, gfx900 through RDNA4 (gfx900/gfx906/gfx90c/gfx1153 best effort). |
sycl-linux-x86-64 |
SYCL (Intel oneAPI, fp16) | Linux x86-64 with Intel GPU (Arc / iGPU) | An Intel oneAPI / Level-Zero runtime. fp16 accumulation, as upstream's SYCL release builds. |
sycl-windows-x86-64 |
SYCL (Intel oneAPI, fp16) | Windows x86-64 with Intel GPU (Arc / iGPU) | The Intel oneAPI / Level-Zero runtime DLLs on PATH. |
openvino-linux-x86-64 |
OpenVINO | Linux x86-64 (Intel GPU / NPU / CPU) | An Intel OpenVINO runtime. |
openvino-windows-x86-64 |
OpenVINO | Windows x86-64 (Intel GPU / NPU / CPU) | The Intel OpenVINO runtime DLLs on PATH. |
Note
No vendor runtime is bundled; it comes from the GPU driver or toolkit on the host. The GPU
module jars are validated build-only in CI (GitHub runners have no GPU): the smoke jobs put
every GPU module of the platform next to the library and see ggml skip each one for lack of a
runtime, so end-to-end GPU inference is verified locally / on self-hosted hardware. A module
whose runtime is installed but finds no device (e.g. the CUDA toolkit without an NVIDIA GPU)
simply registers none. -Dnet.ladenthin.llama.backend=<a>,<b> restricts the modules put in
place to the named ones (cpu for none) and fails loud when a named jar is not on the
classpath; unset, every GPU module jar on the classpath is used.
Note
On the module path each natives jar is an automatic module
(net.ladenthin.llama.natives.<classifier>, with _ for -) that nothing requires, so
resolve them with --add-modules ALL-MODULE-PATH (or keep the natives jars on the
classpath).
Note
Android armeabi-v7a (32-bit ARM) is not published. Only 64-bit
Android binaries are shipped: aarch64 (devices) and x86_64
(emulators, Chromebooks, x86-64 Android hardware) as the cpu-android-*
natives jars and in the llama-android AAR, plus aarch64 as
opencl-android-aarch64. 32-bit Android devices are unsupported
by the released artifacts.
The minimum required Android version is API 28 (Android 9.0 Pie). Devices running Android 8.1 (API 27) or earlier are not supported.
There is no all-in-one jar to download: the natives are modular, and every combination of them is
a classpath Maven resolves. The classes jar's Main-Class is ServerLauncher, so the
embedded server starts straight from the
coordinates with JBang — the library jar of your platform, plus any GPU
module jars, as --deps:
# CPU only (Linux x86-64; cpu-windows-x86-64, metal-macos-aarch64, ... for the others)
jbang --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64 net.ladenthin:llama:5.2.0 -m model.gguf --port 8080
# with a GPU backend -- several may be named, ggml uses the ones whose runtime is installed
jbang --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64,net.ladenthin:llama:5.2.0:vulkan-linux-x86-64 \
net.ladenthin:llama:5.2.0 -m model.gguf --port 8080 -ngl 99Every llama-server flag works, the WebUI included; add --jllama-openai-compat to run the Java
OpenAI-compatible server instead. With Maven, from a checkout,
examples/server/pom.xml depends on llama-platform and carries one
profile per GPU module jar, named after its classifier:
mvn -f examples/server/pom.xml exec:java -Dexec.args="-m model.gguf --port 8080" # CPU
mvn -f examples/server/pom.xml -P cuda13-linux-x86-64 exec:java -Dexec.args="-m model.gguf -ngl 99" # + CUDA
# a local jar-with-dependencies of exactly that combination, for a machine without Maven
mvn -f examples/server/pom.xml -P assembly,cuda13-linux-x86-64 package
java -jar examples/server/target/llama-server-5.2.0-jar-with-dependencies.jar -m model.gguf -ngl 99The GitHub Release of a version attaches the Maven artifacts themselves — the classes jar, every
natives jar, the poms, sources and javadoc, each with the detached GPG .asc that signed it for
Central — and nothing else; a jar of every backend of every platform would be the opposite of
modular natives.
If none of the above listed platforms matches yours, or you want a GPU backend no module jar exists for (e.g. ROCm on Linux aarch64), you have to compile the library yourself.
This consists of two steps: 1) Compiling the libraries and 2) putting them in the right location.
First, have a look at llama.cpp to know which build arguments to use (e.g. for CUDA support).
Any build option of llama.cpp works equivalently for this project.
You then have to run the following commands in the llama/ module directory (the native core lives
there; the repository root is just the Maven reactor aggregator):
cd llama # the native core module
mvn compile # don't forget this line
cmake -B build # add any other arguments for your backend, e.g. -DGGML_CUDA=ON
cmake --build build --config ReleaseTip
Use -DLLAMA_CURL=ON to download models via Java code using ModelParameters#setModelUrl(String).
The library is put in a directory matching your platform and backend, which appears in the cmake output. For example:
-- Backend 'cpu' - installing files to /java-llama.cpp/llama/src/main/natives/net/ladenthin/llama/Linux/x86_64/cpumvn test puts that directory on the test classpath; mvn -P natives package turns it into a
natives jar (in CI, together with every other platform's).
This project has to load a single shared library jllama.
Note, that the file name varies between operating systems, e.g., jllama.dll on Windows, jllama.so on Linux, and jllama.dylib on macOS.
The application will search in the following order in the following locations:
- In net.ladenthin.llama.lib.path: Use this option if you want a custom location for your shared libraries, i.e., set VM option
-Dnet.ladenthin.llama.lib.path=/path/to/directory. - In java.library.path: These are predefined locations for each OS, e.g.,
/usr/java/packages/lib:/usr/lib64:/lib64:/lib:/usr/libon Linux. You can find out the locations usingSystem.out.println(System.getProperty("java.library.path")). Use this option if you want to install the shared libraries as system libraries. - From the natives jars on the classpath: every backend directory found for your platform, in the order described in Choosing the natives jars.
Every net.ladenthin.llama.* system property recognised by the library, deep-scanned from the source. Runtime properties are resolved through LlamaSystemProperties; test-only properties are declared in the test sources (TestConstants) and consumed by individual test classes.
| Property | Default | Scope | Consumer | Description |
|---|---|---|---|---|
net.ladenthin.llama.lib.path |
unset (falls back to java.library.path) |
runtime | LlamaLoader |
Directory containing the native jllama shared library. Checked first, before java.library.path. Set with -Dnet.ladenthin.llama.lib.path=/path/to/dir. |
net.ladenthin.llama.tmpdir |
unset (falls back to java.io.tmpdir) |
runtime | LlamaLoader |
Directory the natives are extracted into, one subdirectory per backend and build (jllama-backend-<backend>-<key>). It is kept across runs: the next start of the same build compares each file with the jar and copies nothing; a directory of a build that is no longer on the classpath is removed by a later start once it is 10 minutes old. Point it elsewhere when the default temp directory is on a slow or policy-restricted volume. |
net.ladenthin.llama.osinfo.architecture |
unset (uses os.arch) |
runtime | OSInfo |
Override for the architecture string used to locate the bundled library inside the JAR. Useful when os.arch reports an unexpected value (e.g. inside dockcross / chrooted environments). |
net.ladenthin.llama.backend |
unset (every GPU module jar on the classpath is put next to the library) | runtime | LlamaLoader |
A comma-separated filter over the GPU module jars (e.g. vulkan, cuda13,vulkan): only the named modules are put in place, cpu names none; a named module whose jar is not on the classpath, or an unknown name, fails the load. The library jar is never a choice — it is the one of the running platform. See Choosing the natives jars. |
net.ladenthin.llama.test.ngl |
43 for the general suite; 0 for ToolCallingIntegrationTest |
test | Model-backed integration tests | Number of GPU layers used during testing. Pin to 0 on CPU-only hosts: mvn test -Dnet.ladenthin.llama.test.ngl=0. The tool test also selects device none at zero layers so Metal/CUDA is not initialized. |
net.ladenthin.llama.text.model |
models/codellama-7b.Q2_K.gguf (tests self-skip if missing) |
test | LlamaModelTest and most model-backed tests (TestConstants.MODEL_PATH) |
Path to the main text-generation GGUF. Most tests only need an instruct-capable model; a few assertions are tuned to the CI model. |
net.ladenthin.llama.draft.model |
models/AMD-Llama-135m-code.Q2_K.gguf (tests self-skip if missing) |
test | Speculative-decoding tests and the tests that need a small, fast model (TestConstants.DRAFT_MODEL_PATH) |
Path to the draft GGUF. |
net.ladenthin.llama.reasoning.model |
models/Qwen3-0.6B-Q4_K_M.gguf (tests self-skip if missing) |
test | Reasoning-budget tests (TestConstants.REASONING_MODEL_PATH) |
Path to a thinking model (one that emits a thinking block). |
net.ladenthin.llama.rerank.model |
models/jina-reranker-v1-tiny-en-Q4_0.gguf (tests self-skip if missing) |
test | Reranking tests (TestConstants.RERANKING_MODEL_PATH) |
Path to a reranker GGUF (must load with enableReranking()). |
net.ladenthin.llama.tool.model |
models/Qwen2.5-1.5B-Instruct-Q4_K_M.gguf (test self-skips if missing) |
test | ToolCallingIntegrationTest |
Path to a tool-capable GGUF used to verify required blocking and streaming tool calls. The default matches the Qwen2.5 model in upstream llama.cpp's tool-call test matrix. |
net.ladenthin.llama.nomic.path |
models/nomic-embed-text-v1.5.f16.gguf (test self-skips if missing) |
test | LlamaEmbeddingsTest#testNomicEmbedLoads |
Path to a Nomic embedding model (nomic-embed-text-v1.5.f16.gguf or a compatible BERT-family encoder). Regression test for upstream issue #98 (BERT-encoder result_output assertion). |
net.ladenthin.llama.vision.model |
models/SmolVLM-500M-Instruct-Q8_0.gguf (test self-skips if missing) |
test | MultimodalIntegrationTest |
Path to a vision-capable model GGUF. Any vision-capable GGUF works; CI default is SmolVLM-500M-Instruct-Q8_0.gguf. |
net.ladenthin.llama.vision.mmproj |
models/mmproj-SmolVLM-500M-Instruct-Q8_0.gguf (test self-skips if missing) |
test | MultimodalIntegrationTest |
Matching mmproj GGUF for the vision model. |
net.ladenthin.llama.vision.image |
llama/src/test/resources/images/test-image.jpg (a CC-BY-4.0 / MIT-granted photo committed to the repo) |
test | MultimodalIntegrationTest |
Visual prompt image. Any png/jpeg/webp/gif works; the extension drives MIME detection. |
net.ladenthin.llama.decision.model |
unset (test self-skips) | test | SystemOneIntegrationTest |
Path to a decision model GGUF (laya, julia-1, lev, openjev, kev, ...) for the /v1/systemone tests of LlamaModel.handleSystemOne; upstream tests with ggml-org/tinylaya-for-testing-gguf. Without it only the rejection of a non-decision model runs. |
net.ladenthin.llama.audio.model |
unset (test self-skips) | test | AudioInputIntegrationTest (llama.cpp discussion #13759) |
Path to an audio-input model GGUF (e.g. Ultravox, Qwen2.5-Omni). |
net.ladenthin.llama.audio.mmproj |
unset (test self-skips) | test | AudioInputIntegrationTest |
Matching audio mmproj (encoder) GGUF. |
net.ladenthin.llama.audio.input |
src/test/resources/audios/sample.wav (committed) |
test | AudioInputIntegrationTest |
.wav/.mp3 audio prompt clip; the extension drives format detection. |
net.ladenthin.llama.tts.model |
models/Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf (test self-skips if missing) |
test | TtsIntegrationTest |
Path to the Qwen3-TTS backbone (text) GGUF. Any Qwen3-TTS-family model works; CI default is Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf. |
net.ladenthin.llama.tts.mmproj |
models/mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0.gguf (test self-skips if missing) |
test | TtsIntegrationTest |
Path to the matching Qwen3-TTS mmproj GGUF (speaker encoder + code predictor + code2wav decoder); CI default is mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0.gguf. |
net.ladenthin.llama.train.model |
models/stories260K.gguf (test self-skips if missing) |
test | LlamaTrainerIntegrationTest |
Path to the model the fine-tuning smoke trains. Must be F32: llama_set_param skips every other tensor type, so a quantized model would train nothing. |
The test-model defaults are exactly CI's model set (.github/models.csv, URLs included), so a model downloaded into models/ is found without any property; the properties only point a test at another file. MultimodalIntegrationTest self-skips when any of the three vision.* paths is missing, so a partial setup (just the vision model + the committed image, no mmproj) lets the test class load without erroring. AudioInputIntegrationTest self-skips the same way over the three audio.* properties. TtsIntegrationTest likewise self-skips unless both tts.model and tts.mmproj paths exist, and LlamaTrainerIntegrationTest unless train.model (default models/stories260K.gguf, an F32 model) does.
This is a short example on how to use this library:
public class Example {
public static void main(String... args) throws IOException {
ModelParameters modelParams = new ModelParameters()
.setModel("models/mistral-7b-instruct-v0.2.Q2_K.gguf")
.setGpuLayers(43);
String system = "This is a conversation between User and Llama, a friendly chatbot.\n" +
"Llama is helpful, kind, honest, good at writing, and never fails to answer any " +
"requests immediately and with precision.\n";
BufferedReader reader = new BufferedReader(new InputStreamReader(System.in, StandardCharsets.UTF_8));
try (LlamaModel model = new LlamaModel(modelParams)) {
System.out.print(system);
String prompt = system;
while (true) {
prompt += "\nUser: ";
System.out.print("\nUser: ");
String input = reader.readLine();
prompt += input;
System.out.print("Llama: ");
prompt += "\nLlama: ";
InferenceParameters inferParams = new InferenceParameters(prompt)
.withTemperature(0.7f)
.withMiroStat(MiroStat.V2)
.withStopStrings("User:");
for (LlamaOutput output : model.generate(inferParams)) {
System.out.print(output);
prompt += output;
}
}
}
}
}Also have a look at the other examples.
There are multiple inference tasks. In general, LlamaModel is stateless, i.e., you have to append the output of the
model to your prompt in order to extend the context. If there is repeated content, however, the library will internally
cache this, to improve performance.
ModelParameters modelParams = new ModelParameters().setModel("/path/to/model.gguf");
InferenceParameters inferParams = new InferenceParameters("Tell me a joke.");
try (LlamaModel model = new LlamaModel(modelParams)) {
// Stream a response and access more information about each output.
for (LlamaOutput output : model.generate(inferParams)) {
System.out.print(output);
}
// Calculate a whole response before returning it.
String response = model.complete(inferParams);
// Returns the hidden representation of the context + prompt.
float[] embedding = model.embed("Embed this");
}Note
Since llama.cpp allocates memory that can't be garbage collected by the JVM, LlamaModel is implemented as an
AutoClosable. If you use the objects with try-with blocks like the examples, the memory will be automatically
freed when the model is no longer needed. This isn't strictly required, but avoids memory leaks if you use different
models throughout the lifecycle of your application.
For chat models, build a list of role/content pairs and let the library apply the model's chat template.
chatComplete() returns the full response, generateChat() streams tokens, and chatCompleteText() returns
just the text content of the assistant message.
List<Pair<String, String>> messages = new ArrayList<>();
messages.add(new Pair<>("user", "Write a haiku about Java."));
InferenceParameters inferParams =
new InferenceParameters("").withMessages("You are a helpful assistant.", messages);
try (LlamaModel model = new LlamaModel(modelParams)) {
// Streaming
for (LlamaOutput output : model.generateChat(inferParams)) {
System.out.print(output);
}
// Or blocking, returns the OpenAI-compatible JSON envelope
String json = model.chatComplete(inferParams);
// Or just the assistant text
String text = model.chatCompleteText(inferParams);
}Reasoning/thinking models can receive custom Jinja template variables via
ModelParameters#setChatTemplateKwargs(Map).
Load a vision-capable GGUF with its matching projector, then place text and image parts in the same user message. Images may come from a file, raw bytes, a data URI, or an HTTP(S) URL:
ModelParameters modelParams = new ModelParameters()
.setModel("models/SmolVLM-500M-Instruct-Q8_0.gguf")
.setMmproj("models/mmproj-SmolVLM-500M-Instruct-Q8_0.gguf");
ChatMessage message = ChatMessage.userMultimodal(
ContentPart.text("Describe this image in one short sentence."),
ContentPart.imageFile(Paths.get("photo.jpg")));
try (LlamaModel model = new LlamaModel(modelParams)) {
String answer = model.chatCompleteText(InferenceParameters.empty()
.withMessages(Collections.singletonList(message))
.withNPredict(64));
System.out.println(answer);
}The same multipart messages[].content shape works through ChatRequest and the embedded
OpenAI-compatible /v1/chat/completions server. For a strictly CPU-only run, use
setDevices("none").setMmprojOffload(false) in addition to setGpuLayers(0); projector offload
has its own upstream default.
On a multi-GPU host the projector can be placed independently of the weights with
setMmprojDevice("CUDA1") (llama.cpp --mmproj-device, added upstream in b10541). Exactly one
device may be named; the literal "none" keeps the projector on the CPU. OpenAiCompatServer's CLI
accepts the same flag as -mmdev/--mmproj-device, and NativeServer forwards it verbatim like
every other llama-server flag.
setMmprojDevice(...) and setMmprojOffload(...) write the same upstream field
(common_params::mmproj_use_gpu), so where they disagree the outcome would depend on argv order —
and the rendered argv comes from a HashMap, whose order is unspecified. The builder therefore
resolves the two genuinely ambiguous combinations by dropping the earlier call, and leaves the rest
alone:
| Combination | Resolves to | Builder behaviour |
|---|---|---|
named device + setMmprojOffload(true) |
(use_gpu=true, device) in either order |
both kept — no clash |
named device + setMmprojOffload(false) |
order-dependent | last call wins |
"none" + setMmprojOffload(true) |
order-dependent | last call wins |
"none" + setMmprojOffload(false) |
(use_gpu=false) in either order |
both kept — no clash |
So a multi-GPU projector pin survives an explicit setMmprojOffload(true); only a call that would
actually contradict the other is dropped. If you need a device after disabling offload, call
setMmprojDevice last.
Video input — decode settings only, so far. mtmd has carried a video path since llama.cpp
b9562 (#24269); b10647 (#24318) added the --video-* CLI flags and the
mtmd_helper_init_opt plumbing that surfaces them. It is compiled into the shipped desktop library
(MTMD_VIDEO is on by default, gated on LLAMA_SUBPROCESS, which upstream force-disables on
Android and iOS). Its decode settings are exposed
as setVideoFps(float), setVideoTimestampInterval(long) and setVideoFfmpegDir(String). The last
one matters most in a JVM: upstream shells out to ffmpeg/ffprobe and resolves them from PATH,
which an application server, an Android app or a JAR-only container frequently does not have them on —
naming the directory is then the only way for video to work at all.
What is not here yet is the content part: upstream's wire type for a video is
{"type":"input_video","input_video":{"data":"<base64>"}} (raw base64, not a data: URI, unlike
image_url), gated server-side on mtmd_helper_support_video. ContentPart has no videoFile(...)
factory emitting that shape, so these knobs currently configure a path this API cannot yet feed
directly. Tracked in TODO.md.
Audio input works identically — load an audio-capable model (Ultravox, Qwen2.5-Omni, …) with its
audio --mmproj and add a ContentPart.audioFile(...) (or inputAudio(bytes, "wav"|"mp3")) part. It
serializes to the OpenAI input_audio content part and routes through the same mtmd pipeline:
ModelParameters modelParams = new ModelParameters()
.setModel("models/ultravox-v0_5-llama-3_2-1b.gguf")
.setMmproj("models/mmproj-ultravox-v0_5-llama-3_2-1b-f16.gguf");
ChatMessage message = ChatMessage.userMultimodal(
ContentPart.text("Transcribe the audio."),
ContentPart.audioFile(Paths.get("speech.wav")));
try (LlamaModel model = new LlamaModel(modelParams)) {
System.out.println(model.supportsAudio()); // true
String answer = model.chatCompleteText(InferenceParameters.empty()
.withMessages(Collections.singletonList(message))
.withNPredict(64));
System.out.println(answer);
}LlamaModel.supportsVision() / supportsAudio() report which modalities the loaded projector enables.
Use a tool-aware instruct model and enable Jinja when loading it. A typed request can either return
the model's tool calls through chat, or execute registered handlers until the model produces a
normal assistant response through chatWithTools:
ToolDefinition weather = new ToolDefinition(
"get_weather",
"Get the current weather for a city",
"{\"type\":\"object\",\"properties\":{\"city\":{\"type\":\"string\"}},"
+ "\"required\":[\"city\"]}");
ChatRequest request = ChatRequest.empty()
.appendMessage("user", "What is the weather in Paris?")
.appendTool(weather)
.withToolChoice("auto")
.withParallelToolCalls(Boolean.FALSE);
Map<String, ToolHandler> handlers = Collections.singletonMap(
"get_weather", argumentsJson -> "{\"temperature_c\":21,\"condition\":\"sunny\"}");
try (LlamaModel model = new LlamaModel(new ModelParameters()
.setModel("models/Qwen2.5-1.5B-Instruct-Q4_K_M.gguf")
.enableJinja())) {
ChatResponse response = model.chatWithTools(request, handlers);
System.out.println(response.getFirstContent());
}tool_choice is the OpenAI-compatible string form (auto, none, or required). Set
parallel_tool_calls to false when handlers should be issued one at a time. Handler failures and
unknown tool names are returned to the model as valid {"error":"..."} tool-result JSON.
You can simply set InferenceParameters#withInputPrefix(String) and InferenceParameters#withInputSuffix(String).
Load the model with enableEmbedding() (or enableReranking()) and call embed(String) to get a sentence
embedding, or rerank(query, documents...) to get relevance scores.
ModelParameters modelParams = new ModelParameters()
.setModel("/path/to/embedding-model.gguf")
.enableEmbedding();
try (LlamaModel model = new LlamaModel(modelParams)) {
float[] embedding = model.embed("Embed this sentence");
// Batch form: one native dispatch for many inputs, results in request order.
List<float[]> embeddings = model.embed(Arrays.asList("First sentence", "Second sentence"));
}Kolibri-1 is Aleph Alpha's 78B-parameter Mixture-of-Experts
reasoning model for German and English (about 3.5B parameters active per token, Apache-2.0). Upstream llama.cpp
does not support its architecture yet (ggml-org/llama.cpp#29922),
so this library carries it as a patch (llama/patches/0016-model-kolibri1.patch) until upstream does. It loads the
community GGUFs of both published converters -- e.g. Hob-forge/Kolibri-1-GGUF
and Eliasfpv28/Kolibri-1-Q3_K_S-GGUF -- which the two
community llama.cpp patches cannot each load from the other. The embedded chat template handles reasoning and
Hermes-style tool calls; for a split GGUF, point the model path at the first part.
ModelParameters params = new ModelParameters()
.setModel("/models/Kolibri-1-Q4_K_M.gguf")
.setCtxSize(32768)
.enableJinja();
try (LlamaModel model = new LlamaModel(params)) {
ChatResponse answer = model.chat(ChatRequest.empty()
.appendMessage("user", "Warum ist der Himmel blau?"));
}The model is large (the Q4_K_M file is 47.5 GB) and runs from system RAM on the CPU, or partly offloaded to a GPU. The architecture is checked numerically against Aleph Alpha's reference on tiny random models in every C++ test run; the real model was not run in this project's CI, and the GPU backends are untested for it.
A decision model answers typed questions about a state without generating text: each question is
evaluated in one forward pass and returns probabilities. handleSystemOne takes and returns
llama.cpp's TypeSafe-compatible /v1/systemone JSON, served by the upstream handler itself (the full
request/response description is in upstream's tools/server/README.md). The native server serves the
same endpoint at POST /v1/systemone, in attach mode too.
try (LlamaModel model = new LlamaModel(new ModelParameters().setModel("/path/to/laya.gguf"))) {
String answers = model.handleSystemOne("{"
+ "\"state\": \"I was charged twice for my order and nobody replied.\","
+ "\"questions\": {"
+ " \"route\": {\"type\": \"choice\", \"instructions\": \"Which team?\","
+ " \"criteria\": {\"billing\": null, \"shipping\": null}},"
+ " \"urgency\": {\"type\": \"score\", \"instructions\": \"How urgent?\","
+ " \"criteria\": [\"can wait\", \"today\", \"right now\"]},"
+ " \"angry\": {\"type\": \"noul\", \"instructions\": \"Is the customer angry?\"}}}");
// {"answers": {"route": {"choice": "billing", "probabilities": {...}, ...}, ...}, "usage": {...}}
}A model that is not a decision model throws a LlamaException ("This model is not a decision model").
Adapters loaded at model-load time (addLoraAdapter(...) / addLoraScaledAdapter(...), optionally
setLoraInitWithoutApply() to start disabled) can be listed and re-scaled at runtime without
reloading the model — the typed counterpart of the upstream GET/POST /lora-adapters endpoints:
ModelParameters modelParams = new ModelParameters()
.setModel("models/base.gguf")
.addLoraScaledAdapter("models/adapter.gguf", 1.0f);
try (LlamaModel model = new LlamaModel(modelParams)) {
List<LoraAdapter> adapters = model.getLoraAdapters(); // [{id=0, path=..., scale=1.0}]
model.setLoraAdapter(0, 0.5f); // re-scale at runtime
model.setLoraAdapters(Collections.emptyMap()); // disable all adapters
}Per the upstream contract, a scale update lists the adapters to keep active — any adapter missing
from the map is set to scale 0 (disabled). The native side clears affected KV caches when the
effective adapter set changes.
TextToSpeech synthesizes audio from text over llama.cpp's upstream Qwen3-TTS pipeline
(mtmd_helper::gen_audio). It is a separate AutoCloseable native type (not a LlamaModel)
because TTS loads its own model pair: a backbone (text) GGUF and an mmproj GGUF bundling the
speaker encoder, code predictor, and code2wav decoder. synthesize(String) returns a
24 kHz mono 16-bit WAV byte stream.
try (TextToSpeech tts = new TextToSpeech(
"models/qwen3-tts-backbone.gguf", "models/qwen3-tts-mmproj.gguf")) {
byte[] wav = tts.synthesize("Hello from llama dot c p p.");
Files.write(Paths.get("out.wav"), wav);
}Add (modelPath, mmprojPath, gpuLayers, threads) to offload to the GPU, or
synthesize(text, maxFrames, topK, seed) for explicit sampling, or the full
synthesize(text, speakerReferenceAudioPath, language, maxFrames, topK, seed) overload for
voice cloning from a reference clip. As with LlamaModel, native memory is not GC-managed — use
try-with-resources or call close().
This replaces the project's earlier two-model OuteTTS + WavTokenizer pipeline, which upstream
#26254 deleted entirely in favor of Qwen3-TTS
(llama.cpp b10270); there is no backward-compatible path for the old model pair.
LlamaQuantizer converts a GGUF to another quantization scheme in-process (llama.cpp's
llama_model_quantize — the llama-quantize tool without the separate binary):
LlamaQuantizer.quantize("model-f16.gguf", "model-q4_k_m.gguf", QuantizationType.Q4_K_M);
// Re-quantizing an already-quantized GGUF degrades quality and must be opted into:
LlamaQuantizer.quantize("model-q8_0.gguf", "model-q4_0.gguf", QuantizationType.Q4_0,
/* threads */ 0, /* allowRequantize */ true);For direct access to the upstream llama.cpp server API, the following methods take a JSON request and return a JSON response, matching the HTTP server's contract:
handleCompletions, handleCompletionsOai, handleChatCompletions, handleInfill,
handleEmbeddings, handleTokenize, handleDetokenize.
Server state is exposed via getMetrics(), eraseSlot(int), saveSlot(int, String),
restoreSlot(int, String), and getModelMeta().
A Session can be snapshotted and branched — the KV-cache slot state and the transcript move
together, so native state and history can never drift apart:
try (Session session = new Session(model, 0, "You are terse.")) {
session.send("My name is Alice.");
SessionCheckpoint cp = session.checkpoint("checkpoints/turn1.bin");
session.send("Tell me a joke.");
session.rewind(cp); // undo everything after the checkpoint
session.send("Tell me a story instead."); // retry from the branch point
// Branch into a second slot (model loaded with setParallel(2)+):
try (Session forked = session.fork(1, "checkpoints/branch.bin")) {
forked.send("Answer as a pirate."); // both sessions continue independently
}
}Checkpoint files are caller-managed (KV dumps grow with context usage) and both operations are
rejected while a stream is in progress. For plain transformer models a rewind is also achievable
cheaply by resending a truncated history with cache_prompt (prefix reuse); checkpoints make the
branch point exact and are the only reliable rollback for recurrent/hybrid models (e.g.
Granite-4), whose state cannot be recomputed from a prefix.
GgufInspector reads a GGUF's header and key/value table without loading the model — pure
Java, no native library, cost independent of file size (parsing stops before the tensor data).
Useful for model pickers and download validators:
GgufMetadata meta = GgufInspector.read(Paths.get("models/Qwen3-0.6B-Q4_K_M.gguf"));
meta.getArchitecture(); // Optional[qwen3]
meta.getModelName(); // Optional[Qwen3 0.6B]
meta.getParameterCount(); // OptionalLong[751632384]
meta.getContextLength(); // OptionalLong[40960] (<arch>.context_length)
meta.getFileType(); // OptionalLong[15] (llama_ftype, cf. QuantizationType)
meta.getChatTemplate(); // Optional[{{- ... }}]
meta.getEntries(); // full decoded key/value tableSupports GGUF v2/v3, little- and big-endian (auto-detected), and fails loud on v1/corrupt files.
For metadata of an already-loaded model use getModelMeta() instead.
Prompt-prefix reuse is enabled by default in llama.cpp and can be controlled per request with
InferenceParameters.withCachePrompt(boolean). withCacheReuse(int) enables non-prefix chunk reuse,
while withSlotId(int) pins a request to a specific server slot. Session applies its slot id to every
request, so generation and save/restore operate on the same KV state.
Typed results expose logical prompt, generated, cached prompt, and evaluated prompt counts through
Usage. Per-request timing also remains available through Timings.getCacheN().
LlamaModel.getMetricsTyped().getSlotMetrics() reports each slot's logical, processed, cached,
decoded, and remaining token counts, and the same ServerMetrics view carries the server-wide
lifetime counters — including cached prompt tokens (getCumulativeCachedPromptTokens()) and the
speculative-decoding tallies (getDraftTokensTotal(), getDraftAcceptedTotal(),
getDraftVerifyStepsTotal(), getDraftAcceptedPerPosition(), plus the derived
getDraftAcceptanceRate()), which upstream otherwise exposes only as Prometheus text.
The embedded HTTP server exposes the same native JSON at authenticated GET /metrics, with the slot
array alone at GET /slots. OpenAI responses preserve
usage.prompt_tokens_details.cached_tokens; Responses API output uses
usage.input_tokens_details.cached_tokens; Anthropic output uses cache_read_input_tokens.
net.ladenthin.llama.server.OpenAiCompatServer turns a loaded model into a local
OpenAI-compatible HTTP endpoint using only the JDK's built-in com.sun.net.httpserver — no extra
dependency and no separate server process. It is embeddable, and runnable via
java -cp <jar> net.ladenthin.llama.server.OpenAiCompatServer … (the classes jar's
Main-Class, ServerLauncher, starts NativeServer by default — see "Native server with the
built-in WebUI" below). It serves:
| Method & path | Backed by |
|---|---|
POST /v1/chat/completions |
LlamaModel.streamChatCompletion (streaming SSE) / chatComplete (blocking) |
POST /v1/completions |
LlamaModel.handleCompletionsOai |
POST /v1/embeddings (requires --embedding) |
LlamaModel.handleEmbeddings |
POST /v1/rerank (requires --reranking) |
LlamaModel.handleRerank (reshaped to results/data) |
POST /infill |
LlamaModel.handleInfill (fill-in-the-middle autocomplete) |
GET /v1/models |
the configured model id |
GET /metrics |
native server and per-slot token/cache counters (JSON) |
GET /slots |
native per-slot token/cache counters (JSON array) |
GET /health |
static {"status":"ok"} (unauthenticated) |
Chat completions support streaming via Server-Sent Events and non-streaming, forwarding
messages/tools verbatim. The streaming path carries delta.tool_calls and (with
stream_options.include_usage) a trailing usage chunk, so agent/tool-calling clients work —
this is the recommended surface for VS Code Copilot agent mode, Cline, Roo Code and Continue.
response_format (json_object / json_schema) is forwarded for structured outputs. Completions,
embeddings, rerank and infill are non-streaming.
Every route is also reachable without the /v1 prefix, the server answers CORS preflight
(OPTIONS) and stamps Access-Control-Allow-Origin (so browser/webview clients work), and
POST /infill is the llama.cpp-native FIM endpoint for local ghost-text autocomplete plugins
(llama.vscode, Twinny, Tabby, Continue's llama.cpp provider). Note: GitHub Copilot's inline
completions cannot be served by any local endpoint — only its chat/agent surfaces — so use one of
those autocomplete plugins for ghost text.
Alternative protocol surfaces. For clients that don't speak OpenAI Chat Completions, the same model is exposed through additional protocols (pure translation over the OpenAI core — no extra inference path), all supporting tools and streaming:
| Surface | Routes | For |
|---|---|---|
| Ollama-native | GET /api/version, GET /api/tags, POST /api/show, POST /api/chat (NDJSON streaming), POST /api/generate (prompt completion / FIM) |
Copilot's built-in Ollama provider; Ollama-hardcoded tools |
| Anthropic Messages | POST /v1/messages (SSE event stream) |
Claude-shaped clients (Claude Code); Copilot messages apiType |
| OpenAI Responses | POST /v1/responses (SSE event stream) |
Copilot responses apiType; Responses-API clients |
/api/show advertises the model's capabilities (tools, insert, and vision when --mmproj is set)
and context length, which Copilot's Ollama provider reads to enable agent mode. The llama.cpp-native
GET /props reports default_generation_settings.n_ctx and a modalities block, which autocomplete
clients such as llama.vscode read to size their context window.
Embed it in your app:
ModelParameters modelParams = new ModelParameters().setModel("models/model.gguf").setParallel(2);
OpenAiServerConfig config = OpenAiServerConfig.builder().port(8080).modelId("local-model").build();
try (LlamaModel model = new LlamaModel(modelParams);
OpenAiCompatServer server = new OpenAiCompatServer(model, config).start()) {
Thread.currentThread().join(); // serve until interrupted
}…or run it standalone. The classes jar's Main-Class is the ServerLauncher dispatcher, so add
--jllama-openai-compat to select this Java server (the launcher strips that flag and forwards the rest);
or name the class explicitly via -cp:
# from Maven Central (JBang resolves the classes jar, the natives jar named and the Java deps)
jbang --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64 net.ladenthin:llama:5.2.0 --jllama-openai-compat \
--model models/Qwen3-0.6B-Q4_K_M.gguf --host 0.0.0.0 --port 8080 --n-gpu-layers 99
# or name the class explicitly, with the jars on the classpath
java -cp 'target/llama-<version>.jar:target/llama-<version>-cpu-linux-x86-64.jar:<deps>' \
net.ladenthin.llama.server.OpenAiCompatServer --model models/model.gguf --port 8080 --model-id local-modelRun with --help for the full option list (-m/--model, --host, -p/--port, -c/--ctx-size,
-b/--batch-size, -ub/--ubatch-size, -ngl/--n-gpu-layers, -t/--threads, -tb/--threads-batch,
-ctk/--cache-type-k, -ctv/--cache-type-v, --jinja, --chat-template-kwargs, --parallel,
--model-id, --api-key, --mmproj, -mmdev/--mmproj-device, --embedding, --reranking). The tuning flags mirror
llama.cpp's server, so an invocation like
--jinja --chat-template-kwargs '{"reasoning_effort":"low"}' -ctk q8_0 -ctv q8_0 -b 4096 -ub 2048
works directly.
Verify with curl (streaming chat):
curl -N http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"local-model","stream":true,"messages":[{"role":"user","content":"hi"}]}'VS Code Copilot setup: Command Palette → Chat: Manage Language Models → Add Models →
Custom Endpoint; enter a group name, a display name and any non-empty API key, and pick API type
Chat Completions. VS Code then opens chatLanguageModels.json — set the model url to your
endpoint (the host/port go here, not in the form):
[
{
"name": "Local llama.cpp",
"vendor": "customendpoint",
"apiKey": "local-dummy-key",
"apiType": "chat-completions",
"models": [
{
"id": "local-model",
"name": "Local model",
"url": "http://127.0.0.1:8080/v1/chat/completions",
"toolCalling": true,
"vision": false,
"maxInputTokens": 6144,
"maxOutputTokens": 2048
}
]
}
]Notes: BYOK powers the chat/agent experience only (inline completions and embeddings still require a
GitHub account). On CPU, prefer a smaller model and a modest context window — the server emits SSE
heartbeats so a long prompt prefill does not trip the client's stream-inactivity timeout. Agent-mode
tool calling depends on the model's own tool-calling quality. Pass --api-key (or
OpenAiServerConfig.apiKey(...)) to require an Authorization: Bearer token; the server binds to
127.0.0.1 by default.
OpenAiCompatServer above is a JSON API server (its / is a 404 — no web page). If you want
the full upstream llama.cpp server, including its bundled Svelte WebUI, use
net.ladenthin.llama.server.NativeServer. It runs the real llama_server inside libjllama over
JNI — no separate llama-server.exe — and forwards the raw llama-server arguments verbatim, so
every flag works exactly as it does for the standalone binary. ServerLauncher (the classes jar's
Main-Class) runs it by default (when --jllama-openai-compat is absent), forwarding its args to
the native server (pass --help for the full llama-server option list):
jbang --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64 net.ladenthin:llama:5.2.0 \
-m models/model.gguf --host 127.0.0.1 --port 8080 -c 65536 --jinja
# then open http://127.0.0.1:8080/ for the WebUIOr embed it:
try (NativeServer server = new NativeServer(
"-m", "gpt-oss-20b-UD-Q4_K_XL.gguf",
"--host", "127.0.0.1", "--port", "8080",
"-c", "65536", "-b", "4096", "-ub", "2048",
"--jinja", "-ngl", "0", "-t", "8", "-tb", "16",
"-ctk", "q8_0", "-ctv", "q8_0",
"--chat-template-kwargs", "{\"reasoning_effort\":\"low\"}",
"--parallel", "1").start()) {
// Open http://127.0.0.1:8080/ in a browser for the WebUI; the OpenAI API is at /v1/... too.
Thread.currentThread().join();
}Differences from OpenAiCompatServer: with the classic constructor it loads its own model from
the arguments (an independent lifecycle, like llama-server.exe), it is single-instance per
process, it serves the WebUI (in released jars — local cmake builds ship the empty-asset
stub, so no UI there), and it is not available on Android (the upstream server needs
posix_spawn). Readiness: poll GET /health. TLS: pass --ssl-key-file and --ssl-cert-file
(BoringSSL is linked statically into every desktop natives jar, nothing to install). The same
build fetches models from https:// URLs (--model-url, -hf), verified against the OS
certificate store -- on Linux /etc/ssl/certs or /etc/ssl/cert.pem, overridable with
SSL_CERT_FILE / SSL_CERT_DIR. Note that --model-url alone starts the router; name the
download target with -m as well.
NativeServer can also attach the full upstream HTTP frontend (routes, WebUI, resumable
streaming) to a LlamaModel you already loaded — one copy of the weights, shared between direct
JNI calls and HTTP:
try (LlamaModel model = new LlamaModel(new ModelParameters().setModel("models/model.gguf"));
NativeServer server = new NativeServer(model, "--host", "127.0.0.1", "--port", "8080").start()) {
// HTTP (incl. WebUI in released jars) and direct Java calls share the same loaded model.
String direct = model.complete(new InferenceParameters("2+2=").withNPredict(4));
Thread.currentThread().join();
}In attach mode the arguments carry only the HTTP-side flags (--host, --port, --api-key,
--ssl-key-file, …; no -m); the routes served are the model's own, so route-level settings — the
endpoint toggles --metrics / --props / --slots, --slot-save-path, the slot count — are set on
the model's ModelParameters (enableMetricsEndpoint(), enablePropsEndpoint(),
enableSlotsEndpoint() / disableSlotsEndpoint()), not in the attach arguments. The server reports
healthy immediately (the model is already loaded), and the caller keeps ownership of the model —
close the server before the model, never the other way around.
Started without a model argument, the upstream server runs in router mode: it lists models
from --models-dir, loads/unloads them on demand (GET /models, POST /models/load,
POST /models/unload, per-request "model" selection) and serves each model from a worker
subprocess. Upstream spawns workers by re-executing its own binary — inside a JVM that binary is
java, so before starting an embedded router you must point the worker spawn at this library's
bootstrap:
String javaBin = System.getProperty("java.home") + File.separator + "bin" + File.separator + "java";
NativeServer.setWorkerCommand(javaBin, "-cp", System.getProperty("java.class.path"),
"net.ladenthin.llama.server.NativeServer");
try (NativeServer router = new NativeServer(
"--host", "127.0.0.1", "--port", "8080", "--models-dir", "models").start()) {
Thread.currentThread().join(); // each loaded model runs as a fresh worker JVM
}Worker-command tokens may not contain whitespace (the value is whitespace-split natively).
Typed model management (RouterClient). Instead of hand-rolling HTTP+JSON against the
management endpoints, use server.RouterClient — a plain-HTTP typed client (works against the
embedded router above or any external llama-server router):
RouterClient client = new RouterClient(8080);
List<RouterModel> models = client.listModels(); // GET /models, typed status per entry
client.loadModel("Qwen3-0.6B-Q4_K_M"); // POST /models/load (non-blocking)
client.awaitModelLoaded("Qwen3-0.6B-Q4_K_M", 240_000L); // poll until LOADED; fails fast if the
// worker died (exit code in the message)
client.unloadModel("Qwen3-0.6B-Q4_K_M"); // POST /models/unloadRouterModel carries the identifier, the lifecycle status
(UNLOADED/LOADING/LOADED/SLEEPING/DOWNLOADING/DOWNLOADED), the router's
failed-worker marker, and the model's input/output modalities (getInputModalities(),
getOutputModalities()), which the router computes without loading the model: isDecisionModel()
picks out a decision model for /v1/systemone before its first load.
A loaded LlamaModel reports the same through getModelMeta(). Chat requests then select a model per
request via the standard "model" field on POST /v1/chat/completions.
Against a router started with --api-key, pass the key — it is sent as
Authorization: Bearer <key> on every call. All of them need it: /models/load and
/models/unload were always gated, and since llama.cpp b10519 the listing endpoints are too.
RouterClient client = new RouterClient(8080, System.getenv("LLAMA_API_KEY"));
// or, for a remote router: new RouterClient("router.internal", 8080, key)Note
awaitModelLoaded waits by polling GET /models, so it cannot observe a model the router
deliberately hides from that listing — a cache model deduplicated by a preset with
dedup-cache-models still loads and still serves by name, but never appears. For those, skip the
await and issue the request directly; with autoload the router waits for the worker itself.
llama.cpp's RPC backend spreads one model over the devices of several machines: every machine
that contributes runs an RPC server, and the machine that loads the model names them with
--rpc. Layers are then distributed over local and remote devices exactly as over several local
GPUs (setGpuLayers, setTensorSplit). Both halves are in every natives jar — CPU and GPU —
with no additional runtime dependency (plain TCP over the system socket
library the library already links).
Serve this machine's devices (every GPU this library found, else the CPU):
try (RpcServer server = RpcServer.startLocal(RpcEndpoint.DEFAULT_PORT)) { // 127.0.0.1:50052
server.awaitTermination();
}or from the command line, from Maven Central:
jbang --main net.ladenthin.llama.RpcServer --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64 \
net.ladenthin:llama:5.2.0 --port 50052Use the servers from a model:
ModelParameters params = new ModelParameters()
.setModel("models/big-model.gguf")
.setGpuLayers(99)
.setRpcServers(RpcEndpoint.parse("10.0.0.2:50052"), RpcEndpoint.parse("10.0.0.3:50052"));The same works for both HTTP servers: --rpc 10.0.0.2:50052,10.0.0.3:50052 is forwarded to the
native server as-is, and OpenAiCompatServer accepts it too. Any upstream rpc-server works as a
server, and this library's RpcServer works for any llama.cpp client.
Warning
The RPC protocol has no authentication and no encryption: whoever reaches the port can use the
devices and read or write the tensors on them. RpcServer.startLocal therefore binds to loopback
only; RpcServer.startOnNetwork(address, …) (or --host on the command line) is the explicit
opt-in for another interface and logs a warning. Across machines, use a trusted network or a
tunnel (SSH, WireGuard).
What to know:
- An unreachable server fails the load with a
LlamaExceptionnaming it, instead of reaching llama.cpp. A server that disappears after the model loaded still terminates the process — llama.cpp has no error path for a device lost mid-inference. - One
RpcServerper process. A secondstartwhile one runs throwsIllegalStateException. - Endpoints are IPv4 addresses or host names (
host:port); llama.cpp's RPC transport has no IPv6.RpcServerbinds to an IPv4 literal (127.0.0.1,0.0.0.0, an interface address). - Registered servers stay registered. llama.cpp keeps RPC devices in a process-wide registry
with no way to remove them; this library therefore gives every later load that does not ask for a
server an explicit device list without it, so a model loaded without
--rpcnever offloads to a server an earlier model used — the same holds for the multimodal projector,TextToSpeechandLlamaTrainer. An explicitsetDevices(...)/--deviceis never overridden. - Android needs the
android.permission.INTERNETpermission for RPC, even over loopback — thellama-androidAAR does not request it, so an app that wants RPC must declare it itself. RpcServer.startLocal(port, threads, cacheDir)enables upstream's tensor cache: a client that loads the same model again sends the large tensors only once.- Choose the served devices when the default is wrong.
RpcServer.startLocal(port, threads, cacheDir, Arrays.asList("CPU"))(or--device CPUon the command line; names as llama.cpp prints them, e.g.CUDA0,Vulkan1,MTL0) replaces the default of every accelerator. It matters because llama.cpp's RPC client treats every operation as supported by the remote device: a served GPU that cannot run one terminates the server process on the first graph that needs it. The paravirtual GPU of a macOS virtual machine is such a device — serveCPUthere.
A separate artifact, net.ladenthin:llama-langchain4j, adapts a LlamaModel to
LangChain4j's ChatModel, StreamingChatModel,
EmbeddingModel and ScoringModel interfaces in-process over JNI — no HTTP hop, no separate
server. It is a separate artifactId (not a classifier of the core) because LangChain4j 1.x
requires Java 17 while the core net.ladenthin:llama stays Java 8; keeping it separate avoids
forcing that floor on every core consumer. It ships and versions in lockstep with the core.
<dependency>
<groupId>net.ladenthin</groupId>
<artifactId>llama-langchain4j</artifactId>
<version>5.1.0</version>
</dependency>From 5.2.0 on, add the natives next to it — llama-platform or the natives jars you need (see
Choosing the natives jars); the core it depends on is classes only.
Each adapter borrows a LlamaModel you already loaded — it never loads or closes the native
model, so you manage its lifecycle (try-with-resources), and one LlamaModel can back several
adapters at once:
try (LlamaModel llama = new LlamaModel(new ModelParameters().setModel("models/qwen3-0.6b.gguf"))) {
ChatModel chat = new JllamaChatModel(llama);
String reply = chat.chat("Write a haiku about lazy senior devs.");
System.out.println(reply);
}| Adapter | LangChain4j interface | java-llama.cpp call |
|---|---|---|
JllamaChatModel |
ChatModel |
LlamaModel.chat(...) |
JllamaStreamingChatModel |
StreamingChatModel |
LlamaModel.generateChat(...) (token streaming) |
JllamaEmbeddingModel |
EmbeddingModel |
LlamaModel.embed(...) (model loaded with enableEmbedding()) |
JllamaScoringModel |
ScoringModel (re-ranking) |
LlamaModel.handleRerank(...) (model loaded with enableReranking()) |
See llama-langchain4j/README.md for streaming/embedding/re-ranking
examples and the current mapping limitations (tool calling, JSON mode, and multimodal input are
not yet forwarded).
Tip
Supported front ends — all on the same agent session (history, slash commands, approval mode):
| Front end | Start with | Where it runs |
|---|---|---|
| Terminal | (default) / --plain |
this console; --plain for piped, logged or line-only sessions |
| Browser | --web |
http://127.0.0.1:8787/?token=… — over SSH: ssh -L 8787:127.0.0.1:8787 user@server |
| IDE via ACP | --acp |
JetBrains IDEs and Zed natively, VS Code through an ACP extension |
llama-atmosphere-agent/ is a copy-and-run general-purpose agent on the JVM — Claude Code /
OpenCode reduced to the essentials, fully offline; it edits files and, with --allow-shell, runs any command on your machine
(docker, git, build tools) — built from Atmosphere's
built-in OpenAI-compatible agent runtime (streaming, tool loop, workspace file tools) driven
headless against this project's OpenAI-compatible server. It is a standalone Maven project (not a
reactor module) published to Maven Central at the core's version, so the quickest way needs only
JDK 21+ and JBang — it resolves the agent and the core with the natives of
every desktop platform:
jbang net.ladenthin:llama-atmosphere-agent:5.2.0 \
--model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --workspace /path/to/projectOr download it from a release (below, JDK 21+ only), or clone the repository and run it from that folder, which needs only JDK 21+ and Maven — the core jar from Maven Central ships the natives:
# get the folder and a tool-capable model (Qwen3-4B-Instruct-2507, 2.3 GB)
git clone --depth 1 https://github.com/bernardladenthin/java-llama.cpp.git
cd java-llama.cpp/llama-atmosphere-agent
curl -L --create-dirs -o models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf \
https://huggingface.co/unsloth/Qwen3-4B-Instruct-2507-GGUF/resolve/main/Qwen3-4B-Instruct-2507-Q4_K_M.gguf
# the agent with the model loaded in-process — a you> prompt appears (/clear, /exit)
mvn -q compile exec:java \
-Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ctx-size 16384 --workspace /path/to/project --allow-shell"Straight from Maven Central instead of cloning. The agent is published as a thin jar whose pom
names llama-platform, so JBang resolves it with the CPU natives of every
desktop platform and starts it; a GPU is one more natives jar on the same line (the module jars are
additive, see Choosing the natives jars):
jbang net.ladenthin:llama-atmosphere-agent:5.2.0 \
--model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ctx-size 16384 --workspace /path/to/project --allow-shell
# e.g. with Vulkan on Linux x86-64 (any current GPU driver), all layers offloaded
jbang --deps net.ladenthin:llama:5.2.0:vulkan-linux-x86-64 net.ladenthin:llama-atmosphere-agent:5.2.0 \
--model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ngl 99 --workspace /path/to/projectThe agent carries no core of its own (started without one it stops with NoClassDefFoundError: net/ladenthin/llama/LlamaModel); CI launches the published thin jar on one classpath with the
published core jars on every run (smoke-agent-linux) before anything is published.
Warning
--allow-shell lets the model run any command with your user's rights. By default every write and
every command is confirmed on the console ([y]es / [n]o / [a]uto); --auto turns that off. Use a
workspace you are willing to hand to the model.
In the REPL, /help lists the commands (/status, /tools, /mode manual|auto, /compact,
/clear, /exit); anything else goes to the model. A status line shows the approval mode and the
context used ([manual · ctx ~3.1k/16k · 9 tools · local-model]), and the answer is rendered with
headings, bullets and code spans.
Also in a browser or an editor. The same agent session has two more front ends. --web serves it
to a browser on 127.0.0.1:8787 (Atmosphere's own AI console on an embedded Jetty, a random access
token in the printed address; from another machine open an SSH tunnel, ssh -L 8787:127.0.0.1:8787 user@server, rather than binding to the network). --acp speaks the
Agent Client Protocol on stdin/stdout, so JetBrains IDEs and Zed —
and VS Code through an ACP extension — run it as their chat agent, with the editor's own permission
dialog for writes and commands:
jbang net.ladenthin:llama-atmosphere-agent:5.2.0 --model model.gguf --allow-shell --web
# JetBrains, ~/.jetbrains/acp.json: {"agent_servers": {"Local llama": {"command": "jbang",
# "args": ["net.ladenthin:llama-atmosphere-agent:5.2.0", "--acp", "--model", "/path/model.gguf"]}}}On Windows PowerShell quote the whole argument ("-Dexec.args=--model models\… --allow-shell"); for
the GPU add e.g. -Dllama.classifier=vulkan-windows-x86-64 and --ngl 99. The agent's
README walks through all of it step by step. Other ways to run it, e.g.
against a java-llama.cpp server that is already running (--jinja is required for tool calling;
see Running the server from Maven Central for a GPU
backend):
# 1. the server, from Maven Central
jbang --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64 net.ladenthin:llama:5.2.0 \
-m models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --jinja --port 8080
# 2. the agent, from the llama-atmosphere-agent/ folder — a you> prompt appears (/clear, /exit)
mvn -q compile exec:java \
-Dexec.args="--base-url http://127.0.0.1:8080/v1 --workspace /path/to/project --allow-shell"
# a single turn instead of the prompt loop
mvn -q compile exec:java \
-Dexec.args="--base-url http://127.0.0.1:8080/v1 --workspace /path/to/project --prompt 'Read the README and summarize it'"
# or without a separate server: load the GGUF in-process
mvn -q compile exec:java \
-Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ngl 99 --workspace /path/to/project"
# everything at once: shell access plus your own system prompt (replaces the built-in one)
mvn -q compile exec:java \
-Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ngl 99 --ctx-size 16384 --workspace /path/to/project --allow-shell --system 'You are a local assistant on this machine with full shell access. run_command executes any command line, including docker, git and build tools. When asked about the system, run a command instead of explaining it. Answer in the language of the user.'"The full streaming tool-calling loop (tools → delta.tool_calls → Java tool → role:"tool" result →
next turn, over several rounds) is verified on every PR against the real OpenAiCompatServer with
no model, and in CI against the Qwen2.5-1.5B tool model — both gate every publish, as does the
release-jar smoke above. See
llama-atmosphere-agent/README.md for the options and the verified
compatibility matrix.
There are two sets of parameters you can configure, ModelParameters and InferenceParameters. Both provide builder
classes to ease configuration. ModelParameters are once needed for loading a model, InferenceParameters are needed
for every inference task. All non-specified options have sensible defaults.
ModelParameters modelParams = new ModelParameters()
.setModel("/path/to/model.gguf")
.addLoraAdapter("/path/to/lora/adapter");
String grammar = """
root ::= (expr "=" term "\\n")+
expr ::= term ([-+*/] term)*
term ::= [0-9]""";
InferenceParameters inferParams = new InferenceParameters("")
.withGrammar(grammar)
.withTemperature(0.8f);
try (LlamaModel model = new LlamaModel(modelParams)) {
model.generate(inferParams);
}LlamaIterable (returned by model.generate(...) and model.generateChat(...))
implements Iterable<LlamaOutput> & AutoCloseable, so every mainstream reactive
library wraps it in a few lines without java-llama.cpp pulling in a runtime
reactive dependency.
Always wrap with the library's resource-management primitive — Flux.using,
Flowable.using, Kotlin use {}, etc. — so that subscription cancellation
flows into LlamaIterable.close() and from there into llama.cpp's native
cancelCompletion. A plain Flux.fromIterable(iterable) or for (x in iter)
loop will NOT close the iterable on cancel; the native task slot stays
occupied until the model is closed.
Flux<LlamaOutput> tokens = Flux.using(
() -> model.generate(params),
Flux::fromIterable,
LlamaIterable::close)
.subscribeOn(Schedulers.boundedElastic());Flowable<LlamaOutput> tokens = Flowable.using(
() -> model.generate(params),
Flowable::fromIterable,
LlamaIterable::close)
.subscribeOn(Schedulers.io());Ready-made: the optional net.ladenthin:llama-kotlin artifact ships
generateFlow/generateChatFlow extensions (close-on-cancellation included) plus suspend
wrappers whose coroutine cancellation is wired to the binding's cooperative CancellationToken:
model.generateChatFlow(params).flowOn(Dispatchers.IO).collect { print(it.text) }Hand-rolled equivalent (no extra dependency):
fun llama(model: LlamaModel, params: InferenceParameters) = flow {
model.generate(params).use { iterable ->
for (output in iterable) emit(output)
}
}.flowOn(Dispatchers.IO)The companion Android sample LLaMAndroid
demonstrates the flow { for (output in model.generate(params)) emit(output) }
shape against the upstream binding. Wrap the for loop in
.use { } if your collector may cancel mid-stream — otherwise the native task
slot will not be released until the model is closed.
val tokens: Source[LlamaOutput, NotUsed] = Source
.fromIterator(() => model.generate(params).iterator())
.async("blocking-io-dispatcher")Why no built-in Publisher? Earlier snapshots of this fork shipped a
hand-rolled LlamaModel.streamPublisher(...) returning a Reactive Streams
Publisher<LlamaOutput>. Since every reactive library bridges blocking
iterables in a few lines via its own resource-management primitive, the binding
now stays free of any reactive runtime dependency — pick whichever library your
app already uses. The pattern is verified end-to-end by
ReactorIntegrationTest in the test sources.
Per default, llama.cpp writes its log as text to stderr (0.00.035.060 I slot … once a model is
loaded): the server's own srv … / slot … lines and, from verbosity 4 on, the llama/ggml lines.
All of it can be intercepted via the static method
LlamaModel.setLogger(LogFormat, BiConsumer<LogLevel, String>): with a callback set, every line goes
to the callback instead of the console (a setLogFile file keeps receiving them). The callback
survives model loads, so set it before new LlamaModel(…) to capture the loading lines too.
LogFormat.TEXT hands over the bare message, LogFormat.JSON one JSON object per line. Passing
null as the callback restores the console output (always llama.cpp's own text format; the format
argument only matters with a callback). Logging can be disabled by passing an empty callback.
Messages arrive asynchronously from llama.cpp's log worker thread; replacing or removing the logger
flushes what is queued to the previous callback first. The verbosity threshold
(ModelParameters.setLogVerbosity(int), llama.cpp's -lv: 1 errors, 2 warnings, 3 info, 4 trace,
5 debug) applies before the callback: 2 keeps warnings and errors and silences the per-request
INFO lines, which is what a console application sharing the terminal with its own output wants.
// Re-direct log messages however you like (e.g. to a logging library)
LlamaModel.setLogger(LogFormat.TEXT, (level, message) -> System.out.println(level.name() + ": " + message));
// Back to llama.cpp's own console output (stderr)
LlamaModel.setLogger(LogFormat.TEXT, null);
// Disable logging by passing a no-op
LlamaModel.setLogger(LogFormat.TEXT, (level, message) -> {});The LogLevel enum values passed to the callback correspond to the native llama.cpp log levels:
| Value | Meaning |
|---|---|
DEBUG |
Verbose diagnostic output |
INFO |
Informational messages about model loading and inference |
WARN |
Non-fatal warnings |
ERROR |
Errors that may affect inference results |
Important
Minimum Android version: API 28 (Android 9.0 Pie). Devices running Android 8.1 (API 27) or earlier are not supported.
One dependency line in Android Studio — no submodule, no NDK build, no manual ProGuard rules:
dependencies {
implementation("net.ladenthin:llama-android:5.2.0")
// additionally, for Qualcomm Adreno GPUs (device must provide an OpenCL ICD) -- the OpenCL
// backend module; it depends on llama-android, so the line above may also be left out:
// implementation("net.ladenthin:llama-android-opencl:5.2.0")
// optional Kotlin coroutines facade (Flow streaming + suspend wrappers):
implementation("net.ladenthin:llama-kotlin:5.2.0")
}The AAR carries the full net.ladenthin:llama Java API, the CI-built native library with ggml's
CPU backend modules (one per ARM feature level, the best for the device loaded at start) for
arm64-v8a (devices) and x86_64 (Android Studio emulator, Chromebooks — app bundles
split per ABI so phones download only arm64), all 16 KB page-size compliant, consumer
R8/ProGuard rules (applied automatically), and a manifest minSdkVersion 28 that AGP
enforces against your app. The OpenCL AAR adds the libggml-opencl.so module, which the same
start loads when the device has an OpenCL ICD and skips otherwise. CI boots an x86_64 emulator
and runs real on-device inference against every AAR build.
Do not also depend on the desktop net.ladenthin:llama JAR in the same app — the AAR
already contains those classes, and the JAR would drag ~70 MB of desktop natives into your
APK. See llama-android/README.md and
llama-kotlin/README.md for details.
Runnable example app — "LLM Service". A minimal, KISS, fully-offline on-device chat app
— pick a GGUF from the file system, then chat with it, tokens streaming into a Jetpack
Compose UI, with a 13-language flag picker and private local save/load — lives in
android-llmservice/ (net.ladenthin.android.llmservice).
It builds with plain Gradle/AGP (no Android Studio required), produces a Play-shaped signed
.aab, and is validated in CI by a real on-device emulator UI test. See its README for the
build, signing/Play, and testing walkthrough.
Use this only if you need to patch the native layer or build for an ABI this project does not ship.
- Add java-llama.cpp as a submodule in your an droid
appproject directory
git submodule add https://github.com/bernardladenthin/java-llama.cpp - Declare the library as a source in your build.gradle
android {
val jllamaLib = file("java-llama.cpp")
// Execute "mvn compile" in the llama/ core module if its target/ doesn't exist
// (the repository root is the Maven reactor aggregator; the native core lives in llama/).
if (!file("$jllamaLib/llama/target").exists()) {
exec {
commandLine = listOf("mvn", "compile")
workingDir = file("java-llama.cpp/llama/")
}
}
...
defaultConfig {
...
externalNativeBuild {
cmake {
// Add an flags if needed
cppFlags += ""
arguments += ""
}
}
}
// Declare c++ sources
externalNativeBuild {
cmake {
path = file("$jllamaLib/CMakeLists.txt")
version = "3.22.1"
}
}
// Declare java sources
sourceSets {
named("main") {
// Add source directory for java-llama.cpp
java.srcDir("$jllamaLib/src/main/java")
}
}
}- Exclude
net.ladenthin.llamain proguard-rules.pro
keep class net.ladenthin.llama.** { *; }
Open work items live in TODO.md.
- Expand PIT mutation-testing scope. PIT is wired in
pom.xmland runs on every CI build (in thetest-java-linux-x86_64job) with<mutationThreshold>100</mutationThreshold>.<targetClasses>currently coversnet.ladenthin.llama.value.*,exception.*,args.*and fourjsonparsers (295 mutations, 100% killed, hermetic — no model or fixture needed); widen it incrementally as additional classes reach mutation-test parity. Final target:<param>net.ladenthin.llama.*</param>matching the streambuffer pattern.
Forward-looking ideas being tracked for this fork:
- Adopt feature ideas from the Kotlin Llama Stack client. Candidates (multimodal image input, typed chat messages, async API, batch inference, typed usage/timings) are inventoried with effort estimates in
docs/feature-investigation-llama-stack-client-kotlin.md, derived fromogx-ai/llama-stack-client-kotlin. - Ship a directly Android-capable artifact — DONE.
net.ladenthin:llama-android/llama-android-opencl(AAR, arm64-v8a, minSdk 28, consumer ProGuard rules, 16 KB page-size compliant) plus the optionalnet.ladenthin:llama-kotlincoroutines façade ship from this repo — see Importing in Android. Typed image input for VLMs is covered byContentPart.imageBytes(...)/imageFile(...)(see the multimodal section), so downstream Android projects can drop their dependency onogx-ai/llama-stack-client-kotlinentirely. A dedicated KISS example app — "LLM Service" (SAF model picker + Compose streaming chat, 13-language flag picker, private local save/load, plain Gradle/AGP, signed.aab, on-device emulator UI test) — ships inandroid-llmservice/. - Resolve all upstream
kherud/java-llama.cppopen issues. All 37 open issues at fork time are catalogued with per-issue verdicts indocs/history/49be664_open_issues.md; fixes land in this fork as they are completed. Vision inputs (issues #103 and #34) are now wired end to end through blocking, typed, streaming, and OpenAI-compatible request surfaces.
A crash of this shape was reported against early releases:
EXCEPTION_ACCESS_VIOLATION (0xc0000005) at pc=0x00007ffa8f4b2f58
C [msvcp140.dll+0x12f58]
This cannot come from this library any more, and the old advice here — deleting
msvcp140.dll from your JDK — is obsolete. Do not do it. The Windows natives link the C++
standard library and vcruntime statically (hybrid CRT), so they do not import msvcp140.dll,
vcruntime140.dll or vcruntime140_1.dll at all; the only non-OS import left is the Universal
CRT, which Windows itself provides and keeps updated. Measured on the current build, the whole
import table of jllama.dll is KERNEL32, ADVAPI32, SHELL32, WS2_32 and the
api-ms-win-crt-* forwarders. CI enforces that per release
(.github/buildcheck/nativedeps.py holds every Windows CPU directory to an exact allowlist).
The report dates from before that: the troubleshooting note was written on 2026-04-04 and the
switch from the DLL runtime (/MD) to a static one landed on 2026-05-13, so releases from
5.0.0 on are unaffected. An older JDK does ship its own outdated msvcp140.dll next to
java.exe, and Windows resolves an import from the already-loaded module list before searching
any directory -- which is why a library that did import it got the JDK's copy. Removing the
dependency fixes that by construction, where moving files around could not.
If you still see a crash inside msvcp140.dll, it originates in another native library loaded
into the same JVM, not in jllama.dll -- check the rest of the hs_err frame list. Please open an
issue with that file rather than editing your JDK installation.
The native libraries are not Authenticode-signed (upstream llama.cpp ships its DLLs unsigned
too), and LlamaLoader extracts them into the temp directory on first use. A client whose
application-control policy (WDAC, AppLocker) blocks unsigned DLLs, or any DLL under %TEMP%,
refuses the load with UnsatisfiedLinkError. The way out is a directory the policy allows: put the
contents of the natives jar's net/ladenthin/llama/Windows/x86_64/cpu/ directory there and start
the JVM with -Dnet.ladenthin.llama.lib.path=<that directory>, which loads from it without
extracting anything. (The project's release-signing key is an OpenPGP key, which signs the jars with detached .asc
files; Windows does not look at those, so the DLLs would need a separate code-signing
certificate.)
llama/pom.xml passes -XDaddTypeAnnotationsToSymbol=true to javac (NullAway's JSpecify mode
needs it below JDK 22). Oracle JDK 21 rejects that flag: mvn compile dies in an Error Prone
IllegalStateException naming it (measured on 21.0.9). Eclipse Temurin 21, which CI uses, compiles
the same tree; mvn -version shows which JDK Maven runs on.
⚠️ DO NOT UPGRADE jqwik past 1.9.3. jqwik 1.10.0 added an anti-AI prompt-injection string to test stdout; the 1.10.1 user guide states the library "is not meant to be used by any 'AI' coding agents at all." 1.9.3 is the last pre-disclosure release and is the pinned version. SeeCLAUDE.mdsection "jqwik prompt-injection in test output" for the full context. Dependabot is configured to ignore allnet.jqwikupdates (every version, including patches) — see theignorerule in.github/dependabot.yml.
Bindings / wrappers
- kherud/java-llama.cpp — the upstream Java binding this project was forked from (see the note at the top of this README); development continues independently here, with the fork-time upstream issues catalogued in
docs/history/49be664_open_issues.md. - llamacpp4j — alternative Java/JNI binding to llama.cpp (SWIG-generated facade); pre-GGUF, dormant since 2023 but historically the other Java JNI option.
- llama-cpp-python — the Python llama.cpp binding; the de-facto feature benchmark among llama.cpp bindings (server mode, multimodal, speculative decoding).
- LLamaSharp — C#/.NET llama.cpp binding with per-backend runtime packages (CPU/CUDA/Vulkan/Metal), the .NET analogue of this project's classifier matrix.
- node-llama-cpp — Node.js/TypeScript llama.cpp binding (prebuilt binaries, JSON-schema-constrained output, function calling).
- LLaMAndroid — Android app demonstrating usage of llama.cpp bindings.
- llama-stack-client-kotlin — Kotlin client for the Llama Stack API with an ExecuTorch-backed local-inference path (the
llama-androidAAR +llama-kotlinfaçade cover the same on-device ground natively). - llama.cpp-android-tutorial — Step-by-step tutorial for running llama.cpp on Android.
Other local inference stacks (no llama.cpp JVM binding)
- Ollama — llama.cpp-based local model runner with its own HTTP API and model registry. This project's OpenAI-compatible server implements the Ollama-native API surface (
/api/version,/api/tags,/api/show,/api/chat,/api/generate), so Ollama-speaking clients (e.g. VS Code Copilot's Ollama provider) work against an in-process jllama model. - ExecuTorch — PyTorch's on-device inference runtime (
.ptemodels, XNNPACK/NPU delegates); the engine behindllama-stack-client-kotlin's local mode and the main non-llama.cpp alternative for Android on-device inference (GGUF is not supported there — different model format ecosystem).
Pure-Java single-model inference (no JNI / no llama.cpp) — Alfonso² Peterssen's *.java family of standalone, dependency-free Java inference runtimes, one per model architecture. Useful when JNI is unavailable (e.g. some sandboxes / GraalVM native-image scenarios) or when you want a single jar with no native side at all. Different design point from this project, which prioritises GGUF compatibility and llama.cpp performance via JNI.
- llama3.java — Llama 3 / 3.1 / 3.2 inference.
- gemma4.java — Gemma 4 (and earlier Gemma 2/3) inference.
- gptoss.java — GPT-OSS architecture inference.
- qwen35.java — Qwen 3.5 inference.
- nemotron3.java — NVIDIA Nemotron-3 inference.
Pure-Java inference engines (no JNI / no llama.cpp)
- Jlama — a full pure-Java LLM inference engine for the JVM (multiple model architectures, quantization, and distributed inference) built on the Java Vector API. A no-native alternative to the JNI approach here; different design point (pure JVM portability vs. GGUF compatibility and llama.cpp performance via JNI).
Frameworks / orchestration
- LangChain4j — LLM-application framework for Java (chat, embeddings, RAG, tool calling, agents) over a unified provider API. This project ships a first-class in-process integration — see the
llama-langchain4jmodule — so a llama.cpp model plugs straight into LangChain4j'sChatModel/StreamingChatModel/EmbeddingModel/ScoringModelwithout an HTTP hop.