Skip to content

About

Java Bindings for llama.cpp - A Port of Facebook's LLaMA model in C/C++

Resources

Code of conduct

Contributing

Security policy

Stars

16 stars

Watchers

0 watching

Forks

 
 

Latest commit

 

History

2,589 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

java-llama.cpp

Note

No IDE or local setup required. This repository is optimized for fully AI-assisted development using Claude Code. No local toolchain, no IDE, nothing to install — everything works completely through Claude.

AI:
Claude

Build:
Java 8+
Platform
llama.cpp b11556
JPMS
JUnit
JSpecify
NullAway
Checker Framework
Error Prone
Maven Enforcer
Lombok
jqwik
ArchUnit
SpotBugs
jcstress
Lincheck
vmlens
JMH
Publish
CodeQL

Build cache:
Build cache by Depot

Coverage:
Coverage Status
codecov
JaCoCo
PIT Mutation

Quality:
Quality Gate
Code Smells
Security Rating

Security:

Known Vulnerabilities
FOSSA Status
Dependencies
OSV-Scanner

Package:
Maven Central
Snapshot
Release Date
Last Commit

License:
License

Community:
OpenSSF Best Practices
Contribute with Gitpod
OpenSSF Scorecard
Dependabot
Conventional Commits
Keep a Changelog
SemVer
REUSE
Maintained?
Issues
Pull Requests
GitHub Stars
Treeware
Stand With Ukraine

Java Bindings for llama.cpp

Forked from kherud/java-llama.cpp: many thanks to @kherud for the great work!

Inference of Meta's LLaMA model (and others) in pure C/C++.

You are welcome to contribute

  1. Features
  2. Quick Start
    2.1 No Setup required
    2.2 Setup required
  3. Documentation
    3.1 Example
    3.2 Inference
    3.3 Chat Completion
    3.4 Infilling
    3.5 Embeddings & Reranking
    3.6 Raw JSON Endpoints
    3.7 Local agent: terminal, browser, IDE via ACP
  4. Android
  5. Feature Ideas

Features

  • Text completion (blocking and streaming) with full control over sampling parameters.
  • OpenAI-compatible chat completion with automatic chat-template application, including streaming and tool/function calling support via the upstream server.
  • Embeddings (single and native-batched via embed(Collection<String>)) and reranking for retrieval pipelines.
  • Runtime LoRA adapter control — list the loaded adapters and change their scales at runtime without reloading the model (getLoraAdapters() / setLoraAdapters(Map)), the typed counterpart of the upstream GET/POST /lora-adapters endpoints.
  • Text-to-speech (TextToSpeech) over llama.cpp's Qwen3-TTS pipeline (mtmd_helper::gen_audio), returning WAV audio.
  • In-JVM GGUF quantization (LlamaQuantizer) over llama.cpp's llama_model_quantize — convert a GGUF to another quantization scheme without shelling out to llama-quantize.
  • Kolibri-1 (Aleph Alpha's German/English reasoning MoE, architecture kolibri1) before upstream llama.cpp supports it, through a carried patch that loads the GGUFs of both community converters. See Kolibri-1.
  • Decision models (handleSystemOne) — llama.cpp's TypeSafe-compatible /v1/systemone API: typed choice / score / noul questions about a state, answered with probabilities in one forward pass, no token generated (laya, julia-1, lev, openjev, kev and the later decision models). See Decision models.
  • Infilling (fill-in-the-middle) for code models.
  • Tokenize / detokenize and JSON-schema → grammar conversion.
  • Raw JSON endpoint handlers mirroring the upstream llama.cpp HTTP server (/completions, /v1/completions, /embeddings, /infill, /tokenize, /detokenize).
  • Two runnable HTTP server modes, one entry point. The classes jar's Main-Class is ServerLauncher, which dispatches on the --jllama-openai-compat flag. Without it, jbang --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64 net.ladenthin:llama:5.2.0 -m model.gguf --port 8080 runs the full upstream llama.cpp server (embedded WebUI, every llama-server flag forwarded) hosted inside libjllama over JNI — no separate llama-server.exe. With it, the same line plus --jllama-openai-compat --model model.gguf --port 8080 runs the Java-transport, zero-extra-dependency OpenAI-compatible server (OpenAiCompatServer, streaming SSE) instead. Both are also runnable directly by class name via java -cp … net.ladenthin.llama.server.{NativeServer,OpenAiCompatServer}; see Running the server from Maven Central.
  • Model metadata access (getModelMeta()) and server management (metrics, slot save/restore, runtime thread reconfiguration).
  • Conversation checkpoints — Session.checkpoint(...) / rewind(...) / fork(...) branch and roll back a chat (KV-cache slot save/restore + transcript snapshot) without re-prefilling.
  • GGUF metadata inspection without loading the model (GgufInspector — pure Java, reads header + key/value table only, big-endian aware).
  • Distributed inference over RPC — offload a model's layers to llama.cpp RPC servers on other machines (ModelParameters.setRpcServers(...) / --rpc host:port), and serve this machine's devices to them with RpcServer (the in-JVM rpc-server). See Distributed inference over RPC.
  • Local agent (llama-atmosphere-agent, release asset, JDK 21+) — a fully offline agent on top of this library that reads and edits files and, if allowed, runs commands. One agent session, three ways to use it: a terminal (full console or line-oriented --plain, fine over SSH/PuTTY), a browser (--web, token-protected, loopback by default — reach it from elsewhere through an SSH tunnel), and IDEs over the Agent Client Protocol (--acp: JetBrains IDEs and Zed natively, VS Code through an ACP extension). See Local agent.
  • Multi-model router mode (--models-dir + per-request model selection, managed via the typed RouterClient) and attach mode (NativeServer(LlamaModel, ...) serves an already-loaded model over the full upstream HTTP frontend — one copy of the weights).
  • HTTPS built in — model downloads from https:// URLs (--model-url, -hf) and a TLS-capable embedded server (--ssl-key-file/--ssl-cert-file) on every desktop platform, with BoringSSL linked statically: no OpenSSL to install, and no dependency on one (Android and s390x ship without SSL).
  • Pre-built native binaries for Linux (x86-64, aarch64, s390x), macOS (arm64, Metal included), Windows (x86-64, x86, arm64) and Android (arm64, x86-64), plus GPU backends (CUDA, Vulkan, OpenCL, ROCm/HIP, SYCL, OpenVINO) — one natives jar each, all loadable side by side with automatic CPU fallback; see Choosing the natives jars. Android additionally ships as the llama-android AAR with the optional llama-kotlin coroutines façade.

Quick Start

Access this library via Maven (released versions on Maven Central). llama-platform brings the classes plus the CPU natives of every desktop platform:

<dependency>
    <groupId>net.ladenthin</groupId>
    <artifactId>llama-platform</artifactId>
    <version>5.2.0</version>
    <type>pom</type>
</dependency>

(<type>pom</type> is required in Maven: llama-platform is a dependency list, not a jar. In Gradle it is just implementation("net.ladenthin:llama-platform:5.2.0").)

Note

This layout starts with 5.2.0. Up to 5.1.0, net.ladenthin:llama was one jar that carried the CPU natives of every platform, and each GPU classifier was a complete replacement for it.

There are multiple examples.

Try it with JBang, no project needed

examples/jbang/Chat.java is a one-file console chat that JBang runs straight from the repository, with any JDK 8+ and a GGUF of an instruction-tuned model:

jbang https://github.com/bernardladenthin/java-llama.cpp/blob/main/examples/jbang/Chat.java model.gguf

Its //DEPS lines name the classes jar and the CPU natives jar of every desktop platform (the jars llama-platform names; JBang treats a pom dependency as a BOM and puts nothing of it on the classpath), and the loader picks this machine's. Copy the file as a starting point; a GPU backend is one more //DEPS line (see Choosing the natives jars).

Snapshot builds

Every push to main publishes a snapshot to the Sonatype Central snapshot repository.

To use the latest snapshot, add the repository and dependency to your pom.xml:

<repositories>
  <repository>
    <id>sonatype-snapshots</id>
    <url>https://central.sonatype.com/repository/maven-snapshots/</url>
    <snapshots><enabled>true</enabled></snapshots>
    <releases><enabled>false</enabled></releases>
  </repository>
</repositories>

<dependency>
    <groupId>net.ladenthin</groupId>
    <artifactId>llama-platform</artifactId>
    <version>5.2.0-SNAPSHOT</version>
    <type>pom</type>
</dependency>

No credentials are required — the repository is publicly readable.

No Setup required

We support CPU inference for the following platforms out of the box:

  • Linux x86-64, aarch64, s390x
  • macOS aarch64 (Apple silicon, with Metal)
  • Windows x86-64, aarch64
  • Android aarch64, x86-64 (see Importing in Android)

If any of these match your platform, you can include the Maven dependency and get started.

Choosing the natives jars

net.ladenthin:llama is the Java classes only. The native code ships as separate jars of the same artifact, selected by a Maven <classifier> of the form <backend>-<os>-<arch>, of two kinds:

  • One library jar per platform (cpu-<os>-<arch>, metal-macos-aarch64): libjllama with llama.cpp, ggml's shared libraries and the CPU backend modules (one per instruction-set level, the best one for the running CPU is loaded at start). Exactly one is loaded, the one of the running platform; llama-platform collects them for every desktop platform.
  • GPU module jars (cuda13-…, rocm-…, sycl-…, vulkan-…, opencl-…, openvino-…): each holds only ggml's backend module for that GPU (libggml-cuda.so, ggml-vulkan.dll, …). They are additive: the loader puts every GPU module jar of the platform next to the library, ggml loads each module whose vendor runtime is installed and registers its devices, and a module whose runtime is missing is skipped with a log line — the model then runs on the CPU, so a GPU jar on a machine without that GPU just works. Several GPU jars together are fine (llama.cpp lists every device they bring, the same GPU seen through two backends once). A module jar alone is not enough: it needs the library jar of its platform, from the same release — the loader compares the build stamp (jllama-build.txt, the llama.cpp tag) of every module jar with the library's and refuses a mismatch before loading anything, which the module's own export check at load time cannot see.

Each jar holds exactly one directory, net/ladenthin/llama/<OS>/<ARCH>/<backend>/, so any combination can share one classpath. The start-up log names the library, the modules put in place and the backends and devices ggml registered.

llama-platform is the classes jar plus the library jars of every desktop platform. For GPU acceleration, add the module jar for your GPU next to it:

<dependency>
    <groupId>net.ladenthin</groupId>
    <artifactId>llama-platform</artifactId>
    <version>5.2.0</version>
    <type>pom</type>
</dependency>
<!-- Add any natives jar from the table below. Example shown: CUDA 13 on Linux x86-64. -->
<dependency>
    <groupId>net.ladenthin</groupId>
    <artifactId>llama</artifactId>
    <version>5.2.0</version>
    <classifier>cuda13-linux-x86-64</classifier>
</dependency>

To ship less, depend on net.ladenthin:llama (the classes) plus only the natives jars of the platforms you target, e.g. cpu-linux-x86-64.

Classifier Backend Target platform Runtime requirement
cpu-linux-x86-64 CPU Linux x86-64 A JDK 8+ JVM; glibc ≥ 2.28 (manylinux_2_28: RHEL 8, Ubuntu 20.04, Debian 10 and later). Ships one CPU backend module per instruction-set level (x86-64 baseline, SSE4.2, AVX, AVX2, AVX-512, AVX-VNNI, AMX), of which ggml loads the best for the running CPU at start-up — a CPU without AVX2 works, and an AVX-512/AMX machine uses its kernels.
cpu-linux-aarch64 CPU Linux aarch64 glibc ≥ 2.28 (RHEL 8, Ubuntu 20.04, Debian 10, Amazon Linux 2023 and later) — built in the manylinux_2_28 image on an arm64 runner. Ships one CPU backend module per ARM feature level (armv8.0 up to armv9.2: dotprod, fp16, SVE, i8mm, SVE2, SME), chosen at start-up like the x86-64 ones.
cpu-linux-s390x CPU Linux s390x (IBM Z, big-endian) A JDK 8+ JVM.
cpu-windows-x86-64 CPU Windows x86-64 A JDK 8+ JVM; Windows 10 or newer. Ships one CPU backend module per instruction-set level (as the Linux x86-64 jar: x86-64 baseline, SSE4.2, AVX, AVX2, AVX-512, AVX-VNNI, AMX), of which ggml loads the best for the running CPU at start-up. Built with clang, hybrid CRT: the STL and vcruntime are static (no msvcp140.dll / vcruntime140.dll, no VC++ redistributable), the Universal CRT is the OS one (ucrtbase.dll), so Microsoft's UCRT security updates apply.
cpu-windows-aarch64 CPU Windows on ARM (Snapdragon X / Surface) A JDK 8+ JVM; Windows 10 or newer. Built natively on windows-11-arm with clang-cl, same hybrid CRT as the x86-64 jar; one CPU backend module (ggml builds no ARM variant set for Windows), so the Vulkan and OpenCL module jars can join it.
metal-macos-aarch64 Metal + CPU macOS aarch64 (Apple silicon) A JDK 8+ JVM.
cpu-android-aarch64 / cpu-android-x86-64 CPU Android For Android use the llama-android AAR; these jars are the same libraries (with the CPU modules, aarch64 one per ARM feature level) for other Android JVM setups.
cuda13-linux-x86-64 CUDA 13 Linux x86-64 with NVIDIA GPU NVIDIA driver + CUDA 13 runtime libraries (libcudart.so.13, libcublas.so.13).
cuda13-windows-x86-64 CUDA 13 Windows x86-64 with NVIDIA GPU NVIDIA driver + CUDA 13 Toolkit (cudart64_13.dll, cublas64_13.dll, cublasLt64_13.dll on PATH).
vulkan-linux-x86-64 Vulkan Linux x86-64 with a Vulkan 1.2+ GPU (NVIDIA / AMD / Intel) A Vulkan runtime (libvulkan.so.1), which current GPU drivers install. The most portable Linux GPU option. glibc ≈ 2.39 (built on ubuntu-latest).
vulkan-linux-aarch64 Vulkan Linux aarch64 with a Vulkan 1.2+ GPU A Vulkan runtime (libvulkan.so.1). glibc ≥ 2.39.
vulkan-windows-x86-64 Vulkan Windows x86-64 with a Vulkan 1.2+ GPU A Vulkan runtime (vulkan-1.dll), which current GPU drivers install. The most portable Windows GPU option.
vulkan-windows-aarch64 Vulkan Windows on ARM (Snapdragon X) with a Vulkan 1.2+ GPU A Vulkan runtime (vulkan-1.dll), which current GPU drivers install. Built natively on windows-11-arm with clang-cl, like upstream's Windows arm64 Vulkan release.
opencl-windows-x86-64 OpenCL Windows x86-64 with an OpenCL 2.0+ GPU A vendor OpenCL ICD (OpenCL.dll). The GGML OpenCL backend is Adreno-tuned; on desktop GPUs CUDA or Vulkan are better supported.
opencl-windows-aarch64 OpenCL (Adreno) Windows on ARM (Snapdragon X) The Adreno driver's OpenCL ICD (OpenCL.dll).
opencl-android-aarch64 OpenCL (Adreno) Android aarch64 with Adreno GPU A device OpenCL ICD (libOpenCL.so); see also the llama-android-opencl AAR.
rocm-linux-x86-64 ROCm / HIP Linux x86-64 with AMD GPU An AMD ROCm 10 runtime (libamdhip64.so, librocblas.so, libhipblas.so) — built against ROCm 10.0 (TheRock), like upstream llama.cpp; every GPU TheRock builds for Linux, Instinct included (gfx900/gfx906/gfx90c/gfx1153 best effort — built, but not release-ready in ROCm 10).
rocm-windows-x86-64 ROCm / HIP Windows x86-64 with AMD GPU The AMD ROCm 10 runtime DLLs (amdhip64.dll, rocblas.dll, hipblas.dll) on PATH; every Radeon target TheRock builds for Windows, gfx900 through RDNA4 (gfx900/gfx906/gfx90c/gfx1153 best effort).
sycl-linux-x86-64 SYCL (Intel oneAPI, fp16) Linux x86-64 with Intel GPU (Arc / iGPU) An Intel oneAPI / Level-Zero runtime. fp16 accumulation, as upstream's SYCL release builds.
sycl-windows-x86-64 SYCL (Intel oneAPI, fp16) Windows x86-64 with Intel GPU (Arc / iGPU) The Intel oneAPI / Level-Zero runtime DLLs on PATH.
openvino-linux-x86-64 OpenVINO Linux x86-64 (Intel GPU / NPU / CPU) An Intel OpenVINO runtime.
openvino-windows-x86-64 OpenVINO Windows x86-64 (Intel GPU / NPU / CPU) The Intel OpenVINO runtime DLLs on PATH.

Note

No vendor runtime is bundled; it comes from the GPU driver or toolkit on the host. The GPU module jars are validated build-only in CI (GitHub runners have no GPU): the smoke jobs put every GPU module of the platform next to the library and see ggml skip each one for lack of a runtime, so end-to-end GPU inference is verified locally / on self-hosted hardware. A module whose runtime is installed but finds no device (e.g. the CUDA toolkit without an NVIDIA GPU) simply registers none. -Dnet.ladenthin.llama.backend=<a>,<b> restricts the modules put in place to the named ones (cpu for none) and fails loud when a named jar is not on the classpath; unset, every GPU module jar on the classpath is used.

Note

On the module path each natives jar is an automatic module (net.ladenthin.llama.natives.<classifier>, with _ for -) that nothing requires, so resolve them with --add-modules ALL-MODULE-PATH (or keep the natives jars on the classpath).

Note

Android armeabi-v7a (32-bit ARM) is not published. Only 64-bit Android binaries are shipped: aarch64 (devices) and x86_64 (emulators, Chromebooks, x86-64 Android hardware) as the cpu-android-* natives jars and in the llama-android AAR, plus aarch64 as opencl-android-aarch64. 32-bit Android devices are unsupported by the released artifacts.

The minimum required Android version is API 28 (Android 9.0 Pie). Devices running Android 8.1 (API 27) or earlier are not supported.

Running the server from Maven Central

There is no all-in-one jar to download: the natives are modular, and every combination of them is a classpath Maven resolves. The classes jar's Main-Class is ServerLauncher, so the embedded server starts straight from the coordinates with JBang — the library jar of your platform, plus any GPU module jars, as --deps:

# CPU only (Linux x86-64; cpu-windows-x86-64, metal-macos-aarch64, ... for the others)
jbang --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64 net.ladenthin:llama:5.2.0 -m model.gguf --port 8080
# with a GPU backend -- several may be named, ggml uses the ones whose runtime is installed
jbang --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64,net.ladenthin:llama:5.2.0:vulkan-linux-x86-64 \
    net.ladenthin:llama:5.2.0 -m model.gguf --port 8080 -ngl 99

Every llama-server flag works, the WebUI included; add --jllama-openai-compat to run the Java OpenAI-compatible server instead. With Maven, from a checkout, examples/server/pom.xml depends on llama-platform and carries one profile per GPU module jar, named after its classifier:

mvn -f examples/server/pom.xml exec:java -Dexec.args="-m model.gguf --port 8080"                      # CPU
mvn -f examples/server/pom.xml -P cuda13-linux-x86-64 exec:java -Dexec.args="-m model.gguf -ngl 99"   # + CUDA
# a local jar-with-dependencies of exactly that combination, for a machine without Maven
mvn -f examples/server/pom.xml -P assembly,cuda13-linux-x86-64 package
java -jar examples/server/target/llama-server-5.2.0-jar-with-dependencies.jar -m model.gguf -ngl 99

The GitHub Release of a version attaches the Maven artifacts themselves — the classes jar, every natives jar, the poms, sources and javadoc, each with the detached GPG .asc that signed it for Central — and nothing else; a jar of every backend of every platform would be the opposite of modular natives.

Setup required

If none of the above listed platforms matches yours, or you want a GPU backend no module jar exists for (e.g. ROCm on Linux aarch64), you have to compile the library yourself.

This consists of two steps: 1) Compiling the libraries and 2) putting them in the right location.

Library Compilation

First, have a look at llama.cpp to know which build arguments to use (e.g. for CUDA support). Any build option of llama.cpp works equivalently for this project. You then have to run the following commands in the llama/ module directory (the native core lives there; the repository root is just the Maven reactor aggregator):

cd llama       # the native core module
mvn compile    # don't forget this line
cmake -B build # add any other arguments for your backend, e.g. -DGGML_CUDA=ON
cmake --build build --config Release

Tip

Use -DLLAMA_CURL=ON to download models via Java code using ModelParameters#setModelUrl(String).

The library is put in a directory matching your platform and backend, which appears in the cmake output. For example:

-- Backend 'cpu' - installing files to /java-llama.cpp/llama/src/main/natives/net/ladenthin/llama/Linux/x86_64/cpu

mvn test puts that directory on the test classpath; mvn -P natives package turns it into a natives jar (in CI, together with every other platform's).

Library Location

This project has to load a single shared library jllama.

Note, that the file name varies between operating systems, e.g., jllama.dll on Windows, jllama.so on Linux, and jllama.dylib on macOS.

The application will search in the following order in the following locations:

  • In net.ladenthin.llama.lib.path: Use this option if you want a custom location for your shared libraries, i.e., set VM option -Dnet.ladenthin.llama.lib.path=/path/to/directory.
  • In java.library.path: These are predefined locations for each OS, e.g., /usr/java/packages/lib:/usr/lib64:/lib64:/lib:/usr/lib on Linux. You can find out the locations using System.out.println(System.getProperty("java.library.path")). Use this option if you want to install the shared libraries as system libraries.
  • From the natives jars on the classpath: every backend directory found for your platform, in the order described in Choosing the natives jars.

System Properties Reference

Every net.ladenthin.llama.* system property recognised by the library, deep-scanned from the source. Runtime properties are resolved through LlamaSystemProperties; test-only properties are declared in the test sources (TestConstants) and consumed by individual test classes.

Property Default Scope Consumer Description
net.ladenthin.llama.lib.path unset (falls back to java.library.path) runtime LlamaLoader Directory containing the native jllama shared library. Checked first, before java.library.path. Set with -Dnet.ladenthin.llama.lib.path=/path/to/dir.
net.ladenthin.llama.tmpdir unset (falls back to java.io.tmpdir) runtime LlamaLoader Directory the natives are extracted into, one subdirectory per backend and build (jllama-backend-<backend>-<key>). It is kept across runs: the next start of the same build compares each file with the jar and copies nothing; a directory of a build that is no longer on the classpath is removed by a later start once it is 10 minutes old. Point it elsewhere when the default temp directory is on a slow or policy-restricted volume.
net.ladenthin.llama.osinfo.architecture unset (uses os.arch) runtime OSInfo Override for the architecture string used to locate the bundled library inside the JAR. Useful when os.arch reports an unexpected value (e.g. inside dockcross / chrooted environments).
net.ladenthin.llama.backend unset (every GPU module jar on the classpath is put next to the library) runtime LlamaLoader A comma-separated filter over the GPU module jars (e.g. vulkan, cuda13,vulkan): only the named modules are put in place, cpu names none; a named module whose jar is not on the classpath, or an unknown name, fails the load. The library jar is never a choice — it is the one of the running platform. See Choosing the natives jars.
net.ladenthin.llama.test.ngl 43 for the general suite; 0 for ToolCallingIntegrationTest test Model-backed integration tests Number of GPU layers used during testing. Pin to 0 on CPU-only hosts: mvn test -Dnet.ladenthin.llama.test.ngl=0. The tool test also selects device none at zero layers so Metal/CUDA is not initialized.
net.ladenthin.llama.text.model models/codellama-7b.Q2_K.gguf (tests self-skip if missing) test LlamaModelTest and most model-backed tests (TestConstants.MODEL_PATH) Path to the main text-generation GGUF. Most tests only need an instruct-capable model; a few assertions are tuned to the CI model.
net.ladenthin.llama.draft.model models/AMD-Llama-135m-code.Q2_K.gguf (tests self-skip if missing) test Speculative-decoding tests and the tests that need a small, fast model (TestConstants.DRAFT_MODEL_PATH) Path to the draft GGUF.
net.ladenthin.llama.reasoning.model models/Qwen3-0.6B-Q4_K_M.gguf (tests self-skip if missing) test Reasoning-budget tests (TestConstants.REASONING_MODEL_PATH) Path to a thinking model (one that emits a thinking block).
net.ladenthin.llama.rerank.model models/jina-reranker-v1-tiny-en-Q4_0.gguf (tests self-skip if missing) test Reranking tests (TestConstants.RERANKING_MODEL_PATH) Path to a reranker GGUF (must load with enableReranking()).
net.ladenthin.llama.tool.model models/Qwen2.5-1.5B-Instruct-Q4_K_M.gguf (test self-skips if missing) test ToolCallingIntegrationTest Path to a tool-capable GGUF used to verify required blocking and streaming tool calls. The default matches the Qwen2.5 model in upstream llama.cpp's tool-call test matrix.
net.ladenthin.llama.nomic.path models/nomic-embed-text-v1.5.f16.gguf (test self-skips if missing) test LlamaEmbeddingsTest#testNomicEmbedLoads Path to a Nomic embedding model (nomic-embed-text-v1.5.f16.gguf or a compatible BERT-family encoder). Regression test for upstream issue #98 (BERT-encoder result_output assertion).
net.ladenthin.llama.vision.model models/SmolVLM-500M-Instruct-Q8_0.gguf (test self-skips if missing) test MultimodalIntegrationTest Path to a vision-capable model GGUF. Any vision-capable GGUF works; CI default is SmolVLM-500M-Instruct-Q8_0.gguf.
net.ladenthin.llama.vision.mmproj models/mmproj-SmolVLM-500M-Instruct-Q8_0.gguf (test self-skips if missing) test MultimodalIntegrationTest Matching mmproj GGUF for the vision model.
net.ladenthin.llama.vision.image llama/src/test/resources/images/test-image.jpg (a CC-BY-4.0 / MIT-granted photo committed to the repo) test MultimodalIntegrationTest Visual prompt image. Any png/jpeg/webp/gif works; the extension drives MIME detection.
net.ladenthin.llama.decision.model unset (test self-skips) test SystemOneIntegrationTest Path to a decision model GGUF (laya, julia-1, lev, openjev, kev, ...) for the /v1/systemone tests of LlamaModel.handleSystemOne; upstream tests with ggml-org/tinylaya-for-testing-gguf. Without it only the rejection of a non-decision model runs.
net.ladenthin.llama.audio.model unset (test self-skips) test AudioInputIntegrationTest (llama.cpp discussion #13759) Path to an audio-input model GGUF (e.g. Ultravox, Qwen2.5-Omni).
net.ladenthin.llama.audio.mmproj unset (test self-skips) test AudioInputIntegrationTest Matching audio mmproj (encoder) GGUF.
net.ladenthin.llama.audio.input src/test/resources/audios/sample.wav (committed) test AudioInputIntegrationTest .wav/.mp3 audio prompt clip; the extension drives format detection.
net.ladenthin.llama.tts.model models/Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf (test self-skips if missing) test TtsIntegrationTest Path to the Qwen3-TTS backbone (text) GGUF. Any Qwen3-TTS-family model works; CI default is Qwen3-TTS-12Hz-1.7B-Base-Q4_K_M.gguf.
net.ladenthin.llama.tts.mmproj models/mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0.gguf (test self-skips if missing) test TtsIntegrationTest Path to the matching Qwen3-TTS mmproj GGUF (speaker encoder + code predictor + code2wav decoder); CI default is mmproj-Qwen3-TTS-12Hz-1.7B-Base-Q8_0.gguf.
net.ladenthin.llama.train.model models/stories260K.gguf (test self-skips if missing) test LlamaTrainerIntegrationTest Path to the model the fine-tuning smoke trains. Must be F32: llama_set_param skips every other tensor type, so a quantized model would train nothing.

The test-model defaults are exactly CI's model set (.github/models.csv, URLs included), so a model downloaded into models/ is found without any property; the properties only point a test at another file. MultimodalIntegrationTest self-skips when any of the three vision.* paths is missing, so a partial setup (just the vision model + the committed image, no mmproj) lets the test class load without erroring. AudioInputIntegrationTest self-skips the same way over the three audio.* properties. TtsIntegrationTest likewise self-skips unless both tts.model and tts.mmproj paths exist, and LlamaTrainerIntegrationTest unless train.model (default models/stories260K.gguf, an F32 model) does.

Documentation

Example

This is a short example on how to use this library:

public class Example {

    public static void main(String... args) throws IOException {
        ModelParameters modelParams = new ModelParameters()
                .setModel("models/mistral-7b-instruct-v0.2.Q2_K.gguf")
                .setGpuLayers(43);

        String system = "This is a conversation between User and Llama, a friendly chatbot.\n" +
                "Llama is helpful, kind, honest, good at writing, and never fails to answer any " +
                "requests immediately and with precision.\n";
        BufferedReader reader = new BufferedReader(new InputStreamReader(System.in, StandardCharsets.UTF_8));
        try (LlamaModel model = new LlamaModel(modelParams)) {
            System.out.print(system);
            String prompt = system;
            while (true) {
                prompt += "\nUser: ";
                System.out.print("\nUser: ");
                String input = reader.readLine();
                prompt += input;
                System.out.print("Llama: ");
                prompt += "\nLlama: ";
                InferenceParameters inferParams = new InferenceParameters(prompt)
                        .withTemperature(0.7f)
                        .withMiroStat(MiroStat.V2)
                        .withStopStrings("User:");
                for (LlamaOutput output : model.generate(inferParams)) {
                    System.out.print(output);
                    prompt += output;
                }
            }
        }
    }
}

Also have a look at the other examples.

Inference

There are multiple inference tasks. In general, LlamaModel is stateless, i.e., you have to append the output of the model to your prompt in order to extend the context. If there is repeated content, however, the library will internally cache this, to improve performance.

ModelParameters modelParams = new ModelParameters().setModel("/path/to/model.gguf");
InferenceParameters inferParams = new InferenceParameters("Tell me a joke.");
try (LlamaModel model = new LlamaModel(modelParams)) {
    // Stream a response and access more information about each output.
    for (LlamaOutput output : model.generate(inferParams)) {
        System.out.print(output);
    }
    // Calculate a whole response before returning it.
    String response = model.complete(inferParams);
    // Returns the hidden representation of the context + prompt.
    float[] embedding = model.embed("Embed this");
}

Note

Since llama.cpp allocates memory that can't be garbage collected by the JVM, LlamaModel is implemented as an AutoClosable. If you use the objects with try-with blocks like the examples, the memory will be automatically freed when the model is no longer needed. This isn't strictly required, but avoids memory leaks if you use different models throughout the lifecycle of your application.

Chat Completion

For chat models, build a list of role/content pairs and let the library apply the model's chat template. chatComplete() returns the full response, generateChat() streams tokens, and chatCompleteText() returns just the text content of the assistant message.

List<Pair<String, String>> messages = new ArrayList<>();
messages.add(new Pair<>("user", "Write a haiku about Java."));

InferenceParameters inferParams =
        new InferenceParameters("").withMessages("You are a helpful assistant.", messages);

try (LlamaModel model = new LlamaModel(modelParams)) {
    // Streaming
    for (LlamaOutput output : model.generateChat(inferParams)) {
        System.out.print(output);
    }
    // Or blocking, returns the OpenAI-compatible JSON envelope
    String json = model.chatComplete(inferParams);
    // Or just the assistant text
    String text = model.chatCompleteText(inferParams);
}

Reasoning/thinking models can receive custom Jinja template variables via ModelParameters#setChatTemplateKwargs(Map).

Vision / Multimodal Chat

Load a vision-capable GGUF with its matching projector, then place text and image parts in the same user message. Images may come from a file, raw bytes, a data URI, or an HTTP(S) URL:

ModelParameters modelParams = new ModelParameters()
        .setModel("models/SmolVLM-500M-Instruct-Q8_0.gguf")
        .setMmproj("models/mmproj-SmolVLM-500M-Instruct-Q8_0.gguf");

ChatMessage message = ChatMessage.userMultimodal(
        ContentPart.text("Describe this image in one short sentence."),
        ContentPart.imageFile(Paths.get("photo.jpg")));

try (LlamaModel model = new LlamaModel(modelParams)) {
    String answer = model.chatCompleteText(InferenceParameters.empty()
            .withMessages(Collections.singletonList(message))
            .withNPredict(64));
    System.out.println(answer);
}

The same multipart messages[].content shape works through ChatRequest and the embedded OpenAI-compatible /v1/chat/completions server. For a strictly CPU-only run, use setDevices("none").setMmprojOffload(false) in addition to setGpuLayers(0); projector offload has its own upstream default.

On a multi-GPU host the projector can be placed independently of the weights with setMmprojDevice("CUDA1") (llama.cpp --mmproj-device, added upstream in b10541). Exactly one device may be named; the literal "none" keeps the projector on the CPU. OpenAiCompatServer's CLI accepts the same flag as -mmdev/--mmproj-device, and NativeServer forwards it verbatim like every other llama-server flag.

setMmprojDevice(...) and setMmprojOffload(...) write the same upstream field (common_params::mmproj_use_gpu), so where they disagree the outcome would depend on argv order — and the rendered argv comes from a HashMap, whose order is unspecified. The builder therefore resolves the two genuinely ambiguous combinations by dropping the earlier call, and leaves the rest alone:

Combination Resolves to Builder behaviour
named device + setMmprojOffload(true) (use_gpu=true, device) in either order both kept — no clash
named device + setMmprojOffload(false) order-dependent last call wins
"none" + setMmprojOffload(true) order-dependent last call wins
"none" + setMmprojOffload(false) (use_gpu=false) in either order both kept — no clash

So a multi-GPU projector pin survives an explicit setMmprojOffload(true); only a call that would actually contradict the other is dropped. If you need a device after disabling offload, call setMmprojDevice last.

Video input — decode settings only, so far. mtmd has carried a video path since llama.cpp b9562 (#24269); b10647 (#24318) added the --video-* CLI flags and the mtmd_helper_init_opt plumbing that surfaces them. It is compiled into the shipped desktop library (MTMD_VIDEO is on by default, gated on LLAMA_SUBPROCESS, which upstream force-disables on Android and iOS). Its decode settings are exposed as setVideoFps(float), setVideoTimestampInterval(long) and setVideoFfmpegDir(String). The last one matters most in a JVM: upstream shells out to ffmpeg/ffprobe and resolves them from PATH, which an application server, an Android app or a JAR-only container frequently does not have them on — naming the directory is then the only way for video to work at all.

What is not here yet is the content part: upstream's wire type for a video is {"type":"input_video","input_video":{"data":"<base64>"}} (raw base64, not a data: URI, unlike image_url), gated server-side on mtmd_helper_support_video. ContentPart has no videoFile(...) factory emitting that shape, so these knobs currently configure a path this API cannot yet feed directly. Tracked in TODO.md.

Audio input works identically — load an audio-capable model (Ultravox, Qwen2.5-Omni, …) with its audio --mmproj and add a ContentPart.audioFile(...) (or inputAudio(bytes, "wav"|"mp3")) part. It serializes to the OpenAI input_audio content part and routes through the same mtmd pipeline:

ModelParameters modelParams = new ModelParameters()
        .setModel("models/ultravox-v0_5-llama-3_2-1b.gguf")
        .setMmproj("models/mmproj-ultravox-v0_5-llama-3_2-1b-f16.gguf");

ChatMessage message = ChatMessage.userMultimodal(
        ContentPart.text("Transcribe the audio."),
        ContentPart.audioFile(Paths.get("speech.wav")));

try (LlamaModel model = new LlamaModel(modelParams)) {
    System.out.println(model.supportsAudio()); // true
    String answer = model.chatCompleteText(InferenceParameters.empty()
            .withMessages(Collections.singletonList(message))
            .withNPredict(64));
    System.out.println(answer);
}

LlamaModel.supportsVision() / supportsAudio() report which modalities the loaded projector enables.

Tool Calling

Use a tool-aware instruct model and enable Jinja when loading it. A typed request can either return the model's tool calls through chat, or execute registered handlers until the model produces a normal assistant response through chatWithTools:

ToolDefinition weather = new ToolDefinition(
        "get_weather",
        "Get the current weather for a city",
        "{\"type\":\"object\",\"properties\":{\"city\":{\"type\":\"string\"}},"
                + "\"required\":[\"city\"]}");

ChatRequest request = ChatRequest.empty()
        .appendMessage("user", "What is the weather in Paris?")
        .appendTool(weather)
        .withToolChoice("auto")
        .withParallelToolCalls(Boolean.FALSE);

Map<String, ToolHandler> handlers = Collections.singletonMap(
        "get_weather", argumentsJson -> "{\"temperature_c\":21,\"condition\":\"sunny\"}");

try (LlamaModel model = new LlamaModel(new ModelParameters()
        .setModel("models/Qwen2.5-1.5B-Instruct-Q4_K_M.gguf")
        .enableJinja())) {
    ChatResponse response = model.chatWithTools(request, handlers);
    System.out.println(response.getFirstContent());
}

tool_choice is the OpenAI-compatible string form (auto, none, or required). Set parallel_tool_calls to false when handlers should be issued one at a time. Handler failures and unknown tool names are returned to the model as valid {"error":"..."} tool-result JSON.

Infilling

You can simply set InferenceParameters#withInputPrefix(String) and InferenceParameters#withInputSuffix(String).

Embeddings & Reranking

Load the model with enableEmbedding() (or enableReranking()) and call embed(String) to get a sentence embedding, or rerank(query, documents...) to get relevance scores.

ModelParameters modelParams = new ModelParameters()
        .setModel("/path/to/embedding-model.gguf")
        .enableEmbedding();
try (LlamaModel model = new LlamaModel(modelParams)) {
    float[] embedding = model.embed("Embed this sentence");
    // Batch form: one native dispatch for many inputs, results in request order.
    List<float[]> embeddings = model.embed(Arrays.asList("First sentence", "Second sentence"));
}

Kolibri-1 (Aleph Alpha)

Kolibri-1 is Aleph Alpha's 78B-parameter Mixture-of-Experts reasoning model for German and English (about 3.5B parameters active per token, Apache-2.0). Upstream llama.cpp does not support its architecture yet (ggml-org/llama.cpp#29922), so this library carries it as a patch (llama/patches/0016-model-kolibri1.patch) until upstream does. It loads the community GGUFs of both published converters -- e.g. Hob-forge/Kolibri-1-GGUF and Eliasfpv28/Kolibri-1-Q3_K_S-GGUF -- which the two community llama.cpp patches cannot each load from the other. The embedded chat template handles reasoning and Hermes-style tool calls; for a split GGUF, point the model path at the first part.

ModelParameters params = new ModelParameters()
        .setModel("/models/Kolibri-1-Q4_K_M.gguf")
        .setCtxSize(32768)
        .enableJinja();
try (LlamaModel model = new LlamaModel(params)) {
    ChatResponse answer = model.chat(ChatRequest.empty()
            .appendMessage("user", "Warum ist der Himmel blau?"));
}

The model is large (the Q4_K_M file is 47.5 GB) and runs from system RAM on the CPU, or partly offloaded to a GPU. The architecture is checked numerically against Aleph Alpha's reference on tiny random models in every C++ test run; the real model was not run in this project's CI, and the GPU backends are untested for it.

Decision models (/v1/systemone)

A decision model answers typed questions about a state without generating text: each question is evaluated in one forward pass and returns probabilities. handleSystemOne takes and returns llama.cpp's TypeSafe-compatible /v1/systemone JSON, served by the upstream handler itself (the full request/response description is in upstream's tools/server/README.md). The native server serves the same endpoint at POST /v1/systemone, in attach mode too.

try (LlamaModel model = new LlamaModel(new ModelParameters().setModel("/path/to/laya.gguf"))) {
    String answers = model.handleSystemOne("{"
            + "\"state\": \"I was charged twice for my order and nobody replied.\","
            + "\"questions\": {"
            + "  \"route\":   {\"type\": \"choice\", \"instructions\": \"Which team?\","
            + "                \"criteria\": {\"billing\": null, \"shipping\": null}},"
            + "  \"urgency\": {\"type\": \"score\", \"instructions\": \"How urgent?\","
            + "                \"criteria\": [\"can wait\", \"today\", \"right now\"]},"
            + "  \"angry\":   {\"type\": \"noul\", \"instructions\": \"Is the customer angry?\"}}}");
    // {"answers": {"route": {"choice": "billing", "probabilities": {...}, ...}, ...}, "usage": {...}}
}

A model that is not a decision model throws a LlamaException ("This model is not a decision model").

Runtime LoRA adapter control

Adapters loaded at model-load time (addLoraAdapter(...) / addLoraScaledAdapter(...), optionally setLoraInitWithoutApply() to start disabled) can be listed and re-scaled at runtime without reloading the model — the typed counterpart of the upstream GET/POST /lora-adapters endpoints:

ModelParameters modelParams = new ModelParameters()
        .setModel("models/base.gguf")
        .addLoraScaledAdapter("models/adapter.gguf", 1.0f);
try (LlamaModel model = new LlamaModel(modelParams)) {
    List<LoraAdapter> adapters = model.getLoraAdapters();      // [{id=0, path=..., scale=1.0}]
    model.setLoraAdapter(0, 0.5f);                             // re-scale at runtime
    model.setLoraAdapters(Collections.emptyMap());             // disable all adapters
}

Per the upstream contract, a scale update lists the adapters to keep active — any adapter missing from the map is set to scale 0 (disabled). The native side clears affected KV caches when the effective adapter set changes.

Text-to-Speech

TextToSpeech synthesizes audio from text over llama.cpp's upstream Qwen3-TTS pipeline (mtmd_helper::gen_audio). It is a separate AutoCloseable native type (not a LlamaModel) because TTS loads its own model pair: a backbone (text) GGUF and an mmproj GGUF bundling the speaker encoder, code predictor, and code2wav decoder. synthesize(String) returns a 24 kHz mono 16-bit WAV byte stream.

try (TextToSpeech tts = new TextToSpeech(
        "models/qwen3-tts-backbone.gguf", "models/qwen3-tts-mmproj.gguf")) {
    byte[] wav = tts.synthesize("Hello from llama dot c p p.");
    Files.write(Paths.get("out.wav"), wav);
}

Add (modelPath, mmprojPath, gpuLayers, threads) to offload to the GPU, or synthesize(text, maxFrames, topK, seed) for explicit sampling, or the full synthesize(text, speakerReferenceAudioPath, language, maxFrames, topK, seed) overload for voice cloning from a reference clip. As with LlamaModel, native memory is not GC-managed — use try-with-resources or call close().

This replaces the project's earlier two-model OuteTTS + WavTokenizer pipeline, which upstream #26254 deleted entirely in favor of Qwen3-TTS (llama.cpp b10270); there is no backward-compatible path for the old model pair.

GGUF Quantization

LlamaQuantizer converts a GGUF to another quantization scheme in-process (llama.cpp's llama_model_quantize — the llama-quantize tool without the separate binary):

LlamaQuantizer.quantize("model-f16.gguf", "model-q4_k_m.gguf", QuantizationType.Q4_K_M);
// Re-quantizing an already-quantized GGUF degrades quality and must be opted into:
LlamaQuantizer.quantize("model-q8_0.gguf", "model-q4_0.gguf", QuantizationType.Q4_0,
        /* threads */ 0, /* allowRequantize */ true);

Raw JSON Endpoints

For direct access to the upstream llama.cpp server API, the following methods take a JSON request and return a JSON response, matching the HTTP server's contract:

handleCompletions, handleCompletionsOai, handleChatCompletions, handleInfill, handleEmbeddings, handleTokenize, handleDetokenize.

Server state is exposed via getMetrics(), eraseSlot(int), saveSlot(int, String), restoreSlot(int, String), and getModelMeta().

Conversation checkpoints: rewind + fork (Session)

A Session can be snapshotted and branched — the KV-cache slot state and the transcript move together, so native state and history can never drift apart:

try (Session session = new Session(model, 0, "You are terse.")) {
    session.send("My name is Alice.");
    SessionCheckpoint cp = session.checkpoint("checkpoints/turn1.bin");

    session.send("Tell me a joke.");
    session.rewind(cp);                     // undo everything after the checkpoint
    session.send("Tell me a story instead."); // retry from the branch point

    // Branch into a second slot (model loaded with setParallel(2)+):
    try (Session forked = session.fork(1, "checkpoints/branch.bin")) {
        forked.send("Answer as a pirate.");   // both sessions continue independently
    }
}

Checkpoint files are caller-managed (KV dumps grow with context usage) and both operations are rejected while a stream is in progress. For plain transformer models a rewind is also achievable cheaply by resending a truncated history with cache_prompt (prefix reuse); checkpoints make the branch point exact and are the only reliable rollback for recurrent/hybrid models (e.g. Granite-4), whose state cannot be recomputed from a prefix.

GGUF metadata inspection (no model load)

GgufInspector reads a GGUF's header and key/value table without loading the model — pure Java, no native library, cost independent of file size (parsing stops before the tensor data). Useful for model pickers and download validators:

GgufMetadata meta = GgufInspector.read(Paths.get("models/Qwen3-0.6B-Q4_K_M.gguf"));
meta.getArchitecture();   // Optional[qwen3]
meta.getModelName();      // Optional[Qwen3 0.6B]
meta.getParameterCount(); // OptionalLong[751632384]
meta.getContextLength();  // OptionalLong[40960]  (<arch>.context_length)
meta.getFileType();       // OptionalLong[15]     (llama_ftype, cf. QuantizationType)
meta.getChatTemplate();   // Optional[{{- ... }}]
meta.getEntries();        // full decoded key/value table

Supports GGUF v2/v3, little- and big-endian (auto-detected), and fails loud on v1/corrupt files. For metadata of an already-loaded model use getModelMeta() instead.

Prompt and KV Cache Reuse

Prompt-prefix reuse is enabled by default in llama.cpp and can be controlled per request with InferenceParameters.withCachePrompt(boolean). withCacheReuse(int) enables non-prefix chunk reuse, while withSlotId(int) pins a request to a specific server slot. Session applies its slot id to every request, so generation and save/restore operate on the same KV state.

Typed results expose logical prompt, generated, cached prompt, and evaluated prompt counts through Usage. Per-request timing also remains available through Timings.getCacheN(). LlamaModel.getMetricsTyped().getSlotMetrics() reports each slot's logical, processed, cached, decoded, and remaining token counts, and the same ServerMetrics view carries the server-wide lifetime counters — including cached prompt tokens (getCumulativeCachedPromptTokens()) and the speculative-decoding tallies (getDraftTokensTotal(), getDraftAcceptedTotal(), getDraftVerifyStepsTotal(), getDraftAcceptedPerPosition(), plus the derived getDraftAcceptanceRate()), which upstream otherwise exposes only as Prometheus text.

The embedded HTTP server exposes the same native JSON at authenticated GET /metrics, with the slot array alone at GET /slots. OpenAI responses preserve usage.prompt_tokens_details.cached_tokens; Responses API output uses usage.input_tokens_details.cached_tokens; Anthropic output uses cache_read_input_tokens.

OpenAI-compatible HTTP server

net.ladenthin.llama.server.OpenAiCompatServer turns a loaded model into a local OpenAI-compatible HTTP endpoint using only the JDK's built-in com.sun.net.httpserver — no extra dependency and no separate server process. It is embeddable, and runnable via java -cp <jar> net.ladenthin.llama.server.OpenAiCompatServer … (the classes jar's Main-Class, ServerLauncher, starts NativeServer by default — see "Native server with the built-in WebUI" below). It serves:

Method & path Backed by
POST /v1/chat/completions LlamaModel.streamChatCompletion (streaming SSE) / chatComplete (blocking)
POST /v1/completions LlamaModel.handleCompletionsOai
POST /v1/embeddings (requires --embedding) LlamaModel.handleEmbeddings
POST /v1/rerank (requires --reranking) LlamaModel.handleRerank (reshaped to results/data)
POST /infill LlamaModel.handleInfill (fill-in-the-middle autocomplete)
GET /v1/models the configured model id
GET /metrics native server and per-slot token/cache counters (JSON)
GET /slots native per-slot token/cache counters (JSON array)
GET /health static {"status":"ok"} (unauthenticated)

Chat completions support streaming via Server-Sent Events and non-streaming, forwarding messages/tools verbatim. The streaming path carries delta.tool_calls and (with stream_options.include_usage) a trailing usage chunk, so agent/tool-calling clients work — this is the recommended surface for VS Code Copilot agent mode, Cline, Roo Code and Continue. response_format (json_object / json_schema) is forwarded for structured outputs. Completions, embeddings, rerank and infill are non-streaming.

Every route is also reachable without the /v1 prefix, the server answers CORS preflight (OPTIONS) and stamps Access-Control-Allow-Origin (so browser/webview clients work), and POST /infill is the llama.cpp-native FIM endpoint for local ghost-text autocomplete plugins (llama.vscode, Twinny, Tabby, Continue's llama.cpp provider). Note: GitHub Copilot's inline completions cannot be served by any local endpoint — only its chat/agent surfaces — so use one of those autocomplete plugins for ghost text.

Alternative protocol surfaces. For clients that don't speak OpenAI Chat Completions, the same model is exposed through additional protocols (pure translation over the OpenAI core — no extra inference path), all supporting tools and streaming:

Surface Routes For
Ollama-native GET /api/version, GET /api/tags, POST /api/show, POST /api/chat (NDJSON streaming), POST /api/generate (prompt completion / FIM) Copilot's built-in Ollama provider; Ollama-hardcoded tools
Anthropic Messages POST /v1/messages (SSE event stream) Claude-shaped clients (Claude Code); Copilot messages apiType
OpenAI Responses POST /v1/responses (SSE event stream) Copilot responses apiType; Responses-API clients

/api/show advertises the model's capabilities (tools, insert, and vision when --mmproj is set) and context length, which Copilot's Ollama provider reads to enable agent mode. The llama.cpp-native GET /props reports default_generation_settings.n_ctx and a modalities block, which autocomplete clients such as llama.vscode read to size their context window.

Embed it in your app:

ModelParameters modelParams = new ModelParameters().setModel("models/model.gguf").setParallel(2);
OpenAiServerConfig config = OpenAiServerConfig.builder().port(8080).modelId("local-model").build();
try (LlamaModel model = new LlamaModel(modelParams);
     OpenAiCompatServer server = new OpenAiCompatServer(model, config).start()) {
    Thread.currentThread().join(); // serve until interrupted
}

…or run it standalone. The classes jar's Main-Class is the ServerLauncher dispatcher, so add --jllama-openai-compat to select this Java server (the launcher strips that flag and forwards the rest); or name the class explicitly via -cp:

# from Maven Central (JBang resolves the classes jar, the natives jar named and the Java deps)
jbang --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64 net.ladenthin:llama:5.2.0 --jllama-openai-compat \
    --model models/Qwen3-0.6B-Q4_K_M.gguf --host 0.0.0.0 --port 8080 --n-gpu-layers 99

# or name the class explicitly, with the jars on the classpath
java -cp 'target/llama-<version>.jar:target/llama-<version>-cpu-linux-x86-64.jar:<deps>' \
  net.ladenthin.llama.server.OpenAiCompatServer --model models/model.gguf --port 8080 --model-id local-model

Run with --help for the full option list (-m/--model, --host, -p/--port, -c/--ctx-size, -b/--batch-size, -ub/--ubatch-size, -ngl/--n-gpu-layers, -t/--threads, -tb/--threads-batch, -ctk/--cache-type-k, -ctv/--cache-type-v, --jinja, --chat-template-kwargs, --parallel, --model-id, --api-key, --mmproj, -mmdev/--mmproj-device, --embedding, --reranking). The tuning flags mirror llama.cpp's server, so an invocation like --jinja --chat-template-kwargs '{"reasoning_effort":"low"}' -ctk q8_0 -ctv q8_0 -b 4096 -ub 2048 works directly.

Verify with curl (streaming chat):

curl -N http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"local-model","stream":true,"messages":[{"role":"user","content":"hi"}]}'

VS Code Copilot setup: Command Palette → Chat: Manage Language Models → Add Models → Custom Endpoint; enter a group name, a display name and any non-empty API key, and pick API type Chat Completions. VS Code then opens chatLanguageModels.json — set the model url to your endpoint (the host/port go here, not in the form):

[
  {
    "name": "Local llama.cpp",
    "vendor": "customendpoint",
    "apiKey": "local-dummy-key",
    "apiType": "chat-completions",
    "models": [
      {
        "id": "local-model",
        "name": "Local model",
        "url": "http://127.0.0.1:8080/v1/chat/completions",
        "toolCalling": true,
        "vision": false,
        "maxInputTokens": 6144,
        "maxOutputTokens": 2048
      }
    ]
  }
]

Notes: BYOK powers the chat/agent experience only (inline completions and embeddings still require a GitHub account). On CPU, prefer a smaller model and a modest context window — the server emits SSE heartbeats so a long prompt prefill does not trip the client's stream-inactivity timeout. Agent-mode tool calling depends on the model's own tool-calling quality. Pass --api-key (or OpenAiServerConfig.apiKey(...)) to require an Authorization: Bearer token; the server binds to 127.0.0.1 by default.

Native server with the built-in WebUI (NativeServer)

OpenAiCompatServer above is a JSON API server (its / is a 404 — no web page). If you want the full upstream llama.cpp server, including its bundled Svelte WebUI, use net.ladenthin.llama.server.NativeServer. It runs the real llama_server inside libjllama over JNI — no separate llama-server.exe — and forwards the raw llama-server arguments verbatim, so every flag works exactly as it does for the standalone binary. ServerLauncher (the classes jar's Main-Class) runs it by default (when --jllama-openai-compat is absent), forwarding its args to the native server (pass --help for the full llama-server option list):

jbang --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64 net.ladenthin:llama:5.2.0 \
    -m models/model.gguf --host 127.0.0.1 --port 8080 -c 65536 --jinja
# then open http://127.0.0.1:8080/ for the WebUI

Or embed it:

try (NativeServer server = new NativeServer(
        "-m", "gpt-oss-20b-UD-Q4_K_XL.gguf",
        "--host", "127.0.0.1", "--port", "8080",
        "-c", "65536", "-b", "4096", "-ub", "2048",
        "--jinja", "-ngl", "0", "-t", "8", "-tb", "16",
        "-ctk", "q8_0", "-ctv", "q8_0",
        "--chat-template-kwargs", "{\"reasoning_effort\":\"low\"}",
        "--parallel", "1").start()) {
    // Open http://127.0.0.1:8080/ in a browser for the WebUI; the OpenAI API is at /v1/... too.
    Thread.currentThread().join();
}

Differences from OpenAiCompatServer: with the classic constructor it loads its own model from the arguments (an independent lifecycle, like llama-server.exe), it is single-instance per process, it serves the WebUI (in released jars — local cmake builds ship the empty-asset stub, so no UI there), and it is not available on Android (the upstream server needs posix_spawn). Readiness: poll GET /health. TLS: pass --ssl-key-file and --ssl-cert-file (BoringSSL is linked statically into every desktop natives jar, nothing to install). The same build fetches models from https:// URLs (--model-url, -hf), verified against the OS certificate store -- on Linux /etc/ssl/certs or /etc/ssl/cert.pem, overridable with SSL_CERT_FILE / SSL_CERT_DIR. Note that --model-url alone starts the router; name the download target with -m as well.

Attach mode — serve an already-loaded LlamaModel

NativeServer can also attach the full upstream HTTP frontend (routes, WebUI, resumable streaming) to a LlamaModel you already loaded — one copy of the weights, shared between direct JNI calls and HTTP:

try (LlamaModel model = new LlamaModel(new ModelParameters().setModel("models/model.gguf"));
     NativeServer server = new NativeServer(model, "--host", "127.0.0.1", "--port", "8080").start()) {
    // HTTP (incl. WebUI in released jars) and direct Java calls share the same loaded model.
    String direct = model.complete(new InferenceParameters("2+2=").withNPredict(4));
    Thread.currentThread().join();
}

In attach mode the arguments carry only the HTTP-side flags (--host, --port, --api-key, --ssl-key-file, …; no -m); the routes served are the model's own, so route-level settings — the endpoint toggles --metrics / --props / --slots, --slot-save-path, the slot count — are set on the model's ModelParameters (enableMetricsEndpoint(), enablePropsEndpoint(), enableSlotsEndpoint() / disableSlotsEndpoint()), not in the attach arguments. The server reports healthy immediately (the model is already loaded), and the caller keeps ownership of the model — close the server before the model, never the other way around.

Router mode — multi-model management

Started without a model argument, the upstream server runs in router mode: it lists models from --models-dir, loads/unloads them on demand (GET /models, POST /models/load, POST /models/unload, per-request "model" selection) and serves each model from a worker subprocess. Upstream spawns workers by re-executing its own binary — inside a JVM that binary is java, so before starting an embedded router you must point the worker spawn at this library's bootstrap:

String javaBin = System.getProperty("java.home") + File.separator + "bin" + File.separator + "java";
NativeServer.setWorkerCommand(javaBin, "-cp", System.getProperty("java.class.path"),
        "net.ladenthin.llama.server.NativeServer");
try (NativeServer router = new NativeServer(
        "--host", "127.0.0.1", "--port", "8080", "--models-dir", "models").start()) {
    Thread.currentThread().join(); // each loaded model runs as a fresh worker JVM
}

Worker-command tokens may not contain whitespace (the value is whitespace-split natively).

Typed model management (RouterClient). Instead of hand-rolling HTTP+JSON against the management endpoints, use server.RouterClient — a plain-HTTP typed client (works against the embedded router above or any external llama-server router):

RouterClient client = new RouterClient(8080);
List<RouterModel> models = client.listModels();          // GET /models, typed status per entry
client.loadModel("Qwen3-0.6B-Q4_K_M");                   // POST /models/load (non-blocking)
client.awaitModelLoaded("Qwen3-0.6B-Q4_K_M", 240_000L);  // poll until LOADED; fails fast if the
                                                         // worker died (exit code in the message)
client.unloadModel("Qwen3-0.6B-Q4_K_M");                 // POST /models/unload

RouterModel carries the identifier, the lifecycle status (UNLOADED/LOADING/LOADED/SLEEPING/DOWNLOADING/DOWNLOADED), the router's failed-worker marker, and the model's input/output modalities (getInputModalities(), getOutputModalities()), which the router computes without loading the model: isDecisionModel() picks out a decision model for /v1/systemone before its first load. A loaded LlamaModel reports the same through getModelMeta(). Chat requests then select a model per request via the standard "model" field on POST /v1/chat/completions.

Against a router started with --api-key, pass the key — it is sent as Authorization: Bearer <key> on every call. All of them need it: /models/load and /models/unload were always gated, and since llama.cpp b10519 the listing endpoints are too.

RouterClient client = new RouterClient(8080, System.getenv("LLAMA_API_KEY"));
// or, for a remote router: new RouterClient("router.internal", 8080, key)

Note

awaitModelLoaded waits by polling GET /models, so it cannot observe a model the router deliberately hides from that listing — a cache model deduplicated by a preset with dedup-cache-models still loads and still serves by name, but never appears. For those, skip the await and issue the request directly; with autoload the router waits for the worker itself.

Distributed inference over RPC

llama.cpp's RPC backend spreads one model over the devices of several machines: every machine that contributes runs an RPC server, and the machine that loads the model names them with --rpc. Layers are then distributed over local and remote devices exactly as over several local GPUs (setGpuLayers, setTensorSplit). Both halves are in every natives jar — CPU and GPU — with no additional runtime dependency (plain TCP over the system socket library the library already links).

Serve this machine's devices (every GPU this library found, else the CPU):

try (RpcServer server = RpcServer.startLocal(RpcEndpoint.DEFAULT_PORT)) {   // 127.0.0.1:50052
    server.awaitTermination();
}

or from the command line, from Maven Central:

jbang --main net.ladenthin.llama.RpcServer --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64 \
    net.ladenthin:llama:5.2.0 --port 50052

Use the servers from a model:

ModelParameters params = new ModelParameters()
        .setModel("models/big-model.gguf")
        .setGpuLayers(99)
        .setRpcServers(RpcEndpoint.parse("10.0.0.2:50052"), RpcEndpoint.parse("10.0.0.3:50052"));

The same works for both HTTP servers: --rpc 10.0.0.2:50052,10.0.0.3:50052 is forwarded to the native server as-is, and OpenAiCompatServer accepts it too. Any upstream rpc-server works as a server, and this library's RpcServer works for any llama.cpp client.

Warning

The RPC protocol has no authentication and no encryption: whoever reaches the port can use the devices and read or write the tensors on them. RpcServer.startLocal therefore binds to loopback only; RpcServer.startOnNetwork(address, …) (or --host on the command line) is the explicit opt-in for another interface and logs a warning. Across machines, use a trusted network or a tunnel (SSH, WireGuard).

What to know:

  • An unreachable server fails the load with a LlamaException naming it, instead of reaching llama.cpp. A server that disappears after the model loaded still terminates the process — llama.cpp has no error path for a device lost mid-inference.
  • One RpcServer per process. A second start while one runs throws IllegalStateException.
  • Endpoints are IPv4 addresses or host names (host:port); llama.cpp's RPC transport has no IPv6. RpcServer binds to an IPv4 literal (127.0.0.1, 0.0.0.0, an interface address).
  • Registered servers stay registered. llama.cpp keeps RPC devices in a process-wide registry with no way to remove them; this library therefore gives every later load that does not ask for a server an explicit device list without it, so a model loaded without --rpc never offloads to a server an earlier model used — the same holds for the multimodal projector, TextToSpeech and LlamaTrainer. An explicit setDevices(...) / --device is never overridden.
  • Android needs the android.permission.INTERNET permission for RPC, even over loopback — the llama-android AAR does not request it, so an app that wants RPC must declare it itself.
  • RpcServer.startLocal(port, threads, cacheDir) enables upstream's tensor cache: a client that loads the same model again sends the large tensors only once.
  • Choose the served devices when the default is wrong. RpcServer.startLocal(port, threads, cacheDir, Arrays.asList("CPU")) (or --device CPU on the command line; names as llama.cpp prints them, e.g. CUDA0, Vulkan1, MTL0) replaces the default of every accelerator. It matters because llama.cpp's RPC client treats every operation as supported by the remote device: a served GPU that cannot run one terminates the server process on the first graph that needs it. The paravirtual GPU of a macOS virtual machine is such a device — serve CPU there.

LangChain4j integration

A separate artifact, net.ladenthin:llama-langchain4j, adapts a LlamaModel to LangChain4j's ChatModel, StreamingChatModel, EmbeddingModel and ScoringModel interfaces in-process over JNI — no HTTP hop, no separate server. It is a separate artifactId (not a classifier of the core) because LangChain4j 1.x requires Java 17 while the core net.ladenthin:llama stays Java 8; keeping it separate avoids forcing that floor on every core consumer. It ships and versions in lockstep with the core.

<dependency>
    <groupId>net.ladenthin</groupId>
    <artifactId>llama-langchain4j</artifactId>
    <version>5.1.0</version>
</dependency>

From 5.2.0 on, add the natives next to it — llama-platform or the natives jars you need (see Choosing the natives jars); the core it depends on is classes only.

Each adapter borrows a LlamaModel you already loaded — it never loads or closes the native model, so you manage its lifecycle (try-with-resources), and one LlamaModel can back several adapters at once:

try (LlamaModel llama = new LlamaModel(new ModelParameters().setModel("models/qwen3-0.6b.gguf"))) {
    ChatModel chat = new JllamaChatModel(llama);
    String reply = chat.chat("Write a haiku about lazy senior devs.");
    System.out.println(reply);
}
Adapter LangChain4j interface java-llama.cpp call
JllamaChatModel ChatModel LlamaModel.chat(...)
JllamaStreamingChatModel StreamingChatModel LlamaModel.generateChat(...) (token streaming)
JllamaEmbeddingModel EmbeddingModel LlamaModel.embed(...) (model loaded with enableEmbedding())
JllamaScoringModel ScoringModel (re-ranking) LlamaModel.handleRerank(...) (model loaded with enableReranking())

See llama-langchain4j/README.md for streaming/embedding/re-ranking examples and the current mapping limitations (tool calling, JSON mode, and multimodal input are not yet forwarded).

Local agent: terminal, browser, IDE via ACP

Tip

Supported front ends — all on the same agent session (history, slash commands, approval mode):

Front end Start with Where it runs
Terminal (default) / --plain this console; --plain for piped, logged or line-only sessions
Browser --web http://127.0.0.1:8787/?token=… — over SSH: ssh -L 8787:127.0.0.1:8787 user@server
IDE via ACP --acp JetBrains IDEs and Zed natively, VS Code through an ACP extension

llama-atmosphere-agent/ is a copy-and-run general-purpose agent on the JVM — Claude Code / OpenCode reduced to the essentials, fully offline; it edits files and, with --allow-shell, runs any command on your machine (docker, git, build tools) — built from Atmosphere's built-in OpenAI-compatible agent runtime (streaming, tool loop, workspace file tools) driven headless against this project's OpenAI-compatible server. It is a standalone Maven project (not a reactor module) published to Maven Central at the core's version, so the quickest way needs only JDK 21+ and JBang — it resolves the agent and the core with the natives of every desktop platform:

jbang net.ladenthin:llama-atmosphere-agent:5.2.0 \
    --model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --workspace /path/to/project

Or download it from a release (below, JDK 21+ only), or clone the repository and run it from that folder, which needs only JDK 21+ and Maven — the core jar from Maven Central ships the natives:

# get the folder and a tool-capable model (Qwen3-4B-Instruct-2507, 2.3 GB)
git clone --depth 1 https://github.com/bernardladenthin/java-llama.cpp.git
cd java-llama.cpp/llama-atmosphere-agent
curl -L --create-dirs -o models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf \
  https://huggingface.co/unsloth/Qwen3-4B-Instruct-2507-GGUF/resolve/main/Qwen3-4B-Instruct-2507-Q4_K_M.gguf

# the agent with the model loaded in-process — a you> prompt appears (/clear, /exit)
mvn -q compile exec:java \
    -Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ctx-size 16384 --workspace /path/to/project --allow-shell"

Straight from Maven Central instead of cloning. The agent is published as a thin jar whose pom names llama-platform, so JBang resolves it with the CPU natives of every desktop platform and starts it; a GPU is one more natives jar on the same line (the module jars are additive, see Choosing the natives jars):

jbang net.ladenthin:llama-atmosphere-agent:5.2.0 \
    --model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ctx-size 16384 --workspace /path/to/project --allow-shell

# e.g. with Vulkan on Linux x86-64 (any current GPU driver), all layers offloaded
jbang --deps net.ladenthin:llama:5.2.0:vulkan-linux-x86-64 net.ladenthin:llama-atmosphere-agent:5.2.0 \
    --model Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ngl 99 --workspace /path/to/project

The agent carries no core of its own (started without one it stops with NoClassDefFoundError: net/ladenthin/llama/LlamaModel); CI launches the published thin jar on one classpath with the published core jars on every run (smoke-agent-linux) before anything is published.

Warning

--allow-shell lets the model run any command with your user's rights. By default every write and every command is confirmed on the console ([y]es / [n]o / [a]uto); --auto turns that off. Use a workspace you are willing to hand to the model.

In the REPL, /help lists the commands (/status, /tools, /mode manual|auto, /compact, /clear, /exit); anything else goes to the model. A status line shows the approval mode and the context used ([manual · ctx ~3.1k/16k · 9 tools · local-model]), and the answer is rendered with headings, bullets and code spans.

Also in a browser or an editor. The same agent session has two more front ends. --web serves it to a browser on 127.0.0.1:8787 (Atmosphere's own AI console on an embedded Jetty, a random access token in the printed address; from another machine open an SSH tunnel, ssh -L 8787:127.0.0.1:8787 user@server, rather than binding to the network). --acp speaks the Agent Client Protocol on stdin/stdout, so JetBrains IDEs and Zed — and VS Code through an ACP extension — run it as their chat agent, with the editor's own permission dialog for writes and commands:

jbang net.ladenthin:llama-atmosphere-agent:5.2.0 --model model.gguf --allow-shell --web
# JetBrains, ~/.jetbrains/acp.json: {"agent_servers": {"Local llama": {"command": "jbang",
#   "args": ["net.ladenthin:llama-atmosphere-agent:5.2.0", "--acp", "--model", "/path/model.gguf"]}}}

On Windows PowerShell quote the whole argument ("-Dexec.args=--model models\… --allow-shell"); for the GPU add e.g. -Dllama.classifier=vulkan-windows-x86-64 and --ngl 99. The agent's README walks through all of it step by step. Other ways to run it, e.g. against a java-llama.cpp server that is already running (--jinja is required for tool calling; see Running the server from Maven Central for a GPU backend):

# 1. the server, from Maven Central
jbang --deps net.ladenthin:llama:5.2.0:cpu-linux-x86-64 net.ladenthin:llama:5.2.0 \
    -m models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --jinja --port 8080

# 2. the agent, from the llama-atmosphere-agent/ folder — a you> prompt appears (/clear, /exit)
mvn -q compile exec:java \
    -Dexec.args="--base-url http://127.0.0.1:8080/v1 --workspace /path/to/project --allow-shell"

# a single turn instead of the prompt loop
mvn -q compile exec:java \
    -Dexec.args="--base-url http://127.0.0.1:8080/v1 --workspace /path/to/project --prompt 'Read the README and summarize it'"

# or without a separate server: load the GGUF in-process
mvn -q compile exec:java \
    -Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ngl 99 --workspace /path/to/project"

# everything at once: shell access plus your own system prompt (replaces the built-in one)
mvn -q compile exec:java \
    -Dexec.args="--model models/Qwen3-4B-Instruct-2507-Q4_K_M.gguf --ngl 99 --ctx-size 16384 --workspace /path/to/project --allow-shell --system 'You are a local assistant on this machine with full shell access. run_command executes any command line, including docker, git and build tools. When asked about the system, run a command instead of explaining it. Answer in the language of the user.'"

The full streaming tool-calling loop (tools → delta.tool_calls → Java tool → role:"tool" result → next turn, over several rounds) is verified on every PR against the real OpenAiCompatServer with no model, and in CI against the Qwen2.5-1.5B tool model — both gate every publish, as does the release-jar smoke above. See llama-atmosphere-agent/README.md for the options and the verified compatibility matrix.

Model/Inference Configuration

There are two sets of parameters you can configure, ModelParameters and InferenceParameters. Both provide builder classes to ease configuration. ModelParameters are once needed for loading a model, InferenceParameters are needed for every inference task. All non-specified options have sensible defaults.

ModelParameters modelParams = new ModelParameters()
        .setModel("/path/to/model.gguf")
        .addLoraAdapter("/path/to/lora/adapter");
String grammar = """
		root  ::= (expr "=" term "\\n")+
		expr  ::= term ([-+*/] term)*
		term  ::= [0-9]""";
InferenceParameters inferParams = new InferenceParameters("")
        .withGrammar(grammar)
        .withTemperature(0.8f);
try (LlamaModel model = new LlamaModel(modelParams)) {
    model.generate(inferParams);
}

Reactive integration (Reactor, RxJava, Kotlin Flow, Akka)

LlamaIterable (returned by model.generate(...) and model.generateChat(...)) implements Iterable<LlamaOutput> & AutoCloseable, so every mainstream reactive library wraps it in a few lines without java-llama.cpp pulling in a runtime reactive dependency.

Always wrap with the library's resource-management primitive — Flux.using, Flowable.using, Kotlin use {}, etc. — so that subscription cancellation flows into LlamaIterable.close() and from there into llama.cpp's native cancelCompletion. A plain Flux.fromIterable(iterable) or for (x in iter) loop will NOT close the iterable on cancel; the native task slot stays occupied until the model is closed.

Project Reactor (Spring WebFlux)

Flux<LlamaOutput> tokens = Flux.using(
        () -> model.generate(params),
        Flux::fromIterable,
        LlamaIterable::close)
    .subscribeOn(Schedulers.boundedElastic());

RxJava 3 (also for RxAndroid)

Flowable<LlamaOutput> tokens = Flowable.using(
        () -> model.generate(params),
        Flowable::fromIterable,
        LlamaIterable::close)
    .subscribeOn(Schedulers.io());

Kotlin Flow (Android / coroutines)

Ready-made: the optional net.ladenthin:llama-kotlin artifact ships generateFlow/generateChatFlow extensions (close-on-cancellation included) plus suspend wrappers whose coroutine cancellation is wired to the binding's cooperative CancellationToken:

model.generateChatFlow(params).flowOn(Dispatchers.IO).collect { print(it.text) }

Hand-rolled equivalent (no extra dependency):

fun llama(model: LlamaModel, params: InferenceParameters) = flow {
    model.generate(params).use { iterable ->
        for (output in iterable) emit(output)
    }
}.flowOn(Dispatchers.IO)

The companion Android sample LLaMAndroid demonstrates the flow { for (output in model.generate(params)) emit(output) } shape against the upstream binding. Wrap the for loop in .use { } if your collector may cancel mid-stream — otherwise the native task slot will not be released until the model is closed.

Akka Streams

val tokens: Source[LlamaOutput, NotUsed] = Source
    .fromIterator(() => model.generate(params).iterator())
    .async("blocking-io-dispatcher")

Why no built-in Publisher? Earlier snapshots of this fork shipped a hand-rolled LlamaModel.streamPublisher(...) returning a Reactive Streams Publisher<LlamaOutput>. Since every reactive library bridges blocking iterables in a few lines via its own resource-management primitive, the binding now stays free of any reactive runtime dependency — pick whichever library your app already uses. The pattern is verified end-to-end by ReactorIntegrationTest in the test sources.

Logging

Per default, llama.cpp writes its log as text to stderr (0.00.035.060 I slot … once a model is loaded): the server's own srv … / slot … lines and, from verbosity 4 on, the llama/ggml lines. All of it can be intercepted via the static method LlamaModel.setLogger(LogFormat, BiConsumer<LogLevel, String>): with a callback set, every line goes to the callback instead of the console (a setLogFile file keeps receiving them). The callback survives model loads, so set it before new LlamaModel(…) to capture the loading lines too. LogFormat.TEXT hands over the bare message, LogFormat.JSON one JSON object per line. Passing null as the callback restores the console output (always llama.cpp's own text format; the format argument only matters with a callback). Logging can be disabled by passing an empty callback. Messages arrive asynchronously from llama.cpp's log worker thread; replacing or removing the logger flushes what is queued to the previous callback first. The verbosity threshold (ModelParameters.setLogVerbosity(int), llama.cpp's -lv: 1 errors, 2 warnings, 3 info, 4 trace, 5 debug) applies before the callback: 2 keeps warnings and errors and silences the per-request INFO lines, which is what a console application sharing the terminal with its own output wants.

// Re-direct log messages however you like (e.g. to a logging library)
LlamaModel.setLogger(LogFormat.TEXT, (level, message) -> System.out.println(level.name() + ": " + message));
// Back to llama.cpp's own console output (stderr)
LlamaModel.setLogger(LogFormat.TEXT, null);
// Disable logging by passing a no-op
LlamaModel.setLogger(LogFormat.TEXT, (level, message) -> {});

The LogLevel enum values passed to the callback correspond to the native llama.cpp log levels:

Value Meaning
DEBUG Verbose diagnostic output
INFO Informational messages about model loading and inference
WARN Non-fatal warnings
ERROR Errors that may affect inference results

Importing in Android

Important

Minimum Android version: API 28 (Android 9.0 Pie). Devices running Android 8.1 (API 27) or earlier are not supported.

Option 1 (recommended): the llama-android AAR from Maven Central

One dependency line in Android Studio — no submodule, no NDK build, no manual ProGuard rules:

dependencies {
    implementation("net.ladenthin:llama-android:5.2.0")
    // additionally, for Qualcomm Adreno GPUs (device must provide an OpenCL ICD) -- the OpenCL
    // backend module; it depends on llama-android, so the line above may also be left out:
    // implementation("net.ladenthin:llama-android-opencl:5.2.0")

    // optional Kotlin coroutines facade (Flow streaming + suspend wrappers):
    implementation("net.ladenthin:llama-kotlin:5.2.0")
}

The AAR carries the full net.ladenthin:llama Java API, the CI-built native library with ggml's CPU backend modules (one per ARM feature level, the best for the device loaded at start) for arm64-v8a (devices) and x86_64 (Android Studio emulator, Chromebooks — app bundles split per ABI so phones download only arm64), all 16 KB page-size compliant, consumer R8/ProGuard rules (applied automatically), and a manifest minSdkVersion 28 that AGP enforces against your app. The OpenCL AAR adds the libggml-opencl.so module, which the same start loads when the device has an OpenCL ICD and skips otherwise. CI boots an x86_64 emulator and runs real on-device inference against every AAR build. Do not also depend on the desktop net.ladenthin:llama JAR in the same app — the AAR already contains those classes, and the JAR would drag ~70 MB of desktop natives into your APK. See llama-android/README.md and llama-kotlin/README.md for details.

Runnable example app — "LLM Service". A minimal, KISS, fully-offline on-device chat app — pick a GGUF from the file system, then chat with it, tokens streaming into a Jetpack Compose UI, with a 13-language flag picker and private local save/load — lives in android-llmservice/ (net.ladenthin.android.llmservice). It builds with plain Gradle/AGP (no Android Studio required), produces a Play-shaped signed .aab, and is validated in CI by a real on-device emulator UI test. See its README for the build, signing/Play, and testing walkthrough.

Option 2 (advanced): build from source inside your app

Use this only if you need to patch the native layer or build for an ABI this project does not ship.

  1. Add java-llama.cpp as a submodule in your an droid app project directory
git submodule add https://github.com/bernardladenthin/java-llama.cpp 
  1. Declare the library as a source in your build.gradle
android {
    val jllamaLib = file("java-llama.cpp")

    // Execute "mvn compile" in the llama/ core module if its target/ doesn't exist
    // (the repository root is the Maven reactor aggregator; the native core lives in llama/).
    if (!file("$jllamaLib/llama/target").exists()) {
        exec {
            commandLine = listOf("mvn", "compile")
            workingDir = file("java-llama.cpp/llama/")
        }
    }

    ...
    defaultConfig {
	...
        externalNativeBuild {
            cmake {
		// Add an flags if needed
                cppFlags += ""
                arguments += ""
            }
        }
    }

    // Declare c++ sources
    externalNativeBuild {
        cmake {
            path = file("$jllamaLib/CMakeLists.txt")
            version = "3.22.1"
        }
    }

    // Declare java sources
    sourceSets {
        named("main") {
            // Add source directory for java-llama.cpp
            java.srcDir("$jllamaLib/src/main/java")
        }
    }
}
  1. Exclude net.ladenthin.llama in proguard-rules.pro
keep class net.ladenthin.llama.** { *; }

TODO

Open work items live in TODO.md.

  • Expand PIT mutation-testing scope. PIT is wired in pom.xml and runs on every CI build (in the test-java-linux-x86_64 job) with <mutationThreshold>100</mutationThreshold>. <targetClasses> currently covers net.ladenthin.llama.value.*, exception.*, args.* and four json parsers (295 mutations, 100% killed, hermetic — no model or fixture needed); widen it incrementally as additional classes reach mutation-test parity. Final target: <param>net.ladenthin.llama.*</param> matching the streambuffer pattern.

Feature Ideas

Forward-looking ideas being tracked for this fork:

  • Adopt feature ideas from the Kotlin Llama Stack client. Candidates (multimodal image input, typed chat messages, async API, batch inference, typed usage/timings) are inventoried with effort estimates in docs/feature-investigation-llama-stack-client-kotlin.md, derived from ogx-ai/llama-stack-client-kotlin.
  • Ship a directly Android-capable artifact — DONE. net.ladenthin:llama-android / llama-android-opencl (AAR, arm64-v8a, minSdk 28, consumer ProGuard rules, 16 KB page-size compliant) plus the optional net.ladenthin:llama-kotlin coroutines façade ship from this repo — see Importing in Android. Typed image input for VLMs is covered by ContentPart.imageBytes(...) / imageFile(...) (see the multimodal section), so downstream Android projects can drop their dependency on ogx-ai/llama-stack-client-kotlin entirely. A dedicated KISS example app — "LLM Service" (SAF model picker + Compose streaming chat, 13-language flag picker, private local save/load, plain Gradle/AGP, signed .aab, on-device emulator UI test) — ships in android-llmservice/.
  • Resolve all upstream kherud/java-llama.cpp open issues. All 37 open issues at fork time are catalogued with per-issue verdicts in docs/history/49be664_open_issues.md; fixes land in this fork as they are completed. Vision inputs (issues #103 and #34) are now wired end to end through blocking, typed, streaming, and OpenAI-compatible request surfaces.

Troubleshooting

Windows: EXCEPTION_ACCESS_VIOLATION with msvcp140.dll

A crash of this shape was reported against early releases:

EXCEPTION_ACCESS_VIOLATION (0xc0000005) at pc=0x00007ffa8f4b2f58
C [msvcp140.dll+0x12f58]

This cannot come from this library any more, and the old advice here — deleting msvcp140.dll from your JDK — is obsolete. Do not do it. The Windows natives link the C++ standard library and vcruntime statically (hybrid CRT), so they do not import msvcp140.dll, vcruntime140.dll or vcruntime140_1.dll at all; the only non-OS import left is the Universal CRT, which Windows itself provides and keeps updated. Measured on the current build, the whole import table of jllama.dll is KERNEL32, ADVAPI32, SHELL32, WS2_32 and the api-ms-win-crt-* forwarders. CI enforces that per release (.github/buildcheck/nativedeps.py holds every Windows CPU directory to an exact allowlist).

The report dates from before that: the troubleshooting note was written on 2026-04-04 and the switch from the DLL runtime (/MD) to a static one landed on 2026-05-13, so releases from 5.0.0 on are unaffected. An older JDK does ship its own outdated msvcp140.dll next to java.exe, and Windows resolves an import from the already-loaded module list before searching any directory -- which is why a library that did import it got the JDK's copy. Removing the dependency fixes that by construction, where moving files around could not.

If you still see a crash inside msvcp140.dll, it originates in another native library loaded into the same JVM, not in jllama.dll -- check the rest of the hs_err frame list. Please open an issue with that file rather than editing your JDK installation.

Windows: unsigned DLLs on a WDAC / AppLocker-managed client

The native libraries are not Authenticode-signed (upstream llama.cpp ships its DLLs unsigned too), and LlamaLoader extracts them into the temp directory on first use. A client whose application-control policy (WDAC, AppLocker) blocks unsigned DLLs, or any DLL under %TEMP%, refuses the load with UnsatisfiedLinkError. The way out is a directory the policy allows: put the contents of the natives jar's net/ladenthin/llama/Windows/x86_64/cpu/ directory there and start the JVM with -Dnet.ladenthin.llama.lib.path=<that directory>, which loads from it without extracting anything. (The project's release-signing key is an OpenPGP key, which signs the jars with detached .asc files; Windows does not look at those, so the DLLs would need a separate code-signing certificate.)

Contributors: build with Temurin 21, not Oracle JDK 21

llama/pom.xml passes -XDaddTypeAnnotationsToSymbol=true to javac (NullAway's JSpecify mode needs it below JDK 22). Oracle JDK 21 rejects that flag: mvn compile dies in an Error Prone IllegalStateException naming it (measured on 21.0.9). Eclipse Temurin 21, which CI uses, compiles the same tree; mvn -version shows which JDK Maven runs on.

Contributors: do not upgrade jqwik past 1.9.3

⚠️ DO NOT UPGRADE jqwik past 1.9.3. jqwik 1.10.0 added an anti-AI prompt-injection string to test stdout; the 1.10.1 user guide states the library "is not meant to be used by any 'AI' coding agents at all." 1.9.3 is the last pre-disclosure release and is the pinned version. See CLAUDE.md section "jqwik prompt-injection in test output" for the full context. Dependabot is configured to ignore all net.jqwik updates (every version, including patches) — see the ignore rule in .github/dependabot.yml.

Similar Projects / Usage

Bindings / wrappers

  • kherud/java-llama.cpp — the upstream Java binding this project was forked from (see the note at the top of this README); development continues independently here, with the fork-time upstream issues catalogued in docs/history/49be664_open_issues.md.
  • llamacpp4j — alternative Java/JNI binding to llama.cpp (SWIG-generated facade); pre-GGUF, dormant since 2023 but historically the other Java JNI option.
  • llama-cpp-python — the Python llama.cpp binding; the de-facto feature benchmark among llama.cpp bindings (server mode, multimodal, speculative decoding).
  • LLamaSharp — C#/.NET llama.cpp binding with per-backend runtime packages (CPU/CUDA/Vulkan/Metal), the .NET analogue of this project's classifier matrix.
  • node-llama-cpp — Node.js/TypeScript llama.cpp binding (prebuilt binaries, JSON-schema-constrained output, function calling).
  • LLaMAndroid — Android app demonstrating usage of llama.cpp bindings.
  • llama-stack-client-kotlin — Kotlin client for the Llama Stack API with an ExecuTorch-backed local-inference path (the llama-android AAR + llama-kotlin façade cover the same on-device ground natively).
  • llama.cpp-android-tutorial — Step-by-step tutorial for running llama.cpp on Android.

Other local inference stacks (no llama.cpp JVM binding)

  • Ollama — llama.cpp-based local model runner with its own HTTP API and model registry. This project's OpenAI-compatible server implements the Ollama-native API surface (/api/version, /api/tags, /api/show, /api/chat, /api/generate), so Ollama-speaking clients (e.g. VS Code Copilot's Ollama provider) work against an in-process jllama model.
  • ExecuTorch — PyTorch's on-device inference runtime (.pte models, XNNPACK/NPU delegates); the engine behind llama-stack-client-kotlin's local mode and the main non-llama.cpp alternative for Android on-device inference (GGUF is not supported there — different model format ecosystem).

Pure-Java single-model inference (no JNI / no llama.cpp) — Alfonso² Peterssen's *.java family of standalone, dependency-free Java inference runtimes, one per model architecture. Useful when JNI is unavailable (e.g. some sandboxes / GraalVM native-image scenarios) or when you want a single jar with no native side at all. Different design point from this project, which prioritises GGUF compatibility and llama.cpp performance via JNI.

Pure-Java inference engines (no JNI / no llama.cpp)

  • Jlama — a full pure-Java LLM inference engine for the JVM (multiple model architectures, quantization, and distributed inference) built on the Java Vector API. A no-native alternative to the JNI approach here; different design point (pure JVM portability vs. GGUF compatibility and llama.cpp performance via JNI).

Frameworks / orchestration

  • LangChain4j — LLM-application framework for Java (chat, embeddings, RAG, tool calling, agents) over a unified provider API. This project ships a first-class in-process integration — see the llama-langchain4j module — so a llama.cpp model plugs straight into LangChain4j's ChatModel / StreamingChatModel / EmbeddingModel / ScoringModel without an HTTP hop.

About

Java Bindings for llama.cpp - A Port of Facebook's LLaMA model in C/C++

Resources

Code of conduct

Contributing

Security policy

Stars

16 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages