Skip to content

Running local LLMs #176

Description

@ndrean

Brief overview of Ollama, vLLM and SGLang

To use open-weight models on your machine, you have three main options: Ollama, vLLM, and SGLang.
Each engine handles requests differently. The diagram below shows the differences and the main techniques behind each engine.

Image
  • Ollama: Ollama is best for local dev, prototyping, and laptop-scale hardware. The architecture is inherently sequential. A local user calls the OpenAI-compatible API, and requests line up in a FIFO queue. Then Ollama runs a pre-quantized GGUF model, a compressed format it pulls, and the response comes back to the user.

  • vLLM: vLLM is best for high-traffic serving, max GPU utilization, and thousands of concurrent requests. Many users hit the server at once, and continuous batching slots new requests into the running batch instead of making them wait for it to finish. PagedAttention stores the KV cache, the memory a model keeps for tokens it has already processed. The PagedAttention maps the OS memory pages to vLLM memory blocks.

  • SGLang: SGLang is best for AI agents and tool loops, multi-turn chats, and JSON/regex outputs. The most commun example is when the workflow involves using repeated context, like a static system prompt or a large RAG documents, during a CI workflow. Agents and multi-turn chats send requests whose prompts overlap heavily. A prefix-aware scheduler routes them through the RadixAttention cache, a radix tree that reuses every shared prefix instead of recomputing it.

Note

This is for the big picture. Needs to be continued with a bit more in deep insights...

Activity

  1. added
    enhancementNew feature or enhancement of existing functionality
    discussShare your constructive thoughts on how to make progress with this issue
    AiAny issue that relates to Ai, LLMs and Machine Learning.
    on Aug 27, 2026
  2. ndrean commented on Sep 5, 2026

    @ndrean
    Author

    Alternative: Magnitude

    They say:

    Why not just have my agent set up Ollama?:
    Your agent would be guessing. It doesn't know your hardware, which quant fits, or how fast it'll run. Magnitude gives it a catalog with recommendations computed for your machine, an onboarding flow that writes your harness config, and inference built for agent workloads. Models load just in time and unload when idle or memory gets tight.

    Image

    Let's see:

    pnpm add -g @magnitudedev/cli

    Then:

    magnitude setup

    An example of what my machine accepts (I selected "smartest" to see).
    .
    Image

    Not sure what "Text Vision" means.. so more models? =>

    Yes. You can download compatible GGUF models from Nvidia/Hugging Face

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    AiAny issue that relates to Ai, LLMs and Machine Learning.discussShare your constructive thoughts on how to make progress with this issueenhancementNew feature or enhancement of existing functionality

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions