KoboldCpp logo
AI · self-hosted

KoboldCpp in one click.

KoboldCpp is a run gguf llms with a built-in web ui you can run on your own machine. Launch it in Even Stacks in one click, serve it over trusted HTTPS, and let your AI agents operate it.

What is KoboldCpp?

KoboldCpp is a single-file application for running GGUF-format language models locally using llama.cpp. It bundles a web UI for chat and story generation with extensive sampler and context configuration options.

CategoryAI
Default port5002
Container imagekoboldai/koboldcpp:latest

Run KoboldCpp in Even

Even Stacks launches KoboldCpp as a managed container on your own machine. No compose files, no manual setup.

  1. Open the Even Stacks panel in Even. The container engine starts on demand.
  2. Find KoboldCpp in the one-click services and click Launch.
  3. Even pulls the image, starts it on port 5002, and serves it over trusted HTTPS at https://koboldcpp.localhost.

Drive KoboldCpp with your AI agents

Even ships an MCP, so any agent you run inside Even (Claude Code, Codex, and more) can operate KoboldCpp directly, at both the UI level and the container level.

Through the Even MCP an agent can pull a model, call the local endpoint, and wire it into a tool, all on your own hardware. It runs commands inside the container, reads its logs, and for web apps opens and clicks through the interface in Even's own browser pane. Ask once, for example "set up KoboldCpp and get it ready", and the agent handles it end to end.

How agents reach it: Chat via the built-in UI at http://localhost:5002 or call the OpenAI-compatible API at /v1/chat/completions. A GGUF model file must be mounted into /models and referenced with the --model flag at container startup. The KoboldAI API is also available at /api/v1/generate.

More AI services

Other one-click services you can launch in Even Stacks.