- Shell 58.7%
- Dockerfile 17.7%
- Python 15.7%
- Jinja 7.9%
| .github | ||
| autocomplete | ||
| coding | ||
| docs | ||
| imagegen | ||
| infra | ||
| logs | ||
| multimodal/audio-sfx/__pycache__ | ||
| scripts | ||
| thunder-llm-monitoring@1cc6668012 | ||
| .env | ||
| .gitignore | ||
| .gitmodules | ||
| CLAUDE.md | ||
| llamacpp-cuda-phi2.Dockerfile | ||
| README.md | ||
Thunder LLM
A self-hosted LLM server instance for use with Pi and other AI coding agents.
Purpose: Reliable, high-throughput local inference for development workflows. Server endpoint:
http://localhost:8080(llama.cpp, ROCm backend)
Server
Current configuration:
| Component | Value |
|---|---|
| Model | Qwen3.6-27B-Q6_K (GGUF) |
| Hardware | AMD RX 9700 (32 GB GDDR6) |
| Backend | llama.cpp (Vulkan, --reasoning on, MTP spec decoding) |
| Build | b9209 |
| Context | 256K |
| Speculative decoding | draft-MTP (MTP speculative decoding) |
| Reasoning | --reasoning on, budget 16384, client-controlled via chat_template_kwargs |
| Server default | --chat-template-kwargs '{"preserve_thinking": true}' (server-side default for thinking preservation) |
| Throughput | ~99 tps decode (Vulkan) |
| Container | Docker, managed via docker compose |
Health: curl -s http://localhost:8080/health
Available Services
Coding (port 8080)
Each compose file in coding/compose/ defines a single service with a permutation name: <backend>-<device>-<model-spec>. Switch between them with ai-server coding <engine>.
| Service | Backend | Model | Notes |
|---|---|---|---|
llamacpp-rocm-qwen36-35b-q5 |
llama.cpp ROCm | Qwen3.6-35B-A3B-UD-Q5_K_XL | Fallback — ~66 tps, no GUI impact |
llamacpp-rocm-qwen36-27b-q6 |
llama.cpp ROCm | Qwen3.6-27B-Q6_K | 27B dense model on ROCm |
llamacpp-rocm-qwen36-27b-q6-mtp |
llama.cpp ROCm (MTP) | Qwen3.6-27B-Q6_K | 27B with draft-MTP speculative decoding |
llamacpp-rocm-qwen36-27b-q4 |
llama.cpp ROCm | Qwen3.6-27B-Q4 | Smallest 27B variant |
llamacpp-vulkan-qwen36-35b-q5 |
llama.cpp Vulkan | Qwen3.6-35B-A3B-UD-Q5_K_XL | Faster decode (~99 tps) but GUI sluggish during heavy PP |
llamacpp-vulkan-qwen36-27b-q6 |
llama.cpp Vulkan | Qwen3.6-27B-Q6_K | 27B dense on Vulkan |
llamacpp-vulkan-qwen36-27b-q6-mtp |
llama.cpp Vulkan (MTP) | Qwen3.6-27B-Q6_K | Default daily driver — ~99 tps, MTP speculative decoding |
vllm-rocm-qwen36-35b-q5 |
vLLM ROCm | Qwen3.6-35B-A3B-UD-Q5_K_XL | Standard vLLM ROCm |
vllm-rocm-qwen36-35b-mxfp4 |
vLLM ROCm (MXFP4) | Qwen3.6-35B-A3B-MXFP4 | RDNA4-specific kernels, custom Dockerfile |
vllm-rocm-qwen36-35b-kyuz0 |
vLLM ROCm (patched) | Qwen3.6-35B-A3B-MXFP4 | kyuz0's gfx1201 patch with AITER unified attention |
vllm-rocm-qwen36-35b-tcclaviger |
vLLM ROCm (patched) | Qwen3.6-35B-A3B-MXFP4 | tcclaviger's 1,980 AMD-specific patches |
Why Vulkan MTP is the default: Vulkan provides faster decode throughput (~99 tps vs ~66 tps ROCm). The MTP speculative decoding further boosts effective speed. ROCm remains available as a fallback when GPU display contention is a concern.
Autocomplete (port 8081)
Ghost-text autocomplete for code editors. Managed independently from the coding server — can run alongside it.
| Service | Backend | Model | Notes |
|---|---|---|---|
llamacpp-rocm-qwen25-coder-15b-q4 |
llama.cpp ROCm | Qwen2.5-Coder-1.5B-Q4_K_M | Small, fast completions (~980MB) |
CLI: ai-server ac <engine> to switch, ai-server ac status to check health.
Image Generation (port 8188)
ComfyUI with GGUF diffusion models. Cannot run simultaneously with the coding server — both compete for the 32 GB VRAM. The switch script stops the coding server automatically.
| Service | Backend | Model | Notes |
|---|---|---|---|
comfyui-rocm |
ComfyUI + ComfyUI-GGUF (ROCm PyTorch) | Qwen-Image-2.1 Q4_K_M GGUF | Primary — uses ROCm PyTorch for diffusion |
comfyui-vulkan |
ComfyUI + ComfyUI-GGUF (CPU PyTorch + Vulkan) | Qwen-Image-2.1 Q4_K_M GGUF | Test — CPU PyTorch, Vulkan via llama.cpp |
Models (stored on Games drive, ~10.5 GB total):
| Model | File | Size | Purpose |
|---|---|---|---|
| Diffusion (DiT) | qwen_image_2.1_Q4_K_M.gguf |
4.0 GB | Core image generation |
| Text Encoder | qwen3vl_8b_w4a8.safetensors |
5.9 GB | Required — DiT embedding space is fixed |
| VAE | qwen_image_2.1_vae_bf16.safetensors |
645 MB | Latent space decode |
CLI: ai-server img <engine> to switch (stops coding server), ai-server img status, ai-server img down.
⚠️ VRAM constraint: Coding server (~14 GB VRAM) + image gen (~15+ GB VRAM) > 32 GB available. Must stop one before starting the other.
Container Health Monitoring
An idle-aware monitor automatically checks TPS and restarts the coding container when performance degrades. Prevents the ~30–40 tps slowdown that can occur after 6–8 hours of continuous uptime.
How it works:
- Runs every 2 hours via cron (
0 */2 * * *) - Benchmarks the container (5 warm-up runs + 3 measured, median TPS)
- Logs TPS on every check — active, idle, fresh, old
- Only restarts when: idle (no requests in 10 min) + uptime ≥ 4 hours + TPS < 70
- Never restarts while you're actively using the service
- Logs:
logs/monitor.log - Monitor binary:
thunder-llm-monitoringsubmodule (AOT-compiled .NET 10.0)
Manual check:
/home/james/.thunder-llm/Thunder.Llm.Monitor tps-monitor
Monitoring
Two subsystems track server health and GPU telemetry:
GPU Wattage Monitor (systemd)
Continuous GPU telemetry — 1-second resolution logging to ~/.thunder-llm/logs/wattage-YYYY-MM-DD.csv:
timestamp, watts, temp_edge_c, temp_junction_c, fan_pct, gpu_use_pct, vram_pct- Service:
monitor-wattage.service(enabled, runs at boot) - Log:
sudo journalctl -u monitor-wattage -f
TPS Health Monitor (cron)
Every 2 hours: benchmarks the LLM server (5 warmup + 3 measured, median TPS). If median TPS < 70 and container is ≥ 4 hours old, restarts the container.
- Log:
~/.thunder-llm/logs/tps-YYYY-MM-DD.csv - Old files are downsampled daily (10-second intervals) to save space
Manual run:
/home/james/.thunder-llm/Thunder.Llm.Monitor tps-monitor
Monitoring Source
The monitor binary is an AOT-compiled .NET 10.0 app in the thunder-llm-monitoring submodule:
cd thunder-llm-monitoring && ./deploy.sh
Bench CLI
The bench CLI has been moved to a dedicated repository: thunder-llm-bench.
See code.thundersizzle.tech/Thunder/thunder-llm-bench for the full CLI, docs, and history.
Quick benchmark (server-side):
# Raw TPS benchmark — no dependencies, just curl against the server
bash docs/work/bench-llamacpp.sh [warmups] [measured]
Quick Start
Global CLI
The ai-server command is available system-wide (symlinked to scripts/ai-server). It dispatches to category-specific scripts:
ai-server status # show status of all categories
ai-server coding <engine> # switch coding engine (with rollback)
ai-server coding --force # force restart coding engine
ai-server coding status # coding health check
ai-server coding logs # follow coding logs
ai-server coding engines # list available coding engines
ai-server coding down # stop all coding engines
ai-server ac <engine> # switch autocomplete engine (with rollback)
ai-server ac --force # force restart autocomplete engine
ai-server ac status # autocomplete health check
ai-server ac logs # follow autocomplete logs
ai-server ac engines # list available autocomplete engines
ai-server ac down # stop all autocomplete engines
ai-server img <engine> # switch image generation engine (stops coding server)
ai-server img --force # force restart image generation engine
ai-server img status # image generation health check
ai-server img logs # follow image generation logs
ai-server img engines # list available image generation engines
ai-server img down # stop all image generation engines
Run the default coding server (ROCm)
ai-server coding llamacpp-rocm-qwen36-35b-q5
Run an alternative coding backend
ai-server coding llamacpp-vulkan-qwen36-35b-q5 # Vulkan backend
ai-server coding vllm-rocm-qwen36-35b-kyuz0 # vLLM RDNA4 experimental
Start autocomplete
ai-server ac llamacpp-rocm-qwen25-coder-15b-q4
Note: Each compose file defines exactly one service. The ai-server command wraps category-specific switch-engine.sh scripts with automatic health checks and rollback on failure.
Run a quick benchmark (server-side)
# Raw TPS benchmark — no build step, just curl against the running server
bash docs/work/bench-llamacpp.sh 5 3
Full bench CLI
For the complete benchmarking suite (TPS, context probes, agentic evaluation, model discovery, etc.), use the dedicated thunder-llm-bench repo:
git clone ssh://git@code.thundersizzle.tech:2222/Thunder/thunder-llm-bench.git
cd thunder-llm-bench
dotnet build -c Release src/Thunder.LlmBench/Thunder.LlmBench.csproj
alias bench="dotnet run --project src/Thunder.LlmBench/ --"
Repository Structure
.
├── infra/ # Shared build assets
│ ├── DockerFiles/ # Dockerfiles (shared across categories)
│ │ ├── llamacpp-stable-roc.Dockerfile
│ │ ├── llamacpp-stable-vulkan.Dockerfile
│ │ ├── llamacpp-rocm.Dockerfile
│ │ ├── llamacpp-vulkan.Dockerfile
│ │ ├── llamacpp-unstable-roc.Dockerfile
│ │ ├── llamacpp-unstable-vulkan.Dockerfile
│ │ ├── vllm-rocm.Dockerfile
│ │ ├── vllm-mxfp4.Dockerfile
│ │ ├── comfyui-rocm.Dockerfile
│ │ └── comfyui-vulkan.Dockerfile
│ └── third-party/ # LLaMA.cpp submodule (build dependency)
│ ├── llama.cpp-stable/
│ └── llama.cpp-unstable/
│
├── coding/ # Main chat inference (port 8080)
│ ├── docker-compose.yaml # name: thunder-llm
│ ├── compose/ # 11 engine configs
│ │ ├── llamacpp-rocm-*.yaml
│ │ ├── llamacpp-vulkan-*.yaml
│ │ └── vllm-rocm-*.yaml
│ ├── scripts/
│ │ ├── compose-wrapper.sh # Resolves GPU device paths
│ │ └── switch-engine.sh # Engine switching with rollback
│ └── qwen3.6/ # Chat template
│
├── autocomplete/ # Ghost-text autocomplete (port 8081)
│ ├── docker-compose.yaml # name: thunder-llm-ac
│ ├── compose/
│ │ └── llamacpp-rocm-qwen25-coder-15b-q4.yaml
│ └── scripts/
│ ├── compose-wrapper.sh
│ └── switch-engine.sh
│
├── imagegen/ # ComfyUI image generation (port 8188)
│ ├── docker-compose.yaml # name: thunder-imagegen
│ ├── compose/
│ │ ├── comfyui-rocm.yaml # ROCm PyTorch backend
│ │ └── comfyui-vulkan.yaml # CPU PyTorch + Vulkan
│ └── scripts/
│ ├── compose-wrapper.sh
│ └── switch-engine.sh
│
├── scripts/ai-server # Global CLI (symlink to /usr/local/bin/ai-server)
├── scripts/load-test.py # Load testing / stress testing
├── thunder-llm-monitoring/ # GPU wattage + TPS monitoring (submodule)
│ ├── deploy.sh # AOT publish, systemd deploy, cron setup
│ └── src/Thunder.Llm.Monitor/ # Monitor source (AOT .NET 10.0)
├── docs/work/ # Milestone docs, benchmarks, status
│ ├── STATUS.md # Current phase/milestone state
│ ├── infrastructure-split-plan.md # Multi-category split plan
│ ├── v1/phase1-.../ # Phase and milestone documentation
│ └── bench-llamacpp.sh # Quick TPS benchmark script
├── logs/monitor.log # Health monitor logs
└── README.md # This file
Note: The
benchCLI (Thunder.LlmBench) lives in a separate repo:thunder-llm-bench. This repo contains only the server infrastructure (Dockerfiles, compose, scripts, and server docs).
Container Management
Each compose file defines exactly one service. Use the ai-server CLI for all operations:
ai-server coding llamacpp-rocm-qwen36-35b-q5 # switch coding to ROCm
ai-server coding logs # view coding logs
ai-server ac llamacpp-rocm-qwen25-coder-15b-q4 # start autocomplete
ai-server img comfyui-rocm # start image generation (stops coding)
Category isolation
Each category (coding, autocomplete) is an independent Docker Compose project:
coding/→ projectthunder-llm(port 8080)autocomplete/→ projectthunder-llm-ac(port 8081)imagegen/→ projectthunder-imagegen(port 8188)
Stopping one category does not affect the other. They share the GPU but have independent lifecycles.
Switching between backends (REQUIRED)
Always use ai-server for switching between backends. It wraps category-specific switch-engine.sh scripts, which auto-detect the running engine, health-check the new one, and roll back on failure:
- Coding rolls back to
llamacpp-rocm-qwen36-35b-q5 - Autocomplete rolls back to
llamacpp-rocm-qwen25-coder-15b-q4
Never call docker compose down/up to switch backends. Use it only for rebuilding containers, config changes, or volume operations — or if the switch script itself fails.
GPU Device Mapping
The AMD GPU's render node (/dev/dri/renderD*) can change number between reboots (e.g., renderD128 ↔ renderD129) depending on kernel device discovery order. Each category has its own compose-wrapper.sh that resolves the device at launch time:
/dev/dri/by-path/pci-0000:03:00.0-render → renderD128 (or whatever it is today)
The wrapper script resolves this symlink before Docker Compose parses the files, so the correct device is always passed regardless of kernel-assigned numbering.
After a reboot: if the renderD numbers swapped, the container's auto-restart (restart: unless-stopped) may attach the wrong device. Fix it with:
ai-server coding --force <current-engine> # or: ai-server ac --force <engine>
This stops and recreates the container with the freshly-resolved device path.
Compose wrapper
Each category has its own compose-wrapper.sh (coding/scripts/, autocomplete/scripts/). It:
- Resolves
AMD_RENDER_DEVICEfrom the PCI-stable symlink - Changes to the category directory
- Passes all arguments through to
docker compose
Both switch-engine.sh and the global ai-server CLI use it internally.
GPU memory
rocm-smi --showmeminfo vram
Reasoning / Thinking Mode
The server runs with --reasoning on and --reasoning-budget 16384 by default. The server also sets --chat-template-kwargs '{"preserve_thinking": true}' as a server-side default, so thinking output is preserved even when clients request non-reasoning mode (clients can still toggle thinking on/off).
How clients control reasoning:
-
Pi harness — uses
thinkingFormat: "qwen-chat-template"in~/.pi/agent/models.json. Specifiesllamacpp/Qwen3.6-35B-A3B-Q5_K_M:high(or:off,:low,:medium,:xhigh) in the model pattern. Automatically injects{enable_thinking: true/false, preserve_thinking: true}. -
Open WebUI (
llm.thundersizzle.tech) — always uses reasoning mode (no client-side toggle). The pi harness hides thinking output viahideThinkingBlock: true. -
Other clients — inject
{"chat_template_kwargs": {"enable_thinking": true, "preserve_thinking": true}}for reasoning, or{"chat_template_kwargs": {"enable_thinking": false}}for non-reasoning mode. The server'spreserve_thinking: truedefault still applies.
Adjusting reasoning budget:
The server default is 16384 tokens. For faster responses on simple queries, reduce in docker-compose.yaml:
--reasoning on
--reasoning-budget 8192 # lower = faster, still good for agentic work
Troubleshooting
TPS degraded to ~30–40
Container state degradation after extended uptime. The idle-aware monitor handles this automatically (every 2 hours, if idle + ≥ 4h uptime + TPS < 70). Manual restart:
ai-server coding --force llamacpp-rocm-qwen36-35b-q5
Shader cache stale after Mesa upgrade
Clear and rebuild:
rm -rf ~/.cache/mesa_shader_cache/
ai-server coding --force llamacpp-vulkan-qwen36-35b-q5
(Only applies to Vulkan backend — ROCm does not use Mesa shader caching.)
Model returning garbage tokens ("0", "&!", "3 years ago")
Caused by --reasoning off combined with --chat-template-kwargs conflict, or corrupted conversation context. Fix:
- Ensure server runs
--reasoning on(not--reasoning off) - The server defaults to
--chat-template-kwargs '{"preserve_thinking": true}'— this is intentional. Clients control reasoning viaenable_thinking. - Start a fresh conversation in the client (corrupted context from previous bad requests propagates)
Container won't start after reboot (GPU device mismatch)
The AMD GPU's render node number can change between reboots. If the container fails to start or returns errors about GPU access:
ai-server coding --force <current-engine> # e.g., ai-server coding --force llamacpp-rocm-qwen36-35b-q5
This resolves the current PCI symlink and recreates the container with the correct device.
vLLM backends failing to start
The vLLM RDNA4 backends are experimental and may fail depending on ROCm driver version. Check logs:
docker logs vllm-rocm-qwen36-35b-kyuz0 -f
docker logs vllm-rocm-qwen36-35b-tcclaviger -f
Previous Hardware
NVIDIA RTX 3090 (24 GB) — v0.x phase. Migration to AMD RX 9700 is complete.
ROCm is the recommended daily driver for the R9700. Vulkan remains available for benchmarking and when maximum decode speed is needed.