[Unit] Description=vLLM — self-hosted coder model for the clone agent (on-prem RTX 5090). Inference appliance only. After=network-online.target Wants=network-online.target [Service] User=ai-admin Environment=HF_HOME=/opt/vllm/hf Environment=CUDA_HOME=/usr/local/cuda Environment=PATH=/opt/vllm/venv/bin:/usr/local/cuda/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin # # HARDWARE: RTX 5090 — Blackwell, sm_120, 32GB GDDR7 (~31.8GB usable), ~1.8TB/s bandwidth. # NOTE this card has LESS VRAM than the 48GB L40S it replaces (32 < 48), but ~2x the bandwidth — # so it is FASTER, not bigger. Nothing larger than a ~30B INT4 fits. A 70B AWQ (~40GB) does NOT. # # MODEL: Qwen3-Coder-30B-A3B-Instruct, AWQ INT4 (~17GB). MoE with only 3B active params => very # fast decode, and it is purpose-built for agentic TOOL CALLING (qwen3_coder parser) — which is the # whole architecture here: the model drives small, targeted edits via agent.py's tools, it never # regenerates a page. A reasoning distill (R1) is the wrong tool for this job. # • --served-model-name exposes BOTH aliases: "devstral" (legacy — keeps .env/openhands config # working unchanged) and "qwen3-coder" (what it actually is). # # CONTEXT: 65536 fits despite the smaller card. Qwen3-30B-A3B uses GQA with only 4 KV heads # (2*48 layers*4 kv*128 dim*2B = ~0.094MB/token), so a full 64k sequence is only ~6GB of KV. # Budget @ 0.90 util: ~28.6GB - ~17GB weights - ~2GB activations/graphs = ~9.6GB KV pool (~100k # tokens). If it OOMs on boot, lower --gpu-memory-utilization first, then --max-model-len. # # BLACKWELL GOTCHA: sm_120 needs a CURRENT vLLM + PyTorch built for CUDA 12.8+. The old # "vllm>=0.6.0" pin will NOT run on this card. Driver 580.x (CUDA 13 capable) is already installed. # FP8 is native on Blackwell, but an FP8 30B (~30GB) leaves no room for KV — so AWQ INT4 it is. # # Bind to 127.0.0.1 ONLY. The box is behind office NAT; Vultr reaches it via the REVERSE autossh # tunnel this box dials out (see onprem-tunnel.service) — never expose 8000 to the internet. ExecStart=/opt/vllm/venv/bin/vllm serve stelterlab/Qwen3-Coder-30B-A3B-Instruct-AWQ --served-model-name devstral qwen3-coder --tool-call-parser qwen3_coder --enable-auto-tool-choice --enable-prefix-caching --trust-remote-code --host 127.0.0.1 --port 8000 --max-model-len 65536 --gpu-memory-utilization 0.90 --max-num-seqs 4 Restart=always RestartSec=15 TimeoutStartSec=1800 StandardOutput=journal StandardError=journal [Install] WantedBy=multi-user.target