llm-watch

apollo · Qwen3.8-27B-UD-IQ4_XS.gguf
generating

Model

Context used
—
Generation
— tok/s
Prompt
— tok/s
Draft accepted
—

GPU

Utilisation
—
VRAM
—
Junction temp
—
Power
—

Host · live, last 3 min

CPU
—
Memory
—
GPU history
—
Power
—

Token usage · the last 24 hours

generated prompt processed prompt from cache

Throughput and cache over the same range

Generation
— tok/s avg
Prompt processing
— tok/s avg
Cache hit
—
Draft accepted
—

All time

llama.cpp log

loading…

Running on apollo