local-ai serving

gpu inference monitor
— —
–
Decode TPSlive
-
tokens/s
avg-
peak-
Prefill TPSlive
-
tokens/s
avg-
peak-
Requestslive
-
running / waiting
GPU activewindow
-
of window

Throughput

tokens/s
decode (generation)prefillleft=decode · right=prefill, auto-scaled

Concurrent streams

recently completed · short-term

Concurrency over time

runningprefillingwaiting

Thermal

°C
GPU core °CCPU package °C

Capacity · KV · context

KV occupancyestimate
-/ 100%
avg context-
KV pool-
GPU powernvidia-smi
-W now
VRAM-
util now-
util trend-
context length distribution
context × decode tok/s (color = ttft)
efficiency · decode share of GPU time
Decode 产出占比decode share of GPU time
-% of device time
Cache 抖动 / Thrashpressure
-events
context · last window
decode share of GPU time · recent trend

Requests

Cache reuse
-
Spec decoding
-

Live queue · errors

Recent errors
none

Recent requests

· newest first
timereqfinishpromptgenttftdec/swallspec

Chat · quick test

proxied to the running ninfer (:8020) · one-shot, no history

  

Raw log

newest first · live tail
gpu service · one 32 GB card, one at a time
profiles
models · download
download model
jobs
none

Serve · in-serve timers

watchdog + log rotation run inside this dashboard process
watchdog · engine autostart (every 60s)
-
log rotation · ninfer log rollover (every 6h)
-
media · async image/video jobs (gen.sh)
-