pat@5090:~$ desk
See what your GPU can actually run.
Choose your card. Get the models that fit, the quant to use, and a command you can paste.
Essential cookies keep the site working. Optional analytics are off unless you allow them.
Review cookie policy/* ------------------------------------------------ */
BMD
// build. manage. deploy.
*/
pat@5090:~$ desk
Choose your card. Get the models that fit, the quant to use, and a command you can paste.
A real shell running in your browser. Type desk to size your GPU, or tap a command. The bench replay streams at the measured rate. It is a replay, not a live model.
228.9 tok/s
llama3.1:8b Q4_K_M · RTX 5090 · measured 2026-07-09
pat@5090:~$ ls tools/
§ 001 / BUILD
13 live / 2 beta from 15 public tools. Open what helps.
$ pip install agentguard47$ pip install agentguard47pat@5090:~$ agentguard --status
§ 002 / MANAGE
AgentGuard stops spend, loops, timeouts, and rate-limit failures before a run touches real work. Four guards, each with one job.
$ pip install agentguard47cost + count
repeats
wall clock
per minute
pat@5090:~$ ps aux | grep agents
Run the model on the target GPU.
Record tokens, latency, VRAM, cost, and failure mode.
Write up the result or ship the tool it required.
pat@5090:~$ tail blog.log
§ 003 / DEPLOY
Loading the latest notes…
pat@5090:~$ mail -s "subscribe" pat@5090
To: the journal // Subj: subscribe
New benchmark results, failed runs, VRAM checks, and tools. I send an email only when there is a real result.
Four copy-ready tools now, then one evidence-backed Local AI Lab Note on Friday when there is something worth sharing.
Get the requested artifact now, then at most one evidence-backed Local AI Lab Note on Friday when there is something worth sharing. One-click unsubscribe. No sponsored placements. Privacy.