[bmdpat]

/* ------------------------------------------------ */

BMD

// build. manage. deploy.

*/

A real shell running in your browser. Type desk to size your GPU, or tap a command. The bench replay streams at the measured rate. It is a replay, not a live model.

228.9 tok/s

llama3.1:8b Q4_K_M · RTX 5090 · measured 2026-07-09

[ Read the 5090 Reports ]

pat@5090:~$ ls tools/

§ 001 / BUILD

Small tools that answer one question each.

pat@5090:~$ agentguard --status

§ 002 / MANAGE

Nothing runs without a ceiling on it.

AgentGuard stops spend, loops, timeouts, and rate-limit failures before a run touches real work. Four guards, each with one job.

$ pip install agentguard47
  • cost + count

    • max_cost_usddollars
    • max_callscount
  • repeats

    • max_repeatscount
  • wall clock

    • max_secondsseconds
  • per minute

    • max_calls_per_minutecount
Open AgentGuard docs

pat@5090:~$ ps aux | grep agents

§ 004 / OPERATING LOOP

One person. Small tools. Agent-assisted ops.

01

Run

Run the model on the target GPU.

02

Measure

Record tokens, latency, VRAM, cost, and failure mode.

03

Publish

Write up the result or ship the tool it required.

pat@5090:~$ tail blog.log

§ 003 / DEPLOY

Everything ships in public, working or not.

§ 005 / BUILD NOTESloading

Build notes.

Loading the latest notes…

pat@5090:~$ mail -s "subscribe" pat@5090

To: the journal // Subj: subscribe

Get the journal by email.

New benchmark results, failed runs, VRAM checks, and tools. I send an email only when there is a real result.

Get the Local AI Field Kit

Four copy-ready tools now, then one evidence-backed Local AI Lab Note on Friday when there is something worth sharing.

Get the requested artifact now, then at most one evidence-backed Local AI Lab Note on Friday when there is something worth sharing. One-click unsubscribe. No sponsored placements. Privacy.

[bmdpat] 0:desk*VRAM 7.8/32.0 GB