The agent auditor's local operator dashboard, showing observe-only mode and a list of recent Claude sessions with their recorded timelines

Auditing the AI Agent That Runs My Homelab

My homelab has picked up a lot of moving parts this year, and several of them are now LLMs. Claude Code runs natively on the box and does real work against real infrastructure — containers, reverse proxy config, monitoring, the lot. That’s genuinely useful. It also creates a trust problem I hadn’t had before. The problem is simple to state: the same agent that makes a change also writes the summary explaining why the change was safe. That’s convenient. It is not independent verification. If the agent quietly skips a validator and then reports “config validated”, I have no signal at all. The report is the evidence, and the thing that wrote the report is the thing being checked. ...

7 August 2026 · 9 min
Langfuse's Tracing view showing a single claude_code.interaction trace with 2 observations and 3.40s latency

Implementing Langfuse to Monitor Claude Code

We all tend to focus on the output of whatever LLM we’re using — did it get the answer right, was it fast, was it useful. What I’d stopped paying attention to was the background: how many tokens a session was actually burning, where they went, and whether I’d have any way of knowing if something had gone quietly wrong. Claude Code runs natively on my homelab now, doing real work against real infrastructure, and I wanted more than “the output looked fine” as my only signal. ...

27 July 2026 · 4 min