Circuit‑Breaking AI Agents: A Local‑First Safety Gate for Shells, MCP, and SQL
About this session
AI agents are being given shell access, database connections, and MCP tool calls but most runtimes still treat “run this command” as a black box. That’s how you get rm -rf /, curl | sh, or DROP TABLE slipping through. I’ll walk through Agent Circuit Breaker (ACB), an open‑source, local‑first runtime that inspects what an agent is about to do and returns a deterministic decision ALLOW, BLOCK, PENDING_APPROVAL, UNKNOWN, or ERROR before execution.
I’ll show how ACB:
normalizes inputs to defeat homoglyph/zero‑width evasion, parses shell commands, filesystem ops, and SQL with dedicated inspectors, runs an ordered rule set (first‑match‑wins) for fast, auditable decisions, and enforces trajectory‑level policies across multi‑step agent runs (scope limits, forbidden targets, egress checks, and secret‑exfil detection). I’ll include real payloads that ACB blocks, show how we plug it into CI via pre‑commit hooks and SARIF, and wrap with a live demo so you can try it on your own agent loops. The goal: you’ll leave with a reusable guardrail you can drop into coding agents, MCP servers, and long‑horizon workflows today.
Speaker
Key takeaways
- How to add a deterministic, pre‑execution safety gate to AI agents without slowing down developer experience (using open‑source ACB as a reference architecture).
- Concrete patterns to inspect and block risky shell/SQL/MCP actions—including homoglyph/normalization tricks—so you’re not relying on prompts alone.
- A ready‑to‑reuse integration pattern: plug ACB into CI via SARIF and pre‑commit hooks, and wire it in front of MCP servers for tool‑call safety.