How agents actually work
part one — eight problems, one loop
hover a card to see where it attaches
{{ n.num }}
{{ n.label }}
{{ c.title }}
{{ c.num }}
{{ c.problem }}
{{ c.src }}
part two — around the loop, not in it: skills·MCP·memory·subagents·prompt injection
an agent's output is the change it made to the world
Once you take that seriously, the transcript stops being the thing you evaluate and becomes a debugging artifact. Refusal becomes a positive outcome a test can assert on. A crash mid-run becomes a question about the filesystem.