Blog
Notes from the workbench.
Codex vs Claude Code: pick by task type, not the benchmark
Codex and Claude Code now score within a point of each other on SWE-bench. Three real-task tests show what actually decides it instead.
AI agent vs skill: I built 104 skills and zero agents
A skill is instructions for one task. An agent owns the loop. After building 104 skills and shipping few agents, here is when each one actually wins.