I crashed four production AI agents mid-action. None of them recorded what already finished.
An engineer stress-tested four published agent codebases by killing them mid-action and retrying: none recorded what had already finished, so retries could repeat real-world side effects. The post argues agents need a durable ledger of completed steps written before tool calls return, plus idempotent tools. A practical reminder that crash recovery is a production requirement, not an edge case.