Technology
Minutes before Apollo 11 touched down on July 20, 1969, the guidance computer started throwing alarms. 1202. Then 1201. Armstrong radioed down asking for a reading on the 1202, because nobody in the cabin knew what it meant on sight.
The computer was overloaded. A radar configuration was feeding it work that had nothing to do with landing, and it was being asked to do more than it had time to do.
What happened next is the part worth studying. The flight software, built at the MIT Instrumentation Lab, had been designed for exactly this. When the job queue overflowed, it restarted, dropped the low-priority work, and kept the guidance running. Mission Control called go. The Eagle landed.
The system did not avoid failure. It knew in advance what to drop.

What actually breaks
I run scheduled AI agents across my companies. One syncs cleaning schedules for the rental side. One drafts invoices every Monday. Others publish blog posts, build carousels, and sort receipts. One of them publishes to this site every day at 11:11.
After months of watching them, the pattern is clear. The model almost never fails at the task. It writes the draft. It formats the invoice. The failures all live at the seams.
Access expires. A login times out, and the publishing agent sits blocked for two runs straight, doing everything right except the one thing that matters.
Quotas run dry. An image tool hit zero credits and stayed there for a week of runs. The agent adapted every time. Nobody topped it up, because nothing was on fire.
Success lies. A download returned an error page, and the upload accepted it as a valid image. Every status code said success.
Timeouts lie too. A request timed out, the agent retried, and the first request had actually landed. Now there are two of everything.
Everything fires at once. Four of my agents are scheduled for 6:00 a.m. That is not a strategy. That is what happens when every agent gets set up on a different day by someone who likes mornings.

The model is rarely the weak link. The handoff is.
The operating rules
1. Stagger the clock. One agent per slot. If two jobs touch the same system, they never share a start time. Collisions are free to prevent and expensive to debug.
2. Rank the work before the overload. Every agent should know what it drops first. A run that cannot finish should ship the essential part and say exactly what it skipped. That is the 1202 lesson, and it is the one most automation ignores.
3. Verify the artifact, not the response. Check the file size. Open the image. Read the live page after publishing. A green status code is a claim, not proof.
4. Assume the timeout worked. Look before you retry. Every retry path needs a check that asks whether the first attempt already landed.

5. Fail loudly. A silent failure is a failure plus a delay. Every agent I run ends by sending one line to my phone when something breaks, and nothing when it does not. Silence has to mean healthy, or it means nothing.
Epictetus drew the line between what is in your control and what is not. Agents make that line operational. The model is outside your control. The schedule, the priorities, the checks and the alarms are entirely inside it.
The real job
Automation does not remove the operator. It moves him. The work shifts from doing the task to designing the seams: who has access, what runs when, what gets dropped, and who hears about it.

The engineers in 1969 did not build a computer that could never be overloaded. They built one that already knew what mattered when it was.
A system is only as reliable as its least-watched handoff.
