Last Monday, my company’s automated weekly review found six real problems. An exposed test server. Safety features shipped switched off. A dashboard that had been silently frozen for three days. The review wrote all six into a beautifully organized document, exactly as designed.
Then nothing happened. Because a document is where problems go to die.
I run Fern Capital alone. No staff. The development team is AI agents, and after two decades operating in B2B payments I can tell you the failure mode above is not an AI problem. Every operating company I have ever seen has the same drawer: the audit finding in a PDF nobody reopens, the exception log that gets appended and never drained, the “we should fix that” from Tuesday that is gone by Friday. Recording a problem feels like accountability. It is usually where accountability ends.
So last week I changed the rule. Problems found anywhere in my systems now land on one visible queue, written in plain English. And three nights a week, at 2:30 in the morning, an AI agent picks the oldest one and fixes it while I sleep.
How the machine actually runs
Everyone asks what AI can automate. That is the wrong question, and it is why most AI adoption stories collapse into demos. The question that matters is where judgment lives.
Here is the shape of it:
Three details carry all the weight.
First, the AI that writes a fix never approves it. A second AI, from a different vendor, reviews every change, and an unresolved objection blocks the merge with no override. On the night I turned this on, the reviewer caught five real design flaws in the fixer itself, including one that would have silently jammed the queue forever. The reviewer caught five flaws in the fixer. That sentence is the entire safety argument.
Second, judgment is never automated. Anything touching money movement, outbound communication, or a genuine business tradeoff is structurally off limits to the machine. It routes to me instead: one plain question or one paste-ready command, sized so I can handle it from my phone.
Third, the whole thing is bounded like an operation, not a demo. One problem per night. A hard 90 minute stop. It refuses to start when the server is short on memory, and a single empty file pauses everything. When the loop itself breaks, it files its own breakage on the same queue it drains.
What the launch night looked like
The part worth sitting with is the morning. I woke up, pasted one staged command, which took about ten seconds, and answered one question. Meanwhile the rest of my systems, following the new convention, had been filing problems on their own: the queue went from six tracked problems to fourteen in a day.
That number going up is the system working. Fourteen visible problems is a healthier company than six invisible ones.
What this did to my job
My job compressed. I do not review fixes. I do not chase status. I answer the questions machines route to me, and every one arrives with its evidence attached: what changed, what the independent reviewer said, how it was verified.
Leadership in an AI-operated company is not supervision. It is queue design: deciding what gets captured, what machines may touch, and which decisions must still cross a human desk. Get those three boundaries right and delegation stops being an act of faith.
It becomes delegation with receipts.
The diligence lens
I wrote a few weeks ago that most payments M&A diligence reads the contracts and the model and never examines the machine that produces the numbers. Here is the sharpest version of that test I now know, and it takes one meeting to run.
Ask where their known problems live.
If the answer is a worked, visible queue with an owner and a drain rate, you can audit the operation in an afternoon, and the forecast sits on something real. If the answer is a scatter of PDFs, inboxes, and tribal knowledge, you are not buying an operation. You are buying a narrative, and the seller may not even know it.
The AI part changes the economics, not the principle. Draining an exception queue used to take headcount, which is why the drawer existed in the first place. It now takes a bounded agent, an independent reviewer, and a founder willing to define where judgment lives. The technology is cheap and getting cheaper. The discipline is the scarce asset, in my one-person company and in every hundred-person portfolio operation I look at.
The queue is the job. Everything else is delegation with receipts.