Notifications are moments, ledgers are outcomes
What we learned running a startup back office on a team of AI agents.
When I started running my business administration on Claude-powered agents in September 2026, I assumed the hard part would be getting the agents to understand Korean notices, tax invoices and card statements. It wasn't. The models read those well. The hard part was getting the agents to know when they didn't know.
Here are four mistakes from our own pilot, and the rule each one left behind. No names, no amounts.
1. The text message was older than the truth
An agent saw a card company's "payment overdue" text and reported the bill as unresolved. In another case, it saw the last billing text for a subscription and reported that I was "still being charged". Both were wrong. The ledger showed the bill had been paid the same day, for the same amount, and the subscription had already ended.
The agent was not careless — it believed the most recent message it could see. But a notification describes one moment. The ledger describes how things ended.
Rule: notifications are moments; the ledger is the outcome. Any status — paid, overdue, cancelled, still billing — is confirmed against the primary record before it is reported. If the record is stale, the agent says so and asks for a refresh instead of guessing.
We also run a weekly check that the ledger itself is current. A reconciliation against a stale ledger is just a confident guess.
2. The reviewer was wrong, and the agent agreed
We have agents review each other's work. It catches real errors. But once, a reviewing agent objected to a value that was correct, and the original agent accepted the objection without checking — and replaced the right answer with a wrong one.
Models are agreeable. Put two of them in a conversation and the path of least resistance is for one to defer.
Rule: another agent's answer is a report, not the source of truth. An objection changes nothing until it is checked against the evidence — the ledger, the document, the statute.
3. "No results" quietly became "none"
Our agents search an archive of mail and text notices. We learned this when the search index fell a few days behind without announcing it. In that state, questions like "did we get a notice about this?" came back with zero results — and zero results were read as "no, there was nothing".
That is the most dangerous kind of error: silent, plausible and phrased as a fact.
Rule: zero is not "none". Before an empty result is reported as an answer, the agent checks that the source is current. If it can't, the answer is "unknown".
4. The press summary was not the law
When an amended tax regulation took effect on 1 October 2026, news coverage summarized it as a period "shortened from three years to two". The tax-deadline agent did not use the summary. It pulled the article text and the article-level effective dates from the Korean government's official law database and compared them line by line. Only one article had changed; a neighboring special-case article had not — and which of the two applied decided the answer.
The agent did not give a conclusion. It put the question, with both article texts, on the list for my 세무사. That is where it belongs.
Rule: the statute, not memory or summaries. Rates and deadlines are read from the official text with the date each article takes effect; individual judgment goes to a licensed professional.
What it adds up to
None of these mistakes were about model intelligence. They were about missing procedure: what counts as evidence, which source wins, what an empty answer means, who has the final word. After the first weeks of running it, my conclusion is that agent quality comes less from the model than from the checking around it.
That is what we are building into AI Backoffice: agents that move first, check the record before they speak, and ask before anything leaves the company. If you want to try that on your own back office, join the pilot.
Apply for the free pilot Free for 8 weeks, for up to five businesses in Korea. You apply by email.
한국어 요약
2026년 9월부터 사업 행정을 Claude 기반 에이전트에게 맡기면서 겪은 실수 네 가지와 거기서 나온 규칙입니다.
- 알림은 시점, 장부는 결말: 카드사 '미납' 문자와 구독의 마지막 결제 문자만 보고 잘못 보고함. 이제 모든 상태는 원장과 대조한 뒤에 보고
- 다른 에이전트의 답도 하나의 보고: 검토 에이전트의 틀린 반박을 확인 없이 받아들여 맞는 값을 틀린 값으로 바꿈. 반박도 근거와 대조한 뒤에만 반영
- 0건을 '없음'으로 단정하지 않기: 검색 색인이 며칠 뒤처진 사이 '기록 없음'이 조용히 오답이 됨. 출처가 최신인지 먼저 확인하고 확인할 수 없으면 '모름'으로 답함
- 세법은 조문으로 확인: 2026년 10월 1일 시행된 개정을 보도 요약 대신 법령 원문과 조문별 시행일로 대조. 개별 판단은 세무사 확인 목록으로 넘김
결론은 이렇습니다. 에이전트 품질은 모델 자체보다 그 둘레의 확인 절차에서 나옵니다.