Incorrect assumptions bottleneck ambitious delegation of non-coding work
I’m currently experimenting with ambitious delegation: asking Claude to handle big projects and ongoing role responsibilities—not just tasks—with minimal support from me.
Experiments (completed):
- Do my company bookkeeping
- Work with my accountant to complete my UK tax return
- Complete a personal finance review
Experiments (ongoing—journal posts to follow):
- Onboard a new client to the TYPE III AUDIO narration service
- Serve as Head of Growth for a low-stakes side project.
With coding tasks, I rarely need to review an agent’s work in detail. But for non-coding tasks, careful review is still required. A key reason is that Claude sometimes makes important and incorrect assumptions, and fails to self-correct or flag them for my review. It’s not terrible at this, but the rate of unacceptable mistakes is too high (e.g. confidently putting the wrong address on my tax return, despite having the correct address in context; making a big mistake about French tax law).
It feels like working with a human colleague who is somewhere between 70th–90th percentile on conscientiousness. But in practice it’s worse because the worst mistakes the agent could make seem much worse than those such a human colleague might make. Jagged.
Some ways to mitigate this:
- Prompting: encourage more exhaustive investigation, fact checks and flagging of important uncertainties.
- Harness: make it easier to investigate key uncertainties, and enable multi-model adversarial review (in a similar spirit to asking Codex for code review on Claude’s work) before things are passed up to me.
- Context: do a better job of providing critical context before work begins.
I expect some of these techniques will remain useful throughout 2026, but a decent amount will just be handled by near-future releases from the AI companies.
