Adventures in Botsitting
A report came out last month with a term that I found pretty funny, and pretty spot on.
“Botsitting.”
It’s from Glean’s Work AI Institute, which surveyed six thousand workers about how AI is going for them, including a decent chunk here in Australia. “Botsitting” is all the work that happens after the AI answers you. The checking and the correcting, and the context you feed it on the second attempt after nearly getting there on the first. The workers surveyed reported that AI saves them around eleven hours a week. They also reported spending six and a half of those hours supervising it. More than half the dividend goes back into minding the thing that produced it. It’s the type of statistic that can be so easily framed for either side of the AI debate. But hear me out on why I think it can be a really positive indication of how things are working.
The report pairs it with a second term, “botshitting,” for what happens when people stop supervising altogether and ship whatever comes out. Sixty-nine per cent of users admitted to some version of it.
(The skeptic in me has to note that Glean sells software that’s a fix for the problem its own research just named, so the exact numbers deserve some space. But the pattern will be familiar to anyone using these tools regardless.)
The instinct on reading this is to assume that “sitting” is as much of a defect as “shitting,” and to assume that if you create better systems then the six and a half hours of sitting would shrink toward zero. For a lot of work that’s right on. If your accounts team is still hand-checking every invoice line the AI extracts three years from now, the problem is the system design, and it’s fixable with some straightforward engineering. That work should be done and those hours should disappear.
But the label is covering more than just that kind of checking.
Take a developer reviewing code an agent wrote before it merges. You could call that botsitting. You could also call it code review, which was part of the job long before the AI arrived, and which nobody suggested eliminating back when the code was written by a graduate. The review was always where the senior engineer earned their keep.
The same goes for advice. A lawyer who reads an AI model’s analysis before signing off on next steps is doing the thing the client pays for. The model getting to 99 per cent doesn’t change who owns the signature.
So there are two different activities travelling under one name. Some of the supervision is compensation for systems that haven’t been engineered properly yet. Some of it is the job itself, the judgement part, with a new label.
The report helps point out what happens when nobody makes that distinction between the two terms. Supervising a model is tiring. You’re working backwards from a wrong answer without knowing which assumption produced it. When that grinds on long enough, people keep using the AI and gradually stop reading what it gives them. That’s where the sixty-nine per cent comes from.
Which brings it back to measurement. If the metric is hours saved, every one of those six and a half hours looks like waste, and the pressure runs toward eliminating all of them. The engineering-shortfall hours deserve that pressure. But the judgement hours are a different matter, and a business that automates them away has traded its expertise for output.
Eleven hours saved, six and a half spent supervising. Before trying to claw those hours back, it’s worth working out which of them you’d want to. Some are a bill for immature systems. The rest are where the expertise lives, and it would be a shame to delete them just because they finally showed up on a timesheet.
Explore our AI services
Brisbane-based AI consulting across strategy, automation, agents, and data.