Case Study · How the work gets done
Everyone says they use AI now. The useful question is what happens when it's wrong — because it will be, and whether you catch it is the entire difference between leverage and liability.
A client sent me a ninety-eight minute recording of a strategy call, transcribed automatically. About twenty-six minutes in, the transcription failed — the model entered a repetition loop and produced forty identical lines of the same half sentence. Three minutes of the conversation were gone, and the section it destroyed was the one about how they get paid.
An earlier analysis had recorded that content as permanently lost. It wasn't. The loop was in the transcription, not the recording — so I pulled the original audio, re-ran just those windows with the setting that causes the loop disabled, and recovered all of it.
What came back changed two recommendations. The client's customers weren't refusing automatic payment out of general reluctance — they were refusing because the amount changes monthly and they can't see how it's calculated. That is a completely different problem with a completely different fix. And a question the earlier analysis had listed as unanswered had in fact been answered on the call.
Nothing about that required unusual technical skill. It required not believing the output.
Four rules, learned the way you'd expect.
Anything a model asserts about the outside world gets verified before it reaches a client. In a recent engagement I ran two research agents across dozens of vendors and every claim came back tagged: verified against the live page, reported secondhand, or uncertain. Roughly a fifth of what came back was wrong — a dead company still listed as operating, a product page that had quietly 404'd, a rate range traceable to a competitor's blog rather than the company itself.
Speed without verification isn't leverage, it's a faster way to be confidently wrong.
These tools fail in characteristic ways: they invent plausible specifics, they agree too readily, they lose the thread in long transcripts, and they will produce a confident answer rather than admit a gap. Knowing the failure modes turns review from a vague worry into a checklist — I check names, numbers, URLs and anything that sounds too neatly aligned with what I wanted to hear.
Most of the value is knowing which sentence to distrust.
Complicated work gets decomposed into stages with defined outputs — research, then verification by a separate pass, then synthesis — rather than asked for in one request. The verification pass is deliberately adversarial: its job is to refute, not confirm. Findings that survive that are worth acting on; the ones that don't never reach the client.
The architecture of the work matters more than the wording of the request.
The tools are extraordinary at producing options, drafts and coverage. They are not good at deciding what matters, what a client can actually absorb, or when the honest answer is that a project shouldn't happen. That's the part I'm accountable for, and it's why the twenty-five years before any of this still counts.
Leverage on execution. No leverage on judgement.
What this buys a client
Not because the tools write the answer — they don't — but because the expensive parts compress. Research that would have been a month of calls runs in an afternoon and comes back sourced. A prototype that would have been a sprint exists before the second meeting. Options that would have been argued about in the abstract get built and compared.
What doesn't compress is deciding what to do. That still takes someone who has run a P&L, sold to an enterprise buyer, and shipped something people used.
Fast where speed is safe, slow where it isn't, and honest about which is which.