The MarginAnalysis

Why AI Isn't Saving You Time: 10 Traps

Ten beliefs about AI that all sound reasonable, which is exactly why they survive. None of them is a problem with the technology. In every one the model is a component, and it is never the component that decided the outcome — it is a decision about your own work that got easier to postpone once there was a machine in the room to point at.

Dark cover plate. An orange Analysis chip, the word One set large in italic serif, and the line reading layer of the work crossed over, the other three stayed and they scale with it. At right, four numbered rows threaded on a spine, 01 Decide, 02 Make lit in orange, 03 Judge, 04 What next.

The work has four layers: deciding what to make, making it, judging whether it is any good, and deciding what happens next. Only the making crossed over. The other three stayed with you — and they scale with the one that left. Ten drafts instead of one means ten things to judge instead of one, and that is why the day feels fuller rather than shorter.

Everything below is a version of that same shape. Ten beliefs, each of which sounds reasonable, none of which is a problem with the technology.

Will AI do the work for you?

It will do one layer of it. Split any piece of work into four: deciding what to make, making it, judging whether it is any good, and deciding what happens next. Only the making crosses over.

The trap is not that the other three are hard. It is that they scale with the layer that left. When making something cost you a day, judging it cost you an hour and you did that once. When making it costs four minutes, you have ten versions by lunchtime and ten judgements to make, and judgement was always the expensive part.

This is why people who adopt these tools enthusiastically report the same thing six months in: more finished, nothing better. The output went up and the queue of decisions went up with it. If you want the version of this with hard measurement attached, the software industry ran the experiment at scale and the results are in is AI making code worse.

Does AI know your business because you told it about your business?

No, and this is the single most common misunderstanding about how these systems work. That paragraph you wrote about your company is not stored anywhere. It sits in front of the model as text and is read again from the beginning every single time, competing for influence with everything the model absorbed from the rest of the world.

Which means describing yourself in adjectives is the weakest possible move. "We're premium but approachable, professional but warm" is three words the model has seen ten million times, attached to ten million other companies.

Supply the artefact, not the adjectives. Your last five emails to a client. The actual proposal that won. The page you are proud of. A real example carries more usable signal than a paragraph of self-description, because it is specific to you and nothing else in the model's world looks exactly like it. That is the practical half of what people call context engineering, and it is worth more than any prompt library.

Do you just need better prompts?

Three things decide what comes back: the instruction, the material it has to work from, and the definition of good it gets checked against. Prompting is the first one. It is also the narrowest, and it is the one everybody spends all their time on.

The pattern is easy to spot in your own work. If output quality swings wildly between attempts with no clear reason, the problem is usually material — the model is inventing what it was not given. If output is consistent but consistently not what you wanted, the problem is the definition of good — you never told it, or yourself, what finished looks like. Neither of those is fixed by a better opening line.

Prompting is still worth learning properly rather than by folklore, and we wrote that up separately in how to prompt Claude. Just do not expect it to carry the other two.

Why does AI sound so confident when it is wrong?

Because fluency and correctness are not the same needle. What comes back is built by choosing the next chunk of text that fits the text so far, and a smooth, well-formed sentence is what that process is for. Correctness is a separate question the process never asked. A confident wrong answer is not the machine going wrong. It is the machine doing exactly what it does.

There is a second reason, and it is about how these systems are graded. The scoring used in training and evaluation gives credit for an answer and gives nothing at all for saying "I do not know" — which is the same rule as an exam where a guess costs you nothing. Under a rule like that, guessing is the correct strategy, and guessing is what you get. That rule was chosen by people. It is not a property of the technology.

So change what you check. Do not ask whether the answer sounds sure, because it always will. Ask what you would have to open in order to prove it wrong, and then open that. If it gave you a number, ask for the row it came from. If it gave you a source, open the source. That habit is the same one that catches reward hacking in agents, and the same one behind when you actually need a human in the workflow.

If you produced more, did you accomplish more?

Not necessarily, and there is a law about it. Speed up one part of a process and the total time you can possibly save is capped by how much of the whole job that part was. If drafting was 20% of the work, making drafting instant saves you 20%, no matter how instant it gets. That is Amdahl's law, written about computers in the 1960s, and it applies unchanged to a Tuesday afternoon.

The practical consequence is that the first question is never "what can AI speed up?" It is "what is the neck?" — the part where work actually queues. In most small businesses the neck is not production at all. It is deciding, approving, or the two days something sat in somebody's inbox. Point a generation tool at a decision bottleneck and you get a bigger pile in front of the same bottleneck.

Do you need more AI tools?

A tool needs three things before it does anything for you: your real data going in, the output landing somewhere the work already runs, and one person who owns it. Without all three, what you have is a subscription.

That is not a moral point, it is an arithmetic one. Most stacks fail on the second condition — the output lands in a tab nobody opens, so somebody has to remember to go and get it, and remembering is the thing that stops happening in week three. The third condition fails almost as often: a tool everybody can use and nobody owns has no one to notice when it quietly stops being used.

We ran the audit version of this on an actual stack in what $100 a month of AI tools is really doing. The cancellations were not the tools that were bad. They were the tools with no owner.

Should you go looking for things to use AI on?

Starting from the tool is the wrong direction, and it fails in a specific way rather than a general one. A capability out looking for a use finds the easy tasks — and the easy tasks were easy because they were small. You end up with six automations that each save four minutes, in a business where the actual cost is a two-week sales cycle.

Start from what costs you. Write down the three things that cost the most time, money or sleep this month, and then ask whether any layer of any of them is a making problem. Often the answer is no, and that is a useful answer, because it stops you spending a fortnight automating something that was never the problem.

If AI can do it, do you still need to learn it?

Yes, and the reason is structural rather than moral. You kept the judging layer — and the judging layer was made out of the doing. You can only reliably see the mistakes you have made yourself. Someone who has never written a contract cannot tell which clause is missing from a generated one, and no amount of confidence in the output changes that.

This has a name and a date: the ironies of automation, Lisanne Bainbridge, 1983. Her observation was that automating a process makes the human operator's remaining job harder, not easier, because what is left is monitoring and exception handling — the two things people are worst at and get least practice in, precisely because the automation removed the practice.

The version that matters for a person rather than a plant: nobody's existing skill evaporates. What stops is the acquisition of the next one. You keep what you had, and you stop adding — which is survivable at twenty years in and a very different proposition at two.

If AI makes everyone faster, does everyone win?

No. The bars all grew and the finish line did not move. When a capability is available to everybody at roughly the same price, it stops being an advantage and becomes a floor — the new minimum standard for being in the conversation at all.

The related pattern is Jevons paradox, written about coal in the nineteenth century: making a resource cheaper to use tends to increase the total amount consumed, not decrease it. Applied here, cheaper writing does not mean less writing. It means far more of it, which means the cost of being read goes up for everyone, including you.

The strategic consequence is the opposite of what the tool vendors imply. When the making gets cheap, the scarce things are the ones that did not get cheaper: judgement, taste, access, accountability, and a reputation somebody can check. Those are also the things nobody can buy at $20 a month, which is exactly why they are worth building.

Is the smartest model the best model?

A benchmark is a fixed set of short problems with known answers. Your week is not that. Speed, cost, whether the output comes back in the same shape every time, whether it can see your data at all, whether it fits the place the work actually happens — the ranking measures none of it.

The practical test is dull and much better: take one real task you actually do, run it through two or three candidates with your own material, and judge the outputs against your own definition of good. That takes an afternoon and beats every leaderboard, because it is the only evaluation performed on your work. There is a related failure mode when the model is doing a job rather than a task, which we covered in how reliable an AI employee actually is.

What all ten have in common

Here is the pattern, and it is the reason this list is ten items rather than ten unrelated complaints.

In every one of the ten, the model is a component, and it is never the component that decided the outcome. What decided it was a definition of good that nobody wrote down, a bottleneck nobody identified, an owner nobody assigned, a skill nobody kept, or a target nobody checked. Each of those is a decision about your own work — and each got much easier to postpone once there was a machine in the room to point at.

If you want to find your own, work backwards from the symptom:

  • More output, nothing better — trap one. Your judging layer is the neck.
  • Generic results no matter how you word it — trap two. You supplied adjectives, not artefacts.
  • Wildly inconsistent results — trap three. It is inventing material you did not give it.
  • Errors that got past you — trap four, and often trap eight underneath it.
  • Faster everywhere and no change to the month — trap five. You sped up the wrong 20%.
  • Tools you pay for and do not open — trap six. No owner.
  • Lots of small automations, same problems — trap seven. You started from the capability.
  • Everyone in your market got better at once — trap nine. It was a floor, not an edge.

The one you found yourself arguing with while reading is usually yours.