IKEA's AI agent Billie, named after the bookcase, now handles 47% of customer service queries. The maths on that is easy to read the wrong way. If a machine is answering nearly half the questions, then nearly half the call center is redundant, and 8,500 people should have shown up as an efficiency line in someone's deck by now.

IKEA read it differently, and I think their read is the correct one. Billie had not removed 47% of the jobs. It had removed 47% of the work in every one of those jobs, and it happened to be the repetitive half: the delivery-status checks, the return questions, and the where-is-my-order calls.

What was left in each of those 8,500 people was the part that made them worth employing in the first place: knowing the product range in their bones and being able to talk a nervous customer through a kitchen. So IKEA retrained them as remote interior design advisers, put them on video calls with people planning renovations, and has attributed something in the region of €1.3 billion of incremental revenue to the service those same people now run.

AI will remove roughly 80% of each job, not 80% of all jobs, and the two produce completely different balance sheets. One shrinks a headcount line, while the other leaves you holding several thousand people who have just been handed back the part of their work they were always best at, along with a lot of newly available capacity, and so the question becomes what you point it at.

Which brings me to a conversation I had this week with the CEO of a large company, one that is currently trying work this out. The CEO keeps asking the CTO when the cost of the engineering team will come down, but the answer still hasn't arrived. Token spend climbs every month, which is itself evidence that the team is building more and building it faster, but headcount has not moved, salaries have not moved, and the project list is both longer and turning over more quickly than it did last year. The value is real, and it is visible; it is showing up entirely in the throughput column, and the CEO is looking for it in the cost column, where nothing whatsoever has happened.

Neither of them is being unreasonable about it. They were handed a promise that was framed wrong at the outset, and almost every company running an AI program right now is heading for the same conversation. The promise was that AI would remove a percentage of jobs.

Companies are finding that their teams are offloading up to 80% of their work to AI, leaving them holding the 20% that survived. The good news is that 20% is usually the part they were best at anyway, the specialist knowledge and accumulated judgment that made them worth hiring.

The bad news is that, for some industries, there is now a new minimum throughput expectation, so you need the same number of people AND you need token use, effectively increasing your baseline cost. Plus, what nobody has trained them to do is let the other 80% go.

The 20% is verification, and it costs what it always did

With generating work becoming close to free, any intern can produce a fifty-page report in an hour. Knowing whether that report is right still takes exactly as much experience as it always did, and no amount of model capability shortens that apprenticeship. This is a large part of why corporate AI programs stall somewhere between pilot and rollout (and the underlying data was never clean enough to begin with).

We have a rule at FOMO.ai that sounds bureaucratic until you see what it prevents. No AI-generated document leaves the person who generated it until they have fact-checked it themselves, using their own expertise rather than another model. An error that survives one hand-off gets quoted in the next document, and the one after that, and within a fortnight it has become a fact inside the company with no traceable origin. Once that happens, you cannot unpick it, because nobody remembers which draft it came from. The check has to sit with the person who has the context, at the moment of creation, or it does not happen at all.

The failure mode has also changed shape, which makes the checking harder rather than easier. AI rarely hallucinates anymore, so the crude errors everyone trained themselves to spot in 2023 have largely gone. What remains are context failures, and they are more dangerous because they read as entirely plausible.

Example: Sarah runs a podcast production company. She ‘coaches podcasters’, and separately, she ‘produces podcasts for coaches’. Two different service lines, two different buyers, and the AI keeps collapsing them into one, no matter how the context is set up (which is somewhat surprising given how smart the AI models are today). In a world where people generate something with AI and just share it without fact-checking, nothing in the output looks wrong. A fluent and confidently structured answer built on a misunderstanding of the business does more damage than an obvious mistake, yet a new (human) hire understands the distinction on their first morning without being told.

You may have seen the politician last week reading his speech into the parliamentary record, including the line from ChatGPT, “here is a simpler version for your audience,” obviously intended only for him! That is an organization with no verification layer, scaled down to one man and a phone. The tooling was fine, but the check was missing.

Three questions nobody has been trained to answer

Handing off 80% of your job to a machine sounds like a relief until you have to do it. There are three separate skills buried in it, and none of them were in anyone's job description two years ago.

The first is knowing what to release, which means being honest about which parts of your work were volume rather than value. That is uncomfortable when the volume is the part you built a reputation on and the part that filled your day, and you are worried about having your job taken away.

The second is choosing where it goes, whether that is a model, an offshore team, a junior hire, or a process you have to design yourself. Most people default to the model because it is the nearest option, when in many cases the right person for the work costs a fraction of what they do.

The third is the hardest, which is working out how to layer your 20% back on top so that it changes the outcome. Doing it well means intervening in specific places rather than reviewing everything, and knowing which places is the actual expertise. A person who reads the whole output line by line has not gained any leverage; they have just changed what they are reading.

I think this is where most organizational design decisions over the next couple of years will be made. Sometimes one person owns both ends, generating the 80% by stewarding the AI and then applying their own 20% on top, which works when the specialist and the operator are the same brain and the volume is manageable. Sometimes the 80% and the 20% belong to the same person, but the checking cannot, because you are a poor auditor of your own output, regardless of how good you are, and a separate verification layer has to sit between generation and release.

We run both patterns at FOMO.ai, depending on the client and the stakes, and getting that choice wrong is the difference between AI making someone faster and AI making them confidently wrong at scale.

Verification needs raw material, and the raw material is observation

There is an argument to be made that proprietary observation will beat proprietary data. For most of the last decade, companies competed on data advantage, on having more of it or cleaner pipes into it. That advantage weakens sharply once AI commoditizes the analysis, because the scarce thing stops being the ability to process information and becomes the ability to notice what has not yet been turned into information at all.

Which means fieldwork comes back.

The opportunity lies in the workaround rather than the dashboard, in the to-do lists, the spreadsheet nobody will admit is mission-critical, the manual reconciliation somebody does every Thursday afternoon, the human exception the software never anticipated. None of that exists in a dataset, so no competitor can prompt their way into understanding it.

The two arguments join up at that point, because you cannot verify what you have never observed. Verification is pattern matching against something you saw with your own eyes, and someone whose entire understanding of a market came from asking a model to summarise it has nothing to match against. All they can assess is whether the output is well written, which it always is.

Why the cost line refuses to move

Back to the CEO and the CTO, because there is a structural reason neither of them can find the savings, and it is visible in our own numbers.

We spend $thousands per month on tokens. When we grow this business 100x, that spend will not grow 20x. The models are already at the capability level we need for what we do; older models keep getting cheaper, and if the economics ever turned against us, we would move to a capable model at a lower price with little drama. Models depreciate. The engine is rented and swappable, and anyone who thinks they have built a moat on top of one is going to find out otherwise.

What compounds is everything wrapped around it. The data, the processes, the judgment about what to feed the machine, the pricing, the new ways of working, and the change management/relationships with clients. That is also why the token bill climbs while the salary bill holds steady. They are paying for two things at once now, the generation and the people who understand the business well enough to check it, and the second of those was never going to get cheaper. It is the thing they are actually buying.

Example: In our world of marketing, anyone can now go plug their credit card into Google PMax, its automated ads platform, and get average results, but if you want to scale bigger or faster, or simply be in the top 30-40% of advertisers on acquisition cost, you still bring in expert humans.

The savings were never the point

Everybody is working towards faster and cheaper, which is a real outcome and shows up quickly, and is also the smallest prize on the table. If the cost of doing something has collapsed, the interesting move is not to do the same volume of work for less money, but to attempt the things that were never feasible before, the projects that failed the business case because the labor cost made them absurd.

IKEA understood that. Retraining a call center worker into an interior design adviser was not a cost decision; it was a service that could not have been staffed at that scale a few years ago, and it produced revenue rather than savings.

So the question to put to your team is not when the team gets cheaper, it is what you can now attempt that you had previously ruled out, which we will dive into more in the next edition of Hard Part First.