AI Scales What You Never Decided
CIOREVIEW >> Artificial Intelligence >> NEWS

Auxiliadora Predial

Marcel Guinther, CTO

AI Scales What You Never Decided

Marcel Guinther, CTO
Marcel Guinther, CTO, Auxiliadora Predial

Marcel Guinther

AI Quality Modernizer

A poorly trained person hesitates. They fail slowly, realize they don't know and escalate. A poorly grounded model answers fluently and wrong and nothing in the answer signals the failure. Fluency is not competence and fluency is what hides the failure until it reaches the customer.

This matters because nearly every enterprise AI initiative lands in the same place, the front line, where the company talks to its customers. That is where the pilot impresses and where the gain shows up fastest. It is also where AI has the least chance, because there it inherits every defect of the process beneath it, undefined states, missing data, policy applied differently by different people. Where the process doesn't decide, the model decides.

Bill Gates wrote the rule in 1996, automation applied to an efficient operation will magnify the efficiency and automation applied to an inefficient operation will magnify the inefficiency. Generative AI raised the price of ignoring it, because the inefficiency now comes out fluent.

The problem starts before the data

The best-known version of this argument is about data. In August 2026, writing on Martin Fowler's site, Pramod Sadalage and Prem Chandrasekaran put it plainly, for thirty years we built data systems for human analysts, who supply the context, judgment and skepticism to work around data that is incomplete or wrong. Autonomous agents supply none of that. They act on whatever they are handed, confidently.

I agree. My only addition is that it starts one step earlier.

You can have data contracts, freshness SLAs and one agreed definition per metric and still have no way to say whether your agent is any good, because the missing piece is judgment nobody ever wrote down.

To improve an agent that talks to customers, you need labeled examples of what a good interaction looks like in your own operation. If human service was never measured against explicit criteria, those examples do not exist. Without them you have no evaluation set, no threshold you can defend and no way to prove the new version is better than the last one.

There is a cheap test for where you stand. Take ten recent service interactions and ask three supervisors which ones were good. If they disagree and they will, the problem sits upstream of any model, "good" was never defined and you were about to ask a machine to reproduce at scale a standard your own organization cannot state.

That is why so many of these projects never leave the pilot, nobody can say whether the model is good enough.

 Protecting that sequence against the pressure to demo is leadership work, not architecture work. 

Most teams discover the complication late. AI is strong on the generic and the shallow and weak precisely where your process is specific and structured. The more your differentiation lives in that specificity, the less a generic model delivers on its own and the more of your own judgment it requires.

The inversion

The alternative is to put AI to work evaluating human performance against explicit criteria, instead of replacing it.

In the operation I lead, that is where we started, AI scoring service interactions against criteria defined together with the people who own the process. The scores mattered less than what they exposed. Our human evaluators disagreed with each other and until we wrote the criteria down nobody knew by how much. That disagreement was the real problem and it had been invisible.

As a by-product of quality management, you gain three things you did not have. You get an operational definition of quality, which until then lived in supervisors' heads and shifted with the supervisor. You get a growing corpus of labeled good and bad examples. And you get a quality signal per person, concrete enough to carry a development conversation.

The risk is asymmetric in your favor. A wrong score costs you an awkward feedback conversation. A wrong answer costs you a customer.

None of this is free. If the evaluator is not calibrated against human judgment and recalibrated often, the team stops believing the score and the mechanism dies within a quarter. Calibration has to continue for as long as the evaluator runs.

Once that exists, the customer-facing agent stops being a bet and becomes engineering,  you have an evaluation set, a threshold you can defend and a human baseline to measure against. That sequence is a technical dependency.

"But the front line is where differentiation lives"

It is the strongest objection to what I have written and it holds. Back-office gains are real, measurable and invisible. No customer switches providers because a competitor's back office improved and whoever moves early on experience reshapes the market before the cautious get clarity.

My answer is that this sequence gets you there faster. The alternative usually plays out as deploying, degrading, reversing and then building the foundation that should have come first. Klarna announced that its AI took on the work of 700 service agents, reversed course in 2025 and settled on a hybrid model of human agents assisted by AI, which is where it should have started. I am not proposing that you wait. I am proposing that you not pay twice.

My argument has a limit. If your process is generic and shallow, AI at the front will solve it quickly and you should go, that is what these models do best. The argument holds where your differentiation lives in the specifics.

Why almost nobody does it this way

Because it is invisible. Nobody demos a quality evaluation pipeline in a board meeting. The front-line bot demos beautifully and fails quietly, three months later, in customer satisfaction.

Protecting that sequence against the pressure to demo is leadership work, not architecture work.

And the question that separates the two paths is simpler and more uncomfortable than choosing a use case, can you state today what good service looks like in your company, with criteria two different people would apply the same way?

If you cannot, AI at the front will not fix that. It will scale what you never decided.

The articles from these contributors are based on their personal expertise and viewpoints, and do not necessarily reflect the opinions of their employers or affiliated organizations.