Read enough post-mortems on public AI failures and a pattern shows up that has nothing to do with the technology itself. It's never really "the AI got something wrong." Models get things wrong constantly, in private, all day, without incident. What actually makes headlines is a company letting an unreviewed output reach a real customer, with no backstop between the model and the person relying on it.

We watch this pattern closely, because it's the single biggest risk in our own category, and it's entirely avoidable. It's worth setting out plainly — not to point at anyone specific, but because the lesson only holds if you're honest about what actually caused the failure.

The pattern behind almost every public AI failure

One version shows up constantly in customer service: a support chatbot, given a direct line to customers with no human checkpoint, states a policy or a discount that doesn't exist. The customer takes it at face value — why wouldn't they, it's the company's own official channel — and acts on it. The company then faces a choice between honoring a commitment it never intended to make, or publicly disowning its own AI system in front of the customer it just misled. Neither outcome is good, and in at least one well-documented tribunal case, the company lost and had to honor the fabricated commitment anyway. The lesson isn't "don't let AI talk to customers." It's "don't let AI make a customer-facing commitment with nobody positioned to catch it before the customer relies on it."

The other version is quieter and takes longer to surface: a business strips out human review to cut cost or increase speed, and quality or compliance erodes gradually rather than all at once. Nobody notices for months, because nothing dramatic happens on any single day — until an audit, a regulator, or a customer complaint surfaces the accumulated gap all at once, and it turns out to have been compounding the entire time nobody was checking.

"First to ship" and "trusted to run your business on" are different races

There's real commercial pressure to look cutting-edge — to be the platform with the newest feature, first. We understand that pressure; we operate in a category that rewards it. But when a feature ships before it's actually reliable, the damage rarely lands on the vendor first. It lands on the client who adopted it — the agency whose candidate got placed on a check that turned out to be stale, the production whose buyer relationship absorbed an avoidable mistake. That's a relationship the client spent years building with their own customers, not something we're entitled to spend on their behalf just to look advanced sooner.

That's the actual trade-off behind every "should we ship this now" decision: whoever adopts a capability first is, functionally, the one who finds out where it breaks. We don't think that should be a client, ever.

Where we draw the line

Every new AI capability we could plausibly adopt goes through a standing internal review before it goes near a live system — evaluated against our own bar, not the market's release schedule. If a capability is available elsewhere and we haven't adopted it, that's a deliberate call, not an oversight, and we'd rather explain that call than explain an incident afterward. When something does clear the bar, it's rolled out with the client's knowledge, not dropped in silently between logins.

None of this makes us immune to getting something wrong eventually — no one building real systems gets to claim that. It just means the place we test is internal, on our own time, against our own standard, before a client ever sees it. Full detail on how human sign-off and independent challenge actually work across every system is on our Trust & Compliance page.

Want to see how the review process works before a change ever reaches your account?

Talk to us