Share
Facebook Facebook icon Twitter Twitter icon LinkedIn LinkedIn icon Email
Small silhouette of a person on a yellow path facing a large gavel, symbolizing pursuit of justice.

Artificial Intelligence

Is your AI a witness for the defense or prosecution?

Published August 11, 2026 in Artificial Intelligence • 13 min read • Audio availableAudio available

Organizations have historically maintained a comfortable distance between corporate rhetoric and operating reality, but all that changed with AI. Boards must move past performative oversight or face severe regulatory and reputational risks, writes  Abraham Hongze Lu.

Rapid read:

  • When AI behavior diverges from public promises, the organization does not just face a technical incident; it faces a credibility event, which is expensive.
  • When scrutiny comes, your AI will not quote your values statement. It will describe its – and that means your organization’s – behavior.
  • AI-enabled scrutiny will either reveal the extent of divergence between proclamations and actions, or help us build more accountable and credible organizations.

When AI risk becomes a board issue

When a European bank rolled out a new AI-driven credit model for retail customers, it looked like a success. Approval times fell. Loss forecasts held within risk appetite. Early scores suggested the model was more accurate than the old rule-based system.

Two weeks later, an internal monitoring script lit up. Applications from a cluster of urban postal codes were being declined at a noticeably higher rate, even after accounting for income and credit history. Those neighborhoods were poorer and more diverse.

One executive stared at the chart: “We are declining more applications from the neighborhoods we just promised regulators we would support.” The room went quiet.

The bank’s last regulatory settlement was barely a year old. What began as a model-monitoring issue was fast becoming a board issue.

On its own, the pattern might have been treated as a technical quirk or a data quality issue. What made it radioactive were two facts everyone in the executive suite knew.

First, the model had learned from the bank’s historical lending decisions, which had never been audited for bias. Second, the bank had spent the last five years telling regulators, investors, and employees that inclusive growth and fair access to credit were central to its identity. Now, a report was telling them: your algorithm has evidence that your commitments and behavior diverge.

AI systems do not just automate decisions and accelerate processes; they log everything. This makes organizational behavior easier to inspect, compare, and challenge.

When scrutiny comes, your AI will not quote your values statement. It will describe its – and that means your organization’s – behavior. We call this the, “algorithmic integrity gap,” the distance between the promises you make as an organization and the patterns of behavior your AI models register and reveal. 

In research published in the Journal of Management, we analyzed 12,555 firm-year observations from 1,704 US public companies and found that disingenuous ethical language, or “cheap talk,” was a reliable predictor of weaker social performance.

However, this connection was weaker when firms faced stronger monitoring. When scrutiny is low, messaging and behavior drift apart more easily. When scrutiny rises, the gap narrows.

AI changes the economics of scrutiny. It makes patterns of behavior measurable to a level of detail not readily seen or available before. Extrapolating from our research, AI-enabled scrutiny will either, like that European bank, reveal the extent of divergence between proclamations and actions, or help us build more accountable and credible organizations.

For boards and CEOs, there is a clear oversight risk. This is not just about ethics; it’s about regulatory exposure and reputational damage. When your AI is asked to testify about your company, how do you guarantee that it supports your story instead of destroying it?

When your AI is asked to testify about your company, how do you guarantee that it supports your story instead of destroying it?

When your story meets the data

For most of the last century, companies could live with a comfortable distance between rhetoric and reality. Time, complexity, and limited visibility blurred the gap between the talk and the walk. AI narrows that distance sharply.

In the past, organizational behavior often disappeared into human discretion and scattered records. Today, company AI systems generate outputs, overrides, and incident histories that make it easier to piece together those patterns.

Any serious model, whether it screens job applicants, sets credit limits, prices insurance, ranks safety incidents, or routes customer complaints, embeds a chain of choices: what data you decide to include, what you leave out, which variables you treat as important, how you define success, where you set thresholds, when humans override the machine, and why. This creates an organizational fingerprint that shows what you value in practice.

If your culture favors speed over safety, your models will learn to prioritize speed, even if your messaging emphasizes safety. If your teams have learned to discount certain categories of complaints, those blind spots will reappear in the data.

If you talk about fairness but pay people to hit targets at any cost, your AI will detect which behaviors get rewarded and optimize for them.

In isolation, any one of these patterns can be written off as “a bug.” Placed alongside your communication materials and CEO speeches, they start to look like evidence.

Regulators are already reading AI outputs the way auditors read ledgers. In November 2025, France’s equalities regulator ruled that Facebook’s job-ad delivery algorithm engaged in indirect gender discrimination, citing differences in who saw ads for specific roles.

The lesson for other executives is not, “Meta got caught,” it’s that regulators did not debate intent, they examined outputs. Your public commitments are claims that can be compared against model behavior.

Once a gap is visible, it becomes a governance problem: do you adjust the system, adjust the incentives around it, or quietly accept the inconsistency and hope nobody notices? If your plan is the latter, the database will beat you.

The good news, however, is that this heightened level of internal scrutiny and record offers an opportunity to ensure that your behavior reflects your values and proclamations.

When a company corrects biased practices, tightens incentives, and builds better oversight, its systems can start learning those improvements too, reinforcing the culture leaders are trying to create – or at least are talking about.

Graphic titled '5 Key Questions to Ask Yourself' listing five questions about AI models, data provenance, model history, escalation patterns, and authority approvals.

Why principles and committees are not enough

Most leadership teams are not ignoring this landscape. They approve AI principles, appoint ethics committees, commission bias training, and add a slide on AI risk to the board’s agenda. The governance deck looks reassuring, but the operating reality often isn’t.

Many organizations have statements that reference fairness, transparency, and human oversight. Yet unless those words are tied to decisions about data, models, and deployment gates, they remain proclamations that look good on a website but fail under cross-examination.

That problem cannot be solved by culture alone. Evidence from our ongoing survey of 533 AI practitioners in four countries, which examines how speaking-up conditions and formal governance shape whether concerns are escalated or absorbed quietly, shows that issues are often noticed early but managed sideways rather than formally raised in companies.

The most common dodge we see is that AI risk is not an oversight issue, but a technical matter that can be delegated. Boards hear: “Our team is on it.” Responsibility is placed with the CIO, chief data officer, or model risk group, while strategic and reputational consequences are not prioritized.

What you get is performative governance; nobody can answer the basic questions. Which models make consequential decisions about customers and employees? Who owns them? What data do they rely on? How are incidents recorded? Who can pause deployment when the evidence turns ugly?

Regulators are building expectations around these questions. Under the EU AI Act, high-risk systems must be documented, monitored, and logged so authorities can reconstruct how a system behaved in practice. That matters even if you never face an audit. It is the difference between, “we think it’s fine” and, “here is the file.”

Five practices of algorithmic integrity as an evidence system

Practice What it looks like Board question to ask
Build a witness list One model register with owner, purpose, where used, and risk tier. Updated regularly. High-stakes models clearly tagged. Which models make consequential decisions about customers or employees, who owns them, and which are high-risk?
Establish chain of custody Data lineage for key datasets, source, coverage gaps, known skews, and how labels were built. Data cards for flagship systems. For one major model, where did the data come from, what’s missing, and how do we test for bias in practice?
Maintain the case file Model cards capturing purpose, assumptions, limits, drift, and major changes. No critical release without an update. If an authority asked for the model’s full history, could we produce one coherent file, end to end?
Create an early warning channel Clear routes to raise concerns, time-bound review, documented outcomes, and visible leadership support for “pause and fix.” How many AI concerns were escalated last year, what patterns emerged, and what did we change because of them?
Give oversight real authority Oversight with access to models, documentation, and logs, plus a direct line to a board committee. Real stop/go authority. Who can stop a high-profile launch, by name, and when was that authority last used?

Five practices that make your AI defensible

If you treat AI as testimony, the board’s job changes. It becomes less about chasing “AI maturity,” scores and more about shaping what your systems can prove under scrutiny.

These five practices will give the board something stronger than reassurance: proof. Each one leaves a record you will want when someone asks how you know a system is fair, safe, and consistent with what the company has promised or declared.

1. Build a witness list

You cannot govern what you cannot list. Yet many executive teams cannot produce a complete inventory of the models they rely on, especially when vendors, business units, and “shadow analytics” teams all deploy quietly.

Start with one enterprise register of significant models: owner, purpose, where it is used, and a simple risk tier. High-stakes systems should be clearly tagged, and the list should be refreshed at a cadence that matches how quickly your teams deploy.

Which models make consequential decisions about customers or employees in your organization, and who owns them? A model you cannot name is a model you cannot defend.

2. Establish a chain of custody

Boards do not need technical detail. They need confidence that the company can explain the data behind a major decision. When a regulator calls, they won’t ask about your principles. They’ll ask where the training data originated from.

Document the lineage of key datasets: source, coverage limits, known skews, and how “ground truth” labels were created. Surface what is missing, not only what is present. Test for proxy variables that smuggle protected characteristics into the model through the back door.

Data cards and similar documentation are not paperwork; they are the chain of custody that allows you to explain outcomes. Without it, you are asking regulators, courts, and the public to take your word for it.

Pick a flagship model. Ask someone to walk you through where the data came from, what is missing, and how bias has been checked in practice.

Regulators are already reading AI outputs the way auditors read ledgers.

3. Maintain the case file

If something goes wrong, the company should be able to show the system’s history without having to reconstruct it from memory.

Models drift. They are retrained, repurposed, and integrated with new systems. Over time, many lose any clear connection to their original assumptions. That is when leadership discovers, too late, that nobody can answer the question, “What did we think this system was doing when we put it into the world?”

Think of model history like an aircraft maintenance log. A single entry doesn’t tell you much, but a history tells you whether the asset has been cared for.

Model cards or equivalent records should capture purpose, assumptions, limitations, performance drift, and material changes. No critical model should go into production or be materially updated without an entry.

If a regulator asked for the full history of a high-stakes model, could you produce one coherent file, end to end?

4. Create an early warning channel

People usually see problems before the board does. Do they have a safe and visible way to raise them? In almost every AI failure we have studied, someone saw the problem earlier, but there was no clear, protected pathway from concern to action.

Recent scrutiny of grocery firm Instacart’s AI-enabled pricing shows how quickly these issues move from internal experimentation to public evidence. In late 2025, an investigation involving 437 volunteers who bought identical baskets through Instacart found wide price variation for the same items from the same stores, sometimes as high as 23%.

Normally, analysts notice anomalies like this early. Without a safe escalation path, many keep it in chat threads until an outsider forces the conversation.

Escalation paths that work share a few traits. They are explicit about who can pause a deployment, how issues are logged, how quickly they are reviewed, and how decisions are communicated.

They are psychologically safe: raising a concern does not create personal career risk. And they are visible; leaders share anonymized examples of issues raised and how they were handled.

How many AI concerns were formally escalated in your organization last year? What patterns did you see? What changed because of them?

5. Give oversight authority

Oversight is only real if someone with enough standing can stop a decision that looks unsafe or inconsistent with the company’s commitments.

Many firms place, “responsible AI” under the same executive who owns AI growth targets. That executive may be committed and capable, but the structure creates an obvious tension.

Oversight works better when it has a direct line to a board committee and the right to slow, reshape, or stop deployment when risks are unacceptable.

Name the role, define the stop/go authority, and require that major exceptions be documented rather than waved through in meetings.

What boards should be asking now

In a recent KPMG survey, 39% of companies said they planned to include AI risk and associated controls in the scope of financial reporting processes over the next year. The same survey also found that 62% wanted their external auditor to use traditional AI for risk mitigation and internal controls.

Those responses suggest companies are starting to treat AI as an exposure that belongs inside assurance and control systems, not just inside innovation teams.

The stakes are not abstract. When AI behavior diverges from public promises, the organization does not just face a technical incident; it faces a credibility event, which is expensive. Such cases raise legal exposure, complicate regulatory relationships, and restrict strategic freedom just when you need it most.

At the European bank highlighted earlier, it was an internal report that spoke first. At your organization, it might be a regulator, a journalist, or a plaintiff’s lawyer asking the same questions your own systems are already answering.

Sooner or later, someone will ask your system to explain itself. It won’t do spin, it will do records. When that moment comes, will it speak as a hostile witness or as your most credible ally?

Authors

Abraham Hongze Lu

Global Board Center Research Fellow at IMD business school

Abraham Hongze Lu is a CAIA charter holder and research fellow at the IMD Global Board Center. Lu has worked with senior executives in many customized and open programs at IMD since 2005. He specializes in quantitative research in governance, strategy, and finance using text analysis and natural language processing techniques.

Related

Learn Brain Circuits

Join us for daily exercises focusing on issues from team building to developing an actionable sustainability plan to personal development. Go on - they only take five minutes.
 
Read more 

Explore Leadership

What makes a great leader? Do you need charisma? How do you inspire your team? Our experts offer actionable insights through first-person narratives, behind-the-scenes interviews and The Help Desk.
 
Read more

Join Membership

Log in here to join in the conversation with the I by IMD community. Your subscription grants you access to the quarterly magazine plus daily articles, videos, podcasts and learning exercises.
 
Sign up

Log in or register to enjoy the full experience

Explore first person business intelligence from top minds curated for a global executive audience