INNOVAREModule 5 · Ethical AI

Case 4: Governance with Teeth

What does AI governance look like when it has enough authority to veto a revenue line? IBM's framework — built after the failures, not before them.

July 2026 · Case 4 of 6
As you read — hold this question

What does it actually look like when an AI ethics framework has enough authority to cancel a product line — and what did it take to build one?

$62M
Watson for Oncology spent before shutdown. Then in 2020, IBM's Ethics Board vetoed its entire facial recognition business — sacrificing a revenue stream on accountability grounds.

IBM's responsible AI framework is widely cited as one of the most developed corporate governance structures in the sector. It features a dedicated AI Ethics Board, Local Focal Points across the organisation, continuous bias auditing, and open-source tooling released for industry-wide use. The case study presents this as a forward-thinking model built from principles. The fuller story is that Watson for Oncology failed at MD Anderson after burning $62 million. IBM was identified as one of the primary subjects of the Gender Shades study, with error rates of up to 34.7% for darker-skinned women. The governance framework was rebuilt after those failures. Whether that makes it more credible or less is a question worth holding.

IBM's three-tier structure
Quiz: IBM Governance

How IBM organises AI accountability across the organisation

IBM's governance structure places accountability at three levels, each with a distinct function. The key design principle is that trust is built into the process — not attached at the end.

Level Body Function Authority
Top AI Ethics Board Sets principles, policy, and accountability requirements across the organisation Veto authority over AI deployments — demonstrated in 2020 by exiting facial recognition entirely
Middle Governance Framework + Risk & Compliance Principles-based oversight; process controls; risk assessment for AI systems in development and deployment Reviews and approvals for systems before deployment; ongoing compliance monitoring
Operational Fairness & Bias Auditing / Explainability Review / Privacy & Security Controls Continuous operational functions — not pre-launch gates. Running audits catch problems that emerge after deployment, not just those predictable before it Flagging and escalation to governance layer; direct authority over specific operational decisions

The Local Focal Points network embeds governance responsibility across IBM's regional and business-unit structure — so the Ethics Board mandate does not remain a head-office position. Whether that mandate reaches delivery teams and client engagements in a large, hierarchical organisation is the unresolved question.

Four pillars

IBM's operational accountability framework

IBM organises its responsible AI practices around four operational pillars. Each corresponds to a specific type of failure that had already occurred — in IBM's own deployments or across the industry.

Explainability

AI systems must be able to give reasons for their outputs. Watson for Oncology could not explain its treatment recommendations in terms oncologists could evaluate. Explainability is the technical prerequisite for human oversight — you cannot review what you cannot understand.

Fairness

Continuous bias auditing — not a pre-launch gate. IBM's facial recognition had error rates of up to 34.7% for darker-skinned women before Gender Shades. The continuous auditing framework came after that finding. It now runs across IBM's AI deployments to catch distributional failures that aggregate accuracy metrics hide.

Robustness

AI systems must perform consistently across their intended deployment range — including edge cases, adversarial inputs, and conditions different from training. Tay failed this completely. Robustness testing must include scenarios where the system is actively stressed, not just standard use cases.

Privacy

Data used to train, operate, and improve AI systems must be governed throughout its lifecycle. This includes consent scope, data residency, access controls, and the question of whether data collected for one purpose is being used for another — a live issue in health AI globally.

Open-source tools

AI Fairness 360 and AI Explainability 360 — governance as infrastructure

IBM open-sourced two toolkits following the Gender Shades and Watson failures. These are now used globally — including by organisations that have no relationship with IBM.

Tool What it does Why it matters
AI Fairness 360 Open-source Python toolkit for detecting and mitigating bias in machine learning datasets and models. Includes 70+ fairness metrics and 10+ bias mitigation algorithms Provides a common, auditable framework for measuring bias across demographic subgroups — addressing the four upstream decisions identified by Google's PAIR team
AI Explainability 360 Open-source toolkit providing 8 explainability methods for different model types and use cases. Covers local (why this decision?) and global (how does the model work?) explanations Operationalises the Transparency principle — giving technical practitioners the tools to make models auditable by humans who need to understand their outputs
The honest read
IBM's governance framework is real — the 2020 Ethics Board decision to exit facial recognition entirely, sacrificing a revenue stream, demonstrates that. But the framework was forged in the fire of Watson for Oncology and Gender Shades. It was not built proactively. IBM is also a top-down, bureaucratic organisation. Whether board-level governance successfully reaches delivery teams and client engagements — the "last mile" of accountability — remains the live question. Governance that exists at the Ethics Board level but doesn't reach the team deploying a model in a client context is not governance. It's documentation.
Take this away

IBM's framework is the most credible corporate AI governance model available as a reference — and it was built after expensive failures, not ahead of them. That history makes it more honest, not less useful. But the test of any governance model is not what the Ethics Board decided. It's whether accountability reaches the person writing the deployment specification for a client project at the end of a long delivery chain.

Quick recall — without looking back

Test yourself on this case

Question 1 of 3

Describe IBM's three-tier AI governance structure — the bodies at each level, their functions, and the authority each holds.

Top level: the AI Ethics Board. Sets principles, policy, and accountability requirements across the organisation. Holds veto authority over AI deployments — exercised in 2020 when it exited IBM's entire facial recognition business on racial profiling and human rights grounds. Middle level: Governance Framework and Risk & Compliance function. Provides principles-based oversight and process controls; reviews and approves AI systems before deployment; monitors ongoing compliance. Operational level: three running functions — Fairness and Bias Auditing (continuous, not a one-time pre-launch check); Explainability and Transparency Review (verifying that systems can give reasons for outputs); Privacy and Security Controls (governing data throughout the AI lifecycle). The Local Focal Points network distributes governance responsibility across IBM's regional and business-unit structure so the mandate does not remain a head-office position.
Question 2 of 3

What are IBM's four operational pillars for responsible AI — and for each, identify the specific failure that demonstrates why it is necessary?

(1) Explainability — Watson for Oncology could not explain its cancer treatment recommendations in terms oncologists could evaluate or interrogate. Human oversight requires that systems can give reasons for their outputs. (2) Fairness — IBM's own facial recognition was identified in the Gender Shades study as having error rates up to 34.7% for darker-skinned women. Continuous bias auditing was implemented after, not before, that finding. (3) Robustness — Tay failed within 24 hours of public launch because no adversarial testing had evaluated how the system would behave when public users tried to break it. Robustness requires testing across the full intended deployment range, not just standard conditions. (4) Privacy — health AI systems globally collect patient data under limited consent for clinical care and use it for commercial model training. Privacy governance must cover data throughout its lifecycle, including secondary uses.
Question 3 of 3

What are AI Fairness 360 and AI Explainability 360 — and what does IBM's decision to open-source them reveal about how serious AI governance actually works?

AI Fairness 360 is an open-source Python toolkit containing 70+ fairness metrics and 10+ bias mitigation algorithms for detecting and reducing bias in ML datasets and models. AI Explainability 360 provides 8 explainability methods covering both local (why this decision?) and global (how does the model work?) explanations for different model types. IBM released both after the Watson and Gender Shades failures. The decision to open-source them reveals two things: first, that IBM treated governance reform as infrastructure that needed to be embedded into the development process across the industry, not just inside IBM; second, that serious governance is operationalised through tools that practitioners can actually use, not through principles documents. The fact that both are now used globally by organisations with no IBM relationship indicates that the tooling filled a real gap in how the industry builds accountability into model development.

Module 5 Videos

Module 5 · Short Video
Module 5 · Long Form · Whose Name Is On That Decision?

Sources

IBM Research
IBM (2021). Everyday Ethics for Artificial Intelligence. IBM Corporation.
Gender Shades
Buolamwini, J. & Gebru, T. (2018). Gender Shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of Machine Learning Research, 81, 1–15.
Watson/MD Anderson
Ross, C. & Swetlitz, I. (2017). IBM pitched its Watson supercomputer as a revolution in cancer care. It's nowhere close. STAT News, 5 September 2017.
Module content
BUSN9049 Module 5 — Ethical Considerations and Responsible AI. Flinders University, 2026.
Innovare Study
Long-form video: Whose Name Is On That Decision? Innovare Study, July 2026. youtu.be/tslQKmfxwk0