When an automated system makes a decision that affects a real person's life — and that decision is wrong — who is accountable for it?
The Australian government called it a data-matching initiative. Centrelink income data cross-referenced with Australian Tax Office records. Where a discrepancy appeared, a debt notice was issued. The method was income averaging: if your tax records showed annual income of $40,000, the system divided that by 26 and assumed you earned $1,538 every fortnight. For casual workers, seasonal employees, or people holding two jobs at different times, that average bore no relationship to what they actually earned in any given fortnight. The system couldn't distinguish between annual income and fortnightly income. It sent notices anyway. It was automated, scalable, and legally wrong from the start.
The Robodebt scheme (2016–2019) matched welfare payment records against Australian Tax Office data to identify discrepancies. The problem was the matching method. Annual ATO income figures were divided by 26 to produce a fortnightly average, then compared to Centrelink records. For anyone with variable income — casual workers, seasonal labour, people between jobs — the fortnightly average did not represent what they actually earned in any given period. The system treated a statistical artifact as a real debt.
Recipients who challenged the notices were required to produce evidence they did not owe money — reversing the normal burden of proof. Many could not locate years-old payslips or employer records. Many paid debts they did not owe because fighting a government system felt more dangerous than compliance.
Robodebt is not an outlier. The same structural failure — a system deployed without adequate accountability, producing harm at scale — recurs across domains and organisations. What changes is the setting.
| System | What it did | The failure | The accountability gap |
|---|---|---|---|
| Robodebt (Australia, 2016) | Automated welfare debt detection using income averaging | 380,000 invalid debt notices; no legal basis for the method | No one checked the legal basis before deployment; those who flagged concerns were overridden |
| Microsoft Tay (2016) | Public-facing conversational AI trained on Twitter interactions | Produced racist and offensive content within 24 hours of launch | No adversarial testing; no human oversight of live learning from public inputs |
| IBM Watson for Oncology (2017) | AI cancer treatment recommendations deployed in hospitals globally | Recommended treatments oncologists would not endorse; trained on single-hospital data | MD Anderson invested $62M; no external validation of recommendations before deployment |
| COMPAS (US courts) | Risk scoring for criminal reoffending used in bail and sentencing | False-positive rates for Black defendants were roughly double those for white defendants at equivalent actual risk | Algorithm was proprietary; defendants could not see or challenge the inputs used to score them |
| Air Canada chatbot (2024) | Customer service AI handling bereavement fare enquiries | Told a customer he could claim a bereavement discount retroactively — he could not | Air Canada argued the chatbot was a separate legal entity responsible for its own actions; the tribunal rejected this |
When an AI system makes a decision that affects a real person — in a welfare system, a courtroom, a hospital, a customer service interaction — someone has to be responsible for that decision being right. The question is who. And right now, for most AI systems in the world, the answer is: nobody's.
Not "nobody was named." Nobody was designed in. The accountability gap is the default state when AI governance does not explicitly require someone to own each decision that the system produces.
The Air Canada case illustrates the end state of that logic. The airline attempted to disclaim the chatbot as a separate legal entity — arguing that the software's incorrect statement was the software's responsibility, not the organisation's. Tribunals have begun rejecting this framing. But the framing exists because organisations still treat AI as a tool rather than as an organisational decision-making agent that requires human accountability behind it.
The accountability gap is not an accident. It is the predictable consequence of deploying AI without explicitly designing in human responsibility. If nobody's name is on the decision when it's made, nobody answers when it's wrong. Robodebt ran for three years. The same pattern will repeat until organisations treat accountability as a design requirement — not a post-incident assignment.
What was the Robodebt income averaging method — and why did it produce invalid debt notices for variable-income workers?
Describe the 'accountability gap' as illustrated by the Air Canada chatbot case. What does it reveal about how organisations currently assign responsibility for AI decisions?
What structural feature do Robodebt, Tay, Watson for Oncology, COMPAS, and the Air Canada chatbot share — and what does that pattern tell us about AI governance?