THE VERTICAL MUSEUM · EXHIBIT 02
The AI that Peaked at Trivia
Welcome back to The Vertical Museum, our recurring series on the winners, losers & ledgers from vertical history. Every two weeks, one exhibit: a system purpose-built to run an industry.
App — Watson for Oncology
Vertical — Cancer Care
Insight — Imitation
Developer — IBM
Date — 2011–2020
Verdict — Loser: never entered clinical use
The artifact is a room, which is awkward for a museum: 90 Power servers stacked across ten racks. IBM named the cluster Watson, after its first CEO.
In February 2011, IBM put Watson on Jeopardy!, where it defeated Ken Jennings and Brad Rutter, two of the show’s best players. Jennings took the loss well. “I, for one, welcome our new computer overlords,” he wrote on his Final Jeopardy screen.
Watson, it turned out, had peaked at trivia.
The Job
After beating Ken Jennings, Watson got a promotion: oncology.
IBM split the medical push between two famous cancer centers. Memorial Sloan Kettering would teach Watson for Oncology to recommend treatments. MD Anderson hired IBM to build a related system called the Oncology Expert Advisor. The assignment sounded simple: ingest the patient information, review the literature, and return a ranked list of options supported by evidence.
Oncology was less cooperative. A treatment decision can turn on tumor type, stage, mutations, prior therapies, other conditions, incomplete tests, available drugs, new evidence, and what the patient is willing to endure. Two reasonable oncologists can see the same case and choose different paths. In contrast, Jeopardy! gives you a complete clue and one correct answer. Patients are less considerate.
The Clever Bit
Much of Watson for Oncology’s later training came from hypothetical cases written by MSK doctors, with the recommended answers attached. It became less an oncology oracle than an extremely expensive impression of one Manhattan cancer center.
That imitation traveled badly. Internal IBM documents recorded multiple “unsafe and incorrect” test recommendations. Doctors outside the US found that some recommended drugs weren’t available in their countries, or that the advice conflicted with local protocols. In one case, Watson suggested a drug with a severe bleeding warning for a patient who was already bleeding. MSK caught it before a real patient received the advice. On Jeopardy!, a wrong answer costs you money. Healthcare has a less forgiving scoring system.
At MD Anderson, the failure was less dramatic but more expensive. Its Oncology Expert Advisor struggled to ingest patient data, integrate with the hospital’s electronic record system, and produce useful recommendations. MD Anderson spent $62.1 million on its state-of-the-art oncology AI. Five years later, after only testing it on old cases, the hospital pulled the plug. It never treated a single patient.
The invoices, at least, were in excellent health.
Why It’s in the Museum
Watson belongs here because answering a question correctly and doing the job are not the same thing.
Many studies reported concordance – how often Watson agreed with doctors. But agreement wasn’t the target outcome. It showed no clinical usefulness, and most of the time during training, doctors would close Watson and keep working. A second opinion that is useful only when it agrees with the first is mostly a very expensive nod.
By the end of 2020, IBM shelved Watson for Oncology. Two years later, it sold much of Watson Health’s data business for parts.
Kushim’s clay tablet survived millennia. Watson just barely survived a decade.
The Catalog So Far
Exhibit 01 — The Beer Ledger of Uruk
Exhibit 02 — The AI that Peaked at Trivia
Next up: the machine that ruined Mark Twain’s life.


