Our research
You don't design the failure in advance. You find it in use.
The question
How do AI systems fail when the people responsible for human life use them, and how do we use them without losing control?
How we do it
There are two different kinds of research. In the lab, you set up a control. You test one thing at a time, under conditions you decide in advance.
What research looks like for me is very different. We publish through our own tools, and AI works in our own database, on our real work. I give it exposure to a lot of different things on purpose, because that is how you understand where it works and where it fails.
What it produces
These use cases, these logs, the plain record of what a tool did when it got something wrong, are what testing produces. They should exist everywhere these tools are being brought in.
We have been collecting failure data for years. We collected it for a reason I did not know yet, and sometimes that's what you do: you collect things until you understand what the relevance is going to be.
63 cases logged · 21 kinds of failure