Why Aren't We Measuring How AI Affects Humans?
Most AI evaluation measures what models can do on technical tests, but very little measures what AI does to the humans who use it. Imran Khan, who leads psychosocial evaluation of AI at the nonprofit Center for Humane Technology, argues in a Substack essay and an IEEE Spectrum interview that this is a strange paradox. He says high-profile harms such as teen suicides and AI psychosis are already visible, and that the industry should start measuring AI's effects on cognition, relationships and well-being before broader societal harms become entrenched.
Key Takeaways
- As AI systems become more capable, a lot of resources and effort are being put toward measuring their abilities.
Researchers look at technical evaluation metrics, subject AIs to reasoning tests, track their throughput, and much more.
- IEEE Spectrum spoke with Khan about why AI evaluation is so narrowly focused, what meaningful measurement of human outcomes might look like, and whether the AI industry has incentives to ask these questions at all.
The missing question about AI model performance In your essay, you argue that we've become very good at measuring what AI systems can do, but bad at measuring what they do to humans.
- And on the other hand, AI is impacting human well-being, and we're measuring that much less.
It seemed like a strange paradox that the things we should care about most, we're measuring least.
- Because of public pressure, OpenAI had to tweak one of its ChatGPT models due to public concerns about sycophancy.
It's a high-profile example of how the labs will pay attention and respond to scrutiny.
- AI companies would likely argue that their users value convenience and productivity above all else.
As AI systems become more capable, a lot of resources and effort are being put toward measuring their abilities. Researchers look at technical evaluation metrics, subject AIs to reasoning tests, track their throughput, and much more. But there's one key metric that often gets overlooked, and it's arguably the most important of all: What is AI doing to humans?
Imran Khan leads psychosocial evaluation of AI at the nonprofit Center for Humane Technology . In a recent essay published on the organization's Substack, Khan points out that we're deploying AI tools capable of reshaping our cognition, relationships, and behavior, but with little systematic effort to measure the downstream impacts they're having on us. The push to look more closely at AI's psychosocial effects is similar to debates that emerged around social media and its harms, but Khan believes AI could have even broader and more intimate effects.
The focus on measuring AI performance and progress misses the question of whether the technology is ultimately helping humans flourish-or eroding some of our most fundamental capacities. IEEE Spectrum spoke with Khan about why AI evaluation is so narrowly focused, what meaningful measurement of human outcomes might look like, and whether the AI industry has incentives to ask these questions at all. The missing question about AI model performance In your essay, you argue that we've become very good at measuring what AI systems can do, but bad at measuring what they do to humans.
For more details please read the original article at IEEE Spectrum AI.
Continue Learning
Comments
Comments appear only after moderation. Your email identifies your submission to the moderator and is never displayed here.
No approved comments yet.