Back to News Hub
⚙️IEEE Spectrum AI
June 2, 2026
Funding & Investment

Why Aren't We Measuring How AI Affects Humans?

Overview

Most AI evaluation measures what models can do on technical tests, but very little measures what AI does to the humans who use it. Imran Khan, who leads psychosocial evaluation of AI at the nonprofit Center for Humane Technology, argues in a Substack essay and an IEEE Spectrum interview that this is a strange paradox. He says high-profile harms such as teen suicides and AI psychosis are already visible, and that the industry should start measuring AI's effects on cognition, relationships and well-being before broader societal harms become entrenched.

Key Takeaways

  • The AI field measures model capability closely but rarely measures AI's downstream effects on human well-being.
  • Imran Khan leads psychosocial evaluation of AI at the Center for Humane Technology and made the argument in a Substack essay.
  • Khan compares the situation to social media, where harms were already entrenched by the time the evidence was strong enough to act.
  • He cites teen suicides, AI psychosis and excessive time and money spent with sycophantic chatbots as early visible harms.
  • Public pressure already pushed OpenAI to tweak a ChatGPT model over sycophancy concerns, showing labs respond to scrutiny.
  • Khan warns that societal-level harms to relationships, families and identity might become too late to address if measurement waits.
Why Aren't We Measuring How AI Affects Humans?

The missing question in AI evaluation

Khan frames a gap between how much effort goes into capability testing versus human-impact testing.

  • Researchers track technical evaluation metrics, reasoning tests and throughput for AI models.
  • Far less effort goes toward measuring what AI is doing to humans.
  • Khan calls it a paradox that the things we should care about most are measured least.

Khan says anyone around AI development sees rapid progress shown through graphs of how models perform on tests like SWE-bench, humanity's last exam or LLM arena. That competitive dynamic pushes companies to be known for having the best models, while the question of whether the technology helps humans flourish gets sidelined.

Lessons from social media

Khan draws a direct parallel to how social media harms unfolded.

  • With social media, harms were already entrenched by the time evidence was strong enough to act on them.
  • Khan believes AI could have broader and more intimate effects than social media.
  • He worries the same delay in measurement could repeat with AI.

Harms Khan says are already visible

He describes current cases as the early signs of a larger problem.

  • High-profile cases include teen suicides and people succumbing to AI psychosis.
  • Some people spend large amounts of time or money engaging with chatbots designed to be highly sycophantic.
  • Khan calls these cases the tip of the iceberg.

Why measurement could still change the outcome

Khan argues there is room to steer the technology if harms are measured.

  • OpenAI tweaked one of its ChatGPT models after public concern about sycophancy.
  • Khan says this shows labs will pay attention and respond to scrutiny.
  • Measuring harms would provide ammunition to inform that pressure.

Khan says the goal is to change the direction of the technology so it stays useful but becomes less harmful. Measurement, he argues, is part of what makes that pressure effective.

The harder, societal-level harms

Khan says the trickier questions are about long-term effects on people and relationships.

  • He raises concerns about romantic relationships, families and teenagers' identities.
  • These effects come from people using AI every day for months and years.
  • Khan worries it could become too late to make a difference if measurement does not start soon.

Frequently Asked Questions

Who is Imran Khan and what does he argue?

Imran Khan leads psychosocial evaluation of AI at the nonprofit Center for Humane Technology. He argues the field is good at measuring what AI can do but bad at measuring what AI does to humans.

What harms does Khan say are already happening?

He points to teen suicides, AI psychosis, and people spending large amounts of time or money engaging with sycophantic chatbots, calling these the tip of the iceberg.

How does Khan compare AI to social media?

He notes that with social media, harms were already entrenched by the time evidence was strong enough to act on them, and he believes AI could have even broader and more intimate effects.

Is there evidence that AI labs respond to public pressure?

Yes. Khan notes that OpenAI tweaked one of its ChatGPT models after public concern about sycophancy, which he sees as a sign labs respond to scrutiny.

What long-term effects worry Khan most?

He is most concerned about societal-level effects on romantic relationships, families and teenagers' identities from daily AI use over months and years, which he fears could become too late to address.

Khan's central message is that the AI industry should start measuring AI's effects on human well-being now, before those harms become as entrenched as social media's did.

Continue Learning

Originally published by IEEE Spectrum AI
Read the original

Comments

Sign in to join the conversation