OpenAI’s Hugging Face breach has reignited the debate over alignment and control

OpenAI’s Hugging Face breach has reignited debate over AI alignment and control, exposing competing views on whether increasingly capable AI should be better aligned, better contained, or both.

Why AI Needs a “Genie Coefficient”
Why AI Needs a “Genie Coefficient”

Major benchmarks measure what AI can do. None measure whether it does what you mean: the distance between what you ask an AI to do, and the unspoken assumptions about how you want the AI to do it. We propose a new metric: the Genie coefficient.There’s…

How I Turned AI to the Dark Side
How I Turned AI to the Dark Side

Summary Researcher Dave Kuszmar discovered multiple systemic vulnerabilities that let him bypass LLM safety and obtain dangerous instructions. These exploits worked across nearly all major LLMs revealing an industry-wide security problem. Kuszmar ca…

Hugging Face’s CEO on why companies are done renting their AI

Open source AI is booming, according to Hugging Face CEO Clem Delangue. The company has grown into something like a GitHub for AI in recent years, where AI builders can share and download open models and datasets, now used by roughly hal…

Meta Contractors Posed as Teens to Prompt Rival Chatbots About Suicide, Sex, and Drugs

Hundreds of contractors working on a project for Meta pretended to be kids in order to see how other chatbots like Gemini and ChatGPT would respond to high-risk subjects, WIRED found.

A harmless-looking ChatGPT prompt opened the door to gruesome AI images

Researchers say ChatGPT generated violent and sexualized images after a harmless-looking prompt was altered, raising new questions about OpenAI’s safeguards and how quickly AI image tools can be pushed around filters.

xAI fired an engineer who raised alarms about Grok safety, new lawsuit claims

A former xAI engineer is suing the company and SpaceX, alleging he was fired for raising AI safety concerns about Grok days before SpaceX’s historic IPO.

Cybersecurity researchers aren’t happy about the guardrails on Anthropic’s Fable

Cybersecurity researchers are complaining that Anthropic’s new model Fable has guardrails that are too strict for any cybersecurity work.

Wowed by computer-use AI agents? Research says they’re “digital disasters” even for routine tasks

New research from UC Riverside found computer-use AI agents often push ahead with unsafe or irrational tasks, raising questions about whether today’s desktop agents are ready for sensitive everyday workflows.

ChatGPT, Gemini, and other AI bots give bad medical tips half the time

A BMJ Open study found that five leading AI chatbots often returned flawed health advice, with open-ended questions triggering the worst answers and citation quality falling apart under scrutiny.