AI ResearchMachine Learning7 min reading time

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

Hugging Face
Read full post
A new paper introduces a method for large language models to refuse only specific harmful subsets within a topic, rather than rejecting entire topics, enabling more precise safety controls tailored to deployment needs.

More in AI Research

Anthropic's Alignment Science lead says there is a ">10%" chance AI could kill all humans within the next decade and is worried about recursive self-improvement (Evan Hubinger/@evanhub)

Covered by 9 sources
AI Research4 min read

Suno trained its v6 AI music models with help from Warner and BMG

Covered by 5 sources

Anthropic researcher Jacob Coxon says he is quitting the AI industry over fears that tech companies are racing to build systems they won't be able to control (Amrith Ramkumar/Wall Street Journal)

Covered by 11 sources