LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes

Hacker News
Read full post
Researchers evaluated large language model (LLM) judges' ability to detect omissions in AI-generated clinical notes by comparing flawed notes with transcripts. They found standard LLM judges struggle to reliably identify missing information but developed new methods that improve omission detection with acceptable false alarm rates. These methods were validated by physicians and released as benchmarks and tools for further research.

More in Healthcare

UK Is Urged to Overhaul Regulation of AI-Medical Devices

Covered by 2 sources
Healthcare6 min read

NYU-DRP AI Model Predicts Five-Year Breast Cancer Risk From 3D Mammograms

Unite.AI

Google DeepMind releases AlphaGenome Atlas, a 1PB dataset of predicted molecular effects for all ~9B possible single-letter DNA changes in the human genome (Google)

Covered by 4 sources