Frame selection is the whole game: notes on making LLMs watch video
Hacker News
Read full postFeeding videos directly to large language models (LLMs) allows them to process raw visual information rather than relying on human-written summaries, which can omit important details. Due to token limits, selecting the most informative video frames is crucial, with scene detection and adaptive sampling methods improving frame selection over uniform sampling.



