AI Models
New multimodal model watches hours of video and answers like it was there
A frontier system can now ingest full-length video with synchronized audio and reason across it, unlocking applications from surgery review to sports analytics.
By Maya Okafor, Senior AI Correspondent — SAN FRANCISCO
SAN FRANCISCO — A newly released multimodal model can ingest hours of continuous video with synchronized audio and answer detailed questions about any moment in it, its developer said this week, citing verified evaluations on long-form comprehension benchmarks.
Early customers are pointed at unglamorous but valuable work: reviewing surgical recordings for training, auditing safety-compliance footage, indexing broadcast archives and generating same-day analytic breakdowns of sports matches.
The release also sharpens familiar concerns: civil-liberties groups warned that fluent machine comprehension of surveillance footage arrives faster than the rules governing it.
Enable JavaScript to read the full story on Neural Daily News.