multimodal-ai
11 posts tagged multimodal-ai.
- OmniScientist argues that AI scientists need eyes, not just workflows
- AMIE’s video consult result is about perception, not replacement
- Video deep research agents need to look before they search
- MODUS brings any-to-any multimodal modeling to decoder-only systems
- Multimodal AI needs a plan for missing inputs
- MIRROR trains vision models by making each modality teach the others
- Sarcasm detection needs the mismatch, not just the meme
- Visual pretraining is a bet against text extraction
- Claude-real-video points to video as an adapter problem
- Speaker recognition is a better agent test than another chat demo
- Multimodal models still change answers when you shuffle the evidence