multimodal-ai
28 posts tagged multimodal-ai.
- Audio models do not automatically share speech and text concepts
- The Modality Gap in Speech Fact-Checking, and Why Retrieval Alone Doesn't Fix It
- DiaVLo Turns Vision-Language Model Failures Into Named Behaviours
- Full-Duplex Voice With Tool Calls: What NemotronLabs VoiceChat Actually Ships
- Qwen-Image-2.1 puts open image editing closer to production work
- AI posters get better when the model stops being the designer
- Paint-Anything makes hex colors a first-class diffusion control
- Camera-free smart glasses would test what AI wearables are really for
- MUSE tests vision-language models where classroom context gets messy
- LACE compresses speech tokens one codec layer at a time
- Slip Detection Is Where Robot Hands Stop Dropping Things
- JPEG XL Is a Workflow Question, Not a Format War
- MindTopo Tests Whether Vision Models Grasp Topology, Not Just Distances
- Instruction vs. Example: How Vision-Language Models Actually Moderate Content
- NOAH models patient records as timelines, not snapshots
- What ChatGPT Images 2.5 Changes for People Who Actually Ship Images
- VBVR-Pro makes visual reasoning a training loop, not a demo
- OmniScientist argues that AI scientists need eyes, not just workflows
- AMIE’s video consult result is about perception, not replacement
- Video deep research agents need to look before they search
- MODUS brings any-to-any multimodal modeling to decoder-only systems
- Multimodal AI needs a plan for missing inputs
- MIRROR trains vision models by making each modality teach the others
- Sarcasm detection needs the mismatch, not just the meme
- Visual pretraining is a bet against text extraction
- Claude-real-video points to video as an adapter problem
- Speaker recognition is a better agent test than another chat demo
- Multimodal models still change answers when you shuffle the evidence