Skip to content
{ ken ashe }
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
  • Building
  • Writing
  • About
  • Newsroom
  • Digest
← Digest / Tags

// tag

multimodal-ai

28 posts tagged multimodal-ai.

  • Audio models do not automatically share speech and text concepts Sep 25, 2026
  • The Modality Gap in Speech Fact-Checking, and Why Retrieval Alone Doesn't Fix It Sep 25, 2026
  • DiaVLo Turns Vision-Language Model Failures Into Named Behaviours Sep 21, 2026
  • Full-Duplex Voice With Tool Calls: What NemotronLabs VoiceChat Actually Ships Sep 21, 2026
  • Qwen-Image-2.1 puts open image editing closer to production work Sep 20, 2026
  • AI posters get better when the model stops being the designer Sep 19, 2026
  • Paint-Anything makes hex colors a first-class diffusion control Sep 18, 2026
  • Camera-free smart glasses would test what AI wearables are really for Sep 17, 2026
  • MUSE tests vision-language models where classroom context gets messy Sep 17, 2026
  • LACE compresses speech tokens one codec layer at a time Sep 16, 2026
  • Slip Detection Is Where Robot Hands Stop Dropping Things Sep 15, 2026
  • JPEG XL Is a Workflow Question, Not a Format War Sep 14, 2026
  • MindTopo Tests Whether Vision Models Grasp Topology, Not Just Distances Sep 11, 2026
  • Instruction vs. Example: How Vision-Language Models Actually Moderate Content Sep 10, 2026
  • NOAH models patient records as timelines, not snapshots Sep 9, 2026
  • What ChatGPT Images 2.5 Changes for People Who Actually Ship Images Sep 9, 2026
  • VBVR-Pro makes visual reasoning a training loop, not a demo Aug 27, 2026
  • OmniScientist argues that AI scientists need eyes, not just workflows Aug 14, 2026
  • AMIE’s video consult result is about perception, not replacement Aug 11, 2026
  • Video deep research agents need to look before they search Aug 5, 2026
  • MODUS brings any-to-any multimodal modeling to decoder-only systems Jul 29, 2026
  • Multimodal AI needs a plan for missing inputs Jul 28, 2026
  • MIRROR trains vision models by making each modality teach the others Jul 24, 2026
  • Sarcasm detection needs the mismatch, not just the meme Jul 20, 2026
  • Visual pretraining is a bet against text extraction Jul 13, 2026
  • Claude-real-video points to video as an adapter problem Jul 4, 2026
  • Speaker recognition is a better agent test than another chat demo Jul 3, 2026
  • Multimodal models still change answers when you shuffle the evidence Jun 25, 2026

Ken Ashe ·AI application builder ·CPA ·PMP

Building with AI in public. No hype, no doom. Receipts only.

hello@kenashe.ai

Explore

  • Building
  • Writing
  • Digest
  • Topics

About & Press

  • About
  • Newsroom
  • Media Kit
  • Lucky Domains

Social

  • LinkedIn
  • X
  • GitHub
  • RSS

Legal

  • Privacy
  • Terms
  • Disclosure

© 2026 Ken Ashe ·Built with AI in public