01WeChat Open-Sources WeMM-Embedding Multimodal Vector Model SuiteRSS · IT HOME · ithome.com · Sep 04, 01:21 PM2d
02Sori-1B: An Audio-Language Model Trained From Scratch Without Text-Only PretrainingREDDIT · /u/Balance- · reddit.com · Aug 31, 02:48 AM6d
03Tencent Releases WeMM-Embedding, a Multimodal Embedding Model FamilyREDDIT · /u/jacek2023 · reddit.com · Aug 25, 09:36 AM12d
04SenseTime Open-Sources SenseNova U1.5 Lite Multimodal Model with Native 4K OutputRSS · IT HOME · ithome.com · Aug 21, 04:20 AM16d
05AI Tennis Coach and ICRA Clothes-Folding Robot Champion UnveiledRSS · 量子位 · mp.weixin.qq.com · Aug 20, 07:56 AM17d
06Giving DeepSeek V4 Flash Vision via a 40M-Parameter ConnectorREDDIT · /u/ButtercupLyn100 · reddit.com · Aug 11, 03:45 AM26d
07ByteDance Launches SeedRealtime, a Native Audio-Video Full-Duplex AI ModelRSS · IT HOME · ithome.com · Aug 05, 04:12 AM32d
08BAAI and Peking University Develop Joint Audio-Video Editing via Single Text InstructionRSS · 量子位 · mp.weixin.qq.com · Aug 04, 09:00 AM33d
09SenseTime Open-Sources SenseNova U1.5-Lite-Preview with Native 4K Image GenerationRSS · IT HOME · ithome.com · Aug 03, 07:34 AM34d
10MiniMax-H3 Omni-Modal Generative Model Released on Hugging FaceREDDIT · /u/Mobile-Pumpkin7944 · reddit.com · Aug 03, 03:06 AM34d
11LMSYS Arena Expands Preference-Based Reward Modeling to Multimodal DomainsTWITTER · arena · x.com · Jul 30, 05:18 PM38d
12Microsoft Releases Mage-VL: An Efficient Codec-Native Streaming Multimodal ModelREDDIT · /u/pmttyji · reddit.com · Jul 28, 06:47 PM40d
13Douyin Upgrades Minor Mode Recommendation Engine with Multimodal AIRSS · IT HOME · ithome.com · Jul 27, 08:41 AM41d
14Black Forest Labs Launches FLUX 3 Multimodal AI ModelRSS · IT HOME · ithome.com · Jul 24, 06:48 AM44d
15Zoom Tool Dramatically Boosts Vision Model Accuracy on Chartography BenchmarkTWITTER · ClaudeDevs · x.com · Jul 23, 11:25 PM45d
16Anthropic Updates Claude "Zoom Tool" Cookbook for High-Resolution Image AnalysisTWITTER · ClaudeDevs · x.com · Jul 23, 11:25 PM45d