Meta releases Muse Glimmer: 30B open-weights multimodal model for local agentic tasks
Meta's Muse Glimmer distills its multimodal foundation to 30B parameters under Apache 2.0, targeting local deployment for coding, document analysis, and agent workflows.
Avatarin Built a 24/7 Retail Agent on GPT-4 Realtime, Reaching 30K Shoppers in Two Weeks
The ANA Holdings spinout deployed voice-first customer service at Yamada Denki, proving multimodal real-time models can replace human sales expertise at scale.
Google DeepMind Launches Gemini Robotics ER 2 With Real-Time Video Understanding and Multi-Robot Coordination
Gemini Robotics ER 2 enables robots to reason about physical tasks, collaborate across multiple units, and adapt in real time using continuous video feeds.
Anthropic Extends Voice Mode to Claude Opus and Sonnet, Adding Multilingual Support
Anthropic's faster models gain voice capabilities, enabling deeper problem-solving through speech in nine new languages.
Anthropic expands Claude voice mode to Opus and Sonnet, adds app integrations
Claude voice mode now supports three model tiers and integrates with Gmail, Slack, and Notion—positioning Anthropic against OpenAI's voice capabilities.
Thinking Machines Lab Launches Inkling, a 975B-Parameter Open-Weight Model Built by OpenAI Defectors
The startup founded by former OpenAI executives releases its flagship multimodal model, challenging the dominance of closed-source systems.
Thinking Machines Releases Inkling, a 1-Trillion-Parameter Multimodal Open Model
Inkling combines native image, audio, and text processing with a 1M-token context window and sparse MoE architecture for efficient multimodal reasoning.
Hugging Face and Cerebras Demonstrate Real-Time Speech-to-Speech with Gemma 4
A modular voice AI pipeline achieves sub-second latency by pairing Google DeepMind's Gemma 4 model with Cerebras inference acceleration, powering conversational robots and assistants.
Lumo 2.0 Brings Image Recognition and Persistent Memory to Proton's Privacy-First Chatbot
Proton's encrypted AI assistant gains multimodal capabilities and 76% faster response times while maintaining zero-access encryption.
Google DeepMind Launches Gemini 3.5 Live Translate with Near-Real-Time Speech-to-Speech Across 70+ Languages
Gemini 3.5 Live Translate enables fluid, continuous speech-to-speech translation without manual language configuration, rolling out to Google Meet, Translate, and the Gemini Live API.
Google DeepMind's Gemma 4 12B Brings Encoder-Free Multimodal AI to Consumer Laptops
Google DeepMind releases Gemma 4 12B, a 12-billion-parameter model with unified vision and audio processing that runs on 16GB consumer hardware.
Google Unveils Gemini Omni Video Generation and Gemini 3.5 Flash for Agentic AI
Google announced Gemini Omni, a multimodal model that generates and edits video through natural language, and Gemini 3.5 Flash, optimized for complex agent workflows.
Google I/O 2026: Gemini Omni, Multimodal Search, and AI Agents debut
Google unveiled Gemini Omni for video generation, Gemini 3.5 Flash for agents and coding, and autonomous Search agents that monitor the web 24/7.
Google launches Gemini 3.5 Flash with agentic benchmarks outpacing Pro, rolls Pro in June
Google unveiled Gemini 3.5 Flash at I/O 2026 as an agent-first model claiming frontier intelligence at sub-flagship latency, with Gemini Omni adding physics-aware video generation.
Google's Gemini Omni blurs the line between text prompt and video simulation
At I/O 2026, Google DeepMind unveiled Gemini Omni, a multimodal family that generates video from combined image, audio, and text inputs, signaling a shift from generative to simulational AI.
Google launches Gemini 3.5 models and Omni multimodal family at I/O 2026
Google unveiled Gemini 3.5 Flash as the new default model, introduced Gemini Omni for text-to-video generation, and previewed always-on agents powered by Gemini Spark.
Google DeepMind Launches Gemini Omni Flash for AI-Powered Video Generation and Editing
Gemini Omni Flash enables users to generate and edit videos through natural language prompts, combining multimodal inputs with real-world knowledge.
Google Search Gets Agentic AI Overhaul With Gemini 3.5 Flash Default
Google debuts AI agents, a redesigned search box, and Gemini 3.5 Flash integration in Search, targeting a billion-user base
Google I/O 2026: Gemini Omni and Agent-First Development Mark Shift Toward Agentic AI
Google unveiled Gemini Omni, a multimodal model capable of video creation, alongside Gemini 3.5 Flash and expanded agent capabilities across Search, Gmail, and shopping.
Elmer Data's Watch Test Exposes a Gap Between Conversational AI and Visual Reasoning
A new analysis shows that large language models excel at language tasks but struggle with seemingly simple visual reasoning—like reading analog clocks.