October 09, 2026 06:31 am (IST)
Follow us:
facebook-white sharing button
twitter-white sharing button
instagram-white sharing button
youtube-white sharing button
‘Bold and inventive’: Canadian poet Anne Carson wins 2026 Nobel Prize in Literature | ‘Attempt to destabilise world’s largest democracy’: 42 retired judges rally behind Gyanesh Kumar over ‘vote chori’ row | ‘Vote chor, gaddi chod’: Dhruv Rathee, Prakash Raj join Bengaluru protest against Gyanesh Kumar | India condemns Houthi attacks on Saudi airports, calls targeting of civilian infrastructure unacceptable | Dalal Street bloodbath: Investors lose Rs. 8 lakh crore as Sensex crashes over 1,000 points | Amazon layoffs hit India, US and UK: Fresh job cuts rock e-commerce giant amid festive shopping season | Bollywood mourns Nana Patekar: Anupam Kher, Akshay Kumar, Jr NTR, Suniel Shetty pay emotional tributes | ‘He was never wavering while expressing his opinions’: PM Modi mourns Nana Patekar | Nana Patekar dies at 75: Veteran actor suffers cardiac arrest at Goa home | RBI shocks borrowers with first repo rate hike since 2023; rates raised to 5.50%
Google
Google logo. Photo: Unsplash

Google drops Gemma 4 12B: A game-changing AI that runs on your laptop

| @indiablooms | Jun 07, 2026, at 05:18 pm

Google has announced the launch of Gemma 4 12B, a dense multimodal model featuring a unified, encoder-free architecture.

Gemma 4 12B marks several key milestones for local AI development. According to Google’s blog post, it introduces a multimodal encoder-free design, eliminating the need for heavy, multi-stage vision and audio encoders. Instead, multimodal inputs are fed directly into the LLM backbone, helping reduce latency in processing images, audio, and other data types.

The company also described it as its first medium-sized model with native audio input. Within the Gemma family, audio capabilities were previously limited to smaller edge-focused models such as E4B. With Gemma 4 12B, Google expands audio understanding to a more capable, general-purpose model.

Positioned as developer-friendly and locally deployable, the model is compact enough to run on laptops equipped with 16GB VRAM or unified memory. To further optimize local inference speed, Google is also releasing a dedicated multi-token prediction (MTP) model.

For the first time, Google is also introducing downloadable macOS desktop applications, enabling developers to experience fully local, real-time multimodal interaction—including voice and visual inputs—on consumer-grade devices.

In its technical overview, Google noted that traditional multimodal systems typically rely on separate, frozen encoders for different modalities, such as vision encoders (150M parameters for edge models and 550M for medium models) and audio encoders (around 300M parameters in smaller variants like E2B and E4B).

Google claims Gemma 4 12B delivers strong performance across a range of capabilities, including automatic speech recognition, agentic reasoning, speaker diarization, video understanding, and coding tasks.

Support Our Journalism

We cannot do without you.. your contribution supports unbiased journalism

IBNS is not driven by any ism- not wokeism, not racism, not skewed secularism, not hyper right-wing or left liberal ideals, nor by any hardline religious beliefs or hyper nationalism. We want to serve you good old objective news, as they are. We do not judge or preach. We let people decide for themselves. We only try to present factual and well-sourced news.

Support objective journalism for a small contribution.