Models model benchmark training release open-source billion

Gemma 4 Architecture Comparison

Reddit r/LocalLLaMAby /u/seraschka https://www.reddit.com/user/seraschkaApril 3, 20262 min read1 views

Flagship open-weight release days are always exciting. Was just reading through the Gemma 4 reports, configs, and code, and here are my takeaways: Architecture-wise, besides multi-model support, Gemma 4 (31B) looks pretty much unchanged compared to Gemma 3 (27B). Link to the comparison page: https://sebastianraschka.com/llm-architecture-gallery/?compare=gemma-3-27b 2Cgemma-4-31b Gemma 4 maintains a relatively unique Pre- and Post-norm setup and remains relatively classic, with a 5:1 hybrid attention mechanism combining a sliding-window (local) layer and a full-attention (global) layer. https://preview.redd.it/7bn493789zsg1.png?width=1444 format=png auto=webp s=4b28421ed276cb0b1ba133e3c325d446d68ea1ef The attention mechanism itself is also classic Grouped Query Attention (GQA). But let’s no

Could not retrieve the full article text.

Read on Reddit r/LocalLLaMA →

Original source

Reddit r/LocalLLaMA

https://www.reddit.com/r/LocalLLaMA/comments/1sbdr75/gemma_4_architecture_comparison/

Was this article helpful?

Ask AI about this article

Ready

Conversation starters

Ask anything about this article…

Daily AI Digest

Get the top 5 AI stories delivered to your inbox every morning.

More about

modelbenchmarktraining

ModelsFresh

DeepSeek V4 draait op Huawei-chips en omzeilt afhankelijkheid Nvidia

Het grote taalmodel DeepSeek V4 draait naar verluidt op chips van Huawei. Tot dusver was het Chinese bedrijf achter de AI-dienst afhankelijk van Nvidia-processors, die onder Amerikaanse exportrestricties vallen. Naar verwachting komt het model deze lente uit.

Tweakers.net

1mabout 4 hours ago

ProductsLive

Studying Human Attitudes Towards Robots Through Experience

Building the next generation of robots for successful integration into our homes, offices, and factories is more than just solving the hardware and software problems – we also need to understand how they will be perceived and how they can work effectively with people in those spaces. aspect_ratio In summer 2025, RAI Institute set up a free popup robot experience in the CambridgeSide mall, designed to let people experience state-of-the-art robotics first hand. While news stories about robots and AI are common, with some being overly critical and some overly optimistic, most people have not encountered robots in the flesh (or metal) as it were. With no direct experience, their opinions are largely shaped by pop culture and social media, both of which are more focused on sensational stories i

IEEE Robotics

9mabout 2 hours ago

ModelsFresh

Higher energy costs from Iran war could threaten fragile economics of AI boom | Heather Stewart

Industry with business model not yet firmly established and investments financed by huge debts is particularly at risk Donald Trump’s most immediate concern in demanding Iran reopen the strait of Hormuz may be rocketing US gasoline prices, but if the conflict drags on, higher energy costs will be felt far beyond the pumps. Systemically higher power prices and fractured supply chains will squeeze industries and consumers worldwide. For the US, one consequence may be to threaten the fragile economics of the AI boom. Continue reading...

The Guardian AI

1mabout 4 hours ago

Knowledge Map

TopicsEntitiesSource

Connected Articles — Knowledge Graph

This article is connected to other articles through shared AI topics and tags.

Knowledge Graph100 articles · 165 connections

Scroll to zoom · drag to pan · click to open

Discussion

No comments yet — be the first to share your thoughts!

Gemma 4 Architecture Comparison

Daily AI Digest

More about

DeepSeek V4 draait op Huawei-chips en omzeilt afhankelijkheid Nvidia

Studying Human Attitudes Towards Robots Through Experience

Higher energy costs from Iran war could threaten fragile economics of AI boom | Heather Stewart

Knowledge Map

Connected Articles — Knowledge Graph

Discussion

More in Models

DeepSeek V4 draait op Huawei-chips en omzeilt afhankelijkheid Nvidia

Exclusive | Pentagon Used Anthropic’s Claude in Maduro Venezuela Raid - WSJ

Show HN: AI-agent instructions to restore Claude Pro/Max in OpenCode

Higher energy costs from Iran war could threaten fragile economics of AI boom | Heather Stewart