[P] GPU friendly lossless 12-bit BF16 format with 0.03% escape rate and 1 integer ADD decode works for AMD & NVIDIA
Hi everyone, I am from Australia : ) I just released a new research prototype It’s a lossless BF16 compression format that stores weights in 12 bits by replacing the 8-bit exponent with a 4-bit group code . For 99.97% of weights , decoding is just one integer ADD . Byte-aligned split storage: true 12-bit per weight, no 16-bit padding waste, and zero HBM read amplification. Yes 12 bit not 11 bit !! The main idea was not just “compress weights more”, but to make the format GPU-friendly enough to use directly during inference : sign + mantissa: exactly 1 byte per element group: two nibbles packed into exactly 1 byte too https://preview.redd.it/qbx94xeeo2tg1.png?width=1536 format=png auto=webp s=831da49f6b1729bd0a0e2d1f075786274e5a7398 1.33x smaller than BF16 Fixed-rate 12-bit per weight , no
Could not retrieve the full article text.
Read on Reddit r/MachineLearning →Reddit r/MachineLearning
https://www.reddit.com/r/MachineLearning/comments/1sbv9jl/p_gpu_friendly_lossless_12bit_bf16_format_with/Sign in to highlight and annotate this article

Conversation starters
Daily AI Digest
Get the top 5 AI stories delivered to your inbox every morning.
More about
llamamistralmodel
Study maps developer frustration over "AI slop" as a "tragedy of the commons" in software development
A qualitative study looks at how developers perceive and push back against low-quality AI content, or "slop," in software development. The critics describe a "tragedy of the commons" where individual productivity gains come at the cost of reviewers and the open-source community. The article Study maps developer frustration over "AI slop" as a "tragedy of the commons" in software development appeared first on The Decoder .

DeepSeek V4 draait op Huawei-chips en omzeilt afhankelijkheid Nvidia
Het grote taalmodel DeepSeek V4 draait naar verluidt op chips van Huawei. Tot dusver was het Chinese bedrijf achter de AI-dienst afhankelijk van Nvidia-processors, die onder Amerikaanse exportrestricties vallen. Naar verwachting komt het model deze lente uit.

Studying Human Attitudes Towards Robots Through Experience
Building the next generation of robots for successful integration into our homes, offices, and factories is more than just solving the hardware and software problems – we also need to understand how they will be perceived and how they can work effectively with people in those spaces. aspect_ratio In summer 2025, RAI Institute set up a free popup robot experience in the CambridgeSide mall, designed to let people experience state-of-the-art robotics first hand. While news stories about robots and AI are common, with some being overly critical and some overly optimistic, most people have not encountered robots in the flesh (or metal) as it were. With no direct experience, their opinions are largely shaped by pop culture and social media, both of which are more focused on sensational stories i
Knowledge Map
Connected Articles — Knowledge Graph
This article is connected to other articles through shared AI topics and tags.
More in Models

DeepSeek V4 draait op Huawei-chips en omzeilt afhankelijkheid Nvidia
Het grote taalmodel DeepSeek V4 draait naar verluidt op chips van Huawei. Tot dusver was het Chinese bedrijf achter de AI-dienst afhankelijk van Nvidia-processors, die onder Amerikaanse exportrestricties vallen. Naar verwachting komt het model deze lente uit.

Higher energy costs from Iran war could threaten fragile economics of AI boom | Heather Stewart
Industry with business model not yet firmly established and investments financed by huge debts is particularly at risk Donald Trump’s most immediate concern in demanding Iran reopen the strait of Hormuz may be rocketing US gasoline prices, but if the conflict drags on, higher energy costs will be felt far beyond the pumps. Systemically higher power prices and fractured supply chains will squeeze industries and consumers worldwide. For the US, one consequence may be to threaten the fragile economics of the AI boom. Continue reading...

![[P] GPU friendly lossless 12-bit BF16 format with 0.03% escape rate and 1 integer ADD decode works for AMD & NVIDIA](https://preview.redd.it/qbx94xeeo2tg1.png?width=140&height=93&auto=webp&s=39ed7f02dad84ccf081f932903c016c7983d4fcd)


Discussion
Sign in to join the discussion
No comments yet — be the first to share your thoughts!