An X post from researcher Jun Song went viral overnight NZ time: 36,000 views and 870-odd likes in its first eight hours, on a claim that Samsung had “just dropped” a paper running a 13-billion-parameter AI model in under 1GB. “If this hits production, local AI changes forever,” it added, offering the perspective that this is “like running Kimi-K3 on a single DGX Spark.” The paper is real, and worth your attention. The tweet’s numbers are not what the paper says.
🔍 THE BOTTOM LINE
Samsung Research’s NanoQuant is a peer-reviewed, ICML 2026-accepted quantisation method that compresses a 70B model from 138GB to 5.35GB — small enough to run on a consumer 8GB graphics card at 20 tokens a second. That is the genuine headline. The viral version inflated it: the paper is eight months old, its own runtime table puts a 13B model at 1.70GB, and the compression’s costs — measurable quality loss and NVIDIA-only kernels — went unmentioned.
What the paper actually demonstrates
NanoQuant is a post-training quantisation method from Samsung Research that pushes model weights below one bit apiece — a frontier where, as the authors note, previous post-training methods simply failed. The verified results from the paper and its open-source release:
- Llama2-70B compressed 25.8×, from 138.04GB to 5.35GB, in 13 hours on a single H100
- That 70B model then runs on a consumer 8GB GPU (an RTX 3050) at up to 20.11 tokens per second
- At the most aggressive 0.55-bit rate, a 7B model peaks at 1.07GB including runtime memory — genuinely under 1GB
- Custom binary CUDA kernels deliver the throughput, memory and energy gains
The same wave as the ternary models we covered running 27B-parameter LLMs in under 6GB, but taken further — past 1-bit quantisation, which is a genuine technical milestone. “The only PTQ framework effectively enabling sub-1-bit compression” while staying quality-competitive, at a new Pareto frontier, is a fair self-description given the benchmark tables.
The arithmetic the tweet skips
The tweet’s core claim — a 13B model “in under 1GB” — is contradicted by the paper’s own runtime table: at 0.55 bits, Llama2-13B peaks at 1.70GB; it is the 7B model that dips under 1GB in practice (1.07GB). Sub-1GB at 13B scale exists only as weight-storage arithmetic — 0.55 bits per parameter times 13 billion, before the scales and low-rank buffers the method adds, and before anything runs.
Two more catches: quality falls with the bit-rate, exactly as expected — the paper’s 1-bit Llama2 result scores just below the best competing method, and sub-1-bit scores fall further. “Staying competitive” is the honest claim, not “matching.” And the kernels are CUDA: this is NVIDIA-first work. Nothing here runs on a phone or an Apple Silicon Mac without ports that don’t yet exist. The open-weight economics shift continues, but on graphics-card hardware.
Old news that arrived eight months late
The paper was submitted on 6 February 2026 and revised in June; the ICML acceptance is part of the record. “Just dropped” is doing heavy lifting — the tweet is a February result dressed in breaking-news clothes, recirculated into a feed primed for local-AI news by the week’s open-weight containment stories and inference price falls. Nothing dishonest about finding a good paper late; the number inflation is the part that doesn’t survive contact.
What it means for New Zealand
The verified version of this story is the relevant one here: sub-bit compression is the difference between “AI needs a data centre” and “AI runs on the machine under someone’s desk.” NZ’s AI adoption problem has always been geography-plus-scale — few big compute budgets, plenty of small businesses that would use a capable local model if it fit. The appliance lane this site’s coverage has tracked — what now runs on consumer GPUs — moves another notch when 70B-class models fit on entry-level cards. Watch for independent replications of the sub-1-bit quality numbers before treating “70B on anything” as settled; the method is open-source, so they will come either quickly or not at all.
❓ FAQ
What is NanoQuant? A Samsung Research post-training quantisation method (arXiv 2602.06694, accepted to ICML 2026) that compresses large language models to binary and sub-1-bit weight precision, with open-source code.
Is the “13B model in under 1GB” claim true? Not as stated. The paper’s runtime table shows a 13B model peaking at 1.70GB at its most aggressive rate; the 7B model is the one that dips near 1GB with runtime included (1.07GB).
Can I run this on a Mac or phone? Not yet — the demonstrated speedups rely on custom CUDA kernels, so it is NVIDIA-GPU work. Ports to other platforms would be separate projects.
Does sub-1-bit compression hurt output quality? Measurably, yes — benchmark scores fall as bits fall; the paper’s claim is competitive quality at a new memory frontier, not parity with full precision.
Why is this tweet going viral if the paper is from February? Viral mechanics, not news: an engaging summary with no links, surfacing into a news week already focused on cheap local inference. The underlying paper deserved the attention eight months ago; the numbers deserved checking before the amplification.
🔍 THE BOTTOM LINE
Samsung’s NanoQuant is a real, peer-reviewed advance that justifies most of the excitement: a 70B model on an 8GB consumer GPU is a genuine threshold, and the code is public. The viral tweet swapped the paper’s demonstrated result for a bigger one the paper doesn’t show, aged an eight-month-old finding into a breaking story, and cut the trade-offs. The correction matters because the honest version — 7B near 1GB in real use, 70B on a desk card, quality visibly traded — is still a big deal.
📰 Sources
- arXiv — NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
- Samsung Research — NanoQuant blog post
- Hugging Face paper page
- SamsungLabs GitHub — NanoQuant
- X — @jun_song post, 8 October 2026