Show HN: Made NZ's member of parliament financial disclosure data searchable (open-register-of-pecuniary-interests.joshmcarthur.com)

[3.3] For saving the KV cache, only the intermediate latent representations need to be stored: [latex] where r is much smaller than nh · dh [n-sub-h, d-sub-h]

[background] In traditional multi-head attention you must cache full key and value matrices of size T x (nh · dh) where T is the token length, nh is the number of attention heads, dh is the dimensionality of each individual head

sounds like a big win for memory constrained environments like local inference

killerstorm · 127d ago

Another paper related to attention distillation, although doing something far more radical: transformer attention is distilled onto RWKV-like model: https://huggingface.co/papers/2505.03005

karmakaze · 127d ago

I'm not "in the field" though I like to read about and use LLMs. This video "How DeepSeek Rewrote the Transformer [MLA]"[0] is really good at explaining MHA, MQA, GQA, and MLA with clear visuals/animations and how DeepSeek MLA is 57x more efficient.

[0] https://www.youtube.com/watch?v=0VLAoVGf_74&t=960s

olq_plo · 128d ago

Very cool idea. Can't wait for converted models on HF.

MichaelMoser123 · 127d ago

deepseek-v2,v3,r1 are all using multi-headed attention.

magicalhippo · 127d ago

I'm just following the field from the sidelines, but this looks interesting to me. Especially the increase in expressiveness that the new model allows for over GQA, at the cost of just ~10% more memory, and the fact that you can convert existing GQA models like LLaMA, Qwen etc with just a bit of fine-tuning.

Perhaps a trivial insight but I feel a lot of progress often comes in the form of generalizations, where existing approaches can be seen as special cases. Here the authors show that Group Query Attention (GQA) and Multi-Query Attention (MQA) falls out as special cases of their new model.

edit:

Adding my own summary, as I understand it.

The key to what they're doing, no pun intended, is to rely on the fact that large, high-dimensional, matrices may contain a lot of redundant information. Thus one may be able to find an good approximation which has less redundant information, by going through an intermediary stage which has fewer dimensions.

A n-by-m matrix M takes n-dimensional vectors and transforms them to m-dimensional vectors. The trick here is to replace matrix A by two matrices, L and R, which are n-by-r and r-by-m respectively, where r is smaller than n and m. This is called a low-rank approximation.

In a sense you're "straining the matrix", by forcing the information to pass through an intermediary, low-dimensional vector.

The memory savings come from the fact that matrix A has n*m entries, while L and R have n*r and r*m entries respectively. Say n = m = 100 and r = 20, that means A has 100*100 = 10k entries, while L and R have just 100*20 + 20*100 = 4k entries in total.

The trick itself is not new, for example it is also used in LoRA where an additional low-rank approximation matrix is used to tweak the output of an existing model. The low rank means there's far fewer the matrix entries, aka parameters, to train than if one had used a regular fully dense matrix.

The extra expressiveness of MLA comes from the fact that in GQA, in order to save memory, some of the matrices are actually built by gluing copies of a narrower matrix together. This means the information in the glued-up matrices are very redundant and fixed in a certain way, and thus are restricted in how they can transform the inputs.

By using the low-rank approximation instead, the information in the full, reconstructed matrices are not fixed in the same way compared to the glued-up result. Thus the inputs can be transformed in a less restrictive way, leading to the increase in expressiveness.

The GQA method saves a bit more memory compared to MLA as the narrower matrices are even smaller than the low-rank matrices in MLA, but at the cost of expressiveness.

wiz21c · 127d ago

Not quite related, but do the mamba models gain ground ?

Answering my own question: https://www.reddit.com/r/MachineLearning/comments/1hpg91o/d_...

kavalg · 127d ago

My (possibly wrong) TLDR: TransMLA is a method to "compress" an already trained GQA model, with the additional option to further fine tune it. Shall make inference faster.

yorwba · 127d ago

It is not a method to compress a Grouped-Query Attention model, but to expand it into an equivalent Multi-head Latent Attention model with the same key-value cache size but larger effective key/value vectors and a correspondingly larger number of trainable parameters. With additional training, you can then obtain a better model that only uses a little bit more memory.

kavalg · 126d ago

Thanks for the clarification.

freeqaz · 127d ago

Also makes models smarter ("expressive")

EGreg · 127d ago

All you need to stop posting titles like that !

AI Companion Futures (osmarks.net)

Secrets of DeepSeek AI model revealed in landmark paper (nature.com)

Google and Microsoft Back NASA Scientist's Food Crisis Hotline (bloomberg.com)

'Dystopian' toilets won't give you loo roll unless you watch an advert first (metro.co.uk)

Hong Kong's Dim Sum Cart 'Aunties' Make Their Final Rounds (nytimes.com)

Show HN: I twisted my foot, so used an old phone's sensors into a recovery alarm

I vibe coded into the top 100 of the berghain puzzle (twitter.com)

Best Calling Apps for PC, Android, and iOS (linkedphone.com)

Doodles.google/Robots.txt (doodles.google)

European ant is the first known animal to clone members of another species (livescience.com)

It's Not Just You: Music Streaming Is Broken Now (youtube.com)

US to invest £150B in UK, promising jobs (bbc.com)

Genomic Integration and Molecular Dysregulation in Cancer Following mRNA Vaxx (zenodo.org)

Evals in 2025: benchmarks to build models people can use (github.com)

Seed Diffusion Preview (seed.bytedance.com)

NASA+ Is Coming to Netflix This Summer (nasa.gov)

BART Cab Cam: Yellow Line from SFO Airport to Antioch (youtube.com)

Ask HN: What makes a great engineering manager?

Mercedes to bring back cabin buttons for current and future models (autocar.co.uk)

Saving Lives by Ignoring the Models (statesmanjournal.com)

Who ~Framed~ Created Pitman? (blog.hardcoregaming101.net)

INapGPU: Text-mode graphics card, using only TTL gates (github.com)

Notes on writing a monovocalic sonnet (muppetlabs.com)

We have outgrown the Process model (sidhion.com)

Three Perspectives on Equivalence Relations (pseudonium.github.io)

Monoids in Public (blog.veritates.love)

The Newton-to-Quantum Mechanics tension between determinism and indeterminism (mathpages.com)

OpenAI says models are programmed to make stuff up (theregister.com)

The Seneca: First Edition $8k PC Keyboard (norbauer.co)

AI bots 'coached' son to suicide: grieving mom delivers testimony [video] (youtube.com)

Why, as a responsible adult, SimCity 2000 hits differently (arstechnica.com)

How to Build Secure AI Coding Agents with Cerebras and Docker Compose (docker.com)

Kubernetes v1.34: Pods Report DRA Resource Health (kubernetes.io)

Show HN: Made NZ's member of parliament financial disclosure data searchable (open-register-of-pecuniary-interests.joshmcarthur.com)

Could a primordial black hole's last burst explain a mysteriously neutrino? (news.mit.edu)

Why Is Xcode So Antagonistic to Reduced Vision? (thecodist.com)

China bans its biggest tech companies from acquiring Nvidia chips, says report (tomshardware.com)

How HR took over British business and got in the way of actual work (thetimes.com)

LLVM: Our AI policy vs. code of conduct and vs. reality (discourse.llvm.org)

Biotech firm announces 'pivotal step' in effort to bring back the dodo (cnn.com)

ABC Pulls Jimmy Kimmel Off Air for Charlie Kirk Comments After FCC Pressure (nytimes.com)

Security Industry turns "Without Warranty" to "Supply Chain Attack, Shame on You (shrimp.starlightnet.work)

Things for OS 26 (culturedcode.com)

WIP: A website for listing / finding free plants in your community (leafrens.com)

China blocks sale of Nvidia AI chips (arstechnica.com)

The "Rainey Street Ripper": An Independent Analysis of the Evidence (digital.library.txst.edu)

Meta Ray-Ban Display: A Breakthrough Category of AI Glasses Quest Blog Store (meta.com)

DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning (nature.com)

Java 25, Ready to Perform to the Limit (hanno.codes)

The Economic Impacts of AI: A Multidisciplinary, Multibook Review [pdf] (kevinbryanecon.com)

TransMLA: Multi-head latent attention is all you need

Comments (32)