Sequence modeling · Efficient attention

RAM-NetLinear-Time Sequence Modeling
with Sparsely Addressable State

A larger fixed-size state, accessed a few slots at a time.

arXiv GitHub Coming soon Hugging Face Coming soon
A large memory, a small access footprintA current token writes to four random slots, then reads them back. Eight groups of sixteen slots form the memory array. This is a conceptual illustration. SPARSELY ADDRESSABLE STATE Many slots. A few selected. current token
The idea, in brief

Mainstream linear attention adds every token into one shared state, where their contributions interfere. Instead, RAM-Net splits the state into many slots. Each token's key selects a few slots to write, and its query selects a few slots to read.

  • Less interference. Each read or write accesses only the selected slots, not the entire state, so tokens interfere only when they select the same slots.
  • Scalable state. More slots add capacity, while the cost of each read or write stays constant.
01Motivation

Larger State, Sparser Access

The cost of history

A sequence model stores information from past tokens and retrieves the relevant part at every step. Its efficiency therefore depends on two costs, the amount of history it stores and the amount it accesses per step.

Full attention keeps one key–value entry per token and compares each query against all of them, so both costs grow with sequence length. To mitigate this, KV compression and sparse attention reduce per-step access, but the cache typically keeps growing.

Instead, mainstream linear attention compresses history into a fixed-size state. However, every step reads and writes the entire state. Consequently, all tokens interfere within the same state, and any increase in state size directly increases per-step access.

To address both issues, RAM-Net divides this fixed-size state into many addressable slots, and each step reads and writes only a few selected ones. Therefore, tokens interfere only when they select the same slots, and the state can grow without increasing per-step access.

MethodHow history is storedWhat each step accesses
Full attentionOne key–value entry per tokenAll entries
KV quantization / MLAOne compressed entry per tokenAll compressed entries
Sparse attentionOne entry per token in most methodsSelected entries
Linear attentionA single fixed-size stateThe entire state in most designs
RAM-NetA fixed-size array of slotsA few selected slots
02Architecture

Reading and Writing a Few Slots

Inside RAM-Net

Each token produces a key to select write slots, a query to select read slots, and a value to store.

RAM-Net processes them in three stages. First, the Address Decoder maps the key and the query to sparse write and read addresses, which select a few slots and weight their contributions. Next, Cyclic Address Positional Embedding (CAPE) can shift these addresses according to token position, making them position-aware. Finally, Gated Sparse Update (GSU) writes the value into the selected write slots and leaves all other slots unchanged. The selected read slots are then combined by their weights, normalized, gated, and projected to produce the output.

Since each step accesses only the selected slots, the size of the state and the per-step access are set separately. The slot count M sets the size of the state, while K sets how many slots each read or write accesses. Neither grows with sequence length, so storage and per-token cost stay constant as the context grows.

03Addressing

Efficient Sparse Addressing

Product Softmax

Select a view or adjust Top-K; drag factor weights to reshape the addresses.

Selecting K slots naively would score all M of them for every token, reintroducing a cost that grows with the state size. Instead, Product Softmax avoids this by composing the large address distribution from a few small ones.

Specifically, Product Softmax factors the slot count as M = dpU, so that each slot index can be written as U digits with dp values each. The key or query is split into U parts, and a softmax over each part gives a distribution over the dp values of one digit. The weight of a slot is then the product of the probabilities assigned to its digits. As a result, U dp scores suffice to address all M slots. For example, RAM-Net uses M = 1024 = 45 with only 20 scores.

The figure above shows how these digit distributions combine into a distribution over slots, from five equivalent perspectives.

Soft Radix Address

In base dp, each slot index is a sequence of U digits. Instead of fixing every digit to select a single slot, Product Softmax gives each digit a distribution over its dp values. Every slot is then weighted by the product of its digit probabilities.

Multilevel Decision Tree

Each digit forms one level of a tree, so every root-to-leaf path corresponds to one slot. A slot's weight is the product of the branch probabilities along its path. Since all nodes at a level share the same digit distribution, the most probable slots can be found level by level, at a cost that grows with the depth U rather than the number of leaves M.

Joint Distribution

Each digit corresponds to one axis of a U-dimensional grid, so each slot is a grid point weighted by the product of its probabilities along every axis. A confident digit confines the weight to a slice of the grid, whereas an uncertain digit spreads it along its axis. Thus, the highest-weight slots lie where these slices intersect, extending only along the axes of uncertain digits. The figure illustrates the case U = dp = 4.

Regroupable Address

Since the weight is a product, any set of digits can be merged into a single coarser digit without changing any slot weight. Splitting the digits into two groups thus arranges the slots as a matrix whose weights form an outer product of two vectors. The figure shows two such groupings, 64 × 4 and 4 × 64, which trace different patterns over the same weights.

Waveform Modulation

Read along the slot index, each digit's distribution becomes a periodic step waveform. Low-order digits change at every slot and repeat quickly, whereas high-order digits change slowly and hold each value over long intervals. The address distribution is the product of these waveforms, with the slow ones setting a coarse envelope and the fast ones modulating it within.

All five views express one property, namely that a slot's weight factorizes into independent per-digit probabilities. This factorization enables an exact and efficient Top-K search. Following the tree, RAM-Net fixes the digits one level at a time and keeps only the K most probable partial addresses at each level. A dropped partial address already trails K others, and since all of them are multiplied by the same remaining digit probabilities, it can never overtake them. The search thus returns the exact Top-K while evaluating only K dp candidates per level, rather than all M slots.

04CAPE

Position-Aware Addressing

Relative position

Toggle CAPE in the diagram; drag read or write weights to compare address overlap.

Product Softmax addresses slots by content alone, so identical content maps to the same slots at every position. Such addresses cannot tell where a token occurs relative to the current one.

Cyclic Address Positional Embedding (CAPE) adds this information by shifting each token's read and write addresses along the slot array according to its position. A read at step t then overlaps a write from step s according to their content and the distance t − s, rather than their absolute positions.

As a result, the same content written at different positions lands on different slots instead of overwriting the same ones. Since the slot array is finite, these shifts wrap around its end, which makes the embedding cyclic.

The shift only permutes slot indices, so each read and write still accesses K slots. RAM-Net applies CAPE to some heads and leaves the others content-only, without any positional shift. The CAPE heads capture relative-position patterns, whereas the content-only heads capture associations that hold regardless of distance.

05GSU

Balancing Retention and New Information

Updating slots

Select a stage, adjust γ, or drag a value to follow the slot update.

Each write must balance a selected slot's history against the incoming value, while leaving all other slots unchanged. Unlike the forgetting gates of mainstream linear attention, which decay the entire state at every step, Gated Sparse Update (GSU) forgets and writes only at the selected slots.

For each slot i, GSU keeps its content St,i together with a write mass mt,i, which measures how much write weight the slot has accumulated. A new write changes the content in proportion to its weight relative to this mass, so a heavily written slot changes slowly. The input-dependent scalar γt, computed from the current token by a learned projection, controls how much of the mass is retained, and thus how readily the slot accepts new values.

Given the write weight wt,i and the incoming value vt, the update proceeds in four steps.

αt,i=σ(−logit(wt,i)−γt) mt,i=αt,imt−1,i+wt,i λt,i=wt,imt,i+ε St,i=St−1,i+λt,i(vt−St−1,i)
  1. Retention (α). A stronger write or a larger γt keeps less of the previous mass.
  2. Mass (m). The new mass adds the current write weight to the retained mass.
  3. Update rate (λ). The rate is the write weight relative to the new mass, with a small ε for numerical stability.
  4. Content (S). The content moves toward vt by a fraction λ of the remaining gap, so a small λ preserves history and a λ near 1 lets the new value dominate.

For an unselected slot, wt,i = 0 gives αt,i = 1 and λt,i = 0, so both its content and mass stay unchanged. For a selected slot, γt ranges the update between two extremes, averaging all past writes as γt → −∞ and overwriting with the newest value as γt → +∞.

06Execution

Parallel Work Across Memory Slots

CUDA · Sparse access on the GPU

Sparse access reduces memory traffic. To turn that saving into efficient GPU execution, RAM-Net organizes work around independent memory slots.

During training and prefill, the model computes Top-K addresses across tokens in parallel, then groups read and write events by slot. Each slot receives its own sequence of access events.

Different slots and their channels can run in parallel, while events within each slot preserve the causal order required by the recurrence. Slots with no events require no work.

During autoregressive decoding, the model directly accesses the slots selected by the current token. It does not scan the full recurrent state.

07Results

Quality and State Access

Evidence & trade-offs

0.4MActive state · 340M Top-826.2M total state elements
72.9S-NIAH average · 340M Top-32Three tasks · 1K–16K contexts
15.90Wiki. perplexity · 1.3B Top-16Lower is better

These tables compare language quality, retrieval, task scores, and active state across model sizes and Top-K settings.

Language Modeling and Reasoning

Language Modeling and Reasoning — 340M and 1.3B models

At 340M, Top-32 achieves the lowest Wiki. perplexity in this comparison at 25.87, while Top-8 leads on Wino at 54.1. At 1.3B, Top-16 lowers perplexity to 15.90 and leads on LMB, ARC-c, COPA, Hella, and Wino. Its normalized ARC-e and ARC-c scores are 64.0 and 39.4. Gated DeltaNet remains ahead on MMLU and SciQ.

↓ Lower is better. ↑ Higher is better. State sizes are in millions of elements.

Bold marks the best value; underlining marks the second-best distinct value within each model size, including ties.

S-NIAH: Retrieval Across Context Lengths

S-NIAH: Retrieval Across Context Lengths — 340M and 1.3B models

Evaluation spans 1K–16K tokens, up to 4× the training context. At 340M, increasing Top-K from 8 to 32 raises the S-NIAH average from 64.7 to 72.9. Top-16 and Top-32 score 100.0 at every tested S-NIAH-1 length. At 1.3B, Top-16 averages 70.1, compared with 65.0 for Top-8. S-NIAH-2 and S-NIAH-3 remain harder, especially at longer contexts.

All scores use ↑ higher is better. Columns within each task give the context length.

Bold marks the best value; underlining marks the second-best distinct value within each model size, including ties.

Performance Across Tasks

Performance Across Tasks — 340M and 1.3B models

At 340M, RAM-Net's average rises from 23.2 with Top-8 to 27.9 with Top-16 and 32.0 with Top-32. Top-32 leads the recurrent baselines on average, while Transformer++ remains ahead at 35.1. At 1.3B, Top-8 averages 40.6 and Top-16 averages 39.0, compared with 39.5 for Gated DeltaNet and 49.8 for Transformer++.

↑ Higher is better.

Bold marks the best value; underlining marks the second-best distinct value within each model size, including ties.

At 340M, Top-8 accesses 0.4M of 26.2M state elements—about 1.5% of its total state and roughly 1/32 of Mamba2’s 12.9M active state. At 1.3B, Top-8 accesses 0.8M of 52.4M elements; the same Top-K accounting gives 1.6M active elements for Top-16. These figures measure state access; decoding throughput is evaluated separately.

What changes with Top-K?

At 340M, moving from Top-8 to Top-32 increases active state from 0.4M to 1.6M elements while total state stays at 26.2M. S-NIAH average improves from 64.7 to 72.9, and the additional-task average from 23.2 to 32.0. The trade-off is better results at a higher per-token access cost.

Beyond these tables

The paper also evaluates fine-grained associative recall on MQAR and longer-context retrieval in a separate S-NIAH-1 setting. With repeated filler, the 340M Top-8 model reaches 99.4% at 128K after training at 4K. That setting is distinct from the 1K–16K comparison above.

Separate 340M-scale decoding benchmarks on a single NVIDIA RTX 5090 report throughput comparable to GLA and higher than the tested Mamba2, DeltaNet, and Gated DeltaNet implementations. Runtime depends on the implementation and hardware as well as state access.

Retrieval still depends on reliable address selection: an unselected slot cannot contribute to the current read. More slots increase capacity; a larger Top-K increases access. Their benefits depend on the task.
08Retrieval Traces

Following a Needle Through Memory

S-NIAH-1 & S-NIAH-3

Where does a fact go, and what remains when the model needs it? These S-NIAH-1 and S-NIAH-3 traces follow needle information from writing to answer-time reading across 1,024 slots. Switch the head or example, or jump to Needle and Read to compare the two moments.

Colors track source contributions under the recorded GSU gains. They illustrate provenance, not the actual hidden vectors. Each head uses a fixed, reordered slot layout to expose spatial patterns. These selected examples help inspect retrieval behavior; they do not estimate its overall success rate.

09Address Patterns

What Do Shared Addresses Reveal?

Word pieces and paired punctuation

Selected heads show shared write–read addresses between adjacent word pieces, or between opening and closing punctuation. The history matrix exposes these connections token by token; the slot view follows how information from the source token is retained and read.

These are exploratory patterns in selected heads and short texts. Address overlap offers a useful window into memory use, but does not by itself establish semantic matching or a causal role in the model’s output.

RAM-Net

More room to remember. Less to access.