19 분 소요

0. Introduction

Paper link

Code link

HOLA-340M model

HOLA+recency-340M model

한 줄 요약: HOLA는 Gated DeltaNet의 고정 크기 recurrent state를 압축 기억으로 유지하면서, state가 제대로 표현하지 못한 token을 delta-rule write magnitude인 $\beta_t|e_t|$로 골라 작은 exact KV cache에 저장하고, 별도의 sharpened cache read로 다시 꺼내는 semiparametric linear attention이다.

이 논문을 지금 볼 가치가 있는 이유는 다음과 같음.

  • Linear attention과 state-space model의 장점인 context-length-independent recurrent memory를 유지하면서 exact recall 부족을 직접 겨냥한다.
  • 단순 sliding window가 아니라 model의 state update 자체를 memory admission signal로 사용한다.
  • 무엇을 저장할지뿐 아니라 exact KV를 어떻게 읽어야 하는지까지 함께 해결한다.
  • 340M same-backbone comparison에서 perplexity, in-context extraction, long-context needle recall을 구분해 method의 작동 영역을 보여준다.
  • Cache를 external retrieval system이 아니라 각 recurrent layer 내부의 bounded test-time memory로 넣는 설계가 명확하다.

Linear attention은 softmax attention의 growing KV cache를 고정 크기 state로 바꾼다. Prefix가 아무리 길어져도 recurrent state 크기는 일정하므로 decoding memory를 context length에 대해 $O(1)$로 유지할 수 있다.

하지만 이 효율성은 prefix를 압축한다는 뜻이기도 하다. State가 모든 과거 key-value association을 정확히 보존하는 것은 아니다. Distinct association이 많아지면 shared key direction에서 interference가 생기고, 오래된 value가 새 write에 의해 덮일 수 있다.

이 문제는 일반적인 language modeling보다 exact retrieval에서 더 선명하게 드러난다.

  • 여러 key-value pair를 동시에 기억하는 associative recall
  • 긴 context 안의 passkey나 needle 찾기
  • 멀리 떨어진 token을 그대로 copy하기
  • Document에서 특정 field나 span을 exact extraction하기

Full softmax attention은 prefix의 모든 KV를 남기기 때문에 이런 문제에서 강하다. 대신 cache memory는 context length에 따라 $O(T)$로 증가하고, training attention compute는 보통 $O(T^2)$다.

HOLA는 두 방식을 하나로 완전히 대체하려 하지 않는다. 대신 서로 다른 역할을 맡긴다.

  1. Recurrent state는 반복적이고 linearly compressible한 구조를 요약한다.
  2. 작은 exact KV cache는 state에 억지로 압축하면 잃기 쉬운 association을 보존한다.

논문은 이를 Complementary Learning Systems 관점으로 설명한다. Neocortex처럼 느리고 압축적인 memory와 hippocampus처럼 빠르고 구체적인 episodic memory를 분리하자는 비유다.

중요한 점은 cache를 추가했다는 사실만이 아니다. Bounded cache는 모든 token을 담을 수 없으므로 두 질문에 답해야 한다.

어떤 token을 exact memory에 남길 것인가?

저장한 exact KV를 soft average가 아니라 실제 retrieval로 어떻게 읽을 것인가?

HOLA의 기여는 이 두 질문을 Gated DeltaNet 내부 signal과 read normalization으로 연결한 데 있다.

1. Problem Setting

1-1. Problem definition

DeltaNet 계열은 history를 matrix state $S_t$로 압축한다. Token $t$의 query, key, value를 $q_t$, $k_t$, $v_t$라고 하고, $q_t$와 $k_t$는 unit L2 norm을 가진다고 하자.

Basic delta-rule update는 다음처럼 쓸 수 있다.

\[S_t = S_{t-1} + \beta_t k_t e_t^\top\] \[e_t = v_t - k_t^\top S_{t-1}\] \[o_t^{\mathrm{state}} = q_t^\top S_t\]

여기서 $e_t$는 current state가 key $k_t$에 대해 value $v_t$를 얼마나 잘 예측하지 못했는지를 나타내는 residual이다. $\beta_t$는 write strength다.

직관은 다음과 같다.

  1. State가 이미 $v_t$를 잘 예측한다면 $e_t$는 작다.
  2. State가 틀리면 $e_t$는 커진다.
  3. Model은 $\beta_t$만큼 residual을 state에 쓴다.
  4. 같은 association이 반복될 때는 이미 state가 잘 예측하므로 불필요한 write가 줄어든다.

이 구조는 online associative memory로 효율적이다. 그러나 $S_t$의 크기와 rank는 고정되어 있다. Context 안의 distinct association 수가 state capacity를 넘으면 새 update가 과거 association에 interference를 만든다.

즉 recurrent state는 다음 두 능력을 동시에 완벽하게 제공하지 못한다.

  • Prefix 전체의 statistical structure를 compact하게 요약하는 능력
  • 특정 과거 token이나 key-value pair를 exact하게 다시 꺼내는 능력

HOLA가 정의하는 문제는 다음과 같다.

Question Meaning
무엇이 손실되는가 Fixed-size recurrent state에 압축된 distant exact association
무엇을 유지할 것인가 GDN의 constant-size state와 sub-quadratic sequence processing
무엇을 추가할 것인가 각 layer의 bounded exact KV cache
어떤 token을 저장할 것인가 State가 가장 크게 수정된 surprising token
어떻게 읽을 것인가 Near-uniform average가 아니라 sharp KV retrieval
무엇과 비교할 것인가 Same-backbone GDN, matched recency cache, full attention, 다른 linear model

1-2. Why previous approaches are insufficient

1) Pure recurrent state는 exact memory가 아니다

GDN의 state는 모든 prefix token을 하나의 fixed-size matrix로 누적한다. Context가 길어져도 memory가 늘지 않는다는 것은 반대로 말하면 token별 exact representation을 모두 유지하지 않는다는 뜻이다.

Language modeling perplexity가 괜찮더라도 다음 failure는 남을 수 있다.

  • 같은 형식의 여러 key가 들어왔을 때 value association이 섞인다.
  • 오래된 passkey가 newer write에 의해 약해진다.
  • Exact span extraction이 paraphrased semantic understanding보다 어렵다.
  • Context length가 training length를 넘어가면 state saturation이 누적된다.

2) Recent-token window는 distant importance를 보존하지 못한다

Efficient hybrid model은 recurrent backbone 옆에 local softmax window를 두는 경우가 많다. 최근 token은 exact하게 보고, 오래된 token은 state에 맡기는 방식이다.

Local dependency에는 합리적이지만 distance와 importance는 같지 않다. Context 초반의 identifier, 계약 조건, entity attribute, passkey가 뒤에서 다시 필요할 수 있다. Sliding window는 중요하더라도 오래되면 반드시 버린다.

따라서 bounded memory에서 필요한 것은 recency가 아니라 selective retention이다.

3) Learned eviction module은 추가 구조와 supervision을 요구한다

Cache admission이나 eviction을 별도 network로 학습할 수 있다. 하지만 다음 문제가 생긴다.

  • 추가 parameter와 compute가 필요하다.
  • Eviction module의 target이 명확하지 않을 수 있다.
  • Backbone state update와 별개로 memory importance를 다시 추정한다.
  • Training과 streaming inference에서 selection semantics가 어긋날 수 있다.

HOLA는 GDN이 이미 계산하는 update magnitude를 재사용한다. 별도의 learned eviction network 없이 state가 가장 어려워한 token을 찾는다.

4) Exact KV를 저장해도 flat read면 exact memory가 아니다

Naive cache는 unit-normalized $q$와 $k$를 그대로 softmax에 넣을 수 있다. 그러나 head dimension이 크고 logit range가 작으면 cache weight가 거의 uniform해진다.

논문의 setting에서 unit-L2 path는 logit scale이 대략 $0.83\cos$ 수준이다. Cache entry가 64개라면 perfect match조차 attention mass를 충분히 독점하지 못한다.

이 경우 exact KV를 저장했어도 output은 여러 value의 soft average가 된다. Storage는 exact하지만 read가 lossy한 모순이 생긴다.

5) Full attention은 exact하지만 memory budget이 다르다

Full softmax attention은 모든 prefix KV를 보존하므로 exact retrieval의 강한 ceiling이다. 하지만 HOLA가 겨냥하는 deployment regime과 resource class가 다르다.

따라서 이 논문의 핵심 비교는 다음 두 축을 분리해 봐야 한다.

  • Same-backbone GDN 대비 bounded cache가 실제로 개선하는가.
  • Bounded memory가 full attention의 exact recall gap을 얼마나 줄이는가.

2. Core Idea

2-1. Main contribution

1) Semiparametric test-time memory

논문은 attention memory를 test-time regression으로 본다. Prefix에서 얻은 KV observation을 이용해 query $q_t$에 맞는 value를 예측하는 문제다.

Pure GDN은 fixed-size state만 사용하는 parametric estimator다.

\[o_t = q_t^\top S_t\]

Full attention은 모든 prefix KV를 유지하는 unbounded non-parametric estimator다.

HOLA는 둘 사이의 bounded semiparametric estimator다.

\[o_t = q_t^\top S_t + \lambda_t g_t(q_t)\]

여기서 역할은 다음과 같다.

  • $q_t^\top S_t$: compressive state estimate
  • $g_t(q_t)$: bounded exact KV set에 대한 non-parametric cache read
  • $\lambda_t$: state와 cache output을 섞는 read-side coefficient

이 관점의 장점은 architecture를 단순 hybrid로 설명하는 데서 끝나지 않는다는 점이다. 어떤 memory가 어떤 function class를 맡는지 분명해진다.

2) State write magnitude를 surprise score로 사용

Delta update를 다음처럼 쓰자.

\[\Delta_t = \beta_t k_t e_t^\top\]

Token이 state에 미친 전체 변화의 크기는 Frobenius norm으로 볼 수 있다.

\[m_t = \|\Delta_t\|_F\]

$k_t$가 unit norm이므로 다음처럼 단순해진다.

\[m_t = \beta_t \|e_t\|\]

이 score는 두 정보를 함께 포함한다.

  • $|e_t|$: State가 해당 value를 얼마나 예측하지 못했는가.
  • $\beta_t$: Model이 그 residual을 state에 얼마나 강하게 쓰기로 했는가.

HOLA는 각 layer에서 지금까지 본 token 중 $m_t$가 큰 top-$w$ exact KV를 유지한다. Main setting은 $w=64$다.

이 선택은 parameter-free다. Token이 들어올 때 score가 정해지므로 online top-k maintenance와 blockwise selection에 같은 의미를 적용할 수 있다.

3) Recency가 아니라 representation failure를 저장 기준으로 사용

Recent window는 position-based memory다. HOLA는 state-failure-based memory다.

둘의 차이는 distant needle에서 분명하다.

  • Recency cache: Needle이 window 밖으로 나가면 exact copy를 잃는다.
  • HOLA cache: Needle이 state에 큰 update를 만들었다면 distance와 무관하게 남을 수 있다.

논문은 HOLA와 cache size, chunk, read path, gate를 같게 두고 eviction rule만 recency로 바꾼 matched control을 제공한다.

4) Cache path만 따로 sharpen한다

State update에는 unit-norm key가 필요하다. Delta rule의 stability 때문이다. 반면 exact cache read에는 larger logit scale이 필요하다.

HOLA는 같은 $q$, $k$ representation을 두 path에서 다르게 normalize한다.

  • State path: L2-normalized unit vector 유지
  • Cache path: Learnable RMSNorm-$\gamma$ 적용

Cache read는 다음처럼 계산된다.

\[o_t^{\mathrm{cache}} = \sum_{j \in \mathcal{V}_t} \mathrm{softmax}_j \left( \frac{\tilde{q}_t^\top \tilde{k}_j}{\sqrt{d}} \right) v_j\] \[\tilde{q} = \mathrm{RMSNorm}_{\gamma}(q)\] \[\tilde{k} = \mathrm{RMSNorm}_{\gamma}(k)\]

RMSNorm path는 norm을 대략 $\sqrt{d}$ scale로 유지해 matching key의 logit 차이를 키운다. State update path는 그대로 두므로 recurrent stability와 cache sharpness를 분리한다.

5) 각 recurrent layer에 작은 exact memory를 부착

HOLA는 몇 개 layer를 full attention으로 바꾸는 inter-layer hybrid가 아니다. 모든 GDN layer에 bounded exact cache를 붙인다.

이는 각 layer의 representation space에서 state가 놓친 association을 해당 layer의 exact KV로 보완한다는 뜻이다.

2-2. Design intuition

HOLA의 가장 좋은 직관은 “state가 스스로 무엇을 기억하지 못했는지 알려준다”는 것이다.

Delta rule은 new value를 그대로 쓰지 않는다. State prediction과의 차이인 residual을 계산하고, write gate를 곱해 state를 수정한다. 큰 $\beta_t|e_t|$는 다음을 뜻한다.

  1. 기존 state가 이 token의 value를 잘 표현하지 못했다.
  2. Model은 이 차이를 중요한 update로 판단했다.
  3. 이 token을 state에만 맡기면 later interference에 취약할 수 있다.
  4. 따라서 exact cache budget을 배정할 가치가 있다.

이 설계는 external salience classifier보다 backbone computation과 직접 맞물린다. Memory selection이 다음 token prediction을 위해 실제로 발생한 state write에 근거하기 때문이다.

Read 쪽 intuition도 중요하다. Exact memory는 단지 exact copy를 보관하는 storage가 아니다. Query가 matching entry를 선택적으로 꺼낼 수 있어야 한다.

따라서 HOLA는 bounded memory problem을 두 개의 independent decision으로 본다.

Decision Failure if wrong HOLA choice
What to write 중요한 distant token이 cache에 없음 Top-$w$ by $\beta|e|$
How to read Exact KV가 uniform average로 섞임 Decoupled RMSNorm-$\gamma$

논문의 ablation은 read sharpening이 단순 세부 구현이 아니라 가장 큰 lever 중 하나임을 보여준다.

3. Architecture / Method

3-1. Overview

Item Description
Goal GDN의 fixed recurrent memory를 유지하면서 exact recall 보강
Backbone Gated DeltaNet
Compressive memory Per-layer recurrent state $S_t$
Exact memory Per-layer bounded KV cache
Cache capacity Persistent top-$w$, main setting $w=64$
Admission score $m_t=\beta_t|e_t|$
Read normalization Cache-only RMSNorm-$\gamma$
Visible cache set Persistent cache + current causal chunk + null sink
Current chunk Main setting $C=256$
Output State read + gated cache read
Main property Context-length-independent bounded inference state

3-2. Module breakdown

1) Gated DeltaNet state

HOLA의 backbone은 GDN이다. GDN은 DeltaNet에 data-dependent decay gate $\alpha_t$를 추가한다.

개념적으로 state prediction과 residual은 다음처럼 바뀐다.

\[\hat{v}_t = \alpha_t k_t^\top S_{t-1}\] \[e_t = v_t - \hat{v}_t\]

$\alpha_t$는 selective forgetting을 제공하고, $\beta_t$는 write strength를 조절한다. HOLA는 이 backbone을 바꾸지 않고 exact cache를 추가한다.

2) Surprise score computation

각 token은 state update와 동시에 score를 얻는다.

\[s_t = \beta_t\|e_t\|\]

이 score는 attention probability나 downstream query frequency가 아니다. Current token이 recurrent state를 얼마나 크게 바꿨는지 나타낸다.

Cache는 history 전체에서 가장 큰 score의 KV pair를 유지한다.

3) Bounded top-$w$ cache

각 layer의 persistent cache는 다음 field를 가진다고 볼 수 있다.

  • Exact key
  • Exact value
  • Fixed surprise score
  • Cache slot metadata

새 token이 들어오면 다음과 같이 처리할 수 있다.

  1. State update에 필요한 residual과 $\beta_t$를 계산한다.
  2. Surprise score $s_t$를 계산한다.
  3. Cache가 비어 있으면 append한다.
  4. Full이면 current minimum score와 비교한다.
  5. 더 크면 minimum entry를 교체한다.

Score가 token write 시점에 고정되므로 history를 다시 평가할 필요가 없다.

4) Current chunk와 null sink

논문의 visible KV set $\mathcal{V}_t$는 persistent top-$w$ cache만 포함하지 않는다.

  • Persistent surprise-selected cache
  • Current processing chunk 안의 causally visible token
  • One null sink

Main configuration에서 $w=64$, $C=256$이므로 visible exact span은 대략 다음과 같다.

\[w + C + 1 = 321\]

Current chunk는 blockwise causal computation에서 local exact interaction을 제공한다. Persistent cache는 chunk boundary를 넘어 selected association을 유지한다. Null sink는 cache를 읽지 않는 선택지를 안정적으로 제공한다.

5) Sharpened cache attention

Unit-L2 state path의 $q$, $k$를 cache softmax에 그대로 넣으면 logit range가 좁다. HOLA는 cache path에만 learnable RMSNorm-$\gamma$를 적용한다.

Head dimension $d=256$이면 $\sqrt{d}=16$이지만 논문의 measured norm discussion과 implementation setting에 따라 effective matching scale은 unit-L2보다 크게 형성된다. 핵심은 absolute number보다 state path와 cache path의 normalization 목적이 다르다는 점이다.

  • State path는 stable rank-1 update가 목적이다.
  • Cache path는 selective near-argmax retrieval이 목적이다.

두 path를 분리하지 않고 state key norm까지 키우면 update operator가 불안정해질 수 있다.

6) Cache gate and output mixing

HOLA layer는 state output과 cache output을 합친다.

\[o_t = o_t^{\mathrm{state}} + \lambda_t o_t^{\mathrm{cache}}\]

Cache gate는 초기값 $-4$에서 시작한다. 초기에는 cache contribution을 작게 두고 training이 필요에 따라 사용량을 키우는 conservative initialization이다.

이 선택은 pretrained GDN을 그대로 retrofit하는 실험은 아니지만, optimization 초기에 새 branch가 backbone computation을 과도하게 흔드는 것을 줄이는 역할을 한다.

7) Parameter and memory overhead

340M configuration은 다음과 같다.

  • 24 layers
  • 4 heads
  • Head dimension 256
  • Model dimension 1024

Cache-specific learned parameter는 Q/K RMSNorm scale, per-head sink, cache gate다.

\[L(2d+2H) = 24(512+8) = 12{,}480\]

이는 전체 model의 $0.004\%$ 미만이다.

Cache state는 parameter가 아니라 inference-time memory다. bf16 기준으로 layer마다 persistent cache와 current chunk KV를 저장한다. 논문은 약 31 MB의 cache state를 계산하고, batch size 1 decoding에서 다음 peak allocation을 보고한다.

Model 32k 128k
GDN 0.72 GB 0.72 GB
HOLA 0.75 GB 0.75 GB

Context length가 늘어도 bounded state이므로 allocation이 flat하게 유지된다. HOLA의 peak memory overhead는 GDN 대비 약 5%다.

8) HOLA와 다른 hybrid의 차이

Method family Exact memory selection Exact memory location Main trade-off
Full attention All prefix tokens Growing KV cache Strong recall, $O(T)$ cache
Sliding-window hybrid Most recent tokens Local attention window Local exactness, distant loss
Inter-layer hybrid Depends on attention layer Selected full/local layers Architecture mixing
Learned eviction cache Learned score Bounded cache Additional module
HOLA Top-$w$ state write magnitude Every GDN layer Bounded selective exactness

4. Training / Data / Recipe

4-1. Data

Main 340M model은 SlimPajama 15.0B tokens로 학습된다.

Item Setting
Corpus SlimPajama
Train tokens 15.0B
Tokenizer Mistral tokenizer
Vocabulary 32,000
Context length 2,048
Epoch 1

Scale experiment는 세 configuration을 사용한다.

Scale Model dim Layers Corpus Train tokens Context
46M 512 12 FineWeb-Edu 0.5B 4,096
170M 1,024 12 SlimPajama 6.22B 2,048
340M 1,024 24 SlimPajama 15.0B 2,048

46M model은 component ablation에 사용된다. 170M과 340M은 GDN recipe family를 따른다.

4-2. Training strategy

340M optimization recipe는 다음과 같다.

Item Value
Optimizer AdamW
Peak learning rate $4\times10^{-4}$
Weight decay 0.01
Schedule Cosine decay
Warmup 1,000 steps
Gradient clipping 1.0
Global batch 0.5M tokens per update
Training length 1 epoch
Hardware 8 x NVIDIA A800

GDN anchor와 HOLA는 backbone과 training recipe를 동일하게 두고 cache만 다르게 한다. 이는 HOLA gain이 parameter count나 base architecture change에서 온 것이 아니라 memory mechanism에서 왔는지 보기 위한 중요한 control이다.

4-3. Cache hyperparameters

Main HOLA cache setting은 다음과 같다.

Hyperparameter Value
Persistent cache window $w$ 64
Processing chunk $C$ 256
Eviction Top-$w$ by $\beta|e|$
Cache normalization RMSNorm-$\gamma$
Initial temperature 1.0
Temperature training Frozen
Gate initialization -4.0
Cache kernel SDPA

Matched HOLA+recency control은 같은 cache size, chunk, read path, gate를 사용하고 eviction만 recent token 유지로 바꾼다.

4-4. Training and inference semantics

Top-$w$ score는 token이 state에 write될 때 결정된다. 따라서 training chunk 안에서 계산한 score와 streaming decoding에서 online으로 계산한 score가 같은 의미를 가진다.

이 점은 cache selection method에서 중요하다. Training에서는 full sequence를 보고 global top-k를 고르지만 inference에서는 future를 모르는 식의 mismatch가 없어야 한다.

HOLA는 score가 causal state update에서 나오므로 다음 구조가 가능하다.

  • Training: Chunkwise score computation and bounded top-k merge
  • Prefill: Blockwise cache update
  • Decode: One-token online admission and eviction

4-5. Engineering notes

1) Cache는 model weight가 아니다

31 MB 수치는 trainable parameter가 아니라 per-sequence runtime state다. Batch size, beam count, sequence concurrency가 늘면 aggregate memory는 함께 늘 수 있다.

2) Persistent cache와 chunk KV를 구분해야 한다

$w=64$만 보고 layer가 64개 token만 exact하게 본다고 해석하면 안 된다. Current chunk 256개와 null sink가 visible set에 더해진다. 반대로 history 전체에서 persistent하게 남는 exact token은 64개다.

3) GDN kernel에 cache path가 추가된다

공개 구현은 Flash Linear Attention codebase 위에 HOLA-specific fused decode와 evaluation path를 추가한다. Theoretical bounded memory와 실제 throughput은 kernel implementation에 영향을 받는다.

4) Recency control evaluation path를 구분해야 한다

공개 repository와 논문은 recency control에 corrected full-recompute evaluation path를 사용한다고 설명한다. Reproduction에서는 model variant뿐 아니라 evaluation path도 맞춰야 한다.

5) 2k training에서 32k test는 extrapolation이다

HOLA가 32k에서 recall을 유지했다는 것은 32k context로 language modeling pretraining했다는 뜻이 아니다. Main model은 2,048 context에서 학습되고 RULER에서 최대 32k를 평가한다.

5. Evaluation

5-1. Main results

1) Language modeling perplexity

340M same-backbone comparison은 다음과 같다.

Model Wikitext PPL LAMBADA PPL
GDN anchor 27.32 30.95
HOLA+recency 25.04 32.33
HOLA 22.92 30.26
Transformer++ full attention 26.88 42.15

HOLA는 GDN anchor 대비 Wikitext perplexity를 27.32에서 22.92로 낮춘다. Relative reduction은 16.1%다. LAMBADA perplexity도 30.95에서 30.26으로 개선된다.

Transformer++보다 Wikitext perplexity가 낮다는 결과는 흥미롭지만, 한 benchmark의 perplexity만으로 bounded memory가 full attention보다 일반적으로 우월하다고 해석해서는 안 된다. Exact extraction에서는 full attention이 여전히 강하다.

2) In-context retrieval and extraction

Task GDN anchor HOLA Relative or absolute change
FDA 11.7 20.1 +72% relative
SWDE 29.0 35.9 +24% relative
SQuAD extraction 32.5 33.8 +1.3 points

Cache가 설계된 목적과 가장 직접적으로 맞는 결과다. Exact span이나 key-value association을 state-only model보다 더 잘 보존한다.

다만 FDA에서 full-attention Transformer++는 46.1이다. HOLA가 linear baseline을 크게 개선해도 pure exact extraction gap을 닫지는 못한다.

3) Commonsense evaluation

Six-task average는 다음 수준이다.

Model Commonsense average
GDN anchor 42.54
HOLA+recency 43.17
HOLA 42.85
KDA 42.75

차이는 약 0.6 point 안쪽이다. 논문도 single-seed noise 범위로 해석한다.

이 결과는 method의 의미를 오히려 분명하게 한다. HOLA는 모든 benchmark를 일괄적으로 올리는 universal capacity expansion이 아니라 perplexity와 retrieval bottleneck을 겨냥한다.

4) Scale consistency

Scale GDN PPL HOLA PPL
46M 71.0 59.5
170M 35.98 30.51
340M 27.32 22.92

세 scale에서 모두 15-16% 수준의 relative Wikitext PPL reduction이 나타난다. 단일 model size에만 맞춘 artifact일 가능성을 줄여준다.

5) Long-context RULER

Main model은 2k context로 학습되지만 RULER에서 2k, 4k, 8k, 16k, 32k를 평가한다.

Headline S-NIAH-1의 32k result는 다음과 같다.

Model 32k recall
GDN 0.14
HOLA+recency 0.24
HOLA 0.58

중요한 comparison은 HOLA와 HOLA+recency다. Same cache budget과 read mechanism을 사용하면서 selection rule만 바뀐다. Distant needle이 recent window 밖으로 밀려나도 surprise-selected cache에는 남을 수 있음을 보여준다.

HOLA도 0.58이므로 perfect recall은 아니다. Bounded cache의 capacity limit가 그대로 드러난다.

6) Multi-task RULER snapshot

8k에서 일부 task는 다음과 같다.

Model S-NIAH-1 S-NIAH-2 Multi-key-1 Multi-value
GDN 0.83 0.09 0.03 0.07
HOLA+recency 0.74 0.04 0.05 0.03
HOLA 0.98 0.35 0.11 0.16

모든 cell에서 압도하는 것은 아니지만, floor가 아닌 difficult retrieval cell에서 HOLA가 대체로 우세하다.

7) Full-attention baseline caveat

Transformer++는 2k training length 안에서 strong ceiling을 제공한다. 그러나 RoPE checkpoint가 4k 이상에서 모든 shown RULER task 0으로 떨어진다.

따라서 4k-32k result에서 HOLA가 Transformer++보다 낫다는 사실을 architecture-level long-context superiority로 해석하면 안 된다. 해당 Transformer checkpoint가 length extrapolation baseline으로 적절하지 않기 때문이다.

5-2. Ablation: what to cache

46M flat-read diagnostic에서 eviction signal을 비교한다.

Eviction score Far needle 4-key capacity Wiki PPL
No cache GDN 0.55 0.41 70.21
Cumulative attention 0.43 0.55 71.45
$|e|$ 0.22 0.48 70.50
$\beta|v|$ 0.42 0.54 70.76
$\beta|e|$ 0.67 0.56 70.10

Residual magnitude만으로는 충분하지 않다. State prediction error가 커도 model이 실제로 강하게 write하지 않았다면 future utility가 낮을 수 있다. 반대로 $\beta|v|$는 write gate를 보지만 state가 이미 잘 표현하는지 반영하지 않는다.

Product인 $\beta|e|$가 두 정보를 결합한다.

5-3. Ablation: how to read

Cache selection을 $\beta|e|$로 고정하고 read normalization을 비교한다.

Read Wiki PPL Far needle 4-key 8-key 16-key
GDN no cache 70.21 0.55 0.41 0.35 0.31
Unit-L2 cache read 70.10 0.67 0.56 0.41 0.31
RMSNorm-$\gamma$ cache read 59.5 0.75 0.77 0.67 0.41

Exact KV를 저장하는 것만으로는 perplexity gain이 거의 없다. RMSNorm-$\gamma$로 read를 sharpen했을 때 큰 변화가 나타난다.

이 ablation은 HOLA의 가장 중요한 실험 중 하나다. Memory research에서 storage policy만 보고 retrieval operator를 가볍게 다루면 안 된다는 점을 보여준다.

5-4. What really matters in the experiments

1) Same-backbone control이 가장 중요하다

Borrowed baseline row는 broader context를 제공하지만, method causality는 own GDN anchor와 HOLA comparison에서 봐야 한다. Backbone, parameter scale, corpus, context, optimizer를 맞춘 상태에서 cache만 추가된다.

2) HOLA+recency가 key ablation이다

Cache 유무만 비교하면 gain이 exact memory 자체인지 selection rule인지 알기 어렵다. HOLA+recency는 same-memory comparison으로 surprise-based retention의 효과를 분리한다.

3) Perplexity와 exact extraction을 함께 봐야 한다

Wikitext PPL은 크게 좋아지지만 FDA에서는 full attention과 gap이 남는다. 이는 HOLA가 state compression을 보완하지만 unbounded token-level access를 완전히 대체하지는 않는다는 뜻이다.

4) Long-context recall은 capacity와 extrapolation을 함께 본다

32k recall 0.58은 강한 relative improvement지만 perfect memory가 아니다. 또한 2k-trained Transformer++ collapse는 fair architecture comparison이 아니라 checkpoint extrapolation failure를 포함한다.

5) Commonsense tie는 negative result가 아니라 scope 확인이다

Method가 targeted bottleneck에만 작동하고 unrelated capability를 크게 흔들지 않는다는 evidence로 읽을 수 있다. 다만 single seed라 작은 차이는 의미를 부여하기 어렵다.

6. Limitations

  1. Exact cache capacity가 작고 고정되어 있다.
    • Persistent memory는 layer당 64 token이다.
    • Current chunk와 null sink를 포함한 visible exact set도 약 321 token이다.
    • Needle이 많거나 여러 distant fact가 모두 중요한 context에서는 필요한 item을 모두 보존할 수 없다.
  2. 32k recall도 완전하지 않다.
    • S-NIAH-1에서 0.58은 GDN 0.14보다 높지만, exact recall system으로는 많은 failure가 남는다.
    • Cache admission score가 중요 token을 놓치거나 capacity competition이 발생할 수 있다.
  3. Full attention의 exact extraction gap을 닫지 못한다.
    • FDA에서 HOLA 20.1은 same-backbone GDN 11.7보다 크게 높다.
    • 그러나 full-attention Transformer++ 46.1과는 큰 차이가 있다.
    • Every-token access가 필요한 task에서는 bounded cache의 구조적 한계가 남는다.
  4. Main-scale result는 340M까지다.
    • 46M, 170M, 340M에서 consistency는 확인된다.
    • Multi-billion parameter LLM에서 cache admission distribution, layer specialization, throughput이 같은지 확인되지 않았다.
  5. Main comparison은 single seed다.
    • Commonsense의 작은 차이는 특히 해석하기 어렵다.
    • 46M cache-read ablation 일부는 multi-seed지만, 340M main table 전반의 variance는 충분히 제시되지 않는다.
  6. Learned eviction과 matched comparison이 없다.
    • 논문은 LTE 같은 learned eviction module과 동일 backbone, cache budget, recipe로 직접 비교하지 않는다.
    • Parameter-free score의 simplicity는 장점이지만 optimality를 입증한 것은 아니다.
  7. Surprise score는 future relevance를 직접 예측하지 않는다.
    • $\beta|e|$는 current state update의 크기다.
    • 현재 state가 어려워한 token과 미래 query에서 중요한 token이 항상 같지는 않다.
    • Low-surprise but later-critical token은 cache에서 밀릴 수 있다.
  8. Cache overhead는 concurrency에 따라 누적된다.
    • Per-sequence state는 작고 context length에 대해 bounded하다.
    • 하지만 serving에서 batch, beam, concurrent request가 늘면 layer별 cache state도 함께 증가한다.
  9. Kernel efficiency는 구현에 의존한다.
    • Peak memory가 flat하다는 것과 wall-clock throughput이 항상 GDN과 비슷하다는 것은 같은 주장이 아니다.
    • Top-k maintenance, cache attention, chunk merge, fused decode quality를 실제 hardware에서 따로 측정해야 한다.
  10. Long-context Transformer comparison은 제한적이다.
    • 2k-trained RoPE Transformer++가 4k 이상에서 collapse한다.
    • Long-context-trained or extrapolation-aware full-attention model과의 fair comparison은 추가로 필요하다.

7. My Take

7-1. Why this matters for my work

HOLA의 가장 큰 가치는 작은 cache를 붙였다는 사실보다, recurrent update를 self-diagnostic memory signal로 바꾼 데 있다.

Model은 이미 매 token마다 다음을 계산한다.

  • State prediction
  • Prediction residual
  • Write gate
  • State correction

HOLA는 이 중 state correction magnitude를 보고 “이 token은 compressed memory가 약하게 표현할 가능성이 높다”고 판단한다. 별도의 salience label 없이 model 내부의 learning dynamics를 runtime memory policy로 재사용한다.

이는 efficient long-context model 설계에서 유용한 방향이다. Memory selection을 external heuristic, recency, query-independent token importance로만 두지 않고, backbone이 실제로 겪는 representation difficulty에 연결할 수 있다.

Document AI나 agent memory에도 비슷한 질문을 적용할 수 있다.

  • OCR document stream에서 state가 크게 수정된 field는 무엇인가.
  • 긴 contract에서 반복 규칙과 one-off exception을 다른 memory에 둘 수 있는가.
  • Search agent trajectory에서 compressed summary가 놓친 observation을 exact buffer에 남길 수 있는가.
  • Streaming model에서 local window 밖의 rare identifier를 bounded cache에 유지할 수 있는가.

HOLA를 그대로 agent memory로 옮기는 것은 아니지만, compressive memory와 exact exception memory를 분리하는 원리는 재사용 가치가 높다.

7-2. Reuse potential

1) Linear-attention language model

GDN, DeltaNet 계열 backbone에 per-layer bounded cache를 추가할 수 있다. 가장 직접적인 재사용이다.

Reproduction에서는 다음을 우선 확인해야 한다.

  • Baseline GDN recipe 재현
  • Surprise score top-k consistency
  • Cache-only RMSNorm path
  • Gate initialization
  • Prefill and decode equivalence
  • Recency matched control

2) Streaming document understanding

긴 document를 chunkwise 처리하면서 recurrent state는 overall structure를 유지하고, exact cache는 rare field, amount, ID, date, exception clause를 보존하도록 설계할 수 있다.

다만 $\beta|e|$가 business-critical field를 항상 선택한다는 보장은 없다. Domain-specific utility signal과 결합할 필요가 있다.

3) Memory audit

어떤 token이 layer별 top-$w$ cache에 남는지 분석하면 model이 무엇을 surprising하게 보는지 관찰할 수 있다.

  • Entity and number concentration
  • Rare token retention
  • Layer별 syntactic vs semantic memory
  • Long-context failure에서 evicted evidence
  • Same token의 layer별 score variation

이 분석은 architecture improvement뿐 아니라 interpretability에도 연결된다.

4) Query-aware extension

HOLA admission은 write-time surprise 기반이다. Future extension으로 query-time relevance를 섞을 수 있다.

예를 들면 다음 score family를 생각할 수 있다.

\[s_i = a\,\beta_i\|e_i\| + b\,r(q_t,k_i)\]

단 query-aware re-ranking을 매 step 수행하면 compute와 cache semantics가 달라진다. HOLA의 parameter-free online simplicity를 잃을 수 있다.

5) Hierarchical exact memory

64-slot per-layer cache를 single pool로 두는 대신 다음 partition을 고려할 수 있다.

  • Surprise pool
  • Recency pool
  • Entity or number pool
  • User-pinned pool

그러나 pool을 늘리면 bounded memory의 simplicity와 fair compute advantage가 약해질 수 있다. 먼저 HOLA처럼 clear single-signal baseline을 두는 것이 좋다.

7-3. Follow-up papers

  • Gated Delta Networks: Improving Mamba2 with Delta Rule
  • Parallelizing Linear Transformers with the Delta Rule over Sequence Length
  • Zoology: Measuring and Improving Recall in Efficient Language Models
  • Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff
  • Repeat After Me: Transformers are Better than State Space Models at Copying
  • Native Hybrid Attention for Efficient Sequence Modeling
  • Artificial Hippocampus Networks for Efficient Long-Context Modeling
  • Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
  • RULER: What’s the Real Context Size of Your Long-Context Language Models?
  • Memorizing Transformers

8. Summary

  • HOLA는 GDN recurrent state를 compressive memory로 유지하고 per-layer bounded exact KV cache를 추가한다.
  • Cache는 최근 token이 아니라 state update magnitude $\beta|e|$가 큰 top-$w$ token을 저장한다.
  • Cache read에는 별도의 RMSNorm-$\gamma$를 적용해 exact KV가 uniform average로 섞이지 않도록 sharpen한다.
  • 340M same-backbone comparison에서 Wikitext PPL 27.32에서 22.92, FDA 11.7에서 20.1, 32k S-NIAH-1 recall 0.14에서 0.58로 개선한다.
  • 핵심 한계는 64-slot persistent cache의 bounded capacity, full attention 대비 exact extraction gap, 340M single-seed validation이다.

댓글남기기