X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Review
0. Introduction
Speech LLM을 가볍게 만드는 가장 직접적인 방법 중 하나는 audio encoder의 layer를 줄이는 것이다. 하지만 block을 통째로 제거하면 FLOPs만 줄어드는 것이 아니다. Decoder가 받아들이던 acoustic embedding의 geometry가 바뀌고, 기존 language backbone이 익숙하지 않은 representation을 보게 된다. 그 결과 transcript 중간을 건너뛰는 deletion error, 너무 일찍 generation을 끝내는 premature EOS, domain별 accuracy 붕괴가 함께 나타날 수 있다.
X-AuT는 이 문제를 단순한 layer pruning으로 보지 않는다. 어떤 layer 조합을 제거할지 짧은 recovery probe로 비교하고, pruning을 한 번에 끝내지 않고 $18 -> 16 -> 14$로 진행한다. 이후 representation alignment, cross-scale teacher distillation, student-generated prefix를 포함한 scheduled supervision, 마지막 gold-transcript finetuning을 순서대로 적용한다.
이 논문을 지금 볼 가치가 있는 이유는 다음과 같음.
- Speech LLM compression을 parameter count가 아니라 decoder-facing representation mismatch 문제로 다룬다.
- Individual layer score를 더하는 대신 pair interaction을 실제 recovery run으로 비교한다.
- Same-scale self-distillation과 larger cross-scale teacher를 동일한 16-layer setting에서 비교한다.
- Accuracy, audio-tower parameter, encoder latency, end-to-end latency를 함께 보고 operating point를 분리한다.
- 공개 checkpoint와 inference code는 제공하지만 full three-stage reproduction 범위는 제한된다는 점도 명확히 드러난다.
한 줄 요약: X-AuT는 Qwen3-ASR-0.6B의 18-layer audio encoder를 behavior-driven probe로 단계적으로 줄이고, cross-scale representation and logit distillation과 student-policy supervision으로 pruning mismatch를 복구해 16-layer accuracy point와 14-layer efficiency point를 만드는 speech LLM compression framework다.
1. Problem Setting
1-1. Speech LLM에서 audio tower를 줄이는 문제
일반적인 speech LLM은 크게 세 부분으로 나눌 수 있다.
- Audio encoder
- Waveform or acoustic feature를 contextual acoustic embedding으로 바꾼다.
- Bridge
- Audio encoder dimension을 language-model hidden dimension에 맞춘다.
- Autoregressive decoder
- Acoustic embedding과 text prefix를 조건으로 transcript token을 생성한다.
Input audio를 $x$, audio encoder를 $E_\theta$, bridge를 $B_\phi$, decoder를 $D_\psi$라고 하면 generation path는 아래처럼 볼 수 있다.
\[h_a = B_\phi(E_\theta(x))\] \[p(y_t \mid y_{<t}, x) = D_\psi(y_{<t}, h_a)\]Audio encoder의 layer를 $N$개에서 $M$개로 줄이면 compute는 감소한다. 하지만 pruned encoder $E_{\theta’}$가 만드는 $h_a’$는 original decoder가 학습한 $h_a$와 다르다.
\[h_a' = B_{\phi'}(E_{\theta'}(x))\]Compression의 실제 목표는 parameter를 줄이면서 decoder가 사용할 수 있는 acoustic information을 유지하는 것이다.
\[\min_{\theta',\phi'} \mathcal{E}(D_\psi, B_{\phi'}(E_{\theta'}(x))) \quad \text{s.t.} \quad M < N\]여기서 $\mathcal{E}$는 CER, WER, deletion, EOS behavior를 포함한 downstream error를 뜻한다.
1-2. 왜 단순 layer removal이 부족한가
1) Layer importance는 독립적이지 않다
Layer 하나를 제거했을 때의 damage가 작아도, 두 layer를 동시에 제거하면 interaction 때문에 error가 크게 늘 수 있다. 반대로 individually important해 보이는 layer 조합이 recovery training 이후에는 더 잘 복구될 수도 있다.
즉 아래 가정은 안전하지 않다.
\[\Delta_{i,j} \approx \Delta_i + \Delta_j\]X-AuT의 probe 결과에서도 individual removal score만 보고 고른 pair가 best pair가 아니었다. 논문은 이 비가산성을 이유로 explicit pair probing을 사용한다.
2) Hidden-state similarity가 transcript behavior를 보장하지 않는다
Pruned encoder output이 cosine similarity 관점에서 teacher와 가까워도, autoregressive decoder가 조금 다른 acoustic evidence를 보고 EOS probability를 크게 올릴 수 있다. Speech recognition은 token-level local similarity보다 sequence behavior에 민감하다.
특히 failure는 다음처럼 나타난다.
- 문장 중간 구간을 건너뛰는 deletion
- 긴 utterance에서 early termination
- English and Chinese domain별 asymmetric degradation
- Clean speech에서는 유지되지만 meeting or noisy speech에서 급격한 하락
3) Frozen decoder는 안정성과 mismatch를 동시에 만든다
Language-model backbone을 freeze하면 parameter drift와 training cost를 줄일 수 있다. 그러나 decoder가 acoustic representation 변화에 적응할 자유도도 제한된다. 따라서 audio encoder, bridge, LoRA adapter, tied embedding 중 어디까지 열어둘지 중요하다.
4) Compression ratio와 end-to-end latency는 같은 지표가 아니다
Audio tower parameter가 20% 줄어도 전체 system latency가 20% 줄지는 않는다. Text decoder가 runtime 대부분을 차지하면 encoder speedup은 end-to-end에서 희석된다. Speech LLM compression은 audio-only metric과 full pipeline metric을 모두 봐야 한다.
2. Core Idea
2-1. Behavior-driven progressive pruning
X-AuT는 $18 -> 14$ layer를 한 번에 만들지 않는다.
- Original 18-layer encoder에서 candidate layer removal을 probe한다.
- Short recovery budget으로 candidate behavior를 비교한다.
- First stage에서 original layer 1 and 18을 제거해 16-layer model을 만든다.
- 16-layer state에서 pair candidates를 다시 probe한다.
- Selected pair인 original layer 5 and 6을 추가로 제거해 14-layer model을 만든다.
최종 14-layer encoder는 original 18-layer stack에서 아래 layer를 유지한다.
[2, 3, 4, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17]
여기서 중요한 점은 adjacent layer removal 자체가 universal rule이라는 뜻이 아니라는 것이다. 논문은 tested candidate interaction이 non-additive였기 때문에, 실제 short recovery probe가 필요하다고 해석한다.
2-2. Compression 이후 mismatch를 세 단계로 복구한다
X-AuT의 recovery recipe는 다음 세 stage로 구성된다.
Stage 0: Representation bridge recovery
초기에는 pruned encoder가 teacher의 intermediate representation과 bridge output을 따라가도록 만든다. Transcript cross-entropy와 logit distillation도 함께 사용한다.
개념적인 total loss는 다음처럼 볼 수 있다.
\[\mathcal{L}_{0} = \lambda_{hid}\mathcal{L}_{hid} + \lambda_{bridge}\mathcal{L}_{bridge} + \lambda_{logit}\mathcal{L}_{logit} + \lambda_{ce}\mathcal{L}_{ce}\]이 stage는 전체 schedule의 앞부분에서 acoustic representation의 급격한 drift를 먼저 줄이는 역할을 한다.
Stage 1: Cross-scale distillation with student-policy contexts
Frozen Qwen3-ASR-1.7B teacher가 pruned student를 supervise한다. Teacher audio hidden width 2048은 learned projection을 통해 student width 1024로 맞춘다.
Teacher-forced transcript prefix만 사용하면 inference 시 student가 만든 error prefix를 보지 못한다. X-AuT는 training 후반 일부 step에서 student-generated context를 사용한다. Teacher와 student는 같은 student prefix 위에서 distribution을 비교한다.
이 설계는 exposure mismatch를 줄이는 방향이다.
Stage 2: Gold-transcript finetuning
마지막에는 낮은 learning rate와 gold transcript cross-entropy로 model을 정리한다. Source mixture도 target domain 쪽으로 다시 조정한다. 이 stage에서는 tied output embedding을 freeze하고, audio encoder와 bridge, decoder attention LoRA를 중심으로 adaptation한다.
2-3. Cross-scale teacher가 필요한 이유
Same-scale self-teacher는 student가 이미 잃은 representation capacity를 충분히 복구하지 못할 수 있다. X-AuT는 1.7B teacher를 사용해 richer acoustic representation and token distribution을 제공한다.
Matched 16-layer Stage 1 recipe에서 논문은 다음 macro mean error를 보고한다.
| Teacher setting | Macro mean error |
|---|---|
| Same-scale self-teacher | 8.45% |
| Cross-scale 1.7B teacher | 5.55% |
이 결과는 larger teacher가 해당 setting에서 더 좋은 supervision을 제공했다는 evidence다. 다만 single-run descriptive comparison이므로 teacher size만이 causal factor라고 단정하면 안 된다.
2-4. Accuracy point와 efficiency point를 분리한다
X-AuT는 하나의 best checkpoint만 제시하지 않는다.
- 16-layer model
- Audio tower parameter를 약 10.35% 줄인다.
- Macro mean error를 5.61%에서 5.27%로 낮춘다.
- Compression과 accuracy가 동시에 좋아지는 operating point다.
- 14-layer model
- Audio tower parameter를 약 20.70% 줄인다.
- Macro mean error는 5.75%다.
- Baseline보다 2.5% relative worse지만 더 큰 latency reduction을 얻는다.
Deployment requirement에 따라 두 point 중 하나를 선택하게 하는 구성이다.
3. Architecture / Method
3-1. Overview
| Item | Description |
|---|---|
| Base model | Qwen3-ASR-0.6B |
| Original audio encoder | 18 Transformer blocks, 186.376M parameters |
| Compressed variants | 16-layer and 14-layer audio encoders |
| Language backbone | Frozen Qwen3 0.6B decoder base |
| Trainable decoder part | Attention projection LoRA, rank 32 |
| Teacher | Frozen Qwen3-ASR-1.7B |
| Recovery | Representation KD, bridge KD, logit KD, CE, student-policy contexts |
| Final checkpoint | 14-layer full model plus inference and limited finetuning code |
3-2. What is actually pruned
X-AuT는 audio tower의 Transformer block만 제거한다.
- ConvStem은 유지한다.
- Bridge는 구조를 유지한다.
ln_post는 유지한다.- Text decoder 28 layers는 pruning하지 않는다.
- Token embedding and LM head는 tied structure를 유지한다.
14-layer model의 audio tower breakdown은 다음과 같다.
| Module | 18-layer baseline | 14-layer X-AuT |
|---|---|---|
| ConvStem | 11.03M | 11.03M |
| Transformer layers | 173.62M | 135.04M |
| Bridge | 1.72M | 1.72M |
| Audio tower total | 186.376M | 147.794M |
Pruning target이 명확하므로 speedup이 어느 component에서 나오는지 해석하기 쉽다.
3-3. Behavior probe
Probe는 candidate architecture를 full schedule로 모두 학습하지 않고, fixed small development and validation subset에서 짧게 recovery한다. 평가 기준은 parameter proxy가 아니라 transcript behavior다.
Probe가 필요한 이유는 다음과 같다.
- Layer pair interaction을 확인한다.
- Deletion and early EOS를 직접 잡는다.
- 같은 compute budget에서 candidate recoverability를 비교한다.
- 18-layer에서 좋은 removal이 16-layer state에서도 좋은지 다시 확인한다.
논문은 first pruning 이후 second pruning을 다시 탐색한다. Progressive process가 direct $18 -> 14$ pruning보다 좋았다는 결과도 제시한다.
| Strategy | 14-layer macro mean error |
|---|---|
| Progressive 18 -> 16 -> 14 | 5.75% |
| Direct 18 -> 14 | 6.73% |
3-4. Cross-layer representation matching
Teacher and student는 layer count와 hidden width가 다르다. X-AuT는 grouped layer mapping과 learned projection을 사용한다.
Teacher feature를 $h_T^{(k)}$, student feature를 $h_S^{(j)}$, projection을 $P$라고 하면 alignment는 아래처럼 볼 수 있다.
\[\mathcal{L}_{hid} = \sum_j \left\| P(h_T^{(m(j))}) - h_S^{(j)} \right\|_2^2\]- $m(j)$는 student layer에 대응하는 teacher layer group을 정한다.
- $P$는 2048-d teacher feature를 1024-d student space로 옮긴다.
- Teacher는 inference 시 제거된다.
3-5. Scheduled student-policy supervision
Autoregressive ASR에서 teacher-forcing only training은 inference distribution을 충분히 반영하지 못한다. X-AuT는 Stage 1 후반에 일정 간격으로 student rollout prefix를 만든다.
- Student가 prefix를 생성한다.
- Teacher and student가 동일한 prefix를 조건으로 다음-token distribution을 계산한다.
- Top-k union support에서 distribution matching을 수행한다.
- Degenerate or empty rollout에는 fallback을 적용한다.
이 방식은 pure on-policy distillation은 아니지만, student가 실제로 방문하는 prefix를 일부 training signal에 포함한다.
4. Training / Data / Recipe
4-1. Transcript-consistency filtering
Source pool은 280k hours를 넘는다. 모든 sample을 동일하게 사용하지 않고, source transcript와 두 offline ASR hypothesis의 agreement를 사용해 confidence class를 만든다.
- Source transcript
- Qwen3-ASR-1.7B hypothesis
- Qwen3.5-Omni hypothesis
세 transcript의 consistency에 따라 9개 class로 나누고, reported main training은 세 stage 모두 class 1만 사용한다. Stage 2에서는 class tier를 넓히는 대신 source weight를 target domain 쪽으로 바꾼다.
따라서 이 논문이 multi-tier curriculum의 효과를 검증했다고 읽으면 안 된다. 공개된 main result는 highest-agreement subset과 later source reweighting의 조합이다.
4-2. Trainable and frozen parameters
Training boundary는 다음과 같다.
- Audio encoder: trainable
- Bridge: trainable
- Decoder base weights: frozen
- Decoder q, k, v, o projections: rank-32 LoRA trainable
- Tied output embedding: distillation stage에서 trainable, final finetuning에서 frozen
- Cross-scale teacher: frozen
- Teacher projection: training-time only
이 구성은 decoder 전체 fine-tuning보다 memory and stability cost를 줄인다. 동시에 decoder가 acoustic representation shift에 적응할 최소 경로를 LoRA and output embedding으로 남긴다.
4-3. Stage-by-stage interpretation
Stage 0
목적은 pruned encoder representation을 빠르게 원래 interface 근처로 돌리는 것이다.
- Intermediate hidden alignment
- Bridge output alignment
- Logit distillation
- Gold transcript CE
Stage 1
목적은 stronger teacher의 acoustic and token knowledge를 student behavior에 연결하는 것이다.
- Cross-scale hidden distillation
- Bridge and logit distillation
- Teacher-forced context
- Scheduled student-policy context
Stage 2
목적은 deployment target distribution에 맞게 최종 transcript behavior를 안정화하는 것이다.
- Gold transcript CE only
- Lower learning rate
- Source reweighting
- Tied embedding freeze
4-4. Engineering notes
1) Probe budget and candidate set을 기록해야 한다
Layer selection은 candidate set과 short recovery budget에 민감하다. 어떤 pair를 탐색했는지, sample subset, optimizer, step 수를 함께 저장해야 한다.
2) EOS and deletion metric을 별도로 봐야 한다
Average CER or WER만 보면 early termination failure가 숨을 수 있다. Utterance length별 deletion rate, output-to-reference length ratio, EOS position을 함께 로그하는 편이 안전하다.
3) Decoder-dominant latency를 사전에 측정해야 한다
Audio tower만 최적화해도 end-to-end speedup이 작을 수 있다. Encoder, bridge, first-token decoder, autoregressive decode 시간을 분리해 profile해야 한다.
4) Teacher projection을 reproduction artifact로 보존해야 한다
Cross-scale KD는 hidden-width projection과 layer grouping에 의존한다. Checkpoint만 공개하고 mapping code가 없으면 Stage 0 and 1을 정확히 재현하기 어렵다.
5) Domain macro mean은 workload distribution과 다를 수 있다
10 benchmark unweighted average는 research comparison에는 편리하지만 actual product traffic mixture를 반영하지 않는다. Deployment에서는 domain-weighted WER/CER와 tail utterance를 다시 측정해야 한다.
5. Evaluation
5-1. Accuracy-compression trade-off
Paper는 Chinese and English 10개 public benchmark에서 CER or WER를 측정하고 unweighted macro mean을 보고한다.
| Model | Audio tower params | Param change | Macro mean error | Relative change |
|---|---|---|---|---|
| Full 18-layer baseline | 186.376M | 0.00% | 5.61% | 0.0% |
| X-AuT 16-layer Stage 2 | 167.085M | -10.35% | 5.27% | -6.1% |
| X-AuT 14-layer Stage 2 | 147.794M | -20.70% | 5.75% | +2.5% |
16-layer model은 mean accuracy도 좋아진다. 이는 mild pruning이 regularization or recovery recipe와 결합해 baseline보다 나은 point를 만들 수 있음을 보여준다.
14-layer model은 benchmark별 결과가 균일하지 않다. LibriSpeech test-clean and CommonVoice zh에서는 baseline을 앞서지만, 나머지 다수 benchmark에서는 degradation이 남는다. 따라서 5.75% 하나만 보고 모든 domain에서 near-lossless라고 부르면 안 된다.
5-2. Inference efficiency
14-layer model의 reported latency change는 다음과 같다.
| Metric | In-vehicle PPU | H800 |
|---|---|---|
| Audio encoder latency | -21.4% | -11.4% |
| End-to-end latency | -4.7% | -2.6% |
| Peak memory | -4.4% | -2.8% |
Audio encoder speedup은 parameter reduction과 비슷한 방향이지만, full pipeline gain은 훨씬 작다. Unpruned text decoder가 전체 inference cost에서 큰 비중을 차지하기 때문이다.
이 결과가 중요한 이유는 compression target을 어디로 잡아야 하는지 보여주기 때문이다.
- Streaming front-end budget이 병목이면 audio encoder pruning이 유효하다.
- End-to-end latency가 핵심이면 decoder quantization, speculative decoding, token reduction 같은 추가 최적화가 필요하다.
- Memory bottleneck이면 audio-only 20.7% reduction이 whole-model에서는 작은 변화일 수 있다.
5-3. What really matters in the experiments
1) Progressive pruning comparison
Direct 14-layer pruning보다 progressive 18 -> 16 -> 14가 나은 결과는 central claim을 직접 지지한다. 다만 candidate selection cost까지 포함한 total research compute도 함께 봐야 한다.
2) Cross-scale versus self-teacher
Matched 16-layer recipe에서 5.55% versus 8.45%는 큰 차이다. 하지만 repeated seeds가 없기 때문에 effect size의 안정성과 teacher checkpoint choice를 추가 검증해야 한다.
3) Stage 2 recovery
14-layer Stage 1 macro mean은 6.48%, Stage 2는 5.75%다. Final CE finetuning and source reweighting이 deployment quality에 중요하다는 뜻이다.
4) Single-run reporting
Public benchmark values는 best checkpoint single run이다. Bold value가 statistical significance를 뜻하지 않는다. Utterance-level confidence interval and repeated-seed variance가 없다.
6. Limitations
- Full training pipeline이 공개되지 않았다.
- Official repository는 inference, 14-layer checkpoint, minimal Stage-2-style LoRA finetuning을 제공한다.
- Behavior probe, Stage 0 and 1 cross-scale distillation, 280k-hour data pipeline은 공개 범위에 포함되지 않는다.
- Main results가 single run이다.
- Seed variation, checkpoint selection variance, utterance-level confidence interval이 없다.
- 16-layer improvement가 항상 재현되는지 추가 검증이 필요하다.
- Training data가 proprietary or internally authorized source를 포함한다.
- Example manifest는 placeholder audio reference만 제공한다.
- Reported data mixture를 그대로 재현하기 어렵다.
- End-to-end gain이 제한적이다.
- 20.7% audio-tower parameter reduction이 full latency에서는 2.6% to 4.7% improvement로 축소된다.
- Product bottleneck이 decoder라면 pruning 우선순위를 다시 정해야 한다.
- License가 commercial deployment를 제한한다.
- Released model and code의 CC BY-NC 4.0 조건과 repository notice를 실제 적용 전에 확인해야 한다.
- Layer selection이 base architecture에 종속될 수 있다.
- Qwen3-ASR에서 찾은 removal pattern이 다른 speech encoder에도 그대로 적용된다는 evidence는 없다.
7. My Take
7-1. Why this matters for my work
X-AuT의 가장 재사용 가치가 큰 부분은 pruning score가 아니라 recovery contract다. Layer를 제거한 뒤 무엇이 깨졌는지 representation, decoder distribution, on-policy prefix, final transcript behavior로 나눠 복구한다.
Compression project에서 다음 순서로 활용할 수 있다.
- Short behavior probe로 candidate를 좁힌다.
- Intermediate feature보다 actual failure mode를 같이 본다.
- Stronger teacher로 representation and token distribution을 복구한다.
- Student-generated state를 일부 training context에 포함한다.
- 마지막에는 deployment distribution의 supervised objective로 정리한다.
7-2. Reuse potential
1) VLM vision tower compression
Vision encoder block을 줄인 뒤 language decoder에 생기는 grounding mismatch도 유사하다. Cross-scale feature projection and student-prefix distillation을 적용할 수 있다.
2) Document audio or multimodal pipeline
ASR front-end 뒤에 document understanding or agent decoder가 붙는 system에서도 early EOS and deletion은 downstream evidence loss로 이어진다. Sequence-level failure metric을 pruning loop에 포함해야 한다.
3) Edge deployment operating points
One-size-fits-all checkpoint보다 16-layer quality point and 14-layer latency point를 분리하는 방식이 현실적이다. Device class별 model routing에도 적합하다.
4) Probe-guided structured pruning
Static norm or saliency만 쓰지 않고 short recovery outcome으로 architecture choice를 결정하는 방식은 MoE expert, attention head, layer-drop 연구에도 확장할 수 있다.
7-3. Follow-up papers
- Qwen3-ASR Technical Report
- Distil-Whisper: robust knowledge distillation for speech recognition
- LayerDrop: structured dropout for Transformer compression
- On-policy distillation for autoregressive sequence models
- Speech LLM serving and streaming-decoding optimization studies
8. Summary
- X-AuT는 speech LLM audio encoder pruning을 decoder-facing representation mismatch 문제로 정의한다.
- Qwen3-ASR-0.6B의 18 layers를 behavior probe로 $18 -> 16 -> 14$ 순서로 줄인다.
- Representation, bridge, logit, transcript loss와 cross-scale 1.7B teacher를 결합한다.
- 16-layer model은 macro mean error를 5.61%에서 5.27%로 낮추며 10.35% parameter reduction을 얻는다.
- 14-layer model은 audio tower를 20.70% 줄이고 5.75% macro mean error를 기록한다.
- Encoder latency는 의미 있게 줄지만 end-to-end gain은 decoder bottleneck 때문에 작다.
- 공개 checkpoint는 유용하지만 full three-stage reproduction package는 아니다.
댓글남기기