Abstract
Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning. The language-model backbone remains frozen, while attention LoRA adapters and the tied output embedding adapt during distillation. Training uses the highest-agreement tier from a transcript-consistency pipeline, followed by source reweighting during finetuning. On ten public Chinese–English benchmarks, compressing Qwen3-ASR-0.6B from 18 to 16 audio-encoder layers reduces macro-average error from 5.61% to 5.27%. The 14-layer model reaches 5.75% with 20.7% fewer audio-tower parameters. Under the matched recipe, the 1.7B teacher yields 5.55% mean error, compared with 8.45% for self-distillation, and progressive 1814 pruning outperforms direct pruning (5.75% vs. 6.73%). These single-run results establish two practical operating points and show that the accuracy effects vary across benchmarks. Project website: https://xpeng-ai.github.io/x-aut.
1 Introduction
Speech large language models typically combine a deep audio encoder, a bridge that maps acoustic features into the language-model embedding space, and an autoregressive text decoder [1, 2, 3]. This architecture is accurate and flexible, but the audio encoder must process every input frame and therefore remains important for first-token latency in streaming, mobile, and in-vehicle systems [4, 5]. Reducing encoder depth is attractive because it removes complete Transformer blocks and produces a regular, deployment-friendly model.
Starting from a strong pretrained model is also substantially cheaper than training a compact speech encoder and realigning it with a decoder from scratch. We therefore study post-training depth reduction of the 18-layer audio Transformer (AuT) in Qwen3-ASR-0.6B [1]. Removing layers changes the audio embeddings inserted into the decoder and can trigger premature end-of-sequence (EOS) predictions and large deletion errors. Our recipe freezes the pretrained language-model weights while allowing decoder attention LoRA adapters and, during distillation, the tied output embedding to adapt. We use “frozen decoder backbone” in this sense throughout the paper.
Prior ASR compression work has explored decoder distillation, low-rank encoder compression, weight sparsity, supernet training, and encoder-layer removal [6, 7, 8, 9, 10]. For an already-pretrained speech LLM, two practical questions remain: which combinations of layers can be removed and recovered under a fixed training budget, and how should recovery address both hidden-state mismatch and errors induced by the student’s own decoding history?
These questions are coupled. A layer that appears redundant in isolation may become important once another layer is removed, because subsequent blocks receive a shifted representation and the bridge must preserve the decoder interface learned during pretraining. Static importance scores therefore cannot fully predict whether a multi-layer candidate can be recovered within a short budget. Selection and recovery must therefore be evaluated together.
X-AuT addresses these questions through progressive pruning and recovery. Before each pruning hop, short behavioral probes compare candidate layer sets under matched initialization, data, and optimization. The selected student then undergoes representation alignment, distillation with scheduled student-policy contexts, and low-rate LoRA finetuning. A Qwen3-ASR-1.7B teacher supplies cross-scale supervision; grouped layer matching and a learned bottleneck projection accommodate the depth and width differences between teacher and student.
Figure 1 compares the Stage 2 16- and 14-layer models with the unpruned baseline across all ten evaluation sets. Each spoke reports accuracy preservation, , where is CER or WER expressed as a fraction. This transformation provides a common visual reference across benchmarks; all quantitative comparisons use the macro-average error and the exact CER/WER values in Tables 2 and 3.
The two contours show different accuracy–efficiency tradeoffs. The 16-layer model remains close to or above the baseline across the suite and reaches 5.27% macro error. Further compression to 14 layers yields 5.75% macro error while reducing audio-tower parameters by 20.7% (186.4M147.8M), with most of the additional loss concentrated on a few English benchmarks.
The benchmark variation also shows why target depth alone is not enough: layer combinations differ in recoverability, and the pruned encoder must be realigned with the decoder under both teacher-forced and student-generated contexts. Figure 2 connects the main steps, from transcript-consistency filtering and behavioral probes to progressive pruning and three-stage recovery. The reported configuration uses class 1 data in all three stages, with source reweighting during Stage 2.
We introduce a progressive recovery framework that combines behavioral candidate screening, cross-scale hidden-state and logit supervision, scheduled student-policy training, and LoRA finetuning while preserving the pretrained decoder backbone.
A matched teacher-scale comparison reduces mean error from 8.45% with self-distillation to 5.55% with the 1.7B teacher, demonstrating the practical value of cross-scale supervision in the tested setting.
Layer-pair probes expose non-additive interactions: the pair combines the two strongest single removals but underperforms by 0.85 pp after matched recovery.
The 14-layer model reaches 5.75% macro error, compared with 5.61% for the baseline, while removing 20.7% of audio-tower parameters and reducing measured encoder latency by 21.4% on the in-vehicle accelerator and 11.4% on H800.
2 Related Work
2.1 Speech Architectures and Encoder Compression
Modern speech systems increasingly couple an acoustic encoder with a pretrained language model. Qwen3-ASR [1] inserts bridged audio representations into the Qwen3 decoder input, while SLAM-ASR [2] examines lightweight connectors between speech encoders and LLMs. Qwen2-Audio [11] similarly integrates acoustic representations with a general-purpose language model. Whisper [3] follows an encoder–decoder architecture rather than an LLM-connector design, but remains an important reference for multilingual ASR and subsequent compression work. In each case, the encoder processes the full acoustic sequence, making its depth a direct contributor to computation and latency.
ASR compression has targeted different parts of this architecture. Distil-Whisper [6] primarily reduces the decoder while retaining the Whisper encoder. LiteASR [7] combines low-rank factorization with distillation for encoder matrices, whereas structured sparsity removes weights or attention heads [8]. LayerDrop [12] trains networks to tolerate variable depth, and Dynamic Encoder Size [9] learns a supernet from which multiple encoder depths can be extracted. These approaches either introduce compression during training or reduce computation within existing blocks. Direct removal of complete blocks offers a regular architecture, but it also changes the representations consumed by downstream modules.
Kolluri et al. [10] study Whisper encoder-layer pruning in an LLM-based SLAM-ASR model and recover the pruned network with LoRA [13]. X-AuT instead starts from a pretrained Qwen3-ASR audio tower, removes layers in successive hops, and uses short post-removal probes to compare candidate layer sets under a shared recovery budget. This design treats recoverability as a property of a layer combination rather than an additive score assigned to individual layers.
2.2 Distillation and Data Selection
Knowledge distillation can transfer teacher output distributions, intermediate representations, or both. Teacher-forced logit distillation evaluates the teacher and student under gold-prefix contexts, which is efficient but differs from inference once the student conditions on its own predictions. On-policy distillation instead supplies supervision on student-generated histories [14]. Ark-ASR [15] develops data-efficient on-policy distillation for ASR, and ASKD-Whisper [16] adjusts the strength of self-distillation during training. X-AuT combines the two regimes: teacher-forced supervision remains the default, while scheduled student-policy batches expose the teacher to the student’s decoding context after an initial stabilization period. Rollout filtering returns degenerate batches to the teacher-forced objective.
Training data also shape recovery. Curriculum and data-selection methods rank, order, or filter examples according to estimated learning value or label reliability [17]. Our preprocessing compares the source transcript with hypotheses from two external ASR systems and assigns a transcript-consistency tier from their pairwise agreement. The reported experiments use the highest-agreement tier throughout recovery and alter source weights during Stage 2 to emphasize target-domain data. The resulting pipeline combines confidence-based filtering with phase-specific resampling without relying on a staged mixture of progressively noisier tiers.
3 Method
X-AuT filters training examples by transcript agreement, identifies recoverable layer combinations through short behavioral probes, and restores each progressively pruned model with a three-stage training schedule (Figure 2).
3.1 Architecture and Objective
An input waveform is encoded by an -layer audio encoder and bridge into audio embeddings . Qwen3-ASR places these embeddings at audio-placeholder positions in the token embedding sequence; the causal language-model decoder then predicts transcription tokens through a tied output projection . This is input-embedding conditioning rather than a separate decoder cross-attention module.
A pruning operation retains an ordered subset and forms from those blocks. We seek a recoverable subset and parameters that minimize aggregate text error rate (TER) under a target depth:
| (1) |
The pretrained weights of remain frozen. LoRA parameters attached to its q/k/v/o attention projections are trainable, and —which shares weights with the token embedding in Qwen3-ASR—is trainable during Stages 0–1 and frozen during Stage 2.
Removing encoder blocks changes the conditioning embeddings presented to the decoder. The recovery objective therefore first aligns intermediate and bridge representations, then adapts token distributions under teacher-forced and student-generated prefixes.
3.2 Transcript-Consistency Filtering
The source pool contains heterogeneous supervision. For each reference-bearing utterance, two strong ASR systems produce offline hypotheses. After language-aware normalization, we compute the three pairwise edit rates among the source transcript and the two hypotheses, using CER for Chinese and WER for English. Their maximum, , measures the largest disagreement within the transcript–hypothesis triplet. Exact agreement, Mandarin homophone agreement, and a consistency vote assign one of the nine tiers in Table 1. Lower tier numbers indicate stronger transcript agreement rather than ground-truth quality.
| Tier | Agreement band | Assignment rule |
|---|---|---|
| 1 | Full agreement | Both model hypotheses exactly match the source transcript. |
| 2 | Partial exact | At least one of the three text pairs matches exactly (i.e., at least two candidates agree). |
| 3 | Homophone | For Mandarin, at least one source–hypothesis pair has zero pinyin CER. |
| 4 | The consistency vote passes and falls in the indicated interval. | |
| 5 | Same voting rule with a small transcript discrepancy. | |
| 6 | Same voting rule with a moderate discrepancy. | |
| 7 | Same voting rule with a relatively large discrepancy. | |
| 8 | Same voting rule with a large discrepancy. | |
| 9 | Failure / | The vote fails, , or an aligned hypothesis is empty. |
The reported configuration uses class 1 for both distillation stages. Stage 2 retains the same consistency threshold but reweights sources toward cockpit and other target-domain data, separating confidence-based filtering from phase-specific source sampling.
3.3 Behavior-Driven Progressive Pruning
We prune in two hops, 181614. The first hop removes original layers . For the second hop, every candidate starts from the same recovered 16-layer checkpoint and is trained with the same 0.3-epoch LoRA warm-up. We first evaluate each remaining layer as a single removal, then evaluate a fixed set of adjacent and non-adjacent layer pairs. Candidate selection uses the lowest aggregate TER on a fixed five-benchmark development suite (Sec. 4). This procedure is more expensive than a static score but directly measures post-removal behavior under the available recovery budget.
The pair probes are necessary because recovery after removing several layers cannot be predicted reliably from the corresponding single-layer scores. The matched comparison selects for the 1614 hop; Section 5.4 reports the candidate-level results.
3.4 Three-Stage Recovery
Each hop uses the same three-stage recipe. The student is the pruned Qwen3-ASR-0.6B model; the teacher is Qwen3-ASR-1.7B with a 24-layer audio encoder. Teacher parameters are frozen and discarded at inference.
3.4.1 Stage 0: Representation Alignment
Stage 0 occupies the first 5% of the distillation epoch and combines intermediate-layer, bridge, logit, and transcript losses:
| (2) |
Both representation losses sum mean-squared error and cosine distance. Teacher layers are divided uniformly into ordered groups, and student layer aligns to the last teacher layer in group . Because the teacher and student hidden widths are 2048 and 1024, respectively, a learned two-layer MLP with a 256-dimensional bottleneck projects teacher hidden and bridge features into the student space. Logit KD uses temperature-scaled KL divergence under gold prefixes.
The pruned audio encoder and bridge are fully trainable. Decoder base weights remain frozen, while rank-32 LoRA adapters on q/k/v/o projections and the tied output embedding are trained with a separate decoder learning rate.
3.4.2 Stage 1: Distillation with Scheduled Student-Policy Contexts
Stage 1 disables intermediate-layer loss and uses bridge alignment, teacher-forced logit KD, and gold-transcript CE:
| (3) |
After 20% of Stage 1 has elapsed, every fifth optimizer step is scheduled for student-policy supervision. The student greedily generates a prefix; student and teacher are then evaluated on the same generated context. Their distributions are compared over the union of each model’s top- support (). Gold CE remains an anchor on these scheduled steps. The implementation uses no confidence reweighting (weight_mode=none).
Rollout safeguards prevent degenerate prefixes from entering the KD loss. Generation enforces min_new_tokens, uses a duration-aware maximum capped at 256 tokens, and rejects budget-exhausted, over-long, or repetitive rollouts. If at least half of a batch is rejected, that microbatch falls back to the teacher-forced objective. Thus, approximately 20% is a scheduling target; the realized on-policy fraction can be lower after filtering.
3.4.3 Stage 2: LoRA Finetuning
Stage 2 initializes from the best Stage 1 checkpoint and optimizes gold-transcript CE for one epoch. The audio encoder, bridge, and decoder LoRA adapters remain trainable at ; the tied lm_head/embedding is frozen. The class-1 data index is reweighted toward target-domain sources. This stage contains no teacher loss:
| (4) |
4 Experimental Setup
4.1 Models and Parameter Accounting
The student starts from Qwen3-ASR-0.6B [1], whose audio tower contains 18 Transformer blocks and a convolutional/bridge frontend. Counting tensors in the released checkpoint gives 186.376M audio-tower parameters: 12.758M outside the Transformer stack and 9.645M per block. The 16- and 14-layer students therefore contain 167.085M and 147.794M audio-tower parameters, corresponding to 10.35% and 20.70% reductions. We report these exact counts rather than inferring compression from rounded labels such as 180M/140M. The cross-scale teacher is Qwen3-ASR-1.7B with a 24-layer audio tower and 2048-dimensional hidden states; the student audio hidden width is 1024.
4.2 Training Data and Quality Labels
The source pool combines public and proprietary multilingual ASR corpora, including AISHELL-1/4/5 [18, 19, 20], CommonVoice [21], Emilia [22], GigaSpeech [23], KeSpeech [24], LibriSpeech [25], WenetSpeech [26], and cockpit-domain speech. The pool exceeds 280k hours before quality filtering. Audio is capped at 40 seconds.
Manifest construction.
Each corpus is converted to a unified JSONL manifest containing an utterance identifier, audio reference, source transcript, language and split tags, duration, sampling rate, and channel count. Text normalization includes Unicode NFKC normalization, width conversion, Traditional-to-Simplified conversion for Mandarin, removal of invisible characters and numeric separators, dash canonicalization, and whitespace normalization. Original and normalized text are retained for traceability. Figure 3 summarizes the complete path from corpus ingestion through transcript agreement and quality ranking.
Reference-bearing utterances are decoded offline by Qwen3-ASR-1.7B and Qwen3.5-Omni [27]. Records without an inference result, a valid audio reference, or nonempty supervision are excluded. The remaining records receive the consistency labels described in Sec. 3.2. The reported distillation runs use the class-1 manifest index; the first-hop log contains approximately 299k weighted target records per epoch and the second-hop log approximately 292k. Stage 2 keeps class 1 but changes corpus weights, increasing AISHELL-4/5 and cockpit-query contributions while dropping several weakly matched web-speech sources. These counts describe the realized loader indices rather than the size of the 280k-hour source pool.
4.3 Selection and Evaluation Suites
Development selection suite.
Behavior probes and checkpoint selection use five fixed development or validation subsets: AISHELL-1 (Mandarin CER), CommonVoice-en (English WER), Fleurs-en (English WER), WenetSpeech-meeting (Mandarin CER), and a proprietary cockpit-query subset (Mandarin CER). Frequent evaluation is capped at 25 utterances per benchmark, and the macro average of the five normalized error rates determines checkpoint selection. The final results are evaluated separately on the full public benchmark suite described below.
Full public evaluation suite.
Final checkpoint results are reported on ten public benchmarks: AISHELL-1 (CER) [18]; Fleurs zh/en (CER/WER) [28]; LibriSpeech test-clean/test-other (WER) [25]; THCHS-30 (CER) [29]; Tedlium (WER) [30]; CommonVoice v15 zh/en (CER/WER) [21]; and WenetSpeech-meeting (CER) [26]. The macro mean weights benchmarks equally, not by utterance count. The proprietary selection subset is excluded from all main-result tables.
4.4 Optimization and Reporting Protocol
The reported 0.6B runs use 32 accelerators, per-device batch size 8, gradient accumulation 2, and global batch size 512. We use AdamW with for the audio tower and for decoder-side LoRA plus the tied output embedding during distillation. Stage 2 uses for all trainable parameters and freezes the tied output embedding. Weight decay is 0.01, gradient clipping is 1.0, and training uses bf16. Stage 0 and Stage 1 occupy 0.05 and 0.95 epoch; Stage 2 runs for one epoch. All configurations use seed 42.
Checkpoints are selected by the lowest observed selection-suite macro TER. Tables report the corresponding single-run checkpoint on the full suite. We did not run repeated seeds or bootstrap utterance-level confidence intervals, so boldface denotes the best observed number in a row and not statistical significance. Full hyperparameters are listed in Appendix A.1.
Efficiency is measured separately on an in-vehicle PPU and an NVIDIA H800 GPU using more than 50 utterances of varying duration. We report descriptive averages from the available benchmark output; run-to-run variability was not retained.
5 Results
5.1 Main Results
| Benchmark | Base (18L) | Prune-16 S1 | Prune-16 S2 |
|---|---|---|---|
| AuT parameters | 186.4M | 167.1M | 167.1M |
| AISHELL-1 (CER) | 3.33% | 3.30% | 3.21% |
| Fleurs-zh (CER) | 2.80% | 3.35% | 3.28% |
| Fleurs-en (WER) | 4.17% | 4.28% | 4.23% |
| LibriSpeech test-clean (WER) | 2.48% | 2.73% | 2.65% |
| THCHS-30 (CER) | 3.87% | 4.10% | 4.06% |
| Tedlium (WER) | 3.35% | 3.92% | 3.79% |
| LibriSpeech test-other (WER) | 5.39% | 5.90% | 5.79% |
| CommonVoice v15 zh (CER) | 9.95% | 8.56% | 8.12% |
| CommonVoice v15 en (WER) | 12.35% | 10.74% | 10.50% |
| WenetSpeech-meeting (CER) | 8.36% | 8.62% | 7.06% |
| Macro mean (%) | 5.61 | 5.55 | 5.27 |
| Relative mean-error change (%) | – |
The 16-layer model lowers macro-average error from 5.61% to 5.27%, an absolute change of pp and a 6.1% relative error reduction (Table 2). It improves four benchmarks: AISHELL-1, CommonVoice zh/en, and WenetSpeech-meeting. The largest gains occur on CommonVoice zh ( pp), CommonVoice en ( pp), and WenetSpeech-meeting ( pp), while the largest degradation is 0.44 pp on Tedlium. Stage 2 improves all ten entries relative to the Stage 1 checkpoint.
| Benchmark | Base (18L) | Prune-14 S1 | Prune-14 S2 |
|---|---|---|---|
| AuT parameters | 186.4M | 147.8M | 147.8M |
| AISHELL-1 (CER) | 3.33% | 3.52% | 3.39% |
| Fleurs-zh (CER) | 2.80% | 3.49% | 3.32% |
| Fleurs-en (WER) | 4.17% | 5.24% | 5.10% |
| LibriSpeech test-clean (WER) | 2.48% | 3.09% | 2.45% |
| THCHS-30 (CER) | 3.87% | 4.23% | 4.17% |
| Tedlium (WER) | 3.35% | 4.07% | 3.95% |
| LibriSpeech test-other (WER) | 5.39% | 7.00% | 5.52% |
| CommonVoice v15 zh (CER) | 9.95% | 9.54% | 8.36% |
| CommonVoice v15 en (WER) | 12.35% | 13.94% | 12.49% |
| WenetSpeech-meeting (CER) | 8.36% | 10.71% | 8.78% |
| Macro mean (%) | 5.61 | 6.48 | 5.75 |
| Relative mean-error change (%) | – |
The 14-layer model contains 147.794M audio-tower parameters, 20.70% fewer than the 186.376M baseline. Its macro error is 5.75%, a 0.14-pp increase over the baseline (Table 3). CommonVoice zh improves by 1.59 pp and LibriSpeech test-clean by 0.03 pp, whereas Fleurs-en has the largest degradation at 0.93 pp. Relative to the 16-layer model, seven benchmark changes remain within 0.3 pp; CommonVoice en ( pp) and WenetSpeech-meeting ( pp) account for most of the macro-average gap. Figure 1 summarizes this benchmark-level variation, and the tables report the corresponding CER/WER values.
The two operating points expose a clear tradeoff. The 16-layer model improves the observed macro average, while the 14-layer model provides a larger structural reduction at a small average cost. Section 6 discusses the uncertainty associated with these single-run comparisons.
5.2 Training Trajectories
Stage 0 TER falls from 10.88% at step 500 to 7.80% at step 2500, followed by 8.38% at the first Stage 1 evaluation near step 4000 (Figure 4). Stage 1 reaches its minimum of 5.76% at step 47,000. Starting from that checkpoint, Stage 2 reduces TER by another 0.40 pp and reaches 5.36% at step 28,000.
The trajectory shows rapid early recovery followed by slower optimization in Stage 1 and a further gain from Stage 2. Because the stage transition changes both the loss and optimizer state, the curve describes the recovery process rather than the contribution of an individual loss term.
5.3 Teacher Scale
A same-scale control replaces the 1.7B teacher with the student’s unpruned 0.6B model while retaining the Stage 0/1 schedule, data, LoRA configuration, and optimization hyperparameters. The cross-scale case additionally requires the learned 20481024 teacher projections. At the Stage 1 best checkpoints, mean full-suite error is 5.55% for the cross-scale teacher and 8.45% for the self-teacher; the cross-scale model is better on all ten benchmarks (Appendix A.3). Relative to the unpruned baseline, the cross-scale checkpoint improves macro error by 1.1%, while the self-teacher checkpoint increases it by 50.6%.
The large gap demonstrates the practical value of the stronger teacher under the implemented recipe. On CommonVoice zh/en, the cross-scale checkpoint also improves over the original 0.6B baseline, suggesting that the larger teacher transfers useful acoustic behavior rather than merely restoring the pruned student to its starting point. The scope of this interpretation is discussed in Section 6.
5.4 Behavior-Driven Layer Selection
The single-layer sweep spans 6.29–8.90% TER. L6 is the strongest single removal at 6.29%, followed by L5 at 6.42% and L8 at 6.52%. The selected pair reaches 6.93%, whereas reaches 7.78%. Defining the interaction penalty as pair TER minus the mean TER of its constituent removals gives 0.58 pp for and 1.38 pp for .
The adjacent candidates , , and rank ahead of the tested non-adjacent candidates and , while the adjacent pair performs worst overall. Thus, adjacency alone is not a reliable selection rule, and pair recoverability cannot be inferred from single-layer scores in isolation.
6 Discussion
6.1 Interpretation
Taken together, the results suggest that successful depth reduction depends on both representation recovery and retained encoder capacity. Removing layers perturbs the audio embeddings presented to the decoder, as reflected in the early TER trajectory and the EOS diagnostics. Stage 0 directly reduces this mismatch at the intermediate and bridge levels, and Stage 1 extends recovery to token distributions under gold and student-generated prefixes. This procedure is sufficient for the 16-layer model to surpass the baseline macro error, but the remaining gap at 14 layers indicates that alignment cannot fully replace the capacity lost through deeper pruning.
The teacher-scale comparison further shows that the source of the recovery signal matters. Under the matched 16-layer recipe, the 1.7B teacher yields 5.55% mean error, compared with 8.45% for self-distillation from the unpruned 0.6B model. The improvement spans all ten benchmarks and is particularly pronounced on CommonVoice zh/en. We interpret this result as useful transfer from the larger teacher within the present training recipe, while recognizing that the comparison does not separate pruning recovery from gains that the same recipe might provide to an unpruned student.
Layer choice and pruning schedule also affect recoverability. Although L8 and L6 are the two strongest single-layer removals, pruning them together performs worse than the selected pair after matched recovery. The pair probes therefore provide information that cannot be inferred by ranking individual layers alone. Similarly, progressive 181614 pruning reaches 5.75% mean error, whereas direct 1814 pruning reaches 6.73% under the same nominal budget. These comparisons favor explicit pair evaluation and progressive pruning for the model and candidates studied here.
6.2 Limitations
The primary results are single runs with seed 42, without repeated-seed variation, paired utterance-level confidence intervals, or significance tests. Checkpoint selection also relies on fixed development subsets capped at 25 utterances per benchmark. This design makes frequent evaluation tractable, but it introduces selection noise and leaves small differences, including the 0.14-pp gap between the 14-layer model and the baseline, as descriptive observations.
The study covers one model family and a limited set of pruning candidates. Layer interactions, projection-based alignment, and EOS behavior may differ in Whisper [3], Qwen2-Audio [11], or other speech–language architectures. The data pipeline likewise supports nine consistency classes, whereas the reported runs use class 1 throughout and change only source weights in Stage 2. Experiments across model families, broader layer combinations, and controlled data mixtures would clarify which parts of the recipe generalize beyond the present setting.
Only the audio tower is compressed, so autoregressive decoding limits the reduction in end-to-end latency. The efficiency measurements are averages from one retained benchmark output and do not include run-to-run uncertainty. A complete deployment study should repeat the hardware measurements under a fixed software stack and report latency distributions alongside average values.
7 Conclusion
X-AuT reduces the Qwen3-ASR-0.6B audio tower from 18 to 14 layers through progressive pruning and cross-scale recovery. The 16-layer model improves macro-average error from 5.61% to 5.27%, while the 14-layer model reaches 5.75% with 20.7% fewer audio-tower parameters. The teacher-scale, pair-selection, and direct-pruning controls show that recovery depends on teacher strength, joint layer selection, and pruning schedule. Because the results come from single runs on one model family, broader models and repeated trials are needed to determine how well this tradeoff generalizes.
References
- [1] X. Shi, X. Wang, Z. Guo, Y. Wang, P. Zhang, X. Zhang, Z. Guo, H. Hao, Y. Xi, B. Yang, J. Xu, J. Zhou, and J. Lin (2026) Qwen3-asr technical report. External Links: 2601.21337 Cited by: §1, §1, §2.1, §4.1.
- [2] Z. Ma, G. Yang, Y. Yang, Z. Gao, J. Wang, Z. Du, F. Yu, Q. Chen, S. Zheng, S. Zhang, and X. Chen (2024) An embarrassingly simple approach for llm with strong asr capacity. External Links: 2402.08846 Cited by: §1, §2.1.
- [3] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever (2022) Robust speech recognition via large-scale weak supervision. arXiv preprint arXiv:2212.04356. Cited by: §1, §2.1, §6.2.
- [4] Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang, Q. Liang, D. Bhatia, Y. Shangguan, B. Li, G. Pandey, K. C. Sim, T. Bagby, S. Chang, K. Rao, and A. Gruenstein (2019) Streaming end-to-end speech recognition for mobile devices. In Proc. IEEE ICASSP, pp. 6381–6385. Cited by: §1.
- [5] S. Han, H. Mao, and W. J. Dally (2016) Deep compression: compressing deep neural networks with pruning, trained quantization and huffman coding. In International Conference on Learning Representations (ICLR), Cited by: §1.
- [6] S. Gandhi, P. von Platen, and A. M. Rush (2023) Distil-whisper: robust knowledge distillation via large-scale pseudo labelling. In arXiv preprint arXiv:2311.00430, Cited by: §1, §2.1.
- [7] K. Kamahori, J. Kasai, N. Kojima, and B. Kasikci (2025) LiteASR: efficient automatic speech recognition with low-rank approximation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), Cited by: §1, §2.1.
- [8] P. K. Mudi, A. Sachan, D. Devapriya, and S. Kalyani (2025) Structured sparsity and weight-adaptive pruning for memory and compute efficient whisper models. External Links: 2510.12666 Cited by: §1, §2.1.
- [9] J. Xu, E. Beck, Z. Yang, and R. Schlüter (2024) Dynamic encoder size based on data-driven layer-wise pruning for speech recognition. In Proc. Interspeech, pp. 4563–4567. Cited by: §1, §2.1.
- [10] G. P. K. B. Kolluri, M. Kampouridis, and R. Shekhar (2026) On the role of encoder depth: pruning whisper and lora fine-tuning in slam-asr. In Proceedings of the SPEAKABLE Workshop, LREC, Cited by: §1, §2.1.
- [11] Qwen Team (2024) Qwen2-audio technical report. External Links: 2407.10759 Cited by: §2.1, §6.2.
- [12] A. Fan, E. Grave, and A. Joulin (2020) Reducing transformer depth on demand with structured dropout. In International Conference on Learning Representations (ICLR), Cited by: §2.1.
- [13] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2022) LoRA: low-rank adaptation of large language models. In International Conference on Learning Representations (ICLR), Cited by: §2.1.
- [14] M. Song and M. Zheng (2026) A survey of on-policy distillation for large language models. External Links: 2604.00626 Cited by: §2.2.
- [15] Y. Lin, Y. Wang, R. Cai, and X. Zeng (2026) Data-efficient on-policy distillation for automatic speech recognition. External Links: 2605.28139 Cited by: §2.2.
- [16] J. Lee, N. Kim, S. Lee, and C. Chun (2026) ASKD-whisper: adaptive self-knowledge distillation for efficient and low-latency automatic speech recognition. External Links: 2601.19919 Cited by: §2.2.
- [17] S. Rampp, M. Milling, A. Triantafyllopoulos, and B. W. Schuller (2024) Does the definition of difficulty matter? scoring functions and their role for curriculum learning. External Links: 2411.00973 Cited by: §2.2.
- [18] H. Bu, J. Du, X. Na, Y. Ji, and F. Zheng (2017) Aishell-1: an open-source mandarin speech corpus and a speech recognition baseline. In Proc. O-COCOSDA, Cited by: §4.2, §4.3.
- [19] Y. Fu, L. Xu, Y. Zhai, Y. Wang, K. Liang, Y. Li, S. Zhang, Y. Yan, and X. Chen (2021) AISHELL-4: an open source dataset for speech recognition in multi-party conference scenario. External Links: 2104.03035 Cited by: §4.2.
- [20] Y. Shi, L. Xu, S. Zhang, and Y. Yan (2023) AISHELL-5: multi-domain mandarin speech recognition corpus with labeled data of 520 hours. In Proc. IEEE ICASSP, Cited by: §4.2.
- [21] R. Ardila, M. Branson, K. Lee, M. Kohler, R. Schumann, L. Sterckx, J. D. Bayron, P. Karunanayake, R. Sanabria, A. Baas, et al. (2020) Common voice: a massively-multilingual speech corpus. In Proc. LREC, Cited by: §4.2, §4.3.
- [22] H. He, Z. Shang, C. Wang, X. Li, Y. Gu, P. Hua, L. Liu, C. Yang, J. Li, P. Shi, Y. Wang, K. Chen, and Z. Wu (2024) Emilia: an extensive, multilingual, and diverse speech dataset for large-scale speech generation. External Links: 2407.05361 Cited by: §4.2.
- [23] G. Chen, S. Chai, G. Wang, J. Du, W. Zhang, C. Weng, D. Lu, D. Povey, J. Trmal, J. Zhang, M. Jin, S. Khudanpur, S. Wang, S. Wu, Y. Yang, Y. Wang, Z. Yu, and Z. Wang (2021) GigaSpeech: an evolving, multi-domain asr corpus with 10,000 hours of transcribed audio. In Proc. Interspeech, Cited by: §4.2.
- [24] Z. Tang, D. Wang, X. Xu, H. Zheng, Y. Lei, J. Li, S. Zhang, Y. Zhu, J. Meng, H. Li, X. Xu, Y. Zheng, and S. Li (2021) KeSpeech: an open source speech dataset of mandarin and its eight subdialects. In Proc. IEEE ASRU, Cited by: §4.2.
- [25] V. Panayotov, G. Chen, D. Povey, and S. Khudanpur (2015) LibriSpeech: an asr corpus based on public domain audio books. In Proc. IEEE ICASSP, Cited by: §4.2, §4.3.
- [26] B. Zhang, H. Lv, P. Guo, Q. Shao, C. Yang, L. Xie, X. Xu, H. Bu, X. Chen, C. Zeng, D. Wu, and Z. Peng (2022) WenetSpeech: a 10,000+ hours multi-domain mandarin corpus for asr. In Proc. IEEE ICASSP, Cited by: §4.2, §4.3.
- [27] Qwen Team (2026) Qwen3.5-omni technical report. External Links: 2604.15804 Cited by: §4.2.
- [28] A. Conneau, A. Bapna, Y. Zhang, M. Ma, P. von Platen, A. Lozhkov, C. Cherry, Y. Jia, C. Rivera, M. Kale, N. Remez, V. Glavcev, S. Gopala, J. Ni, Y. Wu, P. Hsu, J. Liu, A. Sahebi, P. Duquenne, M. Chen, V. Chau, et al. (2022) FLEURS: few-shot learning evaluation of universal representations of speech. In Proc. IEEE SLT, Cited by: §4.3.
- [29] D. Wang and X. Zhang (2015) THCHS-30: a free chinese speech corpus. External Links: 1512.01882 Cited by: §4.3.
- [30] A. Rousseau, P. Deléglise, and Y. Estève (2012) TED-lium: an automatic speech recognition dedicated corpus. In Proc. LREC, Cited by: §4.3.
Appendix
A.1 Full Hyperparameter Configuration
| Parameter | Stage 0 | Stage 1 | Stage 2 |
|---|---|---|---|
| Fraction / epochs | 0.05 epoch | 0.95 epoch | 1 epoch |
| LR (audio tower) | |||
| LR (LoRA / tied head) | /frozen | ||
| Optimizer / weight decay | AdamW / 0.01 | AdamW / 0.01 | AdamW / 0.01 |
| Per-device batch / accumulation | 8 / 2 | 8 / 2 | 8 / 2 |
| Global batch size | 512 | 512 | 512 |
| LoRA /dropout | 32 / 64 / 0.05 | 32 / 64 / 0.05 | 32 / 64 / 0.05 |
| Temperature | 1.5 | 1.0 | – |
| 1.0 | 0.0 | – | |
| 1.0 | 0.5 | – | |
| 0.2 | 0.1 | – | |
| 0.5 | 1.0 | 1.0 | |
| On-policy start / fraction | – | 0.2 / 0.2 | – |
| Union top- / weight mode | – | 512 / none | – |
| Minimum / maximum new tokens | – | 3 / 256 | – |
| Reject-batch threshold | – | 0.5 | – |
A.2 Premature-EOS Safeguards
Immediately after audio-encoder pruning, the model can emit EOS after only a few tokens, producing large deletion errors. Qwen3-ASR does not use a separate decoder cross-attention module: bridge outputs replace audio-placeholder embeddings in the causal decoder input. We hypothesize that pruning shifts these conditioning embeddings, weakening acoustic evidence for the pretrained decoder. We do not directly measure embedding mean/variance or EOS causality, so this account motivates the safeguards rather than constituting a mechanistic proof.
The implementation uses three safeguards. First, during main Stage 0/1 distillation, the tied lm_head/token embedding is trainable at the decoder-side learning rate (; in the EOS ablation). Decoder Transformer weights remain frozen apart from LoRA. Second, on-policy generation enforces a fixed min_new_tokens; its maximum is duration-aware, , where is the longest audio duration in the batch. Third, rollout filters reject budget-exhausted, implausibly long, and repetitive sequences. If the rejected fraction reaches 0.5, the scheduled on-policy microbatch falls back to teacher-forced distillation.
| Configuration | lm_head | min_new | Empty events | TER | Rejections |
|---|---|---|---|---|---|
| A: neither | frozen | 0 | 5 | 6.86 | 20 |
| B: lm_head | trainable | 0 | 3 | 6.83 | 11 |
| C: gating | frozen | 3 | 0 | 6.97 | 2113 |
| D: both | trainable | 3 | 0 | 6.75 | 2162 |
The no-mitigation condition has five windows with a nonzero empty ratio; the maximum is 0.3% and the mean is 0.018%. Training the tied output embedding reduces this to three events, while min_new_tokens removes empty events in both gating conditions. Gating also converts some premature terminations into longer or degenerate rollouts: configurations C and D trigger 2113 and 2162 filter rejections, compared with 20 and 11 for A and B.
The best observed TER is 6.75% for the combined configuration D. The tied-head-only configuration B reaches 6.83% and shows only 0.04-pp best-to-last drift; A drifts from 6.86% to 7.58%. Mean rollout length is 15.06 tokens for D and 11.05 for A. These single runs support using the tied head and gating together in the main recipe, but the 0.08-pp B–D difference is too small to interpret without uncertainty estimates.
A.3 Teacher-Scale Comparison
The self-teacher control uses the unpruned 18-layer Qwen3-ASR-0.6B model. The cross-scale condition uses Qwen3-ASR-1.7B and the required 20481024 bottleneck projections. Other Stage 0/1 settings are matched, and both models are evaluated at their best development checkpoint before Stage 2.
| Benchmark | Base (18L) | Self-teacher | Cross-scale |
|---|---|---|---|
| AISHELL-1 (CER) | 3.33% | 4.29% | 3.30% |
| Fleurs-zh (CER) | 2.80% | 4.17% | 3.35% |
| Fleurs-en (WER) | 4.17% | 6.15% | 4.28% |
| LibriSpeech test-clean (WER) | 2.48% | 6.00% | 2.73% |
| THCHS-30 (CER) | 3.87% | 5.23% | 4.10% |
| Tedlium (WER) | 3.35% | 11.28% | 3.92% |
| LibriSpeech test-other (WER) | 5.39% | 9.89% | 5.90% |
| CommonVoice v15 zh (CER) | 9.95% | 11.71% | 8.56% |
| CommonVoice v15 en (WER) | 12.35% | 14.89% | 10.74% |
| WenetSpeech-meeting (CER) | 8.36% | 10.85% | 8.62% |
| Macro mean (%) | 5.61 | 8.45 | 5.55 |
| Relative mean-error change (%) | – |
The cross-scale checkpoint is better on every public benchmark and lowers mean error from 8.45% to 5.55%. The largest gaps occur on Tedlium (11.28 vs. 3.92), LibriSpeech test-clean (6.00 vs. 2.73), and LibriSpeech test-other (9.89 vs. 5.90). Relative to the original baseline, the cross-scale model also improves CommonVoice zh/en and AISHELL-1, whereas the self-teacher is worse on all ten benchmarks. This pattern is consistent with useful cross-scale transfer, especially on English and higher-error conditions. Because the comparison has one seed and includes projection modules only when dimensions differ, we do not treat it as proof that the teacher creates capabilities absent from every unpruned student.
A.4 Layer-Interaction Details
| Candidate | Type | TER | AISHELL | CV-en | Fleurs-en | Wenet / SC | |
|---|---|---|---|---|---|---|---|
| Adj. | 6.93 | – | 0.63 | 16.97 | 6.85 | 7.30 / 2.91 | |
| Adj. | 7.37 | +0.44 | 0.63 | 16.06 | 6.48 | 8.47 / 5.23 | |
| Adj. | 7.75 | +0.82 | 0.63 | 23.39 | 5.37 | 6.42 / 2.91 | |
| Non-adj. | 7.78 | +0.85 | 0.63 | 20.64 | 6.11 | 8.03 / 3.49 | |
| Non-adj. | 8.12 | +1.19 | 0.63 | 17.89 | 6.48 | 8.03 / 7.56 | |
| Adj. | 9.93 | +3.00 | 0.95 | 23.39 | 11.48 | 8.61 / 5.23 |
The single-layer results used to form the interaction comparison are L8=5.97%, L6=6.11%, and L5=6.15%. Thus, the non-adjacent pair contains the two best constituent removals (mean 6.04%) but reaches 7.78% as a pair. The selected adjacent pair has a slightly worse constituent mean (6.13%) but reaches 6.93%. The descriptive interaction penalties are therefore 1.74 and 0.80 pp, respectively.
All three adjacent candidates in Table 7 rank above the two non-adjacent candidates, while confirms that adjacency alone is not sufficient. One possible account is that removing a contiguous sub-block creates one residual-stream discontinuity whereas dispersed removal creates two. This explanation is untested; causal activation analysis and a larger factorial candidate set are needed.
A.5 Progressive versus Direct Pruning
We compare the reported progressive path with a direct 1814 run that drops the same original layers and uses the same nominal recovery and data budget.
| Benchmark | Base | Direct 1814 | Progressive |
|---|---|---|---|
| AISHELL-1 (CER) | 3.33% | 3.81% | 3.39% |
| Fleurs-zh (CER) | 2.80% | 3.76% | 3.32% |
| Fleurs-en (WER) | 4.17% | 5.94% | 5.10% |
| LibriSpeech test-clean (WER) | 2.48% | 3.42% | 2.45% |
| THCHS-30 (CER) | 3.87% | 4.62% | 4.17% |
| Tedlium (WER) | 3.35% | 5.10% | 3.95% |
| LibriSpeech test-other (WER) | 5.39% | 6.31% | 5.52% |
| CommonVoice v15 zh (CER) | 9.95% | 9.79% | 8.36% |
| CommonVoice v15 en (WER) | 12.35% | 15.22% | 12.49% |
| WenetSpeech-meeting (CER) | 8.36% | 9.29% | 8.78% |
| Macro mean (%) | 5.61 | 6.73 | 5.75 |
Progressive pruning reaches 5.75% mean error, compared with 6.73% for direct pruning, and is better on all ten benchmarks. The largest direct-minus-progressive gaps are CommonVoice en (+2.73 pp), CommonVoice zh (+1.43 pp), and Tedlium (+1.15 pp). This shows that the progressive path is preferable under the tested budget. It does not establish that no alternative schedule or larger budget could improve direct pruning.
A.6 On-Policy Strategy Comparison
Figure 7 compares three Stage 1 runs on prune-16: a plain off-policy baseline, an off-policy control matched to the extra student forward and loss settings used on scheduled steps, and hybrid on-policy training. The dashboard provides the visual trajectory; numerical comparisons use the archived checkpoint logs.
The best observed TER values are 6.98% for plain off-policy, 6.23% for matched off-policy, and 6.41% for hybrid on-policy. Final observations are 7.43%, 6.51%, and 6.99%, respectively. Hybrid training therefore improves over the plain baseline but does not outperform the matched off-policy control in this experiment. We retain the hybrid recipe in the main run because it directly supervises student-generated contexts and works well in the complete pipeline, while recognizing that this ablation does not establish accuracy superiority. More seeds and a sweep over rollout fraction are needed.
A.7 Inference Efficiency
| Model | Encoder (ms) | End-to-end (ms) | RTF | Peak memory (MB) |
| In-vehicle PPU | ||||
| Baseline 18L | 14 | 486 | 0.0786 | 1500.3 |
| Prune-14 | 11 | 463 | 0.0748 | 1434.4 |
| Relative change | ||||
| NVIDIA H800 | ||||
| Baseline 18L | 88 | 1672 | 0.256 | 1500.3 |
| Prune-14 | 78 | 1629 | 0.245 | 1458.2 |
| Relative change | ||||
The 14-layer model reduces encoder time by 21.4% on the in-vehicle PPU and 11.4% on H800. End-to-end reductions are 4.7% and 2.6%, because autoregressive decoding dominates total time. These measurements support encoder pruning as a localized efficiency improvement, especially when encoder and decoder are pipelined, but not as a large end-to-end speedup by itself.