跳到正文
HuggingFace Daily Papers(社区热门论文)
34AI 编辑部评分,满分 100

UniH3:统一分层同质性与异质性的全能医学图像修复框架

2026-09-10 08:00· 1天前
AI 导读

北航与清华团队提出 UniH3 框架,通过分层同质性记忆(H2M)模块从高质量图像中蒸馏任务内与任务间先验,并用同质性引导注意力(HGA)注入修复流程,同时以分层异质性平衡器(H2B)缓解任务内与任务间优化冲突。在 MedIR-2D-500K 和 MedIR-3D-3K 两个大规模基准上,UniH3 在全能及单任务医学图像修复中均达到 SOTA。代码已开源。

Zhiwen Yang

School of Biological Science and Medical Engineering, Beihang University, Beijing 100191, China

Jiayin Li

School of Biological Science and Medical Engineering, Beihang University, Beijing 100191, China

Chengyu Liu

School of Biological Science and Medical Engineering, Beihang University, Beijing 100191, China

Hui Zhang

Department of Biomedical Engineering, Tsinghua University, Beijing 100084, China

Bingzheng Wei

E-mail

upyzwup@buaa.edu.cn xuyan04@gmail.com

Yan Xu

School of Biological Science and Medical Engineering, Beihang University, Beijing 100191, China

Abstract

All-in-One medical image restoration (MedIR) aims to address diverse tasks across modalities and degradation types using a single universal model. Existing methods typically prioritize modeling inter-task heterogeneity (e.g., distinct data distributions and degradation types). However, they largely neglect the inherent homogeneity present in medical images, such as widely shared anatomical structures within and across modalities, which can be leveraged to ease model training and improve generalization. To this end, we propose UniH3, a novel framework that Unifies Hierarchical Homogeneity and Heterogeneity for all-in-one medical image restoration. Specifically, to comprehensively exploit homogeneity, we introduce a Hierarchical Homogeneity Memory (H2M) module that progressively distills intra- and inter-task homogeneity priors from high-quality images during training, and adaptively retrieves the most relevant priors tailored to the input for guided restoration. These retrieved priors are then injected into the restoration pipeline via an efficient Homogeneity-Guided Attention (HGA) mechanism. Furthermore, to comprehensively address heterogeneity, we design a Hierarchical Heterogeneity Balancer (H2B) that mitigates both inter- and intra-task conflicts during optimization, facilitating balanced and effective multi-task learning. Extensive experiments on two large-scale benchmarks—MedIR-2D-500K and MedIR-3D-3K—demonstrate that UniH3 achieves state-of-the-art performance on both all-in-one and single-task medical image restoration. We hope this work establishes a strong benchmark and advances the development of general-purpose medical image restoration models. Code is available at https://github.com/Yaziwel/UniH3.

Keywords: 

Medical Image Restoration All-in-One Universal Model

1 Introduction

Medical image restoration (MedIR) aims to recover a high-quality (HQ) image from a degraded low-quality (LQ) acquisition. Since each medical image modality (e.g., PET, CT, MRI) operates under distinct physical principles and is often studied independently, most MedIR research has focused on a single-task setting, in which researchers train specialized models to address the primary degradation introduced by the imaging physics of each modality. Typical MedIR includes PET image denoising [56, 77, 68, 67], CT image denoising [5, 51, 39], MRI image super-resolution [9, 53, 24, 44]. Despite their success in specific scenarios, single-task models have limited practicality for two main reasons. First, in complex scenarios where multiple MedIR tasks coexist (e.g., multimodal PET/CT and PET/MRI), single-task models trained for one task often underperform on others. Moreover, training separate models for each task leads to inefficiencies in both deployment and maintenance. Second, the single-task paradigm hinders progress toward more general intelligence in MedIR. These limitations motivate interest in a universal model that can handle diverse MedIR tasks.

Recent advances in computer vision have fostered the emergence of All-in-One restoration frameworks [47, 31, 63, 41, 11, 66, 8]. Pioneering research in the medical domain [63, 66, 4, 8], particularly the first work on AMIR [63], has established the feasibility of unified modeling for MedIR. To manage diverse tasks within a single model, existing approaches predominantly focus on modeling task heterogeneity—that is, distinguishing between tasks to apply specialized processing. Techniques such as contrastive learning [31], degradation classification [22], visual prompting [41], and mixture-of-experts (MoE) [63, 69] are widely employed to distinguish between tasks.

However, we argue that current All-in-One methods suffer from two critical limitations. First, they largely overlook the inherent homogeneity of medical images. Compared with natural images, medical images from different modalities and tasks often exhibit more consistent anatomical structures and share stronger biological priors. Neglecting this shared knowledge prevents models from exploiting cross-task synergies, thereby increasing the difficulty of learning as the number of tasks grows. Second, regarding heterogeneity, existing methods typically address only inter-task differences (e.g., different degradation types or modalities) while ignoring intra-task variations (e.g., variations caused by different scanners, centers, or patient demographics). This coarse-grained approach fails to resolve optimization conflicts that arise from subtle intra-task distribution shifts. Therefore, it is imperative to simultaneously model both homogeneity and heterogeneity at a hierarchical level (inter- and intra-task) to achieve robust and effective All-in-One MedIR.

To address these challenges, we propose UniH3, a novel framework that Unifies Hierarchical Homogeneity and Heterogeneity for All-in-One medical image restoration. On the one hand, to fully exploit shared knowledge, we introduce a Hierarchical Homogeneity Memory (H2M) module. This module progressively distills both inter- and intra-task priors from HQ images into a memory bank during training. During inference, it adaptively retrieves the most relevant structural priors tailored to the input, which are then injected into the network via an efficient Homogeneity-Guided Attention (HGA) mechanism to guide restoration. On the other hand, to manage task conflicts comprehensively, we design a Hierarchical Heterogeneity Balancer (H2B). Unlike traditional weighting strategies [26, 58] that only balance loss functions at the task level, H2B dynamically mitigates optimization conflicts at both inter- and intra-task levels, ensuring balanced convergence across diverse data distributions. Finally, to validate the effectiveness of UniH3, we construct a benchmark comprising two large-scale datasets: MedIR-2D-500K, containing 509,200 2D image pairs across seven 2D MedIR tasks, and MedIR-3D-3K, containing 3,522 3D volume pairs across three 3D MedIR tasks. Extensive experiments on this benchmark indicate that UniH3 achieves state-of-the-art (SOTA) performance in both all-in-one and single-task medical image restoration.

  • We propose UniH3, a novel framework that simultaneously models hierarchical inter- and intra-task homogeneity and heterogeneity for effective all-in-one medical image restoration.

  • We present a Hierarchical Homogeneity Memory (H2M) module that can adaptively distill and retrieve homogeneity priors to guide the restoration process. Additionally, an efficient Homogeneity-Guided Attention (HGA) mechanism is introduced to fully exploit the retrieved prior for guided restoration.

  • We develop a Hierarchical Heterogeneity Balancer (H2B), which achieves fine-grained task balancing by resolving optimization conflicts arising from both inter-task distinctions and intra-task variations.

2 Related Work

Single-Task Medical Image Restoration. Because different medical imaging modalities are typically studied independently, most MedIR research focuses on single-task problems that address the primary degradations encountered in each modality. Typical MedIR tasks include PET image denoising [56, 77, 68, 67], CT denoising [5, 51, 39], MRI super-resolution [9, 53, 24], X-ray denoising [46], OCT denoising [13], ultrasound denoising [2], and pathology image super-resolution [30]. With the development of deep learning—especially recent advances in network architectures such as convolutional neural networks (CNNs) [29, 5], Transformers [50, 64], Mamba [18, 39], and RWKV [40, 65]—single-task MedIR methods have made substantial progress. However, these single-task models often suffer large performance drops when applied to other MedIR tasks, which limits their practical applicability in broader contexts such as multi-modal imaging scenarios.

All-in-One Medical Image Restoration. All-in-One image restoration [47, 31, 63, 41, 11, 66, 8]. aims to address multiple degradation types and modalities using a single unified model. Early attempts in computer vision, such as TransWeather [47], relied on task-specific encoder–decoder heads to handle distinct weather conditions, inevitably increasing parameter counts as tasks multiplied. To achieve parameter-efficient unified modeling, AirNet [31] introduced a contrastive learning approach to generate task-specific latent representations, which serve as prompts to guide a shared restoration network. This prompt-based paradigm has become the dominant strategy, with subsequent methods like PromptIR [41], AdaIR [11], and others [66, 8] proposing various mechanisms to learn and inject discriminative prompts for effective task adaptation. In the medical domain, research on All-in-One frameworks is still in its nascent stage [63, 66, 4, 8]. AMIR [63] represents a pioneering effort, utilizing a mixture-of-experts strategy to adapt to three specific medical restoration tasks. While these methods have successfully demonstrated the feasibility of unified restoration, they predominantly focus on modeling inter-task heterogeneity—i.e., distinguishing between different tasks to apply specific processing. Consequently, they largely neglect two critical aspects: the inherent homogeneity of anatomical structures shared across medical modalities, and the fine-grained intra-task heterogeneity arising from variations in scanners and protocols. In contrast, our work formulates a hierarchical learning paradigm that jointly models intra- and inter-task homogeneity and heterogeneity, offering a unified perspective for all-in-one MedIR.

Refer to caption
Figure 1: The framework of UniH3.

3 Method

Fig. 1 illustrates the UniH3 pipeline for all-in-one medical image restoration. UniH3 comprises three main components: a U-shaped restoration backbone (see Fig. 1 (a)) responsible for the basic feature extraction and reconstruction, a Hierarchical Homogeneity Memory (H2M, Fig. 1 (b)) module that employs both intra- and inter-task homogeneity priors to guide the restoration, and a Hierarchical Heterogeneity Balancer (H2B, see Fig. 1 (c)) addresses intra- and inter-task heterogeneity by dynamically balancing task relationships during training. Given a LQ input image ILQH×W×1, UniH3 first applies a 3×3 convolutional input projection to produce shallow features ISH×W×C, where H×W denotes the spatial dimensions and C the number of channels. IS is then processed by a 4-level asymmetric encoder–decoder and transformed into deep features IDH×W×C. Each encoder–decoder level contains multiple Homogeneity-Guided Transformer Blocks (HGATBs, see Fig. 1 (a)) that extract features under the guidance of H2M-generated priors. Considering that both global and local information are important for medical image restoration [19, 65], each HGATB contains two consecutive transformer-layer variants: the first replaces standard self-attention [14] with a Homogeneity-Guided Attention (HGA) to model global interactions, and the second replaces self-attention with a convolution plus Squeeze-and-Excitation (SE) [23] layer to capture local interactions. Finally, ID is projected to a residual image IRH×W×1 by a 3×3 convolution, and the restored HQ output is obtained via the residual connection I^HQ=ILQ+IR. We next introduce our core innovations: the H2M module (Sec. 3.1), the HGA mechanism (Sec. 3.2), and the H2B strategy (Sec. 3.3).

3.1 Hierarchical Homogeneity Memory

We propose the Hierarchical Homogeneity Memory (H2M) to alleviate the escalating learning difficulty associated with the growing number of restoration tasks and imaging domains. Drawing inspiration from multi-task learning [3], which leverages shared knowledge to reduce learning burdens and accelerate convergence, we observe that HQ medical images exhibit rich homogeneous priors at two hierarchical levels: intra-task homogeneity (i.e., consistent anatomical structures among varying patients within the same modality) and inter-task homogeneity (i.e., shared structured representations of the human anatomy across different imaging modalities). To explicitly model these hierarchical properties, H2M establishes two structurally symmetric components: a memory bank M(T+1)L×C and a learnable prototype matrix P(T+1)L×C. Both are organized into T task-specific slots and one task-shared slot (each length of L), as illustrated in Fig. 2. While M is updated via momentum to store distilled HQ anatomical priors, P is a set of learnable parameters that serves as an addressing mechanism, learning how to optimally store and retrieve information from M. The H2M mechanism operates in two phases: Homogeneity Distillation and Homogeneity Retrieval.

Homogeneity Distillation. To acquire compact homogeneity priors that facilitate all-in-one restoration, we distill clean anatomical structures from HQ medical images and progressively archive them into M using an Exponential Moving Average (EMA) during training. Concretely, we project paired LQ–HQ images ILQ, IHQ to the target resolution via pixel-unshuffle downsampling followed by a 3×3 convolution, obtaining paired features FLQ,FHQHW×C. A learnable prototype PL×C with the length of L is then used to query and aggregate crucial HQ priors from FHQ through cross-attention:

CrossAttention(Q,K,V)=Softmax(QK𝖳C)V, (1)
VHQ=CrossAttention(P,FLQ,FHQ), (2)

where VHQ(T+1)L×C. This operation allows P to learn which HQ features are most representative of the clean anatomical structures. Based on the current task index, we select features VShL×C and VSpL×C from VHQ, which are correspondingly stored into the task-shared slot (to store inter-task homogeneity prior) and task-specific slot (to store intra-task homogeneity prior) of M via an EMA strategy:

MSh/SpαMSh/Sp+(1α)VSh/Sp, (3)

where α is the momentum coefficient, and MSh/Sp denotes the corresponding shared or specific slots in M. Initialized as zero, M gradually accumulates generalized intra- and inter-task homogeneity priors from continuous training batches. Note that this distillation procedure (indicated by dashed red arrows in Fig. 2) is performed only during training and discarded at test time.

Homogeneity Retrieval. Once the hierarchical memory M is updated, we retrieve clean homogeneity priors VH tailored to the LQ input by using the LQ feature FLQ as a query to retrieve the relevant clean prior from the memory M via cross attention:

VH=CrossAttention(FLQ,P,M), (4)

where VHHW×C is the retrieved homogeneity prior. The red arrows in Fig. 2 illustrate the HQ information flow from IHQ through M into the resulting homogeneity prior VH. Because the obtained VH is derived from the distilled HQ memory, it is well-suited to compensate for degraded or missing anatomical information in the LQ features.

To facilitate multi-scale guidance, UniH3 incorporates four H2M modules (see Fig. 1(b)) at different levels of the U-shaped restoration backbone so that the retrieved homogeneity priors provide effective restoration guidance across multiple resolutions.

媒体内容 · 前往原文查看
Figure 2: Hierarchical Homogeneity Memory.
媒体内容 · 前往原文查看
Figure 3: Homogeneity-Guided Attention.

3.2 Homogeneity-Guided Attention

To guide the restoration process using homogeneity priors, we propose a novel Homogeneity-Guided Attention (HGA) mechanism. Existing methods typically incorporate restoration guidance via Spatial Feature Transformations (SFT) [55] or cross-attention [11], which treat the LQ features as the basis and the guidance features as supplementary. In contrast, HGA fundamentally shifts the learning paradigm: it anchors the learning starting point on the HQ homogeneity priors rather than the degraded LQ features, thereby substantially reducing the learning difficulty and facilitate model convergence. The design of HGA is detailed below.

HGA is highly flexible and can be implemented on either standard self-attention or transposed self-attention. For clarity of exposition, we formulate it here using standard self-attention. Let the query, key, and value be Q,K,VHW×C, the conventional self-attention output V0O is

V0O=AV,A=Softmax(QK𝖳C). (5)

To incorporate guidance from the homogeneity prior VH, a straightforward variant is to complement the LQ value V with clean VH by direct addition:

V1O=A(V+VH)=AV+AVH. (6)

In Eq. 6, the attention mechanism treats V and VH symmetrically. However, the homogeneity prior VH contains higher-fidelity information than the degraded observation V, the attention mechanism should preferentially exploit the more reliable VH. To encourage such a preference, we introduce a second variant that biases attention away from the LQ value V and toward the homogeneity prior VH by adding identity-based terms to the attention map A:

V2O =(AI)V+(A+I)VH (7)
=A(V+VH)+VHV,

where I denotes the identity matrix. The ±I terms increase the self-contribution of VH while reducing that of V. V+VH denotes the mixed value, and VHV acts as a preference bias that reinforces more reliance on the homogeneity prior VH. To stabilize training and increase model expressivity, the final HGA mechanism is obtained by applying channel-wise learnable weighting parameters λ1,λ2C:

VO=A[(1λ1)V+λ1VH]+λ2(VHV). (8)

When λ1=λ2=0, the HGA reduces to the conventional self-attention. The self-attention–based HGA in Eq. 8 has an analogous form to transposed self-attention (see supplement). Our proposed UniH3 adopts the HGA based on transposed self-attention following Restormer [70]. Fig. 3 illustrates the HGA formulation, which augments attention with simple addition and subtraction operations on the value.

3.3 Hierarchical Heterogeneity Balancer

We propose a Hierarchical Heterogeneity Balancer (H2B) to mitigate inter- and intra-task heterogeneity across diverse MedIR tasks during the optimization process. Heterogeneity among tasks induces gradient conflicts that create an imbalance in optimization: some tasks dominate training while others remain under-trained. Previous work in multi-task learning [26] and all-in-one natural image restoration [58] has shown that uncertainty-based loss balancing is a good way of addressing inter-task heterogeneity by dynamically scaling different task losses for a reasonable optimization route:

UB=1Tt=1T(12σt2rec(t)+logσt), (9)

where T denotes the number of tasks, rec(t) denotes the reconstruction loss, and σt is a learnable scalar that estimates task-level uncertainty. The factor 12σt2 adaptively rescales each task’s contribution while the logσt term regularizes the scaling. When rec(t) increases and tends to dominate the total loss, σt increases to attenuate its contribution, and vice versa. However, this uncertainty balancing is too coarse: a single scalar σt per-task cannot capture intra-task heterogeneity (e.g., scanner/center/anatomy variations), so hard samples still remain insufficiently handled. We therefore introduce a hierarchical uncertainty model. For task t and sample s we define the total uncertainty σt,s as the sum of a global task term σt and a sample-specific correction Δσs:

σt,s=σt+Δσs, (10)

where t indexes tasks and s indexes samples. σt is still a learnable scalar for each task while Δσs is predicted by a lightweight Uncertainty Estimation Block (UEB, see Fig. 1(c)) conditioned on sample-specific signals:

Δσs=UEB(Concat[IsLQ,sg(I^sHQ),IsHQ]), (11)

where IsLQ is the LQ input for sample s, I^sHQ is the model prediction, IsHQ is the HQ ground truth, and sg() denotes stop-gradient to decouple loss balancing from restoration model optimization. The H2B loss then aggregates per-task and per-sample contributions as:

H2B=1TSt=1Ts=1S(12σt,s2rec(t,s)+logσt,s), (12)

where S is the batch size. H2B retains the theoretical foundation of standard uncertainty-based balancing [26] while refining it to capture uncertainty at two hierarchical levels: a task-level term σt to effectively mitigate inter-task heterogeneity, and a sample-level correction Δσs to mitigate intra-task heterogeneity.

4 Experiments

We conduct experiments under two settings, All-in-One and Single-Task, for both 2D and 3D MedIR tasks. In the All-in-One setting, a single universal model is trained to address multiple MedIR tasks within either the 2D or 3D domain. In the Single-Task setting, separate models are trained for each MedIR task. We first describe the experimental setup, including datasets, implementation details, and evaluation. We then present comparative results in Sec. 4.1 and Sec. 4.2, and ablation studies in Sec. 4.3.

Datasets. Most existing MedIR datasets are limited in size and narrowly tailored to specific tasks and modalities. To promote the development of general-purpose MedIR methods, we organize publicly available datasets together with private collections into two datasets, as summarized in Tab. 1: (i) MedIR-2D-500K comprises 509,200 2D LQ-HQ image pairs across seven distinct 2D MedIR tasks: PET image denoising, CT image denoising, MRI image super-resolution, X-ray image denoising, OCT image denoising, ultrasound image denoising, and pathological image super-resolution. (ii) MedIR-3D-3K includes 3,522 3D LQ-HQ volume pairs covering three 3D MedIR tasks: PET image denoising, CT image denoising, and MRI image super-resolution. We expect these two datasets to serve as a useful benchmark for advancing general-purpose MedIR research. More detailed descriptions are shown in the supplement.

媒体内容 · 前往原文查看
Table 1: Overview of the MedIR-2D-500K and MedIR-3D-3K datasets.
Dataset Dimension Modality Training Testing Total Data Source
PET 77,000 8,600 85,600 [61], Private
CT 64,000 8,100 72,100 [36, 37] , Private
MRI 75,000 8,400 83,400 [49, 35]
X-ray 108,000 11,600 119,600 [54, 10, 42, 52]
OCT 32,000 3,600 35,600 [32, 17, 16]
Ultrasound 56,000 5,800 61,800 [20, 38, 48, 28, 62]
Pathology 46,000 5,100 51,100 [15, 27, 12, 43, 1, 45]
MedIR-2D-500K 2D Total 458,000 51,200 509,200 -
PET 1,388 156 1,544 [61], Private
CT 258 30 288 [36, 37], Private
MRI 1,520 170 1,690 [49, 35]
MedIR-3D-3K 3D Total 3,166 356 3,522 -

Implementation. For the UniH3 architecture, the number of HGATBs are N1=2, N2=N3=3, and N4=4. The input channel dimension is C=48. For the H2M module, we set the the number of tasks T=7, memory length L=128 and EMA coefficient α=0.99. During training we use patches of size 128×128 with a batch size of 14. The reconstruction loss rec is defined as the L1 loss. The model is optimized using Muon optimizer [34] for 6×105 iterations, with an initial learning rate 3×104 and annealed to 1×107 using a cosine schedule. To support 3D MedIR, we introduce a 3D variant, UniH3-3D, obtained by replacing each module in UniH3 with its 3D counterpart. For UniH3-3D, the HGATB numbers are N1=N2=1 and N3=N4=5, the input channel number is C=16, and the init learning rate is set as 5×105. The patch size is 64×64×64 and the batch size is 6. All other settings are identical to the original UniH3.

Evaluation. To quantitatively assess image quality, we employ the widely used PSNR and SSIM metrics. In the reported tables, the highest and second-highest scores are indicated in red and blue, respectively.

媒体内容 · 前往原文查看
Table 2: All-in-one MedIR comparison results on the MedIR-2D-500K dataset.
PET CT MRI X-ray OCT Ultrasound Pathology Average
Method #Params (M) FLOPs (G) PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑
SwinIR [33] 11.50 187.93 44.24 0.9866 43.17 0.9338 38.44 0.9453 35.87 0.9275 35.70 0.8892 27.52 0.8089 28.35 0.7941 36.18 0.8979
Uformer [57] 50.47 21.42 44.51 0.9874 43.45 0.9356 39.01 0.9507 36.59 0.9353 35.84 0.8908 27.61 0.8135 28.55 0.8011 36.51 0.9020
Restormer [70] 26.12 35.21 44.47 0.9873 43.44 0.9353 39.05 0.9515 36.46 0.9335 35.84 0.8909 27.66 0.8138 28.53 0.8006 36.49 0.9018
NAFNet [6] 67.89 15.74 44.40 0.9871 43.32 0.9346 38.90 0.9502 36.33 0.9326 35.82 0.8905 27.59 0.8131 28.49 0.7995 36.41 0.9011
Restore-RWKV [65] 27.91 37.46 44.45 0.9872 43.46 0.9352 39.05 0.9514 36.46 0.9333 35.84 0.8909 27.68 0.8149 28.55 0.8011 36.50 0.9020
MambaIR [19] 31.50 34.35 44.50 0.9873 43.47 0.9355 39.07 0.9517 36.57 0.9341 35.83 0.8904 27.69 0.8146 28.53 0.8004 36.52 0.9020
TransWeather [47] 38.05 1.55 43.73 0.9846 41.12 0.9217 37.62 0.9333 35.28 0.9221 35.12 0.8699 27.23 0.8011 28.17 0.7874 35.47 0.8886
AirNet [31] 7.61 230.48 44.32 0.9868 43.34 0.9348 38.81 0.9489 36.38 0.9322 35.77 0.8900 27.62 0.8127 28.46 0.7981 36.39 0.9005
DRMC [68] 0.62 9.92 43.62 0.9841 42.48 0.9281 37.24 0.9301 34.06 0.9085 35.46 0.8862 27.09 0.7993 28.02 0.7830 35.42 0.8885
AMIR [63] 23.54 31.76 44.49 0.9873 43.47 0.9356 39.09 0.9519 36.47 0.9333 35.89 0.8914 27.69 0.8150 28.57 0.8019 36.52 0.9023
PromptIR [41] 35.59 39.49 44.52 0.9874 43.48 0.9355 39.13 0.9524 36.57 0.9341 35.84 0.8909 27.69 0.8152 28.54 0.8006 36.54 0.9023
AdaIR [11] 28.76 36.74 44.55 0.9875 43.49 0.9356 39.17 0.9527 36.60 0.9344 35.86 0.8915 27.69 0.8150 28.56 0.8015 36.56 0.9026
UniH3 (Ours) 28.96 26.33 44.89 0.9883 43.65 0.9368 39.55 0.9564 36.88 0.9368 35.96 0.8921 27.80 0.8179 28.63 0.8035 36.77 0.9045
媒体内容 · 前往原文查看
Table 3: Single-task MedIR comparison results on the MedIR-2D-500K dataset.
PET CT MRI X-ray OCT Ultrasound Pathology Average
Method #Params (M) FLOPs (G) PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑
SwinIR [33] 11.50 187.93 44.08 0.9860 43.53 0.9358 39.19 0.9529 36.30 0.9308 35.80 0.8904 27.68 0.8143 28.50 0.7993 36.44 0.9014
Uformer [57] 50.47 21.42 44.61 0.9876 43.54 0.9362 39.21 0.9526 36.61 0.9363 35.91 0.8915 27.64 0.8151 28.58 0.8022 36.59 0.9031
Restormer [70] 26.12 35.21 44.90 0.9883 43.64 0.9369 39.44 0.9554 36.71 0.9359 35.97 0.8921 27.76 0.8164 28.63 0.8038 36.72 0.9041
NAFNet [6] 67.89 15.74 44.74 0.9880 43.55 0.9362 39.29 0.9540 36.48 0.9343 35.93 0.8919 27.67 0.8139 28.58 0.8024 36.61 0.9030
Restore-RWKV [65] 27.91 37.46 44.93 0.9884 43.64 0.9369 39.57 0.9565 36.77 0.9359 35.97 0.8923 27.76 0.8168 28.62 0.8036 36.75 0.9043
MambaIR [19] 31.50 34.35 44.93 0.9883 43.55 0.9363 39.59 0.9568 36.77 0.9363 35.99 0.8923 27.77 0.8178 28.64 0.8041 36.75 0.9046
UniH3 (Ours) 28.96 26.33 45.11 0.9889 43.77 0.9376 39.78 0.9584 37.05 0.9384 36.04 0.8929 27.86 0.8192 28.69 0.8050 36.90 0.9058
Refer to caption
Figure 4: Visual comparison of methods for all-in-one medical image restoration on the MedIR-2D-500K dataset.
媒体内容 · 前往原文查看
PET CT MRI Average
Method PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑
3D-cGAN [56] 48.22 0.9937 40.82 0.9320 38.57 0.9489 42.54 0.9582
MRDG [53] 48.68 0.9943 42.86 0.9388 38.93 0.9519 43.49 0.9617
DRMC [68] 48.58 0.9933 43.33 0.9391 38.81 0.9514 43.57 0.9613
Spach Transformer [25] 48.91 0.9949 42.21 0.9380 39.04 0.9542 43.39 0.9624
Restore-RWKV-3D [65] 49.08 0.9951 43.35 0.9402 39.43 0.9578 43.95 0.9644
UniH3-3D (Ours) 49.43 0.9955 44.05 0.9429 40.03 0.9637 44.50 0.9674
Table 4: 3D all-in-one MedIR results on the MedIR-3D-3K dataset.
媒体内容 · 前往原文查看
PET CT MRI Average
Method PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑
3D-cGAN [56] 48.47 0.9944 40.86 0.9309 38.76 0.9510 42.70 0.9588
MRDG [53] 48.40 0.9942 43.14 0.9395 39.34 0.9570 43.63 0.9636
DRMC [68] 48.76 0.9947 43.60 0.9404 38.98 0.9540 43.78 0.9630
Spach Transformer [25] 48.71 0.9946 43.57 0.9409 39.16 0.9546 43.81 0.9634
Restore-RWKV-3D [65] 48.67 0.9945 42.48 0.9375 39.15 0.9551 43.43 0.9624
UniH3-3D (Ours) 49.77 0.9958 45.37 0.9448 40.73 0.9689 45.29 0.9698
Table 5: 3D single-task MedIR results on the MedIR-3D-3K dataset.

4.1 All-in-One MedIR Results

2D All-in-One MedIR. We evaluate the 2D All-in-One setting on the MedIR-2D-500K dataset. We compare it to several general image-restoration methods (SwinIR [33], Uformer [57], Restormer [70], NAFNet [6], Restore-RWKV [65], and MambaIR [19]) and to SOTA all-in-one approaches (TransWeather [47], AirNet [31], DRMC [68], AMIR [63], PromptIR [41], and AdaIR [11]). Tab. 2 shows that UniH3 achieves both high efficiency and strong effectiveness. It significantly outperforms all comparison methods across all seven tasks. On average, UniH3 surpasses the second-best AdaIR by 0.21 dB in PSNR, which is an appreciable improvement given that each of the seven task contains thousands of testing images. Visual comparison in Fig. 4 demonstrates that UniH3 best restores different types of medical images with higher structural fidelity and finer detail than competing methods.

3D All-in-One MedIR. We assess the UniH3-3D in the 3D All-in-One setting on the MedIR-3D-3K dataset. We compare it with several SOTA 3D restoration methods, including 3D-cGAN [56], MRDG [53], DRMC [68], Spach Transformer [25], and Restore-RWKV-3D [65]. Tab. 4 shows that UniH3 significantly outperforms all comparison methods across the three 3D MedIR tasks. In particular, UniH3-3D exceeds the second-best Restore-RWKV-3D by an average margin of 0.55 dB in PSNR. Visual comparisons are shown in the supplement.

媒体内容 · 前往原文查看
Table 6: Performance of H2P and H2B on different backbones on the MedIR-2D-500K dataset. denotes applying H2P and H2B to the corresponding backbone.
PET CT MRI X-ray OCT Ultrasound Pathology Average
Method #Params (M) FLOPs (G) PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑
Uformer 50.47 21.42 44.51 0.9874 43.45 0.9356 39.01 0.9507 36.59 0.9353 35.84 0.8908 27.61 0.8135 28.55 0.8011 36.51 0.9021
Uformer 53.15 22.10 44.80 0.9880 43.59 0.9363 39.39 0.9549 36.74 0.9354 35.94 0.8918 27.72 0.8155 28.57 0.8022 36.68 0.9034
Restormer 26.12 35.21 44.47 0.9873 43.44 0.9353 39.05 0.9515 36.46 0.9335 35.84 0.8909 27.66 0.8138 28.53 0.8006 36.49 0.9018
Restormer 27.56 35.83 44.78 0.9880 43.58 0.9362 39.37 0.9548 36.71 0.9352 35.93 0.8919 27.73 0.8159 28.60 0.8027 36.67 0.9035
PromptIR 35.59 39.49 44.52 0.9874 43.48 0.9355 39.13 0.9524 36.57 0.9341 35.84 0.8909 27.69 0.8152 28.54 0.8006 36.54 0.9023
PromptIR 36.15 40.68 44.85 0.9882 43.62 0.9366 39.40 0.9551 36.76 0.9365 35.94 0.8919 27.75 0.8165 28.59 0.8024 36.70 0.9039
AdaIR 28.76 36.74 44.55 0.9875 43.49 0.9356 39.17 0.9527 36.60 0.9344 35.86 0.8915 27.69 0.8150 28.56 0.8015 36.56 0.9026
AdaIR 29.77 37.37 44.93 0.9884 43.64 0.9367 39.53 0.9563 36.82 0.9363 35.95 0.8920 27.76 0.8165 28.61 0.8027 36.75 0.9041
Baseline 27.17 25.50 44.52 0.9874 43.47 0.9355 39.10 0.9521 36.39 0.9330 35.88 0.8914 27.68 0.8149 28.57 0.8018 36.52 0.9023
UniH3 (Ours) 28.96 26.33 44.89 0.9883 43.65 0.9368 39.55 0.9564 36.88 0.9368 35.96 0.8921 27.80 0.8179 28.63 0.8035 36.77 0.9045
媒体内容 · 前往原文查看
Table 7: Component analysis.
H2M H2B #Params (M) FLOPs (G) PSNR↑ SSIM↑
27.17 25.50 36.52 0.9023
28.96 26.33 36.66 0.9034
27.17 25.50 36.64 0.9033
28.96 26.33 36.77 0.9045
媒体内容 · 前往原文查看
Table 8: Ablation studies on HGA.
Method #Params (M) FLOPs (G) PSNR↑ SSIM↑
w/o HGA 27.17 25.50 36.52 0.9023
SFT [55] 37.46 36.99 36.74 0.9041
Cross Attention [11] 29.60 27.43 36.67 0.9035
HGA (Ours) 28.96 26.33 36.77 0.9045
媒体内容 · 前往原文查看
Table 9: Ablation studies on H2M.
Inter-task
Homogeneity
Intra-task
Homogeneity
#Params (M) FLOPs (G) PSNR↑ SSIM↑
28.96 26.33 36.52 0.9023
28.96 26.33 36.61 0.9031
28.96 26.33 36.72 0.9038
28.96 26.33 36.77 0.9045
媒体内容 · 前往原文查看
Table 10: Ablation studies on H2B.
Inter-task
Heterogeneity
Intra-task
Heterogeneity
#Params (M) FLOPs (G) PSNR↑ SSIM↑
28.96 26.33 36.66 0.9034
28.96 26.33 36.70 0.9036
28.96 26.33 36.73 0.9041
28.96 26.33 36.77 0.9045

4.2 Single-Task MedIR Results

2D Single-Task MedIR. We evaluate UniH3 for 2D single-task MedIR on the MedIR-2D-500K dataset, comparing it to five general image restoration methods. As shown in Tab. 3, UniH3 significantly outperforms all comparison methods. On average across seven tasks, UniH3 improves PSNR by 0.15 dB over the second-best MambaIR.

3D Single-Task MedIR. We evaluate 3D single-task MedIR on the MedIR-3D-3K dataset and compare UniH3-3D to five SOTA 3D restoration methods. As shown in Tab. 5, UniH3-3D beats all comparison methods. In particular, it improves average PSNR by 1.48 dB over the second-best Spach Transformer [25] across the three tasks.

4.3 Ablation Studies

To evaluate the effectiveness of individual components, we conduct ablation experiments on the 2D All-in-One MedIR task with the MedIR-2D-500K dataset.

Component Analysis of H2M and H2B. We first perform a component analysis of the H2M module and the H2B strategy by selectively disabling each component. To disable H2M, we remove the H2M module and replace the HGA mechanism with a transposed self-attention layer [70]. The H2B is disabled by substituting the loss term H2B with an L1 loss. Table 8 shows that both components improve model performance with minimal increase in computational cost. This finding is further supported by the visual comparison in Fig. 6, where both components contribute to better preservation of image details. Moreover, we apply the H2M module and H2B strategy to other Transformer-based U-shaped backbones, including Uformer, Restormer, PromptIR, and AdaIR. Results in the Tab. 6 demonstrate significant improvements across all these backbones.

Ablation Studies on H2M. We investigate the effectiveness of both inter- and intra-task homogeneity priors in H2M by selectively disabling the task-specific and task-shared slots in M. By replacing these slots with naive learnable parameters—thus isolating them from the HQ Homogeneity Distillation procedure—we observe a drop in performance in Tab. 10. Results indicate that both priors independently benefit the restoration process, and their combination achieves the best performance. To further understand the internal mechanics of H2M, Fig. 6 visualizes the retrieval attention maps for PET and CT tokens. The distributions confirm that tokens primarily query the task-shared slot and their specific modality slot. Furthermore, attention maps reflect clear semantic correlations: anatomically similar tokens within the same modality share highly similar query patterns (8/10 top-score overlap for two PET spine tokens), and cross-modality similarities are also captured (2/10 overlap for PET and CT spine tokens). In contrast, dissimilar anatomies exhibit distinct query patterns (0/10 overlap for PET spine and lesion tokens). This provides strong visual evidence that H2M successfully maps, stores, and retrieves hierarchical anatomical priors.

Ablation Studies on HGA. We validate the impact of HGA by replacing it with alternative guiding mechanisms, including Spatial Feature Transform (SFT) [55] and cross attention [11]. Tab. 8 shows that the proposed HGA performs the best with minimal computation and parameter increase.

Refer to caption
Figure 5: Visual comparison for component analysis.
Refer to caption
Figure 6: Retrieval attention map in H2M. The Top-10 scores are marked by red rectangles.
媒体内容 · 前往原文查看
Figure 7: Estimated uncertainty distribution.
媒体内容 · 前往原文查看
Figure 8: All-in-One vs. Single-Task.

Ablation Studies on H2B. We validate the roles of intra- and inter-task heterogeneity in H2B in Tab. 10. The results indicate that mitigating both levels of heterogeneity improves overall restoration performance, with their joint optimization achieving the best results. In Fig. 8, we visualize the estimated uncertainty distributions. While the conventional UB [26] estimates an uncertainty σt per task and therefore handles only inter-task heterogeneity, our proposed H2B estimates uncertainty σt,s at both the task and the sample level: each sample receives its own uncertainty while each task exhibits a distinct uncertainty distribution. This enables our H2B to capture finer-grained relationships, i.e., both inter-task and intra-task heterogeneity, facilitating convergence toward a more optimal solution for diverse MedIR tasks.

5 Discussion and Limitation

All-in-one medical image restoration is an emerging research field. The practical value of all-in-one models remains under active discussion. In this paper, experiments on a large-scale dataset show that a single all-in-one model, UniH3, already achieves comparable performance to SOTA single-task MambaIR models across seven MedIR tasks (see Fig. 8). This result strongly supports the practicality of developing all-in-one models instead of separate single-task models for MedIR tasks. Additionally, by fine-tuning the well-trained all-in-one UniH3 model for each task, we are able to transfer its learned knowledge to individual tasks and achieve further improvements (also shown in Fig. 8), indicating that the all-in-one model can serve as a transferable pretrained backbone. Our study has limitations: we focus only on the primary restoration task within each modality and therefore do not cover other tasks or degradation types that may occur in the same modality. Future work should address these gaps and pursue more universal MedIR models to benefit clinical diagnosis and more downstream tasks [75, 76, 59, 60, 21, 7, 72, 74, 71, 73].

6 Conclusion

In this paper, we have presented UniH3, a novel and unified framework for all-in-one medical image restoration. By moving beyond the conventional focus on inter-task heterogeneity, UniH3 effectively leverages the inherent hierarchical homogeneity across diverse medical imaging modalities. The integration of the Hierarchical Homogeneity Memory (H2M) module and the Homogeneity-Guided Attention (HGA) mechanism allows the model to distill and utilize inter- and intra-task homogeneity priors to guide the restoration process. Simultaneously, the Hierarchical Heterogeneity Balancer (H2B) ensures a stable and balanced multi-task learning process by mitigating both inter- and intra-task heterogeneity. Extensive evaluations on the large-scale MedIR-2D-500K and MedIR-3D-3K benchmarks demonstrate that UniH3 significantly outperforms existing methods, achieving state-of-the-art performance in both all-in-one and single-task scenarios. In future work, we plan to expand the task spectrum by incorporating additional modalities and degradation types, moving toward more general MedIR models.

Acknowledgements

This work is supported by the National Natural Science Foundation in China under Grant U23B2063 and 62371016, the Bejing Natural Science Foundation Haidian District Joint Fund in China under Grant L2602042, the Beijing hope run special fund of cancer foundation of China under Grant LC2018L02, the Fundamental Research Funds for the Central University of China from the State Key Laboratory of Software Development Environment in Beihang University in China, the 111 Proiect in China under Grant B13003, the Academic Excellence Foundation of BUAA for PhD Students.

References

  • [1] A. Aksac, D. J. Demetrick, T. Ozyer, and R. Alhajj (2019) BreCaHAD: a dataset for breast cancer histopathological annotation and diagnosis. BMC research notes 12 (1), pp. 82. Cited by: Table I, Table 1.
  • [2] H. Asgariandehkordi, S. Goudarzi, A. Basarab, and H. Rivaz (2023) Deep ultrasound denoising using diffusion probabilistic models. In 2023 IEEE International Ultrasonics Symposium (IUS), pp. 1–4. Cited by: §2.
  • [3] R. Caruana (1997) Multitask learning. Machine learning 28 (1), pp. 41–75. Cited by: §3.1.
  • [4] H. Chen, Z. Yang, H. Hou, H. Zhang, B. Wei, G. Zhou, and Y. Xu (2025) All-in-one medical image restoration with latent diffusion-enhanced vector-quantized codebook prior. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 67–77. Cited by: §1, §2.
  • [5] H. Chen, Y. Zhang, M. K. Kalra, F. Lin, Y. Chen, P. Liao, J. Zhou, and G. Wang (2017) Low-dose ct with a residual encoder-decoder convolutional neural network. IEEE Transactions on Medical Imaging 36 (12), pp. 2524–2535. Cited by: §1, §2.
  • [6] L. Chen, X. Chu, X. Zhang, and J. Sun (2022) Simple baselines for image restoration. In European conference on computer vision, pp. 17–33. Cited by: §4.1, Table 2, Table 3.
  • [7] L. Chen (2026) Beyond external constraints: the missing dimension of ai governance. Available at SSRN 6449738. Cited by: §5.
  • [8] T. Chen, X. Ma, L. Bai, W. Wang, Y. Sun, and L. Zhou (2025) EndoIR: degradation-agnostic all-in-one endoscopic image restoration via noise-aware routing diffusion. arXiv preprint arXiv:2511.05873. Cited by: §1, §2.
  • [9] Y. Chen, F. Shi, A. G. Christodoulou, Y. Xie, Z. Zhou, and D. Li (2018) Efficient and accurate mri super-resolution using a generative adversarial network and 3d multi-level densely connected network. In International conference on medical image computing and computer-assisted intervention, pp. 91–99. Cited by: §1, §2.
  • [10] M. E. Chowdhury, T. Rahman, A. Khandakar, R. Mazhar, M. A. Kadir, Z. B. Mahbub, K. R. Islam, M. S. Khan, A. Iqbal, N. Al Emadi, et al. (2020) Can ai help in screening viral and covid-19 pneumonia?. Ieee Access 8, pp. 132665–132676. Cited by: Table I, Table 1.
  • [11] Y. Cui, S. W. Zamir, S. Khan, A. Knoll, M. Shah, and F. S. Khan (2025) Adair: adaptive all-in-one image restoration via frequency mining and modulation. In 13th International Conference on Learning Representations, ICLR 2025, pp. 57335–57356. Cited by: §1, §2, §3.2, §4.1, §4.3, Table 2, Table 8.
  • [12] Q. Da, X. Huang, Z. Li, Y. Zuo, C. Zhang, J. Liu, W. Chen, J. Li, D. Xu, Z. Hu, et al. (2022) DigestPath: a benchmark dataset with challenge review for the pathological detection and segmentation of digestive-system. Medical image analysis 80, pp. 102485. Cited by: Table I, Table 1.
  • [13] Z. Dong, G. Liu, G. Ni, J. Jerwick, L. Duan, and C. Zhou (2020) Optical coherence tomography image denoising using a generative adversarial network with speckle modulation. Journal of biophotonics 13 (4), pp. e201960135. Cited by: §0.C.2, §2.
  • [14] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929. Cited by: §3.
  • [15] C. R. Drifka, A. G. Loeffler, K. Mathewson, A. Keikhosravi, J. C. Eickhoff, Y. Liu, S. M. Weber, W. J. Kao, and K. W. Eliceiri (2016) Highly aligned stromal collagen is a negative prognostic factor following pancreatic ductal adenocarcinoma resection. Oncotarget 7 (46), pp. 76197. Cited by: Table I, Table 1.
  • [16] L. Fang, S. Li, Q. Nie, J. A. Izatt, C. A. Toth, and S. Farsiu (2012) Sparsity based denoising of spectral domain optical coherence tomography images. Biomedical optics express 3 (5), pp. 927–942. Cited by: Table I, Table 1.
  • [17] M. Geng, X. Meng, L. Zhu, Z. Jiang, M. Gao, Z. Huang, B. Qiu, Y. Hu, Y. Zhang, Q. Ren, et al. (2022) Triplet cross-fusion learning for unpaired image denoising in optical coherence tomography. IEEE Transactions on Medical Imaging 41 (11), pp. 3357–3372. Cited by: Table I, §0.C.2, Table 1.
  • [18] A. Gu and T. Dao (2023) Mamba: linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752. Cited by: §2.
  • [19] H. Guo, J. Li, T. Dai, Z. Ouyang, X. Ren, and S. Xia (2024) Mambair: a simple baseline for image restoration with state-space model. In European conference on computer vision, pp. 222–241. Cited by: §3, §4.1, Table 2, Table 3.
  • [20] Y. Guo, S. Zhou, J. Shi, and Y. Wang (2023) Ultrasound image enhancement challenge 2023. Zenodo. Note: accessed: 2024-02-13 External Links: Link Cited by: Table I, Table 1.
  • [21] R. Hao, B. Jing, H. Yu, and Z. Nie (2026) Styledrive: towards driving-style aware benchmarking of end-to-end autonomous driving. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 4627–4635. Cited by: §5.
  • [22] J. Hu, L. Jin, Z. Yao, and Y. Lu (2025) Universal image restoration pre-training via degradation classification. arXiv preprint arXiv:2501.15510. Cited by: §1.
  • [23] J. Hu, L. Shen, and G. Sun (2018) Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7132–7141. Cited by: §3.
  • [24] J. Huang, Y. Fang, Y. Wu, H. Wu, Z. Gao, Y. Li, J. Del Ser, J. Xia, and G. Yang (2022) Swin transformer for fast mri. Neurocomputing 493, pp. 281–304. Cited by: §1, §2.
  • [25] S. Jang, T. Pan, Y. Li, P. Heidari, J. Chen, Q. Li, and K. Gong (2023) Spach transformer: spatial and channel-wise transformer based on local and global self-attentions for pet image denoising. IEEE Transactions on Medical Imaging. Cited by: §4.1, §4.2, Table 4, Table 5.
  • [26] A. Kendall, Y. Gal, and R. Cipolla (2018) Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7482–7491. Cited by: §1, §3.3, §3.3, §4.3.
  • [27] N. Kumar, R. Verma, D. Anand, Y. Zhou, O. F. Onder, E. Tsougenis, H. Chen, P. Heng, J. Li, Z. Hu, et al. (2019) A multi-organ nucleus segmentation challenge. IEEE transactions on medical imaging 39 (5), pp. 1380–1391. Cited by: Table I, Table 1.
  • [28] S. Leclerc, E. Smistad, J. Pedrosa, A. Østvik, F. Cervenansky, F. Espinosa, T. Espeland, E. A. R. Berg, P. Jodoin, T. Grenier, et al. (2019) Deep learning for segmentation using an open large-scale dataset in 2d echocardiography. IEEE transactions on medical imaging 38 (9), pp. 2198–2210. Cited by: Table I, Table 1.
  • [29] Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel (1989) Backpropagation applied to handwritten zip code recognition. Neural Computation 1 (4), pp. 541–551. Cited by: §2.
  • [30] B. Li, A. Keikhosravi, A. G. Loeffler, and K. W. Eliceiri (2021) Single image super-resolution for whole slide image using convolutional neural networks and self-supervised color normalization. Medical Image Analysis 68, pp. 101938. Cited by: §2.
  • [31] B. Li, X. Liu, P. Hu, Z. Wu, J. Lv, and X. Peng (2022) All-in-one image restoration for unknown corruption. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 17452–17462. Cited by: §1, §2, §4.1, Table 2.
  • [32] M. Li, K. Huang, Q. Xu, J. Yang, Y. Zhang, Z. Ji, K. Xie, S. Yuan, Q. Liu, and Q. Chen (2024) OCTA-500: a retinal dataset for optical coherence tomography angiography study. Medical image analysis 93, pp. 103092. Cited by: Table I, Table 1.
  • [33] J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte (2021) Swinir: image restoration using swin transformer. In Proceedings of the IEEE International Conference on Computer Vision, pp. 1833–1844. Cited by: §4.1, Table 2, Table 3.
  • [34] J. Liu, J. Su, X. Yao, Z. Jiang, G. Lai, Y. Du, Y. Qin, W. Xu, E. Lu, J. Yan, et al. (2025) Muon is scalable for llm training. arXiv preprint arXiv:2502.16982. Cited by: §4.
  • [35] M. LLCIXI dataset(Website) Note: Accessed: 2024-01-15 External Links: Link Cited by: Table I, Table I, Table 1, Table 1.
  • [36] C. H. McCollough, A. C. Bartley, R. E. Carter, B. Chen, T. A. Drees, P. Edwards, D. R. Holmes III, A. E. Huang, F. Khan, S. Leng, et al. (2017) Low-dose ct for the detection and classification of metastatic liver lesions: results of the 2016 low dose ct grand challenge. Medical physics 44 (10), pp. e339–e352. Cited by: Table I, Table I, §0.C.2, Table 1, Table 1.
  • [37] T. R. Moen, B. Chen, D. R. Holmes III, X. Duan, Z. Yu, L. Yu, S. Leng, J. G. Fletcher, and C. H. McCollough (2021) Low-dose ct image and projection dataset. Medical physics 48 (2), pp. 902–911. Cited by: Table I, Table I, Table 1, Table 1.
  • [38] A. Montoya, Hasnin, kaggle446, shirzad, W. Cukierski, and yffud (2016) Ultrasound nerve segmentation. Note: Kaggle Cited by: Table I, Table 1.
  • [39] Ş. Öztürk, O. C. Duran, and T. Çukur (2024) DenoMamba: a fused state-space model for low-dose ct denoising. arXiv preprint arXiv:2409.13094. Cited by: §1, §2.
  • [40] B. Peng, E. Alcaide, Q. Anthony, A. Albalak, S. Arcadinho, H. Cao, X. Cheng, M. Chung, M. Grella, K. K. GV, et al. (2023) Rwkv: reinventing rnns for the transformer era. arXiv preprint arXiv:2305.13048. Cited by: §2.
  • [41] V. Potlapalli, S. W. Zamir, S. H. Khan, and F. Shahbaz Khan (2023) Promptir: prompting for all-in-one image restoration. Advances in Neural Information Processing Systems 36, pp. 71275–71293. Cited by: §1, §2, §4.1, Table 2.
  • [42] T. Rahman, A. Khandakar, Y. Qiblawey, A. Tahir, S. Kiranyaz, S. B. A. Kashem, M. T. Islam, S. Al Maadeed, S. M. Zughaier, M. S. Khan, et al. (2021) Exploring the effect of image enhancement techniques on covid-19 detection using chest x-ray images. Computers in biology and medicine 132, pp. 104319. Cited by: Table I, Table 1.
  • [43] K. Sirinukunwattana, J. P. Pluim, H. Chen, X. Qi, P. Heng, Y. B. Guo, L. Y. Wang, B. J. Matuszewski, E. Bruni, U. Sanchez, et al. (2017) Gland segmentation in colon histology images: the glas challenge contest. Medical image analysis 35, pp. 489–502. Cited by: Table I, Table 1.
  • [44] Z. Song, Z. Qi, X. Wang, X. Zhao, Z. Shen, S. Wang, M. Fei, Z. Wang, D. Zang, D. Chen, et al. (2025) Uni-coal: a unified framework for cross-modality synthesis and super-resolution of mr images. Expert Systems with Applications 270, pp. 126241. Cited by: §1.
  • [45] E. Tekin, Ç. Yazıcı, H. Kusetogullari, F. Tokat, A. Yavariabdi, L. O. Iheme, S. Çayır, E. Bozaba, G. Solmaz, B. Darbaz, et al. (2023) Tubule-u-net: a novel dataset and deep learning-based tubule segmentation framework in whole slide images of breast cancer. Scientific Reports 13 (1), pp. 128. Cited by: Table I, Table 1.
  • [46] D. Thanh P. Surya et al. (2019) A review on ct and x-ray images denoising methods. Informatica 43 (2). Cited by: §0.C.2, §2.
  • [47] J. M. J. Valanarasu, R. Yasarla, and V. M. Patel (2022) Transweather: transformer-based restoration of images degraded by adverse weather conditions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 2353–2363. Cited by: §1, §2, §4.1, Table 2.
  • [48] T. L. van den Heuvel, D. de Bruijn, C. L. de Korte, and B. v. Ginneken (2018) Automated measurement of fetal head circumference using 2d ultrasound images. PloS one 13 (8), pp. e0200412. Cited by: Table I, Table 1.
  • [49] D. C. Van Essen, S. M. Smith, D. M. Barch, T. E. Behrens, E. Yacoub, K. Ugurbil, W. H. Consortium, et al. (2013) The wu-minn human connectome project: an overview. Neuroimage 80, pp. 62–79. Cited by: Table I, Table I, Table 1, Table 1.
  • [50] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin (2017) Attention is all you need. Advances in Neural Information Processing Systems 30. Cited by: §2.
  • [51] D. Wang, F. Fan, Z. Wu, R. Liu, F. Wang, and H. Yu (2023) CTformer: convolution-free token2token dilated vision transformer for low-dose ct denoising. Physics in Medicine & Biology 68 (6), pp. 065012. Cited by: §1, §2.
  • [52] D. Wang, X. Wang, L. Wang, M. Li, Q. Da, X. Liu, X. Gao, J. Shen, J. He, T. Shen, et al. (2023) A real-world dataset and benchmark for foundation model adaptation in medical image classification. Scientific Data 10 (1), pp. 574. Cited by: Table I, Table 1.
  • [53] J. Wang, Y. Chen, Y. Wu, J. Shi, and J. Gee (2020) Enhanced generative adversarial network for 3d brain mri super-resolution. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 3627–3636. Cited by: §1, §2, §4.1, Table 4, Table 5.
  • [54] X. Wang, Y. Peng, L. Lu, Z. Lu, M. Bagheri, and R. M. Summers (2017) Chestx-ray8: hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2097–2106. Cited by: Table I, Table 1.
  • [55] X. Wang, K. Yu, C. Dong, and C. C. Loy (2018) Recovering realistic texture in image super-resolution by deep spatial feature transform. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 606–615. Cited by: §3.2, §4.3, Table 8.
  • [56] Y. Wang, B. Yu, L. Wang, C. Zu, D. S. Lalush, W. Lin, X. Wu, J. Zhou, D. Shen, and L. Zhou (2018) 3D conditional generative adversarial networks for high-quality pet image estimation at low dose. Neuroimage 174, pp. 550–562. Cited by: §1, §2, §4.1, Table 4, Table 5.
  • [57] Z. Wang, X. Cun, J. Bao, W. Zhou, J. Liu, and H. Li (2022) Uformer: a general u-shaped transformer for image restoration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 17683–17693. Cited by: §4.1, Table 2, Table 3.
  • [58] G. Wu, J. Jiang, Y. Wang, K. Jiang, and X. Liu (2025) Debiased all-in-one image restoration with task uncertainty regularization. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, pp. 8386–8394. Cited by: §1, §3.3.
  • [59] Y. Wu, Y. Zhou, J. Saiyin, B. Wei, M. Lai, J. Shou, and Y. Xu (2024) Attriprompter: auto-prompting with attribute semantics for zero-shot nuclei detection via visual-language pre-trained models. IEEE Transactions on Medical Imaging 44 (2), pp. 982–993. Cited by: §5.
  • [60] Y. Wu, Y. Zhou, J. Saiyin, B. Wei, and Y. Xu (2025) Visual textualization for image prompted object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 20900–20910. Cited by: §5.
  • [61] S. Xue, R. Guo, K. P. Bohn, J. Matzke, M. Viscione, I. Alberts, H. Meng, C. Sun, M. Zhang, M. Zhang, et al. (2022) A cross-scanner and cross-tracer deep learning method for the recovery of standard-dose imaging quality from low-dose pet. European journal of nuclear medicine and molecular imaging 49 (6), pp. 1843–1856. Cited by: Table I, Table I, §0.C.2, Table 1, Table 1.
  • [62] J. Yang, X. Ding, Z. Zheng, X. Xu, and X. Li (2023) Graphecho: graph-driven unsupervised domain adaptation for echocardiogram video segmentation. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 11878–11887. Cited by: Table I, Table 1.
  • [63] Z. Yang, H. Chen, Z. Qian, Y. Yi, H. Zhang, D. Zhao, B. Wei, and Y. Xu (2024) All-in-one medical image restoration via task-adaptive routing. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 67–77. Cited by: §0.C.2, Table II, Table II, Appendix 0.D, Appendix 0.D, §1, §2, §4.1, Table 2.
  • [64] Z. Yang, H. Chen, Z. Qian, Y. Zhou, H. Zhang, D. Zhao, B. Wei, and Y. Xu (2024) Region attention transformer for medical image restoration. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 603–613. Cited by: §2.
  • [65] Z. Yang, J. Li, H. Zhang, D. Zhao, B. Wei, and Y. Xu (2025) Restore-rwkv: efficient and effective medical image restoration with rwkv. IEEE Journal of Biomedical and Health Informatics. Cited by: §2, §3, §4.1, §4.1, Table 2, Table 3, Table 4, Table 5.
  • [66] Z. Yang, J. Zhang, Y. Yi, J. Liang, B. Wei, and Y. Xu (2025) TAT: task-adaptive transformer for all-in-one medical image restoration. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 565–575. Cited by: §1, §2.
  • [67] Z. Yang, Y. Zhou, H. Chen, H. Zhang, D. Zhao, B. Wei, and Y. Xu (2026) UniPET: a universal network for high-quality pet image denoising across varied dose reduction factors. Medical Image Analysis, pp. 104059. Cited by: §1, §2.
  • [68] Z. Yang, Y. Zhou, H. Zhang, B. Wei, Y. Fan, and Y. Xu (2023) Drmc: a generalist model with dynamic routing for multi-center pet image synthesis. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 36–46. Cited by: §1, §2, §4.1, §4.1, Table 2, Table 4, Table 5.
  • [69] E. Zamfir, Z. Wu, N. Mehta, Y. Tan, D. P. Paudel, Y. Zhang, and R. Timofte (2025) Complexity experts are task-discriminative learners for any image restoration. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 12753–12763. Cited by: §1.
  • [70] S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M. Yang (2022) Restormer: efficient transformer for high-resolution image restoration. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5728–5739. Cited by: Appendix 0.B, §3.2, §4.1, §4.3, Table 2, Table 3.
  • [71] K. Zhou, R. Cai, Y. Ma, Q. Tan, X. Wang, J. Li, H. P. Shum, F. W. Li, S. Jin, and X. Liang (2023) A video-based augmented reality system for human-in-the-loop muscle strength assessment of juvenile dermatomyositis. IEEE Transactions on Visualization and Computer Graphics 29 (5), pp. 2456–2466. Cited by: §5.
  • [72] K. Zhou, R. Cai, L. Wang, H. P. H. Shum, and X. Liang (2026) A comprehensive survey of action quality assessment: method and benchmark. Pattern Recognition 179, pp. 113933. Cited by: §5.
  • [73] K. Zhou, Z. Hao, L. Wang, and X. Liang (2025) Adaptive score alignment learning for continual perceptual quality assessment of 360-degree videos in virtual reality. IEEE Transactions on Visualization and Computer Graphics 31 (5), pp. 2880–2890. Cited by: §5.
  • [74] K. Zhou, H. P. H. Shum, F. W. B. Li, X. Zhang, and X. Liang (2025) PHI: bridging domain shift in long-term action quality assessment via progressive hierarchical instruction. IEEE Transactions on Image Processing 34, pp. 3718–3732. External Links: ISSN 1057-7149 Cited by: §5.
  • [75] Y. Zhou, Y. Wu, J. Saiyin, B. Wei, M. Lai, E. Chang, and Y. Xu (2024) SDPT: synchronous dual prompt tuning for fusion-based visual-language pre-trained models. In European Conference on Computer Vision, pp. 340–356. Cited by: §5.
  • [76] Y. Zhou, Y. Wu, J. Saiyin, B. Wei, and Y. Xu (2026) SDPT: synchronous dual prompt tuning for visual-language pre-trained models. IEEE Transactions on Pattern Analysis and Machine Intelligence. Cited by: §5.
  • [77] Y. Zhou, Z. Yang, H. Zhang, I. Eric, C. Chang, Y. Fan, and Y. Xu (2022) 3D segmentation guided style-based generative adversarial networks for pet synthesis. IEEE Transactions on Medical Imaging 41 (8), pp. 2092–2104. Cited by: §1, §2.

Appendix 0.A Availability of Code and Data

The code and data are released at https://github.com/Yaziwel/UniH3. We hope this study will contribute to the advancement of general-purpose MedIR methods.

Appendix 0.B Transposed Self-Attention-Based HGA

Let the query, key, and value be Q,K,VHW×C. Following Eqs.5-8 in Sec. 3.2, we derive the transposed self-attention–based [70] homogeneity-guided attention (HGA) as follows:

H0=VA,A=Softmax(Q𝖳KC). (13)
H1=(V+VH)A=VA+VHA. (14)
H2 =V(AI)+VH(A+I) (15)
=(V+VH)A+VHV.
H=[(1λ1)V+λ1VH]A+λ2(VHV). (16)

The finally derived transposed self-attention-based HGA in Eq. 16 can also be illustrated by Fig. 3, which augments attention with simple addition and subtraction operations on the value. Therefore, Fig. 3 illustrates both the self-attention–based HGA and the transposed self-attention–based HGA.

媒体内容 · 前往原文查看
Table I: MedIR-2D-500K and MedIR-3D-3K datasets.
Dataset Dimension Modality Training Testing Total Data Source
73,125 8,000 81,125 [61]
1,850 300 2,150 Private1
PET 2,025 300 2,325 Private2
PET Total 77,000 8,600 85,600 -
4,470 1,466 5,936 [36]
22,995 2,534 25,529 [37]
CT 36,535 4,100 40,635 Private3
CT Total 64,000 8,100 72,100 -
59,600 6,660 66,260 [49]
MRI 15,400 1,740 17,140 [35]
MRI Total 75,000 8,400 83,400 -
100,600 11,000 111,600 [54]
3,000 200 3,200 [10, 42]
X-ray 4,400 400 4,800 [52]
X-ray Total 108,000 11,600 119,600 -
31932 3592 35,524 [32]
33 4 37 [17]
OCT 35 4 39 [16]
OCT Total 32000 3600 35,600 -
1,950 490 2,440 [20]
8,900 2,220 11,120 [38]
1,050 260 1,310 [48]
13,600 1,680 15,280 [28]
Ultrasound 30,500 1,150 31,650 [62]
Ultrasound Total 56,000 5,800 61,800 -
13200 1,830 15,030 [15]
350 90 440 [27]
23100 2,300 25,400 [12]
400 35 435 [43]
7650 695 8,345 [1]
Pathology 1,300 150 1,450 [45]
Pathology Total 46000 5100 51,100 -
MedIR-2D-500K 2D Total 458,000 51,200 509,200
1,233 138 1,371 [61]
74 9 83 Private1
PET 81 9 90 Private2
PET Total 1,388 156 1,544 -
8 2 10 [36]
135 15 150 [37]
CT 115 13 128 Private3
CT Total 258 30 288 -
519 58 577 [49]
MRI 1001 112 1,113 [35]
MRI Total 1,520 170 1,690 -
MedIR-3D-3K 3D Total 3,166 356 3,522 -

Appendix 0.C Additional Dataset Information

The detailed dataset information is shown in Tab. I. We then describe the information of private data and the methods used to simulate LQ–HQ image pairs.

0.C.1 Private Data Source

We collected two private datasets (Private1 and Private2 in Tab. I) for PET, and one private dataset (Private3 in Tab. I) for CT. This study and the experimental procedures involving all three private datasets were approved by the Biological and Medical Ethnics Committee of Beihang University (approval number BM20250008). Informed consent was obtained from all participating patients.

Private1. We collect 83 3D whole-body PET images using PolarStar m660 PET imaging system, whith an average administered dose of 293 MBq of 18F-FDG. 2D images are extracted from slices of 3D images, excluding slices without anatomical content (e.g., air-only regions).

Private2. We collect 90 3D whole-body PET images using PolarStar Flight PET imaging system, whith an average administered dose of 301 MBq of 18F-FDG. 2D images are extracted from slices of 3D images, excluding slices without anatomical content (e.g., air-only regions).

Private3. We collect 128 3D CT images using the Sinovision CT imaging system. Among them, 18 images are of the spine, 50 of the lungs, and 60 of soft tissues. 2D images are extracted from slices of 3D images, excluding slices without anatomical content (e.g., air-only regions).

0.C.2 Methods for Generating LQ-HQ image Pairs

We describe the simulation methods used to generate LQ–HQ image pairs for different tasks in Tab. I. Our focus is primarily on the key degradation affecting each imaging modality.

PET image denoising. Following the paper [61], the original PET data is collected in listmode. To simulate LQ PET images, list mode data are randomly subsampled to achieve a dose reduction factor of 10. Both HQ and LQ PET images undergo reconstruction using the standard OSEM method.

CT Image Denoising. Following the paper [36], we first obtain the original CT projection data. To simulate LQ CT images, Poisson noise is inserted into the projection data for each case to reach a noise level that corresponded to 25% of the full dose. Both HQ and LQ CT images are then generated using standard CT reconstruction applied to their respective projection data.

MRI Image Super-Resolution. Following the paper [63], the LQ image is generated by transforming the HQ image to the frequency domain, retaining only the central 6.25% of frequency data points while zero-filling the high-frequency parts, and then converting it back to the image domain.

X-ray Image Denoising. According to previous studies [46], X-ray images are primarily affected by Poisson noise. Accordingly, we generate LQ images using the following formulation: ILQ=Poisson(λIHQ)λ, where we set the noise level to λ=30.

OCT Image Denosing. According to the paper [13], OCT images suffer from speckle noise, which can be reduced by averaging repeated scans acquired at the same location. Following prior work [17], we averaged five repeated scans to produce a noise-reduced, HQ image, and randomly selected one of the five original scans as the LQ image.

Ultrasound Image Denoising. Ulrasound image often suffer from a multiplicative speckle noise, which can be approximated as: ILQ=IHQ+(IHQ)γϵ, where ϵ is zero-mean Gaussian noise and γ controls the strength of the content-dependent perturbation. We set γ=0.5

Pathological Image Super-Resolution. We perform 4× Y-channel super-resolution. To synthesize LQ pathological images, we first convert the RGB pathlogical images to the YCbCr color space and then downsample all three channels by a factor of four using bicubic interpolation. During reconstruction, only the Y (luminance) channel is processed by the model, while the Cb and Cr channels are restored using bicubic upsampling, since the human visual system is far more sensitive to luminance than to chrominance.

Appendix 0.D Additional Experiments.

2D All-in-One MedIR Results on AMIR Dataset [63]. We also conduct an all-in-one MedIR performance comparison on the dataset provided in the AMIR paper [63]. This dataset includes three tasks: PET image denoising, CT image denoising, and MRI image super-resolution. The results are shown in Tab. II. Our proposed UniH3 consistently outperforms all comparison methods across these three tasks.

媒体内容 · 前往原文查看
Table II: 2D All-in-One MedIR results on the AMIR datatset [63].
PET CT MRI Average
Method PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑ PSNR↑ SSIM↑
Restormer 37.14 0.9473 33.61 0.9177 31.72 0.9362 34.16 0.9337
Eformer 35.11 0.9091 32.44 0.9078 29.19 0.8728 32.25 0.8966
Spach Transformer 37.05 0.9445 33.47 0.9155 31.18 0.9290 33.90 0.9297
DRMC 36.19 0.9376 33.28 0.9153 29.55 0.9032 33.01 0.9187
AirNet 37.17 0.9451 33.62 0.9176 31.39 0.9316 34.06 0.9314
AMIR 37.12 0.9475 33.70 0.9182 32.03 0.9396 34.28 0.9351
UniH3 (Ours) 37.40 0.9478 33.79 0.9203 32.14 0.9405 34.44 0.9362

Appendix 0.E Additional Visualization.

Visual Comparison for 2D All-in-One MedIR. We provide an additional visual comparison for 2D all-in-one medical image restoration in Fig. I and Fig. II. Our UniH3 best preserves structures and details across seven tasks.

Visual Comparison for 3D All-in-One MedIR. We provide an additional visual comparison across the coronal, sagittal, and transverse planes for the 3D all-in-one medical image restoration in Fig. III. Our UniH3-3D best preserves structures and details across three tasks.

Refer to caption
Figure I: Visual comparison for 2D all-in-one MedIR on the MedIR-2D-500K dataset.
Refer to caption
Figure II: Visual comparison for 2D all-in-one MedIR on the MedIR-2D-500K dataset.
Refer to caption
Figure III: Visual comparison for 3D all-in-one MedIR on the MedIR-3D-3K dataset.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org