Zhiwen Yang
School of Biological Science and Medical Engineering, Beihang University, Beijing 100191, China
Jiayin Li
School of Biological Science and Medical Engineering, Beihang University, Beijing 100191, China
Chengyu Liu
School of Biological Science and Medical Engineering, Beihang University, Beijing 100191, China
Hui Zhang
Department of Biomedical Engineering, Tsinghua University, Beijing 100084, China
Bingzheng Wei
upyzwup@buaa.edu.cn xuyan04@gmail.com
Yan Xu
School of Biological Science and Medical Engineering, Beihang University, Beijing 100191, China
Abstract
All-in-One medical image restoration (MedIR) aims to address diverse tasks across modalities and degradation types using a single universal model. Existing methods typically prioritize modeling inter-task heterogeneity (e.g., distinct data distributions and degradation types). However, they largely neglect the inherent homogeneity present in medical images, such as widely shared anatomical structures within and across modalities, which can be leveraged to ease model training and improve generalization. To this end, we propose UniH3, a novel framework that Unifies Hierarchical Homogeneity and Heterogeneity for all-in-one medical image restoration. Specifically, to comprehensively exploit homogeneity, we introduce a Hierarchical Homogeneity Memory (H2M) module that progressively distills intra- and inter-task homogeneity priors from high-quality images during training, and adaptively retrieves the most relevant priors tailored to the input for guided restoration. These retrieved priors are then injected into the restoration pipeline via an efficient Homogeneity-Guided Attention (HGA) mechanism. Furthermore, to comprehensively address heterogeneity, we design a Hierarchical Heterogeneity Balancer (H2B) that mitigates both inter- and intra-task conflicts during optimization, facilitating balanced and effective multi-task learning. Extensive experiments on two large-scale benchmarks—MedIR-2D-500K and MedIR-3D-3K—demonstrate that UniH3 achieves state-of-the-art performance on both all-in-one and single-task medical image restoration. We hope this work establishes a strong benchmark and advances the development of general-purpose medical image restoration models. Code is available at https://github.com/Yaziwel/UniH3.
Keywords:
Medical Image Restoration All-in-One Universal Model
1 Introduction
Medical image restoration (MedIR) aims to recover a high-quality (HQ) image from a degraded low-quality (LQ) acquisition. Since each medical image modality (e.g., PET, CT, MRI) operates under distinct physical principles and is often studied independently, most MedIR research has focused on a single-task setting, in which researchers train specialized models to address the primary degradation introduced by the imaging physics of each modality. Typical MedIR includes PET image denoising [56, 77, 68, 67], CT image denoising [5, 51, 39], MRI image super-resolution [9, 53, 24, 44]. Despite their success in specific scenarios, single-task models have limited practicality for two main reasons. First, in complex scenarios where multiple MedIR tasks coexist (e.g., multimodal PET/CT and PET/MRI), single-task models trained for one task often underperform on others. Moreover, training separate models for each task leads to inefficiencies in both deployment and maintenance. Second, the single-task paradigm hinders progress toward more general intelligence in MedIR. These limitations motivate interest in a universal model that can handle diverse MedIR tasks.
Recent advances in computer vision have fostered the emergence of All-in-One restoration frameworks [47, 31, 63, 41, 11, 66, 8]. Pioneering research in the medical domain [63, 66, 4, 8], particularly the first work on AMIR [63], has established the feasibility of unified modeling for MedIR. To manage diverse tasks within a single model, existing approaches predominantly focus on modeling task heterogeneity—that is, distinguishing between tasks to apply specialized processing. Techniques such as contrastive learning [31], degradation classification [22], visual prompting [41], and mixture-of-experts (MoE) [63, 69] are widely employed to distinguish between tasks.
However, we argue that current All-in-One methods suffer from two critical limitations. First, they largely overlook the inherent homogeneity of medical images. Compared with natural images, medical images from different modalities and tasks often exhibit more consistent anatomical structures and share stronger biological priors. Neglecting this shared knowledge prevents models from exploiting cross-task synergies, thereby increasing the difficulty of learning as the number of tasks grows. Second, regarding heterogeneity, existing methods typically address only inter-task differences (e.g., different degradation types or modalities) while ignoring intra-task variations (e.g., variations caused by different scanners, centers, or patient demographics). This coarse-grained approach fails to resolve optimization conflicts that arise from subtle intra-task distribution shifts. Therefore, it is imperative to simultaneously model both homogeneity and heterogeneity at a hierarchical level (inter- and intra-task) to achieve robust and effective All-in-One MedIR.
To address these challenges, we propose UniH3, a novel framework that Unifies Hierarchical Homogeneity and Heterogeneity for All-in-One medical image restoration. On the one hand, to fully exploit shared knowledge, we introduce a Hierarchical Homogeneity Memory (H2M) module. This module progressively distills both inter- and intra-task priors from HQ images into a memory bank during training. During inference, it adaptively retrieves the most relevant structural priors tailored to the input, which are then injected into the network via an efficient Homogeneity-Guided Attention (HGA) mechanism to guide restoration. On the other hand, to manage task conflicts comprehensively, we design a Hierarchical Heterogeneity Balancer (H2B). Unlike traditional weighting strategies [26, 58] that only balance loss functions at the task level, H2B dynamically mitigates optimization conflicts at both inter- and intra-task levels, ensuring balanced convergence across diverse data distributions. Finally, to validate the effectiveness of UniH3, we construct a benchmark comprising two large-scale datasets: MedIR-2D-500K, containing 509,200 2D image pairs across seven 2D MedIR tasks, and MedIR-3D-3K, containing 3,522 3D volume pairs across three 3D MedIR tasks. Extensive experiments on this benchmark indicate that UniH3 achieves state-of-the-art (SOTA) performance in both all-in-one and single-task medical image restoration.
We propose UniH3, a novel framework that simultaneously models hierarchical inter- and intra-task homogeneity and heterogeneity for effective all-in-one medical image restoration.
We present a Hierarchical Homogeneity Memory (H2M) module that can adaptively distill and retrieve homogeneity priors to guide the restoration process. Additionally, an efficient Homogeneity-Guided Attention (HGA) mechanism is introduced to fully exploit the retrieved prior for guided restoration.
We develop a Hierarchical Heterogeneity Balancer (H2B), which achieves fine-grained task balancing by resolving optimization conflicts arising from both inter-task distinctions and intra-task variations.
2 Related Work
Single-Task Medical Image Restoration. Because different medical imaging modalities are typically studied independently, most MedIR research focuses on single-task problems that address the primary degradations encountered in each modality. Typical MedIR tasks include PET image denoising [56, 77, 68, 67], CT denoising [5, 51, 39], MRI super-resolution [9, 53, 24], X-ray denoising [46], OCT denoising [13], ultrasound denoising [2], and pathology image super-resolution [30]. With the development of deep learning—especially recent advances in network architectures such as convolutional neural networks (CNNs) [29, 5], Transformers [50, 64], Mamba [18, 39], and RWKV [40, 65]—single-task MedIR methods have made substantial progress. However, these single-task models often suffer large performance drops when applied to other MedIR tasks, which limits their practical applicability in broader contexts such as multi-modal imaging scenarios.
All-in-One Medical Image Restoration. All-in-One image restoration [47, 31, 63, 41, 11, 66, 8]. aims to address multiple degradation types and modalities using a single unified model. Early attempts in computer vision, such as TransWeather [47], relied on task-specific encoder–decoder heads to handle distinct weather conditions, inevitably increasing parameter counts as tasks multiplied. To achieve parameter-efficient unified modeling, AirNet [31] introduced a contrastive learning approach to generate task-specific latent representations, which serve as prompts to guide a shared restoration network. This prompt-based paradigm has become the dominant strategy, with subsequent methods like PromptIR [41], AdaIR [11], and others [66, 8] proposing various mechanisms to learn and inject discriminative prompts for effective task adaptation. In the medical domain, research on All-in-One frameworks is still in its nascent stage [63, 66, 4, 8]. AMIR [63] represents a pioneering effort, utilizing a mixture-of-experts strategy to adapt to three specific medical restoration tasks. While these methods have successfully demonstrated the feasibility of unified restoration, they predominantly focus on modeling inter-task heterogeneity—i.e., distinguishing between different tasks to apply specific processing. Consequently, they largely neglect two critical aspects: the inherent homogeneity of anatomical structures shared across medical modalities, and the fine-grained intra-task heterogeneity arising from variations in scanners and protocols. In contrast, our work formulates a hierarchical learning paradigm that jointly models intra- and inter-task homogeneity and heterogeneity, offering a unified perspective for all-in-one MedIR.
3 Method
Fig. 1 illustrates the UniH3 pipeline for all-in-one medical image restoration. UniH3 comprises three main components: a U-shaped restoration backbone (see Fig. 1 (a)) responsible for the basic feature extraction and reconstruction, a Hierarchical Homogeneity Memory (H2M, Fig. 1 (b)) module that employs both intra- and inter-task homogeneity priors to guide the restoration, and a Hierarchical Heterogeneity Balancer (H2B, see Fig. 1 (c)) addresses intra- and inter-task heterogeneity by dynamically balancing task relationships during training. Given a LQ input image , UniH3 first applies a convolutional input projection to produce shallow features , where denotes the spatial dimensions and the number of channels. is then processed by a 4-level asymmetric encoder–decoder and transformed into deep features . Each encoder–decoder level contains multiple Homogeneity-Guided Transformer Blocks (HGATBs, see Fig. 1 (a)) that extract features under the guidance of H2M-generated priors. Considering that both global and local information are important for medical image restoration [19, 65], each HGATB contains two consecutive transformer-layer variants: the first replaces standard self-attention [14] with a Homogeneity-Guided Attention (HGA) to model global interactions, and the second replaces self-attention with a convolution plus Squeeze-and-Excitation (SE) [23] layer to capture local interactions. Finally, is projected to a residual image by a convolution, and the restored HQ output is obtained via the residual connection . We next introduce our core innovations: the H2M module (Sec. 3.1), the HGA mechanism (Sec. 3.2), and the H2B strategy (Sec. 3.3).
3.1 Hierarchical Homogeneity Memory
We propose the Hierarchical Homogeneity Memory (H2M) to alleviate the escalating learning difficulty associated with the growing number of restoration tasks and imaging domains. Drawing inspiration from multi-task learning [3], which leverages shared knowledge to reduce learning burdens and accelerate convergence, we observe that HQ medical images exhibit rich homogeneous priors at two hierarchical levels: intra-task homogeneity (i.e., consistent anatomical structures among varying patients within the same modality) and inter-task homogeneity (i.e., shared structured representations of the human anatomy across different imaging modalities). To explicitly model these hierarchical properties, H2M establishes two structurally symmetric components: a memory bank and a learnable prototype matrix . Both are organized into task-specific slots and one task-shared slot (each length of ), as illustrated in Fig. 2. While is updated via momentum to store distilled HQ anatomical priors, is a set of learnable parameters that serves as an addressing mechanism, learning how to optimally store and retrieve information from . The H2M mechanism operates in two phases: Homogeneity Distillation and Homogeneity Retrieval.
Homogeneity Distillation. To acquire compact homogeneity priors that facilitate all-in-one restoration, we distill clean anatomical structures from HQ medical images and progressively archive them into using an Exponential Moving Average (EMA) during training. Concretely, we project paired LQ–HQ images , to the target resolution via pixel-unshuffle downsampling followed by a convolution, obtaining paired features . A learnable prototype with the length of is then used to query and aggregate crucial HQ priors from through cross-attention:
| (1) |
| (2) |
where . This operation allows to learn which HQ features are most representative of the clean anatomical structures. Based on the current task index, we select features and from , which are correspondingly stored into the task-shared slot (to store inter-task homogeneity prior) and task-specific slot (to store intra-task homogeneity prior) of via an EMA strategy:
| (3) |
where is the momentum coefficient, and denotes the corresponding shared or specific slots in . Initialized as zero, gradually accumulates generalized intra- and inter-task homogeneity priors from continuous training batches. Note that this distillation procedure (indicated by dashed red arrows in Fig. 2) is performed only during training and discarded at test time.
Homogeneity Retrieval. Once the hierarchical memory is updated, we retrieve clean homogeneity priors tailored to the LQ input by using the LQ feature as a query to retrieve the relevant clean prior from the memory via cross attention:
| (4) |
where is the retrieved homogeneity prior. The red arrows in Fig. 2 illustrate the HQ information flow from through into the resulting homogeneity prior . Because the obtained is derived from the distilled HQ memory, it is well-suited to compensate for degraded or missing anatomical information in the LQ features.
To facilitate multi-scale guidance, UniH3 incorporates four H2M modules (see Fig. 1(b)) at different levels of the U-shaped restoration backbone so that the retrieved homogeneity priors provide effective restoration guidance across multiple resolutions.
3.2 Homogeneity-Guided Attention
To guide the restoration process using homogeneity priors, we propose a novel Homogeneity-Guided Attention (HGA) mechanism. Existing methods typically incorporate restoration guidance via Spatial Feature Transformations (SFT) [55] or cross-attention [11], which treat the LQ features as the basis and the guidance features as supplementary. In contrast, HGA fundamentally shifts the learning paradigm: it anchors the learning starting point on the HQ homogeneity priors rather than the degraded LQ features, thereby substantially reducing the learning difficulty and facilitate model convergence. The design of HGA is detailed below.
HGA is highly flexible and can be implemented on either standard self-attention or transposed self-attention. For clarity of exposition, we formulate it here using standard self-attention. Let the query, key, and value be , the conventional self-attention output is
| (5) |
To incorporate guidance from the homogeneity prior , a straightforward variant is to complement the LQ value with clean by direct addition:
| (6) |
In Eq. 6, the attention mechanism treats and symmetrically. However, the homogeneity prior contains higher-fidelity information than the degraded observation , the attention mechanism should preferentially exploit the more reliable . To encourage such a preference, we introduce a second variant that biases attention away from the LQ value and toward the homogeneity prior by adding identity-based terms to the attention map :
| (7) | ||||
where denotes the identity matrix. The terms increase the self-contribution of while reducing that of . denotes the mixed value, and acts as a preference bias that reinforces more reliance on the homogeneity prior . To stabilize training and increase model expressivity, the final HGA mechanism is obtained by applying channel-wise learnable weighting parameters :
| (8) |
When , the HGA reduces to the conventional self-attention. The self-attention–based HGA in Eq. 8 has an analogous form to transposed self-attention (see supplement). Our proposed UniH3 adopts the HGA based on transposed self-attention following Restormer [70]. Fig. 3 illustrates the HGA formulation, which augments attention with simple addition and subtraction operations on the value.
3.3 Hierarchical Heterogeneity Balancer
We propose a Hierarchical Heterogeneity Balancer (H2B) to mitigate inter- and intra-task heterogeneity across diverse MedIR tasks during the optimization process. Heterogeneity among tasks induces gradient conflicts that create an imbalance in optimization: some tasks dominate training while others remain under-trained. Previous work in multi-task learning [26] and all-in-one natural image restoration [58] has shown that uncertainty-based loss balancing is a good way of addressing inter-task heterogeneity by dynamically scaling different task losses for a reasonable optimization route:
| (9) |
where denotes the number of tasks, denotes the reconstruction loss, and is a learnable scalar that estimates task-level uncertainty. The factor adaptively rescales each task’s contribution while the term regularizes the scaling. When increases and tends to dominate the total loss, increases to attenuate its contribution, and vice versa. However, this uncertainty balancing is too coarse: a single scalar per-task cannot capture intra-task heterogeneity (e.g., scanner/center/anatomy variations), so hard samples still remain insufficiently handled. We therefore introduce a hierarchical uncertainty model. For task and sample we define the total uncertainty as the sum of a global task term and a sample-specific correction :
| (10) |
where indexes tasks and indexes samples. is still a learnable scalar for each task while is predicted by a lightweight Uncertainty Estimation Block (UEB, see Fig. 1(c)) conditioned on sample-specific signals:
| (11) |
where is the LQ input for sample , is the model prediction, is the HQ ground truth, and denotes stop-gradient to decouple loss balancing from restoration model optimization. The H2B loss then aggregates per-task and per-sample contributions as:
| (12) |
where is the batch size. H2B retains the theoretical foundation of standard uncertainty-based balancing [26] while refining it to capture uncertainty at two hierarchical levels: a task-level term to effectively mitigate inter-task heterogeneity, and a sample-level correction to mitigate intra-task heterogeneity.
4 Experiments
We conduct experiments under two settings, All-in-One and Single-Task, for both 2D and 3D MedIR tasks. In the All-in-One setting, a single universal model is trained to address multiple MedIR tasks within either the 2D or 3D domain. In the Single-Task setting, separate models are trained for each MedIR task. We first describe the experimental setup, including datasets, implementation details, and evaluation. We then present comparative results in Sec. 4.1 and Sec. 4.2, and ablation studies in Sec. 4.3.
Datasets. Most existing MedIR datasets are limited in size and narrowly tailored to specific tasks and modalities. To promote the development of general-purpose MedIR methods, we organize publicly available datasets together with private collections into two datasets, as summarized in Tab. 1: (i) MedIR-2D-500K comprises 509,200 2D LQ-HQ image pairs across seven distinct 2D MedIR tasks: PET image denoising, CT image denoising, MRI image super-resolution, X-ray image denoising, OCT image denoising, ultrasound image denoising, and pathological image super-resolution. (ii) MedIR-3D-3K includes 3,522 3D LQ-HQ volume pairs covering three 3D MedIR tasks: PET image denoising, CT image denoising, and MRI image super-resolution. We expect these two datasets to serve as a useful benchmark for advancing general-purpose MedIR research. More detailed descriptions are shown in the supplement.
| Dataset | Dimension | Modality | Training | Testing | Total | Data Source |
| PET | 77,000 | 8,600 | 85,600 | [61], Private | ||
| CT | 64,000 | 8,100 | 72,100 | [36, 37] , Private | ||
| MRI | 75,000 | 8,400 | 83,400 | [49, 35] | ||
| X-ray | 108,000 | 11,600 | 119,600 | [54, 10, 42, 52] | ||
| OCT | 32,000 | 3,600 | 35,600 | [32, 17, 16] | ||
| Ultrasound | 56,000 | 5,800 | 61,800 | [20, 38, 48, 28, 62] | ||
| Pathology | 46,000 | 5,100 | 51,100 | [15, 27, 12, 43, 1, 45] | ||
| MedIR-2D-500K | 2D | Total | 458,000 | 51,200 | 509,200 | - |
| PET | 1,388 | 156 | 1,544 | [61], Private | ||
| CT | 258 | 30 | 288 | [36, 37], Private | ||
| MRI | 1,520 | 170 | 1,690 | [49, 35] | ||
| MedIR-3D-3K | 3D | Total | 3,166 | 356 | 3,522 | - |
Implementation. For the UniH3 architecture, the number of HGATBs are , , and . The input channel dimension is . For the H2M module, we set the the number of tasks , memory length and EMA coefficient . During training we use patches of size with a batch size of 14. The reconstruction loss is defined as the L1 loss. The model is optimized using Muon optimizer [34] for iterations, with an initial learning rate and annealed to using a cosine schedule. To support 3D MedIR, we introduce a 3D variant, UniH3-3D, obtained by replacing each module in UniH3 with its 3D counterpart. For UniH3-3D, the HGATB numbers are and , the input channel number is , and the init learning rate is set as . The patch size is and the batch size is 6. All other settings are identical to the original UniH3.
Evaluation. To quantitatively assess image quality, we employ the widely used PSNR and SSIM metrics. In the reported tables, the highest and second-highest scores are indicated in red and blue, respectively.
| PET | CT | MRI | X-ray | OCT | Ultrasound | Pathology | Average | |||||||||||||||||||||
| Method | #Params (M) | FLOPs (G) | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | ||||||||||
| SwinIR [33] | 11.50 | 187.93 | 44.24 | 0.9866 | 43.17 | 0.9338 | 38.44 | 0.9453 | 35.87 | 0.9275 | 35.70 | 0.8892 | 27.52 | 0.8089 | 28.35 | 0.7941 | 36.18 | 0.8979 | ||||||||||
| Uformer [57] | 50.47 | 21.42 | 44.51 | 0.9874 | 43.45 | 0.9356 | 39.01 | 0.9507 | 36.59 | 0.9353 | 35.84 | 0.8908 | 27.61 | 0.8135 | 28.55 | 0.8011 | 36.51 | 0.9020 | ||||||||||
| Restormer [70] | 26.12 | 35.21 | 44.47 | 0.9873 | 43.44 | 0.9353 | 39.05 | 0.9515 | 36.46 | 0.9335 | 35.84 | 0.8909 | 27.66 | 0.8138 | 28.53 | 0.8006 | 36.49 | 0.9018 | ||||||||||
| NAFNet [6] | 67.89 | 15.74 | 44.40 | 0.9871 | 43.32 | 0.9346 | 38.90 | 0.9502 | 36.33 | 0.9326 | 35.82 | 0.8905 | 27.59 | 0.8131 | 28.49 | 0.7995 | 36.41 | 0.9011 | ||||||||||
| Restore-RWKV [65] | 27.91 | 37.46 | 44.45 | 0.9872 | 43.46 | 0.9352 | 39.05 | 0.9514 | 36.46 | 0.9333 | 35.84 | 0.8909 | 27.68 | 0.8149 | 28.55 | 0.8011 | 36.50 | 0.9020 | ||||||||||
| MambaIR [19] | 31.50 | 34.35 | 44.50 | 0.9873 | 43.47 | 0.9355 | 39.07 | 0.9517 | 36.57 | 0.9341 | 35.83 | 0.8904 | 27.69 | 0.8146 | 28.53 | 0.8004 | 36.52 | 0.9020 | ||||||||||
| TransWeather [47] | 38.05 | 1.55 | 43.73 | 0.9846 | 41.12 | 0.9217 | 37.62 | 0.9333 | 35.28 | 0.9221 | 35.12 | 0.8699 | 27.23 | 0.8011 | 28.17 | 0.7874 | 35.47 | 0.8886 | ||||||||||
| AirNet [31] | 7.61 | 230.48 | 44.32 | 0.9868 | 43.34 | 0.9348 | 38.81 | 0.9489 | 36.38 | 0.9322 | 35.77 | 0.8900 | 27.62 | 0.8127 | 28.46 | 0.7981 | 36.39 | 0.9005 | ||||||||||
| DRMC [68] | 0.62 | 9.92 | 43.62 | 0.9841 | 42.48 | 0.9281 | 37.24 | 0.9301 | 34.06 | 0.9085 | 35.46 | 0.8862 | 27.09 | 0.7993 | 28.02 | 0.7830 | 35.42 | 0.8885 | ||||||||||
| AMIR [63] | 23.54 | 31.76 | 44.49 | 0.9873 | 43.47 | 0.9356 | 39.09 | 0.9519 | 36.47 | 0.9333 | 35.89 | 0.8914 | 27.69 | 0.8150 | 28.57 | 0.8019 | 36.52 | 0.9023 | ||||||||||
| PromptIR [41] | 35.59 | 39.49 | 44.52 | 0.9874 | 43.48 | 0.9355 | 39.13 | 0.9524 | 36.57 | 0.9341 | 35.84 | 0.8909 | 27.69 | 0.8152 | 28.54 | 0.8006 | 36.54 | 0.9023 | ||||||||||
| AdaIR [11] | 28.76 | 36.74 | 44.55 | 0.9875 | 43.49 | 0.9356 | 39.17 | 0.9527 | 36.60 | 0.9344 | 35.86 | 0.8915 | 27.69 | 0.8150 | 28.56 | 0.8015 | 36.56 | 0.9026 | ||||||||||
| UniH3 (Ours) | 28.96 | 26.33 | 44.89 | 0.9883 | 43.65 | 0.9368 | 39.55 | 0.9564 | 36.88 | 0.9368 | 35.96 | 0.8921 | 27.80 | 0.8179 | 28.63 | 0.8035 | 36.77 | 0.9045 | ||||||||||
| PET | CT | MRI | X-ray | OCT | Ultrasound | Pathology | Average | |||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | #Params (M) | FLOPs (G) | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | ||||||||||
| SwinIR [33] | 11.50 | 187.93 | 44.08 | 0.9860 | 43.53 | 0.9358 | 39.19 | 0.9529 | 36.30 | 0.9308 | 35.80 | 0.8904 | 27.68 | 0.8143 | 28.50 | 0.7993 | 36.44 | 0.9014 | ||||||||||
| Uformer [57] | 50.47 | 21.42 | 44.61 | 0.9876 | 43.54 | 0.9362 | 39.21 | 0.9526 | 36.61 | 0.9363 | 35.91 | 0.8915 | 27.64 | 0.8151 | 28.58 | 0.8022 | 36.59 | 0.9031 | ||||||||||
| Restormer [70] | 26.12 | 35.21 | 44.90 | 0.9883 | 43.64 | 0.9369 | 39.44 | 0.9554 | 36.71 | 0.9359 | 35.97 | 0.8921 | 27.76 | 0.8164 | 28.63 | 0.8038 | 36.72 | 0.9041 | ||||||||||
| NAFNet [6] | 67.89 | 15.74 | 44.74 | 0.9880 | 43.55 | 0.9362 | 39.29 | 0.9540 | 36.48 | 0.9343 | 35.93 | 0.8919 | 27.67 | 0.8139 | 28.58 | 0.8024 | 36.61 | 0.9030 | ||||||||||
| Restore-RWKV [65] | 27.91 | 37.46 | 44.93 | 0.9884 | 43.64 | 0.9369 | 39.57 | 0.9565 | 36.77 | 0.9359 | 35.97 | 0.8923 | 27.76 | 0.8168 | 28.62 | 0.8036 | 36.75 | 0.9043 | ||||||||||
| MambaIR [19] | 31.50 | 34.35 | 44.93 | 0.9883 | 43.55 | 0.9363 | 39.59 | 0.9568 | 36.77 | 0.9363 | 35.99 | 0.8923 | 27.77 | 0.8178 | 28.64 | 0.8041 | 36.75 | 0.9046 | ||||||||||
| UniH3 (Ours) | 28.96 | 26.33 | 45.11 | 0.9889 | 43.77 | 0.9376 | 39.78 | 0.9584 | 37.05 | 0.9384 | 36.04 | 0.8929 | 27.86 | 0.8192 | 28.69 | 0.8050 | 36.90 | 0.9058 | ||||||||||
| PET | CT | MRI | Average | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | ||||
| 3D-cGAN [56] | 48.22 | 0.9937 | 40.82 | 0.9320 | 38.57 | 0.9489 | 42.54 | 0.9582 | ||||
| MRDG [53] | 48.68 | 0.9943 | 42.86 | 0.9388 | 38.93 | 0.9519 | 43.49 | 0.9617 | ||||
| DRMC [68] | 48.58 | 0.9933 | 43.33 | 0.9391 | 38.81 | 0.9514 | 43.57 | 0.9613 | ||||
| Spach Transformer [25] | 48.91 | 0.9949 | 42.21 | 0.9380 | 39.04 | 0.9542 | 43.39 | 0.9624 | ||||
| Restore-RWKV-3D [65] | 49.08 | 0.9951 | 43.35 | 0.9402 | 39.43 | 0.9578 | 43.95 | 0.9644 | ||||
| UniH3-3D (Ours) | 49.43 | 0.9955 | 44.05 | 0.9429 | 40.03 | 0.9637 | 44.50 | 0.9674 | ||||
| PET | CT | MRI | Average | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | ||||
| 3D-cGAN [56] | 48.47 | 0.9944 | 40.86 | 0.9309 | 38.76 | 0.9510 | 42.70 | 0.9588 | ||||
| MRDG [53] | 48.40 | 0.9942 | 43.14 | 0.9395 | 39.34 | 0.9570 | 43.63 | 0.9636 | ||||
| DRMC [68] | 48.76 | 0.9947 | 43.60 | 0.9404 | 38.98 | 0.9540 | 43.78 | 0.9630 | ||||
| Spach Transformer [25] | 48.71 | 0.9946 | 43.57 | 0.9409 | 39.16 | 0.9546 | 43.81 | 0.9634 | ||||
| Restore-RWKV-3D [65] | 48.67 | 0.9945 | 42.48 | 0.9375 | 39.15 | 0.9551 | 43.43 | 0.9624 | ||||
| UniH3-3D (Ours) | 49.77 | 0.9958 | 45.37 | 0.9448 | 40.73 | 0.9689 | 45.29 | 0.9698 | ||||
4.1 All-in-One MedIR Results
2D All-in-One MedIR. We evaluate the 2D All-in-One setting on the MedIR-2D-500K dataset. We compare it to several general image-restoration methods (SwinIR [33], Uformer [57], Restormer [70], NAFNet [6], Restore-RWKV [65], and MambaIR [19]) and to SOTA all-in-one approaches (TransWeather [47], AirNet [31], DRMC [68], AMIR [63], PromptIR [41], and AdaIR [11]). Tab. 2 shows that UniH3 achieves both high efficiency and strong effectiveness. It significantly outperforms all comparison methods across all seven tasks. On average, UniH3 surpasses the second-best AdaIR by 0.21 dB in PSNR, which is an appreciable improvement given that each of the seven task contains thousands of testing images. Visual comparison in Fig. 4 demonstrates that UniH3 best restores different types of medical images with higher structural fidelity and finer detail than competing methods.
3D All-in-One MedIR. We assess the UniH3-3D in the 3D All-in-One setting on the MedIR-3D-3K dataset. We compare it with several SOTA 3D restoration methods, including 3D-cGAN [56], MRDG [53], DRMC [68], Spach Transformer [25], and Restore-RWKV-3D [65]. Tab. 4 shows that UniH3 significantly outperforms all comparison methods across the three 3D MedIR tasks. In particular, UniH3-3D exceeds the second-best Restore-RWKV-3D by an average margin of 0.55 dB in PSNR. Visual comparisons are shown in the supplement.
| PET | CT | MRI | X-ray | OCT | Ultrasound | Pathology | Average | |||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | #Params (M) | FLOPs (G) | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | ||||||||||
| Uformer | 50.47 | 21.42 | 44.51 | 0.9874 | 43.45 | 0.9356 | 39.01 | 0.9507 | 36.59 | 0.9353 | 35.84 | 0.8908 | 27.61 | 0.8135 | 28.55 | 0.8011 | 36.51 | 0.9021 | ||||||||||
| Uformer | 53.15 | 22.10 | 44.80 | 0.9880 | 43.59 | 0.9363 | 39.39 | 0.9549 | 36.74 | 0.9354 | 35.94 | 0.8918 | 27.72 | 0.8155 | 28.57 | 0.8022 | 36.68 | 0.9034 | ||||||||||
| Restormer | 26.12 | 35.21 | 44.47 | 0.9873 | 43.44 | 0.9353 | 39.05 | 0.9515 | 36.46 | 0.9335 | 35.84 | 0.8909 | 27.66 | 0.8138 | 28.53 | 0.8006 | 36.49 | 0.9018 | ||||||||||
| Restormer | 27.56 | 35.83 | 44.78 | 0.9880 | 43.58 | 0.9362 | 39.37 | 0.9548 | 36.71 | 0.9352 | 35.93 | 0.8919 | 27.73 | 0.8159 | 28.60 | 0.8027 | 36.67 | 0.9035 | ||||||||||
| PromptIR | 35.59 | 39.49 | 44.52 | 0.9874 | 43.48 | 0.9355 | 39.13 | 0.9524 | 36.57 | 0.9341 | 35.84 | 0.8909 | 27.69 | 0.8152 | 28.54 | 0.8006 | 36.54 | 0.9023 | ||||||||||
| PromptIR | 36.15 | 40.68 | 44.85 | 0.9882 | 43.62 | 0.9366 | 39.40 | 0.9551 | 36.76 | 0.9365 | 35.94 | 0.8919 | 27.75 | 0.8165 | 28.59 | 0.8024 | 36.70 | 0.9039 | ||||||||||
| AdaIR | 28.76 | 36.74 | 44.55 | 0.9875 | 43.49 | 0.9356 | 39.17 | 0.9527 | 36.60 | 0.9344 | 35.86 | 0.8915 | 27.69 | 0.8150 | 28.56 | 0.8015 | 36.56 | 0.9026 | ||||||||||
| AdaIR | 29.77 | 37.37 | 44.93 | 0.9884 | 43.64 | 0.9367 | 39.53 | 0.9563 | 36.82 | 0.9363 | 35.95 | 0.8920 | 27.76 | 0.8165 | 28.61 | 0.8027 | 36.75 | 0.9041 | ||||||||||
| Baseline | 27.17 | 25.50 | 44.52 | 0.9874 | 43.47 | 0.9355 | 39.10 | 0.9521 | 36.39 | 0.9330 | 35.88 | 0.8914 | 27.68 | 0.8149 | 28.57 | 0.8018 | 36.52 | 0.9023 | ||||||||||
| UniH3 (Ours) | 28.96 | 26.33 | 44.89 | 0.9883 | 43.65 | 0.9368 | 39.55 | 0.9564 | 36.88 | 0.9368 | 35.96 | 0.8921 | 27.80 | 0.8179 | 28.63 | 0.8035 | 36.77 | 0.9045 | ||||||||||
| H2M | H2B | #Params (M) | FLOPs (G) | PSNR↑ | SSIM↑ | |||||
|---|---|---|---|---|---|---|---|---|---|---|
| 27.17 | 25.50 | 36.52 | 0.9023 | |||||||
| 28.96 | 26.33 | 36.66 | 0.9034 | |||||||
| 27.17 | 25.50 | 36.64 | 0.9033 | |||||||
| 28.96 | 26.33 | 36.77 | 0.9045 |
|
| #Params (M) | FLOPs (G) | PSNR↑ | SSIM↑ | ||||
|---|---|---|---|---|---|---|---|---|---|
| 28.96 | 26.33 | 36.52 | 0.9023 | ||||||
| 28.96 | 26.33 | 36.61 | 0.9031 | ||||||
| 28.96 | 26.33 | 36.72 | 0.9038 | ||||||
| 28.96 | 26.33 | 36.77 | 0.9045 |
|
| #Params (M) | FLOPs (G) | PSNR↑ | SSIM↑ | ||||
|---|---|---|---|---|---|---|---|---|---|
| 28.96 | 26.33 | 36.66 | 0.9034 | ||||||
| 28.96 | 26.33 | 36.70 | 0.9036 | ||||||
| 28.96 | 26.33 | 36.73 | 0.9041 | ||||||
| 28.96 | 26.33 | 36.77 | 0.9045 |
4.2 Single-Task MedIR Results
2D Single-Task MedIR. We evaluate UniH3 for 2D single-task MedIR on the MedIR-2D-500K dataset, comparing it to five general image restoration methods. As shown in Tab. 3, UniH3 significantly outperforms all comparison methods. On average across seven tasks, UniH3 improves PSNR by 0.15 dB over the second-best MambaIR.
3D Single-Task MedIR. We evaluate 3D single-task MedIR on the MedIR-3D-3K dataset and compare UniH3-3D to five SOTA 3D restoration methods. As shown in Tab. 5, UniH3-3D beats all comparison methods. In particular, it improves average PSNR by 1.48 dB over the second-best Spach Transformer [25] across the three tasks.
4.3 Ablation Studies
To evaluate the effectiveness of individual components, we conduct ablation experiments on the 2D All-in-One MedIR task with the MedIR-2D-500K dataset.
Component Analysis of H2M and H2B. We first perform a component analysis of the H2M module and the H2B strategy by selectively disabling each component. To disable H2M, we remove the H2M module and replace the HGA mechanism with a transposed self-attention layer [70]. The H2B is disabled by substituting the loss term with an loss. Table 8 shows that both components improve model performance with minimal increase in computational cost. This finding is further supported by the visual comparison in Fig. 6, where both components contribute to better preservation of image details. Moreover, we apply the H2M module and H2B strategy to other Transformer-based U-shaped backbones, including Uformer, Restormer, PromptIR, and AdaIR. Results in the Tab. 6 demonstrate significant improvements across all these backbones.
Ablation Studies on H2M. We investigate the effectiveness of both inter- and intra-task homogeneity priors in H2M by selectively disabling the task-specific and task-shared slots in . By replacing these slots with naive learnable parameters—thus isolating them from the HQ Homogeneity Distillation procedure—we observe a drop in performance in Tab. 10. Results indicate that both priors independently benefit the restoration process, and their combination achieves the best performance. To further understand the internal mechanics of H2M, Fig. 6 visualizes the retrieval attention maps for PET and CT tokens. The distributions confirm that tokens primarily query the task-shared slot and their specific modality slot. Furthermore, attention maps reflect clear semantic correlations: anatomically similar tokens within the same modality share highly similar query patterns (8/10 top-score overlap for two PET spine tokens), and cross-modality similarities are also captured (2/10 overlap for PET and CT spine tokens). In contrast, dissimilar anatomies exhibit distinct query patterns (0/10 overlap for PET spine and lesion tokens). This provides strong visual evidence that H2M successfully maps, stores, and retrieves hierarchical anatomical priors.
Ablation Studies on HGA. We validate the impact of HGA by replacing it with alternative guiding mechanisms, including Spatial Feature Transform (SFT) [55] and cross attention [11]. Tab. 8 shows that the proposed HGA performs the best with minimal computation and parameter increase.
Ablation Studies on H2B. We validate the roles of intra- and inter-task heterogeneity in H2B in Tab. 10. The results indicate that mitigating both levels of heterogeneity improves overall restoration performance, with their joint optimization achieving the best results. In Fig. 8, we visualize the estimated uncertainty distributions. While the conventional [26] estimates an uncertainty per task and therefore handles only inter-task heterogeneity, our proposed estimates uncertainty at both the task and the sample level: each sample receives its own uncertainty while each task exhibits a distinct uncertainty distribution. This enables our to capture finer-grained relationships, i.e., both inter-task and intra-task heterogeneity, facilitating convergence toward a more optimal solution for diverse MedIR tasks.
5 Discussion and Limitation
All-in-one medical image restoration is an emerging research field. The practical value of all-in-one models remains under active discussion. In this paper, experiments on a large-scale dataset show that a single all-in-one model, UniH3, already achieves comparable performance to SOTA single-task MambaIR models across seven MedIR tasks (see Fig. 8). This result strongly supports the practicality of developing all-in-one models instead of separate single-task models for MedIR tasks. Additionally, by fine-tuning the well-trained all-in-one UniH3 model for each task, we are able to transfer its learned knowledge to individual tasks and achieve further improvements (also shown in Fig. 8), indicating that the all-in-one model can serve as a transferable pretrained backbone. Our study has limitations: we focus only on the primary restoration task within each modality and therefore do not cover other tasks or degradation types that may occur in the same modality. Future work should address these gaps and pursue more universal MedIR models to benefit clinical diagnosis and more downstream tasks [75, 76, 59, 60, 21, 7, 72, 74, 71, 73].
6 Conclusion
In this paper, we have presented UniH3, a novel and unified framework for all-in-one medical image restoration. By moving beyond the conventional focus on inter-task heterogeneity, UniH3 effectively leverages the inherent hierarchical homogeneity across diverse medical imaging modalities. The integration of the Hierarchical Homogeneity Memory (H2M) module and the Homogeneity-Guided Attention (HGA) mechanism allows the model to distill and utilize inter- and intra-task homogeneity priors to guide the restoration process. Simultaneously, the Hierarchical Heterogeneity Balancer (H2B) ensures a stable and balanced multi-task learning process by mitigating both inter- and intra-task heterogeneity. Extensive evaluations on the large-scale MedIR-2D-500K and MedIR-3D-3K benchmarks demonstrate that UniH3 significantly outperforms existing methods, achieving state-of-the-art performance in both all-in-one and single-task scenarios. In future work, we plan to expand the task spectrum by incorporating additional modalities and degradation types, moving toward more general MedIR models.
Acknowledgements
This work is supported by the National Natural Science Foundation in China under Grant U23B2063 and 62371016, the Bejing Natural Science Foundation Haidian District Joint Fund in China under Grant L2602042, the Beijing hope run special fund of cancer foundation of China under Grant LC2018L02, the Fundamental Research Funds for the Central University of China from the State Key Laboratory of Software Development Environment in Beihang University in China, the 111 Proiect in China under Grant B13003, the Academic Excellence Foundation of BUAA for PhD Students.
References
- [1] A. Aksac, D. J. Demetrick, T. Ozyer, and R. Alhajj (2019) BreCaHAD: a dataset for breast cancer histopathological annotation and diagnosis. BMC research notes 12 (1), pp. 82. Cited by: Table I, Table 1.
- [2] H. Asgariandehkordi, S. Goudarzi, A. Basarab, and H. Rivaz (2023) Deep ultrasound denoising using diffusion probabilistic models. In 2023 IEEE International Ultrasonics Symposium (IUS), pp. 1–4. Cited by: §2.
- [3] R. Caruana (1997) Multitask learning. Machine learning 28 (1), pp. 41–75. Cited by: §3.1.
- [4] H. Chen, Z. Yang, H. Hou, H. Zhang, B. Wei, G. Zhou, and Y. Xu (2025) All-in-one medical image restoration with latent diffusion-enhanced vector-quantized codebook prior. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 67–77. Cited by: §1, §2.
- [5] H. Chen, Y. Zhang, M. K. Kalra, F. Lin, Y. Chen, P. Liao, J. Zhou, and G. Wang (2017) Low-dose ct with a residual encoder-decoder convolutional neural network. IEEE Transactions on Medical Imaging 36 (12), pp. 2524–2535. Cited by: §1, §2.
- [6] L. Chen, X. Chu, X. Zhang, and J. Sun (2022) Simple baselines for image restoration. In European conference on computer vision, pp. 17–33. Cited by: §4.1, Table 2, Table 3.
- [7] L. Chen (2026) Beyond external constraints: the missing dimension of ai governance. Available at SSRN 6449738. Cited by: §5.
- [8] T. Chen, X. Ma, L. Bai, W. Wang, Y. Sun, and L. Zhou (2025) EndoIR: degradation-agnostic all-in-one endoscopic image restoration via noise-aware routing diffusion. arXiv preprint arXiv:2511.05873. Cited by: §1, §2.
- [9] Y. Chen, F. Shi, A. G. Christodoulou, Y. Xie, Z. Zhou, and D. Li (2018) Efficient and accurate mri super-resolution using a generative adversarial network and 3d multi-level densely connected network. In International conference on medical image computing and computer-assisted intervention, pp. 91–99. Cited by: §1, §2.
- [10] M. E. Chowdhury, T. Rahman, A. Khandakar, R. Mazhar, M. A. Kadir, Z. B. Mahbub, K. R. Islam, M. S. Khan, A. Iqbal, N. Al Emadi, et al. (2020) Can ai help in screening viral and covid-19 pneumonia?. Ieee Access 8, pp. 132665–132676. Cited by: Table I, Table 1.
- [11] Y. Cui, S. W. Zamir, S. Khan, A. Knoll, M. Shah, and F. S. Khan (2025) Adair: adaptive all-in-one image restoration via frequency mining and modulation. In 13th International Conference on Learning Representations, ICLR 2025, pp. 57335–57356. Cited by: §1, §2, §3.2, §4.1, §4.3, Table 2, Table 8.
- [12] Q. Da, X. Huang, Z. Li, Y. Zuo, C. Zhang, J. Liu, W. Chen, J. Li, D. Xu, Z. Hu, et al. (2022) DigestPath: a benchmark dataset with challenge review for the pathological detection and segmentation of digestive-system. Medical image analysis 80, pp. 102485. Cited by: Table I, Table 1.
- [13] Z. Dong, G. Liu, G. Ni, J. Jerwick, L. Duan, and C. Zhou (2020) Optical coherence tomography image denoising using a generative adversarial network with speckle modulation. Journal of biophotonics 13 (4), pp. e201960135. Cited by: §0.C.2, §2.
- [14] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929. Cited by: §3.
- [15] C. R. Drifka, A. G. Loeffler, K. Mathewson, A. Keikhosravi, J. C. Eickhoff, Y. Liu, S. M. Weber, W. J. Kao, and K. W. Eliceiri (2016) Highly aligned stromal collagen is a negative prognostic factor following pancreatic ductal adenocarcinoma resection. Oncotarget 7 (46), pp. 76197. Cited by: Table I, Table 1.
- [16] L. Fang, S. Li, Q. Nie, J. A. Izatt, C. A. Toth, and S. Farsiu (2012) Sparsity based denoising of spectral domain optical coherence tomography images. Biomedical optics express 3 (5), pp. 927–942. Cited by: Table I, Table 1.
- [17] M. Geng, X. Meng, L. Zhu, Z. Jiang, M. Gao, Z. Huang, B. Qiu, Y. Hu, Y. Zhang, Q. Ren, et al. (2022) Triplet cross-fusion learning for unpaired image denoising in optical coherence tomography. IEEE Transactions on Medical Imaging 41 (11), pp. 3357–3372. Cited by: Table I, §0.C.2, Table 1.
- [18] A. Gu and T. Dao (2023) Mamba: linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752. Cited by: §2.
- [19] H. Guo, J. Li, T. Dai, Z. Ouyang, X. Ren, and S. Xia (2024) Mambair: a simple baseline for image restoration with state-space model. In European conference on computer vision, pp. 222–241. Cited by: §3, §4.1, Table 2, Table 3.
- [20] Y. Guo, S. Zhou, J. Shi, and Y. Wang (2023) Ultrasound image enhancement challenge 2023. Zenodo. Note: accessed: 2024-02-13 External Links: Link Cited by: Table I, Table 1.
- [21] R. Hao, B. Jing, H. Yu, and Z. Nie (2026) Styledrive: towards driving-style aware benchmarking of end-to-end autonomous driving. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 4627–4635. Cited by: §5.
- [22] J. Hu, L. Jin, Z. Yao, and Y. Lu (2025) Universal image restoration pre-training via degradation classification. arXiv preprint arXiv:2501.15510. Cited by: §1.
- [23] J. Hu, L. Shen, and G. Sun (2018) Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7132–7141. Cited by: §3.
- [24] J. Huang, Y. Fang, Y. Wu, H. Wu, Z. Gao, Y. Li, J. Del Ser, J. Xia, and G. Yang (2022) Swin transformer for fast mri. Neurocomputing 493, pp. 281–304. Cited by: §1, §2.
- [25] S. Jang, T. Pan, Y. Li, P. Heidari, J. Chen, Q. Li, and K. Gong (2023) Spach transformer: spatial and channel-wise transformer based on local and global self-attentions for pet image denoising. IEEE Transactions on Medical Imaging. Cited by: §4.1, §4.2, Table 4, Table 5.
- [26] A. Kendall, Y. Gal, and R. Cipolla (2018) Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7482–7491. Cited by: §1, §3.3, §3.3, §4.3.
- [27] N. Kumar, R. Verma, D. Anand, Y. Zhou, O. F. Onder, E. Tsougenis, H. Chen, P. Heng, J. Li, Z. Hu, et al. (2019) A multi-organ nucleus segmentation challenge. IEEE transactions on medical imaging 39 (5), pp. 1380–1391. Cited by: Table I, Table 1.
- [28] S. Leclerc, E. Smistad, J. Pedrosa, A. Østvik, F. Cervenansky, F. Espinosa, T. Espeland, E. A. R. Berg, P. Jodoin, T. Grenier, et al. (2019) Deep learning for segmentation using an open large-scale dataset in 2d echocardiography. IEEE transactions on medical imaging 38 (9), pp. 2198–2210. Cited by: Table I, Table 1.
- [29] Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel (1989) Backpropagation applied to handwritten zip code recognition. Neural Computation 1 (4), pp. 541–551. Cited by: §2.
- [30] B. Li, A. Keikhosravi, A. G. Loeffler, and K. W. Eliceiri (2021) Single image super-resolution for whole slide image using convolutional neural networks and self-supervised color normalization. Medical Image Analysis 68, pp. 101938. Cited by: §2.
- [31] B. Li, X. Liu, P. Hu, Z. Wu, J. Lv, and X. Peng (2022) All-in-one image restoration for unknown corruption. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 17452–17462. Cited by: §1, §2, §4.1, Table 2.
- [32] M. Li, K. Huang, Q. Xu, J. Yang, Y. Zhang, Z. Ji, K. Xie, S. Yuan, Q. Liu, and Q. Chen (2024) OCTA-500: a retinal dataset for optical coherence tomography angiography study. Medical image analysis 93, pp. 103092. Cited by: Table I, Table 1.
- [33] J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte (2021) Swinir: image restoration using swin transformer. In Proceedings of the IEEE International Conference on Computer Vision, pp. 1833–1844. Cited by: §4.1, Table 2, Table 3.
- [34] J. Liu, J. Su, X. Yao, Z. Jiang, G. Lai, Y. Du, Y. Qin, W. Xu, E. Lu, J. Yan, et al. (2025) Muon is scalable for llm training. arXiv preprint arXiv:2502.16982. Cited by: §4.
- [35] M. LLCIXI dataset(Website) Note: Accessed: 2024-01-15 External Links: Link Cited by: Table I, Table I, Table 1, Table 1.
- [36] C. H. McCollough, A. C. Bartley, R. E. Carter, B. Chen, T. A. Drees, P. Edwards, D. R. Holmes III, A. E. Huang, F. Khan, S. Leng, et al. (2017) Low-dose ct for the detection and classification of metastatic liver lesions: results of the 2016 low dose ct grand challenge. Medical physics 44 (10), pp. e339–e352. Cited by: Table I, Table I, §0.C.2, Table 1, Table 1.
- [37] T. R. Moen, B. Chen, D. R. Holmes III, X. Duan, Z. Yu, L. Yu, S. Leng, J. G. Fletcher, and C. H. McCollough (2021) Low-dose ct image and projection dataset. Medical physics 48 (2), pp. 902–911. Cited by: Table I, Table I, Table 1, Table 1.
- [38] A. Montoya, Hasnin, kaggle446, shirzad, W. Cukierski, and yffud (2016) Ultrasound nerve segmentation. Note: Kaggle Cited by: Table I, Table 1.
- [39] Ş. Öztürk, O. C. Duran, and T. Çukur (2024) DenoMamba: a fused state-space model for low-dose ct denoising. arXiv preprint arXiv:2409.13094. Cited by: §1, §2.
- [40] B. Peng, E. Alcaide, Q. Anthony, A. Albalak, S. Arcadinho, H. Cao, X. Cheng, M. Chung, M. Grella, K. K. GV, et al. (2023) Rwkv: reinventing rnns for the transformer era. arXiv preprint arXiv:2305.13048. Cited by: §2.
- [41] V. Potlapalli, S. W. Zamir, S. H. Khan, and F. Shahbaz Khan (2023) Promptir: prompting for all-in-one image restoration. Advances in Neural Information Processing Systems 36, pp. 71275–71293. Cited by: §1, §2, §4.1, Table 2.
- [42] T. Rahman, A. Khandakar, Y. Qiblawey, A. Tahir, S. Kiranyaz, S. B. A. Kashem, M. T. Islam, S. Al Maadeed, S. M. Zughaier, M. S. Khan, et al. (2021) Exploring the effect of image enhancement techniques on covid-19 detection using chest x-ray images. Computers in biology and medicine 132, pp. 104319. Cited by: Table I, Table 1.
- [43] K. Sirinukunwattana, J. P. Pluim, H. Chen, X. Qi, P. Heng, Y. B. Guo, L. Y. Wang, B. J. Matuszewski, E. Bruni, U. Sanchez, et al. (2017) Gland segmentation in colon histology images: the glas challenge contest. Medical image analysis 35, pp. 489–502. Cited by: Table I, Table 1.
- [44] Z. Song, Z. Qi, X. Wang, X. Zhao, Z. Shen, S. Wang, M. Fei, Z. Wang, D. Zang, D. Chen, et al. (2025) Uni-coal: a unified framework for cross-modality synthesis and super-resolution of mr images. Expert Systems with Applications 270, pp. 126241. Cited by: §1.
- [45] E. Tekin, Ç. Yazıcı, H. Kusetogullari, F. Tokat, A. Yavariabdi, L. O. Iheme, S. Çayır, E. Bozaba, G. Solmaz, B. Darbaz, et al. (2023) Tubule-u-net: a novel dataset and deep learning-based tubule segmentation framework in whole slide images of breast cancer. Scientific Reports 13 (1), pp. 128. Cited by: Table I, Table 1.
- [46] D. Thanh P. Surya et al. (2019) A review on ct and x-ray images denoising methods. Informatica 43 (2). Cited by: §0.C.2, §2.
- [47] J. M. J. Valanarasu, R. Yasarla, and V. M. Patel (2022) Transweather: transformer-based restoration of images degraded by adverse weather conditions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 2353–2363. Cited by: §1, §2, §4.1, Table 2.
- [48] T. L. van den Heuvel, D. de Bruijn, C. L. de Korte, and B. v. Ginneken (2018) Automated measurement of fetal head circumference using 2d ultrasound images. PloS one 13 (8), pp. e0200412. Cited by: Table I, Table 1.
- [49] D. C. Van Essen, S. M. Smith, D. M. Barch, T. E. Behrens, E. Yacoub, K. Ugurbil, W. H. Consortium, et al. (2013) The wu-minn human connectome project: an overview. Neuroimage 80, pp. 62–79. Cited by: Table I, Table I, Table 1, Table 1.
- [50] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin (2017) Attention is all you need. Advances in Neural Information Processing Systems 30. Cited by: §2.
- [51] D. Wang, F. Fan, Z. Wu, R. Liu, F. Wang, and H. Yu (2023) CTformer: convolution-free token2token dilated vision transformer for low-dose ct denoising. Physics in Medicine & Biology 68 (6), pp. 065012. Cited by: §1, §2.
- [52] D. Wang, X. Wang, L. Wang, M. Li, Q. Da, X. Liu, X. Gao, J. Shen, J. He, T. Shen, et al. (2023) A real-world dataset and benchmark for foundation model adaptation in medical image classification. Scientific Data 10 (1), pp. 574. Cited by: Table I, Table 1.
- [53] J. Wang, Y. Chen, Y. Wu, J. Shi, and J. Gee (2020) Enhanced generative adversarial network for 3d brain mri super-resolution. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 3627–3636. Cited by: §1, §2, §4.1, Table 4, Table 5.
- [54] X. Wang, Y. Peng, L. Lu, Z. Lu, M. Bagheri, and R. M. Summers (2017) Chestx-ray8: hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2097–2106. Cited by: Table I, Table 1.
- [55] X. Wang, K. Yu, C. Dong, and C. C. Loy (2018) Recovering realistic texture in image super-resolution by deep spatial feature transform. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 606–615. Cited by: §3.2, §4.3, Table 8.
- [56] Y. Wang, B. Yu, L. Wang, C. Zu, D. S. Lalush, W. Lin, X. Wu, J. Zhou, D. Shen, and L. Zhou (2018) 3D conditional generative adversarial networks for high-quality pet image estimation at low dose. Neuroimage 174, pp. 550–562. Cited by: §1, §2, §4.1, Table 4, Table 5.
- [57] Z. Wang, X. Cun, J. Bao, W. Zhou, J. Liu, and H. Li (2022) Uformer: a general u-shaped transformer for image restoration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 17683–17693. Cited by: §4.1, Table 2, Table 3.
- [58] G. Wu, J. Jiang, Y. Wang, K. Jiang, and X. Liu (2025) Debiased all-in-one image restoration with task uncertainty regularization. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, pp. 8386–8394. Cited by: §1, §3.3.
- [59] Y. Wu, Y. Zhou, J. Saiyin, B. Wei, M. Lai, J. Shou, and Y. Xu (2024) Attriprompter: auto-prompting with attribute semantics for zero-shot nuclei detection via visual-language pre-trained models. IEEE Transactions on Medical Imaging 44 (2), pp. 982–993. Cited by: §5.
- [60] Y. Wu, Y. Zhou, J. Saiyin, B. Wei, and Y. Xu (2025) Visual textualization for image prompted object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 20900–20910. Cited by: §5.
- [61] S. Xue, R. Guo, K. P. Bohn, J. Matzke, M. Viscione, I. Alberts, H. Meng, C. Sun, M. Zhang, M. Zhang, et al. (2022) A cross-scanner and cross-tracer deep learning method for the recovery of standard-dose imaging quality from low-dose pet. European journal of nuclear medicine and molecular imaging 49 (6), pp. 1843–1856. Cited by: Table I, Table I, §0.C.2, Table 1, Table 1.
- [62] J. Yang, X. Ding, Z. Zheng, X. Xu, and X. Li (2023) Graphecho: graph-driven unsupervised domain adaptation for echocardiogram video segmentation. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 11878–11887. Cited by: Table I, Table 1.
- [63] Z. Yang, H. Chen, Z. Qian, Y. Yi, H. Zhang, D. Zhao, B. Wei, and Y. Xu (2024) All-in-one medical image restoration via task-adaptive routing. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 67–77. Cited by: §0.C.2, Table II, Table II, Appendix 0.D, Appendix 0.D, §1, §2, §4.1, Table 2.
- [64] Z. Yang, H. Chen, Z. Qian, Y. Zhou, H. Zhang, D. Zhao, B. Wei, and Y. Xu (2024) Region attention transformer for medical image restoration. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 603–613. Cited by: §2.
- [65] Z. Yang, J. Li, H. Zhang, D. Zhao, B. Wei, and Y. Xu (2025) Restore-rwkv: efficient and effective medical image restoration with rwkv. IEEE Journal of Biomedical and Health Informatics. Cited by: §2, §3, §4.1, §4.1, Table 2, Table 3, Table 4, Table 5.
- [66] Z. Yang, J. Zhang, Y. Yi, J. Liang, B. Wei, and Y. Xu (2025) TAT: task-adaptive transformer for all-in-one medical image restoration. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 565–575. Cited by: §1, §2.
- [67] Z. Yang, Y. Zhou, H. Chen, H. Zhang, D. Zhao, B. Wei, and Y. Xu (2026) UniPET: a universal network for high-quality pet image denoising across varied dose reduction factors. Medical Image Analysis, pp. 104059. Cited by: §1, §2.
- [68] Z. Yang, Y. Zhou, H. Zhang, B. Wei, Y. Fan, and Y. Xu (2023) Drmc: a generalist model with dynamic routing for multi-center pet image synthesis. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 36–46. Cited by: §1, §2, §4.1, §4.1, Table 2, Table 4, Table 5.
- [69] E. Zamfir, Z. Wu, N. Mehta, Y. Tan, D. P. Paudel, Y. Zhang, and R. Timofte (2025) Complexity experts are task-discriminative learners for any image restoration. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 12753–12763. Cited by: §1.
- [70] S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M. Yang (2022) Restormer: efficient transformer for high-resolution image restoration. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5728–5739. Cited by: Appendix 0.B, §3.2, §4.1, §4.3, Table 2, Table 3.
- [71] K. Zhou, R. Cai, Y. Ma, Q. Tan, X. Wang, J. Li, H. P. Shum, F. W. Li, S. Jin, and X. Liang (2023) A video-based augmented reality system for human-in-the-loop muscle strength assessment of juvenile dermatomyositis. IEEE Transactions on Visualization and Computer Graphics 29 (5), pp. 2456–2466. Cited by: §5.
- [72] K. Zhou, R. Cai, L. Wang, H. P. H. Shum, and X. Liang (2026) A comprehensive survey of action quality assessment: method and benchmark. Pattern Recognition 179, pp. 113933. Cited by: §5.
- [73] K. Zhou, Z. Hao, L. Wang, and X. Liang (2025) Adaptive score alignment learning for continual perceptual quality assessment of 360-degree videos in virtual reality. IEEE Transactions on Visualization and Computer Graphics 31 (5), pp. 2880–2890. Cited by: §5.
- [74] K. Zhou, H. P. H. Shum, F. W. B. Li, X. Zhang, and X. Liang (2025) PHI: bridging domain shift in long-term action quality assessment via progressive hierarchical instruction. IEEE Transactions on Image Processing 34, pp. 3718–3732. External Links: ISSN 1057-7149 Cited by: §5.
- [75] Y. Zhou, Y. Wu, J. Saiyin, B. Wei, M. Lai, E. Chang, and Y. Xu (2024) SDPT: synchronous dual prompt tuning for fusion-based visual-language pre-trained models. In European Conference on Computer Vision, pp. 340–356. Cited by: §5.
- [76] Y. Zhou, Y. Wu, J. Saiyin, B. Wei, and Y. Xu (2026) SDPT: synchronous dual prompt tuning for visual-language pre-trained models. IEEE Transactions on Pattern Analysis and Machine Intelligence. Cited by: §5.
- [77] Y. Zhou, Z. Yang, H. Zhang, I. Eric, C. Chang, Y. Fan, and Y. Xu (2022) 3D segmentation guided style-based generative adversarial networks for pet synthesis. IEEE Transactions on Medical Imaging 41 (8), pp. 2092–2104. Cited by: §1, §2.
Appendix 0.A Availability of Code and Data
The code and data are released at https://github.com/Yaziwel/UniH3. We hope this study will contribute to the advancement of general-purpose MedIR methods.
Appendix 0.B Transposed Self-Attention-Based HGA
Let the query, key, and value be . Following Eqs.5-8 in Sec. 3.2, we derive the transposed self-attention–based [70] homogeneity-guided attention (HGA) as follows:
| (13) |
| (14) |
| (15) | ||||
| (16) |
The finally derived transposed self-attention-based HGA in Eq. 16 can also be illustrated by Fig. 3, which augments attention with simple addition and subtraction operations on the value. Therefore, Fig. 3 illustrates both the self-attention–based HGA and the transposed self-attention–based HGA.
| Dataset | Dimension | Modality | Training | Testing | Total | Data Source |
| 73,125 | 8,000 | 81,125 | [61] | |||
| 1,850 | 300 | 2,150 | Private1 | |||
| PET | 2,025 | 300 | 2,325 | Private2 | ||
| PET Total | 77,000 | 8,600 | 85,600 | - | ||
| 4,470 | 1,466 | 5,936 | [36] | |||
| 22,995 | 2,534 | 25,529 | [37] | |||
| CT | 36,535 | 4,100 | 40,635 | Private3 | ||
| CT Total | 64,000 | 8,100 | 72,100 | - | ||
| 59,600 | 6,660 | 66,260 | [49] | |||
| MRI | 15,400 | 1,740 | 17,140 | [35] | ||
| MRI Total | 75,000 | 8,400 | 83,400 | - | ||
| 100,600 | 11,000 | 111,600 | [54] | |||
| 3,000 | 200 | 3,200 | [10, 42] | |||
| X-ray | 4,400 | 400 | 4,800 | [52] | ||
| X-ray Total | 108,000 | 11,600 | 119,600 | - | ||
| 31932 | 3592 | 35,524 | [32] | |||
| 33 | 4 | 37 | [17] | |||
| OCT | 35 | 4 | 39 | [16] | ||
| OCT Total | 32000 | 3600 | 35,600 | - | ||
| 1,950 | 490 | 2,440 | [20] | |||
| 8,900 | 2,220 | 11,120 | [38] | |||
| 1,050 | 260 | 1,310 | [48] | |||
| 13,600 | 1,680 | 15,280 | [28] | |||
| Ultrasound | 30,500 | 1,150 | 31,650 | [62] | ||
| Ultrasound Total | 56,000 | 5,800 | 61,800 | - | ||
| 13200 | 1,830 | 15,030 | [15] | |||
| 350 | 90 | 440 | [27] | |||
| 23100 | 2,300 | 25,400 | [12] | |||
| 400 | 35 | 435 | [43] | |||
| 7650 | 695 | 8,345 | [1] | |||
| Pathology | 1,300 | 150 | 1,450 | [45] | ||
| Pathology Total | 46000 | 5100 | 51,100 | - | ||
| MedIR-2D-500K | 2D | Total | 458,000 | 51,200 | 509,200 | |
| 1,233 | 138 | 1,371 | [61] | |||
| 74 | 9 | 83 | Private1 | |||
| PET | 81 | 9 | 90 | Private2 | ||
| PET Total | 1,388 | 156 | 1,544 | - | ||
| 8 | 2 | 10 | [36] | |||
| 135 | 15 | 150 | [37] | |||
| CT | 115 | 13 | 128 | Private3 | ||
| CT Total | 258 | 30 | 288 | - | ||
| 519 | 58 | 577 | [49] | |||
| MRI | 1001 | 112 | 1,113 | [35] | ||
| MRI Total | 1,520 | 170 | 1,690 | - | ||
| MedIR-3D-3K | 3D | Total | 3,166 | 356 | 3,522 | - |
Appendix 0.C Additional Dataset Information
The detailed dataset information is shown in Tab. I. We then describe the information of private data and the methods used to simulate LQ–HQ image pairs.
0.C.1 Private Data Source
We collected two private datasets (Private1 and Private2 in Tab. I) for PET, and one private dataset (Private3 in Tab. I) for CT. This study and the experimental procedures involving all three private datasets were approved by the Biological and Medical Ethnics Committee of Beihang University (approval number BM20250008). Informed consent was obtained from all participating patients.
Private1. We collect 83 3D whole-body PET images using PolarStar m660 PET imaging system, whith an average administered dose of 293 MBq of 18F-FDG. 2D images are extracted from slices of 3D images, excluding slices without anatomical content (e.g., air-only regions).
Private2. We collect 90 3D whole-body PET images using PolarStar Flight PET imaging system, whith an average administered dose of 301 MBq of 18F-FDG. 2D images are extracted from slices of 3D images, excluding slices without anatomical content (e.g., air-only regions).
Private3. We collect 128 3D CT images using the Sinovision CT imaging system. Among them, 18 images are of the spine, 50 of the lungs, and 60 of soft tissues. 2D images are extracted from slices of 3D images, excluding slices without anatomical content (e.g., air-only regions).
0.C.2 Methods for Generating LQ-HQ image Pairs
We describe the simulation methods used to generate LQ–HQ image pairs for different tasks in Tab. I. Our focus is primarily on the key degradation affecting each imaging modality.
PET image denoising. Following the paper [61], the original PET data is collected in listmode. To simulate LQ PET images, list mode data are randomly subsampled to achieve a dose reduction factor of 10. Both HQ and LQ PET images undergo reconstruction using the standard OSEM method.
CT Image Denoising. Following the paper [36], we first obtain the original CT projection data. To simulate LQ CT images, Poisson noise is inserted into the projection data for each case to reach a noise level that corresponded to 25% of the full dose. Both HQ and LQ CT images are then generated using standard CT reconstruction applied to their respective projection data.
MRI Image Super-Resolution. Following the paper [63], the LQ image is generated by transforming the HQ image to the frequency domain, retaining only the central 6.25 of frequency data points while zero-filling the high-frequency parts, and then converting it back to the image domain.
X-ray Image Denoising. According to previous studies [46], X-ray images are primarily affected by Poisson noise. Accordingly, we generate LQ images using the following formulation: , where we set the noise level to .
OCT Image Denosing. According to the paper [13], OCT images suffer from speckle noise, which can be reduced by averaging repeated scans acquired at the same location. Following prior work [17], we averaged five repeated scans to produce a noise-reduced, HQ image, and randomly selected one of the five original scans as the LQ image.
Ultrasound Image Denoising. Ulrasound image often suffer from a multiplicative speckle noise, which can be approximated as: , where is zero-mean Gaussian noise and controls the strength of the content-dependent perturbation. We set
Pathological Image Super-Resolution. We perform 4× Y-channel super-resolution. To synthesize LQ pathological images, we first convert the RGB pathlogical images to the YCbCr color space and then downsample all three channels by a factor of four using bicubic interpolation. During reconstruction, only the Y (luminance) channel is processed by the model, while the Cb and Cr channels are restored using bicubic upsampling, since the human visual system is far more sensitive to luminance than to chrominance.
Appendix 0.D Additional Experiments.
2D All-in-One MedIR Results on AMIR Dataset [63]. We also conduct an all-in-one MedIR performance comparison on the dataset provided in the AMIR paper [63]. This dataset includes three tasks: PET image denoising, CT image denoising, and MRI image super-resolution. The results are shown in Tab. II. Our proposed UniH3 consistently outperforms all comparison methods across these three tasks.
| PET | CT | MRI | Average | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Method | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | PSNR↑ | SSIM↑ | ||||
| Restormer | 37.14 | 0.9473 | 33.61 | 0.9177 | 31.72 | 0.9362 | 34.16 | 0.9337 | ||||
| Eformer | 35.11 | 0.9091 | 32.44 | 0.9078 | 29.19 | 0.8728 | 32.25 | 0.8966 | ||||
| Spach Transformer | 37.05 | 0.9445 | 33.47 | 0.9155 | 31.18 | 0.9290 | 33.90 | 0.9297 | ||||
| DRMC | 36.19 | 0.9376 | 33.28 | 0.9153 | 29.55 | 0.9032 | 33.01 | 0.9187 | ||||
| AirNet | 37.17 | 0.9451 | 33.62 | 0.9176 | 31.39 | 0.9316 | 34.06 | 0.9314 | ||||
| AMIR | 37.12 | 0.9475 | 33.70 | 0.9182 | 32.03 | 0.9396 | 34.28 | 0.9351 | ||||
| UniH3 (Ours) | 37.40 | 0.9478 | 33.79 | 0.9203 | 32.14 | 0.9405 | 34.44 | 0.9362 | ||||
Appendix 0.E Additional Visualization.
Visual Comparison for 2D All-in-One MedIR. We provide an additional visual comparison for 2D all-in-one medical image restoration in Fig. I and Fig. II. Our UniH3 best preserves structures and details across seven tasks.
Visual Comparison for 3D All-in-One MedIR. We provide an additional visual comparison across the coronal, sagittal, and transverse planes for the 3D all-in-one medical image restoration in Fig. III. Our UniH3-3D best preserves structures and details across three tasks.