Research & Publications

MICCAI 2021

UniMiSS: Universal self-supervised learning across 2D and 3D medical imaging

Yutong Xie, Jianpeng Zhang, Yong Xia, Qi Wu

2021

UniMiSS framework combining 2D X-rays and 3D CT scans for universal medical representation learning

Publication summary

UniMiSS addresses a fundamental limitation of medical self-supervised learning: the scarcity of large-scale 3D datasets. The framework leverages abundant 2D medical images, such as chest X-rays, to complement limited 3D CT volumes during pretraining.

To bridge the dimensionality gap between 2D and 3D data, the authors introduce a dimension-free Medical Transformer (MiT) with a Switchable Patch Embedding (SPE) module that dynamically adapts to either 2D or 3D inputs.

Through joint self-supervised learning on 5,022 CT volumes and 108,948 X-ray images, UniMiSS learns transferable representations that improve performance across segmentation and classification tasks in both 2D and 3D medical imaging.

“UniMiSS demonstrates that medical foundation representations can be learned across imaging dimensions, allowing 2D and 3D modalities to reinforce each other during self-supervised training.”

Editorial research summary

OrthoAI research content adaptation

Key findings

113,970

Joint self-supervised pretraining performed using 5,022 CT volumes and 108,948 X-ray images.

+3.15%

Classification improvement over DINO on the RICORD COVID-19 screening benchmark with full labels.

88.11%

Best BCV online benchmark Dice score achieved using the UniMiSS ensemble configuration.

Method

Medical Transformer with switchable patch embedding

UniMiSS introduces the Medical Transformer (MiT), a pyramid U-shaped Transformer architecture capable of processing both 2D and 3D medical images.

The Switchable Patch Embedding module dynamically chooses either 2D or 3D patch embedding based on the incoming data, enabling a unified representation space.

UniMiSS Medical Transformer with switchable patch embedding

Cross-dimensional self-supervised learning

The framework employs a student-teacher self-distillation strategy and alternates training between 2D and 3D datasets.

A volume-slice consistency objective aligns volumetric CT representations with their corresponding 2D slices, strengthening cross-dimensional representation learning.

UniMiSS self-supervised learning pipeline using 2D and 3D medical images

Key contributions

UniMiSS establishes a universal medical self-supervised learning framework capable of bridging dimensionality barriers in medical imaging.

  • Introduces the first self-supervised framework that jointly learns from both 2D and 3D medical images.
  • Proposes the Medical Transformer (MiT) architecture with Switchable Patch Embedding.
  • Breaks the dimensionality barrier between X-rays and CT volumes through a unified Transformer backbone.
  • Introduces volume-slice consistency learning to exploit relationships between 3D volumes and 2D slices.
  • Demonstrates superior transfer learning performance on six downstream medical imaging tasks.
  • Shows strong generalization to unseen imaging modalities including MRI and dermoscopic images.

Similar publications

CVPR 2021 · IEEE/CVF Conference on Computer Vision and Pattern Recognition

DoDNet: Learning to Segment Multi-Organ and Tumors from Multiple Partially Labeled Datasets

Jianpeng Zhang, Yutong Xie, Yong Xia, Chunhua Shen

view

CVPR 2024 · IEEE/CVF Conference on Computer Vision and Pattern Recognition

Continual Self-supervised Learning: Towards Universal Multi-modal Medical Data Representation Learning

Yiwen Ye, Yutong Xie, Jianpeng Zhang, Ziyang Chen, Qi Wu, Yong Xia

view