5
Universal architecture supporting report, x-ray, ct, mri, pathology in a single continual learning pipeline.
Research & Publications
CVPR 2024 · IEEE/CVF Conference on Computer Vision and Pattern Recognition
Yiwen Ye, Yutong Xie, Jianpeng Zhang, Ziyang Chen, Qi Wu, Yong Xia
2024
Yiwen Ye, Yutong Xie, Jianpeng Zhang, Ziyang Chen, Qi Wu, Yong Xia
2024

MedCoSS (MedCoSS) introduces a continual self-supervised learning framework for building universal multi-modal medical representations. Instead of training all modalities jointly in a single stage, the method learns sequentially across report, x-ray, ct, mri, pathology data while preserving knowledge from earlier modalities.
The framework uses a rehearsal buffer with k-means sampling, feature distillation for knowledge retention, and intra-modal mixup augmentation to mitigate catastrophic forgetting as new medical modalities are introduced during training.
Evaluated on downstream benchmarks including PubMed20k, ChestXR, QaTa, RICORD, and others, MedCoSS demonstrates that sequential multi-modal self-supervised learning can produce transferable representations across diverse clinical data types.
"MedCoSS shows that universal medical representation learning does not require simultaneous access to all modalities — continual learning with rehearsal and distillation can build robust cross-modal foundations."
Editorial research summary
OrthoAI research content adaptation
Universal architecture supporting report, x-ray, ct, mri, pathology in a single continual learning pipeline.
Validated across downstream datasets including PubMed20k, ChestXR, QaTa, and more.
Sequential multi-modal self-supervised learning outperforms naive joint training on cross-modal transfer tasks.
MedCoSS trains on medical modalities sequentially rather than jointly, using a universal architecture that supports 1D reports, 2D X-rays and pathology, and 3D CT and MRI volumes.
A rehearsal buffer with k-means sampling replays representative samples from prior modalities, while feature distillation preserves learned representations as new modalities are introduced.

Intra-modal mixup augmentation strengthens representation robustness within each modality during continual training.
The universal backbone adapts to varying input dimensionality while maintaining a shared feature space for downstream medical AI tasks.

Continual self-supervised learning framework for universal multi-modal medical representation learning across reports, X-ray, CT, MRI and pathology data.

CVPR 2024 · IEEE/CVF Conference on Computer Vision and Pattern Recognition
Yutong Xie, Qi Chen, Sinuo Wang, Minh-Son To, Iris Lee, Ee Win Khoo, Kerolos Hendy, Daniel Koh, Yong Xia, Qi Wu

MICCAI 2023
Yutong Xie, Lin Gu, Tatsuya Harada, Jianpeng Zhang, Yong Xia, Qi Wu