
Yanzheng WANG, Yujun WANG, Fengshuo DAI, Xin YANG, Xiaoheng JIANG, Pei LV, Shizhe HU, Mingliang XU. Multi-stage contrastive multi-modal clustering with consistency retained[J]. Journal of Zhejiang University Science C, 2026, 27(8): 1-13.
@article{title="Multi-stage contrastive multi-modal clustering with consistency retained",
author="Yanzheng WANG, Yujun WANG, Fengshuo DAI, Xin YANG, Xiaoheng JIANG, Pei LV, Shizhe HU, Mingliang XU",
journal="Journal of Zhejiang University Science C",
volume="27",
number="8",
pages="1-13",
year="2026",
publisher="Zhejiang University Press & Springer",
doi="10.1631/ENG.ITEE.2025.0185"
}
%0 Journal Article
%T Multi-stage contrastive multi-modal clustering with consistency retained
%A Yanzheng WANG
%A Yujun WANG
%A Fengshuo DAI
%A Xin YANG
%A Xiaoheng JIANG
%A Pei LV
%A Shizhe HU
%A Mingliang XU
%J Frontiers of Information Technology & Electronic Engineering
%V 27
%N 8
%P 1-13
%@ 1869-1951
%D 2026
%I Zhejiang University Press & Springer
%DOI 10.1631/ENG.ITEE.2025.0185
TY - JOUR
T1 - Multi-stage contrastive multi-modal clustering with consistency retained
A1 - Yanzheng WANG
A1 - Yujun WANG
A1 - Fengshuo DAI
A1 - Xin YANG
A1 - Xiaoheng JIANG
A1 - Pei LV
A1 - Shizhe HU
A1 - Mingliang XU
J0 - Frontiers of Information Technology & Electronic Engineering
VL - 27
IS - 8
SP - 1
EP - 13
%@ 1869-1951
Y1 - 2026
PB - Zhejiang University Press & Springer
ER -
DOI - 10.1631/ENG.ITEE.2025.0185
Abstract: Contrastive learning has received extensive attention because it can effectively strengthen inter-modal alignment and information complementarity in multi-modal clustering. However, there are two main problems with existing contrastive multi-modal clustering methods: (1) Current methods usually perform contrastive learning at only one or two stages and have imprecise selection strategies for negative sample pairs. (2) Most methods focus only on contrastive learning and ignore the retention of consistency information. To address the above challenges, we propose a framework of multi-stage contrastive multi-modal clustering with consistency retained. First, we design a systematic contrastive learning framework, which includes early–middle–late contrastive learning between single features, fused features, and clusters, respectively, which improves the semantic integrity and information fidelity of the fused representation. Second, we propose a new negative sample pair selection strategy to discriminately select the correct negative sample pair for each sample. Finally, we add a new loss term to late-stage contrastive learning, which considers the consistency of clustering results of the same sample under different modalities for the first time. Therefore, compared to the baseline model, this module has improved the accuracy by an average of 7.4 percentage points (PPs). In addition, we design a consistency retained module to ensure that the model can learn the information of each modality without being disturbed by high/low quality modalities, which improves the accuracy by an average of 2.2 PPs. Experiments conducted on multiple multi-modal datasets demonstrate that the accuracy of our method is 4.1 PPs higher on average compared to the second-best multi-modal clustering method, proving the excellent potential of our method.
[1]Cai X, Nie F, Huang H, 2013. Multi-view K-means clustering on big data. Proc 23rd Int Joint Conf on Artificial Intelligence, p.2598-2604.
[2]Chua TS, Tang J, Hong R, et al., 2009. NUS-WIDE: a real-world web image database from National University of Singapore. Proc ACM Int Conf on Image Video Retrieval, Article 48.
[3]Cui J, Li Y, Huang H, et al., 2024. Dual contrast-driven deep multi-view clustering. IEEE Trans Image Process, 33:4753-4764.
[4]Dai Z, Wang G, Yuan W, et al., 2022. Cluster contrast for unsupervised person re-identification. 16th Asian Conf on Computer Vision, p.319-337.
[5]Friedman A, 1979. Framing pictures: the role of knowledge in automatized encoding and memory for gist. J Exp Psychol-Gen, 108(3):316-355.
[6]Grubinger M, Clough P, Müller H, et al., 2006. The IAPR TC-12 benchmark: a new evaluation resource for visual information systems. https://www-i6.informatik.rwth-aachen.de/publications/download/34/Grubinger-LREC-2006.pdf [Accessed on Dec. 22, 2025]
[7]Gu WJ, Zhu CM, 2024. Global and local combined contrastive learning for multi-view clustering. Multim Syst, 30(5):307.
[8]Hartigan JA, Wong MA, 1979. Algorithm AS 136: a K-means clustering algorithm. J R Stat Soc, 28(1):100-108.
[9]Herath A, Kobti Z, 2024. Subtype-MMCC: multimodal contrastive clustering approach for cancer subtype discovery with multi-omics data. Proc Comput Sci, 246:696-705.
[10]Hu SZ, Zou GL, Zhang CY, et al., 2023. Joint contrastive triple-learning for deep multi-view clustering. Inform Process Manag, 60(3):103284.
[11]Hu SZ, Zhang CK, Zou GL, et al., 2025a. Deep multiview clustering by pseudo-label guided contrastive learning and dual correlation learning. IEEE Trans Neur Netw Learn Syst, 36(2):3646-3658.
[12]Hu SZ, Fan JH, Zou GL, et al., 2025b. Multi-aspect self-guided deep information bottleneck for multi-modal clustering. Proc 39th AAAI Conf on Artificial Intelligence, p.17314-17322.
[13]Huiskes MJ, Lew MS, 2008. The MIR Flickr retrieval evaluation. Proc 1st ACM Int Conf on Multimedia Information Retrieval, p.39-43.
[14]Joyce JM, 2011. Kullback–Leibler Divergence. Springer Berlin, Heidelberg, p.720-722.
[15]Kampffmeyer M, Løkse S, Bianchi FM, et al., 2019. Deep divergence-based approach to clustering. Neur Netw, 113:91-101. https://doi.org/
[16]Ke GZ, Hong ZY, Zeng ZQ, et al., 2021. CONAN: contrastive fusion networks for multi-view clustering. Proc IEEE Int Conf on Big Data, p.653-660.
[17]Kingma DP, Ba J, 2017. Adam: a method for stochastic optimization.
[18]Kumar A, Rai P, Daumé H III, 2011. Co-regularized multi-view spectral clustering. Proc 25th Int Conf on Neural Information Processing Systems, p.1413-1421.
[19]Lao JH, Huang D, Wang CD, et al., 2024. Towards scalable multi-view clustering via joint learning of many bipartite graphs. IEEE Trans Big Data, 10(1):77-91.
[20]Li FF, Fergus R, Perona P, 2004. Learning generative visual models from few training examples: an incremental Bayesian approach tested on 101 object categories. Proc IEEE Conf on Computer Vision and Pattern Recognition Workshop, p.178.
[21]Lin K, Wang D, Xia FZ, et al., 2018. Device clustering algorithm based on multimodal data correlation in cognitive Internet of Things. IEEE Int Things J, 5(4):2263-2271.
[22]Lou ZZ, Xue H, Wang YZ, et al., 2025. Parameter-free deep multi-modal clustering with reliable contrastive learning. IEEE Trans Image Process, 34:2628-2640.
[23]Lu YD, Lin YJ, Yang MX, et al., 2024. Decoupled contrastive multi-view clustering with high-order random walks. Proc 38th AAAI Conf on Artificial Intelligence, p.14193-14201.
[24]Nie FP, Li J, Li XL, 2017. Self-weighted multiview clustering with multiple graphs. Proc 26th Int Joint Conf on Artificial Intelligence, p.2564-2570.
[25]Pan L, Sun GD, Chang BF, et al., 2023. Visual interactive image clustering: a target-independent approach for configuration optimization in machine vision measurement. Front Inform Technol Electron Eng, 24(3):355-372.
[26]Shen DG, Ip HHS, 1999. Discriminative wavelet shape descriptors for recognition of 2-D patterns. Patt Recog, 32(2):151-165.
[27]Shi JB, Malik J, 2000. Normalized cuts and image segmentation. IEEE Trans Patt Anal Mach Intell, 22(8):888-905.
[28]Tian YL, Krishnan D, Isola P, 2020. Contrastive multiview coding. Proc 16th European Conf on Computer Vision, p.776-794.
[29]Tong Y, 2025. Research on deep clustering and multimodal deep learning service recommendation system for large-scale user data. 3rd Int Conf on Integrated Circuits and Communication Systems, p.1-5.
[30]Trosten DJ, Løkse S, Jenssen R, et al., 2021. Reconsidering representation alignment for multi-view clustering. Proc IEEE/CVF Conf on Computer Vision and Pattern, p.1255-1265.
[31]Wolf L, Hassner T, Taigman Y, 2008. Descriptor based methods in the wild. Workshop on Faces in ‘Real-Life’ Images: Detection, Alignment, and Recognition.
[32]Wu JX, Rehg JM, 2011. CENTRIST: a visual descriptor for scene categorization. IEEE Trans Patt Anal Mach Intell, 33(8):1489-1501.
[33]Wu S, Zheng Y, Ren YZ, et al., 2024. Self-weighted contrastive fusion for deep multi-view clustering. IEEE Trans Multim, 26:9150-9162.
[34]Xu J, Ren YZ, Li GF, et al., 2021. Deep embedded multi-view clustering with collaborative training. Inform Sci, 573:279-290.
[35]Xu J, Tang HY, Ren YZ, et al., 2022. Multi-level feature learning for contrastive multi-view clustering. Proc IEEE/CVF Conf on Computer Vision and Pattern Recognition, p.16030-16039.
[36]Xu J, Chen S, Ren YZ, et al., 2023. Self-weighted contrastive learning among multiple views for mitigating representation degeneration. Proc 37th Int Conf on Neural Information Processing Systems, Article 54.
[37]Yan WQ, Yang TY, Tang C, 2025. Self-supervised semantic soft label learning network for deep multi-view clustering. IEEE Trans Multim, 27:4971-4983.
[38]Yang KL, Zhang TL, Alhuzali H, et al., 2023. Cluster-level contrastive learning for emotion recognition in conversations. IEEE Trans Affect Comput, 14(4):3269-3280.
[39]Yin YJ, Ruan H, Chen Y, et al., 2025. Prototypical clustered federated learning for heart rate prediction. Front Inform Technol Electron Eng, 26(10):1896-1912.
[40]Zhang FD, Kuang K, Chen L, et al., 2023. Federated unsupervised representation learning. Front Inform Technol Electron Eng, 24(8):1181-1193.
[41]Zhou RW, Shen YD, 2020. End-to-end adversarial-attention network for multi-modal clustering. IEEE/CVF Conf on Computer Vision and Pattern Recognition, p.14607-14616.
[42]Zhou SH, Liu XW, Liu JY, et al., 2020. Multi-view spectral clustering with optimal neighborhood Laplacian matrix. Proc 34th AAAI Conf on Artificial Intelligence, p.6965-6972.
CLC number: TP391.41
On-line Access: 2026-06-02
Received: 2025-12-23
Revision Accepted: 2026-06-10
Crosschecked: 2026-06-17
Cited: 0
Clicked: 22
Open peer comments: Debate/Discuss/Question/Opinion
<1>