ENGINEERING Information Technology & Electronic Engineering  2026 Vol.27 No.8 P.1-13

http://doi.org/10.1631/ENG.ITEE.2025.0185


Multi-stage contrastive multi-modal clustering with consistency retained


Author(s):  Yanzheng WANG, Yujun WANG, Fengshuo DAI, Xin YANG, Xiaoheng JIANG, Pei LV, Shizhe HU, Mingliang XU

Affiliation(s):  1. School of Computer Science and Artificial Intelligence, Zhengzhou University, Zhengzhou 450001, China more

Corresponding email(s):   ieshizhehu@zzu.edu.cniexumingliang@zzu.edu.cn

Key Words:  Multi-modal clustering, Contrastive learning, Consistency retention, Negative sample selection


Yanzheng WANG, Yujun WANG, Fengshuo DAI, Xin YANG, Xiaoheng JIANG, Pei LV, Shizhe HU, Mingliang XU. Multi-stage contrastive multi-modal clustering with consistency retained[J]. Journal of Zhejiang University Science C, 2026, 27(8): 1-13.

@article{title="Multi-stage contrastive multi-modal clustering with consistency retained",
author="Yanzheng WANG, Yujun WANG, Fengshuo DAI, Xin YANG, Xiaoheng JIANG, Pei LV, Shizhe HU, Mingliang XU",
journal="Journal of Zhejiang University Science C",
volume="27",
number="8",
pages="1-13",
year="2026",
publisher="Zhejiang University Press & Springer",
doi="10.1631/ENG.ITEE.2025.0185"
}

%0 Journal Article
%T Multi-stage contrastive multi-modal clustering with consistency retained
%A Yanzheng WANG
%A Yujun WANG
%A Fengshuo DAI
%A Xin YANG
%A Xiaoheng JIANG
%A Pei LV
%A Shizhe HU
%A Mingliang XU
%J Frontiers of Information Technology & Electronic Engineering
%V 27
%N 8
%P 1-13
%@ 1869-1951
%D 2026
%I Zhejiang University Press & Springer
%DOI 10.1631/ENG.ITEE.2025.0185

TY - JOUR
T1 - Multi-stage contrastive multi-modal clustering with consistency retained
A1 - Yanzheng WANG
A1 - Yujun WANG
A1 - Fengshuo DAI
A1 - Xin YANG
A1 - Xiaoheng JIANG
A1 - Pei LV
A1 - Shizhe HU
A1 - Mingliang XU
J0 - Frontiers of Information Technology & Electronic Engineering
VL - 27
IS - 8
SP - 1
EP - 13
%@ 1869-1951
Y1 - 2026
PB - Zhejiang University Press & Springer
ER -
DOI - 10.1631/ENG.ITEE.2025.0185


Abstract: 
Contrastive learning has received extensive attention because it can effectively strengthen inter-modal alignment and information complementarity in multi-modal clustering. However, there are two main problems with existing contrastive multi-modal clustering methods: (1) Current methods usually perform contrastive learning at only one or two stages and have imprecise selection strategies for negative sample pairs. (2) Most methods focus only on contrastive learning and ignore the retention of consistency information. To address the above challenges, we propose a framework of multi-stage contrastive multi-modal clustering with consistency retained. First, we design a systematic contrastive learning framework, which includes early–middle–late contrastive learning between single features, fused features, and clusters, respectively, which improves the semantic integrity and information fidelity of the fused representation. Second, we propose a new negative sample pair selection strategy to discriminately select the correct negative sample pair for each sample. Finally, we add a new loss term to late-stage contrastive learning, which considers the consistency of clustering results of the same sample under different modalities for the first time. Therefore, compared to the baseline model, this module has improved the accuracy by an average of 7.4 percentage points (PPs). In addition, we design a consistency retained module to ensure that the model can learn the information of each modality without being disturbed by high/low quality modalities, which improves the accuracy by an average of 2.2 PPs. Experiments conducted on multiple multi-modal datasets demonstrate that the accuracy of our method is 4.1 PPs higher on average compared to the second-best multi-modal clustering method, proving the excellent potential of our method.

一致性保持的多阶段对比多模态聚类算法

王彦铮1,王玉骏1,代丰槊2,杨欣1,姜晓恒1,吕培1,胡世哲1,徐明亮1,3
1郑州大学计算机与人工智能学院,中国郑州市,450001
2郑州大学国际学院,中国郑州市,450001
3山东博算智新信息科技有限公司,中国济南市,250101
摘要:由于能有效强化多模态聚类任务中的跨模态对齐与信息互补特性,对比学习得到广泛关注。然而现有基于对比学习的多模态聚类方法存在两大核心缺陷:(1)现有方法仅能在单一或两个阶段开展对比学习,且负样本对筛选策略精度不足;(2)绝大多数方法仅聚焦于对比学习,却忽略模态一致性信息的留存约束。针对上述两类难题,提出一种一致性保持的多阶段对比多模态聚类框架。首先,构建一套系统的分层对比学习体系,分别在单个表征、融合特征、聚类簇执行前期-中期-后期递进式对比学习,以此提升融合表征的语义完备度与信息保真度。其次,提出一种新的负样本对筛选策略,可为每个样本差异化匹配精准的负样本对。最后,在后期对比学习中引入新的损失项,该损失首次将同一样本在不同模态下聚类结果的一致性纳入优化目标。因此,与基线模型相比,该模块的准确率平均提高了7.4个百分点。此外,设计了一致性保持模块,保障模型能学习各模态信息,规避高/低质量模态的干扰,使准确率平均提高了2.2个百分点。在多组多模态数据集上开展的实验表明,与排名第二的多模态聚类方法相比,所提方法的准确率平均高出4.1个百分点,证明该方法具备良好的应用与研究前景。

关键词:多模态聚类;对比学习;一致性保持;负样本选择

Darkslateblue:Affiliate; Royal Blue:Author; Turquoise:Article

Reference

[1]Cai X, Nie F, Huang H, 2013. Multi-view K-means clustering on big data. Proc 23rd Int Joint Conf on Artificial Intelligence, p.2598-2604.

[2]Chua TS, Tang J, Hong R, et al., 2009. NUS-WIDE: a real-world web image database from National University of Singapore. Proc ACM Int Conf on Image Video Retrieval, Article 48.

[3]Cui J, Li Y, Huang H, et al., 2024. Dual contrast-driven deep multi-view clustering. IEEE Trans Image Process, 33:4753-4764.

[4]Dai Z, Wang G, Yuan W, et al., 2022. Cluster contrast for unsupervised person re-identification. 16th Asian Conf on Computer Vision, p.319-337.

[5]Friedman A, 1979. Framing pictures: the role of knowledge in automatized encoding and memory for gist. J Exp Psychol-Gen, 108(3):316-355.

[6]Grubinger M, Clough P, Müller H, et al., 2006. The IAPR TC-12 benchmark: a new evaluation resource for visual information systems. https://www-i6.informatik.rwth-aachen.de/publications/download/34/Grubinger-LREC-2006.pdf [Accessed on Dec. 22, 2025]

[7]Gu WJ, Zhu CM, 2024. Global and local combined contrastive learning for multi-view clustering. Multim Syst, 30(5):307.

[8]Hartigan JA, Wong MA, 1979. Algorithm AS 136: a K-means clustering algorithm. J R Stat Soc, 28(1):100-108.

[9]Herath A, Kobti Z, 2024. Subtype-MMCC: multimodal contrastive clustering approach for cancer subtype discovery with multi-omics data. Proc Comput Sci, 246:696-705.

[10]Hu SZ, Zou GL, Zhang CY, et al., 2023. Joint contrastive triple-learning for deep multi-view clustering. Inform Process Manag, 60(3):103284.

[11]Hu SZ, Zhang CK, Zou GL, et al., 2025a. Deep multiview clustering by pseudo-label guided contrastive learning and dual correlation learning. IEEE Trans Neur Netw Learn Syst, 36(2):3646-3658.

[12]Hu SZ, Fan JH, Zou GL, et al., 2025b. Multi-aspect self-guided deep information bottleneck for multi-modal clustering. Proc 39th AAAI Conf on Artificial Intelligence, p.17314-17322.

[13]Huiskes MJ, Lew MS, 2008. The MIR Flickr retrieval evaluation. Proc 1st ACM Int Conf on Multimedia Information Retrieval, p.39-43.

[14]Joyce JM, 2011. Kullback–Leibler Divergence. Springer Berlin, Heidelberg, p.720-722.

[15]Kampffmeyer M, Løkse S, Bianchi FM, et al., 2019. Deep divergence-based approach to clustering. Neur Netw, 113:91-101. https://doi.org/

[16]Ke GZ, Hong ZY, Zeng ZQ, et al., 2021. CONAN: contrastive fusion networks for multi-view clustering. Proc IEEE Int Conf on Big Data, p.653-660.

[17]Kingma DP, Ba J, 2017. Adam: a method for stochastic optimization.

[18]Kumar A, Rai P, Daumé H III, 2011. Co-regularized multi-view spectral clustering. Proc 25th Int Conf on Neural Information Processing Systems, p.1413-1421.

[19]Lao JH, Huang D, Wang CD, et al., 2024. Towards scalable multi-view clustering via joint learning of many bipartite graphs. IEEE Trans Big Data, 10(1):77-91.

[20]Li FF, Fergus R, Perona P, 2004. Learning generative visual models from few training examples: an incremental Bayesian approach tested on 101 object categories. Proc IEEE Conf on Computer Vision and Pattern Recognition Workshop, p.178.

[21]Lin K, Wang D, Xia FZ, et al., 2018. Device clustering algorithm based on multimodal data correlation in cognitive Internet of Things. IEEE Int Things J, 5(4):2263-2271.

[22]Lou ZZ, Xue H, Wang YZ, et al., 2025. Parameter-free deep multi-modal clustering with reliable contrastive learning. IEEE Trans Image Process, 34:2628-2640.

[23]Lu YD, Lin YJ, Yang MX, et al., 2024. Decoupled contrastive multi-view clustering with high-order random walks. Proc 38th AAAI Conf on Artificial Intelligence, p.14193-14201.

[24]Nie FP, Li J, Li XL, 2017. Self-weighted multiview clustering with multiple graphs. Proc 26th Int Joint Conf on Artificial Intelligence, p.2564-2570.

[25]Pan L, Sun GD, Chang BF, et al., 2023. Visual interactive image clustering: a target-independent approach for configuration optimization in machine vision measurement. Front Inform Technol Electron Eng, 24(3):355-372.

[26]Shen DG, Ip HHS, 1999. Discriminative wavelet shape descriptors for recognition of 2-D patterns. Patt Recog, 32(2):151-165.

[27]Shi JB, Malik J, 2000. Normalized cuts and image segmentation. IEEE Trans Patt Anal Mach Intell, 22(8):888-905.

[28]Tian YL, Krishnan D, Isola P, 2020. Contrastive multiview coding. Proc 16th European Conf on Computer Vision, p.776-794.

[29]Tong Y, 2025. Research on deep clustering and multimodal deep learning service recommendation system for large-scale user data. 3rd Int Conf on Integrated Circuits and Communication Systems, p.1-5.

[30]Trosten DJ, Løkse S, Jenssen R, et al., 2021. Reconsidering representation alignment for multi-view clustering. Proc IEEE/CVF Conf on Computer Vision and Pattern, p.1255-1265.

[31]Wolf L, Hassner T, Taigman Y, 2008. Descriptor based methods in the wild. Workshop on Faces in ‘Real-Life’ Images: Detection, Alignment, and Recognition.

[32]Wu JX, Rehg JM, 2011. CENTRIST: a visual descriptor for scene categorization. IEEE Trans Patt Anal Mach Intell, 33(8):1489-1501.

[33]Wu S, Zheng Y, Ren YZ, et al., 2024. Self-weighted contrastive fusion for deep multi-view clustering. IEEE Trans Multim, 26:9150-9162.

[34]Xu J, Ren YZ, Li GF, et al., 2021. Deep embedded multi-view clustering with collaborative training. Inform Sci, 573:279-290.

[35]Xu J, Tang HY, Ren YZ, et al., 2022. Multi-level feature learning for contrastive multi-view clustering. Proc IEEE/CVF Conf on Computer Vision and Pattern Recognition, p.16030-16039.

[36]Xu J, Chen S, Ren YZ, et al., 2023. Self-weighted contrastive learning among multiple views for mitigating representation degeneration. Proc 37th Int Conf on Neural Information Processing Systems, Article 54.

[37]Yan WQ, Yang TY, Tang C, 2025. Self-supervised semantic soft label learning network for deep multi-view clustering. IEEE Trans Multim, 27:4971-4983.

[38]Yang KL, Zhang TL, Alhuzali H, et al., 2023. Cluster-level contrastive learning for emotion recognition in conversations. IEEE Trans Affect Comput, 14(4):3269-3280.

[39]Yin YJ, Ruan H, Chen Y, et al., 2025. Prototypical clustered federated learning for heart rate prediction. Front Inform Technol Electron Eng, 26(10):1896-1912.

[40]Zhang FD, Kuang K, Chen L, et al., 2023. Federated unsupervised representation learning. Front Inform Technol Electron Eng, 24(8):1181-1193.

[41]Zhou RW, Shen YD, 2020. End-to-end adversarial-attention network for multi-modal clustering. IEEE/CVF Conf on Computer Vision and Pattern Recognition, p.14607-14616.

[42]Zhou SH, Liu XW, Liu JY, et al., 2020. Multi-view spectral clustering with optimal neighborhood Laplacian matrix. Proc 34th AAAI Conf on Artificial Intelligence, p.6965-6972.

Open peer comments: Debate/Discuss/Question/Opinion

<1>

Please provide your name, email address and a comment





Full Text:   <6>

Summary:  <16>

CLC number: TP391.41

On-line Access: 2026-06-02

Received: 2025-12-23

Revision Accepted: 2026-06-10

Crosschecked: 2026-06-17

Cited: 0

Clicked: 24

Citations:  Bibtex RefMan EndNote GB/T7714

 ORCID:

Yanzheng WANG

0009-0000-4005-7464

Yujun WANG

0009-0007-8684-1300

Fengshuo DAI

0009-0007-7509-9315

Xin YANG

0009-0008-1130-5914

Xiaoheng JIANG

0000-0002-5770-0417

Pei LV

0000-0002-2654-0561

Shizhe HU

0000-0003-1301-2396

Mingliang XU

0000-0002-6885-3451

Journal of Zhejiang University-SCIENCE, 38 Zheda Road, Hangzhou 310027, China
Tel: +86-571-87952783; E-mail: cjzhang@zju.edu.cn
Copyright © 2000 - 2026 Journal of Zhejiang University-SCIENCE