
Lili WU1, Renmin ZHANG1*, Bin ZHANG1, Jincheng ZHANG3, Xi CHEN2, Yingjing QIAN4. Decoupled prompt-guided mixture-of-experts dynamic distillation for multimodal recommendation[J]. Journal of Zhejiang University Science C, 1998, -1(-1): .
@article{title="Decoupled prompt-guided mixture-of-experts dynamic distillation for multimodal recommendation",
author="Lili WU1, Renmin ZHANG1*, Bin ZHANG1, Jincheng ZHANG3, Xi CHEN2, Yingjing QIAN4",
journal="Journal of Zhejiang University Science C",
volume="-1",
number="-1",
pages="",
year="1998",
publisher="Zhejiang University Press & Springer",
doi="10.1631/ENG.ITEE.2026.0160"
}
%0 Journal Article
%T Decoupled prompt-guided mixture-of-experts dynamic distillation for multimodal recommendation
%A Lili WU1
%A Renmin ZHANG1*
%A Bin ZHANG1
%A Jincheng ZHANG3
%A Xi CHEN2
%A Yingjing QIAN4
%J Journal of Zhejiang University SCIENCE C
%V -1
%N -1
%P
%@ 1869-1951
%D 1998
%I Zhejiang University Press & Springer
%DOI 10.1631/ENG.ITEE.2026.0160
TY - JOUR
T1 - Decoupled prompt-guided mixture-of-experts dynamic distillation for multimodal recommendation
A1 - Lili WU1
A1 - Renmin ZHANG1*
A1 - Bin ZHANG1
A1 - Jincheng ZHANG3
A1 - Xi CHEN2
A1 - Yingjing QIAN4
J0 - Journal of Zhejiang University Science C
VL - -1
IS - -1
SP -
EP - 0
%@ 1869-1951
Y1 - 1998
PB - Zhejiang University Press & Springer
ER -
DOI - 10.1631/ENG.ITEE.2026.0160
Abstract: Multimodal recommendation aims to enrich preference modeling by leveraging visual and textual features. However, integrating high-dimensional pretrained features introduces substantial computational overhead. While knowledge distillation provides an effective compression strategy, existing frameworks face three intertwined challenges: rank bottlenecks caused by low-dimensional projections, cross-modal interference induced by shared fusion spaces, and optimization instability under static distillation temperatures. To address these issues, we propose ProMoE-DTS, a decoupled prompt-guided mixture-of-experts framework with dynamic temperature scheduling. Using an asymmetric teacher–student architecture, the teacher model leverages modality-aware soft prompts as semantic anchors to route heterogeneous features into parameter-disjoint expert networks, thereby alleviating cross-modal conflicts and resolving the rank bottlenecks. To ensure stable knowledge transfer, a feedback-driven dynamic temperature scheduler adaptively regulates the distillation intensity based on epoch-wise signals. This asymmetric design confines intensive multimodal operations to the offline teacher, leaving the online student model with a highly efficient, pure identifier-based structure. Extensive experiments on three benchmark datasets demonstrate that ProMoE-DTS improves Recall@20 by 2.24%–3.96% over state-of-the-art baselines, while requiring only 3.28%–3.55% of the teacher’s parameters.
CLC number:
On-line Access: 2026-09-29
Received: 2026-05-26
Revision Accepted: 2026-08-04
Crosschecked: 0000-00-00
Cited: 0
Clicked: 36
Open peer comments: Debate/Discuss/Question/Opinion
<1>