Affiliation(s): 1School of Communication and Electronic Engineering, Jishou University, Jishou 416000, China 2Tencent Inc., Shenzhen 518057, China 3Ant Group Co., Ltd., Changsha 410000, China 4School of Electronic and Information Engineering, Huaihua University, Huaihua 418000, China
Lili WU1, Renmin ZHANG1*, Bin ZHANG1, Jincheng ZHANG3, Xi CHEN2, Yingjing QIAN4. Decoupled prompt-guided mixture-of-experts dynamic distillation for multimodal recommendation[J]. Journal of Zhejiang University Science ,in press.Frontiers of Information Technology & Electronic Engineering,in press.https://doi.org/10.1631/ENG.ITEE.2026.0160
@article{title="Decoupled prompt-guided mixture-of-experts dynamic distillation for multimodal recommendation", author="Lili WU1, Renmin ZHANG1*, Bin ZHANG1, Jincheng ZHANG3, Xi CHEN2, Yingjing QIAN4", journal="Journal of Zhejiang University Science ", year="in press", publisher="Zhejiang University Press & Springer", doi="https://doi.org/10.1631/ENG.ITEE.2026.0160" }
%0 Journal Article %T Decoupled prompt-guided mixture-of-experts dynamic distillation for multimodal recommendation %A Lili WU1 %A Renmin ZHANG1* %A Bin ZHANG1 %A Jincheng ZHANG3 %A Xi CHEN2 %A Yingjing QIAN4 %J Journal of Zhejiang University SCIENCE %P %@ 2095-9184 %D in press %I Zhejiang University Press & Springer doi="https://doi.org/10.1631/ENG.ITEE.2026.0160"
TY - JOUR T1 - Decoupled prompt-guided mixture-of-experts dynamic distillation for multimodal recommendation A1 - Lili WU1 A1 - Renmin ZHANG1* A1 - Bin ZHANG1 A1 - Jincheng ZHANG3 A1 - Xi CHEN2 A1 - Yingjing QIAN4 J0 - Journal of Zhejiang University Science SP - EP - %@ 2095-9184 Y1 - in press PB - Zhejiang University Press & Springer ER - doi="https://doi.org/10.1631/ENG.ITEE.2026.0160"
Abstract: Multimodal recommendation aims to enrich preference modeling by leveraging visual and textual features. However, integrating high-dimensional pretrained features introduces substantial computational overhead. While knowledge distillation provides an effective compression strategy, existing frameworks face three intertwined challenges: rank bottlenecks caused by low-dimensional projections, cross-modal interference induced by shared fusion spaces, and optimization instability under static distillation temperatures. To address these issues, we propose ProMoE-DTS, a decoupled prompt-guided mixture-of-experts framework with dynamic temperature scheduling. Using an asymmetric teacher–student architecture, the teacher model leverages modality-aware soft prompts as semantic anchors to route heterogeneous features into parameter-disjoint expert networks, thereby alleviating cross-modal conflicts and resolving the rank bottlenecks. To ensure stable knowledge transfer, a feedback-driven dynamic temperature scheduler adaptively regulates the distillation intensity based on epoch-wise signals. This asymmetric design confines intensive multimodal operations to the offline teacher, leaving the online student model with a highly efficient, pure identifier-based structure. Extensive experiments on three benchmark datasets demonstrate that ProMoE-DTS improves Recall@20 by 2.24%–3.96% over state-of-the-art baselines, while requiring only 3.28%–3.55% of the teacher’s parameters.
Darkslateblue:Affiliate; Royal Blue:Author; Turquoise:Article
Reference
Open peer comments: Debate/Discuss/Question/Opinion
Open peer comments: Debate/Discuss/Question/Opinion
<1>