










Abstract:Multimodal recommendation combines the user historical behaviors with the modal features of items to capture the tangible user preferences, presenting superior performance compared to the conventional ID-based recommender systems. However, existing methods still encounter two key problems in the representation learning of users and items, respectively: (1) the initialization of multimodal user representations is either agnostic to historical behaviors or contaminated by irrelevant modal noise, and (2) the widely used KNN-based item-item graph contains noisy edges with low similarities and lacks audience co-occurrence relationships. To address such issues, we propose MLLMRec, a novel preference reasoning paradigm with graph refinement for multimodal recommendation. Specifically, on the one hand, the item images are first converted into high-quality semantic descriptions using a multimodal large language model (MLLM), thereby bridging the semantic gap between visual and textual modalities. Then, we construct a behavioral description list for each user and feed it into the MLLM to reason about the purified user preference profiles that contain the latent interaction intents. The reasoned profiles and the multimodal descriptions of items, together with their ID embeddings, are propagated over the user-item interaction graph to absorb the high-order collaborative signals. On the other hand, we develop the threshold-controlled denoising and topology-aware enhancement strategies to refine the suboptimal item-item graph, which are applied to both the multimodal and ID item representations to improve the accuracy of item representation learning. Extensive experiments on three publicly available datasets demonstrate that MLLMRec achieves the state-of-the-art performance. The source code is provided at this https URL.
From: Yuzhuo Dang [view email]
[v1]
Thu, 21 Aug 2025 06:50:00 UTC (1,729 KB)
[v2]
Sat, 24 Jan 2026 15:45:14 UTC (3,395 KB)
[v3]
Thu, 10 Sep 2026 13:20:21 UTC (696 KB)
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。