Where to Go Next for Recommender Systems? ID- vs. Modality-based Recommender Models Revisited

Recommendation models that utilize unique identities (IDs) to represent distinct users and items have been state-of-the-art (SOTA) and dominated the recommender systems (RS) literature for over a decade. Meanwhile, the pre-trained modality encoders, such as BERT and ViT, have become increasingly powerful in modeling the raw modality features of an item, such as text and images. Given this, a natural question arises: can a purely modality-based recommendation model (MoRec) outperforms or matches a pure ID-based model (IDRec) by replacing the itemID embedding with a SOTA modality encoder? In fact, this question was answered ten years ago when IDRec beats MoRec by a strong margin in both recommendation accuracy and efficiency. We aim to revisit this `old' question and systematically study MoRec from several aspects. Specifically, we study several sub-questions: (i) which recommendation paradigm, MoRec or IDRec, performs better in practical scenarios, especially in the general setting and warm item scenarios where IDRec has a strong advantage? does this hold for items with different modality features? (ii) can the latest technical advances from other communities (i.e., natural language processing and computer vision) translate into accuracy improvement for MoRec? (iii) how to effectively utilize item modality representation, can we use it directly or do we have to adjust it with new data? (iv) are there some key challenges for MoRec to be solved in practical applications? To answer them, we conduct rigorous experiments for item recommendations with two popular modalities, i.e., text and vision. We provide the first empirical evidence that MoRec is already comparable to its IDRec counterpart with an expensive end-to-end training method, even for warm item recommendation. Our results potentially imply that the dominance of IDRec in the RS field may be greatly challenged in the future.

Neural CollaborativeFilteringNeural Collaborative FilteringBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingMIND: A Large-scaleDataset for News…MIND: A Large-scale Dataset for News RecommendationZero-Shot RecommenderSystemsZero-Shot Recommender SystemsAn Image is Worth 16x16Words: Transformers for…An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleTiny-NewsRec: Efficientand Effective PLM-based…Tiny-NewsRec: Efficient and Effective PLM-based News RecommendationOne4all UserRepresentation for…One4all User Representation for Recommender Systems in E-commerceSwin Transformer:Hierarchical Vision…Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsTenrec: A Large-scaleMultipurpose Benchmark…Tenrec: A Large-scale Multipurpose Benchmark Dataset for Recommender SystemsTowards UniversalSequence Representation…Towards Universal Sequence Representation Learning for Recommender SystemsScaling Law forRecommendation Models…Scaling Law for Recommendation Models: Towards General-Purpose User RepresentationsTransRec: LearningTransferable…TransRec: Learning Transferable Recommendation from Mixture-of-Modality FeedbackCollaborative Word-basedPre-trained Item…Collaborative Word-based Pre-trained Item Representation for Transferable RecommendationOnlineDistillation-enhanced…Online Distillation-enhanced Multi-modal Transformer for Sequential RecommendationExploring Adapter-basedTransfer Learning for…Exploring Adapter-based Transfer Learning for Recommender Systems: Empirical Studies and Practical InsightsRecommender Systems inthe Era of Large…Recommender Systems in the Era of Large Language Models (LLMs)A Survey on LargeLanguage Models for…A Survey on Large Language Models for RecommendationReLLa:Retrieval-enhanced Larg…ReLLa: Retrieval-enhanced Large Language Models for Lifelong Sequential Behavior Comprehension in RecommendationFLIP: Fine-grainedAlignment between…FLIP: Fine-grained Alignment between ID-based Models and Pretrained Language Models for CTR PredictionLifelong PersonalizedLow-Rank Adaptation of…Lifelong Personalized Low-Rank Adaptation of Large Language Models for RecommendationHow Can RecommenderSystems Benefit from…How Can Recommender Systems Benefit from Large Language Models: A SurveyAn Image Dataset forBenchmarking Recommende…An Image Dataset for Benchmarking Recommender Systems with Raw PixelsA Content-DrivenMicro-Video…A Content-Driven Micro-Video Recommendation Dataset at ScaleFull-Stack OptimizedLarge Language Models…Full-Stack Optimized Large Language Models for Lifelong Sequential Behavior Comprehension in RecommendationWhere to Go Next forRecommender Systems? ID…Where to Go Next for Recommender Systems? ID- vs. Modality-based Recommender Models RevisitedEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.