Automatic Tag Recommendation Algorithms for Social Recommender Systems

被引:127
作者
Song, Yang
Zhang, Lu [1 ]
Giles, C. Lee [1 ]
机构
[1] Penn State Univ, University Pk, PA 16802 USA
关键词
Algorithms; Experimentation; Performance; Tagging system; mixture model; graph partitioning; Gaussian processes; prototype selection; multi-label classification;
D O I
10.1145/1921591.1921595
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
The emergence of Web 2.0 and the consequent success of social network Web sites such as Delicious and Flickr introduce us to a new concept called social bookmarking, or tagging. Tagging is the action of connecting a relevant user-defined keyword to a document, image, or video, which helps the user to better organize and share their collections of interesting stuff. With the rapid growth of Web 2.0, tagged data is becoming more and more abundant on the social network Web sites. An interesting problem is how to automate the process of making tag recommendations to users when a new resource becomes available. In this article, we address the issue of tag recommendation from a machine learning perspective. From our empirical observation of two large-scale datasets, we first argue that the user-centered approach for tag recommendation is not very effective in practice. Consequently, we propose two novel document-centered approaches that are capable of making effective and efficient tag recommendations in real scenarios. The first, graph-based, method represents the tagged data in two bipartite graphs, (document, tag) and (document, word), then finds document topics by leveraging graph partitioning algorithms. The second, prototype-based, method aims at finding the most representative documents within the data collections and advocates a sparse multiclass Gaussian process classifier for efficient document classification. For both methods, tags are ranked within each topic cluster/class by a novel ranking method. Recommendations are performed by first classifying a new document into one or more topic clusters/classes, and then selecting the most relevant tags from those clusters/classes as machine-recommended tags. Experiments on real-world data from Delicious, CiteULike, and BibSonomy examine the quality of tag recommendation as well as the efficiency of our recommendation algorithms. The results suggest that our document-centered models can substantially improve the performance of tag recommendations when compared to the user-centered methods, as well as topic models LDA and SVM classifiers.
引用
收藏
页数:31
相关论文
共 31 条
[1]  
[Anonymous], 661 U CAL BERK DEP S
[2]  
[Anonymous], P COLL WEB TAGG WORK
[3]  
[Anonymous], IEEE T NEURAL NET
[4]  
[Anonymous], WORKSH P LERN WISS A
[5]  
[Anonymous], P 17 ACM C INF KNOWL
[6]  
[Anonymous], COMPUT STAT DATA ANA
[7]  
[Anonymous], P 16 INT C WORLD WID
[8]  
[Anonymous], P ANN INT ACM SIGIR
[9]  
[Anonymous], P EUR C ART INT ECAI
[10]  
[Anonymous], P WORKSH AI STAT