Visual Methods for Analyzing Probabilistic Classification Data

被引:80
作者
Alsallakh, Bilal [1 ]
Hanbury, Allan [1 ]
Hauser, Helwig [2 ]
Miksch, Silvia [1 ]
Rauber, Andreas [1 ]
机构
[1] Vienna Univ Technol, Vienna, Austria
[2] Univ Bergen, N-5020 Bergen, Norway
关键词
Probabilistic classification; confusion analysis; feature evaluation and selection; visual inspection; COMBINING MULTIPLE CLASSIFIERS; LAND-COVER CLASSIFICATIONS; VISUALIZATION; SEPARATION;
D O I
10.1109/TVCG.2014.2346660
中图分类号
TP31 [计算机软件];
学科分类号
081202 ; 0835 ;
摘要
Multi-class classifiers often compute scores for the classification samples describing probabilities to belong to different classes. In order to improve the performance of such classifiers, machine learning experts need to analyze classification results for a large number of labeled samples to find possible reasons for incorrect classification. Confusion matrices are widely used for this purpose. However, they provide no information about classification scores and features computed for the samples. We propose a set of integrated visual methods for analyzing the performance of probabilistic classifiers. Our methods provide insight into different aspects of the classification results for a large number of samples. One visualization emphasizes at which probabilities these samples were classified and how these probabilities correlate with classification error in terms of false positives and false negatives. Another view emphasizes the features of these samples and ranks them by their separation power between selected true and false classifications. We demonstrate the insight gained using our technique in a benchmarking classification dataset, and show how it enables improving classification performance by interactively defining and evaluating post-classification rules.
引用
收藏
页码:1703 / 1712
页数:10
相关论文
共 43 条
  • [31] A Novel Visualization Approach for Data-Mining-Related Classification
    Seifert, Christin
    Lex, Elisabeth
    [J]. INFORMATION VISUALIZATION, IV 2009, PROCEEDINGS, 2009, : 490 - 495
  • [32] Settles B, 2009, ACTIVE LEARNING LIT
  • [33] Shafait F., 2010, RAPIDMINER COMM M C, V9
  • [34] Talbot J, 2009, CHI2009: PROCEEDINGS OF THE 27TH ANNUAL CHI CONFERENCE ON HUMAN FACTORS IN COMPUTING SYSTEMS, VOLS 1-4, P1283
  • [35] Tatu Andrada, 2009, Proceedings of the 2009 IEEE Symposium on Visual Analytics Science and Technology. VAST 2009. Held co-jointly with VisWeek 2009, P59, DOI 10.1109/VAST.2009.5332628
  • [36] Combining multiple classifiers by averaging or by multiplying?
    Tax, DMJ
    van Breukelen, M
    Duin, RPW
    Kittler, J
    [J]. PATTERN RECOGNITION, 2000, 33 (09) : 1475 - 1485
  • [37] Teoh S.T., 2003, 9 ACM SIGKDD INT C K, P667, DOI 10.1145/956750.956837
  • [38] Visualization of cluster structure and separation in multivariate mixed data: A case study of diversity faultlines in work teams
    Tuan Pham
    Metoyer, Ronald
    Bezrukova, Katerina
    Spell, Chester
    [J]. COMPUTERS & GRAPHICS-UK, 2014, 38 : 117 - 130
  • [39] Van de Voorde T, 2007, PHOTOGRAMM ENG REM S, V73, P1017
  • [40] van den Elzen S., 2011, 2011 IEEE Conference on Visual Analytics Science and Technology, P151, DOI 10.1109/VAST.2011.6102453