Textual Analysis of Stock Market Prediction Using Breaking Financial News: The AZFinText System

被引:414
作者
Schumaker, Robert P. [1 ]
Chen, Hsinchun [2 ]
机构
[1] Iona Coll, Dept Informat Syst, New Rochelle, NY 10801 USA
[2] Univ Arizona, Artificial Intelligence Lab, Dept Management Informat Syst, Tucson, AZ 85721 USA
关键词
Algorithms; Design; Experimentation; Measurement; Performance; SVM; stock market; prediction; SUPPORT VECTOR MACHINES;
D O I
10.1145/1462198.1462204
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Our research examines a predictive machine learning approach for financial news articles analysis using several different textual representations: bag of words, noun phrases, and named entities. Through this approach, we investigated 9,211 financial news articles and 10,259,042 stock quotes covering the S&P 500 stocks during a five week period. We applied our analysis to estimate a discrete stock price twenty minutes after a news article was released. Using a support vector machine (SVM) derivative specially tailored for discrete numeric prediction and models containing different stock-specific variables, we show that the model containing both article terms and stock price at the time of article release had the best performance in closeness to the actual future stock price (MSE 0.04261), the same direction of price movement as the future price (57.1% directional accuracy) and the highest return using a simulated trading engine (2.06% return). We further investigated the different textual representations and found that a Proper Noun scheme performs better than the de facto standard of Bag of Words in all three metrics.
引用
收藏
页数:19
相关论文
共 29 条
[1]  
[Anonymous], P 9 INT C INF KNOWL
[2]  
[Anonymous], 1973, RANDOM WALK DOWN WAL
[3]  
Bishop C. M., 2003, BAYESIAN REGRESSION
[4]  
BURNS D, 2005, SCHWAB MISS FORECAST
[5]  
CHO V, 1998, J COMPUTAT INTEL FIN, V26
[6]  
CHO V, 1999, KNOWLEDGE DISCOVERY
[7]   Early user - System interaction for database selection in massive domain-specific online environments [J].
Conrad, JG ;
Claussen, JRS .
ACM TRANSACTIONS ON INFORMATION SYSTEMS, 2003, 21 (01) :94-131
[8]  
Fama E., 1964, The Behavior of Stock Market Prices
[9]  
FUNG GPC, 2002, P PAC AS C KNOWL DIS
[10]   A probabilistic framework for SVM regression and error bar estimation [J].
Gao, JB ;
Gunn, SR ;
Harris, CJ ;
Brown, M .
MACHINE LEARNING, 2002, 46 (1-3) :71-89