Spatiotemporal Pyramid Network for Video Action Recognition

被引：529

作者：

Wang, Yunbo ^{[1
,2
,3
,4
]}

Long, Mingsheng ^{[1
,2
,3
,4
]}

Wang, Jianmin ^{[1
,2
,3
,4
]}

Yu, Philip S. ^{[1
,2
,3
,4
,5
]}

机构：

[1] Tsinghua Univ, MOE, KLiss, Beijing, Peoples R China

[2] Tsinghua Univ, TNList, Beijing, Peoples R China

[3] Tsinghua Univ, NEL BDSS, Beijing, Peoples R China

[4] Tsinghua Univ, Sch Software, Beijing, Peoples R China

[5] Univ Illinois, Chicago, IL 60680 USA

来源：

30TH IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR 2017) | 2017年

基金：

国家重点研发计划;

关键词：

D O I：

10.1109/CVPR.2017.226

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Two-stream convolutional networks have shown strong performance in video action recognition tasks. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, it remains unclear how to model the correlations between the spatial and temporal structures at multiple abstraction levels. First, the spatial stream tends to fail if two videos share similar backgrounds. Second, the temporal stream may be fooled if two actions resemble in short snippets, though appear to be distinct in the long term. We propose a novel spatiotemporal pyramid network to fuse the spatial and temporal features in a pyramid structure such that they can reinforce each other. From the architecture perspective, our network constitutes hierarchical fusion strategies which can be trained as a whole using a unified spatiotemporal loss. A series of ablation experiments support the importance of each fusion strategy. From the technical perspective, we introduce the spatiotemporal compact bilinear operator into video analysis tasks. This operator enables efficient training of bilinear fusion operations which can capture full interactions between the spatial and temporal features. Our final network achieves state-of-the-art results on standard video datasets.

引用

页码：2097 / 2106

页数：10

共 38 条

[1]

[Anonymous], 2016, ICLR

[2]

[Anonymous], 2015, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, DOI DOI 10.1109/CVPR.2015.7299007

[3]

[Anonymous], ARXIV151102274

[4]

Charikar M., 2002, Finding frequent items in data streams, P693

[5] P-CNN: Pose-based CNN Features for Action Recognition [J].

Cheron, Guilhem ;

Laptev, Ivan ;

Schmid, Cordelia .

2015 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2015, :3218-3226

[6]

Deng J, 2009, PROC CVPR IEEE, P248, DOI 10.1109/CVPRW.2009.5206848

[7]

Donahue J, 2015, PROC CVPR IEEE, P2625, DOI 10.1109/CVPR.2015.7298878

[8] Learning Spatiotemporal Features with 3D Convolutional Networks [J].

Du Tran ;

Bourdev, Lubomir ;

Fergus, Rob ;

Torresani, Lorenzo ;

Paluri, Manohar .

2015 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2015, :4489-4497

[9] Convolutional Two-Stream Network Fusion for Video Action Recognition [J].

Feichtenhofer, Christoph ;

Pinz, Axel ;

Zisserman, Andrew .

2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2016, :1933-1941

[10] Compact Bilinear Pooling [J].

Gao, Yang ;

Beijbom, Oscar ;

Zhang, Ning ;

Darrell, Trevor .

2016 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2016, :317-326

← 1 2 3 4 →