Summarizing scientific articles: Experiments with relevance and rhetorical status

被引：262

作者：

Teufel, S

Moens, M

机构：

[1] Univ Cambridge, Comp Lab, Cambridge CB3 0FD, England

[2] Rhetor Syst, Edinburgh EH8 9LS, Midlothian, Scotland

[3] Univ Edinburgh, Edinburgh EH8 9LS, Midlothian, Scotland

来源：

COMPUTATIONAL LINGUISTICS | 2002年 / 28卷 / 04期

关键词：

D O I：

10.1162/089120102762671936

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

In this article we propose a strategy for the summarization of scientific articles that concentrates on the rhetorical status of statements in an article: Material for summaries is selected in such a way that summaries can highlight the new contribution of the source article and situate it with respect to earlier work. We provide a gold standard for summaries of this kind consisting of a substantial corpus of conference articles in computational linguistics annotated with human judgments of the rhetorical status and relevance of each sentence in the articles. We present several experiments measuring our judges' agreement on these annotations. We also present an algorithm that, on the basis of the annotated training material, selects content from unseen articles and classifies it into a fixed set of seven rhetorical categories. The output of this extraction and classification system can be viewed as a single-document summary in its own right; alternatively, it provides starting material for the generation of task-oriented and user-tailored summaries designed to give users an overview of a scientific field.

引用

页码：409 / 445

页数：37

共 49 条

[1]

[Anonymous], IEEE COMPUTER

[2]

[Anonymous], P 22 ANN INT ACM SIG

[3]

[Anonymous], 2000, P 1 N AM CHAPTER ASS

[4]

Barzilay Regina., 1999, Proceedings of the 37th Annual Meeting of the Association for Computational Linguistics, P550, DOI [10.3115/1034678.1034760, DOI 10.3115/1034678.1034760, DOI 10.1115/10146781014760]

[5] MACHINE-MADE INDEX FOR TECHNICAL LITERATURE - AN EXPERIMENT [J].