Document-Level Definition Detection in Scholarly Documents: Existing Models, Error Analyses, and Future Directions

Dongyeop Kang, Andrew Head, Risham Sidhu, Kyle Lo, Daniel Weld, Marti A. Hearst

Abstract

The task of definition detection is important for scholarly papers, because papers often make use of technical terminology that may be unfamiliar to readers. Despite prior work on definition detection, current approaches are far from being accurate enough to use in realworld applications. In this paper, we first perform in-depth error analysis of the current best performing definition detection system and discover major causes of errors. Based on this analysis, we develop a new definition detection system, HEDDEx, that utilizes syntactic features, transformer encoders, and heuristic filters, and evaluate it on a standard sentence-level benchmark. Because current benchmarks evaluate randomly sampled sentences, we propose an alternative evaluation that assesses every sentence within a document. This allows for evaluating recall in addition to precision. HEDDEx outperforms the leading system on both the sentence-level and the document-level tasks, by 12.7 F1 points and 14.4 F1 points, respectively. We note that performance on the high-recall document-level task is much lower than in the standard evaluation approach, due to the necessity of incorporation of document structure as features. We discuss remaining challenges in document-level definition detection, ideas for improvements, and potential issues for the development of reading aid applications.

Anthology ID:: 2020.sdp-1.22
Volume:: Proceedings of the First Workshop on Scholarly Document Processing
Month:: November
Year:: 2020
Address:: Online
Editors:: Muthu Kumar Chandrasekaran, Anita de Waard, Guy Feigenblat, Dayne Freitag, Tirthankar Ghosal, Eduard Hovy, Petr Knoth, David Konopnicki, Philipp Mayr, Robert M. Patton, Michal Shmueli-Scheuer
Venue:: sdp
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 196–206
Language:
URL:: https://aclanthology.org/2020.sdp-1.22
DOI:: 10.18653/v1/2020.sdp-1.22
Bibkey:
Cite (ACL):: Dongyeop Kang, Andrew Head, Risham Sidhu, Kyle Lo, Daniel Weld, and Marti A. Hearst. 2020. Document-Level Definition Detection in Scholarly Documents: Existing Models, Error Analyses, and Future Directions. In Proceedings of the First Workshop on Scholarly Document Processing, pages 196–206, Online. Association for Computational Linguistics.
Cite (Informal):: Document-Level Definition Detection in Scholarly Documents: Existing Models, Error Analyses, and Future Directions (Kang et al., sdp 2020)
Copy Citation:
PDF:: https://aclanthology.org/2020.sdp-1.22.pdf
Optional supplementary material:: 2020.sdp-1.22.OptionalSupplementaryMaterial.zip
Video:: https://slideslive.com/38940724
Code: allenai/scholarphi

PDF Cite Search Code Optional supplementary material Video