Relation Extraction Datasets in the Digital Humanities Domain and their Evaluation with Word Embeddings

Wohlgenannt, Gerhard; Chernyak, Ekaterina; Ilvovsky, Dmitry; Barinova, Ariadna; Mouromtsev, Dmitry

doi:10.1007/978-3-031-23793-5_18

Computer Science > Computation and Language

arXiv:1903.01284v1 (cs)

[Submitted on 4 Mar 2019]

Title:Relation Extraction Datasets in the Digital Humanities Domain and their Evaluation with Word Embeddings

Authors:Gerhard Wohlgenannt, Ekaterina Chernyak, Dmitry Ilvovsky, Ariadna Barinova, Dmitry Mouromtsev

View PDF

Abstract:In this research, we manually create high-quality datasets in the digital humanities domain for the evaluation of language models, specifically word embedding models. The first step comprises the creation of unigram and n-gram datasets for two fantasy novel book series for two task types each, analogy and doesn't-match. This is followed by the training of models on the two book series with various popular word embedding model types such as word2vec, GloVe, fastText, or LexVec. Finally, we evaluate the suitability of word embedding models for such specific relation extraction tasks in a situation of comparably small corpus sizes. In the evaluations, we also investigate and analyze particular aspects such as the impact of corpus term frequencies and task difficulty on accuracy. The datasets, and the underlying system and word embedding models are available on github and can be easily extended with new datasets and tasks, be used to reproduce the presented results, or be transferred to other domains.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:1903.01284 [cs.CL]
	(or arXiv:1903.01284v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1903.01284
Related DOI:	https://doi.org/10.1007/978-3-031-23793-5_18

Submission history

From: Gerhard Wohlgenannt Dr. [view email]
[v1] Mon, 4 Mar 2019 14:46:20 UTC (16 KB)

Computer Science > Computation and Language

Title:Relation Extraction Datasets in the Digital Humanities Domain and their Evaluation with Word Embeddings

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Relation Extraction Datasets in the Digital Humanities Domain and their Evaluation with Word Embeddings

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators