[PDF][PDF] Uniform Density in Linguistic Information Derived from Dependency Structures.

M Richter, MBI Farré, M Kölbl, Y Kyogoku… - ICAART …, 2022 - pdfs.semanticscholar.org
M Richter, MBI Farré, M Kölbl, Y Kyogoku, JN Philipp, T Yousef, G Heyer, NP Himmelmann
ICAART (1), 2022pdfs.semanticscholar.org
This pilot study addresses the question of whether the Uniform Information Density principle
(UID) can be proved for eight typologically diverse languages. The lexical information of
words is derived from dependency structures both in sentences preceding the sentences
and within the sentence in which the target word occurs. Dependency structures are a
realisation of extra-sentential contexts for deriving information as formulated in the surprisal
model. Only subject, object and oblique, ie, the level directly below the verbal root node …
Abstract
This pilot study addresses the question of whether the Uniform Information Density principle (UID) can be proved for eight typologically diverse languages. The lexical information of words is derived from dependency structures both in sentences preceding the sentences and within the sentence in which the target word occurs. Dependency structures are a realisation of extra-sentential contexts for deriving information as formulated in the surprisal model. Only subject, object and oblique, ie, the level directly below the verbal root node, were considered. UID says that in natural language, the variance of information and information jumps from word to word should be small so as not to make the processing of a linguistic message an insurmountable hurdle. We observed cross-linguistically different information distributions but an almost identical UID, which provides evidence for the UID hypothesis and assumes that dependency structures can function as proxies for extrasentential contexts. However, for the dependency structures chosen as contexts, the information distributions in some languages were not statistically significantly different from distributions from a random corpus. This might be an effect of too low complexity of our model’s dependency structures, so lower hierarchical levels (eg phrases) should be considered.
pdfs.semanticscholar.org
Showing the best result for this search. See all results