CoSimLex: A Resource for Evaluating Graded Word Similarity in Context

Santos Armendariz, C; Purver, M; Ulčar, M; Pollak, S; Ljubešić, N; Robnik-Šikonja, M; Granroth-Wilding, M; Vaik, K; Language Resources and Evaluation Conference

dc.contributor.author	Santos Armendariz, C
dc.contributor.author	Purver, M
dc.contributor.author	Ulčar, M
dc.contributor.author	Pollak, S
dc.contributor.author	Ljubešić, N
dc.contributor.author	Robnik-Šikonja, M
dc.contributor.author	Granroth-Wilding, M
dc.contributor.author	Vaik, K
dc.contributor.author	Language Resources and Evaluation Conference
dc.date.accessioned	2020-04-17T09:05:27Z
dc.date.available	2020-02-11
dc.date.available	2020-04-17T09:05:27Z
dc.date.issued	2020
dc.identifier.uri	https://qmro.qmul.ac.uk/xmlui/handle/123456789/63608
dc.description.abstract	State of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Standard tasks and datasets for intrinsic evaluation of embeddings are based on judgements of similarity, but ignore context; standard tasks for word sense disambiguation take account of context but do not provide continuous measures of meaning similarity. This paper describes an effort to build a new dataset, CoSimLex, intended to fill this gap. Building on the standard pairwise similarity task of SimLex-999, it provides context-dependent similarity measures; covers not only discrete differences in word sense but more subtle, graded changes in meaning; and covers not only a well-resourced language (English) but a number of less-resourced languages. We define the task and evaluation metrics, outline the dataset collection methodology, and describe the status of the dataset so far.	en_US
dc.publisher	Language Resources and Evaluation Conference	en_US
dc.rights	This article is distributed under the terms of the CC BY-NC License
dc.title	CoSimLex: A Resource for Evaluating Graded Word Similarity in Context	en_US
dc.type	Conference Proceeding	en_US
dc.rights.holder	© The Author(s) 2020
pubs.author-url	http://www.eecs.qmul.ac.uk/~mpurver/papers/armendariz-et-al20lrec.pdf	en_US
pubs.notes	Not known	en_US
pubs.publication-status	Accepted	en_US
dcterms.dateAccepted	2020-02-11
rioxxterms.funder	Default funder	en_US
rioxxterms.identifier.project	Default project	en_US
qmul.funder	EMBEDDIA: Cross-Lingual Embeddings for Less-Represented Languages in European News Media::European Commission	en_US
qmul.funder	EMBEDDIA: Cross-Lingual Embeddings for Less-Represented Languages in European News Media::European Commission	en_US

Files in this item

Name:: Purver CoSimLex A Resource 2020 ...
Size:: 327.7Kb
Format:: application/
Description:: Published version

View/Open

This item appears in the following Collection(s)

Electronic Engineering and Computer Science [3490]

Show simple item record