HJ-Ky-0.1: an Evaluation Dataset for Kyrgyz Word Embeddings

AmazUtah_NLP at SemEval-2024 Task 9: A MultiChoice Question Answering System for Commonsense Defying Reasoning



arXiv:2411.10724v1 Announce Type: new
Abstract: One of the key tasks in modern applied computational linguistics is constructing word vector representations (word embeddings), which are widely used to address natural language processing tasks such as sentiment analysis, information extraction, and more. To choose an appropriate method for generating these word embeddings, quality assessment techniques are often necessary. A standard approach involves calculating distances between vectors for words with expert-assessed ‘similarity’. This work introduces the first ‘silver standard’ dataset for such tasks in the Kyrgyz language, alongside training corresponding models and validating the dataset’s suitability through quality evaluation metrics.



Source link
lol

By stp2y

Leave a Reply

Your email address will not be published. Required fields are marked *

No widgets found. Go to Widget page and add the widget in Offcanvas Sidebar Widget Area.