Meemi: A simple method for post-processing and integrating cross-lingual word embeddings

被引:0
|
作者
Doval, Yerai [1 ]
Camacho-Collados, Jose [2 ]
Espinosa-Anke, Luis [2 ]
Schockaert, Steven [2 ]
机构
[1] Univ Vigo, Escola Super Enxenaria Informat, Grp COLE, Ourensevigo, Spain
[2] Cardiff Univ, Sch Comp Sci & Informat, Cardiff CF24 3AA, Wales
关键词
D O I
10.1017/S1351324921000280
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Word embeddings have become a standard resource in the toolset of any Natural Language Processing practitioner. While monolingual word embeddings encode information about words in the context of a particular language, cross-lingual embeddings define a multilingual space where word embeddings from two or more languages are integrated together. Current state-of-the-art approaches learn these embeddings by aligning two disjoint monolingual vector spaces through an orthogonal transformation which preserves the structure of the monolingual counterparts. In this work, we propose to apply an additional transformation after this initial alignment step, which aims to bring the vector representations of a given word and its translations closer to their average. Since this additional transformation is non-orthogonal, it also affects the structure of the monolingual spaces. We show that our approach both improves the integration of the monolingual spaces and the quality of the monolingual spaces themselves. Furthermore, because our transformation can be applied to an arbitrary number of languages, we are able to effectively obtain a truly multilingual space. The resulting (monolingual and multilingual) spaces show consistent gains over the current state-of-the-art in standard intrinsic tasks, namely dictionary induction and word similarity, as well as in extrinsic tasks such as cross-lingual hypernym discovery and cross-lingual natural language inference.
引用
收藏
页码:746 / 768
页数:23
相关论文
共 50 条
  • [41] Cross-lingual hate speech detection using domain-specific word embeddings
    Monnar, Ayme Arango
    Rojas, Jorge Perez
    Labra, Barbara Polete
    PLOS ONE, 2024, 19 (07):
  • [42] Evaluating the Impact of Sub-word Information and Cross-lingual Word Embeddings on Mi'kmaq Language Modelling
    Boudreau, Jeremie
    Patra, Akankshya
    Suvarna, Ashima
    Cook, Paul
    PROCEEDINGS OF THE 12TH INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION (LREC 2020), 2020, : 2736 - 2745
  • [43] Why Overfitting Isn't Always Bad: Retrofitting Cross-Lingual Word Embeddings to Dictionaries
    Zhang, Mozhi
    Fujinuma, Yoshinari
    Paul, Michael J.
    Boyd-Graber, Jordan
    58TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2020), 2020, : 2214 - 2220
  • [44] Manipuri-English Cross-lingual Word Embeddings using a Temporally Aligned Comparable Corpus
    Laitonjam, Lenin
    Singh, Sanasam Ranbir
    2021 INTERNATIONAL CONFERENCE ON ASIAN LANGUAGE PROCESSING (IALP), 2021, : 195 - 199
  • [45] Gender Bias in Multilingual Embeddings and Cross-Lingual Transfer
    Zhao, Jieyu
    Mukherjee, Subhabrata
    Hosseini, Saghar
    Chang, Kai-Wei
    Awadallah, Ahmed Hassan
    58TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2020), 2020, : 2896 - 2907
  • [46] A Resource-Free Evaluation Metric for Cross-Lingual Word Embeddings Based on Graph Modularity
    Fujinuma, Yoshinari
    Boyd-Graber, Jordan
    Paul, Michael J.
    57TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2019), 2019, : 4952 - 4962
  • [47] Weakly-Supervised Concept-based Adversarial Learning for Cross-lingual Word Embeddings
    Wang, Haozhou
    Henderson, James
    Merlo, Paola
    2019 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING AND THE 9TH INTERNATIONAL JOINT CONFERENCE ON NATURAL LANGUAGE PROCESSING (EMNLP-IJCNLP 2019): PROCEEDINGS OF THE CONFERENCE, 2019, : 4419 - 4430
  • [48] Hyperpolyglot LLMs: Cross-Lingual Interpretability in Token Embeddings
    Wen-Yi, Andrea W.
    Mimno, David
    2023 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING, EMNLP 2023, 2023, : 1124 - 1131
  • [49] Cross-Lingual Alignment of Contextual Word Embeddings, with Applications to Zero-shot Dependency Parsing
    Schuster, Tal
    Ram, Ori
    Barzilay, Regina
    Globerson, Amir
    2019 CONFERENCE OF THE NORTH AMERICAN CHAPTER OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS: HUMAN LANGUAGE TECHNOLOGIES (NAACL HLT 2019), VOL. 1, 2019, : 1599 - 1613
  • [50] Data Augmentation with Unsupervised Machine Translation Improves the Structural Similarity of Cross-lingual Word Embeddings
    Nishikawa, Sosuke
    Ri, Ryokan
    Tsuruoka, Yoshimasa
    ACL-IJCNLP 2021: THE 59TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS AND THE 11TH INTERNATIONAL JOINT CONFERENCE ON NATURAL LANGUAGE PROCESSING: PROCEEDINGS OF THE STUDENT RESEARCH WORKSHOP, 2021, : 163 - 173