Online Corpus Construction of English Text Collection, Data Cleaning, and Similarity Analysis

被引:0
|
作者
Wang, Huanyu [1 ]
机构
[1] Tangshan Normal Univ, Tangshan 063000, Peoples R China
关键词
TECHNOLOGY; DISCOURSE;
D O I
10.1155/2022/3105790
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Corpora are applied to analyze and study the characteristics of the target language. In language education, corpora are playing an increasingly essential role due to their large capacity, authenticity, rapid and accurate retrieval, as well as quick and easy statistics. At present, a great number of universities are trying to apply the textbook corpus to English teaching. However, most of the existing corpora face the issue of poor sharing. In addition, these corpora may be limited to a specific textbook, which leads to the lack of wide coverage of the retrieval and analysis results. As a result, it is quite necessary to develop a set of English corpora that is highly relevant, well shared, and easy to use by fully integrating existing teaching resources according to the characteristics of English subjects in universities. In recent years, the use of corpus-assisted English language teaching has gained widespread attention and exploration as computers have become more and more popular. After all, a corpus-based teaching model can effectively eliminate the various drawbacks of traditional vocabulary teaching. In fact, the corpus has a large amount of authentic corpus. The authenticity and practicality of the corpus facilitate students' mastery and use of English vocabulary in real contexts. What is more, the new model of corpus-assisted English vocabulary teaching can greatly increase independent learning and cooperative activities, so that students can increase their internal motivation for learning. This study begins with a brief introduction to the concept and characteristics of corpora. To be specific, the advantages of the corpus application in foreign language teaching are explained. At the same time, this research further analyzes the shortcomings of the existing corpus in university English education from the perspective of the current development and application of English corpora as well as clarifies the importance of building a corpus of university English teaching materials. After that, the system's operating environment and main development techniques are determined according to the specific requirements of the corpus for university English textbooks. In other words, the overall design and detailed design of the corpus and its management system were then carried out on the basis of the chosen technology platform. In addition, the structure of the tables in the database is analyzed and the basic components and operation procedures of the system are introduced. Furthermore, the functional modules of the system are designed. At the same time, the automatic word and sentence separation methods of the original corpus, the corpus entry process, the cross-distance search of the corpus, and the statistical analysis of the search results are discussed in detail. In conclusion, this study is based on English text collection and data cleaning techniques to build an online corpus.
引用
收藏
页数:8
相关论文
共 50 条
  • [41] Integrated Framework for Keyword-based Text Data Collection and Analysis
    Cha, Minki
    Kwon, Jung-Hyok
    Lee, Sol-Bee
    Park, Jaehoon
    Youm, Sungkwan
    Kim, Eui-Jik
    SENSORS AND MATERIALS, 2018, 30 (03) : 439 - 445
  • [42] CSR Image Construction of Chinese Construction Enterprises in Africa Based on Data Mining and Corpus Analysis
    Zhong, Yaoping
    Zhu, Wenzhong
    Zhou, Yingying
    MATHEMATICAL PROBLEMS IN ENGINEERING, 2020, 2020
  • [43] Challenges in the Analysis of Online Social Networks: A Data Collection Tool Perspective
    Anuradha Goswami
    Ajey Kumar
    Wireless Personal Communications, 2017, 97 : 4015 - 4061
  • [44] Data and text mining from online reviews: An automatic literature analysis
    Moro, Sergio
    Rita, Paulo
    WILEY INTERDISCIPLINARY REVIEWS-DATA MINING AND KNOWLEDGE DISCOVERY, 2022, 12 (03)
  • [45] Challenges in the Analysis of Online Social Networks: A Data Collection Tool Perspective
    Goswami, Anuradha
    Kumar, Ajey
    WIRELESS PERSONAL COMMUNICATIONS, 2017, 97 (03) : 4015 - 4061
  • [46] Syntactico-semantic realizations of pronouns in the English transitive construction: A corpus-based analysis
    Hwang, Haerim
    CORPUS LINGUISTICS AND LINGUISTIC THEORY, 2022, 18 (01) : 115 - 143
  • [48] Highway Construction Data Collection and Treatment in Preparation for Statistical Regression Analysis
    Williams, Robert C.
    Hildreth, John C.
    Vorster, Michael C.
    JOURNAL OF CONSTRUCTION ENGINEERING AND MANAGEMENT, 2009, 135 (12) : 1299 - 1306
  • [49] Analysis of Students' Behavior in English Online Education Based on Data Mining
    Wang, Chunxia
    MOBILE INFORMATION SYSTEMS, 2021, 2021
  • [50] Exploratory Text Analysis: Data-Driven versus Human Semantic Similarity Judgments
    Lindh-Knuutila, Tiina
    Honkela, Timo
    ADAPTIVE AND NATURAL COMPUTING ALGORITHMS, ICANNGA 2013, 2013, 7824 : 428 - 437